Technical Lead – GPU Infrastructure

🔥 4 minutes ago

🇵🇰 Pakistan – Remote

⏰ Full Time

🟠 Senior

🧑‍💻 Full-stack Engineer

👻 Ghost score 25%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Tether.to

Tether.to

11 - 50 employees

Founded 2014

₿ Crypto

💳 Fintech

💸 Finance

Crypto • Fintech • Finance

Tether. to is a leading digital asset company that pioneers the use of stablecoins in the blockchain space. As the most widely adopted stablecoin, Tether tokens are designed to be pegged 1-to-1 with fiat currencies, offering a stable digital asset option for users. The platform facilitates these token transactions across multiple blockchains, enhancing cross-border transactions while maintaining transparency with daily records of total assets and reserves. Tether's initiatives include educational programs promoting digital asset usage, especially targeting regions like the Middle East, Turkey, and the Philippines. Tether thus positions itself as a disruptor in the traditional financial system by enabling a stable, efficient method of handling transactions in the digital currency world.

📋 Description

• Own Cosmic AC's end-to-end platform architecture through architecture proposals, high-level and low-level designs, reviews, and baseline maintenance • Lead and line-manage approximately twelve distributed engineers across backend, frontend, DevOps, QA, and documentation • Define engineering standards, conduct code and design reviews, manage release gates, hold one-to-ones, and provide growth and performance input • Design, build, and operate a managed Slurm service for research users • Manage Slurm controllers, accounting, partitions, login nodes, node onboarding, driver and CUDA baselines, stalled-job and node-health detection, draining, autohealing, storage visibility, identity, and isolation • Own Kubernetes cluster bootstrap and lifecycle on partner-provided bare metal • Operate NVIDIA GPU Operator and Network Operator, VM-based GPU isolation with KubeVirt and VFIO, upgrades, backup and recovery, and node replacement • Lead managed inference architecture, including serving, multi-GPU and multi-node parallelism, autoscaling, request routing, endpoint reliability, and confidential-compute-capable capacity • Establish metrics, logging, alerting, and SLOs across the control plane, GPU fleet, and application tiers • Lead incident response, post-incident reviews, and a sustainable on-call model • Serve as primary technical interface to infrastructure partners and vendors • Translate requirements into written specifications and acceptance tests, manage escalations, and contribute to capacity planning and hardware sourcing • Work with research, model-training, and product teams to translate workloads into platform requirements and broker capacity • Complete the platform team and establish the technical bar for new engineers

🎯 Requirements

• Eight or more years of hands-on engineering experience • At least three years leading teams that build and operate infrastructure platforms other teams depend on • Bachelor's or Master's degree in computer science or engineering, or equivalent practical experience • Hands-on experience with Slurm at scale, including slurmctld, slurmdbd, partitions, QoS, priority, accounting, prolog and epilog, node health scripting, and upgrades with jobs running • Experience operating an HPC or GPU training cluster for a research population is ideally preferred • Experience operating NVIDIA GPU fleets on bare metal, including driver and CUDA lifecycle, Fabric Manager, NVSwitch, DCGM, MIG, node burn-in, and acceptance • Experience with InfiniBand, subnet configuration, RDMA, SR-IOV, and diagnosing multi-node NCCL performance issues • Deep Linux systems knowledge, including kernel modules, drivers, PCIe passthrough, vfio-pci, cgroups, namespaces, and performance tuning • Production Kubernetes operations, including control plane, upgrades, CNI, CSI, operators, custom controllers, and multi-tenancy • Experience with HPC storage and data movement, including VAST, Lustre, NFS, node-local NVMe caching, and distributing large model weights and datasets • Experience with Prometheus, Grafana, Loki or equivalents, SLOs, incident response, and post-incident review • Working fluency in JavaScript and Node.js sufficient to review control-plane, CLI, and worker services and make architecture decisions • Experience shipping a multi-tenant IaaS, PaaS, or research computing service with isolation, quotas, usage metering, and user-facing API and CLI surfaces • People management across time zones, cross-track review, written architecture decisions, and ability to challenge partners or executives with reasons • Excellent written and spoken English • Availability to work from a location between UTC and UTC+5:30 • Desirable: Slurm operators on Kubernetes or Kubernetes-native schedulers • Desirable: modern serving stacks such as vLLM, SGLang, and TensorRT-LLM • Desirable: VM and container isolation, confidential computing, Cluster API, kubeadm, Cilium, NVSentinel-class autohealing, infrastructure as code, GitOps, GPU cloud/HPC/AI lab platform experience, distributed systems, and hardware-provider partnership experience

🏖️ Benefits

• Fully remote work arrangement • Occasional travel to partner sites and team events

Apply Now

Similar Jobs

🕒 September 9

Smart Working

51 - 200

💼 Consulting

🏥 Healthcare

📣 Marketing

Senior C#/Angular engineer developing scalable backend software for a UK golf-membership business. Solving architecture, integration, performance, and technical-support challenges remotely from Pakistan.

Angular

Azure

JavaScript

Next.js

SOAP

SQL

.NET

🕒 September 8

AIM Qualifications and Assessment Group

51 - 200

📚 Education

🤝 Non-profit

👥 HR Tech

Full Stack Developer building scalable enterprise applications with .NET, Angular, SQL Server, and Azure. Supporting cloud and software engineering projects for Canadian technology consultancy AIM.

Angular

Azure

Cloud

Distributed Systems

Docker

GraphQL

Kubernetes

Microservices

SQL

TypeScript

Vault

.NET

🕒 September 5

CloudPSO

51 - 200

💼 Consulting

📦 Logistics

📣 Marketing

AI frontend lead building real-time voice-agent interfaces for CloudPSO’s enterprise AI platform. Leading React-based architecture, WebRTC/WebSocket integration, dashboards, and frontend team development.

Angular

AWS

Azure

Cypress

Docker

Google Cloud Platform

JavaScript

Jest

Kubernetes

React

Redux

TypeScript

Vue.js

Webpack

🕒 August 5

Idenfo Direct Global

11 - 50

💳 Fintech

📋 Compliance

🔒 Cybersecurity

Full-stack engineer building MERN and OpenAI-integrated web applications for a UK technology client. Deploying scalable, secure digital products on AWS with Next.js.

AWS

Cloud

Docker

EC2

GraphQL

JavaScript

Kubernetes

Microservices

MongoDB

Next.js

Node.js

React

TypeScript

🕒 July 27

Firehorse

2 - 10

🤖 Artificial Intelligence

☁️ SaaS

🛡️ Insurance

Senior Full Stack Developer leading the design and development of web and mobile applications using React, React Native, and Node.js with client interaction and technical leadership.

AWS

Cloud

Docker

JavaScript

Node.js

NoSQL

React

React Native

SQL