Technical Lead – GPU Infrastructure

🔥 2 minutes ago

🇬🇧 United Kingdom – Remote

⏰ Full Time

🟠 Senior

🧑‍💻 Full-stack Engineer

👻 Ghost score 25%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Tether.to

Tether.to

11 - 50 employees

Founded 2014

₿ Crypto

💳 Fintech

💸 Finance

Crypto • Fintech • Finance

Tether. to is a leading digital asset company that pioneers the use of stablecoins in the blockchain space. As the most widely adopted stablecoin, Tether tokens are designed to be pegged 1-to-1 with fiat currencies, offering a stable digital asset option for users. The platform facilitates these token transactions across multiple blockchains, enhancing cross-border transactions while maintaining transparency with daily records of total assets and reserves. Tether's initiatives include educational programs promoting digital asset usage, especially targeting regions like the Middle East, Turkey, and the Philippines. Tether thus positions itself as a disruptor in the traditional financial system by enabling a stable, efficient method of handling transactions in the digital currency world.

📋 Description

• Own the end-to-end architecture of Tether Data's Cosmic AC GPU compute and managed inference platform • Lead architecture proposals, high-level and low-level designs, technical reviews, and maintain the architecture baseline • Lead and line-manage approximately twelve distributed engineers across backend, frontend, DevOps, QA, and documentation • Establish engineering standards, conduct code and design reviews, manage release gates, hold one-to-ones, and provide growth and performance input • Design, build, and operate a managed Slurm service for research users • Own Slurm controller and accounting, partitions and login nodes, node onboarding, driver and CUDA baselines, stalled-job and node-health detection, drain and autohealing, storage visibility, identity, and isolation • Own Kubernetes cluster bootstrap and lifecycle on partner-provided bare metal, NVIDIA GPU Operator and Network Operator, VM-based GPU isolation, upgrades, backup and recovery, and node replacement • Design managed inference serving architecture, multi-GPU and multi-node parallelism, autoscaling, request routing, endpoint reliability, and confidential-compute-capable capacity • Establish metrics, logging, alerting, and SLOs across the control plane, GPU fleet, and application tiers • Lead incident response, post-incident reviews, and design a sustainable on-call model • Serve as the primary technical interface to infrastructure partners and vendors • Translate requirements into written specifications and acceptance tests, manage escalations, and contribute to capacity planning and hardware sourcing • Work with research, model-training, and product teams to translate workloads into platform requirements and broker capacity • Complete the platform team and set the technical bar for new engineers

🎯 Requirements

• Eight or more years of hands-on engineering experience • At least three years leading teams that build and operate infrastructure platforms other teams depend on • Bachelor's or Master's degree in computer science or engineering, or equivalent practical experience • Hands-on experience running Slurm at scale, including slurmctld, slurmdbd, partitions, QoS and priority, accounting, prolog and epilog, node health scripting, and upgrades with jobs on the system • Experience operating an HPC or GPU training cluster for a research population is ideally preferred • GPU fleet operation on bare metal, including NVIDIA driver and CUDA lifecycle, Fabric Manager and NVSwitch behaviour on SXM systems, DCGM-based health and utilisation, MIG, node burn-in and acceptance • High-performance interconnect experience with InfiniBand fabric and subnet configuration, RDMA, SR-IOV, and diagnosing multi-node NCCL performance problems • Linux systems expertise including kernel modules and drivers, PCIe passthrough and vfio-pci, cgroups and namespaces, and performance tuning for compute-heavy workloads • Production Kubernetes operation, including control plane, upgrades, CNI and CSI, operators and custom controllers, and multi-tenancy design • HPC storage and data movement experience with VAST, Lustre, NFS, node-local NVMe caching, and distributing large model weights and datasets across many nodes • Observability and operations experience with Prometheus, Grafana and Loki or equivalents, SLOs, incident response, and post-incident review • Working fluency in JavaScript and Node.js sufficient to review a control plane, CLI and worker services and make architecture decisions • Experience shipping a platform with real users, such as a multi-tenant IaaS or PaaS or research computing service • People management across time zones, cross-track review, written architecture decisions, and ability to communicate decisions to partners or executives • Excellent written and spoken English • Desirable experience with Slurm operators on Kubernetes or Kubernetes-native schedulers • Desirable experience with modern serving stacks such as vLLM, SGLang, or TensorRT-LLM • Desirable experience with VM and container isolation, confidential computing, Cluster API, kubeadm, Cilium, autohealing, infrastructure as code, GitOps, GPU cloud/HPC/AI lab platforms, distributed systems, and hardware-provider partnerships

🏖️ Benefits

• Fully remote work arrangement • Occasional travel to partner sites and team events

Apply Now

Similar Jobs

🔥 3 hours ago

Stark & Stark

201 - 500

💼 Consulting

🛡️ Insurance

⚖️ Legal

Lead software engineer coordinating software for STARK’s AI-enabled autonomous vessels and ground systems. Driving maritime product roadmaps, embedded development, simulation, and defence field testing.

🔥 23 hours ago

MyEdSpace

51 - 200

📚 Education

👥 B2C

Tech Lead scaling MyEdSpace’s education technology platform for accessible learning. Guiding architecture, engineering teams, DevOps, and AI-powered personalized learning.

🕒 Yesterday

Palo Alto Networks

10,000+ employees

🔒 Cybersecurity

🏢 Enterprise

Senior Technical Product Engineer advancing Palo Alto Networks’ cloud-native cybersecurity products. Shaping product roadmaps, automation, and multi-cloud security capabilities through customer and field feedback.

🕒 Yesterday

Mozilla

501 - 1000

👥 B2C

🔒 Cybersecurity

Senior Software Engineer modernizing Mozilla Firefox add-ons infrastructure, moderation systems, and developer tools. Building reliable Python/Django platforms serving millions of Firefox users worldwide.

🕒 2 days ago

AeroVect

11 - 50

📦 Logistics

💼 Consulting

🚀 Aerospace

Senior Staff Perception Engineer architecting multimodal 2D/3D detection systems. Advancing AeroVect’s autonomous ground-handling technology for airlines and service providers.