Technical Lead – GPU Infrastructure

🕒 3 dias atrás

🇬🇧 Reino Unido – Remoto (RU)

⏰ Tempo Integral

🟠 Sênior

🧑‍💻 Engenheiro Full-stack

👻 Score fantasma 25%

infoinfo

🗣️🇺🇸🇬🇧 Inglês obrigatório

Candidatar-se
Encontrar Vagas Remotas Similares

📊 Verifique sua pontuação de currículo para esta vaga

Melhore suas chances de conseguir uma entrevista verificando sua pontuação de currículo antes de se candidatar.

Logo of Tether.to

Tether.to

11 - 50 funcionários

Fundada em 2014

₿ Cripto

💳 Fintech

💸 Finanças

Crypto • Fintech • Finance

Tether. to é uma empresa líder em ativos digitais que pioneiriza o uso de stablecoins no ecossistema blockchain. Como a stablecoin mais amplamente adotada, os tokens Tether são projetados para serem lastreados 1:1 em moedas fiduciárias, oferecendo aos usuários uma opção de ativo digital estável. A plataforma facilita transações desses tokens em múltiplas blockchains, aprimorando operações transfronteiriças enquanto mantém transparência com registros diários do total de ativos e reservas. As iniciativas da Tether incluem programas educacionais que promovem o uso de ativos digitais, com foco especial em regiões como Oriente Médio, Turquia e Filipinas. Assim, a Tether se posiciona como um agente de disrupção no sistema financeiro tradicional ao viabilizar um método estável e eficiente de realizar transações no mundo das moedas digitais.

Descrição

• Own the end-to-end architecture of Tether Data's Cosmic AC GPU compute and managed inference platform • Lead architecture proposals, high-level and low-level designs, technical reviews, and maintain the architecture baseline • Lead and line-manage approximately twelve distributed engineers across backend, frontend, DevOps, QA, and documentation • Establish engineering standards, conduct code and design reviews, manage release gates, hold one-to-ones, and provide growth and performance input • Design, build, and operate a managed Slurm service for research users • Own Slurm controller and accounting, partitions and login nodes, node onboarding, driver and CUDA baselines, stalled-job and node-health detection, drain and autohealing, storage visibility, identity, and isolation • Own Kubernetes cluster bootstrap and lifecycle on partner-provided bare metal, NVIDIA GPU Operator and Network Operator, VM-based GPU isolation, upgrades, backup and recovery, and node replacement • Design managed inference serving architecture, multi-GPU and multi-node parallelism, autoscaling, request routing, endpoint reliability, and confidential-compute-capable capacity • Establish metrics, logging, alerting, and SLOs across the control plane, GPU fleet, and application tiers • Lead incident response, post-incident reviews, and design a sustainable on-call model • Serve as the primary technical interface to infrastructure partners and vendors • Translate requirements into written specifications and acceptance tests, manage escalations, and contribute to capacity planning and hardware sourcing • Work with research, model-training, and product teams to translate workloads into platform requirements and broker capacity • Complete the platform team and set the technical bar for new engineers

🎯 Requisitos

• Eight or more years of hands-on engineering experience • At least three years leading teams that build and operate infrastructure platforms other teams depend on • Bachelor's or Master's degree in computer science or engineering, or equivalent practical experience • Hands-on experience running Slurm at scale, including slurmctld, slurmdbd, partitions, QoS and priority, accounting, prolog and epilog, node health scripting, and upgrades with jobs on the system • Experience operating an HPC or GPU training cluster for a research population is ideally preferred • GPU fleet operation on bare metal, including NVIDIA driver and CUDA lifecycle, Fabric Manager and NVSwitch behaviour on SXM systems, DCGM-based health and utilisation, MIG, node burn-in and acceptance • High-performance interconnect experience with InfiniBand fabric and subnet configuration, RDMA, SR-IOV, and diagnosing multi-node NCCL performance problems • Linux systems expertise including kernel modules and drivers, PCIe passthrough and vfio-pci, cgroups and namespaces, and performance tuning for compute-heavy workloads • Production Kubernetes operation, including control plane, upgrades, CNI and CSI, operators and custom controllers, and multi-tenancy design • HPC storage and data movement experience with VAST, Lustre, NFS, node-local NVMe caching, and distributing large model weights and datasets across many nodes • Observability and operations experience with Prometheus, Grafana and Loki or equivalents, SLOs, incident response, and post-incident review • Working fluency in JavaScript and Node.js sufficient to review a control plane, CLI and worker services and make architecture decisions • Experience shipping a platform with real users, such as a multi-tenant IaaS or PaaS or research computing service • People management across time zones, cross-track review, written architecture decisions, and ability to communicate decisions to partners or executives • Excellent written and spoken English • Desirable experience with Slurm operators on Kubernetes or Kubernetes-native schedulers • Desirable experience with modern serving stacks such as vLLM, SGLang, or TensorRT-LLM • Desirable experience with VM and container isolation, confidential computing, Cluster API, kubeadm, Cilium, autohealing, infrastructure as code, GitOps, GPU cloud/HPC/AI lab platforms, distributed systems, and hardware-provider partnerships

🏖️ Benefícios

• Fully remote work arrangement • Occasional travel to partner sites and team events

Candidatar-se

Vagas Similares

🕒 3 dias atrás

Stark & Stark

201 - 500

💼 Consultoria

🛡️ Seguros

⚖️ Jurídico

Lead software engineer coordinating software for STARK’s AI-enabled autonomous vessels and ground systems. Driving maritime product roadmaps, embedded development, simulation, and defence field testing.

🇬🇧 Reino Unido – Remoto (RU)

⏰ Tempo Integral

🟠 Sênior

🧑‍💻 Engenheiro Full-stack

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 4 dias atrás

MyEdSpace

51 - 200

📚 Educação

👥 B2C

Tech Lead scaling MyEdSpace’s education technology platform for accessible learning. Guiding architecture, engineering teams, DevOps, and AI-powered personalized learning.

🇬🇧 Reino Unido – Remoto (RU)

⏰ Tempo Integral

🟠 Sênior

🧑‍💻 Engenheiro Full-stack

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 5 dias atrás

Mozilla

501 - 1000

👥 B2C

🔒 Cibersegurança

Senior Software Engineer modernizing Mozilla Firefox add-ons infrastructure, moderation systems, and developer tools. Building reliable Python/Django platforms serving millions of Firefox users worldwide.

🇬🇧 Reino Unido – Remoto (RU)

💵 £66.000 - £87.000 / ano

⏰ Tempo Integral

🟠 Sênior

🧑‍💻 Engenheiro Full-stack

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 5 dias atrás

AeroVect

11 - 50

📦 Logística

💼 Consultoria

🚀 Aeroespacial

Senior Staff Perception Engineer architecting multimodal 2D/3D detection systems. Advancing AeroVect’s autonomous ground-handling technology for airlines and service providers.

🇬🇧 Reino Unido – Remoto (RU)

💰 Seed Round em 2022-07

⏰ Tempo Integral

🟠 Sênior

🧑‍💻 Engenheiro Full-stack

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 6 dias atrás

Budibase

51 - 200

💼 Consultoria

📦 Logística

☁️ SaaS

Senior Software Engineer building Budibase’s open-source platform for apps, workflows, integrations, and AI agents. Owning substantial product and technical work from problem through production.

🇬🇧 Reino Unido – Remoto (RU)

💵 £60.000 - £80.000 / ano

💰 $7.000.000 Seed Round em 2022-11

⏰ Tempo Integral

🟠 Sênior

🧑‍💻 Engenheiro Full-stack

🗣️🇺🇸🇬🇧 Inglês obrigatório