Technical Lead – GPU Infrastructure

🔥 il y a 2 heures

🇬🇧 Royaume-Uni – Télétravail

⏰ Temps Plein

🟠 Senior

🧑‍💻 Développeur Full-Stack

👻 Score fantôme 25%

infoinfo

🗣️🇺🇸🇬🇧 Anglais requis

Postuler Maintenant
Trouver des Emplois à Distance Similaires

📊 Vérifiez votre score de CV pour ce poste

Améliorez vos chances d'obtenir un entretien en vérifiant votre score de CV avant de postuler.

Logo of Tether.to

Tether.to

11 - 50 employés

Fondée en 2014

₿ Crypto

💳 Fintech

💸 Finance

Crypto • Fintech • Finance

Tether. to est une entreprise leader dans le domaine des actifs numériques qui est pionnière dans l'utilisation des stablecoins dans l'univers de la blockchain. En tant que stablecoin le plus adopté au monde, les tokens Tether sont conçus pour être indexés à 1 pour 1 avec les monnaies fiduciaires, offrant ainsi une option d'actif numérique stable pour les utilisateurs. La plateforme facilite ces transactions de tokens à travers plusieurs blockchains, améliorant les transactions transfrontalières tout en maintenant la transparence grâce à des relevés quotidiens des actifs et réserves totaux. Les initiatives de Tether incluent des programmes éducatifs promouvant l'utilisation des actifs numériques, ciblant particulièrement des régions comme le Moyen-Orient, la Turquie et les Philippines. Ainsi, Tether se positionne comme un perturbateur du système financier traditionnel en permettant une méthode stable et efficace de gestion des transactions dans le monde de la monnaie numérique.

Description

• Own the end-to-end architecture of Tether Data's Cosmic AC GPU compute and managed inference platform • Lead architecture proposals, high-level and low-level designs, technical reviews, and maintain the architecture baseline • Lead and line-manage approximately twelve distributed engineers across backend, frontend, DevOps, QA, and documentation • Establish engineering standards, conduct code and design reviews, manage release gates, hold one-to-ones, and provide growth and performance input • Design, build, and operate a managed Slurm service for research users • Own Slurm controller and accounting, partitions and login nodes, node onboarding, driver and CUDA baselines, stalled-job and node-health detection, drain and autohealing, storage visibility, identity, and isolation • Own Kubernetes cluster bootstrap and lifecycle on partner-provided bare metal, NVIDIA GPU Operator and Network Operator, VM-based GPU isolation, upgrades, backup and recovery, and node replacement • Design managed inference serving architecture, multi-GPU and multi-node parallelism, autoscaling, request routing, endpoint reliability, and confidential-compute-capable capacity • Establish metrics, logging, alerting, and SLOs across the control plane, GPU fleet, and application tiers • Lead incident response, post-incident reviews, and design a sustainable on-call model • Serve as the primary technical interface to infrastructure partners and vendors • Translate requirements into written specifications and acceptance tests, manage escalations, and contribute to capacity planning and hardware sourcing • Work with research, model-training, and product teams to translate workloads into platform requirements and broker capacity • Complete the platform team and set the technical bar for new engineers

🎯 Exigences

• Eight or more years of hands-on engineering experience • At least three years leading teams that build and operate infrastructure platforms other teams depend on • Bachelor's or Master's degree in computer science or engineering, or equivalent practical experience • Hands-on experience running Slurm at scale, including slurmctld, slurmdbd, partitions, QoS and priority, accounting, prolog and epilog, node health scripting, and upgrades with jobs on the system • Experience operating an HPC or GPU training cluster for a research population is ideally preferred • GPU fleet operation on bare metal, including NVIDIA driver and CUDA lifecycle, Fabric Manager and NVSwitch behaviour on SXM systems, DCGM-based health and utilisation, MIG, node burn-in and acceptance • High-performance interconnect experience with InfiniBand fabric and subnet configuration, RDMA, SR-IOV, and diagnosing multi-node NCCL performance problems • Linux systems expertise including kernel modules and drivers, PCIe passthrough and vfio-pci, cgroups and namespaces, and performance tuning for compute-heavy workloads • Production Kubernetes operation, including control plane, upgrades, CNI and CSI, operators and custom controllers, and multi-tenancy design • HPC storage and data movement experience with VAST, Lustre, NFS, node-local NVMe caching, and distributing large model weights and datasets across many nodes • Observability and operations experience with Prometheus, Grafana and Loki or equivalents, SLOs, incident response, and post-incident review • Working fluency in JavaScript and Node.js sufficient to review a control plane, CLI and worker services and make architecture decisions • Experience shipping a platform with real users, such as a multi-tenant IaaS or PaaS or research computing service • People management across time zones, cross-track review, written architecture decisions, and ability to communicate decisions to partners or executives • Excellent written and spoken English • Desirable experience with Slurm operators on Kubernetes or Kubernetes-native schedulers • Desirable experience with modern serving stacks such as vLLM, SGLang, or TensorRT-LLM • Desirable experience with VM and container isolation, confidential computing, Cluster API, kubeadm, Cilium, autohealing, infrastructure as code, GitOps, GPU cloud/HPC/AI lab platforms, distributed systems, and hardware-provider partnerships

🏖️ Avantages

• Fully remote work arrangement • Occasional travel to partner sites and team events

Postuler Maintenant

Emplois Similaires

🔥 il y a 5 heures

Stark & Stark

201 - 500

💼 Conseil

🛡️ Assurance

⚖️ Juridique

Lead software engineer coordinating software for STARK’s AI-enabled autonomous vessels and ground systems. Driving maritime product roadmaps, embedded development, simulation, and defence field testing.

🇬🇧 Royaume-Uni – Télétravail

⏰ Temps Plein

🟠 Senior

🧑‍💻 Développeur Full-Stack

🗣️🇺🇸🇬🇧 Anglais requis

🕒 il y a 1 jour

MyEdSpace

51 - 200

📚 Éducation

👥 B2C

Tech Lead scaling MyEdSpace’s education technology platform for accessible learning. Guiding architecture, engineering teams, DevOps, and AI-powered personalized learning.

🇬🇧 Royaume-Uni – Télétravail

⏰ Temps Plein

🟠 Senior

🧑‍💻 Développeur Full-Stack

🗣️🇺🇸🇬🇧 Anglais requis

🕒 il y a 1 jour

Palo Alto Networks

10 000+ employés

🔒 Cybersecurity

🏢 Entreprise

Senior Technical Product Engineer advancing Palo Alto Networks’ cloud-native cybersecurity products. Shaping product roadmaps, automation, and multi-cloud security capabilities through customer and field feedback.

🗣️🇺🇸🇬🇧 Anglais requis

🕒 il y a 1 jour

Mozilla

501 - 1000

👥 B2C

🔒 Cybersecurity

Senior Software Engineer modernizing Mozilla Firefox add-ons infrastructure, moderation systems, and developer tools. Building reliable Python/Django platforms serving millions of Firefox users worldwide.

🇬🇧 Royaume-Uni – Télétravail

💵 £66 000 - £87 000 / an

⏰ Temps Plein

🟠 Senior

🧑‍💻 Développeur Full-Stack

🗣️🇺🇸🇬🇧 Anglais requis

🕒 il y a 2 jours

AeroVect

11 - 50

📦 Logistique

💼 Conseil

🚀 Aérospatiale

Senior Staff Perception Engineer architecting multimodal 2D/3D detection systems. Advancing AeroVect’s autonomous ground-handling technology for airlines and service providers.

🇬🇧 Royaume-Uni – Télétravail

💰 Seed Round en 2022-07

⏰ Temps Plein

🟠 Senior

🧑‍💻 Développeur Full-Stack

🗣️🇺🇸🇬🇧 Anglais requis