Technical Lead – GPU Infrastructure

🔥 20 hours ago

🇪🇸 Spain – Remote

⏰ Full Time

🟠 Senior

🧑‍💻 Full-stack Engineer

👻 Ghost score 25%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Tether.to

Tether.to

11 - 50 employees

Founded 2014

₿ Crypto

💳 Fintech

💸 Finance

Crypto • Fintech • Finance

Tether. to is a leading digital asset company that pioneers the use of stablecoins in the blockchain space. As the most widely adopted stablecoin, Tether tokens are designed to be pegged 1-to-1 with fiat currencies, offering a stable digital asset option for users. The platform facilitates these token transactions across multiple blockchains, enhancing cross-border transactions while maintaining transparency with daily records of total assets and reserves. Tether's initiatives include educational programs promoting digital asset usage, especially targeting regions like the Middle East, Turkey, and the Philippines. Tether thus positions itself as a disruptor in the traditional financial system by enabling a stable, efficient method of handling transactions in the digital currency world.

📋 Description

• Own the end-to-end architecture of Cosmic AC, including architecture proposals, high-level and low-level designs, reviews, and maintaining the baseline • Lead and line-manage approximately twelve distributed engineers across backend, frontend, DevOps, QA, and documentation • Set engineering standards, conduct code and design reviews, manage release gates, hold one-to-ones, and provide growth and performance input • Design, build, and operate a managed Slurm service for research users • Own Slurm controller and accounting, partitions, login nodes, node onboarding, driver and CUDA baselines, stalled-job and node-health detection, drain and autohealing, storage visibility, identity, and isolation • Own Kubernetes cluster bootstrap and lifecycle on partner-provided bare metal • Implement NVIDIA GPU Operator and Network Operator, VM-based GPU isolation with KubeVirt and VFIO, upgrades, backup and recovery, and node replacement • Lead managed inference architecture, including serving, multi-GPU and multi-node parallelism, autoscaling, request routing, endpoint reliability, and confidential-compute-capable capacity • Establish metrics, logging, alerting, and SLOs across the control plane, GPU fleet, and application tiers • Lead incident response, post-incident reviews, and a sustainable on-call model • Serve as primary technical interface to infrastructure partners and vendors • Translate requirements into written specifications and acceptance tests, manage escalations, and provide technical input to capacity planning and hardware sourcing • Work with research, model-training, and product teams to translate workloads into platform requirements and broker capacity • Complete the platform team and set the technical bar for new engineers

🎯 Requirements

• Eight or more years of hands-on engineering experience, including at least three leading teams that build and operate infrastructure platforms other teams depend on • Bachelor's or Master's degree in computer science or engineering, or equivalent practical experience • Hands-on experience running Slurm at scale, including slurmctld, slurmdbd, partitions, QoS, priority, accounting, prolog and epilog, node health scripting, and upgrades with jobs on the system • Experience operating an HPC or GPU training cluster for a research population is ideally preferred • Experience operating NVIDIA GPU fleets on bare metal, including NVIDIA driver and CUDA lifecycle, Fabric Manager, NVSwitch, DCGM, MIG, node burn-in, and acceptance • Experience with InfiniBand, subnet configuration, RDMA, SR-IOV, and diagnosing multi-node NCCL performance problems • Deep Linux systems knowledge, including kernel modules and drivers, PCIe passthrough, vfio-pci, cgroups, namespaces, and performance tuning • Production Kubernetes operation, including control plane, upgrades, CNI, CSI, operators, custom controllers, and multi-tenancy design • Experience with HPC storage and data movement, including VAST, Lustre, NFS, node-local NVMe caching, and distributing large model weights and datasets • Experience with Prometheus, Grafana, Loki or equivalents, SLOs, incident response, and post-incident review • Working fluency in JavaScript and Node.js sufficient to review a control plane, CLI, and worker services and make architecture decisions • Experience delivering a platform with real users, such as multi-tenant IaaS/PaaS or a research computing service, including resource isolation, quotas, usage metering, and user-facing API/CLI surfaces • People management across time zones, cross-track review, written architecture decisions, and ability to communicate decisions to partners or executives • Excellent written and spoken English • Based between UTC and UTC+5:30 so the working day overlaps Europe and India • Desirable: Slurm operators on Kubernetes or Kubernetes-native schedulers • Desirable: modern serving stacks such as vLLM, SGLang, and TensorRT-LLM • Desirable: VM and container isolation, confidential computing, Cluster API, kubeadm, Cilium, NVSentinel-class autohealing, infrastructure as code, GitOps, GPU cloud/HPC/AI lab experience, distributed systems, and hardware-provider partnership experience

🏖️ Benefits

• Fully remote work arrangement • Occasional travel to partner sites and team events

Apply Now

Similar Jobs

🔥 21 hours ago

knowmad mood

1001 - 5000

💼 Consulting

🏥 Healthcare

📦 Logistics

Ingeniero Java backend senior desarrollando nuevos proyectos para un cliente bancario. Trabajo 100% remoto con Java, Spring, microservicios y AWS.

🗣️🇪🇸 Spanish Required

AWS

Cloud

Java

Kafka

Spring

Spring Boot

SpringBoot

🔥 23 hours ago

Mitek Systems

201 - 500

🤖 Artificial Intelligence

📋 Compliance

🔐 Security

Full Stack Engineering Team Lead guiding hands-on development of identity authentication and fraud prevention software. Leading architecture, cloud engineering, delivery, and engineer development.

AWS

Cloud

Cypress

Distributed Systems

EC2

Groovy

Java

JavaScript

Microservices

MongoDB

Python

React

Redux

Terraform

TypeScript

Go

🕒 3 days ago

knowmad mood

1001 - 5000

💼 Consulting

🏥 Healthcare

📦 Logistics

Senior Full Stack Engineer construyendo productos digitales desde cero para knowmad mood, líder en transformación digital. Desarrollo fullstack con React, Java e IA generativa.

🗣️🇪🇸 Spanish Required

Java

React

🕒 3 days ago

knowmad mood

1001 - 5000

💼 Consulting

🏥 Healthcare

📦 Logistics

Java Architect proporcionando liderazgo técnico backend para knowmad mood, compañía española de transformación digital. Definiendo arquitecturas, apoyando equipos e impulsando innovación con Java y Spring.

🗣️🇪🇸 Spanish Required

Grafana

Java

Kafka

Spring

🕒 3 days ago

knowmad mood

1001 - 5000

💼 Consulting

🏥 Healthcare

📦 Logistics

Senior Java backend engineer developing cloud-based microservices for knowmad mood’s banking client. Contributing ideas to new digital transformation projects in a fully remote role.

🗣️🇪🇸 Spanish Required

Azure

Cloud

Grafana

Java

Kafka

Prometheus

Spring

Spring Boot

SpringBoot