Technical Lead – GPU Infrastructure

🕒 Setembro 17

🇺🇸 Estados Unidos – Remoto (EUA)

⏰ Tempo Integral

🟠 Sênior

🧑‍💻 Engenheiro Full-stack

👻 Score fantasma 25%

infoinfo

🗣️🇺🇸🇬🇧 Inglês obrigatório

Candidatar-se
Encontrar Vagas Remotas Similares

📊 Verifique sua pontuação de currículo para esta vaga

Melhore suas chances de conseguir uma entrevista verificando sua pontuação de currículo antes de se candidatar.

Logo of Tether.to

Tether.to

11 - 50 funcionários

Fundada em 2014

₿ Cripto

💳 Fintech

💸 Finanças

Crypto • Fintech • Finance

Tether. to é uma empresa líder em ativos digitais que pioneiriza o uso de stablecoins no ecossistema blockchain. Como a stablecoin mais amplamente adotada, os tokens Tether são projetados para serem lastreados 1:1 em moedas fiduciárias, oferecendo aos usuários uma opção de ativo digital estável. A plataforma facilita transações desses tokens em múltiplas blockchains, aprimorando operações transfronteiriças enquanto mantém transparência com registros diários do total de ativos e reservas. As iniciativas da Tether incluem programas educacionais que promovem o uso de ativos digitais, com foco especial em regiões como Oriente Médio, Turquia e Filipinas. Assim, a Tether se posiciona como um agente de disrupção no sistema financeiro tradicional ao viabilizar um método estável e eficiente de realizar transações no mundo das moedas digitais.

Descrição

• Own the end-to-end platform architecture through proposals, high-level and low-level designs, reviews, and baseline maintenance • Lead and line-manage approximately twelve distributed engineers across backend, frontend, DevOps, QA, and documentation • Set engineering standards, conduct code and design reviews, manage release gates, hold one-to-ones, and provide growth and performance input • Design, build, and operate a managed Slurm service for research users • Own Slurm controllers, accounting, partitions, login nodes, node onboarding, driver and CUDA baselines, health detection, autohealing, storage visibility, identity, and isolation • Own Kubernetes cluster bootstrap and lifecycle on partner-provided bare metal, NVIDIA GPU and Network Operators, VM-based GPU isolation, upgrades, backup, recovery, and node replacement • Design managed inference architecture, including multi-GPU and multi-node parallelism, autoscaling, request routing, endpoint reliability, and confidential-compute-capable capacity • Establish metrics, logging, alerting, SLOs, incident response, post-incident reviews, and a sustainable on-call model • Serve as primary technical interface to infrastructure partners and vendors • Translate requirements into written specifications and acceptance tests, manage escalations, and contribute to capacity planning and hardware sourcing • Work with research, model-training, and product teams to translate workloads into platform requirements and broker scarce capacity • Complete the platform team and set the technical bar for new engineers

🎯 Requisitos

• Eight or more years of hands-on engineering experience • At least three years leading teams that build and operate infrastructure platforms other teams depend on • Bachelor's or Master's degree in computer science or engineering, or equivalent practical experience • Hands-on experience running Slurm at scale, including slurmctld, slurmdbd, partitions, QoS, priority, accounting, prolog and epilog, node health scripting, and upgrades with jobs running • Experience operating an HPC or GPU training cluster for a research population is ideally preferred • Experience operating NVIDIA GPU fleets on bare metal, including driver and CUDA lifecycle, Fabric Manager, NVSwitch, DCGM, MIG, node burn-in, and acceptance • Experience with InfiniBand, subnet configuration, RDMA, SR-IOV, and diagnosing multi-node NCCL performance problems • Deep Linux systems knowledge, including kernel modules, drivers, PCIe passthrough, vfio-pci, cgroups, namespaces, and performance tuning • Production Kubernetes operations, including control plane, upgrades, CNI, CSI, operators, custom controllers, and multi-tenancy design • Experience with HPC storage and data movement, including VAST, Lustre, NFS, node-local NVMe caching, and distributing large model weights and datasets • Experience with Prometheus, Grafana, Loki or equivalents, SLOs, incident response, and post-incident review • Working fluency in JavaScript and Node.js sufficient to review and make architecture decisions on control plane, CLI, and worker services • Experience shipping a platform with real users, such as multi-tenant IaaS/PaaS or research computing service • People management across time zones, cross-track review, written architecture decisions, and technical judgment with partners and executives • Excellent written and spoken English • Desirable experience with Slurm operators on Kubernetes or Kubernetes-native schedulers • Desirable experience with modern serving stacks such as vLLM, SGLang, and TensorRT-LLM • Desirable experience with KubeVirt, Kata Containers, QEMU/KVM, Firecracker, confidential computing, Cluster API, kubeadm, Cilium, NVSentinel-class autohealing, infrastructure as code, GitOps, GPU cloud/HPC/AI lab platforms, distributed systems, and hardware-provider relationships

🏖️ Benefícios

• Fully remote work • Occasional travel to partner sites and team events

Candidatar-se

Vagas Similares

🕒 Setembro 17

Ledger Run

51 - 200

🏥 Saúde

💊 Farmacêutico

☁️ SaaS

Tech Lead guiding technical governance, hands-on engineering, and team execution at Ledger Run Inc., a SaaS company. Leading Java and Spring Boot development while mentoring engineers.

🇺🇸 Estados Unidos – Remoto (EUA)

💰 Private equity em 2024-10

⏰ Tempo Integral

🟠 Sênior

🧑‍💻 Engenheiro Full-stack

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Setembro 17

Zscaler

5001 - 10000

🔒 Cibersegurança

☁️ SaaS

🏢 Corporativo

Staff Applied AI Engineer building secure, production-grade AI solutions for Zscaler’s cloud security platform. Advising business leaders, shaping roadmaps, and establishing scalable engineering standards.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $180.000 - $225.000 / ano

💰 Secondary Market em 2017-11

⏰ Tempo Integral

🟠 Sênior

🧑‍💻 Engenheiro Full-stack

🦅 Patrocina Visto H1B

infoinfo

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Setembro 17

Palo Alto Networks

10.000+ funcionários

🔒 Cibersegurança

🏢 Corporativo

Backend alerting architect engineering Chronosphere’s cloud observability platform. Building Go and Kubernetes features that help developers diagnose and resolve incidents.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $183.600 - $297.000 / ano

💰 $1.000.000 Seed Round - Morta Security em 2013-02

⏰ Tempo Integral

🟠 Sênior

🧑‍💻 Engenheiro Full-stack

🦅 Patrocina Visto H1B

infoinfo

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Setembro 17

Humana

10.000+ funcionários

🏥 Saúde

🛡️ Seguros

⚕️ Seguro de Saúde

Senior Software Engineer building secure SSIS workflows and C#/.NET APIs for Humana’s healthcare data integrations. Supporting file transfers, SQL processing, and Record Management System connectivity.

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Setembro 17

LMI

1001 - 5000

📦 Logística

🏥 Saúde

🎖️ Defesa

Full Stack Developer building IronGate’s secure data acquisition platform for U.S. federal agencies. Developing Angular, React, Node.js, PostgreSQL, Kubernetes, and AWS capabilities.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $122.209 - $211.317 / ano

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

🧑‍💻 Engenheiro Full-stack

🦅 Patrocina Visto H1B

infoinfo

🗣️🇺🇸🇬🇧 Inglês obrigatório