Technical Lead – GPU Infrastructure

🔥 19 horas atrás

🇺🇸 Estados Unidos – Remoto (EUA)

⏰ Tempo Integral

🟠 Sênior

🧑‍💻 Engenheiro Full-stack

👻 Score fantasma 25%

infoinfo

🗣️🇺🇸🇬🇧 Inglês obrigatório

Candidatar-se
Encontrar Vagas Remotas Similares

📊 Verifique sua pontuação de currículo para esta vaga

Melhore suas chances de conseguir uma entrevista verificando sua pontuação de currículo antes de se candidatar.

Logo of Tether.to

Tether.to

11 - 50 funcionários

Fundada em 2014

₿ Cripto

💳 Fintech

💸 Finanças

Crypto • Fintech • Finance

Tether. to é uma empresa líder em ativos digitais que pioneiriza o uso de stablecoins no ecossistema blockchain. Como a stablecoin mais amplamente adotada, os tokens Tether são projetados para serem lastreados 1:1 em moedas fiduciárias, oferecendo aos usuários uma opção de ativo digital estável. A plataforma facilita transações desses tokens em múltiplas blockchains, aprimorando operações transfronteiriças enquanto mantém transparência com registros diários do total de ativos e reservas. As iniciativas da Tether incluem programas educacionais que promovem o uso de ativos digitais, com foco especial em regiões como Oriente Médio, Turquia e Filipinas. Assim, a Tether se posiciona como um agente de disrupção no sistema financeiro tradicional ao viabilizar um método estável e eficiente de realizar transações no mundo das moedas digitais.

Descrição

• Own the end-to-end platform architecture through architecture proposals, high-level and low-level designs, reviews, and maintenance of the baseline • Lead and line-manage approximately twelve distributed engineers across backend, frontend, DevOps, QA, and documentation • Set engineering standards, conduct code and design reviews, manage release gates, hold one-to-ones, and provide growth and performance input • Design, build, and operate a managed Slurm service for research users • Own Slurm controllers, accounting, partitions, login nodes, node onboarding, driver and CUDA baselines, stalled-job and node-health detection, drain, autohealing, storage visibility, identity, and isolation • Own Kubernetes cluster bootstrap and lifecycle on partner-provided bare metal • Implement NVIDIA GPU Operator, Network Operator, VM-based GPU isolation with KubeVirt and VFIO, upgrades, backup and recovery, and node replacement • Design managed inference serving architecture, multi-GPU and multi-node parallelism, autoscaling, request routing, endpoint reliability, and confidential-compute-capable capacity • Establish metrics, logging, alerting, and SLOs across the control plane, GPU fleet, and application tiers • Lead incident response, post-incident reviews, and a sustainable on-call model • Serve as primary technical interface to infrastructure partners and vendors • Translate requirements into written specifications and acceptance tests, manage escalations, and contribute to capacity planning and hardware sourcing • Work with research, model-training, and product teams to translate workloads into platform requirements and broker capacity • Complete the platform team and establish the technical bar for new engineers

🎯 Requisitos

• Eight or more years of hands-on engineering experience • At least three years leading teams that build and operate infrastructure platforms other teams depend on • Bachelor's or Master's degree in computer science or engineering, or equivalent practical experience • Hands-on experience with Slurm at scale, including slurmctld, slurmdbd, partitions, QoS, priority, accounting, prolog and epilog, node health scripting, and upgrades with jobs running • Experience operating HPC or GPU training clusters for research users is ideally preferred • NVIDIA driver and CUDA lifecycle experience • Fabric Manager and NVSwitch experience on SXM systems • DCGM-based health and utilization, MIG, node burn-in and acceptance experience • InfiniBand fabric and subnet configuration, RDMA, SR-IOV, and diagnosing multi-node NCCL performance problems • Linux kernel modules and drivers, PCIe passthrough, vfio-pci, cgroups, namespaces, and performance tuning • Production Kubernetes operation, including control plane, upgrades, CNI, CSI, operators, custom controllers, and multi-tenancy design • HPC storage and data movement experience with VAST, Lustre, NFS, node-local NVMe caching, and distributing large model weights and datasets • Prometheus, Grafana, Loki or equivalents, SLOs, incident response, and post-incident review • Working fluency in JavaScript and Node.js sufficient to review and make architecture decisions • Experience with a shipped platform used by real users, such as multi-tenant IaaS, PaaS, or research computing service • People management across time zones, cross-track review, written architecture decisions, and partner/executive communication • Excellent written and spoken English • Desirable experience with Soperator, Slinky, Kueue, Volcano, KAI, Kubeflow Trainer, vLLM, SGLang, TensorRT-LLM, KubeVirt, Kata Containers, QEMU, KVM, Firecracker, confidential computing, Cluster API, kubeadm, Cilium, NVSentinel-class autohealing, infrastructure as code, GitOps, GPU cloud/HPC/AI lab platforms, distributed systems, and hardware-provider contracts

🏖️ Benefícios

• Fully remote work arrangement • Occasional travel to partner sites and team events

Candidatar-se

Vagas Similares

🔥 19 horas atrás

Daktronics

1001 - 5000

💼 Consultoria

📦 Logística

🏭 Manufatura

Senior Principal Software Architect leading cloud, edge, and AI-enabled platform architecture at Daktronics, a digital LED display and audio systems company. Establishing standards, modernization strategy, and technical governance.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $160.000 - $200.000 / ano

⏰ Tempo Integral

🟠 Sênior

🧑‍💻 Engenheiro Full-stack

🗣️🇺🇸🇬🇧 Inglês obrigatório

🔥 19 horas atrás

soonami.io GmbH

1 - 10

🌐 Web 3

🤖 Inteligência Artificial

💳 Fintech

Senior Fullstack Developer advancing LuniNora’s online counseling marketplace launch. Building TypeScript, Firebase, SvelteKit, and Claude Code solutions for accessible counseling.

🇺🇸 Estados Unidos – Remoto (EUA)

⏰ Tempo Integral

🟠 Sênior

🧑‍💻 Engenheiro Full-stack

🗣️🇺🇸🇬🇧 Inglês obrigatório

🔥 20 horas atrás

Embrace Software Inc

201 - 500

💼 Consultoria

💸 Finanças

Full-Stack .NET Developer building scalable .NET and Angular applications for XAP’s career and college planning SaaS platform. Supporting schools, state agencies, and students with education-planning technology.

🇺🇸 Estados Unidos – Remoto (EUA)

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

🧑‍💻 Engenheiro Full-stack

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Ontem

Nagarro

10.000+ funcionários

💼 Consultoria

📣 Marketing

🏥 Saúde

Senior Staff Engineer architecting scalable GenAI and Agentic AI solutions for Nagarro’s digital product engineering business. Designing, coding, and productionizing enterprise AI systems on Azure or AWS.

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Ontem

FedWriters

201 - 500

🤝 B2B

🏛️ Governo

📚 Educação

Senior C++ Software Engineer advising and optimizing NOAA’s MRMS radar precipitation-estimation software. Developing code, architecture strategies, technical documentation, and quarterly status reports.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $81.000 - $85.000 / ano

⏰ Tempo Integral

🟠 Sênior

🧑‍💻 Engenheiro Full-stack

🗣️🇺🇸🇬🇧 Inglês obrigatório

Perl

Python

C++