Technical Lead – GPU Infrastructure

🕒 vor 12 Tagen

🇺🇸 Vereinigte Staaten – Remote

⏰ Vollzeit

🟠 Senior

🧑‍💻 Full-Stack-Entwickler

👻 Geisterscore 25%

infoinfo

🗣️🇺🇸🇬🇧 Englisch erforderlich

Jetzt Bewerben
Ähnliche Remote-Jobs finden

📊 Überprüfen Sie Ihre Lebenslauf-Bewertung für diese Stelle

Verbessern Sie Ihre Chancen auf ein Vorstellungsgespräch, indem Sie Ihre Lebenslauf-Bewertung vor der Bewerbung überprüfen.

Logo of Tether.to

Tether.to

11 - 50 Mitarbeiter

Gegründet 2014

₿ Crypto

💳 Fintech

💸 Finanzen

Crypto • Fintech • Finance

Tether. to ist ein führendes Unternehmen für digitale Vermögenswerte, das die Nutzung von Stablecoins im Blockchain-Bereich vorantreibt. Als der am weitesten verbreitete Stablecoin sind Tether-Token darauf ausgelegt, im Verhältnis 1:1 an Fiatwährungen gekoppelt zu sein, und bieten Nutzern eine stabile digitale Vermögensoption. Die Plattform erleichtert diese Token-Transaktionen über mehrere Blockchains hinweg und fördert grenzüberschreitende Transaktionen, während durch tägliche Aufzeichnungen der Gesamtvermögenswerte und Reserven Transparenz gewahrt wird. Zu den Initiativen von Tether gehören Bildungsprogramme zur Förderung der Nutzung digitaler Vermögenswerte, insbesondere in Regionen wie dem Nahen Osten, der Türkei und den Philippinen. Tether positioniert sich somit als Disruptor des traditionellen Finanzsystems, indem es eine stabile, effiziente Methode für Transaktionen in der digitalen Währungswelt ermöglicht.

Beschreibung

• Own the end-to-end platform architecture through proposals, high-level and low-level designs, reviews, and baseline maintenance • Lead and line-manage approximately twelve distributed engineers across backend, frontend, DevOps, QA, and documentation • Set engineering standards, conduct code and design reviews, manage release gates, hold one-to-ones, and provide growth and performance input • Design, build, and operate a managed Slurm service for research users • Own Slurm controllers, accounting, partitions, login nodes, node onboarding, driver and CUDA baselines, health detection, autohealing, storage visibility, identity, and isolation • Own Kubernetes cluster bootstrap and lifecycle on partner-provided bare metal, NVIDIA GPU and Network Operators, VM-based GPU isolation, upgrades, backup, recovery, and node replacement • Design managed inference architecture, including multi-GPU and multi-node parallelism, autoscaling, request routing, endpoint reliability, and confidential-compute-capable capacity • Establish metrics, logging, alerting, SLOs, incident response, post-incident reviews, and a sustainable on-call model • Serve as primary technical interface to infrastructure partners and vendors • Translate requirements into written specifications and acceptance tests, manage escalations, and contribute to capacity planning and hardware sourcing • Work with research, model-training, and product teams to translate workloads into platform requirements and broker scarce capacity • Complete the platform team and set the technical bar for new engineers

🎯 Anforderungen

• Eight or more years of hands-on engineering experience • At least three years leading teams that build and operate infrastructure platforms other teams depend on • Bachelor's or Master's degree in computer science or engineering, or equivalent practical experience • Hands-on experience running Slurm at scale, including slurmctld, slurmdbd, partitions, QoS, priority, accounting, prolog and epilog, node health scripting, and upgrades with jobs running • Experience operating an HPC or GPU training cluster for a research population is ideally preferred • Experience operating NVIDIA GPU fleets on bare metal, including driver and CUDA lifecycle, Fabric Manager, NVSwitch, DCGM, MIG, node burn-in, and acceptance • Experience with InfiniBand, subnet configuration, RDMA, SR-IOV, and diagnosing multi-node NCCL performance problems • Deep Linux systems knowledge, including kernel modules, drivers, PCIe passthrough, vfio-pci, cgroups, namespaces, and performance tuning • Production Kubernetes operations, including control plane, upgrades, CNI, CSI, operators, custom controllers, and multi-tenancy design • Experience with HPC storage and data movement, including VAST, Lustre, NFS, node-local NVMe caching, and distributing large model weights and datasets • Experience with Prometheus, Grafana, Loki or equivalents, SLOs, incident response, and post-incident review • Working fluency in JavaScript and Node.js sufficient to review and make architecture decisions on control plane, CLI, and worker services • Experience shipping a platform with real users, such as multi-tenant IaaS/PaaS or research computing service • People management across time zones, cross-track review, written architecture decisions, and technical judgment with partners and executives • Excellent written and spoken English • Desirable experience with Slurm operators on Kubernetes or Kubernetes-native schedulers • Desirable experience with modern serving stacks such as vLLM, SGLang, and TensorRT-LLM • Desirable experience with KubeVirt, Kata Containers, QEMU/KVM, Firecracker, confidential computing, Cluster API, kubeadm, Cilium, NVSentinel-class autohealing, infrastructure as code, GitOps, GPU cloud/HPC/AI lab platforms, distributed systems, and hardware-provider relationships

🏖️ Vorteile

• Fully remote work • Occasional travel to partner sites and team events

Jetzt Bewerben

Ähnliche Jobs

🕒 vor 12 Tagen

Ledger Run

51 - 200

🏥 Gesundheitswesen

💊 Pharmazie

☁️ SaaS

Tech Lead guiding technical governance, hands-on engineering, and team execution at Ledger Run Inc., a SaaS company. Leading Java and Spring Boot development while mentoring engineers.

🇺🇸 Vereinigte Staaten – Remote

💰 Private equity im 2024-10

⏰ Vollzeit

🟠 Senior

🧑‍💻 Full-Stack-Entwickler

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 12 Tagen

Zscaler

5001 - 10000

🔒 Cybersecurity

☁️ SaaS

🏢 Unternehmen

Staff Applied AI Engineer building secure, production-grade AI solutions for Zscaler’s cloud security platform. Advising business leaders, shaping roadmaps, and establishing scalable engineering standards.

🇺🇸 Vereinigte Staaten – Remote

💵 $180.000 - $225.000 / Jahr

💰 Secondary Market im 2017-11

⏰ Vollzeit

🟠 Senior

🧑‍💻 Full-Stack-Entwickler

🦅 H1B-Visum-Sponsor

infoinfo

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 12 Tagen

Palo Alto Networks

10.000+ Mitarbeiter

🔒 Cybersecurity

🏢 Unternehmen

Backend alerting architect engineering Chronosphere’s cloud observability platform. Building Go and Kubernetes features that help developers diagnose and resolve incidents.

🇺🇸 Vereinigte Staaten – Remote

💵 $183.600 - $297.000 / Jahr

💰 €1.000.000 Seed Round - Morta Security im 2013-02

⏰ Vollzeit

🟠 Senior

🧑‍💻 Full-Stack-Entwickler

🦅 H1B-Visum-Sponsor

infoinfo

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 12 Tagen

Humana

10.000+ Mitarbeiter

🏥 Gesundheitswesen

🛡️ Versicherung

⚕️ Krankenversicherung

Senior Software Engineer building secure SSIS workflows and C#/.NET APIs for Humana’s healthcare data integrations. Supporting file transfers, SQL processing, and Record Management System connectivity.

🗣️🇺🇸🇬🇧 Englisch erforderlich

🕒 vor 12 Tagen

LMI

1001 - 5000

📦 Logistik

🏥 Gesundheitswesen

🎖️ Verteidigung

Full Stack Developer building IronGate’s secure data acquisition platform for U.S. federal agencies. Developing Angular, React, Node.js, PostgreSQL, Kubernetes, and AWS capabilities.

🇺🇸 Vereinigte Staaten – Remote

💵 $122.209 - $211.317 / Jahr

⏰ Vollzeit

🟡 Mittelstufe

🟠 Senior

🧑‍💻 Full-Stack-Entwickler

🦅 H1B-Visum-Sponsor

infoinfo

🗣️🇺🇸🇬🇧 Englisch erforderlich