Technical Lead – GPU Infrastructure

🔥 21 hours ago

🇺🇸 United States – Remote

⏰ Full Time

🟠 Senior

🧑‍💻 Full-stack Engineer

👻 Ghost score 25%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Tether.to

Tether.to

11 - 50 employees

Founded 2014

₿ Crypto

💳 Fintech

💸 Finance

Crypto • Fintech • Finance

Tether. to is a leading digital asset company that pioneers the use of stablecoins in the blockchain space. As the most widely adopted stablecoin, Tether tokens are designed to be pegged 1-to-1 with fiat currencies, offering a stable digital asset option for users. The platform facilitates these token transactions across multiple blockchains, enhancing cross-border transactions while maintaining transparency with daily records of total assets and reserves. Tether's initiatives include educational programs promoting digital asset usage, especially targeting regions like the Middle East, Turkey, and the Philippines. Tether thus positions itself as a disruptor in the traditional financial system by enabling a stable, efficient method of handling transactions in the digital currency world.

📋 Description

• Own the end-to-end platform architecture through architecture proposals, high-level and low-level designs, reviews, and maintenance of the baseline • Lead and line-manage approximately twelve distributed engineers across backend, frontend, DevOps, QA, and documentation • Set engineering standards, conduct code and design reviews, manage release gates, hold one-to-ones, and provide growth and performance input • Design, build, and operate a managed Slurm service for research users • Own Slurm controllers, accounting, partitions, login nodes, node onboarding, driver and CUDA baselines, stalled-job and node-health detection, drain, autohealing, storage visibility, identity, and isolation • Own Kubernetes cluster bootstrap and lifecycle on partner-provided bare metal • Implement NVIDIA GPU Operator, Network Operator, VM-based GPU isolation with KubeVirt and VFIO, upgrades, backup and recovery, and node replacement • Design managed inference serving architecture, multi-GPU and multi-node parallelism, autoscaling, request routing, endpoint reliability, and confidential-compute-capable capacity • Establish metrics, logging, alerting, and SLOs across the control plane, GPU fleet, and application tiers • Lead incident response, post-incident reviews, and a sustainable on-call model • Serve as primary technical interface to infrastructure partners and vendors • Translate requirements into written specifications and acceptance tests, manage escalations, and contribute to capacity planning and hardware sourcing • Work with research, model-training, and product teams to translate workloads into platform requirements and broker capacity • Complete the platform team and establish the technical bar for new engineers

🎯 Requirements

• Eight or more years of hands-on engineering experience • At least three years leading teams that build and operate infrastructure platforms other teams depend on • Bachelor's or Master's degree in computer science or engineering, or equivalent practical experience • Hands-on experience with Slurm at scale, including slurmctld, slurmdbd, partitions, QoS, priority, accounting, prolog and epilog, node health scripting, and upgrades with jobs running • Experience operating HPC or GPU training clusters for research users is ideally preferred • NVIDIA driver and CUDA lifecycle experience • Fabric Manager and NVSwitch experience on SXM systems • DCGM-based health and utilization, MIG, node burn-in and acceptance experience • InfiniBand fabric and subnet configuration, RDMA, SR-IOV, and diagnosing multi-node NCCL performance problems • Linux kernel modules and drivers, PCIe passthrough, vfio-pci, cgroups, namespaces, and performance tuning • Production Kubernetes operation, including control plane, upgrades, CNI, CSI, operators, custom controllers, and multi-tenancy design • HPC storage and data movement experience with VAST, Lustre, NFS, node-local NVMe caching, and distributing large model weights and datasets • Prometheus, Grafana, Loki or equivalents, SLOs, incident response, and post-incident review • Working fluency in JavaScript and Node.js sufficient to review and make architecture decisions • Experience with a shipped platform used by real users, such as multi-tenant IaaS, PaaS, or research computing service • People management across time zones, cross-track review, written architecture decisions, and partner/executive communication • Excellent written and spoken English • Desirable experience with Soperator, Slinky, Kueue, Volcano, KAI, Kubeflow Trainer, vLLM, SGLang, TensorRT-LLM, KubeVirt, Kata Containers, QEMU, KVM, Firecracker, confidential computing, Cluster API, kubeadm, Cilium, NVSentinel-class autohealing, infrastructure as code, GitOps, GPU cloud/HPC/AI lab platforms, distributed systems, and hardware-provider contracts

🏖️ Benefits

• Fully remote work arrangement • Occasional travel to partner sites and team events

Apply Now

Similar Jobs

🔥 21 hours ago

Daktronics

1001 - 5000

💼 Consulting

📦 Logistics

🏭 Manufacturing

Senior Principal Software Architect leading cloud, edge, and AI-enabled platform architecture at Daktronics, a digital LED display and audio systems company. Establishing standards, modernization strategy, and technical governance.

🇺🇸 United States – Remote

💵 $160k - $200k / year

⏰ Full Time

🟠 Senior

🧑‍💻 Full-stack Engineer

🔥 21 hours ago

soonami.io GmbH

1 - 10

🌐 Web 3

🤖 Artificial Intelligence

💳 Fintech

Senior Fullstack Developer advancing LuniNora’s online counseling marketplace launch. Building TypeScript, Firebase, SvelteKit, and Claude Code solutions for accessible counseling.

🇺🇸 United States – Remote

⏰ Full Time

🟠 Senior

🧑‍💻 Full-stack Engineer

🔥 22 hours ago

Embrace Software Inc

201 - 500

💼 Consulting

💸 Finance

Full-Stack .NET Developer building scalable .NET and Angular applications for XAP’s career and college planning SaaS platform. Supporting schools, state agencies, and students with education-planning technology.

🇺🇸 United States – Remote

⏰ Full Time

🟡 Mid-level

🟠 Senior

🧑‍💻 Full-stack Engineer

🕒 Yesterday

Nagarro

10,000+ employees

💼 Consulting

📣 Marketing

🏥 Healthcare

Senior Staff Engineer architecting scalable GenAI and Agentic AI solutions for Nagarro’s digital product engineering business. Designing, coding, and productionizing enterprise AI systems on Azure or AWS.

AWS

Azure

JavaScript

Neo4j

Python

PyTorch

React

Scikit-Learn

Tensorflow

🕒 Yesterday

FedWriters

201 - 500

🤝 B2B

🏛️ Government

📚 Education

Senior C++ Software Engineer advising and optimizing NOAA’s MRMS radar precipitation-estimation software. Developing code, architecture strategies, technical documentation, and quarterly status reports.

🇺🇸 United States – Remote

💵 $81k - $85k / year

⏰ Full Time

🟠 Senior

🧑‍💻 Full-stack Engineer