Technical Lead – GPU Infrastructure

🔥 12 hours ago

🇮🇳 India – Remote

⏰ Full Time

🟠 Senior

🧑‍💻 Full-stack Engineer

👻 Ghost score 25%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Tether.to

Tether.to

11 - 50 employees

Founded 2014

₿ Crypto

💳 Fintech

💸 Finance

Crypto • Fintech • Finance

Tether. to is a leading digital asset company that pioneers the use of stablecoins in the blockchain space. As the most widely adopted stablecoin, Tether tokens are designed to be pegged 1-to-1 with fiat currencies, offering a stable digital asset option for users. The platform facilitates these token transactions across multiple blockchains, enhancing cross-border transactions while maintaining transparency with daily records of total assets and reserves. Tether's initiatives include educational programs promoting digital asset usage, especially targeting regions like the Middle East, Turkey, and the Philippines. Tether thus positions itself as a disruptor in the traditional financial system by enabling a stable, efficient method of handling transactions in the digital currency world.

📋 Description

• Own Cosmic AC platform architecture end to end through architecture proposals, high-level and low-level designs, reviews, and maintaining the baseline • Lead and line-manage a distributed team of about twelve engineers across backend, frontend, DevOps, QA, and documentation • Set engineering standards, conduct code and design reviews, manage release gates, hold one-to-ones, and provide growth and performance input • Design, build, and operate a managed Slurm service for research users • Own controller and accounting, partitions and login nodes, node onboarding, driver and CUDA baselines and upgrades, stalled-job and node-health detection, drain and autohealing, storage visibility, identity, and isolation • Own Kubernetes cluster bootstrap and lifecycle on partner-provided bare metal • Implement NVIDIA GPU Operator and Network Operator, VM-based GPU isolation with KubeVirt and VFIO, upgrades, backup and recovery, and node replacement • Own managed inference architecture, including serving, multi-GPU and multi-node parallelism, autoscaling, request routing, endpoint reliability, and confidential-compute-capable capacity • Establish metrics, logging, alerting, SLOs, incident response, post-incident reviews, and a sustainable on-call model • Serve as primary technical interface to infrastructure partners and vendors • Translate requirements into written specifications and acceptance tests, manage escalations, and contribute to capacity planning and hardware sourcing • Work with research, model-training, and product teams to translate workloads into platform requirements and broker scarce capacity • Complete the platform team and set the technical bar for new hires

🎯 Requirements

• Eight or more years of hands-on engineering experience • At least three years leading teams that build and operate infrastructure platforms other teams depend on • Bachelor's or Master's degree in computer science or engineering, or equivalent practical experience • Hands-on experience running Slurm at scale, including slurmctld, slurmdbd, partitions, QoS, priority, accounting, prolog and epilog, node health scripting, and upgrades with jobs running • Experience operating an HPC or GPU training cluster for a research population is ideally preferred • Bare-metal NVIDIA GPU fleet operations, including NVIDIA driver and CUDA lifecycle, Fabric Manager, NVSwitch, DCGM, MIG, node burn-in, and acceptance • InfiniBand fabric and subnet configuration, RDMA, SR-IOV, and diagnosing multi-node NCCL performance problems • Linux systems expertise, including kernel modules and drivers, PCIe passthrough, vfio-pci, cgroups, namespaces, and performance tuning • Production Kubernetes operations, including control plane, upgrades, CNI, CSI, operators, custom controllers, and multi-tenancy design • HPC storage and data movement using VAST, Lustre, NFS, node-local NVMe caching, and distribution of large model weights and datasets • Observability and operations with Prometheus, Grafana, Loki or equivalents, SLOs, incident response, and post-incident review • Working fluency in JavaScript and Node.js sufficient to review control-plane, CLI, and worker services and make architecture decisions • Experience shipping a platform with real users, such as multi-tenant IaaS/PaaS or research computing services, including isolation, quotas, usage metering, APIs, and CLIs • People management across time zones, cross-track review, written architecture decisions, and ability to reject partner or executive requests with reasons • Excellent written and spoken English • Ability to work remotely from a location between UTC and UTC+5:30

🏖️ Benefits

• Fully remote work arrangement • Occasional travel to partner sites and team events

Apply Now

Similar Jobs

🔥 22 hours ago

Cyberhaven

51 - 200

🔒 Cybersecurity

☁️ SaaS

🏢 Enterprise

Backend software engineer building scalable, secure systems for Cyberhaven’s data-security platform. Designing Go microservices that process billions of real-time events across enterprise endpoints.

BigQuery

Docker

Kubernetes

Microservices

Redis

Go

🕒 Yesterday

CentralApp

51 - 200

💼 Consulting

📦 Logistics

📣 Marketing

Product Engineer owning ambiguous product problems end to end for CentralApp’s AI-assisted website platform. Building frontend solutions, shaping APIs, and collaborating on Haskell and Node.js backend systems.

AWS

Haskell

JavaScript

Node.js

React

🕒 2 days ago

Kantiv

51 - 200

🏗️ Construction

Senior Software Engineer designing and operating Python chat and agent systems. Improving LLM reliability, observability, coding-agent workflows, and team engineering practices.

Python

🕒 2 days ago

Weekday (YC W21)

11 - 50

💼 Consulting

👥 HR Tech

☁️ SaaS

Java Full Stack Developer building complex cloud applications, web services, and software interfaces for a Weekday client. Leading critical lifecycle work with Java, Vue3, AWS, and Spring Boot.

AWS

Cloud

Java

Spring

Spring Boot

SpringBoot

🕒 2 days ago

Weekday (YC W21)

11 - 50

💼 Consulting

👥 HR Tech

☁️ SaaS

Java Full Stack Engineer building complex cloud applications, interfaces, and web services for a Weekday client. Leading critical software development using Java, Vue3, AWS, and microservices.

AWS

Cloud

Java

Microservices