Technical Lead – GPU Infrastructure

🔥 1 minute ago

🇮🇳 India – Remote

⏰ Full Time

🟠 Senior

🧑‍💻 Full-stack Engineer

👻 Ghost score 25%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Tether.to

Tether.to

11 - 50 employees

Founded 2014

₿ Crypto

💳 Fintech

💸 Finance

Crypto • Fintech • Finance

Tether. to is a leading digital asset company that pioneers the use of stablecoins in the blockchain space. As the most widely adopted stablecoin, Tether tokens are designed to be pegged 1-to-1 with fiat currencies, offering a stable digital asset option for users. The platform facilitates these token transactions across multiple blockchains, enhancing cross-border transactions while maintaining transparency with daily records of total assets and reserves. Tether's initiatives include educational programs promoting digital asset usage, especially targeting regions like the Middle East, Turkey, and the Philippines. Tether thus positions itself as a disruptor in the traditional financial system by enabling a stable, efficient method of handling transactions in the digital currency world.

📋 Description

• Own Cosmic AC's end-to-end platform architecture through architecture proposals, high-level and low-level designs, reviews, and maintenance of the baseline • Lead and line-manage approximately twelve distributed engineers across backend, frontend, DevOps, QA, and documentation • Establish engineering standards, conduct code and design reviews, manage release gates, hold one-to-ones, and provide growth and performance input • Design, build, and operate a managed Slurm service for research users • Manage Slurm controllers, accounting, partitions, login nodes, node onboarding, driver and CUDA baselines, stalled-job and node-health detection, draining, autohealing, storage visibility, identity, and isolation • Own Kubernetes cluster bootstrap and lifecycle on partner-provided bare metal • Implement NVIDIA GPU Operator and Network Operator, VM-based GPU isolation using KubeVirt and VFIO, upgrades, backup and recovery, and node replacement • Design managed inference architecture with multi-GPU and multi-node parallelism, autoscaling, request routing, endpoint reliability, and confidential-compute-capable capacity • Own metrics, logging, alerting, and SLOs across the control plane, GPU fleet, and application tiers • Lead incident response, post-incident reviews, and a sustainable on-call model • Serve as primary technical interface to infrastructure partners and vendors • Translate requirements into written specifications and acceptance tests, manage escalations, and provide input to capacity planning and hardware sourcing • Work with research, model-training, and product teams to translate workloads into platform requirements and broker capacity • Complete the platform team and set the technical bar for new engineers

🎯 Requirements

• Eight or more years of hands-on engineering experience • At least three years leading teams that build and operate infrastructure platforms other teams depend on • Bachelor's or Master's degree in computer science or engineering, or equivalent practical experience • Hands-on experience running Slurm at scale, including slurmctld, slurmdbd, partitions, QoS and priority, accounting, prolog and epilog, node health scripting, and upgrades with jobs on the system • Experience operating HPC or GPU training clusters for research users is ideally desired • Experience operating NVIDIA GPU fleets on bare metal, including NVIDIA driver and CUDA lifecycle, Fabric Manager, NVSwitch, DCGM, MIG, node burn-in, and acceptance • Experience with InfiniBand, subnet configuration, RDMA, SR-IOV, and diagnosing multi-node NCCL performance issues • Deep Linux systems knowledge, including kernel modules and drivers, PCIe passthrough, vfio-pci, cgroups, namespaces, and performance tuning • Production Kubernetes operations experience, including control plane, upgrades, CNI, CSI, operators, custom controllers, and multi-tenancy • Experience with HPC storage and data movement, including VAST, Lustre, NFS, node-local NVMe caching, and distributing large model weights and datasets • Experience with Prometheus, Grafana, Loki or equivalents, SLOs, incident response, and post-incident review • Working fluency in JavaScript and Node.js sufficient to review and make architecture decisions on control plane, CLI, and worker services • Experience delivering a shipped platform with real users, such as multi-tenant IaaS/PaaS or research computing service • People management across time zones, cross-track review, written architecture decisions, and ability to challenge partners or executives with reasons • Excellent written and spoken English • Must be based between UTC and UTC+5:30 • Occasional travel to partner sites and team events • Desirable experience with Soperator, Slinky, Kueue, Volcano, KAI, Kubeflow Trainer, vLLM, SGLang, TensorRT-LLM, KubeVirt, Kata Containers, QEMU/KVM, Firecracker, confidential computing, Cluster API, kubeadm, Cilium, NVSentinel-class autohealing, infrastructure as code, GitOps, GPU cloud/HPC/AI lab platforms, peer-to-peer or distributed systems, and hardware-provider partnerships

🏖️ Benefits

• Fully remote work arrangement • Occasional travel to partner sites and team events • Global, distributed team environment • Line-management and growth/performance development opportunities

Apply Now

Similar Jobs

🔥 4 hours ago

HighLevel

201 - 500

💼 Consulting

📦 Logistics

☁️ SaaS

Software Development Engineer III building scalable frontend, backend, billing, and onboarding systems for HighLevel’s AI-powered business operating system. Driving architecture, reliability, and performance.

Angular

Bootstrap

Distributed Systems

ElasticSearch

Google Cloud Platform

JavaScript

Kafka

Microservices

MobX

MongoDB

Node.js

RabbitMQ

React

Redis

Redux

SDLC

TypeScript

Vue.js

🔥 4 hours ago

HighLevel

201 - 500

💼 Consulting

📦 Logistics

☁️ SaaS

Senior full stack engineer building scalable AI-powered social media systems for HighLevel. Owning distributed systems, data pipelines, media processing, and customer-facing web experiences.

Angular

BigQuery

Distributed Systems

ElasticSearch

ETL

JavaScript

MongoDB

Node.js

React

Redis

TypeScript

Vue.js

🔥 9 hours ago

Empower

10,000+ employees

💸 Finance

💳 Fintech

👥 B2C

Senior data engineer building scalable Snowflake, Redshift, and AWS data platforms for Empower. Leading pipelines, observability, governance, and AI/ML datasets.

Airflow

Amazon Redshift

Apache

AWS

Cloud

ETL

Jenkins

Python

SQL

Terraform

🔥 10 hours ago

PORCH 💚

1 - 10

💼 Consulting

🏥 Healthcare

🏨 Hospitality

Senior software engineer developing property-attribute extraction and backend services for Porch Group's homeowners insurance platform. Remote India role using Scala, SQL, cloud infrastructure, and microservices.

AWS

Cloud

Docker

Google Cloud Platform

Kafka

Kubernetes

Microservices

Postgres

Scala

SQL

🔥 11 hours ago

HighLevel

201 - 500

💼 Consulting

📦 Logistics

☁️ SaaS

Backend engineer owning HighLevel’s CRM Opportunities platform. Building scalable APIs, distributed workflows, and reliable AI-assisted customer features.

Assembly

Distributed Systems

ElasticSearch

JavaScript

MongoDB

Node.js

NoSQL

Postgres

SQL

Vue.js

Go