Technical Lead – GPU Infrastructure

🔥 4 minutes ago

🇦🇲 Armenia – Remote

⏰ Full Time

🟠 Senior

🧑‍💻 Full-stack Engineer

👻 Ghost score 25%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Tether.to

Tether.to

11 - 50 employees

Founded 2014

₿ Crypto

💳 Fintech

💸 Finance

Crypto • Fintech • Finance

Tether. to is a leading digital asset company that pioneers the use of stablecoins in the blockchain space. As the most widely adopted stablecoin, Tether tokens are designed to be pegged 1-to-1 with fiat currencies, offering a stable digital asset option for users. The platform facilitates these token transactions across multiple blockchains, enhancing cross-border transactions while maintaining transparency with daily records of total assets and reserves. Tether's initiatives include educational programs promoting digital asset usage, especially targeting regions like the Middle East, Turkey, and the Philippines. Tether thus positions itself as a disruptor in the traditional financial system by enabling a stable, efficient method of handling transactions in the digital currency world.

📋 Description

• Own the end-to-end architecture of Tether Data's Cosmic AC GPU compute and managed inference platform • Drive architecture proposals, high-level and low-level designs, reviews, and maintenance of the architectural baseline • Lead and line-manage a distributed team of about twelve engineers across backend, frontend, DevOps, QA, and documentation • Set engineering standards, conduct code and design reviews, manage release gates, hold one-to-ones, and provide growth and performance input • Design, build, and operate a managed Slurm service for research users • Own Slurm controller and accounting, partitions and login nodes, node onboarding, driver and CUDA baselines, stalled-job and node-health detection, drain and autohealing, storage visibility, identity, and isolation • Own Kubernetes cluster bootstrap and lifecycle on partner-provided bare metal, NVIDIA GPU Operator and Network Operator, VM-based GPU isolation, and day-2 operations • Design managed inference serving architecture, multi-GPU and multi-node parallelism, autoscaling, request routing, endpoint reliability, and confidential-compute-capable capacity • Establish metrics, logging, alerting, and SLOs across the control plane, GPU fleet, and application tiers • Participate in incident response, post-incident reviews, and creation of a sustainable on-call model • Serve as the primary technical interface to infrastructure partners and vendors • Translate requirements into written specifications and acceptance tests, run escalations to closure, and contribute to capacity planning and hardware sourcing • Work with research, model-training, and product teams to translate workloads into platform requirements and broker capacity • Complete the platform team and set the technical bar for new engineers

🎯 Requirements

• Eight or more years of hands-on engineering experience, including at least three leading teams that build and operate infrastructure platforms other teams depend on • Bachelor's or Master's degree in computer science or engineering, or equivalent practical experience • Hands-on experience operating Slurm at scale, including slurmctld and slurmdbd, partitions, QoS and priority, accounting, prolog and epilog, node health scripting, and upgrades with jobs on the system • Experience operating an HPC or GPU training cluster for a research population is ideally preferred • Experience operating NVIDIA GPU fleets on bare metal, including NVIDIA driver and CUDA lifecycle, Fabric Manager and NVSwitch behaviour, DCGM-based health and utilisation, MIG, node burn-in and acceptance • Experience with InfiniBand fabric and subnet configuration, RDMA, SR-IOV, and diagnosing multi-node NCCL performance problems • Deep Linux systems knowledge, including kernel modules and drivers, PCIe passthrough and vfio-pci, cgroups and namespaces, and performance tuning • Production Kubernetes operation, including control plane, upgrades, CNI and CSI, operators and custom controllers, and multi-tenancy design • Experience with HPC storage and data movement, including VAST, Lustre, NFS, node-local NVMe caching, and distributing large model weights and datasets • Experience with Prometheus, Grafana and Loki or equivalents, SLOs, incident response and post-incident review • Working fluency in JavaScript and Node.js sufficient to review and make architecture decisions on control plane, CLI and worker services • Experience delivering a shipped platform with real users, such as a multi-tenant IaaS/PaaS or research computing service, including resource isolation, quotas, usage metering, and user-facing API and CLI surfaces • People management across time zones, cross-track review, written architecture decisions, and ability to communicate decisions to partners or executives • Excellent written and spoken English • Based between UTC and UTC+5:30 • Desirable experience with Slurm operators on Kubernetes or Kubernetes-native schedulers • Desirable experience with vLLM, SGLang, TensorRT-LLM, GPU isolation, confidential computing, Cluster API, kubeadm, Cilium, NVSentinel-class autohealing, infrastructure as code, GitOps, GPU cloud/HPC/AI lab platforms, distributed systems, and hardware-provider contracts with acceptance tests

🏖️ Benefits

• Fully remote work • Occasional travel to partner sites and team events

Apply Now

Similar Jobs

🕒 September 10

Commit

501 - 1000

🔒 Cybersecurity

Senior full-stack developer building AI-first React and Node.js features for OTOFusion. Deploying production systems with Docker, Kubernetes, and AWS from Armenia.

AWS

Docker

JavaScript

Kubernetes

Node.js

React

React Native

TypeScript

🕒 August 21

Commit

501 - 1000

🔒 Cybersecurity

Senior Edge Software Engineer building high-performance edge-cloud infrastructure for an AI video security platform. Developing real-time video, networking, and on-device AI capabilities.

AWS

Cloud

Distributed Systems

Docker

Google Cloud Platform

Kafka

Rust

TypeScript

Go

🕒 August 12

Resourceful Talent Group

1 - 10

💼 Consulting

🏥 Healthcare

📣 Marketing

Senior Full-Stack Engineer building production React/TypeScript interfaces and Node.js backends for an established fintech-oriented engineering environment. Fully remote across listed international locations, working EST hours.

🗣️🇷🇺 Russian Required

AWS

JavaScript

Node.js

Postgres

React

TypeScript

🕒 August 12

eCom Solutions Inc

201 - 500

💼 Consulting

📦 Logistics

📣 Marketing

Senior Full-Stack Engineer building React/TypeScript interfaces and Node.js backends for production fintech systems. Deploying AWS applications and optimizing PostgreSQL databases remotely.

🗣️🇷🇺 Russian Required

AWS

JavaScript

Node.js

Postgres

React

TypeScript