Technical Lead – GPU Infrastructure

🔥 20 hours ago

🇮🇪 Ireland – Remote

⏰ Full Time

🟠 Senior

🧑‍💻 Full-stack Engineer

👻 Ghost score 25%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Tether.to

Tether.to

11 - 50 employees

Founded 2014

₿ Crypto

💳 Fintech

💸 Finance

Crypto • Fintech • Finance

Tether. to is a leading digital asset company that pioneers the use of stablecoins in the blockchain space. As the most widely adopted stablecoin, Tether tokens are designed to be pegged 1-to-1 with fiat currencies, offering a stable digital asset option for users. The platform facilitates these token transactions across multiple blockchains, enhancing cross-border transactions while maintaining transparency with daily records of total assets and reserves. Tether's initiatives include educational programs promoting digital asset usage, especially targeting regions like the Middle East, Turkey, and the Philippines. Tether thus positions itself as a disruptor in the traditional financial system by enabling a stable, efficient method of handling transactions in the digital currency world.

📋 Description

• Own the end-to-end architecture of Tether Data's Cosmic AC GPU compute and managed inference platform • Create and maintain architecture proposals, high-level designs, and low-level designs through review • Lead and line-manage approximately twelve distributed engineers across backend, frontend, DevOps, QA, and documentation • Set engineering standards, conduct code and design reviews, manage release gates, hold one-to-ones, and provide growth and performance input • Design, build, and operate a managed Slurm service for research users • Own controller and accounting, partitions and login nodes, node onboarding, driver and CUDA baselines, stalled-job and node-health detection, draining, autohealing, storage visibility, identity, and isolation • Own Kubernetes cluster bootstrap and lifecycle on partner-provided bare metal • Manage NVIDIA GPU Operator, Network Operator, VM-based GPU isolation, KubeVirt, VFIO, upgrades, backup and recovery, and node replacement • Own managed inference architecture, multi-GPU and multi-node parallelism, autoscaling, request routing, endpoint reliability, and confidential-compute-capable capacity • Establish metrics, logging, alerting, and SLOs across control plane, GPU fleet, and application tiers • Lead incident response, post-incident reviews, and a sustainable on-call model • Serve as primary technical interface to infrastructure partners and vendors • Translate requirements into written specifications and acceptance tests, manage escalations, and contribute to capacity planning and hardware sourcing • Work with research, model-training, and product teams to translate workloads into platform requirements and broker scarce capacity • Complete the platform team and set the technical bar for new engineers • Own implementation and delivery plans within a fixed first-six-month delivery window

🎯 Requirements

• Eight or more years of hands-on engineering experience • At least three years leading teams that build and operate infrastructure platforms other teams depend on • Bachelor's or Master's degree in computer science or engineering, or equivalent practical experience • Hands-on experience operating Slurm at scale, including slurmctld, slurmdbd, partitions, QoS, priority, accounting, prolog and epilog, node health scripting, and upgrades with jobs running • Experience operating an HPC or GPU training cluster for a research population is ideally desired • Experience operating NVIDIA GPU fleets on bare metal, including driver and CUDA lifecycle, Fabric Manager, NVSwitch, DCGM, MIG, node burn-in, and acceptance • Experience with InfiniBand, subnet configuration, RDMA, SR-IOV, and diagnosing multi-node NCCL performance problems • Deep Linux systems knowledge, including kernel modules and drivers, PCIe passthrough, vfio-pci, cgroups, namespaces, and performance tuning • Production Kubernetes operation, including control plane, upgrades, CNI, CSI, operators, custom controllers, and multi-tenancy design • Experience with HPC storage and data movement, including VAST, Lustre, NFS, node-local NVMe caching, and distributing large model weights and datasets • Experience with Prometheus, Grafana, Loki or equivalents, SLOs, incident response, and post-incident review • Working fluency in JavaScript and Node.js sufficient to review control plane, CLI, and worker services and make architecture decisions • Experience delivering a multi-tenant IaaS, PaaS, or research computing service with resource isolation, quotas, usage metering, and user-facing API and CLI surfaces • People management across time zones, cross-track review, written architecture decisions, and ability to challenge partners or executives with reasons • Excellent written and spoken English • Fully remote location based between UTC and UTC+5:30 • Occasional travel to partner sites and team events • Desirable experience with Slurm operators on Kubernetes or Kubernetes-native schedulers • Desirable experience with modern serving stacks such as vLLM, SGLang, and TensorRT-LLM • Desirable experience with VM and container isolation, confidential computing, Cluster API, kubeadm, Cilium, NVSentinel-class autohealing, infrastructure as code, GitOps, GPU cloud/HPC/AI lab platforms, distributed systems, and hardware-provider relationships

🏖️ Benefits

• Fully remote work • Occasional travel to partner sites and team events

Apply Now

Similar Jobs

🕒 September 4

HubSpot

1001 - 5000

🤝 B2B

☁️ SaaS

📣 Marketing

Staff Technical Lead building AI-powered marketing intelligence for HubSpot’s customer platform. Leading engineers and developing Java backend services, integrations, analytics, and intelligent recommendations.

Java

Kafka

MySQL

🕒 August 28

Arista Networks

1001 - 5000

🏢 Enterprise

📡 Telecommunications

Senior C++ Layer1 engineer developing physical-layer software for Arista’s cloud networking platforms. Building telemetry, automated tests, and hardware control features for high-speed networking.

🇮🇪 Ireland – Remote

💵 $140k - $200k / year

💰 $2.6M Post-IPO Debt on 2015-05

⏰ Full Time

🟠 Senior

🧑‍💻 Full-stack Engineer

Distributed Systems

Linux

Python

Unix

C++

🕒 August 26

Muck Rack

201 - 500

🤝 B2B

📱 Media

Senior fullstack engineer building AI-powered search experiences for Muck Rack’s PR and communications SaaS platform. Developing Vue, Python, Django, and data-intensive search systems.

🇮🇪 Ireland – Remote

💵 €95k - €110k / year

💰 $180M Series A on 2022-09

⏰ Full Time

🟠 Senior

🧑‍💻 Full-stack Engineer

Django

ElasticSearch

JavaScript

Kafka

MySQL

Postgres

Python

TypeScript

Vue.js

🕒 August 26

Zartis

201 - 500

💼 Consulting

📦 Logistics

📣 Marketing

.NET Software Architect modernizing a core banking platform for a private-equity client. Leading C#/.NET re-architecture, Azure integrations, business-rule migration, and AI-assisted delivery at Zartis.

ASP.NET

Azure

Cloud

MS SQL Server

SQL

Vault

.NET

🕒 August 20

Holafly

501 - 1000

✈️ Travel

📡 Telecommunications

👥 B2C

Software Architect shaping resilient, cloud-native systems for Holafly, a global eSIM connectivity provider. Driving event-driven architecture, APIs, security, and platform evolution.

🇮🇪 Ireland – Remote

💰 $112.9k Seed Round - Holafly on 2019-06

⏰ Full Time

🟡 Mid-level

🟠 Senior

🧑‍💻 Full-stack Engineer

Distributed Systems

Django

Flutter

Google Cloud Platform

JavaScript

Microservices

Node.js

Python

React

Terraform