Technical Lead – GPU Infrastructure

🔥 4 minutes ago

🇵🇱 Poland – Remote

⏰ Full Time

🟠 Senior

🧑‍💻 Full-stack Engineer

👻 Ghost score 25%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Tether.to

Tether.to

11 - 50 employees

Founded 2014

₿ Crypto

💳 Fintech

💸 Finance

Crypto • Fintech • Finance

Tether. to is a leading digital asset company that pioneers the use of stablecoins in the blockchain space. As the most widely adopted stablecoin, Tether tokens are designed to be pegged 1-to-1 with fiat currencies, offering a stable digital asset option for users. The platform facilitates these token transactions across multiple blockchains, enhancing cross-border transactions while maintaining transparency with daily records of total assets and reserves. Tether's initiatives include educational programs promoting digital asset usage, especially targeting regions like the Middle East, Turkey, and the Philippines. Tether thus positions itself as a disruptor in the traditional financial system by enabling a stable, efficient method of handling transactions in the digital currency world.

📋 Description

• Own the end-to-end platform architecture through architecture proposals, high-level and low-level designs, reviews, and maintaining the baseline • Lead and line-manage a distributed team of approximately twelve engineers across backend, frontend, DevOps, QA, and documentation • Establish engineering standards, conduct code and design reviews, manage release gates, hold one-to-ones, and provide growth and performance input • Design, build, and operate a managed Slurm service for research users • Manage Slurm controllers, accounting, partitions, login nodes, node onboarding and acceptance, driver and CUDA baselines and upgrades, stalled-job and node-health detection, drain and autohealing, storage visibility, identity, and isolation • Own Kubernetes cluster bootstrap and lifecycle on partner-provided bare metal • Implement NVIDIA GPU Operator and Network Operator, VM-based GPU isolation using KubeVirt and VFIO, upgrades, backup and recovery, and node replacement • Own managed inference architecture, including serving, multi-GPU and multi-node parallelism, autoscaling, request routing, endpoint reliability, and confidential-compute-capable capacity • Establish metrics, logging, alerting, and SLOs across the control plane, GPU fleet, and application tiers • Lead incident response, post-incident reviews, and creation of a sustainable on-call model • Act as primary technical interface to infrastructure partners and vendors • Translate requirements into written specifications and acceptance tests, run escalations to closure, and provide input to capacity planning and hardware sourcing • Work with research, model-training, and product teams to translate workloads into platform requirements and broker capacity • Complete the platform team and set the technical bar for new engineers • Deliver the stack within a fixed delivery window during the first six months

🎯 Requirements

• Eight or more years of hands-on engineering experience • At least three years leading teams that build and operate infrastructure platforms other teams depend on • Bachelor's or Master's degree in computer science or engineering, or equivalent practical experience • Hands-on experience running Slurm at scale, including slurmctld and slurmdbd, partitions, QoS and priority, accounting, prolog and epilog, node health scripting, and upgrades with jobs on the system • Experience operating an HPC or GPU training cluster for a research population is ideally preferred • Experience operating NVIDIA GPU fleets on bare metal, including NVIDIA driver and CUDA lifecycle, Fabric Manager, NVSwitch, DCGM, MIG, node burn-in and acceptance • Experience with InfiniBand fabric and subnet configuration, RDMA, SR-IOV, and diagnosing multi-node NCCL performance problems • Deep Linux systems knowledge, including kernel modules and drivers, PCIe passthrough, vfio-pci, cgroups, namespaces, and performance tuning • Production Kubernetes operations, including control plane, upgrades, CNI, CSI, operators, custom controllers, and multi-tenancy design • Experience with HPC storage and data movement, including VAST, Lustre, NFS, node-local NVMe caching, and distributing large model weights and datasets • Experience with Prometheus, Grafana, Loki or equivalents, SLOs, incident response, and post-incident review • Working fluency in JavaScript and Node.js sufficient to review control plane, CLI, and worker services and make architecture decisions • Experience delivering a multi-tenant IaaS, PaaS, or research computing service with resource isolation, quotas, usage metering, and user-facing API and CLI surfaces • People management across time zones, cross-track review, written architecture decisions, and ability to challenge partners or executives with reasons • Excellent written and spoken English • Availability to work between UTC and UTC+5:30 • Desirable experience with Slurm operators on Kubernetes or Kubernetes-native schedulers • Desirable experience with modern serving stacks such as vLLM, SGLang, and TensorRT-LLM • Desirable experience with VM and container isolation, confidential computing, Cluster API, kubeadm, Cilium, autohealing, infrastructure as code, GitOps, GPU cloud/HPC/AI lab platforms, distributed systems, and hardware-provider partnerships

🏖️ Benefits

• Fully remote work arrangement • Occasional travel to partner sites and team events

Apply Now

Similar Jobs

🔥 3 hours ago

Sphera

1001 - 5000

💼 Consulting

🏥 Healthcare

📦 Logistics

Senior .NET Engineer leading scalable .NET and Azure cloud application delivery. Advancing secure, sustainable ESG software for enterprise customers.

Angular

ASP.NET

Azure

Cloud

Distributed Systems

Docker

Entity Framework

Kubernetes

Microservices

React

SQL

Terraform

Vault

.NET

🕒 Yesterday

Relativity

1001 - 5000

⚖️ Legal

💼 Consulting

🏥 Healthcare

Senior Software Engineer building Azure-native platform services and developer tooling for Relativity’s legal data intelligence platform. Automating infrastructure, CI/CD, observability, and AI-enabled engineering workflows.

Azure

Cloud

Docker

Kubernetes

Terraform

.NET

🕒 Yesterday

TRAX

51 - 200

💼 Consulting

📦 Logistics

🚀 Aerospace

AI Software Engineer developing deep learning, VLM, and AR solutions for FORM’s retail intelligence platform. Deploying optimized computer vision software to edge devices across global stores.

Keras

Python

Tensorflow

🕒 4 days ago

Gramian Consulting

2 - 10

💼 Consulting

📦 Logistics

📣 Marketing

Licensing owned Git repositories for AI training. Reviewing repository eligibility, value, and licensing potential while retaining ownership under non-exclusive terms.

🕒 5 days ago

MetalBear

11 - 50

☁️ SaaS

🤝 B2B

Engineering Team Lead leading cloud integrations for MetalBear’s mirrord Kubernetes development platform. Building queue splitting, DB branching, and related infrastructure features.

Cloud

Kubernetes

Rust

C++