HPC Infrastructure Engineer – GPU Clusters

Job not on LinkedIn

🔥 1 minute ago

🇺🇸 United States – Remote

⏰ Full Time

🟡 Mid-level

🟠 Senior

👷 Infrastructure Engineer

👻 Ghost score 25%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of ElevenLabs

ElevenLabs

1 - 10 employees

🤖 Artificial Intelligence

📱 Media

💰 $19M Series A on 2023-06

Artificial Intelligence • Media

ElevenLabs is a research lab dedicated to exploring new frontiers in voice generation. Their mission focuses on making content universally accessible in any language and voice. Through innovative research and deployment of novel methods in voice AI, ElevenLabs aims to enhance the enjoyment of content for diverse audiences and viewers.

📋 Description

• Operate and improve the GPU fleet end to end, including provisioning, scheduling, monitoring, upgrades, and capacity planning • Build automation for node health checks, automated draining and remediation, and burn-in pipelines • Own the infrastructure stack beneath training code, including OS images, NVIDIA drivers, CUDA, container runtimes, NCCL, and InfiniBand/RoCE networking • Run and tune job scheduling with Slurm or similar systems • Build and maintain high-performance storage for datasets and checkpoints • Investigate and resolve performance problems involving stragglers, degraded links, thermal issues, and faulty GPUs • Evaluate rented GPU capacity by benchmarking providers, validating capacity, and enforcing SLAs • Perform hands-on hardware work, including racking, cabling, and diagnostics • Coordinate with datacenter staff and vendors • Maintain cluster security through access control, network isolation, and secrets management • Work directly with researchers to improve training throughput and researcher velocity

🎯 Requirements

• Production experience running large-scale Linux server or GPU environments • Strong knowledge of NVIDIA drivers, CUDA, NCCL, and DCGM, or deep systems experience with ability to learn hardware stacks quickly • Experience with bare-metal environments, server hardware, and high-speed networking • Proficiency in Python and/or Bash automation • Experience with infrastructure-as-code tools such as Ansible or Terraform • Ability to analyze metrics, logs, and PromQL • Experience supporting ML training workloads from the infrastructure side (nice to have) • Experience evaluating and working with GPU cloud providers (nice to have) • Experience with parallel filesystems such as WEKA or VAST, or large-scale object storage (nice to have) • Experience with BMC/IPMI/Redfish automation and PXE provisioning at scale (nice to have) • Power and cooling awareness for dense GPU deployments (nice to have) • Willingness to perform datacenter trips and hands-on hardware work

🏖️ Benefits

• Annual discretionary professional development stipend • Annual discretionary social travel stipend to meet colleagues • Annual company offsite • Monthly co-working stipend for employees not near a main hub • Remote work flexibility • Ability to work from offices in London, New York, San Francisco, or Warsaw

Apply Now

Similar Jobs

🔥 19 minutes ago

Activision Blizzard

5001 - 10000

🎮 Gaming

👥 B2C

Senior Azure Infrastructure Engineer designing and maintaining Azure, Linux, and hybrid infrastructure for Activision’s global gaming and entertainment business. Driving cloud modernization, automation, security, and resilience.

🇺🇸 United States – Remote

💵 $102.8k - $190.2k / year

💰 $975M Post-IPO Equity - Activision Blizzard on 2021-12

⏰ Full Time

🟠 Senior

👷 Infrastructure Engineer

🕒 Yesterday

Activision

5001 - 10000

🎮 Gaming

👥 B2C

📱 Media

Senior Azure Infrastructure Engineer designing and maintaining Azure and Linux infrastructure. Leading global projects and driving automation and virtualisation initiatives.

🕒 Yesterday

Roboflow

11 - 50

💼 Consulting

📣 Marketing

🤖 Artificial Intelligence

Infrastructure Engineer securing and scaling Roboflow’s AI computer vision platform infrastructure. Operating Kubernetes, cloud systems, ML pipelines, observability, and security.

🕒 Yesterday

Health Care Service Corporation

10,000+ employees

🛡️ Insurance

💼 Consulting

🏥 Healthcare

Infrastructure Engineer improving HCSC’s health insurance company Virtual Desktop experience. Designing Azure Virtual Desktop architectures, automation, deployment patterns, and operational support.

🕒 2 days ago

Tandem Diabetes Care

1001 - 5000

🏥 Healthcare

🏭 Manufacturing

🔧 Hardware

Senior infrastructure engineer managing macOS/Linux runners and mobile CI/CD for Tandem Diabetes Care’s insulin-pump software. Improving reliable iOS and Android builds.