Senior HPC Cluster Administrator – Deep Learning Frameworks

🔥 0 minutes ago

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of NVIDIA

NVIDIA

10,000+ employees

Founded 1993

🤖 Artificial Intelligence

🎮 Gaming

Artificial Intelligence • Gaming • Automotive

NVIDIA is a leading technology company specializing in accelerated computing and artificial intelligence. NVIDIA pioneers advancements in graphical processing units (GPUs), cloud computing, data centers, and virtual reality, with a focus on gaming, automotive, healthcare, and robotics industries. The company's innovations, such as NVIDIA Omniverse, transform traditional digital processes by enabling high-fidelity simulations and rendering tasks. Their applications span various industries, from autonomous vehicles using NVIDIA DRIVE to healthcare solutions with NVIDIA Clara, and AI-driven analytics and workflows.

📋 Description

• Own the full lifecycle of GPU compute clusters — procurement, provisioning, configuration management, monitoring, and deprecation — across heterogeneous Linux environments (DGX, HGX, embedded systems) • Design and scale storage solutions (NFS, Lustre, WekaFS, or equivalent) with a clear roadmap for capacity and performance growth • Lead automation of infrastructure using modern IaC tools (Ansible, Terraform) and CI/CD pipelines (GitLab) • Manage and optimize job scheduling via Slurm, including fair-share policies, reservation management, and MIG/GPU partitioning strategies • Maintain and improve observability stacks (Prometheus, Grafana, DCGM) and drive proactive resolution of hardware and software incidents • Collaborate with ML engineers and software teams to tune cluster configuration for large-scale distributed training workloads • Evaluate and introduce new technologies — networking fabrics (InfiniBand, NVLink, EFA/RDMA), storage tiers, container runtimes — to improve performance and reliability • Mentor junior engineers and contribute to team-wide engineering standards

🎯 Requirements

• BS/MS in CS, EE, CE, or equivalent hands-on experience • 5+ years of experience deploying and administering large-scale HPC or ML training clusters • Deep expertise in Linux systems administration at scale • Strong scripting and automation skills in Python and/or bash • Hands-on experience with Slurm (scheduling, accounting, cgroup configuration) • Proficiency with configuration management and IaC (Ansible required; Terraform a plus) • Experience with container technologies (Docker, Apptainer/Singularity, Kubernetes) • Solid understanding of high-speed networking (InfiniBand, RoCE, RDMA, EFA) • Experience with distributed/parallel filesystems and storage architecture • Ability to own problems end-to-end and communicate clearly with engineering and management stakeholders

Apply Now

Similar Jobs

🕒 5 days ago

Kyndryl

10,000+ employees

🏢 Enterprise

🔒 Cybersecurity

☁️ SaaS

IBMi Systems Administration role at Kyndryl managing environments, ensuring operational continuity, and supporting automation initiatives. Contributing to infrastructure evolution in a mission-critical setting.

🕒 May 5

Relativity

1001 - 5000

🤖 Artificial Intelligence

🏢 Enterprise

☁️ SaaS

Senior Salesforce Application Administrator managing daily operations of Salesforce CPQ platform. Collaborating with various internal teams to deliver scalable CPQ solutions and support internal users.

🕒 April 8

Strada

5001 - 10000

👥 HR Tech

☁️ SaaS

🤝 B2B

Benefits Administration Delivery Analyst overseeing benefits for client’s employees in EMEA region. Collaborating with Strada Colleagues and clients to ensure seamless administration processes.

🗣️🇩🇪 German Required

🕒 April 1

BLIK

51 - 200

💳 Fintech

💸 Finance

Junior IT Systems Administrator assisting with BLIK payment system operations in Warsaw. Responsible for system administration, technical support, and collaboration with stakeholders.

🗣️🇵🇱 Polish Required