Senior HPC Cluster Administrator – Deep Learning Frameworks

🕒 July 22

🌐 Poland, Germany – Remote

infoinfo

💵 zł221.3k - zł507k / year

⏰ Full Time

🟠 Senior

🖥️ Administration

👻 Ghost score 4%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of NVIDIA

NVIDIA

10,000+ employees

Founded 1993

🏥 Healthcare

🏭 Manufacturing

🤖 Artificial Intelligence

Healthcare • Manufacturing • Artificial Intelligence

NVIDIA is a leading technology company specializing in accelerated computing and artificial intelligence. NVIDIA pioneers advancements in graphical processing units (GPUs), cloud computing, data centers, and virtual reality, with a focus on gaming, automotive, healthcare, and robotics industries. The company's innovations, such as NVIDIA Omniverse, transform traditional digital processes by enabling high-fidelity simulations and rendering tasks. Their applications span various industries, from autonomous vehicles using NVIDIA DRIVE to healthcare solutions with NVIDIA Clara, and AI-driven analytics and workflows.

📋 Description

• Own the full lifecycle of GPU compute clusters — procurement, provisioning, configuration management, monitoring, and deprecation — across heterogeneous Linux environments (DGX, HGX, embedded systems) • Design and scale storage solutions (NFS, Lustre, WekaFS, or equivalent) with a clear roadmap for capacity and performance growth • Lead automation of infrastructure using modern IaC tools (Ansible, Terraform) and CI/CD pipelines (GitLab) • Manage and optimize job scheduling via Slurm, including fair-share policies, reservation management, and MIG/GPU partitioning strategies • Maintain and improve observability stacks (Prometheus, Grafana, DCGM) and drive proactive resolution of hardware and software incidents • Collaborate with ML engineers and software teams to tune cluster configuration for large-scale distributed training workloads • Evaluate and introduce new technologies — networking fabrics (InfiniBand, NVLink, EFA/RDMA), storage tiers, container runtimes — to improve performance and reliability • Mentor junior engineers and contribute to team-wide engineering standards

🎯 Requirements

• BS/MS in CS, EE, CE, or equivalent hands-on experience • 5+ years of experience deploying and administering large-scale HPC or ML training clusters • Deep expertise in Linux systems administration at scale • Strong scripting and automation skills in Python and/or bash • Hands-on experience with Slurm (scheduling, accounting, cgroup configuration) • Proficiency with configuration management and IaC (Ansible required; Terraform a plus) • Experience with container technologies (Docker, Apptainer/Singularity, Kubernetes) • Solid understanding of high-speed networking (InfiniBand, RoCE, RDMA, EFA) • Experience with distributed/parallel filesystems and storage architecture • Ability to own problems end-to-end and communicate clearly with engineering and management stakeholders

Apply Now

Similar Jobs

🕒 April 8

Strada

5001 - 10000

💼 Consulting

🏥 Healthcare

📦 Logistics

Benefits Administration Delivery Analyst overseeing benefits for client’s employees in EMEA region. Collaborating with Strada Colleagues and clients to ensure seamless administration processes.

🗣️🇩🇪 German Required

🕒 April 1

BLIK

51 - 200

💳 Fintech

💸 Finance

Junior IT Systems Administrator assisting with BLIK payment system operations in Warsaw. Responsible for system administration, technical support, and collaboration with stakeholders.

🗣️🇵🇱 Polish Required