Senior HPC Cluster Administrator – Deep Learning Frameworks

🕒 Julho 22

🌐 Polônia, Alemanha – Remoto

infoinfo

💵 zł221.250 - zł507.000 / ano

⏰ Tempo Integral

🟠 Sênior

🖥️ Administração

👻 Score fantasma 15%

infoinfo

🗣️🇺🇸🇬🇧 Inglês obrigatório

Candidatar-se
Encontrar Vagas Remotas Similares

📊 Verifique sua pontuação de currículo para esta vaga

Melhore suas chances de conseguir uma entrevista verificando sua pontuação de currículo antes de se candidatar.

Logo of NVIDIA

NVIDIA

10.000+ funcionários

Fundada em 1993

🏥 Saúde

🏭 Manufatura

🤖 Inteligência Artificial

Healthcare • Manufacturing • Artificial Intelligence

A NVIDIA é uma empresa de tecnologia líder, especializada em computação acelerada e inteligência artificial. A companhia é pioneira em avanços em unidades de processamento gráfico (GPUs), computação em nuvem, data centers e realidade virtual, com foco nos setores de games, automotivo, saúde e robótica. As inovações da empresa, como o NVIDIA Omniverse, transformam processos digitais tradicionais ao viabilizar simulações de alta fidelidade e tarefas de renderização. Suas aplicações abrangem diversos setores, desde veículos autônomos com o NVIDIA DRIVE até soluções de saúde com o NVIDIA Clara, além de análises e fluxos de trabalho impulsionados por IA.

Descrição

• Own the full lifecycle of GPU compute clusters — procurement, provisioning, configuration management, monitoring, and deprecation — across heterogeneous Linux environments (DGX, HGX, embedded systems) • Design and scale storage solutions (NFS, Lustre, WekaFS, or equivalent) with a clear roadmap for capacity and performance growth • Lead automation of infrastructure using modern IaC tools (Ansible, Terraform) and CI/CD pipelines (GitLab) • Manage and optimize job scheduling via Slurm, including fair-share policies, reservation management, and MIG/GPU partitioning strategies • Maintain and improve observability stacks (Prometheus, Grafana, DCGM) and drive proactive resolution of hardware and software incidents • Collaborate with ML engineers and software teams to tune cluster configuration for large-scale distributed training workloads • Evaluate and introduce new technologies — networking fabrics (InfiniBand, NVLink, EFA/RDMA), storage tiers, container runtimes — to improve performance and reliability • Mentor junior engineers and contribute to team-wide engineering standards

🎯 Requisitos

• BS/MS in CS, EE, CE, or equivalent hands-on experience • 5+ years of experience deploying and administering large-scale HPC or ML training clusters • Deep expertise in Linux systems administration at scale • Strong scripting and automation skills in Python and/or bash • Hands-on experience with Slurm (scheduling, accounting, cgroup configuration) • Proficiency with configuration management and IaC (Ansible required; Terraform a plus) • Experience with container technologies (Docker, Apptainer/Singularity, Kubernetes) • Solid understanding of high-speed networking (InfiniBand, RoCE, RDMA, EFA) • Experience with distributed/parallel filesystems and storage architecture • Ability to own problems end-to-end and communicate clearly with engineering and management stakeholders

Candidatar-se

Vagas Similares

🕒 Abril 8

Strada

5001 - 10000

💼 Consultoria

🏥 Saúde

📦 Logística

Benefits Administration Delivery Analyst overseeing benefits for client’s employees in EMEA region. Collaborating with Strada Colleagues and clients to ensure seamless administration processes.

🇵🇱 Polônia – Remoto

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

🖥️ Administração

🗣️🇺🇸🇬🇧 Inglês obrigatório

🗣️🇩🇪 Alemão obrigatório

🕒 Abril 1

BLIK

51 - 200

💳 Fintech

💸 Finanças

Junior IT Systems Administrator assisting with BLIK payment system operations in Warsaw. Responsible for system administration, technical support, and collaboration with stakeholders.

🇵🇱 Polônia – Remoto

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

🖥️ Administração

🗣️🇵🇱 Polonês obrigatório