Senior HPC AI Cluster Engineer

🔥 21 minutes ago

🌐 Switzerland, United Kingdom, +3 more countries – Remote

infoinfo

⏰ Full Time

🟠 Senior

🤖 Artificial Intelligence

👻 Ghost score 11%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of NVIDIA

NVIDIA

10,000+ employees

Founded 1993

🏥 Healthcare

🏭 Manufacturing

🤖 Artificial Intelligence

Healthcare • Manufacturing • Artificial Intelligence

NVIDIA is a leading technology company specializing in accelerated computing and artificial intelligence. NVIDIA pioneers advancements in graphical processing units (GPUs), cloud computing, data centers, and virtual reality, with a focus on gaming, automotive, healthcare, and robotics industries. The company's innovations, such as NVIDIA Omniverse, transform traditional digital processes by enabling high-fidelity simulations and rendering tasks. Their applications span various industries, from autonomous vehicles using NVIDIA DRIVE to healthcare solutions with NVIDIA Clara, and AI-driven analytics and workflows.

📋 Description

• Design, implement, and maintain large-scale HPC/AI clusters with monitoring, logging, and alerting • Manage Linux job/workload schedules and orchestration tools • Develop and maintain continuous integration and delivery pipelines • Develop tooling to automate deployment and management of large-scale infrastructure environments • Automate operational monitoring and alerting and enable self-service resource consumption • Deploy monitoring solutions for servers, networks, and storage • Troubleshoot from bare metal through the operating system, software stack, and application level • Develop, redefine, and document standard methodologies for internal teams • Support research and development activities and participate in proofs of concept and proofs of value • Provide at-scale system design and tuning insights for large-scale compute runs • Collaborate with scientific researchers, developers, customers, and HPC, OS, GPU compute, and systems specialists to architect, develop, and bring up large-scale performance platforms

🎯 Requirements

• A degree in Computer Science, Engineering, or a related field • 8+ years of experience • Knowledge of HPC and AI solution technologies from CPUs and GPUs to high-speed interconnects and supporting software • Experience with job scheduling workloads and orchestration tools such as Slurm and Kubernetes • Excellent knowledge of Windows and Linux, including Red Hat/CentOS and Ubuntu • Knowledge of networking, sockets, firewalld, iptables, Wireshark, ACLs, OS-level security protection, TCP, DHCP, and DNS • Experience with storage solutions such as Lustre, GPFS, and Weka.io • Python programming and Bash scripting experience • Experience with Jenkins, Ansible, Puppet, or Chef • Deep knowledge of InfiniBand and Ethernet networking protocols • Experience with virtual systems such as VMware, Hyper-V, KVM, or Citrix • Familiarity with AWS, Azure, or Google Cloud • Preferred: knowledge of CPU and/or GPU architecture • Preferred: knowledge of Kubernetes and container-related microservice technologies • Preferred: experience with GPU-focused hardware/software such as DGX and CUDA • Preferred: experience with RDMA fabrics, including InfiniBand or RoCE

🏖️ Benefits

• Equal opportunity employer • Reasonable accommodation for individuals with disabilities to participate in the application or interview process, perform essential job functions, and receive other benefits and privileges of employment

Apply Now