Senior HPC AI Cluster Engineer

🔥 0 minutes ago

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of NVIDIA

NVIDIA

10,000+ employees

Founded 1993

🏥 Healthcare

🏭 Manufacturing

🤖 Artificial Intelligence

Healthcare • Manufacturing • Artificial Intelligence

NVIDIA is a leading technology company specializing in accelerated computing and artificial intelligence. NVIDIA pioneers advancements in graphical processing units (GPUs), cloud computing, data centers, and virtual reality, with a focus on gaming, automotive, healthcare, and robotics industries. The company's innovations, such as NVIDIA Omniverse, transform traditional digital processes by enabling high-fidelity simulations and rendering tasks. Their applications span various industries, from autonomous vehicles using NVIDIA DRIVE to healthcare solutions with NVIDIA Clara, and AI-driven analytics and workflows.

📋 Description

• Design, implement and maintain large scale HPC/AI clusters with monitoring, logging and alerting • Manage Linux job/workload schedules and orchestration tools • Develop and maintain continuous integration and delivery pipelines • Develop tooling to automate deployment and management of large-scale infrastructure environments, operational monitoring and alerting, and self-service resource consumption • Deploy monitoring solutions for servers, network and storage • Troubleshoot from bare metal through the operating system, software stack, and application level • Develop, redefine, and document standard methodologies for internal teams • Support Research & Development activities and engage in POCs/POVs for future improvements • Provide insights on at-scale system design and tuning mechanisms for large-scale compute runs • Interact with HPC, OS, GPU compute, and systems specialists to architect, develop, and bring up large-scale performance platforms • Work with accelerated computing and deep learning software and hardware platforms, researchers, developers, and customers to improve workflows and develop differentiated solutions

🎯 Requirements

• A degree in Computer Science, Engineering, or a related field • 8+ years of experience • Knowledge of HPC and AI solution technologies spanning CPUs, GPUs, high-speed interconnects, and supporting software • Experience with workload scheduling and orchestration tools such as Slurm and Kubernetes • Excellent knowledge of Windows and Linux, including Red Hat/CentOS and Ubuntu • Knowledge of networking, sockets, firewalld, iptables, Wireshark, networking internals, ACLs, OS-level security protection, TCP, DHCP, DNS, and common protocols • Experience with storage solutions such as Lustre, GPFS, and Weka.io • Python programming and Bash scripting experience • Experience with automation and configuration management tools such as Jenkins, Ansible, Puppet, and Chef • Deep knowledge of networking protocols such as InfiniBand and Ethernet • Deep understanding and experience with virtual systems such as VMware, Hyper-V, KVM, or Citrix • Familiarity with cloud computing platforms such as AWS, Azure, and Google Cloud • Knowledge of CPU and/or GPU architecture • Knowledge of Kubernetes and container-related microservice technologies • Experience with GPU-focused hardware/software such as DGX and CUDA • Experience with RDMA fabrics such as InfiniBand or RoCE

Apply Now

Similar Jobs

🕒 May 29

Swisscom

10,000+ employees

💼 Consulting

📦 Logistics

📣 Marketing

Senior Data, AI & Cloud Consultant supporting clients in data platform development and implementation. Collaborating with teams and coaching juniors while working with modern technologies.

🗣️🇩🇪 German Required