Senior HPC Cluster Engineer – AI, ML

🕒 August 6

🇮🇳 India – Remote

⏰ Full Time

🟠 Senior

🤖 AI Engineer

👻 Ghost score 14%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of NVIDIA

NVIDIA

10,000+ employees

Founded 1993

🏥 Healthcare

🏭 Manufacturing

🤖 Artificial Intelligence

Healthcare • Manufacturing • Artificial Intelligence

NVIDIA is a leading technology company specializing in accelerated computing and artificial intelligence. NVIDIA pioneers advancements in graphical processing units (GPUs), cloud computing, data centers, and virtual reality, with a focus on gaming, automotive, healthcare, and robotics industries. The company's innovations, such as NVIDIA Omniverse, transform traditional digital processes by enabling high-fidelity simulations and rendering tasks. Their applications span various industries, from autonomous vehicles using NVIDIA DRIVE to healthcare solutions with NVIDIA Clara, and AI-driven analytics and workflows.

📋 Description

• Provide leadership in systems administration and service delivery on the AI/HPC fleet by coordinating system upgrades, responding to incidents, and delivering reliability improvements • Collaborate with global teams to deliver a world-class user experience in AI and HPC research • Own day-to-day operations of production AI/HPC clusters, ensuring system health, user satisfaction, and efficient resource utilization • Develop and improve the ecosystem around GPU-accelerated computing, including scalable automation solutions • Build and maintain heterogeneous AI/ML clusters on-premises and in the cloud • Create and cultivate customer and cross-team relationships to meet evolving user needs • Support researchers running workloads, including performance analysis and optimization • Analyze and optimize cluster efficiency, job fragmentation, and GPU waste to meet internal SLA targets • Conduct root cause analysis and suggest corrective actions • Proactively find and fix issues before they occur • Lead SEV triage and postmortems for reliability incidents affecting users or infrastructure • Participate in on-call rotation and incident response for critical production GPU clusters

🎯 Requirements

• Bachelor’s degree in Computer Science, Electrical Engineering or related field or equivalent experience • Minimum 5 years of experience designing and operating large-scale compute infrastructure • Experience with AI/HPC advanced job schedulers such as Slurm, K8s, PBS, RTDA, BCM, or LSF • Proficient in administering CentOS/RHEL and/or Ubuntu Linux distributions • Solid understanding of cluster configuration management tools including BCM, Terraform, Ansible, Puppet, and Salt • Knowledge of container technologies including Docker, Singularity, Podman, Shifter, and Charliecloud • Python programming and Bash scripting • Applied experience with AI/HPC workflows using MPI • Experience analyzing and tuning performance for varied AI/HPC workloads • Passion for continual learning and staying ahead of emerging HPC and AI/ML infrastructure technologies and approaches • Background with NVIDIA GPUs, CUDA programming, NCCL, and MLPerf benchmarking • Experience with AI/ML concepts, algorithms, models, and frameworks such as PyTorch and TensorFlow • Experience with InfiniBand, IPoIB, and RDMA • Understanding of fast, distributed storage systems such as Lustre and GPFS for AI/HPC workloads

Apply Now

Similar Jobs

🕒 August 4

TechTorch

51 - 200

💼 Consulting

📦 Logistics

📣 Marketing

Senior AI engineer at TechTorch building production data foundations and full-stack AI applications for complex business workflows. Delivering client solutions and reusable accelerators remotely from India.

Airflow

AWS

Azure

Cloud

ETL

JavaScript

Next.js

Postgres

Python

SQL

🕒 August 4

MariaDB

201 - 500

🏢 Enterprise

Senior Software Engineer building AI application layers, APIs, and data services for MariaDB’s widely deployed relational database engine. Deploying scalable features across multi-cloud and on-premises environments.

AWS

Azure

Google Cloud Platform

GRPC

Java

MariaDB

Python

SQL

Go

🕒 July 31

Databricks

1001 - 5000

🤖 Artificial Intelligence

🏢 Enterprise

☁️ SaaS

AI engineer building and deploying GenAI applications for Databricks customers. Advising clients and shaping the Data + AI Platform roadmap through professional services engagements.

🇮🇳 India – Remote

💰 $1.6G Series H on 2021-08

⏰ Full Time

🟡 Mid-level

🟠 Senior

🤖 AI Engineer

Apache

AWS

Azure

Google Cloud Platform

Pandas

PyTorch

Scikit-Learn

Spark

🕒 July 30

Cotiviti

5001 - 10000

🏥 Healthcare

💼 Consulting

📦 Logistics

AI Architect building Cotiviti’s centralized agentic AI platform. Designing scalable AI agents, unified APIs, and cloud-native infrastructure for enterprise teams.

AWS

Azure

Cloud

Distributed Systems

Docker

Google Cloud Platform

Kubernetes

Microservices

Python

🕒 July 28

APEXIT

51 - 200

💼 Consulting

🏢 Enterprise

🤝 B2B

Senior Consultant developing AI applications for ApexIT. Designing scalable technical solutions and leading AI engineering efforts to enhance client operations.

Azure

Cloud

JavaScript

Node.js

Oracle

Python