Senior HPC Cluster Engineer – AI, ML

🕒 Agosto 6

🇮🇳 Índia – Remoto

⏰ Tempo Integral

🟠 Sênior

🤖 Engenheiro de IA

🗣️🇺🇸🇬🇧 Inglês obrigatório

Candidatar-se
Encontrar Vagas Remotas Similares

📊 Verifique sua pontuação de currículo para esta vaga

Melhore suas chances de conseguir uma entrevista verificando sua pontuação de currículo antes de se candidatar.

Logo of NVIDIA

NVIDIA

10.000+ funcionários

Fundada em 1993

🏥 Saúde

🏭 Manufatura

🤖 Inteligência Artificial

Healthcare • Manufacturing • Artificial Intelligence

A NVIDIA é uma empresa de tecnologia líder, especializada em computação acelerada e inteligência artificial. A companhia é pioneira em avanços em unidades de processamento gráfico (GPUs), computação em nuvem, data centers e realidade virtual, com foco nos setores de games, automotivo, saúde e robótica. As inovações da empresa, como o NVIDIA Omniverse, transformam processos digitais tradicionais ao viabilizar simulações de alta fidelidade e tarefas de renderização. Suas aplicações abrangem diversos setores, desde veículos autônomos com o NVIDIA DRIVE até soluções de saúde com o NVIDIA Clara, além de análises e fluxos de trabalho impulsionados por IA.

Descrição

• Provide leadership in systems administration and service delivery on the AI/HPC fleet by coordinating system upgrades, responding to incidents, and delivering reliability improvements • Collaborate with global teams to deliver a world-class user experience in AI and HPC research • Own day-to-day operations of production AI/HPC clusters, ensuring system health, user satisfaction, and efficient resource utilization • Develop and improve the ecosystem around GPU-accelerated computing, including scalable automation solutions • Build and maintain heterogeneous AI/ML clusters on-premises and in the cloud • Create and cultivate customer and cross-team relationships to meet evolving user needs • Support researchers running workloads, including performance analysis and optimization • Analyze and optimize cluster efficiency, job fragmentation, and GPU waste to meet internal SLA targets • Conduct root cause analysis and suggest corrective actions • Proactively find and fix issues before they occur • Lead SEV triage and postmortems for reliability incidents affecting users or infrastructure • Participate in on-call rotation and incident response for critical production GPU clusters

🎯 Requisitos

• Bachelor’s degree in Computer Science, Electrical Engineering or related field or equivalent experience • Minimum 5 years of experience designing and operating large-scale compute infrastructure • Experience with AI/HPC advanced job schedulers such as Slurm, K8s, PBS, RTDA, BCM, or LSF • Proficient in administering CentOS/RHEL and/or Ubuntu Linux distributions • Solid understanding of cluster configuration management tools including BCM, Terraform, Ansible, Puppet, and Salt • Knowledge of container technologies including Docker, Singularity, Podman, Shifter, and Charliecloud • Python programming and Bash scripting • Applied experience with AI/HPC workflows using MPI • Experience analyzing and tuning performance for varied AI/HPC workloads • Passion for continual learning and staying ahead of emerging HPC and AI/ML infrastructure technologies and approaches • Background with NVIDIA GPUs, CUDA programming, NCCL, and MLPerf benchmarking • Experience with AI/ML concepts, algorithms, models, and frameworks such as PyTorch and TensorFlow • Experience with InfiniBand, IPoIB, and RDMA • Understanding of fast, distributed storage systems such as Lustre and GPFS for AI/HPC workloads

Candidatar-se

Vagas Similares

🕒 Agosto 4

TechTorch

51 - 200

💼 Consultoria

📦 Logística

📣 Marketing

Senior AI engineer at TechTorch building production data foundations and full-stack AI applications for complex business workflows. Delivering client solutions and reusable accelerators remotely from India.

🇮🇳 Índia – Remoto

⏰ Tempo Integral

🟠 Sênior

🤖 Engenheiro de IA

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Agosto 4

MariaDB

201 - 500

🏢 Corporativo

Senior Software Engineer building AI application layers, APIs, and data services for MariaDB’s widely deployed relational database engine. Deploying scalable features across multi-cloud and on-premises environments.

🇮🇳 Índia – Remoto

⏰ Tempo Integral

🟠 Sênior

🤖 Engenheiro de IA

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Agosto 4

MariaDB

201 - 500

🏢 Corporativo

Senior Developer building production AI platform capabilities on MariaDB’s relational database engine. Integrating agent frameworks, multi-cloud services, secure data models, and scalable APIs.

🇮🇳 Índia – Remoto

⏰ Tempo Integral

🟠 Sênior

🤖 Engenheiro de IA

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Agosto 4

Weekday (YC W21)

11 - 50

💼 Consultoria

👥 RH Tech

☁️ SaaS

AI platform leader driving commercial lending transformation for a financial institution. Leading strategy, architecture, engineering, and complex multi-team delivery through senior governance.

🇮🇳 Índia – Remoto

💵 ₹10.000.000 - ₹15.000.000 / ano

⏰ Tempo Integral

🟠 Sênior

🤖 Engenheiro de IA

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Julho 30

Cotiviti

5001 - 10000

🏥 Saúde

💼 Consultoria

📦 Logística

AI Engineer role involves creating a centralized AI platform for Cotiviti's AI initiatives. Collaborates with cross-functional teams to ensure seamless integration and performance optimization.

🇮🇳 Índia – Remoto

⏰ Tempo Integral

🟠 Sênior

🔴 Especialista

🤖 Engenheiro de IA

🗣️🇺🇸🇬🇧 Inglês obrigatório