Technical Support Engineer – Slurm

🔥 12 minutes ago

🇮🇳 India – Remote

⏰ Full Time

🟡 Mid-level

🟠 Senior

📞 Support Engineer

👻 Ghost score 11%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of NVIDIA

NVIDIA

10,000+ employees

Founded 1993

🏥 Healthcare

🏭 Manufacturing

🤖 Artificial Intelligence

Healthcare • Manufacturing • Artificial Intelligence

NVIDIA is a leading technology company specializing in accelerated computing and artificial intelligence. NVIDIA pioneers advancements in graphical processing units (GPUs), cloud computing, data centers, and virtual reality, with a focus on gaming, automotive, healthcare, and robotics industries. The company's innovations, such as NVIDIA Omniverse, transform traditional digital processes by enabling high-fidelity simulations and rendering tasks. Their applications span various industries, from autonomous vehicles using NVIDIA DRIVE to healthcare solutions with NVIDIA Clara, and AI-driven analytics and workflows.

📋 Description

• Own Slurm support cases from initial investigation through resolution for customers running production AI and HPC clusters • Diagnose complex problems involving slurmctld, slurmd, slurmdbd, job scheduling, node management, resource allocation, accounting, authentication, and high availability • Solve Slurm configuration and policy issues involving partitions, reservations, priorities, fair-share, quality of service, backfill, preemption, GRES/TRES, cgroups, and job constraints • Investigate performance, reliability, and scalability issues using logs, diagnostic data, configuration analysis, reproductions, and source-level debugging when required • Isolate problems across Slurm and dependencies including Linux, MUNGE, databases, networking, parallel storage, containers, GPUs, and cluster-management systems • Advise customers on Slurm configuration, upgrades, operational practices, resource management, and safe recovery from production incidents • Collaborate with engineering teams through clear technical descriptions, reproducible test cases, and well-supported defect reports • Develop guides, knowledge-base articles, diagnostic tools, and internal training to strengthen Slurm expertise across the support organization

🎯 Requirements

• BS degree in Computer Science, Engineering, or a related field, or equivalent practical experience • 5+ years of hands-on experience administering and supporting Slurm in production HPC or AI environments, including business-critical outage incidents • Expert-level understanding of Slurm architecture, daemons, configuration, scheduling behavior, accounting, resource management, and failure modes • Ability to independently identify sophisticated Slurm incidents and guide them to technically sound resolutions • In-depth Linux system-administration and troubleshooting experience, including systemd, cgroups, authentication, networking, and database-backed services • Experience operating Slurm across multi-user clusters with complex scheduling policies and heterogeneous compute resources • Strong analytical and research skills, including distinguishing Slurm defects from configuration, integration, infrastructure, and workload problems • Excellent written and verbal communication skills, including turning detailed technical findings into clear explanations and actionable recommendations • Experience supporting large-scale Slurm environments containing thousands of nodes or GPUs • Experience diagnosing scheduler performance, job-throughput, controller-load, and database-scaling issues • Familiarity with Slurm source code, plugins, SPANK, Lua job-submit plugins, or upstream issue investigation • Experience with containers and HPC integration technologies such as Pyxis, Enroot, Apptainer, or Singularity • Previous experience integrating Slurm with NVIDIA Base Command Manager, Bright Cluster Manager, or another cluster-management platform

🏖️ Benefits

• Diverse, supportive work environment • Opportunities to make a lasting impact on the world • Full-time employment

Apply Now

Similar Jobs

🕒 2 days ago

StarTree

11 - 50

💼 Consulting

📦 Logistics

📣 Marketing

Technical Support Engineer supporting StarTree’s Apache Pinot and cloud analytics platform. Troubleshooting data systems, assisting customers, and improving support operations from India.

Apache

AWS

Azure

Cloud

Hadoop

Kafka

Kubernetes

Linux

Spark

SQL

🕒 3 days ago

Adobe

10,000+ employees

💼 Consulting

📣 Marketing

Technical Support Engineer delivering advanced customer support for Adobe’s creative and productivity platforms. Driving CSAT, retention, ARR contribution, and cross-functional service improvements.

🕒 5 days ago

Harris Computer

10,000+ employees

🏥 Healthcare

💼 Consulting

📦 Logistics

Technical Support Engineer resolving complex escalations across enterprise healthcare applications and secure messaging solutions. Troubleshooting Windows, SQL Server, HL7, IIS, and integration environments.

ASP.NET

AWS

Azure

Citrix

DNS

Docker

Firewalls

ITSM

Linux

MS SQL Server

ServiceNow

SQL

TCP/IP

.NET

🕒 5 days ago

Harris Computer

10,000+ employees

Technical Support Engineer resolving complex escalations across enterprise healthcare applications and secure messaging solutions. Troubleshooting Windows, SQL Server, HL7, IIS, and integration issues.

ASP.NET

AWS

Azure

Citrix

DNS

Docker

Firewalls

ITSM

Linux

MS SQL Server

ServiceNow

SQL

TCP/IP

.NET

🕒 September 4

Elfonze Technologies

201 - 500

💼 Consulting

📦 Logistics

🏭 Manufacturing

RPA Support Engineer monitoring and administering UiPath Orchestrator, robots, queues, and scheduled jobs. Improving automation reliability for an IT services environment.

ITSM

RPA