
10,000+ employees
Founded 1993
🏥 Healthcare
🏭 Manufacturing
🤖 Artificial Intelligence
Healthcare • Manufacturing • Artificial Intelligence
NVIDIA is a leading technology company specializing in accelerated computing and artificial intelligence. NVIDIA pioneers advancements in graphical processing units (GPUs), cloud computing, data centers, and virtual reality, with a focus on gaming, automotive, healthcare, and robotics industries. The company's innovations, such as NVIDIA Omniverse, transform traditional digital processes by enabling high-fidelity simulations and rendering tasks. Their applications span various industries, from autonomous vehicles using NVIDIA DRIVE to healthcare solutions with NVIDIA Clara, and AI-driven analytics and workflows.
🔥 30 minutes ago
🤠 Texas – Remote
💵 $108k - $207k / year
⏰ Full Time
🟠 Senior
📞 Support Engineer
🦅 H1B Visa Sponsor
👻 Ghost score 1%
Improve your chances of getting an interview by checking your resume score before you apply.

10,000+ employees
Founded 1993
🏥 Healthcare
🏭 Manufacturing
🤖 Artificial Intelligence
Healthcare • Manufacturing • Artificial Intelligence
NVIDIA is a leading technology company specializing in accelerated computing and artificial intelligence. NVIDIA pioneers advancements in graphical processing units (GPUs), cloud computing, data centers, and virtual reality, with a focus on gaming, automotive, healthcare, and robotics industries. The company's innovations, such as NVIDIA Omniverse, transform traditional digital processes by enabling high-fidelity simulations and rendering tasks. Their applications span various industries, from autonomous vehicles using NVIDIA DRIVE to healthcare solutions with NVIDIA Clara, and AI-driven analytics and workflows.
• Own Slurm support cases from initial investigation through resolution for customers running production AI and HPC clusters • Diagnose complex problems involving slurmctld, slurmd, slurmdbd, job scheduling, node management, resource allocation, accounting, authentication, and high availability • Solve Slurm configuration and policy issues involving partitions, reservations, priorities, fair-share, quality of service, backfill, preemption, GRES/TRES, cgroups, and job constraints • Investigate performance, reliability, and scalability issues using logs, diagnostic data, configuration analysis, reproductions, and source-level debugging when required • Isolate problems across Slurm and dependencies including Linux, MUNGE, databases, networking, parallel storage, containers, GPUs, and cluster-management systems • Advise customers on Slurm configuration, upgrades, operational practices, system-resource management, and safe recovery from production incidents • Collaborate with engineering teams by producing technical descriptions, reproducible test cases, and defect reports • Develop guides, knowledge-base articles, diagnostic tools, and internal training to strengthen Slurm expertise across the support organization
• BS degree in Computer Science, Engineering, or a related field, or equivalent experience • 5+ years of hands-on experience administering and supporting Slurm in production HPC or AI environments, including business-critical outage incidents • Expert-level understanding of Slurm architecture, daemons, configuration, scheduling behavior, accounting, resource management, and failure modes • Capacity to identify sophisticated Slurm incidents independently and guide them to a technically sound resolution • In-depth Linux system-administration and troubleshooting experience, including systemd, cgroups, authentication, networking, and database-backed services • Experience operating Slurm across multi-user clusters with complex scheduling policies and heterogeneous compute resources • Strong analytical and research skills, including distinguishing Slurm defects from configuration, integration, infrastructure, and workload problems • Excellent written and verbal communication skills • Experience supporting large-scale Slurm environments containing thousands of nodes or GPUs • Experience diagnosing scheduler performance, job-throughput, controller-load, and database-scaling issues • Familiarity with Slurm source code, plugins, SPANK, Lua job-submit plugins, or upstream issue investigation • Experience with containers and HPC integration technologies such as Pyxis, Enroot, Apptainer, or Singularity • Previous experience integrating Slurm with NVIDIA Base Command Manager, Bright Cluster Manager, or another cluster-management platform
• Highly competitive salaries • Comprehensive benefits package • Equity
Apply Now🔥 2 hours ago
Senior Technical Support Engineer serving Vultr’s strategic customers. Resolving complex Linux, Kubernetes, GPU, storage, virtualization, and networking issues for a global cloud infrastructure company.
🇺🇸 United States – Remote
💵 $90k - $110k / year
💰 $329M Debt Financing - Vultr on 2025-06
⏰ Full Time
🟠 Senior
📞 Support Engineer
🔥 3 hours ago
Clinical Support Analyst supporting DaVita’s integrated kidney care model for CKD and ESRD patients. Analyzing clinical support needs within a nationwide healthcare organization.
🇺🇸 United States – Remote
💵 $57.8k - $76k / year
💰 Post-IPO Debt on 2021-02
⏰ Full Time
🟡 Mid-level
🟠 Senior
📞 Support Engineer
🔥 5 hours ago
L1 Technical Support Engineer supporting Abnormal's AI-powered email and SaaS security platform. Troubleshooting enterprise integrations, investigating threats, and improving customer outcomes.
🇺🇸 United States – Remote
💵 $31 - $44 / hour
⏰ Full Time
🟡 Mid-level
🟠 Senior
📞 Support Engineer
🦅 H1B Visa Sponsor
🔥 5 hours ago
Tier 3 Support Engineer investigating backend, data, and integration issues for Hello Heart’s AI heart-health platform. Supporting internal teams with AWS, SQL, MongoDB, and technical troubleshooting.
🔥 7 hours ago
Enterprise Support Engineer troubleshooting OpenRouter’s AI routing infrastructure for enterprise customers. Investigating API incidents, debugging integrations, and improving customer reliability during West Coast hours.