Senior Site Reliability Engineer, BCM – DGX Cloud

🔥 0 minutes ago

🏄 California – Remote

info

💵 $168k - $333.5k / year

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🦅 H1B Visa Sponsor

info
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of NVIDIA

NVIDIA

10,000+ employees

Founded 1993

🏥 Healthcare

🏭 Manufacturing

🤖 Artificial Intelligence

Healthcare • Manufacturing • Artificial Intelligence

NVIDIA is a leading technology company specializing in accelerated computing and artificial intelligence. NVIDIA pioneers advancements in graphical processing units (GPUs), cloud computing, data centers, and virtual reality, with a focus on gaming, automotive, healthcare, and robotics industries. The company's innovations, such as NVIDIA Omniverse, transform traditional digital processes by enabling high-fidelity simulations and rendering tasks. Their applications span various industries, from autonomous vehicles using NVIDIA DRIVE to healthcare solutions with NVIDIA Clara, and AI-driven analytics and workflows.

📋 Description

• Contribute to deployments and daily operations of large-scale next-generation GPU platforms • Handle incidents in GPU clusters, bridging cluster operations and development • Design and implement small features in the Base Command Manager product • Validate complex cluster configurations, including Slurm and Kubernetes orchestrators, for performance, scalability, and resilience • Ensure cluster configurations meet real-world customer scenarios • Support external customers running NVIDIA solution clusters and internal research, operations, and next-generation project clusters

🎯 Requirements

• Bachelor's Degree or equivalent experience in Computer Science or related field • 8+ years of experience in site reliability engineering and/or software development roles • Fluency in Python • In-depth knowledge of Linux and networking • Experience with C++, high-performance computing, Kubernetes, and/or system administration is an asset • Previous experience administering BCM/Bright Cluster Manager/Base Command Manager clusters is a plus • Proficiency with cluster networking, including InfiniBand and Spectrum-X

🏖️ Benefits

• Equity • Benefits • Inclusive and supportive work environment

Apply Now

Similar Jobs

🔥 34 minutes ago

SAIC

10,000+ employees

☁️ SaaS

📣 Marketing

🏢 Enterprise

DevSecOps Engineer securing SAIC’s Windows, Linux, and cloud infrastructure. Managing Active Directory, automation, CI/CD pipelines, and compliance in controlled defense environments.

🇺🇸 United States – Remote

🔥 Funding within the last year

💰 $500M Post-IPO Debt - SAIC on 2025-09

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🔥 1 hour ago

Coupa Software

1001 - 5000

💼 Consulting

📦 Logistics

🏥 Healthcare

Lead Active Directory Site Reliability Engineer architecting secure identity infrastructure for Coupa’s AI-powered spend management platform. Driving automation, cloud integration, observability, and least-privilege administration across global environments.

🔥 3 hours ago

InnoData

2 - 10

🤝 B2B

💼 Consulting

🌍 Social Impact

Application Reliability Engineer supporting GCP and Google App Engine applications for Innodata, a global data engineering and AI services company. Managing incidents, deployments, microservices, and reliability improvements.

🔥 7 hours ago

General Dynamics Information Technology

10,000+ employees

💼 Consulting

🏥 Healthcare

📦 Logistics

DevSecOps Security Engineer automating AWS cloud security, compliance, and vulnerability workflows for GDIT’s U.S. government missions. Managing continuous ATO pipelines, IaC security, and audit evidence at scale.

🕒 2 days ago

TEKsystems

10,000+ employees

💼 Consulting

🎯 Recruiter

🏢 Enterprise

CloudOps Practice Architect designing AWS/GCP SRE platforms and agentic architectures for TEKsystems enterprise clients. Mentoring global teams and accelerating cloud solution delivery from concept to production.