Senior Software Engineer, DGX Cloud Production Engineering

🔥 17 hours ago

🏄 California – Remote

infoinfo

💵 $184k - $356.5k / year

⏰ Full Time

🟠 Senior

🏭 Production Engineer

🦅 H1B Visa Sponsor

infoinfo

👻 Ghost score 1%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of NVIDIA

NVIDIA

10,000+ employees

Founded 1993

🏥 Healthcare

🏭 Manufacturing

🤖 Artificial Intelligence

Healthcare • Manufacturing • Artificial Intelligence

NVIDIA is a leading technology company specializing in accelerated computing and artificial intelligence. NVIDIA pioneers advancements in graphical processing units (GPUs), cloud computing, data centers, and virtual reality, with a focus on gaming, automotive, healthcare, and robotics industries. The company's innovations, such as NVIDIA Omniverse, transform traditional digital processes by enabling high-fidelity simulations and rendering tasks. Their applications span various industries, from autonomous vehicles using NVIDIA DRIVE to healthcare solutions with NVIDIA Clara, and AI-driven analytics and workflows.

📋 Description

• Build and operate automation for large-scale GPU clusters across NVIDIA Cloud Partners and on-prem environments • Develop tools and services for provisioning, validation, upgrades, monitoring, repair, and cluster lifecycle operations • Improve Day 0 / Day 1 / Day 2 workflows for cluster bringup, handoff, and production operations • Reduce manual production touches through APIs, GitOps, automation, and agent-assisted workflows • Participate in on-call, incident response, debugging, and durable follow-up work • Partner with platform, storage, networking, security, and workload teams to make infrastructure production-ready

🎯 Requirements

• 8+ years of experience building or operating production infrastructure • Strong programming skills in Python, Go, or similar • Experience with Linux, Kubernetes, containers, cloud infrastructure, or infrastructure automation • Ability to troubleshoot distributed systems in production • Clear communication and ability to work across teams • BS/MS in Computer Science or equivalent experience • Experience with GPU infrastructure, Kubernetes operators, GitOps, Terraform, ArgoCD, or fleet automation • Experience with SLOs, on-call, incident response, observability, and reliability practices • Exposure to BMaaS, VMaaS, managed Kubernetes, or multi-cloud infrastructure

🏖️ Benefits

• Equity • Benefits

Apply Now

Similar Jobs

🕒 September 2

EXL

10,000+ employees

🏥 Healthcare

🛡️ Insurance

📦 Logistics

Production Engineer developing and supporting secure healthcare applications for EXL, a data analytics and digital operations company. Designing scalable solutions, tuning SSRS reports, and managing end-to-end delivery.

🇺🇸 United States – Remote

💵 $60.1k - $98.7k / year

💰 $2M Venture Round on 2015-01

⏰ Full Time

🟡 Mid-level

🟠 Senior

🏭 Production Engineer

🕒 September 1

Revecore

1001 - 5000

🏥 Healthcare

☁️ SaaS

🤝 B2B

Senior Production Support Engineer stabilizing Revecore’s cloud enterprise application. Leading incident response, Azure observability, automation, and production reliability for hospital revenue recovery.

🕒 August 21

Natera

1001 - 5000

🏥 Healthcare

🧬 Biotechnology

⚕️ Healthcare Insurance

Data engineering manager leading ETL, data products, and delivery systems for Natera’s precision-medicine and genomics testing business. Supporting laboratory operations and clinical teams with reliable data infrastructure.

🕒 August 21

Natera

1001 - 5000

🏥 Healthcare

🧬 Biotechnology

⚕️ Healthcare Insurance

Senior DevOps data science manager leading production engineering for Natera’s precision genomics testing systems. Building reliable cloud, data, and automation platforms for laboratory operations and patient diagnostics.

🕒 August 18

Palo Alto Networks

10,000+ employees

🔒 Cybersecurity

🏢 Enterprise

Engineering Manager leading Developer Infrastructure at Chronosphere, Palo Alto Networks’ cloud observability platform. Managing engineers, reliability, roadmap execution, hiring, and developer tooling for a complex SaaS environment.