Senior Software Engineer – DGX Cloud Production Engineering

🕒 July 7

🏄 California, Colorado, +3 more states – Remote

infoinfo

💵 $184k - $356.5k / year

⏰ Full Time

🟠 Senior

🏭 Production Engineer

🦅 H1B Visa Sponsor

infoinfo

👻 Ghost score 19%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of NVIDIA

NVIDIA

10,000+ employees

Founded 1993

🏥 Healthcare

🏭 Manufacturing

🤖 Artificial Intelligence

Healthcare • Manufacturing • Artificial Intelligence

NVIDIA is a leading technology company specializing in accelerated computing and artificial intelligence. NVIDIA pioneers advancements in graphical processing units (GPUs), cloud computing, data centers, and virtual reality, with a focus on gaming, automotive, healthcare, and robotics industries. The company's innovations, such as NVIDIA Omniverse, transform traditional digital processes by enabling high-fidelity simulations and rendering tasks. Their applications span various industries, from autonomous vehicles using NVIDIA DRIVE to healthcare solutions with NVIDIA Clara, and AI-driven analytics and workflows.

📋 Description

• Build and operate automation for large-scale GPU clusters across NVIDIA Cloud Partners (NCP) and on-prem environments • Develop tools and services for provisioning, validation, upgrades, monitoring, repair, and cluster lifecycle operations • Improve Day 0 / Day 1 / Day 2 workflows for cluster bringup, handoff, and production operations • Reduce manual production touches through APIs, GitOps, automation, and agent-assisted workflows • Participate in on-call, incident response, debugging, and durable follow-up work • Partner with platform, storage, networking, security, and workload teams to make infrastructure production-ready

🎯 Requirements

• 8+ years of experience building or operating production infrastructure • Strong programming skills in Python, Go, or similar • Experience with Linux, Kubernetes, containers, cloud infrastructure, or infrastructure automation • Ability to troubleshoot distributed systems in production • Clear communication and ability to work across teams • BS/MS in Computer Science or equivalent experience • Experience with GPU infrastructure, Kubernetes operators, GitOps, Terraform, ArgoCD, or fleet automation • Experience with SLOs, on-call, incident response, observability, and reliability practices • Exposure to BMaaS, VMaaS, managed Kubernetes, or multi-cloud infrastructure

🏖️ Benefits

• Equity • Benefits

Apply Now

Similar Jobs

🕒 June 12

ProSidian Consulting

11 - 50

📦 Logistics

🏭 Manufacturing

🛡️ Insurance

Production Engineer providing technical due diligence and engineering validation for upstream oil and gas projects. Role involves coordinating with various stakeholders and delivering independent engineering advisory services.

🕒 November 26, 2025

DoubleZero Foundation

1 - 10

💼 Consulting

📦 Logistics

₿ Crypto

SRE role at DoubleZero focused on automation-first reliability systems in Go, ensuring infrastructure's production readiness and performance.