Search Remote Jobs

Senior Software Engineer, DGX Cloud Production Engineering

🔥 0 minutes ago

🏄 California – Remote

infoinfo

💵 $152k - $241.5k / year

⏰ Full Time

🟠 Senior

🏭 Production Engineer

🦅 H1B Visa Sponsor

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of NVIDIA

NVIDIA

10,000+ employees

Founded 1993

🏥 Healthcare

🏭 Manufacturing

🤖 Artificial Intelligence

Healthcare • Manufacturing • Artificial Intelligence

NVIDIA is a leading technology company specializing in accelerated computing and artificial intelligence. NVIDIA pioneers advancements in graphical processing units (GPUs), cloud computing, data centers, and virtual reality, with a focus on gaming, automotive, healthcare, and robotics industries. The company's innovations, such as NVIDIA Omniverse, transform traditional digital processes by enabling high-fidelity simulations and rendering tasks. Their applications span various industries, from autonomous vehicles using NVIDIA DRIVE to healthcare solutions with NVIDIA Clara, and AI-driven analytics and workflows.

📋 Description

• Build the next generation of NVIDIA's Kubernetes platform • Collaborate on core Kubernetes platform capabilities, including cluster management, control plane services, fleet lifecycle, and day-2 operations • Design and build reliable distributed systems and APIs for provisioning, managing, upgrading, and remediating Kubernetes clusters at scale • Define technical requirements, validation criteria, production-readiness practices, and direction for declarative workflows and automation • Create cohesive platform experiences spanning management APIs, lifecycle orchestration, runtime integration, and fleet consistency • Diagnose and resolve complex platform issues across infrastructure, runtime, networking, hardware, and operations • Improve scalability, resilience, and operability for systems supporting large-scale AI deployments • Provide technical leadership across team boundaries and drive cross-functional initiatives to completion

🎯 Requirements

• BS or MS degree in Computer Science, Computer Engineering, or a related field, or equivalent experience • 5+ years of relevant software engineering experience, including building and operating large-scale production systems • Deep expertise in Kubernetes internals, APIs, controllers or operators, and cluster lifecycle management • Strong background in distributed systems design, reliability, scalability, and failure recovery • Experience building platform software, infrastructure control planes, or managed-service foundations • Strong programming skills in one or more systems or cloud-native languages, such as Go, Python, Rust, or C++ • Experience designing clear APIs and abstractions for platform consumers and engineering teams • Ability to provide technical leadership across team boundaries and drive ambiguous, cross-functional initiatives to completion • Excellent communication and collaboration skills, with significant technical contributions and recognized expertise influencing department-level architecture and high-priority company initiatives • Background in Kubernetes platforms or managed Kubernetes services • Experience in fleet management, cluster upgrades, node lifecycle, remediation, or day-2 operations • Experience with declarative infrastructure, Kubernetes controllers, GitOps, or policy-driven platform automation • Familiarity with public-cloud and bare-metal infrastructure environments • Experience supporting AI, GPU, HPC, or other large-scale accelerated computing platforms

🏖️ Benefits

• Competitive salaries • Generous benefits package • Equity • Benefits

Apply Now

Similar Jobs

🔥 40 minutes ago

Palo Alto Networks

10,000+ employees

🔒 Cybersecurity

🏢 Enterprise

Engineering Manager leading Developer Infrastructure at Chronosphere, Palo Alto Networks’ cloud observability platform. Managing engineers, reliability, roadmap execution, hiring, and developer tooling for a complex SaaS environment.

🔥 18 hours ago

EXL

10,000+ employees

🏥 Healthcare

🛡️ Insurance

📦 Logistics

Software Production Engineer building secure, scalable applications for EXL, a data analytics and digital operations company. Delivering healthcare technology from architecture through production support and claims data analysis.

🇺🇸 United States – Remote

💵 $60.1k - $98.7k / year

💰 $2M Venture Round on 2015-01

⏰ Full Time

🟡 Mid-level

🟠 Senior

🏭 Production Engineer

🕒 4 days ago

Vultr

201 - 500

🤖 Artificial Intelligence

🤝 B2B

🔧 Hardware

Infrastructure Production Engineer validating NVIDIA and AMD GPU hardware for Vultr’s global cloud infrastructure. Building Python, Linux, and Ansible automation to improve production readiness and reliability.

🇺🇸 United States – Remote

💵 $60k - $80k / year

💰 $329M Debt Financing - Vultr on 2025-06

⏰ Full Time

🟡 Mid-level

🟠 Senior

🏭 Production Engineer

🕒 July 28

Marqeta

501 - 1000

💼 Consulting

📦 Logistics

💳 Fintech

Senior Production Support Engineer at Marqeta handling technical support and customer satisfaction. Collaborating with teams to resolve issues and improve service delivery.

🕒 July 27

LauraMac

11 - 50

☁️ SaaS

💳 Fintech

🤝 B2B

Software Developer supporting LauraMac’s SaaS platform focused on mortgage efficiency and liquidity. Involves Java and AWS application support in a collaborative team environment.