Search Remote Jobs

Principal Software Engineer, DGX Cloud Production Engineering

🔥 0 minutes ago

🏄 California – Remote

info

đź’µ $272k - $431.3k / year

⏰ Full Time

đź”´ Lead

🏭 Production Engineer

🦅 H1B Visa Sponsor

info
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of NVIDIA

NVIDIA

10,000+ employees

Founded 1993

🏥 Healthcare

🏭 Manufacturing

🤖 Artificial Intelligence

Healthcare • Manufacturing • Artificial Intelligence

NVIDIA is a leading technology company specializing in accelerated computing and artificial intelligence. NVIDIA pioneers advancements in graphical processing units (GPUs), cloud computing, data centers, and virtual reality, with a focus on gaming, automotive, healthcare, and robotics industries. The company's innovations, such as NVIDIA Omniverse, transform traditional digital processes by enabling high-fidelity simulations and rendering tasks. Their applications span various industries, from autonomous vehicles using NVIDIA DRIVE to healthcare solutions with NVIDIA Clara, and AI-driven analytics and workflows.

đź“‹ Description

• Lead the architecture and development of core Kubernetes platform capabilities, including cluster management, control plane services, fleet lifecycle, and day-2 operations • Design and build highly reliable distributed systems and APIs for provisioning, managing, upgrading, and remediating Kubernetes clusters at scale • Define technical requirements, validation criteria, production-readiness practices, and the direction for declarative workflows and automation across the Kubernetes stack • Collaborate across engineering teams to create cohesive platform experiences spanning management APIs, lifecycle orchestration, runtime integration, and fleet consistency • Lead the diagnosis and resolution of complex platform issues spanning infrastructure, runtime, networking, hardware, and operations • Improve the scalability, resilience, and operability of systems supporting large-scale AI deployments • Influence engineering standards, architectural decisions, and long-term platform strategy • Mentor senior engineers and raise the bar for design quality, execution, and engineering rigor across the organization

🎯 Requirements

• BS or MS degree in Computer Science, Computer Engineering, or a related field, or equivalent experience • 15+ years of relevant software engineering experience, including building and operating large-scale production systems • Deep expertise in Kubernetes internals, APIs, controllers or operators, and cluster lifecycle management • Strong background in distributed systems design, reliability, scalability, and failure recovery • Proven experience building platform software, infrastructure control planes, or foundations for managed services • Strong programming skills in one or more systems or cloud-native languages, such as Go, Python, Rust, or C++ • Experience designing clear APIs and abstractions for platform consumers and engineering teams • Ability to provide technical leadership across team boundaries and drive ambiguous, cross-functional initiatives to completion • Excellent communication and collaboration skills, backed by significant technical contributions and recognized expertise influencing department-level architecture and high-priority company initiatives • Experience building Kubernetes platforms or managed Kubernetes services • Expertise in fleet management, cluster upgrades, node lifecycle, remediation, or day-2 operations • Experience with declarative infrastructure, Kubernetes controllers, GitOps, or policy-driven platform automation • Familiarity with public-cloud and bare-metal infrastructure environments • Experience supporting AI, GPU, HPC, or other large-scale accelerated computing platforms

🏖️ Benefits

• Competitive salaries • Generous benefits package • Equity • Benefits

Apply Now

Similar Jobs

đź•’ June 16

Paramount

10,000+ employees

đź’Ľ Consulting

📣 Marketing

📱 Media

Principal Production Engineer at Paramount delivering AI workflow solutions. Leading architecture and integration design across complex systems and ensuring production quality.

đź•’ June 12

ProSidian Consulting

11 - 50

📦 Logistics

🏭 Manufacturing

🛡️ Insurance

Production Engineer providing technical due diligence and engineering validation for upstream oil and gas projects. Role involves coordinating with various stakeholders and delivering independent engineering advisory services.