Distinguished Engineer, Production Engineering, Data Center Automation

Job not on LinkedIn

🔥 0 minutes ago

🏄 California, New York, +2 more states – Remote

infoinfo

💵 $320k - $488.8k / year

⏰ Full Time

🟠 Senior

🔴 Lead

🏭 Production Engineer

🦅 H1B Visa Sponsor

infoinfo

👻 Ghost score 1%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of NVIDIA

NVIDIA

10,000+ employees

Founded 1993

🏥 Healthcare

🏭 Manufacturing

🤖 Artificial Intelligence

Healthcare • Manufacturing • Artificial Intelligence

NVIDIA is a leading technology company specializing in accelerated computing and artificial intelligence. NVIDIA pioneers advancements in graphical processing units (GPUs), cloud computing, data centers, and virtual reality, with a focus on gaming, automotive, healthcare, and robotics industries. The company's innovations, such as NVIDIA Omniverse, transform traditional digital processes by enabling high-fidelity simulations and rendering tasks. Their applications span various industries, from autonomous vehicles using NVIDIA DRIVE to healthcare solutions with NVIDIA Clara, and AI-driven analytics and workflows.

📋 Description

• Define the long-range technical strategy for operating DGX Cloud clusters consistently across on-prem, hyperscalers, and NeoCloud environments • Define the architectural vision and core operational guidelines for cluster lifecycle, runtime delivery, restoration, release readiness, and steady-state operability throughout DGX Cloud resources • Guide the roadmap and execution of critical cross-organizational investments that improve production readiness, operational safety, performance, and cross-team coordination • Make and guide high-impact technical decisions resolving how platform, hardware, provider, and service teams coordinate to operate DGX Cloud resources in production • Develop robust workflows, interfaces, and engineering collaboration across Kubernetes production service, provider and hardware readiness, on-prem and bare-metal infrastructure operations, and service-layer reliability domains • Act as a senior technical leader in the Production Engineering group • Build the architectural direction for cluster operations in DGX Cloud • Set operating standards and direct the evolution of the production model • Drive delivery of cross-organizational capabilities ensuring DGX Cloud resources remain usable, maintainable, and continuously improved at scale • Lead by influence across several teams and critical production results

🎯 Requirements

• BS, MS, or PhD in Computer Science, Electrical Engineering, or a related technical field, or equivalent experience • 18+ years of experience building and operating large-scale distributed systems, infrastructure platforms, or production environments • Confirmed company-level technical leadership at principal, distinguished, or equivalent scope in production engineering, SRE, infrastructure software, or cloud platforms • Confirmed experience in establishing operating models, architectural direction, and engineering standards across various technical domains and organizations • Consistent record leading large, cross-team technical efforts from concept through production, including aligning collaborators, navigating complexity and delivering measurable outcomes • Deep software engineering expertise, system knowledge, and production insight • Experience with Kubernetes service management • Experience with on-premises, hyperscaler, NeoCloud, and bare-metal infrastructure operations • Experience developing automation, workflows, interfaces, APIs, architectures, or operating standards for large-scale infrastructure environments

🏖️ Benefits

• Equity • Benefits

Apply Now

Similar Jobs

🔥 3 hours ago

EXL

10,000+ employees

🏥 Healthcare

🛡️ Insurance

📦 Logistics

Production Engineer developing and supporting secure healthcare applications for EXL, a data analytics and digital operations company. Designing scalable solutions, tuning SSRS reports, and managing end-to-end delivery.

🇺🇸 United States – Remote

💵 $60.1k - $98.7k / year

💰 $2M Venture Round on 2015-01

⏰ Full Time

🟡 Mid-level

🟠 Senior

🏭 Production Engineer

🕒 Yesterday

Revecore

1001 - 5000

🏥 Healthcare

☁️ SaaS

🤝 B2B

Senior Production Support Engineer stabilizing Revecore’s cloud enterprise application. Leading incident response, Azure observability, automation, and production reliability for hospital revenue recovery.

🕒 August 25

EXL

10,000+ employees

🏥 Healthcare

🛡️ Insurance

📦 Logistics

Product Lead supporting EXL’s application infrastructure, servers, databases, and production operations. Managing deployments, workload automation, monitoring, and full-stack incident resolution.

🇺🇸 United States – Remote

💵 $90k - $110k / year

💰 $2M Venture Round on 2015-01

⏰ Full Time

🟠 Senior

🏭 Production Engineer

🕒 August 25

Legion

11 - 50

Director leading AWS DevOps, SRE, and security for Legion’s AI workforce management platform. Building reliable infrastructure and globally distributed engineering teams.

🕒 August 21

Natera

1001 - 5000

🏥 Healthcare

🧬 Biotechnology

⚕️ Healthcare Insurance

Data engineering manager leading ETL, data products, and delivery systems for Natera’s precision-medicine and genomics testing business. Supporting laboratory operations and clinical teams with reliable data infrastructure.