Senior Production Engineer – DGX Cloud

🔥 15 hours ago

🏄 California – Remote

infoinfo

💵 $184k - $356.5k / year

⏰ Full Time

🟠 Senior

🏭 Production Engineer

🦅 H1B Visa Sponsor

infoinfo

👻 Ghost score 1%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of NVIDIA

NVIDIA

10,000+ employees

Founded 1993

🏥 Healthcare

🏭 Manufacturing

🤖 Artificial Intelligence

Healthcare • Manufacturing • Artificial Intelligence

NVIDIA is a leading technology company specializing in accelerated computing and artificial intelligence. NVIDIA pioneers advancements in graphical processing units (GPUs), cloud computing, data centers, and virtual reality, with a focus on gaming, automotive, healthcare, and robotics industries. The company's innovations, such as NVIDIA Omniverse, transform traditional digital processes by enabling high-fidelity simulations and rendering tasks. Their applications span various industries, from autonomous vehicles using NVIDIA DRIVE to healthcare solutions with NVIDIA Clara, and AI-driven analytics and workflows.

📋 Description

• Build and operate production software, automation, and tooling for control plane services, model deployments, and inference and agentic workloads across DGX Cloud environments • Improve the reliability of inference and agentic platforms and services, including NVIDIA Cloud Functions, SGLang- and vLLM-based endpoints, and inference services built with NVIDIA Dynamo • Improve endpoint availability, inference routing, capacity management, and service health • Use infrastructure as code and GitOps to deploy, configure, validate, upgrade, and recover services consistently across environments • Build workflows for service enablement, model releases, handoff, deprecation, and ongoing operations • Define and instrument SLIs and SLOs for inference and control plane services and use error budgets to guide reliability improvements • Participate in on-call and incident response, troubleshoot failures, and turn recurring issues into automation and durable fixes • Collaborate with model, platform, storage, networking, security, and GPU infrastructure teams to design and operate services safely at scale

🎯 Requirements

• 8+ years of experience building or operating production services and large-scale distributed systems, including hands-on automation • Strong programming skills in Python, Go, or a comparable language • Experience developing tools for production operations • Experience with infrastructure as code, configuration management, or GitOps • Experience building automation for repeatable service deployments and changes • Strong knowledge of Linux, Kubernetes, containers, cloud infrastructure, distributed systems, and networking fundamentals • Ability to diagnose failures in production • Understanding of SRE principles, including SLIs, SLOs, error budgets, incident response, and reducing operational toil • Experience instrumenting services and using metrics, logs, and traces to understand system behavior and improve reliability • Clear technical communication and ability to work across engineering teams • BS/MS in Computer Science or equivalent experience

🏖️ Benefits

• Equity • Benefits

Apply Now

Similar Jobs

🕒 September 21

SAIC

10,000+ employees

☁️ SaaS

📣 Marketing

🏢 Enterprise

Systems Engineer supporting SAIC’s VA cloud production systems remotely. Troubleshooting outages, maintaining infrastructure, and ensuring secure, reliable operations.

🕒 September 21

SAIC

10,000+ employees

☁️ SaaS

📣 Marketing

🏢 Enterprise

Systems Engineer supporting SAIC’s VA team with cloud production operations. Troubleshooting outages, maintaining infrastructure, and ensuring secure, reliable system performance.

🕒 September 2

EXL

10,000+ employees

🏥 Healthcare

🛡️ Insurance

📦 Logistics

Production Engineer developing and supporting secure healthcare applications for EXL, a data analytics and digital operations company. Designing scalable solutions, tuning SSRS reports, and managing end-to-end delivery.

🇺🇸 United States – Remote

💵 $60.1k - $98.7k / year

💰 $2M Venture Round on 2015-01

⏰ Full Time

🟡 Mid-level

🟠 Senior

🏭 Production Engineer

🕒 September 1

Revecore

1001 - 5000

🏥 Healthcare

☁️ SaaS

🤝 B2B

Senior Production Support Engineer stabilizing Revecore’s cloud enterprise application. Leading incident response, Azure observability, automation, and production reliability for hospital revenue recovery.

🕒 August 21

Natera

1001 - 5000

🏥 Healthcare

🧬 Biotechnology

⚕️ Healthcare Insurance

Data engineering manager leading ETL, data products, and delivery systems for Natera’s precision-medicine and genomics testing business. Supporting laboratory operations and clinical teams with reliable data infrastructure.