
10,000+ employees
Founded 1993
🏥 Healthcare
🏭 Manufacturing
🤖 Artificial Intelligence
Healthcare • Manufacturing • Artificial Intelligence
NVIDIA is a leading technology company specializing in accelerated computing and artificial intelligence. NVIDIA pioneers advancements in graphical processing units (GPUs), cloud computing, data centers, and virtual reality, with a focus on gaming, automotive, healthcare, and robotics industries. The company's innovations, such as NVIDIA Omniverse, transform traditional digital processes by enabling high-fidelity simulations and rendering tasks. Their applications span various industries, from autonomous vehicles using NVIDIA DRIVE to healthcare solutions with NVIDIA Clara, and AI-driven analytics and workflows.
🔥 0 minutes ago
🏄 California – Remote
💵 $152k - $287.5k / year
⏰ Full Time
🟠 Senior
🏭 Production Engineer
🦅 H1B Visa Sponsor
👻 Ghost score 1%
Improve your chances of getting an interview by checking your resume score before you apply.

10,000+ employees
Founded 1993
🏥 Healthcare
🏭 Manufacturing
🤖 Artificial Intelligence
Healthcare • Manufacturing • Artificial Intelligence
NVIDIA is a leading technology company specializing in accelerated computing and artificial intelligence. NVIDIA pioneers advancements in graphical processing units (GPUs), cloud computing, data centers, and virtual reality, with a focus on gaming, automotive, healthcare, and robotics industries. The company's innovations, such as NVIDIA Omniverse, transform traditional digital processes by enabling high-fidelity simulations and rendering tasks. Their applications span various industries, from autonomous vehicles using NVIDIA DRIVE to healthcare solutions with NVIDIA Clara, and AI-driven analytics and workflows.
• Build automation for bare-metal provisioning, hardware validation, firmware and software upgrades, repair, and cluster lifecycle management • Develop tools that interact with BMC and Redfish interfaces to monitor hardware health, manage server state, and assist recovery workflows • Handle and advance NVIDIA NVL72 systems and BlueField-3 or later DPUs throughout cloud partner and on-premises environments • Diagnose failures across servers, DPUs, GPU systems, CPU systems, networking, Linux, and Kubernetes; turn recurring issues into automated detection and repair • Define validation and handoff criteria so new capacity enters production safely and consistently • Take part in on-call duties, incident response, root-cause analysis, and follow-up to implement permanent solutions • Work with hardware, networking, platform, data center operations, and partner teams to resolve issues across ownership boundaries
• 5+ years building software for or operating production infrastructure, including substantial hands-on bare-metal experience • Strong Go or Python skills, with a record of delivering production automation and services • Direct experience with BMC and Redfish in server provisioning, health inspection, power management, or fault diagnosis • Practical experience working directly with NVIDIA GPU hardware, such as NVL72 systems, and BlueField-3 or newer DPUs • Experience with Linux, firmware and driver management, network boot, and the server lifecycle from initial provisioning through repair • Experience managing production reliability through on-call duties, incident response, observability, and durable solutions • Ability to debug failures across hardware, host operating systems, networking, and distributed services • Clear communication and demonstrated ownership of problems that span multiple teams • BS/MS in Computer Science or equivalent experience in a practical setting
• Equity • Benefits
Apply Now🕒 September 21
Systems Engineer supporting SAIC’s VA team with cloud production operations. Troubleshooting outages, maintaining infrastructure, and ensuring secure, reliable system performance.
🇺🇸 United States – Remote
💰 $500M Post-IPO Debt - SAIC on 2025-09
⏰ Full Time
🟠 Senior
🏭 Production Engineer
🕒 September 21
Systems Engineer supporting SAIC’s VA cloud production systems remotely. Troubleshooting outages, maintaining infrastructure, and ensuring secure, reliable operations.
🇺🇸 United States – Remote
💰 $500M Post-IPO Debt - SAIC on 2025-09
⏰ Full Time
🟠 Senior
🏭 Production Engineer
🕒 September 2
Production Engineer developing and supporting secure healthcare applications for EXL, a data analytics and digital operations company. Designing scalable solutions, tuning SSRS reports, and managing end-to-end delivery.
🇺🇸 United States – Remote
💵 $60.1k - $98.7k / year
💰 $2M Venture Round on 2015-01
⏰ Full Time
🟡 Mid-level
🟠 Senior
🏭 Production Engineer
🕒 September 1
Senior Production Support Engineer stabilizing Revecore’s cloud enterprise application. Leading incident response, Azure observability, automation, and production reliability for hospital revenue recovery.
🕒 August 21
Data engineering manager leading ETL, data products, and delivery systems for Natera’s precision-medicine and genomics testing business. Supporting laboratory operations and clinical teams with reliable data infrastructure.
🇺🇸 United States – Remote
💵 $127.2k - $159k / year
⏰ Full Time
🟡 Mid-level
🟠 Senior
🏭 Production Engineer
🦅 H1B Visa Sponsor