
10,000+ employees
Founded 1993
🏥 Healthcare
🏭 Manufacturing
🤖 Artificial Intelligence
Healthcare • Manufacturing • Artificial Intelligence
NVIDIA is a leading technology company specializing in accelerated computing and artificial intelligence. NVIDIA pioneers advancements in graphical processing units (GPUs), cloud computing, data centers, and virtual reality, with a focus on gaming, automotive, healthcare, and robotics industries. The company's innovations, such as NVIDIA Omniverse, transform traditional digital processes by enabling high-fidelity simulations and rendering tasks. Their applications span various industries, from autonomous vehicles using NVIDIA DRIVE to healthcare solutions with NVIDIA Clara, and AI-driven analytics and workflows.
🔥 7 minutes ago
🏄 California, Massachusetts, +2 more states – Remote
💵 $224k - $431.3k / year
⏰ Full Time
🟠 Senior
🔴 Lead
⛑ DevOps & Site Reliability Engineer (SRE)
🦅 H1B Visa Sponsor
👻 Ghost score 1%
Improve your chances of getting an interview by checking your resume score before you apply.

10,000+ employees
Founded 1993
🏥 Healthcare
🏭 Manufacturing
🤖 Artificial Intelligence
Healthcare • Manufacturing • Artificial Intelligence
NVIDIA is a leading technology company specializing in accelerated computing and artificial intelligence. NVIDIA pioneers advancements in graphical processing units (GPUs), cloud computing, data centers, and virtual reality, with a focus on gaming, automotive, healthcare, and robotics industries. The company's innovations, such as NVIDIA Omniverse, transform traditional digital processes by enabling high-fidelity simulations and rendering tasks. Their applications span various industries, from autonomous vehicles using NVIDIA DRIVE to healthcare solutions with NVIDIA Clara, and AI-driven analytics and workflows.
• Lead a team and own the roadmap for operational processes and platforms, from requirements and delivery through adoption and results • Set technical direction, prioritize work, and guide execution across engineering and operational disciplines • Partner with infrastructure, product, and security teams to establish consistent practices for incident response, maintenance, on-call, issue management, and customer-serving readiness • Hire and develop engineers and technical leads, building a team with clear ownership and accountability • Align priorities across teams, communicate progress and risks, and provide technical leadership during major incidents • Improve reliability and reduce manual work through automation, AI, and lessons from operational events • Build and operate systems supporting chip development
• BS degree or equivalent experience • 10+ overall years of software engineering or related experience • 5+ years of engineering leadership managing teams or complex technical programs • Knowledge of operational processes and supporting platforms, including roadmap, delivery, adoption, and improvement • Strong technical judgment in software architecture, platform integration, and engineering tradeoffs • Clear communication with engineers, cross-functional partners, and executive stakeholders • A record of developing engineers, growing teams, and delivering results under pressure • Established readiness standards covering service ownership, support coverage, and reliability objectives • Experience building, integrating, and scaling platforms pertaining to incident management, maintenance, customer experience management, on-call, and production readiness • Experience applying AI or LLMs to improve triage, knowledge retrieval, incident analysis, or automation • Experience supporting EDA, large-scale compute, or hybrid infrastructure with complex dependencies and demanding availability requirements
• Equity • Benefits
Apply Now🔥 1 hour ago
Senior SRE maintaining reliable, secure AWS production infrastructure for federal financial agency programs. Automating deployments, monitoring systems, and resolving incidents across critical environments.
🔥 2 hours ago
Senior DevOps Engineer building secure AWS infrastructure, CI/CD pipelines, and observability tools. Supporting Bellese’s mission-driven civic healthcare technology solutions.
🇺🇸 United States – Remote
💵 $128.7k - $153.4k / year
⏰ Full Time
🟠 Senior
⛑ DevOps & Site Reliability Engineer (SRE)
🔥 4 hours ago
Director of Site Reliability Engineering leading scalable infrastructure and reliability initiatives. DuckDuckGo protects users online through private browsing, search, subscription, and AI products.
🇺🇸 United States – Remote
💵 $243.8k / year
⏰ Full Time
🔴 Lead
⛑ DevOps & Site Reliability Engineer (SRE)
🔥 4 hours ago
DevSecOps Engineer building secure cloud platforms and automation for Entarian’s U.S. Navy mission systems. Developing CI/CD, Terraform, Kubernetes, and compliance validations.
🔥 5 hours ago
Site Reliability Software Engineer building observability, automation, and resiliency systems for Upstart’s AI lending marketplace. Improving production reliability and incident response across engineering.
🇺🇸 United States – Remote
💵 $142k - $196.6k / year
⏰ Full Time
🟡 Mid-level
🟠 Senior
⛑ DevOps & Site Reliability Engineer (SRE)
🦅 H1B Visa Sponsor