Engineering Manager, Reliability Engineering – EDA Infrastructure

Job not on LinkedIn

🔥 7 minutes ago

🏄 California, Massachusetts, +2 more states – Remote

infoinfo

💵 $224k - $431.3k / year

⏰ Full Time

🟠 Senior

🔴 Lead

⛑ DevOps & Site Reliability Engineer (SRE)

🦅 H1B Visa Sponsor

infoinfo

👻 Ghost score 1%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of NVIDIA

NVIDIA

10,000+ employees

Founded 1993

🏥 Healthcare

🏭 Manufacturing

🤖 Artificial Intelligence

Healthcare • Manufacturing • Artificial Intelligence

NVIDIA is a leading technology company specializing in accelerated computing and artificial intelligence. NVIDIA pioneers advancements in graphical processing units (GPUs), cloud computing, data centers, and virtual reality, with a focus on gaming, automotive, healthcare, and robotics industries. The company's innovations, such as NVIDIA Omniverse, transform traditional digital processes by enabling high-fidelity simulations and rendering tasks. Their applications span various industries, from autonomous vehicles using NVIDIA DRIVE to healthcare solutions with NVIDIA Clara, and AI-driven analytics and workflows.

📋 Description

• Lead a team and own the roadmap for operational processes and platforms, from requirements and delivery through adoption and results • Set technical direction, prioritize work, and guide execution across engineering and operational disciplines • Partner with infrastructure, product, and security teams to establish consistent practices for incident response, maintenance, on-call, issue management, and customer-serving readiness • Hire and develop engineers and technical leads, building a team with clear ownership and accountability • Align priorities across teams, communicate progress and risks, and provide technical leadership during major incidents • Improve reliability and reduce manual work through automation, AI, and lessons from operational events • Build and operate systems supporting chip development

🎯 Requirements

• BS degree or equivalent experience • 10+ overall years of software engineering or related experience • 5+ years of engineering leadership managing teams or complex technical programs • Knowledge of operational processes and supporting platforms, including roadmap, delivery, adoption, and improvement • Strong technical judgment in software architecture, platform integration, and engineering tradeoffs • Clear communication with engineers, cross-functional partners, and executive stakeholders • A record of developing engineers, growing teams, and delivering results under pressure • Established readiness standards covering service ownership, support coverage, and reliability objectives • Experience building, integrating, and scaling platforms pertaining to incident management, maintenance, customer experience management, on-call, and production readiness • Experience applying AI or LLMs to improve triage, knowledge retrieval, incident analysis, or automation • Experience supporting EDA, large-scale compute, or hybrid infrastructure with complex dependencies and demanding availability requirements

🏖️ Benefits

• Equity • Benefits

Apply Now

Similar Jobs

🔥 1 hour ago

Vital Tech Solutions

51 - 200

💼 Consulting

🏥 Healthcare

📦 Logistics

Senior SRE maintaining reliable, secure AWS production infrastructure for federal financial agency programs. Automating deployments, monitoring systems, and resolving incidents across critical environments.

🔥 2 hours ago

Bellese Technologies

51 - 200

🏥 Healthcare

💼 Consulting

⚕️ Healthcare Insurance

Senior DevOps Engineer building secure AWS infrastructure, CI/CD pipelines, and observability tools. Supporting Bellese’s mission-driven civic healthcare technology solutions.

🔥 4 hours ago

DuckDuckGo

51 - 200

💼 Consulting

📣 Marketing

🔒 Cybersecurity

Director of Site Reliability Engineering leading scalable infrastructure and reliability initiatives. DuckDuckGo protects users online through private browsing, search, subscription, and AI products.

🔥 4 hours ago

Entarian

1001 - 5000

🚀 Aerospace

🎖️ Defense

🏛️ Government

DevSecOps Engineer building secure cloud platforms and automation for Entarian’s U.S. Navy mission systems. Developing CI/CD, Terraform, Kubernetes, and compliance validations.

🔥 5 hours ago

Upstart

1001 - 5000

🚘 Automotive

💼 Consulting

🏥 Healthcare

Site Reliability Software Engineer building observability, automation, and resiliency systems for Upstart’s AI lending marketplace. Improving production reliability and incident response across engineering.