Engineering Manager, Reliability Engineering – EDA Infrastructure

🕒 2 days ago

🏄 California, Massachusetts, +2 more states – Remote

infoinfo

💵 $224k - $431.3k / year

⏰ Full Time

🟠 Senior

🔴 Lead

⛑ DevOps & Site Reliability Engineer (SRE)

🦅 H1B Visa Sponsor

infoinfo

👻 Ghost score 1%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of NVIDIA

NVIDIA

10,000+ employees

Founded 1993

🏥 Healthcare

🏭 Manufacturing

🤖 Artificial Intelligence

Healthcare • Manufacturing • Artificial Intelligence

NVIDIA is a leading technology company specializing in accelerated computing and artificial intelligence. NVIDIA pioneers advancements in graphical processing units (GPUs), cloud computing, data centers, and virtual reality, with a focus on gaming, automotive, healthcare, and robotics industries. The company's innovations, such as NVIDIA Omniverse, transform traditional digital processes by enabling high-fidelity simulations and rendering tasks. Their applications span various industries, from autonomous vehicles using NVIDIA DRIVE to healthcare solutions with NVIDIA Clara, and AI-driven analytics and workflows.

📋 Description

• Lead a team and own the roadmap for operational processes and platforms, from requirements and delivery through adoption and results • Set technical direction, prioritize work, and guide execution across engineering and operational disciplines • Partner with infrastructure, product, and security teams to establish consistent practices for incident response, maintenance, on-call, issue management, and customer-serving readiness • Hire and develop engineers and technical leads, building a team with clear ownership and accountability • Align priorities across teams, communicate progress and risks, and provide technical leadership during major incidents • Improve reliability and reduce manual work through automation, AI, and lessons from operational events • Build and operate systems supporting chip development

🎯 Requirements

• BS degree or equivalent experience • 10+ overall years of software engineering or related experience • 5+ years of engineering leadership managing teams or complex technical programs • Knowledge of operational processes and supporting platforms, including roadmap, delivery, adoption, and improvement • Strong technical judgment in software architecture, platform integration, and engineering tradeoffs • Clear communication with engineers, cross-functional partners, and executive stakeholders • A record of developing engineers, growing teams, and delivering results under pressure • Established readiness standards covering service ownership, support coverage, and reliability objectives • Experience building, integrating, and scaling platforms pertaining to incident management, maintenance, customer experience management, on-call, and production readiness • Experience applying AI or LLMs to improve triage, knowledge retrieval, incident analysis, or automation • Experience supporting EDA, large-scale compute, or hybrid infrastructure with complex dependencies and demanding availability requirements

🏖️ Benefits

• Equity • Benefits

Apply Now

Similar Jobs

🕒 2 days ago

Vital Tech Solutions

51 - 200

💼 Consulting

🏥 Healthcare

📦 Logistics

Senior SRE maintaining reliable, secure AWS production infrastructure for federal financial agency programs. Automating deployments, monitoring systems, and resolving incidents across critical environments.

🇺🇸 United States – Remote

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🕒 2 days ago

Bellese Technologies

51 - 200

🏥 Healthcare

💼 Consulting

⚕️ Healthcare Insurance

Senior DevOps Engineer building secure AWS infrastructure, CI/CD pipelines, and observability tools. Supporting Bellese’s mission-driven civic healthcare technology solutions.

🇺🇸 United States – Remote

💵 $128.7k - $153.4k / year

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🕒 2 days ago

Entarian

1001 - 5000

🚀 Aerospace

🎖️ Defense

🏛️ Government

DevSecOps Engineer building secure cloud platforms and automation for Entarian’s U.S. Navy mission systems. Developing CI/CD, Terraform, Kubernetes, and compliance validations.

🇺🇸 United States – Remote

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🕒 2 days ago

Upstart

1001 - 5000

🚘 Automotive

💼 Consulting

🏥 Healthcare

Site Reliability Software Engineer building observability, automation, and resiliency systems for Upstart’s AI lending marketplace. Improving production reliability and incident response across engineering.

🕒 2 days ago

Experian

10,000+ employees

💼 Consulting

📣 Marketing

📦 Logistics

Especialista de SRE garantindo confiabilidade, observabilidade e resiliência dos produtos Identity & Fraud da Experian. Automatizando operações críticas em ambientes distribuídos de grande escala.

🗣️🇧🇷🇵🇹 Portuguese Required