Staff Site Reliability Engineer

🕒 Junho 16

🏄 California – Remoto

info

💵 $200.000 - $230.000 / ano

⏰ Tempo Integral

🔴 Especialista

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🦅 Patrocina Visto H1B

info

🗣️🇺🇸🇬🇧 Inglês obrigatório

Candidatar-se
Encontrar Vagas Remotas Similares

📊 Verifique sua pontuação de currículo para esta vaga

Melhore suas chances de conseguir uma entrevista verificando sua pontuação de currículo antes de se candidatar.

Logo of Domino Data Lab

Domino Data Lab

201 - 500 funcionários

Fundada em 2013

🤖 Inteligência Artificial

🏢 Corporativo

☁️ SaaS

💰 Series F em 2022-06

Artificial Intelligence • Enterprise • SaaS

A Domino Data Lab é uma empresa que capacita empresas impulsionadas por IA a construir e gerenciar IA em escala por meio de sua Plataforma de IA Empresarial. A plataforma oferece uma experiência integrada para desenvolvimento de modelos, MLOps, colaboração e governança, permitindo que empresas globais inovem em diversos setores. A Domino apoia o melhor desenvolvimento de medicamentos, agricultura produtiva e criação de produtos competitivos. Estabelecida em 2013 e apoiada por investidores notáveis como a Sequoia Capital e NVIDIA, a Domino permite que as empresas otimizem o uso da IA de forma eficaz.

Descrição

• Lead the development of Domino's internal AI-assisted reliability tooling, including systems that analyze tickets, logs, traces, and documentation to help teams resolve outages faster with less recurring toil • Improve the observability coverage and signal quality for our most critical customer-facing systems, so engineers have more to work with throughout the development and support lifecycle • Own incident response end-to-end, from detection to remediation, and leave each problem space better documented, better understood, and less likely to recur • Guide the development of customer and user-facing observability tools within our products • Define and mature SLO/SLI frameworks for priority services, turning abstract reliability goals into measurable, actionable standards • Scale cloud operations practices for Domino’s single-tenant SaaS offering, and work with engineering teams to improve the reliability and repeatability of customer deployments and upgrades • Mentor other engineers and shape how SRE is practiced at Domino, including incident response workflows, operational readiness expectations, and post-incident learning culture

🎯 Requisitos

• Deep experience in Site Reliability Engineering, platform engineering, or a software engineering role with genuine, hands-on operational ownership • Fluency with Kubernetes, Linux, cloud platforms, and observability tooling, and the ability to use them to investigate complex, real-world production problems • A strong ability to perceive and close reliability gaps in technical products, tools and processes • Strong software engineering skills in Python or Go, with a track record of building internal tools or services that people actually rely on • Comfort leading technically ambiguous work and influencing direction across teams without needing direct authority to get things done • A history of improving reliability through engineering and automation, not just putting out fires manually • Strong communication skills and real experience mentoring engineers or shaping technical decision-making on your team • Sound judgment about AI/LLM tooling: you know where it genuinely helps in operational workflows and where it adds noise instead of signal • Bonus: Experience with LLM-based systems, retrieval workflows, SaaS platform operations, or building tooling for support or developer teams

🏖️ Benefícios

• equity • company bonus or sales commissions/bonuses • 401(k) plan • medical, dental, and vision benefits • wellness stipends

Candidatar-se

Vagas Similares

🕒 Junho 15

TrueML

51 - 200

💼 Consultoria

🏥 Saúde

📣 Marketing

Sr. Security Engineer leading integration of security across the software development lifecycle at TrueML. Engaging in security automation, cloud security, and innovative AI solutions.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $122.090 - $160.000 / mês

⏰ Tempo Integral

🟠 Sênior

🔴 Especialista

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Junho 15

Stord

1001 - 5000

🍽️ Alimentos e Bebidas

💼 Consultoria

📣 Marketing

Staff Site Reliability Engineer focusing on security to enhance GCP and CI/CD processes. Join Stord in advancing tech solutions for better consumer experiences.

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Junho 12

Mercury Insurance

5001 - 10000

🚘 Automotivo

💼 Consultoria

📦 Logística

SRO Manager at Mercury Insurance leading observability and incident response operations. Focused on proactive detection and collaboration with engineering to improve system resilience.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $118.664 - $230.619 / ano

⏰ Tempo Integral

🟠 Sênior

🔴 Especialista

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🦅 Patrocina Visto H1B

info

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Junho 12

Malwarebytes

501 - 1000

🔒 Cibersegurança

🤝 B2B

👥 B2C

Principal DevOps Engineer at Malwarebytes responsible for AWS infrastructure management and CI/CD automation. Leading security and reliability initiatives while mentoring engineering team members.

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Junho 11

SoluStaff

51 - 200

🏥 Saúde

🎖️ Defesa

🏭 Manufatura

Principal Site Reliability Engineer ensuring reliability, scalability, and performance of a healthcare SaaS platform for U.S. providers. Investigating issues and leading incident response efforts in cloud environments.

🇺🇸 Estados Unidos – Remoto (EUA)

⏰ Tempo Integral

🔴 Especialista

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório