Site Reliability Engineering Team Lead – Principal SRE

🕒 Setembro 17

🌐 Estados Unidos, Canadá – Remoto

infoinfo

💵 $132.000 - $211.400 / ano

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

👻 Score fantasma 11%

infoinfo

🗣️🇺🇸🇬🇧 Inglês obrigatório

Candidatar-se
Encontrar Vagas Remotas Similares

📊 Verifique sua pontuação de currículo para esta vaga

Melhore suas chances de conseguir uma entrevista verificando sua pontuação de currículo antes de se candidatar.

Logo of Cerence Inc.

Cerence Inc.

1001 - 5000 funcionários

Fundada em 2019

💼 Consultoria

📦 Logística

🏭 Manufatura

💰 Grant em 2020-12

Consulting • Logistics • Manufacturing

A Cerence Inc. é uma empresa global focada em fornecer soluções impulsionadas por IA, especialmente na indústria automotiva. Eles se especializam em tecnologias de IA conversacional e generativa que criam interações inteligentes, naturais e personalizadas entre humanos e veículos. Com inovações como seus modelos de linguagem de grande escala automotivos proprietários, a Cerence aprimora as experiências dos usuários em várias formas de transporte, incluindo carros, duas rodas e caminhões. A empresa já possui mais de 500 milhões de veículos entregues com sua tecnologia de IA, atendendo a mais de 80 OEMs e clientes Tier 1 em todo o mundo. A Cerence é dedicada a avanços contínuos em IA, com o objetivo de revolucionar as experiências do usuário no carro por meio da entrega rápida e integração perfeita de suas soluções.

Descrição

• Lead Cerence's Site Reliability Engineering team and own the reliability function for its cloud-native automotive AI platform • Help select, mentor, and technically develop the team across multiple locations • Set technical direction, priorities, and the reliability roadmap across a 2–3 quarter horizon • Design and maintain a sustainable on-call rotation and monitor page load and team health • Define and govern SLI/SLO/SLA frameworks for contracted availability targets up to 99.95% • Serve as Tier 2 technical escalation point for major incidents in partnership with the Global Operations Center • Champion blameless postmortem culture and ensure actionable outcomes • Lead and improve Production Readiness / NFR reviews with development teams • Contribute to root cause analysis and own systemic improvements • Approve high-risk and out-of-window production changes • Set strategic direction for metrics, dashboards, alerting, and automation • Drive CI/CD automation for service deployments, rollbacks, and operational tasks • Partner with DevOps and platform teams to evolve shared infrastructure • Embed reliability into the SDLC through collaboration with development managers and architects • Participate in reliability consulting and architectural reviews • Communicate reliability posture and risk to technical and non-technical stakeholders

🎯 Requisitos

• 8+ years of hands-on experience in site reliability, DevOps, or cloud platform roles, including time leading a team or owning a function • Track record of setting technical direction and holding standards across a team, with or without formal authority • Hands-on experience with Kubernetes, Docker, and Istio • Experience with public cloud platforms, primarily Azure, plus AWS and Google Cloud • Familiarity with observability tooling such as Zabbix, Prometheus, and Grafana • Experience with CI/CD pipelines and infrastructure-as-code practices such as Terraform and Flux • Proficiency in at least one scripting or programming language, such as Python, Go, or Shell • Strong UNIX/Linux background, including system configuration, performance debugging, and network fundamentals (Layer 4/5, DNS, HTTP/S, TLS) • Excellent written and verbal communication skills in English • Previous site reliability leadership experience preferred • Experience leading distributed or multi-site technical teams preferred • Background in high-availability service design preferred • Experience with log aggregation and analytics platforms such as Loki and Thanos preferred • Familiarity with ITSM and project tooling such as Jira and Confluence preferred • Experience in automotive, embedded, or latency-sensitive production environments preferred

🏖️ Benefícios

• Annual bonus opportunity • Insurance coverage (medical, dental, vision, life, and disability) • Paid time off • Paid holidays • Company contribution to the 401(k) plan / RRSP (Canada) • Equity awards for certain positions and levels • Remote and/or hybrid work available depending on the position

Candidatar-se

Vagas Similares

🕒 Setembro 17

Coalfire

1001 - 5000

💼 Consultoria

🏥 Saúde

📦 Logística

Senior Site Reliability Engineer operating FedRAMP-compliant cloud environments for Coalfire, a cybersecurity consulting firm. Owning observability, automation, incident response, recovery, and compliance operations.

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Setembro 17

OnePay

501 - 1000

💳 Fintech

🏦 Bancário

₿ Cripto

SRE Lead building reliable infrastructure for OnePay’s consumer fintech platform. Leading senior engineers while coding, automating operations, and improving incident response.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $250.000 - $280.000 / ano

💰 $300.000.000 Series unknown em 2025-01

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🦅 Patrocina Visto H1B

infoinfo

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Setembro 17

Ad Hoc LLC

501 - 1000

💼 Consultoria

🏥 Saúde

📦 Logística

Senior DevOps Engineer building AWS infrastructure and CI/CD pipelines for Ad Hoc’s Veterans Affairs digital services. Improving security, reliability, developer experience, and software delivery speed.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $130.000 - $140.000 / ano

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Setembro 17

Octus

501 - 1000

💼 Consultoria

⚖️ Jurídico

📚 Educação

Lead DevOps Engineer leading cloud infrastructure, CI/CD, and security for Octus, a global credit intelligence and analytics provider. Mentoring DevOps engineers and ensuring reliable, scalable systems.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $175.000 - $225.000 / ano

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Setembro 17

PhoenixTeam

51 - 200

💳 Fintech

🏠 Imobiliário

🤖 Inteligência Artificial

DevOps Manager modernizing Jenkins-based CI/CD and Fortify quality controls for PhoenixTeam's federal FHA mortgage program. Coordinating releases, documentation, and delivery across development teams.

🇺🇸 Estados Unidos – Remoto (EUA)

⏰ Tempo Integral

🟠 Sênior

🔴 Especialista

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório