Director, Site Reliability Engineering

Vaga não está no LinkedIn

🕒 4 dias atrás

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $187.000 - $243.000 / ano

⏰ Tempo Integral

🔴 Especialista

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

Candidatar-se
Encontrar Vagas Remotas Similares

📊 Verifique sua pontuação de currículo para esta vaga

Melhore suas chances de conseguir uma entrevista verificando sua pontuação de currículo antes de se candidatar.

Logo of Counterpart Health

Counterpart Health

51 - 200 funcionários

Fundada em 2024

🏥 Saúde

🤖 Inteligência Artificial

☁️ SaaS

Healthcare • Artificial Intelligence • SaaS

A Counterpart Health é uma plataforma de capacitação de médicos impulsionada por IA que fornece insights baseados em dados no ponto de atendimento para apoiar cuidados baseados em valor. Incubada na Clover Health como Clover Assistant, seus modelos proprietários de aprendizado de máquina consomem dados de mais de 100 fontes, integram-se com prontuários eletrônicos (EHRs) e apresentam ações clínicas priorizadas para melhorar o acompanhamento pós-hospitalização, o desempenho de qualidade (HEDIS) e a gestão de custos/risco para pagadores, ACOs e práticas de atenção primária. A empresa oferece integrações SaaS baseadas em web, modelos de parceria flexíveis para a transformação de cuidados baseados em valor e tecnologias de ML patenteadas para diagnóstico e gestão de medicamentos.

Descrição

• Lead and grow our SRE team of ~10 engineers, including hiring, retention, career development, and performance management across multiple time zones (US, HK, NZ). • Build strategic partnerships with product engineering pillars — shifting SRE from reactive, ticket-based support to proactive co-ownership of reliability outcomes. • Scale our multi-tenant infrastructure to support new customer onboarding and growing patient populations. • Own cloud cost management and FinOps practices, building frameworks that balance cost control with reliability and performance. • Champion developer self-service and platform engineering. Build self-service capabilities so product teams can manage routine operations without filing SRE tickets. Establish SLOs/SLIs for critical services and improve alert quality so every page is meaningful. • Ensure the SRE team is fully leveraging AI tooling in their workflows — using tools like Claude Code for IaC generation, log analysis, root cause investigation, and automating repetitive work — at the same level as the rest of engineering.

🎯 Requisitos

• You have 6+ years managing an SRE team and 10+ years of hands-on SRE or infrastructure engineering experience. • You're deeply comfortable with our core stack: Kubernetes, GCP (GKE, Cloud SQL, Pub/Sub, GCS), Terraform, Helm, ArgoCD, PostgreSQL, and Prometheus/Grafana. • You have strong programming skills in Python and/or Go, and you're comfortable writing and reviewing infrastructure tooling code — including using AI coding tools to do so. • You have experience with CI/CD pipelines (GitHub Actions) and a track record of building or improving developer tooling and automation. • You have sound build vs. buy judgment — you default to the right answer, not the easiest one, and you're comfortable building internal tooling when existing solutions don't fit. • You have experience leading teams across multiple time zones and a track record of developing engineers into strong technical contributors.

🏖️ Benefícios

• Financial Well-Being: Our commitment to attracting and retaining top talent begins with a competitive base salary and equity opportunities. Additionally, we offer a performance-based bonus program, 401k matching, and regular compensation reviews to recognize and reward exceptional contributions. • Physical Well-Being: We prioritize the health and well-being of our employees and their families by providing comprehensive medical, dental, and vision coverage. Your health matters to us, and we invest in ensuring you have access to quality healthcare. • Mental Well-Being: We understand the importance of mental health in fostering productivity and maintaining work-life balance. To support this, we offer initiatives such as No-Meeting Fridays, monthly company holidays, access to mental health resources, and a generous flexible time-off policy. Additionally, we embrace a remote-first culture that supports collaboration and flexibility, allowing our team members to thrive from any location. • Professional Development: Developing internal talent is a priority for Clover. We offer learning programs, mentorship, professional development funding, and regular performance feedback and reviews. • Additional Perks: Employee Stock Purchase Plan (ESPP) offering discounted equity opportunities • Reimbursement for office setup expenses • Monthly cell phone & internet stipend • Remote-first culture, enabling collaboration with global teams • Paid parental leave for all new parents • And much more!

Candidatar-se

Vagas Similares

🕒 4 dias atrás

Finalsite

201 - 500

📚 Educação

☁️ SaaS

🤝 B2B

Staff Site Reliability Engineer leading Finalsite's infrastructure evolution and operational excellence practices. Collaborating with engineering leadership to enhance CI/CD and multi-cloud reliability.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $180.000 - $250.000 / ano

💰 Debt financing em 2014-12

⏰ Tempo Integral

🔴 Especialista

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 4 dias atrás

Whitespace

11 - 50

🎖️ Defesa

🏛️ Governo

🤖 Inteligência Artificial

Senior DevSecOps Engineer enhancing cybersecurity compliance for federal standards and DoD authorization processes. Leading secure CI/CD implementations and DevSecOps toolchain management for government projects.

🇺🇸 Estados Unidos – Remoto (EUA)

⏰ Tempo Integral

🟠 Sênior

🔴 Especialista

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 4 dias atrás

Martian Wall

11 - 50

🎯 Recrutamento

💼 Consultoria

🤝 B2B

DevOps Architect designing and managing multi-stage CI/CD systems for US-based clients. Strong expertise in cloud and DevOps tool chains is essential.

🇺🇸 Estados Unidos – Remoto (EUA)

⏰ Tempo Integral

🟠 Sênior

🔴 Especialista

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 6 dias atrás

Global Enterprise Services, LLC (GES)

11 - 50

💼 Consultoria

📦 Logística

Reliability Engineer responsible for cloud platform performance and incident response, managing compliance. Requires strong technical expertise and 8 years of experience.

🇺🇸 Estados Unidos – Remoto (EUA)

⏰ Tempo Integral

🟠 Sênior

🔴 Especialista

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Julho 23

RealTime eClinical Solutions

51 - 200

🏥 Saúde

🧬 Biotecnologia

Principal DevOps Architect responsible for cloud platform architecture at RealTime eClinical Solutions. Influencing infrastructure as code adoption and managing multi-tenant clinical-trial services.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $158.000 - $190.000 / ano

💰 Private Equity Round em 2022-01

⏰ Tempo Integral

🔴 Especialista

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório