Senior Site Reliability Engineer

🕒 Setembro 2

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $191.000 - $226.000 / ano

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

👻 Score fantasma 12%

infoinfo

🗣️🇺🇸🇬🇧 Inglês obrigatório

Candidatar-se
Encontrar Vagas Remotas Similares

📊 Verifique sua pontuação de currículo para esta vaga

Melhore suas chances de conseguir uma entrevista verificando sua pontuação de currículo antes de se candidatar.

Logo of Garner Health

Garner Health

51 - 200 funcionários

💼 Consultoria

📦 Logística

🏥 Saúde

Consulting • Logistics • Healthcare

Acreditamos em transparência, tomada de decisão data-driven e experiências do cliente sem fricção. É por isso que estamos construindo soluções para ajudar colaboradores a encontrar médicos de alta qualidade. Nosso time é formado por gestores de saúde, profissionais clínicos, engenheiros e especialistas em benefícios, o que nos permite desenvolver soluções com uma abordagem multidisciplinar.

Descrição

• Own the end-to-end reliability, performance, and resilience of Garner’s AWS and Kubernetes cloud environments, including AI/ML workloads • Define, measure, and uphold SLOs across critical services • Serve in the on-call rotation and lead incident response • Drive root cause analysis, corrective actions, and rigorous infrastructure-change reviews • Build and maintain monitoring, alerting, and observability systems • Translate scaling requirements into automated, composable Terraform infrastructure-as-code deliverables • Implement cloud cost-efficiency and performance improvements across the stack • Reduce operational toil and technical debt using AI tools and automation • Build and maintain deployment and observability standards for engineering teams • Communicate cloud and reliability concepts to technical and non-technical stakeholders • Ensure infrastructure and operations meet security and HIPAA compliance obligations

🎯 Requisitos

• 4+ years of hands-on experience operating production cloud infrastructure at scale in an SRE, DevOps, or platform engineering role • Deep expertise with Kubernetes and Terraform in a cloud-first environment • AWS preferred • Strong production observability experience, including defining SLOs, building monitoring and alerting, leading incident response, and conducting blameless post-incident reviews • Strong software engineering fundamentals in Python or Go, applied to infrastructure automation • Experience driving cloud cost-efficiency and performance optimization across compute, storage, and networking • Fluency with AI tools such as Claude applied to engineering and operations workflows, or strong motivation to build it quickly • Must be unable to require employer sponsorship or transfer of an employment visa • Experience supporting AI/ML or data-intensive workloads in production is a plus • Experience operating in a security-conscious or regulated environment such as HIPAA or SOC 2 is a plus • Experience with Kubernetes APIs is a plus

🏖️ Benefícios

• Equity incentive plan • Flexible PTO • Medical plan options • Dental plan options • Vision plan options • 401(k) with company match • Flexible spending accounts • Teladoc Health

Candidatar-se

Vagas Similares

🕒 Setembro 2

Skimmer

11 - 50

☁️ SaaS

🤝 B2B

⚡ Produtividade

Senior DevOps Engineer building Azure infrastructure, CI/CD, and observability for Skimmer’s pool-service platform. Owning deployment reliability for a new product line.

🇺🇸 Estados Unidos – Remoto (EUA)

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Setembro 2

ICU Medical

5001 - 10000

🏥 Saúde

🍽️ Alimentos e Bebidas

🏭 Manufatura

Senior SRE optimizing AWS infrastructure for ICU Medical, a healthcare technology and IV therapy products company. Supporting HIPAA-compliant production systems, incident response, and Kubernetes modernization.

🇺🇸 Estados Unidos – Remoto (EUA)

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Setembro 2

DriveTime

1001 - 5000

🚘 Automotivo

📦 Logística

📣 Marketing

ITSM engineer building Freshservice ITOM, Asset Management, CMDB, Change, and Release practices. Supporting DriveTime’s used-car sales, financing, and servicing operations.

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Setembro 2

Koniag Government Services

1001 - 5000

🏛️ Governo

🎖️ Defesa

💼 Consultoria

Microsoft DevOps Engineer building Azure, AKS, and CI/CD platforms for federal government digital services. Automating infrastructure, observability, security, and reliable delivery.

🇺🇸 Estados Unidos – Remoto (EUA)

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Setembro 2

Intellum

51 - 200

💼 Consultoria

📣 Marketing

🏥 Saúde

Lead Site Reliability Engineer modernizing cloud, Kubernetes, and deployment infrastructure for Intellum’s corporate education technology platform. Driving reliability, observability, incident response, and Systems Engineering mentorship.

🗣️🇺🇸🇬🇧 Inglês obrigatório