Senior Site Reliability Engineer

🕒 Julho 28

🇺🇸 Estados Unidos – Remoto (EUA)

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

👻 Score fantasma 24%

infoinfo

🗣️🇺🇸🇬🇧 Inglês obrigatório

Candidatar-se
Encontrar Vagas Remotas Similares

📊 Verifique sua pontuação de currículo para esta vaga

Melhore suas chances de conseguir uma entrevista verificando sua pontuação de currículo antes de se candidatar.

Logo of CertifyOS

CertifyOS

201 - 500 funcionários

Fundada em 2021

🏥 Saúde

☁️ SaaS

🤝 B2B

💰 $40.000.000 Series B - Certify em 2025-06

Healthcare • SaaS • B2B

A CertifyOS é uma plataforma de dados de provedores focada em saúde que consolida e verifica informações de provedores para criar uma fonte única de verdade habilitada por IA. Ela substitui planilhas e soluções pontuais por infraestrutura para credenciamento, contratação, pagamentos, conformidade e gerenciamento de dados de provedores, oferecendo serviços como credenciamento, licenciamento, monitoramento e gestão de listas. A plataforma ingere listas e feeds em tempo real, valida registros contra milhares de fontes e fornece um fluxo contínuo de dados limpos acessíveis via uma única API. A CertifyOS se posiciona como uma solução SaaS B2B para planos de saúde, empresas de saúde digital e sistemas de saúde, reivindicando melhorias operacionais mensuráveis (por exemplo, redução de custos de recredenciamento, integração mais rápida) e listando grandes planos de saúde entre seus clientes.

Descrição

• Own the operational lifecycle end-to-end and influence platform architecture, reliability standards, and deployment workflows across systems that matter. • Design for reliability and ship automation; stand behind it in production. • Work across cloud-native infrastructure on systems that process millions of provider records. • Participate in incident response processes, root cause analysis, escalation workflows, and runbooks. • Build and maintain Infrastructure as Code, CI/CD pipelines, and operational tooling that reduce manual work and improve engineering productivity without sacrificing reliability. • Maintain uptime, reduce alert fatigue, and build actionable observability across GKE and Cloud Run without drowning in noise. • Scale infrastructure efficiently, improve autoscaling behavior, resource utilization, and workload efficiency across cloud-native distributed systems.

🎯 Requisitos

• 5+ years in SRE, DevOps, Platform Engineering, or Infrastructure Engineering — operating production systems at scale where your infrastructure is someone else’s dependency and failures have real downstream consequences • Track record of improving reliability end-to-end: you’ve debugged hard production problems, made them not happen again, and built the alerting to prove it • Strong Linux systems administration, incident response, and root cause analysis skills • Comfort influencing operational standards and mentoring teams on reliability practices • Deep hands-on experience with GCP — GKE, Cloud Run, and containerized workloads at scale • Experience building and maintaining Infrastructure as Code with Terraform and/or Pulumi • Fluency across deployment patterns and the judgment to know when each fits: rolling deployments, blue/green, canary — and the rollback story for each • Experience with autoscaling, resource optimization, and infrastructure efficiency for distributed systems • Experience managing infrastructure security, secrets, and access controls in regulated or security-conscious environments • Strong understanding of Golden Signals monitoring — latency, traffic, errors, saturation — and how to make them actionable rather than noisy • Experience designing SLIs, SLOs, error budgets, alerting strategies, dashboards, and escalation workflows • Hands-on experience with observability platforms: Google Cloud Monitoring, Datadog, Grafana, Prometheus, or similar • Strong sense of data platform health: lineage, freshness, and correctness matter as much to you as throughput • Experience building and maintaining CI/CD pipelines using GitHub Actions or similar • Scripting or programming fluency in Python, Bash, Go, or similar — you reduce toil through code, not process • Experience working with Git workflows and modern software delivery practices • Strong written and verbal communication — you can explain an operational risk to an engineer and a product manager in the same conversation • Experience operating systems handling sensitive data or PII in regulated or compliance-adjacent environments • Nice to Have: • - Experience operating large-scale distributed systems or microservices architectures • - Familiarity with healthcare, credentialing, or health-tech environments • - Experience leveraging AI-assisted observability or incident response tooling • - Familiarity with NodeJS, TypeScript, Java, or React application stacks.

🏖️ Benefícios

• Your well-being matters to us. • We provide 100% coverage of health, dental, and vision insurance premiums for employees. • Our US-based team benefits from unlimited PTO, with at least two weeks off each year to recharge. • In India, employees are supported with health insurance, statutory leave benefits, and additional wellness (menstrual) leave for women.

Candidatar-se

Vagas Similares

🕒 Julho 28

The Voleon Group

201 - 500

💸 Finanças

🤖 Inteligência Artificial

Senior Cluster Site Reliability Engineer at Voleon, ensuring robustness and uptime of research compute clusters. Collaborating on systemic improvements and cluster health monitoring.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $205.000 - $235.000 / ano

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Julho 28

STN Incorporated

11 - 50

🏢 Corporativo

🔒 Cibersegurança

🔧 Hardware

Site Reliability Engineer managing reliability and observability for GPU One platform. Leading incident response and ensuring SLOs for customer SLAs.

🇺🇸 Estados Unidos – Remoto (EUA)

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Julho 28

Runpod

51 - 200

🤖 Inteligência Artificial

☁️ SaaS

🤝 B2B

Site Reliability Engineer ensuring the stability and resilience of Runpod's distributed platform. Collaborating with engineering teams on reliability frameworks and preventing incidents.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $150.000 - $200.000 / ano

💰 $20.000.000 Seed em 2024-06

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Julho 28

Thumbtack

1001 - 5000

🏪 Marketplace

☁️ SaaS

Senior Software Engineer designing and maintaining scalable systems to improve reliability and efficiency at Thumbtack. Collaborating with cross-functional teams to optimize platform services.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $179.400 - $232.100 / ano

💰 $75.000.000 Debt Financing - Thumbtack em 2024-07

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Julho 28

ICF

5001 - 10000

💼 Consultoria

🏛️ Governo

🏥 Saúde

DevOps Engineer building healthcare reporting services for ICF. Implementing cloud-based solutions using AWS and fostering collaboration on CI/CD pipeline improvements.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $108.476 - $184.409 / ano

💰 $29.000.000 Grant em 2023-03

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório