Senior Site Reliability Engineer

🕒 3 dias atrás

🇺🇸 Estados Unidos – Remoto (EUA)

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

Candidatar-se
Encontrar Vagas Remotas Similares

📊 Verifique sua pontuação de currículo para esta vaga

Melhore suas chances de conseguir uma entrevista verificando sua pontuação de currículo antes de se candidatar.

Logo of CertifyOS

CertifyOS

201 - 500 funcionários

Fundada em 2021

🏥 Saúde

☁️ SaaS

🤝 B2B

💰 $40.000.000 Series B - Certify em 2025-06

Healthcare • SaaS • B2B

A CertifyOS é uma plataforma de dados de provedores focada em saúde que consolida e verifica informações de provedores para criar uma fonte única de verdade habilitada por IA. Ela substitui planilhas e soluções pontuais por infraestrutura para credenciamento, contratação, pagamentos, conformidade e gerenciamento de dados de provedores, oferecendo serviços como credenciamento, licenciamento, monitoramento e gestão de listas. A plataforma ingere listas e feeds em tempo real, valida registros contra milhares de fontes e fornece um fluxo contínuo de dados limpos acessíveis via uma única API. A CertifyOS se posiciona como uma solução SaaS B2B para planos de saúde, empresas de saúde digital e sistemas de saúde, reivindicando melhorias operacionais mensuráveis (por exemplo, redução de custos de recredenciamento, integração mais rápida) e listando grandes planos de saúde entre seus clientes.

Descrição

• Own the operational lifecycle end-to-end and influence platform architecture, reliability standards, and deployment workflows across systems that matter. • Design for reliability and ship automation; stand behind it in production. • Work across cloud-native infrastructure on systems that process millions of provider records. • Participate in incident response processes, root cause analysis, escalation workflows, and runbooks. • Build and maintain Infrastructure as Code, CI/CD pipelines, and operational tooling that reduce manual work and improve engineering productivity without sacrificing reliability. • Maintain uptime, reduce alert fatigue, and build actionable observability across GKE and Cloud Run without drowning in noise. • Scale infrastructure efficiently, improve autoscaling behavior, resource utilization, and workload efficiency across cloud-native distributed systems.

🎯 Requisitos

• 5+ years in SRE, DevOps, Platform Engineering, or Infrastructure Engineering — operating production systems at scale where your infrastructure is someone else’s dependency and failures have real downstream consequences • Track record of improving reliability end-to-end: you’ve debugged hard production problems, made them not happen again, and built the alerting to prove it • Strong Linux systems administration, incident response, and root cause analysis skills • Comfort influencing operational standards and mentoring teams on reliability practices • Deep hands-on experience with GCP — GKE, Cloud Run, and containerized workloads at scale • Experience building and maintaining Infrastructure as Code with Terraform and/or Pulumi • Fluency across deployment patterns and the judgment to know when each fits: rolling deployments, blue/green, canary — and the rollback story for each • Experience with autoscaling, resource optimization, and infrastructure efficiency for distributed systems • Experience managing infrastructure security, secrets, and access controls in regulated or security-conscious environments • Strong understanding of Golden Signals monitoring — latency, traffic, errors, saturation — and how to make them actionable rather than noisy • Experience designing SLIs, SLOs, error budgets, alerting strategies, dashboards, and escalation workflows • Hands-on experience with observability platforms: Google Cloud Monitoring, Datadog, Grafana, Prometheus, or similar • Strong sense of data platform health: lineage, freshness, and correctness matter as much to you as throughput • Experience building and maintaining CI/CD pipelines using GitHub Actions or similar • Scripting or programming fluency in Python, Bash, Go, or similar — you reduce toil through code, not process • Experience working with Git workflows and modern software delivery practices • Strong written and verbal communication — you can explain an operational risk to an engineer and a product manager in the same conversation • Experience operating systems handling sensitive data or PII in regulated or compliance-adjacent environments • Nice to Have: • - Experience operating large-scale distributed systems or microservices architectures • - Familiarity with healthcare, credentialing, or health-tech environments • - Experience leveraging AI-assisted observability or incident response tooling • - Familiarity with NodeJS, TypeScript, Java, or React application stacks.

🏖️ Benefícios

• Your well-being matters to us. • We provide 100% coverage of health, dental, and vision insurance premiums for employees. • Our US-based team benefits from unlimited PTO, with at least two weeks off each year to recharge. • In India, employees are supported with health insurance, statutory leave benefits, and additional wellness (menstrual) leave for women.

Candidatar-se

Vagas Similares

🕒 3 dias atrás

The Voleon Group

201 - 500

💸 Finanças

🤖 Inteligência Artificial

Senior Cluster Site Reliability Engineer at Voleon, ensuring robustness and uptime of research compute clusters. Collaborating on systemic improvements and cluster health monitoring.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $205.000 - $235.000 / ano

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 3 dias atrás

STN Incorporated

11 - 50

🏢 Corporativo

🔒 Cibersegurança

🔧 Hardware

Site Reliability Engineer managing reliability and observability for GPU One platform. Leading incident response and ensuring SLOs for customer SLAs.

🇺🇸 Estados Unidos – Remoto (EUA)

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 3 dias atrás

Zafran Security

51 - 200

🔐 Segurança

Senior DevOps Engineer at Zafran working on compliance certifications and implementing security controls across infrastructure. Collaborating with teams to enhance security posture and ensure regulatory compliance.

🇺🇸 Estados Unidos – Remoto (EUA)

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 3 dias atrás

Runpod

51 - 200

🤖 Inteligência Artificial

☁️ SaaS

🤝 B2B

Site Reliability Engineer ensuring the stability and resilience of Runpod's distributed platform. Collaborating with engineering teams on reliability frameworks and preventing incidents.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $150.000 - $200.000 / ano

💰 $20.000.000 Seed em 2024-06

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 3 dias atrás

Talkiatry

501 - 1000

🏥 Saúde

👥 B2C

Senior Site Reliability Engineer at Talkiatry, building SRE principles for mental health care. Collaborate with teams to minimize outages and improve reliability for patient services.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $160.000 - $185.000 / ano

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório