Site Reliability Engineer

Vaga não está no LinkedIn

🕒 Julho 28

🤠 Texas – Remoto

infoinfo

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

👻 Score fantasma 25%

infoinfo

🗣️🇺🇸🇬🇧 Inglês obrigatório

Candidatar-se
Encontrar Vagas Remotas Similares

📊 Verifique sua pontuação de currículo para esta vaga

Melhore suas chances de conseguir uma entrevista verificando sua pontuação de currículo antes de se candidatar.

Logo of STN Incorporated

STN Incorporated

11 - 50 funcionários

Fundada em 2016

🏢 Corporativo

🔒 Cibersegurança

🔧 Hardware

Enterprise • Cybersecurity • Hardware

A STN Incorporated é um provedor de infraestrutura de TI gerenciada de nível empresarial e infraestrutura de nuvem que oferece infraestrutura segura e pronta para auditoria para sistemas críticos de negócios e workloads de IA exigentes. A STN opera um modelo operacional gerenciado oferecendo clouds privadas de CPU, infraestrutura GPU One AI, rede e armazenamento seguros, e suporte humano 24/7 com SLAs de alta disponibilidade. Seus serviços incluem infraestrutura gerenciada e operações de nuvem, operações de cibersegurança e resposta a incidentes, gestão de compliance e riscos (SOC 2 Tipo II, pronta para HIPAA), backup e recuperação, e aquisição de tecnologia empresarial e gestão de ciclo de vida. A STN atende a empresas, empresas SaaS de alto crescimento, construtores de IA e desenvolvedores de modelos, empresas de robótica/IA física e indústrias reguladas, como a área de saúde.

Descrição

• Define and operate Service Level Objectives (SLOs) aligned with customer SLAs • Build and maintain the observability stack including metrics, logs, traces, and alerting • Lead incident response and chair post-incident reviews • Drive automation to reduce toil and improve mean-time-to-recover (MTTR) • Author and maintain operational runbooks alongside the NOC • Manage on-call rotation, escalation paths, and incident-management tooling • Coordinate cross-functionally with NOC, Platform Engineering, and Network Engineering • Drive chaos engineering, game days, and reliability testing programs • Produce SLA performance reports in coordination with the SLA Manager • Mentor junior engineers and contribute to engineering culture

🎯 Requisitos

• 5+ years in SRE, DevOps, or production engineering roles • Strong programming skills in Go, Python, or both • Hands-on experience operating Kubernetes-based platforms at scale • Deep familiarity with observability tooling (Prometheus, Grafana, Datadog, OpenTelemetry) • Strong incident management experience including major-incident command

Candidatar-se

Vagas Similares

🕒 Julho 28

Runpod

51 - 200

🤖 Inteligência Artificial

☁️ SaaS

🤝 B2B

Site Reliability Engineer ensuring the stability and resilience of Runpod's distributed platform. Collaborating with engineering teams on reliability frameworks and preventing incidents.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $150.000 - $200.000 / ano

💰 $20.000.000 Seed em 2024-06

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Julho 28

Thumbtack

1001 - 5000

🏪 Marketplace

☁️ SaaS

Senior Software Engineer designing and maintaining scalable systems to improve reliability and efficiency at Thumbtack. Collaborating with cross-functional teams to optimize platform services.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $179.400 - $232.100 / ano

💰 $75.000.000 Debt Financing - Thumbtack em 2024-07

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Julho 28

ICF

5001 - 10000

💼 Consultoria

🏛️ Governo

🏥 Saúde

DevOps Engineer building healthcare reporting services for ICF. Implementing cloud-based solutions using AWS and fostering collaboration on CI/CD pipeline improvements.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $108.476 - $184.409 / ano

💰 $29.000.000 Grant em 2023-03

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Julho 28

ICF

5001 - 10000

💼 Consultoria

🏛️ Governo

🏥 Saúde

Senior DevOps Engineer delivering best in class healthcare reporting services for ICF. Working collaboratively to implement cloud solutions and establish CI/CD pipelines using AWS.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $108.476 - $184.409 / ano

💰 $29.000.000 Grant em 2023-03

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Julho 28

Cisco

10.000+ funcionários

🔧 Hardware

🔐 Segurança

🏢 Corporativo

Lead SRE improving Cisco’s cloud developer platforms and infrastructure for network engineers. Driving scalable automation, reliability, troubleshooting, and operational excellence.

🗣️🇺🇸🇬🇧 Inglês obrigatório