Site Reliability Engineer

Vaga não está no LinkedIn

🕒 3 dias atrás

🗣️🇺🇸🇬🇧 Inglês obrigatório

Candidatar-se
Encontrar Vagas Remotas Similares

📊 Verifique sua pontuação de currículo para esta vaga

Melhore suas chances de conseguir uma entrevista verificando sua pontuação de currículo antes de se candidatar.

Logo of STN Incorporated

STN Incorporated

11 - 50 funcionários

Fundada em 2016

🏢 Corporativo

🔒 Cibersegurança

🔧 Hardware

Enterprise • Cybersecurity • Hardware

A STN Incorporated é um provedor de infraestrutura de TI gerenciada de nível empresarial e infraestrutura de nuvem que oferece infraestrutura segura e pronta para auditoria para sistemas críticos de negócios e workloads de IA exigentes. A STN opera um modelo operacional gerenciado oferecendo clouds privadas de CPU, infraestrutura GPU One AI, rede e armazenamento seguros, e suporte humano 24/7 com SLAs de alta disponibilidade. Seus serviços incluem infraestrutura gerenciada e operações de nuvem, operações de cibersegurança e resposta a incidentes, gestão de compliance e riscos (SOC 2 Tipo II, pronta para HIPAA), backup e recuperação, e aquisição de tecnologia empresarial e gestão de ciclo de vida. A STN atende a empresas, empresas SaaS de alto crescimento, construtores de IA e desenvolvedores de modelos, empresas de robótica/IA física e indústrias reguladas, como a área de saúde.

Descrição

• Define and operate Service Level Objectives (SLOs) aligned with customer SLAs • Build and maintain the observability stack including metrics, logs, traces, and alerting • Lead incident response and chair post-incident reviews • Drive automation to reduce toil and improve mean-time-to-recover (MTTR) • Author and maintain operational runbooks alongside the NOC • Manage on-call rotation, escalation paths, and incident-management tooling • Coordinate cross-functionally with NOC, Platform Engineering, and Network Engineering • Drive chaos engineering, game days, and reliability testing programs • Produce SLA performance reports in coordination with the SLA Manager • Mentor junior engineers and contribute to engineering culture

🎯 Requisitos

• 5+ years in SRE, DevOps, or production engineering roles • Strong programming skills in Go, Python, or both • Hands-on experience operating Kubernetes-based platforms at scale • Deep familiarity with observability tooling (Prometheus, Grafana, Datadog, OpenTelemetry) • Strong incident management experience including major-incident command

Candidatar-se

Vagas Similares

🕒 3 dias atrás

Zafran Security

51 - 200

🔐 Segurança

Senior DevOps Engineer at Zafran working on compliance certifications and implementing security controls across infrastructure. Collaborating with teams to enhance security posture and ensure regulatory compliance.

🇺🇸 Estados Unidos – Remoto (EUA)

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 3 dias atrás

Runpod

51 - 200

🤖 Inteligência Artificial

☁️ SaaS

🤝 B2B

Site Reliability Engineer ensuring the stability and resilience of Runpod's distributed platform. Collaborating with engineering teams on reliability frameworks and preventing incidents.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $150.000 - $200.000 / ano

💰 $20.000.000 Seed em 2024-06

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 3 dias atrás

Talkiatry

501 - 1000

🏥 Saúde

👥 B2C

Senior Site Reliability Engineer at Talkiatry, building SRE principles for mental health care. Collaborate with teams to minimize outages and improve reliability for patient services.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $160.000 - $185.000 / ano

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 4 dias atrás

Multi Media, LLC

51 - 200

💼 Consultoria

📣 Marketing

📱 Mídia

Site Reliability Engineer optimizing infrastructure resilience and performance for a leading live streaming platform. Driving enhancement and automation of cloud-based infrastructure with a global network.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $169.000 - $215.000 / ano

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 4 dias atrás

Thumbtack

1001 - 5000

🏪 Marketplace

☁️ SaaS

Senior Software Engineer designing and maintaining scalable systems to improve reliability and efficiency at Thumbtack. Collaborating with cross-functional teams to optimize platform services.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $179.400 - $232.100 / ano

💰 $75.000.000 Debt Financing - Thumbtack em 2024-07

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório