Site Reliability Engineer

🕒 3 dias atrás

🗣️🇺🇸🇬🇧 Inglês obrigatório

Candidatar-se
Encontrar Vagas Remotas Similares

📊 Verifique sua pontuação de currículo para esta vaga

Melhore suas chances de conseguir uma entrevista verificando sua pontuação de currículo antes de se candidatar.

Logo of Runpod

Runpod

51 - 200 funcionários

Fundada em 2022

🤖 Inteligência Artificial

☁️ SaaS

🤝 B2B

💰 $20.000.000 Seed em 2024-06

Artificial Intelligence • SaaS • B2B

Runpod é uma plataforma de nuvem que fornece computação sob demanda com GPU e infraestrutura gerenciada voltada para desenvolvimento e implantação de IA. A plataforma oferece "Pods" de GPU em 31 regiões globais, endpoints de GPU serverless para inferência de baixa latência, clusters de GPU multinó para treinamento distribuído e um hub para implantar modelos e templates de código aberto. A Runpod destaca-se pelo rápido início (menos de 200ms em iniciações a frio), escalonamento automático de zero a milhares de trabalhadores, suporte para mais de 30 GPU SKUs e ferramentas para todo o ciclo de vida da IA, desde experimentos até a produção, direcionada a desenvolvedores e equipes de IA empresariais.

Descrição

• Define and implement SLIs/SLOs for critical services • Lead incident response and coordinate cross-team mitigation efforts • Conduct blameless postmortems and ensure corrective actions are completed • Perform production readiness reviews for new services and features • Identify systemic risks and drive preventative improvements • Design and improve monitoring, alerting, and dashboards (Prometheus, Grafana, etc.) • Build internal tooling for reliability tracking and reporting • Automate recurring operational workflows • Strengthen CI/CD reliability and release processes • Partner with engineering teams to improve system resilience • Provide guidance on fault tolerance, scalability, and failure handling. • Contribute to architectural discussions with a reliability-first mindset.

🎯 Requisitos

• 5+ years of experience in SRE, Reliability Engineering, or Production Engineering • Strong Linux systems and Networking expertise • Experience managing containerized production systems • Strong understanding of distributed systems and failure modes • Experience defining and managing SLIs/SLOs • Proven incident response and postmortem leadership experience • Strong scripting or programming skills • Experience with monitoring and alerting systems • Excellent written communication skills • Successful completion of a background check. • Preferred: Experience with GPU infrastructure or AI/ML platforms • Experience improving reliability in high-growth or large scale environments • Familiarity with GPU observability tooling • Experience with Infrastructure as Code • Experience working in startup environments • Experience building internal reliability platforms or frameworks.

🏖️ Benefícios

• Meaningful equity in a fast-growing company- everyone on the team receives stock options — your impact drives our growth, and you share in the upside. • Generous medical, dental & vision plans • Flexible PTO- take the time you need to recharge • Most roles are remote work first with an inclusive, collaborative teams utilizing slack as the main form of internal communication • Join a passionate team on the cutting edge of AI infrastructure — where culture, learning, and ownership are at the heart of how we scale.

Candidatar-se

Vagas Similares

🕒 3 dias atrás

Talkiatry

501 - 1000

🏥 Saúde

👥 B2C

Senior Site Reliability Engineer at Talkiatry, building SRE principles for mental health care. Collaborate with teams to minimize outages and improve reliability for patient services.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $160.000 - $185.000 / ano

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 4 dias atrás

Multi Media, LLC

51 - 200

💼 Consultoria

📣 Marketing

📱 Mídia

Site Reliability Engineer optimizing infrastructure resilience and performance for a leading live streaming platform. Driving enhancement and automation of cloud-based infrastructure with a global network.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $169.000 - $215.000 / ano

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 4 dias atrás

Thumbtack

1001 - 5000

🏪 Marketplace

☁️ SaaS

Senior Software Engineer designing and maintaining scalable systems to improve reliability and efficiency at Thumbtack. Collaborating with cross-functional teams to optimize platform services.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $179.400 - $232.100 / ano

💰 $75.000.000 Debt Financing - Thumbtack em 2024-07

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 4 dias atrás

Manulife

10.000+ funcionários

🛡️ Seguros

💸 Finanças

Lead Power Platform Reliability Engineer enhancing enterprise-level solutions through collaboration and mentorship. Shape future data-driven applications and drive cloud integration.

🇺🇸 Estados Unidos – Remoto (EUA)

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 4 dias atrás

ICF

5001 - 10000

💼 Consultoria

🏛️ Governo

🏥 Saúde

DevOps Engineer building healthcare reporting services for ICF. Implementing cloud-based solutions using AWS and fostering collaboration on CI/CD pipeline improvements.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $108.476 - $184.409 / ano

💰 $29.000.000 Grant em 2023-03

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório