Senior Site Reliability Engineer

🕒 Setembro 18

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $210.000 - $275.000 / ano

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🦅 Patrocina Visto H1B

infoinfo

👻 Score fantasma 20%

infoinfo

🗣️🇺🇸🇬🇧 Inglês obrigatório

Candidatar-se
Encontrar Vagas Remotas Similares

📊 Verifique sua pontuação de currículo para esta vaga

Melhore suas chances de conseguir uma entrevista verificando sua pontuação de currículo antes de se candidatar.

Logo of Replit

Replit

51 - 200 funcionários

🤖 Inteligência Artificial

🤝 B2B

Artificial Intelligence • B2B

Um computador multiplayer para criar e compartilhar software. Deixe a inteligência artificial escrever código para você com nossa nova funcionalidade Gerar Código.

Descrição

• Join Replit’s Site Reliability Engineering team to ensure the reliability, scalability, and performance of infrastructure serving millions of developers worldwide • Design and implement comprehensive monitoring, alerting, dashboards, metrics, and logging strategies • Architect and implement infrastructure automation using Terraform, Ansible, or Pulumi • Design and maintain CI/CD pipelines for reliable and consistent deployments • Create self-healing systems that respond automatically to common failure scenarios • Define and implement Service Level Objectives (SLOs) and Service Level Indicators (SLIs) with product and engineering teams • Build systems to track and report reliability metrics • Lead incident response efforts and conduct thorough post-mortems • Develop and maintain runbooks for critical services • Build tools and processes to reduce Mean Time To Recovery (MTTR) • Identify and resolve infrastructure performance bottlenecks • Implement capacity planning strategies and optimize resource utilization • Reduce latency and improve system efficiency across global regions

🎯 Requisitos

• 4-8 years of experience in Site Reliability Engineering or similar roles (DevOps, Systems Engineering, Infrastructure Engineering) • Strong programming skills in languages commonly used for automation (Python, Go, or similar) • Deep understanding of distributed systems • Experience with container orchestration platforms (Kubernetes) and cloud-native technologies • Proven track record of implementing and maintaining monitoring/observability solutions • Strong incident management skills with experience leading incident response • Experience with infrastructure as code and configuration management tools • Experience with Google Cloud Platform (GCP) services and tools (bonus) • Knowledge of modern observability platforms (Prometheus, Grafana, Datadog, etc.) (bonus) • Ability to approach complex operational challenges systematically and devise effective solutions • Capable of working independently while collaborating effectively with cross-functional teams • Strong communication skills to explain complex technical concepts to technical and non-technical audiences • Passion for staying current with industry best practices and new technologies • Strong belief in automating repetitive tasks and building self-healing systems • Legally authorized to work in the United States

🏖️ Benefícios

• Competitive Salary & Equity • 401(k) Program with a 4% match (US Only) • Health, Dental, Vision and Life Insurance • Short Term and Long Term Disability • Paid Parental, Medical, Caregiver Leave • Flexible Time Off (FTO) + Holidays • Commuter Benefits (In-Office & US Only) • Monthly Wellness Stipend • Autonomous Work Environment • In Office Set-Up Reimbursement (In-Office Only) • Quarterly Team Gatherings • In Office Amenities (In-Office Only)

Candidatar-se

Vagas Similares

🕒 Setembro 18

TherapyNotes, LLC

51 - 200

💼 Consultoria

⚖️ Jurídico

🏥 Saúde

Site Reliability Engineer improving reliability, observability, and incident response for TherapyNotes’ behavioral health practice-management and EHR SaaS platform. Designing resilient cloud infrastructure for 24×7 production services.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $110.000 - $150.000 / ano

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Setembro 17

Sphera

1001 - 5000

💼 Consultoria

🏥 Saúde

📦 Logística

Cybersecurity Engineer securing Sphera’s environmental, health, safety, and sustainability software for U.S. government and DoD clients. Managing RMF, STIG, ATO, vulnerability remediation, and DevSecOps compliance.

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Setembro 17

Arista Networks

1001 - 5000

🏢 Corporativo

📡 Telecomunicações

FedRAMP SRE operating Arista Networks’ Kubernetes-native CloudVision networking SaaS. Ensuring reliable, secure, scalable production systems and leading infrastructure projects.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $101.000 - $161.000 / ano

💰 $2.600.000 Post-IPO Debt em 2015-05

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🦅 Patrocina Visto H1B

infoinfo

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Setembro 17

Gainwell Technologies

10.000+ funcionários

💼 Consultoria

📦 Logística

⚕️ Seguro de Saúde

DevOps Engineer automating infrastructure, configuration management, and software releases. Supporting Gainwell’s cloud-based healthcare technology platforms and development teams.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $58.200 - $83.200 / ano

💰 Grant em 2023-06

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🦅 Patrocina Visto H1B

infoinfo

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Setembro 17

Gainwell Technologies

10.000+ funcionários

💼 Consultoria

📦 Logística

⚕️ Seguro de Saúde

DevOps Engineer automating configuration, releases, and cloud infrastructure for Gainwell Technologies’ healthcare platforms. Building scripts, CI processes, and infrastructure as code for SaaS and PaaS products.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $85.000 - $121.400 / ano

💰 Grant em 2023-06

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🦅 Patrocina Visto H1B

infoinfo

🗣️🇺🇸🇬🇧 Inglês obrigatório