Site Reliability Engineer

🕒 Ontem

🗣️🇺🇸🇬🇧 Inglês obrigatório

Candidatar-se
Encontrar Vagas Remotas Similares

📊 Verifique sua pontuação de currículo para esta vaga

Melhore suas chances de conseguir uma entrevista verificando sua pontuação de currículo antes de se candidatar.

Logo of Vynca

Vynca

51 - 200 funcionários

Fundada em 2014

💼 Consultoria

⚕️ Seguro de Saúde

🏥 Saúde

Consulting • Healthcare Insurance • Healthcare

A Vynca é uma empresa de saúde que se especializa em cuidados paliativos para pacientes com doenças graves. A organização foca em aprimorar o gerenciamento de cuidados, oferecendo suporte que inclui gestão da dor, aconselhamento em saúde mental, apoio espiritual e planejamento antecipado de cuidados. A Vynca tem como objetivo auxiliar pacientes a navegarem por condições crônicas e cuidados complexos, garantindo que recebam suporte abrangente de uma equipe dedicada de profissionais da saúde, tanto presencialmente quanto virtualmente. Os serviços geralmente são cobertos por grandes planos de saúde, levando cuidados essenciais diretamente aos pacientes em suas casas ou através de capacidades de telemedicina.

Descrição

• Design, provision, and manage AWS infrastructure using Terraform as the source of truth • Operate, maintain, and scale production workloads running on Kubernetes • Package, deploy, and manage applications using Helm and infrastructure automation tools • Build, operate, and improve distributed and event-driven systems, including event sourcing, partitioning, event ordering, replay, and failure recovery mechanisms • Define, monitor, and maintain Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets to balance reliability and engineering velocity • Develop automation for deployment, scaling, monitoring, incident response, and operational workflows to reduce manual effort and improve system resilience • Own platform observability by implementing and maintaining metrics, logging, tracing, monitoring, and alerting solutions • Lead incident response efforts, facilitate blameless postmortems, and drive long-term corrective actions that improve system reliability • Partner with Product and Engineering teams on capacity planning, performance optimization, and resilient system design • Implement and maintain security best practices to support HIPAA, SOC 2, and other compliance requirements • Participate in an on-call rotation and provide operational support for production systems

🎯 Requisitos

• Three to five (3–5) years of experience in Site Reliability Engineering, DevOps Engineering, Platform Engineering, Cloud Infrastructure Engineering, or similar infrastructure-focused roles • Bachelor’s degree in Computer Science, Information Systems, Software Engineering, or a related technical field; equivalent professional experience will also be considered • Strong hands-on experience operating production workloads within AWS environments • Proven experience managing infrastructure as code using Terraform, including module development, state management, and deployment automation • Experience operating and supporting production Kubernetes environments • Hands-on experience deploying and managing applications using Helm • Experience working with distributed systems, event-driven architectures, or event-sourcing platforms, including concepts such as partitioning, event ordering, replay, and fault tolerance • Experience establishing and managing observability practices including monitoring, logging, tracing, alerting, and incident response • Strong understanding of Linux systems administration, networking, cloud architecture, and distributed systems fundamentals • Experience designing, implementing, and maintaining CI/CD pipelines and deployment automation • Strong problem-solving skills with the ability to troubleshoot complex infrastructure and application issues • Excellent written and verbal communication skills with the ability to collaborate effectively across technical and non-technical teams • High level of ownership, accountability, and initiative with a proactive approach to reliability and operational excellence • Ability and willingness to participate in an on-call rotation supporting production systems

🏖️ Benefícios

• medical, dental, and vision insurance • income protection benefits • flexible PTO • company holidays • 401k • access to other wellness benefits

Candidatar-se

Vagas Similares

🕒 Ontem

Akkadian Labs

51 - 200

☁️ SaaS

🏢 Corporativo

📡 Telecomunicações

DevOps Engineer supporting design, implementation, and maintenance of secure infrastructure. Collaborating with teams to enable reliable deployments and improve system observability at Akkadian Labs.

🇺🇸 Estados Unidos – Remoto (EUA)

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Ontem

Sentara Health

10.000+ funcionários

🏥 Saúde

⚕️ Seguro de Saúde

DevOps Engineer focusing on developing Azure cloud infrastructure and Kubernetes environments for Sentara Health. Managing CI/CD pipelines and collaborating with development teams in a remote setting.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $80.204 - $133.681 / ano

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Ontem

Pekin Insurance

501 - 1000

🏥 Saúde

📦 Logística

🛡️ Seguros

Software Engineer at Pekin Insurance collaborating with product teams to create and support rich applications. Designing, coding, and implementing solutions within an Agile framework.

🗣️🇺🇸🇬🇧 Inglês obrigatório

Guidewire

🕒 Ontem

Michael Baker International

1001 - 5000

💼 Consultoria

🎖️ Defesa

📦 Logística

Senior Cloud DevOps Engineer managing enterprise cloud infrastructure at Michael Baker International. Collaborating on digital transformation and driving CI/CD automation across Azure and AWS.

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Ontem

Pinterest

1001 - 5000

📱 Mídia

👥 B2C

Senior Site Reliability Engineer at Pinterest ensuring reliability of cloud-native platforms on AWS and Kubernetes. Collaborating with teams to improve operational practices and incident responses.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $139.764 - $287.749 / ano

💰 Post IPO equity em 2022-08

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🦅 Patrocina Visto H1B

info

🗣️🇺🇸🇬🇧 Inglês obrigatório