Cloud Systems Engineer – Site Reliability

🕒 Setembro 18

🔔 Pennsylvania – Remoto

infoinfo

💵 $110.000 - $150.000 / ano

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

👻 Score fantasma 10%

infoinfo

🗣️🇺🇸🇬🇧 Inglês obrigatório

Candidatar-se
Encontrar Vagas Remotas Similares

📊 Verifique sua pontuação de currículo para esta vaga

Melhore suas chances de conseguir uma entrevista verificando sua pontuação de currículo antes de se candidatar.

Logo of TherapyNotes, LLC

TherapyNotes, LLC

51 - 200 funcionários

Fundada em 2010

💼 Consultoria

⚖️ Jurídico

🏥 Saúde

Consulting • Legal • Healthcare

TherapyNotes, LLC é um sistema de gestão de práticas abrangente, projetado especificamente para profissionais de saúde comportamental, incluindo psicólogos, terapeutas, psiquiatras e assistentes sociais. A plataforma oferece uma gama de recursos como agendamento, telemedicina, prontuários eletrônicos (EHR), faturamento e um portal do cliente, todos integrados em uma solução de software segura e fácil de usar. O TherapyNotes visa simplificar os fluxos de trabalho clínicos, melhorar o atendimento ao paciente e reduzir as cargas administrativas para práticas de saúde mental, apoiando-as com um serviço de atendimento ao cliente dedicado e inovação contínua baseada no feedback dos usuários.

Descrição

• Own and continuously improve Datadog usage across metrics, logs, traces, dashboards, monitors, alerts, and service-level views • Design, implement, and maintain high-availability, high-throughput, data- and compute-intensive critical systems supporting a growing 24×7 SaaS platform • Partner with service owners to define and improve reliability through SLIs, SLOs, error budgets, actionable alerting, and operational-readiness practices • Participate in and drive incident management, including incident command or technical response, triage, service restoration, escalation, communication, documentation, root cause analysis, and corrective actions • Investigate infrastructure and application issues using metrics, logs, distributed traces, and code-level context • Improve deployment safety and service resilience through automated validation, recovery and rollback capabilities, reliability testing, and failure-mode analysis • Ensure newly introduced systems are supportable and maintainable by development and operations • Provide escalated technical guidance and support to technology teams • Provide on-call coverage for production support and other duties as required • Ensure systems and operational activities comply with security, HIPAA, and operating policies • Eliminate repetitive operational toil using Bash, PowerShell, Python, or Ansible • Manage infrastructure as code using Terraform/OpenTofu and configuration automation using Ansible

🎯 Requisitos

• BS degree in Information Systems, Engineering, or equivalent experience • 5+ years of engineering experience in Systems Engineering, Cloud or Platform Engineering, DevOps, Software Engineering, and/or SRE • Experience designing and operating production systems using cloud-based compute, storage, networking, and containerization technologies; Azure and Kubernetes preferred • Strong Linux systems and networking fundamentals, with experience troubleshooting complex distributed systems in production • Expertise with an observability platform; Datadog experience strongly preferred • Experience with Prometheus, Grafana, New Relic, or equivalent platforms is also valuable • Experience with scripting and operational automation using Bash, PowerShell, or Python, along with infrastructure-as-code and configuration-management practices • Experience participating in production on-call rotations, incident response, root cause analysis, and post-incident improvement • Experience working in Agile/DevOps environments and operating production services using ITSM practices where applicable • Prior software development experience—or experience investigating application behavior through code, logs, and distributed traces—is a plus

🏖️ Benefícios

• Employer sponsored health, dental, vision, life, and disability insurance • Retirement plan with company contribution • Annual company profit sharing • Personal development/training budget • Open, collaborative work environment • Extensive 2-week onboarding plan • Comprehensive mentorship program

Candidatar-se

Vagas Similares

🕒 Setembro 17

Sphera

1001 - 5000

💼 Consultoria

🏥 Saúde

📦 Logística

Cybersecurity Engineer securing Sphera’s environmental, health, safety, and sustainability software for U.S. government and DoD clients. Managing RMF, STIG, ATO, vulnerability remediation, and DevSecOps compliance.

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Setembro 17

Arista Networks

1001 - 5000

🏢 Corporativo

📡 Telecomunicações

FedRAMP SRE operating Arista Networks’ Kubernetes-native CloudVision networking SaaS. Ensuring reliable, secure, scalable production systems and leading infrastructure projects.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $101.000 - $161.000 / ano

💰 $2.600.000 Post-IPO Debt em 2015-05

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🦅 Patrocina Visto H1B

infoinfo

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Setembro 17

Gainwell Technologies

10.000+ funcionários

💼 Consultoria

📦 Logística

⚕️ Seguro de Saúde

DevOps Engineer automating configuration, releases, and cloud infrastructure for Gainwell Technologies’ healthcare platforms. Building scripts, CI processes, and infrastructure as code for SaaS and PaaS products.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $85.000 - $121.400 / ano

💰 Grant em 2023-06

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🦅 Patrocina Visto H1B

infoinfo

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Setembro 17

Gainwell Technologies

10.000+ funcionários

💼 Consultoria

📦 Logística

⚕️ Seguro de Saúde

DevOps Engineer automating infrastructure, configuration management, and software releases. Supporting Gainwell’s cloud-based healthcare technology platforms and development teams.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $58.200 - $83.200 / ano

💰 Grant em 2023-06

⏰ Tempo Integral

🟡 Pleno

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🦅 Patrocina Visto H1B

infoinfo

🗣️🇺🇸🇬🇧 Inglês obrigatório

🕒 Setembro 17

Cerence Inc.

1001 - 5000

💼 Consultoria

📦 Logística

🏭 Manufatura

Principal SRE leading Cerence's cloud-native automotive AI reliability function. Driving SLOs, incident response, observability, automation, and technical team direction.

🇺🇸 Estados Unidos – Remoto (EUA)

💵 $132.000 - $211.400 / ano

💰 Grant em 2020-12

⏰ Tempo Integral

🟠 Sênior

⛑ DevOps & Engenheiro de Confiabilidade do Site (SRE)

🗣️🇺🇸🇬🇧 Inglês obrigatório