Site Reliability Engineer

🔥 16 minutes ago

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Teladoc Health

Teladoc Health

5001 - 10000 employees

🏥 Healthcare

👥 B2C

☁️ SaaS

💰 $80M Post-IPO Debt - Teladoc Health on 2016-07

Healthcare • B2C • SaaS

Teladoc Health is a telehealth company that connects patients with care providers through a virtual care platform offering 24/7 urgent care, primary care, mental health (therapy and psychiatry), chronic condition management (diabetes, hypertension, weight management), specialty consultations, and wellness services. It serves individuals directly and partners with employers, health plans, hospitals and health systems to deliver integrated virtual care solutions and technology at scale.

📋 Description

• Design, implement, and maintain observability solutions across Azure • Define and standardize SLIs/SLOs/SLAs to measure service health • Develop dashboards and automated alerting to identify service degradations • Build and maintain on-call runbooks and playbooks • Drive post-incident “blameless” retrospectives • Develop automation for self-healing systems and monitoring remediation • Contribute to disaster recovery and business continuity planning • Collaborate with security, network, and system engineering teams • Mentor engineering staff in observability tools and incident management

🎯 Requirements

• 3-5 years of experience in Site Reliability Engineering • Expertise in Azure (VMs, AKS, Application Insights, Azure Monitor) • Hands-on experience with enterprise observability tools (Datadog, Dynatrace, Grafana) • Deep understanding of metrics, logs, traces, and distributed system monitoring • Proficiency with Terraform, Bicep, Ansible or similar tools • Knowledge of AKS logging in containerized environments • Experience with AI-driven monitoring or predictive alerting tools • Scripting skills in Python or PowerShell • Experience with resiliency testing tools such as Gremlin or Chaos Mesh

🏖️ Benefits

• Health insurance • Flexible working hours • Professional development opportunities • Paid time off • Remote work options

Apply Now

Similar Jobs

🔥 2 hours ago

Domus Global

501 - 1000

💼 Consulting

📣 Marketing

📦 Logistics

DevOps Engineer designing and optimizing AWS infrastructure and CI/CD processes at Nublit. Responsible for automation and continuous improvement in development teams.

🗣️🇪🇸 Spanish Required

AWS

Cloud

Grafana

Jenkins

Kubernetes

Prometheus

Terraform

🕒 Yesterday

Cognativ

11 - 50

💼 Consulting

🥽 AR/VR

🤖 Artificial Intelligence

Senior Site Reliability Engineer ensuring stability of distributed AI alerting platform. Leading incident response, capacity planning, and observability with a focus on service reliability.

Apache

AWS

Cloud

Grafana

IoT

Java

Kafka

Linux

Postgres

Prometheus

Python

Redis

Terraform

Go

🕒 4 days ago

Software Mind

1001 - 5000

🤖 Artificial Intelligence

☁️ SaaS

📡 Telecommunications

DevOps Engineer designing Infrastructure as Code and automating deployment processes for a multicultural engineering team at Software Mind. Collaborating on CI/CD pipelines while ensuring system observability and security compliance.

AWS

Azure

Cloud

Docker

Grafana

Kubernetes

Prometheus

Python

Terraform

🕒 5 days ago

Sezzle

201 - 500

💳 Fintech

👥 B2C

🛍️ eCommerce

Senior Site Reliability Engineer at Sezzle enhancing infrastructure reliability and scalability through modern tech solutions. Collaborating across teams while driving innovation in the fintech space.

AWS

Distributed Systems

Grafana

Kubernetes

Microservices

MySQL

Postgres

Prometheus

RDBMS

SQL

Go

🕒 6 days ago

intive

1001 - 5000

💼 Consulting

🏥 Healthcare

📣 Marketing

DevOps / Cloud Engineer managing Azure infrastructure and CI/CD pipelines at intive. Join a diverse team driving innovation in technology across various industries.

Azure

Cloud

Docker

Kubernetes

MS SQL Server

SQL

Terraform