Senior Site Reliability Engineer

🔥 12 hours ago

🇦🇷 Argentina – Remote

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

👻 Ghost score 10%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Teladoc Health

Teladoc Health

5001 - 10000 employees

🏥 Healthcare

👥 B2C

☁️ SaaS

💰 $80M Post-IPO Debt - Teladoc Health on 2016-07

Healthcare • B2C • SaaS

Teladoc Health is a telehealth company that connects patients with care providers through a virtual care platform offering 24/7 urgent care, primary care, mental health (therapy and psychiatry), chronic condition management (diabetes, hypertension, weight management), specialty consultations, and wellness services. It serves individuals directly and partners with employers, health plans, hospitals and health systems to deliver integrated virtual care solutions and technology at scale.

📋 Description

• Define, implement, and improve SLIs, SLOs, and error budgets for critical applications and platform services • Partner with application teams to improve reliability, fault tolerance, scalability, and operational readiness • Identify and eliminate recurring reliability issues through root cause analysis, automation, and architectural improvements • Design resilient systems for Azure region, zone, network, dependency, and deployment failures • Participate in production readiness reviews for services, releases, and infrastructure changes • Build observability across applications, infrastructure, networks, and cloud services • Implement monitoring for latency, traffic, errors, and saturation • Develop dashboards, alerts, logs, traces, and metrics using observability and APM platforms • Create service health dashboards for engineering, operations, and leadership • Analyze performance, bottlenecks, saturation trends, and capacity risks • Improve backup, disaster recovery, failover, and business continuity practices • Implement resiliency patterns including retries, circuit breakers, bulkheads, graceful degradation, and queue-based decoupling • Support and improve production workloads running on Microsoft Azure • Collaborate on secure, scalable Azure architecture • Enforce Azure operational standards for tagging, monitoring, backup, recovery, identity, security, and cost awareness • Conduct blameless post-incident reviews and document root causes, contributing factors, corrective actions, and prevention plans • Improve incident response processes and runbooks with Incident Management, NOC, Help Desk, and application teams • Support cloud security and compliance controls for identity, access, encryption, secrets management, vulnerability remediation, logging, and auditability • Partner with software engineering, product, security, and operations leadership • Mentor the SRE team and serve as technical authority for observability and reliability across mission-critical healthcare workloads

🎯 Requirements

• 7+ years in site reliability, including hands-on ownership of mission-critical services, through a combination of applicable work experience, training, military experience, or education • Deep Microsoft Azure experience, including Azure Monitor, Application Insights, AKS, and cloud-native operations across hybrid infrastructure • Proven experience designing and rolling out an SLO program with SLIs, SLOs, and error budget policy in production environments • Hands-on experience with enterprise observability platforms such as Datadog, Dynatrace, Elastic, Grafana, Prometheus, or LogicMonitor • Hands-on experience implementing and configuring Datadog for monitoring, observability, and alerting • Identity and credential verification, live or video interviews, and fraud or misrepresentation screening • Preferred: experience establishing or anchoring an SRE practice across multiple engineering teams • Preferred: strong incident command experience and blameless postmortems • Preferred: healthcare IT experience and familiarity with HIPAA, HITRUST, or equivalent compliance frameworks • Preferred: multi-cloud reliability experience with AWS in addition to Azure • Preferred: chaos engineering and resiliency testing experience • Preferred: Infrastructure as Code expertise with Terraform, Bicep, or Ansible • Preferred: scripting and programming proficiency in Python, PowerShell, or Go • Preferred: experience with security controls, vulnerability management, compliance audits, and cloud governance • Preferred: recognized industry certifications such as Azure Solutions Architect, Google SRE certificate, or CKA

🏖️ Benefits

• Inclusive benefits program centered around employees and their families • Tailored programs addressing employees' unique needs • Meaningful career growth, leadership, and development opportunities • Inclusive workplace and innovative culture • Candidate resources and recruiter support

Apply Now

Similar Jobs

🕒 Yesterday

Wizeline

1001 - 5000

💼 Consulting

🏥 Healthcare

📣 Marketing

Site Reliability Engineer securing multi-cloud infrastructure and automating cloud-risk remediation. Wizeline develops AI-powered digital products and platforms for global clients.

AWS

Azure

Cloud

Google Cloud Platform

Microservices

Python

Terraform

Go

🕒 Yesterday

Commit

501 - 1000

🔒 Cybersecurity

Senior DevOps Engineer modernizing AWS infrastructure and migrating the company’s codebase to GitHub Actions. Strengthening production support, infrastructure automation, observability, security, and cost optimization.

AWS

Docker

EC2

Python

Terraform

🕒 August 25

Helpware

1001 - 5000

📦 Logistics

📣 Marketing

💼 Consulting

Senior DevOps Engineer managing Azure, AKS, Terraform, and CI/CD infrastructure. Supporting secure cloud operations, observability, high availability, and Azure–Google Cloud migrations.

Ansible

Azure

Chef

Cloud

DNS

Docker

Firewalls

Flux

Google Cloud Platform

Grafana

JavaScript

Kubernetes

Microservices

Node.js

Puppet

Python

Redis

SQL

Terraform

Vault

Go

.NET

🕒 August 19

Helpware

1001 - 5000

📦 Logistics

📣 Marketing

💼 Consulting

Senior DevSecOps Engineer securing cloud infrastructure, Kubernetes platforms, and CI/CD pipelines. Automating security guardrails and software supply-chain controls across Argentina, Brazil, and Mexico.

AWS

Cloud

Kubernetes

Terraform

🕒 August 13

Azumo

51 - 200

💼 Consulting

🏥 Healthcare

📦 Logistics

DevSecOps Engineer securing Azumo’s AI-powered financial infrastructure across Latin America. Hardening cloud systems, managing vulnerabilities, and protecting AI agents while supporting compliance and customer security diligence.

AWS

Cloud

Distributed Systems