Senior Site Reliability Engineer

🕒 July 1

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Zipdev

Zipdev

51 - 200 employees

Founded 2017

💼 Consulting

📦 Logistics

📣 Marketing

Consulting • Logistics • Marketing

Zipdev is a company that specializes in providing remote talent solutions, particularly focused on hiring top Latin American professionals. Zipdev enables companies to build high-performing teams in their own time zones while reducing costs compared to local hiring. The company offers a streamlined hiring process that includes defining hiring needs, selecting top candidates, and onboarding new team members, handling payroll and HR overhead. Zipdev fills a variety of technical and professional roles, such as software engineers, project managers, and designers. The company is recognized for its ability to provide cultural alignment and scalable, flexible hiring solutions.

📋 Description

• Support observability tooling implementation (Datadog and/or Azure Monitor/App Insights) and help build SLO definitions, alert rules, and synthetic checks • Participate in a PagerDuty on-call rotation, including escalation handling and incident documentation • Build and maintain operational runbooks for incident response, rollback, and recovery scenarios • Contribute to deployment automation work (blue/green or canary patterns) and Infrastructure as Code • Work across Azure SQL and Cosmos DB environments, supporting performance and cost optimization initiatives • Collaborate closely with US-based engineers during overlapping working hours

🎯 Requirements

• 5+ years in SRE, DevOps, or cloud infrastructure roles • Strong hands-on experience with Microsoft Azure (Azure SQL, Cosmos DB, Container Apps, App Service) • Experience with observability tooling (Datadog, Azure Monitor, or similar) and on- call/incident response • Familiarity with Infrastructure as Code (Terraform preferred) • Strong written and spoken English; you'll be in daily communication with US-based team members and, at times, client stakeholders • **Availability with meaningful overlap with US Eastern or Mountain time zones** • Experience working in HIPAA-regulated environments, including handling PHI under a Business Associate Agreement (BAA) and working within least-privilege, audited access controls • Willingness to complete a healthcare-industry-standard background check prior to production access • **On-Call Expectations** • This role includes participation in a pager-based on-call rotation via PagerDuty, covering SEV- 1/SEV-2 incidents on a shared schedule with the SRE team. This is a core, required part of the role, not an occasional ask.

🏖️ Benefits

• Work remotely • Vacation: 10 business days a year • Holidays: 5 National Holidays a year • Company Holidays: 5 Company Holidays a year (Christmas Eve, Christmas Day, New Year's Eve, New Year's Day, Zipdev Day) • Parental Leave • Health Care Reimbursement • Active Lifestyle Reimbursement • Quarterly Home Office Reimbursement • Payroll Deduction Purchase Plans • Longevity Bonus • Continuous Learning Bonus • Access to Training and Professional Development Platforms • Did we mention it's REMOTE?!!

Apply Now

Similar Jobs

🕒 June 30

Oowlish

51 - 200

💼 Consulting

📣 Marketing

Senior Site Reliability Engineer responsible for maintaining business-critical production systems at Oowlish. Collaborating globally while driving reliability and operational excellence within high-availability environments.

Cloud

Python

TypeScript

Go

🕒 June 27

Visa

10,000+ employees

💳 Fintech

💸 Finance

🏦 Banking

Site Reliability Engineer managing Kubernetes platforms for Visa with focus on reliability and scalability. Collaborating across teams to enhance platform performance and security while implementing Infrastructure-as-Code practices.

Bootstrap

Cloud

Distributed Systems

Kubernetes

Terraform

🕒 June 27

Devexperts

501 - 1000

💳 Fintech

☁️ SaaS

💸 Finance

Site Reliability Engineer ensuring stability of trading platforms at Devexperts. Collaborating with teams to deploy and maintain services effectively while fostering automation.

🗣️🇧🇷🇵🇹 Portuguese Required

Ansible

Apache

Cloud

Docker

ElasticSearch

Firewalls

Grafana

HAProxy

Linux

NGINX

OpenShift

TCP/IP

Terraform

Unix

🕒 June 27

Gorillas Group

51 - 200

💼 Consulting

📣 Marketing

📦 Logistics

SRE Senior role focused on enhancing reliability, observability, and operational processes at Goritek, a tech company in a regulated market.

🗣️🇧🇷🇵🇹 Portuguese Required

Cloud

Kubernetes

🕒 June 25

Airbnb

5001 - 10000

✈️ Travel

📦 Logistics

🛍️ eCommerce

Senior Software Engineer in Reliability Engineering Team at Airbnb. Focusing on developing tools for service reliability and incident management.

AWS

Cloud

Distributed Systems

Docker

Google Cloud Platform

Java

Kubernetes

Python

Go