Senior Site Reliability Engineer

🕒 June 27

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Visa

Visa

10,000+ employees

💳 Fintech

💸 Finance

🏦 Banking

💰 Post IPO equity on 2016-10

Fintech • Finance • Banking

Visa is a global payments technology company that connects consumers, merchants, financial institutions and governments to enable electronic payments and commerce. It operates a worldwide card network for credit, debit and prepaid payments and provides payment processing, security and fraud-mitigation services, data analytics and value-added products for both consumers and businesses. Visa invests in next-generation payment solutions (including AI-driven commerce), supports financial inclusion initiatives through the Visa Foundation, and offers consumer card benefits and merchant services.

📋 Description

• Own the end‑to‑end lifecycle (design, provisioning, upgrades, maintenance, and decommissioning) of core platform components, including: Cloud infrastructure primitives, Kubernetes clusters and cluster services, Networking, ingress, and service discovery, Service Mesh and supporting data‑plane components. • Design platform components to be resilient by default, applying SRE principles such as Fault isolation and graceful degradation, Capacity planning and saturation control, Reduced operational toil and clear failure modes. • Lead the design and implementation of infrastructure bootstrap orchestration, including Automated cluster and environment provisioning, Deterministic, repeatable platform bring‑up and teardown, Dependency‑aware orchestration across cloud, network, and Kubernetes layers. • Drive Infrastructure‑as‑Code and GitOps‑first practices to ensure: Platform components are reproducible and auditable, Changes are automated, testable, and reversible, Manual intervention is minimized or eliminated. • Identify automation gaps and lead initiatives that reduce human effort, onboarding time, and operational risk. • Apply and promote SRE operational excellence practices, including clear ownership and runbooks for platform components, Participation in on‑call rotation as a platform reliability escalation point, Incident response, post‑incident reviews, and problem management. • Improve day‑2 operations by standardizing upgrade/rollback strategies and reducing MTTD/MTTR. • Ensure platform operations align with security, compliance, and internal control requirements. • Collaborate with engineering teams across the organization to influence platform adoption, reliability standards, and cloud‑native best practices.

🎯 Requirements

• Proven experience operating and administering Kubernetes at scale in production environments. • Strong experience with container orchestration platforms and cloud architecture fundamentals (networking, IAM/security concepts, and reliability patterns). • Experience with Infrastructure as Code (Terraform preferred) and automation‑first workflows. • Familiarity with GitOps practices and CI/CD pipelines. • Strong troubleshooting skills for distributed systems, including root‑cause analysis and reliability improvements. • Experience with observability concepts and practices (monitoring, logging, alerting, tracing). • Proficiency in English at B2 level or above (Upper-Intermediate).

🏖️ Benefits

• Health insurance • 401(k) matching • Flexible work hours • Paid time off • Remote work options

Apply Now

Similar Jobs

🕒 June 27

Devexperts

501 - 1000

💳 Fintech

☁️ SaaS

💸 Finance

Site Reliability Engineer ensuring stability of trading platforms at Devexperts. Collaborating with teams to deploy and maintain services effectively while fostering automation.

🗣️🇧🇷🇵🇹 Portuguese Required

Ansible

Apache

Cloud

Docker

ElasticSearch

Firewalls

Grafana

HAProxy

Linux

NGINX

OpenShift

TCP/IP

Terraform

Unix

🕒 June 27

Gorillas Group

51 - 200

💼 Consulting

📣 Marketing

📦 Logistics

SRE Senior role focused on enhancing reliability, observability, and operational processes at Goritek, a tech company in a regulated market.

🗣️🇧🇷🇵🇹 Portuguese Required

Cloud

Kubernetes

🕒 June 25

RHI Magnesita

10,000+ employees

💼 Consulting

📦 Logistics

🏭 Manufacturing

DevOps Analyst SR supporting the digital transformation at RHI Magnesita. Engage in developing technical solutions and guiding teams in cloud infrastructure.

Ansible

Azure

Chef

Cloud

Docker

JavaScript

Kubernetes

Node.js

Terraform

.NET

🕒 June 25

Airbnb

5001 - 10000

✈️ Travel

📦 Logistics

🛍️ eCommerce

Senior Software Engineer in Reliability Engineering Team at Airbnb. Focusing on developing tools for service reliability and incident management.

AWS

Cloud

Distributed Systems

Docker

Google Cloud Platform

Java

Kubernetes

Python

Go

🕒 June 24

Experian

10,000+ employees

💼 Consulting

📣 Marketing

📦 Logistics

SRE Specialist managing cloud operations and automation. Collaborating in an SRE squad to enhance productivity and resolve technical challenges.

🗣️🇧🇷🇵🇹 Portuguese Required

AWS

Azure

Cloud

Grafana

JMeter

Kubernetes

Prometheus

Terraform