Senior Site Reliability Engineer

🕒 July 1

🇮🇳 India – Remote

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

👻 Ghost score 51%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Empower

Empower

10,000+ employees

💸 Finance

💳 Fintech

👥 B2C

Finance • Fintech • B2C

Empower is a leading provider of financial services focused on helping individuals and organizations achieve financial freedom through retirement planning and investment management. Serving over 19 million Americans, Empower offers a comprehensive suite of finance-related services, including smart planning and investment advice, and tools like the Empower Personal Dashboard™ for a complete financial view. The company is renowned as a top retirement plan provider and works closely with personal investors, workplace plan savers, plan sponsors, and financial professionals. Empower is also recognized for initiatives in Diversity, Equity, Inclusion, and has a social commitment that bolsters community impact.

📋 Description

• Design and implement highly available, fault-tolerant systems supporting critical financial transactions • Architect infrastructure solutions using AWS best practices, optimizing for cost, performance, and reliability • Lead complex incident response efforts and coordinate across teams to restore service rapidly • Drive postmortem processes for high-severity incidents and ensure action items are completed • Establish and track SLOs and SLIs for key services • Design and implement disaster recovery strategies and business continuity plans • Build Infrastructure as Code solutions using Terraform, including modules, workspaces, and state management • Architect and optimize multi-cluster EKS environments with pod autoscaling, cluster autoscaling, and resource optimization • Design observability strategies using Datadog and Splunk, including metrics, dashboards, and alerting • Implement progressive delivery mechanisms such as canary and blue-green deployments within GitOps workflows • Build automation frameworks to reduce operational toil and improve team efficiency • Partner with development teams on application reliability, design reviews, and architectural guidance • Mentor junior and intermediate SREs, conduct code reviews, and provide technical coaching • Contribute to architectural decisions affecting platform reliability and scalability • Evangelize SRE best practices across the engineering organization • Participate in on-call rotations and reduce on-call burden • Implement and maintain zero-trust security controls across infrastructure • Ensure systems meet financial services regulatory requirements and internal compliance standards • Conduct security reviews of infrastructure changes and deployment processes • Participate in audit preparations and respond to compliance-related inquiries

🎯 Requirements

• Bachelor's degree in Computer Science, Information Systems or similar emphasis, or equivalent experience • 4-7 years of experience in Site Reliability Engineering (or equivalent), with a track record of operating large-scale production systems • Deep expertise in AWS, with hands-on experience across a broad range of services and architectural patterns • Advanced Kubernetes knowledge, including custom resources, operators, and cluster federation concepts • Expert-level proficiency in Terraform, including module development, state management, and complex workflow orchestration • Strong programming skills in Python and/or Go, with ability to develop production-quality tools and services • Production experience implementing observability at scale using Datadog, Splunk, or similar platforms • Demonstrated experience establishing and maintaining CI/CD pipelines at enterprise scale • Deep understanding of GitOps principles and experience with tools like ArgoCD or Flux • Proven ability to lead complex incident response and conduct thorough postmortems • Strong understanding of networking, security, and infrastructure design patterns • Experience mentoring engineers and conducting technical training • Preferred: Experience in financial services or payments industry • Preferred: Deep knowledge of compliance frameworks (SOC 2, PCI DSS, FINRA) • Preferred: AWS certifications (Solutions Architect Professional, DevOps Engineer Professional) • Preferred: CKA and/or CKAD certifications • Preferred: Experience with service mesh implementations (Istio, Linkerd, Consul) • Preferred: Background in chaos engineering and fault injection testing • Preferred: Experience with FinOps and cloud cost optimization • Preferred: Contributions to open-source projects in the SRE/DevOps space • Preferred: Experience implementing Operational Excellence strategies

🏖️ Benefits

• Flexible work environment • Fluid career paths • Internal mobility opportunities • Well-being support • Work-life balance • Inclusive and welcoming work environment • Volunteering opportunities

Apply Now

Similar Jobs

🕒 June 23

SigNoz

11 - 50

☁️ SaaS

🏢 Enterprise

SRE responsible for the reliability and operability of SigNoz cloud platform while scaling observability systems and ingest pipelines. Work in a fast-paced, remote-first environment with a high-caliber team.

Cloud

Distributed Systems

Kubernetes

Open Source

Go

🕒 June 19

BETSOL

501 - 1000

💼 Consulting

🏥 Healthcare

📦 Logistics

Senior Cloud Engineer at BETSOL building and operating cloud portal workloads across Azure and GCP. Focused on DevOps and DevSecOps with AI-first development practices.

Ansible

Azure

Cloud

Google Cloud Platform

Grafana

JavaScript

Jenkins

Kubernetes

Prometheus

Python

Terraform

TypeScript

Vault

🕒 June 8

RevMind srl

11 - 50

🤖 Artificial Intelligence

💊 Pharmaceuticals

🏢 Enterprise

Senior Cloud DevOps Engineer owning secure, scalable AWS infrastructure for Revmind Labs AI's enterprise AI and analytics systems. Automating deployments, observability, security, and production reliability.

Amazon Redshift

AWS

Cloud

Docker

DynamoDB

EC2

ETL

Flask

Java

JavaScript

Microservices

MySQL

Node.js

NoSQL

Python

React

React Native

Scala

Terraform

🕒 June 1

OpenAI

201 - 500

🤖 Artificial Intelligence

☁️ SaaS

🏢 Enterprise

Partner AI Deployment Engineer responsible for AWS deployment strategies and technical leadership in OpenAI. Guiding enterprise customers from ideation to production while influencing joint account strategy.

AWS

🕒 May 21

Akamai Technologies

5001 - 10000

🔒 Cybersecurity

🏢 Enterprise

📱 Media

Senior Site Reliability Engineer focusing on developing solutions for automation and efficiency with Akamai's Compute products. Enhance reliability and operational excellence in customer-facing applications and infrastructure.

Ansible

AWS

Azure

Cloud

Distributed Systems

Google Cloud Platform

Grafana

Prometheus

Python

SaltStack

Splunk

Terraform

Go