
10,000+ employees
đ¸ Finance
đł Fintech
đĽ B2C
Finance ⢠Fintech ⢠B2C
Empower is a leading provider of financial services focused on helping individuals and organizations achieve financial freedom through retirement planning and investment management. Serving over 19 million Americans, Empower offers a comprehensive suite of finance-related services, including smart planning and investment advice, and tools like the Empower Personal Dashboard⢠for a complete financial view. The company is renowned as a top retirement plan provider and works closely with personal investors, workplace plan savers, plan sponsors, and financial professionals. Empower is also recognized for initiatives in Diversity, Equity, Inclusion, and has a social commitment that bolsters community impact.
đ July 1
đŽđł India â Remote
â° Full Time
đ Senior
â DevOps & Site Reliability Engineer (SRE)
đť Ghost score 51%
Improve your chances of getting an interview by checking your resume score before you apply.

10,000+ employees
đ¸ Finance
đł Fintech
đĽ B2C
Finance ⢠Fintech ⢠B2C
Empower is a leading provider of financial services focused on helping individuals and organizations achieve financial freedom through retirement planning and investment management. Serving over 19 million Americans, Empower offers a comprehensive suite of finance-related services, including smart planning and investment advice, and tools like the Empower Personal Dashboard⢠for a complete financial view. The company is renowned as a top retirement plan provider and works closely with personal investors, workplace plan savers, plan sponsors, and financial professionals. Empower is also recognized for initiatives in Diversity, Equity, Inclusion, and has a social commitment that bolsters community impact.
⢠Design and implement highly available, fault-tolerant systems supporting critical financial transactions ⢠Architect infrastructure solutions using AWS best practices, optimizing for cost, performance, and reliability ⢠Lead complex incident response efforts and coordinate across teams to restore service rapidly ⢠Drive postmortem processes for high-severity incidents and ensure action items are completed ⢠Establish and track SLOs and SLIs for key services ⢠Design and implement disaster recovery strategies and business continuity plans ⢠Build Infrastructure as Code solutions using Terraform, including modules, workspaces, and state management ⢠Architect and optimize multi-cluster EKS environments with pod autoscaling, cluster autoscaling, and resource optimization ⢠Design observability strategies using Datadog and Splunk, including metrics, dashboards, and alerting ⢠Implement progressive delivery mechanisms such as canary and blue-green deployments within GitOps workflows ⢠Build automation frameworks to reduce operational toil and improve team efficiency ⢠Partner with development teams on application reliability, design reviews, and architectural guidance ⢠Mentor junior and intermediate SREs, conduct code reviews, and provide technical coaching ⢠Contribute to architectural decisions affecting platform reliability and scalability ⢠Evangelize SRE best practices across the engineering organization ⢠Participate in on-call rotations and reduce on-call burden ⢠Implement and maintain zero-trust security controls across infrastructure ⢠Ensure systems meet financial services regulatory requirements and internal compliance standards ⢠Conduct security reviews of infrastructure changes and deployment processes ⢠Participate in audit preparations and respond to compliance-related inquiries
⢠Bachelor's degree in Computer Science, Information Systems or similar emphasis, or equivalent experience ⢠4-7 years of experience in Site Reliability Engineering (or equivalent), with a track record of operating large-scale production systems ⢠Deep expertise in AWS, with hands-on experience across a broad range of services and architectural patterns ⢠Advanced Kubernetes knowledge, including custom resources, operators, and cluster federation concepts ⢠Expert-level proficiency in Terraform, including module development, state management, and complex workflow orchestration ⢠Strong programming skills in Python and/or Go, with ability to develop production-quality tools and services ⢠Production experience implementing observability at scale using Datadog, Splunk, or similar platforms ⢠Demonstrated experience establishing and maintaining CI/CD pipelines at enterprise scale ⢠Deep understanding of GitOps principles and experience with tools like ArgoCD or Flux ⢠Proven ability to lead complex incident response and conduct thorough postmortems ⢠Strong understanding of networking, security, and infrastructure design patterns ⢠Experience mentoring engineers and conducting technical training ⢠Preferred: Experience in financial services or payments industry ⢠Preferred: Deep knowledge of compliance frameworks (SOC 2, PCI DSS, FINRA) ⢠Preferred: AWS certifications (Solutions Architect Professional, DevOps Engineer Professional) ⢠Preferred: CKA and/or CKAD certifications ⢠Preferred: Experience with service mesh implementations (Istio, Linkerd, Consul) ⢠Preferred: Background in chaos engineering and fault injection testing ⢠Preferred: Experience with FinOps and cloud cost optimization ⢠Preferred: Contributions to open-source projects in the SRE/DevOps space ⢠Preferred: Experience implementing Operational Excellence strategies
⢠Flexible work environment ⢠Fluid career paths ⢠Internal mobility opportunities ⢠Well-being support ⢠Work-life balance ⢠Inclusive and welcoming work environment ⢠Volunteering opportunities
Apply Nowđ June 23
SRE responsible for the reliability and operability of SigNoz cloud platform while scaling observability systems and ingest pipelines. Work in a fast-paced, remote-first environment with a high-caliber team.
đŽđł India â Remote
đľ âš5M - âš10M / year
â° Full Time
đ Senior
â DevOps & Site Reliability Engineer (SRE)
Cloud
Distributed Systems
Kubernetes
Open Source
Go
đ June 19
Senior Cloud Engineer at BETSOL building and operating cloud portal workloads across Azure and GCP. Focused on DevOps and DevSecOps with AI-first development practices.
Ansible
Azure
Cloud
Google Cloud Platform
Grafana
JavaScript
Jenkins
Kubernetes
Prometheus
Python
Terraform
TypeScript
Vault
đ June 8
Senior Cloud DevOps Engineer owning secure, scalable AWS infrastructure for Revmind Labs AI's enterprise AI and analytics systems. Automating deployments, observability, security, and production reliability.
đŽđł India â Remote
đľ âš3M / year
â° Full Time
đ Senior
â DevOps & Site Reliability Engineer (SRE)
Amazon Redshift
AWS
Cloud
Docker
DynamoDB
EC2
ETL
Flask
Java
JavaScript
Microservices
MySQL
Node.js
NoSQL
Python
React
React Native
Scala
Terraform
đ June 1
Partner AI Deployment Engineer responsible for AWS deployment strategies and technical leadership in OpenAI. Guiding enterprise customers from ideation to production while influencing joint account strategy.
đŽđł India â Remote
â° Full Time
đ Senior
đ´ Lead
â DevOps & Site Reliability Engineer (SRE)
AWS
đ May 21
Senior Site Reliability Engineer focusing on developing solutions for automation and efficiency with Akamai's Compute products. Enhance reliability and operational excellence in customer-facing applications and infrastructure.
Ansible
AWS
Azure
Cloud
Distributed Systems
Google Cloud Platform
Grafana
Prometheus
Python
SaltStack
Splunk
Terraform
Go