Senior Data Reliability Engineer, AWS

🔥 1 minute ago

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Empower

Empower

10,000+ employees

💸 Finance

💳 Fintech

👥 B2C

Finance • Fintech • B2C

Empower is a leading provider of financial services focused on helping individuals and organizations achieve financial freedom through retirement planning and investment management. Serving over 19 million Americans, Empower offers a comprehensive suite of finance-related services, including smart planning and investment advice, and tools like the Empower Personal Dashboard™ for a complete financial view. The company is renowned as a top retirement plan provider and works closely with personal investors, workplace plan savers, plan sponsors, and financial professionals. Empower is also recognized for initiatives in Diversity, Equity, Inclusion, and has a social commitment that bolsters community impact.

📋 Description

• Own the reliability and stability of production data pipelines and data platform services • Diagnose and resolve data pipeline failures, delays, and data quality issues in production environments • Investigate issues across distributed data systems (e.g., Spark/EMR workloads, ingestion pipelines, warehouse performance) • Lead or support incident response, including triage, mitigation, and long-term resolution • Perform root cause analysis (RCA) and implement durable fixes to prevent recurrence • Define and improve data SLAs (freshness, latency, completeness) and ensure adherence • Design and enhance monitoring, alerting, and observability for data systems • Develop automation and tooling to reduce operational toil and improve system resilience • Contribute to disaster recovery (DR) and resiliency planning, including backup validation and recovery workflows • Partner with engineering teams to improve pipeline design, reliability, and operational readiness • Create and maintain runbooks, SOPs, and operational documentation • Participate in occasional off-hours support for production data systems when required

🎯 Requirements

• Minimum 5 years of experience working with production data platforms in AWS environments • Prior experience building data pipelines and seeing them through production, including exposure to real-world failures and operational challenges • Strong experience with Python and SQL in real data systems • Hands-on experience troubleshooting distributed data processing systems (e.g., Spark/EMR, Redshift, streaming systems) • Proven ability to debug and resolve production issues in data pipelines and data platforms • Experience with AWS data services (such as EMR, Redshift, DynamoDB, S3, or similar) • Experience handling production incidents and performing root cause analysis • Strong problem-solving mindset and ability to work through ambiguous production issues.

🏖️ Benefits

• Medical, dental, vision and life insurance • Retirement savings – 401(k) plan with generous company matching contributions (up to 6%), financial advisory services, potential company discretionary contribution, and a broad investment lineup • Tuition reimbursement up to $5,250/year • Business-casual environment that includes the option to wear jeans • Generous paid time off upon hire – including a paid time off program plus ten paid company holidays and three floating holidays each calendar year • Paid volunteer time — 16 hours per calendar year • Leave of absence programs – including paid parental leave, paid short- and long-term disability, and Family and Medical Leave (FMLA) • Business Resource Groups (BRGs) – BRGs facilitate inclusion and collaboration across our business internally and throughout the communities where we live, work and play. BRGs are open to all.

Apply Now

Similar Jobs

🔥 44 minutes ago

Harrods

1001 - 5000

🍽️ Food & Beverage

📣 Marketing

💼 Consulting

DevOps Engineer at Harrods responsible for supporting service and deployment pipelines across digital technology team. Collaborate with various teams to ensure seamless service and improve processes.

Azure

Cloud

ITSM

Terraform

🔥 48 minutes ago

Scientific Games

10,000+ employees

🎮 Gaming

🤝 B2B

Senior DevOps Engineer enhancing infrastructure and deployment reliability for Scientific Games. Collaborating across teams to scale secure, high-performing platforms and improve delivery efficiency.

AWS

Azure

Cloud

Docker

Jenkins

Kubernetes

Linux

Vault

🔥 2 hours ago

Andromeda

11 - 50

🏥 Healthcare

💼 Consulting

🏨 Hospitality

Engineer embedded with teams running large-scale training and inference on GPU clusters in production. Responsible for onboarding, debugging, and improving performance while ensuring reliability.

Ansible

Kubernetes

Linux

Python

Terraform

Go

🔥 4 hours ago

Millennium

201 - 500

💼 Consulting

🎖️ Defense

🔒 Cybersecurity

AWS DevOps Engineer responsible for designing and maintaining AWS cloud environments for Marine Corps IT systems. Collaborating across teams to ensure performance, security, and compliance needs are met.

AWS

Cloud

Docker

Kubernetes

Python

Terraform

🔥 7 hours ago

Smithfield Foods

10,000+ employees

🏭 Manufacturing

🌾 Agriculture

🍽️ Food & Beverage

Senior Utilities Engineer at Smithfield Foods optimizing utility systems for industrial refrigeration and ensuring compliance. Collaborating with facilities teams for operational efficiency and system improvements across various locations.