Senior Data Reliability Engineer, AWS

🕒 August 1

🇺🇸 United States – Remote

💵 $105.7k - $149.3k / year

⏰ Full Time

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🦅 H1B Visa Sponsor

infoinfo

👻 Ghost score 38%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Empower

Empower

10,000+ employees

💸 Finance

💳 Fintech

👥 B2C

Finance • Fintech • B2C

Empower is a leading provider of financial services focused on helping individuals and organizations achieve financial freedom through retirement planning and investment management. Serving over 19 million Americans, Empower offers a comprehensive suite of finance-related services, including smart planning and investment advice, and tools like the Empower Personal Dashboard™ for a complete financial view. The company is renowned as a top retirement plan provider and works closely with personal investors, workplace plan savers, plan sponsors, and financial professionals. Empower is also recognized for initiatives in Diversity, Equity, Inclusion, and has a social commitment that bolsters community impact.

📋 Description

• Own the reliability and stability of production data pipelines and data platform services • Diagnose and resolve data pipeline failures, delays, and data quality issues in production • Investigate distributed data systems, including Spark/EMR workloads, ingestion pipelines, and warehouse performance • Lead or support incident response through triage, mitigation, and long-term resolution • Perform root cause analysis and implement durable fixes • Define and improve data SLAs for freshness, latency, and completeness • Design and enhance monitoring, alerting, and observability for data systems • Develop automation and tooling to reduce operational toil and improve resilience • Contribute to disaster recovery and resiliency planning, including backup validation and recovery workflows • Partner with engineering teams to improve pipeline design, reliability, and operational readiness • Create and maintain runbooks, SOPs, and operational documentation • Participate in occasional off-hours support and on-call rotation for production data systems

🎯 Requirements

• Minimum 5 years of experience working with production data platforms in AWS environments • Prior experience building data pipelines through production, including exposure to real-world failures and operational challenges • Strong experience with Python and SQL in real data systems • Hands-on experience troubleshooting distributed data processing systems, such as Spark/EMR, Redshift, and streaming systems • Proven ability to debug and resolve production issues in data pipelines and data platforms • Experience with AWS data services such as EMR, Redshift, DynamoDB, S3, or similar • Experience handling production incidents and performing root cause analysis • Strong problem-solving mindset and ability to work through ambiguous production issues • Must be authorized to work for any employer in the U.S.; employment visa sponsorship is unavailable, including CPT/OPT • Reliable high-speed internet with a wired connection and suitable home workspace for remote work • Fiber, cable, or DSL internet connection required for remote work • Ability to participate in an on-call rotation and occasional after-hours change windows

🏖️ Benefits

• Flexible work environment • Medical, dental, vision and life insurance • Retirement savings – 401(k) plan with generous company matching contributions (up to 6%), financial advisory services, potential company discretionary contribution, and a broad investment lineup • Tuition reimbursement up to $5,250/year • Business-casual environment with the option to wear jeans • Generous paid time off upon hire, including a paid time off program, ten paid company holidays and three floating holidays each calendar year • Paid volunteer time — 16 hours per calendar year • Leave of absence programs, including paid parental leave, paid short- and long-term disability, and Family and Medical Leave (FMLA) • Business Resource Groups (BRGs) open to all • Bonus program opportunity for non-sales positions • Necessary computer equipment provided • Remote-work high-speed internet and home-workspace requirements

Apply Now

Similar Jobs

🕒 August 1

Scientific Games

10,000+ employees

🎮 Gaming

🤝 B2B

Senior DevOps Engineer enhancing infrastructure and deployment reliability for Scientific Games. Collaborating across teams to scale secure, high-performing platforms and improve delivery efficiency.

🕒 July 31

Andromeda

11 - 50

🏥 Healthcare

💼 Consulting

🏨 Hospitality

Engineer embedded with teams running large-scale training and inference on GPU clusters in production. Responsible for onboarding, debugging, and improving performance while ensuring reliability.

🕒 July 31

Smithfield Foods

10,000+ employees

🏭 Manufacturing

🌾 Agriculture

🍽️ Food & Beverage

Senior Utilities Engineer at Smithfield Foods optimizing utility systems for industrial refrigeration and ensuring compliance. Collaborating with facilities teams for operational efficiency and system improvements across various locations.

🕒 July 31

Rocket.net

11 - 50

☁️ SaaS

🛍️ eCommerce

🏢 Enterprise

Site Reliability Engineer ensuring high standards for servers, services, and customer environments at Rocket.net. Providing advanced technical support and ensuring platform reliability.

🕒 July 31

Hearst Health

1001 - 5000

🏥 Healthcare

⚕️ Healthcare Insurance

☁️ SaaS

Senior DevOps Engineer II at Bring a Trailer modernizing the infrastructure and security of a trusted automotive marketplace. Developing applications to enable engineering teams to work efficiently.