Data Reliability Engineer

🕒 July 15

🇺🇸 United States – Remote

⏰ Full Time

🟡 Mid-level

🟠 Senior

⛑ DevOps & Site Reliability Engineer (SRE)

🦅 H1B Visa Sponsor

infoinfo

👻 Ghost score 11%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Vytalize Health

Vytalize Health

201 - 500 employees

🏥 Healthcare

☁️ SaaS

⚕️ Healthcare Insurance

💰 $100M Series C - Vytalize Health on 2023-02

Healthcare • SaaS • Healthcare Insurance

Vytalize Health is a healthcare technology and services company that helps primary care practices and Accountable Care Organizations (ACOs) transition to value-based care. It combines data-driven analytics, virtual and in-home clinical support, and care management services to improve patient outcomes, enable Medicare-approved remote services for chronic conditions, and help practices earn shared savings under value-based contracts. Vytalize partners with independent PCPs, group practices, community health centers and existing ACOs to deliver clinical enablement, practice-tailored workflows, and performance insights.

📋 Description

• Own and continuously improve the reliability of data pipelines across ingestion, transformation, and delivery layers, ensuring data is accurate, complete, and delivered on schedule. • Establish and maintain data reliability standards, including Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Service Level Agreements (SLAs) for both upstream ingestion and downstream data delivery. • Design, implement, and maintain comprehensive monitoring, logging, and observability frameworks for data pipelines, datasets, and data services with clear visibility into freshness, volume, schema changes, and data quality. • Design and implement data quality testing and validation frameworks — establishing test cases, golden datasets, and regression tests to detect quality issues early. • Lead incident response for data reliability issues, including detection, triage, communication, root cause analysis, and post-incident remediation with documented corrective actions. • Drive improvements in pipeline resiliency through retry strategies, backfills, idempotency, schema enforcement, and safe deployment practices.

🎯 Requirements

• Bachelor's degree in Computer Science, Engineering, Information Systems, or related field, or equivalent professional experience. • 5+ years of experience working with data platforms, data pipelines, or distributed data systems in production environments. • Demonstrated experience improving reliability, observability, or operational quality of data systems with measurable SLI/SLO/SLA improvements. • Hands-on experience supporting both data ingestion pipelines and downstream data consumption or delivery patterns. • 1+ years of hands-on experience with machine learning-based monitoring, anomaly detection, or AI-assisted observability tools. • Demonstrated experience with data quality testing, validation frameworks, and quality metrics definition. • Proficiency in Python and SQL, with experience building or supporting production-grade data pipelines. • Strong understanding of modern data architectures, including data lakehouse patterns and multi-layer (bronze/silver/gold) data models. • Experience with cloud-based data platforms (AWS, Databricks, or similar).

🏖️ Benefits

• Health insurance • 401(k) matching • Flexible work arrangements • Professional development opportunities

Apply Now

Similar Jobs

🕒 July 15

11:11 SYSTEMS

201 - 500

💼 Consulting

🏥 Healthcare

📦 Logistics

Infrastructure Deployment Engineer managing deployment projects across global data centers. Leading cross-functional teams to ensure timely and standardized execution of infrastructure projects.

🕒 July 14

Granicus

501 - 1000

🏛️ Government

☁️ SaaS

📋 Compliance

DevOps Engineer II automating cloud infrastructure, CI/CD, monitoring, and reliability for Granicus, a government technology solutions provider. Applying AI, MCP tools, and AIOps to improve engineering operations.

🕒 July 14

Raya

51 - 200

🌍 Social Impact

👥 B2C

📱 Media

DevSecOps Engineer improving AWS/EKS security and driving collaboration between DevOps and engineering teams at Raya. Focused on hardening systems and closing security findings across the platform.

🕒 July 14

Novellia

1 - 10

🏥 Healthcare

💼 Consulting

⚕️ Healthcare Insurance

Site Reliability Engineer at Novellia, a health tech startup. Responsible for establishing reliability foundations for health data products as part of Platform Engineering.

🕒 July 14

Nymbus

201 - 500

💼 Consulting

📣 Marketing

🏦 Banking

Release Engineer managing production deployments and stability on Nymbus platform. Partnering with Release Coordination to ensure successful deployments in a remote-first environment.