Senior Data Engineer

🕒 April 2

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Machinify

Machinify

1001 - 5000 employees

⚕️ Healthcare Insurance

🤖 Artificial Intelligence

☁️ SaaS

💰 $10M Series A - Machinify on 2018-10

Healthcare Insurance • Artificial Intelligence • SaaS

Machinify is a healthcare-focused AI platform and services company that reshapes healthcare payments and payment integrity. Its AI operating system unifies claims, medical records, contracts, and policies, and uses foundation models and task-specific agents to automate and improve coding, payment accuracy, recoveries, and cost avoidance. Machinify serves health plans (including 18 of the top 20), supports insourced, hybrid, or fully-managed deployments, and emphasizes measurable outcomes — reporting 85+ customers, 270M member lives covered, and $6B+ in annual cost avoidance and recoveries.

📋 Description

• Design and implement robust, production-grade pipelines using Python, Spark SQL, and Airflow to process high-volume file-based datasets (CSV, Parquet, JSON). • Lead efforts to canonicalize raw healthcare data (837 claims, EHR, partner data, flat files) into internal models. • Own the full lifecycle of core pipelines — from file ingestion to validated, queryable datasets — ensuring high reliability and performance. • Onboard new customers by integrating their raw data into internal pipelines and canonical models; collaborate with SMEs, Account Managers, and Product to ensure successful implementation and troubleshooting. • Build resilient, idempotent transformation logic with data quality checks, validation layers, and observability. • Refactor and scale existing pipelines to meet growing data and business needs. • Tune Spark jobs and optimize distributed processing performance. • Implement schema enforcement and versioning aligned with internal data standards. • Collaborate deeply with Data Analysts, Data Scientists, Product Managers, Engineering, Platform, SMEs, and AMs to ensure pipelines meet evolving business needs. • Monitor pipeline health, participate in on-call rotations, and proactively debug and resolve production data flow issues. • Contribute to the evolution of our data platform — driving toward mature patterns in observability, testing, and automation. • Build and enhance streaming pipelines (Kafka, SQS, or similar) where needed to support near-real-time data needs. • Help develop and champion internal best practices around pipeline development and data modeling.

🎯 Requirements

• 6+ years of experience as a Data Engineer (or equivalent), building production-grade pipelines. • Strong expertise in Python, Spark SQL, and Airflow. • Experience processing large-scale file-based datasets (CSV, Parquet, JSON, etc) in production environments. • Experience mapping and standardizing raw external data into canonical models. • Familiarity with AWS (or any cloud), including file storage and distributed compute concepts. • Experience onboarding new customers and integrating external customer data with non-standard formats. • Ability to work across teams, manage priorities, and own complex data workflows with minimal supervision. • Strong written and verbal communication skills — able to explain technical concepts to non-engineering partners. • Comfortable designing pipelines from scratch and improving existing pipelines. • Experience working with large-scale or messy datasets (healthcare, financial, logs, etc.). • Experience building or willingness to learn streaming pipelines using tools such as Kafka or SQS. • Bonus: Familiarity with healthcare data (837, 835, EHR, UB04, claims normalization).

🏖️ Benefits

• Work from anywhere in the US! Machinify is digital-first. • Full Medical/Dental/Vision for employees & their families • Flexible and trusting environment where you’ll feel empowered to do your best work • Unlimited FTO • Competitive salary, equity, 401(k) including employer match

Apply Now

Similar Jobs

🕒 April 1

Collectiv

51 - 200

☁️ SaaS

🏢 Enterprise

🤖 Artificial Intelligence

Data Engineer responsible for designing, implementing, and supporting data platforms using Microsoft Fabric and Databricks. Collaborate with clients to deliver analytics and reporting solutions.

Azure

Cloud

SQL

🕒 April 1

Mark43

201 - 500

🤖 Artificial Intelligence

🏛️ Government

🔒 Cybersecurity

Lead Data Engineer for Mark43 engaged in building scalable data infrastructure and analytics capabilities. Mentoring team members and innovating solutions in a dynamic environment.

Airflow

Apache

AWS

Cloud

MySQL

Python

SQL

Terraform

🕒 April 1

Azure Data Engineer developing scalable data solutions leveraging Azure Data Platform for Eastbanc Technologies. Collaborating with teams to enhance data-driven decision-making and insights.

Python

SQL

🕒 April 1

CRB

1001 - 5000

🤝 B2B

☁️ SaaS

Senior Data Engineer specializing in Microsoft data and reporting tools for innovative solutions in life sciences and food industries. Delivering data initiatives with focus on ETL processes, models, and architectures.

Azure

ERP

ETL

Python

Spark

SQL

SSIS

🕒 April 1

CRB

1001 - 5000

🤝 B2B

☁️ SaaS

Senior Data Engineer at CRB delivering data and business intelligence initiatives. Requires expertise in Microsoft SQL, ETL processes, and mentorship of junior engineers.

Azure

ERP

ETL

Python

Spark

SQL