Senior Data Engineer

🕒 April 2

🇺🇸 United States – Remote

💵 $180k - $220k / year

⏰ Full Time

🟠 Senior

🚰 Data Engineer

🦅 H1B Visa Sponsor

infoinfo

👻 Ghost score 45%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Machinify

Machinify

1001 - 5000 employees

🏥 Healthcare

💼 Consulting

📦 Logistics

💰 $10M Series A - Machinify on 2018-10

Healthcare • Consulting • Logistics

Machinify is a healthcare-focused AI platform and services company that reshapes healthcare payments and payment integrity. Its AI operating system unifies claims, medical records, contracts, and policies, and uses foundation models and task-specific agents to automate and improve coding, payment accuracy, recoveries, and cost avoidance. Machinify serves health plans (including 18 of the top 20), supports insourced, hybrid, or fully-managed deployments, and emphasizes measurable outcomes — reporting 85+ customers, 270M member lives covered, and $6B+ in annual cost avoidance and recoveries.

📋 Description

• Transform raw external data into powerful, trusted datasets that drive payment, product, and operational decisions • Design and implement robust, production-grade pipelines using Python, Spark SQL, and Airflow for high-volume CSV, Parquet, and JSON datasets • Canonicalize raw healthcare data, including 837 claims, EHR, partner data, and flat files, into internal models • Own the full lifecycle of core pipelines from file ingestion to validated, queryable datasets • Onboard new customers by integrating raw data into internal pipelines and canonical models • Collaborate with SMEs, Account Managers, Product, Data Analysts, Data Scientists, Engineering, Platform, and customer teams • Build resilient, idempotent transformation logic with data quality checks, validation layers, and observability • Refactor and scale existing pipelines • Tune Spark jobs and optimize distributed processing performance • Implement schema enforcement and versioning aligned with internal data standards • Monitor pipeline health, participate in on-call rotations, and debug production data flow issues • Contribute to observability, testing, automation, and data platform best practices • Build and enhance streaming pipelines using Kafka, SQS, or similar tools • Develop and champion internal pipeline development and data modeling practices

🎯 Requirements

• 6+ years of experience as a Data Engineer (or equivalent), building production-grade pipelines • Strong expertise in Python, Spark SQL, and Airflow • Experience processing large-scale file-based datasets (CSV, Parquet, JSON, etc.) in production environments • Experience mapping and standardizing raw external data into canonical models • Familiarity with AWS or another cloud, including file storage and distributed compute concepts • Experience onboarding new customers and integrating external customer data with non-standard formats • Ability to work across teams, manage priorities, and own complex data workflows with minimal supervision • Strong written and verbal communication skills; able to explain technical concepts to non-engineering partners • Comfortable designing pipelines from scratch and improving existing pipelines • Experience working with large-scale or messy datasets • Experience building or willingness to learn streaming pipelines using tools such as Kafka or SQS • Familiarity with healthcare data (837, 835, EHR, UB04, claims normalization) is a bonus

🏖️ Benefits

• Work from anywhere in the US; Machinify is digital-first • Full Medical/Dental/Vision for employees and their families • Flexible and trusting environment • Unlimited FTO • Equity • 401(k) including employer match • Meaningful equity • Excellent healthcare • Flexible time off • Other benefits and perks

Apply Now

Similar Jobs

🕒 April 2

Lucenia

1 - 10

🤖 Artificial Intelligence

☁️ SaaS

Senior Software Engineer at Lucenia Inc. specializing in data plane capabilities and distributed indexing operations by collaborating with globally distributed teams.

🕒 April 1

MTC Talent Architecture

1 - 10

🤝 B2B

💼 Consulting

👥 HR Tech

Azure Data Engineer developing scalable data solutions leveraging Azure Data Platform for Eastbanc Technologies. Collaborating with teams to enhance data-driven decision-making and insights.

🕒 April 1

Ex Parte

1 - 10

⚖️ Legal

💼 Consulting

🛡️ Insurance

Senior Data Engineer working with AI and big data solutions at Ex Parte. Collaborating with teams to build products and support a distributed data platform.

🕒 April 1

Ex Parte

1 - 10

⚖️ Legal

💼 Consulting

🛡️ Insurance

Lead Data Engineer responsible for AI self-service portal and microservices at Ex Parte. Collaborate on data platform development and deliver complex client solutions.

🕒 April 1

Scalepex

51 - 200

💼 Consulting

🏥 Healthcare

📦 Logistics

AWS Data Engineer building scalable data solutions for utility datasets. Focused on optimizing data pipelines and ensuring data compliance in a remote role.