Senior Data Engineer

Job not on LinkedIn

🔥 0 minutes ago

🇺🇸 United States – Remote

💵 $133.9k - $178.5k / year

⏰ Full Time

🟠 Senior

🚰 Data Engineer

👻 Ghost score 0%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Gravie

Gravie

201 - 500 employees

Founded 2013

🏥 Healthcare

🛡️ Insurance

🤝 B2B

💰 $150M Private Equity Round - Gravie on 2025-05

Healthcare • Insurance • B2B

Gravie is a benefits and health insurance solutions company that helps small and midsize employers, brokers, and employees access and administer more flexible, affordable health benefits. Its flagship offerings include Comfort (a level-funded health plan), Gravie ICHRA (Individual Coverage Health Reimbursement Arrangement administration), Gravie Pay (support for healthcare costs), and Gravie Care; the company focuses on enabling SMBs to offer individualized, lower-cost coverage and simplifying benefits administration.

📋 Description

• Take full-lifecycle ownership of the streaming platform, including architecture, implementation, and production operations • Run the platform in production and own latency/throughput SLOs • Monitor and alert using tools such as Datadog • Handle replays, per-source backfills, connector-failure recovery, dead-letter triage, and tuning for spiky batch-driven claims load • Build and extend streaming pipelines ingesting CDC events from operational databases and SaaS sources into canonical, contract-validated form • Transform and enrich data in Spark Structured Streaming or dbt-on-Spark micro-batch • Perform cross-stream joins correlating events into unified lifecycle entities • Enable secure, governed access to PHI through classification, row/column controls, and access policies such as ABAC • Design idempotent and replayable pipelines • Implement data-quality validation, runbooks, and observability • Provision streaming, processing, storage, and catalog services on AWS as code using CDK • Implement CI/CD for data pipelines and right-size infrastructure for cost against latency SLOs • Partner with upstream producers on source changes and contracts • Partner with downstream consumers on access and data needs • Translate stakeholder requirements into platform deliverables • Demonstrate Gravie’s core competencies of authenticity, curiosity, creativity, empathy, and outcome orientation

🎯 Requirements

• 6+ years building and operating production data systems, including ownership of streaming or event-driven pipelines • Deep production experience with Apache Kafka, including partitioning, consumer groups, consumer-lag and broker-health troubleshooting, exactly-once/idempotent semantics, schema registry, and replay/backfill under load • Strong hands-on Apache Spark experience (PySpark) for streaming and batch transformation in production • AWS-native data engineering across streaming, processing, storage, and catalog services, including services such as MSK, EMR, Glue, S3, and Athena • Experience with AWS CDK or Terraform, CI/CD for data pipelines, and cost awareness • Comfort debugging distributed data pipelines using observability tooling such as Datadog or CloudWatch • Expert-level SQL and Python • Experience building and consuming REST APIs • Hands-on use of AI-assisted and agentic development tools • Understanding of agentic data consumption patterns, including context management, lineage, provenance, permissions, freshness, and low-latency access • Experience with CDC such as Debezium and open table formats such as Iceberg or Delta Lake • Knowledge of schema evolution, partitioning, and table maintenance • Experience with data contracts, schema governance, cataloging, schema registries, compatibility rules, and dead-letter handling • Familiarity with technical metastores such as Glue Data Catalog or Unity Catalog and governance/discovery catalogs such as Atlan, Alation, or Collibra • Degree in Computer Science, Information Systems, or another quantitative field • Comfort with the command line and a Unix-based OS • Health insurance domain knowledge, including HIPAA Protected Health Information (PHI) and governing access to it • Excellent communication skills and demonstrated success driving results through influence and collaboration • Extra credit: Apache Flink or other stateful stream processors; JVM-based languages such as Kotlin or Java; serving data to AI and agentic consumers; previous venture-backed start-up experience • Must be currently authorized to work in the United States • Visa sponsorship requirement is addressed in the application form

🏖️ Benefits

• Alternative medicine coverage • Flexible PTO • Up to 16 weeks paid parental leave • Paid holidays • 401k program • Transportation perks • Education reimbursement • 2 days of paid paw-ternity leave • Standard health and wellness benefits

Apply Now

Similar Jobs

🔥 47 minutes ago

Peraton

10,000+ employees

💼 Consulting

🏥 Healthcare

📦 Logistics

Senior Data Engineer leading Databricks schema, security, governance, and migration work. Integrating federal marine and terrestrial species data for Peraton’s government program.

🔥 47 minutes ago

Peraton

10,000+ employees

💼 Consulting

🏥 Healthcare

📦 Logistics

Senior Cloud Data Architect leading Databricks lakehouse and ETL architecture for Peraton’s federal ESA species-data integration program. Designing schema harmonization, metadata, and data-quality frameworks in hybrid AWS cloud.

🔥 1 hour ago

Oddball

51 - 200

💼 Consulting

📦 Logistics

🎖️ Defense

Data Engineer building AI-ready financial data pipelines for Oddball’s federal SEC solutions. Supporting semantic search, knowledge graphs, and compliant AWS data infrastructure.

🔥 7 hours ago

InductiveHealth Informatics

51 - 200

🏥 Healthcare

💼 Consulting

📦 Logistics

Senior Data Engineer building Power BI reporting and application integrations. Improving SQL-driven analytics for InductiveHealth’s public health SaaS solutions.

🔥 8 hours ago

Endeavor Media

2 - 10

📱 Media

Senior Director leading EndeavorB2B’s enterprise data platform strategy, architecture, and commercialization. Transforming audience data for analytics, AI, and revenue products at a B2B media and events company.