Principal Software Engineer – Data Platform, Iceberg/Trino

🕒 August 13

🇺🇸 United States – Remote

⏰ Full Time

🔴 Lead

🚰 Data Engineer

🦅 H1B Visa Sponsor

infoinfo

👻 Ghost score 10%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Innovaccer

Innovaccer

1001 - 5000 employees

Founded 2014

🏥 Healthcare

💼 Consulting

📦 Logistics

💰 $150M Series E on 2021-12

Healthcare • Consulting • Logistics

Innovaccer is a healthcare technology company dedicated to transforming how healthcare information is utilized and accessed. They provide a comprehensive Health Cloud that aims to accelerate digital transformation for providers, payers, and life sciences companies by offering unified healthcare data and analytics. Innovaccer leverages artificial intelligence to automate tasks, process medical data, and generate insights to improve clinical and financial outcomes. Their solutions focus on integrating and optimizing healthcare systems, promoting value-based care, improving patient engagement, and reducing the administrative burden on healthcare providers. Innovaccer's platform supports various healthcare stakeholders through advanced data activation, analytics, and streamlined workflows, striving to improve healthcare delivery and patient care globally.

📋 Description

• Own the lakehouse reference architecture, including Iceberg table design, Trino cluster topology, catalog service, Spark transform compute, and object-storage layout • Design on-premise replacements for cloud-managed warehouse capabilities, including change-data-capture streams, scheduled tasks, and write-back paths into operational stores • Run proof-of-concept validation of the catalog and query engine at expected data volumes • Define evidence-based triggers for VM-based versus Kubernetes-native operator placement decisions • Set platform-wide standards for table layout, partitioning, file sizing, and Iceberg maintenance, including compaction, snapshot expiry, and orphan-file cleanup • Lead the SQL dialect strategy for porting existing warehouse workloads to Trino and Spark SQL • Mentor senior engineers across data workstreams and review designs • Partner with platform engineering on storage sizing, resource isolation, and lakehouse capacity planning

🎯 Requirements

• B.E., B.Tech., or M.Sc. degree in Computer Science or a related technical field • 12+ years of industry experience building and operating large-scale data platforms or distributed systems • Deep hands-on expertise with distributed SQL engines, including Trino/Presto or Spark SQL internals, query planning, and performance engineering • Production experience with Apache Iceberg, or Delta Lake/Hudi with willingness to develop deep Iceberg expertise • Knowledge of Iceberg table specifications, merge-on-read versus copy-on-write, and table maintenance at scale • Working knowledge of Iceberg catalog services, including REST catalogs such as Polaris or Nessie, or Hive Metastore • Working knowledge of S3-compatible object storage • Strong understanding of cloud warehouse internals, including Snowflake, BigQuery, or Redshift • Professional software development experience with Java and/or Python • Experience delivering data platforms in on-premise, regulated, or air-gapped environments is a strong plus • Healthcare data experience is a plus

🏖️ Benefits

• 20 days of fixed paid time off per year, in addition to company holidays • Generous parental leave • Monetary incentives and company-wide recognition for impact and dedication • Medical, dental, and vision insurance • 100% company-paid short- and long-term disability insurance • 100% company-paid basic life insurance • Discounted legal aid • Pet insurance

Apply Now

Similar Jobs

🕒 August 13

YPO

201 - 500

🤝 B2B

☁️ SaaS

📣 Marketing

Data Engineer shaping YPO’s internal architecture and member-facing data services. Building pipelines, improving data quality, and supporting analytics across global operations.

🇺🇸 United States – Remote

⏰ Full Time

🟠 Senior

🔴 Lead

🚰 Data Engineer

🕒 August 13

M9 Solutions

51 - 200

💼 Consulting

🎖️ Defense

📦 Logistics

Director leading enterprise integration and data migration for M9 Solutions, an IT services provider to the Federal Government. Owning architecture, migration strategy, sequencing, governance, and stakeholder decisions.

🇺🇸 United States – Remote

💵 $60k - $180k / year

⏰ Full Time

🔴 Lead

🚰 Data Engineer

🕒 August 13

Rackspace Technology

5001 - 10000

💼 Consulting

📦 Logistics

🤖 Artificial Intelligence

Rackspace, a multicloud solutions provider, seeks a Director to grow its Palantir data engineering practice. Leading client delivery, business development, team development, and strategic partner relationships.

🇺🇸 United States – Remote

💵 $185.7k - $272.4k / year

⏰ Full Time

🔴 Lead

🚰 Data Engineer

🕒 August 13

Sparibis

11 - 50

💼 Consulting

🔒 Cybersecurity

🏢 Enterprise

Azure Data Architect developing predictive models and Azure data solutions for Sparibis's HR and defense projects. Leading data science strategy, governance, visualization, and team mentorship.

🇺🇸 United States – Remote

⏰ Full Time

🟠 Senior

🔴 Lead

🚰 Data Engineer

🕒 August 12

TRM Labs

201 - 500

₿ Crypto

📋 Compliance

🤝 B2B

Staff Data Engineer operating StarRocks-backed serving infrastructure for TRM Labs’ AI-powered crime and threat intelligence platforms. Hardening regulated government cloud data pipelines and reducing production incident risk.

🇺🇸 United States – Remote

⏰ Full Time

🔴 Lead

🚰 Data Engineer