Senior Data Engineer

Job not on LinkedIn

🕒 6 days ago

🇺🇸 United States – Remote

⏰ Full Time

🟠 Senior

🚰 Data Engineer

👻 Ghost score 12%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Seamless.AI

Seamless.AI

201 - 500 employees

Founded 2018

🤝 B2B

☁️ SaaS

🤖 Artificial Intelligence

💰 $75M Series A on 2021-06

B2B • SaaS • Artificial Intelligence

Seamless. AI is a B2B software company that provides a sales intelligence and lead-generation platform. It aggregates and enriches contact and company data and uses automation and AI-driven matching to help sales and marketing teams find prospects, build lists, and feed CRMs for outreach and pipeline generation.

📋 Description

• Build and maintain PySpark ETL jobs • Support ML-based scoring models predicting contact-data quality and validity • Write and optimize SQL for large silver/gold lake-table joins • Collaborate on scoring-system architecture, including rule-based gating and ML positioning • Write unit and integration tests for data transformations and model inference • Work within a bronze/silver/gold medallion lake architecture • Identify opportunities for AI and agentic workflows in data acquisition, identity resolution, and profile aggregation • Investigate data-quality issues and analyze scoring-rule outcomes • Own architecture and projects independently from analysis and planning through implementation, testing, and production • Trace data lineage and debug execution across distributed AWS services • Document data-flow findings and architecture decisions

🎯 Requirements

• 5+ years of professional data engineering experience • End-to-end project ownership from design through production • Ability to make and defend architecture decisions with minimal oversight • Strong Python skills, including PySpark DataFrames • Deep hands-on Apache Spark/PySpark experience, including partitioning, shuffles, skew, broadcast joins, predicate/partition pruning, caching, and physical query plan analysis • Complex SQL fluency: multi-way joins, window functions, CTEs, and aggregations at billion-row scale • Experience with AWS Glue, EMR, S3, Step Functions, Lambda, EventBridge, and CloudWatch • Experience writing automated data-pipeline tests using pytest or similar • Ability to trace data lineage across bronze/silver/gold pipeline stages • Orchestration experience with Step Functions, Airflow, or similar DAG-style workflows • Experience modeling and deduplicating heterogeneous upstream data into canonical schemas • Ability to debug distributed AWS services using CloudWatch logs • Exploratory data analysis at scale • Understanding of noisy ground-truth proxies and score-quality implications • Knowledge of classification statistics, including precision/recall, error-cost tradeoffs, and calibration • Backtesting and historical validation experience • Root-cause analysis of data-quality issues • Ability to translate ambiguous goals into testable criteria • Data validation discipline involving row counts, cardinality, output diffs, data profiling, and silent failure modes • Experience with automated data-quality checks • Must be authorized to work in the U.S. • Nice to have: classical ML model development, scikit-learn, contact-data quality, identity resolution, marketing/sales enrichment, CI/CD, EMR, blended rule/ML scoring systems, and LLM/agentic workflows

🏖️ Benefits

• Artificial intelligence and cutting-edge technology work • Equal opportunity employer • Visa sponsorship is not included

Apply Now

Similar Jobs

🕒 6 days ago

SouthState Bank

1001 - 5000

🏦 Banking

💸 Finance

💳 Fintech

Senior data engineer building Snowflake and dbt platforms for SouthState Bank. Leading enterprise pipelines, governance, analytics, and AI/ML data infrastructure.

🕒 6 days ago

Olsson

1001 - 5000

🏗️ Construction

Mechanical piping engineer supporting Olsson’s mission-critical data center and secure-facility designs. Leading calculations, documentation, cost estimates, and multidisciplinary project coordination.

🕒 6 days ago

Cardlytics

201 - 500

💼 Consulting

📦 Logistics

📣 Marketing

Senior Principal Data Engineer building Cardlytics’ Databricks-native MLOps and forecasting platform. Architecting scalable pipelines for campaign projections, lift analysis, budget forecasting, and model deployment.

🕒 6 days ago

Bixal

51 - 200

📣 Marketing

📦 Logistics

🏥 Healthcare

DevOps Engineer securing AWS and Databricks data platforms for Bixal, a technology and human-centered design consultancy. Automating infrastructure, compliance, data governance, and AI/ML workloads.

🕒 6 days ago

SAIC

10,000+ employees

☁️ SaaS

📣 Marketing

🏢 Enterprise

Data Architect modernizing federal supply chain data for SAIC, a defense and mission IT integrator. Building cloud platforms, healthcare interfaces, AI forecasting, and inventory analytics.

🇺🇸 United States – Remote

🔥 Funding within the last year

💰 $500M Post-IPO Debt - SAIC on 2025-09

⏰ Full Time

🟠 Senior

🔴 Lead

🚰 Data Engineer