Senior Data Engineer – AI-Native, Data Layer

Job not on LinkedIn

🔥 0 minutes ago

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Proton.ai

Proton.ai

11 - 50 employees

🤝 B2B

🤖 Artificial Intelligence

☁️ SaaS

💰 $20M Series A on 2022-01

B2B • Artificial Intelligence • SaaS

Proton. ai is an AI-powered CRM platform specifically designed for distribution teams to enhance sales processes, increase efficiency, and provide actionable insights. It centralizes important sales and customer data by integrating with various ERP and eCommerce platforms, enabling distributors to easily access, analyze, and act on information. With features like personalized recommendations, lead tracking, and efficiency-focused tools, Proton helps sales representatives prioritize efforts and anticipate customer needs. It also automates workflows and streamlines processes, allowing teams to focus on customer engagements rather than administrative tasks. Proton aims to foster accountability and performance across sales, marketing, customer service, and management functions within distribution businesses.

📋 Description

• Own the Data Layer end to end: ingestion from file-, event-, and API-based sources; the medallion-style model (raw → refined → curated); and the serving layer that powers the product and the AI brain. • Build and operate the ingestion and transformation pipelines that power the Data Layer, using a modern orchestration framework and cloud data warehouse. • Ingest and reconcile large, messy, real-world data across many source types and shapes — batch files, streaming events, and APIs. • Model data across medallion layers so it's trustworthy, queryable, and stable for downstream teams and the AI. • Help take the Data Layer to the next level — better architecture, better tooling, more scale, more sources — and have a real say in what that looks like. • Operate AI coding agents (Claude Code and similar) at a high level: scope work, structure context, run agents in parallel where it makes sense, and ship reviewed, production-quality output. • Build the systems that make data trustworthy — validation, reconciliation, lineage, backfills, idempotent and incremental loads — so downstream teams and the AI don't inherit silent errors. • Partner with backend, AI, and product engineers (and occasionally customers' IT teams) to define the data contracts they build on.

🎯 Requirements

• 7+ years hands-on as a data engineer with real, demonstrable production ownership — pipelines and data models serving real users at scale. • Strong fundamentals. You understand what your code and your queries are doing and why. You can read a query plan, reason about a slow or expensive pipeline, and debug a data-correctness bug to its root. • Strong programming and SQL skills. You build efficient pipelines, schemas, and queries, and can model data for both transactional and analytical access patterns. • Hands-on orchestration experience, building reliable ingestion/ELT pipelines against messy upstream sources. • Experience with a cloud data warehouse and a major cloud platform. • Experience ingesting from multiple source types: file-based, event/streaming, and API-based. • Solid grasp of data-consistency failure modes — partial loads, late or out-of-order data, idempotency, backfills, schema drift. • Daily, hands-on use of agentic dev tools (Claude Code, Cursor agent mode, Codex, or equivalent) to ship real work. You can talk concretely about how you structure prompts, manage context, parallelize agents, and verify their output. • Ownership and judgment. You take data systems from idea to production and exercise good taste on what to build and what to cut. • Startup mindset and strong communication — pragmatic, fast, biased to ship, and able to explain data decisions to engineers, PMs, and customers in writing. • English at C1 or above.

🏖️ Benefits

• Professional development opportunities

Apply Now

Similar Jobs

🔥 14 minutes ago

PrizePicks

201 - 500

🎮 Gaming

⚽ Sports

Data Engineer building the data infrastructure for PrizePicks' analytics and business decision-making. Entry-level role focused on well-defined pipeline and data tasks under guidance.

🔥 11 hours ago

Spring Venture Group

1001 - 5000

⚕️ Healthcare Insurance

☁️ SaaS

💸 Finance

Senior Data Engineer at Spring Venture Group focused on data pipelines and AI-powered applications. Collaborate with teams to enhance data products and analytics using Snowflake.

🔥 11 hours ago

Blend360

501 - 1000

🤖 Artificial Intelligence

🏢 Enterprise

Lead Data Engineer at BLEND360 focusing on scalable healthcare data solutions. Designing and optimizing data pipelines and mentoring junior engineers in a cloud environment.

🔥 13 hours ago

Accelerant

201 - 500

☁️ SaaS

🤝 B2B

Product Manager enhancing insurance industry's AI-driven Data Platform. Collaborating with teams to deliver innovative, data-driven solutions for Risk Exchange efficiency.

🔥 14 hours ago

Verily

1001 - 5000

🤖 Artificial Intelligence

☁️ SaaS

🤝 B2B

Clinical Informaticist at Verily Health managing complex healthcare data integration. Architecting FHIR-compliant data models and collaborating with teams for analytics and research initiatives.