Senior Data Engineer – Data Platform

🔥 12 hours ago

🌐 India, Vietnam – Remote

infoinfo

⏰ Full Time

🟠 Senior

🚰 Data Engineer

👻 Ghost score 10%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Firmable

Firmable

51 - 200 employees

Founded 2023

🤝 B2B

☁️ SaaS

🤖 Artificial Intelligence

B2B • SaaS • Artificial Intelligence

Firmable is an AI-native B2B sales platform that provides company and contact data, prospect list building, automated buying-signal monitoring, and CRM enrichment. It uses LLMs and agentic AI to assemble and refresh data across hundreds of sources, surface high-intent accounts (role changes, funding, search intent, technology adoption, vertical signals), and generate CRM tasks to guide timely outreach. Firmable integrates with major CRMs (HubSpot, Salesforce, Dynamics, Pipedrive), browser extensions, and other tools, and is positioned as an alternative to legacy data providers like ZoomInfo and Apollo, targeting sales leaders, account executives, SDRs, revenue operations, marketing, and recruiters. The product is offered as a SaaS with terms aimed at smaller and mid-market teams (no enterprise-only contracts or auto-renew traps) and is used by 1,300+ businesses.

📋 Description

• Architect the transformation and warehouse layer turning billions of raw records into an accurate B2B dataset • Design production architecture for LLM-based extraction, enrichment, entity resolution, and semantic validation • Define structured outputs, retry mechanisms, human-review fallbacks, and reusable abstractions • Design labelled evaluation sets, scorers, judge calibration, and regression suites • Establish precision and recall targets for quality checks • Determine when to use deterministic rules versus LLM-based checks • Manage token budgets, model routing, vendor drift detection, and cost ceilings • Build observability logging prompt versions, models, costs, latency, and decisions for every LLM call • Design dbt transformation layers and Snowflake performance, clustering, materialisation, and cost strategies • Define Airflow orchestration, recovery, and cost-aware scheduling patterns • Build AWS infrastructure as code • Develop embeddings and retrieval patterns for company and people entity resolution across 13 markets • Provide technical platform guidance in architecture decisions with sourcing, data quality, product, and analytics teams • Write reference implementations for the team to build upon

🎯 Requirements

• 5+ years building production data platforms in business-critical environments, with end-to-end architecture ownership • Experience working with billions of rows • Production experience shipping LLMs inside data pipelines for extraction, enrichment, or validation • Experience with structured outputs, versioned prompts, and labelled eval sets • Experience designing eval harnesses, including scorers, labelled sets, judge calibration, and regression suites • Strong judgment on rules versus LLMs • Expert Python and SQL skills • Expert dbt and Snowflake skills, including warehouse design, query optimisation, clustering, RBAC, and cost management at scale • Extensive production Airflow experience, including orchestration, dependency management, recovery patterns, and cost optimisation • Solid AWS knowledge: S3, Lambda, Glue, ECS, and RDS • Daily use of AI coding tools such as Claude Code, Cursor, or equivalent, with shipped work to show • Architecture judgment and product mindset • Highly valued: Braintrust, Promptfoo, Inspect, Logfire, OpenTelemetry, embeddings, vector search, fuzzy matching, fine-tuning, distilling small models, dbt Cloud, CI/CD, Spark/PySpark, Kafka, Kinesis, Snowpipe Streaming, B2B data, GDPR, SOC2, and CCPA

🏖️ Benefits

• Competitive base salary • Meaningful equity • Flexible work arrangement / fully remote work • No fixed hours • Small senior teams with minimal process • Weekly releases moving toward daily • Equal opportunity workplace

Apply Now

Similar Jobs

🕒 Yesterday

People10 Technologies Inc.

501 - 1000

🏢 Enterprise

🤝 B2B

🤖 Artificial Intelligence

Data Engineer building scalable cloud data pipelines for People10’s digital transformation solutions. Integrating databases, APIs, SaaS applications, and analytics platforms.

🇮🇳 India – Remote

💰 $1.6M Venture Round - People10 Technologies on 2014-10

⏰ Full Time

🟡 Mid-level

🟠 Senior

🚰 Data Engineer

Airflow

Amazon Redshift

AWS

Azure

BigQuery

Cloud

Docker

ETL

Google Cloud Platform

Python

SQL

🕒 2 days ago

Akamai Technologies

5001 - 10000

🔒 Cybersecurity

Senior Data Engineer II managing Oracle Fusion and FDI data pipelines for Akamai’s intelligent edge platform. Enabling Finance migration, enterprise analytics, and operational intelligence.

Cloud

ETL

Oracle

PySpark

Python

SQL

🕒 4 days ago

Enable Data

51 - 200

💼 Consulting

🏥 Healthcare

🏭 Manufacturing

Senior AWS Data Engineer building enterprise data pipelines and products with Databricks, AWS, and Medallion Architecture. Hands-on coding, optimization, production troubleshooting, and CI/CD implementation.

AWS

PySpark

Python

SQL

Unity

🕒 5 days ago

First Advantage

5001 - 10000

👥 HR Tech

☁️ SaaS

🤝 B2B

Senior Salesforce Data Engineer building scalable pipelines, data models, and enterprise integrations. Supporting First Advantage’s background-screening business through Salesforce analytics, governance, and optimization.

BigQuery

Cloud

ETL

Numpy

Oracle

Pandas

Python

SOAP

SQL

Tableau

🕒 August 26

DigiCert

1001 - 5000

🏥 Healthcare

📦 Logistics

💼 Consulting

Senior Data Engineer building end-to-end Databricks data products for DigiCert, which secures digital interactions worldwide. Developing pipelines, APIs, applications, governance, and production operations.

Apache

AWS

Azure

Cloud

Flask

Google Cloud Platform

JavaScript

Kafka

PySpark

Python

React

Spark

SQL

TypeScript

Unity