Data Engineer – Data Platform

🔥 0 minutes ago

🌐 India, Vietnam – Remote

infoinfo

⏰ Full Time

🟡 Mid-level

🟠 Senior

🚰 Data Engineer

👻 Ghost score 10%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Firmable

Firmable

51 - 200 employees

Founded 2023

🤝 B2B

☁️ SaaS

🤖 Artificial Intelligence

B2B • SaaS • Artificial Intelligence

Firmable is an AI-native B2B sales platform that provides company and contact data, prospect list building, automated buying-signal monitoring, and CRM enrichment. It uses LLMs and agentic AI to assemble and refresh data across hundreds of sources, surface high-intent accounts (role changes, funding, search intent, technology adoption, vertical signals), and generate CRM tasks to guide timely outreach. Firmable integrates with major CRMs (HubSpot, Salesforce, Dynamics, Pipedrive), browser extensions, and other tools, and is positioned as an alternative to legacy data providers like ZoomInfo and Apollo, targeting sales leaders, account executives, SDRs, revenue operations, marketing, and recruiters. The product is offered as a SaaS with terms aimed at smaller and mid-market teams (no enterprise-only contracts or auto-renew traps) and is used by 1,300+ businesses.

📋 Description

• Build pipelines that transform billions of raw records into clean, modelled datasets • Develop LLM-in-the-pipeline systems for extraction, enrichment, entity resolution, and semantic validation • Run structured outputs, retries, and human-review fallbacks over millions of records daily • Build labelled evaluation sets, scorers, and regression suites for prompt and model changes • Track precision and recall for each evaluation check • Apply deterministic checks such as dbt tests, data contracts, and SQL where appropriate • Manage token budgets, model routing, and vendor-model drift detection • Log every LLM call with prompt version, model, cost, latency, and decision • Maintain standard pipeline alerting and observability • Own Airflow orchestration, dbt models from staging through mart, Snowflake performance and cost, and AWS infrastructure as code • Implement embeddings and retrieval patterns for company and people entity resolution across 13 markets • Own data reliability, cost, and downstream consumer trust end to end

🎯 Requirements

• 3+ years in data engineering or data infrastructure • Production pipelines owned end to end • Strong Python and SQL skills • Production-grade, performance-aware engineering at very large scale • Experience shipping LLMs inside data pipelines for extraction, enrichment, or validation • Structured LLM outputs and labelled evaluation sets • Experience building or maintaining LLM evaluation or test harnesses • Judgement on when to use deterministic rules versus LLMs • Production dbt and Airflow experience • Cloud warehouse experience; Snowflake preferred • Schema design, query optimisation, and cost management • Daily use of Claude Code, Cursor, or equivalent AI coding tools, with shipped work to demonstrate • Ownership and systems thinking • Familiarity with LLM evaluation frameworks and tracing such as Logfire or OpenTelemetry • Experience with embeddings, vector search, or fuzzy matching for entity resolution at scale • Experience fine-tuning or distilling small models • AWS at scale, including S3, Lambda, Glue, ECS, and RDS • PostgreSQL • Spark or PySpark • Streaming technologies such as Kafka, Kinesis, or Snowpipe Streaming • Data privacy and compliance knowledge, including GDPR, SOC2, or CCPA

🏖️ Benefits

• Competitive base salary • Meaningful equity • No fixed hours • AI coding tools, automated testing, and AI-assisted review by default • Small senior teams with minimal process • Weekly releases moving toward daily • End-to-end ownership of the stack

Apply Now

Similar Jobs

🔥 5 hours ago

Exavalu

201 - 500

💼 Consulting

🏥 Healthcare

🛡️ Insurance

Azure Data Architect designing secure, scalable Azure warehouses, lakes, and pipelines for IT services company Exavalu. Leading ingestion, governance, and big-data processing architecture.

Azure

Cloud

PySpark

Python

SQL

🕒 Yesterday

People10 Technologies Inc.

501 - 1000

🏢 Enterprise

🤝 B2B

🤖 Artificial Intelligence

Data Engineer building scalable cloud data pipelines for People10’s digital transformation solutions. Integrating databases, APIs, SaaS applications, and analytics platforms.

🇮🇳 India – Remote

💰 $1.6M Venture Round - People10 Technologies on 2014-10

⏰ Full Time

🟡 Mid-level

🟠 Senior

🚰 Data Engineer

Airflow

Amazon Redshift

AWS

Azure

BigQuery

Cloud

Docker

ETL

Google Cloud Platform

Python

SQL

🕒 2 days ago

Akamai Technologies

5001 - 10000

🔒 Cybersecurity

Senior Data Engineer II managing Oracle Fusion and FDI data pipelines for Akamai’s intelligent edge platform. Enabling Finance migration, enterprise analytics, and operational intelligence.

Cloud

ETL

Oracle

PySpark

Python

SQL

🕒 3 days ago

Greenlight Planet

1001 - 5000

⚡ Energy

🌍 Social Impact

👥 B2C

SAP Data Engineer building SAP-to-cloud data pipelines for Sun King’s impact-focused energy business. Supporting BI, FP&A, data quality, and AWS cloud development.

Airflow

Amazon Redshift

AWS

Cloud

EC2

ETL

NoSQL

PySpark

Python

Spark

SQL

🕒 5 days ago

Enable Data

51 - 200

💼 Consulting

🏥 Healthcare

🏭 Manufacturing

Senior AWS Data Engineer building enterprise data pipelines and products with Databricks, AWS, and Medallion Architecture. Hands-on coding, optimization, production troubleshooting, and CI/CD implementation.

AWS

PySpark

Python

SQL

Unity