Senior Data Engineer

🔥 15 minutes ago

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Firmable

Firmable

51 - 200 employees

Founded 2023

🤝 B2B

☁️ SaaS

🤖 Artificial Intelligence

B2B • SaaS • Artificial Intelligence

Firmable is an AI-native B2B sales platform that provides company and contact data, prospect list building, automated buying-signal monitoring, and CRM enrichment. It uses LLMs and agentic AI to assemble and refresh data across hundreds of sources, surface high-intent accounts (role changes, funding, search intent, technology adoption, vertical signals), and generate CRM tasks to guide timely outreach. Firmable integrates with major CRMs (HubSpot, Salesforce, Dynamics, Pipedrive), browser extensions, and other tools, and is positioned as an alternative to legacy data providers like ZoomInfo and Apollo, targeting sales leaders, account executives, SDRs, revenue operations, marketing, and recruiters. The product is offered as a SaaS with terms aimed at smaller and mid-market teams (no enterprise-only contracts or auto-renew traps) and is used by 1,300+ businesses.

📋 Description

• This role sits at the heart of what makes Firmable's data valuable. • As Senior Data Engineer - Sourcing, you'll own the harder extraction and ETL problems — the sites that fight back, the schemas that drift, the LLM pipelines that need to run as production systems, not parsing scripts. • You'll spend your time deep in Python, Airflow, and LLM extraction pipelines — architecting extractors that survive anti-bot defences and schema changes, building the eval scaffolding that holds LLM-based extraction accountable, and shipping ETL systems where rules, LLMs, and humans each do what they're best at. • You'll treat LLMs as production infrastructure — versioned prompts, eval sets, traces, cost ceilings — not as a way to skip writing a parser. • This role is hands-on engineering with sharp judgement on extraction architecture. • You decide when to write a parser, when to ask an LLM, and when to do both.

🎯 Requirements

• 4+ years building production extraction, collection, or ETL pipelines in business-critical environments • Strong Python expertise — pandas, numpy, production-grade code, performance-aware. • Advanced SQL — complex queries, performance optimisation, comfort across large datasets • Extensive Airflow (or equivalent) experience — end-to-end orchestration, dependency management, recovery patterns in production • Shipped real work with agentic IDEs — Claude Code, Cursor, or equivalent. Not "tried it" — built and merged real extraction systems with it. • Deep, demonstrable expertise integrating LLMs into extraction pipelines — explicit prompts with rubrics, structured outputs, eval sets, prompt versioning. • Sharp judgement on rules vs. LLMs — you reach for a parser when the structure allows, and don't default to an LLM because it feels modern • Strong knowledge of web extraction at scale — anti-bot defences, proxy strategy, JS rendering, schema drift handling • A product mindset — you understand that extraction quality directly impacts customer value.

🏖️ Benefits

• Health insurance • 401(k) matching • Flexible work hours • Paid time off • Remote work options

Apply Now

Similar Jobs

🕒 3 days ago

v4c.ai

51 - 200

💼 Consulting

🏥 Healthcare

🏭 Manufacturing

Data Engineer focusing on large-scale data systems and ETL processes. Working with cross-functional teams on big data projects and transitioning into Databricks & AI/ML.

AWS

Azure

Cloud

ETL

Google Cloud Platform

Hadoop

Informatica

Kafka

PySpark

Python

Scala

Spark

SQL

🕒 3 days ago

CLIQHR Recruitment Services

11 - 50

💼 Consulting

📣 Marketing

🎯 Recruiter

Azure Data Engineer responsible for designing, migrating, and optimizing data workloads on Microsoft Azure. Collaborating with teams to implement secure and scalable data solutions.

Azure

Cloud

SQL

🕒 4 days ago

Everest Technologies, Inc

51 - 200

💼 Consulting

📦 Logistics

📣 Marketing

Data Engineer developing scalable data pipelines using Snowflake for a leading IT solutions provider. Collaborating with teams to integrate and transform data for analytics initiatives.

Cloud

ERP

ETL

SQL

🕒 4 days ago

Everest Technologies, Inc

51 - 200

💼 Consulting

📦 Logistics

📣 Marketing

Data Engineer supporting enterprise AI and Generative AI initiatives by building scalable, secure, and AI-ready data platforms at ETech. Collaborating with Data Engineers, AI Engineers, and Data Scientists.

Azure

Cloud

ERP

ETL

PySpark

Python

Spark

SQL

🕒 4 days ago

ITX Corp.

201 - 500

💼 Consulting

📣 Marketing

📦 Logistics

Senior Software Engineer with AI & Data skills for designing scalable chatbot platforms at ITX. Focusing on data pipelines and multi-cloud solutions to enhance user experiences.

Airflow

Amazon Redshift

Apache

AWS

Azure

BigQuery

Cloud

ETL

Google Cloud Platform

Python

Spark