Senior Software Engineer, Data Acquisition

Job not on LinkedIn

🔥 8 minutes ago

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of People Data Labs

People Data Labs

51 - 200 employees

Founded 2015

🔌 API

🤝 B2B

☁️ SaaS

API • B2B • SaaS

People Data Labs is a data infrastructure company that provides large-scale person and company datasets via developer-friendly APIs and data enrichment services. Their platform offers identity resolution, contact and firmographic enrichment, and data delivery tools used by sales, marketing, recruiting, and risk teams. They sell primarily to businesses through a SaaS/API model and emphasize data quality, coverage, and privacy/compliance capabilities.

📋 Description

• Contribute to the architecture and improvement of our data acquisition and processing platform, increasing reliability, throughput, and observability • Use and develop web crawling technologies to capture and catalog data on the internet • Build, operate, and evolve large-scale distributed systems that collect, process, and deliver data from across the web • Design and develop backend services that manage distributed job orchestration, data pipelines, and large-scale asynchronous workloads • Structure and model captured data, ensuring high quality and consistency across datasets • Continuously improve the speed, scalability, and fault-tolerance of our ingestion systems • Partner with data product and engineering teams to design and implement new data products powered by the data you help collect, and enhance and improve upon existing products • Learn and apply domain-specific knowledge in web crawling and data acquisition, with mentorship from experienced teammates and access to existing systems

🎯 Requirements

• 7+ years of professional experience building or operating backend or infrastructure systems at scale • Solid programming experience in Python, Go, Rust, or similar, including experience with async / await, coroutines, or concurrency frameworks • Strong grasp of software architecture and backend fundamentals; you can reason clearly about concurrency, scalability, and fault tolerance • Solid understanding of browser rendering pipeline, web application architecture (auth, cookies, http request / response) • Familiarity with network architecture and debugging (HTTP, DNS, proxies, packet capture and analysis) • Solid understanding of distributed systems concepts: parallelism, asynchronous programming, backpressure, and message-driven design • Experience designing or maintaining resilient data ingestion, API integration, or ETL systems • Proficiency with Linux / Unix command-line tools and system resource management • Familiarity with message queues, orchestration, and distributed task systems (Kafka, SQS, Airflow, etc.) • Experience evaluating and monitoring data quality, ensuring consistency, completeness, and reliability across releases.

🏖️ Benefits

• Stock • Competitive Salaries • Unlimited paid time off • Medical, dental, & vision insurance • Health, fitness, and office stipends • The permanent ability to work wherever and however you want

Apply Now

Similar Jobs

🔥 9 minutes ago

RTX

10,000+ employees

🚀 Aerospace

Data Engineer at RTX designing and developing data solutions for aerospace and defense capabilities. Collaborating with stakeholders to ensure data quality and support decision-making through analytics.

ETL

Matillion

Python

SQL

Vault

🔥 55 minutes ago

Reality Defender (YC W22)

11 - 50

🔐 Security

📱 Media

Data Engineer building and scaling the infrastructure for Reality Defender's data platform. Collaborating with ML engineers on large-scale datasets and pipelines.

Airflow

AWS

ETL

Kubernetes

Python

Ray

Spark

SQL

Go

🔥 3 hours ago

Coinbase

1001 - 5000

₿ Crypto

💸 Finance

💳 Fintech

Senior Software Engineer shaping the API platform connecting client applications to Coinbase's backend services. Owning critical systems that serve API requests at scale to improve performance and developer experience.

Distributed Systems

Java

Microservices

Go

🔥 6 hours ago

Spark Eighteen

51 - 200

🤝 B2B

☁️ SaaS

Data Engineer required to build and scale data platforms for AI-driven analytics. Experience with Apache Airflow, Spark, Kafka, and PostgreSQL is essential.

Airflow

Apache

AWS

Cloud

Kafka

PostGIS

Postgres

Python

Spark

SQL

🔥 17 hours ago

Spring Venture Group

1001 - 5000

⚕️ Healthcare Insurance

☁️ SaaS

💸 Finance

Senior Data Engineer leading modern data platform and AI initiatives. Building robust data systems leveraging Snowflake and Python for a well-established digital marketing company.

AWS

Azure

Cloud

Google Cloud Platform

Python

SQL