Data Engineer

🕒 July 23

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Reality Defender (YC W22)

Reality Defender (YC W22)

11 - 50 employees

🤖 Artificial Intelligence

🔐 Security

📱 Media

Artificial Intelligence • Security • Media

Reality Defender is a company that provides multi-model and multimodal platforms for detecting AI-generated content. It offers services to enterprises, governments, and platforms to detect deepfakes and synthetic media across various modalities such as audio, video, images, and text. The solutions are designed to protect against the rapidly growing threat of AI-generated content used for fraud and disinformation. Reality Defender's tools are probabilistic, meaning they do not rely on watermarks or prior authentication to verify authenticity. The company focuses on industries like media, finance, and government to safeguard against AI-generated threats.

📋 Description

• Design, build, and operate large-scale data processing pipelines handling multi-terabyte and streaming datasets, including audio/video transcoding, feature extraction, and preprocessing workflows. • Deploy, scale, and troubleshoot containerized workloads on Kubernetes and AWS in production environments. • Build and maintain distributed data processing jobs using frameworks such as Spark and Ray. • Design and operate workflow orchestration systems (e.g., Airflow) with dependency management, retries, monitoring, and alerting for production pipelines. • Administer and tune enterprise databases, including performance tuning, backup/recovery, access control, and scaling strategies. • Partner with ML engineers and researchers to support training pipelines, model retraining triggers, feature stores, and other MLOps workflows.

🎯 Requirements

• Hands-on experience with Kubernetes and AWS, including deploying, scaling, and troubleshooting containerized workloads in production environments. • Proficiency with high-performance/distributed computing frameworks such as Spark and Ray for processing large-scale datasets. • Experience with workflow orchestration tools such as Airflow (or comparable systems like Dagster, Prefect, or Luigi) to schedule and manage complex data pipelines. • Strong programming skills in Python and SQL; experience with Golang is a plus. • Demonstrated track record building and operating large-scale data processing pipelines, ideally handling multi-terabyte or streaming datasets. • Experience working with audio or video data at scale is a strong plus (e.g., transcoding, feature extraction, or preprocessing pipelines). • Familiarity with common data transformation patterns applied to large datasets (ETL/ELT, batch and stream processing, data validation and quality checks). • Experience designing and maintaining job orchestration systems, including dependency management, retries, monitoring, and alerting for production pipelines. • Bonus: experience orchestrating machine learning workflows (training pipelines, model retraining triggers, feature stores, or MLOps tooling).

🏖️ Benefits

• Healthcare plans with 100% premium coverage for employees and partial coverage available for dependents • Dental and Vision plans with 100% premium coverage for employees and their dependents • Short/Long-term disability and life insurance plans with 100% premium coverage for employees • FSA/HSA and 401k programs • Equity compensation • 20 days of PTO per year • 12 weeks of Parental Leave • Learning and Development budget • Monthly wellness benefits • Annual company-sponsored offsite • Daily in-office lunch through UberEats • Commuter benefits • Remote Fridays • Happy Hours and other local events

Apply Now

Similar Jobs

🕒 July 23

Coinbase

1001 - 5000

💼 Consulting

₿ Crypto

💸 Finance

Senior Software Engineer shaping the API platform connecting client applications to Coinbase's backend services. Owning critical systems that serve API requests at scale to improve performance and developer experience.

🇺🇸 United States – Remote

💵 $186.1k - $218.9k / year

💰 $21.4M Post-IPO Equity on 2022-11

⏰ Full Time

🟠 Senior

🚰 Data Engineer

🦅 H1B Visa Sponsor

info

🕒 July 23

Spark Eighteen

51 - 200

💼 Consulting

📣 Marketing

🏥 Healthcare

Data Engineer required to build and scale data platforms for AI-driven analytics. Experience with Apache Airflow, Spark, Kafka, and PostgreSQL is essential.

🇺🇸 United States – Remote

⏰ Full Time

🟡 Mid-level

🟠 Senior

🚰 Data Engineer

🕒 July 23

Blue Acorn iCi

201 - 500

💼 Consulting

📣 Marketing

🏥 Healthcare

AEP Data Architect Consultant at Blue Acorn iCi partnering with clients on Adobe Experience Platform data strategies. Leading architectural decisions and optimizing client investments.

🇺🇸 United States – Remote

💵 $135k - $185k / year

⏰ Full Time

🟠 Senior

🔴 Lead

🚰 Data Engineer

🦅 H1B Visa Sponsor

info

🕒 July 23

Risepoint

1001 - 5000

💼 Consulting

📣 Marketing

📚 Education

Senior Data Engineer on Risepoint's Enterprise Data Platform team, building scalable data products for universities. Supporting analytics and machine learning workflows in a remote capacity.

🇺🇸 United States – Remote

⏰ Full Time

🟠 Senior

🚰 Data Engineer

🕒 July 23

LegitScript

201 - 500

💼 Consulting

🏥 Healthcare

📦 Logistics

Senior Data Science Engineer at LegitScript developing ML models for risk detection and operational insights. Involves data engineering, MLOps, and many collaborations with operational teams.

🇺🇸 United States – Remote

⏰ Full Time

🟠 Senior

🚰 Data Engineer