Search Remote Jobs

Data Engineer

Job not on LinkedIn

🔥 11 minutes ago

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Reality Defender (YC W22)

Reality Defender (YC W22)

11 - 50 employees

🔐 Security

📱 Media

AI • Security • Media

Reality Defender is a company that provides multi-model and multimodal platforms for detecting AI-generated content. It offers services to enterprises, governments, and platforms to detect deepfakes and synthetic media across various modalities such as audio, video, images, and text. The solutions are designed to protect against the rapidly growing threat of AI-generated content used for fraud and disinformation. Reality Defender's tools are probabilistic, meaning they do not rely on watermarks or prior authentication to verify authenticity. The company focuses on industries like media, finance, and government to safeguard against AI-generated threats.

📋 Description

• Design, build, and operate large-scale data processing pipelines handling multi-terabyte and streaming datasets, including audio/video transcoding, feature extraction, and preprocessing workflows. • Deploy, scale, and troubleshoot containerized workloads on Kubernetes and AWS in production environments. • Build and maintain distributed data processing jobs using frameworks such as Spark and Ray. • Design and operate workflow orchestration systems (e.g., Airflow) with dependency management, retries, monitoring, and alerting for production pipelines. • Administer and tune enterprise databases, including performance tuning, backup/recovery, access control, and scaling strategies. • Partner with ML engineers and researchers to support training pipelines, model retraining triggers, feature stores, and other MLOps workflows.

🎯 Requirements

• Hands-on experience with Kubernetes and AWS, including deploying, scaling, and troubleshooting containerized workloads in production environments. • Proficiency with high-performance/distributed computing frameworks such as Spark and Ray for processing large-scale datasets. • Experience with workflow orchestration tools such as Airflow (or comparable systems like Dagster, Prefect, or Luigi) to schedule and manage complex data pipelines. • Strong programming skills in Python and SQL; experience with Golang is a plus. • Demonstrated track record building and operating large-scale data processing pipelines, ideally handling multi-terabyte or streaming datasets. • Experience working with audio or video data at scale is a strong plus (e.g., transcoding, feature extraction, or preprocessing pipelines). • Familiarity with common data transformation patterns applied to large datasets (ETL/ELT, batch and stream processing, data validation and quality checks). • Experience designing and maintaining job orchestration systems, including dependency management, retries, monitoring, and alerting for production pipelines. • Bonus: experience orchestrating machine learning workflows (training pipelines, model retraining triggers, feature stores, or MLOps tooling).

🏖️ Benefits

• Healthcare plans with 100% premium coverage for employees and partial coverage available for dependents • Dental and Vision plans with 100% premium coverage for employees and their dependents • Short/Long-term disability and life insurance plans with 100% premium coverage for employees • FSA/HSA and 401k programs • Equity compensation • 20 days of PTO per year • 12 weeks of Parental Leave • Learning and Development budget • Monthly wellness benefits • Annual company-sponsored offsite • Daily in-office lunch through UberEats • Commuter benefits • Remote Fridays • Happy Hours and other local events

Apply Now

Similar Jobs

🔥 3 hours ago

Coinbase

1001 - 5000

₿ Crypto

💸 Finance

💳 Fintech

Senior Software Engineer shaping the API platform connecting client applications to Coinbase's backend services. Owning critical systems that serve API requests at scale to improve performance and developer experience.

🔥 6 hours ago

Spark Eighteen

51 - 200

🤝 B2B

☁️ SaaS

Data Engineer required to build and scale data platforms for AI-driven analytics. Experience with Apache Airflow, Spark, Kafka, and PostgreSQL is essential.

🔥 13 hours ago

Shawmut Services LLC

10,000+ employees

🤝 B2B

🌍 Social Impact

🤝 Non-profit

Data Engineer designing data pipelines for Revolution Field Strategies. Collaborating with leaders on data needs and ensuring data quality through scalable solutions.

🔥 16 hours ago

Spring Venture Group

1001 - 5000

⚕️ Healthcare Insurance

☁️ SaaS

💸 Finance

Senior Data Engineer leading modern data platform and AI initiatives. Building robust data systems leveraging Snowflake and Python for a well-established digital marketing company.

🔥 19 hours ago

Blue Acorn iCi

201 - 500

🛍️ eCommerce

🏢 Enterprise

AEP Data Architect Consultant at Blue Acorn iCi partnering with clients on Adobe Experience Platform data strategies. Leading architectural decisions and optimizing client investments.