Data Engineer – Complex Data Pipelines

🔥 5 minutes ago

🇫🇷 France – Remote

⏰ Full Time

🟡 Mid-level

🟠 Senior

🚰 Data Engineer

👻 Ghost score 12%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of MARSS Group

MARSS Group

51 - 200 employees

📦 Logistics

💼 Consulting

🎖️ Defense

Logistics • Consulting • Defense

MARSS Group is a technology company specializing in the development of advanced security and surveillance systems aimed at enhancing national security. Founded in 2005, the company's expertise is rooted in over 15 years of research and collaboration with EU, NATO, and various defense agencies. MARSS's technological innovations include integrated sensor surveillance, artificial intelligence, and open-source intelligence, which are utilized to protect critical infrastructure, naval assets, special forces, and high-profile individuals globally. Their product lineup features systems such as NiDAR™, advancing command and control functionality, and various solutions designed for counter-unmanned aerial systems and situational awareness in different security contexts.

📋 Description

• Design and build data pipelines from scratch, from data ingestion through processing, transformation, storage and consumption • Design ingestion for distributed recording nodes that are offline most of the time, including local buffering, resumable transfer, and reconciliation of late-arriving or out-of-order data on reconnection • Define, together with the ML team, which data is prioritised during short and bandwidth-limited connection windows • Design pipelines capable of handling structured/tabular data, images, video and temporal/time-series data • Work with sensor-generated sequential data and handle clock drift across nodes to ensure trustworthy downstream timing • Develop robust and scalable data processing solutions using Python and SQL • Design data models and storage approaches, including capacity planning and retention for large volumes of image and video data on self-managed storage • Own workflow definitions in the orchestration layer, including ordering, retry, idempotency and backfill behaviour • Build processes for data ingestion, transformation, validation, quality control and traceability • Develop tools supporting data preparation and availability for machine learning and AI applications • Collaborate with Machine Learning Engineers, DevOps and Software Engineers to understand data requirements and provide solutions • Ensure pipelines are reliable, maintainable and scalable as data volumes and use cases increase • Identify data-quality issues, including gaps and duplicates caused by node outages and retries, and define data-quality and pipeline-health monitoring • Define the overall architecture and technical standards for the in-house data platform • Document pipeline architecture, data flows and technical solutions

🎯 Requirements

• Strong professional experience in a closely related data engineering role • Demonstrated experience designing and implementing data pipelines from zero (not cloud-based), including architectural and technical decisions • Good programming skills in Python & SQL • Professional experience working with several different types of data • Experience building pipelines involving at least some of the following: images, video, sensor data, time-series or other sequential data • Experience with systems that must tolerate unreliable or absent network connectivity and recover gracefully, offline-first, store-and-forward, edge collection or similar architectures • Experience running data infrastructure on bare metal or self-managed servers, rather than exclusively on managed cloud services • Good understanding of data ingestion, transformation, storage, validation and data-quality principles • Experience working with large or complex datasets • Strong Linux skills, including comfort with filesystems, storage, services and network troubleshooting • Good software engineering practices, including Git, testing, code review and documentation • Ability to independently investigate technical problems and propose appropriate architecture and solutions • Fluent English, written and spoken.

🏖️ Benefits

• Full-time employment (CDI)

Apply Now

Similar Jobs

🔥 10 hours ago

Free2move

201 - 500

📦 Logistics

🚘 Automotive

🏥 Healthcare

Data Engineer building real-time pricing pipelines and ML infrastructure for Free2move’s global mobility platform. Owning production data products across France, Spain, or Italy.

Airflow

Apache

AWS

Cloud

Docker

Kubernetes

PySpark

Python

SQL

Tableau

Terraform

Unity

🕒 July 29

BeReal.

51 - 200

👥 B2C

📱 Media

Senior Data Engineer constructing and managing data pipelines for BeReal's growth. Collaborating with engineers and product teams to enhance data-driven solutions.

🕒 July 22

Alan

501 - 1000

🏥 Healthcare

🛡️ Insurance

⚕️ Healthcare Insurance

Senior Platform Engineer managing Data Privacy Orchestrator for 1M+ members in healthcare. Building abstractions and tooling for safe data lifecycle and privacy engineering.

🕒 July 16

Shotgun

11 - 50

📣 Marketing

✈️ Travel

Revenue Data & Operations Engineer optimizing data infrastructure at Shotgun. Focused on data pipelines, analytics, and custom app development to support commercial operations and revenue growth.

🗣️🇫🇷 French Required

JavaScript

React

🕒 July 8

Actian

201 - 500

💼 Consulting

🏥 Healthcare

📦 Logistics

Software Architect designing resilient, cost-efficient cloud data platforms for Actian, a data management company. Improving developer tooling, security, observability, scalability, and cloud infrastructure standards.

AWS

Azure

Cloud

Docker

Google Cloud Platform

Kubernetes

Microservices