Senior Data Engineer

🔥 2 minutes ago

🌐 Uruguay, Argentina, +6 more countries – Remote

infoinfo

⏰ Full Time

🟠 Senior

🚰 Data Engineer

👻 Ghost score 11%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Marvik

Marvik

51 - 200 employees

Founded 2018

💼 Consulting

🏥 Healthcare

📦 Logistics

Consulting • Healthcare • Logistics

Marvik is a technology consultancy that designs, builds, and deploys production-ready artificial intelligence solutions for enterprise customers. They offer end-to-end AI services including strategy and opportunity discovery, data engineering, model development (agents, LLMs, generative AI, computer vision, predictive analytics), robotics and automation, and on-demand senior AI talent and leadership (fractional CAIO). Marvik focuses on delivering scalable AI that drives business impact across industries like retail, e-commerce, logistics, fintech, manufacturing, healthcare, energy, and government.

📋 Description

• Build and operate scalable ingestion, ELT/ETL, and orchestration pipelines, including batch and real-time streaming, within Microsoft Fabric and cloud lakehouse environments • Design and implement low-latency, real-time data ingestion flows for live operational analytics and streaming workloads • Implement layered Bronze/Silver/Gold medallion architectures using PySpark and SQL • Develop idempotent, backfillable, and incrementally loaded jobs • Apply deduplication, normalization, schema validation, and lineage tracking • Deliver feature-ready, curated datasets for business intelligence, analytics, vector search, and AI/ML agentic workloads • Establish testing, monitoring, and pipeline observability for freshness, volume, and schema drift, with clear alerting • Use Claude Code, Copilot, and Cursor to accelerate pipeline development, query tuning, and data transformation scripting

🎯 Requirements

• 5+ years of hands-on data engineering experience building and operating production data pipelines at scale • Strong proficiency in Python, SQL, and PySpark / Apache Spark • Solid software engineering fundamentals, including Git, CI/CD, and unit/integration testing • Hands-on experience implementing real-time data ingestion and streaming pipelines • Proven experience in end-to-end data modeling, schema design, and layered lakehouse architectures (Medallion architecture) • Experience with cloud-native lakehouse platforms; hands-on experience or familiarity with Microsoft Fabric is highly preferred • Strong grasp of data testing frameworks, pipeline monitoring, and data quality enforcement • Active experience leveraging AI-assisted development tools such as Cursor, Copilot, and Claude • Hands-on experience with Microsoft Fabric, Fabric Lakehouse, Data Factory, or Synapse Analytics is a plus • Experience extracting data from MongoDB / MongoDB Atlas and Change Streams / CDC is a plus • Experience with Event Hubs, Kafka, or Spark Structured Streaming is a plus • Exposure to vector embeddings, RAG-ready datasets, or feature stores for AI/ML workloads is a plus • AEC / Construction / MEP domain experience is a plus

🏖️ Benefits

• Fully remote work arrangement • Employment opportunity framework under Law 19.691 on the Promotion of Employment for Persons with Disabilities, including individuals registered in the National Registry of Persons with Disabilities of the Ministry of Social Development

Apply Now

Similar Jobs

🕒 September 17

Blend360

501 - 1000

🏥 Healthcare

🏨 Hospitality

✈️ Travel

Lead Backend Data Engineer building Java distributed systems, concurrent applications and big-data pipelines. Supporting Blend’s AI services through Kubernetes, Spark, Airflow and observability.

Airflow

Apache

Cloud

Distributed Systems

Google Cloud Platform

Grafana

Groovy

Java

Jenkins

Kafka

Kubernetes

Prometheus

Spark

Spring

SQL

Terraform

🕒 September 1

Blend360

501 - 1000

🏥 Healthcare

🏨 Hospitality

✈️ Travel

Lead Backend Data Engineer building scalable Java distributed systems and big-data pipelines. Supporting Blend’s AI services through Kubernetes, observability, and high-performance backend engineering.

Distributed Systems

Grafana

Java

Kafka

Kubernetes

Prometheus

Spring

SQL

🕒 July 27

SunnyData

51 - 200

💼 Consulting

🏥 Healthcare

📦 Logistics

Data Engineer at SunnyData supporting clients in their data journey with scalable data solutions and AI integration. Involves collaborating with teams to provide data insights and solutions.

AWS

Azure

Cloud

Google Cloud Platform

Hadoop

Kafka

Pandas

Scikit-Learn

Spark

SQL

🕒 July 25

Blend360

501 - 1000

🏥 Healthcare

🏨 Hospitality

✈️ Travel

Lead Data Engineer designing ETL pipelines and data transformation for Microsoft Fabric. Collaborating with teams on data quality and mentoring members for effective data integration.

Azure

ERP

ETL

Python

SQL

🕒 July 23

Blend360

501 - 1000

🏥 Healthcare

🏨 Hospitality

✈️ Travel

Lead Data Engineer at Blend, tackling data integration and transformation challenges for AI initiatives. Collaborating across teams to design ETL processes and optimize data quality within Microsoft Fabric.

Azure

ERP

ETL

Python

SQL