
51 - 200 employees
Founded 2018
💼 Consulting
🏥 Healthcare
📦 Logistics
Consulting • Healthcare • Logistics
Marvik is a technology consultancy that designs, builds, and deploys production-ready artificial intelligence solutions for enterprise customers. They offer end-to-end AI services including strategy and opportunity discovery, data engineering, model development (agents, LLMs, generative AI, computer vision, predictive analytics), robotics and automation, and on-demand senior AI talent and leadership (fractional CAIO). Marvik focuses on delivering scalable AI that drives business impact across industries like retail, e-commerce, logistics, fintech, manufacturing, healthcare, energy, and government.
🔥 2 minutes ago
Improve your chances of getting an interview by checking your resume score before you apply.

51 - 200 employees
Founded 2018
💼 Consulting
🏥 Healthcare
📦 Logistics
Consulting • Healthcare • Logistics
Marvik is a technology consultancy that designs, builds, and deploys production-ready artificial intelligence solutions for enterprise customers. They offer end-to-end AI services including strategy and opportunity discovery, data engineering, model development (agents, LLMs, generative AI, computer vision, predictive analytics), robotics and automation, and on-demand senior AI talent and leadership (fractional CAIO). Marvik focuses on delivering scalable AI that drives business impact across industries like retail, e-commerce, logistics, fintech, manufacturing, healthcare, energy, and government.
• Build and operate scalable ingestion, ELT/ETL, and orchestration pipelines, including batch and real-time streaming, within Microsoft Fabric and cloud lakehouse environments • Design and implement low-latency, real-time data ingestion flows for live operational analytics and streaming workloads • Implement layered Bronze/Silver/Gold medallion architectures using PySpark and SQL • Develop idempotent, backfillable, and incrementally loaded jobs • Apply deduplication, normalization, schema validation, and lineage tracking • Deliver feature-ready, curated datasets for business intelligence, analytics, vector search, and AI/ML agentic workloads • Establish testing, monitoring, and pipeline observability for freshness, volume, and schema drift, with clear alerting • Use Claude Code, Copilot, and Cursor to accelerate pipeline development, query tuning, and data transformation scripting
• 5+ years of hands-on data engineering experience building and operating production data pipelines at scale • Strong proficiency in Python, SQL, and PySpark / Apache Spark • Solid software engineering fundamentals, including Git, CI/CD, and unit/integration testing • Hands-on experience implementing real-time data ingestion and streaming pipelines • Proven experience in end-to-end data modeling, schema design, and layered lakehouse architectures (Medallion architecture) • Experience with cloud-native lakehouse platforms; hands-on experience or familiarity with Microsoft Fabric is highly preferred • Strong grasp of data testing frameworks, pipeline monitoring, and data quality enforcement • Active experience leveraging AI-assisted development tools such as Cursor, Copilot, and Claude • Hands-on experience with Microsoft Fabric, Fabric Lakehouse, Data Factory, or Synapse Analytics is a plus • Experience extracting data from MongoDB / MongoDB Atlas and Change Streams / CDC is a plus • Experience with Event Hubs, Kafka, or Spark Structured Streaming is a plus • Exposure to vector embeddings, RAG-ready datasets, or feature stores for AI/ML workloads is a plus • AEC / Construction / MEP domain experience is a plus
• Fully remote work arrangement • Employment opportunity framework under Law 19.691 on the Promotion of Employment for Persons with Disabilities, including individuals registered in the National Registry of Persons with Disabilities of the Ministry of Social Development
Apply Now🕒 September 17
Lead Backend Data Engineer building Java distributed systems, concurrent applications and big-data pipelines. Supporting Blend’s AI services through Kubernetes, Spark, Airflow and observability.
Airflow
Apache
Cloud
Distributed Systems
Google Cloud Platform
Grafana
Groovy
Java
Jenkins
Kafka
Kubernetes
Prometheus
Spark
Spring
SQL
Terraform
🕒 September 1
Lead Backend Data Engineer building scalable Java distributed systems and big-data pipelines. Supporting Blend’s AI services through Kubernetes, observability, and high-performance backend engineering.
Distributed Systems
Grafana
Java
Kafka
Kubernetes
Prometheus
Spring
SQL
🕒 July 27
Data Engineer at SunnyData supporting clients in their data journey with scalable data solutions and AI integration. Involves collaborating with teams to provide data insights and solutions.
AWS
Azure
Cloud
Google Cloud Platform
Hadoop
Kafka
Pandas
Scikit-Learn
Spark
SQL
🕒 July 25
Lead Data Engineer designing ETL pipelines and data transformation for Microsoft Fabric. Collaborating with teams on data quality and mentoring members for effective data integration.
Azure
ERP
ETL
Python
SQL
🕒 July 23
Lead Data Engineer at Blend, tackling data integration and transformation challenges for AI initiatives. Collaborating across teams to design ETL processes and optimize data quality within Microsoft Fabric.
Azure
ERP
ETL
Python
SQL