Data Platform Engineer

🔥 20 hours ago

🌐 Romania, Poland – Remote

infoinfo

⏰ Full Time

🟡 Mid-level

🟠 Senior

🚰 Data Engineer

👻 Ghost score 14%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Software Mind

Software Mind

1001 - 5000 employees

Founded 1999

🤖 Artificial Intelligence

☁️ SaaS

📡 Telecommunications

💰 Private Equity Round on 2020-12

Artificial Intelligence • SaaS • Telecommunications

Software Mind is a technology company that specializes in software development and digital transformation services. With a focus on AI and cloud solutions, the company offers a wide range of services including custom software development, mobile app development, and cloud consulting. Software Mind serves various industries such as financial services, telecom, biotech, and media, providing tailored solutions to accelerate digital transformations and business growth globally.

📋 Description

• Build and operate change-data-capture pipelines from PostgreSQL into Azure using Kafka Connect and Debezium • Configure, deploy and scale connectors end to end, including connector setup, task management, offsets, schema history and snapshot strategy • Run pipelines as stateful workloads on Kubernetes (AKS), covering configuration, secrets, networking and resource tuning • Monitor and troubleshoot the platform in production, including connector failures, task rebalances, restarts, throughput, backpressure, message-size limits, retries and recovery • Automate the platform in Python through configuration-driven onboarding, pipeline orchestration, monitoring and alerting, recovery workflows and automated testing • Integrate CDC streams with Azure Event Hubs, ADLS, Azure PostgreSQL, ADF and Databricks • Manage platform infrastructure as code so environments are reproducible and changes are reviewable • Apply data protection requirements to sensitive data flowing through pipelines, including masking, hashing, access control and retention

🎯 Requirements

• Solid commercial experience as a data or platform engineer, with hands-on work on streaming or CDC pipelines rather than batch reporting alone • Practical Kafka knowledge: topics, partitions, offsets, consumer groups and delivery semantics, including at least one Kafka Connect deployment you ran yourself • Strong SQL and PostgreSQL skills, with working knowledge of WAL, logical replication, replication slots and replication lag • Working understanding of CDC concepts: initial snapshots, inserts, updates and deletes, event ordering, at-least-once delivery, and schema evolution • Confident Python for automation and tooling: orchestration, monitoring, recovery scripts, and automated tests • Hands-on experience with Azure data services, for example Event Hubs, ADLS or Azure PostgreSQL • Comfortable working with Kubernetes as a user: deploying workloads, handling configuration and secrets, reading logs, debugging failing pods • Ability to debug a running pipeline from metrics and logs, distinguishing throughput problems from backpressure, retries or genuine connector failure • Production experience with Debezium specifically: snapshot strategies on large tables, schema history recovery, offset loss, and restoring connectors after failure • Experience operating stateful workloads on AKS: StatefulSets, stable worker identity, and resource tuning under load • Infrastructure-as-code and CI/CD for data platform components, such as Terraform or Bicep • Hands-on work with Databricks and ADF at production scale • Experience implementing data protection controls for sensitive data: masking, hashing, access control and retention policies

🏖️ Benefits

• Flexible employment and remote work • International projects with leading global clients • International business trips • Non-corporate atmosphere • Language classes • Internal & external training • Private healthcare and insurance • Multisport card • Well-being initiatives

Apply Now

Similar Jobs

🕒 Yesterday

AllCloud

201 - 500

💼 Consulting

📦 Logistics

🏥 Healthcare

Data Engineer building AWS lakehouse, ETL, and GenAI platforms. Supporting AllCloud’s cloud migration and managed-services customers through hands-on engineering.

🗣️🇮🇱 Hebrew Required

Airflow

Amazon Redshift

AWS

Cloud

ETL

Kafka

Kubernetes

Microservices

Python

Spark

SQL

🕒 5 days ago

Expleo Group

10,000+ employees

💼 Consulting

🎖️ Defense

📦 Logistics

Data Engineer building Microsoft Fabric pipelines, Lakehouses, and curated datasets. Supporting analytics, reporting, and AI use cases for Expleo, a global engineering and consulting provider.

Azure

Cloud

ERP

ETL

PySpark

Python

Spark

SQL

🕒 5 days ago

Expleo Group

10,000+ employees

💼 Consulting

🎖️ Defense

📦 Logistics

Senior Data Engineer building Microsoft Fabric pipelines, Lakehouses, and trusted datasets. Supporting analytics, reporting, and AI for Expleo’s global engineering and consulting services.

Azure

Cloud

ERP

ETL

PySpark

Python

Spark

SQL

🕒 6 days ago

Sedona Digital

51 - 200

💼 Consulting

🏥 Healthcare

📦 Logistics

Senior Data Engineer migrating enterprise Percona MongoDB environments to MongoDB Atlas for Sedona Digital, a Romanian technology scale-up. Automating migrations, tuning performance, and supporting production cutovers.

AWS

Azure

Cloud

DNS

Firewalls

Kubernetes

Linux

MongoDB

Python

Shell Scripting

Terraform

🕒 August 21

MWDN

51 - 200

💼 Consulting

📦 Logistics

🏥 Healthcare

Senior Data Engineer building scalable real-time and batch pipelines for a fast-growing programmatic advertising company. Optimizing auction analytics, yield models, and ad-tech data infrastructure.

Airflow

Amazon Redshift

AWS

BigQuery

Cloud

ETL

Google Cloud Platform

Kafka

Python

Spark

SQL