Data Engineer, Python, Spark

🔥 0 minutes ago

🇷🇴 Romania – Remote

⏰ Full Time

🟡 Mid-level

🟠 Senior

🔙 Backend Engineer

👻 Ghost score 12%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of principal33

principal33

201 - 500 employees

Founded 2018

principal33 is a fast-growing European IT company helping organizations design and deliver high-quality digital solutions, from software engineering and digital transformation to data, AI, and cloud services.

📋 Description

• Design, build and operate scalable batch, micro-batch and streaming data pipelines using Python, PySpark and Azure Databricks • Integrate internal and external data sources, including APIs, databases, files, event streams and third-party feeds • Develop ingestion solutions for structured, semi-structured and unstructured data • Build web scraping and data acquisition components where APIs are unavailable • Implement pagination, throttling, retries, exponential backoff, checkpointing, schema evolution and recovery patterns • Design idempotent processing, reprocessing, controlled backfills and graceful recovery from partial failures • Develop and maintain batch and streaming patterns for market, weather, fundamental and time-series data • Develop modular, reusable and testable Python and PySpark components • Structure maintainable software projects and write clean, documented code • Build automated tests and incorporate them into delivery pipelines • Conduct code reviews and promote engineering standards • Package reusable functionality as Python modules or wheels • Troubleshoot complex issues across source systems, APIs, processing logic, infrastructure and production environments • Design data models for analysts, traders, reporting solutions and downstream data products • Implement Lakehouse and Medallion architecture across Bronze, Silver and Gold layers • Design solutions for schema evolution, retention, lineage and reproducible processing • Optimise data layouts, partitioning, joins, file sizes, caching and Spark execution plans • Design, schedule and operate workflows using Databricks Workflows and Astronomer • Implement dependency management, parameterisation, environment configuration and controlled promotion • Define runbooks and support diagnosis, recovery and problem management • Monitor pipeline health, freshness, completeness, performance and data quality • Use production-safe release patterns, including controlled rollouts, rollback and validation • Manage production code through Git and pull-request/review practices • Build and maintain CI/CD pipelines using GitHub Actions and/or Azure DevOps • Deploy Databricks jobs, pipelines and application artefacts using Databricks Asset Bundles or equivalent mechanisms • Provision and configure cloud and Databricks resources through Terraform • Apply automated validation, security scanning and testing before production deployment • Contribute reusable pipeline templates, engineering standards and platform automation • Engineer for availability, recoverability, scalability and predictable operations • Optimise Spark workloads and select suitable compute models and cluster configurations • Apply cost-awareness to pipeline design, compute, scheduling, storage and retention • Monitor resource consumption and reduce processing times and cloud costs • Implement automated data validation, schema checks, null checks, referential-integrity controls and business quality rules • Detect and manage schema drift and source-data changes • Use Delta Lake and Unity Catalog for lineage, access control, metadata and governance • Handle secrets and credentials securely using approved mechanisms and managed identities • Maintain technical documentation, metadata and operational information • Collaborate with data governance, architecture, security and platform teams • Work with traders, analysts, data scientists, software engineers, product owners and platform teams to translate requirements into technical solutions • Communicate design decisions, risks, dependencies and trade-offs • Contribute reusable components, templates, documentation and engineering guidelines • Work effectively in a distributed, international and cross-functional environment

🎯 Requirements

• At least 5 years of experience • Ideally, experience in the energy sector and/or trading • Ability to operate fundamental power-price forecasting models for short- to mid-term trading • Ability to work with large amounts of data • Ability to work fully remote in a collaborative environment with interdisciplinary teams • Good communication skills and professional behaviour • Python and PySpark • Azure Databricks • REST APIs, GraphQL, WebSocket and gRPC interfaces • Databases, files, event streams and third-party data feeds • JSON, CSV, Parquet and Delta • Web scraping and data acquisition • Object-oriented and functional design principles • Type hints, interfaces and design patterns • Unit, integration, contract and data-quality testing • Git, GitHub Actions and/or Azure DevOps • Databricks Asset Bundles or equivalent • Terraform • Delta Lake and Unity Catalog • Secure secret-management mechanisms and managed identities

🏖️ Benefits

• Medical insurance: Your health, and your family's, is our top priority. You're fully covered. • Holiday flat in Valencia • Gifts for special occasions, including Easter, Women's Day, Father's Day, and more • Anniversary gifts for 1st, 5th, and 10th work anniversaries • Team-building events • End of year celebrations • Day off on your birthday • Community and social initiatives • Unlimited access to thousands of courses through Udemy • German language courses • Masterclasses with experienced trainers • Investment in courses and training for personal and professional growth

Apply Now

Similar Jobs

🔥 52 minutes ago

Trading 212

201 - 500

💼 Consulting

📣 Marketing

💸 Finance

Senior Backend Engineer building resilient backend services for Trading 212’s mobile trading and investing platform. Owning microservices from design through production operations.

Microservices

🕒 Yesterday

Powerprozesse GmbH

1 - 10

💼 Consulting

📦 Logistics

🤖 Artificial Intelligence

Senior Backend Developer building autonomous, voice, and inbox AI agents for Hausverwalter.ai’s property-management platform. Improving LLM infrastructure, retrieval, evaluations, and agent tooling.

JavaScript

Next.js

React

TypeScript

🕒 3 days ago

AllCloud

201 - 500

💼 Consulting

📦 Logistics

🏥 Healthcare

Backend Developer building Python and AWS generative AI services for AllCloud’s cloud solutions. Designing scalable event-driven systems and integrating AWS Bedrock, LLMs, and RAG workflows.

AWS

Docker

DynamoDB

Kafka

Kubernetes

Microservices

Postgres

Python

Ray

🕒 5 days ago

Smartitory

1 - 10

🏗️ Construction

🏥 Healthcare

🏨 Hospitality

Senior Java developer building secure banking and public-administration systems for Smartitory. Designing REST APIs and microservices while collaborating with testers, analysts, and project teams.

🗣️🇭🇺 Hungarian Required

Angular

AWS

Azure

Cloud

Docker

Java

Kubernetes

React

Spring

Spring Boot

SpringBoot

🕒 August 18

Arista Networks

1001 - 5000

🏢 Enterprise

📡 Telecommunications

Senior C++ Engineer developing SONiC software for Arista Networks’ data-center switches. Designing, testing, debugging, and upstreaming networking features across routing and distributed systems.

Linux

Python

Unix