Senior Databricks Data Engineer

Job not on LinkedIn

🕒 August 1

🇧🇷 Brazil – Remote

⏰ Full Time

🟠 Senior

🚰 Data Engineer

👻 Ghost score 28%

infoinfo

🗣️🇧🇷🇵🇹 Portuguese Required

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Compass

Compass

10,000+ employees

🏠 Real Estate

📱 Media

Real Estate • Media

Compass is a real-estate-focused content and services site that provides detailed market analysis, buying/selling/renting guides, mortgage and financing information, and home improvement and renovation advice. The site offers resources for homebuyers, sellers, renters, agents, and real estate investors — including articles on appraisals, affordable housing, investment strategies, staging and property maintenance. Compass aims to help users make informed decisions across the housing lifecycle through timely market updates and practical how-to content.

📋 Description

• Rebuild legacy data warehouse pipelines in Databricks (PySpark / Spark SQL) based on specifications produced during reverse engineering; • Implement bronze, silver and gold layers following the medallion architecture template, the project's ingestion patterns (CDC/batch) and the defined reconstruction standard; • Execute reconstruction waves by domain, coexisting with the legacy data warehouse until cutover; • Implement business rules and transformations with automated tests in the pipelines; • Perform reconciliation and parity validation of data between the legacy data warehouse and the Lakehouse; • Optimize pipeline performance and cost (partitioning, OPTIMIZE/Z-ORDER, job sizing); • Contribute to technical documentation of migrated rules and support prioritization of reconstruction waves.

🎯 Requirements

• Minimum 5 years of experience in Data Engineering; • Advanced PySpark and Python: development of large-scale batch pipelines with tests and engineering best practices; • Advanced SQL and data modeling: complex transformations, performance tuning and relational/dimensional modeling in a data warehouse context; • Databricks and Delta Lake: jobs, workflows and medallion (bronze, silver, gold) architecture in production; Delta Live Tables / Lakeflow and Asset Bundles; • Migration / reconstruction of ETL pipelines: translating business rules from legacy tools to Spark (not a lift-and-shift), with CDC/batch ingestion from relational databases; • AWS ecosystem: S3, Glue, EMR, Athena, Lambda, DMS and Step Functions; • Azure Synapse: experience with the legacy environment (SQL pools and pipelines) to support reverse engineering; • Data quality and reconciliation: automated tests, DQ gates and legacy vs. new comparisons for parity acceptance; • Git and CI/CD applied to data pipelines. • Knowledge of DataStage (reading jobs for reverse engineering) — important advantage; • Prior experience in data warehouse-to-lakehouse migration programs — important advantage; • Unity Catalog (permissions, lineage) and data contracts; • Experience in the financial or credit industry; • Databricks certifications (Data Engineer Associate / Professional).

Apply Now

Similar Jobs

🕒 July 31

CI&T

5001 - 10000

💼 Consulting

🏥 Healthcare

📣 Marketing

Data Engineer at CI&T helping businesses with scalable data solutions utilizing AI and cloud technologies.

Airflow

Amazon Redshift

AWS

Azure

BigQuery

Cloud

Docker

ETL

Google Cloud Platform

Kafka

Kubernetes

PySpark

Python

SQL

Tableau

Terraform

🕒 July 31

Sigma Software Group

1001 - 5000

💼 Consulting

🏥 Healthcare

🚘 Automotive

Data Engineer building scalable data solutions for high-impact E-commerce project. Collaborating with cross-functional teams on large-scale data flows and modern cloud technologies.

Apache

AWS

Cloud

ETL

Kafka

Python

Spark

🕒 July 30

Sigma Software Group

1001 - 5000

💼 Consulting

🏥 Healthcare

🚘 Automotive

Data Engineer building scalable data solutions for high-traffic E-commerce platforms in Brazil. Collaborating with cross-functional teams on modern cloud technologies and large-scale data architectures.

Apache

AWS

Cloud

ETL

Kafka

Python

Spark

🕒 July 30

Sonepar

10,000+ employees

🤝 B2B

📦 Logistics

Engenheiro de Dados construindo pipelines escaláveis para a Sonepar, líder em distribuição B2B de materiais elétricos. Desenvolvendo soluções com Python, SQL, Spark e Databricks.

🗣️🇧🇷🇵🇹 Portuguese Required

Airflow

Apache

Azure

Cloud

ETL

Python

Spark

SQL

🕒 July 30

Keep IT Simple

11 - 50

📦 Logistics

🏥 Healthcare

🔒 Cybersecurity

Data Engineer with experience in Oracle Data Integrator working on corporate data integration projects remotely for energy commercialization. Seeking a hands-on professional passionate about technology.

🗣️🇧🇷🇵🇹 Portuguese Required

Airflow

AWS

Azure

Cloud

Docker

ETL

Linux

Oracle

Python

SQL