Senior Databricks Data Engineer

Job not on LinkedIn

🔥 0 minutes ago

🗣️🇧🇷🇵🇹 Portuguese Required

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Compass

Compass

10,000+ employees

🏠 Real Estate

📱 Media

Real Estate • Media

Compass is a real-estate-focused content and services site that provides detailed market analysis, buying/selling/renting guides, mortgage and financing information, and home improvement and renovation advice. The site offers resources for homebuyers, sellers, renters, agents, and real estate investors — including articles on appraisals, affordable housing, investment strategies, staging and property maintenance. Compass aims to help users make informed decisions across the housing lifecycle through timely market updates and practical how-to content.

📋 Description

• Rebuild legacy data warehouse pipelines in Databricks (PySpark / Spark SQL) based on specifications produced during reverse engineering; • Implement bronze, silver and gold layers following the medallion architecture template, the project's ingestion patterns (CDC/batch) and the defined reconstruction standard; • Execute reconstruction waves by domain, coexisting with the legacy data warehouse until cutover; • Implement business rules and transformations with automated tests in the pipelines; • Perform reconciliation and parity validation of data between the legacy data warehouse and the Lakehouse; • Optimize pipeline performance and cost (partitioning, OPTIMIZE/Z-ORDER, job sizing); • Contribute to technical documentation of migrated rules and support prioritization of reconstruction waves.

🎯 Requirements

• Minimum 5 years of experience in Data Engineering; • Advanced PySpark and Python: development of large-scale batch pipelines with tests and engineering best practices; • Advanced SQL and data modeling: complex transformations, performance tuning and relational/dimensional modeling in a data warehouse context; • Databricks and Delta Lake: jobs, workflows and medallion (bronze, silver, gold) architecture in production; Delta Live Tables / Lakeflow and Asset Bundles; • Migration / reconstruction of ETL pipelines: translating business rules from legacy tools to Spark (not a lift-and-shift), with CDC/batch ingestion from relational databases; • AWS ecosystem: S3, Glue, EMR, Athena, Lambda, DMS and Step Functions; • Azure Synapse: experience with the legacy environment (SQL pools and pipelines) to support reverse engineering; • Data quality and reconciliation: automated tests, DQ gates and legacy vs. new comparisons for parity acceptance; • Git and CI/CD applied to data pipelines. • Knowledge of DataStage (reading jobs for reverse engineering) — important advantage; • Prior experience in data warehouse-to-lakehouse migration programs — important advantage; • Unity Catalog (permissions, lineage) and data contracts; • Experience in the financial or credit industry; • Databricks certifications (Data Engineer Associate / Professional).

Apply Now

Similar Jobs

🕒 2 days ago

Cyber Tools and Solutions

11 - 50

💼 Consulting

📣 Marketing

🔒 Cybersecurity

Design and implement scalable data architectures and pipelines on GCP (BigQuery). Ensure data governance, quality, GA4 validation and provide technical guidance to Product, BI and Martech teams.

🗣️🇧🇷🇵🇹 Portuguese Required

BigQuery

Cloud

Google Cloud Platform

Python

SQL

🕒 July 20

Vivo (Telefônica Brasil)

10,000+ employees

💼 Consulting

📦 Logistics

📣 Marketing

Data Architect at Vivo defining data and AI architectures ensuring security and scalability for projects. Collaborating on solutions with a focus on operational and strategic initiatives.

🗣️🇧🇷🇵🇹 Portuguese Required

Airflow

AWS

Azure

Docker

ETL

Google Cloud Platform

Hadoop

Kubernetes

MongoDB

Postgres

Redis

Spark

SQL

Terraform

🕒 June 8

Extractta

201 - 500

💼 Consulting

📣 Marketing

☁️ SaaS

Engenheiro(a) de Dados Pleno na Extractta, desenvolvendo soluções de dados para projetos estratégicos e escaláveis. Atuando em engenharia de dados com foco em pipelines, qualidade e governança.

🗣️🇧🇷🇵🇹 Portuguese Required

Airflow

Apache

AWS

Cloud

Kafka

Kubernetes

PySpark

SQL

🕒 June 5

SysMap Solutions

1001 - 5000

💼 Consulting

📣 Marketing

🤖 Artificial Intelligence

Data Architect developing scalable data models for analytics and business transformation at Triggo.ai. Collaborating with modern data architecture and business requirements.

🗣️🇧🇷🇵🇹 Portuguese Required

BigQuery

Cloud

Google Cloud Platform

SQL

Vault