Senior Data Engineer – Databricks, DBT

Job not on LinkedIn

🔥 13 hours ago

🇧🇷 Brazil – Remote

⏰ Full Time

🟠 Senior

🚰 Data Engineer

👻 Ghost score 14%

infoinfo

🗣️🇧🇷🇵🇹 Portuguese Required

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Leega

Leega

201 - 500 employees

Founded 2010

💼 Consulting

📣 Marketing

🔌 API

Consulting • Marketing • API

Leega is a leading technology solutions provider in Latin America, specializing in data analytics and cloud solutions. As the first company in the region certified by Google Cloud for Data Analytics, Leega offers a range of services including application development, machine learning, and risk management analytics. The firm partners with major cloud services such as AWS and Microsoft Azure to help businesses enhance their data management and transition effectively to the cloud, ultimately driving digital transformation and innovation.

📋 Description

• Build and evolve the company’s new data platform, based on DBT running on Databricks • Design and implement the DBT platform architecture on Databricks, including Unity Catalog and medallion layers (bronze/silver/gold) • Structure the DBT project using modeling, naming, layering, macro, testing, and documentation standards • Apply dimensional modeling (Kimball, star schema) and/or Data Vault as needed • Build CI/CD pipelines for DBT through GitHub Actions, including validation, builds, testing, and deployment across dev/staging/prod • Automate provisioning and operations through IaC (Terraform) and Databricks Asset Bundles • Develop and maintain Databricks Workflows for pipeline orchestration • Manage reproducible environments with code versioning using Git, GitFlow, or trunk-based development • Optimize the cost and performance of clusters, SQL Warehouses, Photon, partitioning, and incremental models • Work with Delta Lake, Parquet, and distributed data formats • Leverage Apache Spark and PySpark for use cases beyond SQL when necessary • Define data governance and security, including Unity Catalog, access control, and masking of sensitive data • Ensure compliance with Brazil’s LGPD and data privacy best practices • Mentor the team in analytics engineering best practices • Review code (PRs) and ensure the technical quality of deliverables • Work in agile and collaborative environments

🎯 Requirements

• DBT Core: incremental models, snapshots, macros (Jinja), tests, packages, exposures — minimum 4 years • Databricks: Unity Catalog, SQL Warehouses, clusters, Delta Lake, Workflows — minimum 4 years • Apache Spark and distributed data processing • Python and advanced SQL • GitHub Actions for data pipeline CI/CD • Databricks Asset Bundles for packaging and deployment • Git and collaborative workflows (trunk-based development or GitFlow, PRs, code review) • Cloud computing — preferably GCP (Azure or AWS also accepted) • ETL/ELT and data integration tools • Culture of automated testing, version control, and reproducible deployments • dbt Fusion / migration to dbt Core 2.0 — a plus • Terraform and IaC for environment and secret management — a plus • Orchestration with Airflow or Dagster — preferred • Kubernetes and containers (Docker) — preferred • Data observability: dbt docs, Elementary, Monte Carlo, OpenLineage — preferred • Experience with sensitive data masking / LGPD and compliance — preferred • Kimball dimensional modeling / star schema / Data Vault — preferred

🏖️ Benefits

• Porto Seguro health insurance, with the option to add a spouse and children • Porto Seguro dental insurance for employees and dependents • Profit-Sharing and Results Program (PLR) • Childcare assistance • Alelo meal and food vouchers • Home office allowance • Partnerships with educational institutions, offering discounts and incentives for courses and degree programs • Certification incentives, including cloud certifications (GCP, Azure, AWS, and others) • Livelo points • TotalPass, with discounted gym plans for employees and family members • Mindself, with incentives for meditation and mindfulness • 100% remote work • Ongoing professional development

Apply Now

Similar Jobs

🔥 13 hours ago

Experian

10,000+ employees

💼 Consulting

📣 Marketing

📦 Logistics

Lead Data Engineer evolving Experian’s petabyte-scale AWS data platform. Building scalable pipelines and services for next-generation data-driven marketing solutions.

Airflow

Angular

Apache

AWS

Azure

Cloud

Distributed Systems

Google Cloud Platform

GraphQL

Java

Python

Redis

Scala

Spark

TypeScript

🕒 Yesterday

Stefanini Brasil

10,000+ employees

💼 Consulting

🏥 Healthcare

📦 Logistics

Engenheiro de Dados Especialista apoiando exploração, limpeza e análise de dados na Stefanini. Construção de pipelines Python/PySpark e dashboards para decisões baseadas em dados.

🗣️🇧🇷🇵🇹 Portuguese Required

Airflow

Apache

AWS

Azure

Docker

PySpark

Python

Spark

SQL

SSIS

Terraform

🕒 Yesterday

Compass

10,000+ employees

🏠 Real Estate

📱 Media

Data Engineer building Databricks MLOps pipelines for Compass UOL’s AI-driven digital platforms. Managing scalable data pipelines, model lifecycle automation, and Delta Lake organization.

🗣️🇧🇷🇵🇹 Portuguese Required

Airflow

Apache

AWS

BigQuery

Cassandra

Google Cloud Platform

Kafka

MongoDB

NoSQL

Postgres

PySpark

Spark

SQL

Unity

🕒 Yesterday

Rimini Street

1001 - 5000

💼 Consulting

🏥 Healthcare

📦 Logistics

AI Data Engineer building RAG pipelines, vector infrastructure, and knowledge architecture. Powering Rimini Street’s Agentic ERP Platform for enterprise software support.

Airflow

AWS

Cloud

ERP

ETL

GraphQL

Open Source

Pandas

Postgres

Python

ServiceNow

SQL

🕒 2 days ago

tapouts

11 - 50

📚 Education

🧘 Wellness

Senior Data Engineer building scalable cloud data infrastructure for tapouts, a child-wellbeing mission organization. Developing pipelines, dashboards, lakehouse architectures, and ML data workflows.

Airflow

Amazon Redshift

Apache

AWS

Azure

BigQuery

Cloud

ETL

Google Cloud Platform

Kafka

Python

Spark

SQL