Data Engineer – Databricks, DBT, Streaming, ML, AI

Job not on LinkedIn

🕒 July 15

🇧🇷 Brazil – Remote

⏰ Full Time

🟡 Mid-level

🟠 Senior

🚰 Data Engineer

👻 Ghost score 14%

infoinfo

🗣️🇧🇷🇵🇹 Portuguese Required

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Leega

Leega

201 - 500 employees

Founded 2010

💼 Consulting

📣 Marketing

🔌 API

Consulting • Marketing • API

Leega is a leading technology solutions provider in Latin America, specializing in data analytics and cloud solutions. As the first company in the region certified by Google Cloud for Data Analytics, Leega offers a range of services including application development, machine learning, and risk management analytics. The firm partners with major cloud services such as AWS and Microsoft Azure to help businesses enhance their data management and transition effectively to the cloud, ultimately driving digital transformation and innovation.

📋 Description

• Define architectural paths, standards and guidelines for the data platform focused on AI • Design and build mechanisms for ingestion, processing and provisioning of unstructured data (PDFs, audio, images, videos) • Define the data structures for consumption by AI/ML and generative AI models • Automate, improve and propose changes to existing data processes • Perform proofs of concept and practical tests to validate solutions before scaling them • Serve as a technical reference within an enablement team, supporting multidisciplinary teams • Build and evolve scalable data pipelines integrated with external applications

🎯 Requirements

• Strong experience with Databricks and Apache Spark for large-scale distributed processing • Proficiency in Python and SQL applied to data engineering (Scala is a plus) • Experience processing unstructured data (documents/PDFs, audio, images or video) for AI consumption • Experience defining data architecture and providing technical direction (not only execution) • Familiarity with data preparation for AI/ML and generative AI (feature engineering, chunking, embeddings, model consumption) • Experience integrating data with external systems (APIs, SaaS services, streaming) • Experience with cloud computing, preferably GCP • Knowledge of Delta Lake (Lakehouse architecture, data versioning and governance) • Experience modeling and transforming data with DBT • Practical experience with scalable, robust data pipelines (ETL/ELT) • Experience with code versioning (Git) and CI/CD • Experience in agile environments and collaborating with multidisciplinary teams • Hands-on profile with autonomy to propose, test and improve processes

🏖️ Benefits

• Porto Seguro health insurance • Porto Seguro dental plan • Profit Sharing (PLR) • Childcare assistance • Alelo meal and grocery vouchers • Home office allowance • Partnerships with educational institutions • Support for certifications, including cloud certifications • Livelo points • TotalPass • Mindself

Apply Now

Similar Jobs

🕒 July 14

Compass

10,000+ employees

🏠 Real Estate

📱 Media

AWS Data Engineer responsible for developing and evolving data integrations. Collaborating on AWS services like Databricks, Glue ETL, and Lambda for a leading AI-driven company.

🗣️🇧🇷🇵🇹 Portuguese Required

AWS

EC2

ETL

Java

Postgres

Python

Spark

SQL

Terraform

Unity

🕒 July 13

Dadoteca

51 - 200

💼 Consulting

🏥 Healthcare

📣 Marketing

Senior Data Engineer working on data project development using Azure and Databricks solutions. Collaborating with squads to implement technical solutions and optimize processing routines in Python.

🗣️🇧🇷🇵🇹 Portuguese Required

Azure

MongoDB

Python

🕒 July 10

Extractta

201 - 500

💼 Consulting

📣 Marketing

☁️ SaaS

Senior Data Engineer managing strategic data projects to implement scalable solutions. Collaborating across teams to translate business needs into data engineering solutions.

🗣️🇧🇷🇵🇹 Portuguese Required

Airflow

Apache

AWS

Cloud

Kafka

Kubernetes

PySpark

Python

SQL

🕒 July 10

Extractta

201 - 500

💼 Consulting

📣 Marketing

☁️ SaaS

Data Engineer responsible for developing scalable data pipelines and solutions at Extractta. Collaborating with multidisciplinary teams to translate business needs into engineering solutions.

🗣️🇧🇷🇵🇹 Portuguese Required

Airflow

Apache

AWS

Cloud

Kafka

Kubernetes

PySpark

Python

SQL

🕒 July 10

Dadoteca

51 - 200

💼 Consulting

🏥 Healthcare

📣 Marketing

Engenheiro(a) de Dados Sênior atuando em sustentação de projetos de dados com forte conhecimento técnico e perfil proativo. Contribuindo para a evolução de soluções e melhoria contínua dos processos.

🗣️🇧🇷🇵🇹 Portuguese Required

MongoDB

Python