Senior Data Engineer

Job not on LinkedIn

🔥 13 minutes ago

🗣️🇧🇷🇵🇹 Portuguese Required

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Cadmus Soluções em TI

Cadmus Soluções em TI

1001 - 5000 employees

Founded 1995

💼 Consulting

📣 Marketing

📦 Logistics

Consulting • Marketing • Logistics

Cadmus Soluções em TI is a Brazilian IT services and solutions company that accelerates business transformation by combining human expertise and artificial intelligence. It offers Multi-AI integrations, intelligent squads (multidisciplinary teams), automation of manual processes, digital solutions including custom software development and modernization, and talent-sourcing/capacity services to staff and upskill IT teams. The company operates a Center for Artificial Intelligence and delivers tailored enterprise software, chatbots, assistants and delivery accelerators for clients across large organizations.

📋 Description

• Design, develop and maintain end-to-end batch pipelines (ingestion, transformation, modeling and serving) processing terabytes of data. • Build and optimize PySpark jobs, with attention to partitioning, shuffle, skew, memory usage and execution cost. • Develop, maintain and evolve DAGs in Apache Airflow, ensuring idempotency, failure handling, retries and appropriate SLAs. • Model and implement transformations in dbt, ensuring tests, documentation, lineage and good versioning practices. • Evolve data layers (raw, curated, analytics) following consistent quality standards, contracts and SLAs. • Investigate and resolve incidents in production pipelines, performing root cause analysis and proposing structural improvements. • Implement and maintain data quality, observability and monitoring mechanisms. • Collaborate with analytics and business teams to understand requirements, propose appropriate data models and ensure the reliability of delivered data. • Contribute to architecture decisions, code reviews and the dissemination of best practices within the team.

🎯 Requirements

• Based in São Paulo (city). • Bachelor's degree in Computer Science, Engineering, Mathematics or a related field. • Strong experience (5+ years) in data engineering, with a significant portion in large-scale environments. • Advanced Python skills, with good coding practices, testing and modularization. • Strong production experience with PySpark, including tuning and troubleshooting jobs at scale. • Advanced SQL, with mastery of window functions, CTEs, query optimization and analytical modeling. • Hands-on experience with dbt in medium/large projects (incremental models, tests, macros, exposures). • Solid production experience with Apache Airflow, including development of complex DAGs, custom operators, sensors, dependency management and execution troubleshooting. • Proven experience with AWS data services (S3, Glue, EMR/EMR Serverless, Athena, IAM, Lambda, among others). • Knowledge of open table formats (Iceberg, Delta or Hudi) and their implications for performance and cost. • Ability to discuss and justify architectural trade-offs (cost, latency, complexity, maintainability). • Practical experience applying Data Security and Compliance policies (LGPD/GDPR). • Production experience with Apache Iceberg. • Experience with Infrastructure as Code (Terraform, CDK). • CI/CD applied to data projects (automated tests, DAG deployments, dbt in pipelines). • Knowledge of data contracts, data catalog and governance. • Prior experience with streaming environments (Kinesis, Kafka, Flink) — even though this role focuses on batch. • Experience with DuckDB or equivalent tools for efficient analytical processing. • Contributions to open source projects or published technical content.

Apply Now

Similar Jobs

🕒 Yesterday

Reply

10,000+ employees

🛡️ Insurance

📦 Logistics

📣 Marketing

Engenheiro de Dados construindo e gerenciando pipelines confiáveis para grandes volumes de dados na Reply, multinacional italiana de TI. Integrando fontes heterogêneas e soluções Big Data em AWS para apoiar análises e ciência de dados.

🗣️🇧🇷🇵🇹 Portuguese Required

Amazon Redshift

AWS

ElasticSearch

ETL

Kafka

PySpark

Python

SQL

Terraform

🕒 Yesterday

Localiza&Co

10,000+ employees

🚘 Automotive

📦 Logistics

✈️ Travel

Especialista em engenharia de dados construindo pipelines e soluções analíticas na Localiza&Co, plataforma brasileira de mobilidade sustentável. Trabalhando com AWS, DBT, SQL, Python e Spark.

🗣️🇧🇷🇵🇹 Portuguese Required

Airflow

Apache

AWS

Cloud

ETL

Python

Spark

SQL

🕒 Yesterday

Localiza&Co

10,000+ employees

🚘 Automotive

📦 Logistics

✈️ Travel

Engenheiro de Dados Sênior desenvolvendo pipelines e produtos de dados AWS para a Localiza&Co, plataforma de mobilidade sustentável. Projetando arquiteturas escaláveis, integrações e processamento Spark em tempo real e batch.

🗣️🇧🇷🇵🇹 Portuguese Required

Airflow

Apache

AWS

Cloud

ETL

Kafka

Python

Spark

SQL

🕒 Yesterday

Mycon

51 - 200

💸 Finance

💳 Fintech

👥 B2C

Engenheiro de Dados Sênior construindo pipelines AWS e GCP para o consórcio digital Mycon. Liderando modelagem, FinOps e integrações de IA generativa em produtos de dados.

🗣️🇧🇷🇵🇹 Portuguese Required

Amazon Redshift

AWS

BigQuery

DynamoDB

ETL

Google Cloud Platform

Python

SQL

🕒 Yesterday

Sicredi

10,000+ employees

🛡️ Insurance

📦 Logistics

💼 Consulting

DBRE Oracle sênior administrando bancos de dados críticos do Sicredi, instituição financeira cooperativa brasileira. Implementando automações, infraestrutura como código, monitoramento e alta disponibilidade em ambientes tecnológicos resilientes.

🗣️🇧🇷🇵🇹 Portuguese Required

Cloud

Java

Linux

MongoDB

MySQL

Oracle

Postgres

Python