Senior Data Engineer

Job not on LinkedIn

🔥 23 minutes ago

🗣️🇧🇷🇵🇹 Portuguese Required

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Caju

Caju

501 - 1000 employees

💳 Fintech

👥 HR Tech

☁️ SaaS

💰 $4M Undisclosed on 2025-01

Fintech • HR Tech • SaaS

Caju is a Brazilian payment institution and HR platform that centralizes employee benefits, expense management and HR processes in a single, integrated solution. Authorized by the Banco Central as an institution of payment, Caju offers cards and a mobile app, multibenefits administration (CLT, PAT, flexible), expense and reimbursement workflows, mobility and wellbeing services, and acts as a banking correspondent for payroll lending. The company focuses on B2B customers—serving small to large employers—to simplify HR operations, improve employee engagement and provide financial-services features with regulatory compliance.

📋 Description

• Design and maintain heterogeneous integration pipelines using Apache Airflow, Dagster, Airbyte, AWS DMS, dbt, Kubernetes, REST APIs, SFTP and various connectors for continuous ingestion • Apply advanced partitioning and indexing techniques (Z-Ordering/Liquid Clustering), file layout optimization and read/write strategies on Amazon S3, Databricks and dbt • Eliminate excessive table scans (full table scans), reduce S3 request costs (GET/LIST) and accelerate query consumption • Design and implement real-time, topic- and event-based data ingestion using Apache Kafka, AWS Kinesis and Event Hubs • Reduce direct dependence on queries and bulk loads against relational databases such as PostgreSQL and MySQL • Implement tools, agents and GenAI/LLM-based pipelines to automate engineering operational tasks • Design, build and maintain robust pipelines with dimensional modeling and Medallion architecture • Deliver clean, aggregated and optimized datasets for ingestion and consumption • Ensure compliance (LGPD/GDPR), sensitive data masking, data lineage, granular access control and monitoring of data freshness and data quality • Gather requirements from business areas, document architectures, define data contracts and ensure governance and maintainability of pipelines

🎯 Requirements

• Strong experience in data engineering in high-volume, production-critical environments • Hands-on expertise with partitioning strategies, data layout optimization, compression, Z-Ordering and query optimization • Experience with Airflow, Dagster, Airbyte, AWS DMS, dbt, Kubernetes, REST APIs, SFTP and custom connectors • Experience with event/topic-based ingestion architectures using Kafka, Kinesis, Pulsar or RabbitMQ • Advanced proficiency in Python, PySpark and SQL optimized for large-scale distributed processing • Deep experience with Databricks, Delta Lake, Auto Loader, Structured Streaming, Jobs/Workflows and Medallion Architecture • Experience with AWS data services: S3, IAM, Redshift, DynamoDB, Data Catalog and AWS DMS • Practical knowledge of, or strong interest in, tools/LLMs for automating development and data engineering tasks • Experience with automated testing, Git/GitHub, code reviews (PRs) and CI/CD pipelines • Strong understanding of dimensional data modeling (Star/Snowflake) and security/privacy guidelines (LGPD/GDPR) • Nice to have: experience building pipelines to support AI, chunking, embeddings, vector search, feature stores and Unity Catalog • Nice to have: experience with Unity Catalog, Terraform and Databricks Asset Bundles (DAB) • Nice to have: implementation of data contracts and quality/observability frameworks such as Great Expectations, Soda, dbt tests and Datadog • Nice to have: cloud/Databricks cost management, anomaly alerting and statistical detection of data drift • Nice to have: experience with idempotency, balance reconciliation and financial reprocessing • Nice to have: experience in fintech, payments or healthcare, including compliance and data audits

🏖️ Benefits

• Caju Card, giving you greater flexibility to use your benefits (Meal, Food, Mobility, Health, Home Office, Culture and Education) • Health plan with no copayment (Unimed, Sulamerica or Alice) • Zenklub: online consultations with therapists and coaches to support mental health • Wellhub • Support for language learning through a partnership with Rosetta Stone • Recharge day - extra day off • Conexa Saúde - online medical consultations • Childcare assistance • Partnership with Alura • Remote work — work from anywhere within Brazil • Work equipment provided • Many growth opportunities

Apply Now

Similar Jobs

🔥 18 hours ago

Genial Investimentos

501 - 1000

💼 Consulting

🛡️ Insurance

💸 Finance

Engenheiro de Dados desenvolvendo pipelines batch e streaming para a Genial Investimentos. Integrando bancos, APIs e serviços AWS no mercado financeiro.

🗣️🇧🇷🇵🇹 Portuguese Required

Apache

AWS

PySpark

Python

Spark

SQL

Terraform

Unity

🔥 18 hours ago

Stefanini Brasil

10,000+ employees

💼 Consulting

🏥 Healthcare

📦 Logistics

Engenheiro de Dados Sênior desenvolvendo soluções Hadoop para a Stefanini, empresa global de tecnologia. Criando pipelines escaláveis, protegendo dados e otimizando infraestrutura para análises e ciência de dados.

🗣️🇧🇷🇵🇹 Portuguese Required

Apache

Hadoop

Python

Spark

SQL

🔥 23 hours ago

Combine | Global Recruitment

11 - 50

🎯 Recruiter

🤝 B2B

Data Engineer maintaining Python/Django scrapers, databases, and document-processing pipelines. Supporting civic-tech archives handling court data and millions of digitized records in Brazil.

Django

ElasticSearch

Kubernetes

Postgres

Python

Redis

🕒 Yesterday

GFT Technologies

10,000+ employees

💼 Consulting

🛡️ Insurance

🔒 Cybersecurity

Engenheiro de Dados FICO construindo e operando pipelines robustos para a GFT, empresa global de tecnologia. Garantindo qualidade, governança e disponibilidade de dados para analytics e decisões de negócio.

🗣️🇧🇷🇵🇹 Portuguese Required

Airflow

Cloud

ETL

MongoDB

NoSQL

Python

Spark

SQL

🕒 Yesterday

Leega

201 - 500

💼 Consulting

📣 Marketing

🔌 API

Engenheiro de Dados GCP construindo pipelines e plataformas escaláveis para a consultoria Leega, especializada em Data Analytics, Cloud e IA. Atuando em Build & Run, governança e incidentes de produção.

🗣️🇧🇷🇵🇹 Portuguese Required

Airflow

BigQuery

Cloud

Google Cloud Platform

Python

SQL