Senior Data Developer – Databricks

🔥 31 minutes ago

Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of CI&T

CI&T

5001 - 10000 employees

Founded 1995

💼 Consulting

🏥 Healthcare

📣 Marketing

💰 $5.5M Venture Round on 2014-04

Consulting • Healthcare • Marketing

CI&T is a global tech transformation specialist focusing on helping organizations navigate their technology journey. With services spanning from application modernization and cloud solutions to AI-driven data analytics and customer experience, CI&T empowers businesses to accelerate their growth and maximize operational efficiency. The company emphasizes digital product design, strategy consulting, and immersive experiences, ensuring a robust support system for enterprises in various industries.

📋 Description

• We are seeking a Senior Data Developer (Databricks) to join our team and take ownership of a mission-critical data platform supporting business reporting and analytics for our client. • This is an opportunity to work deep in the Databricks ecosystem — from raw ingestion through a fully modeled dimensional layer — while acting as the technical backbone the rest of the team relies on for platform expertise. • This role blends hands-on execution with technical leadership: you will own and evolve notebooks across the full medallion architecture, design and extend dimensional modeling artifacts that power downstream business intelligence, and serve as the primary reviewer and technical reference point for the team. • You'll operate both strategically — proposing architectural and process improvements — and operationally — troubleshooting production jobs, deploying across environments, and keeping documentation current. • Delivery & Continuous Improvement: Work on assigned tickets and bugs while continuously looking for improvement opportunities beyond the immediate task — proactively identifying tech debt, refactoring opportunities, and automation gaps rather than limiting contributions to what's assigned. • Data Pipeline Ownership: Own and evolve notebooks across the Stage, Bronze, Silver, and Gold layers, from ingestion through the dimensional model consumed by business intelligence reporting tools. • Dimensional Modeling: Design and implement dimensional modeling artifacts — facts, dimensions, and slowly changing dimensions — with a clear understanding of how they enable downstream reporting. • Workflow Management: Make changes to data processing jobs and workflows as needed to support evolving business requirements. • Code Review & Quality: Review pull requests from other developers, enforcing code quality, performance, and architectural consistency across the codebase. • Environment & Deployment Management: Deploy and promote changes across environments (Dev, QA, UAT, PROD), keeping deployment tracking up to date, and support environment operations such as restoring environments or tables from another environment or from a specific point in time. • Production Monitoring & Troubleshooting: Monitor and troubleshoot daily production jobs, investigating failures and performance issues using platform-native diagnostic tools, job logs, and table history. • Testing & Automation: Maintain and improve the automated testing pipeline, including CI/CD workflows and the underlying test framework. • Documentation: Keep technical documentation current so institutional knowledge is not lost as the pipelines and connections the team relies on evolve. • Technical Reference & Mentoring: Act as the go-to technical reference for the team on the data platform — the person others turn to when something needs deep platform expertise — and support other team members on data modeling and development topics. • Stakeholder Collaboration: Propose and recommend architectural and process improvements, collaborating with the client's business and technical stakeholders — including the client's data architecture function — to translate requirements into scalable, well-tested data pipelines, while remaining equally comfortable taking direction from client-side technical leadership.

🎯 Requirements

• Solid experience in data development, with proven hands-on production experience on the Databricks platform • Strong proficiency in PySpark (DataFrame API, Spark SQL, UDFs, window functions) and Databricks SQL (ANSI SQL, MERGE INTO, COPY INTO, CTEs), including performance tuning such as partition pruning, file compaction, skew handling, and query optimization • Solid, practical experience with Delta Lake: MERGE/upsert patterns, ACID transactions, time travel, and table maintenance (OPTIMIZE, VACUUM, ZORDER, liquid clustering, Change Data Feed) • Demonstrated experience implementing Slowly Changing Dimensions (Type 1 and Type 2) and dimensional modeling concepts (star schema, fact/dimension design) — not requiring you to have designed a model from scratch, but requiring the mindset to understand and extend one • Experience with medallion (or comparable layered) architecture in a production data platform, and with Unity Catalog, jobs/workflows, secrets management, and notebook-based development • Experience with Git and Azure DevOps (or equivalent) for version control, pull requests, and CI/CD pipelines, along with Microsoft Azure services (Key Vault, Service Principal/Managed Identity, Data Lake Storage) • Ability to read and navigate a large, established codebase (400+ notebooks), learning and following existing conventions rather than rewriting them, and to ramp up quickly in a business-rule-heavy environment • Advanced English (C1 or above) communication skills, with the ability to work directly with US-based client stakeholders, propose technical recommendations, and align with decisions made by client-side technical leadership • Nice to Have • Experience with Databricks Asset Bundles or other Infrastructure-as-Code approaches for managing jobs, clusters, and permissions as code • Familiarity with Delta Live Tables and with Databricks Genie (AI/BI Genie) for natural-language querying and conversational analytics • Familiarity with data quality frameworks (e.g., Great Expectations, Soda Core, or custom validation frameworks) • Experience with pytest and databricks-connect for automated testing of Spark pipelines outside of manual notebook execution • Familiarity with Pydantic or similar typed-configuration approaches, and experience with schema migration/versioning approaches (e.g., Flyway, Liquibase, or custom frameworks) • Comfortable using AI-assisted development tools (e.g., GitHub Copilot, Cursor, or similar) to accelerate coding, debugging, and code review workflows

🏖️ Benefits

• Health and dental insurance • Meal and food allowance • Childcare assistance • Extended paternity leave • Partnership with gyms and health and wellness professionals via Wellhub (Gympass) TotalPass; • Profit Sharing and Results Participation (PLR); • Life insurance • Continuous learning platform (CI&T University); • Discount club • Free online platform dedicated to physical, mental, and overall well-being • Pregnancy and responsible parenting course • Partnerships with online learning platforms • Language learning platform • And many more!

Apply Now

Similar Jobs

🔥 1 hour ago

Leega

201 - 500

💼 Consulting

📣 Marketing

🔌 API

Desenvolvedor Front Flutter Pleno na Leega, projetando e desenvolvendo aplicativos móveis nativos. Colaborando com equipes de desenvolvimento para soluções inovadoras e eficientes.

🗣️🇧🇷🇵🇹 Portuguese Required

Dart

Flutter

Java

Kotlin

Node.js

SQL

Swift

.NET

🔥 1 hour ago

Compass

10,000+ employees

🏠 Real Estate

📱 Media

Swift Developer Mobile (iOS) developing and maintaining iOS applications for clients. Involved in agile environments, implementing improvements, and refactoring e-commerce frontend.

🗣️🇧🇷🇵🇹 Portuguese Required

iOS

Swift

🔥 6 hours ago

UltraCon Consultoria

11 - 50

💼 Consulting

☁️ SaaS

📚 Education

Desenvolvedor(a) Oracle Retail XStore Sênior para desenvolver e manter soluções de ponto de venda. Necessário 5 anos de experiência com Java e SQL, e espanhol intermediário.

🗣️🇪🇸 Spanish Required

🗣️🇧🇷🇵🇹 Portuguese Required

Java

Oracle

SQL

🔥 6 hours ago

Spassu

1001 - 5000

💼 Consulting

📦 Logistics

📣 Marketing

Senior Developer at Spassu working remote, focusing on software development with Outsystems and Agile methodologies. Leading software architecture and team support for development projects.

🗣️🇧🇷🇵🇹 Portuguese Required

Microservices

SQL

🔥 17 hours ago

Eteg

51 - 200

☁️ SaaS

Full Stack Developer at Eteg Tecnologia Da Informação S/a, integrating in a collaborative squad for end-to-end solutions development. Focusing on systems, APIs, automations, and integrations with bots.

🗣️🇧🇷🇵🇹 Portuguese Required

AWS

Docker

DynamoDB

EC2

MongoDB

Postgres

React

Redis

TypeScript

Webpack