Search Remote Jobs

Senior Databricks Engineer

Job not on LinkedIn

🔥 0 minutes ago

🇺🇸 United States – Remote

đź’µ $100.8k - $245.5k / year

⏳ Contract/Temporary

đźź  Senior

👷🏻‍♀️ Engineer

đź‘» Ghost score 9%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of LouisianaNOW.Jobs

LouisianaNOW.Jobs

51 - 200 employees

🎯 Recruiter

🏪 Marketplace

Recruitment • Marketplace

LouisianaNOW. Jobs is a regional jobs portal and talent attraction site focused on connecting job seekers to employers and economic opportunities across Louisiana. The site highlights regional communities, industry sectors, current employers who are hiring, career advice articles, and job listings, and appears to be operated or supported by LED FastStart (Louisiana Economic Development). It’s positioned as a marketplace that promotes Louisiana’s growing industries and helps match talent with local openings and employer investment announcements.

đź“‹ Description

• Design, develop, test, and deploy scalable Databricks data pipelines and transformation workflows • Reverse engineer existing AWS-based data processing solutions to identify transformations, business rules, dependencies, orchestration, and integration requirements • Build, enhance, and maintain Bronze, Silver, and Gold data layers • Design and implement Databricks Auto Loader ingestion solutions with schema management, checkpointing, incremental processing, backfill/reprocessing, and production-scale file ingestion • Develop and operate Lakeflow Declarative Pipelines for batch and streaming workloads, including data quality, quarantine/error handling, dependencies, monitoring, and recovery • Develop real-time and batch processing solutions using Databricks, Structured Streaming, and related technologies • Implement transformation logic using Apache Spark, PySpark, Spark SQL, and Delta Lake • Design and optimize Delta Lake solutions using MERGE, Change Data Feed, partitioning/clustering, retention, and optimization strategies • Design reliable ingestion, transformation, enrichment, and delivery of high-volume customer data • Implement Unity Catalog governance structures and promote assets across development, test, and production • Develop and support integrations between Databricks, Adobe, AWS services, and downstream systems • Analyze AWS Glue, S3, Redshift, Redshift Spectrum, and supporting AWS services to determine current behavior and target Databricks implementation • Establish reusable Databricks Workflows, deployment-as-code patterns, frameworks, utilities, and engineering standards • Apply best practices for data quality, performance, scalability, observability, reliability, security, and maintainability • Tune Spark workloads, pipelines, queries, streaming processes, and data structures • Troubleshoot complex pipeline, integration, data quality, performance, and production issues; participate in root-cause analysis and remediation • Develop automated validation and testing approaches for data products • Perform legacy-to-target parity validation and reconcile discrepancies • Evaluate legacy logic to preserve business functionality while avoiding unnecessary technical constraints • Participate in design discussions, peer code reviews, technical reviews, and solution refinement • Follow client development, CI/CD, security, governance, and deployment processes • Produce and maintain technical documentation • Provide hands-on knowledge transfer and mentoring to client engineers • Leverage approved AI-assisted development tools where appropriate • Identify and recommend automation and engineering improvements

🎯 Requirements

• 6+ years of data engineering or software engineering experience, including significant experience designing, developing, and supporting enterprise-scale data platforms • 3+ years of hands-on Databricks experience developing and operating production data engineering solutions • Production experience with Databricks Auto Loader, including incremental cloud object storage ingestion, schema inference and evolution, schema hints, rescued data handling, checkpoint/state management, backfill and reprocessing strategies, and scalable file discovery • Production experience with Lakeflow Declarative Pipelines (formerly Delta Live Tables), including batch and streaming pipelines, data quality expectations, failed record/quarantine handling, streaming tables and materialized views, refresh processing, dependency management, monitoring, and alerting • Production experience implementing Unity Catalog, including catalog/schema/volume design, external locations, storage credentials, grants, row- and column-level access controls, lineage, and asset promotion • Hands-on Delta Lake experience, including MERGE patterns, Change Data Feed, time travel, OPTIMIZE, Z ORDER and/or liquid clustering, VACUUM and retention policies, and partitioning strategies • Hands-on Apache Spark, PySpark, and Spark SQL experience, including complex transformations and production performance tuning • Ability to diagnose and optimize Spark workloads using partition sizing, shuffle optimization, join strategy, skew handling, and Spark UI analysis • Production Structured Streaming experience, including watermarking, late-arriving data, stateful processing, checkpointing, and exactly-once processing considerations • Experience designing and implementing Medallion Architecture and personally building Bronze, Silver, and Gold data layers • Experience with Databricks Workflows and deployment as code, including dependencies, retry/failure handling, alerting, and Git-based environment promotion using Databricks Asset Bundles, Terraform, or comparable automation • Strong understanding of distributed data processing, data optimization, scalability, and production pipeline resiliency • Experience with production-grade data quality, error handling, monitoring, logging, observability, and pipeline recovery • Experience with CI/CD, automated deployment, Git-based source control, branching, pull requests, and peer code reviews • Strong hands-on AWS Glue ETL experience using PySpark and Python • Ability to analyze unfamiliar and undocumented AWS Glue code and determine transformations, data movement, dependencies, and business rules • Strong working knowledge of Amazon S3 • Experience with Amazon Redshift, including data structures, distribution and sort strategies, COPY/UNLOAD patterns, and stored procedures • Strong understanding of AWS IAM and data access patterns • Working knowledge of AWS Glue Crawlers and Glue Data Catalog • Working knowledge of Redshift Spectrum and external S3-backed data access patterns • Working knowledge of Lambda, Step Functions, EventBridge, Athena, CloudWatch, and Secrets Manager • Ability to trace end-to-end AWS data pipelines across multiple services • Experience reverse engineering undocumented legacy data pipelines • Experience identifying pipeline behavior and dependencies without original developers or complete documentation • Experience with data reconciliation and parity validation between legacy and rebuilt pipelines • Ability to distinguish business logic from legacy technical workarounds • Ability to troubleshoot complex data engineering and production issues independently • Experience working in Agile delivery environments and delivering against prioritized product backlogs • Strong communication and collaboration skills across engineering, architecture, product, and business teams • Ability to successfully complete a background investigation for U.S. employment

🏖️ Benefits

• Competitive compensation including profit participation program • Comprehensive medical, dental, and vision benefits • Basic life and accidental death & dismemberment insurance • Matching contributions through 401(k) plan • CGI share purchase plan • Paid accrued vacation leave, ranging from 10 to 20 days per year • 10 paid holidays per year • At least 80 consecutive hours of paid sick/safe leave • Paid parental leave, ranging from 20 to 70 consecutive business days • Bereavement leave, ranging from 1 to 7 days per year • Paid jury duty leave, up to time summoned • Learning opportunities and tuition assistance • Wellness and Well-being programs

Apply Now

Similar Jobs

🔥 2 hours ago

e4health

501 - 1000

đź’Ľ Consulting

⚖️ Legal

📦 Logistics

MEDITECH Expanse engineer extracting healthcare data and transforming it into HL7 formats. Supporting e4health’s healthcare migrations, integrations, validation, and data quality initiatives.

đź•’ Yesterday

Genovice

1 - 10

🏥 Healthcare

đź’Ľ Consulting

📦 Logistics

Senior CSV Validation Engineer validating cloud, edge, instrument, and integrated systems for regulated pharma environments. Building risk-based evidence for GxP compliance, data integrity, and audit readiness.

đź•’ 2 days ago

RevStar

51 - 200

đź’Ľ Consulting

🤝 B2B

🤖 Artificial Intelligence

Databricks Engineer building multi-cloud Lakehouse pipelines and MLOps solutions for RevStar, an official Databricks Partner. Optimizing Spark, governance, CI/CD, and AI/ML deployments for enterprise clients.

đź•’ 2 days ago

DAWAR CONSULTING INC

51 - 200

đź’Ľ Consulting

🏥 Healthcare

📦 Logistics

GenAI Engineer building enterprise GenAI orchestration, routing, and reusable skills. Developing production services with Python or TypeScript, AWS, Bedrock, and AgentCore for a biotechnology client.

đź•’ 4 days ago

Mercor

51 - 200

Audio Engineer editing and quality-checking French speech recordings for TTS and speech AI systems. Preparing high-quality datasets for Mercor’s frontier AI model training projects.

🇺🇸 United States – Remote

đź’µ $50 / hour

🔥 Funding within the last year

đź’° $350M Series C - Mercor on 2025-10

⏳ Contract/Temporary

🟡 Mid-level

đźź  Senior

👷🏻‍♀️ Engineer

🗣️🇫🇷 French Required