Data Engineer – DataBricks

Job not on LinkedIn

🕒 2 days ago

🇮🇳 India – Remote

⏰ Full Time

🟡 Mid-level

🟠 Senior

🚰 Data Engineer

👻 Ghost score 15%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Palo Alto Labs

Palo Alto Labs

51 - 200 employees

Founded 2025

💼 Consulting

🔒 Cybersecurity

🤖 Artificial Intelligence

Consulting • Cybersecurity • Artificial Intelligence

Palo Alto Labs is a technology managed-services and innovation firm that partners with businesses to deliver next-generation management technology, digital transformation, and operational excellence. The company provides managed cloud services, cybersecurity, SAP and AI capabilities, finance-as-a-service (FaaS), and data transformation and management, while designing prototypes and processes for new services and experiences. Its focus is on empowering enterprise customers and global capability centers to improve performance, security, and go-to-market readiness.

📋 Description

• Design and implement CI/CD pipelines for data pipelines and transformation projects, including dbt on Databricks, SQL, and notebooks • Orchestrate Databricks jobs and workflows end-to-end, including ingestion, transformation, and quality checks • Integrate automated testing into CI/CD, including schema and contract checks for data models and tables • Implement FinOps best practices for cost monitoring and allocation across the EDP • Automate Databricks platform operations and related services, including workspace and cluster provisioning, library and runtime management, and job deployment/configuration • Implement and maintain identity and access management for Databricks and supporting cloud resources • Manage workspace-, cluster-, table-, and view-level access controls, service principals, groups, roles, RBAC, and TBAC models • Provide Terraform modules, Airflow DAG patterns, and Databricks job templates to accelerate project onboarding • Continuously evaluate and improve tooling, pipelines, and platform architecture for reliability, security, and developer productivity • Define the code promotion process to minimize production impacts across domains • Manage end-to-end orchestration using managed Airflow • Contribute to defining and tracking platform SLA, SLO, and SLI metrics • Participate in incident response, including triage and root cause analysis

🎯 Requirements

• 4–7+ years in DevOps, Cloud Engineering, Site Reliability Engineering, or Platform Engineering • At least 2+ years supporting data/analytics platforms • Hands-on experience with Databricks in a production environment, including workspace and cluster management, jobs/workflows, and integrations with orchestration tools • Strong experience with CI/CD pipelines, such as GitHub Actions, GitLab CI, Azure DevOps, or similar • Experience with Git-based workflows • Strong experience with Infrastructure as Code (IaC) and orchestration tools for provisioning and managing Databricks and cloud infrastructure • Experience implementing automated tests and quality gates in CI/CD pipelines • Experience operating production data workloads, including monitoring, logging, performance tuning, and incident response • Scripting skills in Python, Bash, or PowerShell for automation and integration • Practical knowledge of IT infrastructure technologies, cloud computing, cybersecurity, and disaster recovery • Working knowledge of the Azure ecosystem, including designing, building, and optimizing scalable data pipelines in cloud-native environments • Ability to partner effectively with data engineers, analytics engineers, architects, and security teams • Bachelor’s degree in Computer Science, Information Technology, or related field preferred, or equivalent work experience • Experience in CPG, retail, manufacturing, or distribution environments preferred

🏖️ Benefits

• Remote work arrangement

Apply Now

Similar Jobs

🕒 2 days ago

EXL

10,000+ employees

🏥 Healthcare

🛡️ Insurance

📦 Logistics

Data Architect leading enterprise data modernization with Microsoft Fabric, Azure, and Databricks. Designing secure cloud-native platforms, governance frameworks, and large-scale analytics pipelines.

Apache

Azure

Cloud

ETL

PySpark

Python

Spark

SQL

Unity

Vault

🕒 2 days ago

Firmable

51 - 200

🤝 B2B

☁️ SaaS

🤖 Artificial Intelligence

Data Engineer building LLM-powered data pipelines for Firmable’s B2B sales intelligence platform. Owning evaluation, observability, cost, and reliability across billions of records.

Airflow

AWS

Cloud

Kafka

Postgres

PySpark

Python

Spark

SQL

🕒 3 days ago

Exavalu

201 - 500

💼 Consulting

🏥 Healthcare

🛡️ Insurance

Azure Data Architect designing secure, scalable Azure warehouses, lakes, and pipelines for IT services company Exavalu. Leading ingestion, governance, and big-data processing architecture.

Azure

Cloud

PySpark

Python

SQL

🕒 3 days ago

Firmable

51 - 200

🤝 B2B

☁️ SaaS

🤖 Artificial Intelligence

Data Engineering Lead building AI-powered quality systems for Firmable’s B2B sales intelligence platform. Leading engineers and enforcing trustworthy data across billions of records in 13 markets.

Airflow

AWS

Cloud

Python

SQL

🕒 3 days ago

Firmable

51 - 200

🤝 B2B

☁️ SaaS

🤖 Artificial Intelligence

Senior Data Engineer architecting Firmable’s AI-native data platform for billions of B2B sales intelligence records. Designing production LLM pipelines, evaluation, observability, and Snowflake architecture.

Airflow

AWS

Cloud

Kafka

PySpark

Python

Spark

SQL