Cloud Data Architect, AI Experience

🕒 June 12

🇮🇳 India – Remote

⏰ Full Time

🟠 Senior

🔴 Lead

🚰 Data Engineer

👻 Ghost score 32%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of 3Pillar Global

3Pillar Global

1001 - 5000 employees

💼 Consulting

🏥 Healthcare

🛡️ Insurance

💰 Private Equity Round on 2021-10

Consulting • Healthcare • Insurance

3Pillar Global is a modern application strategy, design, and engineering firm that specializes in delivering strategic software development initiatives for various industries. They offer a range of services, including application technology strategy, digital product engineering, data and analytics, and artificial intelligence development. 3Pillar Global focuses on helping organizations transform their bold ideas into breakthrough solutions by leveraging cutting-edge technologies such as generative and multimodal AI. They work with partners and clients across multiple sectors, including healthcare, financial services, insurance, media, and information services, to solve complex technology challenges and deliver high-performing results.

📋 Description

• Architect and own the enterprise AI data platform — the unified, governed layer that ingests, transforms, stores, and serves all data consumed by AI systems across the organisation. • Design multi-domain data models (lakehouse, data mesh, event-driven) that are structured from day one to serve AI workloads: clean lineage, versioned schemas, well-documented contracts, and low-latency serving APIs. • Own the full data stack: real-time streaming (Kafka, Spark Structured Streaming), batch processing (Databricks, PySpark, Delta Lake), cloud storage and compute (AWS, Azure), and data quality /metadata management. • Ensure this platform is the single, authoritative data source for all downstream consumers — conversational AI, dashboard assistants, autonomous agents, ML models, and reporting — eliminating data silos and conflicting truths. • Drive modernisation of legacy pipelines (on-prem ETL, batch DWH) to cloud-native, AI-ready architectures with measurable improvements in cost, latency, and delivery velocity. • Design the semantic layer that sits above raw data — business-aligned ontologies, entity relationships, domain taxonomies, and knowledge graphs — so AI systems understand context, not just tokens. • Build and maintain knowledge graphs (Neo4j or equivalent) that capture relationships between business entities, policies, KPIs, hierarchies, and domain rules — enabling structured reasoning alongside unstructured retrieval. • Define and govern a feature store and semantic data contracts that serve both classical ML models and LLM-based applications from a single, well-versioned, trusted source. • Own metadata management, data lineage, and audit trails across the semantic layer — ensuring every AI system can trace its outputs back to source data with full accountability. • Design and enforce a comprehensive data governance model that governs access for both human users and AI agents — with role-based access control (RBAC), attribute-based policies, and agent-specific permission scopes that prevent privilege escalation.

🎯 Requirements

• 15+ years of hands-on data engineering and architecture experience, with 3–5+ years building production AI/ML and LLM-era data infrastructure. • Proven experience designing enterprise-scale AI data platforms that serve multiple AI consumers — not just one application or pipeline. • Deep expertise in lakehouse and data mesh architectures: Databricks, Delta Lake, PySpark, Kafka, Spark Structured Streaming, cloud-native data services (AWS, Azure). • Hands-on experience with vector stores, semantic models, knowledge graphs, and retrieval infrastructure in production environments. • Working knowledge of LLMOps: model serving pipelines, MLflow, CI/CD for AI, automated evaluation, and production monitoring. • Strong background in data governance, security, and compliance in regulated industries (financial services, payments, cybersecurity, healthcare). • Experience defining data access controls for AI agents and automated systems — not just human users.

🏖️ Benefits

• Health insurance • Flexible work hours • Professional development opportunities

Apply Now

Similar Jobs

🕒 June 11

NPS Prism

201 - 500

💼 Consulting

📣 Marketing

📦 Logistics

Data Engineer II responsible for developing ETL/ELT workflows and managing data lakes for NPS Prism. Collaborating with teams to design data solutions on cloud platforms like Azure and AWS.

AWS

Azure

Cloud

ETL

PySpark

Python

SQL

Tableau

🕒 June 6

Forbes Advisor

201 - 500

🛡️ Insurance

💼 Consulting

✈️ Travel

Data Engineer building and maintaining data pipelines for marketing analytics and support across business teams. Contributing to data ingestion and modelling from various platforms with a focus on Meta Ads.

Airflow

BigQuery

Cloud

ETL

Microservices

Python

SQL

🕒 June 5

Hillenbrand

5001 - 10000

🍽️ Food & Beverage

💼 Consulting

📦 Logistics

Data Engineer responsible for designing, building, and optimizing scalable data pipelines using Databricks. Collaborating with BI and business teams for high-quality data solutions.

ERP

Python

Spark

SQL

🕒 June 5

Hillenbrand

5001 - 10000

🍽️ Food & Beverage

💼 Consulting

📦 Logistics

Data Engineer designing and optimizing scalable data pipelines using Databricks for Mold-Masters. Collaborating with BI and business teams to ensure reliable data solutions.

Azure

ERP

Python

Spark

SQL

🕒 June 4

Anteriad

201 - 500

💼 Consulting

📦 Logistics

📣 Marketing

Data Engineer at Anteriad optimizing data pipelines using Azure services. Partnering with stakeholders to create scalable data engineering solutions.

AWS

Azure

Cloud

ETL

Google Cloud Platform

MS SQL Server

PySpark

Python

Spark

SQL

SSIS

Vault