Senior Lead AI Engineer, Data

🕒 March 31

🇮🇳 India – Remote

⏰ Full Time

🟠 Senior

🚰 Data Engineer

👻 Ghost score 45%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Coupa Software

Coupa Software

1001 - 5000 employees

Founded 2006

💼 Consulting

📦 Logistics

🏥 Healthcare

Consulting • Logistics • Healthcare

Coupa Software is a leading provider of business spend management solutions. Their platform focuses on optimizing and transforming direct and indirect spend across procurement, finance, supply chain, and IT. Coupa leverages AI and extensive data insights to drive cost efficiencies, manage supplier relationships, and mitigate risks. With products covering areas such as invoicing, payments, expense management, and supply chain collaboration, Coupa serves a wide range of industries including automotive, healthcare, retail, and more. Their comprehensive community and partner ecosystem enable organizations to unlock hidden savings and improve compliance, promoting growth and resilience in a changing economic climate.

📋 Description

• Lead the design and implementation of data pipelines that prepare high-quality training data for AI models. • Build data curation workflows that transform raw enterprise data into labeled, validated datasets. • Design data quality frameworks: validation, profiling, anomaly detection, lineage tracking. • Extend existing anonymized data export pipelines to support AI training workloads. • Implement synthetic data generation pipelines. • Design schema mappings across 197+ enterprise tables for feature extraction. • Collaborate with ML engineers on training data format requirements. • Establish data catalog and metadata management for AI training artifacts.

🎯 Requirements

• 10+ years of software engineering experience, with 5+ years in data engineering. • Strong experience with Apache Spark / PySpark and large-scale data processing. • Experience building ETL/ELT pipelines on cloud infrastructure (managed Spark, object storage, managed ETL, or equivalent). • Knowledge of data quality frameworks and data governance. • Experience with data anonymization and privacy-preserving data processing. • Understanding of ML training data requirements. • Proficiency in Python and SQL. • Experience with data catalog tools and metadata management. • BS/MS in Computer Science or equivalent experience. • Experience in B2B SaaS with multi-tenant data preferred.

🏖️ Benefits

• Pioneering Technology • Collaborative Culture • Global Impact

Apply Now

Similar Jobs

🕒 March 26

Shuru

51 - 200

🤖 Artificial Intelligence

🤝 B2B

🏢 Enterprise

Data Engineer role at Shuru Technologies focusing on scalable data platforms and pipelines. Collaborate with teams to build robust architectures and optimize data workflows.

Azure

ETL

MariaDB

PySpark

Spark

SQL

🕒 March 20

Blend360

501 - 1000

🏥 Healthcare

🏨 Hospitality

✈️ Travel

Data Engineer designing and optimizing data pipelines for Blend360's Media Mix Optimization platform. Collaborating with teams to ensure data quality and operational reliability.

🇮🇳 India – Remote

💰 $100M Private Equity Round on 2022-08

⏰ Full Time

🟡 Mid-level

🟠 Senior

🚰 Data Engineer

Apache

ETL

HDFS

Postgres

PySpark

Python

SQL

🕒 March 18

Codvo.ai

51 - 200

🤖 Artificial Intelligence

🔒 Cybersecurity

☁️ SaaS

Lead design and implementation of enterprise Snowflake data platform solutions driving cloud-native initiatives. Establish scalable analytics capabilities within modern data frameworks.

Cloud

Python

Scala

SQL

🕒 March 18

Codvo.ai

51 - 200

🤖 Artificial Intelligence

🔒 Cybersecurity

☁️ SaaS

Azure Data Platform Lead at Codvo designing and optimizing enterprise-scale Azure data platform solutions. Leading architecture initiatives, implementing data governance, and collaborating with teams.

Apache

Azure

Distributed Systems

ETL

PySpark

Python

Scala

Spark

SQL

🕒 March 11

Smart Working

51 - 200

💼 Consulting

🏥 Healthcare

📣 Marketing

Data Engineer responsible for building and maintaining scalable data pipelines for connecting systems. Collaborating with engineers to develop data architectures and support advanced analytics.

Apache

ETL

Spark

SQL