
1001 - 5000 employees
Founded 2006
💼 Consulting
📦 Logistics
🏥 Healthcare
Consulting • Logistics • Healthcare
Coupa Software is a leading provider of business spend management solutions. Their platform focuses on optimizing and transforming direct and indirect spend across procurement, finance, supply chain, and IT. Coupa leverages AI and extensive data insights to drive cost efficiencies, manage supplier relationships, and mitigate risks. With products covering areas such as invoicing, payments, expense management, and supply chain collaboration, Coupa serves a wide range of industries including automotive, healthcare, retail, and more. Their comprehensive community and partner ecosystem enable organizations to unlock hidden savings and improve compliance, promoting growth and resilience in a changing economic climate.
🕒 March 31
Improve your chances of getting an interview by checking your resume score before you apply.

1001 - 5000 employees
Founded 2006
💼 Consulting
📦 Logistics
🏥 Healthcare
Consulting • Logistics • Healthcare
Coupa Software is a leading provider of business spend management solutions. Their platform focuses on optimizing and transforming direct and indirect spend across procurement, finance, supply chain, and IT. Coupa leverages AI and extensive data insights to drive cost efficiencies, manage supplier relationships, and mitigate risks. With products covering areas such as invoicing, payments, expense management, and supply chain collaboration, Coupa serves a wide range of industries including automotive, healthcare, retail, and more. Their comprehensive community and partner ecosystem enable organizations to unlock hidden savings and improve compliance, promoting growth and resilience in a changing economic climate.
• Lead the design and implementation of data pipelines that prepare high-quality training data for AI models. • Build data curation workflows that transform raw enterprise data into labeled, validated datasets. • Design data quality frameworks: validation, profiling, anomaly detection, lineage tracking. • Extend existing anonymized data export pipelines to support AI training workloads. • Implement synthetic data generation pipelines. • Design schema mappings across 197+ enterprise tables for feature extraction. • Collaborate with ML engineers on training data format requirements. • Establish data catalog and metadata management for AI training artifacts.
• 10+ years of software engineering experience, with 5+ years in data engineering. • Strong experience with Apache Spark / PySpark and large-scale data processing. • Experience building ETL/ELT pipelines on cloud infrastructure (managed Spark, object storage, managed ETL, or equivalent). • Knowledge of data quality frameworks and data governance. • Experience with data anonymization and privacy-preserving data processing. • Understanding of ML training data requirements. • Proficiency in Python and SQL. • Experience with data catalog tools and metadata management. • BS/MS in Computer Science or equivalent experience. • Experience in B2B SaaS with multi-tenant data preferred.
• Pioneering Technology • Collaborative Culture • Global Impact
Apply Now🕒 March 26
Data Engineer role at Shuru Technologies focusing on scalable data platforms and pipelines. Collaborate with teams to build robust architectures and optimize data workflows.
Azure
ETL
MariaDB
PySpark
Spark
SQL
🕒 March 20
Data Engineer designing and optimizing data pipelines for Blend360's Media Mix Optimization platform. Collaborating with teams to ensure data quality and operational reliability.
🇮🇳 India – Remote
💰 $100M Private Equity Round on 2022-08
⏰ Full Time
🟡 Mid-level
🟠 Senior
🚰 Data Engineer
Apache
ETL
HDFS
Postgres
PySpark
Python
SQL
🕒 March 18
Lead design and implementation of enterprise Snowflake data platform solutions driving cloud-native initiatives. Establish scalable analytics capabilities within modern data frameworks.
Cloud
Python
Scala
SQL
🕒 March 18
Azure Data Platform Lead at Codvo designing and optimizing enterprise-scale Azure data platform solutions. Leading architecture initiatives, implementing data governance, and collaborating with teams.
Apache
Azure
Distributed Systems
ETL
PySpark
Python
Scala
Spark
SQL
🕒 March 11
Data Engineer responsible for building and maintaining scalable data pipelines for connecting systems. Collaborating with engineers to develop data architectures and support advanced analytics.
Apache
ETL
Spark
SQL