Principal Data Engineer – RWE

🔥 12 hours ago

🇮🇳 India – Remote

⏰ Full Time

🔴 Lead

🚰 Data Engineer

👻 Ghost score 12%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Veramed

Veramed

501 - 1000 employees

Founded 2012

🏥 Healthcare

💊 Pharmaceuticals

🧬 Biotechnology

Healthcare • Pharmaceuticals • Biotechnology

Veramed is a biometrics-focused contract research organization (CRO) that provides data management, biostatistics, programming, evidence & value generation, and standards/AI services to pharmaceutical and biotechnology clients. The company offers end-to-end biometrics support, Functional Service Provider (FSP) and project-based delivery models, regulatory submission-ready data, advanced automation and AI (Veramed. ai), CDISC compliance, and data visualization to minimize submission issues and audit findings. Veramed positions itself as a trusted partner for biotech and pharma, emphasizing operational governance, scalability, and a people-first culture.

📋 Description

• Develop data processes for automated ongoing generation of patient-level data products for dashboards, reports, and studies • Manipulate datasets onboarded from data vendors and partners into usable structures for RWE studies, dashboards, and other outputs • Transform heterogeneous healthcare datasets into reusable data models for observational research and epidemiology studies • Convert bespoke datasets to OMOP format where appropriate and account for residual data that cannot be converted • Build FAIR data pipelines and semantic data engineering frameworks • Create AI-ready datasets supporting generative AI use cases • Engage with epidemiologists, statisticians, market access and health economists to scope requirements and translate needs into actionable data structures • Collaborate with the RWE programming team to build data structures and support study outputs • Liaise with IT to ensure inbound datasets are fit for agreed purposes • Liaise with technical staff from analysis software vendors such as Databricks • Maintain documentation of data flows, schemas, pipelines, and processes • Design and perform data validation and monitoring to ensure accuracy and reliability • Troubleshoot data loading, extraction, and transformation issues • Collaborate with three other Data Engineering team members and provide workload support as needed

🎯 Requirements

• Strong understanding of Real World Data (RWD) and Real World Evidence (RWE) concepts • Ability to assess business requirements and recommend appropriate real-world healthcare datasets for analytical use cases • Deep understanding of healthcare data models and healthcare data ecosystems • Strong expertise in OMOP CDM v5.4, v6, including extensions • Knowledge of SNOMED CT, RxNorm, ICD-10, LOINC, and HCPCS/CPT • Strong experience building scalable ETL/ELT pipelines • Expertise in Databricks, PySpark, Spark SQL, SQL, and Delta Lake • Experience working with large-scale healthcare and patient-level datasets • Strong understanding of Semantic Data Engineering principles • Experience building FAIR-compliant data pipelines • Experience with cloud-based data platforms and distributed processing frameworks • Strong Power BI development and data modelling skills • Ability to create reusable analytical datasets for dashboards and studies • Experience designing AI-ready datasets and analytics data products • Experience implementing automated data quality frameworks • Strong data profiling, validation, and monitoring skills • Understanding of healthcare data quality assessment methodologies • Excellent stakeholder management and communication skills • Ability to translate complex business requirements into technical solutions • Experience working with cross-functional global teams • Exposure to one or more of: Oncology, Respiratory, Immunology & Inflamation, Infectious Diseases

🏖️ Benefits

• Equal opportunities employer • Inclusive and diverse working environment • B Corp accredited company focused on social and environmental performance, transparency, and accountability • AI-assisted recruitment with human-led assessments, selection decisions, and hiring outcomes

Apply Now

Similar Jobs

🕒 3 days ago

Solvd, Inc.

501 - 1000

💼 Consulting

🏥 Healthcare

📦 Logistics

Snowflake Data Architect designing a workforce intelligence platform as a Snowflake Marketplace native app. Guiding security review, cloud architecture, and compliance documentation for Solvd’s AI consulting business.

AWS

Azure

Cloud

ETL

Google Cloud Platform

🕒 4 days ago

Ellit Groups

51 - 200

💼 Consulting

📣 Marketing

📦 Logistics

Engineering Director leading AI-native platform, infrastructure, and data engineering for Carousell Group’s secondhand marketplace. Scaling reliable systems and developing distributed technical leadership.

Cloud

Distributed Systems

🕒 5 days ago

G-P

1001 - 5000

💼 Consulting

🏥 Healthcare

⚖️ Legal

Data Architect designing Databricks-native architecture, ingestion, governance, and security. Supporting G-P’s SaaS global employment platform serving organizations across 180+ countries.

AWS

Azure

Cloud

Google Cloud Platform

PySpark

Python

SDLC

Spark

SQL

🕒 September 7

Mashreq

1001 - 5000

🏦 Banking

💸 Finance

💳 Fintech

Data Solution Architect leading Mashreq Global Services' Information Management–Data Warehouse department. Designing Databricks Lakehouse solutions, enterprise data models, streaming systems, and governance frameworks.

Cloud

Informatica

🕒 September 7

Mashreq

1001 - 5000

🏦 Banking

💸 Finance

💳 Fintech

Assistant Vice President leading Mashreq Global Services' enterprise data warehouse architecture, operations, and strategy. Guiding specialists, securing platforms, and driving data modernization.

Azure

Kafka

NoSQL