
10,000+ employees
💼 Consulting
🏭 Manufacturing
📣 Marketing
💰 $1G Post-IPO Debt on 2023-05
Consulting • Manufacturing • Marketing
IQVIA is a global leader in data analytics and technology solutions, dedicated to improving health outcomes. The company utilizes its Connected Intelligence platform to harness the power of advanced data analytics and artificial intelligence, facilitating innovation in healthcare. IQVIA focuses on various areas, including clinical research, technology, and consulting services to solve complex healthcare challenges and accelerate the delivery of new therapies to patients.
🔥 15 hours ago
🌐 Spain, Portugal, +4 more countries – Remote
⏰ Full Time
🟡 Mid-level
🟠 Senior
🚰 Data Engineer
👻 Ghost score 10%
Improve your chances of getting an interview by checking your resume score before you apply.

10,000+ employees
💼 Consulting
🏭 Manufacturing
📣 Marketing
💰 $1G Post-IPO Debt on 2023-05
Consulting • Manufacturing • Marketing
IQVIA is a global leader in data analytics and technology solutions, dedicated to improving health outcomes. The company utilizes its Connected Intelligence platform to harness the power of advanced data analytics and artificial intelligence, facilitating innovation in healthcare. IQVIA focuses on various areas, including clinical research, technology, and consulting services to solve complex healthcare challenges and accelerate the delivery of new therapies to patients.
• Design, build, and maintain data infrastructure supporting IQVIA’s AI-enabled COA strategy and COA Accelerator capabilities • Own ingestion, transformation, normalisation, enrichment, indexing, versioning, and governance of public, proprietary, and client-specific data sources • Build ingestion pipelines for structured and unstructured sources including documents, databases, APIs, clinical trial registries, regulatory documents, scientific publications, and internal repositories • Transform source material into standardised, searchable, AI-ready formats for evidence retrieval, source citation, recommendation generation, and expert review • Develop processes for document parsing, OCR, text extraction, metadata enrichment, chunking, deduplication, versioning, indexing, and quality control • Support the platform’s core knowledge layer, including COA metadata, psychometric evidence, therapeutic area mappings, endpoint usage, regulatory precedent, and scientific evidence • Support integration of public evidence sources including clinical trial registries, FDA labels, EMA EPARs, HTA records, scientific literature, FDA guidance, and qualification documents • Support ingestion of proprietary internal knowledge, publications, thought leadership, and expert-authored content • Prepare data for retrieval-augmented generation using chunking, embeddings, indexes, metadata filters, and source reference structures • Collaborate with AI engineers to improve retrieval precision, recall, relevance, and citation accuracy • Implement hybrid retrieval combining semantic search, keyword search, structured database queries, and metadata filtering • Maintain traceability between AI-generated outputs and source documents • Implement data quality controls and maintain audit trails for ingestion, transformation, updates, deletions, access rights, and downstream use • Work with legal, security, compliance, product, and domain stakeholders on contractual, licensing, privacy, intellectual property, and governance requirements • Collaborate across AI Engineering, Product, COA Science, and Software Engineering
• Degree in computer science, data engineering, data science, information systems, bioinformatics, computational biology, engineering, or a related technical field • Experience designing, building, and maintaining data pipelines for structured and unstructured data • Strong Python and SQL skills • Experience with APIs, relational databases, document stores, search indexes, cloud data platforms, and ETL/ELT workflows • Experience handling large volumes of text-heavy documents • Strong understanding of data cleaning, normalisation, metadata management, document parsing, indexing, lineage, versioning, and auditability • Familiarity with data modelling for complex scientific, clinical, regulatory, or healthcare knowledge domains • Ability to translate domain expert requirements into practical data structures, metadata models, retrieval-ready content, and maintainable pipelines • Strong attention to detail and ability to identify data quality issues • Ability to collaborate with AI engineers, product managers, COA scientists, software engineers, security stakeholders, legal teams, and commercial teams • Strong documentation skills • Experience in life sciences, clinical research, healthcare, regulatory data, scientific publishing, HEOR, clinical outcome assessments, patient-reported outcomes, or medical evidence management is strongly preferred • Experience with clinical trial registries, regulatory labels, HTA reports, scientific literature databases, medical knowledge repositories, or similar evidence sources • Experience with vector databases, embeddings, semantic search, Elasticsearch/OpenSearch, Azure AI Search, Pinecone, Weaviate, Milvus, Qdrant, or similar technologies • Experience with cloud data platforms such as Azure, AWS, or GCP • Experience with document AI, OCR, layout-aware parsing, table extraction, metadata enrichment, taxonomy development, or controlled vocabularies is desirable • Familiarity with ontology development, biomedical terminologies, controlled vocabularies, evidence classification, and structured knowledge representation • Understanding of GDPR, data privacy, intellectual property constraints, licensed content, confidential client data, access control, and secure data handling • Ability to work independently in a remote or hybrid environment while collaborating across global teams • Significant experience leveraging AI tools for work • Fluency in English
• Rewarding and progressive career • Training and support • Development and growth opportunities • Hands-on influence in developing and delivering innovative solutions • Multi-cultural, collegial, and collaborative work environment
Apply Now🕒 September 18
Data Engineer developing clustering, churn prediction, and speech analytics models for IRIUM’s banking project. Remote role requiring residence in Spain.
🗣️🇪🇸 Spanish Required
Linux
Python
SQL
🕒 September 8
Senior Data Engineer building Databricks and Airflow pipelines for Leadtech’s web and mobile products. Supporting marketing, payments, reporting, attribution, and LTV data.
Airflow
Apache
BigQuery
Google Cloud Platform
Python
Spark
SQL
Unity
🕒 September 4
Presales Data Architect diseñando arquitecturas Modern Data Stack y AI para Reclut, marketplace de reclutamiento con IA. Liderando preventa, PoCs y conexión entre clientes y equipos de delivery.
🗣️🇪🇸 Spanish Required
AWS
Azure
BigQuery
Cloud
Google Cloud Platform
🕒 September 1
Data Engineer diseñando arquitecturas modernas con Snowflake y dbt para clientes de Reclut. Liderando preventa, consultoría técnica y pruebas de concepto de datos.
🗣️🇪🇸 Spanish Required
AWS
Azure
BigQuery
Cloud
Google Cloud Platform
🕒 August 29
Data Engineer II building high-volume data pipelines for Ookla’s connectivity intelligence platform. Supporting data scientists and analysts with reliable, accurate internet-performance data.
Airflow
Amazon Redshift
AWS
Cloud
ETL
MySQL
Python
Spark
SQL