Senior Python Data Engineer – OCR & Document Processing

🔥 16 hours ago

🌐 Portugal, Romania – Remote

infoinfo

⏰ Full Time

🟠 Senior

🚰 Data Engineer

👻 Ghost score 12%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Inetum

Inetum

10,000+ employees

💼 Consulting

🏥 Healthcare

🛡️ Insurance

💰 Post-IPO Equity on 2007-03

Consulting • Healthcare • Insurance

Inetum is a European leader in digital services, providing technology consulting, solutions, and software innovation to businesses and public sector entities across 19 countries. With a workforce of 28,000 consultants, Inetum focuses on helping clients achieve digital transformation through a broad range of services, including consulting, software integration, and outsourcing, while prioritizing innovation, customer experience, and agility. Inetum partners with major software providers and offers solutions in vertical sectors like public sector, insurance, and healthcare. In 2023, Inetum achieved sales of 2. 5 billion euros, driven by its growth and scale ambitions.

📋 Description

• Design, build, and optimize scalable data ingestion and document processing solutions for large volumes of unstructured insurance data • Design and implement scalable pipelines for PDFs, scans, emails, and Office files • Integrate, configure, and optimize OCR and document extraction technologies • Build automated workflows for document parsing, text cleaning, normalization, semantic chunking, and metadata enrichment • Develop connectors and integrations for SharePoint, email systems, and enterprise repositories • Design and maintain vector database schemas and retrieval mechanisms for RAG solutions and AI applications • Ensure pipelines meet enterprise security, compliance, performance, and availability requirements • Implement monitoring, validation, and quality-control mechanisms for low-confidence OCR and extraction results • Optimize workflows for scalability, reliability, and low-latency operations • Collaborate with AI Engineers, Backend Engineers, and Platform teams on end-to-end AI-powered document processing solutions • Develop and maintain cloud-native data ingestion solutions on public cloud platforms

🎯 Requirements

• 5-10 years of experience in Data Engineering, Data Processing, Document Intelligence, or related fields • Proven experience building scalable data ingestion and processing pipelines • Experience working with large volumes of unstructured and semi-structured data • Experience designing cloud-based data solutions • Strong programming skills in Python • Strong SQL knowledge • Hands-on experience with AWS services, including S3, Step Functions, and CloudWatch • Experience processing unstructured documents such as PDF, Word, Excel, PowerPoint, and email content • Experience building connectors and integrations with enterprise content repositories, such as SharePoint • Experience with OCR and document extraction tools, such as AWS Textract or equivalent • Experience designing and implementing data ingestion and transformation pipelines • Familiarity with vector databases and Retrieval-Augmented Generation (RAG) concepts • Experience with Git, CI/CD, and automated testing • Experience with Vector Databases • Experience with RAG architectures and AI/LLM-based applications • Experience with Azure cloud services • Experience with Databricks • Experience in Insurance, Banking, or other regulated industries

🏖️ Benefits

• Full access to foreign language learning platform • Personalized access to tech learning platforms • Tailored workshops and trainings to sustain your growth • Medical insurance • Meal tickets • Monthly budget to allocate on flexible benefit platform • Access to 7 Card services • Wellbeing activities and gatherings

Apply Now

Similar Jobs

🕒 2 days ago

Keyrus

1001 - 5000

💼 Consulting

Senior Data Engineer migrating Hadoop/Hive workflows into Snowflake for Keyrus, an international AI and data consulting group. Designing scalable cloud data platforms and automated migration processes.

🗣️🇫🇷 French Required

Apache

AWS

Azure

Cloud

ETL

Google Cloud Platform

Hadoop

Informatica

Python

Scala

Spark

SQL

🕒 5 days ago

Southend Pharmacy

201 - 500

🏥 Healthcare

💼 Consulting

📦 Logistics

Data Engineer building reliable pipelines and models for Allia Health’s affordable anti-aging and wellness products. Supporting analytics, reporting, and AI initiatives.

Airflow

AWS

Azure

BigQuery

Cloud

ETL

Google Cloud Platform

Python

SQL

🕒 September 9

Truv

51 - 200

💼 Consulting

🏥 Healthcare

📦 Logistics

Software Engineer building Python cloud services and APIs for Truv’s financial data platform. Developing optimized code, testing integrations, and supporting secure payroll data access.

AWS

Cloud

Django

Docker

Postgres

Python

🕒 August 28

bridge351

51 - 200

🎯 Recruiter

🤝 B2B

Mid Data Platform Engineer supporting data platforms for an international energy provider. Managing pipelines, cloud services, data quality, and production releases at Bridge351.

🗣️🇧🇷🇵🇹 Portuguese Required

Airflow

Apache

AWS

Cloud

ETL

Grafana

Kafka

Python

Scala

Spark

SQL

🕒 August 21

MWDN

51 - 200

💼 Consulting

📦 Logistics

🏥 Healthcare

Senior Data Engineer building scalable real-time data infrastructure for a fast-growing programmatic advertising company. Optimizing auctions, yield, and partner performance with massive-scale ad-tech data.

Airflow

Amazon Redshift

AWS

BigQuery

Cloud

ETL

Google Cloud Platform

Kafka

Python

Spark

SQL