Senior Data & Document Ingestion Engineer – OCR, RAG

🔥 0 minutes ago

🌐 Poland, Czechia, +4 more countries – Remote

infoinfo

⏳ Contract/Temporary

🟠 Senior

👷🏻‍♀️ Engineer

👻 Ghost score 12%

infoinfo
Apply Now
Find Similar Remote Jobs

📊 Check your resume score for this job

Improve your chances of getting an interview by checking your resume score before you apply.

Logo of Gramian Consulting

Gramian Consulting

2 - 10 employees

Founded 2025

💼 Consulting

📦 Logistics

📣 Marketing

Consulting • Logistics • Marketing

Gramian Consulting is a remote-first consulting firm that connects engineering and data/AI talent with organizations through talent augmentation, recruiting, dedicated teams, and contractor management. The firm provides Data & AI services including LLM training and fine-tuning, AI agents and assistants, MLOps, and AI infrastructure, and it offers mentorship and education programs for career readiness, interview preparation, and international market orientation. Rooted in hands-on engineering and recruiting experience, Gramian helps clients scale technical teams and extract business value from AI while developing individual talent.

📋 Description

• Design and build scalable document ingestion pipelines for PDFs, scans, emails, and office documents • Integrate and optimize OCR and document extraction technologies for high-accuracy text and layout extraction • Build workflows for text cleaning, normalization, semantic chunking, and metadata tagging • Process unstructured formats including PDF, Word, Excel, and PowerPoint • Develop connectors for enterprise sources such as SharePoint and email systems • Design data schemas and retrieval mechanisms for downstream AI and RAG use cases • Build validation and monitoring loops to detect low-confidence OCR or extraction results • Ensure ingestion pipelines meet enterprise security, reliability, and latency requirements • Implement logging, testing, and operational monitoring across data-processing workflows • Apply Git, CI/CD, and software-engineering best practices to pipeline development

🎯 Requirements

• Approximately 5–10 years of professional data engineering or backend/data-platform experience • Strong hands-on experience with Python and SQL • Proven experience building data ingestion and document-processing pipelines • Hands-on experience processing unstructured documents such as PDF, Word, Excel, PPT, scans, or emails • Experience with OCR/document extraction tools such as AWS Textract or equivalent • Professional experience building data-processing pipelines on public cloud platforms • Experience with AWS services such as S3, Step Functions, and CloudWatch, or comparable cloud services • Strong development practices including Git, CI/CD, and automated testing • Experience with Azure, AWS, or Databricks in enterprise data environments preferred • Experience with vector databases, embeddings, or RAG architectures preferred • Experience designing connectors to SharePoint, email, or other enterprise content systems preferred • Background in insurance, financial services, or regulated-data environments preferred • Fluent English is required

Apply Now

Similar Jobs

🕒 August 19

Intetics

501 - 1000

💼 Consulting

🏥 Healthcare

📦 Logistics

Senior Blue Yonder Dispatcher Engineer supporting live WMS environments. Developing PL/SQL functionality, reporting, enhancements, and production migrations.

Oracle

SQL

🕒 August 18

Look4IT

11 - 50

💼 Consulting

📣 Marketing

📦 Logistics

Senior AI Engineer building and operating production LLM features for real users at an IT services company. Designing agentic workflows, RAG, evaluation, observability, and cost-efficient cloud services.

Azure

Cloud

🕒 August 11

Work Life Group

11 - 50

🎯 Recruiter

👥 HR Tech

Jira-Confluence Data Center engineer rebuilding a NATO cyber operations platform mock-up. Capturing AQAP requirements and delivering video and live industry demonstrations.

Azure

Groovy

Java

Linux

Python