Senior Data & Document Ingestion Engineer (OCR / RAG)

Gramian Consulting

Poland

On-site

PLN 389,000 - 562,000

Full time

5 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Gramian Consultancy seeks a Senior Data & Document Ingestion Engineer to design and implement scalable ingestion pipelines for high-volume unstructured insurance content. You will focus on OCR, document parsing, and semantics to prepare data for downstream AI systems.

Contractor assignment with potential extensions, remote Europe-based work, fluent English required. Ideal candidates have strong Python/SQL skills and hands-on experience with cloud-based data processing and ETL frameworks.

Qualifications

  • 5–10 years of professional data engineering or backend/data-platform experience.
  • Strong hands-on experience with Python and SQL.
  • Proven experience building data ingestion and document-processing pipelines.
  • Hands-on experience processing unstructured documents (PDF, Word, Excel, PPT, scans, or emails).
  • Experience with OCR/document extraction tools such as AWS Textract or equivalent.
  • Professional experience building data-processing pipelines on public cloud platforms.
  • Experience with AWS services (S3, Step Functions, CloudWatch) or comparable cloud services.
  • Strong development practices including Git, CI/CD, and automated testing.

Responsibilities

  • Design and build scalable document ingestion pipelines for PDFs, scans, emails, and office documents.
  • Integrate and optimize OCR and document extraction technologies for high-accuracy text and layout extraction.
  • Build workflows for text cleaning, normalization, semantic chunking, and metadata tagging.
  • Process unstructured formats including PDF, Word, Excel, and PowerPoint.
  • Develop connectors for enterprise sources such as SharePoint and email systems.
  • Design data schemas and retrieval mechanisms for downstream AI and RAG use cases.
  • Build validation and monitoring loops to detect low-confidence OCR or extraction results.
  • Ensure ingestion pipelines meet enterprise security, reliability, and latency requirements.
  • Implement logging, testing, and operational monitoring across data-processing workflows.
  • Apply Git, CI/CD, and software-engineering best practices to pipeline development.

Skills

Python
SQL
Data pipelines
Git
CI/CD

Tools

AWS Textract
AWS S3
Step Functions
CloudWatch

Job description

About Gramian

Gramian Consultancy is a boutique consultancy specializing in IT professional services and engineering talent solutions. With a strong background in software engineering and leadership, we help companies build high-performing teams by matching them with professionals who truly fit their needs.

About Gramian

Gramian Consultancy is a boutique consultancy specializing in IT professional services and engineering talent solutions. With a strong background in software engineering and leadership, we help companies build high-performing teams by matching them with professionals who truly fit their needs.

About The Role

Our client is a Big4 Consultancy group that works with leading financial institutions on AI-driven transformation, automation, advanced analytics, and financial crime prevention. Their work spans intelligent fraud detection, AML/KYC modernization, autonomous workflows, enterprise AI platforms, and the secure industrialization of AI in highly regulated environments.

We are looking for a Senior Data & Document Ingestion Engineer to build robust pipelines for processing high-volume unstructured insurance content. The role focuses on OCR, document parsing, ingestion pipelines, text normalization, semantic chunking, metadata extraction, and retrieval-ready data preparation for downstream AI systems.

CONTRACT: Contractor assignment, expected October 2026 - July 2027, with extension to other projects (and retention rate)

COMMITMENT: Full-time

LOCATIONS: REMOTE 100%, Europe-based

PROCESS: Initial qualification followed by technical and client interviews

NOTES: Fluent English is required. Must be able to work in EU.

Responsibilities
  • Design and build scalable document ingestion pipelines for PDFs, scans, emails, and office documents
  • Integrate and optimize OCR and document extraction technologies for high-accuracy text and layout extraction
  • Build workflows for text cleaning, normalization, semantic chunking, and metadata tagging
  • Process unstructured formats including PDF, Word, Excel, and PowerPoint
  • Develop connectors for enterprise sources such as SharePoint and email systems
  • Design data schemas and retrieval mechanisms for downstream AI and RAG use cases
  • Build validation and monitoring loops to detect low-confidence OCR or extraction results
  • Ensure ingestion pipelines meet enterprise security, reliability, and latency requirements
  • Implement logging, testing, and operational monitoring across data-processing workflows
  • Apply Git, CI/CD, and software-engineering best practices to pipeline development
Requirements
  • Approximately 5-10 years of professional data engineering or backend/data-platform experience
  • Strong hands-on experience with Python and SQL
  • Proven experience building data ingestion and document-processing pipelines
  • Hands-on experience processing unstructured documents such as PDF, Word, Excel, PPT, scans, or emails
  • Experience with OCR/document extraction tools such as AWS Textract or equivalent
  • Professional experience building data-processing pipelines on public cloud platforms
  • Experience with AWS services such as S3, Step Functions, and CloudWatch, or comparable cloud services
  • Strong development practices including Git, CI/CD, and automated testing
Preferred Qualifications
  • Experience with Azure, AWS, or Databricks in enterprise data environments
  • Experience with vector databases, embeddings, or RAG architectures
  • Experience designing connectors to SharePoint, email, or other enterprise content systems
  • Background in insurance, financial services, or regulated-data environments
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead OCR & RAG Data Ingestion Engineer (Remote)
Lead OCR & RAG Data Ingestion Engineer (Remote)

Gramian Consulting • Poland

On-site
PLN 389,000 - 562,000
Senior Backend / Platform Engineer (AI Infrastructure)
Senior Backend / Platform Engineer (AI Infrastructure)

Gramian Consulting Group • Poland

On-site
PLN 180,000 - 300,000
Senior Business Analyst (Insurance AI Transformation)
Senior Business Analyst (Insurance AI Transformation)

Gramian Consulting • Poland

On-site
PLN 304,000 - 477,000
Senior AI Engineer (Agentic AI / AWS)
Senior AI Engineer (Agentic AI / AWS)

Gramian Consulting • Poland

On-site
PLN 390,000 - 564,000
Senior DevSecOps Engineer (Cloud Security / AI Platforms)
Senior DevSecOps Engineer (Cloud Security / AI Platforms)

Gramian Consulting • Poland

On-site
PLN 250,000 - 350,000
Senior Backend / Platform Engineer (AI Infrastructure)
Senior Backend / Platform Engineer (AI Infrastructure)

Gramian Consulting • Poland

On-site
PLN 165,000 - 248,000
Data Engineer
Data Engineer

Neurons Lab • Poland

On-site
PLN 78,000 - 112,000
Senior Test Automation Engineer (AI Systems)
Senior Test Automation Engineer (AI Systems)

Gramian Consulting • Poland

On-site
PLN 120,000 - 180,000
Senior Software Engineer, AI Agents
Senior Software Engineer, AI Agents

IFS • Województwo małopolskie

Hybrid
PLN 240,000 - 420,000
Hybrid work
Global team
Senior Software Engineer, AI Agents
Senior Software Engineer, AI Agents

IFS • Kraków

Hybrid
PLN 250,000 - 380,000