Senior AI Data Engineer

Billennium

Kraków

Hybrid

PLN 220,000 - 300,000

Full time

6 hours ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Udemy for Business
Private medical care
Multisport card
Veterinary package
Language lessons
Shopping vouchers

Job summary

Billennium is seeking a Senior AI Data Engineer / Data Scientist to transform complex, unstructured enterprise data into AI‑ready assets. You will own the data lifecycle for AI use cases—from discovery and cleaning to enrichment, ingestion and retrieval optimization.

You will work with SharePoint repositories, documents, PDFs, Office files and more, preparing data for RAG systems, AI agents and copilots. Collaborate with AI engineers and stakeholders to ensure accuracy, scalability and

Qualifications

  • 5+ years of professional experience in Data Engineering, Applied Data Science, Analytics Engineering or related field.
  • Proven experience owning and maintaining production‑grade data pipelines.
  • Strong Python skills for data processing, parsing, cleaning, normalization and pipeline development.
  • Hands‑on experience with unstructured enterprise data (documents, PDFs, Office files, wikis, knowledge bases).
  • Experience preparing data for retrieval, NLP or RAG use cases, including embeddings and metadata enrichment.
  • Understanding of data quality engineering, validation, monitoring, lineage and freshness.

Responsibilities

  • Lead data discovery and triage for AI use cases, identifying authoritative, relevant sources.
  • Analyze repositories to identify duplicates, outdated or low‑quality content.
  • Clean, normalize, deduplicate and standardize large data volumes.
  • Define data inclusion/exclusion criteria for AI systems considering relevance and sensitivity.
  • Build and maintain ingestion pipelines for SharePoint and enterprise repositories.
  • Process documents, PDFs, Word/Excel files, wiki pages and other content.
  • Define and maintain document normalization standards and taxonomy.
  • Prepare data for retrieval and RAG applications, including chunking and metadata enrichment.
  • Support retrieval optimization and post‑retrieval techniques such as reranking.
  • Collaborate with AI Engineers on retrieval interfaces and knowledge integration.
  • Establish data quality gates and refresh processes; use frameworks like RAGAS/DeepEval.
  • Ensure traceability, privacy and auditability across data pipelines.
  • Create reusable data cleanup approaches and metadata schemas.

Skills

5+ years experience
Python
production pipelines
unstructured data
retrieval / RAG
data quality
stakeholder collaboration

Tools

Postgres + pgvector
Langfuse
Microsoft Presidio

Job description

We are a global technology company providing IT services and digital solutions to clients around the world. Headquartered in Poland, with offices in Canada, Malaysia, Germany, Switzerland and India, we bring together over 1,700 professionals who share one mission -to harness technology and innovation to make the world run better. At the heart of Billennium is our vibrant and people-centric culture, where everyone’s potential matters. We embrace diversity, inclusion, and flexibility, empowering our employees to grow and thrive. Our values - captured by the acronym TIGER - define who we are and how we work: Trust, Innovation, Growth, Energy, Responsibility – guide us in everything we do.

Role Summary:

We are looking for a Senior AI Data Engineer / Data Scientist to join our team and help transform complex, unstructured enterprise data into high-quality, AI-ready knowledge assets.

In this role, you will take ownership of the entire data lifecycle for AI use cases - from data discovery, cleaning, normalization and enrichment through ingestion, retrieval optimization, evaluation and ongoing refresh cycles.

You will work hands‑on with enterprise content such as SharePoint repositories, documents, PDFs, Office files, wikis, databases and data lakes, preparing them for use in RAG systems, AI agents and enterprise copilots.

This is a senior, hands‑on role for someone who combines strong data engineering skills with an understanding of AI/RAG architectures, data quality and retrieval. You will collaborate closely with AI Engineers, Architects and business stakeholders to determine which data should be used, how it should be structured, and how to ensure it remains accurate, trustworthy and scalable.

What You Will Do:
Enterprise Data Discovery & cleanup
  • Lead data discovery and triage for AI use cases, identifying authoritative, relevant and trustworthy data sources.
  • Analyze enterprise repositories to identify duplicates, outdated, contradictory or low-quality content.
  • Clean, normalize, deduplicate and standardize large volumes of structured, semi-structured and unstructured data.
  • Define data inclusion and exclusion criteria for AI systems, taking into account relevance, quality, freshness and sensitivity.
Unstructured Data Ingestion
  • Build and maintain ingestion pipelines for SharePoint and enterprise document repositories.
  • Process documents, PDFs, Word/Excel files, wiki pages and other enterprise content.
  • Implement text extraction, document parsing, structure recovery and metadata capture.
  • Define and maintain document normalization standards, including naming conventions, taxonomy, metadata and canonical identifiers.
RAG-Ready Data Engineering
  • Prepare enterprise data specifically for retrieval and RAG-based AI applications, going beyond traditional ETL.
  • Design and optimize chunking strategies, metadata enrichment and document structures to improve retrieval quality and efficiency.
  • Prepare and maintain AI‑ready knowledge sets for embedding and serving through Postgres + pgvector.
  • Support retrieval optimization, including filtered retrieval, hybrid search and post‑retrieval techniques such as reranking.
  • Collaborate with AI Engineers on retrieval interfaces and the integration of knowledge assets into AI applications.
Data Quality, Evaluation & Feedback Loops
  • Define and implement data quality gates covering areas such as freshness, completeness, relevance, duplication and metadata coverage.
  • Establish monitoring and refresh processes to ensure AI knowledge remains reliable over time.
  • Work with AI Engineers to evaluate retrieval and RAG performance using approaches and frameworks such as RAGAS and DeepEval.
  • Establish human feedback loops, review queues and targeted data audits where required.
  • Use evaluation results and user feedback to continuously improve the quality and usefulness of AI-ready data.
Governance, Privacy & Auditability
  • Maintain traceability and lineage from the original source through processed data and retrieval corpus to production usage.
  • Apply enterprise data governance, privacy and security requirements to AI data pipelines.
  • Implement PII detection and masking using tools and patterns such as Microsoft Presidio, where required.
  • Ensure data processing is auditable, reproducible and aligned with enterprise standards.
Reusable Data Foundations
  • Create reusable approaches for enterprise data cleanup and RAG readiness.
  • Develop ingestion templates, metadata schemas, chunking playbooks and deduplication strategies.
  • Build repeatable data foundations that can accelerate future AI use cases rather than treating each project as a one‑off implementation.
Requirements:
Must‑Have:
  • 5+ years of professional experience in Data Engineering, Applied Data Science, Analytics Engineering or a related field.
  • Proven experience owning and maintaining production‑grade data pipelines.
  • Strong Python skills, particularly for data processing, parsing, cleaning, normalization and pipeline development.
  • Hands‑on experience working with unstructured enterprise data, including documents, PDFs, Office files, wikis or knowledge bases.
  • Proven experience preparing and managing data for retrieval, NLP or RAG‑based use cases, including embedding preparation, metadata enrichment and corpus management.
  • Strong understanding of data quality engineering, including validation, monitoring, lineage, freshness and refresh cycles.
  • Understanding of RAG and retrieval concepts, including chunking, metadata, embeddings and retrieval optimization.
  • Ability to work effectively with AI Engineers, Architects and business stakeholders to define data requirements and determine what constitutes high‑quality, useful data.
  • Strong analytical and problem‑solving skills, with a hands‑on approach to investigating and resolving data quality issues.
Nice‑to‑Have:
  • Experience with Postgres + pgvector or other vector databases/vector stores.
  • Knowledge of hybrid search, filtered retrieval and reranking techniques.
  • Experience with AI observability and monitoring tools such as Langfuse.
  • Familiarity with RAG evaluation frameworks and metrics such as RAGAS or DeepEval.
  • Experience with PII detection, masking and enterprise privacy workflows, e.g. Microsoft Presidio.
  • Experience working with enterprise data governance, lineage and auditability frameworks.
  • Familiarity with LLM gateway patterns and AI application architectures.
  • Experience collaborating on production AI agents, copilots or RAG systems.
Perks and benefits (our offer):
  • Comprehensive benefits - enjoy Udemy for Business, private medical care, Multisport card, veterinary package, language lessons, and shopping vouchers.
  • Career growth - access opportunities for professional development and learning, including perks related to our official partnerships with global IT giants: Microsoft, AWS, Snowflake, Salesforce & more.
  • Global collaboration - work with a diverse, international team.
  • Innovative environment - be part of a forward‑thinking and growth‑oriented workplace.
  • Engaging community – Work with passionate professionals and participate in team‑building events, hackathons, and CSR initiatives to make an impact beyond work.
  • Team‑building events including our company tradition (annual company event in Mazury).
  • A pleasant surprise to start your journey with us in the form of a welcome pack.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior AI Data Engineer
Senior AI Data Engineer

Billennium • Polska

Hybrid
PLN 180,000 - 260,000
Udemy for Business
Private medical care
Multisport card
+9
Senior AI Engineer - Agentic AI & Copilot
Senior AI Engineer - Agentic AI & Copilot

Billennium • Wrocław

On-site
PLN 180,000 - 240,000
Comprehensive benefits
Career growth
Innovative environment
+3
Senior AI Engineer - Agentic AI & Copilot
Senior AI Engineer - Agentic AI & Copilot

Billennium • Województwo pomorskie

On-site
PLN 260,000 - 420,000
Comprehensive benefits
Career growth
Innovative environment
+3
Senior AI Data Software Engineer
Senior AI Data Software Engineer

EPAM Systems • Poland

Hybrid
PLN 180,000 - 240,000
Hybrid by design
Remote working within Poland
Relocation opportunities
+3
Senior AI Engineer - Agentic AI & Copilot
Senior AI Engineer - Agentic AI & Copilot

Billennium • Warszawa

Hybrid
PLN 240,000 - 360,000
Udemy for Business
Private medical care
Multisport card
+2
Senior AI Engineer
Senior AI Engineer

Billennium • Poland

Hybrid
PLN 180,000 - 240,000
Udemy for Business
Private medical care
Multisport card
+7
Lead AI Data Software Engineer
Lead AI Data Software Engineer

EPAM Systems • Poland

Hybrid
PLN 220,000 - 320,000
Hybrid by design
Remote work within Poland
Career development programs
+2
Senior Data Scientist
Senior Data Scientist

deepsense.ai • Polska

Hybrid
PLN 180,000 - 320,000
Remote work
Flexible hours
Work with AI experts
+2
Senior AI Engineer - Agentic AI & Copilot
Senior AI Engineer - Agentic AI & Copilot

Billennium • Kraków

On-site
PLN 240,000 - 420,000
Udemy for Business
Private medical care
Multisport card
+1
Data Engineer (Spark)
Data Engineer (Spark)

Addepto • Województwo pomorskie

On-site
PLN 120,000 - 180,000
Flexible work arrangements
Remote or office options
Professional development opportunities