Senior Data Engineer (Metadata & Lineage Integration)

Intellias

Slough

On-site

GBP 90,000 - 120,000

Full time

9 days ago
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Intellias is seeking a hands-on senior Python data engineer to design and build the catalogue, semantic, entitlement and analytical layers that transform large on-premise data estates for AI agents. You will craft extraction, enrichment and registration paths, extend scraping to full metadata, and load vendor schema metadata into a central knowledge base.

You will work with event-driven architectures (Kafka), search stores (Elasticsearch/OpenSearch), and OpenLineage-like lineage models, applying

Qualifications

  • 5+ years building production data systems in Python with strong engineering fundamentals and solid SQL.
  • Experience building crawlers, harvesters or connectors extracting inventories, schemas and lineage from databases, filesystems, messages and APIs.
  • Event-driven integration with Kafka or similar, including secure producer patterns and schema-managed topics.
  • Experience with search/document stores backing catalogue platforms (Elasticsearch, OpenSearch, MongoDB).
  • Knowledge of lineage modelling with OpenLineage and readiness to work with in-house event models.
  • Experience applying LLMs to metadata work, drafting descriptions and classifications for human review.
  • Ability to work inside another team's codebase, complete components and contribute via reviews.
  • Fluent English for written and spoken communication with client teams.

Responsibilities

  • Build extraction, enrichment and registration paths populating domain catalogues from live estates (databases, time-series stores, streaming platforms, filesystems, internal and vendor APIs).
  • Extend existing scraping from bare datasets to full metadata: descriptions, field dictionaries, date ranges, asset-class and cadence tags, vendor provenance.
  • Load vendor schema metadata at scale via vendor APIs into a persistent internal knowledge base for human and LLM use.
  • Seed report and dataset inventories from application metadata tables and ETL sources; combine with LLM-drafted descriptions approved by stewards.
  • Integrate lineage into the client's lineage backend across batch, streaming and cross-system report chains.
  • Implement the federation contract defined by the architect: stable identities, ownership, hierarchy, links and availability state exposed by each local catalogue.
  • Keep metadata current through scheduled and event-driven refresh with staleness detection and quality signals.

Skills

Python programming
Data engineering
SQL
Kafka
Data catalog & lineage
Elasticsearch / OpenSearch
LLM integration
English communication

Education

Bachelor's degree in Computer Science or related field

Tools

Kafka
Elasticsearch/OpenSearch
MongoDB
Airflow
DataHub
OpenLineage
Parquet

Job description

Our client is a leading global investment management company headquartered in London. It manages over $228 billion in assets and serves institutional investors, pension funds, wealth managers, and other sophisticated clients worldwide. The firm specializes in quantitative investing, alternative investments, systematic trading strategies, and technology-driven asset management. Data science, machine learning, and AI are core components of its investment and research processes.

As part of our collaboration we will focus on two foundational capabilities required to enable safe and scalable AI adoption across the enterprise: Agentic Security and AI-Ready Data Foundations.

We build the data foundations that make AI useful and safe inside regulated financial firms. The value of AI is capped by the data its agents can reach: if an agent cannot find, interpret, trace or be correctly permissioned against data, the capability is useless, or worse, unsafe. Your job is to close that gap.

This is a hands-on senior role for an excellent Python engineer with strong data-engineering skills who is genuinely comfortable building with AI agents. You will design and build the catalogue, semantic, entitlement and analytical layers that turn large on-premise data estates into something agents can use.

Requirements:

  • 5+ years building production data systems in Python, with strong engineering fundamentals (testing, code review, performance) and solid SQL.
  • Experience building crawlers, harvesters or connector frameworks that extract inventories, schemas, field dictionaries and lineage from databases, filesystems, message platforms and API surfaces, in addition to conventional data pipelines.
  • Event-driven integration with Kafka or similar, including secure producer patterns (mTLS or equivalent) and schema-managed topics.
  • Experience with search and document stores that back catalogue platforms (e.g. Elasticsearch, OpenSearch, MongoDB or similar).
  • Working knowledge of lineage capture and modelling, with OpenLineage or similar as a reference, and readiness to work with proprietary in-house event models.
  • Experience applying LLMs to metadata work, such as drafting descriptions and classifications for human review, including quality evaluation of the generated output.
  • Readiness to work inside another team's codebase, complete components that the team has designed, and contribute through its review process.
  • Fluent English for written and spoken communication with client teams.

Will be a plus

  • Time-series and tick stores (e.g. kdb+ or similar columnar time-series databases), market-data vendor schema APIs, symbology and asset-class concepts (market-data opening).
  • MS SQL Server estates, reporting and BI systems, inventory extraction from application metadata tables (reporting opening).
  • Columnar and lake formats (Parquet or similar), large object stores, orchestration platforms (Airflow or similar).
  • Working-level knowledge of graph databases; data contracts and data quality frameworks; catalogue platforms (DataHub or similar) on the ingestion side.
  • Day-to-day use of AI coding agents; building data services consumed by AI agents.
  • Experience in financial services or other regulated on-premise environments.

Responsibilities:

  • Build extraction, enrichment and registration paths that populate domain catalogues from live estates (databases, time-series stores, streaming platforms, filesystems, internal and vendor APIs) and connect them to the client's central catalogue through its existing mechanisms.
  • Extend an existing scraping capability from bare dataset and symbol inventories to full metadata: descriptions, field-level dictionaries, date ranges, asset-class and cadence tags, vendor provenance.
  • Load vendor schema metadata at scale, through vendor APIs, into a persistent internal knowledge base designed for step-by-step disclosure to humans and LLMs.
  • Seed report and dataset inventories from existing application metadata tables and ETL sources; combine them with LLM-drafted descriptions approved by stewards.
  • Integrate lineage into the client's lineage backend across batch, streaming and cross-system report chains, completing the client's existing registration designs.
  • Implement the federation contract defined by the architect: stable identities, ownership, hierarchy, links and availability state exposed by each local catalogue to the central layer.
  • Keep metadata current through scheduled and event-driven refresh, with explicit staleness detection and quality signals.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Data Engineer (Metadata & Lineage Integration)
Senior Data Engineer (Metadata & Lineage Integration)

Intellias • City Of London

On-site
GBP 100,000 - 150,000
Senior Data Engineer (Metadata & Lineage Integration)
Senior Data Engineer (Metadata & Lineage Integration)

Intellias • Greater London

On-site
GBP 110,000 - 165,000
Senior Data Engineer
Senior Data Engineer

Intellias • Greater London

On-site
GBP 90,000 - 120,000
Data Architect (Metadata, Governance & Semantics)
Data Architect (Metadata, Governance & Semantics)

Intellias • Slough

On-site
GBP 90,000 - 140,000
Data Architect (Metadata, Governance & Semantics)
Data Architect (Metadata, Governance & Semantics)

Intellias • Greater London

On-site
GBP 110,000 - 150,000
Senior Data Engineer - AI-Ready Metadata & Lineage
Senior Data Engineer - AI-Ready Metadata & Lineage

Intellias • Slough

On-site
GBP 90,000 - 120,000
Data Architect (Metadata, Governance & Semantics)
Data Architect (Metadata, Governance & Semantics)

Intellias • City Of London

On-site
GBP 46,000 - 50,000
Senior Data Engineer: AI-Ready Metadata & Lineage
Senior Data Engineer: AI-Ready Metadata & Lineage

Intellias • City Of London

On-site
GBP 100,000 - 150,000
Senior AI Engineer
Senior AI Engineer

Intellias • Slough

On-site
GBP 90,000 - 130,000
Principal AI Engineer
Principal AI Engineer

Intellias • Greater London

On-site
GBP 90,000 - 130,000