Senior Python Data Engineer – AI-Ready Data Foundations & Agentic Security
About the Client
Our client is a leading global investment management company headquartered in London, managing over $228 billion in assets and serving institutional investors, pension funds, wealth managers, and other sophisticated clients worldwide.
The firm specializes in quantitative investing, alternative investments, systematic trading strategies, and technology-driven asset management, with data science, machine learning, and artificial intelligence at the core of its investment and research processes.
As part of our collaboration, we are focusing on two foundational capabilities required to enable safe, scalable AI adoption across the enterprise: Agentic Security and AI-Ready Data Foundations.
Role Overview
We are looking for a Senior Python Data Engineer with strong data-engineering fundamentals and genuine hands‑on experience working with AI agents and LLM‑enabled systems.
The role focuses on building the data foundations that make AI useful, trustworthy, and safe within a highly regulated financial environment.
The effectiveness of enterprise AI is ultimately constrained by the data its agents can securely access. If an agent cannot discover, interpret, trace, or correctly apply permissions to data, the resulting capability is either ineffective or potentially unsafe.
In this hands‑on senior role, you will design and build catalogue, semantic, entitlement, lineage, and analytical layers that transform large on‑premises data estates into trusted, discoverable, AI‑ready data assets.
You will work closely with architects, data engineers, AI specialists, data stewards, and client technology teams to deliver scalable solutions across databases, filesystems, streaming platforms, APIs, and financial data environments.
Key Responsibilities
- Build extraction, enrichment, and registration pipelines that populate domain data catalogues from live enterprise estates, including:
- Databases
- Time‑series and market‑data stores
- Streaming platforms
- Filesystems and object stores
- Internal APIs
- Extend existing scraping and harvesting capabilities beyond basic dataset and symbol inventories to capture comprehensive metadata, including:
- Dataset and field descriptions
- Date ranges and availability
- Asset‑class classifications
- Ownership and stewardship information
- Develop and enhance crawler, harvester, and connector frameworks capable of extracting schemas, inventories, metadata, and lineage from heterogeneous data sources.
- Load vendor schema and market‑data metadata at scale through vendor APIs into a persistent internal knowledge base designed for controlled, step‑by‑step disclosure to both humans and LLM‑powered agents.
- Seed report and dataset inventories using existing application metadata tables, ETL sources, and reporting systems.
- Apply LLMs to metadata enrichment, including generating draft descriptions, classifications, and semantic annotations for review and approval by data stewards.
- Establish processes for evaluating the quality, consistency, and reliability of LLM‑generated metadata.
- Integrate data lineage into the client's existing lineage backend across:
- Batch pipelines
- Cross‑system data flows
- Reports and downstream analytical chains
- Complete and extend existing lineage registration designs using OpenLineage or equivalent concepts, while supporting proprietary internal event models.
- Implement the federation contract defined by the architecture team, ensuring local catalogues expose consistent:
- Stable identities
- Ownership
- Relationships and links
- Availability states
- Maintain metadata freshness through both scheduled and event‑driven refresh mechanisms, including explicit staleness detection and data‑quality signals.
- Integrate with event‑driven architectures using Kafka or similar messaging platforms, including secure producer patterns such as mTLS and schema‑managed topics.
- Work effectively within existing engineering teams and codebases, completing components according to established architectural designs, coding standards, testing practices, and code‑review processes.
- Build data services and platform components that can be consumed securely and reliably by AI agents.
- Contribute to the design and implementation of secure, scalable foundations for enterprise agentic AI.
- 5+ years of professional experience building production‑grade data systems using Python.
- Unit and integration testing
- Code review
- Clean and maintainable architecture
- Performance optimisation
- Debugging and observability
- Strong SQL skills and experience working with relational data systems.
- Proven experience designing and implementing production data pipelines and services.
Data Discovery & Metadata Engineering
- Experience building crawlers, harvesters, connectors, or metadata extraction frameworks.
- Experience extracting and processing:
- Dataset inventories
- Field dictionaries
- Application metadata
- API metadata
- Experience working across heterogeneous data environments rather than only conventional ETL pipelines.
- Hands‑on experience with Kafka or similar event‑streaming platforms.
- Understanding of secure producer and consumer patterns.
- Experience with mTLS or equivalent security mechanisms.
- Experience working with schema‑managed topics and event contracts.
Search & Document Stores
Experience with search or document databases used to support catalogue, metadata, or discovery platforms, such as:
- Elasticsearch
- OpenSearch
- MongoDB
- Similar search/document‑oriented technologies
- Working knowledge of data lineage capture and modelling.
- Familiarity with OpenLineage or similar lineage frameworks.
- Ability to work with proprietary/in‑house lineage event models and registration mechanisms.
AI / LLMs
- Practical experience applying LLMs to metadata and data‑engineering workflows.
- Experience using LLMs to generate or assist with:
- Classifications
- Semantic annotations
- Understanding of how to evaluate generated output for accuracy, consistency, quality, and human‑review requirements.
- Genuine comfort building and working with AI agents and agent‑enabled data services.
- Experience using AI coding agents as part of day‑to‑day software engineering is highly desirable.
- Comfortable working within another team's existing codebase and architecture.
- Ability to implement components designed by architects or senior engineers while contributing constructively through established review processes.
- Strong communication and collaboration skills.
- Fluent English, both written and spoken, with the ability to communicate directly with client technology and data teams.
The following experience would be advantageous:
Financial & Market Data
- Experience with time‑series and tick databases, such as kdb+ or similar columnar time‑series technologies.
- Familiarity with market‑data vendor schema APIs.
- Understanding of:
- Symbology
- Asset classes
- Market‑data structures and conventions
- Experience working with Microsoft SQL Server estates.
- Exposure to reporting and BI systems.
- Experience extracting inventories from application metadata tables.
- Familiarity with reporting and ETL architectures.
Modern Data Platforms
- Experience with Parquet and other columnar/lake formats.
- Experience with large object stores.
- Experience with orchestration platforms such as Apache Airflow or similar.
Data Governance & Platforms
- Working knowledge of graph databases.
- Understanding of data contracts and data‑quality frameworks.
- Experience with catalogue platforms such as DataHub, particularly on the ingestion side.
Regulated Environments
- Experience working in financial services, investment management, banking, or another highly regulated industry.
- Experience working with on‑premise enterprise data estates, security controls, and governance requirements.
Education
- Bachelor's degree in computer science, Engineering, Data Science, Information Technology, Mathematics, or a related technical discipline.
What Makes This Role Unique
This role sits at the intersection of data engineering, artificial intelligence, agentic systems, and financial services.
You will help solve a critical enterprise AI challenge: ensuring that AI agents can discover and reason over trusted, well‑described, traceable, current, and appropriately governed data.
The role provides an opportunity to:
- Take significant technical ownership of foundational AI/data capabilities.
- Work with large‑scale and complex financial data estates.
- Build systems that directly enable enterprise AI agents.
- Combine traditional data engineering with modern LLM and agent architectures.
- Work on metadata, catalogue, lineage, semantic, and entitlement capabilities.
- Operate in a sophisticated and highly regulated financial environment.
- Collaborate closely with architects, engineers, data specialists, and client stakeholders.
- Help shape how a global investment organization safely unlocks value from its data through AI.