Job Purpose
Seeking a hands-on Data Engineer to design, build, and automate data pipelines that bring new datasets onto the platform quickly and reliably as they are identified by the investment team
Seeking a hands-on Data Engineer to design, build, and automate data pipelines that bring new datasets onto the platform quickly and reliably as they are identified by the investment team – Key priority is core data engineering: ingesting, transforming, and organising data so it lands in the right place, in the right shape. Over time, the role will extend into the firm's AI roadmap — preparing large document sets for retrieval-augmented generation (RAG), managing vector databases, and enabling effective querying across integrated datasets to generate insight and reporting. This is not a fixed-scope project role. The engineer will operate as a responsive, standing capability — picking up new pipeline builds as analysts surface new data requirements, while maintaining and improving what is already in place.
Key Responsibilities
- Pipeline development — Design, build, and maintain robust data pipelines to ingest new datasets onto the firm's data platform as demand arises from the investment and analyst teams.
- Automation — Automate ingestion, transformation, and data-quality workflows to reduce manual effort and speed up onboarding of new sources.
- Data platform — Structure and organise datasets so they can be integrated, joined, and queried effectively across the platform.
- AI enablement (roadmap) — Process large volumes of documents into formats suitable for AI consumption; store and index content in vector databases; support retrieval workflows (RAG) that surface the right information reliably.
- Insight & reporting (roadmap) — Enable effective querying across multiple integrated datasets to combine information and support sensible, decision-ready reporting.
- Stakeholder collaboration — Work directly with the hiring manager and analysts to scope new requirements, prioritize the pipeline backlog, and respond quickly as new needs emerge.
Key competencies
Required Skills
- 4–5+ years of hands-on data engineering experience, ideally within financial services or data intensive environments.
- Strong experience building and orchestrating data pipelines with Apache Airflow.
- Solid working knowledge of Apache Spark for large-scale data processing.
- Strong Python and SQL skills, with a focus on clean, maintainable, production-quality code.
- Experience designing data models and organizing datasets for downstream integration and querying.
- Track record of automating data workflows and building for reliability, monitoring, and data quality.
- Comfortable operating with loosely defined, evolving scope — able to self-organize, scope work with stakeholders, and deliver iteratively.
Nice to Have
- Exposure to AI/LLM engineering: document processing at scale, embeddings, and vector databases (e.g. pgvector, Pinecone, Weaviate, or similar).
- Experience building or supporting retrieval-augmented generation (RAG) workflows.
- Familiarity with cloud data platforms and modern data stack tooling.
- Prior experience supporting investment management, hedge fund, or capital markets data environments.
- Experience working as a dedicated remote resource embedded with a client team.