Job Summary
The Principal AI Engineer is responsible for designing, building, and operating the data pipelines and data infrastructure that power AI, machine learning, and agentic applications across the Bank. The role owns the end-to-end data lifecycle from ingestion and integration through transformation, semantic enrichment, and retrieval-ready delivery.
This role partners closely with data, engineering, AI, and platform teams to create scalable, reliable, and governed data foundations that support AI-driven products and intelligent automation. The successful candidate combines deep expertise in large-scale data engineering, distributed processing, semantic modelling, and modern data architecture with the ability to design systems that balance scalability, performance, resilience, and cost.
Key Responsibilities
Strategy
- Drive the design and evolution of the Bank's AI data architecture and data engineering standards.
- Define scalable patterns for ingestion, transformation, enrichment, storage, and retrieval of structured and unstructured data.
- Contribute to enterprise data standards, semantic modelling frameworks, and reusable architecture blueprints supporting AI adoption.
- Partner with architecture, platform, and AI teams to shape the future data ecosystem for AI and agentic applications.
Business
- Design, build, and operate data pipelines supporting AI, machine learning, analytics, and agent-based use cases.
- Deliver trusted, high-quality datasets that support model training, retrieval, feature generation, and operational AI workloads.
- Enable integration across databases, APIs, event streams, cloud platforms, document repositories, and external data sources.
- Collaborate with AI engineers, data scientists, architects, and business stakeholders to deliver scalable data solutions.
Processes
- Develop and maintain batch and streaming data pipelines using distributed processing technologies.
- Design and optimise data workflows running on lakehouse platforms and distributed compute environments.
- Build retrieval-ready and feature-ready datasets for AI and machine learning applications.
- Develop semantic layers, ontologies, taxonomies, and knowledge representations that improve data discovery and AI reasoning.
- Design solutions leveraging graph databases, vector stores, document stores, and other non-relational technologies.
- Implement data quality controls, lineage, observability, monitoring, and operational support processes.
- Support ingestion, parsing, chunking, enrichment, and normalisation of unstructured and multimodal content.
Risk Management
- Apply security, governance, privacy, and data protection controls throughout the data lifecycle.
- Ensure data engineering solutions align with enterprise governance, regulatory requirements, and operational resilience standards.
- Promote strong data quality, traceability, auditability, and reliability practices across data platforms.
Governance
- Contribute to enterprise standards for AI-ready data architecture and engineering practices.
- Maintain documentation, architecture artefacts, engineering standards, and operational runbooks.
- Support architecture reviews and technology governance processes.
Regulatory & Business Conduct
- Display exemplary conduct and live by the Group's Values and Code of Conduct.
- Take personal responsibility for embedding the highest standards of ethics, regulatory compliance, and business conduct.
- Effectively identify, assess, elevate, mitigate, and resolve risk, conduct, and compliance matters associated with data platforms and AI solutions.
- Ensure data engineering solutions align with enterprise governance, regulatory requirements, and operational resilience standards.
- Promote strong data quality, traceability, auditability, and reliability practices across data platforms.
Our Ideal Candidate
- 12+ years of experience in data engineering, software engineering, data platforms, or distributed systems, with a strong track record of delivering enterprise-scale data solutions.
- Proven experience building & operating large-scale batch and streaming data pipelines supporting critical business workloads.
- Strong expertise in Spark, Databricks, or equivalent distributed compute platforms, including performance optimisation, workload tuning, and operational support.
- Experience designing data platforms using lakehouse architectures, object storage, and modern data formats including Parquet, Delta Lake, or Iceberg.
- Strong understanding of data modelling and storage technologies, including graph databases, vector databases, document stores, key-value stores, and wide-column databases.
- Experience building data integration solutions across databases, APIs, event streams, files, and cloud-native platforms.
- Knowledge of ontology development, semantic modelling, taxonomies, RDF, knowledge graphs, or property graph technologies.
- Experience processing an