ML/AI Engineer

Lloyds Banking Group

Manchester

Hybrid

GBP 73,000 - 81,000

Full time

7 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Pension up to 15%
Annual bonus
Share schemes
Discounted shopping
30 days holiday
Wellbeing initiatives
Parental leave

Job summary

Lloyds Banking Group in Manchester is seeking an hands‑on ML/AI Engineer to build, automate, and maintain scalable ML systems across the lifecycle.

You will lead Kubernetes orchestration, CI/CD automation with Harness, GPU optimization, and model deployment in production, with a focus on reliability and observability.

A collaborative hybrid role with 2 days per week in the Manchester office offers opportunities to influence AI governance, fairness, and scalable decisions.

Qualifications

  • Strong Python for automation, tooling, and service development.
  • Deep expertise in Kubernetes, Docker, Helm, operators, node‑pool management, and autoscaling.
  • CI/CD expertise having hands‑on experience with Harness (or similar) building multi‑stage pipelines; experience with GitOps, artefact repositories, and environment promotion.
  • Practical experience with CUDA, TensorRT, Triton, TorchServe, and GPU scheduling/optimisation.
  • Proficiency in Prometheus, Grafana, Dynatrace defining SLIs/SLOs and alert thresholds for ML systems.
  • Experience operating MLflow (or equivalent) for experiment tracking, model bundling, and deployments.

Responsibilities

  • Compose, build, and operate production‑grade Kubernetes clusters for high‑volume model inference and scheduled training jobs.
  • Configure autoscaling, resource quotas, GPU/CPU node pools, service mesh, Helm charts, and custom operators to meet reliability and efficiency targets.
  • Implement GitOps workflows for environment configuration and application releases.
  • Build CI/CD pipelines in Harness (or equivalent) to automate build, test, model packaging, and deployment across environments (dev / pre‑prod / prod).
  • Enable progressive delivery (blue/green, canary) and rollback strategies, integrating quality gates, unit/integration tests, and model‑evaluation checks.
  • Standardise pipelines for continuous training (CT) and continuous monitoring (CM) to keep models fresh and safe in production.
  • Deploy and tune GPU‑backed inference services (e.g., A100), optimise CUDA environments, and leverage TensorRT where appropriate.
  • Operate scalable serving frameworks (NVIDIA Triton, TorchServe) with attention to latency, efficiency, resilience, and cost.
  • Implement end‑to‑end observability for models and pipelines: drift, data quality, fairness signals, latency, GPU utilisation, error budgets, and SLOs/SLIs via Prometheus, Grafana, and Dynatrace.
  • Establish actionable alerting and runbooks for on‑call operations; drive incident reviews and reliability improvements.
  • Operate a model registry (e.g., MLflow) with experiment tracking, versioning, lineage, and environment‑specific artefacts.
  • Enforce audit readiness: model cards, reproducible builds, provenance, and controlled promotion between stages

Skills

Python
Kubernetes
Docker
Harness CI/CD
GPU optimization
Prometheus
Grafana
Dynatrace
MLflow
Git

Tools

Harness
NVIDIA Triton
TorchServe
CUDA
TensorRT
Prometheus
Grafana
Dynatrace
MLflow

Job description

End Date

Friday 09 October 2026

Salary Range

£72,702 - £80,780

We support flexible working – click here for more information on flexible working options

Flexible Working Options

Hybrid Working, Job Share

Job Description

JOB TITLE: ML/AI Engineer

SALARY: £72,702 - £80,000 per annum

LOCATION: Manchester

HOURS: Full-time – 35 hours

WORKING PATTERN: Our work style is hybrid, which involves spending at least two days per week, or 40% of our time, at our Manchester office.

About this opportunity…

Exciting opportunity for a hands‑on ML/AI Engineer to join our Data & AI Engineering team. You’ll build, automate, and maintain scalable systems that support the full machine learning lifecycle. You will lead Kubernetes orchestration, CI/CD automation (including Harness), GPU optimisation, and large‑scale model deployment, owning the path from code commit to reliable, monitored production services

This is a unique opportunity to shape the future of AI by embedding fairness, transparency, and accountability at the heart of innovation. You’ll join us at an exciting time as we move into the next phase of our transformation. We’re looking for curious, passionate engineers who thrive on innovation and want to make a real impact.

About us…

We’re on an exciting journey and there couldn’t be a better time to join us. The investments we’re making in our people, data, and technology are leading to innovative projects, fresh possibilities, and countless new ways for our people to work, learn, and thrive.

What you’ll do…
  • Compose, build, and operate production‑grade Kubernetes clusters for high‑volume model inference and scheduled training jobs.
  • Configure autoscaling, resource quotas, GPU/CPU node pools, service mesh, Helm charts, and custom operators to meet reliability and efficiency targets.
  • Implement GitOps workflows for environment configuration and application releases.
  • Build CI/CD pipelines in Harness (or equivalent) to automate build, test, model packaging, and deployment across environments (dev / pre‑prod / prod).
  • Enable progressive delivery (blue/green, canary) and rollback strategies, integrating quality gates, unit/integration tests, and model‑evaluation checks.
  • Standardise pipelines for continuous training (CT) and continuous monitoring (CM) to keep models fresh and safe in production.
  • Deploy and tune GPU‑backed inference services (e.g., A100), optimise CUDA environments, and leverage TensorRT where appropriate.
  • Operate scalable serving frameworks (NVIDIA Triton, TorchServe) with attention to latency, efficiency, resilience, and cost.
  • Implement end‑to‑end observability for models and pipelines: drift, data quality, fairness signals, latency, GPU utilisation, error budgets, and SLOs/SLIs via Prometheus, Grafana, and Dynatrace.
  • Establish actionable alerting and runbooks for on‑call operations; drive incident reviews and reliability improvements.
  • Operate a model registry (e.g., MLflow) with experiment tracking, versioning, lineage, and environment‑specific artefacts.
  • Enforce audit readiness: model cards, reproducible builds, provenance, and controlled promotion between stages
What you’ll need…
  • Strong Python for automation, tooling, and service development.
  • Deep expertise in Kubernetes, Docker, Helm, operators, node‑pool management, and autoscaling.
  • CI/CD expertise having hands‑on experience with Harness (or similar) building multi‑stage pipelines; experience with GitOps, artefact repositories, and environment promotion.
  • Practical experience with CUDA, TensorRT, Triton, TorchServe, and GPU scheduling/optimisation.
  • Proficiency in Prometheus, Grafana, Dynatrace defining SLIs/SLOs and alert thresholds for ML systems.
  • Experience operating MLflow (or equivalent) for experiment tracking, model bundling, and deployments.
  • Expert use of Git, branching models, protected merges, and code‑review workflows.
It would be great if you had any of the following…
  • Experience with GCP (e.g., GKE, Cloud Run, Pub/Sub, BigQuery) and Vertex AI (Endpoints, Pipelines, Model Monitoring, Feature Store).
  • Hooks for prompt/version management, offline/online evaluation, and human‑in‑the‑loop workflows (e.g., RLHF) to enable continuous improvement.
  • Familiarity with Model Context Protocol (MCP) for tool interoperability, plus Google ADK, LangGraph/LangChain for agent orchestration and multi‑agent patterns.
  • Ray, Kubeflow, or similar frameworks.
  • Experience embedding controls, audit evidence, and governance in regulated environments.
  • Experience with GPU efficiency, autoscaling strategies, and workload right‑sizing.
About working for us…

Our focus is to ensure we're inclusive every day, building an organisation that reflects modern society and celebrates diversity in all its forms. We want our people to feel that they belong and can be their best, regardless of background, identity or culture. And it’s why we especially welcome applications from under‑represent‑ed groups. We’re disability confident. So if you’d like reasonable adjustments to be made to our recruitment processes, just let us know.

We also offer a wide-ranging benefits package, which includes…
  • A generous pension contribution of up to 15%
  • An annual bonus award, subject to Group performance
  • Share schemes including free shares
  • Benefits you can adapt to your lifestyle, such as discounted shopping
  • 30 days’ holiday, with bank holidays on top
  • A range of wellbeing initiatives and generous parental leave policies

At Lloyds Banking Group, we're driven by a clear purpose; to help Britain prosper. Across the Group, our colleagues are focused on making a difference to customers, businesses and communities. With us you'll have a key role to play in shaping the financial services of the future, whilst the scale and reach of our Group means you'll have many opportunities to learn, grow and develop.

We keep your data safe. So, we'll only ever ask you to provide confidential or sensitive information once you have formally been invited along to an interview or accepted a verbal offer to join us which is when we run our background checks. We'll always explain what we need and why, with any request coming from a trusted Lloyds Banking Group person.

We're focused on creating a values-led culture and are committed to building a workforce which reflects the diversity of the customers and communities we serve. Together we’re building a truly inclusive workplace where all of our colleagues have the opportunity to make a real difference.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

ML/AI Engineer
ML/AI Engineer

Lloyds Bank plc • Manchester

On-site
GBP 73,000 - 81,000
Hybrid Working
Job Share
Pension contribution up to 15%
+5
Principal Engineer – Advanced AI Engineering
Principal Engineer – Advanced AI Engineering

Lloyds Banking Group • West of England

Hybrid
GBP 85,000 - 127,000
A generous pension contribution
Annual bonus
Share schemes
+3
Senior Lead Engineer - Advanced Engineering
Senior Lead Engineer - Advanced Engineering

Lloyds Banking Group • West of England

On-site
GBP 85,000 - 127,000
Pension up to 15%
Annual bonus
Share schemes
+3
Machine Learning & AI Engineering Lead
Machine Learning & AI Engineering Lead

Lloyds Banking Group • Greater London

On-site
GBP 120,000 - 180,000
Senior Machine Learning Engineer
Senior Machine Learning Engineer

Lloyds Banking Group • West of England

Hybrid
GBP 73,000 - 81,000
Generous pension with up to 15%
Annual bonus
Share schemes (free shares)
+2
Machine Learning & AI Engineering Lead
Machine Learning & AI Engineering Lead

Lloyds Banking Group • Manchester

On-site
GBP 120,000 - 180,000
Data Science & AI Graduate Scheme (Manchester)
Data Science & AI Graduate Scheme (Manchester)

Lloyds Bank plc • Manchester

On-site
GBP 41,000 - 56,000
Data Science and AI Industrial Placement (Manchester)
Data Science and AI Industrial Placement (Manchester)

Lloyds Bank plc • Manchester

On-site
GBP 26,000 - 32,000
Lead AI & Machine Learning Engineer
Lead AI & Machine Learning Engineer

Lloyds Banking Group • Leeds

On-site
GBP 60,000 - 80,000
Generous pension contribution
Performance-related bonus
Share schemes
+3
Data Science and AI Industrial Placement (Bristol)
Data Science and AI Industrial Placement (Bristol)

Lloyds Bank plc • West of England

On-site
GBP 25,000 - 33,000
Hybrid working policy