ML/AI Engineer — Kubernetes, CI/CD & GPU Inference

Lloyds Banking Group

Manchester

Hybrid

GBP 73,000 - 81,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Pension up to 15%
Annual bonus
Share schemes
Discounted shopping
30 days holiday
Wellbeing initiatives
Parental leave

Job summary

Lloyds Banking Group in Manchester is seeking an hands‑on ML/AI Engineer to build, automate, and maintain scalable ML systems across the lifecycle.

You will lead Kubernetes orchestration, CI/CD automation with Harness, GPU optimization, and model deployment in production, with a focus on reliability and observability.

A collaborative hybrid role with 2 days per week in the Manchester office offers opportunities to influence AI governance, fairness, and scalable decisions.

Qualifications

  • Strong Python for automation, tooling, and service development.
  • Deep expertise in Kubernetes, Docker, Helm, operators, node‑pool management, and autoscaling.
  • CI/CD expertise having hands‑on experience with Harness (or similar) building multi‑stage pipelines; experience with GitOps, artefact repositories, and environment promotion.
  • Practical experience with CUDA, TensorRT, Triton, TorchServe, and GPU scheduling/optimisation.
  • Proficiency in Prometheus, Grafana, Dynatrace defining SLIs/SLOs and alert thresholds for ML systems.
  • Experience operating MLflow (or equivalent) for experiment tracking, model bundling, and deployments.

Responsibilities

  • Compose, build, and operate production‑grade Kubernetes clusters for high‑volume model inference and scheduled training jobs.
  • Configure autoscaling, resource quotas, GPU/CPU node pools, service mesh, Helm charts, and custom operators to meet reliability and efficiency targets.
  • Implement GitOps workflows for environment configuration and application releases.
  • Build CI/CD pipelines in Harness (or equivalent) to automate build, test, model packaging, and deployment across environments (dev / pre‑prod / prod).
  • Enable progressive delivery (blue/green, canary) and rollback strategies, integrating quality gates, unit/integration tests, and model‑evaluation checks.
  • Standardise pipelines for continuous training (CT) and continuous monitoring (CM) to keep models fresh and safe in production.
  • Deploy and tune GPU‑backed inference services (e.g., A100), optimise CUDA environments, and leverage TensorRT where appropriate.
  • Operate scalable serving frameworks (NVIDIA Triton, TorchServe) with attention to latency, efficiency, resilience, and cost.
  • Implement end‑to‑end observability for models and pipelines: drift, data quality, fairness signals, latency, GPU utilisation, error budgets, and SLOs/SLIs via Prometheus, Grafana, and Dynatrace.
  • Establish actionable alerting and runbooks for on‑call operations; drive incident reviews and reliability improvements.
  • Operate a model registry (e.g., MLflow) with experiment tracking, versioning, lineage, and environment‑specific artefacts.
  • Enforce audit readiness: model cards, reproducible builds, provenance, and controlled promotion between stages

Skills

Python
Kubernetes
Docker
Harness CI/CD
GPU optimization
Prometheus
Grafana
Dynatrace
MLflow
Git

Tools

Harness
NVIDIA Triton
TorchServe
CUDA
TensorRT
Prometheus
Grafana
Dynatrace
MLflow

Job description

Lloyds Banking Group in Manchester is seeking an hands‑on ML/AI Engineer to build, automate, and maintain scalable ML systems across the lifecycle.

You will lead Kubernetes orchestration, CI/CD automation with Harness, GPU optimization, and model deployment in production, with a focus on reliability and observability.

A collaborative hybrid role with 2 days per week in the Manchester office offers opportunities to influence AI governance, fairness, and scalable decisions.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

ML/AI Engineer: Scalable ML Ops & GPU Inference
ML/AI Engineer: Scalable ML Ops & GPU Inference

Lloyds Bank plc • Manchester

Hybrid
GBP 73,000 - 81,000
Hybrid Working
Job Share
Pension contribution up to 15%
+5
ML/AI Engineer
ML/AI Engineer

Lloyds Bank plc • Manchester

On-site
GBP 73,000 - 81,000
Hybrid Working
Job Share
Pension contribution up to 15%
+5
Senior ML & AI Engineering Lead — Hybrid (London/Manchester)
Senior ML & AI Engineering Lead — Hybrid (London/Manchester)

Lloyds Banking Group • Greater London

Hybrid
GBP 120,000 - 180,000
ML/AI Engineer
ML/AI Engineer

Lloyds Banking Group • Manchester

Hybrid
GBP 73,000 - 81,000
Pension up to 15%
Annual bonus
Share schemes
+4
AI/ML Engineer
AI/ML Engineer

Queen Square Recruitment Ltd • Manchester

Hybrid
GBP 65,000 - 95,000
Hybrid working model
12-month contract
Enterprise-scale AI projects
+1
AI Platform Lead: ML, LLMs & Responsible AI at Scale
AI Platform Lead: ML, LLMs & Responsible AI at Scale

Lloyds Banking Group • Manchester

On-site
GBP 120,000 - 180,000
Staff Software Engineer, Kubernetes-native GPU Inference
Staff Software Engineer, Kubernetes-native GPU Inference

Together AI • Greater London

Hybrid
GBP 100,000 - 160,000
Senior ML Engineer: Production-Grade AI Systems
Senior ML Engineer: Production-Grade AI Systems

Lloyds Bank plc • West of England, Chester, Manchester

Hybrid
GBP 73,000 - 81,000
A generous pension contribution
Annual performance bonus
Share schemes including free shares
+3
MLOps Engineer: Lead Scalable AI Infra (Hybrid London)
MLOps Engineer: Lead Scalable AI Infra (Hybrid London)

Harnham • Greater London

Hybrid
GBP 75,000 - 85,000
Bonus up to 10%
Private healthcare
Hybrid work London
ML Engineering Manager — Lead Scalable AI Systems
ML Engineering Manager — Lead Scalable AI Systems

Compare the Market • Peterborough

Hybrid
GBP 110,000 - 140,000