Manchester, UK | Hybrid
ROLE PURPOSE
Design, build and operate production-grade AI and machine learning solutions across the full lifecycle. The role combines Generative AI and Agentic AI engineering with MLOps, scalable model serving, cloud infrastructure, production monitoring and responsible engineering controls.
Key responsibilities
- Design enterprise AI applications using Large Language Models, transformer architectures, Retrieval-Augmented Generation (RAG) and Agentic AI patterns such as ReAct.
- Build agents that use tools, reasoning, memory and workflow orchestration; integrate AI capabilities with enterprise platforms through secure APIs and microservices.
- Develop, evaluate and optimise machine learning and deep learning solutions using Python and PyTorch, taking work from experimentation through production.
- Create embedding, indexing, semantic retrieval and ranking pipelines for grounded AI responses and enterprise knowledge use cases.
- Package models using Docker and deploy through KServe, Vertex AI endpoints and Kubernetes-based serving platforms.
- Build CI/CD, continuous training and continuous monitoring workflows, including model evaluation and controlled promotion across environments.
- Operate feature and model registries, model versioning and reproducible release processes aligned to governance and risk controls.
- Monitor model quality, data quality, drift, latency, fairness signals, infrastructure health and service-level objectives.
- Optimise inference performance and cost through autoscaling, GPU scheduling, resource management and cloud-native architecture.
- Enable safe progressive delivery using canary, shadow and blue/green deployments, rollback controls and A/B testing.
- Collaborate with Product, Data Science, Platform, Architecture, Security and Risk teams; establish reusable patterns and engineering standards.
Mandatory skills and experience
Capability
Required experience
Core engineering
Strong production Python engineering, API development, automated testing and asynchronous or distributed service patterns.
Deep learning
Hands-on PyTorch experience and strong understanding of transformer architecture, inference and model evaluation.
Agentic AI
Experience with ReAct or comparable agent patterns, including tool calling, memory, reasoning and orchestration.
RAG & retrieval
Production RAG, embeddings, chunking, indexing, vector search, retrieval/ranking and grounding techniques.
Google Cloud
GCP knowledge with Vertex AI, cloud compute/storage and cloud databases; GKE experience is advantageous.
Containers & serving
Docker plus model packaging and serving using KServe, Vertex AI endpoints or equivalent Kubernetes-native platforms.
MLOps lifecycle
Scale & release
Cost-efficient autoscaling, GPU scheduling, canary and shadow deployment, rollback strategies and A/B testing.
Other experience
- LangChain, LangGraph, LlamaIndex or comparable orchestration frameworks.
- Kubernetes/GKE, Harness or equivalent delivery tooling, GitOps and infrastructure automation.
- Prometheus, Grafana, Dynatrace or similar observability platforms.
- Delivery in a regulated enterprise with security, privacy, model risk and responsible AI controls.
- Technical leadership, architecture reviews, mentoring and cross-functional stakeholder collaboration.
What success looks like
- AI services move safely from experiment to reliable, monitored production with repeatable delivery controls.
- Model serving is secure, resilient, scalable and cost-efficient, with measurable quality and operational performance.
- Reusable patterns improve delivery speed while supporting transparency, governance and risk management.