Senior AI Platform Engineer

Capgemini

Greater London

Hybrid

GBP 90,000 - 130,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Disability Confident Employer
Learning opportunities
Wellbeing programs

Job summary

Capgemini is seeking a production-focused Platform Engineer to build and operate AI infrastructure for regulated industries. You will own platform components like the model gateway, agent runtime, evaluation harness, and guardrail engine, collaborating with product teams to scale self-service services across multiple lines.

You will work on reinforcement learning readiness, model serving, and multi-tenant SaaS capabilities, contributing to a secure, observable, and cost-efficient cloud platform.

Qualifications

  • Strong production engineering: Python plus one systems language (Go or Rust welcome), Kubernetes, infrastructure-as-code, and one major cloud.
  • Hands-on experience with LLM infrastructure: a model gateway pattern (such as LiteLLM or in-house), model serving (such as vLLM or managed endpoints), and vector or retrieval systems.
  • Agent-systems experience: you have built with an agent orchestration framework (such as LangGraph or first-party agent SDKs) and understand tool calling and the Model Context Protocol (MCP)
  • Evaluation engineering: you have built or operated eval harnesses (golden datasets, regression gates in CI, LLM-judge calibration) and can explain why they gate merges
  • Post-training fluency: you have fine-tuned or post-trained a model, built the data pipeline behind one, or reproduced techniques from recent papers in production systems
  • Daily, hands-on use of AI coding assistants as part of your own development workflow

Responsibilities

  • One or more platform components end to end: model gateway (routing, failover, caching, per-line cost attribution), agent runtime and orchestration, evaluation harness, guardrail engine, tool and agent registry, or shared product services
  • The training and adaptation loop: pipelines that turn production traces and evaluation verdicts into reinforcement learning and fine-tuning datasets, and the infrastructure to train, evaluate, and serve NewCo-tuned LLMs and SLMs behind the same gates as vendor models
  • Reliability, latency, and cost of what you build; platform services carry baselines and you hold them
  • The self-service surfaces product engineers use: your components ship with documentation, sane defaults, and no ticket queue
  • Co-building with product teams: new shared services start embedded with a product line and graduate to platform services when proven

Skills

Python
Go
Rust
Kubernetes
IaC
LLM infra
Agent systems
Eval harness
Post-training
AI coding aids

Tools

LiteLLM
vLLM
LangGraph
pgvector
ClickHouse
OpenTelemetry
Grafana

Job description

About the job you're considering

Hybrid working: The places that you work from day to day will vary according to your role, your needs, and those of the business; it will be a blend of Company offices, client sites, and your home; noting that you will be unable to work at home 100% of the time.

If you are successfully offered this position, you will go through a series of pre-employment checks, including, identity, nationality (single or dual) or immigration status, employment history going back 3 continuous years, and unspent criminal record check (known as Disclosure and Barring Service)

About us

We build products, not projects: software for insurance claims, payment operations, and health operations, sold to banks, insurers, and health plans. Three product lines run on one shared platform, built by a deliberately small, senior team. Our engineering model is agentic: engineers author the specifications, tooling, evaluation suites, and guardrails, and AI agents do most of the implementation. Humans own every consequential decision, and in our regulated domains some decisions are human-only by design.

The role

You will build the platform floor that three AI products for banks, insurers, and health plans run on: the model gateway every inference call passes through, the agent runtime that executes planning loops and tool calls, the evaluation infrastructure that gates every release, the guardrail engine, and the shared services (case management, connectors, tenancy, metering) that stop three product teams building the same thing three times. This is production infrastructure for regulated industries, built by a small senior team with heavy AI leverage. The ambition runs past serving frontier models: we close the loop from production feedback through reinforcement learning and fine-tuning, and train our own LLMs and SLMs where evaluations and economics justify it.

What you will own
  • One or more platform components end to end: model gateway (routing, failover, caching, per-line cost attribution), agent runtime and orchestration, evaluation harness, guardrail engine, tool and agent registry, or shared product services
  • The training and adaptation loop: pipelines that turn production traces and evaluation verdicts into reinforcement learning and fine-tuning datasets, and the infrastructure to train, evaluate, and serve NewCo-tuned LLMs and SLMs behind the same gates as vendor models
  • Reliability, latency, and cost of what you build; platform services carry baselines and you hold them
  • The self-service surfaces product engineers use: your components ship with documentation, sane defaults, and no ticket queue
  • Co-building with product teams: new shared services start embedded with a product line and graduate to platform services when proven
What you will need
  • Strong production engineering: Python plus one systems language (Go or Rust welcome), Kubernetes, infrastructure-as-code, and one major cloud
  • Hands-on experience with LLM infrastructure: a model gateway pattern (such as LiteLLM or in-house), model serving (such as vLLM or managed endpoints), and vector or retrieval systems
  • Agent-systems experience: you have built with an agent orchestration framework (such as LangGraph or first-party agent SDKs) and understand tool calling and the Model Context Protocol (MCP)
  • Evaluation engineering: you have built or operated eval harnesses (golden datasets, regression gates in CI, LLM-judge calibration) and can explain why they gate merges
  • Post-training fluency: you have fine-tuned or post-trained a model, built the data pipeline behind one, or reproduced techniques from recent papers in production systems
  • Daily, hands-on use of AI coding assistants as part of your own development workflow

We care about depth in four or five of these areas more than surface familiarity with all of them.

What sets you apart
  • Guardrail-engine or AI-observability experience (NeMo Guardrails, OpenTelemetry GenAI conventions, LangSmith, Braintrust, or equivalent)
  • Reinforcement learning infrastructure (reward modelling, RLHF or GRPO-class pipelines) or SLM distillation experience
  • Multi-tenant SaaS infrastructure: tenancy isolation, metering, usage-based billing
  • Financial services engineering (banks, insurers, or payment providers) is the ideal background; other regulated-industry or security engineering depth also counts
  • Open-source contributions or a portfolio of deployed agents and eval suites (we would rather see this than a CV)
The reference stack

The reference technology stack for this role is our supported paved road: self-hosted LangSmith and LangGraph Platform as the agent runtime and evaluation plane, model providers behind a swappable gateway seam, PostgreSQL with pgvector plus ClickHouse and S3-compatible object storage as the data platform, Neo4j Enterprise as the semantic knowledge graph, an agent memory plane serving episodic and precedent memory over MCP, MCP-native connectors, OpenTelemetry and Grafana for observability, all on CNCF-conformant Kubernetes with Helm and Argo CD, deployable to any hyperscaler or on-prem. A tool-for-tool match is not expected: analogous experience counts fully. If you have built and operated systems of this shape on comparable components (a different orchestration framework, graph engine, evaluation platform, or serving stack), you have what we are looking for.

How we work
  • Engineers write specs, harnesses, evals, and guardrails; AI agents execute the implementation loops. Review, not typing, is where engineering judgment goes.
  • Three human gates govern everything we ship: spec approval, merge, and release. Regulated code paths (money movement, authentication, cryptography, secrets) are always human-owned.
  • Small and senior by design. No separate QA function, no scrum masters; quality comes from evaluation gates and whole-team review rituals.
  • Domain experts (claims practitioners, payment scheme experts, clinicians) are full-time members of the product teams you will serve.
Success in year one
  • A platform component you own is consumed self-service by at least two product lines, with measured adoption
  • Evaluation gates you built block real regressions before clients ever see them
  • Per-line cost attribution from the gateway is trusted enough to drive planning
  • Production traces and evaluation verdicts flow into curated adaptation datasets, and a tuned model serves traffic behind the same gates as vendor models
We are a Disability Confident Employer

Capgemini is proud to be a Disability Confident Employer (Level 2) under the UK Government's Disability Confident scheme.As part of our commitment to inclusive recruitment, we will offer an interview to all candidates who:

  • Declare they have a disability, and
  • Meet the minimum essential criteria for the role.

Please opt in during the application process.

Make it real - what does it mean for you?
  • We realise a Total Reward package should be more than just compensation. At Capgemini we offer range of core and flexible benefits and have a Peer Recognition Portal called Applaud.
  • You'd be joining an accredited Great Place to work for Wellbeing in 2024. Employee wellbeing is vitally important to us as an organisation. We see a healthy and happy workforce a critical component for us to achieve our organisational ambitions. To help support wellbeing we have trained 'Mental Health Champions' across each of our business areas, and we have invested in wellbeing apps such as Thrive and Peppy.
  • You will be empowered to explore, innovate, and progress. You will benefit from Capgemini's 'learning for life' mindset, meaning you will have countless training and development opportunities from thinktanks to hackathons, and access to 250,000 courses with numerous external certifications from AWS, Microsoft, Harvard ManageMentor, Cybersecurity qualifications and much more.

Capgemini. Make it real.

Why you should consider Capgemini

Growing clients' businesses while building a more sustainable, more inclusive future is a tough ask. When you join Capgemini, you'll join a thriving company and become part of a collective of free-thinkers, entrepreneurs and industry experts. We find new ways technology can help us reimagine what's possible. It's why, together, we seek out opportunities that will transform the world's leading businesses, and it's how you'll gain the experiences and connections you need to shape your future. By learning from each other every day, sharing knowledge, and always pushing yourself to do better, you'll build the skills you want. You'll use your skills to help our clients leverage technology to innovate and grow their business. So, it might not always be easy, but making the world a better place rarely is.

About Capgemini

Capgemini is an AI-powered global business and technology transformation partner, delivering tangible business value. We imagine the future of organisations and make it real with AI, technology, and people. With our strong heritage of nearly 60 years, we are a responsible and diverse group of 420,000 team members in more than 50 countries. We deliver end-to-end services and solutions with our deep industry expertise and strong partner ecosystem, leveraging our capabilities across strategy, technology, design, engineering and business operations. The Group reported 2024 global revenues of €22.1 billion.

Make it real

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Platform Engineer
AI Platform Engineer

Capgemini • Greater London

Hybrid
GBP 90,000 - 140,000
Applaud – peer recognition
Wellbeing program
Extensive training and certifications
+1
Principal AI Platform Engineer
Principal AI Platform Engineer

Capgemini • Greater London

On-site
GBP 120,000 - 190,000
AI Engineering Consultant / Senior Consultant
AI Engineering Consultant / Senior Consultant

Capgemini Invent • Manchester

Hybrid
GBP 90,000 - 120,000
AI Engineering Consultant / Senior Consultant
AI Engineering Consultant / Senior Consultant

Capgemini Invent • Greater London

Hybrid
GBP 90,000 - 140,000
GenAI Full Stack Engineer - Consultant / Senior Consultant - Digital Excellence
GenAI Full Stack Engineer - Consultant / Senior Consultant - Digital Excellence

Capgemini Invent • Manchester

Hybrid
GBP 90,000 - 130,000
Automation & Platform Engineer
Automation & Platform Engineer

Capgemini • Abingdon

Hybrid
GBP 70,000 - 100,000
Hybrid working model
Disability Confident Employer
Salesforce AI Engineer
Salesforce AI Engineer

Capgemini • Manchester

Hybrid
GBP 90,000 - 120,000
Hybrid working
Wellbeing support
Training & certifications
+1
AI Engineering Consultant / Senior Consultant
AI Engineering Consultant / Senior Consultant

Capgemini • Greater London

On-site
GBP 75,000 - 110,000
Hybrid working
Manager/Senior Manager - Data Management
Manager/Senior Manager - Data Management

Capgemini Invent • Manchester

Hybrid
GBP 110,000 - 165,000
Hybrid work model
Flexible benefits
Disability Confident Employer
OSS & Autonomous Network Engineer
OSS & Autonomous Network Engineer

CAPGEMINI • City Of London

Hybrid
GBP 90,000 - 120,000
Hybrid working model
Exposure to top global companies
Digital learning platforms