AI Platform & Site Reliability Engineering Managing Consultant

Capgemini

United Kingdom

Hybrid

GBP 90,000 - 130,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Hybrid working
Learning & certification programs
Wellbeing programs

Job summary

Capgemini Invent in the United Kingdom seeks an AI Platform & Site Reliability Engineering Managing Consultant to design, build and scale secure, reliable AI platforms and services. You will lead platform engineering across LLM/ML operations and governance, partnering with senior stakeholders to move from experimentation to production.

You will guide enterprise AI platforms, observability, and reliability while shaping operating models and governance essential for safe, scalable AI-enabled

Qualifications

  • Experience designing and operating cloud-native platforms for AI/ML
  • Strong understanding of LLMOps, MLOps and model lifecycle controls
  • Experience establishing and scaling SRE practices (SLIs/SLOs, incident management)
  • Knowledge of AI governance, regulatory requirements and risk management
  • Experience with observability platforms and data-driven reliability
  • Ability to advise senior stakeholders and lead transformation programmes
  • Experience across Azure, AWS and Google Cloud ecosystems
  • Ability to balance business outcomes with engineering constraints

Responsibilities

  • Design enterprise AI platform architectures including LLM and data platforms
  • Engineer AI platform capabilities: deployment pipelines, model management, observability and guardrails
  • Establish SRE practices: SLIs, SLOs, error budgets, capacity planning and reliability governance
  • Define observability strategies across applications and AI workloads
  • Apply reliability engineering to AI-enabled services with governance and risk controls
  • Collaborate with CIO/CTO/Engineering leaders on strategy and delivery
  • Lead client delivery teams and contribute to growth of Capgemini Invent capability

Skills

Cloud-native platforms
Platform engineering
SRE practices
Observability
LLMOps/MLOps
AI governance
Stakeholder advisory
Hyperscaler ecosystems

Tools

Azure
AWS
GCP
Datadog
Dynatrace
Splunk
Kubernetes

Job description

ABOUT CAPGEMINI

At Capgemini Invent, we believe difference drives change. As inventive transformation consultants, we blend our strategic, creative and scientific capabilities, collaborating closely with clients to deliver cutting-edge solutions. Join us to drive transformation tailored to our client's challenges of today and tomorrow. Informed and validated by science and data. Superpowered by creativity and design. All underpinned by technology created with purpose.

YOUR ROLE

As an AI Platform & Site Reliability Engineering Managing Consultant, you will help clients design, build and scale secure, reliable and operationally effective AI platforms. You will combine expertise in platform engineering, Site Reliability Engineering (SRE), observability and intelligent operations to help organisations move from isolated AI experimentation to production-grade, enterprise-scale AI services. You will work with technology, engineering, operations and business leaders to establish the platforms, operating models, governance and reliability practices required to run AI-enabled services safely, effectively and at scale. Acting as a trusted advisor to senior stakeholders, you will shape client strategy while leading delivery teams and helping grow our AI Platform & Reliability Engineering capability.

This will include:

  • AI Platform Strategy & Architecture: Assess, define and evolve enterprise AI platform architectures, covering LLM and agentic frameworks, AI gateways, model lifecycle management, data platforms, MLOps/LLMOps foundations and integration patterns. Support clients in evaluating build, buy and hybrid approaches aligned to business needs, risk appetite and operational requirements.
  • AI Platform Engineering & LLMOps: Design and implement scalable AI platform capabilities including model deployment pipelines, prompt and model management, evaluation frameworks, AI observability, platform automation and operational guardrails. Enable reliable and repeatable delivery of AI services from experimentation through to production.
  • Reliability Engineering & SRE: Establish SRE practices including SLIs, SLOs, error budgets, capacity planning, resilience engineering and reliability governance. Help clients shift from reactive operations to data-driven reliability management while balancing reliability, innovation and delivery velocity.
  • Observability & Operational Intelligence: Define observability strategies across applications, platforms and AI workloads using metrics, logs, traces and telemetry. Establish operational insight models that support proactive decision-making and enable advanced capabilities including anomaly detection, event intelligence, noise reduction and predictive operational analytics.
  • AI Operations & Service Reliability: Apply reliability engineering principles to AI-enabled services, monitoring AI-specific failure modes such as data quality degradation, hallucination patterns, token consumption, agent reliability and model performance drift. Implement controls, feedback loops and automated guardrails to ensure AI services remain secure, trusted and cost-effective.
  • Responsible AI & Platform Governance: Design and embed AI governance frameworks, model risk controls, compliance measures and responsible AI practices that address security, regulatory and ethical requirements while supporting innovation and adoption at scale.
  • Automation & Operational Efficiency: Identify opportunities to reduce operational complexity and toil through engineering-led automation, intelligent workflows and AI-enhanced operational practices. Help clients improve scalability, consistency and operational performance while reducing manual effort.
  • Client Advisory & Transformation Leadership: Act as a trusted advisor to CIO, CTO, CDO and Engineering leadership stakeholders, shaping platform strategies, operating models, vendor selections and transformation roadmaps. Lead consulting teams and workstreams from assessment and strategy through implementation and scale-up.
YOUR PROFILE

Essential Experience

  • Proven experience designing, delivering and operating cloud-native, platform engineering, AI platform or reliability engineering solutions within complex enterprise environments.
  • Strong understanding of AI platform architectures including LLMOps, MLOps, agentic AI frameworks, model lifecycle management and AI operational controls.
  • Experience establishing and scaling SRE practices including observability, SLIs, SLOs, error budgets, incident management and reliability engineering.
  • Strong understanding of AI governance, responsible AI, regulatory requirements and model risk management.
  • Experience implementing observability strategies using modern monitoring, telemetry and operational analytics platforms.
  • Demonstrated ability to advise senior stakeholders and lead multidisciplinary transformation programmes.
  • Experience working across hyperscaler ecosystems including Azure, AWS and Google Cloud Platform.
  • Proven ability to balance business outcomes, user needs, engineering constraints and operational requirements when shaping platform strategies.

Desirable Experience

  • Experience with AI observability, model monitoring, AI governance tooling or AI platform operations.
  • Experience of platform engineering, DevSecOps, automation and Infrastructure-as-Code practices.
  • Experience developing propositions, leading bids and supporting business growth activities.
  • Active participation in AI, SRE, platform engineering or cloud communities.

Certifications (Desirable)

  • Azure AI Engineer Associate
  • Azure Solutions Architect Expert
  • AWS Machine Learning Specialty
  • Google Professional Cloud Architect
  • Certified Kubernetes Administrator (CKA)
  • SRE Foundation or SRE Practitioner
  • Relevant observability platform certifications (Datadog, Dynatrace, Splunk etc.)
Security Check (SC) Clearance

To be successfully appointed to this role, it is a requirement to obtain Security Check (SC) clearance. (https://www.gov.uk/guidance/united-kingdom-security-vetting-applicant#levels-of-national-security-clearance)

To obtain SC clearance, the successful applicant must have resided continuously within the United Kingdom for the last 5 years, along with other criteria and requirements.

Throughout the recruitment process, you will be asked questions about your security clearance eligibility such as, but not limited to, country of residence and nationality. Some posts are restricted to sole UK Nationals for security reasons; therefore, you may be asked about your citizenship in the application process.

WHAT YOU'LL LOVE ABOUT WORKING HERE
  • Client engagements give you the opportunity to work with our leadership and experienced consulting management, where you can learn from them, challenge them, and accelerate your hands‑on experience, delivery capability and industry insights.
  • You’ll learn how we write compelling client propositions, structure, and lead high‑profile transformation, and gain hands‑on exposure to leading technologies, often taking an idea from a concept to a vision, to strategy and then execution.
  • Our consultants are formally trained from industry experts on management consulting and client delivery.
  • We provide a host of opportunities for learning and certification through internal and partner led programmes from AWS, Google and Microsoft.
  • Les Fontaines: Capgemini Invent has a unique training environment just outside of Paris, where we can immerse ourselves in thought‑leadership, share knowledge and build capabilities which will help us and our clients to succeed.
  • We hold monthly showcases of our initiatives, sharing knowledge and showing off how the power of technology is impacting our clients.
  • There are many opportunities like monthly team drinks to connect face‑to‑face with the wider team over a few drinks in the city. There are regular leadership connect sessions that will give you a different, more relaxed setting to meet up in the office to hear from the leadership, meet colleagues and discuss the trends and insights within the market. Team away days are always a chance to connect with the team, have fun and learn something new.
WHAT YOU NEED TO KNOW

At Capgemini we don’t just believe in inclusion, we actively go out to making it a working reality. Driven by our core values and Inclusive Futures for All campaign, we build environments where you can bring you whole self to work.

We aim to build an environment where employees can enjoy a positive work-life balance.

We embed hybrid working in all that we do and make flexible working arrangements the day‑to‑day reality for our people.

All UK employees are eligible to request flexible working arrangements.

Employee wellbeing is vitally important to us as an organisation.

We see a healthy and happy workforce a critical component for us to achieve our organisational ambitions.

To help support wellbeing we have trained 'Mental Health Champions' across each of our business areas.

We have also invested in wellbeing apps such as Thrive and Peppy.

Whilst you will have London, Manchester or Glasgow as an office base location, you must be fully flexible in terms of assignment location, as these roles may involve periods of time away from home at short notice.

We offer a remuneration package which includes flexible benefits options for you to choose to suit your own personal circumstances and a variable element dependent grade and on company and personal performance.

ABOUT CAPGEMINI

Capgemini is a global business and technology transformation partner, helping organizations to accelerate their dual transition to a digital and sustainable world, while creating tangible impact for enterprises and society. It is a responsible and diverse group of 340,000 team members in more than 50 countries. With its strong over 55-year heritage, Capgemini is trusted by its clients to unlock the value of technology to address the entire breadth of their business needs. It delivers end‑to‑end services and solutions leveraging strengths from strategy and design to engineering, all fuelled by its market leading capabilities in AI, cloud and data, combined with its deep industry expertise and partner ecosystem. The Group reported 2023 global revenues of €22.5 billion.

WE'RE A DISABILITY CONFIDENT EMPLOYER

Capgemini is proud to be a Disability Confident Employer (Level 2) under the UK Government’s Disability Confident scheme. As part of our commitment to inclusive recruitment, we will offer an interview to all candidates who:

  • Declare they have a disability,
  • Meet the minimum essential criteria for the role.

Please opt in during the application process.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Platform & Site Reliability Engineering Consultant/Senior Consultant - Cloud Operating Model
AI Platform & Site Reliability Engineering Consultant/Senior Consultant - Cloud Operating Model

Capgemini • United Kingdom

Hybrid
GBP 75,000 - 110,000
Cloud Operating Model - Managing Consultant
Cloud Operating Model - Managing Consultant

Capgemini • Greater London

Hybrid
GBP 90,000 - 130,000
AI Platform & Site Reliability Engineering Consultant/Senior Consultant - Cloud Operating Model
AI Platform & Site Reliability Engineering Consultant/Senior Consultant - Cloud Operating Model

Capgemini • Greater London

Hybrid
GBP 70,000 - 100,000
Hybrid working
Certifications & training
Wellbeing support
+1
AI Platform & Site Reliability Engineering Consultant/Senior Consultant - Cloud Operating Model
AI Platform & Site Reliability Engineering Consultant/Senior Consultant - Cloud Operating Model

Capgemini Invent • Greater London

Hybrid
GBP 70,000 - 110,000
Learning & development
Hybrid working
Career progression
AI Engineering Consultant / Senior Consultant
AI Engineering Consultant / Senior Consultant

Capgemini • Greater London

On-site
GBP 75,000 - 110,000
Hybrid working
Consultant/Senior Consultant - Automation
Consultant/Senior Consultant - Automation

Capgemini Invent • Greater London

Hybrid
GBP 75,000 - 110,000
Operational Excellence - Consultant/Senior Consultant
Operational Excellence - Consultant/Senior Consultant

Capgemini Invent • Greater London

Hybrid
GBP 65,000 - 90,000
Hybrid working
Managing Consultant - Automation
Managing Consultant - Automation

Capgemini • Greater London

On-site
GBP 90,000 - 130,000
Hybrid working
Disability Confident Employer
Intelligent Industry Consulting Strategy & Transformation - Director
Intelligent Industry Consulting Strategy & Transformation - Director

Capgemini Invent • Greater London

Hybrid
GBP 120,000 - 190,000
Flexible benefits options
Hybrid working
AI Assurance - Senior Consultant
AI Assurance - Senior Consultant

Capgemini Invent • Manchester

Hybrid
GBP 90,000 - 120,000