Data Scientist - AI Evaluation & Benchmarking Manager

Pwc

Leeds

Hybrid

GBP 90,000 - 130,000

Full time

6 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

empowered flexibility
private medical cover
volunteering days

Job summary

PwC Leeds is seeking a Data Scientist - AI Evaluation & Benchmarking Manager to shape model benchmarking and experimentation capabilities. You will drive evidence-based decisions for client engagements and own scalable AI research infrastructure in a fast-moving, collaborative environment.

You will lead end-to-end benchmarking, design workflows, and produce client-ready insights while mentoring others and advancing technical strategy across PwC projects.

Qualifications

  • Hands-on data science or LLM experimentation experience.
  • Proficient in Python and scalable code practices.
  • Experience deploying ML workloads to cloud platforms and containerisation.
  • Knowledge of statistics and experimental design.
  • Ability to lead and influence technical direction.
  • Autonomous, able to manage fast-moving workstreams.

Responsibilities

  • Build and evolve AI benchmarking and experimentation platforms.
  • Influence AI model selection and technical strategy across PwC projects.
  • Design and run end-to-end benchmarking workflows.
  • Build scalable evaluation frameworks and pipelines.
  • Develop, maintain and improve experimentation infrastructure.
  • Provide client-ready insights and support product demos.

Skills

Python
Asynchronous programming
Multithreading
Data science concepts
Experiment design
Leadership
Communication

Tools

Docker
Podman
Azure
AWS
GCP
CI/CD

Job description

Data Scientist - AI Evaluation & Benchmarking Manager

2026-09-17T00:00:00

Leeds

West Yorkshire

GB

LS1 3

Any

2026-12-16T11:39:24

About The Role

You'll join our AI Research team as a Data Scientist - AI Evaluation & Benchmarking Manager, helping to shape and evolve the model benchmarking and experimentation capabilities that underpin AI delivery across PwC and our clients. You'll work within a highly collaborative applied research environment that values curiosity, technical rigour and practical problem solving. In this role, you'll develop frameworks and experimentation workflows used to evaluate emerging AI models, driving evidence-based decisions for client engagements. We're looking for someone who enjoys technical ownership, thrives in fastmoving environments and is motivated by building scalable, secure AI research infrastructure.

You'll join our AI Research team as a Data Scientist - AI Evaluation & Benchmarking Manager, helping to shape and evolve the model benchmarking and experimentation capabilities that underpin AI delivery across PwC and our clients. You'll work within a highly collaborative applied research environment that values curiosity, technical rigour and practical problem solving. In this role, you'll develop frameworks and experimentation workflows used to evaluate emerging AI models, driving evidence-based decisions for client engagements. We're looking for someone who enjoys technical ownership, thrives in fastmoving environments and is motivated by building scalable, secure AI research infrastructure.

What Your Days Will Look Like
  • You’ll play a key role in building and evolving our AI benchmarking and experimentation platforms, enabling robust and repeatable model evaluation.
  • Your work will directly influence AI model selection and technical strategy across PwC projects and client engagements.
  • Design and run end to end benchmarking workflows, from understanding client use cases, designing and running benchmarking strategies, and generating business ready insights
  • Build scalable evaluation frameworks, metrics and pipelines, combining hands‑on engineering with continuous review of academic literature to ensure our evaluation strategies reflect leading research and best practice.
  • Develop, maintain and improve experimentation infrastructure, ensuring robustness and production grade engineering.
  • Produce clear, client ready insights and support technical demos and deep dive sessions.
This Role Is For You If
  • You have strong hands‑on experience in Data Science concepts, or LLM experimentation using structured evaluation frameworks.
  • You are highly proficient in Python, including asynchronous programming, multithreading and writing maintainable code.
  • You have experience deploying ML workloads to cloud platforms (Azure, AWS or GCP), with familiarity in CI/CD and containerisation (Docker/ Podman).
  • You have applied knowledge of statistics and experimental design and can translate findings into actionable recommendations.
  • You are comfortable managing fastmoving workstreams and operating autonomously.
  • You demonstrate emerging leadership behaviours - taking initiative, influencing technical direction, communicating clearly and supporting the development of others.
  • You are motivated by ownership and excited by the opportunity to shape AI research platforms that directly impact client engagements.
What You'll Receive From Us

No matter where you may be in your career or personal life, our benefits are designed to add value and support, recognising and rewarding you fairly for your contributions.

We offer a range of benefits including empowered flexibility and a working week split between office, home and client site; private medical cover and 24/7 access to a qualified virtual GP; six volunteering days a year and much more.

  • empowered flexibility and a working week split between office, home and client site
  • private medical cover and 24/7 access to a qualified virtual GP
  • six volunteering days a year and much more

#J-18808-Ljbffr

#s1-Gen

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Evaluation & Benchmarking Lead — Data Scientist
AI Evaluation & Benchmarking Lead — Data Scientist

Pwc • Leeds

Hybrid
GBP 90,000 - 130,000
empowered flexibility
private medical cover
volunteering days
Data Engineer - Manager
Data Engineer - Manager

PwC UK • Manchester

On-site
GBP 70,000 - 100,000
Private medical cover
Flexible working
Volunteer days
Data Engineer - Manager
Data Engineer - Manager

PwC UK • Birmingham

On-site
GBP 90,000 - 140,000
Advanced Data Analytics
Advanced Data Analytics

PA Consulting • City of Westminster

On-site
GBP 50,000 - 75,000
Private healthcare
25 days annual leave
Generous pension scheme
+2
Data Analytics & AI Consultant - Marketing Expert
Data Analytics & AI Consultant - Marketing Expert

EPAM Systems • Greater London

Hybrid
GBP 90,000 - 130,000
ESPP
Life assurance
Private medical insurance
+4
Data Architect-Manager
Data Architect-Manager

PwC UK • Greater London

Hybrid
GBP 60,000 - 80,000
Private medical cover
Flexibility in working location
Volunteering days
Data Architect-Manager
Data Architect-Manager

PwC UK • City of Edinburgh

Hybrid
GBP 55,000 - 75,000
Flexible work arrangements
Private medical cover
Volunteering days
Data Architect-Manager
Data Architect-Manager

PwC UK • Belfast

Hybrid
GBP 60,000 - 80,000
Private medical cover
Flexible working week
Volunteering days
Data Architect-Manager
Data Architect-Manager

PwC UK • Manchester

Hybrid
GBP 60,000 - 80,000
Private medical cover
Flexible working week
Volunteering days
Data Architect-Manager
Data Architect-Manager

PwC UK • Leeds

Hybrid
GBP 60,000 - 80,000
Private medical cover
24/7 access to a virtual GP
Flexible working week