Data Scientist – AI Evaluation & Benchmarking Manager

Jobtailor

Manchester

On-site

GBP 60,000 - 90,000

Full time

3 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Jobtailor in Manchester is seeking a data science professional to design and run end-to-end benchmarking workflows, build scalable evaluation pipelines and translate findings into actionable insights. You will review literature, deploy ML workloads to cloud platforms, and contribute to production-grade experimentation infrastructure while supporting client demos and strategic technical direction.

The role requires hands-on data science skills, strong Python proficiency, and willingness to travel

Qualifications

  • Strong hands-on experience in Data Science concepts or LLM experimentation using structured evaluation frameworks.
  • Highly proficient in Python, including asynchronous programming, multithreading and writing maintainable code.
  • Experience deploying ML workloads to cloud platforms (Azure, AWS or GCP).
  • Familiarity with CI/CD and containerisation (Docker/ Podman).
  • Applied knowledge of statistics and experimental design.
  • Ability to translate findings into actionable recommendations.
  • Comfortable managing fastmoving workstreams and operating autonomously.
  • Emerging leadership behaviours, including taking initiative, influencing technical direction, communicating clearly and supporting the development of others.
  • Available for work visa sponsorship.
  • Up to 20% travel.

Responsibilities

  • Design and run end-to-end benchmarking workflows from understanding client use cases through generating business-ready insights
  • Build scalable evaluation frameworks, metrics and pipelines
  • Review academic literature to ensure evaluation strategies reflect leading research and best practice
  • Develop, maintain and improve experimentation infrastructure
  • Ensure robustness and production-grade engineering
  • Produce clear, client-ready insights
  • Support technical demos and deep-dive sessions
  • Influence AI model selection and technical strategy across PwC projects and client engagements

Skills

Data Science Concepts
LLM Experimentation
Python Programming
Async Programming
Multithreading
ML Deployment
Statistics
Experimental Design

Tools

Azure
AWS
GCP
Docker
Podman

Job description

  • Design and run end-to-end benchmarking workflows from understanding client use cases through generating business-ready insights
  • Build scalable evaluation frameworks, metrics and pipelines
  • Review academic literature to ensure evaluation strategies reflect leading research and best practice
  • Develop, maintain and improve experimentation infrastructure
  • Ensure robustness and production-grade engineering
  • Produce clear, client-ready insights
  • Support technical demos and deep-dive sessions
  • Influence AI model selection and technical strategy across PwC projects and client engagements
Requirements
  • Strong hands-on experience in Data Science concepts or LLM experimentation using structured evaluation frameworks
  • Highly proficient in Python, including asynchronous programming, multithreading and writing maintainable code
  • Experience deploying ML workloads to cloud platforms (Azure, AWS or GCP)
  • Familiarity with CI/CD and containerisation (Docker/ Podman)
  • Applied knowledge of statistics and experimental design
  • Ability to translate findings into actionable recommendations
  • Comfortable managing fastmoving workstreams and operating autonomously
  • Emerging leadership behaviours, including taking initiative, influencing technical direction, communicating clearly and supporting the development of others
  • Available for work visa sponsorship
  • Up to 20% travel
Core Competencies

Demonstrates expertise in Data Science concepts and LLM experimentation, with a strong focus on building scalable evaluation frameworks and deploying ML workloads on cloud platforms. Proficient in Python and capable of translating complex findings into actionable insights while managing fast-paced workstreams.

Highest-signal resume keywords
  • Data Science Concepts
  • Python Programming
  • ML Workload Deployment
  • CI/CD and Containerisation
  • Experimental Design
ATS Optimization Keywords
Hard Skills
  • Data Science
  • LLM Experimentation
  • Python
  • Asynchronous Programming
  • Multithreading
  • ML Workload Deployment
  • Statistics
  • Experimental Design
Soft Skills
  • Clear Communication
  • Initiative
  • Influencing Technical Direction
  • Supporting Development of Others
Industry Keywords
  • Benchmarking Workflows
  • Evaluation Frameworks
  • Client-Ready Insights
  • Technical Demos
  • Robustness
Tools & Technologies
  • Azure
  • AWS
  • GCP
  • Docker
  • Podman
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Scientist Consultant II
Data Scientist Consultant II

Jobtailor • Belfast City District

On-site
GBP 60,000 - 90,000
Data Scientist
Data Scientist

Jobtailor • City of Edinburgh

On-site
GBP 60,000 - 90,000
Senior AI Data Engineer
Senior AI Data Engineer

Jobtailor • Nottingham

On-site
GBP 70,000 - 110,000
AI Data Expert
AI Data Expert

Eliden • Greater London

On-site
GBP 90,000 - 130,000
Private healthcare and wellness initiatives
26 days annual leave plus bank holidays
Annual performance-based bonus
+1
Customer Facing AI Data Scientist - Systems Integrator
Customer Facing AI Data Scientist - Systems Integrator

Hamilton Barnes Associates Limited • Greater London

On-site
GBP 102,000 - 138,000
Software Engineer II, AI Enablement Team
Software Engineer II, AI Enablement Team

Jobtailor • Greater London

On-site
GBP 60,000 - 90,000
Principal Consultant – Innovation & Technology
Principal Consultant – Innovation & Technology

Jobtailor • Greater London

On-site
GBP 90,000 - 140,000
Lead Machine Learning, AI Engineer
Lead Machine Learning, AI Engineer

Jobtailor • Manchester

On-site
GBP 90,000 - 150,000
AI Evaluation & Benchmarking Lead, Data Science
AI Evaluation & Benchmarking Lead, Data Science

Jobtailor • Manchester

On-site
GBP 60,000 - 90,000
Manager, Analytics & AI, EY-Parthenon
Manager, Analytics & AI, EY-Parthenon

Ernst & Young • City of Westminster

On-site
GBP 90,000 - 130,000