AI Evaluations Engineer

ConnexAI

Manchester

On-site

GBP 50,000 - 70,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

ConnexAI in Manchester is looking for a professional to measure and improve AI systems in production. This role focuses on defining performance metrics for LLMs, ASR, and TTS, and requires strong Python skills. Responsibilities include designing evaluation tests, building datasets, and collaborating with product and engineering teams to ensure AI improvements translate into better user experiences. Ideal candidates should have a solid understanding of AI quality measurement and experience in machine learning environments.

Qualifications

  • Strong Python skills for scripting and data analysis.
  • Experience with machine learning systems, particularly evaluation.
  • Understanding of LLMs or speech systems like ASR and TTS.

Responsibilities

  • Design and run evaluations for AI systems.
  • Build datasets and define metrics for quality evaluation.
  • Work with product teams to improve model performance.

Skills

Strong Python
Experience with ML systems
Understanding of LLMs or speech systems
Ability to design test cases
Comfortable working with engineers and product teams

Job description

This role sits at the centre of how we measure and improve AI systems in production.

You’ll define what good performance means across LLMs, ASR, TTS, and full speech-to-speech pipelines, and build the datasets, metrics, and evaluation systems that make AI quality measurable and comparable in the real world.

You’ll work closely with engineering and product teams to ensure model changes lead to real improvements in user experience, not just better offline benchmarks.

What you’ll do
  • Design and run evaluations across LLM, ASR, TTS, and speech-to-speech systems
  • Build real-world datasets and test cases from production behaviour and edge cases
  • Define metrics and scorecards for model and system quality
  • Benchmark internal models against external and frontier systems
  • Build Python tools to automate evaluation workflows
  • Create internal leaderboards, red-teaming setups, and regression tests
  • Work with engineers and product teams to diagnose system failures
  • Turn vague product goals into measurable evaluation frameworks
What this role is about
  • Defining and measuring AI quality in production systems
  • Turning real user behaviour into structured evaluation signals
  • Ensuring model changes improve real-world performance
  • Understanding why AI systems fail, not just whether they do
What good looks like
  • You can translate improved quality into measurable metrics
  • You think in terms of system impact (before vs after), not just accuracy
  • You’re comfortable working across code, data, and production systems
  • You care about real-world behaviour, not just benchmarks
Core skills
  • Strong Python (scripting, data analysis, tooling)
  • Experience with ML systems, evaluation, or experimentation
  • Understanding of LLMs or speech systems (ASR / TTS)
  • Ability to design test cases and structured datasets
  • Comfortable working with engineers and product teams
Nice to have
  • Experience with LLM evaluation or benchmarking
  • Exposure to speech or multimodal systems
  • Familiarity with production APIs or ML systems
  • Experience with automated testing or CI-style workflows
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

QA Engineer
QA Engineer

iFindTech Ltd • Greater London

On-site
GBP 60,000 - 90,000
AI Quality & Evaluation Engineer: Production Metrics
AI Quality & Evaluation Engineer: Production Metrics

ConnexAI • Manchester

On-site
GBP 50,000 - 70,000
AI QA Evaluation Engineer: Scale Metrics for ML
AI QA Evaluation Engineer: Scale Metrics for ML

Appnovation Technologies • Greater London

On-site
GBP 60,000 - 90,000
AI QA Engineer
AI QA Engineer

WeDoTech • Greater London

Hybrid
GBP 70,000 - 95,000
Employee shares/equity
Private healthcare
Life insurance
+2
Data Scientist
Data Scientist

ConnexAI • Manchester

On-site
GBP 70,000 - 110,000
AI Evaluation Engineer
AI Evaluation Engineer

Summer-Browning Associates • Greater London

Hybrid
GBP 74,000 - 129,000
AI Engineer
AI Engineer

Fruition Group • City Of London

On-site
GBP 120,000 - 180,000
Data Scientist – AI Evaluation & Benchmarking Manager
Data Scientist – AI Evaluation & Benchmarking Manager

Jobtailor • Manchester

On-site
GBP 60,000 - 90,000
AI QA Engineer
AI QA Engineer

WeDo Technology Solutions Limited • Greater London

Hybrid
GBP 70,000 - 95,000
10% bonus
Employee shares/equity
Private healthcare
+3
AI Research Scientist
AI Research Scientist

OxSci • Greater London

On-site
GBP 90,000 - 130,000
Founding equity
Direct line to founders
Open research questions
+1