AI Evaluation Engineer

SmartSourcing Ltd

Greater London

Hybrid

GBP 118,000 - 159,000

Full time

9 days ago
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

SmartSourcing Ltd is seeking a hands-on AI evaluation engineer for a 4-month contract in London, Bristol, or Manchester with two days per week onsite. You will design and deploy AI evaluation tooling for AI/ML and LLM-based products, build evaluation harnesses in Python, and collaborate with technical and non-technical stakeholders to ensure quality, safety, reliability and governance.

Active SC clearance is preferred or eligible; suitable candidates will have deep expertise in AI evaluation

Qualifications

  • Strong AI/ML evaluation expertise with hands-on testing and quality assurance experience.
  • Deep expertise in evaluating AI/ML/LLM-based systems in production.
  • Proven leadership across teams and organisational boundaries.
  • Experience designing AI evaluation frameworks and tools (Python).
  • Active SC clearance preferred / eligible and willing to undergo SC Clearance.

Responsibilities

  • Design, develop and deploy AI evaluation and risk management tooling.
  • Build and maintain evaluation frameworks for AI, ML and LLM-based products.
  • Create and develop code solutions, harnesses and testing capabilities from the ground up using Python.
  • Define and execute evaluation strategies across model, agentic and application layers.
  • Assess AI system quality, performance, reliability, safety and risk.
  • Work closely with technical and non-technical stakeholders across multiple departments and organisations.
  • Provide technical leadership and guidance on AI evaluation best practice.

Skills

AI evaluation
Python development
Leadership
Stakeholder engagement
Quality assurance
ML/LLM evaluation
Production systems

Tools

RAGAS
DeepEval (DP-Eval)

Job description

4 Months 750 per day (Inside IR35)

Location: London, Bristol or Manchester (2 days per week onsite)

Security Clearance: Active SC Clearance preferred / eligible and willing to undergo SC Clearance.

Alternative backgrounds may include SDET, AI Evaluation Engineer, AI Assurance Engineer, Prompt Engineer, Harness Engineer or AI Quality Engineer.

This is a hands‑on engineering role for someone with deep expertise in AI/ML evaluation, testing and quality assurance, rather than a cloud, infrastructure or platform engineer. You will play a key role in establishing robust evaluation frameworks for production‑grade AI systems, helping assure the quality, safety and effectiveness of data‑driven and AI‑enabled products.

Responsibilities
  • Design, develop and deploy AI evaluation and risk management tooling.
  • Build and maintain evaluation frameworks for AI, ML and LLM‑based products.
  • Create and develop code solutions, harnesses and testing capabilities from the ground up using Python.
  • Define and execute evaluation strategies across model, agentic and application layers.
  • Assess AI system quality, performance, reliability, safety and risk.
  • Work closely with technical and non‑technical stakeholders across multiple departments and organisations.
  • Provide technical leadership and guidance on AI evaluation best practice.
Skills Experience
  • Strong commercial experience working with AI, Machine Learning and LLM‑based systems in production environments.
  • Deep understanding of AI evaluation methodologies, frameworks and quality engineering principles.
  • Hands‑on Python development experience with the ability to create and maintain your own codebase.
  • Experience evaluating AI solutions across:
  • Model layer
  • Agentic layer
  • Application layer
  • Knowledge of AI evaluation tooling and frameworks such as:
  • RAGAS
  • DeepEval (DP-Eval)
  • Other AI evaluation and testing frameworks
  • Experience developing automated testing, evaluation harnesses and assurance processes for AI systems.
  • Proven leadership experience, including working across teams, departments and organisational boundaries.
  • Strong stakeholder engagement and communication skills.
Desirable Experience
  • Public sector or government programme experience.
  • Understanding of responsible AI and AI assurance frameworks.
Please Note

This opportunity is specifically seeking candidates with strong AI evaluation and engineering expertise. Experience primarily focused on cloud, infrastructure, platform or DevOps engineering without significant AI evaluation experience are unlikely to be suitable.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Evaluation Engineer — Quality & Risk for AI Systems
AI Evaluation Engineer — Quality & Risk for AI Systems

SmartSourcing Ltd • Greater London

Hybrid
GBP 118,000 - 159,000
AI Evaluation & Assurance Engineer — Production AI Quality
AI Evaluation & Assurance Engineer — Production AI Quality

SmartSourcing Ltd • Greater London

Hybrid
GBP 129,000 - 166,000
AI Quality Engineer
AI Quality Engineer

SoCode Recruitment • Greater London

Hybrid
GBP 70,000 - 90,000
Employee shares
Private health insurance
Unlimited holidays
+1
AI Engineer
AI Engineer

McGregor Boyall • Manchester

On-site
GBP 128,000 - 168,000
AI Engineer
AI Engineer

Datatech Analytics • Greater London

Hybrid
GBP 70,000 - 90,000
AI Implementation Engineer ( Data Science / ML Background)
AI Implementation Engineer ( Data Science / ML Background)

Adria Solutions Ltd. • Manchester

Hybrid
GBP 60,000 - 85,000
Quarterly bonus
Hybrid working
Career progression
AI Engineer
AI Engineer

Peaple Talent • United Kingdom

Remote
AI Engineer
AI Engineer

Xcede • Greater London

Hybrid
GBP 90,000 - 120,000
Lead AI Solutions Architect
Lead AI Solutions Architect

Experis UK • Greater London

Hybrid
GBP 83,000 - 138,000
Generative AI Engineer
Generative AI Engineer

Lorien • Greater London

Hybrid
GBP 65,000 - 110,000