Remote AI Evaluation & Benchmark Engineer

Intergral GmbH

United Kingdom

Remote

GBP 60,000 - 75,000

Full time

8 days ago
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Intergral UK is seeking an AI Evaluation & Benchmarking Engineer to design and implement systems that measure the quality of OpsPilot. You will lead automated evaluations, create synthetic workloads, and benchmark models, prompts, tools, and workflows against repeatable baselines.

You will work with cross-functional teams to extend evaluation across customer journeys, APIs, and backend services, using OpenTelemetry to ensure representative benchmarks and correlate results with metrics, logs, and

Qualifications

  • Practical experience working with LLMs, AI agents or AI evaluation.
  • Strong software engineering skills, particularly Python or a similar language.
  • Experience building automated evaluation, benchmarking, testing or experimentation infrastructure.
  • Experience creating synthetic workloads, datasets or evaluation scenarios.
  • An understanding of non-deterministic evaluation, including repeated measurement, variance and distinguishing meaningful changes from noise.
  • The ability to turn complex system behaviour into measurable criteria.
  • Comfort working across APIs, distributed systems and multiple layers of a software product.

Responsibilities

  • Evaluate the agent and the product it runs on.
  • Across the wider product, extend evaluation across important customer journeys, APIs, backend services and UI.
  • Turn failures and real-world problems into new evaluation scenarios.
  • Identify recurring failure patterns and capability gaps.
  • Test potential improvements against baselines, holdouts and unseen scenarios.
  • Detect regressions, benchmark overfitting and improvements that don\'t generalise.
  • Work directly with AI, software, platform and SRE engineers to investigate findings and improve the product.

Skills

LLM experience
Python
AI evaluation
Automation/testing infrastructure
Synthetic workloads/datasets
Non-deterministic evaluation
APIs and distributed systems

Tools

OpenTelemetry

Job description

Intergral UK is seeking an AI Evaluation & Benchmarking Engineer to design and implement systems that measure the quality of OpsPilot. You will lead automated evaluations, create synthetic workloads, and benchmark models, prompts, tools, and workflows against repeatable baselines.

You will work with cross-functional teams to extend evaluation across customer journeys, APIs, and backend services, using OpenTelemetry to ensure representative benchmarks and correlate results with metrics, logs, and

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Benchmarking Engineer for Agentic AI Systems (London)
Benchmarking Engineer for Agentic AI Systems (London)

Callosum • Greater London

On-site
GBP 90,000 - 140,000
Equity & Ownership
Private healthcare
Visa sponsorship and relocation
+1
Agentic Evaluation Engineer — Benchmarks & Systems
Agentic Evaluation Engineer — Benchmarks & Systems

AI Startups UK • Greater London

Hybrid
GBP 90,000 - 130,000
Competitive salary
Equity ownership
Private healthcare
+1
Senior ML QA Engineer: Benchmark & Validate AI Systems
Senior ML QA Engineer: Benchmark & Validate AI Systems

EngineersOfAI • Bristol

On-site
GBP 60,000 - 80,000
Unlimited annual leave
Up to 5% matched pension
Health cash plan
+2
Remote AI Operations Researcher & Evaluation Designer
Remote AI Operations Researcher & Evaluation Designer

United States Digital Space LLC • Greater London

Hybrid
GBP 408,000 - 1,019,000
Senior AI Compute Systems Engineer – Benchmarking & Metrics
Senior AI Compute Systems Engineer – Benchmarking & Metrics

United States Digital Space LLC • West of England, Greater London

On-site
GBP 70,000 - 110,000
Flexible working
Generous leave
Pension matching
+2
Staff Research Engineer: Agentic Evaluation & Benchmarks
Staff Research Engineer: Agentic Evaluation & Benchmarks

Callosum Technologies Ltd. • Greater London

Hybrid
GBP 90,000 - 140,000
Competitive salary
Equity & Ownership
Private healthcare
+2
AI Evaluation Engineer
AI Evaluation Engineer

SmartSourcing Ltd • Greater London

Hybrid
GBP 118,000 - 159,000
AI Evaluation & Assurance Engineer — Production AI Quality
AI Evaluation & Assurance Engineer — Production AI Quality

SmartSourcing Ltd • Greater London

Hybrid
GBP 129,000 - 166,000
Senior AI Systems Engineer - Scale & Benchmarking
Senior AI Systems Engineer - Scale & Benchmarking

Graphcore • Greater London

On-site
GBP 90,000 - 130,000
Flexible working
Generous leave
Pension
+4
Senior AI Evaluation & Alignment Engineer (Remote)
Senior AI Evaluation & Alignment Engineer (Remote)

Prolific • United Kingdom

Remote
GBP 65,000 - 80,000
Competitive salary
Remote working opportunities
Mission-driven culture