AI Agent Quality Scientist

Lovable

Greater London

Hybrid

GBP 90,000 - 130,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Lovable is seeking a data scientist to own how we measure and improve our AI agent. You’ll define metrics for success, completion, and error rates, and build eval systems and experiments to assess agent changes.

You’ll turn agent traces and telemetry into concrete fixes, collaborating closely with the agent engineering team. You’ll also build automated evaluation tooling for continuous improvement of agent behavior, using strong SQL, Python, and statistics to drive robust results.

Qualifications

  • Must have strong SQL and Python, applied statistics, and experimentation.
  • Experience or strong interest in LLM evaluation and observability: building evals, scoring outputs, tracing agent behavior, and catching regressions.
  • Able to design A/B tests for agent changes where outcomes are noisy.
  • Build systems and agents that produce insights continuously, not one-off analyses.

Responsibilities

  • Define and own the metrics for agent quality: success, completion, error rates, and related behaviors.
  • Build eval systems and experiment framework to decide if an agent change ships (A/B rollout that catches regressions).
  • Turn agent traces and telemetry into concrete fixes with the agent engineering team.
  • Build tooling and agents that produce evaluations continuously as the agent evolves.
  • Set the bar for judging agent behavior where there is no exact answer key.

Skills

SQL
Python
Applied statistics
Experimentation
LLM evaluation & observability
A/B testing

Tools

Braintrust
OTEL tracing
BigQuery
PubSub
Hex
Lovable Apps

Job description

Lovable is seeking a data scientist to own how we measure and improve our AI agent. You’ll define metrics for success, completion, and error rates, and build eval systems and experiments to assess agent changes.

You’ll turn agent traces and telemetry into concrete fixes, collaborating closely with the agent engineering team. You’ll also build automated evaluation tooling for continuous improvement of agent behavior, using strong SQL, Python, and statistics to drive robust results.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Scientist, Agent
Data Scientist, Agent

Lovable • Greater London

Hybrid
GBP 90,000 - 130,000
Staff AI Agent Engineer - Production-Grade General Agents
Staff AI Agent Engineer - Production-Grade General Agents

Scale AI • York and North Yorkshire

On-site
GBP 95,000 - 140,000
Health coverage
Learning & development stipend
Employee community events
+1
QA Engineer
QA Engineer

iFindTech Ltd • Greater London

On-site
GBP 60,000 - 90,000
AI Agent Architect & Orchestration Lead
AI Agent Architect & Orchestration Lead

adm Indicia • Greater London

On-site
GBP 70,000 - 100,000
AI Agent Architect for Financial Services
AI Agent Architect for Financial Services

Data-Intellect • Belfast City District

Hybrid
GBP 90,000 - 150,000
Healthcare cover
Referral scheme
Regular social events
+2
AI QA Evaluation Engineer: Scale Metrics for ML
AI QA Evaluation Engineer: Scale Metrics for ML

Appnovation Technologies • Greater London

On-site
GBP 60,000 - 90,000
Senior Product Engineer - AI Agent Evaluation & Infra
Senior Product Engineer - AI Agent Evaluation & Infra

Prolific • Greater London

On-site
GBP 85,000 - 120,000
Remote working
Benefits
AI Builder & Agent Engineer
AI Builder & Agent Engineer

Appnovation Technologies • Greater London

On-site
GBP 90,000 - 130,000
Agent Product Manager
Agent Product Manager

Eigent AI • Greater London

On-site
GBP 70,000 - 110,000
Production AI Agent Engineer (Reliability & Adoption)
Production AI Agent Engineer (Reliability & Adoption)

Mercor • Greater London

Remote
GBP 65,000 - 95,000