Data Scientist, Agent

Lovable

Greater London

Hybrid

GBP 90,000 - 130,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Lovable is seeking a data scientist to own how we measure and improve our AI agent. You’ll define metrics for success, completion, and error rates, and build eval systems and experiments to assess agent changes.

You’ll turn agent traces and telemetry into concrete fixes, collaborating closely with the agent engineering team. You’ll also build automated evaluation tooling for continuous improvement of agent behavior, using strong SQL, Python, and statistics to drive robust results.

Qualifications

  • Must have strong SQL and Python, applied statistics, and experimentation.
  • Experience or strong interest in LLM evaluation and observability: building evals, scoring outputs, tracing agent behavior, and catching regressions.
  • Able to design A/B tests for agent changes where outcomes are noisy.
  • Build systems and agents that produce insights continuously, not one-off analyses.

Responsibilities

  • Define and own the metrics for agent quality: success, completion, error rates, and related behaviors.
  • Build eval systems and experiment framework to decide if an agent change ships (A/B rollout that catches regressions).
  • Turn agent traces and telemetry into concrete fixes with the agent engineering team.
  • Build tooling and agents that produce evaluations continuously as the agent evolves.
  • Set the bar for judging agent behavior where there is no exact answer key.

Skills

SQL
Python
Applied statistics
Experimentation
LLM evaluation & observability
A/B testing

Tools

Braintrust
OTEL tracing
BigQuery
PubSub
Hex
Lovable Apps

Job description

TL;DR — You own how we measure and improve Lovable's AI agent. You build the eval systems and experiments that tell us whether a change makes the agent better or worse, and you turn agent telemetry into the fixes that raise success rates and cut errors.

At Lovable, data scientists are not isolated model-builders; they sit close to the product, experimenting continuously with how intelligence changes user behavior and product dynamics.

Why Lovable?

Lovable is the software creation platform that gives people the power to act on the problems closest to them. For decades, turning an idea into software required so much capital, technical fluency, and time that many ideas never came to life. Lovable is the counterargument: a platform for all people with ideas, ambition, and problems worth solving. From solopreneurs to small business owners to teams at companies like Adidas and Zendesk, people have built over 60 million projects on Lovable since its launch in November 2024. And we’re just getting started.

We’re building a generational company from Stockholm, with growing teams in London, Boston, New York, and San Francisco. Our team is small, talent-dense, and moving quickly, with a culture rooted in extreme ownership, high velocity, and low-ego collaboration. We look for people who care deeply, ship fast, and are eager to make a dent in the world.

Lovable is one of TIME’s 100 Most Influential Companies and has been recognized on the Forbes AI 50 and CNBC Disruptor 50, reflecting our momentum as one of Europe’s fastest-growing AI companies and one of the most ambitious places to build in this next era of software.

What we're looking for
  • A data scientist who wants to make an AI agent measurably better, not just report on it. You own agent quality metrics and drive them up.

  • Experience or strong interest in LLM evaluation and observability: building evals, scoring outputs, tracing agent behavior, and catching regressions.

  • Strong SQL and Python, applied statistics, and experimentation. Comfortable designing A/B tests for agent changes where outcomes are noisy.

  • You build systems and agents that produce this insight continuously, rather than one-off analyses.

  • Instinct for what "good" looks like in agent behavior (success, error rates, task completion) and how to measure it when there is no clean answer key.

  • Entrepreneurial. Thrives in ambiguity, and works closely with the engineers building the agent.

What you'll do
  • Define and own the metrics for agent quality: success, completion, error rates, and the behaviors that drive them.

  • Build the eval systems and experiment framework that decide whether an agent change ships, like an A/B-tested rollout that catches a change increasing errors before it reaches everyone.

  • Turn agent traces and telemetry into concrete fixes, working directly with the agent engineering team.

  • Build the tooling and agents that produce these evaluations continuously as the agent evolves.

  • Set the bar for how we judge agent behavior where there is no answer key to check against.

Our tech stack

We're building with tools that both humans and AI love:

  • Languages: SQL and Python

  • LLM evaluation & observability: Braintrust, OTEL tracing, many LLM providers

  • Warehouse & events: BigQuery, PubSub

  • Analytics & product: Hex, Lovable Apps

  • Experimentation: A/B and growth testing

  • Cloud: GCP

And always on the lookout for what's next.

We treat all candidates equally.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Scientist
Data Scientist

Lovable • Greater London

Hybrid
GBP 70,000 - 110,000
Data Scientist, Product
Data Scientist, Product

Lovable • Greater London

On-site
GBP 85,000 - 110,000
Data Scientist, Growth
Data Scientist, Growth

Lovable • Greater London

Hybrid
GBP 80,000 - 120,000
Data Scientist, Growth
Data Scientist, Growth

AI Startups UK • Greater London

Hybrid
GBP 90,000 - 130,000
AI Research Engineer, Post-Training
AI Research Engineer, Post-Training

Lovable • Greater London

On-site
GBP 120,000 - 180,000
Data Scientist, Pricing
Data Scientist, Pricing

Lovable • Greater London

Hybrid
GBP 90,000 - 160,000
Data Scientist, Growth
Data Scientist, Growth

Antler • Greater London

On-site
GBP 70,000 - 110,000
Staff / Principal Software Engineer, Product
Staff / Principal Software Engineer, Product

Lovable • Greater London

On-site
GBP 150,000 - 190,000
Software Engineer, Growth
Software Engineer, Growth

Lovable • Greater London

Hybrid
GBP 90,000 - 150,000
AI Agent Quality Scientist
AI Agent Quality Scientist

Lovable • Greater London

Hybrid
GBP 90,000 - 130,000