Data Science Expert - AI Evaluation

Obsidian

Greater London

On-site

GBP 70,000 - 110,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Mercor is seeking experienced data scientists to define what excellent work looks like and to score AI-generated and human work samples accordingly. You will design task-specific grading criteria, provide rigorous written justifications for every score, and ensure our scoring remains reproducible and defensible.

Ideal candidates bring 5+ years in industry data science, strong experiment design and A/B testing skills, and the ability to clearly communicate findings to executive stakeholders.

Qualifications

  • 5+ years of professional data science experience in industry.
  • Background in business operations, product, or growth data science at top-tier technology companies.
  • Deep fluency in experiment design and A/B testing, metric definition, SQL/Python analysis, and communicating findings to executive stakeholders.
  • Exceptionally strong written communication.
  • Detail-oriented, consistent, and comfortable having your judgment reviewed and calibrated against peers.
  • Prior experience with AI training, evaluation, or human-data projects is a strong plus.

Responsibilities

  • Design precise, task-specific grading criteria for real-world data science deliverables (analyses, models, dashboards, experiment readouts, and written recommendations).
  • Score AI-generated and human work samples against those criteria, with detailed written justifications for every score.
  • Apply consistent, evidence-based judgment so that scores are reproducible and defensible.
  • Incorporate structured feedback from senior reviewers and iterate quickly on your work.

Skills

Data science experience
Experiment design & A/B testing
SQL/Python analysis
Executive communication
Written communication
Judgment calibration

Job description

1. Role Overview

Mercor is partnering with a leading AI research organization to engage experienced data scientists for a project focused on evaluating how well AI systems perform real-world data science work. Rather than producing deliverables yourself, you will define what excellent work looks like: designing task-specific grading criteria and scoring completed work samples with rigorous, well-reasoned written justifications.

2. Key Responsibilities
  • Design precise, task-specific grading criteria for real-world data science deliverables (analyses, models, dashboards, experiment readouts, and written recommendations)

  • Score AI-generated and human work samples against those criteria, with detailed written justifications for every score

  • Apply consistent, evidence-based judgment so that scores are reproducible and defensible

  • Incorporate structured feedback from senior reviewers and iterate quickly on your work

3. Ideal Qualifications
  • 5+ years of professional data science experience in industry

  • Background in business operations, product, or growth data science at top-tier technology companies

  • Deep fluency in experiment design and A/B testing, metric definition, SQL/Python analysis, and communicating findings to executive stakeholders

  • Exceptionally strong written communication

  • Detail-oriented, consistent, and comfortable having your judgment reviewed and calibrated against peers

  • Prior experience with AI training, evaluation, or human-data projects is a strong plus

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Science Expert - Evaluation Specialist
Data Science Expert - Evaluation Specialist

Mercor • Greater London

On-site
GBP 70,000 - 110,000
Data Science Expert - AI Evaluation
Data Science Expert - AI Evaluation

Mercor • Greater London

On-site
GBP 90,000 - 130,000
Data Science Evaluation Architect
Data Science Evaluation Architect

Obsidian • Greater London

Remote
GBP 80,000 - 120,000
AI Evaluation Data Scientist
AI Evaluation Data Scientist

Obsidian • Greater London

On-site
GBP 70,000 - 110,000
Data Science Evaluation Architect
Data Science Evaluation Architect

Mercor • Greater London

On-site
GBP 70,000 - 110,000
AI Evaluation Architect for Data Science Quality
AI Evaluation Architect for Data Science Quality

Mercor • Greater London

On-site
GBP 90,000 - 130,000
Data Science Evaluation Specialist
Data Science Evaluation Specialist

Mercor • Greater London

Remote
GBP 90,000 - 130,000
Sales Engineering Expert - Evaluator
Sales Engineering Expert - Evaluator

Mercor • Greater London

On-site
GBP 65,000 - 100,000
Investment Banking Expert - Evaluator
Investment Banking Expert - Evaluator

Obsidian • Greater London

On-site
GBP 60,000 - 90,000
Investment Banking Expert - Evaluator
Investment Banking Expert - Evaluator

Mercor • Greater London

On-site
GBP 75,000 - 120,000