Data Science Evaluation Architect

Obsidian

Greater London

Remote

GBP 80,000 - 120,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Mercor, partnering with a leading AI research organization, seeks experienced data scientists to design task-specific grading criteria and to evaluate AI-generated and human work samples. You will define what excellent work looks like and provide rigorous, written justifications for every score.

Ideal candidates have 5+ years in industry data science, deep expertise in experiment design and A/B testing, and strong SQL/Python analysis, with exceptional written communication to convey findings to

Qualifications

  • 5+ years of professional data science experience in industry.
  • Background in business operations, product, or growth data science at top-tier technology companies.
  • Deep fluency in experiment design and A/B testing, metric definition, SQL/Python analysis, and communicating findings to executive stakeholders.
  • Exceptionally strong written communication.
  • Detail-oriented, consistent, and comfortable having your judgment reviewed and calibrated against peers.
  • Prior experience with AI training, evaluation, or human-data projects is a strong plus.

Responsibilities

  • Design precise, task-specific grading criteria for real-world data science deliverables (analyses, models, dashboards, experiment readouts, and written recommendations).
  • Score AI-generated and human work samples against those criteria, with detailed written justifications for every score.
  • Apply consistent, evidence-based judgment so that scores are reproducible and defensible.
  • Incorporate structured feedback from senior reviewers and iterate quickly on your work.

Skills

5+ years data science
Business operations / product data
Experiment design & A/B testing
SQL/Python analysis
Executive stakeholder communication
Strong written communication
Attention to detail
AI training/evaluation experience

Job description

Mercor, partnering with a leading AI research organization, seeks experienced data scientists to design task-specific grading criteria and to evaluate AI-generated and human work samples. You will define what excellent work looks like and provide rigorous, written justifications for every score.

Ideal candidates have 5+ years in industry data science, deep expertise in experiment design and A/B testing, and strong SQL/Python analysis, with exceptional written communication to convey findings to

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Evaluation Data Scientist
AI Evaluation Data Scientist

Obsidian • Greater London

On-site
GBP 70,000 - 110,000
Data Science Evaluation Specialist
Data Science Evaluation Specialist

Mercor • Greater London

Remote
GBP 90,000 - 130,000
Data Science Expert - AI Evaluation
Data Science Expert - AI Evaluation

Obsidian • Greater London

On-site
GBP 70,000 - 110,000
AI Evaluation Architect for Data Science Quality
AI Evaluation Architect for Data Science Quality

Mercor • Greater London

On-site
GBP 90,000 - 130,000
Data Science Expert - AI Evaluation
Data Science Expert - AI Evaluation

Mercor • Greater London

On-site
GBP 90,000 - 130,000
AI Technical Sales Engineer & Evaluation Specialist
AI Technical Sales Engineer & Evaluation Specialist

Mercor • Greater London

Remote
GBP 90,000 - 120,000
Brand Design Evaluation Specialist
Brand Design Evaluation Specialist

Mercor • Greater London

On-site
GBP 45,000 - 75,000
AI-Driven Sales Engineering Quality Evaluator
AI-Driven Sales Engineering Quality Evaluator

Mercor • Greater London

On-site
GBP 65,000 - 100,000
Brand Design Evaluator & Standards Architect
Brand Design Evaluator & Standards Architect

Mercor • Greater London

Remote
GBP 60,000 - 90,000
Sales Engineering Expert - Evaluator
Sales Engineering Expert - Evaluator

Mercor • Greater London

On-site
GBP 65,000 - 100,000