AI Evaluation Architect for Data Science Quality

Mercor

Greater London

On-site

GBP 90,000 - 130,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Mercor is partnering with a leading AI research organization to engage experienced data scientists for a project focused on evaluating how well AI systems perform real-world data science work. You will define what excellent work looks like: designing task-specific grading criteria and scoring completed work samples with rigorous, well-reasoned written justifications.

Key responsibilities include designing grading criteria, scoring both AI-generated and human samples, applying evidence-based

Qualifications

  • 5+ years of professional data science experience in industry.
  • Deep fluency in experiment design and A/B testing.
  • Strong ability to define metrics and communicate findings to executives.
  • Detailed written communication and clear justification of scores.

Responsibilities

  • Design precise, task-specific grading criteria for real-world data science deliverables.
  • Score AI-generated and human work samples with detailed written justifications.
  • Apply evidence-based judgment to ensure scores are reproducible and defensible.
  • Incorporate feedback from senior reviewers and iterate quickly.

Skills

Data science
Experiment design
A/B testing
SQL
Python
Executive communication

Job description

Mercor is partnering with a leading AI research organization to engage experienced data scientists for a project focused on evaluating how well AI systems perform real-world data science work. You will define what excellent work looks like: designing task-specific grading criteria and scoring completed work samples with rigorous, well-reasoned written justifications.

Key responsibilities include designing grading criteria, scoring both AI-generated and human samples, applying evidence-based

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Evaluation Data Scientist
AI Evaluation Data Scientist

Obsidian • Greater London

On-site
GBP 70,000 - 110,000
Data Science Evaluation Architect
Data Science Evaluation Architect

Obsidian • Greater London

Remote
GBP 80,000 - 120,000
Data Science Expert - AI Evaluation
Data Science Expert - AI Evaluation

Obsidian • Greater London

On-site
GBP 70,000 - 110,000
Data Science Expert - AI Evaluation
Data Science Expert - AI Evaluation

Mercor • Greater London

On-site
GBP 90,000 - 130,000
AI-Driven Sales Engineering Quality Evaluator
AI-Driven Sales Engineering Quality Evaluator

Mercor • Greater London

On-site
GBP 65,000 - 100,000
Data Science Evaluation Specialist
Data Science Evaluation Specialist

Mercor • Greater London

Remote
GBP 90,000 - 130,000
AI Banking Quality Evaluator
AI Banking Quality Evaluator

Mercor • Greater London

On-site
GBP 75,000 - 120,000
AI Technical Sales Engineer & Evaluation Specialist
AI Technical Sales Engineer & Evaluation Specialist

Mercor • Greater London

Remote
GBP 90,000 - 120,000
Brand Design Evaluator & Standards Architect
Brand Design Evaluator & Standards Architect

Mercor • Greater London

Remote
GBP 60,000 - 90,000
Accounting Quality Assessor for AI Evaluation
Accounting Quality Assessor for AI Evaluation

Obsidian • Greater London

On-site
GBP 70,000 - 90,000