Data Science Expert - AI Evaluation

Mercor

Toronto

On-site

CAD 120,000 - 170,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Mercor, partnering with a leading AI research organization, seeks experienced data scientists to help define what excellent work looks like by designing task-specific grading criteria and scoring completed samples with detailed written justifications.

The role emphasizes rigorous evaluation, reproducible scoring, and collaboration with senior reviewers to calibrate judgments across AI training, evaluation, and human-data projects.

Qualifications

  • 5+ years of professional data science experience in industry.
  • Experience in business operations, product, or growth data science at top-tier technology companies.
  • Deep fluency in experiment design, A/B testing, metric definition, SQL/Python analysis, and presenting findings to executives.
  • Exceptionally strong written communication.
  • Detail-oriented with the ability to be calibrated against peers.
  • Prior experience with AI training, evaluation, or human-data projects is a strong plus.

Responsibilities

  • Design precise, task-specific grading criteria for real-world data science deliverables (analyses, models, dashboards, experiment readouts, and written recommendations).
  • Score AI-generated and human work samples against those criteria, with detailed written justifications for every score.
  • Apply consistent, evidence-based judgment so that scores are reproducible and defensible.
  • Incorporate structured feedback from senior reviewers and iterate quickly on your work.

Skills

Data science experience
Experiment design
Written communication
Executive stakeholder communication
Attention to detail
AI training/evaluation experience

Tools

SQL
Python

Job description

1. Role Overview

Mercor is partnering with a leading AI research organization to engage experienced data scientists for a project focused on evaluating how well AI systems perform real-world data science work. Rather than producing deliverables yourself, you will define what excellent work looks like: designing task-specific grading criteria and scoring completed work samples with rigorous, well-reasoned written justifications.

2. Key Responsibilities
  • Design precise, task-specific grading criteria for real-world data science deliverables (analyses, models, dashboards, experiment readouts, and written recommendations)
  • Score AI-generated and human work samples against those criteria, with detailed written justifications for every score
  • Apply consistent, evidence-based judgment so that scores are reproducible and defensible
  • Incorporate structured feedback from senior reviewers and iterate quickly on your work
3. Ideal Qualifications
  • 5+ years of professional data science experience in industry
  • Background in business operations, product, or growth data science at top-tier technology companies
  • Deep fluency in experiment design and A/B testing, metric definition, SQL/Python analysis, and communicating findings to executive stakeholders
  • Exceptionally strong written communication
  • Detail-oriented, consistent, and comfortable having your judgment reviewed and calibrated against peers
  • Prior experience with AI training, evaluation, or human-data projects is a strong plus
4. Application Process
  • Qualified applicants may be asked to complete a brief technical assessment or submit additional information
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Science Quality Architect for AI Evaluation
Data Science Quality Architect for AI Evaluation

Obsidian • Toronto

Remote
CAD 110,000 - 170,000
Data Science Specialist - Fully Remote | Upto $170/hr
Data Science Specialist - Fully Remote | Upto $170/hr

Obsidian • Toronto

Remote
CAD 110,000 - 170,000
Sales Engineering Expert - Evaluator
Sales Engineering Expert - Evaluator

Mercor • Toronto

On-site
CAD 80,000 - 110,000
UI/UX Design Expert - Evaluator
UI/UX Design Expert - Evaluator

Obsidian • Toronto

On-site
CAD 70,000 - 110,000
Design Evaluation Expert - Graphic
Design Evaluation Expert - Graphic

Obsidian • Toronto

On-site
CAD 80,000 - 100,000
Data Engineer - AI Model Evaluation
Data Engineer - AI Model Evaluation

Obsidian • Toronto

On-site
CAD 634,000 - 903,000
UI/UX Design Expert - Evaluator
UI/UX Design Expert - Evaluator

Mercor • Toronto

On-site
CAD 90,000 - 130,000
Technical Sales Consultant - Fully Remote | Upto $150/hr
Technical Sales Consultant - Fully Remote | Upto $150/hr

Obsidian • Toronto

Remote
CAD 90,000 - 130,000
Accounting Expert - CPA Preferred
Accounting Expert - CPA Preferred

Obsidian • Toronto

On-site
CAD 80,000 - 110,000
Financial Analyst - Fully Remote | Upto $160/hr
Financial Analyst - Fully Remote | Upto $160/hr

Obsidian • Toronto

Remote
CAD 90,000 - 130,000