Data Science Expert - Evaluation Specialist

Mercor

Toronto

On-site

CAD 120,000 - 180,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Mercor, in partnership with a leading AI research organization, seeks a seasoned data scientist to define what excellent work looks like by designing task-specific grading criteria and scoring sample results with rigorous, written justifications.

You will evaluate AI-generated and human work, ensure scores are reproducible and defensible, and incorporate structured feedback from senior reviewers to iterate on standards for real-world data science deliverables.

Qualifications

  • 5+ years of professional data science experience in industry.
  • Background in business operations, product, or growth data science at top-tier technology companies.
  • Deep fluency in experiment design and A/B testing, metric definition, SQL/Python analysis, and communicating findings to executive stakeholders.
  • Exceptionally strong written communication.
  • Detail-oriented, consistent, and comfortable having your judgment reviewed and calibrated against peers.
  • Prior experience with AI training, evaluation, or human-data projects is a strong plus.

Responsibilities

  • Design precise, task-specific grading criteria for real-world data science deliverables (analyses, models, dashboards, experiment readouts, and written recommendations).
  • Score AI-generated and human work samples against those criteria, with detailed written justifications for every score.
  • Apply consistent, evidence-based judgment so that scores are reproducible and defensible.
  • Incorporate structured feedback from senior reviewers and iterate quickly on your work.

Skills

Experiment design
A/B testing
Executive communication
Written communication
Detail-oriented
Judgment calibration
AI evaluation

Tools

SQL
Python

Job description

1. Role Overview

Mercor is partnering with a leading AI research organization to engage experienced data scientists for a project focused on evaluating how well AI systems perform real-world data science work. Rather than producing deliverables yourself, you will define what excellent work looks like: designing task-specific grading criteria and scoring completed work samples with rigorous, well-reasoned written justifications.

2. Key Responsibilities
  • Design precise, task-specific grading criteria for real-world data science deliverables (analyses, models, dashboards, experiment readouts, and written recommendations)

  • Score AI-generated and human work samples against those criteria, with detailed written justifications for every score

  • Apply consistent, evidence-based judgment so that scores are reproducible and defensible

  • Incorporate structured feedback from senior reviewers and iterate quickly on your work

3. Ideal Qualifications
  • 5+ years of professional data science experience in industry

  • Background in business operations, product, or growth data science at top-tier technology companies

  • Deep fluency in experiment design and A/B testing, metric definition, SQL/Python analysis, and communicating findings to executive stakeholders

  • Exceptionally strong written communication

  • Detail-oriented, consistent, and comfortable having your judgment reviewed and calibrated against peers

  • Prior experience with AI training, evaluation, or human-data projects is a strong plus

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Science Expert - Evaluation Specialist
Data Science Expert - Evaluation Specialist

Obsidian • Toronto

Hybrid
CAD 90,000 - 140,000
Data Science Expert - AI Evaluation
Data Science Expert - AI Evaluation

Mercor • Toronto

On-site
CAD 120,000 - 170,000
Data Science Specialist - Fully Remote | Upto $170/hr
Data Science Specialist - Fully Remote | Upto $170/hr

Obsidian • Toronto

Remote
CAD 110,000 - 170,000
Data Science Quality Architect for AI Evaluation
Data Science Quality Architect for AI Evaluation

Obsidian • Toronto

Remote
CAD 110,000 - 170,000
AI Evaluation Scientist
AI Evaluation Scientist

Mercor • Toronto

On-site
CAD 120,000 - 180,000
UI/UX Design Expert - Evaluator
UI/UX Design Expert - Evaluator

Obsidian • Toronto

On-site
CAD 70,000 - 110,000
Design Evaluation Expert - Graphic
Design Evaluation Expert - Graphic

Obsidian • Toronto

On-site
CAD 80,000 - 100,000
UI/UX Design Expert - Evaluator
UI/UX Design Expert - Evaluator

Mercor • Toronto

On-site
CAD 90,000 - 130,000
AI Solutions Engineer - Pre-Sales Quality & Scoring
AI Solutions Engineer - Pre-Sales Quality & Scoring

Obsidian • Toronto

Remote
CAD 90,000 - 130,000
Mathematics PhD - AI Evaluation Expert
Mathematics PhD - AI Evaluation Expert

Mercor • Toronto

On-site
CAD 83,000 - 124,000