Data Science Quality Architect for AI Evaluation

Obsidian

Toronto

Remote

CAD 110,000 - 170,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Mercor is partnering with a leading AI research organization to engage experienced data scientists for a project focused on evaluating how well AI systems perform real-world data science work. You will define what excellent work looks like: designing task-specific grading criteria and scoring completed work samples with rigorous, well-reasoned written justifications.

This role emphasizes independent judgment, reproducibility, and clear communication to executive stakeholders, with emphasis on

Qualifications

  • 5+ years of professional data science experience in industry.
  • Background in business operations, product, or growth data science at top-tier technology companies
  • Deep fluency in experiment design and A/B testing, metric definition, SQL/Python analysis, and communicating findings to executive stakeholders
  • Exceptionally strong written communication
  • Detail-oriented, consistent, and comfortable having your judgment reviewed and calibrated against peers
  • Prior experience with AI training, evaluation, or human-data projects is a strong plus

Responsibilities

  • Design precise, task-specific grading criteria for real-world data science deliverables (analyses, models, dashboards, experiment readouts, and written recommendations)
  • Score AI-generated and human work samples against those criteria, with detailed written justifications for every score
  • Apply consistent, evidence-based judgment so that scores are reproducible and defensible
  • Incorporate structured feedback from senior reviewers and iterate quickly on your work

Skills

Data science experience
Experiment design
A/B testing
SQL
Python
Executive communication
Written communication
Attention to detail
AI evaluation

Job description

Mercor is partnering with a leading AI research organization to engage experienced data scientists for a project focused on evaluating how well AI systems perform real-world data science work. You will define what excellent work looks like: designing task-specific grading criteria and scoring completed work samples with rigorous, well-reasoned written justifications.

This role emphasizes independent judgment, reproducibility, and clear communication to executive stakeholders, with emphasis on

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Evaluation Scientist
AI Evaluation Scientist

Mercor • Toronto

On-site
CAD 120,000 - 180,000
Data Science Expert - Evaluation Specialist
Data Science Expert - Evaluation Specialist

Mercor • Toronto

On-site
CAD 120,000 - 180,000
Data Science Expert - Evaluation Specialist
Data Science Expert - Evaluation Specialist

Obsidian • Toronto

On-site
CAD 90,000 - 140,000
AI Solutions Engineer - Pre-Sales Quality & Scoring
AI Solutions Engineer - Pre-Sales Quality & Scoring

Obsidian • Toronto

Remote
CAD 90,000 - 130,000
Data Science Specialist - Fully Remote | Upto $170/hr
Data Science Specialist - Fully Remote | Upto $170/hr

Obsidian • Toronto

Remote
CAD 110,000 - 170,000
Remote CRM Data Quality Evaluator for AI Outputs
Remote CRM Data Quality Evaluator for AI Outputs

Mercor • Toronto

Remote
CAD 34,000 - 55,000
Remote Data Analysis Expert — AI Output Evaluator
Remote Data Analysis Expert — AI Output Evaluator

Mercor • Toronto

On-site
CAD 34,000 - 55,000
Remote AI QA & Competitive Intelligence Analyst
Remote AI QA & Competitive Intelligence Analyst

Mercor • Toronto

Remote
CAD 41,000 - 62,000
Sales Engineering Expert - Evaluator
Sales Engineering Expert - Evaluator

Mercor • Toronto

On-site
CAD 80,000 - 110,000
AI Benchmark Architect for Scientific Computing
AI Benchmark Architect for Scientific Computing

Mercor • Toronto

Remote
CAD 83,000 - 165,000