Data Science Expert - AI Evaluation

Mercor

New York (NY)

On-site

USD 140,000 - 180,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Mercor seeks experienced data scientists to design task-specific grading criteria for real-world deliverables and to score both AI-generated and human work with detailed justifications. You will ensure reproducible, defensible scoring through evidence-based judgment and rapid iteration with senior reviewers.

Ideal candidates have 5+ years in industry data science, plus deep fluency in experiment design, A/B testing, and communicating findings to executives.

Qualifications

  • 5+ years of professional data science experience in industry.
  • Background in business operations, product, or growth data science at top-tier technology companies.
  • Deep fluency in experiment design and A/B testing, metric definition, SQL/Python analysis, and communicating findings to executive stakeholders.
  • Exceptionally strong written communication.
  • Detail-oriented, consistent, and comfortable with peer calibration.

Responsibilities

  • Design precise, task-specific grading criteria for real-world data science deliverables (analyses, models, dashboards, experiment readouts, and written recommendations).
  • Score AI-generated and human work samples against those criteria, with detailed written justifications for every score.
  • Apply consistent, evidence-based judgment so that scores are reproducible and defensible.
  • Incorporate structured feedback from senior reviewers and iterate quickly on your work.

Skills

Experiment design
SQL
Python
Executive communication
Written communication
Attention to detail

Job description

1. Role Overview

Mercor is partnering with a leading AI research organization to engage experienced data scientists for a project focused on evaluating how well AI systems perform real-world data science work. Rather than producing deliverables yourself, you will define what excellent work looks like: designing task-specific grading criteria and scoring completed work samples with rigorous, well-reasoned written justifications.

2. Key Responsibilities
  • Design precise, task-specific grading criteria for real-world data science deliverables (analyses, models, dashboards, experiment readouts, and written recommendations)

  • Score AI-generated and human work samples against those criteria, with detailed written justifications for every score

  • Apply consistent, evidence-based judgment so that scores are reproducible and defensible

  • Incorporate structured feedback from senior reviewers and iterate quickly on your work

3. Ideal Qualifications
  • 5+ years of professional data science experience in industry

  • Background in business operations, product, or growth data science at top-tier technology companies

  • Deep fluency in experiment design and A/B testing, metric definition, SQL/Python analysis, and communicating findings to executive stakeholders

  • Exceptionally strong written communication

  • Detail-oriented, consistent, and comfortable having your judgment reviewed and calibrated against peers

  • Prior experience with AI training, evaluation, or human-data projects is a strong plus

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Science Evaluation Architect
Data Science Evaluation Architect

Dorado • United States

Remote
USD 120,000 - 180,000
Data Science Expert
Data Science Expert

Dorado • United States

Remote
USD 120,000 - 180,000
Data Science Expert - AI Evaluation
Data Science Expert - AI Evaluation

Visa Hunt • Germany (OH)

On-site
USD 165,000 - 234,000
Data Scientist - AI Evaluation
Data Scientist - AI Evaluation

Mercor • United States

On-site
AI Evaluation Architect for Data Science
AI Evaluation Architect for Data Science

Mercor • New York (NY)

On-site
USD 140,000 - 180,000
Remote Data Science Expert — AI Evaluation & Growth Impact
Remote Data Science Expert — AI Evaluation & Growth Impact

Mercor • United States

On-site
Management Consultant - Strategy Expert
Management Consultant - Strategy Expert

Obsidian • San Francisco (CA)

On-site
USD 120,000 - 180,000
Marketing Expert - Evaluation Specialist
Marketing Expert - Evaluation Specialist

Mercor • New York (NY)

On-site
USD 85,000 - 120,000
Sales Engineering Expert - Evaluator
Sales Engineering Expert - Evaluator

Obsidian • San Francisco (CA)

On-site
USD 120,000 - 170,000
Sales Engineering Expert - Evaluator
Sales Engineering Expert - Evaluator

Mercor • San Francisco (CA)

On-site
USD 140,000 - 190,000