ML Challenge Auditor: Rubric-Based Feedback

Mercor

United States

Remote

USD 120,000 - 170,000

Full time

12 days ago
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Mercor seeks an ML evaluation specialist to assess applied/experimental rigor in frontier AI lab models. You will critique experiment design, model selection reasoning, and evaluation methodology, and provide rubric-based feedback to researchers.

This role emphasizes data-quality hygiene, reproducibility, and clear, evidence-based judgments. It is not about building or deploying LLMs or MLOps, but about improving scientific rigor through structured reviews and constructive guidance.

Qualifications

  • 3+ years hands-on applied/experimental ML (experiment design, model selection, hyperparameter tuning, evaluation methodology)
  • Strong grasp of data-quality rigor: leakage detection, metric gaming, and train/test/CV hygiene
  • Proficiency with standard ML frameworks (PyTorch, TensorFlow, scikit-learn, XGBoost)
  • Ability to critique ML claims against evidence and reproduce results

Responsibilities

  • Assess experiment design, model selection reasoning, and evaluation methodology.
  • Provide rubric-based written feedback to researchers.
  • Ensure documented, reproducible evaluation practices and reporting.

Skills

Applied ML
Experiment design
Model selection
Hyperparameter tuning
Evaluation methodology
Data-quality rigor
Leakage detection
Train/test/CV hygiene
PyTorch
TensorFlow
scikit-learn
XGBoost
Reproduce results

Tools

PyTorch
TensorFlow
scikit-learn
XGBoost

Job description

Mercor seeks an ML evaluation specialist to assess applied/experimental rigor in frontier AI lab models. You will critique experiment design, model selection reasoning, and evaluation methodology, and provide rubric-based feedback to researchers.

This role emphasizes data-quality hygiene, reproducibility, and clear, evidence-based judgments. It is not about building or deploying LLMs or MLOps, but about improving scientific rigor through structured reviews and constructive guidance.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Rigor Auditor for Experimental Tasks
ML Rigor Auditor for Experimental Tasks

Mercor • San Francisco (CA)

On-site
USD 130,000 - 190,000
Applied ML Task Auditor & Rubric Evaluator
Applied ML Task Auditor & Rubric Evaluator

Obsidian • San Francisco (CA)

On-site
USD 120,000 - 210,000
ML Challenge Task Auditor Mercor · Remote — United States $70-90/hr →
ML Challenge Task Auditor Mercor · Remote — United States $70-90/hr →

Dorado • Northern (KY)

Hybrid
USD 90,000 - 130,000
ML Task Auditor - Applied ML
ML Task Auditor - Applied ML

Mercor • San Francisco (CA)

On-site
USD 130,000 - 190,000
ML Task Auditor - Applied ML
ML Task Auditor - Applied ML

Obsidian • San Francisco (CA)

On-site
USD 120,000 - 210,000
Applied ML Task Auditor — Experiment Quality Lead
Applied ML Task Auditor — Experiment Quality Lead

Dorado • Northern (KY)

Hybrid
USD 90,000 - 130,000
Applied ML Evaluation Auditor
Applied ML Evaluation Auditor

Obsidian • San Francisco (CA)

Remote
USD 90,000 - 130,000
AI Coding Trace Auditor - Rubric-Based Feedback
AI Coding Trace Auditor - Rubric-Based Feedback

Mercor • San Francisco (CA)

On-site
USD 120,000 - 180,000
ML Evaluation Scientist - LLM Benchmarks
ML Evaluation Scientist - LLM Benchmarks

Scale AI • Seattle (WA)

On-site
USD 181,000 - 226,000
Health coverage
Retirement benefits
L&D stipend
+2
Research Scientist (Remote/US/LATAM)
Research Scientist (Remote/US/LATAM)

Anyone AI • United States

Remote
MXN 2,626,000 - 3,678,000