ML Rigor Auditor for Experimental Tasks

Mercor

San Francisco (CA)

On-site

USD 130,000 - 190,000

Full time

11 days ago
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Mercor is seeking a reviewer to evaluate the quality, correctness, and methodological rigor of applied machine-learning tasks used to train and evaluate frontier AI models. You will assess experiment design, model-selection reasoning, and evaluation methodology, and deliver rubric-based written feedback.

The role emphasizes careful data-quality practices, critique of ML claims against evidence, and the ability to reproduce results across experiments.

Qualifications

  • 3+ years hands-on applied/experimental ML (design, selection, tuning, evaluation)
  • Strong grasp of data-quality rigor: leakage detection, metric gaming, and train/test/CV hygiene
  • Proficiency with standard ML frameworks (PyTorch, TensorFlow, scikit-learn, XGBoost)
  • Ability to critique ML claims against evidence and reproduce results

Responsibilities

  • Evaluate the quality, correctness, and methodological rigor of applied ML tasks used to train and evaluate models at the frontier AI lab.
  • Assess experiment design, model-selection reasoning, and evaluation methodology, and provide rubric-based written feedback.
  • Identify gaps, propose concrete improvements, and communicate findings clearly to technical and non-technical stakeholders

Job description

Mercor is seeking a reviewer to evaluate the quality, correctness, and methodological rigor of applied machine-learning tasks used to train and evaluate frontier AI models. You will assess experiment design, model-selection reasoning, and evaluation methodology, and deliver rubric-based written feedback.

The role emphasizes careful data-quality practices, critique of ML claims against evidence, and the ability to reproduce results across experiments.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Challenge Auditor: Rubric-Based Feedback
ML Challenge Auditor: Rubric-Based Feedback

Mercor • United States

Remote
USD 120,000 - 170,000
Applied ML Task Auditor & Rubric Evaluator
Applied ML Task Auditor & Rubric Evaluator

Obsidian • San Francisco (CA)

On-site
USD 120,000 - 210,000
Applied ML Task Auditor — Experiment Quality Lead
Applied ML Task Auditor — Experiment Quality Lead

Dorado • Northern (KY)

Hybrid
USD 90,000 - 130,000
ML Challenge Task Auditor Mercor · Remote — United States $70-90/hr →
ML Challenge Task Auditor Mercor · Remote — United States $70-90/hr →

Dorado • Northern (KY)

Hybrid
USD 90,000 - 130,000
ML Task Auditor - Applied ML
ML Task Auditor - Applied ML

Mercor • San Francisco (CA)

On-site
USD 130,000 - 190,000
ML Task Auditor - Applied ML
ML Task Auditor - Applied ML

Obsidian • San Francisco (CA)

On-site
USD 120,000 - 210,000
Applied ML Evaluation Auditor
Applied ML Evaluation Auditor

Obsidian • San Francisco (CA)

Remote
USD 90,000 - 130,000
Remote ML Task Auditor & AI Quality Engineer
Remote ML Task Auditor & AI Quality Engineer

Meridial • United States

Remote
USD 96,000 - 138,000
AI Evaluation Scientist Math PhD Frontier Model Benchmark
AI Evaluation Scientist Math PhD Frontier Model Benchmark

Obsidian • San Francisco (CA)

On-site
USD 96,000 - 179,000
AI Coding Trace Auditor - Rubric-Based Feedback
AI Coding Trace Auditor - Rubric-Based Feedback

Mercor • San Francisco (CA)

On-site
USD 120,000 - 180,000