ML Task Auditor - Applied ML

Mercor

San Francisco (CA)

On-site

USD 130,000 - 190,000

Full time

11 days ago
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Mercor is seeking a reviewer to evaluate the quality, correctness, and methodological rigor of applied machine-learning tasks used to train and evaluate frontier AI models. You will assess experiment design, model-selection reasoning, and evaluation methodology, and deliver rubric-based written feedback.

The role emphasizes careful data-quality practices, critique of ML claims against evidence, and the ability to reproduce results across experiments.

Qualifications

  • 3+ years hands-on applied/experimental ML (design, selection, tuning, evaluation)
  • Strong grasp of data-quality rigor: leakage detection, metric gaming, and train/test/CV hygiene
  • Proficiency with standard ML frameworks (PyTorch, TensorFlow, scikit-learn, XGBoost)
  • Ability to critique ML claims against evidence and reproduce results

Responsibilities

  • Evaluate the quality, correctness, and methodological rigor of applied ML tasks used to train and evaluate models at the frontier AI lab.
  • Assess experiment design, model-selection reasoning, and evaluation methodology, and provide rubric-based written feedback.
  • Identify gaps, propose concrete improvements, and communicate findings clearly to technical and non-technical stakeholders

Job description

Evaluate the quality, correctness, and methodological rigor of applied machine-learning tasks used to train and evaluate a frontier AI lab's models. You'll assess experiment design, model-selection reasoning, and evaluation methodology — and provide clear, rubric-based written feedback.

Basic Qualifications
  • 3+ years hands‑on applied/experimental ML (experiment design, model selection, hyperparameter tuning, evaluation methodology)
  • Strong grasp of data-quality rigor: leakage detection, metric gaming, and train/test/CV hygiene
  • Proficiency with standard ML frameworks (PyTorch, TensorFlow, scikit-learn, XGBoost)
  • Ability to critique ML claims against evidence and reproduce results
Preferred Qualifications
  • Competition / benchmark experience (e.g., Kaggle)
  • Graduate research or publication record in applied ML
  • Prior task-grading or peer-review experience

Note: this role evaluates applied/experimental ML rigor — it is not an LLM-application-building or MLOps role.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Task Auditor - Applied ML
ML Task Auditor - Applied ML

Obsidian • San Francisco (CA)

On-site
USD 120,000 - 210,000
ML Challenge Task Auditor Mercor · Remote — United States $70-90/hr →
ML Challenge Task Auditor Mercor · Remote — United States $70-90/hr →

Dorado • Northern (KY)

Hybrid
USD 90,000 - 130,000
Applied ML Task Auditor — Experiment Quality Lead
Applied ML Task Auditor — Experiment Quality Lead

Dorado • Northern (KY)

Hybrid
USD 90,000 - 130,000
Applied ML Evaluation Auditor
Applied ML Evaluation Auditor

Obsidian • San Francisco (CA)

Remote
USD 90,000 - 130,000
ML Challenge Auditor: Rubric-Based Feedback
ML Challenge Auditor: Rubric-Based Feedback

Mercor • United States

Remote
USD 120,000 - 170,000
Applied ML Task Auditor & Rubric Evaluator
Applied ML Task Auditor & Rubric Evaluator

Obsidian • San Francisco (CA)

On-site
USD 120,000 - 210,000
Applied ML Evaluation Specialist (Remote Contract)
Applied ML Evaluation Specialist (Remote Contract)

OpenTrain AI • Northern (KY)

Hybrid
USD 96,000 - 124,000
ML Rigor Auditor for Experimental Tasks
ML Rigor Auditor for Experimental Tasks

Mercor • San Francisco (CA)

On-site
USD 130,000 - 190,000
ML Research Engineer - PhD - AI Trainer
ML Research Engineer - PhD - AI Trainer

Obsidian • Seattle (WA)

Hybrid
USD 100,000 - 150,000
Evaluation & Insights Machine Learning Engineer
Evaluation & Insights Machine Learning Engineer

Apple Inc. • Cupertino (CA)

On-site
USD 150,000 - 230,000