Machine Learning Evaluator - AI Trainer

Mercor

Philadelphia (Philadelphia County)

On-site

USD 120,000 - 180,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Mercor is seeking a researcher to evaluate the quality, correctness, and methodological rigor of applied machine-learning tasks used to train and evaluate frontier AI models. The role involves assessing experiment design, model-selection reasoning, and evaluation methodology and providing rubric-based written feedback.

The candidate should demonstrate 3+ years of hands-on applied/experimental ML work, strong data-quality rigor, and proficiency with PyTorch, TensorFlow, scikit-learn, and XGBoost,

Qualifications

  • 3+ years hands-on applied/experimental ML (design, selection, tuning, evaluation)
  • Strong data-quality rigor: leakage detection, metric gaming, train/test/CV hygiene
  • Proficient with standard ML frameworks (PyTorch, TensorFlow, scikit-learn, XGBoost)
  • Ability to critique ML claims against evidence and reproduce results

Responsibilities

  • Evaluate the quality, correctness, and methodological rigor of applied ML tasks used to train and evaluate models
  • Assess experiment design, model-selection reasoning, and evaluation methodology
  • Provide clear, rubric-based written feedback

Skills

Applied ML experience
Experiment design
Model selection
Hyperparameter tuning
Evaluation methodology
Data-quality rigor
Framework proficiency

Tools

PyTorch
TensorFlow
scikit-learn
XGBoost

Job description

Evaluate the quality, correctness, and methodological rigor of applied machine-learning tasks used to train and evaluate a frontier AI lab's models. You'll assess experiment design, model-selection reasoning, and evaluation methodology — and provide clear, rubric-based written feedback.

Basic Qualifications
  • 3+ years hands-on applied/experimental ML (experiment design, model selection, hyperparameter tuning, evaluation methodology)
  • Strong grasp of data-quality rigor: leakage detection, metric gaming, and train/test/CV hygiene
  • Proficiency with standard ML frameworks (PyTorch, TensorFlow, scikit-learn, XGBoost)
  • Ability to critique ML claims against evidence and reproduce results
Preferred Qualifications
  • Competition / benchmark experience (e.g., Kaggle)
  • Graduate research or publication record in applied ML
  • Prior task-grading or peer-review experience

Note: this role evaluates applied/experimental ML rigor — it is not an LLM-application-building or MLOps role.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

ML Challenge Task Auditor
ML Challenge Task Auditor

DigiNo • Northern (KY)

Hybrid
USD 100,000 - 180,000
ML Task Auditor - Applied ML
ML Task Auditor - Applied ML

Obsidian • San Francisco (CA)

On-site
USD 120,000 - 210,000
ML Task Auditor - Applied ML
ML Task Auditor - Applied ML

Mercor • San Francisco (CA)

On-site
USD 130,000 - 190,000
Applied ML Evaluator & AI Rigor Reviewer
Applied ML Evaluator & AI Rigor Reviewer

Obsidian • Philadelphia

On-site
USD 120,000 - 170,000
AI Evaluation Specialist: ML Task Auditor
AI Evaluation Specialist: ML Task Auditor

HumanitApp • Northern (KY)

Hybrid
USD 96,000 - 124,000
Applied ML Evaluation Auditor
Applied ML Evaluation Auditor

Obsidian • San Francisco (CA)

Remote
USD 90,000 - 130,000
ML Challenge Task Auditor
ML Challenge Task Auditor

HumanitApp • Northern (KY)

Hybrid
USD 96,000 - 124,000
Applied ML Evaluator & Methodology Reviewer
Applied ML Evaluator & Methodology Reviewer

Mercor • Philadelphia

On-site
USD 120,000 - 180,000
Applied ML Task Auditor & Rubric Evaluator
Applied ML Task Auditor & Rubric Evaluator

Obsidian • San Francisco (CA)

On-site
USD 120,000 - 210,000
ML Research Engineer - PhD - AI Trainer
ML Research Engineer - PhD - AI Trainer

Obsidian • Seattle (WA)

Hybrid
USD 100,000 - 150,000