ML Task Auditor - Applied ML

Obsidian

San Francisco (CA)

On-site

USD 120,000 - 210,000

Full time

11 days ago
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Obsidian is seeking an evaluation-focused ML scientist to assess applied ML rigor in frontier AI research settings. You will critique experiment designs, model-selection reasoning, and evaluation methodologies, delivering rubric-based feedback to improve reproducibility and validity.

Responsibilities emphasize identifying data leakage, metric gaming, and CV hygiene issues, ensuring claims are well-supported by evidence and experiments are replicable across datasets and baselines.

Qualifications

  • 3+ years hands-on applied ML, including experiment design and evaluation.
  • Strong data-quality rigor: leakage detection, metric gaming, train/test/CV hygiene.
  • Proficiency with standard ML frameworks: PyTorch, TensorFlow, scikit-learn, XGBoost.
  • Ability to critique ML claims against evidence and reproduce results.

Responsibilities

  • Evaluate the quality, correctness, and rigor of applied ML tasks used to train and evaluate frontier AI models.
  • Assess experiment design, model-selection reasoning, and evaluation methodology.
  • Provide rubric-based written feedback to improve reproducibility and validity.

Skills

Applied ML
Experiment design
Model evaluation

Tools

PyTorch
TensorFlow
scikit-learn
XGBoost

Job description

Evaluate the quality, correctness, and methodological rigor of applied machine-learning tasks used to train and evaluate a frontier AI lab's models. You'll assess experiment design, model-selection reasoning, and evaluation methodology — and provide clear, rubric-based written feedback.

Basic Qualifications
  • 3+ years hands‑on applied/experimental ML (experiment design, model selection, hyperparameter tuning, evaluation methodology)
  • Strong grasp of data-quality rigor: leakage detection, metric gaming, and train/test/CV hygiene
  • Proficiency with standard ML frameworks (PyTorch, TensorFlow, scikit-learn, XGBoost)
  • Ability to critique ML claims against evidence and reproduce results
Preferred Qualifications
  • Competition / benchmark experience (e.g., Kaggle)
  • Graduate research or publication record in applied ML
  • Prior task-grading or peer-review experience

Note: this role evaluates applied/experimental ML rigor — it is not an LLM-application-building or MLOps role.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Task Auditor - Applied ML
ML Task Auditor - Applied ML

Mercor • San Francisco (CA)

On-site
USD 130,000 - 190,000
ML Challenge Task Auditor Mercor · Remote — United States $70-90/hr →
ML Challenge Task Auditor Mercor · Remote — United States $70-90/hr →

Dorado • Northern (KY)

Hybrid
USD 90,000 - 130,000
Applied ML Task Auditor — Experiment Quality Lead
Applied ML Task Auditor — Experiment Quality Lead

Dorado • Northern (KY)

Hybrid
USD 90,000 - 130,000
Applied ML Evaluation Auditor
Applied ML Evaluation Auditor

Obsidian • San Francisco (CA)

Remote
USD 90,000 - 130,000
ML Challenge Auditor: Rubric-Based Feedback
ML Challenge Auditor: Rubric-Based Feedback

Mercor • United States

Remote
USD 120,000 - 170,000
Applied ML Task Auditor & Rubric Evaluator
Applied ML Task Auditor & Rubric Evaluator

Obsidian • San Francisco (CA)

On-site
USD 120,000 - 210,000
Applied ML Evaluation Specialist (Remote Contract)
Applied ML Evaluation Specialist (Remote Contract)

OpenTrain AI • Northern (KY)

Hybrid
USD 96,000 - 124,000
ML Rigor Auditor for Experimental Tasks
ML Rigor Auditor for Experimental Tasks

Mercor • San Francisco (CA)

On-site
USD 130,000 - 190,000
ML Research Engineer - PhD - AI Trainer
ML Research Engineer - PhD - AI Trainer

Obsidian • Seattle (WA)

Hybrid
USD 100,000 - 150,000
Evaluation & Insights Machine Learning Engineer
Evaluation & Insights Machine Learning Engineer

Apple Inc. • Cupertino (CA)

On-site
USD 150,000 - 230,000