Applied ML Evaluation Auditor

Obsidian

San Francisco (CA)

Remote

USD 90,000 - 130,000

Full time

11 days ago
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Obsidian is seeking an applied ML evaluator to assess the quality, correctness, and rigor of experiments used to train frontier AI models. You will review experiment design, model-selection reasoning, and evaluation methodology.

This role emphasizes data-quality hygiene, leakage detection, and reproducibility—no MLOps duties. Experience with PyTorch, TensorFlow, scikit-learn, and XGBoost is expected; prior peer review or benchmarking is a plus.

Qualifications

  • 3+ years hands-on applied ML focusing on experiment design, model selection, tuning, and evaluation.
  • Strong data-quality discipline: leakage detection, metric gaming, and proper train/test/CV hygiene.
  • Proficient with PyTorch, TensorFlow, scikit-learn, and XGBoost.
  • Ability to critique ML claims against evidence and reproduce results.

Skills

Experiment design
Model selection
Hyperparameter tuning
Evaluation methodology

Tools

PyTorch
TensorFlow
scikit-learn
XGBoost

Job description

Obsidian is seeking an applied ML evaluator to assess the quality, correctness, and rigor of experiments used to train frontier AI models. You will review experiment design, model-selection reasoning, and evaluation methodology.

This role emphasizes data-quality hygiene, leakage detection, and reproducibility—no MLOps duties. Experience with PyTorch, TensorFlow, scikit-learn, and XGBoost is expected; prior peer review or benchmarking is a plus.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Applied ML Task Auditor & Rubric Evaluator
Applied ML Task Auditor & Rubric Evaluator

Obsidian • San Francisco (CA)

On-site
USD 120,000 - 210,000
Applied ML Task Auditor — Experiment Quality Lead
Applied ML Task Auditor — Experiment Quality Lead

Dorado • Northern (KY)

Hybrid
USD 90,000 - 130,000
ML Task Auditor - Applied ML
ML Task Auditor - Applied ML

Mercor • San Francisco (CA)

On-site
USD 130,000 - 190,000
ML Task Auditor - Applied ML
ML Task Auditor - Applied ML

Obsidian • San Francisco (CA)

On-site
USD 120,000 - 210,000
Applied ML Evaluation Specialist (Remote Contract)
Applied ML Evaluation Specialist (Remote Contract)

OpenTrain AI • Northern (KY)

Hybrid
USD 96,000 - 124,000
ML Challenge Auditor: Rubric-Based Feedback
ML Challenge Auditor: Rubric-Based Feedback

Mercor • United States

Remote
USD 120,000 - 170,000
ML Challenge Task Auditor Mercor · Remote — United States $70-90/hr →
ML Challenge Task Auditor Mercor · Remote — United States $70-90/hr →

Dorado • Northern (KY)

Hybrid
USD 90,000 - 130,000
Frontier AI ML Engineer — Evaluation & Production Systems
Frontier AI ML Engineer — Evaluation & Production Systems

Obsidian • New York (NY)

Remote
USD 4,000 - 7,000
AI Benchmark Quality Engineer for Task Evaluation
AI Benchmark Quality Engineer for Task Evaluation

Obsidian • San Francisco (CA)

Remote
USD 115,000 - 160,000
ML Rigor Auditor for Experimental Tasks
ML Rigor Auditor for Experimental Tasks

Mercor • San Francisco (CA)

On-site
USD 130,000 - 190,000