AI Evaluation Specialist: ML Task Auditor

HumanitApp

Northern (KY)

Hybrid

USD 96,000 - 124,000

Full time

7 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

HumanitApp is seeking an evaluator to review the quality and rigor of applied ML tasks used to train and assess frontier AI models. You will examine experiment design, model-selection reasoning, and evaluation methodology, delivering clear rubric-based feedback to support research teams.

The role requires 3+ years in AI evaluation or related fields, with emphasis on rigorous methodological critique and precise written communication. Expect a 40-hour work week at a competitive hourly rate.

Qualifications

  • 3+ years of experience in AI evaluation or related field.
  • Ability to assess experiment designs and evaluation methodologies.
  • Strong written communication for rubric-based feedback.

Responsibilities

  • Evaluate quality, correctness, and methodological rigor of applied ML tasks used to train and evaluate models.
  • Assess experiment design, model-selection reasoning, and evaluation methodology.
  • Provide rubric-based written feedback to inform improvements and decision-making.

Skills

AI evaluation
Machine learning
Research

Job description

HumanitApp is seeking an evaluator to review the quality and rigor of applied ML tasks used to train and assess frontier AI models. You will examine experiment design, model-selection reasoning, and evaluation methodology, delivering clear rubric-based feedback to support research teams.

The role requires 3+ years in AI evaluation or related fields, with emphasis on rigorous methodological critique and precise written communication. Expect a 40-hour work week at a competitive hourly rate.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

ML Challenge Task Auditor
ML Challenge Task Auditor

HumanitApp • Northern (KY)

Hybrid
USD 96,000 - 124,000
Applied ML Task Auditor & Rubric Evaluator
Applied ML Task Auditor & Rubric Evaluator

Obsidian • San Francisco (CA)

On-site
USD 120,000 - 210,000
Applied ML Evaluator & AI Rigor Reviewer
Applied ML Evaluator & AI Rigor Reviewer

Obsidian • Philadelphia

On-site
USD 120,000 - 170,000
ML Task Auditor - Applied ML
ML Task Auditor - Applied ML

Obsidian • San Francisco (CA)

On-site
USD 120,000 - 210,000
Machine Learning Evaluator - AI Trainer
Machine Learning Evaluator - AI Trainer

Mercor • Philadelphia

On-site
USD 120,000 - 180,000
ML Task Auditor - Applied ML
ML Task Auditor - Applied ML

Mercor • San Francisco (CA)

On-site
USD 130,000 - 190,000
Kubernetes Task Auditor for AI Evaluation
Kubernetes Task Auditor for AI Evaluation

HumanitApp • Northern (KY)

Hybrid
USD 96,000 - 124,000
AI Benchmark Auditor (Python & Evaluation)
AI Benchmark Auditor (Python & Evaluation)

HumanitApp • Northern (KY)

Hybrid
USD 96,000 - 124,000
Applied ML Evaluation Auditor
Applied ML Evaluation Auditor

Obsidian • San Francisco (CA)

Remote
USD 90,000 - 130,000
AI/ML Evaluator - Human-in-the-Loop (Freelance)
AI/ML Evaluator - Human-in-the-Loop (Freelance)

Welocalize • Town of Texas (WI)

On-site
USD 192,864,000 - 220,416,000