ML Challenge Task Auditor

HumanitApp

Northern (KY)

Hybrid

USD 96,000 - 124,000

Full time

6 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

HumanitApp is seeking an evaluator to review the quality and rigor of applied ML tasks used to train and assess frontier AI models. You will examine experiment design, model-selection reasoning, and evaluation methodology, delivering clear rubric-based feedback to support research teams.

The role requires 3+ years in AI evaluation or related fields, with emphasis on rigorous methodological critique and precise written communication. Expect a 40-hour work week at a competitive hourly rate.

Qualifications

  • 3+ years of experience in AI evaluation or related field.
  • Ability to assess experiment designs and evaluation methodologies.
  • Strong written communication for rubric-based feedback.

Responsibilities

  • Evaluate quality, correctness, and methodological rigor of applied ML tasks used to train and evaluate models.
  • Assess experiment design, model-selection reasoning, and evaluation methodology.
  • Provide rubric-based written feedback to inform improvements and decision-making.

Skills

AI evaluation
Machine learning
Research

Job description

Evaluate the quality, correctness, and methodological rigor of applied machine-learning tasks used to train and evaluate a frontier AI lab's models. You'll assess experiment design, model-selection reasoning, and evaluation methodology

and provide clear, rubric-based written feedback.

Basic Qualifications • 3+ y...

AI evaluation Machine learning Research

Who this role may fit

This opportunity may suit professionals with relevant experience in AI evaluation, Machine learning, Research. Review the official description and requirements before applying.

Compensation context

The listing states $70 - $90 / hour. Confirm the final rate, workload, and payment terms during the official application process.

  • You can show recent, specific work involving AI evaluation, Machine learning, Research.
  • The listed commitment of 40 hours/week fits your schedule.
  • You can communicate your reasoning clearly in the application language.
  • You are comfortable with project availability and hours varying over time.
Not the right fit?

The listing states $70 - $90 / hour. Confirm the final rate, workload, and payment terms during the official application process.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Evaluation Specialist: ML Task Auditor
AI Evaluation Specialist: ML Task Auditor

HumanitApp • Northern (KY)

Hybrid
USD 96,000 - 124,000
AI Developer Trace Task Auditor
AI Developer Trace Task Auditor

HumanitApp • Northern (KY)

Hybrid
USD 96,000 - 124,000
Kubernetes Task Auditor
Kubernetes Task Auditor

HumanitApp • Northern (KY)

Hybrid
USD 96,000 - 124,000
ML Challenge Task Auditor
ML Challenge Task Auditor

DigiNo • Northern (KY)

Hybrid
USD 100,000 - 180,000
SWE-Bench Task Auditor
SWE-Bench Task Auditor

HumanitApp • Northern (KY)

Hybrid
USD 96,000 - 124,000
Machine Learning Evaluator - AI Trainer
Machine Learning Evaluator - AI Trainer

Mercor • Philadelphia

On-site
USD 120,000 - 180,000
ML Task Auditor - Applied ML
ML Task Auditor - Applied ML

Obsidian • San Francisco (CA)

On-site
USD 120,000 - 210,000
ML Task Auditor - Applied ML
ML Task Auditor - Applied ML

Mercor • San Francisco (CA)

On-site
USD 130,000 - 190,000
AWS Serverless & Infrastructure-as-Code Task Auditor
AWS Serverless & Infrastructure-as-Code Task Auditor

HumanitApp • Northern (KY)

Hybrid
USD 96,000 - 124,000
AI Research Scientist (Remote | $140–$150/hr)
AI Research Scientist (Remote | $140–$150/hr)

Synthires • California (MO)

On-site
USD 193,000 - 207,000
Flexible hours
Asynchronous work
Remote-friendly