SWE-Bench Task Auditor

HumanitApp

Northern (KY)

Hybrid

USD 96,000 - 124,000

Full time

7 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

HumanitApp is seeking a role focused on evaluating benchmark tasks used to train and assess frontier AI models, focusing on quality and reproducibility.

You will assess tasks, patches, and harnesses while delivering rubric-based feedback to ensure rigorous evaluation standards.

This remote-friendly opportunity offers compensation ranging from $70 to $90 per hour, with workload and terms clarified during the official application process.

Responsibilities

  • Evaluate the quality, correctness, and reproducibility of software-engineering benchmark tasks used to train and evaluate frontier AI models.
  • Review repository-level tasks, reference patches, test harnesses, and grading integrity.
  • Provide rubric-based written feedback.

Skills

AI evaluation
Python
Software engineering

Job description

Evaluate the quality, correctness, and reproducibility of software-engineering benchmark tasks used to train and evaluate a frontier AI lab's models. You'll assess repository-level tasks, reference patches, test harnesses, and grading integrity

and provide clear, rubric-based written feedback.

Basic Qualificatio...

AI evaluation Python Software engineering

Who this role may fit

This opportunity may suit professionals with relevant experience in AI evaluation, Python, Software engineering. Review the official description and requirements before applying.

Compensation context

The listing states $70 - $90 / hour. Confirm the final rate, workload, and payment terms during the official application process.

  • You can show recent, specific work involving AI evaluation, Python, Software engineering.
  • The listed commitment of 40 hours/week fits your schedule.
  • You can communicate your reasoning clearly in the application language.
  • You are comfortable with project availability and hours varying over time.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Software Engineer - Benchmark Auditor
Software Engineer - Benchmark Auditor

Mercor • San Francisco (CA)

On-site
USD 150,000 - 190,000
Software Engineer - Benchmark Auditor
Software Engineer - Benchmark Auditor

Obsidian • San Francisco (CA)

On-site
USD 140,000 - 190,000
AI Benchmark Auditor (Python & Evaluation)
AI Benchmark Auditor (Python & Evaluation)

HumanitApp • Northern (KY)

Hybrid
USD 96,000 - 124,000
Kubernetes Task Auditor
Kubernetes Task Auditor

HumanitApp • Northern (KY)

Hybrid
USD 96,000 - 124,000
AI Developer Trace Task Auditor
AI Developer Trace Task Auditor

HumanitApp • Northern (KY)

Hybrid
USD 96,000 - 124,000
ML Challenge Task Auditor
ML Challenge Task Auditor

HumanitApp • Northern (KY)

Hybrid
USD 96,000 - 124,000
AWS Serverless & Infrastructure-as-Code Task Auditor
AWS Serverless & Infrastructure-as-Code Task Auditor

HumanitApp • Northern (KY)

Hybrid
USD 96,000 - 124,000
AI Benchmark Quality Engineer for Task Evaluation
AI Benchmark Quality Engineer for Task Evaluation

Obsidian • San Francisco (CA)

Remote
USD 115,000 - 160,000
Remote SWE-Bench AI Task Auditor - Freelance
Remote SWE-Bench AI Task Auditor - Freelance

Meridial • Austin (TX)

On-site
USD 68,000 - 97,000
AI Evaluation Specialist: ML Task Auditor
AI Evaluation Specialist: ML Task Auditor

HumanitApp • Northern (KY)

Hybrid
USD 96,000 - 124,000