Frontier AI Benchmark Task Auditor

Mercor

United States

Remote

USD 110,000 - 170,000

Full time

11 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Mercor seeks a professional to evaluate the quality, correctness, and reproducibility of software-engineering benchmark tasks used to train and evaluate frontier AI models. You will assess repository-level tasks, reference patches, test harnesses, and grading integrity.

You will deliver clear rubric-based written feedback, outlining strengths, gaps, and concrete improvements to ensure reliable benchmarking and reproducibility across teams working with open-source tools.

Qualifications

  • 3+ years of professional software engineering experience.
  • Real open-source contribution or maintainer experience.
  • Ability to audit reference patches, test runners, and Docker isolation.
  • Fluency in Python and at least one of Java / Go / TypeScript / C++.

Responsibilities

  • Evaluate benchmark task quality, correctness, and reproducibility for frontier AI models.
  • Assess repository-level tasks, reference patches, test harnesses, and grading integrity.
  • Provide rubric-based written feedback.

Skills

Software engineering
Open-source contribution
Code review
Docker auditing
Python
Multi-language ecosystem

Tools

Docker

Job description

Mercor seeks a professional to evaluate the quality, correctness, and reproducibility of software-engineering benchmark tasks used to train and evaluate frontier AI models. You will assess repository-level tasks, reference patches, test harnesses, and grading integrity.

You will deliver clear rubric-based written feedback, outlining strengths, gaps, and concrete improvements to ensure reliable benchmarking and reproducibility across teams working with open-source tools.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Software Benchmark Auditor - Code Quality & Integrity
Senior Software Benchmark Auditor - Code Quality & Integrity

Mercor • San Francisco (CA)

On-site
USD 150,000 - 190,000
AI Benchmark Quality Engineer for Task Evaluation
AI Benchmark Quality Engineer for Task Evaluation

Obsidian • San Francisco (CA)

Remote
USD 115,000 - 160,000
Software Benchmark Auditor: Reproducibility & Quality
Software Benchmark Auditor: Reproducibility & Quality

Obsidian • San Francisco (CA)

On-site
USD 140,000 - 190,000
SWE Benchmark Auditor for AI Model Evaluation
SWE Benchmark Auditor for AI Model Evaluation

Dorado • Northern (KY)

Hybrid
USD 100,000 - 150,000
AI Code Trace Auditor — Quality Feedback
AI Code Trace Auditor — Quality Feedback

Mercor • United States

Remote
USD 90,000 - 130,000
AI Benchmark Architect for Scientific Computing
AI Benchmark Architect for Scientific Computing

Obsidian • San Francisco (CA)

Remote
USD 96,000 - 138,000
6-week engagement
Part-time 20+ hrs/week
Immediate start
Software Engineer - Benchmark Auditor
Software Engineer - Benchmark Auditor

Obsidian • San Francisco (CA)

On-site
USD 140,000 - 190,000
Software Engineer - Benchmark Auditor
Software Engineer - Benchmark Auditor

Mercor • San Francisco (CA)

On-site
USD 150,000 - 190,000
SWE-Bench Task Auditor Mercor · Remote — United States $70-90/hr →
SWE-Bench Task Auditor Mercor · Remote — United States $70-90/hr →

Dorado • Northern (KY)

Hybrid
USD 100,000 - 150,000
AI-Driven DevOps Evaluator for Frontier Code Models
AI-Driven DevOps Evaluator for Frontier Code Models

Mercor • San Francisco (CA)

Hybrid
USD 55,000 - 91,000