Senior Software Benchmark Auditor - Code Quality & Integrity

Mercor

San Francisco (CA)

On-site

USD 150,000 - 190,000

Full time

11 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Mercor seeks a software engineering quality evaluator to assess benchmark tasks used to train frontier AI models. You will audit repository-level tasks, reference patches, test harnesses, and grading integrity, providing rubric-based written feedback.

The role requires 3+ years in software engineering, open-source contribution or maintainer experience, and strong ability to audit patches, test runners, and Docker isolation.

Qualifications

  • 3+ years of professional software engineering.
  • Real open-source contribution or maintainer experience (merged PRs, committer / maintainer roles).
  • Strong ability to audit reference patches, test runners, and Docker isolation, and to detect answer leakage / reward hacking.
  • Fluency across common ecosystems (Python and at least one of Java / Go / TypeScript / C++).

Responsibilities

  • Evaluate the quality, correctness, and reproducibility of benchmark tasks used to train frontier AI models.
  • Provide rubric-based written feedback and assess repository-level tasks, reference patches, test harnesses, and grading integrity.

Skills

Software engineering
Open-source contribution
Code review
Audit reference patches
Test harnesses
Python
Java
Go

Tools

Docker
CI/CD

Job description

Mercor seeks a software engineering quality evaluator to assess benchmark tasks used to train frontier AI models. You will audit repository-level tasks, reference patches, test harnesses, and grading integrity, providing rubric-based written feedback.

The role requires 3+ years in software engineering, open-source contribution or maintainer experience, and strong ability to audit patches, test runners, and Docker isolation.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

SWE Benchmark Auditor for AI Model Evaluation
SWE Benchmark Auditor for AI Model Evaluation

Dorado • Northern (KY)

Hybrid
USD 100,000 - 150,000
Software Benchmark Auditor: Reproducibility & Quality
Software Benchmark Auditor: Reproducibility & Quality

Obsidian • San Francisco (CA)

On-site
USD 140,000 - 190,000
Frontier AI Benchmark Task Auditor
Frontier AI Benchmark Task Auditor

Mercor • United States

Remote
USD 110,000 - 170,000
Software Engineer - Benchmark Auditor
Software Engineer - Benchmark Auditor

Obsidian • San Francisco (CA)

On-site
USD 140,000 - 190,000
Software Engineer - Benchmark Auditor
Software Engineer - Benchmark Auditor

Mercor • San Francisco (CA)

On-site
USD 150,000 - 190,000
AI Benchmark Quality Engineer for Task Evaluation
AI Benchmark Quality Engineer for Task Evaluation

Obsidian • San Francisco (CA)

Remote
USD 115,000 - 160,000
SWE-Bench Task Auditor Mercor · Remote — United States $70-90/hr →
SWE-Bench Task Auditor Mercor · Remote — United States $70-90/hr →

Dorado • Northern (KY)

Hybrid
USD 100,000 - 150,000
Remote SWE-Bench Auditor - Part-Time, $70-$90/hr
Remote SWE-Bench Auditor - Part-Time, $70-$90/hr

OpenTrain AI • Northern (KY)

Hybrid
USD 96,000 - 124,000
Remote contract role (US)
40 hours/week target (20+ hours)
Competitive hourly rate $70–$90
SWE-Bench Task Auditor
SWE-Bench Task Auditor

OpenTrain AI • Northern (KY)

Hybrid
USD 96,000 - 124,000
Remote contract role (US)
40 hours/week target (20+ hours)
Competitive hourly rate $70–$90
AI Code Trace Auditor — Quality Feedback
AI Code Trace Auditor — Quality Feedback

Mercor • United States

Remote
USD 90,000 - 130,000