Software Engineer - Benchmark Auditor

Mercor

San Francisco (CA)

On-site

USD 150,000 - 190,000

Full time

11 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Mercor seeks a software engineering quality evaluator to assess benchmark tasks used to train frontier AI models. You will audit repository-level tasks, reference patches, test harnesses, and grading integrity, providing rubric-based written feedback.

The role requires 3+ years in software engineering, open-source contribution or maintainer experience, and strong ability to audit patches, test runners, and Docker isolation.

Qualifications

  • 3+ years of professional software engineering.
  • Real open-source contribution or maintainer experience (merged PRs, committer / maintainer roles).
  • Strong ability to audit reference patches, test runners, and Docker isolation, and to detect answer leakage / reward hacking.
  • Fluency across common ecosystems (Python and at least one of Java / Go / TypeScript / C++).

Responsibilities

  • Evaluate the quality, correctness, and reproducibility of benchmark tasks used to train frontier AI models.
  • Provide rubric-based written feedback and assess repository-level tasks, reference patches, test harnesses, and grading integrity.

Skills

Software engineering
Open-source contribution
Code review
Audit reference patches
Test harnesses
Python
Java
Go

Tools

Docker
CI/CD

Job description

Evaluate the quality, correctness, and reproducibility of software-engineering benchmark tasks used to train and evaluate a frontier AI lab's models. You'll assess repository-level tasks, reference patches, test harnesses, and grading integrity — and provide clear, rubric-based written feedback.

Basic Qualifications
  • 3+ years professional software engineering
  • Real open-source contribution or maintainer experience (merged PRs, committer / maintainer roles)
  • Strong ability to audit reference patches, test runners, and Docker isolation, and to detect answer leakage / reward hacking
  • Fluency across common ecosystems (Python and at least one of Java / Go / TypeScript / C++)
Preferred Qualifications
  • Familiarity with SWE-Bench (Verified) or similar repository benchmarks
  • Maintainer history on major Python OSS (Django, Flask, scikit-learn, sympy, pytest, etc.)
  • Prior code-review or task-grading experience
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Engineer - Benchmark Auditor
Software Engineer - Benchmark Auditor

Obsidian • San Francisco (CA)

On-site
USD 140,000 - 190,000
SWE-Bench Task Auditor Mercor · Remote — United States $70-90/hr →
SWE-Bench Task Auditor Mercor · Remote — United States $70-90/hr →

Dorado • Northern (KY)

Hybrid
USD 100,000 - 150,000
SWE Benchmark Auditor for AI Model Evaluation
SWE Benchmark Auditor for AI Model Evaluation

Dorado • Northern (KY)

Hybrid
USD 100,000 - 150,000
Senior Software Benchmark Auditor - Code Quality & Integrity
Senior Software Benchmark Auditor - Code Quality & Integrity

Mercor • San Francisco (CA)

On-site
USD 150,000 - 190,000
SWE-Bench Task Auditor
SWE-Bench Task Auditor

OpenTrain AI • Northern (KY)

Hybrid
USD 96,000 - 124,000
Remote contract role (US)
40 hours/week target (20+ hours)
Competitive hourly rate $70–$90
Software Benchmark Auditor: Reproducibility & Quality
Software Benchmark Auditor: Reproducibility & Quality

Obsidian • San Francisco (CA)

On-site
USD 140,000 - 190,000
Remote SWE-Bench Auditor - Part-Time, $70-$90/hr
Remote SWE-Bench Auditor - Part-Time, $70-$90/hr

OpenTrain AI • Northern (KY)

Hybrid
USD 96,000 - 124,000
Remote contract role (US)
40 hours/week target (20+ hours)
Competitive hourly rate $70–$90
AI Benchmark Quality Engineer for Task Evaluation
AI Benchmark Quality Engineer for Task Evaluation

Obsidian • San Francisco (CA)

Remote
USD 115,000 - 160,000
Frontier AI Benchmark Task Auditor
Frontier AI Benchmark Task Auditor

Mercor • United States

Remote
USD 110,000 - 170,000
Python Engineer – AI Training
Python Engineer – AI Training

YO HR Consultancy • United States

Remote