Software Engineer - Benchmark Auditor

Obsidian

San Francisco (CA)

On-site

USD 140,000 - 190,000

Full time

11 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Obsidian in San Francisco seeks a reviewer and assessor of benchmark tasks for frontier AI labs. You will evaluate the quality, correctness, and reproducibility of software-engineering benchmark tasks used to train and evaluate models, and review repository-level tasks, reference patches, and test harnesses.

You will provide rubric-based written feedback, detect leakage or reward hacking, and help improve grading integrity across the benchmark suite.

Qualifications

  • 3+ years of professional software engineering experience.
  • Real open-source contribution or maintainer experience with merged PRs.
  • Strong ability to audit reference patches, test runners, and Docker isolation.
  • Fluency in Python and at least one of Java, Go, TypeScript, or C++.

Skills

Software engineering
Open-source contribution
Code auditing
Python
Other languages (Java/Go/TypeScript/C\

Tools

Docker
Test harness tooling

Job description

Evaluate the quality, correctness, and reproducibility of software-engineering benchmark tasks used to train and evaluate a frontier AI lab's models. You'll assess repository-level tasks, reference patches, test harnesses, and grading integrity — and provide clear, rubric-based written feedback.

Basic Qualifications
  • 3+ years professional software engineering
  • Real open-source contribution or maintainer experience (merged PRs, committer / maintainer roles)
  • Strong ability to audit reference patches, test runners, and Docker isolation, and to detect answer leakage / reward hacking
  • Fluency across common ecosystems (Python and at least one of Java / Go / TypeScript / C++)
Preferred Qualifications
  • Familiarity with SWE-Bench (Verified) or similar repository benchmarks
  • Maintainer history on major Python OSS (Django, Flask, scikit-learn, sympy, pytest, etc.)
  • Prior code-review or task-grading experience
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Engineer - Benchmark Auditor
Software Engineer - Benchmark Auditor

Mercor • San Francisco (CA)

On-site
USD 150,000 - 190,000
SWE-Bench Task Auditor Mercor · Remote — United States $70-90/hr →
SWE-Bench Task Auditor Mercor · Remote — United States $70-90/hr →

Dorado • Northern (KY)

Hybrid
USD 100,000 - 150,000
SWE Benchmark Auditor for AI Model Evaluation
SWE Benchmark Auditor for AI Model Evaluation

Dorado • Northern (KY)

Hybrid
USD 100,000 - 150,000
Senior Software Benchmark Auditor - Code Quality & Integrity
Senior Software Benchmark Auditor - Code Quality & Integrity

Mercor • San Francisco (CA)

On-site
USD 150,000 - 190,000
SWE-Bench Task Auditor
SWE-Bench Task Auditor

OpenTrain AI • Northern (KY)

Hybrid
USD 96,000 - 124,000
Remote contract role (US)
40 hours/week target (20+ hours)
Competitive hourly rate $70–$90
Software Benchmark Auditor: Reproducibility & Quality
Software Benchmark Auditor: Reproducibility & Quality

Obsidian • San Francisco (CA)

On-site
USD 140,000 - 190,000
Remote SWE-Bench Auditor - Part-Time, $70-$90/hr
Remote SWE-Bench Auditor - Part-Time, $70-$90/hr

OpenTrain AI • Northern (KY)

Hybrid
USD 96,000 - 124,000
Remote contract role (US)
40 hours/week target (20+ hours)
Competitive hourly rate $70–$90
AI Benchmark Quality Engineer for Task Evaluation
AI Benchmark Quality Engineer for Task Evaluation

Obsidian • San Francisco (CA)

Remote
USD 115,000 - 160,000
Frontier AI Benchmark Task Auditor
Frontier AI Benchmark Task Auditor

Mercor • United States

Remote
USD 110,000 - 170,000
Python Engineer – AI Training
Python Engineer – AI Training

YO HR Consultancy • United States

Remote