Evaluation Engineer: AI Coding Benchmarks & Tests

Mercor

United States

Remote

USD 120,000 - 180,000

Full time

8 days ago
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Mercor is seeking software engineers to build and refine evaluations for frontier models. You’ll transform completed PRs into engineering tasks and use our in‑house framework to benchmark across models such as Claude Code and Codex.

You’ll collaborate with an in‑house research team, write prompts and tests, and help refine grading criteria. Strong programming skills, experience with large codebases, and clear communication are essential to own task development and iteration.

Qualifications

  • Proven ability to work with large, complex codebases.
  • Experience writing tests and evaluation prompts.
  • Excellent communication and independent problem solving.

Responsibilities

  • Identify repositories with substantial work to support evaluations.
  • Write task prompts and build evaluations using tests and prompts.
  • Validate evaluations and refine grading criteria.
  • Run evaluations across multiple models and harnesses.
  • Document findings and iterate on feedback.

Skills

Programming fundamentals
Codebase navigation
AI tooling experience
Testing skills
Git & CI

Tools

Git
Test runners
CI tooling

Job description

Mercor is seeking software engineers to build and refine evaluations for frontier models. You’ll transform completed PRs into engineering tasks and use our in‑house framework to benchmark across models such as Claude Code and Codex.

You’ll collaborate with an in‑house research team, write prompts and tests, and help refine grading criteria. Strong programming skills, experience with large codebases, and clear communication are essential to own task development and iteration.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Applied Engineer, Evaluations Role
Applied Engineer, Evaluations Role

Mercor • United States

Remote
USD 120,000 - 180,000
AI Evaluation Scientist Math PhD Frontier Model Benchmark
AI Evaluation Scientist Math PhD Frontier Model Benchmark

Obsidian • San Francisco (CA)

On-site
USD 96,000 - 179,000
ML Engineer: Frontier AI Coding Evaluator
ML Engineer: Frontier AI Coding Evaluator

Mercor • Miami (FL)

On-site
USD 455,000 - 647,000
AI Evaluation Scientist (PhD) — Math & Frontiers Benchmarking
AI Evaluation Scientist (PhD) — Math & Frontiers Benchmarking

Obsidian • Miami (FL)

On-site
USD 83,000 - 124,000
AI-Driven DevOps Engineer & Model Evaluator
AI-Driven DevOps Engineer & Model Evaluator

Mercor • San Francisco (CA)

On-site
USD 15,000 - 22,000
Senior Software Benchmark Auditor - Code Quality & Integrity
Senior Software Benchmark Auditor - Code Quality & Integrity

Mercor • San Francisco (CA)

On-site
USD 150,000 - 190,000
AI Benchmark Architect: Computational Mathematician
AI Benchmark Architect: Computational Mathematician

Mercor • San Francisco (CA)

Remote
USD 83,000 - 165,000
AI Evaluation Scientist: Math PhD for Frontier Benchmarks
AI Evaluation Scientist: Math PhD for Frontier Benchmarks

Mercor • San Francisco (CA)

On-site
USD 11,021,000 - 13,776,000
AI Evaluations Engineer — Benchmarking Frontiers
AI Evaluations Engineer — Benchmarking Frontiers

Meta • Menlo Park (CA)

On-site
USD 180,000 - 240,000
AI Coding Trace Auditor - Rubric-Based Feedback
AI Coding Trace Auditor - Rubric-Based Feedback

Mercor • San Francisco (CA)

On-site
USD 120,000 - 180,000