Applied Engineer, Evaluations Role

Mercor

United States

Remote

USD 120,000 - 180,000

Full time

8 days ago
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Mercor is seeking software engineers to build and refine evaluations for frontier models. You’ll transform completed PRs into engineering tasks and use our in‑house framework to benchmark across models such as Claude Code and Codex.

You’ll collaborate with an in‑house research team, write prompts and tests, and help refine grading criteria. Strong programming skills, experience with large codebases, and clear communication are essential to own task development and iteration.

Qualifications

  • Proven ability to work with large, complex codebases.
  • Experience writing tests and evaluation prompts.
  • Excellent communication and independent problem solving.

Responsibilities

  • Identify repositories with substantial work to support evaluations.
  • Write task prompts and build evaluations using tests and prompts.
  • Validate evaluations and refine grading criteria.
  • Run evaluations across multiple models and harnesses.
  • Document findings and iterate on feedback.

Skills

Programming fundamentals
Codebase navigation
AI tooling experience
Testing skills
Git & CI

Tools

Git
Test runners
CI tooling

Job description

Mercor is seeking software engineers to build and refine evaluations. You’ll turn completed pull requests into engineering tasks and use our in‑house evaluation framework to multiple frontier coding agents, including Claude Code, Codex and more. We work alongside an in‑house research team and have produced industry leading benchmarks to compare the performance of frontier large language models.

In this role, you will:
  • Identify repositories with enough substantive work to support challenging evaluations, and select suitable completed PRs.
  • Investigate each problem, its reference solution, and the repository’s architecture, tests, and conventions.
  • Write task prompts and build evaluations using automated tests, shell commands, and LLM grading prompts.
  • Validate evaluations against reference solutions and deliberately flawed implementations.
  • Find missing checks, incorrect grades, and criteria that unnecessarily constrain how a problem can be solved.
  • Run evaluations repeatedly across multiple models and harnesses.
  • Investigate whether failures come from the agent’s solution, the environment, or the grading, and establish that tasks expose meaningful weaknesses in agent performance.
  • Refine evaluations through repeated testing and review.
  • Document findings and grading decisions, and work through feedback.
You’ll need:
  • Strong programming fundamentals and practical experience working in substantial and complex codebases.
  • Languages we create evals for include TypeScript/JavaScript, Python, Java, Kotlin, Go, Ruby, PHP, C++ and Rust.
  • The ability to understand unfamiliar code, investigate subtle behavior, and assess whether different implementations solve the same problem correctly.
  • Experience writing meaningful tests, including edge cases and regression coverage.
  • Attention to detail and patience for repeated investigation and refinement.
  • Practical experience using AI coding tools, with the judgment to verify their output and catch mistakes.
  • Confidence using Git, test runners, and CI tooling.
  • Clear and fluent written and verbal communication with the ability to own a task independently while raising questions when requirements are ambiguous.
  • Relevant experience can come from open-source, private, or enterprise repositories.
  • Familiarity with a particular language or ecosystem is helpful; the ability to learn the repository and make sound engineering judgments matters more than its popularity or your public contribution history.

We provide onboarding to the evaluation framework and ongoing review feedback. After onboarding, you’ll be expected to own task selection, evaluation development, and iteration without step‑by‑step direction.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Evaluation Engineer: AI Coding Benchmarks & Tests
Evaluation Engineer: AI Coding Benchmarks & Tests

Mercor • United States

Remote
USD 120,000 - 180,000
ML Engineer - Coding Agent Expert
ML Engineer - Coding Agent Expert

Obsidian • New York (NY)

On-site
USD 275,520 - 826,560
Senior Software Engineer - 35501
Senior Software Engineer - 35501

Turing • San Francisco (CA)

On-site
USD 68,880 - 206,640
Evals Engineer, Offensive Cyber
Evals Engineer, Offensive Cyber

Zealot Labs • New York (NY)

On-site
USD 140,000 - 200,000
ML Engineer - AI Coding Expert
ML Engineer - AI Coding Expert

Mercor • New York (NY)

On-site
USD 220,000 - 551,000
Frontend Engineering AI Evaluator
Frontend Engineering AI Evaluator

AI Trainer Jobs • United States

Remote
USD 55,000 - 96,000
AI-Driven DevOps Engineer & Model Evaluator
AI-Driven DevOps Engineer & Model Evaluator

Mercor • San Francisco (CA)

On-site
USD 15,000 - 22,000
ML Engineer: Frontier AI Coding Evaluator
ML Engineer: Frontier AI Coding Evaluator

Mercor • Miami (FL)

On-site
USD 455,000 - 647,000
Senior Software Engineer - 35501
Senior Software Engineer - 35501

Turing • New York (NY)

On-site
USD 68,880 - 206,640
Agent Evaluation Engineer — Build-Time Framework & Deployment Gates
Agent Evaluation Engineer — Build-Time Framework & Deployment Gates

EPAM Systems Inc • United States

Remote
USD 140,000 - 190,000