AI Benchmarking & Evaluation Engineer

Apply

San Francisco, Northern (CA, KY)

Hybrid

USD 150,000 - 210,000

Full time

5 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Bi-annual bonus
Equity grant
Relocation bonus
Housing bonus
Meals stipend
Equinox membership
Laundry reimbursement
Wellness reimbursement
Health insurance
Dental insurance
Vision insurance

Job summary

Mercor in San Francisco is seeking a Research Engineer to own benchmarking pipelines, evaluation systems, and failure-analysis workflows that directly inform training of frontier language models.

You will design and run evals, build rubrics and scorers, and transform failure analysis into actionable improvements for post-training and data pipelines. Strong collaboration across AI researchers and data teams is essential.

Qualifications

  • Strong applied research background focused on model evaluation, benchmarking, and failure analysis.
  • Hands-on coding experience with ML models and evaluation code.
  • Solid grasp of data structures, algorithms and backend systems.
  • Comfort with APIs, SQL/NoSQL, and cloud platforms for eval results.
  • Ability to reason about model behavior from evals and data quality.
  • Willingness to work in person in San Francisco five days a week.

Responsibilities

  • Design, implement, and maintain benchmarks and metrics for tool usage, agentic behavior, and real-world reasoning.
  • Build end-to-end LLM evaluation systems with dashboards and reporting for scale.
  • Perform systematic failure analyses on model outputs and feed findings into improvements.
  • Create rubrics, evaluators, and scoring frameworks balancing rigor and scalability.
  • Quantify data usability and impact on benchmarks to guide data generation and curation.
  • Collaborate with researchers, AI teams and data producers to align evals with training goals.
  • Own benchmarks and evaluation workflows in a fast-paced environment.

Skills

Applied research
Model evaluation
Benchmarking
Failure analysis
Coding skills
APIs
SQL/NoSQL
Cloud platforms

Tools

Python
ML eval frameworks

Job description

Mercor in San Francisco is seeking a Research Engineer to own benchmarking pipelines, evaluation systems, and failure-analysis workflows that directly inform training of frontier language models.

You will design and run evals, build rubrics and scorers, and transform failure analysis into actionable improvements for post-training and data pipelines. Strong collaboration across AI researchers and data teams is essential.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Benchmarking Engineer — Evaluation & Failure Analysis
AI Benchmarking Engineer — Evaluation & Failure Analysis

Doist • San Francisco (CA)

On-site
USD 150,000 - 210,000
Bi-annual bonus
Equity grant
Relocation bonus
+6
AI Benchmarking Engineer — Evaluations & Failure Analysis
AI Benchmarking Engineer — Evaluations & Failure Analysis

Mercor • San Francisco (CA)

On-site
USD 120,000 - 160,000
Generous equity grant vested over 4 years
$10K housing bonus
$1.5K monthly stipend for meals
+2
Research Engineer – Benchmarking, Evals & Failure Analysis
Research Engineer – Benchmarking, Evals & Failure Analysis

Mercor • San Francisco (CA)

On-site
USD 120,000 - 160,000
Generous equity grant vested over 4 years
$10K housing bonus
$1.5K monthly stipend for meals
+2
Research Engineer – Benchmarking
Research Engineer – Benchmarking

Mercor • San Francisco (CA)

On-site
USD 150,000 - 210,000
Bi-annual bonus
Equity grant
Relocation bonus
+6
AI Research Engineer: Evaluation & Benchmarking
AI Research Engineer: Evaluation & Benchmarking

Aceolution • United States

On-site
USD 120,000 - 180,000
Research Engineer – Benchmarking
Research Engineer – Benchmarking

Apply • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 210,000
Bi-annual bonus
Equity grant
Relocation bonus
+8
AI Benchmarking & Strategy Lead
AI Benchmarking & Strategy Lead

Artificial Analysis • San Francisco (CA)

On-site
USD 180,000 - 260,000
Equity
Competitive compensation
Benchmark Architect for AI Evaluation and Research
Benchmark Architect for AI Evaluation and Research

Vals AI, Inc. • San Francisco (CA)

On-site
USD 150,000 - 260,000
Relocation support
Health insurance
Meals provided
+2
AI Benchmark Designer: Computational Statistics Expert
AI Benchmark Designer: Computational Statistics Expert

Mercor • Seattle (WA)

On-site
USD 60,000 - 120,000
AI Benchmark Engineer — PhD (Remote & Flexible)
AI Benchmark Engineer — PhD (Remote & Flexible)

Mercor • New York (NY)

On-site
USD 55,000 - 110,000