AI Benchmarking Engineer — Evaluation & Failure Analysis

Doist

San Francisco (CA)

On-site

USD 150,000 - 210,000

Full time

5 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Bi-annual bonus
Equity grant
Relocation bonus
Housing allowance
Meals stipend
Equinox membership
Laundry reimbursement
Wellness reimbursement
Healthcare

Job summary

Mercor in San Francisco is seeking a Research Engineer to join the intersection of engineering and applied AI research. You’ll own benchmarking pipelines, evaluation systems, and failure analysis workflows that inform training and improvement of frontier language models.

You’ll design and run evals, build rubrics and scorers, and collaborate with researchers, operators, and data teams to ensure data quality and scalable benchmarks in a fast-paced environment.

Qualifications

  • Strong applied research background in model evaluation, benchmarking, and failure analysis.
  • Strong coding skills and hands-on experience with ML models and evaluation code.
  • Solid grasp of data structures, algorithms, and backend systems.
  • Comfort with APIs, SQL/NoSQL, and cloud platforms for running and storing eval results.
  • Ability to reason about model behavior, experimental results, and data quality from evals and failure analyses.

Responsibilities

  • Benchmarking: Design, implement, and maintain benchmarks and metrics for tool use, agentic behavior, and real-world reasoning; ensure benchmarks scale with training and stay aligned with product and research goals.
  • Evaluation systems: Build and operate LLM evaluation systems end-to-end runs, scoring, dashboards, and reporting, so researchers and applied AI teams can track model performance and compare runs at scale.
  • Failure analysis: Run systematic failure analysis on model outputs (e.g., wrong tool use, reasoning errors, safety/alignment issues); categorize failure modes, quantify prevalence, and feed findings into reward design, data curation, and benchmark design.
  • Rubrics and evaluators: Create and refine rubrics, automated evaluators, and scoring frameworks that drive training and evaluation decisions; balance rigor with scalability (human vs. model-as-judge, calibration, agreement).
  • Data quality and usability: Quantify data usability, quality, and impact on key benchmarks; use evals and failure analysis to guide data generation, augmentation, and curation.
  • Cross-team collaboration: Work with AI researchers, applied AI teams, and data producers to align evals with training objectives and to prioritize benchmarks and failure analyses that matter most.
  • Ownership in a fast-paced environment: Operate in a high-iteration research setting with strong ownership of benchmarks, evals, and failure-analysis workflows.

Job description

Mercor in San Francisco is seeking a Research Engineer to join the intersection of engineering and applied AI research. You’ll own benchmarking pipelines, evaluation systems, and failure analysis workflows that inform training and improvement of frontier language models.

You’ll design and run evals, build rubrics and scorers, and collaborate with researchers, operators, and data teams to ensure data quality and scalable benchmarks in a fast-paced environment.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Benchmarking & Evaluation Engineer
AI Benchmarking & Evaluation Engineer

Apply • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 210,000
Bi-annual bonus
Equity grant
Relocation bonus
+8
AI Benchmarking Engineer — Evaluations & Failure Analysis
AI Benchmarking Engineer — Evaluations & Failure Analysis

Mercor • San Francisco (CA)

On-site
USD 120,000 - 160,000
Generous equity grant vested over 4 years
$10K housing bonus
$1.5K monthly stipend for meals
+2
AI Benchmark Engineer — PhD (Remote & Flexible)
AI Benchmark Engineer — PhD (Remote & Flexible)

Mercor • New York (NY)

On-site
USD 55,000 - 110,000
Research Engineer – Benchmarking, Evals & Failure Analysis
Research Engineer – Benchmarking, Evals & Failure Analysis

Mercor • San Francisco (CA)

On-site
USD 120,000 - 160,000
Generous equity grant vested over 4 years
$10K housing bonus
$1.5K monthly stipend for meals
+2
AI Benchmarking & Strategy Lead
AI Benchmarking & Strategy Lead

Artificial Analysis • San Francisco (CA)

On-site
USD 180,000 - 260,000
Equity
Competitive compensation
AI Research Engineer: Evaluation & Benchmarking
AI Research Engineer: Evaluation & Benchmarking

Aceolution • United States

On-site
USD 120,000 - 180,000
Computational Structural Engineer & AI Benchmark Designer
Computational Structural Engineer & AI Benchmark Designer

Mercor • New York (NY)

Hybrid
USD 110,000 - 170,000
Benchmarking Research Engineer: Frontier Model Evaluations
Benchmarking Research Engineer: Frontier Model Evaluations

Refresh AI • San Francisco (CA)

On-site
USD 120,000 - 150,000
AI Benchmark Designer: Computational Statistics Expert
AI Benchmark Designer: Computational Statistics Expert

Mercor • Seattle (WA)

On-site
USD 60,000 - 120,000
Research Engineer – Benchmarking
Research Engineer – Benchmarking

Mercor • San Francisco (CA)

On-site
USD 150,000 - 210,000
Bi-annual bonus
Equity grant
Relocation bonus
+6