Research Engineer, AI Evaluation & Metrics

Menlo Ventures

New York (NY)

On-site

USD 500,000 - 850,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Anthropic is seeking Research Engineers to design evaluations that quantify Claude's capabilities and safety, turning abstract notions of intelligence into clear, defendable metrics.

You will build and operate the distributed evaluation platform, run experiments across prompts and checkpoints, and partner with researchers to interpret results for leadership and public audiences.

The role supports scalable, well-characterized AI systems and dashboards that monitor health during training.

Qualifications

  • Proficient Python programming for production or research infrastructure.
  • Experience building or operating distributed systems at scale.
  • Clear written and verbal communication, especially with non-specialists.
  • Willingness to work in on-call or production-support capacity.
  • Interest in steering AI to be safe and beneficial.

Responsibilities

  • Design and run new evaluations of Claude's capabilities — reasoning, agentic behavior, knowledge, safety properties — and produce visualizations that make the results legible to researchers and decision-makers
  • Build and harden the distributed eval execution platform so hundreds of evals run reliably against checkpoints throughout production RL training runs
  • Own the dashboards researchers and leadership use to monitor model health during training, improving signal‑to‑noise, reducing latency, and making regressions impossible to miss
  • Debug anomalous eval results mid‑training‑run, determine whether the cause is a model change or an infrastructure issue, and communicate the answer clearly under time pressure
  • Improve the tooling, libraries, and workflows researchers use to implement and iterate on evaluations
  • Partner with research teams across the full lifecycle of a new capability — from defining what to measure to interpreting results as training progresses
  • Run experiments to characterize how prompting, sampling, and scaffolding choices affect results on internal and industry benchmarks
  • Communicate evaluations and their results to internal stakeholders and, where appropriate, external audiences

Skills

Python programming
Distributed systems
Technical communication
Production support
On-call readiness
Societal impact awareness

Education

Bachelor's degree

Job description

Anthropic is seeking Research Engineers to design evaluations that quantify Claude's capabilities and safety, turning abstract notions of intelligence into clear, defendable metrics.

You will build and operate the distributed evaluation platform, run experiments across prompts and checkpoints, and partner with researchers to interpret results for leadership and public audiences.

The role supports scalable, well-characterized AI systems and dashboards that monitor health during training.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Research Engineer: Model Evaluations & Metrics
Research Engineer: Model Evaluations & Metrics

Anthropic • United States

Remote
USD 120,000 - 230,000
Research Engineer, Model Evaluations
Research Engineer, Model Evaluations

Anthropic • United States

Remote
USD 120,000 - 230,000
Research Engineer – AI Evaluation & Metrics
Research Engineer – AI Evaluation & Metrics

EngineersOfAI • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 250,000
Research Engineer: AI Evaluation & Metrics
Research Engineer: AI Evaluation & Metrics

Anthropic • San Francisco (CA)

Hybrid
USD 350,000 - 850,000
Equity donation matching
Generous vacation and parental leave
Flexible working hours
+1
Research Engineer, Model Evaluations
Research Engineer, Model Evaluations

Menlo Ventures • New York (NY)

On-site
USD 500,000 - 850,000
Research Engineer, Model Evaluations
Research Engineer, Model Evaluations

Anthropic • New York (NY), San Francisco (CA)

Hybrid
USD 320,000 - 485,000
Generous vacation and parental leave
Flexible working hours
Lovely office space for collaboration
Research Engineer: AI for Safe, Real-World Software
Research Engineer: AI for Safe, Real-World Software

Anthropic • San Francisco (CA)

On-site
USD 500,000 - 850,000
Research Engineer, Takeoff Intel
Research Engineer, Takeoff Intel

EngineersOfAI • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 250,000
Staff Engineer - AI Evaluation & Metrics Platform
Staff Engineer - AI Evaluation & Metrics Platform

Kindredventures • San Francisco (CA)

On-site
USD 140,000 - 200,000
Research Engineer, AI Computer Use & Vision
Research Engineer, AI Computer Use & Vision

Anthropic • United States

Hybrid
USD 500,000 - 850,000
Competitive compensation
Optional equity donation matching
Generous vacation and parental leave
+2