Eval Engineer for AI Model Evaluations

Neura Market

San Francisco, Northern (CA, KY)

Hybrid

USD 500,000 - 850,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Anthropic is seeking Research Engineers to build evaluations for Claude, turning abstract notions of intelligence into measurable metrics. You will design end-to-end evaluations, implement scalable infrastructure, and partner with researchers to interpret results.

Join a collaborative team focusing on producing defensible evaluations, dashboards, and tooling to monitor progress during training runs and to share findings with stakeholders.

Qualifications

  • Strong Python programming skills, including production or research infrastructure.
  • Experience building or operating distributed systems, data pipelines, or other infrastructure that needs to be reliable at scale.
  • Clear written and verbal communication, especially when explaining technical results to non-specialists.
  • Comfort operating in an on-call or production-support capacity when training runs are live.
  • Care about the societal impacts of your work and an interest in steering powerful AI to be safe and beneficial.

Responsibilities

  • Design and run new evaluations of Claude's capabilities — reasoning, agentic behavior, knowledge, safety properties — and produce visualizations that make the results legible to researchers and decision-makers.
  • Build and harden the distributed eval execution platform so hundreds of evals run reliably against checkpoints throughout production RL training runs.
  • Own the dashboards researchers and leadership use to monitor model health during training, improving signal-to-noise, reducing latency, and making regressions impossible to miss.
  • Debug anomalous eval results mid-training-run, determine whether the cause is a model change or an infrastructure issue, and communicate the answer clearly under time pressure.
  • Improve the tooling, libraries, and workflows researchers use to implement and iterate on evaluations.
  • Partner with research teams across the full lifecycle of a new capability — from defining what to measure to interpreting results as training progresses.
  • Run experiments to characterize how prompting, sampling, and scaffolding choices affect results on internal and industry benchmarks.
  • Communicate evaluations and their results to internal stakeholders and, where appropriate, external audiences

Skills

Python programming
Distributed systems
Communication
On-call / production support
Social impact awareness

Education

Bachelor's degree

Job description

Anthropic is seeking Research Engineers to build evaluations for Claude, turning abstract notions of intelligence into measurable metrics. You will design end-to-end evaluations, implement scalable infrastructure, and partner with researchers to interpret results.

Join a collaborative team focusing on producing defensible evaluations, dashboards, and tooling to monitor progress during training runs and to share findings with stakeholders.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Research Engineer, AI Evaluation & Metrics
Research Engineer, AI Evaluation & Metrics

Menlo Ventures • New York (NY)

On-site
USD 500,000 - 850,000
Research Engineer, Knowledge & Evaluation Systems
Research Engineer, Knowledge & Evaluation Systems

Neura Market • San Francisco (CA), Northern (KY)

Hybrid
USD 350,000 - 850,000
Research Engineer, Model Evaluations
Research Engineer, Model Evaluations

Menlo Ventures • New York (NY)

On-site
USD 500,000 - 850,000
Research Engineer, Model Evaluations
Research Engineer, Model Evaluations

Anthropic • New York (NY), San Francisco (CA)

Hybrid
USD 320,000 - 485,000
Generous vacation and parental leave
Flexible working hours
Lovely office space for collaboration
Staff Software Engineer — AI Safety Evaluation Systems
Staff Software Engineer — AI Safety Evaluation Systems

Menlo Ventures • New York (NY)

Hybrid
USD 320,000 - 485,000
Research Engineer - AI Knowledge Architect
Research Engineer - AI Knowledge Architect

Anthropic • United States

Hybrid
USD 350,000 - 850,000
Research Engineer - AI Knowledge Architect
Research Engineer - AI Knowledge Architect

SignalAI • New York (NY)

Hybrid
USD 350,000 - 850,000
Product Manager: AI Model Launch & Eval Leader
Product Manager: AI Model Launch & Eval Leader

Neura Market • San Francisco (CA)

On-site
USD 305,000 - 460,000
Hybrid Software Engineer — AI Safety Evaluations
Hybrid Software Engineer — AI Safety Evaluations

Anthropic • San Francisco (CA)

Hybrid
USD 320,000 - 485,000
Research Engineer, Model Evaluations - Remote-Friendly Impact
Research Engineer, Model Evaluations - Remote-Friendly Impact

Menlo Ventures • San Francisco (CA)

On-site
USD 320,000 - 485,000
Generous vacation and parental leave
Flexible working hours
Lovely office space for collaboration