Research Engineer: Model Evaluations & Metrics

Anthropic

United States

Remote

USD 120,000 - 230,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Anthropic is seeking Research Engineers to design evaluations that quantify Claude's capabilities, reasoning, safety properties, and alignment with leadership expectations.

You will build scalable eval infrastructure, run experiments across live checkpoints, and present clear results to researchers and decision-makers to push Anthropic toward leadership in well-characterized AI systems.

Qualifications

  • Strong Python programming skills, production or research infra experience.
  • Experience building or operating distributed systems and data pipelines.
  • Clear written and verbal communication to explain technical results to non-specialists or leadership.

Responsibilities

  • Design and run evaluations of Claude's capabilities, safety, and behavior.
  • Build and harden a scalable evaluation platform used during training and testing.
  • Own dashboards and visualizations to monitor model health and evaluation outcomes.
  • Debug anomalous eval results during training and communicate findings clearly under time pressure.
  • Collaborate with research teams across the full lifecycle of new capabilities, from measurement to interpretation.

Skills

Python
Distributed Systems
Communication

Job description

Anthropic is seeking Research Engineers to design evaluations that quantify Claude's capabilities, reasoning, safety properties, and alignment with leadership expectations.

You will build scalable eval infrastructure, run experiments across live checkpoints, and present clear results to researchers and decision-makers to push Anthropic toward leadership in well-characterized AI systems.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Research Engineer, AI Evaluation & Metrics
Research Engineer, AI Evaluation & Metrics

Menlo Ventures • New York (NY)

On-site
USD 500,000 - 850,000
Research Engineer, Model Evaluations
Research Engineer, Model Evaluations

Anthropic • United States

Remote
USD 120,000 - 230,000
Research Engineer – AI Evaluation & Metrics
Research Engineer – AI Evaluation & Metrics

EngineersOfAI • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 250,000
Research Engineer, Model Evaluations
Research Engineer, Model Evaluations

Anthropic • New York (NY), San Francisco (CA)

Hybrid
USD 320,000 - 485,000
Generous vacation and parental leave
Flexible working hours
Lovely office space for collaboration
Research Engineer: AI Evaluation & Metrics
Research Engineer: AI Evaluation & Metrics

Anthropic • San Francisco (CA)

Hybrid
USD 350,000 - 850,000
Equity donation matching
Generous vacation and parental leave
Flexible working hours
+1
Research Engineer, Model Evaluations
Research Engineer, Model Evaluations

Menlo Ventures • New York (NY)

On-site
USD 500,000 - 850,000
Research Engineer, Takeoff Intel
Research Engineer, Takeoff Intel

EngineersOfAI • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 250,000
Research Engineer, Model Evaluations - Remote-Friendly Impact
Research Engineer, Model Evaluations - Remote-Friendly Impact

Menlo Ventures • San Francisco (CA)

On-site
USD 320,000 - 485,000
Generous vacation and parental leave
Flexible working hours
Lovely office space for collaboration
Staff Software Engineer, AI Platforms for Science
Staff Software Engineer, AI Platforms for Science

CDFAM - Computational Design Symposium • San Francisco (CA), Northern (KY)

Hybrid
USD 405,000 - 485,000
Pre-training Distributed Systems Tech Lead / Manager
Pre-training Distributed Systems Tech Lead / Manager

Anthropic • San Francisco (CA)

Hybrid
USD 500,000 - 850,000
Equity donation matching
Generous vacation and parental leave
Flexible working hours
+1