Research Engineer: AI Evaluation & Metrics

Anthropic

San Francisco (CA)

Hybrid

USD 350,000 - 850,000

Full time

3 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Equity donation matching
Generous vacation and parental leave
Flexible working hours
Lovely office space

Job summary

Anthropic is hiring a Research Engineer for the Takeoff Intel team in San Francisco. You will build evaluation instruments, process large-scale telemetry, and help quantify AI capability growth to inform safety and policy work.

The role spans evals infrastructure, data processing, and analysis tooling with an emphasis on shipping reliable results. We value rapid prototyping, accuracy, and the ability to turn vague questions into running instruments.

Qualifications

  • You may be a good fit if you have shipped an evaluation, data product, or research library end to end.
  • Prototype fast and are comfortable throwing code away.
  • Handle messy, large-volume data without over-engineering.
  • Have run experiments on large language models, not just moved their outputs around.
  • Can work from a vague question rather than a spec.
  • Communicate results clearly and collaborate with researchers.

Responsibilities

  • Design, build, and run capability evaluations and measurement instruments at scale
  • Build the data and analysis pipelines that turn large volumes of model outputs and telemetry into reliable metrics
  • Prototype new instruments fast, validate them, and decide what to keep
  • Review and supervise AI-written code as a normal part of the workflow
  • Work closely with research scientists on the team and with partner teams to define what's worth measuring
  • Contribute to internal write-ups and public reporting

Education

Bachelor’s degree or equivalent

Job description

Anthropic is hiring a Research Engineer for the Takeoff Intel team in San Francisco. You will build evaluation instruments, process large-scale telemetry, and help quantify AI capability growth to inform safety and policy work.

The role spans evals infrastructure, data processing, and analysis tooling with an emphasis on shipping reliable results. We value rapid prototyping, accuracy, and the ability to turn vague questions into running instruments.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Research Engineer – AI Evaluation & Metrics
Research Engineer – AI Evaluation & Metrics

EngineersOfAI • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 250,000
Research Engineer, Takeoff Intel — Remote-Friendly
Research Engineer, Takeoff Intel — Remote-Friendly

Anthropic Limited • San Francisco (CA), Northern (KY)

Hybrid
USD 350,000 - 850,000
Research Engineer, Takeoff Intel
Research Engineer, Takeoff Intel

EngineersOfAI • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 250,000
Research Engineer, AI Evaluation & Metrics
Research Engineer, AI Evaluation & Metrics

Menlo Ventures • New York (NY)

On-site
USD 500,000 - 850,000
Research Engineer: Model Evaluations & Metrics
Research Engineer: Model Evaluations & Metrics

Anthropic • United States

Remote
USD 120,000 - 230,000
Research Engineer, Takeoff Intel
Research Engineer, Takeoff Intel

Anthropic • San Francisco (CA)

Hybrid
USD 350,000 - 850,000
Equity donation matching
Generous vacation and parental leave
Flexible working hours
+1
Research Tools Engineer: Build AI Experiment Platforms
Research Tools Engineer: Build AI Experiment Platforms

Anthropic • New York (NY)

Hybrid
USD 300,000 - 405,000
Pre-training Distributed Systems Tech Lead / Manager
Pre-training Distributed Systems Tech Lead / Manager

Anthropic • San Francisco (CA)

Hybrid
USD 500,000 - 850,000
Equity donation matching
Generous vacation and parental leave
Flexible working hours
+1
Research Engineer: AI Safety & Alignment
Research Engineer: AI Safety & Alignment

Anthropic • San Francisco (CA)

Hybrid
USD 350,000 - 500,000
Competitive compensation
Generous vacation and parental leave
Flexible working hours
Staff Engineer - AI Evaluation & Research Execution
Staff Engineer - AI Evaluation & Research Execution

METR • Berkeley (CA)

Hybrid
USD 285,548 - 503,116
Catered lunch and dinner daily
In-office gym and shower
Unlimited PTO
+6