Research Engineer, Takeoff Intel

EngineersOfAI

San Francisco, Northern (CA, KY)

Hybrid

USD 180,000 - 250,000

Full time

5 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Anthropic is hiring a Research Engineer on Takeoff Intel to design and run evaluation and measurement instruments at scale. You will build data pipelines that convert model outputs and telemetry into reliable metrics, prototype fast, and validate approaches before deployment.

You’ll work with research scientists and partner teams to define what’s worth measuring, review AI-written code, and contribute to internal and public reporting.

Qualifications

  • Design, build, and run evaluation and measurement instruments at scale.
  • Prototype fast and validate instruments with rapid iteration.
  • Handle messy, large-volume data without over-engineering.
  • Run experiments on large language models beyond mere data movement.

Responsibilities

  • Design, build, and run capability evaluations and measurement instruments at scale
  • Build data and analysis pipelines turning model outputs into reliable metrics
  • Prototype new instruments quickly, validate, decide what to keep
  • Review and supervise AI-written code as part of typical workflow
  • Collaborate with research scientists and partner teams to define what’s worth measuring
  • Contribute to internal write-ups and public reporting

Skills

End-to-end evaluation
Large-scale data handling
LLM experimentation
Code review & collaboration

Job description

About Anthropic

Anthropic's mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems.

About the team

At Anthropic, we are delegating a growing share of AI development to AI systems themselves. Takeoff Intel is the team that measures this recursion from the inside. We're part of the Anthropic Institute. We design evaluations of AI R&D capabilities, build the internal telemetry Anthropic uses to track how much of its own model development is becoming AI‑assisted, and develop the quantitative methods that turn those signals into a calibrated picture of where capability growth is heading, so that Anthropic and the wider world have accurate situational awareness on this acceleration.

Our work appears in Anthropic's model system cards (we own the AI R&D capability assessments and adapted Epoch's Capabilities Index to our evals); all the data in When AI Builds Itself comes from our team. Internally, our measurements shape research priorities and safety planning; externally, they contribute to Anthropic's public reporting on the pace of AI progress and to collaborations with third‑party evaluators. We're a small team that works closely with pretraining, RL, economics, and policy researchers across the company. If you're passionate about measurement accuracy, and feel urgency about safety and situational awareness, you should consider joining us.

About the role

As a Research Engineer on Takeoff Intel you'll build and run the evaluation and measurement instruments that make this research possible. This is a generalist role on a small team: you'll work across evals infrastructure, large‑scale data processing, and analysis tooling, and you'll prioritize shipping. We build instruments that answer real questions and help set priorities, not dashboards that surface noise. We value working prototypes, rapid iteration, accuracy and good prioritization. We often need to go from a vague research question to a running instrument quickly.

We're hiring at both junior and senior levels.

Responsibilities
  • Design, build, and run capability evaluations and measurement instruments at scale

  • Build the data and analysis pipelines that turn large volumes of model outputs and telemetry into reliable metrics

  • Prototype new instruments fast, validate them, and decide what to keep

  • Review and supervise AI‑written code as a normal part of the workflow

  • Work closely with research scientists on the team and with partner teams to define what's worth measuring

  • Contribute to internal write‑ups and public reporting

You may be a good fit if you
  • Have shipped an evaluation, data product, or research library end to end

  • Prototype fast and are comfortable throwing code away

  • Handle messy, large‑volume data without over‑engineering

  • Have run experiments on large language models, not just moved their outputs around

  • Can work from a vague question rather than a spec

  • Communicate results clearly and collaborate closely with the researchers whose questions your instruments answer

Strong candidates may also have
  • Built evaluation harnesses or benchmark infrastructure for LLMs

  • Experience with large‑scale ML or data infrastructure (self‑driving, observability, or similar) alongside ML exposure

  • Built tools or libraries that other researchers rely on

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Research Engineer, Takeoff Intel
Research Engineer, Takeoff Intel

Anthropic • San Francisco (CA)

Hybrid
USD 350,000 - 850,000
Equity donation matching
Generous vacation and parental leave
Flexible working hours
+1
Research Engineer, Takeoff Intel Remote-Friendly (Travel Required) | San Francisco, CA
Research Engineer, Takeoff Intel Remote-Friendly (Travel Required) | San Francisco, CA

Anthropic Limited • San Francisco (CA), Northern (KY)

Hybrid
USD 350,000 - 850,000
Pre-training Distributed Systems Tech Lead / Manager
Pre-training Distributed Systems Tech Lead / Manager

Anthropic • San Francisco (CA)

Hybrid
USD 500,000 - 850,000
Equity donation matching
Generous vacation and parental leave
Flexible working hours
+1
Research Engineer – AI Evaluation & Metrics
Research Engineer – AI Evaluation & Metrics

EngineersOfAI • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 250,000
Research Engineer: AI Evaluation & Metrics
Research Engineer: AI Evaluation & Metrics

Anthropic • San Francisco (CA)

Hybrid
USD 350,000 - 850,000
Equity donation matching
Generous vacation and parental leave
Flexible working hours
+1
Research Engineer, Model Evaluations
Research Engineer, Model Evaluations

Anthropic • United States

Remote
USD 120,000 - 230,000
Software Engineer, Research Data Platform
Software Engineer, Research Data Platform

Anthropic • New York (NY)

On-site
USD 320,000 - 405,000
Equity donation matching
Generous vacation
Parental leave
+2
Research Engineer (Discovery)
Research Engineer (Discovery)

Anthropic • San Francisco (CA)

On-site
USD 315,000 - 560,000
Competitive compensation
Generous vacation and parental leave
Flexible working hours
Software Engineer, Research Tools
Software Engineer, Research Tools

Anthropic • New York (NY)

Hybrid
USD 300,000 - 405,000
Technical Program Manager, RL Research
Technical Program Manager, RL Research

Anthropic • San Francisco (CA)

Hybrid
USD 365,000 - 435,000