Research Engineer, Interpretability

Anthropic

United States

Hybrid

USD 315,000 - 560,000

Full time

5 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Remote work considered case-by-case
Visa sponsorship may be available
Office in San Francisco

Job summary

Anthropic is seeking an AI Engineer focused on interpretability and safety. The role involves building engineering systems that enable research into how large language models represent information and produce behavior.

You will work across model internals, distributed training and inference, accelerator performance, and researcher-facing tools to translate interpretability methods into dependable safety-audit workflows.

Qualifications

  • 5–10+ years of professional software engineering experience.
  • Proficiency in Python, Rust, Go or Java; ability to work productively in Python.
  • Ability to investigate unfamiliar areas and trace bottlenecks across a system.
  • Comfortable collaborating with researchers and engineers in a fast-moving environment.
  • Genuine interest in interpretability and AI safety; prior interpretability experience not required.

Responsibilities

  • Build and maintain specialized training and inference infrastructure for interpretability research.
  • Support instrumented forward and backward passes, activation extraction, and controlled application of steering vectors.
  • Profile systems, identify scaling constraints, and improve performance and efficiency across hardware and software layers.
  • Create abstractions and platforms that let researchers run experiments quickly without unnecessary engineering friction.
  • Help operationalize interpretability research in production safety audits with strong reliability expectations.
  • Collaborate with infrastructure and research partners while working across the stack from model internals to user-facing tooling.

Skills

Python
Rust
Go
Java

Job description

AI Engineer Remote work considered case-by-case; role based in San Francisco

Job details

Remote work considered case-by-case; role based in San Francisco Eligibility

Lead Experience

Not specified Employment

About this role

Role overview

This role focuses on building the engineering systems that enable research into how large language models represent information and produce behavior. You will work across model internals, distributed training and inference, accelerator performance, and researcher-facing tools, helping translate interpretability methods into dependable safety-audit workflows.

Responsibilities

  • Build and maintain specialized training and inference infrastructure for interpretability research.
  • Support instrumented forward and backward passes, activation extraction, and controlled application of steering vectors.
  • Profile systems, identify scaling constraints, and improve performance and efficiency across hardware and software layers.
  • Create abstractions and platforms that let researchers run experiments quickly without unnecessary engineering friction.
  • Help operationalize interpretability research in production safety audits with strong reliability expectations.
  • Collaborate with infrastructure and research partners while working across the stack from model internals to user-facing tooling.

Requirements

  • Approximately 5-10 or more years of professional software engineering experience, depending on level.
  • Strong proficiency in at least one programming language such as Python, Rust, Go, or Java, with the ability to work productively in Python.
  • Demonstrated ability to investigate unfamiliar technical areas and trace bottlenecks through multiple layers of a system.
  • Sound judgment about prioritization, impact, and trade-offs in an ambiguous, fast-moving environment.
  • Comfortable collaborating closely with researchers and engineers rather than working exclusively independently.
  • Genuine interest in interpretability and the role of technical understanding in AI safety; prior interpretability experience is not required.

Benefits and work setup

The role is based in San Francisco, with remote consideration possible for exceptional candidates. The general expectation is that staff spend at least 25% of their time in an office, though some roles may require more. Visa sponsorship may be available depending on the role and candidate. The listed annual salary range is $315,000-$560,000 USD.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Research Engineer, Interpretability
Research Engineer, Interpretability

Anthropic • San Francisco (CA)

Hybrid
USD 315,000 - 560,000
Interpretability Engineer for AI Safety & Research
Interpretability Engineer for AI Safety & Research

Anthropic • San Francisco (CA)

Hybrid
USD 315,000 - 560,000
Research Engineer - Scalable Interpretability
Research Engineer - Scalable Interpretability

Transluce • San Francisco (CA)

On-site
USD 250,000 - 500,000
Researcher, Alignment Interpretability
Researcher, Alignment Interpretability

OpenAI • San Francisco (CA)

On-site
USD 180,000 - 240,000
Researcher, Interpretability
Researcher, Interpretability

OpenAI • Los Angeles (CA)

On-site
USD 120,000 - 150,000
Mechanistic Interpretability
Mechanistic Interpretability

Acceler8 Talent • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 230,000
Interpretability Engineer, LLM Research & Infra
Interpretability Engineer, LLM Research & Infra

Anthropic • United States

Hybrid
USD 315,000 - 560,000
Remote work considered case-by-case
Visa sponsorship may be available
Office in San Francisco
Infrastructure Engineer, Interpretability & AI Safety
Infrastructure Engineer, Interpretability & AI Safety

Anthropic • San Francisco (CA)

On-site
USD 190,000 - 270,000
Research Scientist, AI Interpretability - SF + Equity
Research Scientist, AI Interpretability - SF + Equity

Goodfire • San Francisco (CA), Northern (KY)

Hybrid
USD 200,000 - 400,000
Research Engineer: Scalable AI Interpretability
Research Engineer: Scalable AI Interpretability

Transluce • San Francisco (CA)

On-site
USD 250,000 - 500,000