Interpretability Engineer, LLM Research & Infra

Anthropic

United States

Hybrid

USD 315,000 - 560,000

Full time

5 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Remote work considered case-by-case
Visa sponsorship may be available
Office in San Francisco

Job summary

Anthropic is seeking an AI Engineer focused on interpretability and safety. The role involves building engineering systems that enable research into how large language models represent information and produce behavior.

You will work across model internals, distributed training and inference, accelerator performance, and researcher-facing tools to translate interpretability methods into dependable safety-audit workflows.

Qualifications

  • 5–10+ years of professional software engineering experience.
  • Proficiency in Python, Rust, Go or Java; ability to work productively in Python.
  • Ability to investigate unfamiliar areas and trace bottlenecks across a system.
  • Comfortable collaborating with researchers and engineers in a fast-moving environment.
  • Genuine interest in interpretability and AI safety; prior interpretability experience not required.

Responsibilities

  • Build and maintain specialized training and inference infrastructure for interpretability research.
  • Support instrumented forward and backward passes, activation extraction, and controlled application of steering vectors.
  • Profile systems, identify scaling constraints, and improve performance and efficiency across hardware and software layers.
  • Create abstractions and platforms that let researchers run experiments quickly without unnecessary engineering friction.
  • Help operationalize interpretability research in production safety audits with strong reliability expectations.
  • Collaborate with infrastructure and research partners while working across the stack from model internals to user-facing tooling.

Skills

Python
Rust
Go
Java

Job description

Anthropic is seeking an AI Engineer focused on interpretability and safety. The role involves building engineering systems that enable research into how large language models represent information and produce behavior.

You will work across model internals, distributed training and inference, accelerator performance, and researcher-facing tools to translate interpretability methods into dependable safety-audit workflows.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Interpretability Research Engineer: Build Tools for Safe AI
Interpretability Research Engineer: Build Tools for Safe AI

Anthropic Limited • San Francisco (CA)

Hybrid
USD 315,000 - 560,000
Equity donation matching
Vacation and parental leave
Flexible working hours
+1
Mechanistic AI Interpretability Scientist
Mechanistic AI Interpretability Scientist

Anthropic • San Francisco (CA)

Hybrid
USD 350,000 - 850,000
Infrastructure Engineer, Interpretability & AI Safety
Infrastructure Engineer, Interpretability & AI Safety

Anthropic • San Francisco (CA)

On-site
USD 190,000 - 270,000
Interpretability Engineer for AI Safety & Research
Interpretability Engineer for AI Safety & Research

Anthropic • San Francisco (CA)

Hybrid
USD 315,000 - 560,000
Research Engineer, Interpretability
Research Engineer, Interpretability

Anthropic • San Francisco (CA)

Hybrid
USD 315,000 - 560,000
Research Engineer, Interpretability
Research Engineer, Interpretability

Anthropic • United States

Hybrid
USD 315,000 - 560,000
Remote work considered case-by-case
Visa sponsorship may be available
Office in San Francisco
Software Engineer, Infrastructure, Interpretability
Software Engineer, Infrastructure, Interpretability

Anthropic • San Francisco (CA)

On-site
USD 190,000 - 270,000
Mechanistic Interpretability
Mechanistic Interpretability

Acceler8 Talent • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 230,000
Researcher, Interpretability
Researcher, Interpretability

OpenAI • Los Angeles (CA)

On-site
USD 120,000 - 150,000
Researcher, Alignment Interpretability
Researcher, Alignment Interpretability

OpenAI • San Francisco (CA)

On-site
USD 180,000 - 240,000