Research Engineer, Interpretability

Anthropic

San Francisco (CA)

Hybrid

USD 315,000 - 560,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Anthropic is looking for a skilled software engineer to join their Interpretability team in San Francisco. The role involves building and maintaining infrastructure for interpretability research while ensuring reliability and efficiency. Candidates should have over 5 years of software development experience and be proficient in at least one programming language. A curiosity about AI safety and the societal impacts of technology is essential. Salary ranges from $315,000 to $560,000, with some remote flexibility.

Qualifications

  • 5-10+ years of experience building software.
  • Highly proficient in at least one programming language.
  • Extremely curious about unfamiliar domains and able to learn quickly.
  • Strong ability to prioritize impactful work amidst ambiguity.
  • Interest in AI safety and interpretability research.

Responsibilities

  • Build and maintain specialized inference and training infrastructure.
  • Resolve scaling and efficiency bottlenecks through optimization.
  • Design tools and abstractions for rapid research experiments.
  • Assist in production safety audits with real deadlines.
  • Work across the stack on model internals and research tooling.

Skills

Proficient in at least one programming language (e.g., Python, Rust, Go, Java)
Experience in building software (5-10+ years)
Curiosity about unfamiliar domains
Ability to prioritize impactful work
Experience with interpretability research

Tools

PyTorch
CUDA
JAX
TPUs

Job description

About the Role

When you see what modern language models are capable of, you may wonder, "How do these things work? How can we trust them?" The Interpretability team at Anthropic is working to reverse‑engineer how trained models work because we believe that a mechanistic understanding is the most robust way to make advanced systems safe. Think of us as doing "neuroscience" of neural networks using "microscopes" we build—or reverse‑engineering neural networks like binary programs.

Responsibilities
  • Build and maintain the specialized inference and training infrastructure that powers interpretability research—including instrumented forward/backward passes, activation extraction, and steering vector application.
  • Resolve scaling and efficiency bottlenecks through profiling, optimization, and close collaboration with peer infrastructure teams.
  • Design tools, abstractions, and platforms that enable researchers to rapidly experiment without hitting engineering barriers.
  • Help bring interpretability research into production safety audits with real deadlines and high reliability expectations.
  • Work across the stack—from model internals and accelerator‑level optimization to user‑facing research tooling.
Candidate Fit
  • Have 5‑10+ years of experience building software.
  • Are highly proficient in at least one programming language (e.g., Python, Rust, Go, Java) and productive with Python.
  • Are extremely curious about unfamiliar domains; can quickly learn and put that knowledge to work, e.g., diving into new layers of the stack to find bottlenecks.
  • Have a strong ability to prioritize the most impactful work and are comfortable operating with ambiguity and questioning assumptions.
  • Prefer fast‑moving collaborative projects to extensive solo efforts.
  • Are curious about interpretability research and its role in AI safety (though no research experience is required!).
  • Care about the societal impacts and ethics of your work.
  • Are comfortable working closely with researchers, translating research needs into engineering solutions.
Strong Candidates May Also Have Experience With
  • Optimizing the performance of large‑scale distributed systems.
  • Language modeling fundamentals with transformers.
  • High Performance LLM optimization: memory management, compute efficiency, parallelism strategies, inference throughput optimization.
  • Working hands‑on in a mainstream ML stack—PyTorch/CUDA on GPUs or JAX/XLA on TPUs.
  • Collaborating closely with researchers and building tooling to support research teams—or directly performed research with complex engineering challenges.
Representative Projects
  • Building Garcon, a tool that allows researchers to easily instrument LLMs to extract internal activations.
  • Designing and optimizing a pipeline to efficiently collect petabytes of transformer activations and shuffle them.
  • Profiling and optimizing ML training jobs, including multi‑GPU parallelism and memory optimization.
  • Building a steered inference system that applies targeted interventions to model internals at scale (conceptually similar to Golden Gate Claude but for safety research).
Annual Salary

$315,000—$560,000 USD

Location

This role is based in the San Francisco office; however, we are open to considering exceptional candidates for remote work on a case‑by‑case basis.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Engineer, Infrastructure, Interpretability
Software Engineer, Infrastructure, Interpretability

United States Digital Space LLC • San Francisco (CA), New York (NY)

On-site
USD 320,000 - 485,000
Software Engineer, Infrastructure, Interpretability
Software Engineer, Infrastructure, Interpretability

Anthropic • San Francisco (CA)

On-site
USD 190,000 - 270,000
Research Engineer - Scalable Interpretability
Research Engineer - Scalable Interpretability

Transluce • San Francisco (CA)

On-site
USD 250,000 - 500,000
Research Engineer - Mechanistic Interpretability
Research Engineer - Mechanistic Interpretability

Acceler8 Talent • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 210,000
Interpretability Research Engineer: Build Tools for Safe AI
Interpretability Research Engineer: Build Tools for Safe AI

Anthropic Limited • San Francisco (CA)

Hybrid
USD 315,000 - 560,000
Equity donation matching
Vacation and parental leave
Flexible working hours
+1
Member of Technical Staff — ML Research, Interpretability
Member of Technical Staff — ML Research, Interpretability

Kindredventures • San Francisco (CA)

On-site
USD 120,000 - 190,000
Research Engineer, Interpretability
Research Engineer, Interpretability

Anthropic Limited • San Francisco (CA)

Hybrid
USD 315,000 - 560,000
Equity donation matching
Vacation and parental leave
Flexible working hours
+1
Researcher, Interpretability
Researcher, Interpretability

OpenAI • Los Angeles (CA)

On-site
USD 120,000 - 150,000
Research Scientist, Interpretability
Research Scientist, Interpretability

Anthropic • San Francisco (CA)

Hybrid
USD 350,000 - 850,000
Mechanistic AI Interpretability Scientist
Mechanistic AI Interpretability Scientist

Anthropic • San Francisco (CA)

Hybrid
USD 350,000 - 850,000