Interpretability Research Engineer: Build Tools for Safe AI

Anthropic Limited

San Francisco (CA)

Hybrid

USD 315,000 - 560,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity donation matching
Vacation and parental leave
Flexible working hours
Office space in SF

Job summary

Anthropic is seeking an experienced software engineer to advance the Interpretability team’s infrastructure for training and inference of safety-focused AI models. You will work with researchers and infra teams to optimize performance, instrument activations, and empower rapid experimentation.

Strong candidates will have 5–10+ years building software, expertise in Python and ML stacks (PyTorch/CUDA or JAX), and a passion for interpretability research and its societal impact.

Qualifications

  • 5-10+ years of software development experience.
  • Proficiency in at least one programming language (Python preferred).
  • Interest in interpretability and AI safety.

Responsibilities

  • Build and maintain inference and training infrastructure for interpretability research.
  • Resolve scaling bottlenecks via profiling and optimization.
  • Design tools and platforms enabling rapid researcher experimentation.
  • Support production safety audits with reliable tooling.
  • Collaborate with researchers and infrastructure teams across the stack.

Skills

Python
Software engineering
5-10 years
Research collaboration
Ambiguity handling
Interpretable AI

Education

Bachelor's degree

Tools

PyTorch
CUDA
JAX

Job description

Anthropic is seeking an experienced software engineer to advance the Interpretability team’s infrastructure for training and inference of safety-focused AI models. You will work with researchers and infra teams to optimize performance, instrument activations, and empower rapid experimentation.

Strong candidates will have 5–10+ years building software, expertise in Python and ML stacks (PyTorch/CUDA or JAX), and a passion for interpretability research and its societal impact.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Infrastructure Engineer, Interpretability & AI Safety
Infrastructure Engineer, Interpretability & AI Safety

Anthropic Limited • New York (NY)

Hybrid
USD 320,000 - 485,000
Interpretability Engineer for AI Safety & Research
Interpretability Engineer for AI Safety & Research

Anthropic • San Francisco (CA)

Hybrid
USD 315,000 - 560,000
Infrastructure Engineer, Interpretability & AI Safety
Infrastructure Engineer, Interpretability & AI Safety

Anthropic • San Francisco (CA)

On-site
USD 190,000 - 270,000
Research Engineer, Interpretability
Research Engineer, Interpretability

Anthropic • San Francisco (CA)

Hybrid
USD 315,000 - 560,000
Software Engineer, Infrastructure, Interpretability
Software Engineer, Infrastructure, Interpretability

Anthropic • San Francisco (CA)

On-site
USD 190,000 - 270,000
Researcher, Interpretability
Researcher, Interpretability

OpenAI • Los Angeles (CA)

On-site
USD 120,000 - 150,000
Mechanistic AI Interpretability Scientist
Mechanistic AI Interpretability Scientist

Anthropic • San Francisco (CA)

Hybrid
USD 350,000 - 850,000
Research Engineer - Mechanistic Interpretability
Research Engineer - Mechanistic Interpretability

Acceler8 Talent • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 210,000
[Expression of Interest] Research Manager, Interpretability
[Expression of Interest] Research Manager, Interpretability

Anthropic • San Francisco (CA)

On-site
USD 340,000 - 425,000
Competitive compensation
Generous vacation leave
Flexible working hours
Research Tools Engineer: Build AI Experiment Platforms
Research Tools Engineer: Build AI Experiment Platforms

Anthropic • New York (NY)

Hybrid
USD 300,000 - 405,000