Mechanistic AI Interpretability Scientist

Anthropic

San Francisco (CA)

Hybrid

USD 350,000 - 850,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Anthropic is seeking researchers and engineers for the Interpretability team to reverse-engineer how trained models work and develop mechanistic understanding. The role focuses on methods to understand LLMs, running robust experiments, and constructing interpretable circuits using features.

Based in San Francisco, exceptional candidates may be considered for remote work. This position offers visa sponsorship and collaboration with Alignment Science and Societal Impacts to enhance model safety.

Qualifications

  • Strong track record of scientific research and interpretability work.
  • Ability to collaborate across teams and communicate results clearly.
  • Comfortable with experimental research and exploring new directions.
  • Proficient in Python for research and tooling.

Responsibilities

  • Develop methods for understanding LLMs by reverse engineering algorithms learned in their weights
  • Design and run robust experiments, scalable to large models
  • Create and analyze interpretability features and circuits to understand model behavior
  • Build infrastructure for running experiments and visualizing results
  • Communicate results internally and publicly

Skills

Scientific research
Interpretability
Team science
Experimental mindset
Coding & experiments
Python

Education

Bachelor's degree or equivalent

Job description

Anthropic is seeking researchers and engineers for the Interpretability team to reverse-engineer how trained models work and develop mechanistic understanding. The role focuses on methods to understand LLMs, running robust experiments, and constructing interpretable circuits using features.

Based in San Francisco, exceptional candidates may be considered for remote work. This position offers visa sponsorship and collaboration with Alignment Science and Societal Impacts to enhance model safety.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Mechanistic Interpretability Research Scientist
Mechanistic Interpretability Research Scientist

Anthropic • California (MO)

Hybrid
USD 350,000 - 850,000
Mechanistic Interpretability
Mechanistic Interpretability

Acceler8 Talent • San Francisco (CA)

On-site
USD 140,000 - 190,000
Research Engineer - Mechanistic Interpretability & AI Insight
Research Engineer - Mechanistic Interpretability & AI Insight

Acceler8 Talent • San Francisco (CA)

On-site
USD 140,000 - 190,000
Research Engineer, Interpretability
Research Engineer, Interpretability

Anthropic • San Francisco (CA)

Hybrid
USD 315,000 - 560,000
Mechanistic Interpretability Researcher for AI Safety
Mechanistic Interpretability Researcher for AI Safety

OpenAI • Los Angeles (CA)

On-site
USD 120,000 - 150,000
Interpretability Research Manager — Lead High-Impact ML
Interpretability Research Manager — Lead High-Impact ML

Anthropic • California (MO)

Hybrid
USD 350,000 - 500,000
Equity donation matching
Generous vacation
Parental leave
+2
Mechanistic Interpretability Researcher
Mechanistic Interpretability Researcher

OpenAI • San Francisco (CA)

On-site
USD 180,000 - 240,000
Mechanistic Interpretability Researcher
Mechanistic Interpretability Researcher

OpenAI • California (MO)

On-site
USD 180,000 - 320,000
Infrastructure Engineer, Interpretability & AI Safety
Infrastructure Engineer, Interpretability & AI Safety

Anthropic • San Francisco (CA)

On-site
USD 190,000 - 270,000
Interpretability Research Engineer: Build Tools for Safe AI
Interpretability Research Engineer: Build Tools for Safe AI

Anthropic Limited • San Francisco (CA)

Hybrid
USD 315,000 - 560,000
Equity donation matching
Vacation and parental leave
Flexible working hours
+1