Mechanistic Interpretability

Acceler8 Talent

San Francisco (CA)

On-site

USD 140,000 - 190,000

Full time

31 hours ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Acceler8 Talent is seeking a Research Engineer focused on Interpretability Systems in the San Francisco area. The role centers on understanding what happens inside large language models, not production ML, MLOps, or large-scale training infra.

You will build experimental systems, develop tooling, and run fast, greenfield research cycles to test ideas and push forward interpretability, alignment, and RL research. This is a hands-on, lab-focused role for curious engineers.

Qualifications

  • Strong software engineering fundamentals.
  • Experience with experimental ML / research systems.
  • Comfort working close to model internals.
  • Interest in interpretability, alignment, RL, or mechanistic understanding.
  • PhD helpful, not required.

Responsibilities

  • Research and develop interpretability tooling and experiments.
  • Conduct activation tracing and mechanistic analysis of language models.
  • Create benchmarks and tools to advance interpretability research.

Skills

Software engineering fundamentals
Experimental ML / research systems
Understanding model internals
Interest in interpretability / RL
PhD helpful

Job description

Research Engineer – Interpretability Systems

An AI research lab working at the frontier of interpretability, alignment, and reinforcement learning is hiring Research Engineers focused on understanding what's happening inside large language models

This role is for engineers who want to build the experimental systems that make interpretability research possible - not production ML, MLOps, or large-scale training infra

You'll work on:
  • Activation tracing & mechanistic analysis
  • Activation-level steering beyond prompting and fine-tuning
  • New benchmarks for model consistency and robustness

The work is fast, experimental, and greenfield: build custom tooling, test research ideas, get results, move on.

Ideal background:
  • Strong software engineering fundamentals
  • Experience with experimental ML / research systems
  • Comfort working close to model internals
  • Interest in interpretability, alignment, RL, or mechanistic understanding
  • PhD helpful, not required

This is not a role for scaling pipelines or maintaining production systems

It's for people who enjoy ambiguous problems, fast research cycles, and building new tools from first principles

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Research Engineer - Mechanistic Interpretability & AI Insight
Research Engineer - Mechanistic Interpretability & AI Insight

Acceler8 Talent • San Francisco (CA)

On-site
USD 140,000 - 190,000
Research Engineer, Interpretability
Research Engineer, Interpretability

Anthropic • San Francisco (CA)

Hybrid
USD 315,000 - 560,000
Mechanistic AI Interpretability Scientist
Mechanistic AI Interpretability Scientist

Anthropic • San Francisco (CA)

Hybrid
USD 350,000 - 850,000
Research Engineer - Scalable Interpretability
Research Engineer - Scalable Interpretability

Transluce • San Francisco (CA)

On-site
USD 250,000 - 500,000
Member of Technical Staff — ML Research, Interpretability
Member of Technical Staff — ML Research, Interpretability

Kindredventures • San Francisco (CA)

On-site
USD 120,000 - 190,000
Mechanistic Interpretability Researcher
Mechanistic Interpretability Researcher

OpenAI • San Francisco (CA)

On-site
USD 180,000 - 240,000
Interpretability Research Engineer: Build Tools for Safe AI
Interpretability Research Engineer: Build Tools for Safe AI

Anthropic Limited • San Francisco (CA)

Hybrid
USD 315,000 - 560,000
Equity donation matching
Vacation and parental leave
Flexible working hours
+1
Researcher, Alignment Interpretability
Researcher, Alignment Interpretability

OpenAI • San Francisco (CA)

On-site
USD 180,000 - 240,000
ML Researcher — Interpretability & Next-Gen Architectures
ML Researcher — Interpretability & Next-Gen Architectures

Tilde Research • San Francisco (CA)

On-site
USD 180,000 - 240,000
Software Engineer, Infrastructure, Interpretability
Software Engineer, Infrastructure, Interpretability

Anthropic • San Francisco (CA)

On-site
USD 190,000 - 270,000