Mechanistic Interpretability

Acceler8 Talent

San Francisco, Northern (CA, KY)

Hybrid

USD 150,000 - 230,000

Full time

9 days ago
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Acceler8 Talent in San Francisco seeks a Research Engineer to build experimental systems for interpretability, alignment, and RL research, focusing on understanding model internals rather than production ML.

You will prototype tooling, run fast experiments, and contribute to new benchmarks for model robustness. PhD welcome but not required; strong software skills and curiosity about interpretability are essential.

Qualifications

  • Strong software engineering fundamentals.
  • Experience with experimental ML / research systems.
  • Comfort working close to model internals.
  • Interest in interpretability, alignment, RL, or mechanistic understanding.

Responsibilities

  • Build experimental tooling and systems to support interpretability research.
  • Create custom RL-style environments for alignment research and testing.
  • Probe internal representations and detect latent concepts such as deception, goals, uncertainty, or hidden objectives.
  • Develop benchmarks for model consistency and robustness and evaluate experimental ideas.
  • Operate in a fast, greenfield research context with rapid iteration.

Skills

Software engineering
Experimental ML
Model internals
Interpretability
PhD helpful

Job description

Research Engineer – Interpretability Systems
San Francisco, CA | Onsite
Early-stage AI research lab | Revenue-generating

An AI research lab working at the frontier of interpretability, alignment, and reinforcement learning is hiring Research Engineers focused on understanding what’s happening inside large language models

This role is for engineers who want to build the experimental systems that make interpretability research possible - not production ML, MLOps, or large-scale training infra

You’ll work on:
  • Activation tracing & mechanistic analysis
  • Custom RL-style environments for alignment research
  • Probing internal representations
  • Detecting latent concepts like deception, goals, uncertainty, or hidden objectives
  • Activation-level steering beyond prompting and fine-tuning
  • New benchmarks for model consistency and robustness

The work is fast, experimental, and greenfield: build custom tooling, test research ideas, get results, move on.

Ideal background:
  • Strong software engineering fundamentals
  • Experience with experimental ML / research systems
  • Comfort working close to model internals
  • Interest in interpretability, alignment, RL, or mechanistic understanding
  • PhD helpful, not required

This is not a role for scaling pipelines or maintaining production systems

It’s for people who enjoy ambiguous problems, fast research cycles, and building new tools from first principles

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Research Engineer – Mechanistic Interpretability Systems
Research Engineer – Mechanistic Interpretability Systems

Acceler8 Talent • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 230,000
Research Engineer, Interpretability
Research Engineer, Interpretability

Anthropic • San Francisco (CA)

Hybrid
USD 315,000 - 560,000
Mechanistic AI Interpretability Scientist
Mechanistic AI Interpretability Scientist

Anthropic • San Francisco (CA)

Hybrid
USD 350,000 - 850,000
Research Engineer, Interpretability
Research Engineer, Interpretability

Anthropic • United States

Hybrid
USD 315,000 - 560,000
Remote work considered case-by-case
Visa sponsorship may be available
Office in San Francisco
ML Researcher — Interpretability & Next-Gen Architectures
ML Researcher — Interpretability & Next-Gen Architectures

Tilde Research • San Francisco (CA)

On-site
USD 180,000 - 240,000
Researcher, Alignment Interpretability
Researcher, Alignment Interpretability

OpenAI • San Francisco (CA)

On-site
USD 180,000 - 240,000
Researcher, Interpretability
Researcher, Interpretability

OpenAI • Los Angeles (CA)

On-site
USD 120,000 - 150,000
Interpretability Engineer, LLM Research & Infra
Interpretability Engineer, LLM Research & Infra

Anthropic • United States

Hybrid
USD 315,000 - 560,000
Remote work considered case-by-case
Visa sponsorship may be available
Office in San Francisco
Researcher, Alignment Interpretability
Researcher, Alignment Interpretability

Triwill Group • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 260,000
ML Researcher
ML Researcher

Tilde Research • San Francisco (CA)

On-site
USD 180,000 - 240,000