Research Engineer – Mechanistic Interpretability Systems

Acceler8 Talent

San Francisco, Northern (CA, KY)

Hybrid

USD 150,000 - 230,000

Full time

9 days ago
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Acceler8 Talent in San Francisco seeks a Research Engineer to build experimental systems for interpretability, alignment, and RL research, focusing on understanding model internals rather than production ML.

You will prototype tooling, run fast experiments, and contribute to new benchmarks for model robustness. PhD welcome but not required; strong software skills and curiosity about interpretability are essential.

Qualifications

  • Strong software engineering fundamentals.
  • Experience with experimental ML / research systems.
  • Comfort working close to model internals.
  • Interest in interpretability, alignment, RL, or mechanistic understanding.

Responsibilities

  • Build experimental tooling and systems to support interpretability research.
  • Create custom RL-style environments for alignment research and testing.
  • Probe internal representations and detect latent concepts such as deception, goals, uncertainty, or hidden objectives.
  • Develop benchmarks for model consistency and robustness and evaluate experimental ideas.
  • Operate in a fast, greenfield research context with rapid iteration.

Skills

Software engineering
Experimental ML
Model internals
Interpretability
PhD helpful

Job description

Acceler8 Talent in San Francisco seeks a Research Engineer to build experimental systems for interpretability, alignment, and RL research, focusing on understanding model internals rather than production ML.

You will prototype tooling, run fast experiments, and contribute to new benchmarks for model robustness. PhD welcome but not required; strong software skills and curiosity about interpretability are essential.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Mechanistic Interpretability
Mechanistic Interpretability

Acceler8 Talent • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 230,000
Research Engineer, Interpretability
Research Engineer, Interpretability

Anthropic • San Francisco (CA)

Hybrid
USD 315,000 - 560,000
Mechanistic AI Interpretability Scientist
Mechanistic AI Interpretability Scientist

Anthropic • San Francisco (CA)

Hybrid
USD 350,000 - 850,000
Research Engineer, Interpretability
Research Engineer, Interpretability

Anthropic • United States

Hybrid
USD 315,000 - 560,000
Remote work considered case-by-case
Visa sponsorship may be available
Office in San Francisco
Interpretable AI Research Scientist
Interpretable AI Research Scientist

World Mechanics • San Francisco (CA)

On-site
USD 250,000 - 400,000
Equity
Research Scientist, AI Interpretability - SF + Equity
Research Scientist, AI Interpretability - SF + Equity

Goodfire • San Francisco (CA), Northern (KY)

Hybrid
USD 200,000 - 400,000
Research Engineer - Scalable Interpretability
Research Engineer - Scalable Interpretability

Transluce • San Francisco (CA)

On-site
USD 250,000 - 500,000
Interpretability Research Engineer: Build Tools for Safe AI
Interpretability Research Engineer: Build Tools for Safe AI

Anthropic Limited • San Francisco (CA)

Hybrid
USD 315,000 - 560,000
Equity donation matching
Vacation and parental leave
Flexible working hours
+1
Mechanistic Interpretability Researcher
Mechanistic Interpretability Researcher

OpenAI • San Francisco (CA)

On-site
USD 180,000 - 240,000
Member of Technical Staff — ML Research, Interpretability
Member of Technical Staff — ML Research, Interpretability

Kindredventures • San Francisco (CA)

On-site
USD 120,000 - 190,000