Researcher—Alignment & Mechanistic Interpretability

Triwill Group

San Francisco, Northern (CA, KY)

Hybrid

USD 180,000 - 260,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

OpenAI seeks a researcher passionate about understanding deep networks, with a strong background in engineering, quantitative reasoning, and the research process. You will develop and carry out a research plan in mechanistic interpretability, working closely with a motivated team.

You will help ensure future models remain safe as they grow in capability and have a meaningful impact on AI safety. This role emphasizes building scalable experiments, publishing results, and collaborating across

Qualifications

  • Ph.D. or research experience in computer science, machine learning, or a related field.
  • 2+ years of research engineering experience.
  • Proficiency in Python or similar languages.

Responsibilities

  • Develop and publish research on techniques for understanding representations of deep networks.
  • Engineer infrastructure for studying model internals at scale.
  • Collaborate across teams to pursue OpenAI-aligned projects.
  • Guide research directions toward demonstrable usefulness and scalability.

Skills

Python
Research engineering
Quantitative reasoning
Collaborative culture

Education

Ph.D. or equivalent research experience

Tools

Python

Job description

OpenAI seeks a researcher passionate about understanding deep networks, with a strong background in engineering, quantitative reasoning, and the research process. You will develop and carry out a research plan in mechanistic interpretability, working closely with a motivated team.

You will help ensure future models remain safe as they grow in capability and have a meaningful impact on AI safety. This role emphasizes building scalable experiments, publishing results, and collaborating across

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Mechanistic Interpretability Researcher
Mechanistic Interpretability Researcher

OpenAI • San Francisco (CA)

On-site
USD 180,000 - 240,000
Researcher, Alignment Interpretability
Researcher, Alignment Interpretability

OpenAI • San Francisco (CA)

On-site
USD 180,000 - 240,000
Researcher, Alignment Interpretability
Researcher, Alignment Interpretability

Triwill Group • San Francisco (CA), Northern (KY)

On-site
USD 180,000 - 260,000
Researcher, Interpretability
Researcher, Interpretability

OpenAI • Los Angeles (CA)

On-site
USD 120,000 - 150,000
Mechanistic Interpretability Researcher for AI Safety
Mechanistic Interpretability Researcher for AI Safety

OpenAI • Los Angeles (CA)

On-site
USD 120,000 - 150,000
Mechanistic AI Interpretability Scientist
Mechanistic AI Interpretability Scientist

Anthropic • San Francisco (CA)

Hybrid
USD 350,000 - 850,000
Researcher, Chain-of-Thought Monitorability & Alignment
Researcher, Chain-of-Thought Monitorability & Alignment

OpenAI • San Francisco (CA)

Hybrid
USD 120,000 - 150,000
Relocation assistance
Hybrid work model
Mechanistic Interpretability
Mechanistic Interpretability

Acceler8 Talent • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 230,000
AI Safety & Alignment Research Scientist (GenAI)
AI Safety & Alignment Research Scientist (GenAI)

DeepMind Technologies Limited • Mountain View (CA)

Hybrid
USD 207,000 - 300,000
Frontier AI Safety Researcher: Mitigations & Alignment
Frontier AI Safety Researcher: Mitigations & Alignment

OpenAI • San Francisco (CA)

On-site
USD 180,000 - 280,000