Interpretability Engineer for AI Safety & Research

Anthropic

San Francisco (CA)

Hybrid

USD 315,000 - 560,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Anthropic is looking for a skilled software engineer to join their Interpretability team in San Francisco. The role involves building and maintaining infrastructure for interpretability research while ensuring reliability and efficiency. Candidates should have over 5 years of software development experience and be proficient in at least one programming language. A curiosity about AI safety and the societal impacts of technology is essential. Salary ranges from $315,000 to $560,000, with some remote flexibility.

Qualifications

  • 5-10+ years of experience building software.
  • Highly proficient in at least one programming language.
  • Extremely curious about unfamiliar domains and able to learn quickly.
  • Strong ability to prioritize impactful work amidst ambiguity.
  • Interest in AI safety and interpretability research.

Responsibilities

  • Build and maintain specialized inference and training infrastructure.
  • Resolve scaling and efficiency bottlenecks through optimization.
  • Design tools and abstractions for rapid research experiments.
  • Assist in production safety audits with real deadlines.
  • Work across the stack on model internals and research tooling.

Skills

Proficient in at least one programming language (e.g., Python, Rust, Go, Java)
Experience in building software (5-10+ years)
Curiosity about unfamiliar domains
Ability to prioritize impactful work
Experience with interpretability research

Tools

PyTorch
CUDA
JAX
TPUs

Job description

Anthropic is looking for a skilled software engineer to join their Interpretability team in San Francisco. The role involves building and maintaining infrastructure for interpretability research while ensuring reliability and efficiency. Candidates should have over 5 years of software development experience and be proficient in at least one programming language. A curiosity about AI safety and the societal impacts of technology is essential. Salary ranges from $315,000 to $560,000, with some remote flexibility.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Interpretability Research Manager — AI Safety Leadership
Interpretability Research Manager — AI Safety Leadership

Menlo Ventures • San Francisco (CA)

Hybrid
USD 340,000 - 425,000
Competitive salary
Generous vacation and parental leave
Flexible working hours
+1
Research Manager, Interpretability — Lead Safe AI Teams
Research Manager, Interpretability — Lead Safe AI Teams

Anthropic • San Francisco (CA)

On-site
USD 340,000 - 425,000
Infrastructure Engineer, Interpretability & AI Safety
Infrastructure Engineer, Interpretability & AI Safety

Anthropic • San Francisco (CA)

On-site
USD 190,000 - 270,000
Interpretability Research Engineer: Build Tools for Safe AI
Interpretability Research Engineer: Build Tools for Safe AI

Anthropic Limited • San Francisco (CA)

Hybrid
USD 315,000 - 560,000
Equity donation matching
Vacation and parental leave
Flexible working hours
+1
Infrastructure Engineer, Interpretability & AI Safety
Infrastructure Engineer, Interpretability & AI Safety

Anthropic Limited • New York (NY)

Hybrid
USD 320,000 - 485,000
Research Engineer, Interpretability
Research Engineer, Interpretability

Anthropic • San Francisco (CA)

Hybrid
USD 315,000 - 560,000
Field Engineer, AI Interpretability & Deployment
Field Engineer, AI Interpretability & Deployment

Goodfire • San Francisco (CA)

On-site
USD 200,000 - 325,000
Market competitive salary
Equity
Competitive benefits
Mechanistic Interpretability Researcher for AI Safety
Mechanistic Interpretability Researcher for AI Safety

OpenAI • Los Angeles (CA)

On-site
USD 120,000 - 150,000
Mechanistic AI Interpretability Scientist
Mechanistic AI Interpretability Scientist

Anthropic • San Francisco (CA)

Hybrid
USD 350,000 - 850,000
Software Engineer, Infrastructure, Interpretability
Software Engineer, Infrastructure, Interpretability

Anthropic • San Francisco (CA)

On-site
USD 190,000 - 270,000