Interpretability Engineer for AI Safety & Research
Anthropic
San Francisco (CA)
Hybrid
USD 315,000 - 560,000
Full time
14 days+
Get more replies from employers
Send a job-specific resume in minutes.
Start fresh or import an existing resume
Job summary
Anthropic is looking for a skilled software engineer to join their Interpretability team in San Francisco. The role involves building and maintaining infrastructure for interpretability research while ensuring reliability and efficiency. Candidates should have over 5 years of software development experience and be proficient in at least one programming language. A curiosity about AI safety and the societal impacts of technology is essential. Salary ranges from $315,000 to $560,000, with some remote flexibility.
Qualifications
5-10+ years of experience building software.
Highly proficient in at least one programming language.
Extremely curious about unfamiliar domains and able to learn quickly.
Strong ability to prioritize impactful work amidst ambiguity.
Interest in AI safety and interpretability research.
Responsibilities
Build and maintain specialized inference and training infrastructure.
Resolve scaling and efficiency bottlenecks through optimization.
Design tools and abstractions for rapid research experiments.
Assist in production safety audits with real deadlines.
Work across the stack on model internals and research tooling.
Skills
Proficient in at least one programming language (e.g., Python, Rust, Go, Java)
Experience in building software (5-10+ years)
Curiosity about unfamiliar domains
Ability to prioritize impactful work
Experience with interpretability research
Tools
PyTorch
CUDA
JAX
TPUs
Job description
Anthropic is looking for a skilled software engineer to join their Interpretability team in San Francisco. The role involves building and maintaining infrastructure for interpretability research while ensuring reliability and efficiency. Candidates should have over 5 years of software development experience and be proficient in at least one programming language. A curiosity about AI safety and the societal impacts of technology is essential. Salary ranges from $315,000 to $560,000, with some remote flexibility.