Research Engineer - Scalable Interpretability

Transluce

San Francisco (CA)

On-site

USD 250,000 - 500,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Transluce, a non-profit research lab in San Francisco, seeks scientists and engineers to advance AI oversight tools. You will help develop interpretable assistants and evaluate model behaviors, ensuring industry standards. Ideal candidates have experience in fine-tuning language models, excellent communication skills, and a curiosity for machine learning.

We offer a competitive salary range of $250,000 - $500,000 annually, in a collaborative, in-person work environment, with visa sponsorship available for international talents.

Qualifications

  • Experience with fine-tuning language models, designing new architectures, and creating evaluations.
  • Reliable results from good experimental design.
  • Strong programming skills and ability to navigate trade-offs between speed and maintainability.

Responsibilities

  • Help develop and train scalable interpretability assistants.
  • Create diverse evaluations that find undesirable model behaviors.
  • Scale up training and inference pipelines for large models.

Skills

Fine-tuning language models
Experimental design
Strong programming ability
Communication skills

Job description

Salary range: $250,000 - $500,000/year + benefits

Description: Transluce is a non‑profit research lab building tools for scalable, end‑to‑end oversight of AI systems. We build world‑class, AI‑backed analysis tools and use these to set industry standards for evaluation. Our tools are integrated with core agent benchmarks like SWE‑bench, while our evaluations are directly underpinning regulation, including our role as EU AI Office’s main evaluation developer for harmful manipulation risks.

About the role:

We are looking for strong scientists and engineers to help advance our vision of scalable end‑to‑end oversight assistants, building on our recent advances such as predictive concept decoders and user model extractors. As part of our highly collaborative team, you will learn and grow quickly, creating technology at the frontier of AI research and with high direct impact.

Core responsibility: Help us develop and train scalable interpretability assistants that can predict and detect unexpected and subtle behaviors from models’ activations. This includes:

  • Creating diverse evaluations that range in difficulty. This involves finding naturally occurring interesting and undesirable behaviors exhibited by open‑source models.
  • Developing novel architectures and objectives for training interpretability assistants.
  • Scaling up the training and inference pipelines to support up to 1T‑scale models.
Qualities of a strong candidate:
  • Experience with fine‑tuning language models, designing new architectures, and creating evaluations.
  • Reliable results: good experimental design, epistemic self‑awareness and transparency
  • Generativeness: coming up with original, productive ideas for unblocking progress
  • Curiosity: a desire to understand ML systems and how they work
  • Strong programming ability, including navigating trade‑offs between prototyping speed and maintainability
  • Strong communication skills, low ego, openness to giving and receiving feedback

We are located in San Francisco and enthusiastic to work together in‑person. We are open to sponsoring international visas.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Research Engineer: Scalable AI Interpretability
Research Engineer: Scalable AI Interpretability

Transluce • San Francisco (CA)

On-site
USD 250,000 - 500,000
Research Engineer, Interpretability
Research Engineer, Interpretability

Anthropic • San Francisco (CA)

Hybrid
USD 315,000 - 560,000
AI Behavior Engineer
AI Behavior Engineer

Transluce • San Francisco (CA)

On-site
USD 310,000 - 500,000
Software Engineer, Infrastructure, Interpretability
Software Engineer, Infrastructure, Interpretability

United States Digital Space LLC • San Francisco (CA), New York (NY)

On-site
USD 320,000 - 485,000
Research Engineer - Mechanistic Interpretability
Research Engineer - Mechanistic Interpretability

Acceler8 Talent • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 210,000
AI Behavior Researcher - Child Safety and Mental Health
AI Behavior Researcher - Child Safety and Mental Health

Transluce • San Francisco (CA)

On-site
USD 250,000 - 450,000
Member of Technical Staff — ML Research, Interpretability
Member of Technical Staff — ML Research, Interpretability

Kindredventures • San Francisco (CA)

On-site
USD 120,000 - 190,000
ML Researcher — Interpretability & Next-Gen Architectures
ML Researcher — Interpretability & Next-Gen Architectures

Tilde Research • San Francisco (CA)

On-site
USD 180,000 - 240,000
Research Engineer, Infrastructure, Inference
Research Engineer, Infrastructure, Inference

Thinking Machines Lab Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 350,000 - 475,000
Health, dental, vision benefits
Unlimited PTO
Parental leave
+1
Research Engineer, Infrastructure, Inference
Research Engineer, Infrastructure, Inference

Thinkingmachines • San Francisco (CA)

On-site
USD 350,000 - 475,000
Generous health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1