AI Behavior Researcher — Impactful AI Evaluation Lead

Transluce

San Francisco (CA)

On-site

USD 250,000 - 450,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Transluce, a fast-moving nonprofit research lab in San Francisco, seeks an AI Behavior Researcher to lead automated evaluations of frontier AI systems, with a focus on wellbeing and user autonomy. You will design pipelines, implement user simulators, and develop robust evaluation rubrics while collaborating with regulators and domain experts.

We value rigorous experimental design and strong Python skills, and we invite applicants across experience levels.

Qualifications

  • Experience designing and validating automated AI evaluation methods, such as LLM-as-a-judge systems or multi-turn benchmarks.
  • Expertise in quantitative generative AI evaluation and measurement; ability to operationalize social concepts.
  • Strong Python proficiency for analysis and tooling.
  • Meticulous experimental design with transparency.
  • Ability to communicate effectively with researchers and decision makers.

Responsibilities

  • Develop novel automated evaluations of AI’s impacts on users, including mental health and decision making.
  • Write code to implement and run automated evaluations, such as user simulators or LLM-as-a-judge pipelines.
  • Design methods to improve ecological validity of evaluations for specific populations.
  • Write and revise rubrics to evaluate model behaviors related to wellbeing and decision making.
  • Collaborate with scientists and research engineers to productionize best practices in AI behavioral evaluation.

Skills

Quantitative AI evaluation
Python proficiency
Experimental design
Communication skills

Tools

LLM-as-a-judge pipelines

Job description

Transluce, a fast-moving nonprofit research lab in San Francisco, seeks an AI Behavior Researcher to lead automated evaluations of frontier AI systems, with a focus on wellbeing and user autonomy. You will design pipelines, implement user simulators, and develop robust evaluation rubrics while collaborating with regulators and domain experts.

We value rigorous experimental design and strong Python skills, and we invite applicants across experience levels.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Frontier AI Behavior Research Scientist
Frontier AI Behavior Research Scientist

Transluce • San Francisco (CA), Northern (KY)

Hybrid
USD 250,000 - 450,000
AI Behavior Researcher - Human Impacts
AI Behavior Researcher - Human Impacts

Transluce • San Francisco (CA)

On-site
USD 250,000 - 450,000
AI Behavior Engineer
AI Behavior Engineer

Transluce • San Francisco (CA)

On-site
USD 310,000 - 500,000
AI Behavior Researcher: Safeguarding Kids & Mental Health
AI Behavior Researcher: Safeguarding Kids & Mental Health

Transluce • San Francisco (CA)

On-site
USD 250,000 - 450,000
Forward‑Deployed AI Behavior Engineer
Forward‑Deployed AI Behavior Engineer

Transluce • San Francisco (CA)

On-site
USD 310,000 - 500,000
AI Behavior Researcher - Child Safety and Mental Health
AI Behavior Researcher - Child Safety and Mental Health

Transluce • San Francisco (CA)

On-site
USD 250,000 - 450,000
AI Behavior Researcher - Agent Alignment
AI Behavior Researcher - Agent Alignment

Transluce • San Francisco (CA), Northern (KY)

Hybrid
USD 250,000 - 450,000
Governance & Policy Fellow
Governance & Policy Fellow

Transluce • San Francisco (CA)

On-site
USD 200,000 - 300,000
Visa sponsorship
Research Scientist, AI Systems & RL for Public Good
Research Scientist, AI Systems & RL for Public Good

Transluce • San Francisco (CA)

On-site
USD 250,000 - 500,000
Research Scientist/Research Engineer
Research Scientist/Research Engineer

Transluce • San Francisco (CA)

On-site
USD 250,000 - 500,000