AI Behavior Researcher - Agent Alignment

Transluce

San Francisco, Northern (CA, KY)

Hybrid

USD 250,000 - 450,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Transluce, a fast-moving nonprofit research lab in San Francisco, seeks an AI Behavior Researcher to lead automated evaluations of frontier AI systems, focusing on honesty and alignment. You will design novel evaluations, write analysis code, and collaborate with governance to drive impactful public-policy work.

You will build environment simulators and LLM-based judge pipelines, contributing to scalable, rigorous research while growing with a highly collaborative team and, when needed, visa

Qualifications

  • Expertise in quantitative AI evaluation and measurement.
  • Experience designing automated AI evaluation methods, such as LLM-as-a-judge systems or multi-turn benchmarks.
  • Proficiency in Python to implement analysis and evaluation tooling.
  • Meticulous, good experimental design, epistemic self-awareness and transparency.
  • Ability to iterate quickly and balance between scrappiness and thoroughness based on impact needs.

Responsibilities

  • Develop novel, valid automated evaluations of AI agents' honesty and alignment.
  • Write code to implement and run automated evaluations, such as environment simulators or LLM-as-a-judge pipelines.
  • Design methods to improve the ecological of automated evaluations, and measure effects as models become more capable.
  • Collaborate with governance to deliver high-impact evaluations for public policy.
  • Collaborate with scientists and research engineers to productionize best practices in AI behavior evaluation.

Skills

AI evaluation
Python
Experimental design
Communication
Epistemic transparency

Tools

LLM-as-a-judge pipelines
Environment simulators

Job description

Salary range:

$250,000 - $450,000/year + benefits

Description:

Transluce is a fast-moving nonprofit research lab building the public tech stack for AI evaluation and oversight. We have contributed foundational research to the study of AI agents and their behaviors, and are using these to study emerging issues in the honesty and alignment of AI agents.

About the role:

As an AI Behavior Researcher, you will lead projects to design and develop automated evaluations of frontier AI systems that are technically sophisticated, scientifically valid, and concretely impactful. You will conduct novel analyses of behaviors related to agentic honesty and alignment. Example behaviors of interest include misreporting results, falsely claiming success, evaluation awareness, and memetic effects within AI swarms.

As an early member of a highly collaborative team, you will learn and grow quickly, and work with our governance and infrastructure teams to scale your impact and technical reach.

Core responsibilities:
  • Develop novel, valid automated evaluations of AI agents' honesty and alignment.
  • Write code to implement and run automated evaluations, such as environment simulators or LLM-as-a-judge pipelines.
  • Design methods to improve the ecological of automated evaluations, and especially to measure which effects are increasing or decreasing as models become more capable.
  • Collaborate with our governance team to deliver high-impact evaluations for public policy.
  • Collaborate with scientists and research engineers to productionize best practices in AI behavior evaluation.
Minimum qualifications:
  • Expertise on quantitative generative AI evaluation and measurement. Good intuition about how to systematize and operationalize complex concepts and to work backwards from possible failure modes of agents.
  • Relevant experience designing and validating automated AI evaluation methods, such as LLM-as-a-judge systems or multi-turn benchmarks.
  • Proficiency in Python to implement analysis and evaluation tooling.
  • Meticulous, good experimental design, epistemic self-awareness and transparency.
  • Ability to iterate quickly and balance between scrappiness and thoroughness based on the impact needs of a project.
  • Strong communication skills, low ego, openness to giving and receiving feedback.
Preferred qualifications (not required):
  • Experience running automated evaluations at scale or in a production context.
  • Experience conducting controlled human subjects experiments to validate automated evaluation methods.
  • Experience in customer-facing, consulting, or forward-deployed roles translating ambiguous stakeholder needs into concrete deliverables.
  • Experience and comfort using AI coding agents at work.

We are located in San Francisco and excited to work together in-person. We are open to sponsoring international visas.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Behavior Researcher - Agent Alignment
AI Behavior Researcher - Agent Alignment

Socket.dev • San Francisco (CA)

On-site
USD 250,000 - 450,000
AI Behavior Engineer
AI Behavior Engineer

Transluce • San Francisco (CA)

On-site
USD 310,000 - 500,000
AI Behavior Researcher - Child Safety and Mental Health
AI Behavior Researcher - Child Safety and Mental Health

Transluce • San Francisco (CA)

On-site
USD 250,000 - 450,000
Frontier AI Behavior Research Scientist
Frontier AI Behavior Research Scientist

Transluce • San Francisco (CA), Northern (KY)

Hybrid
USD 250,000 - 450,000
Research Engineer / Scientist, Alignment
Research Engineer / Scientist, Alignment

Anthropic • San Francisco (CA)

Hybrid
USD 350,000 - 500,000
Competitive compensation
Generous vacation and parental leave
Flexible working hours
Agent Post-Training, Frontier Evals and Environments Research
Agent Post-Training, Frontier Evals and Environments Research

OpenAI • San Francisco (CA)

On-site
USD 380,000 - 500,000
Research Engineer — AI Alignment & Evaluation
Research Engineer — AI Alignment & Evaluation

W3 Sourcing • San Francisco (CA)

Hybrid
USD 140,000 - 210,000
Senior AI Engineer - Agent Team
Senior AI Engineer - Agent Team

FurtherAI Inc • San Francisco (CA)

On-site
USD 120,000 - 150,000
Fully covered health, dental, and vision benefits
Competitive Compensation and stock options
Unlimited PTO
+4
Researcher, Agent Safety, Training and Evaluations
Researcher, Agent Safety, Training and Evaluations

AI Chopping Block • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Hybrid work model
Relocation assistance
Research Engineer - Evals
Research Engineer - Evals

AGI, Inc. • San Francisco (CA)

On-site
USD 120,000 - 160,000
Competitive cash and equity
Top-tier relocation support
In-person work environment