Head of Evals, AI Red Teaming

Trajectory Labs, PBC

Berkeley, Northern (CA, KY)

Hybrid

USD 200,000 - 400,000

Full time

10 days ago
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Health coverage stipend
401(k)
Visa sponsorship
Equity

Job summary

Trajectory Labs, PBC is seeking a senior individual-contributor to own the evaluation pipeline for our prompt injection red-teaming line. You will review evals, design new tasks, and build agent tooling that scales your judgment, while working closely with founders and frontier labs.

As an early-stage startup, expect evolving priorities and significant autonomy to shape the role. Location is Berkeley, CA, with openness to remote for the right candidate and a compensation range of

Qualifications

  • Calibrated eval judgment to identify subtle errors and support evidence-based decisions.
  • Agent-native engineering using LLMs and coding agents to automate workflows.
  • Strong attention to detail across repeated reviews and automation opportunities.
  • Thrives in fast-moving, uncertain startup environments.
  • Proven experience evaluating or owning an eval or grading pipeline for technical products.
  • Experience building LLM judges and rubric-based grading, adapting as models change.
  • Designing eval environments and tooling for evaluation tasks.
  • Expertise in prompt injection, red teaming, or security in eval design.
  • Familiarity with AI safety and the alignment research community.
  • Experience collaborating with frontier labs or demanding technical customers.

Responsibilities

  • Hold the quality bar by reviewing tasks, transcripts, and submissions and deciding what ships to customers.
  • Design the next tests by analyzing model weaknesses and creating targeted tasks and environments.
  • Automate judgment by turning review patterns into agent skills and pipelines.
  • Improve internal workflows by applying eval principles to tooling and processes.

Skills

Calibrated eval judgment
Agent-native engineering
Sustained attention to detail
Early-stage startup drive
Rigorous evaluation experience
LLM judges & rubric-based grading
Evaluation environments
Prompt injection & red teaming
AI safety familiarity
Collaboration with frontier labs

Job description

You'll own the evaluation pipeline for our prompt injection red-teaming line: what we test, how we test it, and what ships to frontier lab customers.

This is a senior individual-contributor role. You'll mostly be directing agents rather than managing people, and you'll split your time between reviewing evals, designing new ones, and building the agent tooling that scales your own judgment.

We're an early-stage startup, so you should expect and enjoy that your responsibilities will grow and priorities change quickly.

About us

Our mission is to automate AI safety , to pave the way for a future where the vast majority of AI safety work is done by AI models.

Frontier models already solve coding problems that take humans days, but a model that can be hijacked by a malicious email or web page can't be trusted to work on its own. Before AI can do the work that matters, including AI safety research itself, models have to be robust to attack. So we build the safety and alignment evals, red-teaming programs, and RL environments that find these failures and train them out.

Frontier labs use our evaluations to make their models robust to prompt injection. That only works if the evals are right: a subtly broken task or a wrong grade teaches the model the wrong lesson. You own that bar.

You’ll:

  • Hold the quality bar. Review tasks, transcripts, and red-teamer submissions, and decide what ships to frontier lab customers.
  • Design what we test next. Study where models struggle and why, then design the task types, methodologies, and environments that target the gaps. The goal is training the failure out, not just finding it.
  • Automate your own judgment. Turn your review patterns into agent skills, checkers, and pipeline automation, so the quality bar scales faster than headcount.
  • Improve how we work. Our internal pipelines are agentic too. Apply the same eval eye to them and keep raising how much one person can do.

You’ll work closely with our founders and the teams developing frontier models, with unusual autonomy to make consequential decisions. The work you review shapes system cards, deployment safeguards, and how much the world can trust the most capable AI systems.

Your primary focus will be in our AI Red Teaming workstream. However, we build evals across several areas, and your responsibilities may expand over time.

What we’re looking for
  • Calibrated eval judgment: you can tell when a task, transcript, grade, or environment is subtly wrong, and explain the evidence behind your call.
  • Agent-native engineering: you use LLMs and coding agents as core tools, and you can independently script and automate your own workflows.
  • Sustained attention to detail: you hold the same bar on the hundredth review as on the first, and you look for ways to automate the repeatable parts.
  • Early-stage startup drive: you enjoy a fast pace, shifting priorities, limited structure, and taking on whatever will have the biggest impact.
  • You’ve evaluated something rigorously: an eval, a benchmark, a grading pipeline, or quality assurance you owned for a technical product. Agentic evals, RL environments, and model-training data are the strongest version
  • Experience building LLM judges and rubric-based grading, and iterating on them as models change
  • Experience designing and building evaluation environments yourself; professional software engineering experience is a strong plus
  • Prompt injection, red teaming, or security experience, especially when paired with eval design or grading judgment
  • Familiarity with AI safety and the alignment research community
  • Experience collaborating with frontier labs or other demanding technical customers

These criteria are a guide, not a checklist. If you want to do your life's work making frontier models safer and this role excites you, we encourage you to apply.

  • Location and workspace: Our team works out of Constellation in Berkeley, CA, and we prefer someone who can work alongside us there. We are open to remote for the right candidate.
  • Compensation: $200,000–$400,000 plus equity. More for exceptional candidates.
  • Health coverage: You'll receive a generous monthly pre-tax allowance to choose the medical, dental, and vision coverage that best fits your needs.
  • 401(k): We offer a 401(k) retirement plan.
  • Visa sponsorship: We sponsor visas, although we cannot successfully sponsor every visa for every role or candidate. If we make you an offer, we will make every reasonable effort to secure the visa you need, with support from an immigration lawyer we retain to guide and coordinate the process.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Head of AI Red Teaming
Head of AI Red Teaming

Trajectory Labs, PBC • Berkeley (CA), Northern (KY)

Hybrid
USD 250,000 - 400,000
Equity
Health coverage
401(k)
+1
Software Engineer, Infrastructure & Platform
Software Engineer, Infrastructure & Platform

10a Labs • United States

Remote
USD 110,000 - 160,000
Fully remote, U.S.-based
Performance-based annual bonus
Professional development support (con-
+1
Research Engineer, Frontier Evals & Environments
Research Engineer, Frontier Evals & Environments

OpenAI • California (MO)

On-site
USD 180,000 - 240,000
Member of Technical Staff (Language Model Evaluations)
Member of Technical Staff (Language Model Evaluations)

Artificial Analysis • San Francisco (CA)

On-site
USD 180,000 - 260,000
Equity
Member of Technical Staff (Language Model Evaluations)
Member of Technical Staff (Language Model Evaluations)

Artificial Analysis, Inc. • San Francisco (CA)

On-site
USD 150,000 - 230,000
Equity
Research Engineer, Agent Systems — Frontier AI Lab
Research Engineer, Agent Systems — Frontier AI Lab

Aionia Group • San Francisco (CA)

On-site
USD 300,000 - 600,000
Meaningful equity
Top-of-market compensation
Collaborative environment with researchers
Senior AI Forward Deployed Engineer
Senior AI Forward Deployed Engineer

Handshake • San Francisco (CA)

Hybrid
USD 120,000 - 160,000
Equity in a fast-growing company
401(k) match
Paid parental leave
+2
RL Environments Engineer
RL Environments Engineer

Bespoke Labs • Mountain View (CA)

On-site
USD 250,000 - 300,000
Health, dental, and vision coverage
401(k)
Daily onsite lunch provided
+2
RL Environments Engineer
RL Environments Engineer

Bespoke Labs Inc. • Mountain View (CA), Northern (KY)

Hybrid
USD 250,000 - 300,000
Health, dental, and vision coverage
401(k)
Daily onsite lunch provided
+2
RL Environments Engineer
RL Environments Engineer

Bespoke-Labs • Mountain View (CA)

On-site
USD 250,000 - 300,000
Health, dental, and vision
401(k)
Daily onsite lunch
+3