Stand out for this role — generate a tailored resume and cover letter in about a minute.
Trajectory Labs, PBC is seeking a senior individual-contributor to own the evaluation pipeline for our prompt injection red-teaming line. You will review evals, design new tasks, and build agent tooling that scales your judgment, while working closely with founders and frontier labs.
As an early-stage startup, expect evolving priorities and significant autonomy to shape the role. Location is Berkeley, CA, with openness to remote for the right candidate and a compensation range of
You'll own the evaluation pipeline for our prompt injection red-teaming line: what we test, how we test it, and what ships to frontier lab customers.
This is a senior individual-contributor role. You'll mostly be directing agents rather than managing people, and you'll split your time between reviewing evals, designing new ones, and building the agent tooling that scales your own judgment.
We're an early-stage startup, so you should expect and enjoy that your responsibilities will grow and priorities change quickly.
Our mission is to automate AI safety , to pave the way for a future where the vast majority of AI safety work is done by AI models.
Frontier models already solve coding problems that take humans days, but a model that can be hijacked by a malicious email or web page can't be trusted to work on its own. Before AI can do the work that matters, including AI safety research itself, models have to be robust to attack. So we build the safety and alignment evals, red-teaming programs, and RL environments that find these failures and train them out.
Frontier labs use our evaluations to make their models robust to prompt injection. That only works if the evals are right: a subtly broken task or a wrong grade teaches the model the wrong lesson. You own that bar.
You’ll:
You’ll work closely with our founders and the teams developing frontier models, with unusual autonomy to make consequential decisions. The work you review shapes system cards, deployment safeguards, and how much the world can trust the most capable AI systems.
Your primary focus will be in our AI Red Teaming workstream. However, we build evals across several areas, and your responsibilities may expand over time.
These criteria are a guide, not a checklist. If you want to do your life's work making frontier models safer and this role excites you, we encourage you to apply.