Stand out for this role — generate a tailored resume and cover letter in about a minute.
Space Executive is seeking a top-tier AI safety researcher to define how the industry measures AI risk. You will release benchmarks every fortnight to three weeks, sometimes publicly and other times privately with labs and developers.
You own taxonomy, evaluation harness, and release quality, working with internal harm-domain experts and freelancers as needed. You will report to the CTO’s office and collaborate with the research lead on public research direction.
This is a role for someone who wants to define how the industry measures AI risk. Every few weeks, you will release a new benchmark targeting a frontier risk area that hasn't yet been properly quantified. Depending on sensitivity, some of this work is published openly, while some is shared privately with model developers. A number of projects are run jointly with major AI labs and academic partners.
You won't be building every eval on your own. For each benchmark, you'll be paired with an internal specialist who owns the relevant harm domain, and you'll have budget to bring in and direct external subject-matter experts. Your ownership covers the taxonomy, the evaluation harness, the quality standard and the final release.
The position reports into the office of the CTO and works closely with the research lead responsible for the company's public research direction. You'll also be able to draw on a large in-house bench of harm-focused researchers whenever a topic calls for extra depth.
Expect to put out a new benchmark roughly every fortnight to three weeks. Scope flexes with the topic: a conversational taxonomy might hold around a hundred evals, whereas an agentic or RL-based benchmark may be closer to twenty, given how costly each item is to review. How sensitive the material is will determine whether it goes public or stays with the labs.
Success looks like this: a frontier lab re-runs your benchmark and reproduces your results. Verifiers are robust, rubrics are unambiguous, the data distribution makes sense, and the lab's own domain experts see the taxonomy as genuinely new. Model-assisted screening is fine, but you'll still go through the evals personally, catch anything that drifts from the taxonomy, and send it back when it isn't right.
You own the plan and the schedule, and you're the one making sure contributing researchers hit their dates. When needed, you'll bring in and steer a small number of freelance experts directly.
Around once a month, you'll join the CTO alongside pod and research leads to review what the research teams are seeing, what customers want, and what's happening across the wider landscape. The outcome is an updated quarterly release plan, aligned with the commercial relationships the business is looking to build.
Roughly a fifth of your time is set aside for the wider ecosystem: keeping up with new research, maintaining relationships with people inside the labs, understanding what's worrying them, and attending a handful of conferences each year. Regular, ideally weekly, conversations with lab contacts are part of the job.
Our client is an established AI safety and security business working with many of the world's leading frontier labs. Its platform protects the technologies people rely on to communicate and create, whether they're interacting with one another or with AI systems.
The company covers the full model lifecycle, from pre-release hardening and red-teaming through to live guardrails and ongoing monitoring, serving frontier labs, enterprises and major consumer platforms. Its team is recognised as one of the strongest in the space, and its work protects users at very large scale.
If you're driven by the challenge of making AI safer and want your research to directly shape how frontier models are tested, we'd love to hear from you.