Artificial Intelligence Researcher

Space Executive

United States

On-site

USD 150,000 - 260,000

Full time

7 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Space Executive is seeking a top-tier AI safety researcher to define how the industry measures AI risk. You will release benchmarks every fortnight to three weeks, sometimes publicly and other times privately with labs and developers.

You own taxonomy, evaluation harness, and release quality, working with internal harm-domain experts and freelancers as needed. You will report to the CTO’s office and collaborate with the research lead on public research direction.

Qualifications

  • Masters or PhD in computer science, ML, or a closely related field, or equivalent industry depth.
  • At least three years designing and operating safety or security evaluations for LLMs in production.
  • A track record of AI safety and security publications, with five+ papers and first authors on two or more.
  • Strong engineering skills: building eval harnesses, distributed inference, vLLM, and codebase debugging.
  • Ability to design a taxonomy from scratch, not just evaluate against an existing one.
  • Experience leading researchers and freelancers through influence rather than formal management.
  • Excellent written and verbal English with strong cross-timezone communication.
  • Genuine curiosity about harms being studied and ability to learn new domains quickly.

Responsibilities

  • Deliver benchmarks with a structured taxonomy, evaluation harness, and quality standards.
  • Own plan and schedule, coordinate internal specialists and external experts.
  • Shaping quarterly roadmap with CTO and research leads; align with commercial goals.
  • Maintain industry connections; stay current with new research and conferences.
  • Collaborate with frontier labs, model providers and academic partners to publish findings.

Skills

Safety evaluation design
Leadership of researchers & freelance
English communication
AI safety publishing
Taxonomy design
Engineering for eval harnesses

Education

Masters or PhD in CS/ML or equivalent

Tools

vLLM

Job description

This is a role for someone who wants to define how the industry measures AI risk. Every few weeks, you will release a new benchmark targeting a frontier risk area that hasn't yet been properly quantified. Depending on sensitivity, some of this work is published openly, while some is shared privately with model developers. A number of projects are run jointly with major AI labs and academic partners.

You won't be building every eval on your own. For each benchmark, you'll be paired with an internal specialist who owns the relevant harm domain, and you'll have budget to bring in and direct external subject-matter experts. Your ownership covers the taxonomy, the evaluation harness, the quality standard and the final release.

The position reports into the office of the CTO and works closely with the research lead responsible for the company's public research direction. You'll also be able to draw on a large in-house bench of harm-focused researchers whenever a topic calls for extra depth.

What You'll Be Doing
Delivering on a fast release rhythm.

Expect to put out a new benchmark roughly every fortnight to three weeks. Scope flexes with the topic: a conversational taxonomy might hold around a hundred evals, whereas an agentic or RL-based benchmark may be closer to twenty, given how costly each item is to review. How sensitive the material is will determine whether it goes public or stays with the labs.

Holding the standard.

Success looks like this: a frontier lab re-runs your benchmark and reproduces your results. Verifiers are robust, rubrics are unambiguous, the data distribution makes sense, and the lab's own domain experts see the taxonomy as genuinely new. Model-assisted screening is fine, but you'll still go through the evals personally, catch anything that drifts from the taxonomy, and send it back when it isn't right.

Keeping delivery on track.

You own the plan and the schedule, and you're the one making sure contributing researchers hit their dates. When needed, you'll bring in and steer a small number of freelance experts directly.

Shaping the roadmap.

Around once a month, you'll join the CTO alongside pod and research leads to review what the research teams are seeing, what customers want, and what's happening across the wider landscape. The outcome is an updated quarterly release plan, aligned with the commercial relationships the business is looking to build.

Staying plugged into the field.

Roughly a fifth of your time is set aside for the wider ecosystem: keeping up with new research, maintaining relationships with people inside the labs, understanding what's worrying them, and attending a handful of conferences each year. Regular, ideally weekly, conversations with lab contacts are part of the job.

What We're Looking For
  • A Masters or PhD in computer science, machine learning or a similar discipline, or comparable depth gained through industry research.
  • At least three years designing and operating safety or security evaluations for LLMs in production, whether at a frontier lab, a model provider or a specialist safety/security research group.
  • A track record of publishing in AI safety and security, with five or more relevant papers and first authorship on at least two.
  • Solid engineering skills: building eval harnesses, distributed inference, vLLM, and the ability to dig into an unfamiliar codebase and fix it.
  • The ability to design a taxonomy from scratch, not just evaluate against an existing one.
  • Comfort leading a researcher and a couple of freelancers through influence rather than formal line management.
  • Excellent written and verbal English; much of the work depends on clear communication across time zones.
  • Genuine curiosity about the harms being studied, and an appetite for getting up to speed on a new domain every few weeks.
Bonus Points
  • Post-training experience across SFT, DPO or GRPO, particularly reward design for subjective or safety-critical objectives.
  • Experience evaluating agentic systems, including tool use, orchestration, permissioning and prompt injection.
  • Papers accepted at leading ML or security venues.
  • Confidence presenting your research to customers and to large or senior audiences.
  • Happy to travel to three or more conferences a year.
About the Company

Our client is an established AI safety and security business working with many of the world's leading frontier labs. Its platform protects the technologies people rely on to communicate and create, whether they're interacting with one another or with AI systems.

The company covers the full model lifecycle, from pre-release hardening and red-teaming through to live guardrails and ongoing monitoring, serving frontier labs, enterprises and major consumer platforms. Its team is recognised as one of the strongest in the space, and its work protects users at very large scale.

If you're driven by the challenge of making AI safer and want your research to directly shape how frontier models are tested, we'd love to hear from you.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Research Lead, Evaluations and Benchmarks
Research Lead, Evaluations and Benchmarks

Alice (Formerly ActiveFence) • New York (NY)

On-site
USD 190,000 - 240,000
Research Lead, Evaluations and Benchmarks
Research Lead, Evaluations and Benchmarks

Alice (Formerly ActiveFence) • United States

On-site
USD 180,000 - 280,000
Researcher, Evaluations and Benchmarks
Researcher, Evaluations and Benchmarks

Alice • New York (NY)

On-site
USD 180,000 - 250,000
Researcher, Evaluations and Benchmarks
Researcher, Evaluations and Benchmarks

Alice • San Francisco (CA)

On-site
USD 180,000 - 230,000
Conference travel
Research Scientist, Frontier Risk Evaluations
Research Scientist, Frontier Risk Evaluations

Scale • San Francisco (CA), New York (NY)

On-site
USD 216,000 - 270,000
Health coverage
Retirement benefits
Learning stipend
+2
Research Scientist, Frontier Risk Evaluations San Francisco, CA Apply →
Research Scientist, Frontier Risk Evaluations San Francisco, CA Apply →

Scale AI, Inc. • New York (NY)

On-site
USD 90,000 - 130,000
Research Scientist, Frontier Risk Evaluations
Research Scientist, Frontier Risk Evaluations

Scale AI • San Francisco (CA)

On-site
USD 197,400 - 246,750
Comprehensive health insurance
Retirement benefits
Generous PTO
+1
Senior Member of Technical Staff - Model Safety
Senior Member of Technical Staff - Model Safety

Xcede • San Francisco (CA)

On-site
USD 180,000 - 240,000
AI Safety Practitioner - Expert Evaluator
AI Safety Practitioner - Expert Evaluator

Obsidian • San Francisco (CA)

On-site
USD 140,000 - 210,000
Strategic Projects Lead, Red Team
Strategic Projects Lead, Red Team

Front Door Defense • New York (NY)

On-site
USD 152,000 - 190,000
Comprehensive health coverage
Retirement benefits
Learning and development stipend
+2