Frontier AI Safety Benchmark Architect

Space Executive

United States

On-site

USD 150,000 - 260,000

Full time

2 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Space Executive is seeking a top-tier AI safety researcher to define how the industry measures AI risk. You will release benchmarks every fortnight to three weeks, sometimes publicly and other times privately with labs and developers.

You own taxonomy, evaluation harness, and release quality, working with internal harm-domain experts and freelancers as needed. You will report to the CTO’s office and collaborate with the research lead on public research direction.

Qualifications

  • Masters or PhD in computer science, ML, or a closely related field, or equivalent industry depth.
  • At least three years designing and operating safety or security evaluations for LLMs in production.
  • A track record of AI safety and security publications, with five+ papers and first authors on two or more.
  • Strong engineering skills: building eval harnesses, distributed inference, vLLM, and codebase debugging.
  • Ability to design a taxonomy from scratch, not just evaluate against an existing one.
  • Experience leading researchers and freelancers through influence rather than formal management.
  • Excellent written and verbal English with strong cross-timezone communication.
  • Genuine curiosity about harms being studied and ability to learn new domains quickly.

Responsibilities

  • Deliver benchmarks with a structured taxonomy, evaluation harness, and quality standards.
  • Own plan and schedule, coordinate internal specialists and external experts.
  • Shaping quarterly roadmap with CTO and research leads; align with commercial goals.
  • Maintain industry connections; stay current with new research and conferences.
  • Collaborate with frontier labs, model providers and academic partners to publish findings.

Skills

Safety evaluation design
Leadership of researchers & freelance
English communication
AI safety publishing
Taxonomy design
Engineering for eval harnesses

Education

Masters or PhD in CS/ML or equivalent

Tools

vLLM

Job description

Space Executive is seeking a top-tier AI safety researcher to define how the industry measures AI risk. You will release benchmarks every fortnight to three weeks, sometimes publicly and other times privately with labs and developers.

You own taxonomy, evaluation harness, and release quality, working with internal harm-domain experts and freelancers as needed. You will report to the CTO’s office and collaborate with the research lead on public research direction.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Frontier AI Safety Benchmark Lead
Frontier AI Safety Benchmark Lead

Alice • San Francisco (CA)

Hybrid
USD 180,000 - 230,000
Conference travel
Frontier AI Safety Evaluator
Frontier AI Safety Evaluator

Obsidian • San Francisco (CA)

Remote
USD 150,000 - 210,000
Frontier AI Safety Evaluator & Alignment Expert
Frontier AI Safety Evaluator & Alignment Expert

Obsidian • New York (NY)

On-site
USD 120,000 - 180,000
AI Safety Benchmark Lead & Evaluation Architect
AI Safety Benchmark Lead & Evaluation Architect

Alice • New York (NY)

On-site
USD 180,000 - 250,000
Frontier AI Safety Evaluator & Policy Expert
Frontier AI Safety Evaluator & Policy Expert

Obsidian • San Francisco (CA)

On-site
USD 140,000 - 210,000
Lead, AI Safety Benchmarks & Evaluations
Lead, AI Safety Benchmarks & Evaluations

Alice (Formerly ActiveFence) • United States

On-site
USD 180,000 - 280,000
Frontier AI Safety Red Team Expert
Frontier AI Safety Red Team Expert

Obsidian • New York (NY)

Remote
USD 140,000 - 210,000
Frontier AI Safety Evaluator & Policy Auditor
Frontier AI Safety Evaluator & Policy Auditor

Obsidian • New York (NY)

On-site
USD 120,000 - 180,000
Frontier AI Safety Evaluator & Alignment Specialist
Frontier AI Safety Evaluator & Alignment Specialist

Obsidian • San Francisco (CA)

On-site
USD 120,000 - 160,000
null
Frontier AI Safety Evaluator
Frontier AI Safety Evaluator

Dorado • United States

Remote
USD 110,000 - 170,000