Frontier AI Safety Benchmark Lead

Alice

San Francisco (CA)

Hybrid

USD 180,000 - 230,000

Full time

8 days ago
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Conference travel

Job summary

Alice is seeking a science-led leader to own and ship AI safety benchmarks. You will collaborate with an in-house researcher for each harm area, with a budget to direct freelancers. The role sits in the CTO office with a large team of researchers.

You will manage the taxonomy, harness, and release schedule, guiding a ~150-person harms team. You will set quarterly release plans in collaboration with 연구 leads and CTO, ensuring alignment with client and media interests.

Qualifications

  • PhD or Masters in computer science, machine learning or a related field, or equivalent depth from industry research.
  • 3+ years building and running safety or security evaluations for language models in production, at an AI lab, a model provider, or a safety and security research organisation.
  • 5+ relevant research publications in AI safety and security including lead author on at least 2 of them
  • Strong engineer with experience in evaluation harnesses, distributed inference, vLLM, reading and fixing codebase
  • You can build a taxonomy, not only score against one
  • You can direct a researcher and two freelancers without managing them formally
  • Strong English, written and spoken; capable of cross-time-zone communication
  • Curiosity about harms; you'll learn a new subject every three weeks

Responsibilities

  • Ship a benchmark cadence roughly every two to three weeks and manage scope based on subject.
  • Own the quality bar: ensure verifications, rubrics, distributions, and taxonomy are sound.
  • Run the process: keep timelines and coordinate two to three freelancers as needed.
  • Set the roadmap with the forum: monthly reviews with CTO, pod, and research leads to revise the release plan.
  • Stay ahead of the curve: engage with labs, read research, and travel to conferences regularly.

Skills

Strong English
Curiosity about harms
Build taxonomy
Direct researchers (no formal mgmt)

Education

PhD or Masters in CS/ML or related field
Equivalent depth from industry research

Tools

vLLM
Distributed inference
Evaluation harnesses

Job description

Alice is seeking a science-led leader to own and ship AI safety benchmarks. You will collaborate with an in-house researcher for each harm area, with a budget to direct freelancers. The role sits in the CTO office with a large team of researchers.

You will manage the taxonomy, harness, and release schedule, guiding a ~150-person harms team. You will set quarterly release plans in collaboration with 연구 leads and CTO, ensuring alignment with client and media interests.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Safety Benchmark Lead & Evaluation Architect
AI Safety Benchmark Lead & Evaluation Architect

Alice • New York (NY)

On-site
USD 180,000 - 250,000
AI Safety Benchmarks Lead — Ship Frontiers & Roadmaps
AI Safety Benchmarks Lead — Ship Frontiers & Roadmaps

Alice (Formerly ActiveFence) • New York (NY)

On-site
USD 190,000 - 240,000
Lead, AI Safety Benchmarks & Evaluations
Lead, AI Safety Benchmarks & Evaluations

Alice (Formerly ActiveFence) • United States

On-site
USD 180,000 - 280,000
Research Lead, Evaluations and Benchmarks
Research Lead, Evaluations and Benchmarks

Alice (Formerly ActiveFence) • New York (NY)

On-site
USD 190,000 - 240,000
Research Lead, Evaluations and Benchmarks
Research Lead, Evaluations and Benchmarks

Alice (Formerly ActiveFence) • United States

On-site
USD 180,000 - 280,000
Researcher, Evaluations and Benchmarks
Researcher, Evaluations and Benchmarks

Alice • New York (NY)

On-site
USD 180,000 - 250,000
Researcher, Evaluations and Benchmarks
Researcher, Evaluations and Benchmarks

Alice • San Francisco (CA)

On-site
USD 180,000 - 230,000
Conference travel
Frontier AI Security Architect - Bay Area & Labs Liaison
Frontier AI Security Architect - Bay Area & Labs Liaison

ActiveFence • San Francisco (CA)

On-site
USD 180,000 - 240,000
Research Lead, AI Safety & Impact
Research Lead, AI Safety & Impact

FAR.AI • United States

Hybrid
USD 170,000 - 270,000
Catered lunch and dinner at Berkeley
Member of the Technical Staff
Member of the Technical Staff

Alice (Formerly ActiveFence) • San Francisco (CA)

On-site
USD 170,000 - 260,000
Bay Area base
Public speaking opportunities