Research Lead, AI Safety Benchmarks & Evaluations

Alice

New York (NY)

On-site

USD 180,000 - 250,000

Full time

8 days ago
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Alice is seeking a Research Lead to drive evaluations and benchmarks in a fast‑moving AI safety context. You will own the taxonomy, harness, and release cycles, coordinating with in-house researchers and a budget for freelancers.

The role sits in the CTO office with ~150 researchers focusing on harms and safety across models and platforms. You will ship biweekly benchmarks, ensure rigorous quality standards, and manage a small pool of SME freelancers.

Qualifications

  • PhD or Masters in computer science, machine learning or related field, or equivalent depth from industry.
  • 3+ years building and running safety or security evaluations for language models in production.
  • 5+ relevant research publications in AI safety and security including lead author on at least 2 of them
  • Strong engineer: evaluation harnesses, distributed inference, vLLM, reading a codebase and fixing it
  • You can build a taxonomy, not only score against one.
  • You can direct a researcher and two freelancers without managing them formally.
  • Strong English, written and spoken across time zones.
  • Curiosity about the harms themselves: learn a new subject every three weeks.

Responsibilities

  • Ship a benchmark cadence every two to three weeks.
  • Own the quality bar: verifiers, rubrics, and taxonomy are clear and robust.
  • Direct two to three freelancers (SMEs) ad-hoc when needed.
  • Set the roadmap with the CTO, pod and research leads; align with accounts.
  • Stay ahead of the curve by reading research, maintaining lab contacts, and attending conferences.

Skills

Strong English
Leadership
Research coordination
Budget management
Freelancer management

Education

PhD or Masters in computer science, ML or related field

Tools

Evaluation harnesses
Distributed inference
vLLM

Job description

Alice is seeking a Research Lead to drive evaluations and benchmarks in a fast‑moving AI safety context. You will own the taxonomy, harness, and release cycles, coordinating with in-house researchers and a budget for freelancers.

The role sits in the CTO office with ~150 researchers focusing on harms and safety across models and platforms. You will ship biweekly benchmarks, ensure rigorous quality standards, and manage a small pool of SME freelancers.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead, AI Safety Benchmarks & Evaluations
Lead, AI Safety Benchmarks & Evaluations

Alice (Formerly ActiveFence) • United States

On-site
USD 180,000 - 280,000
AI Safety Benchmarks Lead — Ship Frontiers & Roadmaps
AI Safety Benchmarks Lead — Ship Frontiers & Roadmaps

Alice (Formerly ActiveFence) • New York (NY)

On-site
USD 190,000 - 240,000
Research Lead, Evaluations and Benchmarks
Research Lead, Evaluations and Benchmarks

Alice (Formerly ActiveFence) • New York (NY)

On-site
USD 190,000 - 240,000
Research Lead, Evaluations and Benchmarks
Research Lead, Evaluations and Benchmarks

Alice • New York (NY)

On-site
USD 180,000 - 250,000
Research Lead, Evaluations and Benchmarks
Research Lead, Evaluations and Benchmarks

Alice (Formerly ActiveFence) • United States

On-site
USD 180,000 - 280,000
Research Lead: Pre-Training Safety & Safe AI
Research Lead: Pre-Training Safety & Safe AI

FAR.AI • United States

On-site
USD 160,000 - 230,000
Research Lead: AI Safety & Pre-Training at Scale
Research Lead: AI Safety & Pre-Training at Scale

FAR.AI • Berkeley (CA)

On-site
USD 180,000 - 240,000
Research Lead, AI Safety & Pre-Training
Research Lead, AI Safety & Pre-Training

AISafety • Northern (KY)

Hybrid
USD 150,000 - 230,000
Senior AI Safety Research Lead
Senior AI Safety Research Lead

Aisafety • San Francisco (CA)

On-site
USD 170,000 - 260,000
Health insurance
401K plan + 4% matching
Unlimited PTO
+2
Research Lead, AI Safety & Impact Architect
Research Lead, AI Safety & Impact Architect

Aisafety • Berkeley (CA)

Hybrid
USD 170,000 - 270,000
Catered lunch and dinner
Visa sponsorship