Head of AI Evaluation & Benchmarks

Vibehackers

San Francisco, Northern (CA, KY)

Hybrid

USD 225,000 - 275,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Relocation and transportation support
Health and dental insurance
Lunch and dinner provided
Unlimited PTO
Housing stipend
401(k) plan
Free snacks/coffee/drinks
Equity
Collaboration with leading AI labs

Job summary

Vals AI in San Francisco seeks a Head of Research to define and advance evaluation methodologies for large language models and AI systems. You will set research direction across the portfolio, publish influential work, and recruit a research team while partnering with enterprise customers and lab partners.

The role emphasizes real‑world deployment impact, collaboration with AI labs, and leadership of a high‑caliber team.

Qualifications

  • PhD in ML/NLP or equivalent industry track record.
  • Deep familiarity with the LLM evaluation landscape, including existing benchmarks and judge‑model approaches.
  • Preference for research that influences real‑world deployments rather than gamed benchmarks.
  • Strong written and verbal communication for publishing, presenting, and customer engagement.
  • Ability to work in‑person in San Francisco.

Responsibilities

  • Develop new paradigms for evaluating long-horizon, real‑world tasks that current benchmarks fail to capture.
  • Oversee and set direction across Vals’ research portfolio and ongoing projects.
  • Publish research and present results to the community, customers, and partners.
  • Recruit, mentor, and grow a high-quality research team alongside the founders.
  • Collaborate closely with enterprise customers and lab partners to solve practical evaluation problems.

Skills

Research leadership
Experimental design
LLM evaluation
Publication & communication
Team building
Mentorship
Customer-facing
Collaboration
Domain expertise
Learning agility
Ownership
Problem solving

Education

PhD in ML/NLP

Tools

Python
AWS
AWS CDK
Django
React

Job description

Vals AI in San Francisco seeks a Head of Research to define and advance evaluation methodologies for large language models and AI systems. You will set research direction across the portfolio, publish influential work, and recruit a research team while partnering with enterprise customers and lab partners.

The role emphasizes real‑world deployment impact, collaboration with AI labs, and leadership of a high‑caliber team.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Chief AI Evaluation & Research
Chief AI Evaluation & Research

Vals AI, Inc. • San Francisco (CA)

On-site
USD 220,000 - 340,000
Relocation support
Health insurance
Meals provided (Lunch/Dinner)
+3
Benchmark Architect for AI Evaluation and Research
Benchmark Architect for AI Evaluation and Research

Vals AI, Inc. • San Francisco (CA)

On-site
USD 150,000 - 260,000
Relocation support
Health insurance
Meals provided
+2
AI Evaluation Lead: Real-World Systems Benchmarking
AI Evaluation Lead: Real-World Systems Benchmarking

SupportFinity™ • San Francisco (CA)

On-site
USD 150,000 - 230,000
Head of Research
Head of Research

Vibehackers • San Francisco (CA), Northern (KY)

Hybrid
USD 225,000 - 275,000
Relocation and transportation support
Health and dental insurance
Lunch and dinner provided
+6
AI Benchmarking & Strategy Lead
AI Benchmarking & Strategy Lead

Artificial Analysis • San Francisco (CA)

On-site
USD 180,000 - 260,000
Equity
Competitive compensation
AI Benchmarks & Evaluations Program Manager
AI Benchmarks & Evaluations Program Manager

Mercor • San Francisco (CA)

On-site
USD 120,000 - 200,000
Performance bonus structure
Equity grant
$15K relocation bonus
+7
GenAI Evaluation Scientist - LLM Benchmarks & Failures
GenAI Evaluation Scientist - LLM Benchmarks & Failures

Scale • San Francisco (CA), Seattle (WA), New York (NY)

On-site
USD 166,000 - 207,000
Health insurance
Dental coverage
Vision coverage
+2
Head of AI Engineering - ML & Knowledge Systems
Head of AI Engineering - ML & Knowledge Systems

Highlight AI Inc. • San Francisco (CA)

On-site
USD 180,000 - 220,000
Competitive salary and generous equity package
Health, dental, and vision insurance
Flexible PTO and parental leave
+2
Member of Technical Staff (Language Model Evaluations)
Member of Technical Staff (Language Model Evaluations)

Artificial Analysis • San Francisco (CA)

On-site
USD 180,000 - 260,000
Equity
Head of Clinical AI Research & Evaluation
Head of Clinical AI Research & Evaluation

Amigo • New York (NY)

On-site
USD 200,000 - 320,000
Health insurance
Catered lunch and dinner
Mental health support
+6