Head of Research

Vibehackers

San Francisco, Northern (CA, KY)

Hybrid

USD 225,000 - 275,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Relocation and transportation support
Health and dental insurance
Lunch and dinner provided
Unlimited PTO
Housing stipend
401(k) plan
Free snacks/coffee/drinks
Equity
Collaboration with leading AI labs

Job summary

Vals AI in San Francisco seeks a Head of Research to define and advance evaluation methodologies for large language models and AI systems. You will set research direction across the portfolio, publish influential work, and recruit a research team while partnering with enterprise customers and lab partners.

The role emphasizes real‑world deployment impact, collaboration with AI labs, and leadership of a high‑caliber team.

Qualifications

  • PhD in ML/NLP or equivalent industry track record.
  • Deep familiarity with the LLM evaluation landscape, including existing benchmarks and judge‑model approaches.
  • Preference for research that influences real‑world deployments rather than gamed benchmarks.
  • Strong written and verbal communication for publishing, presenting, and customer engagement.
  • Ability to work in‑person in San Francisco.

Responsibilities

  • Develop new paradigms for evaluating long-horizon, real‑world tasks that current benchmarks fail to capture.
  • Oversee and set direction across Vals’ research portfolio and ongoing projects.
  • Publish research and present results to the community, customers, and partners.
  • Recruit, mentor, and grow a high-quality research team alongside the founders.
  • Collaborate closely with enterprise customers and lab partners to solve practical evaluation problems.

Skills

Research leadership
Experimental design
LLM evaluation
Publication & communication
Team building
Mentorship
Customer-facing
Collaboration
Domain expertise
Learning agility
Ownership
Problem solving

Education

PhD in ML/NLP

Tools

Python
AWS
AWS CDK
Django
React

Job description

Directly focused on building benchmarks and evaluation frameworks for LLMs and AI models — closely tied to vibe-coding and benchmarking work.

About the Role

Lead Vals AI's research efforts to design and validate new evaluation methodologies and benchmarks for LLMs and other AI systems, set research direction, publish influential work, and build and manage a research team while working closely with enterprise customers and lab partners in San Francisco.

Job Description
Role

Head of Research responsible for defining and advancing evaluation methodologies and benchmarks for large language models and other AI systems. The role sets research direction across Vals’ portfolio, publishes work that moves the field forward, recruits and grows a research team, and partners directly with enterprise customers and AI labs on real‑world evaluation problems.

Key Responsibilities
  • Develop new paradigms for evaluating long-horizon, real‑world tasks that current benchmarks and judge‑model approaches fail to capture.
  • Oversee and set direction across Vals’ research portfolio and ongoing projects.
  • Publish research and present results to the community, customers, and partners.
  • Recruit, mentor, and grow a high-quality research team alongside the founders.
  • Collaborate closely with enterprise customers and lab partners to solve practical evaluation problems.
Requirements
  • PhD in ML/NLP (in progress or completed) or equivalent industry research track record.
  • Deep familiarity with the LLM evaluation landscape, including existing benchmarks, failure modes, judge‑model approaches, and human‑in‑the‑loop methodologies.
  • Preference for research that influences real‑world deployments rather than easily gamed benchmarks.
  • Strong written and verbal communication skills for publishing, presenting, and customer engagement.
  • Ability to work in‑person in San Francisco.
Nice‑to‑Haves
  • A widely cited benchmark or evaluation framework you built or co‑built.
  • Prior experience at a frontier lab (e.g., Anthropic, OpenAI, Google DeepMind, Meta FAIR) or a research‑led startup.
  • Domain depth in verticals such as legal, finance, insurance, or healthcare.
  • Experience leading or mentoring other researchers and maintaining a public research presence (papers, blog posts, talks, OSS contributions).
What We Offer
  • Highly competitive salary and equity.
  • Relocation and transportation support.
  • Health and dental insurance coverage.
  • Lunch and dinner provided, plus free snacks, coffee, and drinks.
  • 401(k) plan.
  • Unlimited PTO.
  • $1,500 housing stipend (within one‑mile radius).
  • Collaboration with leading AI labs and opportunity to work on foundational evaluation research.
  • Python
  • Django
  • React
  • AWS
  • AWS CDK
Compensation
  • Compensation Range: $225,000 - $275,000 (USD)
Skills

Research Leadership Experimental Design LLM Evaluation Publication & Communication Team Building Mentorship Customer‑facing Collaboration Domain Expertise Learning Agility Ownership Problem Solving

Experience Level

USD 225,000 - 275,000/year

Employment Type

Full-time

  • Relocation and transportation support
  • Dental insurance
  • Lunch and dinner provided
  • Free snacks/coffee/drinks
  • $1,500 housing stipend (within one‑mile radius)
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Head of Research
Head of Research

Vals AI, Inc. • San Francisco (CA)

On-site
USD 220,000 - 340,000
Relocation support
Health insurance
Meals provided (Lunch/Dinner)
+3
Evaluations Engineer
Evaluations Engineer

Vibehackers • San Francisco (CA), Northern (KY)

On-site
USD 140,000 - 185,000
Relocation support
Health insurance
Lunch and dinner provided
+6
Member of Technical Staff - Research
Member of Technical Staff - Research

Vibehackers • San Francisco (CA), Northern (KY)

On-site
USD 140,000 - 185,000
Relocation assistance
Housing stipend (within 1 mile)
Health and dental insurance
+2
Member of Technical Staff - Platform
Member of Technical Staff - Platform

Vibehackers • San Francisco (CA), Northern (KY)

Hybrid
USD 140,000 - 185,000
Health insurance
Dental insurance
401K plan
+4
Member of Technical Staff - Platform
Member of Technical Staff - Platform

Vals AI, Inc. • San Francisco (CA)

On-site
USD 150,000 - 230,000
Relocation assistance
Health insurance
Dental insurance
+1
Evaluations Engineer
Evaluations Engineer

Vals AI, Inc. • San Francisco (CA)

On-site
USD 150,000 - 190,000
Relocation and transportation support
Health/dental insurance coverage
Lunch and dinner provided
+2
Member of Technical Staff - Research
Member of Technical Staff - Research

Vals AI, Inc. • San Francisco (CA)

On-site
USD 150,000 - 260,000
Relocation support
Health insurance
Meals provided
+2
Member of ML Technical Staff
Member of ML Technical Staff

Pragmatike • San Francisco (CA)

On-site
USD 200,000 - 350,000
Research Scientist
Research Scientist

Anyone AI Inc. • Northern (KY)

On-site
USD 110,000 - 160,000
AI Engineer - VA
AI Engineer - VA

LawPro.ai • Virginia (MN)

On-site
USD 140,000 - 200,000