LLM Benchmark Architect — Research Lead

Vibehackers

San Francisco, Northern (CA, KY)

Hybrid

USD 140,000 - 185,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Relocation assistance
Housing stipend (within 1 mile)
Health and dental insurance
Meals provided (lunch & dinner)
Unlimited PTO

Job summary

Vibehackers in San Francisco is seeking a Senior Staff Researcher to design and implement novel benchmarks that evaluate real-world LLM capabilities. You will lead research and engineering efforts, publish results, and collaborate with model labs and enterprise partners to shape how foundation models are measured.

The role emphasizes rigorous experimental design, statistical analysis, and scalable benchmark deployment, with relocation support, a competitive salary, health and dental insurance,

Qualifications

  • Advanced research experience (Master’s or PhD in Computer Science, NLP, Machine Learning, or related field)
  • Publication track record in NeurIPS, ICML, ACL, EMNLP
  • Strong understanding of experimental design, statistical analysis, and evaluation frameworks
  • Proficiency in Python for research and experimentation
  • Strong communication skills for technical and non-technical audiences
  • Experience collaborating in research teams and integrating feedback
  • Demonstrated track record of impactful research work / portfolio

Responsibilities

  • Design and implement novel benchmarks that evaluate real-world LLM capabilities
  • Conduct research to ensure benchmarks are valid, reliable, and meaningful
  • Analyze model performance across benchmarks and articulate results
  • Collaborate with foundation model labs, enterprises, and infrastructure teams to implement benchmarks at scale
  • Publish research findings and contribute to the evaluation research community
  • Stay current with developments in LLM capabilities and evaluation methodologies

Skills

Python programming
Strong communication
Research collaboration
Experimental design
Statistical analysis
Data analysis
Project ownership
Rapid learning
Problem solving

Education

Master’s or PhD in CS/NLP/ML

Tools

Django
React
AWS CDK
AWS

Job description

Vibehackers in San Francisco is seeking a Senior Staff Researcher to design and implement novel benchmarks that evaluate real-world LLM capabilities. You will lead research and engineering efforts, publish results, and collaborate with model labs and enterprise partners to shape how foundation models are measured.

The role emphasizes rigorous experimental design, statistical analysis, and scalable benchmark deployment, with relocation support, a competitive salary, health and dental insurance,

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

LLM Benchmark Engineer — Lead Leaderboard Insights
LLM Benchmark Engineer — Lead Leaderboard Insights

Vals AI, Inc. • San Francisco (CA)

On-site
USD 150,000 - 190,000
Relocation and transportation support
Health/dental insurance coverage
Lunch and dinner provided
+2
LLM Benchmark Lead Research Scientist
LLM Benchmark Lead Research Scientist

Vals AI • San Francisco (CA)

On-site
USD 120,000 - 180,000
Relocation support
Health insurance
Lunch and snacks provided
+2
LLM Evaluations Engineer — Benchmark Leaderboards
LLM Evaluations Engineer — Benchmark Leaderboards

Vals AI • San Francisco (CA)

On-site
USD 180,000 - 240,000
Health/dental insurance coverage
Relocation support
Lunch and dinner provided
+2
Member of Technical Staff - Research
Member of Technical Staff - Research

Vibehackers • San Francisco (CA), Northern (KY)

Hybrid
USD 140,000 - 185,000
Relocation assistance
Housing stipend (within 1 mile)
Health and dental insurance
+2
Member of Technical Staff - Research
Member of Technical Staff - Research

Vals AI, Inc. • San Francisco (CA)

On-site
USD 150,000 - 260,000
Relocation support
Health insurance
Meals provided
+2
Platform Engineer for LLM Benchmarking & Tools (SF)
Platform Engineer for LLM Benchmarking & Tools (SF)

Vibehackers • San Francisco (CA), Northern (KY)

Hybrid
USD 140,000 - 185,000
Health insurance
Dental insurance
401K plan
+4
Evaluations Engineer
Evaluations Engineer

Vibehackers • San Francisco (CA), Northern (KY)

Hybrid
USD 140,000 - 185,000
Relocation support
Health insurance
Lunch and dinner provided
+6
Benchmark Architect for AI Evaluation and Research
Benchmark Architect for AI Evaluation and Research

Vals AI, Inc. • San Francisco (CA)

On-site
USD 150,000 - 260,000
Relocation support
Health insurance
Meals provided
+2
Head of Research
Head of Research

Vibehackers • San Francisco (CA), Northern (KY)

Hybrid
USD 225,000 - 275,000
Relocation and transportation support
Health and dental insurance
Lunch and dinner provided
+6
LLM Architecture Researcher
LLM Architecture Researcher

OpenAI • California (MO)

Hybrid
USD 180,000 - 260,000
Relocation assistance
Hybrid work model