LLM Benchmark Engineer — Lead Leaderboard Insights

Vals AI, Inc.

San Francisco (CA)

On-site

USD 150,000 - 190,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Relocation and transportation support
Health/dental insurance coverage
Lunch and dinner provided
401K plan
Unlimited PTO

Job summary

Vals AI, Inc. in San Francisco is seeking strong engineers to own the leaderboards that evaluate real-world tasks by LLMs.

You will test and benchmark new models across domains like law, tax, coding, finance and more, analyzing error modes and collaborating with the communications team to publish results. The role involves working with major labs and institutions, building model library integrations, and maintaining benchmarks infrastructure in a fast, sprint-driven environment.

Qualifications

  • Familiarity with the LLMs and leading models in the space.
  • Strong engineering fundamentals; able to build and ship quickly with high quality.
  • Significant experience in Python in a professional setting.
  • Team collaboration in sprints, Git workflows, and pull request reviews.
  • In-person in San Francisco; relocation support may be provided.

Responsibilities

  • Evaluate new LLM model releases across the Vals AI benchmarks.
  • Work with open-source and closed-source labs to evaluate model performance.
  • Use Docent to analyze common failure modes and patterns in model performance.
  • Collaborate with social media to post findings and results.
  • Add new models and maintain integrations in the model library.
  • Improve and maintain benchmarking infrastructure (agentic and non-agentic).
  • Partner with research to create new benchmarks.

Skills

Python
LLMs familiarity
Engineering fundamentals
Team collaboration

Tools

Docent
Git workflows
Django
React
AWS CDK

Job description

Vals AI, Inc. in San Francisco is seeking strong engineers to own the leaderboards that evaluate real-world tasks by LLMs.

You will test and benchmark new models across domains like law, tax, coding, finance and more, analyzing error modes and collaborating with the communications team to publish results. The role involves working with major labs and institutions, building model library integrations, and maintaining benchmarks infrastructure in a fast, sprint-driven environment.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

LLM Evaluations Engineer — Benchmark Leaderboards
LLM Evaluations Engineer — Benchmark Leaderboards

Vals AI • San Francisco (CA)

On-site
USD 180,000 - 240,000
Health/dental insurance coverage
Relocation support
Lunch and dinner provided
+2
LLM Benchmark Lead Research Scientist
LLM Benchmark Lead Research Scientist

Vals AI • San Francisco (CA)

On-site
USD 120,000 - 180,000
Relocation support
Health insurance
Lunch and snacks provided
+2
Staff Engineer - LLM Benchmark Platform
Staff Engineer - LLM Benchmark Platform

Vals AI, Inc. • San Francisco (CA)

On-site
USD 150,000 - 230,000
Relocation assistance
Health insurance
Dental insurance
+1
Evaluations Engineer
Evaluations Engineer

Vibehackers • San Francisco (CA), Northern (KY)

Hybrid
USD 140,000 - 185,000
Relocation support
Health insurance
Lunch and dinner provided
+6
LLM Benchmark Architect — Research Lead
LLM Benchmark Architect — Research Lead

Vibehackers • San Francisco (CA), Northern (KY)

Hybrid
USD 140,000 - 185,000
Relocation assistance
Housing stipend (within 1 mile)
Health and dental insurance
+2
Platform Engineer for LLM Benchmarking & Tools (SF)
Platform Engineer for LLM Benchmarking & Tools (SF)

Vibehackers • San Francisco (CA), Northern (KY)

Hybrid
USD 140,000 - 185,000
Health insurance
Dental insurance
401K plan
+4
Evaluations Engineer
Evaluations Engineer

Vals AI • San Francisco (CA)

On-site
USD 180,000 - 240,000
Health/dental insurance coverage
Relocation support
Lunch and dinner provided
+2
Benchmark Architect for AI Evaluation and Research
Benchmark Architect for AI Evaluation and Research

Vals AI, Inc. • San Francisco (CA)

On-site
USD 150,000 - 260,000
Relocation support
Health insurance
Meals provided
+2
LLM Evaluation & Benchmarking Engineer
LLM Evaluation & Benchmarking Engineer

Capitolis • San Francisco (CA)

On-site
USD 120,000 - 150,000
ML Evaluation Scientist - LLM Benchmarks
ML Evaluation Scientist - LLM Benchmarks

Scale AI • Seattle (WA)

On-site
USD 181,000 - 226,000
Health coverage
Retirement benefits
L&D stipend
+2