LLM Evaluations Engineer — Benchmark Leaderboards

Vals AI

San Francisco (CA)

On-site

USD 180,000 - 240,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health/dental insurance coverage
Relocation support
Lunch and dinner provided
401K plan
Unlimited PTO

Job summary

Vals AI in San Francisco is seeking engineers to own leaderboards that evaluate LLMs across tasks including law, tax, coding, finance, and more. You will test and benchmark new models as released, analyze error modes, and work with our communications team to publish results.

You will collaborate with open-source and closed-source labs, maintain the model library and integrations, and help improve the benchmarking infrastructure as new models launch in fast sprints.

Qualifications

  • Proficient in Python with professional experience.
  • Strong familiarity with large language models and current benchmarks.
  • Solid engineering fundamentals and ability to ship high-quality code.

Responsibilities

  • Evaluate new LLM releases across Vals AI benchmarks.
  • Assess model performance with open-source and closed-source labs.
  • Use Docent to analyze failure modes and patterns.
  • Collaborate with social media for results publication.
  • Add models and maintain library integrations.
  • Improve benchmark infrastructure for agentic and non-agentic runs.
  • Collaborate with the research team on new benchmarks.
  • Adapt to sprint-driven release cycles.

Skills

Python
LLM familiarity
Engineering fundamentals
Team collaboration
Benchmarking

Tools

Django
AWS
Git
React

Job description

Vals AI in San Francisco is seeking engineers to own leaderboards that evaluate LLMs across tasks including law, tax, coding, finance, and more. You will test and benchmark new models as released, analyze error modes, and work with our communications team to publish results.

You will collaborate with open-source and closed-source labs, maintain the model library and integrations, and help improve the benchmarking infrastructure as new models launch in fast sprints.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

LLM Benchmark Engineer — Lead Leaderboard Insights
LLM Benchmark Engineer — Lead Leaderboard Insights

Vals AI, Inc. • San Francisco (CA)

On-site
USD 150,000 - 190,000
Relocation and transportation support
Health/dental insurance coverage
Lunch and dinner provided
+2
LLM Benchmark Lead Research Scientist
LLM Benchmark Lead Research Scientist

Vals AI • San Francisco (CA)

On-site
USD 120,000 - 180,000
Relocation support
Health insurance
Lunch and snacks provided
+2
Staff Engineer - LLM Benchmark Platform
Staff Engineer - LLM Benchmark Platform

Vals AI, Inc. • San Francisco (CA)

On-site
USD 150,000 - 230,000
Relocation assistance
Health insurance
Dental insurance
+1
Evaluations Engineer
Evaluations Engineer

Vals AI • San Francisco (CA)

On-site
USD 180,000 - 240,000
Health/dental insurance coverage
Relocation support
Lunch and dinner provided
+2
LLM Evaluation & Benchmarking Engineer
LLM Evaluation & Benchmarking Engineer

Capitolis • San Francisco (CA)

On-site
USD 120,000 - 150,000
Senior Model Evaluation Engineer – LLM Benchmarking
Senior Model Evaluation Engineer – LLM Benchmarking

cohere • New York (NY)

Hybrid
USD 150,000 - 210,000
A weekly lunch stipend of $75/£75 or?e
Full health and dental benefits
RRSP matching, 401K, Pension Scheme
+2
Benchmark Architect for AI Evaluation and Research
Benchmark Architect for AI Evaluation and Research

Vals AI, Inc. • San Francisco (CA)

On-site
USD 150,000 - 260,000
Relocation support
Health insurance
Meals provided
+2
Lead AI Engineer — LLM Evaluation & Production Optimizer
Lead AI Engineer — LLM Evaluation & Production Optimizer

LawPro.ai • Kentucky

On-site
USD 150,000 - 190,000
Evaluations Engineer
Evaluations Engineer

Vals AI, Inc. • San Francisco (CA)

On-site
USD 150,000 - 190,000
Relocation and transportation support
Health/dental insurance coverage
Lunch and dinner provided
+2
Remote LLM Evaluation Scientist: Benchmarking Models
Remote LLM Evaluation Scientist: Benchmarking Models

Anyone AI • United States

Remote
MXN 2,626,000 - 3,678,000