Evaluations Engineer

Vibehackers

San Francisco, Northern (CA, KY)

Hybrid

USD 140,000 - 185,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Relocation support
Health insurance
Lunch and dinner provided
Free snacks and drinks
401(k) plan
Unlimited PTO
Housing stipend
Competitive salary and ownership
Tech stack: Django/React

Job summary

Vals AI in San Francisco, CA, seeks a Senior engineer to own leaderboards evaluating LLMs. You will test new model releases against benchmarks, analyze failure modes, maintain integrations, and improve benchmarking infrastructure.

You will collaborate with labs and the communications team to publish findings and support rapid sprints after model launches.

Qualifications

  • Familiarity with the LLM space and leading models.

Responsibilities

  • Evaluate LLM releases across benchmarks and analyze failure modes.
  • Maintain model integrations and benchmarking infrastructure.
  • Collaborate with communications to publish results.
  • Add new models to the library and update integrations.
  • Operate benchmarking infrastructure across workloads.
  • Publish results used by startups, enterprises, and research labs.

Skills

LLM evaluation
Benchmarking
Error analysis
Engineering fundamentals
Collaboration
Git workflows
Pull request reviews
Technical communication
Technical writing
Ownership
Learning velocity
Time management under sprint
Problem solving

Tools

Docent
AWS CDK
Python

Job description

Directly involves LLM benchmarking and leaderboards; linked to vibe-coding benchmark coverage and rapid sprints around model releases.

About the Role

Join Vals AI to own and operate leaderboards that evaluate LLMs: test new model releases against benchmarks, analyze error modes, and publish results used by startups, enterprises, and research labs. Work with foundation model labs, maintain model integrations and benchmarking infrastructure, and collaborate with communications to share findings.

Job Description
Role

You will own the leaderboards on Vals AI by evaluating new LLM releases across a suite of benchmarks, analyzing failure modes and model strengths, maintaining model integrations, and helping to improve the infrastructure used to run benchmarks.

Key Responsibilities
  • Evaluate new LLM model releases across Vals AI benchmarks covering tasks such as law, tax, coding, finance, and social mobility.
  • Work directly with open-source and closed-source foundation model labs to assess model performance.
  • Use tools like Docent to analyze common failure modes and performance patterns.
  • Add new models and maintain integrations in the model library.
  • Maintain and improve benchmarking infrastructure (agentic and non-agentic workloads).
  • Collaborate with the communications/social media team to publish and post results.
  • Operate on a release-driven cadence: expect intensive sprints after major model launches and quieter periods between releases.
Requirements
  • Familiarity with the LLM space, including knowledge of leading models and their relative strengths.
  • Strong engineering fundamentals and a track record of building and shipping significant projects.
  • Significant professional experience with Python.
  • Experience with development sprints, Git workflows, and pull request reviews.
  • Willingness to work long hours during model releases to deliver high-quality results under tight deadlines.
Nice-to-Haves
  • Previous experience benchmarking large language models or creating evaluation benchmarks.
  • Startup experience or experience founding a company.
  • Technical writing ability.
  • Machine learning research experience.
What We Offer
  • Highly competitive salary and meaningful ownership.
  • Relocation and transportation support.
  • Health and dental insurance coverage.
  • Lunch and dinner provided; free snacks, coffee, and drinks.
  • 401(k) plan.
  • Unlimited PTO.
  • $1,500 housing stipend (within one-mile radius).
  • Django backend and React frontend.
  • Infrastructure on AWS using CDK for infrastructure-as-code.
  • Use of Docent for failure-mode analysis.
Location
  • In-person team based in San Francisco; relocation or transportation support available.
Skills

LLM evaluation Benchmarking Error analysis Engineering fundamentals Collaboration Git workflows Pull request reviews Technical communication Technical writing Ownership Learning velocity Time management under sprint conditions Solution-oriented problem solving

Experience Level

Senior

USD 140,000 - 185,000/year

Employment Type

Full-time

  • Relocation and transportation support
  • Lunch and dinner provided
  • Free snacks/coffee/drinks
  • $1,500 housing stipend (within one-mile radius)
  • Highly competitive salary and meaningful ownership
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Member of Technical Staff - Platform
Member of Technical Staff - Platform

Vibehackers • San Francisco (CA), Northern (KY)

Hybrid
USD 140,000 - 185,000
Health insurance
Dental insurance
401K plan
+4
Evaluations Engineer
Evaluations Engineer

Vals AI • San Francisco (CA)

On-site
USD 180,000 - 240,000
Health/dental insurance coverage
Relocation support
Lunch and dinner provided
+2
Head of Research
Head of Research

Vibehackers • San Francisco (CA), Northern (KY)

Hybrid
USD 225,000 - 275,000
Relocation and transportation support
Health and dental insurance
Lunch and dinner provided
+6
Member of Technical Staff - Research
Member of Technical Staff - Research

Vibehackers • San Francisco (CA), Northern (KY)

Hybrid
USD 140,000 - 185,000
Relocation assistance
Housing stipend (within 1 mile)
Health and dental insurance
+2
Evaluations Engineer
Evaluations Engineer

Vals AI, Inc. • San Francisco (CA)

On-site
USD 150,000 - 190,000
Relocation and transportation support
Health/dental insurance coverage
Lunch and dinner provided
+2
Member of Technical Staff - Platform
Member of Technical Staff - Platform

Vals AI, Inc. • San Francisco (CA)

On-site
USD 150,000 - 230,000
Relocation assistance
Health insurance
Dental insurance
+1
Head of Research
Head of Research

Vals AI, Inc. • San Francisco (CA)

On-site
USD 220,000 - 340,000
Relocation support
Health insurance
Meals provided (Lunch/Dinner)
+3
Member of Technical Staff - Research
Member of Technical Staff - Research

Vals AI • San Francisco (CA)

On-site
USD 120,000 - 180,000
Relocation support
Health insurance
Lunch and snacks provided
+2
Member of Technical Staff - Platform
Member of Technical Staff - Platform

Vals AI • San Francisco (CA)

On-site
USD 150,000 - 210,000
Relocation support
Health insurance
Dental insurance
+3
Member of Technical Staff - Research
Member of Technical Staff - Research

Vals AI, Inc. • San Francisco (CA)

On-site
USD 150,000 - 260,000
Relocation support
Health insurance
Meals provided
+2