LLM Benchmark Engineer — Elevate Model Evaluation

Vals AI, Inc.

San Francisco (CA)

Presencial

USD 180 000 - 240 000

Tempo integral

14 dias+
Gerador de candidaturas

Recebe uma resposta deste empregador — um currículo e uma carta de apresentação adaptados exatamente ao que estão a contratar.

Ultrapassa os filtros ATS

Vantagens oferecidas por esta oferta de emprego

Relocation support
Health insurance
Lunch provided
401K
Unlimited PTO
Housing stipend

Resumo da oferta

Vals AI, Inc. is seeking engineers to own and expand our leaderboards, evaluating new LLM models against benchmarks across law, tax, coding, finance, and more. You will analyze error modes, compare strengths and weaknesses, and collaborate with a communications team to share results.

The role emphasizes strong Python engineering, sprint-based work, and close collaboration with both open-source and foundation model labs. Relocation and in-person SF work are supported as needed.

Qualificações

  • Familiarity with the space of large language models and current leading models.
  • Strong engineering fundamentals and ability to ship quickly with high quality.
  • Significant Python experience in a professional setting.
  • Experience working in sprints, Git workflows, and code reviews.
  • Willingness to work long hours during model releases and meet tight deadlines.

Responsabilidades

  • Evaluate new LLM model releases across the Vals AI benchmarks.
  • Collaborate with open-source and closed-source labs to assess model performance.
  • Use Docent to analyze failure modes and patterns in model output.
  • Publish findings with our social media team.
  • Add new models and maintain integrations in the model library.
  • Maintain and improve benchmark infrastructure (agentic and non-agentic).

Conhecimentos

Python
LLM knowledge
Engineering fundamentals
Team collaboration
Git & CI

Ferramentas

Docent
Django
AWS
CDK

Descrição da oferta de emprego

Vals AI, Inc. is seeking engineers to own and expand our leaderboards, evaluating new LLM models against benchmarks across law, tax, coding, finance, and more. You will analyze error modes, compare strengths and weaknesses, and collaborate with a communications team to share results.

The role emphasizes strong Python engineering, sprint-based work, and close collaboration with both open-source and foundation model labs. Relocation and in-person SF work are supported as needed.

Obtém a tua avaliação gratuita e confidencial do currículo.

ou arrasta e larga o ficheiro aqui.

Similar jobs

Ofertas semelhantes que vale a pena comparar

Evaluations Engineer
Evaluations Engineer

Vibehackers • San Francisco (CA), Northern (KY)

Presencial
USD 140 000 - 185 000
Relocation support
Health insurance
Lunch and dinner provided
+6
Senior AI Model Evaluation Scientist (LLM Benchmarks)
Senior AI Model Evaluation Scientist (LLM Benchmarks)

Cohere • Seattle (WA)

Presencial
USD 180 000 - 385 000
Lunch stipend
Health and dental benefits
RRSP matching / 401K / Pension
+5
LLM Evaluation Scientist — Benchmarks & Failure Analysis
LLM Evaluation Scientist — Benchmarks & Failure Analysis

Scale AI, Inc. • Seattle (WA)

Presencial
USD 181 000 - 226 000
Health, dental and vision coverage
Retirement benefits
Learning and development stipend
+3
ML Evaluation Scientist - LLM Benchmarks
ML Evaluation Scientist - LLM Benchmarks

Scale AI • Seattle (WA)

Presencial
USD 181 000 - 226 000
Health coverage
Retirement benefits
L&D stipend
+2
LLM Evaluation & Benchmarking Engineer
LLM Evaluation & Benchmarking Engineer

Capitolis • San Francisco (CA)

Presencial
USD 120 000 - 150 000
GenAI Evaluation Scientist: LLM Benchmarks & Diagnostics
GenAI Evaluation Scientist: LLM Benchmarks & Diagnostics

Scale AI • San Francisco (CA)

Presencial
USD 166 000 - 207 000
Health coverage
Equity
Retirement benefits
+3
LLM Benchmarking Research Scientist
LLM Benchmarking Research Scientist

Anyone AI Inc. • Northern (KY)

Híbrido
USD 110 000 - 160 000
GenAI Evaluation Scientist — Benchmark & Diagnose LLMs
GenAI Evaluation Scientist — Benchmark & Diagnose LLMs

Scale AI, Inc. • New York (NY)

Presencial
USD 181 000 - 226 000
Base salary + equity
Health, dental, vision coverage
Retirement benefits
+3
LLM Benchmark Architect — Research Lead
LLM Benchmark Architect — Research Lead

Vibehackers • San Francisco (CA), Northern (KY)

Híbrido
USD 140 000 - 185 000
Relocation assistance
Housing stipend (within 1 mile)
Health and dental insurance
+2
Staff Engineer, LLM Benchmark Platform
Staff Engineer, LLM Benchmark Platform

Vals AI, Inc. • San Francisco (CA)

Presencial
USD 140 000 - 220 000
Relocation support
Health insurance
Lunch/dinner provided
+3