Directly involves LLM benchmarking and leaderboards; linked to vibe-coding benchmark coverage and rapid sprints around model releases.
About the Role
Join Vals AI to own and operate leaderboards that evaluate LLMs: test new model releases against benchmarks, analyze error modes, and publish results used by startups, enterprises, and research labs. Work with foundation model labs, maintain model integrations and benchmarking infrastructure, and collaborate with communications to share findings.
Job Description
Role
You will own the leaderboards on Vals AI by evaluating new LLM releases across a suite of benchmarks, analyzing failure modes and model strengths, maintaining model integrations, and helping to improve the infrastructure used to run benchmarks.
Key Responsibilities
- Evaluate new LLM model releases across Vals AI benchmarks covering tasks such as law, tax, coding, finance, and social mobility.
- Work directly with open-source and closed-source foundation model labs to assess model performance.
- Use tools like Docent to analyze common failure modes and performance patterns.
- Add new models and maintain integrations in the model library.
- Maintain and improve benchmarking infrastructure (agentic and non-agentic workloads).
- Collaborate with the communications/social media team to publish and post results.
- Operate on a release-driven cadence: expect intensive sprints after major model launches and quieter periods between releases.
Requirements
- Familiarity with the LLM space, including knowledge of leading models and their relative strengths.
- Strong engineering fundamentals and a track record of building and shipping significant projects.
- Significant professional experience with Python.
- Experience with development sprints, Git workflows, and pull request reviews.
- Willingness to work long hours during model releases to deliver high-quality results under tight deadlines.
Nice-to-Haves
- Previous experience benchmarking large language models or creating evaluation benchmarks.
- Startup experience or experience founding a company.
- Technical writing ability.
- Machine learning research experience.
What We Offer
- Highly competitive salary and meaningful ownership.
- Relocation and transportation support.
- Health and dental insurance coverage.
- Lunch and dinner provided; free snacks, coffee, and drinks.
- 401(k) plan.
- Unlimited PTO.
- $1,500 housing stipend (within one-mile radius).
- Django backend and React frontend.
- Infrastructure on AWS using CDK for infrastructure-as-code.
- Use of Docent for failure-mode analysis.
Location
- In-person team based in San Francisco; relocation or transportation support available.
Skills
LLM evaluation Benchmarking Error analysis Engineering fundamentals Collaboration Git workflows Pull request reviews Technical communication Technical writing Ownership Learning velocity Time management under sprint conditions Solution-oriented problem solving
Experience Level
Senior
USD 140,000 - 185,000/year
Employment Type
Full-time
- Relocation and transportation support
- Lunch and dinner provided
- Free snacks/coffee/drinks
- $1,500 housing stipend (within one-mile radius)
- Highly competitive salary and meaningful ownership