Evals Lead

Build AI

San Francisco (CA)

On-site

USD 170,000 - 250,000

Full time

3 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive pay
Medical, dental, and vision packages
Housing subsidy: $2k per month
Relocation assistance (SF or Shenzhen)
Daily meals in office
Wellness benefits
Unlimited compute budget
Unlimited Codex and Claude credits
Travel support

Job summary

Build AI is hiring an Evals Lead to benchmark video and world-model capabilities with and without our data. You will deeply understand evaluation validity, leakage controls, and whether numbers reflect real capability.

You should have previously pushed evals that shifted training, product, or data decisions. The role requires designing robust benchmarks, ensuring honesty in comparisons, and building analytics to reveal where scaling datasets changes capabilities.

Qualifications

  • Must have shipped or driven evals that mattered and influenced decisions.
  • Understanding eval philosophy and validity beyond plotting a curve.
  • Experience with video, robotics, world models, or multimodal evals preferred.

Responsibilities

  • Design and run benchmarks for video and world models using Build data and external data.
  • Ensure evaluation integrity: hold-out tasks, leakage checks, and no number manipulation.
  • Build analytics: model performance, failure patterns, and impact of dataset scaling.

Job description

About Build AI

Build AI is the data hyperscaler for Physical AI. We're vertically integrated across hardware, manufacturing, logistics, collection, and model training to scale the physical labor dataset orders of magnitude faster than anyone in the world.

Job Summary

We’re hiring an Evals Lead to benchmark video and world-model capabilities with and without our data. You need to deeply understand the philosophy of evals: what an eval is allowed to claim, contamination, leakage, construct validity, and whether a number actually corresponds to a capability. More practically, you should have pushed consequential evals before — evals that changed what a lab trained, shipped, or collected, not a weekend leaderboard.

Key Responsibilities
  • Design and run benchmarks for video and world models, with Build data and without it, under the same protocol

  • Make the comparison honest: held-out tasks, contamination and leakage checks, no cooking the numbers

  • Build the analytics layer: model performance, failure patterns, and whether scaling the dataset moves capability — and on which axes

  • Work with research and dataset/quality so collection and evals inform each other

  • Push evals that are consequential enough that people change plans when the number moves

  • Design processes that increase evaluation quality, repeatability, and scale as we add tasks and countries

You may be a good fit if you have (Must-have qualifications)
  • You have shipped or driven evals that mattered: they changed training, hiring, product, or data decisions

  • You understand eval philosophy well enough to argue about validity, not only to plot a curve

  • Ideally video, robotics, world models, or multimodal, but the bar is consequential evals more than a specific domain

  • You will not confuse a pretty dashboard with an eval that is allowed to decide things

Strong candidates may also have experience with (Nice-to-have qualifications)
  • Video, robotics, world models, or multimodal evals

  • You have designed evals used in a paper, a product launch, or a data decision

  • Background in construct validity, contamination, or leakage

  • Exposure to evaluation operations: throughput, failure taxonomy, repeatability

Benefits
  • Competitive pay

  • Medical, dental, and vision packages with generous premium coverage

  • $500 per month credit for waiving medical benefits

  • Housing subsidy of $2k per month for those living within walking distance of the office

  • Relocation support for those moving to San Francisco (Financial District) or Shenzhen (Nanshan)

  • Various wellness benefits covering fitness, mental health, and more

  • Daily lunch and dinner in our office

  • Unlimited compute budget subject to ROI justification

  • Unlimited Codex and Claude credits

  • Travel

How we're different

Build believes in the Bitter Lesson. By taking a general approach of learning from humans, our addressable market is all physical labor.

We are a fully in-person team in San Francisco (Financial District) and Shenzhen (Nanshan), and greatly value engineering skills. We do not have boundaries between engineering and research, and we expect all of our technical staff to contribute to both and work across disciplines as needed.

Build AI is an equal opportunity employer. We review every application. Questions: research@build.ai

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Product Lead
Product Lead

Build AI • San Francisco (CA)

On-site
USD 180,000 - 240,000
Competitive pay
Medical, dental, and vision packages
Housing subsidy for SF/Shenzhen
+5
Software Engineer, Data Visualization
Software Engineer, Data Visualization

Build AI • San Francisco (CA)

On-site
USD 120,000 - 210,000
Medical, dental, and vision packages
Housing subsidy
Relocation support
+4
Head of Global Expansion
Head of Global Expansion

Build AI • San Francisco (CA)

On-site
USD 180,000 - 280,000
Competitive pay
Health coverage
Medical premium credit
+7
Member of Technical Staff
Member of Technical Staff

Build AI • San Francisco (CA)

On-site
USD 200,000 - 280,000
Competitive pay
Medical coverage
Waiver credit $500
+7
Software Engineer, Generalist
Software Engineer, Generalist

Build AI • San Francisco (CA)

On-site
USD 140,000 - 210,000
Competitive pay
Medical, dental, and vision packages
Housing subsidy near office
+3
Software Engineer, Data Infrastructure & Pipelining
Software Engineer, Data Infrastructure & Pipelining

Build AI • San Francisco (CA)

On-site
USD 170,000 - 290,000
Competitive pay
Medical, dental, and vision packages
Housing subsidy
+6
Software Engineer, Scaling
Software Engineer, Scaling

Build AI • San Francisco (CA)

On-site
USD 150,000 - 180,000
Competitive pay
Health insurance (medical/dental/vison
Medical credit $500/mo
+7
Evals Infrastructure Tech Lead / Manager
Evals Infrastructure Tech Lead / Manager

United States Digital Space LLC • San Francisco (CA)

Hybrid
USD 500,000 - 850,000
Software Engineer, Generalist
Software Engineer, Generalist

Worky • San Francisco (CA)

On-site
USD 135,000 - 210,000
Competitive pay
Health coverage
Medical credit for waiving benefits
+6
Talent Recruiting
Talent Recruiting

Build AI • San Francisco (CA)

On-site
USD 110,000 - 160,000
Medical, dental, and vision packages
Housing subsidy of $2k per month for I
Relocation support to SF/Shenzhen
+2