Member of Technical Staff, Evals Lead

Build AI

San Francisco (CA)

On-site

USD 180,000 - 240,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Competitive pay
Medical package
Dental package
Vision package
Housing subsidy
Relocation support
Wellness program
Daily meals in office
Compute budget access
Codex and Claude credits
Travel opportunities

Job summary

Build AI in San Francisco seeks an Evals Lead to benchmark video and world-model capabilities with and without our data. You must deeply understand evals: what an eval is allowed to claim, contamination, leakage, construct validity, and whether a number actually corresponds to a capability.

More practically, you should have pushed consequential evals before — evals that changed what a lab trained, shipped, or collected, not a weekend leaderboard.

Qualifications

  • You have shipped or driven evals that mattered: they changed training, hiring, product, or data decisions.
  • You understand eval philosophy well enough to argue about validity, not only to plot a curve.
  • Ideally video, robotics, world models, or multimodal, but the bar is consequential evals more than a specific domain.
  • You will not confuse a pretty dashboard with an eval that is allowed to decide things.

Responsibilities

  • Design and run benchmarks for video and world models, with Build data and without it, under the same protocol
  • Make the comparison honest: held-out tasks, contamination and leakage checks, no cooking the numbers
  • Build the analytics layer: model performance, failure patterns, and whether scaling the dataset moves capability — and on which axes
  • Work with research and dataset/quality so collection and evals inform each other
  • Push evals that are consequential enough that people change plans when the number moves
  • Design processes that increase evaluation quality, repeatability, and scale as we add tasks and countries

Skills

Eval leadership
Eval philosophy
Benchmark design
Data-driven decisions

Job description

About Build AI

Build AI is the data hyperscaler for Physical AI. We're vertically integrated across hardware, manufacturing, logistics, collection, and model training to scale the physical labor dataset orders of magnitude faster than anyone in the world.

Job Summary

We’re hiring an Evals Lead to benchmark video and world-model capabilities with and without our data. You need to deeply understand the philosophy of evals: what an eval is allowed to claim, contamination, leakage, construct validity, and whether a number actually corresponds to a capability. More practically, you should have pushed consequential evals before — evals that changed what a lab trained, shipped, or collected, not a weekend leaderboard.

Key Responsibilities
  • Design and run benchmarks for video and world models, with Build data and without it, under the same protocol

  • Make the comparison honest: held-out tasks, contamination and leakage checks, no cooking the numbers

  • Build the analytics layer: model performance, failure patterns, and whether scaling the dataset moves capability — and on which axes

  • Work with research and dataset/quality so collection and evals inform each other

  • Push evals that are consequential enough that people change plans when the number moves

  • Design processes that increase evaluation quality, repeatability, and scale as we add tasks and countries

You may be a good fit if you have (Must-have qualifications)
  • You have shipped or driven evals that mattered: they changed training, hiring, product, or data decisions

  • You understand eval philosophy well enough to argue about validity, not only to plot a curve

  • Ideally video, robotics, world models, or multimodal, but the bar is consequential evals more than a specific domain

  • You will not confuse a pretty dashboard with an eval that is allowed to decide things

Strong candidates may also have experience with (Nice-to-have qualifications)
  • Video, robotics, world models, or multimodal evals

  • You have designed evals used in a paper, a product launch, or a data decision

  • Background in construct validity, contamination, or leakage

  • Exposure to evaluation operations: throughput, failure taxonomy, repeatability

Benefits
  • Competitive pay

  • Medical, dental, and vision packages with generous premium coverage

  • $500 per month credit for waiving medical benefits

  • Housing subsidy of $2k per month for those living within walking distance of the office

  • Relocation support for those moving to San Francisco (Financial District) or Shenzhen (Nanshan)

  • Various wellness benefits covering fitness, mental health, and more

  • Daily lunch and dinner in our office

  • Unlimited compute budget subject to ROI justification

  • Unlimited Codex and Claude credits

  • Travel

How we're different

Build believes in the Bitter Lesson. By taking a general approach of learning from humans, our addressable market is all physical labor.

We are a fully in-person team in San Francisco (Financial District) and Shenzhen (Nanshan), and greatly value engineering skills. We do not have boundaries between engineering and research, and we expect all of our technical staff to contribute to both and work across disciplines as needed.

Build AI is an equal opportunity employer. We review every application. Questions: research@build.ai

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Member of Technical Staff
Member of Technical Staff

Build AI • San Francisco (CA)

On-site
USD 230,000 - 260,000
Competitive pay
Medical, dental, and vision packages
Housing subsidy
+6
Software Engineer, Data Visualization
Software Engineer, Data Visualization

Build AI • San Francisco (CA)

On-site
USD 140,000 - 210,000
Competitive pay
Medical, dental, and vision packages
Housing subsidy
+6
Software Engineer, Data Infrastructure & Pipelining
Software Engineer, Data Infrastructure & Pipelining

Build AI • San Francisco (CA)

On-site
USD 170,000 - 290,000
Competitive pay
Medical, dental, and vision packages
Housing subsidy
+6
Lead Engineer, Data Platform
Lead Engineer, Data Platform

Build AI • San Francisco (CA)

On-site
USD 180,000 - 240,000
Competitive pay
Medical, dental, and vision packages
Housing subsidy for SF or Shenzhen
Product Lead
Product Lead

Build AI • San Francisco (CA)

On-site
USD 180,000 - 260,000
Medical,Dental,Vision coverage
Housing subsidy
Relocation support
+4
Software Engineer, Scaling
Software Engineer, Scaling

Build AI • San Francisco (CA)

On-site
USD 180,000 - 230,000
Competitive pay
Medical, dental, and vision packages
Housing subsidy for SF housing near HQ
+6
Head of Global Expansion
Head of Global Expansion

Build AI • San Francisco (CA)

On-site
USD 150,000 - 210,000
Competitive pay
Medical coverage
Dental coverage
+8
Country Lead
Country Lead

Build AI • San Francisco (CA)

On-site
USD 150,000 - 230,000
Competitive pay
Medical coverage
Housing subsidy
+6
Lead Engineer, ML Data Infrastructure & Systems
Lead Engineer, ML Data Infrastructure & Systems

Build AI • San Francisco (CA)

On-site
USD 180,000 - 240,000
Competitive pay
Medical, dental, and vision packages
Housing subsidy for SF office nearby
+6
Software Engineer, Generalist
Software Engineer, Generalist

Worky • San Francisco (CA)

On-site
USD 135,000 - 210,000
Competitive pay
Health coverage
Medical credit for waiving benefits
+6