Founding Machine Learning - Eval Layer

One Robot (YC W26)

San Francisco (CA)

On-site

USD 150,000 - 275,000

Full time

48 hours ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

One Robot in San Francisco is hiring a researcher/engineer to advance its evaluation platform for robot manipulation policies. You will focus on training evaluation models, validating policy behavior on real deployments, and building scalable evaluation layers.

You will work with real customer data and drive improvements from data collection to production feedback, shaping how the eval system predicts and explains policy failures across long-horizon tasks.

Qualifications

  • Strong Python and PyTorch coding skills.
  • Experience training VLMs or LLMs.
  • Experience developing and shipping evals for VLMs/LLMs.

Responsibilities

  • Train evaluation models to classify and verify policy behavior.
  • Build confidence layers translating outputs into actionable signals.
  • Improve grounding of models for physical and spatial scenes.
  • Develop data engine to sharpen eval models with deployments.

Skills

Python
PyTorch
VLM training
LLM training
Eval development

Job description

One Robot builds task-specific world models and an evaluation platform for robot manipulation policies.

Training end-to-end policies for robots is vibes-based today. Teams collect data, train, deploy on a real robot, find out what fails, collect more, retry. We replace the trial-and-error with rigorous validation that tells you where your policy will fail and what data to collect to fix it.

Robotics can't industrialize without an evaluation layer. We're building it.

We're solving challenging technical problems around long-horizon autoregressive generation, world model controllability, and closing the sim-to-real gap. We work with real customer data, real failures, and real deployment pressure.

We're based in San Francisco, backed by Accel, YC, several exited founders, and engineering leaders at leading AI companies.

We're small and deliberately so. Everyone is an IC with deep ownership of a wide surface area. The culture is fast iteration and direct responsibility.

Hemanth Sarabu and Elton Shon co-founded One Robot after leading robot learning together at Industrial Next (YC W22), bringing experience from Google, NASA JPL, and Tesla.

We're building the evaluation layer to understand policy failure modes before they hit production. You'll own modeling work that makes the eval trustworthy.

What you’ll do:
  • Train evaluation models: Develop VLMs that classify and verify policy behavior.
  • Build confidence layers: Convert model outputs into trustworthy signals the customer can act on.
  • Improve model grounding: Make the eval models reason accurately about physical and spatial scenes.
  • Build a self-improving eval layer: Develop data engine that makes the eval models sharper with each customer's deployments and corrections.
Requirements:
  • Very strong coding in Python and PyTorch.
  • VLM/LLM training: Track record in training VLMs or LLMs.
  • Evals experience: Developed and shipped evals for VLMs or LLMs.

Compensation Range: $150K - $275K

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Founding Machine Learning - Eval Layer
Founding Machine Learning - Eval Layer

One Robot • San Francisco (CA)

On-site
USD 140,000 - 210,000
Founding Robot Learning
Founding Robot Learning

One Robot (YC W26) • San Francisco (CA)

On-site
USD 150,000 - 275,000
Founding Robot Learning
Founding Robot Learning

One Robot • San Francisco (CA)

On-site
USD 180,000 - 240,000
Machine Learning Intern
Machine Learning Intern

One Robot (YC W26) • San Francisco (CA)

On-site
USD 25,000 - 39,000
Founding ML Evaluation Layer for Robotic Policy Validation
Founding ML Evaluation Layer for Robotic Policy Validation

One Robot • San Francisco (CA)

On-site
USD 140,000 - 210,000
Founding Robot Learning Engineer — Policy Training & Eval
Founding Robot Learning Engineer — Policy Training & Eval

One Robot • San Francisco (CA)

On-site
USD 180,000 - 240,000
Founding ML Engineer
Founding ML Engineer

a16z-speedrun • San Francisco (CA)

On-site
USD 150,000 - 230,000
Founding ML Engineer - Robotic Policy Evaluation
Founding ML Engineer - Robotic Policy Evaluation

One Robot (YC W26) • San Francisco (CA)

On-site
USD 150,000 - 275,000
Research, Post-Training Evals
Research, Post-Training Evals

Thinking Machines Lab Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 140,000 - 200,000
Research, Post-Training Evals
Research, Post-Training Evals

Mosaic.tech • San Francisco (CA)

On-site
USD 140,000 - 190,000