Research Engineer, Policy Evaluation

Bonfirevc

Palo Alto (CA)

On-site

USD 150,000 - 190,000

Full time

3 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Orbifold AI is seeking a Research Engineer to design evaluation methodologies for embodied AI policies and to run robust, real-world tests. You will work with partner teams to translate failures into data specs and drive the next data collection cycle.

You will build scalable evaluation infrastructure, create fine-grained failure taxonomies, and validate policies on real hardware to demonstrate concrete improvements rather than isolated benchmarks.

Qualifications

  • PhD or equivalent research experience in robot learning or closely related field.
  • Real policies trained and debugged on real robots; ability to distinguish failing policy from miscalibrated robot.
  • Strong PyTorch proficiency; comfortable at scale on Ray; able to work in a codebase not owned by you.

Responsibilities

  • Design evaluation methodology end to end: task suites, held-out slices, scoring, and statistics.
  • Build fine-grained failure taxonomies: occlusion, deformables, long-horizon, distribution shift, edge cases.
  • Train and fine-tune reference policies on curated data to show data moves metrics.

Skills

PhD or equivalent
Real robot policies
PyTorch
Ray
Self-driven

Education

PhD in robotics or ML

Tools

Ray
OpenVLA

Job description

Research Engineer, policy Evaluation

Palo Alto, CA (On-site)

About Orbifold AI

Orbifold AI is building the infrastructure layer for Physical AI. As intelligent systems move beyond language into the physical world, they require a fundamentally new understanding of physics, action, and interaction.

We partner with leading robotics and world model research teams to advance the foundations of embodied intelligence, enabling intelligent systems to perceive, understand, and operate in the real world.

The standards we set, and the infrastructure we build to scale them, will define the next frontier of robotics and Physical AI.

Role Overview

Language models converged in part because everyone agreed how to measure progress. Embodied AI has no equivalent. A robot policy today is graded on a few dozen real-world trials and operator intuition, because nothing can run and check a physical policy a million times. Half a dozen well-funded architectural bets are running in parallel with no shared way to settle which is working.

You will build the thing the field is missing. You take a partner’s trained policy, run it against thousands of real scenes drawn from our corpus, and return not a score but a ranked, repeatable list of the scenes, objects and grasps that break it. Then you close the loop: every failure you characterize becomes a data specification, and the data we collect against it becomes the next evaluation.

This is highly applied work. You will be measured on whether a partner’s real‑world policy improved, not on a benchmark number in isolation.

What You Will Work On
  • Design evaluation methodology end to end: task suites, held‑out slices, scoring, and statistical rigor that survives a skeptical research lead’s scrutiny.
  • Build fine‑grained failure taxonomies: occlusion, transparency, deformables, long‑horizon, contact‑rich, distribution shift, and the edge‑case discovery and long‑tail probing that populate them.
  • Train and fine‑tune reference policies (π0‑class, OpenVLA, ACT, diffusion policy) on our curated data, so we can demonstrate rather than assert that a dataset moves a metric.
  • Build automated critics and judges that approximate human evaluation at scale and correlate with downstream model behavior, the only way this works at corpus scale.
  • Run evaluations on real hardware where it matters, and know precisely when a real‑robot trial is worth its cost versus a corpus replay.
  • Work directly with partner research teams: reproduce their setup, run their model against our data, and be the person who tells them the truth about what you find.
  • Close the loop: translate every evaluation finding into a concrete collection specification, and back again.
What We Are Looking For
  • PhD or equivalent research experience in robot learning, imitation learning, or a closely related field, with first‑author publications or shipped work.
  • You have trained and debugged real policies on real robots, and you know the difference between a policy that is failing and a robot that is miscalibrated.
  • Real opinions about how embodied models should be evaluated, and the statistical literacy to defend them.
  • Strong PyTorch, comfortable at scale on Ray, and fast inside a codebase that is not yours.
  • Self‑driven and high agency, with experience in fast‑paced applied research or startup environments.
  • We index on the quality of the work rather than years served. A recent PhD with strong first‑author publications in a directly relevant area is exactly who we want to talk to.
Nice to Have
  • You have authored an evaluation suite or benchmark that other people use.
  • Multi‑embodiment work: humanoids, bimanual, dexterous hands.
  • Experience training or evaluating multimodal LLMs as critics, judges, or reward models.
  • Published work on generalization, scaling laws, or data quality in robot learning.
  • Teleoperation stacks and real‑robot infrastructure.
  • Frontier lab, autonomous vehicle program, or humanoid company.
Why This Role
  • You sit between frontier labs and the ground truth. Very few people get to run several of the best teams’ models against the same carefully controlled data and see which claims survive.
  • Evaluation is the leverage point. In a field with a demo‑to‑deployment gap this wide, whoever can measure honestly sets the agenda.
  • Your findings become collection programs across multiple partners, the failure modes you characterize are what the industry goes and captures next.
  • Founding scope. The evaluation function at Orbifold is what you decide it is.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Research Engineer, Policy Evaluation
Research Engineer, Policy Evaluation

Orbifold AI, Inc. • Palo Alto (CA), Northern (KY)

Hybrid
USD 180,000 - 230,000
Research Engineer, Physical AI (Robotics, World Models)
Research Engineer, Physical AI (Robotics, World Models)

Orbifold AI • Palo Alto (CA)

On-site
USD 180,000 - 240,000
Member of Technical Staff, Physical AI (Robotics / World Models)
Member of Technical Staff, Physical AI (Robotics / World Models)

Bonfirevc • Palo Alto (CA)

On-site
USD 180,000 - 240,000
Research Engineer, Robot Learning
Research Engineer, Robot Learning

Bonfirevc • Palo Alto (CA)

On-site
USD 150,000 - 230,000
Research Engineer, World Models & Simulation
Research Engineer, World Models & Simulation

Bonfirevc • Palo Alto (CA)

On-site
USD 170,000 - 230,000
Research Engineer, World Models & Simulation
Research Engineer, World Models & Simulation

Orbifold AI, Inc. • Palo Alto (CA)

On-site
USD 120,000 - 180,000
Research Engineer, 3D Perception
Research Engineer, 3D Perception

Bonfirevc • Palo Alto (CA)

On-site
USD 120,000 - 180,000
Embodied AI Policy Evaluation Engineer
Embodied AI Policy Evaluation Engineer

Orbifold AI, Inc. • Palo Alto (CA), Northern (KY)

Hybrid
USD 180,000 - 230,000
Embodied AI Evaluation Engineer
Embodied AI Evaluation Engineer

Bonfirevc • Palo Alto (CA)

On-site
USD 150,000 - 190,000
Founding Machine Learning - Eval Layer
Founding Machine Learning - Eval Layer

One Robot (YC W26) • San Francisco (CA)

On-site
USD 150,000 - 275,000