Research Engineer, Policy Evaluation

Orbifold AI, Inc.

Palo Alto, Northern (CA, KY)

Hybrid

USD 180,000 - 230,000

Full time

5 hours ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Orbifold AI, Inc. is building the infrastructure layer for Physical AI. This role focuses on evaluating robot policies by running trained policies against thousands of real scenes and generating repeatable failure taxonomies.

You will collaborate with partner research teams, translate evaluation findings into concrete data specifications, and help design end-to-end evaluation methodologies that scale to corpus-level testing.

Qualifications

  • PhD or equivalent research background with first-author work.
  • Experience training and debugging real policies on robots.
  • Strong statistical literacy and evaluation design.
  • Proficiency with PyTorch and Ray at scale.

Responsibilities

  • Design end-to-end evaluation methodology.
  • Build fine-grained failure taxonomies.
  • Train and fine-tune reference policies.
  • Build automated critics and judges.
  • Run evaluations on real hardware.
  • Collaborate with partner research teams.
  • Translate findings into concrete data specs.

Skills

PhD or equivalent
Robot learning
PyTorch
Ray
First-author pubs
Self-driven

Education

PhD or equivalent

Tools

Ray
OpenVLA

Job description

Design the harnesses that tell a frontier lab where its robot policy actually breaks. Robot learning or evaluation research background; PyTorch and Ray.

About Orbifold AI

Orbifold AI is building the infrastructure layer for Physical AI. As intelligent systems move beyond language into the physical world, they require a fundamentally new understanding of physics, action, and interaction.

We partner with leading robotics and world model research teams to advance the foundations of embodied intelligence, enabling intelligent systems to perceive, understand, and operate in the real world.

The standards we set, and the infrastructure we build to scale them, will define the next frontier of robotics and Physical AI.

Role Overview

Language models converged in part because everyone agreed how to measure progress. Embodied AI has no equivalent. A robot policy today is graded on a few dozen real-world trials and operator intuition, because nothing can run and check a physical policy a million times. Half a dozen well-funded architectural bets are running in parallel with no shared way to settle which is working.

You will build the thing the field is missing. You take a partner’s trained policy, run it against thousands of real scenes drawn from our corpus, and return not a score but a ranked, repeatable list of the scenes, objects and grasps that break it. Then you close the loop: every failure you characterize becomes a data specification, and the data we collect against it becomes the next evaluation.

This is highly applied work. You will be measured on whether a partner’s real-world policy improved, not on a benchmark number in isolation.

What You Will Work On
  • Design evaluation methodology end to end: task suites, held-out slices, scoring, and statistical rigor that survives a skeptical research lead’s scrutiny.
  • Build fine-grained failure taxonomies: occlusion, transparency, deformables, long-horizon, contact-rich, distribution shift, and the edge-case discovery and long-tail probing that populate them.
  • Train and fine-tune reference policies (π0-class, OpenVLA, ACT, diffusion policy) on our curated data, so we can demonstrate rather than assert that a dataset moves a metric.
  • Build automated critics and judges that approximate human evaluation at scale and correlate with downstream model behavior, the only way this works at corpus scale.
  • Run evaluations on real hardware where it matters, and know precisely when a real-robot trial is worth its cost versus a corpus replay.
  • Work directly with partner research teams: reproduce their setup, run their model against our data, and be the person who tells them the truth about what you find.
  • Close the loop: translate every evaluation finding into a concrete collection specification, and back again.
What We Are Looking For
  • PhD or equivalent research experience in robot learning, imitation learning, or a closely related field, with first-author publications or shipped work.
  • You have trained and debugged real policies on real robots, and you know the difference between a policy that is failing and a robot that is miscalibrated.
  • Real opinions about how embodied models should be evaluated, and the statistical literacy to defend them.
  • Strong PyTorch, comfortable at scale on Ray, and fast inside a codebase that is not yours.
  • Self-driven and high agency, with experience in fast-paced applied research or startup environments.
  • We index on the quality of the work rather than years served. A recent PhD with strong first-author publications in a directly relevant area is exactly who we want to talk to.
Nice to Have
  • You have authored an evaluation suite or benchmark that other people use.
  • Multi-embodiment work: humanoids, bimanual, dexterous hands.
  • Experience training or evaluating multimodal LLMs as critics, judges, or reward models.
  • Published work on generalization, scaling laws, or data quality in robot learning.
  • Teleoperation stacks and real-robot infrastructure.
  • Frontier lab, autonomous vehicle program, or humanoid company.
Why This Role
  • You sit between frontier labs and the ground truth. Very few people get to run several of the best teams’ models against the same carefully controlled data and see which claims survive.
  • Evaluation is the leverage point. In a field with a demo-to-deployment gap this wide, whoever can measure honestly sets the agenda.
  • Your findings become collection programs across multiple partners, the failure modes you characterize are what the industry goes and captures next.
  • Founding scope. The evaluation function at Orbifold is what you decide it is.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Research Engineer, Policy Evaluation
Research Engineer, Policy Evaluation

Bonfirevc • Palo Alto (CA)

On-site
USD 150,000 - 190,000
Research Engineer, Physical AI (Robotics, World Models)
Research Engineer, Physical AI (Robotics, World Models)

Orbifold AI • Palo Alto (CA)

On-site
USD 180,000 - 240,000
Member of Technical Staff, Physical AI (Robotics / World Models)
Member of Technical Staff, Physical AI (Robotics / World Models)

Bonfirevc • Palo Alto (CA)

On-site
USD 180,000 - 240,000
Research Engineer, Robot Learning
Research Engineer, Robot Learning

Bonfirevc • Palo Alto (CA)

On-site
USD 150,000 - 230,000
Research Engineer, World Models & Simulation
Research Engineer, World Models & Simulation

Orbifold AI, Inc. • Palo Alto (CA)

On-site
USD 120,000 - 180,000
Research Engineer, World Models & Simulation
Research Engineer, World Models & Simulation

Bonfirevc • Palo Alto (CA)

On-site
USD 170,000 - 230,000
Research Engineer, 3D Perception
Research Engineer, 3D Perception

Bonfirevc • Palo Alto (CA)

On-site
USD 120,000 - 180,000
Founding Robot Learning
Founding Robot Learning

One Robot • San Francisco (CA)

On-site
USD 180,000 - 240,000
Robotics Researcher — Manipulation (Omakase Zen)
Robotics Researcher — Manipulation (Omakase Zen)

Omakase Robotics • United States

Hybrid
USD 140,000 - 210,000
Member of Technical Staff, ML Engineer (Applied AI Infrastructure)
Member of Technical Staff, ML Engineer (Applied AI Infrastructure)

Bonfirevc • Palo Alto (CA)

On-site
USD 180,000 - 240,000