An application made for this job — a tailored resume and cover letter that speak straight to the posting.
Orbifold AI in Palo Alto is seeking a Member of Technical Staff to build data and evaluation foundations for world models and embodied AI systems. This role centers on real-world system performance and scalable data pipelines, not purely theoretical work.
You will design, build, and iterate data workflows, evaluation frameworks, and tooling with robotics and world-model teams, translating failures into concrete data curation and training actions for partner deployments.
Orbifold AI advances the frontier of physical AI and world model companies through rigorous evaluation and curated, real-world data. We work directly with leading robotics and world model research teams on the field's hardest problem: systematically surfacing where today's foundation models break, and producing the curated multimodal data that closes the gap.
Our work sits at the intersection of data, evaluation, and model training. We design evaluation harnesses that expose a model's true failure modes; each failure becomes a structured data deficit; and our curation pipeline produces the targeted, high-quality, fully verified data that fills it — collected against a specific failure mode, sampled to balance the long tail, annotated to a co-defined taxonomy, and verified before it reaches a training or reinforcement learning run. Each cycle compounds: sharper evaluations expose finer failures, finer failures drive more precise curation, and more precise curation narrows the distance between demo and deployment.
We collaborate with partners end to end, co-designing datasets, evaluation frameworks, and training and RL pipelines that shape how their models learn from real-world signals. The data standard and curation framework we're building will define the next frontier of robotics and world model training.
Location: Palo Alto, CA (On-site)
Orbifold AI advances the frontier of physical AI and world model companies through rigorous evaluation and curated, real-world data. We work directly with leading robotics and world model research teams on the field's hardest problem: systematically surfacing where today's foundation models break, and producing the curated multimodal data that closes the gap.
Our work sits at the intersection of data, evaluation, and model training. We design evaluation harnesses that expose a model's true failure modes; each failure becomes a structured data deficit; and our curation pipeline produces the targeted, high-quality, fully verified data that fills it — collected against a specific failure mode, sampled to balance the long tail, annotated to a co-defined taxonomy, and verified before it reaches a training or reinforcement learning run. Each cycle compounds: sharper evaluations expose finer failures, finer failures drive more precise curation, and more precise curation narrows the distance between demo and deployment.
We collaborate with partners end to end, co-designing datasets, evaluation frameworks, and training and RL pipelines that shape how their models learn from real-world signals. The data standard and curation framework we're building will define the next frontier of robotics and world model training.
We are hiring Members of Technical Staff to build the data and evaluation foundations for world models and embodied AI systems . Today's frontier models look impressive on cherry-picked demos but break in production: they fail on long-tail edge cases, hallucinate, lose temporal coherence, mishandle contact and causality, generalize poorly out of distribution, and produce silent failures that automated metrics don't catch. Closing that gap requires evaluation infrastructure that can systematically surface, categorize, and diagnose failures—and feed those signals directly back into data and training.
In this role, you will work closely with internal teams and external research partners to design, build, and iterate on data pipelines, workflows, and evaluation frameworks that drive model quality. You'll define what "good" means for a given partner, build the harnesses that measure it, and translate failures into the next round of data curation and training. This is a highly applied role focused on real-world system performance , not purely theoretical research.