Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.
XDOF is hiring a Research Engineer / Scientist to lead systems that convert raw egocentric and teleoperation video into high-signal training data for vision-language models and robot learning.
You will design scalable VL pipelines, advance data curation, and collaborate with partner labs to align data quality with policy performance, shaping the future of robotics and AI.
Requires MS or PhD and 3–7 years of relevant experience; deep PyTorch expertise; work across perception, language, and action.
Frontier labs are racing to build general-purpose robots, and the bottleneck isn't compute. It's data. At XDOF, we're building the foundation behind the foundation models: the data collection systems, annotation pipelines, exabyte-scale data infrastructure, and software toolchain that enable our partners to push the field forward.
We're hiring a Research Engineer / Scientist to help lead technical efforts at the intersection of vision-language models and robot learning. You will build systems that turn raw egocentric and teleoperation video into high-signal training data for VLA models, and increasingly, contribute to the models themselves.
Beyond pipelines, you will drive research into what makes robot data useful: discovering new metadata (contact events, affordance labels, implicit reward signals, dynamics priors from video) that unlock capabilities current approaches miss. You'll explore how structured annotations can improve cross-embodiment transfer, automatic curriculum generation, and world models that predict what actually matters for manipulation. The data layer isn't downstream of the research. It is the research.
What You'll Do
Required: