A complete application in a minute — tailored resume and cover letter, ready to send.
Odyssey is an AI lab pioneering general world models. We build large-scale multimodal datasets across video, robotics, and audio to train world models that learn from long-horizon data.
You will design data pipelines, extract signals from raw data, balance training mixes, and collaborate directly with researchers to translate needs into data solutions. We value hands-on data work and curiosity about how data shapes model behavior.
Odyssey https://odyssey.ml is an AI lab pioneering general world models: causal, multimodal systems that learn to predict and interact with the world over long horizons. This foundational technology promises to revolutionize robotics, science, healthcare, education, gaming, defense, and beyond. Odyssey’s founders previously pioneered the most complex application of physical AI: self-driving cars. They’ve now brought together a world-class research team from DeepMind, Tesla, Waymo, Meta, Apple, and Wayve, who have made significant contributions to language models (DeepMind Gemini), video models (DeepMind Veo), world models (Wayve GAIA), and autonomous systems (Tesla FSD). Odyssey has raised significant venture capital from GV, Amazon, AMD, EQT, NVIDIA, Natural Capital, In-Q-Tel, Elad Gil, Jeff Dean, Guillermo Rauch, Garry Tan, Kyle Vogt, and researchers from OpenAI, DeepMind, MSL, Recursive, and Thinking Machines.
Data is fast becoming one of the biggest bottlenecks in building world models. Our models are only as good as the data behind them, and getting that data right is one of the hardest and most important problems we have. We're looking for a data engineer who wants to be the person figuring it out: building and curating the large-scale multimodal datasets our models train on, across video, robotics, and audio. The work spans research and infrastructure. Some days you'll be tuning the platform that processes millions of hours of video and audio. Other days you'll be in the data itself: pulling out signals, improving captions, and working with researchers on what goes into the training mix. We don't treat those as separate jobs, so we want someone comfortable doing both. We care more about how you think about data than about your years of experience or publication record. You might have started in data engineering and moved toward research, or started in research or data science and moved toward engineering. Either way, you like working on data and you've done hands-on work with video and/or audio.
Roughly 1 to 5 years of relevant experience.