Stand out for this role — generate a tailored resume and cover letter in about a minute.
Unknown organization in the United States is seeking a data engineer to curate large-scale multimodal datasets, including video, audio, and robotics content, to train flagship models.
You will work with researchers to optimize captioning, quality scoring, deduplication, and dataset versioning, while tuning infrastructure and storage at petabyte scale.
Our client builds world model systems that learn how the world actually behaves over time and generate it back, interactively, in real time. Seven model releases since December 2024, including the first real-time synchronised audio-and-video world model and a four-player shared world.
Data is fast becoming one of the biggest bottlenecks in building world models. Models are only as good as the data behind them, and getting that data right is one of the hardest and most important problems any research lab has.
Some days you're tuning infrastructure. Some days you're sitting with a researcher, working out why the model learned the wrong thing.
We care more about how you think about data than about your years of experience or your publication record. Data engineering or research background - both work.
you've never touched video, but you've run petabyte-to-exabyte data infrastructure where selection, layout, throughput and cost were the whole problem. That skill transfers. They've hired against it before.
There is no incumbent data platform and no VP of Engineering in post. You'd define versioning, lineage and the quality bar rather than inherit someone else's from three years ago. The scope doesn't usually fit in one role: video, audio, robotics, synthetic data from our adversarial RL work, and licensed corpora. At a larger lab, that's five teams. You'd own a slice of one. And we name every contributor on our model release pages, including infrastructure and data people.