A complete application in a minute — tailored resume and cover letter, ready to send.
Luma in California seeks a researcher to turn its generative video models into world models: interactive, controllable, physically faithful, and useful for embodied reasoning. You will invent architectures (diffusion/transformer/AR hybrids), develop action/view conditioning, define metrics for fidelity and coherence, run scaling studies, and publish toward an open-source release.
This role targets PhD-level expertise in ML, CV, or robotics with strong PyTorch skills and a track record of
You'll turn Luma's industry-leading generative video models into world models: interactive, controllable, physically faithful, and useful as a substrate for embodied reasoning. This is the role at the center of the thesis.
You'll invent next-generation world-model architectures and the controllability that lets an agent step into a generated world, and own the metrics that define success. It fits a researcher with deep generative-modeling or model-based-RL expertise who has trained models to the limits of a multi-node cluster. If you want a narrow, well-scoped research problem, this is broader and more open-ended than that.
One way the first 90 could unfold.
About Luma: Luma's mission is to build unified general intelligence that can generate, understand, and operate in the physical world. We believe multimodality is critical for intelligence — the next step beyond language models comes from vision. Luma is an equal opportunity employer.