Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.
Cantina Labs is building a core team to develop next-generation native video and omni foundation models for multimodal generation and understanding.
You will work across architecture, data, training, evaluation, and deployment, shaping both technical direction and the team from an early stage. Join a collaborative environment focused on creativity and impact.
Cantina Labs is a social AI company, developing a suite of advanced real-time models that push the boundaries of expression, personality, and realism. We bring characters to life, transforming how people tell stories, connect, and create. We build and power ecosystems. Cantina, our flagship social AI platform, is just the beginning.
If you're excited about the potential AI has to shape human creativity and social interactions, join us in building the future!
We are building a core team to develop next-generation native video and omni foundation models for multimodal generation and understanding. Our current focus is large-scale video foundation model development, spanning pre-training, continued training, and post-training for high-quality, controllable, consistent, and efficient generation. Our broader roadmap includes reference- and memory-based generation, multimodal understanding and interaction, and joint audio-video generation.
In this role, you will work on foundational research and large-scale model development across the full model lifecycle, including architecture, data, training, evaluation, post-training, training systems, and efficient inference. You will have the opportunity to shape both the technical direction and the team from an early stage.
The anticipated annual base salary range for this role is between $200,000-$320,000. When determining compensation, a number of factors will be considered, including skills, experience, job scope, location, and competitive compensation market data.