Research Scientist, Video Foundation Models

Cantina

San Francisco (CA)

On-site

USD 200,000 - 320,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Competitive salary
Company equity
Medical insurance
Dental insurance
Vision insurance
Paid time off
Parental leave
Fertility support
401(k) plan
Lunch provided
One Medical membership

Job summary

Cantina Labs is building a core team to advance native video and multimodal foundation models for scalable generation and understanding. You will shape architecture, data, training, and evaluation across the full model lifecycle while collaborating with researchers, engineers, and product teams to drive impactful, real-world capabilities.

The role focuses on pre-training, continued training, and post-training for high-quality, controllable generation with opportunities for publications and

Qualifications

  • Experience with diffusion models, flow matching, DiTs, or related generative methods.
  • Hands-on experience training large-scale image/video or multimodal models.
  • Proven track record through publications, open-source contributions, or production impact.
  • Ability to own ambiguous research problems and collaborate in a team.
  • Specialized depth across foundation model lifecycle: architecture, data, post-training, deployment.

Responsibilities

  • Research, develop, and scale native video and multimodal foundation models.
  • Explore architectures, training objectives, and conditioning for video generation.
  • Build data curation, distributed training, evaluation pipelines for quality control.
  • Design experiments to understand scaling, generation quality, controllability, and efficiency.
  • Collaborate with researchers, engineers, and product teams to shape roadmaps.
  • Contribute to research publications and open-source releases when appropriate.

Skills

Generative modeling
Diffusion models
Video generation
Multimodal models
Distributed training
Research publications

Job description

About Cantina:

Cantina Labs is a social AI company, developing a suite of advanced real-time models that push the boundaries of expression, personality, and realism. We bring characters to life, transforming how people tell stories, connect, and create. We build and power ecosystems. Cantina, our flagship social AI platform, is just the beginning.

If you’re excited about the potential AI has to shape human creativity and social interactions, join us in building the future!

About the Role:

We are building a core team to develop next-generation native video and omni foundation models for multimodal generation and understanding. Our current focus is large-scale video foundation model development, spanning pre-training, continued training, and post-training for high-quality, controllable, consistent, and efficient generation. Our broader roadmap includes reference- and memory-based generation, multimodal understanding and interaction, and joint audio-video generation.

In this role, you will work on foundational research and large-scale model development across the full model lifecycle, including architecture, data, training, evaluation, post-training, training systems, and efficient inference. You will have the opportunity to shape both the technical direction and the team from an early stage.

What You’ll Do:
  • Research, develop, and scale native video and multimodal foundation models, from early prototypes through large-scale pre-training, continued training, and post-training.

  • Explore new model architectures, training objectives, and conditioning mechanisms for video generation, reference- and memory-based generation, multimodal interaction, and joint audio-video generation.

  • Build and improve large-scale data curation, distributed training, evaluation, and post-training pipelines for high-quality and controllable generation.

  • Design systematic experiments to understand model scaling, generation quality, controllability, consistency, robustness, and inference efficiency.

  • Collaborate closely with researchers, engineers, and product teams to help shape the technical roadmap and, where appropriate, translate model advances into real-world capabilities.

  • Contribute to research publications and open-source releases when appropriate.

What You’ll Bring:
  • Strong research and engineering experience in generative modeling, including areas such as diffusion models, flow matching, DiTs, video generation, multimodal models, world models, or related fields.

  • Hands‑on experience training and evaluating large-scale image, video, or unified multimodal models using modern deep learning frameworks and distributed training systems.

  • A strong track record of developing impactful models or systems, demonstrated through research publications, open-source contributions, production impact, or other significant technical work.

  • Ability to independently own ambiguous research problems, move effectively from ideas to experiments, and work well in a highly collaborative environment.

  • Specialized depth in one or more areas across the foundation model lifecycle, such as model architecture, data curation, controllable generation, multimodal understanding and conditioning, post‑training and reward modeling, model acceleration, inference systems, or deployment.

Compensation:

The anticipated annual base salary range for this role is between $200,000-$320,000. When determining compensation, a number of factors will be considered, including skills, experience, job scope, location, and competitive compensation market data.

Benefits:
  • Competitive salary and generous company equity

  • Medical, dental, and vision insurance – 99.99% of premiums covered by Cantina

  • 42 days of paid time off, including:

  • Generous parental leave & fertility support

  • 401(k) retirement savings plan

  • Lifestyle spending account – $500/month to use however you’d like

  • Complimentary lunch and snacks for in-office employees

  • One Medical membership, and more!

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Research Scientist, Video Foundation Models
Research Scientist, Video Foundation Models

Cantina • California (MO)

On-site
USD 200,000 - 320,000
Competitive salary
Equity
Medical / Dental / Vision
+6
Research Scientist, Video Foundation Models
Research Scientist, Video Foundation Models

Cantina Labs • California (MO)

On-site
USD 200,000 - 320,000
Competitive salary and equity
Medical, dental, and vision insurance
42 days of PTO and holidays
+4
Product Director, Video Products
Product Director, Video Products

Cantina, Inc. • California (MO)

On-site
USD 200,000 - 260,000
Medical, dental, and vision insurance
Generous PTO and company holidays
401(k) retirement savings plan
+2
Research Scientist, Multimodal Video AI
Research Scientist, Multimodal Video AI

Cantina • California (MO)

On-site
USD 200,000 - 320,000
Competitive salary
Equity
Medical / Dental / Vision
+6
Video Foundation Models Scientist — Multimodal AI
Video Foundation Models Scientist — Multimodal AI

Cantina • San Francisco (CA)

On-site
USD 200,000 - 320,000
Competitive salary
Company equity
Medical insurance
+8
Staff Software Engineer, Bots
Staff Software Engineer, Bots

Cantina Labs • Los Angeles (CA)

On-site
USD 230,000 - 290,000
Medical, dental, and vision insurance
42 days of paid time off
Generous parental leave & fertility support
+4
Android Engineer (Senior-Staff Levels)
Android Engineer (Senior-Staff Levels)

Cantina • San Francisco (CA), Los Angeles (CA)

On-site
USD 180,000 - 240,000
Medical, dental, and vision insurance – 99.99% of premiums covered
42 days of paid time off
Generous parental leave & fertility support
+2
iOS Engineer (Senior-Staff Levels)
iOS Engineer (Senior-Staff Levels)

Cantina • California (MO)

On-site
USD 180,000 - 230,000
Medical, dental, and vision insurance
Generous parental leave
401(k) retirement savings plan
+2
Senior Staff Software Engineer, Search & Recommendation
Senior Staff Software Engineer, Search & Recommendation

Cantina • United States

Remote
USD 250,000 - 320,000
Competitive salary
Company equity
Health insurance
+6
Research Scientist, World Models Graduate (Intelligent Creation) - Global Frontier Tech Recruit[...]
Research Scientist, World Models Graduate (Intelligent Creation) - Global Frontier Tech Recruit[...]

TikTok • San Jose (CA)

On-site
USD 244,800 - 588,000
Medical, dental, and vision insurance
401(k) savings plan with company match
Paid parental leave
+5