AI Research Scientist, Multimodal: Video & Audio Generation

Meta

Menlo Park (CA)

On-site

USD 180,000 - 280,000

Full time

11 days ago
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Meta's Monetization GenAI Video team is building the next generation of generative AI for video—creating, understanding, and transforming at scale to reach billions of users. We develop foundation models for video generation, video-to-video editing, audio generation, and multimodal video understanding powered by large language models.

The team seeks AI Research Scientists to define the research vision across video and audio generation, drive breakthroughs, and shape the technical direction of

Qualifications

  • Bachelor's degree in CS/engineering or related field.
  • PhD in CS/AI/ML or related field.
  • Experience as technical lead on a team and/or managing complex projects end-to-end.
  • 2+ years of training large language and/or vision models with multimodal LLMs.
  • Research expertise in video generation/understanding, multimodal learning, or diffusion models.
  • Notable industry research contributions with top conferences (e.g., ACL, NeurIPS, ICML, ICLR, AAAI, CVPR).
  • Experience with frontier-quality/state-of-the-art MLLMs.
  • Experience with audio/speech generation or processing.
  • Experience with unified/multi-task foundation model architectures.
  • Industry research lab experience.

Responsibilities

  • Lead end-to-end AI research and model development for video-centric generative AI across Meta's advertising surfaces.
  • Drive advancements in video generation & enhancement.
  • Develop video-to-video & audio generation capabilities.
  • Advance video & visual understanding through novel research.
  • Conduct foundation model research to support generative AI innovation.
  • Define research agendas and pioneer new directions in video/audio generation and multimodal understanding.

Skills

Leadership experience
LLM training
Multimodal learning
Video generation/understanding
Diffusion models
Industry recognition

Education

Bachelor's degree in Computer Science, Computer Engineering or related field
PhD in Computer Science/AI/ML or related field

Job description

Meta's Monetization GenAI Video team is building the next generation of generative AI for video—creating, understanding, and transforming at scale to reach billions of users. We develop foundation models for video generation, video-to-video editing, audio generation, and multimodal video understanding powered by large language models.

The team seeks AI Research Scientists to define the research vision across video and audio generation, drive breakthroughs, and shape the technical direction of

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Research Scientist, Multimodal - Monetization GenAI
AI Research Scientist, Multimodal - Monetization GenAI

Meta • Menlo Park (CA)

On-site
USD 180,000 - 280,000
Senior GenAI Research Scientist - Multimodal AI Innovations
Senior GenAI Research Scientist - Multimodal AI Innovations

Socket.dev • Mountain View (CA)

On-site
USD 174,000 - 252,000
Bonus target
Equity
Benefits
Multimodal AI Research Scientist for Video Editing
Multimodal AI Research Scientist for Video Editing

Descript • San Francisco (CA)

Hybrid
USD 197,000 - 263,000
Equity
Research Scientist, Multimodal Video AI
Research Scientist, Multimodal Video AI

Cantina • California (MO)

On-site
USD 200,000 - 320,000
Competitive salary
Equity
Medical / Dental / Vision
+6
Generative AI Scientist for Video & Media
Generative AI Scientist for Video & Media

Amazon • New York (NY)

On-site
USD 172,000 - 223,000
Health insurance
RSUs
Paid time off
Research Engineer
Research Engineer

Harnham • California (MO)

On-site
USD 120,000 - 180,000
Multimodal AI Researcher: Generative Models, Realtime Vision
Multimodal AI Researcher: Generative Models, Realtime Vision

Socket.dev • Sunnyvale (CA)

Hybrid
USD 150,000 - 230,000
Generative AI Research Engineer - Video & Audio
Generative AI Research Engineer - Video & Audio

Google DeepMind • McLean (VA)

On-site
USD 120,000 - 160,000
Generative Multimodal AI Researcher
Generative Multimodal AI Researcher

Apple Inc. • Sunnyvale (CA)

Hybrid
USD 150,000 - 278,000
Stock programs
Discretionary bonuses
Relocation
+1
Research Scientist: Multi-Modal AI & World Models
Research Scientist: Multi-Modal AI & World Models

Meta • Pittsburgh, Redmond (WA)

On-site
USD 180,000 - 300,000