Research Scientist, Multi-Modal Understanding & Synthesis

Meta

Pittsburgh, Redmond (Allegheny County, WA)

On-site

USD 180,000 - 300,000

Full time

10 days ago
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Meta is seeking a Research Scientist to advance foundational research in multi-modal understanding, synthesis, and world models. The role focuses on building AI systems that perceive, reason, and generate content across vision, language, audio, and other modalities.

You will lead research initiatives, publish in top venues, and collaborate with engineers to translate breakthroughs into Meta's next-generation AI products.

Qualifications

  • PhD in ML, CV, NLP or related field, or equivalent practical experience.
  • 6+ years conducting AI research in multi-modal learning or world models.
  • Experience leading research initiatives from conception to publication or deployment.
  • Proficiency in Python and deep learning frameworks (PyTorch/TensorFlow).
  • Strong publication record in top ML/AI venues.

Responsibilities

  • Lead foundational research in multi-modal understanding, synthesis, and world models.
  • Develop and evaluate large-scale multi-modal systems and models.
  • Publish influential work and present findings to diverse audiences.
  • Collaborate with research and engineering teams to deploy breakthroughs.

Skills

Multi-modal learning
World models
Deep learning
Python
PyTorch
TensorFlow
Research leadership
Publications
Cross-functional comms
Model-based RL

Education

PhD in ML / CV / NLP
Bachelor's in CS/EE or equivalent

Tools

PyTorch
TensorFlow

Job description

Meta is seeking a Research Scientist to drive foundational research in multi-modal understanding, synthesis, and world models. In this role, you will advance the state of the art in building AI systems that perceive, reason across, and generate content spanning vision, language, audio, and other modalities. You will develop world models that learn rich internal representations of human behavior, enabling prediction, planning, and simulation. Collaborating with world-class researchers and engineers, you will define research directions, publish influential work, and translate breakthroughs into technologies that power Meta's next-generation AI products.

  • Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience
  • PhD in Machine Learning, Computer Vision, Natural Language Processing, or a closely related field
  • 6+ years of experience conducting AI research in multi-modal learning, generative models, or world models, including experience leading major research initiatives from conception through publication or production deployment
  • Experience implementing and evaluating multi-modal systems using deep learning frameworks such as PyTorch or TensorFlow, with proficiency in Python
  • Experience publishing original research in peer-reviewed machine learning or AI venues
  • Experience driving cross-functional technical decisions and communicating research findings and trade-offs to both research and engineering audiences through written documents and presentations
  • Experience with techniques spanning multiple modalities such as vision-language models, multi-modal transformers, or cross-modal representation learning
  • Experience developing large-scale multi-modal foundation models or vision-language models
  • Experience with world models, predictive learning, or model-based reinforcement learning for planning and reasoning
  • First-author publications at top-tier venues such as NeurIPS, ICLR, or CVPR demonstrating contributions to multi-modal learning, generative models, or world models
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Research Scientist: Multi-Modal AI & World Models
Research Scientist: Multi-Modal AI & World Models

Meta • Pittsburgh, Redmond (WA)

On-site
USD 180,000 - 300,000
Research Scientist, Multi-Modal Human Understanding
Research Scientist, Multi-Modal Human Understanding

Meta • Pittsburgh, Burlingame (CA)

On-site
USD 150,000 - 190,000
Multi-Modal AI Research Scientist – Human Understanding
Multi-Modal AI Research Scientist – Human Understanding

Meta • Pittsburgh, Burlingame (CA)

On-site
USD 150,000 - 190,000
Research Scientist - Multi-modal AI & Efficient Generative Models
Research Scientist - Multi-modal AI & Efficient Generative Models

Meta • Redmond (WA), Burlingame (CA)

On-site
USD 180,000 - 240,000
Research Scientist, Machine Learning
Research Scientist, Machine Learning

Meta • Sunnyvale (CA)

On-site
USD 150,000 - 210,000
Senior Research Scientist: Multi-Modal AI & Efficient Gen Models
Senior Research Scientist: Multi-Modal AI & Efficient Gen Models

Meta • Redmond (WA), Burlingame (CA)

On-site
USD 180,000 - 240,000
AI Research Scientist, Video Generation and Post Training, FAIR
AI Research Scientist, Video Generation and Post Training, FAIR

Meta • Menlo Park (CA), Seattle (WA), New York (NY)

On-site
USD 180,000 - 240,000
Research Scientist, Robotics - Meta Superintelligence Labs
Research Scientist, Robotics - Meta Superintelligence Labs

Meta • Menlo Park (CA)

On-site
USD 180,000 - 280,000
Research Scientist, AI (Technical Leadership)
Research Scientist, AI (Technical Leadership)

Meta • Burlingame (CA), New York (NY)

On-site
USD 180,000 - 270,000
AI Research Scientist, Multimodal - Monetization GenAI
AI Research Scientist, Multimodal - Monetization GenAI

Meta • Menlo Park (CA)

On-site
USD 180,000 - 280,000