Research Scientist, Multi-Modal Human Understanding

Meta

Pittsburgh, Burlingame (Allegheny County, CA)

On-site

USD 150,000 - 190,000

Full time

12 days ago
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Meta seeks a Research Scientist to advance multi-modal AI technologies for human understanding and synthesis. You will develop Vision-Language Models (VLMs) and video foundation models that enable machines to perceive, interpret, and generate rich representations of human behavior and interaction.

Requirements include a Bachelor's degree in a technical field with 2+ years in multi-modal AI, PyTorch experience, and a track record of research contributions at major venues.

Qualifications

  • Candidate has a BS or higher in a technical field before joining Meta.
  • 2+ years in multi-modal AI research with VLMs, video or human-centric AI.
  • Experience training large-scale neural networks and transformers.
  • Proven ability to design experiments and analyze multi-modal benchmarks.
  • Proficient in Python for research or production code.
  • Publication track record in CVPR/ICCV/NeurIPS is valued.

Responsibilities

  • Develop Vision-Language Models and video foundation models for understanding and generation.
  • Advance multi-modal AI research spanning vision, language, and video.
  • Implement and train large-scale neural networks using PyTorch.
  • Design experiments to evaluate model performance across modalities.
  • Contribute to publications and disseminate research findings.

Skills

Multimodal AI
Vision-Language Models
Video Understanding
Transformers
Python
Research Experience
Scientific Publications

Education

Bachelor's degree in Computer Science / Computer Engineering or related field

Tools

PyTorch

Job description

Meta is seeking a Research Scientist to advance multi-modal AI technologies for human understanding and synthesis. In this role, you will develop Vision-Language Models (VLMs) and video foundation models that enable machines to perceive, interpret, and generate rich representations of human behavior, expression, and interaction. Your research will span multi-modal reasoning, video understanding, and generative synthesis, enabling more natural and intuitive human-computer interaction at scale.

Currently has, or is in the process of obtaining a Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience. Degree must be completed prior to joining Meta

  • 2+ years of experience in multi-modal AI research, including hands-on work with Vision-Language Models, video understanding, or human-centric AI systems
  • 2+ years of experience implementing and training large-scale neural networks using frameworks such as PyTorch, with experience on transformer-based architectures
  • Experience designing and executing experiments to evaluate multi-modal model performance, including quantitative analysis across vision, language, and video benchmarks
  • Experience writing production-quality or research-quality code in Python for multi-modal AI applications
  • Experience developing or fine-tuning Vision-Language Models for human understanding tasks
  • Experience with video foundation models, temporal transformers, or large-scale video pretraining
  • Track record of contributing to published multi-modal AI research at venues such as CVPR, ICCV, or NeurIPS
  • Experience with generative models for human synthesis including diffusion models, GANs, or autoregressive models for video or motion generation
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Multi-Modal AI Research Scientist – Human Understanding
Multi-Modal AI Research Scientist – Human Understanding

Meta • Pittsburgh, Burlingame (CA)

On-site
USD 150,000 - 190,000
AI Research Scientist, Video Generation and Post Training, FAIR
AI Research Scientist, Video Generation and Post Training, FAIR

Meta • Menlo Park (CA), Seattle (WA), New York (NY)

On-site
USD 180,000 - 240,000
Research Scientist - Multi-modal AI & Efficient Generative Models
Research Scientist - Multi-modal AI & Efficient Generative Models

Meta • Redmond (WA), Burlingame (CA)

On-site
USD 180,000 - 240,000
Senior Research Scientist: Multi-Modal AI & Efficient Gen Models
Senior Research Scientist: Multi-Modal AI & Efficient Gen Models

Meta • Redmond (WA), Burlingame (CA)

On-site
USD 180,000 - 240,000
Research Scientist, Computer Vision (PhD)
Research Scientist, Computer Vision (PhD)

Socket.dev • Pittsburgh

On-site
USD 122,000 - 181,000
Bonus
Equity
Benefits
Research Scientist, Machine Learning
Research Scientist, Machine Learning

Meta • Sunnyvale (CA)

On-site
USD 150,000 - 210,000
Senior Research Scientist: Multimodal VLM & Video Understanding
Senior Research Scientist: Multimodal VLM & Video Understanding

techire ai • San Francisco (CA)

On-site
USD 210,000 - 260,000
AI Research Scientist, Computer Vision
AI Research Scientist, Computer Vision

Meta • Menlo Park (CA)

On-site
USD 154,000 - 217,000
Bonus
Equity
Benefits
Research Engineer, Robotics - Meta Superintelligence Labs
Research Engineer, Robotics - Meta Superintelligence Labs

Meta • Menlo Park (CA)

On-site
USD 130,000 - 210,000
Research Engineer (Technical Leadership), FAIR Data - Meta Superintelligence Labs
Research Engineer (Technical Leadership), FAIR Data - Meta Superintelligence Labs

Meta • Menlo Park (CA)

On-site
USD 219,000 - 301,000