Generative AI Research Engineer: Vision & World Models

TikTok

San Jose (CA)

On-site

USD 162,000 - 388,000

Full time

7 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

TikTok, a global leader in short-form video, is building the Vision-Applied Research team in California to advance generative AI and CV/Multimodal understanding. The role focuses on developing next-generation World Models and scalable training pipelines using massive multimodal datasets to deliver intelligent content creation capabilities.

We seek researchers who can push long-horizon consistency, realistic physics, and interactive real-time systems across cross-functional teams and publications.

Qualifications

  • M.S or Ph.D. in Computer Vision, Computer Graphics, Machine Learning, or equivalent experience.
  • Extensive research experiences in broad GenAI, multimodal foundation models, or Embodied AI areas.
  • Demonstrated ability to communicate complex technical concepts and collaborate effectively within cross-functional research teams.
  • Preferred: experience in video generation and synthesis; diffusion models; 3D/physics-based simulation; or RL for agentic environment interaction.
  • Proven first-author publications in CVPR, ICLR, NeurIPS, SIGGRAPH, or ICML.

Responsibilities

  • Develop large-scale, diverse, and interactive multi-modal data generation pipeline.
  • Develop training pipeline for long-context interactive video generation models.
  • Advance video generation models to capture long-horizon temporal consistency, realistic physical dynamics, object interactions, and causal relationships from large-scale multi-modal data.

Skills

M.S./Ph.D. in Computer Vision/Computer
Multimodal AI
Communication skills

Education

M.S./Ph.D. in CV/CG/ML

Job description

TikTok, a global leader in short-form video, is building the Vision-Applied Research team in California to advance generative AI and CV/Multimodal understanding. The role focuses on developing next-generation World Models and scalable training pipelines using massive multimodal datasets to deliver intelligent content creation capabilities.

We seek researchers who can push long-horizon consistency, realistic physics, and interactive real-time systems across cross-functional teams and publications.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Tech Lead — Generative AI & Vision Innovator
Tech Lead — Generative AI & Vision Innovator

TikTok • San Jose (CA)

On-site
USD 254,000 - 480,000
Health insurance
401(k) with match
Paid parental leave
+6
World Model Research Scientist: Creative Generative AI
World Model Research Scientist: Creative Generative AI

TikTok • San Jose (CA)

On-site
USD 162,000 - 317,000
Senior GenAI Research Engineer—Efficient Large Models
Senior GenAI Research Engineer—Efficient Large Models

TikTok • Seattle (WA)

On-site
USD 242,000 - 456,000
Research Scientist, Neural Graphics & World Models
Research Scientist, Neural Graphics & World Models

TikTok • San Jose (CA)

On-site
USD 162,000 - 388,000
Generative AI Research Scientist - Seattle
Generative AI Research Scientist - Seattle

TikTok • Seattle (WA)

On-site
USD 202,160 - 368,220
Medical, dental, and vision insurance.
401(k) savings plan with company match
Paid parental leave
+6
Tech Lead: Neural Graphics & World Models
Tech Lead: Neural Graphics & World Models

TikTok • San Jose (CA)

On-site
USD 254,000 - 588,000
Generative AI Research Scientist: Multimodal Creation
Generative AI Research Scientist: Multimodal Creation

ByteDance • San Jose (CA)

On-site
USD 120,000 - 180,000
GenAI Research Scientist - Multimodal Video & AI
GenAI Research Scientist - Multimodal Video & AI

TikTok • San Jose (CA)

On-site
USD 162,000 - 317,000
Medical, dental, and vision insurance
401(k) with company match
Paid parental leave
+6
GenAI Research Scientist — Multimodal Vision & Audio
GenAI Research Scientist — Multimodal Vision & Audio

ByteDance • San Jose (CA)

On-site
USD 244,000 - 588,000
Medical, dental, and vision insurance
401(k) savings plan with company match
10 paid holidays and 10 sick days
+1
Research Scientist - Video Generation & Multimodal AI
Research Scientist - Video Generation & Multimodal AI

TikTok • San Jose (CA)

On-site
USD 244,000 - 588,000
Medical, dental, and vision insurance
401(k) savings plan with company match
Paid parental leave
+5