Research Scientist, Foundation Model (Video Generation)

Pika

Palo Alto (CA)

Hybrid

USD 190,000 - 250,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Competitive salary
Equity
Health benefits
Hybrid work option

Job summary

Pika is seeking a Research Scientist to lead the development of large-scale multimodal foundation models, focusing on pre-training and mid-training. You will design novel algorithms, curate data, and drive production-ready implementations in collaboration with engineering and product teams.

Ideal candidates have 5+ years of research experience, strong publication records, and hands-on expertise with PyTorch/TensorFlow.

Qualifications

  • 5+ years of research experience in large-scale pre-training/mid-training of multimodal foundation models.
  • First-author publications in top conferences or journals (e.g., NeurIPS, ICML, ICLR).
  • Hands-on experience with design, training, and deployment of large-scale multimodal models.
  • Deep understanding of generative architectures (diffusion, autoregressive, cross-modal).
  • Expertise in scalable data curation and model pipeline optimization.

Responsibilities

  • Lead research on pre-training and mid-training of multimodal foundation models at scale.
  • Design and prototype novel architectures for real-time multimodal synthesis across modalities.
  • Develop scalable data pipelines and training strategies for diverse datasets.
  • Advance state-of-the-art techniques and bring research into production with engineering teams.
  • Publish work in top-tier conferences and journals and communicate findings clearly.

Skills

Multimodal ML
Python
PyTorch
TensorFlow
Research leadership

Tools

Distributed training
GPU clusters

Job description

About the Role

At Pika, we are pioneering the next generation of creative infrastructure built around real-time, multimodal generation and intelligent agentic platforms. We are seeking accomplished Research Scientists in Foundation Models with expertise in pre-training and mid-training large-scale multimodal foundation models to advance our mission of making agentic, real-time generative technology accessible and transformative for millions of creators. This is a staff and lead-level opportunity.

As a key member of our research team, you will design and implement core technologies, develop new methodologies for large-scale multimodal pre-training/mid-training (text, image, audio, and video), and drive innovative approaches for foundational model architecture. You will collaborate closely with engineering and product teams, shaping the future of real-time creative and agentic platforms at scale.

What You’ll Do
  • Lead research and development on pre-training and mid-training of multimodal foundation models at scale.
  • Design and prototype novel algorithms and architectures for high-fidelity, real-time multimodal synthesis and interaction across modalities.
  • Focus on scalable data pipeline curation and model training strategies for broad, diverse, and sensory-rich datasets.
  • Advance state-of-the‑art techniques in diffusion, autoregressive, and other generative models for large-scale pre-training and fine-tuning.
  • Identify, create, and leverage large, high-quality cross-modal datasets.
  • Bring research advancements into production-ready systems in collaboration with engineering and product teams.
  • Publish work in top-tier conferences and journals, and clearly communicate research both internally and externally.
  • Stay at the forefront of foundational model and real-time multimodal AI research.
What We’re Looking For
  • 5+ years of research experience in large-scale pre-training/mid-training of multimodal foundation models (LLMs, VLMs, Audio LMs, or similar), ideally at the staff or lead scientist level.
  • Track record as a first author on major publications in top conferences or journals (e.g., NeurIPS, ICML, ICLR).
  • Extensive hands‑on experience with large-scale multimodal model design, training, and deployment.
  • Deep understanding and implementation experience with generative architectures (diffusion, autoregressive, cross-modal, etc.).
  • Expertise in high-throughput, scalable dataset curation and model pipeline optimization for multimodal applications.
  • Strong programming and prototyping skills (Python, PyTorch, TensorFlow, etc.) and experience deploying research into production systems.
  • Excellent communication and collaboration skills, and a passion for building creative enabling technology.
What We Offer
  • Competitive salary and substantial equity in a high-growth startup
  • Full health benefits + 401k matching and more
  • Collaborative, mission-driven team environment with major growth opportunities
  • Flexible on-site/remote hybrid (HQ in Palo Alto, CA)
About Pika

Pika empowers creators by building state-of-the‑art agentic and multimedia platforms. Our vision is to break down technical barriers to creativity, making real-time generative and intelligent orchestration accessible to all. Join us and help shape the next evolution of creative technology!

If you are a leading researcher excited to build and scale real-time multimodal foundation models, we want to hear from you.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Research Scientist, Data
Research Scientist, Data

Pika • Palo Alto (CA)

Hybrid
USD 120,000 - 160,000
Competitive salary
Full health benefits
401k matching
+1
Lead Research Scientist, Multimodal Foundation Models
Lead Research Scientist, Multimodal Foundation Models

Pika • Palo Alto (CA)

Hybrid
USD 190,000 - 250,000
Competitive salary
Equity
Health benefits
+1
Research Engineer
Research Engineer

Harnham • California (MO)

On-site
USD 120,000 - 180,000
Staff Research Engineer - Multimodal Generative Modelling
Staff Research Engineer - Multimodal Generative Modelling

synthesia • United States

On-site
USD 180,000 - 260,000
Research Scientist – World Modeling, Data
Research Scientist – World Modeling, Data

Institute of Foundation Models • Sunnyvale (CA)

On-site
USD 150,000 - 400,000
Medical benefits
Dental/vision
401K
+5
Machine Learning Researcher / Engineer (Foundational Models)
Machine Learning Researcher / Engineer (Foundational Models)

Pathway Genomics Corporation • Palo Alto (CA)

Remote
USD 100,000 - 140,000
Employee Stock Option Plan
Intellectually stimulating environment
Opportunity to work on impactful research
Member of Technical Staff, Research
Member of Technical Staff, Research

Odyssey • Santa Clara (CA)

On-site
USD 180,000 - 240,000
Research Scientist, Video Foundation Models
Research Scientist, Video Foundation Models

Cantina • California (MO)

On-site
USD 200,000 - 320,000
Competitive salary
Equity
Medical / Dental / Vision
+6
Senior Research Scientist - Foundation Models (Remote, CH)
Senior Research Scientist - Foundation Models (Remote, CH)

careers.bitkraft.vc - Jobboard • Indiana (PA)

On-site
USD 148,000 - 221,000
Research Scientist, Multimodal Generative AI (Intelligent Creation)– Global Frontier Tech Recru[...]
Research Scientist, Multimodal Generative AI (Intelligent Creation)– Global Frontier Tech Recru[...]

TikTok • San Jose (CA)

On-site
USD 244,800 - 588,000
Medical, dental, and vision insurance
401(k) with company match
Paid parental leave
+6