Research Scientist, Post-Training — Video Generation

Sarah Smith Fund

Palo Alto (CA)

Hybrid

USD 180,000 - 260,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Health benefits
401k matching
Equity

Job summary

Pika is seeking Research Scientists with expertise in RL post-training and generative modeling for large-scale video generation. You will own RL-based post-training for video diffusion/flow-matching models and help build robust reward models, collaborating with engineering and product teams.

The role covers evaluating models with human and automated metrics, distilling RL-tuned models for efficient deployment, and advancing state-of-the-art video generation capabilities in a fast-moving startup

Qualifications

  • 2+ years hands-on research in post-training or generative modeling.
  • Experience with RL or preference optimization on generative models.
  • Strong grounding in diffusion/flow-matching models and PyTorch.

Responsibilities

  • Run RL post-training (preference optimization, online RL against learned rewards) for video diffusion/flow-matching models at multi-node scale.
  • Build video reward models: define target evaluation dimensions, design/configure preference data collection workflows, train and validate learned judges, and safeguard against reward hacking.
  • Own post-training evaluation, including human preference studies and their correlation with automated metrics.
  • Distill RL-tuned models to efficient few-step samplers while preserving alignment gains (secondary focus).

Skills

Post-training research
Generative modeling
RL
PyTorch
Distributed training

Tools

Diffusion models
Flow-matching
Video generation

Job description

About the Role

At Pika, we are pioneering the next generation of creative infrastructure built around real-time, multimodal generation and intelligent agentic platforms. We are seeking Research Scientists with expertise in RL post-training and generative modeling for large-scale video generation. The focus is on refining Pika's video generation models using RL alignment and building robust video reward models. This is a staff and lead-level opportunity.

As a key member of our research team, you will own RL-based post-training for video diffusion/flow-matching models, develop state-of-the-art reward models, and lead post-training evaluation across human and automated metrics. You will collaborate closely with engineering and product teams, shaping the frontier of real-time creative and agentic video platforms.

Scope
  • RL alignment of Pika's video generation models and the reward models that drive them.
  • Distillation of RL-tuned models is a secondary focus.
Responsibilities
  • Run RL post-training (preference optimization, online RL against learned rewards) for video diffusion/flow-matching models at multi-node scale.
  • Build video reward models: define target evaluation dimensions, design/configure preference data collection workflows, train and validate learned judges, and safeguard against reward hacking.
  • Own post-training evaluation, including human preference studies and their correlation with automated metrics.
  • Distill RL-tuned models to efficient few-step samplers while preserving alignment gains (secondary focus).
What We’re Looking For
Required
  • 2+ years hands-on research experience in post-training or generative modeling.
  • RL or preference-optimization experience on generative models with evidence of model improvement.
  • Strong grounding in diffusion or flow-matching models, PyTorch, and multi-node distributed training.
Preferred
  • Experience developing reward models for visual generation, including VLM-as-judge or large-scale preference data collection.
  • Distillation expertise (distribution matching, consistency, adversarial approaches), ideally for video models.
  • Familiarity with video-specific failure modes: temporal drift, motion and physics realism.
What We Offer
  • Competitive salary and substantial equity in a high-growth startup
  • Full health benefits + 401k matching and more
  • Collaborative, mission-driven team environment with significant growth opportunities
  • Flexible on-site/remote hybrid (HQ in Palo Alto, CA)
About Pika

Pika empowers creators by building state-of-the-art agentic and multimedia platforms. Our vision is to break down technical barriers to creativity, making real-time generative and intelligent orchestration accessible to all. Join us to shape the next evolution of creative technology!

If you are passionate about advancing RL alignment and generative modeling for video, and want to scale real-time multimodal foundation models, we want to hear from you.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Research Scientist, Post-Training — Video Generation
Research Scientist, Post-Training — Video Generation

pika • United States

Hybrid
USD 180,000 - 280,000
Competitive salary
Full health benefits
401k matching
+1
Staff Research Scientist - RL Post-Training for Video Gen
Staff Research Scientist - RL Post-Training for Video Gen

Sarah Smith Fund • Palo Alto (CA)

On-site
USD 180,000 - 260,000
Health benefits
401k matching
Equity
Lead RL Post-Training Scientist: Video Generation
Lead RL Post-Training Scientist: Video Generation

pika • United States

Hybrid
USD 180,000 - 280,000
Competitive salary
Full health benefits
401k matching
+1
Research Scientist, Data
Research Scientist, Data

pika • United States

Hybrid
USD 180,000 - 240,000
Competitive salary & equity
Full health benefits
Hybrid on-site/remote (HQ Palo Alto)
Research Scientist, Data
Research Scientist, Data

Pika • Palo Alto (CA)

Hybrid
USD 120,000 - 160,000
Competitive salary
Full health benefits
401k matching
+1
Research Member of Technical Staff- Post-training & Robot Learning
Research Member of Technical Staff- Post-training & Robot Learning

Rhoda AI • Mountain View (CA)

On-site
USD 100,000 - 150,000
RESEARCHER, POST-TRAINING
RESEARCHER, POST-TRAINING

MakerMaker.AI • San Francisco (CA)

On-site
USD 120,000 - 160,000
Research Engineer, Infrastructure, RL Systems
Research Engineer, Infrastructure, RL Systems

Thinkingmachines • San Francisco (CA)

On-site
USD 350,000 - 475,000
Health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1
Research Engineer/Scientist - Human Alignment, Consumer Devices
Research Engineer/Scientist - Human Alignment, Consumer Devices

SupportFinity™ • San Francisco (CA)

On-site
USD 100,000 - 180,000
Research, Post-Training
Research, Post-Training

Thinkingmachines • San Francisco (CA)

On-site
USD 350,000 - 475,000
Health benefits
Dental and vision benefits
Unlimited PTO
+2