Lead RL Post-Training Scientist: Video Generation

pika

United States

Hybrid

USD 180,000 - 280,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Competitive salary
Full health benefits
401k matching
Hybrid work model (HQ in Palo Alto)

Job summary

Pika is seeking Research Scientists to lead RL‑driven post‑training for real‑time video generation models and to develop reward models that align with human preferences. You will collaborate across engineering and product to push the frontier of agentic, multimodal video platforms.

This staff/lead role focuses on refining diffusion/flow‑matching models, distilling RL‑tuned models, and evaluating results with both human studies and automated metrics.

Qualifications

  • 2+ years hands‑on research experience in post‑training or generative modeling.
  • RL or preference‑optimization experience on generative models with evidence of model improvement.
  • Strong grounding in diffusion or flow‑matching models, PyTorch, and multi‑node distributed training.

Responsibilities

  • Run RL post-training (preference optimization, online RL against learned rewards) for video diffusion/flow-matching models at multi-node scale.
  • Build video reward models: define target evaluation dimensions, design/data collection workflows, train and validate learned judges, safeguard against reward hacking.
  • Own post-training evaluation, including human preference studies and their correlation with automated metrics.
  • Distill RL-tuned models to efficient few-step samplers while preserving alignment gains (secondary focus).

Skills

Hands-on RL post-training
RL / preference-optimization
Diffusion / flow-matching

Tools

PyTorch

Job description

Pika is seeking Research Scientists to lead RL‑driven post‑training for real‑time video generation models and to develop reward models that align with human preferences. You will collaborate across engineering and product to push the frontier of agentic, multimodal video platforms.

This staff/lead role focuses on refining diffusion/flow‑matching models, distilling RL‑tuned models, and evaluating results with both human studies and automated metrics.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Staff Research Scientist - RL Post-Training for Video Gen
Staff Research Scientist - RL Post-Training for Video Gen

Sarah Smith Fund • Palo Alto (CA)

On-site
USD 180,000 - 260,000
Health benefits
401k matching
Equity
Research Scientist, Post-Training — Video Generation
Research Scientist, Post-Training — Video Generation

pika • United States

Hybrid
USD 180,000 - 280,000
Competitive salary
Full health benefits
401k matching
+1
Research Scientist, Post-Training — Video Generation
Research Scientist, Post-Training — Video Generation

Sarah Smith Fund • Palo Alto (CA)

Hybrid
USD 180,000 - 260,000
Health benefits
401k matching
Equity
Research Member of Technical Staff- Post-training & Robot Learning
Research Member of Technical Staff- Post-training & Robot Learning

Rhoda AI • Mountain View (CA)

On-site
USD 100,000 - 150,000
ML Researcher: Diffusion & RL for Creative AI
ML Researcher: Diffusion & RL for Creative AI

Krea • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 260,000
Health & dental insurance
Flexible PTO
401k with company match
+3
research scientist - RL
research scientist - RL

Cerebro • San Francisco (CA)

On-site
USD 180,000 - 240,000
AI Research Engineer: Post-Training & Agentic RL
AI Research Engineer: Post-Training & Agentic RL

Oho Group • San Francisco (CA)

On-site
USD 160,000 - 230,000
RESEARCHER, POST-TRAINING
RESEARCHER, POST-TRAINING

MakerMaker.AI • San Francisco (CA)

On-site
USD 120,000 - 160,000
Research Engineer
Research Engineer

Oho Group • San Francisco (CA)

On-site
USD 160,000 - 230,000
Member of Technical Staff, Post-training
Member of Technical Staff, Post-training

Hark • San Jose (CA)

On-site
USD 180,000 - 450,000