ML Researcher - Scalable Image & Video Diffusion

Krea

San Francisco, Northern (CA, KY)

Hybrid

USD 140,000 - 240,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Health & wellness
Flexible PTO
401k with 4% company match
Meals in the office
Transit: Ubers covered
Visa sponsorship

Job summary

Krea in San Francisco is seeking an experienced Researcher with engineering skills to advance large-scale image and video model training. You will work on diffusion models, optimize distributed training pipelines, and experiment with architecture, data, and training strategies to push model quality and reliability.

The role requires hands-on experience with PyTorch, distributed training, and data pipelines.

Qualifications

  • Proven track record in working with image or video models at scale (publications or open-source contributions a plus).
  • Strong proficiency in PyTorch and understanding of its inner workings.
  • Strong background in distributed training paradigms such as FSDP, CP, SP, USP, TP, and EP. Knowing how different parallelism strategies work together and their tradeoffs.
  • Experience in profiling and debugging large distributed training. Being comfortable with analyzing traces to identify bottlenecks and look for improvements.
  • Good knowledge of low precision training / inference in FP8, NVFP4, and MXFP8.
  • Solid understanding of diffusion model training pipeline across pretraining, midtraining, preference optimization, and reinforcement learning.
  • Keeping up with the developments in related fields such as LLM, VLM, representation learning, and robotics research.
  • Being comfortable working in a goal-oriented research environment.
  • Having good judgement around when one should explore different training strategies and when it's time to commit to a specific strategy to scale compute and data.
  • Be comfortable getting your hands dirty with data and designing custom data pipelines to improve data quality.

Responsibilities

  • Train diffusion models for image and video generation on large GPU clusters.
  • Fully optimize and profile large distributed training runs across model architectures, kernels, data loading, memory constraints, and communication.
  • Implement and improve various distributed training strategies including FSDP, CP, SP, TP, and EP.
  • Continuously improve model quality and reliability through data, model architecture, training pipeline, structuring experiments, and eval design.
  • Debug distributed training errors and implement fault tolerance solutions, identifying bad GPU, NVLink, Infiniband (IB) components as well as monitoring numerical errors and NCCL issues.
  • Ablate different architecture, attention, optimizer, data, and algorithmic choices to reliably improve efficiency and performance of our models.

Skills

Image models
PyTorch
Distributed training
Profiling
FP8 training
Diffusion models
Experiment design

Tools

GPU clusters

Job description

Krea in San Francisco is seeking an experienced Researcher with engineering skills to advance large-scale image and video model training. You will work on diffusion models, optimize distributed training pipelines, and experiment with architecture, data, and training strategies to push model quality and reliability.

The role requires hands-on experience with PyTorch, distributed training, and data pipelines.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Diffusion AI Researcher — Image & Video at Scale
Diffusion AI Researcher — Image & Video at Scale

AI Chopping Block • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 210,000
Health insurance
401k with 4% company match
Meals in the office
+2
ML Researcher: Diffusion & RL for Creative AI
ML Researcher: Diffusion & RL for Creative AI

Krea • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 260,000
Health & dental insurance
Flexible PTO
401k with company match
+3
ML Engineer — Diffusion Image & Video Models
ML Engineer — Diffusion Image & Video Models

Krea • San Francisco (CA)

On-site
USD 140,000 - 210,000
ML Researcher - Image / Video Diffusion
ML Researcher - Image / Video Diffusion

Krea • San Francisco (CA), Northern (KY)

On-site
USD 140,000 - 240,000
Health & wellness
Flexible PTO
401k with 4% company match
+3
ML Researcher - Image / Video Diffusion
ML Researcher - Image / Video Diffusion

AI Chopping Block • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 210,000
Health insurance
401k with 4% company match
Meals in the office
+2
Video Diffusion AI Researcher - Build from Scratch, Equity
Video Diffusion AI Researcher - Build from Scratch, Equity

Amadeus Search • San Francisco (CA), Northern (KY)

Hybrid
USD 160,000 - 300,000
Senior ML Researcher — Diffusion & RL for Creative AI
Senior ML Researcher — Diffusion & RL for Creative AI

AI Chopping Block • San Francisco (CA), Northern (KY)

Hybrid
USD 190,000 - 230,000
Competitive salary
Equity package
Health & dental insurance
+5
AI Researcher, Video Diffusion
AI Researcher, Video Diffusion

Amadeus Search • San Francisco (CA), Northern (KY)

Hybrid
USD 160,000 - 300,000
Distributed Training Infra Engineer for Large Models
Distributed Training Infra Engineer for Large Models

Kindredventures • San Francisco (CA)

On-site
USD 180,000 - 240,000
Machine Learning Engineer
Machine Learning Engineer

Krea • San Francisco (CA)

On-site
USD 140,000 - 210,000