ML Researcher - Image / Video Diffusion

AI Chopping Block

San Francisco, Northern (CA, KY)

Hybrid

USD 150,000 - 210,000

Full time

10 days ago
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Health insurance
401k with 4% company match
Meals in the office
Transit covered
Visa sponsorship for international位

Job summary

Krea is building next-generation AI creative tools and seeks an experienced Researcher with engineering skills to advance large-scale image and video model training experiments. You will train diffusion models, optimize distributed runs, and push model quality while profiling architectures, loaders, and memory systems.

The role emphasizes hands-on experimentation, fast iteration, and collaboration with a talented team in a research-driven environment.

Qualifications

  • Proven track record in working with image or video models at scale.
  • Strong proficiency in PyTorch.
  • Experience with distributed training paradigms such as FSDP, CP, SP, TP, and EP.
  • Ability to profile and debug large distributed training jobs.
  • Good knowledge of low precision training (FP8, NVFP4, MXFP8).
  • Solid understanding of diffusion model training pipelines across pretraining to reinforcement learning.
  • Keeping up with developments in LLM, VLM, representation learning, and robotics research.

Responsibilities

  • Train diffusion models for image and video generation on large GPU clusters.
  • Fully optimize and profile large distributed training runs across model architectures, kernels, data loading, memory, and communication.
  • Implement and improve various distributed training strategies including FSDP, CP, SP, TP, and EP.
  • Continuously improve model quality and reliability through data, architecture, and training pipeline design.
  • Debug distributed training errors and implement fault tolerance solutions, identifying hardware and software bottlenecks.
  • Ablate architecture, attention, optimizer, data, and algorithm choices to improve efficiency and performance.

Skills

PyTorch
Distributed training
Image/Video models
Profiling & debugging
Low precision training
Diffusion models

Tools

GPU clusters
NVLink
Infiniband
NCCL

Job description

About Krea

At Krea, we are building next-generation AI creative tools.

We're dedicated to making AI intuitive and controllable for creatives - our mission is to build tools that empower human creativity, not replace it. We believe AI is a new medium that allows us to express ourselves through various formats - text, images, video, sound, and even 3D. We're building better, smarter, and more controllable tools to harness this medium. We recently took this a step forward with the launch of Krea 2, our first foundation model, built completely from scratch for aesthetic diversity and stylistic control.

We've raised over $83M and are backed by world-class investors such as a16z, Bain Capital, and Abstract. We work full-time and in-person at our waterfront office in San Francisco. We care about creativity: our team includes musicians, designers, visual artists, and engineers.

We're looking for an experienced Researcher with engineering skills who can work on large-scale image and video models training experiments, with experience training image models at scale.

Our culture
  • We work full-time and in-person at our North Beach office in San Francisco.

  • We believe that demonstrated interest in the creative space is key: our team includes musicians, designers, visual artists and more.

  • Fast iteration and execution speed. Bias towards action, agency, and independence.

What you'll do
  • Train diffusion models for image and video generation on large GPU clusters.

  • Fully optimize and profile large distributed training runs across model architectures, kernels, data loading, memory constraints, and communication.

  • Implement and improve various distributed training strategies including FSDP, CP, SP, TP, and EP.

  • Continuously improve model quality and reliability through data, model architecture, training pipeline, structuring experiments, and eval design.

  • Debug distributed training errors and implement fault tolerance solutions, identifying bad GPU, NVLink, Infiniband (IB) components as well as monitoring numerical errors and NCCL issues.

  • Ablate different architecture, attention, optimizer, data, and algorithmic choices to reliably improve efficiency and performance of our models.

What we're looking for
  • Proven track record in working with image or video models at scale (publications or open-source contributions a plus).

  • Strong proficiency in PyTorch and understanding of its inner workings.

  • Strong background in distributed training paradigms such as FSDP, CP, SP, USP, TP, and EP. Knowing how different parallelism strategies work together and their tradeoffs.

  • Experience in profiling and debugging large distributed training. Being comfortable with analyzing traces to identify bottlenecks and look for improvements.

  • Good knowledge of low precision training / inference in FP8, NVFP4, and MXFP8.

  • Solid understanding of diffusion model training pipeline across pretraining, midtraining, preference optimization, and reinforcement learning.

  • Keeping up with the developments in related fields such as LLM, VLM, representation learning, and robotics research.

  • Being comfortable working in a goal-oriented research environment.

  • Having good judgement around when one should explore different training strategies and when it's time to commit to a specific strategy to scale compute and data.

  • Comfortable working with underspecified goals. We expect every technical member to take an ambiguous research goal and break it down into concrete requirements, plans, experiment plan, and execution items.

  • Good research taste — bias towards simplicity and methods that scale well with compute, data, and minimal human supervision.

  • Ability to iterate rapidly, and propose creative research directions.

  • Be comfortable getting your hands dirty with data and designing custom data pipelines to improve data quality.

What we offer
  • Team: Work alongside a world-class team building the future of AI creative tooling

  • Impact: Significant scope and company-wide impact

  • Competitive compensation: generous salary & equity packages

  • Health & wellness: 100% health & 99% dental/vision insurance premiums covered for employees, health FSA accounts, & long-term disability coverage

  • Time off: Flexible PTO policy

  • Financial planning: 401k with a 4% company-sponsored match

  • Meals in the office: breakfast, lunch, dinner - you name it, we'll cover it

  • Transit: Ubers covered to & from the office

  • Sponsorship: We're open to sponsoring international visas where we can (e.g., STEM OPT, OPT, H-1B, O-1, E-3).

  • And more!

Please note the above benefits & perks are for full-time employees

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Researcher - Image / Video Diffusion
ML Researcher - Image / Video Diffusion

Krea • San Francisco (CA), Northern (KY)

On-site
USD 140,000 - 240,000
Health & wellness
Flexible PTO
401k with 4% company match
+3
ML Researcher - Posttraining
ML Researcher - Posttraining

Krea • San Francisco (CA), Northern (KY)

On-site
USD 180,000 - 260,000
Health & dental insurance
Flexible PTO
401k with company match
+3
ML Researcher - Posttraining
ML Researcher - Posttraining

AI Chopping Block • San Francisco (CA), Northern (KY)

Hybrid
USD 190,000 - 230,000
Competitive salary
Equity package
Health & dental insurance
+5
Fullstack Engineer
Fullstack Engineer

AI Chopping Block • San Francisco (CA), Northern (KY)

Hybrid
USD 140,000 - 200,000
Competitive compensation
Health & dental insurance
401k with company match
+3
Fullstack Engineer
Fullstack Engineer

Krea • San Francisco (CA), Northern (KY)

On-site
USD 180,000 - 260,000
Health & dental coverage
401k with company match
Meals in office
+2
Member of Technical Staff - Research Engineer
Member of Technical Staff - Research Engineer

Black Forest Labs Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 290,000
Member of Technical Staff, Post-Training & Applied Research
Member of Technical Staff, Post-Training & Applied Research

San Francisco Tensor Company • San Francisco (CA)

On-site
USD 275,000 - 315,000
Relocation assistance
Software Engineer, Backend
Software Engineer, Backend

Krea • San Francisco (CA)

On-site
USD 140,000 - 200,000
Team: World-class team
Impact: Company-wide impact
Competitive compensation
+1
AI Researcher, Video Diffusion
AI Researcher, Video Diffusion

Amadeus Search • San Francisco (CA), Northern (KY)

Hybrid
USD 160,000 - 300,000
Member of Technical Staff - Research Engineer
Member of Technical Staff - Research Engineer

Black Forest Labs • San Francisco (CA)

On-site
USD 180,000 - 290,000