Staff Research Engineer: Large-Scale AI Training

Black Forest Labs Inc.

San Francisco, Northern (CA, KY)

Hybrid

USD 180,000 - 290,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Black Forest Labs Inc. is seeking an experienced engineer to join our production training effort for large multimodal models. You will work closely with researchers to optimize training throughput, memory usage, and numerical stability across GPUs and large GPU fleets.

The role emphasizes low-level training code, profiling, and kernel development, with opportunities to influence architecture-level training choices. Base salary ranges for US candidates are highlighted in the ad.

Qualifications

  • Experience with large-scale training systems, preferably with researchers
  • Strong PyTorch fluency and low-level training code familiarity
  • Understanding distributed training concepts (FSDP, tensor/model parallelism, NCCL)
  • Hands-on improvement of training throughput, memory footprint or stability
  • Experience profiling GPU workloads with Nsight or similar tools
  • Ability to verify numerical behavior and ownership of results
  • Knowledge of FP8/FP4-precision training and quantization tradeoffs
  • Comfort operating in ambiguous, fast-moving research-to-production environments

Responsibilities

  • Improve performance, reliability, and stability of production training runs for large multimodal models
  • Profile training steps across model code, attention, kernels, data loading, and memory
  • Implement and validate GPU optimizations: fused kernels, low-precision paths, quantization kernels
  • Advance FP8/FP4-style training and quantization strategies
  • Translate architecture changes into efficient training implementations with researchers
  • Debug distributed training issues (NaNs, memory leaks, NCCL, throughput)
  • Build benchmarking and profiling tools for cross-hardware validation
  • Collaborate with researchers and training team to remove bottlenecks

Skills

PyTorch
Distributed training
Profiling tools
Low-precision training
GPU kernels
NCCL

Tools

Nsight Systems
Nsight Compute
CUDA
Triton
CUTLASS
MXFP8/FP8

Job description

Black Forest Labs Inc. is seeking an experienced engineer to join our production training effort for large multimodal models. You will work closely with researchers to optimize training throughput, memory usage, and numerical stability across GPUs and large GPU fleets.

The role emphasizes low-level training code, profiling, and kernel development, with opportunities to influence architecture-level training choices. Base salary ranges for US candidates are highlighted in the ad.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Research Engineer - Large-Scale Training (Remote + Equity)
Research Engineer - Large-Scale Training (Remote + Equity)

Black Forest Labs • San Francisco (CA)

Hybrid
USD 180,000 - 290,000
AI Training Performance Engineer — Large-Scale GPU Training
AI Training Performance Engineer — Large-Scale GPU Training

Figure • San Jose (CA)

On-site
USD 200,000 - 400,000
ML Research Engineer — Large-Scale Training & Tools
ML Research Engineer — Large-Scale Training & Tools

The Resume Database • Palo Alto (CA)

On-site
USD 120,000 - 160,000
Research Engineer: High-Performance ML Infrastructure
Research Engineer: High-Performance ML Infrastructure

Fleet AI, Inc. • San Francisco (CA)

On-site
USD 180,000 - 240,000
AI Infrastructure Kernel Engineer for Large-Scale Training
AI Infrastructure Kernel Engineer for Large-Scale Training

Thinking Machines Lab Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 350,000 - 475,000
Health, dental, and vision benefits
Unlimited PTO
Parental leave
+1
Senior Model Serving & API Backend Engineer
Senior Model Serving & API Backend Engineer

Black Forest Labs • San Francisco (CA)

Hybrid
USD 180,000 - 300,000
Senior AI Training Performance Engineer (GPU & Scale)
Senior AI Training Performance Engineer (GPU & Scale)

figure.ai • San Jose (CA), Northern (KY)

Hybrid
USD 200,000 - 400,000
Research Engineer, GPU Performance
Research Engineer, GPU Performance

Harnham • California (MO)

On-site
USD 120,000 - 160,000
Research Engineer
Research Engineer

Harnham • United States

On-site
USD 120,000 - 150,000
Lead ML Training Systems Engineer - Multimodal, Large-Scale
Lead ML Training Systems Engineer - Multimodal, Large-Scale

Rhoda AI • Palo Alto (CA)

On-site
USD 210,000 - 320,000