Research Engineer: Large-Scale AI Training Systems

BlackForestLabs

San Francisco (CA)

Hybrid

USD 180,000 - 290,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity

Job summary

Black Forest Labs is recruiting a Member of Technical Staff - Research Engineer to own and optimize large-scale generative-model training systems in a hybrid SF/remote setting. You’ll work closely with researchers to push performance, memory footprint, and stability across GPU clusters.

You’ll implement GPU-level optimizations, profile end-to-end training pipelines, and help translate architectural changes into efficient training code, tooling, and diagnostics for high-quality results.

Qualifications

  • Experience working deeply on large-scale training systems, ideally with researchers
  • Strong PyTorch fluency, including reading and modifying low-level training code
  • Experience with distributed training concepts: FSDP, tensor/model parallelism, NCCL
  • Hands-on experience improving throughput, memory footprint, or stability in real runs
  • Experience profiling GPU workloads with Nsight tools or torch profiler

Responsibilities

  • Improve performance, reliability, and numerical stability of production training runs for large multimodal models
  • Profile full training steps across model code, attention, kernels, data loading, encoders, communication, and checkpointing
  • Implement and validate GPU-level optimizations: fused kernels, low-precision paths, and quantization kernels
  • Debug distributed training failures and build benchmarking/profiling harnesses for trustworthy performance measurements
  • Collaborate with researchers to translate architecture changes into efficient training implementations

Skills

PyTorch fluency
Large-scale training
Distributed training
Profiling GPU workloads
Low-precision training

Tools

Nsight Systems
Nsight Compute
Torch profiler

Job description

Black Forest Labs is recruiting a Member of Technical Staff - Research Engineer to own and optimize large-scale generative-model training systems in a hybrid SF/remote setting. You’ll work closely with researchers to push performance, memory footprint, and stability across GPU clusters.

You’ll implement GPU-level optimizations, profile end-to-end training pipelines, and help translate architectural changes into efficient training code, tooling, and diagnostics for high-quality results.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Research Engineer - Large-Scale Training (Remote + Equity)
Research Engineer - Large-Scale Training (Remote + Equity)

Black Forest Labs • San Francisco (CA)

Hybrid
USD 180,000 - 290,000
Staff Engineer: Large-Scale ML Training Systems
Staff Engineer: Large-Scale ML Training Systems

United States Digital Space LLC • San Francisco (CA)

Hybrid
USD 180,000 - 290,000
Equity
Travel cost coverage
Hybrid/onsite collaboration
Research Engineer: High-Performance ML Infrastructure
Research Engineer: High-Performance ML Infrastructure

Fleet AI, Inc. • San Francisco (CA)

On-site
USD 180,000 - 240,000
Research Engineer, GPU Performance
Research Engineer, GPU Performance

Harnham • California (MO)

On-site
USD 120,000 - 160,000
Research Engineer
Research Engineer

Harnham • United States

On-site
USD 120,000 - 150,000
Senior Model Serving & API Backend Engineer
Senior Model Serving & API Backend Engineer

Black Forest Labs • San Francisco (CA)

Hybrid
USD 180,000 - 300,000
Member of Technical Staff - Research Engineer San Francisco (USA), Freiburg (Germany)
Member of Technical Staff - Research Engineer San Francisco (USA), Freiburg (Germany)

BlackForestLabs • San Francisco (CA)

Hybrid
USD 180,000 - 290,000
Equity
Lead Large-Scale GPU Cluster Engineer for AI Research
Lead Large-Scale GPU Cluster Engineer for AI Research

Linuxcareers • San Francisco (CA)

On-site
USD 120,000 - 180,000
ML Model Serving & High-Performance API Engineer
ML Model Serving & High-Performance API Engineer

Black Forest Labs Inc. • San Francisco (CA)

On-site
USD 180,000 - 260,000
Travel reimbursement
Remote Senior Training Infrastructure Engineer—Multi-GPU AI
Remote Senior Training Infrastructure Engineer—Multi-GPU AI

Luma AI • San Francisco (CA)

Hybrid
USD 187,000 - 395,000