Senior Distributed ML Systems Engineer

Luma

Greater London

On-site

GBP 147,000 - 299,000

Full time

1 hour ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Luma is building distributed training systems for its large-scale multimodal models. This role focuses on PyTorch, CUDA, and advanced parallelism across thousands of GPUs, delivering reliable, scalable infrastructure so researchers can innovate.

You will design and optimize training systems, implement FSDP, Tensor Parallel, Pipeline Parallel, and Expert Parallel, and build monitoring and debugging tools to improve stability and utilization.

Qualifications

  • Extensive distributed PyTorch training and parallelism in foundation-model training.
  • Deep understanding of GPU clusters, networking, and storage systems.
  • Familiarity with communication libraries (NCCL, MPI) and distributed-system optimization.
  • Experience managing training runs across 100+ GPUs.

Responsibilities

  • Design, implement, and optimize distributed training systems for models across thousands of GPUs.
  • Research and implement advanced parallelization (FSDP, Tensor Parallel, Pipeline Parallel, Expert Parallel).
  • Build monitoring, visualization, and debugging tools for large-scale training runs.
  • Optimize training stability, convergence, and resource utilization across massive clusters.

Skills

Distributed PyTorch training
GPU clusters
Networking and storage
NCCL/MPI

Tools

PyTorch
NCCL
MPI

Job description

Luma is building distributed training systems for its large-scale multimodal models. This role focuses on PyTorch, CUDA, and advanced parallelism across thousands of GPUs, delivering reliable, scalable infrastructure so researchers can innovate.

You will design and optimize training systems, implement FSDP, Tensor Parallel, Pipeline Parallel, and Expert Parallel, and build monitoring and debugging tools to improve stability and utilization.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Distributed Training Engineer — Large-Scale GPU Systems
Senior Distributed Training Engineer — Large-Scale GPU Systems

Speedrun Talent Network • Greater London

Hybrid
GBP 120,000 - 180,000
Distributed ML Infrastructure Engineer for Large-Scale Models
Distributed ML Infrastructure Engineer for Large-Scale Models

Lumaai • Greater London

On-site
GBP 120,000 - 180,000
Research Scientist / Engineer – Training Infrastructure
Research Scientist / Engineer – Training Infrastructure

Luma • Greater London

On-site
GBP 147,000 - 299,000
Copy of Research Scientist / Engineer – Training Infrastructure
Copy of Research Scientist / Engineer – Training Infrastructure

Lumaai • Greater London

On-site
GBP 120,000 - 180,000
Copy of Research Scientist / Engineer – Training Infrastructure
Copy of Research Scientist / Engineer – Training Infrastructure

Speedrun Talent Network • Greater London

Hybrid
GBP 120,000 - 180,000
Senior Systems Engineer - Large-Scale AI Inference
Senior Systems Engineer - Large-Scale AI Inference

Speedrun Talent Network • Greater London

Hybrid
GBP 90,000 - 140,000
RL Systems Engineer: Scalable Post-Training Infrastructure
RL Systems Engineer: Scalable Post-Training Infrastructure

Luma • United Kingdom

On-site
GBP 150,000 - 190,000
Senior Software Engineer, Large-Scale Model Inference
Senior Software Engineer, Large-Scale Model Inference

Luma • Greater London

On-site
GBP 147,000 - 261,000
RL Infrastructure Engineer - Scale & Post-Training Systems
RL Infrastructure Engineer - Scale & Post-Training Systems

AItoolnavio • Greater London

Hybrid
GBP 120,000 - 180,000
ML Inference Systems Engineer: Scale GPU Deployments
ML Inference Systems Engineer: Scale GPU Deployments

Luma • United Kingdom

On-site
GBP 90,000 - 150,000