Senior Distributed Training Engineer — Large-Scale GPU Systems

Speedrun Talent Network

Greater London

Hybrid

GBP 120,000 - 180,000

Full time

2 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Luma is building distributed training systems for large-scale multimodal models across thousands of GPUs, enabling researchers to focus on innovation on top of reliable, scalable infrastructure.

This is hard PyTorch, CUDA, and distributed-systems work — advanced parallelism, training stability, and utilization across massive clusters. It fits an engineer who's solved real problems training foundation models at scale.

Qualifications

  • Extensive distributed PyTorch training experience for foundation-model scale.
  • Deep understanding of GPU clusters, networking, and storage systems.
  • Familiarity with communication libraries (NCCL, MPI) and distributed-system optimization.

Responsibilities

  • Design, implement, and optimize efficient distributed training systems for models across thousands of GPUs.
  • Research and implement advanced parallelization (FSDP, Tensor Parallel, Pipeline Parallel, Expert Parallel).
  • Build monitoring, visualization, and debugging tools for large-scale training runs.
  • Optimize training stability, convergence, and resource utilization across massive clusters.

Skills

Distributed PyTorch
GPU clusters
NCCL/MPI
Containerization

Tools

Containerization

Job description

Luma is building distributed training systems for large-scale multimodal models across thousands of GPUs, enabling researchers to focus on innovation on top of reliable, scalable infrastructure.

This is hard PyTorch, CUDA, and distributed-systems work — advanced parallelism, training stability, and utilization across massive clusters. It fits an engineer who's solved real problems training foundation models at scale.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Distributed ML Systems Engineer
Senior Distributed ML Systems Engineer

Luma • Greater London

On-site
GBP 147,000 - 299,000
Distributed ML Infrastructure Engineer for Large-Scale Models
Distributed ML Infrastructure Engineer for Large-Scale Models

Lumaai • Greater London

On-site
GBP 120,000 - 180,000
Research Scientist / Engineer – Training Infrastructure
Research Scientist / Engineer – Training Infrastructure

Luma • Greater London

On-site
GBP 147,000 - 299,000
Copy of Research Scientist / Engineer – Training Infrastructure
Copy of Research Scientist / Engineer – Training Infrastructure

Lumaai • Greater London

On-site
GBP 120,000 - 180,000
Copy of Research Scientist / Engineer – Training Infrastructure
Copy of Research Scientist / Engineer – Training Infrastructure

Speedrun Talent Network • Greater London

Hybrid
GBP 120,000 - 180,000
Senior Software Engineer, Large-Scale Model Inference
Senior Software Engineer, Large-Scale Model Inference

Luma • Greater London

On-site
GBP 147,000 - 261,000
Senior Systems Engineer - Large-Scale AI Inference
Senior Systems Engineer - Large-Scale AI Inference

Speedrun Talent Network • Greater London

Hybrid
GBP 90,000 - 140,000
RL Systems Engineer: Scalable Post-Training Infrastructure
RL Systems Engineer: Scalable Post-Training Infrastructure

Luma • United Kingdom

On-site
GBP 150,000 - 190,000
ML Inference Systems Engineer: Scale GPU Deployments
ML Inference Systems Engineer: Scale GPU Deployments

Luma • United Kingdom

On-site
GBP 90,000 - 150,000
Inference Systems Engineer (GPU/Cluster Scaling)
Inference Systems Engineer (GPU/Cluster Scaling)

AItoolnavio • Greater London

Hybrid
GBP 100,000 - 180,000