Distributed ML Infrastructure Engineer for Large-Scale Models

Lumaai

Greater London

On-site

GBP 120,000 - 180,000

Full time

2 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Luma seeks an engineer to build distributed systems for training large-scale multimodal models across thousands of GPUs, enabling researchers to focus on innovation atop reliable, efficient infrastructure.

You will tackle PyTorch, CUDA, and distributed training challenges, implementing advanced parallelism (FSDP, Tensor Parallel, Pipeline, Expert Parallel) and developing monitoring/tools to keep runs scalable and stable.

Qualifications

  • Extensive distributed PyTorch training on foundation-models.
  • Experience with GPU clusters, networking, and storage systems.
  • Familiarity with NCCL and MPI for distributed communication.

Responsibilities

  • Design, implement, and optimize distributed training systems for models across thousands of GPUs.
  • Research and implement advanced parallelization (FSDP, Tensor Parallel, Pipeline Parallel, Expert Parallel).
  • Build monitoring, visualization, and debugging tools for large-scale training runs.
  • Optimize training stability and resource utilization across massive clusters.

Skills

Distributed PyTorch training
GPU clusters
NCCL/MPI communications

Tools

Containerization
Orchestration
Cloud infrastructure

Job description

Luma seeks an engineer to build distributed systems for training large-scale multimodal models across thousands of GPUs, enabling researchers to focus on innovation atop reliable, efficient infrastructure.

You will tackle PyTorch, CUDA, and distributed training challenges, implementing advanced parallelism (FSDP, Tensor Parallel, Pipeline, Expert Parallel) and developing monitoring/tools to keep runs scalable and stable.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Distributed ML Systems Engineer
Senior Distributed ML Systems Engineer

Luma • Greater London

On-site
GBP 147,000 - 299,000
Senior Distributed Training Engineer — Large-Scale GPU Systems
Senior Distributed Training Engineer — Large-Scale GPU Systems

Speedrun Talent Network • Greater London

Hybrid
GBP 120,000 - 180,000
Research Scientist / Engineer – Training Infrastructure
Research Scientist / Engineer – Training Infrastructure

Luma • Greater London

On-site
GBP 147,000 - 299,000
Copy of Research Scientist / Engineer – Training Infrastructure
Copy of Research Scientist / Engineer – Training Infrastructure

Lumaai • Greater London

On-site
GBP 120,000 - 180,000
Copy of Research Scientist / Engineer – Training Infrastructure
Copy of Research Scientist / Engineer – Training Infrastructure

Speedrun Talent Network • Greater London

Hybrid
GBP 120,000 - 180,000
Senior Systems Engineer - Large-Scale AI Inference
Senior Systems Engineer - Large-Scale AI Inference

Speedrun Talent Network • Greater London

Hybrid
GBP 90,000 - 140,000
Senior Software Engineer, Large-Scale Model Inference
Senior Software Engineer, Large-Scale Model Inference

Luma • Greater London

On-site
GBP 147,000 - 261,000
Inference Systems Engineer (GPU/Cluster Scaling)
Inference Systems Engineer (GPU/Cluster Scaling)

AItoolnavio • Greater London

Hybrid
GBP 100,000 - 180,000
ML Inference Systems Engineer: Scale GPU Deployments
ML Inference Systems Engineer: Scale GPU Deployments

Luma • United Kingdom

On-site
GBP 90,000 - 150,000
Software Engineer, Inference
Software Engineer, Inference

Speedrun Talent Network • Greater London

Hybrid
GBP 90,000 - 140,000