An application made for this job — a tailored resume and cover letter that speak straight to the posting.
Luma seeks an engineer to build distributed systems for training large-scale multimodal models across thousands of GPUs, enabling researchers to focus on innovation atop reliable, efficient infrastructure.
You will tackle PyTorch, CUDA, and distributed training challenges, implementing advanced parallelism (FSDP, Tensor Parallel, Pipeline, Expert Parallel) and developing monitoring/tools to keep runs scalable and stable.
Luma seeks an engineer to build distributed systems for training large-scale multimodal models across thousands of GPUs, enabling researchers to focus on innovation atop reliable, efficient infrastructure.
You will tackle PyTorch, CUDA, and distributed training challenges, implementing advanced parallelism (FSDP, Tensor Parallel, Pipeline, Expert Parallel) and developing monitoring/tools to keep runs scalable and stable.