Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.
Luma is building distributed training systems for its large-scale multimodal models. This role focuses on PyTorch, CUDA, and advanced parallelism across thousands of GPUs, delivering reliable, scalable infrastructure so researchers can innovate.
You will design and optimize training systems, implement FSDP, Tensor Parallel, Pipeline Parallel, and Expert Parallel, and build monitoring and debugging tools to improve stability and utilization.
Luma is building distributed training systems for its large-scale multimodal models. This role focuses on PyTorch, CUDA, and advanced parallelism across thousands of GPUs, delivering reliable, scalable infrastructure so researchers can innovate.
You will design and optimize training systems, implement FSDP, Tensor Parallel, Pipeline Parallel, and Expert Parallel, and build monitoring and debugging tools to improve stability and utilization.