Get more replies from employers
Send a job-specific resume in minutes.
Luma is seeking an engineer to design and optimize distributed training systems for multimodal models across thousands of GPUs. You will tackle advanced parallelism, training stability, and utilization at scale.
You will work with PyTorch, CUDA, NCCL, and Tensor Parallel, building monitoring and tooling for large-scale runs. This role suits someone who has shipped foundation-model training at scale and can improve stability and efficiency on massive clusters.
Luma is seeking an engineer to design and optimize distributed training systems for multimodal models across thousands of GPUs. You will tackle advanced parallelism, training stability, and utilization at scale.
You will work with PyTorch, CUDA, NCCL, and Tensor Parallel, building monitoring and tooling for large-scale runs. This role suits someone who has shipped foundation-model training at scale and can improve stability and efficiency on massive clusters.