A complete application in a minute — tailored resume and cover letter, ready to send.
Luma is building distributed training systems for its large-scale multimodal models. This role focuses on PyTorch, CUDA, and advanced parallelism across thousands of GPUs, delivering reliable, scalable infrastructure so researchers can innovate.
You will design and optimize training systems, implement FSDP, Tensor Parallel, Pipeline Parallel, and Expert Parallel, and build monitoring and debugging tools to improve stability and utilization.
You'll build the distributed systems that train Luma's large-scale multimodal models across thousands of GPUs, so researchers can focus on innovation on top of reliable, efficient, scalable infrastructure.
This is hard PyTorch, CUDA, and distributed-systems work — advanced parallelism, training stability, and utilization across massive clusters. It fits an engineer who's solved real problems training foundation models at scale. If you haven't worked at the level of FSDP and multi-node training, this is the wrong depth.
One way the first 90 could unfold.
About Luma: Luma's mission is to build unified general intelligence that can generate, understand, and operate in the physical world. We believe multimodality is critical for intelligence — the next step beyond language models comes from vision. Luma is an equal opportunity employer.
Compensation Range: $195K - $395K