Turn this role into an interview — a resume and cover letter built around what this employer wants.
Luma AI is hiring for a Performance Optimization role focused on making multimodal models faster and more scalable. You will work closely with research and engineering to optimize training and inference, write high-performance kernels, and push transformer models toward lower latency and higher throughput.
The role emphasizes deep CUDA/Triton expertise, PyTorch proficiency, and the ability to deploy efficient, production-grade optimizations across hardware platforms.
Luma's mission is to build multimodal AI to expand human imagination and capabilities. We believe that multimodality is critical for intelligence. To go beyond language models and build more aware, capable and useful systems, the next step function change will come from vision. So we are working on training and scaling up multimodal foundation models for systems that can see and understand, show and explain, and eventually interact with our world to effect change.
The Performance Optimization team at Luma is dedicated to maximizing the efficiency and performance of our AI models. Working closely with both research and engineering teams, this group ensures that our cutting-edge multimodal models can be trained efficiently and deployed at scale while maintaining the highest quality standards.
Your applications are reviewed by real people.
The base pay range for this role is $187,500 - $395,000 per year.