Research Scientist / Engineer – Performance Optimization

lumalabs-ai

San Francisco (CA)

On-site

USD 237,000 - 395,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Luma AI is hiring for a Performance Optimization role focused on making multimodal models faster and more scalable. You will work closely with research and engineering to optimize training and inference, write high-performance kernels, and push transformer models toward lower latency and higher throughput.

The role emphasizes deep CUDA/Triton expertise, PyTorch proficiency, and the ability to deploy efficient, production-grade optimizations across hardware platforms.

Qualifications

  • Expert-level Triton/CUDA programming experience.
  • Strong PyTorch skills and ability to develop custom operations.
  • Experience with PyTorch kernel development.
  • Proficiency with profiling tools (NVIDIA Nsight, torch profiler, custom tooling).
  • Deep understanding of transformer architectures and attention mechanisms.
  • Preferred: experience with torch.compile, TensorRT, ONNX, XLA.

Responsibilities

  • Profile and optimize GPU/CPU/Accelerator code for maximum utilization and minimal latency.
  • Write high-performance PyTorch, Triton, CUDA, deferring to custom PyTorch operations if necessary.
  • Develop fused kernels and leverage tensor cores and modern hardware features for optimal hardware utilization on different hardware platforms.
  • Optimize model architectures and implementations for distributed multi-node production deployment.
  • Build performance monitoring and analysis tools and automation.
  • Research and implement cutting-edge optimization techniques for transformer model

Skills

Triton/CUDA programming
PyTorch
PyTorch kernel development
Profiling tools
Transformer architectures

Tools

NVIDIA Nsight
torch profiler
custom tooling

Job description

About Luma AI

Luma's mission is to build multimodal AI to expand human imagination and capabilities. We believe that multimodality is critical for intelligence. To go beyond language models and build more aware, capable and useful systems, the next step function change will come from vision. So we are working on training and scaling up multimodal foundation models for systems that can see and understand, show and explain, and eventually interact with our world to effect change.

About the Role

The Performance Optimization team at Luma is dedicated to maximizing the efficiency and performance of our AI models. Working closely with both research and engineering teams, this group ensures that our cutting-edge multimodal models can be trained efficiently and deployed at scale while maintaining the highest quality standards.

Responsibilities
  • Profile and optimize GPU/CPU/Accelerator code for maximum utilization and minimal latency
  • Write high-performance PyTorch, Triton, CUDA, deferring to custom PyTorch operations if necessary
  • Develop fused kernels and leverage tensor cores and modern hardware features for optimal hardware utilization on different hardware platforms
  • Optimize model architectures and implementations for distributed multi-node production deployment
  • Build performance monitoring and analysis tools and automation
  • Research and implement cutting-edge optimization techniques for transformer model
Experience
  • Expert-level proficiency in Triton/CUDA programming and GPU optimization
  • Strong PyTorch skills
  • Experience with PyTorch kernel development and custom operations
  • Proficiency with profiling tools (NVIDIA Nsight, torch profiler, custom tooling)
  • Deep understanding of transformer architectures and attention mechanisms
  • (Preferred) Experience with compilers/exporters such as torch.compile, TensorRT, ONNX, XLA
  • (Preferred) Experience optimizing inference workloads for latency and throughput
  • (Preferred) Experience with Triton compiler and kernel fusion techniques
  • (Preferred) Knowledge of warp-level intrinsics and advanced CUDA optimization

Your applications are reviewed by real people.

Compensation

The base pay range for this role is $187,500 - $395,000 per year.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior AI Performance Engineer - GPU & Triton Optimizations
Senior AI Performance Engineer - GPU & Triton Optimizations

lumalabs-ai • San Francisco (CA)

On-site
USD 237,000 - 395,000
Research Scientist / Engineer – Training Infrastructure
Research Scientist / Engineer – Training Infrastructure

lumalabs-ai • San Francisco (CA)

On-site
USD 188,000 - 395,000
Applied Research Scientist / Engineer
Applied Research Scientist / Engineer

lumalabs-ai • New York (NY)

On-site
USD 200,000 - 450,000
GPU Transformer Performance Engineer (Triton/CUDA)
GPU Transformer Performance Engineer (Triton/CUDA)

Luma AI • United States

Remote
USD 180,000 - 280,000
Research Scientist / Engineer — Multimodal Agent
Research Scientist / Engineer — Multimodal Agent

lumalabs-ai • San Francisco (CA)

On-site
USD 250,000 - 450,000
Research Scientist / Engineer – Training Infrastructure
Research Scientist / Engineer – Training Infrastructure

Luma AI • San Francisco (CA)

Hybrid
USD 187,000 - 395,000
Research Scientist / Engineer – Foundation Model: Core Research
Research Scientist / Engineer – Foundation Model: Core Research

lumalabs-ai • San Francisco (CA)

On-site
USD 250,000 - 450,000
Research Scientist / Engineer - Controllability, Personalization & Productization
Research Scientist / Engineer - Controllability, Personalization & Productization

lumalabs-ai • New York (NY)

On-site
USD 200,000 - 450,000
Research Engineer - Evaluations
Research Engineer - Evaluations

lumalabs-ai • New York (NY)

On-site
USD 190,000 - 375,000
Research Engineer, GPU Performance
Research Engineer, GPU Performance

Harnham • California (MO)

On-site
USD 120,000 - 160,000