GPU Transformer Performance Engineer (Triton/CUDA)

Luma AI

United States

Remote

USD 180,000 - 280,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Luma AI seeks an expert GPU optimization engineer to accelerate multimodal models. You will profile and optimize GPU/CPU/accelerator code for maximal utilization and minimal latency, writing high-performance PyTorch, Triton, and CUDA kernels and developing fused kernels for tensor cores.

Drive distributed multi-node deployments, build monitoring tooling, and apply state-of-the-art transformer optimization techniques.

Qualifications

  • Expert-level Triton/CUDA programming and GPU optimization.
  • Strong PyTorch skills, including kernel development and custom operations.
  • Proficiency with profiling tools (NVIDIA Nsight, torch profiler, custom tooling).
  • Deep understanding of transformer architectures and attention mechanisms.

Responsibilities

  • Profile and optimize GPU/CPU/accelerator code for maximum utilization and minimal latency.
  • Write high-performance PyTorch, Triton, and CUDA, dropping to custom operations when needed.
  • Develop fused kernels and leverage tensor cores and modern hardware features across platforms.
  • Optimize model architectures and implementations for distributed multi-node production deployment.
  • Build performance monitoring and analysis tools and automation.
  • Research and implement cutting-edge optimization techniques for transformer models.

Job description

Luma AI seeks an expert GPU optimization engineer to accelerate multimodal models. You will profile and optimize GPU/CPU/accelerator code for maximal utilization and minimal latency, writing high-performance PyTorch, Triton, and CUDA kernels and developing fused kernels for tensor cores.

Drive distributed multi-node deployments, build monitoring tooling, and apply state-of-the-art transformer optimization techniques.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

GPU Inference Performance Engineer — Equity & Optimization
GPU Inference Performance Engineer — Equity & Optimization

Nvidia Corporation • Santa Clara (CA)

On-site
USD 152,000 - 242,000
GPU ML Infra Intern: Speed Up Training & Profiling
GPU ML Infra Intern: Speed Up Training & Profiling

Plus 2 • Santa Clara (CA)

On-site
USD 150,000 - 210,000
Free lunch
Snacks & drinks
401(k) plan
Research Engineer, GPU Performance
Research Engineer, GPU Performance

Harnham • California (MO)

On-site
USD 120,000 - 160,000
Senior GPU Kernel Engineer - Hybrid Triton AI
Senior GPU Kernel Engineer - Hybrid Triton AI

Advanced Micro Devices • San Jose (CA)

Hybrid
USD 180,000 - 280,000
GPU Systems Engineer — Distributed Training & Inference
GPU Systems Engineer — Distributed Training & Inference

TensorScale AI • San Francisco (CA)

On-site
USD 180,000 - 240,000
TPU Performance Engineer — ML Efficiency & Scale
TPU Performance Engineer — ML Efficiency & Scale

Google • Sunnyvale (CA)

On-site
USD 207,000 - 300,000
GPU Performance Engineer | Experienced Hire
GPU Performance Engineer | Experienced Hire

SIG Susquehanna • Pennsylvania

On-site
USD 120,000 - 150,000
GPU Performance Engineer | Experienced Hire
GPU Performance Engineer | Experienced Hire

Susquehanna International Group • Bala Cynwyd (PA)

On-site
USD 100,000 - 130,000
Senior DL Performance Engineer - LLM/Transformer Optimizations
Senior DL Performance Engineer - LLM/Transformer Optimizations

NVIDIA AI • Santa Clara (CA)

On-site
USD 152,000 - 288,000
Equity
Benefits
Remote GPU Kernel Engineer for AI/ML (PyTorch/Triton)
Remote GPU Kernel Engineer for AI/ML (PyTorch/Triton)

NLB Services • United States

On-site
USD 140,000 - 230,000