GPU Performance Engineer for Multimodal Transformers

AItoolnavio

Greater London

Hybrid

GBP 90,000 - 130,000

Full time

2 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Luma is seeking an expert GPU optimization engineer to accelerate multimodal transformer workloads. You will profile, tune, and deploy kernel- and operator-level improvements across GPUs, CPUs, and accelerators to maximize training throughput and minimize latency.

You will implement fused kernels, leverage tensor cores, and push Triton/CUDA capabilities, while collaborating on distributed multi-node deployment and performance instrumentation.

Qualifications

  • Expert-level Triton/CUDA programming and GPU optimization.
  • Strong PyTorch skills, including kernel development and custom operations.
  • Proficiency with profiling tools (NVIDIA Nsight, torch profiler, custom tooling).
  • Deep understanding of transformer architectures and attention mechanisms.

Responsibilities

  • Profile and optimize GPU/CPU/accelerator code for maximum utilization and minimal latency.
  • Write high-performance PyTorch, Triton, and CUDA, dropping to custom operations when needed.
  • Develop fused kernels and leverage tensor cores and modern hardware features across platforms.
  • Optimize model architectures and implementations for distributed multi-node production deployment.
  • Build performance monitoring and analysis tools and automation.
  • Research and implement cutting-edge optimization techniques for transformer models.
  • Scale and systemize performance improvements for long-term reliability.

Skills

Triton/CUDA programming
PyTorch kernel development
GPU optimization
Profiling tools

Tools

NVIDIA Nsight
Torch profiler
Custom tooling

Job description

Luma is seeking an expert GPU optimization engineer to accelerate multimodal transformer workloads. You will profile, tune, and deploy kernel- and operator-level improvements across GPUs, CPUs, and accelerators to maximize training throughput and minimize latency.

You will implement fused kernels, leverage tensor cores, and push Triton/CUDA capabilities, while collaborating on distributed multi-node deployment and performance instrumentation.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Research Scientist / Engineer – Performance Optimization
Research Scientist / Engineer – Performance Optimization

AItoolnavio • Greater London

Hybrid
GBP 90,000 - 130,000
Principal ML Performance Engineer - GPU Optimization
Principal ML Performance Engineer - GPU Optimization

Sponsor Finder • Boston

On-site
GBP 113,000 - 159,000
Performance Engineer (GPU)
Performance Engineer (GPU)

Anthropic • York and North Yorkshire

On-site
GBP 90,000 - 140,000
Comprehensive health insurance
Fertility benefits
22 weeks parental leave
+1
Senior Distributed Training Engineer — Large-Scale GPU Systems
Senior Distributed Training Engineer — Large-Scale GPU Systems

Speedrun Talent Network • Greater London

Hybrid
GBP 120,000 - 180,000
Senior CUDA Engineer - High-Performance GPU Inference
Senior CUDA Engineer - High-Performance GPU Inference

Fuse Energy, LLC • Greater London

On-site
GBP 120,000 - 180,000
Equity sign-on bonus
Fully expensed tech equipment
Breakfast and dinner allowance
+1
GPU Performance Engineer: Scale Inference & Training
GPU Performance Engineer: Scale Inference & Training

Anthropic • York and North Yorkshire

On-site
GBP 90,000 - 140,000
Comprehensive health insurance
Fertility benefits
22 weeks parental leave
+1
ML Inference Systems Engineer: Scale GPU Deployments
ML Inference Systems Engineer: Scale GPU Deployments

Luma • United Kingdom

On-site
GBP 90,000 - 150,000
Member of Technical Staff, ML Performance
Member of Technical Staff, ML Performance

Odyssey • Greater London

On-site
GBP 70,000 - 90,000
Senior GPU AI/ML Performance Engineer
Senior GPU AI/ML Performance Engineer

Google LLC • Greater London

On-site
GBP 90,000 - 150,000
ML Performance Engineer (GPU Optimization)
ML Performance Engineer (GPU Optimization)

Sponsor Finder • Boston

On-site
GBP 113,000 - 159,000