GPU Kernel & Transformer Performance Scientist

Lumaai

Greater London

On-site

GBP 120,000 - 180,000

Full time

2 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Luma seeks a seasoned performance engineer to accelerate multimodal models across GPU, CPU, and accelerators. You will write high‑performing PyTorch, Triton, and CUDA kernels and push hardware to the limit while preserving model quality.

You’ll own profiling, optimization, and deployment at scale, develop fused kernels, tensor-core aware code, and build monitoring tools for distributed training and inference.

Qualifications

  • Expert-level Triton/CUDA programming and GPU optimization.
  • Strong PyTorch skills, including kernel development and custom operations.
  • Proficiency with profiling tools (NVIDIA Nsight, torch profiler, custom tooling).
  • Deep understanding of transformer architectures and attention mechanisms.
  • Nice to Have: experience with compilers and exporters (torch.compile, TensorRT, ONNX, XLA).

Responsibilities

  • Profile and optimize GPU/CPU/accelerator code for maximum utilization and minimal latency.
  • Write high-performance PyTorch, Triton, and CUDA, dropping to custom operations when needed.
  • Develop fused kernels and leverage tensor cores and modern hardware features across platforms.
  • Optimize model architectures and implementations for distributed multi-node production deployment.
  • Build performance monitoring and analysis tools and automation.
  • Research and implement cutting‑edge optimization techniques for transformer models.

Skills

Triton/CUDA programming
GPU optimization
PyTorch kernel development
Profiling tools

Tools

NVIDIA Nsight
Torch profiler
Custom tooling

Job description

Luma seeks a seasoned performance engineer to accelerate multimodal models across GPU, CPU, and accelerators. You will write high‑performing PyTorch, Triton, and CUDA kernels and push hardware to the limit while preserving model quality.

You’ll own profiling, optimization, and deployment at scale, develop fused kernels, tensor-core aware code, and build monitoring tools for distributed training and inference.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

GPU Performance Engineer for Multimodal Transformers
GPU Performance Engineer for Multimodal Transformers

AItoolnavio • Greater London

Hybrid
GBP 90,000 - 130,000
Research Scientist / Engineer – Performance Optimization
Research Scientist / Engineer – Performance Optimization

AItoolnavio • Greater London

Hybrid
GBP 90,000 - 130,000
Copy of Research Scientist / Engineer – Performance Optimization
Copy of Research Scientist / Engineer – Performance Optimization

Lumaai • Greater London

On-site
GBP 120,000 - 180,000
Performance Engineer (GPU)
Performance Engineer (GPU)

Anthropic • York and North Yorkshire

On-site
GBP 90,000 - 140,000
Comprehensive health insurance
Fertility benefits
22 weeks parental leave
+1
Principal ML Performance Engineer - GPU Optimization
Principal ML Performance Engineer - GPU Optimization

Sponsor Finder • Boston

On-site
GBP 113,000 - 159,000
Performance Kernel Engineer, GPU & ML Inference
Performance Kernel Engineer, GPU & ML Inference

INFERACT SINGAPORE PTE. LTD. • Penarth

On-site
GBP 99,000 - 198,000
Equity
Medical insurance
Dental insurance
+1
Member of Technical Staff, ML Performance
Member of Technical Staff, ML Performance

Odyssey • Greater London

On-site
GBP 70,000 - 90,000
Senior Distributed ML Systems Engineer
Senior Distributed ML Systems Engineer

Luma • Greater London

On-site
GBP 147,000 - 299,000
Distributed ML Infrastructure Engineer for Large-Scale Models
Distributed ML Infrastructure Engineer for Large-Scale Models

Lumaai • Greater London

On-site
GBP 120,000 - 180,000
Research Scientist / Engineer – Training Infrastructure
Research Scientist / Engineer – Training Infrastructure

Luma • Greater London

On-site
GBP 147,000 - 299,000