Principal ML Performance Engineer - GPU Optimization

Sponsor Finder

Boston

Presencial

GBP 113.000 - 159.000

Jornada completa

Hace 3 días
Sé de los primeros/as/es en solicitar esta vacante
Generador de candidaturas

Destaca para este puesto — genera un currículum y una carta de presentación adaptados en cuestión de un minuto.

Supera los filtros ATS

Descripción de la vacante

Proxima is seeking a Principal ML Performance Engineer to optimize training and inference for structural and generative models, with a focus on CUDA, Triton, and distributed training across large GPU clusters. You will profile bottlenecks, write efficient kernels, and drive multi-node scalability.

As a senior leader, you will scale training across 32–64 nodes, reduce inference costs on thousands of GPUs, and mentor engineers while defining technical direction for the team.

Formación

  • 6+ years of experience in ML systems, HPC, or performance engineering.
  • Ability to set technical direction beyond coding and mentor engineers.
  • Deep knowledge of PyTorch internals with profiling experience.
  • Experience with CUDA and Triton and reading Nsight outputs.
  • Experience with distributed training at multi-node scale.
  • Strong Python and C++ proficiency.
  • Ability to name a model they made faster and quantify the improvement.

Responsabilidades

  • Profile and optimize training and inference for models including transformers, diffusion, and geometric deep learning.
  • Write and tune custom kernels (CUDA, Triton) and use compilers (torch.compile, TensorRT, XLA).
  • Scale distributed training across 32-64 nodes with FSDP, DeepSpeed, and mixed precision.
  • Reduce inference cost by memory optimization for large complexes and batching inputs.
  • Manage GPU cluster efficiency on GCP with scheduling and cost reporting.
  • Develop benchmarks and profiling tools for the research team.

Conocimientos

PyTorch internals
CUDA
Triton
Python
C++
Distributed training

Educación

BS/MS/PhD in CS, EE, or related field

Herramientas

Nsight
TensorRT
Torch.compile
XLA
GCP

Descripción del empleo

Proxima is seeking a Principal ML Performance Engineer to optimize training and inference for structural and generative models, with a focus on CUDA, Triton, and distributed training across large GPU clusters. You will profile bottlenecks, write efficient kernels, and drive multi-node scalability.

As a senior leader, you will scale training across 32–64 nodes, reduce inference costs on thousands of GPUs, and mentor engineers while defining technical direction for the team.

Consigue la evaluación confidencial y gratuita de tu currículum.

o arrastra y suelta tu archivo aquí

Similar jobs

Puestos de trabajo similares que vale la pena comparar

ML Performance Engineer (GPU Optimization)
ML Performance Engineer (GPU Optimization)

Sponsor Finder • Boston

Presencial
GBP 113.000 - 159.000
Performance Engineer (GPU)
Performance Engineer (GPU)

Anthropic • York and North Yorkshire

Presencial
GBP 90.000 - 140.000
Comprehensive health insurance
Fertility benefits
22 weeks parental leave
+1
GPU Performance Engineer for Multimodal Transformers
GPU Performance Engineer for Multimodal Transformers

AItoolnavio • Greater London

Híbrido
GBP 90.000 - 130.000
Remote Performance Engineer: ML Training & Kernels
Remote Performance Engineer: ML Training & Kernels

Cohere • Greater London

Presencial
GBP 75.000 - 95.000
Co-working benefit
Daily lunch program
Regular community and social events
Member of Technical Staff, ML Performance
Member of Technical Staff, ML Performance

Odyssey • Greater London

Presencial
GBP 70.000 - 90.000
GPU Performance Engineer: Scale Inference & Training
GPU Performance Engineer: Scale Inference & Training

Anthropic • York and North Yorkshire

Presencial
GBP 90.000 - 140.000
Comprehensive health insurance
Fertility benefits
22 weeks parental leave
+1
Senior GPU AI/ML Performance Engineer
Senior GPU AI/ML Performance Engineer

Google LLC • Greater London

Presencial
GBP 90.000 - 150.000
Machine Learning Performance Engineer
Machine Learning Performance Engineer

Quant Blueprint LLC • Greater London

Presencial
GBP 50.000 - 70.000
ML Systems Performance Engineer
ML Systems Performance Engineer

Quant Blueprint LLC • Greater London

Presencial
GBP 50.000 - 70.000
Performance Engineer
Performance Engineer

Anthropic • York and North Yorkshire

Presencial
GBP 110.000 - 150.000
Health insurance
Fertility benefits
Parental leave 22 weeks
+12