Senior ML Systems Engineer - GPU & HPC Optimizations

Sponsor Finder

Boston

Sur place

GBP 120 000 - 180 000

Plein temps

14 jours+
Générateur de candidature

Obtenez une réponse de cet employeur — un CV et une lettre de motivation adaptés exactement à ce qu’il recherche.

Passez les filtres ATS

Résumé du poste

Proxima is seeking a Principal ML Performance Engineer to optimize training and inference for state-of-the-art models. You will profile PyTorch, write custom kernels in CUDA/Triton, and leverage compilers like torch.compile, TensorRT, and XLA to maximize throughput on large GPU clusters.

You\'ll scale distributed training across 32–64 nodes on GCP, manage memory scaling for large complexes, and build benchmarks and profiling tools for the research team.

Qualifications

  • 6+ years in ML systems, HPC, or performance engineering.
  • Deep knowledge of PyTorch internals and profiling bottlenecks.
  • Experience with CUDA and Triton; reading Nsight outputs.
  • Distributed multi-node training experience.
  • Strong proficiency in Python and C++.
  • Ability to name a model they made materially faster and quantify the improvement.

Responsabilités

  • Profile and optimize training and inference for structural and generative models.
  • Write and tune custom kernels (CUDA, Triton) and use compilers (torch.compile, TensorRT, XLA).
  • Scale distributed training across 32–64 nodes with advanced parallelism and mixed precision.
  • Reduce inference cost by memory scaling and throughput improvements for large GPU workloads.
  • Manage GPU cluster efficiency on cloud, including scheduling and cost reporting.
  • Develop benchmarks and profiling tools for the research team.

Connaissances

ML systems
HPC
Performance engineering
PyTorch internals
CUDA
Triton
Python
C++
Distributed training

Formation

BS/MS/PhD in CS/EE or related field

Outils

Nsight
TensorRT
XLA
torch.compile

Description du poste

Proxima is seeking a Principal ML Performance Engineer to optimize training and inference for state-of-the-art models. You will profile PyTorch, write custom kernels in CUDA/Triton, and leverage compilers like torch.compile, TensorRT, and XLA to maximize throughput on large GPU clusters.

You\'ll scale distributed training across 32–64 nodes on GCP, manage memory scaling for large complexes, and build benchmarks and profiling tools for the research team.

Obtenez votre examen gratuit et confidentiel de votre CV.

ou faites glisser et déposez votre fichier ici.

Similar jobs

Postes similaires à comparer

Principal ML Performance Engineer (GPU Optimization)
Principal ML Performance Engineer (GPU Optimization)

Sponsor Finder • Boston

Sur place
GBP 120 000 - 180 000
Performance Engineer (GPU)
Performance Engineer (GPU)

Anthropic • York and North Yorkshire

Sur place
GBP 90 000 - 140 000
Comprehensive health insurance
Fertility benefits
22 weeks parental leave
+1
Senior ML Systems Engineer — C++ / PyTorch & Performance
Senior ML Systems Engineer — C++ / PyTorch & Performance

CamWebDir • Grande-Bretagne

Hybride
GBP 75 000 - 110 000
Training / AI Infrastructure Engineering & Research London
Training / AI Infrastructure Engineering & Research London

Genesis • Greater London

Hybride
GBP 90 000 - 130 000
Member of Technical Staff, ML Performance
Member of Technical Staff, ML Performance

Odyssey • Greater London

Sur place
GBP 70 000 - 90 000
Remote Performance Engineer: ML Training & Kernels
Remote Performance Engineer: ML Training & Kernels

Cohere • Greater London

Sur place
GBP 75 000 - 95 000
Co-working benefit
Daily lunch program
Regular community and social events
ML Systems Performance Engineer
ML Systems Performance Engineer

Quant Blueprint LLC • Greater London

Sur place
GBP 50 000 - 70 000
Senior ML Engineer - AI Hardware & Scale
Senior ML Engineer - AI Hardware & Scale

EngineersOfAI • Bristol

Sur place
GBP 90 000 - 140 000
Machine Learning Performance Engineer
Machine Learning Performance Engineer

Quant Blueprint LLC • Greater London

Sur place
GBP 50 000 - 70 000
Performance Engineer
Performance Engineer

Anthropic • York and North Yorkshire

Sur place
GBP 110 000 - 150 000
Health insurance
Fertility benefits
Parental leave 22 weeks
+12