ML Performance Engineer (GPU Optimization)

Sponsor Finder

Boston

Presencial

GBP 113 000 - 159 000

Tempo integral

Há 4 dias
Torna-te num dos primeiros candidatos
Gerador de candidaturas

Não envies um currículo genérico — gera um currículo e uma carta de apresentação adaptados a esta função específica.

Ultrapassa os filtros ATS

Resumo da oferta

Proxima is seeking a Principal ML Performance Engineer to optimize training and inference for structural and generative models, with a focus on CUDA, Triton, and distributed training across large GPU clusters. You will profile bottlenecks, write efficient kernels, and drive multi-node scalability.

As a senior leader, you will scale training across 32–64 nodes, reduce inference costs on thousands of GPUs, and mentor engineers while defining technical direction for the team.

Qualificações

  • 6+ years of experience in ML systems, HPC, or performance engineering.
  • Ability to set technical direction beyond coding and mentor engineers.
  • Deep knowledge of PyTorch internals with profiling experience.
  • Experience with CUDA and Triton and reading Nsight outputs.
  • Experience with distributed training at multi-node scale.
  • Strong Python and C++ proficiency.
  • Ability to name a model they made faster and quantify the improvement.

Responsabilidades

  • Profile and optimize training and inference for models including transformers, diffusion, and geometric deep learning.
  • Write and tune custom kernels (CUDA, Triton) and use compilers (torch.compile, TensorRT, XLA).
  • Scale distributed training across 32-64 nodes with FSDP, DeepSpeed, and mixed precision.
  • Reduce inference cost by memory optimization for large complexes and batching inputs.
  • Manage GPU cluster efficiency on GCP with scheduling and cost reporting.
  • Develop benchmarks and profiling tools for the research team.

Conhecimentos

PyTorch internals
CUDA
Triton
Python
C++
Distributed training

Formação académica

BS/MS/PhD in CS, EE, or related field

Ferramentas

Nsight
TensorRT
Torch.compile
XLA
GCP

Descrição da oferta de emprego

Principal ML Performance Engineer (GPU Optimization)
About Proxima

Proxima is a frontier AI and data generation company discovering the next generation of proximity therapeutics by making protein interactions programmable. Our platform brings together foundation-model machine learning, a scalable data generation engine, and a partnership track record exceeding $5B in collaborations across the world's leading biopharma and tech organizations. We've recently closed an oversubscribed seed round with an elite group of VCs including DCVC, NVIDIA's NVentures, AIX, Yosemite among others.

Neo-1 is our all-atom foundation model that combines state-of-the-art structure prediction and molecular generation in a single system. Neo-1 enables rapid exploration of chemical and structural space for high value, previously intractable targets, and in particular unlocks small molecule proximity therapeutics like molecular glues with AI for the first time.

In parallel, we are developing an advanced structural interactomics platform built on proprietary XLMS technology and a lab equipped with next-generation mass spectrometry instrumentation. This platform produces proteome-scale maps of protein interactions and helps identify small molecules that modulate proximity. Together with Neo-1, it creates an integrated system capable of co-folding protein complexes while generating candidate small molecules to influence those interactions.
Proximity-based therapeutics represent one of the most promising frontiers in modern drug discovery with the potential to treat previously intractable diseases and target 'undruggable' proteins. We're building the tech and the team to make that happen. Come join us!

What you'll do
  • Profile and optimize training and inference for structural and generative models, including transformers, diffusion, and geometric deep learning

  • Write and tune custom kernels (CUDA, Triton) and use compilers (torch.compile, TensorRT, XLA) when beneficial

  • Scale distributed training across 32-64 nodes, employing FSDP, DeepSpeed, tensor and pipeline parallelism, and mixed precision

  • Reduce inference cost by optimizing memory scaling for large complexes, improving diffusion sampling efficiency, batching ragged inputs, and maximizing throughput across up to 1000 GPUs

  • Manage GPU cluster efficiency on GCP, focusing on scheduling, utilization, spot strategy, and cost reporting

  • Develop benchmarks and profiling tools for the research team

What we need
  • Minimum of 6+ years experience in ML systems, HPC, or performance engineering, with a BS/MS/PhD in CS, EE, or related field

  • Demonstrated ability to set technical direction beyond coding: selecting infrastructure, influencing research teams, and mentoring engineers

  • Deep knowledge of PyTorch internals with hands-on experience profiling and fixing real bottlenecks

  • Experience with CUDA and Triton, skilled at reading Nsight output, and strong understanding of memory bandwidth and occupancy

  • Experience with distributed training at multi-node scale

  • Strong proficiency in Python and C++

  • Able to name a model they made materially faster and quantify the improvement

Nice to haves
  • Experience in geometric deep learning, equivariant networks, or protein structure models such as AlphaFold, ESM, or RFdiffusion

  • Experience writing kernels for structure-model primitives, including triangle attention, triangle multiplicative updates, cuEquivariance, or FlashAttention for pair bias

  • Experience orchestrating large batch inference and managing Kubernetes GPU scheduling

Obtém a tua avaliação gratuita e confidencial do currículo.

ou arrasta e larga o ficheiro aqui.

Similar jobs

Ofertas semelhantes que vale a pena comparar

Member of Technical Staff
Member of Technical Staff

Sponsor Finder • Boston

Presencial
GBP 90 000 - 130 000
Performance Engineer (GPU)
Performance Engineer (GPU)

Anthropic • York and North Yorkshire

Presencial
GBP 90 000 - 140 000
Comprehensive health insurance
Fertility benefits
22 weeks parental leave
+1
Member of Technical Staff, ML Performance
Member of Technical Staff, ML Performance

Odyssey • Greater London

Presencial
GBP 70 000 - 90 000
Senior ML Infrastructure Engineer (Research Initiatives) - Systems Integrator
Senior ML Infrastructure Engineer (Research Initiatives) - Systems Integrator

Hamilton Barnes Associates Limited • Grã-Bretanha

Presencial
GBP 90 000 - 130 000
Significant stock option packages
Remote-first working setup
Fully paid travel and accommodation
+1
Performance Engineer
Performance Engineer

Anthropic • York and North Yorkshire

Presencial
GBP 110 000 - 150 000
Health insurance
Fertility benefits
Parental leave 22 weeks
+12
Machine Learning Performance Engineer
Machine Learning Performance Engineer

Quant Blueprint LLC • Greater London

Presencial
GBP 50 000 - 70 000
Member of Technical Staff (AI Inference Engineer)
Member of Technical Staff (AI Inference Engineer)

CVFine by Instrovate Technologies • Greater London

Presencial
GBP 70 000 - 90 000
Member of Technical Staff (AI Inference Engineer)
Member of Technical Staff (AI Inference Engineer)

Perplexity • Greater London

Presencial
GBP 80 000 - 120 000
Equity
Senior Performance Engineer | AI Infrastructure | Cambridge (Hybrid)
Senior Performance Engineer | AI Infrastructure | Cambridge (Hybrid)

Pure Resourcing Solutions Limited • Linton

Presencial
GBP 90 000 - 120 000
Founding GPU Engineer
Founding GPU Engineer

Fuse Energy • Greater London

Presencial
GBP 90 000 - 130 000
Equity sign-on bonus
Biannual bonus
Fully expensed tech
+1