Senior ML Systems Engineer – GPU HPC & Distributed Training

Proxima

Boston (MA)

On-site

USD 180,000 - 280,000

Full time

8 days ago
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Proxima is building frontier AI for proximity therapeutics, combining foundation-model ML with a scalable data generation engine. We are seeking a senior ML systems engineer to profile and optimize training and inference for large structural models, implement CUDA/Triton kernels, and scale multi-node distributed training across large GPU clusters.

Ideal candidates will have 6+ years in ML systems or HPC, deep PyTorch internals knowledge, and strong Python/C++ skills.

Qualifications

  • BS/MS/PhD in CS, EE, or related field
  • 6+ years in ML systems, HPC, or performance engineering
  • Deep knowledge of PyTorch internals with profiling experience
  • Experience with CUDA and Triton; reading Nsight output
  • Experience with distributed multi-node training
  • Strong Python and C++ proficiency
  • Ability to name a model they made faster and quantify improvement

Responsibilities

  • Profile and optimize training and inference for structural and generative models, including transformers, diffusion, and geometric deep learning
  • Write and tune custom kernels (CUDA, Triton) and use compilers (torch.compile, TensorRT, XLA) when beneficial
  • Scale distributed training across 32-64 nodes, employing FSDP, DeepSpeed, tensor and pipeline parallelism, and mixed precision
  • Reduce inference cost by optimizing memory scaling for large complexes, improving diffusion sampling efficiency, batching ragged inputs, and maximizing throughput across up to 1000 GPUs
  • Manage GPU cluster efficiency on GCP, focusing on scheduling, utilization, spot strategy, and cost reporting
  • Develop benchmarks and profiling tools for the research team

Skills

Python
C++
PyTorch internals
CUDA
Triton
Distributed training
Benchmarking
Nsight profiling

Education

BS/MS/PhD in CS/EE or related field

Tools

Nsight
TensorRT
Torch.compile

Job description

Proxima is building frontier AI for proximity therapeutics, combining foundation-model ML with a scalable data generation engine. We are seeking a senior ML systems engineer to profile and optimize training and inference for large structural models, implement CUDA/Triton kernels, and scale multi-node distributed training across large GPU clusters.

Ideal candidates will have 6+ years in ML systems or HPC, deep PyTorch internals knowledge, and strong Python/C++ skills.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Principal ML Performance Engineer (GPU Optimization)
Principal ML Performance Engineer (GPU Optimization)

Proxima • Boston (MA)

On-site
USD 180,000 - 280,000
Senior ML Training Systems Engineer - Distributed CUDA
Senior ML Training Systems Engineer - Distributed CUDA

Genesis AI • San Francisco (CA)

On-site
USD 180,000 - 260,000
Senior ML Systems Engineer – Distributed Training
Senior ML Systems Engineer – Distributed Training

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 280,000
Equity
Health benefits
Remote-friendly US culture
+1
Senior ML Systems Engineer — GPU-Accelerated PyTorch & GNNs
Senior ML Systems Engineer — GPU-Accelerated PyTorch & GNNs

NVIDIA AI • Austin (CA)

On-site
USD 140,000 - 190,000
Equity
Benefits
Senior ML Systems Engineer — GPU-Accelerated PyTorch
Senior ML Systems Engineer — GPU-Accelerated PyTorch

NVIDIA • Town of Texas (WI)

On-site
USD 184,000 - 288,000
Senior ML Infra Engineer - Scale GPU Clusters, Remote
Senior ML Infra Engineer - Scale GPU Clusters, Remote

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 320,000 - 500,000
Equity
Medical/Dental/Vision coverage
Unlimited PTO
+1
Senior ML Systems Engineer: Training Frameworks & Tooling
Senior ML Systems Engineer: Training Frameworks & Tooling

Cohere • United States

Remote
USD 150,000 - 230,000
Senior ML Platform Engineer for Distributed GPU Training
Senior ML Platform Engineer for Distributed GPU Training

Xairatherapeutics • Seattle (WA)

On-site
USD 205,000 - 325,000
Equity
Bonus
Senior GPU-Accelerated ML Systems Engineer
Senior GPU-Accelerated ML Systems Engineer

NVIDIA Corporation • Austin (TX)

On-site
USD 152,000 - 242,000
Senior HPC Performance Engineer — AI for Science (Equity)
Senior HPC Performance Engineer — AI for Science (Equity)

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 184,000 - 288,000
Equity
Benefits