ML Systems Engineer: Optimizing Training & GPU Kernels

Jobtailor

Massachusetts

On-site

USD 120,000 - 180,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Jobtailor in Massachusetts seeks an experienced ML engineer to optimize training and inference across large models. You will profile data loading and kernel performance, implement fusion, tiling, and sharding, and collaborate with distributed training teams.

Candidates should excel in Python and PyTorch, have deep ML fundamentals, and demonstrate strong analytical problem-solving with a bias for action. Experience with CUDA or Triton is preferred, with a track record of throughput improvements.

Qualifications

  • Advanced knowledge of Python and PyTorch for ML workloads.
  • Experience optimizing training and inference execution.
  • Strong problem-solving and data-driven approach.

Responsibilities

  • Profile and optimize data loading to maximize throughput.
  • Develop high-performance GPU kernels in CUDA/Triton.
  • Optimize distributed training pipelines with PyTorch Distributed.
  • Implement kernel fusion, sharding, and tiling to reduce step time.

Skills

Python
PyTorch
Data loading
Performance optimization
Analytical thinking

Education

Bachelor's/Master's/PhD in CS/CE or related

Tools

Nsight
PyTorch Profiler
CUDA
Triton

Job description

Jobtailor in Massachusetts seeks an experienced ML engineer to optimize training and inference across large models. You will profile data loading and kernel performance, implement fusion, tiling, and sharding, and collaborate with distributed training teams.

Candidates should excel in Python and PyTorch, have deep ML fundamentals, and demonstrate strong analytical problem-solving with a bias for action. Experience with CUDA or Triton is preferred, with a track record of throughput improvements.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Systems Scientist — High-Throughput Inference & GPUs
ML Systems Scientist — High-Throughput Inference & GPUs

Jobtailor • California (MO)

On-site
USD 150,000 - 210,000
Machine Learning Systems Engineer
Machine Learning Systems Engineer

Jobtailor • Massachusetts

On-site
USD 120,000 - 180,000
Machine Learning Engineer
Machine Learning Engineer

Motional AD Inc. • United States

Hybrid
USD 110,000 - 170,000
ML Systems Engineer - High-Performance Distributed Training
ML Systems Engineer - High-Performance Distributed Training

Motional • Boston (MA)

Hybrid
USD 144,000 - 192,000
Medical, dental, and vision insurance
401k with company match
Life insurance
+1
Remote ML Systems Engineer — Distributed Training & GPU
Remote ML Systems Engineer — Distributed Training & GPU

Motional AD Inc. • United States

Hybrid
USD 110,000 - 170,000
GPU Kernel Engineer — Fast ML Training & Inference
GPU Kernel Engineer — Fast ML Training & Inference

Tilde Research • Palo Alto (CA)

On-site
USD 150,000 - 260,000
Systems ML Engineer - Edge/Cloud Performance Architect
Systems ML Engineer - Edge/Cloud Performance Architect

Transfyr Bio • Cambridge (MA)

On-site
USD 160,000 - 230,000
Member of Technical Staff, Kernels
Member of Technical Staff, Kernels

Inception • San Francisco (CA)

On-site
USD 180,000 - 260,000
Senior Inference Systems Engineer — Low-Latency ML Serving
Senior Inference Systems Engineer — Low-Latency ML Serving

Jobtailor • Palo Alto (CA)

On-site
USD 180,000 - 240,000
LLM Pre-training & Distributed Engineer (AI Infrastructure)
LLM Pre-training & Distributed Engineer (AI Infrastructure)

Hyphen Connect • Boston (MA)

On-site
USD 120,000 - 160,000