Machine Learning Systems Engineer

Jobtailor

Massachusetts

On-site

USD 120,000 - 180,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Jobtailor in Massachusetts seeks an experienced ML engineer to optimize training and inference across large models. You will profile data loading and kernel performance, implement fusion, tiling, and sharding, and collaborate with distributed training teams.

Candidates should excel in Python and PyTorch, have deep ML fundamentals, and demonstrate strong analytical problem-solving with a bias for action. Experience with CUDA or Triton is preferred, with a track record of throughput improvements.

Qualifications

  • Advanced knowledge of Python and PyTorch for ML workloads.
  • Experience optimizing training and inference execution.
  • Strong problem-solving and data-driven approach.

Responsibilities

  • Profile and optimize data loading to maximize throughput.
  • Develop high-performance GPU kernels in CUDA/Triton.
  • Optimize distributed training pipelines with PyTorch Distributed.
  • Implement kernel fusion, sharding, and tiling to reduce step time.

Skills

Python
PyTorch
Data loading
Performance optimization
Analytical thinking

Education

Bachelor's/Master's/PhD in CS/CE or related

Tools

Nsight
PyTorch Profiler
CUDA
Triton

Job description

  • Utilize profiling tools (e.g., Nsight, PyTorch Profiler) to identify bottlenecks in data loading, gradient computation, and communication. Implement optimizations like kernel fusion, sharding, and tiling to improve step time.
  • Optimize distributed training pipelines using frameworks such as PyTorch Distributed.
  • Design and maintain high-performance GPU kernels in Triton or CUDA for state-of-the-art ML workloads.
  • Optimize robust data loading pipelines that maximize training throughput.
Requirements
  • Bachelor’s, Master’s degree, or PhD in Computer Science, Computer Engineering, or a related technical discipline.
  • Strong proficiency in Python.
  • Extensive hands-on experience with PyTorch.
  • Experience optimizing machine learning model execution during training and inference, alongside a strong understanding of fundamental machine learning concepts, architectures, and processes.
  • Exceptional analytical and problem-solving skills, with a bias for action and a data-driven approach to technical challenges.
Core Competencies

Demonstrates expertise in optimizing machine learning workflows through advanced techniques in Python and PyTorch, with a strong foundation in GPU kernel design and data loading strategies to enhance training efficiency.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Machine Learning Engineer
Machine Learning Engineer

Motional AD Inc. • United States

Hybrid
USD 110,000 - 170,000
ML Systems Engineer: Optimizing Training & GPU Kernels
ML Systems Engineer: Optimizing Training & GPU Kernels

Jobtailor • Massachusetts

On-site
USD 120,000 - 180,000
Senior Research Scientist – Machine Learning Systems, Efficiency Engineer
Senior Research Scientist – Machine Learning Systems, Efficiency Engineer

Jobtailor • California (MO)

On-site
USD 150,000 - 210,000
Training / AI Infrastructure
Training / AI Infrastructure

Genesis AI • San Francisco (CA)

On-site
USD 180,000 - 260,000
AI Infrastructure — Training Engineer (Large Model) [33251]
AI Infrastructure — Training Engineer (Large Model) [33251]

Stealth Startup • Menlo Park (CA)

On-site
USD 180,000 - 260,000
AI/ML Engineer – (Next-Generation AI Platforms & Workloads)
AI/ML Engineer – (Next-Generation AI Platforms & Workloads)

VeeAR Projects Inc. • Sunnyvale (CA)

On-site
USD 140,000 - 210,000
Software Engineer, CUDA Deep Learning Systems
Software Engineer, CUDA Deep Learning Systems

Jobtailor • California (MO)

On-site
USD 140,000 - 210,000
LLM Pre-training & Distributed Engineer (AI Infrastructure)
LLM Pre-training & Distributed Engineer (AI Infrastructure)

Hyphen Connect Limited • San Francisco (CA)

On-site
USD 120,000 - 160,000
Machine Learning Performance Engineer
Machine Learning Performance Engineer

Selby Jennings • Chicago (IL)

On-site
USD 140,000 - 210,000
LLM Pre-training & Distributed Engineer (AI Infrastructure)
LLM Pre-training & Distributed Engineer (AI Infrastructure)

Hyphen Connect Limited • Oregon (WI)

On-site
USD 100,000 - 130,000