Remote-Ready ML Systems Engineer: High-Throughput Training

United States Digital Space LLC

United States

Hybrid

USD 120,000 - 160,000

Full time

12 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

AD Inc. is seeking a Machine Learning Systems Engineer to join the ML Acceleration team. You will own core systems enabling researchers to train frontier models at scale, focusing on speed, cost, reliability, and throughput.

The role involves profiling, kernel development, and data-pipeline optimization to push the efficiency of large-scale distributed model training and shorten convergence times.

Qualifications

  • Bachelor's, Master's degree, or PhD in Computer Science, Computer Engineering, or related technical discipline.
  • Strong proficiency in Python.
  • Extensive hands-on experience with PyTorch.
  • Experience optimizing ML model execution during training and inference; solid ML concepts and architectures.

Responsibilities

  • Performance profiling and optimization to reduce step time.
  • Optimize distributed training pipelines using frameworks like PyTorch Distributed.
  • Design and maintain high-performance GPU kernels in Triton or CUDA.
  • Build and optimize robust data loading pipelines for training throughput.

Skills

Python
PyTorch
Problem solving
Analytical thinking

Education

Bachelor's/Master's/PhD in CS/CE

Tools

Nsight
PyTorch Profiler
CUDA
Triton

Job description

AD Inc. is seeking a Machine Learning Systems Engineer to join the ML Acceleration team. You will own core systems enabling researchers to train frontier models at scale, focusing on speed, cost, reliability, and throughput.

The role involves profiling, kernel development, and data-pipeline optimization to push the efficiency of large-scale distributed model training and shorten convergence times.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Systems Engineer - High-Performance Distributed Training
ML Systems Engineer - High-Performance Distributed Training

Motional • Boston (MA)

Hybrid
USD 144,000 - 192,000
Medical, dental, and vision insurance
401k with company match
Life insurance
+1
Remote ML Systems Engineer — Distributed Training & GPU
Remote ML Systems Engineer — Distributed Training & GPU

Motional AD Inc. • United States

Hybrid
USD 110,000 - 170,000
ML Performance Engineer — Scale Distributed Training & Throughput
ML Performance Engineer — Scale Distributed Training & Throughput

applied • Sunnyvale (CA)

On-site
USD 170,000 - 230,000
Senior ML Infrastructure Engineer — High-Throughput AI Research
Senior ML Infrastructure Engineer — High-Throughput AI Research

Preference Model • San Francisco (CA)

On-site
USD 180,000 - 300,000
Cash and equity
Ownership & autonomy
Visa sponsorship
+5
Systems ML Engineer - Edge/Cloud Performance Architect
Systems ML Engineer - Edge/Cloud Performance Architect

Transfyr Bio • Cambridge (MA)

On-site
USD 160,000 - 230,000
Senior ML Performance Engineer: Scale & Throughput
Senior ML Performance Engineer: Scale & Throughput

NLP PEOPLE • Sunnyvale (CA)

On-site
USD 215,000 - 285,000
Lead ML Training Systems Engineer - Multimodal, Large-Scale
Lead ML Training Systems Engineer - Multimodal, Large-Scale

Rhoda AI • Palo Alto (CA)

On-site
USD 210,000 - 320,000
ML Performance Engineer: Scale Training & Throughput
ML Performance Engineer: Scale Training & Throughput

Applied Intuition • Sunnyvale (CA)

On-site
USD 180,000 - 240,000
Systems ML Engineer: Cloud & Edge Performance
Systems ML Engineer: Cloud & Edge Performance

S27a • Cambridge (MA)

On-site
USD 170,000 - 240,000
Competitive compensation
Full benefits
Machine Learning Systems Engineer
Machine Learning Systems Engineer

United States Digital Space LLC • United States

Hybrid
USD 120,000 - 160,000