ML Systems Engineer - Distributed Training on Trainium

Amazon

Cupertino (CA)

On-site

USD 165,000 - 224,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Amazon's Annapurna Labs team at AWS builds Neuron to accelerate DL on Trainium. The role focuses on optimizing distributed training throughput across the Neuron stack, working with PyTorch and JAX to push frontier-scale models.

You will own parallelism strategies, profile workloads, and translate findings into framework requirements, contributing upstream to OSS. The team emphasizes collaboration across hardware and software layers, with opportunities to shape AI acceleration tech across

Qualifications

  • 3+ years of professional software development experience.
  • Experience with distributed training systems and performance optimization.
  • Familiarity with PyTorch or JAX and modern ML frameworks.

Responsibilities

  • Lead efforts to optimize distributed training performance on Trainium.
  • Own parallelism strategies across data, tensor, pipeline, and context.
  • Profile end to end to identify bottlenecks and drive fixes with teams.

Skills

Software development
Distributed training
PyTorch
JAX
Performance tuning

Education

Bachelor's degree in CS

Tools

Python
CUDA

Job description

Amazon's Annapurna Labs team at AWS builds Neuron to accelerate DL on Trainium. The role focuses on optimizing distributed training throughput across the Neuron stack, working with PyTorch and JAX to push frontier-scale models.

You will own parallelism strategies, profile workloads, and translate findings into framework requirements, contributing upstream to OSS. The team emphasizes collaboration across hardware and software layers, with opportunities to shape AI acceleration tech across

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Systems Engineer - Distributed Training on Trainium
ML Systems Engineer - Distributed Training on Trainium

Socket.dev • Seattle (WA)

On-site
USD 144,000 - 194,000
ML Systems Engineer - Distributed Training on Trainium
ML Systems Engineer - Distributed Training on Trainium

Amazon • Cupertino (CA)

On-site
USD 165,000 - 224,000
Health insurance
RSUs
Sign-on bonuses
ML Systems Engineer — AI/GenAI Training & HPC
ML Systems Engineer — AI/GenAI Training & HPC

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 165,000 - 224,000
ML Distributed Training Engineer for Trainium
ML Distributed Training Engineer for Trainium

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 143,000 - 195,000
ML Distributed Training Engineer for Trainium Accelerators
ML Distributed Training Engineer for Trainium Accelerators

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 165,000 - 224,000
Health insurance
401(k) matching
Paid time off
Senior AI/ML Software Engineer — Distributed Training
Senior AI/ML Software Engineer — Distributed Training

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 168,000 - 228,000
RSUs
Health benefits
Sign-on payments
Senior AI/ML Distributed Training Engineer (PyTorch)
Senior AI/ML Distributed Training Engineer (PyTorch)

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 168,100 - 227,400
Health insurance
RSUs
Paid time off
+1
Sr. Software Engineer- AI/ML, AWS Neuron Distributed Training - Performance Optimization
Sr. Software Engineer- AI/ML, AWS Neuron Distributed Training - Performance Optimization

Amazon • Seattle (WA)

On-site
USD 120,000 - 160,000
Health insurance
401(k) matching
Paid time off
+1
Senior AI/ML Distributed Training Engineer
Senior AI/ML Distributed Training Engineer

Amazon • Cupertino (CA)

On-site
USD 193,300 - 261,500
Sr. Software Engineer- AI/ML, AWS Neuron Distributed Training
Sr. Software Engineer- AI/ML, AWS Neuron Distributed Training

Amazon • Cupertino (CA)

On-site
USD 193,300 - 261,500