ML Systems Engineer - Distributed Training on Trainium

Amazon

Cupertino (CA)

On-site

USD 165,000 - 224,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health insurance
RSUs
Sign-on bonuses

Job summary

Amazon Web Services (AWS) is seeking a software engineer for the Distributed Training team in Annapurna Labs to optimize distributed training performance on Trainium. You will work on leveraging PyTorch, JAX, and the Neuron compiler/runtime to scale large models and accelerate training throughput.

The role involves collaboration with chip architects, compiler engineers, and runtime engineers to resolve bottlenecks and push training efficiency across the stack, including memory, communications,

Qualifications

  • 3+ years of non-internship professional software development experience.
  • 2+ years of non-internship design or architecture of new/existing systems.
  • Experience programming with at least one software language.

Responsibilities

  • Lead efforts to optimize distributed training performance on Trainium.
  • Maximize training throughput and model FLOPs utilization.
  • Collaborate across PyTorch, JAX, and Neuron compiler/runtime to enable large-scale training workloads.

Skills

Software development
System design
Programming languages

Education

Bachelor's degree in CS or equivalent

Job description

Amazon Web Services (AWS) is seeking a software engineer for the Distributed Training team in Annapurna Labs to optimize distributed training performance on Trainium. You will work on leveraging PyTorch, JAX, and the Neuron compiler/runtime to scale large models and accelerate training throughput.

The role involves collaboration with chip architects, compiler engineers, and runtime engineers to resolve bottlenecks and push training efficiency across the stack, including memory, communications,

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Systems Engineer - Distributed Training on Trainium
ML Systems Engineer - Distributed Training on Trainium

Socket.dev • Seattle (WA)

On-site
USD 144,000 - 194,000
ML Systems Engineer - Distributed Training on Trainium
ML Systems Engineer - Distributed Training on Trainium

Amazon • Cupertino (CA)

On-site
USD 165,000 - 224,000
ML Distributed Training Engineer for Trainium
ML Distributed Training Engineer for Trainium

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 143,000 - 195,000
ML Systems Engineer — AI/GenAI Training & HPC
ML Systems Engineer — AI/GenAI Training & HPC

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 165,000 - 224,000
ML Distributed Training Engineer for Trainium Accelerators
ML Distributed Training Engineer for Trainium Accelerators

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 165,000 - 224,000
Health insurance
401(k) matching
Paid time off
Senior AI/ML Software Engineer — Distributed Training
Senior AI/ML Software Engineer — Distributed Training

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 168,000 - 228,000
RSUs
Health benefits
Sign-on payments
Senior AI/ML Distributed Training Engineer
Senior AI/ML Distributed Training Engineer

Amazon • Cupertino (CA)

On-site
USD 193,300 - 261,500
Senior AI/ML Distributed Training Engineer (PyTorch)
Senior AI/ML Distributed Training Engineer (PyTorch)

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 168,100 - 227,400
Health insurance
RSUs
Paid time off
+1
Sr. Software Engineer- AI/ML, AWS Neuron Distributed Training - Performance Optimization
Sr. Software Engineer- AI/ML, AWS Neuron Distributed Training - Performance Optimization

Amazon • Seattle (WA)

On-site
USD 120,000 - 160,000
Health insurance
401(k) matching
Paid time off
+1
Sr. Software Engineer- AI/ML, AWS Neuron Distributed Training
Sr. Software Engineer- AI/ML, AWS Neuron Distributed Training

Amazon • Cupertino (CA)

On-site
USD 193,300 - 261,500