ML Distributed Training Engineer for Trainium

Amazon Web Services (AWS)

Seattle (WA)

On-site

USD 143,700 - 194,400

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Annapurna Labs (U.S.) Inc. in Cupertino/Seattle is seeking a software engineer for the Distributed Training team within AWS Neuron.

You will help optimize distributed training on Trainium, focusing on throughput, model flop utilization, and efficiency across the Neuron software stack. You will work with PyTorch, JAX, and the Neuron compiler/runtime to enable and tune large-scale training workloads on the latest Trainium instances.

Qualifications

  • 3+ years of non‑internship professional software development experience.
  • 2+ years of non‑internship design or architecture (design patterns, reliability and scaling) of new and existing systems experience.
  • Experience programming with at least one software programming language.

Responsibilities

  • Lead efforts to optimize distributed training performance on Trainium, maximizing throughput and utilization across the Neuron stack.
  • Collaborate across PyTorch, JAX, and the Neuron compiler and runtime to enable large-scale training workloads on Trainium.

Skills

Software development
System design
Programming languages

Education

Bachelor's degree in CS

Job description

Annapurna Labs (U.S.) Inc. in Cupertino/Seattle is seeking a software engineer for the Distributed Training team within AWS Neuron.

You will help optimize distributed training on Trainium, focusing on throughput, model flop utilization, and efficiency across the Neuron software stack. You will work with PyTorch, JAX, and the Neuron compiler/runtime to enable and tune large-scale training workloads on the latest Trainium instances.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Distributed Training Engineer for Trainium Accelerators
ML Distributed Training Engineer for Trainium Accelerators

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 165,000 - 224,000
Health insurance
401(k) matching
Paid time off
ML Systems Engineer - Distributed Training on Trainium
ML Systems Engineer - Distributed Training on Trainium

Amazon • Cupertino (CA)

On-site
USD 165,000 - 224,000
ML Systems Engineer - Distributed Training on Trainium
ML Systems Engineer - Distributed Training on Trainium

Amazon • Cupertino (CA)

On-site
USD 165,000 - 224,000
Health insurance
RSUs
Sign-on bonuses
ML Systems Engineer - Distributed Training on Trainium
ML Systems Engineer - Distributed Training on Trainium

Socket.dev • Seattle (WA)

On-site
USD 144,000 - 194,000
Senior AI/ML Software Engineer — Distributed Training
Senior AI/ML Software Engineer — Distributed Training

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 168,000 - 228,000
RSUs
Health benefits
Sign-on payments
Senior AI/ML Distributed Training Engineer (PyTorch)
Senior AI/ML Distributed Training Engineer (PyTorch)

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 168,100 - 227,400
Health insurance
RSUs
Paid time off
+1
ML Systems Engineer — AI/GenAI Training & HPC
ML Systems Engineer — AI/GenAI Training & HPC

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 165,000 - 224,000
Senior AI/ML Distributed Training Engineer
Senior AI/ML Distributed Training Engineer

Amazon • Cupertino (CA)

On-site
USD 193,300 - 261,500
Sr. Software Engineer- AI/ML, AWS Neuron Distributed Training
Sr. Software Engineer- AI/ML, AWS Neuron Distributed Training

Amazon • Cupertino (CA)

On-site
USD 193,300 - 261,500
Sr. Software Engineer- AI/ML, AWS Neuron Distributed Training - Performance Optimization
Sr. Software Engineer- AI/ML, AWS Neuron Distributed Training - Performance Optimization

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 168,000 - 228,000
RSUs
Health benefits
Sign-on payments