ML Distributed Training Engineer for Trainium Accelerators

Amazon Web Services (AWS)

Cupertino (CA)

On-site

USD 165,200 - 223,600

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health insurance
401(k) matching
Paid time off

Job summary

Annapurna Labs, part of AWS, is seeking a software engineer to advance distributed training on Trainium and Neuron. You will work with cutting-edge ML models and collaborate with chip architects, compiler engineers, and runtime teams to push performance and efficiency at scale.

Based in Cupertino, CA, this role emphasizes optimization across software stacks and mentorship within a diverse, inclusive team culture. A Bachelor’s degree in CS or equivalent is required.

Qualifications

  • 3+ years of non‑internship professional software development experience.
  • 2+ years design or architecture (design patterns, reliability and scaling) of systems.
  • Experience programming with at least one software programming language.

Responsibilities

  • Lead efforts to optimize distributed training performance on Trainium, maximizing throughput and efficiency.
  • Collaborate across PyTorch, JAX, and the Neuron compiler/runtime to enable large-scale training workloads.
  • Identify and resolve bottlenecks across the stack from communications to kernel performance.

Skills

Software development
System design / architecture
Programming languages

Education

Bachelor's degree in computer science or equivalent

Job description

Annapurna Labs, part of AWS, is seeking a software engineer to advance distributed training on Trainium and Neuron. You will work with cutting-edge ML models and collaborate with chip architects, compiler engineers, and runtime teams to push performance and efficiency at scale.

Based in Cupertino, CA, this role emphasizes optimization across software stacks and mentorship within a diverse, inclusive team culture. A Bachelor’s degree in CS or equivalent is required.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Distributed Training Engineer for Trainium
ML Distributed Training Engineer for Trainium

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 143,000 - 195,000
ML Systems Engineer - Distributed Training on Trainium
ML Systems Engineer - Distributed Training on Trainium

Amazon • Cupertino (CA)

On-site
USD 165,000 - 224,000
ML Systems Engineer — AI/GenAI Training & HPC
ML Systems Engineer — AI/GenAI Training & HPC

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 165,000 - 224,000
ML Systems Engineer - Distributed Training on Trainium
ML Systems Engineer - Distributed Training on Trainium

Amazon • Cupertino (CA)

On-site
USD 165,000 - 224,000
Health insurance
RSUs
Sign-on bonuses
ML Systems Engineer - Distributed Training on Trainium
ML Systems Engineer - Distributed Training on Trainium

Socket.dev • Seattle (WA)

On-site
USD 144,000 - 194,000
Senior AI/ML Software Engineer — Distributed Training
Senior AI/ML Software Engineer — Distributed Training

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 168,000 - 228,000
RSUs
Health benefits
Sign-on payments
Senior AI/ML Distributed Training Engineer
Senior AI/ML Distributed Training Engineer

Amazon • Cupertino (CA)

On-site
USD 193,300 - 261,500
Sr. Software Engineer- AI/ML, AWS Neuron Distributed Training
Sr. Software Engineer- AI/ML, AWS Neuron Distributed Training

Amazon • Cupertino (CA)

On-site
USD 193,300 - 261,500
Senior AI/ML Distributed Training Engineer (PyTorch)
Senior AI/ML Distributed Training Engineer (PyTorch)

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 168,100 - 227,400
Health insurance
RSUs
Paid time off
+1
Applied Scientist II — ML Systems for AI Accelerators
Applied Scientist II — ML Systems for AI Accelerators

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 171,000 - 223,000