AI/ML Systems Engineer — Distributed Training

Amazon Web Services (AWS)

Seattle (WA)

On-site

USD 144,000 - 194,000

Full time

5 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Annapurna Labs (U.S.) Inc. at AWS is seeking engineers to optimize distributed training on Trainium and improve throughput across the Neuron stack. You will own data, tensor, and other parallelism strategies and work with PyTorch and JAX to push performance.

The role requires deep ML, HPC expertise and collaboration with compiler, runtime, and framework teams to land impactful improvements for customers training frontier models.

Qualifications

  • 3+ years of non-internship professional software development experience.
  • 2+ years of non-internship design or architecture experience.
  • Experience programming with at least one programming language.

Responsibilities

  • Optimize distributed training performance on Trainium across the Neuron software stack.
  • Own parallelism strategies: data, tensor, pipeline, expert, and context parallelism.
  • Profile end-to-end workloads to identify bottlenecks and drive fixes with compiler/runtime teams.
  • Translate performance gaps into framework requirements and contribute upstream.

Skills

Software development
Design/Architecture
Programming languages

Education

Bachelor's degree in computer science or equivalent

Job description

Annapurna Labs (U.S.) Inc. at AWS is seeking engineers to optimize distributed training on Trainium and improve throughput across the Neuron stack. You will own data, tensor, and other parallelism strategies and work with PyTorch and JAX to push performance.

The role requires deep ML, HPC expertise and collaboration with compiler, runtime, and framework teams to land impactful improvements for customers training frontier models.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Systems Engineer — AI/GenAI Training & HPC
ML Systems Engineer — AI/GenAI Training & HPC

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 165,000 - 224,000
Senior AI/ML Software Engineer — Distributed Training
Senior AI/ML Software Engineer — Distributed Training

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 168,000 - 228,000
RSUs
Health benefits
Sign-on payments
ML Systems Engineer - Distributed Training on Trainium
ML Systems Engineer - Distributed Training on Trainium

Amazon • Cupertino (CA)

On-site
USD 165,000 - 224,000
ML Systems Engineer - Distributed Training on Trainium
ML Systems Engineer - Distributed Training on Trainium

Amazon • Cupertino (CA)

On-site
USD 165,000 - 224,000
Health insurance
RSUs
Sign-on bonuses
Senior AI/ML Distributed Training Engineer (PyTorch)
Senior AI/ML Distributed Training Engineer (PyTorch)

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 168,100 - 227,400
Health insurance
RSUs
Paid time off
+1
ML Systems Engineer - Distributed Training on Trainium
ML Systems Engineer - Distributed Training on Trainium

Socket.dev • Seattle (WA)

On-site
USD 144,000 - 194,000
ML Distributed Training Engineer for Trainium
ML Distributed Training Engineer for Trainium

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 143,000 - 195,000
Senior AI/ML Distributed Training Engineer
Senior AI/ML Distributed Training Engineer

Amazon • Cupertino (CA)

On-site
USD 193,300 - 261,500
ML Distributed Training Engineer for Trainium Accelerators
ML Distributed Training Engineer for Trainium Accelerators

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 165,000 - 224,000
Health insurance
401(k) matching
Paid time off
Software Engineer, AI Training Collectives
Software Engineer, AI Training Collectives

Amazon • Cupertino (CA)

On-site
USD 165,000 - 224,000