AI/ML Systems Engineer – Distributed Training on Trainium

Amazon Web Services (AWS)

Cupertino (CA)

On-site

USD 165,000 - 224,000

Full time

7 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Annapurna Labs (U.S.) Inc. within AWS is seeking engineers to optimize distributed training on Trainium. You will work across PyTorch and Neuron stack with compiler and runtime teams to improve training throughput on AWS accelerators.

The role focuses on defining parallelism strategies, profiling workloads, and translating performance gaps into framework requirements, contributing upstream to open-source projects.

Qualifications

  • 3+ years of non-internship professional software development experience.
  • 2+ years of design or architecture (design patterns, reliability and scaling) of new and existing systems.
  • Experience programming with at least one software programming language.

Responsibilities

  • Lead optimization of distributed training performance across the Neuron software stack for Trainium.
  • Own parallelism strategies across data, tensor, pipeline, expert, and context parallelism.
  • Profile workloads to determine bottlenecks and drive fixes with compiler, runtime, and collectives teams.

Skills

Software development
System design
Programming languages

Education

Bachelor's degree in CS

Job description

Annapurna Labs (U.S.) Inc. within AWS is seeking engineers to optimize distributed training on Trainium. You will work across PyTorch and Neuron stack with compiler and runtime teams to improve training throughput on AWS accelerators.

The role focuses on defining parallelism strategies, profiling workloads, and translating performance gaps into framework requirements, contributing upstream to open-source projects.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI/ML Software Engineer - Trainium Distributed Training
AI/ML Software Engineer - Trainium Distributed Training

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 168,000 - 227,000
Lead AI/ML Distributed Training Engineer on Trainium
Lead AI/ML Distributed Training Engineer on Trainium

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 193,000 - 262,000
AI/ML Software Engineer - Distributed Training HPC
AI/ML Software Engineer - Distributed Training HPC

Amazon • Cupertino (CA), Northern (KY)

Hybrid
USD 165,000 - 224,000
AI/ML Software Engineer — Trainium Distributed Training
AI/ML Software Engineer — Trainium Distributed Training

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 144,000 - 194,000
Health insurance
401(k) matching
Paid time off
+1
ML Systems Engineer - Distributed Training on Trainium
ML Systems Engineer - Distributed Training on Trainium

Socket.dev • Seattle (WA)

On-site
USD 144,000 - 194,000
Senior AI/ML Distributed Training Engineer
Senior AI/ML Distributed Training Engineer

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 168,000 - 227,000
Health insurance
401(k) matching
Paid time off
+1
Senior AI/ML Distributed Training Engineer
Senior AI/ML Distributed Training Engineer

Amazon • Cupertino (CA)

On-site
USD 193,300 - 261,500
Lead ML Systems Engineer – AI/GenAI Acceleration
Lead ML Systems Engineer – AI/GenAI Acceleration

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 193,000 - 262,000
Health insurance
RSUs (restricted stock units)
401(k) matching
+2
AI/ML Inference Engineer for PyTorch on Trainium
AI/ML Inference Engineer for PyTorch on Trainium

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 150,000 - 210,000
Senior AI/ML Systems Engineer - Distributed Training
Senior AI/ML Systems Engineer - Distributed Training

Amazon • Cupertino (CA)

On-site
USD 193,000 - 262,000
Health insurance
401(k) matching
Paid time off
+2