Lead AI/ML Distributed Training Engineer on Trainium

Amazon Web Services (AWS)

Cupertino (CA)

On-site

USD 193,000 - 262,000

Full time

5 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Amazon Web Services (AWS) Annapurna Labs team seeks engineers to optimize distributed training for Trainium. You will enhance training throughput and model FLOPs utilization while collaborating across PyTorch, JAX, and the Neuron stack to push performance and scalability.

You will mentor peers and shape future architecture within a highly technical, open collaboration culture. You will work on a diverse stack from frameworks to hardware, applying novel parallelism strategies and tuning kernels

Qualifications

  • 5+ years of non-internship professional software development experience.
  • 5+ years of programming with at least one software programming language experience.
  • 5+ years of leading design or architecture (design patterns, reliability and scaling) of new and existing systems experience.
  • 5+ years of full software development life cycle, including coding standards, code reviews, source control management, build processes, testing, and operations experience.
  • Experience as a mentor, tech lead or leading an engineering team.
  • Experience in machine learning, data mining, information retrieval, statistics or natural language processing.

Responsibilities

  • Lead efforts to optimize distributed training performance on Trainium, focusing on throughput, FLOPs utilization, and time to convergence across the Neuron stack.
  • Work across PyTorch, JAX, and the Neuron compiler/runtime to enable large-scale training on Trainium instances and identify missing operators and sharding strategies.
  • Own parallelism strategies like data, tensor, pipeline, expert, and context parallelism, applying reduced-precision formats when beneficial.
  • Profile end-to-end workloads to determine bottlenecks and drive fixes with compiler, runtime, and collectives engineers.
  • Translate performance gaps into requirements influencing future Trainium architecture and contribute upstream to open-source frameworks.

Skills

Programming experience
Leadership experience
Mentor
ML experience
SDLC experience

Education

Bachelor's degree in computer science or equivalent

Tools

PyTorch/JAX/TensorFlow

Job description

Amazon Web Services (AWS) Annapurna Labs team seeks engineers to optimize distributed training for Trainium. You will enhance training throughput and model FLOPs utilization while collaborating across PyTorch, JAX, and the Neuron stack to push performance and scalability.

You will mentor peers and shape future architecture within a highly technical, open collaboration culture. You will work on a diverse stack from frameworks to hardware, applying novel parallelism strategies and tuning kernels

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI/ML Software Engineer - Trainium Distributed Training
AI/ML Software Engineer - Trainium Distributed Training

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 168,000 - 227,000
AI/ML Systems Engineer – Distributed Training on Trainium
AI/ML Systems Engineer – Distributed Training on Trainium

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 165,000 - 224,000
AI/ML Software Engineer - Distributed Training HPC
AI/ML Software Engineer - Distributed Training HPC

Amazon • Cupertino (CA), Northern (KY)

Hybrid
USD 165,000 - 224,000
Senior AI/ML Distributed Training Engineer
Senior AI/ML Distributed Training Engineer

Amazon • Cupertino (CA)

On-site
USD 193,300 - 261,500
AI/ML Software Engineer — Trainium Distributed Training
AI/ML Software Engineer — Trainium Distributed Training

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 144,000 - 194,000
Health insurance
401(k) matching
Paid time off
+1
Senior AI/ML Distributed Training Engineer
Senior AI/ML Distributed Training Engineer

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 168,000 - 227,000
Health insurance
401(k) matching
Paid time off
+1
ML Systems Engineer - Distributed Training on Trainium
ML Systems Engineer - Distributed Training on Trainium

Socket.dev • Seattle (WA)

On-site
USD 144,000 - 194,000
Lead ML Systems Engineer – AI/GenAI Acceleration
Lead ML Systems Engineer – AI/GenAI Acceleration

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 193,000 - 262,000
Health insurance
RSUs (restricted stock units)
401(k) matching
+2
Senior AI/ML Systems Engineer - Distributed Training
Senior AI/ML Systems Engineer - Distributed Training

Amazon • Cupertino (CA)

On-site
USD 193,000 - 262,000
Health insurance
401(k) matching
Paid time off
+2
AI/ML Performance Engineer – Training on Trainium
AI/ML Performance Engineer – Training on Trainium

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 140,000 - 210,000
Health Insurance
Medical Insurance
Dental Insurance
+14