AI/ML Software Engineer - Trainium Distributed Training

Amazon Web Services (AWS)

Seattle (WA)

On-site

USD 168,000 - 227,000

Full time

2 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Annapurna Labs (AWS) is hiring engineers to optimize distributed training for Trainium, focusing on throughput, FLOPs utilization, and convergence time. You’ll work across PyTorch, JAX, and the Neuron stack to enable and tune large-scale workloads, identify missing operators, and design effective parallelism strategies.

You will mentor a team of engineers, collaborate with customers, and contribute to open source ecosystems, shaping performance across the stack from frameworks to kernels.

Qualifications

  • Bachelor's degree in computer science or equivalent
  • 5+ years of non-internship professional software development experience
  • 5+ years of programming with at least one software programming language
  • 5+ years of leading design or architecture of new and existing systems
  • 5+ years of full software development life cycle, including coding standards, code reviews, source control management, build processes, testing, and operations
  • Experience as a mentor, tech lead or leading an engineering team
  • Experience in machine learning, data mining, information retrieval, statistics or natural language processing

Responsibilities

  • Lead efforts to optimize distributed training performance
  • Enable and tune large-scale training workloads across the Neuron stack
  • Own parallelism strategies including data, tensor, pipeline, expert, and context parallelism
  • Profile end-to-end performance to identify bottlenecks and drive fixes
  • Translate performance gaps into requirements influencing Trainium architecture
  • Collaborate with customers to enable model training at scale

Skills

Software development
Mentor / tech lead
ML knowledge
Performance optimization
Distributed systems

Education

Bachelor's degree
Master's degree (preferred)

Tools

PyTorch
JAX
TensorFlow

Job description

Annapurna Labs (AWS) is hiring engineers to optimize distributed training for Trainium, focusing on throughput, FLOPs utilization, and convergence time. You’ll work across PyTorch, JAX, and the Neuron stack to enable and tune large-scale workloads, identify missing operators, and design effective parallelism strategies.

You will mentor a team of engineers, collaborate with customers, and contribute to open source ecosystems, shaping performance across the stack from frameworks to kernels.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI/ML Software Engineer — Trainium Distributed Training
AI/ML Software Engineer — Trainium Distributed Training

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 144,000 - 194,000
Health insurance
401(k) matching
Paid time off
+1
Senior AI/ML Distributed Training Engineer
Senior AI/ML Distributed Training Engineer

Annapurna Labs (U.S.) Inc. • Cupertino (CA)

On-site
USD 180,000 - 240,000
ML Systems Engineer - Distributed Training on Trainium
ML Systems Engineer - Distributed Training on Trainium

Socket.dev • Seattle (WA)

On-site
USD 144,000 - 194,000
Senior AI/ML Distributed Training Engineer
Senior AI/ML Distributed Training Engineer

Amazon • Cupertino (CA)

On-site
USD 193,300 - 261,500
AI/ML Software Engineer for High-Performance Training
AI/ML Software Engineer for High-Performance Training

Amazon • Cupertino (CA)

On-site
USD 165,000 - 224,000
Software Engineer, AI Training Collectives
Software Engineer, AI Training Collectives

Amazon • Cupertino (CA)

On-site
USD 165,000 - 224,000
Sr. Software Engineer- AI/ML, AWS Neuron Distributed Training
Sr. Software Engineer- AI/ML, AWS Neuron Distributed Training

Amazon • Cupertino (CA)

On-site
USD 193,300 - 261,500
AI Training Compute Software Engineer
AI Training Compute Software Engineer

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 165,000 - 224,000
AI/ML Performance Engineer – Training on Trainium
AI/ML Performance Engineer – Training on Trainium

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 140,000 - 210,000
Health Insurance
Medical Insurance
Dental Insurance
+14
Sr. Software Engineer- AI/ML, AWS Neuron Distributed Training - Performance Optimization
Sr. Software Engineer- AI/ML, AWS Neuron Distributed Training - Performance Optimization

Amazon • Seattle (WA)

On-site
USD 120,000 - 160,000
Health insurance
401(k) matching
Paid time off
+1