Senior AI/ML Distributed Training Engineer

Amazon

Cupertino (CA)

On-site

USD 193,300 - 261,500

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Amazon Annapurna Labs is seeking a Sr. Software Engineer for AI/ML distributed training to design and optimize large-scale ML training on Trainium instances. You will enhance distributed frameworks and optimize mixed-precision training, working closely with hardware and runtime teams.

Qualified candidates have 5+ years in software development, leadership experience, and strong PyTorch expertise. On-site in Cupertino, you will collaborate with AWS solution architects and customers to deploy

Qualifications

  • Bachelor's degree in computer science or equivalent.
  • 5+ years of professional software development experience.
  • 5+ years of programming in at least one software programming language.
  • 5+ years of leading design or architecture of new and existing systems (design patterns, reliability, and scaling).
  • 5+ years of full software development life cycle experience, including coding standards, code reviews, source control, build processes, testing, and operations.
  • Experience as a mentor, tech lead, or leading an engineering team.
  • Experience in machine learning, large scale training with LLMs, and expertise in PyTorch.

Responsibilities

  • Design, implement, and optimize distributed training solutions for large scale ML models running on Trainium instances.
  • Extend and optimize popular distributed training frameworks including FSDP (Fully-Sharded Data Parallel), torchtitan and Hugging Face libraries for the Neuron ecosystem.
  • Develop and optimize mixed-precision and low-precision training techniques to maximize training throughput while maintaining model accuracy.
  • Implement precision-aware training strategies, loss scaling techniques, and careful gradient management to ensure training stability across reduced precision formats.
  • Profile, analyze, and tune end-to-end training pipelines to achieve optimal performance on Trainium hardware.
  • Partner with hardware, compiler, and runtime teams to influence system design and unlock new capabilities.
  • Work directly with AWS solution architects and customers to deploy and optimize training workloads at scale.

Skills

Software development experience
Distributed systems design
Mentorship / Tech lead
PyTorch experience

Education

Bachelor's degree in computer science or equivalent
Master's degree in computer science or equivalent

Job description

Amazon Annapurna Labs is seeking a Sr. Software Engineer for AI/ML distributed training to design and optimize large-scale ML training on Trainium instances. You will enhance distributed frameworks and optimize mixed-precision training, working closely with hardware and runtime teams.

Qualified candidates have 5+ years in software development, leadership experience, and strong PyTorch expertise. On-site in Cupertino, you will collaborate with AWS solution architects and customers to deploy

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI/ML Distributed Training Engineer (PyTorch)
Senior AI/ML Distributed Training Engineer (PyTorch)

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 168,100 - 227,400
Health insurance
RSUs
Paid time off
+1
AI/ML Systems Engineer — Distributed Training
AI/ML Systems Engineer — Distributed Training

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 144,000 - 194,000
Senior AI/ML Software Engineer — Distributed Training
Senior AI/ML Software Engineer — Distributed Training

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 168,000 - 228,000
RSUs
Health benefits
Sign-on payments
Sr. Software Engineer- AI/ML, AWS Neuron Distributed Training
Sr. Software Engineer- AI/ML, AWS Neuron Distributed Training

Amazon • Cupertino (CA)

On-site
USD 193,300 - 261,500
ML Systems Engineer - Distributed Training on Trainium
ML Systems Engineer - Distributed Training on Trainium

Amazon • Cupertino (CA)

On-site
USD 165,000 - 224,000
Health insurance
RSUs
Sign-on bonuses
Senior AI/ML Software Engineer - Distributed Training
Senior AI/ML Software Engineer - Distributed Training

Amazon • Seattle (WA)

On-site
USD 120,000 - 160,000
Health insurance
401(k) matching
Paid time off
+1
ML Distributed Training Engineer for Trainium Accelerators
ML Distributed Training Engineer for Trainium Accelerators

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 165,000 - 224,000
Health insurance
401(k) matching
Paid time off
ML Distributed Training Engineer for Trainium
ML Distributed Training Engineer for Trainium

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 143,000 - 195,000
Sr. Software Engineer- AI/ML, AWS Neuron Distributed Training - Performance Optimization
Sr. Software Engineer- AI/ML, AWS Neuron Distributed Training - Performance Optimization

Amazon • Seattle (WA)

On-site
USD 120,000 - 160,000
Health insurance
401(k) matching
Paid time off
+1
ML Systems Engineer - Distributed Training on Trainium
ML Systems Engineer - Distributed Training on Trainium

Socket.dev • Seattle (WA)

On-site
USD 144,000 - 194,000