Senior AI/ML Distributed Training Engineer

Amazon

Cupertino (CA)

On-site

USD 193,300 - 261,500

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Amazon Annapurna Labs is seeking a Sr. Software Engineer for AI/ML distributed training to design and optimize large-scale ML training on Trainium instances. You will enhance distributed frameworks and optimize mixed-precision training, working closely with hardware and runtime teams.

Qualified candidates have 5+ years in software development, leadership experience, and strong PyTorch expertise. On-site in Cupertino, you will collaborate with AWS solution architects and customers to deploy

Qualifications

  • Bachelor's degree in computer science or equivalent.
  • 5+ years of professional software development experience.
  • 5+ years of programming in at least one software programming language.
  • 5+ years of leading design or architecture of new and existing systems (design patterns, reliability, and scaling).
  • 5+ years of full software development life cycle experience, including coding standards, code reviews, source control, build processes, testing, and operations.
  • Experience as a mentor, tech lead, or leading an engineering team.
  • Experience in machine learning, large scale training with LLMs, and expertise in PyTorch.

Responsibilities

  • Design, implement, and optimize distributed training solutions for large scale ML models running on Trainium instances.
  • Extend and optimize popular distributed training frameworks including FSDP (Fully-Sharded Data Parallel), torchtitan and Hugging Face libraries for the Neuron ecosystem.
  • Develop and optimize mixed-precision and low-precision training techniques to maximize training throughput while maintaining model accuracy.
  • Implement precision-aware training strategies, loss scaling techniques, and careful gradient management to ensure training stability across reduced precision formats.
  • Profile, analyze, and tune end-to-end training pipelines to achieve optimal performance on Trainium hardware.
  • Partner with hardware, compiler, and runtime teams to influence system design and unlock new capabilities.
  • Work directly with AWS solution architects and customers to deploy and optimize training workloads at scale.

Skills

Software development experience
Distributed systems design
Mentorship / Tech lead
PyTorch experience

Education

Bachelor's degree in computer science or equivalent
Master's degree in computer science or equivalent

Job description

Amazon Annapurna Labs is seeking a Sr. Software Engineer for AI/ML distributed training to design and optimize large-scale ML training on Trainium instances. You will enhance distributed frameworks and optimize mixed-precision training, working closely with hardware and runtime teams.

Qualified candidates have 5+ years in software development, leadership experience, and strong PyTorch expertise. On-site in Cupertino, you will collaborate with AWS solution architects and customers to deploy

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior AI/ML Systems Engineer - Distributed Training
Senior AI/ML Systems Engineer - Distributed Training

Amazon • Cupertino (CA)

On-site
USD 193,000 - 262,000
Health insurance
401(k) matching
Paid time off
+2
Lead AI/ML Distributed Training Engineer on Trainium
Lead AI/ML Distributed Training Engineer on Trainium

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 193,000 - 262,000
AI/ML Software Engineer - Distributed Training HPC
AI/ML Software Engineer - Distributed Training HPC

Amazon • Cupertino (CA), Northern (KY)

Hybrid
USD 165,000 - 224,000
AI/ML Systems Engineer – Distributed Training on Trainium
AI/ML Systems Engineer – Distributed Training on Trainium

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 165,000 - 224,000
Senior AI/ML Distributed Training Engineer
Senior AI/ML Distributed Training Engineer

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 168,000 - 227,000
Health insurance
401(k) matching
Paid time off
+1
AI/ML Software Engineer - Trainium Distributed Training
AI/ML Software Engineer - Trainium Distributed Training

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 168,000 - 227,000
AI/ML Software Engineer — Trainium Distributed Training
AI/ML Software Engineer — Trainium Distributed Training

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 144,000 - 194,000
Health insurance
401(k) matching
Paid time off
+1
Lead ML Systems Engineer – AI/GenAI Acceleration
Lead ML Systems Engineer – AI/GenAI Acceleration

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 193,000 - 262,000
Health insurance
RSUs (restricted stock units)
401(k) matching
+2
Sr. Software Engineer- AI/ML, AWS Neuron Distributed Training
Sr. Software Engineer- AI/ML, AWS Neuron Distributed Training

Amazon • Cupertino (CA)

On-site
USD 193,300 - 261,500
AI/ML Inference Engineer (Trainium)
AI/ML Inference Engineer (Trainium)

Amazon • Cupertino (CA)

On-site
USD 165,000 - 224,000
Health insurance
401(k) matching
Paid time off
+1