ML Systems Engineer - Distributed Training on Trainium

Socket.dev

Seattle (WA)

On-site

USD 144,000 - 194,000

Full time

8 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Amazon Web Services (AWS) is hiring for a role within the Annapurna Labs team that builds AWS Neuron to accelerate deep learning and GenAI workloads on AWS Trainium. You will optimize distributed training across the Neuron software stack, collaborating with PyTorch and JAX teams to enable large-scale training on the latest Trainium instances.

The role involves leading parallelism strategies, profiling workloads, and driving fixes across compiler, runtime, and collectives teams to improve

Qualifications

  • 3+ years of non-internship professional software development experience.
  • 2+ years of non-internship design or architecture (design patterns, reliability and scaling) of new and existing systems experience
  • Experience programming with at least one software programming language.
  • 3+ years of full software development life cycle, including coding standards, code reviews, source control management, build processes, testing, and operations experience
  • Bachelor's degree in computer science or equivalent

Responsibilities

  • Lead efforts to optimize distributed training performance on Trainium across the Neuron software stack.
  • Own parallelism strategies (data, tensor, pipeline, expert, context) and apply reduced-precision formats when beneficial.
  • Profile workloads end-to-end to identify bottlenecks and drive fixes with compiler, runtime, and collectives teams.
  • Translate performance gaps into framework requirements and contribute upstream to open source frameworks.

Skills

Software development
System design
Programming languages

Education

Bachelor's degree in CS

Job description

Amazon Web Services (AWS) is hiring for a role within the Annapurna Labs team that builds AWS Neuron to accelerate deep learning and GenAI workloads on AWS Trainium. You will optimize distributed training across the Neuron software stack, collaborating with PyTorch and JAX teams to enable large-scale training on the latest Trainium instances.

The role involves leading parallelism strategies, profiling workloads, and driving fixes across compiler, runtime, and collectives teams to improve

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Systems Engineer - Distributed Training on Trainium
ML Systems Engineer - Distributed Training on Trainium

Amazon • Cupertino (CA)

On-site
USD 165,000 - 224,000
ML Systems Engineer - Distributed Training on Trainium
ML Systems Engineer - Distributed Training on Trainium

Amazon • Cupertino (CA)

On-site
USD 165,000 - 224,000
Health insurance
RSUs
Sign-on bonuses
ML Systems Engineer — AI/GenAI Training & HPC
ML Systems Engineer — AI/GenAI Training & HPC

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 165,000 - 224,000
AI/ML Systems Engineer — Distributed Training
AI/ML Systems Engineer — Distributed Training

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 144,000 - 194,000
ML Distributed Training Engineer for Trainium
ML Distributed Training Engineer for Trainium

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 143,000 - 195,000
ML Distributed Training Engineer for Trainium Accelerators
ML Distributed Training Engineer for Trainium Accelerators

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 165,000 - 224,000
Health insurance
401(k) matching
Paid time off
Senior AI/ML Software Engineer — Distributed Training
Senior AI/ML Software Engineer — Distributed Training

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 168,000 - 228,000
RSUs
Health benefits
Sign-on payments
Sr. Software Engineer- AI/ML, AWS Neuron Distributed Training - Performance Optimization
Sr. Software Engineer- AI/ML, AWS Neuron Distributed Training - Performance Optimization

Amazon • Seattle (WA)

On-site
USD 120,000 - 160,000
Health insurance
401(k) matching
Paid time off
+1
Senior AI/ML Distributed Training Engineer (PyTorch)
Senior AI/ML Distributed Training Engineer (PyTorch)

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 168,100 - 227,400
Health insurance
RSUs
Paid time off
+1
Senior AI/ML Distributed Training Engineer
Senior AI/ML Distributed Training Engineer

Amazon • Cupertino (CA)

On-site
USD 193,300 - 261,500