AI/ML Performance Engineer – Training on Trainium

Amazon Web Services (AWS)

Cupertino (CA)

On-site

USD 140,000 - 210,000

Full time

11 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health Insurance
Medical Insurance
Dental Insurance
Vision Insurance
Prescription Insurance
Basic Life & AD&D Insurance
Supplemental Life Plans
EAP
Mental Health Support
Medical Advice Line
Flexible Spending Accounts
Adoption and Surrogacy Reimbursement
401(k) Matching
Paid Time Off
Parental Leave
Sign-on Payments
Restricted Stock Units (RSUs)

Job summary

Amazon Web Services (AWS) is seeking a software professional to lead optimization of distributed training on AWS Trainium, aiming to maximize throughput and model FLOPs utilization. You will diagnose bottlenecks across the stack, from collective communications to compiler optimizations, and drive end-to-end improvements.

Requirements include 3+ years of software development experience and 2+ years in system design or architecture, with proficiency in at least one programming language.

Qualifications

  • Requires 3+ years of professional software development experience and 2+ years of experience in system design or architecture.
  • Proficiency in at least one programming language is required.
  • Preference for candidates with a Bachelor's degree in Computer Science.

Responsibilities

  • Lead efforts to optimize distributed training performance on AWS Trainium.
  • Identify and resolve performance bottlenecks across the software stack, from collective communications to compiler optimizations.

Skills

Distributed Training
Performance Optimization
PyTorch
JAX
AWS Neuron
Machine Learning
Compiler Optimization
Kernel Performance
System Architecture
Software Development Life Cycle
Collective Communications
Memory Utilization

Education

Bachelor's degree in Computer Science

Job description

Amazon Web Services (AWS) is seeking a software professional to lead optimization of distributed training on AWS Trainium, aiming to maximize throughput and model FLOPs utilization. You will diagnose bottlenecks across the stack, from collective communications to compiler optimizations, and drive end-to-end improvements.

Requirements include 3+ years of software development experience and 2+ years in system design or architecture, with proficiency in at least one programming language.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI/ML Software Engineer — Trainium Distributed Training
AI/ML Software Engineer — Trainium Distributed Training

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 144,000 - 194,000
Health insurance
401(k) matching
Paid time off
+1
AI Training Systems Engineer - Collective Compute
AI Training Systems Engineer - Collective Compute

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 130,000 - 170,000
Sign-on payments
RSUs
Health insurance
+11
Senior AI/ML Distributed Training Engineer
Senior AI/ML Distributed Training Engineer

Amazon • Cupertino (CA)

On-site
USD 193,300 - 261,500
AI/ML Software Engineer - Trainium Distributed Training
AI/ML Software Engineer - Trainium Distributed Training

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 168,000 - 227,000
AI/ML Software Engineer for High-Performance Training
AI/ML Software Engineer for High-Performance Training

Amazon • Cupertino (CA)

On-site
USD 165,000 - 224,000
Software Engineer- AI/ML, AWS Neuron Distributed Training - Performance Optimization
Software Engineer- AI/ML, AWS Neuron Distributed Training - Performance Optimization

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 140,000 - 210,000
Health Insurance
Medical Insurance
Dental Insurance
+14
ML Systems Engineer - Distributed Training on Trainium
ML Systems Engineer - Distributed Training on Trainium

Socket.dev • Seattle (WA)

On-site
USD 144,000 - 194,000
Sr. Software Engineer- AI/ML, AWS Neuron Distributed Training - Performance Optimization
Sr. Software Engineer- AI/ML, AWS Neuron Distributed Training - Performance Optimization

Amazon • Seattle (WA)

On-site
USD 120,000 - 160,000
Health insurance
401(k) matching
Paid time off
+1
ML Systems Engineer for Trainium Inference
ML Systems Engineer for Trainium Inference

AI Chopping Block • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 280,000
Software Engineer, AI Training Collectives
Software Engineer, AI Training Collectives

Amazon • Cupertino (CA)

On-site
USD 165,000 - 224,000