Senior AI/ML Distributed Training Engineer

Annapurna Labs (U.S.) Inc.

Cupertino (CA)

On-site

USD 180,000 - 240,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Annapurna Labs (U.S.) Inc. is seeking a Senior Machine Learning Engineer to join the AWS Neuron distributed training team. You will contribute to development, enablement, and performance tuning for large ML models, including GPT and other LLMs, on Trainium and Inferentia.

You'll work with chip architects, compilers, and runtime engineers to implement training support in PyTorch/JAX via XLA, and optimize for peak hardware efficiency.

Qualifications

  • Bachelor's degree in computer science or equivalent
  • 5+ years of non-internship professional software development experience
  • 5+ years of programming with at least one software programming language experience
  • 5+ years of leading design or architecture (design patterns, reliability and scaling) of new and existing systems experience
  • 5+ years of full software development life cycle, including coding standards, code reviews, source control management, build processes, testing

Responsibilities

  • Lead efforts to build distributed training support into PyTorch and JAX using XLA, the Neuron compiler, and runtime stacks.
  • You will optimize models to achieve peak performance and maximize efficiency on AWS Trainium/Inferentia hardware.
  • Collaborate with chip architects, compiler engineers and runtime engineers to create, build and tune distributed training solutions.

Skills

Python
Distributed training
PyTorch
JAX
Software development
System design

Education

Bachelor's degree in computer science or equivalent

Tools

XLA
Neuron compiler
Deepspeed
Nemo
PyTorch
JAX

Job description

Annapurna Labs (U.S.) Inc. is seeking a Senior Machine Learning Engineer to join the AWS Neuron distributed training team. You will contribute to development, enablement, and performance tuning for large ML models, including GPT and other LLMs, on Trainium and Inferentia.

You'll work with chip architects, compilers, and runtime engineers to implement training support in PyTorch/JAX via XLA, and optimize for peak hardware efficiency.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI/ML Software Engineer - Trainium Distributed Training
AI/ML Software Engineer - Trainium Distributed Training

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 168,000 - 227,000
Senior AI/ML Distributed Training Engineer
Senior AI/ML Distributed Training Engineer

Amazon • Cupertino (CA)

On-site
USD 193,300 - 261,500
Senior ML Systems Engineer - AI Inference on AWS Neuron
Senior ML Systems Engineer - AI Inference on AWS Neuron

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 193,000 - 262,000
AI/ML Software Engineer — Trainium Distributed Training
AI/ML Software Engineer — Trainium Distributed Training

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 144,000 - 194,000
Health insurance
401(k) matching
Paid time off
+1
Sr. Software Engineer- AI/ML, AWS Neuron Distributed Training
Sr. Software Engineer- AI/ML, AWS Neuron Distributed Training

Amazon • Cupertino (CA)

On-site
USD 193,300 - 261,500
Sr. Software Engineer- AI/ML, AWS Neuron Distributed Training - Performance Optimization
Sr. Software Engineer- AI/ML, AWS Neuron Distributed Training - Performance Optimization

Amazon • Seattle (WA)

On-site
USD 120,000 - 160,000
Health insurance
401(k) matching
Paid time off
+1
ML Systems Engineer - Distributed Training on Trainium
ML Systems Engineer - Distributed Training on Trainium

Socket.dev • Seattle (WA)

On-site
USD 144,000 - 194,000
Sr. Software Engineer- AI/ML, AWS Neuron Distributed Training
Sr. Software Engineer- AI/ML, AWS Neuron Distributed Training

Annapurna Labs (U.S.) Inc. • Cupertino (CA)

On-site
USD 180,000 - 240,000
Senior ML Compiler Engineer – Neuron
Senior ML Compiler Engineer – Neuron

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 168,000 - 227,000
ML Inference Engineer - AWS Neuron & GenAI
ML Inference Engineer - AWS Neuron & GenAI

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 165,000 - 224,000