Senior ML Software Engineer, AI Accelerator & Inference

Amazon Web Services (AWS)

Seattle (WA)

On-site

USD 168,100 - 227,400

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Annapurna Labs (U.S.) Inc. in Seattle seeks a Senior Software Engineer to shape the evolution of AWS Neuron, AWS's AI accelerator stack. You will drive distributed AI across ML frameworks and hardware, focusing on high-performance inference for state-of-the-art models.

Join a team that bridges PyTorch and JAX with XLA, optimizes Trainium and Inferentia, and creates tools improving LLM accuracy and efficiency at scale.

Qualifications

  • 5+ years of full software development lifecycle experience.
  • 5+ years of programming experience with Python or C++ and PyTorch.
  • Experience with AI acceleration via quantization, parallelism, model compression, batching, KV caching, vllm serving.
  • Experience with accuracy debugging & tooling and performance benchmarking of AI accelerators.
  • Fundamentals of ML and DL models, their architecture, training and inference lifecycles, plus optimization work.

Responsibilities

  • Drive Evolution of Distributed AI at AWS Neuron, bridging ML frameworks and hardware.
  • Architect the bridge between PyTorch, JAX and AI hardware; push scalable AI model execution.
  • Spearhead benchmark methodologies and performance tuning for AI accelerators.
  • Develop ML tools to improve LLM accuracy and efficiency.
  • Collaborate with silicon architects to influence accelerator design.

Skills

Distributed systems
AI acceleration
Performance optimization
Debugging tooling
Inference optimization

Education

Bachelor's degree in CS or equivalent
Master's degree in CS or ML or equivalent

Tools

Python
C++
PyTorch

Job description

Annapurna Labs (U.S.) Inc. in Seattle seeks a Senior Software Engineer to shape the evolution of AWS Neuron, AWS's AI accelerator stack. You will drive distributed AI across ML frameworks and hardware, focusing on high-performance inference for state-of-the-art models.

Join a team that bridges PyTorch and JAX with XLA, optimizes Trainium and Inferentia, and creates tools improving LLM accuracy and efficiency at scale.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI/ML Inference Engineer for Neuron on AWS
Senior AI/ML Inference Engineer for Neuron on AWS

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 193,000 - 262,000
Health insurance
401(k) matching
Paid time off
+1
AI/ML Inference Engineer for AWS Neuron
AI/ML Inference Engineer for AWS Neuron

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 165,000 - 224,000
Senior ML Accelerator Solutions Architect
Senior ML Accelerator Solutions Architect

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 176,000 - 239,000
Senior AI/ML Software Engineer — Distributed Training
Senior AI/ML Software Engineer — Distributed Training

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 168,000 - 228,000
RSUs
Health benefits
Sign-on payments
Senior AI/ML Inference Engineer for LLMs on Neuron
Senior AI/ML Inference Engineer for LLMs on Neuron

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 168,000 - 228,000
AI/ML SDE — High-Performance Inference on Custom Hardware
AI/ML SDE — High-Performance Inference on Custom Hardware

Amazon • Cupertino (CA)

On-site
USD 180,000 - 240,000
Senior Software Engineer - AI/ML, AWS Neuron Inference
Senior Software Engineer - AI/ML, AWS Neuron Inference

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 168,000 - 228,000
Senior AI/ML Distributed Training Engineer (PyTorch)
Senior AI/ML Distributed Training Engineer (PyTorch)

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 168,100 - 227,400
Health insurance
RSUs
Paid time off
+1
Senior AI/ML Systems Engineer - Neuron Inference
Senior AI/ML Systems Engineer - Neuron Inference

Amazon • Cupertino (CA)

On-site
USD 193,300 - 261,500
Applied Scientist II — ML Systems for AI Accelerators
Applied Scientist II — ML Systems for AI Accelerators

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 171,000 - 223,000