AI/ML Systems Engineer - Neuron Inference

Amazon

Seattle (WA)

On-site

USD 144,000 - 194,000

Full time

7 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Amazon's Annapurna Labs team is seeking a senior software engineer to design and optimize ML models and frameworks for Neuron, AWS's ML accelerators. You will work across PyTorch/JAX stacks, implement high-performance kernels, and collaborate with customers to tune inference and training workloads.

You will lead performance profiling, system-level optimizations, and contribute to future architecture. A startup-like, mentorship-rich culture with hands-on coding and cross-team collaboration awaits.

Qualifications

  • 3+ years of non-internship professional software development experience.
  • Bachelor's degree in computer science or equivalent.
  • 3+ years of design or architecture experience in large systems.
  • Fundamentals of ML and LLMs, including model lifecycles and optimizations.
  • Proficiency in C++ and Python; strong debugging and profiling skills.

Responsibilities

  • Design, develop, and optimize ML models and frameworks for custom ML hardware accelerators.
  • Participate in all stages of ML system development including distributed architecture design and production deployment.
  • Build infrastructure to analyze and onboard multiple models with diverse architecture.
  • Design and implement high-performance kernels for ML operations using Neuron.
  • Analyze and optimize system-level performance across Neuron hardware generations.
  • Conduct detailed performance analysis to identify and resolve bottlenecks.
  • Implement optimizations such as fusion, sharding, tiling, and scheduling.
  • Collaborate with customers to enable and optimize ML models on AWS accelerators.

Skills

C++
Python
Machine learning basics
LLMs
Performance optimization
Debugging & profiling
Distributed systems

Education

Bachelor's degree in computer science or equivalent

Tools

PyTorch
JIT compilation
AOT tracing
CUDA kernels
TensorRT

Job description

Amazon's Annapurna Labs team is seeking a senior software engineer to design and optimize ML models and frameworks for Neuron, AWS's ML accelerators. You will work across PyTorch/JAX stacks, implement high-performance kernels, and collaborate with customers to tune inference and training workloads.

You will lead performance profiling, system-level optimizations, and contribute to future architecture. A startup-like, mentorship-rich culture with hands-on coding and cross-team collaboration awaits.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior ML Systems Engineer - AI Inference on AWS Neuron
Senior ML Systems Engineer - AI Inference on AWS Neuron

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 193,000 - 262,000
ML Inference Engineer - AWS Neuron & GenAI
ML Inference Engineer - AWS Neuron & GenAI

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 165,000 - 224,000
Senior AI/ML Software Engineer - Neuron Optimizations
Senior AI/ML Software Engineer - Neuron Optimizations

Annapurna Labs (U.S.) Inc. • Seattle (WA)

On-site
USD 180,000 - 230,000
Senior AI/ML Systems Engineer for Custom Accelerators
Senior AI/ML Systems Engineer for Custom Accelerators

Amazon • Seattle (WA)

On-site
USD 168,000 - 227,000
Health insurance
401(k) matching
Senior Software Engineer — GenAI & ML Acceleration
Senior Software Engineer — GenAI & ML Acceleration

Amazon Web Services (AWS) • New York (NY)

On-site
USD 185,000 - 250,000
Health insurance
401(k) matching
Paid time off
+2
Senior AI/ML Systems Engineer for Accelerator Optimization
Senior AI/ML Systems Engineer for Accelerator Optimization

Socket.dev • Seattle (WA)

On-site
USD 168,000 - 227,000
Senior ML Compiler Engineer – Neuron
Senior ML Compiler Engineer – Neuron

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 168,000 - 227,000
Senior ML Inference Engineer for Trainium & Neuron
Senior ML Inference Engineer for Trainium & Neuron

Artha Nexgen • Cupertino (CA)

Hybrid
USD 193,000 - 262,000
Senior ML Compiler Engineer – Build Next‑Gen ML on AI Accelerators
Senior ML Compiler Engineer – Build Next‑Gen ML on AI Accelerators

Amazon • Cupertino (CA)

On-site
USD 193,000 - 262,000
Health benefits
401(k) matching
Parental leave
Senior Software Development Engineer, AI/ML, AWS Neuron, Model Inference
Senior Software Development Engineer, AI/ML, AWS Neuron, Model Inference

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 193,000 - 262,000