AI/ML Inference Engineer for AWS Neuron

Amazon

Cupertino (CA)

On-site

USD 165,000 - 224,000

Full time

12 hours ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Health insurance
401(k) matching
Paid time off
Parental leave
RSUs

Job summary

Amazon's Annapurna Labs team at AWS is seeking a Software Development Engineer focused on AI/ML for AWS Neuron, the runtime and compiler stack that accelerates models on Inferentia and Trainium. You will design high-performance kernels, optimize inference across large models, and collaborate with scientists, engineers, and customers to push the boundaries of AI acceleration.

This role offers a startup-like culture with mentorship, code reviews, and opportunities to impact state-of-the-art

Qualifications

  • Bachelor's degree in computer science or equivalent.
  • 3+ years of non-internship professional software development experience.
  • 3+ years of non-internship design or architecture experience.
  • Fundamentals of Machine learning and LLMs, their architecture, training and inference lifecycles with optimization experience.
  • Software development in C++, Python (at least one language required).
  • Strong understanding of system performance, memory management, and parallel computing principles.
  • Proficiency in debugging and scalable software engineering practices.

Responsibilities

  • Design, develop, and optimize ML models and frameworks for deployment on custom ML hardware accelerators.
  • Participate in all stages of ML system development lifecycle including distributed architecture design, performance profiling, and production deployment.
  • Build infrastructure to analyze and onboard multiple models with diverse architectures.
  • Design and implement high-performance kernels and features for ML operations leveraging Neuron architecture.
  • Analyze and optimize system-level performance across Neuron hardware generations.
  • Conduct detailed performance analysis using profiling tools to identify bottlenecks.
  • Implement optimizations such as fusion, sharding, tiling, and scheduling.
  • Conduct comprehensive testing including unit and end-to-end model testing with CI/CD pipelines.
  • Work directly with customers to enable and optimize their ML models on AWS accelerators.
  • Collaborate across teams to develop innovative optimization techniques.

Skills

C++
Python
System performance
Debugging
Parallel computing

Education

Bachelor's degree in computer science

Tools

PyTorch
JAX
CUDA kernels
TensorRT
CUTLASS

Job description

Amazon's Annapurna Labs team at AWS is seeking a Software Development Engineer focused on AI/ML for AWS Neuron, the runtime and compiler stack that accelerates models on Inferentia and Trainium. You will design high-performance kernels, optimize inference across large models, and collaborate with scientists, engineers, and customers to push the boundaries of AI acceleration.

This role offers a startup-like culture with mentorship, code reviews, and opportunities to impact state-of-the-art

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI/ML Systems Engineer - Neuron Inference
AI/ML Systems Engineer - Neuron Inference

Amazon • Seattle (WA)

On-site
USD 144,000 - 194,000
ML Inference Engineer - AWS Neuron & GenAI
ML Inference Engineer - AWS Neuron & GenAI

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 165,000 - 224,000
Senior ML Systems Engineer - AI Inference on AWS Neuron
Senior ML Systems Engineer - AI Inference on AWS Neuron

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 193,000 - 262,000
Senior AI/ML Software Engineer - High-Perf Inference
Senior AI/ML Software Engineer - High-Perf Inference

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 168,000 - 227,000
Senior AI/ML Inference Engineer for Neuron SDK
Senior AI/ML Inference Engineer for Neuron SDK

Amazon • Cupertino (CA)

On-site
USD 193,000 - 262,000
Health insurance
RSU program
Senior AI/ML Software Engineer - Neuron Optimizations
Senior AI/ML Software Engineer - Neuron Optimizations

Annapurna Labs (U.S.) Inc. • Seattle (WA)

On-site
USD 180,000 - 230,000
Senior ML Kernel Performance Engineer - AI Accelerator
Senior ML Kernel Performance Engineer - AI Accelerator

Amazon • Cupertino (CA)

On-site
USD 193,000 - 262,000
Senior AI/ML Systems Engineer for Accelerator Optimization
Senior AI/ML Systems Engineer for Accelerator Optimization

Socket.dev • Seattle (WA)

On-site
USD 168,000 - 227,000
Senior ML Accelerator Runtime Engineer
Senior ML Accelerator Runtime Engineer

Amazon • Seattle (WA)

On-site
USD 143,700 - 194,400
Health insurance
Dental
Vision
+4
Senior ML Inference Engineer — LLMs on Neuron
Senior ML Inference Engineer — LLMs on Neuron

Amazon.com Services LLC • Seattle (WA)

On-site
USD 168,000 - 227,000
Health insurance
401(k) matching
Paid time off
+1