Software Development Engineer, AI/ML, AWS Neuron, Model Inference

Amazon

Cupertino (CA)

On-site

USD 180,000 - 240,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Amazon’s Annapurna Labs team, part of AWS, seeks a Software Development Engineer for AI/ML focused on AWS Neuron device acceleration. You will contribute to the Neuron SDK, optimize inference and training workloads on Inferentia/Trainium, and work across compiler, runtime, framework, and hardware teams.

The role emphasizes designing high-performance ML kernels, distributed inference support for PyTorch, and close collaboration with customers to enable efficient model deployment on AWS

Qualifications

  • Bachelor's degree in computer science or equivalent.
  • 5+ years of professional software development experience.
  • 5+ years of design or architecture experience for new and existing systems.
  • Fundamentals of machine learning and large-language models, including architecture, training, and inference lifecycles, and experience optimizing model execution.
  • Software development experience in C++ and Python (experience in at least one language required).
  • Strong understanding of system performance, memory management, and parallel computing principles.
  • Proficiency in debugging, profiling, and implementing best software engineering practices in large-scale systems.

Responsibilities

  • Design, develop, and optimize machine learning models and frameworks for deployment on custom ML hardware accelerators.
  • Participate in all stages of the ML system development lifecycle – distributed computing architecture design, implementation, performance profiling, hardware-specific optimizations, testing, and production deployment.
  • Build infrastructure to systematically analyze and onboard multiple models with diverse architectures.
  • Design and implement high-performance kernels and features for ML operations, leveraging the Neuron architecture and programming models.
  • Analyze and optimize system-level performance across multiple generations of Neuron hardware.
  • Conduct detailed performance analysis using profiling tools to identify and resolve bottlenecks.
  • Implement optimizations such as fusion, sharding, tiling, and scheduling.
  • Conduct comprehensive testing, including unit and end-to-end model testing with continuous deployment and releases through pipelines.
  • Work directly with customers to enable and optimize their ML models on AWS accelerators.
  • Collaborate across teams to develop innovative optimization techniques.

Skills

C++
Python
Large-scale systems
Performance optimization
Debugging & profiling
Parallel computing
ML fundamentals
Architecture design

Education

Bachelor's degree in computer science or equivalent

Tools

CUDA
CUTLASS
FlashInfer
TensorRT
PyTorch
JIT compilation
AOT tracing

Job description

Software Development Engineer, AI/ML, AWS Neuron, Model Inference

The Annapurna Labs team at Amazon Web Services (AWS) builds AWS Neuron, the software development kit that accelerates deep learning and GenAI workloads on Amazon’s custom machine learning accelerators, Inferentia and Trainium. The AWS Neuron SDK is the backbone for accelerating deep learning and GenAI workloads, offering an ML compiler, runtime, and application framework that integrates with popular ML frameworks like PyTorch and JAX to deliver top‑performance inference and training.

The Inference Enablement and Acceleration team focuses on running a wide range of models and supporting new architectures while maximizing performance on AWS’s custom ML accelerators. Working across the stack from PyTorch to the hardware‑software boundary, engineers build systematic infrastructure, innovate new methods, and create high‑performance kernels, ensuring every compute unit is fine‑tuned for optimal performance. This role offers the chance to work at the intersection of machine learning, high‑performance computing, and distributed architectures, shaping the future of AI acceleration technology.

Key responsibilities include leading distributed inference support for PyTorch in the Neuron SDK, tuning these models for highest performance on AWS Trainium and Inferentia silicon, developing low‑level optimizations, and collaborating with compiler, runtime, framework, and hardware teams to optimize machine learning workloads for the global customer base.

Key job responsibilities
  • Design, develop, and optimize machine learning models and frameworks for deployment on custom ML hardware accelerators.
  • Participate in all stages of the ML system development lifecycle – distributed computing architecture design, implementation, performance profiling, hardware‑specific optimizations, testing, and production deployment.
  • Build infrastructure to systematically analyze and onboard multiple models with diverse architectures.
  • Design and implement high‑performance kernels and features for ML operations, leveraging the Neuron architecture and programming models.
  • Analyze and optimize system‑level performance across multiple generations of Neuron hardware.
  • Conduct detailed performance analysis using profiling tools to identify and resolve bottlenecks.
  • Implement optimizations such as fusion, sharding, tiling, and scheduling.
  • Conduct comprehensive testing, including unit and end‑to‑end model testing with continuous deployment and releases through pipelines.
  • Work directly with customers to enable and optimize their ML models on AWS accelerators.
  • Collaborate across teams to develop innovative optimization techniques.
Basic Qualifications
  • Bachelor's degree in computer science or equivalent.
  • 5+ years of professional software development experience.
  • 5+ years of design or architecture experience for new and existing systems.
  • Fundamentals of machine learning and large‑language models, including architecture, training, and inference lifecycles, and experience optimizing model execution.
  • Software development experience in C++ and Python (experience in at least one language required).
  • Strong understanding of system performance, memory management, and parallel computing principles.
  • Proficiency in debugging, profiling, and implementing best software engineering practices in large‑scale systems.
Preferred Qualifications
  • Familiarity with PyTorch, JIT compilation, and AOT tracing.
  • Familiarity with CUDA kernels or equivalent low‑level kernels.
  • Experience with performant kernel development such as CUTLASS or FlashInfer.
  • Knowledge of syntax and tile‑level semantics similar to Triton.
  • Experience with online/offline inference serving using vLLM, SGLang, TensorRT, or similar platforms in production.
  • Deep understanding of computer architecture, operating systems, and parallel computing.
About the team

The Inference Enablement and Acceleration team fosters a builder’s culture where experimentation is encouraged, impact is measurable, and collaboration, technical ownership, and continuous learning are valued. The team provides mentorship, thorough but kind code reviews, and projects that develop engineering expertise, empowering members to handle complex tasks.

Resources

Learn more about Neuron:

https://awsdocs-neuron.readthedocs-hosted.com/en/latest/neuron-guide/neuron-cc/index.html

https://aws.amazon.com/machine-learning/neuron/

https://github.com/aws/aws-neuron-sdk

https://www.amazon.science/how-silicon-innovation-became-the-secret-sauce-behind-awss-success

Equal Opportunity

Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Software Development Engineer, AI/ML, AWS Neuron, Model Inference
Senior Software Development Engineer, AI/ML, AWS Neuron, Model Inference

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 193,000 - 262,000
Health insurance
401(k) matching
Paid time off
+1
Senior Software Development Engineer, AI/ML, AWS Neuron, Model Inference
Senior Software Development Engineer, AI/ML, AWS Neuron, Model Inference

Amazon • Cupertino (CA)

On-site
USD 193,300 - 261,500
Software Development Engineer, AI/ML, AWS Neuron, Model Inference
Software Development Engineer, AI/ML, AWS Neuron, Model Inference

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 165,000 - 224,000
Software Engineer - AI/ML, AWS Neuron Apps
Software Engineer - AI/ML, AWS Neuron Apps

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 120,000 - 160,000
Sr. Software Engineer- AI/ML, AWS Neuron Apps
Sr. Software Engineer- AI/ML, AWS Neuron Apps

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 168,000 - 228,000
Software Development Engineer I, ML Infra Services, Annapurna Labs (AWS)
Software Development Engineer I, ML Infra Services, Annapurna Labs (AWS)

Amazon • Cupertino (CA)

On-site
USD 150,000 - 190,000
Sr. Software Engineer- AI/ML, AWS Neuron Distributed Training - Performance Optimization
Sr. Software Engineer- AI/ML, AWS Neuron Distributed Training - Performance Optimization

Amazon • Seattle (WA)

On-site
USD 120,000 - 160,000
Health insurance
401(k) matching
Paid time off
+1
Software Engineering Manager, ML Kernel Performance, AWS Neuron, Annapurna Labs
Software Engineering Manager, ML Kernel Performance, AWS Neuron, Annapurna Labs

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 212,000 - 288,000
Senior Software Engineer - AI/ML, AWS Neuron Inference
Senior Software Engineer - AI/ML, AWS Neuron Inference

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 168,000 - 228,000
Applied Scientist, AWS Neuron Science team
Applied Scientist, AWS Neuron Science team

Amazon • Cupertino (CA)

On-site
USD 171,600 - 222,200
RSUs
Health insurance
401(k) matching
+1