AI/ML Software Engineer: Neuron Inference & Acceleration

Amazon Inc.

Seattle (WA)

On-site

USD 165,000 - 224,000

Full time

45 hours ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Amazon Inc. in Cupertino, CA is seeking a Software Development Engineer for AI/ML in AWS Neuron, Model Inference.

You will design, implement, and optimize distributed ML inference software for Inferentia and Trainium accelerators, collaborating across compiler, runtime, framework, and hardware teams. The role requires deep knowledge of C++ and Python, strong system performance skills, and hands-on experience with ML frameworks like PyTorch.

Qualifications

  • Bachelor's degree in computer science or equivalent.
  • 3+ years of non-internship professional software development experience.
  • Fundamentals of Machine learning and LLMs, their architecture, training and inference lifecycles.

Responsibilities

  • Design, develop, and optimize machine learning models and frameworks for deployment on custom ML hardware accelerators.
  • Participate in all stages of the ML system development lifecycle including distributed computing based architecture design, implementation, performance profiling, hardware-specific optimizations, testing and production deployment.
  • Build infrastructure to systematically analyze and onboard multiple models with diverse architecture.
  • Design and implement high-performance kernels and features for ML operations, leveraging the Neuron architecture and programming models.
  • Analyze and optimize system-level performance across multiple generations of Neuron hardware.
  • Conduct detailed performance analysis using profiling tools to identify and resolve bottlenecks.
  • Implement optimizations such as fusion, sharding, tiling, and scheduling.
  • Conduct comprehensive testing, including unit and end-to-end model testing with continuous deployment and releases through pipelines.
  • Work directly with customers to enable and optimize their ML models on AWS accelerators.
  • Collaborate across teams to develop innovative optimization techniques.

Skills

Python
C++
Machine learning
System performance
Debugging
Parallel computing

Education

Bachelor's degree in computer science or equivalent
Master's degree or Ph.D. in Computer Science or equivalent

Tools

PyTorch
JIT compilation
AOT tracing
CUDA kernels
TensorRT

Job description

Amazon Inc. in Cupertino, CA is seeking a Software Development Engineer for AI/ML in AWS Neuron, Model Inference.

You will design, implement, and optimize distributed ML inference software for Inferentia and Trainium accelerators, collaborating across compiler, runtime, framework, and hardware teams. The role requires deep knowledge of C++ and Python, strong system performance skills, and hands-on experience with ML frameworks like PyTorch.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior AI/ML Inference Engineer (Neuron)
Senior AI/ML Inference Engineer (Neuron)

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 193,000 - 262,000
Health insurance
401(k) matching
Paid time off
Senior ML Accelerator Runtime Engineer
Senior ML Accelerator Runtime Engineer

Amazon • Seattle (WA)

On-site
USD 143,700 - 194,400
Health insurance
Dental
Vision
+4
Senior ML Systems Engineer - AI Inference on AWS Neuron
Senior ML Systems Engineer - AI Inference on AWS Neuron

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 193,000 - 262,000
ML Inference Engineer - AWS Neuron & GenAI
ML Inference Engineer - AWS Neuron & GenAI

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 165,000 - 224,000
AI/ML Inference Engineer (Trainium)
AI/ML Inference Engineer (Trainium)

Amazon • Cupertino (CA)

On-site
USD 165,000 - 224,000
Health insurance
401(k) matching
Paid time off
+1
Senior AI/ML Software Engineer - High-Perf Inference
Senior AI/ML Software Engineer - High-Perf Inference

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 168,000 - 227,000
AI/ML Systems Engineer for AWS Neuron Inference
AI/ML Systems Engineer for AWS Neuron Inference

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 144,000 - 194,000
Health insurance
401(k) matching
Paid time off
+1
Senior ML Compiler Engineer – Build Next‑Gen ML on AI Accelerators
Senior ML Compiler Engineer – Build Next‑Gen ML on AI Accelerators

Amazon • Cupertino (CA)

On-site
USD 193,000 - 262,000
Health benefits
401(k) matching
Parental leave
Neuron Runtime Engineer - ML Accelerators & Profiling
Neuron Runtime Engineer - ML Accelerators & Profiling

Amazon • Seattle (WA)

On-site
USD 144,000 - 194,000
Health insurance
401(k) matching
Paid time off
+1
Software Engineer - AI/ML, AWS Neuron Apps
Software Engineer - AI/ML, AWS Neuron Apps

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 120,000 - 160,000