AI/ML Inference Engineer (Trainium)

Amazon

Cupertino (CA)

On-site

USD 165,000 - 224,000

Full time

5 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Health insurance
401(k) matching
Paid time off
Parental leave

Job summary

Amazon’s Annapurna Labs/Neuron team seeks a Software Engineer II - AI/ML to optimize inference workloads on AWS Trainium. You will design and tune ML models and kernels, collaborate with compiler, runtime, and framework teams, and enable customers to deploy high‑performance AI workloads.

You will work across PyTorch, kernel development, and distributed systems, improving performance and scalability on cutting‑edge ML accelerators. A strong ML foundation and hands-on software skills are essential.

Qualifications

  • 3+ years of non‑internship professional software development experience.
  • 3+ years of non‑internship design or architecture experience.
  • 1+ years designing and developing large-scale distributed software applications.
  • Bachelor's degree or foreign equivalent in Computer Science, Engineering, Mathematics, or related field.
  • Experience debugging, profiling, and optimizing large-scale systems.
  • Experience with Machine Learning fundamentals, including training/inference lifecycles.
  • Knowledge of Python and/or C++ programming.
  • Strong understanding of system performance, memory management, and parallel computing.

Responsibilities

  • Design, develop, and optimize ML models and frameworks for deployment on custom ML hardware accelerators.
  • Participate in ML system development lifecycle including architecture design, performance profiling, optimization, testing, and deployment.
  • Build infrastructure to analyze and onboard multiple models with diverse architectures.
  • Understand and implement NKI kernels and high-performance ML kernels for Neuron architecture.
  • Analyze system performance across hardware generations to guide optimizations.
  • Conduct detailed profiling to identify bottlenecks and tune performance.
  • Implement optimizations such as fusion, sharding, tiling, and scheduling.
  • Perform comprehensive testing including unit and end-to-end testing with CI/CD pipelines.
  • Collaborate with customers to enable and optimize ML models on AWS accelerators.

Skills

Software development
Distributed systems
Architecture design
Python
C++
ML fundamentals

Education

Bachelor's degree in Computer Science or related field

Tools

CUDA kernels
PyTorch
TensorRT
Triton

Job description

Amazon’s Annapurna Labs/Neuron team seeks a Software Engineer II - AI/ML to optimize inference workloads on AWS Trainium. You will design and tune ML models and kernels, collaborate with compiler, runtime, and framework teams, and enable customers to deploy high‑performance AI workloads.

You will work across PyTorch, kernel development, and distributed systems, improving performance and scalability on cutting‑edge ML accelerators. A strong ML foundation and hands-on software skills are essential.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI/ML Inference Engineer for PyTorch on Trainium
AI/ML Inference Engineer for PyTorch on Trainium

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 150,000 - 210,000
Senior AI/ML Software Engineer: Inference on Trainium
Senior AI/ML Software Engineer: Inference on Trainium

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 193,000 - 262,000
RSUs
Health insurance
401(k) matching
+1
Software Engineer II: AI/ML Inference on Custom Hardware
Software Engineer II: AI/ML Inference on Custom Hardware

Amazon Inc. • Cupertino (CA)

On-site
USD 165,000 - 224,000
Health insurance
RSUs
401(k) matching
+2
AI/ML Systems Engineer for AWS Neuron Inference
AI/ML Systems Engineer for AWS Neuron Inference

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 144,000 - 194,000
Health insurance
401(k) matching
Paid time off
+1
Senior ML Systems Engineer - AI Inference on AWS Neuron
Senior ML Systems Engineer - AI Inference on AWS Neuron

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 193,000 - 262,000
AI/ML Software Engineer — Trainium Distributed Training
AI/ML Software Engineer — Trainium Distributed Training

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 144,000 - 194,000
Health insurance
401(k) matching
Paid time off
+1
Senior AI/ML Inference Engineer (Neuron)
Senior AI/ML Inference Engineer (Neuron)

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 193,000 - 262,000
Health insurance
401(k) matching
Paid time off
AI/ML Systems Engineer – Distributed Training on Trainium
AI/ML Systems Engineer – Distributed Training on Trainium

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 165,000 - 224,000
Senior ML Kernel Performance Engineer - AI Accelerator
Senior ML Kernel Performance Engineer - AI Accelerator

Amazon • Cupertino (CA)

On-site
USD 193,000 - 262,000
AI/ML Software Engineer - Distributed Training HPC
AI/ML Software Engineer - Distributed Training HPC

Amazon • Cupertino (CA), Northern (KY)

Hybrid
USD 165,000 - 224,000