AI/ML Systems Engineer for AWS Neuron Inference

Amazon Web Services (AWS)

Seattle (WA)

On-site

USD 144,000 - 194,000

Full time

27 hours ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Health insurance
401(k) matching
Paid time off
Parental leave

Job summary

Annapurna Labs (U.S.) Inc. in Seattle seeks an experienced software engineer to advance inference for Generative AI on AWS accelerators. You will work on distributed inference with PyTorch in the Neuron SDK, tuning models for Trainium and Inferentia, and collaborating across compiler, runtime, and hardware teams.

You will design high‑performance kernels, optimize memory and parallel computing, and contribute to a startup‑like culture that values experimentation and mentorship.

Qualifications

  • 3+ years of non‑internship professional software development.
  • Bachelor's degree in computer science or equivalent.
  • Experience with C++, Python and ML model acceleration.
  • Strong knowledge of system performance, memory management and parallel computing.
  • Experience debugging and profiling large scale systems.

Responsibilities

  • Design, develop, and optimize ML models and frameworks for AWS accelerators.
  • Collaborate across compiler, runtime, framework and hardware teams.
  • Tune models for highest performance on Trainium and Inferentia.
  • Implement high‑performance kernels and ML operations leveraging Neuron.
  • Analyze system performance and drive end‑to‑end optimizations.

Skills

C++
Python
PyTorch
LLMs
Machine learning basics
Profiling
Debugging
Performance optimization
Parallel computing
Memory management

Education

Bachelor's degree in CS or equivalent

Tools

CUDA kernels
CUTLASS
JIT compilation
TensorRT

Job description

Annapurna Labs (U.S.) Inc. in Seattle seeks an experienced software engineer to advance inference for Generative AI on AWS accelerators. You will work on distributed inference with PyTorch in the Neuron SDK, tuning models for Trainium and Inferentia, and collaborating across compiler, runtime, and hardware teams.

You will design high‑performance kernels, optimize memory and parallel computing, and contribute to a startup‑like culture that values experimentation and mentorship.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI/ML Inference Engineer for AWS Neuron
AI/ML Inference Engineer for AWS Neuron

Amazon • Cupertino (CA)

On-site
USD 165,000 - 224,000
Health insurance
401(k) matching
Paid time off
+2
AI/ML Systems Engineer - Neuron Inference
AI/ML Systems Engineer - Neuron Inference

Amazon • Seattle (WA)

On-site
USD 144,000 - 194,000
Senior ML Systems Engineer - AI Inference on AWS Neuron
Senior ML Systems Engineer - AI Inference on AWS Neuron

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 193,000 - 262,000
Senior AI/ML Software Engineer - High-Perf Inference
Senior AI/ML Software Engineer - High-Perf Inference

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 168,000 - 227,000
ML Inference Engineer - AWS Neuron & GenAI
ML Inference Engineer - AWS Neuron & GenAI

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 165,000 - 224,000
Senior AI/ML Inference Engineer for Neuron SDK
Senior AI/ML Inference Engineer for Neuron SDK

Amazon • Cupertino (CA)

On-site
USD 193,000 - 262,000
Health insurance
RSU program
Senior AI/ML Distributed Training Engineer
Senior AI/ML Distributed Training Engineer

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 168,000 - 227,000
Health insurance
401(k) matching
Paid time off
+1
Senior ML Compiler Engineer – Neuron
Senior ML Compiler Engineer – Neuron

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 168,000 - 227,000
GenAI ML Systems Engineer
GenAI ML Systems Engineer

Amazon Web Services (AWS) • New York (NY)

On-site
USD 158,000 - 214,000
Senior Software Engineer — GenAI & ML Acceleration
Senior Software Engineer — GenAI & ML Acceleration

Amazon Web Services (AWS) • New York (NY)

On-site
USD 185,000 - 250,000
Health insurance
401(k) matching
Paid time off
+2