ML Inference Systems Engineer — High-Performance Serving

Annapurna Labs (U.S.) Inc.

Seattle (WA)

On-site

USD 144,000 - 194,000

Full time

2 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

RSUs
401(k) matching
Paid time off

Job summary

Amazon in Seattle is seeking engineers to advance the Neuron-based serving stack, enabling fast, scalable model inference across a wide range of models. You will contribute core features to open-source frameworks, optimize performance on Neuron hardware, and broaden model support with robust tooling and validation.

The role emphasizes collaboration across model development, performance, and runtime teams, with a focus on end-to-end efficiency, reliability, and release readiness in production

Qualifications

  • 3+ years of non-internship professional software development experience
  • 3+ years of design or architecture experience for large-scale systems
  • 2+ years designing multi-tiered, distributed software applications
  • Bachelor's degree or foreign equivalent in CS/Engineering/Math or related field
  • Knowledge of ML and LLM fundamentals, transformer architectures, and optimization techniques

Responsibilities

  • Contribute core serving features to open-source inference frameworks such as vLLM and SGLang, implementing and upstreaming support for continuous batching, paged attention, quantization, and distributed inference on Neuron
  • Broaden the range of models supported out of the box and build tooling that shortens the path from a new model to production deployment
  • Improve model development and shipping velocity by reducing enablement time and strengthening test, benchmarking, and release workflows
  • Collaborate with model development, performance, compiler, and runtime engineers to deliver end-to-end model performance across a broad range of models and workloads
  • Apply strong engineering practices — code reviews, testing, and operational excellence — to ship reliable, high-performance inference

Skills

C++
Java
C#
Perl
Distributed systems
Machine learning basics

Education

Bachelor's degree in Computer Science or related field

Tools

vLLM
SGLang
TensorRT
CUDA
Triton

Job description

Amazon in Seattle is seeking engineers to advance the Neuron-based serving stack, enabling fast, scalable model inference across a wide range of models. You will contribute core features to open-source frameworks, optimize performance on Neuron hardware, and broaden model support with robust tooling and validation.

The role emphasizes collaboration across model development, performance, and runtime teams, with a focus on end-to-end efficiency, reliability, and release readiness in production

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior ML Inference Engineer – High-Performance Serving
Senior ML Inference Engineer – High-Performance Serving

Amazon Inc. • Seattle (WA)

On-site
USD 168,000 - 227,000
Senior ML Inference Engineer — Open-Source LLM Serving
Senior ML Inference Engineer — Open-Source LLM Serving

Amazon • Seattle (WA), Northern (KY)

Hybrid
USD 210,000 - 320,000
Open-Source ML Inference Engineer for Neuron
Open-Source ML Inference Engineer for Neuron

Amazon Inc. • Seattle (WA)

On-site
USD 144,000 - 194,000
RSUs
Health insurance
401(k) matching
+2
Senior ML Inference Engineer — High-Performance LLM Serving
Senior ML Inference Engineer — High-Performance LLM Serving

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 168,000 - 227,000
Health insurance
401(k) matching
Paid time off
+2
AI/ML Systems Engineer - High-Performance Inference
AI/ML Systems Engineer - High-Performance Inference

Amazon Inc. • Cupertino (CA)

On-site
USD 165,000 - 224,000
Health insurance
401(k) matching
Paid time off
+1
ML Inference Engineer - AWS Neuron & GenAI
ML Inference Engineer - AWS Neuron & GenAI

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 165,000 - 224,000
Sr. Software Development Engineer, Inference Team - AWS Neuron
Sr. Software Development Engineer, Inference Team - AWS Neuron

Amazon • Seattle (WA), Northern (KY)

On-site
USD 210,000 - 320,000
Senior ML Systems Engineer - AI Inference on AWS Neuron
Senior ML Systems Engineer - AI Inference on AWS Neuron

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 193,000 - 262,000
Senior SDE: Neuron Inference for LLM Acceleration
Senior SDE: Neuron Inference for LLM Acceleration

Amazon Inc. • Seattle (WA)

On-site
USD 144,000 - 194,000
Health insurance
401(k) matching
RSUs
Software Engineer-AI/ML, Inference Team - AWS Neuron
Software Engineer-AI/ML, Inference Team - AWS Neuron

Annapurna Labs (U.S.) Inc. • Seattle (WA)

On-site
USD 144,000 - 194,000
RSUs
401(k) matching
Paid time off