Senior ML Inference Engineer – High-Performance Serving

Amazon Inc.

Seattle (WA)

On-site

USD 168,000 - 227,000

Full time

3 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Amazon Inc. in Seattle, WA is seeking a Sr. Software Development Engineer for the Inference Team on AWS Neuron to deliver high-performance model inference for customer workloads on Inferentia- and Trainium-powered instances.

You will lead customization of open-source inference frameworks (vLLM, SGLang) and drive kernel-level and framework optimizations, collaborating across model development, compiler, and runtime teams to ensure scalable, production-ready performance.

Qualifications

  • 5+ years of non-internship professional software development experience.
  • 5+ years of programming with at least one software programming language experience.
  • 5+ years of leading design or architecture (design patterns, reliability and scaling) of new and existing systems experience.
  • 5+ years of full software development life cycle, including coding standards, code reviews, source control management, build processes, testing, and operations experience.
  • Experience as a mentor, tech lead or leading an engineering team
  • Fundamentals of Machine learning models, their architecture, training and inference lifecycles along with work experience on some optimizations for improving the model performance.

Responsibilities

  • Lead customization and optimization of open-source inference frameworks (vLLM, SGLang) on AWS Neuron—including core framework logic such as scheduling and model execution—applying state-of-the-art techniques in kernel development, parallel computation, distributed KV cache, speculative decoding, and systems engineering to deliver best-in-class LLM serving performance.
  • Influence the team's technical roadmap by evaluating emerging inference research and translating it into production-ready features; provide design leadership and mentorship.
  • Collaborate with model development, compiler, runtime, and performance engineering teams to ensure end-to-end model performance and scalable deployments.

Skills

Software development
Programming experience
Design/architecture leadership
SDLC experience
Mentor / tech lead
ML inference basics

Education

Bachelor's degree in CS or equivalent

Tools

PyTorch/JAX
vLLM/SGLang
CUDA/Triton

Job description

Amazon Inc. in Seattle, WA is seeking a Sr. Software Development Engineer for the Inference Team on AWS Neuron to deliver high-performance model inference for customer workloads on Inferentia- and Trainium-powered instances.

You will lead customization of open-source inference frameworks (vLLM, SGLang) and drive kernel-level and framework optimizations, collaborating across model development, compiler, and runtime teams to ensure scalable, production-ready performance.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior ML Inference Engineer — Open-Source LLM Serving
Senior ML Inference Engineer — Open-Source LLM Serving

Amazon • Seattle (WA), Northern (KY)

Hybrid
USD 210,000 - 320,000
Senior ML Inference Engineer — High-Performance LLM Serving
Senior ML Inference Engineer — High-Performance LLM Serving

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 168,000 - 227,000
Health insurance
401(k) matching
Paid time off
+2
AI/ML Systems Engineer - High-Performance Inference
AI/ML Systems Engineer - High-Performance Inference

Amazon Inc. • Cupertino (CA)

On-site
USD 165,000 - 224,000
Health insurance
401(k) matching
Paid time off
+1
Sr. Software Development Engineer, Inference Team - AWS Neuron
Sr. Software Development Engineer, Inference Team - AWS Neuron

Amazon • Seattle (WA), Northern (KY)

On-site
USD 210,000 - 320,000
Senior AI/ML Systems Engineer, Neuron Inference
Senior AI/ML Systems Engineer, Neuron Inference

Amazon Inc. • Cupertino (CA)

On-site
USD 193,000 - 262,000
Health insurance
401(k) matching
Paid time off
+1
Senior ML Systems Engineer - AI Inference on AWS Neuron
Senior ML Systems Engineer - AI Inference on AWS Neuron

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 193,000 - 262,000
Senior AI/ML Software Engineer - High-Perf Inference
Senior AI/ML Software Engineer - High-Perf Inference

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 168,000 - 227,000
Senior ML Inference Systems Engineer - Custom Accelerator
Senior ML Inference Systems Engineer - Custom Accelerator

Amazon • Cupertino (CA)

On-site
USD 193,000 - 262,000
Health insurance
RSUs
401(k) matching
ML Inference Engineer - AWS Neuron & GenAI
ML Inference Engineer - AWS Neuron & GenAI

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 165,000 - 224,000
Senior AI/ML Software Engineer: Inference on Trainium
Senior AI/ML Software Engineer: Inference on Trainium

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 193,000 - 262,000
RSUs
Health insurance
401(k) matching
+1