Senior ML Inference Systems Engineer

Annapurna Labs (U.S.) Inc.

Seattle (WA)

On-site

USD 180,000 - 260,000

Full time

6 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Amazon AWS Neuron is the complete software stack for Inferentia and Trainium, delivering high-performance model inference for customer workloads on AWS. This senior software engineering role focuses on building and optimizing large-scale inference solutions within the ML Inference Applications team.

The engineer will lead core serving tech within open-source frameworks, drive performance across frameworks like vLLM and SGLang, and collaborate with model development, compiler, runtime, and

Qualifications

  • 5+ years of non-internship professional software development experience.
  • 5+ years of programming with at least one software programming language experience.
  • 5+ years of leading design or architecture of new and existing systems.
  • 5+ years of full software development life cycle, including coding standards, code reviews, source control management, build processes, testing, and operations experience.
  • Experience as a mentor, tech lead or leading an engineering team.
  • Fundamentals of Machine learning models, their architecture, training and inference lifecycles along with work experience on some optimizations for improving the model performance.

Responsibilities

  • Lead the customization and optimization of open-source inference frameworks (vLLM, SGLang) on AWS Neuron— including core framework logic such as scheduling and model execution— applying state-of-the-art kernel development, parallel computation, distributed KV cache, speculative decoding, and systems engineering to deliver best-in-class LLM serving performance.
  • Influence the team's technical roadmap by evaluating emerging inference research and translating it into production-ready features, and raise the bar through design leadership and mentorship.
  • Collaborate across model development, compiler, runtime, and performance engineering teams to ensure end-to-end model performance and production-ready accuracy, scalability, and efficiency.

Skills

5+ years professional software
Mentor / tech lead
ML fundamentals
CS fundamentals
Full software lifecycle

Education

Bachelor's degree in CS or equivalent

Tools

PyTorch
JAX
CUDA
Triton

Job description

Amazon AWS Neuron is the complete software stack for Inferentia and Trainium, delivering high-performance model inference for customer workloads on AWS. This senior software engineering role focuses on building and optimizing large-scale inference solutions within the ML Inference Applications team.

The engineer will lead core serving tech within open-source frameworks, drive performance across frameworks like vLLM and SGLang, and collaborate with model development, compiler, runtime, and

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior ML Inference Engineer — Open-Source LLM Serving
Senior ML Inference Engineer — Open-Source LLM Serving

Amazon • Seattle (WA), Northern (KY)

Hybrid
USD 210,000 - 320,000
Senior ML Inference Engineer — LLMs on Neuron
Senior ML Inference Engineer — LLMs on Neuron

Amazon.com Services LLC • Seattle (WA)

On-site
USD 168,000 - 227,000
Health insurance
401(k) matching
Paid time off
+1
Senior ML Systems Engineer - AI Inference on AWS Neuron
Senior ML Systems Engineer - AI Inference on AWS Neuron

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 193,000 - 262,000
Sr. Software Development Engineer, Inference Team - AWS Neuron
Sr. Software Development Engineer, Inference Team - AWS Neuron

Amazon • Seattle (WA), Northern (KY)

Hybrid
USD 210,000 - 320,000
AI/ML Inference Engineer for AWS Neuron
AI/ML Inference Engineer for AWS Neuron

Amazon • Cupertino (CA)

On-site
USD 165,000 - 224,000
Health insurance
401(k) matching
Paid time off
+2
Senior AI/ML Inference Engineer for Neuron SDK
Senior AI/ML Inference Engineer for Neuron SDK

Amazon • Cupertino (CA)

On-site
USD 193,000 - 262,000
Health insurance
RSU program
ML Inference Engineer - AWS Neuron & GenAI
ML Inference Engineer - AWS Neuron & GenAI

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 165,000 - 224,000
Sr. Software Development Engineer, Inference Team - AWS Neuron
Sr. Software Development Engineer, Inference Team - AWS Neuron

Annapurna Labs (U.S.) Inc. • Seattle (WA)

On-site
USD 180,000 - 260,000
AI/ML Systems Engineer - Neuron Inference
AI/ML Systems Engineer - Neuron Inference

Amazon • Seattle (WA)

On-site
USD 144,000 - 194,000
Senior ML Compiler Engineer — Flexible, High-Impact AI
Senior ML Compiler Engineer — Flexible, High-Impact AI

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 180,000 - 250,000