Senior ML Inference Engineer — High-Performance LLM Serving

Amazon Web Services (AWS)

Seattle (WA)

On-site

USD 168,000 - 227,000

Full time

5 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Health insurance
401(k) matching
Paid time off
Parental leave
RSUs

Job summary

Annapurna Labs (U.S.) Inc. in Seattle seeks a senior software engineer to advance AWS Neuron, the complete software stack for AWS Inferentia and Trainium.

You will deliver high-performance model inference for customer workloads on Inferentia- and Trainium-powered instances and lead core serving tech within open-source frameworks such as vLLM and SGLang. You will collaborate with model development, compiler, runtime, and performance teams to ensure production-ready accuracy, scalability, and

Qualifications

  • 5+ years of non-internship professional software development experience.
  • 5+ years of programming with at least one software programming language experience.
  • 5+ years of leading design or architecture (design patterns, reliability and scaling) of new and existing systems experience.
  • 5+ years of full software development life cycle, including coding standards, code reviews, source control management, build processes, testing, and operations experience.
  • Experience as a mentor, tech lead or leading an engineering team.
  • Fundamentals of Machine learning models, their architecture, training and inference lifecycles along with work experience on some optimizations for improving the model performance.

Responsibilities

  • Lead the customization and optimization of open-source inference frameworks (vLLM, SGLang) on AWS Neuron, including core framework logic like scheduling and model execution.
  • Apply state-of-the-art kernel development, parallel computation, distributed KV cache, speculative decoding, and systems engineering to deliver best-in-class LLM serving performance.
  • Influence the team's technical roadmap by evaluating emerging inference research and translating it into production-ready features.
  • Collaborate closely with model development, compiler, runtime, and performance engineering teams to ensure end-to-end model performance and production readiness.

Skills

5+ years dev exp
5+ years programming
5+ years design/architecture
5+ years SDLC
Mentor / tech lead
ML model lifecycle

Education

Bachelor's degree in Computer Science

Tools

PyTorch
JAX
CUDA
Triton
AWS Neuron

Job description

Annapurna Labs (U.S.) Inc. in Seattle seeks a senior software engineer to advance AWS Neuron, the complete software stack for AWS Inferentia and Trainium.

You will deliver high-performance model inference for customer workloads on Inferentia- and Trainium-powered instances and lead core serving tech within open-source frameworks such as vLLM and SGLang. You will collaborate with model development, compiler, runtime, and performance teams to ensure production-ready accuracy, scalability, and

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior ML Inference Engineer — Open-Source LLM Serving
Senior ML Inference Engineer — Open-Source LLM Serving

Amazon • Seattle (WA), Northern (KY)

Hybrid
USD 210,000 - 320,000
Senior ML Systems Engineer - AI Inference on AWS Neuron
Senior ML Systems Engineer - AI Inference on AWS Neuron

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 193,000 - 262,000
ML Inference Engineer - AWS Neuron & GenAI
ML Inference Engineer - AWS Neuron & GenAI

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 165,000 - 224,000
Senior AI/ML Software Engineer - High-Perf Inference
Senior AI/ML Software Engineer - High-Perf Inference

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 168,000 - 227,000
Senior AI/ML Inference Engineer (Neuron)
Senior AI/ML Inference Engineer (Neuron)

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 193,000 - 262,000
Health insurance
401(k) matching
Paid time off
Senior AI/ML Software Engineer: Inference on Trainium
Senior AI/ML Software Engineer: Inference on Trainium

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 193,000 - 262,000
RSUs
Health insurance
401(k) matching
+1
Sr. Software Development Engineer, Inference Team - AWS Neuron
Sr. Software Development Engineer, Inference Team - AWS Neuron

Amazon • Seattle (WA), Northern (KY)

Hybrid
USD 210,000 - 320,000
Senior ML Compiler Engineer – Neuron
Senior ML Compiler Engineer – Neuron

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 168,000 - 227,000
AI/ML Systems Engineer for AWS Neuron Inference
AI/ML Systems Engineer for AWS Neuron Inference

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 144,000 - 194,000
Health insurance
401(k) matching
Paid time off
+1
Sr. Software Development Engineer, Inference Team - AWS Neuron
Sr. Software Development Engineer, Inference Team - AWS Neuron

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 168,000 - 227,000
Health insurance
401(k) matching
Paid time off
+2