Senior ML Systems Engineer - AI Inference on AWS Neuron

Amazon Web Services (AWS)

Cupertino (CA)

On-site

USD 193,000 - 262,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Annapurna Labs (U.S.) Inc. is seeking a seasoned software engineer to design, develop, and optimize ML models and frameworks for deployment on AWS Neuron hardware accelerators (Inferentia and Trainium).

You will collaborate across teams with PyTorch and JAX, implement high‑performance kernels, analyze system performance, and work directly with customers to enable scalable AI inference on Neuron.

Qualifications

  • Bachelor's degree in computer science or equivalent.
  • 5+ years of professional software development experience.
  • Experience in design/architecture of scalable systems.
  • Knowledge of ML concepts and LLMS; optimization experience.
  • Proficiency in C++ and Python.
  • Strong understanding of system performance and memory management.
  • Experience with debugging and profiling.

Responsibilities

  • Design, develop, and optimize ML models and frameworks for deployment on custom ML hardware accelerators.
  • Participate in all stages of the ML system development lifecycle including distributed computing based architecture design, implementation, performance profiling, hardware-specific optimizations, testing and production deployment.
  • Build infrastructure to systematically analyze and onboard multiple models with diverse architecture.
  • Design and implement high-performance kernels and features for ML operations, leveraging the Neuron architecture and programming models
  • Analyze and optimize system-level performance across multiple generations of Neuron hardware
  • Conduct detailed performance analysis using profiling tools to identify and resolve bottlenecks
  • Implement optimizations such as fusion, sharding, tiling, and scheduling
  • Conduct comprehensive testing, including unit and end-to-end model testing with continuous deployment and releases through pipelines.
  • Work directly with customers to enable and optimize their ML models on AWS accelerators
  • Collaborate across teams to develop innovative optimization techniques

Skills

C++
Python
Parallel computing
Debugging
Profiling
Performance tuning

Education

Bachelor's degree in computer science
Master's degree in computer science

Tools

PyTorch
JIT compilation
AOT tracing
CUDA kernels
TensorRT
CUTLASS
FlashInfer
vLLM
SGLang

Job description

Annapurna Labs (U.S.) Inc. is seeking a seasoned software engineer to design, develop, and optimize ML models and frameworks for deployment on AWS Neuron hardware accelerators (Inferentia and Trainium).

You will collaborate across teams with PyTorch and JAX, implement high‑performance kernels, analyze system performance, and work directly with customers to enable scalable AI inference on Neuron.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Inference Engineer - AWS Neuron & GenAI
ML Inference Engineer - AWS Neuron & GenAI

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 165,000 - 224,000
Senior ML Compiler Engineer – Neuron
Senior ML Compiler Engineer – Neuron

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 168,000 - 227,000
Senior AI/ML Distributed Training Engineer
Senior AI/ML Distributed Training Engineer

Annapurna Labs (U.S.) Inc. • Cupertino (CA)

On-site
USD 180,000 - 240,000
GenAI ML Systems Engineer
GenAI ML Systems Engineer

Amazon Web Services (AWS) • New York (NY)

On-site
USD 158,000 - 214,000
Applied Scientist, ML Systems for AWS Neuron
Applied Scientist, ML Systems for AWS Neuron

Amazon • Cupertino (CA)

On-site
USD 171,600 - 222,200
RSUs
Health insurance
401(k) matching
+1
Senior Software Engineer — GenAI & ML Acceleration
Senior Software Engineer — GenAI & ML Acceleration

Amazon Web Services (AWS) • New York (NY)

On-site
USD 185,000 - 250,000
Health insurance
401(k) matching
Paid time off
+2
Senior ML Accelerator Runtime Engineer
Senior ML Accelerator Runtime Engineer

Amazon • Seattle (WA)

On-site
USD 143,700 - 194,400
Health insurance
Dental
Vision
+4
Senior ML Compiler Engineer – Build Next‑Gen ML on AI Accelerators
Senior ML Compiler Engineer – Build Next‑Gen ML on AI Accelerators

Amazon • Cupertino (CA)

On-site
USD 193,000 - 262,000
Health benefits
401(k) matching
Parental leave
Senior Software Development Engineer, AI/ML, AWS Neuron, Model Inference
Senior Software Development Engineer, AI/ML, AWS Neuron, Model Inference

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 193,000 - 262,000
Sr. Software Engineer- AI/ML, AWS Neuron Apps
Sr. Software Engineer- AI/ML, AWS Neuron Apps

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 168,000 - 227,000