Senior AI/ML Inference Engineer for Neuron on AWS

Amazon Web Services (AWS)

Cupertino (CA)

On-site

USD 193,300 - 261,500

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health insurance
401(k) matching
Paid time off
Parental leave

Job summary

Annapura Labs (U.S.) Inc. in Cupertino, CA, is seeking a senior software engineer to advance ML model acceleration on AWS Neuron with Trainium and Inferentia. You will lead design and optimization of distributed inference, collaborating with hardware, runtime, and framework teams to push performance.

You’ll build high-performance kernels, profile systems, and contribute to open-source integration while mentoring engineers and shaping the future of AI acceleration at AWS.

Qualifications

  • Bachelor's degree in computer science or equivalent.
  • 5+ years of professional software development experience.
  • Experience with C++ and Python.
  • Fundamentals of ML/LLMs and model inference.
  • Strong understanding of system performance and memory management.

Responsibilities

  • Design, develop, and optimize ML models and frameworks for deployment on custom ML hardware accelerators.
  • Participate in ML system lifecycle including distributed architecture design, performance profiling, hardware-specific optimizations, testing and production deployment.
  • Build infrastructure to analyze and onboard multiple models.
  • Design and implement high-performance kernels and features for ML operations leveraging the Neuron architecture.
  • Analyze and optimize system-level performance across generations of Neuron hardware.
  • Conduct performance analysis using profiling tools to identify bottlenecks.
  • Implement optimizations such as fusion, sharding, tiling, and scheduling.
  • Conduct comprehensive testing including unit and end-to-end model testing with pipelines.
  • Work directly with customers to enable and optimize ML models on AWS accelerators.
  • Collaborate across teams to develop innovative optimization techniques.

Skills

C++
Python
ML fundamentals
Debugging
Profiling
Performance optimization
Distributed systems

Education

Bachelor's degree in computer science or equivalent

Tools

PyTorch
JIT compilation
CUDA kernels
Triton

Job description

Annapura Labs (U.S.) Inc. in Cupertino, CA, is seeking a senior software engineer to advance ML model acceleration on AWS Neuron with Trainium and Inferentia. You will lead design and optimization of distributed inference, collaborating with hardware, runtime, and framework teams to push performance.

You’ll build high-performance kernels, profile systems, and contribute to open-source integration while mentoring engineers and shaping the future of AI acceleration at AWS.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI/ML Inference Engineer for AWS Neuron
AI/ML Inference Engineer for AWS Neuron

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 165,000 - 224,000
Senior ML Software Engineer, AI Accelerator & Inference
Senior ML Software Engineer, AI Accelerator & Inference

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 168,000 - 228,000
Senior AI/ML Systems Engineer - Neuron Inference
Senior AI/ML Systems Engineer - Neuron Inference

Amazon • Cupertino (CA)

On-site
USD 193,300 - 261,500
Senior SDE — AI/ML Inference on Trainium/Neuron
Senior SDE — AI/ML Inference on Trainium/Neuron

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 193,000 - 262,000
RSUs
Comprehensive benefits
401(k)
Applied Scientist II — ML Systems for AI Accelerators
Applied Scientist II — ML Systems for AI Accelerators

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 171,000 - 223,000
AI/ML SDE — High-Performance Inference on Custom Hardware
AI/ML SDE — High-Performance Inference on Custom Hardware

Amazon • Cupertino (CA)

On-site
USD 180,000 - 240,000
Applied Scientist, AWS Neuron Science team
Applied Scientist, AWS Neuron Science team

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 133,000 - 222,000
Senior AI/ML Software Engineer — Distributed Training
Senior AI/ML Software Engineer — Distributed Training

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 168,000 - 228,000
RSUs
Health benefits
Sign-on payments
Software Development Engineer, AI/ML, AWS Neuron, Model Inference
Software Development Engineer, AI/ML, AWS Neuron, Model Inference

Amazon • Cupertino (CA)

On-site
USD 180,000 - 240,000
Senior ML Accelerator Solutions Architect
Senior ML Accelerator Solutions Architect

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 176,000 - 239,000