Senior AI/ML Systems Engineer, Neuron Inference

Amazon Inc.

Cupertino (CA)

On-site

USD 193,000 - 262,000

Full time

5 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Health insurance
401(k) matching
Paid time off
Parental leave

Job summary

Amazon’s Annapurna Labs team is seeking a Sr. Software Development Engineer for AI/ML on AWS Neuron to accelerate GenAI workloads on Inferentia/Trainium. You will architect features, mentor engineers, and optimize ML models and kernels for high-performance inference across clusters.

You will collaborate with cross-functional teams to push the boundaries of AI acceleration, participate in performance tuning, and work closely with customers to enable and optimize models on AWS accelerators.

Qualifications

  • Bachelor's degree in computer science or equivalent.
  • 5+ years of professional software development experience.
  • 5+ years of programming in at least one language.
  • 5+ years of leading design or architecture (design patterns, reliability and scaling) of new and existing systems experience.
  • Experience as a mentor, tech lead or leading an engineering team.
  • Fundamentals of Machine learning and LLMs, their architecture, training and inference lifecycles along with work experience on some optimizations for improving the model execution.
  • Software development experience in C++, Python (experience in at least one language is required).
  • Strong understanding of system performance, memory management, and parallel computing principles.
  • Deep understanding of computer architecture, operation systems level software and working knowledge of parallel computing.

Responsibilities

  • Design, develop, and optimize machine learning models and frameworks for deployment on custom ML hardware accelerators.
  • Participate in all stages of the ML system development lifecycle including distributed computing based architecture design, implementation, performance profiling, hardware-specific optimizations, testing and production deployment.
  • Build infrastructure to systematically analyze and onboard multiple models with diverse architecture.
  • Design and implement high-performance kernels and features for ML operations, leveraging the Neuron architecture and programming models
  • Analyze and optimize system-level performance across multiple generations of Neuron hardware
  • Conduct detailed performance analysis using profiling tools to identify and resolve bottlenecks
  • Implement optimizations such as fusion, sharding, tiling, and scheduling
  • Conduct comprehensive testing, including unit and end-to-end model testing with continuous deployment and releases through pipelines.
  • Work directly with customers to enable and optimize their ML models on AWS accelerators
  • Collaborate across teams to develop innovative optimization techniques

Skills

C++
Python
Debugging
Profiling
ML fundamentals
LLMs
Mentorship
Parallel computing

Education

Bachelor's degree in CS
Master's degree in CS

Tools

CUDA kernels
PyTorch
JIT compilation
TensorRT

Job description

Amazon’s Annapurna Labs team is seeking a Sr. Software Development Engineer for AI/ML on AWS Neuron to accelerate GenAI workloads on Inferentia/Trainium. You will architect features, mentor engineers, and optimize ML models and kernels for high-performance inference across clusters.

You will collaborate with cross-functional teams to push the boundaries of AI acceleration, participate in performance tuning, and work closely with customers to enable and optimize models on AWS accelerators.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior AI/ML Inference Engineer (Neuron)
Senior AI/ML Inference Engineer (Neuron)

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 193,000 - 262,000
Health insurance
401(k) matching
Paid time off
AI/ML Systems Engineer - High-Performance Inference
AI/ML Systems Engineer - High-Performance Inference

Amazon Inc. • Cupertino (CA)

On-site
USD 165,000 - 224,000
Health insurance
401(k) matching
Paid time off
+1
Senior ML Systems Engineer - AI Inference on AWS Neuron
Senior ML Systems Engineer - AI Inference on AWS Neuron

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 193,000 - 262,000
Senior AI/ML Software Engineer - High-Perf Inference
Senior AI/ML Software Engineer - High-Perf Inference

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 168,000 - 227,000
AI/ML Systems Engineer for AWS Neuron Inference
AI/ML Systems Engineer for AWS Neuron Inference

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 144,000 - 194,000
Health insurance
401(k) matching
Paid time off
+1
Senior AI/ML Software Engineer: Inference on Trainium
Senior AI/ML Software Engineer: Inference on Trainium

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 193,000 - 262,000
RSUs
Health insurance
401(k) matching
+1
ML Inference Engineer - AWS Neuron & GenAI
ML Inference Engineer - AWS Neuron & GenAI

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 165,000 - 224,000
Senior ML Kernel Performance Engineer - AI Accelerator
Senior ML Kernel Performance Engineer - AI Accelerator

Amazon • Cupertino (CA)

On-site
USD 193,000 - 262,000
AI/ML Inference Engineer (Trainium)
AI/ML Inference Engineer (Trainium)

Amazon • Cupertino (CA)

On-site
USD 165,000 - 224,000
Health insurance
401(k) matching
Paid time off
+1
Senior Software Engineer — GenAI & ML Acceleration
Senior Software Engineer — GenAI & ML Acceleration

Amazon Web Services (AWS) • New York (NY)

On-site
USD 185,000 - 250,000
Health insurance
401(k) matching
Paid time off
+2