Sr. ML Kernel Performance Engineer, AWS Neuron, Annapurna Labs

Amazon

Cupertino (CA)

On-site

USD 193,300 - 261,500

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Amazon is seeking a Sr. ML Kernel Performance Engineer for AWS Neuron at Annapurna Labs to design and optimize high‑performance kernels for ML operations on Neuron accelerators. You will collaborate with hardware, compiler, runtime, and framework teams to maximize throughput and efficiency.

Responsibilities include profiling, identifying bottlenecks, implementing fusion, tiling, and scheduling, and enabling customers to deploy optimized models on AWS accelerators.

Qualifications

  • 5+ years of non‑internship professional software development experience.
  • 5+ years of programming in at least one software programming language.
  • 5+ years of leading design or architecture of new and existing systems (design patterns, reliability, scaling).
  • 5+ years of full software development life cycle experience, including coding standards, code reviews, source control, build processes, testing, and operations.
  • Experience as a mentor, tech lead, or leading an engineering team.
  • Experience with GPU kernel optimization and GPGPU computing.

Responsibilities

  • Design and implement high‑performance compute kernels for ML operations on Neuron hardware
  • Analyze and optimize kernel‑level performance across multiple generations of Neuron accelerators
  • Conduct detailed performance analysis using profiling tools to identify and resolve bottlenecks
  • Implement compiler optimizations such as fusion, sharding, tiling, and scheduling
  • Work directly with customers to enable and optimize their ML models on AWS accelerators
  • Collaborate across teams to develop innovative kernel optimization techniques

Skills

5+ years software development
5+ years programming
5+ years design/architecture
5+ years SDLC experience
Mentor / tech lead
Parallel programming

Education

Bachelor’s degree in computer science or equivalent

Tools

CUDA
NVIDIA PTX
LLVM/MLIR
GPGPU backends
OpenCL
Triton

Job description

Sr. ML Kernel Performance Engineer, AWS Neuron, Annapurna Labs

We are building AWS Neuron, an SDK that accelerates deep learning and GenAI workloads on Amazon’s custom machine learning accelerators, Inferentia and Trainium. As part of the Acceleration Kernel Library team, you will craft high‑performance kernels for ML functions, optimizing performance across hardware, compiler, runtime, and framework layers.

Key Responsibilities
  • Design and implement high‑performance compute kernels for ML operations on Neuron hardware
  • Analyze and optimize kernel‑level performance across multiple generations of Neuron accelerators
  • Conduct detailed performance analysis using profiling tools to identify and resolve bottlenecks
  • Implement compiler optimizations such as fusion, sharding, tiling, and scheduling
  • Work directly with customers to enable and optimize their ML models on AWS accelerators
  • Collaborate across teams to develop innovative kernel optimization techniques
Basic Qualifications
  • 5+ years of non‑internship professional software development experience
  • 5+ years of programming in at least one software programming language
  • 5+ years of leading design or architecture of new and existing systems (design patterns, reliability, scaling)
  • 5+ years of full software development life cycle experience, including coding standards, code reviews, source control, build processes, testing, and operations
  • Experience as a mentor, tech lead, or leading an engineering team
Preferred Qualifications
  • Bachelor’s degree in computer science or equivalent
  • 6+ years of full software development experience
  • Expertise in accelerator architectures for ML or HPC (GPUs, CPUs, FPGAs, or custom)
  • Experience with GPU kernel optimization and GPGPU computing (CUDA, NKI, Triton, OpenCL, SYCL, or ROCm)
  • Demonstrated experience with NVIDIA PTX and/or AMD GPU ISA
  • Experience developing high‑performance libraries for HPC applications
  • Proficiency in low‑level performance optimization for GPUs
  • Experience with LLVM/MLIR backend development for GPUs
  • Knowledge of ML frameworks (PyTorch, TensorFlow) and their GPU backends
  • Experience with parallel programming and optimization techniques
  • Understanding of GPU memory hierarchies and optimization strategies

Amazon is an equal‑opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status.

Base salary range: $193,300.00 – $261,500.00 annually (location‑specific).

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Engineering Manager, ML Kernel Performance, AWS Neuron, Annapurna Labs
Software Engineering Manager, ML Kernel Performance, AWS Neuron, Annapurna Labs

Amazon • Cupertino (CA)

On-site
USD 212,700 - 287,700
Health insurance
401(k) matching
Parental leave
+2
Sr. ML Kernel Performance Engineer, AWS Neuron, Annapurna Labs
Sr. ML Kernel Performance Engineer, AWS Neuron, Annapurna Labs

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 193,000 - 262,000
ML Kernel Performance Engineer, AWS Neuron, Annapurna Labs
ML Kernel Performance Engineer, AWS Neuron, Annapurna Labs

Amazon • Cupertino (CA)

On-site
USD 140,000 - 210,000
ML Kernel Performance Engineer, AWS Neuron, Annapurna Labs
ML Kernel Performance Engineer, AWS Neuron, Annapurna Labs

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 165,000 - 224,000
Health insurance
RSUs
401(k) matching
+1
Senior ML Kernel Performance Architect
Senior ML Kernel Performance Architect

Amazon • Cupertino (CA)

On-site
USD 193,000 - 262,000
ML Kernel Performance Engineer for Neuron Accelerators
ML Kernel Performance Engineer for Neuron Accelerators

Amazon • Cupertino (CA)

On-site
USD 140,000 - 210,000
Software Engineering Manager, ML Kernel Performance, AWS Neuron, Annapurna Labs
Software Engineering Manager, ML Kernel Performance, AWS Neuron, Annapurna Labs

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 212,000 - 288,000
Software Development Engineer, AI/ML, AWS Neuron, Model Inference
Software Development Engineer, AI/ML, AWS Neuron, Model Inference

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 165,000 - 224,000
Senior Software Development Engineer, AI/ML, AWS Neuron, Model Inference
Senior Software Development Engineer, AI/ML, AWS Neuron, Model Inference

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 193,000 - 262,000
Health insurance
401(k) matching
Paid time off
+1
ML Kernel Performance Engineer for AI Accelerators
ML Kernel Performance Engineer for AI Accelerators

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 165,000 - 224,000
Health insurance
RSUs
401(k) matching
+1