Senior AI/ML Inference Engineer (Neuron)

Amazon Web Services (AWS)

Cupertino (CA)

On-site

USD 193,000 - 262,000

Full time

2 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Health insurance
401(k) matching
Paid time off

Job summary

Annapurna Labs (U.S.) Inc. is seeking an experienced software engineer to design, develop, and optimize ML models for AWS Neuron on Inferentia and Trainium accelerators. You will work across frameworks and kernels, collaborating with compiler, runtime, and customer teams to maximize performance.

You will lead high-performance kernel development, analyze system-wide bottlenecks, and mentor engineers in a startup-like environment with a strong focus on performance and scale.

Qualifications

  • Bachelor's degree in computer science or equivalent.
  • 5+ years of non-internship professional software development experience.
  • 5+ years of programming with at least one software language (C++ or Python).
  • 5+ years of leading design or architecture of new and existing systems.
  • Experience mentoring, tech lead, or leading an engineering team.

Responsibilities

  • Design, develop, and optimize ML models and frameworks for deployment on custom ML hardware accelerators.
  • Participate in all stages of ML system development lifecycle including distributed computing design, performance profiling, hardware optimizations, testing and deployment.
  • Build infrastructure to analyze and onboard multiple models with diverse architecture.
  • Design and implement high-performance kernels and features for ML operations leveraging Neuron architecture.
  • Analyze and optimize system-level performance across Neuron hardware generations.
  • Conduct detailed performance analysis using profiling tools to identify bottlenecks.
  • Implement optimizations such as fusion, sharding, tiling, and scheduling.
  • Perform comprehensive testing including unit and end-to-end model testing with CI/CD pipelines.
  • Work with customers to enable and optimize their ML models on AWS accelerators.
  • Collaborate across teams to develop innovative optimization techniques.

Skills

C++
Python
ML concepts
Profiling & debugging
Parallel computing

Education

Bachelor's degree in computer science or equivalent
Master's degree in CS (preferred)

Tools

PyTorch
JIT compilation
AOT tracing
CUDA kernels
TensorRT

Job description

Annapurna Labs (U.S.) Inc. is seeking an experienced software engineer to design, develop, and optimize ML models for AWS Neuron on Inferentia and Trainium accelerators. You will work across frameworks and kernels, collaborating with compiler, runtime, and customer teams to maximize performance.

You will lead high-performance kernel development, analyze system-wide bottlenecks, and mentor engineers in a startup-like environment with a strong focus on performance and scale.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior ML Systems Engineer - AI Inference on AWS Neuron
Senior ML Systems Engineer - AI Inference on AWS Neuron

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 193,000 - 262,000
ML Inference Engineer - AWS Neuron & GenAI
ML Inference Engineer - AWS Neuron & GenAI

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 165,000 - 224,000
Senior AI/ML Software Engineer - High-Perf Inference
Senior AI/ML Software Engineer - High-Perf Inference

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 168,000 - 227,000
AI/ML Systems Engineer - Neuron Inference
AI/ML Systems Engineer - Neuron Inference

Amazon • Seattle (WA)

On-site
USD 144,000 - 194,000
AI/ML Systems Engineer for AWS Neuron Inference
AI/ML Systems Engineer for AWS Neuron Inference

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 144,000 - 194,000
Health insurance
401(k) matching
Paid time off
+1
Senior AI/ML Software Engineer: Inference on Trainium
Senior AI/ML Software Engineer: Inference on Trainium

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 193,000 - 262,000
RSUs
Health insurance
401(k) matching
+1
Senior ML Kernel Performance Engineer - AI Accelerator
Senior ML Kernel Performance Engineer - AI Accelerator

Amazon • Cupertino (CA)

On-site
USD 193,000 - 262,000
Neuron Runtime Engineer — AI Inference & Profiling
Neuron Runtime Engineer — AI Inference & Profiling

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 144,000 - 194,000
Senior ML Compiler Engineer – Neuron
Senior ML Compiler Engineer – Neuron

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 168,000 - 227,000
Senior ML Kernel Optimizer for AWS Neuron
Senior ML Kernel Optimizer for AWS Neuron

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 193,000 - 262,000