AI/ML Inference Engineer for AWS Neuron

Amazon Web Services (AWS)

Cupertino (CA)

On-site

USD 165,000 - 224,000

Full time

3 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Annapurna Labs (U.S.) Inc. in Cupertino, California, is seeking an experienced software engineer to design and optimize ML models and frameworks for deployment on custom hardware accelerators.

You will work across ML system lifecycle, implement high-performance kernels, optimize system-wide bottlenecks, and collaborate with customers to enable model acceleration on Neuron and Trainium/Inferentia. Strong fundamentals in C++, Python, ML, and distributed systems are required.

Qualifications

  • Bachelor's degree in computer science or equivalent.
  • 3+ years of professional software development experience.
  • Experience in design or architecture of scalable systems.
  • Fundamentals of ML, LLMs, and inference lifecycles.

Responsibilities

  • Design, develop, and optimize ML models and frameworks for deployment on custom ML hardware accelerators.
  • Participate in the ML system development lifecycle including distributed computing based architecture design, implementation, performance profiling, hardware-specific optimizations, testing and production deployment.
  • Build infrastructure to analyze and onboard multiple models with diverse architecture.
  • Design and implement high-performance kernels and features for ML operations, leveraging the Neuron architecture and programming models.
  • Analyze and optimize system-level performance across multiple generations of Neuron hardware.
  • Conduct detailed performance analysis using profiling tools to identify and resolve bottlenecks.
  • Implement optimizations such as fusion, sharding, tiling, and scheduling.
  • Conduct comprehensive testing, including unit and end-to-end model testing with pipelines.

Skills

Fundamentals of ML & LLMs
Performance optimization
Profiling & debugging
Parallel computing
Large-scale distributed systems

Education

Bachelor's degree in computer science or equivalent

Tools

C++
Python

Job description

Annapurna Labs (U.S.) Inc. in Cupertino, California, is seeking an experienced software engineer to design and optimize ML models and frameworks for deployment on custom hardware accelerators.

You will work across ML system lifecycle, implement high-performance kernels, optimize system-wide bottlenecks, and collaborate with customers to enable model acceleration on Neuron and Trainium/Inferentia. Strong fundamentals in C++, Python, ML, and distributed systems are required.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI/ML Inference Engineer for Neuron on AWS
Senior AI/ML Inference Engineer for Neuron on AWS

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 193,000 - 262,000
Health insurance
401(k) matching
Paid time off
+1
Applied Scientist II — ML Systems for AI Accelerators
Applied Scientist II — ML Systems for AI Accelerators

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 171,000 - 223,000
AI/ML SDE — High-Performance Inference on Custom Hardware
AI/ML SDE — High-Performance Inference on Custom Hardware

Amazon • Cupertino (CA)

On-site
USD 180,000 - 240,000
Senior ML Kernel Optimization Engineer
Senior ML Kernel Optimization Engineer

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 193,000 - 262,000
Senior AI/ML Systems Engineer - Neuron Inference
Senior AI/ML Systems Engineer - Neuron Inference

Amazon • Cupertino (CA)

On-site
USD 193,300 - 261,500
Senior ML Software Engineer, AI Accelerator & Inference
Senior ML Software Engineer, AI Accelerator & Inference

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 168,000 - 228,000
ML Kernel Performance Engineer for AI Accelerators
ML Kernel Performance Engineer for AI Accelerators

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 165,000 - 224,000
Health insurance
RSUs
401(k) matching
+1
Senior ML Kernel Performance Architect
Senior ML Kernel Performance Architect

Amazon • Cupertino (CA)

On-site
USD 193,000 - 262,000
Senior ML Accelerator Solutions Architect
Senior ML Accelerator Solutions Architect

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 176,000 - 239,000
Engineering Manager, ML Kernel Performance
Engineering Manager, ML Kernel Performance

Amazon • Cupertino (CA)

On-site
USD 212,700 - 287,700
Health insurance
401(k) matching
Parental leave
+2