ML Inference Engineer - AWS Neuron & GenAI

Amazon Web Services (AWS)

Cupertino (CA)

On-site

USD 165,000 - 224,000

Full time

3 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Annapurna Labs (U.S.) Inc. in Cupertino, CA, seeks engineers to design and optimize ML models for AWS Neuron on Inferentia and Trainium accelerators, bridging software and hardware for high-performance inference and training.

You will work across frameworks, kernels and compilers, implement high-performance ML kernels, profile performance, and collaborate with customers to enable their models on AWS accelerators, in a fast-paced startup‑like culture.

Qualifications

  • Bachelor's degree in computer science or equivalent.
  • 3+ years of professional software development experience.
  • 3+ years of design or architecture (design patterns, reliability and scaling) of new and existing systems experience.
  • Fundamentals of Machine learning and LLMs, their architecture, training and inference lifecycles along with work experience on some optimizations for improving the model execution.
  • Software development experience in C++, Python (experience in at least one language is required).
  • Strong understanding of system performance, memory management, and parallel computing principles.
  • Proficiency in debugging, profiling, and implementing best software engineering practices in large-scale systems.

Responsibilities

  • Design, develop, and optimize machine learning models and frameworks for deployment on custom ML hardware accelerators.
  • Participate in all stages of the ML system development lifecycle including distributed computing based architecture design, implementation, performance profiling, hardware-specific optimizations, testing and production deployment.
  • Build infrastructure to systematically analyze and onboard multiple models with diverse architecture.
  • Design and implement high-performance kernels and features for ML operations, leveraging the Neuron architecture and programming models
  • Analyze and optimize system-level performance across multiple generations of Neuron hardware
  • Conduct detailed performance analysis using profiling tools to identify and resolve bottlenecks
  • Implement optimizations such as fusion, sharding, tiling, and scheduling
  • Conduct comprehensive testing, including unit and end-to-end model testing with continuous deployment and releases through pipelines.
  • Work directly with customers to enable and optimize their ML models on AWS accelerators
  • Collaborate across teams to develop innovative optimization techniques

Skills

C++
Python
Performance profiling
Debugging
Parallel computing

Education

Bachelor's degree in computer science or equivalent

Tools

PyTorch
JIT compilation
TensorRT
CUDA kernels

Job description

Annapurna Labs (U.S.) Inc. in Cupertino, CA, seeks engineers to design and optimize ML models for AWS Neuron on Inferentia and Trainium accelerators, bridging software and hardware for high-performance inference and training.

You will work across frameworks, kernels and compilers, implement high-performance ML kernels, profile performance, and collaborate with customers to enable their models on AWS accelerators, in a fast-paced startup‑like culture.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior ML Systems Engineer - AI Inference on AWS Neuron
Senior ML Systems Engineer - AI Inference on AWS Neuron

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 193,000 - 262,000
ML Kernel Performance Engineer for Neuron
ML Kernel Performance Engineer for Neuron

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 165,000 - 224,000
Engineering Manager, ML Kernel Performance
Engineering Manager, ML Kernel Performance

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 213,000 - 288,000
Health insurance
RSUs
401(k) matching
+2
ML Kernel Performance Engineer for Neuron Accelerators
ML Kernel Performance Engineer for Neuron Accelerators

Amazon • Cupertino (CA)

On-site
USD 140,000 - 210,000
Senior ML Compiler Engineer – Neuron
Senior ML Compiler Engineer – Neuron

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 168,000 - 227,000
Engineering Manager, ML Kernel Performance
Engineering Manager, ML Kernel Performance

Amazon • Cupertino (CA)

On-site
USD 212,700 - 287,700
Health insurance
401(k) matching
Parental leave
+2
GenAI ML Systems Engineer
GenAI ML Systems Engineer

Amazon Web Services (AWS) • New York (NY)

On-site
USD 158,000 - 214,000
Senior Software Engineer — GenAI & ML Acceleration
Senior Software Engineer — GenAI & ML Acceleration

Amazon Web Services (AWS) • New York (NY)

On-site
USD 185,000 - 250,000
Health insurance
401(k) matching
Paid time off
+2
SDE I, ML Infra & AI Accelerators
SDE I, ML Infra & AI Accelerators

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 127,000 - 185,000
Senior Software Development Engineer, AI/ML, AWS Neuron, Model Inference
Senior Software Development Engineer, AI/ML, AWS Neuron, Model Inference

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 193,000 - 262,000