Senior ML Inference Engineer for Custom Hardware

Amazon Web Services (AWS)

Cupertino (CA)

On-site

USD 193,000 - 262,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Health insurance
401(k) matching
Paid time off
Parental leave

Job summary

Annapurna Labs (U.S.) Inc. is seeking a Senior Software Development Engineer to own the design and implementation of the inference data plane for large models.

You will build software for efficient execution on custom hardware, covering model execution, memory management, data movement, and serving integration. Responsibilities include developing compute kernels, validating end-to-end LLM architectures, and integrating backends into ML serving frameworks.

Qualifications

  • Bachelor’s degree in computer science or equivalent.
  • 7+ years of full software development life cycle experience.
  • Strong proficiency in C/C++.
  • Strong Linux systems knowledge.
  • Experience developing compute kernels for GPUs, DSPs, or custom accelerators.
  • Proven track record of owning and delivering complex software features end-to-end.
  • Knowledge of ML fundamentals, including transformer architectures.
  • Familiarity with ML frameworks including PyTorch, JAX, vLLM, SGLang, Dynamo, TorchXLA, and TensorRT.

Responsibilities

  • Develop and optimize compute kernels for a custom ML accelerator architecture, targeting production-level performance for large language model inference.
  • Implement and validate LLM architectures end-to-end—from PyTorch model definition through distributed execution on custom hardware.
  • Integrate custom accelerator backends into open-source ML serving frameworks (vLLM, PyTorch), including scheduler extensions, memory management, and model parallelism.
  • Build and maintain test infrastructure for model correctness validation across CPU, GPU, simulator, and hardware targets.
  • Profile and optimize inference workloads—identify bottlenecks, instrument critical paths, and drive latency and throughput improvements from simulation through hardware bringup.
  • Own features end-to-end: from design through implementation, testing, and integration into the broader software stack.
  • Contribute to CI/CD pipelines that gate model and kernel changes on correctness and performance regressions.
  • Mentor engineers, drive design reviews, and raise the engineering bar across the team.

Skills

C/C++
Linux
Compute kernels
Software development lifecycle
ML frameworks (PyTorch, JAX, vLLM)
ML fundamentals

Education

Bachelor's degree in Computer Science or equivalent

Tools

PyTorch
JAX
vLLM
Dynamo
TorchXLA
TensorRT

Job description

Annapurna Labs (U.S.) Inc. is seeking a Senior Software Development Engineer to own the design and implementation of the inference data plane for large models.

You will build software for efficient execution on custom hardware, covering model execution, memory management, data movement, and serving integration. Responsibilities include developing compute kernels, validating end-to-end LLM architectures, and integrating backends into ML serving frameworks.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

LLM Inference Systems Engineer — Custom Accelerator
LLM Inference Systems Engineer — Custom Accelerator

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 165,000 - 224,000
Health insurance
RSUs
401(k) matching
+2
Senior ML Inference Systems Engineer - Custom Accelerator
Senior ML Inference Systems Engineer - Custom Accelerator

Amazon • Cupertino (CA)

On-site
USD 193,000 - 262,000
Health insurance
RSUs
401(k) matching
Senior ML Systems Engineer - AI Inference on AWS Neuron
Senior ML Systems Engineer - AI Inference on AWS Neuron

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 193,000 - 262,000
Senior AI/ML Inference Engineer (Neuron)
Senior AI/ML Inference Engineer (Neuron)

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 193,000 - 262,000
Health insurance
401(k) matching
Paid time off
Senior ML Inference Engineer — High-Performance LLM Serving
Senior ML Inference Engineer — High-Performance LLM Serving

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 168,000 - 227,000
Health insurance
401(k) matching
Paid time off
+2
ML Inference Engineer - AWS Neuron & GenAI
ML Inference Engineer - AWS Neuron & GenAI

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 165,000 - 224,000
ML Inference Engineer - Custom Accelerator Data Plane
ML Inference Engineer - Custom Accelerator Data Plane

Amazon • Cupertino (CA)

On-site
USD 165,000 - 224,000
Senior ML Acceleration Hardware Engineer
Senior ML Acceleration Hardware Engineer

Amazon Web Services (AWS) • Austin (TX)

On-site
USD 136,000 - 184,000
Software Engineer II: AI/ML Inference on Custom Hardware
Software Engineer II: AI/ML Inference on Custom Hardware

Amazon Inc. • Cupertino (CA)

On-site
USD 165,000 - 224,000
Health insurance
RSUs
401(k) matching
+2
Senior Embedded Software Engineer - ML Hardware & Firmware
Senior Embedded Software Engineer - ML Hardware & Firmware

Amazon Inc. • Austin (TX), Northern (KY)

Hybrid
USD 168,000 - 227,000
Health insurance
RSUs
401(k) matching
+1