Senior ML Inference Systems Engineer - Custom Accelerator

Amazon

Cupertino (CA)

On-site

USD 193,000 - 262,000

Full time

3 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health insurance
RSUs
401(k) matching

Job summary

Amazon is seeking a Senior Software Development Engineer to own the design and implementation of the inference data plane for large models on custom hardware. You will build high-performance kernels, validate architectures end-to-end, and drive optimizations across CPU, GPU, and hardware targets.

You will integrate backends into ML serving frameworks, contribute to CI/CD, and mentor engineers while advancing the frontier of model inference at scale in a ground-up environment.

Qualifications

  • Bachelor's degree in computer science or equivalent.
  • 7+ years of full software development life cycle experience.
  • Strong proficiency in C/C++.
  • Strong Linux systems knowledge.
  • Experience developing compute kernels for GPUs, DSPs, or custom accelerators.
  • Proven track record of owning and delivering complex software features end-to-end.

Responsibilities

  • Develop and optimize compute kernels for a custom ML accelerator architecture.
  • Implement and validate LLM architectures end-to-end—from PyTorch model definition to distributed execution on custom hardware.
  • Integrate custom accelerator backends into open-source ML serving frameworks (vLLM, PyTorch).
  • Build and maintain test infrastructure for model correctness across CPU, GPU, simulator, and hardware targets.
  • Profile and optimize inference workloads to reduce latency and increase throughput across the stack.
  • Own features end-to-end from design through implementation, testing, and integration.
  • Contribute to CI/CD pipelines ensuring correctness and performance.
  • Mentor engineers and elevate engineering standards across the team.

Skills

C/C++
Linux
Compute kernels
End-to-end ownership
Model serving
Distributed systems

Education

B.S. CS or equivalent

Tools

PyTorch
vLLM
JAX
TensorRT
TorchXLA
Dynamo

Job description

Amazon is seeking a Senior Software Development Engineer to own the design and implementation of the inference data plane for large models on custom hardware. You will build high-performance kernels, validate architectures end-to-end, and drive optimizations across CPU, GPU, and hardware targets.

You will integrate backends into ML serving frameworks, contribute to CI/CD, and mentor engineers while advancing the frontier of model inference at scale in a ground-up environment.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior ML Inference Engineer - Custom Hardware, LLMs
Senior ML Inference Engineer - Custom Hardware, LLMs

Socket.dev • Cupertino (CA)

On-site
USD 193,000 - 262,000
Senior AI/ML Software Engineer for Hardware Accelerators
Senior AI/ML Software Engineer for Hardware Accelerators

Amazon • Seattle (WA)

On-site
USD 168,000 - 227,000
Health insurance
401(k) matching
RSUs
Senior ML Software Engineer, Data Plane
Senior ML Software Engineer, Data Plane

Socket.dev • Cupertino (CA)

On-site
USD 193,000 - 262,000
Senior ML Software Engineer, Data Plane (AWS)
Senior ML Software Engineer, Data Plane (AWS)

Amazon • Cupertino (CA)

On-site
USD 193,000 - 262,000
Health insurance
RSUs
401(k) matching
Senior SoC Systems Software Engineer – ML Accelerators
Senior SoC Systems Software Engineer – ML Accelerators

Amazon • Cupertino (CA)

On-site
USD 193,000 - 261,500
RSUs
Health insurance
401(k) matching
+2
Senior ML Accelerator Runtime Engineer
Senior ML Accelerator Runtime Engineer

Amazon • Seattle (WA)

On-site
USD 143,700 - 194,400
Health insurance
Dental
Vision
+4
Senior ML Acceleration Hardware Engineer — PCIe/SerDes
Senior ML Acceleration Hardware Engineer — PCIe/SerDes

Amazon • Cupertino (CA)

On-site
USD 143,000 - 248,000
Engineering Manager, ML Kernel Performance
Engineering Manager, ML Kernel Performance

Amazon • Cupertino (CA)

On-site
USD 212,700 - 287,700
Health insurance
401(k) matching
Parental leave
+2
Senior ML Systems Engineer - AI Inference on AWS Neuron
Senior ML Systems Engineer - AI Inference on AWS Neuron

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 193,000 - 262,000
LLM Inference Systems Engineer — Custom Accelerator
LLM Inference Systems Engineer — Custom Accelerator

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 165,000 - 224,000
Health insurance
RSUs
401(k) matching
+2