LLM Inference Systems Engineer — Custom Accelerator

Amazon Web Services (AWS)

Cupertino (CA)

On-site

USD 165,000 - 224,000

Full time

3 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health insurance
RSUs
401(k) matching
Paid time off
Parental leave

Job summary

Annapurna Labs (U.S.) Inc. in Cupertino, CA is seeking a Software Development Engineer to own the design and implementation of our inference data plane for large models running on custom hardware.

You will shape model execution, memory management, data movement, and serving integration across the full inference path. You will write and optimize low-level code, validate architectures end-to-end, and build profiling and test infra to drive performance across the stack, from simulation to hardware

Qualifications

  • Bachelor's degree or equivalent.
  • 4+ years of full software development lifecycle.
  • Knowledge of computer architecture, operating systems, and parallel computing.
  • Knowledge of Linux fundamentals.
  • Strong proficiency in C/C++.
  • Experience developing compute kernels for GPUs, DSPs, or custom accelerators.
  • Proven track record of owning and delivering complex software features.

Responsibilities

  • Develop and optimize compute kernels for a custom ML accelerator architecture, targeting production-level performance for large language model inference.
  • Implement and validate LLM architectures end-to-end - from PyTorch model definition through distributed execution on custom hardware.
  • Integrate custom accelerator backends into open-source ML serving frameworks (vLLM, PyTorch), including scheduler extensions, memory management, and model parallelism.
  • Build and maintain test infrastructure for model correctness validation across CPU, GPU, simulator, and hardware targets.
  • Profile and optimize inference workloads - identify bottlenecks, instrument critical paths, and drive latency and throughput improvements from simulation through hardware bring-up.
  • Own features end-to-end: from design through implementation, testing, and integration into the broader software stack.
  • Contribute to CI/CD pipelines that gate model and kernel changes on correctness and performance regressions.

Skills

C/C++
Linux fundamentals
Computer architecture
Performance optimization
GPU/accelerator kernels

Education

Bachelor's degree or equivalent

Job description

Annapurna Labs (U.S.) Inc. in Cupertino, CA is seeking a Software Development Engineer to own the design and implementation of our inference data plane for large models running on custom hardware.

You will shape model execution, memory management, data movement, and serving integration across the full inference path. You will write and optimize low-level code, validate architectures end-to-end, and build profiling and test infra to drive performance across the stack, from simulation to hardware

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

LLM Inference Runtime Architect for AI Accelerator
LLM Inference Runtime Architect for AI Accelerator

United States Digital Space LLC • United States

Remote
USD 180,000 - 250,000
Senior ML Inference Engineer - Custom Hardware, LLMs
Senior ML Inference Engineer - Custom Hardware, LLMs

Socket.dev • Cupertino (CA)

On-site
USD 193,000 - 262,000
ML Inference Engineer - AWS Neuron & GenAI
ML Inference Engineer - AWS Neuron & GenAI

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 165,000 - 224,000
SoC Systems Software Engineer — ML Accelerators
SoC Systems Software Engineer — ML Accelerators

Amazon Web Services (AWS) • Austin (TX)

On-site
USD 144,000 - 194,000
Health insurance
401(k) matching
Paid time off
+1
Senior ML Inference Systems Engineer - Custom Accelerator
Senior ML Inference Systems Engineer - Custom Accelerator

Amazon • Cupertino (CA)

On-site
USD 193,000 - 262,000
Health insurance
RSUs
401(k) matching
SoC Systems Software Engineer - Hardware/Software Co-Design
SoC Systems Software Engineer - Hardware/Software Co-Design

Amazon • Cupertino (CA)

On-site
USD 165,200 - 223,600
RSUs
Health insurance
401(k) matching
+1
Lead SDE: Hardware/Software Co-Design for ML Acceleration
Lead SDE: Hardware/Software Co-Design for ML Acceleration

Socket.dev • Cupertino (CA)

On-site
USD 193,000 - 262,000
Health insurance
401(k) matching
Paid time off
+1
Senior SDE: C/C++ for HW/SW Co-Design & ML Acceleration
Senior SDE: C/C++ for HW/SW Co-Design & ML Acceleration

Amazon • Cupertino (CA)

On-site
USD 193,300 - 261,500
Health insurance
401(k) matching
Paid time off
+1
ML Software Engineer, Data Plane
ML Software Engineer, Data Plane

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 165,000 - 224,000
Health insurance
RSUs
401(k) matching
+2
Senior LLM Inference Engineer - End-to-End, Equity
Senior LLM Inference Engineer - End-to-End, Equity

Entrada Ventures • Santa Clara (CA)

On-site
USD 180,000 - 280,000
Competitive compensation
Equity opportunities
Benefits in Santa Clara, CA