LLM Inference Software Engineer - Data Plane

Amazon Web Services (AWS)

Cupertino (CA)

On-site

USD 165,000 - 224,000

Full time

12 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health insurance
RSUs
401(k) matching
Paid time off

Job summary

Annapurna Labs (U.S.) Inc. in Cupertino, CA is seeking a Software Development Engineer to own the design and implementation of the inference data plane for large models on custom hardware.

You will drive model execution, memory management, data movement, and integration with serving engines for production-scale performance. The role focuses on low-level optimization, end-to-end feature ownership, and contributions to scalable test and profiling infrastructure across CPU, GPU, simulator, and

Qualifications

  • Bachelor's degree or equivalent required.
  • 4+ years of full software development life cycle experience including coding standards, reviews, source control, build, testing and operations.
  • Knowledge of computer architecture, operating systems, and parallel computing.
  • Knowledge of Linux fundamentals.
  • Strong proficiency in C/C++.
  • Experience developing compute kernels for GPUs, DSPs, or custom accelerators.
  • Proven track record of owning and delivering complex software features end-to-end.

Responsibilities

  • Develop and optimize compute kernels for a custom ML accelerator architecture, targeting production-level performance for large language model inference.
  • Implement and validate LLM architectures end-to-end - from PyTorch model definition through distributed execution on custom hardware.
  • Integrate custom accelerator backends into open-source ML serving frameworks (vLLM, PyTorch), including scheduler extensions, memory management, and model parallelism.
  • Build and maintain test infrastructure for model correctness validation across CPU, GPU, simulator, and hardware targets.
  • Profile and optimize inference workloads - identify bottlenecks, instrument critical paths, and drive latency and throughput improvements from simulation through hardware bring-up.
  • Own features end-to-end: from design through implementation, testing, and integration into the broader software stack.
  • Contribute to CI/CD pipelines that gate model and kernel changes on correctness and performance regressions.

Skills

C/C++
Parallel computing
Linux
Computer architecture
GPU compute kernels

Education

Bachelor's degree

Tools

PyTorch
JAX
vLLM
TensorRT
TorchXLA

Job description

Annapurna Labs (U.S.) Inc. in Cupertino, CA is seeking a Software Development Engineer to own the design and implementation of the inference data plane for large models on custom hardware.

You will drive model execution, memory management, data movement, and integration with serving engines for production-scale performance. The role focuses on low-level optimization, end-to-end feature ownership, and contributions to scalable test and profiling infrastructure across CPU, GPU, simulator, and

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Inference Engineer - Custom Hardware Data Plane
ML Inference Engineer - Custom Hardware Data Plane

Socket.dev • Cupertino (CA)

On-site
USD 165,000 - 224,000
RSUs
Health benefits
401(k) matching
Senior ML Inference Engineer - Custom Hardware, LLMs
Senior ML Inference Engineer - Custom Hardware, LLMs

Socket.dev • Cupertino (CA)

On-site
USD 193,000 - 262,000
Senior AI/ML Inference Engineer for LLMs on Neuron
Senior AI/ML Inference Engineer for LLMs on Neuron

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 168,000 - 228,000
AI/ML Inference Engineer for AWS Neuron
AI/ML Inference Engineer for AWS Neuron

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 165,000 - 224,000
Cloud-Scale Physical Design Engineer – ML Inference
Cloud-Scale Physical Design Engineer – ML Inference

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 157,000 - 213,000
Health insurance
401(k) matching
Paid time off
+1
Senior LLM Inference Engineer - End-to-End, Equity
Senior LLM Inference Engineer - End-to-End, Equity

Entrada Ventures • Santa Clara (CA)

On-site
USD 180,000 - 280,000
Competitive compensation
Equity opportunities
Benefits in Santa Clara, CA
Senior SDE — AI/ML Inference on Trainium/Neuron
Senior SDE — AI/ML Inference on Trainium/Neuron

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 193,000 - 262,000
RSUs
Comprehensive benefits
401(k)
Senior LLM Inference Architect — Heterogeneous Hardware
Senior LLM Inference Architect — Heterogeneous Hardware

d-Matrix inc. • Santa Clara (CA)

On-site
USD 130,000 - 170,000
Competitive compensation
Equity
Inclusive work environment
Software Development Manager: LLM Inference Enablement
Software Development Manager: LLM Inference Enablement

Amazon • Cupertino (CA)

On-site
USD 213,000 - 288,000
RSUs (Restricted Stock Units)
Health insurance
401(k) plan
Cloud-Scale Physical Design Engineer for ML Inference
Cloud-Scale Physical Design Engineer for ML Inference

Amazon Web Services (AWS) • Austin (TX)

On-site
USD 136,000 - 184,000
Health insurance
401(k) matching
Paid time off
+2