Systems Engineer, High‑Speed AI Inference Networking

Amazon

Cupertino (CA)

On-site

USD 165,000 - 224,000

Full time

10 days ago
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Health insurance
RSUs / stock options
401(k) matching
Paid time off
Parental leave

Job summary

Annapurna Labs in Cupertino, CA, part of AWS, is seeking a Software Development Engineer focusing on AI/ML networking disaggregated inference. You will build low-level data movement software across accelerators and servers to minimize latency in LLM serving.

You will profile workloads, identify bottlenecks, and push components toward hardware limits while collaborating across chips, runtimes, and models teams. Prior AI/ML experience not required.

Qualifications

  • Strong C/C++ and a genuine interest in low-level, performance-critical systems — solid command of Linux, memory, and writing fast code.
  • Exposure to high-speed networking or HPC interconnects (RDMA, InfiniBand, libfabric, UCX, NCCL, MPI) is a strong plus; embedded-systems experience is welcome.
  • Prior AI/ML experience is not required — if you're a strong systems engineer eager to learn, we'll teach you the ML side.

Responsibilities

  • Build and optimize the low-level data‑movement software that transfers KV cache and activations across accelerators, servers, and heterogeneous memory — over AWS's highest-performance network fabric.
  • Profile real workloads, find the true bottleneck, and close the gap between "it works" and "it runs fast" — pushing components toward the hardware's limit.
  • Work across the stack — from network transport up to the inference frameworks — learning from the teams building the chips, runtime, and models.
  • Deliver features that ship to our largest clusters, for our largest customers, serving the largest AI models in production.

Skills

C/C++
Linux
Networking concepts

Education

Bachelor's degree in computer science or equivalent

Job description

Annapurna Labs in Cupertino, CA, part of AWS, is seeking a Software Development Engineer focusing on AI/ML networking disaggregated inference. You will build low-level data movement software across accelerators and servers to minimize latency in LLM serving.

You will profile workloads, identify bottlenecks, and push components toward hardware limits while collaborating across chips, runtimes, and models teams. Prior AI/ML experience not required.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Engineer - AI/ML Disaggregated Inference
Software Engineer - AI/ML Disaggregated Inference

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 165,000 - 224,000
Senior Systems Engineer, AI Inference & HPC Networking
Senior Systems Engineer, AI Inference & HPC Networking

Amazon • Cupertino (CA)

On-site
USD 193,000 - 262,000
RSUs
Health insurance
401(k) matching
+1
Senior AI/ML Systems Engineer - Flexible Hours & Networking
Senior AI/ML Systems Engineer - Flexible Hours & Networking

Amazon • Cupertino (CA)

On-site
USD 193,000 - 262,000
SDE, AI/ML Networking — Disaggregated Inference Expert
SDE, AI/ML Networking — Disaggregated Inference Expert

Amazon • Cupertino (CA)

On-site
USD 165,000 - 224,000
Health insurance
RSUs
Paid time off
Senior AI/ML Systems Engineer - Disaggregated Inference
Senior AI/ML Systems Engineer - Disaggregated Inference

Energy Jobline ZR • Cupertino (CA)

On-site
USD 193,000 - 262,000
Health insurance
401(k) matching
Paid time off
+1
AI/ML Network Infrastructure Engineer I
AI/ML Network Infrastructure Engineer I

Amazon • Cupertino (CA)

On-site
USD 127,000 - 185,000
Senior AI/ML Cloud Hardware Systems Engineer
Senior AI/ML Cloud Hardware Systems Engineer

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 174,000 - 235,000
ML Inference Engineer - AWS Neuron & GenAI
ML Inference Engineer - AWS Neuron & GenAI

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 165,000 - 224,000
Senior AI/ML Software Engineer - High-Perf Inference
Senior AI/ML Software Engineer - High-Perf Inference

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 168,000 - 227,000
Senior ML Systems Engineer - AI Inference on AWS Neuron
Senior ML Systems Engineer - AI Inference on AWS Neuron

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 193,000 - 262,000