Software Development Engineer – AI/ML Networking Disaggregated Inference, Annapurna Labs , Elastic Collectives

Amazon Inc.

Cupertino (CA)

On-site

USD 180,000 - 240,000

Full time

6 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Annapurna Labs in Cupertino, part of AWS, is seeking a Software Development Engineer focused on AI/ML networking disaggregated inference. You will build data-m movement software across accelerators and servers, profile workloads, and collaborate across the stack from network transport to inference frameworks.

Strong C/C++ and Linux skills are required, with a passion for performance and the ability to push hardware limits.

Qualifications

  • Strong C/C++ and Linux experience in performance-critical systems.
  • Interest in low-level software and ability to measure and optimize performance.
  • Exposure to high-speed networking or HPC interconnects is a strong plus; embedded-systems experience welcome.
  • Prior AI/ML experience is not required; strong systems engineer willing to learn the ML side.

Responsibilities

  • Build and optimize low-level data-movement software transferring KV cache and activations across accelerators and servers over AWS's high-performance fabric.
  • Profile workloads, identify bottlenecks, and push components toward hardware limits.
  • Work across the stack—from network transport to inference frameworks—collaborating with chip, runtime, and model teams.
  • Deliver features to large-scale clusters for major customers and AI models in production.

Skills

C/C++
Low-level systems
Linux proficiency
Performance-minded
High-speed networking exposure

Tools

libfabric
UCX
NCCL
MPI
InfiniBand/RDMA

Job description

Software Development Engineer – AI/ML Networking Disaggregated Inference, Annapurna Labs , Elastic Collectives

Every token a large language model generates depends on data reaching the right accelerator at the right moment. As AI models outgrow any single chip, the network between accelerators becomes the bottleneck that decides how fast — and how affordably — the world's largest models can serve real users. That network layer is what our team builds.

We're looking for an engineer to work at the frontier of disaggregated inference: splitting LLM serving into separate prefill and decode pools and moving the model's KV cache between them at the limit of what the hardware allows. Get it right and users get answers in milliseconds; get it wrong and the fastest accelerators in the world sit idle waiting on data. You'll build components of the high-speed transfer path that make that difference, and you'll learn to measure success in how close we run to the theoretical peak of the machine.

In this role you will:
  • Build and optimize the low-level data-movement software that transfers KV cache and activations across accelerators, servers, and heterogeneous memory — over AWS's highest-performance network fabric.
  • Profile real workloads, find the true bottleneck, and close the gap between "it works" and "it runs fast" — pushing components toward the hardware's limit.
  • Work across the stack — from network transport up to the inference frameworks — learning from the teams building the chips, runtime, and models.
  • Deliver features that ship to our largest clusters, for our largest customers, serving the largest AI models in production.
What we're looking for:
  • Strong C/C++ and a genuine interest in low-level, performance-critical systems — solid command of Linux, memory, and writing fast code.
  • The instinct to ask "how fast could this go?" and the discipline to measure it.
  • Exposure to high-speed networking, HPC interconnects, or GPU/accelerator systems (RDMA, InfiniBand, libfabric, UCX, NCCL, MPI) is a strong plus; embedded-systems experience is welcome.
  • Prior AI/ML experience is not required — if you're a strong systems engineer eager to learn, we'll teach you the ML side.
  • If you like solving genuinely hard problems, working alongside HPC and ML customers, iterating fast, and shipping at a scale few places can offer, come join us. You'll work alongside senior engineers and Principal Engineers who've built this layer from the ground up, with real room to grow your scope and technical depth — on a team at the leading edge of AI/ML infrastructure.
About the team

About the team: You'd be joining Annapurna Labs, an integral part of AWS. Annapurna designs the hardware and software building blocks behind EC2 — every EC2 instance runs on hardware we designed. We specialize in the chips, systems, and software that optimize the AWS customer experience, and this team sits where AI meets the silicon and the network underneath it.

A day in the life

Annapurna Labs, a crucial part of AWS, is responsible for developing hardware and software components for EC2 infrastructure. Our team focuses on building networking solutions that for Machine Learning (ML) and High-Performance Computing (HPC) workloads on AWS.

We have mixed discipline orgs, you’d be working side by side with infrastructure experts, hardware engineers, RTL engineers, scientists & architects. Our workforce spans the globe and is truly international, you’ll find yourself working side by

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Sr. Software Development Engineer – AI/ML Networking Disaggregated Inference, Annapurna Labs, Annapurna Labs
Sr. Software Development Engineer – AI/ML Networking Disaggregated Inference, Annapurna Labs, Annapurna Labs

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 193,000 - 262,000
Health insurance
Stock options
Sr. Software Development Engineer – AI/ML Networking Disaggregated Inference, Annapurna Labs, Annapurna Labs
Sr. Software Development Engineer – AI/ML Networking Disaggregated Inference, Annapurna Labs, Annapurna Labs

Amazon • Cupertino (CA)

On-site
USD 193,000 - 262,000
RSUs
Health insurance
401(k) matching
+1
Sr. Software Development Engineer – AI/ML Networking Disaggregated Inference, Annapurna Labs, Annapurna Labs (AWS)
Sr. Software Development Engineer – AI/ML Networking Disaggregated Inference, Annapurna Labs, Annapurna Labs (AWS)

Amazon Inc. • Cupertino (CA)

On-site
USD 193,000 - 262,000
Senior Systems Engineer, AI Inference & HPC Networking
Senior Systems Engineer, AI Inference & HPC Networking

Amazon • Cupertino (CA)

On-site
USD 193,000 - 262,000
RSUs
Health insurance
401(k) matching
+1
Software Engineer — AI/ML Networking for Inference
Software Engineer — AI/ML Networking for Inference

Amazon Inc. • Cupertino (CA)

On-site
USD 180,000 - 240,000
Sr. Software Development Engineer, AI/ML, AWS Neuron, Model Inference
Sr. Software Development Engineer, AI/ML, AWS Neuron, Model Inference

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 193,000 - 262,000
Health insurance
401(k) matching
Paid time off
Lead Software Engineer, ML Network Stack - Annapurna Labs
Lead Software Engineer, ML Network Stack - Annapurna Labs

Amazon Inc. • Seattle (WA)

On-site
USD 168,000 - 227,000
SoC Systems Software Engineer, Annapurna Labs Machine Learning Accelerators, AWS
SoC Systems Software Engineer, Annapurna Labs Machine Learning Accelerators, AWS

Amazon Web Services (AWS) • Austin (TX)

On-site
USD 144,000 - 194,000
Health insurance
401(k) matching
Paid time off
+1
Software Development Engineer, AI/ML, AWS Neuron, Model Inference
Software Development Engineer, AI/ML, AWS Neuron, Model Inference

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 165,000 - 224,000
Sign-on payments
RSUs
Health insurance
+3
SoC Systems Software Engineer, Annapurna Labs Machine Learning Accelerators, AWS
SoC Systems Software Engineer, Annapurna Labs Machine Learning Accelerators, AWS

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 165,000 - 224,000
Health insurance
401(k)
Paid time off
+1