Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.
Annapurna Labs in Cupertino, CA, part of AWS, is seeking a Software Development Engineer focusing on AI/ML networking disaggregated inference. You will build low-level data movement software across accelerators and servers to minimize latency in LLM serving.
You will profile workloads, identify bottlenecks, and push components toward hardware limits while collaborating across chips, runtimes, and models teams. Prior AI/ML experience not required.
Every token a large language model generates depends on data reaching the right accelerator at the right moment. As AI models outgrow any single chip, the network between accelerators becomes the bottleneck that decides how fast — and how affordably — the world's largest models can serve real users. That network layer is what our team builds.
We're looking for an engineer to work at the frontier of disaggregated inference: splitting LLM serving into separate prefill and decode pools and moving the model's KV cache between them at the limit of what the hardware allows. Get it right and users get answers in milliseconds; get it wrong and the fastest accelerators in the world sit idle waiting on data. You'll build components of the high-speed transfer path that make that difference, and you'll learn to measure success in how close we run to the theoretical peak of the machine.
In this role you will:
What we're looking for:
About the team: You'd be joining Annapurna Labs, an integral part of AWS. Annapurna designs the hardware and software building blocks behind EC2 — every EC2 instance runs on hardware we designed. We specialize in the chips, systems, and software that optimize the AWS customer experience, and this team sits where AI meets the silicon and the network underneath it.
A day in the life: Annapurna Labs, a crucial part of AWS, is responsible for developing hardware and software components for EC2 infrastructure. Our team focuses on building networking solutions that for Machine Learning (ML) and High‑Performance Computing (HPC) workloads on AWS.
We have mixed discipline orgs, you’d be working side by side with infrastructure experts, hardware engineers, RTL engineers, scientists & architects. Our workforce spans the globe and is truly international, you’ll find yourself working side by side with individuals from numerous countries. We take mentorship seriously, you can both expect senior mentorship and will be expected to mentor new and junior engineers.
The pace is fast as we work on the latest advancements of AI/ML, but we take the time to bond as a team and enjoy the successes. We offer flexibility in working hours, and respect WLB as a core org tenet. The team enjoys working with numerous principal‑level engineers and closely with directors, career growth opportunities are certainly available. This is a role where you will always be encouraged to keep learning, the AI/ML field is fast moving and constantly evolving.
About the team: Our team is dedicated to supporting new members. We have a broad mix of experience levels and tenures, and we’re building an environment that celebrates knowledge‑sharing and mentorship. Our senior members enjoy one‑on‑one mentoring and thorough, but kind, code reviews. We care about your career growth and strive to assign projects that help our team members develop your engineering expertise so you feel empowered to take on more complex tasks in the future.
Diverse Experiences: AWS values diverse experiences. Even if you do not meet all of the qualifications and skills listed in the job description, we encourage candidates to apply. If your career is just starting, hasn’t followed a traditional path, or includes alternative experiences, don’t let it stop you from applying.
About AWS: Amazon Web Services (AWS) is the world’s most comprehensive and broadly adopted cloud platform. We pioneered cloud computing and never stopped innovating — that’s why customers from the most successful startups to Global 500 companies trust our robust suite of products and services to power their businesses.
Inclusive Team Culture: Here at AWS, it’s in our nature to learn and be curious. Our employee‑led affinity groups foster a culture of inclusion that empower us to be proud of our differences. Ongoing events and learning experiences, including our Conversations on Race and Ethnicity (CORE) and AmazeCon (diversity) conferences, inspire us to never stop embracing our uniqueness.
Work/Life Balance: We value work‑life harmony. Achieving success at work should never come at the expense of sacrifices at home, which is why we strive for flexibility as part of our working culture. When we feel supported in the workplace and at home, there’s nothing we can’t achieve in the cloud.
Mentorship & Career Growth: We’re continuously raising our performance bar as we strive to become Earth’s Best Employer. That’s why you’ll find endless knowledge‑sharing, mentorship and other career‑advancing resources here to help you develop into a better‑rounded professional.
The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at https://amazon.jobs/en/benefits .
USA, CA, Cupertino - 165,200.00 - 223,600.00 USD annually
Important FAQs for current Government employees Before proceeding, please review the following FAQs https://www.amazon.jobs/en/faqs#faqs-for-us-government-employees
Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status.