Senior AI/ML Systems Engineer - Disaggregated Inference

Energy Jobline ZR

Cupertino (CA)

On-site

USD 193,000 - 262,000

Full time

11 days ago
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Health insurance
401(k) matching
Paid time off
Parental leave

Job summary

Annapurna Labs, an integral part of AWS, is seeking a systems-focused engineer to advance disaggregated inference for large AI models. You will work on high-speed data movement across accelerators and memory, optimizing the path that powers real-time ML at scale.

In this role, you will collaborate across kernel, network, runtime, and models teams, mentoring others and delivering features for the world's largest clusters.

Qualifications

  • 5+ years of non professional software development experience
  • 5+ years of programming with at least one software programming experience
  • 5+ years of leading design or architecture (design patterns, reliability and scaling) of new and existing systems experience
  • 5+ years of full software development life cycle, including coding standards, code reviews, source control management, build processes, testing, and operations experience
  • Experience as a mentor, tech lead or leading an engineering team
  • Must have C/C++ Coding Experience

Responsibilities

  • Build and optimize the low-level data-movement software across accelerators, servers, and heterogeneous memory — over AWS's highest-performance network fabric.
  • Push performance to the hardware roofline: profile, find the real bottleneck, and close the gap between "it works" and "it runs as fast as physics permits."
  • Work across the stack — from kernel and network transport up to the inference frameworks — and partner with teams building the chips, runtime, and models.
  • Deliver features that run on our largest clusters, for our largest customers, serving the largest AI models in production.

Skills

C/C++ coding
Linux kernel
Performance profiling
High-speed networking
Tech leadership

Education

Bachelor's degree in computer science or equivalent

Tools

NCCL
NIXL
NVSHMEM

Job description

Annapurna Labs, an integral part of AWS, is seeking a systems-focused engineer to advance disaggregated inference for large AI models. You will work on high-speed data movement across accelerators and memory, optimizing the path that powers real-time ML at scale.

In this role, you will collaborate across kernel, network, runtime, and models teams, mentoring others and delivering features for the world's largest clusters.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Engineer - AI/ML Disaggregated Inference
Software Engineer - AI/ML Disaggregated Inference

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 165,000 - 224,000
Senior Systems Engineer, AI Inference & HPC Networking
Senior Systems Engineer, AI Inference & HPC Networking

Amazon • Cupertino (CA)

On-site
USD 193,000 - 262,000
RSUs
Health insurance
401(k) matching
+1
SDE, AI/ML Networking — Disaggregated Inference Expert
SDE, AI/ML Networking — Disaggregated Inference Expert

Amazon • Cupertino (CA)

On-site
USD 165,000 - 224,000
Health insurance
RSUs
Paid time off
Systems Engineer, High‑Speed AI Inference Networking
Systems Engineer, High‑Speed AI Inference Networking

Amazon • Cupertino (CA)

On-site
USD 165,000 - 224,000
Health insurance
RSUs / stock options
401(k) matching
+2
Senior ML Systems Engineer - AI Inference on AWS Neuron
Senior ML Systems Engineer - AI Inference on AWS Neuron

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 193,000 - 262,000
Senior AI/ML Systems Engineer - Flexible Hours & Networking
Senior AI/ML Systems Engineer - Flexible Hours & Networking

Amazon • Cupertino (CA)

On-site
USD 193,000 - 262,000
Senior AI/ML Software Engineer - High-Perf Inference
Senior AI/ML Software Engineer - High-Perf Inference

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 168,000 - 227,000
AI/ML Inference Engineer for AWS Neuron
AI/ML Inference Engineer for AWS Neuron

Amazon • Cupertino (CA)

On-site
USD 165,000 - 224,000
Health insurance
401(k) matching
Paid time off
+2
ML Inference Engineer - AWS Neuron & GenAI
ML Inference Engineer - AWS Neuron & GenAI

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 165,000 - 224,000
Senior AI/ML Cloud Hardware Systems Engineer
Senior AI/ML Cloud Hardware Systems Engineer

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 174,000 - 235,000