AI/ML Network Infrastructure Engineer I

Amazon

Cupertino (CA)

On-site

USD 127,000 - 185,000

Full time

31 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Annapurna Labs (U.S.) Inc. in Cupertino, CA is seeking an early-career engineer to own the network stack for EC2 distributed AI/ML systems. You will contribute to software enabling training of world-scale AI models across massive GPU clusters, supporting NCCL, NVSHMEM, and NIXL.

This role sits at the intersection of HPC networking and ML infrastructure, building systems powering the largest AI workloads in the cloud, with opportunities to influence cross-stack design and performance optimization.

Qualifications

  • Bachelor's degree or above in Computer Science, Computer Engineering, or related fields.
  • Strong proficiency in C/C++.
  • Familiarity with Linux development environments and toolchains.
  • Knowledge of operating systems, parallel computer architecture, and distributed systems is a plus.

Responsibilities

  • Write high-performance C/C++ code for network communication libraries running on custom AWS hardware
  • Build and maintain infrastructure that monitors functionality and performance of large-scale AI/ML workloads
  • Develop automation using Python and AWS tools (CI/CD, Grafana, Athena) to test, benchmark, and deliver software to customers
  • Design mechanisms to detect functional and performance regressions before they reach production
  • Work across many instance types, software stacks, and Linux environments

Skills

C/C++
Linux development environments
Network programming exposure
GPU programming concepts

Education

Bachelor's degree or above in Computer Science, Computer Engineering, or related fields

Tools

NCCL
NVSHMEM
NIXL

Job description

Annapurna Labs (U.S.) Inc. in Cupertino, CA is seeking an early-career engineer to own the network stack for EC2 distributed AI/ML systems. You will contribute to software enabling training of world-scale AI models across massive GPU clusters, supporting NCCL, NVSHMEM, and NIXL.

This role sits at the intersection of HPC networking and ML infrastructure, building systems powering the largest AI workloads in the cloud, with opportunities to influence cross-stack design and performance optimization.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior ML Networking & HPC Stack Lead
Senior ML Networking & HPC Stack Lead

Amazon • Cupertino (CA)

On-site
USD 193,000 - 262,000
Senior ML Networking Lead for Scalable AI Clusters
Senior ML Networking Lead for Scalable AI Clusters

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 193,000 - 262,000
RSUs
Comprehensive benefits
Health insurance
Engineering Manager - ML Network Stack & HPC Systems
Engineering Manager - ML Network Stack & HPC Systems

Socket.dev • Seattle (WA)

On-site
USD 185,000 - 250,000
Software Development Engineer I – AI/ML Network Infrastructure, Annapurna Labs
Software Development Engineer I – AI/ML Network Infrastructure, Annapurna Labs

Amazon • Cupertino (CA)

On-site
USD 127,000 - 185,000
AI Cloud Systems Engineer - Backend & Automation
AI Cloud Systems Engineer - Backend & Automation

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 129,000 - 175,000
AI/ML Inference Engineer for AWS Neuron
AI/ML Inference Engineer for AWS Neuron

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 165,000 - 224,000
Engineering Manager, ML Network Stack for EC2
Engineering Manager, ML Network Stack for EC2

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 185,000 - 250,000
Health insurance
401(k) matching
Paid time off
+1
Senior Manager, ML Network Stack & HPC Systems
Senior Manager, ML Network Stack & HPC Systems

Amazon • Seattle (WA)

On-site
USD 220,000 - 298,000
Health insurance
401(k) matching
Paid time off
+1
Senior Network Engineer — AI Infra & HPC Fabric Expert
Senior Network Engineer — AI Infra & HPC Fabric Expert

Nscale • Houston (TX)

On-site
USD 150,000 - 210,000
Competitive benefits package
Flexible paid time off
Parental leave
+1
Applied Scientist II — ML Systems for AI Accelerators
Applied Scientist II — ML Systems for AI Accelerators

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 171,000 - 223,000