AI/ML Network Infrastructure Engineer I

Amazon Inc.

Cupertino (CA)

On-site

USD 126,000 - 171,000

Full time

6 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Amazon Inc. in Cupertino, CA is seeking an early-career software engineer to join the team that owns the network stack for EC2 distributed AI/ML systems, building high-performance C/C++ code for communication libraries on AWS hardware.

You will develop tools to monitor, benchmark, and deliver software for massive GPU clusters, and collaborate across Linux environments and diverse stacks. This role blends HPC networking with ML infrastructure, offering hands-on experience with CUDA, MPI, and

Qualifications

  • Bachelor's degree or above in Computer Science, Computer Engineering, or related fields.
  • Strong proficiency in C/C++.
  • Solid coursework or project experience in: Operating Systems (Linux internals, kernel concepts, memory management); Parallel Computer Architecture (multithreading, SIMD, GPU programming, cache coherence); Distributed Systems (consensus, message passing, fault tolerance, scalability).
  • Familiarity with Linux development environments and toolchains.

Responsibilities

  • Write high-performance C/C++ code for network communication libraries running on custom AWS hardware.
  • Build and maintain infrastructure that monitors functionality and performance of large-scale AI/ML workloads.
  • Develop automation using Python and AWS tools (CI/CD, Grafana, Athena) to test, benchmark, and deliver software to customers.
  • Design mechanisms to detect functional and performance regressions before they reach production.
  • Work across many instance types, software stacks, and Linux environments.

Skills

C/C++
Linux internals
Parallel computing
Distributed systems
Linux toolchains

Education

Bachelor's degree in CS/CE

Tools

CUDA
MPI
RDMA

Job description

Amazon Inc. in Cupertino, CA is seeking an early-career software engineer to join the team that owns the network stack for EC2 distributed AI/ML systems, building high-performance C/C++ code for communication libraries on AWS hardware.

You will develop tools to monitor, benchmark, and deliver software for massive GPU clusters, and collaborate across Linux environments and diverse stacks. This role blends HPC networking with ML infrastructure, offering hands-on experience with CUDA, MPI, and

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Software Engineer — AI/ML Networking for Inference
Software Engineer — AI/ML Networking for Inference

Amazon Inc. • Cupertino (CA)

On-site
USD 180,000 - 240,000
Senior ML Network Stack Engineer — Scale EC2 AI
Senior ML Network Stack Engineer — Scale EC2 AI

Amazon Inc. • Seattle (WA)

On-site
USD 168,000 - 227,000
Lead ML Network Stack Engineer for Scalable EC2 AI
Lead ML Network Stack Engineer for Scalable EC2 AI

Amazon • Cupertino (CA)

On-site
USD 193,000 - 262,000
Health insurance
RSUs and sign-on options
401(k) matching
+2
Lead ML Network Stack Engineer (RDMA, CUDA, NCCL)
Lead ML Network Stack Engineer (RDMA, CUDA, NCCL)

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 180,000 - 240,000
Sign-on payments
RSUs
Health Insurance
+14
ML Network Stack Lead Engineer
ML Network Stack Lead Engineer

Amazon • Cupertino (CA), Northern (KY)

Hybrid
USD 193,000 - 262,000
Health insurance
401(k) matching
Paid time off
+1
Senior AI/ML Systems Engineer - Flexible Hours & Networking
Senior AI/ML Systems Engineer - Flexible Hours & Networking

Amazon • Cupertino (CA)

On-site
USD 193,000 - 262,000
Senior Systems Engineer, AI Inference & HPC Networking
Senior Systems Engineer, AI Inference & HPC Networking

Amazon • Cupertino (CA)

On-site
USD 193,000 - 262,000
RSUs
Health insurance
401(k) matching
+1
AI/ML Server Hardware Architect in High-Performance Systems
AI/ML Server Hardware Architect in High-Performance Systems

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 183,000 - 248,000
Health insurance
401(k) matching
RSU equity
Senior ML Network Stack Engineering Manager
Senior ML Network Stack Engineering Manager

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 220,000 - 298,000
Health insurance
401(k) matching
Paid time off
+1
Senior Manager, ML Network Stack & HPC Systems
Senior Manager, ML Network Stack & HPC Systems

Amazon • Seattle (WA)

On-site
USD 220,000 - 298,000
Health insurance
401(k) matching
Paid time off
+1