Software Development Engineer I – AI/ML Network Infrastructure, Annapurna Labs

Amazon

Cupertino (CA)

On-site

USD 127,000 - 185,000

Full time

7 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Annapurna Labs (U.S.) Inc. in Cupertino, CA is seeking an early-career engineer to own the network stack for EC2 distributed AI/ML systems. You will contribute to software enabling training of world-scale AI models across massive GPU clusters, supporting NCCL, NVSHMEM, and NIXL.

This role sits at the intersection of HPC networking and ML infrastructure, building systems powering the largest AI workloads in the cloud, with opportunities to influence cross-stack design and performance optimization.

Qualifications

  • Bachelor's degree or above in Computer Science, Computer Engineering, or related fields.
  • Strong proficiency in C/C++.
  • Familiarity with Linux development environments and toolchains.
  • Knowledge of operating systems, parallel computer architecture, and distributed systems is a plus.

Responsibilities

  • Write high-performance C/C++ code for network communication libraries running on custom AWS hardware
  • Build and maintain infrastructure that monitors functionality and performance of large-scale AI/ML workloads
  • Develop automation using Python and AWS tools (CI/CD, Grafana, Athena) to test, benchmark, and deliver software to customers
  • Design mechanisms to detect functional and performance regressions before they reach production
  • Work across many instance types, software stacks, and Linux environments

Skills

C/C++
Linux development environments
Network programming exposure
GPU programming concepts

Education

Bachelor's degree or above in Computer Science, Computer Engineering, or related fields

Tools

NCCL
NVSHMEM
NIXL

Job description

Description

We're looking for a talented early-career engineer to join our team that owns the network stack for EC2 distributed AI/ML systems. You'll work on software that enables the world's largest AI models to train across massive GPU clusters, developing support for communication libraries and frameworks like NCCL, NVSHMEM, and NIXL.

Description

We're looking for a talented early-career engineer to join our team that owns the network stack for EC2 distributed AI/ML systems. You'll work on software that enables the world's largest AI models to train across massive GPU clusters, developing support for communication libraries and frameworks like NCCL, NVSHMEM, and NIXL.

This is a ground-floor opportunity to work at the intersection of high-performance computing, networking, and machine learning infrastructure - building the systems that power the largest AI workloads in the cloud.

Key job responsibilities
  • Write high-performance C/C++ code for network communication libraries running on custom AWS hardware
  • Build and maintain infrastructure that monitors functionality and performance of large-scale AI/ML workloads
  • Develop automation using Python and AWS tools (CI/CD, Grafana, Athena) to test, benchmark, and deliver software to customers
  • Design mechanisms to detect functional and performance regressions before they reach production
  • Work across many instance types, software stacks, and Linux environments
Basic Qualifications
  • Bachelor's degree or above in Computer Science, Computer Engineering, or related fields
  • Strong proficiency in C/C++
  • Solid coursework or project experience in: 1/ Operating Systems (Linux internals, kernel concepts, memory management) 2/ Parallel Computer Architecture (multi-threading, SIMD, GPU programming, cache coherence) 3/ Distributed Systems (consensus, message passing, fault tolerance, scalability)
  • Familiarity with Linux development environments and toolchains
Preferred Qualifications
  • Internship experience in ML communications, HPC networking, or RDMA/high-speed interconnects
  • Exposure to network programming (sockets, MPI, collective communication patterns)
  • Experience with performance profiling and optimization
  • Familiarity with GPU programming (CUDA) or hardware-software co-design
  • Contributions to open-source projects in systems, networking, or HPC

Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status.

Los Angeles County applicants: Job duties for this position include: work safely and cooperatively with other employees, supervisors, and staff; adhere to standards of excellence despite stressful conditions; communicate effectively and respectfully with employees, supervisors, and staff to ensure exceptional customer service; and follow all federal, state, and local laws and Company policies. Criminal history may have a direct, adverse, and negative relationship with some of the material job duties of this position. These include the duties and responsibilities listed above, as well as the abilities to adhere to company policies, exercise sound judgment, effectively manage stress and work safely and respectfully with others, exhibit trustworthiness and professionalism, and safeguard business operations and the Company’s reputation. Pursuant to the Los Angeles County Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records.

Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit https://amazon.jobs/content/en/how-we-hire/accommodations for more information. If the country/region you’re applying in isn’t listed, please contact your Recruiting Partner.

The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at https://amazon.jobs/en/benefits.

USA, CA, Cupertino - 127,100.00 - 185,000.00 USD annually

Company - Annapurna Labs (U.S.) Inc.

Job ID: A10490741

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Sr. Software Development Engineer, HPC/ML Networking Engineer, Annapurna Labs
Sr. Software Development Engineer, HPC/ML Networking Engineer, Annapurna Labs

Socket.dev • Cupertino (CA)

On-site
USD 193,300 - 261,500
Lead Software Engineer, ML Network Stack - Annapurna Labs (AWS)
Lead Software Engineer, ML Network Stack - Annapurna Labs (AWS)

Amazon • Cupertino (CA)

On-site
USD 193,000 - 262,000
Health insurance
RSUs and sign-on options
401(k) matching
+2
Sr. Software Development Engineer, HPC/ML Networking Engineer, Annapurna Labs
Sr. Software Development Engineer, HPC/ML Networking Engineer, Annapurna Labs

Amazon • Cupertino (CA)

On-site
USD 150,000 - 210,000
Sr. Software Development Engineer, HPC/ML Networking Engineer, Annapurna Labs (AWS)
Sr. Software Development Engineer, HPC/ML Networking Engineer, Annapurna Labs (AWS)

Amazon • Cupertino (CA)

On-site
USD 180,000 - 230,000
Health insurance
401(k) matching
Paid time off
+2
Sr. Manager, Software Development, ML Network Stack - Annapurna Labs
Sr. Manager, Software Development, ML Network Stack - Annapurna Labs

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 220,000 - 298,000
Health insurance
401(k) matching
Paid time off
+1
Sr. Manager, Software Development, ML Network Stack - Annapurna Labs
Sr. Manager, Software Development, ML Network Stack - Annapurna Labs

Amazon • Seattle (WA)

On-site
USD 220,000 - 298,000
Health insurance
401(k) matching
Paid time off
+1
Manager, Software Development, ML Network Stack - Annapurna Labs
Manager, Software Development, ML Network Stack - Annapurna Labs

Socket.dev • Seattle (WA)

On-site
USD 185,000 - 250,000
Software Engineer II, Annapurna Labs ML Acceleration Systems Software
Software Engineer II, Annapurna Labs ML Acceleration Systems Software

Amazon Web Services (AWS) • Austin (TX)

On-site
USD 144,000 - 194,000
Software Development Engineer I, ML Infra Services, Annapurna Labs
Software Development Engineer I, ML Infra Services, Annapurna Labs

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 127,000 - 185,000
Software Engineer II, Annapurna Labs ML Acceleration System Software
Software Engineer II, Annapurna Labs ML Acceleration System Software

Amazon Web Services (AWS) • Austin (TX)

On-site
USD 144,000 - 194,000