SDE II, ML Infra Services, Annapurna Labs (AWS)

Amazon

Seattle (WA)

On-site

USD 144,000 - 194,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health insurance
401(k) matching
Paid time off
Parental leave

Job summary

Amazon is hiring a Software Engineer to lead the development of ML infrastructure tools for running, optimizing, and analyzing ML workloads. You will guide ML tool projects from architecture through delivery, with deep profiling, resource management, scheduling, and code generation.

The role involves building scalable infra for ML accelerators and coordinating with ML scientists, engineers, and hardware teams. Hybrid onsite with remote options when needed.

Job description

Job ID: 10463313 | Amazon Development Center U.S., Inc.

Annapurna Labs was a startup company acquired by AWS in 2015, and is now fully integrated. If AWS is an infrastructure company, then think Annapurna Labs as the infrastructure provider of AWS. Our org covers multiple disciplines including silicon engineering, hardware design and verification, software, and operations. AWS Nitro, ENA, EFA, Graviton and F1 EC2 Instances, AWS Neuron, Inferentia and Trainium ML Accelerators, and in storage with scalable NVMe, are some of the products we have delivered, over the last few years.

AWS Neuron is the complete software stack for the AWS Inferentia and Trainium cloud-scale machine learning accelerators and the Trn1 and Inf1 servers that use them. This position is for a Software Engineer that will lead the development of machine learning tools to run, optimize, and analyze machine learning workloads. This candidate must have had experience leading machine learning tool projects, preferably starting from architecture through several generations of delivery to customers. Deep knowledge of profiling and optimization, resource management, scheduling, code generation are needed. The ideal candidate will have worked on new instruction set architectures, which may include CPU, NPU, GPU and other forms of compute.

Key job responsibilities

  • This engineer will lead the design and implementation of ML infrastructure platform, building systems for capacity management, workload scheduling, and fleet orchestration across ML accelerators.
  • They will work with ML scientists, training infrastructure engineers, hardware teams, and internal customers to ensure the ML Infra service delivers seamless ML Accelerator access with low wait times, high utilization, and zero-config deployment from various environments.

A day in the life

As you design and code solutions to help our team drive efficiencies in software architecture, you’ll create metrics, implement automation and other improvements, and resolve the root cause of software defects. You’ll also:

  • Build high-impact solutions to deliver to our large customer base.
  • Participate in design discussions, code review, and communicate with internal and external stakeholders.
  • Work cross-functionally to help drive business decisions with your technical input.
  • Work in a startup-like development environment, where you’re always working on the most important stuff.
About the team
  • High-impact, high-visibility: You'll directly accelerate every Neuron team's ability to ship — your work multiplies the output of 100+ engineers
  • Greenfield opportunities: We're actively building new capabilities with significant design ownership for SDEs
  • Small, senior team: where every person owns major components and drives architectural decisions
  • AI infrastructure: Work at the intersection of Kubernetes, custom silicon, and large-scale ML workloads
Diverse Experiences

We value diverse experiences and non-traditional career paths. If your career is just starting or includes alternative experiences, we encourage you to apply.

Inclusive Team Culture

Our employee-led affinity groups foster inclusion. Events like CORE and AmazeCon inspire us to embrace our uniqueness.

Work/Life Balance

We strive for flexibility as part of our working culture, supporting you both at work and at home.

Mentorship & Career Growth

We offer knowledge-sharing, mentorship, and one-on-one code reviews to help you grow as a professional.

Hybrid Work

This role is onsite, with flexibility to work remotely when you’re unable to make it into the office.

Basic Qualifications
  • 3+ years of non-internship professional software development experience
  • 2+ years of non-internship design or architecture (design patterns, reliability and scaling) of new and existing systems experience
  • Experience programming with at least one software programming language
Preferred Qualifications
  • 3+ years of full software development life cycle, including coding standards, code reviews, source control management, build processes, testing, and operations experience
  • Bachelor's degree in computer science or equivalent
  • Experience taking a leading role in building complex software or computing infrastructure that has been successfully delivered to customers
  • Experience with AWS Services including EC2, Lambda, S3, DynamoDB, SQS
  • Experience in Kubernetes, Docker or containers ecosystem, or experience managing full application stacks from the OS up through custom applications and experience in any Bigdata architecture
  • Experience with version control systems and CI/CD pipeline implementation
  • Strong proficiency in Go/Java, Python, and Javascript/Typescript
  • Application and kernel performance profiling and optimization
  • Proficiency in integrated software/hardware performance analysis and optimization
  • Experience designing and operating production services

Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status.

Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit https://amazon.jobs/content/en/how-we-hire/accommodations for more information. If the country/region you’re applying in isn’t listed, please contact your Recruiting Partner.

Preferred Qualifications
  • 3+ years of full software development life cycle, including coding standards, code reviews, source control management, build processes, testing, and operations experience
  • Bachelor's degree in computer science or equivalent
  • Experience taking a leading role in building complex software or computing infrastructure that has been successfully delivered to customers
  • Experience with AWS Services including EC2, Lambda, S3, DynamoDB, SQS
  • Experience in Kubernetes, Docker or containers ecosystem, or experience managing full application stacks from the OS up through custom applications and experience in any Bigdata architecture
  • Experience with version control systems and CI/CD pipeline implementation
  • Strong proficiency in Go/Java, Python, and Javascript/Typescript
  • Application and kernel performance profiling and optimization
  • Proficiency in integrated software/hardware performance analysis and optimization
  • Experience designing and operating production services

Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status.

Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit https://amazon.jobs/content/en/how-we-hire/accommodations for more information. If the country/region you’re applying in isn’t listed, please contact your Recruiting Partner.

The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at https://amazon.jobs/en/benefits .

USA, WA, Seattle - 143,700.00 - 194,400.00 USD annually

Important FAQs for current Government employees. Before proceeding, please review the following FAQs https://www.amazon.jobs/en/faqs#faqs-for-us-government-employees

Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

SDE II, Neuron Infra Services (AWS)
SDE II, Neuron Infra Services (AWS)

Amazon • Cupertino (CA)

On-site
USD 165,000 - 224,000
Health insurance
401(k) matching
Paid time off
+1
Software Development Engineer I, ML Infra Services, Annapurna Labs
Software Development Engineer I, ML Infra Services, Annapurna Labs

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 127,000 - 185,000
Software Development Engineer, ML Systems, Annapurna Labs (AWS)
Software Development Engineer, ML Systems, Annapurna Labs (AWS)

Amazon • New York (NY)

On-site
USD 158,000 - 214,000
Software Development Engineer III, Annapurna Labs
Software Development Engineer III, Annapurna Labs

Amazon • New York (NY)

On-site
USD 185,000 - 250,000
Health insurance
401(k) matching
Paid time off
+2
Lead Software Engineer, ML Network Stack - Annapurna Labs (AWS)
Lead Software Engineer, ML Network Stack - Annapurna Labs (AWS)

Amazon • Cupertino (CA)

On-site
USD 193,000 - 262,000
Health insurance
RSUs and sign-on options
401(k) matching
+2
Software Engineer - ML and Distributed Systems, Amazon Personalize
Software Engineer - ML and Distributed Systems, Amazon Personalize

Amazon • Mountain View (CA)

On-site
USD 165,200 - 223,600
Health insurance
401(k) matching
Paid time off
Software Engineering Manager, ML Kernel Performance, AWS Neuron, Annapurna Labs
Software Engineering Manager, ML Kernel Performance, AWS Neuron, Annapurna Labs

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 213,000 - 288,000
Health insurance
RSUs
401(k) matching
+2
Software Engineer- AI/ML, Amazon Neuron Training
Software Engineer- AI/ML, Amazon Neuron Training

Amazon • Cupertino (CA)

On-site
USD 165,000 - 224,000
Neuron Runtime Software Development Engineer , Neuron Runtime (AWS)
Neuron Runtime Software Development Engineer , Neuron Runtime (AWS)

Amazon • Cupertino (CA)

On-site
USD 165,000 - 224,000
Software Development Engineer, ML Systems, Annapurna Labs
Software Development Engineer, ML Systems, Annapurna Labs

Amazon Web Services (AWS) • New York (NY)

On-site
USD 158,000 - 214,000