SDE II, ML Infra Services, Annapurna Labs

Amazon.com Services LLC

Seattle (WA)

On-site

USD 180,000 - 240,000

Full time

5 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Amazon.com Services LLC is seeking an engineer to lead the design and implementation of ML infrastructure platforms, enabling capacity management, workload scheduling, and fleet orchestration across ML accelerators.

You will collaborate with ML scientists, infrastructure engineers, hardware teams, and internal customers to ensure low wait times, high utilization, and zero-config deployment across environments.

Qualifications

  • 3+ years of non-internship professional software development experience.
  • 2+ years of non-internship design or architecture experience of new and existing systems.
  • Experience programming with at least one software programming language.

Responsibilities

  • Lead the design and implementation of ML infrastructure platform for capacity management, workload scheduling, and fleet orchestration.
  • Collaborate with ML scientists, training infrastructure engineers, hardware teams, and internal customers to deliver seamless ML Infra access.
  • Build high-impact solutions, participate in design discussions, code reviews, and communicate with stakeholders.
  • Work in a startup-like development environment, focusing on the most important tasks.

Skills

Go
Java
Python
JavaScript/TypeScript

Education

Bachelor's degree in computer science or equivalent

Tools

Kubernetes
Git

Job description

Annapurna Labs was a startup company acquired by AWS in 2015, and is now fully integrated. If AWS is an infrastructure company, then think Annapurna Labs as the infrastructure provider of AWS. Our org covers multiple disciplines including silicon engineering, hardware design and verification, software, and operations. AWS Nitro, ENA, EFA, Graviton and F1 EC2 Instances, AWS Neuron, Inferentia and Trainium ML Accelerators, and in storage with scalable NVMe, are some of the products we have delivered, over the last few years.

Key job responsibilities

This engineer will lead the design and implementation of ML infrastructure platform, building systems for capacity management, workload scheduling, and fleet orchestration across ML accelerators. They will work with ML scientists, training infrastructure engineers, hardware teams, and internal customers to ensure the ML Infra service delivers seamless ML Accelerator access with low wait times, high utilization, and zero-config deployment from various environments.

A day in the life

As you design and code solutions to help our team drive efficiencies in software architecture, you'll create metrics, implement automation and other improvements, and resolve the root cause of software defects. You'll also:

  • Build high-impact solutions to deliver to our large customer base.
  • Participate in design discussions, code review, and communicate with internal and external stakeholders.
  • Work cross-functionally to help drive business decisions with your technical input.
  • Work in a startup-like development environment, where you're always working on the most important stuff.
About the team
  • High-impact, high-visibility: You'll directly accelerate every Neuron team's ability to ship - your work multiplies the output of 100+ engineers
  • Greenfield opportunities: We're actively building new capabilities with significant design ownership for SDEs
  • Small, senior team: where every person owns major components and drives architectural decisions
  • AI infrastructure: Work at the intersection of Kubernetes, custom silicon, and large-scale ML workloads
Diverse Experiences

We value diverse experiences and non-traditional career paths. If your career is just starting or includes alternative experiences, we encourage you to apply.

Inclusive Team Culture

Our employee-led affinity groups foster inclusion. Events like CORE and AmazeCon inspire us to embrace our uniqueness.

Work/Life Balance

We strive for flexibility as part of our working culture, supporting you both at work and at home.

Mentorship & Career Growth

We offer knowledge-sharing, mentorship, and one-on-one code reviews to help you grow as a professional.

Basic Qualifications:
  • 3+ years of non-internship professional software development experience
  • 2+ years of non-internship design or architecture (design patterns, reliability and scaling) of new and existing systems experience
  • Experience programming with at least one software programming language
Preferred Qualifications:
  • 3+ years of full software development life cycle, including coding standards, code reviews, source control management, build processes, testing, and operations experience
  • Bachelor's degree in computer science or equivalent
  • 2+ years of building large-scale machine-learning infrastructure for online recommendation, ads ranking, personalization or search experience
  • Strong proficiency in Go/Java, Python and working knowledge Javascript/TypeScript
  • Experience building and operating large-scale distributed systems on Kubernetes
  • Experience designing, deploying, and maintaining production services at scale, including on-call ownership
  • Experience with machine learning infrastructure - orchestration, scheduling, or resource management at scale
  • Proficiency in application and kernel-level performance profiling and optimization
  • Experience with integrated software/hardware performance analysis in heterogeneous compute environments
  • Proficiency in observability and telemetry - instrumentation, metrics collection, alarming, dashboarding, and monitoring
  • Experience debugging complex issues in large-scale distributed systems and driving best practices

Amazon is an equal opportunity employer and does not discriminate on the basis

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

SDE II, ML Infra Services, Annapurna Labs
SDE II, ML Infra Services, Annapurna Labs

Amazon • Seattle (WA)

Hybrid
USD 144,000 - 194,000
Health insurance
401(k) matching
Paid time off
+1
SDE II, Neuron Infra Services
SDE II, Neuron Infra Services

Amazon • Cupertino (CA), Northern (KY)

Hybrid
USD 140,000 - 190,000
Software Development Engineer I, ML Infra Services, Annapurna Labs
Software Development Engineer I, ML Infra Services, Annapurna Labs

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 127,000 - 185,000
Sr. Solutions Architect, Annapurna ML
Sr. Solutions Architect, Annapurna ML

Amazon • Seattle (WA)

On-site
USD 190,000 - 230,000
SDE II, Neuron Infra Services
SDE II, Neuron Infra Services

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 165,000 - 224,000
Health insurance
401(k) matching
Paid time off
ML Kernel Performance Engineer, AWS Neuron, Annapurna Labs
ML Kernel Performance Engineer, AWS Neuron, Annapurna Labs

Annapurna Labs (U.S.) Inc. - D63 • Cupertino (CA)

On-site
USD 180,000 - 240,000
Sr. Software Development Engineer, Annapurna Labs
Sr. Software Development Engineer, Annapurna Labs

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 180,000 - 280,000
AI Hardware Systems Engineer, Annapurna Labs, Trainium Machine Learning Fleet Operations
AI Hardware Systems Engineer, Annapurna Labs, Trainium Machine Learning Fleet Operations

Amazon Web Services (AWS) • Austin (TX)

On-site
USD 136,000 - 184,000
Health insurance
RSUs / restricted stock units
401(k) match
+2
Lead Software Engineer, ML Network Stack - Annapurna Labs (AWS)
Lead Software Engineer, ML Network Stack - Annapurna Labs (AWS)

Amazon • Cupertino (CA)

On-site
USD 193,000 - 262,000
Health insurance
RSUs and sign-on options
401(k) matching
+2
Software Development Engineer, ML Systems, Annapurna Labs
Software Development Engineer, ML Systems, Annapurna Labs

Amazon • Northern (KY), New York (NY)

Hybrid
USD 158,000 - 214,000
Health insurance
RSUs
401(k) matching
+2