Software Dev Mgr, ML Infrastructure, Edge AI Platform

Amazon Inc.

Bellevue (WA)

On-site

USD 185,000 - 250,000

Full time

21 hours ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Health insurance
401(k) matching
Paid time off
Parental leave
RSU/stock options

Job summary

Amazon.com Services LLC is seeking a Software Development Manager for the Edge AI ML Platform and Infrastructure to lead a distributed training team across multi-node GPU clusters. You will own capacity planning, CI/CD, and observability, guiding engineers and scientists to ship models at scale, with a strong emphasis on reliability and performance.

You will recruit, mentor, and develop leaders, shape architecture decisions, and translate complex scientific needs into a practical, prioritized

Qualifications

  • 3+ years of engineering team management experience.
  • 7+ years of working directly within engineering teams experience.
  • 3+ years of designing or architecting (design patterns, reliability and scaling) of new and existing systems experience.
  • Knowledge of engineering practices and patterns for the full software/hardware/networks development life cycle, including coding standards, code reviews, source control management, build processes, testing, certification, and livesite operations.
  • Experience partnering with product or program management teams.
  • Experience managing a team of high calibre Software Engineers developing complex, world class, scalable software systems that have been successfully delivered to customers.

Responsibilities

  • Build, lead, and grow a team of software and ML infrastructure engineers: recruit and hire, set clear goals, coach for growth, and manage performance across the team.
  • Own the roadmap for ML infrastructure—distributed training, GPU capacity, workflow orchestration, CI/CD, and observability—balancing near-term deliveries with long-term platform health.
  • Drive the architecture of distributed training capabilities (data, tensor, pipeline, and model parallelism) for large language and multimodal models, partnering with senior engineers and applied scientists.
  • Establish operational excellence for production platform services, including metrics, alarms, runbooks, on‑call processes, and root‑cause correction of recurring issues, while owning GPU fleet efficiency, capacity planning, and cost optimization.
  • Partner with applied science, compiler, runtime, hardware, security, and product teams to align requirements, manage dependencies, and deliver cross‑team programs.

Skills

Team management
Software engineering
Cross-functional collaboration
Product/PM partnership
Site reliability and ops
Leadership coaching

Tools

MXNet
TensorFlow
Caffe
PyTorch

Job description

Software Dev Mgr, ML Infrastructure, Edge AI Platform

Job ID: 10568253 | Amazon.com Services LLC

Amazon Devices (Lab126) builds products and services that delight millions of customers globally. The Edge AI ML Platform and Infrastructure team is building the platform that enables Amazon teams to train, optimize, evaluate, and deploy generative AI models on devices and in the cloud.

Today, optimizing a large model for a new hardware target requires experts to connect model onboarding, distributed training, compression, evaluation, compilation, and deployment systems by hand. We are turning that work into a repeatable, self-service workflow. Our platform supports large language, vision, audio, multimodal, and mixture-of-experts models, and gives scientists and engineers the tools to move new optimization techniques from research code into reliable production workflows.

We are looking for a Software Development Manager to build and lead the ML infrastructure team behind this platform. You will own distributed training on multi-node GPU clusters, compute capacity and utilization, CI/CD, observability, and operational reliability for GPU-intensive workloads. You will hire and develop a team of software and ML infrastructure engineers, set its technical direction and roadmap, and deliver platform capabilities that scientists and product teams depend on to ship models with hundreds of billions of parameters.

This role combines people leadership with deep technical judgment. You will grow engineers and managers-in-the-making, drive architecture decisions with your senior engineers, turn ambiguous science and product needs into a prioritized plan, and hold a high bar for delivery and operational excellence.

Key job responsibilities
  • Build, lead, and grow a team of software and ML infrastructure engineers: recruit and hire, set clear goals, coach for growth, and manage performance across the team.
  • Own the roadmap for ML infrastructure—distributed training, GPU capacity, workflow orchestration, CI/CD, and observability—balancing near-term deliveries with long-term platform health.
  • Drive the architecture of distributed training capabilities (data, tensor, pipeline, and model parallelism) for large language and multimodal models, partnering with senior engineers and applied scientists.
  • Establish operational excellence for production platform services, including metrics, alarms, runbooks, on‑call processes, and root‑cause correction of recurring issues, while owning GPU fleet efficiency, capacity planning, and cost optimization.
  • Partner with applied science, compiler, runtime, hardware, security, and product teams to align requirements, manage dependencies, and deliver cross‑team programs.
A day in the life

You will move between people, planning, and technology. A typical day might include a 1:1 with an engineer on a growth plan, a design review for a new training-orchestration capability, triage of a failed multi-node training run, a capacity review against upcoming model deliveries, and a planning session with science leads on next quarter's priorities.

You will use performance, reliability, cost, and developer-productivity data to decide where the team invests. You will deliver incrementally while protecting long‑term architecture, and make sure the team fixes recurring problems at their root.

About the team

The Edge AI ML Platform and Infrastructure team brings together software engineers, ML infrastructure engineers, and GPU performance specialists. We build reusable model training, optimization, and deployment capabilities for Amazon product teams, working closely with applied scientists across Edge AI. Our customers need to adapt rapidly changing model architectures to constrained hardware and production workloads without rebuilding the toolchain for every model.

The team owns the platform foundations that connect model development to deployment. Because our scope runs end to end, we can improve training, compression, evaluation, and deployment as one system. We value clear interfaces, measurable performance, automated quality gates, and direct collaboration between science and engineering.

Basic Qualifications
  • 3+ years of engineering team management experience
  • 7+ years of working directly within engineering teams experience
  • 3+ years of designing or architecting (design patterns, reliability and scaling) of new and existing systems experience
  • Knowledge of engineering practices and patterns for the full software/hardware/networks development life cycle, including coding standards, code reviews, source control management, build processes, testing, certification, and livesite operations
  • Experience partnering with product or program management teams
  • Experience managing a team of high calibre Software Engineers developing complex, world class, scalable software systems that have been successfully delivered to customers
Preferred Qualifications
  • Experience in communicating with users, other technical teams, and senior leadership to collect requirements, describe software product features, technical designs, and product strategy
  • Experience in recruiting, hiring, mentoring/coaching and managing teams of Software Engineers to improve their skills, and make them more effective, product software engineers
  • 7+ years of full software development life cycle, including coding standards, code reviews, source control management, build processes, testing, and operations experience
  • Experience with Machine and Deep Learning toolkits such as MXNet, TensorFlow, Caffe and PyTorch
  • Experience building or operating distributed systems or high-performance computing systems
  • Experience leading teams that build distributed ML training, inference, evaluation, or data platforms using frameworks such as PyTorch, JAX, NeMo, or Megatron
  • Experience managing GPU cluster capacity, utilization, and cost at scale
  • Experience with model compression, quantization, knowledge distillation, model compilation, or edge deployment

Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status.

Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit https://amazon.jobs/content/en/how-we-hire/accommodations for more information. If the country/region you’re applying in isn’t listed, please contact your Recruiting Partner.

The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at https://amazon.jobs/en/benefits .

USA, WA, Bellevue - 184,900.00 - 250,200.00 USD annually

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Software Dev Mgr, ML Infrastructure, Edge AI Platform
Software Dev Mgr, ML Infrastructure, Edge AI Platform

Amazon Inc. • Factoria (WA)

On-site
USD 185,000 - 250,000
Health insurance
401(k) matching
Paid time off
+1
SDE II, ML Infra Services, Annapurna Labs
SDE II, ML Infra Services, Annapurna Labs

Amazon • Seattle (WA)

Hybrid
USD 144,000 - 194,000
Health insurance
401(k) matching
Paid time off
+1
Senior Manager, EC2 Nitro, EC2 Nitro
Senior Manager, EC2 Nitro, EC2 Nitro

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 220,000 - 298,000
Health insurance
401(k) matching
Software Development Engineer, SageMaker HyperPod Data Plane (AWS)
Software Development Engineer, SageMaker HyperPod Data Plane (AWS)

Amazon Inc. • Santa Clara (CA)

On-site
USD 165,000 - 224,000
Health insurance
401(k) matching
Paid time off
+1
Senior Manager, EC2 Nitro, EC2 Nitro (AWS)
Senior Manager, EC2 Nitro, EC2 Nitro (AWS)

Amazon • Seattle (WA)

On-site
USD 220,000 - 298,000
Health insurance
RSUs
401(k) matching
+1
Principal AI/ML HPC Specialist Technical Account Manager (STAM) , AWS Enterprise Support, NAMER-Sp
Principal AI/ML HPC Specialist Technical Account Manager (STAM) , AWS Enterprise Support, NAMER-Sp

Amazon Web Services (AWS) • Herndon (VA)

On-site
USD 183,000 - 247,000
Software Engineer II - AI/ML, Neuron Inference
Software Engineer II - AI/ML, Neuron Inference

Amazon • Cupertino (CA)

On-site
USD 165,000 - 224,000
Health insurance
401(k) matching
Paid time off
+1
Worldwide Specialist Solutions Architect - GenAI, Data & AI GTM
Worldwide Specialist Solutions Architect - GenAI, Data & AI GTM

Amazon Web Services (AWS) • Mountain View (CA)

On-site
USD 177,000 - 239,000
Health insurance
RSUs
Worldwide Specialist Solutions Architect - GenAI, Data & AI GTM
Worldwide Specialist Solutions Architect - GenAI, Data & AI GTM

Amazon Web Services (AWS) • Austin (TX)

On-site
USD 154,000 - 208,000
Software Development Engineer II, AWS SageMaker AI
Software Development Engineer II, AWS SageMaker AI

Amazon Web Services (AWS) • Bellevue (WA)

On-site
USD 144,000 - 194,000
RSUs
Health insurance
401(k) matching
+1