Lead ML Infrastructure for Edge AI Platform

Amazon Inc.

Bellevue (WA)

On-site

USD 185,000 - 250,000

Full time

21 hours ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Health insurance
401(k) matching
Paid time off
Parental leave
RSU/stock options

Job summary

Amazon.com Services LLC is seeking a Software Development Manager for the Edge AI ML Platform and Infrastructure to lead a distributed training team across multi-node GPU clusters. You will own capacity planning, CI/CD, and observability, guiding engineers and scientists to ship models at scale, with a strong emphasis on reliability and performance.

You will recruit, mentor, and develop leaders, shape architecture decisions, and translate complex scientific needs into a practical, prioritized

Qualifications

  • 3+ years of engineering team management experience.
  • 7+ years of working directly within engineering teams experience.
  • 3+ years of designing or architecting (design patterns, reliability and scaling) of new and existing systems experience.
  • Knowledge of engineering practices and patterns for the full software/hardware/networks development life cycle, including coding standards, code reviews, source control management, build processes, testing, certification, and livesite operations.
  • Experience partnering with product or program management teams.
  • Experience managing a team of high calibre Software Engineers developing complex, world class, scalable software systems that have been successfully delivered to customers.

Responsibilities

  • Build, lead, and grow a team of software and ML infrastructure engineers: recruit and hire, set clear goals, coach for growth, and manage performance across the team.
  • Own the roadmap for ML infrastructure—distributed training, GPU capacity, workflow orchestration, CI/CD, and observability—balancing near-term deliveries with long-term platform health.
  • Drive the architecture of distributed training capabilities (data, tensor, pipeline, and model parallelism) for large language and multimodal models, partnering with senior engineers and applied scientists.
  • Establish operational excellence for production platform services, including metrics, alarms, runbooks, on‑call processes, and root‑cause correction of recurring issues, while owning GPU fleet efficiency, capacity planning, and cost optimization.
  • Partner with applied science, compiler, runtime, hardware, security, and product teams to align requirements, manage dependencies, and deliver cross‑team programs.

Skills

Team management
Software engineering
Cross-functional collaboration
Product/PM partnership
Site reliability and ops
Leadership coaching

Tools

MXNet
TensorFlow
Caffe
PyTorch

Job description

Amazon.com Services LLC is seeking a Software Development Manager for the Edge AI ML Platform and Infrastructure to lead a distributed training team across multi-node GPU clusters. You will own capacity planning, CI/CD, and observability, guiding engineers and scientists to ship models at scale, with a strong emphasis on reliability and performance.

You will recruit, mentor, and develop leaders, shape architecture decisions, and translate complex scientific needs into a practical, prioritized

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Lead ML Infrastructure Manager – Edge AI Platform
Lead ML Infrastructure Manager – Edge AI Platform

Amazon Inc. • Factoria (WA)

On-site
USD 185,000 - 250,000
Health insurance
401(k) matching
Paid time off
+1
Senior Manager, ML Compute Platform
Senior Manager, ML Compute Platform

Amazon • Seattle (WA)

On-site
USD 220,000 - 298,000
Health insurance
RSUs
401(k) matching
+1
Senior Manager, AI Compute Platform & ML Systems
Senior Manager, AI Compute Platform & ML Systems

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 220,000 - 298,000
Health insurance
401(k) matching
Software Dev Mgr, ML Infrastructure, Edge AI Platform
Software Dev Mgr, ML Infrastructure, Edge AI Platform

Amazon Inc. • Bellevue (WA)

On-site
USD 185,000 - 250,000
Health insurance
401(k) matching
Paid time off
+2
Lead ML Network Stack Engineer for Scalable EC2 AI
Lead ML Network Stack Engineer for Scalable EC2 AI

Amazon • Cupertino (CA)

On-site
USD 193,000 - 262,000
Health insurance
RSUs and sign-on options
401(k) matching
+2
Senior ML Engineer: Scale AI Platforms & GenAI
Senior ML Engineer: Scale AI Platforms & GenAI

Amazon Inc. • Factoria (WA)

On-site
USD 168,000 - 227,000
Health insurance
401(k) matching
Paid time off
Senior ML Architect: Large-Scale AI Systems Lead
Senior ML Architect: Large-Scale AI Systems Lead

Amazon.com Services LLC • Bellevue (WA)

On-site
USD 350,000 - 650,000
Sign-on payments
Restricted stock units (RSUs)
Health insurance (medical, dental, and
+4
Senior ML Platform Architect - Scale AI Systems & MLOps
Senior ML Platform Architect - Scale AI Systems & MLOps

Amazon • Bellevue (WA)

On-site
USD 168,000 - 227,000
RSUs
401(k) matching
Paid time off
+2
Lead ML Platform Engineer for Scalable GPU Training
Lead ML Platform Engineer for Scalable GPU Training

Next Frontier Capital • Palo Alto (CA)

On-site
USD 180,000 - 245,000
Senior AI Infrastructure Engineer — Edge-Cloud ML + Equity
Senior AI Infrastructure Engineer — Edge-Cloud ML + Equity

Anduril Industries • Washington

On-site
USD 191,000 - 253,000