Lead, ML Hardware Fleet Operations & Automation

Amazon

Austin (TX)

On-site

USD 175,000 - 237,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Amazon Annapurna Labs' MLA Fleet Operations team seeks a Manager to lead 5–10 engineers across multiple ML server platforms, owning fleet health, automation, and data infrastructure at scale. You will define a 6–12 month roadmap, influence priorities, and drive operational excellence across cutting-edge ML hardware.

In this role you will mentor engineers, partner with hardware/software/product teams, and represent fleet operations in executive reviews with data-driven narratives.

Qualifications

  • Bachelor's degree in computer science, electrical engineering, or related field.
  • 2+ years of engineering team management experience.
  • Knowledge of and proficiency in Python scripting language.
  • Experience with general troubleshooting/debugging of hardware.
  • Experience designing, building, operating, and managing large-scale distributed systems or web services.
  • 7+ years of experience in systems engineering, platform engineering, SRE, or hardware operations.

Responsibilities

  • Build, hire, mentor, and grow a team of platform development engineers responsible for ML fleet operations across multiple accelerator platforms
  • Define team roadmap and technical strategy for fleet health, automation, and data infrastructure — balancing near-term operational demands against long-term engineering investments
  • Drive operational excellence by establishing metrics, SLAs, and processes that maximize platform sellability and customer experience
  • Partner with hardware engineering, software engineering, and product teams to prioritize debug efforts and translate fleet learnings into permanent design fixes
  • Own escalation paths for critical fleet incidents and lead cross-functional war rooms to resolution
  • Influence org-level priorities by surfacing fleet-wide patterns and advocating for systemic improvements across the ML hardware portfolio
  • Raise the bar on team software practices — ensuring automation is maintainable, tested, documented, and reusable at scale
  • Represent fleet operations in executive reviews, providing data-driven narratives on platform health and roadmap

Skills

Engineering management
Python scripting
Hardware debugging
Distributed systems
SRE experience

Education

Bachelor's degree in CS/EE or related field

Job description

Amazon Annapurna Labs' MLA Fleet Operations team seeks a Manager to lead 5–10 engineers across multiple ML server platforms, owning fleet health, automation, and data infrastructure at scale. You will define a 6–12 month roadmap, influence priorities, and drive operational excellence across cutting-edge ML hardware.

In this role you will mentor engineers, partner with hardware/software/product teams, and represent fleet operations in executive reviews with data-driven narratives.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Hardware Fleet Operations Manager
ML Hardware Fleet Operations Manager

Amazon Web Services (AWS) • Austin (TX)

On-site
USD 175,000 - 237,000
Health insurance
401(k) matching
Paid time off
+1
AI Hardware Systems Manager, Annapurna Labs, Trainium Machine Learning Fleet Operations
AI Hardware Systems Manager, Annapurna Labs, Trainium Machine Learning Fleet Operations

Amazon Web Services (AWS) • Austin (TX)

On-site
USD 175,000 - 237,000
Health insurance
401(k) matching
Paid time off
+1
ML Fleet Operations Manager — AI Hardware Platforms
ML Fleet Operations Manager — AI Hardware Platforms

Amazon • Austin (TX)

On-site
USD 175,100 - 236,900
Health insurance
401(k) matching
Sign-on payments
AI Hardware Systems Manager, Annapurna Labs, Trainium Machine Learning Fleet Operations
AI Hardware Systems Manager, Annapurna Labs, Trainium Machine Learning Fleet Operations

Amazon • Austin (TX)

On-site
USD 175,100 - 236,900
Health insurance
401(k) matching
Sign-on payments
Senior Manager, ML Compute Platform
Senior Manager, ML Compute Platform

Amazon • Seattle (WA)

On-site
USD 220,000 - 298,000
Health insurance
RSUs
401(k) matching
+1
Senior Systems Dev Engineer — AI/ML Fleet Health
Senior Systems Dev Engineer — AI/ML Fleet Health

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 151,000 - 205,000
Lead, Inference Infrastructure & Fleet
Lead, Inference Infrastructure & Fleet

Anthropic • San Francisco (CA)

Hybrid
USD 405,000 - 625,000
Competitive compensation
Equity donation matching
Generous vacation and parental leave
+2
ML Ops Platform Lead
ML Ops Platform Lead

Veho • United States

Remote
USD 160,000 - 230,000
Senior AI/ML Fleet Automation Engineer
Senior AI/ML Fleet Automation Engineer

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 174,000 - 235,000
401(k) matching
Paid time off
Parental leave
+1
Senior Manager, AI Compute Platform & ML Systems
Senior Manager, AI Compute Platform & ML Systems

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 220,000 - 298,000
Health insurance
401(k) matching