Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.
Amazon in Austin, Texas is seeking an AI Hardware Systems Manager to lead the ML fleet operations team. You will mentor engineers and drive operational excellence by establishing metrics to maximize platform health and customer experience.
The ideal candidate holds a Bachelor's in Computer Science or Electrical Engineering and has over 7 years of experience in systems engineering or hardware operations. Benefits include a comprehensive package with health insurance and stock units.
Annapurna Labs designs silicon and software that accelerates innovation. Customers choose us to create cloud solutions that solve challenges that were unimaginable a short time ago, even yesterday. Our custom chips, accelerators, and software stacks enable us to take on technical challenges that have never been seen before, and deliver results that help our customers change the world.
As a Manager on the MLA Fleet Operations team, you set the direction for how your team keeps the world's most advanced ML accelerators healthy at scale. You start each day with your people — holding 1:1s, coaching engineers through ambiguous technical problems, removing blockers, and ensuring the team is focused on the highest‑impact work. From there, you review fleet health with the team, understanding which issues are trending, which investigations need unblocking, and where to allocate engineering effort for maximum customer impact. You partner with hardware design teams to advocate for fleet‑informed design changes and with service teams to align on deployment schedules. You balance long‑term automation investments against near‑term operational demands, and you represent your team's work to senior leadership with clear data and crisp narratives. When critical incidents arise, you lead the response — marshaling the right people, driving root cause, and ensuring corrective actions land.
The MLA Fleet Operations team was formed to maintain an exceptionally high quality bar for our fleet of advanced machine learning accelerators and server products. We perfect the customer experience by developing scalable software for rapid incident response times and data visualization as well as diving deep into hardware issues as they arise.
Salary Range: 175,100.00-236,900.00 USD annually. Your Amazon package will include sign‑on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and optional Supplemental life plans), EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage, 401(k) matching, paid time off, and parental leave.
Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status.