ML Hardware Systems Engineer – Fleet Automation & Debugging

Amazon

Austin (TX)

On-site

USD 136,000 - 184,000

Full time

3 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Health insurance
401(k) matching
Paid time off
RSUs

Job summary

Annapurna Labs ( Amazon ) is seeking an AI Hardware Systems Engineer to own ML server platforms within a fleet, focusing on health, reliability, and customer experience. The role involves debugging hardware-software interactions, developing scalable automation, and leading end-to-end testing across the product lifecycle.

Your work will drive autonomous software for monitoring, remediation, and optimization of cutting-edge ML accelerators and servers used worldwide in AWS.

Qualifications

  • 2+ years of non-internship software development experience.
  • 1+ years of designing or architecting (design patterns, reliability and scaling) of new and existing systems experience.
  • 1+ years of administrative experience in networking, storage systems, operating systems and hands-on systems engineering experience.
  • Knowledge of systems engineering fundamentals (networking, storage, operating systems).
  • Experience programming with at least one modern language such as C++, C#, Java, Python, Golang, PowerShell, Ruby.
  • Experience with Linux/Unix.
  • Bachelor's degree in Computer Science, Computer Engineering, or Electrical Engineering.

Responsibilities

  • Member of a team responsible for system remediation, operational excellence, and customer experience on bleeding edge ML products.
  • Utilize data to root cause hardware failures and identify live trends on the most complex systems in AWS.
  • Implement and improve system level testing across the product lifecycle.
  • Develop software which can be maintained, improved upon, documented, tested, and reused.
  • Dive deep on issues at the intersection of hardware and software.

Skills

Software development
Systems design
Networking basics
Linux/Unix
Programming languages
Debugging

Education

Bachelor's degree in CS/CE/EE
Master's degree (preferred)

Tools

C++
Python
Java
Go
PowerShell

Job description

Annapurna Labs ( Amazon ) is seeking an AI Hardware Systems Engineer to own ML server platforms within a fleet, focusing on health, reliability, and customer experience. The role involves debugging hardware-software interactions, developing scalable automation, and leading end-to-end testing across the product lifecycle.

Your work will drive autonomous software for monitoring, remediation, and optimization of cutting-edge ML accelerators and servers used worldwide in AWS.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Hardware Systems Engineer: Fleet Automation
ML Hardware Systems Engineer: Fleet Automation

JobCubby • Austin (TX), Northern (KY)

Hybrid
USD 136,000 - 184,000
Health insurance
401(k) matching
RSUs
+1
ML Hardware Platform Engineer
ML Hardware Platform Engineer

Socket.dev • Austin (TX)

On-site
USD 136,000 - 184,000
Machine Learning Hardware Platform Engineer
Machine Learning Hardware Platform Engineer

Amazon Web Services (AWS) • Austin (TX)

On-site
USD 136,000 - 184,000
Health insurance
RSUs / restricted stock units
401(k) match
+2
AI Hardware Systems Engineer, Annapurna Labs, Trainium Machine Learning Fleet Operations
AI Hardware Systems Engineer, Annapurna Labs, Trainium Machine Learning Fleet Operations

Amazon Web Services (AWS) • Austin (TX)

On-site
USD 136,000 - 184,000
Health insurance
RSUs / restricted stock units
401(k) match
+2
AI Hardware Systems Engineer, Annapurna Labs, Trainium Machine Learning Fleet Operations
AI Hardware Systems Engineer, Annapurna Labs, Trainium Machine Learning Fleet Operations

Annapurna Labs (U.S.) Inc. • Austin (TX)

On-site
USD 120,000 - 180,000
Health insurance
RSUs
AI Hardware Systems Engineer, Annapurna Labs, Trainium Machine Learning Fleet Operations
AI Hardware Systems Engineer, Annapurna Labs, Trainium Machine Learning Fleet Operations

Socket.dev • Austin (TX)

On-site
USD 136,000 - 184,000
Senior Systems Dev Engineer — AI/ML Fleet Health
Senior Systems Dev Engineer — AI/ML Fleet Health

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 151,000 - 205,000
ML Hardware Fleet Operations Manager
ML Hardware Fleet Operations Manager

Amazon Web Services (AWS) • Austin (TX)

On-site
USD 175,000 - 237,000
Health insurance
401(k) matching
Paid time off
+1
Lead, ML Hardware Fleet Operations & Automation
Lead, ML Hardware Fleet Operations & Automation

Amazon • Austin (TX)

On-site
USD 175,000 - 237,000
ML Platform Engineer: Scale & Debug GPU Server Fleet
ML Platform Engineer: Scale & Debug GPU Server Fleet

Annapurna Labs (U.S.) Inc. • Austin (TX)

On-site
USD 120,000 - 180,000
Health insurance
RSUs