Senior AI/ML Systems Automation Engineer

Amazon Web Services (AWS)

Cupertino (CA)

On-site

USD 174,000 - 235,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health insurance
401(k) matching
Paid time off
Parental leave

Job summary

Amazon Web Services, Inc. seeks an experienced engineer to own end-to-end health of AI/ML accelerator fleets. You will build automation that prevents failures, correlate telemetry, and drive proactive diagnostics across PCIe, Linux drivers, and GPU subsystems.

You will design predictive failure detection, develop monitoring dashboards, and work across hardware, firmware, and software teams to improve fleet reliability at scale.

Qualifications

  • Bachelor's degree in CS/CE/EE or equivalent work experience.
  • 6+ years of professional software development experience.
  • 6+ years of designing or architecting (design patterns, reliability and scaling) of new and existing systems experience.

Responsibilities

  • Build and own the automation infrastructure for accelerator fleet health at a large scale of servers, driving toward zero-touch operations that detect, diagnose, triage, and remediate faults without human intervention.
  • Design and develop test frameworks, test coverage strategies, and diagnostic tooling to validate hardware functionality, detect faults, and ensure qualification coverage across the platform lifecycle.
  • Design predictive failure detection using telemetry, sensor data, error trending, and log correlation to identify degrading components before customer impact.
  • Develop monitoring dashboards and alerting for real-time fleet health visibility across manufacturing, lab, and production environments.

Skills

C++
Python
Linux
System design
Kernel debugging

Education

Bachelor's degree in CS/CE/EE or equivalent

Tools

Linux drivers
PCIe topology
GPU debugging

Job description

Amazon Web Services, Inc. seeks an experienced engineer to own end-to-end health of AI/ML accelerator fleets. You will build automation that prevents failures, correlate telemetry, and drive proactive diagnostics across PCIe, Linux drivers, and GPU subsystems.

You will design predictive failure detection, develop monitoring dashboards, and work across hardware, firmware, and software teams to improve fleet reliability at scale.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Systems Dev Engineer, AI/ML Fleet Automation
Senior Systems Dev Engineer, AI/ML Fleet Automation

Amazon • Seattle (WA)

On-site
USD 151,000 - 205,000
Senior AI Hardware Test & Automation Engineer
Senior AI Hardware Test & Automation Engineer

Amazon Web Services (AWS) • Salt Lake City (UT)

On-site
USD 136,000 - 184,000
Health insurance
401(k) matching
Paid time off
+1
Senior Edge AI/ML Accelerator Systems Engineer
Senior Edge AI/ML Accelerator Systems Engineer

Amazon Web Services (AWS) • Austin (TX)

On-site
USD 151,200 - 204,600
Senior Systems Engineer — AI/ML Accelerator Servers
Senior Systems Engineer — AI/ML Accelerator Servers

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 151,000 - 205,000
Health insurance
401(k) matching
Parental leave
ML Hardware Systems Engineer – Fleet Platform Automation
ML Hardware Systems Engineer – Fleet Platform Automation

Amazon • Austin (TX)

On-site
USD 136,000 - 184,000
Health insurance
401(k) matching
Paid time off
+1
Senior Edge & AI/ML Server Systems Engineer
Senior Edge & AI/ML Server Systems Engineer

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 173,000 - 236,000
Health insurance
401(k) matching
Paid time off
+1
Systems Engineer, AI-Powered Infrastructure Monitoring
Systems Engineer, AI-Powered Infrastructure Monitoring

Amazon • Nashville (TN)

On-site
USD 123,000 - 166,000
Health insurance
401(k) matching
Paid time off
+1
Senior AI & ML Server Hardware Engineer
Senior AI & ML Server Hardware Engineer

Amazon Web Services (AWS) • Austin (TX)

On-site
USD 159,000 - 216,000
Sr Systems Development Engineer, AWS Hardware Engineering Services, AI UltraServers
Sr Systems Development Engineer, AWS Hardware Engineering Services, AI UltraServers

Amazon Web Services (AWS) • Austin (TX)

On-site
USD 151,200 - 204,600
Sr Systems Development Engineer, AWS Hardware Engineering Services, AI UltraServers
Sr Systems Development Engineer, AWS Hardware Engineering Services, AI UltraServers

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 174,000 - 235,000
Health insurance
401(k) matching
Paid time off
+1