Senior Systems Dev Engineer, AI/ML Fleet Automation

Amazon

Seattle (WA)

On-site

USD 151,000 - 205,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Amazon is seeking a Sr Systems Development Engineer for the AWS Hardware Engineering Services AI UltraServers in Seattle. You will design automation, diagnostics, and fleet health infrastructure across hardware, firmware, and software to keep our AI accelerator fleet running at scale.

You will own end-to-end systems, build telemetry pipelines, and collaborate with ODMs and datacenter teams to prevent customer-impacting failures while enabling rapid deployment and reliability improvements.

Qualifications

  • 6+ years of programming in at least one modern language (C++, C#, Java, Python, Golang, PowerShell, Ruby).
  • 6+ years of non-internship professional software development experience.
  • 6+ years of designing or architecting (design patterns, reliability and scaling) of new and existing systems.
  • 5+ years of deploying and operating in a Linux/Unix environment.
  • 6+ years of systems design, software development, operations, automation, and process improvement.
  • Bachelor's degree in Computer Science, Computer Engineering, Electrical Engineering, or equivalent work experience.

Responsibilities

  • Build and own the automation infrastructure for accelerator (AI/ML) fleet health at a large scale of servers, driving toward zero-touch operations that detect, diagnose, triage, and remediate faults without human intervention.
  • Design and develop test frameworks, test coverage strategies, and diagnostic tooling to validate hardware functionality, detect faults, and ensure qualification coverage across the platform lifecycle.
  • Design predictive failure detection using telemetry, sensor data, error trending, and log correlation to identify degrading components before customer impact.
  • Develop monitoring dashboards and alerting for real-time fleet health visibility across manufacturing, lab, and production environments.
  • Define and track fleet health metrics: failure rates, mean time to detect and resolve issues, first-time fix rate, test dwell time, and predictive accuracy.
  • Debug complex system-level issues across compute, GPU, and networking in production — including Linux boot/runtime failures, PCIe, power, NIC, NVMe, and GPU subsystems on x86 and ARM.
  • Perform root cause analysis correlating across firmware, kernel, driver, and physical layer; feed findings into manufacturing quality and design improvements.
  • Design scalable test automation for hardware bring-up, regression, and qualification — reducing manufacturing test cycle times without sacrificing coverage through intelligent test sequencing and parallel execution.
  • Build data pipelines correlating test results, sensor telemetry, and component-level data to identify systemic yield issues and drive upstream fixes.
  • Develop and maintain Linux device drivers on ARM and x86; work with OS internals and accelerator/GPU software stacks.
  • Build and manage tests covering all functional aspects of the system and CI/CD pipelines for rapid deployment to manufacturing lines and production fleet.
  • Work across engineering teams and internal customers to ensure new accelerated compute hardware meets data path, control path, and onboarding requirements.
  • Engage with ODMs and design partners on testability, diagnostic coverage, and automation requirements during hardware design and bring-up phases — influencing functional and performance readiness of the platform.
  • Partner with datacenter operations to close the loop between field failures, manufacturing escapes, and design improvements.
  • Participate in post-incident reviews, identify contributing causes and drive permanent fixes that eliminate whole classes of risk.
  • Produce clear, maintainable documentation for systems, runbooks, and automation to enable others to operate and extend your work.
  • Drive process improvements that increase team agility — reducing development friction, eliminating unnecessary gates, and improving delivery velocity.

Skills

C++
C#
Java
Python
Golang
PowerShell
Ruby

Education

Bachelor's degree in Computer Science/Computer Engineering/Electrical Engineering

Tools

Linux
BMC/IPMI

Job description

Amazon is seeking a Sr Systems Development Engineer for the AWS Hardware Engineering Services AI UltraServers in Seattle. You will design automation, diagnostics, and fleet health infrastructure across hardware, firmware, and software to keep our AI accelerator fleet running at scale.

You will own end-to-end systems, build telemetry pipelines, and collaborate with ODMs and datacenter teams to prevent customer-impacting failures while enabling rapid deployment and reliability improvements.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI/ML Systems Automation Engineer
Senior AI/ML Systems Automation Engineer

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 174,000 - 235,000
Health insurance
401(k) matching
Paid time off
+1
Fleet Automation Systems Engineer - Hardware
Fleet Automation Systems Engineer - Hardware

Amazon Web Services (AWS) • Seattle (WA)

Hybrid
USD 129,000 - 175,000
Health insurance
401(k) matching
Paid time off
+1
Senior Edge & AI/ML Server Systems Engineer
Senior Edge & AI/ML Server Systems Engineer

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 173,000 - 236,000
Health insurance
401(k) matching
Paid time off
+1
Senior Systems Engineer — AI/ML Accelerator Servers
Senior Systems Engineer — AI/ML Accelerator Servers

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 151,000 - 205,000
Health insurance
401(k) matching
Parental leave
ML Hardware Systems Engineer – Fleet Platform Automation
ML Hardware Systems Engineer – Fleet Platform Automation

Amazon • Austin (TX)

On-site
USD 136,000 - 184,000
Health insurance
401(k) matching
Paid time off
+1
Senior SDE: SSD Fleet Automation & Reliability
Senior SDE: SSD Fleet Automation & Reliability

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 168,000 - 228,000
Senior AI/ML Server Hardware Architect
Senior AI/ML Server Hardware Architect

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 159,000 - 216,000
Senior Edge AI/ML Accelerator Systems Engineer
Senior Edge AI/ML Accelerator Systems Engineer

Amazon Web Services (AWS) • Austin (TX)

On-site
USD 151,200 - 204,600
Senior Systems Development Manager - AI Automation
Senior Systems Development Manager - AI Automation

Amazon • Town of Westboro (WI)

On-site
USD 198,000 - 268,000
Senior AI Hardware Test & Automation Engineer
Senior AI Hardware Test & Automation Engineer

Amazon Web Services (AWS) • Salt Lake City (UT)

On-site
USD 136,000 - 184,000
Health insurance
401(k) matching
Paid time off
+1