AI Compute Fleet Engineer – GPU & Server Automation

Amazon Web Services (AWS)

Cupertino (CA)

On-site

USD 149,000 - 201,000

Full time

6 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Health insurance
401(k) matching
Paid time off

Job summary

Amazon Development Center U.S., Inc. in Cupertino is seeking a Systems Development Engineer to own the health and development of server platforms at worldwide fleet scale.

You will develop automation, analyze hardware telemetry across tens of thousands of hosts, and build tooling that directly determines whether capacity is available to customers. This role spans the full stack—from hardware monitoring interfaces, to health diagnostics up through fleet-wide data pipelines and dashboards—working

Qualifications

  • 2+ years of non-internship professional software development experience.
  • 1+ years of designing or architecting scalable systems.
  • 3+ years of administrative experience in networking, storage systems, operating systems and hands-on systems engineering.
  • Knowledge of systems engineering fundamentals.
  • Experience programming with at least one modern language such as C++, C#, Java, Python, Golang, PowerShell, Ruby

Responsibilities

  • Analyze hardware failure patterns using fleet telemetry and logs to identify root causes and quantify customer impact.
  • Contribute to predictive failure detection using sensor data, error trending, and log correlation.
  • Build and maintain operational dashboards and metrics for platform fleet health.
  • Develop automation for hardware test, firmware qualification, and capacity recovery workflows.

Skills

Software development
System design
Systems engineering
Networking/storage/OS basics
Programming languages (C++, C#, Java,/

Job description

Amazon Development Center U.S., Inc. in Cupertino is seeking a Systems Development Engineer to own the health and development of server platforms at worldwide fleet scale.

You will develop automation, analyze hardware telemetry across tens of thousands of hosts, and build tooling that directly determines whether capacity is available to customers. This role spans the full stack—from hardware monitoring interfaces, to health diagnostics up through fleet-wide data pipelines and dashboards—working

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Infra Systems Engineer: GPU & Accelerator Servers
AI Infra Systems Engineer: GPU & Accelerator Servers

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 129,000 - 175,000
Health insurance
401(k) matching
Paid time off
+1
Systems Engineer, GPU & AI Accelerator Servers
Systems Engineer, GPU & AI Accelerator Servers

Amazon • Cupertino (CA)

On-site
USD 149,000 - 201,000
Health insurance (medical, dental, and
Stock options or RSUs
401(k) matching
Fleet-Scale Systems Engineer for GPU & AI Accelerators
Fleet-Scale Systems Engineer for GPU & AI Accelerators

Amazon • Denver (CO)

On-site
USD 129,000 - 175,000
Health insurance
401(k) matching
Paid time off
+1
GPU Server Hardware Engineer for AI/ML
GPU Server Hardware Engineer for AI/ML

Amazon • Cupertino (CA)

On-site
USD 126,000 - 185,000
RSUs
Health benefits
401(k) matching
Senior AI/ML Fleet Automation Engineer
Senior AI/ML Fleet Automation Engineer

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 174,000 - 235,000
401(k) matching
Paid time off
Parental leave
+1
Cloud AI Systems Engineer - Fleet Health & Automation
Cloud AI Systems Engineer - Fleet Health & Automation

Amazon Web Services (AWS) • Austin (TX)

On-site
USD 129,000 - 175,000
Health insurance
RSUs (restricted stock units)
401(k) matching
+1
Senior Cloud Hardware Architect — AI/ML Server Systems
Senior Cloud Hardware Architect — AI/ML Server Systems

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 183,000 - 248,000
Health insurance
401(k) matching
Paid time off
+2
Cloud Hardware Engineer — AI/ML Server Fleet
Cloud Hardware Engineer — AI/ML Server Fleet

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 159,000 - 215,000
Cloud AI Hardware Architect — GPU Server Systems
Cloud AI Hardware Architect — GPU Server Systems

Amazon Web Services (AWS) • Denver (CO)

On-site
USD 159,000 - 215,000
AI/ML Server Hardware Architect in High-Performance Systems
AI/ML Server Hardware Architect in High-Performance Systems

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 183,000 - 248,000
Health insurance
401(k) matching
RSU equity