Senior AI Accelerator & GPU Fleet Architect

Amazon Web Services (AWS)

Cupertino (CA)

On-site

USD 183,000 - 248,000

Full time

2 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Health insurance
401(k) matching
Paid time off
Parental leave

Job summary

Amazon Data Services, Inc. is seeking a Senior Hardware Development Engineer to own the GPU and accelerator lifecycle from definition to field operations. You will collaborate with vendor teams to shape hardware, firmware and diagnostic requirements and translate fleet-scale failure data into actionable improvements.

You will lead root-cause analyses, set health metrics, and present GPU lifecycle status to senior leadership while mentoring engineers in failure analysis methodologies.

Qualifications

  • Bachelor's degree in electrical engineering, computer engineering, or equivalent.
  • Experience in developing functional specifications, design verification plans and functional test procedures.
  • 7+ years of hardware design and development experience for server, compute, or large-scale infrastructure platforms.

Responsibilities

  • Own GPU component lifecycle strategy across AI accelerator platforms, including hardware, firmware and diagnostics requirements.
  • Lead fleet-scale failure analysis using telemetry, event logs, and vendor diagnostics.
  • Define GPU fleet health metrics, thresholds and dashboards for leadership review.
  • Engage with vendors and cross-team engineering to translate failure data into action items and roadmaps.

Skills

Data analysis
Vendor engagement
Executive communication
Cross-team leadership
GPU/AI accelerator knowledge

Education

Bachelors in electrical/computer engineering

Tools

Firmware lifecycle tools
Telemetry / diagnostics tooling

Job description

Amazon Data Services, Inc. is seeking a Senior Hardware Development Engineer to own the GPU and accelerator lifecycle from definition to field operations. You will collaborate with vendor teams to shape hardware, firmware and diagnostic requirements and translate fleet-scale failure data into actionable improvements.

You will lead root-cause analyses, set health metrics, and present GPU lifecycle status to senior leadership while mentoring engineers in failure analysis methodologies.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior GPU Fleet Engineer — Hardware & Failure Analysis
Senior GPU Fleet Engineer — Hardware & Failure Analysis

Amazon Web Services (AWS) • Denver (CO)

On-site
USD 159,000 - 215,000
Senior GPU Fleet Engineer – AI Accelerators
Senior GPU Fleet Engineer – AI Accelerators

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 159,000 - 215,000
Senior Hardware Engineering Lead, Accelerated Servers
Senior Hardware Engineering Lead, Accelerated Servers

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 208,000 - 282,000
Health insurance
401(k) matching
Paid time off
+2
Senior Hardware Engineering Lead - Accelerated GPU
Senior Hardware Engineering Lead - Accelerated GPU

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 240,000 - 324,000
Health insurance
RSUs
401(k) matching
+2
Senior GPU Cloud Hardware Architect for AI Training
Senior GPU Cloud Hardware Architect for AI Training

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 159,000 - 215,000
Health insurance
401(k) matching
Paid time off
+1
Senior Hardware Development Manager, AWS Accelerator Servers
Senior Hardware Development Manager, AWS Accelerator Servers

Amazon • Nashville (TN)

On-site
USD 120,000 - 150,000
Senior Accelerator Engineer, Cloud AI/ML server team
Senior Accelerator Engineer, Cloud AI/ML server team

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 183,000 - 248,000
Health insurance
401(k) matching
Paid time off
+1
Senior Accelerator Engineer, Cloud AI/ML server team
Senior Accelerator Engineer, Cloud AI/ML server team

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 159,000 - 215,000
Senior AI/ML GPU Server Manufacturing Engineer
Senior AI/ML GPU Server Manufacturing Engineer

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 183,000 - 248,000
Health insurance
401(k) matching
Paid time off
+1
Senior Manufacturing Hardware Engineer — GPU AI Accelerator
Senior Manufacturing Hardware Engineer — GPU AI Accelerator

Amazon Web Services (AWS) • Denver (CO)

On-site
USD 159,000 - 215,000