Lead GPU-Compute Server Architect for AI Scale

Amazon Web Services (AWS)

Austin (TX)

On-site

USD 159,000 - 215,000

Full time

7 hours ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Amazon is seeking a Cloud Hardware Development Engineer to define server architectures for GPU-accelerated AI training platforms and drive validation from PCBA bring-up through fleet deployment. You will partner with ODM/JDM teams, own hardware debug across EVT/DVT/PVT, and maintain fleet quality metrics post-launch.

You will lead cross-functional teams across thermal, mechanical, power delivery, and signal integrity domains, ensuring scalable, debuggable, and producible designs for large-scale

Qualifications

  • Bachelor's degree in electrical engineering, computer engineering, or equivalent.
  • Experience in developing functional specifications, design verification plans and functional test procedures.
  • 7+ years of hardware design and development experience for server, compute, or large-scale infrastructure platforms.
  • Experience in one or more server technologies: thermal/mechanical design, power delivery, high-speed signal integrity, or accelerator subsystems.
  • Experience leading hardware development through full product lifecycle (concept through production ramp).
  • Experience working with ODMs through the product development and manufacturing lifecycle (EVT, DVT, PVT).
  • In-depth expertise in high-speed bus design, signal integrity, or power delivery for GPU/accelerator platforms.
  • Experience owning fleet quality metrics and driving design improvements based on operational failure data.
  • Experience with thermal/mechanical design for high-power-density compute platforms (liquid cooling, air cooling, or hybrid).
  • Track record of defining engineering standards and design best practices adopted across teams or partner organizations.

Responsibilities

  • Define server architectures based on workload demand, translating them into detailed designs and component specifications that enable high-performance AI training and inference at scale.
  • Work with interdisciplinary teams of component, firmware, test, qualification, and integration engineers to deliver cohesive designs.
  • Drive design reviews with ODM/JDM partners covering schematic, layout, BOM, and manufacturing DFx (Design for Test, Design for Manufacturing).
  • Define and execute validation strategies from PCBA bring-up through server and rack integration — covering power sequencing, signal integrity, thermal characterization, and accelerator interconnect performance.
  • Own hardware debug during EVT/DVT/PVT builds, correlating failures across PCIe, power rails, memory channels, and GPU subsystems.
  • Triage hardware issues at both ODM facilities and datacenters, conduct root cause analysis, and implement corrective actions.
  • Own fleet quality metrics post-launch: server-level annualized failure rates and component-level failure modes.
  • Monitor operational telemetry to identify systemic issues and drive design or process changes for current and future platforms.
  • Partner with test and automation teams to improve manufacturing yield and reduce test dwell times.
  • Work with EC2 architecture teams to align on instance definitions, workload requirements, and platform trade-offs.
  • Drive ODM/JDM design partners through development milestones and production ramp.
  • Collaborate with firmware, software, and operations teams to ensure designs are debuggable, serviceable, and automation-ready.

Skills

Hardware design
Server technologies
Power delivery
Signal integrity
Thermal design
ODM collaboration
Fleet quality data
HW bring-up
Debug & root cause
Design for test/manufacturing

Education

Bachelor's degree in electrical engineering or computer engineering
Master's degree in electrical engineering or related field

Job description

Amazon is seeking a Cloud Hardware Development Engineer to define server architectures for GPU-accelerated AI training platforms and drive validation from PCBA bring-up through fleet deployment. You will partner with ODM/JDM teams, own hardware debug across EVT/DVT/PVT, and maintain fleet quality metrics post-launch.

You will lead cross-functional teams across thermal, mechanical, power delivery, and signal integrity domains, ensuring scalable, debuggable, and producible designs for large-scale

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior GPU Cloud Hardware Architect for AI Training
Senior GPU Cloud Hardware Architect for AI Training

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 159,000 - 215,000
Health insurance
401(k) matching
Paid time off
+1
Cloud AI Hardware Architect — GPU Server Systems
Cloud AI Hardware Architect — GPU Server Systems

Amazon Web Services (AWS) • Denver (CO)

On-site
USD 159,000 - 215,000
Cloud AI Hardware Architect
Cloud AI Hardware Architect

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 159,000 - 215,000
Health insurance
RSU & sign-on bonuses
Parental leave
GPU Server Hardware Engineer for AI Systems
GPU Server Hardware Engineer for AI Systems

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 110,000 - 160,000
Senior AI Accelerator & GPU Fleet Architect
Senior AI Accelerator & GPU Fleet Architect

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 183,000 - 248,000
Health insurance
401(k) matching
Paid time off
+1
Hardware Development Engineer for AI/ML GPU Servers
Hardware Development Engineer for AI/ML GPU Servers

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 126,000 - 185,000
Senior Cloud Hardware Architect — AI/ML Server Systems
Senior Cloud Hardware Architect — AI/ML Server Systems

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 183,000 - 248,000
Health insurance
401(k) matching
Paid time off
+2
GPU Server Hardware Engineer for AI/ML
GPU Server Hardware Engineer for AI/ML

Amazon • Cupertino (CA)

On-site
USD 126,000 - 185,000
RSUs
Health benefits
401(k) matching
Lead AI Accelerator Hardware Engineer - Cloud-Scale
Lead AI Accelerator Hardware Engineer - Cloud-Scale

Amazon • Austin (TX)

On-site
USD 159,000 - 215,000
Health insurance
401(k) matching
Paid time off
+1
AI Infra Systems Engineer: GPU & Accelerator Servers
AI Infra Systems Engineer: GPU & Accelerator Servers

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 129,000 - 175,000
Health insurance
401(k) matching
Paid time off
+1