Sr Cloud Hardware Dev Engineer, AWS Generative AI & ML Servers

Amazon Development Center U.S., Inc.

Cupertino (CA)

On-site

USD 183,000 - 248,000

Full time

5 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Health insurance
RSUs
401(k) matching
Paid time off

Job summary

Amazon Development Center U.S., Inc. is seeking a hardware architect to define server architectures for AI training workloads.

You will drive validation from silicon to fleet deployment, triage failures across PCIe, power, memory, and interconnects, and guide design improvements. This role involves cross-team collaboration with EC2 architecture, firmware, software, and operations to ensure debuggable, serviceable, and scalable designs.

Qualifications

  • Define server architectures based on workload demand and customer requirements.
  • Translate designs into detailed specifications enabling AI training at scale.
  • Collaborate with firmware, test, qualification, and integration engineers.

Responsibilities

  • Define server architectures based on workload demand and customer requirements.
  • Work with interdisciplinary teams to deliver cohesive designs.
  • Drive design reviews with ODM/JDM partners covering schematic, layout, BOM, and DFx.
  • Define and execute validation strategies from PCBA bring-up through server and rack integration.
  • Own hardware debug during EVT/DVT/PVT builds and correlate failures.
  • Triage hardware issues at ODM facilities and datacenters; implement corrective actions.
  • Own fleet quality metrics post-launch and drive design or process improvements.

Skills

Server architecture
Power delivery
Thermal design
Signal integrity
Root cause analysis
Cross-functional collaboration
Hardware bring-up
ODM/JDM coordination

Tools

PCIe
DFx
EVT/DVT/PVT

Job description

What You Will Do

You will define the hardware that runs the world's largest AI training workloads. Your designs span thermal, mechanical, power, and signal integrity across GPU-accelerated platforms. You will drive validation from first silicon through fleet-scale deployment, triage failures correlating across PCIe, power delivery, memory, and accelerator interconnects, and feed root cause findings back into design improvements. When a new server platform launches at a large scale, the architecture, component choices, and quality gates are yours.

Key job responsibilities
Architecture & Design
  • Define server architectures based on workload demand and customer requirements, translating them into detailed designs and component specifications that enable high-performance AI training and inference at scale
  • Work with interdisciplinary teams of component, firmware, test, qualification, and integration engineers to deliver cohesive designs
  • Drive design reviews with ODM/JDM partners covering schematic, layout, BOM, and manufacturing DFx (Design for Test, Design for Manufacturing)
Validation & Bring-up
  • Define and execute validation strategies from PCBA bring-up through server and rack integration - covering power sequencing, signal integrity, thermal characterization, and accelerator interconnect performance
  • Own hardware debug during EVT/DVT/PVT builds, correlating failures across PCIe, power rails, memory channels, and GPU subsystems
  • Triage hardware issues at both ODM facilities and datacenters, conduct root cause analysis, and implement corrective actions
Fleet Quality & Continuous Improvement
  • Own fleet quality metrics post-launch: server-level annualized failure rates and component-level failure modes
  • Monitor operational telemetry to identify systemic issues and drive design or process changes for current and future platforms
  • Partner with test and automation teams to improve manufacturing yield and reduce test dwell times
Cross-Team Collaboration
  • Work with EC2 architecture teams to align on instance definitions, workload requirements, and platform trade-offs
  • Drive ODM/JDM design partners through development milestones and production ramp
  • Collaborate with firmware, software, and operations teams to ensure designs are debuggable, serviceable, and automation-ready

May require occasional ( https://amazon.jobs/content/en/how-we-hire/accommodations for more information. If the country/region you're applying in isn't listed, please contact your Recruiting Partner.

The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at https://amazon.jobs/en/benefits .

USA, CA, Cupertino - 183,000.00 - 247,600.00 USD annually

USA, TX, Austin - 159,200.00 - 215,300.00 USD annually

USA, WA, Seattle - 159,200.00 - 215,300.00 USD annually

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Sr Cloud Hardware Dev Engineer, AWS Generative AI & ML Servers
Sr Cloud Hardware Dev Engineer, AWS Generative AI & ML Servers

Amazon Development Center U.S., Inc. • Seattle (WA)

On-site
USD 159,000 - 215,000
Health insurance
RSU program
Sign-on bonus
+1
Sr Cloud Hardware Dev Engineer, AWS Generative AI & ML Servers
Sr Cloud Hardware Dev Engineer, AWS Generative AI & ML Servers

Amazon • Seattle (WA)

On-site
USD 159,000 - 215,000
Health insurance
401(k) matching
Paid time off
Senior Hardware Development Engineer, Cloud AI/ML Server Team
Senior Hardware Development Engineer, Cloud AI/ML Server Team

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 159,000 - 215,000
Health insurance
RSU & sign-on bonuses
Parental leave
Senior Hardware Development Engineer, Cloud AI/ML Server Team
Senior Hardware Development Engineer, Cloud AI/ML Server Team

Amazon Web Services (AWS) • Denver (CO)

On-site
USD 159,000 - 215,000
Senior Hardware Development Engineer, Cloud AI/ML Server Team
Senior Hardware Development Engineer, Cloud AI/ML Server Team

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 183,000 - 248,000
Health insurance
401(k) matching
Paid time off
+2
Sr Hardware Development Engineer, High Performance AI & ML Servers
Sr Hardware Development Engineer, High Performance AI & ML Servers

Amazon Web Services (AWS) • Austin (TX)

On-site
USD 159,000 - 215,000
Health insurance
401(k) matching
Paid time off
+1
Sr Hardware Development Engineer, High Performance AI & ML Servers
Sr Hardware Development Engineer, High Performance AI & ML Servers

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 183,000 - 248,000
Health insurance
401(k) matching
RSU equity
Sr Hardware Development Engineer, High Performance AI & ML Servers
Sr Hardware Development Engineer, High Performance AI & ML Servers

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 159,000 - 215,000
Cloud Hardware Dev Engineer (AWS Generative AI & ML Servers), AWS Hardware Engineering Services
Cloud Hardware Dev Engineer (AWS Generative AI & ML Servers), AWS Hardware Engineering Services

Amazon • Cupertino (CA)

On-site
USD 157,000 - 213,000
Cloud Hardware Dev Engineer (AWS Generative AI & ML Servers), AWS Generative AI & ML Servers
Cloud Hardware Dev Engineer (AWS Generative AI & ML Servers), AWS Generative AI & ML Servers

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 183,000 - 248,000