Sr Cloud Hardware Dev Engineer, AWS Generative AI & ML Servers

Amazon Web Services (AWS)

Austin (TX)

On-site

USD 159,000 - 215,000

Full time

5 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Health insurance
401(k) matching
Paid time off
Parental leave

Job summary

Amazon Development Center U.S., Inc. seeks a Cloud Hardware Development Engineer to define server architectures for AI workloads and drive validation from silicon to fleet deployment.

You will lead ODM/JDM partners through development, triage hardware issues across datacenters, and own fleet quality metrics post-launch. You will work across thermal, mechanical, power delivery, and signal integrity disciplines, collaborating with firmware, software, and operations teams to ensure end-to-end

Qualifications

  • Bachelor's degree in Electrical or Computer Engineering or equivalent.
  • 7+ years hardware design and development for server or large-scale infra.
  • Experience leading hardware development through full product lifecycle.

Responsibilities

  • Define hardware for AI training workloads across thermal, mechanical, power and signal integrity.
  • Drive validation from PCBA bring-up to fleet deployment; triage issues across PCIe, power, memory and interconnects.
  • Own EVT/DVT/PVT hardware debug; implement corrective actions.
  • Triage issues at ODM facilities and datacenters; monitor fleet quality metrics and drive design improvements.
  • Collaborate with cross-functional teams and ODM/JDM partners to ensure debuggable, serviceable designs.

Skills

Hardware design
Verification plans
Functional testing
Lifecycle management
Thermal design
Power delivery
Signal integrity
ODM collaboration

Education

Bachelor's degree in Electrical/Computer Engineering
Master's or PhD in EE/CE or related field

Job description

Description
AWS operates the world's largest fleet of GPU-accelerated servers powering AI/ML training and inference at cloud scale. Our team defines the server architectures, drives the hardware designs, and owns the fleet quality for these platforms — from component selection through datacenter operations. If you want to shape the physical hardware that frontier models train on, this is the role.

Description
AWS operates the world's largest fleet of GPU-accelerated servers powering AI/ML training and inference at cloud scale. Our team defines the server architectures, drives the hardware designs, and owns the fleet quality for these platforms — from component selection through datacenter operations. If you want to shape the physical hardware that frontier models train on, this is the role.

We are seeking a Cloud Hardware Development Engineer to define server architectures based on workload demand, translate them into detailed component specifications, and drive validation from PCBA bring-up through rack integration. You will lead ODM design partners through development and production, triage hardware issues across manufacturing and datacenters, and own fleet quality metrics post-launch.

What You Will Do

You will define the hardware that runs the world's largest AI training workloads. Your designs span thermal, mechanical, power, and signal integrity across GPU-accelerated platforms. You will drive validation from first silicon through fleet-scale deployment, triage failures correlating across PCIe, power delivery, memory, and accelerator interconnects, and feed root cause findings back into design improvements. When a new server platform launches at a large scale, the architecture, component choices, and quality gates are yours.

Key job responsibilities
Architecture & Design
  • Define server architectures based on workload demand and customer requirements, translating them into detailed designs and component specifications that enable high-performance AI training and inference at scale
  • Work with interdisciplinary teams of component, firmware, test, qualification, and integration engineers to deliver cohesive designs
  • Drive design reviews with ODM/JDM partners covering schematic, layout, BOM, and manufacturing DFx (Design for Test, Design for Manufacturing)
Validation & Bring-up
  • Define and execute validation strategies from PCBA bring-up through server and rack integration — covering power sequencing, signal integrity, thermal characterization, and accelerator interconnect performance
  • Own hardware debug during EVT/DVT/PVT builds, correlating failures across PCIe, power rails, memory channels, and GPU subsystems
  • Triage hardware issues at both ODM facilities and datacenters, conduct root cause analysis, and implement corrective actions
Fleet Quality & Continuous Improvement
  • Own fleet quality metrics post-launch: server-level annualized failure rates and component-level failure modes
  • Monitor operational telemetry to identify systemic issues and drive design or process changes for current and future platforms
  • Partner with test and automation teams to improve manufacturing yield and reduce test dwell times
Cross-Team Collaboration
  • Work with EC2 architecture teams to align on instance definitions, workload requirements, and platform trade-offs
  • Drive ODM/JDM design partners through development milestones and production ramp
  • Collaborate with firmware, software, and operations teams to ensure designs are debuggable, serviceable, and automation-ready

May require occasional (<10%) regional and international travel to Design and Manufacturing Partner sites.

The Ideal Candidate

You think across the full hardware stack — from silicon packaging and power delivery to rack-level thermal and mechanical design. You are as comfortable reviewing a schematic as you are analyzing fleet failure data. You drive quality through data, not assumption, and you hold design partners to the same standard you hold yourself. You mentor and develop junior engineers, contribute to hiring, and share your expertise to make the team stronger.

Why You Will Love It

The world's most advanced frontier models train on the hardware you design. You will see your architecture decisions scale to a large fleet of servers. The team is deeply technical and high-trust — you own platforms end to end from architecture definition through fleet operations.

A day in the life

You start the day reviewing thermal and power validation data from an EVT build at your ODM partner. Mid-morning, you join a design review to close signal integrity findings on a high-speed accelerator interconnect. In the afternoon, you triage a fleet quality signal — correlating component-level failure data with manufacturing lot information to identify a systemic issue. You end the day aligning with architecture teams on requirements for the next-generation platform.

About The Team

The Hardware Engineering AI/ML UltraServer platform team is a group of engineers and technical program managers directly responsible for launching GPU-accelerated servers into the AWS fleet. Located in Seattle, Austin, and Cupertino, we collaborate with global development teams and ODM partners to deliver next-generation AI/ML infrastructure deployed in datacenters worldwide. We move fast with small, empowered teams delivering end-to-end — from server conception through fleet-scale operations.

Basic Qualifications
  • Bachelor's degree in electrical engineering, computer engineering, or equivalent
  • Experience in developing functional specifications, design verification plans and functional test procedures
  • 7+ years of hardware design and development experience for server, compute, or large-scale infrastructure platforms
  • Experience in one or more server technologies: thermal/mechanical design, power delivery, high-speed signal integrity, or accelerator subsystems
  • Experience leading hardware development through full product lifecycle (concept through production ramp)
Preferred Qualifications
  • Master's degree or PhD in Electrical Engineering, Computer Engineering, or a related field
  • 5+ years of experience working with ODMs through the product development and manufacturing lifecycle (EVT, DVT, PVT)
  • In-depth expertise in high-speed bus design, signal integrity analysis, or power delivery for GPU/accelerator platforms
  • 5+ years of experience with hardware bring-up, debug, and root cause analysis across PCIe, NVMe, memory, and accelerator interconnects
  • Experience owning fleet quality metrics and driving design improvements based on operational failure data
  • Experience with thermal/mechanical design for high-power-density compute platforms (liquid cooling, air cooling, or hybrid)
  • Experience working in large-scale datacenter or cloud environments
  • Track record of defining engineering standards and design best practices adopted across teams or partner organizations

Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status.

Los Angeles County applicants: Job duties for this position include: work safely and cooperatively with other employees, supervisors, and staff; adhere to standards of excellence despite stressful conditions; communicate effectively and respectfully with employees, supervisors, and staff to ensure exceptional customer service; and follow all federal, state, and local laws and Company policies. Criminal history may have a direct, adverse, and negative relationship with some of the material job duties of this position. These include the duties and responsibilities listed above, as well as the abilities to adhere to company policies, exercise sound judgment, effectively manage stress and work safely and respectfully with others, exhibit trustworthiness and professionalism, and safeguard business operations and the Company’s reputation. Pursuant to the Los Angeles County Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records.

Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit https://amazon.jobs/content/en/how-we-hire/accommodations for more information. If the country/region you’re applying in isn’t listed, please contact your Recruiting Partner.

The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at https://amazon.jobs/en/benefits.

USA, CA, Cupertino - 183,000.00 - 247,600.00 USD annually

USA, TX, Austin - 159,200.00 - 215,300.00 USD annually

USA, WA, Seattle - 159,200.00 - 215,300.00 USD annually

Company - Amazon Development Center U.S., Inc.

Job ID: A10533788

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Sr Cloud Hardware Dev Engineer, AWS Generative AI & ML Servers
Sr Cloud Hardware Dev Engineer, AWS Generative AI & ML Servers

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 159,000 - 215,000
Sr Cloud Hardware Dev Engineer, AWS Generative AI & ML Servers
Sr Cloud Hardware Dev Engineer, AWS Generative AI & ML Servers

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 183,000 - 248,000
Health insurance
RSUs
Senior Hardware Development Engineer, Cloud AI/ML Server Team
Senior Hardware Development Engineer, Cloud AI/ML Server Team

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 159,000 - 215,000
Health insurance
RSU & sign-on bonuses
Parental leave
Senior Hardware Development Engineer, Cloud AI/ML Server Team
Senior Hardware Development Engineer, Cloud AI/ML Server Team

Amazon Web Services (AWS) • Denver (CO)

On-site
USD 159,000 - 215,000
Senior Hardware Development Engineer, Cloud AI/ML Server Team
Senior Hardware Development Engineer, Cloud AI/ML Server Team

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 183,000 - 248,000
Health insurance
401(k) matching
Paid time off
+2
Senior Hardware Development Engineer, Cloud AI/ML Server Team (AWS)
Senior Hardware Development Engineer, Cloud AI/ML Server Team (AWS)

Amazon Inc. • Cupertino (CA)

On-site
USD 183,000 - 248,000
Health insurance
401(k) matching
Paid time off
+1
Sr Systems Development Engineer, AWS AI/ML Servers
Sr Systems Development Engineer, AWS AI/ML Servers

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 174,000 - 235,000
Health insurance
401(k) matching
Paid time off
+1
Sr Hardware Development Engineer, High Performance AI & ML Servers
Sr Hardware Development Engineer, High Performance AI & ML Servers

Amazon Web Services (AWS) • Austin (TX)

On-site
USD 159,000 - 215,000
Health insurance
401(k) matching
Paid time off
+1
Sr Cloud Hardware Dev Engineer, AWS Generative AI & ML Servers
Sr Cloud Hardware Dev Engineer, AWS Generative AI & ML Servers

Amazon Development Center U.S., Inc. • Cupertino (CA)

On-site
USD 183,000 - 248,000
Health insurance
RSUs
401(k) matching
+1
Sr Systems Development Engineer, AWS AI/ML Servers
Sr Systems Development Engineer, AWS AI/ML Servers

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 151,000 - 205,000
Health insurance
401(k) matching
Paid time off
+1