Cloud Hardware Development Engineer, Cloud AI/ML/storage server teams

Amazon

Denver (CO)

On-site

USD 157,300 - 212,800

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health insurance
401(k) matching
Paid time off
Parental leave

Job summary

Amazon is looking for a Cloud Hardware Development Engineer located in Denver, Colorado. You will own the full lifecycle of AI/ML/GPU server platforms, including design, development, and operational excellence. The role requires strong technical skills, particularly in server technologies, with a focus on improving performance and reliability. A bachelor's degree in electrical or computer engineering and over 5 years of experience are required. Amazon offers comprehensive benefits and a competitive salary range of $157,300 to $212,800 annually.

Qualifications

  • Experience in developing functional specifications, design verification plans, and functional test procedures.
  • 5+ years of professional work experience.
  • Bachelor's degree or above in electrical engineering, computer engineering, or equivalent.

Responsibilities

  • Own the NPI lifecycle for server platforms from architecture to launch.
  • Lead technical solutions for server challenges.
  • Design predictive failure detection systems.

Skills

Developing functional specifications
Design and validation experience
Communication skills
Knowledge of operating systems
Server technologies

Education

Bachelor's degree in electrical/computer engineering

Tools

Telemetry
Predictive failure detection systems

Job description

Overview

As a Cloud Hardware Development Engineer, you will be an end‑to‑end owner of storage and/or accelerator (AI/ML/GPU) server platforms — from New Product Introduction (NPI) through fleet health in production. You own the full lifecycle: design, development, qualification, launch, and ongoing operational excellence of servers running at scale in the AWS fleet.

You will work closely with internal customers to understand their technical needs and business goals, leveraging your experience with server design and the knowledge of various teams to architect solutions we deploy at scale. To deliver your products, you will work with an interdisciplinary team of component, firmware, power, mechanical, electrical, test, qualification, and manufacturing engineers, and lead our ODM (design and manufacturing partners) to bring these servers to the data center. After launch, you own the fleet — monitoring quality, driving reliability improvements, and ensuring servers continue to meet customer requirements throughout their operational life.

This role demands deep technical curiosity and the willingness to jump in and personally solve the hardest problems. When a complex system failure occurs — whether during NPI qualification or in a production fleet of hundreds of thousands of servers — you roll up your sleeves, dive into the details across hardware, firmware, software, and physical layers, and drive to root cause. You don't wait for someone else to figure it out.

You will own end‑to‑end system reliability — proactively identifying deficiencies and driving toward zero‑touch operations where automation detects, diagnoses, and resolves issues before customer impact. You will decompose complex server system problems (testability, reliability, diagnostics) into deliverable tasks and features, leading delivery yourself and through others in parallel.

This is a fast‑paced, intellectually challenging position. You'll work with thought leaders in multiple technology areas, hold high standards for yourself and everyone you work with, and constantly look for ways to improve your products' performance, quality, and cost. We're changing an industry, and we want individuals who are ready for this challenge and want to reach beyond what is possible today.

Key Job Responsibilities
  • Own the end‑to‑end NPI lifecycle for storage and/or accelerator (AI/ML/GPU) server platforms — from architecture definition through design, qualification, manufacturing ramp, and launch.

  • Lead technical solutions for complex server and rack system architectural challenges.

  • Work with ODM/manufacturing partners to develop, validate, and manufacture server products at scale.

  • Develop functional specifications, design verification plans, and test procedures.

  • Drive qualification and readiness milestones, ensuring new platforms meet performance, reliability, and cost targets before fleet deployment.

  • Identify and resolve technical risks early in the development cycle — don't let problems reach production.

Fleet Health, Diagnostics & Automation
  • Own fleet health for the server platforms you launch — reliability doesn't end at ship.

  • Design and implement predictive failure detection systems using telemetry, sensor data, error trending, and log correlation to identify hardware issues before they cause customer impact.

  • Drive toward zero‑touch operations — help build detection, diagnosis, and remediation of faults without human intervention.

  • Debug complex system failures in time‑sensitive settings — personally diving deep when the problem demands it.

  • Perform root cause analysis correlating across firmware, kernel, driver, thermal, power, and physical layers.

Systems Design & Technical Depth
  • Apply expertise across hardware, software, system design, x86 architecture, processes, and operations (compute, storage, network, GPU).

  • Design and implement solutions to address system‑level issues at large scale.

  • Decompose complex server system problems (testability, reliability, diagnostics) into deliverable tasks and features.

  • Collaborate with hardware, software, manufacturing, supply chain, and product management teams.

Cross‑Team Collaboration
  • Work closely with internal customers to ensure new server hardware meets data path and control path requirements.

  • Identify early any potential problems onboarding new servers into customer ecosystems.

  • Collaborate across Hardware Engineering, component, firmware, test, qualification, and integration teams.

  • Partner with datacenter operations to close the loop between field failures and design improvements.

A Day in the Life

Your day‑to‑day responsibilities include interfacing with internal and external customers to understand product requirements and facilitate system development on top of your server designs. You will learn operational challenges facing our existing fleet with the goal of improving the current customer experience and developing improved systems for future designs. You will work directly with vendors and ODM (manufacture partners) to scale your product. Some days you're reviewing a new platform design with your ODM; other days you're deep in logs and telemetry data chasing a failure mode across the fleet. You thrive on that range.

Basic Qualifications
  • Experience in developing functional specifications, design verification plans and functional test procedures.

  • Bachelor's degree or above in electrical engineering, computer engineering, or equivalent.

  • Experience in English‑language communication skills, both written and verbal.

  • Experience with design & innovation and research & development.

  • Knowledge of operating systems, hardware, storage, network, security, database administration and cloud infrastructure.

  • Experience in server technologies such as thermal, mechanical, power, and signal integrity.

  • 5+ years of professional work (non‑internship) experience.

Preferred Qualifications
  • 5+ years of hardware design and validation of components, subsystems and systems experience.

  • Experience in server technologies: board design, high‑speed bus design and signal integrity, failure analysis, server components (CPU, GPU, SSDs, memory), BIOS, BMC, and networking.

  • Experience developing and executing test procedures for mechanical or electrical systems/components.

  • Experience working with ODMs/manufacturers through the product development and manufacturing lifecycle.

  • Experience building predictive failure detection or proactive remediation systems at fleet scale.

  • Experience with storage/compute/GPU/accelerator platforms including integration, diagnostics, or performance validation.

  • Familiarity with PCIe topology, NVLink, NVMe, and accelerator interconnects.

  • Experience with large‑scale datacenter or cloud environments.

EEO Statement

Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status.

Benefits

Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave.

Compensation

USA, CA, Cupertino - $157,300.00 - $212,800.00 USD annually

USA, WA, Seattle - $136,000.00 - $184,000.00 USD annually

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Manufacturing hardware engineer, Cloud AI/ML/storage server teams
Manufacturing hardware engineer, Cloud AI/ML/storage server teams

Amazon Web Services (AWS) • Denver (CO)

On-site
USD 136,000 - 184,000
Health insurance
401(k) matching
Paid time off
+1
Senior Cloud Hardware Development Engineer, Cloud AI/ML/storage server teams
Senior Cloud Hardware Development Engineer, Cloud AI/ML/storage server teams

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 159,000 - 215,000
Health insurance
RSUs
401(k) matching
Senior Cloud Hardware Development Engineer, Cloud AI/ML/storage server teams
Senior Cloud Hardware Development Engineer, Cloud AI/ML/storage server teams

Amazon Web Services (AWS) • Denver (CO)

On-site
USD 159,000 - 216,000
Health insurance
401(k) matching
Stock options
Senior Hardware Development Engineer AWS AI & ML, Accelerator Servers
Senior Hardware Development Engineer AWS AI & ML, Accelerator Servers

Amazon Web Services (AWS) • Austin (TX)

On-site
USD 159,000 - 216,000
Health insurance
401(k) matching
Paid time off
+1
Sr Hardware Development Engineer, High Performance AI & ML Servers
Sr Hardware Development Engineer, High Performance AI & ML Servers

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 183,000 - 248,000
Health insurance
401(k) matching
RSU equity
Cloud Hardware Dev Engineer (AWS Generative AI & ML Servers), AWS Hardware Engineering Services
Cloud Hardware Dev Engineer (AWS Generative AI & ML Servers), AWS Hardware Engineering Services

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 157,000 - 213,000
Health insurance
401(k) matching
Paid time off
+1
Cloud Hardware Dev Engineer (AWS Generative AI & ML Servers), AWS Hardware Engineering Services
Cloud Hardware Dev Engineer (AWS Generative AI & ML Servers), AWS Hardware Engineering Services

Amazon • Cupertino (CA)

On-site
USD 157,000 - 213,000
Sr. Technical Program Manager, Hardware NPI, Edge & High Performance Accelerator Servers for AI/ML
Sr. Technical Program Manager, Hardware NPI, Edge & High Performance Accelerator Servers for AI/ML

Amazon • Austin (TX)

On-site
USD 149,000 - 201,000
Senior Hardware Development Engineer AWS AI & ML, Accelerator Servers
Senior Hardware Development Engineer AWS AI & ML, Accelerator Servers

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 159,000 - 216,000
Health insurance
401(k) matching
Paid time off
+1
Cloud Hardware Dev Engineer (AWS Generative AI & ML Servers), AWS Hardware Engineering Services
Cloud Hardware Dev Engineer (AWS Generative AI & ML Servers), AWS Hardware Engineering Services

Amazon Web Services (AWS) • Austin (TX)

On-site
USD 136,000 - 184,000