Hardware Development Engineer, AI/ML Server Development

Amazon Web Services (AWS)

Cupertino (CA)

On-site

USD 126,000 - 185,000

Full time

2 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Amazon Web Services (AWS) is seeking a Hardware Development Engineer in Cupertino, CA to join our GPU server platforms team. You will investigate hardware issues, analyze failure data, and work with senior engineers to understand server architecture, contributing to reliability and fleet operations.

You will learn end-to-end system behavior across hardware, firmware, software, and operations, with mentorship and opportunities to influence design improvements that power AI workloads for millions

Qualifications

  • Bachelor's degree in electrical or computer engineering or equivalent.
  • Hands-on Linux experience in coursework, labs or projects.
  • Understanding of computer architecture (CPU, memory, buses, I/O).
  • Ability to communicate technical concepts clearly and visually.
  • Desire to work in a fast-paced, learning-intensive environment.

Responsibilities

  • Investigate and root case hardware issues on GPU server platforms with guidance.
  • Analyze failure data and trends to identify systemic problems and fixes.
  • Work with senior engineers to understand server architecture, GPU subsystems, and fleet operations.
  • Participate in design reviews for new server platforms with a reliability focus.
  • Collaborate with cross-functional teams to close the loop between field failures and design improvements.
  • Document findings, contribute to runbooks, and share knowledge with the team.

Skills

Python
Java
C
C++
Linux
Communication
Hardware debugging

Education

Bachelor's degree in Electrical or Computer Engineering

Tools

Linux

Job description

Description

Application deadline: Sep 16, 2026

Are you curious about what happens inside the servers that power AI? Do you enjoy debugging problems and figuring out why something broke? Are you excited to learn fast, work on real systems, and see your contributions matter from day one? If that sounds like you, we'd love to meet you.

AWS Hardware Engineering builds and maintains the servers behind the world's largest cloud. Our team focuses on GPU-accelerated platforms - the infrastructure that runs generative AI, machine learning training, and large language models at global scale. It's complex, it's fast-moving, and it's an incredible place to start your career.

We're looking for a Hardware Development Engineer to join us and grow. You don't need to know everything on day one - you need to be curious, willing to learn, and energized by solving problems. You'll be supported by experienced engineers who' ll mentor you as you ramp up on real fleet challenges: understanding why a GPU server misbehaves, building tools to catch issues automatically, and helping improve the reliability of infrastructure used by millions.

You’ll learn how servers work end-to-end - hardware, firmware, software, and operations - and you’ll contribute to a team that’s pushing toward a future where systems heal themselves. Your ideas will matter, your work will ship, and you’ll develop skills that are hard to get anywhere else.

Key job responsibilities
  • Investigate and root cause hardware issues on GPU server platforms with support of an experienced engineer - you’ll learn to trace problems across electrical, firmware, and software layers
  • Analyze failure data and trends to identify systemic problems and propose fixes
  • Work with senior engineers to understand server architecture, GPU subsystems, and fleet operations
  • Participate in design reviews for new server platforms, contributing a reliability and serviceability perspective
  • Collaborate with cross-functional teams (firmware, software, manufacturing, datacenter operations) to close the loop between field failures and design improvements
  • Document findings, contribute to runbooks, and share knowledge with the team
A day in the life
  • You’ll work on infrastructure at a scale that doesn't exist anywhere else
  • You’ll develop rare cross-stack expertise (hardware + firmware + software + operations) early in your career
  • You’ll have mentorship from senior engineers and a team that invests in your growth
  • You’ll see your work directly improve the systems powering the next generation of AI
Basic Qualifications
  • Bachelor's degree or above in electrical engineering, computer engineering, or equivalent
  • Experience communicating technical concepts and processes using clear, simple language and visuals
  • Understanding of computer architecture (CPU, memory, buses, I/O)
  • Hands-on experience with Linux environments (coursework, labs, or personal projects)
  • Strong organizational skills and the ability to work well with cross-functional teams
  • Desire and energy to work in a fast-paced, learning-intensive environment
Preferred Qualifications
  • Experience programming or scripting language like Python, Java, C or C++
  • co-op program at your university in engineering or equivalent, or experience from previous technical internship(s) or demonstrated project experience
  • 1+ years of industry work experience
  • Coursework or projects in embedded systems, computer architecture, or hardware-software interaction
  • Experience debugging hardware or software issues in lab or project settings

Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status.

Los Angeles County applicants: Job duties for this position include: work safely and cooperatively with other employees, supervisors, and staff; adhere to standards of excellence despite stressful conditions; communicate effectively and respectfully with employees, supervisors, and staff to ensure exceptional customer service; and follow all federal, state, and local laws and Company policies. Criminal history may have a direct, adverse, and negative relationship with some of the material job duties of this position. These include the duties and responsibilities listed above, as well as the abilities to adhere to company policies, exercise sound judgment, effectively manage stress and work safely and respectfully with others, exhibit trustworthiness and professionalism, and safeguard business operations and the Company’s reputation. Pursuant to the Los Angeles County Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records.

Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit https://amazon.jobs/content/en/how-we-hire/accommodations for more information. If the country/region you're applying in isn't listed, please contact your Recruiting Partner.

The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at https://amazon.jobs/en/benefits.

USA, CA, Cupertino - 126,100.00 - 185,000.00 USD annually

USA, CO, Denver - 109,700.00 - 160,000.00 USD annually

USA, WA, Seattle - 109,700.00 - 160,000.00 USD annually

Company - Amazon Data Services, Inc.

Job ID: A10482698

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Hardware Development Engineer, AI/ML Server Development
Hardware Development Engineer, AI/ML Server Development

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 110,000 - 160,000
Senior Hardware Development Engineer, Cloud AI/ML Server Team
Senior Hardware Development Engineer, Cloud AI/ML Server Team

Amazon Web Services (AWS) • Denver (CO)

On-site
USD 159,000 - 215,000
Senior Hardware Development Engineer, Cloud AI/ML Server Team
Senior Hardware Development Engineer, Cloud AI/ML Server Team

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 159,000 - 215,000
Health insurance
RSU & sign-on bonuses
Parental leave
Senior Hardware Development Engineer, Cloud AI/ML Server Team
Senior Hardware Development Engineer, Cloud AI/ML Server Team

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 183,000 - 248,000
Health insurance
401(k) matching
Paid time off
+2
Hardware Development Engineer, AI/ML Server Development (AWS)
Hardware Development Engineer, AI/ML Server Development (AWS)

Amazon • Cupertino (CA)

On-site
USD 126,000 - 185,000
RSUs
Health benefits
401(k) matching
Systems Development Engineer, GPU & AI Accelerator Servers, AWS Hardware Engineering
Systems Development Engineer, GPU & AI Accelerator Servers, AWS Hardware Engineering

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 129,000 - 175,000
Health insurance
401(k) matching
Paid time off
+1
Systems Development Engineer, GPU & AI Accelerator Servers, AWS Hardware Engineering
Systems Development Engineer, GPU & AI Accelerator Servers, AWS Hardware Engineering

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 149,000 - 201,000
Health insurance
401(k) matching
Paid time off
Senior Accelerator Engineer, Cloud AI/ML server team
Senior Accelerator Engineer, Cloud AI/ML server team

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 183,000 - 248,000
Health insurance
401(k) matching
Paid time off
+1
Senior Accelerator Engineer, Cloud AI/ML server team
Senior Accelerator Engineer, Cloud AI/ML server team

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 159,000 - 215,000
Systems Development Engineer, AWS Generative AI & ML Servers
Systems Development Engineer, AWS Generative AI & ML Servers

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 149,000 - 201,000