Systems Development Engineer, GPU & AI Accelerator Servers, AWS Hardware Engineering

Amazon

Cupertino (CA)

On-site

USD 149,000 - 201,000

Full time

25 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health insurance (medical, dental, and
Stock options or RSUs
401(k) matching

Job summary

Amazon Development Center U.S., Inc. seeks a Systems Development Engineer to own health and development of server platforms at worldwide fleet scale.

You will build automation, analyze telemetry across thousands of hosts, and create tooling that determines capacity for customers. You will work across hardware monitoring interfaces to fleet-wide data pipelines and dashboards, spanning Linux on ARM/x86, PCIe, Power, NIC, NVMe, and GPU subsystems.

Qualifications

  • 2+ years of non-internship professional software development experience
  • 1+ years of designing or architecting (design patterns, reliability and scaling) of new and existing systems experience
  • 3+ years of administrative experience in networking, storage systems, operating systems and hands-on systems engineering experience
  • Knowledge of systems engineering fundamentals (networking, storage, operating systems)
  • Experience programming with at least one modern language such as C++, C#, Java, Python, Golang, PowerShell, Ruby

Responsibilities

  • Fleet Health & Data Analysis
  • Analyze hardware failure patterns using fleet telemetry, system event logs, and datacenter tooling to identify root causes and quantify customer impact
  • Contribute to predictive failure detection using sensor data, error trending, and log correlation
  • Build and maintain operational dashboards and metrics for platform fleet health
  • Build tooling to track component lifecycle (firmware versions, part revisions, supply chain status) across large-scale fleets
  • Develop and maintain automation for hardware test, firmware qualification, and capacity recovery workflows
  • Develop diagnostic tools for Linux on ARM and x86 architectures
  • Debug and resolve Linux boot and runtime issues across processor architectures - PCIe, Power, NIC, NVMe, and GPU subsystems
  • Build automation solutions using Python, Java, or similar languages with focus on scalability and operational durability
  • Collaborate with software, hardware, manufacturing, networking, and vendor teams to validate and qualify new compute solutions
  • Troubleshoot complex system-level issues in production environments, correlating across firmware, operating systems, drivers, and physical layers
  • Participate in sprint-based planning and oncall rotation for platform-level escalations
  • A day in the life: you work with hardware engineers, firmware teams, datacenter operations, and vendor partners – driving quality and reliability from manufacturing through steady-state operations

Skills

Non-internship software development
System design/architecture
Administrative experience in networks,
Systems engineering fundamentals
Programming languages (C++, C#, Java,

Tools

PowerShell
Python
Ruby
Java

Job description

Systems Development Engineer, GPU & AI Accelerator Servers, AWS Hardware Engineering

Job ID: 10517016 | Amazon Development Center U.S., Inc.

Application deadline: Sep 1, 2026

Do you want to build the infrastructure that keeps Artificial Intelligence compute capacity available to Generative AI customers? Do you want to solve problems at the boundary between physical hardware and software - at cloud scale?

AWS Hardware Engineering is looking for a Systems Development Engineer to own the health and development of server platforms at worldwide fleet scale. You will develop automation, analyze hardware telemetry across tens of thousands of hosts, and build tooling that directly determines whether capacity is available to customers. Your work spans the full stack – from hardware monitoring interfaces, to health diagnostics up through fleet-wide data pipelines and operational dashboards.

Key job responsibilities
  • Fleet Health & Data Analysis
  • Analyze hardware failure patterns using fleet telemetry, system event logs, and datacenter tooling to identify root causes and quantify customer impact
  • Contribute to predictive failure detection using sensor data, error trending, and log correlation
  • Build and maintain operational dashboards and metrics for platform fleet health.
  • Build tooling to track component lifecycle (firmware versions, part revisions, supply chain status) across large-scale fleets
Systems Development & Automation
  • Develop and maintain automation for hardware test, firmware qualification, and capacity recovery workflows
  • Develop diagnostic tools for Linux on ARM and x86 architectures
  • Debug and resolve Linux boot and runtime issues across processor architectures - PCIe, Power, NIC, NVMe, and GPU subsystems
  • Build automation solutions using Python, Java, or similar languages with focus on scalability and operational durability
Cross-Team Collaboration
  • Collaborate with software, hardware, manufacturing, networking, and vendor teams to validate and qualify new compute solutions
  • Troubleshoot complex system-level issues in production environments, correlating across firmware, operating systems, drivers, and physical layers
  • Participate in sprint-based planning and oncall rotation for platform-level escalations
A day in the life

Some days you are deep in system event logs chasing a failure pattern across thousands of hosts; other days you are writing automation that eliminates a manual triage workflow entirely. You work with hardware engineers, firmware teams, datacenter operations, and vendor partners – driving quality and reliability from manufacturing through steady-state operations. Located in Cupertino, Seattle, or Denver, you work with global development teams on servers deployed in datacenters worldwide.

About the team

AWS Hardware Engineering designs and delivers next-generation cloud infrastructure - the servers, accelerators, and storage platforms that power AWS. Our team builds custom systems for AI training, inference, and compute workloads at global scale. We are directly responsible for launching and maintaining server hardware in the fleet, working across internal development teams, and design partners.

We value work-life harmony, inclusive culture, and continuous learning. Even if you do not meet all preferred qualifications listed below, we encourage you to apply.

Basic Qualifications
  • 2+ years of non-internship professional software development experience
  • 1+ years of designing or architecting (design patterns, reliability and scaling) of new and existing systems experience
  • 3+ years of administrative experience in networking, storage systems, operating systems and hands-on systems engineering experience
  • Knowledge of systems engineering fundamentals (networking, storage, operating systems)
  • Experience programming with at least one modern language such as C++, C#, Java, Python, Golang, PowerShell, Ruby
Preferred Qualifications
  • Experience with PowerShell (preferred), Python, Ruby, or Java
  • Experience working in an Agile environment using the Scrum methodology

Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status.

Los Angeles County applicants: Job duties for this position include: work safely and cooperatively with other employees, supervisors, and staff; adhere to standards of excellence despite stressful conditions; communicate effectively and respectfully with employees, supervisors, and staff to ensure exceptional customer service; and follow all federal, state, and local laws and Company policies. Criminal history may have a direct, adverse, and negative relationship with some of the material job duties of this position. These include the duties and responsibilities listed above, as well as the abilities to adhere to company policies, exercise sound judgment, effectively manage stress and work safely and respectfully with others, exhibit trustworthiness and professionalism, and safeguard business operations and the Company’s reputation. Pursuant to the Los Angeles County Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records.

Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit https://amazon.jobs/content/en/how-we-hire/accommodations for more information. If the country/region you’re applying in isn't listed, please contact your Recruiting Partner.

The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at https://amazon.jobs/en/benefits .

  • USA, CA, Cupertino - 148,700.00 - 201,200.00 USD annually
  • USA, CO, Denver - 129,200.00 - 174,800.00 USD annually
  • USA, WA, Seattle - 129,200.00 - 174,800.00 USD annually
Important FAQs for current Government employees

Before proceeding, please review the following FAQs

https://www.amazon.jobs/en/faqs#faqs-for-us-government-employees

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Cloud Hardware Development Engineer II, AI/ML/Storage Server Team (AWS)
Cloud Hardware Development Engineer II, AI/ML/Storage Server Team (AWS)

Amazon • Cupertino (CA)

On-site
USD 157,000 - 213,000
Comprehensive benefits including RSU/"
Systems Development Engineer, AWS Generative AI & ML Servers
Systems Development Engineer, AWS Generative AI & ML Servers

Amazon Web Services (AWS) • Austin (TX)

On-site
USD 129,000 - 175,000
Health insurance
RSUs (restricted stock units)
401(k) matching
+1
Systems Development Engineer, AWS Generative AI & ML Servers
Systems Development Engineer, AWS Generative AI & ML Servers

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 129,000 - 175,000
Health insurance
RSUs
401(k) matching
+1
Systems Development Eng (AWS Generative AI & ML Servers), AWS Hardware Engineering Accelerators
Systems Development Eng (AWS Generative AI & ML Servers), AWS Hardware Engineering Accelerators

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 149,000 - 201,000
Health insurance
401(k) matching
Paid time off
Systems Development Engineer (AWS Generative AI & ML Servers), AWS HW Engineering
Systems Development Engineer (AWS Generative AI & ML Servers), AWS HW Engineering

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 149,000 - 201,000
Sr. System Development Engineer, High-Performance Accelerator Servers for AI/ML
Sr. System Development Engineer, High-Performance Accelerator Servers for AI/ML

Amazon Web Services (AWS) • Austin (TX)

On-site
USD 151,000 - 205,000
Health insurance
401(k) matching
Paid time off
+1
Hardware Development Engineer, AI/ML Server Development (AWS)
Hardware Development Engineer, AI/ML Server Development (AWS)

Amazon • Cupertino (CA)

On-site
USD 126,000 - 185,000
RSUs
Health benefits
401(k) matching
Sr. System Development Engineer, Edge & High Performance Accelerator Servers for AI/ML
Sr. System Development Engineer, Edge & High Performance Accelerator Servers for AI/ML

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 174,000 - 235,000
Sr. System Development Engineer, Edge & High Performance Accelerator Servers for AI/ML
Sr. System Development Engineer, Edge & High Performance Accelerator Servers for AI/ML

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 151,000 - 205,000
RSUs
Health insurance
401(k) matching
Systems Development Eng (AWS Generative AI & ML Servers), AWS Hardware Engineering Accelerators
Systems Development Eng (AWS Generative AI & ML Servers), AWS Hardware Engineering Accelerators

Amazon Web Services (AWS) • Austin (TX)

On-site
USD 129,000 - 175,000