Fleet-Scale Systems Engineer for GPU & AI Accelerators

Amazon

Denver (CO)

On-site

USD 129,000 - 175,000

Full time

10 days ago
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Health insurance
401(k) matching
Paid time off
Parental leave

Job summary

Amazon’s AWS Hardware Engineering team seeks a Systems Development Engineer to own the health and development of server platforms at fleet scale, spanning hardware and software stacks. You will build automation, analyze telemetry across thousands of hosts, and create tooling that informs capacity availability for customers.

You will collaborate with software, hardware, manufacturing, networking, and vendor teams to validate compute solutions and resolve complex production issues.

Qualifications

  • 2+ years of non-internship professional software development experience.
  • 1+ years of designing or architecting of new and existing systems experience.
  • 3+ years of administrative experience in networking, storage systems, operating systems and hands-on systems engineering experience.
  • Knowledge of systems engineering fundamentals (networking, storage, operating systems).
  • Experience programming with at least one modern language such as C++, C#, Java, Python, Golang, PowerShell, Ruby.

Responsibilities

  • Fleet Health & Data Analysis: Analyze hardware failure patterns using fleet telemetry, system event logs, and datacenter tooling to identify root causes and quantify customer impact.
  • Systems Development & Automation: Develop automation for hardware test, firmware qualification, and capacity recovery workflows.
  • Cross-Team Collaboration: Collaborate with software, hardware, manufacturing, networking, and vendor teams to validate and qualify new compute solutions.
  • A day in the life: Participate in sprint-based planning and oncall rotation for platform-level escalations.

Skills

C++
C#
Java
Python
Golang
PowerShell
Ruby

Job description

Amazon’s AWS Hardware Engineering team seeks a Systems Development Engineer to own the health and development of server platforms at fleet scale, spanning hardware and software stacks. You will build automation, analyze telemetry across thousands of hosts, and create tooling that informs capacity availability for customers.

You will collaborate with software, hardware, manufacturing, networking, and vendor teams to validate compute solutions and resolve complex production issues.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Systems Engineer, GPU & AI Accelerator Servers
Systems Engineer, GPU & AI Accelerator Servers

Amazon • Cupertino (CA)

On-site
USD 149,000 - 201,000
Health insurance (medical, dental, and
Stock options or RSUs
401(k) matching
Cloud AI Systems Engineer - Fleet Health & Automation
Cloud AI Systems Engineer - Fleet Health & Automation

Amazon Web Services (AWS) • Austin (TX)

On-site
USD 129,000 - 175,000
Health insurance
RSUs (restricted stock units)
401(k) matching
+1
AI Infra Systems Engineer: GPU & Accelerator Servers
AI Infra Systems Engineer: GPU & Accelerator Servers

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 129,000 - 175,000
Health insurance
401(k) matching
Paid time off
+1
Senior AI/ML Fleet Automation Engineer
Senior AI/ML Fleet Automation Engineer

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 174,000 - 235,000
401(k) matching
Paid time off
Parental leave
+1
Senior Systems Dev Engineer — AI/ML Fleet Health
Senior Systems Dev Engineer — AI/ML Fleet Health

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 151,000 - 205,000
Senior GPU Fleet Engineer – AI Accelerators
Senior GPU Fleet Engineer – AI Accelerators

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 159,000 - 215,000
Senior GPU Fleet Engineer — Hardware & Failure Analysis
Senior GPU Fleet Engineer — Hardware & Failure Analysis

Amazon Web Services (AWS) • Denver (CO)

On-site
USD 159,000 - 215,000
AI Compute Fleet Engineer – GPU & Server Automation
AI Compute Fleet Engineer – GPU & Server Automation

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 149,000 - 201,000
Health insurance
401(k) matching
Paid time off
Senior Systems Dev Engineer, AI/ML Fleet Automation
Senior Systems Dev Engineer, AI/ML Fleet Automation

Amazon • Seattle (WA)

On-site
USD 151,000 - 205,000
Automation Engineer, AI/ML Server Fleet Health & Predictive
Automation Engineer, AI/ML Server Fleet Health & Predictive

Amazon Data Services, Inc. • Cupertino (CA)

On-site
USD 140,000 - 190,000