Systems Engineer, GPU & AI Accelerator Servers

Amazon

Cupertino (CA)

On-site

USD 149,000 - 201,000

Full time

11 days ago
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Health insurance (medical, dental, and
Stock options or RSUs
401(k) matching

Job summary

Amazon Development Center U.S., Inc. seeks a Systems Development Engineer to own health and development of server platforms at worldwide fleet scale.

You will build automation, analyze telemetry across thousands of hosts, and create tooling that determines capacity for customers. You will work across hardware monitoring interfaces to fleet-wide data pipelines and dashboards, spanning Linux on ARM/x86, PCIe, Power, NIC, NVMe, and GPU subsystems.

Qualifications

  • 2+ years of non-internship professional software development experience
  • 1+ years of designing or architecting (design patterns, reliability and scaling) of new and existing systems experience
  • 3+ years of administrative experience in networking, storage systems, operating systems and hands-on systems engineering experience
  • Knowledge of systems engineering fundamentals (networking, storage, operating systems)
  • Experience programming with at least one modern language such as C++, C#, Java, Python, Golang, PowerShell, Ruby

Responsibilities

  • Fleet Health & Data Analysis
  • Analyze hardware failure patterns using fleet telemetry, system event logs, and datacenter tooling to identify root causes and quantify customer impact
  • Contribute to predictive failure detection using sensor data, error trending, and log correlation
  • Build and maintain operational dashboards and metrics for platform fleet health
  • Build tooling to track component lifecycle (firmware versions, part revisions, supply chain status) across large-scale fleets
  • Develop and maintain automation for hardware test, firmware qualification, and capacity recovery workflows
  • Develop diagnostic tools for Linux on ARM and x86 architectures
  • Debug and resolve Linux boot and runtime issues across processor architectures - PCIe, Power, NIC, NVMe, and GPU subsystems
  • Build automation solutions using Python, Java, or similar languages with focus on scalability and operational durability
  • Collaborate with software, hardware, manufacturing, networking, and vendor teams to validate and qualify new compute solutions
  • Troubleshoot complex system-level issues in production environments, correlating across firmware, operating systems, drivers, and physical layers
  • Participate in sprint-based planning and oncall rotation for platform-level escalations
  • A day in the life: you work with hardware engineers, firmware teams, datacenter operations, and vendor partners – driving quality and reliability from manufacturing through steady-state operations

Skills

Non-internship software development
System design/architecture
Administrative experience in networks,
Systems engineering fundamentals
Programming languages (C++, C#, Java,

Tools

PowerShell
Python
Ruby
Java

Job description

Amazon Development Center U.S., Inc. seeks a Systems Development Engineer to own health and development of server platforms at worldwide fleet scale.

You will build automation, analyze telemetry across thousands of hosts, and create tooling that determines capacity for customers. You will work across hardware monitoring interfaces to fleet-wide data pipelines and dashboards, spanning Linux on ARM/x86, PCIe, Power, NIC, NVMe, and GPU subsystems.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Infra Systems Engineer: GPU & Accelerator Servers
AI Infra Systems Engineer: GPU & Accelerator Servers

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 129,000 - 175,000
Health insurance
401(k) matching
Paid time off
+1
Fleet-Scale Systems Engineer for GPU & AI Accelerators
Fleet-Scale Systems Engineer for GPU & AI Accelerators

Amazon • Denver (CO)

On-site
USD 129,000 - 175,000
Health insurance
401(k) matching
Paid time off
+1
AI Compute Fleet Engineer – GPU & Server Automation
AI Compute Fleet Engineer – GPU & Server Automation

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 149,000 - 201,000
Health insurance
401(k) matching
Paid time off
Automation Engineer, AI/ML Server Fleet Health & Predictive
Automation Engineer, AI/ML Server Fleet Health & Predictive

Amazon Data Services, Inc. • Cupertino (CA)

On-site
USD 140,000 - 190,000
Systems Development Engineer, AWS Generative AI & ML Servers
Systems Development Engineer, AWS Generative AI & ML Servers

Amazon Data Services, Inc. • Cupertino (CA)

On-site
USD 140,000 - 190,000
GPU Server Hardware Engineer for AI Systems
GPU Server Hardware Engineer for AI Systems

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 110,000 - 160,000
Senior AI/ML Fleet Automation Engineer
Senior AI/ML Fleet Automation Engineer

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 174,000 - 235,000
401(k) matching
Paid time off
Parental leave
+1
Cloud AI Systems Engineer - Fleet Health & Automation
Cloud AI Systems Engineer - Fleet Health & Automation

Amazon Web Services (AWS) • Austin (TX)

On-site
USD 129,000 - 175,000
Health insurance
RSUs (restricted stock units)
401(k) matching
+1
Manufacturing Systems Engineer: GPU/AI Server Automation
Manufacturing Systems Engineer: GPU/AI Server Automation

Amazon • Cupertino (CA)

On-site
USD 149,000 - 201,000
Health insurance
RSUs
401(k) matching
GPU Server Hardware Engineer for AI/ML
GPU Server Hardware Engineer for AI/ML

Amazon • Cupertino (CA)

On-site
USD 126,000 - 185,000
RSUs
Health benefits
401(k) matching