AI Infra Systems Engineer: GPU & Accelerator Servers

Amazon Web Services (AWS)

Seattle (WA)

On-site

USD 129,000 - 175,000

Full time

9 days ago
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Health insurance
401(k) matching
Paid time off
Parental leave

Job summary

Amazon Development Center U.S., Inc. is seeking a Systems Development Engineer to own health and development of server platforms at worldwide fleet scale.

You will build automation, analyze telemetry across thousands of hosts, and create tooling that determines capacity availability for AI workloads. Responsibilities span hardware monitoring interfaces through fleet-wide data pipelines and dashboards, collaborating with software, hardware, and vendor teams to deliver scalable, reliable compute

Qualifications

  • 2+ years of professional software development experience
  • 1+ years of designing or architecting systems
  • 3+ years of hands-on systems engineering experience in networking, storage, OS
  • Knowledge of systems engineering fundamentals (networking, storage, OS)
  • Experience with at least one modern language (C++, C#, Java, Python, Golang, PowerShell, Ruby)

Responsibilities

  • Analyze fleet telemetry and logs to identify failure patterns and quantify customer impact
  • Develop automation for hardware test, firmware qualification, and capacity recovery workflows
  • Debug and resolve Linux boot/run issues across architectures (PCIe, Power, NIC, NVMe, GPU)
  • Build scalable automation solutions using Python/Java or similar languages

Skills

Software development
System design
OS and kernel knowledge
Programming languages

Tools

PowerShell
Python
Ruby
Linux

Job description

Amazon Development Center U.S., Inc. is seeking a Systems Development Engineer to own health and development of server platforms at worldwide fleet scale.

You will build automation, analyze telemetry across thousands of hosts, and create tooling that determines capacity availability for AI workloads. Responsibilities span hardware monitoring interfaces through fleet-wide data pipelines and dashboards, collaborating with software, hardware, and vendor teams to deliver scalable, reliable compute

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Systems Engineer, GPU & AI Accelerator Servers
Systems Engineer, GPU & AI Accelerator Servers

Amazon • Cupertino (CA)

On-site
USD 149,000 - 201,000
Health insurance (medical, dental, and
Stock options or RSUs
401(k) matching
Fleet-Scale Systems Engineer for GPU & AI Accelerators
Fleet-Scale Systems Engineer for GPU & AI Accelerators

Amazon • Denver (CO)

On-site
USD 129,000 - 175,000
Health insurance
401(k) matching
Paid time off
+1
AI Compute Fleet Engineer – GPU & Server Automation
AI Compute Fleet Engineer – GPU & Server Automation

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 149,000 - 201,000
Health insurance
401(k) matching
Paid time off
GPU Server Hardware Engineer for AI Systems
GPU Server Hardware Engineer for AI Systems

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 110,000 - 160,000
Cloud AI Hardware Architect — GPU Server Systems
Cloud AI Hardware Architect — GPU Server Systems

Amazon Web Services (AWS) • Denver (CO)

On-site
USD 159,000 - 215,000
Cloud AI Systems Engineer - Fleet Health & Automation
Cloud AI Systems Engineer - Fleet Health & Automation

Amazon Web Services (AWS) • Austin (TX)

On-site
USD 129,000 - 175,000
Health insurance
RSUs (restricted stock units)
401(k) matching
+1
GPU Server Hardware Engineer for AI/ML
GPU Server Hardware Engineer for AI/ML

Amazon • Cupertino (CA)

On-site
USD 126,000 - 185,000
RSUs
Health benefits
401(k) matching
Senior AI/ML Fleet Automation Engineer
Senior AI/ML Fleet Automation Engineer

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 174,000 - 235,000
401(k) matching
Paid time off
Parental leave
+1
Cloud Hardware Engineer — AI/ML Server Fleet
Cloud Hardware Engineer — AI/ML Server Fleet

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 159,000 - 215,000
Cloud Hardware Engineer for Accelerated AI Servers
Cloud Hardware Engineer for Accelerated AI Servers

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 110,000 - 160,000
Health insurance
401(k) matching
Paid time off