GPU Server Manufacturing Engineer - Automation & Debug

Amazon Web Services (AWS)

Seattle (WA)

On-site

USD 129,000 - 175,000

Full time

24 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health insurance
401(k) matching
Paid time off
Parental leave

Job summary

Amazon Data Services, Inc. in Seattle is seeking a Manufacturing Systems Development Engineer to own manufacturing test software, diagnostic tooling, and hardware debug for GPU-based server platforms. You will develop test automation, diagnose complex line failures, and build pipelines to speed fault isolation.

You will work across firmware, kernel, drivers, PCIe, and power subsystems, bridging design intent with manufacturing execution. Approximately 25% travel is required.

Qualifications

  • 2+ years of non-internship professional software development experience.
  • 1+ years of designing or architecting (design patterns, reliability and scaling) of new and existing systems.
  • Experience programming with at least one modern language such as C++, C#, Java, Python, Golang, PowerShell, Ruby.
  • 5+ years of software development experience with at least one modern language (Python, C/C++, Java).
  • 3+ years of experience debugging hardware systems — server, accelerator, storage, or high-tech platforms.
  • Experience with Linux/Unix systems including boot flow, kernel, drivers, and OS-level diagnostics.
  • Hands-on experience troubleshooting hardware failures at a manufacturing line or lab environment.
  • Experience working with ODMs/CMs through product development and manufacturing lifecycle.
  • Willingness to travel domestically and internationally (~25%), including extended on-site manufacturing support.

Responsibilities

  • Debug complex system-level failures at the manufacturing line across compute, storage, GPU, networking, power, and thermal domains.
  • Perform root cause analysis correlating across firmware, kernel, driver, PCIe, signal integrity, and physical layers to isolate faults.
  • Troubleshoot Linux boot and runtime failures across x86 and ARM architectures, including NVMe, GPU, NIC, and accelerator subsystems.
  • Drive Root Cause Corrective Action (RCCA) for yield detractors, test escapes, and recurring manufacturing failures.
  • Provide on-site ODM/CM support during critical builds, EVT/DVT/PVT phases, and production ramp.
  • Design, develop, and maintain manufacturing test software and diagnostic tools deployed at ODM/CM lines.
  • Build automation that reduces manual triage — enabling faster fault isolation and higher first-pass yield.
  • Develop and optimize system-level test flows (BFT, functional test, stress test, burn-in) for GPU accelerator platforms.

Skills

Python
C++
Java
Linux
C#
Golang
PowerShell
Ruby
ODM/CM collaboration
Debugging hardware systems

Job description

Amazon Data Services, Inc. in Seattle is seeking a Manufacturing Systems Development Engineer to own manufacturing test software, diagnostic tooling, and hardware debug for GPU-based server platforms. You will develop test automation, diagnose complex line failures, and build pipelines to speed fault isolation.

You will work across firmware, kernel, drivers, PCIe, and power subsystems, bridging design intent with manufacturing execution. Approximately 25% travel is required.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

GPU Server Manufacturing Debug & Test Automation Engineer
GPU Server Manufacturing Debug & Test Automation Engineer

Amazon Web Services (AWS) • Denver (CO)

On-site
USD 129,200 - 174,800
GPU Server Manufacturing Systems Engineer – Automation & Debugging
GPU Server Manufacturing Systems Engineer – Automation & Debugging

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 148,000 - 202,000
Health insurance
401(k) matching
Paid time off
+1
Manufacturing System Development Engineer, Cloud AI/ML/storage server teams
Manufacturing System Development Engineer, Cloud AI/ML/storage server teams

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 148,000 - 202,000
Health insurance
401(k) matching
Paid time off
+1
Manufacturing System Development Engineer, Cloud AI/ML/storage server teams
Manufacturing System Development Engineer, Cloud AI/ML/storage server teams

Amazon Web Services (AWS) • Denver (CO)

On-site
USD 129,200 - 174,800
Manufacturing System Development Engineer, Cloud AI/ML/storage server teams
Manufacturing System Development Engineer, Cloud AI/ML/storage server teams

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 129,000 - 175,000
Health insurance
401(k) matching
Paid time off
+1
Senior Manufacturing Hardware Engineer, GPU Accelerators
Senior Manufacturing Hardware Engineer, GPU Accelerators

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 159,000 - 215,000
Senior GPU Accelerator Manufacturing Engineer (NPI & Debug)
Senior GPU Accelerator Manufacturing Engineer (NPI & Debug)

Amazon • Cupertino (CA)

On-site
USD 183,000 - 248,000
RSUs
Health insurance
401(k) matching
GPU Server Test Automation Engineer
GPU Server Test Automation Engineer

HireReady Partners • Fort Worth (TX)

On-site
USD 85,000 - 125,000
Competitive salary
Health, dental, and vision insurance
401(k) retirement plan with company匹配
+2
NPI Manufacturing Engineer - AI/ML GPU Servers
NPI Manufacturing Engineer - AI/ML GPU Servers

Amazon • Cupertino (CA)

On-site
USD 183,000 - 248,000
Sr. Manufacturing Hardware Engineer, AI/ML Server Development
Sr. Manufacturing Hardware Engineer, AI/ML Server Development

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 159,000 - 215,000