GPU Server Manufacturing Systems Engineer – Automation & Debugging

Amazon Web Services (AWS)

Cupertino (CA)

On-site

USD 148,700 - 201,200

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health insurance
401(k) matching
Paid time off
Parental leave

Job summary

Amazon Data Services, Inc. in Cupertino is seeking a Manufacturing Systems Development Engineer to own manufacturing test software, diagnostics, and hardware debugging for GPU server platforms at ODM/CM sites.

You will debug line failures, automate tests, and build diagnostic tools to enable rapid fault isolation and high first‑pass yield. You will bridge hardware design intent with manufacturing execution, tackle firmware, kernel, drivers, PCIe, power, and GPU subsystems, and travel up to 25%

Qualifications

  • 2+ years of non-internship professional software development experience.
  • 1+ year of designing or architecting (design patterns, reliability and scaling) of new and existing systems experience.
  • Experience programming with at least one modern language such as C++, C#, Java, Python, Golang, PowerShell, Ruby.
  • 5+ years of software development experience with at least one modern language (Python, C/C++, Java).
  • 3+ years of experience debugging hardware systems — server, accelerator, storage, or high-tech platforms.
  • Experience with Linux/Unix systems including boot flow, kernel, drivers, and OS-level diagnostics.
  • Hands‑on experience troubleshooting hardware failures at a manufacturing line or lab environment.
  • Experience working with ODMs/CMs through product development and manufacturing lifecycle.
  • Willingness to travel domestically and internationally (~25%), including extended on‑site manufacturing support.

Responsibilities

  • Debug complex system-level failures at the manufacturing line across compute, storage, GPU, networking, power, and thermal domains.
  • Perform root cause analysis correlating across firmware, kernel, driver, PCIe, signal integrity, and physical layers to isolate faults.
  • Troubleshoot Linux boot and runtime failures across x86 and ARM architectures, including NVMe, GPU, NIC, and accelerator subsystems.
  • Drive Root Cause Corrective Action (RCCA) for yield detractors, test escapes, and recurring manufacturing failures.
  • Provide on-site ODM/CM support during critical builds, EVT/DVT/PVT phases, and production ramp.
  • Design, develop, and maintain manufacturing test software and diagnostic tools deployed at ODM/CM lines.
  • Build automation that reduces manual triage—enabling faster fault isolation and higher first-pass yield.
  • Develop and optimize system-level test flows (BFT, functional test, stress test, burn-in) for GPU accelerator platforms.
  • Build, manage, and deploy CI/CD pipelines for rapid deployment of test code to manufacturing environments.
  • Write scalable, robust code in Python, C/C++, or Java to solve manufacturing test and debug challenges.

Skills

Python
C/C++
Linux
Debugging
Manufacturing experience
Travel 25%

Tools

PCIe analysis tools
Oscilloscope

Job description

Amazon Data Services, Inc. in Cupertino is seeking a Manufacturing Systems Development Engineer to own manufacturing test software, diagnostics, and hardware debugging for GPU server platforms at ODM/CM sites.

You will debug line failures, automate tests, and build diagnostic tools to enable rapid fault isolation and high first‑pass yield. You will bridge hardware design intent with manufacturing execution, tackle firmware, kernel, drivers, PCIe, power, and GPU subsystems, and travel up to 25%

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

GPU Server Manufacturing Debug & Test Automation Engineer
GPU Server Manufacturing Debug & Test Automation Engineer

Amazon Web Services (AWS) • Denver (CO)

On-site
USD 129,000 - 175,000
Manufacturing Systems Engineer - GPU Server Debug
Manufacturing Systems Engineer - GPU Server Debug

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 129,000 - 175,000
Senior Manufacturing Hardware Engineer - GPU Accelerators
Senior Manufacturing Hardware Engineer - GPU Accelerators

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 183,000 - 248,000
Senior GPU Accelerator Manufacturing Engineer (NPI & Debug)
Senior GPU Accelerator Manufacturing Engineer (NPI & Debug)

Amazon • Cupertino (CA)

On-site
USD 183,000 - 248,000
RSUs
Health insurance
401(k) matching
Manufacturing System Development Engineer, Cloud AI/ML/storage server teams
Manufacturing System Development Engineer, Cloud AI/ML/storage server teams

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 148,000 - 202,000
Health insurance
401(k) matching
Paid time off
+1
Manufacturing System Development Engineer, Cloud AI/ML/storage server teams
Manufacturing System Development Engineer, Cloud AI/ML/storage server teams

Amazon Web Services (AWS) • Denver (CO)

On-site
USD 129,000 - 175,000
Manufacturing Hardware Engineer, GPU Accelerator Systems
Manufacturing Hardware Engineer, GPU Accelerator Systems

Amazon Web Services (AWS) • Denver (CO)

On-site
USD 136,000 - 184,000
GPU Accelerator Manufacturing Engineer – NPI & Fleet Quality
GPU Accelerator Manufacturing Engineer – NPI & Fleet Quality

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 157,000 - 213,000
GPU Accelerator Manufacturing Engineer - Line Debug & Scale
GPU Accelerator Manufacturing Engineer - Line Debug & Scale

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 136,000 - 184,000
Health insurance
401(k) matching
Manufacturing System Development Engineer, Cloud AI/ML/storage server teams
Manufacturing System Development Engineer, Cloud AI/ML/storage server teams

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 129,000 - 175,000