Manufacturing System Development Engineer, Cloud AI/ML/storage server teams

Amazon Web Services (AWS)

Seattle (WA)

On-site

USD 129,200 - 174,800

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Amazon Web Services (AWS) is seeking a Manufacturing Systems Development Engineer to own test software, tooling, and hardware debug for GPU-based server platforms at ODM/CM sites. You will debug line faults, automate tests, and build diagnostic tools to improve yield and throughput, bridging hardware design intent with manufacturing execution.

You will work across firmware, kernel, drivers, PCIe, power, and GPU subsystems, and travel ~25% to support on-site builds and ramp phases.

Qualifications

  • 2+ years of non-internship professional software development experience.
  • 1+ years designing or architecting new and existing systems.
  • Experience programming in at least one modern language (C++, C#, Java, Python, Go, PowerShell, Ruby).
  • 5+ years of software development experience in at least one modern language.
  • 3+ years debugging hardware systems on manufacturing lines.
  • Experience with Linux/Unix boot flow, kernel, drivers, OS diagnostics.
  • Willingness to travel ~25% domestically and internationally.

Responsibilities

  • Debug complex system-level failures on the manufacturing line across compute, storage, GPU, networking, power, and thermal domains.
  • Perform RCCA for yield detractors and test escapes.
  • Provide on-site ODM/CM support during critical builds and ramp.
  • Design and maintain manufacturing test software and diagnostic tools.
  • Build automation to reduce manual triage and improve first-pass yield.

Skills

Python
C++
C#
Java
Golang
PowerShell
Linux debugging
Hardware debugging
ODM/CM collaboration
CI/CD pipelines

Education

Bachelor's degree or equivalent

Tools

Oscilloscope
Logic Analyzer
CI/CD toolchains
NVMe & PCIe debugging

Job description

Description

Application deadline: Jul 27, 2026

Amazon Web Services (AWS) Hardware Engineering designs and delivers next-generation cloud infrastructure. Our team builds custom accelerator systems that power AI, machine learning, and compute workloads at global scale.

We are seeking a Manufacturing Systems Development Engineer to own manufacturing test software, diagnostic tooling, and hardware debug for GPU-based server platforms at our ODM/CM manufacturing sites. In this role, you will be on the manufacturing floor debugging complex system failures, developing automation to improve yield and throughput, and building diagnostic tools that enable root cause identification at the line. You will bridge the gap between hardware design intent and manufacturing execution — ensuring our platforms are testable, diagnosable, and launch with exceptional quality.

This role requires someone equally comfortable writing code and debugging hardware. You will develop test automation, build diagnostic frameworks, and personally troubleshoot failures spanning firmware, kernel, drivers, PCIe, power, and GPU subsystems — all in a fast-paced manufacturing environment. When something fails at the line, you are the person who figures out why.

Domestic and international travel (~25%)

Key job responsibilities

Manufacturing Debug & Root Cause Analysis

  • Debug complex system-level failures at the manufacturing line across compute, storage, GPU, networking, power, and thermal domains
  • Perform root cause analysis correlating across firmware, kernel, driver, PCIe, signal integrity, and physical layers to isolate faults
  • Troubleshoot Linux boot and runtime failures across x86 and ARM architectures, including NVMe, GPU, NIC, and accelerator subsystems
  • Drive Root Cause Corrective Action (RCCA) for yield detractors, test escapes, and recurring manufacturing failures
  • Provide on-site ODM/CM support during critical builds, EVT/DVT/PVT phases, and production ramp
Test Software & Automation Development
  • Design, develop, and maintain manufacturing test software and diagnostic tools deployed at ODM/CM lines
  • Build automation that reduces manual triage — enabling faster fault isolation and higher first-pass yield
  • Develop and optimize system-level test flows (BFT, functional test, stress test, burn-in) for GPU accelerator platforms
  • Build, manage, and deploy CI/CD pipelines for rapid deployment of test code to manufacturing environments
  • Write scalable, robust code in Python, C/C++, or Java to solve manufacturing test and debug challenges
Manufacturing Process & Quality
  • Define and improve manufacturing test strategy including coverage, duration, fixture requirements, and pass/fail criteria
  • Analyze test data and yield trends to identify systemic issues; drive design and process improvements
  • Collaborate on DFx reviews (DFT/DFM) to ensure new designs are testable and diagnosable at the manufacturing line
  • Develop diagnostic tooling requirements for ODM/CM enablement — ensuring partners can effectively screen and debug at scale
  • Research and implement automation techniques to improve manufacturing efficiency and reduce human intervention
Cross-Team Collaboration
  • Work across hardware design, firmware, qualification, and manufacturing engineering teams to close the loop between line failures and design improvements
  • Engage with ODMs and design partners on testability, diagnostic, and automation requirements during NPI
  • Collaborate with internal teams on GPU module integration, test coverage, and manufacturing debug procedures
  • Partner with fleet health teams to ensure manufacturing diagnostics align with production monitoring and field failure analysis
Basic Qualifications

  • 2+ years of non-internship professional software development experience
  • 1+ years of designing or architecting (design patterns, reliability and scaling) of new and existing systems experience
  • Experience programming with at least one modern language such as C++, C#, Java, Python, Golang, PowerShell, Ruby
  • 5+ years of software development experience with at least one modern language (Python, C/C++, Java)
  • 3+ years of experience debugging hardware systems — server, accelerator, storage, or high-tech platforms
  • Experience with Linux/Unix systems including boot flow, kernel, drivers, and OS-level diagnostics
  • Hands-on experience troubleshooting hardware failures at a manufacturing line or lab environment
  • Experience working with ODMs/CMs through product development and manufacturing lifecycle
  • Willingness to travel domestically and internationally (~25%), including extended on-site manufacturing support
Preferred Qualifications

  • Experience with GPU-based server or accelerator platform manufacturing and debug
  • Familiarity with server hardware architecture: PCIe topology, NVMe, BMC/IPMI, power delivery, thermal
  • Experience developing manufacturing test automation or diagnostic frameworks at scale
  • Experience with board-level debug (oscilloscope, logic analyzer)
  • Knowledge of firmware, BIOS, BMC, and their interaction with manufacturing test flows
  • Experience with manufacturing yield analysis, test optimization, and throughput improvement
  • Experience building CI/CD pipelines for test software deployment
  • Familiarity with telemetry, log correlation, and failure pattern analysis in manufacturing environments
Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status.

Los Angeles County applicants: Job duties for this position include: work safely and cooperatively with other employees, supervisors, and staff; adhere to standards of excellence despite stressful conditions; communicate effectively and respectfully with employees, supervisors, and staff to ensure exceptional customer service; and follow all federal, state, and local laws and Company policies. Criminal history may have a direct, adverse, and negative relationship with some of the material job duties of this position. These include the duties and responsibilities listed above, as well as the abilities to adhere to company policies, exercise sound judgment, effectively manage stress and work safely and respectfully with others, exhibit trustworthiness and professionalism, and safeguard business operations and the Company’s reputation. Pursuant to the Los Angeles County Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records.

Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit https://amazon.jobs/content/en/how-we-hire/accommodations for more information. If the country/region you’re applying in isn’t listed, please contact your Recruiting Partner.

The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at https://amazon.jobs/en/benefits.

USA, CA, Cupertino - 148,700.00 - 201,200.00 USD annually

USA, CO, Denver - 129,200.00 - 174,800.00 USD annually

USA, WA, Seattle - 129,200.00 - 174,800.00 USD annually


Company - Amazon Data Services, Inc.

Job ID: A10477150
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Manufacturing System Development Engineer, Cloud AI/ML/storage server teams
Manufacturing System Development Engineer, Cloud AI/ML/storage server teams

Amazon Web Services (AWS) • Denver (CO)

On-site
USD 129,000 - 175,000
Manufacturing System Development Engineer, Cloud AI/ML/storage server teams
Manufacturing System Development Engineer, Cloud AI/ML/storage server teams

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 148,000 - 202,000
Health insurance
401(k) matching
Paid time off
+1
Manufacturing hardware engineer, Cloud AI/ML/storage server teams
Manufacturing hardware engineer, Cloud AI/ML/storage server teams

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 136,000 - 184,000
Health insurance
401(k) matching
Sr. Manufacturing Hardware Engineer, AI/ML Server Development
Sr. Manufacturing Hardware Engineer, AI/ML Server Development

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 159,000 - 216,000
Sr. Manufacturing Hardware Engineer, AI/ML Server Development
Sr. Manufacturing Hardware Engineer, AI/ML Server Development

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 183,000 - 248,000
Sr. Manufacturing Hardware Engineer, AI/ML Server Development (AWS)
Sr. Manufacturing Hardware Engineer, AI/ML Server Development (AWS)

Amazon • Cupertino (CA)

On-site
USD 183,000 - 248,000
RSUs
Health insurance
401(k) matching
Manufacturing hardware engineer, Cloud AI/ML/storage server teams
Manufacturing hardware engineer, Cloud AI/ML/storage server teams

Amazon Web Services (AWS) • Denver (CO)

On-site
USD 136,000 - 184,000
Manufacturing hardware engineer, Cloud AI/ML/storage server teams
Manufacturing hardware engineer, Cloud AI/ML/storage server teams

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 157,000 - 213,000
Sr Mfg NPI Product Engineer, AI/ML Server Development - Manufacturing team (AWS)
Sr Mfg NPI Product Engineer, AI/ML Server Development - Manufacturing team (AWS)

Amazon • Cupertino (CA)

On-site
USD 183,000 - 248,000
Sr Mfg NPI Product Engineer, AI/ML Server Development - Manufacturing team
Sr Mfg NPI Product Engineer, AI/ML Server Development - Manufacturing team

Amazon Web Services (AWS) • Denver (CO)

On-site
USD 159,000 - 216,000