Lead Engineer for Manufacturing and Datacenter Lab, Trainium Manufacturing, Quality and Reliability

Amazon Inc.

Austin (TX)

On-site

USD 159,000 - 215,000

Full time

7 hours ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Health insurance
401(k) matching
RSUs
Paid time off
Parental leave

Job summary

Amazon seeks a Lead Engineer for Manufacturing and Datacenter Lab in Austin to bridge manufacturing outcomes with datacenter performance. You will build a preparedness lab, define assembly/repair recipes, and drive data-driven improvements across ODM/CM ecosystems.

The role requires deep experience in manufacturing design, reliability engineering, and cross-functional leadership, with a focus on ML-driven analytics and robust test strategies for scalable server systems.

Qualifications

  • Experience carrying design concepts through exploration, development, and into deployment or mass production.
  • Experience communicating requirements and designs to customers, technical teams and management.
  • Experience presenting results to senior leadership and influencing decisions.
  • BS or MS in Electrical/Mechanical/Computer/Industrial Engineering or related field.
  • 8+ years in Manufacturing/ Test/ Quality/ Reliability or Datacenter Infrastructure Engineering.
  • 7+ years in cross-functional engineering teams.
  • Experience with AI/ML acceleration systems, HPC servers, or complex multi-rack systems.
  • Proven track record of stable, performant hardware meeting cost/quality targets.
  • Experience with System Mechanical & Thermal design for air/liquid cooling.
  • Strong problem solving to resolve cross-domain issues.

Responsibilities

  • Own operational production performance of Trainium systems across lifecycle.
  • Design and build preparedness lab replicating datacenter conditions.
  • Define and drive assembly/repair recipes in the lab.
  • Ensure all test flows are regressed in the lab before deployment.
  • Influence hardware design strategy for DFM/DFR/DFT based on field data.
  • Establish analytics tying test data to datacenter performance using ML.
  • Build and mentor cross-functional team across manufacturing, test, quality, reliability.
  • Collaborate with AWS teams to translate field learnings into process improvements.
  • Drive continuous improvement to reduce failure rates and degradation.
  • Develop or adapt manufacturing process at ODM/CM with fixture and test specs.

Skills

8+ years in Mfg Eng / Reliability
Cross-functional collaboration
Problem-solving
Python
Linux
Data analysis
Root cause analysis
DFM/DFR/DFT
Design for manufacturability
AI/ML awareness

Education

BS or MS in Electrical/Mechanical/Computer/Industrial Eng
Masters degree preferred

Tools

Python
Bash
Linux

Job description

Lead Engineer for Manufacturing and Datacenter Lab, Trainium Manufacturing, Quality and Reliability

Within the Trainium Manufacturing Quality & Reliability (TRN MQR) organization, we are establishing a critical new function that bridges manufacturing outcomes with datacenter operational performance. We are seeking a talented and motivated Manufacturing & Datacenter Preparedness Lab Leader to build and lead this strategic capability in Austin, Texas.

This role will report to the leader of Trainium Manufacturing Quality & Reliability and serve as the essential feedback loop between our ODM/JDM/CM manufacturing operations and AWS datacenter fleet performance. You will establish and operate a specialized preparedness lab focused on analyzing datacenter performance of manufactured Trainium systems to identify root causes of field rework and repairs, feeding critical insights back into manufacturing processes, test strategies, and design improvements.

You will participate in the early phase of manufacturing line development for our next generation servers and racks to improve our manufacturing flows informing system design, manufacturing, and fleet operations. You will manage early lifecycle changes, identify initial product quality improvements, and drive to technical root cause in supplier quality activities. The candidate will have experience in design or manufacturing and is capable of making wide-ranging business decisions on behalf of the organization.

You'll join a diverse team working across Manufacturing Engineering, Manufacturing Test Engineering, and Quality & Reliability Engineering. You'll collaborate with people across AWS Data Center Engineering, Hardware Design, ODM/JDM/CM partners, and datacenter operations teams to help us deliver the highest standards for safety and reliability while providing seemingly infinite capacity at the lowest possible cost for our customers. And you'll experience an inclusive culture that welcomes bold ideas and empowers you to own them to completion.

Key job responsibilities
  • Own operational production performance of Trainium systems across entire product lifecycle from manufacturing through datacenter deployment and fleet operations
  • Design and build preparedness lab replicating datacenter conditions for assembly, repair and system testing
  • Define and drive assembly and repair recipes in the manufacturing lab as the baseline prior to high volume manufacturing and datacenter deployment.
  • Ensure all manufacturing and datacenter test flows are regressed in the manufacturing lab prior to deployment.
  • Influence hardware design strategy for Design for Manufacturing (DFM), Design for Reliability (DFR), and Design for Test (DFT) based on field failure analysis.
  • Establish data-driven analytics frameworks connecting manufacturing test data to datacenter performance, leveraging ML techniques to predict field failures.
  • Build and mentor cross-functional team spanning manufacturing, test, quality, and reliability engineering; perform technical promotion assessments as force multiplier.
  • Collaborate with AWS datacenter operations teams to understand failure modes, repair patterns, and operational challenges firsthand; translate operator insights and field learnings into actionable manufacturing process improvements and design changes.
  • Drive continuous improvement reducing failure rates and lifecycle degradation through rapid feedback loops.
  • Develop or adapt manufacturing process at the ODM and CM, including defining fixture requirements, critical assembly requirements, test methodology, signal integrity, power and heat management requirement.
About the team

Annapurna Labs is a wholly owned subsidiary of AWS, focused on developing custom silicon and servers including the Nitro(K2), Graviton, Inferentia, and Trainium families of processors.

Machine Learning Annapurna functions as a vertically integrated team including software, firmware, hardware, and silicon design in a single organization.

We are the Trainium Servers and Systems organization under MLA focused on Hardware Development, Software Development, Fleet Ops Systems, and Manufacturing, Quality, and Reliability.

This position is in the Manufacturing, Quality and Reliability team.

Basic Qualifications
  • Experience carrying design concepts through exploration, development, and into deployment or mass production
  • Experience communicating with customers, technical, regulatory, business teams, and management to collect requirements, describe product features, and technical designs
  • Experience communicating results to senior leadership, or experience communicating complex information and solutions to senior stakeholders and influencing decisions
  • BS or MS degree in Electrical Engineering, Mechanical Engineering, Computer Engineering, Industrial Engineering, or related technical fields
  • 8+ years industry experience in one or more of the following: Manufacturing Engineering, Test Engineering, Quality Engineering, Reliability Engineering, or Datacenter Infrastructure Engineering
  • 7+ years working directly with engineering teams in cross-functional environments
  • Experience with AI/ML acceleration systems, high-performance computing servers, or complex multi-rack systems
  • Demonstrated track record delivering stable, performant hardware solutions meeting cost and quality targets
  • Experience with System Mechanical & Thermal design for air-cooled and liquid-cooled systems
  • Strong problem-solving capabilities to isolate, define, and resolve complex problems spanning manufacturing quality and field reliability
  • Experience with root cause analysis methodologies (8D, 5-Why, Fishbone, FMEA) and implementing corrective/preventive actions
  • Proficiency in data analysis tools, statistical methods, and programming (Python, Bash, Shell script, Linux)
  • Experience working with ODMs, JDMs, component vendors, and internal design teams on cross-boundary triaging, debugging, and resolving issues
  • Experience in Design for Manufacturing (DFM), also known as Design for Manufacturability, a product design approach that focuses on optimizing the ease and cost of manufacturing a product
  • Can be given complex hardware engineering problem to solve and design project strategy that splits work appropriately for parallel development
Preferred Qualifications
  • Experience in management of datacenter operations, facility engineering operations, information technology critical environment facilities, advanced high volume manufacturing, datacenter build-outs and scaling, or similar fields
  • Knowledge of AWS services including compute, storage, networking, security, databases, machine learning, and serverless technologies
  • Knowledge of the electrical and mechanical systems involved in critical data center operations including systems such as feeders, transformers, generators, switchgear, UPS systems, ATS units, PDU units, chillers, pumps, air handling units, and CRAC units
  • Masters Degree in Electrical Engineering, Mechanical Engineering, Computer Engineering, Industrial Engineering, or related technical fields
  • Hands-on design experience with enterprise hardware design, sled level design, and rack level designs
  • Track record of implementing data-driven process improvements that measurably reduced field failure rates or improved manufacturing yield
  • Experience with liquid cooling systems, direct-to-chip cooling, coolant distribution units (CDUs), and thermal management for AI/ML workloads

Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status.

Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit https://amazon.jobs/content/en/how-we-hire/accommodations for more information. If the country/region you’re applying in isn't listed, please contact your Recruiting Partner.

Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status.

Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit https://amazon.jobs/content/en/how-we-hire/accommodations for more information. If the country/region you’re applying in isn't listed, please contact your Recruiting Partner.

USA, TX, Austin - 159,200.00 - 215,300.00 USD annually

The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at https://amazon.jobs/en/benefits .

The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at https://amazon.jobs/en/benefits .

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Sr PCBA Manufacturing Engineer, Trainium Manufacturing, Quality and Reliability (AWS)
Sr PCBA Manufacturing Engineer, Trainium Manufacturing, Quality and Reliability (AWS)

Amazon • Austin (TX)

On-site
USD 159,000 - 215,000
Sr Manager, AI Systems Quality & Reliability , Annapurna AI Servers and Systems (AWS)
Sr Manager, AI Systems Quality & Reliability , Annapurna AI Servers and Systems (AWS)

Amazon • Austin (TX)

On-site
USD 208,000 - 282,000
Sr Manager, AI Systems Quality & Reliability , Annapurna AI Servers and Systems
Sr Manager, AI Systems Quality & Reliability , Annapurna AI Servers and Systems

Amazon Web Services (AWS) • Seattle (WA)

On-site
USD 208,000 - 282,000
Sr Manager, AI Systems Quality & Reliability , Annapurna AI Servers and Systems
Sr Manager, AI Systems Quality & Reliability , Annapurna AI Servers and Systems

Amazon Web Services (AWS) • Cupertino (CA)

On-site
USD 240,000 - 324,000
Health insurance
RSUs
Sr. Quality & Reliability Engineer, Annapurna AI Servers and Systems (AWS)
Sr. Quality & Reliability Engineer, Annapurna AI Servers and Systems (AWS)

Amazon • Austin (TX)

On-site
USD 159,000 - 215,000
Health insurance
RSUs
Paid time off
Sr. Manufacturing Engineer – Test Support , Hardware Engineering - Manufacturing
Sr. Manufacturing Engineer – Test Support , Hardware Engineering - Manufacturing

Amazon Web Services (AWS) • Salt Lake City (UT)

On-site
USD 151,000 - 205,000
Quality & Reliability Engineer, Trainium Manufacturing, Quality & Reliability
Quality & Reliability Engineer, Trainium Manufacturing, Quality & Reliability

Amazon • Cupertino (CA), Northern (KY)

Hybrid
USD 157,000 - 213,000
Sr. Manufacturing Engineer – Test Support , Hardware Engineering - Manufacturing (AWS)
Sr. Manufacturing Engineer – Test Support , Hardware Engineering - Manufacturing (AWS)

Amazon Inc. • Hebron Estates (KY)

On-site
USD 151,000 - 205,000
Quality & Reliability Engineer, Trainium Manufacturing, Quality & Reliability
Quality & Reliability Engineer, Trainium Manufacturing, Quality & Reliability

Amazon Web Services (AWS) • Austin (TX)

On-site
USD 136,000 - 184,000
System Development Engineer, Mfg Bring Up and Test, AWS Mainstream Compute
System Development Engineer, Mfg Bring Up and Test, AWS Mainstream Compute

RiseMe • Seattle (WA)

On-site
USD 129,000 - 175,000
Health insurance
RSUs / Stock options
401(k) matching
+1