Sr. Infrastructure Reliability Engineer, Infrastructure Reliability & Quality
Job ID: 10455491 | Amazon Data Services, Inc.
AWS Infrastructure Services owns the design, planning, delivery, and operation of all AWS global infrastructure. We support all AWS data centers and all of the servers, storage, networking, power, and cooling equipment that ensure our customers have continual access to the innovation they rely on. You’ll join a diverse team of software, hardware, and network engineers, supply chain specialists, security experts, operations managers, and other vital roles to deliver safety and security while providing high capacity at low cost.
Responsibilities
- Proactively drive reliability risk identification, assessment, and mitigation for datacenter infrastructure equipment (e.g., LV Generator, MV Transformers, UPS, HV Transformers).
- Perform root cause analysis of critical equipment failures and drive continuous improvements to increase datacenter availability.
- Collaborate with internal and external partners, including suppliers, to drive product specification, risk identification plans, and execution.
- Use physics‑of‑failure approaches to develop and implement analytical and empirical methods for product quality and reliability risk identification during design, manufacture, and deployment.
- Conduct lifecycle environmental and operational stress‑driven risk analysis, including thermal, electrical, chemical, and mechanical stresses.
- Evaluate product design, manufacturing process, and electronics manufacturing quality‑related issues.
- Use statistical techniques and models to analyze test and field data.
- Identify critical components and lead vendor selection and qualification.
- Develop datacenter system‑level reliability models and perform reliability quantification and risk analysis for configuration optimization.
- Monitor product performance during the sustaining stage, perform root cause analysis of failures, and drive corrective and preventive actions.
- Execute vendor audits and quarterly review processes to improve datacenter availability.
- Travel within the US and internationally as required.
Basic Qualifications
- 6+ years of data center design, construction, operations, or facility maintenance experience.
- 6+ years of industrial or commercial engineering in mission‑critical facilities such as data centers, power generation, or oil and gas facilities.
- Knowledge of critical data center mechanical and electrical equipment.
- Experience in data center design, construction, operations, or facility maintenance.
- Bachelor’s degree in Electrical Engineering, Mechanical Engineering, or a related field.
Preferred Qualifications
- Experience carrying design concepts through exploration, development, and deployment or mass production.
- Experience reading, interpreting, and creating construction drawings, specifications, and submittal documents.
- Master’s degree in Reliability Engineering, Physics, Electrical, Mechanical, or Materials Engineering, or a related field.
- 8+ years of work experience in reliability risk identification and assessment from component to system level, applying analytical, experimental, and statistical approaches.
- Experience with proactive and effective reliability approaches in a cost‑effective manner throughout product design, manufacture, and deployment stages.
Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status.
Benefits
The base salary range for this position is $136,600.00 - $184,800.00 USD annually. The package includes sign‑on payments and restricted stock units (RSUs). Benefits include health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance, Supplemental life plans), EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage, 401(k) matching, paid time off, and parental leave. Learn more at https://amazon.jobs/en/benefits.