Sr. Infrastructure Reliability Engineer, Infrastructure Reliability & Quality

Amazon Inc.

Singapore

On-site

SGD 150,000 - 190,000

Full time

9 hours ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Amazon Asia-Pacific Resources Private Limited in Singapore seeks a Senior Infrastructure Reliability Engineer to drive reliability risk identification for datacenter equipment and lead root-cause analysis of critical failures.

You will collaborate with suppliers and internal teams to define risk plans, develop system-level reliability models, and monitor field performance, implementing corrective and preventive actions. Travel internationally may be required.

Qualifications

  • Bachelor's degree in engineering or related field.
  • Experience in data center design, construction, operations, or facility maintenance.
  • Experience in mission-critical facilities including data centers or power generation.

Responsibilities

  • Drive reliability risk identification, assessment and mitigation for datacenter infrastructure equipment.
  • Perform root cause analysis of critical equipment failures and drive continuous improvements.
  • Collaborate with suppliers and internal partners on product specifications and risk plans.
  • Develop datacenter system-level reliability models and related risk analyses.
  • Monitor field performance and drive corrective/preventive actions; oversee vendor audits.
  • Travel internationally as needed.

Skills

Data center design
Operations
Facility maintenance
Engineering

Education

Bachelor's degree in engineering
Master's degree or above

Job description

Sr. Infrastructure Reliability Engineer, Infrastructure Reliability & Quality

Job ID: 10574836 | Amazon Asia-Pacific Resources Private Limited (Singapore)

AWS Infrastructure Services owns the design, planning, delivery, and operation of all AWS global infrastructure. In other words, we’re the people who keep the cloud running. We support all AWS data centers and all of the servers, storage, networking, power, and cooling equipment that ensure our customers have continual access to the innovation they rely on. We work on the most challenging problems, with thousands of variables impacting the supply chain — and we’re looking for talented people who want to help.

You’ll join a diverse team of software, hardware, and network engineers, supply chain specialists, security experts, operations managers, and other vital roles. You’ll collaborate with people across AWS to help us deliver the highest standards for safety and security while providing seemingly infinite capacity at the lowest possible cost for our customers. And you’ll experience an inclusive culture that welcomes bold ideas and empowers you to own them to completion.

As a Senior Infrastructure Reliability Engineer you will be proactively driving the reliability risk identification, assessment and mitigation for datacenter infrastructure equipment (Example: Busduct, Tapoff Box, MV Transformers, LV SWGR, Breakers, UPS, Power Distribution Cabinets etc.). You will also be responsible for root cause analysis of critical equipment failures and drive the continuous improvements to improve datacenter availability for AWS customers. You will work closely with both internal and outside partners including suppliers to drive key aspects of product specification, risk identification plan and execution. You must be ownership minded, independent, action and results oriented to succeed in an open collaborative environment.

The candidate should have experience in using Physics‑of‑Failure based approach to develop and implement both analytical and empirical approaches for product quality/reliability risk identification and assessment during product design, manufacture as well as deployment stages. The individual should be able to drive AWS application‑specific requirements in carrying out both lifecycle environmental and operational stress driven risk analysis, including thermal, electrical, chemical and mechanical stresses so to identify overstress and fatigue‑related product weaknesses. Candidate should be capable of evaluating not only product design quality/reliability risks, but also have the skills and experiences in assessing electronics manufacture process related quality/reliability issues. Knowledge of statistical techniques and models is required to analyze test as well as field data.

At the component level, the individual will drive critical component identification and the associated vendor selection and qualification requirements. The candidate will be expected to use knowledge of process capability for electronic component production as well as system‑level performance requirements to establish critical to quality and reliability metrics.

At the system level, the individual will develop datacenter system level reliability model and related reliability quantification and risk analysis for datacenter configuration optimization. The candidate will be expected to be familiar with system reliability engineering tools, such as reliability block diagram, statistical modeling and data analytics.

During sustaining stage, candidate will be responsible for monitoring product performance in the field and will be responsible to drive root cause analysis of any critical failures and the associated corrective and preventive actions. The individual should also be able to drive effective vendor auditing and quarterly review process to drive the continuous improvements of datacenter availability.

The successful candidate should be considered as an expert in the reliability engineering field and have a proven track record of success in not only product reliability leadership, as well as business negotiations and program management. Strong skill‑set in problem analysis and solving as well as communication and vendor management are necessary. Candidates should also be able to travel internationally.

About the team

AWS values diverse experiences. Even if you do not meet all of the preferred qualifications and skills listed in the job description, we encourage candidates to apply. If your career is just starting, hasn’t followed a traditional path, or includes alternative experiences, don’t let it stop you from applying.

Amazon Web Services (AWS) is the world’s most comprehensive and broadly adopted cloud platform. We pioneered cloud computing and never stopped innovating — that’s why customers from the most successful startups to Global 500 companies trust our robust suite of products and services to power their businesses.

AWS values diverse experiences. Even if you do not meet all of the preferred qualifications and skills listed in the job description, we encourage candidates to apply. If your career is just starting, hasn’t followed a traditional path, or includes alternative experiences, don’t let it stop you from applying.

We value work‑life harmony. Achieving success at work should never come at the expense of sacrifices at home, which is why we strive for flexibility as part of our working culture. When we feel supported in the workplace and at home, there’s nothing we can’t achieve in the cloud.

Here at AWS, it’s in our nature to learn and be curious. Our employee‑led affinity groups foster a culture of inclusion that empower us to be proud of our differences. Ongoing events and learning experiences, including our Conversations on Race and Ethnicity (CORE) and AmazeCon (gender diversity) conferences, inspire us to never stop embracing our uniqueness.

We’re continuously raising our performance bar as we strive to become Earth’s Best Employer. That’s why you’ll find endless knowledge‑sharing, mentorship and other career‑advancing resources here to help you develop into a better‑rounded professional.

Basic Qualifications
  • 6+ years of data center design, construction, operations, or facility maintenance experience
  • 6+ years of industrial or commercial engineering in mission critical facilities including but not limited to: data centers, power generation or oil and gas facilities experience
  • Knowledge of critical data center mechanical and electrical equipment
  • Bachelor's degree in engineering or a related technical field
  • Experience in industrial or commercial engineering in mission critical facilities including but not limited to: data centers, power generation or oil and gas facilities
  • Experience in data center design, construction, operations, or facility maintenance
Preferred Qualifications
  • Experience carrying design concepts through exploration, development, and into deployment or mass production
  • Experience reading and interpreting construction drawings and specifications, or experience reading and writing procedures, technical documents, and engineering drawings
  • Master's degree or above in Electrical Engineering, Mechanical Engineering, Technology, or Reliability Engineering

Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit https://amazon.jobs/content/en/how-we-hire/accommodations for more information. If the country/region you’re applying in isn’t listed, please contact your Recruiting Partner.

Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status. Veterans, military spouses, and people with disabilities are encouraged to apply.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Data Center Engineering Ops Engineer, DCEO
Data Center Engineering Ops Engineer, DCEO

Amazon • Singapore

On-site
SGD 80,000 - 120,000
Senior Technical Program Manager, Infrastructure Reliability and Quality (IRQ)
Senior Technical Program Manager, Infrastructure Reliability and Quality (IRQ)

Amazon Inc. • Singapore

On-site
SGD 180,000 - 280,000
DCO Tech 3
DCO Tech 3

Amazon Inc. • Singapore

On-site
SGD 32,000 - 52,000
Data Center Operation Technician, Data Center Operations
Data Center Operation Technician, Data Center Operations

Amazon Inc. • Singapore

On-site
SGD 42,000 - 72,000
Technical Infra Program Manager, Data Center Delivery (AWS)
Technical Infra Program Manager, Data Center Delivery (AWS)

Amazon Inc. • Singapore

On-site
SGD 150,000 - 230,000
Physical Security Architect, Data Center Engineering
Physical Security Architect, Data Center Engineering

Amazon Web Services (AWS) • Singapore

On-site
SGD 110,000 - 170,000
Technical Infra Program Manager, Data Center Delivery
Technical Infra Program Manager, Data Center Delivery

Sourceo • Singapore

On-site
SGD 140,000 - 200,000
DCEO Engineer, Data Center Engineering Operations
DCEO Engineer, Data Center Engineering Operations

Amazon Web Services (AWS) • Singapore

On-site
SGD 90,000 - 120,000
Physical Security Architect, Data Center Engineering
Physical Security Architect, Data Center Engineering

Amazon • Singapore

On-site
SGD 120,000 - 180,000
Controls Design Engineer (EPMS), CDE
Controls Design Engineer (EPMS), CDE

Amazon Web Services (AWS) • Singapore

On-site
SGD 120,000 - 180,000