System Hardware Reliability Engineer

Google Inc.

Sunnyvale (CA)

On-site

USD 188,000 - 274,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Google Sunnyvale is hiring a System Hardware Reliability Engineer to own physics-based health management and predictive analytics for high-density compute hardware. You will define reliability standards, design tests, and oversee analysis while collaborating with product and design teams.

The role emphasizes physics-of-failure modeling integrated with ML, forecasting remaining useful life and optimizing cooling and power strategies across the global fleet.

Qualifications

  • Bachelor’s degree in Reliability Engineering or related field.
  • 8+ years applying Design for Reliability in consumer electronics.
  • 8+ years hardware reliability engineering and predictive analytics.

Responsibilities

  • Define standards, specify tests, and supervise execution and failure analysis.
  • Develop PHM models and physics-informed ML for reliability forecasting.
  • Lead risk-benefit trade-offs between cooling, capacity, and lifespan.

Skills

Design for Reliability
Hardware reliability engineering
Predictive analytics
Physics-informed ML

Education

Bachelor's degree (Reliability Engineering)
Master's degree (Reliability Engineering)

Tools

PyTorch
TensorFlow
PoF modeling

Job description

Share System Hardware Reliability Engineer

corporate_fare Google place Sunnyvale, CA, USA

Advanced

Experience owning outcomes and decision making, solving ambiguous problems and influencing stakeholders;deep expertise in domain.

  • Bachelor’s degree in Reliability Engineering, Data Science, Mechanical/Electrical Engineering, Applied Physics, or equivalent practical experience.
  • 8 years of experience in applying Design for Reliability techniques, and working on multiple consumer electronics products.
  • 8 years of experience in hardware reliability engineering, physics of failure, and predictive analytics.
Preferred qualifications:
  • Master’s degree in Reliability Engineering, Data Science, Mechanical/Electrical Engineering, Applied Physics, or equivalent practical experience.
  • 7 years of experience with experimental design and execution of complex hardware products.
  • 6 years of statistical analysis experience.
  • 6 years of experience working with manufacturing partners.
  • 5 years of experience with failure analysis techniques.
  • Expertise in physics-informed machine learning, Bayesian analysis, and utilizing frameworks like PyTorch or TensorFlow for reliability forecasting.
About the job

As a Reliability Engineer, you will play a key role in creating new consumer electronic products that meet a high bar for reliability and performance. You will work closely with the product management and design engineering teams to define standards, specify tests, and then supervise test execution and failure analysis. A broad engineering background and command of statistical methods will help to inform the design of new products. Your strong people management and communication skills will be key to ensuring adoption of your technical recommendations.

As a System Hardware Reliability Engineer, you will serve as the principal technical authority on hardware reliability, prognostics, and predictive analytics under dynamic thermal and environmental operating profiles. You will lead the development of sophisticated health monitoring models to evaluate the impact of elevated coolant temperatures, ambient air excursions, and dynamic workloads on the degradation and failure rates of compute accelerators, high-density servers, power electronics, and energy storage systems.

You will bridge classic Physics-of-Failure (PoF) modeling with machine learning to establish advanced Prognostics and Health Management (PHM) frameworks for our infrastructure. By developing algorithms that forecast remaining useful life and detect early-warning anomalies, you will perform system-level risk-benefit trade-offs between capacity efficiency and hardware lifespan. Your data-driven prognostic models will shape advanced cooling architectures and operational control strategies across our global computing footprint.

The AI and Infrastructure team is redefining what’s possible. We empower Google customers with breakthrough capabilities and insights by delivering AI and Infrastructure at unparalleled scale, efficiency, reliability and velocity. Our customers include Googlers, Google Cloud customers, and billions of Google users worldwide.

We’re the driving team behind Google’s groundbreaking innovations, empowering the development of our cutting-edge AI models, delivering unparalleled computing power to global services, and providing the essential platforms that enable developers to build the future. From software to hardware our teams are shaping the future of world-leading hyperscale computing, with key teams working on the development of our TPUs, Vertex AI for Google Cloud, Google Global Networking, Data Center operations, systems research, and much more.

Individual pay is determined by factors including job-related skills, experience, and relevant education or training.

About the job

As a Reliability Engineer, you will play a key role in creating new consumer electronic products that meet a high bar for reliability and performance. You will work closely with the product management and design engineering teams to define standards, specify tests, and then supervise test execution and failure analysis. A broad engineering background and command of statistical methods will help to inform the design of new products. Your strong people management and communication skills will be key to ensuring adoption of your technical recommendations.

As a System Hardware Reliability Engineer, you will serve as the principal technical authority on hardware reliability, prognostics, and predictive analytics under dynamic thermal and environmental operating profiles. You will lead the development of sophisticated health monitoring models to evaluate the impact of elevated coolant temperatures, ambient air excursions, and dynamic workloads on the degradation and failure rates of compute accelerators, high-density servers, power electronics, and energy storage systems.

You will bridge classic Physics-of-Failure (PoF) modeling with machine learning to establish advanced Prognostics and Health Management (PHM) frameworks for our infrastructure. By developing algorithms that forecast remaining useful life and detect early-warning anomalies, you will perform system-level risk-benefit trade-offs between capacity efficiency and hardware lifespan. Your data-driven prognostic models will shape advanced cooling architectures and operational control strategies across our global computing footprint.

The AI and Infrastructure team is redefining what’s possible. We empower Google customers with breakthrough capabilities and insights by delivering AI and Infrastructure at unparalleled scale, efficiency, reliability and velocity. Our customers include Googlers, Google Cloud customers, and billions of Google users worldwide.

We’re the driving team behind Google’s groundbreaking innovations, empowering the development of our cutting-edge AI models, delivering unparalleled computing power to global services, and providing the essential platforms that enable developers to build the future. From software to hardware our teams are shaping the future of world-leading hyperscale computing, with key teams working on the development of our TPUs, Vertex AI for Google Cloud, Google Global Networking, Data Center operations, systems research, and much more.

Individual pay is determined by factors including job-related skills, experience, and relevant education or training.

US: $188000 - $274000 (USD) + 20% bonus target + equity + benefits

Learn more about benefits at Google .

  • Design and implement Prognostics and Health Management (PHM) algorithms using physics-informed machine learning to forecast hardware degradation and Remaining Useful Life (RUL).
  • Develop stochastic degradation models to predict the impact of dynamic thermal and power envelopes on fleet reliability.
  • Build health state monitoring and anomaly detection frameworks leveraging massive fleet telemetry data to enable predictive maintenance.
  • Partner with software and controls teams to integrate predictive health models into automated load-management and thermal capping mechanisms.
  • Lead comprehensive reliability assessments and PoF modeling for silicon, interconnects, optics, thermal solutions and power delivery/battery systems. Drive Failure Modes and Effects Analyses (FMEAs) to identify vulnerabilities under extreme environmental operating conditions.

Google is proud to be an equal opportunity and affirmative action employer. We are committed to building a workforce that is representative of the users we serve, creating a culture of belonging, and providing an equal employment opportunity regardless of race, creed, color, religion, gender, sexual orientation, gender identity/expression, national origin, disability, age, genetic information, veteran status, marital status, pregnancy or related condition (including breastfeeding), expecting or parents-to-be, criminal histories consistent with legal requirements, or any other basis protected by law. See also Google's EEO Policy , Know your rights: workplace discrimination is illegal , Belonging at Google , and How we hire .

Google is a global company and, in order to facilitate efficient collaboration and communication globally, English proficiency is a requirement for all roles unless stated otherwise in the job posting.

Equity is granted exclusively and discretionarily by Alphabet Inc. on the basis of an agreement concluded between you and Alphabet Inc. Alphabet Inc. is your sole contractual partner with respect to equity grants. GSU grants are not guaranteed, are discretionary, are subject to approval by the Alphabet Inc. board of directors or its delegate, the terms of the relevant Alphabet Inc. stock plan, and your grant agreement. They have no impact on statutory payments. Current or past grants do not confer an acquired right.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Hardware Reliability Engineer, Global Hardware Reliability Engineering
Hardware Reliability Engineer, Global Hardware Reliability Engineering

Google Inc. • Austin (TX)

On-site
USD 144,000 - 209,000
Senior Hardware Engineering Manager, Emergent AI Infrastructure
Senior Hardware Engineering Manager, Emergent AI Infrastructure

Google Inc. • Sunnyvale (CA), Kirkland (WA)

On-site
USD 236,000 - 329,000
Health Insurance
Dental Insurance
Vision Insurance
+8
Staff Hardware Design Engineer, Platforms Infrastructure Engineering
Staff Hardware Design Engineer, Platforms Infrastructure Engineering

Google Inc. • Sunnyvale (CA)

On-site
USD 188,000 - 274,000
Staff Software Engineer, Data Center Resource Modeling
Staff Software Engineer, Data Center Resource Modeling

Google Inc. • Sunnyvale (CA)

On-site
USD 210,000 - 300,000
Senior Hardware Engineering Manager, Emergent AI Infrastructure
Senior Hardware Engineering Manager, Emergent AI Infrastructure

Google • United States

On-site
USD 236,000 - 329,000
Health insurance
401(k) with match
Paid time off
+1
Hardware Validation Engineer, Data Center Engineering Labs
Hardware Validation Engineer, Data Center Engineering Labs

Google Inc. • Sunnyvale (CA)

On-site
USD 132,000 - 189,000
Senior Staff Software Engineer, GPU System Software
Senior Staff Software Engineer, GPU System Software

Google Inc. • Sunnyvale (CA)

On-site
USD 262,000 - 364,000
Senior Software Engineer, Machine Health
Senior Software Engineer, Machine Health

Google • United States

On-site
USD 174,000 - 252,000
Senior Hardware Systems Design Engineer, Platforms Infrastructure
Senior Hardware Systems Design Engineer, Platforms Infrastructure

Google • United States

On-site
USD 159,000 - 230,000
Senior Staff Software Engineer, GPU System Software
Senior Staff Software Engineer, GPU System Software

Google • Sunnyvale (CA)

On-site
USD 262,000 - 364,000
Equity grants
Bonus target