Get more replies from employers
Send a job-specific resume in minutes.
Google Sunnyvale is hiring a System Hardware Reliability Engineer to own physics-based health management and predictive analytics for high-density compute hardware. You will define reliability standards, design tests, and oversee analysis while collaborating with product and design teams.
The role emphasizes physics-of-failure modeling integrated with ML, forecasting remaining useful life and optimizing cooling and power strategies across the global fleet.
Share System Hardware Reliability Engineer
corporate_fare Google place Sunnyvale, CA, USA
Experience owning outcomes and decision making, solving ambiguous problems and influencing stakeholders;deep expertise in domain.
As a Reliability Engineer, you will play a key role in creating new consumer electronic products that meet a high bar for reliability and performance. You will work closely with the product management and design engineering teams to define standards, specify tests, and then supervise test execution and failure analysis. A broad engineering background and command of statistical methods will help to inform the design of new products. Your strong people management and communication skills will be key to ensuring adoption of your technical recommendations.
As a System Hardware Reliability Engineer, you will serve as the principal technical authority on hardware reliability, prognostics, and predictive analytics under dynamic thermal and environmental operating profiles. You will lead the development of sophisticated health monitoring models to evaluate the impact of elevated coolant temperatures, ambient air excursions, and dynamic workloads on the degradation and failure rates of compute accelerators, high-density servers, power electronics, and energy storage systems.
You will bridge classic Physics-of-Failure (PoF) modeling with machine learning to establish advanced Prognostics and Health Management (PHM) frameworks for our infrastructure. By developing algorithms that forecast remaining useful life and detect early-warning anomalies, you will perform system-level risk-benefit trade-offs between capacity efficiency and hardware lifespan. Your data-driven prognostic models will shape advanced cooling architectures and operational control strategies across our global computing footprint.
The AI and Infrastructure team is redefining what’s possible. We empower Google customers with breakthrough capabilities and insights by delivering AI and Infrastructure at unparalleled scale, efficiency, reliability and velocity. Our customers include Googlers, Google Cloud customers, and billions of Google users worldwide.
We’re the driving team behind Google’s groundbreaking innovations, empowering the development of our cutting-edge AI models, delivering unparalleled computing power to global services, and providing the essential platforms that enable developers to build the future. From software to hardware our teams are shaping the future of world-leading hyperscale computing, with key teams working on the development of our TPUs, Vertex AI for Google Cloud, Google Global Networking, Data Center operations, systems research, and much more.
Individual pay is determined by factors including job-related skills, experience, and relevant education or training.
As a Reliability Engineer, you will play a key role in creating new consumer electronic products that meet a high bar for reliability and performance. You will work closely with the product management and design engineering teams to define standards, specify tests, and then supervise test execution and failure analysis. A broad engineering background and command of statistical methods will help to inform the design of new products. Your strong people management and communication skills will be key to ensuring adoption of your technical recommendations.
As a System Hardware Reliability Engineer, you will serve as the principal technical authority on hardware reliability, prognostics, and predictive analytics under dynamic thermal and environmental operating profiles. You will lead the development of sophisticated health monitoring models to evaluate the impact of elevated coolant temperatures, ambient air excursions, and dynamic workloads on the degradation and failure rates of compute accelerators, high-density servers, power electronics, and energy storage systems.
You will bridge classic Physics-of-Failure (PoF) modeling with machine learning to establish advanced Prognostics and Health Management (PHM) frameworks for our infrastructure. By developing algorithms that forecast remaining useful life and detect early-warning anomalies, you will perform system-level risk-benefit trade-offs between capacity efficiency and hardware lifespan. Your data-driven prognostic models will shape advanced cooling architectures and operational control strategies across our global computing footprint.
The AI and Infrastructure team is redefining what’s possible. We empower Google customers with breakthrough capabilities and insights by delivering AI and Infrastructure at unparalleled scale, efficiency, reliability and velocity. Our customers include Googlers, Google Cloud customers, and billions of Google users worldwide.
We’re the driving team behind Google’s groundbreaking innovations, empowering the development of our cutting-edge AI models, delivering unparalleled computing power to global services, and providing the essential platforms that enable developers to build the future. From software to hardware our teams are shaping the future of world-leading hyperscale computing, with key teams working on the development of our TPUs, Vertex AI for Google Cloud, Google Global Networking, Data Center operations, systems research, and much more.
Individual pay is determined by factors including job-related skills, experience, and relevant education or training.
US: $188000 - $274000 (USD) + 20% bonus target + equity + benefits
Learn more about benefits at Google .
Google is proud to be an equal opportunity and affirmative action employer. We are committed to building a workforce that is representative of the users we serve, creating a culture of belonging, and providing an equal employment opportunity regardless of race, creed, color, religion, gender, sexual orientation, gender identity/expression, national origin, disability, age, genetic information, veteran status, marital status, pregnancy or related condition (including breastfeeding), expecting or parents-to-be, criminal histories consistent with legal requirements, or any other basis protected by law. See also Google's EEO Policy , Know your rights: workplace discrimination is illegal , Belonging at Google , and How we hire .
Google is a global company and, in order to facilitate efficient collaboration and communication globally, English proficiency is a requirement for all roles unless stated otherwise in the job posting.
Equity is granted exclusively and discretionarily by Alphabet Inc. on the basis of an agreement concluded between you and Alphabet Inc. Alphabet Inc. is your sole contractual partner with respect to equity grants. GSU grants are not guaranteed, are discretionary, are subject to approval by the Alphabet Inc. board of directors or its delegate, the terms of the relevant Alphabet Inc. stock plan, and your grant agreement. They have no impact on statutory payments. Current or past grants do not confer an acquired right.