Hardware Reliability Engineer, Global Hardware Reliability Engineering

Google

Austin (TX)

On-site

USD 144,000 - 209,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Google's Hardware Operations team in Austin seeks an experienced Hardware Reliability Engineer. You will manage the hardware reliability of machine learning, server, networking, and storage products, perform early configuration analysis, and support reliability initiatives.

The role involves collaborating with system architects, defining reliability plans, and executing tests to ensure rugged performance in data centers and challenging environments.

Qualifications

  • Bachelor's degree in a relevant engineering field and 8 years of reliability engineering experience.
  • Experience with system-level reliability tools and simulation environments.
  • Strong root-cause analysis and data-driven problem solving.

Responsibilities

  • Lead analysis of system hardware designs to enable proactive design evaluations and de-risk at an early stage of development.
  • Lead system reliability efforts by defining reliability goals and plans and securing required resources.
  • Develop mission profiles for chassis and racks from integration sites to field (data centers) to predict field reliability.
  • Implement reliability plans and lead efforts to assess and mitigate risk during new product introductions (NPI).
  • Drive reliability test plans and analyze data to verify design reliability goals.

Skills

Reliability engineering
Statistical analysis
Root cause analysis
Project coordination

Education

Bachelor's degree in Reliability, Electrical, Industrial, or Mechanical Engineering (or equivalent practical experience)

Tools

RBDs (Reliability Block Diagrams)
MCF (Mean Cumulative Function)
HPP/NHPP
Simulation tools
JMP

Job description

Minimum qualifications
  • Bachelor's degree in Reliability, Electrical, Industrial or Mechanical Engineering, or equivalent practical experience.
  • 8 years of experience with reliability engineering.
Preferred qualifications
  • Master's degree or PhD in Reliability, Electrical, Industrial, or Mechanical Engineering.
  • Experience with system level reliability tools such as Reliability Block Diagrams (RBDs), Mean Cumulative Function (MCF), Homogeneous and Non-Homogeneous Poisson Processes (HPP, NHPP), and simulation tools.
  • Experience in accelerated life testing, reliability modeling, reliability stats, mission profile development.
  • Experience with failure analysis and fault isolation techniques and their application to find root causes of failure.
  • Experience in statistics and JMP.
  • Understanding of physics of failure and reliability physics.
About The Job

Google isn't just a software company. The Hardware Operations team is responsible for monitoring the state-of-the-art physical infrastructure behind Google's powerful search technology. As an Operations Technician, you'll install, configure, test, troubleshoot and maintain hardware (like servers and its components) and server software (like Google's Linux cluster). You'll also take on the configuration of more complex components such as networks, routers, hubs, bridges, switches and networking protocols. You'll participate in or lead small project teams on larger installations and develop project contingency plans. A typical day involves manual movement and installation of racks, and while it can sometimes be physically demanding, you are excited to work with infrastructure that is at the cutting-edge of computer technology.

As a Hardware Reliability Engineer, you will manage the hardware reliability of new machine learning, server, networking, and storage products. You will also perform early system configuration analysis and simulations, to assess reliability capability of the design and plan of record (POR) configurations. You will perform DFR, Reliability test development for hardware targeted to survive rugged environments. You will support field excursion issues and statistical analysis needs for reliability. You will conduct Mechanical, Environmental and Operational accelerated stress test development to cover transportation, operation in challenging usage environments.

In this role, you will perform early engagement with the system architect and product teams to drive the selection of best design options, materials, and sub-components/modules/sub-systems. You will define optimal reliability plans for new products, securing resources for required samples and testing needs. You will define and implement an Ongoing Reliability program for the product at the contract manufacturer.

The AI and Infrastructure team is redefining what’s possible. We empower Google customers with breakthrough capabilities and insights by delivering AI and Infrastructure at unparalleled scale, efficiency, reliability and velocity. Our customers include Googlers, Google Cloud customers, and billions of Google users worldwide.

We're the driving force behind Google's groundbreaking innovations, empowering the development of our cutting-edge AI models, delivering unparalleled computing power to global services, and providing the essential platforms that enable developers to build the future. From software to hardware our teams are shaping the future of world-leading hyperscale computing, with key teams working on the development of our TPUs, Vertex AI for Google Cloud, Google Global Networking, Data Center operations, systems research, and much more.

Individual pay is determined by factors including job-related skills, experience, and relevant education or training.

US: $144000 - $209000 (USD) + 15% bonus target + equity + benefits

Responsibilities

Learn more about benefits at Google.

  • Lead analysis of system hardware designs to enable proactive design evaluations and product de-risk at an early stage of development.
  • Lead system reliability efforts by working with other organizations to define reliability goals and reliability plans, securing the resources needed to execute the plan.
  • Develop mission profiles for chasis, rack from integration sites to field (data centers) that help predict field reliability.
  • Implement the reliability plan and lead all efforts to assess and mitigate risk of failure early during new product introduction (NPI).
  • Drive reliability test plans and collect, analyze, and synthesize the test data to enable verification of the design reliability goals.

Google is proud to be an equal opportunity workplace and is an affirmative action employer. We are committed to equal employment opportunity regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity or Veteran status. We also consider qualified applicants regardless of criminal histories, consistent with legal requirements. See also Google's EEO Policy and EEO is the Law. If you have a disability or special need that requires accommodation, please let us know by completing our Accommodations for Applicants form .

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Hardware Reliability Engineer, Global Hardware Reliability Engineering
Hardware Reliability Engineer, Global Hardware Reliability Engineering

Google Inc. • Austin (TX)

On-site
USD 144,000 - 209,000
Hardware Data and Failure Analysis Engineer
Hardware Data and Failure Analysis Engineer

Google • Mountain View (CA)

On-site
USD 132,000 - 190,000
Senior Staff Technical Lead, Platforms Hardware Validation and Infrastructure
Senior Staff Technical Lead, Platforms Hardware Validation and Infrastructure

Google • Sunnyvale (CA)

On-site
USD 236,000 - 330,000
Product Engineer, ARM Servers, Storage and ML Systems
Product Engineer, ARM Servers, Storage and ML Systems

Google • Sunnyvale (CA)

On-site
USD 144,000 - 209,000
Hardware Validation Engineer, University Graduate, Platforms Infrastructure
Hardware Validation Engineer, University Graduate, Platforms Infrastructure

Google • Sunnyvale (CA)

On-site
USD 105,000 - 151,000
Equity options
Comprehensive benefits
Work-life balance initiatives
Hardware Validation Engineer, ML Products, Google Cloud
Hardware Validation Engineer, ML Products, Google Cloud

Google • Sunnyvale (CA)

On-site
USD 105,000 - 151,000
Manager, Hardware Engineering
Manager, Hardware Engineering

Google • San Jose (CA)

Hybrid
USD 283,000 - 330,000
Equity
Benefits
Data Center Area Lead, Server Operations
Data Center Area Lead, Server Operations

Google • Midlothian (TX)

On-site
USD 205,000 - 285,000
Equity
Benefits
Hardware Validation Engineer, Data Center Engineering Labs
Hardware Validation Engineer, Data Center Engineering Labs

Google • Sunnyvale (CA)

On-site
USD 132,000 - 189,000
Hardware Test Engineer, Verification Engineering
Hardware Test Engineer, Verification Engineering

Google • The Dalles (OR)

On-site
USD 113,000 - 161,000