Hardware Reliability Engineer, Global Hardware Reliability Engineering

Socket.dev

Austin (TX)

On-site

USD 144,000 - 208,000

Full time

5 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Google is seeking an Hardware Reliability Engineer to manage the hardware reliability of new machine learning, server, networking, and storage products. You will perform early system configuration analysis and simulations to assess reliability capability and plan configurations for rugged environments.

You will define and implement ongoing reliability programs, engage with system architects, and drive reliability plans across product development and contract manufacturing, shaping the future of

Qualifications

  • Bachelor's degree in Reliability, Electrical, Industrial or Mechanical Engineering, or equivalent practical experience.
  • 8 years of experience with reliability engineering.

Responsibilities

  • Lead analysis of system hardware designs to enable proactive design evaluations and product de-risk at an early stage of development.
  • Lead system reliability efforts by defining reliability goals and reliability plans, securing the resources needed to execute the plan.
  • Develop mission profiles for chassis, rack from integration sites to field (data centers) that help predict field reliability.
  • Implement the reliability plan and lead all efforts to assess and mitigate risk of failure early during new product introduction (NPI).
  • Drive reliability test plans and collect, analyze, and synthesize the test data to enable verification of the design reliability goals.

Skills

Reliability engineering
RBDs (Reliability Block Diagrams)
MCF (Mean Cumulative Function)
HPP/NHPP (Poisson processes)
Simulation tools
Accelerated life testing
Reliability modeling
Reliability statistics
Mission profile development
Failure analysis
Statistical analysis (JMP)

Education

Bachelor's degree in Reliability, Electrical, Industrial or Mechanical Engineering, or equivalent practical experience.

Tools

JMP

Job description

Minimum qualifications:
  • Bachelor's degree in Reliability, Electrical, Industrial or Mechanical Engineering, or equivalent practical experience.
  • 8 years of experience with reliability engineering.
Preferred qualifications:
  • Master's degree or PhD in Reliability, Electrical, Industrial, or Mechanical Engineering.
  • Experience with system level reliability tools such as Reliability Block Diagrams (RBDs), Mean Cumulative Function (MCF), Homogeneous and Non-Homogeneous Poisson Processes (HPP, NHPP), and simulation tools.
  • Experience in accelerated life testing, reliability modeling, reliability stats, mission profile development.
  • Experience with failure analysis and fault isolation techniques and their application to find root causes of failure.
  • Experience in statistics and JMP.
  • Understanding of physics of failure and reliability physics.
About the job:

Google isn't just a software company. The Hardware Operations team is responsible for monitoring the state-of-the-art physical infrastructure behind Google's powerful search technology. As an Operations Technician, you'll install, configure, test, troubleshoot and maintain hardware (like servers and its components) and server software (like Google's Linux cluster). You'll also take on the configuration of more complex components such as networks, routers, hubs, bridges, switches and networking protocols. You'll participate in or lead small project teams on larger installations and develop project contingency plans. A typical day involves manual movement and installation of racks, and while it can sometimes be physically demanding, you are excited to work with infrastructure that is at the cutting-edge of computer technology.

As a Hardware Reliability Engineer, you will manage the hardware reliability of new machine learning, server, networking, and storage products. You will also perform early system configuration analysis and simulations, to assess reliability capability of the design and plan of record (POR) configurations. You will perform DFR, Reliability test development for hardware targeted to survive rugged environments. You will support field excursion issues and statistical analysis needs for reliability. You will conduct Mechanical, Environmental and Operational accelerated stress test development to cover transportation, operation in challenging usage environments.

In this role, you will perform early engagement with the system architect and product teams to drive the selection of best design options, materials, and sub-components/modules/sub-systems. You will define optimal reliability plans for new products, securing resources for required samples and testing needs. You will define and implement an Ongoing Reliability program for the product at the contract manufacturer.

The AI and Infrastructure team is redefining what’s possible. We empower Google customers with breakthrough capabilities and insights by delivering AI and Infrastructure at unparalleled scale, efficiency, reliability and velocity. Our customers include Googlers, Google Cloud customers, and billions of Google users worldwide.

We're the driving force behind Google's groundbreaking innovations, empowering the development of our cutting-edge AI models, delivering unparalleled computing power to global services, and providing the essential platforms that enable developers to build the future. From software to hardware our teams are shaping the future of world-leading hyperscale computing, with key teams working on the development of our TPUs, Vertex AI for Google Cloud, Google Global Networking, Data Center operations, systems research, and much more.

Individual pay is determined by factors including job-related skills, experience, and relevant education or training.

US: $144000 - $208000 (USD) + 15% bonus target + equity + benefits

Learn more about benefits at Google.

Responsibilities:
  • Lead analysis of system hardware designs to enable proactive design evaluations and product de-risk at an early stage of development.
  • Lead system reliability efforts by working with other organizations to define reliability goals and reliability plans, securing the resources needed to execute the plan.
  • Develop mission profiles for chasis, rack from integration sites to field (data centers) that help predict field reliability.
  • Implement the reliability plan and lead all efforts to assess and mitigate risk of failure early during new product introduction (NPI).
  • Drive reliability test plans and collect, analyze, and synthesize the test data to enable verification of the design reliability goals.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Hardware Reliability Engineer, Global Hardware Reliability Engineering
Hardware Reliability Engineer, Global Hardware Reliability Engineering

Google Inc. • Austin (TX)

On-site
USD 144,000 - 209,000
System Hardware Reliability Engineer
System Hardware Reliability Engineer

Google Inc. • Sunnyvale (CA)

On-site
USD 188,000 - 274,000
Software Engineer, Google Cloud Platform, Fault Management
Software Engineer, Google Cloud Platform, Fault Management

Google • Sunnyvale (CA)

On-site
USD 180,000 - 300,000
Software Engineer, Google Cloud Platform, Fault Management
Software Engineer, Google Cloud Platform, Fault Management

Socket.dev • Sunnyvale (CA)

On-site
USD 207,000 - 300,000
Hardware Reliability Engineer – System & NPI Focus
Hardware Reliability Engineer – System & NPI Focus

Socket.dev • Austin (TX)

On-site
USD 144,000 - 208,000
Data Center Operations Manager
Data Center Operations Manager

Socket.dev • Fort Wayne (IN)

On-site
USD 122,000 - 173,000
Staff Hardware Design Engineer, Platforms Infrastructure Engineering
Staff Hardware Design Engineer, Platforms Infrastructure Engineering

Socket.dev • Sunnyvale (CA)

On-site
USD 188,000 - 274,000
Equity
Bonus potential
Benefits
Hardware Validation Engineer, University Graduate, Platforms Infrastructure
Hardware Validation Engineer, University Graduate, Platforms Infrastructure

Google • Sunnyvale (CA)

On-site
USD 105,000 - 151,000
Equity options
Comprehensive benefits
Work-life balance initiatives
Hardware Test Engineer, Verification Engineering
Hardware Test Engineer, Verification Engineering

Google • The Dalles (OR)

On-site
USD 113,000 - 161,000
Staff Software Engineer, Network Health
Staff Software Engineer, Network Health

Google • Sunnyvale (CA)

On-site
USD 207,000 - 300,000