Hardware Reliability Engineer

Meta

Menlo Park, Northern (CA, KY)

Hybrid

USD 160,000 - 260,000

Full time

3 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Meta Infrastructure's Hardware Product Integrity team seeks a seasoned hardware reliability engineer to help design and sustain next-generation data center hardware.

You will apply FMEA, Weibull, MTBF modeling, analyze field failures, and drive root-cause investigations with suppliers, influencing server, storage, and networking hardware across Meta's global data centers.

Qualifications

  • Bachelor's degree in Electrical or Mechanical Engineering.
  • 6+ years of hardware reliability engineering experience.
  • Experience applying FMEA, HALT, ALT, Weibull, MTBF to infrastructure hardware.
  • Experience analyzing field failure data and corrective actions.
  • Experience collaborating with suppliers and contract manufacturers on reliability.
  • Ability to communicate reliability findings to engineers and operations.
  • MSc preferred in related disciplines.
  • Familiarity with data center environments and server rack hardware is a plus.

Responsibilities

  • Drive reliability engineering across server, storage, and networking hardware in data centers.
  • Apply failure analysis, accelerated life testing, and reliability modeling to reduce field failures.
  • Collaborate with suppliers and contract manufacturers to enforce qualification standards.
  • Translate findings into root-cause investigations and corrective actions.
  • Communicate complex reliability findings to engineering and operations stakeholders.

Skills

Hardware reliability engineering
Failure analysis
FMEA
HALT
ALT
Weibull analysis
MTBF modeling
Root cause analysis
Data analysis
Communication to stakeholders
Silicon reliability (plus)

Education

Bachelor's degree in Electrical Engineering or Mechanical Engineering
MSc in Mechanical or Electrical Engineering

Job description

Job Description

As a member of Meta Infrastructure's Hardware Product Integrity team you will work on next-generation data center hardware. You will be a part of futuristic projects including HW that will serve as the backbone for Meta's AGI vision, be in a position to influence HW technology that serves to connect billions of people across the world! In this role, you will drive reliability engineering efforts across server, storage, and networking hardware deployed in Meta's data centers, applying failure analysis, accelerated life testing, and reliability modeling to reduce field failures and improve hardware reliability.

Qualifications and Responsibilities
  • Bachelor's degree in Electrical Engineering or Mechanical Engineering or a related discipline
  • 6+ years of experience in hardware reliability engineering, including failure analysis and reliability testing of infrastructure hardware
  • Experience applying reliability engineering methodologies such as FMEA, HALT, ALT, Weibull analysis, and MTBF modeling to infrastructure hardware
  • Experience analyzing field failure data and translating findings into actionable root cause investigations and corrective actions
  • Experience collaborating with hardware suppliers and contract manufacturers to evaluate component reliability and enforce qualification standards
  • Experience communicating complex reliability findings and technical trade-offs to engineering and operations stakeholders through written reports and presentations
  • Experience in silicon reliability and working on custom silicon is a plus
  • MSc in Mechanical or Electrical Engineering or related disciplines
  • Familiarity with data center environments is beneficial
  • First-hand knowledge of server rack hardware is preferred
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Hardware Reliability Engineer
Hardware Reliability Engineer

Meta • Fremont (CA)

On-site
USD 144,000 - 204,000
Senior Hardware Reliability Engineer — Data Center Systems
Senior Hardware Reliability Engineer — Data Center Systems

Meta • Fremont (CA)

On-site
USD 144,000 - 204,000
Reliability Engineer
Reliability Engineer

Meta • Sunnyvale (CA), Redmond (WA), Seattle (WA)

On-site
USD 130,000 - 165,000
Hardware Reliability Role for MLB
Hardware Reliability Role for MLB

OSI Engineering • Austin (TX)

Remote
Senior Reliability Engineer
Senior Reliability Engineer

Insight Global • Redmond (WA)

On-site
USD 110,000 - 170,000
Hardware Engineer
Hardware Engineer

Quanta Manufacturing Fremont • Fremont (CA)

On-site
USD 90,000 - 150,000
Hardware Reliability Engineer - Wearables & VR
Hardware Reliability Engineer - Wearables & VR

Meta • Sunnyvale (CA), Redmond (WA), Seattle (WA)

On-site
USD 130,000 - 165,000
Data Center - MLB Reliability Engineer
Data Center - MLB Reliability Engineer

Socket.dev • Austin (TX)

On-site
USD 140,000 - 180,000
Reliability Engineer - AI Hardware & Data Center RAS
Reliability Engineer - AI Hardware & Data Center RAS

Intel • Boxborough (MA)

On-site
USD 122,000 - 232,000
Stock bonuses
Health benefits
Retirement plans
+1
Hardware Engineer, AI Accelerator Module Design
Hardware Engineer, AI Accelerator Module Design

Meta • Menlo Park (CA)

On-site
USD 140,000 - 180,000