Hardware Reliability Engineer

Meta

Fremont (CA)

On-site

USD 144,000 - 204,000

Full time

2 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Meta Infrastructure's Hardware Product Integrity team is seeking a skilled Hardware Reliability Engineer to help build next-generation data center hardware that underpins Meta's AI initiatives. You will drive reliability across server, storage, and networking hardware deployed in data centers, applying failure analysis, accelerated life testing, and lifetime reliability modeling to reduce field failures.

In this role you will lead DFMEA, work with ODMs, translate test results into life metrics,

Qualifications

  • 6+ years of hardware reliability engineering experience in infrastructure hardware.
  • Expertise in FMEA, HALT, ALT, Weibull analysis, MTBF modeling.
  • Experience analyzing field failure data and driving corrective actions.
  • Experience working with hardware suppliers/contract manufacturers to evaluate component reliability.
  • Ability to communicate complex reliability findings to engineering and operations teams.

Responsibilities

  • Lead DFMEA activities across AI, compute and storage platforms.
  • Develop reliability tests to reveal design weaknesses in compute, storage, server hardware, and networking modules.
  • Establish Design Verification tests to meet life-time reliability metrics.
  • Collaborate with ODMs to ensure tests run as planned and implement improvements.
  • Translate test results into product life metrics and highlight unmet targets.
  • Use reliability statistics to inform decisions and quantify risk.
  • Develop internal reliability test infrastructure to support designs and experiments.
  • Collaborate cross-functionally with Hardware Engineering, Release to Production, Thermal, and Failure Analysis teams to de-risk issues.

Skills

DFMEA & reliability testing
Weibull analysis
MTBF modeling
Failure analysis
Cross-functional collaboration
Data analysis
Root cause investigations

Education

MSc in Mechanical or Electrical Engineering

Job description

Summary:

As a member of Meta Infrastructure's Hardware Product Integrity team you will work on next-generation data center hardware. You will be a part of futuristic projects including HW that will serve as the backbone for Meta's AGI vision, be in a position to influence HW technology that serves to connect billions of people across the world! In this role, you will drive reliability engineering efforts across server, storage, and networking hardware deployed in Meta's data centers, applying failure analysis, accelerated life testing, and reliability modeling to reduce field failures and improve hardware reliability.

Required Skills:
Hardware Reliability Engineer Responsibilities:
  1. Lead DFR activities such as DFMEA, derating across various AI, compute and storage platforms

  2. Understanding technology that drives compute, storage, server hardware, and networking modules to develop reliability tests to bring out design weaknesses

  3. Establish Design Verification tests, to bring out environmental stress weaknesses in server design and ensure designs meet Meta's lifetime reliability metrics

  4. Work closely with ODMs to ensure and oversee tests are being executed as planned, suggest necessary improvements based on lessons learned from previous platforms

  5. Translate test results into meaningful product life metrics and highlight shortcomings in any metrics that are not met

  6. Utilize reliability statistics to help with decision making and quantifying risk and

  7. Lead the development of internal reliability test infrastructure to support initiatives and design of experiments

  8. Collaborate cross-functionally with Hardware Engineering, Release To Production, Thermal, and Failure Analysis teams to de-risk design issues

Minimum Qualifications:

Minimum Qualifications:

  1. 6+ years of experience in hardware reliability engineering, including failure analysis and reliability testing of infrastructure hardware

  2. Experience applying reliability engineering methodologies such as FMEA, HALT, ALT, Weibull analysis, and MTBF modeling to infrastructure hardware

  3. Experience analyzing field failure data and translating findings into actionable root cause investigations and corrective actions

  4. Experience collaborating with hardware suppliers and contract manufacturers to evaluate component reliability and enforce qualification standards

  5. Experience communicating complex reliability findings and technical trade-offs to engineering and operations stakeholders through written reports and presentations

Preferred Qualifications:

Preferred Qualifications:

  1. Experience in silicon reliability and working on custom silicon is a plus

  2. MSc in Mechanical or Electrical Engineering or related disciplines

  3. Familiarity with data center environments is beneficial

  4. First-hand knowledge of server rack hardware is preferred

Public Compensation:

$144,000/year to $204,000/year + bonus + equity + benefits

Industry:

Internet

Equal Opportunity:

Meta is proud to be an Equal Employment Opportunity and affirmative action employer. We do not discriminate based upon race, religion, color, national origin, sex (including pregnancy, childbirth, or related medical conditions), sexual orientation, gender, gender identity, gender expression, transgender status, sexual stereotypes, age, status as a protected veteran, status as an individual with a disability, or other applicable legally protected characteristics. We also consider qualified applicants with criminal histories, consistent with applicable federal, state and local law. Meta participates in the E-Verify program in certain locations, as required by law. Please note that Meta may leverage artificial intelligence and machine learning technologies in connection with applications for employment.

Meta is committed to providing reasonable accommodations for candidates with disabilities in our recruiting process. If you need any assistance or accommodations due to a disability, please let us know at accommodations-ext@meta.com.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Hardware Reliability Engineer
Hardware Reliability Engineer

Meta • Menlo Park (CA), Northern (KY)

Hybrid
USD 160,000 - 260,000
Production Systems Engineer, AI Systems
Production Systems Engineer, AI Systems

Meta • Austin (TX)

On-site
USD 144,000 - 204,000
Senior Hardware Reliability Engineer — Data Center Systems
Senior Hardware Reliability Engineer — Data Center Systems

Meta • Fremont (CA)

On-site
USD 144,000 - 204,000
Product Quality Engineer
Product Quality Engineer

Meta • Austin (TX)

On-site
USD 144,000 - 204,000
Data Center Production Operations Engineer
Data Center Production Operations Engineer

Meta • El Paso (TX)

On-site
USD 84,000 - 130,000
Mechanical Engineer, Infrastructure
Mechanical Engineer, Infrastructure

Meta • Menlo Park (CA)

On-site
USD 118,000 - 170,000
Bonus
Equity
Benefits
Technical Program Manager, AI Infrastructure
Technical Program Manager, AI Infrastructure

Meta • Bellevue (WA)

On-site
USD 168,000 - 234,000
Equity
Bonus
Benefits
Reliability Engineer
Reliability Engineer

Meta • Sunnyvale (CA), Redmond (WA), Seattle (WA)

On-site
USD 130,000 - 165,000
Strategic Partnerships Director, Infrastructure Hardware
Strategic Partnerships Director, Infrastructure Hardware

Meta • Providence (RI)

On-site
USD 260,000 - 319,000
Technical Lead, Infrastructure Silicon Validation
Technical Lead, Infrastructure Silicon Validation

Meta • Sunnyvale (CA)

On-site
USD 212,000 - 294,000
Bonus
Equity
Benefits