Senior Hardware Reliability Engineer, Data Center & AI

NVIDIA Corporation

Santa Clara (CA)

On-site

USD 168,000 - 265,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

NVIDIA is seeking a System Reliability Engineer to join the Reliability Engineering team, focusing on GPUs, data center servers, and related products. You will define reliability tests, lead failure analyses, and drive corrective actions to improve design and manufacturing quality.

The role requires a Bachelor’s or Master’s degree in engineering and 8+ years in hardware reliability. Expect collaboration across engineering and suppliers, with opportunities to influence product reliability

Qualifications

  • Bachelor’s or Master’s degree in Electrical or Mechanical Engineering or equivalent experience.
  • 8+ years of hardware reliability experience in datacenter, systems, or computer industries.
  • Hands-on experience with theoretical and practical reliability concepts for high-tech electronic products.
  • Strong understanding of statistical concepts as they relate to product reliability and life analysis.
  • Excellent verbal and written communication skills at a high level.
  • Strong project management skills with ability to balance multiple projects during development.

Responsibilities

  • Represent Product Reliability Engineering in development teams.
  • Develop and execute reliability test plans for product qualification.
  • Collaborate with peer reliability groups for successful implementation.
  • Define product reliability tests for products from embedded to server and cluster.
  • Lead reliability testing, failure analysis, and root cause investigations; drive corrective actions.

Skills

Reliability engineering
Statistical concepts
Verbal written communication
Project management

Education

Bachelor’s or Master’s degree in Electrical/Mechanical Engineering

Tools

FMEA
DoE

Job description

NVIDIA is seeking a System Reliability Engineer to join the Reliability Engineering team, focusing on GPUs, data center servers, and related products. You will define reliability tests, lead failure analyses, and drive corrective actions to improve design and manufacturing quality.

The role requires a Bachelor’s or Master’s degree in engineering and 8+ years in hardware reliability. Expect collaboration across engineering and suppliers, with opportunities to influence product reliability

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Head of Product Reliability - Hardware & AI
Head of Product Reliability - Hardware & AI

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 256,000 - 385,000
Senior Director, Hardware Reliability & AI-Driven Strategy
Senior Director, Hardware Reliability & AI-Driven Strategy

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 332,000 - 500,000
Senior System Reliability Engineer
Senior System Reliability Engineer

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 168,000 - 265,000
Head of Failure Analysis & Reliability Innovation
Head of Failure Analysis & Reliability Innovation

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 232,000 - 368,000
GPU Field Engineer - Manufacturing & Reliability
GPU Field Engineer - Manufacturing & Reliability

NVIDIA Corporation • Tennessee

Hybrid
USD 132,000 - 207,000
Equity compensation
Benefits
Failure Analysis Leader – Hardware Reliability
Failure Analysis Leader – Hardware Reliability

NVIDIA • Santa Clara (CA)

On-site
USD 232,000 - 368,000
Equity
Benefits
Senior Director, Reliability Engineering
Senior Director, Reliability Engineering

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 332,000 - 500,000
Senior Manager, Product Reliability
Senior Manager, Product Reliability

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 256,000 - 385,000
Lead Systems Quality & Reliability Engineer
Lead Systems Quality & Reliability Engineer

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 168,000 - 311,000
Equity
Benefits
Senior Hardware Quality & Reliability Leader — Equity
Senior Hardware Quality & Reliability Leader — Equity

NVIDIA • Santa Clara (CA)

On-site
USD 168,000 - 258,750
Equity
Comprehensive benefits package