Senior Manager, Product Reliability

NVIDIA Corporation

Santa Clara (CA)

On-site

USD 256,000 - 385,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

NVIDIA in Santa Clara, CA is seeking a Manager of Product Reliability to lead a team of reliability engineers and define end-to-end reliability across the product lifecycle. You will drive test plans, root-cause analyses, and cross-functional efforts with engineering, validation, manufacturing, and field teams to ensure robust hardware platforms.

The role emphasizes data-driven risk assessment, collaboration with ODM partners and labs, and delivering reliability predictions while aligning with

Qualifications

  • 10+ years of reliability, qualification, or hardware engineering experience in datacenter or computer sectors.
  • 5+ years of people management experience.
  • Strong foundation in reliability engineering principles, including accelerated testing, failure analysis, and statistical modeling.
  • Experience with system-level hardware: servers, racks, or complex electronic systems.
  • Excellent cross-functional collaboration and communication skills.

Responsibilities

  • Lead a team of Product Reliability Engineers.
  • Drive end-to-end reliability strategy across the product lifecycle (concept - qualification - production - field).
  • Coordinate creation and implementation of reliability test plans and references.
  • Perform failure analysis and root cause investigations to mitigate risk.
  • Collaborate with engineering, validation, manufacturing, operations, and field teams to improve product robustness.

Skills

Reliability engineering
Leadership
Statistical modeling

Education

Electrical/Mechanical Eng degree

Tools

Test equipment

Job description

Joining NVIDIA means becoming part of a legacy of innovation that has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. As a Manager, Product Reliability, you will be at the forefront of defining the next era of computing, driven by our groundbreaking AI technologies. This outstanding opportunity allows you to lead and build the future of our engineering projects, making a lasting impact on the world. Our team in Santa Clara, CA, is dedicated to collaboration, innovation, and excellence. We are looking for passionate individuals who thrive in a dynamic environment and are eager to contribute to world-class projects. If you are ambitious and ready to take on new challenges, this role is perfect for you!

What you’ll be doing:

Lead a team of Product Reliability Engineers. Drive end-to-end reliability strategy throughout the product lifecycle (concept - qualification - production - field). Coordinate the creation and implementation of the reliability test plan. Establish reliability test plan references, test methodologies, and qualification processes. Drive failure analysis and root cause investigations, ensuring timely resolution and risk mitigation. Partner cross-functionally with engineering, validation, manufacturing, operations, and field teams to resolve issues and improve product robustness. Build positive relationships with ODM partners and third-party labs to scale testing capabilities. Provide reliability predictions. Use data-driven approaches to assess risk, make informed decisions, and communicate reliability status to collaborators and leadership.

What we need to see:

Bachelor’s or Master’s degree in Electrical Engineering, Mechanical Engineering, or a related field (or equivalent experience). 10+ overall years of experience in reliability, qualification, or hardware engineering positions within datacenter or computer sectors. 5+ years of experience in managing personnel. Strong foundation in reliability engineering principles, including accelerated testing, failure analysis, and statistical modeling. Experience with system-level hardware: servers, racks, or complex electronic systems. Proven ability to drive cross-functional alignment and execution in a global environment. Excellent written and verbal communication skills, with the ability to present complex technical concepts to diverse audiences.

Widely considered to be one of the technology world’s most desirable employers, NVIDIA offers highly competitive salaries and a comprehensive benefits package. As you plan your future, see what we can offer to you and your family www.nvidiabenefits.com/. Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 256,000 USD - 385,250 USD. You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until July 13, 2026.

This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes. NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law. NVIDIA pioneered accelerated computing. Today, our AI infrastructure powers global intelligence, transforming every industry. Learn more about NVIDIA.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Director, Reliability Engineering
Senior Director, Reliability Engineering

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 332,000 - 500,000
Senior System Reliability Engineer
Senior System Reliability Engineer

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 168,000 - 265,000
Systems Quality and Reliability Lead - LPU
Systems Quality and Reliability Lead - LPU

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 168,000 - 311,000
Equity
Benefits
Engineering Manager, Reliability Engineering Flywheel - EDA Infrastructure
Engineering Manager, Reliability Engineering Flywheel - EDA Infrastructure

NVIDIA • California (MO)

On-site
USD 224,000 - 431,000
Equity
Benefits
Manager, Mechanical Engineering
Manager, Mechanical Engineering

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 168,000 - 265,000
Equity
Benefits package
Senior Manager, Post-Silicon Bring-Up
Senior Manager, Post-Silicon Bring-Up

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 232,000 - 368,000
Equity
Benefits package
Senior Manager, Failure Analysis
Senior Manager, Failure Analysis

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 232,000 - 368,000
Manager, System Design Tools and Methodology
Manager, System Design Tools and Methodology

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 200,000 - 380,000
Lead Hardware Design Validation Engineer
Lead Hardware Design Validation Engineer

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 168,000 - 322,000
Senior Manager, Test, Manufacturability, Reliability and Quality
Senior Manager, Test, Manufacturability, Reliability and Quality

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 232,000 - 368,000