Principal Software Engineer - Fleet Reliability & Equity

NVIDIA Corporation

Santa Clara, Northern (CA, KY)

Hybrid

USD 272,000 - 431,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

NVIDIA Corporation is seeking a Principal Software Engineer to lead CSP-facing reliability initiatives for fleet-scale software and firmware. You will coordinate cross-team efforts to ensure MTBI goals are achieved in production, leveraging telemetry, failure analysis, and predictive models across GPU/accelerator deployments.

You will drive burn-in and certification frameworks, integrate health monitoring with CSP workflows, and translate fleet insights into engineering priorities for hardware,

Qualifications

  • 15+ years of experience in systems software at datacenter scale or reliability engineering with at-scale focus
  • BS or MS in Computer Science, Electrical Engineering, Statistics, or related field (or equivalent experience)
  • Deep expertise in multi-NUMA, rack-scale system software and firmware
  • Statistical failure analysis methods: MTBF/MTBI calculation, Pareto analysis, root cause classification
  • Experience with fleet-level telemetry and observability systems: time-series databases, anomaly detection, health scoring, event correlation
  • Understanding of hardware failure modes in large-scale GPU/accelerator deployments — classification across compute, interconnect, memory, power, and thermal domains
  • Experience defining or operating burn-in, stress testing, or certification frameworks for complex hardware systems
  • Familiarity with predictive maintenance or anomaly detection approaches applied to fleet health data
  • Customer obsession and ability to translate fleet reliability challenges into actionable priorities
  • Strong communication to both technical audiences and executive leadership
  • Demonstrated success driving cross-functional improvements across hardware, firmware, and software teams

Responsibilities

  • Drive reliability work streams with CSP engineering teams — shared MTBI measurement, failure classification, and health monitoring architecture
  • Gather and synthesize CSP fleet reliability data — identify cross-customer failure patterns and drive improvements to firmware, driver, and hardware teams
  • Define MTBI measurement methodology that works across CSP monitoring environments and practices
  • Conduct fleet-scale failure pattern analysis using Pareto, survival analysis, and Weibull methods
  • Drive fleet health monitoring integration architecture — align NVIDIA health agents, telemetry, and reporting with CSP workflows
  • Define burn-in reliability test environment and cluster certification criteria with quality teams
  • Collaborate with CSPs to ensure reliability-related integration work is complete before at-scale launches
  • Develop predictive failure models using fleet telemetry and validate them in customer environments

Skills

Systems software at scale
Reliability engineering
Statistical failure analysis
Fleet telemetry and observability
MTBI/MTBF measurement
Hardware failure modes (GPU/accelerate
Burn-in, stress testing, certification
Cross-functional collaboration
Predictive maintenance
Communication to execs and technical

Education

BS or MS in Computer Science, Electrical Engineering, Statistics, or related field

Tools

Time-series databases
Observability systems
Telemetry pipelines
Health scoring

Job description

NVIDIA Corporation is seeking a Principal Software Engineer to lead CSP-facing reliability initiatives for fleet-scale software and firmware. You will coordinate cross-team efforts to ensure MTBI goals are achieved in production, leveraging telemetry, failure analysis, and predictive models across GPU/accelerator deployments.

You will drive burn-in and certification frameworks, integrate health monitoring with CSP workflows, and translate fleet insights into engineering priorities for hardware,

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Principal Software Engineer — Fleet MTBI (Equity)
Principal Software Engineer — Fleet MTBI (Equity)

NVIDIA • Santa Clara (CA)

On-site
USD 272,000 - 431,250
Equity and benefits
Senior GPU Firmware & Fleet-Scale Systems Engineer
Senior GPU Firmware & Fleet-Scale Systems Engineer

NVIDIA • California (MO)

On-site
USD 272,000 - 432,000
Fleet Reliability Architect for CSPs
Fleet Reliability Architect for CSPs

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 272,000 - 432,000
Equity and benefits
Principal Software Engineer, At-Scale Reliability and Fleet Intelligence — CSP Engagements
Principal Software Engineer, At-Scale Reliability and Fleet Intelligence — CSP Engagements

NVIDIA • Santa Clara (CA)

On-site
USD 272,000 - 431,250
Equity and benefits
Principal Software Engineer, At-Scale Reliability and Fleet Intelligence — CSP Engagements
Principal Software Engineer, At-Scale Reliability and Fleet Intelligence — CSP Engagements

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 272,000 - 431,250
Equity and benefits
Principal Software Engineer, At-Scale Reliability and Fleet Intelligence — CSP Engagements
Principal Software Engineer, At-Scale Reliability and Fleet Intelligence — CSP Engagements

NVIDIA Corporation • Santa Clara (CA), Northern (KY)

Hybrid
USD 272,000 - 431,000
Principal GPU Firmware & System Software Architect - Fleet
Principal GPU Firmware & System Software Architect - Fleet

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 272,000 - 431,250
Equity and benefits
Senior Software Engineer: Fleet Intelligence & GPU Telemetry
Senior Software Engineer: Fleet Intelligence & GPU Telemetry

NVIDIA • New York (NY)

On-site
USD 152,000 - 242,000
Senior Software Engineer, Fleet Intelligence & Telemetry
Senior Software Engineer, Fleet Intelligence & Telemetry

Nvidia Corporation in • Santa Clara (CA)

On-site
USD 152,000 - 242,000
Equity
Benefits
Principal Software Engineer, Rack-Scale System Software - CSP Engagements
Principal Software Engineer, Rack-Scale System Software - CSP Engagements

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 272,000 - 432,000
Equity
Comprehensive benefits package