Fleet-Scale Debuggability Engineer

NVIDIA Corporation

Santa Clara (CA)

On-site

USD 184,000 - 356,500

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity
Benefits
Equal opportunity

Job summary

NVIDIA Corporation in Santa Clara, CA seeks a Senior Engineer to drive fleet-scale debugging infrastructure for multi-rack log collection and analysis. You will architect end-to-end tooling to normalize and time-align logs from kernel, drivers, and BMC interfaces, enabling rapid root-cause diagnoses across GPUs, CPUs, and network products.

You will design scalable solutions, own open-source contributions, and collaborate with developers, SWQA, and product teams while maintaining production

Qualifications

  • 10+ years in the software industry with specialization in system software and/or firmware development.
  • BS, MS, or PhD in CS, CE, EE, or related field — or equivalent experience.
  • Proven track record of shipping scalable server products or fleet-wide experience.
  • Experience with SCM (e.g., Git, Perforce) and project-management tools like Jira.
  • Strong, demonstrable skills in Python or Rust.
  • Deep Linux systems experience: kernel and driver logs, syslog, journald, and the realities of debugging on server platforms.
  • Hands-on experience with out-of-band management and platform interfaces—BMC, Redfish, IPMI, SEL.

Responsibilities

  • Architect, design, and build fleet-wide log collection and analysis solutions that aggregate signals across components, trays, and racks.
  • Develop tooling to collect, normalize, and time-align logs from heterogeneous sources — kernel and driver logs, syslog, Redfish event logs, SEL, firmware and BMC logs — over both in-band and out-of-band channels.
  • Build and maintain a log catalog and taxonomy that maps raw log signatures to fault classes, severity, and remediation guidance, making triage repeatable rather than tribal knowledge.
  • Develop debug and root-case tooling that turns high-volume fleet logs into ranked, actionable diagnoses for hardware, firmware, and platform faults.
  • Drive the design for collecting and analyzing logs at fleet scale while keeping overhead on production compute nodes low.
  • Partner with all matrixed organizations—developers, SWQA, and product engineering—in a fast-moving environment with end-to-end logging solutions, event schemas, and the contract between log producers and your tooling.
  • Steward the project’s open-source release: keep internal and public code paths clean, review community contributions, and represent the tooling in upstream discussions.
  • Write design docs and own end-to-end delivery, working across teams from definition through implementation, debugging, testing, and early customer support.
  • Perform code reviews and partner with development and QA to strengthen unit testing, integration coverage, and test plans.
  • Track work through Jira and bug-management tools and build a realistic end-to-end execution plan in collaboration with other engineers and managers.

Job description

NVIDIA Corporation in Santa Clara, CA seeks a Senior Engineer to drive fleet-scale debugging infrastructure for multi-rack log collection and analysis. You will architect end-to-end tooling to normalize and time-align logs from kernel, drivers, and BMC interfaces, enabling rapid root-cause diagnoses across GPUs, CPUs, and network products.

You will design scalable solutions, own open-source contributions, and collaborate with developers, SWQA, and product teams while maintaining production

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Fleet-Scale Systems Engineer (Logging & Debugging)
Senior Fleet-Scale Systems Engineer (Logging & Debugging)

NVIDIA AI • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior Fleet Debuggability Architect
Senior Fleet Debuggability Architect

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior Systems Software Engineer - Fleet Debuggability
Senior Systems Software Engineer - Fleet Debuggability

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior Systems Software Engineer - Fleet Debuggability
Senior Systems Software Engineer - Fleet Debuggability

NVIDIA AI • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior Systems Software Engineer - Fleet Debuggability
Senior Systems Software Engineer - Fleet Debuggability

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 184,000 - 356,500
Equity
Benefits
Equal opportunity
Senior GPU Datacenter Debug Engineer
Senior GPU Datacenter Debug Engineer

NVIDIA • Santa Clara (CA)

On-site
USD 200,000 - 322,000
Equity
Comprehensive benefits package
NVLink Systems Software Engineer – Debug & AI Tools (Equity)
NVLink Systems Software Engineer – Debug & AI Tools (Equity)

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 152,000 - 288,000
Senior System Software Engineer - Hardware Diagnostics & AI
Senior System Software Engineer - Hardware Diagnostics & AI

NVIDIA • Durham (NC)

On-site
USD 180,000 - 240,000
Equity
Benefits package
Senior Software Engineer - Factory & Data Center Automation
Senior Software Engineer - Factory & Data Center Automation

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior Debug Systems Engineer — Datacenter GPUs
Senior Debug Systems Engineer — Datacenter GPUs

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 168,000 - 259,000
Comprehensive benefits package
Equity options