Senior Systems Software Engineer - Fleet Debuggability

NVIDIA Corporation

Santa Clara (CA)

On-site

USD 184,000 - 356,500

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity
Benefits
Equal opportunity

Job summary

NVIDIA Corporation in Santa Clara, CA seeks a Senior Engineer to drive fleet-scale debugging infrastructure for multi-rack log collection and analysis. You will architect end-to-end tooling to normalize and time-align logs from kernel, drivers, and BMC interfaces, enabling rapid root-cause diagnoses across GPUs, CPUs, and network products.

You will design scalable solutions, own open-source contributions, and collaborate with developers, SWQA, and product teams while maintaining production

Qualifications

  • 10+ years in the software industry with specialization in system software and/or firmware development.
  • BS, MS, or PhD in CS, CE, EE, or related field — or equivalent experience.
  • Proven track record of shipping scalable server products or fleet-wide experience.
  • Experience with SCM (e.g., Git, Perforce) and project-management tools like Jira.
  • Strong, demonstrable skills in Python or Rust.
  • Deep Linux systems experience: kernel and driver logs, syslog, journald, and the realities of debugging on server platforms.
  • Hands-on experience with out-of-band management and platform interfaces—BMC, Redfish, IPMI, SEL.

Responsibilities

  • Architect, design, and build fleet-wide log collection and analysis solutions that aggregate signals across components, trays, and racks.
  • Develop tooling to collect, normalize, and time-align logs from heterogeneous sources — kernel and driver logs, syslog, Redfish event logs, SEL, firmware and BMC logs — over both in-band and out-of-band channels.
  • Build and maintain a log catalog and taxonomy that maps raw log signatures to fault classes, severity, and remediation guidance, making triage repeatable rather than tribal knowledge.
  • Develop debug and root-case tooling that turns high-volume fleet logs into ranked, actionable diagnoses for hardware, firmware, and platform faults.
  • Drive the design for collecting and analyzing logs at fleet scale while keeping overhead on production compute nodes low.
  • Partner with all matrixed organizations—developers, SWQA, and product engineering—in a fast-moving environment with end-to-end logging solutions, event schemas, and the contract between log producers and your tooling.
  • Steward the project’s open-source release: keep internal and public code paths clean, review community contributions, and represent the tooling in upstream discussions.
  • Write design docs and own end-to-end delivery, working across teams from definition through implementation, debugging, testing, and early customer support.
  • Perform code reviews and partner with development and QA to strengthen unit testing, integration coverage, and test plans.
  • Track work through Jira and bug-management tools and build a realistic end-to-end execution plan in collaboration with other engineers and managers.

Job description

NVIDIA’s invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined modern computer graphics, and revolutionized parallel computing. More recently, GPU deep learning ignited modern deep learning — the next era of computing — with the GPU acting as the brain of computers, robots, and self-driving cars that can perceive and understand the world. Today, NVIDIA is increasingly known as "the AI computing company."

We are the Datacenter System Software team and are looking for a highly motivated, creative Senior Engineer to drive Fleet‑Scale Debuggability end to end. The role focuses on designing, architecting, and building infrastructure, tooling, and analytics to collect multi‑rack‑scale logs. Solutions must normalize, correlate, and reason over logs from multiple components, trays, or racks including NVIDIA’s GPUs, CPUs, and Network products, fetched in‑band or out‑of‑band to help triage fleet‑level issues seen by our customers. Your work directly shortens the path from a raw, noisy log stream to actionable root causes.

Responsibilities
  • Architect, design, and build fleet‑wide log collection and analysis solutions that aggregate signals across components, trays, and racks.
  • Develop tooling to collect, normalize, and time‑align logs from heterogeneous sources — kernel and driver logs, syslog, Redfish event logs, SEL, firmware and BMC logs — over both in‑band and out‑of‑band channels.
  • Build and maintain a log catalog and taxonomy that maps raw log signatures to fault classes, severity, and remediation guidance, making triage repeatable rather than tribal knowledge.
  • Develop debug and root‑cause tooling that turns high‑volume fleet logs into ranked, actionable diagnoses for hardware, firmware, and platform faults.
  • Drive the design for collecting and analyzing logs at fleet scale while keeping overhead on production compute nodes low.
  • Partner with all matrixed organizations—developers, SWQA, and product engineering—in a fast‑moving environment with end‑to‑end logging solutions, event schemas, and the contract between log producers and your tooling.
  • Steward the project’s open‑source release: keep internal and public code paths clean, review community contributions, and represent the tooling in upstream discussions.
  • Write design docs and own end‑to‑end delivery, working across teams from definition through implementation, debugging, testing, and early customer support.
  • Perform code reviews and partner with development and QA to strengthen unit testing, integration coverage, and test plans.
  • Track work through Jira and bug‑management tools and build a realistic end‑to‑end execution plan in collaboration with other engineers and managers.
Qualifications
  • 10+ years in the software industry with specialization in system software and/or firmware development.
  • BS, MS, or PhD in CS, CE, EE, or a related technical field — or equivalent experience.
  • Proven track record of shipping scalable server products or fleet‑wide experience.
  • A self‑starter who loves finding creative solutions to complicated problems, with excellent written and oral communication skills—including executive‑level reporting—strong work ethic, and dedication to teamwork.
  • Flexibility to work and communicate effectively across teams, partners, and time zones.
  • Experience with SCM (e.g., Git, Perforce) and project‑management tools like Jira.
  • Strong, demonstrable skills in Python or RUST.
  • Deep Linux systems experience: kernel and driver logs, syslog, journald, and the realities of debugging on server platforms.
  • Hands‑on experience with out‑of‑band management and platform interfaces—BMC, Redfish, IPMI, SEL—and an understanding of in‑band vs. out‑of‑band trade‑offs.
  • Strong skills in log parsing, normalization, and structured logging, and comfort designing schemas and taxonomies for machine‑readable events.
Ways to Stand Out
  • Experience leading debuggability solutions on sophisticated rack‑scale compute architectures like GB200/GB300 NVL72.
  • Familiarity with log and telemetry analytics stacks (e.g., OpenSearch/ELK, Loki, Prometheus, Grafana, PagerDuty) and time‑series databases.
  • Hands‑on experience with x86/ARM system architecture and coding (C/C++, Python).
  • Track record of integrating AI/LLM tooling into engineering workflows— for triage, validation, log analysis, or test generation.
  • Experience standing up follow‑the‑sun support organizations with measurable response SLAs.
  • Experience contributing to or maintaining open‑source projects, including managing the boundary between internal and public code.
Salary & Benefits

Base salary ranges:
•Level4:$184,000–$287,500
•Level5:$224,000–$356,500

You will also be eligible for equity and benefits. Applications for this job will be accepted at least until July24, 2026.

NVIDIA is considered one of the technology world’s most desirable employers. We have some of the most forward‑thinking and hardworking people on the planet working for us. If you are creative and autonomous, we want to hear from you!

NVIDIA uses AI tools in its recruiting processes. NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

Learn more about NVIDIA.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Systems Software Engineer - Fleet Debuggability
Senior Systems Software Engineer - Fleet Debuggability

NVIDIA AI • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior Systems Software Engineer - Fleet Debuggability
Senior Systems Software Engineer - Fleet Debuggability

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Benefits
Principal Software Engineer, Rack-Scale System Software — CSP Engagements
Principal Software Engineer, Rack-Scale System Software — CSP Engagements

NVIDIA Corporation • Santa Clara (CA), Northern (KY)

Hybrid
USD 272,000 - 431,000
Equity
Benefits
System Software Engineer – Data Center Compute Diagnostics
System Software Engineer – Data Center Compute Diagnostics

NVIDIA Corporation • Durham (NC)

On-site
USD 152,000 - 242,000
Equity
Benefits package
Global Factory Systems Engineering Manager - Diagnostics
Global Factory Systems Engineering Manager - Diagnostics

NVIDIA • Santa Clara (CA)

On-site
USD 224,000 - 356,500
System Software Engineer – Data Center Compute Diagnostics
System Software Engineer – Data Center Compute Diagnostics

NVIDIA Gruppe • Durham (NC)

On-site
USD 152,000 - 242,000
Equity
Benefits package
Principal Platform Software Engineer - RAS
Principal Platform Software Engineer - RAS

2100 NVIDIA USA • Santa Clara (CA)

On-site
USD 272,000 - 431,000
Equity
Benefits
Senior System Software Engineer – Data Center Compute Diagnostics
Senior System Software Engineer – Data Center Compute Diagnostics

NVIDIA AI • Durham (NC)

On-site
USD 224,000 - 357,000
Equity
Benefits
Senior System Software Engineer, Enterprise MODS
Senior System Software Engineer, Enterprise MODS

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 184,000 - 288,000
Equity
Benefits
Systems Software Engineer - Infrastructure
Systems Software Engineer - Infrastructure

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 152,000 - 242,000
Equity
Comprehensive benefits