Principal System Software Engineer - Data Center MODS

NVIDIA Gruppe

Santa Clara (CA)

On-site

USD 272,000 - 431,250

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

NVIDIA Gruppe is hiring a Principal Engineer in Santa Clara, CA to architect and scale diagnostic systems for Cloud Service Providers. This role involves defining technical strategies and leading development teams to ensure robust diagnostic frameworks for AI products.

The ideal candidate will have over 15 years of experience in distributed systems, programming skills in C++ or Python, and a strong background in x86/ARM architectures and Linux internals. A competitive salary package, including equity and benefits, is offered.

Qualifications

  • 15+ years of system software experience with distributed systems.
  • Deep systems knowledge of x86/ARM architectures.
  • Consistent track record of technical leadership.

Responsibilities

  • Define technical strategy for NVIDIA’s Data Center diagnostic systems.
  • Mentor and grow engineering teams while providing technical leadership.
  • Drive root-cause analysis of systemic failures across hardware and software.

Skills

Distributed systems
Programming in C++
Programming in Python
Technical leadership
Software testing methodologies

Education

Bachelor's degree in Computer Science/Engineering or related field

Tools

Linux OS internals
UEFI/BIOS
Redfish
HMC
BMC protocols

Job description

The Data Center MODS organization seeks a Principal Engineer to architect and scale next‑generation L10 and L11 diagnostic systems for Cloud Service Providers (CSPs). In this high‑impact role, you will define the technical roadmap and lead multi‑functional development to deploy robust diagnostic frameworks for AI accelerator products. Proficiency in distributed systems and hardware/software interfaces is essential for success.

Responsibilities
  • Define technical strategy and development of NVIDIA’s Data Center diagnostic systems, orchestrating large‑scale stress testing for CPUs, GPUs, networking, memory, and high‑speed interconnects.
  • Mentor and grow engineering teams, providing technical leadership and encouraging a culture of innovation and excellence.
  • Drive the root‑cause analysis of systemic failures that intersect multiple hardware and software domains.
  • Partner with CSPs to diagnose and address scalability challenges within their unique data center infrastructures.
Qualifications
  • Bachelor’s degree in Computer Science/Engineering, Electrical Engineering, or a related field (or equivalent experience).
  • 15+ years of system software experience working on highly resilient distributed systems with programming experience in C++ or Python.
  • Deep systems knowledge of x86/ARM architectures, Linux OS internals, firmware (UEFI/BIOS), Redfish, HMC, BMC protocols and platform security.
  • Consistent track record demonstrating technical leadership leading project teams and setting technical direction.
  • Expertise in software testing methodologies with an automation‑led, AI‑first approach to ensuring software quality.

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 272,000 USD - 431,250 USD.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until March 10, 2026.

NVIDIA uses AI tools in its recruiting processes.

NVIDIA is committed to fostering a diverse work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Principal System Software Engineer - Data Center MODS
Principal System Software Engineer - Data Center MODS

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 272,000 - 432,000
Senior System Software Engineer, Enterprise MODS
Senior System Software Engineer, Enterprise MODS

NVIDIA AI • Eugene (OR)

On-site
USD 184,000 - 288,000
Equity
Benefits
Senior System Software Engineer, Enterprise MODS
Senior System Software Engineer, Enterprise MODS

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 184,000 - 288,000
Equity
Benefits
Principal Software Engineer – CSP Engagements
Principal Software Engineer – CSP Engagements

NVIDIA • Santa Clara (CA)

On-site
USD 272,000 - 431,250
Equity
Benefits package
Senior System Software Engineer – Data Center Compute Diagnostics
Senior System Software Engineer – Data Center Compute Diagnostics

NVIDIA AI • Durham (NC)

On-site
USD 224,000 - 357,000
Equity
Benefits
System Software Engineer – Data Center Compute Diagnostics
System Software Engineer – Data Center Compute Diagnostics

NVIDIA Gruppe • Durham (NC)

On-site
USD 152,000 - 242,000
Equity
Benefits package
System Software Engineer – Data Center Compute Diagnostics
System Software Engineer – Data Center Compute Diagnostics

NVIDIA Corporation • Durham (NC)

On-site
USD 152,000 - 242,000
Equity
Benefits package
Senior Systems Software Engineer, Data Center Platform Enablement
Senior Systems Software Engineer, Data Center Platform Enablement

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 184,000 - 356,500
System Software Engineer – Data Center Compute Diagnostics
System Software Engineer – Data Center Compute Diagnostics

NVIDIA • Durham (NC)

On-site
USD 152,000 - 242,000
Equity
Benefits package
Senior System Software Engineer – Data Center Compute Diagnostics
Senior System Software Engineer – Data Center Compute Diagnostics

NVIDIA • Durham (NC)

On-site
USD 180,000 - 240,000
Equity
Benefits package