Senior System Software Engineer, Enterprise MODS

NVIDIA Gruppe

Santa Clara (CA)

On-site

USD 184,000 - 288,000

Full time

12 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity
Benefits

Job summary

NVIDIA is seeking a visionary Diagnostics Architect to lead development and validation of diagnostic systems for data center platforms in Santa Clara, CA. You will shape how we validate, debug, and optimize server hardware and software across ODMs, CSP deployments, and field operations.

The role requires 8+ years of experience at the SW/HW interface, deep knowledge of x86/ARM, Linux/Windows internals, firmware, BMC, and strong leadership.

Qualifications

  • Proven experience architecting diagnostics for complex server systems at the SW/HW interface.
  • Deep systems knowledge: x86/ARM, Linux/Windows OS internals, firmware, BMC, and platform security.
  • Strong collaboration with customers and multi-disciplinary teams to drive optimal solutions.
  • Hands-on with C, C++, and Python for tool development and automation.

Responsibilities

  • Develop diagnostic systems for NVIDIA data center platforms, including hardware and software tools to stress CPUs, GPUs, memory, storage, and interconnects.
  • Lead platform bring-up and integration, ensuring diagnostics are embedded early across the server lifecycle.
  • Drive hardware validation strategy with architecture and hardware teams, crafting robust validation plans.
  • Analyze root causes of complex failures and provide scalable solutions across the stack.
  • Develop diagnostics software to ensure quality and performance at scale across ODM and partner production lines.
  • Mentor and grow engineering teams, providing technical leadership and fostering innovation.
  • Influence long-term strategy by developing diagnostic architecture and roadmaps for upcoming NVIDIA products.

Skills

C
C++
Python
Diagnostics
Embedded systems
Leadership

Education

BS in Computer Science or Electrical Engineering
MS in CS or EE

Tools

Linux
Windows
UEFI/BIOS
BMC
PCIe
NVLink

Job description

At NVIDIA, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work.

The data center platforms like GB200 NVL72 by NVIDIA are redefining AI, HPC, and cloud computing. To accommodate leading workloads globally, our diagnostic systems need to evolve across diverse hardware technologies. We’re in search of a visionary technical leader to engineer and propel innovation in diagnostics for NVIDIA's partner ecosystem. This role is essential in crafting how we validate, debug, and optimize complex server platforms across ODM factories, Cloud Service Provider (CSP) deployments, and field operations.

What You’ll Be Doing:
  • Develop diagnostic systems for NVIDIA data center platforms, which involve hardware and software tools to develop the worst case stress workloads for CPUs, GPUs, memory, storage, and interconnects.
  • Lead platform bring-up and integration, ensuring diagnostics are embedded early and effectively across the server lifecycle.
  • Drive hardware validation strategy in collaboration with architecture and hardware teams, crafting robust validation plans for new server generations.
  • Analyze root causes of complex failures, acting as a Level 2 engineering contact for critical issues and offering scalable solutions across the stack.
  • Develop diagnostics software to ensure quality and performance at scale across ODM and partner production lines.
  • Mentor and grow engineering teams, providing technical leadership and encouraging a culture of innovation and excellence.
  • Influence the long-term strategy by developing diagnostic architecture and roadmaps for the upcoming products of NVIDIA and its partners.
What we need to see:
  • Proven experience architecting diagnostics for complex server systems, especially at the SW/HW interface.
  • Deep systems knowledge: x86/ARM architectures, Linux/Windows OS internals, firmware (UEFI/BIOS), BMC, and platform security.
  • Ability to weigh tradeoffs in system development and drive the most optimum solutions with customers and multi-disciplinary teams
  • Expertise in programming languages like C, C++, and Python for tool development and automation.
  • Familiarity with high-speed interconnects such as PCIe, Infiniband, NVLink, and Ethernet.
  • Strong communication skills to engage with technical and executive team.
  • BS/MS or equivalent experience in Computer Science, Electrical Engineering, or related field.
  • 8+ years of engineering experience in diagnostics, embedded systems, or cloud platforms.
Ways to stand out from the crowd:
  • Experience driving diagnostics across rack-level or cluster-level deployments.
  • Background in cloud-scale infrastructure and partner engagement.
  • Demonstrated success in influencing product direction and vendor roadmaps.
  • Passion for mentoring and building high-performing teams.

NVIDIA is at the forefront of AI, HPC, and visualization. Our diagnostics are the nervous system of our platforms—ensuring reliability, performance, and innovation at scale. If you’re a creative, driven architect ready to shape the future of diagnostics, we want to hear from you.

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until August 15, 2026.

This posting is for an existing vacancy.

NVIDIA uses AI tools in its recruiting processes.

NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior System Software Engineer, Enterprise MODS
Senior System Software Engineer, Enterprise MODS

NVIDIA AI • Eugene (OR)

On-site
USD 184,000 - 288,000
Equity
Benefits
System Software Engineer – Data Center Compute Diagnostics
System Software Engineer – Data Center Compute Diagnostics

NVIDIA Gruppe • Durham (NC)

On-site
USD 152,000 - 242,000
Equity
Benefits package
System Software Engineer – Data Center Compute Diagnostics
System Software Engineer – Data Center Compute Diagnostics

NVIDIA Corporation • Durham (NC)

On-site
USD 152,000 - 242,000
Equity
Benefits package
Senior System Software Engineer – Data Center Compute Diagnostics
Senior System Software Engineer – Data Center Compute Diagnostics

NVIDIA AI • Durham (NC)

On-site
USD 224,000 - 357,000
Equity
Benefits
Senior System Software Engineer – Data Center Compute Diagnostics
Senior System Software Engineer – Data Center Compute Diagnostics

NVIDIA • Durham (NC)

On-site
USD 180,000 - 240,000
Equity
Benefits package
System Software Engineer - Tegra
System Software Engineer - Tegra

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 124,000 - 196,000
Equity
Benefits
Principal System Software Engineer - Data Center MODS
Principal System Software Engineer - Data Center MODS

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 272,000 - 431,250
System Software Engineer – Data Center Compute Diagnostics
System Software Engineer – Data Center Compute Diagnostics

NVIDIA • Durham (NC)

On-site
USD 152,000 - 242,000
Equity
Benefits package
Global Factory Systems Engineering Manager - Diagnostics
Global Factory Systems Engineering Manager - Diagnostics

NVIDIA • Santa Clara (CA)

On-site
USD 224,000 - 356,500
Distinguished Resiliency and Safety Architect, GPU Diagnostics
Distinguished Resiliency and Safety Architect, GPU Diagnostics

NVIDIA • Santa Clara (CA)

On-site
USD 320,000 - 489,000
Equity
Benefits