Principal Software Engineer, Rack-Scale System Software - CSP Engagements

NVIDIA Gruppe

Santa Clara (CA)

On-site

USD 272,000 - 431,250

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity
Comprehensive benefits package

Job summary

NVIDIA Gruppe is seeking a Principal Software Engineer for the CSP Engagements team in Santa Clara, California. The role involves being the technical focal point for rack-scale system software and firmware, ensuring reliable deployment and operation of complex systems at scale.

Ideal candidates will have over 15 years of experience in system software and distributed systems. Responsibilities include driving SW/FW architecture alignment, collaborating across teams, and synthesizing engineering feedback. Exceptional communication and customer obsession are essential to understanding CSP operations.

Qualifications

  • 15+ years of experience in system software, platform firmware, or large-scale distributed systems engineering.
  • Deep understanding of multi-component coordination, error propagation, and reliability.
  • Experience with telemetry systems and fleet-level observability.

Responsibilities

  • Drive SW/FW architecture alignment across CSP engagements.
  • Collaborate with multi-functional teams for operational requirements.
  • Capture and synthesize engineering feedback on system software.

Skills

System software engineering
Distributed systems
Fabric management software
Error handling
Health monitoring
Technical leadership
Communication

Education

BS or MS in Computer Science, Electrical Engineering, or related field

Job description

We're looking for a Principal Software Engineer to join our CSP Engagements team as the technical focal point for rack‑scale system SW/FW, working with CSP engineering teams to ensure they can deploy, monitor, and operate these systems reliably at fleet scale.

Responsibilities
  • Drive rack‑scale SW/FW architecture alignment across CSP engagements — including fabric management software, link health monitoring, GPU/NVSwitch error handling, SW/FW serviceability features (e.g., hot‑plug support, component isolation, firmware‑driven recovery), and multi‑component firmware orchestration.
  • Drive technical work streams with CSP engineering teams on rack‑scale system software — ensuring they deeply understand fabric management, NVSwitch behavior, error handling and recovery policies, health telemetry APIs, and SW/FW‑controlled recovery operation.
  • Capture and synthesize CSP engineering feedback on rack‑scale system software — health monitoring APIs, SW‑driven serviceability workflows, firmware update orchestration, and error recovery behavior — champion that feedback into NVIDIA’s architecture decisions.
  • Collaborate with multi‑functional teams to ensure customer operational requirements are reflected in system software and firmware development.
  • Identify cross‑CSP patterns in rack‑scale SW/FW issues, error handling behavior, and system configuration practices — drive documentation, tooling, and test strategy improvements as a result.
  • Collaborate with execution teams on left‑shift strategy — ensuring customer‑side SW/FW integration work is identified early and completed ahead of hardware availability.
  • Make critical technical decisions on rack‑scale system SW/FW tradeoffs and mitigate execution risks through early engagement with CSP engineering teams.
Qualifications
  • 15+ years of experience in system software, platform firmware, or large‑scale distributed systems engineering. BS or MS in Computer Science, Electrical Engineering, or related field (or equivalent experience).
  • Deep understanding of rack‑scale system software challenges: multi‑component coordination, error propagation, health monitoring, and serviceability / reliability.
  • Experience with fabric management software, cluster management, or system‑level orchestration frameworks. Familiarity with firmware architectures and update lifecycle management (multi‑component update sequencing, rollback, recovery).
  • Understanding of error handling and recovery design patterns in distributed systems — fault isolation, retry policies, graceful degradation.
  • Experience with health monitoring and telemetry systems: health scoring, event correlation, API design for fleet‑level observability.
  • Understanding of GPU or accelerator system software (drivers, device management, power management) is a strong plus.
  • Customer obsession — genuine passion for understanding how CSPs operate sophisticated systems at fleet scale and simplifying their experience.
  • Proven success providing technical leadership across organizational boundaries and influencing system software design without direct authority. Strong communication — ability to translate complex system software architecture into actionable mentorship for customer engineering teams.
Ways to Stand Out
  • Experience with NVIDIA NVSwitch, NVOS, or GPU fabric management software.
  • Background in system software for large‑scale clusters at a hyperscaler (cluster management, fleet orchestration, health platforms).
  • Experience crafting error handling and recovery frameworks for multi‑component systems (hundreds or thousands of coordinating devices).
  • Familiarity with GPU or accelerator fleet operations — driver lifecycle, firmware rollout strategies, health‑based scheduling.
  • Understanding of how system software decisions impact serviceability, availability, and operational cost at fleet scale.

Salary: 272,000 USD - 431,250 USD.

Employees will also be eligible for equity and benefits.

We are an equal opportunity employer. NVIDIA is committed to fostering an inclusive work environment and is proud to be an equal opportunity employer. We do not discriminate on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Principal Software Engineer, Rack-Scale System Software — CSP Engagements
Principal Software Engineer, Rack-Scale System Software — CSP Engagements

NVIDIA • Austin (TX)

On-site
USD 272,000 - 432,000
Principal Software Engineer, Rack-Scale System Software — CSP Engagements
Principal Software Engineer, Rack-Scale System Software — CSP Engagements

NVIDIA • Oregon (WI)

On-site
USD 272,000 - 432,000
Equity
Benefits
Principal Software Engineer, Rack-Scale System Software — CSP Engagements
Principal Software Engineer, Rack-Scale System Software — CSP Engagements

NVIDIA • California (MO)

On-site
USD 272,000 - 432,000
Equity
Comprehensive benefits package
Principal Software Engineer, Rack-Scale System Software — CSP Engagements
Principal Software Engineer, Rack-Scale System Software — CSP Engagements

NVIDIA Corporation • Santa Clara (CA), Northern (KY)

Hybrid
USD 272,000 - 431,000
Equity
Benefits
Principal Software Engineer – Rack-Scale System Software
Principal Software Engineer – Rack-Scale System Software

Jobtailor • California (MO)

On-site
USD 180,000 - 260,000
Principal Software Engineer, GPU Firmware and GPU System Software - CSP Engagements
Principal Software Engineer, GPU Firmware and GPU System Software - CSP Engagements

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 272,000 - 432,000
Principal Software Engineer, GPU Firmware and GPU System Software — CSP Engagements
Principal Software Engineer, GPU Firmware and GPU System Software — CSP Engagements

2100 NVIDIA USA • Santa Clara (CA)

On-site
USD 272,000 - 431,250
Equity and benefits
Principal Software Engineer, GPU Firmware and GPU System Software — CSP Engagements
Principal Software Engineer, GPU Firmware and GPU System Software — CSP Engagements

NVIDIA • California (MO)

On-site
USD 272,000 - 432,000
Principal Software Engineer - Rack Scale Systems Infrastructure
Principal Software Engineer - Rack Scale Systems Infrastructure

2100 NVIDIA USA • Santa Clara (CA)

On-site
USD 272,000 - 431,250
Principal Software Engineer, At-Scale Reliability and Fleet Intelligence — CSP Engagements
Principal Software Engineer, At-Scale Reliability and Fleet Intelligence — CSP Engagements

NVIDIA • Santa Clara (CA)

On-site
USD 272,000 - 431,250
Equity and benefits