Distinguished Engineer – Data Center System Software Architect

NVIDIA Corporation

Santa Clara (CA)

On-site

USD 320,000 - 489,000

Full time

8 days ago
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Equity
Benefits

Job summary

NVIDIA Corporation seeks a strong technical architect to own end-to-end architecture of data center products at the system software level. You will oversee firmware, kernel drivers, operating systems, and user-mode drivers.

You will lead discussions with major customers, define KPIs, gather requirements, and drive adoption of new technologies and protocols across hyperscalers.

Qualifications

  • Deep expertise in scalable server architecture with HW/SW interfaces.
  • Extensive system software experience for accelerators (GPUs/DPUs/FPGAs).
  • Firmware/embedded systems and Linux kernel internals mastery (OpenBMC, SBIOS).
  • Proficiency with management protocols (MCTP, PLDM, SPDM, RDE) and Redfish/IPMI.
  • Strong networking knowledge (TCP/IP, Ethernet, InfiniBand) and routing.
  • Ability to balance security and usability with security experts.
  • Proven cross-functional project leadership without direct authority.
  • Experience applying left-shift risk-reduction strategies.
  • BS or MS in CS/EE or related field; 20+ years in architecture.

Responsibilities

  • Lead end-to-end architecture of data center products at system software level.
  • Serve as primary technical contact for major customers and define KPIs.
  • Architect collaborations with hyperscalers to evolve product roadmaps.
  • Align NVIDIA's roadmap with customer requirements through engagement.
  • Develop and drive adoption of new technologies and protocols.

Skills

System architecture
HW/SW interfaces
GPUs/DPUs/FPGAs
Linux kernel internals
OpenBMC/SBIOS
MCTP/PLDM/SPDM
Redfish/IPMI
Networking (TCP/IP, InfiniBand)
Security tradeoffs
Cross-functional leadership
Left-shift risk strategies

Education

BS or MS in CS/EE or related field

Tools

OpenBMC
SBIOS
Redfish

Job description

NVIDIA data center systems, such as DGX and HGX, have become core to NVIDIA's rapidly growing enterprise and cloud provider businesses. These platforms bring together the full power of NVIDIA GPUs, NVIDIA NVLink, NVIDIA InfiniBand networking, NVIDIA Grace CPUs, and a fully optimized NVIDIA AI and HPC software stack. We’re looking for a strong technical architect to own the end-to-end architecture of these products, at the system software level. Including firmware, kernel drivers, operating systems, and user mode drivers.

What You’ll Be Doing
  • Serve as the primary technical point of contact for major customers, leading technological discussions, defining KPIs, gathering requirements, and addressing complex technical queries.
  • Lead technical innovation and strategic collaborations with major hyperscalers to architect next-generation data center products.
  • Align NVIDIA's roadmap with major customers' requirements through direct engagement.
  • Develop and drive adoption of new technologies and protocols.
  • Make critical technical decisions in ambiguous situations, mitigating risks through left-shift strategies.
What We Need to See
  • Deep expertise in scalable and performant server system architecture, focusing on SW/HW interfaces.
  • Extensive experience with complex system software for accelerators (GPUs, DPUs, FPGAs).
  • Mastery of system firmware (SBIOS, OpenBMC), embedded systems, and Linux kernel internals.
  • Proficiency in Out-of-Band and In-Band management architectures, device management protocols (e.g., MCTP, PLDM, SPDM, RDE) and system management protocols (Redfish, IPMI).
  • Extensive knowledge of networking technologies and protocols, including TCP/IP, Ethernet, InfiniBand, as well as advanced switching and routing concepts.
  • Experience collaborating with platform security experts to define tradeoffs between security and ease of use.
  • Demonstrated success in leading complex, cross-functional projects to completion, showcasing the ability to influence and achieve results without direct authority in large-scale, collaborative environments.
  • Demonstrable experience in implementing left shift strategy to de-risk program execution.
  • BS or MS degree in Computer Science, Electrical Engineering or related field (or equivalent experience).
  • 20+ years in system architecture and design.
Ways to Stand Out From the Crowd
  • Knowledge of cloud and cluster level deployment and management systems.
  • Participation and contributions in standards bodies such as OCP and DMTF.
  • Familiarity with NVIDIA HPC programming models and libraries (CUDA, cuDNN, DOCA).
  • Knowledge of enterprise storage architectures and distributed parallel processing paradigms.

Base salary range: $320,000 - $488,750 USD per year. Eligible for equity and benefits.

We are an equal opportunity employer. We do not discriminate on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Distinguished Engineer – Data Center System Software Architect
Distinguished Engineer – Data Center System Software Architect

Segment (Twilio) • Santa Clara (CA)

On-site
USD 320,000 - 488,750
Equity
Benefits
Distinguished Engineer – Data Center System Software Architect
Distinguished Engineer – Data Center System Software Architect

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 320,000 - 489,000
Equity options
Comprehensive benefits package
Principal Firmware Engineer – Server Manageability and Observability
Principal Firmware Engineer – Server Manageability and Observability

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 272,000 - 432,000
Equity options
Comprehensive benefits
Principal Firmware Engineer – Server Manageability and Observability
Principal Firmware Engineer – Server Manageability and Observability

Segment (Twilio) • Santa Clara (CA)

On-site
USD 272,000 - 432,000
Equity
Benefits package
Distinguished Engineer, System Software Integration
Distinguished Engineer, System Software Integration

NVIDIA Gruppe • California (MO)

On-site
USD 320,000 - 489,000
Equity
Benefits package
Senior Software Architect - Data Center Systems
Senior Software Architect - Data Center Systems

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 224,000 - 356,500
Equity
Senior Data Center System Architect
Senior Data Center System Architect

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Benefits
Distinguished Engineer - Rack Scale Architecture
Distinguished Engineer - Rack Scale Architecture

NVIDIA • Durham (NC)

On-site
USD 320,000 - 489,000
Principal Firmware Engineer – Server Manageability and Observability
Principal Firmware Engineer – Server Manageability and Observability

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 272,000 - 431,000
Distinguished Software Engineer - NVLink Fusion Software
Distinguished Software Engineer - NVLink Fusion Software

NVIDIA • Hillsboro (OR)

On-site
USD 320,000 - 489,000
Equity options
Comprehensive benefits