Senior Systems Performance Engineer

NVIDIA Gruppe

Santa Clara (CA)

On-site

USD 136,000 - 212,750

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

NVIDIA Gruppe in Santa Clara is seeking a Senior Validation Engineer for the DGX Server Product Engineering Team. In this role, you will work closely with HW/SW engineers to develop automated test plans for leading GPU computing products. Responsibilities include system architecture design and developing stress testing strategies.

The ideal candidate will have over 5 years of experience, a BSEE or BSCE, and strong skills in Dynamo, TensorRT, and Python programming. The position offers competitive salaries based on experience and eligibility for equity.

Qualifications

  • 5+ years of experience in validating and debugging complex systems.
  • Developing/running real-world ML/LLM workload.
  • Experience with datacenter products including system management, security, networking, and storage.

Responsibilities

  • Develop and implement complex automated test plans for GPU accelerated computing products.
  • Conduct system architecture, design, and performance modelling.
  • Enable GPU SKU bring up, validation, and model enablement.
  • Develop system-level stress and performance testing strategies using AI applications.

Skills

Dynamo
TensorRT
Slurm
BCM
Cuda
Cublas
Cutlass
Python programming
Understanding of computing architectures

Education

BSEE or BSCE or equivalent experience

Job description

Senior Validation Engineer – DGX Server Product Engineering Team

In this role you will be working with a team of HW/SW engineers to develop and implement complex automated test plans for our industry leading GPU accelerated computing products.

What you will be doing
  • System architecture, design, performance modelling, estimation across new models and new packages.
  • Enable GPU SKU bring up, validation and model enablement.
  • Develop system level stress and performance testing strategies using industry leading Deep Learning/AI applications.
What We Need to See
  • Ability to work on site in hardware lab environment 5 days a week
  • BSEE or BSCE or equivalent experience
  • 5+ years or more of experience in validating and debugging complex systems
  • Developing/running real world ML/LLM workload
  • Dynamo, TensorRT, Slurm, BCM skills mandatorily required
  • Knowledge of vLLM, SG Lang preferred
  • Proficiency in Cuda, Cublas and Cutlass
  • Deep understanding of computing architectures
  • Coding experience with python programming, running simulators
  • Experience with datacenter products including system management, security, networking, and storage
Ways to Stand Out
  • Background with x86/Arm server architectures and accelerated GPU computing
  • Track record of continuous process improvement with a passion for tools and automation
Benefits and Compensation

NVIDIA offers highly competitive salaries and a comprehensive benefits package. Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is $136,000 USD - $212,750 USD for Level 3, and $168,000 USD - $258,750 USD for Level 4. You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until April 5, 2026. This posting is for an existing vacancy.

NVIDIA uses AI tools in its recruiting processes.

EEO Statement

NVIDIA is committed to fostering a diverse work environment and proud to be an equal opportunity employer. We do not discriminate on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Systems Performance Engineer
Senior Systems Performance Engineer

NVIDIA • Santa Clara (CA)

On-site
USD 136,000 - 259,000
Senior Systems Software Engineer - GPU Performance at Scale
Senior Systems Software Engineer - GPU Performance at Scale

NVIDIA AI • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior Performance Engineer - DGX Cloud
Senior Performance Engineer - DGX Cloud

NVIDIA AI • Eugene (OR)

On-site
USD 224,000 - 432,000
Equity
Benefits
Senior Systems Software Engineer - GPU Performance at Scale
Senior Systems Software Engineer - GPU Performance at Scale

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 184,000 - 288,000
Equity
Benefits
Senior Systems Software Engineer - GPU Performance at Scale
Senior Systems Software Engineer - GPU Performance at Scale

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 184,000 - 356,500
Senior Systems Software Engineer, Data Center Platform Enablement
Senior Systems Software Engineer, Data Center Platform Enablement

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 184,000 - 356,500
Senior Software Development Engineer in Test - Datacenter Server OS
Senior Software Development Engineer in Test - Datacenter Server OS

Nvidia Corporation in • Santa Clara (CA)

On-site
USD 140,000 - 224,000
Equity
Benefits
Senior Software QA Test Development Engineer - Diagnostics
Senior Software QA Test Development Engineer - Diagnostics

NVIDIA • Santa Clara (CA)

On-site
USD 140,000 - 270,000
Equity
Benefits
Senior Full Stack Software Engineer - DGX Cloud
Senior Full Stack Software Engineer - DGX Cloud

Nvidia Corporation • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Senior Full Stack Software Engineer - DGX Cloud
Senior Full Stack Software Engineer - DGX Cloud

NVIDIA Gruppe • North Carolina

On-site
USD 224,000 - 357,000