Senior Systems Software Engineer, AI Stack and Performance - DGX Station

Nvidia Corporation

Santa Clara (CA)

On-site

USD 224,000 - 356,500

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Nvidia Corporation is looking for an experienced professional to optimize AI workloads at DGX Station (Galaxy) in Santa Clara, California. This role involves ensuring production readiness of AI applications, profiling multi-GPU environments, and collaborating to enhance NVIDIA's software stack.

The ideal candidate has over 12 years of experience in systems software engineering and strong expertise with deep learning frameworks. The position offers an attractive salary with equity and benefits.

Qualifications

  • 12+ years in systems software engineering with hands-on experience.
  • Strong proficiency with deep learning frameworks.
  • Experience profiling and optimizing GPU workloads.

Responsibilities

  • Define and validate production readiness of AI applications.
  • Profile and optimize workloads for multi-GPU architectures.
  • Collaborate to improve kernel fusion and memory management.

Skills

AI/ML workload optimization
GPU performance analysis
Deep learning frameworks (PyTorch, TensorFlow, JAX)
C/C++
CUDA
Python

Education

BS or MS in Computer Science or related field

Tools

Nsight Systems
Nsight Compute
CUPTI

Job description

About DGX Station (Galaxy)

DGX Station (Galaxy) is a workstation-class AI computer built on GB300 Blackwell GPUs with NVLink interconnect, delivering data‑center‑grade AI compute in a deskside form factor. It ships with a complete software and firmware GA release that includes DGX BaseOS, GPU drivers, CUDA toolkit, DCGM, and DOCA/OFED. AI applications such as NemoClaw, LLM inference via NIM, Hermes agents, and deep learning frameworks run production‑ready out of the box, optimized for this multi‑GPU, high‑bandwidth architecture.

Responsibilities
  • Own production readiness of AI applications on DGX Station, define "ready to ship" criteria, run validation, and close every gap between "it runs" and "it runs well" across single‑GPU and multi‑GPU configurations.
  • Profile and optimize LLM and deep learning workloads (PyTorch, TensorFlow, JAX) for training and inference on the GB300 Blackwell multi‑GPU architecture. Characterize performance across model sizes, batch sizes, precision modes (FP16, INT8, FP8), and GPU scaling to establish benchmarks and identify regressions.
  • Identify bottlenecks in GPU compute, NVLink bandwidth, host memory, PCIe, and CPU‑GPU communication. Implement or drive optimizations across the stack—including kernel tuning, memory placement, NVLink utilization, data pipeline efficiency, and scheduling—to increase throughput on DGX Station’s multi‑GPU topology.
  • Collaborate with NVIDIA’s framework, compiler (TensorRT, NVCC, Triton), and GPU architecture teams to improve kernel fusion, graph execution, operator scheduling, and memory management for Blackwell GPUs. Translate platform‑specific constraints into actionable optimization requests for upstream teams.
  • Validate multi‑user and concurrent workload scenarios, ensuring reliable performance as a shared workstation with MIG or time‑slicing for resource isolation.
  • Validate the full NVIDIA AI software stack on DGX Station, ensuring version compatibility, functional correctness, and performance parity with reference data‑center configurations.
  • Build and maintain performance benchmarking infrastructure for DGX Station, automating regression tracking across key models, framework versions, and driver updates. Make performance data actionable for GA release decisions.
  • Work with product management and OEM/OSV partners to understand target use cases and support customer deployment readiness and field critical issues.
Qualifications
  • BS or MS or equivalent experience in Computer Science, Electrical Engineering, or related field.
  • 12+ years in systems software engineering with hands‑on experience in AI/ML workload optimization, GPU performance analysis, or deep learning infrastructure.
  • Strong proficiency with deep learning frameworks—PyTorch, TensorFlow, or JAX—including internals: graph execution, operator dispatch, memory management, and custom kernel integration.
  • Experience profiling and optimizing GPU workloads using Nsight Systems, Nsight Compute, CUPTI, or equivalent; ability to read GPU traces and translate observations into actionable optimizations.
  • Strong understanding of GPU architecture: compute units, memory hierarchy, NVLink, multi‑GPU scaling, and their impact on AI workload performance.
  • Experience with inference optimization: quantization (INT8/FP8), model compilation (TensorRT, torch.compile), batching strategies, and serving frameworks.
  • Proficiency in C/C++, CUDA, and Python; comfortable reading and modifying GPU kernels.
Ways to Stand Out
  • Experience optimizing LLM training or inference on multi‑GPU NVIDIA systems such as DGX, HGX, or multi‑GPU workstations.
  • Contributions to open‑source AI frameworks, CUDA libraries, or inference engines.
  • Experience with multi‑GPU communication optimization—NCCL tuning, NVLink utilization, collective operations, and parallel training strategies.
  • Track record of collaborating with compiler and hardware architecture teams to drive kernel fusion, graph optimization, or hardware‑specific performance improvements.
  • Experience shipping AI‑powered products where application performance on specific hardware was a hard shipping requirement.

Salary: 224,000 USD – 356,500 USD for the base, with equity and benefits available.

Applications will be accepted until June 5, 2026.

NVIDIA is committed to fostering an inclusive work environment and is proud to be an equal opportunity employer. We do not discriminate on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Systems Software Engineer, Windows and Linux Enablement - DGX Station
Senior Systems Software Engineer, Windows and Linux Enablement - DGX Station

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 224,000 - 357,000
Equity
Comprehensive benefits
Manager, Distinguished Engineer - DGX Systems Software
Manager, Distinguished Engineer - DGX Systems Software

2100 NVIDIA USA • Santa Clara (CA)

On-site
USD 320,000 - 489,000
Senior Full Stack Software Engineer - DGX Cloud
Senior Full Stack Software Engineer - DGX Cloud

NVIDIA AI • Santa Clara (CA)

On-site
USD 184,000 - 288,000
Equity
Benefits
Senior Performance Engineer - DGX Cloud
Senior Performance Engineer - DGX Cloud

NVIDIA AI • Eugene (OR)

On-site
USD 224,000 - 432,000
Equity
Benefits
Senior Software Engineer, Distributed Systems Engineer - DGX Cloud
Senior Software Engineer, Distributed Systems Engineer - DGX Cloud

Nvidia Corporation • Santa Clara (CA)

On-site
USD 152,000 - 287,500
Equity
Benefits
Senior Systems Software Engineer, Windows and Linux Enablement - DGX Station
Senior Systems Software Engineer, Windows and Linux Enablement - DGX Station

2100 NVIDIA USA • Santa Clara (CA)

On-site
USD 224,000 - 357,000
Senior Software Engineer, Distributed Systems Engineer - DGX Cloud
Senior Software Engineer, Distributed Systems Engineer - DGX Cloud

NVIDIA AI • Santa Clara (CA)

On-site
USD 152,000 - 288,000
Senior Systems Software Engineer, Accelerated Kubernetes Performance and Scale - DGX Cloud
Senior Systems Software Engineer, Accelerated Kubernetes Performance and Scale - DGX Cloud

NVIDIA • Seattle (WA)

Hybrid
USD 184,000 - 357,000
Equity
Benefits
Senior Customer Success Engineer - DGX Cloud
Senior Customer Success Engineer - DGX Cloud

NVIDIA AI • United States Virgin Islands

On-site
USD 200,000 - 322,000
Equity
Benefits
Senior Full-Stack Software Engineer
Senior Full-Stack Software Engineer

Jobtailor • California (MO)

On-site
USD 224,000 - 357,000
Comprehensive benefits package
Equity options