Senior System Software Engineer - GPU Performance

NVIDIA Gruppe

Santa Clara (CA)

On-site

USD 152,000 - 241,500

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity
Benefits

Job summary

NVIDIA Gruppe is seeking a motivated Performance Engineer to influence the roadmap of our communication libraries. The role involves conducting in-depth performance characterization on large multi-GPU and multi-node clusters and studying the interaction of our libraries with hardware and software components.

The ideal candidate will have an M.S. or Ph.D. in Computer Science, experience in parallel programming and communication runtimes, and a passion for learning new areas and tools. The base salary ranges from 152,000 to 241,500 USD based on level and experience.

Qualifications

  • 3+ years of experience with parallel programming and communication runtime.
  • Experience conducting performance benchmarking on large-scale HPC clusters.
  • Ability to debug performance issues across hardware and software.

Responsibilities

  • Conduct in-depth performance characterization on multi-GPU clusters.
  • Study interaction of libraries with hardware and software components.
  • Evaluate proof-of-concepts and conduct trade-off analysis.

Skills

Performance benchmarking
Parallel programming
Troubleshooting network issues
Scripting (Python)
HPC experience

Education

M.S. or Ph.D. in Computer Science or related field

Tools

MPI
NCCL
CUDA
Docker
Kubernetes

Job description

We are the GPU Communications Libraries and Networking team at NVIDIA and are looking for a motivated Performance Engineer to influence the roadmap of our communication libraries. The DL and HPC applications of today have a huge compute demand and run at scales that reach tens of thousands of GPUs.

What you will be doing:
  • Conduct in-depth performance characterization and analysis on large multi‑GPU and multi‑node clusters.
  • Study the interaction of our libraries with all hardware (GPU, CPU, networking) and software components in the stack.
  • Evaluate proof‑of‑concepts and conduct trade‑off analysis when multiple solutions are available.
  • Triage and root‑cause performance issues reported by our customers.
  • Collect a lot of performance data; build tools and infrastructure to visualize and analyze the information.
  • Collaborate with a very dynamic team across multiple time zones.
What we need to see:
  • M.S. (or equivalent experience) or Ph.D. in Computer Science, or a related field with relevant performance engineering and HPC experience.
  • 3+ years of experience with parallel programming and at least one communication runtime (MPI, NCCL, UCX, NVSHMEM).
  • Experience conducting performance benchmarking and triage on large‑scale HPC clusters.
  • Good understanding of computer system architecture, hardware–software interactions, and operating systems principles.
  • Implement micro‑benchmarks in C/C++, read and modify the code base when required.
  • Ability to debug performance issues across the entire hardware–software stack; proficient in a scripting language, preferably Python.
  • Familiarity with containers, cloud provisioning, and scheduling tools (Kubernetes, SLURM, Ansible, Docker).
  • Adaptability and passion to learn new areas and tools; flexibility to work and communicate effectively across different teams and time zones.
Ways to stand out from the crowd:
  • Practical experience with Infiniband/Ethernet networks in areas such as RDMA, topologies, and congestion control.
  • Experience debugging network issues in large‑scale deployments.
  • Familiarity with CUDA programming and/or GPUs.
  • Experience with deep learning frameworks such as PyTorch or TensorFlow.

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary ranges are 152,000 USD – 241,500 USD for Level 3 and 184,000 USD – 287,500 USD for Level 4.

You will also be eligible for equity and benefits.

NVIDIA is committed to fostering a diverse work environment and is proud to be an equal opportunity employer. We do not discriminate on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status, or any other characteristic protected by law.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior System Software Engineer - GPU Performance
Senior System Software Engineer - GPU Performance

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 152,000 - 288,000
Equity
Benefits
Senior HPC Performance Engineer
Senior HPC Performance Engineer

NVIDIA • Indiana (PA)

On-site
USD 221,000 - 384,000
Diversity and inclusion
Extensive benefits package
Senior Systems Software Engineer - GPU Performance at Scale
Senior Systems Software Engineer - GPU Performance at Scale

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 184,000 - 356,500
Senior GPU Performance Engineer - HPC & Networking Equity
Senior GPU Performance Engineer - HPC & Networking Equity

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 152,000 - 242,000
Equity
Benefits
Senior Systems Software Engineer - GPU Performance at Scale
Senior Systems Software Engineer - GPU Performance at Scale

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 184,000 - 288,000
Equity
Benefits
Principal Software Engineer, E2E Performance and Goodput - CSP Engagements
Principal Software Engineer, E2E Performance and Goodput - CSP Engagements

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 272,000 - 432,000
Equity
Comprehensive benefits
Senior Systems Software Engineer - GPU Performance at Scale
Senior Systems Software Engineer - GPU Performance at Scale

NVIDIA AI • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior Systems Software Engineer - GPU Performance at Scale
Senior Systems Software Engineer - GPU Performance at Scale

NVIDIA • Town of Texas (WI)

On-site
USD 184,000 - 288,000
Equity
Benefits
Senior Data Center Performance Engineer - Benchmarking and Optimization
Senior Data Center Performance Engineer - Benchmarking and Optimization

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 184,000 - 287,500
Equity
Benefits
Senior Systems Software Engineer - GPU Performance at Scale
Senior Systems Software Engineer - GPU Performance at Scale

NVIDIA • California (MO)

On-site
USD 184,000 - 288,000
Equity
Comprehensive benefits