Senior HPC Performance Engineer: Multi-GPU Networking

NVIDIA

California (MO)

On-site

USD 184,000 - 288,000

Full time

3 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Equity
Benefits

Job summary

NVIDIA in the United States is seeking a Performance Engineer to advance the state-of-the-art in HPC libraries and multi-GPU communication. You will drive performance analysis across large clusters and study HW/SW interactions to optimize overall system efficiency.

You will characterize workloads, develop micro-benchmarks, and build tooling to visualize data, collaborating with teams across multiple time zones to push the cutting edge of GPU communications.

Qualifications

  • MS or PhD in Computer Science or related field with perf/HPC experience.
  • 3+ years of experience with parallel programming and at least one communication runtime (MPI, NCCL, UCX, NVSHMEM).
  • Experience benchmarking and triage on large-scale HPC clusters.
  • Good understanding of computer architecture and systems software fundamentals.
  • Implement micro-benchmarks in C/C++, read/modify code when required.
  • Proficient in Python scripting; able to debug across HW/SW stack.
  • Familiar with containers and orchestration: Kubernetes, SLURM, Ansible, Docker.
  • Adaptability and ability to collaborate across time zones.

Responsibilities

  • Conduct in-depth performance characterization on large multi-GPU and multi-node clusters.
  • Study interaction of libraries with HW and SW components in the stack.
  • Evaluate proofs-of-concepts and trade-offs when multiple solutions exist.
  • Triage and root-cause performance issues reported by customers.
  • Collect performance data and build tools to visualize and analyze it.
  • Collaborate with a dynamic team across multiple time zones.

Skills

Parallel programming
Performance benchmarking
Systems software fundamentals
C/C++ micro-benchmarks
Python scripting
Container orchestration
HW/SW stack debugging

Education

MS or PhD in Computer Science

Tools

Kubernetes
SLURM
Ansible
Docker

Job description

NVIDIA in the United States is seeking a Performance Engineer to advance the state-of-the-art in HPC libraries and multi-GPU communication. You will drive performance analysis across large clusters and study HW/SW interactions to optimize overall system efficiency.

You will characterize workloads, develop micro-benchmarks, and build tooling to visualize data, collaborating with teams across multiple time zones to push the cutting edge of GPU communications.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior HPC Middleware Engineer - High-Perf Networking
Senior HPC Middleware Engineer - High-Perf Networking

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 170,000 - 280,000
Senior HPC Middleware Engineer — Equity
Senior HPC Middleware Engineer — Equity

NVIDIA • California (MO)

On-site
USD 184,000 - 287,500
Senior HPC Middleware Engineer - MPI/RDMA & Performance
Senior HPC Middleware Engineer - MPI/RDMA & Performance

NVIDIA • Boulder (CO)

On-site
USD 152,000 - 287,500
Equity
Benefits
Senior HPC Middleware Engineer (Remote)
Senior HPC Middleware Engineer (Remote)

NVIDIA Corporation • Santa Clara (CA)

Hybrid
USD 184,000 - 288,000
Equity
Benefits
Senior HPC Middleware Engineer (MPI/RDMA & Networking)
Senior HPC Middleware Engineer (MPI/RDMA & Networking)

NVIDIA • Colorado

On-site
USD 152,000 - 287,500
Equity
Benefits
Senior GPU Performance Engineer - HPC & Networking Equity
Senior GPU Performance Engineer - HPC & Networking Equity

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 152,000 - 241,500
Equity
Benefits
Senior HPC Middleware Engineer - MPI/RDMA & Equity
Senior HPC Middleware Engineer - MPI/RDMA & Equity

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 288,000
Senior HPC Performance Engineer for AI-Driven Science
Senior HPC Performance Engineer for AI-Driven Science

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 287,500
Equity
Senior HPC Middleware Engineer – Equity & Performance
Senior HPC Middleware Engineer – Equity & Performance

NVIDIA • Illinois

On-site
USD 184,000 - 287,500
Equity
Benefits
Senior System Software Engineer - GPU Performance
Senior System Software Engineer - GPU Performance

NVIDIA • California (MO)

On-site
USD 184,000 - 288,000
Equity
Benefits