HPC Performance Engineer — Large-Scale GPU Clusters

NVIDIA AI

Val-de-Travers

Vor Ort

CHF 120.000 - 180.000

Vollzeit

Vor 8 Tagen

Erhalte mehr Antworten von Arbeitgebern

Versende in nur wenigen Minuten einen passgenauen Lebenslauf.

Zusammenfassung

NVIDIA AI in Switzerland seeks a performance engineer to drive optimization of HPC communication libraries and large-scale GPU workloads. You will characterize performance on multi-GPU clusters, study hardware-software interactions, and guide roadmap choices.

You will implement C/C++ micro-benchmarks, debug across HW/SW, and script with Python. Containers, Kubernetes, and SLURM enable scalable testing in a collaborative, global team.

Qualifikationen

  • MS or PhD in Computer Science or related field with performance engineering and HPC experience.
  • 3+ years of experience with parallel programming and a communication runtime (MPI, NCCL, UCX, NVSHMEM).
  • Experience benchmarking and triaging performance on large HPC clusters.
  • Strong understanding of computer system architecture and OS principles.
  • Proficiency in C/C++ micro-benchmarks and code modification when required.
  • Proficient in Python scripting and debugging performance issues across HW/SW.
  • Familiar with containers and scheduling tools (Kubernetes, SLURM, Docker, Ansible).
  • Willingness to learn new areas and collaborate across time zones.

Aufgaben

  • Conduct in-depth performance characterization on large multi-GPU and multi-node clusters.
  • Study interactions of libraries with hardware and software components in the stack.
  • Evaluate concepts and perform trade-offs when multiple solutions exist.
  • Triage and root-cause performance issues reported by users.
  • Collect and analyze performance data; build visualization tools.
  • Collaborate with a dynamic team across multiple time zones.

Kenntnisse

Parallel programming
MPI
NCCL
UCX
NVSHMEM
C/C++
Python scripting
Docker
Kubernetes
SLURM
Ansible

Ausbildung

MS or PhD in CS or related field

Tools

MPI
NCCL
UCX
NVSHMEM
C/C++
Python
Docker
Kubernetes
SLURM
Ansible

Jobbeschreibung

NVIDIA AI in Switzerland seeks a performance engineer to drive optimization of HPC communication libraries and large-scale GPU workloads. You will characterize performance on multi-GPU clusters, study hardware-software interactions, and guide roadmap choices.

You will implement C/C++ micro-benchmarks, debug across HW/SW, and script with Python. Containers, Kubernetes, and SLURM enable scalable testing in a collaborative, global team.

Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.
oder ziehe deine Datei hierhin.
Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Senior HPC AI Cluster Architect
Senior HPC AI Cluster Architect

CH01 NVIDIA Switzerland AG • Schweiz

Vor Ort
CHF 120.000 - 180.000
Senior HPC Performance Engineer — GPU Optimization
Senior HPC Performance Engineer — GPU Optimization

name • Schweiz

Vor Ort
CHF 140.000 - 230.000
Competitive salaries
Generous benefits package
Equal opportunity employer
HPC-AI Cluster Architect for Next-Gen Systems
HPC-AI Cluster Architect for Next-Gen Systems

NVIDIA • Zürich

Vor Ort
CHF 170.000 - 210.000
Senior HPC Performance Engineer
Senior HPC Performance Engineer

NVIDIA AI • Val-de-Travers

Vor Ort
CHF 120.000 - 180.000
Remote Senior Performance Engineer - AI & HPC Systems
Remote Senior Performance Engineer - AI & HPC Systems

NVIDIA Corporation • Zürich

Vor Ort
CHF 140.000 - 210.000
Senior HPC‑AI Systems Architect
Senior HPC‑AI Systems Architect

NVIDIA AI • Zürich

Vor Ort
CHF 180.000 - 260.000
Senior HPC AI Cluster Engineer
Senior HPC AI Cluster Engineer

CH01 NVIDIA Switzerland AG • Schweiz

Vor Ort
CHF 120.000 - 180.000
Senior HPC AI Cluster Engineer
Senior HPC AI Cluster Engineer

NVIDIA • Zürich

Vor Ort
CHF 170.000 - 210.000
Senior Performance Engineer
Senior Performance Engineer

NVIDIA • Schweiz

Vor Ort
CHF 150.000 - 210.000
HPC & AI Cluster Engineer: Scale & Automate
HPC & AI Cluster Engineer: Scale & Automate

NVIDIA Corporation • Zürich

Vor Ort
CHF 120.000 - 190.000