Senior HPC Engineer, GPU Compute

Meyandy LLC

Berlin

Vor Ort

EUR 120.000 - 180.000

Vollzeit

14 Tage+
Bewerbungsgenerator

Mach aus dieser Rolle ein Bewerbungsgespräch — ein Lebenslauf und ein Anschreiben, die genau auf das zugeschnitten sind, was dieser Arbeitgeber sucht.

Schaffe es an den ATS-Filtern vorbei

Zusammenfassung

Nebius in Berlin is seeking a Senior HPC Cluster Engineer to advance our hyperscaler platform. You will focus on GPU computing, InfiniBand networks, and the KVM/QEMU stack within a multi-GPU HPC environment.

You will analyze performance, troubleshoot root causes, integrate new hardware, and automate fault detection and resolution to keep systems fast, secure and reliable.

Qualifikationen

  • 5+ years of system-level software development focused on performance optimization.
  • 3+ years of hands-on Linux systems administration, troubleshooting, and performance tuning.
  • Deep understanding of server architecture including PCIe devices, NICs, Linux OS/kernel and HPC systems.
  • Strong proficiency in performance-oriented languages (C/C++, Go, Python).

Aufgaben

  • Tuning the performance of GPU clusters and InfiniBand networks to ensure optimal operation in HPC and GPU-based environments.
  • Analyzing and troubleshooting the root cause of issues related to GPUs and InfiniBand networks, and proposing corrective actions.
  • Integrating new hardware into the existing infrastructure, including support for new GPU hardware through software stacks like Kubernetes, QEMU, and KVM.
  • Enhancing automation systems for proactive monitoring, detecting, and resolving issues in GPU and InfiniBand environments.
  • Configuring and managing GPU devices and InfiniBand fabrics, ensuring efficient and reliable operation.

Kenntnisse

System-level software
Linux systems
Server architecture
C/C++/Go/Python

Jobbeschreibung

About Nebius:

Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure.

Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI.

Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D.

The role

We’re looking for aSenior HPC Cluster Engineerto join our team and play a key role in the development of our cutting-edge hyperscaler platform. TheGPU & InfiniBand teamis responsible for enhancing and optimizing the core components of our Cloud platform, with a specific focus onGPU computing,InfiniBand networks, and theKVM/QEMU stack. You’ll work closely with hardware virtualization and device emulation technologies, ensuring high performance and security in multi-GPU, HPC environments. The role involves analyzing, troubleshooting, and improving infrastructure to support new hardware, fine-tuning system performance, and automating fault detection and resolution in a complex system.

In this position, you will be responsible for:

  • Tuning the performance of GPU clusters and InfiniBand networks toensure optimal operation in HPC and GPU-based environments.
  • Analyzing and troubleshooting the root cause of issues related to GPUs and InfiniBand networks, and proposing corrective actions.
  • Integrating new hardware into the existing infrastructure, includingsupport for new GPU hardware through software stacks like Kubernetes, QEMU, and KVM.
  • Enhancing automation systems for proactive monitoring, detecting, and resolving issues in GPU and InfiniBand environments.
  • Configuring and managing GPU devices and InfiniBand fabrics, ensuring efficient and reliable operation.

We expect you to have:

  • 5+ years of professional experience insystem-level software development(focused on performance optimization, low-level programming).
  • 3+ years of hands-on experience withLinux systems(administration, troubleshooting, and performance tuning).
  • In-depth understanding of server architecture, including PCIe devices, NICs, Linux OS/Kernel, and high-performance computing (HPC) systems.
  • Strong proficiency in one or moreperformance-oriented programming languages(C/C++, Go, Python).

It would be a plus if you have:

  • Experience withGPU end-to-end testingin acluster environmentusing InfiniBand networking.
  • Proven track record of analyzing and optimizing the performance ofHPC workloads(e.g., simulations, data analysis, AI/ML workloads).
  • Familiarity withRDMA, RoCE, and InfiniBandprotocols for high-performance communication.
Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.

oder ziehe deine Datei hierhin.

Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Senior HPC Engineer, GPU Compute arbeitnow Nebius Berlin · 9/28/2026
Senior HPC Engineer, GPU Compute arbeitnow Nebius Berlin · 9/28/2026

Primetime • Berlin

Hybrid
EUR 90.000 - 130.000
Competitive pay
Career growth
Flexible work
+3
Senior HPC AI Cluster Engineer
Senior HPC AI Cluster Engineer

NVIDIA • Deutschland

Vor Ort
EUR 120.000 - 160.000
Senior Solutions Architect, HPC and AI
Senior Solutions Architect, HPC and AI

NVIDIA • Berlin

Vor Ort
EUR 110.000 - 170.000
Senior Cloud Infrastructure and DevOps Solutions Architect
Senior Cloud Infrastructure and DevOps Solutions Architect

NVIDIA • Deutschland

Vor Ort
EUR 120.000 - 180.000
Senior Solutions Architect, HPC and AI
Senior Solutions Architect, HPC and AI

NVIDIA Corporation • Berlin

Vor Ort
EUR 120.000 - 170.000
Senior HPC Cluster Administrator - Deep Learning Frameworks Infrastructure
Senior HPC Cluster Administrator - Deep Learning Frameworks Infrastructure

NVIDIA • Deutschland

Vor Ort
USD 58.306 - 101.063
Head of Compute Engineering
Head of Compute Engineering

Impossible Cloud GmbH • Hamburg

Vor Ort
EUR 80.000 - 100.000
Competitive salary
ESOP
Subsidized gym membership
+1
Senior HPC GPU Cluster Lead for Deep Learning Infra
Senior HPC GPU Cluster Lead for Deep Learning Infra

NVIDIA • Deutschland

Vor Ort
USD 58.306 - 101.063
GPU Cluster Engineer (human)
GPU Cluster Engineer (human)

NEURA Robotics • Deutschland

Vor Ort
USD 140.000 - 195.000
ML Infrastructure Engineer
ML Infrastructure Engineer

Nebius B.V. • Deutschland

Vor Ort
EUR 90.000 - 120.000
Competitive compensation
Career growth and learning
Flexibility and ownership
+3