Erhalte mehr Antworten von Arbeitgebern
Versende in nur wenigen Minuten einen passgenauen Lebenslauf.
Nebius in Berlin is seeking a Senior HPC Cluster Engineer to advance our hyperscaler platform. You will focus on GPU computing, InfiniBand networks, and the KVM/QEMU stack within a multi-GPU HPC environment.
You will analyze performance, troubleshoot root causes, integrate new hardware, and automate fault detection and resolution to keep systems fast, secure and reliable.
About Nebius:
Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure.
Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI.
Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D.
We’re looking for aSenior HPC Cluster Engineerto join our team and play a key role in the development of our cutting-edge hyperscaler platform. TheGPU & InfiniBand teamis responsible for enhancing and optimizing the core components of our Cloud platform, with a specific focus onGPU computing,InfiniBand networks, and theKVM/QEMU stack. You’ll work closely with hardware virtualization and device emulation technologies, ensuring high performance and security in multi-GPU, HPC environments. The role involves analyzing, troubleshooting, and improving infrastructure to support new hardware, fine-tuning system performance, and automating fault detection and resolution in a complex system.
In this position, you will be responsible for:
We expect you to have:
It would be a plus if you have: