Senior Systems Software Engineer

Nebius B.V.

Deutschland

Remote

EUR 146.000 - 258.000

Vollzeit

Vor 4 Tagen
Sei unter den ersten Bewerbenden

Erhalte mehr Antworten von Arbeitgebern

Versende in nur wenigen Minuten einen passgenauen Lebenslauf.

Zusammenfassung

Nebius B.V. is seeking a Senior Software Systems Engineer to advance our hyperscaler platform, focusing on GPU computing, InfiniBand networks, and the KVM/QEMU stack. You will work with hardware virtualization and device emulation to push performance and security in multi-GPU HPC environments.

Responsibilities include tuning, troubleshooting, and integrating new hardware, plus expanding automated fault detection and resolution within a complex system.

Qualifikationen

  • 5+ years in system-level software development focusing on performance.
  • 3+ years Linux administration, troubleshooting and performance tuning.
  • Deep understanding of PCIe devices, NICs, Linux kernel, and HPC systems.
  • Strong proficiency in C/C++, Go, or Python.

Aufgaben

  • Tune GPU clusters and InfiniBand networks for HPC and GPU-based environments.
  • Analyze and troubleshoot root causes, proposing corrective actions.
  • Integrate new hardware including GPU support via Kubernetes, QEMU, and KVM.
  • Enhance automation for proactive monitoring, detecting, and resolving issues.
  • Configure and manage GPU devices and InfiniBand fabrics for reliability.

Kenntnisse

System-level development
Linux systems
Performance optimization
C/C++
Go
Python

Tools

KVM/QEMU
Kubernetes

Jobbeschreibung

We’re looking for a Senior Software Systems Engineer to join our team and play a key role in the development of our cutting-edge hyperscaler platform. The GPU & InfiniBand team is responsible for enhancing and optimizing the core components of our Cloud platform, with a specific focus on GPU computing, InfiniBand networks, and the KVM/QEMU stack. You’ll work closely with hardware virtualization and device emulation technologies, ensuring high performance and security in multi-GPU, HPC environments. The role involves analyzing, troubleshooting, and improving infrastructure to support new hardware, fine-tuning system performance, and automating fault detection and resolution in a complex system.

In this position, you will be responsible for:
  • Tuning the performance of GPU clusters and InfiniBand networks to ensure optimal operation in HPC and GPU-based environments.
  • Analyzing and troubleshooting the root cause of issues related to GPUs and InfiniBand networks, and proposing corrective actions.
  • Integrating new hardware into the existing infrastructure, including support for new GPU hardware through software stacks like Kubernetes, QEMU, and KVM.
  • Enhancing automation systems for proactive monitoring, detecting, and resolving issues in GPU and InfiniBand environments.
  • Configuring and managing GPU devices and InfiniBand fabrics, ensuring efficient and reliable operation.
We expect you to have:
  • 5+ years of professional experience in system-level software development (focused on performance optimization, low-level programming).
  • 3+ years of hands-on experience with Linux systems (administration, troubleshooting, and performance tuning).
  • In-depth understanding of server architecture, including PCIe devices, NICs, Linux OS/Kernel, and high-performance computing (HPC) systems.
  • Strong proficiency in one or more performance-oriented programming languages (C/C++, Go, Python).
It would be a plus if you have:
  • Experience with GPU end-to-end testing in a cluster environment using InfiniBand networking.
  • Proven track record of analyzing and optimizing the performance of HPC workloads (e.g., simulations, data analysis, AI/ML workloads).
  • Familiarity with RDMA, RoCE, and InfiniBand protocols for high-performance communication.
  • Background in Software-Defined Networking (SDN) and experience with HPC cluster networking.
  • Understanding of QEMU/KVM virtualization and managing virtualized environments.
  • Experience with deep learning frameworks such as PyTorch and TensorFlow, and their integration with HPC systems.
  • Familiarity with collective communication libraries like MPI and NCCL for distributed computing.

We offer competitive salaries ranging from $170k-$300k + equity based on your experience.

We conduct coding interviews as part of the process.

Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.
oder ziehe deine Datei hierhin.
Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Senior HPC Engineer, GPU Compute
Senior HPC Engineer, GPU Compute

Meyandy LLC • Berlin

Hybrid
EUR 120.000 - 180.000
Hardware Engineer
Hardware Engineer

Nebius B.V. • Deutschland

Remote
EUR 90.000 - 130.000
Senior HPC GPU Cluster Lead for Deep Learning Infra
Senior HPC GPU Cluster Lead for Deep Learning Infra

NVIDIA • Deutschland

Vor Ort
USD 58.306 - 101.063
Senior Solutions Architect, Infiniband and Networking Ethernet
Senior Solutions Architect, Infiniband and Networking Ethernet

NVIDIA Gruppe • München

Vor Ort
EUR 75.000 - 100.000
Senior HPC Cluster Administrator - Deep Learning Frameworks Infrastructure
Senior HPC Cluster Administrator - Deep Learning Frameworks Infrastructure

NVIDIA • Berlin

Vor Ort
EUR 120.000 - 180.000
Infrastructure/Systems Engineer
Infrastructure/Systems Engineer

DUDE CHEM • Berlin

Hybrid
EUR 90.000 - 130.000
Remote/Hybrid Berlin
Home office budget
30 days vacation
+1
Senior HPC Cluster Administrator - Deep Learning Frameworks Infrastructure
Senior HPC Cluster Administrator - Deep Learning Frameworks Infrastructure

NVIDIA • Deutschland

Vor Ort
USD 58.306 - 101.063
Software Python Engineer (GPU Cloud)
Software Python Engineer (GPU Cloud)

Talanto • Deutschland

Hybrid
EUR 90.000 - 140.000
Remote/Hybrid work
Private medical insurance
Paid sick leave
+4
Senior HPC Performance Engineer: CPU/GPU Optimizations
Senior HPC Performance Engineer: CPU/GPU Optimizations

NVIDIA • Deutschland

Vor Ort
USD 120.000 - 180.000
Senior System Engineer (Munich, Germany)
Senior System Engineer (Munich, Germany)

Remotestar • München

Hybrid
EUR 80.000 - 110.000
Indefinite contract
Equal pay guaranteed
Variable performance bonus
+8