Senior HPC GPU Cluster Engineer

Nebius Group

Madrid

Presencial

EUR 80.000 - 120.000

Jornada completa

14 días+
Generador de candidaturas

Destaca en este puesto — crea un currículum adaptado y una carta de presentación en aproximadamente un minuto.

Supera los filtros ATS

Ventajas ofrecidas por este puesto de trabajo

Competitive compensation
Career growth
Flexible work
Collaborative culture
Impactful AI projects
International environment

Descripción de la vacante

Nebius is hiring a Senior HPC Cluster Engineer to advance our hyperscaler platform. You will optimize GPU clusters, work with InfiniBand networks, and collaborate on virtualization stacks like QEMU/KVM and Kubernetes to push performance and reliability at scale.

The role focuses on analyzing, troubleshooting, and automating fault detection in multi-GPU HPC environments, with close collaboration across hardware and software teams.

Formación

  • 5+ years of professional experience in system-level software development, focused on performance optimization.
  • 3+ years of hands-on experience with Linux systems (administration, troubleshooting, and performance tuning).
  • In-depth understanding of server architecture, including PCIe devices, NICs, Linux OS/Kernel, and HPC systems.
  • Strong proficiency in performance-oriented programming languages (C/C++, Go, Python).

Responsabilidades

  • Tune the performance of GPU clusters and InfiniBand networks for HPC and GPU-based environments.
  • Analyze and troubleshoot root causes of issues related to GPUs and InfiniBand networks; propose corrective actions.
  • Integrate new hardware into the existing infrastructure, including support for new GPU hardware via software stacks like Kubernetes, QEMU, and KVM.
  • Enhance automation systems for proactive monitoring, detecting, and resolving issues in GPU/InfiniBand environments.
  • Configure and manage GPU devices and InfiniBand fabrics for efficient, reliable operation.

Conocimientos

System-level dev
Linux systems
Server architecture
C/C++
Go
Python
MPI
NCCL
RDMA/InfiniBand

Herramientas

KVM
QEMU
Kubernetes

Descripción del empleo

Nebius is hiring a Senior HPC Cluster Engineer to advance our hyperscaler platform. You will optimize GPU clusters, work with InfiniBand networks, and collaborate on virtualization stacks like QEMU/KVM and Kubernetes to push performance and reliability at scale.

The role focuses on analyzing, troubleshooting, and automating fault detection in multi-GPU HPC environments, with close collaboration across hardware and software teams.

Consigue la evaluación confidencial y gratuita de tu currículum.

o arrastra y suelta tu archivo aquí

Similar jobs

Puestos de trabajo similares que vale la pena comparar

Senior HPC Cluster Engineer
Senior HPC Cluster Engineer

Nebius Group • Madrid

Presencial
EUR 80.000 - 120.000
Competitive compensation
Career growth
Flexible work
+3
Senior HPC Systems Administrator — Linux/Cluster
Senior HPC Systems Administrator — Linux/Cluster

Atos SE • Madrid

Presencial
EUR 42.000 - 65.000
Senior Cloud Infrastructure and DevOps Solutions Architect
Senior Cloud Infrastructure and DevOps Solutions Architect

NVIDIA • España

Presencial
EUR 90.000 - 130.000
Senior Cloud Infra & DevOps Architect for AI/HPC
Senior Cloud Infra & DevOps Architect for AI/HPC

NVIDIA • España

Presencial
EUR 90.000 - 130.000
Remote HPC Network Engineer: InfiniBand & AI Infra
Remote HPC Network Engineer: InfiniBand & AI Infra

Mirantis • Barcelona

Presencial
EUR 60.000 - 90.000
Senior HPC Networking Architect - InfiniBand & Fortinet
Senior HPC Networking Architect - InfiniBand & Fortinet

Mirantis • Barcelona

Presencial
EUR 90.000 - 120.000
Remote HPC Architect: Scalable Clusters & Cloud Networking
Remote HPC Architect: Scalable Clusters & Cloud Networking

Quantori • España

Híbrido
EUR 90.000 - 130.000
Remote or office work
Healthcare benefits
Professional development
Network Engineer
Network Engineer

European Tech Recruit • España

Presencial
EUR 90.000 - 130.000
Principal HPC Network Engineer (remote in the EU)
Principal HPC Network Engineer (remote in the EU)

Mirantis • Barcelona

Presencial
EUR 90.000 - 120.000
HPC Administrator
HPC Administrator

Atos SE • Madrid

Presencial
EUR 42.000 - 65.000