Senior GPU HPC Engineer: InfiniBand & KVM Optimization

Nebius

Greater London

On-site

GBP 90,000 - 130,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive compensation
Career growth
Flexibility and ownership
Collaborative culture
Impactful AI projects
International teams

Job summary

Nebius is seeking a Senior HPC Cluster Engineer to advance our hyperscaler platform. You will optimize GPU clusters, InfiniBand networks and the KVM/QEMU stack, collaborating with hardware virtualization and device emulation teams to ensure top performance and security in multi-GPU HPC environments.

The role emphasizes troubleshooting, integration of new hardware, and automation to sustain high-throughput AI workloads across Europe and the UK, with a fast-moving, innovative culture and strong

Qualifications

  • 5+ years of professional experience in system-level software development with a focus on performance optimization.
  • 3+ years hands-on experience with Linux systems (administration, troubleshooting, performance tuning).
  • In-depth understanding of server architecture including PCIe devices, NICs, Linux OS/Kernel, and HPC systems.
  • Strong proficiency in one or more performance-oriented programming languages (C/C++, Go, Python).

Responsibilities

  • Tune the performance of GPU clusters and InfiniBand networks for HPC and GPU-based environments.
  • Analyze and troubleshoot root causes related to GPUs and InfiniBand networks; propose corrective actions.
  • Integrate new hardware into existing infrastructure, including GPU hardware support via Kubernetes, QEMU, and KVM.
  • Enhance automation systems for proactive monitoring and fault detection in GPU and InfiniBand environments.
  • Configure and manage GPU devices and InfiniBand fabrics for efficient and reliable operation.

Skills

System-level development
Linux administration
C/C++
Go
Python

Tools

KVM/QEMU
RDMA/InfiniBand
MPI/NCCL

Job description

Nebius is seeking a Senior HPC Cluster Engineer to advance our hyperscaler platform. You will optimize GPU clusters, InfiniBand networks and the KVM/QEMU stack, collaborating with hardware virtualization and device emulation teams to ensure top performance and security in multi-GPU HPC environments.

The role emphasizes troubleshooting, integration of new hardware, and automation to sustain high-throughput AI workloads across Europe and the UK, with a fast-moving, innovative culture and strong

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior HPC Engineer, GPU Compute
Senior HPC Engineer, GPU Compute

Nebius • Greater London

On-site
GBP 90,000 - 130,000
Competitive compensation
Career growth
Flexibility and ownership
+3
HPC & AI Platform Engineer – GPU Clusters
HPC & AI Platform Engineer – GPU Clusters

Era4 • United Kingdom

Hybrid
GBP 95,000 - 130,000
Data Center IT Manager - Lead GPU Cloud Ops & IT Infra
Data Center IT Manager - Lead GPU Cloud Ops & IT Infra

Nebius • Greater London

On-site
GBP 70,000 - 100,000
Competitive compensation
Career growth and learning
Flexibility and ownership
+3
Data Center IT Manager - GPU & Infra Ops
Data Center IT Manager - GPU & Infra Ops

Nebius • Greater London

On-site
GBP 50,000 - 70,000
Competitive compensation
Career growth and learning opportunities
Flexibility and work-life balance
+3
Data Center IT Ops Lead — GPU Clusters & Reliability
Data Center IT Ops Lead — GPU Clusters & Reliability

Nebius • Newport

On-site
GBP 55,000 - 85,000
Competitive compensation
Career growth
Flexibility and ownership
+3
Senior GPU Systems Engineer - Clusters & Platform Infra
Senior GPU Systems Engineer - Clusters & Platform Infra

Radley James • Greater London

On-site
GBP 20,000 - 40,000
Senior GPU & AI Infra Architect — Remote, 4-Day Week
Senior GPU & AI Infra Architect — Remote, 4-Day Week

Civo Ltd • United Kingdom

Hybrid
GBP 110,000 - 170,000
4-day week
Uncapped holidays
Remote work environment
HPC & AI Platform Engineer (GPU/Networking)
HPC & AI Platform Engineer (GPU/Networking)

Carbon3ai Limited. • United Kingdom

Hybrid
GBP 90,000 - 120,000
Technical Solutions Architect – Investors
Technical Solutions Architect – Investors

Hamilton Barnes Associates Limited • Greater London

On-site
GBP 110,000 - 150,000
HPC Networking Architect for GPU Clusters
HPC Networking Architect for GPU Clusters

Fuse Energy • Greater London

On-site
GBP 90,000 - 120,000
Competitive salary and an equity sign‑
Biannual bonus scheme
Fully expensed tech to match your need
+1