Senior HPC Engineer - GPU Compute & InfiniBand

Nebius

United States

Remote

USD 150,000 - 230,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Nebius is building a cutting-edge AI cloud platform for developers and enterprises, blending GPU orchestration with InfiniBand networking. We seek a Senior HPC Cluster Engineer to optimize GPU clusters, InfiniBand fabrics, and virtualization stacks, enabling high-performance, secure multi-GPU HPC environments.

In this role you will integrate new hardware, automate fault detection, and drive performance improvements across Kubernetes, QEMU, and KVM stacks, collaborating with a global,

Qualifications

  • 5+ years of system-level software development focused on performance optimization.
  • 3+ years of hands-on Linux system administration, troubleshooting, and tuning.
  • Deep understanding of server architectures, PCIe devices, NICs, Linux OS/kernel, and HPC systems.
  • Proficiency in C/C++, Go, and Python.

Responsibilities

  • Tune the performance of GPU clusters and InfiniBand networks for HPC and GPU environments.
  • Analyze and troubleshoot the root cause of GPU/InfiniBand issues and propose corrective actions.
  • Integrate new hardware into the infrastructure, supporting GPU hardware via Kubernetes, QEMU, and KVM.
  • Enhance automation for proactive monitoring and fault detection in GPU/InfiniBand systems.
  • Configure and manage GPU devices and InfiniBand fabrics for reliable operation.

Skills

C/C++
Go
Python
Linux
PCIe
HPC

Tools

KVM
QEMU
Kubernetes

Job description

Nebius is building a cutting-edge AI cloud platform for developers and enterprises, blending GPU orchestration with InfiniBand networking. We seek a Senior HPC Cluster Engineer to optimize GPU clusters, InfiniBand fabrics, and virtualization stacks, enabling high-performance, secure multi-GPU HPC environments.

In this role you will integrate new hardware, automate fault detection, and drive performance improvements across Kubernetes, QEMU, and KVM stacks, collaborating with a global,

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior HPC Systems Engineer: GPU Clusters & AI Infra
Senior HPC Systems Engineer: GPU Clusters & AI Infra

Nebius • United States

Remote
USD 180,000 - 240,000
Competitive pay
Career growth
Flexibility and ownership
+3
Senior AI Systems Engineer — GPU/HPC Troubleshooter
Senior AI Systems Engineer — GPU/HPC Troubleshooter

Nebius • United States

On-site
USD 180,000 - 224,000
Health insurance
401(k) match
Parental leave
+2
Senior HPC Support Engineer: InfiniBand & NVLink | Equity
Senior HPC Support Engineer: InfiniBand & NVLink | Equity

NVIDIA Corporation • Redmond (WA)

On-site
USD 108,000 - 207,000
Equity
Benefits package
Senior HPC & InfiniBand Support Engineer (Equity)
Senior HPC & InfiniBand Support Engineer (Equity)

NVIDIA • Santa Clara (CA)

On-site
USD 108,000 - 207,000
Equity
Benefits package
Senior HPC Support Engineer – InfiniBand/NVLink Expert
Senior HPC Support Engineer – InfiniBand/NVLink Expert

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 120,000 - 230,000
Equity
Benefits package
Senior HPC Deployment Lead - AI Data Centers & InfiniBand
Senior HPC Deployment Lead - AI Data Centers & InfiniBand

NVIDIA • North Carolina

On-site
USD 216,000 - 397,000
Equity
Benefits
Senior HPC Support Engineer (InfiniBand) - Equity Eligible
Senior HPC Support Engineer (InfiniBand) - Equity Eligible

NVIDIA • New York (NY)

On-site
USD 108,000 - 207,000
Equity
Comprehensive benefits
Senior HPC Network Engineer – InfiniBand & NVLink
Senior HPC Network Engineer – InfiniBand & NVLink

NVIDIA Corporation • New York (NY), Northern (KY)

Hybrid
USD 108,000 - 207,000
Equity
Benefits package
Senior Network Engineer — AI Infra & HPC Fabric Expert
Senior Network Engineer — AI Infra & HPC Fabric Expert

Nscale • Houston (TX)

On-site
USD 150,000 - 210,000
Competitive benefits package
Flexible paid time off
Parental leave
+1
Senior Network Solutions Engineer for AI HPC & InfiniBand
Senior Network Solutions Engineer for AI HPC & InfiniBand

NVIDIA • Santa Clara (CA)

On-site
USD 168,000 - 322,000
Equity
Benefits