Senior HPC Engineer - GPU Compute & InfiniBand

Nebius

United States

Remote

USD 150,000 - 230,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Nebius is building a cutting-edge AI cloud platform for developers and enterprises, blending GPU orchestration with InfiniBand networking. We seek a Senior HPC Cluster Engineer to optimize GPU clusters, InfiniBand fabrics, and virtualization stacks, enabling high-performance, secure multi-GPU HPC environments.

In this role you will integrate new hardware, automate fault detection, and drive performance improvements across Kubernetes, QEMU, and KVM stacks, collaborating with a global,

Qualifications

  • 5+ years of system-level software development focused on performance optimization.
  • 3+ years of hands-on Linux system administration, troubleshooting, and tuning.
  • Deep understanding of server architectures, PCIe devices, NICs, Linux OS/kernel, and HPC systems.
  • Proficiency in C/C++, Go, and Python.

Responsibilities

  • Tune the performance of GPU clusters and InfiniBand networks for HPC and GPU environments.
  • Analyze and troubleshoot the root cause of GPU/InfiniBand issues and propose corrective actions.
  • Integrate new hardware into the infrastructure, supporting GPU hardware via Kubernetes, QEMU, and KVM.
  • Enhance automation for proactive monitoring and fault detection in GPU/InfiniBand systems.
  • Configure and manage GPU devices and InfiniBand fabrics for reliable operation.

Skills

C/C++
Go
Python
Linux
PCIe
HPC

Tools

KVM
QEMU
Kubernetes

Job description

Nebius is building a cutting-edge AI cloud platform for developers and enterprises, blending GPU orchestration with InfiniBand networking. We seek a Senior HPC Cluster Engineer to optimize GPU clusters, InfiniBand fabrics, and virtualization stacks, enabling high-performance, secure multi-GPU HPC environments.

In this role you will integrate new hardware, automate fault detection, and drive performance improvements across Kubernetes, QEMU, and KVM stacks, collaborating with a global,

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior HPC Systems Engineer: GPU/InfiniBand & KVM
Senior HPC Systems Engineer: GPU/InfiniBand & KVM

Nebius • United States

On-site
USD 170,000 - 300,000
Competitive compensation
Career growth
Flexible and ownership culture
+1
Senior HPC Systems Engineer: GPU Clusters & AI Infra
Senior HPC Systems Engineer: GPU Clusters & AI Infra

Nebius • United States

Remote
USD 180,000 - 240,000
Competitive pay
Career growth
Flexibility and ownership
+3
Senior HPC Cloud Engineer — GPU/InfiniBand AI Infra
Senior HPC Cloud Engineer — GPU/InfiniBand AI Infra

Jobgether • Germany (OH)

On-site
USD 81,000 - 105,000
Career development opportunities
Flexible working arrangements
Collaborative engineering environment
+1
Senior Systems Software Engineer, GPU Compute
Senior Systems Software Engineer, GPU Compute

Nebius • United States

On-site
USD 170,000 - 300,000
Competitive compensation
Career growth
Flexible and ownership culture
+1
Senior HPC & Infiniband Network Engineer
Senior HPC & Infiniband Network Engineer

Nscale • New York (NY)

On-site
USD 150,000 - 210,000
Medical insurance
Dental insurance
Vision insurance
+3
Senior InfiniBand & HPC Networking Engineer
Senior InfiniBand & HPC Networking Engineer

NVIDIA • Redmond (WA)

On-site
USD 108,000 - 173,000
Equity
Benefits
Senior Network Engineer — AI Infra & HPC Fabric Expert
Senior Network Engineer — AI Infra & HPC Fabric Expert

Nscale • Houston (TX)

On-site
USD 150,000 - 210,000
Competitive benefits package
Flexible paid time off
Parental leave
+1
Senior HPC Support Engineer: InfiniBand & NVLink
Senior HPC Support Engineer: InfiniBand & NVLink

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 108,000 - 207,000
Senior HPC Networking Engineer: InfiniBand & NVLink
Senior HPC Networking Engineer: InfiniBand & NVLink

Nvidia Corporation in • Westford (MA)

On-site
USD 108,000 - 207,000
Equity
Benefits package
Senior HPC Support Engineer (InfiniBand) - Equity Eligible
Senior HPC Support Engineer (InfiniBand) - Equity Eligible

NVIDIA • New York (NY)

On-site
USD 108,000 - 207,000
Equity
Comprehensive benefits