Senior HPC Systems Engineer: GPU/InfiniBand & KVM

Nebius

United States

On-site

USD 170,000 - 300,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive compensation
Career growth
Flexible and ownership culture
Opportunity to work on impactful AI

Job summary

Nebius is seeking a Senior Software Systems Engineer to advance our hyperscaler platform. You will work on GPU computing, InfiniBand networks, and the KVM/QEMU stack, improving performance and security in multi‑GPU HPC environments.

Role involves analyzing, troubleshooting, and automating fault detection within a complex system, with a focus on hardware virtualization and device emulation. Collaboration with hardware teams is essential.

Qualifications

  • 5+ years of professional experience in system‑level software development with focus on performance optimization.
  • 3+ years of hands‑on Linux system administration and tuning.
  • Deep understanding of server architecture, PCIe devices, NICs, Linux OS/Kernel and HPC systems.
  • Proficiency in C/C++, Go or Python for performance‑oriented programming.
  • Experience with GPU/InfiniBand in clusters and virtualization stacks (KVM/QEMU).

Responsibilities

  • Tune GPU clusters and InfiniBand networks for high‑performance operation in HPC environments.
  • Analyze and troubleshoot GPU/InfiniBand issues, propose corrective actions.
  • Integrate new hardware into existing infra, including GPU support via Kubernetes, QEMU and KVM.
  • Enhance automation for monitoring and proactive fault detection in GPU/InfiniBand setups.
  • Configure and manage GPU devices and InfiniBand fabrics for reliability and performance.

Tools

Kubernetes
QEMU
KVM

Job description

Nebius is seeking a Senior Software Systems Engineer to advance our hyperscaler platform. You will work on GPU computing, InfiniBand networks, and the KVM/QEMU stack, improving performance and security in multi‑GPU HPC environments.

Role involves analyzing, troubleshooting, and automating fault detection within a complex system, with a focus on hardware virtualization and device emulation. Collaboration with hardware teams is essential.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior HPC Engineer - GPU Compute & InfiniBand
Senior HPC Engineer - GPU Compute & InfiniBand

Nebius • United States

Remote
USD 150,000 - 230,000
Senior HPC Systems Engineer: GPU Clusters & AI Infra
Senior HPC Systems Engineer: GPU Clusters & AI Infra

Nebius • United States

Remote
USD 180,000 - 240,000
Competitive pay
Career growth
Flexibility and ownership
+3
Senior Systems Software Engineer, GPU Compute
Senior Systems Software Engineer, GPU Compute

Nebius • United States

On-site
USD 170,000 - 300,000
Competitive compensation
Career growth
Flexible and ownership culture
+1
Senior HPC Support Engineer: InfiniBand & NVLink
Senior HPC Support Engineer: InfiniBand & NVLink

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 108,000 - 207,000
Senior InfiniBand & HPC Networking Engineer
Senior InfiniBand & HPC Networking Engineer

NVIDIA • Redmond (WA)

On-site
USD 108,000 - 173,000
Equity
Benefits
Lead GPU Performance Engineer — HPC Systems
Lead GPU Performance Engineer — HPC Systems

Nebius • United States

On-site
USD 170,000 - 300,000
Health insurance
401(k) plan
Parental leave
+2
Senior HPC & Infiniband Network Engineer
Senior HPC & Infiniband Network Engineer

Nscale • New York (NY)

On-site
USD 150,000 - 210,000
Medical insurance
Dental insurance
Vision insurance
+3
Senior HPC Networking Engineer: InfiniBand & NVLink
Senior HPC Networking Engineer: InfiniBand & NVLink

Nvidia Corporation in • Westford (MA)

On-site
USD 108,000 - 207,000
Equity
Benefits package
Senior HPC Support Engineer: InfiniBand & NVLink Expert
Senior HPC Support Engineer: InfiniBand & NVLink Expert

NVIDIA • Durham (NC)

On-site
USD 120,000 - 207,000
Equity
Comprehensive benefits package
Senior HPC Support Engineer (InfiniBand) - Equity Eligible
Senior HPC Support Engineer (InfiniBand) - Equity Eligible

NVIDIA • New York (NY)

On-site
USD 108,000 - 207,000
Equity
Comprehensive benefits