Senior HPC Systems Engineer: GPU Clusters & AI Infra

Nebius

United States

Remote

USD 180,000 - 240,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive pay
Career growth
Flexibility and ownership
Collaborative culture
Impactful AI projects
International teams

Job summary

Nebius is seeking a Senior Systems HPC Engineer to optimize large-scale GPU clusters and run across the full stack from hardware to software. You will identify bottlenecks, validate performance improvements, and work with teams across infra, software, and hardware vendors.

Ideal candidates bring 5+ years in system-level software performance, Linux expertise, and strong C/C++/Go/Python skills. The role offers an international, fast-moving environment with opportunities to influence AI-scale

Qualifications

  • 5+ years in system-level software development focused on performance optimization.
  • 3+ years with Linux systems (administration, troubleshooting, performance tuning).
  • Deep understanding of server architecture including PCIe devices, NICs, Linux OS/Kernel, and HPC systems.
  • Proficiency in performance-oriented languages: C/C++, Go, Python.

Responsibilities

  • Analyze system behavior across multiple layers to identify bottlenecks and drive improvements for GPU clusters.
  • Troubleshoot performance issues under real workloads (training and inference).
  • Evaluate and integrate new hardware, configurations, and tuning approaches via software stack.
  • Collaborate with infra, software, and hardware vendor teams (NVIDIA, Mellanox, Intel) on performance.
  • Contribute to hardware and cluster qualification ensuring systems meet performance targets.
  • Support performance escalations from internal teams and customers.

Skills

System-level software
Linux systems
Performance optimization
C/C++
Python

Tools

KVM/QEMU
MPI/NCCL

Job description

Nebius is seeking a Senior Systems HPC Engineer to optimize large-scale GPU clusters and run across the full stack from hardware to software. You will identify bottlenecks, validate performance improvements, and work with teams across infra, software, and hardware vendors.

Ideal candidates bring 5+ years in system-level software performance, Linux expertise, and strong C/C++/Go/Python skills. The role offers an international, fast-moving environment with opportunities to influence AI-scale

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior HPC Engineer - GPU Compute & InfiniBand
Senior HPC Engineer - GPU Compute & InfiniBand

Nebius • United States

Remote
USD 150,000 - 230,000
Lead GPU Performance Engineer — HPC Systems
Lead GPU Performance Engineer — HPC Systems

Nebius • United States

On-site
USD 170,000 - 300,000
Health insurance
401(k) plan
Parental leave
+2
Senior HPC Systems Engineer: GPU/InfiniBand & KVM
Senior HPC Systems Engineer: GPU/InfiniBand & KVM

Nebius • United States

On-site
USD 170,000 - 300,000
Competitive compensation
Career growth
Flexible and ownership culture
+1
Senior Data Center GPU Infrastructure Engineer
Senior Data Center GPU Infrastructure Engineer

Nebius • New Jersey

On-site
USD 85,000 - 140,000
Career growth opportunities
Flexible work environment
Collaborative culture
+1
Senior HPC-AI Systems Architect (Equity)
Senior HPC-AI Systems Architect (Equity)

Nvidia Corporation in • Santa Clara (CA)

On-site
USD 176,000 - 334,000
Senior HPC Cloud Engineer — GPU/InfiniBand AI Infra
Senior HPC Cloud Engineer — GPU/InfiniBand AI Infra

Jobgether • Germany (OH)

On-site
USD 81,000 - 105,000
Career development opportunities
Flexible working arrangements
Collaborative engineering environment
+1
Senior Systems Software Engineer, GPU Compute
Senior Systems Software Engineer, GPU Compute

Nebius • United States

On-site
USD 170,000 - 300,000
Competitive compensation
Career growth
Flexible and ownership culture
+1
Senior HPC Architect - GPU Compute, Equity Eligible
Senior HPC Architect - GPU Compute, Equity Eligible

NVIDIA • California (MO)

On-site
USD 184,000 - 288,000
Equity
Inclusive work environment
Comprehensive benefits
Lead Software Systems Engineer - GPU Performance
Lead Software Systems Engineer - GPU Performance

Nebius • United States

On-site
USD 170,000 - 300,000
Health insurance
401(k) plan
Parental leave
+2
Senior GPU HPC Cluster Engineer — Equity Eligible
Senior GPU HPC Cluster Engineer — Equity Eligible

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 152,000 - 242,000