Senior AI Systems Engineer — GPU/HPC Troubleshooter

Nebius

United States

On-site

USD 180,000 - 224,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Health insurance
401(k) match
Parental leave
Remote work reimbursement
Disability & life insurance

Job summary

Nebius is seeking a senior systems engineer to diagnose and resolve complex hardware and platform failures in high-performance AI infrastructure. You will work across Linux, GPU servers, PCIe, firmware, power and cooling to determine root causes and build repeatable paths to resolution.

With 5+ years in systems engineering, you will leverage deep hardware and software knowledge, performance analysis, and debugging in cloud/HPC environments to deliver reliable, scalable AI infrastructure.

Qualifications

  • Minimum 5 years of hands-on systems engineering across Linux, server hardware, firmware, PCIe, and GPU/HPC platforms.
  • Extensive Linux experience for hardware debugging, system-level troubleshooting and platform investigation.
  • Strong knowledge of NVIDIA GPU platforms, diagnostic tooling, and related firmware interactions.

Responsibilities

  • Form technical hypotheses, design targeted tests and analyze system logs to identify root causes.
  • Diagnose hardware and platform failures across Linux, GPU servers, PCIe, and firmware.
  • Develop repeatable troubleshooting procedures and evidence-based conclusions for complex issues.

Skills

Linux
Server hardware
Firmware
PCIe
NVIDIA GPU platforms
NVLink/NVSwitch
Kernel debugging
Troubleshooting complex systems
Scripting (Python/Go)
Benchmarking

Tools

nvidia-smi
DCGM
NCCL testing
GSP/driver behavior analysis

Job description

Nebius is seeking a senior systems engineer to diagnose and resolve complex hardware and platform failures in high-performance AI infrastructure. You will work across Linux, GPU servers, PCIe, firmware, power and cooling to determine root causes and build repeatable paths to resolution.

With 5+ years in systems engineering, you will leverage deep hardware and software knowledge, performance analysis, and debugging in cloud/HPC environments to deliver reliable, scalable AI infrastructure.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior HPC Systems Engineer: GPU Clusters & AI Infra
Senior HPC Systems Engineer: GPU Clusters & AI Infra

Nebius • United States

Remote
USD 180,000 - 240,000
Competitive pay
Career growth
Flexibility and ownership
+3
Senior HPC Engineer - GPU Compute & InfiniBand
Senior HPC Engineer - GPU Compute & InfiniBand

Nebius • United States

Remote
USD 150,000 - 230,000
Senior HPC Systems Engineer: GPU/InfiniBand & KVM
Senior HPC Systems Engineer: GPU/InfiniBand & KVM

Nebius • United States

On-site
USD 170,000 - 300,000
Competitive compensation
Career growth
Flexible and ownership culture
+1
Senior Technical Program Manager - AI Infra & Data Centers
Senior Technical Program Manager - AI Infra & Data Centers

nebius • United States

Remote
USD 115,000 - 275,000
Healthcare coverage
401(k) with company match
Parental leave (20 weeks primary, 12–?
+3
Senior Data Center Network Architect for AI Cloud
Senior Data Center Network Architect for AI Cloud

Nebius • United States

On-site
USD 125,000 - 180,000
Health insurance
401(k) plan with company contribution
Paid time off
GPU Benchmark Engineer for AI Cloud Infra
GPU Benchmark Engineer for AI Cloud Infra

Nebius • United States

Remote
USD 150,000 - 190,000
Senior Compute Solutions Engineer - GPU, AI & Linux
Senior Compute Solutions Engineer - GPU, AI & Linux

NVIDIA Corporation • Santa Clara (CA)

Hybrid
USD 140,000 - 270,000
Equity
Comprehensive benefits
Senior GPU Solutions Engineer — AI, HPC & Multi-GPU
Senior GPU Solutions Engineer — AI, HPC & Multi-GPU

NVIDIA • Westford (MA)

On-site
USD 140,000 - 270,000
Equity
Benefits
System Engineer
System Engineer

Acceler8 Talent • Fremont (CA), Northern (KY)

Hybrid
USD 135,000 - 165,000
Senior Systems Software Engineer, GPU Compute
Senior Systems Software Engineer, GPU Compute

Nebius • United States

On-site
USD 170,000 - 300,000
Competitive compensation
Career growth
Flexible and ownership culture
+1