GPU Systems Engineer for AI Inference & Performance

Nebius

United States

Remote

USD 180,000 - 260,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive compensation
Career growth
Flexible ownership & autonomy
Collaborative culture
Impactful AI projects
International environment

Job summary

Nebius is building Token Factory within Nebius Cloud, a world-scale GPU cloud for AI inference. The team develops and optimizes low-level kernels and runtime components to support fast, reliable deployment of foundation models across diverse hardware.

The role emphasizes performance tuning, profiling, and collaboration with ML and backend engineers to improve end-to-end execution on GPU platforms, with strong memory hierarchy and system-level expertise required.

Qualifications

  • Strong proficiency in C++ and memory management for high-performance systems.
  • Experience in GPU programming or systems-level software development.
  • Hands-on profiling and debugging to optimize CPU/GPU performance.
  • Knowledge of CPU/GPU architecture and memory hierarchy.

Responsibilities

  • Develop and optimize low-level kernels and runtime components for AI inference.
  • Improve performance of inference engines on GPU platforms.
  • Profile and debug system-level and hardware-level performance issues.
  • Integrate support for new hardware architectures (e.g., Hopper, Blackwell, Rubin).
  • Collaborate with ML and backend teams to optimize end-to-end execution.

Skills

C++ programming
GPU programming
Kernel modules
Performance profiling

Tools

Perf tools
Nsight
VTune
ROCm profiler

Job description

Nebius is building Token Factory within Nebius Cloud, a world-scale GPU cloud for AI inference. The team develops and optimizes low-level kernels and runtime components to support fast, reliable deployment of foundation models across diverse hardware.

The role emphasizes performance tuning, profiling, and collaboration with ML and backend engineers to improve end-to-end execution on GPU platforms, with strong memory hierarchy and system-level expertise required.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior ML Engineer - GPU Inference & Large-Scale AI Cloud
Senior ML Engineer - GPU Inference & Large-Scale AI Cloud

Nebius • United States

Remote
USD 180,000 - 250,000
GPU Benchmark Engineer for AI Cloud Infra
GPU Benchmark Engineer for AI Cloud Infra

Nebius • United States

Remote
USD 150,000 - 190,000
GPU ML Benchmarking Engineer for Next-Gen AI Infra
GPU ML Benchmarking Engineer for Next-Gen AI Infra

Nebius • Amsterdam (VA)

On-site
USD 130,000 - 190,000
Competitive compensation
Career growth and learning
Flexibility and ownership
+2
Senior HPC Systems Engineer: GPU Clusters & AI Infra
Senior HPC Systems Engineer: GPU Clusters & AI Infra

Nebius • United States

Remote
USD 180,000 - 240,000
Competitive pay
Career growth
Flexibility and ownership
+3
GPU Systems Performance Engineer for AI Inference
GPU Systems Performance Engineer for AI Inference

Yoh Services LLC • California (MO)

On-site
USD 250,000 - 300,000
Medical benefits
Dental & Vision
401K Retirement
GPU Inference Performance Engineer — Equity & Optimization
GPU Inference Performance Engineer — Equity & Optimization

Nvidia Corporation • Santa Clara (CA)

On-site
USD 152,000 - 242,000
Senior Systems Software Engineer, GPU Compute
Senior Systems Software Engineer, GPU Compute

Nebius • United States

On-site
USD 170,000 - 300,000
Competitive compensation
Career growth
Flexible and ownership culture
+1
Senior HPC Engineer - GPU Compute & InfiniBand
Senior HPC Engineer - GPU Compute & InfiniBand

Nebius • United States

Remote
USD 150,000 - 230,000
Lead GPU Performance Engineer — HPC Systems
Lead GPU Performance Engineer — HPC Systems

Nebius • United States

On-site
USD 170,000 - 300,000
Health insurance
401(k) plan
Parental leave
+2
Compute Node Systems Engineer - GPU/VM Orchestration
Compute Node Systems Engineer - GPU/VM Orchestration

Nebius • United States

Remote
USD 150,000 - 210,000
Competitive compensation
Career growth and learning
Flexibility and ownership
+2