GPU Benchmark Engineer for AI Cloud Infra

Nebius

United States

Remote

USD 150,000 - 190,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Nebius is seeking a highly skilled ML/AI Engineer to lead benchmarking of GPU platforms for ML and AI workloads. You will profile GPU performance at system and kernel levels and evaluate across platforms, architectures, and software stacks, including CUDA and ROCm.

You will debug and optimise workloads and run experiments to understand interconnect impacts. Join a team defining performance standards for next‑gen AI hardware, with dashboards and tooling to visualize trends.

Qualifications

  • Strong understanding of ML theory and fundamentals.
  • Deep knowledge of performance aspects of large neural networks training and inference.
  • Hands-on experience with PyTorch, JAX, Megatron-LM, Tensort-LLM.
  • Solid grasp of the GPU stack: CUDA, NCCL, drivers, libraries.
  • Experience with containerized environments (Docker, Kubernetes).
  • Excellent communication and ability to work independently.

Responsibilities

  • Profile and analyze GPU performance at system and kernel levels with hardware teams.
  • Evaluate and compare GPU performance across platforms, architectures and stacks (CUDA, ROCm).
  • Debug and optimize ML workloads to run efficiently on GPU hardware and resolve bottlenecks.
  • Perform acceptance testing for new GPU clusters ensuring performance and compatibility for AI workloads.
  • Experiment with diverse GPU configurations to assess interconnect strategies and system-level optimisations.
  • Develop tools and dashboards to visualize performance metrics, bottlenecks and trends.
  • Contribute to internal tooling, frameworks and best practices.

Skills

ML theory
Large-scale model perf
PyTorch
JAX
Megatron-LM
Tensor-LLM
CUDA
NCCL
Docker
Kubernetes
Python
Performance profiling
Communication
Independent work

Tools

CUDA
NCCL
Nsight
nvprof
Docker
Kubernetes
perf
TensorRT

Job description

Nebius is seeking a highly skilled ML/AI Engineer to lead benchmarking of GPU platforms for ML and AI workloads. You will profile GPU performance at system and kernel levels and evaluate across platforms, architectures, and software stacks, including CUDA and ROCm.

You will debug and optimise workloads and run experiments to understand interconnect impacts. Join a team defining performance standards for next‑gen AI hardware, with dashboards and tooling to visualize trends.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

GPU ML Benchmarking Engineer for Next-Gen AI Infra
GPU ML Benchmarking Engineer for Next-Gen AI Infra

Nebius • Amsterdam (VA)

On-site
USD 130,000 - 190,000
Competitive compensation
Career growth and learning
Flexibility and ownership
+2
GPU Systems Engineer for AI Inference & Performance
GPU Systems Engineer for AI Inference & Performance

Nebius • United States

Remote
USD 180,000 - 260,000
Competitive compensation
Career growth
Flexible ownership & autonomy
+3
Senior GPU Benchmarking & Optimization Engineer
Senior GPU Benchmarking & Optimization Engineer

Webhosting • United States

On-site
USD 140,000 - 150,000
Health insurance
401(k) plan with matching
Professional development reimbursement
+4
ML Infrastructure Engineer
ML Infrastructure Engineer

Nebius • Amsterdam (VA)

On-site
USD 130,000 - 190,000
Competitive compensation
Career growth and learning
Flexibility and ownership
+2
Senior ML Engineer - GPU Inference & Large-Scale AI Cloud
Senior ML Engineer - GPU Inference & Large-Scale AI Cloud

Nebius • United States

Remote
USD 180,000 - 250,000
GPU Inference Performance Engineer — Equity & Optimization
GPU Inference Performance Engineer — Equity & Optimization

Nvidia Corporation • Santa Clara (CA)

On-site
USD 152,000 - 242,000
Senior HPC Systems Engineer: GPU Clusters & AI Infra
Senior HPC Systems Engineer: GPU Clusters & AI Infra

Nebius • United States

Remote
USD 180,000 - 240,000
Competitive pay
Career growth
Flexibility and ownership
+3
GPU Systems Performance Engineer for AI Inference
GPU Systems Performance Engineer for AI Inference

Yoh Services LLC • California (MO)

On-site
USD 250,000 - 300,000
Medical benefits
Dental & Vision
401K Retirement
GPU AI Performance & Benchmark Engineer
GPU AI Performance & Benchmark Engineer

Yoh, A Day & Zimmermann Company • California (MO)

On-site
USD 250,000 - 300,000
Medical+Vision
HSA
Life & Disability
+5
AI Benchmark Architect—Systems & GPUs
AI Benchmark Architect—Systems & GPUs

Sandisk • California (MO)

On-site
USD 120,000 - 160,000
Comprehensive benefits package
Paid vacation and sick leave
Employee Stock Purchase Plan