GPU Systems Performance Engineer for AI Inference

Yoh Services LLC

California (MO)

On-site

USD 250,000 - 300,000

Full time

10 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Medical benefits
Dental & Vision
401K Retirement

Job summary

Yoh Services LLC is seeking a hands-on Performance/ Benchmark Engineer to characterize and optimize AI workloads on large-scale NVIDIA GPU infrastructure, including DGX and B200/B300 systems. You will focus on GPU compute performance, AI inference, and system benchmarking while collaborating across architecture, compute, and networking teams.

The role requires deep hands-on GPU benchmarking experience, strong Python scripting, and proficiency with CUDA/NCCL.

Qualifications

  • Deep hands-on experience with NVIDIA GPU compute platforms and AI/ML performance benchmarking.
  • Strong understanding of AI inference, model performance, workload characterization, and GPU architecture.
  • Experience with NVIDIA DGX, B200/B300, H100/H200, Blackwell, Hopper, or comparable GPU systems.
  • Experience analyzing performance metrics including latency, throughput, GPU utilization, memory bandwidth, and multi-GPU scaling.
  • Strong scripting and automation skills using Python or similar languages.

Responsibilities

  • Develop and execute performance benchmarks for AI inference and machine learning workloads across NVIDIA GPU systems.
  • Characterize performance on platforms including NVIDIA DGX and B200/B300-based systems, analyzing throughput, latency, utilization, memory behavior, and scaling efficiency.
  • Evaluate AI models and workload configurations to identify performance bottlenecks and recommend system or architecture improvements.
  • Build benchmarking methodologies, automation, and reporting frameworks to produce repeatable performance results.
  • Collaborate with architecture, compute, networking, and software teams to optimize end-to-end AI cluster performance.

Skills

NVIDIA GPU
AI benchmarking
Python scripting
Performance analysis

Education

BS in CS/EE

Tools

CUDA
NCCL
TensorRT
PyTorch
Nsight

Job description

Yoh Services LLC is seeking a hands-on Performance/ Benchmark Engineer to characterize and optimize AI workloads on large-scale NVIDIA GPU infrastructure, including DGX and B200/B300 systems. You will focus on GPU compute performance, AI inference, and system benchmarking while collaborating across architecture, compute, and networking teams.

The role requires deep hands-on GPU benchmarking experience, strong Python scripting, and proficiency with CUDA/NCCL.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior GPU Performance & Benchmark Engineer – AI Inference
Senior GPU Performance & Benchmark Engineer – AI Inference

Yoh,-A-Day- • Santa Clara (CA)

On-site
USD 250,000 - 300,000
Medical, Prescription, Dental & Vision
401K Retirement Savings Plan
Direct Deposit & weekly epayroll
+1
AI Inference Performance Engineer for NVIDIA GPUs
AI Inference Performance Engineer for NVIDIA GPUs

YOH Services LLC • Santa Clara (CA)

On-site
USD 250,000 - 300,000
Medical benefits
Health Savings Account (HSA)
401K Retirement Savings Plan
+2
GPU AI Performance & Benchmark Engineer
GPU AI Performance & Benchmark Engineer

Yoh, A Day & Zimmermann Company • California (MO)

On-site
USD 250,000 - 300,000
Medical+Vision
HSA
Life & Disability
+5
Performance/ Benchmark Engineer - NVIDIA GPU Systems
Performance/ Benchmark Engineer - NVIDIA GPU Systems

Yoh, A Day & Zimmermann Company • California (MO)

On-site
USD 250,000 - 300,000
Medical+Vision
HSA
Life & Disability
+5
GPU Inference Performance Engineer — Equity & Optimization
GPU Inference Performance Engineer — Equity & Optimization

Nvidia Corporation • Santa Clara (CA)

On-site
USD 152,000 - 242,000
Inference Performance Engineer: AI GPU Optimization&Equity
Inference Performance Engineer: AI GPU Optimization&Equity

NVIDIA • Santa Clara (CA)

Hybrid
USD 124,000 - 242,000
Equity
Benefits package
Performance/ Benchmark Engineer - NVIDIA GPU Systems
Performance/ Benchmark Engineer - NVIDIA GPU Systems

YOH Services LLC • Santa Clara (CA)

On-site
USD 250,000 - 300,000
Medical benefits
Health Savings Account (HSA)
401K Retirement Savings Plan
+2
Senior Inference Performance Engineer — Equity & Hybrid
Senior Inference Performance Engineer — Equity & Hybrid

NVIDIA Gruppe • Santa Clara (CA)

Hybrid
USD 124,000 - 242,000
Senior AI Inference Optimization Engineer
Senior AI Inference Optimization Engineer

Nvidia Corporation in • Santa Clara (CA)

Hybrid
USD 124,000 - 196,000
Equity
Benefits package
Hybrid work model
Senior AI Inference Performance Engineer — Scale GPUs
Senior AI Inference Performance Engineer — Scale GPUs

NVIDIA • California (MO)

On-site
USD 124,000 - 196,000
Equity eligibility