Senior GPU Performance & Benchmark Engineer – AI Inference

Yoh,-A-Day-

Santa Clara (CA)

On-site

USD 250,000 - 300,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Medical, Prescription, Dental & Vision
401K Retirement Savings Plan
Direct Deposit & weekly epayroll
Certification and training

Job summary

Yoh is seeking a hands-on Performance and Benchmark Engineer to characterize AI workloads on NVIDIA GPU systems in Santa Clara. You will scope, run, and optimize benchmarks across DGX and related platforms, focusing on latency, throughput, and scaling.

Collaborate with architecture, compute, and networking teams to drive end-to-end performance improvements while delivering repeatable results and clear reporting.

Qualifications

  • Deep hands-on experience with NVIDIA GPU compute platforms and AI/ML performance benchmarking.
  • Strong understanding of AI inference, model performance, workload characterization, and GPU architecture.
  • Experience with NVIDIA DGX, B200/B300, H100/H200, Blackwell, Hopper, or comparable GPU systems.
  • Experience analyzing performance metrics including latency, throughput, GPU utilization, memory bandwidth, and multi-GPU scaling.
  • Strong scripting and automation skills using Python or similar languages.

Responsibilities

  • Develop and execute performance benchmarks for AI inference and machine learning workloads across NVIDIA GPU systems.
  • Characterize performance on platforms including NVIDIA DGX and B200/B300-based systems, analyzing throughput, latency, utilization, memory behavior, and scaling efficiency.
  • Evaluate AI models and workload configurations to identify performance bottlenecks and recommend system or architecture improvements.
  • Build benchmarking methodologies, automation, and reporting frameworks to produce repeatable performance results.
  • Collaborate with architecture, compute, networking, and software teams to optimize end-to-end AI cluster performance.

Skills

NVIDIA GPU compute platforms
AI/ML performance benchmarking
Python scripting
Performance metrics

Tools

NVIDIA DGX
B200/B300 systems
H100/H200
Nsight

Job description

Yoh is seeking a hands-on Performance and Benchmark Engineer to characterize AI workloads on NVIDIA GPU systems in Santa Clara. You will scope, run, and optimize benchmarks across DGX and related platforms, focusing on latency, throughput, and scaling.

Collaborate with architecture, compute, and networking teams to drive end-to-end performance improvements while delivering repeatable results and clear reporting.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

GPU Systems Performance Engineer for AI Inference
GPU Systems Performance Engineer for AI Inference

Yoh Services LLC • California (MO)

On-site
USD 250,000 - 300,000
Medical benefits
Dental & Vision
401K Retirement
GPU AI Performance & Benchmark Engineer
GPU AI Performance & Benchmark Engineer

Yoh, A Day & Zimmermann Company • California (MO)

On-site
USD 250,000 - 300,000
Medical+Vision
HSA
Life & Disability
+5
AI Inference Performance Engineer for NVIDIA GPUs
AI Inference Performance Engineer for NVIDIA GPUs

YOH Services LLC • Santa Clara (CA)

On-site
USD 250,000 - 300,000
Medical benefits
Health Savings Account (HSA)
401K Retirement Savings Plan
+2
GPU Inference Performance Engineer — Equity & Optimization
GPU Inference Performance Engineer — Equity & Optimization

Nvidia Corporation • Santa Clara (CA)

On-site
USD 152,000 - 242,000
Performance/ Benchmark Engineer - NVIDIA GPU Systems
Performance/ Benchmark Engineer - NVIDIA GPU Systems

Yoh, A Day & Zimmermann Company • California (MO)

On-site
USD 250,000 - 300,000
Medical+Vision
HSA
Life & Disability
+5
Senior Inference Performance Engineer — Equity & Hybrid
Senior Inference Performance Engineer — Equity & Hybrid

NVIDIA Gruppe • Santa Clara (CA)

Hybrid
USD 124,000 - 242,000
Senior Performance Engineer: AI Workload Optimization
Senior Performance Engineer: AI Workload Optimization

NVIDIA • Redmond (WA)

On-site
USD 224,000 - 432,000
Equity
Benefits
Inference Performance Engineer: AI GPU Optimization&Equity
Inference Performance Engineer: AI GPU Optimization&Equity

NVIDIA • Santa Clara (CA)

Hybrid
USD 124,000 - 242,000
Equity
Benefits package
Senior AI Systems Performance Engineer
Senior AI Systems Performance Engineer

NVIDIA • Austin (TX)

On-site
USD 272,000 - 432,000
Equity
Benefits
Senior GPU Inference Performance Architect
Senior GPU Inference Performance Architect

Advanced Micro Devices • Santa Clara (CA)

On-site
USD 180,000 - 280,000