Senior AI Inference Performance Engineer

NVIDIA AI

Santa Clara (CA)

On-site

USD 140,000 - 230,000

Full time

12 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

NVIDIA is seeking a Senior Inference Performance Engineer to push performance limits on large-scale AI inference benchmarks using modern GPUs. You will optimize autonomous benchmarking and profiling workflows, enabling agents to run experiments, tune configurations, and report credible wins.

You will collaborate across TensorRT-LLM, vLLM, and CUDA teams, exercise rigorous experiments, and shape patches to serving frameworks.

Qualifications

  • Experience optimizing AI model execution for throughput-latency tradeoffs.
  • Hands-on benchmarking with Nsight Systems, Nsight Compute, CUPTI, or PyTorch profiler.
  • Strong Python skills and ability to modify large C++/CUDA codebases.
  • Rigorous experimental methodology with reproducible benchmarks.
  • Experience with quantization and KV cache considerations.
  • Ability to explain tradeoffs clearly in writing and speech.

Responsibilities

  • Distill performance instincts into reusable skills and methodologies for autonomous AI agents.
  • Improve AI inference workloads to increase throughput-per-GPU and user interactivity.
  • Measure and optimize serving architectures across TensorRT-LLM, SGLang, vLLM, and Dynamo.
  • Profile workloads with Nsight tools and analyze results to drive fixes.
  • Land upstream improvements to serving frameworks and kernels.
  • Collaborate with TensorRT-LLM, SGLang, vLLM, kernel, benchmarking, and GPU teams.

Skills

Python
C++/CUDA
Nsight profiling
PyTorch profiler
GPU benchmarking
MoE serving
Quantization
Kernel optimization
Kernel profiling

Education

BS/MS/PhD in CS/CE/EE

Tools

Nsight Systems
Nsight Compute
CUPTI
TensorRT-LLM
vLLM

Job description

NVIDIA is seeking a Senior Inference Performance Engineer to push performance limits on large-scale AI inference benchmarks using modern GPUs. You will optimize autonomous benchmarking and profiling workflows, enabling agents to run experiments, tune configurations, and report credible wins.

You will collaborate across TensorRT-LLM, vLLM, and CUDA teams, exercise rigorous experiments, and shape patches to serving frameworks.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Inference Optimization Engineer
Senior AI Inference Optimization Engineer

Nvidia Corporation in • Santa Clara (CA)

Hybrid
USD 124,000 - 196,000
Equity
Benefits package
Hybrid work model
GPU Inference Performance Engineer — Equity & Optimization
GPU Inference Performance Engineer — Equity & Optimization

Nvidia Corporation • Santa Clara (CA)

On-site
USD 152,000 - 242,000
Senior Inference Performance Engineer — Equity & Hybrid
Senior Inference Performance Engineer — Equity & Hybrid

NVIDIA Gruppe • Santa Clara (CA)

Hybrid
USD 124,000 - 242,000
Senior AI Inference Performance Engineer — Scale GPUs
Senior AI Inference Performance Engineer — Scale GPUs

NVIDIA • California (MO)

On-site
USD 124,000 - 196,000
Equity eligibility
AI Inference Performance Engineer - New College Grad 2026
AI Inference Performance Engineer - New College Grad 2026

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 120,000 - 160,000
Inference Performance Engineer, Agent Driven Inference Optimization
Inference Performance Engineer, Agent Driven Inference Optimization

NVIDIA • California (MO)

On-site
USD 124,000 - 196,000
Equity eligibility
Senior AI Inference Performance Product Manager
Senior AI Inference Performance Product Manager

Nvidia Corporation in • Santa Clara (CA)

On-site
USD 208,000 - 328,000
Inference Performance Engineer, AI Inference Configuration Optimization
Inference Performance Engineer, AI Inference Configuration Optimization

Nvidia Corporation • Santa Clara (CA)

On-site
USD 152,000 - 242,000
Inference Performance Engineer, Agent Driven Inference Optimization
Inference Performance Engineer, Agent Driven Inference Optimization

NVIDIA Gruppe • Santa Clara (CA)

Hybrid
USD 124,000 - 242,000
AI Inference Performance Engineer for NVIDIA GPUs
AI Inference Performance Engineer for NVIDIA GPUs

YOH Services LLC • Santa Clara (CA)

On-site
USD 250,000 - 300,000
Medical benefits
Health Savings Account (HSA)
401K Retirement Savings Plan
+2