Inference Performance Engineer: AI GPU Optimization&Equity

NVIDIA

Santa Clara (CA)

Hybrid

USD 124,000 - 242,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity
Benefits package

Job summary

NVIDIA is seeking a Senior Inference Performance Engineer to push performance limits on large-scale AI inference benchmarks. You will optimize AI model execution and profiling workflows, enabling autonomous optimization across TensorRT-LLM, vLLM, and disaggregated serving architectures.

The role emphasizes measurable gains, reproducible experiments, and close collaboration with GPU and software teams. Responsibilities include distilling performance methods, benchmarking with Nsight tools, and

Qualifications

  • Must have extensive knowledge of AI model execution efficiency and optimization.
  • Hands-on benchmarking and profiling GPU workloads with Nsight Systems, Nsight Compute, CUPTI, or PyTorch profiler.
  • Strong Python and the ability to navigate/modify large C++/CUDA codebases.
  • Rigorous experimental methodology with controlled single-variable comparisons and reproducible benchmarks.
  • Excellent written and verbal communication of performance tradeoffs.

Responsibilities

  • Distill performance instincts into reusable skills, workflows, and evidence-backed methodologies for autonomous AI agents.
  • Improve throughput-per-GPU and interactivity via configuration options, batching, and memory management.
  • Profile workloads using Nsight tools and analyze with roofline/latency metrics to drive fixes.
  • Land upstream improvements: serving framework patches, optimized kernels, and deployment recipes.
  • Collaborate with TensorRT-LLM, SGLang, vLLM, kernel, benchmarking, and GPU teams to convert profiling insights into gains.
  • Communicate tradeoffs clearly to humans and documentation for autonomous systems.

Skills

Python
C++/CUDA
Performance optimization
Experimental methodology

Education

BS/MS/PhD in CS/CE/EE or related field

Tools

Nsight Systems
Nsight Compute
CUPTI
PyTorch profiler
TensorRT-LLM

Job description

NVIDIA is seeking a Senior Inference Performance Engineer to push performance limits on large-scale AI inference benchmarks. You will optimize AI model execution and profiling workflows, enabling autonomous optimization across TensorRT-LLM, vLLM, and disaggregated serving architectures.

The role emphasizes measurable gains, reproducible experiments, and close collaboration with GPU and software teams. Responsibilities include distilling performance methods, benchmarking with Nsight tools, and

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

GPU Inference Performance Engineer — Equity & Optimization
GPU Inference Performance Engineer — Equity & Optimization

Nvidia Corporation • Santa Clara (CA)

On-site
USD 152,000 - 242,000
Senior AI Inference Optimization Engineer
Senior AI Inference Optimization Engineer

Nvidia Corporation in • Santa Clara (CA)

Hybrid
USD 124,000 - 196,000
Equity
Benefits package
Hybrid work model
Senior Inference Performance Engineer — Equity & Hybrid
Senior Inference Performance Engineer — Equity & Hybrid

NVIDIA Gruppe • Santa Clara (CA)

Hybrid
USD 124,000 - 242,000
Senior AI Inference Performance Engineer — Scale GPUs
Senior AI Inference Performance Engineer — Scale GPUs

NVIDIA • California (MO)

On-site
USD 124,000 - 196,000
Equity eligibility
Inference Performance Engineer, Agent Driven Inference Optimization
Inference Performance Engineer, Agent Driven Inference Optimization

NVIDIA • California (MO)

On-site
USD 124,000 - 196,000
Equity eligibility
Inference Performance Engineer, AI Inference Configuration Optimization
Inference Performance Engineer, AI Inference Configuration Optimization

Nvidia Corporation • Santa Clara (CA)

On-site
USD 152,000 - 242,000
Inference Performance Engineer, AI Inference Configuration Optimization
Inference Performance Engineer, AI Inference Configuration Optimization

Nvidia Corporation in • Santa Clara (CA)

Hybrid
USD 124,000 - 196,000
Equity
Benefits package
Hybrid work model
AI Inference Performance Engineer - New College Grad 2026
AI Inference Performance Engineer - New College Grad 2026

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 120,000 - 160,000
AI Inference Performance Engineer for NVIDIA GPUs
AI Inference Performance Engineer for NVIDIA GPUs

YOH Services LLC • Santa Clara (CA)

On-site
USD 250,000 - 300,000
Medical benefits
Health Savings Account (HSA)
401K Retirement Savings Plan
+2
Inference Performance Engineer, AI Inference Configuration Optimization
Inference Performance Engineer, AI Inference Configuration Optimization

NVIDIA • Santa Clara (CA)

Hybrid
USD 124,000 - 242,000
Equity
Benefits package