Senior AI Inference Performance Engineer

NVIDIA

Santa Clara (CA)

On-site

USD 184,000 - 357,000

Full time

13 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

NVIDIA seeks a Senior Software Engineer – AI Inference Performance to push LLM/VLM workloads toward practical performance limits on NVIDIA GPUs. You will lead end-to-end analysis, build performance models, and optimize latency, throughput, and energy efficiency across models, serving software, and distributed runtimes.

The role spans profiling, tuning, and developing high-performance kernels with CUDA, Triton, and CUTLASS while collaborating with model, kernel, networking, and GPU architecture

Qualifications

  • 6+ years of experience in full-stack LLM/VLM inference performance.
  • Strong programming skills in Python, Rust and/or C++, with CUDA experience.
  • Expertise in speed-of-light analysis, roofline models, microbenchmarks, and profiling tools.

Responsibilities

  • Lead end-to-end analysis of LLM/VLM inference processes and optimize latency, throughput, and KV-cache.
  • Build performance models and translate profiling data into actionable optimizations.
  • Profile workloads with Nsight Systems, Nsight Compute, and PyTorch Profiler; remove bottlenecks in host, CUDA kernels, memory, and scheduling.
  • Tune serving parameters (batching, KV-cache management, quantization, speculative decoding, CUDA Graphs, model parallelism).
  • Develop performance-critical kernels (attention, matrix multiply, mixture-of-experts routing) using CUDA/CUTLASS/Triton.
  • Establish benchmarks, run records, and regression gates; collaborate across teams to upgrade TensorRT-LLM, vLLM, SGLang, etc.

Skills

Python
C++
Rust
CUDA

Education

BS or MS in CS/CE or related field

Tools

NVIDIA Nsight Systems
Nsight Compute
PyTorch Profiler
CUDA

Job description

NVIDIA seeks a Senior Software Engineer – AI Inference Performance to push LLM/VLM workloads toward practical performance limits on NVIDIA GPUs. You will lead end-to-end analysis, build performance models, and optimize latency, throughput, and energy efficiency across models, serving software, and distributed runtimes.

The role spans profiling, tuning, and developing high-performance kernels with CUDA, Triton, and CUTLASS while collaborating with model, kernel, networking, and GPU architecture

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Inference Performance Engineer — Equity Eligible
Senior AI Inference Performance Engineer — Equity Eligible

Nvidia Corporation in • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior AI Inference Performance Engineer (CUDA/LLM/VLM)
Senior AI Inference Performance Engineer (CUDA/LLM/VLM)

NVIDIA AI • Santa Clara (CA)

On-site
USD 180,000 - 300,000
Equity
Generous Benefits Package
Senior Software Engineer, LLM Inference & Performance
Senior Software Engineer, LLM Inference & Performance

NVIDIA AI • Town of Santa Clara (NY)

On-site
USD 150,000 - 230,000
Equity
Health Insurance
Senior Software Engineer - AI Inference Performance
Senior Software Engineer - AI Inference Performance

NVIDIA AI • Santa Clara (CA)

On-site
USD 180,000 - 300,000
Equity
Generous Benefits Package
Senior DL Inference Engineer - GPU/LLM Performance & Equity
Senior DL Inference Engineer - GPU/LLM Performance & Equity

NVIDIA • Washington

On-site
USD 184,000 - 288,000
Equity
Benefits
Senior Software Engineer - AI Inference Performance
Senior Software Engineer - AI Inference Performance

Nvidia Corporation in • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior AI Inference Engineer: GPU Kernels & LLM Runtimes
Senior AI Inference Engineer: GPU Kernels & LLM Runtimes

NVIDIA • Redmond (WA)

On-site
USD 184,000 - 288,000
Equity
Benefits
Senior AI Inference Platform Product Lead
Senior AI Inference Platform Product Lead

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 208,000 - 328,000
Equity
Comprehensive benefits
Inclusive culture
Senior DL Inference Engineer — GPU-Optimized LLMs (Remote)
Senior DL Inference Engineer — GPU-Optimized LLMs (Remote)

NVIDIA Corporation • Northern (KY)

Hybrid
USD 152,000 - 288,000
Senior AI Inference Systems Engineer | GPU Kernels & Runtime
Senior AI Inference Systems Engineer | GPU Kernels & Runtime

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 288,000