Senior AI Inference Performance Engineer

NVIDIA Corporation

Santa Clara (CA)

On-site

USD 184,000 - 357,000

Full time

3 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Benefits package
Equity

Job summary

NVIDIA is seeking a Senior Software Engineer – AI Inference Performance to advance LLM and VLM inference on GPU-accelerated systems. You will lead end-to-end analysis, define workloads, and push latency and throughput improvements across models, serving software, and distributed runtimes.

The role requires hands-on coding in Python/C++/Rust, deep GPU architecture knowledge, and experience profiling with Nsight tools.

Qualifications

  • 6+ years in full-stack LLM/VLM inference performance.
  • Strong Python, Rust and/or C++ with GPU profiling experience.
  • Experience with Nsight tools and profiling workflows.
  • Deep understanding of GPU architecture and memory hierarchy.
  • Experience optimizing inference servers and deployment.

Responsibilities

  • Lead end-to-end analysis of LLM/VLM inference processes.
  • Define representative prefill and decode workloads.
  • Optimize latency, throughput, and KV-cache usage.
  • Profile workloads using Nsight Systems, Nsight Compute, PyTorch Profiler, and custom tools.
  • Eliminate bottlenecks in host code, CUDA kernels, memory, and scheduling.
  • Tune batching, KV-cache management, quantization, and model parallelism.

Skills

Python
Rust
C++
Performance analysis

Education

BS/MS in CS/CE

Tools

CUDA
Nsight Systems
Nsight Compute
PyTorch Profiler
Triton
CUTLASS
TensorRT-LLM
vLLM
NCCL
CUDA Graphs

Job description

NVIDIA is seeking a Senior Software Engineer – AI Inference Performance to advance LLM and VLM inference on GPU-accelerated systems. You will lead end-to-end analysis, define workloads, and push latency and throughput improvements across models, serving software, and distributed runtimes.

The role requires hands-on coding in Python/C++/Rust, deep GPU architecture knowledge, and experience profiling with Nsight tools.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Inference Performance Engineer — Equity Eligible
Senior AI Inference Performance Engineer — Equity Eligible

Nvidia Corporation in • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior AI Inference Performance Engineer
Senior AI Inference Performance Engineer

NVIDIA AI • Santa Clara (CA)

On-site
USD 140,000 - 230,000
Senior Software Engineer, LLM Inference & Performance
Senior Software Engineer, LLM Inference & Performance

NVIDIA AI • Town of Santa Clara (NY)

On-site
USD 150,000 - 230,000
Equity
Health Insurance
GPU Inference Performance Engineer — Equity & Optimization
GPU Inference Performance Engineer — Equity & Optimization

Nvidia Corporation • Santa Clara (CA)

On-site
USD 152,000 - 242,000
Senior AI Inference Performance Engineer — Scale GPUs
Senior AI Inference Performance Engineer — Scale GPUs

NVIDIA • California (MO)

On-site
USD 124,000 - 196,000
Equity eligibility
Senior Software Engineer - AI Inference Performance
Senior Software Engineer - AI Inference Performance

Nvidia Corporation in • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior AI Inference Engineer: GPU Kernels & LLM Runtimes
Senior AI Inference Engineer: GPU Kernels & LLM Runtimes

NVIDIA • Redmond (WA)

On-site
USD 184,000 - 288,000
Equity
Benefits
Senior DL Inference Engineer — GPU-Accelerated AI, Equity
Senior DL Inference Engineer — GPU-Accelerated AI, Equity

NVIDIA Gruppe • California (MO)

On-site
USD 152,000 - 288,000
Senior AI Inference Platform Product Lead
Senior AI Inference Platform Product Lead

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 208,000 - 328,000
Equity
Comprehensive benefits
Inclusive culture
Senior Software Engineer - AI Inference
Senior Software Engineer - AI Inference

NVIDIA AI • Town of Santa Clara (NY)

On-site
USD 150,000 - 230,000
Equity
Health Insurance