Senior AI Inference Performance Engineer — Equity Eligible

Nvidia Corporation in

Santa Clara (CA)

On-site

USD 184,000 - 357,000

Full time

2 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity
Benefits

Job summary

NVIDIA is seeking a Senior Software Engineer to advance AI inference performance on GPU-accelerated systems. You will optimize LLM/VLM workloads, profile with Nsight tools, and contribute to open-source inference engines while collaborating with model, kernel, and networking teams.

The role emphasizes end-to-end analysis, batching, KV-cache management, and quantization, with a focus on lower latency, higher throughput, and scalable deployments across NVIDIA platforms.

Qualifications

  • BS or MS in Computer Science, Computer Engineering, or a related field, or equivalent experience.
  • Strong programming skills in Python, C++, and/or Rust; hands-on CUDA or GPU programming experience.
  • Experience with speed-of-light analysis, roofline models, microbenchmarks, profiling tools.

Responsibilities

  • Lead end-to-end analysis of LLM/VLM inference processes and optimize latency, throughput, and KV cache usage.
  • Build speed-of-light and roofline models to quantify performance headroom across GPUs.
  • Profile workloads using Nsight Systems/Compute, PyTorch Profiler, and custom instrumentation to remove bottlenecks.
  • Tune serving hyperparameters and techniques (batching, caching, quantization, speculative decoding, CUDA Graphs).
  • Develop performance-critical kernels in CUDA, CUTLASS, Triton, or related technologies.
  • Establish benchmarks, regressions gates, and collaborate across teams to upgrade inference software (TensorRT-LLM, vLLM, SGLang).

Skills

Python
C++
Rust
CUDA
GPU programming

Education

BS or MS in CS/CE or related field

Tools

Nsight Systems
Nsight Compute

Job description

NVIDIA is seeking a Senior Software Engineer to advance AI inference performance on GPU-accelerated systems. You will optimize LLM/VLM workloads, profile with Nsight tools, and contribute to open-source inference engines while collaborating with model, kernel, and networking teams.

The role emphasizes end-to-end analysis, batching, KV-cache management, and quantization, with a focus on lower latency, higher throughput, and scalable deployments across NVIDIA platforms.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Inference Performance Engineer
Senior AI Inference Performance Engineer

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Benefits package
Equity
Senior AI Inference Performance Engineer
Senior AI Inference Performance Engineer

NVIDIA AI • Santa Clara (CA)

On-site
USD 140,000 - 230,000
GPU Inference Performance Engineer — Equity & Optimization
GPU Inference Performance Engineer — Equity & Optimization

Nvidia Corporation • Santa Clara (CA)

On-site
USD 152,000 - 242,000
Senior AI Inference Performance Engineer — Scale GPUs
Senior AI Inference Performance Engineer — Scale GPUs

NVIDIA • California (MO)

On-site
USD 124,000 - 196,000
Equity eligibility
Senior Software Engineer, LLM Inference & Performance
Senior Software Engineer, LLM Inference & Performance

NVIDIA AI • Town of Santa Clara (NY)

On-site
USD 150,000 - 230,000
Equity
Health Insurance
Senior DL Inference Engineer — GPU-Accelerated AI, Equity
Senior DL Inference Engineer — GPU-Accelerated AI, Equity

NVIDIA Gruppe • California (MO)

On-site
USD 152,000 - 288,000
Senior AI Systems Engineer: GPU Kernels & Inference Equity
Senior AI Systems Engineer: GPU Kernels & Inference Equity

NVIDIA AI • Michigan

On-site
USD 150,000 - 190,000
Equity
Health Insurance
Senior Software Engineer - AI Inference Performance
Senior Software Engineer - AI Inference Performance

Nvidia Corporation in • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior Inference Engineer: GPU Kernel Optimizations + Equity
Senior Inference Engineer: GPU Kernel Optimizations + Equity

Nvidia Corporation • Santa Clara (CA)

On-site
USD 184,000 - 288,000
Equity
Comprehensive benefits
Senior AI Inference Engineer: GPU Kernels & LLM Runtimes
Senior AI Inference Engineer: GPU Kernels & LLM Runtimes

NVIDIA • Redmond (WA)

On-site
USD 184,000 - 288,000
Equity
Benefits