Senior AI Inference Performance Engineer (CUDA/LLM/VLM)

NVIDIA AI

Santa Clara (CA)

On-site

USD 180,000 - 300,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Equity
Generous Benefits Package

Job summary

Lead the end-to-end performance analysis and optimization of LLM and VLM inference on NVIDIA GPU-accelerated systems. Develop performance-critical kernels and contribute improvements to open-source inference engines to reduce latency and increase throughput.

Requires over 6 years of experience in full-stack AI inference performance with strong programming skills in Python, C++, or Rust and expertise in CUDA.

Qualifications

  • Requires 6+ years in full-stack AI inference performance.
  • Strong programming skills in Python, C++, or Rust and expert CUDA.
  • Degree in Computer Science or Computer Engineering required.
  • Deep knowledge of GPU architecture and profiling tools.

Responsibilities

  • Lead end-to-end performance analysis and optimization of LLM and VLM inference on NVIDIA GPU-accelerated systems.
  • Develop performance-critical kernels and contribute improvements to open-source inference engines to reduce latency and increase throughput.

Skills

Full-stack AI
Python
C++
Rust
CUDA
GPU Architecture
Profiling Tools
Nsight Systems
Nsight Compute
PyTorch Profiler
TensorRT-LLM
vLLM
Triton
CUTLASS
Distributed Systems
Quantization

Education

Bachelor's degree in Computer Science or Computer Engineering

Tools

Nsight Systems
Nsight Compute
PyTorch Profiler
TensorRT-LLM
vLLM
Triton
CUTLASS

Job description

Lead the end-to-end performance analysis and optimization of LLM and VLM inference on NVIDIA GPU-accelerated systems. Develop performance-critical kernels and contribute improvements to open-source inference engines to reduce latency and increase throughput.

Requires over 6 years of experience in full-stack AI inference performance with strong programming skills in Python, C++, or Rust and expertise in CUDA.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Software Engineer - AI Inference Performance
Senior Software Engineer - AI Inference Performance

NVIDIA AI • Santa Clara (CA)

On-site
USD 180,000 - 300,000
Equity
Generous Benefits Package
Senior AI Inference Performance Engineer
Senior AI Inference Performance Engineer

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Senior Software Engineer, LLM Inference & Performance
Senior Software Engineer, LLM Inference & Performance

NVIDIA AI • Town of Santa Clara (NY)

On-site
USD 150,000 - 230,000
Equity
Health Insurance
Senior AI Inference Performance Engineer — Equity Eligible
Senior AI Inference Performance Engineer — Equity Eligible

Nvidia Corporation in • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Benefits
LLM Inference Performance Engineer
LLM Inference Performance Engineer

Intel • Santa Clara (CA)

Hybrid
USD 171,000 - 315,000
Senior Software Engineer - AI Inference
Senior Software Engineer - AI Inference

NVIDIA AI • Town of Santa Clara (NY)

On-site
USD 150,000 - 230,000
Equity
Health Insurance
Senior DL Inference Engineer - GPU/LLM Performance & Equity
Senior DL Inference Engineer - GPU/LLM Performance & Equity

NVIDIA • Washington

On-site
USD 184,000 - 288,000
Equity
Benefits
Senior Software Engineer - AI Inference Performance
Senior Software Engineer - AI Inference Performance

Nvidia Corporation in • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Benefits
LLM Inference Performance Engineer - GPU Kernel Optimizer
LLM Inference Performance Engineer - GPU Kernel Optimizer

Intel • Folsom (CA)

Hybrid
USD 171,000 - 315,000
Stock bonuses
Health benefits
Hybrid work model
Senior AI Inference Engineer: GPU Kernels & LLM Runtimes
Senior AI Inference Engineer: GPU Kernels & LLM Runtimes

NVIDIA • Redmond (WA)

On-site
USD 184,000 - 288,000
Equity
Benefits