Senior Inference Engineer: AI-Driven GPU Kernel Optimization

NVIDIA Corporation

Santa Clara (CA)

On-site

USD 184,000 - 288,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

NVIDIA Corporation in Santa Clara, CA, seeks a Sr. Inference Engineer to accelerate LLM inference through GPU kernel optimization. You will lead kernel benchmarking, model-level performance analysis, and AI-driven optimization workflows across silicon and software stacks.

The role requires strong Python and C++, hands-on GPU profiling (CUPTI/NSYS/NCU), and experience with TRT-LLM, SGLang, or vLLM. Collaboration with compiler, hardware, kernel, and framework teams is essential to deliver

Qualifications

  • Master's or PhD in CS/CE or equivalent experience.
  • 6+ years of relevant industry experience.
  • Experience building or directing agentic AI systems—code generation, automated optimization, or multi-step reasoning workflows.
  • Strong Python and C++ skills with proven software engineering fundamentals.
  • Hands-on GPU profiling with CUPTI, NSYS, and NCU.
  • Direct experience with LLM inference frameworks such as TRT-LLM, SGLang, or vLLM.
  • Working knowledge of GPU kernel optimization — CUDA, CUTLASS, Triton, or equivalent.

Responsibilities

  • Drive GPU kernel microbenchmarking across configurations with real-silicon fidelity.
  • Perform end-to-end model performance analysis for production inference deployments.
  • Develop agentic kernel optimization using AI-driven analysis and silicon-verified validation.
  • Collaborate with compiler, hardware, kernel, and framework teams to deliver production-grade gains.

Skills

Python
C++
Kernel optimization
Agentic AI systems
Multi-step reasoning workflows

Education

Master's or PhD in Computer Science or Computer Engineering

Tools

CUPTI
NSYS
NCU
TRT-LLM
SGLang
vLLM
CUDA
CUTLASS
Triton

Job description

NVIDIA Corporation in Santa Clara, CA, seeks a Sr. Inference Engineer to accelerate LLM inference through GPU kernel optimization. You will lead kernel benchmarking, model-level performance analysis, and AI-driven optimization workflows across silicon and software stacks.

The role requires strong Python and C++, hands-on GPU profiling (CUPTI/NSYS/NCU), and experience with TRT-LLM, SGLang, or vLLM. Collaboration with compiler, hardware, kernel, and framework teams is essential to deliver

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior LLM Inference: GPU Kernel Optimization
Senior LLM Inference: GPU Kernel Optimization

NVIDIA • Austin (TX)

On-site
USD 184,000 - 287,500
Equity
Benefits
Senior GPU Kernel Optimizer for LLM Inference
Senior GPU Kernel Optimizer for LLM Inference

NVIDIA • Seattle (WA)

On-site
USD 184,000 - 287,500
Senior Inference Engineer, GPU Kernel Optimization
Senior Inference Engineer, GPU Kernel Optimization

NVIDIA • Austin (TX)

On-site
USD 184,000 - 287,500
Equity
Benefits
Senior Inference Engineer, GPU Kernel Optimization
Senior Inference Engineer, GPU Kernel Optimization

NVIDIA • Seattle (WA)

On-site
USD 184,000 - 287,500
Senior Inference Engineer, GPU Kernel Optimization
Senior Inference Engineer, GPU Kernel Optimization

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 184,000 - 288,000
Senior AI Inference Performance Engineer
Senior AI Inference Performance Engineer

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Senior DL Inference Engineer - GPU/LLM Performance & Equity
Senior DL Inference Engineer - GPU/LLM Performance & Equity

NVIDIA • Washington

On-site
USD 184,000 - 288,000
Equity
Benefits
Senior DL Inference Engineer — GPU-Optimized LLMs (Remote)
Senior DL Inference Engineer — GPU-Optimized LLMs (Remote)

NVIDIA Corporation • Northern (KY)

Hybrid
USD 152,000 - 288,000
Senior DL Inference Engineer - GPU & LLM Performance
Senior DL Inference Engineer - GPU & LLM Performance

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior AI Inference Performance Engineer (CUDA/LLM/VLM)
Senior AI Inference Performance Engineer (CUDA/LLM/VLM)

NVIDIA AI • Santa Clara (CA)

On-site
USD 180,000 - 300,000
Equity
Generous Benefits Package