Senior GPU Kernel Optimizer for LLM Inference

NVIDIA

Seattle (WA)

On-site

USD 184,000 - 287,500

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

NVIDIA is seeking a Sr. Inference Engineer to push GPU kernel optimization for LLM inference. The role focuses on silicon-measured kernel benchmarking, model-level performance projection, and agentic optimization systems that improve kernels at the assembly level.

You will collaborate with compiler, hardware, kernel, and framework teams to surface bottlenecks and deliver production-grade performance gains, with a base salary range clearly stated in the posting.

Qualifications

  • Master's or PhD in Computer Science, Computer Engineering, or a related field, or equivalent experience.
  • 6+ years of relevant industry experience.
  • Experience building or directing agentic AI systems — code generation, automated optimization, or multi-step reasoning workflows.
  • Strong Python and C++ skills with proven software engineering fundamentals.
  • Hands-on GPU profiling with CUPTI, NSYS, and NCU; proven track record to attribute bottlenecks across kernel execution, compiler decisions, and runtime scheduling.
  • Direct experience with LLM inference frameworks such as TRT-LLM, SGLang, or vLLM and clear understanding of how kernel selection drives model-level throughput and latency.
  • Working knowledge of GPU kernel optimization — CUDA, CUTLASS, Triton, or equivalent — and the ability to read PTX or SASS output.

Responsibilities

  • Drive GPU kernel microbenchmarking for LLM inference across configurations.
  • Analyze end-to-end model performance and surface optimization opportunities.
  • Apply agentic optimization to diagnose and validate kernel performance improvements.

Skills

Agentic AI systems
Python
C++
GPU profiling
LLM inference frameworks
Kernel optimization
Reading PTX/SASS

Education

Master's or PhD in CS/CE or related

Job description

NVIDIA is seeking a Sr. Inference Engineer to push GPU kernel optimization for LLM inference. The role focuses on silicon-measured kernel benchmarking, model-level performance projection, and agentic optimization systems that improve kernels at the assembly level.

You will collaborate with compiler, hardware, kernel, and framework teams to surface bottlenecks and deliver production-grade performance gains, with a base salary range clearly stated in the posting.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior LLM Inference: GPU Kernel Optimization
Senior LLM Inference: GPU Kernel Optimization

NVIDIA • Austin (TX)

On-site
USD 184,000 - 287,500
Equity
Benefits
Senior Inference Engineer: AI-Driven GPU Kernel Optimization
Senior Inference Engineer: AI-Driven GPU Kernel Optimization

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 184,000 - 288,000
Senior Inference Engineer, GPU Kernel Optimization
Senior Inference Engineer, GPU Kernel Optimization

NVIDIA • Seattle (WA)

On-site
USD 184,000 - 287,500
Senior Inference Engineer, GPU Kernel Optimization
Senior Inference Engineer, GPU Kernel Optimization

NVIDIA • Austin (TX)

On-site
USD 184,000 - 287,500
Equity
Benefits
Senior Inference Engineer, GPU Kernel Optimization
Senior Inference Engineer, GPU Kernel Optimization

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 184,000 - 288,000
Senior AI Inference Performance Engineer
Senior AI Inference Performance Engineer

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Senior DL Inference Engineer - GPU/LLM Performance & Equity
Senior DL Inference Engineer - GPU/LLM Performance & Equity

NVIDIA • Washington

On-site
USD 184,000 - 288,000
Equity
Benefits
Senior DL Inference Engineer - GPU & LLM Performance
Senior DL Inference Engineer - GPU & LLM Performance

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Benefits
Kernel Engineer: GPU Performance & Inference
Kernel Engineer: GPU Performance & Inference

Acceler8 Talent • San Francisco (CA)

On-site
USD 180,000 - 240,000
Senior GPU Kernel & Performance Engineer
Senior GPU Kernel & Performance Engineer

Designworks Talent LLC • Bellevue (WA)

Hybrid
USD 180,000 - 240,000
Medical insurance
Dental insurance
Vision insurance
+2