Senior GPU Kernel Optimizer for LLM Inference

NVIDIA

Seattle (WA)

On-site

USD 184,000 - 287,500

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

NVIDIA is seeking a Sr. Inference Engineer to push GPU kernel optimization for LLM inference. The role focuses on silicon-measured kernel benchmarking, model-level performance projection, and agentic optimization systems that improve kernels at the assembly level.

You will collaborate with compiler, hardware, kernel, and framework teams to surface bottlenecks and deliver production-grade performance gains, with a base salary range clearly stated in the posting.

Qualifications

  • Master's or PhD in Computer Science, Computer Engineering, or a related field, or equivalent experience.
  • 6+ years of relevant industry experience.
  • Experience building or directing agentic AI systems — code generation, automated optimization, or multi-step reasoning workflows.
  • Strong Python and C++ skills with proven software engineering fundamentals.
  • Hands-on GPU profiling with CUPTI, NSYS, and NCU; proven track record to attribute bottlenecks across kernel execution, compiler decisions, and runtime scheduling.
  • Direct experience with LLM inference frameworks such as TRT-LLM, SGLang, or vLLM and clear understanding of how kernel selection drives model-level throughput and latency.
  • Working knowledge of GPU kernel optimization — CUDA, CUTLASS, Triton, or equivalent — and the ability to read PTX or SASS output.

Responsibilities

  • Drive GPU kernel microbenchmarking for LLM inference across configurations.
  • Analyze end-to-end model performance and surface optimization opportunities.
  • Apply agentic optimization to diagnose and validate kernel performance improvements.

Skills

Agentic AI systems
Python
C++
GPU profiling
LLM inference frameworks
Kernel optimization
Reading PTX/SASS

Education

Master's or PhD in CS/CE or related

Job description

NVIDIA is seeking a Sr. Inference Engineer to push GPU kernel optimization for LLM inference. The role focuses on silicon-measured kernel benchmarking, model-level performance projection, and agentic optimization systems that improve kernels at the assembly level.

You will collaborate with compiler, hardware, kernel, and framework teams to surface bottlenecks and deliver production-grade performance gains, with a base salary range clearly stated in the posting.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior LLM Inference: GPU Kernel Optimization
Senior LLM Inference: GPU Kernel Optimization

NVIDIA • Austin (TX)

On-site
USD 184,000
Equity
Benefits
Senior Inference Engineer: GPU Kernel Optimizations + Equity
Senior Inference Engineer: GPU Kernel Optimizations + Equity

Nvidia Corporation • Santa Clara (CA)

On-site
USD 184,000 - 288,000
Equity
Comprehensive benefits
Senior Inference Engineer, GPU Kernel Optimization
Senior Inference Engineer, GPU Kernel Optimization

NVIDIA • Austin (TX)

On-site
USD 184,000
Equity
Benefits
Senior Inference Engineer, GPU Kernel Optimization
Senior Inference Engineer, GPU Kernel Optimization

NVIDIA • Seattle (WA)

On-site
USD 184,000 - 288,000
Senior Inference Engineer, GPU Kernel Optimization
Senior Inference Engineer, GPU Kernel Optimization

Nvidia Corporation • Santa Clara (CA)

On-site
USD 184,000 - 288,000
Equity
Comprehensive benefits
Senior GPU Kernel Engineer: Inference Performance
Senior GPU Kernel Engineer: Inference Performance

Neura Market • Sunnyvale (CA), Northern (KY)

Hybrid
USD 182,000 - 242,000
Medical, dental, and vision insurance
Company-paid Life Insurance
Tuition Reimbursement
+2
Senior GPU Kernel Engineer: Inference Throughput
Senior GPU Kernel Engineer: Inference Throughput

CoreWeave • Bellevue (WA)

On-site
USD 182,000 - 242,000
Medical, dental, and vision insurance
Company-paid Life Insurance
401(k) with employer match
+3
LLM Inference Performance Engineer - GPU Kernel Optimizer
LLM Inference Performance Engineer - GPU Kernel Optimizer

Intel • Folsom (CA)

Hybrid
USD 171,000 - 315,000
Stock bonuses
Health benefits
Hybrid work model
Senior GPU Kernel Engineer for High-Performance AI Inference
Senior GPU Kernel Engineer for High-Performance AI Inference

CoreWeave • Sunnyvale (CA)

On-site
USD 182,000 - 242,000
Medical, dental, and vision insurance
401(k) with employer match
ESPP
+2
GPU Inference Performance Engineer — Equity & Optimization
GPU Inference Performance Engineer — Equity & Optimization

Nvidia Corporation • Santa Clara (CA)

On-site
USD 152,000 - 242,000