Senior LLM Inference: GPU Kernel Optimization

NVIDIA

Austin (TX)

On-site

USD 184,000 - 287,500

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity
Benefits

Job summary

NVIDIA is seeking a Sr. Inference Engineer to push LLM inference performance through GPU kernel optimization.

You will help develop silicon-measured benchmarking, model-level performance projection tooling, and agentic optimization systems, collaborating across compiler, hardware, and framework teams to surface bottlenecks and deliver measurable gains. You will work on GPU kernel microbenchmarking, end-to-end model performance analysis, and agentic optimization, shaping production inference

Qualifications

  • Master's or PhD in CS/CE or related field, or equivalent experience.
  • 6+ years of relevant industry experience.
  • Experience building or directing agentic AI systems — code generation, automated optimization, or multi-step reasoning workflows.
  • Strong Python and C++ skills with proven software engineering fundamentals.
  • Hands-on GPU profiling with CUPTI, NSYS, and NCU; attribution of bottlenecks across kernel execution, compiler decisions, and runtime scheduling.
  • Direct experience with LLM inference frameworks such as TRT-LLM, SGLang, or vLLM and understanding of how kernel-driven throughput and latency.
  • Working knowledge of GPU kernel optimization — CUDA, CUTLASS, Triton; ability to read PTX or SASS output.

Responsibilities

  • Develop silicon-measured kernel benchmarking infrastructure.
  • Build model-level performance projection tooling.
  • Create agentic optimization systems to improve GPU kernels at the assembly level.
  • Drive GPU kernel microbenchmarking across configuration space.
  • Perform end-to-end model performance analysis and derive optimization policies.
  • Collaborate with compiler, hardware, kernel, and framework teams to deliver production-grade gains.

Skills

Python
C++
GPU profiling
LLM frameworks (TRT-LLM, SGLang, vLLM)
Kernel optimization (CUDA, CUTLASS, Tr
Agentic AI systems
Multi-agent orchestration

Education

Master's or PhD in Computer Science/Computer Engineering or related field
Equivalent experience

Tools

CUPTI
NSYS
NCU

Job description

NVIDIA is seeking a Sr. Inference Engineer to push LLM inference performance through GPU kernel optimization.

You will help develop silicon-measured benchmarking, model-level performance projection tooling, and agentic optimization systems, collaborating across compiler, hardware, and framework teams to surface bottlenecks and deliver measurable gains. You will work on GPU kernel microbenchmarking, end-to-end model performance analysis, and agentic optimization, shaping production inference

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior GPU Kernel Optimizer for LLM Inference
Senior GPU Kernel Optimizer for LLM Inference

NVIDIA • Seattle (WA)

On-site
USD 184,000 - 288,000
Senior Inference Engineer: GPU Kernel Optimizations + Equity
Senior Inference Engineer: GPU Kernel Optimizations + Equity

Nvidia Corporation • Santa Clara (CA)

On-site
USD 184,000 - 288,000
Equity
Comprehensive benefits
LLM Inference Performance Engineer - GPU Kernel Optimizer
LLM Inference Performance Engineer - GPU Kernel Optimizer

Intel • Folsom (CA)

Hybrid
USD 171,000 - 315,000
Stock bonuses
Health benefits
Hybrid work model
Senior GPU Kernel Engineer: Inference Performance
Senior GPU Kernel Engineer: Inference Performance

Neura Market • Sunnyvale (CA), Northern (KY)

Hybrid
USD 182,000 - 242,000
Medical, dental, and vision insurance
Company-paid Life Insurance
Tuition Reimbursement
+2
Senior LLM Inference Engineer: Performance & Optimization
Senior LLM Inference Engineer: Performance & Optimization

Confidential • United States

On-site
USD 180,000 - 240,000
Senior LLM Inference Architect — Performance & Benchmarking
Senior LLM Inference Architect — Performance & Benchmarking

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity and benefits
Senior GPU Kernel Engineer: Inference Throughput
Senior GPU Kernel Engineer: Inference Throughput

CoreWeave • Bellevue (WA)

On-site
USD 182,000 - 242,000
Medical, dental, and vision insurance
Company-paid Life Insurance
401(k) with employer match
+3
Senior Inference Engineer, GPU Kernel Optimization
Senior Inference Engineer, GPU Kernel Optimization

Nvidia Corporation • Santa Clara (CA)

On-site
USD 184,000 - 288,000
Equity
Comprehensive benefits
Senior Inference Engineer, GPU Kernel Optimization
Senior Inference Engineer, GPU Kernel Optimization

NVIDIA • Austin (TX)

On-site
USD 184,000
Equity
Benefits
Senior Inference Engineer, GPU Kernel Optimization
Senior Inference Engineer, GPU Kernel Optimization

NVIDIA • Seattle (WA)

On-site
USD 184,000 - 288,000