Senior AI Inference Engineer: GPU Kernels & LLM Runtimes

NVIDIA

Redmond (WA)

On-site

USD 184,000 - 288,000

Full time

2 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity
Benefits

Job summary

NVIDIA is seeking outstanding AI systems engineers to advance inference software, building libraries, code generators, and GPU kernel technologies for our hardware architecture. You will design abstractions for LLM serving engines and just-in-time compilers to accelerate large language models and agents.

You will work with NVIDIA teams on deep learning frameworks and kernels, contributing to open source ecosystems like FlashInfer, vLLM, and SGLang while shaping high-performance AI workloads

Qualifications

  • Masters degree in Computer Science, Electrical Engineering, or related field; PhD preferred.
  • 6+ years experience with ML/DL systems development.
  • Strong experience with deep learning frameworks (e.g. PyTorch, JAX, TensorFlow, ONNX) and inference engines/runtimes (vLLM, SGLang, MLC).
  • Strong Python and C/C++ programming skills.
  • Strong GPU kernel development and performance optimization experience (CUDA C/C++, cuTile, Triton) with hands-on Matrix Multiplication

Responsibilities

  • Innovate and develop new AI systems technologies for efficient inference.
  • Design, implement, and optimize kernels for high-impact AI workloads.
  • Create extensible abstractions for LLM serving engines.
  • Build just-in-time compilers and runtimes for ML workloads.
  • Collaborate with NVIDIA engineers across frameworks, libraries, kernels, and GPU teams.
  • Contribute to open source projects like FlashInfer, vLLM, and SGLang

Skills

Python
C/C++
GPU kernel development
CUDA
Triton
Matrix Multiplication

Education

Masters degree
PhD preferred

Tools

CUDA
Triton

Job description

NVIDIA is seeking outstanding AI systems engineers to advance inference software, building libraries, code generators, and GPU kernel technologies for our hardware architecture. You will design abstractions for LLM serving engines and just-in-time compilers to accelerate large language models and agents.

You will work with NVIDIA teams on deep learning frameworks and kernels, contributing to open source ecosystems like FlashInfer, vLLM, and SGLang while shaping high-performance AI workloads

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Systems Engineer - Inference & GPU Kernels
Senior AI Systems Engineer - Inference & GPU Kernels

NVIDIA AI • California (MO)

On-site
USD 184,000 - 288,000
Senior AI Inference Systems Engineer | GPU Kernels & Runtime
Senior AI Inference Systems Engineer | GPU Kernels & Runtime

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 288,000
Senior AI Inference Systems Engineer – GPU Kernels
Senior AI Inference Systems Engineer – GPU Kernels

NVIDIA • Seattle (WA)

On-site
USD 184,000 - 288,000
Senior AI Inference Systems Engineer (GPU & HPC)
Senior AI Inference Systems Engineer (GPU & HPC)

NVIDIA • Santa Clara (CA)

Hybrid
USD 184,000 - 357,000
Equity
Benefits
Senior AI Systems Engineer - Inference & GPU Kernels
Senior AI Systems Engineer - Inference & GPU Kernels

NVIDIA • Westford (MA)

On-site
USD 184,000 - 288,000
Equity
Benefits
Senior AI Inference Engineer — GPU-Accelerated DL Systems
Senior AI Inference Engineer — GPU-Accelerated DL Systems

NVIDIA Corporation • California (MO)

On-site
USD 152,000 - 287,500
Equity
Benefits
GPU Inference Engineer — Deep Learning
GPU Inference Engineer — Deep Learning

2100 NVIDIA USA • California (MO)

On-site
USD 124,000 - 242,000
Equity
Benefits
Senior AI Inference & Kernel Engineer
Senior AI Inference & Kernel Engineer

NVIDIA • Durham (NC)

On-site
USD 184,000 - 288,000
Equity
Benefits
Senior AI Inference & Kernel Engineer
Senior AI Inference & Kernel Engineer

Intel • Austin (TX)

Hybrid
USD 189,000 - 315,000
Stock bonuses
Health benefits
Vacation
Engineering Manager, AI Inference — GPU-Accelerated DL
Engineering Manager, AI Inference — GPU-Accelerated DL

NVIDIA • California (MO)

On-site
USD 272,000 - 432,000