Senior AI Inference Systems Engineer – GPU Kernels

NVIDIA

Seattle (WA)

On-site

USD 184,000 - 288,000

Full time

6 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

NVIDIA in Seattle, WA seeks outstanding AI systems engineers to advance the inference software stack for AI workloads. You will develop libraries, code generators, and GPU kernel technologies for NVIDIA hardware, including new abstractions and runtimes for large language models.

The role emphasizes implementing high-performance kernels, JIT compilers, and collaboration across frameworks. Strong CUDA/C/C++ and Python skills are essential, with 6+ years in ML/DL systems.

Qualifications

  • Master's degree in Computer Science, Electrical Engineering, or related field (or equivalent experience); PhD are preferred.
  • 6+ years (academic/ industry) experience with ML/DL systems development preferable
  • Strong experience in developing or using deep learning frameworks (e.g. PyTorch, JAX, TensorFlow, ONNX, etc) and ideally inference engines and runtimes such as vLLM, SGLang, and MLC.
  • Strong Python and C/C++ programming skills
  • Strong experience in GPU kernel development and performance optimizations (especially using CUDA C/C++, cuTile, Triton, or similar) with hands-on experience with Matrix Multiplication

Responsibilities

  • Innovating and developing new AI systems technologies for efficient inference
  • Designing, implementing, and optimizing kernels for high impact AI workloads
  • Designing and implementing extensible abstractions for LLM serving engines
  • Building efficient just-in-time domain specific compilers and runtimes
  • Collaborating closely with other engineers at NVIDIA across deep learning frameworks, libraries, kernels, and GPU arch teams
  • Contributing to open source communities like FlashInfer, vLLM, and SGLang

Skills

Python
C/C++
GPU kernel development
CUDA
Deep learning frameworks
ML/DL systems

Education

Master's degree in CS/EE (or related)
PhD preferred

Tools

vLLM
SGLang
MLC
FlashInfer
Flash Attention

Job description

NVIDIA in Seattle, WA seeks outstanding AI systems engineers to advance the inference software stack for AI workloads. You will develop libraries, code generators, and GPU kernel technologies for NVIDIA hardware, including new abstractions and runtimes for large language models.

The role emphasizes implementing high-performance kernels, JIT compilers, and collaboration across frameworks. Strong CUDA/C/C++ and Python skills are essential, with 6+ years in ML/DL systems.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Inference Systems Engineer | GPU Kernels & Runtime
Senior AI Inference Systems Engineer | GPU Kernels & Runtime

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 288,000
Senior AI Systems Engineer - Inference & GPU Kernels
Senior AI Systems Engineer - Inference & GPU Kernels

NVIDIA • Westford (MA)

On-site
USD 184,000 - 288,000
Equity
Benefits
Senior AI Inference Engineer: GPU Kernels & LLM Runtimes
Senior AI Inference Engineer: GPU Kernels & LLM Runtimes

NVIDIA • Redmond (WA)

On-site
USD 184,000 - 288,000
Equity
Benefits
Senior AI Systems Engineer - Inference & GPU Kernels
Senior AI Systems Engineer - Inference & GPU Kernels

NVIDIA AI • California (MO)

On-site
USD 184,000 - 288,000
Senior AI Inference & Kernel Engineer
Senior AI Inference & Kernel Engineer

NVIDIA • Durham (NC)

On-site
USD 184,000 - 288,000
Equity
Benefits
Senior AI Inference Systems Engineer (GPU & HPC)
Senior AI Inference Systems Engineer (GPU & HPC)

NVIDIA • Santa Clara (CA)

Hybrid
USD 184,000 - 357,000
Equity
Benefits
Senior AI Inference Engineer — GPU-Accelerated DL Systems
Senior AI Inference Engineer — GPU-Accelerated DL Systems

NVIDIA Corporation • California (MO)

On-site
USD 152,000 - 287,500
Equity
Benefits
GPU Kernel Engineer for High-Performance AI Inference
GPU Kernel Engineer for High-Performance AI Inference

Baseten • San Francisco (CA)

On-site
USD 180,000 - 360,000
Competitive compensation, including equity
100% coverage of medical, dental, and vision insurance
Flexible PTO policy
+3
Senior GPU AI Inference Engineer
Senior GPU AI Inference Engineer

Nvidia Corporation • Santa Clara (CA)

Hybrid
USD 224,000 - 432,000
Equity
Benefits
AI Inference GPU Systems Engineer
AI Inference GPU Systems Engineer

Vast.ai Inc. • San Francisco (CA)

On-site
USD 120,000 - 160,000
Comprehensive health, dental, vision, and life insurance
401(k) with company match
Early-stage equity
+2