Senior AI Inference Systems Engineer | GPU Kernels & Runtime

NVIDIA

Santa Clara (CA)

On-site

USD 184,000 - 288,000

Full time

8 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

NVIDIA in Santa Clara, CA is seeking outstanding AI systems engineers to advance the inference software stack. You will design and optimize kernels, build new abstractions for LLM serving engines, and contribute to accelerators and runtimes that power large language models and AI workloads.

You will collaborate across teams, work on CUDA C/C++, Triton, and cutting-edge MLIR-based tooling, and participate in open source projects.

Qualifications

  • Masters degree in CS, EE, or related field; PhD preferred.
  • 6+ years experience in ML/DL systems development (academic/industry).
  • Strong experience with deep learning frameworks and inference engines.
  • Strong Python and C/C++ programming skills.
  • Experience in GPU kernel development and performance optimizations (CUDA, cuTile, Triton).

Responsibilities

  • Innovating and developing AI systems technologies for efficient inference.
  • Designing, implementing, and optimizing kernels for high impact AI workloads.
  • Designing and implementing abstractions for LLM serving engines.
  • Building JIT domain specific compilers and runtimes.
  • Collaborating with NVIDIA teams across frameworks, libraries, kernels, and GPU arch teams.

Skills

Python
C/C++
GPU kernel development
ML/DL systems development
Deep learning frameworks experience

Education

Master's degree in CS/EE or related field
PhD preferred

Tools

CUDA C/C++
cuTile
Triton
PyTorch
JAX
TensorFlow
ONNX

Job description

NVIDIA in Santa Clara, CA is seeking outstanding AI systems engineers to advance the inference software stack. You will design and optimize kernels, build new abstractions for LLM serving engines, and contribute to accelerators and runtimes that power large language models and AI workloads.

You will collaborate across teams, work on CUDA C/C++, Triton, and cutting-edge MLIR-based tooling, and participate in open source projects.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Systems Engineer - Inference & GPU Kernels
Senior AI Systems Engineer - Inference & GPU Kernels

NVIDIA AI • California (MO)

On-site
USD 184,000 - 288,000
Senior AI Inference Engineer: GPU Kernels & LLM Runtimes
Senior AI Inference Engineer: GPU Kernels & LLM Runtimes

NVIDIA • Redmond (WA)

On-site
USD 184,000 - 288,000
Equity
Benefits
Senior AI Inference Systems Engineer – GPU Kernels
Senior AI Inference Systems Engineer – GPU Kernels

NVIDIA • Seattle (WA)

On-site
USD 184,000 - 288,000
Senior AI Inference & Kernel Engineer
Senior AI Inference & Kernel Engineer

NVIDIA • Durham (NC)

On-site
USD 184,000 - 288,000
Equity
Benefits
Senior AI Inference Systems Engineer (GPU & HPC)
Senior AI Inference Systems Engineer (GPU & HPC)

NVIDIA • Santa Clara (CA)

Hybrid
USD 184,000 - 357,000
Equity
Benefits
Senior AI Systems Engineer - Inference & GPU Kernels
Senior AI Systems Engineer - Inference & GPU Kernels

NVIDIA • Westford (MA)

On-site
USD 184,000 - 288,000
Equity
Benefits
Senior AI Inference Engineer — GPU-Accelerated DL Systems
Senior AI Inference Engineer — GPU-Accelerated DL Systems

NVIDIA Corporation • California (MO)

On-site
USD 152,000 - 287,500
Equity
Benefits
Senior AI Inference & Kernel Engineer
Senior AI Inference & Kernel Engineer

Intel • Austin (TX)

Hybrid
USD 189,000 - 315,000
Stock bonuses
Health benefits
Vacation
Senior AI Kernel & Inference Engineer
Senior AI Kernel & Inference Engineer

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 184,000 - 288,000
Senior System Software Engineer — GPU AI Inference (Triton)
Senior System Software Engineer — GPU AI Inference (Triton)

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 152,000 - 242,000
Equity
Benefits