Senior AI Inference Systems Engineer (GPU & HPC)

NVIDIA

Santa Clara (CA)

Hybrid

USD 184,000 - 357,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity
Benefits

Job summary

NVIDIA is seeking highly skilled software engineers to build AI inference systems that scale across multi-GPU, multi-node, and multi-cloud deployments. You will architect high-performance inference stacks, optimize kernels and compilers, and collaborate across inference, compiler, scheduling, and performance teams to push the frontier of accelerated AI computing.

Responsibilities include contributing features to vLLM, benchmarking, and driving MLPerf Inference submissions, with emphasis on

Qualifications

  • Bachelor’s degree in CS/CE/SE with 7+ years of experience or Master’s with 5+ years, or PhD with publications in ML Systems/HPC.
  • Strong programming in Python and C/C++; Go or Rust is a plus; solid CS fundamentals.
  • Knowledge of performance engineering in ML frameworks (PyTorch) and inference engines (vLLM/SGLang).
  • Familiarity with GPU programming (CUDA), memory hierarchies, Nsight tools; experience with profiling.

Responsibilities

  • Contribute features to vLLM and optimize the inference framework for NVIDIA GPUs.
  • Develop, optimize, and benchmark GPU kernels; extend compiler infrastructure and DSLs.
  • Define benchmarking methodologies; contribute to MLPerf Inference submissions.
  • Architect containerized large-scale inference deployments across clouds.
  • Publish original research and bridge ideas into NVIDIA software products.

Skills

Python
C/C++
Go or Rust
Algorithms & data structures
Distributed systems
GPU programming
CUDA
Performance tuning
PyTorch
NCCL

Education

Bachelor's degree in CS/CE/SE
Master's degree in CS/CE/SE
PhD in ML Systems / GPU architecture / HPC

Tools

Docker
Kubernetes
Slurm
Nsight Systems/Compute
MLIR/LLVM
Triton
TorchDynamo/Inductor
CUTLASS

Job description

NVIDIA is seeking highly skilled software engineers to build AI inference systems that scale across multi-GPU, multi-node, and multi-cloud deployments. You will architect high-performance inference stacks, optimize kernels and compilers, and collaborate across inference, compiler, scheduling, and performance teams to push the frontier of accelerated AI computing.

Responsibilities include contributing features to vLLM, benchmarking, and driving MLPerf Inference submissions, with emphasis on

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Systems Engineer - Inference & GPU Kernels
Senior AI Systems Engineer - Inference & GPU Kernels

NVIDIA AI • California (MO)

On-site
USD 184,000 - 288,000
Senior AI Inference Systems Engineer | GPU Kernels & Runtime
Senior AI Inference Systems Engineer | GPU Kernels & Runtime

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 288,000
Senior AI Inference Engineer: GPU Kernels & LLM Runtimes
Senior AI Inference Engineer: GPU Kernels & LLM Runtimes

NVIDIA • Redmond (WA)

On-site
USD 184,000 - 288,000
Equity
Benefits
Senior AI Inference Engineer — GPU-Accelerated DL Systems
Senior AI Inference Engineer — GPU-Accelerated DL Systems

NVIDIA Corporation • California (MO)

On-site
USD 152,000 - 287,500
Equity
Benefits
GPU Inference Engineer — Deep Learning
GPU Inference Engineer — Deep Learning

2100 NVIDIA USA • California (MO)

On-site
USD 124,000 - 242,000
Equity
Benefits
Senior AI Inference Systems Engineer – GPU Kernels
Senior AI Inference Systems Engineer – GPU Kernels

NVIDIA • Seattle (WA)

On-site
USD 184,000 - 288,000
Senior AI Systems Engineer - Inference & GPU Kernels
Senior AI Systems Engineer - Inference & GPU Kernels

NVIDIA • Westford (MA)

On-site
USD 184,000 - 288,000
Equity
Benefits
Senior AI Inference Optimization Engineer
Senior AI Inference Optimization Engineer

Nvidia Corporation in • Santa Clara (CA)

Hybrid
USD 124,000 - 196,000
Equity
Benefits package
Hybrid work model
GPU Inference Performance Engineer — Equity & Optimization
GPU Inference Performance Engineer — Equity & Optimization

Nvidia Corporation • Santa Clara (CA)

On-site
USD 152,000 - 242,000
Senior DL Inference Engineer — GPU-Optimized LLMs (Remote)
Senior DL Inference Engineer — GPU-Optimized LLMs (Remote)

NVIDIA Corporation • Northern (KY)

Hybrid
USD 152,000 - 288,000