AI Inference Software Manager — GPU-Accelerated, Equity

NVIDIA AI

Santa Clara (UT)

On-site

USD 280,000 - 355,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity
Benefits

Job summary

NVIDIA is seeking an exceptional Manager, Deep Learning Inference Software, to lead a world-class engineering team focused on deploying AI models with GPU acceleration. You will shape software powering large language models and multimodal AI, leveraging open-source frameworks like vLLM, SGLang, and FlashInfer.

You will drive strategy, mentor engineers, oversee CUDA, Triton, and CUTLASS optimization, and collaborate with compiler and research teams to deliver end-to-end inference pipelines across

Qualifications

  • MS, PhD, or equivalent experience in Computer Science, Electrical/Computer Engineering, or related field.
  • 6+ years of overall software development experience with 3+ years in technical leadership or engineering management.
  • Strong background in C/C++ software design; Python is a plus.
  • Hands-on experience with GPU programming (CUDA, Triton, CUTLASS) and performance optimization.
  • Proven record of deploying or optimizing deep learning models in production environments.
  • Experience leading teams using Agile or collaborative software development practices.

Responsibilities

  • Lead, mentor, and scale a high-performing engineering team focused on deep learning inference and GPU-accelerated software.
  • Drive strategy, roadmap, and execution of NVIDIA’s inference frameworks engineering, focusing on Client AI.
  • Partner with internal compiler, libraries, and research teams to deliver end-to-end optimized inference pipelines across NVIDIA accelerators.
  • Oversee performance tuning, profiling, and optimization of large-scale models for LLM, multimodal, and generative AI applications.
  • Guide engineers in adopting best practices for CUDA, Triton, CUTLASS, and multi-GPU communications (NIXL, NCCL, NVSHMEM).
  • Represent the team in roadmap and planning discussions, ensuring alignment with broader AI and software strategies.

Skills

C/C++ design
Python
GPU programming
Agile leadership
team mentorship

Education

MS/PhD in CS/EE

Tools

CUDA
Triton
CUTLASS
NCCL/NVSHMEM
vLLM/SGLang/TensorRT-LLM

Job description

NVIDIA is seeking an exceptional Manager, Deep Learning Inference Software, to lead a world-class engineering team focused on deploying AI models with GPU acceleration. You will shape software powering large language models and multimodal AI, leveraging open-source frameworks like vLLM, SGLang, and FlashInfer.

You will drive strategy, mentor engineers, oversee CUDA, Triton, and CUTLASS optimization, and collaborate with compiler and research teams to deliver end-to-end inference pipelines across

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Head of GPU-Accelerated AI Inference
Head of GPU-Accelerated AI Inference

NVIDIA • Washington

On-site
USD 224,000 - 432,000
Equity
Benefits package
Engineering Manager, AI Inference — GPU-Accelerated DL
Engineering Manager, AI Inference — GPU-Accelerated DL

NVIDIA • California (MO)

On-site
USD 272,000 - 432,000
Head of GPU-Accelerated AI Inference
Head of GPU-Accelerated AI Inference

Nvidia Corporation in • Santa Clara (CA)

On-site
USD 224,000 - 431,000
Equity
Benefits
Engineering Manager: GPU-Accelerated AI Inference
Engineering Manager: GPU-Accelerated AI Inference

NVIDIA • Illinois

On-site
USD 224,000 - 432,000
Equity
Benefits
AI Inference Platform Lead – Equity Eligible
AI Inference Platform Lead – Equity Eligible

NVIDIA • Washington

On-site
USD 184,000 - 357,000
Equity
Benefits package
Engineering Manager, Deep Learning Inference — GPU AI
Engineering Manager, Deep Learning Inference — GPU AI

NVIDIA AI • Santa Clara (CA)

On-site
USD 224,000 - 431,000
Equity compensation
Benefits
Engineering Manager - Deep Learning Inference on GPUs
Engineering Manager - Deep Learning Inference on GPUs

NVIDIA • Massachusetts

On-site
USD 224,000 - 431,000
Equity compensation
Comprehensive benefits
Engineering Manager, GPU AI Inference & Open-Source
Engineering Manager, GPU AI Inference & Open-Source

Nvidia Corporation • Santa Clara (CA)

On-site
USD 272,000 - 431,250
Equity
Benefits package
Engineering Manager - GPU AI Inference & Frameworks
Engineering Manager - GPU AI Inference & Frameworks

NVIDIA • Town of Texas (WI)

On-site
USD 184,000 - 357,000
Equity
Benefits
Engineering Manager, Deep Learning Inference & GPU
Engineering Manager, Deep Learning Inference & GPU

NVIDIA • Georgia

On-site
USD 224,000 - 432,000
Equity
Benefits