Senior AI Inference Platform PM — Performance & Scale

NVIDIA

Santa Clara (CA)

On-site

USD 208,000 - 328,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

NVIDIA is seeking a highly technical Product Manager to own AI inference optimization on NVIDIA hardware, from single-GPU workstations to large data centers. You will translate deep optimization techniques into capabilities that customers can adopt across the inference stack.

The role spans framework integration with TensorRT-LLM, vLLM, SGLang, and NVIDIA Dynamo, plus benchmarking, release readiness, and cross-functional collaboration to deliver measurable performance gains and compelling

Qualifications

  • 12+ years in product management at a tech company or comparable senior leadership experience.
  • Deep expertise in AI inference optimization: caching, quantization, speculative decoding, disaggregated serving.
  • Familiarity with inference/orchestration frameworks: TensorRT-LLM, vLLM, SGLang, NVIDIA Dynamo.
  • Proven ability to work independently from ambiguous problem spaces to shipped outcomes.
  • Operational experience with live product releases, quality/regression processes, and customer issues.
  • Ability to translate low-level capabilities into business value (lower TCO, faster responses, better GPU utilization).
  • BS/MS/PhD in CS/CE or equivalent experience.

Responsibilities

  • Own the inference performance roadmap; define model representation, memory/state management, scheduling, and token generation.
  • Build platforms with generalization across models, topologies, and customer sizes for easy adoption.
  • Define performance strategy for agentic workloads and multi-turn sessions.
  • Shape framework integration across TensorRT-LLM, vLLM, SGLang, and Dynamo; partner with open-source and internal teams.
  • Own benchmarking, measurement methods, and credible performance claims.
  • Manage day-to-day product operations: release readiness, quality bars, regression tracking, and customer feedback loops.

Skills

AI inference optimization
Product management leadership
Independent decision-making
Cross-functional collaboration
Operational excellence
Communication

Education

BS/MS/PhD in Computer Science or Computer Engineering

Tools

TensorRT-LLM
vLLM
SGLang
Dynamo

Job description

NVIDIA is seeking a highly technical Product Manager to own AI inference optimization on NVIDIA hardware, from single-GPU workstations to large data centers. You will translate deep optimization techniques into capabilities that customers can adopt across the inference stack.

The role spans framework integration with TensorRT-LLM, vLLM, SGLang, and NVIDIA Dynamo, plus benchmarking, release readiness, and cross-functional collaboration to deliver measurable performance gains and compelling

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Inference Performance Product Manager
Senior AI Inference Performance Product Manager

Nvidia Corporation in • Santa Clara (CA)

On-site
USD 208,000 - 328,000
Senior PM, AI Inference Platform - Equity & Impact
Senior PM, AI Inference Platform - Equity & Impact

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 168,000 - 328,000
Equity
Benefits
Senior AI Inference Platform Product Lead
Senior AI Inference Platform Product Lead

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 208,000 - 328,000
Equity
Comprehensive benefits
Inclusive culture
Senior PM, AI Inference Platform & Tools
Senior PM, AI Inference Platform & Tools

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 168,000 - 328,000
Equity
Benefits
Senior AI Inference Performance Engineer
Senior AI Inference Performance Engineer

NVIDIA AI • Santa Clara (CA)

On-site
USD 140,000 - 230,000
Lead Product Manager, AI Inference Platform
Lead Product Manager, AI Inference Platform

NVIDIA • Santa Clara (CA)

On-site
USD 208,000 - 328,000
Equity
Benefits
Senior AI Inference Performance Engineer — Scale GPUs
Senior AI Inference Performance Engineer — Scale GPUs

NVIDIA • California (MO)

On-site
USD 124,000 - 196,000
Equity eligibility
Senior Product Manager - AI Platform Inference
Senior Product Manager - AI Platform Inference

NVIDIA • Santa Clara (CA)

On-site
USD 208,000 - 328,000
Equity
Benefits
Senior Product Manager - AI Inference Performance
Senior Product Manager - AI Inference Performance

Nvidia Corporation in • Santa Clara (CA)

On-site
USD 208,000 - 328,000
GPU Inference Performance Engineer — Equity & Optimization
GPU Inference Performance Engineer — Equity & Optimization

Nvidia Corporation • Santa Clara (CA)

On-site
USD 152,000 - 242,000