Senior AI Inference Platform Product Manager

NVIDIA

United States

On-site

USD 180,000 - 260,000

Full time

6 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

NVIDIA is seeking a senior product leader to own the inference performance roadmap across AI models and serving stacks. You will orchestrate strategy for memory/state management, request scheduling, and token generation while coordinating across TensorRT-LLM, vLLM, SGLang, and NVIDIA Dynamo.

You will build scalable platforms rather than one-offs, drive cross-functional execution, and ensure credible benchmarking and production readiness for diverse deployments.

Qualifications

  • 12+ years in product management at a technology company or equivalent leadership experience.
  • Deep knowledge of AI inference optimization including KV caching, quantization, and speculative decoding.
  • Familiarity with inference/orchestration frameworks: TensorRT-LLM, vLLM, SGLang, NVIDIA Dynamo.
  • Proven track record of independent strategy development and shipping outcomes.
  • Experience running a live product: release management and customer support processes.

Responsibilities

  • Own the inference performance roadmap. Set direction across the stack: models, memory/state management, scheduling, and token generation.
  • Build platforms, not one-offs; deliver capabilities that generalize across model families and deployments.
  • Define performance strategy for agentic and multi-turn workloads with cross-turn cache reuse and prioritization.
  • Define framework/ecosystem strategy across TensorRT-LLM, vLLM, SGLang, and Dynamo; partner with OSS and internal teams.
  • Own benchmarking methodology and credible performance figures; publish and reproduce results.
  • Manage day-to-day product life cycle: release readiness, quality bars, and production feedback loops.

Skills

Product management
AI inference optimization
Cross-functional leadership
Strategy development

Education

BS/MS/PhD in CS/CE or related field

Tools

TensorRT-LLM
vLLM
SGLang
NVIDIA Dynamo

Job description

NVIDIA is seeking a senior product leader to own the inference performance roadmap across AI models and serving stacks. You will orchestrate strategy for memory/state management, request scheduling, and token generation while coordinating across TensorRT-LLM, vLLM, SGLang, and NVIDIA Dynamo.

You will build scalable platforms rather than one-offs, drive cross-functional execution, and ensure credible benchmarking and production readiness for diverse deployments.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Product Manager — AI Inference Platform
Senior Product Manager — AI Inference Platform

NVIDIA Corporation • Santa Clara (CA)

Hybrid
USD 208,000 - 380,000
Equity
Benefits
Lead Product Manager, AI Inference Platform
Lead Product Manager, AI Inference Platform

NVIDIA • Santa Clara (CA)

On-site
USD 208,000 - 328,000
Equity
Benefits
Senior PM, AI Inference Platform - Equity & Impact
Senior PM, AI Inference Platform - Equity & Impact

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 168,000 - 328,000
Equity
Benefits
Senior Product Manager - AI Inference Performance
Senior Product Manager - AI Inference Performance

NVIDIA • United States

On-site
USD 180,000 - 260,000
Senior Product Manager - AI Platform Inference
Senior Product Manager - AI Platform Inference

NVIDIA • Santa Clara (CA)

On-site
USD 208,000 - 328,000
Equity
Benefits
Senior PM, AI Inference Platform & Tools
Senior PM, AI Inference Platform & Tools

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 168,000 - 328,000
Equity
Benefits
AI Inference Platform Lead – Equity Eligible
AI Inference Platform Lead – Equity Eligible

NVIDIA • Washington

On-site
USD 184,000 - 357,000
Equity
Benefits package
Senior AI Inference Performance Engineer
Senior AI Inference Performance Engineer

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Senior Product Manager - AI Platform Inference
Senior Product Manager - AI Platform Inference

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 168,000 - 328,000
Equity
Benefits
Engineering Manager, AI Inference & GPU Scaling
Engineering Manager, AI Inference & GPU Scaling

NVIDIA Corporation • Santa Clara (CA)

Hybrid
USD 184,000 - 357,000
Equity
Comprehensive benefits
Hybrid work model