Senior Product Manager - AI Inference Performance

NVIDIA

United States

On-site

USD 180,000 - 260,000

Full time

6 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

NVIDIA is seeking a senior product leader to own the inference performance roadmap across AI models and serving stacks. You will orchestrate strategy for memory/state management, request scheduling, and token generation while coordinating across TensorRT-LLM, vLLM, SGLang, and NVIDIA Dynamo.

You will build scalable platforms rather than one-offs, drive cross-functional execution, and ensure credible benchmarking and production readiness for diverse deployments.

Qualifications

  • 12+ years in product management at a technology company or equivalent leadership experience.
  • Deep knowledge of AI inference optimization including KV caching, quantization, and speculative decoding.
  • Familiarity with inference/orchestration frameworks: TensorRT-LLM, vLLM, SGLang, NVIDIA Dynamo.
  • Proven track record of independent strategy development and shipping outcomes.
  • Experience running a live product: release management and customer support processes.

Responsibilities

  • Own the inference performance roadmap. Set direction across the stack: models, memory/state management, scheduling, and token generation.
  • Build platforms, not one-offs; deliver capabilities that generalize across model families and deployments.
  • Define performance strategy for agentic and multi-turn workloads with cross-turn cache reuse and prioritization.
  • Define framework/ecosystem strategy across TensorRT-LLM, vLLM, SGLang, and Dynamo; partner with OSS and internal teams.
  • Own benchmarking methodology and credible performance figures; publish and reproduce results.
  • Manage day-to-day product life cycle: release readiness, quality bars, and production feedback loops.

Skills

Product management
AI inference optimization
Cross-functional leadership
Strategy development

Education

BS/MS/PhD in CS/CE or related field

Tools

TensorRT-LLM
vLLM
SGLang
NVIDIA Dynamo

Job description

NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It's a unique legacy of innovation that's fueled by great technology-and amazing people. Today, we're tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what's never been done before takes vision, innovation, and the world's best talent. As an NVIDIAN, you'll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world.

What You'll Be Doing:
  • Own the inference performance roadmap. Set direction across the stack: how models are represented, how memory and state are managed, how requests are scheduled and served, and how tokens get generated. The techniques change fast. Judge which ones matter, then decide what we build, what we adopt, and what we retire.
  • Build platforms, not one-offs. Deliver capabilities that generalize across model families, deployment topologies, and customer sizes. Build for easy adoption, sane defaults, and extensibility.
  • Agentic and Multi-Turn Workloads: Define the performance strategy for agentic applications, where long-running sessions, tool-call stalls, and unpredictable output lengths break the assumptions built into single-turn serving. Drive capabilities around cross-turn cache reuse, request prioritization, and efficient handling of idle time in agent loops.
  • Framework & Ecosystem Strategy: Define how our optimizations land across TensorRT-LLM, vLLM, SGLang, and NVIDIA Dynamo. Partner with open-source communities and internal engineering teams so customers get great performance on NVIDIA hardware.
  • Benchmarking & Performance Claims: Own how performance is measured, published, and reproduced. Define the benchmark methodology, the metrics that matter (TTFT, ITL, throughput per GPU, cost per million tokens), and the guardrails that keep our numbers credible.
  • Run the product day to day. Own release readiness, quality bars, regression tracking, customer blocking issues, and the feedback loop from production deployments back into the roadmap.
What We Need to See:
  • 12+ years in product management at a technology company, or comparable time as a founder, engineering lead, or technical product owner.
  • Depth in AI inference optimization: KV caching and reuse, quantization, speculative decoding, disaggregated serving. Know how each one moves accuracy, latency, and cost.
  • Familiarity with the inference and orchestration frameworks customers use: TensorRT-LLM, vLLM, SGLang, NVIDIA Dynamo, and the surrounding serving and scheduling ecosystem.
  • Proven track record of working independently - you can take an ambiguous problem space, define the strategy, and drive it to a shipped outcome without waiting to be told what to do next.
  • Operational experience running a live product: release management, quality and regression rigor, customer issues, and support processes.
  • Skill at translating low-level capability into business value - lower TCO, faster response, better GPU utilization - for engineers and executives alike.
  • BS, MS, or PhD in Computer Science, Computer Engineering, or another relevant area of study (or equivalent experience).
Ways to Stand Out From the Crowd:
  • Engineering experience with LLM inference performance: profiling, kernel-level or serving-level optimization, or building a serving stack!
  • Open-source contributions or product leadership in vLLM, SGLang, TensorRT-LLM, Triton Inference Server, or Dynamo. Production experience at scale counts too: capacity planning, autoscaling, SLA management, or stateful multi-turn applications.
  • A habit of read
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Product Manager - AI Platform Inference
Senior Product Manager - AI Platform Inference

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 168,000 - 328,000
Equity
Benefits
Senior Product Manager - AI Platform Inference
Senior Product Manager - AI Platform Inference

NVIDIA • Santa Clara (CA)

On-site
USD 208,000 - 328,000
Equity
Benefits
AI Senior Engineer - Fulltime
AI Senior Engineer - Fulltime

IMR Soft Llc • Plano (TX)

On-site
USD 140,000 - 190,000
Senior Product Manager, Inference Platform
Senior Product Manager, Inference Platform

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 208,000 - 380,000
Equity
Benefits
Senior Product Manager - AI Platform Inference
Senior Product Manager - AI Platform Inference

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 168,000 - 328,000
Equity
Benefits
Senior AI Inference Platform Product Manager
Senior AI Inference Platform Product Manager

NVIDIA • United States

On-site
USD 180,000 - 260,000
Engineering Manager, Inference Benchmarking — AI Perf
Engineering Manager, Inference Benchmarking — AI Perf

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 224,000 - 357,000
Engineering Manager, Deep Learning Inference
Engineering Manager, Deep Learning Inference

NVIDIA Corporation • Santa Clara (CA)

Hybrid
USD 184,000 - 357,000
Equity
Comprehensive benefits
Hybrid work model
Principal Product Manager, AI Frameworks
Principal Product Manager, AI Frameworks

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 240,000 - 380,000
Staff Applied AI Inference Engineer
Staff Applied AI Inference Engineer

Crusoe Energy Systems • San Francisco (CA)

On-site
USD 180,000 - 240,000
Health benefits
401(k) match
Paid time off
+1