Product Manager, Inference Platform

Baseten

San Francisco (CA)

On-site

USD 210,000 - 280,000

Full time

5 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive equity
Medical, dental, vision insurance
Flexible PTO including Winter Break
Paid parental leave
Fertility stipend
401(k)
Exposure to ML startups

Job summary

Baseten in San Francisco builds a platform to ship AI models at scale. We’re looking for a senior product manager to own infrastructure-focused features like deployment autoscaling, traffic routing, and multi-region failover.

You’ll work with engineers and customers to define end-to-end product experiences—from backend surface area to UX observations. You’ll drive cross-team roadmaps, own end-to-end capabilities, and help define Baseten’s roadmap as the product evolves in a fast-growing AI

Qualifications

  • 8+ years in product management with infra, distributed systems, or ML serving.
  • Ability to reason about scaling, routing, failover, and cost/performance frontier.
  • End-to-end ownership from backend to UX.
  • Cross-team roadmap leadership and defining new categories.

Responsibilities

  • Own the surface for deployment scale: autoscale, traffic routing, failover, and multi-cluster workloads.
  • Define and drive the roadmap for infrastructure and product teams.
  • Ensure production inference is reliable by default with region-aware routing and multi-region deployments.
  • Build release mechanics for canary/shadow/A/B testing, warm-ups, drains, and probes.
  • Push the cost/performance frontier for serving AI at scale with measurable MTTR reductions.

Skills

PM with infra focus
Scaling & routing
End-to-end ownership
Cross-team roadmaps

Tools

Kubernetes
ML serving

Job description

Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production. We're growing quickly and recently raised our $1.5B Series F, led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to to ship AI products.

PRODUCT AT BASETEN

Product at Baseten is a nascent function. Our company today has a strong engineering culture, is heavily customer-obsessed, and moves fast. We're building the product function now, and you'd be one of the people who defines it. You'll work directly with our founders and with some of the best systems and AI engineers and you'll set the standard for what product looks like here. PMs at Baseten don't sit above engineers - you earn ownership by being technical, finding the truth in front of customers, building great cross-functional relationships, and just shipping great product experiences.

THE ROLE

Once a model is deployed, keeping it fast, reliable, and economical at scale is where production inference is won or lost. You'll own the surface that makes that happen: how deployments autoscale, how traffic is routed, how the system fails over, and how workloads scale across clusters and regions. You'll own these as products end to end - both how they work under the hood and how customers configure and observe them - and you'll help set and define the roadmap that infrastructure and product teams alike can build towards. This space is largely still evolving - think Cloud Infrastructure in mid-2000s. Your job is to make it 10x easier to reliably scale and serve AI models in production and set the market standard.

IMPACT
  • You will own how workloads scale and where they land — autoscaling to demand (up under load, down to zero when idle) and a single placement policy expressing region, compliance regime, and capacity preference, with compliance-bound workloads given right-of-way on sensitive capacity.
  • You will make production inference reliable by default — every request reaches a healthy replica, rolling deploys never drop traffic, region-aware routing with multi-region / active-active and fallback as first-class policy, and health-aware recovery from stuck or bad replicas.
  • You will build the release engine beneath safe rollouts — the traffic-shifting that powers canary/shadow/A/B, warm-ups, drain, and probes.
  • You will push the cost/performance frontier for serving AI at scale — latency, throughput, uptime, and cost-efficiency, plus a measurable decline in MTTR through self-serve incident management.
QUALIFICATIONS
  • 8+ years in product management, including deep experience with infrastructure, distributed systems, or ML serving.
  • You reason fluently about scaling, routing, failover, and the cost/performance frontier — and you earn the respect of staff engineers doing it.
  • You’ve owned capabilities end to end, backend through UX, rather than a single slice.
  • You drive cross-team roadmaps and the dependencies beneath them, and you're at your best defining a category that doesn't fully exist yet.
BONUS

Hands-on with GPU infrastructure, Kubernetes, or serving frameworks like vLLM, TensorRT-LLM, or SGLang.

NOT FOR YOU IF
  • You prefer only strategy, UX, or writing great docs over doing whatever it takes to ship great products for customers.
  • You lean toward applied AI over building platform, systems, GPUs, models, and scaling platforms and infrastructure.
  • You're not interested in the foundational, sometimes unglamorous work of making AI systems reliable and scalable at the infrastructure level.
BENEFITS
  • Competitive compensation, including meaningful equity.
  • 100% coverage of medical, dental, and vision insurance for employee and dependents
  • Flexible PTO policy including company wide Winter Break (our offices are closed from Christmas Eve to New Year's Day!)
  • Paid parental leave
  • Fertility and family-building stipend through Carrot
  • Company-facilitated 401(k)
  • Exposure to a variety of ML startups, offering unparalleled learning and networking opportunities.

At Baseten, we are committed to fostering a diverse and inclusive workplace. We provide equal employment opportunities to all employees and applicants without regard to race, color, religion, gender, sexual orientation, gender identity or expression, national origin, age, genetic information, disability, or veteran status.

We are an Equal Opportunity Employer and will consider qualified applicants with criminal histories in a manner consistent with applicable law (by example, the requirements of the San Francisco Fair Chance Ordinance, where applicable).

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Product Manager, Inference Platform
Product Manager, Inference Platform

The Consensus • New York (NY)

On-site
USD 120,000 - 160,000
Competitive compensation including equity
100% coverage of medical, dental, and vision insurance
Flexible PTO including Winter Break
+4
Product Manager, Inference Platform
Product Manager, Inference Platform

Baseten • San Francisco (CA)

On-site
USD 120,000 - 160,000
Competitive compensation with equity
100% coverage of health insurance
Generous PTO policy
+3
Product Manager - Dedicated Inference
Product Manager - Dedicated Inference

Baseten • New York (NY)

On-site
USD 200,000 - 285,000
Competitive compensation including equity
100% coverage of medical, dental, and vision insurance
Flexible PTO including company-wide Winter Break
+4
Product Manager, Developer Experience
Product Manager, Developer Experience

Baseten • San Francisco (CA)

On-site
USD 130,000 - 160,000
Competitive compensation including equity
100% coverage of medical, dental, and vision insurance
Flexible PTO policy
+4
Product Manager, Developer Experience
Product Manager, Developer Experience

The Consensus • San Francisco (CA)

On-site
USD 120,000 - 160,000
Competitive compensation including equity
100% health coverage
Flexible PTO
+2
Product Manager, Developer Experience
Product Manager, Developer Experience

Baseten • New York (NY)

On-site
USD 165,000 - 330,000
Equity
Medical, dental, vision coverage
Flexible PTO
+4
AI Inference Engineer
AI Inference Engineer

The Consensus • San Francisco (CA)

On-site
USD 130,000 - 190,000
Competitive equity and compensation
100% health, dental, vision insurance
Flexible PTO including Winter Break
+4
Product Manager, Enterprise
Product Manager, Enterprise

The Consensus • New York (NY)

On-site
USD 170,000 - 210,000
Competitive compensation
Equity
Medical, dental, vision
+4
AI Inference Engineer
AI Inference Engineer

BaseTen • New York (NY), San Francisco (CA)

Hybrid
USD 140,000 - 210,000
Equity
Healthcare coverage
Flexible PTO & Winter Break
+4
Product Manager, Enterprise
Product Manager, Enterprise

Baseten • San Francisco (CA)

On-site
USD 180,000 - 250,000
Equity
Health insurance
Flexible PTO
+3