Senior Observability Platform Product Manager (GPU/AI)

Kforce Inc

West Palm Beach (FL)

On-site

USD 150,000 - 190,000

Full time

20 hours ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Kforce has a client in West Palm Beach, FL that is seeking a Senior Technical Product Manager Observability. You will own the end-to-end observability platform roadmap across telemetry ingestion, querying, visualization, alerting, and retention for large-scale GPU clusters and multi-tenant cloud environments.

Work closely with engineering on tradeoffs across metrics agents, data models, telemetry pipelines, and APIs.

Qualifications

  • Bachelor's degree in Computer Science, Engineering, or a related field
  • 7+ years of product management experience in cloud infrastructure, observability, monitoring, or developer platforms
  • Deep understanding of observability and monitoring systems, including metrics, logging, tracing, alerting, and telemetry pipeline architecture
  • Experience defining product strategy and roadmaps for platform or infrastructure products at scale
  • Strong technical background - ability to engage with engineering on telemetry agents, data models, query engines, retention, and distributed systems
  • Experience with GPU, AI/ML, or HPC infrastructure monitoring and the unique observability challenges of training and inference workloads
  • Track record of shipping developer- and operator-facing products with measurable impact on reliability, time-to-detect, or operational efficiency
  • Experience working across cross-functional teams (engineering, design, marketing, sales) in a fast-paced environment
  • Excellent written and verbal communication skills, with the ability to translate complex technical concepts for diverse audiences

Responsibilities

  • Own the end-to-end Observability Platform roadmap across telemetry ingestion, querying, visualization, alerting, and retention for large-scale GPU clusters and multi-tenant cloud environments
  • Define Vultr's observability strategy across bare metal, VMs, Kubernetes, and managed services, aligned to infrastructure roadmap, reliability goals, and customer experience
  • Drive the customer-facing observability surface across dashboards, APIs, telemetry pipelines, and topology-aware insights
  • Translate low-level signals across GPU, CPU, memory, storage, and network into actionable health views, alerts, and debugging workflows for customers
  • Work closely with engineering on technical tradeoffs across metrics agents, collectors, data models, telemetry pipelines, APIs, and retention architecture
  • Build products for distributed AI environments by understanding how training and inference workloads behave across nodes, clusters, schedulers, and network fabrics

Skills

Product management
Observability
Cloud infrastructure
Developer platforms
Cross-functional collaboration
Technical leadership

Education

Bachelor's degree in CS/Engineering

Tools

Telemetry agents
Query engines
Data models
Retention architecture

Job description

Kforce has a client in West Palm Beach, FL that is seeking a Senior Technical Product Manager Observability. You will own the end-to-end observability platform roadmap across telemetry ingestion, querying, visualization, alerting, and retention for large-scale GPU clusters and multi-tenant cloud environments.

Work closely with engineering on tradeoffs across metrics agents, data models, telemetry pipelines, and APIs.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Technical Product Manager Observability
Senior Technical Product Manager Observability

Kforce Inc • West Palm Beach (FL)

On-site
USD 150,000 - 190,000
Senior Observability Product Manager, GPU Fleet
Senior Observability Product Manager, GPU Fleet

Nscale • United States

On-site
USD 200,000 - 280,000
Competitive benefits package
Flexible paid time off
Parental leave
Product Manager, GPUaaS & OE Telemetry Platform
Product Manager, GPUaaS & OE Telemetry Platform

GMI Cloud • Mountain View (CA)

On-site
USD 120,000 - 160,000
Senior Observability Product Manager – AI Infra Platform
Senior Observability Product Manager – AI Infra Platform

Vultr • United States

Remote
USD 130,000 - 165,000
100% company-paid insurance premiums
401(k) with matching
Professional Development Reimbursement of $2,500
+5
Senior Observability Platform Engineer for AI/GPU Infra
Senior Observability Platform Engineer for AI/GPU Infra

Nscale • Seattle (WA)

On-site
USD 180,000 - 240,000
Senior Technical Product Manager - AI GPU Observability
Senior Technical Product Manager - AI GPU Observability

Mirantis • United States

On-site
USD 140,000 - 210,000
Product Manager - GPUaaS and OE Telemetry
Product Manager - GPUaaS and OE Telemetry

GMI Cloud • Mountain View (CA)

On-site
USD 120,000 - 160,000
Product Manager, AI Infra: GPU Clusters & Observability
Product Manager, AI Infra: GPU Clusters & Observability

Together • San Francisco (CA)

On-site
USD 175,000 - 220,000
Startup equity
Health insurance
Competitive benefits
Senior Observability Engineer: Telemetry for GPU Cloud
Senior Observability Engineer: Telemetry for GPU Cloud

Submer - Datacenters That Make Sense • United States

Remote
USD 104,000 - 174,000
Hybrid-friendly approach
International team diversity
Flexible work environment
Senior Observability Platform Engineer – AI GPU Scale
Senior Observability Platform Engineer – AI GPU Scale

Nscale • United States

On-site
USD 160,000 - 230,000
Medical, dental, vision insurance
Flexible paid time off (PTO)
Parental leave
+1