Technical Product Manager (Observability)

Mirantis

United States

On-site

USD 140,000 - 210,000

Full time

5 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Mirantis is seeking a Technical Product Manager to own observability for k0rdent AI, Mirantis’ control plane for GPU infrastructure and AI workloads. You will define the observability strategy, roadmap, and feature priorities to give operators visibility into health, performance, and resource utilization of GPU clusters running large-scale training and inference.

You will collaborate with engineering to shape technical requirements, with marketing to define positioning, and with customers to

Qualifications

  • Experience owning observability strategy, roadmap, and priorities for large-scale AI infrastructure.
  • Strong fluency with Kubernetes observability, cloud-native monitoring, and metrics/alerting pipelines.
  • Knowledge of OpenTelemetry, Prometheus, distributed tracing, and log aggregation.
  • Exposure to GPU observability metrics and performance profiling in AI workloads.

Responsibilities

  • Own the observability vision and backlog for k0rdent AI across GPU compute, fabric, storage, and telemetry.
  • Define integration strategies for vendor telemetry sources and partner with engineering to evaluate trade-offs.
  • Collaborate with marketing and field teams on positioning, briefs, and reference architectures.
  • Represent Mirantis with customers, analysts, and ecosystem partners to shape go-to-market success.

Skills

Observability product management
Kubernetes observability
OpenTelemetry
Prometheus
Distributed tracing
Stakeholder collaboration
AI infrastructure stack

Tools

Prometheus
OpenTelemetry
Jaeger
Tempo
Loki
Elasticsearch/OpenSearch
DCGM metrics
InfiniBand
RoCE
SLURM

Job description

  • Mirantis is looking for a Technical Product Manager to own observability for k0rdent AI, our control plane for GPU infrastructure and distributed AI workloads. In this role, you will define the observability strategy, roadmap, and feature priorities that determine how operators gain visibility into the health, performance, and resource utilization of GPU clusters running large-scale training and inference
  • You will shape how k0rdent AI handles everything from GPU-level metrics and distributed tracing across AI workloads, to multi-tenant log aggregation and intelligent alerting — powered by the OpenTelemetry ecosystem, and Prometheus-compatible metrics pipelines
  • You will work directly with engineering to shape requirements, with marketing to define positioning, and with customers to help ensure their success
  • Own the vision, roadmap, and priorities for k0rdent AI observability across the full stack: GPU compute, east-west fabric (InfiniBand, RoCE), high-performance storage, DPU/SmartNIC telemetry, workload schedulers, inference serving, and data services
  • Translate requirements from NeoClouds, GPU clouds, telcos, sovereign clouds, and enterprise platform teams into clear product direction; partner with engineering to define requirements and evaluate trade-offs
  • Manage the observability backlog using feedback from production deployments and design partners to refine priorities
  • Track and shape our response to emerging observability standards and technologies, including OpenTelemetry (OTel), DCGM GPU metrics, InfiniBand/RoCE fabric counters, storage platform telemetry APIs, and AI workload profiling
  • Define integration strategies for vendor telemetry sources across the ecosystem — NVIDIA compute and BlueField DPUs, storage platforms (VAST, Weka, DDN), workload managers (SLURM), inference stacks, and vector and relational databases — into a unified, operator-facing observability plane
  • Partner with product marketing and field teams on positioning, technical briefs, and reference architectures; represent Mirantis with customers, analysts, and ecosystem partners
  • Build the observability foundation for the AI cloud era, working directly with leading GPU cloud operators, NeoClouds, sovereign clouds, and AI-first enterprises
  • Collaborate with a world-class, distributed team committed to openness and technical excellence
  • Shape the product narrative and influence go-to-market success

The ideal candidate brings strong technical fluency in observability tooling and the AI infrastructure stackFluency in Kubernetes observability, cloud-native monitoring, or metrics and alerting pipeline architectureAbility to work directly with engineering on technical trade-offs and with field teams in competitive GPU cloud and NeoCloud deals5+ years in product management or a senior technical role owning an observability product or operating large-scale monitoring infrastructureWorking knowledge of Prometheus, OpenTelemetry, distributed tracing (Jaeger, Tempo), and log aggregation (Loki, Elasticsearch/OpenSearch)Exposure to GPU observability, including DCGM metrics, AI workload profiling and performance analysisFamiliarity with east-west fabric telemetry - InfiniBand counters, RoCEv2 congestion metrics (ECN, PFC, DCQCN), or switch-level fabric healthExperience with high-performance storage telemetry from platforms such as VAST Data, Weka, or DDN, including IOPS, latency, and throughput instrumentation at scaleFamiliarity with NVIDIA BlueField DPU telemetry, SR-IOV, or offload pipeline observabilityExposure to workload-level visibility for SLURM job scheduling, inference serving stacks (vLLM, Triton, TensorRT-LLM), or data service telemetry from vector databases (Milvus, Qdrant) and relational databases in AI pipelines

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Technical Product Manager, Observability – remote in the US
Technical Product Manager, Observability – remote in the US

Mirantis • Northern (KY)

Hybrid
USD 120,000 - 160,000
Senior Technical Product Manager - AI GPU Observability
Senior Technical Product Manager - AI GPU Observability

Mirantis • United States

On-site
USD 140,000 - 210,000
Remote Technical PM, AI Observability & GPU Infra
Remote Technical PM, AI Observability & GPU Infra

Mirantis • Northern (KY)

Hybrid
USD 120,000 - 160,000
Technical Product Manager (Kubernetes Services)
Technical Product Manager (Kubernetes Services)

Mirantis • Austin (TX)

On-site
USD 140,000 - 210,000
Technical Product Manager, Kubernetes Services – Remote (US)
Technical Product Manager, Kubernetes Services – Remote (US)

Mirantis • Northern (KY)

Hybrid
USD 140,000 - 190,000
Advanced AI infra
NVIDIA GPUs
Open standards
Senior Site Reliability Engineer (Golang / Kubernetes)
Senior Site Reliability Engineer (Golang / Kubernetes)

JobCubby • Northern (KY)

Hybrid
USD 140,000 - 180,000
Competitive compensation package
Professional development
Conference attendance
+2
Head of Solutions Architecture - AI Infrastructure
Head of Solutions Architecture - AI Infrastructure

Mirantis • United States

Remote
USD 180,000 - 280,000
Senior Site Reliability Engineer (Golang / Kubernetes)
Senior Site Reliability Engineer (Golang / Kubernetes)

Mirantis • Northern (KY)

Hybrid
USD 140,000 - 210,000
Competitive compensation
Professional development
Conference attendance
+1
Senior Site Reliability Engineer (Golang / Kubernetes)
Senior Site Reliability Engineer (Golang / Kubernetes)

Mirantis • United States

On-site
USD 140,000 - 230,000
Competitive compensation package
Strong benefits plan
Professional development and training
+1
Product Manager - GPUaaS and OE Telemetry
Product Manager - GPUaaS and OE Telemetry

GMI Cloud • Mountain View (CA)

On-site
USD 120,000 - 160,000