An application made for this job — a tailored resume and cover letter that speak straight to the posting.
Mirantis is seeking a Technical Product Manager to own observability for k0rdent AI, Mirantis’ control plane for GPU infrastructure and AI workloads. You will define the observability strategy, roadmap, and feature priorities to give operators visibility into health, performance, and resource utilization of GPU clusters running large-scale training and inference.
You will collaborate with engineering to shape technical requirements, with marketing to define positioning, and with customers to
The ideal candidate brings strong technical fluency in observability tooling and the AI infrastructure stackFluency in Kubernetes observability, cloud-native monitoring, or metrics and alerting pipeline architectureAbility to work directly with engineering on technical trade-offs and with field teams in competitive GPU cloud and NeoCloud deals5+ years in product management or a senior technical role owning an observability product or operating large-scale monitoring infrastructureWorking knowledge of Prometheus, OpenTelemetry, distributed tracing (Jaeger, Tempo), and log aggregation (Loki, Elasticsearch/OpenSearch)Exposure to GPU observability, including DCGM metrics, AI workload profiling and performance analysisFamiliarity with east-west fabric telemetry - InfiniBand counters, RoCEv2 congestion metrics (ECN, PFC, DCQCN), or switch-level fabric healthExperience with high-performance storage telemetry from platforms such as VAST Data, Weka, or DDN, including IOPS, latency, and throughput instrumentation at scaleFamiliarity with NVIDIA BlueField DPU telemetry, SR-IOV, or offload pipeline observabilityExposure to workload-level visibility for SLURM job scheduling, inference serving stacks (vLLM, Triton, TensorRT-LLM), or data service telemetry from vector databases (Milvus, Qdrant) and relational databases in AI pipelines