Technical Product Manager, Observability – remote in the US

Mirantis

Northern (KY)

Hybrid

USD 120,000 - 160,000

Full time

8 hours ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Mirantis, an IREN company, seeks a Technical Product Manager to own observability for k0rdent AI in GPU infrastructure. Define the observability strategy, roadmap, and priorities enabling operators to monitor health, performance, and resource use across distributed AI workloads.

You will work with engineering and marketing, shaping requirements, integrations, and partnerships to ensure successful deployments and customer outcomes in a modern cloud-native environment.

Qualifications

  • 5+ years in product management or a senior technical role owning an observability product.
  • Working knowledge of Prometheus, OpenTelemetry, distributed tracing, and log aggregation.
  • Fluency in Kubernetes observability, cloud-native monitoring, or metrics and alerting pipeline architecture.

Responsibilities

  • Own the vision, roadmap, and priorities for k0rdent AI observability across the full stack: GPU compute, fabric, storage, telemetry, schedulers, inference serving, and data services.
  • Translate requirements from NeoClouds, GPU clouds, telcos, sovereign clouds, and enterprise teams into clear product direction; partner with engineering to define requirements and evaluate trade-offs.
  • Manage the observability backlog using feedback from deployments and design partners to refine priorities.
  • Track and shape our response to emerging observability standards and technologies, including OpenTelemetry, DCGM metrics, InfiniBand/RoCE fabric counters, and AI workload profiling.
  • Define integration strategies for vendor telemetry sources into a unified, operator-facing observability plane.
  • Partner with product marketing and field teams on positioning, technical briefs, and reference architectures.

Skills

Observability product ownership
Stakeholder collaboration
Kubernetes observability
OpenTelemetry
Prometheus
Distributed tracing
Log aggregation

Tools

Jaeger
Tempo
Loki
Elasticsearch/OpenSearch
NVIDIA DCGM
SLURM
vLLM/Triton (AI workloads)

Job description

Technical Product Manager, Observability – remote in the US
  • Full-time

About Mirantis

Mirantis, an IREN company, is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. By combining open source innovation with deep expertise in Kubernetes orchestration, Mirantis empowers platform engineering teams to deliver composable, production-ready developer platforms across any environment—on-premises, in the cloud, at the edge, or in sovereign data centers. As enterprises navigate the growing complexity of AI-driven workloads, Mirantis delivers the automation, GPU orchestration, and policy-driven control needed to manage infrastructure with confidence and agility. Committed to open standards and freedom from lock-in, Mirantis ensures that customers retain full control of their infrastructure strategy. https://www.mirantis.com/

Job Summary

Mirantis is looking for a Technical Product Manager to own observability for k0rdent AI, our control plane for GPU infrastructure and distributed AI workloads. In this role, you will define the observability strategy, roadmap, and feature priorities that determine how operators gain visibility into the health, performance, and resource utilization of GPU clusters running large-scale training and inference. You will shape how k0rdent AI handles everything from GPU-level metrics and distributed tracing across AI workloads, to multi-tenant log aggregation and intelligent alerting - powered by the OpenTelemetry ecosystem, and Prometheus-compatible metrics pipelines.

The ideal candidate brings strong technical fluency in observability tooling and the AI infrastructure stack. You will work directly with engineering to shape requirements, with marketing to define positioning, and with customers to help ensure their success.

Responsibilities

Own the vision, roadmap, and priorities for k0rdent AI observability across the full stack: GPU compute, east-west fabric (InfiniBand, RoCE), high-performance storage, DPU/SmartNIC telemetry, workload schedulers, inference serving, and data services

Translate requirements from NeoClouds, GPU clouds, telcos, sovereign clouds, and enterprise platform teams into clear product direction; partner with engineering to define requirements and evaluate trade-offs

Manage the observability backlog using feedback from production deployments and design partners to refine priorities

Track and shape our response to emerging observability standards and technologies, including OpenTelemetry (OTel), DCGM GPU metrics, InfiniBand/RoCE fabric counters, storage platform telemetry APIs, and AI workload profiling

Define integration strategies for vendor telemetry sources across the ecosystem - NVIDIA compute and BlueField DPUs, storage platforms (VAST, Weka, DDN), workload managers (SLURM), inference stacks, and vector and relational databases - into a unified, operator-facing observability plane

Partner with product marketing and field teams on positioning, technical briefs, and reference architectures; represent Mirantis with customers, analysts, and ecosystem partners

5+ years in product management or a senior technical role owning an observability product or operating large-scale monitoring infrastructure

Working knowledge of Prometheus, OpenTelemetry, distributed tracing (Jaeger, Tempo), and log aggregation (Loki, Elasticsearch/OpenSearch)

Fluency in Kubernetes observability, cloud-native monitoring, or metrics and alerting pipeline architecture

Ability to work directly with engineering on technical trade-offs and with field teams in competitive GPU cloud and NeoCloud deals

Strongly Preferred:

Exposure to GPU observability, including DCGM metrics, AI workload profiling and performance analysis

Familiarity with east-west fabric telemetry - InfiniBand counters, RoCEv2 congestion metrics (ECN, PFC, DCQCN), or switch-level fabric health

Experience with high-performance storage telemetry from platforms such as VAST Data, Weka, or DDN, including IOPS, latency, and throughput instrumentation at scale

Familiarity with NVIDIA BlueField DPU telemetry, SR-IOV, or offload pipeline observability

Exposure to workload-level visibility for SLURM job scheduling, inference serving stacks (vLLM, Triton, TensorRT-LLM), or data service telemetry from vector databases (Milvus, Qdrant) and relational databases in AI pipelines

Why you'll love Mirantis

Build the observabilityfoundation for the AI cloud era, working directly with leading GPU cloud operators, NeoClouds, sovereign clouds, and AI-first enterprises

Collaborate with a world-class, distributed team committed to openness and technical excellence

Shape the product narrative and influence go-to-market success

It is understood that Mirantis, Inc. may use automated decision-making technology (ADMT) for specific employment-related decisions. Opting out of ADMT use is requested for decisions about evaluation and review connected with the specific employment decision for the position applied for. You also have the right to appeal any decisions made by ADMT by sending your request to [emailprotected]

By submitting your resume, you consent to the processing and storage of your personal data in accordance with applicable data protection laws, for the purposes of considering your application for current and future job opportunities.

By clicking the link above or any third-party link within this posting, you are leaving this site and going to a third-party website where the third-party website's terms and privacy policy apply

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Technical Product Manager (Observability)
Technical Product Manager (Observability)

Mirantis • United States

On-site
USD 140,000 - 210,000
Technical Product Manager, Kubernetes Services – Remote (US)
Technical Product Manager, Kubernetes Services – Remote (US)

Mirantis • Northern (KY)

Hybrid
USD 140,000 - 190,000
Advanced AI infra
NVIDIA GPUs
Open standards
Product Manager - AI Inference & Model Serving
Product Manager - AI Inference & Model Serving

Mirantis • Austin (TX)

On-site
USD 120,000 - 160,000
Professional development and training
Customized workstation
Competitive compensation package
Remote Technical PM, AI Observability & GPU Infra
Remote Technical PM, AI Observability & GPU Infra

Mirantis • Northern (KY)

Hybrid
USD 120,000 - 160,000
Senior Technical Product Manager - AI GPU Observability
Senior Technical Product Manager - AI GPU Observability

Mirantis • United States

On-site
USD 140,000 - 210,000
Technical Product Marketer, k0rdent AI - remote in the US
Technical Product Marketer, k0rdent AI - remote in the US

Mirantis • United States

On-site
USD 90,000 - 150,000
AI Infrastructure & Platform Operations Engineer (remote in the US)
AI Infrastructure & Platform Operations Engineer (remote in the US)

Mirantis • United States

On-site
USD 120,000 - 180,000
Professional development and training
Conferences and working groups
Company outings and social events
+1
Senior Site Reliability Engineer (Golang / Kubernetes)
Senior Site Reliability Engineer (Golang / Kubernetes)

Mirantis • United States

On-site
USD 140,000 - 230,000
Competitive compensation package
Strong benefits plan
Professional development and training
+1
AI Infrastructure & Platform Operations Engineer (remote in the EU)
AI Infrastructure & Platform Operations Engineer (remote in the EU)

Mirantis, Inc. • Union (NJ)

Remote
USD 60,000 - 67,000
Remote Technical Product Marketer, k0rdent AI – remote in the US
Remote Technical Product Marketer, k0rdent AI – remote in the US

Mirantis Inc. • Town of Florida (NY)

Hybrid
USD 120,000 - 180,000
Competitive compensation package
Strong benefits plan
Open-source innovation