Product Manager - GPUaaS and OE Telemetry

GMI Cloud

Mountain View (CA)

On-site

USD 120,000 - 160,000

Full time

5 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

GMI Cloud is seeking an experienced Product Manager to lead the GPUaaS platform and Observability (OE) platform strategy, roadmap, and execution. You will define products that enable customers to consume GPU infrastructure while equipping engineering teams with telemetry across infrastructure and AI workloads.

Preferred Location: Remote, USA. The role requires collaboration with Infra Eng, SRE, and Platform teams to ensure scalability, reliability, and success metrics across GPU provisioning,

Qualifications

  • Bachelor's degree in CS/Engineering or related field.
  • 5+ years of Product Management in cloud infra or enterprise platforms.
  • Strong understanding of GPU computing and Kubernetes.
  • Experience with monitoring: Prometheus, Grafana, OpenTelemetry.
  • Ability to translate customer requirements into technical specs.
  • Experience building AI cloud or GPU cloud platforms.
  • Knowledge of multi-tenant IaaS platforms.
  • Experience with Agile/Scrum.

Responsibilities

  • Drive strategy and roadmap for the GPUaaS and OE Telemetry platform.
  • Translate customer and engineering requirements into roadmaps.
  • Define requirements, user stories, acceptance criteria, and success metrics.
  • Lead product planning, roadmap reviews, and release planning.
  • Measure adoption, impact, and business outcomes using KPIs.
  • Define capabilities for GPU provisioning, lifecycle, scheduling, APIs, multi-tenancy, billing, access management.
  • Define telemetry capabilities across GPU, compute, storage, networking, Kubernetes, Slurm, AI workloads.
  • Drive adoption of metrics, logs, traces, events, OpenTelemetry instrumentation.
  • Partner with Infra, SRE, Platform, Data Center Ops to improve observability and reliability.

Skills

Product management
Cloud infrastructure
Telemetry / observability
Kubernetes
Slurm
Linux
APIs
Agile / Scrum
Data-driven decision making

Education

Bachelor's degree in Computer Science, Engineering, or related technical field

Tools

Prometheus
Grafana
OpenTelemetry
Elasticsearch/OpenSearch

Job description

GMI Cloud is a fast-growing, AI-native infrastructure company delivering high-performance GPU compute, inference services, and infrastructure for AI agents.

Following 8x ARR growth, GMI Cloud continues to scale rapidly across the U.S. and APAC. As a Reference Platform NVIDIA Cloud Partner (NCP) and a validated leading NCP across both markets, we power production AI for leading AI-native companies including Fireworks AI, Cartesia, Reflection, and OpenRouter.

From large-scale compute to optimized inference and agentic workloads, GMI Cloud gives AI teams the infrastructure they need to build, deploy, and scale on one unified cloud.

One cloud for compute, inference, and agents.

The Role:

We are seeking an experienced Product Manager to lead the strategy, roadmap, and execution of our GPU-as-a-Service (GPUaaS) platform and Observability (OE) platform. This role is responsible for defining products that enable customers to seamlessly consume GPU infrastructure while empowering engineering and operations teams with comprehensive observability across infrastructure, platform services, and AI workloads through telemetry, monitoring, analytics, and automation.

Preferred Location: Remote, USA.

Responsibilities
  • Drive the product strategy and roadmap for the GPUaaS and OE Telemetry platform in collaboration with the Infrastructure Engineering team.
  • Translate customer and engineering requirements into prioritized product roadmaps.
  • Define product requirements, user stories, acceptance criteria, and success metrics.
  • Lead product planning, roadmap reviews, and release planning.
  • Measure product adoption, operational impact, and business outcomes using data-driven KPIs.
  • Define capabilities requirements for GPU provisioning, lifecycle management, scheduling, orchestration, self-service portal, APIs, multi-tenancy, billing, metering, quotas, and access management.
  • Define telemetry and observability capabilities requirements across GPU, compute, storage, networking, Kubernetes, Slurm, AI workloads, and supporting infrastructure.
  • Drive capabilities adoption including (but not limited to):
  • Metrics, logs, traces, and events collection
  • OpenTelemetry adoption and instrumentation
  • Real-time dashboards and visualization
  • Intelligent alerting and incident detection
  • Service health and dependency mapping
  • Distributed tracing
  • Root cause analysis
  • AI-driven anomaly detection and predictive insights
  • SLO/SLI measurement and reliability reporting
  • Partner with Infra Engineering teams, SRE, Platform Developers, Data Center Operations to improve observability, reliability, scalability and operation experience for GPUaaS platform.
Qualifications
  • Bachelor's degree in Computer Science, Engineering, or a related technical field.
  • 5+ years of Product Management experience in cloud infrastructure, AI infrastructure, or enterprise platforms.
  • Strong understanding of GPU computing (NVIDIA H100/H200/B200/B300 or equivalent), Kubernetes, Slurm, Linux, Cloud infrastructure, APIs, Telemetry and observability platforms
  • Experience with monitoring technologies such as Prometheus, Grafana, OpenTelemetry, Elasticsearch/OpenSearch, or similar.
  • Experience translating customer requirements into technical product specifications.
  • Strong analytical and communication skills.
  • Experience building AI cloud or GPU cloud platforms.
  • Knowledge of NVIDIA AI Enterprise and related software ecosystems NVSentinel, Fleet Intelligence, etc.
  • Experience with multi-tenant IaaS platforms.
  • Experience with Agile/Scrum product development.
  • Meeting every qualification is not required—if you’re excited about this role, we’d love to hear from you. We believe diverse perspectives and experiences strengthen our team.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Product Manager, GPUaaS & OE Telemetry Platform
Product Manager, GPUaaS & OE Telemetry Platform

GMI Cloud • Mountain View (CA)

On-site
USD 120,000 - 160,000
Product Manager (AI Infrastructure) - Hosting
Product Manager (AI Infrastructure) - Hosting

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 225,000 - 275,000
Stock Options
Company Bonus
Senior Product Manager – Cloud Infrastructure & GPU Platforms
Senior Product Manager – Cloud Infrastructure & GPU Platforms

Voktio • Cambridge (MA)

On-site
USD 170,000 - 210,000
Senior Product Manager – Cloud Infrastructure & GPU Platforms
Senior Product Manager – Cloud Infrastructure & GPU Platforms

MaxIT Consulting - Max Corporate Group • Massachusetts

On-site
USD 140,000 - 190,000
Senior Product Manager – GPU Products
Senior Product Manager – GPU Products

MaxIT Consulting - Max Corporate Group • Massachusetts

On-site
USD 140,000 - 210,000
Healthcare benefits
401(k) savings plan
Annual bonus or incentive
+7
Accelerated Computing GPU Product Manager
Accelerated Computing GPU Product Manager

Nvidia Corporation in • Santa Clara (CA)

On-site
USD 168,000 - 328,000
Senior Solution Architect – AI / GPU Cloud
Senior Solution Architect – AI / GPU Cloud

GMI Cloud • Mountain View (CA)

On-site
USD 190,000 - 260,000
Influence product roadmap
Career growth opportunities
Work with advanced AI organizations
Accelerated Computing Portfolio Lead – AI & HPC Services
Accelerated Computing Portfolio Lead – AI & HPC Services

MaxIT Consulting - Max Corporate Group • Massachusetts

On-site
USD 180,000 - 260,000
Healthcare benefits
401(k) plan
Annual bonus
+7
Technical Product Manager - AI Cloud Networking
Technical Product Manager - AI Cloud Networking

Summit Group Solutions, LLC • United States

On-site
USD 140,000 - 190,000
Stock options
Competitive compensation
Professional development
+1
Product Manager - SiliconMark
Product Manager - SiliconMark

Silicon Data • United States

On-site
USD 120,000 - 180,000
Competitive salary
Meaningful equity
Remote-friendly work environment
+1