Technical Program Manager – AI Infrastructure / GPU Clusters

GMI Cloud

United States

On-site

USD 140,000 - 210,000

Full time

31 hours ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

GMI Cloud is seeking a Technical Program Manager to drive end-to-end deployment of AI GPU clusters and oversee production readiness across data center infrastructures. You will coordinate with solution architects, engineers, vendors, and contractors to ensure timely delivery of scalable GPU infrastructure.

The role focuses on hardware BOM design, server architecture, and network integration, with responsibility for cross-region collaboration and field deployment management in a fast-paced AI

Qualifications

  • 5+ years experience in technical program or infrastructure project delivery.
  • Experience with GPU cluster deployments or high-performance computing environments.
  • Familiarity with GPU server architecture and distributed computing infrastructure.
  • Experience working with Infrastructure Solution Architects to define system architecture and hardware BOM.
  • Experience managing data center hardware deployments and system integration.

Responsibilities

  • Lead the end-to-end deployment of AI GPU clusters, from infrastructure planning through production launch.
  • Drive coordination across Infrastructure Solution Architects, network engineers, hardware vendors, and data center teams.
  • Manage delivery timelines covering hardware deployment, network integration, cluster bring-up, and production readiness.

Skills

Technical Program Management
HPC infrastructure
GPU cluster deployments
Cross-region coordination

Job description

GMI Cloud is building next-generation AI infrastructure designed for large-scale GPU training and inference workloads. Our platform supports high-density GPU clusters deployed in modern data centers across multiple regions.

We are looking for a Technical Program Manager (TPM) to drive the deployment and delivery of GPU cluster infrastructure. This role will work at the intersection of AI hardware platforms, high-performance networking, and data center infrastructure, coordinating across solution architects, engineering teams, vendors, and contractors to deliver production-ready AI clusters.

Responsibilities
  • Lead the end-to-end deployment of AI GPU clusters, from infrastructure planning through production launch.
  • Drive coordination across Infrastructure Solution Architects, network engineers, hardware vendors, and data center teams.
  • Manage delivery timelines covering hardware deployment, network integration, cluster bring-up, and production readiness.
Infrastructure Architecture Collaboration
  • Work closely with Infrastructure Solution Architects (SA) to define:
  • GPU server platform selection
  • Network architecture for distributed GPU clusters
  • Storage integration and cluster infrastructure design
  • Support development of the cluster Bill of Materials (BOM) including compute, networking, storage, and supporting infrastructure components.
  • Ensure architecture decisions align with data center constraints such as power density, cooling capacity, and rack layout.
System Integration
  • Drive system integration for large-scale GPU clusters, including:
  • GPU server deployment and configuration
  • High-speed network topology implementation
  • Power and cooling readiness
  • Ensure deployments align with vendor reference architectures and validated cluster designs.
Contractor & Field Deployment Management
  • Work closely with General Contractors (GC) and system integrators to manage on-site infrastructure implementation.
  • Lead contractor onboarding, including SOW development, scope definition, and delivery milestone alignment.
  • Coordinate and oversee field deployment activities such as:
  • Structured cabling installation
  • Rack installation and equipment mounting
  • Network and power connectivity preparation
  • Hardware staging and deployment logistics
Cluster Validation & Performance Testing
  • Coordinate cluster bring-up and validation activities including:
  • Single-node GPU validation
  • Drive cluster benchmarking, stress testing, and performance verification before production release.
Operational Readiness
  • Ensure deployed GPU clusters are fully ready for production workloads by driving:
  • Hardware and network validation
  • Monitoring and telemetry integration
  • Operational documentation and runbooks
  • Handover to operations teams
Required Qualifications
  • 5+ years experience in Technical Program Management, Infrastructure Program Management, or HPC infrastructure delivery
  • Experience with GPU cluster deployments or high-performance computing environments
  • Familiarity with GPU server architecture and distributed computing infrastructure
  • Experience working with Infrastructure Solution Architects to define system architecture and hardware BOM
  • Experience managing data center hardware deployments and system integration
  • Ability to coordinate multi-vendor infrastructure projects across regions
Preferred Qualifications
  • Experience deploying large-scale AI infrastructure or GPU clusters
  • Familiarity with:
  • InfiniBand / RoCE / high-speed Ethernet networking
  • rack elevation and high-density rack deployment
  • Experience with cluster validation and performance benchmarking
  • Background as Systems Engineer, HPC Engineer, or Infrastructure Architect
  • Experience working in AI infrastructure, cloud infrastructure, or hyperscale data centers
Nice to Have
  • Experience deploying liquid-cooled GPU clusters or high-power racks
  • Experience working with NVIDIA AI infrastructure platforms
  • Familiarity with AI training environments and distributed workloads
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Infra DevOps and Backend Engineer
Infra DevOps and Backend Engineer

GMI Cloud • Mountain View (CA)

On-site
USD 160,000 - 210,000
Technical Program Manager
Technical Program Manager

GMI Cloud • Mountain View (CA)

On-site
USD 150,000 - 230,000
Site Reliability Lead
Site Reliability Lead

GMI Cloud • United States

On-site
USD 120,000 - 180,000
Site Reliability Engineer
Site Reliability Engineer

GMI Cloud • United States

On-site
USD 110,000 - 170,000
Technical Program Manager, Data Center Infrastructure Delivery
Technical Program Manager, Data Center Infrastructure Delivery

GMI Cloud • United States

On-site
USD 170,000 - 230,000
Cluster Engineer
Cluster Engineer

STN Inc • San Francisco (CA)

On-site
USD 180,000 - 240,000
Senior Solution Architect – AI / GPU Cloud
Senior Solution Architect – AI / GPU Cloud

GMI Cloud • Mountain View (CA)

On-site
USD 190,000 - 260,000
Influence product roadmap
Career growth opportunities
Work with advanced AI organizations
Artificial Intelligence Engineer
Artificial Intelligence Engineer

Calance • Costa Mesa (CA)

Hybrid
USD 180,000 - 240,000
Head of AI Data Center Infrastructure Platforms and Software
Head of AI Data Center Infrastructure Platforms and Software

Summit Group Solutions, LLC • United States

On-site
USD 150,000 - 350,000
Cluster Design
Cluster Design

Blue Signal Search • San Francisco (CA)

On-site
USD 150,000 - 230,000