GMI Cloud is building next-generation AI infrastructure designed for large-scale GPU training and inference workloads. Our platform supports high-density GPU clusters deployed in modern data centers across multiple regions.
We are looking for a Technical Program Manager (TPM) to drive the deployment and delivery of GPU cluster infrastructure. This role will work at the intersection of AI hardware platforms, high-performance networking, and data center infrastructure, coordinating across solution architects, engineering teams, vendors, and contractors to deliver production-ready AI clusters.
Responsibilities
- Lead the end-to-end deployment of AI GPU clusters, from infrastructure planning through production launch.
- Drive coordination across Infrastructure Solution Architects, network engineers, hardware vendors, and data center teams.
- Manage delivery timelines covering hardware deployment, network integration, cluster bring-up, and production readiness.
Infrastructure Architecture Collaboration
- Work closely with Infrastructure Solution Architects (SA) to define:
- GPU server platform selection
- Network architecture for distributed GPU clusters
- Storage integration and cluster infrastructure design
- Support development of the cluster Bill of Materials (BOM) including compute, networking, storage, and supporting infrastructure components.
- Ensure architecture decisions align with data center constraints such as power density, cooling capacity, and rack layout.
System Integration
- Drive system integration for large-scale GPU clusters, including:
- GPU server deployment and configuration
- High-speed network topology implementation
- Power and cooling readiness
- Ensure deployments align with vendor reference architectures and validated cluster designs.
Contractor & Field Deployment Management
- Work closely with General Contractors (GC) and system integrators to manage on-site infrastructure implementation.
- Lead contractor onboarding, including SOW development, scope definition, and delivery milestone alignment.
- Coordinate and oversee field deployment activities such as:
- Structured cabling installation
- Rack installation and equipment mounting
- Network and power connectivity preparation
- Hardware staging and deployment logistics
Cluster Validation & Performance Testing
- Coordinate cluster bring-up and validation activities including:
- Single-node GPU validation
- Drive cluster benchmarking, stress testing, and performance verification before production release.
Operational Readiness
- Ensure deployed GPU clusters are fully ready for production workloads by driving:
- Hardware and network validation
- Monitoring and telemetry integration
- Operational documentation and runbooks
- Handover to operations teams
Required Qualifications
- 5+ years experience in Technical Program Management, Infrastructure Program Management, or HPC infrastructure delivery
- Experience with GPU cluster deployments or high-performance computing environments
- Familiarity with GPU server architecture and distributed computing infrastructure
- Experience working with Infrastructure Solution Architects to define system architecture and hardware BOM
- Experience managing data center hardware deployments and system integration
- Ability to coordinate multi-vendor infrastructure projects across regions
Preferred Qualifications
- Experience deploying large-scale AI infrastructure or GPU clusters
- Familiarity with:
- InfiniBand / RoCE / high-speed Ethernet networking
- rack elevation and high-density rack deployment
- Experience with cluster validation and performance benchmarking
- Background as Systems Engineer, HPC Engineer, or Infrastructure Architect
- Experience working in AI infrastructure, cloud infrastructure, or hyperscale data centers
Nice to Have
- Experience deploying liquid-cooled GPU clusters or high-power racks
- Experience working with NVIDIA AI infrastructure platforms
- Familiarity with AI training environments and distributed workloads