Remote TPM: AI/GPU Cluster Deployment Lead

5C Group

Springfield (OH)

On-site

USD 155,000 - 175,000

Full time

5 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

5C DATA CENTERS is seeking a Technical Project Manager to lead planning, coordination, and execution of large-scale GPU cluster deployments for AI and HPC initiatives. You will partner with engineering, datacenter operations, and external vendors to manage schedules, budgets, and readiness across compute, networking, storage, power, and cooling.

The role requires strong program leadership in hyperscale infrastructure and AI deployment lifecycles, plus the ability to drive repeatable deployment

Qualifications

  • Lead end-to-end project management for large-scale AI/GPU cluster deployments.
  • Coordinate with engineering teams and managers for project tracking and milestones.
  • Align with facilities engineering and datacenter operations for power and cooling readiness.

Responsibilities

  • Lead end-to-end project management for large-scale AI/GPU cluster deployments, including multi-rack GPU compute platforms (NVIDIA DGX and similar), InfiniBand and Ethernet GPU fabrics, and high-performance storage environments (e.g., VAST Data).
  • Partner with engineering teams and managers to establish and drive consistent project tracking, milestone reporting, and status updates across the team and systems (e.g., Jira, Confluence).
  • Coordinate procurement, rack-and-stack sequencing, cabling schedules, network deployment timelines, burn-in testing, cluster validation, and operational handoff.
  • Coordinate with facilities engineering and datacenter operations to align MEP readiness (power distribution, cooling capacity, floor layout, containment) with deployment schedules.
  • Communicate project status, risks, milestones, and dependencies to stakeholders at all levels including executive leadership.
  • Prepare and present regular program reviews, steering committee updates, and ad-hoc project analyses.
  • Maintain centralized project documentation and ensure consistent, accurate data.
  • Contribute to improving and documenting repeatable deployment methodologies, scalable operational standards, and project management best practices.
  • Track deployment KPIs (schedule variance, budget adherence, quality metrics) and drive continuous improvement through data-driven retrospectives.

Skills

Project management
AI infrastructure
GUPs deployments
Jira/Confluence
Vendor coordination
MEP coordination

Tools

Ansible
Python
Shell
SQL
DCIM
BMS
NVLink/NVSwitch

Job description

5C DATA CENTERS is seeking a Technical Project Manager to lead planning, coordination, and execution of large-scale GPU cluster deployments for AI and HPC initiatives. You will partner with engineering, datacenter operations, and external vendors to manage schedules, budgets, and readiness across compute, networking, storage, power, and cooling.

The role requires strong program leadership in hyperscale infrastructure and AI deployment lifecycles, plus the ability to drive repeatable deployment

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Remote TPM for AI/GPU Infrastructure Deployments
Remote TPM for AI/GPU Infrastructure Deployments

5C Group • United States

On-site
USD 155,000 - 175,000
Staff TPM: AI Infra Deployments & GPU Cloud
Staff TPM: AI Infra Deployments & GPU Cloud

Crusoe • San Francisco (CA)

On-site
USD 200,000 - 240,000
Equity
Restricted Stock Units
Paid time off & holidays
+12
AI Cluster Infra TPM — Hardware-Scale AI Programs
AI Cluster Infra TPM — Hardware-Scale AI Programs

AMD • United States

On-site
USD 140,000 - 180,000
Staff TPM: AI Infrastructure Deployments & GPU Cloud
Staff TPM: AI Infrastructure Deployments & GPU Cloud

Crusoe • Bellevue (WA)

On-site
USD 200,000 - 240,000
Equity
Paid time off
Health insurance
+2
AI Cluster TPM: Hyperscale AI Infrastructure
AI Cluster TPM: Hyperscale AI Infrastructure

AMD • Austin (TX)

On-site
USD 140,000 - 210,000
AMD benefits
Senior AI Infrastructure TPM — Scale Global GPU Cloud
Senior AI Infrastructure TPM — Scale Global GPU Cloud

Together • San Francisco (CA)

Hybrid
USD 225,000 - 265,000
Health insurance
Equity
Data Center Deployment Lead — GPU Clusters
Data Center Deployment Lead — GPU Clusters

Introl • New York (NY)

On-site
USD 98,000 - 118,000
Paid Family Leave
Unlimited PTO
Sick Time Off
+5
Staff TPM, AI Infrastructure Deployment & Scale
Staff TPM, AI Infrastructure Deployment & Scale

Crusoe • United States

On-site
USD 200,000 - 240,000
Equity packages
Paid time off
Health insurance
+2
Cloud Infrastructure TPM — Data Center & GPU Deployments
Cloud Infrastructure TPM — Data Center & GPU Deployments

Nebius • United States

Hybrid
USD 115,000 - 225,000
100% company-paid health insurance
401(k) Plan with company match
20 weeks parental leave
+2
Technical Program Manager - AI Hardware & Clusters
Technical Program Manager - AI Hardware & Clusters

Advanced Micro Devices • Austin (TX)

On-site
USD 140,000 - 190,000