Role Description
UST Job Title: Lead II - Cloud Infrastructure Services
Who We Are
At UST, we help the world s best organizations grow and succeed through transformation. Bringing together the right talent, tools, and ideas, we co-create lasting change with our clients. With over 26,000 employees in 25 countries, we build for boundless impact touching billions of lives worldwide. Visit us at UST.com.
The Opportunity
We are looking for a Technical Program Manager (TPM) to drive the deployment and delivery of GPU cluster infrastructure. This role will work at the intersection of AI hardware platforms, high-performance networking, and data center infrastructure, coordinating across solution architects, engineering teams, vendors, and contractors to deliver production-ready AI clusters.
Key Roles & Responsibilities
- Lead the end-to-end deployment of AI GPU clusters, from infrastructure planning through production launch.
- Drive coordination across Infrastructure Solution Architects, network engineers, hardware vendors, and data center teams.
- Manage delivery timelines covering hardware deployment, network integration, cluster bring-up, and production readiness.
- Work closely with Infrastructure Solution Architects (SA) to define GPU server platform selection, Network architecture for distributed GPU clusters, Storage integration and cluster infrastructure design.
- Support development of the cluster Bill of Materials (BOM) including compute, networking, storage, and supporting infrastructure components.
- Ensure architecture decisions align with data center constraints such as power density, cooling capacity, and rack layout.
- Drive system integration for large-scale GPU clusters, Rack elevation planning, GPU server deployment and configuration, High-speed network topology implementation, Power and cooling readiness.
- Ensure deployments align with vendor reference architectures and validated cluster designs.
- Work closely with General Contractors (GC) and system integrators to manage on-site infrastructure implementation.
- Lead contractor onboarding, including SOW development, scope definition, and delivery milestone alignment.
- Coordinate and oversee field deployment activities such as Structured cabling installation, Rack installation and equipment mounting, Network and power connectivity preparation, Hardware staging and deployment logistics.
- Coordinate cluster bring-up and validation activities such as Single-node GPU validation, Multi-node cluster deployment, GPU interconnect validation (P2P, RDMA).
- Drive cluster benchmarking, stress testing, and performance verification before production release.
- Ensure deployed GPU clusters are fully ready for production workloads by driving Hardware and network validation, Monitoring and telemetry integration, Operational documentation and runbooks, Handover to operations teams.
Required Qualifications
- 5+ years experience in Technical Program Management, Infrastructure Program Management, or HPC infrastructure delivery
- Experience with GPU cluster deployments or high-performance computing environments
- Familiarity with GPU server architecture and distributed computing infrastructure
- Experience working with Infrastructure Solution Architects to define system architecture and hardware BOM
- Experience managing data center hardware deployments and system integration
- Ability to coordinate multi-vendor infrastructure projects across regions
Preferred Qualifications
- Experience deploying large-scale AI infrastructure or GPU clusters
- Familiarity with InfiniBand / RoCE / high-speed Ethernet networking, GPU interconnect validation (P2P / RDMA), rack elevation and high-density rack deployment
- Experience with cluster validation and performance benchmarking, Background as Systems Engineer, HPC Engineer, or Infrastructure Architect
- Experience working in AI infrastructure, cloud infrastructure, or hyperscale data centers
- Nice to Have: Experience deploying liquid-cooled GPU clusters or high-power racks, Experience working with NVIDIA AI infrastructure platforms, Familiarity with AI training environments and distributed workloads
At UST, our culture is built on our core values:
Humility
We listen, learn, and act selflessly.
Humanity
We strive to improve the lives of others through our work.
Integrity
We honor commitments and act responsibly.
We celebrate diversity, inclusion, and innovation placing people at the center of everything we do.
Equal Employment Opportunity Statement UST is an Equal Opportunity Employer. All employment decisions are made without regard to age, race, creed, color, religion, sex, national origin, disability status, veteran status, sexual orientation, gender identity or expression, genetic information, or any other protected characteristic under applicable law.
Skills
Project Management, Process Documentation, Data Center Operations, Network Architecture, Project Management, RDMA, Storage Systems, Performance Testing, Observability, NVIDIA GPU