- Lead day-to-day and strategic operations of SenseTime-supported AIDC facilities in Hong Kong.
- Own AIDC operational KPIs including availability, GPU utilization, PUE, PFLOPS/MW, capacity utilization and SLA performance.
- Establish operating procedures for GPU clusters, high-density racks, liquid cooling, network, storage and facility infrastructure.
- Coordinate across SenseCore, facility operators, GPU/server vendors, network vendors, utilities, MEP contractors and managed service partners.
- Lead customer workload onboarding, capacity allocation and production readiness processes.
- Monitor GPU cluster health and work with engineering teams to resolve compute, network, storage and thermal bottlenecks.
- Manage preventive maintenance, incident response, change management, escalation and disaster-recovery procedures.
- Drive operational efficiency across power consumption, cooling, rack utilization and infrastructure capacity.
- Support commercial teams with capacity planning, technical service definition, SLA commitments and customer proposals.
- Develop standard operating models that can be replicated across future Hong Kong, Macau and regional AIDC projects.
This operating capability is particularly important in AIDC strategy includes large scale facilities where the partner is expected to provide cloud provisioning, AI software platforms and operational know how for building and running AIDCs.
Must-Haves
- 10+ years of experience in data-center, HPC, cloud infrastructure or large-scale IT operations.
- Strong experience operating mission critical infrastructure with 24x7 service requirements.
- Experience with large scale GPU/HPC environments or AI infrastructure strongly preferred.
- Good understanding of GPU servers, high speed networking, storage, power systems and cooling infrastructure.
- Familiarity with InfiniBand/RoCE, Linux, Kubernetes and/or Slurm.
- Experience managing operational KPIs, SLAs, incidents, vendors and capacity planning.
- Strong knowledge of high density data center operations, including liquid cooling and PUE optimization.
- Proven ability to lead cross functional technical and operational teams.
- Strong stakeholder management skills across customers, vendors, facility operators and senior management.
- Fluent English; Cantonese and/or Mandarin preferred.
Nice-to-Haves
Experience with NVIDIA DGX/HGX, Huawei Ascend, large GPU clusters, 100kW+ racks, Direct Liquid Cooling, DCIM/BMS/EPMS, 400G/800G networks, AI cloud operations, HPC benchmarking or >10 MW data center environments.