AIDC Operations Leader

SenseTime 商汤科技

Hong Kong

On-site

HKD 1,800,000 - 2,800,000

Full time

46 hours ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

SenseTime 商汤科技 in Hong Kong seeks an experienced Operations Lead for AIDC facilities to drive day-to-day and strategic operations. You will own KPIs like availability, GPU utilization, PUE, and SLA performance, and establish procedures for GPU clusters, liquid cooling, and network infrastructure.

You will coordinate across SenseCore, vendors, utilities, and MEP contractors, leading onboarding and production readiness while monitoring cluster health and driving efficiency.

Qualifications

  • 10+ years of experience in data-center, HPC, cloud infrastructure or large-scale IT operations.
  • Strong experience operating mission critical infrastructure with 24x7 service requirements.
  • Experience with large scale GPU/HPC environments or AI infrastructure strongly preferred.
  • Good understanding of GPU servers, high speed networking, storage, power systems and cooling infrastructure.
  • Familiarity with InfiniBand/RoCE, Linux, Kubernetes and/or Slurm.
  • Experience managing operational KPIs, SLAs, incidents, vendors and capacity planning.
  • Strong knowledge of high density data center operations, including liquid cooling and PUE optimization.
  • Proven ability to lead cross functional technical and operational teams.
  • Strong stakeholder management skills across customers, vendors, facility operators and senior management.
  • Fluent English; Cantonese and/or Mandarin preferred.

Responsibilities

  • Lead day-to-day and strategic operations of SenseTime-supported AIDC facilities in Hong Kong.
  • Own AIDC operational KPIs including availability, GPU utilization, PUE, PFLOPS/MW, capacity utilization and SLA performance.
  • Establish operating procedures for GPU clusters, high-density racks, liquid cooling, network, storage and facility infrastructure.
  • Coordinate across SenseCore, facility operators, GPU/server vendors, network vendors, utilities, MEP contractors and managed service partners.
  • Lead customer workload onboarding, capacity allocation and production readiness processes.
  • Monitor GPU cluster health and work with engineering teams to resolve compute, network, storage and thermal bottlenecks.
  • Manage preventive maintenance, incident response, change management, escalation and disaster-recovery procedures.
  • Drive operational efficiency across power consumption, cooling, rack utilization and infrastructure capacity.
  • Support commercial teams with capacity planning, technical service definition, SLA commitments and customer proposals.
  • Develop standard operating models that can be replicated across future Hong Kong, Macau and regional AIDC projects.

Skills

Data center operations
HPC environments
Cloud infrastructure
GPU/HPC management
Kubernetes
Slurm
InfiniBand/RoCE
Linux

Tools

InfiniBand/RoCE
Linux
Kubernetes
Slurm
DCIM/BMS/EPMS
400G/800G networks

Job description

  • Lead day-to-day and strategic operations of SenseTime-supported AIDC facilities in Hong Kong.
  • Own AIDC operational KPIs including availability, GPU utilization, PUE, PFLOPS/MW, capacity utilization and SLA performance.
  • Establish operating procedures for GPU clusters, high-density racks, liquid cooling, network, storage and facility infrastructure.
  • Coordinate across SenseCore, facility operators, GPU/server vendors, network vendors, utilities, MEP contractors and managed service partners.
  • Lead customer workload onboarding, capacity allocation and production readiness processes.
  • Monitor GPU cluster health and work with engineering teams to resolve compute, network, storage and thermal bottlenecks.
  • Manage preventive maintenance, incident response, change management, escalation and disaster-recovery procedures.
  • Drive operational efficiency across power consumption, cooling, rack utilization and infrastructure capacity.
  • Support commercial teams with capacity planning, technical service definition, SLA commitments and customer proposals.
  • Develop standard operating models that can be replicated across future Hong Kong, Macau and regional AIDC projects.

This operating capability is particularly important in AIDC strategy includes large scale facilities where the partner is expected to provide cloud provisioning, AI software platforms and operational know how for building and running AIDCs.

Must-Haves
  • 10+ years of experience in data-center, HPC, cloud infrastructure or large-scale IT operations.
  • Strong experience operating mission critical infrastructure with 24x7 service requirements.
  • Experience with large scale GPU/HPC environments or AI infrastructure strongly preferred.
  • Good understanding of GPU servers, high speed networking, storage, power systems and cooling infrastructure.
  • Familiarity with InfiniBand/RoCE, Linux, Kubernetes and/or Slurm.
  • Experience managing operational KPIs, SLAs, incidents, vendors and capacity planning.
  • Strong knowledge of high density data center operations, including liquid cooling and PUE optimization.
  • Proven ability to lead cross functional technical and operational teams.
  • Strong stakeholder management skills across customers, vendors, facility operators and senior management.
  • Fluent English; Cantonese and/or Mandarin preferred.
Nice-to-Haves

Experience with NVIDIA DGX/HGX, Huawei Ascend, large GPU clusters, 100kW+ racks, Direct Liquid Cooling, DCIM/BMS/EPMS, 400G/800G networks, AI cloud operations, HPC benchmarking or >10 MW data center environments.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AIDC Operations Leader: GPU/HPC Data Center, 24x7
AIDC Operations Leader: GPU/HPC Data Center, 24x7

SenseTime 商汤科技 • Hong Kong

On-site
HKD 1,800,000 - 2,800,000
Head of Data Centre Operations
Head of Data Centre Operations

Rise Associates Asia Limited • Hong Kong

On-site
HKD 1,800,000 - 2,400,000
Data Centre Facilities Engineer
Data Centre Facilities Engineer

Janestreet • Hong Kong

On-site
HKD 600,000 - 900,000
Head of AI Infrastructure & Data Center Operations
Head of AI Infrastructure & Data Center Operations

Michael Page International (HK) Ltd • Hong Kong

On-site
HKD 1,800,000 - 3,200,000
Data Centre Facilities Engineer
Data Centre Facilities Engineer

Jane Street • Hong Kong

On-site
HKD 600,000 - 1,000,000
Senior Data Center & AI Compute Programs Leader
Senior Data Center & AI Compute Programs Leader

Leadingnation • Hong Kong

On-site
HKD 900,000 - 1,300,000
Competitive annual leave
MPF top-up
Medical benefits from Day 1
+3
AI Infrastructure Manager
AI Infrastructure Manager

China Mobile International Limited • Hong Kong

On-site
HKD 800,000 - 1,200,000
Data Centre Operations Lead for HPC
Data Centre Operations Lead for HPC

Hong Kong Cyberport Management Co Ltd • Hong Kong

On-site
HKD 800,000 - 1,100,000
Head of AI Infrastructure & Data Center Operations
Head of AI Infrastructure & Data Center Operations

Michael Page International (Hong Kong) Limited • Hong Kong

On-site
HKD 1,800,000 - 2,400,000
Platform Manager (Working location: Shenzhen)
Platform Manager (Working location: Shenzhen)

Goldhorse Securities • Hong Kong

On-site
HKD 780,000 - 1,100,000