Head – AI Infrastructure/Data Center Operations

LinkCxO (The CxO's Marketplace)

Chennai District

On-site

INR 4,000,000 - 6,500,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

LinkCxO (The CxO's Marketplace) seeks an experienced technology leader to head the operations of a large-scale AI Cloud and HPC infrastructure in Chennai, India. The role focuses on ensuring availability, reliability, scalability, and operational excellence across GPU, cloud, storage, and networking platforms supporting enterprise AI workloads.

You will lead 24x7 operations, drive ITIL-based incident and change management, oversee lifecycle upgrades, and build high-performing Cloud Operations

Qualifications

  • 18+ years in Cloud Infrastructure or HPC operations.
  • Proven experience leading large enterprise or hyperscale ops.
  • Strong GPU infra, Kubernetes/OpenShift, and ITIL service mgmt.

Responsibilities

  • Lead 24x7 AI cloud and HPC operations.
  • Ensure HA, performance, capacity planning across compute/storage/networking.
  • Drive ITIL aligned incident, problem, change and availability management.
  • Oversee infrastructure lifecycle, upgrades, and technology refresh.
  • Build and lead Cloud Operations, Infrastructure, and NOC teams.
  • Manage vendor and OEM relationships.
  • Drive automation and observability for continuous improvement.

Skills

GPU infrastructure
Kubernetes/OpenShift
Enterprise Linux
ITIL practices
Cloud operations
NOC leadership
Vendor management
Automation
Observability
Capacity planning

Education

Bachelor's or Master's in Engineering/CS

Tools

NVIDIA GPU Platforms
InfiniBand networking
Cloud platforms

Job description

Role Summary

We are seeking an experienced technology leader to head the operations of a large-scale AI Cloud and High-Performance Computing (HPC) infrastructure. The role will be responsible for ensuring the availability, reliability, scalability, and operational excellence of mission-critical GPU, cloud, storage, and networking platforms supporting enterprise AI workloads.

Key Responsibilities
  • Lead 24x7 operations of AI cloud, GPU infrastructure, and HPC environments.
  • Ensure high availability, performance, capacity planning, and operational excellence across compute, storage, networking, and cloud platforms.
  • Drive Incident, Problem, Change, and Availability Management in line with ITIL best practices.
  • Lead infrastructure lifecycle management, including upgrades, patching, capacity expansion, and technology refresh.
  • Build and lead high-performing Cloud Operations, Infrastructure Operations, and NOC teams.
  • Manage strategic relationships with OEMs, technology partners, and data center service providers.
  • Drive automation, observability, operational governance, and continuous service improvement.
Candidate Profile
  • 18+ years of experience in Cloud Infrastructure, Data Center Operations, AI Infrastructure, or HPC environments.
  • Proven experience managing large-scale enterprise or hyperscale infrastructure operations.
  • Strong expertise in GPU infrastructure, Kubernetes/OpenShift, Enterprise Linux, high-performance networking, enterprise storage, cloud platforms, and ITIL-based service management.
  • Experience leading large operations teams, managing vendors, and driving operational transformation.
Preferred Skills
  • AI Infrastructure
  • NVIDIA GPU Platforms
  • High-Performance Computing (HPC)
  • Cloud Operations
  • Kubernetes / OpenShift
  • InfiniBand Networking
  • Enterprise Storage
  • ITIL Service Management
  • Infrastructure Automation
  • Capacity Planning
  • Vendor Management
  • Incident & Problem Management
Education:

Bachelor's or Master's degree in Engineering, Computer Science, or a related discipline.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Compute Engineer
Senior AI Compute Engineer

Neysa • Mumbai

On-site
INR 3,500,000 - 6,000,000
Solution Architect – AI Infrastructure & Private Cloud
Solution Architect – AI Infrastructure & Private Cloud

BayOne Solutions • Bengaluru

On-site
INR 1,500,000 - 2,000,000
Infrastructure Engineer
Infrastructure Engineer

Emergys • Pune District

On-site
INR 900,000 - 1,500,000
Lead HPC Engineer
Lead HPC Engineer

Clovertex • Hyderabad

On-site
INR 2,000,000 - 3,000,000
NOC Technical Lead (L3)
NOC Technical Lead (L3)

Larsen & Toubro • Chennai District

On-site
INR 4,500,000 - 7,500,000
Solutions Architect - AI Infrastructure & Private Cloud
Solutions Architect - AI Infrastructure & Private Cloud

At Dawn Technologies • Bengaluru, Pune District

On-site
INR 2,500,000 - 3,500,000
Senior AI Infrastructure Engineer - DGX Cloud, Senior AI Infrastructure Engineer - DGX Cloud
Senior AI Infrastructure Engineer - DGX Cloud, Senior AI Infrastructure Engineer - DGX Cloud

NVIDIA • Bengaluru

On-site
INR 1,500,000 - 2,500,000
PCAI AI Factory
PCAI AI Factory

Talworx Solutions • Bengaluru

Hybrid
INR 250,000 - 450,000
Head AI Cloud Solution Architect
Head AI Cloud Solution Architect

Larsen & Toubro • Mumbai

On-site
INR 4,000,000 - 6,000,000
DevOps Engineer
DevOps Engineer

Hermes Corporate • India

On-site
INR 900,000 - 1,400,000