GPU Data Center Operations Lead

Nava

Jakarta Pusat

On-site

IDR 350,000,000 - 550,000,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Nava is seeking an Infrastructure Operations Lead to oversee operational performance for our high-density GPU data center footprint. This on-site role manages vendor performance, enforces SLAs, and coordinates MSP/Colo/provider interactions to sustain reliability across the estate.

The role requires leadership in incident response, root-cause analysis, and technical escalation for GPU/HPC environments, with a focus on maintaining power, cooling, and service boundaries.

Qualifications

  • 5+ years in data center infrastructure operations with vendor management.
  • Crisis management experience guiding incident recovery.
  • Strong communication across vendors, MSPs, Colo providers, and enterprise clients.
  • RCA expertise using frameworks like 5-Whys or Fishbone.
  • Knowledge of enterprise GPU/HPC hardware and high-density cooling.

Responsibilities

  • Oversee vendor/service governance and enforce hardware repair SLAs, ticket response times, and spare parts/RMA workflows.
  • Act as On-Site Incident Commander during high-severity outages; coordinate recovery across vendor teams and internal stakeholders.
  • Provide concise, real-time updates to executive leadership and clients during major outages.
  • Lead RCA investigations after incidents and ensure permanent corrective actions are implemented.
  • Serve as SME for complex GPU/HPC escalations and maintain operational standards.

Skills

Vendor management
SLA enforcement
Crisis management
Stakeholder communication
RCA investigations
GPU/HPC hardware
Data center operations
Power & cooling knowledge

Education

Bachelor’s degree in CS/Engineering

Job description

Nava is seeking an Infrastructure Operations Lead to oversee operational performance for our high-density GPU data center footprint. This on-site role manages vendor performance, enforces SLAs, and coordinates MSP/Colo/provider interactions to sustain reliability across the estate.

The role requires leadership in incident response, root-cause analysis, and technical escalation for GPU/HPC environments, with a focus on maintaining power, cooling, and service boundaries.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Data Centre Field Operations Engineer
Data Centre Field Operations Engineer

Nava • Jakarta Pusat

On-site
IDR 350,000,000 - 550,000,000
GPU Data Center Engineer — 24/7 Infra Fault Specialist
GPU Data Center Engineer — 24/7 Infra Fault Specialist

Lintasarta • Jakarta Pusat

On-site
IDR 279,000,000 - 502,200,000
Senior GPU Infra & OEM Escalation Engineer
Senior GPU Infra & OEM Escalation Engineer

Lintasarta • Jakarta Pusat

On-site
IDR 140,000,000 - 320,000,000
L3 Data Center Engineer
L3 Data Center Engineer

Lintasarta • Jakarta Pusat

On-site
IDR 140,000,000 - 320,000,000
Data Center Engineer (L2)
Data Center Engineer (L2)

Lintasarta • Jakarta Pusat

On-site
IDR 279,000,000 - 502,200,000
DC Network Lead
DC Network Lead

PT Leap Digital Indonesia • Jakarta Pusat

On-site
IDR 700,000,000 - 1,000,000,000
GPU Data Center Network Lead - High-Perf Fabric Architect
GPU Data Center Network Lead - High-Perf Fabric Architect

PT Leap Digital Indonesia • Jakarta Pusat

On-site
IDR 700,000,000 - 1,000,000,000
Data Center Engineer (L1)
Data Center Engineer (L1)

Lintasarta • Jakarta Pusat

On-site
IDR 89,280,000 - 156,240,000
Data Center Engineer/Lead
Data Center Engineer/Lead

Atomic Recruitment SEA • Jakarta Pusat

On-site
IDR 111,600,000 - 189,720,000
Cloud & Data Center Engineer (GPU & OpenStack)
Cloud & Data Center Engineer (GPU & OpenStack)

Lintasarta • Kota Medan ᯔᯩᯑᯉ᯲

On-site
IDR 120,000,000 - 200,000,000