Data Centre Field Operations Engineer

Nava

Jakarta Pusat

On-site

IDR 420,000,000 - 620,000,000

Full time

3 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Nava is seeking an Infrastructure Operations Lead to oversee vendor governance and incident management for a high-density GPU data center footprint in Jakarta. You will be the on-site anchor, ensuring MSPs and co-location providers meet contracts and SLAs while coordinating with enterprise clients.

The role requires 5+ years in data center infra, strong crisis management skills, and the ability to translate complex failures into operational impact for executives.

Qualifications

  • 5+ years in data center infrastructure operations with focus on vendor management
  • Proven crisis management and incident recovery leadership
  • Strong communication to bridge vendor teams, internal stakeholders and clients
  • Experience auditing vendor RCAs and driving permanent corrective actions
  • Solid understanding of high-density GPU/HPC hardware and data center facilities

Responsibilities

  • Oversee vendor/service governance to enforce SLAs, repair and spare parts workflows
  • Act as On-Site Incident Commander during outages and coordinate recovery
  • Provide concise real-time updates to leadership and clients during outages
  • Lead RCA investigations and audit vendor-supplied RCAs for systemic issues
  • Serve as SME for escalations beyond standard vendor runbooks and maintain standards

Skills

Vendor management
SLA enforcement
Incident response
Stakeholder communication
RCA

Education

Bachelor’s degree in Computer Science or Engineering

Tools

RoCE
Infiniband

Job description

About The Role

We are seeking an Infrastructure Operations Lead to oversee operational performance across our high-density GPU data center footprint.

This is an on-site vendor governance and technical escalation role. You will act as our primary operational anchor—managing vendor performance and holding our Managed Service Provider (MSP), Co-Location Facility Provider, and Enterprise Customers accountable to their operational standards, service contracts, and SLAs.

Key Responsibilities
  • 360° Vendor & Service Governance: Oversee third-party service delivery to enforce hardware repair SLAs, ticket response times, and spare parts/RMA workflows. Ensure the co-location provider meets power and high-density cooling guarantees, while keeping enterprise customers within agreed operating boundaries.
  • Crisis Management & Incident Command: Act as On-Site Incident Commander during high-severity outages or infrastructure degradation. Lead recovery efforts by coordinating across vendor engineering teams and internal stakeholders.
  • Incident & Executive Communication: Provide concise, real-time updates to executive leadership and enterprise clients during major outages, translating complex technical failures into clear operational impact.
  • Root Cause Analysis (RCA) & Post-Mortems: Lead technical investigations following major incidents. Audit and challenge technical RCAs provided by vendors (MSP/Colo) to identify systemic hardware, environmental, or workflow issues, ensuring permanent corrective actions are executed.
  • Technical Escalation & Standards: Serve as the Subject Matter Expert (SME) for complex GPU/HPC hardware escalations that exceed standard vendor runbooks, and maintain ownership of operational standards.
Requirements
  • 5+ years in data center infrastructure operations, with a strong focus on third-party vendor management, SLA enforcement, and service delivery for high-density GPU/HPC platforms.
  • Crisis management: Proven ability to drive incident recovery and direct third-party vendor teams during critical data center outages under high-pressure conditions.
  • Stakeholder Communication: Outstanding verbal and written communication skills to bridge technical vendor teams, internal stakeholders, and enterprise client representatives.
  • Root Cause Analysis (RCA) Expertise: Demonstrated capability in methodical troubleshooting, post-incident investigations, and auditing vendor-supplied RCAs (using frameworks like 5-Whys or Fishbone) to drive long-term infrastructure reliability.
  • Technical & Facility Knowledge: Solid understanding of enterprise server platforms (NVIDIA Blackwell GPU architectures, high-speed networking RoCE,infiniband) and data center facility constraints (high-density power distribution, liquid cooling).
  • Education: Bachelor’s degree in Computer Science, Engineering, or equivalent practical experience.

Skills: infrastructure,oems,critical infrastructure,commissioning,skills,maintenance,field execution,operations,contractors,data

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Center Engineer (L2)
Data Center Engineer (L2)

Lintasarta • Jakarta Pusat

On-site
IDR 279,000,000 - 502,200,000
Data Center Field Ops Lead - GPU/HPC & Vendor Governance
Data Center Field Ops Lead - GPU/HPC & Vendor Governance

Nava • Jakarta Pusat

On-site
IDR 420,000,000 - 620,000,000
Datacenter Operations Manager
Datacenter Operations Manager

DAMAC Digital • Jakarta Pusat

On-site
IDR 400,000,000 - 700,000,000
Project Manager (Data Center)
Project Manager (Data Center)

amIT Global Solutions Sdn Bhd • Indonesia

On-site
IDR 2,100,105,000 - 3,150,158,000
Data Center Engineering Operations, AWS DCEO
Data Center Engineering Operations, AWS DCEO

Amazon • Jakarta Pusat

On-site
IDR 600,000,000 - 800,000,000
Data Center Engineer
Data Center Engineer

Esha Parama Technology • Daerah Khusus Ibukota Jakarta

On-site
IDR 180,000,000 - 280,000,000
Data Center Operation Manager
Data Center Operation Manager

Michael Page • Daerah Khusus Ibukota Jakarta

On-site
IDR 900,000,000 - 1,500,000,000
GPU Data Center Engineer — 24/7 Infra Fault Specialist
GPU Data Center Engineer — 24/7 Infra Fault Specialist

Lintasarta • Jakarta Pusat

On-site
IDR 279,000,000 - 502,200,000
Senior Server Administrator
Senior Server Administrator

Digital Edge Data Center • Jakarta Pusat

On-site
IDR 420,000,000 - 620,000,000
Asst Manager Operations Support
Asst Manager Operations Support

Digital Edge Data Center • Indonesia

On-site
IDR 420,000,000 - 640,000,000