Data Centre Field Operations Engineer

Nava

Metro Manila

On-site

PHP 1,200,000 - 2,000,000

Full time

2 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Nava seeks an Infrastructure Operations Lead to oversee performance across its high-density GPU data center footprint. This on-site role anchors vendor governance, incident command, and SLA adherence across MSP, Colo providers, and enterprise clients.

You will drive RCA investigations, audit vendor RCAs, and coordinate across vendor engineering teams to enforce standards and ensure rapid recovery from outages. Strong leadership and communication are essential.

Qualifications

  • Bachelor’s degree in CS, engineering, or equivalent practical experience.
  • 5+ years in data center infrastructure operations with vendor management.
  • Experience enforcing SLAs and managing high-density GPU/HPC platforms.
  • Proven crisis management and incident response leadership.
  • Strong stakeholder communication and RCA audit skills.
  • Knowledge of enterprise server platforms and data center facilities.

Responsibilities

  • Govern vendor and service delivery to enforce SLAs and repair workflows.
  • Act as On-Site Incident Commander during outages and coordinate recovery.
  • Provide real-time updates to executives and enterprise clients during major outages.
  • Lead RCA investigations and audit vendor RCAs to drive permanent fixes.
  • Serve as SME for complex GPU/HPC escalations and enforce standards.

Skills

Vendor management
SLA enforcement
Crisis management
Stakeholder communication
RCA expertise
GPU/HPC infrastructure
Data center operations
Technical escalation

Education

Bachelor’s degree in CS/Engineering

Tools

NVIDIA Blackwell GPUs
RoCE
InfiniBand

Job description

About The Role

We are seeking an Infrastructure Operations Lead to oversee operational performance across our high-density GPU data center footprint.

About The Role

We are seeking an Infrastructure Operations Lead to oversee operational performance across our high-density GPU data center footprint.

This is an on-site vendor governance and technical escalation role. You will act as our primary operational anchor—managing vendor performance and holding our Managed Service Provider (MSP), Co-Location Facility Provider, and Enterprise Customers accountable to their operational standards, service contracts, and SLAs.

Key Responsibilities
  • 360° Vendor & Service Governance: Oversee third-party service delivery to enforce hardware repair SLAs, ticket response times, and spare parts/RMA workflows. Ensure the co-location provider meets power and high-density cooling guarantees, while keeping enterprise customers within agreed operating boundaries.
  • Crisis Management & Incident Command: Act as On-Site Incident Commander during high-severity outages or infrastructure degradation. Lead recovery efforts by coordinating across vendor engineering teams and internal stakeholders.
  • Incident & Executive Communication: Provide concise, real-time updates to executive leadership and enterprise clients during major outages, translating complex technical failures into clear operational impact.
  • Root Cause Analysis (RCA) & Post-Mortems: Lead technical investigations following major incidents. Audit and challenge technical RCAs provided by vendors (MSP/Colo) to identify systemic hardware, environmental, or workflow issues, ensuring permanent corrective actions are executed.
  • Technical Escalation & Standards: Serve as the Subject Matter Expert (SME) for complex GPU/HPC hardware escalations that exceed standard vendor runbooks, and maintain ownership of operational standards.
Requirements
  • 5+ years in data center infrastructure operations, with a strong focus on third-party vendor management, SLA enforcement, and service delivery for high-density GPU/HPC platforms.
  • Crisis management: Proven ability to drive incident recovery and direct third-party vendor teams during critical data center outages under high‑pressure conditions.
  • Stakeholder Communication: Outstanding verbal and written communication skills to bridge technical vendor teams, internal stakeholders, and enterprise client representatives.
  • Root Cause Analysis (RCA) Expertise: Demonstrated capability in methodical troubleshooting, post-incident investigations, and auditing vendor‑supplied RCAs (using frameworks like 5‑Whys or Fishbone) to drive long‑term infrastructure reliability.
  • Technical & Facility Knowledge: Solid understanding of enterprise server platforms (NVIDIA Blackwell GPU architectures, high‑speed networking RoCE,infiniband) and data center facility constraints (high‑density power distribution, liquid cooling).
  • Education: Bachelor’s degree in Computer Science, Engineering, or equivalent practical experience.

Skills: oems,infrastructure,commissioning,data,operations,field execution,skills,critical infrastructure,maintenance,contractors

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Centre Field Operations Engineer
Data Centre Field Operations Engineer

Hammerjack Pty Ltd • Philippines

On-site
PHP 1,200,000 - 2,000,000
GPU Data Center Ops Lead | On-Site Incident Command
GPU Data Center Ops Lead | On-Site Incident Command

Nava • Metro Manila

On-site
PHP 1,200,000 - 2,000,000
Assistant Manager – Critical Facilities
Assistant Manager – Critical Facilities

Our Clients • Mandaluyong

On-site
PHP 1,200,000 - 2,000,000
Data Center Site Operations Lead - Global Facility Operations Center
Data Center Site Operations Lead - Global Facility Operations Center

IBEX Global Solutions (Philippines) Inc. • Mandaluyong

On-site
PHP 420,000 - 700,000
Project Manager - Data Center and Network
Project Manager - Data Center and Network

Iron Systems, Inc • Cebu City

On-site
PHP 1,000,000 - 1,700,000
GPU HPC Data Center Operations Lead
GPU HPC Data Center Operations Lead

Hammerjack Pty Ltd • Philippines

On-site
PHP 1,200,000 - 2,000,000
Data Center Infrastructure Operations Technician
Data Center Infrastructure Operations Technician

SPD Jobs, Inc. • Metro Manila

On-site
PHP 391,000 - 614,000
Senior Staff Site Reliability Engineer
Senior Staff Site Reliability Engineer

NVIDIA Corporation • Hinoba-an

Hybrid
PHP 2,000,000 - 4,500,000
Hybrid work model
Data Center Operations Support (North East / Shift up to 8PM)
Data Center Operations Support (North East / Shift up to 8PM)

HCL Singapore Pte Ltd • Santo Niño 1st

On-site
PHP 357,000 - 580,000
Data Center Technician , Data Center Operations (DCO)
Data Center Technician , Data Center Operations (DCO)

Amazon • Hinoba-an

On-site
PHP 350,000 - 550,000