Incident Response Analyst (Data Centre Facilities Management)

ASTREYA ASIA PACIFIC PTE. LIMITED

Penarth

On-site

GBP 25,000 - 60,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

ASTREYA ASIA PACIFIC PTE. LIMITED is seeking a Real-Time Infrastructure Monitoring specialist to oversee critical facility systems and data center operations.

The role involves 24x7 monitoring, incident triage, ticketing, and coordination with internal teams and vendors to resolve issues promptly. You will work across EPMS, BMS, DCIM and related platforms, maintain logs and reports, and contribute to continuous improvement in monitoring coverage and operational governance.

Qualifications

  • Associate degree or higher in Engineering, Information Technology, Facilities Management, or related disciplines.
  • Minimum 2 years of experience in data center operations, facility monitoring, NOC, command center, or mission-critical environments.
  • Working knowledge of Electrical, Mechanical, HVAC and cooling infrastructure, Fire detection and suppression systems, Building Management Systems (BMS), Electrical Power Monitoring Systems (EPMS), DCIM or centralized monitoring platforms.
  • Experience with incident management and escalation procedures.
  • Strong communication and coordination skills.
  • Ability to work in a 24x7 rotating shift environment.
  • Fluent in English; Chinese proficiency preferred for alarm messages and communications.

Responsibilities

  • Perform 24x7 monitoring of critical facility systems across global data centers, including Electrical power, Mechanical, HVAC, Fire detection, and Water systems.
  • Continuously monitor EPMS, BMS, DCIM, and centralized monitoring platforms.
  • Detect abnormal operating conditions and alarms; acknowledge and investigate promptly; track incidents to closure.
  • Provide first-level incident triage and technical assessment; execute escalation procedures.
  • Coordinate with internal teams, site personnel, vendors, and regional stakeholders to ensure timely issue resolution.
  • Maintain monitoring platform master data, asset records, logs, and operational documentation.
  • Analyze operational data, prepare reports, and provide recommendations to improve reliability and efficiency.

Skills

Fluent English
Excellent communication
24x7 shift work
Problem solving
Team coordination

Education

Associate Degree or higher in Engineering, IT, Facilities Management

Tools

DCIM
EPMS
BMS
CMMS
Ticketing platforms

Job description

Key Responsibilities
Real-Time Infrastructure Monitoring
  • Perform 24x7 monitoring of critical facility systems across global data centers, including: Electrical power systems, Mechanical systems, HVAC and cooling infrastructure, Fire detection and suppression systems, Water systems and supporting infrastructure
  • Continuously monitor EPMS, BMS, DCIM, and centralized monitoring platforms.
  • Detect abnormal operating conditions and alarms.
  • Acknowledge and investigate alarms promptly.
  • Track incidents and issues through to closure.
  • Identify monitoring gaps and recommend improvements to monitoring coverage.
Incident Response and Coordination
  • Provide first-level incident triage and technical assessment.
  • Respond to facility alarms and operational events in real time.
  • Execute escalation procedures according to defined protocols.
  • Coordinate with internal teams, site personnel, vendors, and regional stakeholders to ensure timely issue resolution.
  • Support major incident management activities for events such as: Utility power failures, UPS and generator events, Cooling/HVAC failures, Fire alarm activations, Water leakage events, Security and environmental alerts
  • Maintain end-to-end ownership of incidents until resolution.
Ticket Management and Change Coordination
  • Create, update, and manage event tickets within established SLA targets.
  • Process work orders and monitor completion quality.
  • Track maintenance activities and change requests.
  • Support change management processes and ensure operational compliance.
  • Maintain accurate records of facility maintenance activities and change windows.
Compliance and Operational Governance
  • Monitor and follow up on preventive maintenance activities and routine operational changes.
  • Review technical documentation submitted by vendors and service providers, including: Method of Procedure (MOP), Risk Assessment (RA), Standard Operating Procedure (SOP).
  • Ensure maintenance activities comply with operational standards and freeze-period requirements.
  • Support risk management and operational audit activities.
Monitoring Platform and Data Administration
  • Maintain monitoring platform master data and infrastructure records.
  • Ensure the accuracy, completeness, and timeliness of asset and alarm information.
  • Support platform optimization and continuous improvement initiatives.
  • Maintain facility logs, event records, and operational documentation.
Reporting and Data Analysis
  • Analyze facility operational data and identify trends or recurring issues.
  • Prepare operational reports and performance summaries.
  • Provide recommendations to improve reliability and operational efficiency.
  • Maintain records required for audit, compliance, and management reporting.
Operational Support and Continuous Improvement
  • Participate in after-hours support and emergency escalations.
  • Provide remote support for overseas data center operations when required.
  • Support centralized cross-regional operations and collaboration.
  • Contribute to process improvements and monitoring platform enhancements.
  • Perform other duties as assigned to support business continuity and operational excellence.
Minimum Qualifications
  • Associate Degree, Diploma, or higher in Engineering, Information Technology, Facilities Management, or related disciplines.
  • Minimum 2 years of experience in data center operations, facility monitoring, NOC, command center, or mission-critical environments.
  • Working knowledge of: Electrical systems, Mechanical systems, HVAC and cooling infrastructure, Fire detection and suppression systems, Building Management Systems (BMS), Electrical Power Monitoring Systems (EPMS), DCIM or centralized monitoring platforms
  • Experience working with incident management and escalation procedures.
  • Strong communication and coordination skills.
  • Ability to work in a 24x7 rotating shift environment.
  • Ability to manage multiple priorities in high-pressure situations.
  • Fluent in English.
  • Chinese language proficiency (reading, writing, and verbal communication) is preferred to support Chinese alarm messages, documentation, and communications.
Preferred Qualifications
  • Experience in: Network Operations Center (NOC), Facility Operations Center (FOC), Data Center Operations, Critical Environment Operations, Mission Critical Facilities
  • Experience supporting global or cross-regional operations.
  • Familiarity with structured incident, change, and problem management processes.
  • Understanding of data center capacity management (space, power, cooling).
  • Experience working with CMMS, DCIM, EPMS, BMS, or ticketing platforms.
  • Ability to perform root cause analysis and drive issue resolution.
Desired Competencies
  • Strong sense of ownership and urgency.
  • Excellent communication and stakeholder management skills.
  • Detail-oriented with strong documentation practices.
  • Analytical and problem-solving mindset.
  • Ability to learn quickly and adapt to changing operational environments.
  • Team-oriented with a proactive and customer-focused attitude.
Preferred Certifications

Candidates with the following certifications will have an advantage:

  • CDCP – Certified Data Centre Professional
  • CDCS – Certified Data Centre Specialist
  • FSM – Facilities Systems Management
  • Uptime Institute ATD
  • ITIL Foundation
  • DCCA or DCT certifications
  • Electrical or Mechanical engineering certifications
Shift Requirements
  • Must be willing to work a 24x7 rotating shift schedule.
  • Participate in weekends, public holidays, and on-call duty rotations when required.
  • Support emergency response activities and major incidents.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Centre Technician (24/7 Shift)
Data Centre Technician (24/7 Shift)

Allscreens Nationwide Ltd • Dundee

On-site
GBP 26,000 - 38,000
Data Center Facility Manager - United Kingdom
Data Center Facility Manager - United Kingdom

Alibaba Cloud • Greater London

On-site
GBP 65,000 - 100,000
Data Center Facility Manager / Senior Facility Engineer-London, UK
Data Center Facility Manager / Senior Facility Engineer-London, UK

Alibaba Cloud • Greater London

On-site
GBP 60,000 - 90,000
Shift Maintenance Engineer
Shift Maintenance Engineer

Pure Data Centres Group Ltd • Greater London

On-site
GBP 42,000 - 64,000
Data Centre Technician
Data Centre Technician

Intercontinental Exchange Holdings, Inc. • Basildon

On-site
GBP 29,000 - 42,000
Critical Facilities Maintenance Lead Engineer
Critical Facilities Maintenance Lead Engineer

NTT Global Data Centers • Greater London

On-site
GBP 50,000 - 70,000
Critical Facilities Maintenance Shift Lead Engineer
Critical Facilities Maintenance Shift Lead Engineer

NTT Global Data Centers • Hemel Hempstead

On-site
GBP 70,000 - 90,000
Shift Manager
Shift Manager

Bimplus • United Kingdom

On-site
GBP 45,000 - 60,000
Shift Engineer
Shift Engineer

PRS • Woking

On-site
GBP 35,000 - 52,000
Data Centre Security Ops Specialist
Data Centre Security Ops Specialist

Confidential Company • City Of London

On-site
GBP 60,000 - 80,000