Job Description
We are looking for a proactive and detail-oriented Data Center NOC (Network Operations Center) Engineer to join our Command Center operations. This role is critical to the real-time monitoring, incident escalation, and hands-on support of our data center infrastructure, including power, cooling, environmental, and network systems.
The engineer will also contribute to data-driven analysis, support the development of operational policies and procedures, and perform other assigned tasks as directed by management. You will be a key enabler in maintaining the reliability, security, and efficiency of mission-critical systems in a 24x7 environment.
Key Responsibilities
Infrastructure Monitoring & Surveillance
- Continuously monitor facility and IT systems via BMS, DCIM, EMS, and NMS platforms.
- Track and log key metrics such as power usage, temperature, humidity, network traffic, and system health.
- Maintain detailed event logs, dashboards, and audit trails to support compliance and reporting.
Incident Management & Escalation
- Detect, validate, and log infrastructure or network anomalies.
- Perform basic diagnostic checks to verify alarm legitimacy.
- Escalate incidents promptly to Engineering, Network, or Vendor teams following SOPs and SLAs.
- Assist with post-incident investigations, timelines, and Root Cause Analysis (RCA) documentation.
Data Analysis & Insights
- Analyze operational data from monitoring systems to identify trends, anomalies, and performance bottlenecks.
- Generate reports and dashboards that inform capacity planning, energy efficiency, and system health.
- Provide recommendations to improve monitoring effectiveness, system reliability, and operational readiness.
- Collaborate with engineering teams to refine alarm thresholds, monitoring logic, and alert policies based on historical data.
Remote Hands Support
- Execute physical tasks under directions, including:
- Power cycling equipment
- Cable tracing, labeling, and organization
- Visual inspections and validation of hardware conditions
- Perform simple logical actions (e.g., console resets) only when clearly instructed.
- Ensure all interventions are logged and compliant with audit requirements.
Rack & Capacity Management Support
- Maintain up-to-date rack layout documentation, power budgets, and port usage maps.
- Support cabling organization and structured layout in collaboration with Network teams.
- Identify and escalated capacity constraints or environmental risks (power, cooling, port utilization).
Reporting, Documentation & Policy Support
- Prepare shift turnover reports, highlighting incidents, alerts, and ongoing issues.
- Update incident and ticketing platforms with complete and timely information.
- Contribute to weekly/monthly Command Center performance and capacity reports.
- Assist in the development or revision of operational procedures, SOPs, and escalation workflows, under the guidance of superiors or compliance requirements.
Compliance, Safety & Standards
- Follow all SOPs, work instructions, and escalation paths based on ITIL or ISO frameworks.
- Ensure operations align with ISO 27001 (Information Security), ISO 20000 (Service Management), and TIA-942(Data Center Standards).
- Maintain strict adherence to safety, security, and access control policies.
Teamwork & Communication
- Work closely with Data Center Engineers, Network Engineers, Facilities, and Security teams.
- Ensure smooth shift handovers with thorough documentation and verbal updates.
- Support internal audits, testing drills, and compliance reviews as needed.
- Foster a disciplined, reliable, and collaborative NOC culture.
Other Duties
- Perform additional tasks and responsibilities as directed by superiors or management, including support for new tools, workflows, or special projects.
Qualifications
Education
- Bachelor’s Degree in IT, Computer Science, Electrical/Electronic Engineering OR
- Technical diploma with relevant industry certifications and experience.
Experience
- Entry-Level: 1–2 years in NOC monitoring, data center operations, or IT support.
- Mid-Level: 3–5 years in mission-critical infrastructure monitoring, with escalation and analysis responsibilities.
Technical Skills
- Familiar with enterprise monitoring tools (BMS, DCIM, EMS, NMS).
- Understanding of data center infrastructure (UPS, CRAC, PDUs) and basic networking (IP, cabling, ports).
- Hands-on knowledge of IT support or Remote Hands procedures.
- Comfortable using ticketing systems (e.g., ServiceNow, Remedy, Jira).
- Basic data analysis using Excel, dashboards, or monitoring exports.
Soft Skills
- High attention to detail and operational discipline.
- Strong written and verbal communication for logs, coordination, and escalation.
- Calm and responsive during critical incidents or high-pressure situations.
- Flexible and adaptable to 24x7 rotating shift environments.
- Team player with a service-oriented mindset.
Preferred Certifications (Optional)
- ITIL Foundation
- CompTIA Network+ / A+
- Data Center Certified Professional (CDCP) or similar
What We Offer
- Exposure to cutting-edge data center technologies and global standards
- High-impact role in a mission-critical operations environment
- Collaborative, structured, and process-driven work culture
- Opportunities to learn, grow, and influence operational excellence