BB Wave Inc. is seeking a dedicated operations support specialist in Metro Manila, Taguig. The role involves 24/7 monitoring of production systems and dashboards, proactive management of system alerts, and routine inspections. Candidates must possess a diploma in Computer Science or related fields, and familiarity with monitoring tools like Zabbix and Prometheus is ideal. The position also requires strong log analysis skills and an understanding of networking concepts, along with a willingness to work in rotating shifts.
Qualifications
13 years of experience in Operations, NOC, or similar roles.
Basic understanding of DNS and load balancing concepts.
Familiar with ITSM/ticketing systems.
Responsibilities
Provide 24/7 monitoring support for production systems.
Respond to alerts and perform initial troubleshooting.
Create and maintain operational manuals and SOP documentation.
Skills
Linux commands
Log analysis
TCP/IP
Monitoring platforms
Shell scripting
Communication skills
Education
Diploma in Computer Science or related fields
Tools
Zabbix
Prometheus
Grafana
Docker
Kubernetes
Job description
Responsibilities
Provide 24/7 monitoring support for production systems and dashboards (Zabbix, Prometheus, Grafana, Huawei Cloud Monitoring).
Continuously monitor key business metrics (QPS, error rate, response time, success rate) and infrastructure health indicators (CPU, memory, disk, and network usage).
Perform routine inspections of online systems, data centre environments, and network connectivity.
Maintain shift logs and inspection reports according to operational standards.
Respond immediately to alerts received via phone, SMS, Microsoft Teams, etc., and perform initial troubleshooting based on SOPs.
Identify alert severity levels and execute corresponding response procedures.
Independently resolve common issues such as restarting services, cleaning up disk space, resource scaling, and traffic switching.
Escalate unresolved issues to L2 support or development teams following the escalation process, and track incidents until resolution.
Continuously optimised alert rules to reduce false positives, missed alerts, and alert storms.
Handle daily operational requests through the ticketing system, including account creation, permission requests, resource allocation, firewall policy changes, and domain/SSL certificate applications.
Support deployment and release processes, including version releases, rollbacks, and configuration changes.
Assist SRE, DBA, network, and security teams with routine tasks such as backup verification, slow query checks, and vulnerability scan follow‑ups.
Maintain CMDB asset information to ensure accuracy of servers, IPs, applications, and ownership records.
Act as the first responder during incidents, coordinating communication channels, incident bridges, and status updates.
Participate in incident review and post‑mortem meetings, helping document timelines and improve SOPs and knowledge base articles.
Create and maintain operational manuals, emergency response plans, and SOP documentation.
Prepare weekly/monthly operational reports, including top alerts, incident counts, resolution times, and other operational metrics.
Job Requirements
Diploma or above in Computer Science, Communications, Networking, or related fields.
13 years of experience in Operations, NOC, or system monitoring roles (fresh graduates highly encouraged).
Familiar with common Linux commands and able to independently perform log analysis, process troubleshooting, network connectivity testing, and disk and memory issue investigation.
Basic understanding of TCP/IP, HTTP, DNS, and load balancing concepts.
Ability to interpret outputs from tools such as ping, telnet, curl, and tcpdump.
Familiar with at least one monitoring platform: Zabbix, Prometheus + Grafana, Nagios, Alibaba Cloud / Tencent Cloud Monitoring.
Basic operational knowledge of services and middleware: Nginx, Tomcat, Redis, MySQL (checking status, restarting services, reviewing logs).
Basic shell‑scripting skills; able to read and make small modifications to scripts.
Strong sense of responsibility, attention to detail, and ability to work under pressure.
Willingness to work rotating shifts, including night shifts and holiday coverage.
Good communication skills and ability to stay calm and organized during incidents.
Preferred Qualifications
Knowledge of ITIL processes or related certifications.
Familiar with ITSM/ticketing systems such as Jira Service Management, ServiceNow, ONES, and in‑house ticketing systems.
Experience with basic Docker/Kubernetes operations: checking pod status, viewing logs, restarting services.
Experience supporting large‑scale systems (e‑commerce, finance, gaming, live streaming, etc.).
Ability to write small automation tools using Python.
Relevant certifications such as RHCSA, HCIA, Huawei/Cisco networking certifications, and cloud provider certifications.