Get more replies from employers
Send a job-specific resume in minutes.
MyRepublic Broadband Pte Ltd in Singapore is seeking a hands-on Network Operations and Incident Management professional to monitor and maintain the health of our network, voice, cloud, and managed service platforms. The role emphasizes proactive monitoring, SLA adherence, and seamless service delivery.
You will lead incident resolution, perform root-cause analysis, and coordinate with carriers and data centers.
Monitor and maintain the availability, performance, and health of network, systems, voice, cloud, and managed service platforms.
Perform proactive monitoring, incident detection, fault isolation, service restoration, and operational troubleshooting.
Ensure services meet defined Service Level Agreements (SLAs) and operational targets.
Perform configuration, validation, and operational support of core network infrastructure and/or customer edge devices.
Coordinate with local and international carriers, data centre providers, vendors, and partners for service delivery and fault resolution.
Participate in 24×7 operational support through rotating shift and on-call duties
Own incidents from initial detection through restoration and closure.
Perform root cause analysis for recurring operational issues.
Participate in change implementation, validation, rollback planning, and post-change verification.
Maintain accurate operational documentation, knowledge articles, and network records.
Provide technical support for internal teams including Customer Service, Service Delivery, Project Management, and Field Operations.
Act as technical escalation for operational issues affecting customer services.
Validate new service activations and ensure successful customer onboarding.
Support customer maintenance activities and planned service migrations.
Operate and maintain monitoring, alerting, logging, and observability platforms.
Analyse operational events, alarms, logs, and performance metrics to identify service degradation before customer impact.
Produce operational reports covering incidents, availability, utilization, service trends, and capacity planning.
Continuously improve monitoring coverage, alert quality, and operational dashboards.
Develop and maintain scripts to automate repetitive operational tasks, health checks, reporting, and service validation.
Utilize Python, PowerShell, Bash, or similar scripting languages to improve operational efficiency.
Identify opportunities to simplify operational workflows through automation and standardization.
Strong analytical and troubleshooting skills.
Able to remain calm and make sound technical decisions during major incidents.
Customer-focused with strong communication and stakeholder management skills.
Demonstrates ownership and accountability for operational outcomes.
Passionate about continuous improvement and operational automation.
Willing to learn new technologies across networking, systems, cloud, and automation.
Networking: TCP/IP, IPv4 / IPv6, OSPF, MPLS, VLAN, VPN, DNS, DHCP, NAT
Infrastructure: Linux administration, Virtualization technologies, Storage fundamentals
Monitoring & Observability: Grafana, Prometheus, Zabbix, Graylog, ELK / OpenSearch
Automation: Python, PowerShell, Bash, REST APIs
Network Platforms: Nokia, Juniper, Huawei