I. PURPOSE
The NOC Lead of the 24/7 Network Team is responsible for leading the day-to-day operations, overseeing Level 1 (L1) and Level 2 (L2) engineers in the monitoring, detection, triage, and resolution of incidents across all infrastructure services in the company's Managed Services Catalogue:
- Managed Network Infrastructure
- Managed IT Compute & Server Management
This role ensures that NOC teams deliver timely responses, effective escalations, and high-quality service in compliance with SLAs, while fostering operational excellence and continuous improvement in a 24/7 shiftbased environment.
DUTIES AND RESPONSIBILITIES
- Accomplish all assigned tasks by the Shift Manager in a timely and effective manner as deemed necessary for the betterment of the organization.
- Follow effective and efficient processes and comply with escalation protocols.
- Report significant events to the Shift Manager and participate in shift turnovers.
- Contribute to the knowledge and information relevant to Service Operations.
- Collaborate with other team members to improve workflows, documentations, standards, and processes.
- Participate in activities promoting a harmonious working environment such as demonstrating trust and respect and practicing open communication.
- Comply with company policies, guidelines, standards, and procedures.
- Perform all other duties and tasks as assigned by the Shift Manager and Technical Manager.
NOC Operations & Service Monitoring
- Oversee real-time monitoring of network, systems, security, cloud, and application infrastructure.
- Ensure timely detection, logging, and categorization of events and incidents.
- Direct L1 and L2 teams in troubleshooting and incident containment, ensuring minimal downtime.
- Enforce proactive monitoring processes to detect anomalies before they become service-impacting.
Incident & Escalation Management
- Act as the escalation point for unresolved L1/L2 issues, ensuring appropriate handover to L3 teams when required.
- Manage major incident response calls, coordinating technical teams, telco providers, and stakeholders.
- Oversee communication updates during incidents for both internal teams and clients.
Team Leadership & Performance Management
- Lead L1 and L2 engineers across multiple shifts to maintain 24/7 service coverage.
- Conduct shift handovers and ensure accurate documentation of ongoing incidents.
- Mentor, train, and coach team members to improve technical skills and incident handling.
- Monitor team KPIs such as response times, resolution times, and SLA compliance.
Process & Continuous Improvement
- Implement and maintain ITIL v4-aligned processes for incident, event, and request management.
- Work with Service Managers and L3 leads to refine escalation procedures and knowledge base articles.
- Identify operational gaps and recommend tools, automations, and workflow improvements.
Stakeholder & Client Interaction
- Coordinate with Service Managers for incident reports, service updates, and performance reviews.
- Support client onboarding by ensuring proper monitoring configurations and alert workflows are in place.
- Participate in service review meetings and operational governance discussions.
III. QUALIFICATIONS
Minimum Education
Must be a graduate of any IT related bachelor's degree such as:
Minimum Experience/Training
- Have at least 5 years of working experience in a Network Operations Center (NOC) or IT operations, with at least 2 years in a leadership role, overseeing L1/L2 Engineers.
- Experience with enterprise monitoring tools (e.g., SolarWinds, ManageEngine, PRTG, Zabbix, etc.).
- Knowledge of network, server, cloud, and security fundamentals.
- Familiarity with ITIL v4 processes for incident, event, and request management.
C. Competency
- Operational Awareness – Strong situational awareness in a live monitoring environment.
- Leadership & Mentorship – Skilled at guiding and developing technical teams across shifts.
- Problem-Solving – Quick decision‑making for incident containment and resolution.
- Communication – Clear, concise updates during incident calls and escalations.
- Process Orientation – Consistent application of structured workflows and SLAs.
Performance Metrics
- SLA and KPI compliance for incident response and resolution.
- Accuracy and timeliness of incident documentation and updates.
- Mean Time to Detect (MTTD) and Mean Time to Resolve (MTTR).
- Team productivity and escalation efficiency.
- Client satisfaction and positive feedback in service reviews.