Job Description:
Description
Position Summary
The NOC Engineer is part of a newly established Network Operations Center that moves IT operations from a reactive support model toward proactive, centralized monitoring across our US and India environment. The role spans the full operational picture — platform and service health, business applications, data ingestion and integration jobs, and security alerting — not infrastructure alone.
This is not a help desk role that waits for tickets. The NOC Engineer detects issues through monitoring and alerts, assesses business impact, performs Tier 1 triage and resolution, escalates the rest to the appropriate engineering, application, data, or security teams, and retains ownership until service is restored. Desktop and end-user support is carried in parallel with these duties on the same shift.
The NOC is the operational front door. It does not replace specialized engineering teams: infrastructure, network, cloud, security, application, and data teams retain ownership of advanced engineering work and permanent corrective actions.
Key Responsibilities
Monitoring and Alert Response
- Monitor platform, service-health, and application dashboards throughout the shift; validate every alert and classify it by severity and business impact before acting.
- Respond to alerts covering availability, performance, CPU, memory, disk, connectivity, service health, and application errors.
- Write and modify log queries (KQL) to investigate incidents and surface operational trends; tune alert thresholds and routing rules to reduce noise and false positives.
- Monitor connectivity and latency between sites, VPN paths, and cloud services; monitor virtual desktop session health and user experience.
- Monitor backups, endpoint health, patching, device compliance, and scheduled operational jobs; perform documented daily health checks for anything not yet covered by automated monitoring.
- Create and maintain the service-health dashboards and workbooks the team works from, and keep them current as the environment changes.
- Map what is monitored today, identify critical systems and services not yet covered, and onboard them into monitoring so that manual checks are progressively replaced by automated alerting
Requirements
Tier 1 Application and Data Ingestion Support
- Monitor business-critical applications for availability, response time, error rates, and failed transactions; perform Tier 1 triage and resolution within documented procedures.
- Monitor data ingestion, integration, and interface jobs — scheduled transfers, feeds, and pipelines — for failures, delays, backlogs, and record-count or reconciliation mismatches.
- Restart failed jobs, re-trigger transfers, and clear known error conditions where a runbook exists; escalates anything outside the runbook to the application or data teams with full diagnostic detail.
- Track ingestion issues through to confirmed data delivery, not just to job restart.
Security Alert Monitoring (Tier 1)
- Monitor and triage security alerts — suspicious sign-ins, account lockouts, MFA anomalies, endpoint protection detections, and reported phishing — as first line of review.
- Validate alerts against documented criteria, capture supporting evidence, and elevate to the Security team promptly with a clear summary; deeper investigation and remediation remain with Security.
- Follow defined containment steps (such as user or device isolation) only where explicitly authorised by runbook.
Incident Response and Ownership
- Perform first-level triage, determine business impact, resolve within established procedures, and elevate complex incidents to the Network, Server, Cloud, Application, Data, or Security teams as appropriate.
- Create and update incidents in Freshservice with accurate categorisation, timelines, and diagnostic detail.
- Own escalated incidents through to service restoration; support incident coordination, communication, and stakeholder updates during major incidents.
- Maintain a consistent shift-handover log and end-of-shift summaries in Teams and Freshservice.
- Identify recurring incidents, contribute to root-cause analysis and post-incident reviews rather than repeatedly resolving the same symptoms.
- Maintain targeted watchlists for known problem areas and review operational trends to spot issues before they generate alerts.
Desktop and End-User Support (carried in parallel)
- Provide day-to-day desktop and end-user support — desktops, laptops, peripherals, printers, conference rooms, and business applications — alongside monitoring duties on the same shift.
- Manage the Freshservice queue to agreed response and resolution targets, prioritising between desktop requests and NOC alerts by business impact.
- Support endpoint lifecycle activities: imaging, deployment, onboarding and offboarding, compliance, patching, and asset accuracy.
- Provide the Help Desk with clear pass/fail checks, documented escalation steps, and defined NOC tasks they can support during each shift.
- Train and cross-train Help Desk staff to acknowledge, validate, and elevate alerts correctly, and act as a technical point of reference for them during the shift.
Documentation, Automation, and Reporting
- Develop and maintain runbooks, SOPs, troubleshooting guides, and escalation paths — every alert should map to a documented action.
- Build automation with PowerShell and Azure Automation to reduce repetitive monitoring and remediation work.
- Contribute to operational reporting on service availability, alert trends, recurring issues, response times, and incident resolution, giving leadership clear visibility into the health of the technology environment.
Required Qualifications
- Bachelor’s degree in a technical discipline, or equivalent practical experience.
- 2–4 years in IT support or NOC monitoring, with exposure beyond desktop-only troubleshooting.
- Sound networking fundamentals — TCP/IP, subnetting, DNS, DHCP, VPN, and VLAN concepts — with the ability to isolate where in the path a fault sits using ping, traceroute, nslookup, and port checks. Device configuration depth is not required.
- Working knowledge of cloud monitoring on Microsoft Azure — the portal, Azure Monitor, Log Analytics, Resource Health, and Service Health: reading metrics and logs, understanding alert rules and action groups, and basic KQL. Demonstrable KQL familiarity will be prioritised.
- Ability to read application and job logs to identify failures, and to distinguish an application fault from a platform fault before escalating.
- Awareness of common security alert types — suspicious sign-in, account lockout, malware detection, phishing — and the discipline to elevate rather than investigate beyond defined scope.
- Strong Windows endpoint and server administration, including Active Directory, Entra ID, and Microsoft 365, plus proven hands-on desktop support experience.
- Experience with an ITSM ticketing platform such as Freshservice or ServiceNow, and disciplined documentation of every action taken.
- A methodical approach to troubleshooting, and good spoken and written English – able to explain a technical problem clearly and simply, both on calls and in written updates.
- Availability for a 24×7 rotational shift, including nights, weekends, and holidays on rotation.
Preferred Qualifications
- Experience building alert rules, workbooks, or dashboards from scratch; intermediate KQL (joins, aggregations, time-series operators).
- PowerShell scripting and Azure Automation.
- Exposure to Microsoft Sentinel, Microsoft Defender for Cloud, Microsoft Intune, or Defender for Endpoint.
- Exposure to Power BI or Microsoft Fabric for operational reporting.
- Exposure to application or integration monitoring, scheduled job orchestration, SQL basics, or API and file-transfer troubleshooting.
- Prior experience supporting a healthcare or other regulated environment.
Certifications such as AZ-900, AZ-104, SC-900, MS-900, CCNA, or ITIL Foundation will be an added advantage