As a NOC Operations & Enablement Engineer, you will report to the NOC Lead and work closely with CBTW Program Leadership to operationalize our “single pane of glass” philosophy — where every alert that reaches an operator is signal, every action is backed by a runbook, and every P1/P2 incident has a commander (EWS’s Application Commander).
We don’t need someone who just watches dashboards. We need a relentless noise-hunter with strong operational fundamentals — someone who can triage a daily landscape of roughly 267 alerts across the NOC down to what actually matters, write the procedures everyone else follows, and stay calm in command when the environment is on fire.
Responsibilities
Observability & Alert Engineering
- Operate and tune monitoring/alerting platforms (Splunk, Zabbix, AppDynamics, and Honeycomb), the PagerDuty alerting hub, and ITSM systems (ServiceNow) as your daily working environment.
- Triage at scale. Group alerts by source, frequency, and component. Discovery already found that 259 alert signatures — under 5% of all signatures — account for half of all alert volume; isolate patterns like these and drive consolidation.
- Define and validate severity. Pre-classify alerts against the existing P1–P5 matrix and the 76-application severity tiering already built during Discovery, confirm assignments with EWS SMEs, and keep the framework consistent across the environment.
- Hunt blind spots. Cross-reference environment architecture against active alerts — building on the 15 confirmed blind spots already identified during Discovery (9 of them in Tier-1 applications) — to find remaining coverage gaps and specify the monitors needed.
- Engineer signal. Build correlation / parent-child rules, automated health updates, and status-page integrations that remove manual overhead.
NOC Operations & Runbook Discipline
- Execute runbooks precisely and consistently for actionable alerts, with full adherence to documented procedures.
- Author and maintain runbooks. Translate actionable alerts into step-by-step, validated L1 procedures. 659 runbooks already exist, but only 183 currently match Discovery’s 642 confirmed actionable alert types — and many of those have thin or missing remediation steps. Close that gap and test procedures in a live-system context.
- Operate from one clean view. Help shape and then work from a redesigned, unambiguous single-pane-of-glass NOC dashboard — EWS leadership has already greenlit a custom interim pane (React + ECharts) as the near-term direction.
- Own shift handoffs. Maintain clean shift logs and handover notes so coverage across Americas/APAC is seamless.
- Set the standard. Define documentation, naming, and quality standards that sustain severity and runbook coverage targets.
Incident Command & Escalation
- Respond first. Acknowledge P1/P2 alerts within target MTTA (5 minutes for P1, 10 minutes for P2), classify correctly, and act or elevate per the matrix without hesitation.
- Command incidents. Run P1/P2 bridges as Application Commander (EWS’s current term for incident commander) — coordinate responders, drive communications, and push resolution toward MTTR targets (60 minutes for P1, 240 minutes for P2).
- Standardize coordination. Deploy standardized incident bridges and resolve external dialing friction for customers.
- Build the IR process. Define and continuously improve the escalation matrix — mapping severity tiers to response teams, SLAs, and communication channels.
- Close the loop. Run post-incident learning, capture root cause, and feed fixes back into monitors and runbooks.
Visibility, Automation & Continuous Improvement
- Stand up visibility. Deploy and maintain a centralized status page providing real-time environment health — this directly supports EWS leadership’s top backlog priority (status-page automation), so customers subscribe to updates instead of EWS pushing manual comms.
- Automate the manual. Replace error-prone manual status updates with automated health rules.
- Drive the metrics. Push measurable progress against the stabilization exit criteria — alert rationalization (Discovery modeled a 56–68% cut in daily NOC alert volume), severity/runbook coverage, blind-spot elimination, and MTTA/MTTR.
Qualifications
Must-Haves
- Operational Foundation: 3+ years of hands‑on NOC, SRE, or IT operations experience in a 24/7 monitored environment.
- Monitoring & ITSM Fluency: working depth in at least one observability platform used at EWS (Splunk, Zabbix, AppDynamics, or Honeycomb) and one ITSM tool (ServiceNow or equivalent).
- Incident Composure: a demonstrated history of triaging and resolving incidents under pressure while following — and improving — documented procedures.
- Noise-Reduction Instinct: you don’t tolerate alert chaos. You instinctively consolidate, correlate, and rationalize until what’s left is actionable.
- Communication & Clarity: ability to write clear runbooks and incident updates, and to explain operational status to non-technical stakeholders in professional English (U.S.-facing engagement).
- Shift Flexibility: able to work coverage aligned to Americas night / APAC daytime.
Nice-to-Haves
- Scripting/automation skills for health checks, status pages, or alert tuning (Python, shell, or platform‑native).
- Incident command (Application Commander) experience on P1/P2 bridges; familiarity with ITIL incident management.
- Experience with status-page tooling and automated health reporting.
- Exposure to correlation / AIOps tooling for alert noise reduction (EWS is currently evaluating xMatters).
- Relevant certifications (ITIL, platform‑specific monitoring, or cloud).