Key Responsibilities
1. Fleet Health & Observability
- Architect the monitoring stack the NOC runs on scrape architecture, alert rules, dashboards, and SLOs rather than only consuming it.
- Own end‑to‑end visibility of the connected fleet across data center, edge Kubernetes, Linux, and network layers.
- Continuously tune alerting to reduce noise and catch degradation early.
2. Data Center & Infrastructure Operations
- Oversee the health and availability of data center and colocation infrastructure compute, storage, network, power, and cooling dependencies — supporting the fleet.
- Coordinate with data center operators, colocation providers, and remote-hands teams for physical interventions, maintenance windows, and capacity changes.
- Manage hardware fault handling, RMA workflows, and site‑level incident response across distributed data center and edge sites.
- Coordinate with telecommunications providers, colocation partners, cloud vendors, and infrastructure service providers to resolve service‑impacting issues and maintain operational continuity.
3. Incident Management & Response
- Serve as incident commander on the hardest escalations, including Sev1/P1 incidents, and drive them to resolution within SLA.
- Perform deep diagnosis and root‑cause analysis on the running system using logs, metrics, and telemetry.
- Diagnose complex issues across:
- Network connectivity, routing, NAT/CGNAT, and tunnels
- Degraded or intermittent links to field‑deployed hardware
- Production Kubernetes and Linux platform dependencies at the edge
- Assume end‑to‑end ownership of major incidents, service disruptions, and customer escalations until resolution and formal closure.
- Act as the highest operational escalation point within the NOC for complex technical and service‑impacting incidents.
- Provide timely updates to internal stakeholders, leadership teams, and customers during major incidents and service outages.
4. Standards, Runbooks & Operator Certification
- Codify diagnostic procedures, escalation paths, and operational standards that make the NOC repeatable and scalable.
- Own the operator‑certification layer that qualifies operators to run and support the fleet.
- Develop and continuously improve runbooks and operational procedures inherited by L1/L2 teams.
5. Escalation & Coordination
- Maintain escalation hygiene: reserve L4 product engineering for genuinely complex issues requiring code or firmware changes.
- Clearly define problem scope, business impact, and affected systems when escalating, with supporting logs and evidence.
- Facilitate technical communication across engineering and operations teams during major incidents, providing timely stakeholder updates.
6. Continuous Service Improvement
- Lead post‑incident reviews and drive corrective actions to closure.
- Identify recurring incidents and contribute preventive improvements and monitoring optimization.
7. Client & Stakeholder Engagement
- Act as a primary technical point of contact for customers, partners, vendors, and internal stakeholders on operational and service‑related matters.
- Participate in customer meetings, operational reviews, service reviews, and technical discussions as a subject matter expert.
- Provide clear, professional, and executive‑level communication during major incidents, service disruptions, and operational escalations.
- Build and maintain strong working relationships with customer technical teams while driving accountability and service excellence.
8. Leadership & Team Development
- Provide technical leadership, mentoring, and guidance to L1 and L2 engineers.
- Drive knowledge sharing, operational maturity, and continuous improvement initiatives across the NOC function.
- Support the development and maintenance of operator certification programs and competency frameworks.
- Participate in recruitment, onboarding, training, and capability development activities as the NOC organization scales.
Required Skills & Qualifications
- 8+ years in SRE, production engineering, or senior NOC roles with strong platform depth (not generalist IT‑NOC).
- Production Kubernetes and Linux experience, ideally in edge or distributed environments rather than pure cloud.
- Ownership‑level experience with Prometheus and Grafana — designing scrape architectures and alert rules, not just reading dashboards.
- Hands‑on experience with data center operations — running production infrastructure across data center, colocation, or distributed edge sites.
- Strong practical networking: TCP/IP, NAT/CGNAT, tunnels, and degraded‑link diagnosis.
- Demonstrated on‑call leadership and Sev1 incident command in production environments.
- Proven ability to codify standards, runbooks, and certification programs for operations teams.
- Strong analytical, troubleshooting, and documentation skills, with excellent communication across engineering and business stakeholders.
- Strong ownership mindset with the ability to independently drive incidents, projects, and operational improvements to completion.
- Ability to remain calm, structured, and decisive during high‑severity incidents and customer escalations.
- Excellent stakeholder management and communication skills, including the ability to communicate complex technical concepts to both technical and non‑technical audiences.
- Experience interacting directly with enterprise customers, vendors, and senior stakeholders in operational environments.
Category
Tools
Observability & Telemetry
Prometheus, Grafana, and related monitoring/alerting platforms
Platform
Kubernetes, Linux (edge/distributed deployments)
Automation
Infrastructure-as-Code tooling (advantage)
Incident & Collaboration
Incident management, ticketing, and on‑call platforms
Preferred Qualifications
- Greenfield NOC build experience — standing up a NOC, runbooks, and on‑call from scratch (strongest positive signal).
- Prior NOC/SRE experience supporting a hardware or field‑deployed product.
- Exposure to OT/BMS environments and comfort bridging IT and industrial systems.
- Relevant certifications (Kubernetes, Linux, networking, or cloud) are an advantage.
- Highly desired: Proven experience building or scaling a Network Operations Center (NOC) from the ground up, including monitoring strategy, runbooks, escalation models, operational processes, and team development.
Work Environment
- Operates within a 247 Network Operations Center supporting a live fleet of field‑deployed systems and distributed data center infrastructure.
- May require participation in rotational shifts, weekend coverage, planned maintenance activities, and on‑call support arrangements.
- Serves as a key escalation point for critical incidents and major customer‑impacting events.
- Acts as a primary technical point of contact for customers, partners, and internal stakeholders during operational reviews and service incidents.
- Expected to perform effectively in high‑pressure operational environments while maintaining clear communication, leadership, and decision‑making capabilities.
- Provides technical leadership, mentoring, and operational guidance to L1 and L2 support teams.