NOC / SRE Lead – NOC Operations

Phykon

Thiruvananthapuram

On-site

INR 4,200,000 - 6,400,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Phykon is seeking an experienced NOC/SRE Lead in Thiruvananthapuram to own fleet health, observability, and incident response for a distributed data center infrastructure.

You will lead deep diagnosis, runbooks, and operator certification while acting as escalation point for the most complex incidents. This role blends hands-on technical work with NOC leadership and aims for growth into senior management.

Qualifications

  • 8+ years in SRE, production engineering, or senior NOC roles with platform depth.
  • Production Kubernetes and Linux experience, ideally in edge or distributed environments.
  • Ownership-level experience with Prometheus and Grafana — designing scrape architectures and alert rules.
  • Hands‑on experience with data center operations across data center, colocation, or distributed edge sites.
  • Strong practical networking: TCP/IP, NAT/CGNAT, tunnels, and degraded-link diagnosis.
  • Demonstrated on-call leadership and Sev1 incident command in production environments.
  • Proven ability to codify standards, runbooks, and certification programs for operations teams.
  • Strong ownership mindset with ability to drive incidents, projects, and improvements to completion.
  • Excellent stakeholder communication across engineering and business teams.

Responsibilities

  • Own fleet health and observability, designing monitoring stacks and SLOs for the NOC.
  • Oversee data center and infrastructure operations, coordinating with providers and remote-hands teams.
  • Codify standards, runbooks, and operator certification for scalable NOC operations.
  • Lead incident response, including Sev1/P1 escalations, and drive root cause analysis and post-incident actions.
  • Engage with customers and partners as a technical point of contact during operational events.

Skills

NOC Lead
SRE Experience
Kubernetes
Linux
Incident Command
On-call Leadership
Stakeholder Management
Root Cause Analysis

Tools

Prometheus
Grafana
IaC tooling

Job description

Location: Thiruvananthapuram, Kerala, India (Onsite)

Experience: 8+ Years

Employment Type: Full-Time

Department: Network Operations Center (NOC)

Reports To: [Head of Operations / Engineering – to be confirmed]

About the Role

We are seeking an experienced NOC / SRE Lead to own the health, observability, and incident response for a connected fleet of field-deployed systems and the data center infrastructure that supports it.

This is the keystone L3 role in our Network Operations Center (NOC). You will own deep diagnosis and root-cause analysis on the running system, command the hardest escalations, and build the standards, runbooks, observability architecture, and operator-certification program that the rest of the NOC inherits.

As the L3 Lead, you will act as the escalation point above L2 support teams, taking ownership of the most complex incidents while keeping escalations to L4 product engineering limited to genuine code and firmware issues — protecting both fleet uptime and engineering velocity.

This is a greenfield opportunity. In the near term, the role combines a hands-on L3 technical function with NOC leadership. As the NOC scales, the role is designed to grow into a Principal SRE or NOC Manager track.

Key Responsibilities

1. Fleet Health & Observability

  • Architect the monitoring stack the NOC runs on — scrape architecture, alert rules, dashboards, and SLOs — rather than only consuming it.
  • Own end-to-end visibility of the connected fleet across data center, edge Kubernetes, Linux, and network layers.
  • Continuously tune alerting to reduce noise and catch degradation early.

2. Data Center & Infrastructure Operations

  • Oversee the health and availability of data center and colocation infrastructure — compute, storage, network, power, and cooling dependencies — supporting the fleet.
  • Coordinate with data center operators, colocation providers, and remote-hands teams for physical interventions, maintenance windows, and capacity changes.
  • Manage hardware fault handling, RMA workflows, and site-level incident response across distributed data center and edge sites.
  • Coordinate with telecommunications providers, colocation partners, cloud vendors, and infrastructure service providers to resolve service-impacting issues and maintain operational continuity.
  • Serve as incident commander on the hardest escalations, including Sev1/P1 incidents, and drive them to resolution within SLA.
  • Perform deep diagnosis and root-cause analysis on the running system using logs, metrics, and telemetry.
  • Diagnose complex issues across:
    • Network connectivity, routing, NAT/CGNAT, and tunnels
    • Degraded or intermittent links to field-deployed hardware
    • Production Kubernetes and Linux platform dependencies at the edge
  • Assume end-to-end ownership of major incidents, service disruptions, and customer escalations until resolution and formal closure.
  • Act as the highest operational escalation point within the NOC for complex technical and service- impacting incidents.
  • Provide timely updates to internal stakeholders, leadership teams, and customers during major incidents and service outages.

4. Standards, Runbooks & Operator Certification

  • Codify diagnostic procedures, escalation paths, and operational standards that make the NOC repeatable and scalable.
  • Own the operator-certification layer that qualifies operators to run and support the fleet.
  • Develop and continuously improve runbooks and operational procedures inherited by L1/L2 teams.
  • Maintain escalation hygiene: reserve L4 product engineering for genuinely complex issues requiring code or firmware changes.
  • Clearly define problem scope, business impact, and affected systems when escalating, with supporting logs and evidence.
  • Facilitate technical communication across engineering and operations teams during major incidents, providing timely stakeholder updates.
  • Lead post-incident reviews and drive corrective actions to closure.
  • Identify recurring incidents and contribute preventive improvements and monitoring optimization.

7. Client & Stakeholder Engagement

  • Act as a primary technical point of contact for customers, partners, vendors, and internal stakeholders on operational and service-related matters.
  • Participate in customer meetings, operational reviews, service reviews, and technical discussions as a subject matter expert.
  • Provide clear, professional, and executive-level communication during major incidents, service disruptions, and operational escalations.
  • Build and maintain strong working relationships with customer technical teams while driving accountability and service excellence.
  • Provide technical leadership, mentoring, and guidance to L1 and L2 engineers.
  • Drive knowledge sharing, operational maturity, and continuous improvement initiatives across the NOC function.
  • Support the development and maintenance of operator certification programs and competency frameworks.
  • Participate in recruitment, onboarding, training, and capability development activities as the NOC organization scales.
Required Skills & Qualifications
  • 8+ years in SRE, production engineering, or senior NOC roles with strong platform depth (not generalist IT-NOC).
  • Production Kubernetes and Linux experience, ideally in edge or distributed environments rather than pure cloud.
  • Ownership-level experience with Prometheus and Grafana — designing scrape architectures and alert rules, not just reading dashboards.
  • Hands‑on experience with data center operations — running production infrastructure across data center, colocation, or distributed edge sites.
  • Strong practical networking: TCP/IP, NAT/CGNAT, tunnels, and degraded-link diagnosis.
  • Demonstrated on-call leadership and Sev1 incident command in production environments.
  • Proven ability to codify standards, runbooks, and certification programs for operations teams.
  • Strong analytical, troubleshooting, and documentation skills, with excellent communication across engineering and business stakeholders.
  • Strong ownership mindset with the ability to independently drive incidents, projects, and operational improvements to completion.
  • Ability to remain calm, structured, and decisive during high‑severity incidents and customer escalations.
  • Excellent stakeholder management and communication skills, including the ability to communicate complex technical concepts to both technical and non‑technical audiences.
  • Experience interacting directly with enterprise customers, vendors, and senior stakeholders in operational environments.
Tools & Technologies

Observability & Telemetry: Prometheus, Grafana, and related monitoring/alerting platforms

Automation: Infrastructure-as-Code tooling (advantage)

Incident & Collaboration: Incident management, ticketing, and on-call platforms

Preferred Qualifications
  • Greenfield NOC build experience — standing up a NOC, runbooks, and on‑call from scratch (strongest positive signal).
  • Prior NOC/SRE experience supporting a hardware or field-deployed product.
  • Exposure to OT/BMS environments and comfort bridging IT and industrial systems.
  • Relevant certifications (Kubernetes, Linux, networking, or cloud) are an advantage.
  • Highly desired: Proven experience building or scaling a Network Operations Center (NOC) from the ground up, including monitoring strategy, runbooks, escalation models, operational processes, and team development.
  • Operates within a 24×7 Network Operations Center supporting a live fleet of field-deployed systems and distributed data center infrastructure.
  • May require participation in rotational shifts, weekend coverage, planned maintenance activities, and on-call support arrangements.
  • Serves as a key escalation point for critical incidents and major customer-impacting events.
  • Acts as a primary technical point of contact for customers, partners, and internal stakeholders during operational reviews and service incidents.
  • Expected to perform effectively in high-pressure operational environments while maintaining clear communication, leadership, and decision-making capabilities.
  • Provides technical leadership, mentoring, and operational guidance to L1 and L2 support teams.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

SRE Lead (NOC Operations)
SRE Lead (NOC Operations)

Phykon Solutions • Thiruvananthapuram

On-site
INR 1,800,000 - 2,400,000
Noc Engineer
Noc Engineer

Alike Thoughts • Bengaluru

Hybrid
INR 1,400,000 - 2,200,000
Network Operations Team Lead
Network Operations Team Lead

Availity • Bengaluru

On-site
INR 1,200,000 - 1,800,000
Noc Specialist / SRE engineer
Noc Specialist / SRE engineer

Bridgenext • Pune District

On-site
INR 1,400,000 - 2,200,000
Network Operations Center - NOC
Network Operations Center - NOC

YO IT Consulting • Hyderabad

Hybrid
INR 800,000 - 1,200,000
Opportunity to work with enterprise-level network infrastructure
Career growth opportunities in Network Operations and IT Infrastructure
Comprehensive training on tools and operational processes
Network Operations Manager
Network Operations Manager

Futurism Technologies, INC. • Pune District

On-site
INR 4,000,000 - 6,000,000
Noc Manager
Noc Manager

Zybisys Consulting Services • Bengaluru

On-site
INR 1,500,000 - 2,000,000
NOC Manager
NOC Manager

Futurism Technologies • Mhalunge

On-site
INR 4,000,000 - 7,000,000
NOC Manager
NOC Manager

Capgemini • Bengaluru, Gurugram District, Bangalore Rural

On-site
INR 1,200,000 - 1,800,000
NOC Operation Engineer
NOC Operation Engineer

YO IT Consulting • Hyderabad

Hybrid
INR 600,000 - 800,000
Opportunity to work with enterprise-level network infrastructure
Comprehensive training on tools and processes
Collaborative work environment