Sr. NOC Engineer (NOC, SRO, Scripting, Observability, AWS, Monitoring Tools)

Vertafore Career Center

Hyderabad

On-site

INR 2,000,000 - 3,000,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Vertafore Career Center in Hyderabad seeks a Senior Reliability Operations Engineer to own day-to-day production operations, service continuity, and restoration excellence across AWS, hybrid data centers, and customer-hosted environments. You will lead incident response, readiness reviews, and automation initiatives with cross-functional partners.

The role emphasizes operational governance, runbook maturity, and continuous improvement, partnering with SRE, engineering, security, and product

Responsibilities

  • Own operational execution for day-to-day service health, including monitoring, triage, escalation, restoration coordination, operational risk tracking, and service continuity across cloud and hybrid environments.
  • Lead operational readiness reviews for new services, releases, infrastructure changes, and production onboarding; ensure support models, runbooks, dashboards, alerts, escalation paths, rollback procedures, access, and validation checks are ready before handoff.
  • Operate and tune monitoring, alerting, dashboards, and operational observability practices aligned with the Four Golden Signals; partner with SRE and engineering on telemetry and instrumentation standards.
  • Track SLO attainment, SLA risk, error budget burn, operational KPIs, recurring incidents, and reliability trends; elevate service risk and recommend operational guardrails when thresholds are breached.
  • Monitor service performance, infrastructure utilization, capacity, saturation, and customer-impacting degradation to proactively identify and elevate operational risk.
  • Maintain an operational toil backlog and reduce repetitive manual work through automation, scripting, AI-assisted operations, self-healing workflows, AIOps capabilities and process simplification.
  • Plan, execute, coordinate, and validate production changes including patching, certificate renewals, software releases, infrastructure updates, and maintenance using standardized change governance, risk assessment, rollback readiness, stakeholder communication, and post-change verification.
  • Investigate and troubleshoot complex production issues across applications, infrastructure, middleware, databases, and platform services; restore service while partnering with SRE/engineering on permanent remediation.
  • Develop and improve runbooks, SOPs, escalation frameworks, recovery procedures, production validation checks and knowledge documentation to ensure consistent support execution.
  • Identify recurring operational failure patterns and partner with SRE, engineering, platform, and application teams to drive preventive controls and permanent corrective actions.
  • Act as the operational incident lead during major incidents, managing bridges, stakeholder communications, escalation paths, restoration coordination, event timelines, and post-incident follow-up.
  • Facilitate blameless post-incident reviews and own corrective-action tracking for remediation items, runbook updates, automation opportunities, preventive controls, and operational improvements.
  • Own incident, problem, and service reliability trend reporting, operational dashboards, SLA/SLO tracking, error budget burn visibility and leadership reviews.
  • Support the GCC operational model by driving globally standardized operational practices, governance, runbook maturity, escalation frameworks, and service management processes.
  • Collaborate with globally distributed SRE, engineering, cloud operations, security, support, product, and business teams to align operational priorities with reliability risks and service-continuity needs.
  • Mentor junior engineers and promote knowledge sharing while maintaining a customer-first mindset focused on service reliability, responsiveness, operational quality, and business continuity.

Job description

Role Summary

We are seeking a Senior Reliability Operations Engineer to own day-to-day production operations, service continuity, operational readiness, and restoration excellence for critical production services. This role is responsible for monitoring effectiveness, incident coordination, production change execution, operational governance, runbook maturity, automation adoption, and continuous operational improvement across AWS, hybrid data centers, and customer-hosted environments.

As part of the Global Command Center (GCC), this role partners with SRE, engineering, platform, security, product, support, and business teams to ensure services are observable, supportable, resilient, and operationally mature. The role focuses on operational execution and service reliability outcomes while partnering with engineering on systemic reliability improvements.

Key Responsibilities
Service Reliability & Production Operations
  • Own operational execution for day-to-day service health, including monitoring, triage, escalation, restoration coordination, operational risk tracking, and service continuity across cloud and hybrid environments.
  • Lead operational readiness reviews for new services, releases, infrastructure changes, and production onboarding; ensure support models, runbooks, dashboards, alerts, escalation paths, rollback procedures, access, and validation checks are ready before handoff.
  • Operate and tune monitoring, alerting, dashboards, and operational observability practices aligned with the Four Golden Signals; partner with SRE and engineering on telemetry and instrumentation standards.
  • Track SLO attainment, SLA risk, error budget burn, operational KPIs, recurring incidents, and reliability trends; elevate service risk and recommend operational guardrails when thresholds are breached.
  • Monitor service performance, infrastructure utilization, capacity, saturation, and customer-impacting degradation to proactively identify and elevate operational risk.
Operational Excellence & Automation
  • Maintain an operational toil backlog and reduce repetitive manual work through automation, scripting, AI-assisted operations, self-healing workflows, AIOps capabilities and process simplification.
  • Plan, execute, coordinate, and validate production changes including patching, certificate renewals, software releases, infrastructure updates, and maintenance using standardized change governance, risk assessment, rollback readiness, stakeholder communication, and post-change verification.
  • Investigate and troubleshoot complex production issues across applications, infrastructure, middleware, databases, and platform services; restore service while partnering with SRE/engineering on permanent remediation.
  • Develop and improve runbooks, SOPs, escalation frameworks, recovery procedures, production validation checks and knowledge documentation to ensure consistent support execution.
  • Identify recurring operational failure patterns and partner with SRE, engineering, platform, and application teams to drive preventive controls and permanent corrective actions.
Incident Management, Reporting & GCC Collaboration
  • Act as the operational incident lead during major incidents, managing bridges, stakeholder communications, escalation paths, restoration coordination, event timelines, and post-incident follow-up.
  • Facilitate blameless post-incident reviews and own corrective-action tracking for remediation items, runbook updates, automation opportunities, preventive controls, and operational improvements.
  • Own incident, problem, and service reliability trend reporting, operational dashboards, SLA/SLO tracking, error budget burn visibility and leadership reviews.
  • Support the GCC operational model by driving globally standardized operational practices, governance, runbook maturity, escalation frameworks, and service management processes.
  • Collaborate with globally distributed SRE, engineering, cloud operations, security, support, product, and business teams to align operational priorities with reliability risks and service-continuity needs.
  • Mentor junior engineers and promote knowledge sharing while maintaining a customer-first mindset focused on service reliability, responsiveness, operational quality, and business continuity.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Noc Specialist / SRE engineer
Noc Specialist / SRE engineer

Bridgenext • Pune District

On-site
INR 1,400,000 - 2,200,000
SRE Lead (NOC Operations)
SRE Lead (NOC Operations)

Phykon Solutions • Thiruvananthapuram

On-site
INR 1,800,000 - 2,400,000
VS01700 - SRE & Production Reliability Engineer
VS01700 - SRE & Production Reliability Engineer

E4 Software Services Pvt Ltd. • India

On-site
INR 2,000,000 - 4,000,000
Production Support Lead
Production Support Lead

Cloudxtreme • Hyderabad

On-site
INR 3,200,000 - 6,000,000
Noc Engineer
Noc Engineer

Alike Thoughts • Bengaluru

Hybrid
INR 1,400,000 - 2,200,000
NOC / SRE Lead – NOC Operations
NOC / SRE Lead – NOC Operations

Phykon • Thiruvananthapuram

On-site
INR 4,200,000 - 6,400,000
Associate Network Ops Manager
Associate Network Ops Manager

HealthEdge • Hyderabad

On-site
INR 1,800,000 - 2,400,000
Forward Deployment Engineer (SRE)
Forward Deployment Engineer (SRE)

PwC Acceleration Centers • Bengaluru

On-site
INR 400,000 - 700,000
Forward Deployment Engineer (SRE)
Forward Deployment Engineer (SRE)

PwC Acceleration Centers • Hyderabad

On-site
INR 2,500,000 - 4,000,000
Site Reliability Engineering Lead (Application SRE Lead)
Site Reliability Engineering Lead (Application SRE Lead)

Hirexa Solutions • Bengaluru

Hybrid
INR 3,500,000 - 7,000,000