Manager of Network Reliability and Resiliency

ServiceNow

Toronto

On-site

CAD 140,000 - 200,000

Full time

5 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Generous family leave
Matched donations
Annual learning stipends
Flexible PTO
Competitive retirement plan
Paid volunteer time

Job summary

ServiceNow seeks a Manager, Network Reliability and Resiliency to lead a team responsible for the reliability and day-to-day operation of production network services supporting ServiceNow’s cloud platform. This is a technical people-manager role focused on coaching engineers, managing priorities, and guiding complex troubleshooting and high-severity incidents.

You will apply SRE principles to network operations, using SLIs/SLOs, error budgets, observability, and automation to improve

Qualifications

  • Five+ years of experience in network engineering, reliability, cloud infrastructure, or large-scale production operations.
  • Experience leading engineers, coaching, and delivering with accountability.
  • Strong written and verbal communication with stakeholder management.

Responsibilities

  • Lead and develop a network reliability team through goals, feedback, and career development.
  • Define priorities for operational work, reliability initiatives, and technical debt.
  • Provide technical and incident leadership during major outages and customer escalations.
  • Improve observability, SLIs/SLOs, monitoring, and post-incident improvements.
  • Collaborate with network engineering, SRE, security, and platform teams to improve service reliability.

Skills

Leadership
SRE Principles
Incident Management
Networking Concepts
Cross-Functional Collaboration

Job description

  • We are seeking a Manager, Network Reliability and Resiliency to lead a team responsible for the reliability and day-to-day operation of production network services supporting ServiceNow’s cloud platform
  • This is a technical people-manager role
  • You will develop engineers and manage team priorities while staying actively engaged in complex troubleshooting, high-severity incidents, customer escalations, operational readiness, and reliability improvement
  • You will apply SRE principles to network operations by using service indicators and objectives, error-budget thinking, observability, post-incident learning, and automation to improve availability, reduce operational toil, and make execution safer and more consistent
  • While this is not an individual contributor role, you must have the technical depth and judgment to guide investigations, challenge assumptions, make risk-based decisions, and help the team reach durable solutions
  • Lead and develop the team:
  • Manage, coach, and develop network reliability engineers through clear goals, regular feedback, performance reviews, and career development
  • Set priorities and ownership for operational work, reliability initiatives, technical debt, and project commitments
  • Build sustainable on-call and escalation practices and promote calm, accountable execution during high-pressure events
  • Hire and onboard new team members and ensure they gain the technical context, operating practices, and support needed to succeed
  • Provide technical and incident leadership:
  • Actively engage in complex production troubleshooting and customer-impacting escalations by reviewing evidence, guiding technical hypotheses, identifying risk, and coordinating the right subject-matter experts
  • Lead or support major incident response, including mitigation decisions, stakeholder communication, escalation management, and restoration of service
  • Ensure post-incident reviews identify contributing factors and result in clear, prioritized, and completed preventive actions
  • Review high-risk changes and operational plans for technical soundness, rollback readiness, monitoring coverage, and customer impact
  • Improve reliability through SRE practices:
  • Partner with engineering and service owners to define and use meaningful SLIs and SLOs for network services
  • Use error budgets, incident trends, capacity signals, and operational data to balance service reliability, delivery pace, and risk
  • Improve observability, alert quality, dashboards, runbooks, and operational readiness so the team can detect and resolve issues efficiently
  • Track practical reliability outcomes such as availability, recurring incidents, change success, alert effectiveness, and time to detect and recover
  • Embed automation in daily operations:
  • Create a strong automation mindset across the team and identify repetitive, error-prone, or slow operational activities that should be eliminated or automated
  • Prioritize automation that improves change safety, validation, triage, remediation, reporting, and operational consistency
  • Work with engineering and automation partners to move useful tools and workflows into production with clear ownership, documentation, monitoring, and support models
  • Measure whether automation reduces toil and operational risk rather than treating automation delivery alone as the outcome
  • Partner across the organization:
  • Collaborate with network engineering, SRE, security, platform, data center, customer support, and other partner teams to resolve issues and improve service reliability
  • Represent the team’s technical assessment, customer impact, risks, dependencies, and recovery plan clearly to technical and business stakeholders
  • Ensure new technologies, services, and automations meet operational acceptance criteria before the team assumes production ownership
  • Improve incident, change, problem-management, and escalation processes based on operational evidence and team feedback
Benefits
  • Generous family leave
  • Matched donations
  • Annual learning stipends
  • Flexible PTO
  • Competitive retirement plan
  • Paid volunteer time

Strong written and verbal communication skills, sound judgment under pressure, and consistent attention to detailWorking knowledge of networking concepts and technologies such as TCP/IP, routing, DNS, load balancing or ADCs, firewalls, cloud networking, and network observability. Deep expertise in every area is not requiredSufficient hands‑on technical background to guide production troubleshooting across Linux‑based systems and network services. You can interpret logs, metrics, alerts, and packet‑level evidence and make sound operational decisionsWorking knowledge of SRE practices, including SLIs, SLOs, error budgets, monitoring and alerting, incident management, and post‑incident improvementExperience leading or coordinating significant incidents and customer‑impacting escalations in an always‑on service environmentFive or more years of relevant experience in network engineering, network reliability, cloud infrastructure, SRE, or large‑scale production operationsExperience using or evaluating AI‑assisted tools to improve analysis, decision‑making, automation, or team workflows, with appropriate attention to accuracy, security, and operational riskExperience working with geographically distributed teams and cross‑functional partners in software, platform, infrastructure, or cloud servicesAn automation mindset and experience using scripting, workflow automation, or engineering partnerships to reduce manual operational work and improve consistencyExperience managing or formally leading engineers, including prioritization, coaching, performance feedback, and delivery accountabilityDue to Government of Canada regulatory requirements, this position requires the successful completion of a Government of Canada Reliability Status screening as a condition of employment. The screening process requires 5 years of verifiable background history. This includes identity verification, education verification, a criminal record check, and a credit check. Candidates must be eligible to obtain and maintain Reliability Status, which generally requires Canadian citizenship or Canadian permanent resident status. Employment is contingent upon successful completion and maintenance of the required screeningExperience operating networking for a global SaaS, large enterprise, cloud provider, or similarly complex production environmentFamiliarity with BGP or OSPF, data center fabrics, load balancers or ADCs, DDoS protection, VPNs, firewalls, or public‑cloud networkingExperience with IT service management practices, including incident, change, and problem managementExperience improving observability, change safety, capacity management, or operational readiness for production servicesRelevant certifications such as CCNA, CCNP, Azure/AWS/GCP related

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Manager, Network Reliability and Resiliency
Manager, Network Reliability and Resiliency

Servicenow • Toronto

On-site
CAD 126,000 - 220,000
Health plans
RRSP plan with company match
ESPP
+3
Spécialiste Réseaux (Senior)
Spécialiste Réseaux (Senior)

Groupe SII • Montreal (administrative region)

On-site
CAD 90,000 - 120,000
Spécialiste Réseaux (Senior)
Spécialiste Réseaux (Senior)

Groupe SII • Canada

On-site
CAD 90,000 - 130,000
Network Reliability & Resiliency Lead
Network Reliability & Resiliency Lead

Servicenow • Toronto

Hybrid
CAD 126,000 - 220,000
Health plans
RRSP plan with company match
ESPP
+3
Spécialiste Réseau / Network Specialist
Spécialiste Réseau / Network Specialist

eStruxture Data Centers • Montreal (administrative region)

On-site
CAD 85,000 - 105,000
Site Reliability Manager, Data Center Networking, SRE
Site Reliability Manager, Data Center Networking, SRE

Google Canada • Southwestern Ontario

On-site
CAD 216,000 - 221,000
Site Reliable Engineer (Canada - Remote)
Site Reliable Engineer (Canada - Remote)

AXON Networks • Toronto

On-site
CAD 110,000 - 152,000
Principal Network Systems Engineer
Principal Network Systems Engineer

VSG Recruitment • Hamilton

On-site
CAD 120,000 - 160,000
Manager, Systems Engineering
Manager, Systems Engineering

ServiceNow • Toronto

On-site
CAD 126,000 - 220,000
Health plans
RRSP with company match
ESPP
+3
Gestionnaire, Fiabilité et résilience du réseau
Gestionnaire, Fiabilité et résilience du réseau

Servicenow • Montreal (administrative region)

On-site
CAD 120,000 - 180,000
Certifications industrielles (RHCE/CCN
ITIL