- We are looking for a Director of Site Reliability Engineering to lead the next phase of our reliability transformation as ServiceNow modernizes toward a cloud-agnostic, cloud-ready production platform
- This leader will own key elements of the SRE operating model across Reliability Engineering, Service Enablement, Service Registry, SLI/SLO standards, reliability governance, automation, AI-enabled operations, and production readiness
- The role will lead a global engineering organization and partner across Product Engineering, Infrastructure, Architecture, Security, Release Engineering, and Customer Support to establish consistent reliability practices across ServiceNow products and services
- The Director will play a critical role in evolving the organization from reactive operations toward an engineering-led SRE model focused on prevention, automation, resilience, and continuous improvement
- Define and execute the SRE strategy and operating model across reliability engineering, service enablement, observability, automation, incident learning, and production readiness
- Lead and develop a global organization of engineering managers, technical leaders, and SREs
- Establish enterprise reliability standards for service ownership, tiering, golden signals, SLIs/SLOs, error budgets, alerting, on-call practices, and service health reviews
- Lead the Service Enablement strategy by establishing minimum reliability requirements and maturity standards for critical services
- Own the Service Registry strategy, improving service ownership, dependency visibility, maturity tracking, and impact-aware operational decision-making
- Drive adoption of SLIs, SLOs, error budgets, and burn-rate alerting across critical services, ensuring teams consistently use reliability signals to manage customer impact
- Build a culture of engineering away toil by turning recurring operational work and incident patterns into automation, self-service, and systemic fixes
- Establish the AI-enabled SRE roadmap, including change-risk assessment, operational insights, remediation recommendations, and policy-driven automation
- Drive reliability and production-readiness strategy across AWS, Azure, and GCP by establishing cloud-agnostic patterns while addressing hyperscaler-specific operational requirements
- Partner with product and platform engineers to design, launch, and operate reliable services throughout the production lifecycle
- Establish launch and production-readiness practices that validate availability, latency, performance, capacity, dependencies, rollback, and recovery before customer impact
- Drive sustainable operations by scaling self-service capabilities, automation platforms, and systemic reliability improvements across engineering teams
- Lead incident response, blameless postmortems, and corrective actions that convert production failures into lasting reliability improvements
- Measure reliability through SLIs, SLOs, error budgets, golden signals, change failure rate, MTTR, capacity health, and toil reduction
- Influence architecture and platform direction to simplify operating models and improve reliability across ServiceNow’s global infrastructure
- Partner with executive and engineering leaders to prioritize reliability investments and drive adoption beyond the direct SRE organization
Benefits
- Generous family leave
- Matched donations
- Annual learning stipends
- Flexible PTO
- Competitive retirement plan
- Paid volunteer time
Proven success leading managers and senior technical leaders across geographically distributed engineering organizationsExperience with service catalogs, service registries, service ownership models, Backstage, CMDB, dependency mapping, or service topologyAbility to use incident, reliability, and operational data to prioritize engineering work and drive systemic improvementsStrong understanding of SLIs/SLOs, error budgets, observability, incident management, reliability governance, and on-call practicesExperience driving automation through orchestration, Infrastructure as Code, self-service platforms, and auto-remediationExperience establishing production-readiness practices for releases, resilience, disaster recovery, infrastructure changes, and cloud migrationsStrong cross-functional influence and executive communication skills12 years of significant leadership experience in Site Reliability Engineering, Production Engineering, Platform Engineering, Cloud Infrastructure, or large-scale distributed systems with a Bachelor’s degree; or 8 years and a Master’s degree; or a PhD with 5 years experience; or equivalent experienceStrong background in cloud infrastructure and modernization across AWS, Azure, and/or GCPDemonstrated success leading SRE, infrastructure, or reliability transformation at scaleUnderstanding of Kubernetes, distributed systems, networking, databases, infrastructure automation, and cloud-native architectureAbility to operate effectively through ambiguity, organizational transformation, and large-scale technical changeFamiliarity with AI-assisted operations, autonomous remediation, or agentic technologies is highly desirableExperience in leveraging or critically thinking about how to integrate AI into work processes, decision-making, or problem-solving. This may include using AI-powered tools, automating workflows, analyzing AI-driven insights, or exploring AI’s potential impact on the function or industry