Director, Site Reliability Engineering - Incident Management

Yum! Brands

Plano (TX)

On-site

USD 180,000 - 240,000

Full time

6 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Yum! Brands is seeking a Director, Site Reliability Engineering - Incident Management to own the enterprise Reliability practice across Byte, KFC and Taco Bell digital platforms, including strategy, governance and delivery.

This leader also owns the reliability platform, tooling, and capabilities that enable engineers and markets to achieve high availability and resilient services. The role partners across Engineering, Product, Infrastructure and Brand leadership to drive operational excellence,

Responsibilities

  • Provide strategic leadership for Incident Management across Byte, KFC, and Taco Bell digital platforms, ensuring high availability and operational excellence.
  • Establish vision and governance for enterprise Incident Management, driving standardized processes and best practices across brands.
  • Drive operational readiness for major product launches and promotions through change management, risk assessment, and cross-functional planning.
  • Define and monitor SLIs/SLAs, MTTR, incident trends, and executive KPIs to improve platform resilience.
  • Own the governance framework for enterprise Incident Management and continuous improvement across all brands.
  • Own observability strategy and platform health measurement in partnership with Platform Engineering.
  • Own the strategy and roadmap for the reliability platform and its self-service capabilities.
  • Provide executive-level communications and updates to Digital & Technology leadership.

Job description

Purpose Of The Position

The Director, Site Reliability Engineering - Incident Management owns the enterprise Reliability practice across Byte, KFC, and Taco Bell digital platforms, including its strategy, standards, governance, and delivery. Incident Management is the most visible part of that practice and sits with this role exclusively. This leader also owns the reliability platform, meaning the tooling, products, and capabilities that make reliability real for engineers and for markets, and serves as the accountable face to brands and markets for reliability process implementation, reporting, and the operational relationship. This leader establishes the strategy, governance, and operational excellence required to ensure highly available, resilient, and customer-centric technology services while developing high-performing teams and partnering across Engineering, Product, Infrastructure, and Brand leadership to continuously improve reliability and business continuity. This leader operates with exceptional diligence and care, bridging deep technical detail with human and business context so that executives, engineers, and restaurant teams experience clear, calm, and trustworthy communication in the moments that matter most.

The Director, Site Reliability Engineering - Incident Management owns the enterprise Reliability practice across Byte, KFC, and Taco Bell digital platforms, including its strategy, standards, governance, and delivery. Incident Management is the most visible part of that practice and sits with this role exclusively. This leader also owns the reliability platform, meaning the tooling, products, and capabilities that make reliability real for engineers and for markets, and serves as the accountable face to brands and markets for reliability process implementation, reporting, and the operational relationship. This leader establishes the strategy, governance, and operational excellence required to ensure highly available, resilient, and customer-centric technology services while developing high-performing teams and partnering across Engineering, Product, Infrastructure, and Brand leadership to continuously improve reliability and business continuity. This leader operates with exceptional diligence and care, bridging deep technical detail with human and business context so that executives, engineers, and restaurant teams experience clear, calm, and trustworthy communication in the moments that matter most.

Scope And Magnitude

~30+ Teams - Byte, KFC and Taco Bell

Global Team - Vietnam, India, Colombia and US

Platform Team - Responsible for all products (Edge, Commerce, POS, KDS, Menu, Portal, etc.)

Position Functions
Strategic Leadership
  • Provide strategic leadership for Incident Management and Technical Operations across Byte, KFC, and Taco Bell digital platforms, ensuring high availability, operational excellence, and a consistent customer experience.
  • Establish the vision, governance, and operating model for enterprise Incident Management, driving standardized processes, tooling, and best practices across multiple brands and technology organizations.
  • Drive enterprise operational readiness for major product launches, restaurant initiatives, promotions, and seasonal events through effective change management, risk assessment, and cross-functional planning.
  • Own the enterprise framework and targets for service level objectives and indicators, set in partnership with the engineering teams that own the services, which remain accountable for achieving them. Define and monitor MTTR, incident trends, service health, operational maturity, and executive KPIs, using data to prioritize investments and improve platform resilience.
  • Own the enterprise Incident Management governance framework, ensuring consistent execution, accountability, and continuous improvement across all brands.
  • Own enterprise observability strategy, setting the standard and direction for how platform health is measured and seen, with Platform Engineering partnering on the underlying platform and instrumentation.
  • Own the strategy and roadmap for the reliability platform, including the tooling, products, and self-service capabilities that deliver reliability to engineering teams and to markets.
  • Provide executive-level communications and operational updates to Digital & Technology leadership,
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Director, Site Reliability Engineering - Incident Management
Director, Site Reliability Engineering - Incident Management

KFC Corporation • Plano (TX)

On-site
USD 180,000 - 240,000
Director, Global Reliability & Incident Management
Director, Global Reliability & Incident Management

KFC Corporation • Plano (TX)

On-site
USD 180,000 - 240,000
Director, SRE Incident Management - Enterprise Reliability
Director, SRE Incident Management - Enterprise Reliability

Yum! Brands • Plano (TX)

On-site
USD 180,000 - 240,000
Senior Director – Observability | SRE
Senior Director – Observability | SRE

FashionUnited • Coppell (TX)

On-site
USD 180,000 - 240,000
Director, Site Reliability Engineering
Director, Site Reliability Engineering

Hirebridge • Denver (CO)

On-site
USD 175,000 - 220,000
Senior Manager, AI Reliability Engineering -Kroger Technology & Digital (P2498)
Senior Manager, AI Reliability Engineering -Kroger Technology & Digital (P2498)

84.51˚ • Cincinnati (OH)

On-site
USD 190,000 - 270,000
Site Reliability Engineering Manager
Site Reliability Engineering Manager

O.C. Tanner • Salt Lake City (UT)

On-site
USD 180,000 - 240,000
Site Reliability Engineering Manager
Site Reliability Engineering Manager

Harvey Nash • Charlotte (NC)

On-site
USD 100,000 - 130,000
Director, Site Reliability Engineering
Director, Site Reliability Engineering

Jobtailor • Minnesota

On-site
USD 170,000 - 210,000
Mgr IT Site Reliability Eng
Mgr IT Site Reliability Eng

Kforce, Inc • Town of Florida (NY)

On-site
USD 110,000 - 140,000