Senior Director – Observability | SRE

FashionUnited

Coppell (TX)

On-site

USD 180,000 - 240,000

Full time

11 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Gap Inc. seeks a Senior Director – Observability and SRE to lead reliability of enterprise technology, overseeing Observability, SRE, and Live Sight Insights.

You will partner with engineering, infrastructure, cybersecurity, and product teams to shape the future of operational reliability. This role drives strategy, sets standards for availability and performance, and builds a culture of data-driven improvement, automation, and proactive monitoring across ecommerce, stores, and fulfillment

Qualifications

  • Proven strategic leader with 10+ years transforming operations at scale.
  • Deep expertise in ITIL, SRE, observability practices.
  • Strong technical understanding of infrastructure, cloud, automation, and service management.
  • Ability to translate reliability concepts into business value across levels.
  • Passion for developing people and fostering ownership and continuous improvement.
  • Experience leading large, diverse teams delivering reliability and performance gains.
  • Executive presence with ability to align organizations around reliability goals.

Responsibilities

  • Define and execute enterprise Observability and SRE strategy aligned with business goals.
  • Lead transformation of Unified Observability to reduce MTTR via critical path anomaly detection.
  • Partner with senior leaders to embed reliability metrics into product development and planning.
  • Lead SRE practices across platforms driving automation and proactive monitoring.
  • Establish standards for availability, latency, scalability, and efficiency.
  • Champion reliability by design with observability, capacity planning, and chaos testing.
  • Oversee Mission Control for real-time system monitoring across ecommerce, fulfillment centers, stores.
  • Drive adoption of Live Sight Insights for predictive intelligence on service health.
  • Enable enterprise visibility with dashboards and alerting models.
  • Lead platform governance focusing on reliability, scalability, and usability.
  • Build and develop a global Observability and SRE team with accountability and innovation.
  • Foster data-driven decision making and continuous improvement.
  • Mentor emerging leaders and raise organizational reliability.

Skills

Strategic leadership
ITIL expertise
SRE principles
Cloud operations
Observability

Education

Bachelor's degree in Computer Science or related field

Tools

AWS
Kubernetes

Job description

About the Role

The Senior Director – Observability and SRE, is a strategic leader accountable for ensuring the reliability, availability, and performance of the enterprise technology ecosystem. This role oversees Observability, Site Reliability Engineering (SRE), Visibility and Live Sight Insights. This leader drives operational excellence through a proactive strategy that combines process discipline, automation, observability, and real-time insights. They will partner closely with engineering, infrastructure, cybersecurity, and product teams to build and sustain systems that power Gap Inc.’s digital and in-store experiences. As a thought leader, the Sr. Director will shape the long-term vision for operational reliability, defining modern capabilities, optimizing service performance, and establishing an innovation-driven reliability culture.


What You'll Do

Strategic Leadership & Vision



  • Define and execute the enterprise Observability and SRE strategy, ensuring alignment with businessobjectivesand technology roadmaps.


  • Lead transformation of Unified Observability for end-to-end visibility of systems to actively and proactively reduce mean time to resolve through critical path anomaly detection.


  • Partner with senior technology and business leaders to embed reliability and performance metrics into product development and operational planning.



Operational Excellence & Reliability Engineering



  • Lead Site Reliability Engineering (SRE) practices across platforms and services driving automation, self-healing capabilities, and proactive monitoring to achieve measurable service resiliency improvements.


  • Establishstandards for availability, latency, scalability, and operational efficiency through engineering-driven reliability principles.


  • Champion reliability by design ensuring observability, capacity planning, and chaos testing are core to delivery processes.



Mission Control &Live SightInsights



  • Oversee the Mission Control organization responsible for real time system monitoring, across ecommerce, fulfillment centers, and stores.


  • Drive adoption of Live Sight Insights to create predictive and actionable intelligence on service health and performance trends.


  • Enable enterprise visibility of key metrics through intuitive dashboards and business-impact-based alerting models.


  • Lead a platformgovernancemindset focusing on reliability, scalability, and ease of use.



People Leadership & Culture



  • Build, inspire, and develop a high-performing global Observability and SRE team that embodies accountability, collaboration, and innovation.


  • Foster a culture of data driven decision making, continuous learning, and operational excellence.


  • Serve as a mentor and coach toemergingleaders raising the organizational bar for reliability engineering and service leadership.



Cross-Functional Partnership



  • Work closely with Software Engineering, Infrastructure, Cybersecurity, and Business Technology teams to ensure reliabilityobjectivesare integrated end-to-end.


  • Partner with Enterprise Architecture and Program Management to align technology investments with reliability outcomes.


  • Act as a trusted advisor to executive leadership on reliability strategy, risk posture, andenterprise service health and performance.



Who You Are


  • Proven strategic leader with success driving operational transformation at scale in global, complex environments for more than 10 years.


  • Deepexpertisein ITIL frameworks, SRE principles, and architecture, and modern observability and SRE practices.


  • Strong technical understanding across infrastructure, cloud operations, automation, and service management ecosystems.


  • Exceptional ability to influence at all levels translating technical reliability concepts into business impact and strategic value.


  • Passionate about developing people and creating a culture of ownership, reliability, and continuous improvement.


  • Demonstratedtrack recordof leading large, diverse teams and delivering measurable improvements in service reliability, performance, and user satisfaction.


  • A high performing leader operatingwith strategic agility, executive presence, and the ability to build organizational alignment through clarity, accountability, and purpose.


Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Director - Observability & SRE Strategy
Senior Director - Observability & SRE Strategy

Gap Inc. • Coppell (TX)

On-site
USD 180,000 - 280,000
Merchandise discount
Paid time off
401(k) matching
+3
Senior Director, Observability & SRE Strategy
Senior Director, Observability & SRE Strategy

FashionUnited • Coppell (TX)

On-site
USD 180,000 - 240,000
Senior Director – Observability | SRE
Senior Director – Observability | SRE

Gap Inc. • San Francisco (CA)

On-site
USD 250,000 - 420,000
Merchandise discount (50% on brands)
Competitive PTO plans
Volunteer hours (up to 5 per month)
+3
Senior Director, Observability & SRE: Drive Reliability
Senior Director, Observability & SRE: Drive Reliability

Gap Inc. • San Francisco (CA)

On-site
USD 250,000 - 420,000
Merchandise discount (50% on brands)
Competitive PTO plans
Volunteer hours (up to 5 per month)
+3
Site Reliability Engineering Manager
Site Reliability Engineering Manager

O.C. Tanner • Salt Lake City (UT)

On-site
USD 180,000 - 260,000
Senior Director – Observability | SRE
Senior Director – Observability | SRE

Gap Inc. • Coppell (TX)

On-site
USD 180,000 - 280,000
Merchandise discount
Paid time off
401(k) matching
+3
Director, Site Reliability Engineering
Director, Site Reliability Engineering

Jobtailor • California (MO)

On-site
USD 180,000 - 260,000
Site Reliability Engineer – Lead
Site Reliability Engineer – Lead

Jobtailor • Arizona

On-site
USD 140,000 - 230,000
Senior Manager SRE
Senior Manager SRE

Expedite Talent Solutions • United States

Hybrid
USD 130,000 - 160,000
Observability SRE Manager, Apple Services Engineering
Observability SRE Manager, Apple Services Engineering

Socket.dev • Seattle (WA)

On-site
USD 180,000 - 260,000