Director, Site Reliability Engineering

Walgreens

Deerfield (IL)

On-site

USD 150,000 - 241,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Walgreens is seeking a Director of Site Reliability Engineering to lead a global SRE organization. You will set strategy, drive reliability across cloud and on‑premise platforms, and partner with Engineering, Security, and Product teams to reduce risk and improve service quality.

You will establish SLIs/SLOs, promote automation, and foster a culture of ownership and continuous improvement, influencing architecture and operational practices at scale.

Qualifications

  • Bachelor’s degree in Computer Science, Engineering, or related field OR equivalent work experience.
  • 10+ years of experience in software engineering, infrastructure, SRE, or operations roles with increasing scope and complexity.
  • 5+ years of people leadership experience, including leading managers or senior technical leaders.
  • Experience operating and scaling highly available, distributed systems in cloud or hybrid environments.
  • Demonstrated expertise in incident management, reliability engineering practices, and production operations.
  • At least 5 years of experience contributing to financial decisions in the workplace.
  • Willing to travel up to/at least 10% of the time for business purposes (within state and out of state).

Responsibilities

  • Leads and develops a team of Site Reliability Engineers, including managers and senior technical leaders.
  • Defines and executes the enterprise SRE strategy, aligning reliability goals with business priorities.
  • Establishes clear ownership models, engagement patterns, and operating rhythms between SRE, Engineering, and Infrastructure teams.
  • Builds organizational capability through hiring, coaching, succession planning, and skills development.
  • Oversees the reliability, availability, and performance of mission critical platforms and services across cloud and hybrid environments.
  • Drives the definition, adoption, and monitoring of SLIs, SLOs, and error budgets.
  • Leads improvements in incident response, root cause analysis, and post-incident learning.
  • Ensures effective on-call models, escalation paths, and operational readiness practices.

Skills

Site Reliability Engineering
People leadership
Cloud/hybrid architectures
Incident management
Observability

Education

Bachelor's degree in CS/Engineering
Master's degree (preferred)

Tools

CI/CD tooling
Automation frameworks
Observability platforms

Job description

Job Summary

The Director, Site Reliability Engineering (SRE) is a senior technical and people leader responsible for ensuring the reliability, availability, scalability, and performance of critical enterprise platforms and services. This role leads the SRE organization, setting the strategy and operating model for reliability engineering practices across cloud and on-premise environments. This leader partners closely with Engineering, Infrastructure, Architecture, Security, and Product teams to reduce operational risk, improve system resilience, and enable teams to deliver high-quality services at scale. The Director drives adoption of SRE principles including service level objectives (SLOs), error budgets, observability, automation, and incident management excellence, while building a strong culture of ownership, learning, and continuous improvement. Outcomes directed have a moderate to significant impact on the organization’s short- and long-term results, customer experience, and operational stability. Decisions have moderate to significant impact across technology platforms and services.

Job Responsibilities
  • Leads and develops a team of Site Reliability Engineers, including managers and senior technical leaders, fostering a culture of accountability, learning, and operational excellence.
  • Defines and executes the enterprise SRE strategy, aligning reliability goals with business priorities and technology roadmaps.
  • Establishes clear ownership models, engagement patterns, and operating rhythms between SRE, Engineering, and Infrastructure teams.
  • Builds organizational capability through hiring, coaching, succession planning, and skills development.
  • Oversees the reliability, availability, and performance of mission critical platforms and services across cloud and hybrid environments.
  • Drives the definition, adoption, and monitoring of service level indicators (SLIs), service level objectives (SLOs), and error budgets.
  • Leads efforts to improve incident response, root cause analysis, and post incident learning to reduce repeat issues and operational toil.
  • Ensures effective on call models, escalation paths, and operational readiness practices are in place.
  • Champions automation to reduce manual work, improve recovery times, and increase system scalability and resilience.
  • Oversees observability capabilities, including monitoring, logging, tracing, alerting, and dashboards, to proactively detect and resolve issues.
  • Partners with Engineering and Architecture to influence system design for reliability, scalability, and fault tolerance.
  • Drives continuous improvement in deployment safety, capacity planning, and change management practices.
  • Collaborates with Product, Engineering, Infrastructure, Security, and Vendor partners to balance innovation velocity with operational stability.
  • Provides executive level visibility into reliability posture, risks, trends, and improvement initiatives.
  • Influences standards, policies, and best practices related to reliability, availability, and operational excellence.
About Walgreens

Founded in 1901, Walgreens (www.walgreens.com) has a storied heritage of caring for communities for generations and proudly serves nearly 9 million customers and patients each day across its approximately 8,500 stores throughout the U.S. and Puerto Rico, and leading omni channel platforms. Walgreens has approximately 220,000 team members, including nearly 90,000 healthcare service providers, and is committed to being the first choice for retail pharmacy and health services, building trusted relationships that create healthier futures for customers, patients, team members and communities.

Basic Qualifications
  • Bachelor’s degree in Computer Science, Engineering, or a related field OR equivalent work experience.
  • 10+ years of experience in software engineering, infrastructure, SRE, or operations roles with increasing scope and complexity.
  • 5+ years of people leadership experience, including leading managers or senior technical leaders.
  • Demonstrated experience operating and scaling highly available, distributed systems in cloud or hybrid environments.
  • Strong background in incident management, reliability engineering practices, and production operations.
  • At least 5 years of experience contributing to financial decisions in the workplace.
  • At least 5 years of direct leadership, indirect leadership and/or cross functional team leadership.
  • Willing to travel up to/at least 10% of the time for business purposes (within state and out of state).
Preferred Qualifications
  • Master’s degree in Computer Science, Engineering, or related field.
  • Experience implementing SRE practices at enterprise scale.
  • Deep knowledge of cloud platforms (e.g., Azure), containerization, CI/CD, and infrastructure automation.
  • Experience with modern observability platforms and automation frameworks.
  • Proven ability to influence senior leaders and drive cultural change toward reliability and operational ownership.

We will consider employment of qualified applicants with arrest and conviction records.

The Salary below is being provided to promote pay transparency and equal employment opportunities at Walgreens. The actual hourly salary within this range that you will be offered will depend on a variety of factors including geography, skills and abilities, education, experience and other relevant factors. This role will remain open until filled.

Salary Range: $150000 - $240,625.00 / Salaried

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Director of Reliability & SRE Strategy
Director of Reliability & SRE Strategy

Walgreens • Deerfield (IL)

On-site
USD 150,000 - 241,000
Lead, Site Reliability Engineer
Lead, Site Reliability Engineer

CardWorks • Pittsburgh

Hybrid
USD 146,000 - 163,000
Competitive Pay
Medical, Dental, and Vision Benefits
401(k) Plan with Company Match
+1
Senior Software Engineer - SRE, Retail and Pharmacy
Senior Software Engineer - SRE, Retail and Pharmacy

CVS Health • Richardson (TX)

On-site
USD 93,000 - 204,000
Principal Site Reliability Engineer
Principal Site Reliability Engineer

ViziRecruiter,LLC. • Quincy (MA)

Hybrid
USD 146,000 - 221,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

ViziRecruiter,LLC. • Salisbury (NC)

Hybrid
USD 125,000 - 188,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

ViziRecruiter,LLC. • Chicago (IL)

Hybrid
USD 125,000 - 188,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

ViziRecruiter,LLC. • Quincy (MA)

Hybrid
USD 125,000 - 188,000
Senior Site Reliability Engineer (SRE)
Senior Site Reliability Engineer (SRE)

Practice by Numbers • United States

On-site
USD 120,000 - 160,000
High ownership and autonomy
Strong engineering culture
Impactful work on healthcare infrastructure
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Hobbsnews • Jersey City (NJ)

On-site
USD 153,000 - 192,000
Benefits eligible
Discretionary incentive plan
Staff Engineer - SRE, Retail and Pharmacy
Staff Engineer - SRE, Retail and Pharmacy

CVS Health • Woonsocket (RI)

On-site
USD 118,000 - 261,000
Bonus program
Equity awards
Comprehensive benefits