Senior SRE Transformation / Reliability Engineering Leader
We are seeking an experienced SRE Transformation / Reliability Engineering Leader to lead the evolution of a traditional production support organization into a modern, data-driven reliability engineering function. This is a transformation leadership role—not an SRE Engineer position. The ideal candidate will have a proven track record of implementing SRE practices at enterprise scale, establishing SLI/SLO and error-budget frameworks, reducing operational toil, improving reliability, and influencing senior leadership on how reliability is measured, managed, and funded.
Responsibilities
- Lead the transformation of traditional Production Support/Operations organizations into a mature Site Reliability Engineering and Reliability Engineering function.
- Define and implement enterprise-level SLI, SLO, and error-budget frameworks across applications, platforms, and business services.
- Establish reliability metrics and reporting that provide clear visibility into availability, performance, incidents, MTTR, operational risk, and customer impact.
- Partner with engineering, application support, infrastructure, DevOps, and business stakeholders to embed reliability engineering practices into day-to-day operations.
- Identify opportunities to reduce operational toil, recurring incidents, manual processes, and reliability risk through automation and engineering improvements.
- Develop and execute strategies that measurably improve MTTR, incident frequency, system resilience, availability, and operational efficiency.
- Establish reliability maturity models, governance, standards, and best practices across a large enterprise environment.
- Use SLOs, error budgets, incident trends, and operational data to drive engineering and business decisions.
- Present reliability performance, transformation progress, risks, and investment recommendations to senior executives and leadership teams.
- Influence organizational leadership to change how reliability is measured, prioritized, and funded.
- Drive cross-functional initiatives involving engineering, operations, architecture, infrastructure, and application teams.
- Establish a culture focused on engineering reliability rather than reactive production support.
- Define and communicate measurable business outcomes from the SRE transformation.
Required Qualifications
- 8+ years of experience in technology, engineering, operations, production support, SRE, or related disciplines.
- Demonstrated experience leading an SRE or Reliability Engineering transformation within a large, complex enterprise.
- Proven experience implementing SLI/SLO and error-budget frameworks at scale.
- Strong understanding of SRE principles, reliability engineering, incident management, observability, automation, and operational excellence.
- Demonstrated ability to deliver measurable improvements in areas such as:
- MTTR
- Operational toil
- Automation
- Resilience
- Operational risk
- Experience transforming traditional Production Support, Application Support, or Operations organizations into engineering-focused reliability functions.
- Strong executive communication and stakeholder-management skills, with the ability to influence senior leadership and drive organizational change.
- Excellent written and verbal communication skills.
- Ability to translate complex technical reliability concepts into business outcomes and investment decisions.
Preferred Qualifications
- Experience within a large regulated enterprise, preferably financial services, banking, insurance, healthcare, or another highly regulated industry.
- Experience working with enterprise-scale application portfolios and multiple technology teams.
- Experience establishing reliability governance and maturity frameworks across a business unit or enterprise.
- Experience with observability and monitoring platforms such as Splunk, Dynatrace, AppDynamics, Prometheus, Grafana, or similar technologies.
- Experience driving automation and DevOps initiatives as part of an SRE transformation.
- Experience managing or influencing distributed engineering and operations teams.
What We're Looking For
The strongest candidate will be someone who can demonstrate:
"I took an organization that primarily measured success by keeping production running, transformed it into an engineering-driven reliability organization, established SLOs and error budgets, reduced incidents and toil, improved MTTR and resilience, and demonstrated those improvements to executive leadership."
This position is best suited for a transformation leader with substantial SRE experience, rather than a hands-on SRE Engineer whose primary focus has been operating production systems.