A leading technology services firm in Charlotte, NC, is seeking a professional to lead the expansion of Site Reliability Engineering (SRE) practices. This role involves evaluating operational workflows, executing proactive processes for reliability and incident management, and upskilling team members through tailored training programs. The ideal candidate will foster a high-performing team culture while collaborating with various stakeholders to enhance operational performance.
Qualifications
Experience in leading SRE practices and team management.
Strong knowledge in operational workflows and incident management.
Ability to prepare training programs and workshops.
Responsibilities
Lead the expansion of SRE practices globally.
Evaluate operational workflows and identify areas for improvement.
Execute a roadmap for transitioning to proactive operational activities.
Job description
Lead the expansion of SRE practices from a small and high performing team to a larger global function incorporating on-premise infrastructure technologies.
Evaluate current operational workflows and RACIs, identify toil and complete assessment of skills across the global team.
Execute a comprehensive roadmap to transition reactive operational day to day activities into proactive, SRE-aligned processes with a focus on reliability, automation, observability, and incident management.
Upskill team members through tailored training programs on SRE principles, cloud operations and automation tools.
Collaborate with architects, platform engineering, ServiceNow developers and application teams to define and implement an observability framework in order to enhance proactive incident detection and reduce MTTR.
Define and implement an automation framework to ensure sustainable, responsible, and effective use of automation to reduce toil and risk.
Define and regularly review SLIs, SLOs, SLAs, error budgets, and incident response processes.
Oversee recruitment, orientation, and professional development of the global SRE team.
Foster a high-performing team culture.
Build strong relationships with internal and external stakeholders.
Prepare and present reports on operational performance.
Oversee incident response and post-incident analysis processes and drive a culture of blameless post-mortems across multiple teams.