AVP Site Reliability Engineer - SRE/Infrastructure/Python/Powershell/AWS/Observability/ITIL - PERM - Financial Services
Job Purpose:
The role is primarily responsible for developing SRE methodologies and ensuring they are applied to the Cloud hosted environment. In addition, the role will act as a central point of expertise for SRE automation across the Platform Operations team.
- Responsible for driving the implementation of SRE methodologies, collaborating closely with other infrastructure teams to optimize infrastructure and deployment processes, focusing on automation and operational excellence.
- Drives continuous improvement in system observability, alerting, and capacity planning through the definition and implementation of SLA, SLOs & SLIs
- Define and enhance frameworks for Toil identification, analysis & remediation to identify opportunities to eliminate or automate remediation of recurring tasks and issues
- Develops secure high-quality production code, and reviews and debugs code written by others.
- Build out and enhance GitOps capabilities for use in the Cloud hosted environments using tools such as Terraform and Ansible Automation Platform
- Provide on-call support and escalation for Cloud & Automation related issues ensuring that Production stability is the primary requirement.
- Ensure risks and stability issues in the cloud hosted environment are understood and addressed where possible through SRE best practices as part of any incident postmortems.
Minimum Job-Related Experience Required:
- Must have strong technical operational support experience within an infrastructure services team performing on-call duties such as handling tickets, owning incidents & investigating their root cause
- Minimum of 2 years experience applying SRE methodologies within a support team and an understanding of Service Level metrics associated with this.
- Strong knowledge of at least 1 Scripting language, preferably either Python or Ansible. PowerShell would also be a positive
- Experience with supporting and building multi environment, multi region platforms with cloud providers such as AWS/GCP and managing them through Infrastructure as Code and GitOps methodologies
- Experience of Observability/APM tools (eg Grafana/Datadog/Dynatrace).
Permanent Role based in Canary Wharf - Hybrid Working