AVP Site Reliability Engineer - SRE/Infrastructure/Python/Powershell/AWS/Observability/ITIL - PERM

Scope AT Limited

Greater London

On-site

GBP 90,000 - 130,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Hybrid working
Canary Wharf location

Job summary

Scope AT Limited is seeking an AVP Site Reliability Engineer to lead SRE practices in a Cloud hosted environment. The role focuses on automation, IaC, and production resilience across multi-region platforms.

You will drive SRE methodologies, implement observability improvements, and collaborate with platform operations to ensure stability and optimal deployment processes.

Qualifications

  • Minimum two years applying SRE methodologies within a support team.
  • Experience with on-call duties, incident ownership and root-cause analysis.
  • Strong knowledge of at least one scripting language (Python or Ansible); PowerShell is a plus.
  • Cloud experience with AWS/GCP and multi-environment platforms.
  • Experience with Observability/APM tools (Grafana, Datadog, Dynatrace).

Responsibilities

  • Drive implementation of SRE methodologies and automation to optimize deployment processes.
  • Lead efforts to improve observability, alerting, and capacity planning with SLA/SLO/SLI definitions.
  • Develop secure production code and review peers' code.
  • Enhance GitOps capabilities using Terraform and Ansible Automation Platform.
  • Provide on-call support for cloud and automation issues to maintain production stability.

Skills

SRE methodologies
Python
Ansible
PowerShell
AWS
Observability
GitOps

Tools

Terraform
Ansible Automation Platform
Grafana
Datadog
Dynatrace

Job description

AVP Site Reliability Engineer - SRE/Infrastructure/Python/Powershell/AWS/Observability/ITIL - PERM - Financial Services

Job Purpose:
The role is primarily responsible for developing SRE methodologies and ensuring they are applied to the Cloud hosted environment. In addition, the role will act as a central point of expertise for SRE automation across the Platform Operations team.

  • Responsible for driving the implementation of SRE methodologies, collaborating closely with other infrastructure teams to optimize infrastructure and deployment processes, focusing on automation and operational excellence.
  • Drives continuous improvement in system observability, alerting, and capacity planning through the definition and implementation of SLA, SLOs & SLIs
  • Define and enhance frameworks for Toil identification, analysis & remediation to identify opportunities to eliminate or automate remediation of recurring tasks and issues
  • Develops secure high-quality production code, and reviews and debugs code written by others.
  • Build out and enhance GitOps capabilities for use in the Cloud hosted environments using tools such as Terraform and Ansible Automation Platform
  • Provide on-call support and escalation for Cloud & Automation related issues ensuring that Production stability is the primary requirement.
  • Ensure risks and stability issues in the cloud hosted environment are understood and addressed where possible through SRE best practices as part of any incident postmortems.

Minimum Job-Related Experience Required:

  • Must have strong technical operational support experience within an infrastructure services team performing on-call duties such as handling tickets, owning incidents & investigating their root cause
  • Minimum of 2 years experience applying SRE methodologies within a support team and an understanding of Service Level metrics associated with this.
  • Strong knowledge of at least 1 Scripting language, preferably either Python or Ansible. PowerShell would also be a positive
  • Experience with supporting and building multi environment, multi region platforms with cloud providers such as AWS/GCP and managing them through Infrastructure as Code and GitOps methodologies
  • Experience of Observability/APM tools (eg Grafana/Datadog/Dynatrace).

Permanent Role based in Canary Wharf - Hybrid Working

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer (SRE) - Cloud Engineer - GCP & Azure
Site Reliability Engineer (SRE) - Cloud Engineer - GCP & Azure

Deloitte - Recruitment • Glasgow

Hybrid
GBP 60,000 - 90,000
Site Reliability Engineer
Site Reliability Engineer

DNS INFO LTD • City Of London

On-site
GBP 70,000 - 95,000
SRE / DevOps Engineers
SRE / DevOps Engineers

HCLTech • Greater London

On-site
GBP 60,000 - 80,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Xpertise Recruitment • West Drayton

On-site
GBP 60,000 - 80,000
Windows SRE (PowerShell)
Windows SRE (PowerShell)

Bonhill Partners • England

Hybrid
GBP 85,000 - 100,000
SRE
SRE

Technopride Ltd • Hove

Hybrid
GBP 60,000 - 80,000
Site Reliability Engineer (Sre)
Site Reliability Engineer (Sre)

Focus Cloud • Greater London

Hybrid
GBP 118,000 - 159,000
Site Reliability Engineer
Site Reliability Engineer

Reward Gateway • Greater London

Hybrid
GBP 70,000 - 110,000
Life assurance
Pension
Employee Share Plan
+3
Director of Site Reliability Engineering
Director of Site Reliability Engineering

EPAM Systems • Greater London

Hybrid
GBP 140,000 - 200,000
ESPP
Life assurance
Income protection
+11
Site Reliability Engineer (SRE) / Platform Engineer
Site Reliability Engineer (SRE) / Platform Engineer

Adecco • City Of London

Hybrid
Hybrid work arrangement
Competitive day rate
London-based contract