Lead SRE - Azure and GCP

Hackajob Ltd

Glasgow

On-site

GBP 90,000 - 110,000

Full time

2 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

JPMorganChase is seeking a Lead Site Reliability Engineer to oversee global Google Cloud environments and drive SRE excellence.

You will apply mastery of application, data, and infrastructure disciplines, using enterprise AI to accelerate incident response while upholding security and sensitivity requirements. This role supports a follow-the-sun model and cross-functional collaboration.

Qualifications

  • Expertise in Google/Azure cloud in production environments.
  • Strong container tech: Docker, Kubernetes, Helm, GKE.
  • Programming in Python, Shell, or Go; REST APIs knowledge.
  • Hands-on with monitoring/ops tools: Datadog, Prometheus, Grafana, etc.
  • Experience applying enterprise AI capabilities to SRE workflows with validation.
  • Ability to set guardrails for AI-assisted operations and ensure security.
  • Knowledge of Google Cloud governance and cost management.
  • Experience with Agile, CI/CD, IaC (Terraform), Git.
  • Google Cloud certification or equivalent experience.
  • Familiar with Windows and Linux (Red Hat/Ubuntu).
  • Familiar with AI/ML frameworks for AIOps.

Responsibilities

  • Lead and implement SRE frameworks for Google Cloud environments.
  • Maintain high SLOs through operational excellence.
  • SRE governance of AI-assisted triage and post-incident analysis.
  • Develop and improve engineering documentation.
  • Provide technical supervision and problem resolution.
  • Champion DevOps model for automation across platforms.
  • Collaborate with teams across the firm to achieve goals.

Skills

Google Cloud
Azure
Kubernetes
Docker
Python
CI/CD
Terraform
Jenkins
Git
Observability
LLM/AI

Tools

GKE
Helm
Prometheus
Grafana
Datadog
Splunk
Elasticsearch
Agentic AI SDKs
GitHub Copilot

Job description

Salary: £100,000 - 100,000 per year

Requirements:
  • Google and Azure cloud expertise in a mission-critical production environment
  • Strong understanding of container technologies such as Docker, Kubernetes, GKE, and Helm
  • Programming experience in Python, shell scripting, or Go, with a good understanding of REST APIs
  • Hands-on experience with cloud-based technologies and tools for deployment, monitoring, and operations, such as Google Observability, Azure Monitor, Datadog, Prometheus, Splunk, Elasticsearch, and Grafana
  • Demonstrated experience using enterprise-authorized AI capabilities in the work environment to improve SRE workflows, with strong validation habits and awareness of data sensitivity
  • Ability to evaluate AI-assisted operational recommendations for correctness and risk, define appropriate guardrails for team usage, and ensure outcomes align with resiliency and security expectations
  • Strong understanding of Google Cloud governance, compliance, and cost management
  • Strong working knowledge of modern development technologies and tools such as Agile, CI/CD, Git, Infrastructure as Code, Terraform, and Jenkins
  • Google Cloud certification or equivalent technical experience in the public cloud
  • Good understanding of Agentic AI SDKs and GitHub Copilot skills
  • Good understanding of operating systems such as Windows and Linux (Red Hat/Ubuntu)
  • Good understanding of LLM and other AI/ML frameworks that can be used in AIOps
Responsibilities:
  • Lead and implement SRE frameworks to support global Google Cloud environments and ensure the highest level of SLOs through operational excellence
  • Apply mastery of application, data, infrastructure, and Agentic AI disciplines
  • Use enterprise-authorized AI capabilities within our work environment to accelerate major-incident triage, troubleshooting, and post-incident analysis, validating outputs and handling operational data according to sensitivity and security requirements
  • Provide support to develop and improve the quality of technical engineering documentation
  • Provide technical supervision, oversight, and problem resolution for engineering activities
  • Champion a DevOps model so that services are automated and elastic across all platforms
  • Work in partnership with colleagues throughout the firm and lead collaborative teams to achieve common goals
Technologies:
  • Agentic AI
  • AI
  • Azure
  • CI/CD
  • Cloud
  • Copilot
  • Datadog
  • DevOps
  • Docker
  • ElasticSearch
  • Git
  • GitHub
  • Grafana
  • Helm
  • Support
  • Jenkins
  • Kubernetes
  • LLM
  • Linux
  • Prometheus
  • Python
  • REST
  • Security
  • Splunk
  • Terraform
  • Ubuntu
  • Windows
  • Network
More:

We are JPMorganChase, a global leader in financial services providing strategic advice and products to the worlds most prominent corporations, governments, wealthy individuals, and institutional investors. We are hiring for a Lead Site Reliability Engineer role within our Google Cloud Site Reliability Engineering team, part of our Infrastructure Platform - Cloud Foundational Services SRE organization, operating in a global follow-the-sun support model. Our Corporate Technology team develops applications and provides tech support across our corporate functions, supporting areas such as Global Finance, Corporate Treasury, Risk Management, Human Resources, Compliance, Legal, and the Corporate Administrative Office. We value diversity and inclusion, support reasonable accommodations, and strive to build trusted, long-term partnerships with our clients.

last updated 36 week of 2026

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead SRE- Azure & GCP
Lead SRE- Azure & GCP

JPMorganChase • Glasgow

On-site
GBP 90,000 - 130,000
Lead SRE- Azure & GCP
Lead SRE- Azure & GCP

Fairygodboss • Glasgow

On-site
GBP 120,000 - 170,000
Lead SRE- Azure & GCP
Lead SRE- Azure & GCP

Next Frontier Capital • Glasgow

On-site
GBP 90,000 - 140,000
Lead SRE: Azure & GCP with AI-Driven Ops
Lead SRE: Azure & GCP with AI-Driven Ops

Hackajob Ltd • Glasgow

On-site
GBP 90,000 - 110,000
Lead SRE - AWS Platform
Lead SRE - AWS Platform

Hackajob Ltd • Glasgow

On-site
GBP 90,000 - 110,000
Lead SRE: Azure & GCP Cloud Reliability
Lead SRE: Azure & GCP Cloud Reliability

JPMorganChase • Glasgow

On-site
GBP 90,000 - 130,000
Lead SRE - AWS Platform
Lead SRE - AWS Platform

JP Morgan Chase • Glasgow

On-site
GBP 62,000 - 102,000
Lead SRE: Google Cloud & Azure (Global)
Lead SRE: Google Cloud & Azure (Global)

Fairygodboss • Glasgow

On-site
GBP 120,000 - 170,000
Global Cloud SRE Lead - Google Cloud & DevOps Strategy
Global Cloud SRE Lead - Google Cloud & DevOps Strategy

Next Frontier Capital • Glasgow

On-site
GBP 90,000 - 140,000
Lead Site Reliability Engineer - Glasgow
Lead Site Reliability Engineer - Glasgow

JP Morgan Chase • Glasgow

On-site
GBP 62,000 - 102,000