Senior SRE, Observability & Cloud Reliability

United States Digital Space LLC

Greater London

Hybrid

GBP 120,000 - 170,000

Full time

8 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Hybrid work up to 3 days per week

Job summary

the company is seeking a Principal Site Reliability Engineer, Infrastructure Observability to guide a team of SREs focused on observability, reliability, and scalable cloud/on‑prem solutions. The role requires hands‑on expertise and collaboration with diverse partners to drive measurable improvements.

The ideal candidate has extensive cloud experience, DevOps/SRE leadership, and strong automation skills, with a track record in designing resilient systems and implementing effective monitoring and

Qualifications

  • Bachelor's degree or equivalent with 10+ years cloud infrastructure experience
  • 5+ years AWS experience
  • 5+ years DevOps/SRE experience
  • Experience with chaos engineering at scale
  • Experience implementing SLOs/SLIs and error budgets
  • Proficient in scripting and automation
  • Strong cross-team communication skills

Responsibilities

  • Lead design of technology solutions to prevent or minimize service disruptions
  • Drive automation to reduce technology failures and outages
  • Mentor and grow the SRE team and promote blameless post-mortems
  • Oversee observability strategy across the technology stack including dashboards and logging
  • Ensure 24x7 monitoring and incident response readiness
  • Collaborate with stakeholders to define target state architecture and roadmaps
  • Maintain awareness of industry trends and best practices
  • Balance strategic goals with pragmatic execution

Skills

Python
Java
Go
.NET Core
SQL Databases

Education

Bachelor's degree or equivalent

Tools

New Relic
Elastic Stack
Prometheus
Grafana
Splunk
AWS
Ansible
Terraform

Job description

the company is seeking a Principal Site Reliability Engineer, Infrastructure Observability to guide a team of SREs focused on observability, reliability, and scalable cloud/on‑prem solutions. The role requires hands‑on expertise and collaboration with diverse partners to drive measurable improvements.

The ideal candidate has extensive cloud experience, DevOps/SRE leadership, and strong automation skills, with a track record in designing resilient systems and implementing effective monitoring and

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior SRE – Infrastructure Observability (Hybrid)
Senior SRE – Infrastructure Observability (Hybrid)

T. Rowe Price • Greater London

Hybrid
GBP 120,000 - 180,000
Senior SRE: Cloud Reliability & Observability Lead
Senior SRE: Cloud Reliability & Observability Lead

Pulse Recruit • Greater London

Hybrid
GBP 65,000 - 85,000
Principal Site Reliability Engineer, Infrastructure Observability
Principal Site Reliability Engineer, Infrastructure Observability

United States Digital Space LLC • Greater London

Hybrid
GBP 120,000 - 170,000
Hybrid work up to 3 days per week
Senior SRE: Cloud Reliability & Observability Leader
Senior SRE: Cloud Reliability & Observability Leader

Omilia • United Kingdom

On-site
GBP 90,000 - 130,000
Fixed compensation
Long-term employment
Professional growth
+1
SRE: Cloud Reliability, DevSecOps & Observability
SRE: Cloud Reliability, DevSecOps & Observability

SR2 REC LTD • Greater London

Hybrid
GBP 70,000 - 110,000
SRE Architect: Reliability, Observability & Automation Lead
SRE Architect: Reliability, Observability & Automation Lead

Hitachi Digital Services • Greater London

On-site
GBP 90,000 - 150,000
Senior Cloud Platform SRE & Observability Lead
Senior Cloud Platform SRE & Observability Lead

NICE • Southampton

On-site
GBP 60,000 - 80,000
Senior Site Reliability Engineer: Observability & Cloud
Senior Site Reliability Engineer: Observability & Cloud

Infinity Quest • United Kingdom

On-site
GBP 85,000 - 120,000
Senior SRE: Observability & Platform Reliability
Senior SRE: Observability & Platform Reliability

Selby Jennings • Greater London

On-site
GBP 70,000 - 90,000
SRE
SRE

Technopride Ltd • Hove

Hybrid
GBP 60,000 - 80,000