SRE Lead: AI-Driven Observability & Auto-Healing

Falconsmartit

Hove

Hybrid

GBP 90,000 - 130,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Falconsmartit in the United Kingdom invites an experienced Site Reliability Engineer to join our hybrid team in Hove. You will drive modernization of IT operations by implementing observability practices, reducing toil, and leading automation initiatives across cloud, containers, and AI‑driven monitoring.

The role requires deep SRE knowledge, hands‑on expertise with Dynatrace and Datadog, and proficiency in Python and Ansible, plus AWS and Azure experience.

Qualifications

  • Experience implementing SRE and observability in large-scale environments.
  • Strong automation and scripting capabilities.
  • Certifications in cloud platforms and observability areas preferred.

Responsibilities

  • Collaborate with Product Engineering to modernize IT operations and reduce toil.
  • Architect and deploy observability platforms to monitor health, performance and reliability.
  • Develop AI-driven alerting and anomaly detection to reduce MTTR/MTTD.
  • Define SLOs, SLIs and error budgets and enforce SRE best practices.
  • Create an AIOPS roadmap to improve operational efficiency.
  • Automate repetitive tasks and incident responses for autonomous operations.
  • Lead incident management and root cause analysis with automated tooling.
  • Partner with teams to enable shift-left reliability and guide SRE adoption.
  • Mentor teams on SRE principles and promote culture of reliability.

Skills

SRE principles
Observability
Automation scripting
Python
Ansible
AWS
Azure
Docker
Kubernetes
AI/ML for reliability
CI/CD
Incident management

Tools

Dynatrace
Datadog
Gremlin
Chaos Monkey
Docker
Kubernetes
AI/ML tooling

Job description

Falconsmartit in the United Kingdom invites an experienced Site Reliability Engineer to join our hybrid team in Hove. You will drive modernization of IT operations by implementing observability practices, reducing toil, and leading automation initiatives across cloud, containers, and AI‑driven monitoring.

The role requires deep SRE knowledge, hands‑on expertise with Dynatrace and Datadog, and proficiency in Python and Ansible, plus AWS and Azure experience.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

SRE
SRE

Technopride Ltd • Hove

Hybrid
GBP 60,000 - 80,000
Site Reliability Engineer
Site Reliability Engineer

Falconsmartit • Hove

Hybrid
GBP 90,000 - 130,000
Senior SRE: AWS, Dynatrace & Observability Lead (Remote)
Senior SRE: AWS, Dynatrace & Observability Lead (Remote)

SF Partners • Birmingham

On-site
GBP 99,000 - 121,000
Electric company car scheme
6% private pension
Death in service / income protection
+2
SRE Architect
SRE Architect

Hitachi • Greater London

On-site
GBP 42,000 - 70,000
Senior SRE Architect: Reliability & Self-Healing at Scale
Senior SRE Architect: Reliability & Self-Healing at Scale

Hitachi • Greater London

On-site
GBP 42,000 - 70,000
SRE & Reliability Lead — AI-Ops & Observability
SRE & Reliability Lead — AI-Ops & Observability

LexisNexis Risk Solutions • Carshalton

On-site
GBP 90,000 - 130,000
SRE Architect: Cloud & Data Reliability Leader
SRE Architect: Cloud & Data Reliability Leader

Hitachids • Greater London

On-site
GBP 90,000 - 140,000
Dynatrace SRE & Observability Lead
Dynatrace SRE & Observability Lead

SF Partners • United Kingdom

On-site
GBP 70,000 - 110,000
SRE Lead: Observability & Platform Reliability
SRE Lead: Observability & Platform Reliability

Camwebdir • United Kingdom

Remote
GBP 90,000 - 120,000
23 days holiday
Birthday day off
Private medical insurance
+5
Cloud SRE — Observability, CI/CD & Automation
Cloud SRE — Observability, CI/CD & Automation

Viasat • Greater London

On-site
GBP 70,000 - 100,000