SaaS Platform Lead: Reliability, Automation & AI Ops

Jobtailor

Greater London

Hybrid

GBP 110,000 - 140,000

Full time

5 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

ITRS is seeking a Senior SaaS Operations Leader to guide the day-to-day reliability and continuous improvement of its Analytics platform. You will direct a hands-on engineering team, set standards for automation, and drive incident response, capacity planning, and observability.

Collaboration with security, product and customer-facing teams is essential to maintain resilience. The role focuses on reducing toil, advancing runbooks, and applying Agentic AI to operational workflows, while ensuring

Qualifications

  • Proven experience leading SaaS hosting teams in hands-on technical leadership roles.
  • Track record of building, leading and improving operational teams.
  • Deep practical knowledge of Kubernetes operations (deployments, upgrades, networking, storage, ingress, secrets, autoscaling, troubleshooting).
  • Strong cloud operations experience across AWS, Azure or GCP.
  • Hands-on experience with Infrastructure as Code, CI/CD, GitOps or similar automation.
  • Experience reducing operational toil through automation, self-service, runbooks and observability.
  • Understanding of Agentic AI or intelligent automation in operations.
  • Technical depth to troubleshoot production issues and contribute directly.
  • Understanding of security/compliance expectations for SaaS operations.
  • Experience collaborating with Forward Deployed Engineering or customer-embedded teams.

Responsibilities

  • Lead day-to-day management, reliability and continuous improvement of the Analytics SaaS platform.
  • Provide technical leadership for a hands-on SaaS engineering team.
  • Set standards, coach engineers and foster ownership, automation and reliability.
  • Resolve production issues, Kubernetes troubleshooting, incident response and deployment automation.
  • Drive automation to reduce toil and manual hand-offs.
  • Explore Agentic AI for operational decision support and runbook execution.
  • Own SaaS operational standards including runbooks, monitoring, and change control.
  • Collaborate across Engineering, Product, Security and Customer-facing teams on operability and security.
  • Manage and develop the Forward Deployed Engineering function.
  • Ensure practices meet regulated-industry expectations for access control and continuity.
  • Report to the Global Head of Platform Engineering.

Skills

SaaS Operations Leadership
Kubernetes Operations
Cloud Operations
Infrastructure as Code
CI/CD
GitOps
Automation
Observability
Incident Management
Change Management
Production Troubleshooting
Security & Compliance understanding
Forward Deployed Engineer experience
ISO 27001 / SOC 2 awareness
Agentic AI / Intelligent Automation

Education

ISO 27001
SOC 2

Tools

Amazon EKS
Azure AKS
Google GKE
Argo CD
Flux

Job description

ITRS is seeking a Senior SaaS Operations Leader to guide the day-to-day reliability and continuous improvement of its Analytics platform. You will direct a hands-on engineering team, set standards for automation, and drive incident response, capacity planning, and observability.

Collaboration with security, product and customer-facing teams is essential to maintain resilience. The role focuses on reducing toil, advancing runbooks, and applying Agentic AI to operational workflows, while ensuring

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Director of Platform Engineering - SaaS & Observability
Director of Platform Engineering - SaaS & Observability

ITRS • Greater London

Hybrid
GBP 150,000 - 190,000
Health Insurance and Dental Coverage
Pension
Flexible Hybrid Working
+6
Platform Engineering Director - SaaS & Observability
Platform Engineering Director - SaaS & Observability

Itrs Insights • Greater London

Hybrid
GBP 140,000 - 180,000
Health Insurance
Pension
Flexible Hybrid Working
+7
SRE & Operations Manager — AI-Driven Reliability
SRE & Operations Manager — AI-Driven Reliability

LexisNexis Risk Solutions • Sutton

Hybrid
GBP 90,000 - 130,000
Global SVP, Platform Operations & AI-Driven Reliability
Global SVP, Platform Operations & AI-Driven Reliability

Trading Technologies International • Greater London

Hybrid
GBP 180,000 - 280,000
Hybrid in-office 3 days/week
25 days PTO per year
Volunteer day
+3
SRE & Operations Leader - Reliability at Scale
SRE & Operations Leader - Reliability at Scale

LexisNexis Risk Solutions • United Kingdom

Remote
GBP 90,000 - 130,000
AI Platform Tech Lead - Observability & SRE
AI Platform Tech Lead - Observability & SRE

Cisco Systems Inc • City of Westminster

On-site
GBP 110,000 - 165,000
GenAI Platform SRE Lead: Reliability, Automation & Scale
GenAI Platform SRE Lead: Reliability, Automation & Scale

Aviva plc • United Kingdom

Hybrid
GBP 65,000 - 74,000
Bonus opportunity
Generous pension up to 14%
29 days holiday + bank holidays (buy/s
+5
SRE & Reliability Lead — AI-Ops, Observability, & On-Call
SRE & Reliability Lead — AI-Ops, Observability, & On-Call

RELX Group • City of Westminster

On-site
GBP 90,000 - 120,000
Ops Team Lead: Reliability & Incident Management
Ops Team Lead: Reliability & Incident Management

Complexio • Hawton

On-site
GBP 90,000 - 120,000
Senior AI Platform SRE: Reliability & Automation
Senior AI Platform SRE: Reliability & Automation

CloudFactory • Reading

On-site
GBP 70,000 - 110,000