Tech Lead: AI-Driven SRE for Scalable SaaS

CISCO Systems

Greater London

On-site

GBP 39,000 - 79,000

Full time

5 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Cisco Systems in Greater London is seeking a technically seasoned leader to shape the AI-assisted observability platform. You will define the technical roadmap and architecture for automated incident response and self-healing cloud infrastructure across multi‑cluster environments.

You will build production‑grade AI agents, MCP tool integrations, and deterministic evaluation pipelines to support operational decisions.

Qualifications

  • Bachelor's degree plus 8 years of related experience in Computer Science, Software Engineering, or a related field.
  • Proven record as Technical Lead or Lead SRE delivering distributed, high-availability SaaS platforms at scale.
  • Strong proficiency in Python, Go, Java, or C++ with experience designing microservices, APIs, and production automation.
  • Deep experience with Kubernetes, Docker, and container orchestration in large-scale multi-cluster environments.
  • Proven background in SRE practices, including SLI/SLO design, observability platforms, incident management, and automated RCA.

Responsibilities

  • Define the technical roadmap and architecture for AI-assisted observability, automated incident response, and self-healing cloud infrastructure.
  • Design and build production-grade AI agents, MCP tool integrations, and deterministic evaluation pipelines for automated operational decision support.
  • Architect ingestion and correlation pipelines across distributed logs, metrics, OpenTelemetry traces, change events, and runbooks to accelerate MTTD and MTTR.
  • Develop proactive anomaly detection and Human-in-the-Loop remediation workflows with safety, security, and quality guardrails.
  • Partner with application and infrastructure teams to define SLIs/SLOs, handle error budgets, and lead deep-dive post-incident reviews.
  • Mentor senior and mid-level engineers, establish engineering best practices, and drive alignment across global development and operations teams.
  • Manage priorities and deadlines, communicate progress clearly, and work across teams to turn production needs into reliable software and AI-assisted capabilities.

Skills

Python
Go
Java
C++
Kubernetes
Docker
SRE practices
AI/ML pipelines
OpenTelemetry
CI/CD

Education

Bachelor's degree + related experience

Tools

Terraform
Jenkins
GitHub Actions
AWS
GCP
Azure
Kafka
PostgreSQL
Elasticsearch
Grafana
Prometheus
OpenTelemetry
Redis

Job description

Cisco Systems in Greater London is seeking a technically seasoned leader to shape the AI-assisted observability platform. You will define the technical roadmap and architecture for automated incident response and self-healing cloud infrastructure across multi‑cluster environments.

You will build production‑grade AI agents, MCP tool integrations, and deterministic evaluation pipelines to support operational decisions.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Engineering Tech Lead (SRE + AI)
Software Engineering Tech Lead (SRE + AI)

CISCO Systems • Greater London

On-site
GBP 39,000 - 79,000
Software Engineering Tech Lead (SRE + AI)
Software Engineering Tech Lead (SRE + AI)

Cisco Systems, Inc. • Greater London

On-site
GBP 130,000 - 180,000
AI-Powered Production Reliability Tech Lead
AI-Powered Production Reliability Tech Lead

624 Cisco International Limited • Greater London

On-site
GBP 110,000 - 150,000
Software Engineering Tech Lead (SRE + AI)
Software Engineering Tech Lead (SRE + AI)

624 Cisco International Limited • Greater London

On-site
GBP 110,000 - 150,000
AI-Driven Site Reliability Engineer - Scalable Cloud Infra
AI-Driven Site Reliability Engineer - Scalable Cloud Infra

Cisco • Greater London

On-site
GBP 90,000 - 130,000
AI-Powered Production Intelligence Tech Lead
AI-Powered Production Intelligence Tech Lead

Cisco Systems, Inc. • Greater London

On-site
GBP 130,000 - 180,000
SRE & Operations Manager — AI-Driven Reliability
SRE & Operations Manager — AI-Driven Reliability

LexisNexis Risk Solutions • Sutton

Hybrid
GBP 90,000 - 130,000
SRE & Reliability Lead — AI-Ops & Observability
SRE & Reliability Lead — AI-Ops & Observability

LexisNexis Risk Solutions • Carshalton

On-site
GBP 90,000 - 130,000
SRE & Reliability Manager — AI-Ops, Observability
SRE & Reliability Manager — AI-Ops, Observability

RELX Group • Carshalton

On-site
GBP 90,000 - 120,000
Senior SRE: Kubernetes, GitOps & AI-Driven Reliability
Senior SRE: Kubernetes, GitOps & AI-Driven Reliability

Cisco Systems, Inc. • City Of London

Hybrid
GBP 90,000 - 130,000