Software Engineering Tech Lead (SRE + AI)

CISCO Systems

Greater London

On-site

GBP 39,000 - 79,000

Full time

6 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Cisco Systems in Greater London is seeking a technically seasoned leader to shape the AI-assisted observability platform. You will define the technical roadmap and architecture for automated incident response and self-healing cloud infrastructure across multi‑cluster environments.

You will build production‑grade AI agents, MCP tool integrations, and deterministic evaluation pipelines to support operational decisions.

Qualifications

  • Bachelor's degree plus 8 years of related experience in Computer Science, Software Engineering, or a related field.
  • Proven record as Technical Lead or Lead SRE delivering distributed, high-availability SaaS platforms at scale.
  • Strong proficiency in Python, Go, Java, or C++ with experience designing microservices, APIs, and production automation.
  • Deep experience with Kubernetes, Docker, and container orchestration in large-scale multi-cluster environments.
  • Proven background in SRE practices, including SLI/SLO design, observability platforms, incident management, and automated RCA.

Responsibilities

  • Define the technical roadmap and architecture for AI-assisted observability, automated incident response, and self-healing cloud infrastructure.
  • Design and build production-grade AI agents, MCP tool integrations, and deterministic evaluation pipelines for automated operational decision support.
  • Architect ingestion and correlation pipelines across distributed logs, metrics, OpenTelemetry traces, change events, and runbooks to accelerate MTTD and MTTR.
  • Develop proactive anomaly detection and Human-in-the-Loop remediation workflows with safety, security, and quality guardrails.
  • Partner with application and infrastructure teams to define SLIs/SLOs, handle error budgets, and lead deep-dive post-incident reviews.
  • Mentor senior and mid-level engineers, establish engineering best practices, and drive alignment across global development and operations teams.
  • Manage priorities and deadlines, communicate progress clearly, and work across teams to turn production needs into reliable software and AI-assisted capabilities.

Skills

Python
Go
Java
C++
Kubernetes
Docker
SRE practices
AI/ML pipelines
OpenTelemetry
CI/CD

Education

Bachelor's degree + related experience

Tools

Terraform
Jenkins
GitHub Actions
AWS
GCP
Azure
Kafka
PostgreSQL
Elasticsearch
Grafana
Prometheus
OpenTelemetry
Redis

Job description

Salary: £39,000 - 79,000 per year

Requirements:
  • Bachelors degree plus 8 years of related experience, Masters degree plus 6 years, or PhD plus 3 years in Computer Science, Software Engineering, or a related technical field.
  • Proven record as a Technical Lead or Lead SRE/Software Engineer delivering distributed, high-availability SaaS platforms at scale.
  • Strong proficiency in Python, Go, Java, or C++ with experience designing microservices, APIs, and production automation.
  • Deep experience with Kubernetes, Docker, and container orchestration in large-scale multi-cluster environments.
  • Proven background in SRE practices, including SLI/SLO design, observability platforms, incident management, and automated RCA.
  • Preferred experience building LLM pipelines, AI agents, Model Context Protocol (MCP) servers/clients, RAG architectures, and evaluation frameworks.
  • Experience with OpenTelemetry (OTel), Prometheus, Grafana, Splunk, ThousandEyes, or distributed tracing systems.
  • Expertise in public cloud providers (AWS, GCP, Azure), Terraform/IaC, and GitOps/CI/CD pipelines such as Jenkins or GitHub Actions.
  • Experience implementing responsible AI guardrails, deterministic fallback logic, and policy-driven remediation engines.
  • Experience with streaming and data platforms such as Kafka, Redis, PostgreSQL, and Elasticsearch or vector databases.
Responsibilities:
  • Define the technical roadmap and architecture for AI-assisted observability, automated incident response, and self-healing cloud infrastructure.
  • Design and build production-grade AI agents, MCP tool integrations, and deterministic evaluation pipelines for automated operational decision support.
  • Architect ingestion and correlation pipelines across distributed logs, metrics, OpenTelemetry traces, change events, and runbooks to accelerate MTTD and MTTR.
  • Develop proactive anomaly detection and Human-in-the-Loop remediation workflows with safety, security, and quality guardrails.
  • Partner with application and infrastructure teams to define SLIs/SLOs, handle error budgets, and lead deep-dive post-incident reviews.
  • Mentor senior and mid-level engineers, establish engineering best practices, and drive alignment across global development and operations teams.
  • Manage priorities and deadlines, communicate progress clearly, and work across teams to turn production needs into reliable software and AI-assisted capabilities.
Technologies:
  • AI
  • AI Agents
  • AWS
  • Architect
  • Azure
  • CI/CD
  • Cloud
  • Cisco
  • Docker
  • ElasticSearch
  • GCP
  • GitHub
  • GitOps
  • Grafana
  • Incident Management
  • Support
  • Java
  • Jenkins
  • Kafka
  • Kubernetes
  • LLM
  • MCP
  • OpenTelemetry
  • PostgreSQL
  • Prometheus
  • Python
  • RAG
  • Redis
  • Security
  • Splunk
  • Terraform
  • microservices
  • Agentic AI
  • Network
More:

We are the Collaboration Technology Group, building and continuously improving the platform services behind Ciscos collaboration products to connect people effortlessly across devices, locations, and time zones. We operate at global scale across numerous datacentres and are a passionate, collaborative team focused on reliability, innovation, and engineering excellence. This full-time Product and Engineering role is part of Cisco, where we are innovating for the AI era and beyond, with a global team, strong collaboration culture, and opportunities to build impactful solutions across the digital world.

last updated 36 week of 2026

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Engineering Tech Lead (SRE + AI)
Software Engineering Tech Lead (SRE + AI)

Cisco Systems, Inc. • Greater London

On-site
GBP 130,000 - 180,000
Software Engineering Tech Lead (SRE + AI)
Software Engineering Tech Lead (SRE + AI)

Cisco • Greater London

On-site
GBP 120,000 - 180,000
Software Engineering Tech Lead (SRE + AI)
Software Engineering Tech Lead (SRE + AI)

624 Cisco International Limited • Greater London

On-site
GBP 110,000 - 150,000
Tech Lead: AI-Driven SRE for Scalable SaaS
Tech Lead: AI-Driven SRE for Scalable SaaS

CISCO Systems • Greater London

On-site
GBP 39,000 - 79,000
Lead Software Engineer
Lead Software Engineer

CISCO Systems • Greater London

On-site
GBP 39,000 - 79,000
Software Engineer – SRE
Software Engineer – SRE

Jobtailor • Greater London

On-site
GBP 90,000 - 120,000
SRE Architect
SRE Architect

Hitachi • Greater London

On-site
GBP 42,000 - 70,000
Solutions Engineer
Solutions Engineer

Cisco • Greater London

On-site
GBP 90,000 - 130,000
AI-Powered Observability Tech Lead
AI-Powered Observability Tech Lead

Cisco • Greater London

On-site
GBP 120,000 - 180,000
Software Engineer AI Engineering
Software Engineer AI Engineering

JP Morgan Chase • Greater London

On-site
GBP 62,000 - 102,000