Software Engineer

Cisco Systems, Inc

Greater London

Híbrido

GBP 90.000 - 150.000

Jornada completa

Hace 11 días
Generador de candidaturas

No envíes un currículum genérico — crea un currículum y una carta de presentación adaptados a este puesto concreto.

Supera los filtros ATS

Descripción de la vacante

Cisco Systems, Inc is seeking a Software Engineer to design and deploy AI-powered production intelligence capabilities at scale. You will combine Site Reliability Engineering with agentic AI to monitor, diagnose, and auto-remediate global SaaS infrastructure.

You will build AI agents, integrate MCP tooling, and implement end-to-end telemetry across distributed systems, driving reliability and rapid incident response.

Formación

  • Experience delivering distributed, high-availability SaaS platforms at scale.
  • Strong proficiency in Python, Go, Java, or C++ with microservices experience.
  • Deep Kubernetes/Docker experience in large multi-cluster environments.

Responsabilidades

  • Design and implement high-resilience software for AI-assisted observability and self-healing infra.
  • Build AI agents, MCP tooling, and deterministic evaluation pipelines.
  • Implement ingestion/correlation across logs, metrics, and traces to speed MTTD/MTTR.
  • Develop safe production automation with guardrails and HITL workflows.
  • Partner with teams to define SLIs/SLOs and lead PIRs.

Conocimientos

Python
Go
Java
C++
Kubernetes
Docker
SRE practices
OpenTelemetry
ML/AI tooling

Educación

Bachelor's degree in Computer Science or related field
Master's/PhD preferred for senior roles

Herramientas

Prometheus
Grafana
Kafka
Redis
Elasticsearch

Descripción del empleo

Meet the Team

The Collaboration Technology Group is redefining the future of teamwork, building services that connect people effortlessly across devices, locations and time zones.

Our team builds, runs and continuously improves the platform services behind Cisco’s collaboration products, operating at global scale across numerous datacentres. We’re a passionate, collaborative team focused on reliability, innovation and engineering excellence.


Your Impact

As a Software Engineer, you will design, build, and deploy the core capabilities of our next-generation AI-powered Production Intelligence platform. You will combine Site Reliability Engineering practices with modern agentic AI (reusable Skills, Model Context Protocol (MCP), and LLM tooling) to transform how engineering and leadership teams monitor, diagnose, and auto-remediate global SaaS infrastructure.


What You'll Do
  • Technical Design & Architecture: Design and implement high-resilience software systems for AI-assisted observability, automated incident response, and self-healing cloud infrastructure.
  • Agentic Workflows & Tooling: Design, build, and maintain production-grade AI agents, MCP tool integrations, and deterministic evaluation pipelines for automated operational decision support.
  • Telemetry & Insights: Implement ingestion and correlation pipelines across distributed logs, metrics, OpenTelemetry traces, change events, and runbooks to accelerate Mean Time to Detection (MTTD) and Resolution (MTTR).
  • Safe Production Automation: Develop proactive anomaly detection and Human-in-the-Loop (HITL) remediation workflows with rigorous safety, security, and quality guardrails.
  • Reliability & Scalability Engineering: Partner with application and infrastructure teams to define SLIs/SLOs, handle error budgets, and lead deep-dive post-incident reviews (PIRs).
  • Mentorship & Best Practices: Mentor mid-level and junior engineers, conduct thorough code reviews, establish engineering guidelines, and drive operational excellence across global development and operations teams.
  • Cross-Team Delivery: Manage priorities and deadlines, communicate progress clearly, and work across teams to turn production needs into reliable software and AI-assisted capabilities.

Minimum Qualifications
  • Bachelor’s degree + 8 years of related experience, Master’s + 6 years, or PhD + 3 years in Computer Science, Software Engineering, or a related technical field.
  • Experience as a Senior / Lead SRE or Software Engineer delivering distributed, high-availability SaaS platforms at scale.
  • Strong proficiency in Python, Go, Java, or C++ with experience designing microservices, APIs, and production automation.
  • Deep experience with Kubernetes, Docker, and container orchestration in large-scale multi-cluster environments.
  • Proven background in SRE practices: SLI/SLO design, observability platforms (metrics/logs/traces), incident management, and automated root cause analysis (RCA).

Preferred Qualifications
  • AI & Agentic Systems: Experience building LLM pipelines, AI Agents, Model Context Protocol (MCP) servers/clients, RAG architectures, and evaluation frameworks.
  • Observability & Telemetry: Experience with OpenTelemetry (OTel), Prometheus, Grafana, Splunk, ThousandEyes, or distributed tracing systems.
  • Cloud & Infrastructure: Expertise in public cloud providers (AWS, GCP, Azure), Terraform/IaC, and GitOps/CI/CD pipelines (Jenkins, GitHub Actions, Harness).
  • Safe Automation & Guardrails: Experience implementing responsible AI guardrails, deterministic fallback logic, and policy-driven remediation engines.
  • Data & Messaging: Experience with streaming and data platforms (Kafka, Redis, PostgreSQL, Elasticsearch/Vector DBs).
  • CollabHiring
Why Cisco?

At Cisco, we’re revolutionizing how data and infrastructure connect and protect organizations in the AI era – and beyond. We’ve been innovating fearlessly for 40 years to create solutions that power how humans and technology work together across the physical and digital worlds. These solutions provide customers with unparalleled security, visibility, and insights across the entire digital footprint.

Fueled by the depth and breadth of our technology, we experiment and create meaningful solutions. Add to that our worldwide network of doers and experts, and you’ll see that the opportunities to grow and build are limitless. We work as a team, collaborating with empathy to make really big things happen on a global scale. Because our solutions are everywhere, our impact is everywhere.

We are Cisco, and our power starts with you.

Consigue la evaluación confidencial y gratuita de tu currículum.

o arrastra y suelta tu archivo aquí

Similar jobs

Puestos de trabajo similares que vale la pena comparar

Software Engineering Tech Lead (SRE + AI)
Software Engineering Tech Lead (SRE + AI)

Cisco • Greater London

Presencial
GBP 120.000 - 180.000
Software Engineering Tech Lead (SRE + AI)
Software Engineering Tech Lead (SRE + AI)

Cisco Systems, Inc. • Greater London

Presencial
GBP 130.000 - 180.000
Senior Software Engineer - ThousandEyes
Senior Software Engineer - ThousandEyes

624 Cisco International Limited • Greater London

Presencial
GBP 60.000 - 110.000
Principal Software Engineer
Principal Software Engineer

Cisco Systems, Inc. • Greater London

Presencial
GBP 110.000 - 170.000
AI Solutions Engineer
AI Solutions Engineer

624 Cisco International Limited • Greater London

Presencial
GBP 90.000 - 130.000
Site Reliability Engineer, Infrastructure - ThousandEyes
Site Reliability Engineer, Infrastructure - ThousandEyes

Cisco • Greater London

Presencial
GBP 90.000 - 130.000
Customer Program Manager
Customer Program Manager

624 Cisco International Limited • Greater London

Presencial
GBP 90.000 - 140.000
Site Reliability Engineer
Site Reliability Engineer

Cisco Systems, Inc • Greater London

Híbrido
GBP 85.000 - 125.000
Principal Software Engineer
Principal Software Engineer

Cisco • Greater London

Presencial
GBP 90.000 - 130.000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Cisco • Greater London

Híbrido
GBP 90.000 - 150.000