Senior SRE - Observability & Cloud Platform Lead

Cisco

Nashville (TN)

On-site

USD 168,000 - 245,000

Full time

6 hours ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Cisco seeks an experienced Site Reliability Engineer to design, deploy, and operate enterprise observability platforms across logs, metrics, and traces. You will build Splunk infrastructure, ELK-based tooling, and end-to-end telemetry pipelines using Grafana Tempo, OpenTelemetry, and Kafka.

The role requires on-call participation, collaboration with cross-functional teams, and driving reliability improvements.

Qualifications

  • 7+ years in SRE/Platform or DevOps or related field.
  • Experience administering Splunk Enterprise or Splunk Cloud in production.
  • Strong SPL, dashboards, alerts, analytics experience.
  • Experience designing and operating Elasticsearch/ELK platforms.
  • Familiar with Prometheus, Grafana, Grafana Tempo, OpenTelemetry, Kafka.
  • Terraform and IaC familiarity; Linux and networking knowledge.
  • Programming or scripting in Python, Go, Ruby, Bash.
  • Experience in on-call and incident response; strong communication.

Responsibilities

  • Design, deploy, operate, and continuously improve enterprise observability platforms.
  • Build and maintain Splunk infrastructure, indexers, clusters, deployments.
  • Operate Elasticsearch clusters for logs, search, and troubleshooting.
  • Design, deploy telemetry via Grafana Tempo and OpenTelemetry.
  • Develop dashboards, alerts, analytics, trace visualizations.
  • Define data quality, retention, and observability patterns.
  • Scale Prometheus, Grafana, Kafka, Tempo, OpenTelemetry; monitor systems.
  • Automate provisioning and config with Terraform and tooling.
  • Participate in capacity planning, DR, and production readiness reviews.
  • Troubleshoot distributed systems and lead incident root-cause analyses.
  • Collaborate with platform, app, security, DB, and network teams.

Skills

SRE
Splunk
Splunk SPL
Elasticsearch
Prometheus
Grafana
OpenTelemetry
Kafka
Kubernetes
Terraform
Linux
Python
Go
Bash
On-call
Communication

Tools

Splunk Enterprise
Splunk Cloud
Elasticsearch
Kibana
Prometheus
Grafana
Grafana Tempo
OpenTelemetry
Kafka
Kubernetes
Terraform
Ansible
AWS
Azure
GCP
Python
Go
Bash

Job description

Cisco seeks an experienced Site Reliability Engineer to design, deploy, and operate enterprise observability platforms across logs, metrics, and traces. You will build Splunk infrastructure, ELK-based tooling, and end-to-end telemetry pipelines using Grafana Tempo, OpenTelemetry, and Kafka.

The role requires on-call participation, collaboration with cross-functional teams, and driving reliability improvements.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior SRE: Observability & Cloud Platform Engineer
Senior SRE: Observability & Cloud Platform Engineer

Cisco • San Francisco (CA)

On-site
USD 168,000 - 245,000
Medical, dental, and vision insurance
401(k) with Cisco matching
Paid parental leave
+3
Senior SRE - Observability & Cloud Reliability
Senior SRE - Observability & Cloud Reliability

Cisco Systems, Inc. • San Francisco (CA)

On-site
USD 168,000 - 245,000
Medical benefits
401(k) matching
Parental leave
+1
Senior Observability & Reliability Engineer
Senior Observability & Reliability Engineer

Cisco • Boise (ID)

On-site
USD 120,000 - 170,000
Senior Observability SRE – Kubernetes & Cloud
Senior Observability SRE – Kubernetes & Cloud

Cisco • San Francisco (CA)

On-site
USD 168,000 - 245,000
Senior Site Reliability Engineer – Observability
Senior Site Reliability Engineer – Observability

Cisco • Boise (ID)

On-site
USD 120,000 - 170,000
Senior Site Reliability Engineer – Observability & Cloud
Senior Site Reliability Engineer – Observability & Cloud

Cosm Inc. • El Segundo (CA), Northern (KY)

Hybrid
USD 110,000 - 145,000
Senior Observability Platform Engineer - Splunk & ELK
Senior Observability Platform Engineer - Splunk & ELK

Tata Consultancy Services • San Jose (CA)

On-site
USD 94,000 - 130,000
SRE Lead: Drive Resilience, Scale, and Observability
SRE Lead: Drive Resilience, Scale, and Observability

Iscale Solutions • United States

Remote
USD 120,000 - 160,000
Health coverage
Vacation and sick leave credits
Professional development
+5
Senior SRE: Observability, Splunk & Automation (Hybrid)
Senior SRE: Observability, Splunk & Automation (Hybrid)

ISO New England Inc. • Holyoke (MA)

Hybrid
USD 134,000 - 170,000
Hybrid work environment (3 days/week)
Senior SRE: Observability & Automation Leader (Hybrid)
Senior SRE: Observability & Automation Leader (Hybrid)

ISO New England • Holyoke (MA)

Hybrid
USD 134,000 - 170,000
Hybrid work environment
Relocation assistance
Tuition reimbursement
+6