Senior SRE (Hybrid) - AI Resilience Platform

020 Cisco Systems, Inc.

San Jose (CA)

Hybrid

USD 168,000 - 245,000

Full time

2 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Cisco Systems, Inc. is seeking a Senior Site Reliability Engineer to build, operate, and continuously improve the reliability and scalability of Splunk Agent Observability's deployment platform and production infrastructure.

You will own the operational backbone supporting cloud and air-gapped deployments, automate operational tasks, and partner with engineering teams to deliver secure and scalable systems.

Qualifications

  • 7+ years’ experience in Site Reliability Engineering, Platform/Cloud/Infrastructure Engineering, or related fields.
  • 3+ years’ operating Kubernetes in production; experience with Helm.

Responsibilities

  • Operate and improve Kubernetes-based production infrastructure and deployment systems.
  • Own customer deployments across cloud and air-gapped environments, including installation, upgrades, troubleshooting, and lifecycle management.
  • Build and improve deployment observability, monitoring, logging, and alerting.
  • Improve reliability, scalability, and operational efficiency through automation and performance optimization.
  • Participate in production incident response, root cause analysis, and reliability improvements.
  • Tune infrastructure components—including databases and services—to improve performance and resilience.
  • Design and develop internal tooling using Python and/or Go.
  • Debug complex production issues spanning Kubernetes, networking, storage, and application layers.
  • Manage infrastructure using Terraform or similar Infrastructure as Code tools.
  • Collaborate with software engineers and customers to design secure, scalable, and reliable deployment architectures.

Skills

Kubernetes operations
CI/CD
Automation
Observability
Troubleshooting
Python
Go
Cloud fundamentals

Education

Bachelor's degree
Master's degree

Tools

Kubernetes
Helm
Terraform
Python
Go
AWS
GCP

Job description

Cisco Systems, Inc. is seeking a Senior Site Reliability Engineer to build, operate, and continuously improve the reliability and scalability of Splunk Agent Observability's deployment platform and production infrastructure.

You will own the operational backbone supporting cloud and air-gapped deployments, automate operational tasks, and partner with engineering teams to deliver secure and scalable systems.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior SRE: Observability, Splunk & Automation (Hybrid)
Senior SRE: Observability, Splunk & Automation (Hybrid)

ISO New England Inc. • Holyoke (MA)

Hybrid
USD 134,000 - 170,000
Hybrid work environment (3 days/week)
Senior SRE - Hybrid, Observability & Reliability
Senior SRE - Hybrid, Observability & Reliability

Early Warning Services LLC • Chicago (IL)

Hybrid
USD 106,000 - 130,000
Healthcare Coverage
401(k) Plan with match
PTO and Holidays
+1
Senior SRE - Hybrid, AWS & Observability
Senior SRE - Hybrid, AWS & Observability

Early Warning • San Francisco (CA)

Hybrid
USD 128,000 - 156,000
Healthcare coverage
401(k) plan
Paid time off
+2
Senior SRE: Observability, Automation & Scalable Systems
Senior SRE: Observability, Automation & Scalable Systems

Replit • Northern (KY)

Hybrid
USD 140,000 - 190,000
Competitive Salary & Equity
401(k) 4% match (US)
Health, Dental, Vision & Life
+7
Senior SRE: AI-Driven, Cloud-Native Reliability (Hybrid)
Senior SRE: AI-Driven, Cloud-Native Reliability (Hybrid)

OutSystems • San Francisco (CA)

Hybrid
USD 140,000 - 210,000
Hybrid work model
Senior SRE: Scale Reliability, Observability & Resilience
Senior SRE: Scale Reliability, Observability & Resilience

Early Warning Services LLC • Scottsdale (AZ)

Hybrid
USD 106,000 - 130,000
Healthcare Coverage
401(k) Retirement Plan
Paid Time Off
+2
Senior SRE: Automate Reliability & Observability
Senior SRE: Automate Reliability & Observability

United States Digital Space LLC • Charlotte (TX)

On-site
USD 153,000 - 192,000
Discretionary incentive eligible
Benefits package
Senior SRE - Multi-Cloud Reliability & AI-Driven Ops
Senior SRE - Multi-Cloud Reliability & AI-Driven Ops

Satsuma AI, Inc. • Austin (TX), Northern (KY)

Hybrid
USD 140,000 - 210,000
Unlimited PTO
401(K)
Healthcare Stipend
+1
Senior SRE: Scalable Infra, Observability & Automation
Senior SRE: Scalable Infra, Observability & Automation

Early Warning • Scottsdale (AZ)

Hybrid
USD 106,000 - 156,000
Healthcare Coverage
401(k) Company Match
Paid Time Off
+2
Senior SRE - AI-Powered Cloud Reliability (Remote)
Senior SRE - AI-Powered Cloud Reliability (Remote)

ServiceTitan, Inc. • Northern (KY)

Hybrid
USD 148,000 - 221,000