Senior SRE - AI Resilience & Kubernetes Platform

Cisco

San Francisco (CA)

Hybrid

USD 168,000 - 245,000

Full time

25 hours ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Medical insurance
401(k) with company match
Paid parental leave
Disability coverage
Life insurance
Paid holidays and vacation

Job summary

Cisco is seeking a Senior Site Reliability Engineer to build and operate the Splunk Agent Observability deployment platform, ensuring reliability, scalability, and security across cloud and air-gapped environments. The role is hybrid with about two on-site days per week at Cisco offices in San Francisco, San Jose, or New York City.

You will own deployments, automate operations, and work with engineers to deliver secure, scalable architectures.

Qualifications

  • 7+ years’ experience in SRE, platform or cloud engineering.
  • 3+ years’ operating Kubernetes in production; experience with Helm.
  • Experience building and maintaining CI/CD platforms and deployment automation.
  • Experience working with AWS, GCP, or similar cloud platforms.

Responsibilities

  • Operate and improve Kubernetes-based production infrastructure and deployment systems.
  • Own customer deployments across cloud and air-gapped environments, including installation, upgrades, troubleshooting, and lifecycle management.
  • Build and improve deployment observability, monitoring, logging, and alerting.
  • Improve reliability, scalability, and operational efficiency through automation and performance optimization.
  • Participate in production incident response, root cause analysis, and reliability improvements.
  • Tune infrastructure components—including databases and services—to improve performance and resilience.
  • Design and develop internal tooling using Python and/or Go.
  • Debug complex production issues spanning Kubernetes, networking, storage, and application layers.
  • Manage infrastructure using Terraform or similar Infrastructure as Code tools.
  • Collaborate with software engineers and customers to design secure, scalable, and reliable deployment architectures.

Skills

Kubernetes
Helm
CI/CD
Cloud platforms

Education

Bachelor's degree

Tools

Terraform
Python
Go

Job description

Cisco is seeking a Senior Site Reliability Engineer to build and operate the Splunk Agent Observability deployment platform, ensuring reliability, scalability, and security across cloud and air-gapped environments. The role is hybrid with about two on-site days per week at Cisco offices in San Francisco, San Jose, or New York City.

You will own deployments, automate operations, and work with engineers to deliver secure, scalable architectures.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior SRE - AI Resilience & Cloud Platform
Senior SRE - AI Resilience & Cloud Platform

Cisco • San Francisco (CA)

Hybrid
USD 187,000 - 268,000
Health insurance
401(k) with Cisco match
Paid parental leave
+2
Senior SRE (Hybrid) - AI Resilience Platform
Senior SRE (Hybrid) - AI Resilience Platform

020 Cisco Systems, Inc. • San Jose (CA)

Hybrid
USD 168,000 - 245,000
Senior SRE: AI Resilience & Platform Reliability Lead
Senior SRE: AI Resilience & Platform Reliability Lead

Cisco • San Jose (CA)

Hybrid
USD 170,000 - 308,000
Senior SRE - AI Platform Reliability (Hybrid)
Senior SRE - AI Platform Reliability (Hybrid)

Cisco • New York (NY)

Hybrid
USD 187,000 - 308,000
Medical insurance
Dental insurance
Vision insurance
+5
Senior AI Reliability Engineer
Senior AI Reliability Engineer

Cisco • New York (NY)

Hybrid
USD 149,000 - 282,000
Medical, dental, vision insurance
401(k) with Cisco matching
Paid parental leave
+1
Senior SRE — AI Resilience & Cloud Platforms
Senior SRE — AI Resilience & Cloud Platforms

Cisco • San Jose (CA)

Hybrid
USD 168,000 - 245,000
Senior SRE: Cloud-Native Infra & Scale (Hybrid)
Senior SRE: Cloud-Native Infra & Scale (Hybrid)

Cisco Systems, Inc. • San Francisco (CA)

Hybrid
USD 168,000 - 245,000
Senior SRE: AI-Driven Kubernetes Reliability at Scale
Senior SRE: AI-Driven Kubernetes Reliability at Scale

fal - Features & Labels • San Francisco (CA)

On-site
USD 180,000 - 240,000
Health insurance
Dental insurance
Vision insurance
+1
Senior SRE — Flexible, AI-Driven Reliability
Senior SRE — Flexible, AI-Driven Reliability

Salesforce, Inc. • San Francisco (CA)

Hybrid
USD 148,000 - 224,000
Senior SRE: Observability, Splunk & Automation (Hybrid)
Senior SRE: Observability, Splunk & Automation (Hybrid)

ISO New England Inc. • Holyoke (MA)

Hybrid
USD 134,000 - 170,000
Hybrid work environment (3 days/week)