Senior AI Reliability Engineer

Cisco

New York (NY)

Hybrid

USD 149,000 - 282,000

Full time

21 hours ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Medical, dental, vision insurance
401(k) with Cisco matching
Paid parental leave
Paid holidays and vacation

Job summary

Cisco is seeking a Senior Site Reliability Engineer to architect, operate, and evolve the Splunk Agent Observability deployment platform. You will manage cloud and air-gapped environments, automate operations, and partner with engineering to deliver reliable, secure systems.

The role emphasizes Kubernetes, CI/CD, Terraform, and programming with Python/Go across hybrid on-site/off-site environments. It requires extensive SRE experience and cloud proficiency.

Qualifications

  • 7+ years of experience with a Bachelor's degree or 4+ years with a Masters or 1 year with a PhD, or equivalent; at least 4 years in SRE/Platform/Cloud/Infra Engineering
  • 3+ years operating Kubernetes in production; Helm experience
  • Experience building and maintaining CI/CD platforms and deployment automation
  • Experience with AWS, GCP, or similar cloud platforms

Responsibilities

  • Operate and improve Kubernetes-based production infrastructure and deployment systems
  • Own deployments across cloud and air-gapped environments including installation, upgrades, troubleshooting, and lifecycle management
  • Build and improve deployment observability, monitoring, logging, and alerting
  • Improve reliability, scalability, and efficiency through automation and performance optimization
  • Participate in production incident response, RCA, and reliability improvements
  • Tune databases and services to improve performance and resilience
  • Design and develop internal tooling using Python and/or Go
  • Debug production issues across Kubernetes, networking, storage, and application layers
  • Manage infrastructure using Terraform or similar IaC tools
  • Collaborate with engineers and customers to design secure, scalable deployment architectures

Skills

Kubernetes
Helm
CI/CD
Python
Go
Terraform
AWS/GCP
Troubleshooting

Education

Bachelor's degree OR 4+ yrs with a Masters OR 1 year with a PhD

Tools

Kubernetes
Helm
Terraform
AWS
GCP

Job description

Cisco is seeking a Senior Site Reliability Engineer to architect, operate, and evolve the Splunk Agent Observability deployment platform. You will manage cloud and air-gapped environments, automate operations, and partner with engineering to deliver reliable, secure systems.

The role emphasizes Kubernetes, CI/CD, Terraform, and programming with Python/Go across hybrid on-site/off-site environments. It requires extensive SRE experience and cloud proficiency.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior SRE: AI Resilience & Platform Reliability Lead
Senior SRE: AI Resilience & Platform Reliability Lead

Cisco • San Jose (CA)

Hybrid
USD 170,000 - 308,000
Senior SRE (Hybrid) - AI Resilience Platform
Senior SRE (Hybrid) - AI Resilience Platform

020 Cisco Systems, Inc. • San Jose (CA)

Hybrid
USD 168,000 - 245,000
Senior SRE - AI Platform Reliability (Hybrid)
Senior SRE - AI Platform Reliability (Hybrid)

Cisco • New York (NY)

Hybrid
USD 187,000 - 308,000
Medical insurance
Dental insurance
Vision insurance
+5
Senior SRE - AI Resilience & Kubernetes Platform
Senior SRE - AI Resilience & Kubernetes Platform

Cisco • San Francisco (CA)

Hybrid
USD 168,000 - 245,000
Medical insurance
401(k) with company match
Paid parental leave
+3
Senior SRE - AI Resilience & Cloud Platform
Senior SRE - AI Resilience & Cloud Platform

Cisco • San Francisco (CA)

Hybrid
USD 187,000 - 268,000
Health insurance
401(k) with Cisco match
Paid parental leave
+2
Senior SRE — AI Resilience & Cloud Platforms
Senior SRE — AI Resilience & Cloud Platforms

Cisco • San Jose (CA)

Hybrid
USD 168,000 - 245,000
Senior SRE: AI Cloud Reliability & Observability (Remote)
Senior SRE: AI Cloud Reliability & Observability (Remote)

SDI International • Chicago (IL)

Hybrid
USD 130,000 - 180,000
Senior Site Reliability Engineer: Observability & Cloud
Senior Site Reliability Engineer: Observability & Cloud

VBeyond Corporation • Jersey City (NJ)

On-site
USD 100,000 - 260,000
AI Platform SRE: Reliability, Observability & Scale
AI Platform SRE: Reliability, Observability & Scale

Schonfeld • New York (NY)

On-site
USD 175,000 - 225,000
Senior SRE - AI-Powered Cloud Reliability (Remote)
Senior SRE - AI-Powered Cloud Reliability (Remote)

ServiceTitan, Inc. • Northern (KY)

Hybrid
USD 148,000 - 221,000