Software Engineering Technical Leader - SRE + Kubernetes

Cisco

Bengaluru

On-site

INR 300,000 - 550,000

Full time

12 days ago
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Cisco is hiring a Senior Site Reliability Engineer to help scale and secure collaboration services across cloud and hybrid environments. You will work with global software engineers and SREs to ensure reliability, performance, and cost efficiency for Webex services.

The role emphasizes on-call reliability, incident management, and continuous improvement of deployment pipelines, with a strong focus on observability and automation.

Qualifications

  • Bachelor's degree in Computer Science, Engineering, or related field (or equivalent experience).
  • Proficiency in Python, Go, or Bash for automation.
  • Experience operating production services with Docker and Kubernetes in cloud or hybrid environments.
  • Experience with monitoring, observability and on-call incident reviews.

Responsibilities

  • Own deployment and operation of critical collaboration services across cloud and hybrid environments.
  • Design, evolve, and optimize CI/CD pipelines and automation.
  • Lead incident response for complex production issues and perform root cause analysis.
  • Use observability data to guide capacity planning and resource optimization.
  • Define and champion operational best practices and runbooks.

Skills

Python
Go
Bash
On-call

Education

Bachelor's degree in CS/Engineering or related field

Tools

Docker
Kubernetes
Git

Job description

Job Summary

Cisco s Collaboration Business Unit empowers people and organizations worldwide to connect, communicate, and innovate seamlessly.

You will collaborate with a global team of software engineers and SREs responsible for delivering extraordinary collaboration experiences at scale. Our team supports backend services deployed worldwide and works closely with development, product, and operations partners to ensure reliability and performance.

Webex is powering the shift to the hybrid workforce, helping people stay connected in a rapidly evolving digital world. We cultivate a startup-like culture that values innovation, ownership, and collaboration, while offering the scale and impact of a global technology leader.

Responsibilities

From a reliability standpoint, this role involves evaluating the scalability, resiliency, performance, and security properties and techniques used in production environments. It supports the uptime of production services through an On-Call rotation, which includes monitoring and alerting to meet internal Service Level Objectives (SLOs) and customer-facing Service Level Agreements (SLAs). Ensuring reliable incident processes is achieved by conducting Disaster Recovery drills.

The role also focuses on improving reliability through incident management by investigating incidents, implementing remediation strategies, and learning from past incidents to make improvements. It involves determining the reliability and security requirements of components and systems to meet the reliability objectives of the company, customers, and any relevant governmental agencies. Additionally, the role aims to reduce operational expenses through automation, by identifying and mitigating failure points, and automating repetitive and resource-intensive tasks. It also involves developing new acceleration techniques and analytical tools to ensure the early identification of potential issues with new products, packaging, processes, and overall product reliability.

  • Own the deployment and operation of critical collaboration services across cloud and hybrid environments, driving reliability and scalability.
  • Design, evolve, and optimize CI/CD pipelines and automation, including AI-first tooling for deployment, monitoring, and incident response.
  • Lead incident response for complex production issues, perform root cause analysis, and drive systemic reliability and performance improvements.
  • Use observability data to guide capacity planning, scaling strategies, and resource optimization across services.
  • Define and champion operational best practices, documentation standards, and a culture of reliability and operational excellence.
Minimum Qualifications
  • Bachelor s degree in Computer Science, Engineering, or related field (or equivalent experience) with 7-13 years in Site Reliability Engineering, Cloud Operations, or Systems Engineering.
  • Strong hands-on experience operating production services using Docker and Kubernetes in cloud or hybrid environments.
  • Proficiency in one or more programming or scripting languages (e.g., Python, Go, Bash) to build automation and operational tooling.
  • Experience with monitoring, observability, and incident response in production environments, including on-call participation and post-incident reviews.
  • Working knowledge of Linux systems, networking, distributed systems, CI/CD pipelines, infrastructure-as-code, and Git-based workflows.
Preferred Qualifications
  • Experience operating large-scale, globally distributed SaaS platforms.
  • Familiarity with hybrid cloud environments and multi-region deployments.
  • Experience applying AI-assisted or automation-first approaches to SRE tooling and workflows.
  • Strong written communication skills for creating clear operational documentation and runbooks.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Engineering Technical Leader - SRE + Kubernetes + Cloud + Automation + Incident Management + AI-first (10-14 Years)
Software Engineering Technical Leader - SRE + Kubernetes + Cloud + Automation + Incident Management + AI-first (10-14 Years)

123 Cisco Systems (India) Private Limited • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Software Engineering Technical Leader - SRE + Kubernetes + Cloud + Automation + Incident Management + AI-first (10-14 Years)
Software Engineering Technical Leader - SRE + Kubernetes + Cloud + Automation + Incident Management + AI-first (10-14 Years)

Engg • Bengaluru

On-site
INR 4,200,000 - 6,000,000
Software Engineering Technical Leader - SRE + Kubernetes + Cloud + Automation + Incident Management + AI-first (10-14 Years)
Software Engineering Technical Leader - SRE + Kubernetes + Cloud + Automation + Incident Management + AI-first (10-14 Years)

Cisco • Bengaluru

On-site
INR 1,800,000 - 3,200,000
Site Reliability Engineer
Site Reliability Engineer

Cisco Systems, Inc. • Bengaluru

On-site
INR 1,200,000 - 1,800,000
Forward Deployment Engineer (SRE)
Forward Deployment Engineer (SRE)

PwC Acceleration Centers • Hyderabad

On-site
INR 2,500,000 - 4,000,000
Software Engineering Technical Leader - SRE | Security Architect, Kubernetes,AWS,Terraform | 13+ years | Bangalore
Software Engineering Technical Leader - SRE | Security Architect, Kubernetes,AWS,Terraform | 13+ years | Bangalore

Engg • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Lead SRE
Lead SRE

Cvent, Inc. • Gurugram District

On-site
INR 4,000,000 - 8,000,000
Site Reliability Engineer-Vice President
Site Reliability Engineer-Vice President

Citi Bank • Pune District

On-site
INR 1,800,000 - 3,200,000
Lead Site Reliability Engineer (SRE)
Lead Site Reliability Engineer (SRE)

Skillventory • Kamrup Metropolitan

On-site
INR 3,500,000 - 7,000,000
Lead SRE
Lead SRE

Cvent, Inc. • India

On-site
INR 2,500,000 - 4,500,000