Turn this role into an interview — a resume and cover letter built around what this employer wants.
Cisco Systems is seeking an experienced Site Reliability Engineer to join our Collaboration Business Unit in Bengaluru. You will own deployment and operation across cloud and hybrid environments, build robust CI/CD pipelines, and lead incident responses to ensure service reliability and performance at scale.
A strong automation mindset and experience with AI-assisted tooling are valued. Ideal candidates bring 7–13 years in SRE or Cloud Operations, expertise with Docker/Kubernetes, and
Meet the Team Cisco's Collaboration Business Unit empowers people and organizations worldwide to connect, communicate, and innovate seamlessly. You will collaborate with a global team of software engineers and SREs responsible for delivering extraordinary collaboration experiences at scale. Our team supports backend services deployed worldwide and works closely with development, product, and operations partners to ensure reliability and performance. Webex is powering the shift to the hybrid workforce, helping people stay connected in a rapidly evolving digital world. We cultivate a startup-like culture that values innovation, ownership, and collaboration, while offering the scale and impact of a global technology leader.
From a reliability standpoint, this role involves evaluating the scalability, resiliency, performance, and security properties and techniques used in production environments. It supports the uptime of production services through an On-Call rotation, which includes monitoring and alerting to meet internal Service Level Objectives (SLOs) and customer-facing Service Level Agreements (SLAs). Ensuring reliable incident processes is achieved by conducting Disaster Recovery drills. The role also focuses on improving reliability through incident management by investigating incidents, implementing remediation strategies, and learning from past incidents to make improvements. It involves determining the reliability and security requirements of components and systems to meet the reliability objectives of the company, customers, and any relevant governmental agencies. Additionally, the role aims to reduce operational expenses through automation, by identifying and mitigating failure points, and automating repetitive and resource-intensive tasks. It also involves developing new acceleration techniques and analytical tools to ensure the early identification of potential issues with new products, packaging, processes, and overall product reliability.
At Cisco, we're revolutionizing how data and infrastructure connect and protect organizations in the AI era – and beyond. We've been innovating fearlessly for 40 years to create solutions that power how humans and technology work together across the physical and digital worlds. These solutions provide customers with unparalleled security, visibility, and insights across the entire digital footprint. Fueled by the depth and breadth of our technology, we experiment and create meaningful solutions. Add to that our worldwide network of doers and experts, and you'll see that the opportunities to grow and build are limitless. We work as a team, collaborating with empathy to make really big things happen on a global scale. Because our solutions are everywhere, our impact is everywhere. We are Cisco, and our power starts with you.
Cisconians power the future. We make impact as a team, innovating fast and fearlessly to create meaningful solutions on a large scale. The depth and breadth of our technology doesn't just benefit our customers – it also means limitless opportunities for us to experiment and learn. We understand the power each of our unique backgrounds bring when we work together. Because of that, we have a global network of thinkers, doers, experts, and curious creators who help one another do their life's best work.