Stand out for this role — generate a tailored resume and cover letter in about a minute.
Cisco is hiring a Senior Site Reliability Engineer to help scale and secure collaboration services across cloud and hybrid environments. You will work with global software engineers and SREs to ensure reliability, performance, and cost efficiency for Webex services.
The role emphasizes on-call reliability, incident management, and continuous improvement of deployment pipelines, with a strong focus on observability and automation.
Cisco s Collaboration Business Unit empowers people and organizations worldwide to connect, communicate, and innovate seamlessly.
You will collaborate with a global team of software engineers and SREs responsible for delivering extraordinary collaboration experiences at scale. Our team supports backend services deployed worldwide and works closely with development, product, and operations partners to ensure reliability and performance.
Webex is powering the shift to the hybrid workforce, helping people stay connected in a rapidly evolving digital world. We cultivate a startup-like culture that values innovation, ownership, and collaboration, while offering the scale and impact of a global technology leader.
From a reliability standpoint, this role involves evaluating the scalability, resiliency, performance, and security properties and techniques used in production environments. It supports the uptime of production services through an On-Call rotation, which includes monitoring and alerting to meet internal Service Level Objectives (SLOs) and customer-facing Service Level Agreements (SLAs). Ensuring reliable incident processes is achieved by conducting Disaster Recovery drills.
The role also focuses on improving reliability through incident management by investigating incidents, implementing remediation strategies, and learning from past incidents to make improvements. It involves determining the reliability and security requirements of components and systems to meet the reliability objectives of the company, customers, and any relevant governmental agencies. Additionally, the role aims to reduce operational expenses through automation, by identifying and mitigating failure points, and automating repetitive and resource-intensive tasks. It also involves developing new acceleration techniques and analytical tools to ensure the early identification of potential issues with new products, packaging, processes, and overall product reliability.