Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.
Cisco’s Collaboration Business Unit seeks a seasoned Site Reliability Engineer to own deployment and operation of critical collaboration services across cloud and hybrid environments. You will drive reliability, design scalable pipelines, and lead incident responses with a focus on performance and security.
You will use observability data to guide capacity planning and resource optimization, while promoting best practices and automation.
Cisco’s Collaboration Business Unit empowers people and organizations worldwide to connect, communicate, and innovate seamlessly. You will collaborate with a global team of software engineers and SREs responsible for delivering extraordinary collaboration experiences at scale. Our team supports backend services deployed worldwide and works closely with development, product, and operations partners to ensure reliability and performance. Webex is powering the shift to the hybrid workforce, helping people stay connected in a rapidly evolving digital world. We cultivate a startup‑like culture that values innovation, ownership, and collaboration, while offering the scale and impact of a global technology leader. Your impact From a reliability standpoint, this role involves evaluating the scalability, resiliency, performance, and security properties and techniques used in production environments. It supports the uptime of production services through an On-Call rotation, which includes monitoring and alerting to meet internal Service Level Objectives (SLOs) and customer‑facing Service Level Agreements (SLAs). Ensuring reliable incident processes is achieved by conducting Disaster Recovery drills. The role also focuses on improving reliability through incident management by investigating incidents, implementing remediation strategies, and learning from past incidents to make improvements. It involves determining the reliability and security requirements of components and systems to meet the reliability objectives of the company, customers, and any relevant governmental agencies. Additionally, the role aims to reduce operational expenses through automation, by identifying and mitigating failure points, and automating repetitive and resource‑intensive tasks. It also involves developing new acceleration techniques and analytical tools to ensure the early identification of potential issues with new products, packaging, processes, and overall product reliability.
Bachelor’s degree in Computer Science, Engineering, or related field (or equivalent experience) with 7 –13 years in Site Reliability Engineering, Cloud Operations, or Systems Engineering. Strong hands‑on experience operating production services using Docker and Kubernetes in cloud or hybrid environments. Proficiency in one or more programming or scripting languages (e.g., Python, Go, Bash) to build automation and operational tooling. Experience with monitoring, observability, and incident response in production environments, including on‑call participation and post‑incident reviews. Working knowledge of Linux systems, networking, distributed systems, CI/CD pipelines, infrastructure‑as‑code, and Git‑based workflows.
Experience operating large‑scale, globally distributed SaaS platforms. Familiarity with hybrid cloud environments and multi‑region deployments. Experience applying AI‑assistant first, ...??? strange mapping randomness...
At Cisco, we’re revolutionizing how data and infrastructure connect and protect organizations in the AI era – and beyond. We’ve been innovating fearlessly for 40 years to create solutions that power how humans and technology work together across the physical and digital worlds. These solutions provide customers with unparalleled security, visibility, and insights across the entire digital footprint. Fueled by the depth and breadth of our technology, we experiment and create meaningful solutions. Add to that our worldwide network of doers and experts, and you’ll see that the opportunities to grow and build are limitless. We work as a team, collaborating with empathy to make really big things happen on a global scale. Because our solutions are everywhere, our impact is everywhere. We are Cisco, and our power starts with you.