An application made for this job — a tailored resume and cover letter that speak straight to the posting.
Cisco is seeking an AI Operations Engineering Technical Leader to drive the operationalization of complex autonomous agent architectures into secure, scalable production environments. You will define robust MLOps practices, pipelines, and model serving architectures to bring Agentic AI concepts into customer-facing products.
Design telemetry to monitor token usage, control compute costs, and ensure secure model performance.
The team operates at the cutting edge of AI, functioning within the broader Customer Experience Engineering organization as a specialized Customer Reliability Engineering (XRE) group to support Cisco’s customer experience (CX) platform. Our primary mission is to ensure seamless, highly reliable production systems by fiercely resolving complex customer reliability issues, managing critical escalations, and ensuring AI-driven solutions are secure and performant. We are a collaborative, agile group of MLOps experts, reliability engineers, and software developers who thrive on troubleshooting and stabilizing complex technical ecosystems. Working closely with design, data science, and product management, our efforts directly protect Cisco's CX product roadmap and ensure long-term customer success. What’s most exciting is the opportunity to shape the reliability of a rapidly evolving AI landscape, directly impacting the customer experience by tackling complex challenges in hybrid cloud environments, LLM orchestration, and inference optimization every day.
As an AI Operations Engineering Technical Leader focusing on Agentic AI and MLOps, you will drive the operationalization of complex autonomous agent architectures into secure, scalable, and high-performing production environments. Define and establish robust MLOps practices, foundational ML pipelines, and model serving architectures to bring innovative Agentic AI concepts out of the lab and into customer-facing products. Design and implement robust telemetry monitoring tools to track token usage, control compute costs, monitor agent reasoning paths, and ensure optimal model performance and security. Proactively resolve complex infrastructure challenges by managing containerized models on GPU clusters and addressing model inference issues. Guide architectural choices based on deep market knowledge and mentor peers, ensuring our Agentic AI platforms remain reliable, scalable, and ahead of the curve.