Responsible for the design, implementation, administration, integration, and automation of enterprise Monitoring, Observability, ITSM, Event Management, and Ticketing platforms supporting Cloud, Data Center, Network, Security, Managed Services, and AI/GPU infrastructure environments. The role will drive end-to-end service visibility, operational efficiency, event intelligence, automated ticket lifecycle management, auto-remediation, and AIOps capabilities across enterprise and cloud services.
- Design, deploy, configure, and maintain enterprise monitoring platforms.
- Create monitoring standards, templates, and governance frameworks.
- Implement service health dashboards and executive operational reporting.
- Define monitoring thresholds, alerting policies, and escalation frameworks.
- Configure event correlation, suppression, enrichment, and intelligent alerting workflows.
- Define ticket lifecycle workflows and SLA management mechanisms.
- Automate incident creation, assignment, escalation, and closure processes.
- Implement workflow orchestration across multiple operational domains.
- Design self-service automation and service request fulfillment processes.
- Automate approval workflows and change management processes.
- Develop chatbot and virtual agent integrations.
- Integrate monitoring platforms with ITSM ticketing systems.
- Enable bi-directional synchronization between monitoring and ticketing platforms.
- Develop and maintain automation scripts and workflows.
- Build self-healing and auto-remediation capabilities.
- Automate operational runbooks and standard procedures.
- Develop reusable automation modules and APIs.
- Support AIOps initiatives across cloud and managed services environments.
- Support machine learning-based event correlation.
- Translate business and operational requirements into tooling solutions.
- Conduct architecture reviews and operational readiness assessments.
Experience & Educational Requirement
BE/B-Tech or equivalent with Computer Science or Electronics & Communication
RELEVANT EXPERIENCE
- Minimum 8-10 years of experience in Monitoring, ITSM, Service Assurance, Tool Administration, Platform Operations, Automation Engineering, or Service Operations.
- Experience supporting large-scale enterprise, cloud, telecom, managed service, or data center environments.
- Hands‑on experience with Grafana, Prometheus, Zabbix, Nagios, SolarWinds, ManageEngine, ServiceNow etc.
- Excellent troubleshooting and analytical skills.
- Strong stakeholder management capabilities.
- Effective communication and presentation skills.
- Strong documentation and process management capabilities.
- Preferred Certifications ITIL v4 Foundation / Managing Professional /ServiceNow System Administrator/ ServiceNow Implementation Specialist