Location: Makati City | Hybrid 2-3x RTO per week (M-F)
Employment type: Full-time | Day Shift
Role Overview
Architect the intelligent heartbeat of our enterprise ecosystem! We are looking for a Principal Site Reliability Engineer (SRE) to join our Monitoring and Observability Design Team. In this role, you will lead the architectural strategy for enterprise event management, telemetry, and automated operations. You will partner with cross-functional engineering leads to design resilient, proactive, and self-healing systems across a modern multi-cloud landscape.
Key Responsibilities
- Observability Architecture: Direct the design and rollout of scalable event management, telemetry pipelines, and automated incident response frameworks.
- Platform & Alert Integration: Lead the integration of modern observability streams into enterprise ITOM and AIOps platforms to drive rapid event correlation and remediation.
- Cross-Domain Collaboration: Partner with domain leads across cloud infrastructure, application performance monitoring (APM), and network design to build unified observability standards.
- Hybrid & Cloud Solutions: Architect end-to-end monitoring strategies for workloads across SaaS, IaaS, and PaaS environments.
- Innovation & Engineering: Drive automation strategies, software upgrade plans, and proof-of-concept initiatives to continuously advance enterprise AIOps capabilities.
Qualifications
- SRE & Architecture Expertise: Extensive background leading enterprise monitoring, observability, or IT architecture initiatives.
- Event Management & AIOps: Demonstrated mastery in alert integrations and AIOps platform implementation (specifically ServiceNow ITOM Event Management).
- Cloud & Telemetry: Hands-on experience with cloud telemetry (Azure Monitor, Log Analytics, KQL), synthetic transaction monitoring, and proactive performance management.
- Automation & Scripting: Strong coding skills in Python, PowerShell, JavaScript, C#, or SQL, along with API-driven automation and configuration tools (Ansible, DevOps workflows).
- Technical Fundamentals: Deep understanding of Application Performance Management (APM), network architecture (WAN/LAN, TCP/IP, PKI), and ITSM governance frameworks.
- Education: Bachelor’s degree in Computer Science, Information Technology, Engineering, or equivalent practical experience.
Technology Stack
- ITOM & AIOps: ServiceNow ITOM, Event Management, Automated Incident Correlation
- Observability & APM Tools: Azure Monitor, Grafana, Prometheus, AppDynamics, ThousandEyes, Riverbed, Nobl9
- Scripting & DevOps: Python, PowerShell, JavaScript, .NET/C#, REST APIs, Ansible
RH-TT