Turn this role into an interview — a resume and cover letter built around what this employer wants.
Ascendion in Metro Manila seeks a seasoned SRE / incident response professional to monitor system health across applications, infrastructure, APIs and services. You will respond to alerts in real time and expand observability through dashboards and runbooks.
You'll perform initial triage using logs and metrics, coordinate with on-call responders during high-priority events, and maintain clear shift handoff notes while supporting a 24/7 rotation, including nights and weekends.
Monitor system health across applications, infrastructure, APIs, and services using Splunk, Dynatrace, and Grafana
Review and respond to alerts, dashboards, and metrics in real time
Create, expand, and maintain dashboards and alerting to improve observability coverage
Perform initial triage using logs, traces, and metrics; identify symptoms and potential root causes
Execute runbooks/SOPs for common production issues (restarts, validation checks, health checks, etc.)
Create incidents, document findings, and escalat to L2/L3 engineering teams
Coordinate with on-call responders during high-priority events
Perform routine health checks and proactive monitoring tasks every shift
Provide clear communication and shift-to-shift handoff notes
Operate within a 24/7 shift rotation, including nights/weekends/holidays