Site Realibility Engineer

Gratitude Philippines

Manila

On-site

PHP 600,000 - 1,200,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Gratitude Philippines is seeking an experienced SRE/Monitoring professional to raise reliability across platforms with emphasis on Azure Monitor and ITOM Event Management. You’ll collaborate with IT Operations, platform, cyber, and product teams to implement SLO/SLI-driven operations and reduce alert noise.

The role requires hands-on expertise in Azure Monitor, KQL, ServiceNow Event Management, and observability tools, plus strong written and verbal communication to enable enterprise standards.

Qualifications

  • Bachelor’s degree in IT /Computer science/Engineering, or related field.
  • 3+ years in monitoring/observability/SRE roles with hands‑on experience in Azure Monitor/App Insights (KQL) and ServiceNow Event Management.
  • Strong knowledge in Azure Log Analytics, KQL, Telemetry, APM implementations.
  • Demonstrated ability to collaborate across IT Operations team, platform, cyber, network, and product teams, strong written verbal communication for standards and enablement.
  • 5+ years of experience with SRE role and deep understanding of monitoring and application performance management.
  • Knowledge of SLO platforms (e.g., Nobl9) and experience contributing to standards/governance artifacts.
  • Knowledge of proactive monitoring using Azure monitor services, telemetry, and synthetic transactions.
  • Understanding of network architecture and security: WAN/LAN, TCP/IP, PKI.
  • Familiarity with ITSM processes and tools (e.g., ServiceNow), and compliance processes
  • Have AIOps vision and awareness

Responsibilities

  • You will design and define standards, patterns, and automations opportunities that elevate monitoring and reliability across platforms and applications, with a strong focus on Azure Monitor, ServiceNow ITOM Event Management, Grafana, and APM/Synthetics tooling
  • You’ll partner with product teams to implement SLO/SLI‑driven operations, reduce alert noise, accelerate incident response, and embed self‑healing where it matters most.
  • Engineer enterprise monitoring & event patterns by authoring and maintaining reference architectures, runbooks, and event management models (alert → event → incident) with actionable alerts and incidents routing.
  • Contribute to Monitoring and Observability & Event Management Strategy and tooling intake/governance checkpoints and coach product teams
  • Excellent communication skills to drive continuous improvement by reducing alert noise, shorten MTTR, and improve change success by embedding postmortem learnings into patterns, rules, and pipelines.

Skills

Cloud Observability (Azure Monitor)
Grafana
Prometheus
AppDynamics
ServiceNow Event Management
SLI/SLOs and postmortems

Education

Bachelor’s degree in IT/CS/Engineering

Tools

ServiceNow
Grafana
Prometheus
AppDynamics
ThousandEyes

Job description

Qualifications
  • Bachelor’s degree in IT /Computer science/Engineering, or related field.
  • 3+ years in monitoring/observability/SRE roles with hands‑on experience in Azure Monitor/App Insights (KQL) and ServiceNow Event Management.
  • Strong knowledge in Azure Log Analytics, KQL, Telemetry, APM implementations
  • Demonstrated ability to collaborate across IT Operations team, platform, cyber, network, and product teams, strong written verbal communication for standards and enablement.
  • 5+ years of experience with SRE role and deep understanding of monitoring and application performance management
  • Knowledge of SLO platforms (e.g., Nobl9) and experience contributing to standards/governance artifacts.
  • Knowledge of proactive monitoring using Azure monitor services, telemetry, and synthetic transactions.
  • Understanding of network architecture and security: WAN/LAN, TCP/IP, PKI.
  • Familiarity with ITSM processes and tools (e.g., ServiceNow), and compliance processes
  • Have AIOps vision and awareness
Responsibilities
  • You will design and define standards, patterns, and automations opportunities that elevate monitoring and reliability across platforms and applications, with a strong focus on Azure Monitor, ServiceNow ITOM Event Management, Grafana, and APM/Synthetics tooling
  • You’ll partner with product teams to implement SLO/SLI‑driven operations, reduce alert noise, accelerate incident response, and embed self‑healing where it matters most.
  • Engineer enterprise monitoring & event patterns by authoring and maintaining reference architectures, runbooks, and event management models (alert → event → incident) with actionable alerts and incidents routing.
  • Contribute to Monitoring and Observability & Event Management Strategy and tooling intake/governance checkpoints and coach product teams
  • Excellent communication skills to drive continuous improvement by reducing alert noise, shorten MTTR, and improve change success by embedding postmortem learnings into patterns, rules, and pipelines.
Must have Skills
  • Cloud Observability: Azure Monitor/App Insights/Log Analytics (KQL)
  • Knowledge of Grafana, Prometheus, App Dynamics, ThousandEyes
  • Communication & Teaming – Able to translate complex reliability patterns into consumable standards and coach IT operations team via office hours/CoP sessions.
  • Technical Depth in Monitoring and Observability Stack – Hands‑on in ServiceNow Event Management, Azure Monitor/KQL, and automation.
  • Analytical & Systems Thinking – Uses SLI/SLOs, postmortems, and CMDB context to reduce noise, drive self‑healing, and measurably improve MTTR and KPIs.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer- (Monitoring and Event Management)
Site Reliability Engineer- (Monitoring and Event Management)

Pan Asia Resources • Philippines

On-site
PHP 1,200,000 - 2,000,000
Site Reliability Engineer | Hybrid - Centris/Makati
Site Reliability Engineer | Hybrid - Centris/Makati

TASQ • Makati

Hybrid
PHP 1,000,000 - 1,800,000
Site Realibility Engineer
Site Realibility Engineer

Gratitude Philippines • Quezon

Hybrid
PHP 914,000 - 1,095,000
SRE: Observability & Incident-Driven Reliability
SRE: Observability & Incident-Driven Reliability

Pan Asia Resources • Philippines

On-site
PHP 1,200,000 - 2,000,000
Azure Site Reliability Engineer
Azure Site Reliability Engineer

GSS HR Solutions Private Limted • Quezon City

Hybrid
PHP 600,000 - 900,000
IT Solutions Architect – ServiceNow ITOM
IT Solutions Architect – ServiceNow ITOM

Avensys Consulting • Makati

On-site
PHP 1,200,000 - 2,100,000
Now Hiring: Site Reliability Engineer | Hybrid Quezon City
Now Hiring: Site Reliability Engineer | Hybrid Quezon City

Gratitude Philippines • Manila

On-site
PHP 900,000 - 1,200,000
IT Solution Architect (ServiceNow ITOM) | Hybrid - Centris/Makati
IT Solution Architect (ServiceNow ITOM) | Hybrid - Centris/Makati

TASQ • Makati

Hybrid
PHP 1,800,000 - 3,200,000
Monitoring, Observability & Event Management Architect
Monitoring, Observability & Event Management Architect

Gratitude Philippines • Quezon

Hybrid
PHP 2,100,000 - 3,600,000
Observability & ITOM Architect — Hybrid Role
Observability & ITOM Architect — Hybrid Role

Lewis Personnel Management • Makati

Hybrid
PHP 1,800,000 - 2,400,000