Site Reliability Engineer

Gratitude Philippines

Quezon City

Hybrid

PHP 1,200,000 - 1,800,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Hybrid work arrangement
Night shift

Job summary

Gratitude Philippines is seeking an experienced Site Reliability Engineer to drive monitoring, observability, and reliability across platforms. You will leverage Azure Monitor, Application Insights, Log Analytics, KQL, and ServiceNow ITOM to improve MTTR and reduce alert noise.

You will collaborate with IT Operations, Platform, Cybersecurity, Network, and Product teams, applying AIOps concepts to automate patterns and self-healing initiatives. Hybrid work in Quezon City is offered.

Qualifications

  • Bachelor's degree in IT/CS/Engineering or related field.
  • Hands-on experience in SRE, Monitoring or Observability roles.
  • Experience with Azure Monitor, Application Insights, Azure Log Analytics and KQL.
  • Hands-on experience with ServiceNow Event Management / ITOM.
  • Strong understanding of telemetry and APM.
  • Experience with proactive monitoring and synthetic transactions.
  • Collaboration across IT Operations, Platform, Cybersecurity, Network, and Product teams.
  • Knowledge of SLO/SLI concepts and reliability practices.

Responsibilities

  • Design and implement enterprise monitoring standards and automation.
  • Develop observability solutions using Azure tools, Grafana, and APM/Synthetics tooling.
  • Implement proactive monitoring, telemetry, and synthetic transactions.
  • Engineer event management patterns and incident routing with ServiceNow.
  • Partner with product teams to drive SLO/SLI‑driven operations.
  • Apply AIOps concepts to improve monitoring and MTTR.

Skills

SRE
Monitoring & Observability
Telemetry
APM
Collaboration & Stakeholders
Problem Solving

Education

Bachelor's degree in IT/CS/Engineering

Tools

Azure Monitor
Application Insights
Azure Log Analytics
KQL
ServiceNow Event Management
ITOM
Grafana
Prometheus
AppDynamics
ThousandEyes
Nobl9
AIOps
Self-Healing Automation
Event Correlation

Job description

Job Overview

We are looking for an experienced Site Reliability Engineer (SRE) to drive enterprise monitoring, observability, application performance, and reliability initiatives across platforms and applications.

The ideal candidate will have strong hands-on experience with Azure Monitor, Application Insights, Azure Log Analytics, KQL, ServiceNow Event Management, Grafana, APM, and observability tools. This role will focus on improving reliability, reducing alert noise, accelerating incident response, strengthening SLO/SLI-driven operations, and enabling proactive monitoring and automation.

Key Responsibilities
Monitoring, Observability & Reliability
  • Design and define enterprise standards, patterns, and automation opportunities to improve monitoring and reliability across platforms and applications.

  • Develop and maintain monitoring and observability solutions using Azure Monitor, Application Insights, Log Analytics, KQL, Grafana, and APM/Synthetics tooling.

  • Implement proactive monitoring using Azure monitoring services, telemetry, and synthetic transactions.

  • Improve application performance monitoring and overall system reliability.

  • Contribute to observability and event management strategy, tooling intake, and governance activities.

Event Management & Incident Response
  • Engineer enterprise monitoring and event management patterns.

  • Develop and maintain reference architectures, runbooks, and event management models covering alert → event → incident workflows.

  • Design actionable alerts and ensure appropriate incident routing.

  • Work with ServiceNow Event Management to improve event correlation, monitoring, and operational response.

SLO / SLI & Reliability Engineering
  • Partner with product teams to implement SLO/SLI-driven operations.

  • Use service-level objectives, service-level indicators, postmortems, and operational data to identify reliability improvements.

  • Support self-healing initiatives and automation opportunities.

  • Incorporate postmortem learnings into monitoring patterns, rules, standards, and pipelines.

Collaboration & Enablement
  • Collaborate with IT Operations, Platform, Cybersecurity, Network, and Product teams.

  • Translate complex monitoring and reliability concepts into practical and consumable standards.

Continuous Improvement & AIOps
  • Identify opportunities to automate monitoring, event management, and operational processes.

  • Support proactive reliability engineering and self-healing capabilities.

  • Apply AIOps concepts and awareness to improve monitoring, event correlation, and operational efficiency.

  • Use analytics and systems thinking to improve reliability, MTTR, alert quality, and service performance.

Required Qualifications & Experience
  • Bachelor's degree in Information Technology, Computer Science, Engineering, or a related field.

  • Strong experience in Monitoring, Observability, or Site Reliability Engineering (SRE) roles.

  • Hands‑on experience with Azure Monitor, Application Insights, Azure Log Analytics, and KQL.

  • Hands‑on experience with ServiceNow Event Management / ITOM.

  • Strong understanding of telemetry and Application Performance Management (APM).

  • Experience with proactive monitoring and synthetic transactions.

  • Strong understanding of monitoring and application performance management.

  • Experience collaborating across IT Operations, Platform, Cybersecurity, Network, and Product teams.

  • Knowledge of SLO/SLI concepts and reliability engineering practices.

Mandatory / Core Skills
  • Site Reliability Engineering (SRE)

  • Monitoring & Observability

  • Azure Monitor

  • Azure Application Insights

  • Azure Log Analytics

  • KQL

  • Telemetry

  • Application Performance Monitoring (APM)

  • ServiceNow Event Management

  • ServiceNow ITOM

Preferred Skills
  • Grafana

  • Prometheus

  • AppDynamics

  • ThousandEyes

  • Nobl9

  • AIOps

  • Self‑Healing Automation

  • Event Correlation

Soft Skills
  • Strong analytical and systems‑thinking abilities.

  • Excellent written and verbal communication.

  • Strong collaboration and stakeholder management skills.

  • Ability to translate complex technical concepts into practical standards.

  • Coaching and knowledge‑sharing capabilities.

  • Strong problem‑solving and troubleshooting skills.

Work Details
  • Work Location: Eton Centris, Quezon Avenue, Quezon City, Manila

  • Work Setup: Hybrid – 2 Days Work From Home & 3 Days Return to Office

  • Work Schedule: Night Shift

  • Employment Type: Full‑Time

  • Start Date: ASAP

  • Headcount: 1

Eligibility Requirements
  • Bachelor's degree in IT, Computer Science, Engineering, or a related field.

  • Candidates should have relevant experience in SRE, Monitoring, or Observability.

  • Must have hands‑on experience with Azure Monitor / Application Insights / KQL and ServiceNow Event Management.

  • Must be amenable to a hybrid work setup in Quezon City.

  • Must be willing to work on a night shift schedule.

  • Candidates should demonstrate stable employment history and career progression.

  • Candidates who are current or former employees of Wipro are not eligible for this opportunity.

Benefits & Other Information
  • Hybrid work arrangement with 2 WFH and 3 RTO days.

  • Opportunity to work on enterprise-level monitoring, observability, and reliability initiatives.

  • Exposure to Azure cloud monitoring and ServiceNow ITOM/Event Management.

Recruitment Process
  1. Paper Screening – Profile is endorsed to the Operations team to determine whether the candidate meets the basic requirements and qualifications for the position.

  2. L1 Interview – Interview with the Practice / Operations Team.

  3. L2 Interview – Optional, depending on the recruitment process.

  4. Technical Assessment / Final Interview – Conducted by the customer's Operations team.

Pre‑Screening Questions
  1. What is your highest educational attainment?

  2. How many years of relevant experience do you have in Monitoring, Observability, or SRE roles, with hands‑on experience in Azure Monitor / Application Insights (KQL) and ServiceNow Event Management?

  3. How many years of relevant experience do you have in SRE roles with a strong understanding of monitoring and Application Performance Management (APM)?

  4. What was your last drawn salary?

  5. What is your salary expectation?

  6. Are you amenable to working on a hybrid setup in Quezon City with a night shift schedule?

  7. When are you available to start once hired?

  8. Are you currently or formerly employed by Wipro?

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Realibility Engineer
Site Realibility Engineer

Gratitude Philippines • Quezon City

Hybrid
PHP 914,000 - 1,095,000
Site Reability Engineer
Site Reability Engineer

Stafflink Express • Philippines

Hybrid
PHP 914,000 - 1,095,000
Site Reliability Engineer | Hybrid - Centris/Makati
Site Reliability Engineer | Hybrid - Centris/Makati

TASQ • Makati

Hybrid
PHP 900,000 - 1,500,000
Monitoring, Observability & Event Management Architect
Monitoring, Observability & Event Management Architect

Gratitude Philippines • Quezon City

Hybrid
PHP 3,000,000 - 5,000,000
Azure SRE: Reliability Engineer (Hybrid, Night Shift, Manila)
Azure SRE: Reliability Engineer (Hybrid, Night Shift, Manila)

Stafflink Express • Philippines

Hybrid
PHP 914,000 - 1,095,000
Network Engineer
Network Engineer

Gratitude Philippines • Quezon City

Hybrid
PHP 1,897,000 - 1,964,000
Site Reliability Engineer
Site Reliability Engineer

EROAD • Manila

On-site
PHP 1,200,000 - 1,800,000
Site Reliability Engineer
Site Reliability Engineer

Coretex • Manila

On-site
PHP 1,200,000 - 2,000,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Procter & Gamble • Philippines

On-site
PHP 1,200,000 - 1,500,000
Performance bonus (STAR)
Flexible work schedule with work-from-
Health insurance
+2
Site Reliability Engineering | Onsite in Manila - ATCP-1429180-S424201(CL8,9,10)
Site Reliability Engineering | Onsite in Manila - ATCP-1429180-S424201(CL8,9,10)

Gratitude Philippines • Manila

On-site
PHP 725,000 - 1,674,000