Site Reliability Engineer (SRE)

Gratitude Philippines

Quezon City

Hybrid

PHP 1,000,000 - 1,800,000

Full time

5 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Gratitude Philippines is seeking a Site Reliability Engineer (SRE) to design, implement, and govern monitoring and reliability standards across Azure-based systems. You will work to reduce alert noise, accelerate incident response, and establish enterprise monitoring practices.

The role requires hands-on experience with Azure Monitor, Application Insights, Log Analytics, KQL, Grafana, and ITOM/Event Management, and offers a hybrid work arrangement in Quezon City on a night shift.

Qualifications

  • Bachelor's degree in IT, Computer Science, Engineering, or a related field is required.
  • 3+ years of experience in Monitoring, Observability, or SRE roles.
  • Hands-on experience with Azure Monitor, Application Insights/KQL, and ServiceNow Event Management.
  • Strong knowledge of Azure Log Analytics, telemetry, APM, and proactive monitoring.
  • 5+ years of SRE experience is preferred for deep monitoring expertise.
  • Knowledge of SLO/SLI platforms such as Nobl9 is an advantage.
  • Hands-on knowledge of ITSM processes and ServiceNow.
  • Understanding of network architecture and security, including WAN/LAN, TCP/IP, and PKI.
  • Awareness of AIOps concepts and vision.
  • Strong communication and collaboration skills.
  • Must be willing to work hybrid in Quezon City on a night shift.
  • Must not be a current or former Wipro employee.
  • Stable employment history; candidates should not be frequent job hoppers.

Responsibilities

  • Design and define monitoring, observability, reliability standards, patterns, and automation opportunities across platforms and applications.
  • Work extensively with Azure Monitor, ServiceNow ITOM Event Management, Grafana, APM, and synthetic monitoring tools.
  • Partner with product teams to implement SLI/SLO-driven operations.
  • Reduce alert noise and improve incident response and overall system reliability.
  • Develop and maintain reference architectures, runbooks, monitoring standards, and event‑management models.
  • Establish actionable alert and incident routing using the alert → event → incident model.
  • Contribute to Monitoring, Observability, and Event Management strategy and governance.
  • Coach product and IT operations teams on monitoring and reliability best practices.
  • Use postmortem learnings to improve monitoring rules, patterns, and automation pipelines.
  • Support initiatives focused on reducing MTTR and improving change success and operational KPIs.
  • Identify opportunities for self‑healing and proactive monitoring.

Skills

Azure Monitor
Application Insights
Azure Log Analytics
KQL
ServiceNow Event Management / ITOM
Cloud Observability
Monitoring & Application Performance
Grafana
Prometheus
AppDynamics
ThousandEyes
Telemetry and APM
Synthetic Monitoring
SLI/SLO concepts
ITSM processes
Monitoring automation
AIOps awareness

Education

Bachelor's degree in IT, Computer Science, Engineering, or a related field

Tools

Azure Monitor
Application Insights
Azure Log Analytics
KQL
ServiceNow ITOM
Grafana
Prometheus
AppDynamics
ThousandEyes
Telemetry and APM
Synthetic Monitoring

Job description

Site Reliability Engineer (SRE)
JOB SUMMARY

We are looking for a Site Reliability Engineer (SRE) with strong experience in monitoring, observability, application performance management, and Azure monitoring technologies. The role will focus on improving reliability, reducing alert noise, accelerating incident response, and establishing enterprise monitoring and event-management standards.

The ideal candidate should have hands‑on experience with Azure Monitor, Application Insights, Log Analytics, KQL, ServiceNow Event Management, Grafana, and APM/Synthetic monitoring tools.

IMPORTANT ELIGIBILITY REQUIREMENTS
  • Bachelor's degree in IT, Computer Science, Engineering, or a related field.

  • 3+ years of relevant experience in Monitoring, Observability, or SRE roles.

  • Hands‑on experience with Azure Monitor, Application Insights, KQL, and ServiceNow Event Management.

  • Strong knowledge of Azure Log Analytics, telemetry, APM, and proactive monitoring.

  • 5+ years of SRE experience is preferred for candidates with deep monitoring and application performance expertise.

  • Knowledge of SLO/SLI platforms such as Nobl9 is an advantage.

  • Hands‑on knowledge of ITSM processes and ServiceNow.

  • Understanding of network architecture and security, including WAN/LAN, TCP/IP, and PKI.

  • Awareness of AIOps concepts and vision.

  • Strong communication and collaboration skills.

  • Must be willing to work hybrid in Quezon City on a night shift.

  • Must not be a current or former Wipro employee.

  • Stable employment history; candidates should not be frequent job hoppers.

KEY RESPONSIBILITIES
  • Design and define monitoring, observability, reliability standards, patterns, and automation opportunities across platforms and applications.

  • Work extensively with Azure Monitor, ServiceNow ITOM Event Management, Grafana, APM, and synthetic monitoring tools.

  • Partner with product teams to implement SLI/SLO-driven operations.

  • Reduce alert noise and improve incident response and overall system reliability.

  • Develop and maintain reference architectures, runbooks, monitoring standards, and event‑management models.

  • Establish actionable alert and incident routing using the alert → event → incident model.

  • Contribute to Monitoring, Observability, and Event Management strategy and governance.

  • Coach product and IT operations teams on monitoring and reliability best practices.

  • Use postmortem learnings to improve monitoring rules, patterns, and automation pipelines.

  • Support initiatives focused on reducing MTTR and improving change success and operational KPIs.

  • Identify opportunities for self‑healing and proactive monitoring.

MUST-HAVE SKILLS
  • Azure Monitor

  • Application Insights

  • Azure Log Analytics

  • KQL

  • ServiceNow Event Management / ITOM

  • Cloud Observability

  • Monitoring & Application Performance Management

  • Grafana

  • Prometheus

  • AppDynamics

  • ThousandEyes

  • Telemetry and APM

  • Synthetic Monitoring

  • SLI/SLO concepts

  • ITSM processes

  • Monitoring automation

  • AIOps awareness

KEY COMPETENCIES
  • Communication & Teamwork: Ability to explain complex reliability concepts through standards, documentation, office hours, and community-of-practice sessions.

  • Technical Depth: Strong hands‑on knowledge of monitoring, observability, ServiceNow Event Management, Azure Monitor/KQL, and automation.

  • Analytical & Systems Thinking: Ability to use SLI/SLOs, postmortems, and CMDB context to reduce alert noise, improve self‑healing, and enhance MTTR and operational KPIs.

RECRUITMENT PROCESS
  1. Paper Screening – Profile review by the Operations team against basic requirements.

  2. L1 Interview – Practice / Operations Team.

  3. L2 Interview – Optional.

  4. Technical Assessment / Final Interview – Customer's Operations Team.

PRE-SCREENING QUESTIONS
  1. What is your highest educational attainment?

  2. How many years of relevant experience do you have in Monitoring, Observability, or SRE roles?

  3. How many years of hands‑on experience do you have with Azure Monitor, Application Insights/KQL, and ServiceNow Event Management?

  4. How many years of SRE experience do you have with monitoring and application performance management?

  5. What is your current/last drawn salary?

  6. What is your expected salary?

  7. Are you amenable to working in a hybrid setup in Quezon City, with 2 WFH and 3 RTO days, on a night shift?

  8. When are you available to start if selected?

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

Gratitude Philippines • Quezon City

Hybrid
PHP 1,200,000 - 1,800,000
Hybrid work arrangement
Night shift
Site Realibility Engineer
Site Realibility Engineer

Gratitude Philippines • Quezon City

Hybrid
PHP 914,000 - 1,095,000
(ACTIVE) Site Reliability Engineer | Hybrid Makati city
(ACTIVE) Site Reliability Engineer | Hybrid Makati city

Gratitude Philippines • Manila

On-site
PHP 1,800,000 - 3,000,000
Gratitude Philippines benefits
Site Reability Engineer
Site Reability Engineer

Stafflink Express • Philippines

Hybrid
PHP 914,000 - 1,095,000
Site Reliability Engineer | Hybrid - Centris/Makati
Site Reliability Engineer | Hybrid - Centris/Makati

TASQ • Makati

Hybrid
PHP 900,000 - 1,500,000
Azure Monitoring & Reliability SRE — Hybrid Night Shift
Azure Monitoring & Reliability SRE — Hybrid Night Shift

Gratitude Philippines • Quezon City

Hybrid
PHP 1,000,000 - 1,800,000
Azure SRE: Reliability Engineer (Hybrid, Night Shift, Manila)
Azure SRE: Reliability Engineer (Hybrid, Night Shift, Manila)

Stafflink Express • Philippines

Hybrid
PHP 914,000 - 1,095,000
Azure SRE: Observability, Monitoring & Incident Response
Azure SRE: Observability, Monitoring & Incident Response

Gratitude Philippines • Quezon City

Hybrid
PHP 1,200,000 - 1,800,000
Hybrid work arrangement
Night shift
Azure Observability & SRE Engineer
Azure Observability & SRE Engineer

Gratitude Philippines • Quezon City

Hybrid
PHP 914,000 - 1,095,000
SRE: Cloud, Observability & Automation
SRE: Cloud, Observability & Automation

Trinity Workforce Solutions, Inc. • Makati

On-site