SRE Manager

Quality Ai

Hinoba-an

On-site

PHP 2,653,000 - 4,642,000

Full time

4 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

QualityAI is seeking a hands-on SRE Lead to own reliability across the NewLeaf Azure application platform, covering websites, APIs, and integrations.

You will drive incident management, self-healing automation, on-call readiness, and operational reporting to build a resilient product ecosystem.

You will collaborate with engineering, platform, and product teams to reduce operational risk, enforce runbooks, and raise overall reliability maturity.

Qualifications

  • Hands-on SRE leader with ownership of production systems.
  • Strong incident management including postmortem rigor.
  • Observability across logs, metrics, traces, dashboards and alerts.
  • Ability to build or influence automation for detection, triage and remediation.
  • Experience shaping on-call practices and escalation paths.
  • Comfort operating across services, integrations, async processing, and releases.
  • Clear communication translating risk into operational language.
  • Ability to coach engineers and raise team maturity.

Responsibilities

  • Define the reliability strategy for the NewLeaf platform aligned with business priorities.
  • Own incident detection, triage, escalation, communication and post-incident follow-through.
  • Build operational dashboards showing service health, user impact and reliability trends.
  • Design self-healing and auto-remediation workflows for recurring failures.
  • Create and maintain runbooks, playbooks and a living reliability rulebook.
  • Design and support a sustainable on-call rotation model for engineering teams.
  • Drive root-cause analysis and convert recurring issues into permanent fixes or automation.
  • Partner with application, integration and platform teams to reduce operational risk.

Skills

Ownership of production systems
Incident management
Observability
Automation
On-call practices
Release workflows
Clear communication
Coaching engineers

Education

Bachelor’s degree in Computer Science or related

Tools

Azure
Azure Monitor
App Insights
Log Analytics
Terraform/Bicep

Job description

SRE Manager

Date: 4 Sept 2026 Company: QualityAI Country/Region: IN


SRE Lead, Platform Reliability and OperationsNewLeaf Azure application platform


This is the role specification for a senior reliability leader who can combine SRE discipline, platform operations, and automation-first thinking across the NewLeaf Azure application estate.


Role Summary


  • We are hiring a hands-on SRE Lead to own reliability across the full NewLeaf application platform on Azure.

  • This role covers customer-facing websites, backend services, APIs, integrations, event-driven workflows, observability, incident management, self-healing automation, on-call operations, and operational reporting.

  • The goal is to build a resilient operating model for the entire product ecosystem, not just the underlying cloud infrastructure.


What \"Platform\" Means Here


  • Public and internal websites

  • Backend services and APIs

  • Azure-hosted application workloads

  • Integrations and event processing

  • Monitoring, alerting, dashboards, and reporting

  • Incident response, postmortems, and reliability improvements

  • Runbooks, operational standards, and self-healing automation

  • On-call rotation design and engineering readiness


Primary Responsibilities


  • Define the reliability strategy for the NewLeaf platform and keep it aligned with business priorities.

  • Own incident detection, triage, escalation, communication, and post-incident follow-through.

  • Build operational dashboards that show service health, user impact, and reliability trends.

  • Design self-healing and auto-remediation workflows for recurring failure modes.

  • Create and maintain runbooks, playbooks, and a living reliability rulebook.

  • Design and support a sustainable on-call rotation model for engineering teams.

  • Drive root-cause analysis and convert recurring issues into permanent fixes or automation.

  • Partner with application, integration, and platform teams to reduce operational risk across boundaries.


Required Skill Set


  • Strong SRE or platform engineering background with direct ownership of production systems.

  • Deep incident management experience, including severity classification, incident command, and postmortem rigor.

  • Practical observability expertise across logs, metrics, traces, dashboards, alerting, and service health reporting.

  • Ability to build or influence automation for detection, triage, remediation, and follow-up tasks.

  • Experience shaping on-call practices, escalation paths, and team readiness.

  • Comfort operating across application services, integrations, asynchronous processing, and release workflows.

  • Clear communication skills for translating technical risk into operational and executive language.

  • Ability to coach engineers and raise the maturity of a team without creating unnecessary process overhead.


Azure and Platform Stack Familiarity


  • Microsoft Azure operations and governance.

  • Azure Monitor, Application Insights, and Log Analytics.

  • Azure Service Bus, Event Grid, Event Hub, and timer-driven/background processing patterns.

  • Azure Functions, App Service, and other cloud-hosted application runtimes.

  • Key Vault, Storage, and identity-aware service configuration.

  • CI/CD pipelines and safe deployment practices, including release validation and rollback readiness.

  • Infrastructure-as-code tooling such as Bicep or Terraform, where used by the team.

  • Kusto Query Language and operational reporting from telemetry data.


Agentic Operations and Automation


  • Use agents or automation to assist with analysis, incident summarization, runbook execution, and follow-up work.

  • Treat automation as a reliability control, not a novelty layer.

  • Apply safety checks, approvals, logging, and rollback paths before any automated remediation is allowed to act on production systems.

  • Continuously improve detection and remediation based on real incident patterns and post-incident learning.


Leadership and Collaboration


  • Partner with engineering leaders to improve reliability without slowing delivery unnecessarily.

  • Lead cross-team incident coordination when production issues span multiple systems.

  • Establish practical standards for uptime, alert quality, runbook quality, and operational readiness.

  • Help build a culture where reliability is owned by every team, not delegated to a single hero.


Nice-to-Have Experience


  • Experience introducing self-healing or auto-remediation in production.

  • Experience with executive dashboards and operational scorecards.

  • Experience in multi-team or multi-product environments with complex dependencies.

  • Background in Azure-native environments, distributed systems, and event-driven architectures.

  • Familiarity with AI-assisted operations or agent-based workflows in a safe enterprise setting.


What Success Looks Like


  • Incidents are detected faster and resolved faster.

  • Recurring reliability issues are systematically eliminated or automated away.

  • On-call is sustainable, documented, and well supported.

  • Engineers have clear runbooks and confidence during outages.

  • Leadership has accurate visibility into platform health and risk.

  • Reliability becomes a repeatable operating discipline instead of a hero exercise.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior SRE Lead: Platform Reliability & Automation
Senior SRE Lead: Platform Reliability & Automation

Quality Ai • Hinoba-an

On-site
PHP 2,653,000 - 4,642,000
Software Engineer- AI-Driven SRE & Cloud SRE
Software Engineer- AI-Driven SRE & Cloud SRE

Keka Technologies Private Limited • Mexico

On-site
PHP 2,215,000 - 3,322,000
Site Reliability Engineer
Site Reliability Engineer

IDEMIA • Philippines

On-site
PHP 900,000 - 1,500,000
Technical Lead - Site Reliability Engineering
Technical Lead - Site Reliability Engineering

LSEG • Taguig

On-site
PHP 4,914,000 - 7,372,000
Healthcare
Retirement planning
Paid volunteering days
+1
Site Reliability Engineer
Site Reliability Engineer

IDEMIA PHILIPPINES INC. • Philippines

On-site
PHP 900,000 - 1,350,000
Reliability Operations Engineer (Philippines)
Reliability Operations Engineer (Philippines)

Serve Robotics • Philippines

On-site
PHP 900,000 - 1,500,000
Site Reliability Engineer
Site Reliability Engineer

AgileEngine • Mexico

Hybrid
PHP 8,734,000 - 13,100,000
Professional growth
Competitive USD-based pay
Exciting projects
+1
Reliability Operations Engineer (Malaysia)
Reliability Operations Engineer (Malaysia)

Industrious Ventures • Philippines

On-site
PHP 1,200,000 - 1,800,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

8x8 • Manila, Hinoba-an

On-site
PHP 1,000,000 - 2,000,000
VS01700 - SRE & Production Reliability Engineer
VS01700 - SRE & Production Reliability Engineer

E4 Software Services Pvt Ltd. • Hinoba-an

On-site
PHP 893,000 - 1,674,000