Incident Response Lead

Permhunt

Cebu City

On-site

PHP 900,000 - 1,700,000

Full time

47 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Permhunt in the Philippines is seeking an experienced Incident & Problem Manager to lead end-to-end incidents across a high-tech, multi-cloud environment. You will triage, coordinate, and communicate in real time, acting as Incident Commander and driving rapid restoration.

You will improve MTTR metrics, develop post-incident reviews, and partner with SREs and engineering to tighten runbooks and automation. Requires ITIL familiarity, 3+ years of leadership in 24x7 ops, and a degree.

Qualifications

  • Bachelor’s degree in CS/IT/Engineering or related field.
  • 3+ years in service delivery, incident response, or operations leadership in enterprise-scale, 24×7 environments.
  • Strong grounding in ITSM / ITIL principles (Incident & Problem Management).

Responsibilities

  • Lead end-to-end incident response, triage, communication, and resolution in real time.
  • Act as Service Operations & Reliability Incident Commander for high-impact events across a global environment.
  • Track and improve metrics like MTTD, MTTR, and MTBM.
  • Champion blameless PIRs and translate learnings into long-term system and process improvements.

Skills

Incident leadership
Communication under pressure
Stakeholder management
Analytical thinking

Education

Bachelor's degree in CS/IT/Engineering

Tools

PagerDuty
Datadog
Grafana
Splunk
ServiceNow

Job description

Our clientis a technical consulting company specializing in operational services for the high-tech industry. They specialize inhelping platform and infrastructure teams operate multi-cloud environments, execute complex migrations, and enable seamless app deployments.

Your Role
Incident & Problem Management
  • Lead end-to-end incident response, triage, communication, and resolution in real time.
  • Act as Service Operations & Reliability Incident Commander for high-impact events across a global environment.
  • Track and improve metrics like MTTD, MTTM, and MTTR.
  • Champion blameless Post-Incident Reviews (PIRs) and translate learnings into long-term system and process improvements.
Service Operations & Reliability
  • Oversee daily service health, capacity, and reliability across all supported environments.
  • Ensure compliance with operational KPIs through proactive planning and improvement.
  • Balance demand vs. capacity and manage shift coverage to prevent burnout.
  • Partner with engineering teams to maintain runbooks, knowledge bases, and escalation paths.
  • Drive automation and workflow optimization to reduce manual overhead.
  • Use data insights to guide decisions and improvements.
Strategic & Cross-Functional Impact
  • Represent in customer reviews, operational syncs, and briefings.
  • Collaborate with SREs, product owners, and partner engineers to align priorities and reliability goals.
  • Contribute to frameworks and governance initiatives.
  • Lead service onboarding/off-boarding and strengthen operational readiness checkpoints.
  • Identify and close systemic operational gaps through process and tool improvements.
Your Qualifications
  • Bachelor's degree in Computer Science, Information Technology, Engineering, or a related discipline.
  • 3+ years in Service Delivery, Incident Response, or Operations Leadership within enterprise-scale, 24×7 environments.
  • Proven experience managing technical teams, driving performance, and leading through critical situations.
  • Strong grounding in ITSM / ITIL principles (Incident & Problem Management).
  • Familiarity with cloud, distributed systems, or enterprise infrastructure.
  • Skilled in monitoring, alerting, and ticketing tools (e.g., PagerDuty, Datadog, Grafana, Splunk, ServiceNow).
Core Competencies
  • Incident Command and Escalation Management
  • Analytical and Problem-Solving Skills
  • Communication and Decision-Making Under Pressure
  • Root Cause and Post-Incident Analysis
  • Operational Planning and Service Governance
  • Stakeholder and Partner Management
  • IT Service Management (Incident & Problem Management)
  • Observability, Monitoring, and Automation Tools
Plus points if you have:
  • ITIL V3 or V4 certification or Incident Management Certifications
  • Familiarity in SRE practices and operational frameworks that promote reliability and automation
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Incident Manager
Incident Manager

Talentium Inc. • Muntinlupa

On-site
Incident & Change Lead — IT Operations & SLA Expert
Incident & Change Lead — IT Operations & SLA Expert

The AM3 Global Singapore/USA • Muntinlupa

On-site
PHP 600,000 - 800,000
Incident Manager
Incident Manager

The AM3 Global Singapore/USA • Muntinlupa

On-site
PHP 600,000 - 800,000
Coordinator, Support Ops
Coordinator, Support Ops

Kroll Global Solutions Inc. • Manila

Hybrid
PHP 420,000 - 780,000
Consulting_Cyber Detection & Response IRR Senior
Consulting_Cyber Detection & Response IRR Senior

EY • Taguig

On-site
PHP 900,000 - 1,200,000
Health and wellness packages
Opportunities for continuous learning
Access to cutting-edge technologies
Incident and Problem Management Dept. Head/Manager | Networld Capital Ventures, Inc.
Incident and Problem Management Dept. Head/Manager | Networld Capital Ventures, Inc.

pj lhuillier group of companies • Philippines

On-site
PHP 1,800,000 - 3,000,000
Incident Manager|Hybrid - Alabang
Incident Manager|Hybrid - Alabang

Hunter's Hub Inc. • Muntinlupa

On-site
PHP 700,000 - 900,000
Site Reliability Engineer
Site Reliability Engineer

Infojini Inc • Mexico

Hybrid
MXN 900,000 - 1,300,000
IT.IT Quality.Problem Management.Analyst
IT.IT Quality.Problem Management.Analyst

Alorica • Taguig

On-site
PHP 900,000 - 1,300,000
Systems Engineer
Systems Engineer

Cognizant • Manila

On-site
PHP 800,000 - 1,000,000