A complete application in a minute — tailored resume and cover letter, ready to send.
Get past ATS filters
Job summary
A leading pharmaceutical company is seeking a Major Incident Manager to lead incident management and enhance service reliability. Reporting to senior management, you will coordinate with technical and business teams during critical outages and ensure adherence to SLAs. The ideal candidate will have 15+ years of experience and relevant certifications in ITIL, cloud technologies, and automation. Strong communication and technical skills are essential for this role.
Qualifications
Experience leading Major Incident Management (P1/P2).
Certifications in ITIL or Service Operations preferred.
Familiarity with AIOps tools for automation.
Responsibilities
Lead end-to-end management of Major Incidents.
Enhance service reliability and observability.
Automate operational tasks for improved service recovery.
Skills
Incident Management
Reliability Engineering
Continuous Improvement
Communication
Automation
Education
15+ years' experience
SRE Foundation / Practitioner certification
Cloud certifications (AWS, Azure, GCP)
Tools
Cloud Technologies
Database
Networking
Containerization
Infrastructure as code
Job description
**Main responsibilities:** * **Incident Management*** Lead the end-to-end management of Major Incidents (P1/P2), ensuring timely resolution and effective stakeholder communication.* Act as command centre lead during critical outages, coordinating across technical and business teams.* Ensure accurate and detailed incident documentation, including root cause, timeline and resolution steps.* Drive post-incident-reviews and ensure action items are implemented to prevent recurrence.* Maintain consistent communication and escalation processes aligned with ITSM best practices (e.g. ITIL)* **Reliability Engineering*** Collaborate with service owners and platform teams to enhance service reliability, observability, and fault tolerance.* Implement proactive monitoring, alerting, and automated recovery mechanisms.* Analyse incident trends and develop reliability improvement plans.* Participate in capacity planning, change reviews, and failure mode analysis to anticipate and mitigate risks.* Develop and track SLOs/SLIs/SLAs to measure service health and performance.* **Continuous Improvement*** Partner with problem management to identify recurring issues and lead root cause elimination initiatives.* Automate operational tasks and enhance service recovery using scripts, runbooks, and AIOps tools.* Contribute to the evolution of the Major Incident Process, ensuring best practices are embedded across the organization.* **Key Performance Indicators*** Mean Time to Resolve (MTTR) and Mean Time to Detect (MTTD).* Reduction in number and impact of recurring incidents.* Adherence to SLA/SLO targets.* Completion rate of post-incident actions.* Stakeholder satisfaction and transparency during incidents.* 15+ years' experience.* **Preferred Certifications:*** ITIL v4 or Service Operations certification.* SRE Foundation / Practitioner certification.* Cloud certifications (AWS, Azure, or GCP).* Incident Command System (ICS) or equivalent leadership training in crisis response.* **Soft skills**:* Communication (verbal and written).* **Technical skills**:* Virtualization* Cloud Technologies* Database* Networking* Containerization* Automation* Middleware/Scheduling* Infrastructure as code* **Languages**:* English**Experience**: