D
Major Incident Manager
Riyadh, Riyadh Province, Saudi Arabia · Full Time
Be the first to apply
- Experience
- Any
- Salary
- —
- Openings
- 1
- Posted
- vor 1 Stunde
- Work mode
- In office
- Resume
- Required to apply
Where you'll work
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
Role Overview
As a Major Incident Manager, you will lead 24/7 incident management efforts to minimize disruptions and business impact, ensuring swift resolution and communication during critical service incidents.
Key Responsibilities
- Conduct round-the-clock incident management to reduce service interruptions and mitigate business effects.
- Log and oversee incidents caused by service degradation, outages, or operational alerts.
- Set up and handle communication channels and Incident Bridges for critical issues.
- Coordinate technical troubleshooting, restoration activities, and stakeholder engagement efficiently.
- Facilitate quick decisions and clear operational blockages during incident handling.
- Lead efforts to restore services promptly while following Incident Management procedures.
- Organize and moderate Major Incident Review (MIR) meetings post-incident resolution.
- Record incident timelines, triage steps, recovery measures, lessons learned, and improvement possibilities.
- Manage stakeholder involvement in review meetings and finalize MIR reports.
- Oversee Root Cause Analysis (RCA) for major and recurring incidents involving internal teams, OEMs, third parties, and stakeholders.
- Identify proactively and create problem records based on incident trends and risks.
- Perform in-depth problem analysis and impact evaluations.
- Create and handle problem records documenting root causes, workarounds, corrective actions, and permanent fixes.
- Monitor ongoing problems and track resolution progress within agreed service levels.
- Lead investigative efforts to uncover underlying incident causes.
- Collaborate with technical teams, vendors, OEMs, and regulatory stakeholders to implement lasting solutions.
- Champion proactive remediation initiatives to prevent recurring incidents.
- Ensure implementation and validation of corrective and preventive measures.
- Coordinate temporary fixes to maintain service continuity during permanent solution development.
- Validate the success of workarounds and permanent resolutions.
- Generate regular KPI reports including problem trends, RCA status, resolution timelines, recurring issues, incident counts, and SLA compliance.
- Ensure adherence to Incident and Problem Management policies, operational procedures, and regulatory standards.
- Identify opportunities for process enhancements and automation, providing recommendations to the CSI team.
- Drive initiatives focused on minimizing incident recurrence and boosting service availability.
- Update knowledge bases with validated root causes, known errors, workarounds, and permanent solutions.
- Ensure documentation and communication of lessons learned from major incidents across support teams.
Required Technical Skills
- Expertise in Major Incident Management
- Proficiency in Root Cause Analysis (RCA) techniques
- Strong understanding of Incident and Problem Management practices
- Knowledge of ITIL Service Management Framework
- Ability to analyze trends, report metrics, and perform operational analytics
- Understanding of infrastructure, applications, cloud services, networks, databases, and security operations
- Familiarity with SLAs, KPIs, and service performance management
Performance Indicators
- Decrease in recurring incidents
- High percentage of major incidents with complete RCA
- Timely closure of problems and major incidents within SLA targets
- Reduced Mean Time to Identify Root Cause (MTTRCA)
- Number of proactive problems identified and resolved
- Reduction in incident volume through permanent resolutions
Industry
IT Services & Consulting