This page was automatically translated and may contain errors. View in English.
malomatia

Engineer – Incident & Problem Management

malomatia

Qatar · À temps plein

Soyez le premier à postuler

Expérience
4 à 6 ans
Salaire
Ouvertures
1
Publié
il y a 1 heure
Mode de travail
Au bureau
Éducation
Bachelor's degree in IT or related discipline
CV
Candidature requise

Votre lieu de travail

Description de l'emploi

Overview

This role focuses on enhancing the effectiveness, coordination, and continual improvement of Incident Management, Major Incident Management, and both Reactive and Proactive Problem Management for the assigned services. The engineer ensures timely and accurate handling of incidents—capturing, prioritizing, escalating, communicating, restoring, validating, and closing incidents within agreed service levels. It also involves investigating recurring and critical issues via root-cause analysis and corrective actions to improve service reliability, customer experience, and operational resilience, all under the supervision of senior leadership.

Key Responsibilities

  • Assist with daily administration and consistent execution of processes related to Incident Management, Major Incident Management, and Problem Management, including managing workflows, templates, escalation procedures, and quality controls.
  • Validate incident records for accuracy in logging, correct categorization, impact and urgency assessments, priority settings, assignments, timestamps, configuration links, diagnostic data, resolution information, validation, and proper closure documentation.
  • Monitor critical incidents that are high-priority, aging, breached SLAs, reopened, sensitive, or dependent on suppliers; coordinate follow-up with responsible teams and escalate delays or risks as needed.
  • Support declaration of major incidents, coordinate bridge calls, activate relevant participants, log timelines and actions, track restoration, communicate with stakeholders, and assist in official incident closure under guidance.
  • Coordinate among service desk, technical, cloud, infrastructure, application, cybersecurity, supplier, customer, and business units to facilitate quick and safe service restoration, ensuring emergency changes, workarounds, failovers, and recovery steps are authorized and documented.
  • Gather logs, monitoring data, decision records, impact details, timelines, and evidence required for major-incident reports, post-incident reviews, customer reports, and audits; document lessons learned and monitor follow-up actions.
  • Identify potential problem candidates through analysis of recurring incidents, trends, failed changes, post-release issues, supplier performance, capacity or availability alarms, and operational risks.
  • Support structured root-cause analyses using approved methodologies; document findings and monitor corrective and preventive measures to ensure closure.
  • Maintain comprehensive records including problem backlogs, known errors, workarounds, fix statuses, ownerships, deadlines, risks, dependencies, and links across ITSM systems including incident, problem, change, knowledge, supplier, service, and configuration management.
  • Ensure known errors and verified workarounds are published and kept current for service desk and support teams to enhance first-line resolutions.
  • Generate operational reports and dashboards detailing incident trends, SLA metrics, response and resolution timelines, backlog ages, reopen rates, major/repeat incident stats, problem statuses, and closure progress.
  • Prepare documentation such as meeting materials, minutes, decision logs, and action trackers for incident review meetings, problem backlog discussions, supplier engagements, and service improvement gatherings.
  • Contribute to enhancements in monitoring, event management, automation, runbooks, knowledge management, shift-left practices, swarming techniques, and self-healing processes that expedite detection, diagnosis, restoration, and prevention.
  • Maintain controlled documentation for procedures, priority models, communication templates, escalation contact lists, major-incident checklists, root-cause analysis templates, and knowledge articles.
  • Assist with quality reviews, compliance audits, supplier coordination, preparation of customer or audit evidence, and closure of identified gaps and corrective actions.
  • Identify recurring process weaknesses, data quality problems, bottlenecks, supplier delays, and automation opportunities; actively support practical process improvements.

Required Education, Certifications, and Knowledge

  • Bachelor’s degree in information technology, computer science, software engineering, engineering, information systems, or related discipline.
  • ITIL 4 Foundation certification mandatory; additional training in Incident Management, Major Incident Management, Problem Management, or related ITSM modules preferred.
  • Additional certifications or training in Site Reliability Engineering, SIAM, COBIT, ISO/IEC 20000, Lean Six Sigma, risk management, business continuity, project management, or information security are advantages.
  • Comprehensive understanding of IT service management, including operation, transition, service reliability, customer experience management, service levels, supplier obligations, and continuous improvement.

Technical Expertise

  • In-depth working knowledge of incident and problem life cycles, roles and responsibilities, priority frameworks, escalation protocols, governance, and documentation standards.
  • Proficiency in reviewing incident and problem records for completeness, consistency, SLA adherence, proper traceability, and compliant escalation and closure.
  • Hands-on experience coordinating major incidents, managing bridge calls, documenting timelines and actions, updating stakeholders, monitoring restoration progress, and conducting post-incident documentation.
  • Competency in facilitating root-cause analysis techniques, distinguishing symptoms from causes, and managing corrective and preventive action tracking.
  • Familiarity with enterprise ITSM tools, monitoring and observability platforms, knowledge repositories, dashboards, workflow automation, CMDBs, service mapping, and audit trail maintenance.
  • Good understanding of infrastructure technologies spanning cloud, networks, telecommunications, cybersecurity, databases, middleware, applications, and end-user computing to effectively coordinate multi-disciplinary teams.
  • Ability to develop operational dashboards, incident trend reports, root-cause analysis documents, meeting records, management presentations, and audit-compliant supporting evidence.
  • Knowledge of analytical methods such as Pareto charts, 5 Whys, fishbone diagrams, chronological analysis, causal mapping, and foundational data analysis techniques.

Soft Skills and Behavioural Attributes

  • Exceptional verbal and written communication skills that can translate technical issues into clear business impact, risk, and resolution information.
  • Strong facilitation, coordination, and stakeholder management abilities across technical teams, customers, service managers, and external suppliers.
  • Excellent organizational skills with the ability to prioritize tasks, manage multiple simultaneous incidents, problems, and deadlines effectively.
  • Analytical mind with attention to detail capable of working under uncertain or rapidly changing conditions.
  • Customer-centric approach emphasizing collaboration, expectation management, empathetic communication, timely escalation, and thorough follow-up.
  • Maintains composure, structure, responsiveness, and credibility during high-pressure major incidents and competing priorities.
  • Demonstrates ownership, accountability, integrity, confidentiality, and strict adherence to established procedures.
  • Balances the trade-offs between fast restoration, customer impact, technical risks, security, compliance, and sustainability in solutions.
  • Displays curiosity, persistence, critical thinking, adaptability, and a commitment to resolving root causes rather than accepting temporary fixes.
  • Encourages collaborative working environments fostering swarming, transparency, shared learning, psychological safety, and accountability.

Experience Requirements

  • Between 4 to 6 years of relevant professional experience in IT service operations, Incident Management, Major Incident Management, Problem Management, service desk or service reliability roles.
  • Minimum of 2 years of hands-on experience in coordinating priority incidents, managing major incidents, maintaining problem records, supporting root-cause investigations, and tracking corrective actions.
  • Experience using enterprise ITSM platforms and associated tools such as monitoring or observability systems, dashboards, knowledge bases, CMDBs, and automation workflows.
  • Exposure to coordinating with cross-functional teams including infrastructure, cloud, application, network, cybersecurity, service desk, customers, and suppliers in operational or 24x7 environments.
  • Experience supporting operational reporting, customer communications, audits, conducting post-incident reviews, managing problem backlogs, known-error documentation, and continuous improvement initiatives is a plus.

Laissez ce message si vous souhaitez une réponse — nous ne l'utiliserons à aucune autre fin.

Cliquez pour parcourir, glisser-déposer, ou coller une capture d'écran

PNG, JPG, GIF, MP4, WebM, MOV · 20 Mo maximum par fichier · Jusqu'à 5 fichiers

🤖
En ligne · Aide IA instantanée