Contact

Community

Integrated ITSM Framework in the AI Era (Part 2)
  • Date2026.08.04

In [Part 1], we explored three core frameworks of the next-generation ITSM that must evolve in step with the AI era: from the upstream planning stage of 'Service Strategy and Value Management', to the frontline touchpoint of 'Service Support', and 'Change and Release Management', which facilitates safe expansion.


If the frameworks introduced in [Part 1] focused on setting the direction of IT services and safely implementing them, this [Part 2] will cover the 'value chain of operations and innovation.' This ensures the implemented services run flawlessly 24/7, provide optimal user experiences, and self-evolve through AI technology. I will introduce the remaining four core frameworks of the practical ITSM operating model that organically connects the entire process from service strategy to ensuring continuity.




'Operational Stability Management' for Uninterrupted Business


No matter how brilliantly a service is planned and deployed, if an outage or performance degradation occurs during actual operations, it deals a fatal blow to the business. Operational Stability Management goes beyond the passive activity of staring at monitoring tool screens and watching for error logs; it is a core area that proactively maintains optimal operating conditions so that services can be seamlessly delivered at promised levels.


In the field, management often remains limited to fragmented, infrastructure-centric metrics like CPU, memory, and disk utilization. However, true operational stability must be evaluated based on comprehensive, service-oriented metrics such as response time, transaction success rates, and actual user impact. When AI and AIOps technologies are integrated into this, they analyze massive amounts of system logs and event metrics to detect subtle abnormal patterns that humans might miss, pre-emptively blocking the possibility of failures.



For successful Operational Stability Management, the following activities must be organically aligned:


Monitoring and Event Management: Monitors system status 24/7, analyzes the severity and business impact of detected events, and links them to optimal response procedures.


Availability Management: Goes beyond simple system uptime; it consistently maintains agreed-upon service levels with the goal of ensuring a state where users can smoothly utilize actual services.


Capacity and Performance Management: Predicts future increases in service demand to prevent resource shortages or performance degradation, securing necessary resources in advance.


Backup and Recovery Management: Perfectly protects data against security incidents, failures, or data loss situations, and regularly verifies actual recoverability beyond just executing routine backups.




'Experience and Access Management': Beyond SLAs to Perceived Quality


Past IT operators believed that service quality was good as long as technical metrics (SLAs) like system uptime or incident resolution time were met. However, today's users are different. Just as important as how fast a request was processed is how intuitive the application process was and whether they can transparently track the progress of their request—these are the true measures of service quality.


Experience and Access Management is an area that maximizes accessibility and perceived satisfaction, going beyond normal system operation to ensure users can utilize necessary services timely, conveniently, and safely. By leveraging AI, it can analyze a user's past request patterns to recommend necessary services proactively, or use intelligent chatbots to kindly guide them through complex application procedures.



The main detailed processes for innovating user experience are as follows:


Self-Service Portal: Linked with the service catalog, it provides an intuitive, single-pane-of-glass interface where users can easily search, request, and track necessary IT services themselves.


User Privilege and Access Management: Balances convenience and security by controlling access based on the principle of least privilege, granting only the necessary permissions tailored to job duties and roles.


SLA Management: Defines clear expectations between service providers and users, objectively measuring them to serve as a baseline for performance management.


Customer Experience and Satisfaction Management: Conducts in-depth analysis of satisfaction surveys and user feedback data to evaluate service quality strictly from the user's perspective and derive areas for improvement.




'AI-Based Intelligence Management': From Simple Tool to Subject of Control


AI-based Intelligence Management is an area that maximizes efficiency by integrating generative AI, automation, and data analysis technologies into the IT service operation process. The scope of AI application is boundless, ranging from automated ticket classification at the service desk to anomaly detection based on operational data, and predictive analysis of change risks.


The most important thing to note is that AI should not be regarded merely as a simple automation tool. Organizations must be able to prove the logic behind AI's decision-making and clearly define who holds accountability when it makes a wrong judgment. Therefore, AI models, training data, prompts, and automation rules must become subjects of strict control under Change Management and Configuration Management. This is also a core requirement demanded by international standards such as ISO/IEC 38507 (AI Governance) and ISO/IEC 42001 (AI Management System).



Specifically, this area consists of the following control and utilization processes:


AIOps and Predictive Analysis: Utilizes machine learning on massive operational data to detect early signs of anomalies, supporting operators to take proactive measures rather than reactive ones.


Intelligent Automation: Optimizes resources by fully automating repetitive, rule-based operational tasks such as account processing, notification dispatch, and standard requests.


AI-Driven Decision Support: Provides powerful, data-driven insights for determining root causes of failures, resource allocation, and investment priorities. However, final approval authority must remain with human operators.


Continuous Improvement and Innovation: Continuously evaluates the effects of AI adoption—such as automation rates, accuracy improvements, and processing time reductions—and discovers new areas for intelligence.




'Continuity and Resilience Management' to Overcome the Worst Crises


Amid unexpected disaster situations like natural disasters, large-scale cloud outages, or ransomware attacks, even a few hours of IT service interruption directly translates into massive losses that threaten corporate survival. Continuity and Resilience Management refers to the capability to seamlessly sustain core services—the heart of the business—even in such critical crises, or to recover them at incredible speed if they do collapse.


In a mature ITSM framework, BCP (Business Continuity Plan), DRP (Disaster Recovery Plan), and backup systems are not scattered in silos. Organizations can only respond to a real crisis when they precisely define the Recovery Time Objective (RTO) and Recovery Point Objective (RPO) of core services through Business Impact Analysis (BIA), and ensure these operate in perfect tandem with Configuration Information (CMDB) and the crisis response organizational structure.



The detailed processes of resilience management that increase organizational survivability are as follows:


Continuity Planning: Analyzes business impacts to set clear recovery goals and establishes robust recovery strategies capable of sustaining core services under any circumstances.


Crisis Response Management: Clearly defines rapid decision-making authority, roles, reporting lines, and internal/external communication procedures in advance so as not to miss the 'golden hour' when a disaster strikes.


Recovery Execution and Testing: Conducts regular simulated drills to ensure plans do not remain just on paper, fiercely verifying that the recovery framework operates perfectly even in actual crisis situations.


Disaster Recovery and Service Restoration: Technically restores systems according to predefined priorities and rapidly activates user notifications and workaround procedures to return to a normal state.




So far, we have examined the seven core frameworks of next-generation ITSM. For this massive value chain to operate flawlessly without fragmentation, it must be supported by the essential backbone of strong Governance and a Continual Service Improvement (CSI) framework.


Organically combining the requirements of ISO/IEC 20000 (IT Service Management) with the principles of ISO/IEC 38507/42001 (AI Governance) secures not only outstanding service quality but also the ethics and reliability of intelligent technology utilization. Furthermore, all performance metrics (KPIs) derived during the operational process—such as availability, MTTR (Mean Time To Recovery), change success rate, and user satisfaction—must not be consumed as mere numbers for reporting. When a virtuous cycle is established where performance is transparently measured, causes are analyzed, and this leads directly to the next stage's improvement plan, the organization's ITSM can self-evolve and mature every single day.



@STEG Luke Kim