From "Monitoring" to "Correction", Taking the First Step of ITSM

· Insights

The value stream of “from monitoring to correction” discovers information system anomalies through monitoring, records them as “incident tickets” and defines the “causes” of incidents, then eliminates the “causes” through remediation and changes. Relying on the technical mechanism of unified monitoring of information systems, it connects incident management, problem management, and remediation management processes, forming a closed loop between monitoring and incidents, and triggering a closed loop of defining and eliminating “causes.”

Purpose: Relying on the technical mechanism of unified monitoring of information systems, it connects incident management, problem management, and remediation management processes, forming a closed loop between monitoring and incidents, and triggering “problem management” and “remediation management” to form a closed loop of defining and eliminating “causes.”

Scope: The value stream of “from monitoring to correction” includes: system monitoring, event management, incident management, problem management, change control, configuration management, and remediation management, connected to form a cross-process end-to-end perspective.

Performance monitoring is the continuous monitoring of performance metrics of various software and hardware in information systems, including: networks, storage, server hardware, virtualization platforms, virtual hosts, database systems, middleware systems, and web services. Unified monitoring of information systems has the following values:

  1. Detect hidden risks or locate incidents as early as possible;
  2. Quantify the usage of system resources;
  3. Reduce patrol content and ease manual workload.

“Monitoring” is the process of continuously collecting performance status data of each IT resource through software. Through “monitoring,” the performance of each IT device is quantified, and the allocation and usage of all resources are quantified, providing a basis for the IT department to increase or decrease resource allocation.

Performance Monitoring: Collection and Alerting of Performance Metrics

Performance monitoring of information systems is not just about performance data collection and alerting; more importantly, it is about computing, utilizing, and presenting data. For example, forming early warning mechanisms through feature analysis, building visual response plans through logical relationship models, and analyzing potential hidden risks through capacity analysis using historical performance data.

Monitoring technical means: To avoid additional installation and maintenance costs and overhead, software systems should primarily monitor through standard protocols. Data center managers need to count resource allocation and virtual machine growth in units of “clouds” and “resource pools.” Control resource allocation and utilization rates to quantify physical resource construction needs in advance.

Virtualized Compute Resource Management

  1. Operate enterprise private cloud resources: automatically inventory resources and usage of each resource pool, as well as usage trends.
  2. Discover potential risks of insufficient resources: automatically inventory over-allocation of each resource pool.

Performance Monitoring: Capacity Analysis

While “performance monitoring” may bring some convenience in usage and management, the most important thing is to conduct correlation analysis of monitoring data from various software and hardware. Centering on “business systems” as the “subject,” perform correlation analysis of performance data of associated software and hardware.

An “IT service” is an organism where multiple application systems and multiple layers of software and hardware interact and support each other. Isolated attention to the performance of a single software or hardware is insufficient to describe the overall performance of an “IT service.”

Only by correlating performance metrics at all layers around the same “subject” (IT service) and observing them comprehensively under a unified time dimension can we discover, among these correlated metrics, which ones share common performance fluctuations, and which of those are running under high load and may be bottlenecks. Whether these commonly fluctuating metrics have changed — high load and changes are potential hidden risks.

Changes in business access and system processing will be reflected at all layers supporting the IT service system. Therefore, based on the dependency relationship model in the “Configuration Management Database (CMDB),” around a “business system,” correlate the data of metrics at different layers supporting the same business system to form a cross-device performance metric correlation analysis.

System capacity changes with business changes, and also changes due to software updates or hardware replacements. Therefore, regular analysis of business system capacity trends and changes is needed.

Business changes often do not manifest in the short term, and the changes reflected in data may not be obvious. However, capacity trend changes can be seen in long-term (quarterly) data statistics. Therefore, capacity analysis requires three different cycle perspectives: short-term, medium-term, and long-term, to observe whether the system\'s capacity trend shows early signs of gradually trending toward collapse in long-term development. For example, short-term in days and weeks, medium-term in months, and long-term as trend statistics over three months or more.

Want to see how NI-System v7.0 solves your IT management challenges?

Book a demo, we tailor communication based on your industry and scenario