JPH0433035A - Error monitoring system - Google Patents

Error monitoring system

Info

Publication number
JPH0433035A
JPH0433035A JP2134418A JP13441890A JPH0433035A JP H0433035 A JPH0433035 A JP H0433035A JP 2134418 A JP2134418 A JP 2134418A JP 13441890 A JP13441890 A JP 13441890A JP H0433035 A JPH0433035 A JP H0433035A
Authority
JP
Japan
Prior art keywords
error
erp
recovery procedure
code
statistical information
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
JP2134418A
Other languages
Japanese (ja)
Inventor
Kaori Takahashi
かおり 高橋
Kunio Yajima
矢島 邦夫
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Fujitsu Ltd
Original Assignee
Fujitsu Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Fujitsu Ltd filed Critical Fujitsu Ltd
Priority to JP2134418A priority Critical patent/JPH0433035A/en
Publication of JPH0433035A publication Critical patent/JPH0433035A/en
Pending legal-status Critical Current

Links

Landscapes

  • Debugging And Monitoring (AREA)

Abstract

PURPOSE:To easily correct a fault point after separating it by providing a means to discriminate a sub system, for which an error generating device belongs, based on a system configuration definition table, error recovery process code and threshold value control table for the unit of the sub system, and a means to control the number of times for generating the faults for the unit of the doubtful point. CONSTITUTION:In the form of tables 1 and 2, the doubtful point weighted corresponding to whether the device is in a high-order or low-order, and a threshold value to instruct the exchange of the doubt full point are provided for each ERP code generated and outputted to a hardware when the error is generated during an input / output operation. On the other hand, the number of times for generating the error for each doubtful point is counted and controlled as a statistical information table 3, and when the error is generated, the doubtful device can be specified based on the outputted ERP code and the address of the error generating device. Further, when it is recognized that the number of times for generating the error in the doubtful device up to the moment exceeds the threshold value, it is judged to exchange the device, and it is automatically instruct to exchange the doubtful device. Thus, the fault point can be easily separated and the fault can be corrected without fail.

Description

【発明の詳細な説明】 〔概要〕 入出力動作時に発生した装置の誤りに対して、その解析
と回復を試みるエラー回復手順(ERP)機構を備えた
計算機システムで発生したエラーを監視する方式に関し
、 人手を介することなく、容易に障害箇所を切り分けるこ
とができ、且つ、計算機システムの故障。
[Detailed Description of the Invention] [Summary] This invention relates to a method for monitoring errors that occur in a computer system equipped with an error recovery procedure (ERP) mechanism that attempts to analyze and recover from device errors that occur during input/output operations. , The failure point can be easily isolated without human intervention, and the computer system can fail.

障害を的確に修復することができるエラー監視方式を提
供することを目的とし、 システム構成定義テーブルと、該計算機システムのハー
ドウェアが作成するエラー回復手順コード(ERPコー
ド)ごとに、被疑箇所の重み付けを行い、重み付けの大
きい方を上位の次元とし、該重み付けの小さい方を下位
の次元として管理すると共に、該重み付けされた被疑箇
所に対応して、エラー監視時間を複数設定し、該設定し
た複数個のエラー監視時間に対応して、閾値制御を行う
為のエラー発生回数を設定した閾値制御テーブルと、上
記重み付けされた被疑箇所の個数分のエラー発生の回数
を、重み付け別に記憶する統計情報テーブルとを設けて
、入出力動作時にエラーが発生したとき、上記エラー回
復手順(ERP)機構が生成。
The aim is to provide an error monitoring method that can accurately repair failures, and weights suspect locations based on the system configuration definition table and the error recovery procedure code (ERP code) created by the computer system hardware. The dimension with the larger weight is treated as the upper dimension, and the dimension with the smaller weight is managed as the lower dimension. In addition, multiple error monitoring times are set corresponding to the weighted suspect locations, and the set multiple error monitoring times are A threshold control table that sets the number of error occurrences to perform threshold control in accordance with the error monitoring time; and a statistical information table that stores the number of error occurrences for the number of suspect points weighted above for each weight. When an error occurs during an input/output operation, the error recovery procedure (ERP) mechanism is generated.

出力したエラー回復手順(ERP)コードと、エラー発
生装置アドレスと、上記システム構成定義テーブルとに
基づいてエラー発生装置を識別し、上記エラー回復手順
(ERP)機構が生成、出力したエラー回復手順(ER
P)コードと、エラー発生装置アドレスとをキーとして
、上記閾値制御テーブルと。
The error recovery procedure (ERP) mechanism generated and outputted by the error recovery procedure (ERP) mechanism identifies the error generation device based on the output error recovery procedure (ERP) code, the error generation device address, and the system configuration definition table. E.R.
P) The above threshold control table using the code and the error generating device address as keys.

統計情報テーブルとを参照して、上記統計情報テーブル
の対応する箇所のエラー回数を更新し、該更新されたエ
ラー回数が、上記閾値制御テーブルに設定されている閾
値を越えている場合には、エラー通知を行うように構成
する。
Update the number of errors in the corresponding part of the statistical information table by referring to the statistical information table, and if the updated number of errors exceeds the threshold set in the threshold control table, Configure for error notification.

〔産業上の利用分野〕[Industrial application field]

本発明は、入出力動作時に発生した装置の誤りに対して
、その解析と回復を試みるエラー回復手順(ERP)機
構を備えた計算機システムで発生したエラーを監視する
方式に関する。
The present invention relates to a method for monitoring errors that occur in a computer system that includes an error recovery procedure (ERP) mechanism that attempts to analyze and recover from errors in devices that occur during input/output operations.

信転性の高い計算機システムのニーズが高まっている近
年では、システムの故障・障害を未然に防ぐこと、即ち
、間欠障害の発生回数がトータルして規定回数以上発生
する場合、回復が可能であっても、致命的な故障に至る
前に保守を行うことが重要になってきている。
In recent years, as the need for highly reliable computer systems has increased, it has become increasingly important to prevent system failures and failures, that is, to ensure that recovery is possible if the total number of intermittent failures occurs exceeds a specified number of times. However, it is becoming increasingly important to perform maintenance before catastrophic failure occurs.

又、発生したエラー事象から、保守員に対して被疑部品
(ユニット、装置等)の通知を行うことで障害の修復時
間を短縮することができる。
Further, by notifying maintenance personnel of the suspected component (unit, device, etc.) based on the error event that has occurred, the time required to repair the fault can be shortened.

このようなことから、入出力動作に関連したエラーの発
生に対して、効果的に、保守員に対して被疑部品の通知
を行うことができるエラー監視方式が必要とされる。
For this reason, there is a need for an error monitoring system that can effectively notify maintenance personnel of suspect parts when errors related to input/output operations occur.

〔従来の技術と発明が解決しようとする課題〕第3図は
従来のエラー監視方式を説明する図である。
[Prior art and problems to be solved by the invention] FIG. 3 is a diagram illustrating a conventional error monitoring system.

従来、O3(オペレーティング・システム)稼働中に、
入出力動作に関連するハードウェアエラーが発生すると
、保守員がテストプログラムを起動してエラーリストを
解析したり、出力されたログ編集リストを目で見て障害
箇所の切り分けを行い部品交換の判断を行っていた。
Traditionally, during O3 (operating system) operation,
When a hardware error related to input/output operations occurs, maintenance personnel can start a test program, analyze the error list, visually isolate the fault by looking at the output log edit list, and decide whether to replace the part. was going on.

このため、エラー発生要因と被疑箇所が1対n。Therefore, there is a ratio of 1:n between the cause of the error and the suspected location.

すなわち一つのエラー発生要因に対して複数の被疑箇所
が考えられる場合、被疑箇所を特定することは困難であ
った。
That is, when a plurality of suspected locations are considered for one error occurrence factor, it is difficult to identify the suspected location.

又、ソフトウェアの回復処理によって訂正が可能であり
、その発生回数を監視する必要のあるエラーの発生頻度
は、短時間で多数発生する場合や一定間隔を置いて発生
する場合など様々であり、特定の一定間隔(例えば、3
0分)に、閾値として定められている閾値を越えるエラ
ー発生回数を検出して、被疑箇所を特定するという、所
謂、しきい値制御をすることが困難であった。
In addition, errors can be corrected by software recovery processing, and the frequency of occurrence of errors that requires monitoring the number of occurrences varies, such as cases in which many errors occur in a short period of time or cases in which they occur at regular intervals. at regular intervals (for example, 3
It has been difficult to perform so-called threshold control, in which the number of error occurrences exceeding a predetermined threshold is detected during the 0 minute period, and the suspected location is identified.

本発明は上記従来の欠点に鑑み、入出力動作時に発生し
た装置の誤りに対して、その解析と回復を試みるエラー
回復手順(ERP)機構を備えた計算機システムで発生
したエラーに対して、人手を介することなく、容易に障
害箇所を切り分けることができ、且つ、計算機システム
の故障、障害を的確に修復することができるエラー監視
方式を提供することを目的とするものである。
In view of the above-mentioned conventional drawbacks, the present invention provides a manual system for errors that occur in a computer system equipped with an error recovery procedure (ERP) mechanism that attempts to analyze and recover from errors in devices that occur during input/output operations. It is an object of the present invention to provide an error monitoring method that can easily isolate the location of a fault without going through the steps, and can accurately repair failures and failures in a computer system.

〔課題を解決するための手段〕[Means to solve the problem]

第1図は本発明の原理説明図であって、(a)は原理構
成図を示し、(b)はエラー監視処理フローの概要を示
している。
FIG. 1 is a diagram explaining the principle of the present invention, in which (a) shows a diagram of the basic configuration, and (b) shows an outline of an error monitoring process flow.

上記の問題点は下記の如くに構成したエラー監視方式に
よって解決される。
The above problems are solved by an error monitoring system configured as follows.

入出力動作時に発生した装置のエラーに対して、その解
析と回復を試みるエラー回復手順(ERP)機構を備え
た計算機システムで発生したエラーを監視する方式であ
って、 システム構成定義テーブルlと。
This is a system for monitoring errors that occur in a computer system that is equipped with an error recovery procedure (ERP) mechanism that attempts to analyze and recover from device errors that occur during input/output operations, the system configuration definition table l.

該計算機システムのハードウェアが作成するエラー回復
手順コード(ERPコード)ごとに、被疑箇所の重み付
けを行い、重み付けの大きい方を上位の次元とし、該重
み付けの小さい方を下位の次元として管理すると共に、
該重み付けされた被疑箇所に対応して、エラー監視時間
を複数設定し、該設定した複数個のエラー監視時間に対
応して、閾値制御を行う為のエラー発生回数を設定した
閾値制御テーブル2と。
For each error recovery procedure code (ERP code) created by the hardware of the computer system, the suspected locations are weighted, the one with the larger weight is treated as the upper dimension, and the one with the smaller weight is managed as the lower dimension. ,
A threshold control table 2 in which a plurality of error monitoring times are set corresponding to the weighted suspect locations, and a number of error occurrences for performing threshold control is set in accordance with the plurality of set error monitoring times. .

上記重み付けされた被疑箇所の個数分のエラー発生の回
数を、重み付け別に記憶する統計情報テーブル3,3a
+3bとを設けて、 入出力動作時にエラーが発生したとき、上記エラー回復
手順(ERP)機構が生成、出力したエラー回復手順(
ERP)コードと、エラー発生装置アドレスと、上記シ
ステム構成定義テーブル1とに基づいてエラー発生装置
を識別し、 上記エラー回復手順(ERP)機構が生成、出力したエ
ラー回復手順(ERP)コードと、エラー発生装置アド
レスとをキーとして、上記HM制御テーブル2と、統計
情報テーブル3.3a、3bとを参照して、上記統計情
報テーブル3,3a、3bの対応する箇所のエラー回数
を更新し、 該更新されたエラー回数が、上記閾値制御テーブル2に
設定されている閾値を越えている場合には、エラー通知
を行うように構成する。
Statistical information tables 3, 3a that store the number of error occurrences corresponding to the number of weighted suspect locations, according to weighting.
+3b is provided, so that when an error occurs during input/output operation, the error recovery procedure (ERP) generated and output by the above error recovery procedure (ERP) mechanism is
An error recovery procedure (ERP) code generated and output by the error recovery procedure (ERP) mechanism that identifies the error generation device based on the error generation device address, the error generation device address, and the system configuration definition table 1; Using the error generating device address as a key, refer to the HM control table 2 and the statistical information tables 3.3a and 3b, and update the number of errors in the corresponding parts of the statistical information tables 3, 3a and 3b, If the updated number of errors exceeds the threshold set in the threshold control table 2, the system is configured to issue an error notification.

〔作用] 即ち、本発明によれば、第1図(a)に示したように、
入出力動作時に発生した装置の誤りに対して、その解析
と回復を試みるエラー回復手順(ERP) m構を備え
た計算機システムで発生したエラーを監視するのに、シ
ステム構成定義テーブルと。
[Operation] That is, according to the present invention, as shown in FIG. 1(a),
An error recovery procedure (ERP) that attempts to analyze and recover from a device error that occurs during input/output operations.A system configuration definition table is used to monitor errors that occur in a computer system equipped with an error recovery procedure (ERP).

エラー回復手順(ERP)コード(以下、ERPコード
という)と9例えば、上位装置か/下位装置かに対応し
て重み付けが施された被疑箇所と複数監視時間単位の閾
値が定義されているサブシステム単位の閾値制御テーブ
ルを基に、エラー発生装置がどのサブシステムに属する
かを判別する手段と。
Error recovery procedure (ERP) code (hereinafter referred to as ERP code) and 9. For example, a subsystem in which suspect points are weighted depending on whether they are upper or lower devices and thresholds for multiple monitoring time units are defined. means for determining to which subsystem the error generating device belongs based on a unit threshold value control table;

上記重み(次元)の個数分の統計情報テーブルを作成し
、各被疑箇所単位に障害発生回数等を管理する手段と、
該指定監視時間内のエラー発生頻度をチエツクする手段
とを設けて、以下のようにエラー監視を行う。
Means for creating statistical information tables for the number of weights (dimensions) mentioned above and managing the number of failure occurrences, etc. for each suspect location;
A means for checking the frequency of error occurrence within the specified monitoring time is provided, and error monitoring is performed as follows.

即ち、(b)図に示したように、先ず、O3稼働中にハ
ードウェア障害が発生すると、O3はエラーログ情報を
組み立て、これをエラーロギングファイルに記録する。
That is, as shown in Figure (b), first, when a hardware failure occurs during O3 operation, O3 assembles error log information and records it in an error logging file.

この事象を契機に制御部はERPコードとエラー発生装
置アドレスを引数としてログ解析部を呼び出す。
Taking this event as a trigger, the control unit calls the log analysis unit using the ERP code and the error generating device address as arguments.

ログ解析部は、システム構成定義テーブルを基に、この
装置がどのサブシステムに属するかを判別し、統計ファ
イルから統計情報(多重次元)をメモリ上に読み込む。
The log analysis unit determines which subsystem this device belongs to based on the system configuration definition table, and reads statistical information (multidimensional) from the statistical file onto the memory.

そして、制御部から引き渡されたERPコードとエラー
発生装置アドレスをキーとして、当該サブシステムの閾
値制御テーブルと、統計情報テーブル(多重次元)のデ
ータを基に時間監視制御および被疑箇所の重み付けを特
徴とした閾値制御を行い、統計情報テーブルを更新し、
これを統計ファイルに書き込む。
Using the ERP code and error generating device address handed over from the control unit as keys, time monitoring control and weighting of suspect locations are performed based on the data of the threshold control table and statistical information table (multidimensional) of the relevant subsystem. Perform threshold control and update the statistical information table.
Write this to the statistics file.

装置で発生したERPコード別のエラー発生回数が、上
記閾値制御テーブルに定義されている閾値を超えた場合
はメツセージを出力し、オペレータに被疑箇所(装置)
の交換を依願する。又、該閾値制御を必要としない「無
条件交換」のエラーに対しても、同様に、被疑装置のメ
ツセージを出力し、該閾値制御テーブルが、複数個の被
疑箇所を指示している場合には、ハードウェアの生成す
る装置アドレスに応じて、次に重みの大きい(即ち。
If the number of errors for each ERP code that occurs in the device exceeds the threshold defined in the threshold control table above, a message will be output and the operator will be notified of the suspected location (device).
Request a replacement. Also, for "unconditional replacement" errors that do not require threshold control, a message from the suspect device is output in the same way, and if the threshold control table indicates multiple suspect locations, has the next highest weight (i.e., ) depending on the device address generated by the hardware.

二次の)被疑箇所を求めるように動作する。It operates to find the suspected location (secondary).

即ち、本発明においては、入出力動作中にエラーが発生
したとき、ハードウェアが生成して出力するERPコー
ド毎に、該装置が上位にあるか、下位にあるかに応じて
重み付けされた被疑箇所と、該被疑箇所を交換した方が
よいことを指示する閾値(監視時間対応)とがテーブル
の形で用意されており、且つ、該被疑箇所毎のエラー発
生回数(上記監視時間対応)を計数して、統計情報テー
ブルとして管理されているので、エラーが発生したとき
、該出力されたERPコードと、エラー発生装置アドレ
スを基に、被疑装置を特定でき、且つ、該被疑装置の今
迄のエラー発生回数が閾値を越えていると認識されたと
き、該被疑装置の交換が適当と判断して、自動的に、該
被疑装置の交換を保守者に指示することができ、システ
ムの故障・障害を、予防保全の形で的確に修復すること
ができるという効果がある。
That is, in the present invention, when an error occurs during an input/output operation, for each ERP code generated and output by the hardware, a suspect weight is assigned depending on whether the device is located at a higher level or lower level. The location and the threshold value (corresponding to the monitoring time) indicating that it is better to replace the suspect location are prepared in the form of a table, and the number of error occurrences for each suspect location (corresponding to the above monitoring time) is prepared. It is counted and managed as a statistical information table, so when an error occurs, the suspect device can be identified based on the output ERP code and the address of the device where the error occurred, and the suspect device's past history can be identified. When it is recognized that the number of error occurrences exceeds a threshold, it is determined that it is appropriate to replace the suspect device and automatically instructs maintenance personnel to replace the suspect device, thereby preventing system failure.・It has the effect of being able to accurately repair failures in the form of preventive maintenance.

〔実施例〕〔Example〕

以下本発明の実施例を図面によって詳述する。 Embodiments of the present invention will be described in detail below with reference to the drawings.

前述の第1図は本発明の詳細な説明する図であり、第2
図は本発明の一実施例を示した図であって、(a)は閾
値制御テーブルの構成例を示し、(bl)〜(b3)は
統計情報テーブルの構成例を示し、(C1) 、 (C
2)は磁気テープ(MT)サブシステムにおける閾値制
御の詳細処理フローを示している。
The above-mentioned FIG. 1 is a diagram for explaining the present invention in detail, and FIG.
The figure shows an example of the present invention, in which (a) shows an example of the configuration of a threshold control table, (bl) to (b3) show examples of the configuration of a statistical information table, (C1), (C
2) shows a detailed processing flow of threshold control in the magnetic tape (MT) subsystem.

本発明においては、システム構成定義テーブル1と、 
ERPコード毎に閾値条件と、上位装置、下位装置に応
じて重み付けされた被疑箇所を指示する閾値制御テーブ
ル2と、 ERPコード毎に、且つ、上記閾値条件毎に
、被疑箇所に発生したエラーの回数を記録する統計情報
テーブル3,3a、3bとを設けて、入出力動作中にエ
ラーが発生したとき、上記システム構成定義テーブル1
を参照して、被疑箇所(サブシステム)を特定し、更に
、閾値制御テーブル2と、統計情報テーブル3,3a、
3bを参照し、統計情報テーブル3.3a、3b上に記
録されているエラー発生の回数が、上記閾値制御テーブ
ル2が指示している閾値を越えている被疑箇所、或いは
、無条件交換の被疑箇所に対して交換を指示する手段が
本発明を実施するのに必要な手段である。尚、全図を通
して同じ符号は同じ対象物を示している。
In the present invention, a system configuration definition table 1,
Threshold control table 2 that specifies threshold conditions for each ERP code and suspect locations weighted according to upper and lower devices; Statistical information tables 3, 3a, and 3b are provided to record the number of times, and when an error occurs during input/output operation, the above system configuration definition table 1 is provided.
With reference to
3b, the number of error occurrences recorded in the statistical information tables 3.3a and 3b exceeds the threshold specified by the threshold control table 2, or the suspected unconditional exchange. A means for instructing replacement of a part is a necessary means for carrying out the present invention. Note that the same reference numerals indicate the same objects throughout the figures.

以下、第1図を参照しながら、第2図によって、本発明
のエラー監視方式を説明する。
Hereinafter, the error monitoring system of the present invention will be explained with reference to FIG. 2 while referring to FIG.

本実施例においては、磁気テープ(MT)サブシステム
を例にしているが、これに限定されるものでないことは
いう迄もないことである。
In this embodiment, a magnetic tape (MT) subsystem is taken as an example, but it goes without saying that the present invention is not limited to this.

先ず、(a)図に示した閾値制御テーブル2は、ERP
コード毎に、閾値条件と1図示されているように、上位
装置(MTU) =>下位装置(TAPE)に対応して
重み付けされた被疑箇所がテーブルの形で示されている
First, the threshold control table 2 shown in FIG.
For each code, a threshold condition and, as shown in the figure, suspect locations weighted in correspondence with upper-level equipment (MTU) => lower-level equipment (TAPE) are shown in the form of a table.

該閾値条件としては、無条件交換の場合と、監視期間を
定めて、例えば、30分間隔、或いは、1ケ月間隔で計
数したエラー回数の閾値を定義し、二の閾値を越えるエ
ラーがあると、該被疑箇所は、交換した方がよいとする
ものである。
The threshold conditions include the case of unconditional exchange and the monitoring period, for example, defining a threshold of the number of errors counted at 30 minute intervals or monthly intervals, and if there is an error exceeding the second threshold. , it is recommended that the suspected part be replaced.

(bl)〜(b3)図に示した統計情報テーブル3,3
a。
(bl) to (b3) Statistical information tables 3, 3 shown in the figures
a.

3bは、上記閾値制御テーブル2で重み付けされた被疑
箇所に対応して、後述する閾値制御で、現在のエラーを
加算するように構成されている。
3b is configured to add the current error in accordance with the suspect location weighted in the threshold control table 2 using threshold control, which will be described later.

先ず、第1図(b)の概略動作フローに示されているよ
うに、O3稼働中にハードウェア障害が発生すると、O
8はエラーログ情報を組み立て、これをエラーロギング
ファイルに記録する。この事象を契機に制御部はERP
コードとエラー発生装置アドレスを引数としてログ解析
部を呼び出す。
First, as shown in the schematic operation flow in Figure 1(b), when a hardware failure occurs during O3 operation, the O3
8 assembles error log information and records it in an error logging file. Taking this event as an opportunity, the control unit started the ERP
Call the log analysis section with the code and error generating device address as arguments.

ログ解析部は、システム構成定義テーブル1を基に、こ
の装置がどのサブシステムに属するかを判別し、統計フ
ァイルから統計情報(多重次元)テーブル3.3a、3
bの内容を、図示されていないメモリ上に読み込む。
The log analysis unit determines which subsystem this device belongs to based on the system configuration definition table 1, and extracts statistical information (multidimensional) tables 3.3a, 3 from the statistical file.
The contents of b are read into a memory (not shown).

そして、制御部から引き渡されたERPコードと、エラ
ー発生装置アドレスをキーとして、当該サブシステム(
本実施例では、MTサブシステム)の閾値制御、テーブ
ル2と、統計情報テーブル(多重次元) 3.3a、3
bのデータを基に、時間監視制御。
Then, the subsystem (
In this embodiment, threshold control of the MT subsystem), table 2, and statistical information table (multidimensional) 3.3a, 3
Time monitoring control based on the data of b.

および、被疑箇所の重み付けを特徴とした閾値制御を行
い、該統計情報テーブル3.3a、 3bのエラー回数
を更新し、これを統計ファイルに書き込む。
Then, threshold control is performed, characterized by weighting of suspected locations, the number of errors in the statistical information tables 3.3a and 3b is updated, and this is written in the statistical file.

該統計情報テーブル3,3a、3bに記録されているエ
ラー発生回数が、上記閾値制御テーブル2に定められて
いる閾値を超えた場合はメツセージを出力し、オペレー
タに装置の交換を依願する。
If the number of errors recorded in the statistical information tables 3, 3a, and 3b exceeds the threshold set in the threshold control table 2, a message is output to request the operator to replace the device.

以下、第2図(cl) 、 (C2) 、 (C3)に
示した動作フローにより上記閾値制御の詳細動作を説明
する。
The detailed operation of the threshold value control will be explained below using the operation flows shown in FIGS. 2(cl), (C2), and (C3).

制御部から出力されたERPコード、エラー発生装置ア
ドレスをキーとして、先ず、−次元の統計情報テーブル
3を参照したとき、該統計情報テーブル3に、該当のE
RPコードと9重み付けが施された後のアドレスが一致
する項目があるかどうかが調べられ、なければ、該当項
目を新設するが、あれば、該当項目について、閾値制御
テーブル2を参照し、閾値制御の為のパラメータ (閾
値)■の有無を見て、無ければ、即ち、「無条件交換」
が指示されている場合には、該閾値制御テーブル2が指
示している被疑箇所■を保守者(オペレータ)に通知す
る。(第2図(cl)のステップ1O111,12,2
0参照) 若し、上記閾値パラメータ■が指示されている場合には
、該エラーの発生した時刻について、監視開始時刻(前
に設定された監視開始時刻に、チエツク範囲時間(例え
ば、30分とか、1月等)を、定期的に加算した時刻)
■に、チエツク範囲時間を足した時刻を経過しているか
どうかが調べられる。
When first referring to the −-dimensional statistical information table 3 using the ERP code and error generating device address output from the control unit as keys, the corresponding E
It is checked whether there is an item that matches the RP code and the address after 9 weighting. If not, a new corresponding item is created, but if there is, the threshold control table 2 is referred to for the corresponding item and the threshold value Check the presence or absence of the parameter (threshold value) for control, and if there is none, that is, "unconditional exchange"
If specified, the maintenance person (operator) is notified of the suspected location (3) specified by the threshold control table 2. (Steps 1O111, 12, 2 in Figure 2 (cl)
0) If the above threshold parameter ■ is specified, the monitoring start time (previously set monitoring start time) is set to the check range time (for example, 30 minutes) for the time when the error occurred. , January, etc.).
It is checked whether the time obtained by adding the check range time to (2) has passed.

ここで、該エラー発生時刻がチエツク範囲時間を足した
時刻を越えていなければ、該エラーは定期的なエラーと
認識され、該統計情報テーブル3の現在のエラー回数に
+1″されるが、該チエツク範囲時間を足した時刻を越
えていると、上記のエラーは一時的なエラーとして、そ
れまでに計数されていた、該当チエツク範囲時間に対応
したエラー回数はクリアされ、且つ、その時刻を上記監
視開始時刻■に設定して、その時刻を監視開始時刻■と
して、新たに、定期的なエラーの監視を行うようにする
。(第2図(C1)のステップ13.14゜15参照) このようにして、該統計情報テーブル3の更新されたエ
ラー回数を、上記閾値制御テーブル2に指示されている
閾値と比較し、該閾値を越えている場合には、該被疑箇
所は、エラーが定期的に起こっており、いずれダウンす
る可能性がある箇所と判断され、保守者(オペレータ)
に、該被疑箇所を交換するように通知する。(第2図(
cl)のステップ16.21参照) 上記−次元の統計情報テーブル3に、二次元テーブル3
aがあることが指示されている場合には、上記と同じ手
順によって、該二次元テーブル3aに対して、上記−次
元テーブル3と同じ処理を実行する。
Here, if the error occurrence time does not exceed the sum of the check range time, the error is recognized as a regular error, and the current number of errors in the statistical information table 3 is added by 1", but If the time exceeds the sum of the check range time, the above error is treated as a temporary error, and the error count corresponding to the check range time that has been counted up to that point is cleared, and the time is Set the monitoring start time ■, and use that time as the monitoring start time ■ to perform new periodic error monitoring. (Refer to steps 13, 14, and 15 in Figure 2 (C1).) The updated number of errors in the statistical information table 3 is compared with the threshold specified in the threshold control table 2, and if it exceeds the threshold, the suspected location is determined to have regular errors. It has been determined that this is a location that may go down someday, and the maintenance personnel (operator)
and notify them to replace the suspect part. (Figure 2 (
cl) step 16.21) Add two-dimensional table 3 to the statistical information table 3 of the - dimension above.
If it is specified that there is a, the same process as for the above-mentioned -dimensional table 3 is performed on the two-dimensional table 3a using the same procedure as above.

同様にして、該二次元テーブル3aに、三次元テーブル
3bがあることが指示されている場合には、上記と同じ
手順によって、該三次元テーブル3bに対して、上記−
次元テーブル3と同じ処理を実行する。(第2図(C2
) 、 (C3)参照)このように、本発明は、システ
ム構成定義テーブル1と、 ERPコード毎に閾値条件
と、上位装置→下位装置に対応して重み付けされた被疑
箇所を指示する閾値制御テーブル2と、 ERPコード
毎に、且つ、上記閾値条件毎に、被疑箇所に発生したエ
ラーの回数を記録する多重次元の統計情報テーブル3.
3a、3bとを設けて、入出力動作中にエラーが発生し
たとき、上記システム構成定義テーブル1を参照して、
被疑箇所を特定し、更に、閾値制御テーブル2と、統計
情報テーブル3.3a、3bを参照し、統計情報テーブ
ル3,3a、3b上に記録されているエラー回数を更新
し、該更新後のエラー発生の回数が、上記閾値制御テー
ブル2が指示している閾値を越えている被疑箇所等に対
して交換をオペレータに指示するようにした所に特徴が
ある。
Similarly, if it is specified that the two-dimensional table 3a has a three-dimensional table 3b, the above-mentioned -
Execute the same process as for dimension table 3. (Figure 2 (C2
), (C3)) As described above, the present invention includes the system configuration definition table 1, the threshold conditions for each ERP code, and the threshold control table that indicates suspected locations that are weighted according to the order of upper device → lower device. 2. A multi-dimensional statistical information table that records the number of errors occurring at the suspect location for each ERP code and for each of the above threshold conditions.3.
3a and 3b, and when an error occurs during input/output operation, refer to the system configuration definition table 1 above,
Identify the suspected location, refer to the threshold control table 2 and the statistical information tables 3.3a and 3b, update the number of errors recorded on the statistical information tables 3, 3a and 3b, and A feature of this system is that the operator is instructed to replace suspected locations where the number of error occurrences exceeds the threshold specified by the threshold control table 2.

〔発明の効果〕〔Effect of the invention〕

以上、詳細に説明したように、本発明のエラー監視方式
は、入出力動作時に発生した装置の誤りに対して、その
解析と回復を試みるエラー回復手順(ERP)機構を備
えた計算機システムにおいて、システム構成定義テーブ
ルと、該計算機システムのハードウェアが作成するエラ
ー回復手順コード(1!RPコード)ごとに、被疑箇所
の重み付けを行い、重み付けの大きい方を上位の次元と
し、該重み付けの小さい方を下位の次元として管理する
すると共に、該重み付けされた被疑箇所に対応して、エ
ラー監視時間を複数設定し、該設定した複数個のエラー
監視時間に対応して、閾値制御を行う為のエラー発生回
数を設定した閾値@御テーブルと、上記重み付けされた
被疑箇所の個数分のエラー発生の回数を、重み付け別に
記憶する統計情報テーブルとを設けて、入出力動作時に
エラー発生したとき、上記エラー回復手順(ERP)機
構が生成、出力したエラー回復手順(ERP)コードと
、エラー発生装置アドレスと、上記システム構成定義テ
ーブルとに基づいてエラー発生装置を識別し、上記エラ
ー回復手順(ERP)機構が生成、出力したエラー回復
手順(ERP)コードと、エラー発生装置アドレスとを
キーとして、上記閾値制御テーブルと、多重構成の統計
情報テーブルとを参照して、上記統計情報テーブルの対
応する箇所のエラー回数を更新し、該更新されたエラー
回数が、上記閾値制御テーブルに設定されている閾値を
越えている場合には、エラー通知を行うようにしたもの
であるので、一つのエラー発生要因に対して複数の被疑
箇所が考えられる場合でも、人手を介入することなく容
易に障害箇所の切り分けができ、又、システムの故障・
障害を、予防保全の形で的確に修復できるという効果が
ある。
As described above in detail, the error monitoring method of the present invention is applicable to a computer system equipped with an error recovery procedure (ERP) mechanism that attempts to analyze and recover from errors in devices that occur during input/output operations. Suspect points are weighted for each error recovery procedure code (1!RP code) created by the system configuration definition table and the hardware of the computer system, and the one with the larger weight is considered the upper dimension, and the one with the smaller weight is managed as a lower dimension, multiple error monitoring times are set corresponding to the weighted suspect locations, and error threshold control is performed in response to the set multiple error monitoring times. A threshold@control table that sets the number of occurrences and a statistical information table that stores the number of error occurrences for the number of suspected points weighted above according to weighting are provided, so that when an error occurs during input/output operation, the above error The error recovery procedure (ERP) mechanism identifies the error generation device based on the error recovery procedure (ERP) code generated and output by the error recovery procedure (ERP) mechanism, the error generation device address, and the system configuration definition table, and executes the error recovery procedure (ERP) mechanism. Using the error recovery procedure (ERP) code generated and output by The number of errors is updated, and if the updated number of errors exceeds the threshold set in the threshold control table, an error notification is issued. In contrast, even when multiple suspect locations are considered, the fault location can be easily isolated without human intervention, and system failures and
This has the effect of accurately repairing failures in the form of preventive maintenance.

【図面の簡単な説明】[Brief explanation of drawings]

第1図は本発明の原理説明図。 第2図は本発明の一実施例を示した図。 第3図は従来のエラー監視方式を説明する図。 である。 図面において、 1はシステム構成定義テーブル。 2は閾値制御テーブル。 3.3a、3bは統計情報テーブル。 10〜17.20.21は処理ステップ。 をそれぞれ示す。 第1圓(その2) (b2) (b3) 本発明の=一実施例を示した図 第 図 (その2) (bl) 本発明の一実施例を示した図 第 図 (そのl) 第 図 (その3) 第 図 (その4) 第 図 (その5) FIG. 1 is a diagram explaining the principle of the present invention. FIG. 2 is a diagram showing an embodiment of the present invention. FIG. 3 is a diagram explaining a conventional error monitoring method. It is. In the drawing, 1 is the system configuration definition table. 2 is a threshold control table. 3.3a and 3b are statistical information tables. 10-17.20.21 are processing steps. are shown respectively. The first circle (part 2) (b2) (b3) A diagram showing an embodiment of the present invention No. figure (Part 2) (bl) A diagram showing an embodiment of the present invention No. figure (Part 1) No. figure (Part 3) No. figure (Part 4) No. figure (Part 5)

Claims (1)

【特許請求の範囲】 入出力動作時に発生した装置のエラーに対して、その解
析と回復を試みるエラー回復手順(ERP)機構を備え
た計算機システムで発生したエラーを監視する方式であ
って、 システム構成定義テーブル(1)と、 該計算機システムのハードウェアが作成するエラー回復
手順コード(ERPコード)ごとに、被疑箇所の重み付
けを行い、重み付けの大きい方を上位の次元とし、該重
み付けの小さい方を下位の次元として管理すると共に、
該重み付けされた被疑箇所に対応して、エラー監視時間
を複数設定し、該設定した複数個のエラー監視時間に対
応して、閾値制御を行う為のエラー発生回数を設定した
閾値制御テーブル(2)と、 上記重み付けされた被疑箇所の個数分のエラー発生の回
数を、重み付け別に記憶する統計情報テーブル(3、3
a、3b)とを設けて、 入出力動作時にエラーが発生したとき、上記エラー回復
手順(ERP)機構が生成、出力したエラー回復手順(
ERP)コードと、エラー発生装置アドレスと、上記シ
ステム構成定義テーブル(1)とに基づいてエラー発生
装置を識別し、 上記エラー回復手順(ERP)機構が生成、出力したエ
ラー回復手順(ERP)コードと、エラー発生装置アド
レスとをキーとして、上記閾値制御テーブル(2)と、
統計情報テーブル(3、3a、3b)とを参照して、上
記統計情報テーブル(3、3a、3b)の対応する箇所
のエラー回数を更新し、 該更新されたエラー回数が、上記閾値制御テーブル(2
)に設定されている閾値を越えている場合には、エラー
通知を行うことを特徴するエラー監視方式。
[Claims] A method for monitoring errors occurring in a computer system equipped with an error recovery procedure (ERP) mechanism that attempts to analyze and recover from errors in devices occurring during input/output operations, the system comprising: For each configuration definition table (1) and the error recovery procedure code (ERP code) created by the hardware of the computer system, weight the suspect locations, and consider the one with the higher weight as the upper dimension, and the one with the lower weight. In addition to managing it as a lower dimension,
A threshold control table (2) is created in which a plurality of error monitoring times are set corresponding to the weighted suspect locations, and the number of error occurrences for performing threshold control is set corresponding to the plurality of set error monitoring times. ), and a statistical information table (3, 3
a, 3b), and when an error occurs during input/output operation, the error recovery procedure (ERP) generated and output by the above error recovery procedure (ERP) mechanism is
The error recovery procedure (ERP) code generated and output by the error recovery procedure (ERP) mechanism identifies the error generation device based on the error generation device address, the error generation device address, and the system configuration definition table (1) above. and the error generating device address as keys, the threshold control table (2),
Update the number of errors in the corresponding part of the statistical information table (3, 3a, 3b) by referring to the statistical information table (3, 3a, 3b), and the updated error number is applied to the threshold control table. (2
), an error monitoring method is characterized in that an error notification is issued if the threshold value set in the error monitoring method is exceeded.
JP2134418A 1990-05-24 1990-05-24 Error monitoring system Pending JPH0433035A (en)

Priority Applications (1)

Application Number Priority Date Filing Date Title
JP2134418A JPH0433035A (en) 1990-05-24 1990-05-24 Error monitoring system

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
JP2134418A JPH0433035A (en) 1990-05-24 1990-05-24 Error monitoring system

Publications (1)

Publication Number Publication Date
JPH0433035A true JPH0433035A (en) 1992-02-04

Family

ID=15127926

Family Applications (1)

Application Number Title Priority Date Filing Date
JP2134418A Pending JPH0433035A (en) 1990-05-24 1990-05-24 Error monitoring system

Country Status (1)

Country Link
JP (1) JPH0433035A (en)

Similar Documents

Publication Publication Date Title
KR101856543B1 (en) Failure prediction system based on artificial intelligence
JPH02105947A (en) Computer surrounding subsystem and exception event automatic detecting analyzing method
CN101201786B (en) A fault log monitoring method and device
US6598179B1 (en) Table-based error log analysis
US7401263B2 (en) System and method for early detection of system component failure
KR100579956B1 (en) Change monitoring system of computer system
JP4318643B2 (en) Operation management method, operation management apparatus, and operation management program
Murphy et al. Measuring system and software reliability using an automated data collection process
CN102110485B (en) Automated periodic surveillance testing method and apparatus in digital reactor protection system
WO1992014206A1 (en) Knowledge based machine initiated maintenance system
CN109062723A (en) The treating method and apparatus of server failure
CN115794588A (en) Memory fault prediction method, device and system and monitoring server
CN115098306A (en) Embedded fault-tolerant self-healing structure, method and system applied to power industrial control terminal
CN110659147B (en) Self-repairing method and system based on module self-checking behavior
AU674231B2 (en) Fault-tolerant computer systems
CN121166472A (en) Intelligent operation and maintenance method and system for server
CN120822221A (en) Secure boot time optimization method, device, terminal and storage medium
JP2008198123A (en) Fault detection system and fault detection program
CN119065228A (en) A safe and redundant PLC communication control system
CN118735251A (en) An automated inspection planning method based on abnormal feedback data
JP7534700B2 (en) Apparatus for generating correct data, method for generating correct data, and program for generating correct data
JPH04257035A (en) Fault information processing system under virtual computer system
CN114492068A (en) Fault processing test method, system, device, medium, and program product
JP2008181432A (en) Health check device, health check method and program
JPH04127247A (en) Preventive maintenance support system