JPH10320236A - Maintenance deterrent device for preventive maintenance - Google Patents

Maintenance deterrent device for preventive maintenance

Info

Publication number
JPH10320236A
JPH10320236A JP9131211A JP13121197A JPH10320236A JP H10320236 A JPH10320236 A JP H10320236A JP 9131211 A JP9131211 A JP 9131211A JP 13121197 A JP13121197 A JP 13121197A JP H10320236 A JPH10320236 A JP H10320236A
Authority
JP
Japan
Prior art keywords
maintenance
disk
control device
disk controller
application
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Granted
Application number
JP9131211A
Other languages
Japanese (ja)
Other versions
JP3961617B2 (en
JPH10320236A5 (en
Inventor
Tomoyuki Kato
智之 加藤
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Hitachi Ltd
Original Assignee
Hitachi Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Hitachi Ltd filed Critical Hitachi Ltd
Priority to JP13121197A priority Critical patent/JP3961617B2/en
Publication of JPH10320236A publication Critical patent/JPH10320236A/en
Publication of JPH10320236A5 publication Critical patent/JPH10320236A5/ja
Application granted granted Critical
Publication of JP3961617B2 publication Critical patent/JP3961617B2/en
Anticipated expiration legal-status Critical
Expired - Fee Related legal-status Critical Current

Links

Landscapes

  • Debugging And Monitoring (AREA)

Abstract

(57)【要約】 【課題】 保守中に、ディスク制御装置の状態情報と保
守用アプリケーションの操作情報をチェックすることに
より、早い段階で障害発生を検知し、保守を抑止するこ
と。 【解決手段】 二重化されたプロセッサとメモリを有す
るディスク制御装置202と、ディスク制御装置の保守
を実施するためにディスク制御装置に具備された保守用
コンピュータ203と、を備え、保守用コンピュータ
は、それに搭載された保守用アプリケーション107の
作動に関する操作情報105と、ディスク制御装置20
2の状態情報101と、を監視102し、監視の結果、
ディスク制御装置に障害が発生したことが検知されたと
き、保守用アプリケーションの操作情報の来歴105に
基づいて保守直前の状態にディスク制御装置を復旧させ
るリカバリ処理108を実施すること。
(57) [Summary] [PROBLEMS] To detect the occurrence of a failure at an early stage and suppress maintenance by checking status information of a disk control device and operation information of a maintenance application during maintenance. SOLUTION: A disk controller 202 having a duplicated processor and a memory, and a maintenance computer 203 provided in the disk controller for performing maintenance of the disk controller, are provided. Operation information 105 relating to the operation of the installed maintenance application 107 and the disk controller 20
2 and the status information 101, and monitor 102, and as a result of the monitoring,
When it is detected that a failure has occurred in the disk controller, a recovery process 108 for recovering the disk controller to a state immediately before maintenance is performed based on the history 105 of the operation information of the maintenance application.

Description

【発明の詳細な説明】DETAILED DESCRIPTION OF THE INVENTION

【0001】[0001]

【発明の属する技術分野】本発明は、プログラムで動作
する複数のプロセッサ及びメモリからなる制御装置の保
守操作に関する。
BACKGROUND OF THE INVENTION 1. Field of the Invention The present invention relates to a maintenance operation of a control device including a plurality of processors and memories operated by a program.

【0002】[0002]

【従来の技術】プログラムで動作する複数のプロセッサ
及びメモリからなる制御装置の保守を、装置に具備する
保守専用機(以下「サービスプロセッサSVP」と略称
する。)で行う際、サービスプロセッサが例えばパソコ
ンと呼ばれるコンピュータのアプリケーションによって
実施する場合の技術に関して以下説明する。
2. Description of the Related Art When a control device composed of a plurality of processors and a memory operated by a program is maintained by a dedicated maintenance machine (hereinafter abbreviated as "service processor SVP") provided in the device, the service processor is, for example, a personal computer. The following is a description of a technique implemented by a computer application referred to as “computer”.

【0003】制御装置がホストマシンと連動して無停止
状態で使用している場合の保守作業は、使用状態下で実
施しなければならなく、保守操作でのミスは許されな
い。このようなことから、制御装置の保守実施の直前
に、制御装置より状態情報(例えば、ディスク制御装置
のチャネルアダプタが保守閉塞している状態である等)
を取得して、その保守の続行により制御装置のシステム
ダウン、データロスト及び性能低下をさせないかチェッ
クを行なっている。また、保守を中断する際には、メッ
セージを表示することで保守員に指示または注意を促し
ている。
[0003] When the control device is used in a non-stop state in conjunction with the host machine, the maintenance work must be performed under the use state, and errors in the maintenance operation are not allowed. For this reason, immediately before the maintenance of the control device is performed, status information is transmitted from the control device (for example, a state in which the channel adapter of the disk control device is closed for maintenance).
And checks whether the system maintenance, data loss and performance degradation of the control device are not caused by the continuation of the maintenance. Further, when the maintenance is interrupted, a message is displayed to alert the maintenance staff of an instruction or a caution.

【0004】制御装置における障害の監視・検知及び、
障害解析・回復に関する技術は既に公知である。例え
ば、障害監視技術は、特開平5−324377号公報の
「プロセッサ監視システム」で示されている。また、障
害回復については、特開平6−282451号公報の
「マイクロプロセッサを備えた装置のシステムダウン復
旧方法」で示されてる。
[0004] Monitoring and detection of faults in the control device,
Techniques for failure analysis and recovery are already known. For example, a fault monitoring technique is disclosed in Japanese Patent Application Laid-Open No. 5-324377, entitled "Processor Monitoring System". The recovery from a failure is described in Japanese Patent Application Laid-Open No. 6-282451, entitled "Method of Restoring System Down of Device with Microprocessor".

【0005】これらの障害に対する復旧技術は、自動的
に行う点に特徴がある。
[0005] The feature of the restoration technique for these faults is that they are automatically performed.

【0006】[0006]

【発明が解決しようとする課題】ところが、制御装置の
状態は常に移行している。しかし、前記従来技術では装
置状態の移行に伴ったチェックはされない。つまり、保
守実施中における制御装置の障害は認識されない。この
ようなことから、保守実施中に障害が発生した場合、そ
の障害部位が実施中の保守によりシステムダウン、デー
タロスト及び制御装置性能低下等の2次障害を発生させ
る可能性がある。
However, the state of the control device is constantly shifting. However, in the above-described conventional technology, a check is not performed in accordance with the transition of the device state. That is, the failure of the control device during the maintenance is not recognized. For this reason, if a failure occurs during maintenance, the failure may cause secondary failures such as system down, data loss, and deterioration of control device performance due to the maintenance being performed.

【0007】例えば、先に示した公知例では、自動的に
障害監視から回復まで可能であるが、その回復プロセス
中に、保守員が別の保守操作を行った場合に、2次障害
を発生させる可能性がある。また、回復中に保守員の操
作を抑止する事も可能だが、障害にも重大なものから些
細なものまであり、抑止させる判断が難しく、これを解
決する従来技術はない。
[0007] For example, in the above-mentioned known example, it is possible to automatically perform from failure monitoring to recovery. However, during the recovery process, when a maintenance worker performs another maintenance operation, a secondary failure occurs. There is a possibility. It is also possible to suppress the operation of maintenance personnel during recovery, but there are serious and trivial failures, and it is difficult to judge the failure, and there is no conventional technology to solve this.

【0008】本発明は、保守作業による制御装置のシス
テムダウン、データロスト及び、性能低下を未然に防ぐ
ことが課題であり、制御装置において保守実施中の制御
装置状態をいかにチェックし抑止するかが課題である。
SUMMARY OF THE INVENTION An object of the present invention is to prevent a system failure, data loss, and performance degradation of a control device due to maintenance work, and how the control device checks and suppresses the status of the control device during maintenance. It is an issue.

【0009】[0009]

【課題を解決するための手段】前記課題を解決するため
に、主として本発明は次のような構成を採用する。
Means for Solving the Problems In order to solve the above problems, the present invention mainly employs the following constitution.

【0010】二重化されたプロセッサとメモリを有する
ディスク制御装置と、二重化されたディスク駆動装置
と、前記ディスク制御装置と前記ディスク駆動装置の保
守を実施するために前記ディスク制御装置に具備された
保守用コンピュータと、を備え、前記保守用コンピュー
タは、前記保守用コンピュータに搭載された保守用アプ
リケーションの作動に関する操作情報と、前記ディスク
制御装置および前記ディスク駆動装置の状態情報と、を
監視し、前記監視の結果、前記ディスク制御装置または
前記ディスク駆動装置に障害が発生したことが検知され
たとき、操作されている該当の保守用アプリケーション
を強制的に終了させて保守を自動的に中断させる保守抑
止装置。
A disk controller having a duplicated processor and a memory, a duplicated disk drive, and a maintenance device provided in the disk controller for performing maintenance of the disk controller and the disk drive. A computer, wherein the maintenance computer monitors operation information relating to operation of a maintenance application mounted on the maintenance computer, and status information of the disk control device and the disk drive device. As a result, when it is detected that a failure has occurred in the disk control device or the disk drive device, a maintenance suppression device for forcibly terminating the corresponding maintenance application being operated and automatically interrupting maintenance. .

【0011】また、二重化されたプロセッサとメモリを
有するディスク制御装置と、前記ディスク制御装置の保
守を実施するために前記ディスク制御装置に具備された
保守用コンピュータと、を備え、前記保守用コンピュー
タは、前記保守用コンピュータに搭載された保守用アプ
リケーションの作動に関する操作情報と、前記ディスク
制御装置の状態情報と、を監視し、前記監視の結果、前
記ディスク制御装置に障害が発生したことが検知された
とき、前記保守用アプリケーションの操作情報の来歴に
基づいて保守直前の状態に前記ディスク制御装置を復旧
させるリカバリ処理を実施する保守抑止装置。
A disk controller having a duplicated processor and a memory, and a maintenance computer provided in the disk controller for performing maintenance of the disk controller are provided. Monitoring operation information relating to the operation of the maintenance application installed in the maintenance computer and status information of the disk controller, and as a result of the monitoring, it is detected that a failure has occurred in the disk controller. And performing a recovery process for recovering the disk control device to a state immediately before maintenance based on a history of operation information of the maintenance application.

【0012】[0012]

【発明の実施の形態】本発明の実施形態について、図面
を用いて以下詳細に説明する。
Embodiments of the present invention will be described in detail below with reference to the drawings.

【0013】図1と図2に本発明の実施形態を示す。こ
こにおいて、101はディスク制御装置の状態情報の取
得機能、102は操作情報とディスク制御装置状態情報
の相互チェックの実施機能、103は操作情報、ディス
ク制御装置及び対処アプリケーションの対応を示したテ
ーブル、104はアプリケーションのインターフェー
ス、105は操作来歴を格納するテーブル、106はデ
ィスク制御装置とアプリケーションI/Fとのやり取り
を行うタスク、107は保守用のアプリケーション、1
08はリカバリ用のアプリケーション、109はメッセ
ージ表示用のアプリケーション、120はCRT、12
1はキーボード、122はマウス、123はフロッピー
ディスクドライブ、201はホストマシン、202はデ
ィスク制御装置、203はサービスプロセッサ、204
は共通バス、205,206はチャネルアダプタ、20
7,208はディスクアダプタ、209は不揮発性キャ
ッシュメモリ、210は内部バス、220はディスク駆
動装置、221はハードディスクユニット、222はデ
ィスク制御装置及び駆動装置に電力を提供するための電
源、223はSCSIケーブル、をそれぞれ表す。
FIGS. 1 and 2 show an embodiment of the present invention. In this case, 101 is a function for acquiring the status information of the disk control device, 102 is a function for performing a mutual check between the operation information and the disk control device status information, 103 is a table showing the correspondence between the operation information, the disk control device and the application to be handled, 104, an application interface; 105, a table for storing operation history; 106, a task for exchanging a disk controller with an application I / F; 107, a maintenance application;
08 is a recovery application, 109 is a message display application, 120 is a CRT, 12
1 is a keyboard, 122 is a mouse, 123 is a floppy disk drive, 201 is a host machine, 202 is a disk controller, 203 is a service processor, 204
Is a common bus, 205 and 206 are channel adapters, 20
7, 208 are disk adapters, 209 is a non-volatile cache memory, 210 is an internal bus, 220 is a disk drive, 221 is a hard disk unit, 222 is a power supply for supplying power to the disk controller and drive, and 223 is a SCSI. Cable, respectively.

【0014】図2にディスクサブシステム装置を示す。
ディスクサブシステム装置は、ディスク制御装置202
(以下「DKC」と略称する。)とディスク駆動装置2
20(以下「DKU」と略称する。)により構成されて
いる。
FIG. 2 shows a disk subsystem device.
The disk subsystem device includes a disk control device 202
(Hereinafter abbreviated as “DKC”) and the disk drive 2
20 (hereinafter abbreviated as “DKU”).

【0015】DKC202には、ホスト201間の情報
を制御するチャネルアダプタ205,206(以下「C
HA」と略称する。)、DKU220間の情報を制御す
るディスクアダプタ207,208(以下「DKA」と
略称する。)、ディスクアクセスの高速化を目的とした
不揮発性キャッシュメモリ209(以下「CM」と略称
する。)によって構成されている。
The DKC 202 has channel adapters 205 and 206 (hereinafter referred to as “C”) for controlling information between the hosts 201.
HA ”. ), Disk adapters 207 and 208 (hereinafter abbreviated as "DKA") for controlling information between DKUs 220, and a non-volatile cache memory 209 (hereinafter abbreviated as "CM") for speeding up disk access. It is configured.

【0016】また、これらは全て二重化されており、内
部バス210により接続され通信可能となっている。C
HA205,206とDKA207,208にはマイク
ロプロセッサ(以下「MP」と略称する。)が搭載され
ている。また、DKC202には保守を目的としたサー
ビスプロセッサSVP203が搭載されており、CHA
205,206/DKA207,208/CM209は
共通バス204により接続されて通信可能となってい
る。
These are all duplicated, and are connected by the internal bus 210 to enable communication. C
The HAs 205 and 206 and the DKAs 207 and 208 are equipped with a microprocessor (hereinafter abbreviated as “MP”). The DKC 202 is equipped with a service processor SVP 203 for maintenance.
205, 206 / DKA 207, 208 / CM 209 are connected by a common bus 204 to enable communication.

【0017】DKU220には、顧客情報を格納するデ
ィスク群221(以下「HDU」と略称する。)及び、
DKUとDKCに電力を供給する電源222が搭載され
ている。また、HDU221とDKA207,208を
SCSIケーブル223により接続されている。
The DKU 220 has a disk group 221 (hereinafter abbreviated as "HDU") for storing customer information, and
A power supply 222 for supplying power to the DKU and DKC is mounted. The HDU 221 and the DKAs 207 and 208 are connected by a SCSI cable 223.

【0018】図1に本発明の実施形態の構成を示す。図
1の点線で囲まれた左側部分はサービスプロセッサ20
3の内部構成を示し、その左側部分はディスク制御装置
202を示している。
FIG. 1 shows the configuration of an embodiment of the present invention. The left part surrounded by the dotted line in FIG.
3 shows the internal configuration, and the left part thereof shows the disk control device 202.

【0019】保守員によりキーボード121またはマウ
ス122により、アプリケーションAP107を起動し
保守を実行する。その後、AP107は、入力された具
体的な保守手順を分割してコマンド単位(保守閉塞と
か、または交換とか、またはリカバリとかのコマンド単
位)をAP I/F104へ発行する。AP I/F10
4は受け取ったコマンド単位を来歴105への格納とD
KC状態情報取得101へ発行をする。すなわち、コマ
ンド単位毎にDKC状態情報を取得できるようにトリガ
ーをかけているのである。
The maintenance worker starts the application AP 107 by using the keyboard 121 or the mouse 122 to execute maintenance. After that, the AP 107 divides the input specific maintenance procedure and issues a command unit (a command unit such as maintenance blockage, replacement, or recovery) to the API I / F 104. AP I / F10
4 stores the received command unit in the history 105 and D
It is issued to the KC status information acquisition 101. That is, a trigger is issued so that DKC status information can be acquired for each command unit.

【0020】DKC状態情報取得は現在のDKC状態を
DKC通信106より取得し、操作&状態(操作情報と
装置状態)の相互チェック102(カレントな情報が、
図3に示す条件式に当てはまるかどうかを突き合わせる
こと)を起動(トリガー)する。
In the DKC status information acquisition, the current DKC status is acquired from the DKC communication 106, and the operation & status (operation information and device status) mutual check 102 (current information is
(Triggering whether or not the condition expression shown in FIG. 3 is satisfied) is started (triggered).

【0021】操作&状態の相互チェック102は、操作
情報をAPI/F104を介して来歴105より取得す
る。また、チェック条件として条件式103を取得し、
その現在のDKC状態情報、操作情報及び条件式より相
互間のチェックを102が実施し、次処理の指示をAP
I/F104へ行う。
The operation & status mutual check 102 acquires operation information from the history 105 via the API / F 104. Also, a conditional expression 103 is acquired as a check condition,
Based on the current DKC status information, operation information, and conditional expression, the 102 performs a mutual check, and issues an instruction for the next process to the AP.
Perform to I / F104.

【0022】API/F104は、その指示によりDK
C通信106へのコマンド発行、リカバリ用APへのコ
マンド発行、メッセージ表示用AP109へのコマンド
発行、AP107へのコマンド発行を行う。DKC通信
106へのコマンド発行の場合、コマンドを受け取った
DKC通信106は共通バス204を通して該当MP2
05〜208へコマンドを発行する。
The API / F 104 issues a DK
It issues a command to the C communication 106, issues a command to the recovery AP, issues a command to the message display AP 109, and issues a command to the AP 107. In the case of issuing a command to the DKC communication 106, the DKC communication 106 that has received the command
Issue commands to 05-208.

【0023】このように、本発明では、制御装置の保守
は、制御装置に搭載されたサービスプロセッサから各マ
イクロプロセッサへコマンドを発行することで実施され
る為に、サービスプロセッサによりこれらのコマンド発
行単位に制御装置の状態のチェックを実施する。このチ
ェックを実施するによって、装置の状態情報を随時取得
し、操作情報及び装置状態情報により次の処理を決定す
る為の条件式が必要となる。これらを用いて、保守用ア
プリケーションから発行されたコマンドをアプリケーシ
ョンI/Fにより各タスクへ振り分ける。そうすること
により、コマンド発行単位での制御装置の状態チェック
を実現させ、制御装置のシステムダウン、データロスト
及び性能低下を未然に防ぐこととする。
As described above, in the present invention, the maintenance of the control device is performed by issuing a command from the service processor mounted on the control device to each of the microprocessors. Check the status of the control device. By executing this check, a condition expression for acquiring the state information of the apparatus as needed and determining the next processing based on the operation information and the apparatus state information is required. Using these, the command issued from the maintenance application is distributed to each task by the application I / F. By doing so, it is possible to check the status of the control device in units of command issuance, and to prevent a system down, data loss, and performance degradation of the control device.

【0024】図3に条件式を示し、都合の悪い状態を予
めテーブル化しておいてそれに準拠して処理を行い、前
記テーブルに当てはまらないときには保守作業を続行し
てもよいことを意味している。
FIG. 3 shows a conditional expression, which means that an inconvenient state is tabulated in advance, processing is performed in accordance with the table, and maintenance work can be continued when the condition does not apply to the table. .

【0025】条件式とは、操作情報とDKC状態の対応
表で、この2つの状態により次処理を決定する。この条
件式には、操作情報、DKCの装置状態、対処AP、の
大きく3つに分けられる。
The conditional expression is a correspondence table between the operation information and the DKC state, and the next processing is determined based on the two states. This conditional expression is roughly divided into three items: operation information, DKC device state, and coping AP.

【0026】操作情報には、保守を実施したアプリケー
ション(AP)名称(オペレータが保守に対応してAP
を選択するもの)、保守の種類、対象の保守部位、対象
保守の詳細部位、対象部位の状態を記載する。ここにお
いて、例えば、パッケージボードとなっているCHA1
を交換する場合、当該CHA1は二重化されているの
で、交換すべきボードを交換という保守のために保守閉
塞の状態にして、交換しようとする。
The operation information includes the name of the application (AP) that has performed the maintenance (the
), The type of maintenance, the target maintenance part, the detailed part of the target maintenance, and the state of the target part. Here, for example, CHA1 which is a package board
Is replaced, the CHA1 is duplicated, so that the board to be replaced is placed in a maintenance closed state for maintenance such as replacement, and replacement is attempted.

【0027】また、DKCの装置状態には、他資源の部
位、他資源の詳細部位、他資源部位の状態を記載する。
ここにおいて、例えば、前記当該CHA1の一方のボー
ド、即ち、他資源の状態を把握しておかないと、二重化
されているといえどもシステム全体がダウンしてしまう
おそれがある。
The device status of the DKC describes a portion of another resource, a detailed portion of another resource, and a status of another resource portion.
Here, for example, if the status of one board of the CHA 1, that is, the status of other resources is not grasped, the whole system may be down even if it is duplexed.

【0028】対処APにはリカバリまたはメッセージ表
示時の対象APを記載する。ここにおいて、例えば、シ
ステムダウンしそうな状況がでてくれば、保守閉塞して
いるボードをリカバリさせるべくリカバリAP1を動作
させるということになる。
The target AP at the time of recovery or message display is described in the coping AP. In this case, for example, if the system is going to go down, the recovery AP 1 is operated to recover the board whose maintenance is blocked.

【0029】図3の下段に示した[例]は、操作情報は
「交換APによるCHAの1番目の交換」である。ま
た、DKC状態は「DKAの3番目が障害閉塞」であ
る。その結果対処APとして「メッセージ表示用AP」
が起動され、保守員に対してメッセージにより警告を表
示する。このように、条件式は、操作情報、装置状態、
対処APにおけるそれぞれの項目が特定されて多数の考
えられ得る条件式が出来上がるものである。即ち、考え
られ得る全ての状況を把握しその対処の仕方(メッセー
ジを表示するとか、適宜のリカバリ用APを動作させる
等)までも前記条件式に含ませてしまうものである。
[Example] shown in the lower part of FIG. 3 is that the operation information is "first exchange of CHA by exchange AP". In addition, the DKC state is “the third DKA is a fault blockage”. As a result, the "AP for displaying messages"
Is started and a warning is displayed to the maintenance staff by a message. As described above, the conditional expression is composed of the operation information, the device state,
Each item in the coping AP is specified, and many possible conditional expressions are completed. In other words, the condition expression includes all possible situations and how to deal with them (such as displaying a message or operating an appropriate recovery AP).

【0030】図4はAPI/F104のフローチャート
を示す。
FIG. 4 shows a flowchart of the API / F 104.

【0031】API/F104として、保守用APから
のコマンド有無401、DKC通信からのコマンド有無
402、相互チェックからのコマンド有無403、の大
きく3つからなる。
The API / F 104 is roughly composed of three commands: a command 401 from the maintenance AP, a command 402 from the DKC communication, and a command 403 from the mutual check.

【0032】まず、保守用APからのコマンド401が
あった場合、そのコマンドの操作来歴を格納410す
る。その後、保守用APからの受信411だった場合、
DKC通信へコマンドを発行413する。また、リカバ
リAPからの受信412だった場合には、回答先を判断
414し、DKC通信への回答の場合、DKC通信へコ
マンドを発行415する。
First, when there is a command 401 from the maintenance AP, the operation history of the command is stored 410. After that, if it is received 411 from the maintenance AP,
A command 413 is issued to the DKC communication. In the case of the reception 412 from the recovery AP, the response destination is determined 414, and in the case of the response to the DKC communication, a command is issued 415 to the DKC communication.

【0033】次に、DKC通信からのコマンド402が
あった場合、そのコマンドの操作来歴を格納420す
る。その後、保守用APからの受信421だった場合、
DKC通信へコマンドを発行423する。また、リカバ
リAPからの受信422だった場合には、リカバリ用A
Pへコマンドを発行424する。
Next, when there is a command 402 from the DKC communication, the operation history of the command is stored 420. After that, if it is received 421 from the maintenance AP,
A command 423 is issued to the DKC communication. In the case of the reception 422 from the recovery AP, the recovery A
Issue command 424 to P.

【0034】次に、相互チェックからのコマンド403
があった場合、そのコマンドの操作来歴を格納430す
る。その後、操作を抑止するか判断431し、抑止する
場合には対象のリカバリAPを検索433し、その対象
リカバリ用APへコマンドを発行434する。
Next, the command 403 from the mutual check
If there is, the operation history of the command is stored 430. Thereafter, it is determined 431 whether to suppress the operation. If the operation is to be suppressed, the target recovery AP is searched 433 and a command is issued 434 to the target recovery AP.

【0035】また、抑止しない場合、対象の保守用AP
へ正常報告432を行う。
If not suppressed, the target maintenance AP
A normal report 432 is made.

【0036】これらの3つのチェックをループすること
により、常に監視する。
The three checks are looped to constantly monitor.

【0037】図5はDKC状態情報取得101のフロー
チャートを示す。DKC状態情報取得101では、AP
I/Fからのコマンド501があった場合、DKC状態
情報をDKCより取得511し、操作&DKC状態の相
互チェックへ指示を発行511する。この処理をループ
することにより、常に最新のDKC状態情報を取得す
る。
FIG. 5 shows a flowchart of the DKC status information acquisition 101. In the DKC status information acquisition 101, the AP
When there is a command 501 from the I / F, the DKC status information is acquired 511 from the DKC, and an instruction is issued 511 to the operation & DKC status mutual check. By looping this process, the latest DKC state information is always obtained.

【0038】図6は操作&状態の相互チェック102の
フローチャートを示す。操作&状態の相互チェック10
2では、相互チェック実施の指示601があった場合、
対象の条件式を検索及び取得611する。その後、条件
式とDKC状態と操作来歴のチェック612を行なう。
その結果613、保守を抑止する場合には、API/F
へリカバリ処理の指示を発行615する。また、保守を
抑止しない場合には、AP I/Fへ正常報告を発行6
14する。この処理をループすることにより、常に操作
&状態の相互チェックを実施可能な状態とする。
FIG. 6 shows a flowchart of the operation & status mutual check 102. Mutual check of operation & status 10
In 2, when the mutual check execution instruction 601 is given,
The target conditional expression is searched and acquired 611. Thereafter, a check 612 of the conditional expression, the DKC state and the operation history is performed.
As a result 613, if maintenance is to be suppressed, API / F
Then, an instruction 615 for recovery processing is issued. If the maintenance is not suppressed, a normal report is issued to the AP I / F.
14 By looping this process, it is possible to always perform the operation and state mutual check.

【0039】以上の説明では、装置状態の情報として、
主として、ディスク制御装置の状態を対象としていた
が、これに加えて、ディスク駆動装置の状態も含み得る
ことができるものである。
In the above description, the information on the device state is
Although the status of the disk control device has been mainly targeted, the status of the disk drive device can also be included in addition to this.

【0040】[0040]

【発明の効果】本発明によって、制御装置は無停止状態
での保守が要求されており、保守による2次障害を事前
及び保守中に回避することが可能となる。また、メッセ
ージ表示による操作性の向上やリカバリ処理の自動化に
より保守性及び信頼性が向上する効果がある。
According to the present invention, the control device is required to be maintained in a non-stop state, and a secondary failure due to the maintenance can be avoided in advance and during the maintenance. Further, there is an effect that maintainability and reliability are improved by improving operability by displaying a message and automating recovery processing.

【図面の簡単な説明】[Brief description of the drawings]

【図1】本発明の実施形態1を示す構成図である。FIG. 1 is a configuration diagram showing a first embodiment of the present invention.

【図2】本発明の実施形態を実施するハードウェア構成
図で、特に磁気ディスク装置のハードウェア概略図であ
る。
FIG. 2 is a hardware configuration diagram for implementing an embodiment of the present invention, and particularly a hardware schematic diagram of a magnetic disk device.

【図3】操作情報、制御装置及び対処アプリケーション
の対応を示したテーブルを示す図である。
FIG. 3 is a diagram illustrating a table indicating correspondence between operation information, a control device, and a countermeasure application;

【図4】本発明の実施形態を実現するためのフローチャ
ートである。
FIG. 4 is a flowchart for implementing an embodiment of the present invention.

【図5】本発明の実施形態を実現するためのフローチャ
ートである。
FIG. 5 is a flowchart for realizing an embodiment of the present invention.

【図6】本発明の実施形態を実現するためのフローチャ
ートである。
FIG. 6 is a flowchart for implementing an embodiment of the present invention.

【符号の説明】[Explanation of symbols]

101:ディスク制御装置の状態情報の取得機能 102:操作情報とディスク制御装置状態情報の相互チ
ェックの実施機能 103:操作情報、ディスク制御装置及び対処アプリケ
ーションの対応を示したテーブル 104:アプリケーションのインターフェース 105:操作来歴を格納するテーブル 106:ディスク制御装置とアプリケーションI/Fと
のやり取りを行うタスク 107:保守用のアプリケーション 108:リカバリ用のアプリケーション 109:メッセージ表示用のアプリケーション 120:CRT 121:キーボード 122:マウス 123:フロッピーディスクドライブ 201:ホストマシン 202:ディスク制御装置 203:サービスプロセッサ 204:共通バス 205,206:チャネルアダプタ 207,208:ディスクアダプタ 209:不揮発性キャッシュメモリ 210:内部バス 220:ディスク駆動装置 221:ハードディスクユニット 222:ディスク制御装置及び駆動装置に電力を提供す
るための電源 223:SCSIケーブル
101: Function for acquiring status information of a disk controller 102: Function for performing mutual check between operation information and disk controller status information 103: Table showing correspondence between operation information, disk controller and corresponding application 104: Interface of application 105 : Table for storing operation history 106: task for exchanging disk controller and application I / F 107: application for maintenance 108: application for recovery 109: application for message display 120: CRT 121: keyboard 122: Mouse 123: Floppy disk drive 201: Host machine 202: Disk controller 203: Service processor 204: Common bus 205, 206: Channel adapter 207, 08: Disk Adapter 209: non-volatile cache memory 210: internal bus 220: a disk drive 221: Hard Disk Unit 222: power supply for providing power to the disk controller and the drive unit 223: SCSI cable

Claims (3)

【特許請求の範囲】[Claims] 【請求項1】 二重化されたプロセッサとメモリを有す
るディスク制御装置と、二重化されたディスク駆動装置
と、前記ディスク制御装置と前記ディスク駆動装置の保
守を実施するために前記ディスク制御装置に具備された
保守用コンピュータと、を備え、 前記保守用コンピュータは、前記保守用コンピュータに
搭載された保守用アプリケーションの作動に関する操作
情報と、前記ディスク制御装置および前記ディスク駆動
装置の状態情報と、を監視し、 前記監視の結果、前記ディスク制御装置または前記ディ
スク駆動装置に障害が発生したことが検知されたとき、
操作されている該当の保守用アプリケーションを強制的
に終了させて保守を自動的に中断させることを特徴とす
る保守抑止装置。
A disk controller having a duplicated processor and a memory; a duplicated disk drive; and a disk controller provided for performing maintenance of the disk controller and the disk drive. A maintenance computer, wherein the maintenance computer monitors operation information related to operation of a maintenance application mounted on the maintenance computer, and status information of the disk control device and the disk drive device, As a result of the monitoring, when it is detected that a failure has occurred in the disk control device or the disk drive device,
A maintenance suppression device for forcibly terminating a corresponding maintenance application being operated and automatically suspending maintenance.
【請求項2】 二重化されたプロセッサとメモリを有す
るディスク制御装置と、前記ディスク制御装置の保守を
実施するために前記ディスク制御装置に具備された保守
用コンピュータと、を備え、 前記保守用コンピュータは、前記保守用コンピュータに
搭載された保守用アプリケーションの作動に関する操作
情報と、前記ディスク制御装置の状態情報と、を監視
し、 前記監視の結果、前記ディスク制御装置に障害が発生し
たことが検知されたとき、前記保守用アプリケーション
の操作情報の来歴に基づいて保守直前の状態に前記ディ
スク制御装置を復旧させるリカバリ処理を実施すること
を特徴とする保守抑止装置。
2. A disk control device having a duplicated processor and a memory, and a maintenance computer provided in the disk control device for performing maintenance of the disk control device, wherein the maintenance computer is Monitoring operation information related to the operation of the maintenance application mounted on the maintenance computer and status information of the disk controller, and as a result of the monitoring, it is detected that a failure has occurred in the disk controller. And performing a recovery process for recovering the disk control device to a state immediately before maintenance based on a history of operation information of the maintenance application.
【請求項3】 二重化されたプロセッサとメモリを有す
るディスク制御装置と、前記ディスク制御装置の保守を
実施するために前記ディスク制御装置に具備された保守
用コンピュータと、を備え、 前記保守用コンピュータは、前記保守用コンピュータに
搭載された保守用アプリケーションの作動に関する操作
情報と、前記ディスク制御装置の状態情報と、 を監視するともとに、 前記操作情報と前記状態情報と障害時の対処アプリケー
ションとの組み合わせからなる種々の障害対処テーブル
を予め具備し、 前記操作情報と前記状態情報とを監視してこれらの情報
と前記テーブルと照合し、 前記テーブルに該当するときには、保守直前の状態に前
記ディスク制御装置を復旧させるリカバリ処理を実施す
ることを特徴とする保守抑止装置。
3. A disk controller having a duplicated processor and a memory, and a maintenance computer provided in the disk controller to perform maintenance of the disk controller, wherein the maintenance computer is Monitoring the operation information on the operation of the maintenance application mounted on the maintenance computer and the status information of the disk control device, and monitors the operation information, the status information, and the failure handling application. Various failure handling tables comprising a combination are provided in advance, and the operation information and the status information are monitored and compared with the information and the table. A maintenance suppressing device for performing a recovery process for recovering the device.
JP13121197A 1997-05-21 1997-05-21 Maintenance deterrence device for preventive maintenance Expired - Fee Related JP3961617B2 (en)

Priority Applications (1)

Application Number Priority Date Filing Date Title
JP13121197A JP3961617B2 (en) 1997-05-21 1997-05-21 Maintenance deterrence device for preventive maintenance

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
JP13121197A JP3961617B2 (en) 1997-05-21 1997-05-21 Maintenance deterrence device for preventive maintenance

Publications (3)

Publication Number Publication Date
JPH10320236A true JPH10320236A (en) 1998-12-04
JPH10320236A5 JPH10320236A5 (en) 2005-03-17
JP3961617B2 JP3961617B2 (en) 2007-08-22

Family

ID=15052645

Family Applications (1)

Application Number Title Priority Date Filing Date
JP13121197A Expired - Fee Related JP3961617B2 (en) 1997-05-21 1997-05-21 Maintenance deterrence device for preventive maintenance

Country Status (1)

Country Link
JP (1) JP3961617B2 (en)

Also Published As

Publication number Publication date
JP3961617B2 (en) 2007-08-22

Similar Documents

Publication Publication Date Title
EP2733611B1 (en) Internal fault handling method, device and system for virtual machine
US7236454B2 (en) Loop diagnosis system and method for disk array apparatuses
JP2003345531A (en) Storage system, management server, and application management method
US20120266027A1 (en) Storage apparatus and method of controlling the same
US8117501B2 (en) Virtual library apparatus and method for diagnosing physical drive
JP2000132413A (en) Error retry method, error retry system and its recording medium
JPH10320236A (en) Maintenance deterrent device for preventive maintenance
CN119806745A (en) Cloud platform virtual machine operating system anomaly detection and recovery method, device and medium
JP3420919B2 (en) Information processing device
JPH03179538A (en) Data processing system
JPH0541706A (en) Network automatic monitor and control system
US7962781B2 (en) Control method for information storage apparatus, information storage apparatus and computer readable information recording medium
CN120560894B (en) Memory fault management system, method, server and electronic equipment
JP2559771B2 (en) Line logging automatic stop control method
JP2011159234A (en) Fault handling system and fault handling method
KR100257162B1 (en) Monitoring method and device of counterpart system in redundant system
JP3334174B2 (en) Fault handling verification device
JP2695552B2 (en) Failure handling method
JPH09259050A (en) Computer peripheral device control device error reporting method and peripheral device control device
JP2025008095A (en) Control device and control program
JPH05274093A (en) Volume fault prevention control system
JPH02230338A (en) Fault report system
JPH1153217A (en) Distributed object management system
CN115373801A (en) SAN storage-based method and system for automatically recovering virtual machine to prevent split brain
JPH02293937A (en) Fault reporting system

Legal Events

Date Code Title Description
A621 Written request for application examination

Free format text: JAPANESE INTERMEDIATE CODE: A621

Effective date: 20040220

A521 Written amendment

Free format text: JAPANESE INTERMEDIATE CODE: A523

Effective date: 20040416

A977 Report on retrieval

Free format text: JAPANESE INTERMEDIATE CODE: A971007

Effective date: 20061012

A131 Notification of reasons for refusal

Free format text: JAPANESE INTERMEDIATE CODE: A131

Effective date: 20061024

A521 Written amendment

Free format text: JAPANESE INTERMEDIATE CODE: A523

Effective date: 20061213

A131 Notification of reasons for refusal

Free format text: JAPANESE INTERMEDIATE CODE: A131

Effective date: 20070306

A521 Written amendment

Free format text: JAPANESE INTERMEDIATE CODE: A523

Effective date: 20070329

TRDD Decision of grant or rejection written
A01 Written decision to grant a patent or to grant a registration (utility model)

Free format text: JAPANESE INTERMEDIATE CODE: A01

Effective date: 20070508

A61 First payment of annual fees (during grant procedure)

Free format text: JAPANESE INTERMEDIATE CODE: A61

Effective date: 20070517

R150 Certificate of patent or registration of utility model

Free format text: JAPANESE INTERMEDIATE CODE: R150

FPAY Renewal fee payment (event date is renewal date of database)

Free format text: PAYMENT UNTIL: 20110525

Year of fee payment: 4

FPAY Renewal fee payment (event date is renewal date of database)

Free format text: PAYMENT UNTIL: 20110525

Year of fee payment: 4

FPAY Renewal fee payment (event date is renewal date of database)

Free format text: PAYMENT UNTIL: 20120525

Year of fee payment: 5

FPAY Renewal fee payment (event date is renewal date of database)

Free format text: PAYMENT UNTIL: 20120525

Year of fee payment: 5

FPAY Renewal fee payment (event date is renewal date of database)

Free format text: PAYMENT UNTIL: 20130525

Year of fee payment: 6

FPAY Renewal fee payment (event date is renewal date of database)

Free format text: PAYMENT UNTIL: 20130525

Year of fee payment: 6

LAPS Cancellation because of no payment of annual fees