JPH01217664A - Fault detecting system for multiprocessor system - Google Patents
Fault detecting system for multiprocessor systemInfo
- Publication number
- JPH01217664A JPH01217664A JP63042316A JP4231688A JPH01217664A JP H01217664 A JPH01217664 A JP H01217664A JP 63042316 A JP63042316 A JP 63042316A JP 4231688 A JP4231688 A JP 4231688A JP H01217664 A JPH01217664 A JP H01217664A
- Authority
- JP
- Japan
- Prior art keywords
- processor
- health signal
- timer device
- signal output
- monitoring timer
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
- 238000012544 monitoring process Methods 0.000 claims description 38
- 238000001514 detection method Methods 0.000 claims description 10
- 230000005856 abnormality Effects 0.000 abstract description 3
- 238000010586 diagram Methods 0.000 description 3
- 241000862969 Stella Species 0.000 description 1
Landscapes
- Multi Processors (AREA)
- Debugging And Monitoring (AREA)
Abstract
Description
【発明の詳細な説明】
〔産業上の利用分野〕
本発明はマルチプロセッサシステムの障害検出方式に関
し、特にプロセラ、す異常の検出方式に関する。DETAILED DESCRIPTION OF THE INVENTION [Field of Industrial Application] The present invention relates to a fault detection method for a multiprocessor system, and more particularly to a fault detection method for a processor.
従来のマルチプロセッサシステムでの障害検出方式では
、外部装置として障害監視タイマ装置をプロセッサの個
数分持ち、それぞれのプロセッサと障害監視タイマ装置
を1対1に対応させて、各プロセッサが対応する障害監
視タイマ装置に対して一定時間間隔で出力する信号をプ
ロセッサが正常であることを示すヘルス信号として出力
し、ヘルス信号が途絶えた場合にプロセッサ異常として
障害監視タイマ装置が割込みを起こすことで異常検出を
行っていた。In conventional fault detection methods in multiprocessor systems, fault monitoring timer devices are provided as external devices for the number of processors, and each processor and fault monitoring timer device are in one-to-one correspondence. Anomaly detection is done by outputting a signal to the timer device at regular intervals as a health signal indicating that the processor is normal, and when the health signal is interrupted, the fault monitoring timer device generates an interrupt as a processor error. I was going.
上述した従来の障害検出方式では、プロセッサの個数分
障害監視タイマ装置が必要となるうえに。In the conventional fault detection method described above, fault monitoring timer devices are required for each processor.
各々プロセッサがヘルス信号を出力するためにシステム
全体としての障害検出処理のオーバヘットが大きく、さ
らには障害監視タイマ装置が異常を検出した場合の割込
みが、障害を起こしたプロセッサに対して発生するため
、障害監視タイマ装置からの割込み処理が正常に行われ
るとは限らないという欠点があった。Since each processor outputs a health signal, the overhead of fault detection processing for the entire system is large, and furthermore, when the fault monitoring timer device detects an abnormality, an interrupt is generated for the faulty processor. There is a drawback that interrupt processing from the fault monitoring timer device is not always performed normally.
本発明によれば、主記憶と、該主記憶にバスを介して接
続された複数のプロセッサと、該プロセッサと独立に動
作し一定時間内に前記プロセッサからヘルス信号を受信
しない時には最後にヘルス信号を出力したプロセッサに
対して割込みを発生する障害監視タイマ装置と、前記プ
ロセッサと独立に動作し時刻を計時する外部タイマ装置
とを備エタマルチプロセッサシステムに於いて。According to the present invention, a main memory, a plurality of processors connected to the main memory via a bus, and a plurality of processors that operate independently of the processor, and when a health signal is not received from the processor within a certain period of time, a health signal is sent at the end. A multiprocessor system is provided with a fault monitoring timer device that generates an interrupt to a processor that outputs a signal, and an external timer device that operates independently of the processor and measures time.
前記主記憶は、前記外部タイマ装置により更新されるシ
ステム時刻領域と。The main memory includes a system time area updated by the external timer device.
前記障害監視タイマ装置に対してヘルス信号を出力する
プロセッサを記憶するヘルス信号出力プロセッサ記憶領
域とを有し。and a health signal output processor storage area for storing a processor that outputs a health signal to the failure monitoring timer device.
前記各プロセッサは、前記システム時刻領域から時刻を
読み出しいずれのプロセッサが前記障害監視タイマ装置
にヘルス信号を出方するかを選択するヘルス信号出力選
択手段と、前記障害監視タイマ装置に対してヘルス信号
を出力するヘルス信号出力手段と、いずれかのプロセッ
サが障害により前記障害監視タイマ装置が割込みを発生
した場合の割込みを受付ける障害監視タイマ装置割込み
受付は手段と、前記ヘルス信号出方プロセッサ記憶領域
から障害を発生したプロセッサを判別する障害プロセッ
サ判別手段とを有することを特徴とするマルチプロセッ
サシステム障害検出方式が得られる。Each of the processors includes a health signal output selection unit that reads time from the system time area and selects which processor will output a health signal to the fault monitoring timer device, and health signal output means for outputting a health signal, a fault monitoring timer device for accepting an interrupt when the fault monitoring timer device generates an interrupt due to a fault in one of the processors; and a means for accepting an interrupt from the health signal output processor storage area. A multiprocessor system failure detection method is obtained, which is characterized by having a failure processor determination means for determining a processor in which a failure has occurred.
ヘルス信号出力時点に於いては、前記ヘルス信号出力選
択手段により前記システム時刻領域から得られたシステ
ム時刻と前記ヘルス信号出方プロセッサ記憶領域からヘ
ルス信号を出力するプロセッサを選択し。At the time of outputting the health signal, the health signal output selection means selects the system time obtained from the system time area and the processor that outputs the health signal from the health signal output processor storage area.
前記プロセッサいずれかの異常発生時における前記障害
監視タイマ装置からの割込み時点に於いては、前記障害
プロセッサ判別手段にょシ障害発生のプロセッサを検出
することで、単一障害監視タイマ装置により複数のプロ
セッサの障害を検出できる。At the time of an interrupt from the fault monitoring timer device when an abnormality occurs in any of the processors, the faulty processor determining means detects the faulty processor, and the single fault monitoring timer device detects the faulty processor. failures can be detected.
次に2本発明について図面を参照して詳細に説明する。 Next, two aspects of the present invention will be described in detail with reference to the drawings.
第1図はプロセッサがN個の場合の一実施例であり、プ
ロセッサ1〜N、障害監視タイマ装置10゜外部タイマ
装置20.及び主記憶3oを有する。FIG. 1 shows an embodiment in which there are N processors, including processors 1 to N, a fault monitoring timer device 10°, an external timer device 20. and a main memory 3o.
主記憶30は、外部タイマ装置20により計時された時
刻を記憶するシステム時刻領域4oと。The main memory 30 includes a system time area 4o that stores the time measured by the external timer device 20.
障害監視タイマ装置10に対してヘルス信号を出力する
プロセッサを記憶するヘルス信号出力プロセッサ記憶領
域50とを有する。各プロセッサは。It has a health signal output processor storage area 50 that stores processors that output health signals to the fault monitoring timer device 10. Each processor.
システム時刻領域40とヘルス信号出力プロセッサ記憶
領域50からプロセッサ1〜Nのいずれがヘルス信号を
出力するかを選択するヘルス信号出力選択手段60.障
害監視タイマ装置10に対してヘルス信号を出力するヘ
ルス信号出力手段70゜プロセッサ1〜Nのいずれかの
障害にょシ障害監視タイマ装置10が割込みを発生した
場合の割込みを受付ける障害監視タイマ装置割込み受付
は手段80.及びヘルス信号出力プロセッサ記憶領域5
0から障害が発生したプロセッサを判別する障害プロセ
ッサ判別手段9oを有する。なお9図中には、これら手
段60,70,80.及び9oは。Health signal output selection means 60 for selecting which of the processors 1 to N will output a health signal from the system time area 40 and the health signal output processor storage area 50. A health signal output means 70 for outputting a health signal to the fault monitoring timer device 10; a fault monitoring timer device interrupt that accepts an interrupt when the fault monitoring timer device 10 generates an interrupt in the event of a fault in any of the processors 1 to N; Reception is by means 80. and health signal output processor storage area 5
It has a faulty processor determining means 9o for determining a processor in which a fault has occurred from zero. Note that these means 60, 70, 80 . and 9o.
主記憶30にプログラムとして格納されているように示
されている。It is shown as being stored as a program in the main memory 30.
障害監視タイマ装置10は、一定時間ヘルス信号を受信
しないと最後にヘルス信号を出力したプロセッサに対し
て割込みを起こすが、この一定時間を監視時間Tとする
。まだ各プロセッサ1〜Nが障害監視タイマ装置10に
対してヘルス信号を出力する時間間隔をヘルス信号出力
時間間隔tとすると2両者の間には
t<T〈2t ・・・・・・・・・・・・(1)の関
係がある。If the fault monitoring timer device 10 does not receive a health signal for a certain period of time, it will cause an interrupt to the processor that last outputted the health signal, and this certain period of time is defined as a monitoring time T. If the time interval at which each processor 1 to N outputs a health signal to the fault monitoring timer device 10 is the health signal output time interval t, then t<T<2t between the two. ...There is the relationship (1).
第2図は、全てのプロセッサ1〜Nが正常な場合に、障
害監視タイマ装置10に対してヘルス信号を出力してい
く動作を時系列に表した図である。FIG. 2 is a diagram chronologically showing the operation of outputting a health signal to the fault monitoring timer device 10 when all the processors 1 to N are normal.
まず第2図を参照しながら、全てのプロセッサが正常な
場合の本発明の動作を時系列に説明する。First, with reference to FIG. 2, the operation of the present invention when all processors are normal will be explained in chronological order.
システム時刻Tnにおいて、プロセッサnがヘルス信号
を出力する場合の動作を第2図及び第4図を参照しなが
ら詳細に説明する。The operation when processor n outputs a health signal at system time Tn will be described in detail with reference to FIGS. 2 and 4.
プロセッサnに於いて、システム時刻Tnの時点でヘル
ス信号出力選択手段60によりヘルス信号の出力選択が
行われる。In processor n, the health signal output selection means 60 selects the health signal output at system time Tn.
第4図を参照すると、この選択は先ず、システム′時刻
領域40から時刻Tnが読み出されヘルス信号出力タイ
ミングと判断される(ステップ61)。Referring to FIG. 4, in this selection, first, time Tn is read out from the system' time area 40 and determined to be the health signal output timing (step 61).
次にヘルス信号出力プロセッサ記憶領域50から最後に
ヘルス信号を出力したプロセッサn−1が読み出される
(ステップ62)。そこで次にヘルス信号を出力するプ
ロセッサn = ((n−1)+1 )が選択される(
ステップ63)。最後にヘルス信号出力プロセッサ記憶
領域50が選択されたプロセッサnに更新される(ステ
ップ64)。選択されたプロセッサnは、ヘルス信号出
力手段70により障害監視タイマ装置10に対してヘル
ス信号を出力する。Next, the processor n-1 that last outputted the health signal is read from the health signal output processor storage area 50 (step 62). Therefore, the processor n = ((n-1)+1) that will output the health signal next is selected (
Step 63). Finally, the health signal output processor storage area 50 is updated to the selected processor n (step 64). The selected processor n outputs a health signal to the fault monitoring timer device 10 by the health signal output means 70.
次に時間を経過後、つまり時刻Tn+1に再度ヘルス信
号出力選択手段60により、ヘルス信号出力プロセッサ
としてプロセッサn+1が選択され。Next, after a lapse of time, that is, at time Tn+1, the health signal output selection means 60 again selects the processor n+1 as the health signal output processor.
ヘルス信号出力プロセッサ記憶領域50にはn+1が記
憶され、ヘルス信号出力手段70によりプロセッサn
+ 1からヘルス信号が障害監視タイマ装置10に出力
される。n+1 is stored in the health signal output processor storage area 50, and the health signal output means 70 outputs the processor n.
+1, a health signal is output to the fault monitoring timer device 10.
以上述べてきたように、全てのプロセッサ1〜Nが正常
な場合には、7°ロセツサ1から順にプロセッサ2,3
・・・Nとヘルス信号をヘルス信号出力時間間隔tで出
力する。As described above, when all processors 1 to N are normal, processors 2 and 3 are loaded in order from 7° processor 1.
. . . N and a health signal are output at a health signal output time interval t.
第3図は、プロセッサn +1が障害となった場合の動
作を時系列にあられした図である。FIG. 3 is a chronological diagram showing the operation when processor n+1 becomes a failure.
システム時刻Tn’に於いてプロセッサn + 1が障
害となった場合の障害検出動作を第3図及び第5図を参
照しながら詳細に説明する。The fault detection operation when processor n+1 becomes faulty at system time Tn' will be described in detail with reference to FIGS. 3 and 5.
システム時刻Tnにおいて、ヘルス信号出力選択手段6
0によりプロセッサnが選択され、ヘルス信号出力プロ
セッサ記憶領域50にプロセッサnが更新され、ヘルス
信号出力手段70により障害監視タイマ装置10に対し
てプロセッサnがヘルス信号を出力する。At system time Tn, the health signal output selection means 6
0 selects the processor n, the health signal output processor storage area 50 is updated with the processor n, and the health signal output means 70 outputs a health signal to the fault monitoring timer device 10.
次にシステム時刻Tn’に於いてプロセッサn + 1
が障害を起こしたとする。ここでシステム時刻Tn+1
に於いてプロセッサn + 1がヘルス信号出力プロセ
ッサとして選択されるはずだが、プロセッサn + 1
は障害を起こしている為ヘルス信号は出力されない。Next, at system time Tn', processor n + 1
Suppose that a problem occurs. Here, system time Tn+1
Processor n + 1 should be selected as the health signal output processor in
is in trouble, so no health signal is output.
そこでシステム時刻Tn+1に於いて障害監視タイマ装
置10は最後にヘルス信号を出力したプロセッサnに対
して割込みを起こす。プロセッサnは障害監視タイマ装
置割込み受付は手段80により、障害監視タイマ装置1
0からの割込みを受付ける。最後に障害プロセッサ判別
手段90により障害が発生したプロセッサが判別される
。Therefore, at system time Tn+1, the fault monitoring timer device 10 causes an interrupt to the processor n that last outputted the health signal. The processor n accepts the fault monitoring timer device interrupt by the means 80, and receives the fault monitoring timer device 1.
Accepts interrupts from 0. Finally, the faulty processor determining means 90 identifies the processor in which the fault has occurred.
第5図を参照すると、この判別は先ず、ヘルス信号出力
プロセッサ記憶領域50から、最後にヘルス信号を出力
したプロセッサnが読み出され(ステラ7’91)、、
障害が発生したプロセッサはプロセッサn + 1と判
別される(ステップ92)。Referring to FIG. 5, in this determination, first, the processor n that last outputted the health signal is read out from the health signal output processor storage area 50 (Stella 7'91),
The processor in which the failure has occurred is determined to be processor n + 1 (step 92).
以上のような手順により、障害が発生したプロセッサが
検出される。Through the steps described above, a processor in which a failure has occurred is detected.
以上説明したように本発明は、プロセッサの個数分だけ
障害監視タイマ装置を持つことなく単一の障害監視タイ
マ装置で各プロセッサの障害を検出できるという効果が
ある。As described above, the present invention has the advantage that a single fault monitoring timer device can detect a fault in each processor without having to provide fault monitoring timer devices corresponding to the number of processors.
第1図は本発明の一実施例の構成を示すブロック図、第
2図は全てのプロセッサが正常な場合のヘルス信号出力
動作を時系列に示すタイムチャート、第3図はあるプロ
セッサが障害を起こした場合の障害検出動作を時系列に
示すタイムチャート。
第4図は第1図中のヘルス信号出力選択手段60の動作
を説明するための流れ図、第5図は第1図中の障害プロ
セッサ判別手段90の動作を説明するだめの流れ図であ
る。
1〜N・・・プロセッサ、10・・・障害監視タイマ装
置、20・・・外部タイマ装置、30・・・主記憶、4
0・・・システム時刻領域、50・・・ヘルス信号出力
プロセッサ記憶領域、60・・・ヘルス信号出力選択手
段。
70・・・ヘルス信号出力手段、80・・・障害監視タ
イマ装置割込み受付は手段、90・・・障害プロセッサ
第1図
第2図
第3図
第4図
第5図
ロコFig. 1 is a block diagram showing the configuration of an embodiment of the present invention, Fig. 2 is a time chart showing the health signal output operation in chronological order when all processors are normal, and Fig. 3 is a time chart showing the health signal output operation when all processors are normal. A time chart showing the failure detection operation in chronological order when a failure occurs. FIG. 4 is a flowchart for explaining the operation of the health signal output selection means 60 in FIG. 1, and FIG. 5 is a flowchart for explaining the operation of the faulty processor determining means 90 in FIG. 1. 1 to N... Processor, 10... Fault monitoring timer device, 20... External timer device, 30... Main memory, 4
0...System time area, 50...Health signal output processor storage area, 60...Health signal output selection means. 70... Health signal output means, 80... Fault monitoring timer device interrupt reception means, 90... Fault processor (Figure 1, Figure 2, Figure 3, Figure 4, Figure 5)
Claims (1)
のプロセッサと、該プロセッサと独立に動作し一定時間
内に前記プロセッサからヘルス信号を受信しない時には
最後にヘルス信号を出力したプロセッサに対して割込み
を発生する障害監視タイマ装置と、前記プロセッサと独
立に動作し時刻を計時する外部タイマ装置とを備えたマ
ルチプロセッサシステムに於いて、 前記主記憶は、前記外部タイマ装置により更新されるシ
ステム時刻領域と、前記障害監視タイマ装置に対してヘ
ルス信号を出力するプロセッサを記憶するヘルス信号出
力プロセッサ記憶領域とを有し、 前記各プロセッサは、前記システム時刻領域から時刻を
読み出しいずれのプロセッサが前記障害監視タイマ装置
にヘルス信号を出力するかを選択するヘルス信号出力選
択手段と、前記障害監視タイマ装置に対してヘルス信号
を出力するヘルス信号出力手段と、いずれかのプロセッ
サが障害により前記障害監視タイマ装置が割込みを発生
した場合の割込みを受付ける障害監視タイマ装置割込み
受付け手段と、前記ヘルス信号出力プロセッサ記憶領域
から障害を発生したプロセッサを判別する障害プロセッ
サ判別手段とを有することを特徴とするマルチプロセッ
サシステム障害検出方式。[Scope of Claims] 1. A main memory, a plurality of processors connected to the main memory via a bus, and a processor that operates independently of the processors and, when a health signal is not received from the processors within a certain period of time, In a multiprocessor system comprising a fault monitoring timer device that generates an interrupt to a processor that outputs a health signal, and an external timer device that operates independently of the processor and measures time, the main memory includes the a system time area that is updated by an external timer device; and a health signal output processor storage area that stores processors that output health signals to the fault monitoring timer device; Health signal output selection means for reading time and selecting which processor outputs a health signal to the fault monitoring timer device; and health signal output means for outputting a health signal to the fault monitoring timer device. fault monitoring timer device interrupt accepting means for accepting an interrupt when the fault monitoring timer device generates an interrupt due to a fault in the processor; and faulty processor determining means for determining the faulty processor from the health signal output processor storage area. A multiprocessor system failure detection method characterized by having the following.
Priority Applications (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP63042316A JPH01217664A (en) | 1988-02-26 | 1988-02-26 | Fault detecting system for multiprocessor system |
Applications Claiming Priority (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP63042316A JPH01217664A (en) | 1988-02-26 | 1988-02-26 | Fault detecting system for multiprocessor system |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| JPH01217664A true JPH01217664A (en) | 1989-08-31 |
Family
ID=12632613
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| JP63042316A Pending JPH01217664A (en) | 1988-02-26 | 1988-02-26 | Fault detecting system for multiprocessor system |
Country Status (1)
| Country | Link |
|---|---|
| JP (1) | JPH01217664A (en) |
-
1988
- 1988-02-26 JP JP63042316A patent/JPH01217664A/en active Pending
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| JPH01293450A (en) | Troubled device specifying system | |
| US20070271486A1 (en) | Method and system to detect software faults | |
| JPH11337655A (en) | Earthquake warning monitoring control processor | |
| JPH08307438A (en) | Token ring type transmission system | |
| JPS6115239A (en) | Processor diagnosis system | |
| JP2512325B2 (en) | Fan failure detection device | |
| CN115421948B (en) | Method for detecting memory data faults and related equipment thereof | |
| JPS6259435A (en) | Data transfer supervisory equipment | |
| JPH0581080A (en) | Microprocessor runaway monitoring device | |
| JPH0264745A (en) | Interface controller | |
| JPH03123230A (en) | Early alarm detector relating to network monitor system | |
| JP2531372B2 (en) | Fault detection circuit | |
| JP2580311B2 (en) | Mutual monitoring processing method of multiplex system | |
| JPH04352041A (en) | Device and method for collecting log information for information processing system | |
| JPH02281368A (en) | Trouble detecting mechanism for controller | |
| JPH07334433A (en) | Bus controller | |
| JP2606160B2 (en) | Failure detection method for parity check circuit | |
| JPH0898278A (en) | Digital control system | |
| JP2007042017A (en) | Fault diagnostic system, fault diagnostic method, and fault diagnostic program | |
| JPS6244847A (en) | System supervising method | |
| JPH05292068A (en) | Signal switching system | |
| JPH0237433A (en) | Monitor method for multiprocessor system | |
| JPH04273741A (en) | Package comprising device | |
| JPH07253995A (en) | Transmission device with self-diagnosis function | |
| JPH0247947A (en) | Common bus fault processor detection system |