JPH02216577A - Fault detecting system in multi-processor system - Google Patents
Fault detecting system in multi-processor systemInfo
- Publication number
- JPH02216577A JPH02216577A JP1037431A JP3743189A JPH02216577A JP H02216577 A JPH02216577 A JP H02216577A JP 1037431 A JP1037431 A JP 1037431A JP 3743189 A JP3743189 A JP 3743189A JP H02216577 A JPH02216577 A JP H02216577A
- Authority
- JP
- Japan
- Prior art keywords
- processor
- operation monitoring
- processors
- monitoring signal
- pair
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
- 238000012544 monitoring process Methods 0.000 claims abstract description 28
- 238000001514 detection method Methods 0.000 claims description 18
- 230000005540 biological transmission Effects 0.000 abstract description 2
- 238000010586 diagram Methods 0.000 description 5
- 238000000034 method Methods 0.000 description 2
Landscapes
- Debugging And Monitoring (AREA)
- Multi Processors (AREA)
Abstract
Description
【発明の詳細な説明】
〔産業上の利用分野〕
本発明はマルチプロセッサシステムにおける障害検出方
式に関する。DETAILED DESCRIPTION OF THE INVENTION [Field of Industrial Application] The present invention relates to a failure detection method in a multiprocessor system.
(従来の技術〕
従来、この種のマルチプロセッサシステムにおける障害
検出方式は、システム内の各プロセッサの動作状態を管
理する管理用プロセッサからすべてのプロセッサに対し
て動作監視信号を順次送信して、返送されてくる応答信
号を監視し、所定の時間内に応答信号が受信できなけれ
ばそのプロセッサは障害であると判断することによりシ
ステム内の全プロセッサの障害発生を検出していた。(Prior Art) Conventionally, a failure detection method in this type of multiprocessor system involves sequentially transmitting operation monitoring signals to all processors from a management processor that manages the operating status of each processor in the system, and then returning the signals. The occurrence of a failure in all processors in the system is detected by monitoring incoming response signals and determining that the processor is at fault if the response signal is not received within a predetermined time.
第3図は従来のマルチプロセッサシステムにおける障害
検出方式の一例を示すシステムブロック図である。FIG. 3 is a system block diagram showing an example of a failure detection method in a conventional multiprocessor system.
本例のマルチプロセッサシステムは1個の管理用プロセ
ッサ21と3個のプロセッサ31.〜33をリング形の
バス11で結合した形で構成されている、管理用プロセ
ッサ21はプロセッサ31.〜33の動作状態を管理し
ており、プロセッサ31に対して動作監視信号Slを送
信して所定時間内にその応答信号S3が受信されること
を監視し、その結果をプロセッサ31の動作状態として
管理し、引続きプロセッサ32.プロセッサ33に対し
て順次同様の手順を繰り返すことにより各プロセッサ3
1.〜33の障害検出を行っている。The multiprocessor system of this example includes one management processor 21 and three processors 31. The management processor 21 is composed of processors 31 . ~ 33, transmits an operation monitoring signal Sl to the processor 31, monitors that the response signal S3 is received within a predetermined time, and uses the result as the operating status of the processor 31. The processor 32. By sequentially repeating the same procedure for the processors 33, each processor 3
1. ~33 failures have been detected.
上述した従来のマルチプロセッサシステムにおける障害
検出方式は、管理用プロセッサから他の全プロセッサに
対して順次動作監視信号を送信してその応答信号を監視
する方式のため、システム内のプロセッサ数が多くなる
と、動作監視信号の送出数が増えて管理用プロセッサの
負荷が増大するという欠点がある。一方、個々のプロセ
ッサから見ると動作監視信号を受信する間隔が大きくな
って、障害検出までの時間が長くなるという欠点がある
。The fault detection method in the conventional multiprocessor system described above is a method in which the management processor sequentially sends an operation monitoring signal to all other processors and monitors the response signals, so when the number of processors in the system increases, However, the disadvantage is that the number of operation monitoring signals to be sent increases, increasing the load on the management processor. On the other hand, from the point of view of individual processors, there is a drawback that the interval at which the operation monitoring signal is received becomes long, and the time until a failure is detected becomes long.
本発明のマルチプロセッサシステムにおける障害検出方
式は、複数のプロセッサとこの各プロセッサの動作状態
を管理する1個の管理用プロセッサとをバスによって結
合して分散処理を行うマルチプロセッサシステムにおい
て、前記複数のプロセッサの2個ずつを1つのベアとし
この1つのベア内の前記2個のプロセッサ間で動作監視
信号を交互に送受信する送受信手段と、一方の前記プロ
セッサが前記動作監視信号を送信した後所定の時間内に
相手プロセッサからの前記動作監視信号を受信できない
ときは前記一方のプロセッサは前記管理用プロセッサに
前記相手プロセッサが障害である旨を通知する通知手段
を備えることを特徴とする。A failure detection method in a multiprocessor system according to the present invention is a multiprocessor system in which a plurality of processors and one management processor that manages the operating state of each processor are connected via a bus to perform distributed processing. transmitting/receiving means for alternately transmitting and receiving operation monitoring signals between the two processors in one bear, each of which has two processors; The one processor is characterized in that, when the operation monitoring signal cannot be received from the partner processor within a time, the one processor includes notification means for notifying the management processor that the partner processor is in trouble.
次に本発明について第1図、第2図を参照して説明する
。Next, the present invention will be explained with reference to FIGS. 1 and 2.
第1図は本発明のマルチプロセッサシステムにおける障
害検出方式の一実施例を示すシステムブロック図であり
、(a)は通常状態での障害検出動作を、(b)は障害
発生時の動作を、(c)は相手プロセッサが障害中で管
理用プロセッサによって代行される障害検出動作を示し
ている。また第2図は第1図における一連の障害検出動
作のシーケンスの一例を示すシーケンス図である。FIG. 1 is a system block diagram showing an embodiment of the fault detection method in the multiprocessor system of the present invention, in which (a) shows the fault detection operation in a normal state, and (b) shows the operation when a fault occurs. (c) shows a failure detection operation performed by the management processor when the other processor is in failure. Further, FIG. 2 is a sequence diagram showing an example of a sequence of a series of failure detection operations in FIG. 1.
始めに第1図を参照して本実施例の障害検出動作につい
て説明する。First, the fault detection operation of this embodiment will be explained with reference to FIG.
本実施例におけるマルチプロセッサシステムは、管理用
プロセッサ21と一般のプロセッサ318〜36がリン
グ形のバス11によって結合された形で構成されている
。プロセッサ31とプロセッサ32でベア41、プロセ
ッサ33とプロセッサ34でベア42、プロセッサ35
とプロセッサ36でベア43の計3つのベアを形成し、
第1図(a)に示すように各ベア内の2つのプロセッサ
間で互いに動作監視信号S1を送受信している。この状
態で例えばプロセッサ32に障害が発生すると、第1図
(b)に示すようにベア41の中で送受信されていた動
作監視信号S1はストップし、それを相手プロセッサ3
1によって検出し、プロセッサ31は相手障害通知信号
S2を管理用プロセッサ21に送信する。この相手障害
通知信号S2を受信した管理用プロセッサ21は、第1
図(c)に示すようにプロセッサ32が障害中の期間は
プロセッサ31の仮のベアの相手となってプロセッサ3
1との間で動作監視信号S1の送受信を交互に行う、な
おこの間も、他の2つのベア42.43は正常なので、
通常通り各ベア内のプロセッサ間で交互に動作監視信号
S1の送受を行っている。The multiprocessor system in this embodiment is configured such that a management processor 21 and general processors 318 to 36 are connected by a ring-shaped bus 11. Bear 41 for processor 31 and processor 32, Bear 42 for processor 33 and processor 34, Processor 35
and the processor 36 form a total of three bears, the bear 43,
As shown in FIG. 1(a), the two processors in each bear mutually transmit and receive operation monitoring signals S1. If, for example, a failure occurs in the processor 32 in this state, the operation monitoring signal S1 that has been sent and received within the bear 41 will stop, as shown in FIG.
1, and the processor 31 transmits a partner failure notification signal S2 to the management processor 21. The management processor 21 that received this partner failure notification signal S2
As shown in FIG. 3(c), during the period when the processor 32 is in failure, the processor 32 becomes a temporary bare partner of the processor 31.
During this period, the other two bears 42 and 43 are normal, so
As usual, the operation monitoring signal S1 is sent and received alternately between the processors in each bear.
次に第2図を参照して障害検出動作シーケンスについて
説明する。Next, the failure detection operation sequence will be explained with reference to FIG.
第2図ではプロセッサ31とプロセッサ32で形成され
たベア41における障害検出動作シーケンスを示してい
る0通常時にはプロセッサ31とプロセッサ32との間
でそれぞれS1送出タイミングアウトTOを契機に動作
監視信号S1を交互に送信し合い、次に相手プロセッサ
から受信される動作監視信号S1を監視している。この
状態でプロセッサ32に障害が発生すると、相手プロセ
ッサS1待ちタイミングアウタTlによりプロセッサ3
1はプロセッサ32の障害を検出する。この時プロセッ
サ31は相手障害通知信号S2を管理用プロセッサ21
に送信する。相手障害通知信号S2を受信した管理用プ
ロセッサ21は障害が発生したプロセッサ32の動作状
態を変更した後、プロセッサ31に対して動作監視信号
S1を送信する。プロセッサ31はこの動作監視信号S
1を受信し、それ以降は管理用プロセッサ21との間で
動作監視信号S1を交互に受信する。FIG. 2 shows the failure detection operation sequence in the bare 41 formed by the processor 31 and the processor 32. In normal operation, the operation monitoring signal S1 is sent between the processor 31 and the processor 32 at the timing of S1 transmission timing out. The operation monitoring signals S1 which are sent to each other alternately and then received from the partner processor are monitored. If a failure occurs in the processor 32 in this state, the processor 32 is
1 detects a failure of the processor 32. At this time, the processor 31 sends the other party failure notification signal S2 to the management processor 21.
Send to. The management processor 21 that has received the partner failure notification signal S2 changes the operating state of the processor 32 in which the failure has occurred, and then transmits the operation monitoring signal S1 to the processor 31. The processor 31 receives this operation monitoring signal S.
1, and thereafter receives the operation monitoring signal S1 alternately with the management processor 21.
以上説明したように、本発明のマルチプロセッサシステ
ムにおける障害検出方式は、システム内の管理用プロセ
ッサ以外の各プロセッサを2つずつベアにし、ベアごと
に独立に各ベア内の2つのプロセッサの間で交互に動作
監視信号を送受信して相手プロセッサの動作状態を監視
し、所定時間内に相手プロセッサからの動作監視信号を
受信できないときには相手プロセッサ障害発生と判断し
て管理用プロセッサに通知し、通知を受けた管理用プロ
セッセではそれ以降障害通知元のプロセッサとの間で動
作監視信号の送受信を実行することにより、システムの
正常動作状態では管理用プロセッサは障害検出に介入し
ないので、マルチプロセッサシステム内のプロセッサ数
の増加に伴う管理用プロセッサの負荷の増加および各プ
ロセッサの動作監視間隔の増大がなくなり、信頼性の高
い大規模なマルチプロセッサシステムが実現できる効果
がある。As explained above, the failure detection method in the multiprocessor system of the present invention makes each processor other than the management processor in the system bear two, and independently communicates between the two processors in each bear for each bear. It alternately sends and receives operation monitoring signals to monitor the operating state of the other processor, and if it cannot receive the operation monitoring signal from the other processor within a predetermined time, it determines that a failure has occurred in the other processor and notifies the management processor. The management processor that receives the fault then sends and receives operation monitoring signals to and from the processor that sent the fault notification, and the management processor does not intervene in fault detection under normal operating conditions of the system. This eliminates an increase in the load on the management processor and an increase in the operation monitoring interval of each processor due to an increase in the number of processors, and has the effect of realizing a highly reliable large-scale multiprocessor system.
第1図(a)、(b)、(C)は本発明のマルチプロセ
ッサシステムにおける障害検出方式の一実施例を示すシ
ステムブロック図、第2図は第1図における一連の障害
検出動作の一例を示すシーケンス図、第3図は従来のマ
ルチプロセッサシステムにおける障害検出方式の一例を
示すシステムブロック図である。
11・・・バス、21・・・管理用プロセッサ、31゜
32.33.34.35.36・・・プロセッサ、41
.42.43・・・ベア、Sl・・・動作監視信号、S
2・・・相手障害通知信号。
第 1 国 (α)FIGS. 1(a), (b), and (C) are system block diagrams showing an embodiment of a fault detection method in a multiprocessor system of the present invention, and FIG. 2 is an example of a series of fault detection operations in FIG. 1. FIG. 3 is a system block diagram showing an example of a failure detection method in a conventional multiprocessor system. 11... Bus, 21... Management processor, 31° 32.33.34.35.36... Processor, 41
.. 42.43...Bear, Sl...Operation monitoring signal, S
2: Other party failure notification signal. 1st country (α)
Claims (1)
する1個の管理用プロセッサとをバスによって結合して
分散処理を行うマルチプロセッサシステムにおいて、前
記複数のプロセッサの2個ずつを1つのペアとしこの1
つのペア内の前記2個のプロセッサ間で動作監視信号を
交互に送受信する送受信手段と、一方の前記プロセッサ
が前記動作監視信号を送信した後所定の時間内に相手プ
ロセッサからの前記動作監視信号を受信できないときは
前記一方のプロセッサは前記管理用プロセッサに前記相
手プロセッサが障害である旨を通知する通知手段とを備
えることを特徴とするマルチプロセッサシステムにおけ
る障害検出方式。In a multiprocessor system that performs distributed processing by connecting a plurality of processors and one management processor that manages the operating state of each processor via a bus, two of the plurality of processors each form one pair, and this one
transmitting/receiving means for alternately transmitting and receiving an operation monitoring signal between the two processors in a pair; A failure detection method in a multiprocessor system, characterized in that the one processor includes notification means for notifying the management processor that the other processor is in failure when reception is not possible.
Priority Applications (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP1037431A JPH02216577A (en) | 1989-02-16 | 1989-02-16 | Fault detecting system in multi-processor system |
Applications Claiming Priority (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP1037431A JPH02216577A (en) | 1989-02-16 | 1989-02-16 | Fault detecting system in multi-processor system |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| JPH02216577A true JPH02216577A (en) | 1990-08-29 |
Family
ID=12497327
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| JP1037431A Pending JPH02216577A (en) | 1989-02-16 | 1989-02-16 | Fault detecting system in multi-processor system |
Country Status (1)
| Country | Link |
|---|---|
| JP (1) | JPH02216577A (en) |
Cited By (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP2012137877A (en) * | 2010-12-24 | 2012-07-19 | Toshiba Corp | Secondary battery device, processor, monitoring program and vehicle |
-
1989
- 1989-02-16 JP JP1037431A patent/JPH02216577A/en active Pending
Cited By (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP2012137877A (en) * | 2010-12-24 | 2012-07-19 | Toshiba Corp | Secondary battery device, processor, monitoring program and vehicle |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US5712856A (en) | Method and apparatus for testing links between network switches | |
| JPH02216577A (en) | Fault detecting system in multi-processor system | |
| JPH01217666A (en) | Fault detecting system for multiprocessor system | |
| JPS62174838A (en) | Processor fault detection method in multiprocessor system | |
| JPH055418B2 (en) | ||
| JPH02279040A (en) | Fault detection system for multi-processor system | |
| JPH0348997A (en) | Monitoring system | |
| JPH0223740A (en) | Communication network system | |
| JPH0341838A (en) | Loop bus diagnostic system | |
| JPH02281368A (en) | Trouble detecting mechanism for controller | |
| JPS58211268A (en) | Multi-processor system | |
| JPH0253351A (en) | Subscriber communication system | |
| JPH03177191A (en) | Remote supervisory system | |
| JPH02112398A (en) | Preferential transmitting system in cyclic digital transmission | |
| JPH10207745A (en) | Method for confirming inter-processor existence | |
| JPH01303836A (en) | Data transmitting system | |
| JPS63228849A (en) | distributed transmission equipment | |
| JPH04167142A (en) | Fault detection system for information processor | |
| JPH02277397A (en) | Remote supervisory system | |
| JPS59127160A (en) | Fault detecting system | |
| JPH01205299A (en) | Fire alarm device | |
| JPH05257913A (en) | Health check system | |
| JPH02308638A (en) | Diagnostic equipment for duplex transmission line | |
| JPH0481903B2 (en) | ||
| JPS6316742A (en) | Communication control system |