JPH02216577A - Fault detecting system in multi-processor system - Google Patents

Fault detecting system in multi-processor system

Info

Publication number
JPH02216577A
JPH02216577A JP1037431A JP3743189A JPH02216577A JP H02216577 A JPH02216577 A JP H02216577A JP 1037431 A JP1037431 A JP 1037431A JP 3743189 A JP3743189 A JP 3743189A JP H02216577 A JPH02216577 A JP H02216577A
Authority
JP
Japan
Prior art keywords
processor
operation monitoring
processors
monitoring signal
pair
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
JP1037431A
Other languages
Japanese (ja)
Inventor
Kazuo Nishidai
西大 和男
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
NEC Corp
Original Assignee
NEC Corp
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by NEC Corp filed Critical NEC Corp
Priority to JP1037431A priority Critical patent/JPH02216577A/en
Publication of JPH02216577A publication Critical patent/JPH02216577A/en
Pending legal-status Critical Current

Links

Landscapes

  • Debugging And Monitoring (AREA)
  • Multi Processors (AREA)

Abstract

PURPOSE:To decrease an increase of a load of a management use processor and an operation monitoring interval of each processor and to enhance the reliability by forming two pieces each of each processor except the management use processor to a pair, and transmitting and receiving alternately an operation monitoring signal between two processors. CONSTITUTION:Between two processors in each pair, an operation monitoring signal S1 is transmitted and received each other. For instance, when a fault is occurred in a processor 32, the operation monitoring signal S1 transmitted and received in a pair 41 stops, it is detected by the other processor 31, and the processor 31 transmits the other processor fault informing signal to a management use processor 21. The management use processor 21 executes alternately transmission and reception of the operation monitoring signal S1 to and from the processor 31 in a period of a fault of the processor 32. In this regard, during this time, as well, other two pairs 42, 43 execute a normal operation, and transmit and receive alternately the operation monitoring signal S1 between the processors in each pair.

Description

【発明の詳細な説明】 〔産業上の利用分野〕 本発明はマルチプロセッサシステムにおける障害検出方
式に関する。
DETAILED DESCRIPTION OF THE INVENTION [Field of Industrial Application] The present invention relates to a failure detection method in a multiprocessor system.

(従来の技術〕 従来、この種のマルチプロセッサシステムにおける障害
検出方式は、システム内の各プロセッサの動作状態を管
理する管理用プロセッサからすべてのプロセッサに対し
て動作監視信号を順次送信して、返送されてくる応答信
号を監視し、所定の時間内に応答信号が受信できなけれ
ばそのプロセッサは障害であると判断することによりシ
ステム内の全プロセッサの障害発生を検出していた。
(Prior Art) Conventionally, a failure detection method in this type of multiprocessor system involves sequentially transmitting operation monitoring signals to all processors from a management processor that manages the operating status of each processor in the system, and then returning the signals. The occurrence of a failure in all processors in the system is detected by monitoring incoming response signals and determining that the processor is at fault if the response signal is not received within a predetermined time.

第3図は従来のマルチプロセッサシステムにおける障害
検出方式の一例を示すシステムブロック図である。
FIG. 3 is a system block diagram showing an example of a failure detection method in a conventional multiprocessor system.

本例のマルチプロセッサシステムは1個の管理用プロセ
ッサ21と3個のプロセッサ31.〜33をリング形の
バス11で結合した形で構成されている、管理用プロセ
ッサ21はプロセッサ31.〜33の動作状態を管理し
ており、プロセッサ31に対して動作監視信号Slを送
信して所定時間内にその応答信号S3が受信されること
を監視し、その結果をプロセッサ31の動作状態として
管理し、引続きプロセッサ32.プロセッサ33に対し
て順次同様の手順を繰り返すことにより各プロセッサ3
1.〜33の障害検出を行っている。
The multiprocessor system of this example includes one management processor 21 and three processors 31. The management processor 21 is composed of processors 31 . ~ 33, transmits an operation monitoring signal Sl to the processor 31, monitors that the response signal S3 is received within a predetermined time, and uses the result as the operating status of the processor 31. The processor 32. By sequentially repeating the same procedure for the processors 33, each processor 3
1. ~33 failures have been detected.

〔発明が解決しようとする課題〕[Problem to be solved by the invention]

上述した従来のマルチプロセッサシステムにおける障害
検出方式は、管理用プロセッサから他の全プロセッサに
対して順次動作監視信号を送信してその応答信号を監視
する方式のため、システム内のプロセッサ数が多くなる
と、動作監視信号の送出数が増えて管理用プロセッサの
負荷が増大するという欠点がある。一方、個々のプロセ
ッサから見ると動作監視信号を受信する間隔が大きくな
って、障害検出までの時間が長くなるという欠点がある
The fault detection method in the conventional multiprocessor system described above is a method in which the management processor sequentially sends an operation monitoring signal to all other processors and monitors the response signals, so when the number of processors in the system increases, However, the disadvantage is that the number of operation monitoring signals to be sent increases, increasing the load on the management processor. On the other hand, from the point of view of individual processors, there is a drawback that the interval at which the operation monitoring signal is received becomes long, and the time until a failure is detected becomes long.

〔課題を解決するための手段〕[Means to solve the problem]

本発明のマルチプロセッサシステムにおける障害検出方
式は、複数のプロセッサとこの各プロセッサの動作状態
を管理する1個の管理用プロセッサとをバスによって結
合して分散処理を行うマルチプロセッサシステムにおい
て、前記複数のプロセッサの2個ずつを1つのベアとし
この1つのベア内の前記2個のプロセッサ間で動作監視
信号を交互に送受信する送受信手段と、一方の前記プロ
セッサが前記動作監視信号を送信した後所定の時間内に
相手プロセッサからの前記動作監視信号を受信できない
ときは前記一方のプロセッサは前記管理用プロセッサに
前記相手プロセッサが障害である旨を通知する通知手段
を備えることを特徴とする。
A failure detection method in a multiprocessor system according to the present invention is a multiprocessor system in which a plurality of processors and one management processor that manages the operating state of each processor are connected via a bus to perform distributed processing. transmitting/receiving means for alternately transmitting and receiving operation monitoring signals between the two processors in one bear, each of which has two processors; The one processor is characterized in that, when the operation monitoring signal cannot be received from the partner processor within a time, the one processor includes notification means for notifying the management processor that the partner processor is in trouble.

〔実施例〕〔Example〕

次に本発明について第1図、第2図を参照して説明する
Next, the present invention will be explained with reference to FIGS. 1 and 2.

第1図は本発明のマルチプロセッサシステムにおける障
害検出方式の一実施例を示すシステムブロック図であり
、(a)は通常状態での障害検出動作を、(b)は障害
発生時の動作を、(c)は相手プロセッサが障害中で管
理用プロセッサによって代行される障害検出動作を示し
ている。また第2図は第1図における一連の障害検出動
作のシーケンスの一例を示すシーケンス図である。
FIG. 1 is a system block diagram showing an embodiment of the fault detection method in the multiprocessor system of the present invention, in which (a) shows the fault detection operation in a normal state, and (b) shows the operation when a fault occurs. (c) shows a failure detection operation performed by the management processor when the other processor is in failure. Further, FIG. 2 is a sequence diagram showing an example of a sequence of a series of failure detection operations in FIG. 1.

始めに第1図を参照して本実施例の障害検出動作につい
て説明する。
First, the fault detection operation of this embodiment will be explained with reference to FIG.

本実施例におけるマルチプロセッサシステムは、管理用
プロセッサ21と一般のプロセッサ318〜36がリン
グ形のバス11によって結合された形で構成されている
。プロセッサ31とプロセッサ32でベア41、プロセ
ッサ33とプロセッサ34でベア42、プロセッサ35
とプロセッサ36でベア43の計3つのベアを形成し、
第1図(a)に示すように各ベア内の2つのプロセッサ
間で互いに動作監視信号S1を送受信している。この状
態で例えばプロセッサ32に障害が発生すると、第1図
(b)に示すようにベア41の中で送受信されていた動
作監視信号S1はストップし、それを相手プロセッサ3
1によって検出し、プロセッサ31は相手障害通知信号
S2を管理用プロセッサ21に送信する。この相手障害
通知信号S2を受信した管理用プロセッサ21は、第1
図(c)に示すようにプロセッサ32が障害中の期間は
プロセッサ31の仮のベアの相手となってプロセッサ3
1との間で動作監視信号S1の送受信を交互に行う、な
おこの間も、他の2つのベア42.43は正常なので、
通常通り各ベア内のプロセッサ間で交互に動作監視信号
S1の送受を行っている。
The multiprocessor system in this embodiment is configured such that a management processor 21 and general processors 318 to 36 are connected by a ring-shaped bus 11. Bear 41 for processor 31 and processor 32, Bear 42 for processor 33 and processor 34, Processor 35
and the processor 36 form a total of three bears, the bear 43,
As shown in FIG. 1(a), the two processors in each bear mutually transmit and receive operation monitoring signals S1. If, for example, a failure occurs in the processor 32 in this state, the operation monitoring signal S1 that has been sent and received within the bear 41 will stop, as shown in FIG.
1, and the processor 31 transmits a partner failure notification signal S2 to the management processor 21. The management processor 21 that received this partner failure notification signal S2
As shown in FIG. 3(c), during the period when the processor 32 is in failure, the processor 32 becomes a temporary bare partner of the processor 31.
During this period, the other two bears 42 and 43 are normal, so
As usual, the operation monitoring signal S1 is sent and received alternately between the processors in each bear.

次に第2図を参照して障害検出動作シーケンスについて
説明する。
Next, the failure detection operation sequence will be explained with reference to FIG.

第2図ではプロセッサ31とプロセッサ32で形成され
たベア41における障害検出動作シーケンスを示してい
る0通常時にはプロセッサ31とプロセッサ32との間
でそれぞれS1送出タイミングアウトTOを契機に動作
監視信号S1を交互に送信し合い、次に相手プロセッサ
から受信される動作監視信号S1を監視している。この
状態でプロセッサ32に障害が発生すると、相手プロセ
ッサS1待ちタイミングアウタTlによりプロセッサ3
1はプロセッサ32の障害を検出する。この時プロセッ
サ31は相手障害通知信号S2を管理用プロセッサ21
に送信する。相手障害通知信号S2を受信した管理用プ
ロセッサ21は障害が発生したプロセッサ32の動作状
態を変更した後、プロセッサ31に対して動作監視信号
S1を送信する。プロセッサ31はこの動作監視信号S
1を受信し、それ以降は管理用プロセッサ21との間で
動作監視信号S1を交互に受信する。
FIG. 2 shows the failure detection operation sequence in the bare 41 formed by the processor 31 and the processor 32. In normal operation, the operation monitoring signal S1 is sent between the processor 31 and the processor 32 at the timing of S1 transmission timing out. The operation monitoring signals S1 which are sent to each other alternately and then received from the partner processor are monitored. If a failure occurs in the processor 32 in this state, the processor 32 is
1 detects a failure of the processor 32. At this time, the processor 31 sends the other party failure notification signal S2 to the management processor 21.
Send to. The management processor 21 that has received the partner failure notification signal S2 changes the operating state of the processor 32 in which the failure has occurred, and then transmits the operation monitoring signal S1 to the processor 31. The processor 31 receives this operation monitoring signal S.
1, and thereafter receives the operation monitoring signal S1 alternately with the management processor 21.

〔発明の効果〕〔Effect of the invention〕

以上説明したように、本発明のマルチプロセッサシステ
ムにおける障害検出方式は、システム内の管理用プロセ
ッサ以外の各プロセッサを2つずつベアにし、ベアごと
に独立に各ベア内の2つのプロセッサの間で交互に動作
監視信号を送受信して相手プロセッサの動作状態を監視
し、所定時間内に相手プロセッサからの動作監視信号を
受信できないときには相手プロセッサ障害発生と判断し
て管理用プロセッサに通知し、通知を受けた管理用プロ
セッセではそれ以降障害通知元のプロセッサとの間で動
作監視信号の送受信を実行することにより、システムの
正常動作状態では管理用プロセッサは障害検出に介入し
ないので、マルチプロセッサシステム内のプロセッサ数
の増加に伴う管理用プロセッサの負荷の増加および各プ
ロセッサの動作監視間隔の増大がなくなり、信頼性の高
い大規模なマルチプロセッサシステムが実現できる効果
がある。
As explained above, the failure detection method in the multiprocessor system of the present invention makes each processor other than the management processor in the system bear two, and independently communicates between the two processors in each bear for each bear. It alternately sends and receives operation monitoring signals to monitor the operating state of the other processor, and if it cannot receive the operation monitoring signal from the other processor within a predetermined time, it determines that a failure has occurred in the other processor and notifies the management processor. The management processor that receives the fault then sends and receives operation monitoring signals to and from the processor that sent the fault notification, and the management processor does not intervene in fault detection under normal operating conditions of the system. This eliminates an increase in the load on the management processor and an increase in the operation monitoring interval of each processor due to an increase in the number of processors, and has the effect of realizing a highly reliable large-scale multiprocessor system.

【図面の簡単な説明】[Brief explanation of the drawing]

第1図(a)、(b)、(C)は本発明のマルチプロセ
ッサシステムにおける障害検出方式の一実施例を示すシ
ステムブロック図、第2図は第1図における一連の障害
検出動作の一例を示すシーケンス図、第3図は従来のマ
ルチプロセッサシステムにおける障害検出方式の一例を
示すシステムブロック図である。 11・・・バス、21・・・管理用プロセッサ、31゜
32.33.34.35.36・・・プロセッサ、41
.42.43・・・ベア、Sl・・・動作監視信号、S
2・・・相手障害通知信号。 第 1 国 (α)
FIGS. 1(a), (b), and (C) are system block diagrams showing an embodiment of a fault detection method in a multiprocessor system of the present invention, and FIG. 2 is an example of a series of fault detection operations in FIG. 1. FIG. 3 is a system block diagram showing an example of a failure detection method in a conventional multiprocessor system. 11... Bus, 21... Management processor, 31° 32.33.34.35.36... Processor, 41
.. 42.43...Bear, Sl...Operation monitoring signal, S
2: Other party failure notification signal. 1st country (α)

Claims (1)

【特許請求の範囲】[Claims] 複数のプロセッサとこの各プロセッサの動作状態を管理
する1個の管理用プロセッサとをバスによって結合して
分散処理を行うマルチプロセッサシステムにおいて、前
記複数のプロセッサの2個ずつを1つのペアとしこの1
つのペア内の前記2個のプロセッサ間で動作監視信号を
交互に送受信する送受信手段と、一方の前記プロセッサ
が前記動作監視信号を送信した後所定の時間内に相手プ
ロセッサからの前記動作監視信号を受信できないときは
前記一方のプロセッサは前記管理用プロセッサに前記相
手プロセッサが障害である旨を通知する通知手段とを備
えることを特徴とするマルチプロセッサシステムにおけ
る障害検出方式。
In a multiprocessor system that performs distributed processing by connecting a plurality of processors and one management processor that manages the operating state of each processor via a bus, two of the plurality of processors each form one pair, and this one
transmitting/receiving means for alternately transmitting and receiving an operation monitoring signal between the two processors in a pair; A failure detection method in a multiprocessor system, characterized in that the one processor includes notification means for notifying the management processor that the other processor is in failure when reception is not possible.
JP1037431A 1989-02-16 1989-02-16 Fault detecting system in multi-processor system Pending JPH02216577A (en)

Priority Applications (1)

Application Number Priority Date Filing Date Title
JP1037431A JPH02216577A (en) 1989-02-16 1989-02-16 Fault detecting system in multi-processor system

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
JP1037431A JPH02216577A (en) 1989-02-16 1989-02-16 Fault detecting system in multi-processor system

Publications (1)

Publication Number Publication Date
JPH02216577A true JPH02216577A (en) 1990-08-29

Family

ID=12497327

Family Applications (1)

Application Number Title Priority Date Filing Date
JP1037431A Pending JPH02216577A (en) 1989-02-16 1989-02-16 Fault detecting system in multi-processor system

Country Status (1)

Country Link
JP (1) JPH02216577A (en)

Cited By (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2012137877A (en) * 2010-12-24 2012-07-19 Toshiba Corp Secondary battery device, processor, monitoring program and vehicle

Cited By (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2012137877A (en) * 2010-12-24 2012-07-19 Toshiba Corp Secondary battery device, processor, monitoring program and vehicle

Similar Documents

Publication Publication Date Title
US5712856A (en) Method and apparatus for testing links between network switches
JPH02216577A (en) Fault detecting system in multi-processor system
JPH01217666A (en) Fault detecting system for multiprocessor system
JPS62174838A (en) Processor fault detection method in multiprocessor system
JPH055418B2 (en)
JPH02279040A (en) Fault detection system for multi-processor system
JPH0348997A (en) Monitoring system
JPH0223740A (en) Communication network system
JPH0341838A (en) Loop bus diagnostic system
JPH02281368A (en) Trouble detecting mechanism for controller
JPS58211268A (en) Multi-processor system
JPH0253351A (en) Subscriber communication system
JPH03177191A (en) Remote supervisory system
JPH02112398A (en) Preferential transmitting system in cyclic digital transmission
JPH10207745A (en) Method for confirming inter-processor existence
JPH01303836A (en) Data transmitting system
JPS63228849A (en) distributed transmission equipment
JPH04167142A (en) Fault detection system for information processor
JPH02277397A (en) Remote supervisory system
JPS59127160A (en) Fault detecting system
JPH01205299A (en) Fire alarm device
JPH05257913A (en) Health check system
JPH02308638A (en) Diagnostic equipment for duplex transmission line
JPH0481903B2 (en)
JPS6316742A (en) Communication control system