JPH0423149A - Duplex data preserving device - Google Patents

Duplex data preserving device

Info

Publication number
JPH0423149A
JPH0423149A JP2128324A JP12832490A JPH0423149A JP H0423149 A JPH0423149 A JP H0423149A JP 2128324 A JP2128324 A JP 2128324A JP 12832490 A JP12832490 A JP 12832490A JP H0423149 A JPH0423149 A JP H0423149A
Authority
JP
Japan
Prior art keywords
area
input
cluster
clusters
shared memory
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Granted
Application number
JP2128324A
Other languages
Japanese (ja)
Other versions
JP2716571B2 (en
Inventor
Hitoshi Sugiyama
仁志 杉山
Kazunori Hiraishi
平石 壽徳
Takeshi Kumano
熊野 剛
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Fujitsu Ltd
Original Assignee
Fujitsu Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Fujitsu Ltd filed Critical Fujitsu Ltd
Priority to JP2128324A priority Critical patent/JP2716571B2/en
Priority to DE69124285T priority patent/DE69124285T2/en
Priority to EP91107879A priority patent/EP0457308B1/en
Publication of JPH0423149A publication Critical patent/JPH0423149A/en
Priority to US08/249,046 priority patent/US5548743A/en
Priority to US08/430,315 priority patent/US5568609A/en
Application granted granted Critical
Publication of JP2716571B2 publication Critical patent/JP2716571B2/en
Anticipated expiration legal-status Critical
Expired - Fee Related legal-status Critical Current

Links

Landscapes

  • Techniques For Improving Reliability Of Storages (AREA)
  • Hardware Redundancy (AREA)

Abstract

PURPOSE:To surely switch a duplex operating state to a single operating state at a high speed when one of two shared memories has a fault by providing an input/output path cutting mechanism which cuts physically the input/output paths which connect the shared memories and the clusters. CONSTITUTION:For instance, a cluster 12-1 separates a shared memory 10-1 due to the detection of a fault occurred in an area 1 of the memory 10-1. Under such conditions, an input/output path cut processing part 13 cuts physically the input/output paths P11 - P31 which connect the memory 10-1 and all clusters by means of an input/output patch cutting mechanism 11-1. When the clusters 12-2 and 12-3 have accesses to an area 1 of the memory 10-1, a program check state occurs owing to the physical separation. Thus the separation of the area 1 can be recognized with no performance of the inter-cluster communication. Thus it is possible to surely switch a duplex operating state to a single operating state at a high speed when one of the shared memories has a fault.

Description

【発明の詳細な説明】 (概要] プロセッサを存する複数のクラスタと、入出力パスを介
して各クラスタに接続される二重化された共用メモリと
を備えた複合計算機システムにおける二重化データ保全
装置に関し 一方の共用メモリで障害が発生した場合に、高速かつ確
実に二重化運用状態から片肺運用状態に移行できるよう
にすることを目的とし 共用メモリとクラスタとを結ぶ入出力パスを物理的に切
断する入出力パス切断機構を備えるとともに、各クラス
タは、二重化された共用メモリの各領域の有効/無効を
管理する制御情報部と、共用メモリ上の領域を切り離す
ときに、制御情報部にその領域が無効であることを設定
し、その領域と各クラスタとを結ぶ全入出力パスを切断
する処理を行う入出力パス切断処理部と、入出力パスの
切断により、アクセスが異常終了した場合に、制御情報
部にその領域が無効であることを設定し。
[Detailed Description of the Invention] (Summary) This invention relates to a redundant data maintenance device in a multi-purpose computer system comprising a plurality of clusters including processors and a redundant shared memory connected to each cluster via an input/output path. An input/output method that physically disconnects the input/output path connecting the shared memory and the cluster, with the aim of quickly and reliably transitioning from a duplex operation state to a single-lung operation state in the event of a shared memory failure. In addition to being equipped with a path disconnection mechanism, each cluster has a control information section that manages the validity/invalidity of each area of the duplicated shared memory, and a control information section that manages the validity/invalidity of each area of the shared memory. An input/output path disconnection processing unit that sets a certain value and disconnects all input/output paths connecting that area and each cluster, and a control information unit that handles abnormal access termination due to disconnection of input/output paths. to set that area to be invalid.

以後、その領域に対するアクセスを禁止するアクセス禁
止処理部とを備えるように構成する。
Thereafter, the device is configured to include an access prohibition processing unit that prohibits access to that area.

[産業上の利用分野〕 本発明は5二重化された共用メモリを持つ複合計算機シ
ステムにおいて、障害発生または使用者からの指定によ
り、一方のクラスタの切り離しを行うクラスタが、切り
離す側のメモリと全クラスタとの入出力パスを切断する
ことにより、二重化データの保全を行うようにした二重
化データ保全装置に関するものである。
[Industrial Application Field] The present invention is applicable to a compound computer system having a 5-duplex shared memory, in which when a failure occurs or a user specifies, one cluster is disconnected, and the memory of the disconnected side and all clusters are disconnected. The present invention relates to a duplex data maintenance device that maintains duplex data by cutting off an input/output path to the duplex data.

近年のコンピュータシステムでは、単一プロセッサの能
力の伸びが鈍化していることや、信頼性向上の強いニー
ズがあることなどの理由から、複数クラスタを共用メモ
リで接続した複合計算機システムを構築することが一般
的になりつつある。
In modern computer systems, the growth in the power of single processors is slowing down, and there is a strong need for improved reliability, so it is becoming increasingly important to build complex computer systems that connect multiple clusters with shared memory. is becoming common.

共用メモリには、複数のクラスタで共用される重要なデ
ータが置かれるため、信較性を高める目的で2共用メモ
リの二重化を行うことが多い。
Since the shared memory stores important data that is shared by multiple clusters, two shared memories are often duplicated for the purpose of increasing reliability.

このような二重化された共用メモリの一方に障害が発生
した場合や5使用者からの切り離しの指示があった場合
には、全クラスタとの整合性を保ちながら、高速に他の
共用メモリのみを使用する片肺運転に移行する必要があ
る。
If a failure occurs in one of these duplexed shared memories or if there is an instruction to disconnect from 5 users, only the other shared memory can be quickly used while maintaining consistency with all clusters. It is necessary to shift to single-lung driving.

〔従来の技術〕[Conventional technology]

第5図は従来技術の例を示す。 FIG. 5 shows an example of the prior art.

従来、二重化された共用メモリ10−1.10−2の一
方を切り離す場合に、第5図(イ)または第5図(ロ)
に示すような方式が用いられている。
Conventionally, when disconnecting one of the duplicated shared memories 10-1 and 10-2, the process shown in FIG. 5(a) or FIG. 5(b)
The method shown in is used.

第5図(イ)に示す方式では、二重化された共用メモリ
10−1.10−2へのアクセスを、共用メモリ内の制
御表50を用いて制御する。すなわち、制御表50内に
は、リードするとき、どちらの共用メモリから読み込む
のか、またライトするときには、どちらの共用メモリに
書き込むかなどの情報が格納され管理されている。
In the system shown in FIG. 5(a), access to the duplicated shared memories 10-1 and 10-2 is controlled using a control table 50 in the shared memory. That is, the control table 50 stores and manages information such as which shared memory to read from when reading, and which shared memory to write to when writing.

一方の共用メモリ、例えば共用メモリ10−1の領域I
で障害が発生した場合、その障害を検出したクラスタ1
2aは、その領域lが使用不可能状態である旨を、共用
メモリ内の制御表50に書き込む。
One shared memory, for example, area I of shared memory 10-1
If a failure occurs in Cluster 1 that detected the failure,
2a writes to the control table 50 in the shared memory that the area l is unusable.

残りのクラスタ12bは、共用メモリ10−110−2
のアクセスに先立ち、この制御表50を参照し、使用可
能状態であるかどうかを調べる。
The remaining cluster 12b has shared memory 10-110-2.
Before accessing the control table 50, it is checked whether the control table 50 is available or not.

共用メモリ10−1の領域lは、使用可能状態でないの
で、その領域1に対するアクセスを禁止し二重化された
他方の領域2へのアクセスのみを行う。
Since area 1 of shared memory 10-1 is not available, access to area 1 is prohibited and only access to the other duplicated area 2 is performed.

このように5共用メモリ内の制御表50を用いることに
より、全クラスタは、共用メモリ101へのアクセスを
禁止し、残った一方の共用メモリ10−2へのアクセス
だけを行う。
By using the control table 50 in the 5 shared memories in this manner, all clusters prohibit access to the shared memory 101 and only access the remaining shared memory 10-2.

第5図(ロ)に示す方式では、二重化された共用メモリ
10−1.10−2へのアクセスを、各クラスタ12a
、12bが固有に持つ制御表50150−2で制御し、
障害発生時には、クラスタ間の通信を行うことにより、
障害発生を通知する。
In the system shown in FIG. 5(b), each cluster 12a
, 12b has its own control table 50150-2,
In the event of a failure, by communicating between clusters,
Notify of failure occurrence.

各制御表50−1.50−2内には、リードするとき、
どちらの共用メモリから読み込むのか。
In each control table 50-1, 50-2, when reading,
Which shared memory should be read from?

またライトするときには、どちらの共用メモリに書き込
むかなどの情報が格納され管理されている。
Also, when writing, information such as which shared memory to write to is stored and managed.

第5図(イ)に示す前述の方式に比べて5クラスタ内の
ローカルメモリへのアクセスでよいので。
Compared to the above-described method shown in FIG. 5(a), it is only necessary to access local memory within five clusters.

1回のアクセス時間が短くなる。The time required for one access is shortened.

一方の共用メモリ、例えば共用メモリ10−1の領域1
で障害発生した場合、その障害を検出したクラスタ12
aは、自クラスタ内の制御表50−1に、その領域1は
使用不可能状態である旨を記録する。さらに、 SI域
1上で障害発生した旨を他クラスタ12bに通知する。
One shared memory, for example, area 1 of shared memory 10-1
If a failure occurs in the cluster 12 that detected the failure,
A records in the control table 50-1 in its own cluster that area 1 is unusable. Further, it notifies other clusters 12b that a failure has occurred on SI area 1.

通知を受けたクラスタ12bは1 クラスタ12aと同
様に、領域1が使用不可能状態である旨を制御表50−
2に記録する。
The cluster 12b that received the notification sends a notification to the control table 50-1 that area 1 is unavailable, similar to the cluster 12a.
Record in 2.

このとき、全クラスタでの制御表の情報更新が完了する
まで、全クラスタの通常の処理を停止する必要がある。
At this time, it is necessary to stop normal processing in all clusters until information updating of control tables in all clusters is completed.

以上の処理により、全クラスタを通じ一方の共用メモリ
10−1が使用不可能状態であることを認識し、全クラ
スタが残った他方の共用メモリ1O−2をアクセスする
ように制御する。
Through the above processing, all clusters recognize that one shared memory 10-1 is in an unusable state, and all clusters are controlled to access the remaining shared memory 10-2.

(発明が解決しようとする課題〕 第5図(イ)に示す方式には、以下の問題がある。(Problem to be solved by the invention) The method shown in FIG. 5(a) has the following problems.

(a)  各クラスタ12a、12bから共用メモリ1
0−1.10−2をアクセスする際に、−旦5共用メモ
リ内の制御表50を参照し、アクセスが可能であるかを
判断する処理が必要になり一回のアクセスに時間がかか
る。
(a) Shared memory 1 from each cluster 12a, 12b
When accessing 0-1 and 10-2, it is necessary to refer to the control table 50 in the shared memory 5 once to determine whether access is possible, and it takes time for one access.

(+))障害発生により使用不可能とした領域との入出
力パスが実際には存在するため、クラスタが誤動作した
場合に、その領域へアクセスする危険性がある。
(+)) Since there actually exists an input/output path to an area that has become unusable due to a failure, there is a risk of accessing that area if the cluster malfunctions.

(C)  共用メモリ内の制御表50の排他制御が煩雑
である。
(C) Exclusive control of the control table 50 in the shared memory is complicated.

(d)共用メモリ内の制御表域で障害が発生した場合、
または設定されている情報に矛盾が生じた場合5 シス
テムが誤動作する可能性がある。
(d) If a failure occurs in the control table area in shared memory,
Or if a contradiction occurs in the set information 5. The system may malfunction.

また5第5図(ロ)に示す方式には、以下の問題がある
Furthermore, the method shown in FIG. 5(b) has the following problems.

(a)  障害を検出したクラスタで、障害が発生した
旨を他のクラスタに通知する処理が必要であり。
(a) The cluster that has detected the failure must perform processing to notify other clusters that the failure has occurred.

処理が複雑となる。Processing becomes complicated.

(ハ)障害発生により使用不可能とした領域との入出力
パスが実際には存在するため、クラスタが誤動作した場
合に、その領域へアクセスする危険性がある。
(c) Since there actually exists an input/output path to an area that has become unusable due to a failure, there is a risk of accessing that area if the cluster malfunctions.

(C)  共用メモリ内の制御表域で障害が発生した場
合、または設定されている情報に矛盾が生じた場合、シ
ステムが誤動作する可能性がある。
(C) If a failure occurs in the control table area in the shared memory, or if an inconsistency occurs in the set information, the system may malfunction.

本発明は5以上のような通常のメモリアクセス時間の低
下、クラスタ間の通信に伴う処理の複雑化、クラスタ誤
動作による共用メモリのデータ破壊といった従来技術の
問題点を解決し、一方の共用メモリで障害が発生した場
合に、高速かつ確実に二重化運用状態から片肺運用状態
に移行できるようにすることを目的としている。
The present invention solves the problems of the conventional technology such as reduction in normal memory access time such as 5 or more, complication of processing due to communication between clusters, and data corruption in shared memory due to cluster malfunction. The purpose of this system is to enable a rapid and reliable transition from duplex operation to single-lung operation in the event of a failure.

〔課題を解決するための手段〕[Means to solve the problem]

第1図は本発明の原理説明図である。 FIG. 1 is a diagram explaining the principle of the present invention.

第1図において、10−1.10−2は共用メモリであ
って二重化されているもの、11−111−2は入出力
パス切断機構、12−1ないし12−3は各々プロセッ
サを備えたクラスタ、13は入出力パスの物理的な切断
処理を行う入出力パス切断処理部、14は切り乱された
共用メモリに対するアクセスを事前に禁止するアクセス
禁止処理部、15は二重化された共用メモリに対するア
クセス管理情報を持つ制御情報部、pH−P32は各ク
ラスタと共用メモリ間のデータ転送に用いられる入出力
パスを表す。
In Figure 1, 10-1 and 10-2 are shared memories that are duplicated, 11-111-2 is an input/output path disconnection mechanism, and 12-1 to 12-3 are clusters each equipped with a processor. , 13 is an input/output path disconnection processing unit that physically disconnects the input/output path, 14 is an access prohibition processing unit that prohibits access to the shredded shared memory in advance, and 15 is an access to the duplicated shared memory. A control information section containing management information, pH-P32, represents an input/output path used for data transfer between each cluster and the shared memory.

本発明では、二重化された共用メモリ10−110−2
と、各クラスタ12−1〜12−3とを結ぶ入出力パス
pH〜P32について、物理的に切断するハードウェア
による入出力パス切断機構I11.1m−2が設けられ
る。
In the present invention, the dual shared memory 10-110-2
A hardware-based input/output path disconnection mechanism I11.1m-2 is provided to physically disconnect the input/output paths pH to P32 connecting the clusters 12-1 to 12-3.

入出力パス切断機構11−1または11−2によって切
り離された共用メモリにアクセスすると入出力パス切断
機構11−1または11−2を含むハードウェアによる
制御部は1回復可能なエラーを発生させ、その事象をア
クセス元クラスタのソフトウェアに、プログラムチェン
ク割込みなどにより通知する。
When accessing the shared memory separated by the input/output path disconnection mechanism 11-1 or 11-2, the hardware control unit including the input/output path disconnection mechanism 11-1 or 11-2 generates a recoverable error; The event is notified to the software of the access source cluster using a program change interrupt or the like.

各クラスタ12−1〜12−3は7 自クラスタのロー
カルメモリ内に、共用メモリのどちらの領域をアクセス
するかを決定するための制御情報を管理する制御情報部
15を持つ。
Each of the clusters 12-1 to 12-3 has a control information unit 15 in its own local memory that manages control information for determining which area of the shared memory is to be accessed.

例えばクラスタ12−1が、共用メモリ1〇−1内の領
域1の障害発生検出により、または使用者からの切り離
しの指示により、共用メモリ101を切り離す場合、入
出力パス切断処理部13は、自クラスタ内の制御情報部
15に領域1について使用不可能状態を示す情報を設定
しく第1図■)、入出力パス切断機構11−1により、
共用メモリ10−1と全クラスタとを結ぶ入出力パスp
H,P21.P31を物理的に切断する。
For example, when the cluster 12-1 disconnects the shared memory 101 due to the detection of a failure in area 1 in the shared memory 10-1 or due to a disconnection instruction from the user, the input/output path disconnection processing unit 13 automatically disconnects the shared memory 101. Information indicating that area 1 is in an unusable state is set in the control information unit 15 in the cluster (Fig. 1), and the input/output path disconnection mechanism 11-1
Input/output path p connecting shared memory 10-1 and all clusters
H, P21. Physically disconnect P31.

他のクラスタ12−2.12〜3が、切断された入出力
パスl”21.P31を使って、共用メモU 10−1
内の領域1にアクセスすると、物理的に切り離されてい
るため1回復可能なプログラムチエツクが発生する。こ
こで1回復可能とは、エラーの割込み処理などをjテっ
だ後に1元の処理に復帰できることを意味する。
Other clusters 12-2.12 to 3 use the disconnected input/output path l"21.P31 to access the shared memo U10-1.
When area 1 is accessed, a program check that can be recovered by 1 occurs because it is physically separated. Here, 1-recovery means that it is possible to return to 1-original processing after performing error interrupt processing and the like.

クラスタ12−2および12−3のアクセス禁止処理部
14は1プログラムヂエツクを検出することにより、領
域lが使用不可能状態であることを認識し、以後2その
領域lへのアクセスを行わないようにするために、クラ
スタ内に存在する制御情報部15に、領域1が使用不可
能状態であることを設定する(第1図■、■)。
The access prohibition processing units 14 of clusters 12-2 and 12-3 recognize that area l is unusable by detecting 1 program check, and do not access area 1 from now on. In order to do this, it is set in the control information section 15 existing in the cluster that area 1 is in an unusable state ((2), (2) in FIG. 1).

結果的に、共用メモリ10−2上の領域2に対してだけ
、全クラスタがアクセスするようになり。
As a result, all clusters access only area 2 on the shared memory 10-2.

二重化データの保全が実現される。The integrity of duplicated data is achieved.

〔作用〕[Effect]

本発明では、切り離された領域1への入出力パスpHP
21.P31が、全クラスタを通して物理的に切断され
た状態となる。
In the present invention, the input/output path pHP to the separated area 1
21. P31 is physically disconnected throughout all clusters.

障害検出クラスタ12−1以外のクラスタ122 12
−3では5切り離された領域2に対してアクセスするが
1 その領域1と自クラスタの入出力パスP21.P3
1が存在しないため、プログラムチエツクが発生する。
Clusters 122 other than failure detection cluster 12-1 12
-3 accesses 5 separated area 2, but 1 input/output path P21 between that area 1 and the own cluster. P3
Since 1 does not exist, a program check occurs.

プログラムチエツクが発生することで1各クラスタは、
クラスタ間通信などを用いずに、領域1が切り離されて
いることを認識することができる。
When a program check occurs, each cluster is
It is possible to recognize that area 1 is separated without using inter-cluster communication or the like.

このように物理的に切断することにより、各クラスタは
、領域1に対してアクセスしようとしても アクセスが
不可能な状態になり、接続されている側の領域2だけを
各クラスタがアクセスするよ・)になる。
By physically disconnecting in this way, each cluster will be unable to access Area 1 even if it attempts to do so, and each cluster will only access Area 2 on the connected side. )become.

クラスタが誤動作した場合にも、切り離された側の領域
1をアクセスすることがなく、二重化データの保全が実
現できる。また、各入出力パスの物理的切断により、ク
ラスタ間での通信を使用ゼずに、他クラスタに障害発生
メモリが使用不可能状態である旨を知らせることができ
る。各クラスタごとに制御情報部15を設けているため
、共用メモリの二重化データに対するアクセスを高速に
行うことができる。
Even if the cluster malfunctions, the area 1 on the separated side will not be accessed, and the duplexed data can be maintained. Furthermore, by physically disconnecting each input/output path, it is possible to notify other clusters that the faulty memory is unusable without using communication between clusters. Since the control information unit 15 is provided for each cluster, the duplexed data in the shared memory can be accessed at high speed.

高速化できる理由は、ローカルメモリのアクセスは共用
メモリのアクセスよりも高速であるので制御情報を共用
メモリ内に設定した場合に比べて。
The reason for this is that accessing local memory is faster than accessing shared memory, compared to when control information is set in shared memory.

制御情報を高速に参照できること、制御情報に関する複
雑な排他側’<711等を行う必要がないごとなどであ
る。
The control information can be referred to at high speed, and there is no need to perform complicated exclusion side '<711 etc. regarding the control information.

〔実施例〕〔Example〕

第2図は本発明の一実施例による状態遷移の例。 FIG. 2 is an example of state transition according to an embodiment of the present invention.

第3図は本発明の一実施例処理フロー、第4図は本発明
の一実施例で用いる入出力パス切断機構の説明図を示す
FIG. 3 shows a processing flow of an embodiment of the present invention, and FIG. 4 shows an explanatory diagram of an input/output path cutting mechanism used in an embodiment of the present invention.

以下、説明を簡単にするために、クラスタが2つの場合
を例に説明するが、3以上の場合にも同様に適用できる
Hereinafter, in order to simplify the explanation, a case where there are two clusters will be explained as an example, but it can be similarly applied to a case where there are three or more clusters.

第2図(イ)は2通常の二重化された共用メ;Lす10
−1.10−2を介した複数クラスタ12a、12bに
よるシステムの運用状態を示している。領域1が主系で
あり1領域2が従系である。
Figure 2 (a) shows 2 normal duplex shared systems;
-1.10-2 shows the operational status of a system using multiple clusters 12a and 12b. Area 1 is the main system, and area 1 is the slave system.

主系はり一ト/ライ)・の対象となり、従系はライト・
アクセスのみ行われる。
The main system is subject to light
Access only.

第2図(ロ)に示すように、クラスタ12aが領域lに
アクセスし、障害を検出したとする。
As shown in FIG. 2(b), it is assumed that the cluster 12a accesses area l and detects a failure.

クラスタ12aは、領域lの障害を検出すると第2図(
ハ)に示すように、内部の制御情報中で主系を領域2に
書き換え、領域1については使用不可能状態とする。そ
れとともに、クラスタ12aから領域1への入出力パス
およびクラスタ12bから領域1への入出力パスを切断
する。この状態では、クラスタ12b中の制御情報は1
元のままである。
When cluster 12a detects a failure in area l,
As shown in c), the main system is rewritten to area 2 in the internal control information, and area 1 is rendered unusable. At the same time, the input/output path from cluster 12a to area 1 and the input/output path from cluster 12b to area 1 are cut off. In this state, the control information in cluster 12b is 1
It remains as it was.

クラスタ12bにおいて、二重化データに対するアクセ
ス要求のため、第2図(ニ)に示すように、領域1にア
クセスしたとする。
Assume that in cluster 12b, area 1 is accessed as shown in FIG. 2(d) in order to request access to duplicated data.

クラスタ】21〕が領域1にアクセスすると、領域1に
対する入出力パスが切断されているため。
When cluster ]21] accesses area 1, the input/output path to area 1 is disconnected.

第2図(ボ)に示すように、プログラムチエツクが発生
する。
As shown in FIG. 2 (BO), a program check occurs.

クラスタ12bは、プログラムチエツクが発生すると、
領域1が切り離されていることを認識し内部の制御情報
中で、主系を領域2に書き換え領域1については使用不
可能状態とする。
When a program check occurs in the cluster 12b,
It recognizes that area 1 has been separated and rewrites the main system to area 2 in the internal control information, making area 1 unusable.

これにより、以後、領域1について全クラスタのアクセ
スが行われないようになり、領域2だけによる片肺運転
が行われるようになる。
As a result, from now on, all clusters will not be accessed for area 1, and single-lung operation will be performed using only area 2.

領域1を障害により切り離す場合の例について説明した
が1オペレータ等の指示により、領域Iを切り離す場合
も同様である。
Although the example in which area 1 is separated due to a failure has been described, the same applies to the case in which area I is separated by an instruction from an operator or the like.

処理の流れは1例えば第3図に示ず■〜■のようになる
The processing flow is as shown in (1) to (2), not shown in FIG. 3, for example.

■ 例えばクラスタ12aが二重化された一方の領域で
障害発生を検出したとする。
(2) For example, assume that the cluster 12a detects a failure in one of the redundant areas.

■ 領域lで障害が発生したか領域2で障害が発生した
かを判定する。
■ Determine whether a failure has occurred in area l or area 2.

■ 領域1で障害が発生した場合、内部制御表の更新を
行った後、領域1と全クラスタとの入出力パスを切断す
る。なお、領域2で障害が発生した場合には、領域2と
全クラスタとの入出力パスを切断する。
■ If a failure occurs in area 1, after updating the internal control table, disconnect the input/output paths between area 1 and all clusters. Note that if a failure occurs in area 2, the input/output paths between area 2 and all clusters are disconnected.

■ クラスタ12aが入出力パスの切断を行った後、障
害を知らない他のクラスタ12bが5領域アクセスを行
ったとする。
(2) Suppose that after the cluster 12a disconnects the input/output path, another cluster 12b, which is unaware of the failure, accesses 5 areas.

■ 一方の領域でプログラムチエツクが発生する。■ A program check occurs in one area.

■ 領域1のアクセスでプログラムチエツクが発生した
か、領域2のアクセスでプログラムチエツクが発生した
かを判定する。
(2) Determine whether a program check occurs when accessing area 1 or when accessing area 2.

■ 領域Iでプログラムチエツクが発生した場合内部の
制御表中で領域1を使用不可能状態にする。領域2でプ
ログラムチエツクが発生した場合には5領域2に対して
同様の処理を行う。
■ When a program check occurs in area I, area 1 is made unusable in the internal control table. If a program check occurs in area 2, similar processing is performed for five areas 2.

以上の処理により、一方の領域の切り離しが行われると
、以後、他方の領域だけによる片肺運転が行われること
になる。
Once one region is separated by the above processing, one-lung operation will be performed using only the other region.

第1図に示す入出力パス切断機構11−1.11−2は
、スイッチその他により、各クラスタから共用メモリに
対するアクセスを物理的に不可能にすることができるも
のであれば、どのようなイ〕のでもよい。
The input/output path disconnection mechanism 11-1 and 11-2 shown in FIG. ].

本実施例で用いている入出力パス切断機構等は第4図に
示すような構造になっている。
The input/output path cutting mechanism used in this embodiment has a structure as shown in FIG.

共用メモリ10−1.10−2は、第4図に示すように
、データを格納する記憶機構40と、共用メモリ全体の
制御またはクラスタ12との通信を司る制′42j機構
41に分かれている。
As shown in FIG. 4, the shared memory 10-1 and 10-2 are divided into a storage mechanism 40 for storing data and a control mechanism 41 for controlling the entire shared memory or communicating with the cluster 12. .

各クラスタ12との通信は、制御機構41にあるボート
43を介して行われる。各クラスタ12ごとに、1つの
ボート43が固定的に割り当てられる。
Communication with each cluster 12 is performed via a boat 43 in the control mechanism 41. One boat 43 is fixedly assigned to each cluster 12.

ボート43には、有効と無効の2つの状態が存在し、そ
の状態制御のために、各ボート43と1対1に対応する
1ポートにつき1ビットの制御メモリ42が、制御機構
41内に存在する。この制御メモリ42は、記憶機構4
0内のメモリとは別のものである。
The boat 43 has two states, valid and invalid, and in order to control the state, a control memory 42 of 1 bit per port exists in the control mechanism 41 and corresponds to each boat 43 on a one-to-one basis. do. This control memory 42 includes the storage mechanism 4
This is different from the memory in 0.

この制御メモリ42のビットが“″1パのとき対応する
ボート43の状態は有効であり、そのホト43に割り当
てられているクラスタI2は共用メモリとの通信が可能
である。この状態ではクラスタ12が共用メモリとのデ
ータ転送を行なえるだけでなく、制御メモリ全体の内容
の変更も可能である。すなわら、有効状態のボート43
につながっているクラスタ12は、他のボート43の状
態を変更することも可能である。
When the bit of the control memory 42 is "1", the state of the corresponding port 43 is valid, and the cluster I2 assigned to that port 43 can communicate with the shared memory. In this state, the cluster 12 can not only transfer data to and from the shared memory, but also change the contents of the entire control memory. That is, the boat 43 in the active state
The cluster 12 connected to the boat 43 can also change the status of other boats 43.

無効状態のボート43につながっているクラスタ12は
、データ転送を行えないばかりでなく。
A cluster 12 connected to an invalid boat 43 is not only unable to transfer data.

制御メモリの変更も行えない。Control memory cannot be changed either.

以上の機構により、障害が発生した共用メモリに関する
全クラスタ12のボート43を無効状態とすることで、
各クラスタ12からデータ転送を行うことを物理的に抑
止することができる。これにより、クラスタ12の誤動
作によるデータの破壊を避けることも可能となる。
By using the above mechanism, the ports 43 of all clusters 12 related to the shared memory where the failure has occurred are made invalid.
Data transfer from each cluster 12 can be physically inhibited. This also makes it possible to avoid data destruction due to malfunction of the cluster 12.

[発明の効果] 本発明による効果は以下のとおりである。[Effect of the invention] The effects of the present invention are as follows.

(a)  データを保証した状態で、共用メモリの切り
離し7を、高速かつ簡単に実現できる。クラスタ間の通
信は不要である。
(a) The shared memory can be separated 7 quickly and easily while data is guaranteed. No communication between clusters is required.

α))切り離された共用メモリは、単にラフ1−ウェア
によりアクセスを禁止するだけでなく、物理的にもアク
セスできない状態になるので、システムの誤動作の危険
がない。
α)) Access to the separated shared memory is not only prohibited by rough hardware, but also physically inaccessible, so there is no risk of system malfunction.

(0)共用メモリ内に使用可否を管理する情報を持つ必
要がないので、共用メモリの二重化データに対するアク
セスを高速に行うことができる。
(0) Since there is no need to have information for managing availability in the shared memory, it is possible to access duplexed data in the shared memory at high speed.

(d)  共用メモリの切り離しに関する処理について
(d) Processing related to separation of shared memory.

クラスタごとの独立性が強いため、信顛性が高High reliability due to strong independence of each cluster

【図面の簡単な説明】[Brief explanation of the drawing]

第1図は本発明の原理説明図。 第2図は本発明の一実施例による状態遷移の側梁3図は
本発明の一実施例処理フロー 第4図は本発明の一実施例で用いる入出力パス切断機構
の説明図。 第5図は従来技術の例を示す。 図中、10−1.1(1−2は共用メモリ、111.1
1−2は入出力パス切断機構、12−1〜12−3はク
ラスタ、13は入出力パス切断処理部、14はアクセス
禁止処理部、15は制御情報部、P11〜P32は人出
力パスを表す。
FIG. 1 is a diagram explaining the principle of the present invention. FIG. 2 is a diagram showing a side beam of a state transition according to an embodiment of the present invention. FIG. 4 is an explanatory diagram of an input/output path cutting mechanism used in an embodiment of the present invention. FIG. 5 shows an example of the prior art. In the figure, 10-1.1 (1-2 is shared memory, 111.1
1-2 is an input/output path disconnection mechanism, 12-1 to 12-3 are clusters, 13 is an input/output path disconnection processing unit, 14 is an access prohibition processing unit, 15 is a control information unit, and P11 to P32 are human output paths. represent.

Claims (1)

【特許請求の範囲】 プロセッサを有する複数のクラスタ(12−1、12−
2、・・・)と、入出力パス(P11、P12、・・・
)を介して各クラスタに接続される二重化された共用メ
モリ(10−1、10−2)とを備えた複合計算機シス
テムにおいて、共用メモリとクラスタとを結ぶ入出力パ
スを物理的に切断する入出力パス切断機構(11−1、
11−2)を備えるとともに、 前記各クラスタは、 二重化された共用メモリの各領域の有効/無効を管理す
る制御情報部(15)と、 共用メモリ上の領域を、障害または外部からの指定によ
り切り離すときに、前記制御情報部にその領域が無効で
あることを設定し、前記入出力パス切断機構により、そ
の領域と各クラスタとを結ぶ全入出力パスを切断する処
理を行う入出力パス切断処理部(13)と、 切断された入出力パスを使用することにより、アクセス
が異常終了した場合に、前記制御情報部にその領域が無
効であることを設定し、以後、その領域に対するアクセ
スを禁止するアクセス禁止処理部(14)とを備えたこ
とを特徴とする二重化データ保全装置。
[Claims] A plurality of clusters (12-1, 12-
2,...) and input/output paths (P11, P12,...
) In a multicomputer system equipped with duplicated shared memories (10-1, 10-2) connected to each cluster via Output path cutting mechanism (11-1,
11-2), and each cluster includes: a control information unit (15) that manages the validity/invalidity of each area of the duplicated shared memory; When disconnecting, the input/output path disconnection process sets that the area is invalid in the control information section and uses the input/output path disconnection mechanism to disconnect all input/output paths connecting the area and each cluster. By using the processing unit (13) and the disconnected input/output path, when an access terminates abnormally, the area is set to be invalid in the control information unit, and access to that area is disabled from now on. 1. A duplex data protection device comprising: an access prohibition processing unit (14) that prohibits access.
JP2128324A 1990-05-18 1990-05-18 Redundant data security device Expired - Fee Related JP2716571B2 (en)

Priority Applications (5)

Application Number Priority Date Filing Date Title
JP2128324A JP2716571B2 (en) 1990-05-18 1990-05-18 Redundant data security device
DE69124285T DE69124285T2 (en) 1990-05-18 1991-05-15 Data processing system with an input / output path separation mechanism and method for controlling the data processing system
EP91107879A EP0457308B1 (en) 1990-05-18 1991-05-15 Data processing system having an input/output path disconnecting mechanism and method for controlling the data processing system
US08/249,046 US5548743A (en) 1990-05-18 1994-05-24 Data processing system with duplex common memory having physical and logical path disconnection upon failure
US08/430,315 US5568609A (en) 1990-05-18 1995-04-28 Data processing system with path disconnection and memory access failure recognition

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
JP2128324A JP2716571B2 (en) 1990-05-18 1990-05-18 Redundant data security device

Publications (2)

Publication Number Publication Date
JPH0423149A true JPH0423149A (en) 1992-01-27
JP2716571B2 JP2716571B2 (en) 1998-02-18

Family

ID=14981964

Family Applications (1)

Application Number Title Priority Date Filing Date
JP2128324A Expired - Fee Related JP2716571B2 (en) 1990-05-18 1990-05-18 Redundant data security device

Country Status (1)

Country Link
JP (1) JP2716571B2 (en)

Cited By (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
WO2011121666A1 (en) 2010-03-31 2011-10-06 富士通株式会社 Multi-cluster system

Cited By (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
WO2011121666A1 (en) 2010-03-31 2011-10-06 富士通株式会社 Multi-cluster system

Also Published As

Publication number Publication date
JP2716571B2 (en) 1998-02-18

Similar Documents

Publication Publication Date Title
US6266785B1 (en) File system filter driver apparatus and method
US5548743A (en) Data processing system with duplex common memory having physical and logical path disconnection upon failure
US7318138B1 (en) Preventing undesired trespass in storage arrays
US8498967B1 (en) Two-node high availability cluster storage solution using an intelligent initiator to avoid split brain syndrome
JP2994070B2 (en) Method and apparatus for determining the state of a pair of mirrored data storage units in a data processing system
US6732231B1 (en) System and method for management of mirrored storage devices storing device serial numbers
JP4475598B2 (en) Storage system and storage system control method
US7650467B2 (en) Coordination of multiprocessor operations with shared resources
JPH0420493B2 (en)
US7774640B2 (en) Disk array apparatus
US6785840B1 (en) Call processor system and methods
JP4132322B2 (en) Storage control device and control method thereof
JP3987241B2 (en) Inter-system information communication system
JP2000181887A5 (en)
US6408399B1 (en) High reliability multiple processing and control system utilizing shared components
US6490662B1 (en) System and method for enhancing the reliability of a computer system by combining a cache sync-flush engine with a replicated memory module
CN112445652B (en) Remote copy system
US7472221B1 (en) Mirrored memory
US7302526B1 (en) Handling memory faults for mirrored memory
JP2716571B2 (en) Redundant data security device
JPH0460750A (en) Cluster stop device
JP2006260141A (en) Storage system control method, storage system, storage control device, storage system control program, and information processing system
JPH02294723A (en) Duplex control method for auxiliary memory device
JP3015537B2 (en) Redundant computer system
JP3015538B2 (en) Redundant computer system

Legal Events

Date Code Title Description
FPAY Renewal fee payment (event date is renewal date of database)

Free format text: PAYMENT UNTIL: 20081107

Year of fee payment: 11

FPAY Renewal fee payment (event date is renewal date of database)

Free format text: PAYMENT UNTIL: 20081107

Year of fee payment: 11

FPAY Renewal fee payment (event date is renewal date of database)

Free format text: PAYMENT UNTIL: 20091107

Year of fee payment: 12

LAPS Cancellation because of no payment of annual fees