JPS5949062A - Information storing system - Google Patents

Information storing system

Info

Publication number
JPS5949062A
JPS5949062A JP57158125A JP15812582A JPS5949062A JP S5949062 A JPS5949062 A JP S5949062A JP 57158125 A JP57158125 A JP 57158125A JP 15812582 A JP15812582 A JP 15812582A JP S5949062 A JPS5949062 A JP S5949062A
Authority
JP
Japan
Prior art keywords
storage
matching
information
text
document
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
JP57158125A
Other languages
Japanese (ja)
Inventor
Hiromichi Fujisawa
藤沢 浩道
Masaaki Kurosu
黒須 正明
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Hitachi Ltd
Original Assignee
Hitachi Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Hitachi Ltd filed Critical Hitachi Ltd
Priority to JP57158125A priority Critical patent/JPS5949062A/en
Publication of JPS5949062A publication Critical patent/JPS5949062A/en
Pending legal-status Critical Current

Links

Classifications

    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N1/00Scanning, transmission or reproduction of documents or the like, e.g. facsimile transmission; Details thereof

Landscapes

  • Engineering & Computer Science (AREA)
  • Multimedia (AREA)
  • Signal Processing (AREA)
  • Information Retrieval, Db Structures And Fs Structures Therefor (AREA)
  • Storing Facsimile Image Data (AREA)

Abstract

(57)【要約】本公報は電子出願前の出願データであるた
め要約のデータは記録されません。
(57) [Summary] This bulletin contains application data before electronic filing, so abstract data is not recorded.

Description

【発明の詳細な説明】 〔発明の利用分野〕 本発明は情報記憶装置の記憶方式に係り、特に文書ファ
イルに適した記憶方式に関する。
DETAILED DESCRIPTION OF THE INVENTION [Field of Application of the Invention] The present invention relates to a storage system for an information storage device, and particularly to a storage system suitable for document files.

〔従来技術〕[Prior art]

近年、「オフィスオートメーションJ ’ktJ+に英
文や日本文のテキストヶ扱うことが増えてきた。
In recent years, ``Office Automation J'ktJ+'' has increasingly included English and Japanese texts.

これらは一般に文書処理と呼ばれている。従来は文書処
理の中心課題はワードプロセシングにあったが、そこで
作成や偏集のされた大量の文書全記憶・管理することが
現在要望されている。
These are generally called document processing. Traditionally, the central issue in document processing has been word processing, but there is now a need to store and manage all the large amounts of documents that have been created and collected.

これに対して、高記憶密度?保持する光デイスク記憶装
置が、文書全記憶する−ところの文書ファイル装置とし
て注目されている。しかるVこ、光ディスクの特徴の一
つは一度記録(記1.は)した情報は消去したり書替え
たシできないことであり、これは犬htの文書?記憶す
るに当っては欠点となる。
On the other hand, high storage density? Optical disk storage devices are attracting attention as document file devices that store entire documents. However, one of the characteristics of optical discs is that once the information is recorded (note 1.), it cannot be erased or rewritten.Is this a dog's document? This is a drawback when it comes to memorization.

その理由は、重子的な文書処理の最大の利点の一つが、
編集、修正、′校正などの書替えが非常に効率よくでき
るようになることであり、はとんどの文書がその書替え
の対象になるからである。すなわち、文1″ファイル装
置の中には編集や修正などが行われた非常に似た文書、
つまり多数の版が記憶されることになる。
The reason for this is that one of the biggest advantages of multiple document processing is
This is because rewriting such as editing, correction, and proofreading can be done very efficiently, and most documents can be rewritten. In other words, there are very similar documents that have been edited or modified in the document 1'' file device.
In other words, a large number of versions will be stored.

一部分しか変更されていない多数の頁からなる大きな文
書においては特に問題は大きく、重複の多い記1意とな
ってしまう。
This problem is particularly severe in large documents consisting of many pages that have only been partially changed, resulting in a large number of overlapping entries.

したがって、従来は大きな文書の場合は復戚の部分、た
とえば章や節に分割し、それぞれヶ一つの記憶情報単位
、すなわちファイルとして記1.ハ・管理していた。こ
の場合は、それぞrしの章f節の版の煩雑な管理ケ人間
が行なわねばならないという欠点がある。
Therefore, in the past, large documents were divided into parts, such as chapters and sections, and each part was recorded as a storage information unit, that is, a file. Ha-I was managing it. In this case, there is a drawback that the complicated management of the versions of each chapter and section f must be done by a person.

このように、個々の記憶情報単位?互いに独立に記憶す
る従来の記憶方式では、重複して記憶することにより記
憶効率が下る、また分割して記I、依する場合には多数
の版の管理が煩雑になる、という欠点があった。
In this way, individual memory information units? Conventional storage methods that store files independently of each other have the drawbacks that storage efficiency decreases due to redundant storage, and that managing multiple versions becomes complicated if the files are stored in separate parts. .

上記重複した記憶による記憶効率低下の欠点については
、従来も似た状況としてオンラインシステムの各時刻に
おける全状態量の保存という課題があるっオンラインシ
ステム、たとえば銀行システムでは万一のシステム事故
に備えて各時刻の状態を保存しておく必要がある。文字
通りこれら全状態量?保存することは膨大な記憶容量が
必要なために不可能である。したがって従来は、定期的
(たとえば1日に1回ンに全状態量とマスタファイルと
して記憶・保存し、各期間の各時刻の状態については、
全てのトランズアクション< trans actio
n ) lc記憶しておくことにょシ、万一事故が発生
したときにマスタファイルとトランズアクションから再
生するという方法金とっている。
Regarding the drawback of decreased storage efficiency due to redundant storage mentioned above, there is a similar problem in the past, in which there is a problem of preserving the entire state quantity at each time in an online system. It is necessary to save the state at each time. Literally all these state quantities? Storing it is impossible because it requires huge storage capacity. Therefore, in the past, all state quantities and master files were stored and saved periodically (for example, once a day), and the state at each time in each period was stored and saved as a master file.
All transactions
n) It is important to memorize the LC, but in the unlikely event that an accident occurs, there is a way to play it back from the master file and transaction.

この方法の要点は一つの規準とそれとの差異(相違部分
)と全記憶することにある。
The key point of this method is to memorize one standard and all of its differences.

前記文科ファイルにおける問題も基本的にはこの考え方
で解決することができる。規準となるファイルはある文
書の第1版であり、第2版との差異は編集プログラム(
エディタ)又はワードプロセッサに対する編集命令の列
として表現できる。
The problem with the literature file mentioned above can basically be solved using this idea. The standard file is the first version of a document, and the differences from the second version are determined by the editing program (
(editor) or a word processor.

しかし、このままの方法では次のような問題点がある。However, this method has the following problems.

すなわち、一般にはエディタやワードプロセッサには数
多くの種類があり、文書ファイル装置は記憶した文書ケ
編集命令の列から再生するために、前月己のエディタや
ワードプロセッサの1重類ケ記憶していることと、記憶
したエディタやワードプロセ・ツサと同等の機能ケもっ
ていることが必要である。
In other words, in general, there are many types of editors and word processors, and in order to play back from the stored sequence of document editing commands, the document file device remembers the single category of the editor or word processor that was used last month. It is necessary to have the same functionality as the editor, word processor, or editor you have memorized.

今後の文書ファイル装置ば単独で閉じた機能葡もち、他
の機器、たとえばワードプロセッサ、プリンタ、あるい
はパーソナルコンピュータなどとはネットワークで繋が
るようになる。したがって文書ファイル装置は接続され
るであろうすべてのワードプロセッサなどの機能は持つ
ことができない。
In the future, document file devices will have independent functions and will be connected to other devices such as word processors, printers, or personal computers via a network. Therefore, the document file device cannot have all the functions of a word processor to which it may be connected.

すなわち、オンラインシステムにおける方法はそのまま
適用できない。
In other words, the method used in the online system cannot be applied as is.

〔発明の目的〕[Purpose of the invention]

したがって、本発明の目的は上記欠点を改善し大量の文
書?すくない記憶容量で記憶可能にした記憶方式盆提供
することである。
Therefore, the purpose of the present invention is to improve the above-mentioned drawbacks and solve the problem of large amount of documents. To provide a storage tray that can store data with a small storage capacity.

〔発明の概要〕[Summary of the invention]

この目的?達成するため本発明においては、記憶すべき
暖数の情報単位について、特定の情報単位?除く他の情
報単位は特定の情報単位との重複部分音自動的に検出す
ることによシ冗長性ケ除去して相違部分のみ全記憶する
点に特徴がある。本発明の方式によると、記憶容量ヶす
くなくできるのみでなく重複を自動的に検出するから大
きな情報単位?細く分割して管理する必要性をなくすこ
とができる。
This purpose? In order to achieve this, in the present invention, regarding the information unit of the warm number to be stored, a specific information unit? The system is characterized by automatically detecting overlapping partials with the specific information unit for other information units, thereby eliminating redundancy and storing only the different parts. According to the method of the present invention, not only the storage capacity can be reduced, but also duplication is automatically detected, so it is possible to use large information units. It is possible to eliminate the need to manage the information by dividing it into small pieces.

〔発明め実施例〕[Embodiment of the invention]

以下、本発明V実施例にもとづいて詳細に説明する。 Hereinafter, a detailed explanation will be given based on embodiment V of the present invention.

第1図は本発明方式音用いる情報記憶装置のシステム構
成図である。本装置はコンピュータ401゜402 、
 CRT (Cathod Iもay Tube )デ
ィスプレイ403とキーボード404からなる端末;3
種類の副記憶装置410,411,412.およびロー
カルエリ、アネットワークとの接続葡する通信制個装@
421からなっている。なお、同図において461がロ
ーカルエリアネットワークの1言号バス、462が本装
置の内部データバスである。
FIG. 1 is a system configuration diagram of an information storage device using sound according to the present invention. This device includes computers 401, 402,
A terminal consisting of a CRT (Cathode I ay Tube) display 403 and a keyboard 404; 3
Types of secondary storage devices 410, 411, 412. And local area, communication system individual package for connection with network.
It consists of 421. In the figure, 461 is a one-word bus of the local area network, and 462 is an internal data bus of this device.

つぎに、第2図〜第8図荀用いて本発明の詳細な説明す
る。いま、説明葡簡明にするため、二つの文書ファイル
の中身はそれぞれ第2図(1)のテキスト(文字列)S
lおよびS2であるとする。同図は英文の例であるが日
本文の場合も呟く同じに説明できる。違いは英文の一芋
は1バイトで表現しうるのに対して日本文のそれは2バ
イト?要する点のみである。なお同図で記号C4Lは改
行コードを意味する。
Next, the present invention will be explained in detail with reference to FIGS. 2 to 8. Now, to simplify the explanation, the contents of the two document files are the text (character string) S shown in Figure 2 (1).
1 and S2. The figure shows an example of English text, but Japanese text can also be explained in the same way. The difference is that ``ichiimo'' in English can be expressed in 1 byte, while in Japanese it can be expressed in 2 bytes. Only the essential points are covered. Note that in the figure, the symbol C4L means a line feed code.

SIk第1版のテキスト、82ケ第2版のテキストとす
ると、後者は前者に対して以下のような処理?”ノーる
ことによシ得られる。
Assuming the text of SIk 1st edition and the text of 82 ke 2nd edition, the latter is processed as follows with respect to the former? ``You can get away with saying no.

CI)S、の最初の2文字外その′ま゛ま引用。CI) Exactly quotes except for the first two letters of S.

C2)次の1α文字?削除。C2) Next 1α character? delete.

(シ3)次の6文字を引用。(C3) Quote the next 6 characters.

C4)9文字” Xvatcbing△n′t−挿入。C4) 9 characters "Xvatcbing△n't-insertion.

(ここで△は空白?意味する。) C5)引続き7文字を引用。(Here, △ means blank?) C5) Continue to quote the 7 characters.

C6)次の10文字t2文字”TV”と置換。C6) Replace with the next 10 characters t2 characters “TV”.

(2文字“TV”r挿入して10文字削除) C7)続いて12文字引用。(Insert 2 characters “TV”r and delete 10 characters) C7) Followed by 12 character quotations.

C8)次の2文字金6文字” aZ\dark”と@ 
A、1C9)埴後の7文字を旬月。
C8) Next 2 letters 6 letters gold “aZ\dark” and @
A, 1C9) The 7 characters of Hanago are shungetsu.

したがって、Szkそのまま記憶する代シに、cl)〜
C8)の情報を記憶す几ばよい。同情報は第2図(2)
のように表現することができる。同図において1は第1
版の文書のファイル名称、ここではSlであり、2〜4
が挿入するところの新し。
Therefore, instead of memorizing Szk as is, cl)~
All you have to do is memorize the information in C8). The same information is shown in Figure 2 (2)
It can be expressed as: In the same figure, 1 is the first
The file name of the version document, here Sl, 2 to 4
The new where is inserted.

い文字列、10〜】3が二つの文字外、の差異の状態を
表わしている。ここで、記号C,D、r。
The character string 10~]3 represents the difference between two characters. Here, the symbols C, D, r.

EOFは次の意味全もつ、。EOF has the following meanings:

C・・・引用(Cite) D ・・・削除(1)elete) ■・・・挿入(In5ert ) E OF ・・・ファイル終了(End of p I
 Ic  )また、これらの記号に続く数値、例えば1
0のCに続く数値2は、文字数を示す。具体的には(C
,2)は2文字の引用ケ表わす。
C...Cite D...Delete (1) delete) ■...Insert (In5ert) E OF...End of file (End of p I)
Ic) Also, the numbers following these symbols, e.g. 1
The number 2 following 0C indicates the number of characters. Specifically (C
, 2) represents a two-character quotation.

第2図(3)に記号10,12.13の具体的な表現方
法の一例?示す。第2図(2)で示した記号(第2図(
3)では左側31)は具体的には同図(3)の右側32
のように表現できる。符7832は8ピツトで上位2ビ
ツトが(C,D、I、 EOF)の区別葡示し、下位6
ピツトがその長さtをバイナリで示す。tがθ〜63(
2’−1)までは1つの符号で表わせる。、tが64以
上のときは同−再の符月全連続して311シベ、tτ1
2ピッ′ト、18ビツト、・・・で表わす。例えば、同
図記号33のようにt=689のときは1.2語の符号
34.35で[l1689]7表わす。符号36は符号
34.35’の意味全等価的に表わしたもので、第1符
号34が下位、第2符号35が上位?示す。
An example of how to express symbols 10, 12, and 13 in Figure 2 (3)? show. Symbols shown in Figure 2 (2) (Figure 2 (
In 3), the left side 31) is specifically the right side 32 of the same figure (3).
It can be expressed as The code 7832 has 8 pits, the upper 2 bits indicate (C, D, I, EOF), and the lower 6 bits indicate the distinction between (C, D, I, EOF).
A pit indicates its length t in binary. t is θ~63(
2'-1) can be represented by one code. , when t is 64 or more, all the same-recurring sign months are consecutively 311 times, tτ1
It is expressed as 2 pits, 18 bits, etc. For example, when t=689, as shown by symbol 33 in the same figure, it is represented by [l1689]7 with the symbol 34.35 of 1.2 words. The code 36 is a fully equivalent representation of the meanings of the codes 34 and 35', with the first code 34 being the lower one and the second code 35 being the higher one? show.

この方式によれば、ファイル名称1を16バイト、挿入
文字列2〜4を各文字1バイトとして合計17バイト、
また差異を示す記号10〜13を各1バイトの合計12
バイトで表現することができ、結局51バイトの長さ全
もつテキストs2は16+17+12=45バイトで表
現することができ、6バイトだけ記憶に必要な量を削減
したことになる。
According to this method, file name 1 is 16 bytes, insertion character strings 2 to 4 are each 1 byte, and a total of 17 bytes.
Also, symbols 10 to 13 indicating differences are each 1 byte, totaling 12
It can be expressed in bytes, and as a result, text s2 having a total length of 51 bytes can be expressed in 16+17+12=45 bytes, reducing the amount required for storage by 6 bytes.

実際ノ文書においては書替えのある部分はテキスト全体
の小さな部分であり、記憶量削減の効果はもつと大きい
In an actual document, the part that is rewritten is a small part of the entire text, and the effect of reducing the amount of memory is large.

さて本発明は要するに第2図(1)K線で示した2つの
テキストの対応関係を自動的に検出し、前記CI)〜C
9)の情報を抽出する方法を与えるものである。つぎに
この方法について説明する。
Now, in short, the present invention automatically detects the correspondence between the two texts shown by the K line in FIG. 2 (1), and
9) provides a method for extracting information. Next, this method will be explained.

テキスト同志の対応関係は第3図のように示すとよシ分
シやすい。縦軸がSlで横軸が82である。ここで、・
印は引用0.X印は削除(ト)、0印は挿入(I) x
表わしている。これらの記号の列は同メツシュ上の2次
元領域の左下の端点から右上の端点まで連続していなけ
ればならない。この道程に?l−こてはパスと呼ぶこと
にする。
It is easier to understand the correspondence between texts as shown in Figure 3. The vertical axis is Sl and the horizontal axis is 82. here,·
The mark is 0. X marks delete (G), 0 marks insert (I) x
It represents. These strings of symbols must be continuous from the lower left end point to the upper right end point of the two-dimensional area on the same mesh. On this journey? The l-trowel will be called a pass.

したがって、対応関係を検出することはこの領域におい
てこのパスを、左下端から順次探索していくことに等し
い。
Therefore, detecting the correspondence is equivalent to sequentially searching for this path in this area starting from the lower left end.

第4図を用いて探索方法を説明する。いま一般的に、 s t = (a(i) )T、t         
 (1)S 2 = (b(j) )1゜1(2)と書
くことにする。
The search method will be explained using FIG. Now generally, s t = (a(i))T, t
(1) Let us write S 2 = (b(j) )1゜1(2).

探索はa(1)とb(1)の比較(マツチングという)
K始よる。a(1)とb(1)は等しいのでa(1+1
)’とb(i+i)に進み、更にa(3)とb(3)の
マツチングに進み失敗する(a(3)〜b (3) )
。第4図では目印で示す、 ここで次の一致点を探す過程に入る。同図で+印で探索
領域ケ示す。探索ば○印で囲った番号の順に進む。すな
わち、 a(3)b (3)−+ a (3) b(4)→a 
(4) b (,3)−a (3) b(5)−にa 
(4) b (4)−+ a (,3) b(5)−+
a(3)b((31→−・−(3)このときa (G)
 b (3)のマツチングが成功する(a(6)=b(
3))。
Search is a comparison of a(1) and b(1) (called matching)
K begins. Since a(1) and b(1) are equal, a(1+1
)' and b(i+i), then proceed to matching a(3) and b(3) and fail (a(3) to b(3))
. In Figure 4, we begin the process of searching for the next matching point, which is indicated by a landmark. In the figure, the search area is indicated by a + mark. When searching, proceed in the order of the numbers circled. That is, a(3)b (3)-+ a (3) b(4)→a
(4) b (,3)-a (3) b(5)-to a
(4) b (4)−+ a (,3) b(5)−+
a(3) b((31→-・-(3) then a (G)
b Matching of (3) is successful (a(6)=b(
3)).

この時点で、この点から以降T文字のマツチングが連続
して成功するか否か孕、 + 8 (6+t ) 、 b (3+t ) 1.、
、      (4)のマツチング?続行することによ
り検定する。第4図の例ではa (6+1 ) b (
3+1 )のマツチングが失敗し、先の探索過程へ戻り
探索を続行する。
At this point, whether or not matching of T characters succeeds continuously from this point onwards, + 8 (6+t), b (3+t) 1. ,
, Matching (4)? Verify by continuing. In the example of Figure 4, a (6+1) b (
3+1) fails and returns to the previous search process to continue the search.

結局66回目のマツチングa (13) b[3)がT
2に対して成功して、[メ降同様の過程全繰返兄す。
In the end, the 66th matching a (13) b[3] is T
If you succeed against 2, repeat the same process again.

探索の開始点がa (3) b (3]で終了点(マツ
チングが成功した点)がa(、t3)’b(3]である
ことから、文字部分列a(3)〜a(13−1)が削除
であることが分る。 一般(C開始点ケミ(+1)b(
)・ )、終了点をa(j2)bN2 )とすると、文
字部分列a(il)〜a(i21):削除   (5)
文字部分列b(j+)〜’)(j2t):挿入   (
G)であり、マツチングの成功した文字部分列が引用と
なる。ただし、il≧12−1のときは削除なし、j1
≧32−1のときは挿入沈しである。
Since the starting point of the search is a (3) b (3) and the ending point (the point where matching was successful) is a(,t3)'b(3], the character substrings a(3) to a(13 -1) is a deletion. General (C starting point chemistry (+1) b(
)・ ), and the end point is a(j2)bN2 ), character substring a(il) to a(i21): Delete (5)
Character substring b(j+)~')(j2t): Insert (
G), and the character substrings that are successfully matched become quotations. However, when il≧12-1, no deletion occurs, j1
When ≧32-1, it is insertion sinking.

この過程?最後まで続けたときの様子(r第5図に示す
。探索領域は斜線で示した直角二等辺三角形の領域であ
る。探索の/こめのマツチングの回数は約2・10回で
ある。ここでT−:2である。
This process? What it looks like when continued to the end (r shown in Figure 5. The search area is a right-angled isosceles triangle area indicated by diagonal lines. The number of matching operations in the search is approximately 2.10 times. Here, T-:2.

パラメータ゛Pはマツチングの長さであり、小さく選び
ずらゐと局、・目的に最適なパスが見い出さ&’L′I
ヨ体として正しい対応関係が・べられないことがある3
、第6図はT−1のそのような場合の例である。
The parameter ゛P is the matching length, and if it is chosen to be small, the optimum path for the purpose will be found &'L'I
Sometimes it is not possible to find the correct correspondence as a Yo body 3
, FIG. 6 is an example of such a case for T-1.

このような誤った対応は正I7い情報の記1意には影響
ケ与え役いが、記憶効率ケ最適な場合に比して小さくす
る。二つのテキストの一致部分が小さい状況(第1図の
ように)では逆に原デキスト【記憶さするより多くの記
憶量を要求する可能性がある。しかし、rの値全適当な
大きな値にしておけばこの問題は確率的に小さい。シス
テム的な対策としては、原テキストの長さと本方式(C
よる記・:、(マ長(量)と孕比較して、短い方?選択
する方法が考えられる。
Although such erroneous correspondence has an effect on the memory of correct information, it makes the storage efficiency smaller than in the optimal case. In situations where the matching portion of the two texts is small (as in Figure 1), this may require more storage than the original dexterity. However, if the value of r is set to a suitably large value, this problem will be reduced in terms of probability. As a systematic measure, the length of the original text and this method (C
By comparison: (Comparing macho (amount) and pregnancy, which one is shorter? There are ways to choose.

以上説明した二つのテキストの対応関係ケとる方法?ハ
ターンマッチングという。ここで、本パターンマツチン
グのアルゴリズムの流れ図ケ第7図(])〜(7)に示
す。本アルゴリズムは2本のテキスト(文字コード列)
(1)式(2)式を入力し、第2図に示すような差異全
表わすコード列?出力する。ここでテキストS1が規準
で、差異はSlからのS・の1晶差である。
How to find the correspondence between the two texts explained above? It's called Hatern matching. Here, a flowchart of the pattern matching algorithm is shown in FIGS. 7(] to 7). This algorithm uses two texts (character code strings)
(1) Input the formula (2) and enter the code string that represents all the differences as shown in Figure 2? Output. Here, the text S1 is the standard, and the difference is one crystal difference of S from S1.

ここで、第7図(1)−(7)(r若干説明する。Here, FIGS. 7(1)-(7)(r) will be briefly explained.

第7図(1)において、才ずステップ201で初朗化ケ
行う。ここでi、Jはそルぞれテキス)S++82に対
t−るポインタ、m、nはテキストの一致する+qtt
分の端点および不一致の始まる点オ記憶するためのポイ
ンタである。また、Rは差異を表わすコード列であり、
規準となるテキスト名称(ファイル名称)f(Sl )
に初期化される。
In FIG. 7(1), the initial transformation is performed in step 201. Here, i, J are respectively text) pointers to t- to S++82, m, n are text matching +qtt
This is a pointer for storing the end point of the minute and the starting point of the discrepancy. Further, R is a code string representing a difference,
Standard text name (file name) f (Sl)
is initialized to .

ラベル101から102は一致している部分テキストr
固定する処理である。ステップ202゜203において
テキストの終端(’EOF)k検知した場合は、そnぞ
れ終了処理のステップ261゜262へ飛ぶ。ステップ
204において二つの文字コードa(i)とb(Dが一
致しているか否かケ判定し、一致しているときはステッ
プ205でポインタケ進めて、次の文字コードヶ比較す
る。一致していないとき・Jl、一致していた部分全表
現するコード7作り几に追加する(202)。ステップ
202において、■記号は付加(apl〕end)する
ことを表わす。記号[Ci7’]は第2図のコード2?
表わす。具体的には、長さtだけテキスト名称用(Ci
te)する、すヱわちtだけ一致していたことケ意味す
る。
Labels 101 to 102 are matching partial texts r
This is a fixing process. If the end of text ('EOF) k is detected in steps 202 and 203, the process jumps to steps 261 and 262 of the end process, respectively. In step 204, it is determined whether or not the two character codes a(i) and b(D) match. If they match, the pointer is advanced in step 205 and the next character code is compared. When Jl, add to the code 7 creation process that expresses all the matching parts (202). In step 202, the ■ symbol represents addition (apl] end). The symbol [Ci7'] is shown in Figure 2. Code 2?
represent. Specifically, for the text name (Ci
te), which means that only t matched.

第7図(2)において、ラベル102以降は文字コード
の不一致が見つかった後に、次の一致点ケ捜すところの
探索部分である。途中、ステップ211゜212におい
てテキストの終端ゲ検ノ、口した場合は、終了処理のス
テップ263,264へ飛ぶ。ステップ213で始まり
213に戻るループで第4図に示した三角形の領域の探
索?実現する。ステツプ215以降の処理では、T文字
だけ連続して部分テキストが一致するか否が全判定する
。一致しないときは探索ケ続行し7、一致するときはラ
ベル103へ飛ぶ。
In FIG. 7(2), the area after label 102 is a search portion where the next matching point is searched for after a character code mismatch is found. If the end of the text is detected during steps 211 and 212, the process jumps to steps 263 and 264 of the end process. Search for the triangular area shown in FIG. 4 in a loop starting at step 213 and returning to 213? Realize. In the processing from step 215 onwards, it is determined whether or not the partial texts match by consecutive T characters. If there is no match, the search continues 7, and if there is a match, the process jumps to label 103.

第7図(3)において、ラベル103以降不一致部分に
対するコードr出力する。ステップ231ではRに記号
[I ; e](長さtだけ挿入ン勿追加し、更にステ
ップ232において挿入する部分テキスト(b(n)、
・・・・・・、b(j+β−1))全追加する。
In FIG. 7(3), a code r is output for the mismatched portion after label 103. In step 231, the symbol [I;
......, b(j+β-1)) Add all.

第7図のアルゴリズムの以降の部分については以上の説
明から理解できるので、説明音名P1hする。
Since the subsequent parts of the algorithm shown in FIG. 7 can be understood from the above explanation, the pitch name will be explained as P1h.

さて、上記アルゴリズム(iその−1,までは第5図か
らも理解されるように、長い部分が削除されたシ挿入さ
t’l−fc 、jl)すると、探索領域が大きくなり
パターンマツチングに要する時間が長くなるという問題
がある。
Now, as can be understood from Fig. 5, the above algorithm (up to i-1, the long part is deleted and inserted t'l-fc, jl), the search area becomes larger and the pattern matching There is a problem in that it takes a long time.

次にこの間碩wWI決するための拡張アルゴリズムr説
明する。
Next, an extended algorithm r for determining the current WWI will be explained.

拡張アルゴリズムの原理?第8図に示す。その原理はマ
ツチング全行う単位紮、文字単位よりも大きくすること
である。第8図の場合は空白(文字コードの一種)で区
切られる部分の単位でマツチングす・へ。その晰位は単
語とは限られず、より一般的に設定することができる。
Principle of expansion algorithm? It is shown in FIG. The principle is to make the unit that performs all matching larger than the character unit. In the case of Figure 8, matching is performed in units of parts separated by spaces (a type of character code). The position is not limited to words, and can be set more generally.

より大きい1夕1]としては、改行コード(C/几記号
で示す)で区切られる「行」の単位である− この上う圧大きな単位?1回のマツチングで比較するた
めには、文字コード列から何らかの簡単に比較できる特
徴?抽出することが必要である。
1] is a unit of "line" separated by a line feed code (indicated by the C/几 symbol) - Is this a larger unit? In order to compare in one matching, is there any feature that can be easily compared from the character code string? It is necessary to extract.

有効な特徴の一つけ第8図の場合に各四辺形の内側によ
る数字で示さ1.るように、−[;記名区分の長さであ
る。各区分の長さ孕、2本のテキストに対して、それぞ
n先と同様に(a<1> )”−1,(b(j)l ’
、−1と書けば先に示したアルゴリズム、を全くそのま
ま利用して一致する単語又は行の候補孕二本のテキスト
から探し出すことができる。
In the case of Figure 8, valid features are indicated by numbers inside each quadrilateral.1. This is the length of the name segment, as shown in the following. The length of each segment is (a<1>)"-1, (b(j)l'
, -1, the algorithm shown above can be used exactly as is to search for matching words or lines from two texts.

この大きな単位でのパター/マツチングバ一致しない部
分に遭遇したとき、すなわち第7図のアルゴリズムの探
索過程(同図xo2>−e行えばより0つfシ一致部分
の同定は文字コード単位でマツチング7行い、一致しな
い部分に遭、J1シたときに、単語又は1斤・\の分割
?行い長、!ヲ計測しながら、次の一致する単語又は行
r探索する。
When a pattern/matching match is encountered in this large unit, in the search process of the algorithm shown in Figure 7 (xo2>-e in the same figure, it is more Then, when you encounter a non-matching part and press J1, search for the next matching word or line while measuring the word or division of 1 loaf/\?action length, !.

すなわら、拡張アルゴリズムは前述の単純なパターンマ
ツヂングアルゴリズムヶ階層的にす:ねて用いるもので
ある。
In other words, the extended algorithm is a layered version of the simple pattern matching algorithm described above.

第8図を用いて若干具体的に説明する。まず2本のテキ
ス)S+ 、82は最初の2文字゛■Δ″(Δは空白?
意味する)が一致し、次のS″と”e″が一致しない5
したがって探索過程に入って、まず空白で区uJつた区
間(ハ5語)の長さを到ると、10と6で一致しない。
This will be explained in more detail using FIG. 8. First, two texts) S+, 82 is the first two characters ゛■Δ'' (Δ is a blank?
) matches, and the next S″ and “e” do not match 5
Therefore, in the search process, when we first reach the length of the interval (c) divided by blanks (c), 10 and 6 do not match.

そこで引続いて以前と同様に三角形の探紫領域ケ順次展
開していく。
Then, as before, we will sequentially expand the triangular purple detection area.

次はlOと9で一致しない。その次け6と6で長゛さは
一致する。そこで本当に一致しているか否かケ文手コー
ドのレベルで同定し、”enjoy”と”enjoy”
が一致していることが分り、一致部同定部にもどる。次
に再びa″と“W″が一致しa″ctz:+す・ 2[
”CEl (7J T’J ’jib iM ’f”L
 ICA A・11同様に進行する。
Next, IO and 9 do not match. The lengths of the next 6 and 6 are the same. Therefore, we identify whether or not they really match at the Bunte code level, and distinguish between "enjoy" and "enjoy".
It turns out that they match, and we return to the matching part identification section. Next, a″ and “W” match again and a″ctz:+su・2[
”CEl (7J T'J 'jib iM 'f'L
Proceed in the same way as ICA A.11.

このように拡張アルゴリズムでは単純なアルゴリズムに
比して短い探索処理でパターンマツチングを行うことが
できる。一般にN文字の削除、挿入、又はM文字とN文
字の買戻が行われN2Mであるとすると、探索過程での
処理量は N・(N+1)/2          (7)に比例
する。いま、単語又は行の平均の長さがL文字であると
すると、拡張アルゴリズムでの探索処理量は に比例する。明らかに単位の長さが大きい程処理量は少
なくてすむ。
In this way, the extended algorithm can perform pattern matching with a shorter search process than the simple algorithm. In general, if N2M is obtained by deleting or inserting N characters or redeeming M and N characters, the amount of processing in the search process is proportional to N·(N+1)/2 (7). Now, assuming that the average length of a word or line is L characters, the amount of search processing in the extended algorithm is proportional to . Obviously, the larger the length of the unit, the smaller the amount of processing required.

ちなみに、第5図(単純なアルゴリズム)の場合は、探
索処理量は240のオーダで、第8図の場合は16のオ
ーダである。但し、仁こでは(8)式と違って、三角形
の面積ではなく実際のマツチング回数?計数した。
Incidentally, in the case of FIG. 5 (simple algorithm), the search throughput is on the order of 240, and in the case of FIG. 8, it is on the order of 16. However, unlike Equation (8), in Niko, it is not the area of the triangle but the actual number of matchings. I counted.

つぎに、本発明の記憶方式r第1図の装置に適用する場
合について説明する。
Next, the case where the storage method of the present invention is applied to the apparatus shown in FIG. 1 will be explained.

第1図の装置における3種の副記憶装置41o。Three types of secondary storage devices 41o in the device shown in FIG.

411,412tまそれぞれ、)Y、ディスク装置、固
にヘッド磁気ディスク装置、およびフロッピ磁気ディス
ク装置である。光デイスク装置は以後変更の起らない凍
結した恒久情報単位(ファイル)の記憶、又は画(酸デ
ータのように多量な情報量ケもつ情報単位の記憶rする
。固定ヘッド磁気ディスク装置は多数の、版に分れる情
報単位のうち最も新しい版、すなわち変更の起りうる凍
結さnなり情報単位の記憶と、光デイスク装置に記i意
さ11.ている情報単位のカタログの記憶と、一時的な
記憶なトケスる。また、装置全体7制御するシステムプ
ログラムや、ワードプロセシングなどケ行う処理プログ
ラムもae1意する。第3の副記憶装置である70ソピ
磁気デイスク装置は、他のスタンドアロンの機器、たと
えばワードプロセッサなどとの情報交換のために存在す
る。たとえばネットワークにつながらないワードプロセ
ッサで大量多種の文′9I:ヲ作成する場合は、それら
の記憶・管理金本装野で行い、本装置から文書)゛fイ
ル全フロッピディスクに読出してワードプロセッサへ運
ヒ、処理力り多丁したとき本装置・にもどすことができ
る。
411 and 412t, respectively) Y, disk device, head magnetic disk device, and floppy magnetic disk device. Optical disk devices store frozen permanent information units (files) that will never be changed, or store information units with a large amount of information, such as images (acid data).Fixed head magnetic disk devices store a large number of , storage of the latest version of information units divided into versions, that is, frozen information units that are subject to change; storage of catalogs of information units stored in optical disk devices; There is also a system program that controls the entire device and a processing program that performs word processing, etc.The third secondary storage device, a 70-segment magnetic disk device, is used to store other stand-alone devices. exists for exchanging information with, for example, a word processor.For example, when creating a large number of various types of text with a word processor that is not connected to a network, it is stored and managed by Kanemoto Sono, and the documents are transferred from this device. The file can be read out onto all floppy disks and sent to a word processor, and returned to the present device when the processing power is too high.

本装置はローカルエリアネットワーク上の7アイリング
ステーシヨンとしての没利ケもっと同時に、ワードプロ
セシング等の処理機能にも持つ。
This device not only functions as a seven-way station on a local area network, but also has processing functions such as word processing.

ここでは発明の中心であるファイリングステーションと
しての基本的々役割である記1意についてのみ説明し、
他の機能については公知技術により実現できるので説明
ケ省略する。
Here, we will only explain the basic role of the filing station, which is the center of the invention.
The other functions can be realized using known techniques, so their explanation will be omitted.

本装置自身から、あるいはネットワーク全弁しての記憶
要求は第9図に示す構造のデータで表現する。同図にお
いて、第1記録501は要求内容、第2記録502は記
憶又は読出し要求時のそのファイル名称、第3記録は同
ファイルの属性ケ表わす。第4および第5記録は記憶時
に存在して、それぞれ旧フアイル名称(処理會する母体
となったファイル)および記憶すべきデータの本体であ
る。
A storage request from this device itself or from the entire network is expressed by data having the structure shown in FIG. In the figure, a first record 501 shows the request content, a second record 502 shows the file name at the time of a storage or read request, and a third record shows the attributes of the same file. The fourth and fifth records exist at the time of storage, and are the old file name (the file that became the base file for processing) and the main body of the data to be stored, respectively.

EOF記号506はデータ本体505の末尾に付方iさ
nている。
The EOF symbol 506 is placed at the end of the data body 505.

本装置に出された記憶要求は−J↓副記憶装置411上
のスプール(SpOOt)に記憶さ几、その後要求内容
の解析とその実行の実1j’1件のチェック1行う。要
求内イに誤りがなければ同要求ケ受理した旨ケ内部状態
表に憚へ込み、同時に要求元にその旨?伝達する。内部
状態表は同記憶情報単位が通常の記憶領域には存在せず
、まだスプール上にあることケ示している。システム的
に装置の外部から眺めたときに、′ま、記憶情報単位が
具体的にどこであるかは見えないように干る。すなわち
、この状態で外部より同記・、#、H報単位の読出し要
求があった場舒は、通常と全く同様に1、ノ″tみ出す
ことができる。
A storage request issued to this device is stored in the spool (SpOOt) on the -J↓ secondary storage device 411, and then the content of the request is analyzed and its execution is checked 1j'. If there is no error in the request, the request will be acknowledged in the internal status table, and at the same time, the requester will be informed that the request has been accepted. introduce. The internal status table indicates that the same storage information unit is not in normal storage, but is still on the spool. When viewed from outside the device in terms of the system, it is difficult to see the specific location of the storage information unit. That is, in this state, if there is an external read request for the same, #, or H information, the data can be read out in exactly the same way as normal.

スプール上にある情報単位は本文で説明した拡張アルゴ
リズムケ用いて冗長性ケ除いた後にスプール上から本来
の主記憶領域に移動する。本装置では上記冗長性ケ除く
処理と、次の要求紫受理・解析する処理と?並列して行
う。W数のタスクヶ並列して実行する技術については公
知の技術であるので説明を省略する。
The information units on the spool are moved from the spool to the original main storage area after redundancy is removed using the expansion algorithm described in the main text. In this device, there is a process to remove the redundancy mentioned above, and a process to accept and analyze the next request. Do it in parallel. Since the technique of executing W tasks in parallel is a well-known technique, a description thereof will be omitted.

さて、記憶全完了するまでの処理金より詳しく第10図
?用いて説明する。同図(a)は記憶要求の受理が完了
した状況、(b)は記憶処理全体が完了した状況である
。外部から眺めたときV′1(a)は記憶完了と見える
Now, what about Figure 10 in more detail than the processing fee until the memory is completely completed? I will explain using (a) of the figure shows a situation in which the storage request has been accepted, and (b) shows a situation in which the entire storage process has been completed. When viewed from the outside, V'1(a) appears to have completed storage.

副記憶装置411は第10図において記憶領域611全
もち、それは4つの副領域621,622゜623.6
24に分かnる。そjしそれ副記憶装置410のための
カタログ、副記憶装置411自身のカタログ、前記スプ
ール、および主記憶領域である。また副記憶装置410
は記憶領域610にもつ。
The sub storage device 411 has the entire storage area 611 in FIG. 10, which is four sub areas 621, 622, 623,
It's about 24 minutes. These are the catalog for the secondary storage device 410, the catalog of the secondary storage device 411 itself, the spool, and the main storage area. Also, the secondary storage device 410
is also stored in the storage area 610.

いま、記憶要求の前には第1版の文書Text、1と第
2版’l”ext、2が記憶されていたとする。前者6
53は後者よシ古いので領域610に、後者652はそ
の時点で最新版であるので領域624に記憶されている
3、 第3版’ll’exe、3の記憶要求があったとすると
、その受理直後は第1O図(a)のように、’l’ex
t、3の本体651はスプール領域623にある。
Now, it is assumed that the first version of the document Text, 1 and the second version 'l"ext, 2 were stored before the storage request. The former 6
53 is older than the latter, so it is stored in the area 610, and the latter 652 is the latest version at that time, so it is stored in the area 624. If there is a request to store 3, 3rd edition 'll'exe, then the request will be accepted. Immediately after, as shown in Figure 1O (a), 'l'ex
The body 651 of t,3 is in the spool area 623.

スプール内の情報単位の存在音検知し、記憶内容の書替
え(冗長性を除くための)処理?開始する。同処理結果
が第10図(b)である。
Detects the existence of information units in the spool and rewrites the stored contents (to remove redundancy)? Start. The processing result is shown in FIG. 10(b).

第2版Text、2の第3版Text、 3からの差異
(相違部分)r抽出し、情報単位655として領域61
0に追加記憶し、もとの情報単位652は削除する。カ
タログ621,622内のカタログ情報はそれに応じて
書替える。次に、スプール内の情報単位651の中のテ
キスト本体505(第9図)?情報単位654として領
域624に書込む。またそれに応じてカタログ622紫
書替える。
Differences (differences) r from the 2nd edition Text, 2 and the 3rd edition Text, 3 are extracted and the area 61 is extracted as an information unit 655.
0 and the original information unit 652 is deleted. Catalog information in catalogs 621 and 622 is rewritten accordingly. Next, the text body 505 (FIG. 9) in the information unit 651 in the spool? It is written in the area 624 as an information unit 654. Catalog 622 will also be rewritten accordingly.

したがって副記憶装置410(光ディスク)の中には冗
長性紮除いた変化分653,655が記憶されることに
なり、効率的な記憶ケ実現する。
Therefore, the changes 653 and 655, excluding redundancy, are stored in the secondary storage device 410 (optical disk), realizing efficient storage.

更に、アクセス頻度の高い最新版654は原形のまま記
憶さn1アクセス時間は従来通り短)、−1゜古い版?
アクセスするときは、最新版から古い版を差異情報から
順次復元する。したがって、最新版よシもアクセス時間
が長くなる。復元の方法は行単位のテキスト編集ケ行う
エディタに用いられる方法と同じで、公知であるので説
明は省略する。
Furthermore, the latest version 654, which is frequently accessed, is stored in its original form (n1 access time is short as before), -1゜older version?
When accessing, restore sequentially from the latest version to the older version based on the difference information. Therefore, the access time will be longer even for the latest version. The restoration method is the same as the method used in editors that edit text line by line, and is well known, so the explanation will be omitted.

非常に古い版のアクセス時間が極端に遅くならないよう
に、本装置ではに版毎(Kはパラメータとして指定可能
)に原形のまま記憶する。たとえばI(=3とすると、
Text、3 、 Text、6 、 ・−は光デイス
クファイルに冗長性を除去せずに記憶する。
In order to prevent extremely slow access times for very old versions, this device stores each version (K can be specified as a parameter) in its original form. For example, if I (=3),
Text, 3, Text, 6, . . . are stored in the optical disk file without removing redundancy.

これにより、長い復元処理の連鎖ケ作らずにすむ。This eliminates the need to create a long chain of restoration processes.

以上のように本実施例によれば、版の異る重Fiの多い
文書?、自動的に相違部分を抽出することにより冗長性
金除去した形で記憶し、結果的に従来に比して多量の文
書を同一の記憶容量で記憶させることかり能である。
As described above, according to this embodiment, documents with many heavy files with different versions? By automatically extracting different parts, the document is stored in a form with redundancy removed, and as a result, it is possible to store a larger amount of documents in the same storage capacity than before.

なお、本実施例では副記憶装@410は■替え′ができ
ない光ディスクであったが、光ディスクを用いずに磁気
ディスクを用いてもよいし、副記憶装#410は副記憶
装置411と一体となった磁気ディスク装・dであって
もよい。
In this embodiment, the secondary storage #410 is an optical disk that cannot be replaced, but a magnetic disk may be used instead of an optical disk, and the secondary storage #410 may be integrated with the secondary storage 411. It may also be a magnetic disk drive.

更に、本記憶方式は磁気テープなどの他の記憶装置にも
そのまま適用できることは言うまでもない。
Furthermore, it goes without saying that this storage method can be applied as is to other storage devices such as magnetic tape.

また文書は日本文でも欧文でもよい。相違は文字コード
の長さく前者は2バイト、後者は1バイト)と、拡張パ
ターンマッチングケ行う際の分割用の文字コードが異る
のみである。文書が日本文か欧文かは、第9図の第3記
録503のファイル属性の中に記録されている。分割7
行うための記号としては、空白や改行の他に、日本文で
は「。」や「、」、あるいはそnらの集合であってもよ
い。
Furthermore, the document may be in Japanese or European language. The only difference is the length of the character code (the former is 2 bytes, the latter is 1 byte) and the character code used for division when performing extended pattern matching. Whether the document is in Japanese or European is recorded in the file attributes of the third record 503 in FIG. division 7
In addition to blank spaces and line breaks, symbols for doing this may also be ".", ",", or a set of these in Japanese.

また本実施例では2本のテキストの差異?すべて自動的
に求めたが、記憶要求元、たとえばある棟のエディタか
ら補足情報?ヒントとして得てもよい。たとえば、長い
文書の場合にそれをブロックに分け、変更のあったブロ
ックにその旨knすフラッグ?立てることが考えら几る
。このような拡張も本発明に含まれる。
Also, in this example, is there a difference between the two texts? Everything was requested automatically, but supplementary information from the source of the memory request, for example, an editor in a certain building? You can take it as a hint. For example, if you have a long document, how about dividing it into blocks and flagging blocks that have changed? I can't even think of standing up. Such extensions are also included in the present invention.

さらに、文書は必ずしも文字コードげか9ではなく、図
形や画像が混在している場合もある。更に、オンライン
タブレットなどから入力したコメントなどの筆跡データ
が混在している場合もある。
Furthermore, a document does not necessarily have a character code of 9, but may also contain a mixture of figures and images. Furthermore, handwriting data such as comments input from an online tablet may also be included.

しかし、これらの場合、異種データはそれぞれ同種のデ
ータ毎にグループ化され、データエンベロープというブ
ロックに入れられる。このような場合には、各ブロック
毎に本発明方式?適用することができる。
However, in these cases, different types of data are grouped into the same type of data and put into blocks called data envelopes. In such a case, should the method of the present invention be used for each block? Can be applied.

また、この場合に、デルタエンベロープ毎に変更があっ
たか否か?同定して、データエンベロープr単位として
冗長性を除くこともoJ能である。
Also, in this case, was there a change for each delta envelope? It is also possible to identify and remove redundancy as a data envelope r unit.

たとえば、第1版の文書に対してオンラインタブレット
からコメン)k加筆した第2版の文書は、同コメントと
本体のためのポインタのみを記憶すれば゛よいうこγL
らの拡張もすべて本発明に含まれる。
For example, if you add comments to the first version of a document from an online tablet, you can create a second version of the document by remembering only the comments and the pointer for the main body.
All such extensions are also included in the present invention.

〔発明の効果〕〔Effect of the invention〕

本発明方式によれば、重複の粕い記憶情報単位の中から
重複部分と相違部分とを自動的に抽出して、冗長性金除
去して記憶するので、一定の記1、ホ容量で従来よシも
多量の・清報を記憶することができる。
According to the method of the present invention, overlapping portions and different portions are automatically extracted from duplicate storage information units, and redundancy is removed before storage. Yoshi is also able to memorize a large amount of news.

特に文書?扱う場合には、多数の異る版は数%しか互い
に相違していないことも多い。仮に10%相違している
100の長さの文書が5版あるとすると、従来の記憶方
式では500、本発明方式では(K−5として)140
→−αのl己1.ホ量ですみ・約3倍の記1意?するこ
とができる。
Especially documents? In many cases, a large number of different versions differ from each other by only a few percent. Assuming that there are 5 versions of a document with a length of 100 that differs by 10%, the conventional storage method would store 500 copies, and the inventive method would store 140 copies (as K-5).
→-α's self 1. It's only a small amount, about 3 times as much? can do.

この効果は、特に光ディスクのように書替えができない
記1意装置において大きい。
This effect is particularly great in recording devices that cannot be rewritten, such as optical discs.

才だ副次的効果として、従来記憶界;i二の制限から文
書作成の過程である古い版の記憶はN複が多いゆえに行
っていなかったが、本方式の採用によりすべての過程を
残しておくことができる。この効果は定量的に計測する
ことは難しいが、文書ケ媒体として進める仕事の「質」
?格段に向上させることができる。
As a side effect, in the past, due to the limitations of the memory world, the memorization of old versions, which is the process of document creation, was not done because there were many N-folds, but by adopting this method, all processes can be preserved. You can leave it there. Although it is difficult to measure this effect quantitatively, it does affect the quality of work carried out as a document management medium.
? It can be improved significantly.

【図面の簡単な説明】[Brief explanation of drawings]

第1図は本発明の方式?適用する記憶装置のシステム借
成図、第2図(1)は本発明の基本原理の説明の図、第
2図(2)i3)は2つのテキストの相違の表現方法?
示す図、第3図から第6図はそれぞれ、自動的にテキス
トの一致/相違部分?抽出する方法の説明図、第7図は
四方法?実現するアルゴリズムの流れ図、’iR8図は
拡張アルゴリズム全説明データの構造ケ示す図、第10
図は第11gの装置の記・謹傾城の関係ケ示す図である
。 臀 3 図 、S2 第 4 図 子 7 口(3) 第7 図(4) 冗 7 図(S) 第 7 図(につ
Is Figure 1 the method of the present invention? Figure 2 (1) is a diagram explaining the basic principle of the present invention, and Figure 2 (2) i3) is a diagram of the system of the storage device to which it is applied.
The figures shown in Figures 3 to 6 each automatically match/difference text? An explanatory diagram of the extraction method, Figure 7 shows the four methods? Flowchart of the algorithm to be realized, 'iR8 diagram is a diagram showing the structure of the extended algorithm complete explanatory data, No. 10
The figure is a diagram showing the relationship between the 11th device's description and the leaning castle. Buttocks 3 Figure, S2 Figure 4 Child 7 Mouth (3) Figure 7 (4) Red Figure 7 (S) Figure 7 (Nitsu)

Claims (1)

【特許請求の範囲】[Claims] 記憶すべき第1の“情報単位と、すてに記憶さルている
第2の情報単位との相違部分?抽出し、該相違部分荀前
記第1の情報単位の代りに第3の情報単位として記憶さ
せることケ特徴とする情報記憶方式。
A difference between a first information unit to be stored and a second information unit that has already been stored is extracted, and a third information unit is used instead of the first information unit. An information storage method characterized by the ability to store information as
JP57158125A 1982-09-13 1982-09-13 Information storing system Pending JPS5949062A (en)

Priority Applications (1)

Application Number Priority Date Filing Date Title
JP57158125A JPS5949062A (en) 1982-09-13 1982-09-13 Information storing system

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
JP57158125A JPS5949062A (en) 1982-09-13 1982-09-13 Information storing system

Publications (1)

Publication Number Publication Date
JPS5949062A true JPS5949062A (en) 1984-03-21

Family

ID=15664834

Family Applications (1)

Application Number Title Priority Date Filing Date
JP57158125A Pending JPS5949062A (en) 1982-09-13 1982-09-13 Information storing system

Country Status (1)

Country Link
JP (1) JPS5949062A (en)

Cited By (8)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPS61193241A (en) * 1985-02-21 1986-08-27 Hitachi Ltd Update history recording method
JPS6289134A (en) * 1985-10-16 1987-04-23 Nippon Steel Corp Character string difference extracting method and its device
JPS62128364A (en) * 1985-11-30 1987-06-10 Toshiba Corp Picture file device
JPS6376031A (en) * 1986-09-19 1988-04-06 Fujitsu Ltd File difference calculation processing system
JPS63184850A (en) * 1987-01-27 1988-07-30 Alps Electric Co Ltd History management system
JPS63305439A (en) * 1987-06-08 1988-12-13 Nippon Steel Corp Compressive storing method for similar data file and its restoring method
JPH02181224A (en) * 1988-09-30 1990-07-16 Yokogawa Electric Corp Software development system
JPH04168569A (en) * 1990-10-31 1992-06-16 Chubu Nippon Denki Software Kk Generation managing system for document file

Citations (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPS5697144A (en) * 1979-12-29 1981-08-05 Fujitsu Ltd File comparison system

Patent Citations (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPS5697144A (en) * 1979-12-29 1981-08-05 Fujitsu Ltd File comparison system

Cited By (8)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPS61193241A (en) * 1985-02-21 1986-08-27 Hitachi Ltd Update history recording method
JPS6289134A (en) * 1985-10-16 1987-04-23 Nippon Steel Corp Character string difference extracting method and its device
JPS62128364A (en) * 1985-11-30 1987-06-10 Toshiba Corp Picture file device
JPS6376031A (en) * 1986-09-19 1988-04-06 Fujitsu Ltd File difference calculation processing system
JPS63184850A (en) * 1987-01-27 1988-07-30 Alps Electric Co Ltd History management system
JPS63305439A (en) * 1987-06-08 1988-12-13 Nippon Steel Corp Compressive storing method for similar data file and its restoring method
JPH02181224A (en) * 1988-09-30 1990-07-16 Yokogawa Electric Corp Software development system
JPH04168569A (en) * 1990-10-31 1992-06-16 Chubu Nippon Denki Software Kk Generation managing system for document file

Similar Documents

Publication Publication Date Title
EP1406181B1 (en) Document revision support
US5355472A (en) System for substituting tags for non-editable data sets in hypertext documents and updating web files containing links between data sets corresponding to changes made to the tags
US7673235B2 (en) Method and apparatus for utilizing an object model to manage document parts for use in an electronic document
US7617444B2 (en) File formats, methods, and computer program products for representing workbooks
US5140521A (en) Method for deleting a marked portion of a structured document
US6901418B2 (en) Data archive recovery
WO2004057494A1 (en) Building one or more indexes on data concurrent with manipulation of data
CN112347765B (en) Entity labeling method, module and device based on dictionary matching
CN112395851A (en) Text comparison method and device, computer equipment and readable storage medium
WO2020119143A1 (en) Database deleted record recovery method and system
US6631385B2 (en) Efficient recovery method for high-dimensional index structure employing reinsert operation
CN116090416B (en) Standard writing method, system, equipment and medium based on standard knowledge graph
CN115995087B (en) Method and system for intelligent generation of document catalog based on fusion of visual information
JPH02297284A (en) document processing system
CN114546886A (en) Space recovery method of value log system
US6357002B1 (en) Automated extraction of BIOS identification information for a computer system from any of a plurality of vendors
JPH0588957A (en) Directory format
CN116185711A (en) Data backup and recovery method and device
JP2822869B2 (en) Library file management device
CN118643660B (en) Terminal block splicing method and system based on XML parsing
JP4167578B2 (en) Backup system, backup method and program
JPS6370372A (en) document processing device
CN1987802A (en) Basic input and output system information acquisition and editing method and system
CN121503433A (en) A document parsing method, system, and storage medium based on open-source components
JPH1165837A (en) Data exception detecting method for external file data and storage medium where data exception detection program for external file data is recorded