JPH10254894A - Similar document search device, similar document search method, and similar document search storage medium - Google Patents
Similar document search device, similar document search method, and similar document search storage mediumInfo
- Publication number
- JPH10254894A JPH10254894A JP9056723A JP5672397A JPH10254894A JP H10254894 A JPH10254894 A JP H10254894A JP 9056723 A JP9056723 A JP 9056723A JP 5672397 A JP5672397 A JP 5672397A JP H10254894 A JPH10254894 A JP H10254894A
- Authority
- JP
- Japan
- Prior art keywords
- word
- document
- comparison
- destination
- calculating
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Landscapes
- Information Retrieval, Db Structures And Fs Structures Therefor (AREA)
Abstract
(57)【要約】
【課題】 高精度の類似度計算処理を実現して文書デー
タの自動分類の効率向上を図る。
【解決手段】 比較元となる任意の文書及び外部記憶装
置に記憶された任意の比較先の文書から各々単語を抽出
し、抽出された各単語を、比較元の文書及び比較先の文
書の中での出現位置情報とともに各々個別に単語情報格
納バッファ4n、4pに格納し、比較元、比較先の双方
の文書に対して共通の特定の単語が存在するか否かを判
定し、特定の単語に関して比較元、比較先の双方の文書
に対して特定の単語から一定の距離内にある当該単語に
相当する単語を検索し、重み係数決定部4cにより、検
索した特定の単語から一定の距離内にある当該単語に相
当する単語に対して前記双方の文書における特定の単語
からの距離に応じた重み係数を算出し、類似度算出部4
bにより、重み係数決定部4cの算出結果を基にして、
前記比較元の文書に対する比較先の文書の単語の出現位
置に応じた類似度を算出するものである。
(57) [Summary] [PROBLEMS] To improve the efficiency of automatic classification of document data by realizing highly accurate similarity calculation processing. A word is extracted from an arbitrary document as a comparison source and an arbitrary comparison destination document stored in an external storage device, and each extracted word is extracted from the comparison source document and the comparison destination document. Are stored individually in the word information storage buffers 4n and 4p together with the appearance position information in the document, and it is determined whether or not there is a specific word common to both the comparison source and the comparison destination documents. With respect to both the source and destination documents, a word corresponding to the word within a certain distance from the specific word is searched for, and the weight coefficient determining unit 4c searches for a word within a certain distance from the searched specific word. , A weight coefficient corresponding to the distance from the specific word in both documents is calculated for the word corresponding to the word in
b, based on the calculation result of the weight coefficient determining unit 4c,
The similarity is calculated according to the appearance position of the word of the comparison destination document with respect to the comparison source document.
Description
【0001】[0001]
【発明の属する技術の分野】本発明は、類似文書検索装
置、類似文書検索方法及び類似文書検索用記憶媒体に関
する。[0001] 1. Field of the Invention [0002] The present invention relates to a similar document search device, a similar document search method, and a similar document search storage medium.
【0002】[0002]
【従来の技術】近年、大量の電子化された文書データが
流通するようになり、自動分類等を行なう目的で、文書
データベース中から指定された文書に類似するものを検
索する装置が実用化されてきている。この装置では指定
された文書(これを比較元文書と称する)の類似文書を
検索するにあたって、比較元文書と、データベース中の
各文書(これらを比較先文書と称する)との間で類似度
を計算し、この値の大きなものを検索結果として出力す
る。2. Description of the Related Art In recent years, a large amount of electronic document data has been distributed, and a device for searching a document database for a document similar to a designated document has been put to practical use for the purpose of automatic classification and the like. Is coming. In this apparatus, when searching for a similar document of a designated document (hereinafter referred to as a comparison source document), the similarity between the comparison source document and each document in the database (these are referred to as comparison destination documents) is determined. Calculate and output the one with the larger value as the search result.
【0003】類似度の計算方式としては、比較元文書、
比較先文書が共通した単語を多く含むほど類似度を大き
くするベクトル空間法が一般的に用いられている。[0003] As a method of calculating the similarity, a comparison source document,
A vector space method is generally used in which the similarity increases as the comparison destination document includes more common words.
【0004】このようなベクトル空間法では、まず、比
較元文書と比較先文書の双方から単語を抽出し、これら
から各単語の重みを要素としたベクトルデータを作成す
る。In such a vector space method, first, words are extracted from both a comparison source document and a comparison destination document, and vector data is created from them by using the weight of each word as an element.
【0005】各単語に対する重みとしては、各単語の文
書内での出現頻度を用い、その単語が存在する場合には
正の値が、存在しない場合には0のが設定される。As the weight for each word, the frequency of occurrence of each word in the document is used. When the word exists, a positive value is set, and when the word does not exist, 0 is set.
【0006】そして、ベクトル空間法では、比較元文書
と比較先文書の双方から抽出した単語双方の各ベクトル
データを計算した後、これらの内積を計算し、内積の値
を類似度としている。In the vector space method, vector data of both words extracted from both the source document and the destination document are calculated, and then the inner product is calculated, and the value of the inner product is used as the similarity.
【0007】[0007]
【発明が解決しようとする課題】上述したように、既存
のベクトル空間法による類似度計算方式によると、文書
の類似度を計算する際に、各単語について重み付けを行
なう必要があるが、例えば、単語Aと単語Bの両方が比
較元文書と比較先文書の双方に存在し、比較元の文書の
テキスト中で、単語Aと単語Bとが隣接あるいは接近し
ている場合において、比較先の文書のテキスト中にも単
語Aと単語Bとが隣接あるいは接近して存在したとして
も、これら単語Aと単語Bに対して、単語Aと単語Bと
が独立して存在している場合と同様の重み付けが行われ
ている。As described above, according to the similarity calculation method using the existing vector space method, it is necessary to weight each word when calculating the similarity of a document. When both word A and word B exist in both the source document and the destination document, and word A and word B are adjacent or close to each other in the text of the source document, the destination document Even if the word A and the word B are adjacent or close to each other in the text of the above, the word A and the word B are similar to the case where the word A and the word B exist independently. Weighting has been done.
【0008】このような複数の文書において、隣接ある
いは接近して現れる単語の組は文書の内容を決定づける
大きな特徴であるが、既存の類似度計算処理では、この
ような単語に対して重みを格別に大きく与えるような処
理を行なっておらず、単に画一化された出現頻度を重み
付けの要素としているのみであり、この結果、類似度計
算処理で十分な精度を得ることが困難となり、多数の文
書データの自動分類を行う際のユーザの負担軽減を図る
ことが難しいという課題があった。[0008] In such a plurality of documents, a set of words that appear adjacent or close to each other is a major feature that determines the content of the document. Is not performed, and only the uniformized appearance frequency is used as a weighting element. As a result, it becomes difficult to obtain sufficient accuracy in the similarity calculation processing, and There is a problem that it is difficult to reduce the burden on the user when performing automatic classification of document data.
【0009】本発明は上記事情に鑑みてなされたもので
あり、その目的とするところは、類似度を計算する過程
で、比較元文書で互いに近傍に存在する単語の組みを比
較先文書が含んでいる場合に、それらの単語の重みを大
きく設定できる機能を具備し、高精度の類似度計算処理
を実現して多数の文書データの自動分類の効率向上を図
ることができる類似文書検索装置、類似文書検索方法及
び類似文書検索用記憶媒体を提供することにある。The present invention has been made in view of the above circumstances. It is an object of the present invention to include, in a process of calculating a similarity, a set of words existing near each other in a comparison source document in a comparison destination document. A similar document search device that has a function of setting the weight of those words to be large, realizes highly accurate similarity calculation processing, and can improve the efficiency of automatic classification of a large number of document data, A similar document search method and a similar document search storage medium are provided.
【0010】[0010]
【課題を解決するための手段】請求項1記載の発明に係
る類似文書検索装置は、比較元となる任意の文書及び文
書記憶手段に記憶された任意の比較先の文書から各々単
語を抽出する抽出手段と、前記抽出手段により比較元の
文書及び比較先の文書から抽出された各単語を、比較元
の文書及び比較先の文書の中での出現位置情報とともに
各々個別に格納する単語位置情報格納手段と、この単語
位置情報格納手段を参照して比較元、比較先の双方の文
書に対して共通の特定の単語が存在するか否かを判定す
る判定手段と、この判定手段により判定された特定の単
語に関して単語位置情報格納手段を参照して比較元、比
較先の双方の文書に対して特定の単語から一定の距離内
にある当該単語に相当する単語を検索する検索手段と、
この検索手段により検索した特定の単語から一定の距離
内にある当該単語に相当する単語に対して比較元、比較
先の双方の文書における特定の単語からの距離に応じた
重み係数を算出する重み係数算出手段と、この重み係数
算出手段の算出結果を基に、前記比較元の文書に対する
比較先の文書の単語の出現位置に応じた類似度を算出す
る類似度算出手段とを有することを特徴とするものであ
る。According to a first aspect of the present invention, there is provided a similar document search apparatus which extracts words from an arbitrary document as a comparison source and an arbitrary comparison destination document stored in a document storage unit. Extraction means, and word position information for individually storing each word extracted from the comparison source document and the comparison destination document by the extraction means together with the appearance position information in the comparison source document and the comparison destination document A storage unit, a determination unit that determines whether a specific word common to both the comparison source and the comparison destination documents exists with reference to the word position information storage unit, Searching means for searching for a word corresponding to the specific word within a certain distance from the specific word with respect to both the comparison source and the comparison destination documents by referring to the word position information storage means for the specific word,
A weight for calculating a weight coefficient corresponding to a distance from a specific word in both the comparison source and the comparison destination documents for a word corresponding to the specific word within a certain distance from the specific word searched by the search means. Coefficient calculating means, and similarity calculating means for calculating, based on the calculation result of the weighting coefficient calculating means, a similarity corresponding to the word appearance position of the word of the comparison destination document with respect to the comparison source document. It is assumed that.
【0011】この発明によれば、抽出手段が比較元とな
る任意の文書及び文書記憶手段に記憶された任意の比較
先の文書から各々単語を抽出し、前記比較元の文書及び
比較先の文書から抽出された各単語を、比較元の文書及
び比較先の文書の中での出現位置情報とともに各々個別
に単語位置情報格納手段に格納し、判定手段が単語位置
情報格納手段を参照して比較元、比較先の双方の文書に
対して共通の特定の単語が存在するか否かを判定し、検
索手段が判定された特定の単語に関して単語位置情報格
納手段を参照して比較元、比較先の双方の文書に対して
特定の単語から一定の距離内にある当該単語に相当する
単語を検索し、重み係数算出手段により、検索した特定
の単語から一定の距離内にある当該単語に相当する単語
に対して比較元、比較先の双方の文書における特定の単
語からの距離に応じた重み係数を算出し、類似度算出手
段により、重み係数算出手段の算出結果を基にして、前
記比較元の文書に対する比較先の文書の単語の出現位置
に応じた類似度を算出するものである。According to the present invention, the extracting means extracts words from an arbitrary document to be compared and an arbitrary document to be compared stored in the document storage means, and the extraction source document and the compared document are extracted. Are stored individually in the word position information storage unit together with the appearance position information in the comparison source document and the comparison destination document, and the determination unit compares the words with reference to the word position information storage unit. It is determined whether or not a specific word common to both the source document and the comparison destination document exists, and the search unit refers to the word position information storage unit for the determined specific word, and compares the comparison source and the comparison destination. In both documents, a word corresponding to the word within a certain distance from the specific word is searched for, and the weight coefficient calculating means corresponds to the word corresponding to the word within a certain distance from the searched specific word. Source for words, A weight coefficient corresponding to the distance from a specific word in both documents of the comparison target is calculated, and the similarity calculation means calculates the weight coefficient according to the weight coefficient calculation means based on the calculation result of the comparison source document with respect to the comparison source document. Is calculated based on the appearance position of the word.
【0012】従って、比較元文書で互いに一定の距離内
に存在する単語の組みを比較先文書が含んでいる場合
に、それらの単語の重み係数の算出により文脈の類似度
まで考慮した精度の高い類似文書検索が可能となる。Therefore, when a set of words existing within a certain distance from each other in the comparison source document is included in the comparison destination document, the calculation of the weighting coefficients of those words allows for a high degree of accuracy taking into account the similarity of the context. Similar document search can be performed.
【0013】請求項2記載の発明に係る類似文書検索装
置は、請求項1記載の発明に係る類似文書検索装置にお
ける前記重み係数算出手段は、比較元、比較先の双方の
文書に対して特定の単語からの当該単語に相当する単語
の距離が小さいほど大きい値の重み係数とすることを特
徴とするものである。According to a second aspect of the present invention, there is provided a similar document search apparatus according to the first aspect, wherein the weighting factor calculating means specifies both the comparison source and the comparison destination documents. The weighting coefficient is set to a larger value as the distance of the word corresponding to the word from the word is smaller.
【0014】この発明によれば、特定の単語間との距離
が小さいほど大きい値の重み係数を付加して高精度の類
似文書検索が可能となる。According to the present invention, the smaller the distance from a specific word is, the larger the weighting factor is added, and a high-precision similar document search becomes possible.
【0015】請求項3記載の発明に係る類似文書検索装
置は、請求項1記載の発明に係る類似文書検索装置にお
いて、前記重み係数算出手段により得られた最大の重み
係数を用いて空間ベクトル法による単語ベクトル間の内
積の一つの項データを計算する項データ計算手段と、前
記最大の重み係数を用いて内積データの正規化を行う内
積データ正規化手段とを更に有し、項データ計算手段、
内積データ正規化手段の処理において前記重み係数算出
手段により得られた最大の重み係数を前記特定の単語の
出現回数に乗じる毎に、前記重み係数算出手段による重
み係数の算出処理を行うことを特徴とするものである。A similar document search device according to a third aspect of the present invention is the similar document search device according to the first aspect of the present invention, wherein the maximum weighting factor obtained by the weighting factor calculating means is used for a space vector method. Term data calculation means for calculating one term data of an inner product between word vectors according to the above, and inner product data normalizing means for normalizing the inner product data using the maximum weighting coefficient. ,
Each time the maximum weight coefficient obtained by the weight coefficient calculating means is multiplied by the number of appearances of the specific word in the processing of the inner product data normalizing means, the weight coefficient calculating means performs weight coefficient calculating processing. It is assumed that.
【0016】この発明によれば、上述したような類似度
を算出する際に、重み係数として従来例のような出現回
数のみを用いた場合と同様な処理が可能となり、簡略、
かつ、正確に内積演算及び正規化演算を行なうことが可
能となる。According to the present invention, when calculating the similarity as described above, the same processing as in the case where only the number of appearances as in the conventional example is used as the weighting factor can be performed.
In addition, the inner product operation and the normalization operation can be accurately performed.
【0017】請求項4記載の発明に係る類似文書検索方
法は、比較元となる任意の文書及び文書記憶手段に記憶
された任意の比較先の文書から各々単語を抽出し、前記
比較元の文書及び比較先の文書から抽出された各単語
を、比較元の文書及び比較先の文書の中での出現位置情
報とともに各々個別に単語位置情報格納手段に格納し、
この単語位置情報格納手段を参照して比較元、比較先の
双方の文書に対して共通の特定の単語が存在するか否か
を判定し、判定された特定の単語に関して単語位置情報
格納手段を参照して比較元、比較先の双方の文書に対し
て特定の単語から一定の距離内にある当該単語に相当す
る単語を検索し、重み係数算出手段により、検索した特
定の単語から一定の距離内にある当該単語に相当する単
語に対して比較元、比較先の双方の文書における特定の
単語からの距離に応じた重み係数を算出し、類似度算出
手段により、重み係数算出手段の算出結果を基にして、
前記比較元の文書に対する比較先の文書の単語の出現位
置に応じた類似度を算出することを特徴とするものであ
る。According to a fourth aspect of the present invention, in the similar document search method, a word is extracted from an arbitrary document to be compared and an arbitrary document to be compared stored in the document storage unit, and the document of the comparison source is extracted. And each word extracted from the comparison destination document is separately stored in the word location information storage means together with the appearance position information in the comparison source document and the comparison destination document,
By referring to the word position information storage means, it is determined whether or not a specific word common to both the comparison source document and the comparison destination document exists, and the word position information storage means is determined for the determined specific word. By referring to both the source and destination documents, a word corresponding to the word within a certain distance from the specific word is searched for, and a certain distance from the searched specific word is calculated by the weight coefficient calculating means. A weight coefficient corresponding to the distance from a specific word in both the comparison source document and the comparison destination document is calculated for a word corresponding to the word in the calculation result, and the similarity calculation means calculates the weight result of the weight coefficient calculation means. Based on
It is characterized in that a similarity is calculated in accordance with the appearance position of the word of the comparison destination document with respect to the comparison source document.
【0018】この発明によれば、請求項1記載の発明に
係る類似文書検索装置の構成を用いて、比較元文書で互
いに一定の距離内に存在する単語の組みを比較先文書が
含んでいる場合に、それらの単語の重み係数の算出によ
り文脈の類似度まで考慮した精度の高い類似文書検索の
方法を確立できる。According to this invention, using the configuration of the similar document search device according to the first aspect of the present invention, the comparison destination document includes a set of words existing within a certain distance from each other in the comparison source document. In this case, it is possible to establish a highly accurate similar document search method that takes into account the similarity of the context by calculating the weighting coefficients of those words.
【0019】請求項5記載の発明に係る類似文書検索用
記憶媒体は、比較元となる任意の文書及び文書記憶手段
に記憶された任意の比較先の文書から各々単語を抽出す
る手順と、前記比較元の文書及び比較先の文書から抽出
された各単語を、比較元の文書及び比較先の文書の中で
の出現位置情報とともに各々個別に単語位置情報格納手
段に格納する手順と、前記単語位置情報格納手段を参照
して比較元、比較先の双方の文書に対して共通の特定の
単語が存在するか否かを判定する手順と、判定された特
定の単語に関して単語位置情報格納手段を参照して比較
元、比較先の双方の文書に対して特定の単語から一定の
距離内にある当該単語に相当する単語を検索する手順
と、検索した特定の単語から一定の距離内にある当該単
語に相当する単語に対して比較元、比較先の双方の文書
における特定の単語からの距離に応じた重み係数を算出
する手順と、重み係数を算出する手順による算出結果を
基にして、前記比較元の文書に対する比較先の文書の単
語の出現位置に応じた類似度を算出する手順とからなる
プログラムを格納したことを特徴とするものである。A similar-document storage medium according to a fifth aspect of the present invention comprises the steps of: extracting a word from an arbitrary document as a comparison source and an arbitrary comparison-destination document stored in the document storage means; Storing each word extracted from the comparison source document and the comparison destination document together with the appearance position information in the comparison source document and the comparison destination document in the word position information storage unit; A procedure of referring to the position information storage means to determine whether or not a specific word common to both the document of the comparison source and the document of the comparison destination exists; and a step of determining the word position information storage means for the determined specific word. A procedure for searching for a word corresponding to the word within a certain distance from a specific word for both the reference source and the comparison target documents, and a method for searching for a word within a certain distance from the specific word searched. To the word equivalent to the word And calculating the weighting factor according to the distance from the specific word in the document of both the comparison source and the comparison destination, and comparing the comparison source document based on the calculation result by the procedure of calculating the weighting factor. A program for calculating a degree of similarity according to the appearance position of a word in the preceding document.
【0020】この発明によれば、この記憶媒体を既存の
コンピュータ装置等により読み込んで活用することによ
り、請求項1記載の発明に係る類似文書検索装置の場合
と同様にして、既存のコンピュータ装置等を、比較元文
書で互いに一定の距離内に存在する単語の組みを比較先
文書が含んでいる場合に、それらの単語の重み係数の算
出により文脈の類似度まで考慮した精度の高い類似文書
検索を行う装置として活用できる。According to the present invention, this storage medium is read and used by an existing computer device or the like, so that the existing computer device or the like can be used in the same manner as the similar document search device according to the first aspect of the present invention. When the comparison target document includes a set of words existing within a certain distance from each other in the comparison source document, a highly accurate similar document search that takes into account the similarity of the context by calculating the weighting coefficients of those words It can be used as a device for performing
【0021】[0021]
【発明の実施の形態】以下、図面を参照して、本発明の
実施の形態について説明する。Embodiments of the present invention will be described below with reference to the drawings.
【0022】図1は、本実施の形態装置のハードウェア
構成図である。図1に示すように本実施の形態装置は、
表示装置1、入力装置2、外部記憶装置3、制御装置
4、メモリ5、通信装置6から構成されており、各装置
はバスを介して接続されている。FIG. 1 is a hardware configuration diagram of the present embodiment. As shown in FIG.
It comprises a display device 1, an input device 2, an external storage device 3, a control device 4, a memory 5, and a communication device 6, and each device is connected via a bus.
【0023】表示装置1は、例えば、カラー液晶ディス
プレイ及びそのコントローラから構成されており、検索
結果の表示等を行なう。The display device 1 is composed of, for example, a color liquid crystal display and its controller, and displays search results and the like.
【0024】入力装置2は、キーボードやマウス等から
なり、制御装置4に対して各種のデータ及び命令の入力
を実行するものである。The input device 2 includes a keyboard, a mouse, and the like, and executes input of various data and commands to the control device 4.
【0025】外部記憶装置3は、ハードディスク及びコ
ントローラからなり、検索対象となる文書データや各文
書に含まれる単語のインデックスが格納される。The external storage device 3 comprises a hard disk and a controller, and stores document data to be searched and indexes of words included in each document.
【0026】本実施の形態装置の外部記憶装置3に格納
されている文書データ及び単語情報データの格納形式を
図2に示す。FIG. 2 shows the storage format of the document data and word information data stored in the external storage device 3 of this embodiment.
【0027】各文書データはタイトルデータ、作成日時
データ、作成者データ等の文書の各性を表わすヘッダ部
とテキストデータ部からなっており、これらがID番号
順に格納されている。Each document data is composed of a header part and a text data part representing titles, creation date / time data, creator data and the like of each document, and these are stored in the order of ID numbers.
【0028】また、単語情報データは各文書に対応付け
て格納されており、文書データのテキストデータに対し
て形態素解析を行ない名詞を抽出した結果が格納されて
いる。図4に一つの文書から作成した単語情報データの
例を元テキストと対応付けて示す。The word information data is stored in association with each document, and the result of extracting a noun by performing a morphological analysis on the text data of the document data is stored. FIG. 4 shows an example of word information data created from one document in association with the original text.
【0029】図4に示すように、単語情報データ中に
は、各単語のテキスト中での出現回数及び各出現位置が
格納される。As shown in FIG. 4, the number of appearances of each word in the text and each appearance position are stored in the word information data.
【0030】制御装置4は、CPUから構成されるもの
で、以上の各ハードウェア装置とバスにより接続されて
おり、各装置の制御、装置間のデータの転送等の処理を
行なうものである。The control device 4 is composed of a CPU, and is connected to each of the above hardware devices via a bus, and controls each device and performs processing such as data transfer between the devices.
【0031】メモリ5は、ダイナミックRAMからな
り、図3に示すように制御装置5が各種制御や処理を実
行するためのプログラムを格納するプログラム部と、処
理の際に必要なデータを格納するためのバッファ部から
なっている。The memory 5 is composed of a dynamic RAM, and as shown in FIG. 3, a program section for storing programs for the control device 5 to execute various controls and processes, and for storing data necessary for the processes. It consists of a buffer section.
【0032】メモリ5のプログラム部は、処理全体の制
御を行うメイン処理部4aの他、メイン処理部4aで呼
び出されるサブルーチンを処理するために、類似度算出
部4b、重み係数決定部4c,項データ加算部4d、内
積データ正規化部4e、文書一覧表示部4f、文書選択
部4g,文書内容表示部4hを具備している。The program section of the memory 5 includes a similarity calculating section 4b, a weight coefficient determining section 4c, and a term for processing a subroutine called by the main processing section 4a in addition to a main processing section 4a for controlling the entire processing. It has a data adder 4d, an inner product data normalizer 4e, a document list display 4f, a document selector 4g, and a document content display 4h.
【0033】また、バッファ部は、比較元文書データ格
納バッファ4m、比較元単語情報格納バッファ4n、比
較先単語情報格納バッファ4p、類似度格納バッファ4
qを具備している。The buffer unit includes a comparison source document data storage buffer 4m, a comparison source word information storage buffer 4n, a comparison destination word information storage buffer 4p, and a similarity storage buffer 4.
q.
【0034】ここで、比較元文書データ格納バッファ4
mは、図2に示す文書データを格納できる構造となって
いる。比較元単語情報格納バッファ4n及び比較先単語
情報格納バッファ4pは、図4で示す形式の単語情報デ
ータを格納できる構造となっている。Here, the comparison source document data storage buffer 4
m has a structure capable of storing the document data shown in FIG. The comparison source word information storage buffer 4n and the comparison destination word information storage buffer 4p have a structure capable of storing word information data in the format shown in FIG.
【0035】類似度格納バッファ4qは、各文書毎の類
似度を文書タイトル毎に格納できる構造となっている。
バッファ部は、この他各作業用変数のための領域4sを
具備している。The similarity storage buffer 4q has a structure in which the similarity of each document can be stored for each document title.
The buffer unit has an area 4s for each work variable.
【0036】この領域4s内には、後に述べる処理で用
いられる文書IDカウント用変数idDoc、カウント
用変数iSrc、iDst、kSrc、kDst、重み
係数変数coefSrc等が格納されている。This area 4s stores a document ID counting variable idDoc, a counting variable iSrc, iDst, kSrc, kDst, a weight coefficient variable coefSrc, and the like, which are used in the processing described later.
【0037】通信装置6は、通信回線を介して外部とデ
ータの送受を行なう装置であり、例えば、LAN回線と
LANコントローラ等から構成される。The communication device 6 transmits and receives data to and from the outside via a communication line, and includes, for example, a LAN line and a LAN controller.
【0038】次に、上記の構成要素を有する本実施の形
態装置の処理の流れについて、図5を参照して説明す
る。Next, the flow of processing of the apparatus having the above-described components according to this embodiment will be described with reference to FIG.
【0039】ここで、予め前記外部記憶装置3中には、
図2に示すように、先に示した形式の文書データ及び単
語情報データが、ID(識別)番号0からN一1のもの
まで合計N個格納されているものとする。Here, in the external storage device 3 in advance,
As shown in FIG. 2, it is assumed that a total of N pieces of document data and word information data in the above-described format from ID (identification) number 0 to N-11 are stored.
【0040】本実施の形態装置の全体の制御は、前記メ
モリ5のプログラム部に格納されたメイン処理部4aが
担当する。The overall control of the apparatus of this embodiment is performed by the main processing section 4a stored in the program section of the memory 5.
【0041】まず、ユーザが入力装置2より入力した文
書データ又は通信装置6により受信した文書データが、
比較元文書データ格納バッファ4mに格納される(ステ
ップ5a)。比較元文書データは、タイトル、作成日
時、作成者の文書の各属性を表わすヘッダ部とテキスト
データ部とからなっている。First, the document data input by the user from the input device 2 or the document data received by the communication device 6 is
It is stored in the comparison source document data storage buffer 4m (step 5a). The comparison source document data is composed of a header portion representing each attribute of the title, creation date and time, and the document of the creator, and a text data portion.
【0042】次に、比較元文書データ格納バッファ4m
に格納した比較元文書データのテキストデータ部に対し
て、メイン処理部4aが予め格納しているプログラムに
基づいて形態素解析を行ない、テキストデータ部に含ま
れ名詞を抽出し(ステップ5b)、比較元単語情報デー
タを作成し、比較元単語情報格納バッファ4nに格納す
る(ステップ5c)。Next, the comparison source document data storage buffer 4m
The main processing unit 4a performs a morphological analysis on the text data part of the comparison source document data stored in the text data part based on a program stored in advance, and extracts nouns included in the text data part (step 5b). Original word information data is created and stored in the comparison source word information storage buffer 4n (step 5c).
【0043】比較元単語情報データの構造は、先に図4
に示した単語情報データの構造と同様であり、各単語の
出現回数及びテキスト中での出現位置の各情報が格納さ
れている。The structure of the comparison source word information data is as shown in FIG.
The information has the same structure as that of the word information data shown in (a), and stores information on the number of appearances of each word and the appearance position in the text.
【0044】次に、メイン処理部4aは、文書IDカウ
ント用変数idDocに0を格納する(ステップ5
d)。続いて、文書IDカウント用変数idDocに対
応する外部記憶装置3中の単語情報データを比較先単語
情報格納バッファ4pに格納する(ステップ5e)。Next, the main processing section 4a stores 0 in the document ID counting variable idDoc (step 5).
d). Subsequently, the word information data in the external storage device 3 corresponding to the document ID counting variable idDoc is stored in the comparison target word information storage buffer 4p (step 5e).
【0045】続いて類似度算出部4bが起動し、比較元
単語情報格納バッファ4nに格納されている比較元文書
データの内容と、比較先単語情報格納バッファ4pに格
納されている比較元文書データとを参照して類似度を算
出する(ステップ5f)。Subsequently, the similarity calculation unit 4b is activated, and the contents of the comparison source document data stored in the comparison source word information storage buffer 4n and the comparison source document data stored in the comparison destination word information storage buffer 4p Is calculated with reference to (step 5f).
【0046】以下にステップ5eの類似度算出部4bで
の処理の詳細を、図6を参照して説明する。The details of the processing performed by the similarity calculating section 4b in step 5e will be described below with reference to FIG.
【0047】ここで、比較元文書データに格納されてい
るi番目の単語の見出し語をwordSrc[i]、そ
の出現回数をnApSrc[i]、そのj番目の出現位
置をposSrc[i][j]、比較先の単語情報デー
タに格納されているi番目の単語の見出し語をword
Dst[i]、その出現回数をnApDst[i]、そ
のj番目の出現位置をposDst[i][j]と表わ
すことにする。Here, the headword of the i-th word stored in the comparison source document data is wordSrc [i], the number of occurrences is nApSrc [i], and the j-th occurrence position is posSrc [i] [j ], The headword of the i-th word stored in the word information data of the comparison destination is word
Dst [i], the number of appearances thereof will be represented by nApDst [i], and the jth occurrence position thereof will be represented by posDst [i] [j].
【0048】また、比較元の各単語ベクトル毎の重み係
数を、実数変数cofeSrc[iSrc]、比較元文
書データに格納されている単語の総数をnSrC、比較
先の単語情報データに格納されている単語の総数をnD
stで表わすことにする。The weight coefficient for each word vector of the comparison source is stored in the real variable coffSrc [iSrc], the total number of words stored in the comparison source document data is nSrC, and the word information data of the comparison destination is stored. ND the total number of words
It is represented by st.
【0049】さらに、カウント用変数iSrc、iDs
t、kSrc、kDstをメモリ5の作業用変数のため
の領域4s内に用意し、以下の処理を実行する。Further, count variables iSrc, iDs
t, kSrc, and kDst are prepared in an area 4s for a working variable in the memory 5, and the following processing is executed.
【0050】まず、類似度格納変数Sに0を格納し初期
化する(ステップ6a)。First, 0 is stored in the similarity storage variable S and initialized (step 6a).
【0051】以下、ステップ6bからステップ6dまで
の処理で、比較元文書及び比較先文書に共通に含まれる
単語を見つける。Hereinafter, words commonly included in the comparison source document and the comparison destination document are found in the processing from step 6b to step 6d.
【0052】まず、変数iSrcに0を代入する(ステ
ップ6b)。次に、実数変数cofeSrc[iSr
c]に1.0を代入する(ステップ6c)。次に、変数
iDstに0を代入する(ステップ6d)。次に、実数
変数cofeDst[iDst]に1.0を代入する
(ステップ6e)。First, 0 is substituted for the variable iSrc (step 6b). Next, the real variable cofeSrc [iSr
c] is substituted with 1.0 (step 6c). Next, 0 is substituted for the variable iDst (step 6d). Next, 1.0 is substituted for the real variable cofeDst [iDst] (step 6e).
【0053】次に、類似度算出部4bはwordSrc
[iSrc]とwordSt[iDst]が等しいか否
かその文字列内容を比較する(ステップ6f)。両者が
等しかった場合にはステップ6gに制御が移り、等しく
なかったなら、ステップ6iに制御が移る。Next, the similarity calculation unit 4b uses wordSrc
The character string content is compared to determine whether [iSrc] is equal to wordSt [iDst] (step 6f). If they are equal, control is transferred to step 6g, and if they are not equal, control is transferred to step 6i.
【0054】ステップ6gでは、重み係数決定部4cが
起動し、比較先の単語ベクトルの一つつの要素に係る重
み係数変数coefSrc[iSrc]を決定する。こ
こでの重み係数決定処理の詳細については後に図7を参
照して説明する。In step 6g, the weighting factor determination unit 4c is activated, and determines a weighting factor variable coefSrc [iSrc] related to one element of the word vector to be compared. The details of the weight coefficient determination process will be described later with reference to FIG.
【0055】次に、項データ加算部4dが起動し、ステ
ップ6gで決定したcoefSrc[iSrc]の値を
参照して類似度格納変数Sに単語ベクトルの内積の一つ
の項データの加算を行なう(ステップ6h)。Next, the term data adder 4d is activated, and adds one term data of the inner product of the word vector to the similarity degree storage variable S with reference to the value of coefSrc [iSrc] determined in step 6g ( Step 6h).
【0056】具体的には、項データ加算部4dでの加算
処理は下記数1の計算式に従って行なうものである。More specifically, the addition processing in the term data addition section 4d is performed according to the following equation (1).
【0057】[0057]
【数1】 (Equation 1)
【0058】ここで、nApSrc[iSrc]は、比
較元単語の出現回数、nApDst[iDst]は、比
較先単語の出現回数であり、これらは一般的な内積計算
で重みとして用いられているものである。Here, nApSrc [iSrc] is the number of appearances of the comparison source word, and nApDst [iDst] is the number of appearances of the comparison target word, which are used as weights in general inner product calculation. is there.
【0059】これらの積に、先に求めた重み係数変数c
oefSrc[iSrc]を乗じている。The product of these is added to the weight coefficient variable c obtained earlier.
oefSrc [iSrc].
【0060】加算処理を終えた後、処理はステップ6i
に移る。ステップ6iでは、変数iDstの値に、1を
加える。ここで変数iDstの値が総数nDStより小
さければ(ステップ6j)、ステップ6fに戻り、処理
を繰り返す。変数iDstの値が総数nDStより小さ
くなければ、変数iSrcの値に1を加える(ステップ
6k)。After the addition process is completed, the process proceeds to step 6i.
Move on to In step 6i, 1 is added to the value of the variable iDst. Here, if the value of the variable iDst is smaller than the total number nDSt (step 6j), the process returns to step 6f and repeats the processing. If the value of the variable iDst is not smaller than the total number nDSt, 1 is added to the value of the variable iSrc (step 6k).
【0061】次に、変数iSrcの値が総数nSrcよ
り小さければステップ6eに戻り処理を繰り返す。ステ
ップ6kでの処理で変数iSrcの値が総数nSrc以
上であれば、内積データ正規化部4eが起動し、類似度
格納変数Sの値を比較元、比較先の各単語ベクトルの大
きさの積で割る(ステップ6m)正規化処理を行う(ス
テップ6m)。Next, if the value of the variable iSrc is smaller than the total number nSrc, the process returns to step 6e and repeats the processing. If the value of the variable iSrc is equal to or greater than the total number nSrc in the processing in step 6k, the inner product data normalizing unit 4e is activated, and the value of the similarity storage variable S is calculated as the product of the sizes of the word vectors of the comparison source and the comparison destination. (Step 6m) to perform a normalization process (Step 6m).
【0062】具体的に内積データ正規化部4eでの処理
で用いる計算式を下記数2に示す。The formula used in the process of the inner product data normalizing section 4e is shown in the following equation (2).
【0063】[0063]
【数2】 (Equation 2)
【0064】比較元の単語ベクトルの大きさは、比較元
の単語の出現回数nApSrc[iSrc]と重み係数
変数coefSrc[iSrc]との積の前記変数iS
rc分の和の平方根、比較先の単語ベクトルの大きさ
は、比較先の単語の出現回数nApDst[iDst]
の2乗の前記変数iDst分の和の平方根で求められ
る。The magnitude of the word vector of the comparison source is determined by the variable iS of the product of the number of appearances nApSrc [iSrc] of the word of the comparison source and the weight coefficient variable coefSrc [iSrc].
The square root of the sum of rc and the size of the word vector of the comparison destination are represented by the number of appearances nApDst [iDst] of the word of the comparison destination
Is obtained by the square root of the sum of the square of the variable iDst.
【0065】上述した内積データ正規化部4eでの処理
を終了した後、図5のステップ5fで示す類似度計算処
理を終える(ステップ6n)。After finishing the processing in the inner product data normalizing section 4e described above, the similarity calculation processing shown in step 5f of FIG. 5 is finished (step 6n).
【0066】次に、前記重み係数決定部4cの重み係数
決定処理(ステップ7a)について図7を参照して説明
する。Next, the weight coefficient determination processing (step 7a) of the weight coefficient determination section 4c will be described with reference to FIG.
【0067】まず、以下に説明するステップ7bからス
テップ7pまでの処理で比較元の単語wordSrc
[iSrc]の近傍に存在し、かつ、比較先の単語wo
rdDst[iDst]の近傍にも存在している単語が
あるかどうか調べ、存在している場合には、比較元の単
語の重み係数変数coefSrc[iSrc]の値を最
小となるように再設定する。First, in the processing from step 7b to step 7p described below, the word wordSrc of the comparison source is used.
Word wo that exists near [iSrc] and is compared
It is checked whether there is a word existing near rdDst [iDst], and if it exists, the value of the weight coefficient variable coefSrc [iSrc] of the word to be compared is reset so as to be minimum. .
【0068】即ち、まず、比較元の単語を表わす変数k
Srcに0を代入する(ステップ7b)。この後、変数
iSrcと変数kSrcとが同じなら(ステップ7
c)、同一単語を表わすので制御はステップ7sに移
る。変数iSrcと変数kSrcとが異なる場合には、
単語wordSrc[iSrc]と単語wordSrc
[kSrc]とが等しいか否か比較する(ステップ7
d)。That is, first, the variable k representing the word to be compared
Substitute 0 for Src (step 7b). Thereafter, if the variable iSrc and the variable kSrc are the same (step 7
c) Since the same word is represented, the control moves to step 7s. When the variable iSrc and the variable kSrc are different,
Word wordSrc [iSrc] and word wordSrc
Compare whether [kSrc] is equal or not (step 7)
d).
【0069】両者が等しかったら、制御はステップ7e
に移り、等しくなかったらステップ7sに移る。If both are equal, control is passed to step 7e.
If not, the process proceeds to step 7s.
【0070】ステップ7eでは、テキスト中に現れる単
語wordSrc[iSrc]と単語wordSrc
[kSrc]のすべての出現位置を参照して、その出現
位置の差の絶対値(最小値)minSrcを求める(ス
テップ7e)。In step 7e, the words wordSrc [iSrc] appearing in the text and the words wordSrc
With reference to all the appearance positions of [kSrc], the absolute value (minimum value) minSrc of the difference between the appearance positions is obtained (step 7e).
【0071】ここでの具体的な処理については、後に図
8を参照して説明する。The specific processing here will be described later with reference to FIG.
【0072】この後、ステップ7eで求めた最小値が、
例えば80(80文字)以下なら(ステップ7f)、ス
テップ7gに制御を移し、そうでなければステップ7s
に制御を移す。続くステップ7gでは比較元の単語を表
わすkDstに0を代入する(ステップ7g)。Thereafter, the minimum value obtained in step 7e is
For example, if it is equal to or less than 80 (80 characters) (step 7f), the control is shifted to step 7g, and if not, step 7s
Transfer control to. In the following step 7g, 0 is substituted for kDst representing the word of the comparison source (step 7g).
【0073】この後、iDstとkDstが同じなら
(ステップ7h)、同一単語を表わすので制御はステッ
プ7qに移る。iDStとkDStが異なるなら(ステ
ップ7h)、単語wordDst[iDst]と単語w
ordDst[kDst]とが等しいか比較する(ステ
ップ7i)。Thereafter, if iDst and kDst are the same (step 7h), they represent the same word, and therefore control transfers to step 7q. If iDSt and kDSt are different (step 7h), the word wordDst [iDst] and the word w
It is compared whether ordDst [kDst] is equal (step 7i).
【0074】両者が等しかったら制御はステップ7jに
移り、等しくなかったらステップ7qに移る。If they are equal, the control moves to step 7j, and if they are not equal, the control moves to step 7q.
【0075】ステップ7jでは、テキスト中に現れる単
語wordDst[iDst]と単語wordst[k
Dst]の全ての出現位置を参照して、その出現位置の
差の絶対値(最小値)minDstを求める(ステップ
7j)。In step 7j, the words wordDst [iDst] and the words wordst [k
Dst], the absolute value (minimum value) minDst of the difference between the appearance positions is determined (step 7j).
【0076】ここでの具体的な処理については後に図9
を参照して説明する。The specific processing here will be described later with reference to FIG.
This will be described with reference to FIG.
【0077】この後、ステップ7jで求めた最小値mi
nDstが80以下なら(ステップ7k)、ステップ7
mに制脚を移し、そうでなければステップ7qに制脚を
移す。Thereafter, the minimum value mi obtained in step 7j
If nDst is 80 or less (step 7k), step 7
m, otherwise move to step 7q.
【0078】ステップ7mでは、先に求めたminSr
c及びminDstの値を参照して、minSrc及び
minDstの値が小さいほど大きくなるように重み係
数の実数変数coefの値を計算する(ステップ7
m)。In step 7m, the minSr obtained earlier
Referring to the values of c and minDst, the value of the real number variable coef of the weight coefficient is calculated so that the smaller the values of minSrc and minDst, the larger the value (step 7).
m).
【0079】ここで具体的には、以下の数3の計算式に
より実数変数coefの値を求める。Here, specifically, the value of the real number variable coef is obtained by the following equation (3).
【0080】[0080]
【数3】 (Equation 3)
【0081】尚、実数変数coefの値は、1以下では
不可であり、従って、数3のminSrc/80、mi
nDst/80の値は、各々最大でも1に設定される。It should be noted that the value of the real number variable coef cannot be less than 1; therefore, minSrc / 80, mi
The value of nDst / 80 is set to 1 at the maximum.
【0082】次に、ここで求めた実数変数coefの値
が、既に求めてある前記変数coefSrc[iSr
c]より大きいか否か比較する(ステップ7n)。Next, the value of the real number variable coef obtained here is compared with the previously obtained variable coefSrc [iSr
c] is compared (step 7n).
【0083】実数変数coefの値が、既に求めてある
変数coefSrc[iSrc]より大きい場合には、
新たな最大のcoefSrc[iSrc]の値としてス
テップ7mで求めたcoefの値を採用する(ステップ
7p)。続いて、ステップ7qでは、変数kDstの値
をインクリメントし(ステップ7q)、その値が比較先
文書の単語総数nDstより小さいか比較し(ステップ
7r)、小さかったならステップ7hからの処理を繰り
返し、そうでなければステップ7sに制御を移す。If the value of the real number variable coef is larger than the already obtained variable coefSrc [iSrc],
The value of coef obtained in step 7m is adopted as the new maximum value of coefSrc [iSrc] (step 7p). Subsequently, in step 7q, the value of the variable kDst is incremented (step 7q), and it is compared whether the value is smaller than the total number of words nDst of the comparison target document (step 7r). If smaller, the processing from step 7h is repeated. Otherwise, control is transferred to step 7s.
【0084】ステップ7sでは、変数kSrcの値をイ
ンクリメントし、その値が比較先文書の単語総数nSr
cより小さいか否か比較し(ステップ7t)、小さい場
合には、ステップ7cからの処理を繰り返し。そうでな
ければ重み係数決定処理を終える(ステップ7u)。In step 7s, the value of the variable kSrc is incremented, and the value is set to the total number of words nSr in the comparison target document.
It is compared whether it is smaller than c (step 7t), and if it is smaller, the processing from step 7c is repeated. Otherwise, the weight coefficient determination processing ends (step 7u).
【0085】次に、上記のステップ7eの処理で、比較
元のテキスト中に現れる単語wordSrc[iSr
c]と、単語wordSrc[kSrc]の全ての出現
位置を参照して、その出現位置の差の絶対値minSr
cを求める処理の詳細について、図8を参照して説明す
る。Next, in the process of step 7e, the word wordSrc [iSr
c] and the absolute value minSr of the difference between the appearance positions with reference to all the appearance positions of the word wordSrc [kSrc].
The details of the process for obtaining c will be described with reference to FIG.
【0086】この処理では、まず、j0=0とし(ステ
ップ8a)、ステップ8bからステップ8jまでのルー
プで、j0の値を0から、nApSrc[iSrc]ま
で変化させ、このループの内部のステップ8cからステ
ップ8hまでのループでj1の値を0からnApSrc
[kSrc]まで変化させ、j0とj1の全ての組み合
わせに対して、kSrcで表わされる単語の出現位置p
osSrc[kSrc][jo]、iSrcで表わされ
る単語の出現位置posSrc[iSrc][j1]の
距離、つまり、下記数4で表される差の絶対値dSrc
を計算する(ステップ8c)。In this process, first, j0 is set to 0 (step 8a), and in the loop from step 8b to step 8j, the value of j0 is changed from 0 to nApSrc [iSrc]. In the loop from to 8h, the value of j1 is changed from 0 to nApSrc
[KSrc], and for all combinations of j0 and j1, the appearance position p of the word represented by kSrc
osSrc [kSrc] [jo], the distance of the occurrence position posSrc [iSrc] [j1] of the word represented by iSrc, that is, the absolute value dSrc of the difference represented by the following equation (4)
Is calculated (step 8c).
【0087】[0087]
【数4】 (Equation 4)
【0088】さらに、ステップ8dからステップ8fの
処理によって、初期状態の場合(j0=0、かつ、j1
=0)(ステップ8d)、あるいは以前の処理で既に求
めた距離の最小値minSrcよりステップ8cで求め
たdSrcの値が小さい場合(ステップ8e)には、新
たなminSrcの値としてステップ8cで求めたdS
rcの値を採用する(ステップ8f)。Further, in the case of the initial state (j0 = 0 and j1
= 0) (step 8d), or when the value of dSrc obtained in step 8c is smaller than the minimum value minSrc already obtained in the previous processing (step 8e), a new minSrc value is obtained in step 8c. DS
The value of rc is adopted (step 8f).
【0089】これらの処理が終了したら、処理を終了す
る(ステップ8k)。When these processes are completed, the process ends (step 8k).
【0090】次に、上記のステップ7jで、比較先のテ
キスト中に現れる単語wordDst[iDst]と単
語wordDst[kDst]の全ての出現位置を参照
して、その出現位置の差の絶対値minDstを求める
処理の詳細について図9を参照して説明する。Next, in step 7j, the absolute value minDst of the difference between the appearance positions is referred to by referring to all the appearance positions of the words wordDst [iDst] and the word wordDst [kDst] appearing in the text to be compared. The details of the process to be determined will be described with reference to FIG.
【0091】この処理では、まず、j0=0とし(ステ
ップ9a)、ステップ9bからステップ9jまでのルー
プで、j0の値を0からnApDst[iDst]まで
変化させ、このループの内部のステップ9cからステッ
プ9hまでのループでj1の値を0からnApDst
[kDst]まで変化させ、j0とj1の全ての組み合
わせに対して、kDstで表わされる単語の出現位置p
osDst[kDst][j0]、iDstで表わされ
る単語の出現位置posDst[iDst][j1]の
距離、つまり、下記数5で表される差の絶対値を計算す
る(ステップ9c)。In this processing, first, j0 = 0 (step 9a), the value of j0 is changed from 0 to nApDst [iDst] in a loop from step 9b to step 9j, and step 9c inside this loop is executed. In the loop up to step 9h, the value of j1 is changed from 0 to nApDst
[KDst], and for all combinations of j0 and j1, the appearance position p of the word represented by kDst
The distance between the occurrence position posDst [iDst] [j1] of the word represented by osDst [kDst] [j0] and iDst, that is, the absolute value of the difference represented by the following equation 5 is calculated (step 9c).
【0092】[0092]
【数5】 (Equation 5)
【0093】さらに、ステップ9dからステップ9fの
処理によって、初期状態の場合(j0=0、かつ、j1
=0)(ステップ9d)、あるいは以前の処理で既に求
めた距離の最小値minDstよりステップ9cで求め
たdDstの値が小さい場合(ステップ9eでの比較が
成功した場合)には、新たなminDstの値としてス
テップ9cで求めたdDstの値を採用する(ステップ
9f)。Further, by the processing of steps 9d to 9f, the case of the initial state (j0 = 0 and j1
= 0) (step 9d), or when the value of dDst obtained in step 9c is smaller than the minimum value minDst of the distance already obtained in the previous processing (when the comparison in step 9e is successful), a new minDst The value of dDst obtained in step 9c is adopted as the value of (step 9f).
【0094】これらの処理が終了したら、処理を終了す
る(ステップ9k)。When these processes are completed, the process ends (step 9k).
【0095】次に、上述したようにステップ5fにおい
て類似度を算出した後、図5に示すように、ここで得た
類似度を類似度格納バッファ4qのidDocに対応す
る位置に格納する(ステップ5g)。Next, after calculating the similarity in step 5f as described above, as shown in FIG. 5, the obtained similarity is stored in the similarity storage buffer 4q in a position corresponding to idDoc (step 5f). 5g).
【0096】次にidDocの値をインクリメントする
(ステップ5h)。さらに、idDocの値と外部記憶
装置3中に格納されている文書の総数Nとを比較し、i
dDocが総数Nより小さい場合には、ステップ5eか
らの処理を繰り返し、そうでなければステップ5jに制
御を移す(ステップ5i)。Next, the value of idDoc is incremented (step 5h). Further, the value of idDoc is compared with the total number N of documents stored in the external storage device 3, and i
If dDoc is smaller than the total number N, the processing from step 5e is repeated, otherwise, control is transferred to step 5j (step 5i).
【0097】ステップ5jでは、前記文書一覧表示部4
fが起動する。文書一覧表示部4fは、類似度格納バッ
ファ4qの内容及び各文書データ中のタイトルデータを
参照して、類似度の大きな文書から順(類似度の値が小
さいものから類似度の値が大きいものの順)に、その文
書タイトルの一覧を対応する類似度(類似度1、類似度
2、…)情報とともに前記表示装置1の画面上に表示す
る(ステップ5j)。In step 5j, the document list display section 4
f is activated. The document list display unit 4f refers to the contents of the similarity storage buffer 4q and the title data in each document data, and starts from the document with the highest similarity (from the one with the lowest similarity to the one with the highest similarity). Then, the list of the document titles is displayed on the screen of the display device 1 together with the corresponding similarity (similarity 1, similarity 2,...) Information (step 5j).
【0098】この場合の文書タイトルー覧表示状態の表
示装置1の画面の状況を図10に示す。FIG. 10 shows the state of the screen of the display device 1 in the document title-list display state in this case.
【0099】続いて、ステップ5kにおいて、文書選択
部4gが起動する。文書選択部4gは、入力装置2を用
いて、画面上に表示されている文書タイトルの一つ又は
複数をユーザに選択させる。Subsequently, in step 5k, the document selection section 4g is activated. The document selection unit 4g allows the user to select one or a plurality of document titles displayed on the screen using the input device 2.
【0100】次に文書内容表示部4hが起動し、ステッ
プ5kで選択された文書タイトルに対応する文書の内容
が表示される(ステップ5m)。これにより、ユーザ
は、自己の選択に応じて、所望の類似度を持った文書の
内容を画面上に表示して確認できる。以上により処理が
終了する(ステップ5n)。Next, the document content display section 4h is activated, and the content of the document corresponding to the document title selected in step 5k is displayed (step 5m). Thereby, the user can display and confirm the contents of the document having the desired similarity on the screen according to his / her selection. Thus, the process ends (step 5n).
【0101】以上が、本実施の形態装置での処理の流れ
である。The above is the flow of processing in the present embodiment.
【0102】尚、本発明は上記の実施の形態に限定され
るものではない。The present invention is not limited to the above embodiment.
【0103】例えば、項データ加算部4dのステップ6
fの加算処理では、結果に重み係数変数の値が反映され
ていれば実施の形態に示した計算式を用いなくても良
い。For example, step 6 of the term data adder 4d
In the addition processing of f, the calculation formula shown in the embodiment may not be used as long as the value of the weight coefficient variable is reflected in the result.
【0104】また、重み係数決定部4cのステップ7m
で用いる計算式としては、minSrcあるいはmin
Dstの値が小さいほど重み係数が大きくなる性質を持
つものであれば本実施の形態で示したものでなくても良
い。Step 7m of the weight coefficient determining unit 4c
The calculation formula used in minSrc or minSrc
The present invention need not be the one shown in the present embodiment as long as the weight coefficient increases as the value of Dst decreases.
【0105】さらに、図7に示すステップ7f及び7h
では距離の比較の際に、80という値を用いたが、これ
も他の適当な値に設定することが可能である。Further, steps 7f and 7h shown in FIG.
In the above, the value of 80 was used in comparing the distances, but this value can also be set to another appropriate value.
【0106】この他、本発明の趣旨を逸脱しない範囲で
種々の変形が可能である。In addition, various modifications can be made without departing from the spirit of the present invention.
【0107】[0107]
【発明の効果】以上説明した本発明によれば、以下の効
果を奏する。According to the present invention described above, the following effects can be obtained.
【0108】請求項1記載の発明によれば、比較元文書
で互いに一定の距離内に存在する単語の組みを比較先文
書が含んでいる場合に、それらの単語の重み係数の算出
により文脈の類似度まで考慮した精度の高い類似文書検
索が可能な類似文書検索装置を提供することができる。According to the first aspect of the present invention, in the case where a set of words existing within a certain distance from each other in a comparison source document is included in a comparison destination document, the weight coefficient of those words is calculated to determine the context. A similar document search device capable of performing a similar document search with high accuracy in consideration of similarity can be provided.
【0109】請求項2記載の発明によれば、特定の単語
間との距離が小さいほど大きい値の重み係数を付加して
高精度の類似文書検索が可能な類似文書検索装置を提供
することができる。According to the second aspect of the present invention, it is possible to provide a similar document search apparatus capable of performing a high-accuracy similar document search by adding a larger weighting factor as the distance between specific words becomes smaller. it can.
【0110】請求項3記載の発明によれば、類似度を算
出する際に、重み係数として従来例のような出現回数の
みを用いた場合と同様な処理が可能となり、簡略、か
つ、正確に内積演算及び正規化演算を行なうことが可能
な類似文書検索装置を提供することができる。According to the third aspect of the present invention, when calculating the similarity, it is possible to perform the same processing as in the case where only the number of appearances as in the conventional example is used as the weighting coefficient, and it is simple and accurate. A similar document search device capable of performing inner product operation and normalization operation can be provided.
【0111】請求項4記載の発明によれば、請求項1記
載の発明に係る類似文書検索装置の構成を用いて、比較
元文書で互いに一定の距離内に存在する単語の組みを比
較先文書が含んでいる場合に、それらの単語の重み係数
の算出により文脈の類似度まで考慮した精度の高い類似
文書検索の方法を確立できる類似文書検索方法を提供す
ることができる。According to the fourth aspect of the present invention, by using the configuration of the similar document search apparatus according to the first aspect of the present invention, a set of words existing within a certain distance from each other in the comparison source document is compared with the comparison destination document. , It is possible to provide a similar document search method capable of establishing a highly accurate similar document search method that takes into account the similarity of context by calculating weighting coefficients of those words.
【0112】請求項5記載の発明によれば、既存のコン
ピュータ装置等を、比較元文書で互いに一定の距離内に
存在する単語の組みを比較先文書が含んでいる場合に、
それらの単語の重み係数の算出により文脈の類似度まで
考慮した精度の高い類似文書検索を行う装置として活用
できる類似文書検索用記憶媒体を提供することができ
る。According to the fifth aspect of the present invention, when an existing computer device or the like includes a set of words existing within a certain distance from each other in a comparison source document, the comparison destination document includes:
It is possible to provide a similar document search storage medium that can be used as a device for performing a highly accurate similar document search that takes into account the similarity of context by calculating the weighting coefficients of those words.
【図1】本実施の形態装置のハードウェアの構成を示し
たブロック図である。FIG. 1 is a block diagram illustrating a hardware configuration of a device according to an embodiment.
【図2】本実施の形態装置の文書データ及び単語情報デ
ータの格納形式を示した説明図である。FIG. 2 is an explanatory diagram showing a storage format of document data and word information data of the present embodiment.
【図3】本実施の形態装置のメモリの構成を示した図で
ある。FIG. 3 is a diagram showing a configuration of a memory of the present embodiment device.
【図4】本実施の形態装置の単語情報データの例を元テ
キストと対応付けて示した図である。FIG. 4 is a diagram showing an example of word information data of the present embodiment device in association with an original text.
【図5】本実施の形態装置の全体の処理の流れを示した
フローチャートである。FIG. 5 is a flowchart showing a flow of overall processing of the apparatus according to the embodiment.
【図6】本実施の形態装置の類似度算出部での処理の詳
細を示したフローチャートである。FIG. 6 is a flowchart illustrating details of a process performed by a similarity calculation unit of the embodiment.
【図7】本実施の形態装置の重み係数決定部の処理を示
したフローチャートである。FIG. 7 is a flowchart showing a process of a weight coefficient determination unit of the present embodiment.
【図8】本実施の形態装置の比較元テキスト中に現れる
単語間の距離を求める処理を示したフローチャートであ
る。FIG. 8 is a flowchart illustrating a process of calculating a distance between words appearing in a comparison source text according to the present embodiment.
【図9】本実施の形態装置の比較先テキスト中に現れる
単語間の距離を求める処理を示したフローチャートであ
る。FIG. 9 is a flowchart illustrating a process of calculating a distance between words appearing in a comparison target text according to the present embodiment.
【図10】本実施の形態装置のタイトルー覧表示の場合
の表示装置の画面を示す図である。FIG. 10 is a diagram showing a screen of the display device in the case of title-view display of the present embodiment device.
1 表示装置 2 入力装置 3 外部記憶装置、 4 制御装置 4a メイン処理部 4b 類似度算出部 4c 重み係数決定部 4d 項データ加算部 4e 内積データ正規化部 4f 文書一覧表示部 4g 文書選択部 4h 文書内容表示部 4m 比較元文書データ格納バッファ 4n 比較元単語情報格納バッファ 4p 比較先単語情報格納バッファ 4q 類似度格納バッファ 4s 各種作業用変数のための領域 5 メモリ 6 通信装置 Reference Signs List 1 display device 2 input device 3 external storage device, 4 control device 4a main processing unit 4b similarity calculation unit 4c weight coefficient determination unit 4d term data addition unit 4e inner product data normalization unit 4f document list display unit 4g document selection unit 4h document Content display section 4m Comparison source document data storage buffer 4n Comparison source word information storage buffer 4p Comparison destination word information storage buffer 4q Similarity storage buffer 4s Area for various work variables 5 Memory 6 Communication device
───────────────────────────────────────────────────── フロントページの続き (72)発明者 中本 幸夫 東京都青梅市新町1381番地1 東芝コンピ ュータエンジニアリング株式会社内 (72)発明者 仁科 卓哉 東京都青梅市新町1381番地1 東芝コンピ ュータエンジニアリング株式会社内 (72)発明者 久保田 直秀 東京都青梅市新町1381番地1 東芝コンピ ュータエンジニアリング株式会社内 ──────────────────────────────────────────────────続 き Continuing from the front page (72) Inventor Yukio Nakamoto 1381-1, Shinmachi, Ome-shi, Tokyo Toshiba Computer Engineering Co., Ltd. (72) Takuya Nishina 1381-1, Shinmachi, Ome-shi, Tokyo Toshiba Computer (72) Inventor Naohide Kubota 1381 Shinmachi, Ome-shi, Tokyo Toshiba Computer Engineering Co., Ltd.
Claims (5)
段に記憶された任意の比較先の文書から各々単語を抽出
する抽出手段と、 前記抽出手段により比較元の文書及び比較先の文書から
抽出された各単語を、比較元の文書及び比較先の文書の
中での出現位置情報とともに各々個別に格納する単語位
置情報格納手段と、 この単語位置情報格納手段を参照して比較元、比較先の
双方の文書に対して共通の特定の単語が存在するか否か
を判定する判定手段と、 この判定手段により判定された特定の単語に関して単語
位置情報格納手段を参照して比較元、比較先の双方の文
書に対して特定の単語から一定の距離内にある当該単語
に相当する単語を検索する検索手段と、 この検索手段により検索した特定の単語から一定の距離
内にある当該単語に相当する単語に対して比較元、比較
先の双方の文書における特定の単語からの距離に応じた
重み係数を算出する重み係数算出手段と、 この重み係数算出手段の算出結果を基に、前記比較元の
文書に対する比較先の文書の単語の出現位置に応じた類
似度を算出する類似度算出手段と、 を有することを特徴とする類似文書検索装置。An extracting means for extracting a word from an arbitrary document to be compared and an arbitrary document to be compared stored in a document storing means; and extracting means for extracting a word from the document to be compared and the document to be compared by the extracting means. Word position information storage means for individually storing each extracted word together with the appearance position information in the comparison source document and the comparison destination document; and the comparison source and comparison with reference to the word position information storage means Determining means for determining whether or not a specific word common to both of the preceding documents exists; and comparing the specific word determined by the determining means with reference to the word position information storage means for comparison and comparison. Searching means for searching for a word corresponding to the word within a certain distance from a specific word in both of the preceding documents; and searching for the word within a certain distance from the specific word searched by the searching means. Equivalent Weighting factor calculating means for calculating a weighting factor corresponding to the distance from a specific word in both the comparison source and the comparison target documents for the word to be compared, and the comparison source based on the calculation result of the weighting factor calculation means A similarity calculating unit configured to calculate a similarity according to an appearance position of a word of a document to be compared with the document of the similar document.
先の双方の文書に対して特定の単語からの当該単語に相
当する単語の距離が小さいほど大きい値の重み係数とす
ることを特徴とする請求項1記載の類似文書検索装置。2. The weighting factor calculation means according to claim 2, wherein the weighting factor is set to a larger value as the distance of a word corresponding to the word from a specific word is shorter for both the source document and the destination document. The similar document search device according to claim 1, wherein
て、 前記重み係数算出手段により得られた最大の重み係数を
用いて空間ベクトル法による単語ベクトル間の内積の一
つの項データを計算する項データ計算手段と、 前記最大の重み係数を用いて内積データの正規化を行う
内積データ正規化手段とを更に有し、 項データ計算手段、内積データ正規化手段の処理におい
て前記重み係数算出手段により得られた最大の重み係数
を前記特定の単語の出現回数に乗じる毎に、前記重み係
数算出手段による重み係数の算出処理を行うことを特徴
とする類似文書検索装置。3. The similar document search device according to claim 1, wherein one term data of an inner product between word vectors is calculated by a space vector method using a maximum weighting factor obtained by the weighting factor calculation means. Data calculating means, further comprising inner product data normalizing means for normalizing inner product data using the maximum weighting coefficient, term processing means, in the processing of the inner product data normalizing means by the weighting coefficient calculating means A similar document search device, wherein the weight coefficient calculating means performs a weight coefficient calculating process each time the obtained maximum weight coefficient is multiplied by the number of appearances of the specific word.
段に記憶された任意の比較先の文書から各々単語を抽出
し、 前記比較元の文書及び比較先の文書から抽出された各単
語を、比較元の文書及び比較先の文書の中での出現位置
情報とともに各々個別に単語位置情報格納手段に格納
し、 この単語位置情報格納手段を参照して比較元、比較先の
双方の文書に対して共通の特定の単語が存在するか否か
を判定し、 判定された特定の単語に関して単語位置情報格納手段を
参照して比較元、比較先の双方の文書に対して特定の単
語から一定の距離内にある当該単語に相当する単語を検
索し、 重み係数算出手段により、検索した特定の単語から一定
の距離内にある当該単語に相当する単語に対して比較
元、比較先の双方の文書における特定の単語からの距離
に応じた重み係数を算出し、 類似度算出手段により、重み係数算出手段の算出結果を
基にして、前記比較元の文書に対する比較先の文書の単
語の出現位置に応じた類似度を算出すること、 を特徴とする類似文書検索方法。4. Extracting words from an arbitrary document serving as a comparison source and an arbitrary comparison destination document stored in a document storage unit, and extracting each word extracted from the comparison source document and the comparison destination document. Are stored individually in the word position information storage means together with the appearance position information in the comparison source document and the comparison destination document, and are referred to both the comparison source and comparison destination documents by referring to the word position information storage means. It is determined whether or not there is a specific word that is common to the specific word, and the determined specific word is referred to the word position information storage means, and the specific word is fixed for both the source and destination documents. The word corresponding to the word within the distance of the search is searched for, and the weighting factor calculating means compares the word corresponding to the word within a certain distance from the searched specific word with both the comparison source and the comparison destination. From a specific word in the document Calculating a weight coefficient corresponding to the distance, and calculating, by the similarity calculating means, a similarity corresponding to an appearance position of a word of the document of the comparison destination with respect to the document of the comparison source based on the calculation result of the weight coefficient calculating means. A similar document search method, characterized in that:
段に記憶された任意の比較先の文書から各々単語を抽出
する手順と、 前記比較元の文書及び比較先の文書から抽出された各単
語を、比較元の文書及び比較先の文書の中での出現位置
情報とともに各々個別に単語位置情報格納手段に格納す
る手順と、 前記単語位置情報格納手段を参照して比較元、比較先の
双方の文書に対して共通の特定の単語が存在するか否か
を判定する手順と、 判定された特定の単語に関して単語位置情報格納手段を
参照して比較元、比較先の双方の文書に対して特定の単
語から一定の距離内にある当該単語に相当する単語を検
索する手順と、 検索した特定の単語から一定の距離内にある当該単語に
相当する単語に対して比較元、比較先の双方の文書にお
ける特定の単語からの距離に応じた重み係数を算出する
手順と、 重み係数を算出する手順による算出結果を基にして、前
記比較元の文書に対する比較先の文書の単語の出現位置
に応じた類似度を算出する手順と、 からなるプログラムを格納したことを特徴とする類似文
書検索用記憶媒体。5. A procedure for extracting words from an arbitrary document serving as a comparison source and an arbitrary comparison destination document stored in a document storage unit, and a procedure for extracting words from the comparison source document and the comparison destination document. Storing the words individually in the word position information storage unit together with the appearance position information in the comparison source document and the comparison destination document; and referring to the word position information storage unit, A procedure for determining whether or not a specific word common to both documents exists; and referring to the word position information storage means for the determined specific word, for both the source and destination documents. Searching for a word corresponding to the word within a certain distance from the specific word, and comparing a word corresponding to the word within a certain distance from the searched specific word with a comparison source and a comparison destination. Specific in both documents Calculating a weighting factor according to the distance from the word, and calculating the similarity according to the word appearance position of the word of the comparison target document with respect to the comparison source document based on the calculation result obtained by the procedure of calculating the weighting factor. A storage medium for similar document search, characterized by storing a calculation procedure and a program consisting of:
Priority Applications (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP9056723A JPH10254894A (en) | 1997-03-11 | 1997-03-11 | Similar document search device, similar document search method, and similar document search storage medium |
Applications Claiming Priority (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP9056723A JPH10254894A (en) | 1997-03-11 | 1997-03-11 | Similar document search device, similar document search method, and similar document search storage medium |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| JPH10254894A true JPH10254894A (en) | 1998-09-25 |
Family
ID=13035422
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| JP9056723A Pending JPH10254894A (en) | 1997-03-11 | 1997-03-11 | Similar document search device, similar document search method, and similar document search storage medium |
Country Status (1)
| Country | Link |
|---|---|
| JP (1) | JPH10254894A (en) |
Cited By (6)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP2001249951A (en) * | 2000-03-06 | 2001-09-14 | Kddi Corp | Document set characterization method and document set search method and apparatus using the same |
| KR20030016799A (en) * | 2001-08-22 | 2003-03-03 | 장경진 | System and method for internet-based documents comparison |
| JP2008511081A (en) * | 2004-08-23 | 2008-04-10 | トムソン グローバル リソーシーズ | Duplicate document detection and display function |
| US7561685B2 (en) | 2001-12-21 | 2009-07-14 | Research In Motion Limited | Handheld electronic device with keyboard |
| US7973765B2 (en) | 2004-06-21 | 2011-07-05 | Research In Motion Limited | Handheld wireless communication device |
| US7982712B2 (en) | 2004-06-21 | 2011-07-19 | Research In Motion Limited | Handheld wireless communication device |
-
1997
- 1997-03-11 JP JP9056723A patent/JPH10254894A/en active Pending
Cited By (6)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP2001249951A (en) * | 2000-03-06 | 2001-09-14 | Kddi Corp | Document set characterization method and document set search method and apparatus using the same |
| KR20030016799A (en) * | 2001-08-22 | 2003-03-03 | 장경진 | System and method for internet-based documents comparison |
| US7561685B2 (en) | 2001-12-21 | 2009-07-14 | Research In Motion Limited | Handheld electronic device with keyboard |
| US7973765B2 (en) | 2004-06-21 | 2011-07-05 | Research In Motion Limited | Handheld wireless communication device |
| US7982712B2 (en) | 2004-06-21 | 2011-07-19 | Research In Motion Limited | Handheld wireless communication device |
| JP2008511081A (en) * | 2004-08-23 | 2008-04-10 | トムソン グローバル リソーシーズ | Duplicate document detection and display function |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US8375027B2 (en) | Search supporting apparatus and method utilizing exclusion keywords | |
| US7130849B2 (en) | Similarity-based search method by relevance feedback | |
| US7769771B2 (en) | Searching a document using relevance feedback | |
| US8891860B2 (en) | Color name determination device, color name determination method, information recording medium, and program | |
| US7065521B2 (en) | Method for fuzzy logic rule based multimedia information retrival with text and perceptual features | |
| US7949644B2 (en) | Method and apparatus for constructing a compact similarity structure and for using the same in analyzing document relevance | |
| US20070244881A1 (en) | System, method and user interface for retrieving documents | |
| US20020174120A1 (en) | Relevance maximizing, iteration minimizing, relevance-feedback, content-based image retrieval (CBIR) | |
| JP2002519751A (en) | User profile driven information retrieval based on context | |
| US7363311B2 (en) | Method of, apparatus for, and computer program for mapping contents having meta-information | |
| CN118747293A (en) | Document writing intelligent recall method and device and document generation method and device | |
| KR102648613B1 (en) | Method, apparatus and computer-readable recording medium for generating product images displayed in an internet shopping mall based on an input image | |
| US20240370665A1 (en) | Intelligent document processing and information extraction using artificial intelligence | |
| JP4194680B2 (en) | Data processing apparatus and method, and storage medium storing the program | |
| JP2000132554A (en) | Image retrieval apparatus and image retrieval method | |
| JP2000222418A (en) | Database search method and apparatus | |
| JPH10269235A (en) | Similar document search device and similar document search method | |
| CN112650869B (en) | Image retrieval reordering method and device, electronic equipment and storage medium | |
| JP4065470B2 (en) | Information retrieval apparatus and control method thereof | |
| JPH11110395A (en) | Similar document search device and similar document search method | |
| JPH11272709A (en) | File search method | |
| JPH10289245A (en) | Image processing apparatus and control method thereof | |
| JP7614705B2 (en) | Information processing system, information processing method, and program | |
| JPH1131156A (en) | Document search apparatus and method | |
| JP3862059B2 (en) | Search expression expansion method and search system |