JP2000331027A - Similar document search device and similar document search method - Google Patents

Similar document search device and similar document search method

Info

Publication number
JP2000331027A
JP2000331027A JP11142448A JP14244899A JP2000331027A JP 2000331027 A JP2000331027 A JP 2000331027A JP 11142448 A JP11142448 A JP 11142448A JP 14244899 A JP14244899 A JP 14244899A JP 2000331027 A JP2000331027 A JP 2000331027A
Authority
JP
Japan
Prior art keywords
document
search
item
key
search target
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Withdrawn
Application number
JP11142448A
Other languages
Japanese (ja)
Inventor
Shigemi Nakazato
茂美 中里
Yukio Nakamoto
幸夫 中本
Takeshi Matsukuma
剛 松隈
Takuya Nishina
卓哉 仁科
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Toshiba Corp
Toshiba Computer Engineering Corp
Original Assignee
Toshiba Corp
Toshiba Computer Engineering Corp
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Toshiba Corp, Toshiba Computer Engineering Corp filed Critical Toshiba Corp
Priority to JP11142448A priority Critical patent/JP2000331027A/en
Publication of JP2000331027A publication Critical patent/JP2000331027A/en
Withdrawn legal-status Critical Current

Links

Landscapes

  • Information Retrieval, Db Structures And Fs Structures Therefor (AREA)

Abstract

(57)【要約】 【課題】 検索キー文書に異なる内容の項目が存在する
場合にその項目毎の内容に的を絞った類似文書検索を実
現する。 【解決手段】 検索キー文書および検索対象文書から項
目の単位の文書を切り出し、検索キー文書/検索対象文
書間の類似度をベクトル空間法などを用いて前記項目の
単位でそれぞれ算出し、この算出結果に基づいて類似文
書の検索結果(例えば、文書ID)を判別して出力す
る。
(57) [Summary] [Problem] To realize a similar document search focused on the content of each item when an item of different content exists in a search key document. A document in a unit of an item is cut out from a search key document and a search target document, and a similarity between the search key document and the search target document is calculated in the unit of the item using a vector space method or the like. Based on the result, a search result (for example, a document ID) of a similar document is determined and output.

Description

【発明の詳細な説明】DETAILED DESCRIPTION OF THE INVENTION

【0001】[0001]

【発明の属する技術分野】本発明は、任意の文書を検索
のキーとして、このキー文書と類似したものを複数の検
索対象文書の中から自動検索する類似文書検索装置およ
び類似文書検索方法に関する。
BACKGROUND OF THE INVENTION 1. Field of the Invention The present invention relates to a similar document search apparatus and a similar document search method for automatically searching a document similar to the key document from a plurality of search target documents, using an arbitrary document as a search key.

【0002】[0002]

【従来の技術】近年、電子化された大量の文書データが
流通するようになり、これら大量の文書データを一定の
規則に従い分類して利用性を高めることが重要となって
きている。文書データを分類するために、ある一文書を
検索キーとし、その文書と内容が類似した文書を検索対
象文書データベースから抽出する類似文書検索装置があ
る。
2. Description of the Related Art In recent years, a large amount of digitized document data has been distributed, and it has become important to classify such a large amount of document data according to a certain rule to enhance the usability. In order to classify document data, there is a similar document search apparatus that uses a certain document as a search key and extracts a document whose content is similar to the document from a search target document database.

【0003】この類似文書検索装置は、検索キーとする
文書(以降、検索キー文書と呼ぶ。)を構成する単語
と、検索対象文書データベース内の個々の文書(以降、
検索対象文書)を構成する単語とを比較し、この比較結
果を基に検索キー文書と各検索対象文書との類似度を算
出し、その類似度により複数の検索対象文書の中からの
類似文書の抽出を行っている。単語比較結果を基に類似
度を算出する方法としては、検索キー文書と各検索対象
文書との間に共通に出現する単語の種類、出現回数、出
現場所などからベクトル空間法により算出する方法があ
る。
[0003] This similar document search apparatus uses words constituting a document as a search key (hereinafter, referred to as a search key document) and individual documents (hereinafter, referred to as a search key document database) in a search target document database.
The search document is compared with the words constituting the search document, and the similarity between the search key document and each search target document is calculated based on the comparison result, and the similarity is calculated from the plurality of search target documents based on the similarity. Is being extracted. As a method of calculating the similarity based on the word comparison result, there is a method of calculating the similarity between the search key document and each search target document by a vector space method from the type, the number of appearances, the appearance place, etc. of the word commonly appearing. is there.

【0004】[0004]

【発明が解決しようとする課題】上記従来の類似文書検
索方式では、文書間の類似度、つまり検索キー文書全体
と検索対象文書全体との類似度から類似文書を抽出して
いるため、異なる複数の内容が記述されている一文書を
検索キーとした場合は、その複数の内容を同時に含んだ
検索対象文書が類似文書として抽出される。このこと
は、検索キー文書に含まれる個々の内容毎に類似した文
書、言い換えると検索キー文書の一つの内容のみに類似
した文書を検索できないという制限があることを意味す
る。
In the conventional similar document search method, similar documents are extracted from similarities between documents, that is, similarities between the entire search key document and the entire search target document. If one document in which the contents are described is used as a search key, a search target document including the plurality of contents at the same time is extracted as a similar document. This means that there is a limitation that a document similar to each content included in the search key document, in other words, a document similar to only one content of the search key document cannot be searched.

【0005】本発明は、このような課題を解決するため
のもので、文書を構成する項目の単位で、検索キー文書
と類似した文書を複数の検索対象文書の中から検索する
ことのできる類似文書検索装置と類似文書検索方法の提
供を目的とする。
[0005] The present invention is to solve such a problem. A similar document which can search for a document similar to a search key document from a plurality of search target documents in units of items constituting the document. A document search device and a similar document search method are provided.

【0006】すなわち、本発明は、検索キー文書に異な
る内容の項目が存在する場合に、その項目毎の内容に的
を絞った類似文書検索を行うことのできる類似文書検索
装置と類似文書検索方法を提供することを目的とする。
That is, the present invention provides a similar document search apparatus and a similar document search method capable of performing a similar document search focused on the content of each item when an item having different contents exists in the search key document. The purpose is to provide.

【0007】[0007]

【課題を解決するための手段】上記の目的を達成するた
めに、本発明の類似文書検索装置は、検索キー文書と類
似した文書を複数の検索対象文書の中から検索する類似
文書検索装置において、前記検索キー文書および前記検
索対象文書から項目の単位の文書を切り出す項目切り出
し手段と、前記検索キー文書と前記検索対象文書との類
似度を前記切り出された項目の単位で算出し、その算出
結果に基づいて類似文書検索結果を出力する計算手段と
を具備することを特徴とするものである。
In order to achieve the above object, a similar document search apparatus according to the present invention provides a similar document search apparatus for searching a document similar to a search key document from a plurality of search target documents. An item extracting unit for extracting a document in an item unit from the search key document and the search target document, and calculating a similarity between the search key document and the search target document in the unit of the extracted item, and calculating the same. Calculating means for outputting a similar document search result based on the result.

【0008】この発明は、検索キー文書と検索対象文書
との類似度を各文書を構成する項目の単位で求めること
によって、検索キー文書に異なる内容の項目が存在する
場合に、その項目毎の内容に的を絞った類似文書検索を
行うことができる。
According to the present invention, the similarity between a search key document and a search target document is obtained in units of items constituting each document. It is possible to perform similar document search focused on the content.

【0009】また、項目毎に優先度を設定する手段を付
加し、この設定された優先度を加味して検索キー文書と
検索対象文書との項目単位の類似度を算出するように構
成することによって、たとえば、検索対象文書毎に全項
目の類似度の総和を求めその結果を検索結果として出力
する場合に、文書の重要部分の類似度に重みを加えたよ
り最適な類似文書検索を実現することができる。
[0009] Further, a means for setting a priority for each item is added, and the similarity of the retrieval key document and the retrieval target document in item units is calculated in consideration of the set priority. For example, when a total sum of similarities of all items is obtained for each search target document and the result is output as a search result, a more optimal similar document search in which a similarity of an important part of the document is weighted is realized. Can be.

【0010】さらに、項目切り出し手段によって切り出
された検索キー文書または検索対象文書の各項目の文構
造を統一化させる手段をさらに付加し、このように文構
造を統一化された各項目について検索キー文書と検索対
象文書との類似度を項目単位で算出するように構成する
ことで、各項目の文構造の違いが類似度に影響する度合
を軽減することができ、より算出される類似度の妥当性
を高めることができる。
Further, means for unifying the sentence structure of each item of the search key document or the document to be searched extracted by the item extracting means is further added, and the search key is unified for each item having such unified sentence structure. By configuring the similarity between the document and the search target document on an item-by-item basis, the degree to which the difference in the sentence structure of each item affects the similarity can be reduced. Relevance can be increased.

【0011】[0011]

【発明の実施の形態】以下、本発明の一実施形態を図面
を参照しながら説明する。
An embodiment of the present invention will be described below with reference to the drawings.

【0012】図1に、本実施形態に係る類似文書検索装
置のハードウェア構成を示す。同図に示すように、この
類似文書検索装置は、CPU、メモリから構成される制
御装置1、キーボードなどの入力装置2、類似検索結果
などの表示する表示装置3、検索データなどを格納する
外部記憶装置4といったコンピュータシステム環境にお
いて構築される。
FIG. 1 shows a hardware configuration of a similar document search apparatus according to this embodiment. As shown in FIG. 1, the similar document search device includes a control device 1 including a CPU and a memory, an input device 2 such as a keyboard, a display device 3 for displaying similar search results, and an external device for storing search data and the like. It is constructed in a computer system environment such as the storage device 4.

【0013】図2に、類似文書検索装置の制御部の構成
を機能別にブロック化して示す。このように類似文書検
索装置の制御部は処理部11とメモリ部12とからな
る。
FIG. 2 is a block diagram showing the configuration of the control unit of the similar document search apparatus by function. As described above, the control unit of the similar document search device includes the processing unit 11 and the memory unit 12.

【0014】制御部11は各種の制御や処理のための演
算を実行する部分であり、メイン処理部200、初期化
部201、入力部202、出力部203、検索対象文書
読み出し部204、検索対象文書項目切り出し部20
5、検索対象文書項目生成部206、検索対象文書単語
抽出部207、検索対象単語出現頻度算出部208、検
索対象単語情報算出部209、検索キー文書入力部21
0、検索キー文書項目優先度入力部211、検索キー文
書項目切り出し部212、検索キー文書項目生成部21
3、検索キー単語抽出部214、検索キー単語出現頻度
算出部215、検索条件設定部216、共通単語抽出部
217、類似度算出部218、検索結果出力部219な
どからなる。
The control section 11 is a section for executing calculations for various controls and processes. The main processing section 200, an initialization section 201, an input section 202, an output section 203, a search target document reading section 204, a search target Document item extraction unit 20
5. Search target document item generation unit 206, search target document word extraction unit 207, search target word appearance frequency calculation unit 208, search target word information calculation unit 209, search key document input unit 21
0, search key document item priority input unit 211, search key document item cutout unit 212, search key document item generation unit 21
3, a search key word extraction unit 214, a search key word appearance frequency calculation unit 215, a search condition setting unit 216, a common word extraction unit 217, a similarity calculation unit 218, a search result output unit 219, and the like.

【0015】また、メモリ部12は、検索条件設定バッ
ファ部229、検索対象文書格納バッファ部230、検
索対象文書項目格納バッファ部231、検索対象文書項
目生成格納バッファ部232、検索対象単語情報格納バ
ッファ部233、検索キー文書項目優先度バッファ部2
34、検索キー文書格納バッファ部235、検索キー文
書項目格納バッファ部236、検索キー文書項目生成格
納バッファ部237、検索キー単語情報格納バッファ部
238、共通単語情報格納バッファ部239、類似度格
納バッファ部240、検索結果出力バッファ部241、
作業バッファ部242などからなる。
The memory unit 12 includes a search condition setting buffer unit 229, a search target document storage buffer unit 230, a search target document item storage buffer unit 231, a search target document item generation storage buffer unit 232, and a search target word information storage buffer. Section 233, search key document item priority buffer section 2
34, search key document storage buffer 235, search key document item storage buffer 236, search key document item generation storage buffer 237, search key word information storage buffer 238, common word information storage buffer 239, similarity storage buffer Unit 240, search result output buffer unit 241,
It comprises a work buffer unit 242 and the like.

【0016】処理部11の各部の詳細について説明す
る。
The details of each section of the processing section 11 will be described.

【0017】初期化部201は、メモリ部12内の各バ
ッファ部229〜242の初期化を行う。
The initialization section 201 initializes each of the buffer sections 229 to 242 in the memory section 12.

【0018】入力部202は、ユーザからの各種設定の
ための情報や検索キーとなる文書などの入力を処理す
る。
The input unit 202 processes input of information for various settings from a user and a document serving as a search key.

【0019】出力部203は、検索キー文書や類似検索
結果、さらには各種の設定情報などを表示装置3を通じ
てユーザに表示する処理を行う。
The output unit 203 performs a process of displaying a search key document, a similar search result, various setting information, and the like to the user through the display device 3.

【0020】検索対象文書読み出し部204は、外部記
憶装置4に格納されてい検索対象文書をデータベース化
して検索対象文書格納バッファ部230に格納する。
The search target document reading unit 204 converts the search target documents stored in the external storage device 4 into a database and stores them in the search target document storage buffer unit 230.

【0021】検索対象文書項目切り出し部205は、検
索対象文書格納バッファ部230に格納されている検索
対象文書の文構造を解析し、その解析結果を基に検索対
象文書から定型の項目の文書を切り出して検索対象文書
項目格納バッファ部231に格納する。
The search target document item cutout unit 205 analyzes the sentence structure of the search target document stored in the search target document storage buffer unit 230 and, based on the analysis result, extracts a document of a fixed item from the search target document. It is cut out and stored in the search target document item storage buffer unit 231.

【0022】検索対象文書項目生成部206は、必要に
応じて、検索対象文書項目格納バッファ部231に格納
された項目の文書の再生成を行い、再生成された項目の
文書を検索対象文書項目生成格納バッファ部232に格
納する。
The search target document item generation unit 206 regenerates the document of the item stored in the search target document item storage buffer unit 231 as necessary, and replaces the regenerated item document with the search target document item. The data is stored in the generation storage buffer unit 232.

【0023】検索対象文書単語抽出部207は、検索対
象文書格納バッファ部230あるいは検索対象文書項目
格納バッファ部231に格納されている文書から単語を
切り出し、切り出された単語群の中からその文書(ある
いは項目)の内容を表す上でキーとなる単語を抽出し、
抽出した単語を検索対象単語情報格納バッファ部233
に格納する。
The search target document word extraction unit 207 cuts out a word from a document stored in the search target document storage buffer unit 230 or the search target document item storage buffer unit 231, and selects the document (from the cut-out word group). Or a key word to represent the content of
The extracted words are stored in the search target word information storage buffer unit 233.
To be stored.

【0024】検索対象単語出現頻度算出部208は、検
索対象文書単語抽出部207により抽出されたキー単語
の、検索対象文書項目格納部231あるいは検索対象文
書項目生成格納バッファ部232に格納されている検索
対象文書中での出現頻度を単語種毎に算出し、その結果
を検索対象単語情報格納バッファ部233に格納する。
The search target word appearance frequency calculation unit 208 stores the key words extracted by the search target document word extraction unit 207 in the search target document item storage unit 231 or the search target document item generation storage buffer unit 232. The appearance frequency in the search target document is calculated for each word type, and the result is stored in the search target word information storage buffer unit 233.

【0025】検索対象単語情報算出部209は、検索対
象単語情報格納バッファ部233に格納されている検索
対象文書中の単語種毎の出現頻度を外部記憶装置4に格
納する。
The search target word information calculation unit 209 stores the appearance frequency of each word type in the search target document stored in the search target word information storage buffer unit 233 in the external storage device 4.

【0026】検索キー文書入力部210は、入力部20
2を通じてユーザより入力された検索キーの文書を検索
キー文書格納バッファ部235に格納する。
The search key document input unit 210 is connected to the input unit 20
2 stores the document of the search key input by the user in the search key document storage buffer unit 235.

【0027】検索キー文書項目優先度入力部211は、
検索キー文書の各項目の優先度をユーザからの入力によ
り検索キー文書項目優先度バッファ部234に設定す
る。
The search key document item priority input unit 211
The priority of each item of the search key document is set in the search key document item priority buffer unit 234 by input from the user.

【0028】検索キー文書項目切り出し部212は、検
索キー文書格納バッファ部235に格納された検索キー
文書の文構造を解析し、その解析結果を基に検索キー文
書から定型の項目の文書を切り出して検索キー文書項目
格納バッファ部236に格納する。
The search key document item cutout unit 212 analyzes the sentence structure of the search key document stored in the search key document storage buffer unit 235, and cuts out a document of fixed items from the search key document based on the analysis result. Then, it is stored in the search key document item storage buffer unit 236.

【0029】検索キー文書項目生成部213は、必要に
応じて、検索キー文書項目格納バッファ部236に格納
されている項目の文書の再生成を行い、再生成された項
目の文書を検索キー文書項目生成格納バッファ部237
に格納する。
The search key document item generation unit 213 regenerates the document of the item stored in the search key document item storage buffer unit 236 as necessary, and converts the regenerated item document into the search key document. Item generation storage buffer unit 237
To be stored.

【0030】検索キー文書単語抽出部214は、検索キ
ー文書項目格納バッファ部236あるいは検索キー文書
項目生成格納バッファ部237に格納されている検索キ
ー文書から単語を切り出し、切り出された単語群の中か
らその検索キー文書(あるいは項目)の内容を表す上で
キーとなる単語を抽出し、抽出された単語を検索キー単
語情報格納バッファ部238に格納する。
The search key document word extraction unit 214 cuts out words from the search key document stored in the search key document item storage buffer unit 236 or the search key document item generation storage buffer unit 237, and outputs the words from the extracted word group. , A key word representing the contents of the search key document (or item) is extracted, and the extracted word is stored in the search key word information storage buffer unit 238.

【0031】検索キー単語出現頻度算出部215は、検
索キー単語抽出部214により抽出されたキー単語の、
検索キー文書格納バッファ部238に格納されている検
索キー文書中での出現頻度を単語種毎に算出し、その結
果を検索キー単語情報格納バッファ部238に格納す
る。
The search key word appearance frequency calculation unit 215 calculates the key word extracted by the search key word extraction unit 214
The appearance frequency in the search key document stored in the search key document storage buffer unit 238 is calculated for each word type, and the result is stored in the search key word information storage buffer unit 238.

【0032】検索条件設定部216は、類似文書を算出
する際の検索条件をユーザからの入力により設定する。
検索条件としては、検索キー文書の項目の生成処理の有
無、検索対象文書の項目の切り出し処理の有無、検索対
象文書の項目の生成処理の有無などがある。これら検索
条件の設定内容は検索条件設定バッファ部229に格納
される。
The search condition setting section 216 sets search conditions for calculating similar documents based on input from a user.
The search conditions include the presence / absence of a process of generating a search key document item, the presence / absence of a process of extracting a search target document item, and the presence / absence of a process of generating a search target document item. The setting contents of these search conditions are stored in the search condition setting buffer unit 229.

【0033】共通単語抽出部217は、検索キー単語情
報格納バッファ部238に格納されている検索キー文書
と検索対象単語情報格納バッファ部233に格納されて
いる検索対象文書とに共通の単語を抽出し、この共通単
語の情報と該共通単語毎の頻度情報を共通単語情報格納
バッファ部239に格納する。
The common word extracting section 217 extracts a word common to the search key document stored in the search key word information storage buffer section 238 and the search target document stored in the search target word information storage buffer section 233. Then, the common word information and the frequency information for each common word are stored in the common word information storage buffer unit 239.

【0034】類似度算出部218は、検索キー単語情報
格納バッファ部238に格納されている検索キー文書中
の単語種毎の出現頻度、検索対象単語情報格納バッファ
部233に格納されている検索対象文書中の単語種毎の
出現頻度、そして共通単語情報格納バッファ部239に
格納されている共通単語種毎の出現頻度から、ベクトル
空間法等によって検索キー文書と個々の検索対象文書と
の類似度をそれぞれ算出し、その類似度値を類似度格納
バッファ部240に格納する。このとき検索キー文書項
目優先度バッファ部234に検索キー文書項目の優先度
が設定されている場合は、該当する項目の類似度値に優
先度に応じた重みを付与し、この重みが付与された類似
度値を類似度格納バッファ部240に再格納する。
The similarity calculation unit 218 includes an appearance frequency for each word type in the search key document stored in the search key word information storage buffer unit 238, and a search target stored in the search target word information storage buffer unit 233. Based on the appearance frequency of each word type in the document and the appearance frequency of each common word type stored in the common word information storage buffer unit 239, the similarity between the search key document and each search target document by a vector space method or the like. Is calculated, and the similarity value is stored in the similarity storage buffer unit 240. At this time, when the priority of the search key document item is set in the search key document item priority buffer unit 234, a weight corresponding to the priority is assigned to the similarity value of the corresponding item, and this weight is assigned. The stored similarity value is stored in the similarity storage buffer unit 240 again.

【0035】検索結果出力部219は、類似度格納バッ
ファ部240に格納されている検索対象文書毎の類似度
値を基に検索結果を検索結果出力バッファ部241に格
納する。そして検索結果出力部219は、検索結果出力
バッファ部239の内容を表示装置3に出力する。
The search result output unit 219 stores a search result in the search result output buffer unit 241 based on the similarity value for each search target document stored in the similarity storage buffer unit 240. Then, the search result output unit 219 outputs the contents of the search result output buffer unit 239 to the display device 3.

【0036】次に、本実施形態の類似文書検索装置の動
作を説明する。
Next, the operation of the similar document search apparatus according to this embodiment will be described.

【0037】図3および図4に、本実施形態の類似文書
検索処理の流れを示す。まず、初期化部201が起動さ
れメモリ部12のクリア等の初期化が行われる(ステッ
プ300)。
FIGS. 3 and 4 show the flow of a similar document search process according to this embodiment. First, the initialization unit 201 is activated and initialization such as clearing of the memory unit 12 is performed (step 300).

【0038】続いて検索条件設定部216が起動され
る。検索条件設定部216は類似文書を算出する際の検
索条件を入力装置2からの入力により設定し、その設定
データを検索条件設定バッファ部229に格納する(ス
テップ301)。図4に、検索モードの設定例を示す。
同図に示すように、検索モードの設定項目には 検索キー文書の項目生成=する/しない 検索対象文書の項目切り出し=する/しない 検索対象文書の項目生成=する/しない などがあり、これら設定項目毎にユーザは任意のモード
を選択することができる。この例では、 検索キー文書の項目生成=する 検索対象文書の項目切り出し=する 検索対象文書の項目生成=する がそれぞれ選択されたものとする。これらの検索モード
の設定情報は検索条件設定バッファ部229に格納され
る。図6に検索条件設定バッファ部229への検索モー
ド設定情報の格納形態を示す。
Subsequently, the search condition setting section 216 is activated. The search condition setting unit 216 sets search conditions for calculating a similar document by inputting from the input device 2, and stores the set data in the search condition setting buffer unit 229 (step 301). FIG. 4 shows a setting example of the search mode.
As shown in the figure, the setting items of the search mode include the item generation of the search key document = Yes / No, the item extraction of the search target document = Yes / No, the item generation of the search target document = Yes / No, and the like. The user can select an arbitrary mode for each item. In this example, it is assumed that “Generate item of search key document = Yes” and “Cut out item of search target document = Yes” are selected. The setting information of these search modes is stored in the search condition setting buffer unit 229. FIG. 6 shows a storage mode of the search mode setting information in the search condition setting buffer unit 229.

【0039】次に、検索キー文書の項目優先度を入力す
るかどうかを判断する(ステップ302)、項目優先度
を入力する場合は検索キー文書項目優先度入力部211
が起動される。検索キー文書項目優先度入力部211は
入力部202を通じてユーザより、検索キー文書中の項
目毎の優先度を入力し、その項目別優先度の情報を検索
キー文書項目優先度バッファ部234に格納する(ステ
ップ303)。
Next, it is determined whether or not to input the item priority of the search key document (step 302). When the item priority is input, the search key document item priority input section 211 is input.
Is started. The search key document item priority input unit 211 inputs the priority of each item in the search key document from the user via the input unit 202, and stores information on the priority for each item in the search key document item priority buffer unit 234. (Step 303).

【0040】図7は検索キー文書中の定型5項目それぞ
れの重要度を設定するときの設定画面を示している。同
図では各項目の優先度(=重要度)を10段階に表すも
のとし、項目番号1の優先度を「10」、項目番号2の
優先度を「5」、項目番号3の優先度を「3」、項目番
号4,5の優先度を「1」に設定している。
FIG. 7 shows a setting screen for setting the importance of each of the five standard items in the search key document. In the figure, the priority (= importance) of each item is represented in 10 levels, the priority of item number 1 is “10”, the priority of item number 2 is “5”, and the priority of item number 3 is The priority of “3” and the item numbers 4 and 5 are set to “1”.

【0041】この検索キー文書の項目優先度の設定後、
あるいはステップ304で検索キー文書の項目優先度を
入力しないこととした場合はステップ304に進む。
After setting the item priority of the search key document,
Alternatively, if it is determined in step 304 that the item priority of the search key document is not input, the process proceeds to step 304.

【0042】ステップ304では検索キー文書入力部2
10が起動される。ここでユーザより検索キーとなる文
書が入力されることで、検索キー文書入力部210はそ
の入力された検索キー文書を検索キー文書格納バッファ
部235に格納する(ステップ304)。図8に、その
入力された検索キー文書の例を示す。ここで、入力さ
In step 304, the retrieval key document input unit 2
10 is activated. When the user inputs a document serving as a search key, the search key document input unit 210 stores the input search key document in the search key document storage buffer unit 235 (step 304). FIG. 8 shows an example of the input search key document. Where

【0043】れた検索キー文書は複数の定型項目(例え
The retrieved search key document has a plurality of fixed items (for example,

【請求項 】で表記される項目)の文書からなるものと
する。
).

【0044】この後、検索キー文書項目切り出し部21
2が起動される。検索キー文書項目切り出し部212
は、検索キー文書格納バッファ部235に格納されてい
る検索キー文書の文構造を解析して、この解析結果を基
に当該検索キー文書から定型項目の文書部分をすべて切
り出し、切り出した項目単位の文書を検索キー文書項目
格納バッファ部236に格納する(ステップ305)。
図9に、検索キー文書から切り出された項目単位の文書
を示す。
Thereafter, the retrieval key document item cutout unit 21
2 is activated. Search key document item cutout unit 212
Analyzes the sentence structure of the search key document stored in the search key document storage buffer unit 235, cuts out all the document portions of the standard items from the search key document based on this analysis result, and The document is stored in the search key document item storage buffer unit 236 (step 305).
FIG. 9 shows an item-by-item document cut out from the search key document.

【0045】次に、検索キー文書項目生成部213は、
検索条件設定バッファ部229を参照し、ここに「検索
キー文書の項目生成=する」が設定されているかどうか
を調べる(ステップ306)。「検索キー文書の項目生
成=する」が設定されていれば、検索キー文書項目生成
部213は、検索キー文書項目格納バッファ部236に
格納されている定型項目の文書の再生成を行う(ステッ
プ307)。再生成された定型項目の文書は検索キー文
書項目生成格納バッファ部237に格納される。
Next, the search key document item generation unit 213
With reference to the search condition setting buffer unit 229, it is checked whether or not “item generation of search key document = Yes” is set here (step 306). If “Generate item of search key document = Yes” is set, the search key document item generation unit 213 regenerates the document of the standard item stored in the search key document item storage buffer unit 236 (step). 307). The regenerated document of the standard item is stored in the search key document item generation storage buffer unit 237.

【0046】図10に、この定型項目の文書の再生成の
例を示す。定型項目の文書の再生成は所定の規則に則っ
て行われる。この定型項目の文書の再生成は項目文書間
の文構造や表記の違いを吸収する目的で行われる。例え
ば、図10に示すように、項目番号2以降の項目文書に
含まれる「請求項X」(ただしXは1以上の整数)とい
う記述部分は、項目番号Xの項目文書に置き換えれる。
FIG. 10 shows an example of regenerating the document of the fixed item. Regeneration of the document of the standard item is performed according to a predetermined rule. The regeneration of the document of the standard item is performed for the purpose of absorbing the difference in the sentence structure and the notation between the item documents. For example, as shown in FIG. 10, the description portion of “Claim X” (where X is an integer of 1 or more) included in the item documents after item number 2 is replaced with the item document of item number X.

【0047】このような定型項目の文書の再生成を行っ
た後、あるいはステップ306で「検索キー文書の項目
生成=しない」が設定されている場合は次にステップ3
08が実行される。
After the re-creation of the document of such a fixed item, or when “Generation of item of search key document = not set” is set in step 306, then step 3 is executed.
08 is executed.

【0048】ステップ308では検索キー単語抽出部2
14が起動される。検索キー単語抽出部214は、検索
キー文書項目格納バッファ部236あるいは検索キー文
書項目生成格納バッファ部237に格納されている検索
キー文書から単語を切り出し、切り出した単語群からそ
の検索キー文書(あるいは項目)の内容を反映するキー
単語を抽出し、抽出した単語を検索キー単語情報格納バ
ッファ部238に格納する。ここで単語の切り出しは形
態素解析などにより行われ、そのキー単語は単語の品詞
に基づき決定することができる。例えば「名詞」や「サ
変名詞」の単語をキー単語として判別するようにする。
このようにして抽出された単語の情報は、図11に示す
ように、検索キー単語情報格納バッファ部238に格納
される。
In step 308, the retrieval key word extraction unit 2
14 is activated. The search key word extracting unit 214 cuts out a word from the search key document stored in the search key document item storage buffer unit 236 or the search key document item generation storage buffer unit 237, and extracts the search key document (or Key words that reflect the content of the item are extracted, and the extracted words are stored in the search key word information storage buffer unit 238. Here, the extraction of the word is performed by morphological analysis or the like, and the key word can be determined based on the part of speech of the word. For example, words such as "noun" and "sa-noun" are determined as key words.
The word information extracted in this way is stored in the search key word information storage buffer unit 238 as shown in FIG.

【0049】続いて、検索キー単語出現頻度算出部21
5が起動される。検索キー単語出現頻度算出部215
は、検索キー単語情報格納バッファ部238に格納され
た単語種について、検索キー文書項目生成格納バッファ
部237または検索キー文書項目格納バッファ部236
に格納されている項目文書中の出現頻度を算出し、図1
7に示すように、その結果を検索キー単語情報格納バッ
ファ部238に格納する(ステップ309)。図12に
おいて、「文書データベース=1」は「文書データベー
ス」という単語が1回出現していることを示す。
Subsequently, the search key word appearance frequency calculation unit 21
5 is activated. Search key word appearance frequency calculation unit 215
Is a search key document item generation storage buffer unit 237 or a search key document item storage buffer unit 236 for the word type stored in the search key word information storage buffer unit 238.
1 is calculated in the item document stored in FIG.
As shown in FIG. 7, the result is stored in the search key word information storage buffer unit 238 (step 309). In FIG. 12, "document database = 1" indicates that the word "document database" appears once.

【0050】次に、検索対象文書読み出し部204が起
動される。検索対象文書読み出し部204は、外部記憶
装置4にまだ処理を終えてない検索対象文書あるか否か
を判断し(ステップ310)、もし検索対象文書があれ
ば一つの検索対象文書を検索対象文書格納バッファ部2
30に格納する(ステップ311)。図13にその検索
対象文書の例を示す。
Next, the retrieval target document reading section 204 is started. The search target document reading unit 204 determines whether there is a search target document which has not been processed yet in the external storage device 4 (step 310). If there is a search target document, one search target document is searched. Storage buffer unit 2
30 (step 311). FIG. 13 shows an example of the search target document.

【0051】続いて、検索対象文書項目切り出し部20
5は、検索条件設定バッファ部229を参照し、検索対
象文書の項目抽出を行うかどうかを判断する(ステップ
312)。項目抽出を行う場合、検索対象文書項目切り
出し部205は、検索対象文書格納バッファ部230に
格納されている検索対象文書の文構造を解析し、この解
析結果を基に当該検索対象文書から定型項目の文書部分
をすべて切り出し、切り出した項目単位の文書を検索対
象文書項目格納バッファ部231に格納する(ステップ
313)。図14に検索対象文書から切り出された項目
単位の文書を示す。
Subsequently, the search target document item cutout unit 20
5 refers to the search condition setting buffer unit 229 and determines whether to extract items of the search target document (step 312). When performing the item extraction, the search target document item cutout unit 205 analyzes the sentence structure of the search target document stored in the search target document storage buffer unit 230, and based on the analysis result, extracts the fixed-form item from the search target document. Is cut out, and the cut out document in item units is stored in the search target document item storage buffer unit 231 (step 313). FIG. 14 shows an item-by-item document cut out from the search target document.

【0052】次に、検索対象文書項目生成部206は、
検索条件設定バッファ部229を参照し、ここに「検索
対象文書の項目生成=する」が設定されているかどうか
を調べる(ステップ314)。「検索対象文書の項目生
成=する」が設定されていれば、検索対象文書項目生成
部206は、検索対象文書項目格納バッファ部231に
格納されている定型項目の文書の再生成を行う(ステッ
プ315)。
Next, the search target document item generation unit 206
With reference to the search condition setting buffer unit 229, it is checked whether or not "create item of search target document = Yes" is set here (step 314). If “Generate item of search target document = Yes” is set, the search target document item generation unit 206 regenerates the document of the standard item stored in the search target document item storage buffer unit 231 (step). 315).

【0053】図15に、この定型項目の文書の再生成の
例を示す。この検索対象文書の項目再再生は前述した検
索キー文書の項目再生成と同様に行われる。再生成され
た定型項目の文書は検索対象文書項目生成格納バッファ
部232に格納される。
FIG. 15 shows an example of regenerating the document of the fixed item. The item reproduction of the search target document is performed in the same manner as the above-described item regeneration of the search key document. The regenerated standard item document is stored in the search target document item generation storage buffer unit 232.

【0054】このような定型項目の文書の再生成を行っ
た後、あるいはステップ314で「検索対象文書の項目
生成=しない」が設定されている場合はステップ316
が実行される。
After regenerating the document of such a fixed item, or when “create item of search target document = not set” is set in step 314, step 316 is executed.
Is executed.

【0055】ステップ316では検索対象単語抽出部2
09が起動される。検索対象単語抽出部209は、検索
対象文書項目格納バッファ部231あるいは検索対象文
書項目生成格納バッファ部232に格納されている検索
対象文書から単語を切り出し、切り出した単語群からそ
の検索対象文書(あるいは項目)の内容を反映するキー
単語を抽出し、抽出した単語を検索対象単語情報格納バ
ッファ部233に格納する。この検索対象文書からのキ
ー単語の切り出しは前述した検索キー文書からのキー単
語の切り出しと同様に形態素解析などによって行われ
る。このようにして抽出された単語の情報は、図16に
示すように、検索対象単語情報格納バッファ部233に
格納される。
In step 316, the search target word extracting unit 2
09 is started. The search target word extracting unit 209 cuts out a word from the search target document stored in the search target document item storage buffer unit 231 or the search target document item generation storage buffer unit 232, and extracts the search target document (or Key words that reflect the contents of the item are extracted, and the extracted words are stored in the search target word information storage buffer unit 233. The extraction of the key word from the search target document is performed by morphological analysis or the like, similarly to the extraction of the key word from the search key document. The word information thus extracted is stored in the search target word information storage buffer unit 233 as shown in FIG.

【0056】続いて、検索対象単語出現頻度算出部20
8が起動される。検索対象単語出現頻度算出部208
は、検索対象単語情報格納バッファ部233に格納され
ている単語について、検索対象文書項目格納バッファ部
231または検索対象文書項目生成格納バッファ部23
2に格納されている項目文書中の出現頻度を算出し、図
17に示すように、その結果を検索対象単語情報格納バ
ッファ部233に格納する(ステップ317)。
Subsequently, the search target word appearance frequency calculation unit 20
8 is activated. Search target word appearance frequency calculation unit 208
Are the search target document item storage buffer unit 231 or the search target document item generation storage buffer unit 23 for the words stored in the search target word information storage buffer unit 233.
2 is calculated, and the result is stored in the search target word information storage buffer unit 233 as shown in FIG. 17 (step 317).

【0057】この後、共通単語抽出部217が起動され
る。共通単語抽出部217は、検索対象単語情報格納バ
ッファ部233および検索キー単語情報格納バッファ部
238の同一番号の項目毎に、共通に格納されている単
語を検索し、検索した項目毎の共通単語を共通単語情報
格納バッファ部239に格納する(ステップ318)。
図18にこの共通単語の検索結果の例を示す。
Thereafter, the common word extraction unit 217 is activated. The common word extraction unit 217 searches for a commonly stored word for each item of the same number in the search target word information storage buffer unit 233 and the search key word information storage buffer unit 238, and searches for a common word for each searched item. Is stored in the common word information storage buffer unit 239 (step 318).
FIG. 18 shows an example of a search result of the common word.

【0058】次に、類似度算出部218が起動される。
類似度算出部218は、検索対象単語情報格納バッファ
部233、検索キー単語情報格納バッファ部238およ
び共通単語情報格納バッファ部239の格納情報を基に
ベクトル空間法などを用いて検索キー文書/検索対象文
書間の項目別の類似度を算出し、その類似度値を類似度
格納バッファ部240に格納する。図19にその項目別
の類似度値の例を示す。
Next, the similarity calculating section 218 is started.
The similarity calculation unit 218 uses a search key document / search based on the storage information of the search target word information storage buffer unit 233, the search key word information storage buffer unit 238, and the common word information storage buffer unit 239 using a vector space method or the like. The similarity for each item between the target documents is calculated, and the similarity value is stored in the similarity storage buffer unit 240. FIG. 19 shows an example of the similarity value for each item.

【0059】また、この類似度値の計算の際に類似度算
出部218は、検索キー文書項目優先度バッファ部23
4を参照する。図20に検索キー文書項目優先度バッフ
ァ部234に設定された優先度の例を示す。ここに優先
度が設定されていれば、その優先度が設定されている項
目についての類似度に優先度を加味する。例えば、「優
先度=10」であれば、その項目の類似度を2倍に、
「優先度=5」であればその項目の類似度を1.5倍に
するなどして項目別の類似度に優先度による重みを付与
し、この結果を類似度格納バッファ部240に再格納す
る(ステップ319)。図21に、優先度による重みが
付与された項目別の類似度を示す。
When calculating the similarity value, the similarity calculating unit 218 searches the search key document item priority buffer unit 23
Refer to FIG. FIG. 20 shows an example of the priorities set in the search key document item priority buffer unit 234. If a priority is set here, the priority is added to the similarity for the item for which the priority is set. For example, if “priority = 10”, the similarity of the item is doubled,
If “priority = 5”, the similarity of the item is weighted by priority by, for example, multiplying the similarity of the item by 1.5, and the result is stored in the similarity storage buffer unit 240 again. (Step 319). FIG. 21 shows the similarity for each item to which a weight is assigned according to the priority.

【0060】その後、ステップ310に戻る。ステップ
310で外部記憶装置4に処理を終えてない検索対象文
書がないことが判断されると検索結果出力部219が起
動される。検索結果出力部219は、図22に示すよう
に、検索キー単語の項目別に最も高い類似度を持つ検索
対象文書を判別し、その文書情報(例えば、文書ID)
を検索結果出力バッファ部240に格納する。そして、
検索結果出力バッファ部241の内容を表示装置3に出
力する(ステップ320)。図23に出力結果の例を示
す。
Thereafter, the flow returns to step 310. If it is determined in step 310 that there is no search target document in the external storage device 4 that has not been processed, the search result output unit 219 is activated. As shown in FIG. 22, the search result output unit 219 determines the search target document having the highest similarity for each item of the search key word, and determines its document information (for example, document ID).
Is stored in the search result output buffer unit 240. And
The contents of the search result output buffer unit 241 are output to the display device 3 (Step 320). FIG. 23 shows an example of the output result.

【0061】このように本実施形態によれば、検索キー
文書の個々の項目毎にこれに類似した項目を有する検索
対象文書を検索することができる。
As described above, according to the present embodiment, it is possible to search for a search target document having an item similar to each item of the search key document.

【0062】また、図24に示すように、検索キー文書
の項目毎に、検索対象文書の各項目の類似度の総和を求
め、そして図25に示すように、その結果を検索結果出
力バッファ部240に格納して類似度の総和とともに出
力するようにしても構わない。このとき、類似度の総和
が高い順に検索対象文書のIDを並べて表示するように
してもよい。
Further, as shown in FIG. 24, for each item of the retrieval key document, the sum of the similarities of the respective items of the retrieval target document is obtained, and as shown in FIG. 25, the result is stored in the retrieval result output buffer section. 240 and may be output together with the sum of the similarities. At this time, the IDs of the search target documents may be displayed side by side in descending order of the sum of the similarities.

【0063】さらに、図26に示すように、検索対象文
書毎に全項目の類似度の総和を求め、図27に示すよう
に、その結果を類似度の総和とともに出力するようにし
てもよく、この場合も、類似度の総和が高い順に検索対
象文書のIDを並べて表示することが好ましい。
Further, as shown in FIG. 26, the sum of similarities of all items may be obtained for each document to be searched, and the result may be output together with the sum of similarities as shown in FIG. Also in this case, it is preferable to display the IDs of the search target documents in the descending order of the total sum of the similarities.

【0064】以上、説明した類似文書検索の機能は、コ
ンピュータが読み取り可能なCD−ROMやその他の記
憶媒体にコンピュータ上で実行可能なプログラムとして
記憶して提供することが可能である。
The above-described similar document search function can be provided by storing it as a computer-executable program on a computer-readable CD-ROM or other storage medium.

【0065】[0065]

【発明の効果】以上説明したように本発明によれば、検
索キー文書と検索対象文書との類似度を各文書を構成す
る項目の単位で求めることによって、検索キー文書に異
なる内容の項目が存在する場合に、その項目毎の個別の
内容に的を絞った類似文書検索を行うことができる。
As described above, according to the present invention, the similarity between a search key document and a search target document is obtained in units of items constituting each document. If there is, a similar document search can be performed that focuses on the individual content of each item.

【0066】また、項目毎に優先度を設定する手段を付
加し、この設定された優先度を加味して検索キー文書と
検索対象文書との項目単位の類似度を算出するように構
成することによって、たとえば、検索対象文書毎に全項
目の類似度の総和を求めその結果を検索結果として出力
する場合に、文書の重要部分の類似度に重みを加えたよ
り最適な類似文書検索を実現することができる。
A means for setting a priority for each item is added, and the similarity of the retrieval key document and the retrieval target document is calculated in item units in consideration of the set priority. For example, when a total sum of similarities of all items is obtained for each search target document and the result is output as a search result, a more optimal similar document search in which a similarity of an important part of the document is weighted is realized. Can be.

【0067】さらに、検索キー文書または検索対象文書
の各項目の文構造を統一化させる手段をさらに付加し、
このように文構造を統一化された各項目について検索キ
ー文書と検索対象文書との類似度を項目単位で算出する
ように構成することで、各項目の文構造の違いが類似度
に影響する度合を軽減することができ、より算出される
類似度の妥当性を高くすることができる。
Further, means for unifying the sentence structure of each item of the retrieval key document or the retrieval target document is further added.
By configuring such that the similarity between the search key document and the search target document is calculated for each item whose sentence structure is unified in this way, the difference in the sentence structure of each item affects the similarity. The degree can be reduced, and the validity of the calculated similarity can be increased.

【図面の簡単な説明】[Brief description of the drawings]

【図1】本発明の一実施形態に係る類似文書検索装置の
ハードウェア構成を示す図。
FIG. 1 is a diagram showing a hardware configuration of a similar document search device according to an embodiment of the present invention.

【図2】図1の類似文書検索装置の制御部の構成を機能
別にブロック化して示す図。
FIG. 2 is a block diagram showing a configuration of a control unit of the similar document search apparatus of FIG.

【図3】図1の類似文書検索装置の処理の流れを示すフ
ローチャート。
FIG. 3 is a flowchart showing the flow of processing of the similar document search device of FIG. 1;

【図4】同じく図1の類似文書検索装置の処理の流れを
示すフローチャート。
FIG. 4 is a flowchart showing the flow of processing of the similar document search device of FIG. 1;

【図5】検索モードの設定画面を示す図。FIG. 5 is a diagram showing a search mode setting screen.

【図6】検索モード設定情報の格納例を示す図。FIG. 6 is a diagram showing a storage example of search mode setting information.

【図7】検索キー文書中の各項目の重要度を設定する画
面を示す図。
FIG. 7 is a diagram showing a screen for setting the importance of each item in a search key document.

【図8】検索キー文書の例を示す図。FIG. 8 is a diagram showing an example of a search key document.

【図9】図8の検索キー文書から切り出された項目単位
の文書を示す図。
9 is a diagram showing a document in item units cut out from the search key document of FIG. 8;

【図10】図9の項目文書から再生成された文書の例を
示す図。
FIG. 10 is a view showing an example of a document regenerated from the item document shown in FIG. 9;

【図11】図10に示す各項目の文書から抽出された単
語の情報の例を示す図。
11 is a diagram showing an example of word information extracted from the document of each item shown in FIG. 10;

【図12】図10に示す各項目の文書から抽出された単
語の出現頻度の情報の例を示す図。
FIG. 12 is a view showing an example of information on the frequency of appearance of words extracted from the document of each item shown in FIG. 10;

【図13】検索対象文書の例を示す図。FIG. 13 is a diagram showing an example of a search target document.

【図14】図13の検索対象文書から切り出された項目
単位の文書を示す図。
FIG. 14 is a diagram showing an item-based document cut out from the search target document in FIG. 13;

【図15】図14の項目文書から再生成された文書の例
を示す図。
FIG. 15 is a view showing an example of a document regenerated from the item document shown in FIG. 14;

【図16】図15に示す各項目の文書から抽出された単
語の情報の例を示す図。
FIG. 16 is a view showing an example of word information extracted from the document of each item shown in FIG. 15;

【図17】図15に示す各項目の文書から抽出された単
語の出現頻度の情報の例を示す図。
FIG. 17 is a diagram showing an example of information on the frequency of appearance of words extracted from the document of each item shown in FIG. 15;

【図18】共通単語の検索結果の例を示す図。FIG. 18 is a diagram illustrating an example of a search result of a common word.

【図19】検索キー文書/検索対象文書間の項目別の類
似度の算出結果を示す図。
FIG. 19 is a view showing a calculation result of similarity for each item between a search key document and a search target document.

【図20】項目別に設定された優先度の例を示す図。FIG. 20 is a diagram showing an example of priorities set for each item.

【図21】優先度による重みが付与された項目別の類似
度を示す図。
FIG. 21 is a diagram showing the similarity for each item to which a weight is assigned according to the priority;

【図22】項目別の類似文書検索結果を示す図。FIG. 22 is a view showing a similar document search result for each item.

【図23】図22の項目別の類似文書検索結果の出力例
を示す図。
FIG. 23 is a diagram showing an output example of a similar document search result for each item in FIG. 22;

【図24】検索キー文書の項目別に検索対象文書の各項
目の類似度の総和を求めた類似文書検索結果を示す図。
FIG. 24 is a view showing a similar document search result obtained by calculating the sum of similarities of the items of the search target document for each item of the search key document.

【図25】図24の類似文書検索結果の出力例を示す
図。
FIG. 25 is a diagram showing an output example of a similar document search result of FIG. 24;

【図26】検索対象文書毎に全項目の類似度の総和を求
めた類似文書検索結果を示す図。
FIG. 26 is a diagram illustrating a similar document search result obtained by calculating the sum of the similarities of all items for each search target document.

【図27】図26の類似文書検索結果の出力例を示す
図。
FIG. 27 is a view showing an output example of a similar document search result of FIG. 26;

【符号の説明】[Explanation of symbols]

1…制御装置 2…入力装置 3…表示装置 4…外部記憶装置 200…メイン処理部 201…初期化部 202…入力部 203…出力部 204…検索対象文書読み出し部 205…検索対象文書項目切り出し部 206…検索対象文書項目生成部 207…検索対象文書単語抽出部 208…検索対象単語出現頻度算出部 209…検索対象単語情報算出部 210…検索キー文書入力部 211…検索キー文書項目優先度入力部 212…検索キー文書項目切り出し部 213…検索キー文書項目生成部 214…検索キー単語抽出部 215…検索キー単語出現頻度算出部 216…検索条件設定部 217…共通単語抽出部 218…類似度算出部 219…検索結果出力部 229…検索条件設定バッファ部 230…検索対象文書格納バッファ部 231…検索対象文書項目格納バッファ部 232…検索対象文書項目生成格納バッファ部 233…検索対象単語情報格納バッファ部 234…検索キー文書項目優先度バッファ部 235…検索キー文書格納バッファ部 236…検索キー文書項目格納バッファ部 237…検索キー文書項目生成格納バッファ部 238…検索キー単語情報格納バッファ部 239…共通単語情報格納バッファ部 240…類似度格納バッファ部 241…検索結果出力バッファ部 242…作業バッファ部 REFERENCE SIGNS LIST 1 control device 2 input device 3 display device 4 external storage device 200 main processing unit 201 initialization unit 202 input unit 203 output unit 204 search document reading unit 205 search target document item cutout unit 206: search target document item generation unit 207: search target document word extraction unit 208: search target word appearance frequency calculation unit 209: search target word information calculation unit 210: search key document input unit 211: search key document item priority input unit 212 ... search key document item cutout unit 213 ... search key document item generation unit 214 ... search key word extraction unit 215 ... search key word appearance frequency calculation unit 216 ... search condition setting unit 217 ... common word extraction unit 218 ... similarity calculation unit 219 search result output section 229 search condition setting buffer section 230 search target document storage buffer section 231 Search target document item storage buffer unit 232 ... search target document item generation storage buffer unit 233 ... search target word information storage buffer unit 234 ... search key document item priority buffer unit 235 ... search key document storage buffer unit 236 ... search key document item Storage buffer unit 237 search key document item generation storage buffer unit 238 search key word information storage buffer unit 239 common word information storage buffer unit 240 similarity storage buffer unit 241 search result output buffer unit 242 work buffer unit

───────────────────────────────────────────────────── フロントページの続き (72)発明者 中本 幸夫 東京都青梅市新町3丁目3番地の1 東芝 コンピュータエンジニアリング株式会社内 (72)発明者 松隈 剛 東京都青梅市新町3丁目3番地の1 東芝 コンピュータエンジニアリング株式会社内 (72)発明者 仁科 卓哉 東京都青梅市新町3丁目3番地の1 東芝 コンピュータエンジニアリング株式会社内 Fターム(参考) 5B075 ND03 NK02 NK31 PP24 PQ36 PR04 PR08 QM08  ──────────────────────────────────────────────────続 き Continuation of the front page (72) Inventor Yukio Nakamoto 1-3-3 Shinmachi, Ome-shi, Tokyo Inside Toshiba Computer Engineering Co., Ltd. (72) Inventor Tsuyoshi Matsukuma 3-3-1 Shinmachi, Ome-shi, Tokyo Toshiba Computer Engineering Co., Ltd. (72) Inventor Takuya Nishina 3-3-1, Shinmachi, Ome-shi, Tokyo F-term within Toshiba Computer Engineering Co., Ltd. 5B075 ND03 NK02 NK31 PP24 PQ36 PR04 PR08 QM08

Claims (8)

【特許請求の範囲】[Claims] 【請求項1】 検索キー文書と類似した文書を複数の検
索対象文書の中から検索する類似文書検索装置におい
て、 前記検索キー文書および前記検索対象文書から項目の単
位の文書を切り出す項目切り出し手段と、 前記検索キー文書と前記検索対象文書との類似度を前記
切り出された項目の単位で算出し、その算出結果に基づ
いて類似文書検索結果を出力する計算手段とを具備する
ことを特徴とする類似文書検索装置。
1. A similar document search apparatus for searching a document similar to a search key document from a plurality of search target documents, comprising: an item cutout unit that cuts out a document of an item unit from the search key document and the search target document; Calculating means for calculating a similarity between the search key document and the search target document in units of the cut-out items, and outputting a similar document search result based on the calculation result. Similar document search device.
【請求項2】 前記項目毎に優先度を設定する手段をさ
らに有し、 前記計算手段は、前記設定された優先度を加味して前記
検索キー文書と前記検索対象文書との項目単位の類似度
を算出することを特徴とする請求項1記載の類似文書検
索装置。
2. The method according to claim 1, further comprising: setting a priority for each of the items, wherein the calculation unit considers the similarity of the search key document and the search target document in item units in consideration of the set priority. The similar document search apparatus according to claim 1, wherein the degree is calculated.
【請求項3】 前記項目切り出し手段によって切り出さ
れた前記検索キー文書または前記検索対象文書の各項目
の文構造を統一化させる手段をさらに有し、 前記計算手段は、前記文構造を統一化された各項目につ
いて、前記検索キー文書と前記検索対象文書との類似度
を前記項目単位で算出することを特徴とする請求項1記
載の類似文書検索装置。
3. The apparatus according to claim 1, further comprising a unit configured to unify a sentence structure of each item of the search key document or the search target document extracted by the item extraction unit, wherein the calculation unit unifies the sentence structure. 2. The similar document search apparatus according to claim 1, wherein a similarity between the search key document and the search target document is calculated for each item.
【請求項4】 前記計算手段は、 前記検索キー文書より切り出された項目からキー単語を
抽出し、抽出されたキー単語の項目毎の出現頻度を求め
る手段と、 前記検索対象文書より切り出された項目からキー単語を
抽出し、抽出されたキー単語の項目毎の出現頻度を求め
る手段と、 前記検索キー文書と前記検索対象文書に共通のキー単語
を検索する手段と、 前記検索キー文書および前記検索対象文書よりそれぞれ
前記抽出されたキー単語と該キー単語の項目毎の出現頻
度と前記検索された共通キー単語とから前記検索キー文
書と前記検索対象文書との項目単位の類似度を算出する
手段とを有することを特徴とする請求項1記載の類似文
書検索装置。
4. The calculation means includes: means for extracting a key word from an item cut out from the search key document, obtaining an appearance frequency for each item of the extracted key word, and information obtained by cutting out the search target document. Means for extracting a key word from an item, and calculating an appearance frequency of each item of the extracted key word; means for searching for a key word common to the search key document and the search target document; The degree of similarity of each of the search key document and the search target document is calculated based on the extracted key words from the search target document, the appearance frequency of each item of the key word, and the searched common key words. 2. A similar document search apparatus according to claim 1, further comprising:
【請求項5】 前記計算手段は、前記検索キー文書の項
目別に、最も類似度の高い検索対象文書を検索結果とし
て出力することを特徴とする請求項1記載の類似文書検
索装置。
5. The similar document search apparatus according to claim 1, wherein the calculation unit outputs a search target document having the highest similarity as a search result for each item of the search key document.
【請求項6】 前記計算手段は、前記検索キー文書の項
目別に、検索対象文書の各項目の類似度の総和を求め、
その結果を検索結果として出力することを特徴とする請
求項1記載の類似文書検索装置。
6. The calculation means calculates a sum of similarities of respective items of a search target document for each item of the search key document,
2. The similar document search device according to claim 1, wherein the result is output as a search result.
【請求項7】 前記計算手段は、検索対象文書毎に全項
目の類似度の総和を求め、その結果を検索結果として出
力することを特徴とする請求項1または2記載の類似文
書検索装置。
7. The similar document search device according to claim 1, wherein said calculation means calculates a total sum of similarities of all items for each search target document and outputs the result as a search result.
【請求項8】 検索キー文書と類似した文書を複数の検
索対象文書の中から検索する類似文書検索方法におい
て、 前記検索キー文書および前記検索対象文書から項目の単
位の文書を切り出す段階と、 前記検索キー文書と前記検索対象文書との類似度を前記
切り出された項目の単位で算出する段階と、 前記算出結果に基づいて類似文書検索結果を出力する段
階とを有することを特徴とする類似文書検索方法。
8. A similar document search method for searching a document similar to a search key document from a plurality of search target documents, wherein a document of an item unit is cut out from the search key document and the search target document. A similar document, comprising: calculating a similarity between a search key document and the search target document in units of the cut-out items; and outputting a similar document search result based on the calculation result. retrieval method.
JP11142448A 1999-05-21 1999-05-21 Similar document search device and similar document search method Withdrawn JP2000331027A (en)

Priority Applications (1)

Application Number Priority Date Filing Date Title
JP11142448A JP2000331027A (en) 1999-05-21 1999-05-21 Similar document search device and similar document search method

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
JP11142448A JP2000331027A (en) 1999-05-21 1999-05-21 Similar document search device and similar document search method

Publications (1)

Publication Number Publication Date
JP2000331027A true JP2000331027A (en) 2000-11-30

Family

ID=15315557

Family Applications (1)

Application Number Title Priority Date Filing Date
JP11142448A Withdrawn JP2000331027A (en) 1999-05-21 1999-05-21 Similar document search device and similar document search method

Country Status (1)

Country Link
JP (1) JP2000331027A (en)

Cited By (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2003281186A (en) * 2001-11-13 2003-10-03 Posco Example-based search method and search system for similarity determination
JP2009151746A (en) * 2007-12-20 2009-07-09 Inst For Information Industry Information resource collaborative tagging system and method
JP2011008334A (en) * 2009-06-23 2011-01-13 Nippon Hoso Kyokai <Nhk> Related content display device and computer program
US8045228B2 (en) 2007-03-19 2011-10-25 Ricoh Company, Ltd. Image processing apparatus

Cited By (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2003281186A (en) * 2001-11-13 2003-10-03 Posco Example-based search method and search system for similarity determination
US8045228B2 (en) 2007-03-19 2011-10-25 Ricoh Company, Ltd. Image processing apparatus
JP2009151746A (en) * 2007-12-20 2009-07-09 Inst For Information Industry Information resource collaborative tagging system and method
JP2011008334A (en) * 2009-06-23 2011-01-13 Nippon Hoso Kyokai <Nhk> Related content display device and computer program

Similar Documents

Publication Publication Date Title
JPH1145241A (en) Kana-kanji conversion system and computer-readable recording medium storing a program for causing a computer to function as each means of the system
JPH08305730A (en) Automatic method for selection of key phrase from document of machine-readable format to processor
JP2011018330A (en) System and method for transforming kanji into vernacular pronunciation string by statistical method
US7684975B2 (en) Morphological analyzer, natural language processor, morphological analysis method and program
JP2004318510A (en) Bilingual information creation device, bilingual information creating program, bilingual information creating method, bilingual information searching device, bilingual information searching program, and bilingual information searching method
JP2011065255A (en) Data processing apparatus, data name generation method and computer program
JPH11259515A (en) Similar document retrieval device and method and recording medium recording similar document retrieval program
JP2000331027A (en) Similar document search device and similar document search method
JPH1173415A (en) Similar document search device and similar document search method
JP3881638B2 (en) Document search apparatus, document search method, and document search program
JP3614765B2 (en) Concept dictionary expansion device
JP4524640B2 (en) Information processing apparatus and method, and program
JP4754849B2 (en) Document search device, document search method, and document search program
JP2001318947A (en) Information integration system and information integration method, and recording medium storing the program
JPH1145254A (en) Document retrieval apparatus and computer-readable recording medium recording a program for causing a computer to function as the apparatus
JP4304146B2 (en) Dictionary registration device, dictionary registration method, and dictionary registration program
JP2001249921A (en) Compound word analysis method and apparatus, and recording medium recording compound word analysis program
JP2001290826A (en) Document classification device, document classification method, and recording medium recording document classification program
JP2004062806A (en) Similar document search device and similar document search method
KR20220041336A (en) Graph generation system of recommending significant keywords and extracting core documents and method thereof
JP4682627B2 (en) Document retrieval apparatus and method
JP2002304407A (en) Program and information processing device
JP2002215672A (en) Search expression expansion method, search system, and search expression expansion computer program
JP2004234175A (en) Content search device and program thereof
JP2003242446A (en) Character string prediction apparatus and method, and computer-executable program embodying the method

Legal Events

Date Code Title Description
A300 Application deemed to be withdrawn because no request for examination was validly filed

Free format text: JAPANESE INTERMEDIATE CODE: A300

Effective date: 20060801