JPH11203309A - Method for preparing retrieval expression and device therefor - Google Patents

Method for preparing retrieval expression and device therefor

Info

Publication number
JPH11203309A
JPH11203309A JP10005129A JP512998A JPH11203309A JP H11203309 A JPH11203309 A JP H11203309A JP 10005129 A JP10005129 A JP 10005129A JP 512998 A JP512998 A JP 512998A JP H11203309 A JPH11203309 A JP H11203309A
Authority
JP
Japan
Prior art keywords
document
words
word
search
compound
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
JP10005129A
Other languages
Japanese (ja)
Inventor
Hiroyuki Nakajima
浩之 中島
Tsuyoshi Kitani
強 木谷
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
NTT Data Group Corp
Original Assignee
NTT Data Corp
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by NTT Data Corp filed Critical NTT Data Corp
Priority to JP10005129A priority Critical patent/JPH11203309A/en
Publication of JPH11203309A publication Critical patent/JPH11203309A/en
Pending legal-status Critical Current

Links

Landscapes

  • Machine Translation (AREA)
  • Information Retrieval, Db Structures And Fs Structures Therefor (AREA)

Abstract

(57)【要約】 【課題】 複数の単語から構成される複合語を考慮
して検索精度を一定値以上に維持可能な検索式を作成す
ることができる検索式作成装置を提供する。 【解決手段】 形態素解析部31、キーワード抽出部3
2、複合語処理部11、文書集合分割部33、検索式作
成部34の各機能を備えて検索式作成装置10を構成す
る。キーワード抽出部32より抽出された複数の単語に
おいて名詞句が連続する場合、対応する単語を複合語処
理部11において結合して複合語とし、これを、キーワ
ード抽出部32で出力される文書集合及び文書集合分割
部33における文書集合の分割に用いる。文書集合分割
の際に単語よりも複合語が有効であれば、この複合語を
検索キーワードとして決定し、検索式作成部34で作成
される検索式に反映させる。
(57) [Summary] [PROBLEMS] To provide a search formula creation device capable of creating a search formula capable of maintaining a search accuracy at a certain value or more in consideration of a compound word composed of a plurality of words. A morphological analysis unit and a keyword extraction unit.
2. The retrieval formula creation device 10 is provided with the functions of the compound word processing unit 11, the document set division unit 33, and the retrieval formula creation unit 34. When the noun phrases are continuous in a plurality of words extracted by the keyword extracting unit 32, the corresponding words are combined into a compound word in the compound word processing unit 11, and the combined words are combined into a document set output by the keyword extracting unit 32 and It is used for dividing a document set in the document set dividing unit 33. If a compound word is more effective than a word at the time of document set division, this compound word is determined as a search keyword, and is reflected in the search formula created by the search formula creation unit 34.

Description

【発明の詳細な説明】DETAILED DESCRIPTION OF THE INVENTION

【0001】[0001]

【発明の属する技術分野】本発明は、例えば大量に蓄積
された電子文書から特定の情報を索出する文書データベ
ースや、予め蓄積された電子文書例等を文書作成や発想
展開の支援のために利用する各種支援システム等に適用
される文書検索技術に係り、特に、電子文書中から抽出
したキーワードを用いて、検索者が関心のある文書の索
出を効率的に行うための検索式を試行錯誤的に作成する
手法に関する。
BACKGROUND OF THE INVENTION 1. Field of the Invention The present invention relates to a document database for retrieving specific information from a large amount of stored electronic documents, and a pre-stored example of an electronic document for supporting document creation and idea development. Related to the document search technology applied to various support systems to be used, in particular, using a keyword extracted from an electronic document, a search formula for a searcher to efficiently search for a document of interest is tried. It relates to a method of making mistakes.

【0002】[0002]

【従来の技術】検索対象となる電子文書を蓄積した文書
データベースから単語を抽出し、この単語を試行錯誤的
に組み合わて所要の検索式を作成する検索式作成装置が
知られている。図3は、従来のこの種の検索式作成装置
の機能構成図である。この検索式作成装置30は、コン
ピュータ装置が所定のプログラムを読み込んで実行する
ことにより形成される、形態素解析部31、キーワード
抽出部32、文書集合分割部33、及び検索式作成部3
4の機能ブロックを備えている。なお、文書には、それ
ぞれ検索者が関心のある必要文書か、関心のない不要文
書かを表す必要・不要の指定情報が付与されている。
2. Description of the Related Art There is known a retrieval formula creation apparatus which extracts words from a document database in which electronic documents to be searched are stored, and combines the words by trial and error to create a required search formula. FIG. 3 is a functional configuration diagram of this type of conventional search expression creating apparatus. The search formula creation device 30 is formed by a computer device reading and executing a predetermined program, and is formed by a morphological analysis unit 31, a keyword extraction unit 32, a document set division unit 33, and a search formula creation unit 3.
4 functional blocks. Note that each document is provided with necessary / unnecessary designation information indicating whether the document is a necessary document that the searcher is interested in or an unnecessary document that is not interested.

【0003】形態素解析部31は、複数の入力文書から
文書毎に形態素解析を行うものである。図中、符号31
Bは、形態素解析部31の出力例を示したものである。
なお、以下の説明では、下記内容の5つの文書が入力さ
れた場合を想定する。
[0003] The morphological analysis unit 31 performs morphological analysis for each document from a plurality of input documents. In the figure, reference numeral 31
B shows an output example of the morphological analysis unit 31.
In the following description, it is assumed that five documents having the following contents are input.

【0004】 文書番号1(必要文書)テレホーダイ(注)等の電話サ
ービス(注:日本電信電話株式会社の商標) 文書番号2(不要文書)地下鉄切符で美術館の入館料無
料 文書番号3(必要文書)様々な電話割引サービスが人気 文書番号4(不要文書)証券会社が電話で債券販売 文書番号5(必要文書)テレホーダイの加入者増加
Document No. 1 (necessary document) Telephone service such as Telehodai (note) (note: a trademark of Nippon Telegraph and Telephone Corporation) Document No. 2 (unnecessary document) Free admission to museums by subway ticket Document No. 3 (necessary document) Various telephone discount services are popular. Document number 4 (unnecessary document) Securities companies sell bonds by telephone Document number 5 (necessary document) Telehodai's subscribers increase

【0005】キーワード抽出部32は、個々の文書毎の
形態素群をもとにキーワードとして使用可能な単語を抽
出する。さらに、個々の文書における単語の出現の有無
を表す判別情報及び当該文書が必要文書か不要文書かを
判別するための指定情報を、文書名や文書番号等の文書
識別子と共に文書集合として出力する。符号32Bは、
キーワード抽出部32から出力される文書集合を例示し
たものであり、“1”〜“5”を文書識別子、“必要”
/“不要”が指定情報、“○”/“×”が判別情報であ
る。
[0005] A keyword extraction unit 32 extracts words that can be used as keywords based on a morpheme group for each document. Further, it outputs determination information indicating presence / absence of a word in each document and designation information for determining whether the document is a necessary document or an unnecessary document as a document set together with a document identifier such as a document name and a document number. Reference numeral 32B is
This is an example of a document set output from the keyword extraction unit 32. “1” to “5” are document identifiers, and “necessary”
/ "Unnecessary" is designation information, and "o" / "x" is discrimination information.

【0006】文書集合分割部33は、上記文書集合を判
別情報“○”/“×”に基づいて段階的に分割し、文書
検索に用いる検索式を作成する場合の基礎となる検索キ
ーワードを、抽出された単語群の中から決定する。この
場合、出来るだけ少数のキーワードの判別情報によって
文書集合を分割していくことで、必要文書と不要文書と
を区別した検索者の意図の抽出が可能となる。文書集合
分割部33で決定した検索キーワードは、検索式作成部
34において論理演算子“and”または“or”で結
合され、検索式として後続処理に出力される。
The document set division unit 33 divides the document set in stages based on the discrimination information “O” / “X”, and generates a search keyword as a basis for creating a search formula used for document search. It is determined from the extracted word group. In this case, by dividing the document set by the discrimination information of as few keywords as possible, it is possible to extract the intention of the searcher who distinguishes the necessary documents from the unnecessary documents. The search keywords determined by the document set division unit 33 are combined by the logical operator “and” or “or” in the search expression creation unit 34, and are output to the subsequent processing as a search expression.

【0007】文書集合分割部33における文書集合の分
割処理は、例えば良く知られたMDL(Minimum Descri
ption Length:最小記述長)原理に基づいて行われる。
このMDL原理は、「より多くの必要文書と不要文書と
をできるだけ少ないキーワードの組み合わせ(検索式)
で区別することにより、人間(検索者)の意図をより正
確に表現できる」とするヒューリスティックな手法であ
るが、このMDL原理を厳密に実現するには多くの処理
量が必要となるため、実際には処理量の軽減を図るため
に近似的に実現するのが一般的である。MDL原理を近
似的に実現する手法としては、例えば、決定木(論理式
を木構造で表現したもの)学習アルゴリズムである「I
D3」が知られている。この「ID3」については、
「知識獲得と学習シリーズ1:知識獲得入門」(Mic
halski,R.S.他編、共立出版)に詳細に記載
されている。
The document set dividing process in the document set dividing section 33 is performed, for example, by using a well-known MDL (Minimum Descri
ption Length: minimum description length).
This MDL principle is based on the idea that “more necessary documents and unnecessary documents are combined with as few keywords as possible (search expression).
Can be more accurately expressed the intention of a human (searcher) by discriminating with the "." However, since strict realization of the MDL principle requires a large amount of processing, Is generally realized approximately in order to reduce the processing amount. As a method of approximately realizing the MDL principle, for example, a decision tree (a logical expression represented by a tree structure) learning algorithm “I
D3 "is known. About this "ID3",
"Knowledge Acquisition and Learning Series 1: Introduction to Knowledge Acquisition" (Mic
halski, R .; S. Other editions, Kyoritsu Shuppan) are described in detail.

【0008】以下、この決定木学習アルゴリズム「ID
3」による文書集合の分割処理の概要を図4を参照して
説明する。まず、キーワード抽出部32から送られた文
書集合を初期文書集合Set0とする(ステップS20
1)。次に、初期文書集合Set0の“未分割”のフラ
グをオンにする(ステップS202)。これをSeti
とする(ステップS203)。次に、この文書集合Se
i中の必要文書、不要文書に含まれる各キーワードtj
(1≦j≦N)について、文書全体の情報量に対する個
別文書の情報量の相対関係を表す相互情報量I(tj)を
算出する(ステップS204)。相互情報量I(tj)
は、具体的には、未分割の文書集合についての情報量H
からキーワードtjが含まれた文書集合及び含まれない
文書集合についての情報量H(tj)を差し引いた以下の
式(1)で表される。 I(tj)=H−H(tj) (1)
Hereinafter, this decision tree learning algorithm “ID
The outline of the document set division process by “3” will be described with reference to FIG. First, the document set sent from the keyword extraction unit 32 is set as an initial document set Set 0 (step S20).
1). Next, the “undivided” flag of the initial document set Set 0 is turned on (step S202). This is Set i
(Step S203). Next, this document set Se
Necessary documents in t i , each keyword t j included in unnecessary documents
For (1 ≦ j ≦ N), a mutual information amount I (t j ) representing the relative relationship between the information amount of the individual document and the information amount of the entire document is calculated (step S204). Mutual information I (t j )
Is, specifically, the information amount H about the undivided document set.
Is subtracted from the information amount H (t j ) of the document set including the keyword t j and the document set not including the keyword t j from the following expression (1). I (t j ) = H−H (t j ) (1)

【0009】但し、式(1)におけるパラメータは下記
のようになる。 pi:Seti中の必要文書数、 ni:Seti中の不要文書数、 si:pi+ni、i(tj):Seti中でキーワードtjを含む必要文書
数、 ni(tj):Seti中でキーワードtjを含む不要文書
数、 si(tj):pi(tj)+ni(tj)、 pi not(tj):Seti中でキーワードtjを含まない
必要文書数、 ni not(tj):Seti中でキーワードtjを含まない
不要文書数、 si not(tj):pi not(tj)+ni not(tj)、 h(a,b,c):-{a/c・log2(a/c)+b/c・log2(b/c)}
However, the parameters in equation (1) are as follows. p i: Set i need the number of documents in, n i: Set i unnecessary number of documents in, s i: p i + n i, p i (t j): necessary number of documents that contain the keyword t j in Set i , n i (t j): Set i unnecessary number of documents that contain the keyword t j in, s i (t j): p i (t j) + n i (t j), p i not (t j): Set i need the number of documents that do not contain the keyword t j in, n i not (t j) : Set i unnecessary number of documents that do not contain the keyword t j in, s i not (t j) : p i not (t j) + N i not (t j ), h (a, b, c):-{a / c · log 2 (a / c) + b / c · log 2 (b / c)}

【0010】また、各情報量H及びH(tj)は、各々下
記の式(2)、式(3)で表される。
The information amounts H and H (t j ) are expressed by the following equations (2) and (3), respectively.

【0011】[0011]

【数1】 (Equation 1)

【0012】次に、複数のキーワードtjから相互情報
量I(tk)の値を最大にすることが可能なキーワードt
kを選択し、これを検索キーワードとする(ステップS
205)。この相互情報量I(tk)が正の有限値(>
0)の場合(ステップS206)、検索キーワードtk
を含む文書の番号からなる文書集合をSeti′、検索キ
ーワードtkを含まない文書の番号からなる文書集合を
Seti″として分割し、分割したそれぞれの文書集合
の“未分割”のフラグをオンにする(ステップS207
〜S210)。i′,i″は、既に文書集合Seti′、
Seti″が存在しなければ任意の値で良い。一方、相
互情報量I(tk)がゼロ値(=0)の場合は文書集合の
分割を行わない(ステップS206)。
Next, a keyword t that can maximize the value of the mutual information I (t k ) from a plurality of keywords t j
k is selected and set as a search keyword (step S
205). This mutual information I (t k ) has a positive finite value (>
In the case of 0) (step S206), the search keyword t k
Set i 'a document set consisting of number of documents containing the flag of the search keyword set of documents consisting of number of documents that do not contain t k Set i "divided as, for each document set divided" undivided " Turn on (step S207)
To S210). i ′, i ″ are already document sets Set i ′,
If Set i ″ does not exist, any value may be used.On the other hand, if the mutual information I (t k ) is a zero value (= 0), the document set is not divided (step S206).

【0013】その後、集合Setiの“未分割”のフラ
グをオフにする(ステップS211)。“未分割”のフ
ラグがオンの文書集合がある場合はステップS103に
戻り(ステップS212:Yes)、“未分割”のフラグ
がオンの文書集合がなくなるまで処理を繰り返す。そし
て、すべての文書集合についての“未分割”のフラグが
オフになった時点で処理を終える(ステップS212:
No)。
Thereafter, the flag of “undivided” of the set Set i is turned off (step S211). If there is a document set with the “undivided” flag on, the process returns to step S103 (step S212: Yes), and the process is repeated until there is no document set with the “undivided” flag on. Then, the process ends when the “undivided” flag is turned off for all the document sets (step S212:
No).

【0014】また、上記アルゴリズム「ID3」は、例
えば、公知のアルゴリズムである「C4.5」等による
代用も可能である。なお、「C4.5」のアルゴリズム
については、「C4.5 Programs for Machine Learning」
(Quinlan、J.R.著、Morgan Kaufmann Publishers 刊)の
記載を参考にすることができる。
The above-mentioned algorithm "ID3" can be replaced with, for example, a known algorithm such as "C4.5". The algorithm of “C4.5” is described in “C4.5 Programs for Machine Learning”.
(Quinlan, JR, published by Morgan Kaufmann Publishers).

【0015】図5は、上記検索式作成装置30におい
て、一つの文書集合が複数の文書集合に分割され、検索
式が試行錯誤的に作成されていく過程を示した図であ
る。以下、図5を参照して、従来の検索式作成手法の概
要を説明する。まず、キーワード抽出部32で生成され
る初期文書集合Set0から(符号32B参照)、決定
木学習アルゴリズム「ID3」に基づいて相互情報量が
最大となるキーワードを決定し、これを検索キーワード
とする。ここでは、検索キーワード「テレホーダイ」が
決定されたとする。
FIG. 5 is a diagram showing a process in which one document set is divided into a plurality of document sets and the search formula is created by trial and error in the search formula creating apparatus 30. Hereinafter, an outline of a conventional search formula creation method will be described with reference to FIG. First, from the initial document set Set 0 generated by the keyword extraction unit 32 (see reference numeral 32B), a keyword having the maximum mutual information is determined based on the decision tree learning algorithm “ID3”, and this is set as a search keyword. . Here, it is assumed that the search keyword “Telehodai” has been determined.

【0016】文書集合分割部33は、この検索キーワー
ド「テレホーダイ」によって、初期文書集合Set0
を、当該検索キーワードを含む必要文書の集合Set1
と、含まない必要文書及び不要文書の集合Set2とに
分割する。文書集合Set1は、これ以上の分割は不可
能である。一方、文書集合Set2はさらなる分割が可
能である。そこで、この文書集合Set2において相互
情報量が最大となる検索キーワード「割引」を決定し、
当該検索キーワードによって文書集合Set2を、検索
キーワード「割引」を含まない不要文書の集合Set3
と、含む必要及び不要文書の集合Set4とに分割す
る。
The document set division unit 33 uses the search keyword “Telehodai” to generate an initial document set Set0.
Is a set of required documents Set 1 including the search keyword.
And a set Set 2 of necessary documents and unnecessary documents that are not included. The document set Set 1 cannot be further divided. On the other hand, the document set Set 2 can be further divided. Therefore, a search keyword “discount” that maximizes the mutual information in the document set Set 2 is determined,
A document set Set 2 is set by the search keyword, and a set Set 3 of unnecessary documents not including the search keyword “discount”.
And a set Set 4 of necessary and unnecessary documents to be included.

【0017】文書集合Set4は、さらなる分割が可能
なので、この文書集合Set4において相互情報量が最
大となるキーワード「電話」を検索キーワードとして決
定し、当該検索キーワードを含む必要文書の集合Set
5と、含まない文書の集合Set6とに分割する。文書集
合Set5及びSet6は、共にこれ以上の分割が不可能
であるため、分割処理を終える。
Since the document set Set 4 can be further divided, a keyword “telephone” having the maximum mutual information in the document set Set 4 is determined as a search keyword, and a set Set of necessary documents including the search keyword is set.
5 and a set Set 6 of documents not included. Since the document sets Set 5 and Set 6 cannot be further divided, the division processing ends.

【0018】上記分割処理において決定された各検索キ
ーワード「テレホーダイ」、「割引」、及び「電話」
は、逐次図示しない記憶手段に保持され、分割処理が終
了した時点で検索式作成部34に渡される。検索式作成
部34では、文書集合分割部33の結果である各検索キ
ーワードを、論理演算子“and”、“or”、“no
t”により結合して検索式queryを作成する。図3
の符号34Bは、検索式作成部34から出力される検索
式を例示したものである。
Each of the search keywords "tele-hodai", "discount", and "telephone" determined in the above-described division processing.
Are sequentially stored in a storage unit (not shown), and are passed to the search formula creation unit 34 when the division processing is completed. In the search formula creation unit 34, each search keyword that is the result of the document set division unit 33 is defined by the logical operators "and", "or", "no".
The search expression "query" is created by combining with "t". FIG.
Is an example of a search formula output from the search formula creating unit.

【0019】[0019]

【発明が解決しようとする課題】ところで、上述の従来
の検索式作成装置30では、形態素解析処理に基づいて
抽出された単語群を文書の属性として用いている。その
ため、形態素解析に起因して、例えば、複数の単語から
なる語句について特定の意味が想起される場合(以下、
このような語句を複合語と称する)であってもそれが個
々の構成単語に分割してしまったり、一つの単語が複数
の単語に誤って分割されてしまうことがあり、検索キー
ワードとしての有益性が損なわれるという問題があっ
た。このことを、下記内容の3つの文書が入力された場
合を例に挙げて説明する。 文書番号1(必要文書)電話の設置には施設負担金が必
要だ・・・ 文書番号2(不要文書)競技施設設置の負担金の支払い
を電話で催促された・・・ 文書番号3(必要文書)電話設置に施設負担金を必要と
しない
By the way, in the above-mentioned conventional retrieval formula creation device 30, the word group extracted based on the morphological analysis processing is used as the attribute of the document. Therefore, for example, when a specific meaning is recalled for a phrase including a plurality of words due to the morphological analysis (hereinafter, referred to as a phrase).
Even if such a phrase is called a compound word), it may be divided into individual constituent words, or one word may be erroneously divided into a plurality of words, which is useful as a search keyword. There was a problem that the property was impaired. This will be described with an example in which three documents having the following contents are input. Document number 1 (necessary document) Facility installation fee is required to set up the telephone ... Document number 2 (unnecessary document) Payment of the payment of the competition facility installation was prompted by telephone ... Document number 3 (necessary) Document) Telephone installation does not require facility contribution

【0020】この例では、キーワード抽出部32から図
6のような内容の文書集合が得られる。この場合、文書
番号1,3の文書中に「施設負担金」のような複合語が
存在しているが、形態素解析によって得られる単語は
「施設」と「負担金」であり、これらを単に組み合わせ
るだけでは適切な検索式が得られない。また、この例で
は、抽出された各単語がすべての文書中に含まれること
から、検索キーワードを決定して必要文書及び不要文書
を区別する検索式を迅速に作成することは困難となる。
In this example, a set of documents having contents as shown in FIG. 6 is obtained from the keyword extracting unit 32. In this case, a compound word such as “facility contribution” exists in the documents of document numbers 1 and 3, but the words obtained by the morphological analysis are “facility” and “payment”. An appropriate search formula cannot be obtained just by combining them. Further, in this example, since the extracted words are included in all the documents, it is difficult to determine a search keyword and quickly create a search formula for distinguishing necessary documents and unnecessary documents.

【0021】そこで本発明の課題は、文書検索等におけ
る検索精度を一定値以上に維持するとともに、複合語を
分割することなく検索キーワードの決定及び検索式の作
成を迅速に行うことができる、改良された検索式作成方
法を提供することにある。本発明の他の課題は、上記検
索式作成方法の実施に適した検索式作成装置を提供する
ことにある。
SUMMARY OF THE INVENTION An object of the present invention is to improve the search accuracy in a document search or the like at a fixed value or more, and to quickly determine a search keyword and create a search expression without dividing a compound word. It is an object of the present invention to provide a search method creating method. Another object of the present invention is to provide a search formula creation device suitable for implementing the above search formula creation method.

【0022】[0022]

【課題を解決するための手段】上記課題を解決する本発
明の検索式作成方法は、コンピュータ装置を用いた検索
式作成方法であって、検索式作成のために入力された指
定文書群に形態素解析を施して文書単位で単語を抽出す
るとともに、名詞句に相当する単語が連続する場合の当
該単語群を一意の複合語として特定する過程と、抽出さ
れた個々の単語及び特定された前記複合語について、こ
れらの単語または複合語を含む文書群及び含まない文書
群の情報量と前記指定文書群の総情報量との差分で表さ
れる相互情報量を算出し、この相互情報量を最大にする
個々の単語または複合語を検索キーワードとして決定す
る過程とを含み、決定した検索キーワードを要素とする
検索式を作成することを特徴とする。
According to a first aspect of the present invention, there is provided a method for creating a search formula using a computer, wherein a specified document group input for creating the search formula includes a morpheme. Analyzing and extracting words in document units, specifying a group of words as unique compound words when words corresponding to noun phrases are continuous, and extracting the extracted individual words and the specified compound For a word, a mutual information amount represented by a difference between the information amount of the document group including and not including the word or the compound word and the total information amount of the specified document group is calculated, and the mutual information amount is set to the maximum. Deciding an individual word or compound word to be used as a search keyword, and creating a search formula having the determined search keyword as an element.

【0023】また、上記他の課題を解決する本発明の検
索式作成装置は、検索者にとって関心のある必要文書及
び関心のない不要文書を含む文書群の文書検索に用いる
検索式を作成する装置であって、前記必要文書または不
要文書を識別するための指定情報が付与された指定文書
群から文書毎に単語抽出を行う単語抽出手段と、抽出さ
れた個々の単語の品詞を判定し、名詞句に相当する単語
が連続する場合の当該単語群を一意の複合語として特定
する複合語処理手段と、個々の単語及び前記複合語が文
書中に含まれるか否かを表す判別情報及び文書に付与さ
れた前記指定情報を文書識別情報と共に集合させた文書
集合を生成する文書集合生成手段と、前記抽出された単
語及び前記特定された前記複合語のうち、これらの単語
または複合語を含む文書群及び含まない文書群の情報量
と前記指定文書群の総情報量との差分で表される相互情
報量を最大にする個々の単語または複合語を検索キーワ
ードとして決定するとともに、決定した検索キーワード
を用いて一つの文書集合を複数の文書集合に分割する文
書集合分割手段とを備え、前記文書集合の分割を繰り返
す度に決定された検索キーワードを論理式で結合して前
記検索式を作成することを特徴とする。
According to another aspect of the present invention, there is provided a search formula generating apparatus for generating a search formula used for document search of a document group including necessary documents of interest to a searcher and unnecessary documents of no interest. Word extracting means for extracting a word for each document from a specified document group to which specified information for identifying the required document or the unnecessary document has been added, and determining the part of speech of each extracted word, Compound word processing means for specifying a group of words as a unique compound word in the case where words corresponding to a phrase are continuous; discriminating information indicating whether each word and the compound word are included in the document; A document set generating means for generating a document set in which the assigned designation information is collected together with document identification information; and a document set including these words or compound words among the extracted words and the specified compound words. Individual words or compound words that maximize the mutual information represented by the difference between the information amount of the document group and the document group not included and the total information amount of the specified document group are determined as search keywords, and the determined search is performed. Document set dividing means for dividing one document set into a plurality of document sets by using a keyword, wherein the search keywords determined each time the division of the document set is repeated are combined with a logical expression to create the search expression It is characterized by doing.

【0024】この検索式作成装置において、前記複合語
処理手段を、抽出された単語を予め設定された複合語作
成基準に基づいて所定順に結合し、これにより得られた
単語群を一意の複合語として特定するように構成しても
良い。
[0024] In the retrieval formula creating apparatus, the compound word processing means combines the extracted words in a predetermined order based on a preset compound word creating criterion, and converts the obtained word group into a unique compound word. It may be configured to be specified as

【0025】前記文書集合分割手段は、例えば、所定の
最小記述長原理に基づいて前記相互情報量を最大とする
単語または複合語を前記検索キーワードとして逐次決定
するように構成する。
The document set dividing means is configured to successively determine, as the search keyword, a word or a compound word that maximizes the mutual information based on, for example, a predetermined minimum description length principle.

【0026】[0026]

【発明の実施の形態】以下、本発明の実施の形態を詳細
に説明する。図1は、上記検索式の作成方法の実施に適
した検索式作成装置を示す機能構成図であり、図3で説
明した従来の検索式作成装置30と同一機能の構成要素
については、同一符号を付して重複説明を省略する。ま
た、説明の便宜上、本装置に入力される文書群には、利
用者等にとって必要文書か不要文書かを表す必要・不要
の指定情報が予め付与されているものとする。
Embodiments of the present invention will be described below in detail. FIG. 1 is a functional block diagram showing a search formula creation device suitable for implementing the above search formula creation method. Components having the same functions as those of the conventional search formula creation device 30 described with reference to FIG. And a duplicate description is omitted. For convenience of explanation, it is assumed that a document group input to the present apparatus is given in advance with necessary / unnecessary designation information indicating whether the document is necessary or unnecessary for a user or the like.

【0027】本実施形態の検索式作成装置10は、コン
ピュータ装置が所定のプログラムを読み込んで実行する
ことにより形成される、形態素解析部31、キーワード
抽出部32、複合語処理部11、文書集合分割部33、
検索式作成部34の各機能を備えて構成される。上記プ
ログラムは、通常、コンピュータ装置の内部記憶装置あ
るいは外部記憶装置に格納され、随時読み取られて実行
されるようになっているが、コンピュータ装置とは分離
可能な記録媒体、例えばCD−ROMやFD等の可搬性
記録媒体、あるいは当該コンピュータ装置と構内ネット
ワークに接続されたプログラムサーバ等に格納され、使
用時に上記内部記憶装置または外部記憶装置にインスト
ールされて随時実行に供されるものであってもよい。
The retrieval formula creation device 10 of the present embodiment includes a morphological analysis unit 31, a keyword extraction unit 32, a compound word processing unit 11, and a document set division formed by reading and executing a predetermined program by a computer device. Part 33,
It is provided with each function of the search formula creation unit 34. The above program is usually stored in an internal storage device or an external storage device of the computer device, and is read and executed as needed. However, a recording medium separable from the computer device, for example, a CD-ROM or FD Etc., or stored in a program server or the like connected to the computer device and the local network, and installed in the internal storage device or the external storage device at the time of use, and provided for execution at any time. Good.

【0028】複合語処理部11は、形態素解析部31及
びキーワード抽出部32において抽出された単語群の品
詞を判定し、一意な複合語の特定を行うものである。こ
の複合語の特定は、例えば、品詞判定の結果、名詞句に
相当する単語群が連続して抽出される場合に、対応する
複数の単語を結合することにより行われるようにする。
あるいは、所定の複合語作成基準をシステムパラメータ
等で設定しておき、当該複合語作成基準に基づいて対応
する各単語を所定の順序で連続して結合するようにす
る。この場合の複合語作成基準としては種々の形態が考
えられるが、一例としては、予め複合語として使用する
予定の単語の組み合わせ手順を設定しておき、この手順
に則って単語を組み合わせるようにする。あるいは名詞
句が連続するかどうかに関わらず、所定個数の名詞句に
相当する単語を組み合わせるようにしても良い。このよ
うにして特定された複合語は、文書集合の際に用いられ
る。
The compound word processing unit 11 determines the part of speech of the word group extracted by the morphological analysis unit 31 and the keyword extraction unit 32, and specifies a unique compound word. The specification of the compound word is performed, for example, by combining a plurality of corresponding words when a word group corresponding to a noun phrase is continuously extracted as a result of the part-of-speech determination.
Alternatively, a predetermined compound word creation criterion is set by a system parameter or the like, and corresponding words are successively combined in a predetermined order based on the compound word creation criterion. In this case, various forms can be considered as a compound word creation criterion. For example, a combination procedure of words to be used as a compound word is set in advance, and words are combined in accordance with this procedure. . Alternatively, words corresponding to a predetermined number of noun phrases may be combined regardless of whether the noun phrases are continuous. The compound words specified in this way are used in document collection.

【0029】この複合語が文書集合分割部33における
必要文書と不要文書とを区別する際に有効となる場合、
つまり前述した相互情報量が大きくなる場合には、当該
複合語が検索キーワードとして決定され、検索式作成部
34で作成される検索式に反映される。複数の単語また
は複合語が検索キーワードの候補となるような場合は、
例えば、予め保持した個々の単語の文書中における頻度
情報に基づいて、文書数がより小さくなる、即ち出現頻
度が小さくなる単語または複合語を検索キーワードとし
て決定するように適宜構成する。
When this compound word is effective in distinguishing a necessary document from an unnecessary document in the document set dividing unit 33,
That is, when the mutual information amount described above increases, the compound word is determined as a search keyword, and is reflected in the search formula created by the search formula creating unit 34. If multiple words or compound words are suggested search terms,
For example, based on frequency information of individual words held in a document in advance, a word or compound word having a smaller number of documents, that is, a word having a lower appearance frequency is appropriately determined as a search keyword.

【0030】次に、上記構成の検索式作成装置10を用
いた検索式作成方法を図2を参照して説明する。ここで
は、便宜上、下記内容の3つの文書が入力されたとす
る。 文書番号1(必要文書)電話の設置には施設負担金が必
要だ・・・ 文書番号2(不要文書)競技施設設置の負担金の支払い
を電話で催促された・・・ 文書番号3(必要文書)電話設置に施設負担金を必要と
しない
Next, a description will be given of a method for creating a retrieval formula using the retrieval formula producing device 10 having the above configuration with reference to FIG. Here, it is assumed that three documents having the following contents are input for convenience. Document number 1 (necessary document) Facility installation fee is required to set up the telephone ... Document number 2 (unnecessary document) Payment of the payment of the competition facility installation was prompted by telephone ... Document number 3 (necessary) Document) Telephone installation does not require facility contribution

【0031】上記文書が入力されると(ステップS10
1)、検索式作成装置10は、入力文書に対して形態素
解析部31で形態素解析を施し、文書毎の形態素群を抽
出する(ステップS102)。この形態素解析の結果を
示したのが図1の符号31Aである。キーワード抽出部
32は、文書毎の形態素群をもとにキーワードとして使
用可能な単語を抽出する(ステップS103)。また、
名詞句が連続する場合に(ステップS104:Yes)、
対応する単語を結合して複合語とする(ステップS10
5)。本例では、「施設」、「負担金」のように、連続
して抽出された単語の結合が複合語「施設負担金」とし
て特定される。キーワード抽出部32では、また、個々
の文書における単語、及びステップS105で生成した
複合語の出現の有無を表す判別情報及び当該文書が必要
文書か不要文書かを表す識別情報を、文書名や文書番号
等の文書識別子と共に集合させ、文書集合を作成する
(ステップS106)。図1の符号32Aは、キーワー
ド抽出部32から出力される文書集合の内容を例示した
ものである。この文書集合では、複合語「施設負担金」
が必要文書と不要文書とを区別するうえで有効なキーワ
ードとなっていることがわかる。
When the above document is input (step S10)
1) The search formula creation device 10 performs morphological analysis on the input document by the morphological analysis unit 31 and extracts a morpheme group for each document (step S102). Reference numeral 31A in FIG. 1 shows the result of the morphological analysis. The keyword extraction unit 32 extracts words that can be used as keywords based on the morpheme group of each document (step S103). Also,
If the noun phrases are consecutive (step S104: Yes),
The corresponding words are combined into a compound word (step S10
5). In this example, a combination of words extracted consecutively, such as “facility” and “contribution”, is specified as a compound word “facility contribution”. The keyword extraction unit 32 also stores a word in each document, discrimination information indicating whether or not the compound word generated in step S105 appears, and identification information indicating whether the document is a required document or an unnecessary document. The documents are collected together with document identifiers such as numbers to create a document set (step S106). Reference numeral 32A in FIG. 1 exemplifies the contents of a document set output from the keyword extracting unit 32. In this document set, the compound term "facility contribution"
It can be seen that is an effective keyword for distinguishing a necessary document from an unnecessary document.

【0032】その後、文書集合分割部33において、ス
テップS106で作成された文書集合を前述の相互情報
量に基づいて段階的に分割するとともに(ステップS1
07)、検索式作成部34で、この分割処理の過程にお
いて逐次決定される検索キーワードに基づく検索式を作
成する(ステップS108)。図1の符号34Aは、検
索式作成部34から出力される検索式を例示したもので
ある。この例では、検索キーワードとして決定された複
合語「施設負担金」のみで検索式が作成されることを表
している。新規に検索式作成対象となる文書の入力があ
る場合にはステップS101に戻り(ステップS10
9:Yes)、上記一連の処理を繰り返す。他の入力文書
がない場合は処理を終了する(ステップS109:N
o)。
Thereafter, the document set division unit 33 divides the document set created in step S106 stepwise based on the above-mentioned mutual information (step S1).
07), the search formula creation unit 34 creates a search formula based on the search keywords sequentially determined in the process of the division process (step S108). Reference numeral 34A in FIG. 1 exemplifies a search formula output from the search formula creation unit 34. This example shows that a search formula is created using only the compound word “facility contribution” determined as a search keyword. If there is an input of a new document for which a search expression is to be created, the process returns to step S101 (step S10).
9: Yes), the above series of processing is repeated. If there is no other input document, the process ends (step S109: N
o).

【0033】このように、本実施形態の検索式作成装置
10では、連続して抽出される名詞句に相当する単語の
結合により一意な複合語を特定し、これを文書集合の作
成及び検索キーワードに用いるようにしたので、特定の
単語が複数の単語に分割されてしまうという形態素解析
に起因する問題を回避することができ、検索語としての
キーワードの有益性が保証される。
As described above, the retrieval formula creating apparatus 10 of the present embodiment specifies a unique compound word by combining words corresponding to noun phrases that are successively extracted, and generates a unique compound word by creating a document set and a search keyword. , It is possible to avoid a problem caused by morphological analysis in which a specific word is divided into a plurality of words, and the usefulness of a keyword as a search word is guaranteed.

【0034】また、従来、分割された単語を組み合わせ
るだけでは得られなかった適切な検索式が、複合語を検
索式に用いることで容易に取得可能となり、より少ない
検索キーワードによって必要文書と不要文書との区別が
できるようになった。このことから、「より少ない検索
語で必要文書と不要文書とを区別する検索式ほど人間
(検索者)の意図を正確に表現する検索式である」とい
う前述のMDL原理の仮定により、取得される検索式
は、従来手法と比較して正確なものとなる。さらに、作
成された検索式を検索処理に適用することにより検索精
度の高い結果が取得可能となった。
In addition, an appropriate search formula that could not be obtained by simply combining divided words can be easily obtained by using a compound word as a search formula, and required documents and unnecessary documents can be obtained by using fewer search keywords. And can be distinguished. From this, it is obtained based on the above-mentioned assumption of the MDL principle that "a search formula that distinguishes necessary documents from unnecessary documents with fewer search words is a search formula that accurately expresses the intention of a human (searcher)." The retrieval formula is more accurate than the conventional method. Furthermore, by applying the created search formula to the search processing, a result with high search accuracy can be obtained.

【0035】[0035]

【発明の効果】以上の説明から明らかなように、本発明
によれば、文書データベース全体を分割することなくよ
り少ない検索キーワードによる検索式の作成が可能とな
る効果がある。また、本発明により得られる検索式を用
いることで、文書の検索精度を一定値以上に維持するこ
とが可能となり、検索処理の効率が大幅に向上するとい
う効果もある。
As is apparent from the above description, according to the present invention, it is possible to create a search formula with fewer search keywords without dividing the entire document database. Further, by using the retrieval formula obtained by the present invention, it is possible to maintain the retrieval accuracy of the document at a certain value or more, and there is also an effect that the efficiency of the retrieval processing is greatly improved.

【図面の簡単な説明】[Brief description of the drawings]

【図1】本発明の一実施形態に係る検索式作成装置の実
施形態を表す機能ブロック図。
FIG. 1 is a functional block diagram illustrating an embodiment of a search formula creation device according to an embodiment of the present invention.

【図2】本実施形態の検索式作成装置における処理手順
図。
FIG. 2 is a processing procedure diagram in the search expression creating apparatus of the embodiment.

【図3】従来の検索式作成装置の機能ブロック図。FIG. 3 is a functional block diagram of a conventional search expression creation device.

【図4】従来の検索式作成装置における処理手順説明
図。
FIG. 4 is an explanatory diagram of a processing procedure in a conventional search expression creation device.

【図5】従来の分割処理過程で得られる情報の模式図。FIG. 5 is a schematic diagram of information obtained in a conventional dividing process.

【図6】従来の入力文書群に対応する文書集合の作成結
果。
FIG. 6 shows a document set creation result corresponding to a conventional input document group.

【符号の説明】[Explanation of symbols]

10,30 検索式作成装置 11 複合語処理部 31 形態素解析部 32 キーワード抽出部 33 文書集合分割部 34 検索式作成部 10, 30 retrieval formula creation device 11 compound word processing unit 31 morphological analysis unit 32 keyword extraction unit 33 document set division unit 34 search formula creation unit

Claims (6)

【特許請求の範囲】[Claims] 【請求項1】 検索式作成のために入力された指定文書
群に形態素解析を施して文書単位で単語を抽出するとと
もに、名詞句に相当する単語が連続する場合の当該単語
群を一意の複合語として特定する過程と、 抽出された個々の単語及び特定された前記複合語につい
て、これらの単語または複合語を含む文書群及び含まな
い文書群の情報量と前記指定文書群の総情報量との差分
で表される相互情報量を算出し、この相互情報量を最大
にする個々の単語または複合語を検索キーワードとして
決定する過程とを含み、 決定した検索キーワードを要素とする検索式を作成する
ことを特徴とする、 コンピュータ装置を用いた検索式作成方法。
1. A morphological analysis is performed on a designated document group input for creating a retrieval formula to extract words in document units, and a word group in the case where words corresponding to a noun phrase are continuous is uniquely compounded. Specifying the word as a word, and for the extracted individual words and the specified compound word, the information amount of the document group including and not including the word or compound word and the total information amount of the specified document group Calculating the mutual information represented by the difference between the two, and determining individual words or compound words that maximize the mutual information as search keywords, and creating a search formula using the determined search keywords as elements. A method for creating a retrieval formula using a computer device.
【請求項2】 前記指定文書群が、検索者にとって関心
のある必要文書または関心のない不要文書を識別するた
めの指定情報が付与された文書群であり、個々の文書に
付与された前記指定情報が前記相互情報量に反映されて
いることを特徴とする請求項1記載の検索式作成方法。
2. The designated document group is a document group to which designated information for identifying a necessary document of interest or an unnecessary document not of interest to a searcher is added, and the designated document assigned to each document is provided. 2. The method according to claim 1, wherein information is reflected in the mutual information amount.
【請求項3】 前記複合語は、予め設定された複合語作
成基準に基づいて抽出された単語群を所定順に結合した
ものであることを特徴とする請求項1記載の検索式作成
方法。
3. The method according to claim 1, wherein the compound word is obtained by combining word groups extracted based on a predetermined compound word creation criterion in a predetermined order.
【請求項4】 検索者にとって関心のある必要文書及び
関心のない不要文書を含む文書群の文書検索に用いる検
索式を作成する装置であって、 前記必要文書または不要文書を識別するための指定情報
が付与された指定文書群から文書毎に単語抽出を行う単
語抽出手段と、 抽出された個々の単語の品詞を判定し、名詞句に相当す
る単語が連続する場合の当該単語群を一意の複合語とし
て特定する複合語処理手段と、 個々の単語及び前記複合語が文書中に含まれるか否かを
表す判別情報及び文書に付与された前記指定情報を文書
識別情報と共に集合させた文書集合を生成する文書集合
生成手段と、 前記抽出された単語及び前記特定された前記複合語のう
ち、これらの単語または複合語を含む文書群及び含まな
い文書群の情報量と前記指定文書群の総情報量との差分
で表される相互情報量を最大にする個々の単語または複
合語を検索キーワードとして決定するとともに、決定し
た検索キーワードを用いて一つの文書集合を複数の文書
集合に分割する文書集合分割手段とを備え、 前記文書集合の分割を繰り返す度に決定された検索キー
ワードを論理式で結合して前記検索式を作成することを
特徴とする検索式作成装置。
4. An apparatus for creating a retrieval formula for use in document retrieval of a group of documents including a required document of interest and an unnecessary document of no interest to a searcher, wherein a specification for identifying the required document or the unnecessary document is provided. A word extraction means for extracting words for each document from a specified document group to which information is added; determining a part of speech of each extracted word; and, when words corresponding to a noun phrase are continuous, uniquely identify the word group. A compound word processing means for specifying a compound word, a document set in which individual words, discrimination information indicating whether or not the compound word is included in the document, and the designation information given to the document together with document identification information A set of document generating means for generating, from among the extracted words and the specified compound words, the information amount of a document group including and not including the words or compound words and the specified document group Individual words or compound words that maximize the mutual information represented by the difference from the total information amount are determined as search keywords, and one document set is divided into a plurality of document sets using the determined search keywords. A search formula creating apparatus, comprising: a document set dividing unit, wherein the search formula is created by combining search keywords determined every time the document set is divided by a logical formula.
【請求項5】 検索者にとって関心のある必要文書及び
関心のない不要文書を含む文書群の文書検索に用いる検
索式を作成する装置であって、 前記必要文書または不要文書を識別するための指定情報
が付与された指定文書群から文書毎に単語抽出を行う単
語抽出手段と、 抽出された単語を予め設定された複合語作成基準に基づ
いて所定順に結合し、これにより得られた単語群を一意
の複合語として特定する複合語処理手段と、 個々の単語及び前記複合語が文書中に含まれるか否かを
表す判別情報及び文書に付与された前記指定情報を文書
識別情報と共に集合させた文書集合を生成する文書集合
生成手段と、 前記抽出された単語及び前記特定された前記複合語のう
ち、これらの単語または複合語を含む文書群及び含まな
い文書群の情報量と前記指定文書群の総情報量との差分
で表される相互情報量を最大にする個々の単語または複
合語を検索キーワードとして決定するとともに、決定し
た検索キーワードを用いて一つの文書集合を複数の文書
集合に分割する文書集合分割手段とを備え、 前記文書集合の分割を繰り返す度に決定された検索キー
ワードを論理式で結合して前記検索式を作成することを
特徴とする検索式作成装置。
5. An apparatus for creating a retrieval formula for use in document retrieval of a group of documents including a required document of interest and an unnecessary document of no interest to a searcher, wherein a specification for identifying the required document or the unnecessary document is provided. Word extraction means for extracting words for each document from a specified document group to which information has been added; and combining the extracted words in a predetermined order based on a preset compound word creation criterion. Compound word processing means for specifying a unique compound word, identification information indicating whether each word and the compound word are included in the document, and the designation information given to the document are collected together with document identification information. A document set generation unit for generating a document set; and, among the extracted words and the identified compound words, information amounts of documents including and not including the words or the compound words, Individual words or compound words that maximize the mutual information represented by the difference from the total information amount of the specified document group are determined as search keywords, and one document set is divided into multiple documents using the determined search keywords. A search formula creating apparatus, comprising: a document set dividing unit that divides a document into sets, wherein the search keywords determined each time the document set is divided repeatedly are combined with a logical formula to create the search formula.
【請求項6】 前記文書集合分割手段は、所定の最小記
述長原理に基づいて前記相互情報量を最大とする単語ま
たは複合語を前記検索キーワードとして逐次決定するよ
うに構成されることを特徴とする請求項3記載の検索式
作成装置。
6. The document set dividing means is configured to sequentially determine a word or compound word that maximizes the mutual information amount as the search keyword based on a predetermined minimum description length principle. The retrieval formula creation device according to claim 3, wherein
JP10005129A 1998-01-13 1998-01-13 Method for preparing retrieval expression and device therefor Pending JPH11203309A (en)

Priority Applications (1)

Application Number Priority Date Filing Date Title
JP10005129A JPH11203309A (en) 1998-01-13 1998-01-13 Method for preparing retrieval expression and device therefor

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
JP10005129A JPH11203309A (en) 1998-01-13 1998-01-13 Method for preparing retrieval expression and device therefor

Publications (1)

Publication Number Publication Date
JPH11203309A true JPH11203309A (en) 1999-07-30

Family

ID=11602716

Family Applications (1)

Application Number Title Priority Date Filing Date
JP10005129A Pending JPH11203309A (en) 1998-01-13 1998-01-13 Method for preparing retrieval expression and device therefor

Country Status (1)

Country Link
JP (1) JPH11203309A (en)

Cited By (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
WO2002077867A1 (en) * 2001-03-27 2002-10-03 Mitsubishi Space Software Co., Ltd. Web monitoring system and method
WO2002078275A1 (en) * 2001-03-27 2002-10-03 Mitsubishi Space Software Co., Ltd. Electronic mail monitoring system and method

Cited By (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
WO2002077867A1 (en) * 2001-03-27 2002-10-03 Mitsubishi Space Software Co., Ltd. Web monitoring system and method
WO2002078275A1 (en) * 2001-03-27 2002-10-03 Mitsubishi Space Software Co., Ltd. Electronic mail monitoring system and method
JP2002290469A (en) * 2001-03-27 2002-10-04 Mitsubishi Space Software Kk Email audit system and method
JP2002288173A (en) * 2001-03-27 2002-10-04 Mitsubishi Space Software Kk Web audit system and method

Similar Documents

Publication Publication Date Title
KR100721406B1 (en) Product search system and method using category search logic
CN114238573A (en) Information pushing method and device based on text countermeasure sample
JP3438781B2 (en) Database dividing method, program storage device storing program, and recording medium
JP2000348041A (en) Document retrieval method and apparatus, and machine-readable recording medium recording program
JP3577972B2 (en) Similarity determination method, document search device, document classification device, storage medium storing document search program, and storage medium storing document classification program
CN108287901A (en) Method and apparatus for generating information
CN112417996A (en) Information processing method and device for industrial drawing, electronic equipment and storage medium
JP5915274B2 (en) Information search method, program, and information search apparatus
JP2000172722A (en) Automatic product information indexing method and system on online store
KR101355945B1 (en) On line context aware advertising apparatus and method
WO2008062822A1 (en) Text mining device, text mining method and text mining program
KR100835290B1 (en) Document classification system and document classification method
CN111737523B (en) Video tag, generation method of search content and server
JP2019128925A (en) Event presentation system and event presentation device
CN114117047A (en) Method and system for classifying illegal voice based on C4.5 algorithm
CN117972025B (en) Massive text retrieval matching method based on semantic analysis
KR20010006632A (en) Information Processing System
KR20220041337A (en) Graph generation system of updating a search word from thesaurus and extracting core documents and method thereof
JP5644087B2 (en) Component highlighting apparatus, program, and method
JP3598738B2 (en) Information extraction device, information retrieval method and information extraction method
JP3772401B2 (en) Document classification device
KR20220041336A (en) Graph generation system of recommending significant keywords and extracting core documents and method thereof
JP3314720B2 (en) String search device
JPH10320403A (en) Method and device for generating retrieval expression, and record medium
JPH1040253A (en) Method and apparatus for generating viewpoints of words in sentences