JPH01112331A - Automatic evaluation device for significance of key word - Google Patents
Automatic evaluation device for significance of key wordInfo
- Publication number
- JPH01112331A JPH01112331A JP62270014A JP27001487A JPH01112331A JP H01112331 A JPH01112331 A JP H01112331A JP 62270014 A JP62270014 A JP 62270014A JP 27001487 A JP27001487 A JP 27001487A JP H01112331 A JPH01112331 A JP H01112331A
- Authority
- JP
- Japan
- Prior art keywords
- word
- words
- dictionary
- keywords
- keyword
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Granted
Links
Landscapes
- Information Retrieval, Db Structures And Fs Structures Therefor (AREA)
Abstract
Description
【発明の詳細な説明】
〔産業上の利用分野〕
本発明は、キーワード重要度自動評価装置に係り、詳し
くは、新聞記事データベース等の検索のために、個々の
記事からキーワードを自動的に抽出し、かつ、それらの
キーワードのもとの記事中における統計的、構文的、意
味的な重要度を評価し、キーワードを統合的な重要度の
順に順位付けする装置に関する。[Detailed Description of the Invention] [Industrial Application Field] The present invention relates to an automatic keyword importance evaluation device, and more specifically, to a device for automatically extracting keywords from individual articles for searching newspaper article databases, etc. The present invention also relates to a device that evaluates the statistical, syntactic, and semantic importance of these keywords in articles, and ranks the keywords in order of integrated importance.
従来、新聞記事等からキーワ等を自動的に抽出する方式
としてはフリーターム方式と統制キーワード方式が知ら
れている。Conventionally, the free-term method and the controlled keyword method are known as methods for automatically extracting key words from newspaper articles and the like.
フリーターム方式では、まず対象新聞記事等の分かち書
きを、漢字、ひらがな等の字種の変わり目、あるいは、
「、」、「。」等の区切り記号に着目してキーワード抽
出を行い、さらに分かち書き用の辞書を用いて語を品詞
単位に分割する。次に、接頭語、接尾語を登録した辞書
との照合により、分かち書きされた語から接頭語、接尾
語を取り去り、さらに、複合語の分割を、最小単位の単
語を登録した語い辞書を利用して、例えば「情報検索」
を「情報」と「検索」のように分割する。In the free-term method, first, divide the target newspaper article, etc. into characters at the transition points of kanji, hiragana, etc., or
Keywords are extracted by focusing on delimiters such as ",", ".", etc., and words are further divided into parts of speech using a separation dictionary. Next, prefixes and suffixes are removed from the separated words by comparing them with a dictionary in which prefixes and suffixes are registered, and compound words are further divided using a word dictionary in which minimum unit words are registered. For example, "information search"
Divide into "information" and "search".
次に、数字の単位語を登録した単位語辞書、並びに「昨
日」、「傾向」、「いま」のような不要語あるいはスト
ップワードなどと称するひらがな列・漢字列から成る語
であって一般的でキーワードとはならない語を登録した
不要語辞書を作成しておき、これらの辞書と分かち書き
された語との照合を行い、数字の単位語、並びにストッ
プワードを取り除き、あわせて数字も取り除いて、残っ
た語の中で名詞をキーワードとする。Next, there is a unit word dictionary in which numerical unit words are registered, as well as common words consisting of hiragana and kanji strings called unnecessary words or stop words such as ``yesterday'', ``trend'', and ``now''. Create unnecessary word dictionaries in which words that are not keywords are registered, check these dictionaries with the separated words, remove numerical unit words and stop words, and also remove numbers. Use nouns as keywords among the remaining words.
統制キーワード方式は、上記フリーターム方式の処理に
おいてキーワードとされた語について、キーワードとす
る語を登録した辞書と照合を行いキーワードを選択する
方式である。The controlled keyword method is a method in which words that are used as keywords in the free-term method processing are checked against a dictionary in which words that are keywords are registered, and keywords are selected.
上記従来技術のフリーターム方式と統制キーワード方式
は、いずれもキーワード抽出だけのためのものであり、
キーワードの記事中における統計的、構文的、意味的な
重要度までも評価して出力するものではなかった。その
結果、新聞記事等に対してインデクサと呼ばれるキーワ
ード付けの専門家が付けるキーワードの数は通常5〜6
個であるのに対して、従来技術によると、20個以上も
のキーワードが付けられることになり、このため、新聞
記事データベース等をキーワード検索する際に多数の不
必要な記事がキーワード検索に適合して、精度が低く能
率が悪いとか、データベース中に不必要なキーワードの
ための記憶スペースを大量に確保しなければならないと
いう欠点を有していた。Both the free term method and controlled keyword method of the above-mentioned prior art are only for keyword extraction.
It did not even evaluate and output the statistical, syntactic, and semantic importance of keywords in articles. As a result, the number of keywords assigned to newspaper articles by keyword experts called indexers is usually 5 to 6.
However, according to the conventional technology, more than 20 keywords are assigned, and for this reason, when keyword searching a newspaper article database, etc., a large number of unnecessary articles match the keyword search. However, these methods have drawbacks such as low accuracy and inefficiency, and the need to reserve a large amount of storage space for unnecessary keywords in the database.
本発明の目的は、キーワード検索を高精度、高能率なも
のにするために、個々の新聞記事等からのキーワード抽
出において、該抽出されたキーワードの重要度を評価し
て重要なキーワードによる検索を可能ならしめるキーワ
ード重要度自動評価装置を提供することに有る。An object of the present invention is to perform a search using important keywords by evaluating the importance of the extracted keywords when extracting keywords from individual newspaper articles, etc., in order to make keyword searches highly accurate and efficient. An object of the present invention is to provide an automatic keyword importance evaluation device that makes it possible to evaluate the importance of keywords.
〔問題点を解決するための手段及び作用〕本発明のキー
ワード重要度自動評価装置は、入力処理部1名詞抽出部
、接辞・数詞削除部、不要語削除部、シソーラス・重要
語辞書照合部、並立語認定部、上中位語認定部、出現位
置認定部、出現頻度認定部1語重要度評価部及び接頭語
辞書、接尾語辞書、「昨日」、「傾向」などの−船釣な
語でキーワードにはならない語を登録した不要語辞書、
キーワードになり得る語を9.録し、さらにそれらの語
の相互関係として、同義語、上位語。[Means and operations for solving the problem] The automatic keyword importance evaluation device of the present invention includes an input processing unit 1, a noun extraction unit, an affix/numerical deletion unit, an unnecessary word deletion unit, a thesaurus/important word dictionary collation unit, Parallel word recognition section, Upper middle word recognition section, Occurrence position recognition section, Appearance frequency recognition section, Single word importance evaluation section, Prefix dictionary, Suffix dictionary, -Fune fishing words such as "yesterday" and "trend" An unnecessary word dictionary that registers words that cannot be used as keywords,
9. Words that can be keywords. In addition, synonyms and hypernyms are used to record the interrelationships between these words.
下位語、関連語といった語関係を示したシソーラス辞書
、特に重要な語であるとしてキーワードとしたい固有名
、地名等を9.録した重要語辞書などから構成される。9. A thesaurus dictionary that shows word relationships such as hyponyms and related words, as well as proper names, place names, etc. that are particularly important words that you want to use as keywords. It consists of a dictionary of important words that have been recorded.
入力処理部では、磁気記憶装置等に記録されている新聞
記事データベース等から記事を読み込み、名詞抽出部で
は、読み込まれた記事中から、「は」、「が」、「を」
等の助詞の直前の漢字カタカナ列を名詞として抽出し、
それらを抽出名詞テーブルに登録する。接辞・数詞削除
部では、抽出名詞テーブルの中の個々の語に対して接頭
語辞書、接尾語辞書と照合を行って個々の語の中の接頭
語、接尾語、助数詞を削除し、かつ個々の語の中の数詞
も削除し、抽出名詞テーブルを更新する。不要語削除部
では、抽出名詞テーブルの語に対して、不要語辞書と照
合を行って照合した不要語を削除し、抽出名詞テーブル
を更新する。The input processing unit reads articles from a newspaper article database recorded on a magnetic storage device, etc., and the noun extraction unit extracts "ha", "ga", and "wo" from the read articles.
Extract the kanji katakana sequence immediately before the particle such as as a noun,
Register them in the extracted noun table. The affix and numeral deletion section deletes prefixes, suffixes, and numerals from each word by comparing each word in the extracted noun table with a prefix dictionary and a suffix dictionary. The number word in the word is also deleted and the extracted noun table is updated. The unnecessary word deletion unit compares the words in the extracted noun table with an unnecessary word dictionary, deletes the checked unnecessary words, and updates the extracted noun table.
シソーラス・重要語辞書照合部では、更新された抽出名
詞テーブル中の語に対して、シソーラス及び重要語辞書
と照合を行って照合した語をキーワード候補としてキー
ワード候補テーブルに登録する。The thesaurus/key word dictionary collation unit collates the words in the updated extracted noun table with the thesaurus and key word dictionary and registers the collated words as keyword candidates in the keyword candidate table.
並立語認定部では、キーワード候補テーブルの語で、も
との記事中において「AやBJ、rAとB」、rA、B
Jのように並立に表現されている語を並立語として認定
し、上中位語認定部では、キーワード候補テーブルの語
について、シソーラスにおいて下位語が有る語を上中位
語として認定し、出現位置認定部では、キーワード候補
テーブルの語について、もとの記事中での出現位置が文
の最初から所定文字目まで\あるかを認定し、出現頻度
認定部では、キーワード候補テーブルの語について、も
との記事中で全部で何回出現しているかをカウントする
。In the parallel word recognition department, the words in the keyword candidate table are "A, BJ, rA and B", rA, B in the original article.
Words that are expressed in parallel, such as J, are recognized as parallel words, and the upper-middle word recognition department recognizes words in the keyword candidate table as upper-middle words that have hyponyms in the thesaurus. The position verification section verifies whether the words in the keyword candidate table appear in the original article from the beginning of the sentence to the predetermined character, and the appearance frequency verification section verifies the words in the keyword candidate table. Count the total number of times it appears in the original article.
これらの各認定部の認定結果を語特徴認定テーブルに登
録し、語重要度評価部では、語特徴認定テーブルの結果
に基づいて、上記の各認定部において認定された語に各
認定項目ごとに固有の評価点を与えて、その後、個々の
語について評価点を合計し総合計の順に語の重要度を決
める。The certification results of each of these certification departments are registered in the word feature certification table, and the word importance evaluation section uses the results of the word feature certification table to assign each certification item to the words certified by each certification department above. A unique evaluation score is given, and then the evaluation points are totaled for each word and the importance of the words is determined in order of the total sum.
以下、本発明の一実施例について図面により説明する。 An embodiment of the present invention will be described below with reference to the drawings.
第1図は本発明のキーワード重要度自動評価装置の一実
施例の基本構成図である。1はキーボード、電算写植等
の入力装置である。2は入力装置1によって読み込まれ
、磁気記憶装置等に文字コードの形式で記録されている
データベースで、こ\では新聞記事データベースとする
。3は新聞記事データベース2からの読み込みを行う入
力処理部である。FIG. 1 is a basic configuration diagram of an embodiment of the keyword importance automatic evaluation device of the present invention. Reference numeral 1 denotes an input device such as a keyboard or computer phototypesetting. Reference numeral 2 denotes a database read by the input device 1 and recorded in the form of character codes in a magnetic storage device, etc., which is herein referred to as a newspaper article database. 3 is an input processing unit that reads from the newspaper article database 2;
4は読み込まれた新聞記事中から、「は」、「が」、「
を」等の助詞の直前に位置する漢字カタカナ列を名詞と
して抽出する名詞抽出部である。4 is "ha", "ga", " from the loaded newspaper article.
This is a noun extraction unit that extracts the kanji-katakana sequence located immediately before a particle such as "wo" as a noun.
5は名詞抽出部4で抽出された名詞が9.録される抽出
名詞テーブルである。5 is the noun extracted by the noun extraction unit 4 is 9. This is the extracted noun table that is recorded.
6は抽出名詞テーブル5の中の個々の語に対して接頭語
辞書7、接尾語辞書8との照合を行って個々の中の接頭
語、接尾語、助数詞を削除し、かつ個々の語の中の数詞
も削除し、抽出名詞テーブル5を更新する接辞・数詞削
除部である。7,8はそれぞれ接頭語辞書(助数詞を含
む)、接尾語辞書(助数詞も含む)である。6 compares each word in the extracted noun table 5 with the prefix dictionary 7 and suffix dictionary 8, deletes the prefix, suffix, and particle in each word, and This is an affix and numeral deletion unit that also deletes the numerals inside and updates the extracted noun table 5. 7 and 8 are a prefix dictionary (including classifiers) and a suffix dictionary (also including classifiers), respectively.
9は更新された抽出名詞テーブル5の中の個々の語に対
して、不要語辞書10と照合を行って、照合した不要語
を削除し、抽出名詞テーブル5を更新する不要語削除部
である。10は「昨日」、「傾向」などの−船釣な語で
キーワードにはならないものを登録した不要語辞書であ
る。Reference numeral 9 denotes an unnecessary word deletion unit that compares each word in the updated extracted noun table 5 with the unnecessary word dictionary 10, deletes the checked unnecessary words, and updates the extracted noun table 5. . Reference numeral 10 is an unnecessary word dictionary in which words such as ``yesterday'' and ``trend'' that are commonplace and cannot be used as keywords are registered.
11は更新された抽出名詞テーブル5の中の個々の語に
対して、シソーラス辞書12並びに重要語辞書13と照
合を行うシソーラス・重要語照合部である。12はシソ
ーラス辞書で、これはキーワードになる得る語を登録し
、さらにそれらの語の相互関係として、同義語、上位語
、下位語、関連語といった語関係を示したものである。Reference numeral 11 denotes a thesaurus/key word matching unit that matches each word in the updated extracted noun table 5 with the thesaurus dictionary 12 and key word dictionary 13. 12 is a thesaurus dictionary which registers words that can be used as keywords and further shows word relationships such as synonyms, hypernyms, hyponyms, and related words as mutual relationships between these words.
13は特に重要な語であるとして、キーワードとしたい
固有名、地名等を登録した重要語辞書である。14はシ
ソーラス・重要語辞書照合部11で照合のとれた語がキ
ーワード候補語として登録されるキーワード候補テーブ
ルである。13 is an important word dictionary in which proper names, place names, etc. that are considered to be particularly important words are registered as keywords. 14 is a keyword candidate table in which words matched by the thesaurus/key word dictionary matching unit 11 are registered as keyword candidate words.
15はキーワード候補テーブル14中の語について、も
との新聞記事中に並立に表現されているか否かを認定す
る並立語認定部である。16はキーワード候補テーブル
14中の語について、シソーラス辞書12で下位語が有
る語を上中位語として認定する上中位語認定部である。Reference numeral 15 denotes a parallel word recognition unit that determines whether or not words in the keyword candidate table 14 are expressed concurrently in the original newspaper article. Reference numeral 16 denotes an upper-middle word recognition unit that certifies words in the keyword candidate table 14 that have lower-order words in the thesaurus dictionary 12 as upper-middle words.
17はキーワード候補テーブル14中の語について、も
との新聞記事中での出現位置が文の最初から所定文字目
まで\あるかを認定する出現位置認定部である。Reference numeral 17 denotes an appearance position determination unit that determines whether the word in the keyword candidate table 14 appears at a position in the original newspaper article from the beginning of the sentence to a predetermined character.
18はキーワード候補テーブル14中の語について、も
との新聞記事中で全部で何回出現しているかをカウント
する出現頻度認定部である。19は各認定部15〜18
で認定した結果が登録される諸特徴認定テーブルである
。Reference numeral 18 denotes an appearance frequency recognition unit that counts how many times a word in the keyword candidate table 14 appears in the original newspaper article. 19 is each certification department 15-18
This is a feature certification table in which the certification results are registered.
20は諸特徴認定テーブル19に基づいて、上記の各認
定部15〜18において認定された個々の語に対して各
認定項目ごとに固有の評価点を与え、その後、個々の語
について評価点を合計して、総合計の順に語の重要度を
決める語重要度評価部である。21は語重要度評価部2
0の結果を出力する印字装置、22は同じく語重要度評
価部20の結果を登録する結果ファイルである。20 gives a unique evaluation point for each certification item to each word certified in each of the above-mentioned certification sections 15 to 18 based on the various characteristic certification table 19, and then assigns an evaluation point to each word. This is a word importance evaluation unit that determines the importance of words in order of the total sum. 21 is word importance evaluation unit 2
A printing device 22 outputs a result of 0, and a result file 22 registers the results of the word importance evaluation section 20.
まず、キーワード抽出の対象となる新聞記事がキーボー
ド、電算写植等の入力装置1から読み込まれ、磁気記憶
装置等に記録されて新聞記事データベース2となる。こ
の新聞記事データベース2からキーワード抽出対象新聞
記事が入力処理部3によって入力される。名詞抽出部4
は、この処理対象新聞記事中から、「は」、「が」、「
を」等の助詞の直前に位置する漢字カタカナ列を名詞と
して抽出し、それらが抽出名詞テーブル5に登録される
。第2図(イ)に抽出名詞テーブル5に登録された抽出
名詞の内容の一部を示す。First, a newspaper article that is a target for keyword extraction is read from an input device 1 such as a keyboard or computer typesetting, and is recorded in a magnetic storage device or the like to form a newspaper article database 2. From this newspaper article database 2, the input processing section 3 inputs newspaper articles to be extracted for keywords. Noun extraction part 4
``ha'', ``ga'', ``from among the newspaper articles to be processed.
Kanji-katakana sequences located immediately before particles such as "wo" are extracted as nouns and are registered in the extracted noun table 5. FIG. 2(A) shows part of the contents of extracted nouns registered in the extracted noun table 5.
次に、接辞・数詞削除部6は、抽出名詞テーブル5に登
録されている語に対して接頭語辞#(助数詞も含む)7
、接尾語辞書(助数詞も含む)8と照合を行って個々の
語の中の接頭語、接尾語。Next, the affix/number deletion unit 6 deletes the prefix # (including the particle) 7 from the word registered in the extracted noun table 5.
, prefixes and suffixes in individual words by checking with a suffix dictionary (including suffixes) 8.
助数詞を削除し、かつ個々の語の中の数詞も削除し、抽
出名詞テーブル5を更新する。第2図(ロ)に、この接
辞・数詞が削除された抽出名詞テーブル5の一部を示す
。次に、不要語削除部9は、更新された抽出名詞テーブ
ル5の中の個々の語に対して、不要語辞書10と照合を
行って、照合のとれた「ts査」、「昨日」、[傾向J
なとの一般的な語でキーワードにはならい不要語を削除
し、抽出名詞テーブル5を更新する。第2図(ハ)に、
この不要語が削除された抽出名詞テーブル5の一部を示
す。The extracted noun table 5 is updated by deleting the particle and also deleting the numeral in each word. FIG. 2(b) shows a part of the extracted noun table 5 from which this affix/number has been deleted. Next, the unnecessary word deletion unit 9 compares each word in the updated extracted noun table 5 with the unnecessary word dictionary 10, and the words "ts-shu", "yesterday", [Trend J
The extracted noun table 5 is updated by deleting unnecessary words using the common word ``nato'' as a keyword. In Figure 2 (c),
A part of the extracted noun table 5 from which unnecessary words have been deleted is shown.
次に、シソーラス・重要語辞書照合部11は、更新され
た抽出名詞テーブル5の中の個々の語に対して、シソー
ラス辞書12及び重要語辞書13と照合を行って、照合
のとれた語をキーワード候補としてキーワード候補テー
ブル14に登録する。Next, the thesaurus/key word dictionary matching unit 11 matches each word in the updated extracted noun table 5 with the thesaurus dictionary 12 and key word dictionary 13, and selects matched words. It is registered in the keyword candidate table 14 as a keyword candidate.
第2図(ニ)に、このようにしてキーワード候補テーブ
ル14に9.録された語の一部を示す。As shown in FIG. 2(d), 9. is added to the keyword candidate table 14 in this way. Shows some of the words recorded.
次に、並立語認定部15はキーワード候補テーブル14
中の語について、それが新聞記事データベース2のもと
の新聞記事中で、「AやB」、「AとB」、rA、BJ
のA、Bのように並立に表現されているか否かを認定し
、その結果を諸特徴認定テーブル19に登録する1次に
上中位語認定部16はキーワード候補テーブル14中の
語について、シソーラスで下位語が有る語を上中位語と
して認定してその結果を諸特徴認定テーブル19に登録
する。次に、出現位置認定部17はキーワード候補テー
ブル14中の語について、もとの新聞記事中での出現位
置が文の最初から予め定めた文字位置までNであるかを
認定して、その結果を諸特徴認定テーブル19に登録す
る。なお、実験では文の最初から80〜90文字目程度
が最適で、それより小さくても、あるいは大きくてもあ
まり意味がないことが確められた。Next, the parallel word recognition unit 15 uses the keyword candidate table 14
Regarding the words in the middle, they are found in the newspaper articles in the newspaper article database 2, such as "A and B", "A and B", rA, BJ.
For words in the keyword candidate table 14, the upper middle word recognition unit 16 determines whether or not they are expressed in parallel, such as A and B, and registers the results in the feature recognition table 19. A word with a lower term in the thesaurus is recognized as an upper middle term, and the result is registered in a characteristic recognition table 19. Next, the appearance position recognition unit 17 determines whether the appearance position of the word in the keyword candidate table 14 in the original newspaper article is N from the beginning of the sentence to a predetermined character position, and the result is are registered in the characteristics recognition table 19. In addition, experiments have confirmed that the optimum value is around the 80th to 90th character from the beginning of a sentence, and that there is little meaning in smaller or larger characters.
次に、出現頻度認定部18はキーワード候補テーブル1
4中の語について、もとの新聞記事中で全部で何回出現
しているかをカウントしてその結果を諸特徴認定テーブ
ル19に登録する。Next, the appearance frequency recognition unit 18 uses the keyword candidate table 1
The total number of times the words in 4 appear in the original newspaper article is counted and the results are registered in the feature recognition table 19.
第3図は諸特徴認定テーブル19の内容例で。FIG. 3 shows an example of the contents of the various feature recognition table 19.
キーワード候補テーブル14中の各語に対する上記各認
定部15〜18での認定の有無を、有の場合は[0」、
無の場合は無印で示したものである。Whether each word in the keyword candidate table 14 has been certified by each of the certification units 15 to 18, if yes, [0];
If there is no item, it is shown with no mark.
次に、語重要度評価部20は諸特徴認定テーブル19に
基づいて、上記各認定部15〜18において認定された
個々の語に対して各認定項目ごとに固有の評価点を与え
、その後、個々の語について評価点を合計して、総合計
の順し二語の重要度を決め、印字装置21へ結果を出力
し、また磁気記憶装置などの結果ファイル22に登録す
る。第4図は語の重要度評価結果の一例を示したもので
、語が評価された重要度の順に並べられている。Next, the word importance evaluation section 20 gives a unique evaluation point for each certification item to each word certified by each of the above-mentioned certification sections 15 to 18 based on the various feature certification table 19, and then, The evaluation scores for each word are totaled, the importance of the two words is determined based on the total sum, and the result is output to the printing device 21 and registered in a result file 22 such as a magnetic storage device. FIG. 4 shows an example of the results of evaluating the importance of words, in which words are arranged in the order of their evaluated importance.
キーワードの重要度の総合的順位付けの精度は実験によ
って確認されていて、一般新聞紙から無作為に選んだ2
00記事を実験サンプルとして。The accuracy of the overall ranking of keyword importance has been confirmed through experiments.
00 articles as experimental samples.
この200記事中の必要なキーワードの95%までが、
各記事での重要度の上位10位の語群に中に含まれてい
る。従って、例えば本装置の出力結果の上位10個をキ
ーワードとすることにより、従来の技術では個々の新聞
記事に対して20個以上のキーワードが付けられていた
のに対して、入力新聞記事につけるキーワードの数を1
/2以下にでき、その結果、新聞記事データベースのキ
ーワードによる検索を高精度かつ高能率にし、またデー
タベース中のキーワードのための記憶容量も1/2以下
にできること\なる。Up to 95% of the necessary keywords in these 200 articles are
It is included in the top 10 words of importance in each article. Therefore, for example, by using the top 10 keywords in the output results of this device, keywords can be added to the input newspaper article, whereas in the conventional technology, 20 or more keywords are attached to each newspaper article. The number of keywords is 1
/2 or less, and as a result, keyword searches in newspaper article databases can be performed with high accuracy and efficiency, and the storage capacity for keywords in the database can also be reduced to 1/2 or less.
以上説明したように1本発明のキーワード重要度自動評
価装置は、従来の技術に加えて、並立語認定部、上中位
認定部、出現位置認定部、出現頻度認定部、語重要度評
価部などを備え、並立語認定部ではキーワード候補語に
ついて、並立に表現されているかどうかを認定し、上中
位語認定部ではキーワード候補語について、その語がシ
ソーラスにおいて上中位語であるかどうかを認定し、出
現位置認定部では、キーワード候補語について、もとの
新聞記事中での出現位置が文の最初から所定文字位置ま
で\あるかを認定し、出現頻度認定部では、キーワード
候補語について、もとの新聞記事中で全部で何回出現し
ているかをカウントし、語重要度評価部では、上記の各
認定部において認定された個々の語に対して各認定部ご
とに固有の評価点を与え、その後、個々の語について評
価点を合計して、総合計の順に語の重要度を精度良く決
めるものである。As explained above, in addition to the conventional technology, the keyword importance automatic evaluation device of the present invention includes a parallel word recognition section, an upper middle recognition section, an appearance position recognition section, an appearance frequency recognition section, and a word importance evaluation section. The parallel word certification section certifies whether keyword candidate words are expressed in parallel, and the upper middle word certification section certifies whether keyword candidate words are upper middle words in the thesaurus. The appearance position recognition section certifies whether the keyword candidate word appears in the original newspaper article from the beginning of the sentence to a predetermined character position.The appearance frequency recognition section certifies the keyword candidate word. The word importance evaluation department counts the total number of times a word appears in the original newspaper article, and then assigns a unique rating to each word recognized in each of the abovementioned departments. An evaluation point is given, and then the evaluation points are totaled for each word, and the importance of the words is determined with high precision in the order of the total sum.
このため、従来の技術では、個々の新聞記事等に対して
キーワードを抽出するだけで、しかも20個以上ものキ
ーワードが付けられていて、その中に不適切なキーワー
ドも多数含まれていて、これらのキーワードをキーワー
ド検索で使用すると多数の不適切な記事が抽出されるな
ど、検索の精度が低く、かつ非能率的であったのに対し
て1本装置はキーワードを抽出するだけでなく、抽出さ
れたキーワードを、もとの記事中での統計的、構文的、
意味的な総合的な重要度の順に出力することができるこ
とにより、例えば本装置の出力結果の上位10個をキー
ワードとすることにより、入力新聞記事等につけるキー
ワードの数を1/2以下にでき、その結果記事データベ
ースのキーワードによる検索を高精度かつ高能率にし、
またデータベース中のキーワードのための記憶容量も1
/2以下にできる利点が有る。For this reason, conventional technology only extracts keywords from individual newspaper articles, etc., but more than 20 keywords are attached, including many inappropriate keywords. When keywords were used in a keyword search, many inappropriate articles were extracted, resulting in low search accuracy and inefficiency.In contrast, this device not only extracts keywords, but also extracts a large number of inappropriate articles. keywords in the original article, statistically, syntactically,
By being able to output in order of overall semantic importance, for example, by using the top 10 output results of this device as keywords, it is possible to reduce the number of keywords attached to input newspaper articles, etc. by half or less. As a result, keyword searches in the article database are made highly accurate and efficient.
Also, the storage capacity for keywords in the database is 1
There is an advantage that it can be made less than /2.
第1図は本発明のキーワード重要度自動評価装置の一実
施例の基本構成図、第2図は第1図の抽出名詞テーブル
の内容の遷移及びキーワード候補テーブルの内容の一例
を示す図、第3図は第1図の諸特徴認定テーブルの内容
の一例を示す図、第4図はキーワード候補テーブル中の
語の重要度評価結果の一例を示す図である。
1・・・入力装置、 2・・・新聞記事データベース、
3・・・入力処理部、 4・・・名詞抽出部、5・・・
抽出名詞テーブル、
6・・・接辞・数詞削除部、 7・・・接頭語辞書。
8・・・接尾語辞書、 9・・・不要語削除部、10・
・・不要語辞書、
11・・・シソーラス・重要語辞書照合部。
12・・・シソーラス辞書、 13・・・重要語辞書
、14・・・キーワード候補テーブル、
15・・・並立語認定部、 16・・・上中位語認定
部、17・・・出現位置認定部。
18・・・出現頻度認定部、
19・・・諸特徴認定テーブル、
20・・・語重要度評価部、 21・・・印字装置。
22・・・結果ファイル。
第2
の/滲p −7°fしつ一改
P(ハ) (ニ)−’7
”lしの、舎pFIG. 1 is a basic configuration diagram of an embodiment of the automatic keyword importance evaluation device of the present invention, FIG. FIG. 3 is a diagram showing an example of the contents of the feature recognition table shown in FIG. 1, and FIG. 4 is a diagram showing an example of the results of evaluating the importance of words in the keyword candidate table. 1... Input device, 2... Newspaper article database,
3... Input processing unit, 4... Noun extraction unit, 5...
Extracted noun table, 6... Affix/number deletion unit, 7... Prefix dictionary. 8... Suffix dictionary, 9... Unnecessary word deletion section, 10.
...Unnecessary word dictionary, 11...Thesaurus/important word dictionary collation unit. 12... Thesaurus dictionary, 13... Important word dictionary, 14... Keyword candidate table, 15... Parallel word recognition section, 16... Upper middle word recognition section, 17... Occurrence position recognition Department. 18... Appearance frequency recognition section, 19... Feature recognition table, 20... Word importance evaluation section, 21... Printing device. 22...Result file. 2nd / 滲p -7°f Shitsuichi Kai P (c) (d) -'7
"l Shino, Shap
Claims (1)
し、それらのキーワードの記事中における統計的、構文
的、意味的な重要度を自動的に評価するキーワード重要
度自動評価装置において、記事データベース、抽出名詞
テーブル、キーワード候補テーブル、語特徴認定テーブ
ルと、接頭接尾語辞書、キーワードにならない一般的語
を登録した不要語辞書、同義語、上位語、下位語、関連
語等の語の相互関係を示すシソーラス辞書、特にキーワ
ードとしたい重要な語を登録した重要語辞書と、 前記記事データベースから記事を読み込む入力処理部と
、 前記読み込まれた記事中から名詞を抽出して前記抽出名
詞テーブルに登録する名詞抽出部と、前記抽出名詞テー
ブル中の個々の語に対して前記接頭接尾語辞書と照合を
行って、接頭語、接尾語、助数詞、数詞等を削除し、該
抽出名詞テーブルを更新する接辞・数詞削除部と、 前記抽出名詞テーブル中の語に対して、前記不要語辞書
と照合を行って照合した不要語を削除し、該抽出名詞テ
ーブルを更新する不要語削除部と、 前記更新された抽出名詞テーブル中の語に対して、前記
シソーラス辞書及び重要語辞書と照合を行って照合した
語をキーワード候補として前記キーワード候補テーブル
に登録するシソーラス・重要語辞書照合部と、 前記キーワード候補テーブルの各語について、前記記事
データベースのもとの記事を参照して並立語、上中位語
、出現位置、出現頻度等を認定して前記語特徴認定テー
ブルに登録する認定部と、 前記諸特徴認定テーブルの結果に基づいて、前記認定部
において認定された語に認定項目ごとの固有の評価点を
与え、その総合計の順に語の重要度を決める語重要度評
価部と、 を有することを特徴とするキーワード重要度自動評価装
置。(1) In an automatic keyword importance evaluation device that automatically extracts keywords from individual newspaper articles, etc., and automatically evaluates the statistical, syntactic, and semantic importance of those keywords in the article, Database, extraction noun table, keyword candidate table, word feature recognition table, prefix and suffix dictionary, unnecessary word dictionary with common words that cannot be keywords, synonyms, hypernyms, hyponyms, related words, etc. A thesaurus dictionary showing relationships, especially an important word dictionary that registers important words to be used as keywords, an input processing unit that reads articles from the article database, and extracts nouns from the read articles and stores them in the extracted noun table. The registered noun extraction unit compares each word in the extracted noun table with the prefix and suffix dictionary, deletes prefixes, suffixes, particles, number words, etc., and updates the extracted noun table. an affix/numerical deletion unit that performs a check on the words in the extracted noun table with the unnecessary word dictionary, deletes the checked unnecessary words, and updates the extracted noun table; a thesaurus/key word dictionary matching unit that matches the words in the updated extracted noun table with the thesaurus dictionary and the key word dictionary and registers the matched words as keyword candidates in the keyword candidate table; and the keywords. a recognition unit that refers to the original article in the article database to identify parallel words, upper middle words, appearance positions, appearance frequencies, etc. for each word in the candidate table, and registers the results in the word feature recognition table; a word importance evaluation unit that assigns unique evaluation points for each certification item to the words certified by the certification unit based on the results of the various characteristic certification tables, and determines the importance of the words in order of the total sum; An automatic keyword importance evaluation device characterized by:
Priority Applications (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP62270014A JPH0740275B2 (en) | 1987-10-26 | 1987-10-26 | Keyword automatic evaluation system |
Applications Claiming Priority (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP62270014A JPH0740275B2 (en) | 1987-10-26 | 1987-10-26 | Keyword automatic evaluation system |
Publications (2)
| Publication Number | Publication Date |
|---|---|
| JPH01112331A true JPH01112331A (en) | 1989-05-01 |
| JPH0740275B2 JPH0740275B2 (en) | 1995-05-01 |
Family
ID=17480345
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| JP62270014A Expired - Fee Related JPH0740275B2 (en) | 1987-10-26 | 1987-10-26 | Keyword automatic evaluation system |
Country Status (1)
| Country | Link |
|---|---|
| JP (1) | JPH0740275B2 (en) |
Cited By (15)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JPH03135669A (en) * | 1989-06-29 | 1991-06-10 | Tokyo Electric Power Co Inc:The | Automatic key word extracting system |
| JPH03244080A (en) * | 1990-02-22 | 1991-10-30 | Teremateiiku Kokusai Kenkyusho:Kk | Description integration processor |
| JPH04133173A (en) * | 1990-09-25 | 1992-05-07 | Teremateiiku Kokusai Kenkyusho:Kk | Information retrieving device |
| JPH04262460A (en) * | 1991-02-15 | 1992-09-17 | Ricoh Co Ltd | Information retrieval device |
| JPH05120345A (en) * | 1991-05-31 | 1993-05-18 | Teremateiiku Kokusai Kenkyusho:Kk | Keyword extracting device |
| JPH06251072A (en) * | 1993-02-27 | 1994-09-09 | Omron Corp | Device and method for processing document |
| JPH06314297A (en) * | 1993-04-30 | 1994-11-08 | Omron Corp | Device and method for processing of document and device and method for retrieving data base |
| JPH0778182A (en) * | 1993-06-18 | 1995-03-20 | Hitachi Ltd | Keyword assignment system |
| JPH0785101A (en) * | 1993-09-20 | 1995-03-31 | Fujitsu F I P Kk | Keyword extraction processor |
| JPH07114573A (en) * | 1993-10-18 | 1995-05-02 | Atr Tsushin Syst Kenkyusho:Kk | Picture retrieval device |
| JPH08340519A (en) * | 1995-06-13 | 1996-12-24 | Matsushita Electric Ind Co Ltd | Information extracting device and teletext receiving device with information extracting function |
| JPH09269951A (en) * | 1996-04-03 | 1997-10-14 | Matsushita Electric Ind Co Ltd | English summarization device |
| JP2003308324A (en) * | 2002-04-12 | 2003-10-31 | Yomiuri Shimbun | Search word processor, and device for retrieving document |
| JP2014191550A (en) * | 2013-03-27 | 2014-10-06 | Intelligent Wave Inc | Content search server, content search device, and content search method |
| JP2016122398A (en) * | 2014-12-25 | 2016-07-07 | 日本放送協会 | Subject word extraction device and program |
Families Citing this family (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| KR101254362B1 (en) * | 2007-05-18 | 2013-04-12 | 엔에이치엔(주) | Method and system for providing keyword ranking using common affix |
Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JPS57137965A (en) * | 1981-02-20 | 1982-08-25 | Nippon Kagaku Gijutsu Joho Center | Automatic key word extraction system of sentence consisting of chinese character and "kana"(japanese syllabary) |
| JPS57182279A (en) * | 1981-05-02 | 1982-11-10 | Canon Inc | Character processor |
| JPS5850071A (en) * | 1979-12-28 | 1983-03-24 | インタ−ナショナル ビジネス マシ−ンズ コ−ポレ−ション | Document excerpt memory |
| JPS608981A (en) * | 1983-06-28 | 1985-01-17 | Fujitsu Ltd | Semantics extracting device of natural language |
| JPS61262924A (en) * | 1985-05-17 | 1986-11-20 | Canon Inc | electronic file device |
-
1987
- 1987-10-26 JP JP62270014A patent/JPH0740275B2/en not_active Expired - Fee Related
Patent Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JPS5850071A (en) * | 1979-12-28 | 1983-03-24 | インタ−ナショナル ビジネス マシ−ンズ コ−ポレ−ション | Document excerpt memory |
| JPS57137965A (en) * | 1981-02-20 | 1982-08-25 | Nippon Kagaku Gijutsu Joho Center | Automatic key word extraction system of sentence consisting of chinese character and "kana"(japanese syllabary) |
| JPS57182279A (en) * | 1981-05-02 | 1982-11-10 | Canon Inc | Character processor |
| JPS608981A (en) * | 1983-06-28 | 1985-01-17 | Fujitsu Ltd | Semantics extracting device of natural language |
| JPS61262924A (en) * | 1985-05-17 | 1986-11-20 | Canon Inc | electronic file device |
Cited By (15)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JPH03135669A (en) * | 1989-06-29 | 1991-06-10 | Tokyo Electric Power Co Inc:The | Automatic key word extracting system |
| JPH03244080A (en) * | 1990-02-22 | 1991-10-30 | Teremateiiku Kokusai Kenkyusho:Kk | Description integration processor |
| JPH04133173A (en) * | 1990-09-25 | 1992-05-07 | Teremateiiku Kokusai Kenkyusho:Kk | Information retrieving device |
| JPH04262460A (en) * | 1991-02-15 | 1992-09-17 | Ricoh Co Ltd | Information retrieval device |
| JPH05120345A (en) * | 1991-05-31 | 1993-05-18 | Teremateiiku Kokusai Kenkyusho:Kk | Keyword extracting device |
| JPH06251072A (en) * | 1993-02-27 | 1994-09-09 | Omron Corp | Device and method for processing document |
| JPH06314297A (en) * | 1993-04-30 | 1994-11-08 | Omron Corp | Device and method for processing of document and device and method for retrieving data base |
| JPH0778182A (en) * | 1993-06-18 | 1995-03-20 | Hitachi Ltd | Keyword assignment system |
| JPH0785101A (en) * | 1993-09-20 | 1995-03-31 | Fujitsu F I P Kk | Keyword extraction processor |
| JPH07114573A (en) * | 1993-10-18 | 1995-05-02 | Atr Tsushin Syst Kenkyusho:Kk | Picture retrieval device |
| JPH08340519A (en) * | 1995-06-13 | 1996-12-24 | Matsushita Electric Ind Co Ltd | Information extracting device and teletext receiving device with information extracting function |
| JPH09269951A (en) * | 1996-04-03 | 1997-10-14 | Matsushita Electric Ind Co Ltd | English summarization device |
| JP2003308324A (en) * | 2002-04-12 | 2003-10-31 | Yomiuri Shimbun | Search word processor, and device for retrieving document |
| JP2014191550A (en) * | 2013-03-27 | 2014-10-06 | Intelligent Wave Inc | Content search server, content search device, and content search method |
| JP2016122398A (en) * | 2014-12-25 | 2016-07-07 | 日本放送協会 | Subject word extraction device and program |
Also Published As
| Publication number | Publication date |
|---|---|
| JPH0740275B2 (en) | 1995-05-01 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| JPH01112331A (en) | Automatic evaluation device for significance of key word | |
| Robertson et al. | Applications of n‐grams in textual information systems | |
| US5937422A (en) | Automatically generating a topic description for text and searching and sorting text by topic using the same | |
| JP2742115B2 (en) | Similar document search device | |
| JPH09259140A (en) | Information retrieval method and device therefor, and medium for storing information retrieval program | |
| JPS6330648B2 (en) | ||
| JP2001034623A (en) | Information retrieval method and information retrieval device | |
| Chen et al. | Named entity extraction for information retrieval | |
| JP2572314B2 (en) | Keyword extraction device | |
| JPH01217623A (en) | Automatic key word generating device | |
| JPH04205560A (en) | Information retrieval processing system | |
| JP3544749B2 (en) | Keyword automatic extraction device | |
| Zahoranský et al. | Text search of surnames in some slavic and other morphologically rich languages using rule based phonetic algorithms | |
| JPS63244259A (en) | Keyword extraction device | |
| Masuyama et al. | Automatic construction of Japanese KATAKANA variant list from large corpus | |
| JPH06208588A (en) | Document retrieving system | |
| KR20020054254A (en) | Analysis Method for Korean Morphology using AVL+Trie Structure | |
| JPH10149370A (en) | Document retrieval method and device using context information | |
| JPH04340164A (en) | Information retrieval processing system | |
| JPS63136224A (en) | Automatic key word extracting device | |
| JPH06325091A (en) | Similarity evaluation type database search device | |
| JP2004280323A (en) | Question document summarization device, question answer retrieval device, question document summarization program | |
| JPH04340165A (en) | Information retrieval processing system | |
| JPH0228769A (en) | Automatic key word generating device | |
| JPS63192130A (en) | Automatic key word extracting device |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| LAPS | Cancellation because of no payment of annual fees |