JPH0749872A - Automatic key word extraction system - Google Patents

Automatic key word extraction system

Info

Publication number
JPH0749872A
JPH0749872A JP4252434A JP25243492A JPH0749872A JP H0749872 A JPH0749872 A JP H0749872A JP 4252434 A JP4252434 A JP 4252434A JP 25243492 A JP25243492 A JP 25243492A JP H0749872 A JPH0749872 A JP H0749872A
Authority
JP
Japan
Prior art keywords
phrase
word
speech
dictionary
information
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Withdrawn
Application number
JP4252434A
Other languages
Japanese (ja)
Inventor
Noriko Otsuki
紀子 大槻
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
NEC Solution Innovators Ltd
Original Assignee
NEC Solution Innovators Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by NEC Solution Innovators Ltd filed Critical NEC Solution Innovators Ltd
Priority to JP4252434A priority Critical patent/JPH0749872A/en
Publication of JPH0749872A publication Critical patent/JPH0749872A/en
Withdrawn legal-status Critical Current

Links

Landscapes

  • Information Retrieval, Db Structures And Fs Structures Therefor (AREA)

Abstract

PURPOSE:To narrow down key-word extracted words by recognizing a phrase which is important when subjects are represented, term by term, and extracting a noun and an 's'-series inflection verb stem included therein as key words. CONSTITUTION:A word dividing means l divides a sentence into word units first by using a prepared dictionary 2 for word division. Then a part-d-speech recognizing means 3 recognizes the parts of speech of the respective words by utilizing a dictionary 4 for part-of-speech recognition. On the basis of the part-of-speech information recognized by the phrase recognizing means 3, nouns which succeed are recognized as one combined noun from the data divided into the word units and a particle and an auxiliary verb are recognized as one phrase by putting other words which are present right before together. Lastly, a key word extracting means 6 recognizes phrases which are used together with terms, term by term, and important when, specially, the subject is represented, in order by utilizing the obtained phrase information and a term dictionary 7, and the stems of the noun and 's'-series inflection verb in the phrases as stems that are suitable as key words.

Description

【発明の詳細な説明】Detailed Description of the Invention

【0001】[0001]

【産業上の利用分野】本発明はキーワード自動抽出方式
に関し、特にコンピュータによる情報検索システムで、
日本語の自然語文の文書テキストから、その文書テキス
トを検索する場合に有効なキーワードを抽出するキーワ
ード自動抽出方式に関する。
BACKGROUND OF THE INVENTION 1. Field of the Invention The present invention relates to an automatic keyword extraction method, and more particularly to an information retrieval system using a computer,
The present invention relates to an automatic keyword extraction method for extracting effective keywords from a document text of a Japanese natural language sentence when searching the document text.

【0002】[0002]

【従来の技術】従来のキーワード自動抽出方式は、文書
テキスト中の日本語の自然語文を、まず、単語分割用辞
書に基づいて単語分割を行った後、キーワードの候補と
して名詞とサ変動詞の語幹とを抽出し、あらかじめ用意
した不要語辞書の情報に基づいて、候補として抽出した
名詞およびサ変動詞語幹の中から不要語を除去したもの
をキーワードとしていた。なお、不要語辞書とは、利用
者が使用経験に基づいてキーワードとして不適当と思わ
れる単語を登録蓄積するものである。
2. Description of the Related Art In a conventional automatic keyword extraction method, a Japanese natural language sentence in a document text is first word-divided based on a word-division dictionary, and then the stems of a noun and a verb are used as keyword candidates. Based on the information in the unnecessary word dictionary prepared in advance, and were extracted as candidates, the unnecessary words were removed from the nouns and syllables stems extracted as candidates. It should be noted that the unnecessary word dictionary is used to register and accumulate words that are considered inappropriate as keywords by the user based on experience.

【0003】[0003]

【発明が解決しようとする課題】上述した従来のキーワ
ード自動抽出方式では、文書テキスト中の日本語の自然
語文の中から、文の主題とは関係なく、すべての名詞お
よびサ変動詞語幹を候補として抽出してしまうため、不
要語辞書を用意して不要語を蓄積する不便がある上に、
それでもなおキーワードとしては適当でない語を数多く
含むという欠点があった。
According to the above-described conventional automatic keyword extraction method, all nouns and sa variation stems are selected as candidates from the Japanese natural language sentences in the document text regardless of the subject of the sentence. Since it will be extracted, there is the inconvenience of preparing an unnecessary word dictionary and accumulating unnecessary words.
Nevertheless, there was a drawback that it contained many words that were not suitable as keywords.

【0004】本発明の目的は、上述した欠点を除去し、
文の主題と関係のある語のみを候補として抽出できるキ
ーワード自動抽出方式を提供することにある。
The object of the present invention is to eliminate the above-mentioned drawbacks,
It is to provide an automatic keyword extraction method that can extract only words related to the subject of a sentence as candidates.

【0005】[0005]

【課題を解決するための手段】本発明のキーワード自動
抽出方式は、コンピュータを用いた情報検索システムに
より、日本語の自然語文の文書テキストを単語分割用辞
書を参照して単語単位に分割した後に、その文書テキス
トを検索する場合に有効なキーワードを抽出するキーワ
ード自動抽出方式において、各単語の品詞情報が格納さ
れている品詞認識用辞書と、単語単位に分割されたデー
タに前記品詞認識用辞書を参照して品詞情報を付加する
品詞認識手段と、前記品詞認識手段により付加された品
詞情報を参照して連続した複数の名詞単語をまとめて1
語とした後に助詞,助動詞を直前の単語と結合して文節
単位にまとめる文節認識手段と、用言と共に用いられて
文を構成する助詞とその用言との関係情報を格納した用
言辞書と、前記文節認識手段により文節にまとめられた
データから用言を含む文節に注目して前記用言辞書を参
照しキーワードを抽出するキーワード抽出手段とを備え
て構成されている。
According to an automatic keyword extraction method of the present invention, an information retrieval system using a computer divides a document text of a Japanese natural language sentence into word units by referring to a word division dictionary. In the keyword automatic extraction method for extracting effective keywords when searching the document text, a part-of-speech recognition dictionary storing part-of-speech information of each word and the part-of-speech recognition dictionary for data divided into words Referring to the part-of-speech recognition means for adding the part-of-speech information, and referring to the part-of-speech information added by the part-of-speech recognition means, a plurality of consecutive noun words are collectively referred to as 1
A phrase recognition means for combining particles and auxiliary verbs with words immediately preceding them into words and combining them into phrases, and a phrase dictionary that is used together with a phrase to store the relational information between the particles forming the sentence and the phrases. , And a keyword extracting means for extracting a keyword by referring to the phrase dictionary and focusing on a phrase including a phrase from the data collected in the phrase by the phrase recognition means.

【0006】[0006]

【実施例】次に、本発明の実施例について図面を参照し
て説明する。
Embodiments of the present invention will now be described with reference to the drawings.

【0007】図1は本発明の一実施例の構成を示すブロ
ック図である。
FIG. 1 is a block diagram showing the configuration of an embodiment of the present invention.

【0008】本実施例のキーワード自動抽出方式は、図
1に示すように、対象となる日本語の自然語文の文書テ
キストを単語単位に分割する単語分割手段1と、単語分
割を行うために必要な情報が格納されている単語分割用
辞書2と、単語単位に分割されたデータに対して品詞情
報を認識し付加する品詞認識手段3と、品詞認識に必要
な情報が格納されている品詞認識用辞書4と、品詞情報
が付加された単語単位のデータを文節ごとにまとめる文
節認識手段5と、文節にまとめられたデータから用言に
注目してキーワードを抽出するキーワード抽出手段6
と、キーワード抽出に必要な用言に関する情報が格納さ
れている用言辞書7とを備えている。
As shown in FIG. 1, the automatic keyword extraction method of the present embodiment is necessary for word division means 1 for dividing the document text of the target Japanese natural language sentence into word units, and for performing word division. Dictionary for storing word-division information 2, part-of-speech recognition means 3 for recognizing and adding part-of-speech information to data divided into words, and part-of-speech recognition for storing information necessary for part-of-speech recognition Dictionary 4, phrase recognition means 5 for collecting word-by-word data to which part-of-speech information is added for each phrase, and keyword extraction means 6 for extracting keywords from the data collected in the phrase by paying attention to the phrase.
And a vocabulary dictionary 7 in which information about vocabulary necessary for keyword extraction is stored.

【0009】図2に、本実施例の動作を説明するための
日本語の自然語文の例文を示す。以下、図2の例文を用
いて本実施例の動作について説明する。
FIG. 2 shows an example sentence of a Japanese natural language sentence for explaining the operation of this embodiment. The operation of this embodiment will be described below with reference to the example sentence of FIG.

【0010】まず、単語分割手段1で、従来と同一の方
法による単語分割処理を行う。すなわち、指定された文
書テキスト中の日本語の自然語文を入力データとして、
あらかじめ用意されている単語分割用辞書2を利用して
単語単位に分割する。単語分割手段1により分割された
結果を図3に示す。図3において、“△”は各単語の区
切りを表している。
First, the word dividing means 1 performs the word dividing process by the same method as the conventional method. That is, the Japanese natural language sentence in the specified document text is used as input data,
It is divided into word units by using the word division dictionary 2 prepared in advance. The result of division by the word dividing means 1 is shown in FIG. In FIG. 3, “Δ” represents the division of each word.

【0011】次に、品詞認識手段3では、品詞認識用辞
書4を利用して各単語の品詞を認識する。ここで品詞と
は、通常の日本語文法で利用する動詞,サ変動詞,形容
詞,形容動詞,名詞,副詞,連体詞,接続詞,感動詞,
助動詞,助詞の11品詞である。なお、サ変動詞は動詞
の一部であるが、名詞が動詞化されたものでありキーワ
ードとなる単語を含むため、他の動詞とは別扱いとす
る。又、以降の説明で使用する「用言」とは、動詞,サ
変動詞,形容詞,形容動詞の4品詞をまとめたものであ
る。
Next, the part-of-speech recognition means 3 uses the part-of-speech recognition dictionary 4 to recognize the part-of-speech of each word. Here, the part-of-speech is a verb used in normal Japanese grammar, a sa verb, an adjective, an adjective, a noun, an adverb, a conjunction, a conjunction, a verb,
These are the 11 parts of speech of auxiliary verbs and particles. It should be noted that although the sa verb is a part of the verb, it is treated separately from other verbs because it is a verbized noun and includes words that are keywords. In addition, the “defective” used in the following description is a collection of four parts of speech: verb, sa verb, adjective, and adjective verb.

【0012】品詞認識用辞書4には、各単語の表記と品
詞の2情報のほかに、動詞,サ変動詞,形容詞,形容動
詞,助動詞の5品詞のような語尾変化する単語について
は、基本的な表現である基本形と、表記が変化しない部
分である語幹との2情報が格納されている。図2の例文
に必要な品詞認識用辞書4の情報を図4に示す。
The part-of-speech recognition dictionary 4 basically includes not only the notation of each word and two pieces of information on the part-of-speech, but also about words that change their endings such as five parts-parts of verbs, sa verbs, adjectives, adjectives and auxiliary verbs. Two types of information are stored: a basic form, which is an expression, and a stem, which is a part where the notation does not change. Information of the part-of-speech recognition dictionary 4 necessary for the example sentence of FIG. 2 is shown in FIG.

【0013】例を用いて説明すると、図3中の単語「同
定」(参照番号31)については、図4中の「同定」
(参照番号41)の情報から品詞が「名詞」であること
を認識する。同様に、図3中の単語「比較し」(参照番
号32)については、図4中の「比較し」(参照番号4
2)の情報から品詞が「サ変動詞」であり、その基本形
が「比較する」であること、及び語幹が「比較」である
ことを認識する。
Explaining with an example, the word "identification" (reference numeral 31) in FIG. 3 is "identification" in FIG.
It is recognized from the information (reference numeral 41) that the part of speech is a "noun". Similarly, for the word “compare” (reference numeral 32) in FIG. 3, “compare” (reference numeral 4) in FIG.
It is recognized from the information in 2) that the part-of-speech is "sa verb," its basic form is "compare," and the stem is "comparison."

【0014】図2の例文に対し、単語分割手段1及び品
詞認識手段3の処理を行った結果を図5に示す。図5に
おいて、“△”は各単語の区切りを表し、参照番号Aで
示した欄は認識した品詞情報の格納場所である。
FIG. 5 shows the result of processing the word segmentation means 1 and the part-of-speech recognition means 3 on the example sentence of FIG. In FIG. 5, “Δ” indicates the division of each word, and the column indicated by the reference number A is the storage location of the recognized part-of-speech information.

【0015】次に、文節認識手段5では、品詞認識手段
3で認識した品詞情報と、以下に示す二つのルールとに
基づき、単語単位に分割されたデータから文節を認識す
る。文節認識手段5で利用するルールは以下のとおりで
ある。 (1)名詞が連続している場合、連続している名詞を合
わせて1名詞とする。 (2)助詞,助動詞は、直前に存在する他の単語とまと
めて1文節とする。
Next, the phrase recognition unit 5 recognizes a phrase from the data divided into words based on the part-of-speech information recognized by the part-of-speech recognition unit 3 and the following two rules. The rules used by the phrase recognition means 5 are as follows. (1) When nouns are consecutive, the consecutive nouns are combined into one noun. (2) Particles and auxiliary verbs are grouped together with other words that exist immediately before into one phrase.

【0016】例を用いて説明すると、図5中の単語「同
定」(参照番号51)及び「結果」(参照番号52)
は、名詞が連続しているので、前述のルール(1)によ
り「同定結果」という1名詞と認識する。又、図5中の
単語「と」(参照番号53)は助詞であるため、前述の
ルール(2)により、直前に存在する名詞「同定結果」
とまとめて「同定結果△と」という1文節と認識する。
同様にして、図5中の単語「て」(参照番号55)も助
詞であるため、前述のルール(2)により、直前に存在
するサ変動詞「比較し」(参照番号54)とまとめて
「比較し△て」という1文節と認識する。
Explaining with an example, the words "identification" (reference numeral 51) and "result" (reference numeral 52) in FIG.
Since the noun is continuous, it is recognized as one noun “identification result” according to the above-mentioned rule (1). Further, since the word “to” (reference numeral 53) in FIG. 5 is a postpositional particle, the noun “identification result” immediately preceding the noun existing according to the above rule (2)
Recognized as one clause "identification result △ and" collectively.
Similarly, since the word “te” (reference numeral 55) in FIG. 5 is also a particle, it is collectively referred to as the immediately preceding sa verb “comparison” (reference numeral 54) according to the above-mentioned rule (2). Recognize it as one phrase, "Compare and △".

【0017】図2の文例に対して単語分割手段1,品詞
認識手段3,文節認識手段5の処理を行った結果を図6
に示す。図6において、“△”は各単語の区切りを、
“★”は各文節の区切りを表し、参照番号Aで示した欄
は認識した品詞情報の格納場所である。
FIG. 6 shows the result obtained by processing the word segmentation means 1, the part-of-speech recognition means 3, and the clause recognition means 5 on the example sentence of FIG.
Shown in. In FIG. 6, “Δ” indicates the division of each word,
"★" represents a delimiter of each clause, and the column indicated by reference number A is the storage location of the recognized part-of-speech information.

【0018】最後に、キーワード抽出手段6について説
明する。キーワード抽出手段6は、文節認識手段5によ
って得られた文節情報と用言辞書7とを利用して、各用
言ごとに用言と共に用いられて文を構成するために必要
な文節と、それらの文節の中で特に主題を表す場合に重
要となる文節を順次認識し、主題を表す場合に重要とな
る文節中の名詞,サ変動詞の語幹をキーワードとしてふ
さわしいものとして抽出する。なお、この処理に利用す
る用言辞書7には、各用言ごとに、用言と共に用いられ
て文を構成するために必要な文節中に含まれる助詞の情
報と、特に各用言が主題を構成する場合に重要と判断さ
れる文節中に含まれる助詞の情報とが格納されている。
図2の例文に関する用言辞書7の情報を図7に示す。
Finally, the keyword extracting means 6 will be described. The keyword extraction unit 6 uses the phrase information obtained by the phrase recognition unit 5 and the vocabulary dictionary 7, and the vocabulary necessary for constructing a sentence by using the verb with each vocabulary and those phrases. Among the bunsetsus, the bunsetsus that are particularly important when expressing the subject are sequentially recognized, and the stems of the nouns and sa verbs that are important when expressing the subject are extracted as appropriate keywords. It should be noted that the denotation dictionary 7 used for this processing contains information about particles, which is included in a clause necessary for constructing a sentence together with a denotation, and in particular, each subject is a subject. The information of the particle included in the clause which is judged to be important when constructing is stored.
Information of the vocabulary dictionary 7 regarding the example sentence of FIG. 2 is shown in FIG.

【0019】例を用いて説明すると、図6中の用言(サ
変動詞)を含む文節「比較し△て」(参照番号65)中
の単語「比較し」は、図4の「比較し」(参照番号4
2)の情報により基本形が「比較する」であることが得
られるので、基本形である「比較する」を用いて図7の
「比較する」(参照番号71)の情報を参照する。つま
り、「比較する」については、助詞が「が」「を」
「と」のいずれかを含む文節が文を構成するためには必
要であり、その中の助詞「が」「を」については、主題
を表す場合に重要となる文節中に含まれるという情報で
ある。従って、「比較し△て」から文の先頭方向に向か
って助詞が「が」「を」「と」のいずれかを含む文節を
捜していき、「比較する」の用言が文を構成するために
必要な文節は、図6中の「同定結果△と」(参照番号6
2)と、「話題同定結果△を」(参照番号61)の文節
であると判断する。図7の「比較する」(参照番号7
1)の情報から、助詞「と」を含む文節は主題を表す場
合に重要とならないことが分かるので、図6中の助詞
「と」を含む文節「同定結果△と」(参照番号62)中
にある名詞「同定結果」はキーワードとして抽出しない
と判断する。逆に、助詞「を」を含む文節は主題を表す
場合に重要であるので、助詞「を」を含む文節「話題同
定結果△を」(参照番号64)中にある名詞「話題同定
結果」はキーワードとして抽出すると判断する。
Explaining with an example, the word "comparing" in the phrase "comparing Δ te" (reference numeral 65) containing the adjective (sa verb) in FIG. 6 is "comparing" in FIG. (Reference number 4
Since the basic form is "comparing" from the information of 2), the information of "compare" (reference numeral 71) in FIG. 7 is referred to by using the basic form "compare". In other words, for "comparing", the particles are "ga" and "wo".
A phrase containing either "to" is necessary to compose a sentence, and the particles "ga" and "wo" in that phrase are included in the phrase that is important when expressing the subject. is there. Therefore, from “comparison △ te” toward the beginning of the sentence, we search for a phrase in which the particle contains either “ga”, “wo” or “to”, and the phrase “comparison” composes the sentence. The phrase necessary for this is "identification result △ and" in FIG.
2), it is determined that the phrase is “topic identification result Δ” (reference number 61). “Compare” in FIG. 7 (reference numeral 7
From the information in 1), it can be seen that the phrase containing the particle "to" is not important when expressing the subject. Therefore, in the phrase "identification result Δto" (reference numeral 62) containing the particle "to" in FIG. It is determined that the noun "identification result" in is not extracted as a keyword. On the contrary, since the phrase containing the particle "o" is important when representing the subject, the noun "topic identification result" in the phrase "topic identification result △ wo" (reference number 64) containing the particle "o" is It is determined to be extracted as a keyword.

【0020】同様に図6中の用言(動詞)を含む文節
「行っ△た」(参照番号68)中の単語「行っ」につい
ても判断する。まず、図4の「行っ」(参照番号43)
の情報から基本形が「行う」であることが得られるの
で、次に基本形である「行う」を用いて図7の「行う」
(参照番号72)の情報を参照する。つまり、「行う」
については助詞が「が」「を」のいずれかである文節が
文を構成するためには必要であり、助詞が「が」「を」
のいずれかである文節が主題を表す場合に重要となると
いう情報である。従って、「行っ△た」から文の先頭方
向に向かって助詞が「が」「を」である文節を捜してい
き、「行う」の用言が文を構成するために必要な文節
は、図6中の「評価△を」(参照番号67)の文節であ
ると判断する。このとき、図6中の「話題同定結果△
を」(参照番号64)も助詞「を」を含む文節である
が、先に「比較する」の用言で文を構成するために必要
な文節と判断されているので、重複して用言「行う」が
必要とする文節とは判断しない。前述したとおり、図7
の「行う」(参照番号72)の情報から、助詞「を」を
含む文節は主題を表す場合に重要な文節であるので、
「行う」の用言が文を構成するために必要な文節である
図6中の「評価△を」(参照番号67)中の名詞「評
価」をキーワードとして抽出するものと判断する。
Similarly, the word “go” in the phrase “go Δta” (reference numeral 68) including the adjective (verb) in FIG. 6 is also judged. First, “Go” in FIG. 4 (reference numeral 43)
Since the basic form "do" can be obtained from the information of "do", the basic form "do" is used to "do" in FIG.
Refer to the information of (reference number 72). In other words, "do"
For, a clause whose particle is either "ga" or "wo" is necessary to compose a sentence, and the particle is "ga" or "wo".
Information that is important when a clause that is one of Therefore, the phrase necessary for composing a sentence is searched for the phrase in which the particle is “ga” “wo” from the direction of “go △” toward the beginning of the sentence. It is judged that it is a clause of “evaluation Δ” in 6 (reference numeral 67). At this time, the “topic identification result Δ in FIG.
”” (Reference numeral 64) is also a phrase that includes the particle “”, but since it was previously determined to be a phrase necessary to compose a sentence by the “comparison” phrase, it is a duplicate phrase. Do not judge that the phrase is necessary for "do". As described above, FIG.
From the information of “do” (reference number 72) of, the phrase including the particle “o” is an important phrase when expressing the subject,
It is determined that the noun “evaluation” in “evaluation Δ” in FIG. 6 (reference numeral 67), which is a clause necessary to compose a sentence, is extracted as a keyword.

【0021】なお、図6の中の助詞「による」を含む文
節である「話題同定結果△による」(参照番号61),
「人手△による」(参照番号63)及び助詞「の」を含
む文節である「話題同定モデル△の」(参照番号66)
にいては、「比較する」「行う」のいずれの用言に対し
ても図7の「比較する」(参照番号72)と「行う」
(参照番号72)との情報より、文を構成する場合に必
要な文節とは判断されない。従って、これら3文節、す
なわち図6中の「話題同定結果△による」(参照番号6
1),「人手△による」(参照番号63),「話題同定
モデル△の」(参照番号66)からはキーワードを抽出
しない。
Note that the phrase "by topic identification result Δ" (reference numeral 61), which is a clause including the particle "by" in FIG.
“Topic identification model Δno” (reference number 66), which is a clause including “by hand Δ” (reference number 63) and a particle “no”
In this case, “compare” (reference numeral 72) and “perform” in FIG. 7 are applied to any of the terms “compare” and “perform”.
Based on the information (reference numeral 72), it is not determined that the clause is necessary when composing a sentence. Therefore, these three clauses, that is, “according to the topic identification result Δ” in FIG. 6 (reference numeral 6
No keywords are extracted from 1), “by human Δ” (reference number 63), and “of topic identification model Δ” (reference number 66).

【0022】図2の例文に対し、単語分割手段1,品詞
認識手段3,文節認識手段5,キーワード抽出手段6の
各処理を行った結果を、従来の方法による抽出結果と併
せて図8に示す。図8において、“△”は各単語の区切
りを、“★”は各文節の区切りを示し、参照番号Aの欄
には認識した品詞情報を、参照番号Bの欄には本実施例
によるキーワードとして抽出するか否かの判断を、参照
番号Cの欄には従来方式による判断を示してある。図8
を参照すると、本発明の方式と従来の方式の間には、
「同定話題モデル」(参照番号81),「同定結果」
(参照番号82),「人手」(参照番号83),「話題
同定モデル」(参照番号84)に関して相違が認めら
れ、本発明の方式では不必要なキーワード候補が多数抽
出されないことが分かる。
FIG. 8 shows the results obtained by performing the respective processes of the word dividing means 1, the part-of-speech recognition means 3, the clause recognition means 5, and the keyword extraction means 6 on the example sentence of FIG. 2 together with the extraction results by the conventional method. Show. In FIG. 8, “Δ” indicates the division of each word, “★” indicates the division of each clause, the column of reference number A shows the recognized part-of-speech information, and the column of reference number B shows the keyword according to this embodiment. The determination by the conventional method is shown in the column of reference numeral C. Figure 8
Referring to, between the method of the present invention and the conventional method,
"Identification topic model" (reference number 81), "identification result"
(Reference number 82), “manpower” (reference number 83), and “topic identification model” (reference number 84), differences are recognized, and it can be seen that a large number of unnecessary keyword candidates are not extracted by the method of the present invention.

【0023】[0023]

【発明の効果】以上説明したように、本発明のキーワー
ド自動抽出方式は、コンピュータによる情報検索システ
ムにより日本語の自然語文の文書テキストからキーワー
ドを抽出する際、文節単位にまとめた後、用言に注目
し、用言が自然語文の中で必要とする文節と、用言が主
題を表す場合に重要となる文節を、文の意味解析を行わ
ずに表記方法から判断し、各用言ごとに主題を表す際に
重要となる文節を認識し、その中に含まれる名詞,サ変
動詞語幹をキーワードとして抽出することにより、キー
ワードとして抽出する語の絞り込みを行える効果があ
る。
As described above, according to the automatic keyword extraction method of the present invention, when a keyword is extracted from a document text of a Japanese natural language sentence by an information retrieval system using a computer, the keywords are grouped into phrases and then Pay attention to the phrase, which phrase is needed in the natural language sentence, and the phrase that is important when the phrase is representative of the subject, are judged from the notation method without performing the semantic analysis of the sentence, and By recognizing the clauses that are important when expressing the subject and extracting the nouns and sa verbs included in them as keywords, there is an effect that the words extracted as keywords can be narrowed down.

【図面の簡単な説明】[Brief description of drawings]

【図1】本発明の一実施例の構成を示すブロック図であ
る。
FIG. 1 is a block diagram showing the configuration of an embodiment of the present invention.

【図2】本実施例を説明するための文書テキストの例文
を示した説明図である。
FIG. 2 is an explanatory diagram showing an example sentence of a document text for explaining the present embodiment.

【図3】図2の例文に対して単語分割処理を行った結果
の説明図である。
FIG. 3 is an explanatory diagram of a result of performing word division processing on the example sentence of FIG.

【図4】図2の例文に対する品詞認識用辞書の情報を示
した説明図である。
FIG. 4 is an explanatory diagram showing information in a part-of-speech recognition dictionary for the example sentence of FIG.

【図5】図2の例文に対する品詞認識処理結果を示した
説明図である。
5 is an explanatory diagram showing a part-of-speech recognition processing result for the example sentence of FIG. 2. FIG.

【図6】図2の例文に対する文節認識処理結果を示した
説明図である。
FIG. 6 is an explanatory diagram showing a phrase recognition processing result for the example sentence of FIG. 2;

【図7】図2の例文にに対する用言辞書の情報を示した
説明図である。
FIG. 7 is an explanatory diagram showing information of a vocabulary dictionary for the example sentence of FIG. 2;

【図8】図2の例文に対するキーワード抽出結果を従来
と比較した説明図である。
FIG. 8 is an explanatory diagram comparing a keyword extraction result for the example sentence of FIG. 2 with a related art.

【符号の説明】[Explanation of symbols]

1 単語分割手段 2 単語分割用辞書 3 品詞認識手段 4 品詞認識用辞書 5 文節認識手段 6 キーワード抽出手段 7 用言辞書 1 word dividing means 2 word dividing dictionary 3 part-of-speech recognition means 4 part-of-speech recognition dictionary 5 phrase recognition means 6 keyword extraction means 7 vocabulary dictionary

Claims (1)

【特許請求の範囲】[Claims] 【請求項1】 コンピュータを用いた情報検索システム
により、日本語の自然語文の文書テキストを単語分割用
辞書を参照して単語単位に分割した後に、その文書テキ
ストを検索する場合に有効なキーワードを抽出するキー
ワード自動抽出方式において、各単語の品詞情報が格納
されている品詞認識用辞書と、単語単位に分割されたデ
ータに前記品詞認識用辞書を参照して品詞情報を付加す
る品詞認識手段と、前記品詞認識手段により付加された
品詞情報を参照して連続した複数の名詞単語をまとめて
1語とした後に助詞,助動詞を直前の単語と結合して文
節単位にまとめる文節認識手段と、用言と共に用いられ
て文を構成する助詞とその用言との関係情報を格納した
用言辞書と、前記文節認識手段により文節にまとめられ
たデータから用言を含む文節に注目して前記用言辞書を
参照しキーワードを抽出するキーワード抽出手段とを備
えたことを特徴とするキーワード自動抽出方式。
1. An information retrieval system using a computer divides a document text of a Japanese natural language sentence into word units by referring to a word segmentation dictionary, and then identifies effective keywords when retrieving the document text. In the keyword automatic extraction method for extracting, a part-of-speech recognition dictionary in which part-of-speech information of each word is stored, and a part-of-speech recognition unit that adds part-of-speech information to data divided into words by referring to the part-of-speech recognition dictionary. A phrase recognition unit that refers to the part-of-speech information added by the part-of-speech recognition unit to group a plurality of consecutive noun words into one word, and then combines the particle and auxiliary verb with the immediately preceding word to group them into phrase units; A phrase is stored from the data collected in the phrase by the phrase recognition means and the phrase dictionary that stores the relational information between the particle and the phrase that are used together with the phrase to compose the sentence. An automatic keyword extraction method, comprising: keyword extraction means for extracting a keyword by referring to the vocabulary including the phrase.
JP4252434A 1992-09-22 1992-09-22 Automatic key word extraction system Withdrawn JPH0749872A (en)

Priority Applications (1)

Application Number Priority Date Filing Date Title
JP4252434A JPH0749872A (en) 1992-09-22 1992-09-22 Automatic key word extraction system

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
JP4252434A JPH0749872A (en) 1992-09-22 1992-09-22 Automatic key word extraction system

Publications (1)

Publication Number Publication Date
JPH0749872A true JPH0749872A (en) 1995-02-21

Family

ID=17237321

Family Applications (1)

Application Number Title Priority Date Filing Date
JP4252434A Withdrawn JPH0749872A (en) 1992-09-22 1992-09-22 Automatic key word extraction system

Country Status (1)

Country Link
JP (1) JPH0749872A (en)

Cited By (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
KR100375746B1 (en) * 2000-02-09 2003-03-10 이만균 Method and system for processing internet command language and thereof program products
US9679166B2 (en) 2014-03-27 2017-06-13 Panasonic Intellectual Property Management Co., Ltd. Settlement terminal device
CN116977998A (en) * 2023-07-27 2023-10-31 江苏苏力机械股份有限公司 Workpiece feeding visual identification system and method for coating production line

Cited By (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
KR100375746B1 (en) * 2000-02-09 2003-03-10 이만균 Method and system for processing internet command language and thereof program products
US9679166B2 (en) 2014-03-27 2017-06-13 Panasonic Intellectual Property Management Co., Ltd. Settlement terminal device
CN116977998A (en) * 2023-07-27 2023-10-31 江苏苏力机械股份有限公司 Workpiece feeding visual identification system and method for coating production line

Similar Documents

Publication Publication Date Title
Cussens Part-of-speech tagging using Progol
Lita et al. Truecasing
US6345253B1 (en) Method and apparatus for retrieving audio information using primary and supplemental indexes
JP2583386B2 (en) Keyword automatic extraction device
Glavitsch et al. A system for retrieving speech documents
Zechner Automatic generation of concise summaries of spoken dialogues in unrestricted domains
US20020099744A1 (en) Method and apparatus providing capitalization recovery for text
WO2008107305A2 (en) Search-based word segmentation method and device for language without word boundary tag
WO1997004405A9 (en) Method and apparatus for automated search and retrieval processing
KR20030056655A (en) Similar sentence retrieval method for translation aid
JPWO2018097091A1 (en) Model creation device, text search device, model creation method, text search method, data structure, and program
Echeverry-Correa et al. Topic identification techniques applied to dynamic language model adaptation for automatic speech recognition
CN106294460A (en) A kind of Chinese speech keyword retrieval method based on word and word Hybrid language model
JP2572314B2 (en) Keyword extraction device
Ihm et al. Skip-gram-KR: Korean word embedding for semantic clustering
Yohannes et al. Amharic document clustering using semantic information from neural word embedding and encyclopedic knowledge
Kim et al. Question answering considering semantic categories and co-occurrence density
Psutka et al. System for fast lexical and phonetic spoken term detection in a Czech cultural heritage archive
JP3794597B2 (en) Topic extraction method and topic extraction program recording medium
Fujii et al. A method for open-vocabulary speech-driven text retrieval
Isaev et al. Towards Kyrgyz stop words
JPH10149370A (en) Document retrieval method and device using context information
JP2005025555A (en) Thesaurus construction system, thesaurus construction method, program for executing the method, and storage medium storing the program
Liu et al. Cross-language information matching technology based on term extraction
JPH11250063A (en) Search device and search method

Legal Events

Date Code Title Description
A300 Withdrawal of application because of no request for examination

Free format text: JAPANESE INTERMEDIATE CODE: A300

Effective date: 19991130