JPH1097286A - Word / phrase classification processing method, collocation extraction method, word / phrase classification processing device, speech recognition device, machine translation device, collocation extraction device, and word / phrase storage medium - Google Patents

Word / phrase classification processing method, collocation extraction method, word / phrase classification processing device, speech recognition device, machine translation device, collocation extraction device, and word / phrase storage medium

Info

Publication number
JPH1097286A
JPH1097286A JP9167243A JP16724397A JPH1097286A JP H1097286 A JPH1097286 A JP H1097286A JP 9167243 A JP9167243 A JP 9167243A JP 16724397 A JP16724397 A JP 16724397A JP H1097286 A JPH1097286 A JP H1097286A
Authority
JP
Japan
Prior art keywords
word
class
words
text data
classes
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Granted
Application number
JP9167243A
Other languages
Japanese (ja)
Other versions
JP3875357B2 (en
Inventor
Akira Shioda
明 潮田
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Fujitsu Ltd
Original Assignee
Fujitsu Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Fujitsu Ltd filed Critical Fujitsu Ltd
Priority to JP16724397A priority Critical patent/JP3875357B2/en
Publication of JPH1097286A publication Critical patent/JPH1097286A/en
Application granted granted Critical
Publication of JP3875357B2 publication Critical patent/JP3875357B2/en
Anticipated expiration legal-status Critical
Expired - Fee Related legal-status Critical Current

Links

Landscapes

  • Machine Translation (AREA)

Abstract

(57)【要約】 【課題】 単語と連語とをまとめて自動的に分類する。 【解決手段】 テキストデータにおいて出現する確率が
所定値以上の単語クラス列にトークンを付与し、テキス
トデータの単語・トークン列に含まれる単語とトークン
とが混在する集合を、テキストデータの単語・トークン
列の生成確率が最大になるように分割し、トークンをテ
キストデータに存在する連語に置換する。
(57) [Summary] [Problem] Automatically classify words and collocations. Kind Code: A1 A token is assigned to a word class string having a probability of occurrence in text data that is equal to or greater than a predetermined value, and a set in which words and tokens included in the word / token string of text data are mixed is defined as a word / token of text data. The sequence is divided so that the generation probability of the sequence is maximized, and the token is replaced with a collocation existing in the text data.

Description

【発明の詳細な説明】DETAILED DESCRIPTION OF THE INVENTION

【0001】[0001]

【発明の属する技術分野】本発明は、単語・連語分類処
理方法、連語抽出方法、単語・連語分類処理装置、音声
認識装置、機械翻訳装置、連語抽出装置及び単語・連語
記憶媒体に関し、特に、テキストデータの中から連語を
自動的に抽出し、単語及び連語を自動的に分類する場合
に好適なものである。
BACKGROUND OF THE INVENTION 1. Field of the Invention The present invention relates to a word / phrase classification processing method, a collocation extraction method, a word / phrase classification processing device, a speech recognition device, a machine translation device, a collocation extraction device, and a word / phrase storage medium. This is suitable for automatically extracting collocations from text data and automatically classifying words and collocations.

【0002】[0002]

【従来の技術】従来の単語分類処理装置には、例えば、
「Brown,P.,Della Pietra,
V.,deSouza,P.,Lai,J.,Merc
er,R.(1992)“Class−Based n
−gram Models ofNatural La
nguage”.Computational Lin
guistics,Vol.18,No4,pp.46
7−479」に記載されているように、テキストデータ
の中で使用されている単独の単語を統計的に処理するこ
とにより、単独の単語を自動的に分類するものがあり、
この単独の単語の分類結果を用いて音声認識や機械翻訳
を行っていた。
2. Description of the Related Art Conventional word classification processing apparatuses include, for example,
"Brown, P., Della Pietra,
V. , DeSouza, P .; , Lai, J. et al. , Merc
er, R .; (1992) "Class-Based n
-Gram Models of Natural La
nguage ".Computational Lin
guistics, Vol. 18, No4, pp. 46
7-479 ", a single word used in text data is statistically processed to automatically classify the single word.
Speech recognition and machine translation were performed using the classification result of the single word.

【0003】[0003]

【発明が解決しようとする課題】しかしながら、従来の
単語分類処理装置は、単語と連語とをまとめて自動的に
分類することができず、単語と連語あるいは連語と連語
の対応関係や類似度を用いて、音声認識や機械翻訳を行
うことがきないため、音声認識や機械翻訳を正確に実行
することができないという問題があった。
However, the conventional word classification processing device cannot automatically classify words and collocations together, and cannot determine the correspondence and similarity between words and collocations or collocations and collocations. However, since speech recognition and machine translation cannot be performed using this method, speech recognition and machine translation cannot be performed accurately.

【0004】そこで、本発明の第1の目的は、単語と連
語とをまとめて自動的に分類することが可能な単語・連
語分類処理方法及び単語・連語分類処理装置を提供する
ことである。
Accordingly, a first object of the present invention is to provide a word / phrase classification processing method and a word / phrase classification processing device that can automatically classify words and collocations collectively.

【0005】また、本発明の第2の目的は、大量のテキ
ストデータから高速に連語を抽出することが可能な連語
抽出装置を提供することである。また、本発明の第3の
目的は、単語と連語あるいは連語と連語の対応関係や類
似度を用いることにより、正確な音声認識が可能な音声
認識装置を提供することである。
A second object of the present invention is to provide a collocation extracting apparatus capable of extracting a collocation from a large amount of text data at high speed. A third object of the present invention is to provide a speech recognition device capable of performing accurate speech recognition by using the correspondence and similarity between words and collocations or between collocations and collocations.

【0006】また、本発明の第4の目的は、単語と連語
あるいは連語と連語の対応関係や類似度を用いることに
より、正確な機械翻訳が可能な機械翻訳装置を提供する
ことである。
A fourth object of the present invention is to provide a machine translation apparatus capable of performing accurate machine translation by using the correspondence and similarity between words and collocations or between collocations and collocations.

【0007】[0007]

【課題を解決するための手段】上述した第1の目的を達
成するために、本発明によれば、テキストデータに含ま
れる単語と連語とを一緒に分類して、単語と連語とが混
在するクラスを生成するようにしている。
According to the present invention, in order to achieve the first object described above, words and collocations are classified together and words and collocations are mixed. Class is generated.

【0008】このことにより、単語と単語とをまとめて
分類するだけでなく、単語と連語あるいは連語と連語と
をまとめて一緒に分類することができ、単語と連語ある
いは連語と連語との対応関係や類似度を容易に判別する
ことができる。
[0008] This makes it possible not only to classify words and words collectively, but also to classify words and collocations or collocations and collocations together. And similarity can be easily determined.

【0009】また、本発明の一態様によれば、単語を分
類した単語クラスをテキストデータの単語の一次元列に
マッピングして単語クラスの一次元列を生成し、テキス
トデータの単語クラスの一次元列において、隣接する単
語クラス間の粘着度が全て所定値以上の単語クラス列を
抽出してその単語クラス列にトークンを付与し、単語と
トークンとを一緒に分類してから、トークンに対応する
単語クラス列をその単語クラス列に属する連語で置換す
るようにしている。
According to another aspect of the present invention, a word class obtained by classifying words is mapped to a one-dimensional sequence of words in text data to generate a one-dimensional sequence of word classes, and a primary class of the word class in text data is generated. In the original sequence, a word class sequence in which the degree of adhesion between adjacent word classes is all greater than or equal to a predetermined value is extracted, a token is added to the word class sequence, words and tokens are classified together, and Is replaced by a collocation belonging to the word class sequence.

【0010】このことにより、単語クラス列にトークン
を付与してその単語クラス列を1つの単語とみなし、テ
キストデータに含まれる単語とトークンを付与された単
語クラス列とを同等に取り扱って単語と連語との区別な
く分類処理を行うことができる。また、単語を分類した
単語クラスをテキストデータの単語の一次元列にマッピ
ングして単語クラスの一次元列を生成し、隣接する単語
クラス間の粘着度に基づいて連語を抽出することによ
り、テキストデータからの連語の抽出を高速に行うこと
ができる。
Thus, a token is assigned to the word class string, the word class string is regarded as one word, and the word included in the text data and the word class string to which the token is assigned are treated equally, and Classification processing can be performed without distinction from collocations. In addition, by mapping a word class obtained by classifying words to a one-dimensional sequence of words in text data to generate a one-dimensional sequence of word classes, and extracting collocations based on the degree of adhesion between adjacent word classes, Extraction of collocations from data can be performed at high speed.

【0011】また、上述した第2の目的を達成するため
に、本発明によれば、単語を分類した単語クラスをテキ
ストデータの単語の一次元列にマッピングして単語クラ
スの一次元列を生成し、テキストデータの単語クラスの
一次元列において、隣接する単語クラス間の粘着度が全
て所定値以上の単語クラス列を抽出し、単語クラス列を
構成する個々の単語クラスから、テキストデータに隣接
して存在する個々の単語を別々に取り出して連語を抽出
するようにしている。
According to the present invention, in order to achieve the second object, a word class in which words are classified is mapped to a one-dimensional string of words in text data to generate a one-dimensional string of word classes. Then, in the one-dimensional sequence of the word classes of the text data, a word class sequence in which the degree of adhesion between adjacent word classes is all equal to or greater than a predetermined value is extracted, and the individual word classes constituting the word class sequence are adjacent to the text data. Each word that exists is extracted separately to extract collocations.

【0012】このことにより、単語クラス列に基づいて
連語を抽出することができ、テキストデータに存在する
異なる単語の数よりも、それらの単語を分類した単語ク
ラスの数のほうが少ないので、テキストデータの単語ク
ラスの一次元列において、隣接する単語クラス間の粘着
度が所定値以上の単語クラス列を抽出するほうが、テキ
ストデータの単語の一次元列において、隣接する単語間
の粘着度が所定値以上の単語列を抽出する場合に比べ
て、演算量及びメモリ容量を少なくすることができ、連
語の抽出処理を高速に行うことができるとともに、メモ
リ資源を節約できる。なお、単語クラス列には、テキス
トデータの単語の一次元列に存在しない単語列が含まれ
ている場合があるので、単語クラス列を構成する個々の
単語クラスから、テキストデータに隣接して存在する個
々の単語を別々に取り出して連語としている。
[0012] Thus, collocations can be extracted based on the word class sequence, and the number of word classes in which those words are classified is smaller than the number of different words existing in the text data. In the one-dimensional sequence of word classes, it is better to extract a word class sequence in which the degree of adhesion between adjacent word classes is equal to or more than a predetermined value. Compared with the case of extracting the above word strings, the amount of calculation and the memory capacity can be reduced, the collocation extraction processing can be performed at high speed, and memory resources can be saved. Note that the word class string may include a word string that does not exist in the one-dimensional string of the words in the text data. The individual words to be taken are separately taken out as collocations.

【0013】また、上述した第3の目的を達成するため
に、本発明によれば、所定のテキストデータに含まれる
単語と連語とを、単語と連語とが混在するクラスに分類
して格納している単語・連語辞書を参照することによ
り、発音音声を音声認識するようにしている。
According to the present invention, in order to achieve the third object, words and collocations included in predetermined text data are classified and stored in a class in which words and collocations are mixed. By referring to the word / syllable dictionary, the pronunciation voice is recognized.

【0014】このことにより、単語と連語あるいは連語
と連語の対応関係や類似度を用いながら音声認識を行う
ことができ、正確な処理が可能になる。また、上述した
第4の目的を達成するために、本発明によれば、所定の
テキストデータに含まれる単語と連語とを、単語と連語
とが混在するクラスに分類して格納している単語・連語
辞書に基づいて、用例文集に格納されている用例原文と
入力された原文とを対応させるようにしている。
Thus, speech recognition can be performed using the correspondence and similarity between words and collocations or between collocations and collocations, and accurate processing can be performed. According to the present invention, in order to achieve the fourth object described above, words and collocations included in predetermined text data are classified and stored in a class in which words and collocations are mixed and stored. -Based on the collocation dictionary, the example original sentence stored in the example sentence collection is made to correspond to the input original sentence.

【0015】このことにより、用例文集に格納されてい
る用例原文の単語が連語に置き換わった原文が入力され
た場合においても、入力された原文に用例原文を適用し
て機械翻訳を行うことができ、単語と連語あるいは連語
と連語の対応関係や類似度を用いた正確な機械翻訳が可
能になる。
Thus, even when an original sentence in which words of the example original sentence stored in the example sentence collection are replaced with collocations is input, machine translation can be performed by applying the example original sentence to the input original sentence. In addition, accurate machine translation using the correspondence and similarity between words and collocations or between collocations and collocations becomes possible.

【0016】[0016]

【発明の実施の形態】以下、本発明の一実施例に係わる
単語・連語分類処理装置について図面を参照しながら説
明する。この実施例は、所定のテキストデータに含まれ
る単語と連語とを、単語と連語とが混在するクラスに分
類するものである。
DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS Hereinafter, a word / phrase classification processing apparatus according to an embodiment of the present invention will be described with reference to the drawings. In this embodiment, words and collocations included in predetermined text data are classified into classes in which words and collocations are mixed.

【0017】図1は、本発明の一実施例に係わる単語・
連語分類処理装置の機能的な構成を示すブロック図であ
る。図1において、単語分類手段1は、テキストデータ
の単語の一次元列から互いに異なる単語を抽出し、抽出
された単語の集合を分割して単語クラスを生成する。
FIG. 1 is a diagram showing words and words according to an embodiment of the present invention.
It is a block diagram which shows the functional structure of a collocation classification processing apparatus. In FIG. 1, a word classifying unit 1 extracts words different from each other from a one-dimensional string of words in text data, and divides a set of the extracted words to generate a word class.

【0018】図2は、単語分類手段1の処理を説明する
もので、テキストデータに含まれるT個の単語よりなる
単語の一次元列(w1 2 3 4 ・・・wT )から、
テキストデータでの出現頻度順に並べたV個のボキャブ
ラリーとしての単語{v1 、v2 、v3 、v4 、・・
・、vV }を生成し、このテキストデータのボキャブラ
リーとしての単語{v1 、v2 、v3 、v4 、・・・、
V }のそれぞれに初期化クラスを割り当てる。ここ
で、単語の個数T個は、例えば、5000万個であり、
ボキャブラリーの個数V個は、例えば、7000個であ
る。
FIG. 2 explains the processing of the word classifying means 1 and is a one-dimensional sequence of words (w 1 w 2 w 3 w 4 ... W T ) composed of T words contained in the text data. From
V vocabulary words {v 1 , v 2 , v 3 , v 4 ,... Arranged in the order of appearance frequency in text data
, V V }, and the words {v 1 , v 2 , v 3 , v 4 ,..., As a vocabulary of the text data.
v V } is assigned an initialization class. Here, the number T of words is, for example, 50 million,
The number V of vocabularies is, for example, 7,000.

【0019】図2の例では、テキストデータでの出現頻
度が高い、例えば、“the”、“a”、“in”、
“of”が、それぞれボキャブラリーとしての単語
1 、v2、v3 、v4 に対応している。初期化クラス
を割り当てられたV個のボキャブラリーとしての単語
{v1 、v2 、v3 、v4 、・・・、vV }は、クラス
タリングによりC個の単語クラス{C1 、C2 、C3
4 、・・・、CC }に分割される。ここで、単語クラ
スの個数C個は、例えば、500個である。
In the example of FIG. 2, the frequency of appearance in the text data is high, for example, "the", "a", "in",
“Of” corresponds to words v 1 , v 2 , v 3 , and v 4 as vocabularies, respectively. The words {v 1 , v 2 , v 3 , v 4 ,..., V V } as V vocabularies to which the initialization class is assigned are converted into C word classes {C 1 , C 2 , C 3 ,
C 4 ,..., C C }. Here, the number C of the word classes is, for example, 500.

【0020】また、図2では、例えば、“spea
k”、“say”、“tell”、“talk”・・・
が単語クラスC1 に分類され、“he”、“she”、
“it”・・・が単語クラスC5 に分類され、“ca
r”、“track”、“wagon”・・・が単語ク
ラスC32に分類され、“Toyota”、“Nissa
n”、“GM”・・・が単語クラスC300 に分類されて
いる例を示している。
In FIG. 2, for example, "spea"
k "," say "," tell "," talk "...
There are classified as word class C 1, "he", " she",
"She" ··· is classified in the word class C 5, "ca
r "," track "," wagon "··· are classified in the word class C 32," Toyota "," Nissa
n "," GM "··· is an example that has been classified as word class C 300.

【0021】このV個のボキャブラリーとしての単語
{v1 、v2 、v3 、v4 、・・・、vV }よりなる単
語の分類は、例えば、テキストデータに存在する2つの
単語がおのおの属する2つの単語クラスをマージした場
合、元のテキストデータの生成確率の減少が最も少なく
なるものを同一の単語クラスに統合することにより行
う。ここで、元のテキストデータのクラスバイモデルに
よる生成確率は、平均相互情報量AMIを用いて表現す
ることができ、この平均相互情報量AMIは以下の式に
より表すことができる。
The classification of the words composed of the V words {v 1 , v 2 , v 3 , v 4 ,..., V V } as V vocabulary is, for example, such that two words existing in text data are When two word classes that belong to each other are merged, the one that minimizes the decrease in the generation probability of the original text data is integrated into the same word class. Here, the generation probability of the original text data by the class-by model can be expressed by using the average mutual information AMI, and the average mutual information AMI can be expressed by the following equation.

【0022】[0022]

【数1】 (Equation 1)

【0023】ここで、Pr(Ci )は、テキストデータ
の単語の一次元列(w1 2 3 4 ・・・wT )をそ
の単語が属する単語クラスで置き換えた場合、そのテキ
ストデータの単語クラスの一次元列でのクラスCi の出
現確率、Pr(Cj )は、テキストデータの単語の一次
元列(w1 2 3 4 ・・・wT )をその単語が属す
る単語クラスで置き換えた場合、そのテキストデータの
単語クラスの一次元列でのクラスCj の出現確率、Pr
(Ci 、Cj )は、テキストデータの単語の一次元列
(w1 2 3 4・・・wT )を、その単語が属する
単語クラスで置き換えた場合、そのテキストデータの単
語クラスの一次元列での単語クラスCi の次に隣接して
単語クラスC j が出現する確率である。
Here, Pr (Ci) Is text data
One-dimensional sequence of words (w1wTwowThreewFour... wT)
Replaced by the word class to which the word belongs, the text
Class C in a one-dimensional sequence of word classesiOut of
Current probability, Pr (Cj) Is the primary word of the text data
Original column (w1wTwowThreewFour... wT) To which the word belongs
If you replace it with a word class,
Class C in a one-dimensional sequence of word classesjProbability of occurrence of Pr
(Ci, Cj) Is a one-dimensional sequence of words in the text data
(W1wTwowThreewFour... wT) To which the word belongs
When replaced with a word class, the text data
Word class C in one-dimensional sequence of word classesiNext to
Word class C jIs the probability that appears.

【0024】図3は、図1の単語分類手段1の機能的な
構成の一例を示すブロック図である。図3において、初
期化クラス設定部10は、テキストデータの単語の一次
元列{w1 2 3 4 ・・・wT }から互いに異なる
単語を抽出し、所定の出現頻度を有する単語{v1 、v
2 、v3 、v4 、・・・、vV }のそれぞれに固有の単
語クラス{C1 、C2 、C3 、C4 、・・・、CV }を
割り当てる。
FIG. 3 is a block diagram showing an example of a functional configuration of the word classifying means 1 of FIG. 3, the initialization class setting unit 10 extracts different words from each other from the one-dimensional string of words in the text data {w 1 w 2 w 3 w 4 ··· w T}, words having a predetermined frequency {V 1 , v
2, v 3, v 4, ···, v V each unique word class {C 1 of}, C 2, C 3, C 4, assigns ..., and C V}.

【0025】仮マージ部11は、単語クラスの集合{C
1 、C2 、C3 、C4 、・・・、C M }から2つの単語
クラス{Ci 、Cj }を取り出して仮マージする。平均
相互情報量算出部12は、テキストデータの仮マージさ
れた単語クラス{C1 、C2 、C3 、C4 、・・・、C
M-1 }についての平均相互情報量AMIを(1)式によ
り算出する。この場合、M個の単語クラスの集合
{C1 、C2 、C 3 、C4 、・・・、CM }から2つの
単語クラス{Ci 、Cj }を取り出だす取り出しかた
は、M(M−1)/2個だけ存在するので、M(M−
1)/2回の平均相互情報量AMIの計算を行う必要が
ある。
The temporary merging unit 11 generates a set of word classes {C
1, CTwo, CThree, CFour, ..., C MTwo words from}
Class II Ci, Cj} Is taken out and temporarily merged. average
The mutual information calculation unit 12 calculates the temporary merging of the text data.
Word class {C1, CTwo, CThree, CFour, ..., C
M-1The average mutual information AMI for} is given by equation (1).
Calculated. In this case, a set of M word classes
{C1, CTwo, C Three, CFour, ..., CMTwo from}
Word class $ Ci, CjHow to remove}
Is M (M−1) / 2, so M (M−
1) It is necessary to calculate the average mutual information AMI twice.
is there.

【0026】本マージ部13は、仮マージにより計算さ
れたM(M−1)/2個の平均相互情報量AMIの基づ
いて、平均相互情報量AMIを最大とする2つの単語ク
ラス{Ci 、Cj }を単語クラスの集合{C1 、C2
3 、C4 、・・・、CM }から取り出して本マージす
る。このことにより、本マージされたいずれかの単語ク
ラス{Ci 、Cj }に属する単語は、同一の単語クラス
に分類される。
Based on the M (M-1) / 2 pieces of average mutual information AMI calculated by the provisional merge, the merging unit 13 generates two word classes {C i, which maximize the average mutual information AMI. , C j } to a set of word classes {C 1 , C 2 ,
C 3 , C 4 ,..., C M } Thus, the words belonging to any of the merged word classes {C i , C j } are classified into the same word class.

【0027】図1の単語クラス列生成手段2は、テキス
トデータの単語列(w1 2 3 4 ・・・wT )を構
成する個々の単語を、単語が属する単語クラス{C1
2、C3 、C4 、・・・、CV }で置換することによ
り、テキストデータの単語クラス列を生成する。
The word class sequence generating means 2 shown in FIG.
Word string (w1wTwowThreew Four... wT)
Each word to be formed is represented by a word class {C} to which the word belongs.1,
CTwo, CThree, CFour, ..., CVBy replacing with}
Then, a word class sequence of the text data is generated.

【0028】図4は、テキストデータの単語クラスの一
次元列の一例を示す図である。図4において、単語分類
手段1によりC個の単語クラス{C1 、C2 、C3 、C
4 、・・・、CC }が生成されているものとし、例え
ば、単語クラスC1 には、ボキャブラリーv1 、v37
・・・が属しており、単語クラスC2 には、ボキャブラ
リーv3 、v15、・・・が属しており、単語クラスC3
には、ボキャブラリーv2 、v4 、・・・が属してお
り、単語クラスC4 には、ボキャブラリーv 7 、v9
・・・が属しており、単語クラスC5 には、ボキャブラ
リーv6 、v 8 、v26、vV 、・・・が属しており、単
語クラスC6 には、ボキャブラリーv 6 、v23、・・・
が属しており、単語クラスC7 には、ボキャブラリーv
5 、v 10、・・・が属しているものとする。
FIG. 4 shows a word class of text data.
It is a figure showing an example of a dimension sequence. In FIG. 4, word classification
By means 1, C word classes {C1, CTwo, CThree, C
Four, ..., CC} Has been generated, and
For example, word class C1The vocabulary v1, V37,
.. Belong to the word class CTwoThe vocabulary
Lee vThree, VFifteen, ... belong to the word class CThree
The vocabulary vTwo, VFour... belong to
Word class CFourThe vocabulary v 7, V9,
.. Belong to the word class CFiveThe vocabulary
Lee v6, V 8, V26, VV, ... belong to
Word Class C6The vocabulary v 6, Vtwenty three...
Belongs to the word class C7The vocabulary v
Five, V Ten, ... belong.

【0029】また、テキストデータの単語の一次元列
(w1 2 3 4 ・・・wT )において、例えば、単
語w1 が示すボキャブラリーとしての単語がv15、単語
2 が示すボキャブラリーとしての単語がv2 、単語w
3 が示すボキャブラリーとしての単語がv23、単語w4
が示すボキャブラリーとしての単語がv4 、単語w5
示すボキャブラリーとしての単語がv5 、単語w6 が示
すボキャブラリーとしての単語がv15、単語w7 が示す
ボキャブラリーとしての単語がv5 、単語w8 が示すボ
キャブラリーとしての単語がv26、単語w9 が示すボキ
ャブラリーとしての単語がv37、単語w10が示すボキャ
ブラリーとしての単語がv2 、・・・、単語wT が示す
ボキャブラリーとしての単語がv8 であるとする。
In the one-dimensional string (w 1 w 2 w 3 w 4 ... W T ) of the words of the text data, for example, the word as the vocabulary indicated by the word w 1 is represented by v 15 and the word w 2 is represented by w 15 The word as vocabulary is v 2 , the word w
The vocabulary indicated by 3 is the word v 23 and the word w 4
Words as vocabulary indicated is v 4, word w 5 is word v 5 as vocabulary shown, the words in the vocabulary word w 6 are indicated v 15, word w 7 words v 5 as vocabulary indicated, word The word as the vocabulary indicated by w 8 is v 26 , the word as the vocabulary indicated by word w 9 is v 37 , the word as the vocabulary indicated by word w 10 is v 2 ,..., the vocabulary indicated by word w T word is assumed to be v 8.

【0030】この場合、ボキャブラリーv15は単語クラ
スC2 に属しているので、単語w1は単語クラスC2
マッピングされ、ボキャブラリーv2 は単語クラスC3
に属しているので、単語w2 は単語クラスC3 にマッピ
ングされ、ボキャブラリーv 23は単語クラスC6 に属し
ているので、単語w3 は単語クラスC6 にマッピングさ
れ、ボキャブラリーv4 は単語クラスC3 に属している
ので、単語w4 は単語クラスC3 にマッピングされ、ボ
キャブラリーv5 は単語クラスC7 に属しているので、
単語w5 は単語クラスC7 にマッピングされ、ボキャブ
ラリーv15は単語クラスC2 に属しているので、単語w
6 は単語クラスC2 にマッピングされ、ボキャブラリー
5 は単語クラスC7 に属しているので、単語w7 は単
語クラスC7 にマッピングされ、ボキャブラリーv26
単語クラスC5 に属しているので、単語w8 は単語クラ
スC5 にマッピングされ、ボキャブラリーv37は単語ク
ラスC1 に属しているので、単語w9 は単語クラスC1
にマッピングされ、ボキャブラリーv2 は単語クラスC
3 に属しているので、単語w10は単語クラスC3 にマッ
ピングされ、・・・、ボキャブラリーv8 は単語クラス
5 に属しているので、単語wT は単語クラスC5 にマ
ッピングされる。
In this case, the vocabulary vFifteenIs the word kura
STwoThe word w1Is the word class CTwoTo
Mapped, vocabulary vTwoIs the word class CThree
The word wTwoIs the word class CThreeMapi
Vocabulary v twenty threeIs the word class C6Belongs to
The word wThreeIs the word class C6Mapped to
Vocabulary vFourIs the word class CThreeBelongs to
So the word wFourIs the word class CThreeIs mapped to
CabrioFiveIs the word class C7Belongs to
Word wFiveIs the word class C7Mapped to the vocabulary
Rally vFifteenIs the word class CTwoThe word w
6Is the word class CTwoMapped to vocabulary
vFiveIs the word class C7The word w7Is simply
Word Class C7Is mapped to the vocabulary v26Is
Word class CFiveThe word w8Is the word kura
SFiveIs mapped to the vocabulary v37Is the word
Las C1The word w9Is the word class C1
Is mapped to the vocabulary vTwoIs the word class C
ThreeThe word wTenIs the word class CThreeTo
Pinged, ..., vocabulary v8Is a word class
CFiveThe word wTIs the word class CFiveNima
Be pinged.

【0031】すなわち、テキストデータの単語の一次元
列(w1 2 3 4 ・・・wT )が、C個の単語クラ
ス{C1 、C2 、C3 、C4 、・・・、CC }によりマ
ッピングされた結果として、テキストデータの単語クラ
スの一次元列(C2 3 63 7 2 7 5 1
3 ・・・C5 )が1対1対応で生成される。
That is, the one-dimensional sequence (w 1 w 2 w 3 w 4 ... W T ) of the words of the text data is represented by C word classes {C 1 , C 2 , C 3 , C 4 ,. ., As a result of mapping by C C }, a one-dimensional sequence of word classes of text data (C 2 C 3 C 6 C 3 C 7 C 2 C 7 C 5 C 1)
C 3 ... C 5 ) are generated in a one-to-one correspondence.

【0032】図1の単語クラス列抽出手段3は、テキス
トデータの単語クラスの一次元列においての単語クラス
間の粘着度が全て所定値以上の単語クラス列を、テキス
トデータの単語クラスの一次元列から抽出する。ここ
で、単語クラス間の粘着度は、単語クラス列を構成する
単語クラス間のつながりの強さを示す指標であり、この
粘着度を表現するものとして、例えば、相互情報量M
I、相関係数、コサインメジャー、liklihood
ratioなどがある。
The word class string extracting means 3 shown in FIG. 1 converts a word class string in which the degree of adhesion between word classes in the one-dimensional string of the word class of the text data is all equal to or more than a predetermined value into one-dimensional word class of the text data. Extract from a column. Here, the degree of adhesion between the word classes is an index indicating the strength of connection between the word classes constituting the word class sequence. As an expression of the degree of adhesion, for example, the mutual information M
I, correlation coefficient, cosine measure, liklihood
ratio.

【0033】以下の説明では、単語クラス間の粘着度と
して、相互情報量MIを用いることにより、テキストデ
ータの単語クラスの一次元列から単語クラス列を抽出す
る場合を例にとる。
In the following description, a case where a word class string is extracted from a one-dimensional string of a word class of text data by using the mutual information MI as the degree of adhesion between word classes will be described as an example.

【0034】図5は、単語クラス列抽出手段3により抽
出された単語クラス列の一例を示す図である。図5にお
いて、テキストデータの単語の一次元列(w1 2 3
4 5 67 ・・・wT )に対してマッピングされ
た結果として、テキストデータの単語クラスの一次元列
(C2 3 6 3 7 2 7 ・・・C5 )が1対1
対応で生成されているものとする。このテキストデータ
の単語クラスの一次元列(C23 6 3 7 2
7 ・・・C5 )から、隣接する2つの単語クラス
(Ci、Cj )を順次に取り出し、隣接する2つの単語
クラス(Ci 、Cj )についての相互情報量MI
(Ci 、Cj )を、以下の(2)式により計算する。
FIG. 5 is a diagram showing an example of a word class string extracted by the word class string extracting means 3. In FIG. 5, a one-dimensional string (w 1 w 2 w 3
.. w T ), the one-dimensional sequence (C 2 C 3 C 6 C 3 C 7 C 2 C 7 ...) of the word class of the text data is obtained as a result of mapping to w 4 w 5 w 6 w 7. C 5 ) is one-on-one
It is assumed that it is generated in correspondence. One-dimensional sequence (C 2 C 3 C 6 C 3 C 7 C 2 C) of the word class of this text data
From 7 ··· C 5), two adjacent word classes (C i, C j) sequentially taken out, two adjacent word classes (C i, mutual information MI about C j)
(C i , C j ) is calculated by the following equation (2).

【0035】 MI(Ci 、Cj ) =log{Pr(Ci 、Cj )/(Pr(Ci )Pr(Cj ))} ・・・(2) そして、隣接する2つの単語クラス(Ci 、Cj )につ
いての相互情報量MI(Ci 、Cj )が所定のしきい値
TH以上の場合、これら隣接する2つの単語クラス(C
i 、Cj )をクラスチェーンで結んで互いに関連づけ
る。
MI (C i , C j ) = log {Pr (C i , C j ) / (Pr (C i ) Pr (C j ))} (2) and two adjacent word classes (C i, C j) mutual information MI (C i, C j) for the case of more than a predetermined threshold value TH, 2 single word class thereof adjacent (C
i , C j ) are connected by a class chain.

【0036】例えば、図5において、隣接する2つの単
語クラス(C2 、C3 )についての相互情報量MI(C
2 、C3 )、隣接する2つの単語クラス(C3 、C6
についての相互情報量MI(C3 、C6 )、隣接する2
つの単語クラス(C6 、C3)についての相互情報量M
I(C6 、C3 )、隣接する2つの単語クラス(C3
7 )についての相互情報量MI(C3 、C7 )、隣接
する2つの単語クラス(C7 、C2 )についての相互情
報量MI(C7 、C2 )、隣接する2つの単語クラス
(C2 、C7 )についての相互情報量MI(C2
7 )、・・・を(2)式により順次に計算する。
For example, in FIG. 5, the mutual information MI (C 2 ) for two adjacent word classes (C 2 , C 3 )
2 , C 3 ), two adjacent word classes (C 3 , C 6 )
Mutual information MI (C 3 , C 6 ) for
Mutual information M about two word classes (C 6 , C 3 )
I (C 6 , C 3 ), two adjacent word classes (C 3 ,
C 7) mutual information MI about (C 3, C 7), mutual information MI about two adjacent word classes (C 7, C 2) ( C 7, C 2), 2 single word class adjacent (C 2, C 7) mutual information MI (C 2 for,
C 7 ),... Are sequentially calculated by equation (2).

【0037】そして、相互情報量MI(C2 、C3 )、
相互情報量MI(C3 、C7 )、相互情報量MI
(C7 、C2 )、・・・がしきい値TH以上で、相互情
報量MI(C3 、C6 )、相互情報量MI(C6
3 )、相互情報量MI(C2 、C7 )、・・・がしき
い値THより小さい場合、隣接する2つの単語クラス
(C2 、C 3 )、(C3 、C7 )、(C7 、C2 )、・
・・をそれぞれクラスチェーンで結ぶことにより、単語
クラス列C2 −C3 、C3 −C7 −C2 、・・・を抽出
する。
Then, the mutual information MI (CTwo, CThree),
Mutual information MI (CThree, C7), Mutual information MI
(C7, CTwo), ... are greater than or equal to the threshold value TH
Information MI (CThree, C6), Mutual information MI (C6,
CThree), Mutual information MI (CTwo, C7), ...
If the value is less than TH, two adjacent word classes
(CTwo, C Three), (CThree, C7), (C7, CTwo),
・ ・ By connecting each with a class chain, the word
Class column CTwo-CThree, CThree-C7-CTwoExtract ...
I do.

【0038】図6は、図1の単語クラス列抽出手段3の
機能的な構成の一例を示すブロック図である。図6にお
いて、単語クラス取出部30は、テキストデータの単語
クラスの一次元列から、隣接して存在する2つの単語ク
ラス(Ci 、Cj )を順次に取り出す。
FIG. 6 is a block diagram showing an example of a functional configuration of the word class string extracting means 3 of FIG. In FIG. 6, the word class extracting unit 30 sequentially extracts two adjacent word classes (C i , C j ) from a one-dimensional sequence of the word classes of the text data.

【0039】相互情報量算出部31は、単語クラス取出
部30により取り出した2つの単語クラス(Ci
j )の相互情報量MI(Ci 、Cj )を(2)式によ
り算出する。
The mutual information calculator 31 calculates the two word classes (C i ,
Mutual information MI (C i of C j), is calculated by the C j) (2) expression.

【0040】クラスチェーン結合部32は、相互情報量
MI(Ci 、Cj )が所定のしきい値以上の2つの単語
クラス(Ci 、Cj )をクラスチェーンで結ぶ。図1の
トークン付与手段4は、単語クラス列抽出手段3により
クラスチェーンで結ばれた単語クラス列にトークンを付
与する。
The class chain connecting unit 32 connects two word classes (C i , C j ) whose mutual information MI (C i , C j ) is equal to or larger than a predetermined threshold value by a class chain. The token assigning means 4 in FIG. 1 assigns a token to the word class strings connected by the class chain by the word class string extracting means 3.

【0041】図7は、トークン付与手段4により付与さ
れたトークンの一例を示す図である。図7において、ク
ラスチェーンで結ばれた単語クラス列は、例えば、C1
−C 3 、C1 −C7 、・・・、C2 −C3 、C2
11、・・・、C300 −C32、・・・、C1 −C3 −C
80、C1 −C4 −C5 、C3 −C7 −C2 、・・・、C
1−C9 −C11−C32、・・・とする。この場合、単語
クラス列C1 −C3 に対してトークンt1 を付与し、単
語クラス列C1 −C7 に対してトークンt2 を付与し、
・・・、単語クラス列C2 −C3 に対してトークンt3
を付与し、単語クラス列C2 −C11に対してトークンt
4 を付与し、・・・、単語クラス列C300 −C32に対し
てトークンt5 を付与し、、・・・、単語クラス列C1
−C3 −C80に対してトークンt6 を付与し、単語クラ
ス列C1 −C4 −C5 に対してトークンt7 を付与し、
単語クラス列C3 −C7 −C2 に対してトークンt8
付与し、・・・、単語クラス列C1 −C9 −C11−C32
に対してトークンt9 を付与する。
FIG. 7 shows the state of the token provided by the token providing means 4.
FIG. 6 is a diagram illustrating an example of a token obtained. In FIG.
A word class sequence connected by a lath chain is, for example, C1
-C Three, C1-C7, ..., CTwo-CThree, CTwo
C11, ..., C300-C32, ..., C1-CThree-C
80, C1-CFour-CFive, CThree-C7-CTwo, ..., C
1-C9-C11-C32, ... In this case, the word
Class column C1-CThreeFor the token t1And simply
Word class sequence C1-C7For the token tTwo, And
..., word class sequence CTwo-CThreeFor the token tThree
And the word class sequence CTwo-C11For the token t
Four, ..., word class sequence C300-C32Against
T tokenFive,..., Word class sequence C1
-CThree-C80For the token t6Grant the word club
Row C1-CFour-CFiveFor the token t7, And
Word class sequence CThree-C7-CTwoFor the token t8To
…, Word class sequence C1-C9-C11-C32
For the token t9Is given.

【0042】図1の単語・トークン列生成手段5は、テ
キストデータの単語の一次元列(w 1 2 3 4 5
6 7 ・・・wT )のうち、単語クラス列抽出手段4
により抽出された単語クラス列に属する単語列をトーク
ンで置換することにより、テキストデータの単語・トー
クンの一次元列を生成する。
The word / token string generation means 5 in FIG.
One-dimensional sequence of words in the text data (w 1wTwowThreewFourwFive
w6w7... wT), Word class string extracting means 4
Talks word strings belonging to word class strings extracted by
By replacing with words, words and words in the text data
Generate a one-dimensional sequence of kuns.

【0043】図8は、テキストデータの単語・トークン
の一次元列の一例を示す図である。図8において、テキ
ストデータの単語の一次元列(w1 2 3 4 5
67 ・・・wT )に対してマッピングされた結果とし
て、テキストデータの単語クラスの一次元列(C2 3
6 3 7 2 7 ・・・C5 )が1対1対応で生成
されているものとし、クラスチェーンで結ばれた単語ク
ラス列C2 −C3 、C3 −C7 −C2 、・・・に対し
て、図7に示すように、トークンt3 、t8 、・・・が
付与されているものとする。
FIG. 8 is a diagram showing an example of a one-dimensional sequence of words and tokens of text data. In FIG. 8, a one-dimensional string (w 1 w 2 w 3 w 4 w 5 w
6 w 7 ... W T ), the result is a one-dimensional sequence (C 2 C 3 ) of the word class of the text data.
C 6 C 3 C 7 C 2 C 7 ... C 5 ) are generated in a one-to-one correspondence, and word class strings C 2 -C 3 and C 3 -C 7- connected by class chains. It is assumed that tokens t 3 , t 8 ,... Are given to C 2 ,.

【0044】この場合、クラスチェーンで結ばれた単語
クラス列C2 −C3 に属するテキストデータの単語列
(w1 2 )をトークンt3 で置き換え、クラスチェー
ンで結ばれた単語クラス列C3 −C7 −C2 に属するテ
キストデータの単語列(w4 5 6 )をトークンt8
で置き換えることにより、テキストデータの単語・トー
クンの一次元列(t3 3 8 7 ・・・wT )を生成
する。
In this case, the words connected by the class chain
Class column CTwo-CThreeWord string of text data belonging to
(W1wTwo) To the token tThreeReplace with
Word class sequence CThree-C7-CTwoBelongs to
Word string of text data (wFourw Fivew6) To the token t8
Can be replaced by words / toes in text data.
One-dimensional sequence of kung (tThreewThreet8w7... wT)Generate a
I do.

【0045】図9は、テキストデータの単語・トークン
の一次元列の一例を英文を例にとって示す図である。図
9(b)のテキストデータの単語の一次元列(w1 2
3 4 5 6 7 8 9 1011121314
15)として、図9(a)の“He wentto th
e apartment by bus and sh
e went to New York by pla
ne”が対応しているものとし、この単語の一次元列
(w1 2 3 4 5 6 7 8 9 101112
13 1415)に1対1で対応する単語クラスの一次元
列が図9(c)の(C5 90 3 2118101 32
2 5 903 6328101 32)で与えられるもの
とする。
FIG. 9 shows words and tokens of text data.
FIG. 5 is a diagram showing an example of a one-dimensional sequence in English. Figure
9 (b) one-dimensional string of words in the text data (w1wTwo
wThreewFourwFivew6w 7w8w9wTenw11w12w13w14w
Fifteen) As “Hewent to th” in FIG.
e apartment by bus and sh
e sent to New York by pla
ne ”corresponds to the one-dimensional sequence of this word.
(W1wTwowThreewFourwFivew6w7w8w9wTenw11w12
w13w 14wFifteen) One-dimensional one-to-one correspondence of word classes
The column is (C) in FIG.FiveC90C ThreeCtwenty oneC18C101C32C
TwoCFiveC90CThreeC63C28C101C32What is given in)
And

【0046】この単語クラスの一次元列(C5 903
2118101 322 5 90 3 6328101
32)において、隣接する2つの単語クラス(Ci
j )の相互情報量MI(Ci 、Cj )を計算し、相互
情報量MI(C63、C28)が所定のしきい値TH以上、
相互情報量MI(C5 、C90)、MI(C90、C3 )、
MI(C3 、C21)、MI(C21、C18)、MI
(C18、C101 )、MI(C101、C32)、MI
(C32、C2 )、MI(C2 、C5 )、MI(C5 、C
90)、MI(C90、C3 )、MI(C3 、C63)、MI
(C28、C101)及びMI(C101 、C32)が所定のしき
い値THより小さい場合、隣接する2つの単語クラス
(C63、C28)が、図9(d)に示すように、クラスチ
ェーンで結ばれる。
The one-dimensional sequence (CFiveC90CThree
Ctwenty oneC18C101C32CTwoCFiveC90C ThreeC63C28C101C
32), Two adjacent word classes (Ci,
Cj) Mutual information MI (Ci, Cj) Calculate and mutual
Information amount MI (C63, C28) Is greater than or equal to a predetermined threshold TH,
Mutual information MI (CFive, C90), MI (C90, CThree),
MI (CThree, Ctwenty one), MI (Ctwenty one, C18), MI
(C18, C101), MI (C101, C32), MI
(C32, CTwo), MI (CTwo, CFive), MI (CFive, C
90), MI (C90, CThree), MI (CThree, C63), MI
(C28, C101) And MI (C101 , C32) Is the prescribed threshold
If the value is less than TH, two adjacent word classes
(C63, C28), As shown in FIG.
Are tied together.

【0047】このクラスチェーンで結ばれた2つの単語
クラス(C63、C28)はトークンt 1 に置き換えられ、
図9(e)に示すように、単語・トークンの一次元列
(w12 3 4 5 6 7 8 9 10111
1415)が生成される。
Two words connected by this class chain
Class (C63, C28) Is the token t 1Is replaced by
As shown in FIG. 9E, a one-dimensional sequence of words and tokens
(W1wTwowThreewFourwFivew6w7w8w9wTenw11t1
w14wFifteen) Is generated.

【0048】図1の単語・トークン分類手段6は、テキ
ストデータの単語・トークンの一次元列のN個の単語の
集合{w1 、w2 、w3 、w4 、・・・、wN }又はL
個のトークンの集合{t1 、t2 、t3 、t4 、・・
・、tL }を分割することにより、単語とトークンとが
混在して存在するD個の単語・トークンクラス{T1
2 、T3 、T4 、・・・、TD }を生成する。
The word / token classification means 6 in FIG. 1 performs a set of N words {w 1 , w 2 , w 3 , w 4 ,. } Or L
Set of tokens {t 1 , t 2 , t 3 , t 4 ,.
, T L }, so that D words and tokens {T 1 ,
T 2 , T 3 , T 4 ,..., T D } are generated.

【0049】この単語・トークン分類手段6では、トー
クンを付与された単語クラス列が1つの単語のようにみ
なされ、テキストデータに含まれる単語{w1 、w2
3、w4 、・・・、wN }とトークン{t1 、t2
3 、t4 、・・・、tL }とを同等に取り扱うことが
できるので、単語{w1 、w2 、w3 、w4 、・・・、
N }とトークン{t1 、t2 、t3 、t4 、・・・、
L }との区別なく分類処理を行うことができる図10
は、図1の単語・トークン分類手段6の機能的な構成を
示すブロック図である。
In the word / token classification means 6, the word class sequence to which the token has been assigned is regarded as one word, and the words {w 1 , w 2 ,
w 3 , w 4 ,..., w N } and tokens {t 1 , t 2 ,
Since t 3 , t 4 ,..., t L } can be treated equivalently, the words {w 1 , w 2 , w 3 , w 4 ,.
w N } and tokens {t 1 , t 2 , t 3 , t 4 ,.
FIG. 10 that can perform classification processing without distinction from t L
FIG. 3 is a block diagram showing a functional configuration of the word / token classification means 6 of FIG.

【0050】図10において、初期化クラス設定部40
は、テキストデータの単語・トークン列から互いに異な
る単語と互いに異なるトークンとを抽出し、所定の出現
頻度を有するN個の単語{w1 、w2 、w3 、w4 、・
・・、wN }とL個のトークン{t1 、t2 、t3 、t
4 、・・・、tL }とのそれぞれに固有の単語・トーク
ンクラス{T1 、T2 、T3 、T4 、・・・、TY }を
割り当てる。
In FIG. 10, an initialization class setting section 40
Extracts different words and different tokens from the word / token sequence of the text data, and generates N words {w 1 , w 2 , w 3 , w 4 ,.
.., w N } and L tokens {t 1 , t 2 , t 3 , t
4, allocates ..., each unique word token class {T 1 and t L}, T 2, T 3, T 4, ···, the T Y}.

【0051】仮マージ部41は、単語・トークンクラス
の集合{T1 、T2 、T3 、T4 、・・・、TM }から
2つの単語・トークンクラス{Ti 、Tj }を取り出し
て仮マージする。
The temporary merging unit 41 generates two word / token classes {T i , T j } from the set of word / token classes {T 1 , T 2 , T 3 , T 4 ,..., T M }. Take out and temporarily merge.

【0052】平均相互情報量算出部42は、テキストデ
ータの仮マージされた単語・トークンクラス{T1 、T
2 、T3 、T4 、・・・、TM-1 }についての平均相互
情報量AMIを(1)式により算出する。この場合、M
個の単語クラス・トークンクラスの集合{T1 、T2
3 、T4 、・・・、TM }から、2つの単語・トーク
ンクラス{Ti 、Tj }を取り出だす取り出しかたは、
M(M−1)/2個だけ存在するので、M(M−1)/
2回の平均相互情報量AMIの計算を行う必要がある。
The average mutual information calculation section 42 calculates the word / token class {T 1 , T
The average mutual information AMI for 2 , T 3 , T 4 ,..., T M-1 } is calculated by equation (1). In this case, M
Set of word classes and token classes {T 1 , T 2 ,
From T 3 , T 4 ,..., T M }, two words / token classes {T i , T j } are extracted.
Since only M (M-1) / 2 exist, M (M-1) /
It is necessary to calculate the average mutual information AMI twice.

【0053】本マージ部43は、仮マージにより計算さ
れたM(M−1)/2個の平均相互情報量AMIの基づ
いて、平均相互情報量AMIを最大とする2つの単語・
トークンクラス{Ti 、Tj }を単語クラス・トークン
クラスの集合{T1 、T2 、T3 、T4 、・・・、
M }から取り出して本マージする。このことにより、
本マージされたいずれかの単語・トークンクラス
{Ti 、Tj }に属する単語及びトークンは、同一の単
語クラス・トークンクラスに分類される。
Based on the M (M−1) / 2 average mutual information AMI calculated by the temporary merge, the main merge unit 43 generates two words / maximum average mutual information AMI.
A token class {T i , T j } is a set of word classes and token classes {T 1 , T 2 , T 3 , T 4 ,.
Take out from T Mす る and perform full merging. This allows
The words and tokens belonging to any of the merged word / token classes {T i , T j } are classified into the same word class / token class.

【0054】図1の連語置換手段7は、単語・トークン
クラスの中のトークンを、単語・トークン列生成手段5
により置換された単語列に逆置換して連語を生成する。
図11は、クラスチェーンと連語との関係を説明する図
である。
The collocation unit 7 in FIG. 1 converts the tokens in the word / token class into word / token sequence generation units 5
Is reversely replaced with the word string replaced by, to generate a collocation.
FIG. 11 is a diagram for explaining the relationship between class chains and collocations.

【0055】図11において、例えば、単語クラスC
300 と単語クラスC32とがクラスチェーンで結ばれ、こ
のクラスチェーンで結ばれた単語クラス列C300 −C32
にトークンt5 が付与されているとする。また、単語
“Toyota”、“Nissan”、“GM”・・・
などのA個の単語が単語クラスC300 に属し、単語“c
ar”、“track”、“wagon”・・・などの
B個の単語が単語クラスC 32に属しているものとする。
In FIG. 11, for example, word class C
300And word class C32Are connected by a class chain,
Class string C connected by the class chain300-C32
To the token tFiveIs given. Also the word
"Toyota", "Nissan", "GM" ...
A words like word class C300Belong to the word "c"
ar "," track "," wagon "...
B words are in word class C 32It belongs to

【0056】この場合、連語の候補として、図11
(b)に示すように、“Toyotacar”、“To
yota track”、“Toyota wago
n”、“Nissan car”、“Nissan t
rack”、“Nissanwagon”、“GM c
ar”、“GM track”、“GM wago
n”、・・・など、単語クラスC300 に属するA個の単
語と単語クラスC32に属するB個の単語との順列の数A
×Bだけ連語の候補が生成される。この連語の候補の中
にはテキストデータに存在しない連語も含まれているの
で、テキストデータをスキャンすることにより、これら
の連語の候補からテキストデータに存在する連語のみを
抽出する。例えば、テキストデータには、“Nissa
n track”及び“Toyota wagon”は
存在するが、“Toyota car”、“Toyot
a track”、 “Nissan car”、“N
issan wagon”、“GM car”、“GM
track”及び“GM wagon”は存在しない
場合、図11(c)に示すように、“Nissan t
rack”及び“Toyota wagon”のみが連
語としてテキストデータから抽出される。
In this case, as a candidate for the collocation,
As shown in (b), “Toyotacar”, “Toyotacar”
yota track ”,“ Toyota wago ”
n "," Nissant car "," Nissant t "
rack "," Nissanwagon "," GM c
ar "," GM track "," GM wago "
n ”,..., the number of permutations of A words belonging to word class C 300 and B words belonging to word class C 32
Only × B of collocation candidates are generated. Since the collocation candidates include collocations that do not exist in the text data, the text data is scanned to extract only the collocations present in the text data from these collocation candidates. For example, text data includes “Nissa
“n track” and “Toyota wagon” exist, but “Toyota car”, “Toyota wagon”
a track "," Nissan car "," N
issan wagon ”,“ GM car ”,“ GM
When “track” and “GM wagon” do not exist, as shown in FIG. 11C, “Nissant t”
Only “track” and “Toyota wagon” are extracted from the text data as collocations.

【0057】図12は、C個の単語クラス{C1
2 、C3 、C4 、・・・、CC }、D個の単語・トー
クンクラス{T1 、T2 、T3 、T4 、・・・、TD
及びD個の単語・連語クラス{R1 、R2 、R3
4 、・・・、RD }の一例を示す図である。
FIG. 12 shows C word classes {C 1 ,
C 2, C 3, C 4 , ···, C C}, D number word token classes {T 1, T 2, T 3, T 4, ···, T D}
And D word / syllable classes {R 1 , R 2 , R 3 ,
R 4, ···, is a diagram showing an example of R D}.

【0058】図12(a)において、C個の単語クラス
{C1 、C2 、C3 、C4 、・・・、CC }が、図1の
単語分類手段1により生成され、例えば、“he”、
“she”、“it”・・・などの単語が単語クラスC
5 に属し、“York”、“London”・・・など
の単語が単語クラスC28に属し、“car”、“tra
ck”、“wagon”・・・などの単語が単語クラス
32に属し、“new”、“old”・・・などの単語
が単語クラスC63に属し、“Toyota”、“Nis
san”、“GM”・・・などの単語が単語クラスC
300 に属しているものとする。また、テキストデータに
は、“New York”、“Nissantrac
k”及び“Toyota wagon”の連語が多数存
在しているものとする。
In FIG. 12A, C word classes {C 1 , C 2 , C 3 , C 4 ,..., C C } are generated by the word classifying means 1 of FIG. "He",
Words such as “she”, “it”,.
Belongs to 5, "York", "London " words such as ... belong to the word class C 28, "car", " tra
ck "," wagon "words such as ... belong to the word class C 32," new "," old " words such as ... belong to the word class C 63," Toyota "," Nis
words such as “san”, “GM”,.
It belongs to 300 . The text data includes “New York” and “Nissantrac”.
It is assumed that a number of collocations of "k" and "Toyota wagon" exist.

【0059】このC個の単語クラス{C1 、C2
3 、C4 、・・・、CC }をテキストデータの単語の
一次元列(w1 2 3 4 ・・・wT )に1対1対応
でマッピングした単語クラスの一次元列において、図1
の単語クラス列抽出手段3は、“new”が属する単語
クラスC63と“York”が属する単語クラスC28との
粘着度が大きいと判断し、単語クラスC63と単語クラス
28とをクラスチェーンで結ぶ。また、単語クラス列抽
出手段3は、“Toyota”及び“Nissan”が
属する単語クラスC300 と“track”及び“wag
on”が属する単語クラスC32との粘着度が大きいと判
断し、単語クラスC300 と単語クラスC32とをクラスチ
ェーンで結ぶ。
The C word classes {C 1 , C 2 ,
C 3, C 4, ···, C C} one-dimensional string of words in the text data (w 1 w 2 w 3 w 4 ··· w T) in a one-dimensional word classes mapped one-to-one correspondence Figure 1
Determines that the degree of adhesion between the word class C 63 to which “new” belongs and the word class C 28 to which “York” belongs is large, and classifies the word class C 63 and the word class C 28 into classes. Connect with a chain. In addition, the word class string extracting means 3 determines that the word class C 300 to which “Toyota” and “Nissan” belong and “track” and “wag”
It determines that the tackiness of the word class C 32 which on "belongs is large, connecting the word class C 300 and the word class C 32 in the class chain.

【0060】トークン付与手段4は、単語クラス列C63
−C28にトークンt1 を付与し、単語クラス列C300
32にトークンt5 を付与する。単語・トークン列生成
手段5は、テキストデータの単語の一次元列(w1 2
3 4 ・・・wT )に存在する“New York”
をトークンt1 で置き換え、テキストデータの単語の一
次元列(w1 2 3 4 ・・・wT )に存在する“N
issan track”及び“Toyota wag
on”をトークンt5 で置き換えた単語・トークンの一
次元列を生成する。
The token assigning means 4 generates the word class string C 63
The token t 1 is given to -C 28, word class sequence C 300 -
The token t 5 given to the C 32. The word / token string generation means 5 generates a one-dimensional string (w 1 w 2
w 3 w 4 ··· w T) to present "New York"
The replaced by a token t 1, exist in a one-dimensional string of words of text data (w 1 w 2 w 3 w 4 ··· w T) "N
issan track ”and“ Toyota wag ”
to generate a one-dimensional string of words or token by replacing the on "in the token t 5.

【0061】単語・トークン分類手段6は、この単語・
トークンの一次元列に存在する“he”、“she”、
“it”、“London”、“car”、“trac
k”、“wagon”・・・などの単語及び“t1 ”、
“t5 ”などのトークンについての分類処理を行い、図
12(b)のD個の単語・トークンクラス{T1
2 、T3 、T4 、・・・、TD }を生成する。
The word / token classification means 6
"He", "she",
“It”, “London”, “car”, “trac”
k ”,“ wagon ”... and“ t ”1”,
"TFiveClassification processing for tokens such as "
12 (b) D words / token class @T1,
T Two, TThree, TFour, ..., TDGenerate}.

【0062】単語・トークンクラス{T1 、T2
3 、T4 、・・・、TD }において、例えば、“h
e”、“she”、“it”・・・などの単語やトーク
ンが単語・トークンクラスT5 に属し、“t1 ”、“L
ondon”・・・などの単語やトークンが単語・トー
クンクラスT28に属し、“car”、“track”、
“wagon”、“t5 ”・・・などの単語やトークン
が単語・トークンクラスT32に属し、“new”、“o
ld”・・・などの単語やトークンが単語・トークンク
ラスT63に属し、“Toyota”、“Nissa
n”、“GM”・・・などの単語やトークンが単語・ト
ークンクラスT300 に属している。このように、単語・
トークンクラス{T1 、T2 、T3 、T4 、・・・、T
D }には、単語とトークンとの区別なく、単語とトーク
ンとが混在して分類されている。
The word / token class {T 1 , T 2 ,
In T 3 , T 4 ,..., T D }, for example, “h
e "," she "," it " word or token, such as ... belong to the word token class T 5," t 1 ", " L
ondon "word or token, such as ... belong to the word token class T 28," car "," track ",
Words or tokens such as “wagon”, “t 5 ”... Belong to the word / token class T 32 , and “new”, “o”
ld "word or token, such as ... belong to the word token class T 63," Toyota "," Nissa
n "," GM "word or token, such as ... are belong to the word token class T 300. In this way, the words,
Token class {T 1 , T 2 , T 3 , T 4 ,..., T
In D単 語, words and tokens are classified in a mixed manner without distinction between words and tokens.

【0063】連語置換手段7は、図12(b)の単語・
トークンクラス{T1 、T2 、T3、T4 、・・・、T
D }に存在する“t1 ”、“t5 ”などのトークンを、
テキストデータの単語の一次元列に存在する連語で逆置
換することにより、図12(c)の単語・連語クラス
{R1 、R2 、R3 、R4 、・・・、RD }を生成す
る。例えば、単語・トークンクラスT28に属しているト
ークンt1 は、 単語・トークン列生成手段5により、
テキストデータの単語の一次元列に存在する“New
York”と置換されたものなので、このトークンt1
を“New York”で逆置換することにより、単語
・連語クラスR28を生成し、単語・トークンクラスT32
に属しているトークンt5 は、単語・トークン列生成手
段5により、テキストデータの単語の一次元列に存在す
る“Nissan track”及び“Toyota
wagon”と置換されたものなので、このトークンt
5 を“Nissan track”及び“Toyota
wagon”で逆置換することにより、単語・連語ク
ラスR32を生成する。
The collocation unit 7 converts the word / word of FIG.
Token class {T 1 , T 2 , T 3 , T 4 ,..., T
Tokens such as “t 1 ” and “t 5 ” existing in D
By reversely permuting words and collocation classes {R 1 , R 2 , R 3 , R 4 ,..., R D } in FIG. Generate. For example, the token t 1 that belongs to the word token class T 28 is the word token string generating means 5,
"New" existing in a one-dimensional string of words of text data
This token t 1
Is replaced with “New York” to generate a word / collocation class R 28 and a word / token class T 32
The token t 5 that belong to, by the word token string generating means 5, exist in a one-dimensional string of words in the text data "Nissan track" and "Toyota
wagon ", the token t
5 for “Nissan track” and “Toyota
By inverse replaced by wagon in ", to generate the word-phrase class R 32.

【0064】図13は、図1の単語・連語分類処理装置
を実現するシステム構成を示すブロック図である。図1
3において、単語・連語分類処理部41のメモリインタ
ーフェース42、46、CPU43、ROM44、ワー
クRAM45、RAM47、ドライバ71及び通信イン
タフェース72はバス48を介して互いに接続され、テ
キストデータ40が単語・連語分類処理部41に入力さ
れると、ROM44に格納されているプログラムに従っ
て、CPU43はテキストデータ40を処理し、テキス
トデータ40の単語及び連語の分類処理を行う。テキス
トデータ40の単語及び連語の分類処理結果は、単語・
連語辞書49に格納される。なお、テキストデータ40
や単語及び連語の分類処理結果を通信インタフェース7
2から通信ネットワーク73を介して送信したり、受信
したりすることも可能である。
FIG. 13 is a block diagram showing a system configuration for realizing the word / phrase classification processing apparatus of FIG. FIG.
3, the memory interfaces 42, 46, the CPU 43, the ROM 44, the work RAM 45, the RAM 47, the driver 71, and the communication interface 72 of the word / phrase classification processing unit 41 are connected to each other via a bus 48, and the text data 40 When input to the processing unit 41, the CPU 43 processes the text data 40 according to a program stored in the ROM 44, and performs a classification process of words and collocations of the text data 40. The classification processing result of the words and collocations in the text data 40 is
It is stored in the collocation dictionary 49. The text data 40
Interface processing results of classification of words and words and collocations
2 can also be transmitted and received via the communication network 73.

【0065】また、単語及び連語の分類処理を行うプロ
グラムを、ハードディスク74、ICメモリカード7
5、磁気テープ76、フロッピーディスク77またはC
D−ROMやDVD−ROMなどの光ディスク78によ
る記憶媒体からRAM47にロードした後、このプログ
ラムをCPU43で実行させるようにしてもよい。
A program for classifying words and collocations is stored in the hard disk 74 and the IC memory card 7.
5. Magnetic tape 76, floppy disk 77 or C
The program may be loaded into the RAM 47 from a storage medium such as a D-ROM or a DVD-ROM, such as an optical disk 78, and then executed by the CPU 43.

【0066】さらに、単語及び連語の分類処理を行うプ
ログラムを、通信インタフェース72を介して通信ネッ
トワーク73から取り出すこともできる。通信インタフ
ェース72と接続される通信ネットワーク73として、
例えば、LAN(LocalArea Networ
k)、WAN(Wide Area Networ
k)、インターネット、アナログ電話網、デジタル電話
網(ISDN:Integral Service D
igital Network)、PHS(パーソナル
ハンディシステム)や衛星通信などの無線通信網などを
用いることが可能である。
Further, a program for classifying words and collocations can be extracted from the communication network 73 via the communication interface 72. As a communication network 73 connected to the communication interface 72,
For example, LAN (Local Area Network)
k), WAN (Wide Area Network)
k), the Internet, an analog telephone network, a digital telephone network (ISDN: Integral Service D)
It is possible to use a wireless communication network such as digital network (PTE), PHS (Personal Handy System) and satellite communication.

【0067】図14は、図1の単語・連語分類処理装置
の動作を示すフローチャートである。図14において、
まず、ステップS1に示すように、単語クラスタリング
処理を行う。この単語クラスタリング処理では、複数の
単語の一次元列(w1 2 3 4 ・・・wT )として
のテキストデータから、互いに異なるV個の単語
{v 1 、v2 、v3 、v4 、・・・、vV }を抽出し、
V個の単語の集合{v1 、v 2 、v3 、v4 、・・・、
V }をC個の単語クラス{C1 、C2 、C3 、C4
・・・、CC }に分割する第1のクラスタリング処理を
行う。
FIG. 14 shows the word / phrase classification processing apparatus of FIG.
6 is a flowchart showing the operation of the first embodiment. In FIG.
First, as shown in step S1, word clustering
Perform processing. In this word clustering process,
One-dimensional sequence of words (w1wTwow ThreewFour... wTAs)
V words that are different from each other
{V 1, VTwo, VThree, VFour, ..., vVExtract},
Set of V words {v1, V Two, VThree, VFour, ...,
vVLet} be C word classes {C1, CTwo, CThree, CFour,
..., CCThe first clustering process of dividing into}
Do.

【0068】ここで、V個の単語{v1 、v2 、v3
4 、・・・、vV }それぞれに単語クラス{C1 、C
2 、C3 、C4 、・・・、CV }を割り当ててから、V
個の単語クラス{C1 、C2 、C3 、C4 、・・・、C
V }についてマージ処理を行うことにより、V個の単語
クラス{C1 、C2 、C3 、C4 、・・・、CV }の個
数を1つずつ減らしてC個の単語クラス{C1 、C2
3 、C4 、・・・、CC }を生成する場合、Vが70
00もの数となって大きなものとなるときは、マージ処
理を行うための(1)式の平均相互情報量AMIの計算
回数が莫大なものとなり、現実的ではなくなる。このた
め、ウィンドウ処理を行って、マージ処理を行う単語ク
ラスの数を減らすようにする。
Here, V words {v 1 , v 2 , v 3 ,
v 4 ,..., v V }, respectively, the word class {C 1 , C
2 , C 3 , C 4 ,..., C V }
Word classes {C 1 , C 2 , C 3 , C 4 ,..., C
V }, the number of V word classes {C 1 , C 2 , C 3 , C 4 ,..., C V } is reduced by one, and C word classes {C 1, C 2,
When generating C 3 , C 4 ,..., C C }, V is 70
When the number is 00 and becomes large, the number of calculations of the average mutual information AMI of the equation (1) for performing the merge processing becomes enormous, which is not realistic. Therefore, window processing is performed to reduce the number of word classes to be merged.

【0069】図15は、ウィンドウ処理を説明する図で
ある。図15(a)において、テキストデータのV個の
単語{v1 、v2 、v3 、v 4 、・・・、vV }それぞ
れに割り当てられたV個の単語クラス{C1 、C2 、C
3 、C4 、・・・、CV }のうち、テキストデータでの
出現頻度の大きい単語に割り当てられたC+1個の単語
クラス{C1 、C2 、C3 、C4 、・・・、C C 、C
C+1 }を取り出し、このC+1個の単語クラス{C1
2 、C3 、C4、・・・、CC 、CC+1 }についての
マージ処理を行う。
FIG. 15 is a diagram for explaining window processing.
is there. In FIG. 15A, V data of text data
The word $ v1, VTwo, VThree, V Four, ..., vV} Each
V word classes assigned to them1, CTwo, C
Three, CFour, ..., CV} Of the text data
C + 1 words assigned to words with high appearance frequency
Class II C1, CTwo, CThree, CFour, ..., C C, C
C + 1、, and C + 1 word classes {C1,
CTwo, CThree, CFour, ..., CC, CC + 1}about
Perform merge processing.

【0070】ここで、図15(b)に示すように、M個
の単語クラス{C1 、C2 、C3 、C4 、・・・、
M }は、ウィンドウ内のC+1個の単語クラス
{C1 、C2 、C3 、C4 、・・・、CC 、CC+1 }に
ついてのマージ処理を行った場合、M個の単語クラス
{C1 、C2 、C3 、C4 、・・・、CM }の数が1つ
減ってM−1個の単語クラス{C1 、C2 、C3
4 、・・・、CM-1 }となるとともに、ウィンドウ内
のC+1個の単語クラス{C1 、C2 、C3 、C4 、・
・・、C C 、CC+1 }の数も1つ減ってC個の単語クラ
ス{C1 、C2 、C3 、C4 、・・・、CC }となる。
Here, as shown in FIG.
Word class {C1, CTwo, CThree, CFour, ...,
CM} Is C + 1 word classes in the window
{C1, CTwo, CThree, CFour, ..., CC, CC + 1Puni
When the merge process is performed, M word classes
{C1, CTwo, CThree, CFour, ..., CMOne 数
Reduced M-1 word classes {C1, CTwo, CThree,
CFour, ..., CM-1} And in the window
C + 1 word classes {C1, CTwo, CThree, CFour,
・ ・ 、 C C, CC + 1The number of} is also reduced by one to C word classes.
{C1, CTwo, CThree, CFour, ..., CCIt becomes}.

【0071】この場合、図15(c)に示すように、ウ
ィンドウ外の単語クラス{CC+1 、・・・、CM-1 }の
うち、テキストデータでの出現頻度が最も大きい単語ク
ラスCC+1 をウィンドウ内に入れ、ウィンドウ内の単語
クラスの数が一定に保たれるようにする。
[0071] In this case, as shown in FIG. 15 (c), the window outside the word class {C C + 1, ···, C M-1} of the largest word classes the frequency of occurrence of a text data Put CC + 1 in the window so that the number of word classes in the window is kept constant.

【0072】そして、ウィンドウ外に単語クラスがなく
なり、図15(d)のC個の単語クラス{C1 、C2
3 、C4 、・・・、CC }が生成された時に、単語ク
ラスタリング処理を終了する。
Then, there are no word classes outside the window, and the C word classes {C 1 , C 2 ,
When C 3 , C 4 ,..., C C } are generated, the word clustering process ends.

【0073】なお、上述した実施例では、ウィンドウ内
の単語クラスの個数をC+1個に設定したが、C+1個
以外のV個未満の数でもよく、また、途中で変化させる
ようにしてもよい。
In the above embodiment, the number of word classes in the window is set to C + 1. However, the number of word classes may be less than V other than C + 1, or may be changed in the middle.

【0074】図16は、ステップS1の単語クラスタリ
ング処理を示すフローチャートである。図16におい
て、まず、ステップS10に示すように、T個の単語の
一次元列(w1 2 3 4 ・・・wT )としてのテキ
ストデータに基づいて、重複を除いた全てのV個の単語
{v1 、v2 、v3 、v4 、・・・、vV }の出現頻度
を調べ、これらのV個の単語{v1 、v2 、v3
4 、・・・、vV }を出現頻度の高い単語から順に並
べて、これらのV個の単語{v1 、v2 、v3 、v4
・・・、vV }のそれぞれをV個の単語クラス{C1
2 、C3 、C4 、・・・、CV }に割り当てる。
FIG. 16 is a flowchart showing the word clustering process in step S1. 16, first, as shown in step S10, based on the text data as the T dimensional row of words (w 1 w 2 w 3 w 4 ··· w T), all except the overlapping The frequency of appearance of the V words {v 1 , v 2 , v 3 , v 4 ,..., V V } is checked, and these V words {v 1 , v 2 , v 3 ,
v 4 ,..., v V } are arranged in descending order of the frequency of occurrence, and these V words {v 1 , v 2 , v 3 , v 4 ,.
, V V } are represented by V word classes {C 1 ,
C 2 , C 3 , C 4 ,..., C V }.

【0075】次に、ステップS11に示すように、V個
の単語クラス{C1 、C2 、C3 、C4 、・・・、
V }の単語のうち、出現頻度の高い単語クラスの単語
から、V個未満のC+1個の単語クラスの単語を1つの
ウィンドウ内の単語クラスの単語とする。
Next, as shown in step S11, V word classes {C 1 , C 2 , C 3 , C 4 ,.
Of the words of C V }, less than V words of the C + 1 word class from words of the word class having a high frequency of occurrence are defined as words of the word class in one window.

【0076】次に、ステップS12に示すように、1つ
のウィンドウ内の単語クラスの単語の中で、全ての組み
合わせの仮ペアを作り、各仮ペアを仮マージした時の平
均相互情報量AMIを(1)式により計算する。
Next, as shown in step S12, among the words of the word class in one window, provisional pairs of all combinations are created, and the average mutual information AMI when the provisional pairs are provisionally merged is calculated. It is calculated by equation (1).

【0077】次に、ステップS13に示すように、全て
の組み合わせの仮ペアについての平均相互情報量AMI
のうち、最大となる平均相互情報量AMIを有する仮ペ
アを本マージすることにより、単語クラスを1つだけ減
らし、本マージ後の1つのウィンドウ内の単語クラスの
単語を更新する。
Next, as shown in step S13, the average mutual information AMI for the temporary pairs of all the combinations is determined.
Among them, the temporary pair having the maximum average mutual information amount AMI is main-merged to reduce the word class by one, and the word of the word class in one window after the main merge is updated.

【0078】次に、ステップS14に示すように、ウィ
ンドウ外の単語クラスはなくなり、かつ、ウィンドウ内
の単語クラスはC個になったかどうかを判断し、この条
件が成り立たない場合、ステップS15に進み、現在の
ウィンドウよりも外側にあり、最大の出現頻度を有する
クラスの単語をウィンドウ内に入れ、ステップS12に
戻り、以上の処理を繰り返すことにより、単語クラスの
数を減少させる。
Next, as shown in step S14, it is determined whether there are no more word classes outside the window and there are C word classes in the window. If this condition is not satisfied, the flow advances to step S15. The words of the class having the maximum appearance frequency outside the current window are put in the window, and the process returns to step S12 to repeat the above processing to reduce the number of word classes.

【0079】一方、ステップS14の条件が成り立ち、
ウィンドウ外に単語クラスがなくなり、単語クラスの数
がC個となった場合、ステップS16に進み、ウィンド
ウ内のC個の単語クラス{C1 、C2 、C3 、C4 、・
・・、CC }をメモリに記憶する。
On the other hand, the condition of step S14 holds,
If there are no word classes outside the window and the number of word classes becomes C, the process proceeds to step S16, and the C word classes {C 1 , C 2 , C 3 , C 4 ,.
.., C C } are stored in the memory.

【0080】次に、図14のステップS2に示すよう
に、クラスチェーン抽出処理を行う。このクラスチェー
ン抽出処理では、ステップS1の第1のクラスタリング
処理に基づいて生成されたテキストデータの単語クラス
の一次元列において、所定のしきい値以上の相互情報量
を有する隣接する2つの単語クラスをチェーンで結ぶこ
とにより、チェーンで結ばれた単語クラス列の集合を抽
出する。
Next, as shown in step S2 of FIG. 14, a class chain extraction process is performed. In this class chain extraction processing, in a one-dimensional sequence of word classes of text data generated based on the first clustering processing in step S1, two adjacent word classes having a mutual information amount equal to or more than a predetermined threshold value are included. Are connected by a chain to extract a set of word class strings connected by the chain.

【0081】図17は、ステップS2のクラスチェーン
抽出処理の第1実施例を示すフローチャートである。図
17において、まず、ステップS20に示すように、テ
キストデータの単語クラスの一次元列から、互いに隣接
する2つの単語クラス(Ci 、Cj )を取り出す。
FIG. 17 is a flowchart showing a first embodiment of the class chain extracting process in step S2. In FIG. 17, first, as shown in step S20, two adjacent word classes (C i , C j ) are extracted from the one-dimensional sequence of the word classes of the text data.

【0082】次に、ステップS21に示すように、ステ
ップS20で取り出した2つの単語クラス(Ci
j )についての相互情報量MI(Ci 、Cj )を
(2)式により計算する。
Next, as shown in step S21, the two word classes (C i ,
Mutual information MI (C i for C j), is calculated by the C j) (2) expression.

【0083】次に、ステップS22に示すように、ステ
ップS21で計算した相互情報量MI(Ci 、Cj )が
所定のしきい値TH以上であるかどうかを判断し、相互
情報量MI(Ci 、Cj )が所定のしきい値TH以上で
ある場合、ステップS23に進んで、ステップS20で
取り出した2つの単語クラス(Ci 、Cj )をクラスチ
ェーンで結んでメモリに格納し、相互情報量MI
(Ci 、Cj )が所定のしきい値THより小さい場合、
ステップS23をスキップする。
Next, as shown in step S22, it is determined whether or not the mutual information amount MI (C i , C j ) calculated in step S21 is equal to or greater than a predetermined threshold value TH, and the mutual information amount MI (C i , C j ) is determined. If C i , C j ) is equal to or greater than the predetermined threshold value TH, the process proceeds to step S23, where the two word classes (C i , C j ) extracted in step S20 are connected in a class chain and stored in a memory. , Mutual information MI
When (C i , C j ) is smaller than a predetermined threshold TH,
Step S23 is skipped.

【0084】次に、ステップS24に示すように、メモ
リに格納されているクラスチェーンで結ばれた単語クラ
スにおいて、単語クラスCi で終了しているクラスチェ
ーンが存在するかどうかを判断し、単語クラスCi で終
了しているクラスチェーンが存在する場合、ステップS
25に進んで、単語クラスCi で終了しているクラスチ
ェーンに単語クラスCj をつなぐ。
[0084] Next, as shown in step S24, the word class connected by class chain stored in the memory, to determine whether the class chain ending with the word classes C i is present, the words If the class chain ending with the class C i is present, step S
Proceed to 25, connecting the word class C j in class chain that ends with the word class C i.

【0085】一方、ステップS24において、単語クラ
スCi で終了しているクラスチェーンが存在しない場
合、ステップS25をスキップする。次に、ステップS
26に示すように、テキストデータの単語クラスの一次
元列から、互いに隣接する2つの単語クラス(Ci 、C
j )を全て取り出したかどうかを判断し、互いに隣接す
る2つの単語クラス(Ci 、Cj )を全て取り出した場
合、クラスチェーン抽出処理を終了し、互いに隣接する
2つの単語クラス(C i 、Cj )を全て取り出していな
い場合、ステップS20に戻って以上の処理を繰り返
す。
On the other hand, in step S24, the word
SiIf there is no class chain ending with
In this case, step S25 is skipped. Next, step S
As shown in FIG. 26, the primary
From the source column, two adjacent word classes (Ci, C
jJudge whether all of them have been taken out and
Two word classes (Ci, CjPlace where all of them were taken out
End the class chain extraction process and
Two word classes (C i, Cj) Has not been taken out
If not, return to step S20 and repeat the above processing.
You.

【0086】図18は、ステップS2のクラスチェーン
抽出処理の第2実施例を示すフローチャートである。図
18において、まず、ステップS201に示すように、
テキストデータの単語クラスの一次元列から、互いに隣
接する2つの単語クラス(Ci 、Cj )を順次に取り出
す。そして、取り出した2つの単語クラス(Ci
j )について、相互情報量MI(Ci 、Cj )を
(2)式により計算することにより、長さ2の全てのク
ラスチェーンをテキストデータの単語クラスの一次元列
から抽出する。
FIG. 18 is a flowchart showing a second embodiment of the class chain extracting process in step S2. In FIG. 18, first, as shown in step S201,
Two adjacent word classes (C i , C j ) are sequentially extracted from the one-dimensional sequence of the word classes of the text data. Then, the extracted two word classes (C i ,
For C j ), the mutual information MI (C i , C j ) is calculated by equation (2) to extract all the class chains of length 2 from the one-dimensional sequence of the word classes of the text data.

【0087】次に、ステップS202に示すように、長
さ2の全てのクラスチェーンをそれぞれオブジェクトで
置き換える。ここで、オブジェクトは、上述したトーク
ンと同じものを表しているが、長さ2のクラスチェーン
に付与されたトークンを、特に、オブジェクトと呼ぶ。
Next, as shown in step S202, all class chains of length 2 are replaced with objects. Here, the object represents the same token as the above-mentioned token, but the token given to the class chain of length 2 is particularly called an object.

【0088】次に、ステップS203に示すように、テ
キストデータのクラスの一次元列に対し、ステップS2
02でオブジェクトが付与された長さ2のクラスチェー
ンをオブジェクトで置き換え、テキストデータのクラス
とオブジェクトの一次元列を生成する。
Next, as shown in step S203, the one-dimensional sequence of the text data class is processed in step S2.
02, the class chain of length 2 to which the object is assigned is replaced with the object, and a class of the text data and a one-dimensional sequence of the object are generated.

【0089】次に、ステップS204に示すように、テ
キストデータのクラスとオブジェクトの一次元列の中に
存在する1つのオブジェクトを1つのクラスとみなし、
2つのクラス(Ci 、Cj )についての相互情報量MI
(Ci 、Cj )を(2)式により計算する。すなわち、
テキストデータのクラスとオブジェクトの一次元列にお
いての相互情報量MI(Ci 、Cj )は、互いに隣接す
る1つのクラスと1つのクラスとの間で算出される場
合、互いに隣接する1つのクラスと1つのオブジェクト
(長さ2のクラスチェーン)との間で算出される場合、
及び互いに隣接する1つのオブジェクト(長さ2のクラ
スチェーン)と1つのオブジェクト(長さ2のクラスチ
ェーン)との間で算出される場合がある。
Next, as shown in step S204, one object existing in the one-dimensional sequence of the text data class and the object is regarded as one class,
Mutual information MI about two classes (C i , C j )
(C i , C j ) is calculated by equation (2). That is,
When the mutual information MI (C i , C j ) in the one-dimensional sequence of the class of the text data and the object is calculated between one adjacent class and one class, one adjacent class is calculated. And between one object (a class chain of length 2)
And one object (class chain of length 2) and one object (class chain of length 2) adjacent to each other.

【0090】次に、ステップS205に示すように、ス
テップS204で計算した相互情報量MI(Ci
j )が所定のしきい値TH以上であるかどうかを判断
し、相互情報量MI(Ci 、Cj )が所定のしきい値T
H以上である場合、ステップS26に進んで、ステップ
S204で取り出した互いに隣接する2つのクラス、又
は互いに隣接する1つのクラスと1つのオブジェクト、
又は互いに隣接する2つのオブジェクトをクラスチェー
ンで結び、相互情報量MI(Ci 、Cj )が所定のしき
い値THより小さい場合、ステップS206をスキップ
する。
Next, as shown in step S205, the mutual information MI (C i ,
C j ) is determined to be greater than or equal to a predetermined threshold TH, and the mutual information MI (C i , C j ) is determined to be equal to or smaller than the predetermined threshold T
If not less than H, the process proceeds to step S26, where two classes adjacent to each other extracted in step S204, or one class and one object adjacent to each other,
Alternatively, when two objects adjacent to each other are connected by a class chain and the mutual information amount MI (C i , C j ) is smaller than a predetermined threshold value TH, step S206 is skipped.

【0091】図19は、テキストデータのクラスとオブ
ジェクトの一次元列において抽出されたクラスチェーン
を示す図である。図19において、互いに隣接する1つ
のクラスと1つのクラスとの間でクラスチェーンが抽出
された場合、長さ2のクラスチェーン(オブジェクト)
が生成され、互いに隣接する1つのクラスと1つのオブ
ジェクトとの間でクラスチェーンが抽出された場合、長
さ3のクラスチェーンが生成され、互いに隣接する1つ
のオブジェクトと1つのオブジェクトとの間でクラスチ
ェーンが抽出された場合、長さ4のクラスチェーンが生
成される。
FIG. 19 is a diagram showing a class of text data and a class chain extracted in a one-dimensional sequence of objects. In FIG. 19, when a class chain is extracted between one class and one class adjacent to each other, a class chain (object) having a length of 2
Is generated, and a class chain is extracted between one class and one object adjacent to each other, a class chain of length 3 is generated, and a class chain between one adjacent object and one object is generated. When a class chain is extracted, a class chain having a length of 4 is generated.

【0092】次に、図18のステップS207に示すよ
うに、クラスチェーン抽出処理が所定の回数行われたか
どうかを判断し、所定の回数行われていない場合は、ス
テップS202に戻って以上の処理を繰り返す。
Next, as shown in step S207 of FIG. 18, it is determined whether or not the class chain extraction processing has been performed a predetermined number of times. If the predetermined number of times has not been performed, the flow returns to step S202 to perform the above processing. repeat.

【0093】このように、長さ2のクラスチェーンをオ
ブジェクトに置き換えて、相互情報量MI(Ci
j )を算出することを繰り返すことにより、任意の長
さのクラスチェーンを抽出することができる。
As described above, the class chain of length 2 is replaced with an object, and the mutual information MI (C i ,
By repeatedly calculating C j ), a class chain having an arbitrary length can be extracted.

【0094】次に、図14のステップS3に示すよう
に、トークン置換処理を行う。このトークン置換処理で
は、ステップS2のクラスチェーン抽出処理で抽出され
た単語クラス列に固有のトークンを対応させ、この単語
クラス列に属する単語列をテキストデータの単語の一次
元列から検索し、テキストデータの単語列を対応するト
ークンで置換することにより、テキストデータについて
の単語とトークンとの一次元列を生成する。
Next, as shown in step S3 of FIG. 14, a token replacement process is performed. In this token replacement processing, a unique token is made to correspond to the word class string extracted in the class chain extraction processing in step S2, and a word string belonging to this word class string is searched for from a one-dimensional string of words in the text data. By replacing the word string of the data with the corresponding token, a one-dimensional string of words and tokens for the text data is generated.

【0095】図20は、ステップS3のトークン置換処
理を示すフローチャートである。図20において、ま
ず、ステップS30に示すように、抽出されたクラスチ
ェーンを重複を除いて所定の規則でソートし、それぞれ
のクラスチェーンにトークンを対応させて、クラスチェ
ーンに名前を付ける。ここで、クラスチェーンのソート
は、例えば、ASCIIコード順で行う。
FIG. 20 is a flowchart showing the token replacement processing in step S3. In FIG. 20, first, as shown in step S30, the extracted class chains are sorted according to a predetermined rule except for duplication, tokens are made to correspond to the respective class chains, and names are given to the class chains. Here, the sorting of the class chains is performed, for example, in ASCII code order.

【0096】次に、ステップS31に示すように、トー
クンに対応させたクラスチェーンを1つ取り出す。次
に、ステップS32に示すように、テキストデータの単
語の一次元列の中にクラスチェーンで結ばれた単語クラ
ス列に属する単語列が存在するかどうかを判断し、クラ
スチェーンで結ばれた単語クラス列に属する単語列が存
在する場合、ステップS33に進み、テキストデータの
対応する単語列を1つのトークンで置き換え、クラスチ
ェーンで結ばれた単語クラス列に属する単語列がテキス
トデータの単語の一次元列の中に存在しなくなるまで以
上の処理を繰り返す。
Next, as shown in step S31, one class chain corresponding to the token is extracted. Next, as shown in step S32, it is determined whether or not there is a word string belonging to the word class string connected by the class chain in the one-dimensional string of words of the text data, and the word connected by the class chain is determined. If there is a word string belonging to the class string, the process proceeds to step S33, where the corresponding word string in the text data is replaced with one token, and the word string belonging to the word class string connected by the class chain is the primary word of the text data. The above processing is repeated until there is no longer existing in the original sequence.

【0097】一方、クラスチェーンで結ばれた単語クラ
ス列に属する単語列が存在しない場合、ステップS34
に進み、ステップS30でトークンに対応させた全ての
クラスチェーンについての連語・トークン置換処理が終
了したかどうかを判断し、全てのクラスチェーンについ
ての連語・トークン置換処理が終了してない場合、ステ
ップS31に戻って、新たなクラスチェーンを1つ取り
出して、以上の処理を繰り返す。
On the other hand, if there is no word string belonging to the word class string connected by the class chain, step S34
It is determined whether the collocation / token replacement process has been completed for all the class chains corresponding to the tokens in step S30. If the collocation / token replacement process has not been completed for all the class chains, the process proceeds to step S30. Returning to S31, one new class chain is taken out, and the above processing is repeated.

【0098】次に、図14のステップS4に示すよう
に、単語・トークンクラスタリング処理を行う。この単
語・トークンクラスタリング処理では、テキストデータ
についての単語とトークンとの一次元列において、互い
に異なる単語と互いに異なるトークンとを抽出し、単語
とトークンとが混在する集合を単語・トークンクラス
{T1 、T2 、T3 、T4 、・・・、TD }に分割する
第2のクラスタリング処理を行う。
Next, as shown in step S4 of FIG. 14, a word / token clustering process is performed. In the word / token clustering process, different words and different tokens are extracted from a one-dimensional sequence of words and tokens of text data, and a set in which words and tokens coexist is defined as a word / token class {T 1 , T 2 , T 3 , T 4 ,..., T D }.

【0099】図21は、ステップS4の単語・トークン
クラスタリング処理を示すフローチャートである。図2
1において、ステップS40に示すように、ステップS
3で得られたテキストデータの単語・トークンの一次元
列を入力データとして、ステップS1の第1の単語クラ
スタリング処理と同一の方法でクラスタリングを行うこ
とより、単語・トークンクラス{T1 、T2 、T3 、T
4 、・・・、TD }を生成する。この第2のクラスタリ
ング処理では、単語とトークンは区別せず、トークンは
1つの単語として扱われる。また、生成されたそれぞれ
の単語・トークンクラス{T 1 、T2 、T3 、T4 、・
・・、TD }は、その要素として単語とトークンを含ん
でいる。
FIG. 21 shows the word / token in step S4.
It is a flowchart which shows a clustering process. FIG.
In step S1, as shown in step S40,
One-dimensional words and tokens of the text data obtained in step 3
Using the sequence as input data, the first word class
Perform clustering in the same way as the stalling process.
And more words / token class @T1, TTwo, TThree, T
Four, ..., TDGenerate}. This second cluster
The tokening process does not distinguish between words and tokens,
Treated as one word. Also, each generated
Word / token class @T 1, TTwo, TThree, TFour,
・ ・ 、 TD含 ん contains words and tokens as its elements
In.

【0100】次に、図14のステップS5に示すよう
に、データ出力処理を行う。このデータ出力処理では、
テキストデータの単語の一次元列に存在する単語列のう
ち、トークンに対応するものを連語として抽出し、単語
・トークンクラス{T1 、T2、T3 、T4 、・・・、
D }の中のトークンを連語で置換することにより、単
語と連語とが混在する集合を単語・連語クラス{R1
2 、R3 、R4 、・・・、RD }に分割する第3のク
ラスタリング処理を行う。
Next, a data output process is performed as shown in step S5 of FIG. In this data output process,
Of the word strings existing in the one-dimensional string of words of the text data, those corresponding to the tokens are extracted as collocations, and the word / token class {T 1 , T 2 , T 3 , T 4 ,.
By replacing the tokens in T D } with collocations, a set in which words and collocations are mixed is converted into a word / composition class {R 1 ,
A third clustering process for dividing into R 2 , R 3 , R 4 ,..., R D } is performed.

【0101】図22は、ステップS5のデータ出力処理
を示すフローチャートである。図22において、まず、
ステップS50に示すように、1つの単語・トークンク
ラスTi から1つのトークンtK を取り出す。
FIG. 22 is a flowchart showing the data output processing in step S5. In FIG. 22, first,
As shown in step S50, one token t K is extracted from one word / token class T i .

【0102】次に、ステップS51に示すように、テキ
ストデータの単語の一次元列をスキャンし、ステップS
52において、ステップS50で取り出したトークンt
K に対応するクラスチェーンで結ばれた単語クラス列に
属する単語列が存在するかどうかを判断する。そして、
トークンtK に対応するクラスチェーンで結ばれた単語
クラス列に属する単語列がテキストデータの単語の一次
元列に存在する場合、ステップS53に進んで、この単
語列を連語とみなす処理を繰り返し、テキストデータの
単語の一次元列をスキャンすることにより得られたこれ
らの連語でトークンtK を置き換える。
Next, as shown in step S51, a one-dimensional sequence of words in the text data is scanned, and step S51 is executed.
In 52, the token t extracted in step S50
It is determined whether there is a word string belonging to the word class string connected by the class chain corresponding to K. And
When the word string belonging to the word class string connected by the class chain corresponding to the token t K exists in the one-dimensional string of the word of the text data, the process proceeds to step S53, and the process of regarding this word string as a collocation is repeated. replacing the token t K in these collocation obtained by scanning a one-dimensional string of words in the text data.

【0103】一方、トークンtK に対応するクラスチェ
ーンで結ばれた単語クラス列に属する単語列がテキスト
データの単語の一次元列に存在しない場合、ステップS
54に進んで、全てのトークンについて処理が終了した
かどうかを判断し、全てのトークンについて処理が終了
していない場合、ステップS50に進んで、以上の処理
を繰り返す。
On the other hand, if the word string belonging to the word class string connected by the class chain corresponding to the token t K does not exist in the one-dimensional string of the word of the text data, step S
Proceeding to 54, it is determined whether or not processing has been completed for all tokens. If processing has not been completed for all tokens, processing has proceeded to step S50, and the above processing is repeated.

【0104】例えば、ステップS3のトークン置換処理
において、テキストデータの単語の一次元列(w1 2
3 4 ・・・wT )のうち、単語列(w1 2 )、
(w1314)、・・・がトークンt1 で置換され、単語
列(w4 5 6 )、(w17 18)、・・・がトークン
2 で置換されたとすると、トークンt1 に対応する連
語として、{w1 −w2 、w13−w14、・・・}がテキ
ストデータから抽出され、トークンt2 に対応する連語
として、{w4 −w5 −w6 、w17−w18、・・・}が
テキストデータから抽出される。
For example, the token replacement processing in step S3
, A one-dimensional string (w1wTwo
wThreewFour... wT), The word string (w1wTwo),
(W13w14), ... are tokens t1Replaced by the word
Column (wFourwFivew6), (W17w 18), ... are tokens
tTwoIs replaced by the token t1Corresponding to
As a word,1-WTwo, W13-W14...
Extracted from the dataTwoCollocations corresponding to
As {wFour-WFive-W6, W17-W18,···}But
Extracted from text data.

【0105】1つの単語・トークンクラスTi が単語の
集合Wi とトークンの集合Ji ={ti1、ti2、・・・
in}からなり、トークンクラスTi が{Wi ∪Ji
により表され、、トークンの集合Ji の中の1つのトー
クンtimが、連語の集合Vim={vim (1) 、vim (2)
・・・}に逆トークン置換されたとすると、1つの単語
・連語クラスRi は、
One word / token class T i is a set of words W i and a set of tokens J i = {t i1 , t i2,.
t in }, and the token class T i is {W i {J i }
And one token t im in the set of tokens J i is represented by a set of collocations V im = {v im (1) , v im (2) ,
... If the inverse token replacement is made to {}, one word / syllable class R i becomes

【0106】[0106]

【数2】 (Equation 2)

【0107】で与えられる。以上説明したように、本発
明の一実施例による単語・連語分類処理装置によれば、
単語と連語とを区別することなく分類することができ
る。
Is given by As described above, according to the word / phrase classification processing device according to one embodiment of the present invention,
Words and collocations can be classified without distinction.

【0108】次に、本発明の一実施例による音声認識装
置について説明する。図23は、図1の単語・連語分類
処理装置により得られた単語・連語分類処理結果を利用
して音声認識を行う音声認識装置の構成を示すブロック
図である。
Next, a speech recognition apparatus according to an embodiment of the present invention will be described. FIG. 23 is a block diagram showing the configuration of a speech recognition device that performs speech recognition using the word / phrase classification processing result obtained by the word / phrase classification processing device of FIG.

【0109】図23において、所定のテキストデータ4
0に含まれる単語と連語とが、単語・連語分類処理部4
1により単語と連語とが混在するクラスに分類され、こ
の分類された単語と連語とが単語・連語辞書49に格納
されている。
In FIG. 23, predetermined text data 4
0, the word and the collocation classification processing unit 4
1, the words and collocations are classified into classes in which words and collocations coexist, and the classified words and collocations are stored in the word / composition dictionary 49.

【0110】一方、複数の単語と連語とからなる発音音
声は、マイクロフォン50によりアナログ音声信号に変
換された後、A/D変換器51でデジタル音声信号に変
換され、特徴抽出部52に入力される。特徴抽出部52
は、デジタル音声信号に対して、例えば、LPC分析を
行い、ケプストラム係数や対数パワーなどの特徴パラメ
ータを抽出する。特徴抽出部52で抽出された特徴パラ
メータは、音声認識部54に出力され、音素隠れマルコ
フモデルなどの言語モデル55を参照するとともに、単
語・連語辞書49に格納されている単語と連語との分類
結果を参照しながら、単語及び連語ごとに音声認識を行
う。
On the other hand, a pronunciation sound composed of a plurality of words and collocations is converted into an analog sound signal by a microphone 50, converted into a digital sound signal by an A / D converter 51, and input to a feature extraction unit 52. You. Feature extraction unit 52
Performs, for example, LPC analysis on a digital audio signal and extracts characteristic parameters such as cepstrum coefficients and logarithmic power. The feature parameters extracted by the feature extraction unit 52 are output to the speech recognition unit 54, refer to a language model 55 such as a phoneme hidden Markov model, and classify words and collocations stored in the word and collocation dictionary 49. While referring to the result, speech recognition is performed for each word and collocation.

【0111】図24は、単語・連語分類処理結果を利用
して音声認識を行う場合の例を示す図である。図24に
おいて、「本日は晴天なり」と発声された発音音声がマ
イクロフォン50に入力され、この発音音声に対して音
声モデルを適用するとにより、例えば、「本日は晴天な
り」という認識結果と「本日は静電なり」という認識結
果とが得られる。これらの音声モデルによる認識結果に
対し、言語モデルによる処理を行って単語・連語辞書4
9の参照を行い、「晴天なり」という連語が単語・連語
辞書49に登録されている場合、「本日は晴天なり」と
いう認識結果に対しては高い確率が与えられ、「本日は
静電なり」という認識結果に対しては低い確率が与えら
れる。
FIG. 24 is a diagram showing an example in which speech recognition is performed using the result of the word / phrase classification processing. In FIG. 24, a pronunciation voice uttered “Today is fine weather” is input to the microphone 50, and a speech model is applied to this pronunciation voice, for example, the recognition result “Today is fine weather” and “Today is fine weather” Is a static electricity ". The recognition results of these speech models are processed by a language model to obtain a word / syllable dictionary 4
9 is performed, and if the collocation word “sunny weather” is registered in the word / syllable dictionary 49, a high probability is given to the recognition result “today is clear weather”, and Is given a low probability.

【0112】以上説明したように、本発明の一実施例に
よる音声認識装置によれば、単語・連語辞書49を参照
して音声認識を行うことにより、より正確な認識処理が
可能になる。
As described above, according to the speech recognition apparatus according to one embodiment of the present invention, by performing speech recognition with reference to the word / syllable dictionary 49, more accurate recognition processing becomes possible.

【0113】次に、本発明の一実施例による機械翻訳装
置について説明する。図25は、図1の単語・連語分類
処理装置により得られた単語・連語分類処理結果を利用
して機械翻訳を行う機械翻訳装置の構成を示すブロック
図である。
Next, a machine translation apparatus according to one embodiment of the present invention will be described. FIG. 25 is a block diagram showing a configuration of a machine translation apparatus that performs machine translation using the word / phrase classification processing result obtained by the word / phrase classification processing apparatus of FIG.

【0114】図25において、所定のテキストデータ4
0に含まれる単語と連語とが、単語・連語分類処理部4
1により単語と連語とが混在するクラスに分類され、こ
の分類された単語と連語とが単語・連語辞書49に格納
されている。また、用例原文とその用例原文に対する用
例訳文とが、それぞれ対応させて用例文集60に格納さ
れている。
In FIG. 25, predetermined text data 4
0, the word and the collocation classification processing unit 4
1, the words and collocations are classified into classes in which words and collocations coexist, and the classified words and collocations are stored in the word / composition dictionary 49. The example original sentence and the example translated sentence corresponding to the example original sentence are stored in the example sentence collection 60 in association with each other.

【0115】用例検索部61に原文が入力されると、単
語・連語辞書49を参照しながら入力された原文の単語
が属するクラスを検索し、そのクラスと同一のクラスに
属する単語又は連語により構成される用例原文を用例文
集60から検索する。用例文集60から検索された用例
原文及びその用例訳文は、用例適用部62に入力され、
用例訳文の中の訳語を、入力された原文の単語に対する
訳語に置換することにより、入力された原文に対する訳
文を生成する。
When the original sentence is input to the example search section 61, the class to which the word of the input original sentence belongs is searched with reference to the word / syllable dictionary 49, and is composed of the words or the collocations belonging to the same class as the class. The example sentence to be used is searched from the example sentence collection 60. The example original sentence retrieved from the example sentence collection 60 and its example translation are input to the example application unit 62,
A translation for the input original text is generated by replacing the translation in the example translation text with a translation for the word of the input original text.

【0116】図26は、単語・連語分類処理結果を利用
して音声認識を行う場合の例を示す図である。図26に
おいて、“Toyota”と“Kohlberg Kr
avis Robert & Co.”とは同一のクラ
スに属し、“gained”と“lost”とは同一の
クラスに属し、“2”と“1”とは同一のクラスに属
し、“30 1/4”と“80 1/2”とは同一のク
ラスに属しているものとする。
FIG. 26 is a diagram showing an example in which speech recognition is performed using the result of the word / phrase classification processing. In FIG. 26, “Toyota” and “Kohlberg Kr”
avis Robert & Co. "Belong to the same class," gained "and" lost "belong to the same class," 2 "and" 1 "belong to the same class," 30 1/4 "and" 80 1 / "2" belongs to the same class.

【0117】原文として、“Toyota gaine
d 2 to 30 1/4.”が入力されると、用例
原文として、用例文集60から“Kohlberg K
ravis Robert & Co. lost 1
to 80 1/2.”が検索されるとともに、その
用例原文に対する用例訳文「Kohlberg Kra
vis Robert & Co.社は、1ドル値を下
げて終値80 1/2ドルだった。」も検索される。
As an original, “Toyota gain”
d 2 to 30 4. Is input from the example sentence collection 60 as "Kohlberg K" as an example original sentence.
ravis Robert & Co. lost 1
to 80 1/2. Is searched, and the example translation sentence “Kohlberg Kra” for the example original sentence is searched.
vis Robert & Co. The company has dropped $ 1 to close $ 1/20. Is also searched.

【0118】次に、用例原文の原語“Kohlberg
Kravis Robert &Co.”と同一のク
ラスに属している入力原文の原語“Toyota”に対
する訳語「トヨタ」で、用例訳文の訳語「Kohlbe
rg Kravis Robert & Co.社」を
置き換え、用例原文の原語“lost”と同一のクラス
に属している入力原文の原語“gained”に対する
訳語「上げて」で、用例訳文の訳語「下げて」を置き換
え、用例訳文の数値“1”を“2”で置き換え、用例訳
文の数値“80 1/2”を“30 1/4”で置き換
えることにより、入力原文に対する訳文「トヨタは、2
ドル値を上げて終値30 1/2ドルだった。」を出力
する。
Next, the original word “Kohlberg” in the original example text is used.
Kravis Robert & Co. "Toyota" for the original word "Toyota" of the input source text belonging to the same class as "" and the translation "Kohlbe" for the example translation
rg Kravis Robert & Co. Is replaced with the translation of the input source text "gained" of the input source text "gained" that belongs to the same class as the source text "lost" of the example source text, and the translation "down" of the example source translation. By replacing “1” with “2” and replacing the numerical value “80 1/2” of the example translation with “30 1/4”, the translated sentence “Toyota
The dollar rose to a closing price of $ 301/2. Is output.

【0119】以上説明したように、本発明の一実施例に
よる機械翻訳装置によれば、単語・連語辞書49を参照
して機械翻訳を行うことにより、より正確な翻訳処理が
可能になる。
As described above, according to the machine translation apparatus of one embodiment of the present invention, by performing machine translation with reference to the word / syllable dictionary 49, more accurate translation processing becomes possible.

【0120】以上、本発明の一実施例について説明した
が、本発明は上述した実施例に限定されるものではな
く、本発明の技術的思想の範囲内で他の様々な変更が可
能である。例えば、上述した実施例では、単語・連語分
類処理装置を音声認識装置及び機械翻訳装置に適用した
場合について説明したが、単語・連語分類処理装置を文
字認識装置に用いるようにしてもよい。また、上述した
実施例では、単語と連語とを混在される分類する場合に
ついて説明したが、連語のみを抽出し、この抽出した連
語を分類するようにしてもよい。
As described above, one embodiment of the present invention has been described. However, the present invention is not limited to the above-described embodiment, and various other modifications are possible within the technical idea of the present invention. . For example, in the above-described embodiment, a case has been described in which the word / phrase classification processing device is applied to a speech recognition device and a machine translation device. However, the word / phrase classification processing device may be used for a character recognition device. Further, in the above-described embodiment, the case of classifying a mixture of words and collocations has been described. However, only collocations may be extracted and the extracted collocations may be classified.

【0121】[0121]

【発明の効果】以上説明したように、本発明の単語・連
語分類処理装置によれば、テキストデータに含まれる単
語と連語とを一緒に分類して、単語と連語とが混在する
クラスを生成することにより、単語と単語とをまとめて
分類するだけでなく、単語と連語あるいは連語と連語と
をまとめて分類することができ、単語と連語あるいは連
語と連語との対応関係や類似度を容易に判別することが
できる。
As described above, according to the word / phrase classification processing device of the present invention, words and collocations included in text data are classified together to generate a class in which words and collocations are mixed. By doing so, it is possible not only to classify words and words collectively, but also to classify words and collocations or collocations and collocations, and to easily correlate words and collocations or collocations and collocations and to measure similarities. Can be determined.

【0122】また、本発明の一態様によれば、テキスト
データの単語クラス列にトークンを付与して単語クラス
列を1つの単語とみなし、テキストデータに含まれる単
語とトークンを付与された単語クラス列とを同等に取り
扱ってこれらを分類してから、テキストデータに存在す
る単語列で対応する単語クラス列を置き換えるようにし
たので、単語と連語との区別なく分類処理を行うことが
できるとともに、テキストデータからの連語の抽出を高
速に行うことができる。
Further, according to one aspect of the present invention, a token is assigned to the word class string of the text data, the word class string is regarded as one word, and the word included in the text data and the word class assigned with the token are added. Since the strings are treated equally and classified, and then the corresponding word class strings are replaced with the word strings existing in the text data, the classification process can be performed without distinction between words and collocations. Extraction of collocations from text data can be performed at high speed.

【0123】また、本発明の連語抽出装置によれば、テ
キストデータの単語列を構成する個々の単語を、その単
語が属する単語クラスで置換し、テキストデータにおい
て出現する確率が所定値以上の単語クラス列を抽出して
から、テキストデータに存在する連語を抽出することに
より、連語を高速に抽出することができる。
Further, according to the collocation extraction device of the present invention, each word constituting the word string of the text data is replaced with the word class to which the word belongs, and the word whose probability of occurrence in the text data is not less than a predetermined value is obtained. By extracting the collocations existing in the text data after extracting the class sequence, the collocations can be extracted at high speed.

【0124】また、本発明の音声認識装置によれば、単
語と連語あるいは連語と連語の対応関係や類似度を用い
ながら音声認識を行うことができ、正確な処理が可能に
なる。
Further, according to the speech recognition apparatus of the present invention, speech recognition can be performed using the correspondence and similarity between words and collocations or collocations and collocations, and accurate processing becomes possible.

【0125】また、本発明の機械翻訳装置によれば、用
例文集に格納されている用例原文の単語が連語に置き換
わった原文が入力された場合においても、入力された原
文に用例原文を適用して機械翻訳を行うことができ、単
語と連語あるいは連語と連語の対応関係や類似度を用い
た正確な機械翻訳が可能になる。
According to the machine translation apparatus of the present invention, even when an original sentence in which words of the example original sentence stored in the example sentence collection are replaced with collocations is input, the example original sentence is applied to the input original sentence. Machine translation, and accurate machine translation using the correspondence and similarity between words and collocations or collocations and collocations becomes possible.

【図面の簡単な説明】[Brief description of the drawings]

【図1】本発明の一実施例に係わる単語・連語分類処理
装置の機能的な構成を示すブロック図である。
FIG. 1 is a block diagram showing a functional configuration of a word / phrase classification processing device according to an embodiment of the present invention.

【図2】本発明の一実施例に係わる単語・連語分類処理
装置の単語クラスタリング処理を説明する図である。
FIG. 2 is a diagram illustrating a word clustering process of the word / phrase classification processing device according to one embodiment of the present invention.

【図3】図1の単語分類手段の機能的な構成を示すブロ
ック図である。
FIG. 3 is a block diagram illustrating a functional configuration of a word classification unit in FIG. 1;

【図4】本発明の一実施例に係わる単語・連語分類処理
装置の単語クラス列生成処理を説明する図である。
FIG. 4 is a diagram illustrating a word class sequence generation process of the word / phrase classification processing device according to one embodiment of the present invention.

【図5】本発明の一実施例に係わる単語・連語分類処理
装置のクラスチェーン抽出処理を説明する図である。
FIG. 5 is a diagram illustrating a class chain extraction process of the word / phrase classification processing device according to one embodiment of the present invention.

【図6】図1の単語クラス列抽出手段の機能的な構成を
示すブロック図である。
FIG. 6 is a block diagram illustrating a functional configuration of a word class string extracting unit in FIG. 1;

【図7】本発明の一実施例に係わる単語・連語分類処理
装置によるクラスチェーンとトークンとの関係を示す図
である。
FIG. 7 is a diagram showing a relationship between a class chain and a token by the word / phrase classification processing device according to one embodiment of the present invention.

【図8】本発明の一実施例に係わる単語・連語分類処理
装置のトークン置換処理を説明する図である。
FIG. 8 is a diagram illustrating token replacement processing of the word / phrase classification processing device according to one embodiment of the present invention.

【図9】本発明の一実施例に係わる単語・連語分類処理
装置によるトークン置換処理の英文例を示す図である。
FIG. 9 is a diagram showing an example of an English sentence in a token replacement process by the word / phrase classification processing device according to one embodiment of the present invention.

【図10】図1の単語・トークン分類手段の機能的な構
成を示すブロック図である。
FIG. 10 is a block diagram showing a functional configuration of the word / token classification means of FIG. 1;

【図11】本発明の一実施例に係わる単語・連語分類処
理装置によるトークンと連語の関係を示す図である。
FIG. 11 is a diagram showing the relationship between tokens and collocations by the word / phrase classification processing device according to one embodiment of the present invention.

【図12】本発明の一実施例に係わる単語・連語分類処
理装置による単語・連語分類処理結果を示す図である。
FIG. 12 is a diagram showing a result of a word / phrase classification process performed by the word / phrase classification processing device according to one embodiment of the present invention.

【図13】本発明の一実施例に係わる単語・連語分類処
理装置のシステム構成を示すブロック図である。
FIG. 13 is a block diagram illustrating a system configuration of a word / phrase classification processing device according to an embodiment of the present invention.

【図14】本発明の一実施例に係わる単語・連語分類処
理装置の単語・連語分類処理を示すフローチャートであ
る。
FIG. 14 is a flowchart showing a word / phrase classification process of the word / phrase classification processing device according to one embodiment of the present invention.

【図15】本発明の一実施例に係わる単語・連語分類処
理装置のウインドウ処理を説明する図である。
FIG. 15 is a diagram illustrating window processing of the word / phrase classification processing device according to one embodiment of the present invention.

【図16】本発明の一実施例に係わる単語・連語分類処
理装置の単語クラスタリング処理を示すフローチャート
である。
FIG. 16 is a flowchart showing a word clustering process of the word / phrase classification processing device according to one embodiment of the present invention.

【図17】本発明に係わる単語・連語分類処理装置のク
ラスチェーン抽出処理の第1実施例を示すフローチャー
トである。
FIG. 17 is a flowchart showing a first embodiment of a class chain extraction process of the word / phrase classification processing device according to the present invention.

【図18】本発明に係わる単語・連語分類処理装置のク
ラスチェーン抽出処理の第2実施例を示すフローチャー
トである。
FIG. 18 is a flowchart showing a second embodiment of the class chain extraction processing of the word / phrase classification processing device according to the present invention.

【図19】本発明に係わる単語・連語分類処理装置のク
ラスチェーン抽出処理の第2実施例を説明する図であ
る。
FIG. 19 is a diagram illustrating a second embodiment of the class chain extraction processing of the word / phrase classification processing device according to the present invention.

【図20】本発明の一実施例に係わる単語・連語分類処
理装置のトークン置換処理を示すフローチャートであ
る。
FIG. 20 is a flowchart showing token replacement processing of the word / phrase classification processing device according to one embodiment of the present invention.

【図21】本発明の一実施例に係わる単語・連語分類処
理装置の単語・トークンクラスタリング処理を示すフロ
ーチャートである。
FIG. 21 is a flowchart showing word / token clustering processing of the word / phrase classification processing device according to one embodiment of the present invention.

【図22】本発明の一実施例に係わる単語・連語分類処
理装置のデータ出力処理を示すフローチャートである。
FIG. 22 is a flowchart showing data output processing of the word / phrase classification processing device according to one embodiment of the present invention.

【図23】本発明の一実施例に係わる音声認識装置の機
能的な構成を示すブロック図である。
FIG. 23 is a block diagram showing a functional configuration of a speech recognition device according to one embodiment of the present invention.

【図24】本発明の一実施例に係わる音声認識方法を説
明する図である。
FIG. 24 is a diagram illustrating a voice recognition method according to an embodiment of the present invention.

【図25】本発明の一実施例に係わる機械翻訳装置の機
能的な構成を示すブロック図である。
FIG. 25 is a block diagram illustrating a functional configuration of a machine translation apparatus according to an embodiment of the present invention.

【図26】本発明の一実施例に係わる機械翻訳方法を説
明する図である。
FIG. 26 is a diagram illustrating a machine translation method according to an embodiment of the present invention.

【符号の説明】[Explanation of symbols]

1 単語分類手段 2 単語クラス列生成手段 3 単語クラス列抽出手段 4 トークン付与手段 5 単語・トークン列生成手段 6 単語・トークン分類手段 7 連語置換手段 40 テキストデータ 41 単語・連語分類処理部 42、46 メモリインターフェイス 43 CPU 44 ROM 45 ワークRAM 47 RAM 48 バス 49 単語・連語辞書 50 マイクロフォン 51 A/D変換器 52 特徴抽出部 53 バッファメモリ 54 音声認識部 55 言語モデル 60 用例文集 61 用例検索部 62 用例適用部 REFERENCE SIGNS LIST 1 word classification means 2 word class string generation means 3 word class string extraction means 4 token assignment means 5 word / token string generation means 6 word / token classification means 7 word replacement means 40 text data 41 word / word word classification processing units 42, 46 Memory interface 43 CPU 44 ROM 45 Work RAM 47 RAM 48 Bus 49 Word / phrase dictionary 50 Microphone 51 A / D converter 52 Feature extraction unit 53 Buffer memory 54 Voice recognition unit 55 Language model 60 Example sentence collection 61 Example search unit 62 Example application Department

Claims (17)

【特許請求の範囲】[Claims] 【請求項1】 複数の単語の一次元列としてのテキスト
データから、互いに異なるV個の単語を抽出し、前記V
個の単語の集合をC個の単語クラスに分割した第1のク
ラスタリングを生成するステップと、 前記第1のクラスタリングに基づいて生成された前記テ
キストデータの単語クラスの一次元列において、隣接す
る単語クラス間の粘着度が全て所定値以上の単語クラス
列の集合を抽出するステップと、 前記単語クラス列に固有のトークンを対応させ、前記単
語クラス列に属する単語列を前記テキストデータから検
索し、前記テキストデータの単語列を対応するトークン
で置換することにより、前記テキストデータについての
単語とトークンとの一次元列を生成するステップと、 前記テキストデータについての単語とトークンとの一次
元列において、互いに異なる単語と互いに異なるトーク
ンとを抽出し、前記単語と前記トークンとが混在する集
合を単語・トークンクラスに分割した第2のクラスタリ
ングを生成するステップと、 前記テキストデータに存在する単語列のうち、前記トー
クンに対応するものを連語として抽出し、前記単語・ト
ークンクラスの中のトークンを前記連語で置換すること
により、前記単語と前記連語とが混在する集合を単語・
連語クラスに分割した第3のクラスタリングを生成する
ステップとを備えることを特徴とする単語・連語分類処
理方法。
1. A method for extracting V words different from each other from text data as a one-dimensional string of a plurality of words,
Generating a first clustering obtained by dividing a set of words into C word classes; and in the one-dimensional sequence of word classes of the text data generated based on the first clustering, adjacent words A step of extracting a set of word class strings in which the degree of adhesion between the classes is all equal to or greater than a predetermined value, and causing a unique token to correspond to the word class string, and searching a word string belonging to the word class string from the text data Generating a one-dimensional sequence of words and tokens for the text data by replacing the word sequence of the text data with corresponding tokens; anda one-dimensional sequence of words and tokens for the text data. A set in which different words and different tokens are extracted, and the words and the tokens are mixed. Generating a second clustering divided into words / token classes; extracting word strings corresponding to the tokens among word strings existing in the text data as collocations; and extracting tokens in the word / token classes. By replacing the word with the collocation, a set in which the word and the collocation are mixed is defined as a word / word.
Generating a third clustering divided into collocation classes.
【請求項2】 前記第1のクラスタリングは、前記単語
クラスの平均相互情報量に基づいて生成されることを特
徴とする請求項1に記載の単語・連語分類処理方法。
2. The word / phrase classification method according to claim 1, wherein the first clustering is generated based on an average mutual information amount of the word classes.
【請求項3】 前記第2のクラスタリングは、前記単語
・トークンクラスの平均相互情報量に基づいて生成され
ることを特徴とする請求項1に記載の単語・連語分類処
理方法。
3. The method according to claim 1, wherein the second clustering is generated based on an average mutual information of the word and token classes.
【請求項4】 テキストデータに含まれる単語を分類し
た単語クラスを生成するステップと、 前記単語クラスを前記テキストデータの単語の一次元列
にマッピングして単語クラスの一次元列を生成するステ
ップと、 前記テキストデータの単語クラスの一次元列において、
隣接する単語クラス間の粘着度が全て所定値以上の単語
クラス列を、前記テキストデータの単語クラスの一次元
列から抽出するステップと、 前記テキストデータに含まれる単語と前記単語クラス列
とを一緒に分類するステップと、 前記単語クラス列を構成する個々の単語クラスから、前
記テキストデータに隣接して存在する個々の単語を別々
に取り出して連語を抽出するステップと、 前記単語クラス列を前記単語クラス列に属する連語で置
換するステップとを備えることを特徴とする単語・連語
分類処理方法。
Generating a word class in which words included in the text data are classified; and generating a one-dimensional string of the word class by mapping the word class to a one-dimensional string of the words in the text data. In a one-dimensional sequence of word classes of the text data,
Extracting, from the one-dimensional sequence of the word classes of the text data, a word class sequence in which the degree of adhesion between adjacent word classes is all equal to or greater than a predetermined value; and combining the words included in the text data with the word class sequence. And extracting individual words adjacent to the text data from the individual word classes constituting the word class sequence to extract collocations; and extracting the word class sequence into the words Replacing with a collocation belonging to a class sequence.
【請求項5】 テキストデータに含まれる単語を分類し
た単語クラスを生成するステップと、 前記単語クラスを前記テキストデータの単語の一次元列
にマッピングして単語クラスの一次元列を生成するステ
ップと、 前記テキストデータの単語クラスの一次元列において、
隣接する単語クラス間の粘着度が全て所定値以上の単語
クラス列を、前記テキストデータの単語クラスの一次元
列から抽出するステップと、 前記単語クラス列を構成する個々の単語クラスから、前
記テキストデータに隣接して存在する個々の単語を別々
に取り出して連語を抽出するステップとを備えることを
特徴とする連語抽出方法。
5. A step of generating a word class in which words included in text data are classified; and a step of mapping the word class to a one-dimensional string of words in the text data to generate a one-dimensional string of word classes. In a one-dimensional sequence of word classes of the text data,
Extracting, from a one-dimensional sequence of the word classes of the text data, a word class sequence in which the degree of adhesion between adjacent word classes is all equal to or greater than a predetermined value; Separately extracting individual words existing adjacent to the data to extract a collocation word.
【請求項6】 テキストデータの単語列から互いに異な
る単語を抽出し、抽出された前記単語の集合を分割して
単語クラスを生成する単語分類手段と、 前記テキストデータの単語の一次元列を構成する個々の
単語を、前記単語が属する前記単語クラスで置換するこ
とにより、前記テキストデータの単語クラスの一次元列
を生成する単語クラス列生成手段と、 前記テキストデータの単語クラスの一次元列において、
隣接する単語クラス間の粘着度が全て所定値以上の単語
クラス列を、前記テキストデータの単語クラスの一次元
列から抽出する単語クラス列抽出手段と、 前記単語クラス列抽出手段により抽出された各単語クラ
ス列にトークンを付与するトークン付与手段と、 前記テキストデータの単語の一次元列のうち、前記単語
クラス列抽出手段により抽出された単語クラス列に属す
る単語列を前記トークンで置換することにより、前記テ
キストデータの単語・トークンの一次元列を生成する単
語・トークン列生成手段と、 前記テキストデータの単語・トークンの一次元列に含ま
れる単語とトークンとが混在する集合を分割して単語・
トークンクラスを生成する単語・トークン分類手段と、 前記単語・トークンクラスの中のトークンを、前記単語
・トークン列生成手段により置換された単語列に逆置換
して連語を生成する連語置換手段とを備えることを特徴
とする単語・連語分類処理装置。
6. A word classifying means for extracting words different from each other from a word string of text data, generating a word class by dividing the set of extracted words, and forming a one-dimensional string of words of the text data A word class sequence generating means for generating a one-dimensional sequence of the word class of the text data by replacing each word to be performed with the word class to which the word belongs; ,
A word class sequence extracting unit that extracts a word class sequence in which the degree of adhesion between adjacent word classes is all equal to or greater than a predetermined value from a one-dimensional sequence of the word classes of the text data; A token assigning means for assigning a token to a word class string, and replacing the word string belonging to the word class string extracted by the word class string extracting means with the token in the one-dimensional string of words of the text data. A word / token sequence generating means for generating a one-dimensional sequence of words / tokens of the text data; and dividing a set in which words and tokens contained in the one-dimensional sequence of words / tokens of the text data are mixed to form a word・
Word / token classifying means for generating a token class; and collocation replacing means for generating a collocation by inversely replacing the token in the word / token class with the word string replaced by the word / token string generating means. A word / phrase classification processing device, comprising:
【請求項7】 前記単語分類手段は、 前記テキストデータの単語の一次元列から互いに異なる
単語を抽出し、所定の出現頻度を有する単語のそれぞれ
に固有の単語クラスを割り当てる初期化クラス設定部
と、 単語クラスの集合から2つの単語クラスを取り出して仮
マージする仮マージ部と、 前記テキストデータの仮マージされた単語クラスについ
ての平均相互情報量を算出する平均相互情報量算出部
と、 前記単語クラスの集合のうち、前記平均相互情報量が最
大である2つの単語クラスを本マージする本マージ部と
を備えることを特徴とする請求項6に記載の単語・連語
分類処理装置。
7. An initialization class setting unit that extracts words different from each other from a one-dimensional string of words of the text data, and assigns a unique word class to each word having a predetermined appearance frequency. A temporary merging unit that extracts two word classes from a set of word classes and temporarily merges them; an average mutual information calculating unit that calculates an average mutual information amount of the temporarily merged word classes of the text data; 7. The word / phrase classification processing device according to claim 6, further comprising: a main merge unit that performs a main merge of two word classes having the largest average mutual information amount among a set of classes.
【請求項8】 前記単語クラス列抽出手段は、 前記テキストデータの単語クラスの一次元列から、隣接
して存在する2つの単語クラスを順次に取り出す単語ク
ラス取出部と、 前記単語クラス取出部により取り出した2つの単語クラ
スの相互情報量を算出する相互情報量算出部と、 前記相互情報量が所定のしきい値以上の2つの単語クラ
スをクラスチェーンで結ぶクラスチェーン結合部とを備
えることを特徴とする請求項6に記載の単語・連語分類
処理装置。
8. A word class extracting section, comprising: a word class extracting section for sequentially extracting two adjacent word classes from a one-dimensional string of the word class of the text data; and a word class extracting section. A mutual information calculating unit that calculates a mutual information amount of the two extracted word classes; and a class chain connecting unit that connects two word classes having the mutual information amount equal to or more than a predetermined threshold value with a class chain. 7. The word / phrase classification processing device according to claim 6, wherein:
【請求項9】 前記単語・トークン分類手段は、 前記テキストデータの単語・トークンの一次元列から互
いに異なる単語と互いに異なるトークンとを抽出し、所
定の出現頻度を有する単語とトークンとのそれぞれに固
有の単語・トークンクラスを割り当てる初期化クラス設
定部と、 単語・トークンクラスの集合から2つの単語・トークン
クラスを取り出して仮マージする仮マージ部と、 前記テキストデータの仮マージされた単語・トークンク
ラスについての平均相互情報量を算出する平均相互情報
量算出部と、 前記単語・トークンクラスの集合のうち、前記平均相互
情報量が最大である2つの単語・トークンクラスを本マ
ージする本マージ部とを備えることを特徴とする請求項
6に記載の単語・連語分類処理装置。
9. The word / token classification means extracts a different word and a different token from a one-dimensional string of words / tokens of the text data, and extracts each of the words and tokens having a predetermined appearance frequency. An initialization class setting unit for allocating a unique word / token class; a temporary merging unit for taking out two words / token classes from a set of word / token classes; and temporarily merging the words / token classes; An average mutual information calculation unit for calculating an average mutual information amount for a class; and a main merge unit for performing a main merge of two word / token classes having the maximum average mutual information amount among the set of the word / token classes. 7. The word / phrase classification processing device according to claim 6, comprising:
【請求項10】 テキストデータから連語を抽出する連
語抽出手段と、 前記テキストデータに含まれる単語と連語とを一緒に分
類して、単語と連語とが混在するクラスを生成する単語
・連語分類手段とを備えることを特徴とする単語・連語
分類処理装置。
10. A collocation extracting means for extracting a collocation from text data, and a word / composition classification means for classifying together words and collocations included in the text data to generate a class in which words and collocations are mixed. A word / phrase classification processing device comprising:
【請求項11】 前記クラスは、前記クラスの平均相互
情報量に基づいて生成されることを特徴とする請求項1
0に記載の単語・連語分類処理装置。
11. The method according to claim 1, wherein the class is generated based on an average mutual information amount of the class.
0, a word / syllable classification processing device.
【請求項12】 テキストデータに含まれる単語を分類
して単語クラスを生成する単語分類手段と、 前記テキストデータの単語の一次元列を構成する個々の
単語を、前記単語が属する前記単語クラスで置換するこ
とにより、前記テキストデータの単語クラスの一次元列
を生成する単語クラス列生成手段と、 前記テキストデータの単語クラスの一次元列において、
隣接する単語クラス間の粘着度が全て所定値以上の単語
クラス列を、前記テキストデータの単語クラスの一次元
列から抽出する単語クラス列抽出手段と、 前記単語クラス列を構成する個々の単語クラスから、前
記テキストデータに隣接して存在する個々の単語を別々
に取り出して連語を抽出する連語抽出手段とを備えるこ
とを特徴とする連語抽出装置。
12. A word classifying means for classifying a word included in text data to generate a word class, wherein each word constituting a one-dimensional sequence of words of the text data is represented by the word class to which the word belongs. A word class string generating means for generating a one-dimensional string of the word class of the text data by substituting;
Word class string extracting means for extracting, from a one-dimensional string of the word classes of the text data, a word class string in which the degree of adhesion between adjacent word classes is all equal to or greater than a predetermined value; individual word classes constituting the word class string A collocation extracting means for separately extracting each word existing adjacent to the text data and extracting a collocation from the text data.
【請求項13】 前記単語クラスは、前記単語クラスの
平均相互情報量に基づいて生成されることを特徴とする
請求項12に記載の連語抽出装置。
13. The collocation extraction apparatus according to claim 12, wherein the word class is generated based on an average mutual information amount of the word class.
【請求項14】 所定のテキストデータに含まれる単語
と連語とを、単語と連語とが混在するクラスに分類して
格納している単語・連語辞書と、 前記単語・連語辞書と所定の隠れマルコフモデルとを参
照することにより、発音音声を音声認識する音声認識手
段とを備えることを特徴とする音声認識装置。
14. A word / concatenation dictionary storing words and collocations included in predetermined text data in a class in which words and concatenations are mixed, and the word / concatenation dictionary and a predetermined hidden Markov dictionary. And a voice recognition unit for recognizing the pronunciation voice by referring to the model.
【請求項15】 所定のテキストデータに含まれる単語
と連語とを、単語と連語とが混在するクラスに分類して
格納している単語・連語辞書と、 用例原文と前記用例原文に対する用例訳文とを対応させ
て格納している用例文集と、 入力された原文の単語が属するクラスと同一のクラスに
属する単語又は連語により構成される用例原文を前記用
例文集から検索する用例検索手段と、 前記用例原文に対する用例訳文の中の訳語を、入力され
た原文の単語に対する訳語に置換することにより、前記
入力された原文に対する訳文を生成する用例適用手段と
を備えることを特徴とする機械翻訳装置。
15. A word / syllable dictionary that stores words and collocations included in predetermined text data in a class in which words and collocations are mixed, an example original sentence, and an example translation sentence for the example original sentence. An example sentence collection in which an example original sentence composed of words or collocations belonging to the same class as the class to which the input original word belongs is searched from the example sentence collection; and A machine translation apparatus comprising: an example application unit configured to generate a translation of the input original text by replacing a translation in an example translation of the original text with a translation of a word of the input original text.
【請求項16】 所定のテキストデータに含まれる単語
と連語とを、単語と連語とが混在するクラスに分類して
格納している単語・連語記憶媒体であって、 前記クラスは、前記クラスの平均相互情報量に基づいて
生成されていることを特徴とする単語・連語記憶媒体。
16. A word / syllable storage medium that stores words and collocations included in predetermined text data in a class in which words and collocations are mixed, wherein the class is a class of the class. A word / syllable storage medium characterized by being generated based on an average mutual information amount.
【請求項17】 テキストデータの単語の一次元列から
互いに異なる単語を抽出し、抽出された前記単語の集合
を分割して単語クラスを生成する機能と、 前記テキストデータの単語の一次元列を構成する個々の
単語を、前記単語が属する前記単語クラスで置換するこ
とにより、前記テキストデータの単語クラスの一次元列
を生成する機能と、 前記テキストデータの単語クラスの一次元列から、隣接
する単語クラス間の粘着度が全て所定値以上の単語クラ
ス列を抽出する機能と、 前記単語クラス列にトークンを付与する機能と、 前記テキストデータの単語の一次元列のうち、前記単語
クラス列に属する単語列を前記トークンで置換すること
により、前記テキストデータの単語・トークンの一次元
列を生成する機能と、 前記テキストデータの単語・トークンの一次元列に含ま
れる単語とトークンとが混在する集合を分割して単語・
トークンクラスを生成する機能と、 前記単語・トークンクラスの中のトークンを、前記テキ
ストデータに存在する単語列に逆置換して連語を生成す
る機能とをコンピュータに実行させるプログラムを格納
したコンピュータ読み取り可能な記憶媒体。
17. A function of extracting different words from a one-dimensional string of words in text data, dividing the set of extracted words to generate a word class, and A function of generating a one-dimensional column of the word class of the text data by replacing each constituent word with the word class to which the word belongs; A function of extracting a word class string in which the degree of adhesion between word classes is all equal to or greater than a predetermined value; a function of assigning a token to the word class string; and a one-dimensional string of words of the text data, A function of generating a one-dimensional string of words / tokens of the text data by replacing the word string to which the text data belongs with the token; - and the words and token that is included in the one-dimensional string of tokens by dividing the set of mixed word -
A computer-readable program storing a program for causing a computer to execute a function of generating a token class and a function of generating a collocation by reversely replacing a token in the word / token class with a word string existing in the text data Storage media.
JP16724397A 1996-08-02 1997-06-24 Word / collocation classification processing method, collocation extraction method, word / collocation classification processing device, speech recognition device, machine translation device, collocation extraction device, and word / collocation storage medium Expired - Fee Related JP3875357B2 (en)

Priority Applications (1)

Application Number Priority Date Filing Date Title
JP16724397A JP3875357B2 (en) 1996-08-02 1997-06-24 Word / collocation classification processing method, collocation extraction method, word / collocation classification processing device, speech recognition device, machine translation device, collocation extraction device, and word / collocation storage medium

Applications Claiming Priority (3)

Application Number Priority Date Filing Date Title
JP20498696 1996-08-02
JP8-204986 1996-08-02
JP16724397A JP3875357B2 (en) 1996-08-02 1997-06-24 Word / collocation classification processing method, collocation extraction method, word / collocation classification processing device, speech recognition device, machine translation device, collocation extraction device, and word / collocation storage medium

Publications (2)

Publication Number Publication Date
JPH1097286A true JPH1097286A (en) 1998-04-14
JP3875357B2 JP3875357B2 (en) 2007-01-31

Family

ID=26491346

Family Applications (1)

Application Number Title Priority Date Filing Date
JP16724397A Expired - Fee Related JP3875357B2 (en) 1996-08-02 1997-06-24 Word / collocation classification processing method, collocation extraction method, word / collocation classification processing device, speech recognition device, machine translation device, collocation extraction device, and word / collocation storage medium

Country Status (1)

Country Link
JP (1) JP3875357B2 (en)

Cited By (8)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US7269789B2 (en) 2003-04-10 2007-09-11 Mitsubishi Denki Kabushiki Kaisha Document information processing apparatus
EP1551007A4 (en) * 2002-10-08 2008-05-21 Matsushita Electric Industrial Co Ltd LANGUAGE MODEL CREATION / CREATION DEVICE, VOICE RECOGNITION DEVICE, LANGUAGE MODEL CREATION METHOD, AND VOICE RECOGNITION METHOD
JP2013083897A (en) * 2011-10-12 2013-05-09 Fujitsu Ltd Recognition device, recognition program, recognition method, generation device, generation program and generation method
US9524295B2 (en) 2006-10-26 2016-12-20 Facebook, Inc. Simultaneous translation of open domain lectures and speeches
US9753918B2 (en) 2008-04-15 2017-09-05 Facebook, Inc. Lexicon development via shared translation database
CN111159409A (en) * 2019-12-31 2020-05-15 腾讯科技(深圳)有限公司 Text classification method, device, equipment and medium based on artificial intelligence
CN111768023A (en) * 2020-05-11 2020-10-13 国网冀北电力有限公司电力科学研究院 A probabilistic peak load estimation method based on smart city energy meter data
US11222185B2 (en) 2006-10-26 2022-01-11 Meta Platforms, Inc. Lexicon development via shared translation database

Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPS6154562A (en) * 1984-08-24 1986-03-18 Nec Corp Japanese input device
JPH03179498A (en) * 1989-12-08 1991-08-05 Nippon Telegr & Teleph Corp <Ntt> Voice japanese conversion system
JPH05189481A (en) * 1991-07-25 1993-07-30 Internatl Business Mach Corp <Ibm> Computor operating method for translation, term- model forming method, model forming method, translation com-putor system, term-model forming computor system and model forming computor system
JPH06274546A (en) * 1993-03-19 1994-09-30 A T R Jido Honyaku Denwa Kenkyusho:Kk Information quantity matching degree calculation system
JPH06301722A (en) * 1993-04-13 1994-10-28 Matsushita Electric Ind Co Ltd Morpheme analyzing device and keyword extracting device

Patent Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPS6154562A (en) * 1984-08-24 1986-03-18 Nec Corp Japanese input device
JPH03179498A (en) * 1989-12-08 1991-08-05 Nippon Telegr & Teleph Corp <Ntt> Voice japanese conversion system
JPH05189481A (en) * 1991-07-25 1993-07-30 Internatl Business Mach Corp <Ibm> Computor operating method for translation, term- model forming method, model forming method, translation com-putor system, term-model forming computor system and model forming computor system
JPH06274546A (en) * 1993-03-19 1994-09-30 A T R Jido Honyaku Denwa Kenkyusho:Kk Information quantity matching degree calculation system
JPH06301722A (en) * 1993-04-13 1994-10-28 Matsushita Electric Ind Co Ltd Morpheme analyzing device and keyword extracting device

Non-Patent Citations (2)

* Cited by examiner, † Cited by third party
Title
柏岡秀紀他: ""相互情報量を用いた単語の分類における出現頻度の低い単語の処理手法"", 情報処理学会第49回(平成6年後期)全国大会講演論文集(3), vol. 1994年9月,7G-5, JPNX006049922, pages 185 - 186, ISSN: 0000784930 *
柏岡秀紀他: "相互情報量を用いた単語の分類における出現頻度の低い単語の処理方法", 情報処理学会第49回(平成6年後期)全国大会講演論文集(3), vol. 7G-5, JPN4005003875, 20 September 1994 (1994-09-20), pages 185 - 186, ISSN: 0000751010 *

Cited By (13)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
EP1551007A4 (en) * 2002-10-08 2008-05-21 Matsushita Electric Industrial Co Ltd LANGUAGE MODEL CREATION / CREATION DEVICE, VOICE RECOGNITION DEVICE, LANGUAGE MODEL CREATION METHOD, AND VOICE RECOGNITION METHOD
US7269789B2 (en) 2003-04-10 2007-09-11 Mitsubishi Denki Kabushiki Kaisha Document information processing apparatus
US11222185B2 (en) 2006-10-26 2022-01-11 Meta Platforms, Inc. Lexicon development via shared translation database
US9524295B2 (en) 2006-10-26 2016-12-20 Facebook, Inc. Simultaneous translation of open domain lectures and speeches
US9830318B2 (en) 2006-10-26 2017-11-28 Facebook, Inc. Simultaneous translation of open domain lectures and speeches
US11972227B2 (en) 2006-10-26 2024-04-30 Meta Platforms, Inc. Lexicon development via shared translation database
US9753918B2 (en) 2008-04-15 2017-09-05 Facebook, Inc. Lexicon development via shared translation database
US9082404B2 (en) 2011-10-12 2015-07-14 Fujitsu Limited Recognizing device, computer-readable recording medium, recognizing method, generating device, and generating method
JP2013083897A (en) * 2011-10-12 2013-05-09 Fujitsu Ltd Recognition device, recognition program, recognition method, generation device, generation program and generation method
CN111159409A (en) * 2019-12-31 2020-05-15 腾讯科技(深圳)有限公司 Text classification method, device, equipment and medium based on artificial intelligence
CN111159409B (en) * 2019-12-31 2023-06-02 腾讯科技(深圳)有限公司 Text classification method, device, equipment and medium based on artificial intelligence
CN111768023A (en) * 2020-05-11 2020-10-13 国网冀北电力有限公司电力科学研究院 A probabilistic peak load estimation method based on smart city energy meter data
CN111768023B (en) * 2020-05-11 2024-04-09 国网冀北电力有限公司电力科学研究院 A probabilistic peak load estimation method based on smart city electric energy meter data

Also Published As

Publication number Publication date
JP3875357B2 (en) 2007-01-31

Similar Documents

Publication Publication Date Title
US6178396B1 (en) Word/phrase classification processing method and apparatus
CN110364171B (en) Voice recognition method, voice recognition system and storage medium
CN112397054B (en) Power dispatching voice recognition method
JP6171544B2 (en) Audio processing apparatus, audio processing method, and program
US20110131038A1 (en) Exception dictionary creating unit, exception dictionary creating method, and program therefor, as well as speech recognition unit and speech recognition method
KR20030018073A (en) Voice recognition apparatus and voice recognition method
CN107562760A (en) A kind of voice data processing method and device
CN112201275B (en) Voiceprint segmentation method, voiceprint segmentation device, voiceprint segmentation equipment and readable storage medium
CN113051923B (en) Data verification method and device, computer equipment and storage medium
US8423354B2 (en) Speech recognition dictionary creating support device, computer readable medium storing processing program, and processing method
CN114818649A (en) Service consultation processing method and device based on intelligent voice interaction technology
CN112259083A (en) Audio processing method and device
CN117059076A (en) Dialect voice recognition method, device, equipment and storage medium
KR20230066970A (en) Method for processing natural language, method for generating grammar and dialogue system
CN109961775A (en) Dialect recognition method, device, equipment and medium based on HMM model
CN115391506A (en) Method and device for detecting standardization of question and answer content for multi-segment replies
JP3875357B2 (en) Word / collocation classification processing method, collocation extraction method, word / collocation classification processing device, speech recognition device, machine translation device, collocation extraction device, and word / collocation storage medium
CN117292680A (en) A speech recognition method for power transmission and inspection based on small sample synthesis
CN113990288B (en) A method for automatically generating and deploying a speech synthesis model for voice customer service
Imperl et al. Clustering of triphones using phoneme similarity estimation for the definition of a multilingual set of triphones
JP3911178B2 (en) Speech recognition dictionary creation device and speech recognition dictionary creation method, speech recognition device, portable terminal, speech recognition system, speech recognition dictionary creation program, and program recording medium
JP2000221991A (en) Appropriate word string estimation device
JPH06266393A (en) Voice recognizer
Liu et al. Supra-Segmental Feature Based Speaker Trait Detection.
Liao et al. Towards the development of automatic speech recognition for Bikol and Kapampangan

Legal Events

Date Code Title Description
A977 Report on retrieval

Free format text: JAPANESE INTERMEDIATE CODE: A971007

Effective date: 20050630

A131 Notification of reasons for refusal

Free format text: JAPANESE INTERMEDIATE CODE: A131

Effective date: 20050705

A521 Request for written amendment filed

Free format text: JAPANESE INTERMEDIATE CODE: A523

Effective date: 20050823

A02 Decision of refusal

Free format text: JAPANESE INTERMEDIATE CODE: A02

Effective date: 20060704

AA91 Notification that invitation to amend document was cancelled

Free format text: JAPANESE INTERMEDIATE CODE: A971091

Effective date: 20060725

A131 Notification of reasons for refusal

Free format text: JAPANESE INTERMEDIATE CODE: A131

Effective date: 20060905

A521 Request for written amendment filed

Free format text: JAPANESE INTERMEDIATE CODE: A523

Effective date: 20060928

TRDD Decision of grant or rejection written
A01 Written decision to grant a patent or to grant a registration (utility model)

Free format text: JAPANESE INTERMEDIATE CODE: A01

Effective date: 20061024

A61 First payment of annual fees (during grant procedure)

Free format text: JAPANESE INTERMEDIATE CODE: A61

Effective date: 20061026

R150 Certificate of patent or registration of utility model

Free format text: JAPANESE INTERMEDIATE CODE: R150

FPAY Renewal fee payment (event date is renewal date of database)

Free format text: PAYMENT UNTIL: 20101102

Year of fee payment: 4

FPAY Renewal fee payment (event date is renewal date of database)

Free format text: PAYMENT UNTIL: 20101102

Year of fee payment: 4

FPAY Renewal fee payment (event date is renewal date of database)

Free format text: PAYMENT UNTIL: 20111102

Year of fee payment: 5

FPAY Renewal fee payment (event date is renewal date of database)

Free format text: PAYMENT UNTIL: 20111102

Year of fee payment: 5

FPAY Renewal fee payment (event date is renewal date of database)

Free format text: PAYMENT UNTIL: 20121102

Year of fee payment: 6

FPAY Renewal fee payment (event date is renewal date of database)

Free format text: PAYMENT UNTIL: 20121102

Year of fee payment: 6

FPAY Renewal fee payment (event date is renewal date of database)

Free format text: PAYMENT UNTIL: 20131102

Year of fee payment: 7

LAPS Cancellation because of no payment of annual fees