JPH0431118B2 - - Google Patents
Info
- Publication number
- JPH0431118B2 JPH0431118B2 JP59058172A JP5817284A JPH0431118B2 JP H0431118 B2 JPH0431118 B2 JP H0431118B2 JP 59058172 A JP59058172 A JP 59058172A JP 5817284 A JP5817284 A JP 5817284A JP H0431118 B2 JPH0431118 B2 JP H0431118B2
- Authority
- JP
- Japan
- Prior art keywords
- phoneme
- dictionary
- likelihood
- word
- phonemes
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Expired
Links
- 238000000034 method Methods 0.000 claims description 8
- 230000011218 segmentation Effects 0.000 description 21
- 238000010586 diagram Methods 0.000 description 4
- 238000000605 extraction Methods 0.000 description 4
- 238000007796 conventional method Methods 0.000 description 1
- 239000000284 extract Substances 0.000 description 1
- 230000007704 transition Effects 0.000 description 1
Description
(産業上の利用分野)
本発明は、入力音声と、音素表記された単語辞
書を照合して単語を認識する音声認識方法に関す
るものである。
(従来例の構成とその問題点)
第1図は従来の音声認識方法の一例及び本発明
の音声認識方法の実施例等を実行するための装置
の機能ブロツク図である。従来例を第1図ととも
に説明する。第1図において、1は入力音声から
パラメータの時系列を作成するパラメータ抽出
部、2は音素毎のセグメンテーシヨン、尤度計算
および類似度計算等を行なう単語認識部、3は認
識すべき全単語を音素単位の記号列で表記した単
語辞書が記憶されている単語辞書部である。その
単語辞書は、例えば単語「サツポロ」、「トーキヨ
ー」、「トヨナカ」、「ヤマカタ」はそれぞれ
「SAQPORO」、「TOOKJOO」、
「TOJONAKA」、「JAMAGATA」等と表記さ
れている。
次に上記従来例の動作について説明する。先ず
入力音声をパラメータ抽出部1で10msのフレー
ム毎に分析し、パラメータを抽出して、パラメー
タ時系列を作成する。パラメータ時系列は、以後
の処理で通共的に用いるパラメータを予め計算し
ておくものである。次に単語認識部2において単
語辞書部3を照合して各辞書項目毎に類似度を求
めるのであるが、この類似度計時時に、その辞書
項目を構成する辞書音素系列に従つて音素のセグ
メンテーシヨンを行ない、そのセグメンテーシヨ
ンされた音声区間がその音素を発声したものであ
る確からしさを表わす尺度である尤度を計算し、
その辞書項目における各音素の尤度の平均値とし
て類似度を求め、類似度が最大となる辞書項目を
もつて認識単語とする。本従来例において、辞書
音素系列におけるi番目の音素の尤度liは次式で
計算される。
li=li0−lip1 ……
ここでli0は、セグメンテーシヨンされた区間中
の各フレームにおけるパラメータが、予め用意さ
れたその音素の標準パタンにどれだけ合致するか
を表わす尺度として5000点満点で計算される。ま
たlip1は、セグメンテーシヨンされた区間の時間
長が、予め定められたスレツシヨルド以下の時は
lip1=0、スレツシヨルドを越えた場合合には1
から5000間の値となり尤度に対する減点となる。
しかしながら上記従来例においては、下記のよ
うな欠点があつた。この例を第2図とともに説明
する。第2図は入力音声がTOJONAKA(トヨナ
カ)である時、時刻を右向きにとつて、辞書項目
TOJONAKAとTAMAGATAとにおけるセグメ
ンテーシヨン結果の対応関係を示す図である。こ
の例において、辞書項目TOJONAKAの場合の
セグメンテーシヨン、尤度計算は正常になされ
た。一方、JAMAGATAの場合のセグメンテー
シヨンは、TOJ−−J、A−AGAと、2ケ所誤
つた対応を含んでいたが、尤度計算においては、
入力のTOJの部分をJと見なしてもパラメータ
上むじゆんなく、またGとセグメンテーシヨンさ
れた区間はAからKへ移行する音声の不安定な部
分であるため小さなパワデイツプが存在するため
にGとセグメンテーシヨンされたのであるが、し
かもパラメータがGの標準パタンと良く合致し、
またlip1の減点も無かつたため高い尤度が得られ
てしまい、類似度も大となつた。このため、本例
に示す入力音声はJAMAGATAであると誤認識
されていた。本例に示す辞書項目JAMAGAHA
におけるセグメンテーシヨンにおいて、Gとセグ
メンテーシヨンされた区間は2フレーム、次のA
とセグメンテーシヨンされた区間は1フレームの
みであつた。ある音素をセグメンテーシヨンした
時、その区間の時間長が1、2フレームと短いも
のは、発声において、その音素の性質が弱く、そ
の音素と隣の音素との間の移行部分が隣の音素の
区間にセグメンテーシヨンされた場合が多く、従
つて、短い時間長のセグメンテーシヨンが連続す
ることは少い。よつて、本従来例においては、第
2図に示すJAMAGATAの例のように、2音素
GAで3フレームという実際にはあり得ないセグ
メンテーシヨンを行ないながら、類似度は大とな
つて、単語を誤認識するという欠点があつた。ま
た上記とは反対の理由で、非常に長い時間のセグ
メンテーシヨンが連続することも少いが、そのよ
うな場合に対しても、本従来例では類似度が大と
なることがあるという欠点があつた。
(発明の目的)
本発明は上記従来例の欠点を除去するものであ
り、上記のようにあり得ない、あるいはまれなセ
グメンテーシヨンが生じた時、尤度に減点を加
え、それにより単語認識率を向上させることを目
的とする。
(発明の構成)
本発明は、入力音声を単語辞書の各辞書項目と
照合し、各辞書項目を構成する辞書音素系列に従
い各音素毎に入力音声をセグメンテーシヨンし、
そのセグメンテーシヨンされた音声区間が、その
音素を発声したものである確からしさを示す尺度
である尤度を求め、この尤度の値を用いて各辞書
項目と入力音声の類似度を求めて入力単語を認識
するにあたり、前記目的を達成するために、音素
の尤度計算において、その音素に対応してセグメ
ンテーシヨンされた区間の時間長に、その音素の
1つ、又はそれ以上前の音素の時間長を加えて得
られた2音素又はそれ以上の音素の時間長に対
し、長過ぎ又は短過ぎのスレツシヨルドを設け、
スレツシヨルドの範囲外の場合には尤度を減点す
ることにより、セグメンテーシヨンの確からしさ
を類似度に反映させ、高い単語認識率を得る効果
を得るものである。
(実施例の説明)
以下に本発明の一実施例について、図面ととも
に説明する。本実施例の方法を実施するための基
本構成は、前記従来例と同様に、第1図により示
される。第1図において、単語辞書は前記従来例
と同様である。
本実施例の動作について説明する。先ずパラメ
ータ抽出部1において、入力音声を10msのフレ
ーム毎に分析し、パラメータを抽出してパラメー
タ時系列を作成する。ここ迄は前記従来例と同様
である。次にこれを単語辞書部3の単語辞書と照
合し、各辞書項目毎に、単語認識部2においてそ
の辞書項目を構成する辞書音素系列に従つて音素
のセグメンテーシヨンを行ない、そのセグメンテ
ーシヨンされた音声区間がその音素を発声したも
のである確からしさを表わす尺度である尤度を計
算する。ここで本実施例において、辞書音素系列
におけるi番目の音素の尤度liは次式で計算され
る。
li=li0−lip1−lip2 ……
ここで、li0,lip1は前記従来例の式における
ものと同様である。lip2は、このi番目の音素に
対しセグメンテーシヨンされた区間の時間長τiと
(i−1)番目の音素の時間長τi-1の和が、予め
定められたスレツシヨルドの範囲内の場合には
lip2=0、範尉外の場合には1〜5000の間の値と
なり、尤度に対する減点となる。例えば(j−
1)番目の音素がG、i番目の音素がAの場合、
スレツシヨルドとlip2の関係を第1表に示す。第
1表において、スレツシヨルドは長過ぎ、短過ぎ
とも2段階とし、lip2も夫々に応じ、値を変えて
いる。
(Industrial Application Field) The present invention relates to a speech recognition method for recognizing words by comparing input speech with a word dictionary in which phonemes are expressed. (Constitution of Conventional Example and Problems thereof) FIG. 1 is a functional block diagram of an apparatus for executing an example of a conventional speech recognition method and an embodiment of the speech recognition method of the present invention. A conventional example will be explained with reference to FIG. In Fig. 1, 1 is a parameter extraction unit that creates a time series of parameters from input speech, 2 is a word recognition unit that performs segmentation for each phoneme, likelihood calculation, similarity calculation, etc., and 3 is a total number of words to be recognized. This is a word dictionary section that stores a word dictionary in which words are expressed as symbol strings in units of phonemes. For example, the words "Satsuporo", "Tokyo", "Toyonaka", and "Yamakata" are written as "SAQPORO", "TOOKJOO", and "Yamakata" respectively.
It is written as "TOJONAKA", "JAMAGATA", etc. Next, the operation of the above conventional example will be explained. First, the input audio is analyzed by the parameter extraction unit 1 every 10 ms frame, parameters are extracted, and a parameter time series is created. The parameter time series is one in which parameters commonly used in subsequent processing are calculated in advance. Next, the word recognition unit 2 compares the word dictionary unit 3 to determine the degree of similarity for each dictionary item. When measuring this similarity, the phoneme segmentation is performed according to the dictionary phoneme sequence that constitutes the dictionary item. calculate the likelihood, which is a measure of the probability that the segmented speech interval is the one that uttered the phoneme,
The degree of similarity is determined as the average value of the likelihood of each phoneme in the dictionary item, and the dictionary item with the maximum degree of similarity is determined as a recognized word. In this conventional example, the likelihood l i of the i-th phoneme in the dictionary phoneme sequence is calculated by the following equation. l i = l i0 − l ip1 ... Here, l i0 is 5000 as a measure of how well the parameters in each frame in the segmented section match the standard pattern of the phoneme prepared in advance. It is calculated on a full point basis. In addition, l ip1 is used when the time length of the segmented section is less than a predetermined threshold.
l ip1 = 0, 1 if threshold is exceeded
It becomes a value between 5000 and 5000, and points are deducted from the likelihood. However, the above conventional example had the following drawbacks. This example will be explained with reference to FIG. Figure 2 shows when the input voice is TOJONAKA, the time is set to the right, and the dictionary entry is
FIG. 3 is a diagram showing the correspondence of segmentation results between TOJONAKA and TAMAGATA. In this example, the segmentation and likelihood calculation for the dictionary item TOJONAKA were successfully performed. On the other hand, the segmentation in the case of JAMAGATA included two incorrect correspondences, TOJ--J and A-AGA, but in the likelihood calculation,
Even if the TOJ part of the input is regarded as J, there is no problem in terms of parameters, and since the section segmented as G is an unstable part of the voice transitioning from A to K, there is a small power dip, so Moreover, the parameters matched well with the standard pattern of G,
In addition, since there was no deduction for l ip1 , a high likelihood was obtained, and the degree of similarity was also large. For this reason, the input voice shown in this example was erroneously recognized as JAMAGATA. Dictionary entry JAMAGAHA shown in this example
In the segmentation in G, the segmented section is 2 frames, and the next A
The segmented section was only one frame. When segmenting a certain phoneme, if the time length of the segment is as short as 1 or 2 frames, the characteristics of that phoneme are weak in utterance, and the transition between that phoneme and the next phoneme is similar to that of the next phoneme. In many cases, the segmentation is performed in intervals of 1 to 2, and therefore, segmentations of short time length are rarely continuous. Therefore, in this conventional example, as in the JAMAGATA example shown in Figure 2, two phonemes are used.
Although the GA performed segmentation using three frames, which is actually impossible, the degree of similarity became large, resulting in words being misrecognized. Furthermore, for the opposite reason to the above, it is rare for segmentations to continue for a very long time, but even in such cases, this conventional example has the disadvantage that the degree of similarity may be large. It was hot. (Objective of the Invention) The present invention eliminates the drawbacks of the above-mentioned conventional examples, and when an improbable or rare segmentation occurs as described above, it adds deduction points to the likelihood, thereby making it possible to recognize words. The aim is to improve the rate. (Structure of the Invention) The present invention collates input speech with each dictionary entry in a word dictionary, and segments the input speech for each phoneme according to the dictionary phoneme series that constitutes each dictionary entry.
The likelihood, which is a measure of the certainty that the segmented speech interval is the one that uttered the phoneme, is calculated, and this likelihood value is used to calculate the degree of similarity between each dictionary item and the input speech. In order to achieve the above objective in recognizing an input word, in calculating the likelihood of a phoneme, one or more previous phonemes are added to the time length of the segmented interval corresponding to that phoneme. Setting a too long or too short threshold for the time length of two or more phonemes obtained by adding the time lengths of the phonemes,
By subtracting likelihood points when the threshold is outside the range, the certainty of segmentation is reflected in the similarity degree, thereby achieving the effect of obtaining a high word recognition rate. (Description of Embodiment) An embodiment of the present invention will be described below with reference to the drawings. The basic configuration for carrying out the method of this embodiment is shown in FIG. 1, as in the conventional example. In FIG. 1, the word dictionary is the same as in the conventional example. The operation of this embodiment will be explained. First, the parameter extraction unit 1 analyzes input audio every 10 ms frame, extracts parameters, and creates a parameter time series. The process up to this point is the same as the conventional example. Next, this is compared with the word dictionary in the word dictionary unit 3, and for each dictionary item, phoneme segmentation is performed in the word recognition unit 2 according to the dictionary phoneme series that constitutes that dictionary item. The likelihood is calculated, which is a measure of the probability that the phoneme was uttered in the voiced segment. In this embodiment, the likelihood l i of the i-th phoneme in the dictionary phoneme sequence is calculated by the following equation. l i =l i0 -l ip1 -l ip2 ... Here, l i0 and l ip1 are the same as in the formula of the conventional example. l ip2 is the sum of the time length τ i of the segmented interval for this i-th phoneme and the time length τ i-1 of the (i-1)th phoneme is within a predetermined threshold. In Case of
If l ip2 = 0, the value will be between 1 and 5000, and points will be deducted from the likelihood. For example (j−
1) If the th phoneme is G and the ith phoneme is A,
Table 1 shows the relationship between the threshold and l ip2 . In Table 1, the threshold is set to two levels, one for too long and one for too short, and the value of l ip2 is changed accordingly.
【表】
即わち、実際にはありえないような(τi+τi-1)
に対しては、尤度を大きく感じて、そのようなセ
グメンテーシヨンにより大きな類似度が得られな
いようにしている。このようにして尤度計算を行
なつた後、類似度を求めることは従来と同様であ
る。
本実施例における効果を例とともに述べる。第
2図に示す。前記従来例と同様な入力において、
辞書項目がTOJONAKAの場合、セグメンテー
シヨン及び尤度計算は、前記従来例と同様、正常
になされた。またTOJONAKAの全ての音素に
対しlip2=0であつた。辞書項目がJAMAGATA
の場合、語頭からG(i=5)迄のセグメンテー
シヨン、尤度計算、及び次のA(i=6)のセグ
メンテーシヨンは従来と同様になされたが、前記
の通りτ5=2フレーム、τ6=1フレームであるの
で表1に示すようにl6p2=4000となり、このAの
尤度は従来より4000点低いものとなつた。その結
果JAMAGATAの類似度も小さなものとなり、
入力は正しくTOJONAKAと認識された。この
ように本実施例においては、明らかに正しくな
い、あるいはまれにしか存在しないようなセグメ
ンテーシヨンがなされた時、その音素の尤度に減
点を加えて類似度を減少させ、そのようなセグメ
ンテーシヨンがなされた辞書項目を認識単語とす
る可能性を小さくし、単語の誤認識を減少させる
ことができる利点がある。
なお本実施例では、1単語のみを発声した入力
単語の例を示したが連続単語、文章中の単語にお
いても全く同様の効果がある。
(発明の効果)
本発明は上記のような構成であり、以下に示す
効果が得られるものである。
音素の尤度計算時に、その音素に対応してセグ
メンテーシヨンされた区間の時間長に、その音素
の1つ、又はそれ以上前の音素の時間長を加えて
得られた2音素、又はそれ以上の音素の時間長に
対し、長過ぎ、又は短過ぎのスレツシヨルドを設
け、そのスレツシヨルドの範囲外となつた時には
尤度の値を感じ、そのような実際にはありえな
い、あるいはまれなセグメンテーシヨンにより大
きな類似度が得られないように単語の誤認識を減
少させ、単語認識率を向上させることができる。[Table] In other words, (τ i + τ i-1 ) that is actually impossible
, the likelihood is felt to be large, and such segmentation prevents a large degree of similarity from being obtained. After performing the likelihood calculation in this manner, the similarity is determined in the same way as in the conventional method. The effects of this embodiment will be described with examples. Shown in Figure 2. In the same input as the conventional example,
When the dictionary item was TOJONAKA, segmentation and likelihood calculation were performed normally as in the conventional example. Furthermore, l ip2 = 0 for all phonemes in TOJONAKA. Dictionary entry is JAMAGATA
In the case of , the segmentation from the beginning of the word to G (i=5), the likelihood calculation, and the segmentation of the next A (i=6) were performed in the same way as before, but as mentioned above, τ 5 =2 frame, τ 6 = 1 frame, so l 6p2 = 4000 as shown in Table 1, and the likelihood of A is 4000 points lower than before. As a result, the similarity of JAMAGATA is also small,
The input was correctly recognized as TOJONAKA. In this example, when a segmentation that is clearly incorrect or rarely exists, points are deducted from the likelihood of the phoneme to reduce the similarity, and such segmentation is performed. This method has the advantage that it is possible to reduce the possibility that a dictionary item that has been edited is used as a recognized word, and to reduce misrecognition of words. In this embodiment, an example of an input word in which only one word is uttered is shown, but the same effect can be obtained for continuous words or words in a sentence. (Effects of the Invention) The present invention has the above-described configuration, and provides the following effects. When calculating the likelihood of a phoneme, the two phonemes obtained by adding the time length of one or more preceding phonemes to the time length of the segmented interval corresponding to that phoneme, or two phonemes. We set a threshold that is too long or too short for the above phoneme time length, and when it falls outside the range of the threshold, we feel the value of the likelihood, and such segmentation is impossible or rare. It is possible to reduce misrecognition of words so that a greater degree of similarity is not obtained, and improve the word recognition rate.
第1図は従来例、及び本発明の実施例における
音声認識方法を実施するための装置の基本的構成
を示す図。第2図は従来例、及び本発明の実施例
におけるセグメンテーシヨンを示す図である。
1……パラメータ抽出部、2……単語認識部、
3……単語辞書部。
FIG. 1 is a diagram showing the basic configuration of an apparatus for implementing a speech recognition method in a conventional example and an embodiment of the present invention. FIG. 2 is a diagram showing segmentation in a conventional example and an embodiment of the present invention. 1...Parameter extraction unit, 2...Word recognition unit,
3...Word dictionary section.
Claims (1)
た単語辞書の各辞書項目と照合し、各辞書項目を
構成する辞書音素系列に従い、各一音素毎に入力
音声をセグメンテーシヨンし、そのセグメンテー
シヨンされた音声の区間がその音素を発声したも
のである確からしさを示す尺度である尤度を計算
し、この尤度の値を用いて各辞書項目と入力音声
の類似度を求めて入力単語を認識するにあたり、
音素の尤度計算時に、その音素の区間の時間長
に、その音素の1つ又はそれ以上前の音素の時間
長を加えて得られた2音素又はそれ以上の音素の
時間長に対し、長過ぎ又は短過ぎのスレツシヨル
ドを設け、そのスレツシヨルドの範囲を越えた場
合には前記音素の尤度を減ずることを特徴とする
音声認識方法。1. The input speech is checked against each dictionary entry in a word dictionary in which the word to be recognized is expressed in phonemes, and the input speech is segmented for each phoneme according to the dictionary phoneme series that constitutes each dictionary entry. The likelihood, which is a measure of the probability that the segment of voiced speech is the one that uttered the phoneme, is calculated, and this likelihood value is used to calculate the similarity between each dictionary item and the input speech and input it. In recognizing words,
When calculating the likelihood of a phoneme, calculate the length for the time length of two or more phonemes obtained by adding the time length of the phoneme interval to the time length of one or more phonemes before the phoneme. A speech recognition method characterized by setting a threshold that is too high or too short, and reducing the likelihood of the phoneme when the range of the threshold is exceeded.
Priority Applications (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP59058172A JPS60202493A (en) | 1984-03-28 | 1984-03-28 | Voice recognition |
Applications Claiming Priority (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP59058172A JPS60202493A (en) | 1984-03-28 | 1984-03-28 | Voice recognition |
Publications (2)
| Publication Number | Publication Date |
|---|---|
| JPS60202493A JPS60202493A (en) | 1985-10-12 |
| JPH0431118B2 true JPH0431118B2 (en) | 1992-05-25 |
Family
ID=13076576
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| JP59058172A Granted JPS60202493A (en) | 1984-03-28 | 1984-03-28 | Voice recognition |
Country Status (1)
| Country | Link |
|---|---|
| JP (1) | JPS60202493A (en) |
Families Citing this family (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JPS617894A (en) * | 1984-06-22 | 1986-01-14 | 松下通信工業株式会社 | Voice recognition |
-
1984
- 1984-03-28 JP JP59058172A patent/JPS60202493A/en active Granted
Also Published As
| Publication number | Publication date |
|---|---|
| JPS60202493A (en) | 1985-10-12 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| JPH07146699A (en) | Speech recognition method | |
| JP3069531B2 (en) | Voice recognition method | |
| JPH067357B2 (en) | Voice recognizer | |
| JPS60202493A (en) | Voice recognition | |
| JPH0431117B2 (en) | ||
| JP3291073B2 (en) | Voice recognition method | |
| JPH0458636B2 (en) | ||
| JPS6147999A (en) | Voice recognition system | |
| JP3007357B2 (en) | Dictionary update method for speech recognition device | |
| JPH0469959B2 (en) | ||
| JPS63161499A (en) | Voice recognition equipment | |
| JPH0431116B2 (en) | ||
| JPS60147795A (en) | Voice recognition | |
| JPS62255999A (en) | Word voice recognition equipment | |
| JPS60149099A (en) | Voice recognition | |
| JPS607492A (en) | Monosyllable voice recognition system | |
| JPS58159598A (en) | Monosyllabic voice recognition system | |
| JPH045397B2 (en) | ||
| JPH0431114B2 (en) | ||
| JPH045391B2 (en) | ||
| JPH045400B2 (en) | ||
| JPH0155477B2 (en) | ||
| JPH0413719B2 (en) | ||
| JPH0336439B2 (en) | ||
| JPH045392B2 (en) |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| EXPY | Cancellation because of completion of term |