JPH0431117B2 - - Google Patents
Info
- Publication number
- JPH0431117B2 JPH0431117B2 JP59058171A JP5817184A JPH0431117B2 JP H0431117 B2 JPH0431117 B2 JP H0431117B2 JP 59058171 A JP59058171 A JP 59058171A JP 5817184 A JP5817184 A JP 5817184A JP H0431117 B2 JPH0431117 B2 JP H0431117B2
- Authority
- JP
- Japan
- Prior art keywords
- phoneme
- dictionary
- word
- segmentation
- speech
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Expired
Links
- 230000011218 segmentation Effects 0.000 claims description 22
- 238000000034 method Methods 0.000 claims description 9
- 238000010586 diagram Methods 0.000 description 4
- 230000000694 effects Effects 0.000 description 4
- 238000007796 conventional method Methods 0.000 description 3
- 238000000605 extraction Methods 0.000 description 3
- 230000007704 transition Effects 0.000 description 2
- 239000000284 extract Substances 0.000 description 1
Description
【発明の詳細な説明】
(産業上の利用分野)
本発明は、入力音声と、音素表記された単語辞
書を照合して単語を認識する音声認識方法に関す
るものである。DETAILED DESCRIPTION OF THE INVENTION (Field of Industrial Application) The present invention relates to a speech recognition method for recognizing words by comparing input speech with a word dictionary in which phonemes are expressed.
(従来例の構成とその問題点)
第1図は従来の音声認識方法の一例及び本発明
の音声認識方法の実施例等を実行するための装置
の機能ブロツク図である。従来例を第1図ととも
に説明する。第1図において、1は入力音声から
パラメータの時系列を作成するパラメータ抽出
部、2は音素毎のセグメンテーシヨン、尤度計算
および類似度計算等を行なう単語認識部、3は認
識すべき全単語を音素単位の記号列で表記した単
語辞書が記憶されている単語辞書部である。その
単語辞書は、例えば単語「サツポロ」、「トーキヨ
ー」、「トヨナカ」、「ヤマガタ」はそれぞれ
「SAQPORO」、「TOOKYOO」、
「TOYONAKA」、「JAMAGTA」等と表記され
ている。(Constitution of Conventional Example and Problems thereof) FIG. 1 is a functional block diagram of an apparatus for executing an example of a conventional speech recognition method and an embodiment of the speech recognition method of the present invention. A conventional example will be explained with reference to FIG. In Fig. 1, 1 is a parameter extraction unit that creates a time series of parameters from input speech, 2 is a word recognition unit that performs segmentation for each phoneme, likelihood calculation, similarity calculation, etc., and 3 is a total number of words to be recognized. This is a word dictionary section that stores a word dictionary in which words are expressed as symbol strings in units of phonemes. In the word dictionary, for example, the words "Satsuporo", "Tokyo", "Toyonaka", and "Yamagata" are respectively "SAQPORO" and "TOOKYOO".
It is written as "TOYONAKA", "JAMAGTA", etc.
次に上記従来例の動作について説明する。先ず
入力音声をパラメータ抽出部1で10msのフレー
ム毎に分析し、パラメータを抽出して、パラメー
タ時系列を作成する。パラメータ時系列は、以後
の処理で共通的に用いるパラメータを予め計算し
ておくものである。次に、単語認識部2において
単語辞書部3を照合して各辞書項目毎に類似度を
求めるのであるが、この類似度計算時に、その辞
書項目を構成する辞書音素系列に従つて音素のセ
グメンテーシヨンを行ない、そのセグメンテーシ
ヨンされた音声区間がその音素を発声したもので
ある確からしさを表わす尺度で尤度を計算し、そ
の辞書項目における各音素の尤度の平均値として
類似度を求め、類似度が最大となる辞書項目をも
つて認識単語とする。ここで、ある音素のセグメ
ンテーシヨンを行なうとは具体的には、〔(その音
素の前音素の後端のフレーム番号)+1〕をその
音素の始端フレームとして、そこからその音素の
後端フレームを探して見つけることである。こ
の、ある音素に対しセグメンテーシヨンされる音
声の区間の時間長は、自然な発声をする限り当然
一定の範囲内にある。従つて前記の音素の後端フ
レームを探すにあたつては、ある限られた範囲の
みでよい。本従来例においては、この範囲を1〜
30フレーム(10〜300ms)としていたが、実際の
音声認識において、この値は適当であつた。 Next, the operation of the above conventional example will be explained. First, the input audio is analyzed in every 10 ms frame by the parameter extractor 1, parameters are extracted, and a parameter time series is created. The parameter time series is one in which parameters commonly used in subsequent processing are calculated in advance. Next, the word recognition unit 2 compares the word dictionary unit 3 to determine the degree of similarity for each dictionary item. When calculating this degree of similarity, the phoneme segment is determined according to the dictionary phoneme sequence that constitutes the dictionary item. The likelihood is calculated as a measure of the probability that the segmented speech interval is the one that uttered the phoneme, and the similarity is calculated as the average of the likelihoods of each phoneme in the dictionary entry. The dictionary entry with the highest degree of similarity is selected as the recognized word. Specifically, performing segmentation for a certain phoneme means that [(frame number of the rear end of the phoneme before the phoneme) + 1] is set as the start frame of the phoneme, and from there, the rear end frame of the phoneme is segmented. It is about searching and finding. The time length of the segmented speech segment for a certain phoneme is naturally within a certain range as long as the speech is natural. Therefore, when searching for the last frame of the phoneme, only a limited range is required. In this conventional example, this range is from 1 to
It was set at 30 frames (10 to 300 ms), but this value was appropriate for actual speech recognition.
しかしながら上記従来例においては、下記のよ
うな欠点があつた。これの例を第2図とともに説
明する。第2図は、入力音声がTOJONAKA(ト
ヨナカ)である時、時刻を右向きにとつて、辞書
項目TOJONAKAとJAMAGATAとにおけるセ
グメンテーシヨン結果の対応関係を示す図であ
る。この例において、辞書項目TOJONAKAの
場合のセグメンテーシヨンは正しかつた。一方
JAMAGATAの場合のセグメンテーシヨンは、
TOJ−J,A−AGAと2ケ所誤つた対応を含ん
でいたが、尤度計算においては、入力のTOJの
部分をJと見なしてもパラメータ上にむじゆんな
く、またGとセグメンテーシヨンされた区間はA
からKへ移行する発声の不安定な部分であるため
小さなパワデイツプが存在し、しかもパラメータ
がGJしさを示すため高い尤度が得られてしまい、
類似度も大となつた。このため、本例に示す入力
音声は、JAMAGATAであると誤認識されてい
た。本例に示す辞書項目JAMAGATAにおける
セグメンテーシヨンにおいて、Gとセグメンテー
シヨンされた区間は2フレーム、次のAとセグメ
ンテーシヨンされた区間は1フレームのみであつ
た。ある音素をセグメンテーシヨンした時、その
区間の時間長が1,2フレームと短いものは、発
声において、その音素の性質が弱く、その音素と
隣の音素との間の移行部分が、隣の音素の区間に
セグメンテーシヨンされた場合が多く、従つて、
短い時間長のセグメンテーシヨンが連続すぬこと
は実際にはあり得ない。よつて、本従来において
は、第2図に示すJAMAGATAの例のように、
実際にはあり得ないセグメンテーシヨンを行ない
ながら、類似度は大となつて、単語を誤認識する
という欠点があつた。 However, the above conventional example had the following drawbacks. An example of this will be explained in conjunction with FIG. FIG. 2 is a diagram showing the correspondence of segmentation results between dictionary items TOJONAKA and JAMAGATA when the input voice is TOJONAKA (Toyonaka), with time oriented toward the right. In this example, the segmentation for the dictionary entry TOJONAKA was correct. on the other hand
Segmentation in the case of JAMAGATA is
There were two incorrect correspondences, TOJ-J and A-AGA, but in the likelihood calculation, even if the TOJ part of the input was considered as J, there was no problem on the parameters, and it was segmented as G. The section is A
Since this is an unstable part of the vocalization that transitions from
The degree of similarity also increased. For this reason, the input voice shown in this example was erroneously recognized as JAMAGATA. In the segmentation for the dictionary item JAMAGATA shown in this example, the section segmented with G was two frames, and the section segmented with the next A was only one frame. When segmenting a certain phoneme, if the time length of the segment is as short as 1 or 2 frames, the characteristics of that phoneme are weak in the utterance, and the transition between that phoneme and the next phoneme is similar to that of the next phoneme. It is often segmented into phoneme intervals, and therefore,
It is actually impossible for segmentations of short time duration to be continuous. Therefore, in this conventional method, as in the JAMAGATA example shown in Figure 2,
Although the method performed segmentation that would not be possible in reality, the degree of similarity increased, resulting in words being misrecognized.
(発明の目的)
本発明は上記従来例の欠点を除去するものであ
り、上記のように明らかにあり得ないセグメンテ
ーシヨンを排除し、それにより単語認識率を向上
させることを目的とする。(Objective of the Invention) The present invention eliminates the drawbacks of the conventional example described above, and aims to eliminate the clearly impossible segmentation as described above, thereby improving the word recognition rate.
(発明の構成)
本発明は、入力音声を単語辞書の各辞書項目と
照合し、各辞書項目を構成する辞書音素系列に従
い各音素毎に入力音声をセグメンテーシヨンし、
そのセグメンテーシヨンされた音声区間が、その
音素を発声したものである確からしさを示す尺度
である尤度を求め、この尤度の値を用いて各辞書
項目と入力音声の類似度を求めて入力単語を認識
するにあたり、前記目的を達成するために、音素
のセグメンテーシヨンにおいて、その音素の区間
の時間長に、その音素の1つ、又はそれ以上前の
音素の時間長を加えて得られた2音素又はそれ以
上の音素の時間長に対し、長過ぎ又は短過ぎの制
限を行ない、明らかに正しくないセグメンテーシ
ヨンを排除し、高い単語認識率を得る効果を得る
ものである。(Structure of the Invention) The present invention collates input speech with each dictionary entry in a word dictionary, and segments the input speech for each phoneme according to the dictionary phoneme series that constitutes each dictionary entry.
The likelihood, which is a measure of the certainty that the segmented speech interval is the one that uttered the phoneme, is calculated, and this likelihood value is used to calculate the degree of similarity between each dictionary item and the input speech. In recognizing an input word, in order to achieve the above purpose, in phoneme segmentation, the time length of the interval of the phoneme is obtained by adding the time length of one or more previous phonemes. This method limits the time length of two or more phonemes to be too long or too short, eliminates clearly incorrect segmentation, and obtains the effect of obtaining a high word recognition rate.
(実施例の説明)
以下に発明の一実施例について、図面とともに
説明する。本実施例の方法を実施するための装置
の基本構成は、前記従来例と同様に、第1図に示
される。第1図において、単語辞書は前記従来例
と同様である。(Description of Embodiment) An embodiment of the invention will be described below with reference to the drawings. The basic configuration of an apparatus for carrying out the method of this embodiment is shown in FIG. 1, as in the conventional example. In FIG. 1, the word dictionary is the same as in the conventional example.
本実施例の動作について説明する。先ず、パラ
メータ抽出部1において、入力音声を10msのフ
レーム毎に分析し、パラメータを抽出してパラメ
ータ時系列を作成する。ここ迄は前記従来例と同
様である。次にこれを単語辞書部2内の単語辞書
と照合し、各辞書項目毎に、その辞書項目を構成
する辞書音素系列に従つて音素のセグメンテーシ
ヨンを行なう。ここで本実施例において、ある音
素の後端を探す範囲を、従来と同様に1〜30フレ
ームに限定すると同時に、1つ前の音素に対しセ
グメンテーシヨンされた区間の時間長と合わせ
て、2音素の時間長がある一定の範囲になるよう
に限定する。例えばGAの場合には5〜44フレー
ムの範囲としている。セグメンテーシヨン後に尤
度計算を行ない類似度を求めることは従来と同様
である。 The operation of this embodiment will be explained. First, the parameter extraction unit 1 analyzes input audio every 10 ms frame, extracts parameters, and creates a parameter time series. The process up to this point is the same as the conventional example. Next, this is compared with the word dictionary in the word dictionary section 2, and phoneme segmentation is performed for each dictionary item according to the dictionary phoneme series that constitutes that dictionary item. Here, in this embodiment, the range for searching for the end of a certain phoneme is limited to 1 to 30 frames as in the conventional case, and at the same time, in conjunction with the time length of the segmented section for the previous phoneme, The time length of two phonemes is limited to a certain range. For example, in the case of GA, the range is 5 to 44 frames. After segmentation, likelihood calculation is performed to obtain similarity as in the conventional method.
本実施例における効果を例とともに述べる。第
2図に示す、前記従来例と同様な入力において、
辞書項目がTOJONAKAの場合、セグメンテー
シヨンは前記従来例と同様、正常になされた。辞
書項目がJAMAGATAの場合、語頭からG迄は
従来と同様なセグメンテーシヨンであつたが、G
が2フレームであるため、次のAの後端は、A長
さが3〜30フレームとなる範囲で探すことにな
り、従来と同様なセグメンテーシヨンはなされな
い。この例において、Aの後端を探す範囲は、A
の次のKの区間の無音部分(Kの破裂の前の閉鎖
区間)にかかつてしまい、Aのセグメンテーシヨ
ンは不能となり、JAMAGATAは入力単語では
あり得ないとい判断がなされた。これにより入力
は、正しくTOJONAKAと認識された。このよ
うに本実施例においては、明らかに正しくないセ
グメンテーシヨンを排除することにより、単語の
誤認識を減少させることができる利点がある。 The effects of this embodiment will be described with examples. In the same input as the conventional example shown in FIG.
When the dictionary entry was TOJONAKA, segmentation was performed normally as in the conventional example. When the dictionary entry was JAMAGATA, the segmentation from the beginning of the word to G was the same as before, but G
Since the length of A is 2 frames, the rear end of the next A is searched within the range where the length of A is 3 to 30 frames, and segmentation as in the conventional method is not performed. In this example, the range to search for the trailing edge of A is
It was found that the segmentation of A was impossible, and it was determined that JAMAGATA could not be an input word. As a result, the input was correctly recognized as TOJONAKA. As described above, this embodiment has the advantage of being able to reduce misrecognition of words by eliminating clearly incorrect segmentations.
なお本実施例では、1単語のみを発声した入力
単語の例を示したが、連続単語、文章中の単語に
おいても全く同様の効果がある。 Although this embodiment shows an example of an input word in which only one word is uttered, the same effect can be obtained for continuous words or words in a sentence.
本発明は上記のような構成であり、以下に示す
効果が得られるものである。 The present invention has the above-described configuration, and provides the following effects.
音素のセグメンテーシヨン時に、その音素の区
間の時間長に、その音素の1つ、又はそれ以上前
の音素の時間長を加えて得られた2音素、又はそ
れ以上の音素の時間長に対し、長過ぎ、又は短過
ぎの制限を行ない、その音素の後端位置を限定す
ることにより、実際にはあり得ない、正しくない
セグメンテーシヨンを排除して、単語の誤認識を
減少させ、単語認識率を向上させることができ
る。 During phoneme segmentation, for the time length of two or more phonemes obtained by adding the time length of the phoneme interval to the time length of one or more previous phonemes. , too long or too short, and by limiting the rear end position of the phoneme, we can eliminate incorrect segmentation that is actually impossible, reduce misrecognition of words, and The recognition rate can be improved.
第1図は従来例、及び本発明の実施例における
音声認識方法を実施するための装置の基本的構成
を示す図。第2図は、従来例における、セグメン
テーシヨンの説明図である。
1……パラメータ抽出部、2……単語認識部、
3……単語辞書部。
FIG. 1 is a diagram showing the basic configuration of an apparatus for implementing a speech recognition method in a conventional example and an embodiment of the present invention. FIG. 2 is an explanatory diagram of segmentation in a conventional example. 1...Parameter extraction unit, 2...Word recognition unit,
3...Word dictionary section.
Claims (1)
た単語辞書の各辞書項目とを照合し、各辞書項目
を構成する辞書音素系列に従い、各一音素毎に入
力音声をセグメンテーシヨンし、そのセグメンテ
ーシヨンされた音声の区間がその音素を発声した
ものである確からしさを示す尺度である尤度を計
算し、この尤度の値を用いて各辞書項目と入力音
声の類似度を求めて入力単語を認識するにあた
り、音素のセグメンテーシヨン時に、その音素の
区間の時間長に、その音素の1つ、又はそれ以上
前の音素の時間長を加えて得られた2音素又はそ
れ以上の音素の時間長に対し、長過ぎ又は短過ぎ
の制限を行ない、その音素の後端位置を限定する
ことを特徴とする音声認識方法。1. The input speech is compared with each dictionary entry in a word dictionary in which the word to be recognized is expressed in phonemes, and the input speech is segmented for each phoneme according to the dictionary phoneme series that constitutes each dictionary entry. The likelihood, which is a measure of the probability that the segmented speech segment is the one that uttered the phoneme, is calculated, and this likelihood value is used to determine the degree of similarity between each dictionary entry and the input speech. When recognizing an input word, during phoneme segmentation, two or more phonemes are obtained by adding the time length of the interval of that phoneme to the time length of one or more previous phonemes. A speech recognition method characterized by limiting the time length of a phoneme to be too long or too short, and limiting the rear end position of the phoneme.
Priority Applications (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP59058171A JPS60202492A (en) | 1984-03-28 | 1984-03-28 | Voice recognition |
Applications Claiming Priority (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP59058171A JPS60202492A (en) | 1984-03-28 | 1984-03-28 | Voice recognition |
Publications (2)
| Publication Number | Publication Date |
|---|---|
| JPS60202492A JPS60202492A (en) | 1985-10-12 |
| JPH0431117B2 true JPH0431117B2 (en) | 1992-05-25 |
Family
ID=13076548
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| JP59058171A Granted JPS60202492A (en) | 1984-03-28 | 1984-03-28 | Voice recognition |
Country Status (1)
| Country | Link |
|---|---|
| JP (1) | JPS60202492A (en) |
Families Citing this family (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JPS617894A (en) * | 1984-06-22 | 1986-01-14 | 松下通信工業株式会社 | Voice recognition |
-
1984
- 1984-03-28 JP JP59058171A patent/JPS60202492A/en active Granted
Also Published As
| Publication number | Publication date |
|---|---|
| JPS60202492A (en) | 1985-10-12 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| JPH11191000A (en) | Method for aligning text and voice signal | |
| JPS62217295A (en) | Voice recognition system | |
| JPH10254475A (en) | Voice recognition method | |
| JPS60202492A (en) | Voice recognition | |
| JPH0431118B2 (en) | ||
| JP3291073B2 (en) | Voice recognition method | |
| KR20040092572A (en) | Speech recognition method of processing silence model in a continous speech recognition system | |
| JPH0431116B2 (en) | ||
| JPH0412479B2 (en) | ||
| JPH0458636B2 (en) | ||
| Elghonemy et al. | Speaker independent isolated Arabic word recognition system | |
| JPH0695684A (en) | Sound recognizing system | |
| JPH05303391A (en) | Speech recognition device | |
| JPH0431114B2 (en) | ||
| JPH0155477B2 (en) | ||
| JPH045392B2 (en) | ||
| JPS58159598A (en) | Monosyllabic voice recognition system | |
| JPH09274496A (en) | Speech recognition device | |
| JPS60149099A (en) | Voice recognition | |
| JPH045395B2 (en) | ||
| JPH0469959B2 (en) | ||
| JPH0412480B2 (en) | ||
| Gao et al. | Telephone Conversation Speaker Recogniton System Based on Speech Purify | |
| JPS6147992A (en) | Voice recognition system | |
| JPH045391B2 (en) |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| EXPY | Cancellation because of completion of term |