JPH0568717B2 - - Google Patents
Info
- Publication number
- JPH0568717B2 JPH0568717B2 JP5843784A JP5843784A JPH0568717B2 JP H0568717 B2 JPH0568717 B2 JP H0568717B2 JP 5843784 A JP5843784 A JP 5843784A JP 5843784 A JP5843784 A JP 5843784A JP H0568717 B2 JPH0568717 B2 JP H0568717B2
- Authority
- JP
- Japan
- Prior art keywords
- input
- audio
- frame number
- standard pattern
- dissimilarity
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Expired - Lifetime
Links
- 238000000034 method Methods 0.000 claims description 26
- 238000012545 processing Methods 0.000 description 10
- 230000008569 process Effects 0.000 description 8
- 238000010586 diagram Methods 0.000 description 6
- 238000010606 normalization Methods 0.000 description 5
- 238000001228 spectrum Methods 0.000 description 3
- 238000003708 edge detection Methods 0.000 description 2
- 230000000694 effects Effects 0.000 description 2
- 230000014509 gene expression Effects 0.000 description 2
- 230000003595 spectral effect Effects 0.000 description 2
- 230000008901 benefit Effects 0.000 description 1
- 238000012790 confirmation Methods 0.000 description 1
- 230000008602 contraction Effects 0.000 description 1
- 238000001514 detection method Methods 0.000 description 1
- 230000006872 improvement Effects 0.000 description 1
- 230000035484 reaction time Effects 0.000 description 1
- 238000012827 research and development Methods 0.000 description 1
- 230000004044 response Effects 0.000 description 1
- 230000000717 retained effect Effects 0.000 description 1
- 230000009466 transformation Effects 0.000 description 1
- 210000001260 vocal cord Anatomy 0.000 description 1
Description
【発明の詳細な説明】
(技術分野)
本発明は音声認識方法に関し、具体的には単語
入力音声の終端の確認を待たないで、入力音声の
始端検出から認識動作を開始するようにした音声
認識方法に関する。[Detailed Description of the Invention] (Technical Field) The present invention relates to a speech recognition method, and more specifically, the present invention relates to a speech recognition method, and specifically, a speech recognition method in which a recognition operation is started from detecting the beginning of input speech without waiting for confirmation of the end of word input speech. Regarding recognition methods.
(背景技術)
音声認識方法の一形式として、各標準音声に対
応して周波数成分のフレーム時系列として標準パ
ターンを記憶しておき、入力音声から同じく周波
数成分のフレーム時系列として入力パターンを抽
出し、入力パターンと各標準パターンとの非類似
度を計算し、その非類似度に基づいて入力音声を
識別する方法が知られている。(Background Art) As a form of speech recognition method, a standard pattern is stored as a frame time series of frequency components corresponding to each standard voice, and the input pattern is extracted from the input speech as a frame time series of frequency components. A known method is to calculate the degree of dissimilarity between an input pattern and each standard pattern, and to identify input speech based on the degree of dissimilarity.
一例として、沖研究開発第118号第53頁、昭和
57年12月、に開示されている。このような方法に
おける標準又は入力の音声パターンは、通常、規
則的にフレームを設定して周波数分析し、対数変
換と最小自乗近似直線を用いた声帯振動特性等の
正規化とを経て、周波数成分のフレーム時系列と
して表現したものを用いる。 For example, Oki Research and Development No. 118, page 53, Showa
It was disclosed in December 1957. In such a method, the standard or input speech pattern is usually frequency-analyzed by setting frames regularly, and then subjected to logarithmic transformation and normalization of vocal fold vibration characteristics using a least squares approximation straight line, and then the frequency components are determined. , expressed as a frame time series.
また、入力パターンと各標準パターンとの非類
似度を計算するためにマツチングパスを設定する
方法としては、動的計画法を用いたDPマツチン
グ法と前記文献に見られるような本質的に線形な
マツチング法とが知られている。構成の簡易化の
観点からは、線形マツチング法が有利であるが、
第1図に例示した如く、単語の発声速度は変動が
きわめて大きく、個人差があると共に心理状態や
状況によつても変動し、標準という感覚のもとで
すら20%〜40%の発声長のばらつきが見られ、何
等かの工夫が必要である。線形マツチング法には
種々の形式が提案されているが、前記文献に限ら
ず、そこでは入力音声の終端を検出したのち、マ
ツチングパスを設定していて、認識応答の面では
問題があり、入力パターンも始端から終端を確認
するまで記憶しておく必要がある。 In addition, methods for setting matching paths to calculate the degree of dissimilarity between the input pattern and each standard pattern include the DP matching method using dynamic programming and the essentially linear matching method as seen in the above-mentioned literature. The law is known. From the perspective of simplifying the configuration, the linear matching method is advantageous, but
As illustrated in Figure 1, the utterance speed of a word varies extremely widely, with individual differences as well as changes depending on the psychological state and situation. There is some variation in the results, and some kind of improvement is required. Various types of linear matching methods have been proposed, but not only in the above-mentioned literature, in which a matching path is set after detecting the end of input speech, there are problems in terms of recognition response, and there are It is also necessary to memorize the starting point until the ending point is confirmed.
(発明の目的)
本発明の目的は、入力音声の始端を検出して直
ちにマツチング動作を開始させることによつて認
識速度を高め、且つ発声速度変動を予想してマツ
チングパスを設定することによつて発声速度の変
動を吸収することにある。(Objective of the Invention) An object of the present invention is to increase the recognition speed by detecting the beginning of the input voice and immediately starting the matching operation, and by setting the matching path in anticipation of variations in the speaking rate. The purpose is to absorb fluctuations in speaking speed.
(発明の概要)
本発明の第1の特徴は、入力パターンの各フレ
ーム毎にマツチング処理を行ない、各フレーム毎
に各標準パターンの各マツチングパスに対応した
非類似度を更新記憶するようにしたことにある。(Summary of the Invention) The first feature of the present invention is that matching processing is performed for each frame of the input pattern, and the degree of dissimilarity corresponding to each matching pass of each standard pattern is updated and stored for each frame. It is in.
まず、音声の有音状態の検出は音声パワーを用
いる方法を用いることができる。この場合音声の
始端検出はフレーム電力P(j)(但しjは入力パタ
ーンのフレーム番号)があらかじめ定められた閾
値を越えた時点を始端と考える。但し外部からの
雑音などにより音声入力が行なわれていなくとも
電力P(j)が閾値を越えてしまい、誤つた始端とす
る場合がある。そのため、ともかくフレーム電力
P(j)が閾値を越えたフレームを始端と考え認識処
理を開始するものの連続して3フレーム以上フレ
ーム電力が閾値を越えなければその入力フレーム
を音声の始端とは考えず認識処理を中断し始端検
出のための処理へともどる。但し、フレーム長を
16msecとしている。ここで音声の始端からフレ
ーム電力が閾値を越えたフレームの番号付けを定
義しh番目の音声フレームと称し、単なる入力フ
レーム番号とは区別する。すなわち、音声フレー
ム番号hの音声フレームは、有音区間でh番目の
入力フレームに対応する。 First, a method using audio power can be used to detect the presence of audio. In this case, the start end of the audio is detected when the frame power P(j) (where j is the frame number of the input pattern) exceeds a predetermined threshold. However, due to external noise or the like, the power P(j) may exceed the threshold value even if no voice input is being performed, and the starting point may be incorrectly determined. Therefore, although the frame whose frame power P(j) exceeds the threshold is considered to be the starting point and recognition processing is started, if the frame power exceeds the threshold for three or more consecutive frames, the input frame is not considered to be the starting point of the audio. The recognition process is interrupted and the process returns to the start edge detection process. However, the frame length
It is set to 16 msec. Here, the numbering of the frame whose frame power exceeds the threshold from the start of the audio is defined and is called the h-th audio frame, which is distinguished from a simple input frame number. That is, the audio frame with audio frame number h corresponds to the h-th input frame in the voiced period.
発声速度の正規化を行なうマツチング処理を音
声の始端から開始し、音声分析部出力が得られる
周期(フレーム周期)ごとに行なえれば音声分析
部のデータを始端からすべて格納しておく必要も
なく、また、反応時間も速くなる。 If the matching process that normalizes the speech rate can be started from the beginning of the voice and performed every cycle (frame period) in which the output of the voice analyzer is obtained, there is no need to store all the data from the voice analyzer from the start. , and the reaction time is also faster.
本発明の第2の特徴は発声がおそく行なわれた
場合、標準的に行なわれた場合、はやく行なわれ
た場合を想定したマツチングパスを設定しそれぞ
れのマツチングパス上でのマツチング処理を行な
うことにある。音声の始端検出時点では今から入
力される単語の発声速度は不明である。そこで発
声がおそく行なわれた場合、標準的に行なわれた
場合、はやく行なわれた場合を想定したマツチン
グパスを設定し、それぞれのマツチングパス上で
マツチング処理を行なえば終端検出前からでもマ
ツチング処理が開始可能となる。もちろん、この
場合、入力の終端と標準パターンの終端が一致す
るパスが存在する可能性は少ないが、入力の終端
と標準パターンの終端が最も一致しているパス上
での非類似度が最小となることが予想される。 A second feature of the present invention is to set matching paths assuming cases in which utterances are performed slowly, normally, and quickly, and perform matching processing on each matching path. At the time of detecting the beginning of speech, the speech rate of the word that is about to be input is unknown. Therefore, by setting matching paths assuming that the utterance is performed slowly, normally, or quickly, and performing matching processing on each matching path, matching processing can be started even before the end is detected. becomes. Of course, in this case, there is a small possibility that there exists a path where the end of the input matches the end of the standard pattern, but the dissimilarity on the path where the end of the input most matches the end of the standard pattern is minimal. It is expected that
一方、単語には「イチ」の「イ」と「チ」の間
のように単語内にフレーム電力が閾値に満たない
部分すなわち無音状態のフレームを持つ単語があ
る。のような部分を「パワーデイツプ」と称す
る。このパワーデイツプの長さは単語によつて異
なるが通常30フレーム長を越えることはほとんど
ない。音声の始端を検出後、あるフレーム時間点
においてそのフレーム電力が閾値未満すなわち無
音状態となつた場合、そのフレーム時間点はパワ
ーデイツプの始まりなのか、音声の終端なのかは
判断がつかない。この判定は通常、その時点から
30フレームの間に音声の始端条件(3フレーム以
上連続してフレーム電力が閾値以上)を満足する
フレームが存在するかしないかによつて行なうた
め最大30フレーム後でなければ判断が下されな
い。従つて、フレーム電力が閾値未満となつた場
合のマツチング結果は何らかの形で保留されなけ
ればならない。本発明では、音声フレーム番号の
更新を停止して、フレーム電力が閾値未満すなわ
ち無音状態となつたフレームに対してはマツチン
グ処理を停止することによりこの問題を解決す
る。 On the other hand, some words have a portion where the frame power is less than the threshold, such as between the "i" and "chi" in "ichi", that is, a silent frame. This part is called the "power dip." The length of this power dip varies depending on the word, but it usually rarely exceeds 30 frames. After detecting the start of audio, if the frame power at a certain frame time point is less than the threshold, that is, the frame becomes silent, it is difficult to determine whether that frame time point is the beginning of a power dip or the end of audio. This determination is usually made from that point onwards.
This is done depending on whether there is a frame that satisfies the audio start condition (3 or more frames in a row where the frame power is above the threshold) within 30 frames, so a decision cannot be made until 30 frames have passed at most. Therefore, the matching results when the frame power becomes less than the threshold must be suspended in some way. In the present invention, this problem is solved by stopping the updating of the audio frame number and stopping the matching process for frames whose frame power is less than a threshold value, that is, in a silent state.
第2図aは本発明による音声認識方法における
入力パターンと標準パターンとのマツチングを行
なう複数のマツチングパス例を示した図、第2図
bは入力パターンのフレーム電力例を示した図、
第2図cは入力パターンと標準パターンとの各マ
ツチングパスにおける非類似度Dn(j)、D′n(j)、
D″n(j)の例を示した図である。 FIG. 2a is a diagram showing an example of a plurality of matching passes for matching an input pattern and a standard pattern in the speech recognition method according to the present invention, FIG. 2b is a diagram showing an example of frame power of the input pattern,
Figure 2c shows the dissimilarity Dn(j), D′n(j), in each matching pass between the input pattern and the standard pattern,
FIG. 3 is a diagram showing an example of D″n(j).
第2図aにおいては発声速度の範囲を例えば±
20%と考え、マツチングパスを3本設定した場合
を示している。第2図aにおいて、横軸は入力パ
ターンのフレーム番号を表わす。また、縦軸は標
準パターンのフレーム番号を表わし、n番目の標
準パターンSnを例として考え、そのフレーム長
をSL(n)とする。101は発声を20%遅く発声し
た場合を想定したパス、102は標準的な発声を
想定したパス、103は発声を20%速く想定した
場合のパスを示す。j番目の入力フレームの電力
が閾値以上の場合、3本のパス上での標準パター
ンSnとの距離を次式によつて与える。但し、h
はj番目の入力フレーム番号に対応した音声フレ
ーム番号であり、W(i、j)は入力フレーム番
号がjでチヤンネル番号がi(但し、i=1〜8)
の入力パターンの成分であり、Sn(i、k)はフ
レーム番号がkでチヤンネル番号がiの標準パタ
ーンの成分である。 In Figure 2 a, the range of speech rate is, for example, ±
The figure shows the case where three matching paths are set, assuming that the ratio is 20%. In FIG. 2a, the horizontal axis represents the frame number of the input pattern. Further, the vertical axis represents the frame number of the standard pattern, and taking the nth standard pattern Sn as an example, let its frame length be SL(n). Reference numeral 101 indicates a path assuming that the speech is uttered 20% slower, 102 a path that assumes standard speech, and 103 a path that assumes that the speech is uttered 20% faster. When the power of the j-th input frame is greater than or equal to the threshold, the distance from the standard pattern Sn on the three paths is given by the following equation. However, h
is the audio frame number corresponding to the j-th input frame number, and W(i, j) is when the input frame number is j and the channel number is i (where i = 1 to 8).
Sn (i, k) is a component of the standard pattern with frame number k and channel number i.
パス101に対する距離
dn(j)=8
〓i=1
|W(i、j)−Sn(i、k)|
…第1式
但し
k=〔1/1.2h〕 〔1/1.2h〕≦SL(n)
SL(n) 〔1/1.2h〕>SL(n)
パス102に対する距離
d′n(j)=8
〓i=1
|W(i、j)−Sn(i、k′)|
…第2式
但し
k′=h
SL(n) h≦SL(n)
h>SL(n)
パス103に対する距離
d″n(j)=8
〓i=1
|W(i、j)−Sn(i、k″)|
…第3式
但し
k″=〔1/0.8h〕 〔1/0.8h〕≦SL(n)
SL(n) 〔1/0.8h〕>SL(n)
尚、〔 〕はガウス記号を示す。 Distance to path 101 dn(j)= 8 〓 i=1 |W(i,j)−Sn(i,k)|
...Equation 1, where k = [1/1.2h] [1/1.2h]≦SL(n) SL(n) [1/1.2h]>SL(n) Distance to path 102 d′n(j)= 8 〓 i=1 | W (i, j) − Sn (i, k′) |
...Second formula where k'=h SL(n) h≦SL(n) h>SL(n) Distance to path 103 d″n(j)= 8 〓 i=1 | W(i, j)−Sn (i, k″) |
...Third formula where k″=[1/0.8h] [1/0.8h]≦SL(n) SL(n) [1/0.8h]>SL(n) Note that [ ] indicates a Gaussian symbol.
前記の式によればパス101においては入力パ
ターンのj番目の入力音声フレームと標準パター
ンのk番目のフレームの間の距離計算を行なう。
パス102においては入力パターンj番目の入力
フレームと標準パターンのk′番目のフレームの間
の距離計算を行ない、パス103においては入力
パターンのj番目の入力フレームと標準パターン
のk″番目のフレームの間の距離計算が行なわれ
る。但し、標準パターンのフレーム番号を示す
k、k′、k″はその標準パターンの長さSL(n)より
大きくなる場合にはSL(n)に制限される。 According to the above equation, in path 101, the distance between the j-th input audio frame of the input pattern and the k-th frame of the standard pattern is calculated.
In pass 102, the distance between the j-th input frame of the input pattern and the k'-th frame of the standard pattern is calculated, and in the pass 103, the distance between the j-th input frame of the input pattern and the k''-th frame of the standard pattern is calculated. However, if k, k', k'' indicating the frame number of the standard pattern is larger than the length SL(n) of the standard pattern, it is limited to SL(n).
一方、j番目の入力フレームの電力が閾値未満
の場合、それぞれのパス上での距離dn(j)、d′n
(j)、d″n(j)を強制的に
dn(j)=0 …第4式
d′n(j)=0 …第5式
d″n(j)=0 …第6式
とすることによりフレーム電力が閾値以下の場合
の非類似度計算を事実上加算しない処理を行な
う。またこのために、標準パターンもパワーデイ
ツプ対応のフレームすなわち無音状態に対応する
フレームを除いた形で蓄積する。 On the other hand, if the power of the j-th input frame is less than the threshold, the distances dn(j), d′n on each path
(j), d″n(j) are forced to dn(j)=0…4th equation d′n(j)=0…5th equation d″n(j)=0…6th equation As a result, processing is performed in which the dissimilarity calculation when the frame power is less than the threshold value is not actually added. For this purpose, the standard pattern is also stored in a form excluding frames corresponding to power dips, that is, frames corresponding to the silent state.
このように、パワーデイツプや標準パターンの
終端以後でのマツチングのように、非類似度とし
て重要でないフレームでは距離を0としているけ
れども、本発明では本質的に線形なマツチングで
ある。 In this way, although the distance is set to 0 for frames that are not important as dissimilarities, such as power dips or matching after the end of a standard pattern, the matching is essentially linear in the present invention.
次に入力パターンのj番目の入力フレームまで
の非類似度Dn(j)、D′n(j)、D″n(j)が計算される。 Next, the dissimilarities Dn(j), D′n(j), and D″n(j) of the input pattern up to the j-th input frame are calculated.
パス101の非類似度
Dn(j)=dn(j)+Dn(j−1) …第7式
パス102の非類似度
D′n(j)=d′n(j)+D′n(j−1) …第8式
パス103の非類似度
D″n(j)=d″n(j)+D″n(j−1) …第9式
すなわち、それぞれのパス上でのj番目のフレ
ームの非類似度の算出は各チヤンネルごとの距離
(例えば|W(i、j)−Sn(i、k)|)をチヤン
ネル分、j−1番目のフレームに対する非類似度
値(たとえばDn(j−1))に加えることによつ
て得られる。これらの演算はj番目のフレームの
入力がなされた時点で行なわれる。j番目の入力
フレームに対する非類似度の算出にあたつてはj
番目のフレームの入力パターンデータとそれぞれ
のパスに相当する標準パターンのデータおよび1
フレーム前のj−1番目の入力フレームの非類似
度データのみが必要であつて2フレーム以上前の
入力パターンデータは不必要である。そのため、
終端を検出するまでの入力パターンを格納してお
かなければならない線形伸縮マツチング法に比較
しても記憶領域が小さくなる効果が生じる。 Dissimilarity of path 101 Dn(j)=dn(j)+Dn(j-1) ...Equation 7 Dissimilarity of path 102 D'n(j)=d'n(j)+D'n(j- 1) ...Equation 8 Dissimilarity of path 103 D″n(j)=d″n(j)+D″n(j−1) ...Equation 9 In other words, the j-th frame on each path The dissimilarity is calculated by dividing the distance for each channel (for example, |W (i, j) - Sn (i, k) 1)). These operations are performed when the j-th frame is input. When calculating the dissimilarity for the j-th input frame,
The input pattern data of the th frame, the standard pattern data corresponding to each pass, and 1
Only the dissimilarity data of the j-1th input frame before the frame is necessary, and the input pattern data of two or more frames before is unnecessary. Therefore,
This method also has the effect of reducing the storage area compared to the linear expansion/contraction matching method in which input patterns must be stored until the end is detected.
第2図cは入力パターンと標準パターンとの各
マツチングパターンでの非類似度Dn(j)、D′n(j)、
D″n(j)を示したものであるが、第2図cに見られ
るようにフレーム電力が閾値以下となつたとき距
離値を強制的に0にすることにより非類似度Dn
(j)、D′n(j)、D″n(j)は保持される。従つて、終端
における非類似度と終端から30フレームへだてた
入力フレーム(この時点で初めて終端が検出され
る)における非類似度は等しい。 Figure 2c shows the dissimilarity Dn(j), D′n(j),
D″n(j), but as shown in Figure 2c, by forcing the distance value to 0 when the frame power is below the threshold, the dissimilarity Dn
(j), D′n(j), and D″n(j) are retained. Therefore, the dissimilarity at the end and the input frame 30 frames from the end (the end is detected for the first time at this point) The dissimilarities at are equal.
次に、音声の終端を検出した時点(音声の終端
から30フレーム後)から各標準パターンごとに得
られた非類似度によつてカテゴリーの判定が行な
われる。終端検出時点の入力フレーム番号をJ、
音声フレーム番号をHとするとn番目の標準パタ
ーンに対する各パスの非類似度Dn(J)、D′n(J)、
D″n(J)で与えられる。これらの非類似度の組が標
準パターンの数(Nとする)だけ存在する。これ
らの非類似度を用いてカテゴリー判定を行なう。
すなわち、各標準パターンごとに求められた各非
類似度Dn(J)、D′n(J)、D″n(J)全てに対して最小値
を求める。この最小値を与える標準パターンに付
加されたカテゴリが認識結果となる。 Next, the category is determined based on the degree of dissimilarity obtained for each standard pattern from the time when the end of the audio is detected (30 frames after the end of the audio). The input frame number at the time of end detection is J,
Letting the audio frame number be H, the dissimilarity of each path to the nth standard pattern is Dn(J), D′n(J),
It is given by D″n(J). There are as many sets of these dissimilarities as there are standard patterns (assumed to be N). Category determination is performed using these dissimilarities.
In other words, find the minimum value for all of the dissimilarities Dn(J), D′n(J), and D″n(J) found for each standard pattern. The selected category becomes the recognition result.
(実施例)
第3図は本発明におけるマツチング処理と判定
処理を行なう回路構成を示した一実施例である。
以下、その動作について詳細に説明する。(Embodiment) FIG. 3 is an embodiment showing a circuit configuration for performing matching processing and determination processing in the present invention.
The operation will be explained in detail below.
第3図において、52は始端からの音声フレー
ム数をカウントする音声フレームカウンタで始端
検出時はリセツトパルス50によつてその内容は
0となり以後入力フレームの電力が閾値を越えた
ときに入力されるカウントパルス51によつてカ
ウントアツプ動作を行なう。入力フレームの電力
が閾値未満の場合にはカウントパルス51は付加
されず、音声フレームカウンタ52の出力は保持
される。音声フレームカウンタ52の出力である
音声フレーム番号をhとする。53はマツチング
の際の数本のパスの種類を表わす信号でパスの数
は3本なので0〜2の値をとる。54はh番目の
音声フレームにおいて各パス上で対応する標準パ
ターンのフレーム番号を与えるROMである。
ROM54には〔1/0.8h〕、h、〔1/1.2h〕に相当
する値が格納されている。ROM54の出力をl
とする。55は標準パターンの番号を与える標準
パターン番号信号であつてnとする。標準パター
ンの総数がNのとき0〜N−1の値をとる。56
はn番目の標準パターンに対してその標準パター
ンの長さSL(n)を格納するROMである。57は標
準パターンフレーム番号lを出力するROM54
の出力と標準パターンの長さSL(n)を出力する
ROM56の内容を比較してl≦SL(n)ならば
“1”を、l>SL(n)ならば“0”を出力するコン
パレータである。58はコンパレータ57の出力
が“1”のときはROM54の出力を、コンパレ
ータ57の出力が“0”のときはROM56の出
力を選択するセレクタである。コンパレータ57
とセレクタ58によつて1≦SL(n)ならばlが、
l>SL(n)ならばSL(n)がセレクタ58より出力さ
れる動作が行なわれる。セレクタ58の出力をk
とする。59はチヤンネル番号iを与える信号で
ある。60はチヤンネル番号iとセレクタ58の
出力例えばkと標準パターン番号信号nによつて
アドレツシングされ標準パターンの各成分Sn
(i、k)を出力する標準パターンのメモリであ
る。61はスペクトル正規化を行なつた1フレー
ム分の入力パターンの各成分W(i、j)を格納
しておくメモリでチヤンネル番号信号59によつ
てアドレスが与えられる。62はメモリ61の入
力端子であり、図示しないスペクトル正規化部で
スペクトル正規化された入力データW(i、j)
が入力される。63はメモリ61の出力W(i、
j)と標準パターンROM60の出力Sn(i、k)
の間でコントロール信号CONTによつて以下の
値を出力する演算器である。 In Fig. 3, reference numeral 52 is an audio frame counter that counts the number of audio frames from the start edge, and when the start edge is detected, its content becomes 0 by the reset pulse 50, and is input thereafter when the power of the input frame exceeds the threshold. A count up operation is performed by the count pulse 51. If the power of the input frame is less than the threshold, the count pulse 51 is not added and the output of the audio frame counter 52 is held. Let h be the audio frame number output from the audio frame counter 52. 53 is a signal representing the type of several paths during matching, and since the number of paths is three, it takes a value from 0 to 2. 54 is a ROM that gives frame numbers of the corresponding standard pattern on each path in the h-th audio frame.
The ROM 54 stores values corresponding to [1/0.8h], h, and [1/1.2h]. The output of ROM54 is
shall be. 55 is a standard pattern number signal which gives the number of the standard pattern, and is assumed to be n. When the total number of standard patterns is N, it takes a value from 0 to N-1. 56
is a ROM that stores the length SL(n) of the nth standard pattern. 57 is a ROM 54 that outputs the standard pattern frame number l.
output and the standard pattern length SL(n)
This is a comparator that compares the contents of the ROM 56 and outputs "1" if l≦SL(n), and outputs "0" if l>SL(n). A selector 58 selects the output of the ROM 54 when the output of the comparator 57 is "1", and selects the output of the ROM 56 when the output of the comparator 57 is "0". Comparator 57
According to the selector 58, if 1≦SL(n), then l is
If l>SL(n), an operation is performed in which SL(n) is output from the selector 58. The output of selector 58 is k
shall be. 59 is a signal giving channel number i. 60 is addressed by the channel number i, the output of the selector 58, for example k, and the standard pattern number signal n, and each component Sn of the standard pattern is
This is a standard pattern memory that outputs (i, k). Reference numeral 61 denotes a memory for storing each component W(i, j) of an input pattern for one frame which has undergone spectrum normalization, and is given an address by a channel number signal 59. 62 is an input terminal of the memory 61, which receives input data W(i,j) whose spectrum has been normalized by a spectrum normalization unit (not shown).
is input. 63 is the output W(i,
j) and the output Sn (i, k) of the standard pattern ROM60
This is an arithmetic unit that outputs the following values according to the control signal CONT between
CONT=1のとき
CONT=0のとき |W(i、j)−Sn(i、k)|
0 …第13式
CONT信号はフレーム電力が閾値以上のとき
は“1”を、閾値未満のときは“0”となる信号
である。64は加算器、65はパス信号53と標
準パターン番号信号55の値をアドレスとする
RAMであり非類似度Dn(j)、D′n(j)、D″n(j)が格
納されている。70はRAM65の出力である非
類似度と後で述べるレジスタ71の出力を比較し
て2つの信号を出力する最小値選択回路である。
1つは比較した結果のうち小さい方の値を与える
信号であり、この信号はレジスタ71に格納され
る。もう一方の信号は比較の結果RAM65の出
力の方が小さければ発するクロツクでありレジス
タ72の入力クロツクとなる。レジスタ71は非
類似度の最小値を与えるレジスタであり、フレー
ム周期の始めに非類似度の最大値がセツトされ
る。レジスタ72は、最小値選択回路70の出力
パルスによつて標準パターン番号信号を格納する
レジスタで、非類似度最小値を与える標準パター
ンの番号が格納されている。73は出力端子であ
る。 When CONT = 1 When CONT = 0 |W (i, j) - Sn (i, k) | 0 ...Equation 13 CONT signal is "1" when the frame power is above the threshold, and when it is below the threshold is a signal that becomes "0". 64 is an adder, and 65 is an address that uses the values of the pass signal 53 and standard pattern number signal 55.
It is a RAM and stores dissimilarities Dn(j), D′n(j), and D″n(j). 70 compares the dissimilarity that is the output of the RAM 65 and the output of the register 71, which will be described later. This is a minimum value selection circuit that outputs two signals.
One is a signal that gives the smaller value of the comparison results, and this signal is stored in the register 71. The other signal is a clock that is generated if the output of RAM 65 is smaller as a result of the comparison, and becomes the input clock of register 72. Register 71 is a register that provides the minimum value of dissimilarity, and the maximum value of dissimilarity is set at the beginning of a frame period. The register 72 is a register that stores a standard pattern number signal based on the output pulse of the minimum value selection circuit 70, and stores the number of the standard pattern that gives the minimum value of dissimilarity. 73 is an output terminal.
第3図は以上のごとく構成されており、以下動
作について説明する。 FIG. 3 is constructed as described above, and its operation will be explained below.
各処理はフレーム電力p(j)が閾値以上となつた
時点から開始されるが、3フレーム以上連続して
フレーム電力P(j)が閾値以上でなければ処理はリ
セツトされる。音声の始端フレーム前はカウンタ
52はリセツトパルス50によつてリセツト状態
にある。また、メモリ65の値はすべてリセツト
されている。以後、始端検出後の1フレーム周期
内の処理を順次説明する。但し、説明のため入力
フレーム番号はjとする。j番目の入力フレーム
のフレーム電力が閾値を越えた場合、カウントパ
ルス51がカウンタ52に印加され、カウンタ5
2はカウントアツプし音声フレーム番号hを出力
する。音声フレーム番号に対応する標準パターン
のフレーム番号はROM54とROM56とコン
パレータ57とセレクタ58によつて出力され
る。n番目の標準パターンのk番目のフレームの
iチヤンネルのデータSn(i、k)はROM60
によつて出力される。一方、メモリ61には前段
のスペクトル正規化部(図示せず)より出力され
るj番目の入力フレームのスペクトル正規化後の
入力データW(i、j)が入力端子62より入力
され格納されている。ROM60の出力Sn(i、
k)とメモリ61の出力W(i、j)はチヤンネ
ル番号信号59に同期して出力され演算器63に
おいて第13式に与えられる演算を行なう。演算器
63の出力とメモリ65の間で第1式〜第9式に
相当する演算が実行される。実際は第1式〜第9
式の演算は統合された次の形式で行なわれる。 Each process is started when the frame power p(j) becomes equal to or greater than the threshold value, but the process is reset unless the frame power P(j) is equal to or greater than the threshold value for three or more consecutive frames. The counter 52 is in the reset state by the reset pulse 50 before the start frame of the audio. Also, all values in the memory 65 have been reset. Hereinafter, the processing within one frame period after the start edge detection will be sequentially explained. However, for the sake of explanation, the input frame number is assumed to be j. If the frame power of the jth input frame exceeds the threshold, a count pulse 51 is applied to the counter 52;
2 counts up and outputs the audio frame number h. The standard pattern frame number corresponding to the audio frame number is output by ROM 54, ROM 56, comparator 57, and selector 58. The i-channel data Sn (i, k) of the k-th frame of the n-th standard pattern is stored in the ROM60.
is output by . On the other hand, input data W(i, j) after spectral normalization of the j-th input frame outputted from the previous-stage spectral normalization unit (not shown) is input to the memory 61 from the input terminal 62 and stored therein. There is. ROM60 output Sn(i,
k) and the output W(i, j) of the memory 61 are outputted in synchronization with the channel number signal 59, and the arithmetic unit 63 performs the calculation given by equation (13). Calculations corresponding to the first to ninth expressions are executed between the output of the arithmetic unit 63 and the memory 65. Actually, formulas 1 to 9
The expression operations are performed in the following unified form.
(メモリ65)←(メモリ65)
+|W(i、j)−Sn(i、k)|
0
次に最小値選択回路70によつてレジスタ71
に格納されている非類似度とRAM65より出力
される非類似度のうち小さい方がレジスタ71に
格納される。と同時にRAM65の出力の方が小
さければパルスがレジスタ72に加えられそのと
きの標準パターン番号がレジスタ72に格納され
る。この処理をすべての標準パターンについて行
なえばそのときの最小非類似度を与える標準パタ
ーン番号がレジスタ72に格納されることにな
る。 (Memory 65)←(Memory 65) +|W(i,j)−Sn(i,k)|0 Next, the minimum value selection circuit 70 selects the register 71.
The smaller of the dissimilarity stored in the RAM 65 and the dissimilarity output from the RAM 65 is stored in the register 71. At the same time, if the output of the RAM 65 is smaller, a pulse is applied to the register 72 and the standard pattern number at that time is stored in the register 72. If this process is performed for all standard patterns, the standard pattern number giving the minimum dissimilarity at that time will be stored in the register 72.
以上の処理は1フレーム周期ごとに行なわれ終
端が検出された時点におけるレジスタ72の結果
が最終的な認識結果となり、出力端子73を通し
て出力される。 The above processing is performed for each frame period, and the result of the register 72 at the time when the end is detected becomes the final recognition result, and is outputted through the output terminal 73.
なお標準パターン番号が同じ時でもパス信号が
違えば、パスの各々に対応する非類似度が異なる
ので、前と同じ標準パターン番号がレジスタ72
にセツトされることは超り得る。 Note that even when the standard pattern number is the same, if the path signals are different, the degree of dissimilarity corresponding to each path will be different.
It can be exceeded.
(発明の効果)
本発明は以上説明したように、入力フレームご
とに距離計算、非類似度計算を行なうため終端検
出後、1フレーム以内に認識結果が出る利点があ
る。また、非類似度計算のためには前の入力フレ
ームに対する非類似度値を現入力フレームに対す
る距離値との累算を行なうだけでよく、始端から
終端までの入力データを格納する必要がない。(Effects of the Invention) As described above, the present invention has the advantage that the recognition result is obtained within one frame after the termination is detected because distance calculation and dissimilarity calculation are performed for each input frame. Furthermore, in order to calculate the dissimilarity, it is only necessary to accumulate the dissimilarity value for the previous input frame with the distance value for the current input frame, and there is no need to store input data from the start to the end.
さらに、回路構成の簡易化を目的とした方式で
あるためLSI化が容易であり、ゲート数の少ない
安価な音声認識用LSIチツプを供給すると同時に
汎用マイクロプロセツサのソフト処理によつても
実現され得るものである。 Furthermore, since it is a method aimed at simplifying the circuit configuration, it is easy to implement on an LSI, and it can be realized by providing an inexpensive LSI chip for speech recognition with a small number of gates, and at the same time by using software processing on a general-purpose microprocessor. It's something you get.
第1図は発声長変動を説明するための図、第2
図は本発明のマツチングパスの概要を説明するた
めに示した図、第3図は本発明の一実施例を示す
ブロツク図、
52……音声フレーム番号hのカウンタ、54
……標準パターンのフレーム番号相当のものを発
生させるためのROM、56……標準パターンの
長さを記憶しているROM、57……コンパレー
タ、58……セレクタ、60……標準パターンの
メモリ、61……入力パターンのメモリ、63…
…距離の演算器、64……加算器、65……非類
似度格納用RAM、70……最小値選択回路、7
1……最小非類似度のメモリ、72……認識結果
としての標準パターン番号のメモリ。
Figure 1 is a diagram for explaining utterance length variation, Figure 2
The figure is a diagram shown to explain the outline of the matching path of the present invention, and FIG. 3 is a block diagram showing an embodiment of the present invention. 52...Counter of audio frame number h, 54
...ROM for generating something equivalent to the frame number of the standard pattern, 56...ROM storing the length of the standard pattern, 57...comparator, 58...selector, 60...memory for the standard pattern, 61...Input pattern memory, 63...
... Distance calculator, 64 ... Adder, 65 ... RAM for storing dissimilarity, 70 ... Minimum value selection circuit, 7
1...Memory of minimum dissimilarity, 72...Memory of standard pattern number as recognition result.
Claims (1)
時系列として表現され且つ無音状態に対応するフ
レームを除去した形で表現された標準パターンを
記憶しておき、 (a) 入力音声から周波数成分のフレーム時系列と
して入力パターンを抽出し、 (b) 入力音声の始端を検出して入力パターンのフ
レームの計数を開始し、有音状態を検出してい
る間は音声フレーム番号を順次更新し、一方無
音状態を検出している間は音声フレーム番号の
更新を停止し、入力音声の終端を確認する以前
に再び有音状態を検出すると当該音声フレーム
番号の更新を再開し、 (c) 音声フレーム番号の更新毎に、その音声フレ
ーム番号に本質的に線形な関係で標準パターン
の複数のフレーム番号を発生させることによつ
て入力パターンと各標準パターンとの間に複数
のマツチングパスを設定し、 (d) 音声フレーム番号の更新毎に、前記各マツチ
ングパスで対応づけられたフレーム間で入力パ
ターンと各標準パターンとの距離を計算し、 (e) 入力音声の始端から任意の音声フレーム番号
までの前記マツチングパスに沿つた前記距離の
累算値を非類似度として、音声フレーム番号の
更新毎に、直前の非類似度と当該フレーム番号
での距離とを加算して一旦記憶することによつ
て、各標準パターン毎の各マツチングパスに対
応した非類似度を更新記憶し、 (f) 全ての前記非類似度のうちで最小値を示す非
類似度に対応した標準パターンのコードを前記
音声フレーム番号の更新毎に更新記憶し、 (g) 入力音声の終端を確認した時点で、入力音声
の終端の音声フレーム番号に対応して記憶され
ている最小値を示す非類似度に対応した前記標
準パターンのコードを入力音声のカテゴリとし
て認識することを特徴とした音声認識方法。[Claims] 1. A standard pattern expressed as a frame time series of frequency components corresponding to each standard voice and expressed in a form in which frames corresponding to a silent state are removed is stored, and (a) input Extract the input pattern from the audio as a frame time series of frequency components, (b) Detect the start of the input audio and start counting the frames of the input pattern, and keep the audio frame number while detecting the active state. The audio frame number is updated sequentially, and while a silent state is detected, updating of the audio frame number is stopped, and when a sound state is detected again before the end of the input audio is confirmed, updating of the audio frame number is resumed, ( c) creating multiple matching passes between the input pattern and each standard pattern by generating multiple frame numbers of the standard pattern in an essentially linear relationship to that audio frame number for each update of the audio frame number; (d) Every time the audio frame number is updated, calculate the distance between the input pattern and each standard pattern between the frames associated with each matching pass, and (e) calculate the distance between the input pattern and each standard pattern from the start of the input audio. The accumulated value of the distance along the matching path up to the number is taken as the dissimilarity, and each time the audio frame number is updated, the immediately preceding dissimilarity and the distance at the frame number are added and temporarily stored. Therefore, the degree of dissimilarity corresponding to each matching pass for each standard pattern is updated and stored, and (f) the code of the standard pattern corresponding to the degree of dissimilarity showing the minimum value among all the degrees of dissimilarity is updated and stored in the voice. (g) When the end of the input audio is confirmed, the dissimilarity value corresponding to the minimum value stored corresponding to the audio frame number at the end of the input audio is updated and stored each time the frame number is updated. A speech recognition method characterized by recognizing standard pattern codes as categories of input speech.
Priority Applications (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP5843784A JPS60203993A (en) | 1984-03-28 | 1984-03-28 | Voice recognition |
| US06/716,154 US4868879A (en) | 1984-03-27 | 1985-03-26 | Apparatus and method for recognizing speech |
Applications Claiming Priority (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP5843784A JPS60203993A (en) | 1984-03-28 | 1984-03-28 | Voice recognition |
Publications (2)
| Publication Number | Publication Date |
|---|---|
| JPS60203993A JPS60203993A (en) | 1985-10-15 |
| JPH0568717B2 true JPH0568717B2 (en) | 1993-09-29 |
Family
ID=13084374
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| JP5843784A Granted JPS60203993A (en) | 1984-03-27 | 1984-03-28 | Voice recognition |
Country Status (1)
| Country | Link |
|---|---|
| JP (1) | JPS60203993A (en) |
Families Citing this family (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JPS61252593A (en) * | 1985-05-02 | 1986-11-10 | 沖電気工業株式会社 | Voice recognition equipment |
-
1984
- 1984-03-28 JP JP5843784A patent/JPS60203993A/en active Granted
Also Published As
| Publication number | Publication date |
|---|---|
| JPS60203993A (en) | 1985-10-15 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US5621849A (en) | Voice recognizing method and apparatus | |
| JPS58134699A (en) | Continuous word string recognition method and apparatus | |
| US6029130A (en) | Integrated endpoint detection for improved speech recognition method and system | |
| US4868879A (en) | Apparatus and method for recognizing speech | |
| JPH0247760B2 (en) | ||
| US5974381A (en) | Method and system for efficiently avoiding partial matching in voice recognition | |
| JPH0568716B2 (en) | ||
| JP3477751B2 (en) | Continuous word speech recognition device | |
| JPH0313600B2 (en) | ||
| JPS61133994A (en) | Voice recognition | |
| JPS60203993A (en) | Voice recognition | |
| JPS61170799A (en) | Voice recognition | |
| JPH0567037B2 (en) | ||
| JP3007357B2 (en) | Dictionary update method for speech recognition device | |
| JPH0313599B2 (en) | ||
| JPH0449954B2 (en) | ||
| JPS6344699A (en) | Voice recognition equipment | |
| JPS61235899A (en) | Voice recognition equipment | |
| JPH10143190A (en) | Voice recognition device | |
| JPH0262879B2 (en) | ||
| JPS5872996A (en) | Word voice recognition | |
| JPS61138298A (en) | voice recognition device | |
| JPS6131878B2 (en) | ||
| JPS595292A (en) | Word voice recognition method | |
| JPS60182499A (en) | voice recognition device |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| EXPY | Cancellation because of completion of term |