JPH0134400B2 - - Google Patents
Info
- Publication number
- JPH0134400B2 JPH0134400B2 JP56208791A JP20879181A JPH0134400B2 JP H0134400 B2 JPH0134400 B2 JP H0134400B2 JP 56208791 A JP56208791 A JP 56208791A JP 20879181 A JP20879181 A JP 20879181A JP H0134400 B2 JPH0134400 B2 JP H0134400B2
- Authority
- JP
- Japan
- Prior art keywords
- digit
- dissimilarity
- word
- input pattern
- time
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Expired
Links
- 230000006870 function Effects 0.000 claims description 9
- 239000013598 vector Substances 0.000 claims description 9
- 230000001186 cumulative effect Effects 0.000 claims 1
- 238000010586 diagram Methods 0.000 description 15
- 238000000034 method Methods 0.000 description 13
- 230000000875 corresponding effect Effects 0.000 description 4
- 238000010606 normalization Methods 0.000 description 3
- 230000003247 decreasing effect Effects 0.000 description 2
- 230000001276 controlling effect Effects 0.000 description 1
- 230000002596 correlated effect Effects 0.000 description 1
- 238000012544 monitoring process Methods 0.000 description 1
Description
【発明の詳細な説明】
本発明は1個以上の単語を連続して発声した連
続音声を自動的に認識する連続音声認識装置に関
する。DETAILED DESCRIPTION OF THE INVENTION The present invention relates to a continuous speech recognition device that automatically recognizes continuous speech in which one or more words are successively uttered.
音声認識の手段としては従来から種々の方法が
試みられている。それらの中で最も簡単、かつ有
効な方法としてパタンマツチング法があげられ
る。 Various methods have been tried in the past as voice recognition means. Among them, the pattern matching method is the simplest and most effective method.
この方法は、認識すべき語の各単語に標準的な
パタン(以下単語標準パタンと称する)を用意し
ておき、入力された未知の音声パタン(以下入力
パタンと称する)との間で比較操作(すなわちパ
タンマツチング)を行つて相互で異なる度合を表
わす量(以下相異度と称する)を算出し、最も相
異の少ないすなわち相異度が最小になる単語標準
パタンと同じ単語に属すると判定する方法であ
る。 In this method, a standard pattern (hereinafter referred to as word standard pattern) is prepared for each word of the word to be recognized, and a comparison operation is performed between it and an unknown input speech pattern (hereinafter referred to as input pattern). (in other words, pattern matching) to calculate the amount representing the degree of mutual difference (hereinafter referred to as the degree of dissimilarity). This is a method of determining.
同一出願人の特許出願(昭和56年12月9日特許
願(23))には上記パタンマツチング法を基礎と
して動作する連続音声認識装置の動作原理が記載
されている。この原理は大略次のようである。す
なわち、単語標準パタンBvl=(〓vl 1、〓vl 2…、〓vl o
、
…、〓vl Nvl)を複数個あらゆる順列で接続すること
によつて得られるパタンを連続音声の標準パタン
(以下連続音声標準パタンと称す)C=Bv1、
Bv2、…、Bvl、…、BvLmaxと考えて、入力パタン
A=〓1、〓2、…、〓n、…、〓M)とのマツチン
グを行う。全体としての相異度S(A、C)が最
小となるように単語標準パタンの順列を定めるこ
とによつて認識を行なう。 A patent application filed by the same applicant (December 9, 1981 Patent Application (23)) describes the operating principle of a continuous speech recognition device that operates based on the above pattern matching method. The principle is roughly as follows. That is, word standard pattern B vl = (〓 vl 1 , 〓 vl 2 ..., 〓 vl o
,
..., 〓 vl Nvl ) in any permutation, the standard pattern of continuous speech (hereinafter referred to as continuous speech standard pattern) C=B v1 ,
Considering B v2 , ..., B vl , ..., B vLmax , matching with the input pattern A=〓 1 , 〓 2 , ..., 〓 n , . Recognition is performed by determining the permutation of word standard patterns so that the overall degree of difference S(A, C) is minimized.
入力パタンAと連続音声標準パタンC=Bv1、
Bv2、…、Bvl、…、BvLmaxとの相異度は次のよう
にして求める。入力パタンの時間点mと連続音声
標準パタンの時間点nを第1図に示したような最
適な単調増加で非線形関数n=n(m)(以下時間
正規化関数という)にて対応づけを行い、その対
応づけられた時間点における特徴ベクトル間の距
離d(m、n)を時間正規化関数に沿つて加算し
たものを相異度S(A、C)と定義する。 Input pattern A and continuous voice standard pattern C = B v1 ,
The degree of difference from B v2 , ..., B vl , ..., B vLmax is determined as follows. The time point m of the input pattern and the time point n of the continuous speech standard pattern are correlated using an optimal monotonically increasing nonlinear function n=n(m) (hereinafter referred to as time normalization function) as shown in Figure 1. The sum of the distances d(m, n) between the feature vectors at the associated time points along the time normalization function is defined as the degree of dissimilarity S(A, C).
S(A、C)=
minn(m)
M
〓m=1
d(m、n(m)) ……(1)
n=n(m) ……(2)
ここで距数d(m、n)は例えば(3)式にて求める
ことができる。S(A, C) = min n(m) M 〓 m=1 d(m, n(m)) ……(1) n=n(m) ……(2) Here, the distance number d(m, n) can be determined using equation (3), for example.
d(m、n)=Dis(〓n、〓v o)
=R
〓r=1
|anr−v or| ……(3)
ただし
〓n=(an1、an2、…、anR)
〓v o=(bv o1、bv o2、…、bv oR)
(1)式の最小化を次のような動的計画の手法で行
う。ここでvは単語、Vは単語数、Nvは第v番
目の単語標準パタンの終端、Mは入力パタンの終
端、lは桁、Lmaxは最大桁数である。d(m, n)=Dis(〓 n , 〓 v o )
= R 〓 r=1 |a nr − v or | ...(3) where 〓 n = (a n1 , a n2 , ..., a nR ) 〓 v o = (b v o1 , b v o2 , ..., b v oR ) Equation (1) is minimized using the following dynamic programming method. Here, v is a word, V is the number of words, Nv is the end of the vth word standard pattern, M is the end of the input pattern, l is a digit, and Lmax is the maximum number of digits.
初期条件
D(l、v、n)=〓 ……(4)
l=1〜Lmax、v=1〜V、n=1〜Nv
DB(l、m)=〓 ……(5)
l=0〜Lmax、m=0〜M
DB(o、o)=0 ……(6)
をもとに、m=1よりMまで順次以下に示す(7)(8)
式による設定値により(9)、(10)、(11)式の漸化式を求
める。すなわち
D(l、v、o)=DB(l-1、m-1) ……(7)
l=1〜Lmax、v=1〜V
F(l、v、o)=m−1 ……(8)
l=1〜Lmax、v=1〜V
を設定し、漸化式
D(l、v、n)=d(n)+D(l、v、n^)
……(9)
F(l、v、n)=F(l、v、n^)……(10)
ただし
n^=argmin〔D(l、v、n′)〕 ……(11)
n−2n′n
をn=1〜Nv、v=1〜V、l=1〜Lmaxにつ
いてすなわち第3図の斜線部分で示した縦1列に
ついて求める。Initial conditions D(l, v, n)=〓 ……(4) l=1~Lmax, v=1~V, n=1~N v DB(l, m)=〓 ……(5) l= 0 ~ Lmax, m = 0 ~ M DB (o, o) = 0 ... (6) Based on, m = 1 to M are shown below in order (7) (8)
Find the recurrence formulas of equations (9), (10), and (11) using the set values from equations. That is, D(l, v, o)=DB(l-1, m-1)...(7) l=1~Lmax, v=1~V F(l, v, o)=m-1... (8) Set l = 1 ~ Lmax, v = 1 ~ V, and use the recurrence formula D (l, v, n) = d (n) + D (l, v, n^)
...(9) F (l, v, n) = F (l, v, n^) ... (10) where n^ = argmin [D (l, v, n')] ... (11) n -2n'n is determined for n=1 to Nv , v=1 to V, l=1 to Lmax, that is, for one vertical column shown in the shaded area in FIG.
ここで argmin xEXyはxEXの条件のもとでyを最小
とするxを意味している。すなわち(11)式はn−2
n′nのもとでD(m−1、n′)を最小とする
n′をn^としている。また、(9)式は第2図に示す3
つの経路より最小を選択することを示しており、
許される経路を3つに制限したのは時間正規化関
数による対応づけが必要以上に歪むことを防ぐた
めである。ここで相異度を求める時用いた最小値
を選択した経路をマツチング経路と呼び、(10)式の
F(l、v、n)を経路情報と呼ぶ。 Here, argmin xEX y means x that minimizes y under the condition of xEX. In other words, equation (11) is n-2
Minimize D(m-1, n') under n'n
n' is set to n^. Also, equation (9) is expressed as 3 shown in Figure 2.
indicates that the smallest path is selected from among the two paths,
The reason for limiting the number of allowed routes to three is to prevent the correspondence by the time normalization function from being unnecessarily distorted. Here, the route from which the minimum value used when calculating the degree of dissimilarity is selected is called a matching route, and F(l, v, n) in equation (10) is called route information.
つづいて、単語標準パタンの終端Nvにおいて
vに関して最小の相異度を求める。すなわち
ただし
n^=argmin〔D(l、v、Nv)〕 ……(15)
1vV
を求める。ここでDB(l、m)を桁相異度、FB
(l、m)を桁経路情報、W(l、m)を桁認識カ
テゴリと呼ぶ。 Next, the minimum degree of difference with respect to v at the terminal end Nv of the word standard pattern is determined. i.e. However, n^=argmin[D(l,v, Nv )]...(15) Find 1vV. Here, DB (l, m) is the order of magnitude difference, FB
(l, m) is called digit path information, and W(l, m) is called digit recognition category.
このように、(7)〜(15)式に示した縦1列の相
異度計算をmを増加させながら進め、入力パタン
の終端Mにおいてlについて最小の相異度を求め
ることにより(1)式の相異度S(A、C)が得られ
る。また、入力パタンの認識結果は次のようにし
て求める。入力パタンの終端Mにおける各桁の桁
相異度DB(l、M)より許された桁すなわち
Lmin桁よりLmax桁の間で最小値を求め、最小
値の得られた桁Lが入力パタンの桁数である。さ
らに第L桁目の認識結果R(L)をW(L、M)より
得、また桁経路情報FB(L、M)より第L−1桁
目の終端を得る。前記操作を順にくり返すことに
よつて各桁での認識結果R(l)が得られる。 In this way, by proceeding with the calculation of dissimilarity in one vertical column shown in equations (7) to (15) while increasing m, and finding the minimum dissimilarity with respect to l at the terminal M of the input pattern, (1 ) is obtained. Furthermore, the recognition result of the input pattern is obtained as follows. The digits allowed by the digit dissimilarity DB(l, M) of each digit at the terminal M of the input pattern, that is,
The minimum value is found between the Lmin digit and the Lmax digit, and the digit L where the minimum value is obtained is the number of digits of the input pattern. Furthermore, the recognition result R(L) of the Lth digit is obtained from W(L,M), and the termination of the L-1st digit is obtained from the digit path information FB(L,M). By repeating the above operations in order, the recognition result R(l) for each digit can be obtained.
しかしながら、前述の特許出願では、相異度D
(l、v、n)と経路情報F(l、v、n)の記憶
量は桁数lに比例して大きくなる。また、(9)、
(10)、(11)式の漸化式の計算量も桁数lに比例して大
きくなる。一方、連続して発声された数字列を認
識する場合など、桁数の制限を必要としない場合
が多い。 However, in the aforementioned patent application, the degree of dissimilarity D
The storage capacity of (l, v, n) and the route information F(l, v, n) increases in proportion to the number of digits l. Also, (9),
The amount of calculation for the recurrence formulas (10) and (11) also increases in proportion to the number of digits l. On the other hand, in many cases, such as when recognizing a string of consecutively uttered numbers, there is no need to limit the number of digits.
本発明の目的は、入力音声に含まれる桁数の制
限を取り除くことにより、入力音声の最大桁数を
Lmaxとした場合記憶量、計算量ともに従来の
1/Lmaxに減少させる連続音声認識装置を提供
することである。 The purpose of the present invention is to increase the maximum number of digits of input audio by removing the restriction on the number of digits included in input audio.
It is an object of the present invention to provide a continuous speech recognition device that reduces both the storage amount and calculation amount to 1/Lmax of the conventional one when Lmax is used.
次に本発明の原理を第4図を用いて説明する。
今、全体の最小相異度を求めた時得られたマツチ
ング経路(m、n(m))上のある点(m1、n1)
において、始端よりそのマツチング経路に沿つて
その点(m1、n1)まで得られた部分相異度は、
その点(m1、n1)を通るすべてマツチング経路
に沿つて得られる部分相異度の最小値である。す
なわちある点(m1、n1)を通るすべてのマツチ
ング経路に沿つて得られる全体相異度の最小値は
始端よりその点(m1、n1)までの部分相異度と
その点(m1、n1)より終端までの部分相異度の
それぞれの最小値の和で与えられる。 Next, the principle of the present invention will be explained using FIG. 4.
Now, a certain point (m 1 , n 1 ) on the matching path (m, n (m)) obtained when calculating the overall minimum dissimilarity
The partial dissimilarity obtained from the starting point to the point (m 1 , n 1 ) along the matching path is
It is the minimum value of the partial dissimilarities obtained along all matching paths passing through the point (m 1 , n 1 ). In other words, the minimum value of the overall dissimilarity obtained along all matching paths passing through a certain point (m 1 , n 1 ) is the partial dissimilarity from the starting edge to that point (m 1 , n 1 ) and that point ( m 1 , n 1 ) to the terminal end.
すなわち、 S(A、C)= minn(m) M 〓m=1 d(m、n) = minn(m) n 〓m=1 d(m、n)+minn(m) M 〓m=m1+1 d(m、n) ……(16) ただし(m1、n1)はn(m)上の点である。 That is, S(A, C) = min n(m) M 〓 m=1 d(m, n) = min n(m) n 〓 m=1 d(m, n)+min n(m) M 〓 m =m1+1 d(m, n)...(16) However, (m 1 , n 1 ) is a point on n(m).
これにより、ある点(m1、n1)より終端まで
の最小化は、始端より点(m1、n1)までの最小
化と独立に行うことができる。今、第4図に示す
点Xlより終端El+Lまでの最小化を考える。この部
分相異度Sl L(A′、Cl L)は部分入力パタンA′(〓n+
1、〓n1+2、…〓M)と連続音声標準パタンの部分
パタンであるCl L=Bvl、Bvl+1、…、Bvl+Lとの間で
求められる。すなわち
Sl L(A′、Cl L)=
minn(m)
M
〓m=m1+1
d(m、n) ……(17)
にて求められる。ここで桁数の制限がないとすれ
ば、Lを変化させ、部分相異度Sl L(A′、Cl L)の最
小値を求め、最小値の得られた桁数L^が部分入力
パタンA′の桁数である。すなわち
で与えられる。 Thereby, the minimization from a certain point (m 1 , n 1 ) to the terminal can be performed independently of the minimization from the starting end to the point (m 1 , n 1 ). Now, consider the minimization from the point X l to the terminal E l+L shown in Figure 4. This partial dissimilarity S l L (A′, C l L ) is the partial input pattern A′ (〓 n+
1 , 〓 n1+2 , . . . 〓 M ) and C l L =B vl , B vl+1 , . . . , B vl+L , which are partial patterns of the continuous speech standard pattern. That is, S l L (A′, C l L )=min n(m) M 〓 m=m1+1 d(m, n) (17). Here, if there is no limit on the number of digits, change L and find the minimum value of the partial dissimilarity S l L (A′, C l L ), and the number of digits L^ for which the minimum value is obtained is the partial It is the number of digits of input pattern A′. i.e. is given by
次に点Xl+1より終端El+1+Lまでの最小化を考え
る。この部分相異度Sl+1 L(A′、Cl+1 L)は部分入力
パタン(〓n1+1、〓n1+2、…〓M)と連続音声標
準パタンであるCl+1 L=B〓l+1、B〓l+2、…、B〓l+1+L
との間で求められる。ここでCl+1 Lは単語標準パタ
ンをL個あるゆる順列で接続したパタンであり
Cl Lと同じものである。よつて
Sl+1 L(A′、Cl+1 L)=Sl+1 L(A′、Cl L)=Sl L(A′、Cl L)
……(21)
となり、点Xl+1より終端El+1+Lまでの最小化は、
点Xlより終端El+Lまでの最小化と同一の結果とな
る。 Next, consider the minimization from the point X l+1 to the terminal E l+1+L . This partial dissimilarity S l+1 L (A′, C l+1 L ) is the partial input pattern (〓 n1+1 , 〓 n1+2 ,...〓 M ) and the continuous speech standard pattern C l+1 L = B〓 l+1 , B〓 l+2 , ..., B〓 l+1+L
required between. Here, C l+1 L is a pattern in which L word standard patterns are connected in any permutation.
It is the same as C l L. Therefore, S l+1 L (A′, C l+1 L )=S l+1 L (A′, C l L )=S l L (A′, C l L )
...(21) Then, the minimization from the point X l+1 to the terminal E l+1+L is
The result is the same as minimizing from point X l to terminal E l+L .
このように桁数の制限がないとすれば、各点
X1,X2,…,Xl,…より終端までの部分相異度
はすべて同一で(17)、(18)、(19)式で与えら
れ、その時得られた標準パタン系列は
(20)式のCl
∧
Lで与えられる。 If there is no limit on the number of digits, each point
The partial dissimilarities from X 1 , X 2 , ..., ) is given by Cl ∧ L in the equation.
一方、全体の相異度は(16)式に示すように始
端より点Xlまでの部分相異度と点Xlより終端まで
の部分相異度の和
で与えられる。このため、入力パタンの時刻m1
において点X1、X2、…Xl、…の相異度の最小値
を求め、その最小値とm1+1より終端までの
相異度Sl
∧
L(A′、Cl
∧
L)の和により全体の相
異度S(A、C)を求めることができる。 On the other hand, the overall dissimilarity is the sum of the partial dissimilarity from the starting point to point X l and the partial dissimilarity from point X l to the terminal, as shown in equation (16). is given by Therefore, the input pattern time m 1
Find the minimum value of the dissimilarity of the points X 1 , X 2 , ... The overall degree of dissimilarity S(A, C) can be determined by:
次に本発明の基本動作である相異度S(A、C)
の計算手順を第5図、第6図を用いて説明する。
相異度計算は、初期条件のもとで動的計画の漸化
式を入力パタンの時間軸mの順に求めることであ
る。 Next, the degree of dissimilarity S(A, C) which is the basic operation of the present invention
The calculation procedure will be explained using FIGS. 5 and 6.
The dissimilarity calculation is to obtain the recurrence formula of the dynamic program in the order of the time axis m of the input pattern under the initial conditions.
初期条件は
D(v、n)=〓 ……(23)
v=1〜V、n=1〜Nv
DB(o)=o ……(24)
DB(m)=〓 ……(25)
m=1〜M
であり、第6図のブロツク1で行われる。次に入
力パタンの時間点mにおけるn軸に平行な縦1列
の相異度計算は以下のように行われる。初期値を
D(v、o)=DB(m−1) ……(26)
F(v、o)=m−1 ……(27)
として(第6図のブロツク2で行われる。)
漸化式
をnを減少させる方向で計算する(第6図のブロ
ツク4で行われる)。ここでd(〓m、〓v o)は入
力パタンの時間点mの特徴ベクトル〓mと第v番
目の単語標準パタン〓v oとのベクトル間距離を(3)
式により求める(第6図のブロツク3で行われ
る)。 The initial conditions are D(v, n)=〓 ……(23) v=1~V, n=1~N v DB(o)=o ……(24) DB(m)=〓 ……(25) m=1 to M, and is performed in block 1 of FIG. Next, the dissimilarity calculation for one vertical column parallel to the n-axis at time point m of the input pattern is performed as follows. Set the initial value as D(v,o)=DB(m-1)...(26) F(v,o)=m-1...(27) (This is done in block 2 of Figure 6) Gradually formula is calculated in the direction of decreasing n (performed in block 4 of FIG. 6). Here, d(〓m, 〓 v o ) is the distance between the feature vector 〓 m at time point m of the input pattern and the v-th word standard pattern 〓 v o (3)
(carried out in block 3 of FIG. 6).
第2図に示すように(m、n)点の計算は(m
−1、n)、(m−1、n−1)、(m−1、n−
2)の3点の相異度より求められる。次の(m、
n−1)点の計算は(m−1、n−1)、(m−
1、n−2)、(m−1、n−3)の3点の相異度
より求められ(m−1、n)点の相異度は使用し
ないので(m、n)点の計算結果を(m−1、
n)点へ記憶しても(m、n−1)点の形算に影
響を与えない。ゆえにnを減少させる方向で計算
を進めれば、m−1点の相異度とm点の相異度の
記憶エリアを共有することができる。上記の漸化
式計算を縦1列実行した後、単語標準パタンの終
端Nvにおける相異度D(v、Nv)とそれまで計
算された最小の単語相異度である桁相異度DB
(m)と比較し、得られた相異度D(v、Nv)の
方が小さい場合は、その相異度D(v、Nv)を桁
相異度DB(m)とし、その単語標準パタンの属
するカテゴリvを桁認識カテゴリW(m)とし、
その相異度D(v、Nv)が得られたマツチング経
路情報F(v、Nv)を桁経路情報FB(m)とする
(第6図のブロツク5で行われる)。このようにし
て行われる縦1列の相異度計算(第6図のブロツ
ク2,3,4,5の計算)をV個の標準単語パタ
ンについて実行する。次に入力パタンの時間点m
を1つ増加して同様の縦1列の相異度計算をV個
の標準単語パタンについて実行し、入力パタンの
終端Mまで求める。最後に桁経路情報FB(m)と
桁認識カテゴリW(m)より入力パタンの判定を
行う。この判定の方法は、第5図に示すように、
まず、入力パタンの終端Mにおける認識結果F(L)
をW(M)より得、つづいて桁経路情報FB(M)
より第L−1桁目の終端を得る(第6図のブロツ
クで行われる)。前記操作を順にくり返すことに
よつて各桁lでの認識結果R(l)が得られる。 As shown in Figure 2, the calculation at (m, n) point is (m
-1, n), (m-1, n-1), (m-1, n-
It is obtained from the degree of dissimilarity of the three points in 2). Next (m,
The calculation for point n-1) is (m-1, n-1), (m-
1, n-2), (m-1, n-3), and the dissimilarity of the (m-1, n) point is not used, so the (m, n) point is calculated. The result is (m-1,
Even if it is stored at point n), it does not affect the calculation at point (m, n-1). Therefore, if the calculation proceeds in the direction of decreasing n, the storage area for the dissimilarity of the m-1 point and the dissimilarity of the m point can be shared. After executing the above recurrence formula calculation in one column, we obtain the dissimilarity D (v, N v ) at the end N v of the word standard pattern and the digit dissimilarity that is the minimum word dissimilarity calculated so far. D.B.
(m), if the obtained dissimilarity D(v, N v ) is smaller, that dissimilarity D(v, N v ) is taken as the digit dissimilarity DB(m), and Let the category v to which the standard word pattern belongs be the digit recognition category W(m),
The matching path information F (v, N v ) from which the degree of dissimilarity D (v, N v ) is obtained is defined as the digit path information FB (m) (this is carried out in block 5 of FIG. 6). The dissimilarity calculations in one vertical column (calculations in blocks 2, 3, 4, and 5 in FIG. 6) performed in this way are performed for V standard word patterns. Next, the time point m of the input pattern
is incremented by one, and the same dissimilarity calculation in one vertical column is performed for V standard word patterns to obtain the terminal end M of the input pattern. Finally, the input pattern is determined based on the digit path information FB(m) and the digit recognition category W(m). This determination method is as shown in Figure 5.
First, the recognition result F(L) at the end M of the input pattern
is obtained from W(M), and then the digit route information FB(M)
Then, the end of the L-1st digit is obtained (this is done in the block of FIG. 6). By repeating the above operations in order, a recognition result R(l) for each digit l can be obtained.
本発明の連続音声認識装置は前記の相異度計算
手順を実行する装置であるから次のような各部を
必要とする。すなわち、入力パタンA=〓1、〓
2、…、〓n、…、〓Mと、あらかじめ記憶されて
いるV個の単語標準パタンBv=〓v 1、〓v 2、…、〓
v o、…、〓v Nv(v=1、2、…、V)と、
入力パタンの時間点を示す信号mを1よりMま
で変化させ、各mに関して単語を示す信号vを1
よりVまで変化させ、さらに各vに関して標準パ
タンの時間点を示す信号nを1よりNvまで変化
させて与える制御部と、
上記制御部の信号v、nによつて番地指定され
る相異度メモリ部D(v、n)と、経路情報メモ
リ部F(v、n)とを有し、前記制御部より指定
されるm、v、nにおいて入力パタン〓nと単語
vの単語標準パタン〓v oとのベクトル間距離d(〓
n、〓v o)を求める距離計算部と、
各時間点mにおいて各単語vに関して最初に初
期条件を時間点m−1の結果である単語相異度
DB(m−1)と時間点m−1により与え前記距
離d(〓n、〓v o)と時間点m−1における相異度
D(v、n)と経路情報F(v、n)とを参照して
動的計画の漸化式を計算し時間点mにおける相異
度D(v、n)と経路情報F(v、n)を順次求め
る漸化式計算部と、
各時間点mにおいて前記漸化式計算部で求めら
れた各単語の終端での相異度D(v、Nv)の中よ
り最小を求めこれを桁相異度DB(m)としこれ
に対応した経路情報F(v、Nv)を桁経路情報
FB(m)とし最小値が得られた単語名vを桁認識
カテゴリW(m)とする桁相異度計算部と、これ
らを記憶するための桁相異度メモリ部DB(m)
と、桁経路情報メモリ部FB(m)と、桁認識カテ
ゴリメモリ部W(m)と、桁経路情報FB(m)と
桁認識カテゴリW(m)に基づいて逆順に入力パ
タンの各桁のカテゴリを判定し出力する判定部と
を有している。 Since the continuous speech recognition device of the present invention is a device that executes the above-described dissimilarity calculation procedure, it requires the following sections. That is, input pattern A=〓 1 , 〓
2 , ..., 〓 n , ..., 〓 M and V pre-memorized word standard patterns B v = 〓 v 1 , 〓 v 2 , ..., 〓
v o , ..., 〓 v Nv (v = 1, 2, ..., V), the signal m indicating the time point of the input pattern is changed from 1 to M, and the signal v indicating the word for each m is set to 1.
a control section that varies the signal n from 1 to N v, and further varies the signal n indicating the time point of the standard pattern for each v from 1 to N v , and a difference address specified by the signals v and n of the control section; It has a route information memory section D (v, n) and a route information memory section F (v, n), and at m, v, n specified by the control section, input pattern 〓 n and word standard pattern of word v. 〓 Intervector distance d(〓
n , 〓 v o ), and the initial condition for each word v at each time point m is the word dissimilarity which is the result of the time point m-1.
The distance d (〓 n , 〓 v o ) given by DB (m-1) and time point m-1, the degree of dissimilarity D (v, n) at time point m-1, and the route information F (v, n) a recurrence formula calculation unit that sequentially calculates the degree of dissimilarity D (v, n) and route information F (v, n) at time point m by calculating a recurrence formula of the dynamic program with reference to m, find the minimum among the dissimilarities D(v, N v ) at the end of each word obtained by the recurrence formula calculation unit, and use this as the digit dissimilarity DB(m), and calculate the path corresponding to this. The information F(v, Nv ) is the digit route information
A digit dissimilarity calculating unit which takes the word name v for which the minimum value is obtained as FB(m) as a digit recognition category W(m), and a digit dissimilarity memory unit DB(m) for storing these.
, digit path information memory section FB (m), digit recognition category memory section W (m), digit path information FB (m) and digit recognition category W (m) in reverse order for each digit of the input pattern. and a determination section that determines and outputs the category.
本発明では、入力パターンの桁数の制限をなく
すことによつて相異度計算を各桁ごとに行わず、
一箇所にまとめて計算することが可能である。す
なわち、相異度計算においては、最大化原理によ
り入力パターンのある時刻を通る最小値を求める
計算を前半の最小化と後半の最小化に分けること
ができる。ここで、入力パターンの桁制限をなく
すと、前半部の桁数が1からXまでにおけるそれ
ぞれの相違度とは無関係に後半部の最小化を行う
ことができる。この後半分の最小化は相違度が最
小となるよう桁数が求められる。この後半分の最
小化は前半分の桁数にかかわらず同一であり、よ
つて後半分の最小化を一箇所にまとめて計算する
ことができる。この原理を第1桁より適用すれば
全体の相違度計算を1つの桁にまとめることが可
能となる。 In the present invention, by eliminating the limit on the number of digits of the input pattern, the degree of dissimilarity calculation is not performed for each digit,
It is possible to perform calculations all in one place. That is, in dissimilarity calculation, calculation for finding the minimum value of an input pattern passing through a certain time can be divided into the first half of minimization and the second half of minimization based on the maximization principle. Here, if the digit limit of the input pattern is eliminated, the latter half can be minimized regardless of the degree of difference between the number of digits in the first half from 1 to X. In minimizing this latter half, the number of digits is determined so that the degree of difference is minimized. The minimization of this latter half is the same regardless of the number of digits in the first half, and therefore the minimization of the latter half can be calculated all at once. By applying this principle starting from the first digit, it becomes possible to consolidate the entire dissimilarity calculation into one digit.
このように本発明では、入力パタンの桁数の制
限をなくすことによつて相異度計算を各桁ごとに
行わずに、一ケ所でまとめて計算することが可能
となり、計算量、記憶量ともに従来の1/Lmax
(Lmaxは入力音声の最大桁数)に減少すること
ができる。 In this way, in the present invention, by eliminating the limit on the number of digits of the input pattern, it is possible to calculate the degree of dissimilarity all at one place without having to perform the calculation for each digit, which reduces the amount of calculation and memory. Both are conventional 1/Lmax
(Lmax is the maximum number of digits of the input audio).
次に本発明の装置の具体的構成を図面を参照し
ながら説明する。第7図は、本発明の一構成例を
示すブロツク図であり、第8図は制御指令信号の
タイムチヤートである。制御部10は、m1,n
1,v1などの制御指令信号を第8図に示すよう
に発することによつて、他の各部を制御する機能
を持つが、その詳細は他の各部の動作に関連して
その都度説明する。 Next, the specific configuration of the apparatus of the present invention will be explained with reference to the drawings. FIG. 7 is a block diagram showing a configuration example of the present invention, and FIG. 8 is a time chart of control command signals. The control unit 10 has m1, n
It has a function of controlling other parts by issuing control command signals such as 1 and v1 as shown in FIG. 8, but the details will be explained each time in relation to the operation of the other parts.
入力部11は、信号Speech inで与えられる入
力部を分析し一定時間ごとに特徴ベクトルを出力
する。この連続分析は例えば、多チヤンネルのフ
イルタより構成されるフイルタバンクによる周波
数分析などがある。また入力部11には入力音声
のレベルを監視し、音声の始端、終端を検出する
機能を持ち、その検出した時点を制御部10へ信
号SPにより伝える。入力パタンバツフア12は、
音声の始端が検出された後、信号m3に従つて入
力部11より与えられる特徴ベクトルamを記憶
する。信号m3は入力パタンの時間点mに対応し
た信号である。標準パタンメモリ部13は、V個
の単語標準パタンB1,B2,…Bvを記憶し、標準
パタン長メモリ部14は単語標準パタンBvの長
さNvを記憶している。 The input section 11 analyzes the input section given by the signal Speech in, and outputs a feature vector at regular intervals. This continuous analysis includes, for example, frequency analysis using a filter bank composed of multi-channel filters. The input unit 11 also has a function of monitoring the level of input audio and detecting the start and end of the audio, and transmits the detected time to the control unit 10 by a signal SP. The input pattern buffer 12 is
After the start of the voice is detected, the feature vector am given from the input section 11 is stored in accordance with the signal m3. Signal m3 is a signal corresponding to time point m of the input pattern. The standard pattern memory unit 13 stores V word standard patterns B 1 , B 2 , . . . B v , and the standard pattern length memory unit 14 stores the length N v of the word standard pattern B v .
制御部10は標準パタンの単語vを指示する信
号v1を標準パタン長メモリ部14へ発し、単語
標準パタンBvの長さNvを読み出し、単語標準パ
タンの時間点nに対応する信号n1を発生する。
信号n1に従つて入力パタンバツフア12より入
力パタンの特徴ベクトル〓mが読み出され、標準
パタンメモリ部より〓v oが順次読み出され距離計
算部15において(3)式が計算される。 The control unit 10 issues a signal v1 instructing the word v of the standard pattern to the standard pattern length memory unit 14, reads out the length Nv of the word standard pattern Bv , and outputs the signal n1 corresponding to the time point n of the word standard pattern. Occur.
According to the signal n1, the feature vector 〓m of the input pattern is read out from the input pattern buffer 12, 〓vo is sequentially read out from the standard pattern memory section, and the distance calculation section 15 calculates the equation (3).
距離計算部15において第11図に示すように
初めに信号Cl153にてアキユムレータ153が
クリヤされ、入力パタンバツフア12と標準パタ
ンメモリ部13より信号r1に従つてR個のデー
タが読み込まれ、絶対値回路151にて絶対値を
求め、加算器152にて加算され、(3)式の距離d
(〓n、〓v n)がアキユムレータ153に求まり、
この距離が漸化式計算部17へ入力される。 In the distance calculation section 15, as shown in FIG. 11, the accumulator 153 is first cleared by the signal Cl153, R pieces of data are read from the input pattern buffer 12 and the standard pattern memory section 13 according to the signal r1, and the absolute value circuit The absolute value is obtained in step 151 and added in adder 152, and the distance d in equation (3) is obtained.
(〓 n , 〓 v n ) is found in the accumulator 153,
This distance is input to the recurrence formula calculation section 17.
相異度メモリ部18、経路情報メモリ部19は
第9図に示す2次元の構成であり、桁相異度メモ
リ部21、桁認識カテゴリメモリ部22、桁経路
情報メモリ部23は第10図に示す1次元の構成
である。 The dissimilarity memory section 18 and the route information memory section 19 have a two-dimensional configuration as shown in FIG. This is a one-dimensional configuration shown in .
漸化式計算の初期セツトは音声の入力される前
に制御部10の信号CLにより行われ、相異度メ
モリ部18、桁相異度メモリ21へ(23)、(24)、
(25)式で示した値がセツトされる。 The initial set of the recurrence formula calculation is performed by the signal CL from the control unit 10 before the voice is input, and the data is sent to the dissimilarity memory unit 18, digit dissimilarity memory 21 (23), (24),
The value shown by equation (25) is set.
漸化式計算部17は第6図のブロツク4を行う
部分であり、漸化式(28)、(29)、(30)を実行す
る。すなわち、漸化式計算部17は、第12図に
示すように3つの相異度レジスタD1,D2,D
3と、その3つのレジスタD1,D2,D3の最
小値を計算する比較回路171と、加算器172
と、3つの経路レジスタF1,F2,F3より構
成される。制御部10より発せられた信号n1,
n11,n12によつて相異度メモリ部18と経
路メモリ部19より3つの相異度D(v、n)、D
(v、n−1)、D(v、n−2)と3つの経路情
報F(v、n)、F(v−n−1)、F(v、n−2)
を読み出しそれぞれ相異度レジスタD1,D2,
D3と経路レジスタF1,F2,F3へ格納す
る。比較回路171は3つの相異度レジスタD
1,D2,D3より最小値を検出し、その最小値
が得られた相異度Dn^(n^は1、2、3のどれか)
に対応した経路レジスタFn^を選択するゲート信
号n^を発する。前記ゲートn^により選択された経
路レジスタFn^の内容が経路メモリ部19のF
(v、n)へ格納される。また比較回路171よ
り出力された相異度の最小値D(v、n^)は、距
離計算部15より出力された距離d(〓m、〓v o)
と加算器172によつて加算され、相異度メモリ
部18へ格納される。 The recurrence formula calculation unit 17 is a part that performs block 4 in FIG. 6, and executes recurrence formulas (28), (29), and (30). That is, the recurrence formula calculation unit 17 operates on three difference degree registers D1, D2, and D as shown in FIG.
3, a comparison circuit 171 that calculates the minimum value of the three registers D1, D2, and D3, and an adder 172.
and three route registers F1, F2, and F3. The signal n1 issued from the control unit 10,
Three dissimilarities D(v, n), D
(v, n-1), D(v, n-2) and three route information F(v, n), F(v-n-1), F(v, n-2)
are read out and the difference registers D1, D2,
D3 and path registers F1, F2, and F3. The comparison circuit 171 has three dissimilarity registers D.
The minimum value is detected from 1, D2, and D3, and the degree of dissimilarity Dn^ (where n^ is 1, 2, or 3) is the minimum value obtained.
A gate signal n^ is generated that selects the path register Fn^ corresponding to the path register Fn^. The contents of the route register Fn^ selected by the gate n^ are stored in the route memory section 19 F.
(v, n). Further, the minimum value D(v, n^) of the degree of dissimilarity outputted from the comparison circuit 171 is the distance d(〓m, 〓 v o ) outputted from the distance calculation section 15.
is added by the adder 172 and stored in the dissimilarity memory unit 18.
この漸化式計算がn=1よりNvまで算出され、
この結果である相異度D(v、Nv)が各vに対し
て算出される。 This recurrence formula calculation is calculated from n=1 to N v ,
The degree of dissimilarity D(v, N v ) which is the result is calculated for each v.
桁相異度計算部20は、第6図のブロツク5を
行う部分であり、V個の相異度D(v、Nv)の最
小値を逐次求める。すなわち、桁相異度計算部2
0は第13図に示すように、比較回路201と、
相異度D(v、Nv)を保持するレジスタ202
と、単語標準パタンの属するカテゴリvを保持す
るレジスタ203と、経路情報F(v、Nv)を保
持するレジスタ204より構成される。制御部1
0より発せられた信号n1に従い、相異度メモリ
部18と経路メモリ部19より相異度D(v、
Nv)と経路情報F(v、Nv)が読み出され、そ
れぞれレジスタ202と204へ格納され、単語
標準パタンの属するカテゴリvをレジスタ203
へ格納される。一方、比較回路201は前記相異
度D(v、Nv)と桁相異度メモリ部21より読み
出された桁相異度DB(m)と比較し、相異度D
(v、Nv)がより小さいと判定するゲート信号v^
を発生する。ゲート信号v^に従つてレジスタ20
2,203,204に保持されていた相異度D
(v、Nv)、カテゴリv、経路情報F(v、Nv)
がそれぞれ桁相異度メモリ部21のDB(m)、桁
認識カテゴリメモリ部22W(m)桁経路メモリ
部23FB(m)へ格納される。さらに制御部10
より信号Cl2によつて第5図のブロツク2にて行
われる部分である縦1列の相異度計算部の(26)、
(27)式に示した初期セツトが行われる。すなわ
ち桁相異度メモリ部21よりDB(m−1)が読
み出され、相異度メモリ部18のD(v、o)へ
格納され、経路メモリ部19のF(v、o)へm
−1が格納される。判定部24は、第5図のブロ
ツク6を行う部分であり、桁経路情報FB(m)と
桁認識カテゴリW(m)より入力パタンの各桁の
認識結果R(l)を出力する。すなわち、判定部24
は第14図に示すように、桁経路情報F(m)を
保持するレジスタ244と認識結果を保持するレ
ジスタ245より構成される。音声の終端が検出
されると入力部11より信号SPによつて制御部
10に通知され、つづいて制御部10は判定部2
4へ信号m1を発し、判定部24は処理を開始す
る。判定制御部246は信号m1を受けた後、m
=Mとしてアドドレス信号m2を桁経路メモリ部
23と桁認識カテゴリメモリ部22へ発し、FB
(M)とW(M)が読み出され、レジスタ244と
レジスタ245へ格納される。レジスタ245の
内容が認識結果として出力される。さらに判定制
御部246はm=(レジスタ244の内容)とし
てアドレス信号m2を桁経路メモリ部23と桁認
識カテゴリメモリ部22へ発し、FB(m)とW
(m)が読み出されレジスタ244とレジスタ2
45へ格納される。この処理を順次m=Mよりm
=0となるまで繰り返すことにより認識結果がレ
ジスタRより出力される。 The digit dissimilarity calculation unit 20 is a part that performs block 5 in FIG. 6, and sequentially finds the minimum value of V dissimilarities D(v, N v ). In other words, the digit difference calculation unit 2
0, as shown in FIG. 13, the comparison circuit 201,
Register 202 that holds the degree of dissimilarity D (v, N v )
, a register 203 that holds the category v to which the standard word pattern belongs, and a register 204 that holds the route information F(v, N v ). Control part 1
According to the signal n1 emitted from 0, the dissimilarity memory section 18 and the route memory section 19 calculate the dissimilarity degree D(v,
N v ) and route information F (v, N v ) are read out and stored in registers 202 and 204, respectively, and the category v to which the word standard pattern belongs is stored in register 203.
is stored in On the other hand, the comparison circuit 201 compares the degree of dissimilarity D(v, Nv ) with the degree of digit dissimilarity DB(m) read out from the digit dissimilarity memory section 21, and calculates the degree of dissimilarity D
Gate signal v^ that determines that (v, N v ) is smaller
occurs. Register 20 according to gate signal v^
Dissimilarity D held at 2,203,204
(v, N v ), category v, route information F (v, N v )
are stored in the DB(m) of the digit dissimilarity memory unit 21, the digit recognition category memory unit 22W(m), and the digit path memory unit 23FB(m), respectively. Furthermore, the control unit 10
(26) of the dissimilarity calculation unit in one vertical column, which is the part performed in block 2 of FIG. 5 by the signal Cl2.
The initial set shown in equation (27) is performed. That is, DB(m-1) is read from the digit dissimilarity memory section 21, stored in D(v, o) of the dissimilarity memory section 18, and stored in F(v, o) of the path memory section 19.
-1 is stored. The determining unit 24 is a part that performs block 6 in FIG. 5, and outputs the recognition result R(l) of each digit of the input pattern from the digit path information FB(m) and the digit recognition category W(m). That is, the determination unit 24
As shown in FIG. 14, it is composed of a register 244 that holds digit path information F(m) and a register 245 that holds recognition results. When the end of the audio is detected, the input unit 11 notifies the control unit 10 by the signal SP, and then the control unit 10
The determination unit 24 starts processing. After receiving the signal m1, the determination control unit 246 determines m
= M, and sends the address signal m2 to the digit path memory section 23 and digit recognition category memory section 22, and
(M) and W(M) are read and stored in register 244 and register 245. The contents of the register 245 are output as the recognition result. Furthermore, the determination control unit 246 issues an address signal m2 as m=(content of the register 244) to the digit path memory unit 23 and the digit recognition category memory unit 22, and outputs FB(m) and W
(m) is read and register 244 and register 2
45. This process is performed sequentially from m=M to m
By repeating this until =0, the recognition result is output from register R.
以上、本発明の原理とその一構成例を説明した
が、これらの記載は本発明の範囲を限定するもの
ではない。特に、入力パタン〓nと標準パタン〓v o
との距離を(3)式のような距離尺度を用いて説明し
たが、このかわりに(31)式のようなユークリツ
ド距離、(32)式のような内積等を用いてよい。 Although the principle of the present invention and one configuration example thereof have been explained above, these descriptions do not limit the scope of the present invention. In particular, input pattern 〓 n and standard pattern 〓 v o
Although the distance between
d(m、n)=R
〓r=1
(anr−bv or)2 ……(31)
d(m、n)=R
〓r=1
(anr×bv nr) ……(32)
また、相異度を計算するための漸化式は(28)、
(29)、(30)式の形の他にも種々考えられ、この
(28)、(29)、(30)式の代わりに特公告56−28278
号に記載されている形も使用できることは明白で
ある。d(m, n)= R 〓 r=1 (a nr −b v or ) 2 …(31) d(m, n)= R 〓 r=1 (a nr ×b v nr ) ……(32 ) Also, the recurrence formula for calculating the degree of dissimilarity is (28),
Various forms other than formulas (29) and (30) can be considered, and instead of formulas (28), (29), and (30), Japanese Patent Publication No. 56-28278
It is clear that the forms described in this issue can also be used.
第1図はマツチング経路の例を示した図であ
り、第2図は漸化式において許されているマツチ
ング経路を示した図であり、第3図は相異度計算
順序を示した図であり、第4図は本発明の原理を
示した図であり、第5図は、判定処理の計算順序
を示した図であり、第6−1および6−2図は本
発明の計算手順を示すフローチヤートであり、第
7図は本発明の一実施例の構成図であり、第8図
は本発明の実施例の動作を説明するためのタイム
チヤートであり、第9図は相異度メモリ部、経路
情報メモリ部の構成図であり、第10図は単語相
異度メモリ部、単語経路情報メモリ部、単語認識
カテゴリメモリ部の構成図であり、第11図は本
発明の一構成要素の一つである距離計算部の構成
図であり、第12図は漸化式計算部の構成図であ
り、第13図は桁相異度計算部の構成図であり、
第14図は判定部の構成図である。第7図、第1
1図、第12図、第13図、第14図において、
10……制御部、11……入力部、12……入
力パタンバツフア、13……標準パタンメモリ
部、14……標準パタン長メモリ部、15……距
離計算部、17……漸化式計算部、18……相異
度メモリ部、19……経路情報メモリ部、20…
…桁相異度計算部、21……桁相異度メモリ部、
22……桁認識カテゴリメモリ部、23……桁経
路情報メモリ部、24……判定部、151……絶
対値回路、152……加算器、153……アキユ
ムレータ、171……比較回路、172……加算
器、D1,D2,D3……相異度を保持するレジ
スタ、F1,F2,F3……経路を保持するレジ
スタ、201……比較回路、202……単語相異
度を保持するレジスタ、203……カテゴリを保
持するレジスタ、204……経路情報を保持する
レジスタ、244……単語経路情報を保持するレ
ジスタ、245……認識結果を保持し出力するレ
ジスタ、246……判定制御部。
Figure 1 is a diagram showing an example of matching paths, Figure 2 is a diagram showing matching paths allowed in the recurrence formula, and Figure 3 is a diagram showing the order of dissimilarity calculation. Fig. 4 is a diagram showing the principle of the present invention, Fig. 5 is a diagram showing the calculation order of the judgment process, and Figs. 6-1 and 6-2 are diagrams showing the calculation procedure of the present invention. FIG. 7 is a configuration diagram of an embodiment of the present invention, FIG. 8 is a time chart for explaining the operation of the embodiment of the present invention, and FIG. 9 is a flowchart showing the difference degree. FIG. 10 is a configuration diagram of a memory unit and a route information memory unit; FIG. 10 is a configuration diagram of a word dissimilarity memory unit, a word route information memory unit, and a word recognition category memory unit; FIG. 11 is a configuration diagram of one configuration of the present invention. FIG. 12 is a configuration diagram of a distance calculation section which is one of the elements, FIG. 12 is a configuration diagram of a recurrence formula calculation section, and FIG. 13 is a configuration diagram of a digit dissimilarity calculation section.
FIG. 14 is a configuration diagram of the determination section. Figure 7, 1st
1, 12, 13, and 14, 10...control section, 11...input section, 12...input pattern buffer, 13...standard pattern memory section, 14...standard pattern length memory section , 15... Distance calculation section, 17... Recurrence equation calculation section, 18... Dissimilarity degree memory section, 19... Route information memory section, 20...
... Digit difference calculation unit, 21... Digit difference memory unit,
22... Digit recognition category memory unit, 23... Digit path information memory unit, 24... Judgment unit, 151... Absolute value circuit, 152... Adder, 153... Accumulator, 171... Comparison circuit, 172... ...Adder, D1, D2, D3...Register that holds the degree of dissimilarity, F1, F2, F3...Register that holds the path, 201...Comparison circuit, 202...Register that holds the degree of word dissimilarity, 203...Register for holding categories, 204...Register for holding route information, 244...Register for holding word route information, 245...Register for holding and outputting recognition results, 246...Determination control unit.
Claims (1)
よりなる入力パタンA=〓1、〓2、…、〓n 、…、
〓Mとあらかじめ記憶されているV個の単語標準
パタンBv=〓v 1、〓v 2、…、〓v o、…、〓v Nv(v=
1、2、…、V)をあらゆる順序で接続した連続
音声標準パタンC=Bv1、Bv2、…、Bvl、…、
BvLmaxとの間で入力パタンの時間軸mと連続音声
標準パタンの時間軸nを対応させる時間関数n
(m)の上の入力パタン〓nと連続音声標準パタン
〓oのベクトル間距離d(〓n、〓o)の和として定
義される相異度の最小値を求める操作に関し、入
力パタンの時間点mを単語の終端とした入力パタ
ンの部分と単語標準パタンとの間における最適な
時間関数n(m)によつて与えられるベクトル間
距離の最小累積量を示す桁相異度DB(m)と、
この時間関数の先頭の時間点を示す桁経路情報
FB(m)と、桁認識カテゴリW(m)とを、入力
パタンの時間点mに対して順次求め、最後に入力
パタンの桁数および各桁の認識結果を判定する連
続音声認識装置において、入力パタンの時間点を
示す信号mを1からMまで変化させ、各mに関し
て単語を示す信号vを1からVまで変化させ、さ
らに各vに関して標準パタンの時間点を示す信号
nを1からNvまで変化させて与える制御部と、 前記制御部の信号v、nによつて番地指定され
る相異度メモリ部D(v、n)と、経路情報メモ
リ部F(v、n)とを有し、前記制御部より指定
されるm、v、nにおいて入力パタン〓nと単語
vの単語標準パタン〓v oとのベクトル間距離d(〓
n、〓v o)を求める距離計算部と、 各時間点mにおいて各単語vに関して最初に初
期条件を時間点m−1の結果である単語相異度
DB(m−1)と時間点m−1により与え前記距
離d(〓n、〓v oと時間点m−1における相異度D
(v、n)と経路情報F(v、n)とを参照して動
的計画の漸化式を計算し時間点mにおける相異度
D(v、n)と経路情報F(v、n)を順次求める
漸化式計算部と、 各時間mにおいて前記漸化式計算部で求められ
た各単語の終端での相異度D(v、Nv)の中より
最小を求めこれを桁相異度DB(m)としこれに
対応した経路情報F(v、Nv)を桁経路情報FB
(m)とし最小値が得られた単語名vを桁認識カ
テゴリW(m)とする桁相異度計算部と、これら
を記憶するための桁相異度メモリ部DB(m)と、
桁経路情報メモリ部FB(m)と、桁認識カテゴリ
メモリ部W(m)と、桁経路情報FB(m)と桁認
識カテゴリW(m)に基づいて逆順に入力パタン
の各桁のカテゴリを判定し出力する判定部とを持
つことを特徴とする連続音声認識装置。 [Claims] 1. Input pattern A consisting of one or more words that is a time series of feature vectors = 1 , 2 , ..., n , ...,
〓 M and V word standard patterns stored in advance B v = 〓 v 1 , 〓 v 2 , ..., 〓 v o , ..., 〓 v N v (v =
1, 2, ..., V) connected in any order C = B v1 , B v2 , ..., B vl , ...,
B A time function n that makes the time axis m of the input pattern correspond to the time axis n of the continuous audio standard pattern between vLmax
Regarding the operation to find the minimum value of the degree of dissimilarity defined as the sum of the vector distance d (〓 n , 〓 o ) between the input pattern 〓 n and the continuous speech standard pattern 〓 o on (m), the time of the input pattern Digit dissimilarity DB(m) indicating the minimum cumulative amount of distance between vectors given by the optimal time function n(m) between the part of the input pattern with point m as the end of the word and the word standard pattern and,
Digit path information indicating the first time point of this time function
In a continuous speech recognition device that sequentially obtains FB(m) and digit recognition category W(m) for time points m of an input pattern, and finally determines the number of digits of the input pattern and the recognition result of each digit, The signal m indicating the time point of the input pattern is varied from 1 to M, the signal v indicating the word is varied from 1 to V for each m, and the signal n indicating the time point of the standard pattern is varied from 1 to N for each v. a control section that changes the information up to v , a dissimilarity memory section D ( v , n) whose address is specified by the signals v and n of the control section, and a route information memory section F (v, n). and at m, v, and n specified by the control unit , the vector distance d (〓
n , 〓 v o ), and the initial condition for each word v at each time point m is the word dissimilarity which is the result of the time point m-1.
Given by DB(m-1) and time point m-1, the distance d(〓 n , 〓 v o and the degree of dissimilarity D at time point m-1
(v, n) and the route information F(v, n), calculate the recurrence formula of the dynamic program, and calculate the dissimilarity degree D(v, n) at time point m and the route information F(v, n). ), and a recurrence formula calculation unit that sequentially calculates the difference D(v, N v ) at the end of each word calculated by the recurrence formula calculation unit at each time m. Set the degree of dissimilarity DB (m) to the corresponding route information F (v, N v ) as the digit route information FB.
(m), and a digit dissimilarity calculation unit which sets the word name v for which the minimum value was obtained as a digit recognition category W(m), and a digit dissimilarity memory unit DB(m) for storing these;
The category of each digit of the input pattern is determined in reverse order based on the digit path information memory section FB (m), the digit recognition category memory section W (m), the digit path information FB (m), and the digit recognition category W (m). A continuous speech recognition device characterized by having a determination unit that makes a determination and outputs the result.
Priority Applications (5)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP56208791A JPS58126599A (en) | 1981-12-23 | 1981-12-23 | Continuous voice recognition equipment |
| US06/447,829 US4592086A (en) | 1981-12-09 | 1982-12-08 | Continuous speech recognition system |
| CA000417329A CA1193013A (en) | 1981-12-09 | 1982-12-09 | Continuous speech recognition system |
| EP82306577A EP0081390B1 (en) | 1981-12-09 | 1982-12-09 | Continuous speech recognition system |
| DE8282306577T DE3267835D1 (en) | 1981-12-09 | 1982-12-09 | Continuous speech recognition system |
Applications Claiming Priority (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP56208791A JPS58126599A (en) | 1981-12-23 | 1981-12-23 | Continuous voice recognition equipment |
Publications (2)
| Publication Number | Publication Date |
|---|---|
| JPS58126599A JPS58126599A (en) | 1983-07-28 |
| JPH0134400B2 true JPH0134400B2 (en) | 1989-07-19 |
Family
ID=16562167
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| JP56208791A Granted JPS58126599A (en) | 1981-12-09 | 1981-12-23 | Continuous voice recognition equipment |
Country Status (1)
| Country | Link |
|---|---|
| JP (1) | JPS58126599A (en) |
-
1981
- 1981-12-23 JP JP56208791A patent/JPS58126599A/en active Granted
Also Published As
| Publication number | Publication date |
|---|---|
| JPS58126599A (en) | 1983-07-28 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US4592086A (en) | Continuous speech recognition system | |
| JP2595495B2 (en) | Pattern matching device | |
| JPS62133500A (en) | Method and apparatus for identifying electrical signal introduced from acoustic signal | |
| JPH0123798B2 (en) | ||
| EP0079578A1 (en) | Continuous speech recognition method and device | |
| JPH0159600B2 (en) | ||
| EP0144689B1 (en) | Pattern matching system | |
| EP0086081B1 (en) | Continuous speech recognition system | |
| US4426551A (en) | Speech recognition method and device | |
| EP0162255B1 (en) | Pattern matching method and apparatus therefor | |
| JP2980026B2 (en) | Voice recognition device | |
| JPH0247760B2 (en) | ||
| JP2964881B2 (en) | Voice recognition device | |
| JPH0134399B2 (en) | ||
| JP2712856B2 (en) | Voice recognition device | |
| JPH0251519B2 (en) | ||
| JP2738403B2 (en) | Voice recognition device | |
| JPH022159B2 (en) | ||
| JPH0436400B2 (en) | ||
| JPS61275896A (en) | Pattern zoning apparatus | |
| JPS5926960B2 (en) | Time series pattern matching device | |
| JPS58126599A (en) | Continuous voice recognition equipment | |
| JPH0247755B2 (en) | ||
| JPH044600B2 (en) | ||
| JPH08248984A (en) | Voice recognition method |