JPH01204172A - Dynamic neural network having learning mechanism - Google Patents
Dynamic neural network having learning mechanismInfo
- Publication number
- JPH01204172A JPH01204172A JP63029675A JP2967588A JPH01204172A JP H01204172 A JPH01204172 A JP H01204172A JP 63029675 A JP63029675 A JP 63029675A JP 2967588 A JP2967588 A JP 2967588A JP H01204172 A JPH01204172 A JP H01204172A
- Authority
- JP
- Japan
- Prior art keywords
- learning
- time
- output
- layer
- input
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
- 230000007246 mechanism Effects 0.000 title claims abstract description 6
- 238000013528 artificial neural network Methods 0.000 title claims description 26
- 230000008878 coupling Effects 0.000 claims abstract description 35
- 238000010168 coupling process Methods 0.000 claims abstract description 35
- 238000005859 coupling reaction Methods 0.000 claims abstract description 35
- 238000000034 method Methods 0.000 claims abstract description 11
- 230000001537 neural effect Effects 0.000 claims description 2
- 238000005516 engineering process Methods 0.000 claims 1
- 238000012937 correction Methods 0.000 abstract description 13
- 238000004364 calculation method Methods 0.000 description 12
- 239000013598 vector Substances 0.000 description 8
- 230000006870 function Effects 0.000 description 7
- 230000001186 cumulative effect Effects 0.000 description 5
- 238000010586 diagram Methods 0.000 description 5
- 230000010365 information processing Effects 0.000 description 4
- 238000012545 processing Methods 0.000 description 4
- 230000004044 response Effects 0.000 description 3
- 230000008602 contraction Effects 0.000 description 2
- 230000000694 effects Effects 0.000 description 2
- 238000003062 neural network model Methods 0.000 description 2
- 210000002569 neuron Anatomy 0.000 description 2
- 238000010606 normalization Methods 0.000 description 2
- 238000003909 pattern recognition Methods 0.000 description 2
- 238000005316 response function Methods 0.000 description 2
- 238000012935 Averaging Methods 0.000 description 1
- 101100136092 Drosophila melanogaster peng gene Proteins 0.000 description 1
- 230000004913 activation Effects 0.000 description 1
- 238000004422 calculation algorithm Methods 0.000 description 1
- 238000011156 evaluation Methods 0.000 description 1
- 230000004048 modification Effects 0.000 description 1
- 238000012986 modification Methods 0.000 description 1
- 210000000653 nervous system Anatomy 0.000 description 1
- 238000005457 optimization Methods 0.000 description 1
- 230000009466 transformation Effects 0.000 description 1
Landscapes
- Image Analysis (AREA)
Abstract
Description
【発明の詳細な説明】
(産業上の利用分野)
本発明は音声等の時系列パターンの認識に用いるパター
ン学習機構を有するダイナミック・ニューラル・ネット
ワークに関する。DETAILED DESCRIPTION OF THE INVENTION (Field of Industrial Application) The present invention relates to a dynamic neural network having a pattern learning mechanism used for recognizing time-series patterns such as speech.
(従来の技術)
ニューラル・ネットワークは生体の脳神経系が比較的単
純な動作特性を有する神経細胞とその間の多数の結合か
ら構成されている情報処理システムであることを参考に
して考案された情報処理モデルで、神経細胞に相当する
処理ユニt7 ト(以下ユニットと略す)とその間を結
ぶユニット間結合を有する。このユニット間結合の係数
を変えることによってシステムはさまざまな情報処理動
作を行なう。(Prior Art) A neural network is an information processing system devised based on the fact that the biological nervous system is an information processing system composed of neurons with relatively simple operating characteristics and numerous connections between them. The model has processing units t7 (hereinafter abbreviated as units) corresponding to neurons and inter-unit connections connecting them. By changing the coefficients of this inter-unit coupling, the system performs various information processing operations.
このニューラル・ネットワーク・モデルは情報処理シス
テムとして特に画像や音声等のパターン認識処理に有効
であろうと期待されており、その詳細に関しては「日経
エレクトロニクス」誌、第427号の第115頁く昭和
62年8月10日発行)「ニューラル・ネットをパター
ン認識、信号処理、知識処理に使う」に解説されている
。(以下、文献1と称する。)
上記文献1によるとニューラル・ネットワークは第2図
に示すように、入力層、中間層、出力層と呼ばれる階層
構造を有しており、各層は複数のユニットから構成され
ている。またユニット間結合は隣接する層の間にだけ許
され、層内でのユニット間結合は禁止されている。認識
時にはネッワークは入力層の各ユニットの活性度として
入力データを与えられ、ユニット間結合を通じて順次隣
接する中間層へ情報を伝達し、最後に出力層にまで到達
する。こうして入力データに対するネットワークの応答
結果が出力層のユニットの活性度のパターンとして得ら
れる。This neural network model is expected to be effective as an information processing system, especially for pattern recognition processing of images and sounds, etc. For details, please refer to "Nikkei Electronics" magazine, No. 427, page 115. (Published August 10, 2016) ``Using Neural Nets for Pattern Recognition, Signal Processing, and Knowledge Processing.'' (Hereinafter referred to as Document 1.) According to Document 1, a neural network has a hierarchical structure called an input layer, a middle layer, and an output layer, as shown in Figure 2, and each layer consists of multiple units. It is configured. Further, inter-unit coupling is allowed only between adjacent layers, and inter-unit coupling within a layer is prohibited. During recognition, the network is given input data as the activation level of each unit in the input layer, transmits information sequentially to adjacent intermediate layers through inter-unit connections, and finally reaches the output layer. In this way, the response result of the network to the input data is obtained as a pattern of the activity levels of the units in the output layer.
ネットワークが指定した動作を行なうようにユニット間
結合を定める為には教師付き学習と呼ばれる手法を用い
る。即ち、入力層に学習させたいパターンを提示し、出
力層には対応して出力すべき教師信号を提示して、出力
層での教師信号と実際の出力値との差異を小さくするよ
うに結合係数を決定する。上記のような構成のニューラ
ル・ネットワークの場合には、この出力誤差最小化学習
はパックプロパゲーション学習と呼ばれており、その詳
細なアルゴリズムに関しては文献1に詳しい。A method called supervised learning is used to determine connections between units so that the network performs specified operations. In other words, a pattern to be learned is presented to the input layer, a corresponding teacher signal to be output is presented to the output layer, and the combination is made to reduce the difference between the teacher signal and the actual output value in the output layer. Determine the coefficients. In the case of a neural network configured as described above, this output error minimization learning is called pack propagation learning, and its detailed algorithm can be found in Reference 1.
(発明が解決しようとする問題点)
このようなニューラル・ネットワークを音声認識に用い
ることができれば、音声パターンの有する多様性を学習
によって吸収して、良好な認識性能を実現できる可能性
があるが、実際に上記のニューラル・ネットワークを音
声認識に用いる為には、いくつかの解決しなければなら
ない問題が存在する。(Problem to be solved by the invention) If such a neural network can be used for speech recognition, it may be possible to absorb the diversity of speech patterns through learning and achieve good recognition performance. In order to actually use the above neural network for speech recognition, there are several problems that must be solved.
第一に音声は同じカテゴリ(例えば単語〉のパターンで
も発声の度に、或は話者毎にその継続時間長が異なるの
で、長さの異なる音声パターンを同じニューラル・ネッ
トワークの入力層に提示する為の工夫が必要となる。First, speech patterns of the same category (for example, words) have different durations each time they are uttered or for each speaker, so speech patterns of different lengths are presented to the input layer of the same neural network. It is necessary to devise ways to do so.
第二に長さの異なる音声パターンをニューラル・ネット
ワークの入力に提示できたときに、ネットワークが期待
する認識動作を行なうようにユニット間結合を定める学
習方法を確立しなければならない。Second, when speech patterns of different lengths can be presented as input to a neural network, a learning method must be established to determine connections between units so that the network performs the expected recognition operations.
本発明は固定時間長の特徴パラメータ時系列を入力でき
る入力層を持つニューラル・ネットワークに長さの異な
る音声パターンを提示する為に認識時は出力層の出力が
最大になるように入力層の時間軸と入力音声時系列との
対応付けを行い、ユニット間結合係数を定める学習時に
は提示するパターンを固定継続時間長に正規化してネッ
トワークに提示して出力層での誤差を最小にする教師付
きの学習機構を有するダイナミック・ニューラル・ネッ
トワークを提供しようとするものである。In order to present speech patterns of different lengths to a neural network having an input layer that can input feature parameter time series with a fixed time length, the time of the input layer is adjusted so that the output of the output layer is maximized during recognition. A supervised system that associates the axes with the input audio time series and determines the inter-unit coupling coefficient, normalizes the presented pattern to a fixed duration length and presents it to the network to minimize the error in the output layer. This paper attempts to provide a dynamic neural network with a learning mechanism.
(問題点を解決するための手段)
本発明は音声等の時系列パターンを認識するニューラル
・ネットワークで、入力・出力層と複数の中間層から構
成される階層構造を有し、更に入力層と中間層が時間軸
に対応する時系列的構造を有し、認識時には動的計画法
によって入力時系列パターンの時間軸をニューラル・ネ
ットワークの出力が最大になるように入力層の持つ時間
軸と対応付けを行い、その時の出力層の出力を認識結果
とするダイナミック・ニューラル・ネットワークに於て
、その各階層間のユニット間結合係数を学習するに際し
て、入力層の時間長と同じ一定の継続時間長に線形伸縮
によって正規化して入力層に提示し、出力層には対応し
て出力すべき教師信号を提示して、出力層での教師信号
と実際の出力値の差異を小さくするように結合係数を決
定する教師付き学習を行なう機構を有することを特徴と
する。(Means for Solving the Problems) The present invention is a neural network that recognizes time-series patterns such as speech, and has a hierarchical structure consisting of an input/output layer and a plurality of intermediate layers. The middle layer has a time-series structure corresponding to the time axis, and during recognition, dynamic programming is used to match the time axis of the input time-series pattern with the time axis of the input layer so that the output of the neural network is maximized. In a dynamic neural network where the recognition result is the output of the output layer at that time, when learning the coupling coefficient between units between each layer, a constant time length that is the same as the time length of the input layer is used. is normalized by linear expansion/contraction and presented to the input layer, and the corresponding teaching signal to be output is presented to the output layer, and the coupling coefficient is set so as to reduce the difference between the teaching signal and the actual output value in the output layer. It is characterized by having a mechanism that performs supervised learning to determine.
(作用)
本発明の詳細な説明を簡単のために中間層を1層にした
3層構造のモデルを用いて行なう。中間層が2層以上の
場合にも同様に適用できることは言うまでもない。(Function) For the sake of simplicity, the present invention will be described in detail using a three-layer structure model with one intermediate layer. Needless to say, the present invention can be similarly applied to cases where there are two or more intermediate layers.
モデルの入力層はP次元の特徴ベクトルの時系列(長さ
J)を受は取ることができるようにJ×P個のユニッか
ら構成されている。この入カニニットの出力値をy ”
’J(P) (j=1〜J、 p・1〜P)とする。一
般には入力層の時間軸の長さJと認識時に入力される入
力時系列パターンat(p>(i・1〜1.p・1〜P
)の長さIは異なるので、入力時系列の時間軸になんら
かの伸縮変換を施して長さJに揃えなければならない。The input layer of the model is composed of J×P units so that it can receive a time series (length J) of P-dimensional feature vectors. The output value of this input crab unit is y”
'J(P) (j=1~J, p・1~P). In general, the length J of the time axis of the input layer and the input time series pattern at(p>(i・1~1.p・1~P
) are different, so the time axis of the input time series must be subjected to some expansion/contraction transformation to make it equal to the length J.
入力層の時間軸jと入力時系列パターンの時間軸iで構
成される平面(i、j)上での対応関係を次式で表わす
。The correspondence relationship on the plane (i, j) formed by the time axis j of the input layer and the time axis i of the input time series pattern is expressed by the following equation.
c(k) ・ (i(kン、 j(k)) 、
(k4〜K) =111但し、
c(1) ・ (1,1)、 c(k) = <
1.J)i(k)≧1(k−1>、 j(k)≧j(k
−1) ・・・+21この関係を用いて入カニ
ニットの出力値y(1)、。+(p) (J・1〜J
、 p=l〜P〉はy(1ゝjt k+(p)
” az k+(+)) 、
=i3)と表わされる。即ち、入カニニッ
トは時間軸を整合して入力されたデータをそのまま次の
層へ伝達することになる。c(k) ・(i(kn, j(k)),
(k4~K) = 111 However, c(1) ・ (1, 1), c(k) = <
1. J) i(k)≧1(k-1>, j(k)≧j(k
-1) ...+21Using this relationship, the output value y(1) of the input crab unit. +(p) (J・1~J
, p=l~P> is y(1ゝjt k+(p)
” az k+(+)),
=i3). That is, the input unit aligns the time axis and transmits the input data as is to the next layer.
中間層はJXM個のユニット(隠れユニットと呼ぶ)か
ら構成され、各ユニットへの入力値X ”’J(1)
(j・1〜J、 m・1〜M)は入カニニットの出力
値y + 1 ) j(p)と入カニニットと隠れユニ
ットの間の結合係数βOt(m、p)、 β’J(I
IP)を用いて次式のように与えられる。The middle layer is composed of JXM units (called hidden units), and the input value to each unit is
(j・1~J, m・1~M) is the output value y + 1) j(p) of the input crab unit and the coupling coefficient between the input crab unit and the hidden unit βOt(m, p), β'J(I
IP) is given as follows.
X J(鵬):Σ くβ jくLp)y
j(k)(p)+β jD+1
<jp)3””j+に−t+<p)ン
azb−t+(P)) ・
・・4)このようにj(k)番目の隠れユニットは入力
層の1(k)番目と1(k−1)番目のユニットからだ
け情報を受は取るようにユニット間結合を制限したニュ
ーラル・ネットワークの構造を時系列構造と呼ぶことに
する。このようなネットワークの構造は音声パターン等
のようにデータ自体が時系列的な構造を持っている場合
には、完全結合くすべての入カニニットとすべての隠れ
ユニットを結ぶ)に比べて少ないユニット間結合でモデ
ルが構成できるので、認識・学習時の計算量を大幅に削
減することができる。4)式で与えられる入力に対する
隠れユニットの応答は次のようになる。X J (Peng): Σ ku β j ku Lp) y
j(k)(p)+β jD+1 <jp)3""j+to -t+<p) azb-t+(P)) ・
...4) In this way, the j(k)th hidden unit is a neural network that restricts the connections between units so that it only receives information from the 1(k)th and 1(k-1)th units in the input layer.・The structure of the network will be called the time-series structure. When the data itself has a time-series structure, such as a speech pattern, the structure of such a network requires fewer connections between units than a completely connected network (connecting all incoming units and all hidden units). Since a model can be constructed by combining, the amount of calculation during recognition and learning can be significantly reduced. 4) The response of the hidden unit to the input given by equation is as follows.
y (2’>(m)・f < x (2’t(
m)−θ (2)、く■)) ・・・(9
f (x> =1z’ (1+e−X)
=161ここでθ(21j(1)は隠れユニット(
j、m)が持つ閾値である。式6から明らかなように隠
れユニットは一種の閾値論理の働きをしている。y (2'>(m)・f < x (2't(
m)-θ (2),ku■)) ...(9
f (x>=1z' (1+e-X)
= 161 where θ(21j(1) is the hidden unit (
j, m) has a threshold value. As is clear from Equation 6, the hidden unit functions as a kind of threshold logic.
出力層は認識対象となるN個のカテゴリに対応するN個
のユニットから構成されている。n番目の出カニニット
への入力値x ”’(n>(n=l〜N)は隠れユニッ
トの出力値y”’J(1)と隠れユニットと出カニニッ
トの間の結合係数α″(j、i+)を用いて次式のよう
に与えられる。The output layer is composed of N units corresponding to N categories to be recognized. The input value x ''(n>(n=l~N) to the n-th output unit is the output value y''J(1) of the hidden unit and the coupling coefficient α''(j , i+) is given as follows.
出カニニットの入出力の応答間係は式2と同じである。The response relationship between the input and output of the output unit is the same as Equation 2.
y(3’(n)” f (x (3)(n)−θ(3J
n)) 、、−Uここでθ(3′(n)は出カ
ニニットnの持つ同値である。y(3'(n)" f (x (3)(n)-θ(3J
n)) ,, -U where θ(3'(n) is the equivalent value of output unit n.
こうして得られるネットワークの出力値y(31(n)
は式1で与えられている入力時系列の時間軸と入カニニ
ット層の時間軸の対応関係(C(k))に依存している
。最終的なカテゴリnのネットワークによる認識結果は
(c (k> )に関して最適化された(最大化された
)出力値Onとして得られる。The output value of the network obtained in this way y(31(n)
depends on the correspondence relationship (C(k)) between the time axis of the input time series and the time axis of the input crabnit layer given by Equation 1. The final recognition result by the network for category n is obtained as an optimized (maximized) output value On with respect to (c (k>)).
ここで8式は単調関数なので(9)式は+β1.〈曹)
al、−口))1 ・−・(1o)と置き
換えても同じである。ここでfNの中の特徴ベクトルの
成分pに関する和は省略した。Here, since equation 8 is a monotone function, equation (9) is +β1. (Cao)
It is the same even if it is replaced with al, -口))1 .--(1o). Here, the sum of the component p of the feature vector in fN is omitted.
(10)式の()の中の式を
7 (c(k)、 c(k−1))= (−)
=・くll)と定
義すると、(10)式は
o、=max[Σγ(c(k)、 c(k−1>)]
−(12)lcfkll k冨l
となり、この最適化は良く知られた動的計画法を用いて
解くことができることが分かる。即ち、γ(c(k)、
c(k−1))の累積和をg(k)として、次の漸化式
を計算してo 、=g(k)を求めればよい。(10) Expression in parentheses is 7 (c(k), c(k-1))= (-)
=・kull), then equation (10) becomes o, =max[Σγ(c(k), c(k-1>)]
-(12) lcfkll ktl , and it can be seen that this optimization can be solved using the well-known dynamic programming method. That is, γ(c(k),
The cumulative sum of c(k-1)) is set as g(k), and the following recurrence formula is calculated to obtain o,=g(k).
g<k)= max[7(c(k>、c(k−1))+
g(k−1)] =−(13)c(k−11
次にニューラル・ネットワーク・モデルのパラメータで
あるユニット間結合係数(β0.(■、p)。g<k)=max[7(c(k>,c(k-1))+
g(k-1)] =-(13)c(k-11) Next, the inter-unit coupling coefficient (β0.(■, p), which is a parameter of the neural network model.
β1.(―・p)、α″(j、m)と閾値(θ(21,
<II)lθ(3)<n))を決定する学習法について
説明する。β1. (--p), α″(j, m) and threshold value (θ(21,
<II) A learning method for determining lθ(3)<n)) will be explained.
カテゴリnの学習に用いる特徴ベクトルの時系列の組を
A”’q・ (a ’q、 1(p) )とする。ここ
でqは同じカテゴリ内の複数の時系列パターンを区別す
る添字、iは時系列の時間軸を表わす添字、pは各時刻
での特徴ベクトルの成分を表わす添字である。各添字の
範囲は
n =1〜N、 q −1〜Q”、 i =1〜I
’、 p=1〜P ・=<14)ネットワークにこの
データA (n +9を提示する為には時系列の長さI
Qをネッワークの入力層の時間軸の長さJに正規化しな
ければならない。学習時にはモデルのパラメータが最適
化されていないので、認識時のように動的計画法を用い
ることは難しい。Let the set of time series of feature vectors used for learning category n be A"'q (a 'q, 1(p)). Here, q is a subscript that distinguishes multiple time series patterns within the same category. i is a subscript that represents the time axis of the time series, and p is a subscript that represents the component of the feature vector at each time.The range of each subscript is n = 1 to N, q -1 to Q'', i = 1 to I
', p = 1 ~ P ・= < 14) In order to present this data A (n + 9) to the network, the time series length I
Q must be normalized to the length J of the time axis of the input layer of the network. Since the model parameters are not optimized during learning, it is difficult to use dynamic programming as during recognition.
そこで学習の為にはカテゴリnのデータの集合A3″″
、(q・1〜Q″)の中から代表となる時系列パターン
ALfilqoを選び出し、それ以外のデータA(II
)q(q≠qo)の時間軸をDPマツチングによって記
代表パターンの時間軸に対応付ける。その方法を次に示
す。代表パターンA+I11.oの時間をj(j・1〜
J)、時間軸の対応付け(正規化)を行ないたいデータ
A”’、(q≠qo)の時間軸をi(i・1〜■)とす
る。このとき2つのパターンをDPマツチングすること
によって2つのパターンの時間軸の間の対応関係(歪関
係)i=iりj)が得られる。Therefore, for learning, we need a set of data of category n A3''''
, (q・1~Q″), the representative time series pattern ALfilqo is selected, and the other data A(II
) q (q≠qo) is associated with the time axis of the representative pattern by DP matching. The method is shown below. Representative pattern A+I11. Let the time of o be j(j・1~
J) Let i (i・1~■) be the time axis of data A"' (q≠qo) for which you want to perform time axis matching (normalization). At this time, perform DP matching of the two patterns. A correspondence relationship (distortion relationship) i=i r j) between the time axes of the two patterns can be obtained.
DPマツチングと歪関数に関しては「日経エレクトロニ
クス」誌、第329号の第171頁(昭和58年11月
7日発光)に詳しく解説されている(以下、文献2と呼
ぶ)。この歪関数i<j)によって代表パターンの時間
軸jには学習データの時間軸1=i(j)のフレーム・
ベクトルa’Q、li)を対応付ければ良いことが分か
る。この歪み関数はDPマツチングに用いる局所的な径
路の制限の仕方によっては1=j(i)のような形にな
り、あるjに対応するフレーム・ベクトルが複数存在す
ることが起こるが、このような場合にも対応するフレー
ム・ベクトルを平均化することによって同様の時間軸対
応付けが行える。DP matching and distortion functions are explained in detail in "Nikkei Electronics" magazine, No. 329, page 171 (published on November 7, 1982) (hereinafter referred to as Document 2). Due to this distortion function i<j), the time axis j of the representative pattern has a frame with the time axis 1=i(j) of the learning data.
It can be seen that it is sufficient to associate the vectors a'Q, li). This distortion function takes the form 1=j(i) depending on how the local path used for DP matching is restricted, and there may be multiple frame vectors corresponding to a certain j. Even in such cases, similar time axis correspondence can be achieved by averaging the corresponding frame vectors.
この結果、データ毎にばらついていた時間長IQが一定
の長さI″。に正規化される。ネットワークの入力層の
時間長JはこのIQOに等しく設定する。As a result, the time length IQ, which varies from data to data, is normalized to a constant length I''. The time length J of the input layer of the network is set equal to this IQO.
ここでカテゴリーnの代表パターンの選び方としては様
々な方法が考えられるが、例えばカテゴリnのパターン
集合の中でパターン間のDPマツチングによる累積距離
d (A、0. Aq>をパターン間距離として、次式
で与えられる量Δ、を最小にするようなQoとする。こ
のq。はすべての9・1〜Q″をQoと仮定してΔを計
算する総当たり法によって容易に求めることができる。Here, various methods can be considered to select the representative pattern of category n, but for example, the cumulative distance d (A, 0. Aq> is the distance between patterns) by DP matching between patterns in the set of patterns of category n. Let Qo be such that the quantity Δ given by the following equation is minimized. This q can be easily determined by the brute force method that calculates Δ assuming all 9.1 to Q″ as Qo. .
この池にも任意の1パターンを代表にすることも可能で
ある。It is also possible to make any one pattern representative of this pond.
こうして時間軸を長さJに正規化した入力学習データを
A(fi’ q= (α″q、 +(p) ) (i・
1〜J)とする。また、同じ長さJに正規化された他の
カテゴリの学習データをB(″’−・(b −r、 +
(p) ) (r・1〜R)とする(以後このBを反学
習データと呼ぶ)。In this way, the input learning data with the time axis normalized to length J is A(fi' q= (α″q, +(p) ) (i・
1 to J). In addition, the learning data of other categories normalized to the same length J are B(″'−・(b −r, +
(p) ) (r・1~R) (hereinafter, this B will be referred to as anti-learning data).
このときq番目の学習データに対するネットワークの出
力値をy+s+、、(n)、望ましい出力値をZQ(n
)(・1.0)、r番目の反学習データに対する第nユ
ニットの出力値をy(3’r(n) 、望ましい出力値
をz 、(n)(・0.0)とすると、出カニニット層
に於ける出力値の誤差Eは
E=1/2Σ [3’ 、<n)−z q(n)
]q=1
+1/2Σ [3’ 、(n)−z r(n
)] −・−(16)r=1
で与えられる。この誤差量Eは学習によって決定しなけ
ればならないユニット間結合係数(β0」(m、p)、
β’>(m、p)、α”(j、m))と閾値(θ(21
,(■)。At this time, the output value of the network for the qth learning data is y+s+,,(n), and the desired output value is ZQ(n
)(・1.0), the output value of the nth unit for the rth unlearning data is y(3'r(n), and the desired output value is z, (n)(・0.0), then the output is The error E of the output value in the crab knit layer is E=1/2Σ[3',<n)-z q(n)
]q=1 +1/2Σ [3', (n)-z r(n
)] −・−(16) r=1. This error amount E is determined by the inter-unit coupling coefficient (β0'(m, p)), which must be determined by learning.
β′>(m, p), α”(j, m)) and threshold value (θ(21
, (■).
θ”(n))の関数と考えられるのでEを評価関数とし
て最小化するようにこれらのパラメータを決定すればよ
い。またユニットの閾値は常に1を出力するユニットを
仮想的に考えて、そのユニットとの結合係数と考えれば
ユニット間結合と同じように学習することができる。そ
こで隣接する2層、第n層のユニットiと第n+1層の
ユニットjを結ぶユニット間結合係数をω″′1とする
と、このωn、Jに関するEの微係数を用いてとすれば
、必ず、
E (t+1)≦E(t) ・
・・(18)となる。ここでtは繰り返し学習のステッ
プを表わす整数値、εは修正の程度を決める定数である
。結局、Eを小さくするようにωfl。を繰り返し修正
することがパラメータの学習になるのである。ここでω
0目と前記モデルのユニット間結合係数(β0J(m、
P)β’J(1,P)、α”(j、l)、θ(21j(
1)、θ”(nNとは例えば次のように対応付ければよ
い。θ”(n)), so these parameters can be determined to minimize E as the evaluation function.Also, the threshold value of the unit is determined by hypothetically considering a unit that always outputs 1. If you think of it as a coupling coefficient with a unit, you can learn it in the same way as inter-unit coupling.Then, the inter-unit coupling coefficient that connects the unit i of the adjacent 2nd layer, the nth layer, and the unit j of the n+1th layer is ω''' 1, and if we use the differential coefficient of E with respect to ωn and J, we will definitely have E (t+1)≦E(t) ・
...(18). Here, t is an integer value representing the step of iterative learning, and ε is a constant that determines the degree of correction. In the end, ωfl is used to reduce E. Parameter learning is achieved by repeatedly modifying the parameters. Here ω
Unit coupling coefficient (β0J(m,
P) β'J(1, P), α''(j, l), θ(21j(
1), θ” (nN may be associated with each other as follows, for example.
Eの微係数は解析的な計算の結果次式のようになること
が分かる。As a result of analytical calculation, the differential coefficient of E is found to be as shown in the following equation.
ここでδ″+1ゝ1.Qはq番目の学習(または反学習
)データを入力層に提示した場合の第n+1層のユニツ
iの入力値に換算された誤差で、y(n)I、9はq番
目の学習データに対する第n層のユニッ1− jの出力
値である。δ(n l 、 、qは次のような漸化式を
用いて計算することができる。Here, δ″+1ゝ1.Q is the error converted to the input value of unit i in the n+1th layer when the qth learning (or anti-learning) data is presented to the input layer, and y(n)I, 9 is the output value of unit 1-j of the n-th layer for the q-th learning data. δ(n l , , q can be calculated using the following recurrence formula.
δ 3”ゝ31.. ・ Σ δ (o++1.、
、Qω ’ml (df/dx)、=x”′lここで
f (x)は式6で与えられるユニットの入出力応答関
数で、x311は第n層のユニットiへ人力値、Zlは
第N層(出力層)のユニットiがとるべき値で学習の時
には1.0で反学習の時には0.0である。この式21
に基づいて、各ユニットに換算された誤差量δを求める
計算が出力層から入力層の方向に進むので、この学習法
は逆伝播学習法(バック・プロパゲーション学習法)と
呼ばれている(詳細は文献1を参照のこと)。δ 3”ゝ31.. ・Σ δ (o++1.,
, Qω 'ml (df/dx), = x'''l where f (x) is the input/output response function of the unit given by Equation 6, x311 is the human input value to unit i in the nth layer, and Zl is the input/output response function of the unit given by Equation 6. The value that unit i of the N layer (output layer) should take is 1.0 during learning and 0.0 during anti-learning.Equation 21
This learning method is called a back-propagation learning method because the calculation to find the converted error amount δ for each unit proceeds from the output layer to the input layer based on For details, see Reference 1).
結局、ユニット間結合係数に任意の初期値を与えたモデ
ルから出発して、複数の学習・反学習データを提示して
、各ユニ・ソト間結合に関して上記の繰り返し訂正学習
を行なえば、出力層での誤差を極小化するユニ・ソト間
結合の組を得ることができる。After all, if we start from a model in which arbitrary initial values are given to the inter-unit coupling coefficients, present multiple learning/unlearning data, and perform the above-mentioned iterative correction learning for each uni/sotho coupling, the output layer We can obtain a set of uni-sotho connections that minimize the error in .
(実施例)
以下に式13の漸化式計算の為の(i、j)平面上での
時間軸対応付は規則(c(k)とc(k−1)の相対位
置関係)として第3図のような規則を用いた場合の本発
明の詳細な説明する。第3図の場合はc(k)=(i、
j)とするとc(k−1)としては(i−1,j) 、
(i−1,j−1) 、(i−1,j−2>の3点だけ
が可能になる。このような対応付は規則の場合にはニュ
ーラル・ネ・ソトワークの出力を決める(12)、(1
3)式は次のように書ける。(Example) The time axis correspondence on the (i, j) plane for calculating the recurrence formula of Equation 13 is as follows as a rule (relative positional relationship between c(k) and c(k-1)). The present invention will be described in detail when the rules as shown in FIG. 3 are used. In the case of Fig. 3, c(k)=(i,
j), then c(k-1) is (i-1, j),
Only three points (i-1, j-1) and (i-1, j-2> are possible. In the case of rules, such a correspondence determines the output of the neural neural network (12 ), (1
3) The formula can be written as follows.
γ” (i、j) ・ Σ α (j、曹)f
(β i<m)at+β 。γ” (i, j) ・Σ α (j, Cao) f
(β i < m)at+β.
+!1為l
(m) at−1)
−・・(23)g ” <i、j> =7
”(i、j)+max [g ’(i−1,j)。+! 1 for l (m) at-1)
−...(23) g ” <i, j> = 7
”(i,j)+max[g'(i-1,j).
g ’ (i−1,j−1)、 g fl(i−1,j
−2)] ・・・(24)第1図は(22)
〜り24)式に基づいて本発明を実現した一実施例を示
したブロック図である。分析部10は入力された音声波
形データを分析して特徴ベクトルの時系列に変換して、
パターンバッファ部20に記憶する。パターンバ・ソフ
ァ部20には学習動作時には学習用時系列データが記憶
され、認識動作時には未知発声の分析データが記憶され
る。続く切り替えスイッチによって学習動作と認識動作
の切り替えを行なう。g' (i-1, j-1), g fl (i-1, j
-2)] ...(24) Figure 1 is (22)
FIG. 24 is a block diagram showing an embodiment of the present invention based on formulas 24) to 24). The analysis unit 10 analyzes the input audio waveform data and converts it into a time series of feature vectors.
The data is stored in the pattern buffer section 20. The pattern bass section 20 stores learning time series data during a learning operation, and stores analysis data of unknown utterances during a recognition operation. The following changeover switch switches between learning operation and recognition operation.
時間軸整合部30は学習データ群中の各カテゴリの代表
パターンを式15に基づいて決定して、他の学習データ
の時間軸を代表パターンへDPマツチングすることによ
って整合し、すべての学習データの時間長を長さJへ規
格化する。修正量計算部40は時間軸整合部30から送
られた学習データとユニット間結合係数記憶部50に蓄
えられた結合係数を用いて、式17.20.21に基づ
いて結合係数ω″1.の修正量Δω″’IJを算出して
、結合係数修正部60に送る。結合係数修正部60はユ
ニット間結合係数記憶部50に蓄えられた結合係数に前
記修正量Δω″。を加えて、書き戻す。修正量計算部4
0はすべての結合係数に対する修正量Δω″IJが予め
定められた閾値より小さくなるまでか、あるいはIC正
回数が予め定められた回数を越えるまで、この修正動作
を繰り返す。The time axis matching unit 30 determines a representative pattern for each category in the learning data group based on Equation 15, performs DP matching on the time axes of other learning data to the representative pattern, and matches all the learning data. Standardize the time length to length J. The correction amount calculation unit 40 uses the learning data sent from the time axis alignment unit 30 and the coupling coefficients stored in the inter-unit coupling coefficient storage unit 50 to calculate the coupling coefficient ω″1. A correction amount Δω″'IJ is calculated and sent to the coupling coefficient correction unit 60. The coupling coefficient correction unit 60 adds the correction amount Δω″ to the coupling coefficient stored in the inter-unit coupling coefficient storage unit 50 and writes it back. Correction amount calculation unit 4
0 repeats this correction operation until the correction amount Δω''IJ for all coupling coefficients becomes smaller than a predetermined threshold, or until the number of correct ICs exceeds a predetermined number.
格子点計算部70はパターンバッファ部20から送られ
た未知発声データとユニット間結合係数記憶部50に蓄
えられた結合係数を用いて、式23に基づいて格子点デ
ータγ’(i、j)(i・1〜I、j・1〜J、 n・
1〜N)を計算する。計算された格子点データは格子点
記憶部80に格納される。漸化式計算部90は格子点記
憶部80に蓄えられた格子点データを用いて、式24に
基づく漸化式計算を行なって累積値g ”(1,J)を
作業用記憶部100に格納する。作業用記憶部100は
漸化式計算途中にもg’(i、j>の記憶に用いられる
。認識判定部110は作業用記憶部100に格納された
累積値g ’<1゜J)の中から最大の累積値を与える
nの値を認識結果として出力する。The lattice point calculation section 70 uses the unknown utterance data sent from the pattern buffer section 20 and the coupling coefficient stored in the inter-unit coupling coefficient storage section 50 to calculate the lattice point data γ'(i, j) based on Equation 23. (i・1~I, j・1~J, n・
1 to N). The calculated grid point data is stored in the grid point storage section 80. The recurrence formula calculation section 90 uses the grid point data stored in the grid point storage section 80 to perform recurrence formula calculation based on equation 24, and stores the cumulative value g''(1, J) in the working storage section 100. The working storage unit 100 is used to store g'(i, j> even during recurrence formula calculation. The recognition determining unit 110 stores the cumulative value g'<1° stored in the working storage unit 100. J), the value of n that gives the maximum cumulative value is output as the recognition result.
(発明の効果)
以上述べたように、本発明によれば認識動作時に未知音
声データの発声時間長の変動を動的計画法によって正規
化してニューラル・ネ・ソトワークに入力することがで
きる時間軸の正規化能力を有するニューラル・ネッワー
クを提供できる。(Effects of the Invention) As described above, according to the present invention, the time axis that allows the fluctuation of the utterance time length of unknown speech data to be normalized by dynamic programming and input into the neural network during the recognition operation. It is possible to provide a neural network with a normalization ability of
このように本発明のニューラル・ネットワークは認識動
作時に時間軸正規化能力を有するので、学習動作時には
音声データの発声毎の特徴パラメータの変動を少数の学
習データ(発声時間長の変動による多様性を持たなくて
よい)を用いて学習することによって、良好な認識装置
を提供することができる。In this way, the neural network of the present invention has the ability to normalize the time axis during the recognition operation, so during the learning operation, it is possible to reduce the variation in the feature parameters for each utterance of the audio data by using a small number of learning data (the diversity due to the variation in the utterance duration). It is possible to provide a good recognition device by learning using the following information.
第1図は本発明の一実施例を示すブロック図、第2図は
ニューラル・ネッワークの階層構造を表わす図、第3図
は漸化式計算の為の(i、n平面上での時間軸整合部は
規則の例を表わす図である。
図に於いて、10は分析部、20はパターンバッファ部
、30は時間軸整合部、40は修正量計算部、50はユ
ニッ1〜間結合係数記憶部、60は結合係数修正部、7
0は格子点計算部、80は格子点記憶部、90は漸化式
計算部、100は作業用記憶部、110は認識判定部で
ある。Figure 1 is a block diagram showing an embodiment of the present invention, Figure 2 is a diagram showing the hierarchical structure of a neural network, and Figure 3 is a diagram showing the time axis on the (i, n plane) for calculating recurrence formulas. The matching section is a diagram showing an example of rules. In the figure, 10 is an analysis section, 20 is a pattern buffer section, 30 is a time axis matching section, 40 is a modification amount calculation section, and 50 is a coupling coefficient between units 1 to 1. storage unit, 60 is a coupling coefficient correction unit, 7
0 is a lattice point calculation section, 80 is a lattice point storage section, 90 is a recurrence formula calculation section, 100 is a working storage section, and 110 is a recognition determination section.
Claims (3)
ネットワークで、入力・出力層と複数の中間層から構成
される階層構造を有し、更に入力層と中間層が時間軸に
対応する時系列的構造を有し、認識時には動的計画法に
よって入力時系列パターンの時間軸をニューラル・ネッ
トワークの出力が最大になるように入力層の持つ時間軸
と対応付けを行い、その時の出力層の出力を認識結果と
するダイナミック・ニューラル・ネットワークに於て、
その各階層間のユニット間結合係数を学習するに際して
、入力層の時間長と同じ一定の継続時間長に正規化した
学習用時系列パターンを入力層の時間長と同じ一定の継
続時間長に正規化した学習用時系列パターンを入力層に
提示し、出力層には対応して出力すべき教師信号を提示
して、出力層での教師信号と実際の出力値との差異を小
さくするように結合係数を決定する教師付き学習を行な
う機構を有するダイナミック・ニューラル・ネットワー
ク。(1) Neural technology that recognizes time-series patterns such as voices, etc.
The network has a hierarchical structure consisting of input/output layers and multiple intermediate layers, and the input layer and intermediate layers have a time-series structure corresponding to the time axis. During recognition, input is input using dynamic programming. In a dynamic neural network, the time axis of the time series pattern is associated with the time axis of the input layer so that the output of the neural network is maximized, and the output of the output layer at that time is the recognition result.
When learning the inter-unit coupling coefficient between each layer, the training time series pattern is normalized to a constant duration length that is the same as the time length of the input layer. The trained time series pattern is presented to the input layer, and the corresponding teacher signal to be output is presented to the output layer to reduce the difference between the teacher signal and the actual output value in the output layer. A dynamic neural network that has a mechanism for supervised learning to determine coupling coefficients.
キの正規化を、代表パターンへのDPマッチングによっ
て行なうことを特徴とする特許請求範囲第(1)項記載
のダイナミック・ニューラル・ネットワーク。(2) The dynamic neural network according to claim (1), wherein the variation in the duration of the learning time-series pattern is normalized by DP matching to a representative pattern.
クプロパゲーション学習法によって実現することを特徴
とする特許請求範囲第(1)項記載のダイナミック・ニ
ューラル・ネットワーク。(3) The dynamic neural network according to claim (1), wherein the supervised learning of the inter-unit coupling coefficients is realized by a backpropagation learning method.
Priority Applications (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP63029675A JPH01204172A (en) | 1988-02-09 | 1988-02-09 | Dynamic neural network having learning mechanism |
Applications Claiming Priority (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP63029675A JPH01204172A (en) | 1988-02-09 | 1988-02-09 | Dynamic neural network having learning mechanism |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| JPH01204172A true JPH01204172A (en) | 1989-08-16 |
Family
ID=12282687
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| JP63029675A Pending JPH01204172A (en) | 1988-02-09 | 1988-02-09 | Dynamic neural network having learning mechanism |
Country Status (1)
| Country | Link |
|---|---|
| JP (1) | JPH01204172A (en) |
Cited By (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JPH0581226A (en) * | 1991-09-19 | 1993-04-02 | A T R Shichiyoukaku Kiko Kenkyusho:Kk | Method for learning neural circuit network and device using the method |
| JPH06112931A (en) * | 1992-09-30 | 1994-04-22 | Victor Co Of Japan Ltd | Pre-processing method for digital signal |
-
1988
- 1988-02-09 JP JP63029675A patent/JPH01204172A/en active Pending
Cited By (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JPH0581226A (en) * | 1991-09-19 | 1993-04-02 | A T R Shichiyoukaku Kiko Kenkyusho:Kk | Method for learning neural circuit network and device using the method |
| JPH06112931A (en) * | 1992-09-30 | 1994-04-22 | Victor Co Of Japan Ltd | Pre-processing method for digital signal |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| EP0342630B1 (en) | Speech recognition with speaker adaptation by learning | |
| Li et al. | Robust automatic speech recognition: a bridge to practical applications | |
| EP0510632B1 (en) | Speech recognition by neural network adapted to reference pattern learning | |
| US5461696A (en) | Decision directed adaptive neural network | |
| KR100306848B1 (en) | A selective attention method using neural networks | |
| Lee et al. | Ensemble of jointly trained deep neural network-based acoustic models for reverberant speech recognition | |
| JPH06102899A (en) | Voice recognizer | |
| US5758021A (en) | Speech recognition combining dynamic programming and neural network techniques | |
| US5181256A (en) | Pattern recognition device using a neural network | |
| JP2000298663A (en) | Recognition device using neural network and learning method thereof | |
| Watrous | Speaker normalization and adaptation using second-order connectionist networks | |
| JPH0540497A (en) | Speaker adaptive speech recognizer | |
| JPH01241667A (en) | Dynamic neural network to have learning mechanism | |
| JPH01204171A (en) | Dynamic neural network having learning mechanism | |
| US5581650A (en) | Learning dynamic programming | |
| Levin et al. | Time-warping network: a hybrid framework for speech recognition | |
| Bedworth et al. | Comparison of neural and conventional classifiers on a speech recognition problem | |
| Zhang et al. | End-to-end models with auditory attention in multi-channel keyword spotting | |
| JPH01241668A (en) | Dynamic neural network to have learning mechanism | |
| Salmela et al. | Isolated spoken number recognition with hybrid of self-organizing map and multilayer perceptron | |
| Makino et al. | Recognition of phonemes in continuous speech using a modified LVQ2 method | |
| Yuk et al. | Robust speech recognition using maximum likelihood neural networks and continuous density hidden Markov models | |
| Byorick et al. | Isolated vowel recognition using linear predictive features and neural network classifier fusion | |
| KR0185755B1 (en) | Voice recognition system using neural net | |
| Castro et al. | The use of multilayer perceptrons in isolated word recognition |