JPH03157697A - Word recognizing system - Google Patents
Word recognizing systemInfo
- Publication number
- JPH03157697A JPH03157697A JP1298502A JP29850289A JPH03157697A JP H03157697 A JPH03157697 A JP H03157697A JP 1298502 A JP1298502 A JP 1298502A JP 29850289 A JP29850289 A JP 29850289A JP H03157697 A JPH03157697 A JP H03157697A
- Authority
- JP
- Japan
- Prior art keywords
- word
- neural network
- threshold
- additional learning
- word recognition
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Granted
Links
Abstract
Description
【発明の詳細な説明】
[産業上の利用分野コ
本発明は、家電機器を音声操作する等に好適な単語認識
システムに関する。DETAILED DESCRIPTION OF THE INVENTION [Field of Industrial Application] The present invention relates to a word recognition system suitable for voice operation of home appliances.
[従来の技術]
従来の単語認識システムは、特開昭63−229496
に記載される如く、以下の手順による。[Prior art] A conventional word recognition system is disclosed in Japanese Patent Application Laid-Open No. 63-229496.
According to the following procedure as described in .
■入力音声に含まれる単語に関する特徴を抽出する。■Extract features related to words contained in input speech.
■予め上記■と同様の方法で抽出しておいた辞書との距
離を計算する。■Calculate the distance to the dictionary extracted in advance using the same method as in (■) above.
■上記■の計算結果より、入力音声が辞書に登録してお
いたどの単語か判定する。■ Based on the calculation result of (■) above, determine which word the input voice is registered in the dictionary.
■今回入力話者に認識単語を知らせる。入力話者は、単
語認識システムによる誤認識の有無をチエツクする。■Inform the input speaker of the recognized word. The input speaker checks whether there is any misrecognition by the word recognition system.
■上記■のチエツクにより、誤認識が判明した場合、今
回入力話者の手入力により、今回誤認識した単語の特徴
量を新たな標準パターンとして追加登録する。■If misrecognition is found through the check in (■) above, the feature amount of the word that was misrecognized this time is additionally registered as a new standard pattern by manual input by the input speaker.
これにより、1つの登録単語について、既登録の標準パ
ターンに加え、追加登録の標準パターンが並存すること
になる。As a result, for one registered word, additionally registered standard patterns coexist in addition to already registered standard patterns.
〔発明か解決しようとする課M]
然しながら、上記従来の単語認識システムでは、下記■
〜■の問題点がある。[Section M to be invented or solved] However, in the above conventional word recognition system, the following ■
There are ~■ problems.
■標準パターン作成時から時間が経過するにつれ、認識
率が劣化する0例えば、3ケ月経過により、認識率は1
00.0%から85.0%に劣化する。■The recognition rate deteriorates as time passes from the time the standard pattern was created.For example, after 3 months have passed, the recognition rate deteriorates to 1.
It deteriorates from 00.0% to 85.0%.
■実時間処理が困難である。即ち、一定以上の認識率を
確保するためには複雑な特徴量を用いる必要があるが、
複雑な特徴量を抽出するためには複雑な処理装置が必要
であり、処理時間も多大となる。■Real-time processing is difficult. In other words, in order to ensure a recognition rate above a certain level, it is necessary to use complex features;
In order to extract complex feature quantities, a complex processing device is required, and the processing time is also large.
■誤認識の再発防止のため、誤認識に係る登録単語につ
いてその標準パターンを追加登録する毎に、標準パター
ンのメモリ容量が必要になる。- In order to prevent recurrence of misrecognition, memory capacity for the standard pattern is required each time a standard pattern is additionally registered for a registered word related to misrecognition.
■上記■にて標準パターンを追加登録する結果、認識時
に距離計算するパターン数が増え、認識処理時間が増大
化する。■As a result of additionally registering standard patterns in the above (■), the number of patterns to be calculated for distance during recognition increases, and the recognition processing time increases.
本発明は、経時的な認識率の劣化が極めて少なく、かつ
容易に実時間処理できる単語認識システムを得ることを
目的とする。An object of the present invention is to obtain a word recognition system that exhibits extremely little deterioration in recognition rate over time and that can be easily processed in real time.
又、本発明は、認識不能に係る登録単語を追加学習して
認識不能の再発を防止するに際し、メモリ容量の増加と
認識処理時間の増大化を必要とすることなく・、かつ容
易に実時間処理できる単語認識システムを得ることを目
的とする。In addition, the present invention does not require an increase in memory capacity or an increase in recognition processing time when additionally learning registered words related to unrecognizability to prevent recurrence of unrecognizability, and can easily be performed in real time. The purpose is to obtain a word recognition system that can process.
[課題を解決するための手段]
請求項1に記載の本発明は、ニューラルネットワークを
用いた単語認識システムであって、登録単語に対応する
出力ユニットの出力値に対し、単語認識用しきい値と追
加学習用しきい値とを設定し、上記出力値が単語認識用
しきい値より大なることを条件に、今回の入力音声単語
を登録単語と判定し、上記出力値が単語認識用しきい値
より大、かつ追加学習用しきい値より小なることを条件
に、今回の入力音声データを用いてニューラルネットワ
ークの追加学習を行なうようにしたものである。[Means for Solving the Problems] The present invention according to claim 1 is a word recognition system using a neural network, in which a word recognition threshold value is set for an output value of an output unit corresponding to a registered word. and a threshold for additional learning, and on the condition that the above output value is greater than the threshold for word recognition, the current input voice word is determined to be a registered word, and the above output value is determined as a registered word. Additional learning of the neural network is performed using the current input audio data, provided that it is greater than the threshold and smaller than the additional learning threshold.
請求項2に記載の本発明は、ニューラルネットワークを
用いた単語認識システムであフて、登録単語に対応する
出力ユニットの出力値に対し、単語認識用しきい値を設
定し、上記出力値が単語認識用しきい値より大なること
を条件に、今回の入力音声単語を登録単語と判定し、上
記出力値が単語認識用しきい値より小なることを条件に
、今回の入力音声単語を認識てきなかったことを表示し
、この表示に対して今回の話者により教示される今回の
入力単語データと今回の入力音声データとを用いてニュ
ーラルネットワークの追加学習を行なうようにしたもの
である。The present invention according to claim 2 is a word recognition system using a neural network, in which a word recognition threshold is set for an output value of an output unit corresponding to a registered word, and the output value is On the condition that the output value is greater than the word recognition threshold, the current input audio word is determined to be a registered word, and on the condition that the output value is smaller than the word recognition threshold, the current input audio word is determined as a registered word. What has not been recognized is displayed, and in response to this display, the neural network performs additional learning using the current input word data taught by the current speaker and the current input voice data. .
請求項3に記載の本発明は、ニューラルネットワークを
用いた単語認識システムであって、登録単語に対応する
出力ユニットの出力値に対し、単語認識用しきい値と追
加学習用しきい値とを設定し、上記出力値が単語認識用
しきい値より大なることを条件に、今回の入力音声単語
を登録単語と判定し、(A)上記出力値が単語認識用し
きい値より小なることを条件に、今回の入力音声単語を
認識できなかったことを表示し、この表示に対して今回
の話者により教示される今回の入力単語データと今回の
入力音声データとを用いてニューラルネットワークの追
加学習を行ない、(B)上記出力値が単語認識用しきい
値より大、かつ追加学習用しきい値より小なることを条
件に、今回の入力音声データを用いてニューラルネット
ワークの追加学習を行なうようにしたものである。The present invention according to claim 3 is a word recognition system using a neural network, which sets a word recognition threshold value and an additional learning threshold value to an output value of an output unit corresponding to a registered word. (A) The output value is smaller than the word recognition threshold, and (A) the output value is smaller than the word recognition threshold. Under the condition, it is displayed that the current input voice word could not be recognized, and in response to this display, the neural network is constructed using the current input word data taught by the current speaker and the current input voice data. Perform additional learning, and (B) perform additional learning of the neural network using the current input audio data, on the condition that the above output value is greater than the word recognition threshold and smaller than the additional learning threshold. This is what I decided to do.
請求項4に記載の本発明は、前記ニューラルネットワー
クへの入力として、
■音声の周波数特性の時間的変化、
■音声の平均的な線形予測係数、
■音声の平均的なPARCOR係数、
■音声の平均的な周波数特性、及びピッチ周波数、
■高域強調を施された音声波形の平均的な周波数特性、
並びに
■音声の平均的な周波数特性
のうちの1つ以上を使用するようにしたものである。The present invention according to claim 4 provides, as inputs to the neural network, (1) temporal changes in the frequency characteristics of the voice, (2) average linear prediction coefficients of the voice, (2) average PARCOR coefficients of the voice, and (2) the average PARCOR coefficient of the voice. Average frequency characteristics and pitch frequency, ■Average frequency characteristics of high-frequency emphasized audio waveforms,
and (1) one or more of the average frequency characteristics of voice is used.
請求項5に記載の本発明は、前記ニューラルネットワー
クが階層的なニューラルネットワークであるようにした
ものである。According to a fifth aspect of the present invention, the neural network is a hierarchical neural network.
[作用]
請求項1に記載の本発明によれば、下記(1)〜(5)
の作用効果がある。[Function] According to the present invention as set forth in claim 1, the following (1) to (5)
It has the function and effect of
(1)経時的な認識率の劣化が極めて少ない、このこと
は、後述する実験結果により確認されていることである
が、ニューラルネットワークが音声の時期差による変動
の影響を受けにくい構造をとることが可能なためと推定
される。(1) The deterioration of the recognition rate over time is extremely small. This is confirmed by the experimental results described below, and this is because the neural network has a structure that is not easily affected by fluctuations due to differences in speech timing. This is presumed to be because it is possible.
(2)ニューラルネットワークを構成する、登録単語に
対応する出力ユニットの出力値に対し、単語認識用しき
い値の他に、追加学習用しきい値を設けた。即ち、上記
出力値が単語認識用しきい値を超えて大なるものであり
、入力音声単語を登録単語と判定できるものであっても
、該出力値が該1i35認識用しきい値より大なる追加
学習用しきい値を超えるものでない場合には、今回の入
力音声データを用いてニューラルネットワークの追加学
習を行なう、これにより、ニューラルネットワークの認
識率が劣化する前に、常に該ニューラルネットワークを
更新し、結果として、経時的な認識率の劣化が極めて少
ない単語認識システムを構成できる。(2) In addition to the word recognition threshold, an additional learning threshold is provided for the output value of the output unit corresponding to the registered word, which constitutes the neural network. In other words, even if the output value is greater than the word recognition threshold and the input speech word can be determined as a registered word, the output value is greater than the 1i35 recognition threshold. If the additional training threshold is not exceeded, additional training is performed on the neural network using the current input audio data. This allows the neural network to be constantly updated before the recognition rate of the neural network deteriorates. As a result, it is possible to construct a word recognition system in which the deterioration of the recognition rate over time is extremely small.
(3)ニューラルネットワークは、原理的に、ネットワ
ーク全体の演算処理が単純且つ迅速である。(3) In principle, in a neural network, the calculation processing of the entire network is simple and quick.
(4)ニューラルネットワークは、原理的に、それを構
成している各ユニットが独立に動作しており、並列的な
演算処理が可能である。従って、演算処理が迅速である
。(4) In principle, each unit constituting a neural network operates independently, and parallel arithmetic processing is possible. Therefore, calculation processing is quick.
(5)上記(3)〜(4)により、単語認識システムを
複雑な処理装置によることなく容易に実時間処理できる
。(5) With the above (3) and (4), the word recognition system can be easily processed in real time without using a complicated processing device.
又、請求項2に記載の本発明によれば、上記(1)
(3)〜(5)の作用効果に加えて、下記(6)の作
用効果がある。Further, according to the present invention as set forth in claim 2, the above (1)
In addition to the effects (3) to (5), there is the effect (6) below.
(6)ニューラルネットワークにて今回の入力音声単語
が認識不能である時、入力話者の助力により今回の入力
単語データを教示されて該ニューラルネットワークの追
加学習を行ない、認識不能の再発を防止できる。この際
、追加学習は、二二一ラルネットワークの各ユニットの
変換関数及び重みを修正することによりなされるもので
あるため、メモリ容量の増加や認識処理時間の増大化を
必要とすることがない。(6) When the current input voice word is unrecognizable by the neural network, the current input word data is taught with the help of the input speaker, and additional learning is performed on the neural network, thereby preventing recurrence of unrecognizability. . At this time, additional learning is performed by modifying the conversion function and weight of each unit of the 221 ral network, so there is no need to increase memory capacity or recognition processing time. .
請求項3に記載の本発明によれば、上記(1)〜(6)
の作用効果がある。According to the present invention according to claim 3, the above (1) to (6)
It has the function and effect of
請求項4に記載の本発明によれば、上記(1)〜(6)
の作用効果に加えて、下記(7)の作用効果がある。According to the present invention according to claim 4, the above (1) to (6)
In addition to the above effects, there is the following effect (7).
(7)ニューラルネットワークへの入力として、請求項
4に記載の■〜■の各要素のうちの1つ以上を用いるか
ら、入力を得るための前処理が、従来の複雑な特徴量抽
出に対して、単純となり、この前処理に要する時間が短
くて足りる。(7) Since one or more of the elements ① to ② described in claim 4 are used as input to the neural network, the preprocessing for obtaining the input is different from conventional complex feature extraction. Therefore, the process is simple, and the time required for this preprocessing is short.
請求項5に記載の本発明によれば、上記(1)〜(7)
の作用効果に加えて、下記(8)の作用効果がある。According to the present invention according to claim 5, the above (1) to (7)
In addition to the effects described above, there is the effect described in (8) below.
(8)階層的なニューラルネットワークにあっては、現
在、後述する如くの簡単な学習アルゴリズム(パックプ
ロパゲーション)が確立されており、高い認識率を実現
できるニューラルネットワークを容易に形成できる。(8) Regarding hierarchical neural networks, a simple learning algorithm (pack propagation) as described later has been established, and it is possible to easily create a neural network that can achieve a high recognition rate.
[実施例]
第1図は本発明が適用された単語認識システムの一例を
示す模式図、第2図は音声処理部とニューラルネットワ
ークの一例を示す模式図、第3図は入力音声を示す模式
図、第4図はバンドパスフィルタの出力を示す模式図、
第5図はニューラルネットワークを示す模式図、第6図
は階層的なニューラルネットワークを示す模式図、第7
図はユニットの構造を示す模式図である。[Example] Fig. 1 is a schematic diagram showing an example of a word recognition system to which the present invention is applied, Fig. 2 is a schematic diagram showing an example of a speech processing unit and a neural network, and Fig. 3 is a schematic diagram showing an input speech. Figure 4 is a schematic diagram showing the output of the bandpass filter,
Figure 5 is a schematic diagram showing a neural network, Figure 6 is a schematic diagram showing a hierarchical neural network, and Figure 7 is a schematic diagram showing a hierarchical neural network.
The figure is a schematic diagram showing the structure of the unit.
本発明の具体的実施例の説明に先立ち、ニューラルネッ
トワークの構成、学習アルゴリズムについて説明する。Prior to describing specific embodiments of the present invention, the configuration of the neural network and the learning algorithm will be described.
U)ニューラルネットワークは、その構造から、第5図
(A)に示す階層的ネットワークと第5図(B)に示す
相互結合ネットワークの2種に大別できる0本発明は、
両ネットワークのいずれを用いて構成するものであって
も良いが、階層的ネットワークは後述する如くの簡単な
学習アルゴリズムか確立されているためより有用である
。U) Neural networks can be roughly divided into two types based on their structure: the hierarchical network shown in FIG. 5(A) and the interconnected network shown in FIG. 5(B).
Although it may be configured using either of the two networks, the hierarchical network is more useful because it has a simple learning algorithm as described below and has been established.
(2)ネットワークの構造
階層的ネットワークは、第6図に示す如く、入力層、中
間層、出力層からなる階層構造をとる。(2) Network Structure A hierarchical network has a hierarchical structure consisting of an input layer, an intermediate layer, and an output layer, as shown in FIG.
各層は1以上のユニットから構成される。結合は、入力
層→中間層→出力層という前向きの結合たけて、各層内
での結合はない。Each layer is composed of one or more units. The connections are made in the forward direction from the input layer to the middle layer to the output layer, and there are no connections within each layer.
(3)ユニットの構造
ユニットは第7図に示す如く脳のニューロンのモデル化
であり構造は簡単である。他のユニットから入力を受け
、その総和をとり一定の規則(変換関数)で変換し、結
果を出力する。他のユニットとの結合には、それぞれ結
合の強さを表わす可変の重みを付ける。(3) Structure of the unit The unit is a model of a neuron in the brain and has a simple structure as shown in FIG. It receives input from other units, sums it up, transforms it using a certain rule (conversion function), and outputs the result. Each connection with another unit is given a variable weight that represents the strength of the connection.
(4)学習(パックプロパゲーション)ネットワークの
学習とは、実際の出力を目標値(望ましい出力)に近づ
けることであり、−船釣には第7図に示した各ユニット
の変換関数及び重みを変化させて学習を行なう。(4) Learning (pack propagation) Network learning means bringing the actual output closer to the target value (desired output). Learn by making changes.
又、学習のアルゴリズムとしては、例えば、Rumel
hart、 D、E、’、McC1elland、
J、L、 and thePDP Re5ea
rch Group、 PARALLEL DI
STRIBUTEDPROCESSING、 the
MIT Press、 1986.に記載されているパ
ックプロパゲーションを用いることができる。Further, as a learning algorithm, for example, Rumel
hart, D.E.', McCelland,
J, L, and thePDP Re5ea
rch Group, PARALLEL DI
STRIBUTED PROCESSING, the
MIT Press, 1986. Pack propagation as described in .
以下、本発明の具体的な実施例について説明する。Hereinafter, specific examples of the present invention will be described.
単語認識システム10は、第1図に示す如く、音声入力
部11、音声処理部12、ニューラルネットワーク13
、判定部14、メモリ部15、ネットワーク制御部16
、機器制御部17、表示制御部18、表示部19、教示
部20を有して構成される。As shown in FIG. 1, the word recognition system 10 includes a voice input section 11, a voice processing section 12, and a neural network 13.
, determination unit 14, memory unit 15, network control unit 16
, a device control section 17, a display control section 18, a display section 19, and a teaching section 20.
(1)音声入力部11に登録音声を入力する。(1) Input the registered voice to the voice input section 11.
この時、学習単語を、「ショウメイ」、「エアコン」、
「カーテン」、「テレビ」、「ドア」の5単語とする。At this time, learn words such as "shoumei", "air conditioner",
The five words are "curtain,""television," and "door."
又、入力単語を、「ショウメイ」、「エアコン」、「カ
ーテン」、「テレビ」、「ドア」の5単話とする。Furthermore, input words are assumed to be five single words: "shoumei", "air conditioner", "curtain", "television", and "door".
(2)音声処理部12て、上記(1)の入力音声に簡単
な前処理を施す。(2) The audio processing unit 12 performs simple preprocessing on the input audio in (1) above.
前処理結果は、今回の単語認識のためにニューラルネッ
トワーク13に転送されるとともに、追加学習の可能性
に備えて、メモリ部15に転送される。The preprocessing results are transferred to the neural network 13 for the current word recognition, and are also transferred to the memory unit 15 in preparation for the possibility of additional learning.
(3)ニューラルネットワーク13は、下記■の学習動
作と下記■の評価動作を行なう。(3) The neural network 13 performs the following learning operation (2) and the following evaluation operation (2).
■学習
′Q録単語に対応する出力ユニットの目標出力値を(1
)、対応しない出力ユニットの目標出力値を特徴とする
特定話者の入力音声に、音声処理部12による前処理を
施し、この前処理結果をニューラルネットワーク13に
入力する。そして、ニューラルネットワーク13の出力
値(出力層を構成する各出力ユニットの出力値)が上記
目標値に近づくように、ニューラルネットワーク13の
各ユニットの変換関数及び重みを修正する。■Learning' Set the target output value of the output unit corresponding to the Q-recorded word to (1
), the input speech of a specific speaker characterized by the target output value of the uncorresponding output unit is subjected to preprocessing by the speech processing unit 12, and the preprocessing result is input to the neural network 13. Then, the conversion function and weight of each unit of the neural network 13 are corrected so that the output value of the neural network 13 (the output value of each output unit constituting the output layer) approaches the target value.
この学習動作を例えば1000回くり返す。This learning operation is repeated, for example, 1000 times.
■評価
今回話者の入力音声に前処理を施し、この前処理を施し
た音声をニューラルネットワーク13に入力し、ニュー
ラルネットワークの出力値を得る。■Evaluation This time, preprocessing is performed on the input speech of the speaker, and the preprocessed speech is input to the neural network 13 to obtain the output value of the neural network.
そして、ニューラルネットワーク13の各登録単語に対
応する出力ユニットの出力値(X)が判定部14に転送
される。Then, the output value (X) of the output unit corresponding to each registered word of the neural network 13 is transferred to the determination unit 14.
(4)判定部14は、ニューラルネットワーク13の出
力値(X)に対し、しきい値θ1、θ2(θ1〉θ2)
を設ける。(4) The determination unit 14 sets threshold values θ1 and θ2 (θ1>θ2) for the output value (X) of the neural network 13.
will be established.
Olは追加学習用しきい値、02は単語認識用しきい値
である。Ol is a threshold for additional learning, and 02 is a threshold for word recognition.
判定部14は、上記しきい値を用いて、下記■〜■の判
定動作を行なう。The determination unit 14 performs the following determination operations (1) to (2) using the above threshold value.
■[X>θ2]
であることを条件に、判定部14は、今回の入力音声単
語を登録単語と判定し、この登録単語判定信号を機器制
御部17に出力する。(2) Under the condition that [X>θ2], the determining unit 14 determines the current input voice word as a registered word, and outputs this registered word determination signal to the device control unit 17.
■[X<02]
であることを条件に、判定部14は、今回の入力音声単
語を認識できなかった旨の認識不能信号を表示制御部1
8に出力する。■ [X<02] On the condition that
Output to 8.
■上記■の登録単語判定時に限り、判定部14は、更に
次の(a) 、 (b)の処理を行なう。(2) Only when determining registered words in (2) above, the determination unit 14 further performs the following processes (a) and (b).
(a)[X<θ1]
であることを条件に、判定部14は、今回の入力音声デ
ータを用いてニューラルネットワーク13の追加学習を
行なうべく、ネットワーク制御部16に追加学習実行信
号を出力する。(a) On the condition that [X<θ1], the determination unit 14 outputs an additional learning execution signal to the network control unit 16 in order to perform additional learning of the neural network 13 using the current input audio data. .
(b)[X>01] である時、判定部14は何もしない。(b) [X>01] When this is the case, the determination unit 14 does nothing.
(5)機器制御部17は、判定部14による上記■の判
定結果に基づく登録単語判定信号により、機器を制御す
る。(5) The device control unit 17 controls the device using the registered word determination signal based on the determination result of the above-mentioned (2) by the determination unit 14.
この機器は、例えば照明器具であり、上記登録単語判定
信号に基づいて点灯制御を行なう。This device is, for example, a lighting fixture, and performs lighting control based on the registered word determination signal.
(6)ネットワーク制御部16は、判定部14による上
記■の判定結果に基づく追加学習実行信号により、ニュ
ーラルネットワーク13の追加学習を行なうことを判断
する。この時、ネットワーク制御部16は、メモリ部1
5より、今回の入力音声データを取出し、この入力音声
データをニューラルネットワーク13に再入力し、この
入力に対するニューラルネットワーク13の出力値(X
)が今回認識済の入力単語に対応する前述(3)■の登
録単語についての目標値(1)に近づくように、ニュー
ラルネットワーク13の各ユニットの変換関数及び重み
を修正する。ネットワーク制御部16は、この追加学習
動作を例えば1000回くり返す。(6) The network control unit 16 determines to perform additional learning of the neural network 13 based on the additional learning execution signal based on the determination result of the above-mentioned (2) by the determining unit 14. At this time, the network control unit 16 controls the memory unit 1
5, extract the current input audio data, re-input this input audio data to the neural network 13, and calculate the output value of the neural network 13 for this input (X
) is closer to the target value (1) for the registered word in (3) (3) above, which corresponds to the currently recognized input word. The network control unit 16 repeats this additional learning operation, for example, 1000 times.
(7)表示制御部18は、判定部14による上記■の判
定結果に基づく認識不能信号により、表示部19を駆動
し、今回の入力音声単語を認識てきなかったことを表示
し、これを話者に知らしめる。(7) The display control unit 18 drives the display unit 19 based on the unrecognized signal based on the determination result of the above-mentioned (■) by the determination unit 14 to display that the current input voice word has not been recognized, and to display the unrecognized word. make the person aware.
同時に、表示制御部18は、表示部19を駆動し、全登
録単語(前述の5単語)を順次表示する。これに対し、
今回の話者は、教示部20を手操作して、自らが今回入
力して認識不能であった入力gt語を表示部19の表示
単語データから特定し、この入力単語データを表示制御
部18に送信する。At the same time, the display control section 18 drives the display section 19 to sequentially display all registered words (the five words mentioned above). In contrast,
The present speaker manually operates the teaching unit 20 to identify the input gt word that he or she has input this time and is unrecognizable from the displayed word data on the display unit 19, and transfers this input word data to the display control unit 18. Send to.
表示制御部18は、教示部20から送信された入力単語
データを判断し、この認識不能であった単語についての
追加学習実行信号をネットワーク制御部16に出力する
。The display control unit 18 determines the input word data transmitted from the teaching unit 20 and outputs an additional learning execution signal for the unrecognized word to the network control unit 16.
(8)ネットワーク制御部16は、表示制御部18によ
る上記(7)の制御結果に基づく追加学習実行信号によ
り、ニューラルネットワーク13の追加学習を行なうこ
とを判断する。この時、ネットワーク制御部16は、今
回の話者により教示された今回の入力単語データと、メ
モリ部15より取出した今回の入力音声データとに基づ
き、今回の人力音声データをニューラルネットワーク1
3に再入力し、この入力に対するニューラルネットワー
ク13の出力値(X)が今回認識不能であった入力単語
に対応する前述(3)■の登録単語についての目標値(
1)に近づくように、ニューラルネットワーク13の各
ユニットの変換関数及び重みを修正する。ネットワーク
制御部16は、この追加学習動作を例えば1000回く
り返す。(8) The network control unit 16 determines to perform additional learning of the neural network 13 based on the additional learning execution signal based on the control result of the above (7) by the display control unit 18. At this time, the network control unit 16 transfers the current human voice data to the neural network 1 based on the current input word data taught by the current speaker and the current input voice data retrieved from the memory unit 15.
3, and the output value (X) of the neural network 13 for this input is set to the target value (
The conversion function and weight of each unit of the neural network 13 are modified so as to approach 1). The network control unit 16 repeats this additional learning operation, for example, 1000 times.
以下、第2図に示す如く、階層的なニューラルネットワ
ーク13を用い、ニューラルネットワーク13の入力と
して音声の一定時間内における平均的な周波数特性の時
間的変化を用いた場合の具体的実施例について説明する
。Hereinafter, as shown in FIG. 2, a specific example will be described in which a hierarchical neural network 13 is used and temporal changes in the average frequency characteristics of audio within a certain period of time are used as input to the neural network 13. do.
尚、音声処理部12は、第2図に示す如く、ローパスフ
ィルタ21、バンドパスフィルタ22、平均化回路23
の結合にて構成される。Note that the audio processing section 12 includes a low-pass filter 21, a band-pass filter 22, and an averaging circuit 23, as shown in FIG.
It is composed of the combination of
■入力音声の音声信号の高域成分を、ローパスフィルタ
21にてカットする。そして、この入力音声を第3図に
示す如く、4つのブロックに時間的に等分割する。(2) The high-frequency components of the audio signal of the input audio are cut by the low-pass filter 21. Then, this input audio is temporally equally divided into four blocks as shown in FIG.
■音声波形を、第2図に示す如く、複数(n個)チャン
ネルのバンドパスフィルタ22に通し、各ツロック即ち
各一定時間毎に第4図(A)〜(D)のそれぞれに示す
如くの周波数特性を得る。■As shown in Fig. 2, the audio waveform is passed through a band pass filter 22 of multiple (n) channels, and the sound waveform is passed through the band pass filter 22 of multiple (n) channels, and the waveform is filtered at each block, that is, at each fixed time interval, as shown in Figs. 4 (A) to (D). Obtain frequency characteristics.
この時、バンドパスフィルタ22の出力信号は、平均化
回路23にて、各ブロック毎、即ち一定時間て平均化さ
れる。At this time, the output signal of the bandpass filter 22 is averaged by the averaging circuit 23 for each block, that is, for a certain period of time.
以上の前処理により、「音声の一定時間内における平均
的な周波数特性の時間的変化」が得られた。Through the above preprocessing, the "temporal change in the average frequency characteristics of audio within a certain period of time" was obtained.
平均化回路23の出力は、直接的にニューラルネットワ
ーク13に転送され、或いはメモリ部15を経由して間
接的にニューラルネットワーク13に転送される。The output of the averaging circuit 23 is transferred directly to the neural network 13, or indirectly transferred to the neural network 13 via the memory section 15.
■ニューラルネットワーク13は、3層の階層的なニュ
ーラルネットワークにて構成される。入力Ji131は
、前処理の4ブロツク、nチャンネルに対応する4Xn
ユニツトにて構成される。出力層32は、前述した登録
単語数(5単語)に対応する5ユニツトにて構成される
。■The neural network 13 is composed of a three-layer hierarchical neural network. Input Ji131 is 4Xn corresponding to 4 blocks of preprocessing and n channels.
It is composed of units. The output layer 32 is composed of 5 units corresponding to the number of registered words (5 words) described above.
出力132の目標値は、登録単語「ショウメイ」に対応
する出力ユニットについては(lO,O,0,0)、「
エアコン」に対応する出力ユニットについては(0、1
、00、0)、「カーテン」に対応する出力ユニットに
ついては(0、0、1、0、0) rテレビ」に対
応する出力ユニットについては(0,0,0,1゜0)
、「ドア」に対応する出力ユニットについては(0,0
,0,0,1)である。The target value of the output 132 is (lO, O, 0, 0) for the output unit corresponding to the registered word "Shoumei", "
For the output unit corresponding to "air conditioner" (0, 1
, 00, 0), for the output unit corresponding to "Curtain" (0, 0, 1, 0, 0), for the output unit corresponding to "r TV" (0, 0, 0, 1゜0)
, for the output unit corresponding to "door" (0,0
,0,0,1).
実験
(1)学習時の話者である特定話者(1人)による入力
時、認識率は100%であった。Experiment (1) When input was made by a specific speaker (one person) who was the speaker during learning, the recognition rate was 100%.
(2)不特定話者(20人)による入力時、前述の追加
学習なしの場合、認識率は92.0%であった。(2) When input by unspecified speakers (20 people), the recognition rate was 92.0% without the above-mentioned additional learning.
(3)不特定話者(20人)による入力時、前述の追加
学習ありの場合、認識率は99.0%であった。(3) When input by unspecified speakers (20 people), the recognition rate was 99.0% with the above-mentioned additional learning.
次に、上記実施例の作用について説明する。Next, the operation of the above embodiment will be explained.
(1)経時的な認識率の劣化が極めて少ない、このこと
は、後述する実験結果により確認されていることである
が、ニューラルネットワーク13が音声の時期差による
変動の影響を受けにくい構造をとることが可能なためと
推定される。(1) The deterioration of the recognition rate over time is extremely small. This is confirmed by the experimental results described later, and the neural network 13 has a structure that is not easily affected by fluctuations due to differences in the timing of speech. It is presumed that this is because it is possible.
(2)ニューラルネットワーク13を構成する、登録単
語に対応する出力ユニットの出力値に対し、単語認識用
しきい値の他に、追加学習用しきい値を設けた。即ち、
上記出力値が単語認識用しきい値を超えて大なるもので
あり、入力音声単語を登録単語と判定できるものてあっ
ても、該出力値が該単語認識用しきい値より大なる追加
学習用しきい値を超えるものでない場合には、今回の入
力音声データを用いてニューラルネットワーク13の追
加学習を行なう。これにより、ニューラルネットワーク
13の認wt率が劣化する前に、常に該ニューラルネッ
トワーク13を更新し、結果どして、経時的な認識率の
劣化か極めて少ない単語認識システムを構成できる。(2) In addition to the word recognition threshold, an additional learning threshold is provided for the output value of the output unit corresponding to the registered word constituting the neural network 13. That is,
Even if the above output value is larger than the word recognition threshold and the input speech word can be determined as a registered word, additional learning is performed if the output value is larger than the word recognition threshold. If the input audio data does not exceed the threshold value, the neural network 13 performs additional learning using the current input audio data. As a result, the neural network 13 is constantly updated before the recognition wt rate of the neural network 13 deteriorates, and as a result, a word recognition system can be constructed in which the recognition rate deteriorates very little over time.
(3)ニューラルネットワーク13は、原理的に、ネッ
トワーク全体の演算処理が単純且つ迅速である。(3) In principle, the neural network 13 has simple and quick arithmetic processing for the entire network.
(4)ニューラルネットワーク13は、原理的に、それ
を構成している各ユニットか独立に動作しており、並列
的な演算処理が可能である。従って、演算処理か迅速で
ある。(4) In principle, the neural network 13 operates independently of each of its constituent units, and is capable of parallel arithmetic processing. Therefore, calculation processing is quick.
(5)上記(3)〜(4)により、単語認識システム1
0を複雑な処理装置によることなく容易に実時間処理で
きる。(5) According to (3) to (4) above, the word recognition system 1
0 can be easily processed in real time without using a complicated processing device.
(6)ニューラルネットワーク13にて今回の入力音声
単語が認識不能である時、入力話者の助力により今回の
入力単語データを教示されて該ニューラルネットワーク
13の追加学習を行ない、認識不能の再発を防止できる
。この際、追加学習は、ニューラルネットワーク13の
各ユニットの変換関数及び重みを修正することによりな
されるものであるため、メモリ容量の増加や認識処理時
間の増大化を必要とすることかない。(6) When the current input speech word is unrecognizable in the neural network 13, the current input word data is taught with the help of the input speaker, and the neural network 13 performs additional learning to prevent the recurrence of unrecognizability. It can be prevented. At this time, since the additional learning is performed by modifying the conversion function and weight of each unit of the neural network 13, there is no need to increase memory capacity or recognition processing time.
(7)ニューラルネットワーク13への入力として、「
音声の周波数特性の時間的変化」を用いたから、入力を
得るための前処理が従来の複雑な特徴量抽出に比して、
単純となりこの前処理に要する時間が短くて足りる。(7) As an input to the neural network 13, “
Because we use "temporal changes in the frequency characteristics of audio," the preprocessing to obtain input is more complex than conventional feature extraction.
It is simple and the time required for this preprocessing is short.
この時、上記ニューラルネットワークへの入力として、
更に、「音声の一定時間内における平均的な周波数特性
の時間的変化」を用いたから、ニューラルネットワーク
13における処理が単純となり、この処理に要する時間
がより短くて足りる。At this time, as an input to the above neural network,
Furthermore, since the "temporal change in the average frequency characteristic within a certain period of time" is used, the processing in the neural network 13 is simple, and the time required for this processing is shorter.
(8)階層的なニューラルネットワーク13を用いたか
ら、現在、既に確立している簡単な学習アルゴリズム(
パックプロパゲーション)を用いて、高い認識率を達成
できる。(8) Since the hierarchical neural network 13 is used, a simple learning algorithm that has already been established (
Pack propagation) can be used to achieve high recognition rates.
尚、本発明の実施においては、ニューラルネットワーク
への入力として、
■音声の周波数特性の時間的変化、
■音声の平均的な線形予測係数、
■音声の平均的なPARCOR係数、
■音声の平均的な周波数特性、及びピッチ周波数、
■高域強調を施された音声波形の平均的な周波数特性、
並びに
■音声の平均的な周波数特性
のうちの1つ以上を使用できる。In the implementation of the present invention, as inputs to the neural network, ■temporal changes in the frequency characteristics of audio, ■average linear prediction coefficients of audio, ■average PARCOR coefficients of audio, ■average audio frequency characteristics and pitch frequency, ■Average frequency characteristics of high-frequency emphasized audio waveforms,
and ■ one or more of the average frequency characteristics of voice can be used.
そして、上記■の要素が更に「音声の一定時間内におけ
る平均的な周波数特性の時間的変化」として用いられた
ように、上記■の要素は「音声の一定時間内における平
均的な線形予測係数の時間的変化」、上記■の要素は「
音声の一定時間内における平均的なPARCOR係数の
時間的変化」、上記■の要素は「音声の一定時間内にお
ける平均的な周波数特性、及びピッチ周波数の時間的変
化」、上記■の要素は、「高域強調を施された音声波形
の一定時間内における平均的な周波数特性の時間的変化
」として用いることができる。Then, just as the element (■) above was further used as "temporal change in the average frequency characteristics within a certain period of time", the element (■) above is also used as "the average linear prediction coefficient within a certain period of time". "Temporal change in
``Temporal change in the average PARCOR coefficient within a certain time period of audio'', the above element (■) is ``the average frequency characteristic and temporal change in pitch frequency within a certain time period of audio'', and the above element (■) is: It can be used as a "temporal change in the average frequency characteristics within a certain period of time of a voice waveform that has been subjected to high-frequency emphasis."
尚、上記■の線形予測係数は、以下の如く定義される。Incidentally, the linear prediction coefficient of (2) above is defined as follows.
即ち、音声波形のサンプル値(χ。)の間には、−mに
高い近接相関があることが知られている。That is, it is known that -m has a high proximity correlation between sample values (χ) of audio waveforms.
そこで次のような線形予測か可能であると仮定する。Therefore, assume that the following linear prediction is possible.
線形予測値 χt=−Σα1χt−1 ・・・(1
)線形予測誤差 ε(=χ先−χ(・・・(2)ここて
、χに時刻tにおける音声波形のサンプル値、(αム)
(1=1.・・・、p): (9次の)線形予測係数
さて、本発明の実施においては、線形予測誤差ε、の2
乗平均値が最小となるように線形予測係数(a、)を求
める。Linear predicted value χt=-Σα1χt-1 ...(1
)Linear prediction error ε(=χ ahead−χ(...(2) Here, χ is the sample value of the audio waveform at time t, (αm)
(1=1...., p): (9th order) linear prediction coefficient Now, in the implementation of the present invention, the linear prediction error ε, 2
The linear prediction coefficient (a,) is determined so that the root mean value is minimized.
具体的には (ε )2を求め、その時間平均を(ε、
)2と表わして、θ(εt)” /aa (=O,i=
]、2.・・・、pとおくことによって、次の式から(
al)が求められる。Specifically, find (ε)2 and calculate the time average as (ε,
)2, θ(εt)”/aa (=O, i=
], 2. ..., p, from the following equation (
al) is required.
ΣQ 1vll−Jl = L j= 1 +
2 + ・・・、p ・・・(3)又、上記■の
PARCOR係数は以下の如く定義される。ΣQ 1vll−Jl = L j= 1 +
2 + . . . p .
即ち、[k、](n=1.・・・、p)を(9次の)P
AR(:OR係数(偏自己相関係数)とする時、PAR
COR係数k。、lは、線形予測による前向き残差εt
l)ど後向き残差εL−+n+11 (b’間の正規化
相関係数として、次の式によって定義される。That is, [k,] (n=1..., p) is (9th order) P
When AR (:OR coefficient (partial autocorrelation coefficient)), PAR
COR coefficient k. , l is the forward residual εt due to linear prediction
l) Backward residual εL-+n+11 (defined as the normalized correlation coefficient between b' by the following equation).
ε%f1.ε、−(。。、、 (bl
・・・ (4)
ここで、8t(f)=χえ−′t αlχt−1(α、
):前向き予測係数、
ε1−(川、ll=χ(−(川じりj・χt−J・(β
ノ):後向き予測係数
又、上記■の音声のピッチ周波数とは、声帯波の繰り返
し周期(ピッチ周期)の逆数である。ε%f1. ε, −(..,, (bl ... (4) Here, 8t(f)=χe−′t αlχt−1(α,
): forward prediction coefficient, ε1−(river, ll=χ(−(river edge j・χt−J・(β
C): Backward prediction coefficient Also, the pitch frequency of the voice mentioned in (2) above is the reciprocal of the repetition period (pitch period) of the vocal cord wave.
尚、ニューラルネットワークへの入力として、個人差が
ある声帯の基本的なパラメータであるピッチ周波数を付
加したから、特に大人/小人、男性/女性間の話者の認
識率を向上することができる。Furthermore, since the pitch frequency, which is a basic parameter of the vocal cords that differs among individuals, was added as an input to the neural network, it is possible to improve the recognition rate of speakers, especially between adults/dwarfs and male/female. .
又、上記■の高域強調とは、音声波形のスペクトルの平
均的な傾きを補償して、低域にエネルギが集中すること
を防止することである。然るに、音声波形のスペクトル
の平均的な傾きは話者に共通のものであり、話者の認識
には無関係である。Furthermore, the above-mentioned high frequency enhancement (2) is to compensate for the average slope of the spectrum of the audio waveform to prevent concentration of energy in the low frequency range. However, the average slope of the spectrum of the speech waveform is common to all speakers and is unrelated to the speaker's recognition.
ところが、このスペクトルの平均的な傾きが補償されて
いない音声波形をそのままニューラルネットワークへ入
力する場合には、ニューラルネットワークか学習する時
にスペクトルの平均的な傾きの特徴の方を抽出してしま
い、話者の認識に必要なスペクトルの山と谷を抽出する
のに時間がかかる。これに対し、ニューラルネットワー
クへの入力を高域強調する場合には、話者に共通で、認
識には無関係でありながら、学習に影響を及ぼすスペク
トルの平均的な傾きを補償できるため、学習速度が速く
なるのである。However, when inputting an audio waveform that has not been compensated for the average slope of the spectrum to a neural network as is, the neural network will extract the feature of the average slope of the spectrum during learning, and the speech will be distorted. It takes time to extract the peaks and valleys of the spectrum necessary for human recognition. On the other hand, when high-frequency emphasis is applied to the input to a neural network, it is possible to compensate for the average slope of the spectrum that is common to all speakers and is unrelated to recognition, but that affects learning, which speeds up the learning process. becomes faster.
[発明の効果]
以上のように本発明によれば、経時的な認識率の劣化が
極めて少なく、かつ容易に実時間処理できる単語認識シ
ステムを得ることができる。[Effects of the Invention] As described above, according to the present invention, it is possible to obtain a word recognition system in which the deterioration of the recognition rate over time is extremely small and which can easily perform real-time processing.
又、本発明によれば、認識不能に係る登録単語を追加学
習して認識不能の再発を防止するに際し、メモリ容量の
増加と認識処理時間の増大化を必要とすることなく、か
つ容易に実時間処理できる単語認識システムを得ること
ができる。Further, according to the present invention, when additionally learning registered words related to unrecognizability to prevent the recurrence of unrecognizability, it is possible to easily implement it without requiring an increase in memory capacity or an increase in recognition processing time. A word recognition system capable of time processing can be obtained.
第1図は本発明が適用された単語認識システムの一例を
示す模式図、第2図は音声処理部とニューラルネットワ
ークの一例を示す模式図、第3図は入力音声を示す模式
図、第4図はバンドパスフィルタの出力を示す模式図、
第5図はニューラルネットワークを示す模式図、第6図
は階層的なニューラルネットワークを示す模式図、第7
図はユニットの構造を示す模式図である。
10・・・話者認識システム、
11・・・音声入力部、
12・・・音声処理部、
13・・・ニューラルネットワーク、
14・・・判定部、
15・・・メモリ部、
16・・・ネットワーク制御部、
17・・・機器制御部、
18・・・表示制御部、
19・・・表示部、
2o・・・教示部。FIG. 1 is a schematic diagram showing an example of a word recognition system to which the present invention is applied, FIG. 2 is a schematic diagram showing an example of a speech processing unit and a neural network, FIG. 3 is a schematic diagram showing input speech, and FIG. The figure is a schematic diagram showing the output of a bandpass filter.
Figure 5 is a schematic diagram showing a neural network, Figure 6 is a schematic diagram showing a hierarchical neural network, and Figure 7 is a schematic diagram showing a hierarchical neural network.
The figure is a schematic diagram showing the structure of the unit. DESCRIPTION OF SYMBOLS 10...Speaker recognition system, 11...Speech input unit, 12...Speech processing unit, 13...Neural network, 14...Determination unit, 15...Memory unit, 16... Network control unit, 17... Equipment control unit, 18... Display control unit, 19... Display unit, 2o... Teaching unit.
Claims (5)
ムであって、登録単語に対応する出力ユニットの出力値
に対し、単語認識用しきい値と追加学習用しきい値とを
設定し、上記出力値が単語認識用しきい値より大なるこ
とを条件に、今回の入力音声単語を登録単語と判定し、
上記出力値が単語認識用しきい値より大、かつ追加学習
用しきい値より小なることを条件に、今回の入力音声デ
ータを用いてニューラルネットワークの追加学習を行な
う単語認識システム。(1) A word recognition system using a neural network, in which a word recognition threshold and an additional learning threshold are set for the output value of an output unit corresponding to a registered word, and the output value is The current input voice word is determined to be a registered word on the condition that it is greater than the word recognition threshold,
A word recognition system that performs additional learning of a neural network using current input voice data on the condition that the output value is greater than a word recognition threshold and smaller than an additional learning threshold.
ムであって、登録単語に対応する出力ユニットの出力値
に対し、単語認識用しきい値を設定し、上記出力値が単
語認識用しきい値より大なることを条件に、今回の入力
音声単語を登録単語と判定し、上記出力値が単語認識用
しきい値より小なることを条件に、今回の入力音声単語
を認識できなかったことを表示し、この表示に対して今
回の話者により教示される今回の入力単語データと今回
の入力音声データとを用いてニューラルネットワークの
追加学習を行なう単語認識システム。(2) A word recognition system using a neural network, in which a word recognition threshold is set for the output value of an output unit corresponding to a registered word, and the output value is greater than the word recognition threshold. On the condition that , a word recognition system that performs additional learning of a neural network using the current input word data taught by the current speaker and the current input voice data for this display.
ムであって、登録単語に対応する出力ユニットの出力値
に対し、単語認識用しきい値と追加学習用しきい値とを
設定し、上記出力値が単語認識用しきい値より大なるこ
とを条件に、今回の入力音声単語を登録単語と判定し、
(A)上記出力値が単語認識用しきい値より小なること
を条件に、今回の入力音声単語を認識できなかったこと
を表示し、この表示に対して今回の話者により教示され
る今回の入力単語データと今回の入力音声データとを用
いてニューラルネットワークの追加学習を行ない、(B
)上記出力値が単語認識用しきい値より大、かつ追加学
習用しきい値より小なることを条件に、今回の入力音声
データを用いてニューラルネットワークの追加学習を行
なう単語認識システム。(3) A word recognition system using a neural network, in which a word recognition threshold and an additional learning threshold are set for the output value of the output unit corresponding to the registered word, and the output value is The current input voice word is determined to be a registered word on the condition that it is greater than the word recognition threshold,
(A) On the condition that the above output value is smaller than the threshold for word recognition, it is displayed that the current input speech word could not be recognized, and in response to this display, the current input voice word taught by the current speaker is displayed. Additional learning of the neural network is performed using the input word data of
) A word recognition system that performs additional learning of a neural network using the current input voice data, provided that the output value is greater than a word recognition threshold and smaller than an additional learning threshold.
性、並びに [6]音声の平均的な周波数特性 のうちの1つ以上を使用する請求項1〜3のいずれかに
記載の単語認識システム。(4) As inputs to the neural network, [1] Temporal changes in the frequency characteristics of speech, [2] Average linear prediction coefficients of speech, [3] Average PARCOR coefficients of speech, [4] speech using one or more of the following: average frequency characteristics and pitch frequency; [5] average frequency characteristics of high-frequency emphasized audio waveform; and [6] average frequency characteristics of audio. The word recognition system according to any one of claims 1 to 3.
ルネットワークである請求項1〜4のいずれかに記載の
単語認識システム。(5) The word recognition system according to any one of claims 1 to 4, wherein the neural network is a hierarchical neural network.
Priority Applications (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP1298502A JP2543603B2 (en) | 1989-11-16 | 1989-11-16 | Word recognition system |
Applications Claiming Priority (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP1298502A JP2543603B2 (en) | 1989-11-16 | 1989-11-16 | Word recognition system |
Publications (2)
| Publication Number | Publication Date |
|---|---|
| JPH03157697A true JPH03157697A (en) | 1991-07-05 |
| JP2543603B2 JP2543603B2 (en) | 1996-10-16 |
Family
ID=17860545
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| JP1298502A Expired - Lifetime JP2543603B2 (en) | 1989-11-16 | 1989-11-16 | Word recognition system |
Country Status (1)
| Country | Link |
|---|---|
| JP (1) | JP2543603B2 (en) |
Cited By (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JPH08259124A (en) * | 1995-03-28 | 1996-10-08 | Fujitec Co Ltd | Voice recognizing device for elevator |
| JP2002502992A (en) * | 1998-02-03 | 2002-01-29 | デテモビール ドイチェ テレコム モビールネット ゲーエムベーハー | Method and apparatus for improving recognition rate in speech recognition system |
| WO2019235191A1 (en) * | 2018-06-05 | 2019-12-12 | 日本電信電話株式会社 | Model learning device, method and program |
Families Citing this family (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP3322491B2 (en) | 1994-11-25 | 2002-09-09 | 三洋電機株式会社 | Voice recognition device |
| JP3322536B2 (en) | 1995-09-13 | 2002-09-09 | 三洋電機株式会社 | Neural network learning method and speech recognition device |
Citations (8)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JPS5837700A (en) * | 1981-08-31 | 1983-03-04 | トヨタ自動車株式会社 | Voice recognition unit |
| JPS5852696A (en) * | 1981-09-25 | 1983-03-28 | 大日本印刷株式会社 | Voice recognition unit |
| JPS605960A (en) * | 1983-06-25 | 1985-01-12 | 産業振興株式会社 | Wall remodeling method of existing building |
| JPS6014300A (en) * | 1983-07-06 | 1985-01-24 | シャープ株式会社 | Feature extraction system for voice |
| JPS6047600A (en) * | 1983-08-09 | 1985-03-14 | ロバート マイケル グランバーグ | Audio image forming device |
| JPS61114299A (en) * | 1984-11-09 | 1986-05-31 | 日本電気株式会社 | Voice recognition system |
| JPS62149000A (en) * | 1985-12-23 | 1987-07-02 | 日本電気株式会社 | Voice analyzer |
| JPS63261400A (en) * | 1987-04-20 | 1988-10-28 | 富士通株式会社 | Voice recognition system |
-
1989
- 1989-11-16 JP JP1298502A patent/JP2543603B2/en not_active Expired - Lifetime
Patent Citations (8)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JPS5837700A (en) * | 1981-08-31 | 1983-03-04 | トヨタ自動車株式会社 | Voice recognition unit |
| JPS5852696A (en) * | 1981-09-25 | 1983-03-28 | 大日本印刷株式会社 | Voice recognition unit |
| JPS605960A (en) * | 1983-06-25 | 1985-01-12 | 産業振興株式会社 | Wall remodeling method of existing building |
| JPS6014300A (en) * | 1983-07-06 | 1985-01-24 | シャープ株式会社 | Feature extraction system for voice |
| JPS6047600A (en) * | 1983-08-09 | 1985-03-14 | ロバート マイケル グランバーグ | Audio image forming device |
| JPS61114299A (en) * | 1984-11-09 | 1986-05-31 | 日本電気株式会社 | Voice recognition system |
| JPS62149000A (en) * | 1985-12-23 | 1987-07-02 | 日本電気株式会社 | Voice analyzer |
| JPS63261400A (en) * | 1987-04-20 | 1988-10-28 | 富士通株式会社 | Voice recognition system |
Cited By (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JPH08259124A (en) * | 1995-03-28 | 1996-10-08 | Fujitec Co Ltd | Voice recognizing device for elevator |
| JP2002502992A (en) * | 1998-02-03 | 2002-01-29 | デテモビール ドイチェ テレコム モビールネット ゲーエムベーハー | Method and apparatus for improving recognition rate in speech recognition system |
| WO2019235191A1 (en) * | 2018-06-05 | 2019-12-12 | 日本電信電話株式会社 | Model learning device, method and program |
| JP2019211627A (en) * | 2018-06-05 | 2019-12-12 | 日本電信電話株式会社 | Model learning device, method and program |
| US12057107B2 (en) | 2018-06-05 | 2024-08-06 | Nippon Telegraph And Telephone Corporation | Model learning apparatus, method and program |
Also Published As
| Publication number | Publication date |
|---|---|
| JP2543603B2 (en) | 1996-10-16 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| Barnard et al. | Pitch detection with a neural-net classifier | |
| CN109243494B (en) | Children emotion recognition method based on multi-attention mechanism long-time memory network | |
| CN103325381B (en) | A kind of speech separating method based on fuzzy membership functions | |
| CN102800316B (en) | Optimal codebook design method for voiceprint recognition system based on nerve network | |
| CN110379441A (en) | A kind of voice service method and system based on countering type smart network | |
| CN101527141A (en) | Method of converting whispered voice into normal voice based on radial group neutral network | |
| Zhu et al. | Contribution of modulation spectral features on the perception of vocal-emotion using noise-vocoded speech | |
| Wei et al. | Iifc-net: A monaural speech enhancement network with high-order information interaction and feature calibration | |
| CN108447470A (en) | An Emotional Speech Conversion Method Based on Vocal Tract and Prosodic Features | |
| Gandhiraj et al. | Auditory-based wavelet packet filterbank for speech recognition using neural network | |
| JP2543603B2 (en) | Word recognition system | |
| JPH03157698A (en) | Speaker recognizing system | |
| EP0369485B1 (en) | Speaker recognition system | |
| JPH05257496A (en) | Word recognizing system | |
| JPH03175498A (en) | Speaker collating system | |
| JPH02275996A (en) | Word recognition system | |
| Lashkari et al. | NMF-based cepstral features for speech emotion recognition | |
| Li et al. | Cross-modal mask fusion and modality-balanced audio-visual speech recognition | |
| CN115862636B (en) | A method of Internet man-machine verification based on speech recognition technology | |
| JPH02273798A (en) | Speaker recognition system | |
| Drgas | Speech intelligibility prediction based on syllable tokenizer | |
| Zhao et al. | Speech enhancement based on dual-path cross-parallel conformer network | |
| JP2518939B2 (en) | Speaker verification system | |
| JPH02304498A (en) | Word recognition system | |
| JPH0415700A (en) | Speaker recognition system |