JPH08146996A - Speech recognition device - Google Patents

Speech recognition device

Info

Publication number
JPH08146996A
JPH08146996A JP6291725A JP29172594A JPH08146996A JP H08146996 A JPH08146996 A JP H08146996A JP 6291725 A JP6291725 A JP 6291725A JP 29172594 A JP29172594 A JP 29172594A JP H08146996 A JPH08146996 A JP H08146996A
Authority
JP
Japan
Prior art keywords
voice
pattern
input
neural network
section
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Granted
Application number
JP6291725A
Other languages
Japanese (ja)
Other versions
JP3322491B2 (en
Inventor
Hiroya Murao
浩也 村尾
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Sanyo Electric Co Ltd
Original Assignee
Sanyo Electric Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Sanyo Electric Co Ltd filed Critical Sanyo Electric Co Ltd
Priority to JP29172594A priority Critical patent/JP3322491B2/en
Publication of JPH08146996A publication Critical patent/JPH08146996A/en
Application granted granted Critical
Publication of JP3322491B2 publication Critical patent/JP3322491B2/en
Anticipated expiration legal-status Critical
Expired - Fee Related legal-status Critical Current

Links

Abstract

PURPOSE: To improve the recognition precision by using, as an input pattern, what causes a misrecognition when it is inputted to a neural network after initial learning to be subjected to speech recognition, and making the neural network to perform additional learning by using unsupervised data. CONSTITUTION: A neural network arithmetic part 4 generates standard speech patterns for initial learning based upon a suitable speech section and standard speech patterns for initial learning based upon speech sections different from the suitable speech section by speeches to be recognized. The neural network initially learns the standard speech patterns for initial learning as input patterns and speech discrimination data representing the speeches corresponding to the respective input patterns as tutor data. When a standard speech pattern for additional learning is inputted to the neural network after the initial learning and speech recognition is performed, a misrecognized pattern is regarded as input data and the neural network performs the additional learning by using the unsupervised data.

Description

【発明の詳細な説明】Detailed Description of the Invention

【0001】[0001]

【産業上の利用分野】この発明は、音声によりデータを
入力するための音声認識装置に関し、たとえば、録画番
組の予約が音声入力によって行われる録画装置等に利用
される音声認識装置に関する。
BACKGROUND OF THE INVENTION 1. Field of the Invention The present invention relates to a voice recognition device for inputting data by voice, for example, a voice recognition device used for a recording device or the like in which reservation of a recorded program is made by voice input.

【0002】[0002]

【従来の技術】図6は、従来の音声認識装置の構成を示
している。
2. Description of the Related Art FIG. 6 shows the configuration of a conventional voice recognition device.

【0003】音声分析部101は、入力音声の音声パワ
ー信号と、入力音声に対する音声スペクトルとを生成す
る。入力音声の音声パワー信号は、音声区間検出部10
2に送られる。入力音声に対する音声スペクトルは、音
声パターン作成部103に送られる。
The voice analysis unit 101 generates a voice power signal of an input voice and a voice spectrum for the input voice. The voice power signal of the input voice is the voice section detection unit 10
Sent to 2. The voice spectrum for the input voice is sent to the voice pattern creating unit 103.

【0004】音声区間検出部102は、音声検出部11
1および音声区間切出し部112とを備えている。音声
検出部111は、図7に示すように、音声検出用しきい
値αを用いて、音声パワー信号中の音声部分を検出す
る。
The voice section detection unit 102 includes a voice detection unit 11
1 and a voice section cutout unit 112. As shown in FIG. 7, the voice detection unit 111 detects the voice portion in the voice power signal using the voice detection threshold value α.

【0005】音声区間切出し部112は、図7に示すよ
うに、切出し用しきい値βを用いて、音声認識に有効な
音声区間Lを求める。切出し用しきい値βは、音声検出
部111によって検出された音声部分より所定時間前の
雑音パワーに基づいて決定される。
As shown in FIG. 7, the voice section cutout unit 112 obtains a voice section L effective for voice recognition using a cutout threshold value β. The cut-out threshold value β is determined based on the noise power of a predetermined time before the voice portion detected by the voice detection unit 111.

【0006】音声パターン作成部103は、音声区間切
出し部112によって求められた音声区間Lに対する音
声スペクトルに基づいて、音声パターンを作成する。作
成された音声パターンは、学習済のニューラルネットワ
ーク104に入力される。
The voice pattern creating section 103 creates a voice pattern based on the voice spectrum for the voice section L obtained by the voice section cutting section 112. The created voice pattern is input to the learned neural network 104.

【0007】このニューラルネットワーク104の学習
は、次のように行なわれる。まず、各認識対象音声に対
する標準音声パターンを、予め収集した音声を用いてそ
れぞれ求める。そして、各標準音声パターンを入力パタ
ーンとし、各入力パターンに対応する音声を表す音声識
別データを教師データとして、ニューラルネットワーク
104を学習させる。
The learning of the neural network 104 is performed as follows. First, a standard voice pattern for each recognition target voice is obtained by using voices collected in advance. Then, the neural network 104 is trained by using each standard voice pattern as an input pattern and voice identification data representing a voice corresponding to each input pattern as teacher data.

【0008】学習済のニューラルネットワーク104
に、音声パターンが入力されることにより、入力された
音声パターンに対応する出力パターンが得られる。この
出力パターンは、認識結果判定部105に送られる。認
識結果判定部105は、送られてきた出力パターンに基
づいて当該音声検出部分の音声を認識し、その認識結果
を出力する。
The learned neural network 104
By inputting the voice pattern to the input device, an output pattern corresponding to the input voice pattern is obtained. This output pattern is sent to the recognition result determination unit 105. The recognition result determination unit 105 recognizes the voice of the voice detection portion based on the sent output pattern and outputs the recognition result.

【0009】[0009]

【発明が解決しようとする課題】このような音声認識装
置では、音声認識に有効な音声区間を設定するための切
出し用しきい値βは1つであるため、雑音が音声区間に
含まれることによって誤認識が発生したり、音声パワー
の小さい語尾等が音声区間から脱落することによって誤
認識が発生したりする可能性が高い。
In such a voice recognition device, since there is one cut-out threshold β for setting a voice section effective for voice recognition, noise is included in the voice section. There is a high possibility that erroneous recognition may occur, or a false ending may occur due to a word ending or the like having low voice power dropping from the voice section.

【0010】そこで、本出願人は、次のような音声認識
方法を開発した。つまり、図5に示すように、複数のし
きい値β1、β2、β3およびβ4を用いて、複数の音
声区間L1、L2、L3およびL4を設定する。各音声
区間L1〜L4それぞれに対して、音声パターンを作成
する。ニューラルネットワークに各音声パターンを入力
して、各音声パターンごとに出力パターンを得る。そし
て、得られたこれらの複数の出力パターンに基づいて、
音声を認識する。
Therefore, the present applicant has developed the following voice recognition method. That is, as shown in FIG. 5, a plurality of voice sections L1, L2, L3, and L4 are set using a plurality of threshold values β1, β2, β3, and β4. A voice pattern is created for each of the voice sections L1 to L4. Each voice pattern is input to the neural network, and an output pattern is obtained for each voice pattern. And, based on these multiple output patterns obtained,
Recognize voice.

【0011】各認識対象音声を表す音声識別データは、
ニューラルネットワークの出力層の各ユニットに対応し
た数のデータから構成されているものとする。そして、
その1つのみが”1”で他が全て”0”のデータで構成
され、データ”1”の位置が各音声識別データごとに異
なっているものとする。
The voice identification data representing each recognition target voice is
It is assumed that it is composed of a number of data corresponding to each unit in the output layer of the neural network. And
It is assumed that only one of them is composed of data of "1" and the other is composed of "0", and the position of the data "1" is different for each voice identification data.

【0012】このような音声認識方法では、図5の各音
声区間L1〜L2の認識結果は、たとえば、次のように
なることがある。すなわち、音声区間L1での認識結果
は”しち”で、出力最大値(ニューラルネットワークの
出力層のユニットの出力のうちの最大値)が0.90で
ある。音声区間L2での認識結果は”に”で、出力最大
値が0.85である。音声区間L3での認識結果は”
に”で、出力最大値が0.91である。音声区間L4で
の認識結果は”に”で、出力最大値が0.88である。
In such a voice recognition method, the recognition result of each voice section L1 to L2 in FIG. 5 may be as follows, for example. That is, the recognition result in the voice section L1 is "shichi", and the maximum output value (the maximum value of the outputs of the units in the output layer of the neural network) is 0.90. The recognition result in the voice section L2 is "ni", and the maximum output value is 0.85. The recognition result in the voice section L3 is "
The output maximum value is 0.91 in the case of "ni". The recognition result in the voice section L4 is "ni" and the maximum output value is 0.88.

【0013】このような場合には、最終認識結果として
は、出力最大値が”1”に最も近い音声区間L3での認
識結果”に”が、入力音声の認識結果として選択され、
本来”しち”と認識されるべきところが、”に”と誤認
識されてしまう。
In such a case, as the final recognition result, the recognition result "in" in the voice section L3 whose output maximum value is closest to "1" is selected as the recognition result of the input voice,
Where it should have been recognized as "shichi", it is incorrectly recognized as "ni".

【0014】この発明は、認識精度の向上が図れる音声
認識装置を提供することを目的とする。
It is an object of the present invention to provide a voice recognition device capable of improving recognition accuracy.

【0015】[0015]

【課題を解決するための手段】この発明による第1の音
声認識装置は、入力音声に対して音声区間を設定する音
声区間設定手段、音声区間の特徴に基づいて、音声区間
の音声パターンを作成する音声パターン作成手段、およ
び音声パターンが入力されるニューラルネットワークを
有しかつニューラルネットワークの出力に基づいて入力
音声を認識する音声認識手段を備えており、各認識対象
音声ごとに、好適な音声区間に基づく初期学習用標準音
声パターンと、好適な音声区間とは異なる音声区間に基
づく追加学習用標準音声パターンとが作成され、初期学
習用標準音声パターンを入力パターンとし、各入力パタ
ーンに対応する音声を表す音声識別データを教師データ
として、ニューラルネットワークが初期学習され、追加
学習用標準音声パターンのうち、初期学習済のニューラ
ルネットワークにそれが入力されて音声認識が行なわれ
たときに、誤認識が生じたものを入力パターンとし、反
教師データを用いてニューラルネットワークが追加学習
されていることを特徴とする。上記音声区間の特徴とし
ては、たとえば、音声スペクトルが挙げられる。
A first voice recognition apparatus according to the present invention creates a voice pattern of a voice section based on a feature of the voice section and a voice section setting means for setting a voice section for an input voice. And a voice recognition means for recognizing the input voice based on the output of the neural network, and a suitable voice section for each recognition target voice. And a standard speech pattern for additional learning based on a speech section different from a suitable speech section are created, and the standard speech pattern for initial learning is used as an input pattern, and a speech corresponding to each input pattern is generated. The neural network is initially trained using the voice identification data that represents When the speech recognition is performed by inputting it to the neural network that has already been trained, the input pattern is the one in which misrecognition occurs, and the neural network is additionally learned using the anti-teaching data. It is characterized by being The features of the voice section include, for example, a voice spectrum.

【0016】反教師データは、各音声識別データがニュ
ーラルネットワークの出力層の各ユニットに対応した数
のデータから構成されており、その1つのみが”1”で
他が全て”0”のデータで構成され、データ”1”の位
置が各音声識別データごとに異なっている場合には、全
て”0”のデータから構成される。
The anti-teacher data is composed of a number of data in which each voice identification data corresponds to each unit of the output layer of the neural network, and only one of them is "1" and the others are "0". When the position of the data "1" is different for each voice identification data, it is composed of all the data "0".

【0017】各音声識別データがニューラルネットワー
クの出力層の各ユニットに対応した数のデータから構成
されており、その1つのみが”0”で他が全て”1”の
データで構成され、データ”0”の位置が各音声識別デ
ータごとに異なっている場合には、反教師データは、全
て”1”のデータから構成される。
Each voice identification data is composed of a number of data corresponding to each unit of the output layer of the neural network, only one of which is "0" and the other is all "1". When the position of "0" is different for each voice identification data, the counter teacher data is composed of all "1" data.

【0018】この発明による第2の音声認識装置は、入
力音声に対して複数の音声区間を設定する音声区間設定
手段、各音声区間の特徴に基づいて、各音声区間ごとの
音声パターンをそれぞれ作成する音声パターン作成手
段、および各音声区間ごとの音声パターンがそれぞれ入
力されるニューラルネットワークを有しかつ各音声区間
ごとの音声パターンに対するニューラルネットワークの
出力に基づいて入力音声を認識する音声認識手段を備え
ており、各認識対象音声ごとに、好適な音声区間に基づ
く初期学習用標準音声パターンと、好適な音声区間とは
異なる音声区間に基づく追加学習用標準音声パターンと
が作成され、初期学習用標準音声パターンを入力パター
ンとし、各入力パターンに対応する音声を表す音声識別
データを教師データとして、ニューラルネットワークが
初期学習され、追加学習用標準音声パターンのうち、初
期学習済のニューラルネットワークにそれが入力されて
音声認識が行なわれたときに、誤認識が生じたものを入
力パターンとし、反教師データを用いてニューラルネッ
トワークが追加学習されていることを特徴とする。上記
音声区間の特徴としては、たとえば、音声スペクトルが
挙げられる。
The second voice recognition apparatus according to the present invention creates a voice pattern for each voice section on the basis of the voice section setting means for setting a plurality of voice sections for the input voice and the characteristics of each voice section. And a voice recognition means for recognizing the input voice based on the output of the neural network for the voice pattern for each voice section. For each recognition target speech, a standard speech pattern for initial learning based on a suitable speech section and a standard speech pattern for additional learning based on a speech section different from the suitable speech section are created, and a standard speech pattern for initial learning is created. The voice pattern is used as an input pattern, and the voice identification data representing the voice corresponding to each input pattern is used as teacher data. Then, the neural network is initially learned, and among the standard speech patterns for additional learning, when it is input to the already learned neural network and speech recognition is performed, the one that causes misrecognition is taken as the input pattern. , The neural network is additionally learned by using the anti-teacher data. The features of the voice section include, for example, a voice spectrum.

【0019】[0019]

【作用】この発明による第1の音声認識装置では、入力
音声に対して、音声区間が設定される。音声区間の特徴
に基づいて、音声区間の音声パターンが作成される。音
声パターンがニューラルネットワークに入力される。そ
して、ニューラルネットワークの出力に基づいて入力音
声が認識される。
In the first voice recognition apparatus according to the present invention, the voice section is set for the input voice. A voice pattern of the voice section is created based on the characteristics of the voice section. The voice pattern is input to the neural network. Then, the input voice is recognized based on the output of the neural network.

【0020】この発明による第2の音声認識装置では、
入力音声に対して、複数の音声区間が設定される。各音
声区間の特徴に基づいて、各音声区間ごとの音声パター
ンがそれぞれ作成される。各音声区間ごとの音声パター
ンがニューラルネットワークにそれぞれ入力される。各
音声区間ごとの音声パターンに対するニューラルネット
ワークの出力に基づいて入力音声が認識される。
In the second voice recognition device according to the present invention,
A plurality of voice sections are set for the input voice. A voice pattern for each voice section is created based on the characteristics of each voice section. The voice pattern for each voice section is input to the neural network. The input voice is recognized based on the output of the neural network for the voice pattern for each voice section.

【0021】この発明による第1または第2の音声認識
装置のニューラルネットワークの学習は、次のように行
なわれている。
Learning of the neural network of the first or second speech recognition apparatus according to the present invention is performed as follows.

【0022】つまり、各認識対象音声ごとに、好適な音
声区間に基づく初期学習用標準音声パターンと、好適な
音声区間とは異なる音声区間に基づく追加学習用標準音
声パターンとが作成され、初期学習用標準音声パターン
を入力パターンとし、各入力パターンに対応する音声を
表す音声識別データを教師データとして、ニューラルネ
ットワークが初期学習される。
That is, a standard speech pattern for initial learning based on a suitable speech section and a standard speech pattern for additional learning based on a speech section different from the preferred speech section are created for each speech to be recognized, and initial learning is performed. The neural network is initially learned using the standard voice pattern for input as an input pattern and the voice identification data representing the voice corresponding to each input pattern as the teacher data.

【0023】また、追加学習用標準音声パターンのう
ち、初期学習済のニューラルネットワークにそれが入力
されて音声認識が行なわれたときに、誤認識が生じたも
のを入力パターンとし、反教師データを用いてニューラ
ルネットワークが追加学習される。
Further, among the standard voice patterns for additional learning, the one that is erroneously recognized when the voice is recognized by inputting it to the neural network which has been initially learned is used as the input pattern, and the anti-teaching data is set. The neural network is additionally trained using this.

【0024】[0024]

【実施例】以下、図1〜図5を参照して、この発明の実
施例について説明する。
Embodiments of the present invention will be described below with reference to FIGS.

【0025】図1は、音声認識装置の構成を示してい
る。
FIG. 1 shows the configuration of a voice recognition device.

【0026】音声認識装置は、音声分析部1、音声区間
検出部2、音声パターン作成部3、ニューラルネットワ
ーク演算部4、認識結果記憶部5および認識結果判定部
6を備えている。音声区間検出部2は、音声検出部2
1、音声区間切出し部22および切出し位置記憶部23
を備えている。
The voice recognition device comprises a voice analysis unit 1, a voice section detection unit 2, a voice pattern creation unit 3, a neural network operation unit 4, a recognition result storage unit 5 and a recognition result determination unit 6. The voice section detection unit 2 is a voice detection unit 2
1. Voice segment cutout unit 22 and cutout position storage unit 23
It has.

【0027】図2は、ニューラルネットワーク演算部4
に設けられているニューラルネットワークの構造の一例
を示している。
FIG. 2 shows a neural network operation unit 4
3 shows an example of the structure of the neural network provided in the.

【0028】このニューラルネットワークは、入力層4
1、中間層42および出力層43からなる。入力層41
は、たとえば、128個(16channel ×8frame ) の
入力ユニットから構成されている。中間層42は、入力
層41の各入力ユニットと相互に結合された、たとえ
ば、50個の中間ユニットから構成されている。出力層
43は、中間層42の各中間ユニットと相互に結合され
た、たとえば、20個の出力ユニットから構成されてい
る。
This neural network has an input layer 4
1, an intermediate layer 42 and an output layer 43. Input layer 41
Is composed of, for example, 128 (16 channel × 8 frame) input units. The intermediate layer 42 is composed of, for example, 50 intermediate units that are mutually coupled to the input units of the input layer 41. The output layer 43 is composed of, for example, 20 output units that are mutually coupled to the respective intermediate units of the intermediate layer 42.

【0029】ここでは、認識対象音声は20個あるもの
とする。各認識対象音声を表す音声識別データは、出力
ユニットに対応した20個のデータからなり、その1つ
のみが”1”で他が全て”0”のデータで構成されてい
るものとする。そして、データ”1”の位置が、各音声
識別データごとに異なっている。
Here, it is assumed that there are 20 speeches to be recognized. It is assumed that the voice identification data representing each recognition target voice is composed of 20 pieces of data corresponding to the output unit, only one of which is "1" and the other is all "0". The position of the data "1" is different for each voice identification data.

【0030】図3は、ニューラルネットワークの学習方
法を示している。各認識対象音声ごとに、初期学習用標
準音声パターンと追加学習用標準音声パターンとが作成
される(ステップ1)。
FIG. 3 shows a learning method of the neural network. A standard speech pattern for initial learning and a standard speech pattern for additional learning are created for each recognition target speech (step 1).

【0031】つまり、たとえば、図4に示すように、所
定の音声、たとえば「しち」の標準音声信号に対する音
声パワー信号を生成する。そして、好適なしきい値δ1
を用いて、音声区間R1を設定する。また、他の1また
は複数のしきい値δ2、δ3…δn(この例では、δ
2、δ3、δ4)を用いて、音声区間R2、R3…Rn
(この例では、R2、R3、R4)を設定する。
That is, for example, as shown in FIG. 4, a sound power signal for a predetermined sound, for example, a standard sound signal of "shichi" is generated. And the preferred threshold δ1
Is used to set the voice section R1. Further, another one or a plurality of threshold values δ2, δ3 ... δn (in this example, δ
2, δ3, δ4), the voice sections R2, R3 ... Rn
(R2, R3, R4 in this example) are set.

【0032】そして、各音声区間R1〜Rnに対する標
準音声パターンが作成される。音声区間R1に対する標
準音声パターンが初期学習用標準音声パターンであり、
音声区間R2〜Rnに対する標準音声パターンが追加学
習用標準音声パターンである。各標準音声パターンとし
ては、対応する音声区間を8等分した各区間それぞれの
平均スペクトルが用いられている。また、各区間の音声
スペクトルは、予め定められた16の周波数帯域に対す
る音声スペクトルから構成されている。
Then, a standard voice pattern is created for each voice section R1 to Rn. The standard voice pattern for the voice section R1 is the standard voice pattern for initial learning,
The standard voice pattern for the voice sections R2 to Rn is the standard voice pattern for additional learning. As each standard speech pattern, an average spectrum of each section obtained by dividing the corresponding speech section into eight equal parts is used. The voice spectrum of each section is composed of voice spectra for 16 predetermined frequency bands.

【0033】このようにして、全ての認識対象音声に対
する初期学習用標準音声パターンおよび追加学習用標準
音声パターンとが作成されると、初期学習が行なわれる
(ステップ2)。
When the standard voice pattern for initial learning and the standard voice pattern for additional learning have been created in this way for all recognition target voices, initial learning is performed (step 2).

【0034】つまり、各認識対象音声に対する初期学習
用標準音声パターンを入力パターンとし、各入力パター
ンに対応する音声を表す音声識別データを教師データと
して、バックプロパゲーション法により、ニューラルネ
ットワークを学習させる。
That is, the neural network is trained by the back propagation method using the initial learning standard voice pattern for each recognition target voice as the input pattern and the voice identification data representing the voice corresponding to each input pattern as the teacher data.

【0035】次に、追加学習用の入力パターンの選択処
理が行なわれる(ステップ3)。
Next, a process of selecting an input pattern for additional learning is performed (step 3).

【0036】つまり、各認識対象音声に対する追加学習
用標準音声パターンを、初期学習済のニューラルネット
ワークに順次入力し、その出力に基づいて音声認識結果
を得る。追加学習用標準音声パターンのうち、誤認識が
発生したものを、追加学習用の入力パターンとして選択
する。
That is, the standard voice patterns for additional learning for the respective voices to be recognized are sequentially input to the neural network which has been initially learned, and the voice recognition result is obtained based on the output. Of the standard voice patterns for additional learning, the one in which erroneous recognition has occurred is selected as the input pattern for additional learning.

【0037】たとえば、図4に示す音声区間R2、R3
およびR4に対する追加学習用標準音声パターンを初期
学習済のニューラルネットワークに順次入力して音声認
識を行なった場合に、各追加学習用標準音声パターンに
対して本来”しち”と認識されるべきところが、”に”
と誤認識されたとする。このような場合には、音声区間
R2、R3およびR4に対する追加学習用標準音声パタ
ーンは、追加学習用の入力パターンとして選択される。
For example, the voice sections R2 and R3 shown in FIG.
When voice recognition is performed by sequentially inputting additional learning standard speech patterns for R4 and R4 to a neural network that has been initially learned, there is a place where each additional learning standard speech pattern should be recognized as "shichi". , "To"
It is assumed that it was mistakenly recognized as. In such a case, the standard voice pattern for additional learning for the voice sections R2, R3, and R4 is selected as the input pattern for additional learning.

【0038】次に、追加学習が行なわれる(ステップ
4)。
Next, additional learning is performed (step 4).

【0039】つまり、ステップ3で追加学習用の入力パ
ターンとして選択された各追加学習用標準音声パターン
と、ステップ1で作成された初期学習用標準音声パター
ンとを入力パターンとして、初期学習済のニューラルネ
ットワークを追加学習させる。この際、各追加学習用標
準音声パターンに対する教師データとしては、全て0の
データを用いる。また、初期学習用標準音声パターンに
対する教師データとしては、各初期学習用標準音声パタ
ーンに対応する音声を表す音声識別データが用いられ
る。
That is, the initial learned standard speech pattern selected in step 3 as the additional learning input standard pattern and the initial learning standard speech pattern created in step 1 are used as the input patterns. Train additional networks. At this time, all 0 data is used as teacher data for each standard voice pattern for additional learning. Also, as the teacher data for the standard voice pattern for initial learning, voice identification data representing the voice corresponding to each standard voice pattern for initial learning is used.

【0040】図4を例にとると、音声区間R2、R3、
R4に対する追加学習用標準音声パターンが入力パター
ンとされ、全て0の教師データを用いて、追加学習が行
なわれる。
Taking FIG. 4 as an example, the voice sections R2, R3,
The standard voice pattern for additional learning for R4 is used as an input pattern, and additional learning is performed using the teacher data of all 0s.

【0041】図1の音声認識装置の動作について説明す
る。
The operation of the voice recognition apparatus of FIG. 1 will be described.

【0042】音声分析部1は、入力音声の音声パワー信
号と、入力音声に対する音声スペクトルとを生成する。
入力音声の音声パワー信号は、音声区間検出部2に送ら
れる。入力音声に対する音声スペクトルは、音声パター
ン作成部3に送られる。
The voice analysis unit 1 generates a voice power signal of the input voice and a voice spectrum for the input voice.
The voice power signal of the input voice is sent to the voice section detection unit 2. The voice spectrum for the input voice is sent to the voice pattern creating unit 3.

【0043】音声検出部21は、図5に示すように、音
声検出用しきい値αを用いて、入力された音声パワー信
号中の音声部分を検出する。
As shown in FIG. 5, the voice detecting section 21 detects the voice portion in the input voice power signal by using the voice detecting threshold value α.

【0044】音声区間切出し部22は、図5に示すよう
に、複数の切出し用しきい値β1、β2、β3、β4を
用いて、複数の音声区間を設定する。この例では、第1
から第4の音声区間L1、L2、L3、L4を設定す
る。そして、設定した各音声区間L1〜L4の開始点と
終了点とを、各音声区間L1〜L4に対応させて、切出
し位置記憶部23に格納する。
As shown in FIG. 5, the voice section cutout unit 22 sets a plurality of voice sections using a plurality of cutout thresholds β1, β2, β3, β4. In this example, the first
To the fourth voice section L1, L2, L3, L4. Then, the set start points and end points of the respective voice sections L1 to L4 are stored in the cutout position storage unit 23 in association with the respective voice sections L1 to L4.

【0045】各切出し用しきい値β1、β2、β3、β
4は、たとえば、次のようにして設定される。まず、最
小の切出し用しきい値β1が、音声検出部21によって
検出された音声部分の開始位置より所定時間前の雑音パ
ワーに基づいて決定される。そして、決定された最小の
切出し用しきい値β1に、定数γが加算されることによ
りしきい値β2が求められ、しきい値β2に定数γが加
算されることによりしきい値β3が求められ、しきい値
β3に定数γが加算されることによりしきい値β4が求
められる。
Threshold values β1, β2, β3, β for cutting out
4 is set as follows, for example. First, the minimum clipping threshold β1 is determined based on the noise power detected by the voice detection unit 21 a predetermined time before the start position of the voice portion. Then, a threshold β2 is obtained by adding a constant γ to the determined minimum clipping threshold β1, and a threshold β3 is obtained by adding a constant γ to the threshold β2. Then, the threshold value β4 is obtained by adding the constant γ to the threshold value β3.

【0046】音声パターン作成部3は、音声区間切出し
部22によって求められた各音声区間L1〜L4に対す
る音声スペクトルに基づいて、各音声区間L1〜L4ご
とに音声パターンを作成して、ニューラルネットワーク
演算部4に入力させる。
The voice pattern creating section 3 creates a voice pattern for each of the voice sections L1 to L4 based on the voice spectrum for each of the voice sections L1 to L4 obtained by the voice section cutout section 22, and calculates the neural network. Input to the section 4.

【0047】つまり、切出し位置記憶部23に格納され
ている第1の音声区間L1の開始点と終了点とに基づい
て、当該音声区間L1に対する音声パターン(P1)を
作成する。この音声パターンとしては、当該音声区間を
8等分した各区間それぞれの平均スペクトルが用いられ
ている。そして、各区間の音声スペクトルパターンは、
予め定められた16の周波数帯域に対する音声スペクト
ルから構成されている。作成された第1の音声パターン
(P1)は、学習済のニューラルネットワークに入力さ
れる。
That is, the voice pattern (P1) for the voice section L1 is created based on the start point and the end point of the first voice section L1 stored in the cut-out position storage section 23. As the voice pattern, the average spectrum of each of the eight voice segments is used. And the speech spectrum pattern of each section is
It is composed of speech spectra for 16 predetermined frequency bands. The created first voice pattern (P1) is input to the learned neural network.

【0048】学習済のニューラルネットワークに、第1
の音声パターン(P1)が入力されることにより、第1
の音声パターン(P1)に対応する出力パターンが得ら
れる。そして、得られた出力パターンに基づいて、認識
結果と出力最大値(20個の出力のうちの最大値)と
が、第1認識結果として認識結果記憶部5に記憶され
る。
In the learned neural network, the first
By inputting the voice pattern (P1) of
An output pattern corresponding to the voice pattern (P1) is obtained. Then, based on the obtained output pattern, the recognition result and the output maximum value (the maximum value of the 20 outputs) are stored in the recognition result storage unit 5 as the first recognition result.

【0049】次に、切出し位置記憶部13に格納されて
いる第2の音声区間L2の開始点と終了点とに基づい
て、当該音声区間L2に対する音声パターン(P2)が
作成され、作成された第2の音声パターン(P2)が学
習済のニューラルネットワークに入力される。これによ
り、第2の音声パターン(P2)に対応する出力パター
ンが得られる。得られた出力パターンに基づいて、認識
結果と出力最大値が、第2認識結果として認識結果記憶
部5に記憶される。
Next, a voice pattern (P2) for the voice section L2 is created and created based on the start point and end point of the second voice section L2 stored in the cut-out position storage section 13. The second voice pattern (P2) is input to the learned neural network. As a result, an output pattern corresponding to the second voice pattern (P2) is obtained. The recognition result and the maximum output value are stored in the recognition result storage unit 5 as the second recognition result based on the obtained output pattern.

【0050】次に、第3の音声区間L3の開始点と終了
点とに基づいて、当該音声区間L3に対する音声パター
ン(P3)が作成されて、学習済のニューラルネットワ
ークに入力される。これにより、第3の音声パターン
(P3)に対応する出力パターンが得られる。得られた
出力パターンに基づいて、認識結果と出力最大値が、第
3認識結果として認識結果記憶部5に記憶される。
Next, a voice pattern (P3) for the voice section L3 is created based on the start point and end point of the third voice section L3, and is input to the learned neural network. As a result, an output pattern corresponding to the third voice pattern (P3) is obtained. Based on the obtained output pattern, the recognition result and the maximum output value are stored in the recognition result storage unit 5 as the third recognition result.

【0051】次に、第4の音声区間L4の開始点と終了
点とに基づいて、当該音声区間L4に対する音声パター
ン(P4)が作成されて、学習済のニューラルネットワ
ークに入力される。これにより、第4の音声パターン
(P4)に対応する出力パターンが得られる。得られた
出力パターンに基づいて、認識結果と出力最大値が、第
4認識結果として認識結果記憶部5に記憶される。
Next, a voice pattern (P4) for the voice section L4 is created based on the start point and the end point of the fourth voice section L4 and is input to the learned neural network. As a result, an output pattern corresponding to the fourth voice pattern (P4) is obtained. The recognition result and the maximum output value are stored in the recognition result storage unit 5 as the fourth recognition result based on the obtained output pattern.

【0052】このようにして、第1〜第4の音声パター
ン(P1〜P4)に対する第1〜第4の認識結果が得ら
れると、認識結果判定部6は、出力パターン記憶部5に
記憶されている第1〜第4の認識結果のうち、出力最大
値が”1”に最も近い音声認識結果を、当該検出音声部
分の音声認識結果として選択して出力する。つまり、音
声識別データ(教師データ)に類似度が最も高い出力パ
ターンに基づいて、入力音声が認識される。
In this way, when the first to fourth recognition results for the first to fourth voice patterns (P1 to P4) are obtained, the recognition result determination section 6 is stored in the output pattern storage section 5. Among the first to fourth recognition results, the voice recognition result whose output maximum value is closest to "1" is selected and output as the voice recognition result of the detected voice portion. That is, the input voice is recognized based on the output pattern having the highest degree of similarity to the voice identification data (teacher data).

【0053】上記実施例では、1つの音声検出部分に対
して、複数の切出し用しきい値β1〜β4によって得ら
れた複数の音声区間L1〜L4が設定されている。そし
て、各音声区間ごとの音声パターンに基づいて、当該音
声検出部分の音声が認識されているので、雑音が音声区
間に含まれることによって誤認識が発生したり、音声パ
ワーの小さい語尾等が音声区間から脱落することによっ
て誤認識が発生したりするといったことが防止される。
この結果、音声認識精度が向上する。
In the above embodiment, a plurality of voice sections L1 to L4 obtained by a plurality of clipping thresholds β1 to β4 are set for one voice detection portion. Then, since the voice of the voice detection portion is recognized based on the voice pattern for each voice section, erroneous recognition occurs due to noise being included in the voice section, and the ending of a voice with a small voice power is voiced. It is possible to prevent erroneous recognition from occurring due to dropping from the section.
As a result, the voice recognition accuracy is improved.

【0054】また、上記実施例では、各認識対象音声に
対して、複数のしきい値によって標準音声パターンを作
成し、それらの標準音声パターンのうち、他の音声と誤
認識される可能性のあるものについては、それらを入力
パターンとし、全て0の教師データを用いて、初期学習
済のニューラルネットワークが追加学習されている。こ
のため、音声パターンが初期学習用標準音声パターンに
近いときのみ、ニューラルネットワークから高感度の出
力パターンが得られる。この結果、認識精度が向上す
る。
In the above embodiment, a standard voice pattern is created for each voice to be recognized with a plurality of threshold values, and there is a possibility that the standard voice pattern may be erroneously recognized as another voice. For some, the initial learned neural network is additionally learned by using them as input patterns and using teacher data of all zeros. Therefore, a highly sensitive output pattern can be obtained from the neural network only when the voice pattern is close to the standard voice pattern for initial learning. As a result, the recognition accuracy is improved.

【0055】上記実施例では、入力音声に対して複数の
しきい値β1〜β4によって複数の音声区間が設定され
ているが、入力音声に対して1つのしきい値によって1
の音声区間のみ設定するようにしてもよい。
In the above embodiment, a plurality of voice sections are set for the input voice by a plurality of thresholds β1 to β4, but one threshold is set for the input voice by one threshold.
It is also possible to set only the voice section.

【0056】上記実施例では、音声区間は、入力音声の
音声パワーと、切出し用しきい値とに基づいて設定され
ているが、音声パワー以外の音声区間判定用のパラメー
タと、そのパラメータに応じたしきい値とに基づいて音
声区間を設定してもよい。音声区間判定用のパラメータ
としては、音声パワー以外に、パワーの傾き、広域パワ
ー、低域パワー等がある。
In the above embodiment, the voice section is set based on the voice power of the input voice and the cut-out threshold value. However, the voice section determination parameters other than the voice power and the parameters are set according to the parameters. The voice section may be set based on the threshold value. Besides the voice power, the parameters for the voice section determination include a power slope, a wide range power, a low range power, and the like.

【0057】また、各音声区間ごとの音声パターンをそ
れぞれ作成するための、音声区間の特徴としては、音声
スペクトルの他、音声スペクトルの傾き、音声パワー等
を用いてもよい。
Further, as the characteristics of the voice section for creating the voice pattern for each voice section, the inclination of the voice spectrum, the voice power, etc. may be used in addition to the voice spectrum.

【0058】[0058]

【発明の効果】この発明によれば、認識精度の向上が図
れる。
According to the present invention, the recognition accuracy can be improved.

【図面の簡単な説明】[Brief description of drawings]

【図1】音声認識装置の構成を示すブロック図である。FIG. 1 is a block diagram showing a configuration of a voice recognition device.

【図2】図1のニューラルネットワーク演算部に設けら
れているニューラルネットワークの構造を示す模式図で
ある。
FIG. 2 is a schematic diagram showing a structure of a neural network provided in a neural network operation unit of FIG.

【図3】ニューラルネットワークの学習方法を説明する
ためのフローチャートである。
FIG. 3 is a flowchart for explaining a learning method of a neural network.

【図4】ニューラルネットワークの初期学習用標準音声
パターンと、追加学習用標準音声パターンとを作成する
方法を説明するためのタイムチャートである。
FIG. 4 is a time chart for explaining a method of creating a standard speech pattern for initial learning and a standard speech pattern for additional learning of a neural network.

【図5】図1の音声認識装置において、複数の切出し用
しきい値に基づいて複数の音声区間が設定されることを
示すタイムチャートである。
5 is a time chart showing that a plurality of voice sections are set on the basis of a plurality of clipping thresholds in the voice recognition device of FIG. 1. FIG.

【図6】従来の音声認識装置の構成を示すブロック図で
ある。
FIG. 6 is a block diagram showing a configuration of a conventional voice recognition device.

【図7】図6の音声認識装置において、1つの切出し用
しきい値に基づいて1つの音声区間が設定されることを
示すタイムチャートである。
7 is a time chart showing that one voice section is set based on one clipping threshold value in the voice recognition device of FIG. 6. FIG.

【符号の説明】[Explanation of symbols]

1 音声分析部 2 音声区間検出部 3 音声パターン作成部 4 ニューラルネットワーク演算部 5 認識結果記憶部 6 認識結果判定部 21 音声検出部 22 音声区間切出し部 23 切出し位置記憶部 1 voice analysis unit 2 voice section detection unit 3 voice pattern creation unit 4 neural network operation unit 5 recognition result storage unit 6 recognition result determination unit 21 voice detection unit 22 voice section cutout unit 23 cutout position storage unit

Claims (3)

【特許請求の範囲】[Claims] 【請求項1】 入力音声に対して音声区間を設定する音
声区間設定手段、音声区間の特徴に基づいて、音声区間
の音声パターンを作成する音声パターン作成手段、およ
び音声パターンが入力されるニューラルネットワークを
有しかつニューラルネットワークの出力に基づいて入力
音声を認識する音声認識手段を備えており、 各認識対象音声ごとに、好適な音声区間に基づく初期学
習用標準音声パターンと、好適な音声区間とは異なる音
声区間に基づく追加学習用標準音声パターンとが作成さ
れ、初期学習用標準音声パターンを入力パターンとし、
各入力パターンに対応する音声を表す音声識別データを
教師データとして、ニューラルネットワークが初期学習
され、追加学習用標準音声パターンのうち、初期学習済
のニューラルネットワークにそれが入力されて音声認識
が行なわれたときに、誤認識が生じたものを入力パター
ンとし、反教師データを用いてニューラルネットワーク
が追加学習されている音声認識装置。
1. A voice section setting means for setting a voice section for an input voice, a voice pattern creating means for producing a voice pattern of a voice section based on the characteristics of the voice section, and a neural network to which the voice pattern is input. And a voice recognition means for recognizing an input voice based on the output of the neural network. For each recognition target voice, a standard voice pattern for initial learning based on a suitable voice segment, and a suitable voice segment Is created as a standard speech pattern for additional learning based on different speech sections, and the standard speech pattern for initial learning is used as an input pattern,
The neural network is initially learned by using the voice identification data representing the voice corresponding to each input pattern as the teacher data, and it is input to the already learned neural network among the additional learning standard voice patterns for voice recognition. A speech recognition device in which a neural network is additionally learned by using anti-teaching data by using an input pattern that is erroneously recognized.
【請求項2】 入力音声に対して複数の音声区間を設定
する音声区間設定手段、各音声区間の特徴に基づいて、
各音声区間ごとの音声パターンをそれぞれ作成する音声
パターン作成手段、および各音声区間ごとの音声パター
ンがそれぞれ入力されるニューラルネットワークを有し
かつ各音声区間ごとの音声パターンに対するニューラル
ネットワークの出力に基づいて入力音声を認識する音声
認識手段を備えており、 各認識対象音声ごとに、好適な音声区間に基づく初期学
習用標準音声パターンと、好適な音声区間とは異なる音
声区間に基づく追加学習用標準音声パターンとが作成さ
れ、初期学習用標準音声パターンを入力パターンとし、
各入力パターンに対応する音声を表す音声識別データを
教師データとして、ニューラルネットワークが初期学習
され、追加学習用標準音声パターンのうち、初期学習済
のニューラルネットワークにそれが入力されて音声認識
が行なわれたときに、誤認識が生じたものを入力パター
ンとし、反教師データを用いてニューラルネットワーク
が追加学習されている音声認識装置。
2. A voice section setting means for setting a plurality of voice sections for an input voice, based on characteristics of each voice section,
Based on the output of the neural network having a voice pattern creating means for creating a voice pattern for each voice section and a neural network to which the voice pattern for each voice section is respectively input, and for the voice pattern for each voice section A voice recognition means for recognizing an input voice is provided, and for each recognition target voice, a standard voice pattern for initial learning based on a suitable voice section and a standard voice for additional learning based on a voice section different from the suitable voice section. The pattern and is created, the standard voice pattern for initial learning as the input pattern,
The neural network is initially learned by using the voice identification data representing the voice corresponding to each input pattern as the teacher data, and is input to the initially learned neural network of the additional learning standard voice patterns for voice recognition. A speech recognition device in which a neural network is additionally learned by using anti-teaching data by using an input pattern that is erroneously recognized.
【請求項3】 音声区間の特徴が音声スペクトルである
請求項1および2のいずれかに記載の音声認識装置。
3. The voice recognition device according to claim 1, wherein the feature of the voice section is a voice spectrum.
JP29172594A 1994-11-25 1994-11-25 Voice recognition device Expired - Fee Related JP3322491B2 (en)

Priority Applications (1)

Application Number Priority Date Filing Date Title
JP29172594A JP3322491B2 (en) 1994-11-25 1994-11-25 Voice recognition device

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
JP29172594A JP3322491B2 (en) 1994-11-25 1994-11-25 Voice recognition device

Publications (2)

Publication Number Publication Date
JPH08146996A true JPH08146996A (en) 1996-06-07
JP3322491B2 JP3322491B2 (en) 2002-09-09

Family

ID=17772592

Family Applications (1)

Application Number Title Priority Date Filing Date
JP29172594A Expired - Fee Related JP3322491B2 (en) 1994-11-25 1994-11-25 Voice recognition device

Country Status (1)

Country Link
JP (1) JP3322491B2 (en)

Cited By (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2013503393A (en) * 2009-08-31 2013-01-31 シマンテック コーポレーション System and method for using multiple in-line heuristics to reduce false positives
JP2016161823A (en) * 2015-03-03 2016-09-05 株式会社日立製作所 Acoustic model learning support device and acoustic model learning support method
WO2020162239A1 (en) * 2019-02-08 2020-08-13 日本電信電話株式会社 Paralinguistic information estimation model learning device, paralinguistic information estimation device, and program
JP2022150777A (en) * 2021-03-26 2022-10-07 日本放送協会 Section extraction device and program

Families Citing this family (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP3474949B2 (en) 1994-11-25 2003-12-08 三洋電機株式会社 Voice recognition device

Family Cites Families (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2543603B2 (en) 1989-11-16 1996-10-16 積水化学工業株式会社 Word recognition system

Cited By (6)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2013503393A (en) * 2009-08-31 2013-01-31 シマンテック コーポレーション System and method for using multiple in-line heuristics to reduce false positives
JP2015215906A (en) * 2009-08-31 2015-12-03 シマンテック コーポレーションSymantec Corporation System and method for using multiple in-line heuristics to reduce false positives
JP2016161823A (en) * 2015-03-03 2016-09-05 株式会社日立製作所 Acoustic model learning support device and acoustic model learning support method
WO2020162239A1 (en) * 2019-02-08 2020-08-13 日本電信電話株式会社 Paralinguistic information estimation model learning device, paralinguistic information estimation device, and program
JP2020129051A (en) * 2019-02-08 2020-08-27 日本電信電話株式会社 Paralinguistic information estimation model learning device, paralinguistic information estimation device, and program
JP2022150777A (en) * 2021-03-26 2022-10-07 日本放送協会 Section extraction device and program

Also Published As

Publication number Publication date
JP3322491B2 (en) 2002-09-09

Similar Documents

Publication Publication Date Title
US8140330B2 (en) System and method for detecting repeated patterns in dialog systems
EP2486562B1 (en) Method for the detection of speech segments
EP0435282B1 (en) Voice recognition apparatus
US8145486B2 (en) Indexing apparatus, indexing method, and computer program product
US6134527A (en) Method of testing a vocabulary word being enrolled in a speech recognition system
EP0623914A1 (en) Speaker independent isolated word recognition system using neural networks
WO1996013828A1 (en) Method and system for identifying spoken sounds in continuous speech by comparing classifier outputs
Awotunde et al. Speech segregation in background noise based on deep learning
CN117292688A (en) A control method based on intelligent voice mouse and intelligent voice mouse
Yarra et al. A mode-shape classification technique for robust speech rate estimation and syllable nuclei detection
Prabavathy et al. An enhanced musical instrument classification using deep convolutional neural network
JP3322491B2 (en) Voice recognition device
EP0109140B1 (en) Recognition of continuous speech
US5727121A (en) Sound processing apparatus capable of correct and efficient extraction of significant section data
JP3322536B2 (en) Neural network learning method and speech recognition device
JP3474949B2 (en) Voice recognition device
Migel et al. Speech Recognition System for Ukrainian Language
JP3357752B2 (en) Pattern matching device
KR20260001132A (en) Conversation method and system for operating conversation models based on embedding information and the distribution of related knowledge about user utterance
JPH03269500A (en) Speech recognition device
JP2891259B2 (en) Voice section detection device
Ananthapadmanabha et al. Relative occurrences and difference of extrema for detection of transitions between broad phonetic classes
JPH06110491A (en) Speech recognition device
HK40002006B (en) Speech marking method and device, and equipment
HK40002006A (en) Speech marking method and device, and equipment

Legal Events

Date Code Title Description
LAPS Cancellation because of no payment of annual fees