JPH06167992A - Voice pattern creating device and standard pattern registration device using the same - Google Patents
Voice pattern creating device and standard pattern registration device using the sameInfo
- Publication number
- JPH06167992A JPH06167992A JP4341224A JP34122492A JPH06167992A JP H06167992 A JPH06167992 A JP H06167992A JP 4341224 A JP4341224 A JP 4341224A JP 34122492 A JP34122492 A JP 34122492A JP H06167992 A JPH06167992 A JP H06167992A
- Authority
- JP
- Japan
- Prior art keywords
- pattern
- standard pattern
- standard
- voice
- new
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Abstract
(57)【要約】
【目的】 任意の言葉の標準パターンを登録して音声認
識させる場合、誤認識が生ずるのを著しく低減すること
が可能である。
【構成】 標準パターン作成部6は、音声区間検出部4
から出力された特徴変換結果の音声に係わる部分を標準
パターンとしてそのままの形で辞書5に登録する登録部
11と、そのままの形で辞書5に登録されている複数の
標準パターンを例えばワーキングエリアに読み出し、各
標準パターンをそれぞれ複数の部分に分割する分割部1
2と、複数の特徴パターンをそれぞれ複数の部分に分割
したときにこれらを組合せ連結して新たな音声の標準パ
ターンを生成し、登録部11によりそのままの形で辞書
5に登録されている標準パターンと併せて、新たな標準
パターンの全てを,または一部を辞書5に登録するパタ
ーン生成部13とを備えている。
(57) [Abstract] [Purpose] When registering a standard pattern of an arbitrary word for voice recognition, it is possible to significantly reduce the occurrence of erroneous recognition. [Structure] The standard pattern creation unit 6 includes a voice section detection unit 4
The registration unit 11 for registering the voice-related part of the feature conversion result output from the dictionary 5 as a standard pattern in the dictionary 5 as it is, and the plurality of standard patterns registered in the dictionary 5 as they are in the working area, for example, in the working area. Dividing unit 1 for reading and dividing each standard pattern into a plurality of portions
2 and a plurality of characteristic patterns are divided into a plurality of parts, respectively, these are combined and connected to generate a new standard pattern of voice, and the standard pattern registered in the dictionary 5 as it is by the registration unit 11. In addition, a pattern generation unit 13 that registers all or some of the new standard patterns in the dictionary 5 is provided.
Description
【0001】[0001]
【産業上の利用分野】本発明は、音声認識等に用いられ
る音声パターンを作成する音声パターン作成装置および
それを用いた標準パターン登録装置に関する。BACKGROUND OF THE INVENTION 1. Field of the Invention The present invention relates to a voice pattern creating apparatus for creating a voice pattern used for voice recognition and the like and a standard pattern registration apparatus using the same.
【0002】[0002]
【従来の技術】一般に、音声認識装置においては、実際
の音声認識を行なうに先立って、種々の音声の標準的な
特徴パターン,すなわち標準パターンを辞書に予め登録
しておき、音声認識時には、話者によって発声された未
知の音声の特徴パターンを辞書に登録されている種々の
標準パターンと照合してこれらの間の類似度を求め、種
々の標準パターンのうちで最も大きな類似度を与えた標
準パターンに対応した言葉(例えば単語)を認識結果と
して出力するようになっている。2. Description of the Related Art Generally, in a voice recognition apparatus, standard feature patterns of various voices, that is, standard patterns are registered in a dictionary in advance before actual voice recognition is performed. The characteristic pattern of the unknown voice uttered by the person is collated with various standard patterns registered in the dictionary to find the similarity between them, and the standard that gives the largest similarity among the various standard patterns. A word (for example, a word) corresponding to the pattern is output as a recognition result.
【0003】ところで、この種の音声認識装置では、標
準パターンとして、例えば、「足立」,「井上」,「宇
野」,「遠藤」,「小野」,「加藤」,「木下」,「日
下部」,「小林」,「山田」の10人の名前の単語に対
応したパターンが登録されている時に、話者が「佐藤」
と発声すると、上記10個の標準パターンの中で「佐
藤」に最も類似している「加藤」が認識結果として出力
されてしまう。By the way, in this type of voice recognition device, as a standard pattern, for example, "Adachi", "Inoue", "Uno", "Endo", "Ono", "Kato", "Kinoshita", "Kusenbe" , "Kobayashi", "Yamada" when the pattern corresponding to the words of 10 names is registered, the speaker is "Sato"
"Kato", which is the most similar to "Sato" among the above 10 standard patterns, is output as a recognition result.
【0004】このような誤認識を防止するため、従来で
は、特開昭61−133996号,特開昭63−295
394号に開示されているように、類似度に閾値を設け
る技術が提案されている。これによれば、種々の標準パ
ターンのうち、最も大きな類似度を与える標準パターン
が選択されたときにも、この類似度が閾値以下である場
合には、これをリジェクトすることができ、誤認識がな
されるのを防止できる。また、特開昭58−86598
号に開示されているような標準パターンに重み付けをし
て、上記のような誤認識を防止する技術も提案されてい
る。In order to prevent such erroneous recognition, conventionally, Japanese Patent Laid-Open Nos. 61-133996 and 63-295 have been used.
As disclosed in Japanese Patent No. 394, a technique of setting a threshold value for the similarity has been proposed. According to this, even when the standard pattern that gives the highest similarity among various standard patterns is selected, if this similarity is less than or equal to the threshold value, it can be rejected, resulting in erroneous recognition. Can be prevented. Also, JP-A-58-86598
There is also proposed a technique for preventing the erroneous recognition as described above by weighting the standard pattern as disclosed in No.
【0005】[0005]
【発明が解決しようとする課題】しかしながら、上述し
た従来技術において、類似度に対しどの程度の閾値を設
定するかは非常に難かしく、誤認識等を確実に防止する
には限度があった。具体的に説明すると、上述した例に
おいて、「加藤」,「佐藤」の場合は、それぞれ/ka
too/,/satoo/と発音するから、両者は/k
/と/s/が違うだけで他はすべて同じである。この結
果、「加藤」,「佐藤」にはもともと大きな差がなく、
閾値を決めにくいという問題があった。また、この場
合、厳密に閾値を決めると、「加藤」と正しく発声して
いるのにリジェクトされてしまい、認識結果が得られな
いという事態が起きるという欠点があった。However, in the above-mentioned prior art, it is very difficult to set a threshold value for the similarity, and there is a limit to surely prevent erroneous recognition and the like. To be more specific, in the above example, in the case of “Kato” and “Sato”, / ka respectively
Both are / k because they are pronounced as too /, / satoo /
Everything else is the same except that / and / s / are different. As a result, there was no big difference between "Kato" and "Sato" originally,
There was a problem that it was difficult to determine the threshold. Further, in this case, if a strict threshold value is determined, it is rejected even though "Kato" is correctly uttered, and a recognition result cannot be obtained.
【0006】また、標準パターンへの重み付けは、登録
される言葉が予め決まっている場合や、特定話者方式の
時に限られるという問題があった。具体的には、「佐
藤」と「加藤」との違いを強調するため、標準パターン
の「加藤」の/k/の部分に重みを付けておくというこ
とも考えられるが、これは、「加藤」と「佐藤」の違い
を強調するには十分であっても、入力される全ての言葉
に対して十分とはいえない。すなわち、例えば、「加
納」という名前が入力された場合に、「加藤」と異常に
高い類似度を示し、誤認識を生じさせる事態が生ずる。Further, there is a problem that the weighting to the standard pattern is limited when the words to be registered are predetermined or in the case of the specific speaker system. Specifically, in order to emphasize the difference between "Sato" and "Kato", it is possible to weight the / k / part of "Kato" in the standard pattern. It is enough to emphasize the difference between "and" Sato, but not enough for every word entered. That is, for example, when the name "Kano" is input, an abnormally high degree of similarity with "Kato" is displayed, resulting in erroneous recognition.
【0007】このように、上述した従来技術では、任意
の言葉の標準パターンを登録して音声認識させる場合、
誤認識が生ずるのを著しく低減するには限度があり、汎
用的かつ実用的な音声認識装置を提供するのは非常に難
しいという欠点があった。例えば、音声認識の利用者が
任意の言葉を登録して使用するような場合、(特に、あ
る決められたキ−ワ−ドが認識されたとき動作するよう
なアプリケ−ションの場合)、特定の言葉が入力された
か否かを判定できず、このため、利用できる範囲に限界
があった。As described above, in the above-mentioned conventional technique, when a standard pattern of an arbitrary word is registered for voice recognition,
There is a limit to significantly reducing the occurrence of erroneous recognition, and it is very difficult to provide a general-purpose and practical voice recognition device. For example, when a user of voice recognition registers and uses an arbitrary word (especially, in an application that operates when a certain keyword is recognized), it is specified. It was not possible to determine whether or not the word was input, and therefore, there was a limit to the usable range.
【0008】本発明は、任意の言葉の標準パターンを登
録して音声認識させる場合、誤認識が生ずるのを著しく
低減することの可能な汎用的かつ実用的な音声認識装置
を実現することの可能な音声パタ−ン作成装置および標
準パタ−ン登録装置を提供することを目的としている。The present invention can realize a general-purpose and practical voice recognition device capable of remarkably reducing the occurrence of erroneous recognition when a standard pattern of an arbitrary word is registered for voice recognition. The purpose of the present invention is to provide a voice pattern creating device and a standard pattern registration device.
【0009】[0009]
【課題を解決するための手段および作用】上記目的を達
成するために、請求項1記載の音声パタ−ン作成装置
は、複数の音声の特徴パターンのそれぞれを複数の部分
に分割する分割手段と、分割手段により分割された複数
の特徴パターンの各部分を組合せ連結して新たな音声の
特徴パターンを作成するパターン作成手段とを有してい
ることを特徴としている。これにより、他の言葉に対す
る音声の特徴パターンを簡単に作成することができる。In order to achieve the above object, a voice pattern creating apparatus according to a first aspect of the present invention comprises a dividing means for dividing each of a plurality of voice characteristic patterns into a plurality of portions. , And a pattern creating means for creating a new voice characteristic pattern by combining and connecting the respective portions of the plurality of characteristic patterns divided by the dividing means. This makes it possible to easily create a voice characteristic pattern for another word.
【0010】また、請求項2記載の標準パターン登録装
置は、請求項1記載の音声パタ−ン作成装置において、
複数の音声の特徴パターンを標準パターンとしてそのま
まの形で辞書に登録する登録手段をさらに有し、前記分
割手段は、前記標準パターンを複数の部分に分割し、前
記パターン作成手段は、分割された複数の標準パターン
の各部分を組合せ連結して新たな特徴パターンを生成
し、前記登録手段によりそのままの形で登録されている
前記標準パターンと併せて、該新たな特徴パターンを新
たな標準パターンとして辞書に登録することを特徴とし
ている。これにより、任意の言葉の標準パターンを登録
して音声認識させる場合、誤認識が生ずるのを著しく低
減することができる。A standard pattern registration device according to a second aspect is the voice pattern creation device according to the first aspect.
It further has a registration means for registering a plurality of voice characteristic patterns as a standard pattern in the dictionary as it is, the dividing means divides the standard pattern into a plurality of parts, and the pattern creating means divides the divided parts. A new characteristic pattern is generated by combining and connecting each part of a plurality of standard patterns, and the new characteristic pattern is used as a new standard pattern together with the standard pattern registered as it is by the registration means. It is characterized by registering in a dictionary. As a result, when a standard pattern of an arbitrary word is registered for voice recognition, the occurrence of erroneous recognition can be significantly reduced.
【0011】また、請求項9記載の音声パタ−ン作成装
置は、複数の音声の特徴パターンのそれぞれを、周波数
軸方向に一定量だけ高域へおよび/または低域へずら
し、新たな音声の特徴パターンを作成するパターン作成
手段を有していることを特徴としている。これにより、
元の標準パターンと同じ言葉であるが、話者が相違する
場合に対応した音声の標準パターンを簡単に作成するこ
とができる。According to a ninth aspect of the present invention, a voice pattern creating apparatus shifts each of a plurality of voice characteristic patterns to a high frequency band and / or a low frequency band in a frequency axis direction to generate a new voice pattern. It is characterized by having a pattern creating means for creating a characteristic pattern. This allows
Although it is the same word as the original standard pattern, it is possible to easily create a standard pattern of voice corresponding to the case where the speakers are different.
【0012】また、請求項10記載の標準パターン登録
装置は、請求項9記載の音声パターン作成装置におい
て、複数の音声の特徴パターンを標準パターンとしてそ
のままの形で辞書に登録する登録手段をさらに有し、前
記パターン作成手段は、複数の標準パターンのそれぞれ
を周波数軸方向に一定量だけ高域へおよび/または低域
へずらし、新たな音声パターンを作成したとき、前記登
録手段によりそのままの形で登録されている前記標準パ
ターンと併せて、該新たな音声パターンの全て,または
一部を標準パターンとして辞書に登録するようになって
いることを特徴としている。これにより、不特定話者用
の辞書を容易に作成することができ、不特定話者用の音
声認識装置に適用することができる。The standard pattern registration apparatus according to a tenth aspect of the present invention is the voice pattern creation apparatus according to the ninth aspect, further comprising registration means for registering a plurality of voice characteristic patterns as standard patterns in the dictionary as they are. However, the pattern creating means shifts each of the plurality of standard patterns to a high band and / or a low band by a certain amount in the frequency axis direction, and when a new voice pattern is created, the registering means keeps the same form as it is. Along with the registered standard pattern, all or part of the new voice pattern is registered in the dictionary as a standard pattern. Accordingly, the dictionary for the unspecified speaker can be easily created, and the dictionary can be applied to the voice recognition device for the unspecified speaker.
【0013】[0013]
【実施例】以下、本発明の実施例を図面に基づいて説明
する。図1は本発明に係る音声パターン作成装置が適用
された音声認識装置の構成例を示す図である。図1を参
照すると、この音声認識装置は、マイクロフォン等の音
声入力部1と、音声入力部1から入力された音声信号を
アナログ−デジタル変換するA/D変換器2と、デジタ
ル変換された音声信号に対し特徴変換を施す特徴変換部
3と、特徴変換結果から音声に係わる部分を音声区間と
して検出する音声区間検出部4と、種々の標準パターン
が登録される辞書5と、標準パターンを作成し登録する
場合と実際の音声認識を行なう場合とで切り替えがなさ
れるスイッチSWと、標準パターンの作成,登録時に、
音声区間検出部4から出力された特徴変換結果の音声に
係わる部分に基づき音声の標準パターンを作成し、辞書
5に登録する標準パターン作成部6と、音声認識時に、
音声区間検出部4から出力される特徴変換結果の音声に
係わる部分,すなわち未知のパターンを辞書5に登録さ
れている種々の標準パターンと照合して類似度を求め、
最も大きな類似度を与える標準パターンに対応した言葉
(例えば単語)を認識結果として出力する認識部7とを
備えている。なお、上記特徴変換部3は、例えばFFT
(高速フーリエ変換器)により構成されており、この場
合には、音声信号を周波数領域に変換し、特徴量として
音声のスペクトルを得るようになっている。また、この
場合、A/D変換器2としては、16ビット,16KH
Z程度の特性のものが用いられる。Embodiments of the present invention will be described below with reference to the drawings. FIG. 1 is a diagram showing a configuration example of a voice recognition device to which a voice pattern creation device according to the present invention is applied. Referring to FIG. 1, the voice recognition device includes a voice input unit 1 such as a microphone, an A / D converter 2 for analog-digital converting a voice signal input from the voice input unit 1, and a digitally converted voice. A feature conversion unit 3 that performs feature conversion on a signal, a voice section detection unit 4 that detects a portion related to voice from the feature conversion result as a voice section, a dictionary 5 in which various standard patterns are registered, and a standard pattern is created. Switch SW, which is switched between the case of registering and the case of actual voice recognition, and the time of creating and registering the standard pattern.
A standard pattern creating unit 6 that creates a standard pattern of a voice based on a part related to the voice of the feature conversion result output from the voice section detecting unit 4 and registers it in the dictionary 5,
A portion related to the voice of the feature conversion result output from the voice section detection unit 4, that is, an unknown pattern is collated with various standard patterns registered in the dictionary 5 to obtain a similarity,
The recognition unit 7 outputs a word (for example, a word) corresponding to the standard pattern that gives the highest similarity as a recognition result. The feature conversion unit 3 may be, for example, an FFT.
(Fast Fourier Transform), and in this case, the voice signal is transformed into the frequency domain, and the spectrum of the voice is obtained as a feature amount. In this case, the A / D converter 2 has 16 bits and 16 KH.
The one with a characteristic of about Z is used.
【0014】また、標準パターン作成部6は、音声区間
検出部4から出力された特徴変換結果の音声に係わる部
分を標準パターンとしてそのままの形で辞書5に登録す
る登録部11と、そのままの形で辞書5に登録されてい
る複数の標準パターンを例えばワーキングエリアに読み
出し、各標準パターンをそれぞれ複数の部分に分割する
分割部12と、複数の特徴パターンをそれぞれ複数の部
分に分割したときにこれらを組合せ連結して新たな音声
の標準パターンを生成し、登録部11によりそのままの
形で辞書5に登録されている標準パターンと併せて、新
たな標準パターンの全てを,または一部を辞書5に登録
するパターン生成部13とを備えている。Further, the standard pattern creating section 6 registers the part relating to the voice of the feature conversion result output from the voice section detecting section 4 as a standard pattern in the dictionary 5 as it is, and the registering section 11 as it is. When a plurality of standard patterns registered in the dictionary 5 are read into, for example, a working area and each standard pattern is divided into a plurality of parts, and a plurality of characteristic patterns are divided into a plurality of parts, respectively, Are combined and concatenated to generate a new standard pattern of voice, and all or a part of the new standard pattern is combined with the standard pattern registered in the dictionary 5 by the registration unit 11 as it is. And a pattern generation unit 13 that registers the pattern.
【0015】次にこのような構成の音声認識装置の動作
を図2のフローチャートを用いて説明する。実際の音声
認識動作を行なうに先立って、標準パターンの作成処理
が選択されるようにスイッチSWを切り替える。しかる
後、話者は、ある言葉,例えば単語に対応した音声の標
準パターンを辞書5に登録するため、この単語を発声す
る。話者の発声した単語音声が音声入力部1から入力さ
れると(ステップS1)、この入力音声信号は、特徴変
換部3において、特徴量(例えばスペクトル)に変換さ
れ(ステップS2)、音声区間検出部4において、特徴
変換結果のうち音声に係わる部分が音声区間として検出
され(ステップS3)、スイッチSWを介して標準パタ
ーン作成部6に入力する。Next, the operation of the speech recognition apparatus having such a configuration will be described with reference to the flowchart of FIG. Prior to the actual voice recognition operation, the switch SW is switched so that the standard pattern creation process is selected. Then, the speaker utters this word in order to register a standard pattern of voice corresponding to a word, for example, the word in the dictionary 5. When the word voice uttered by the speaker is input from the voice input unit 1 (step S1), the input voice signal is converted into a feature amount (for example, spectrum) in the feature conversion unit 3 (step S2), and the voice section The detection unit 4 detects a voice-related portion of the feature conversion result as a voice section (step S3) and inputs it to the standard pattern creation unit 6 via the switch SW.
【0016】標準パターン作成部6では、先ず、登録部
11において、音声区間検出部4から出力された特徴変
換結果の音声に係わる部分を標準パターンとしてそのま
まの形で言葉(例えば単語)と対応させて辞書5に登録
する(ステップS4)。ステップS1乃至S4の処理を
繰り返し、例えば所定種類の単語音声の標準パターンが
そのままの形で辞書5に登録されたとき(ステップS
5)、さらに、分割部12では、辞書5にそのままの形
で登録された個々の標準パターンを複数の部分に分割す
る(ステップS6)。次いで、パターン生成部13で
は、複数種類の標準パターンの各分割部分を組合せ連結
して、新たな音声の標準パターンとして生成し(ステッ
プS7)、この新たな音声の標準パターンをも辞書5に
併せて登録する(ステップS8)。このようにして、辞
書5には、音声区間検出部4から出力されたそのままの
形での標準パターンと併せて、上記新たな標準パターン
が登録される。In the standard pattern creating section 6, first, in the registering section 11, the part relating to the voice of the feature conversion result output from the voice section detecting section 4 is made to correspond to a word (for example, a word) as a standard pattern as it is. And register it in the dictionary 5 (step S4). When the processes of steps S1 to S4 are repeated and, for example, a standard pattern of a predetermined type of word voice is registered in the dictionary 5 as it is (step S
5) Further, the dividing unit 12 divides each standard pattern registered in the dictionary 5 as it is into a plurality of portions (step S6). Next, in the pattern generation unit 13, the divided portions of the plurality of types of standard patterns are combined and connected to generate a new voice standard pattern (step S7), and the new voice standard pattern is also combined in the dictionary 5. To register (step S8). In this way, the new standard pattern is registered in the dictionary 5 together with the standard pattern output from the voice section detection unit 4 as it is.
【0017】辞書5に上記のような標準パターンが登録
された後、認識処理が選択されるようスイッチSWを切
り替える。この場合には、話者の発声が音声入力部1か
ら入力されると、この入力音声信号は、特徴変換部3で
特徴量(例えばスペクトル)に変換され、音声区間検出
部4,スイッチSWを介して認識部7に入力する。認識
部7では、スイッチSWを介して入力した未知音声のパ
ターンを前述のようにして辞書5に登録されている種々
の標準パターンと照合して類似度を求め、最も大きな類
似度を与える標準パターンに対応した言葉を認識結果と
して出力する。このようにして、一連の音声認識処理を
行なうことができる。After the standard pattern as described above is registered in the dictionary 5, the switch SW is switched so that the recognition process is selected. In this case, when the utterance of the speaker is input from the voice input unit 1, the input voice signal is converted into a feature amount (for example, spectrum) by the feature conversion unit 3, and the voice section detection unit 4 and the switch SW are turned on. It is input to the recognition unit 7 via the. In the recognition unit 7, the unknown voice pattern input via the switch SW is collated with various standard patterns registered in the dictionary 5 as described above to obtain the similarity, and the standard pattern giving the highest similarity is obtained. The word corresponding to is output as a recognition result. In this way, a series of voice recognition processing can be performed.
【0018】ところで、本実施例では、上記のように、
標準パターン作成部6に、分割部12とパターン生成部
13とがさらに設けられていることを特徴としており、
分割部12,パターン生成部13における処理の態様を
より詳細に説明する。By the way, in this embodiment, as described above,
The standard pattern creation unit 6 is further provided with a division unit 12 and a pattern generation unit 13,
The mode of processing in the division unit 12 and the pattern generation unit 13 will be described in more detail.
【0019】分割部12における分割の仕方として、辞
書5に登録された標準パターンを時間軸方向に例えばn
等分する態様が考えられる。より具体的には、上記ステ
ップS5の処理において、辞書5に例えば図3(a)に
示すように、3つの標準パターンA,B,C(それぞ
れ、単語「日下部」,「加藤」,「小林」に対応)が登
録された場合、分割部12では、図4に示すように、こ
れら3つの標準パターンA,B,Cを例えばワーキング
エリアにコピーし、しかる後、これらを時間軸方向に3
等分(n=3)し、それぞれの部分をA1,A2,A
3,B1,B2,B3,C1,C2,C3とするよう分
割することができる。As a method of division in the division unit 12, the standard pattern registered in the dictionary 5 is, for example, n in the time axis direction.
A mode of equally dividing is possible. More specifically, in the process of step S5, as shown in the dictionary 5 for example, as shown in FIG. (Corresponding to “.”) Is registered, the dividing unit 12 copies these three standard patterns A, B, and C into, for example, a working area, and thereafter, these three standard patterns A, B, and C are copied in the time axis direction.
Divide into equal parts (n = 3) and divide each part into A1, A2, A
3, B1, B2, B3, C1, C2, C3.
【0020】また、パターン生成部13は、分割部12
において分割された複数種類の標準パターンの各部分を
組合せて新たな標準パターンを生成するが、この際に一
定の組合せ規則が設定されている。すなわち、パターン
生成部13は、分割前の標準パターンにおいて、先頭に
あった部分は連結後も先頭に位置させるような規則を有
している。この規則によって、パターン生成部13は、
組合せを行なうときに、例えば、分割前のパターンで先
頭にあった部分を特別なグループとして扱い、先頭以外
の部分を連結してから先頭の付加を行なうようになって
いる。Further, the pattern generation unit 13 includes a division unit 12
A new standard pattern is generated by combining the respective parts of the plurality of types of standard patterns that have been divided in 1. At this time, a certain combination rule is set. That is, the pattern generation unit 13 has a rule that the portion at the beginning of the standard pattern before division is positioned at the beginning even after the connection. According to this rule, the pattern generation unit 13
When combining, for example, the portion at the beginning in the pattern before division is treated as a special group, and the portions other than the beginning are connected and then the beginning is added.
【0021】また、パターン生成部13は、分割前の標
準パターンにおいて、末尾にあった部分は連結後も末尾
に位置させるような規則を有している。この規則によっ
て、パターン生成部13は、組合せを行なうときに、分
割前のパターンで末尾にあった部分を特別なグループと
して扱い、末尾以外の部分を連結してからこれを末尾に
付加するようになっている。Further, the pattern generation unit 13 has a rule that the portion at the end of the standard pattern before division is positioned at the end even after connection. According to this rule, the pattern generation unit 13 treats the part at the end of the pattern before division as a special group, combines the parts other than the end, and adds this to the end when combining. Has become.
【0022】換言すれば、このような規則は、図4の例
において、分割された部分を組合せるときに、An,B
n,Cnの添え字nの順番を守ることを意味する。これ
によって、パターンの組替え時に常識的でない標準パタ
ーンが作成されるのを防止することができる。すなわ
ち、実際の音声認識処理においては、「佐藤」,「加
藤」の例からも明らかなように、先頭が違う音で、末尾
が同じ音のような場合が問題になるのであって、「佐
藤」と「藤加」のような先頭と末尾が入れ替わっている
ような言葉に関しては、一般には問題とはならず、従っ
て、上記規則に従った組合せにより作成される標準パタ
ーンのみが重要なものとなる。In other words, such a rule is that when combining the divided parts in the example of FIG.
It means to keep the order of the subscript n of n and Cn. As a result, it is possible to prevent a standard pattern that is not common sense from being created when the patterns are rearranged. That is, in the actual speech recognition process, as is clear from the examples of "Sato" and "Kato", the problem is that the sound has different beginnings and the same sound at the end. "," And "Fujika", such as words with the beginning and the end interchanged, are generally not a problem, and therefore only standard patterns created by combinations according to the above rules are important. Become.
【0023】このような規則に従って図4に示したよう
な3つの標準パターンA,B,Cの3分割された各部分
A1,A2,A3,B1,B2,B3,C1,C2,C
3に対して組合せ連結処理を行なうと、例えば図5に示
すような新しい標準パターンD(この例ではA2+B2
+B3)を作ることができ、この新たな標準パターンD
(A2+B2+B3)を図3(b)に示すように、標準
パターンA,B,Cと併せて辞書5に登録することがで
きる。この場合、組合せ連結によって作り出した新たな
パターンについても長さなどの情報を作りこれを付加し
て登録することができる。但し、図3(b)の例では、
辞書5への登録時に、そのままの形で登録される標準パ
ターンA,B,Cには、それぞれ言葉(単語),すなわ
ち「日下部」,「加藤」,「小林」が対応付けられる
が、新たに作成された標準パターンについては、所定の
言葉(単語)への対応付けがなされない。In accordance with such a rule, three standard patterns A, B, C as shown in FIG. 4 are divided into three parts A1, A2, A3, B1, B2, B3, C1, C2, C.
When the combination and connection process is performed on No. 3, a new standard pattern D (A2 + B2 in this example) as shown in FIG.
+ B3) can be made, and this new standard pattern D
As shown in FIG. 3B, (A2 + B2 + B3) can be registered in the dictionary 5 together with the standard patterns A, B, and C. In this case, it is possible to create information such as the length of the new pattern created by the combination connection and add the information to register it. However, in the example of FIG.
At the time of registration in the dictionary 5, the standard patterns A, B, and C that are registered as they are are associated with words (words), that is, “Kusek”, “Kato”, and “Kobayashi”, respectively. The created standard pattern is not associated with a predetermined word (word).
【0024】いま例えば、A,B,Cの各標準パターン
がそれぞれ「日下部」,「加藤」,「小林」であった場
合、新たな標準パターンD(A2+B2+B3)は「佐
藤」のパターンに非常に近い物となる。従って、標準パ
ターンに重み付けなどをせずとも、類似の標準パターン
が自動的に作成されるので、類似音の入力音声に対して
誤った認識がなされるのを容易に防止することができ
る。すなわち、類似の音声入力に対しては、組合せによ
って新たに作り出された標準パターンの方が大きな類似
性をもつので、認識結果をリジェクトとして出力した
り、あるいは所定のメッセージを出力したりすることが
できる。すなわち、上記の例において、話者が「佐藤」
と発声すると、この入力音声は、標準パターンA,B,
C,Dのうち、新たな標準パターンDに最も類似し、従
って、「加藤」と誤認識しない。また、図3(b)の例
では、新たな標準パターンDに最も類似する場合でも、
この標準パターンDには単語との対応付けがなされてい
ないので、認識結果は、特定の単語ではなく、リジェク
トとなる。これにより、例えば、ある決められたキ−ワ
−ドが認識されたとき動作するようなアプリケ−ション
の場合、簡単な手法で、ある特定の言葉が発声されたか
否かを判定することができる。なお、図3(c)に示す
ように、新たな標準パターンDに対しても言葉(単語)
との対応付けを行なって辞書5に登録することも可能で
あり、この場合には、リジェクトではなく、正しい認識
結果を出力させることができる。For example, if the standard patterns of A, B, and C are "Kusekbe", "Kato", and "Kobayashi", the new standard pattern D (A2 + B2 + B3) is very similar to the pattern of "Sato". It will be close. Therefore, since a similar standard pattern is automatically created without weighting the standard pattern, it is possible to easily prevent erroneous recognition of an input voice having a similar sound. In other words, for similar voice input, the standard pattern newly created by the combination has greater similarity, so it is possible to output the recognition result as a reject or to output a predetermined message. it can. That is, in the above example, the speaker is "Sato".
, The input voice is the standard pattern A, B,
Of C and D, it is the most similar to the new standard pattern D, and thus is not erroneously recognized as "Kato". Further, in the example of FIG. 3B, even when the new standard pattern D is most similar,
Since this standard pattern D is not associated with a word, the recognition result is not a specific word but a reject. Thus, for example, in the case of an application that operates when a certain predetermined keyword is recognized, it is possible to determine with a simple method whether or not a specific word is uttered. . Note that, as shown in FIG. 3C, words are also included in the new standard pattern D.
It is also possible to correlate with and register in the dictionary 5, and in this case, a correct recognition result can be output instead of rejecting.
【0025】また、上記例では、標準パターンの時間軸
方向の分割数nを“3”としたが、nを3分割以外の分
割数とすることもできるし、また、等分割である必要も
ない。Further, in the above example, the division number n of the standard pattern in the time axis direction is "3", but n may be a division number other than three, and it is also necessary that the division is even division. Absent.
【0026】例えば、通常、単語の音声認識の場合、1
0〜20m秒ごとにスペクトルを求める処理が行なわ
れ、また、単語の音声の短かいものは「2/ni/」,
「5/go/」のように単音節で構成され、継続時間が
200m秒程度である。従って、これを多数の部分に分
割しても、それぞれの部分の時間長は短かく、連結した
後の新しいパターン上での影響力は小さい。さらに、単
音節であれば、継続部は何分割しても同じ母音の部分片
となるので、分割することに意味がないという事実も考
慮し、図3(a)のように登録された複数の音声の標準
パターンA,B,Cのそれぞれを、時間軸に対して冒頭
近傍と、末尾近傍に2分し、Aの冒頭近傍にBの末尾近
傍を継ぎ足して新たな音声パターンとして登録し、Bの
冒頭近傍にCの末尾近傍を継ぎ足すことで新たな音声パ
ターンを作成して登録し、この過程を繰り返すことによ
り作成されたパターンの全て,または一部を新たな標準
パターンとして登録するようにすることもできる。For example, normally, in the case of speech recognition of words, 1
A process for obtaining a spectrum is performed every 0 to 20 msec, and a short word speech is "2 / ni /",
It is composed of a single syllable like "5 / go /", and the duration is about 200 msec. Therefore, even if this is divided into a large number of parts, the time length of each part is short and the influence on the new pattern after connection is small. Further, in the case of a single syllable, the continuation part becomes a partial piece of the same vowel no matter how many times it is divided, so in consideration of the fact that there is no point in dividing it, a plurality of registered vowels as shown in FIG. Each of the standard patterns A, B, and C of the voice is divided into two parts near the beginning and the end with respect to the time axis, and the vicinity of the end of B is added to the beginning of A and registered as a new voice pattern. A new voice pattern is created and registered by adding the vicinity of the end of C to the vicinity of the beginning of B, and all or part of the created pattern is registered as a new standard pattern by repeating this process. You can also
【0027】図6はこのような処理の流れを示すフロー
チャートである。図6を参照すると、分割部12は、辞
書5に登録されているm個の標準パターンを3分割する
かわりに2分割して、パターン生成部13に与える(ス
テップS11)。パターン生成部13では、先ず、本来
の言葉として登録されたm個の標準パターンに番号を付
け(ステップS12)、2分割された部分を各番号のパ
ターンの冒頭部,末尾部とする。次いで、カウンタiを
初期化して“1”にし(ステップS13)、i番目のパ
ターンの冒頭部と(i+1)番目のパターンの末尾部と
を連結して新しい標準パターンを形成する(ステップS
15)。iを“1”ずつ歩進して(ステップS16)、
上記処理を繰り返し行ない、カウンタiがmになったと
き(ステップS14)、m番目のパターンの冒頭に最初
の1番目のパターンの末尾を付け加えて(ステップS1
7)、処理を終了する。この結果、短かい音声にとって
も効果的な、また、無駄のないパターン連結を行なうこ
とができる。なお、上記処理例では、カウンタiを
“1”ずつ歩進し、パターン番号を1番ずつずらすよう
にしたが、1番ずつである必要もないし、また、ランダ
ムに継ぎ合わせるようにしても良い。FIG. 6 is a flow chart showing the flow of such processing. Referring to FIG. 6, the dividing unit 12 divides the m standard patterns registered in the dictionary 5 into two instead of dividing them into two, and gives the divided patterns to the pattern generating unit 13 (step S11). In the pattern generation unit 13, first, the m standard patterns registered as the original words are numbered (step S12), and the two-divided portions are used as the beginning and the end of each numbered pattern. Next, the counter i is initialized to "1" (step S13) and the beginning of the i-th pattern and the end of the (i + 1) -th pattern are connected to form a new standard pattern (step S13).
15). i is incremented by 1 (step S16),
The above process is repeated, and when the counter i reaches m (step S14), the end of the first 1st pattern is added to the beginning of the mth pattern (step S1).
7), the process ends. As a result, it is possible to perform pattern concatenation that is effective for short voices and has no waste. In the above processing example, the counter i is incremented by "1" and the pattern number is shifted by one, but the pattern number does not have to be one by one and may be randomly spliced. .
【0028】このように、登録されている複数種類の標
準パターンを時間軸方向にn分割して組合せを行ない、
新たな標準パターンを作成したが、その際、上述の例,
例えば2分割の例では、一定の時間長(例えば2等分な
ど)での分割を行なっている。しかしながら、このよう
な分割では、厳密には、単語中のある音韻をその真ん中
で分割してしまうことが避けられない。従って、音韻の
切れ目とパターンの分割部とをできる限り一致させるの
が望ましい。このために、例えば、複数の音声の標準パ
ターンA,B,Cのそれぞれについて、時間軸方向にパ
ターンの変化の大きな部分を検出し、この部分で、末尾
近傍に分割し、Aの冒頭近傍にBの末尾近傍を継ぎ足し
て新たな音声パターンとして登録し、Bの冒頭近傍にC
の末尾近傍を継ぎ足すことで新たな音声パターンを作成
して登録し、この過程を繰り返すことにより作成された
パターンの全て、または一部を新たな標準パターンとし
て登録するようにすることもできる。In this way, a plurality of types of registered standard patterns are divided into n in the time axis direction and combined.
I created a new standard pattern, with the above example,
For example, in the case of two divisions, division is performed for a fixed time length (for example, equally divided into two). However, with such division, strictly speaking, it is inevitable that a phoneme in a word is divided in the middle. Therefore, it is desirable to match the phoneme breaks and the pattern divisions as much as possible. For this purpose, for example, for each of the standard patterns A, B, and C of a plurality of voices, a portion with a large change in the pattern in the time axis direction is detected, and at this portion, it is divided near the end and near the beginning of A A new voice pattern is added by adding the vicinity of the end of B, and C is added near the beginning of B.
It is also possible to create and register a new voice pattern by adding the vicinity of the end of, and to register all or part of the created pattern as a new standard pattern by repeating this process.
【0029】図7は分割部12における上記のような処
理の流れを示すフローチャートである。図7の処理で
は、時間をi,周波数をjとし、音声の標準パターンを
a(i,j)とし、また標準パターンの全長をnで表わ
すとき、先ず、時間カウンタiを“1”に初期設定する
(ステップS21)。次いで、時間iにおけるパターン
全体に対して、次式のように、パターンの時間変化率d
iを求める(ステップS22)。FIG. 7 is a flow chart showing the flow of the above processing in the dividing unit 12. In the process of FIG. 7, when the time is i, the frequency is j, the standard voice pattern is a (i, j), and the total length of the standard pattern is n, first, the time counter i is initialized to "1". It is set (step S21). Then, with respect to the entire pattern at time i, the time change rate d of the pattern is expressed by the following equation.
i is calculated (step S22).
【0030】[0030]
【数1】 [Equation 1]
【0031】iを“1”ずつ歩進して(ステップS2
4)、上記処理を繰り返し、各iごとにパターンの時間
変化率diを求め、カウンタiが(n−1)になったと
きに(ステップS23)、パターン全長nの中から、d
iが最大となるようなiを求めて、これをinとする(ス
テップS25)。次いで、標準パターンを1〜inとi
n+1からnまでとに2分割する(ステップS25)。こ
のような分割によって、パターンが音韻の中央で不連続
に連結されるような事態を防止することができる。I is incremented by "1" (step S2
4) The above process is repeated to obtain the time change rate d i of the pattern for each i, and when the counter i reaches (n-1) (step S23), d is selected from the total pattern length n
The i that maximizes the i is obtained and is set as i n (step S25). Then, the reference pattern 1 to i n a i
It is divided into n + 1 to n (step S25). Such division can prevent a situation in which patterns are discontinuously connected at the center of the phoneme.
【0032】あるいは、上記処理のかわりに、複数の音
声の標準パターンA,B,Cのそれぞれについて、時間
軸方向に音声エネルギーの極小部分を検出し、この部分
で末尾近傍に分割し、Aの冒頭近傍にBの末尾近傍を継
ぎ足して新たな音声パターンとして登録し、Bの冒頭近
傍にCの末尾近傍を継ぎ足すことで新たな音声パターン
を作成して登録し、この過程を繰り返すことにより作成
されたパターンの全てまたは一部を新たな標準パターン
として登録するようにすることもできる。Alternatively, instead of the above processing, a minimum portion of the voice energy is detected in the time axis direction for each of a plurality of voice standard patterns A, B, and C, and this portion is divided into the vicinity of the tail end. Create a new voice pattern by adding the vicinity of the end of B to the beginning and registering it as a new voice pattern, and adding the vicinity of the end of C to the beginning of B to create and register a new voice pattern and repeating this process. It is also possible to register all or part of the created pattern as a new standard pattern.
【0033】図8は分割部12における上記のような処
理の流れを示すフローチャートである。この処理では、
先ず、時間カウンタiを“1”に初期設定し(ステップ
S31)、次いで、a(i,j)をjについて合計する
ことによって、時間iでのエネルギーpiを求める(ス
テップS32)。しかる後、pi-1とpiとの差を求め、
これをDiとする(ステップS33)。次いで、iを
“1”ずつ歩進して(ステップS35)、上記処理を繰
り返し、各iごとにエネルギーpi並びに差Diを求め、
カウンタiがnとなったときに(ステップS34)、D
iに基づき、エネルギーの極小部分の検出を行なう。す
なわち、この検出処理では、piの極小値が、Di-1とD
iの符号が異なり、かつDiが正であることにより求めら
れるので、この点を求めてi’とすることによりなされ
る(ステップS36)。しかる後、パターンを1〜i’
とi’〜nとに分割する(ステップS37)。このよう
な分割によって、図7の処理例と同様に、パターンが音
韻の中央で不連続に連結されるような事態を防止するこ
とができる。なお、1つの音声中でこのようなi’が複
数個ある時は、その時のDiが小さい方を選んだり、あ
るいは先頭に近い(または遠い)ものを選んだり、さら
には中央に近いものを選ぶなどの処置を行なうことがで
きる。FIG. 8 is a flow chart showing the flow of the above processing in the dividing section 12. In this process,
First, the time counter i is initialized to "1" (step S31), and then the energy p i at the time i is obtained by summing a (i, j) for j (step S32). Then, the difference between p i-1 and p i is calculated,
This is designated as D i (step S33). Next, i is incremented by "1" (step S35), the above process is repeated, and the energy p i and the difference D i are obtained for each i,
When the counter i reaches n (step S34), D
Based on i , the minimum energy is detected. That is, in this detection process, the minimum value of p i is D i−1 and D i.
Since i is obtained when the sign is different and D i is positive, this point is obtained and set as i ′ (step S36). After that, the patterns 1 to i '
And i ′ to n (step S37). By such division, as in the processing example of FIG. 7, it is possible to prevent a situation in which the patterns are discontinuously connected at the center of the phoneme. When there are a plurality of i's in one voice, the one with a smaller D i at that time is selected, the one near (or far from) the head is selected, and the one near the center is selected. You can take actions such as choosing.
【0034】上述した各例のように、標準パターンを時
間軸方向に分割し、これらを組合せるという処理は、元
の音声に類似した別の言葉を合成することを直感的に把
握することができ、非常に考え易いが、パターン間の類
似ということを考えれば、時間軸方向に分割するだけで
なく、周波数軸方向に分割しても同様の効果を得ること
ができる。従って、分割部12としては、例えば、図9
に示すように、複数の音声の特徴パターンA,B,Cの
それぞれを、周波数軸に対して高域近傍と低域近傍とに
2分して、それぞれA1,A2,B1,B2,C1,C
2とし、図10に示すように、Aの低域近傍A1にBの
高域近傍B2を継ぎ足して新たな音声パターンDとして
登録し、Bの低域近傍B1にCの高域近傍C2を継ぎ足
すことで新たな音声パターンEを作成して登録するとい
うような過程を繰り返して作成されたパターンD,E,
…の全て,または一部を新たな標準パターンとして登録
するようにすることもできる。As in each of the above-described examples, the process of dividing the standard pattern in the time axis direction and combining these can be intuitively understood to synthesize another word similar to the original voice. Although it is possible and very easy to think, considering that the patterns are similar to each other, the same effect can be obtained not only by dividing in the time axis direction but also by dividing in the frequency axis direction. Therefore, as the dividing unit 12, for example, FIG.
As shown in FIG. 3, each of the plurality of voice characteristic patterns A, B, and C is divided into a high-frequency neighborhood and a low-frequency neighborhood with respect to the frequency axis, and A1, A2, B1, B2, C1, respectively. C
As shown in FIG. 10, the high-frequency neighborhood A1 of A is added to the high-frequency neighborhood B2 of B to be registered as a new voice pattern D, and the high-frequency neighborhood C2 of C is connected to the low-frequency neighborhood B1 of B. Patterns D, E, created by repeating the process of creating and registering a new voice pattern E by adding
It is also possible to register all or part of ... as a new standard pattern.
【0035】あるいは、分割部12における分割処理で
はなく、複数の音声の標準パターンA,B,Cのそれぞ
れを、周波数軸に対して一定量だけ高域へシフト(ずら
す)ような操作を施して作成された新たな音声パターン
の全て,または一部を新たな標準パターンとして登録す
るようにすることもできる。Alternatively, instead of performing the division processing in the division unit 12, each of the plurality of standard patterns A, B and C of the voices is operated by shifting (shifting) to a high frequency range by a certain amount with respect to the frequency axis. It is also possible to register all or part of the created new voice pattern as a new standard pattern.
【0036】図11は1つの標準パターンa(i,j)
を、周波数軸に対して一定量だけ高域へずらせる操作を
施して、新たな標準パターンを生成する処理の流れを示
すフローチャートである。この処理では、先ず、時間カ
ウンタiを“1”に初期設定する(ステップS41)。
次いで、周波数カウンタjを“1”に設定する(ステッ
プS42)。しかる後、a(i,j)を周波数方向へ例
えば2サンプル分シフトし、b(i,j+2)とする
(ステップS43)。jの範囲を1からJとするとき、
jを“1”ずつ歩進して(ステップS45)、j=1〜
J−2の範囲でこのシフト処理を繰り返し行ない、jが
J−2となったときに(ステップS44)、シフトの結
果、作成されたb(i,j)をa(i,j)に再び写像
する。すなわち、先ず、jを“1”に初期設定し(ステ
ップS46)、しかる後、b(i,j)をa(i,j)
に写像する(ステップS47)。jを“1”ずつ歩進し
て(ステップS49)、j=3〜Jの範囲でこの写像処
理を繰り返し行ない、jがJ−2となったときに(ステ
ップS48)、この写像処理を終了する。この写像処理
の結果、a(i,j)のj=1とj=2の部分は元のま
まのデータが残り、j=3〜Jの部分は、元の1からJ
−2までのデータがシフトされたことになる。このよう
にして、1つの時間iにおいて高域へのデータシフトを
行なった後、iを“1”ずつ歩進して(ステップS5
1)、上記と同様の処理を繰り返し行ない、iがn(全
パターン長)となったときに(ステップS50)、1つ
の標準パターンに対するシフト処理を完了する。なお、
この例では、周波数方向に2サンプルずらしたが、シフ
ト量は2サンプルに限らず任意に設定でき、サンプリン
グの周波数に応じて、これを変化させることができる。FIG. 11 shows one standard pattern a (i, j).
Is a flow chart showing the flow of processing for generating a new standard pattern by performing an operation for shifting the frequency axis by a certain amount to the high frequency band. In this process, first, the time counter i is initially set to "1" (step S41).
Next, the frequency counter j is set to "1" (step S42). Then, a (i, j) is shifted in the frequency direction by, for example, 2 samples to be b (i, j + 2) (step S43). When the range of j is 1 to J,
Stepping j by "1" (step S45), j = 1 to
This shift process is repeated in the range of J-2, and when j becomes J-2 (step S44), the b (i, j) created as a result of the shift is again converted into a (i, j). Map. That is, first, j is initially set to "1" (step S46), and then b (i, j) is set to a (i, j).
(Step S47). Stepping j by "1" (step S49), this mapping process is repeated within the range of j = 3 to J, and when j becomes J-2 (step S48), this mapping process ends. To do. As a result of this mapping processing, the original data remains in the part of a (i, j) where j = 1 and j = 2, and the part where j = 3 to J is from the original 1 to J.
The data up to -2 has been shifted. In this way, after the data shift to the high frequency band at one time i, i is incremented by "1" (step S5).
1) The same processing as described above is repeated, and when i becomes n (total pattern length) (step S50), the shift processing for one standard pattern is completed. In addition,
In this example, two samples are shifted in the frequency direction, but the shift amount is not limited to two samples and can be set arbitrarily, and this can be changed according to the sampling frequency.
【0037】また、上記例では、高域方向へシフトさせ
たが、低域方向へシフトさせることもできる。すなわ
ち、複数の音声の標準パターンA,B,Cのそれぞれ
を、周波数軸に対して一定量だけ低域へシフト(ずら
す)ような操作を施して作成された新たな音声パターン
の全て,または一部を新たな標準パターンとして登録す
るようにすることもできる。この処理は図11の処理と
基本的にほぼ同じになされ、図11のステップS43に
おいて、a(i,j+2)を周波数方向へ2サンプル分
シフトして、b(i,j)とする処理だけが相違する。
このようにして、1つの標準パターンに対して周波数軸
に対して2サンプル分だけ低域へシフトすると、a
(i,j)のj=J−1とJには元のままのデータが残
り、j=1〜J−2には、元のj=3からJのデータが
シフトされたことになる。Further, in the above example, the shift is performed in the high band direction, but it may be shifted in the low band direction. That is, all or one of the new voice patterns created by performing an operation of shifting (shifting) the standard patterns A, B, and C of a plurality of voices to the low frequency range by a certain amount with respect to the frequency axis. It is also possible to register a copy as a new standard pattern. This process is basically the same as the process of FIG. 11, and in step S43 of FIG. 11, only a process of shifting a (i, j + 2) by 2 samples in the frequency direction to obtain b (i, j). Is different.
In this way, when one standard pattern is shifted to the low frequency range by 2 samples with respect to the frequency axis, a
The original data remains in j = J-1 and J of (i, j), and the data of J is shifted from j = 3 in j = 1 to J-2.
【0038】また、高域へのシフトと低域へのシフトと
の両方を施すこともできる。すなわち、複数の音声の標
準パターンA,B,Cのそれぞれを、周波数軸に対して
一定量だけ高域へずらせるような操作を施して作成され
た新たな音声パターンの全て,または一部を標準パター
ンとして登録し、さらに、元の標準パターンA,B,C
のそれぞれを周波数軸に対して一定量だけずらせるよう
な操作を施して作成された新たな音声パターンの全て,
または一部をも標準パターンとして登録するようにする
こともできる。It is also possible to perform both a shift to the high frequency band and a shift to the low frequency band. That is, all or a part of a new voice pattern created by performing an operation of shifting each of a plurality of voice standard patterns A, B, and C to a high frequency range by a certain amount with respect to the frequency axis. Registered as a standard pattern, and further, the original standard patterns A, B, C
All of the new voice patterns created by performing an operation that shifts each of the
Alternatively, a part of them may be registered as a standard pattern.
【0039】上記のような標準パターンの高域,低域へ
のシフトは、異なる発声者のパターンを作るような操作
に匹敵する。従って、高域または低域へのシフトを行な
って作成された新たな標準パターンを辞書5に登録する
ことによって、登録者(特定話者)の声よりも高い声に
対してまたは低い声に対して誤認識がなされないよう防
御することができ、さらに、この新たな標準パターンに
所定の単語を対応させれば、辞書5を不特定話者用のも
のにすることができる。The above-described shift of the standard pattern to the high frequency range and the low frequency range is comparable to an operation for creating patterns of different utterers. Therefore, by registering a new standard pattern created by shifting to the high frequency region or the low frequency region in the dictionary 5, the voices higher or lower than the voice of the registrant (specific speaker) are registered. It is possible to prevent erroneous recognition from being made, and further, by making a predetermined word correspond to this new standard pattern, the dictionary 5 can be made for an unspecified speaker.
【0040】また、高域へのシフトと低域へのシフトと
の両方を施すことによって、登録者(特定話者)の声よ
りも高い声に対しても、また、低い声に対しても防御す
ることができ、さらに、この新たな標準パターンに所定
の単語を対応させれば、辞書5を不特定話者用のものに
容易にすることができる。Further, by performing both the shift to the high frequency range and the shift to the low frequency range, both a voice higher than the voice of the registrant (specific speaker) and a voice lower than the voice of the registrant (specific speaker) can be obtained. It is possible to protect, and further, by making a predetermined word correspond to this new standard pattern, it is possible to make the dictionary 5 easy for an unspecified speaker.
【0041】上述した各例では、標準パターンを時間軸
方向または周波数軸方向に分割して、違うもの同士を連
結して新たな標準パターンを作り、これを併せて辞書5
に登録するようにしており、これにより、例えば「加
藤」,「佐藤」のような類似語間での誤認識を防止でき
るが、例えば、登録すべき言葉の中に、「佐藤」,「佐
治」のように冒頭が同じ言葉や、「坂本」,「橋本」の
ように末尾が同じ言葉であるような場合、「佐藤」,
「佐治」の例において冒頭,末尾に分割して組合せる
と、組合せた結果にも、元の「佐藤」,「佐治」が生成
されることがある。「坂本」,「橋本」の場合にも同様
である。このような元の単語と同一のものが新たな標準
パターンとして登録されてしまうと、本来の正しい認識
ができなくなる。従って、このような新たな標準パター
ンは辞書5に登録されるべきではない。In each of the above-described examples, the standard pattern is divided in the time axis direction or the frequency axis direction, and different patterns are connected to each other to create a new standard pattern.
It is possible to prevent erroneous recognition between similar words such as "Kato" and "Sato". However, for example, in the words to be registered, "Sato" When the beginning of the word is the same, or when the ending is the same, such as “Sakamoto” and “Hashimoto”, “Sato”,
In the example of "Saji", if the beginning and the end are divided and combined, the original "Sato" and "Saji" may be generated in the combined result. The same applies to "Sakamoto" and "Hashimoto". If the same original word as this is registered as a new standard pattern, the original correct recognition cannot be performed. Therefore, such a new standard pattern should not be registered in the dictionary 5.
【0042】そこで、このような問題を回避するため、
さらに、辞書5に登録されている元の標準パターン,新
たな標準パターンの全ての標準パターン間の類似性を求
め、類似性が決められた値よりも大きな標準パターンに
ついてはこれを消去するような処理を行なうこともでき
る。Therefore, in order to avoid such a problem,
Further, the similarity between all the standard patterns of the original standard pattern and the new standard pattern registered in the dictionary 5 is calculated, and the standard pattern having the similarity larger than the determined value is deleted. Processing can also be performed.
【0043】図12はこのような処理の一例を示すフロ
ーチャートである。なお、この処理例では、元の標準パ
ターンに基づき新たな標準パターンを作成し、この新た
な標準パターンが元の標準パターンと併せて辞書5に登
録されたとき、先ず、辞書5に登録されている標準パタ
ーン(元の標準パターンの数W,並びに新たな標準パタ
ーンの数(L−W))の全て(合計L個の標準パター
ン)に、T1,T2,…,TLのように番号付けをする。
この際、例えば、元の標準パターンから順番に番号付け
をし、(T1〜TW)元の標準パターンの番号に続いて新
たな標準パターンの番号付けを行なう(TW+1〜TL)。
また、2つの標準パターン間の類似度を求めるために、
便宜上、2つの標準パターンのうちの一方の標準パター
ンの番号をrとし、他方の標準パターンの番号をkとす
る。FIG. 12 is a flow chart showing an example of such processing. In this processing example, a new standard pattern is created based on the original standard pattern, and when this new standard pattern is registered in the dictionary 5 together with the original standard pattern, it is first registered in the dictionary 5. All of the existing standard patterns (the number of original standard patterns W and the number of new standard patterns (L-W)) (total of L standard patterns) are represented by T 1 , T 2 , ..., T L. Number them.
In this case, for example, the numbered sequentially from the original standard pattern, performs numbering of new reference pattern following the number (T 1 ~T W) original reference pattern (T W + 1 ~T L ).
Moreover, in order to obtain the similarity between two standard patterns,
For convenience, let us say that the number of one of the two standard patterns is r and the number of the other standard pattern is k.
【0044】全ての標準パターンに対して番号付けを行
なった後、rを“1”に初期設定する(ステップS6
1)。次いで、kを“W+1”に設定する(ステップS
62)。しかる後、標準パターンTrと標準パターンTk
との間の類似度Srkを求める(ステップS63)。次い
で、rがkよりも小さいか否かを判断する(ステップS
64)。なお、この判断は、rとkとが同じである場合
(2つの標準パターンTr,Tkが同一のものである場
合)、およびrとkとが逆の場合(Tr,Tkに対し
Tk,Trの場合)を除外するためになされるものであ
る。すなわち、図13に示すようなテーブルにおいて、
新たな標準パタ−ンTW+1〜TLに関し、対角線DG上を
も含めてこの対角線DGよりも下の部分については以後
の処理を行なわせないためになされるものである。この
ような判断によりrがkよりも小さいとき、すなわち図
13のテーブルで対角線DGよりも上の部分の関係にあ
るとき、ステップS63で求めた類似度Srkが所定の閾
値S’よりも小さいか否かを判断する(ステップS6
5)。SrkがS’よりも小さいときには、2つの標準パ
ターンTr,Tk間の類似度が小さいと判断し、消去処理
を行なわない。これに対し、SrkがS’よりも小さくな
いときには、2つの標準パターンTr,Tk間の類似度が
大きいと判断し、一方の標準パターンTkを辞書5から
消去する(ステップS66)。After all standard patterns are numbered, r is initialized to "1" (step S6).
1). Then, k is set to "W + 1" (step S
62). After that, the standard pattern T r and the standard pattern T k
Then, the similarity S rk between and is obtained (step S63). Then, it is determined whether r is smaller than k (step S
64). This judgment is made when r and k are the same (when the two standard patterns T r and T k are the same) and when r and k are opposite (in T r and T k) . (In the case of T k and T r ), it is done for the purpose of exclusion. That is, in the table as shown in FIG.
With respect to the new standard patterns T W + 1 to T L , this is done so as not to perform subsequent processing on the portion below the diagonal line DG, including on the diagonal line DG. When r is smaller than k by such a judgment, that is, when there is a relationship above the diagonal line DG in the table of FIG. 13, the similarity S rk obtained in step S63 is smaller than a predetermined threshold S ′. It is determined whether or not (step S6)
5). When S rk is smaller than S ′, it is determined that the similarity between the two standard patterns T r and T k is small, and the erasing process is not performed. On the other hand, when S rk is not smaller than S ′, it is determined that the similarity between the two standard patterns T r and T k is large, and one standard pattern T k is deleted from the dictionary 5 (step S66). .
【0045】このようにして、ある2つの標準パターン
Tr,Tk間の類似度Srkを求め、これに基づき所定の処
理を行なった後、kを“1”ずつ歩進して(ステップS
68)、同様の処理を繰り返し行ない、kがLとなった
ときに(ステップS67)、rを“1”ずつ歩進して
(ステップS70)、同様の処理を繰り返し行ない、r
がL−1となったときに(ステップS69)、処理を終
了する。In this way, the degree of similarity S rk between two certain standard patterns T r and T k is obtained, and a predetermined process is performed based on the degree of similarity, and then k is incremented by "1" (step S
68), the same processing is repeated. When k becomes L (step S67), r is incremented by "1" (step S70) and the same processing is repeated, r
When L becomes L-1 (step S69), the process ends.
【0046】上記のような処理は、換言すれば、kをW
+1からLまで順番に変化させ、またrを1からL−1
まで順番に変化させて、図13に示すようなTkとTrと
の間の類似度表を作成し、この中で、大きな類似度とな
った2つの標準パターンのうちの一方のパターンTkを
消去することにある。また、このような処理(具体的に
はステップS64の処理)において、新たな標準パター
ンの番号が元の標準パターンの番号よりも常に大きいの
で、消去対象となるのは、新たな標準パターンのみであ
り、元のパターンについてはこれを消去させずに残すこ
とができる。なお、この時の閾値S’の決め方について
は特に限定されないが、例えば、これを次のように決め
ることができる。すなわち、この標準パターンを使って
正しい認識が行なわれた時に、正解となる標準パターン
が獲得する類似度の90%程度に閾値S’を設定するこ
とができる。この方法によれば、まぎらわしい言葉の中
で、混乱を引き起こすような標準パターンを取り除くこ
とができる。In other words, the above-mentioned processing is performed by setting k to W.
Change from +1 to L in order, and r from 1 to L-1
13 to create a similarity table between T k and T r as shown in FIG. 13, in which one of the two standard patterns T having a large similarity is used. to eliminate k . Further, in such a process (specifically, the process of step S64), since the number of the new standard pattern is always larger than the number of the original standard pattern, only the new standard pattern is to be erased. Yes, the original pattern can be left without being erased. The method of determining the threshold value S ′ at this time is not particularly limited, but can be determined as follows, for example. That is, when correct recognition is performed using this standard pattern, the threshold value S ′ can be set to about 90% of the similarity acquired by the standard pattern that is the correct answer. This method removes confusional standard patterns in misleading words.
【0047】また、本発明の本来の目的は、類似した別
の言葉を区別するために、類似の標準パターンを準備す
ることにある。従って、全く類似性のない標準パターン
や、分割の都合で大した類似性がない標準パターンが生
成された場合には、これらを辞書5に登録する必要がな
く、また、これらを登録しないようにすることで、辞書
5の節約を行なうことができる。すなわち、連結により
作成された新たな標準パターンと元の標準パターンとの
間の類似性を調べ、元の標準パターン全てとの類似性が
決められた値よりも小さな標準パターンについてはこれ
を消去するようにすることができる。なお、これと同様
の方法は、特開昭61−292696号公報にも示され
ている。この公報に開示の方法は、予め、1つの音声に
対して複数の標準パターンが登録されている時、複数の
パターン間で類似度を求め、似ているものは一方を消去
するものであるが、この方法では、その際にどちらのパ
ターンを消去してもよく、この方法を本発明に適用する
と、元のパターンが消去されてしまうことがあり、この
結果正しい認識ができなくなってしまうという問題が生
じる。従って、本発明への適用においては、類似度が所
定閾値よりも低い標準パターンを消去するに際して、元
のパターンについてはこれが消去されず、新たなパター
ンのみが消去されるような処理がなされる必要がある。The original purpose of the present invention is to prepare a similar standard pattern for distinguishing another similar word. Therefore, when a standard pattern having no similarity or a standard pattern having little similarity due to division is generated, it is not necessary to register these in the dictionary 5, and these should not be registered. By doing so, the dictionary 5 can be saved. That is, the similarity between the new standard pattern created by concatenation and the original standard pattern is checked, and the standard pattern whose similarity to all the original standard patterns is smaller than a predetermined value is deleted. You can A method similar to this is also disclosed in Japanese Patent Laid-Open No. 61-292696. According to the method disclosed in this publication, when a plurality of standard patterns are registered in advance for one voice, the degree of similarity between the plurality of patterns is obtained, and one of the similar patterns is deleted. In this method, either pattern may be erased at that time, and if this method is applied to the present invention, the original pattern may be erased, and as a result, correct recognition cannot be performed. Occurs. Therefore, in the application to the present invention, when erasing a standard pattern whose degree of similarity is lower than a predetermined threshold value, it is necessary to perform processing such that the original pattern is not erased and only a new pattern is erased. There is.
【0048】図14はこのような処理の一例を示すフロ
ーチャートである。この処理では、元の標準パターンの
数をL1、連結操作等で作った新たな標準パターンの数
をL2とし、元の標準パターンの各々に、1〜L1の番号
付けをし、また、新たな特徴パターンの各々にL1+1
〜L1+L2の番号付けをしておく。また、変数kが元の
標準パターンを示し、rが操作で作り出した標準パター
ンを示す。FIG. 14 is a flow chart showing an example of such processing. In this process, the number of original standard pattern L 1, the number of new reference pattern made by the connecting operation or the like and L 2, each of the original standard pattern, and the numbering of 1 to L 1, also , L 1 +1 for each new feature pattern
Number them to L 1 + L 2 . The variable k represents the original standard pattern, and r represents the standard pattern created by the operation.
【0049】先ず、rにL+1を初期設定する(ステッ
プS81)。次いで、kに“1”を設定し(ステップS
82)、類似度の最大値mxを“0”に初期設定する
(ステップS83)。しかる後、番号kの元の標準パタ
ーンTkと番号rの新たな標準パターンTrとの間の類似
度Srkを求め(ステップS84)、この類似度Srkがm
xよりも大きいか否かを判断する(ステップS85)。
mxよりも大きい場合には、mxにこの類似度Srkの値
を設定する(ステップS86)。次いで、kを“1”ず
つ歩進して(ステップS88)、同様の処理を繰り返し
行ない、kがL1となったとき(ステップS87)、こ
のときのmxが所定閾値S’よりも大きいか否かを判断
する(ステップS89)。すなわち、上記の処理は、番
号kが1〜L1までのL1個の元の標準パターンTkのそ
れぞれと番号rの新たな標準パターンTrとの間のL1個
の類似度を求め、これらの類似度のうちで最も大きな類
似度が所定の閾値S’よりも大きいか否かを判断するも
のである。この結果、最も大きな類似度mxが所定閾値
S’よりも大きくないときには、番号rの新たな標準パ
ターンTrは、L1個の元の標準パターンTkのいずれと
も類似度が小さいので、これを消去する(ステップS9
0)。次いで、rを“1”ずつ歩進して(ステップS9
2)、L2個の新たな標準パターンTrのそれぞれに対し
て上記と同様の処理を繰り返し行ない、rがL1+L2と
なったときに(ステップS91)、処理を終了する。こ
のようにして、L2個の新たな標準パターンのうちで、
L1個の元の標準パターンのいずれとも類似度の小さい
ものを消去することができる。なお、この処理におい
て、閾値S’の決め方としては、これを例えば全く類似
音を含まない言葉同士を比較した時の類似度と同程度に
選んでおけば良い。この結果、全ての登録すべき言葉と
類似性が乏しいようなもの、すなわち、標準パターンの
ノイズとなるようなものを消去することができる。従っ
て、辞書5の容量を節約して本来の効果を得ることがで
きる。First, L + 1 is initialized to r (step S81). Next, k is set to "1" (step S
82), the maximum value mx of the similarity is initialized to "0" (step S83). Then, the similarity S rk between the original standard pattern T k with the number k and the new standard pattern T r with the number r is calculated (step S84), and this similarity S rk is m.
It is determined whether it is larger than x (step S85).
If it is larger than mx, the value of this similarity S rk is set in mx (step S86). Then, k is incremented by "1" (step S88), the same process is repeated, and when k becomes L 1 (step S87), whether mx at this time is larger than a predetermined threshold value S '. It is determined whether or not (step S89). That is, the above-described processing obtains L 1 similarity between each of the L 1 original standard patterns T k with numbers k 1 to L 1 and the new standard pattern T r with number r. Of these similarities, it is determined whether or not the largest similarity is larger than a predetermined threshold value S ′. As a result, when the largest similarity mx is not larger than the predetermined threshold value S ′, the new standard pattern T r with the number r has a small similarity to any of the L 1 original standard patterns T k. Is erased (step S9
0). Then, r is incremented by "1" (step S9
2), the same processing as described above is repeated for each of the L 2 new standard patterns T r , and when r becomes L 1 + L 2 (step S91), the processing ends. In this way, among the L 2 new standard patterns,
Any of the L 1 original standard patterns having a small degree of similarity can be erased. In this process, the threshold value S ′ may be determined in the same degree as the degree of similarity when words having no similar sounds are compared with each other. As a result, it is possible to erase a word having a low similarity to all the words to be registered, that is, a word that becomes noise of the standard pattern. Therefore, the capacity of the dictionary 5 can be saved and the original effect can be obtained.
【0050】また、上記例では1〜L1を元の標準パタ
ーンとしたが、これとは逆に、1〜L2を新たな標準パ
ターンとし、L2+1からL1+L2を元の標準パターン
とし、1〜L2を消去対象として、同様の処理を行なう
こともできる。[0050] Further, in the above example was with the original standard pattern 1 to L 1, on the contrary, a 1 to L 2 as a new reference pattern, based on the standard L 1 + L 2 from L 2 +1 The same process can be performed by setting a pattern and erasing 1 to L 2 .
【0051】[0051]
【発明の効果】以上に説明したように、請求項1記載の
発明によれば、複数の音声の特徴パターンのそれぞれを
複数の部分に分割し、分割手段により分割された複数の
特徴パターンの各部分を組合せ連結して新たな音声の特
徴パターンを作成するようにしているので、他の言葉に
対する音声の特徴パターンを簡単に作成することができ
る。As described above, according to the invention of claim 1, each of a plurality of voice characteristic patterns is divided into a plurality of parts, and each of the plurality of characteristic patterns divided by the dividing means is divided. Since the parts are combined and connected to create a new voice characteristic pattern, it is possible to easily create a voice characteristic pattern for another word.
【0052】また、請求項2記載の発明によれば、請求
項1記載の音声パターン作成装置を用いた標準パターン
登録装置であって、複数の音声の特徴パターンを標準パ
ターンとしてそのままの形で辞書に登録する登録手段を
さらに有し、上記分割手段は、上記標準パターンを複数
の部分に分割し、上記パターン作成手段は、分割された
複数の標準パターンの各部分を組合せ連結して新たな特
徴パターンを生成し、登録手段によりそのままの形で登
録されている標準パターンと併せて、該新たな特徴パタ
ーンを新たな標準パターンとして辞書に登録するので、
任意の言葉の標準パターンを登録して音声認識させる場
合、誤認識が生ずるのを著しく低減することができる。According to a second aspect of the present invention, there is provided a standard pattern registration apparatus using the voice pattern creating apparatus according to the first aspect, wherein the dictionary is a feature pattern of a plurality of voices as a standard pattern as it is. The dividing means divides the standard pattern into a plurality of parts, and the pattern creating means combines and connects the respective parts of the plurality of divided standard patterns to create a new feature. Since the pattern is generated and the new feature pattern is registered in the dictionary as a new standard pattern together with the standard pattern registered in the same form by the registration means,
When a standard pattern of an arbitrary word is registered for voice recognition, it is possible to significantly reduce the occurrence of erroneous recognition.
【0053】また、請求項3乃至請求項5記載の発明に
よれば、パターン作成手段は、複数の標準パターンのそ
れぞれを複数の部分に分割し、これらを組合せ連結して
新たな標準パターンを作成する際に、分割前の標準パタ
ーンにおいて先頭にあった部分は連結後の新たな標準パ
ターンにおいても先頭に位置させ、また、分割前の標準
パターンにおいて末尾にあった部分は連結後の新たな標
準パターンにおいても末尾に位置させるように組合せる
ので、常識的でないパターンが生成,登録されるのを防
止することができる。According to the third to fifth aspects of the invention, the pattern creating means divides each of the plurality of standard patterns into a plurality of parts, and combines and connects them to create a new standard pattern. In this case, the part at the beginning of the standard pattern before division is positioned at the beginning of the new standard pattern after concatenation, and the part at the end of the standard pattern before division is the new standard after concatenation. Since the patterns are also combined so that they are positioned at the end, it is possible to prevent the generation and registration of uncommon patterns.
【0054】また、請求項6乃至請求項8記載の発明に
よれば、分割手段は、複数の音声の標準パターンのそれ
ぞれを時間軸方向に分割し、パターン作成手段は、時間
軸方向に分割された各標準パターンの各部分を所定の規
則に従って組合せ連結して新たな音声パターンとして作
成し、この過程を繰り返すことにより作成された音声パ
ターンの全て,または一部を標準パターンとして登録す
るようになっているので、他の言葉に対する音声の特徴
パターンを容易に作成し、登録することができる。According to the sixth to eighth aspects of the invention, the dividing means divides each of the plurality of standard voice patterns in the time axis direction, and the pattern forming means divides in the time axis direction. Each part of each standard pattern is combined and connected according to a predetermined rule to create a new voice pattern, and all or part of the created voice pattern is registered as a standard pattern by repeating this process. Therefore, it is possible to easily create and register a voice characteristic pattern for another word.
【0055】また、請求項9記載の発明によれば、複数
の音声の特徴パターンのそれぞれを、周波数軸方向に一
定量だけ高域へおよび/または低域へずらし、新たな音
声の特徴パターンを作成するので、元の標準パターンと
同じ言葉であるが、話者が相違する場合に対応した音声
の標準パターンを簡単に作成することができる。According to the ninth aspect of the present invention, each of a plurality of voice characteristic patterns is shifted to the high frequency band and / or the low frequency band by a certain amount in the frequency axis direction to obtain a new voice characteristic pattern. Since it is created, it is the same word as the original standard pattern, but it is possible to easily create a standard pattern of voice corresponding to the case where the speakers are different.
【0056】また、請求項10記載の発明によれば、請
求項9記載の音声パターン作成装置を用いた標準パター
ン登録装置であって、複数の音声の特徴パターンを標準
パターンとしてそのままの形で辞書に登録する登録手段
をさらに有し、パターン作成手段は、複数の標準パター
ンのそれぞれを周波数軸方向に一定量だけ高域へおよび
/または低域へずらし、新たな音声パターンを作成した
とき、登録手段によりそのままの形で登録されている標
準パターンと併せて、該新たな音声パターンの全て,ま
たは一部を標準パターンとして辞書に登録するようにな
っているので、不特定話者用の辞書を容易に作成するこ
とができ、不特定話者用の音声認識装置に適用すること
ができる。According to a tenth aspect of the present invention, there is provided a standard pattern registration apparatus using the voice pattern creating apparatus according to the ninth aspect, wherein the characteristic patterns of a plurality of voices are used as standard patterns in the dictionary as they are. Further, the pattern creating means shifts each of the plurality of standard patterns to a high frequency band and / or a low frequency band in the frequency axis direction by a certain amount to register a new voice pattern. The whole or a part of the new voice pattern is registered as a standard pattern in the dictionary together with the standard pattern registered as it is by the means. It can be easily created and can be applied to a voice recognition device for an unspecified speaker.
【0057】また、請求項11記載の発明によれば、そ
のままの形で辞書に登録された元の標準パターンと前記
パターン作成手段により作成された新たな標準パターン
との間の類似性を求め、元の標準パターンとの類似性が
決められた値よりも大きな新たな標準パターンについて
は、これを辞書から消去するようになっているので、音
声認識においてまぎらわしい言葉の中で、混乱を生じさ
せるような標準パターンを取り除くことができる。According to the eleventh aspect of the invention, the similarity between the original standard pattern registered in the dictionary as it is and the new standard pattern created by the pattern creating means is calculated, For new standard patterns whose similarity to the original standard pattern is larger than a predetermined value, it is deleted from the dictionary, so that it may cause confusion in confusing words in speech recognition. You can remove various standard patterns.
【0058】また、請求項12記載の発明によれば、そ
のままの形で登録された元の標準パターンと新たな標準
パターンとの間の類似性を求め、元の標準パターンとの
類似性が決められた値よりも小さな新たな標準パターン
については、これを辞書から消去するようになっている
ので、音声認識に不要な標準パターンを取り除くことが
できる。According to the twelfth aspect of the invention, the similarity between the original standard pattern registered as it is and the new standard pattern is obtained, and the similarity with the original standard pattern is determined. The new standard pattern smaller than the given value is deleted from the dictionary, so that the standard pattern unnecessary for voice recognition can be removed.
【0059】また、請求項13記載の発明によれば、登
録手段は、複数の音声の標準パターンをそのままの形で
登録するに際し、該標準パターンを所定の言葉と対応付
けして辞書に登録する一方、パターン作成手段は、新た
な標準パターンを言葉との対応付けを行なわずに辞書に
登録するようになっているので、元の標準パターンに対
応した音声と多少類似しているが、これとは異なる音声
が入力したとき、この音声が新たな標準パターンに最も
類似する場合に、これをリジェクトすることができ、誤
認識がなされるのを有効に防止することができる。According to the thirteenth aspect of the present invention, the registration means, when registering the standard patterns of a plurality of voices as they are, registers the standard patterns in the dictionary in association with predetermined words. On the other hand, the pattern creating means is designed to register a new standard pattern in the dictionary without associating it with a word, and thus is somewhat similar to the voice corresponding to the original standard pattern. When a different voice is input, the voice can be rejected when the voice is most similar to the new standard pattern, and it is possible to effectively prevent erroneous recognition.
【0060】また、請求項14記載の発明によれば、登
録手段は、複数の音声の標準パターンをそのままの形で
登録するに際し、該標準パターンを所定の言葉と対応付
けして辞書に登録し、また、パターン作成手段も、新た
な標準パターンを所定の言葉と対応付けして辞書に登録
するようになっているので、元の標準パターンに対応し
た音声と多少類似しているが、これとは異なる音声が入
力したとき、この音声が新たな標準パターンに最も類似
する場合には、これをも正しく認識させることができ
る。According to the fourteenth aspect of the invention, the registration means, when registering the standard patterns of a plurality of voices as they are, registers the standard patterns in the dictionary in association with predetermined words. Also, the pattern creating means is also adapted to register a new standard pattern in a dictionary in association with a predetermined word, so that it is somewhat similar to the voice corresponding to the original standard pattern. When a different voice is input, the voice can be correctly recognized even if the voice is most similar to the new standard pattern.
【図1】本発明に係る音声パターン作成装置,標準パタ
ーン登録装置が適用された音声認識装置の構成例を示す
図である。FIG. 1 is a diagram showing a configuration example of a voice recognition device to which a voice pattern creation device and a standard pattern registration device according to the present invention are applied.
【図2】図1の音声認識装置の動作を示すフローチャー
トである。FIG. 2 is a flowchart showing an operation of the voice recognition device of FIG.
【図3】辞書への標準パターンの登録例を示す図であ
る。FIG. 3 is a diagram showing an example of registration of standard patterns in a dictionary.
【図4】標準パターンA,B,Cのそれぞれを時間軸方
向に3等分する様子を示す図である。FIG. 4 is a diagram showing how standard patterns A, B, and C are each divided into three equal parts in the time axis direction.
【図5】図4のように時間軸方向に3等分した標準パタ
ーンを組合せて作成される新たな標準パターンの一例を
示す図である。5 is a diagram showing an example of a new standard pattern created by combining standard patterns divided into three equal parts in the time axis direction as shown in FIG.
【図6】複数の標準パターンのそれぞれを時間軸方向に
分割して組合せ新たな標準パターンを作成する処理の流
れを示すフローチャートである。FIG. 6 is a flowchart showing a flow of processing of dividing each of a plurality of standard patterns in the time axis direction to create a new standard pattern by combining them.
【図7】複数の標準パターンのそれぞれを時間軸方向に
分割して組合せ新たな標準パターンを作成する処理の流
れを示すフローチャートである。FIG. 7 is a flowchart showing a flow of a process of dividing each of a plurality of standard patterns in the time axis direction and creating a new standard pattern by combining them.
【図8】複数の標準パターンのそれぞれを時間軸方向に
分割して組合せ新たな標準パターンを作成する処理の流
れを示すフローチャートである。FIG. 8 is a flowchart showing a flow of a process of dividing each of a plurality of standard patterns in the time axis direction and combining them to create a new standard pattern.
【図9】標準パターンA,B,Cのそれぞれを周波数軸
方向に2等分する様子を示す図である。FIG. 9 is a diagram showing how standard patterns A, B, and C are each divided into two equal parts in the frequency axis direction.
【図10】図9のように周波数軸方向に2等分した標準
パターンを組合せて作成される新たな標準パターンの一
例を示す図である。FIG. 10 is a diagram showing an example of a new standard pattern created by combining standard patterns bisected in the frequency axis direction as shown in FIG.
【図11】1つの標準パターンa(i,j)を、周波数
軸に対して一定量だけ高域へずらせる操作を施して、新
たな標準パターンを生成する処理の流れを示すフローチ
ャートである。FIG. 11 is a flowchart showing a flow of processing for generating a new standard pattern by performing an operation of shifting one standard pattern a (i, j) to a high frequency band with respect to the frequency axis.
【図12】元の標準パターン,新たな標準パターンの全
ての標準パターン間の類似性を求め、類似性が決められ
た値よりも大きな標準パターンについてはこれを消去す
る処理の流れを示すフローチャートである。FIG. 12 is a flow chart showing the flow of processing for obtaining the similarity between all the standard patterns of the original standard pattern and the new standard pattern, and erasing the standard pattern whose similarity is larger than the determined value. is there.
【図13】標準パターン間の類似度表の一例を示す図で
あるFIG. 13 is a diagram showing an example of a similarity table between standard patterns.
【図14】元の標準パターンと新たな標準パターンとの
間の類似性を調べ、元の標準パターン全てとの類似性が
決められた値よりも小さな標準パターンを消去する処理
の流れを示すフローチャートである。FIG. 14 is a flowchart showing the flow of processing for checking the similarity between the original standard pattern and the new standard pattern and erasing the standard pattern whose similarity to all the original standard patterns is smaller than a predetermined value. Is.
1 音声入力部 2 A/D変換器 3 特徴変換部 4 音声区間検出部 5 辞書 6 標準パターン作成部 7 認識部 11 登録部 12 分割部 13 パターン生成部 SW スイッチ 1 voice input unit 2 A / D converter 3 feature conversion unit 4 voice section detection unit 5 dictionary 6 standard pattern creation unit 7 recognition unit 11 registration unit 12 division unit 13 pattern generation unit SW switch
Claims (14)
複数の部分に分割する分割手段と、分割手段により分割
された複数の特徴パターンの各部分を組合せ連結して新
たな音声の特徴パターンを作成するパターン作成手段と
を有していることを特徴とする音声パターン作成装置。1. A new voice feature pattern is created by combining and connecting a dividing unit that divides each of a plurality of voice feature patterns into a plurality of parts, and the respective portions of the plurality of feature patterns that are divided by the dividing unit. And a pattern creating means for performing the same.
用いた標準パターン登録装置であって、複数の音声の特
徴パターンを標準パターンとしてそのままの形で辞書に
登録する登録手段をさらに有し、前記分割手段は、前記
標準パターンを複数の部分に分割し、前記パターン作成
手段は、分割された複数の標準パターンの各部分を組合
せ連結して新たな特徴パターンを生成し、前記登録手段
によりそのままの形で登録されている前記標準パターン
と併せて、該新たな特徴パターンを新たな標準パターン
として辞書に登録することを特徴とする標準パターン登
録装置。2. A standard pattern registration device using the voice pattern creation device according to claim 1, further comprising registration means for registering a plurality of voice feature patterns as standard patterns in the dictionary as they are, The dividing means divides the standard pattern into a plurality of parts, and the pattern creating means combines and connects the respective parts of the divided plurality of standard patterns to generate a new characteristic pattern, which is directly registered by the registering means. A standard pattern registration device characterized in that the new feature pattern is registered in the dictionary as a new standard pattern together with the standard pattern registered in the form.
おいて、前記パターン作成手段は、複数の標準パターン
のそれぞれを複数の部分に分割し、これらを組合せ連結
して新たな標準パターンを作成する際に、分割前の標準
パターンにおいて先頭にあった部分は連結後の新たな標
準パターンにおいても先頭に位置させるように組合せる
ことを特徴とする標準パターン登録装置。3. The standard pattern registration apparatus according to claim 2, wherein the pattern creating means divides each of the plurality of standard patterns into a plurality of parts, and combines and connects these parts to create a new standard pattern. In addition, the standard pattern registration device is characterized in that the portion at the beginning of the standard pattern before division is combined so as to be positioned at the beginning of the new standard pattern after connection.
おいて、前記パターン作成手段は、複数の標準パターン
のそれぞれを複数の部分に分割し、これらを組合せ連結
して新たな標準パターンを作成する際に、分割前の標準
パターンにおいて末尾にあった部分は連結後の新たな標
準パターンにおいても末尾に位置させるように組合せる
ことを特徴とする標準パターン登録装置。4. The standard pattern registration device according to claim 2, wherein the pattern creating means divides each of the plurality of standard patterns into a plurality of parts, and combines and connects them to create a new standard pattern. In addition, the standard pattern registration device is characterized in that the part at the end of the standard pattern before division is combined so as to be positioned at the end of the new standard pattern after connection.
おいて、前記パターン作成手段は、複数の標準パターン
のそれぞれを複数の部分に分割し、これらを組合せ連結
して新たな標準パターンを作成する際に、分割前の標準
パターンにおいて先頭にあった部分は連結後の新たな標
準パターンにおいても先頭に位置させ、また、分割前の
標準パターンにおいて末尾にあった部分は連結後の新た
な標準パターンにおいても末尾に位置させるように組合
せることを特徴とする標準パターン登録装置。5. The standard pattern registration device according to claim 2, wherein the pattern creating means divides each of the plurality of standard patterns into a plurality of parts, and combines and connects them to create a new standard pattern. In addition, the part at the beginning of the standard pattern before division is positioned at the beginning of the new standard pattern after concatenation, and the part at the end of the standard pattern before division is in the new standard pattern after concatenation. A standard pattern registration device characterized in that it is also combined so that it is located at the end.
おいて、前記分割手段は、複数の音声の標準パターンの
それぞれを時間軸方向に分割し、前記パターン作成手段
は、時間軸方向に分割された各標準パターンの各部分を
所定の規則に従って組合せ連結して新たな音声パターン
として作成し、この過程を繰り返すことにより作成され
た音声パターンの全て,または一部を標準パターンとし
て登録するようになっていることを特徴とする標準パタ
ーン登録装置。6. The standard pattern registration device according to claim 2, wherein said dividing means divides each of a plurality of standard patterns of voice in a time axis direction, and said pattern creating means is divided in a time axis direction. By combining and connecting each part of each standard pattern according to a predetermined rule to create a new voice pattern, by repeating this process, all or part of the created voice pattern can be registered as a standard pattern. A standard pattern registration device characterized in that
おいて、前記分割手段は、標準パターンを時間軸方向に
等分割するか、または、時間軸方向において標準パター
ンの変化の大きな部分を検出し、この部分で標準パター
ンを分割するか、または、時間軸方向において音声エネ
ルギーの極小部分を検出し、この部分で標準パターンを
分割するようになっていることを特徴とする標準パター
ン登録装置。7. The standard pattern registration device according to claim 6, wherein the dividing means divides the standard pattern into equal parts in the time axis direction, or detects a portion in which the standard pattern changes largely in the time axis direction. A standard pattern registration device characterized in that the standard pattern is divided at this portion, or a minimum portion of voice energy in the time axis direction is detected, and the standard pattern is divided at this portion.
おいて、前記分割手段は、複数の音声の標準パターンの
それぞれを周波数軸方向に分割し、前記パターン作成手
段は、周波数軸方向に分割された各標準パターンの各部
分を所定の規則に従って組合せ連結して新たな音声パタ
ーンとして作成し、この過程を繰り返すことにより作成
された音声パターンの全て,または一部を標準パターン
として登録するようになっていることを特徴とする標準
パターン登録装置。8. The standard pattern registration device according to claim 2, wherein the dividing unit divides each of the plurality of standard patterns of voice in the frequency axis direction, and the pattern creating unit divides in the frequency axis direction. By combining and connecting each part of each standard pattern according to a predetermined rule to create a new voice pattern, by repeating this process, all or part of the created voice pattern can be registered as a standard pattern. A standard pattern registration device characterized in that
を、周波数軸方向に一定量だけ高域へおよび/または低
域へずらし、新たな音声の特徴パターンを作成するパタ
ーン作成手段を有していることを特徴とする音声パター
ン作成装置。9. A pattern creating means for creating a new audio feature pattern by shifting each of a plurality of audio feature patterns to a high frequency band and / or a low frequency band in the frequency axis direction by a fixed amount. A voice pattern creating device characterized by the above.
を用いた標準パターン登録装置であって、複数の音声の
特徴パターンを標準パターンとしてそのままの形で辞書
に登録する登録手段をさらに有し、前記パターン作成手
段は、複数の標準パターンのそれぞれを周波数軸方向に
一定量だけ高域へおよび/または低域へずらし、新たな
音声パターンを作成したとき、前記登録手段によりその
ままの形で登録されている前記標準パターンと併せて、
該新たな音声パターンの全て,または一部を標準パター
ンとして辞書に登録するようになっていることを特徴と
する標準パターン登録装置。10. A standard pattern registration device using the voice pattern creation device according to claim 9, further comprising registration means for registering a plurality of voice characteristic patterns as standard patterns in the dictionary as they are, The pattern creating means shifts each of the plurality of standard patterns to a high frequency band and / or a low frequency band in the frequency axis direction by a certain amount, and when a new voice pattern is created, is registered in the same form by the registration means. In addition to the standard pattern that is
A standard pattern registration device characterized in that all or part of the new voice pattern is registered in a dictionary as a standard pattern.
パターン登録装置において、そのままの形で辞書に登録
された元の標準パターンと前記パターン作成手段により
作成された新たな標準パターンとの間の類似性を求め、
元の標準パターンとの類似性が決められた値よりも大き
な新たな標準パターンについては、これを辞書から消去
するようになっていることを特徴とする標準パターン登
録装置。11. The standard pattern registration device according to claim 2 or 10, wherein an original standard pattern registered in the dictionary as it is and a new standard pattern created by the pattern creating means are provided. Seeking similarity,
A standard pattern registration device characterized in that a new standard pattern whose similarity to the original standard pattern is larger than a predetermined value is deleted from the dictionary.
パターン登録装置において、そのままの形で登録された
元の標準パターンと新たな標準パターンとの間の類似性
を求め、元の標準パターンとの類似性が決められた値よ
りも小さな新たな標準パターンについては、これを辞書
から消去するようになっていることを特徴とする標準パ
ターン登録装置。12. The standard pattern registration device according to claim 2 or 10, wherein the similarity between the original standard pattern registered as it is and the new standard pattern is calculated to obtain the original standard pattern. The standard pattern registration device is characterized in that a new standard pattern whose similarity is smaller than a predetermined value is deleted from the dictionary.
ン登録装置において、前記登録手段は、複数の音声の標
準パターンをそのままの形で登録するに際し、該標準パ
ターンを所定の言葉と対応付けして辞書に登録する一
方、前記パターン作成手段は、新たな標準パターンを言
葉との対応付けを行なわずに辞書に登録するようになっ
ていることを特徴とする標準パターン登録装置。13. The standard pattern registration apparatus according to claim 2 or 10, wherein said registration means associates the standard pattern with a predetermined word when registering the standard patterns of a plurality of voices as they are. The standard pattern registration device is characterized in that, while registering in a dictionary, the pattern creating means registers a new standard pattern in the dictionary without associating it with a word.
ン登録装置において、前記登録手段は、複数の音声の標
準パターンをそのままの形で登録するに際し、該標準パ
ターンを所定の言葉と対応付けして辞書に登録し、ま
た、前記パターン作成手段も、新たな標準パターンを所
定の言葉と対応付けして辞書に登録するようになってい
ることを特徴とする標準パターン登録装置。14. The standard pattern registration device according to claim 2 or 10, wherein said registration means associates the standard pattern with a predetermined word when registering the standard patterns of a plurality of voices as they are. A standard pattern registration device, characterized in that the standard pattern registration means registers in a dictionary, and the pattern creating means also registers a new standard pattern in the dictionary in association with a predetermined word.
Priority Applications (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP4341224A JPH06167992A (en) | 1992-11-27 | 1992-11-27 | Voice pattern creating device and standard pattern registration device using the same |
Applications Claiming Priority (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP4341224A JPH06167992A (en) | 1992-11-27 | 1992-11-27 | Voice pattern creating device and standard pattern registration device using the same |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| JPH06167992A true JPH06167992A (en) | 1994-06-14 |
Family
ID=18344347
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| JP4341224A Pending JPH06167992A (en) | 1992-11-27 | 1992-11-27 | Voice pattern creating device and standard pattern registration device using the same |
Country Status (1)
| Country | Link |
|---|---|
| JP (1) | JPH06167992A (en) |
Cited By (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| KR100441181B1 (en) * | 1995-04-07 | 2005-04-06 | 소니 가부시끼 가이샤 | Voice recognition method and device |
| CN112334975A (en) * | 2018-06-29 | 2021-02-05 | 索尼公司 | Information processing apparatus, information processing method and program |
-
1992
- 1992-11-27 JP JP4341224A patent/JPH06167992A/en active Pending
Cited By (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| KR100441181B1 (en) * | 1995-04-07 | 2005-04-06 | 소니 가부시끼 가이샤 | Voice recognition method and device |
| CN112334975A (en) * | 2018-06-29 | 2021-02-05 | 索尼公司 | Information processing apparatus, information processing method and program |
| US12067971B2 (en) | 2018-06-29 | 2024-08-20 | Sony Corporation | Information processing apparatus and information processing method |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| JP3078205B2 (en) | Speech synthesis method by connecting and partially overlapping waveforms | |
| US20110131038A1 (en) | Exception dictionary creating unit, exception dictionary creating method, and program therefor, as well as speech recognition unit and speech recognition method | |
| JPH07334184A (en) | Calculating device for acoustic category mean value and adapting device therefor | |
| JP2876861B2 (en) | Automatic transcription device | |
| JPS60158498A (en) | pattern matching device | |
| US6594631B1 (en) | Method for forming phoneme data and voice synthesizing apparatus utilizing a linear predictive coding distortion | |
| US5970454A (en) | Synthesizing speech by converting phonemes to digital waveforms | |
| JPH0854891A (en) | Device and method for acoustic classification process and speaker classification process | |
| CN118379983A (en) | A method, system and storage medium for cross-speaker emotional speech synthesis | |
| JPH07319495A (en) | Synthetic unit data generation method and method for speech synthesizer | |
| US20240144934A1 (en) | Voice Data Generation Method, Voice Data Generation Apparatus And Computer-Readable Recording Medium | |
| JPH02232696A (en) | Voice recognition device | |
| JP2980382B2 (en) | Speaker adaptive speech recognition method and apparatus | |
| JPH04324499A (en) | Speech recognition device | |
| US6502074B1 (en) | Synthesising speech by converting phonemes to digital waveforms | |
| JP3514481B2 (en) | Voice recognition device | |
| JP2757356B2 (en) | Word speech recognition method and apparatus | |
| JP3348735B2 (en) | Pattern matching method | |
| JPH0556519B2 (en) | ||
| KR100236962B1 (en) | Method for speaker dependent allophone modeling for each phoneme | |
| JPS6060697A (en) | Voice standard feature pattern generation processing system | |
| Mikkilineni et al. | A procedure to generate training sequences for a connected word recognizer using the segmental k-means training algorithm. | |
| Zalewski | Text dependent speaker recognition in noise. | |
| JPS60175098A (en) | Voice recognition equipment | |
| JPH04347898A (en) | Voice recognizing method |