JPH053596B2 - - Google Patents
Info
- Publication number
- JPH053596B2 JPH053596B2 JP59269755A JP26975584A JPH053596B2 JP H053596 B2 JPH053596 B2 JP H053596B2 JP 59269755 A JP59269755 A JP 59269755A JP 26975584 A JP26975584 A JP 26975584A JP H053596 B2 JPH053596 B2 JP H053596B2
- Authority
- JP
- Japan
- Prior art keywords
- standard pattern
- pattern group
- section
- similarity
- voice
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Expired - Lifetime
Links
Description
産業上の利用分野
本発明は音声の内容を自動的に認識するための
音声認識装置に関するものである。
従来の技術
最近音声入力装置は荷物の仕分け、銀行預金の
問い合わせ等の分野で盛んに利用されるようにな
つてきた。この音声認識装置は例えば三輪譲二他
により「音声スペクトルの概略形とその動特性を
利用した単語音声認識システム」日本音響学会誌
34・3(1978)に記載されている構成が知られて
いる。以下、第3図を参照して、従来の音声認識
装置について説明する。
まずあらかじめ多数話者の音声データをクラス
タリング手法等を用いてグループ分けし、音素単
位で標準パタンを作成し、標準パタン群1,2,
……として格納部59に格納しておく。ここでは
説明のため、格納部59中の標準パタン群1は男
声で、標準パタン群2は女声で作成したものが格
納されているとする。又、説明を簡単にするため
格納部59中には前記標準パタン群1,2のみ用
意されているとする。
マイク51からの入力音声はA/D変換器52
で変換した後に一方は信号処理回路53へ送られ
プリエンフアシス、窓計算を行つて特徴抽出部5
4へ転送される。他方、セグメンテーシヨン部5
6は帯域パワー計算、音声区間の検出、有音無声
判定、子音のセグメンテーシヨンを行い、結果を
メインメモリ57に転送する。類似度計算部55
は、特徴抽出部54で得た帯域フイルタ群の出力
をパラメータとして、格納部59内の標準パタン
との類似度計算を行なう。まず標準パタン群1と
の間の類似度計算を行い、結果をメインメモリ5
7に転送する。同様にして標準パタン群2につい
ても行い、メインプロセツサ58はメインメモリ
57中の最も類似度の高かつた標準パタンに相当
する音素を認識結果として採用する。さらにセグ
メンテーシヨン部56の結果を用いて音素の時系
列を作成する。
最後に前記音素の時系列を単語マツチング部6
1に送り、単語辞書52と、用意されている単一
のコンフユージヨンマトリクスとを照合して単語
類似度を求めることによつて単語認識を行い、メ
インプロセツサ58は最も単語類似度の高かつ単
語を出力部63に出力する。
発明が解決しようとする問題点
しかし、以上のような構成では次の問題点を有
していた。
(1) マイクの種類、騒音の内容、話者の種類によ
つて音素の認識、置換、付加、脱落等の認識傾
向が異なるため、単一のコンフユージヨンマト
リクスでは前記認識傾向との間にずれが生じる
ことにより、単語認識率が低下する。
(2) 複数組の標準パタン群に対し常に類似度計算
を行う必要があり、計算量が多くなつてしまう
とともに、標準パタン数の増大によつて認識時
における互いの混同が多くなり、音素認識率を
低下させる。
本発明は上記問題を解決するもので、未知入力
音声を用いて、入力環境や話者の種類に最も適し
た音素の標準パタン群およびコンフユージヨンマ
トリクスを自動選択して用いることにより、入力
環境や話者の種類の変化に対して安定で高い単語
認識率を得る音声認識装置を提供することを目的
とする。
問題点を解決するための手段
本発明は、入力音声を音素毎に分割するセグメ
ンテーシヨン部と、同一の音声入力環境又は話者
の種類によつて作成した標準パタン群とコンフユ
ージヨンマトリクスの対を複数組用意して格納す
る格納部と、前記音素標準パタン群を用いて類似
度を算出する類似度計算部と、セグメンテーシヨ
ン部と類似度計算部の結果を用いて音素系列を作
成し、その中の母音・鼻音の中心位置を求め、か
つ装置全体の制御を行うメインプロセツサと、前
記メインプロセツサの結果を用いて入力音声に最
も類似した標準パタン群を自動選択する選択部
と、前記選択された標準パタン群の対となるコン
フユージヨンマトリクスを用いて単語認識を行う
単語マツチング部を主な要素として構成された音
声認識装置により、上記目的を達成するものであ
る。
作 用
本発明は予め同一の音声入力環境又は話者の種
類によつて作成した標準パタン群とコンフユージ
ヨンマトリクスを対として準備しておき、音声の
入力環境、話者の種類が変化した場合、その使用
条件に最も適した音素標準パタン群及びコンフユ
ージヨンマトリクスを自動的に選択するので、類
似度計算における計算量を増大させることなく、
認識時の混同をさけ高い音素認識率を達成するこ
とができると同時に、単語認識率の向上をも図る
ことができる。
実施例
本発明は不特定話者の音声を、周囲の騒音の種
類やマイクの種類等の入力環境に影響されること
なく、又性別、年令等の話者の違いに影響される
ことなく、安定に認識できる音声認識装置を提供
するものである。
そのために、音声を認識するための標準パタン
群を騒音の種類、マイクの種類等の入力環境によ
つて、又性別、年令等の話者の種類によつて分け
て作成しておく。又、同一条件にて多数話者の音
声からコンフユージヨンマトリクスを作成し、上
記標準パタン群と対にして用意しておく。コンフ
ユージヨンマトリクス(以下CMと略す)とは、
音素同士の置換の割合、各音素の脱落する割合、
音素間に他の音素の付加する割合を表わしたもの
である。
実際に入力される音声はどういう入力環境、あ
るいは話者の種類であるかは不明であるが、本方
法を用いることによつて、どういう入力環境か、
あるいは話者の種類であるかを自動的に決定し、
入力音声に最も適した標準パタン群とCMを使用
して安定した認識を実現することができる。
以下の実施例では成人男女を対象に、男声と女
声の2つの種類で認識を行う場合の方法について
説明する。本発明の一実施例における音声認識装
置の機能ブロツク図を第1図に示す。
まず格納部9に入れる内容を説明する。まず男
声の音素iのLPCケプストラム係数をパラメー
タとし、その平均値を
Mi (1)=(Mi1 (1)、Mi2 (1)、……、Mip (1))
とする。pは使用パラメータ数である。これを母
音|a|、|i|、|u|、|e|、|o|と鼻音の
6個について求める。
同様に女声に対しても6個求め、計12個とす
る。この12個の音素に共通の共分数行列を|Rと
し、逆行列を
〓-1とする。〓-1の(j、j′)要素をijj′とする
と、パラメータのj次に対する重み係数は例えば
男声の音素iに対し
aij(1)=2p
〓j
′=1rjj′Mij′(1) (1)
で求める。又平均値に対する距離di(1)は
di(1)=〓i(1)t〓−1〓i(1) (2)
で求める。
このaij(1)、di(1)を音素毎に求め、格納部9の標
準パタン群1に入れる。同時にこの標準パタン群
1には、男声の他の音素の標準パタン(子音・半
母音等)も共存させておく。
同様に、標準パタン群2には女声の各種標準パ
タンを格納しておく。
更に格納部9には上記各標準パタン群と対をな
すコンフユージヨンマトリクスCM1、CM2、…
…、CMoを格納しておく。この場合コンフユー
ジヨンマトリクスは音素標準パタン群と同一条件
にて多数話者の音声から作成されていれば良く、
格納場所は別の箇所に格納されても良いことはも
ちろんである。
次に、類似度計算部5の動作について説明す
る。未知入力音声がマイク1から入力されると、
A/D変換器2でA/D変換し、信号処理回路3
でプリエンフアシス、窓処理を行つた後、特徴抽
出部4にてLPCケプストラム係数Cj(j=1、
2、……、p)を求める。類似度計算部5はこの
Cjと、選択部10を通して転送された格納部9中
の標準パタンを用いて、類似度計算を行なう。標
準パタン群1の音素iに対する類似度li(1)は
li(1)=p
〓j=1
aij(1)Cj−di(1) (3)
で求める。標準パタン群2に対しては
li(2)=p
〓j=1
aij(2)Cj−di(2) (4)
で求め、メインメモリ7へ転送する。
次に選択部10の動作とメインプロセツサ8と
の関係について述べる。メインプロセツサ8は、
メインメモリ7に転送された類似度計算部5の結
果とセグメンテーシヨン部6の結果を参照して、
母音・鼻音の位置を求め、その中の中心位置を決
定する。さらに、その位置における母音・鼻音の
全ての類似度(標準パターン群ごとに6個づつ、
標準パタン群1,2では計12個)の中で最大の類
似度を持つものを選ぶ。選ばれた標準パタンが標
準パタン群1に属する場合は、選ばれた回数N(1)
を1増やす。標準パタン群2に属する場合は回数
N(2)を1増やす。この操作を各母音・鼻音の中心
位置毎にくり返し、N(1)又はN(2)をカウントアツ
プしていく。
N(1)、N(2)を用いて入力単語毎に信頼度Reを次
式で求める。
Re(1)=N(1)/N(2) (5)
Re(2)=N(2)/N(4) (6)
このRe(1)又はRe(2)があらかじめ定められた閾
値を超えた時点で閾値を超えた方の標準パタン群
を以降使用することを決定し、学習を打切る。
閾値を超えない間はN(1)とN(2)を比較し、大き
い方の標準パタン群を使用するよう、メインプロ
セツサに指示する。
同時に選択された標準パタン群と対の関係にあ
るCMを単語マツチングに使用するよう、単語マ
ツチング部11に指示する。
単語マツチングの方法について説明する。
単語辞書12の各項目を〓とし、セグメンテー
シヨン部6の結果に基づきメインプロセツサ8で
作成された音素系列を〓とし、両者の類似度をS
(〓、〓)で表わす。音素系列に脱落と付加がそ
れぞれ3音素以上連続しないことと、脱落と付加
が連続しないことを仮定すると、S(〓、〓)は
次式で計算する。
S(〓,〓)
=g(I+1、J+1)/(I+J) (7)
g(i、j)=2l(i、j)
+max{L、La、Laa、Lo、Loo} (8)
L=g(i−1、j−1)
La=g(i−1、j−2)+la(j−1)
La=g(i−1、j−2)+la(j−1)
Laa=g(i−1、j−3)+laa(j−2)+laa(j
−1)
Lo=g(i−2、j−1)+lo(j−1)
Lo=g(i−2、j−1)+lo(j−1)
Loo=g(i−3、j−1)−loo(i−2)+loo(i
−1)(9)
ここでIは〓の音素数、Jは〓の音素数、l
(i、j)は〓のi番目音素と〓のj番目音素の
尤度関数、la(j)、lo(i)は1音素の付加と脱落の尤
度関数、laa(j)、loo(i)は2音素連続の付加と脱落
の尤度関数を表わす。
得られたS(〓、〓の中で最大類似度となる〓
を単語認識結果として、メインプロセツサ8を経
由して出力部13へ出力する。
従つてCMの内容は前記l(i、j)、la(j)、lo
(i)、laa(j)、loo(i)に相当し、その内容は入力環
境、話者の種類によつて分類して用意しておくこ
とにより、使用条件にあつた音素認識の誤り傾向
を反映させた単語マツチングを行うことができ
る。
本実施例による処理の流れを第2図に示す。最
初の音声が入力された(判断イ)後、プリエンフ
アシス、窓計算、LPCケプストラム係数の計算
等の音響分析を行う(ステツプロ)。次に、セグ
メンテーシヨンと類似度計算(1)を行い(ステツプ
ハ)、母音・鼻音の中心位置において、全ての標
準パタン群の中で最大類似度のパタンを選出する
(ステツプニ)。これによつて前記N(1)、N(2)を求
めておき、(5)式、(6)式で信頼度を計算する(ステ
ツプホ)。この信頼度が閾値以上であつた場合
(判断ヘ)は、閾値を超えた方の標準パタン群を
使用することに決定する。又、この標準パタン群
と対にして用意されたCMを使用することに決定
する(ステツプト)。閾値以下であつた場合は上
記N(1)とN(2)を比較し、大きい方の標準パタン群
およびそれと対に用意されたCMを使用する。
以下指示された標準パタン群を用いて音素認識
を行い(ステツプチ)、指示されたCMを用いて
単語認識を行い(ステツプリ)、結果を出力して
(ステツプヌ)もとにもどる。
次に音声が入力されたら音響分析(ステツプ
ロ)の後、選択終了か否かを調べ(判断ル)、さ
れてなければ最初の音声と同様な処理をくり返
す。されていれば、すでに選択された標準パタン
群を用いてセグメンテーシヨン、類似度計算(2)を
行い(ステツプオ)、音素認識を行つた(ステツ
プチ)後、すでに選択されているCMを用いて単
語認識を行い(ステツプリ)、その結果を出力す
る(ステツプヌ)。
本実施例は標準パタン群が男声と女声の場合に
ついて述べたが、話者の種類でなく、音声の入力
環境(外部騒音の種類、マイクの種類等)によつ
て分けて標準パタン群を作り、同一の入力環境に
てコンフユージヨンマトリクスを標準パタン群と
対にして用意すると、本方法はより有効である。
又、格納部は1つでなくても良く、複数個に分け
て格納しておいても良いのは当然である。
又、標準パタン群が3組以上の場合にも同様な
方法で選択作業が出来、(5)式、(6)式のN(1)、N(2)
にN(3)、N(4)……が加わつてその中で最も大きい
2つのNを用いてその比を計算すれば信頼度Re
を求めることができる。
このように、本方法は音声の入力環境、話者の
種類によつて複数組の標準パタン群を用意し、又
同一条件にてコンフユージヨンマトリクスを標準
パタン群と対に用意しておき、母音・鼻音の中心
付近の位置における類似度を求め、最も類似度の
高い標準パタンの属する標準パタン群の使用回数
を用いて選択のための信頼度を求め、信頼度が閾
値を超えた時点で選択作業を終了することを特徴
とし、特に複雑な演算、処理を要することなく実
現することができる。
本実施例の音声認識装置について成人男女100
名を対象に、212単語中の最初の10単語を用いて、
選択に必要な単語数を人数で累計して評価した結
果を第1表に示す。
FIELD OF THE INVENTION The present invention relates to a speech recognition device for automatically recognizing the content of speech. 2. Description of the Related Art Recently, voice input devices have come into widespread use in fields such as sorting luggage and inquiring about bank deposits. This speech recognition device was developed, for example, by Joji Miwa et al. in ``Word speech recognition system using the outline form of the speech spectrum and its dynamic characteristics,'' Journal of the Acoustical Society of Japan.
34.3 (1978) is known. Hereinafter, a conventional speech recognition device will be explained with reference to FIG. First, the speech data of many speakers is divided into groups using clustering methods, etc., and standard patterns are created for each phoneme.Standard pattern groups 1, 2,
... is stored in the storage unit 59. For the sake of explanation, it is assumed here that the standard pattern group 1 in the storage section 59 is created for a male voice, and the standard pattern group 2 is stored for a female voice. Further, for the sake of simplicity, it is assumed that only the standard pattern groups 1 and 2 are prepared in the storage section 59. Input audio from the microphone 51 is sent to the A/D converter 52
After being converted by
Transferred to 4. On the other hand, the segmentation section 5
6 performs band power calculation, voice section detection, voiced/unvoiced determination, and consonant segmentation, and transfers the results to the main memory 57. Similarity calculation unit 55
calculates the degree of similarity with the standard pattern in the storage unit 59 using the output of the band filter group obtained by the feature extraction unit 54 as a parameter. First, calculate the similarity with the standard pattern group 1, and store the result in the main memory 5.
Transfer to 7. The same process is performed for the standard pattern group 2, and the main processor 58 adopts the phoneme corresponding to the standard pattern with the highest degree of similarity in the main memory 57 as the recognition result. Furthermore, a time series of phonemes is created using the results of the segmentation unit 56. Finally, word matching unit 6
1, word recognition is performed by comparing the word dictionary 52 with a single prepared confusion matrix to determine word similarity, and the main processor 58 selects the words with the highest word similarity. And the word is output to the output section 63. Problems to be Solved by the Invention However, the above configuration has the following problems. (1) Since the recognition tendency of phoneme recognition, substitution, addition, omission, etc. differs depending on the type of microphone, the content of the noise, and the type of speaker. Due to the deviation, the word recognition rate decreases. (2) It is necessary to always perform similarity calculations for multiple sets of standard patterns, which increases the amount of calculation, and as the number of standard patterns increases, they are often confused with each other during recognition, making it difficult to recognize phonemes. reduce the rate. The present invention solves the above problem, and uses unknown input speech to automatically select and use a standard phoneme pattern group and conflation matrix that are most suitable for the input environment and type of speaker. It is an object of the present invention to provide a speech recognition device that obtains a stable and high word recognition rate even when the type of speaker changes. Means for Solving the Problems The present invention includes a segmentation unit that divides input speech into phonemes, and a standard pattern group and confusion matrix created based on the same speech input environment or type of speaker. A storage section that prepares and stores a plurality of pairs, a similarity calculation section that calculates similarity using the phoneme standard pattern group, and a phoneme sequence created using the results of the segmentation section and similarity calculation section. a main processor that determines the center position of vowels and nasal sounds therein and controls the entire device; and a selection section that automatically selects a group of standard patterns most similar to the input speech using the results of the main processor. The above object is achieved by a speech recognition device configured mainly of a word matching unit that performs word recognition using a fusion matrix that is a pair of the selected standard pattern group. Function The present invention prepares in advance a standard pattern group and a conflation matrix created for the same voice input environment or type of speaker as a pair, and when the voice input environment or type of speaker changes. , automatically selects the phoneme standard pattern group and conflation matrix that are most suitable for the usage conditions, without increasing the amount of calculation in similarity calculation.
It is possible to avoid confusion during recognition and achieve a high phoneme recognition rate, and at the same time, it is possible to improve the word recognition rate. Embodiment The present invention allows the voice of unspecified speakers to be heard without being affected by the input environment such as the type of surrounding noise or the type of microphone, and without being affected by differences between speakers such as gender, age, etc. , to provide a speech recognition device that can stably recognize speech. For this purpose, a group of standard patterns for recognizing speech are created separately according to the input environment such as the type of noise and the type of microphone, and according to the type of speaker such as gender and age. Also, a conflation matrix is created from the voices of multiple speakers under the same conditions, and is prepared as a pair with the standard pattern group. What is Confusion Matrix (hereinafter abbreviated as CM)?
The rate of substitution between phonemes, the rate of dropout of each phoneme,
This represents the rate at which other phonemes are added between phonemes. Although it is unknown what kind of input environment or type of speaker the actual input voice is, by using this method, it is possible to
or automatically determine the type of speaker,
Stable recognition can be achieved by using the standard pattern group and CM that are most suitable for the input voice. In the following embodiment, a method for recognizing two types of voices, male and female voices, will be described for adult men and women. FIG. 1 shows a functional block diagram of a speech recognition device according to an embodiment of the present invention. First, the contents to be stored in the storage section 9 will be explained. First, the LPC cepstrum coefficient of male voice phoneme i is taken as a parameter, and its average value is set as M i (1) = (M i1 (1) , M i2 (1) , . . . , M ip (1) ). p is the number of parameters used. This is calculated for six vowels: |a|, |i|, |u|, |e|, |o|, and nasal sounds. Similarly, find 6 pieces for the female voice, making a total of 12 pieces. Let the co-fraction matrix common to these 12 phonemes be |R, and let the inverse matrix be 〓 -1 . If the (j, j′) element of 〓 -1 is ijj′, then the weighting coefficient for the j-th parameter is, for example, aij (1) = 2 p 〓 j ′ =1 rjj′Mij′ (1 ) (1). Also, the distance di (1) to the average value is determined by di (1) =〓i (1)t 〓−1〓i (1) (2). These aij (1) and di (1) are obtained for each phoneme and stored in the standard pattern group 1 in the storage section 9. At the same time, in this standard pattern group 1, standard patterns of other phonemes (consonants, semi-vowels, etc.) for male voices are also allowed to coexist. Similarly, standard pattern group 2 stores various standard patterns for female voices. Furthermore, the storage section 9 stores conflation matrices CM 1 , CM 2 , . . . that are paired with each of the above standard pattern groups.
..., store CM o . In this case, the conflation matrix only needs to be created from the voices of multiple speakers under the same conditions as the phoneme standard pattern group.
Of course, the storage location may be stored in another location. Next, the operation of the similarity calculation section 5 will be explained. When unknown input audio is input from microphone 1,
A/D converter 2 performs A/D conversion, and signal processing circuit 3
After performing pre-emphasis and window processing, the feature extraction unit 4 extracts the LPC cepstral coefficients Cj (j=1,
2, ..., p) is found. The similarity calculation unit 5 uses this
Similarity calculation is performed using Cj and the standard pattern in the storage section 9 transferred through the selection section 10. The similarity li (1) of standard pattern group 1 to phoneme i is determined by li (1) = p 〓 j=1 aij (1) Cj−di (1) (3). For the standard pattern group 2, li (2) = p 〓 j=1 aij (2) Cj−di (2) (4) is obtained and transferred to the main memory 7. Next, the operation of the selection section 10 and its relationship with the main processor 8 will be described. The main processor 8 is
With reference to the results of the similarity calculation unit 5 and the results of the segmentation unit 6 transferred to the main memory 7,
Find the positions of vowels and nasal sounds, and determine the center position among them. Furthermore, all similarities of vowels and nasals at that position (6 for each standard pattern group,
(12 patterns in total for standard pattern groups 1 and 2) The one with the highest degree of similarity is selected. If the selected standard pattern belongs to standard pattern group 1, the number of times it was selected N (1)
Increase by 1. If it belongs to standard pattern group 2, the number of times
Increase N (2) by 1. Repeat this operation for each vowel/nasal center position and count up N (1) or N (2) . Reliability Re is calculated for each input word using N (1) and N (2) using the following formula. Re (1) =N (1) /N (2) (5) Re (2) =N (2) /N (4) (6) This Re (1) or Re (2) is a predetermined threshold When the threshold value is exceeded, the standard pattern group that exceeds the threshold value is decided to be used from now on, and learning is terminated. As long as the threshold is not exceeded, N (1) and N (2) are compared and the main processor is instructed to use the larger standard pattern group. The word matching unit 11 is instructed to use the CM in a pairing relationship with the simultaneously selected standard pattern group for word matching. The word matching method will be explained. Let each item of the word dictionary 12 be 〓, let the phoneme sequence created by the main processor 8 based on the result of the segmentation unit 6 be 〓, and let the similarity between the two be S.
Represented by (〓, 〓). S (〓, 〓) is calculated by the following formula, assuming that dropouts and additions do not occur consecutively for three or more phonemes in the phoneme sequence, and that dropouts and additions do not occur consecutively. S(〓,〓) =g(I+1, J+1)/(I+J) (7) g(i, j)=2l(i, j) +max{L, La, Laa, Lo, Loo} (8) L= g(i-1, j-1) La=g(i-1, j-2)+la(j-1) La=g(i-1, j-2)+la(j-1) Laa=g( i-1, j-3) + laa (j-2) + laa (j
-1) Lo=g(i-2, j-1)+lo(j-1) Lo=g(i-2, j-1)+lo(j-1) Loo=g(i-3, j-1 )−loo(i−2)+loo(i
−1)(9) Here, I is the number of phonemes of 〓, J is the number of phonemes of 〓, l
(i, j) is the likelihood function of the i-th phoneme of 〓 and the j-th phoneme of 〓, la(j), lo(i) are the likelihood functions of adding and dropping one phoneme, laa(j), loo( i) represents the likelihood function of addition and omission of two consecutive phonemes. Obtained S(〓, maximum similarity among 〓〓
is output to the output unit 13 via the main processor 8 as a word recognition result. Therefore, the content of the CM is the above l(i, j), la(j), lo
(i), laa(j), and loo(i), and their contents can be categorized and prepared according to the input environment and type of speaker to determine the error tendency of phoneme recognition that meets the usage conditions. It is possible to perform word matching that reflects the FIG. 2 shows the flow of processing according to this embodiment. After the first voice is input (judgment A), acoustic analysis such as pre-emphasis, window calculation, and calculation of LPC cepstral coefficients is performed (STUTSUPRO). Next, segmentation and similarity calculation (1) are performed (step 1), and the pattern with the maximum similarity among all the standard pattern groups at the center position of the vowel/nasal sound is selected (step 2). With this, N (1) and N (2) are obtained, and the reliability is calculated using equations (5) and (6) (step 4). If the reliability is greater than or equal to the threshold (determination), it is determined to use the standard pattern group that exceeds the threshold. Also, it is decided to use the CM prepared in pair with this standard pattern group (step). If it is below the threshold, compare N (1) and N (2) above, and use the larger standard pattern group and the CM prepared for its pair. Next, phoneme recognition is performed using the specified standard pattern group (step), word recognition is performed using the specified commercial (step), the results are output (step), and the process returns to the beginning. When the next voice is input, after acoustic analysis (STEP PRO), it is checked whether the selection has been completed (JUDGE), and if not, the same process as for the first voice is repeated. If so, perform segmentation and similarity calculation (2) using the already selected standard pattern group (step), perform phoneme recognition (step), and then perform segmentation and similarity calculation (2) using the already selected standard pattern group. Performs word recognition (Stepply) and outputs the result (Stepnu). In this example, the standard pattern groups are for male and female voices, but the standard pattern groups are created based on the voice input environment (type of external noise, type of microphone, etc.) rather than the type of speaker. This method is more effective if the fusion matrix is prepared in pairs with a group of standard patterns in the same input environment.
Furthermore, it is natural that the number of storage units does not need to be one, and that the storage unit may be divided into a plurality of units. Also, when there are three or more standard pattern groups, selection can be done in the same way, and N (1) and N (2) in equations (5) and (6)
If we add N (3) , N (4) ... and calculate the ratio using the two largest Ns, the reliability Re
can be found. In this way, in this method, multiple sets of standard patterns are prepared depending on the voice input environment and the type of speaker, and a conflation matrix is prepared in pairs with the standard pattern group under the same conditions. The degree of similarity at the position near the center of the vowel/nasal sound is determined, and the degree of confidence for selection is determined using the number of times the standard pattern group to which the standard pattern with the highest degree of similarity belongs is used, and when the degree of confidence exceeds the threshold, It is characterized by completing the selection work, and can be realized without requiring particularly complex calculations or processing. Regarding the voice recognition device of this example, 100 adult men and women
Using the first 10 words out of 212 words,
Table 1 shows the results of the cumulative evaluation of the number of words required for selection based on the number of people.
【表】
すなわち、4単語を用いれば100人中98人まで
正しく標準パタン群の選択を行うこができる。誤
つた1名は女声を男声に誤つた場合であるが、こ
の話者は音声の性質が男声と女声の中間に位置
し、誤つても母音・鼻音認識率の低下は極めて少
ない。
このように本方法を用いれば高い確度で最適な
標準パタン群の選択を行なうことができる。
男女20名を対象に、5母音・鼻音の平均認識率
をフレーム単位で評価、比較した結果を第2表に
示す。フレーム認識率を%で示し、( )で認識
率のバラツキを標準偏差で示す。
従来法の男女の区別なしに比べ、本方法による
男女の区別ありでは認識率は大きく向上し、バラ
ツキは減少して、本方法の有効性を示している。
また、前記標準パタン群が選択されることによ
つて、対となるCMを選択することができ、選択
されたCMを用いることによつて高い単語認識率
を得ることができる。CMを男女別に作成した場
合と男女共通に作成した場合の単語認識率の比較
を第2表に示す。40人の212単語で評価した。[Table] In other words, if four words are used, up to 98 out of 100 people can correctly select the standard pattern group. One person who made a mistake mistook a female voice for a man's voice, but the nature of this speaker's voice is between that of a man's voice and that of a woman's voice, so even if he makes a mistake, the rate of vowel/nasal sound recognition will be extremely small. In this way, by using this method, it is possible to select the optimal standard pattern group with high accuracy. Table 2 shows the results of evaluating and comparing the average recognition rate of the five vowels and nasal sounds on a frame-by-frame basis for 20 men and women. The frame recognition rate is shown in %, and the variation in recognition rate is shown in parentheses as standard deviation. Compared to the conventional method, which does not distinguish between men and women, the recognition rate of the present method with the distinction of men and women is greatly improved and the variation is reduced, demonstrating the effectiveness of the present method. Further, by selecting the standard pattern group, a paired CM can be selected, and a high word recognition rate can be obtained by using the selected CM. Table 2 shows a comparison of word recognition rates when commercials were created for both men and women and when commercials were created for both men and women. It was evaluated using 212 words written by 40 people.
【表】
すなわち、その入力環境又は話者の種類に最も
合うCMを用いることによつて平均認識率は向上
し、単語認識率の話者によるバラツキも減少し、
安定な結果が得られる。
発明の効果
以上述べたように本発明は、入力環境や話者の
種類によつて標準パタン群を別々に用意し、同時
に同一条件にてCMを作成して前記標準パタン群
と対にして用意しておき、未知入力音声を用いて
その音声に最も適した標準パタン群およびCMを
自動選択する機能を持たせることにより、話者に
負担をかけることなく、最も適した標準パタン群
およびCMを用いて音声を認識することができ、
種々の入力環境、話者の種類に対して安定した高
い単語認識率を得ることが可能となるという利点
を有する。[Table] In other words, by using the commercial that best matches the input environment or type of speaker, the average recognition rate improves, and the variation in word recognition rate among speakers decreases.
Stable results are obtained. Effects of the Invention As described above, the present invention separately prepares standard pattern groups depending on the input environment and type of speaker, simultaneously creates commercials under the same conditions, and prepares them in pairs with the standard pattern group. By providing a function that automatically selects the most suitable standard pattern group and CM for the unknown input voice using unknown input voice, it is possible to select the most suitable standard pattern group and CM without putting a burden on the speaker. can be used to recognize speech,
This method has the advantage that it is possible to obtain a stable and high word recognition rate for various input environments and types of speakers.
第1図は本発明の一実施例における音声認識装
置の機能ブロツク図、第2図は本発明の処理手順
を示すフローチヤート、第3図は従来例の音声認
識装置の機能ブロツク図である。
5……類似度計算部、8……メインプロセツ
サ、9……格納部、10……選択部、11……単
語マツチング部、12……単語辞書、55……類
似度計算部、59……格納部、60……コンフユ
ージヨンマトリクス、61……単語マツチング
部。
FIG. 1 is a functional block diagram of a speech recognition device according to an embodiment of the present invention, FIG. 2 is a flowchart showing a processing procedure of the present invention, and FIG. 3 is a functional block diagram of a conventional speech recognition device. 5...Similarity calculation unit, 8...Main processor, 9...Storage unit, 10...Selection unit, 11...Word matching unit, 12...Word dictionary, 55...Similarity calculation unit, 59... . . . Storage section, 60 . . . Confusion matrix, 61 . . . Word matching section.
Claims (1)
部と、入力音声を音素毎に分割するセグメンテー
シヨン部と、互いに対をなす音素標準パタン群と
コンフユージヨンマトリクスを複数組格納する1
個又は複数個の格納部と、前記音素標準パタン群
を用いて前記特徴パラメータの類似度を算出する
類似度計算部と、前記セグメンテーシヨン部と類
似度計算部の結果を用いて音素系列を作成し、そ
の中の母音・鼻音の中心位置を求め、かつ装置全
体の制御を行うメインプロセツサと、前記メイン
プロセツサの処理結果を用いて入力音声に最も類
似した標準パタン群を自動選択する選択部と、音
素系列で表記された単語辞書部と、前記選択され
た標準パタン群の対となるコンフユージヨンマト
クリスを用いて単語辞書部の内容との類似度を算
出する単語マツチング部とを少なくとも具備する
音声認識装置。 2 前記選択部における標準パタン群の選択と、
最大類似度となる標準パタンを得る回数の総和を
標準パタン群毎に求め、最大回数と2番目に大き
い回数の比又は差があらかじめ定めた閾値を超え
た時点で打切ることによつて行うことを特徴とす
る特許請求の範囲第1項記載の音声認識装置。 3 前記格納部の標準パタン群とコンフユージヨ
ンマトリクスの対が、同一の音声入力環境又は同
一種類の性別、年令範囲によつて作成したもので
あることを特徴とする特許請求の範囲第1項記載
の音声認識装置。[Scope of Claims] 1. A feature extraction unit that obtains feature parameters from input speech, a segmentation unit that divides input speech into phonemes, and a plurality of pairs of phoneme standard pattern groups and confusion matrices that are stored. Do 1
one or more storage units, a similarity calculation unit that calculates the similarity of the feature parameters using the phoneme standard pattern group, and a phoneme sequence using the results of the segmentation unit and similarity calculation unit. a main processor that determines the center position of the vowels and nasals in the voice and controls the entire device; and automatically selects a group of standard patterns that are most similar to the input voice using the processing results of the main processor. a selection section, a word dictionary section written in a phoneme sequence, and a word matching section that calculates the degree of similarity between the content of the word dictionary section using a confusion matrix that is a pair of the selected standard pattern group. A voice recognition device comprising at least the following. 2. Selection of a standard pattern group in the selection section;
Calculate the total number of times a standard pattern with the maximum similarity is obtained for each standard pattern group, and terminate when the ratio or difference between the maximum number and the second largest number exceeds a predetermined threshold. A speech recognition device according to claim 1, characterized in that: 3. Claim 1, characterized in that the pair of the standard pattern group and the confusion matrix in the storage section are created under the same voice input environment or the same type of gender and age range. Speech recognition device described in section.
Priority Applications (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP59269755A JPS61147299A (en) | 1984-12-20 | 1984-12-20 | Voice recognition equipment |
Applications Claiming Priority (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP59269755A JPS61147299A (en) | 1984-12-20 | 1984-12-20 | Voice recognition equipment |
Publications (2)
| Publication Number | Publication Date |
|---|---|
| JPS61147299A JPS61147299A (en) | 1986-07-04 |
| JPH053596B2 true JPH053596B2 (en) | 1993-01-18 |
Family
ID=17476697
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| JP59269755A Granted JPS61147299A (en) | 1984-12-20 | 1984-12-20 | Voice recognition equipment |
Country Status (1)
| Country | Link |
|---|---|
| JP (1) | JPS61147299A (en) |
Family Cites Families (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JPS56119199A (en) * | 1980-02-26 | 1981-09-18 | Sanyo Electric Co | Voice identifying device |
| JPS57104193A (en) * | 1980-12-19 | 1982-06-29 | Matsushita Electric Industrial Co Ltd | Voice recognizer |
| JPS5872996A (en) * | 1981-10-28 | 1983-05-02 | 電子計算機基本技術研究組合 | Word voice recognition |
| JPS58129497A (en) * | 1982-01-28 | 1983-08-02 | 電子計算機基本技術研究組合 | Word voice recognition |
-
1984
- 1984-12-20 JP JP59269755A patent/JPS61147299A/en active Granted
Also Published As
| Publication number | Publication date |
|---|---|
| JPS61147299A (en) | 1986-07-04 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US10176811B2 (en) | Neural network-based voiceprint information extraction method and apparatus | |
| Kamble et al. | Emotion recognition for instantaneous Marathi spoken words | |
| Hunt | Speaker adaptation for word‐based speech recognition systems | |
| JP2980382B2 (en) | Speaker adaptive speech recognition method and apparatus | |
| JPH0455518B2 (en) | ||
| JPH0252278B2 (en) | ||
| Chakraborty et al. | Speech recognition of isolated words using a new speech database in sylheti | |
| JPS61147299A (en) | Voice recognition equipment | |
| KR20010000054A (en) | Dtw based isolated-word recognization system employing voiced/unvoiced/silence information | |
| JPS60164800A (en) | Voice recognition equipment | |
| JPH0424720B2 (en) | ||
| JPS6336678B2 (en) | ||
| JPH0619497A (en) | Speech recognition method | |
| JPS60147797A (en) | Voice recognition equipment | |
| JPH042197B2 (en) | ||
| JP2000242292A (en) | Speech recognition method, apparatus for implementing the method, and storage medium storing program for executing the method | |
| JPS62133499A (en) | voice recognition device | |
| JPS62111295A (en) | Voice recognition equipment | |
| JPH03201161A (en) | Sound recognizing device | |
| JPH04220699A (en) | Voice recognition method | |
| JPS5958498A (en) | Voice recognition equipment | |
| JPH0344320B2 (en) | ||
| HK1248396B (en) | Voiceprint information extraction method based on neural network and apparatus thereof | |
| JPS6069694A (en) | Segmentation of head consonant | |
| JPS6247100A (en) | Voice recognition equipment |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| EXPY | Cancellation because of completion of term |