JPH0225899A - Voice recognizing device - Google Patents

Voice recognizing device

Info

Publication number
JPH0225899A
JPH0225899A JP63176704A JP17670488A JPH0225899A JP H0225899 A JPH0225899 A JP H0225899A JP 63176704 A JP63176704 A JP 63176704A JP 17670488 A JP17670488 A JP 17670488A JP H0225899 A JPH0225899 A JP H0225899A
Authority
JP
Japan
Prior art keywords
dictionary
orthogonalized
orthogonal
axis
learning
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
JP63176704A
Other languages
Japanese (ja)
Inventor
Tsuneo Nitta
恒雄 新田
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Toshiba Corp
Original Assignee
Toshiba Corp
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Toshiba Corp filed Critical Toshiba Corp
Priority to JP63176704A priority Critical patent/JPH0225899A/en
Publication of JPH0225899A publication Critical patent/JPH0225899A/en
Pending legal-status Critical Current

Links

Abstract

PURPOSE:To easily prepare an orthogonalized dictionary to deal with a speaker by making the dictionary of an orthogonal axis obtained first from the learning patterns of plural speakers into a basis and successively dictionary-registering an orthogonal vector according to the learning pattern obtained from the individual speaker. CONSTITUTION:First, according to the learning patterns to be respectively obtained from plural speakers, the orthogonal axis made into the basis is decided, a first dictionary is prepared, and it is dictionary-registered to an orthogonalized dictionary 9. Next, as to a dictionary preparation according to the learning pattern obtained from the individual speaker, the dictionary of the axis to be respectively orthogonal to each axis of the dictionary to be already registered is obtained, whether to register the dictionary or not is successively decided, it is additionally registered to the orthogonalized dictionary 9, and thereby, the orthogonalized dictionary 9 is constructed. Thus, the orthogonalized dictionary with a high recognizing performance can be efficiently prepared.

Description

【発明の詳細な説明】 [発明の目的] (産業上の利用分野) 本発明は特定のグループ構成員により共通に利用でき、
上記グループ構成員からの少ない学習パターンで高い認
識性能を得ることのできる音声認識装置に関する。
[Detailed Description of the Invention] [Object of the Invention] (Field of Industrial Application) The present invention can be commonly used by members of a specific group,
The present invention relates to a speech recognition device that can obtain high recognition performance with a small number of learning patterns from group members.

(従来の技術) 音声による情報の入出力は人間にとって自然性が高く、
マン◆マシン・インターフェースとして優れた利点を有
することから従来より種々研究されている。現在、実用
化されている音声認識装置の殆んどは単語音声を認識す
る方式のもので、一般的には第4図に示すように構成さ
れている。
(Conventional technology) Inputting and outputting information through voice is highly natural for humans;
Since it has excellent advantages as a man-machine interface, it has been studied in various ways. Most of the speech recognition devices currently in practical use are of a type that recognizes word speech, and are generally configured as shown in FIG.

この装置は発声入力された音声を電気信号に変換して取
込み、バンド・バス・フィルタ等からなる音響分析部1
にて音響分析し、始端・終端検出部2にてその単語音声
区間を検出する。そして入力音声の上記単語音声区間の
音響分析データ(特徴情報;音声パターン)と、標準パ
ターン辞書3に予め登録されている認識対象単語の各標
準パターンとの類似度や距離等をパターン・マツチング
部4にて計算し、その計算結果を判定部5にて判定して
、例えば類似度値の最も高い標準パターンのカテゴリ名
を前記入力音声に対する認識結果として求めるものとな
っている。
This device converts the voice input into an electrical signal and captures it.
The sound analysis section 2 performs acoustic analysis, and the start/end detection section 2 detects the speech section of the word. Then, a pattern matching unit calculates the similarity and distance between the acoustic analysis data (feature information; speech pattern) of the word speech section of the input speech and each standard pattern of recognition target words registered in advance in the standard pattern dictionary 3. 4, and the result of the calculation is determined by the determining unit 5, to obtain, for example, the category name of the standard pattern with the highest similarity value as the recognition result for the input voice.

しかしこのようにパターン・マツチング法による音声認
識では入力音声パターンと予め登録されている標準パタ
ーンとの時間軸方向のずれ(パターン変形)が問題とな
る。そこで従来では、専ら線形伸縮や動的計画法(DP
)に代表される非線形伸縮等により、上述した時間軸方
向のずれに対する課題を解消している。
However, in speech recognition using the pattern matching method, a problem arises in that the input speech pattern and the standard pattern registered in advance are misaligned in the time axis direction (pattern deformation). Therefore, in the past, only linear stretching and dynamic programming (DP) were used.
) The above-mentioned problem with the deviation in the time axis direction is solved by non-linear expansion and contraction as typified by.

一方、このようなパターン・マツチング法とは別に、予
め収集された学習パターンから直交化辞書を作成し、こ
の直交化辞書を用いて音声認識する方式(部分空間法)
が提唱されている。この方式は第5図にその構成例を示
すように、音響分析されて音声区間検出された音声パタ
ーンから、標本点抽出部6にて上記音声区間を等分割し
た所定点数の標本点を抽出して[特徴ベクトルの数X標
本点数]で示される標本パターンを求める。このような
標本パターンを認識対象とするカテゴリ毎に所定数ずつ
収集してパターン蓄積部7に格納する。そしてグラム・
シュミツI−(GS)直交化部8において、上記パター
ン蓄積部7に収集された所定数(3個以上)の標本パタ
ーンを用い、以下に示す手順で直交化辞書9を作成する
On the other hand, apart from such a pattern matching method, there is a method (subspace method) in which an orthogonalized dictionary is created from learning patterns collected in advance and speech recognition is performed using this orthogonalized dictionary.
has been proposed. As shown in FIG. 5, an example of the configuration of this method is such that a sample point extraction unit 6 extracts a predetermined number of sample points obtained by equally dividing the speech section from a speech pattern that has been acoustically analyzed and detected a speech section. Then, a sample pattern represented by [number of feature vectors x number of sample points] is obtained. A predetermined number of such sample patterns are collected for each category to be recognized and stored in the pattern storage section 7. And Gram
In the Schmidts I-(GS) orthogonalization unit 8, an orthogonalization dictionary 9 is created using the predetermined number (three or more) of sample patterns collected in the pattern storage unit 7 in the following procedure.

即ち、上記直交化辞書9の作成は、各カテゴリ毎にその
カテゴリのm回目の学習パターンをal、lとし、3回
発声された学習パターンを用いる場合には、 ■ 1回目の学習データa1を第1軸の辞書b1とし、 b 1−a t               ・・・
(1)これを直交化辞書9に登録する。
That is, to create the above-mentioned orthogonalized dictionary 9, for each category, the m-th learning pattern of that category is set as al, l, and when using the learning pattern that has been uttered three times, ■ the first learning data a1 is Let the dictionary b1 be the first axis, and b 1-a t...
(1) Register this in the orthogonalization dictionary 9.

■ 2回目の学習データa2からグラム・シュミットの
直交化式を用い、 なる計算を行い、1lb211が一定値より大きい場合
、これを第2軸の辞書b2として前記直交化辞書9に登
録する。但し、(・)は内積、1111はノ゛ルムを示
す。
(2) Using the Gram-Schmidt orthogonalization formula from the second learning data a2, perform the following calculation, and if 1lb211 is larger than a certain value, register this in the orthogonalization dictionary 9 as the second axis dictionary b2. However, (.) indicates the inner product, and 1111 indicates the norm.

■ そして3回目の学習データa3から、なる計算を行
い、1lb311が一定値より大きい場合、これを第3
軸の辞書b3として前記直交化辞書9に登録する。但し
、第2軸の辞書が求められていない場合には、上記(2
)式の計算を行う。
■ Then, from the third learning data a3, perform the following calculation, and if 1lb311 is larger than a certain value, use this as the third
It is registered in the orthogonalization dictionary 9 as the axis dictionary b3. However, if the second axis dictionary is not required, the above (2)
) calculates the formula.

以上の■〜■の処理を各カテゴリについて繰返し実行し
て直交化辞書9を予め形成しておく。
The orthogonalized dictionary 9 is formed in advance by repeatedly performing the above processes (1) to (2) for each category.

類似度計算部lOは上述した如く作成された直交化部@
9と、人力音声パターンXとの間でとして、カテゴリi
の直交化辞書b  との間の工・r 類似度を計算するものである。これらの各カテゴリiに
ついて求められた類似度値に従って上記入力音声パター
ンXが認識される。尚、上記カテゴリiの直交化辞書b
  は予め正規化されたもの1、「 であり、K1はカテゴリiの辞書の個数(軸数)を示し
ている。
The similarity calculation unit IO is an orthogonalization unit created as described above.
9 and human voice pattern X, category i
This is to calculate the degree of similarity between the orthogonalized dictionary b and the orthogonalized dictionary b. The input speech pattern X is recognized according to the similarity value determined for each of these categories i. In addition, the orthogonalized dictionary b of the above category i
is pre-normalized 1, ``, and K1 indicates the number of dictionaries (number of axes) of category i.

このようなGS直交化を用いることにより、その認識性
能の大幅な向りが図られている。また微分フィルタを用
いて時間軸方向および周波数方向の変動を吸収した直交
化辞書を作成し、更にその認識性能の向上を図ることも
試みられている。
By using such GS orthogonalization, the recognition performance is greatly improved. Furthermore, attempts have been made to create orthogonalized dictionaries that absorb fluctuations in the time and frequency directions using differential filters, and to further improve their recognition performance.

ところがこの種の装置にあっては、専ら特定の話者に対
して標準音声辞書の作成が行なわれる。
However, in this type of device, a standard speech dictionary is created exclusively for a specific speaker.

この為、別の話者が上記音声認識装置を利用17ようと
する場合には、その都度、音声辞書を変更する必要が生
じた。そこで多数の話者から数多くの学習パターンを収
集して直交化辞書を作成することが考えられているが、
その辞書作成が徒に複雑化し、認識性能の高い辞書を得
ることが困難化する等の不具合が生じた。
For this reason, when another speaker attempts to use the speech recognition device 17, it is necessary to change the speech dictionary each time. Therefore, it has been considered to collect many learning patterns from many speakers and create an orthogonal dictionary.
Problems such as creating a dictionary became unnecessarily complicated and making it difficult to obtain a dictionary with high recognition performance occurred.

(発明が解決しようとする問題点) このように従来の直交化辞書を用いた部分空間法による
音声認識にあっては、複数の話者から収集された学習パ
ターンから如何にして性能の高い直交化辞書を効率良く
作成するかと云う点で課迦が残されている。また直交化
辞書の作成に必要な複数の話者の学習パターンを如何に
して効率良く収集し、直交化辞書を作成するかと云う点
でも問題があった。
(Problems to be Solved by the Invention) In speech recognition using the subspace method using conventional orthogonalized dictionaries, it is difficult to determine how to obtain high-performance orthogonal dictionaries from learning patterns collected from multiple speakers. There remains work to be done on how to efficiently create a dictionary. There is also a problem in how to efficiently collect the learning patterns of multiple speakers necessary for creating an orthogonal dictionary and create an orthogonal dictionary.

本発明はこのような事情を考慮してなされたもので、そ
の目的とするところは、複数の話者から収集される少な
い学習パターンにて認識性能の高い直交化辞書を効率的
に作成I2、複数の利用者にて共通に利用可能な認識性
能の高い音声認識装置を提供することにある。
The present invention has been made in consideration of these circumstances, and its purpose is to efficiently create an orthogonalized dictionary with high recognition performance using a small number of learning patterns collected from a plurality of speakers. An object of the present invention is to provide a speech recognition device with high recognition performance that can be commonly used by a plurality of users.

[発明の構成] (問題点を解決するための手段) 本発明は入力音声を分析処理して求められる入力音声パ
ターンと、予め収集された学習パターンに基いて作成さ
れて1.する直交化辞書との間で類似度を計算して上記
入力音声を認識する音声認識装置において、 上記直交化辞書として複数話者の学習パターンから基本
となる直交軸を決定して基準となる辞書を作成して辞書
登録した後、個別話者の学習パターンからの辞書作成に
ついては、既に登録されている辞書の幀ど直交する新た
な軸を決定しながら、この新たな軸の辞書を追加辞書登
録するか否かを、例えばその軸のノルムの値から判定し
、前記直交化辞書を順次構築していくようにしたことを
特徴とするものである。
[Structure of the Invention] (Means for Solving the Problems) The present invention is constructed based on input speech patterns obtained by analyzing input speech and learning patterns collected in advance. In the speech recognition device that recognizes the input speech by calculating the similarity between the input speech and the orthogonalized dictionary, the orthogonalized dictionary determines a basic orthogonal axis from the learning patterns of multiple speakers and serves as a reference dictionary. After creating a dictionary and registering it in a dictionary, when creating a dictionary from the learning patterns of individual speakers, determine a new axis that is orthogonal to the width of the already registered dictionary, and add the dictionary of this new axis to the dictionary. The feature is that whether or not to register is determined based on, for example, the value of the norm of the axis, and the orthogonalized dictionary is sequentially constructed.

(作用) 本発明によれば、複数の話者からそれぞれ求められた学
習パターンから直交化辞書の基本となる直交軸が決定さ
れて辞書の作成が行なわれ、その辞書登録がなされた後
、個別話者からの学習パターンに基づく辞書作成に際し
ては、既に作成されて辞書登録されている辞書の軸と直
交する軸が求められ、この新たな軸についての辞書が上
記個別話者の学習パターンから求められる。そしてその
ノルムの値を判定することによ・)で辞書に追加登録す
るか否かが調べられ、辞書として有用な場合にのみ前記
直交化辞書・\の追加辞書登録が行なわれる。
(Operation) According to the present invention, orthogonal axes, which are the basis of an orthogonalized dictionary, are determined from learning patterns obtained from a plurality of speakers, a dictionary is created, and after the dictionary is registered, individual When creating a dictionary based on learning patterns from speakers, an axis that is perpendicular to the axis of the dictionary that has already been created and registered in the dictionary is found, and a dictionary about this new axis is created from the learning patterns of the individual speakers. It will be done. Then, by determining the value of the norm, it is checked whether or not it should be additionally registered in the dictionary, and the orthogonalized dictionary is additionally registered in the dictionary only when it is useful as a dictionary.

この結果、複数の話者の学習パターンから、そのパター
ン変動要素を効率良く表現した直交化辞書を構築してい
くことが可能となり、認識性能の高い直交化辞書を得る
ことが可能となる。しかも基本となる直交軸の辞書に対
して、個別話者の変動パターンを直交ベクトルの組に効
率良く組入れて辞書表現することが可能となるので、そ
の計算量を少なくし、簡易に効率良く辞書を作成してい
くことが可能となる。
As a result, it becomes possible to construct an orthogonalized dictionary that efficiently expresses the pattern variation elements from the learning patterns of a plurality of speakers, and it becomes possible to obtain an orthogonalized dictionary with high recognition performance. Moreover, since it is possible to efficiently incorporate the variation patterns of individual speakers into a set of orthogonal vectors and express the dictionary in relation to the basic dictionary of orthogonal axes, the amount of calculation can be reduced and the dictionary can be easily and efficiently used. It becomes possible to create.

(実施例) 以下、図面を3照して本発明の一実施例につき説明する
(Example) Hereinafter, one example of the present invention will be described with reference to the drawings.

第1図は本発明の一実施例に係る音声認識装置の概略構
成図で、第5図に示した従来装置と同一部分には同一符
号を付して示しである。
FIG. 1 is a schematic configuration diagram of a speech recognition device according to an embodiment of the present invention, and the same parts as those of the conventional device shown in FIG. 5 are denoted by the same reference numerals.

この実施例装置が特徴とするところは、パターン蓄積部
7に蓄積された学習パターンを用いて直交化辞書9を作
成する手段として、直交ベクトル計算部8a、直交ベク
トル登録判定部8b、および残差ノルムメモリ8cとか
らなる直交化辞書作成部8を設け、第2図にこの直交化
辞書作成部8における処理概念を模式的に示すように、
先ず複数の話者からそれぞれ求められた学習パターンに
従って、基本となる直交軸を決定して最初の辞書を作成
して直交化部w9に辞書登録した後、個別話者から求め
られる学習パターンに従う辞書作成については、上記基
本軸に直交する軸(既に登録されている辞書の各軸にそ
れぞれ直交する軸)の辞書を求め、この辞書を登録する
か否かを逐次判定しながら前言コ直交化辞書9に追加登
録して行くことで、認識性能の高い直交化辞書9を構築
していくようにした点を特徴としている。
This embodiment device is characterized by an orthogonal vector calculation unit 8a, an orthogonal vector registration determination unit 8b, and a residual An orthogonalized dictionary creation unit 8 is provided which includes a norm memory 8c, and the processing concept in this orthogonalized dictionary creation unit 8 is schematically shown in FIG.
First, the basic orthogonal axes are determined according to the learning patterns obtained from a plurality of speakers, the first dictionary is created, and the dictionary is registered in the orthogonalization unit w9, and then the dictionary is created according to the learning patterns obtained from each individual speaker. To create a dictionary, find a dictionary with axes orthogonal to the basic axis (orthogonal to each axis of the already registered dictionaries), and create an orthogonalized dictionary by sequentially determining whether to register this dictionary or not. The feature is that an orthogonalized dictionary 9 with high recognition performance is constructed by adding additional registrations to the dictionary 9.

この直交化辞書作成部8における直交化辞書の作成につ
いて、第3図に示す処理手続きに従って更に詳しく説明
する。
The creation of the orthogonalized dictionary in the orthogonalized dictionary creation section 8 will be explained in more detail according to the processing procedure shown in FIG.

尚、ここではパターン蓄積部7に収集される学習パター
ンとしては、例えばj  (−1,2,〜8)で示され
る6点の音響分析された特徴ベクトルからなり、その音
声区間をk (−0,1,2,〜l l)とし5て11
等分する12個の標本点に亙って採取したデータ系列と
して与えられるものとして説明する。
In this case, the learning pattern collected in the pattern storage unit 7 is composed of six acoustically analyzed feature vectors indicated by, for example, j (-1, 2, ~8), and the speech interval is defined as k (- 0,1,2,~l l) and 5 and 11
The explanation will be given assuming that it is given as a data series collected over 12 equally divided sampling points.

前記直交化辞書作成部8は、先ず辞書登録対象とするカ
テゴリiについて複数(L人)の話者からそれぞれ3個
づつ学習パターンを収集する(ステップa)。しかる後
、これらの複数話者からそれぞれ収集した学習パターン
中のm番目([T1−1゜2.3.〜M ; M −3
x L)の学習パターンをam(j、k)とし、たとき
、基本となる直交化辞書9を次のようにして作成してい
る。
The orthogonalized dictionary creation unit 8 first collects three learning patterns from each of a plurality of (L) speakers for category i to be registered in the dictionary (step a). After that, the m-th learning pattern ([T1-1゜2.3.~M; M-3
x L) is defined as am(j, k), then the basic orthogonalization dictionary 9 is created as follows.

■ 先ず、カテゴリiの学習パターンam(j−k)か
ら、その平均パターンA   を (j、k) [j−1,2,〜1B、に−0,1,2,〜17]とし
て求める(ステップb)。
■ First, from the learning pattern am(j-k) of category i, find its average pattern A as (j, k) [j-1, 2, ~1B, -0, 1, 2, ~17] ( Step b).

■ し5かる後、上述した如くして求めた平均パターン
A(j、k)を用いて、 −A      +2*A     +Abl(j、k
)   (j、に−1)     (j、k)    
(j、に+1)[j=1,2.〜16.   k−L2
.〜16コ             ・・・(6)な
る演算にて第1軸の辞書bl(j、k)を求め(ステッ
プc)、これを直交化辞書9に登録する(ステップd)
。この辞書b   は前記平均パターン10、k) A(j、k)を時間軸方向に平滑化したものとして求め
られ、直交化辞書9の基準となる第1軸の辞書データと
なる。
■ After that, using the average pattern A(j, k) obtained as described above, -A +2*A +Abl(j, k
) (j, ni-1) (j, k)
(j,+1) [j=1,2. ~16. k-L2
.. ~16th... Find the dictionary bl(j, k) of the first axis by the calculation (6) (step c) and register it in the orthogonalization dictionary 9 (step d)
. This dictionary b is obtained by smoothing the average pattern 10, k) A(j, k) in the time axis direction, and becomes dictionary data on the first axis that serves as a reference for the orthogonalized dictionary 9.

■ I2かる後、前記平均パターンA(j=k)を用い
、謬−A     +A b2(j、k)    (j、に−1)   (j、に
+1)[j=1,2.〜1B、 k−1,2,〜16]
      ・・・(7)なる演算に゛C第2軸の辞書
b2(j、k)を求め(ステップe)、これを正規化し
た後に前記直交化辞書9に登録する(ステップf)。こ
の第2軸の辞書b2(j、k)は前記平均パターンA(
j、k)を時間軸方向に微分したものとしC求められる
■ After I2, using the average pattern A (j=k), calculate -A +A b2(j, k) (j, -1) (j, +1) [j=1,2. ~1B, k-1, 2, ~16]
. . . In the calculation (7), a dictionary b2(j, k) of the second axis of C is obtained (step e), and after being normalized, it is registered in the orthogonalization dictionary 9 (step f). This second axis dictionary b2(j,k) is the average pattern A(
C is obtained by differentiating j, k) in the time axis direction.

尚、このようにして4算される第2軸の辞書b2(j、
k)は、前記第1軸の辞書b1(j、k)に対して完全
には直交していないことから、 ”2U、k)”b2(j、k) (b2(j、k)   1(、i、k))bl(j、k
)争 b なる再直交化処理を施し、この再直交化された辞書デー
タB2(j、k)を正規化後、新たな第2軸の辞書b 
  として前記直交化辞書9に登録するよ2(j、k) うにしても良い。
Incidentally, the second axis dictionary b2(j,
k) is not completely orthogonal to the first axis dictionary b1(j,k), so "2U,k)"b2(j,k) (b2(j,k) 1( , i, k)) bl(j, k
) Conflict b After performing re-orthogonalization processing and normalizing this re-orthogonalized dictionary data B2 (j, k), a new second axis dictionary b
2(j,k) may be registered in the orthogonalization dictionary 9 as 2(j,k).

またここでは第2軸まで作成する例を示したが、更に2
次微分を行なう等して3軸以降の辞書を基本軸の直交化
辞書とし、て作成することも勿論可能である。
Also, here we have shown an example of creating up to the second axis, but there are also two
Of course, it is also possible to create dictionaries for the third and subsequent axes as orthogonalized dictionaries for the basic axes by performing second-order differentiation or the like.

■ しかる後、上述した如く求められた直交化辞書を基
本とし、直交ベクトル計算部8aに”C前記パターン蓄
積部7に格納されている複数の話者の個々の学習パター
ンを順に抽出しくステップg)、その学習パターンに従
って上記直交化辞書に直交する付加辞書を次のように1
7で作成する。
■ Thereafter, based on the orthogonalized dictionary obtained as described above, the orthogonal vector calculation section 8a sequentially extracts the individual learning patterns of a plurality of speakers stored in the pattern storage section 7 (step g). ), and according to the learning pattern, an additional dictionary that is orthogonal to the above orthogonalized dictionary is created as follows.
Create with 7.

即ち、この付加辞書の作成は、前記パターン蓄積部7に
収集された学習パターンal(j、k)について、既に
求められている直交化辞書の軸数をPとしたとき [n = 1.2.〜p  、  m −1,2,〜M
lなるグラムシュミットの直交化式を演算して行われる
(ステップh)。そしてこの新しく求められた個々の話
者の特徴的変動を表現する直交ベクトル(付加辞書)b
  を直交ベクトル登録判定部P十踵 8bに4え、そのノルムllb   IIが所定値より
もhi 大きいか否かを判定す、る(ステップi)。そしてその
ノルム値が所定値よりも大きい場合、これを付加辞書と
してパターン正規化処理を施した後に前記直交化辞書9
に登録する(ステップj)。この際、上記ノルムllb
   11の値を残差ツルムチpm −ツル8cに登録する(ステップk)。
That is, the creation of this additional dictionary is based on the learning pattern al(j, k) collected in the pattern storage section 7, where P is the number of axes of the orthogonalized dictionary already obtained [n = 1.2 .. ~p, m-1,2,~M
This is performed by calculating the Gram-Schmidt orthogonalization formula l (step h). Then, the orthogonal vector (additional dictionary) b expressing this newly found characteristic variation of each individual speaker
is inputted to the orthogonal vector registration determination unit P 8b, and it is determined whether the norm llbII is greater than a predetermined value (step i). If the norm value is larger than a predetermined value, this is used as an additional dictionary to perform pattern normalization processing, and then the orthogonalized dictionary 9
(Step j). In this case, the above norm llb
The value of 11 is registered in the residual error pm - 8c (step k).

以上の■〜■の処理を複数の話者から求められた個々の
学習パターン毎に繰返し実行することによってカテゴリ
iについての直交化辞書9が作成される。
The orthogonalized dictionary 9 for category i is created by repeatedly performing the above processes 1 to 2 for each individual learning pattern obtained from a plurality of speakers.

尚、新たに求められた軸の辞書の前記直交化部lf9へ
の登録に際しては、直交化辞書9として予め定められて
いる軸数を越えることがある。このような場合、新たな
軸の辞書登録を中止すると、その辞書を得た話者に対す
る認識性能が劣化する虞れがある。そこでこのような場
合には、前記残差ノルムメモリ8cからそのカテゴリi
についての各軸での残差ノルムllb   IIをそれ
ぞれ読出し、P++11 新たな軸の残差ノルムの値と比較する。そして既に登録
された辞書の中で、その残差ノルムの値が小さいものが
あれば、その残差ノルムに対応する辞書(直交ベクトル
)を前記直交化辞書9から抹消し、代わりに前述した新
しく求められた辞書(直交ベクトル)を辞書登録する。
Note that when registering a dictionary of newly obtained axes in the orthogonalization unit lf9, the number of axes that is predetermined as the orthogonalization dictionary 9 may be exceeded. In such a case, if dictionary registration of a new axis is stopped, there is a risk that the recognition performance for the speaker who obtained the dictionary will deteriorate. Therefore, in such a case, the category i is stored from the residual norm memory 8c.
The residual norm llb II on each axis for P++11 is read out and compared with the value of the residual norm on the new axis. If there is a dictionary whose residual norm value is small among the already registered dictionaries, the dictionary (orthogonal vector) corresponding to that residual norm is deleted from the orthogonalized dictionary 9 and replaced with the above-mentioned new dictionary. The obtained dictionary (orthogonal vector) is registered in the dictionary.

この場合、残差ノルムメモリ8cにおける対応ノルムの
値も書替えることは勿論のことである。
In this case, it goes without saying that the value of the corresponding norm in the residual norm memory 8c is also rewritten.

以l二のよ・うにして複数の話者の学習パターンから最
初に求められる直交軸の辞書を基本として、個々の話者
から求められる学習パターンに従う直交ベクトルを順次
辞書登録して直交化辞書9を構築していく。この結果、
一定の人数範囲内であれば、その全ての登録話者の入力
音声パターンに対して認識性能の高い直交化辞書9を得
ることが可能となり、その認識性能の向」ニを図ること
が可能となる。
Based on the dictionary of orthogonal axes that is first obtained from the learning patterns of multiple speakers as described below, orthogonal vectors according to the learning patterns obtained from each individual speaker are sequentially registered in the dictionary to create an orthogonalized dictionary. 9 will be built. As a result,
If the number of speakers is within a certain range, it is possible to obtain an orthogonalized dictionary 9 with high recognition performance for the input speech patterns of all registered speakers, and it is possible to improve the recognition performance. Become.

まh−上述したように簡単な演算処理によって新たな軸
の辞書を逐次作成していくので、その処理負担が非常に
軽く、複数の話者に適応し得る直交化辞書9を効率的に
作成することが可能となる等の効果が奏せられる。
As mentioned above, new axis dictionaries are created one after another through simple arithmetic processing, so the processing load is very light, and the orthogonalized dictionary 9 that can be adapted to multiple speakers can be created efficiently. Effects such as making it possible to do the following can be achieved.

次表は男性5名2女性3名から数字音声と人名からなる
30語の音声データをそれぞれ13回に亙って収集し、
そのうちの3回分を学習用、残り10回分を認忠性能評
価に用いた実験例を示すものである。
The following table shows the voice data of 30 words consisting of numeric sounds and personal names collected from 5 men, 2 women, and 3 women over 13 times each.
This shows an experimental example in which 3 of the tests were used for learning and the remaining 10 tests were used for recognition performance evaluation.

表 尚、この表における話者Aは比較的性能の悪い話者であ
り、話者Bは性能の良い話者である。またこれらの結果
は、10名の話者の全てが辞書登録を終えた時点での直
交化辞書セ・ノドを用いたときの認識性能を示している
。尚、参考として上記話者A、Bが単独で、所謂特定話
者で直交化辞書(4軸)を作成したときの認識性能は、
それぞれ92.5%、 98.3%であった。
Note that speaker A in this table is a speaker with relatively poor performance, and speaker B is a speaker with good performance. These results also show the recognition performance when using the orthogonalized dictionary SENODO at the time when all 10 speakers have completed dictionary registration. For reference, the recognition performance when the above speakers A and B create an orthogonalized dictionary (4 axes) with so-called specific speakers is as follows.
They were 92.5% and 98.3%, respectively.

この実験データに示されるように、本方式によれば10
名程度の登録話者に対して上述した如く直交化辞書9を
作成することで、その登録順序に拘らず全ての登録話者
に対して安定に、また比較的高い性能で音声認識し得る
ことが明らかとなった。
As shown in this experimental data, according to this method, 10
By creating the orthogonalized dictionary 9 as described above for registered speakers of approximately 500,000 registered speakers, it is possible to perform speech recognition stably and with relatively high performance for all registered speakers regardless of the registration order. became clear.

尚、本発明は上述17た実施例に限定されるものではな
い。ここでは最初に複数の登録話者から2軸の直交化辞
書を作成する例について説明lまたが、更に多くの軸数
の基本直交化辞書を作成することも可能である。この場
合、直交化フィルタの係数としては幾つかのバリエーシ
ョンが考えられるが、要は学習パターンを平滑、1次微
分、2次微分。
Note that the present invention is not limited to the 17 embodiments described above. First, an example will be described in which a two-axis orthogonalized dictionary is created from a plurality of registered speakers, but it is also possible to create a basic orthogonalized dictionary with an even larger number of axes. In this case, several variations can be considered for the coefficients of the orthogonalization filter, but the key is to smooth the learning pattern, differentiate it to the first order, and differentiate it to the second order.

・・・すれば良いものであり、種々変形して実施するこ
とができる。また学習パターンの次元数等も特に限定さ
れるものでもない。更には新たに作成する辞書の軸数も
学習パターン数に応じて定めれば良く、グラムシュミッ
ト以外の直交化法を用いて辞書を作成することも可能で
ある。その他、本発明はその要旨を逸脱しない範囲で変
形して実施可能である。
. . . and can be implemented with various modifications. Furthermore, the number of dimensions of the learning pattern is not particularly limited either. Furthermore, the number of axes of a newly created dictionary may be determined according to the number of learning patterns, and it is also possible to create a dictionary using an orthogonalization method other than Gram-Schmidt. In addition, the present invention can be modified and implemented without departing from the gist thereof.

[発明の効果] 以上説明したように本発明によれば複数の話者から収集
した学習パターンを用いて、これらの話者に対応可能な
直交化辞書を簡易に、nつ性能良く生成していくことが
可能なので、少ない学習パターンでパターンの変動を効
果的に表現した辞書を得ることができ、その認識性能の
向上を図り得る等の実用上多大なる効果を奏する。
[Effects of the Invention] As explained above, according to the present invention, learning patterns collected from a plurality of speakers are used to easily generate n orthogonalized dictionaries compatible with these speakers with good performance. Therefore, it is possible to obtain a dictionary that effectively expresses pattern variations with a small number of learning patterns, and this has great practical effects such as improving recognition performance.

【図面の簡単な説明】[Brief explanation of the drawing]

第1図は本発明の一実施例に係る音声認識装置の概略構
成図、第2図は実施例装置における直交化辞書作成の概
念を模式的に示す図、第3図は直交化辞書作成の処理手
続きの例を示す図、第4図および第5図はそれぞれ従来
の音声認識装置の概略構成を示す図である。 ■・・・音響分析部、2・・・始端・終端検出部、5・
・・判定部、6・・・標本点抽出部、7・・・パターン
蓄積部、8・・・直交化辞書作成部、9・・・直交化辞
書、10・・・類似度演算部、8a・・・直交ベクトル
計算部、8b・・・直交ベク トル登録判定部、 8c・・・残差ノルムメ そり。
FIG. 1 is a schematic configuration diagram of a speech recognition device according to an embodiment of the present invention, FIG. 2 is a diagram schematically showing the concept of orthogonal dictionary creation in the embodiment device, and FIG. 3 is a diagram showing the concept of orthogonal dictionary creation in the embodiment device. FIGS. 4 and 5 are diagrams showing an example of a processing procedure, and each is a diagram showing a schematic configuration of a conventional speech recognition device. ■...Acoustic analysis section, 2...Start/end detection section, 5.
. . . Judgment unit, 6 . . Sample point extraction unit, 7 . ... Orthogonal vector calculation unit, 8b... Orthogonal vector registration determination unit, 8c... Residual norm measurement.

Claims (1)

【特許請求の範囲】 入力音声を分析処理して求められる入力音声パターンと
、予め収集された複数話者の学習パターンに基いて作成
されている直交化辞書との間で類似度を計算して上記入
力音声を認識する音声認識装置において、 複数話者の学習パターンから上記直交化辞書としての基
本となる直交軸を決定した後、個別話者の学習パターン
からの辞書作成は、既に登録されている辞書の軸と直交
する新たな軸を決定し、この新たな軸の辞書を登録する
か否かを判定して前記直交化辞書を構築することを特徴
とする音声認識装置。
[Claims] The similarity is calculated between an input speech pattern obtained by analyzing input speech and an orthogonal dictionary created based on learning patterns of multiple speakers collected in advance. In the speech recognition device that recognizes the input speech, after determining the orthogonal axes that are the basis of the orthogonalized dictionary from the learning patterns of multiple speakers, dictionary creation from the learning patterns of individual speakers is performed using the already registered A speech recognition device characterized in that the orthogonalized dictionary is constructed by determining a new axis orthogonal to an axis of the existing dictionary, and determining whether or not to register the dictionary of this new axis.
JP63176704A 1988-07-15 1988-07-15 Voice recognizing device Pending JPH0225899A (en)

Priority Applications (1)

Application Number Priority Date Filing Date Title
JP63176704A JPH0225899A (en) 1988-07-15 1988-07-15 Voice recognizing device

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
JP63176704A JPH0225899A (en) 1988-07-15 1988-07-15 Voice recognizing device

Publications (1)

Publication Number Publication Date
JPH0225899A true JPH0225899A (en) 1990-01-29

Family

ID=16018291

Family Applications (1)

Application Number Title Priority Date Filing Date
JP63176704A Pending JPH0225899A (en) 1988-07-15 1988-07-15 Voice recognizing device

Country Status (1)

Country Link
JP (1) JPH0225899A (en)

Similar Documents

Publication Publication Date Title
Yogesh et al. A new hybrid PSO assisted biogeography-based optimization for emotion and stress recognition from speech signal
JP2739950B2 (en) Pattern recognition device
Tuncer et al. Automatic voice based disease detection method using one dimensional local binary pattern feature extraction network
Bharali et al. Speech recognition with reference to Assamese language using novel fusion technique
JPH04369696A (en) Voice recognizing method
JPH02165388A (en) Pattern recognition system
JPH0225898A (en) Voice recognizing device
Abdullaeva et al. Formant set as a main parameter for recognizing vowels of the Uzbek language
Safie Spoken digit recognition using convolutional neural network
JPH0225899A (en) Voice recognizing device
Telembici et al. Optimizing Audio Recognition for Assistive Robotics with Feature Optimization, Machine Learning and Data Augmentation
Suryawanshi et al. Hardware implementation of speech recognition using mfcc and euclidean distance
JPH01277297A (en) Sound recognizing device
JPH0194396A (en) Voice recognition system
Alex et al. Performance analysis of SOFM based reduced complexity feature extraction methods with back propagation neural network for multilingual digit recognition
JP2856429B2 (en) Voice recognition method
JPH0194394A (en) Voice recognition system
JPH0194397A (en) Voice recognition system
Pentapati et al. Log-melspectrum and excitation features based speaker identification using deep learning
Besbes et al. Classification of speech under stress based on cepstral features and one-class SVM
JPH054678B2 (en)
ALTAF et al. ELEVATING VOICE DIAGNOSTICS: SAVA UNLEASHES NEW FRONTIERS IN HEALTHY AND PATHOLOGICAL VOICE DETECTION
JPH0194395A (en) Voice recognition system
Singh et al. Broad Acoustic Classification of Spoken Hindi Hybrid Paired Words using Artificial Neural Networks
JPS60147797A (en) Voice recognition equipment