JPH02226200A - Voice recognition device - Google Patents

Voice recognition device

Info

Publication number
JPH02226200A
JPH02226200A JP4588389A JP4588389A JPH02226200A JP H02226200 A JPH02226200 A JP H02226200A JP 4588389 A JP4588389 A JP 4588389A JP 4588389 A JP4588389 A JP 4588389A JP H02226200 A JPH02226200 A JP H02226200A
Authority
JP
Japan
Prior art keywords
speaker
template
similar
section
matching
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
JP4588389A
Other languages
Japanese (ja)
Inventor
Haruyuki Hayashi
晴之 林
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
NEC Corp
Original Assignee
NEC Corp
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by NEC Corp filed Critical NEC Corp
Priority to JP4588389A priority Critical patent/JPH02226200A/en
Publication of JPH02226200A publication Critical patent/JPH02226200A/en
Pending legal-status Critical Current

Links

Abstract

PURPOSE:To improve the recognition rate by obtaining template information which is most similar to an input pattern according to the output of a matching part and selecting the template of a corresponding similar speaker cluster by referring to two tables. CONSTITUTION:A standard pattern part 4 is provided with a similar speaker cluster table stored with plural other speaker clusters whose voices are most similar to an actual speaker cluster and the actual speaker cluster as standard patterns which are predetermined correspondingly by speaker clusters. A template selection part 5 obtains the template information which is most similar to an input pattern according to the output template 17 of a matching part 2 and refers to a category and a similar speaker cluster table to select and supply the template 16 of the similar speaker clusters to a matching part 2. Consequently, the quality of a limited amount of templates can be increased as much as possible adaptively to the speaker, and the recognition rate can be improved.

Description

【発明の詳細な説明】 〔産業上の利用分野〕 本発明は、ディジタル音声処理の音声認識装置に利用す
る。特に、マルチテンプレート方式による不特定話者用
および多数話者用の音声認識装置に関するものである。
DETAILED DESCRIPTION OF THE INVENTION [Industrial Application Field] The present invention is applied to a speech recognition device for digital speech processing. In particular, the present invention relates to a speech recognition device for non-specific speakers and for multiple speakers using a multi-template method.

〔概要〕〔overview〕

本発明は音声認識装置において、 1回前のマツチング処理の結果から得られた人力パタン
に最も類似したテンプレートの情報に基づいて類似話者
クラスタを推定し次にマッチング処理を行う際に用いる
テンプレートを選択することにより、 限られた量のテンプレートの利用価値を話者に適応して
最大限に高めることができ、認識率を向上するようにし
たものである。
The present invention provides a speech recognition device that estimates a cluster of similar speakers based on information about a template that is most similar to a human pattern obtained from the results of the previous matching process, and selects a template to be used when performing the next matching process. By making this selection, the utility value of the limited amount of templates can be maximized by adapting them to the speaker, thereby improving the recognition rate.

〔従来の技術〕[Conventional technology]

従来、不特定話者用および多数話者用の音声認識装置は
、マルチテンプレート方式を用いるものが多いが、この
方式はカテゴリごとにあらかじめ用意された複数のテン
プレートを常時すべてマツチングに使用していた。
Conventionally, speech recognition devices for non-specific speakers and for multiple speakers often use a multi-template method, but this method always uses multiple templates prepared in advance for each category for matching. .

〔発明が解決しようとする問題点〕[Problem that the invention seeks to solve]

しかし、このような音声tRm装置では、テンプレート
の質と量とによって認識性能が大きく左右される。この
うちテンプレートの量に関してはハードウェアの制限を
受けるために、限られたテンプレートの量で多くの話者
に適用させる場合に、認識性能に限界が生じる欠点があ
った。すなわち、相異度の高い話者がそれぞれ発声した
別々の内容の音声が高い類似度を示す場合がある。その
ために別々の内容の音声が同じ識別結果となったり、逆
に同じ内容の音声が別々の識別結果となる場合があり、
結果として認識性能の低下を招く欠点があった。
However, in such a voice tRm device, recognition performance is greatly influenced by the quality and quantity of templates. Since the amount of templates is limited by hardware, there is a drawback in that recognition performance is limited when a limited amount of templates is applied to many speakers. That is, voices of different contents uttered by speakers with a high degree of dissimilarity may exhibit a high degree of similarity. As a result, voices with different content may have the same identification result, or conversely, voices with the same content may have different identification results.
As a result, there was a drawback that the recognition performance deteriorated.

本発明は上記の欠点を解決するもので、限られた量のテ
ンプレートの質を話者に適応して最大限に高めることが
でき、認識率を向上できる音声δ忍識装置を提供するこ
とを目的とする。
The present invention solves the above-mentioned drawbacks, and aims to provide a speech delta intelligence device that can maximize the quality of a limited amount of templates by adapting them to the speaker and improve the recognition rate. purpose.

〔問題点を解決するための手段〕[Means for solving problems]

本発明は、あらかじめ定められた標準パタンの音声を発
声した話者に対してクラスタリングした話者クラスタご
とに対応し所定のカテゴリごとに分割されたテンプレー
トが格納されたカテゴリテーブルを含む標準パタン部と
、入力信号を分析して人力パタンに変換する分析部と、
この分析部の出力と上記カテゴリテーブルの内容とのマ
ツチング処理を行うマツチング部とを備えた音声認識装
置において、上記標準パタン部は、上記話者クラスタご
とに対応して上記あらかじめ定められた標準パタンとし
て発声した音声が自話者クラスタに最も類似した他の複
数の話者クラスタおよび自話者クラスタが格納された類
似話者クラスタテーブルを設けておき、上記マツチング
部の出力に基づいて上記人力パタンに最も類似したテン
プレート情報を得て上記二つのテーブルを参照して該当
する類似話者クラスタのテンプレートを選択して上記マ
ツチング処理部に与えるテンプレート選択部を備えたこ
とを特徴とする。
The present invention includes a standard pattern section including a category table storing templates divided into predetermined categories and corresponding to each speaker cluster that is clustered for speakers who have uttered sounds in a predetermined standard pattern. , an analysis section that analyzes input signals and converts them into human patterns;
In a speech recognition device that includes a matching unit that performs a matching process between the output of the analysis unit and the contents of the category table, the standard pattern unit generates the predetermined standard pattern for each speaker cluster. A similar speaker cluster table is provided in which multiple other speaker clusters and own speaker clusters whose uttered voice is most similar to the own speaker cluster are provided, and the above human pattern is determined based on the output of the matching section. The present invention is characterized by comprising a template selection unit that obtains template information most similar to the above, refers to the two tables, selects a template of a corresponding similar speaker cluster, and supplies the template to the matching processing unit.

〔作用〕[Effect]

標準パタン部に話者クラスタごとに対応してあらかじめ
定められた標準パタンとして発声した音声が自話者クラ
スタに最も類似した他の複数の話者クラスタおよび自話
者クラスタが格納された類似話者タラスタテーブルを設
ける。テンプレート選択部はマツチング部の出力に基づ
いて入力パタンに最も類似したテンプレート情報を得て
カテゴリテーブルおよび類似話者クラスタテーブルを参
照して該当する類似話者クラスタのテンプレートを選択
してマツチング処理部に与える。以上の動作により限ら
れた量のテンプレートの質を話者に適応して最大限に高
めることができ、認識率を向上できる。
The standard pattern section stores multiple other speaker clusters and self-speaker clusters whose voice uttered as a predetermined standard pattern corresponding to each speaker cluster is most similar to the self-speaker cluster. A Tarastar table will be provided. The template selection section obtains template information most similar to the input pattern based on the output of the matching section, refers to the category table and the similar speaker cluster table, selects the template of the corresponding similar speaker cluster, and sends the template to the matching processing section. give. Through the above operations, the quality of the limited amount of templates can be maximized by adapting them to the speaker, and the recognition rate can be improved.

〔実施例〕〔Example〕

本発明の実施例について図面を参照して説明する。第1
図は本発明一実施例音声認識装置のブロック構成図であ
る。第1図において、音声認識装置は、あらかじめ定め
られた標準パタンとして音声を発声した話者に対してク
ラスタリングした話者クラスタごとに対応し所定のカテ
ゴリごとに分割されたテンプレートが格納されたカテゴ
リテーブルを含む標準パタン部4と、入力信号11を分
析して人力パタンに変換する分析部1と、この分析部1
の出力人力パタン12と上記カテゴリテーブルの内容と
のマツチング処理を行うマツチング部2と、マツチング
部2のマツチング結果に基づいて識別結果14を出力す
る識別部3とを備える。
Embodiments of the present invention will be described with reference to the drawings. 1st
The figure is a block diagram of a speech recognition device according to an embodiment of the present invention. In FIG. 1, the speech recognition device stores templates divided into predetermined categories corresponding to each speaker cluster, which is clustered for speakers who have uttered speech according to a predetermined standard pattern. an analysis section 1 that analyzes the input signal 11 and converts it into a manual pattern;
The matching unit 2 includes a matching unit 2 that performs a matching process between the output human pattern 12 and the content of the category table, and an identification unit 3 that outputs an identification result 14 based on the matching result of the matching unit 2.

ここで本発明の特徴とするところは、標準パタン部4は
、上記話者クラスタごとに対応して上記あらかじめ定め
られた標準パタンとして発声した音声が自話者クラスタ
に最も類似した他の複数の話者クラスタおよび白話者ク
ラスタが格納された類似話者クラスタテーブルを設けて
おき、マツチング部2の出力テンプレーH7に基づいて
上記入力パタンに最も類似したテンプレート情報を得て
上記二つのテーブルを参照して該当する類似話者クラス
タのテンプレート16を選択してマツチング処理部2に
与えるテンプレート選択部5を備えたことを特徴とする
。
Here, the feature of the present invention is that the standard pattern section 4 is configured to select a plurality of other patterns in which the voice uttered as the predetermined standard pattern corresponds to each speaker cluster and is most similar to the own speaker cluster. A similar speaker cluster table in which speaker clusters and white speaker clusters are stored is provided, and template information most similar to the input pattern is obtained based on the output template H7 of the matching section 2, and the two tables are referred to. The present invention is characterized in that it includes a template selection section 5 that selects a template 16 of a corresponding similar speaker cluster and supplies it to the matching processing section 2.

このような構成の音声認識装置の動作について説明する
。第1表は本発明の音声認識装置の類似話者クラスタテ
ーブルである。第2表は本発明の音声認識装置のカテゴ
リテーブルである。第2図は本発明の音声認識装置のテ
ンプレート選択部の動作を示すフローチャートである。
The operation of the speech recognition device having such a configuration will be explained. Table 1 is a similar speaker cluster table of the speech recognition device of the present invention. Table 2 is a category table for the speech recognition device of the present invention. FIG. 2 is a flowchart showing the operation of the template selection section of the speech recognition device of the present invention.

(以下本頁余白) 第1表 第2表 まず、分析部1は、入力信号11を入力パタン12に変
換する。マツチング部2では、この入カバターン12と
マツチングに用いるテンプレート16とのマツチング処
理を行う。識別部3では、そのマツチング結果13から
識別結果14を出力する。
(Hereinafter, this page margin) Table 1 Table 2 First, the analysis section 1 converts the input signal 11 into an input pattern 12. The matching section 2 performs a matching process between this input cover pattern 12 and a template 16 used for matching. The identification section 3 outputs the identification result 14 from the matching result 13.

また、テンプレート選択部5では、前回のマツチング処
理の結果入カバターン12に最も類似したテンプレート
17を受取り、その情報から次にマツチングを行う際に
用いるテンプレートを選択する。
Further, the template selection unit 5 receives the template 17 most similar to the input cover pattern 12 as a result of the previous matching process, and selects a template to be used for the next matching based on that information.

選択されたテンプレート15を標準パターン部4から受
取り次にマツチングに用いるテンプレート16として出
力する。
The selected template 15 is received from the standard pattern section 4 and then output as a template 16 used for matching.

ここでこのテンプレート選択部5での処理をさらに詳し
く説明する。第1表は話者クラスタに対応する類似話者
クラスタを示す話者クラスタテーブルである。まず話者
クラスタ (St 、S2、S、)とはあらかじめ定め
られた標準パタンとして用いる音声を発声した話者に対
してクラスタリングしたものである。クラスタ、リング
するためのデータは、たとえば各話者が発声した5個の
母音を用いたり、より多くの孤立発声した音素データを
用いたり、またはカテゴリとなる全単語(または音素等
)のデータを用いたりする方法がある。
Here, the processing in the template selection section 5 will be explained in more detail. Table 1 is a speaker cluster table showing similar speaker clusters corresponding to speaker clusters. First, the speaker cluster (St, S2, S,) is a clustering of speakers who have uttered voices used as predetermined standard patterns. The data for clustering and ringing may be, for example, using five vowels uttered by each speaker, using more phoneme data of isolated utterances, or using data for all words (or phonemes, etc.) that form a category. There are ways to use it.

クラスタリングされた各話者クラスタ間の類似度をクラ
スタリングに使用したデータを用いて求める。たとえば
各話者クラスタについて最も類似した他の話者クラスタ
を数個選び、自分自身を含めた数個の類似話者クラスタ
を求める。この個数は全話者クラスタにおいて同じにす
る必要はない。
The degree of similarity between each clustered speaker cluster is determined using the data used for clustering. For example, select several other speaker clusters that are most similar to each speaker cluster, and find several similar speaker clusters including the speaker cluster itself. This number does not need to be the same for all speaker clusters.

このテーブル例では、S、に対する類似話者クラスタは
、s、 、s、、Sbであり、S2に対しては32 、
Sc、Saである。
In this example table, the similar speaker clusters for S, are s, , s, , Sb, and 32 for S2,
Sc, Sa.

次に第2表は、各話者クラスタ(St、・−1S15、
Sl)における各カテゴリ (W+ 、−1WJ、・、
Wo)のテンプレート(T10、・・・・、Tl、、・
・T□)を示すカテゴリテーブルである。この例では各
話者クラスタにおける各カテゴリのテンプレートは1゛
個であるが、複数の場合もある。
Next, Table 2 shows each speaker cluster (St, -1S15,
SL) for each category (W+, -1WJ,...
Wo) template (T10,..., Tl,...
・T□). In this example, there is one template for each category in each speaker cluster, but there may be more than one template.

第2図において、まず、マツチング部2からマツチング
処理の結果入カバターンと最も類似したテンプレート1
7としてテンプレートT1」の情報を受取る(Sl)。
In FIG. 2, first, the matching section 2 selects a template 1 that is most similar to the input cover pattern as a result of the matching process.
7, the information on the template T1 is received (Sl).

これから第2表に示すカテゴリテーブルを参照して話者
クラスタSiを見つける(S2)。次に第1表に示す話
者クラスタテーブルを参照して類似話者クラスタS1、
So、Sfを見つける(S3)。最後にもう一度第2表
に示すテーブルを参照して次のマツチング処理に用いる
テンプレート15としてテンプレートTi0、TI□、
% Tin%Tel、T、2、  、Tan5 T’r
+、Tf21、Tいを選択する。そしてこの選択された
テンプレートをテンプレート16としてマツチング部2
へ出力する。(S4)。
From now on, the speaker cluster Si is found by referring to the category table shown in Table 2 (S2). Next, referring to the speaker cluster table shown in Table 1, similar speaker cluster S1,
Find So and Sf (S3). Finally, referring to the table shown in Table 2 again, templates Ti0, TI□,
%Tin%Tel,T,2, ,Tan5 T'r
+, Tf21, and Tf21. Then, the matching section 2 uses this selected template as the template 16.
Output to. (S4).

〔発明の効果〕〔Effect of the invention〕

以上説明したように、本発明は、限られた量のテンプレ
ートの質を話者に適応して最大限に高めることができ、
認識率を高くできる優れた効果がある。
As explained above, the present invention can maximize the quality of a limited amount of templates by adapting them to the speaker.
It has an excellent effect of increasing the recognition rate.

構成図。Diagram.

第2図は本発明の音声認識装置のテンプレート選択部の
動作を示すフローチャート。
FIG. 2 is a flowchart showing the operation of the template selection section of the speech recognition device of the present invention.

1・・・分析部、2・・・マツチング部、3・・・識別
部、4・・・標準パタン部、5・・・テンプレート選択
部、11・・・入力信号、12・・・入力パタン、13
・・・マツチング結果、14・・・識別結果、15・・
・選択されたテンプレート、16・・・マツチングに用
いるテンプレート、17・・・マツチング処理の結果入
カバターンと最も類似したテンプレート。
DESCRIPTION OF SYMBOLS 1... Analysis section, 2... Matching section, 3... Identification section, 4... Standard pattern section, 5... Template selection section, 11... Input signal, 12... Input pattern , 13
...Matching result, 14...Identification result, 15...
- Selected template, 16...Template used for matching, 17...Template most similar to the cover pattern entered as a result of matching processing.

代理人  弁理士 井 出 直 孝Agent: Patent attorney Naotaka Ide

【図面の簡単な説明】[Brief explanation of the drawing]

第1図は本発明一実施例音声認識装置ブロック実施例 第1図 実施例 テンプレート選択部のフローチャート第2図 FIG. 1 is an embodiment of a speech recognition device block according to one embodiment of the present invention. Figure 1 Example: Flowchart of template selection section Fig. 2

Claims (1)

【特許請求の範囲】 1、あらかじめ定められた標準パタンの音声を発声した
話者に対してクラスタリングした話者クラスタごとに対
応し所定のカテゴリごとに分割されたテンプレートが格
納されたカテゴリテーブルを含む標準パタン部と、 入力信号を分析して入力パタンに変換する分析部と、 この分析部の出力と上記カテゴリテーブルの内容とのマ
ッチング処理を行うマッチング部とを備えた音声認識装
置において、 上記標準パタン部は、上記話者クラスタごとに対応して
上記あらかじめ定められた標準パタンとして発声した音
声が自話者クラスタに最も類似した他の複数の話者クラ
スタおよび自話者クラスタが格納された類似話者クラス
タテーブルを設けておき、 上記マッチング部の出力に基づいて上記入力パタンに最
も類似したテンプレート情報を得て上記二つのテーブル
を参照して該当する類似話者クラスタのテンプレートを
選択して上記マッチング処理部に与えるテンプレート選
択部 を備えたことを特徴とする音声認識装置。
[Claims] 1. Contains a category table storing templates divided into predetermined categories corresponding to each speaker cluster obtained by clustering speakers who have uttered a predetermined standard pattern of speech. A speech recognition device comprising a standard pattern section, an analysis section that analyzes an input signal and converts it into an input pattern, and a matching section that performs a matching process between the output of this analysis section and the contents of the category table described above. The pattern section stores a plurality of other speaker clusters and a self-speaker cluster whose voice uttered according to the predetermined standard pattern is most similar to the self-speaker cluster corresponding to each of the above-mentioned speaker clusters. A speaker cluster table is provided, template information most similar to the input pattern is obtained based on the output of the matching section, the template of the corresponding similar speaker cluster is selected by referring to the two tables, and the above is performed. A speech recognition device comprising a template selection section for providing a template to a matching processing section.
JP4588389A 1989-02-27 1989-02-27 Voice recognition device Pending JPH02226200A (en)

Priority Applications (1)

Application Number Priority Date Filing Date Title
JP4588389A JPH02226200A (en) 1989-02-27 1989-02-27 Voice recognition device

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
JP4588389A JPH02226200A (en) 1989-02-27 1989-02-27 Voice recognition device

Publications (1)

Publication Number Publication Date
JPH02226200A true JPH02226200A (en) 1990-09-07

Family

ID=12731634

Family Applications (1)

Application Number Title Priority Date Filing Date
JP4588389A Pending JPH02226200A (en) 1989-02-27 1989-02-27 Voice recognition device

Country Status (1)

Country Link
JP (1) JPH02226200A (en)

Cited By (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2010145784A (en) * 2008-12-19 2010-07-01 Casio Computer Co Ltd Voice recognizing device, acoustic model learning apparatus, voice recognizing method, and program

Cited By (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2010145784A (en) * 2008-12-19 2010-07-01 Casio Computer Co Ltd Voice recognizing device, acoustic model learning apparatus, voice recognizing method, and program

Similar Documents

Publication Publication Date Title
EP1185976B1 (en) Speech recognition device with reference transformation means
JPS597998A (en) Continuous voice recognition equipment
JPH02232696A (en) Voice recognition device
Chiba et al. A speaker-independent word-recognition system using multiple classification functions
JPH09179578A (en) Syllable recognition device
JP2561553B2 (en) Standard speaker selection device
KR19990015122A (en) Speech recognition method
JP2000207166A (en) Device and method for voice input
JPH04324499A (en) Speech recognition device
JP3446666B2 (en) Apparatus and method for speaker adaptation of acoustic model for speech recognition
JP3536380B2 (en) Voice recognition device
JPH0430598B2 (en)
JPH02109100A (en) Voice input device
JPS638798A (en) voice recognition device
JPS6073592A (en) Voice recognition equipment for specific speaker
JPH04271397A (en) Voice recognizer
JPS59214900A (en) voice recognition device
JPS6287993A (en) voice recognition device
JPH01161399A (en) Method of suiting voice recognition apparatus to speaker
JPH02251999A (en) Production of standard pattern
JPS59212900A (en) voice recognition device
JPS61105599A (en) Continuous sound recognition equipment
JPS60241097A (en) Speech recognition application device
JPS63218999A (en) voice recognition device
JPS59176791A (en) Voice registration system