JPH02220099A - Word voice recognition device - Google Patents

Word voice recognition device

Info

Publication number
JPH02220099A
JPH02220099A JP1042191A JP4219189A JPH02220099A JP H02220099 A JPH02220099 A JP H02220099A JP 1042191 A JP1042191 A JP 1042191A JP 4219189 A JP4219189 A JP 4219189A JP H02220099 A JPH02220099 A JP H02220099A
Authority
JP
Japan
Prior art keywords
codebook
vectors
vector
input
speech
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
JP1042191A
Other languages
Japanese (ja)
Inventor
Sadahiro Furui
古井 貞▲ひろ▼
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
NTT Inc
Original Assignee
Nippon Telegraph and Telephone Corp
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Nippon Telegraph and Telephone Corp filed Critical Nippon Telegraph and Telephone Corp
Priority to JP1042191A priority Critical patent/JPH02220099A/en
Publication of JPH02220099A publication Critical patent/JPH02220099A/en
Pending legal-status Critical Current

Links

Abstract

(57)【要約】本公報は電子出願前の出願データであるた
め要約のデータは記録されません。
(57) [Summary] This bulletin contains application data before electronic filing, so abstract data is not recorded.

Description

【発明の詳細な説明】 〔産業上の利用分野〕 この発明は、不特定多数の発声者に対して認識能力を向
上した単語音声認識装置に関する。
DETAILED DESCRIPTION OF THE INVENTION [Field of Industrial Application] The present invention relates to a word speech recognition device that has improved recognition ability for an unspecified number of speakers.

〔従来の技術〕[Conventional technology]

不特定多数の話者が発声した単語音声の認識においては
、従来、確率的なモデルによって音声の変動をモデル化
し吸収する方法、各単語毎に複数の話者の発声に対応す
る標準パターンを用意しておき、入力音声に最も近い標
準パターンを選択する方法、などが用いられているが、
いずれの場合でも、標準からはずれた話者に関しては高
い認識性能を得ることが難しい、このため、発声者に認
識装置を適応化する方法が試みられているが(例えば文
献、中村・鹿狩:セパレートベクトル量子化を用いたス
ペクトログラムの正規化、日本音響学会誌、44. 8
. P、595.1988参照)、これを行うためには
、発声者が変わるたびにあらかじめ決められた学習用の
多数の単語や文章を発声する必要があり、極めて不便で
あった。あらかじめ決められた単語や文章ではなく、認
識すべき未知の入力音声を用いて適応化する方法もこれ
までに試みられているが(例えば文献、杉山:母音の教
師なし話者適応における各種の方法の比較、電子情報通
信学会論文誌、70−D、  5. P、958.19
87参照)、高い適応化性能が得られないという問題が
あった。
Conventionally, in the recognition of word sounds uttered by an unspecified number of speakers, methods have been used to model and absorb speech fluctuations using a probabilistic model, and to prepare standard patterns that correspond to the utterances of multiple speakers for each word. A method is used in which the standard pattern closest to the input audio is selected.
In either case, it is difficult to obtain high recognition performance for speakers who deviate from the standard.For this reason, methods of adapting the recognition device to the speaker have been attempted (for example, in the literature, Nakamura and Shikagari: Spectrogram normalization using separate vector quantization, Journal of the Acoustical Society of Japan, 44.8
.. In order to do this, it was necessary to utter a large number of predetermined words and sentences for learning each time the speaker changed, which was extremely inconvenient. Adaptation methods using unknown input speech to be recognized rather than predetermined words or sentences have been attempted (for example, in the literature, Sugiyama: Various methods for unsupervised speaker adaptation of vowels). Comparison of IEICE Transactions, 70-D, 5. P, 958.19
87), there was a problem that high adaptation performance could not be obtained.

この発明は、上記に鑑みてなされたもので、その目的と
しては、発声者の任意の音声を用いて、その発声者の音
声と標準的な音声との対応関係を算出し、これに基づい
てその発声者の音声を標準的な音声に自動的に適応化し
、あるいは標準的な音声をその発声者の音声に適応化す
ることにより、不特定話者の音声に対して高い認識性能
が得られるようにした音声認識装置を提供することにあ
る。
This invention was made in view of the above, and its purpose is to use any voice of a speaker, calculate the correspondence between the voice of the speaker and a standard voice, and based on this, By automatically adapting the speaker's voice to a standard voice, or by adapting the standard voice to the voice of the speaker, high recognition performance can be obtained for the voice of an unspecified speaker. An object of the present invention is to provide a speech recognition device that does the following.

〔課題を解決するための手段〕[Means to solve the problem]

上記目的を達成するため、この発明は、単語の音声信号
の周波数スペクトルおよびパワーの時間的変化を示すパ
ラメータベクトルを算出するパラメータベクトル算出手
段と、 一人または複数の標準話者が発声した多数の単語の音声
信号から算出した多数のパラメータベクトルをクラスタ
化して、複数の代表的なベクトル値を符号帳に蓄える符
号帳作成手段と、各認識対象l!鴬を、符号帳に蓄えら
れているベクトル値の、一つまたは複数の時系列で表現
して単語辞書に蓄える単語辞書作成手段と、入力音声信
号から算出された多数の人力パラメータベクトルと、符
号帳の要素ベクトルを、階層的にクラスタに分割化する
手段と、 分割化された各クラスタごとに入力パラメータベクトル
を符号帳に適応化するための、区分移動方向ベクトルを
決定する手段と、 これらの区分移動方向ベクトルの加重平均を用いてすべ
ての入力パラメータベクトルを適応化する手段と、 これらの適応化された入力パラメータベクトルと、符号
帳の各要素との距離を算出する手段と、これらの算出さ
れた距離と各単語辞書を用いて、動的計画法あるいは隠
れマルコフモデルによって入力音声と各認識対象語彙と
の距離を算出する手段とを有することを要旨とする。あ
るいは入力パラメータベクトルを符号帳に適応化する代
りに、符号帳を入力パラメータに適応して、その適応化
された符号帳要素と入力パラメータベクトルとの距離を
算出する。
In order to achieve the above object, the present invention provides a parameter vector calculation means for calculating a parameter vector indicating temporal changes in the frequency spectrum and power of a word audio signal; codebook creation means for clustering a large number of parameter vectors calculated from audio signals of l! and storing a plurality of representative vector values in a codebook; A word dictionary creation means for expressing the word "Umugi" as one or more time series of vector values stored in a codebook and storing the expression in a word dictionary, a large number of human parameter vectors calculated from input audio signals, and a code. means for hierarchically dividing the element vectors of the codebook into clusters; means for determining segmental movement direction vectors for adapting input parameter vectors to the codebook for each divided cluster; means for adapting all input parameter vectors using a weighted average of segmented movement direction vectors; means for calculating distances between these adapted input parameter vectors and each element of the codebook; The object of the present invention is to have a means for calculating the distance between the input speech and each recognition target vocabulary by dynamic programming or hidden Markov model using the distance and each word dictionary. Alternatively, instead of adapting the input parameter vector to the codebook, the codebook is adapted to the input parameters and the distance between the adapted codebook element and the input parameter vector is calculated.

従来技術とは、発声内容が限定されない任意の言葉を発
声した音声を用いて、パラメータベクトルと符号帳のク
ラスタ化を行い、これに基づいて入力音声パラメータベ
クトル又は符号帳を適応化する手段を有することが異な
る。
The conventional technology includes means for clustering parameter vectors and codebooks using speech uttering arbitrary words whose utterance content is not limited, and adapting input speech parameter vectors or codebooks based on this clustering. Things are different.

〔実施例〕〔Example〕

以下、図面を用いてこの発明の詳細な説明する。 Hereinafter, the present invention will be explained in detail using the drawings.

第1図は、この発明に係る単語音声認識装置の回路ブロ
ックを示す図である。
FIG. 1 is a diagram showing a circuit block of a word speech recognition device according to the present invention.

同図において、1は例えばマイクロホン等に接続され、
単語の音声信号(以下単に「音声信号」と呼ぶ)を入力
して、次段の音声区間検出回路2に供給する音声入力端
子である。
In the figure, 1 is connected to, for example, a microphone,
This is an audio input terminal for inputting a word audio signal (hereinafter simply referred to as "audio signal") and supplying it to the next stage audio section detection circuit 2.

音声区間検出回路2は、音声信号が一般に主に雑音で構
成される無音の部分とそうでない実際の音声の部分を含
むので、設定時間(例えば10m5)毎のパワーを演算
し、それに基づいて無音の部分と音声の部分を判別する
回路である0判別方法としては、例えば設定時間ごとの
パワーの絶対値が所定レベルを越えている部分を音声部
分と判別する方法、設定時間ごとのパワーについて所定
のレベルを越える状態が所定時間継続すればこれを音声
部分と判別する方法等、種々の周知の方法が通用できる
。音声区間検出回路2はパラメータベクトル算出回路3
に接続されている。音声区間検出回路2で演算されたパ
ワー値のうち、音声区間と判別された区間の値は、パラ
メータベクトル算出回路3に入力される。
The voice section detection circuit 2 calculates the power for each set time (for example, 10 m5) and detects silence based on the power since the voice signal generally includes a silent part mainly consisting of noise and an actual voice part that is not. For example, a method of determining a part where the absolute value of the power for each set time exceeds a predetermined level is a sound part, and a method for determining a part where the absolute value of the power for each set time exceeds a predetermined level is a circuit that distinguishes between the part and the audio part. Various known methods can be used, such as a method in which if the level exceeds the level for a predetermined period of time, it is determined that this is a voice part. The voice section detection circuit 2 is a parameter vector calculation circuit 3
It is connected to the. Among the power values calculated by the speech section detection circuit 2, the values of the sections determined to be speech sections are input to the parameter vector calculation circuit 3.

パラメータベクトル算出回路3は、音声区間検出回路2
で検出された音声部分の信号を周波数スペクトルおよび
パワーの時間的変化を示すパラメータベクトル系列に変
換処理する回路である。この変換処理については、すで
に公知の方法、例えば音声信号を線形予測ケプストラム
(以下単にrケプストラム」と呼ぶ)の時系列に変換す
る方法を用いる。これは、音声信号をまず線形予測係数
の時系列に変換し、次にこれをケプストラムに変換する
ことによって行う。
The parameter vector calculation circuit 3 is the voice section detection circuit 2.
This is a circuit that converts the signal of the audio part detected in the above process into a parameter vector series that indicates temporal changes in frequency spectrum and power. For this conversion process, a known method is used, for example, a method of converting an audio signal into a linear predictive cepstrum (hereinafter simply referred to as r-cepstrum) time series. This is done by first converting the audio signal into a time series of linear prediction coefficients, and then converting this into a cepstrum.

音声信号から線形予測係数への変換処理の概要(例えば
、文献、機素・斉H:統計的手法による音声スペクトル
密度とホルマント周波数の推定、電子通信学会論文誌、
53−A、 1. P、35.1970参照)は、次の
通りである。基本的にはまず低域通過フィルタに通した
のち、標本化及び量子化を行い、一定時間(例えばio
ms)ごとに短区間の波形を切り出してハミング窓等を
乗じ、積和の演算によって相関係数を計算する。その相
関係数から、繰り返し演算処理によって代数方程式を解
くことにより、容易に線形予測係数が抽出されるのであ
る。
Overview of the conversion process from audio signals to linear prediction coefficients (e.g., literature, Kimoto-Sai H: Estimation of audio spectral density and formant frequency using statistical methods, Journal of the Institute of Electronics and Communication Engineers,
53-A, 1. P, 35.1970) is as follows. Basically, it is first passed through a low-pass filter, then sampled and quantized, and then for a certain period of time (for example,
ms), a short section waveform is cut out, multiplied by a Hamming window, etc., and a correlation coefficient is calculated by calculating the sum of products. From the correlation coefficients, linear prediction coefficients can be easily extracted by solving algebraic equations through repeated arithmetic processing.

線形予測係数からケプストラムへの変換処理の詳11(
例えば文献、斉藤・中日:音声情報処理の基礎、オーム
社、第7章、P、102 、1981参照)は省略する
が、線形予測係数を用いた再帰式を演算することにより
処理できる。この変換処理で得られたケプストラムは、
音声区間検出回路2から入力されたパワーと組み合わさ
れて、パラメータベクトルとされ、パラメータベクトル
算出回路3の出力段に接続されている符号帳作成回路4
に供給される。
Details of the conversion process from linear prediction coefficients to cepstrum 11 (
For example, see the literature, Saito and Chunichi: Fundamentals of Speech Information Processing, Ohmsha, Chapter 7, P, 102, 1981), but it can be processed by calculating a recursive formula using linear prediction coefficients. The cepstrum obtained by this conversion process is
A codebook creation circuit 4 which is combined with the power input from the speech interval detection circuit 2 to form a parameter vector and which is connected to the output stage of the parameter vector calculation circuit 3.
is supplied to

符号帳作成回路4は、一人または複数の標準話者が発声
した多数の単語の音声信号から算出された多数のパラメ
ータベクトルをクラスタ化するものである。このクラス
タ化は、多数のパラメータベクトルの組を、あらかじめ
定められた一定数の代表的なベクトル値の組にまとめる
ことである。
The codebook creation circuit 4 clusters a large number of parameter vectors calculated from audio signals of a large number of words uttered by one or more standard speakers. This clustering is to group a large number of parameter vector sets into a predetermined number of representative vector value sets.

例えば4名の話者が発声した100種類の単語音声から
10m!毎に11次元のベクトル(10次元のケプスト
ラムとパワー)が抽出されているとすると、単語音声の
長さが平均して500s+sであるとすれば、全部で5
0 X 4 X 100−20.000種類の11次元
ベクトルが与えられる。これを例えば1 、024種類
の代表的11次元ベクトルにまとめるには、公知の方法
(文献、!、 Linde+ A、Buzo andR
,M、  Gray  :  An  algorit
lv  for  vector  quanti−z
ation+  111111  TraIIs、  
Commun、、  vol、  C0M−28+  
pp。
For example, 10 meters from 100 different words uttered by 4 speakers! Assuming that an 11-dimensional vector (10-dimensional cepstrum and power) is extracted for each word, and the average length of word speech is 500 s + s, a total of 5
0.times.4.times.100-20.000 types of 11-dimensional vectors are given. For example, to summarize this into 1,024 types of representative 11-dimensional vectors, a known method (Reference, !, Linde + A, Buzo and R
,M.Gray: An algorithm
lv for vector quanti-z
ation+ 111111 TraIIs,
Commun,, vol, C0M-28+
pp.

84−95.1980)を用いることがてきる。この方
法では、amしているベクトルはまとめて一つの平均値
で代表させ、もとの20.006種類のすべてのベクト
ルを1,024種類の代表値のうちの最も近いもので置
き換えたときの、置き換えによる誤差が全体として最も
小さくなるように、代表値が決定される。このようにし
て決定された1、024種類のそれぞれ11次元のベク
トル代表値は、符号帳要素として、符号帳蓄積部5に蓄
積される。符号帳蓄積部5の出力端子と、パラメータベ
クトル算出回路3の出力端子は、単語辞書作成回路6に
接続されている。
84-95.1980) can be used. In this method, vectors that are am are collectively represented by one average value, and when all the original 20.006 types of vectors are replaced with the closest one among the 1,024 types of representative values, , the representative value is determined so that the error due to replacement is minimized as a whole. The 1,024 kinds of 11-dimensional vector representative values determined in this manner are stored in the codebook storage unit 5 as codebook elements. The output terminal of the codebook storage section 5 and the output terminal of the parameter vector calculation circuit 3 are connected to a word dictionary creation circuit 6.

単語辞書作成回路6は、一人または複数の話者が発声し
たすべての認識対象語案の音声信号から算出されたパラ
メータベクトル系列を、符号帳蓄積部5に蓄えられてい
る符号帳要素の時系列に変換するものである。この方法
は、各単語の例えば10■3毎のパラメータベクトルと
、すべての符号帳要素との距離を算出して、最も距離の
小さい符号帳要素を選択し、こうして得られた符号帳要
素を示す番号の時系列に変換することによって行う。
The word dictionary creation circuit 6 converts the parameter vector series calculated from the audio signals of all recognition target words uttered by one or more speakers into a time series of codebook elements stored in the codebook storage unit 5. It is converted into . This method calculates the distance between the parameter vector of each word, for example every 10×3, and all codebook elements, selects the codebook element with the smallest distance, and indicates the codebook element obtained in this way. This is done by converting it into a time series of numbers.

こうして決定された各単語ごとに一つまたは複数の番号
系列は、単語辞書として単語辞書蓄積部7に蓄えられる
One or more number sequences for each word thus determined are stored in the word dictionary storage section 7 as a word dictionary.

パラメータベクトル算出回路3の出力端子は、学習ベク
トル蓄積部8に接続されている。学習ベクトル蓄積部8
は、認識すべき音声の話者が発声した複数の単語の音声
信号から算出された多数のパラメータベクトル(以下、
「学習用パラメータベクトル」と呼ぶ)の組を蓄えるも
のである。
The output terminal of the parameter vector calculation circuit 3 is connected to the learning vector storage section 8. Learning vector storage unit 8
is a large number of parameter vectors (hereinafter referred to as
It stores a set of parameters (referred to as "learning parameter vectors").

符号帳蓄積部5の出力端子と、学習ベクトル蓄積部8の
出力端子は、階層的クラスタ化回路9に接続されている
0階層的クラスタ化回路9は、符号帳蓄積部5に蓄えら
れているすべての符号帳の要素ベクトル(以下、「符号
帳ベクトル」と呼ぶ)と、学習ベクトル蓄積部8に蓄え
られているすべての学習用パラメータベクトルを階層的
にクラスタ化するものである。この方法は、まずクラス
タ数を1とし、全符号帳ベクトルの平均ベクトル(以下
、「符号帳セントロイド」と呼ぶ)を算出する。同時に
、全学習用パラメータベクトルの平均ベクトル(以下、
「学習音声セントロイド」と呼ぶ)を算出する。この両
者のセントロイドは、階1葡クラスタ化回路9に接続さ
れている区分移動ベクトル算出回路lOに供給される。
The output terminal of the codebook storage section 5 and the output terminal of the learning vector storage section 8 are connected to the hierarchical clustering circuit 9. The zero hierarchical clustering circuit 9 is stored in the codebook storage section 5. All codebook element vectors (hereinafter referred to as "codebook vectors") and all learning parameter vectors stored in the learning vector storage unit 8 are hierarchically clustered. In this method, first, the number of clusters is set to 1, and the average vector of all codebook vectors (hereinafter referred to as "codebook centroid") is calculated. At the same time, the average vector of all learning parameter vectors (hereinafter,
(referred to as the “learning speech centroid”). Both centroids are supplied to a segment movement vector calculation circuit IO connected to the floor 1 clustering circuit 9.

区分移動ベクトル算出回路10は、符号帳セントロイド
から学習音声セントロイドを減算することにより、区分
移動ベクトルを求める回路である。
The segmented movement vector calculation circuit 10 is a circuit that calculates segmented movement vectors by subtracting the learning speech centroid from the codebook centroid.

このようにして得られた区分移動ベクトルは、学習ベク
トル蓄積部8の内容とともに学習音声適応化回路11に
供給される。
The segmented movement vectors thus obtained are supplied to the learning speech adaptation circuit 11 together with the contents of the learning vector storage section 8.

学習音声適応化回路11は、各学習用パラメータベクト
ルに区分移動ベクトルを加算することによって、学習用
パラメータベクトルを符号帳ベクトルに近付ける適応化
処理を行うものである。適応化処理された学習用パラメ
ータベクトルは、−旦学習ベクトル蓄積部8に蓄えられ
た後、再び階層的クラスタ化回路9に供給される。この
際、学習ベクトル蓄積部8に蓄えられている学習用パラ
メータベクトルの初期値は消去されず、適応化処理され
たベクトルが別に蓄えられる。
The learning speech adaptation circuit 11 performs an adaptation process that brings the learning parameter vector closer to the codebook vector by adding a segmentation movement vector to each learning parameter vector. The learning parameter vectors subjected to the adaptation process are stored in the learning vector storage section 8 and then supplied to the hierarchical clustering circuit 9 again. At this time, the initial values of the learning parameter vectors stored in the learning vector storage section 8 are not deleted, and the vectors subjected to the adaptation process are stored separately.

階層的クラスタ化回路9では、クラスタ数を2倍に増や
し、全符号帳ベクトルをクラスタ化して、各クラスタの
平均ベクトルすなわち符号帳セントロイドを算出する0
次に、各学習用パラメータベクトルについて、すべての
符号帳セントロイドとのI!離を計算し、最も距離の小
さい符号帳セントロイドに対応付ける。すべての学習用
パラメータベクトルについて、同じ符号帳セントロイド
に対応付けられたベクトルの平均ベクトルすなわち学習
音声セントロイドを算出する。すべての符号帳セントロ
イドと、対応するすべての学習音声セントロイドは、階
層的クラスタ化回路9に接続されている区分移動ベクト
ル算出回路10に供給される。
The hierarchical clustering circuit 9 doubles the number of clusters, clusters all codebook vectors, and calculates the average vector of each cluster, that is, the codebook centroid.
Next, for each learning parameter vector, I! with all codebook centroids. The distance is calculated and associated with the codebook centroid with the smallest distance. For all learning parameter vectors, the average vector of vectors associated with the same codebook centroid, that is, the learning speech centroid is calculated. All codebook centroids and all corresponding learning speech centroids are fed to a partitioned movement vector calculation circuit 10 which is connected to a hierarchical clustering circuit 9.

区分移動ベクトル算出回路10では、各符号帳セントロ
イドから対応する学習音声セントロイドを減算すること
により、各セントロイドに対応する区分移動ベクトルが
算出される。このようにして得られた各区分移動ベクト
ルは、学習ベクトル蓄積部8の内容とともに学習音声適
応化回路11に供給される。
The segmental movement vector calculation circuit 10 calculates the segmental movement vector corresponding to each centroid by subtracting the corresponding learning speech centroid from each codebook centroid. Each segmental movement vector obtained in this way is supplied to the learning speech adaptation circuit 11 together with the contents of the learning vector storage section 8.

学習音声適応化回路11では、各学習用パラメータベク
トルに区分移動ベクトルの重み付き平均値(以下、「適
応化ベクトル」と呼ぶ)を加算することによって、学習
用パラメータベクトルを符号帳ベクトルに近付ける適応
化処理が行われる。
The learning speech adaptation circuit 11 performs adaptation to bring the learning parameter vector closer to the codebook vector by adding a weighted average value of segmented movement vectors (hereinafter referred to as "adaptation vector") to each learning parameter vector. processing is performed.

適応化ベクトルは、次のようにして算出される。The adaptation vector is calculated as follows.

二こで、a、は1番目の学習用パラメータベクトルに加
算する適応化ベクトル、p、はm番目の区分移動ベクト
ル、Mは区分移動ベクトルの総数(クラスタの数)%W
111は重み係数で、W1+e−1/  l  I  
C!  −u−1l             (2)
にょう計算される。ここで、clは1番目の学習用パラ
メータベクトル、U、はm番目の学習音声セントロイド
、II   IIはベクトルのノルム(大きさ)の計算
である。適応化処理された学習用パラメータベクトルは
、−旦学習ベクトル蓄積部8に蓄えられた後、再び階層
的クラスタ化回路9に供給される。
Here, a is the adaptation vector to be added to the first learning parameter vector, p is the m-th segmental movement vector, and M is the total number of segmental movement vectors (number of clusters) %W
111 is a weighting coefficient, W1+e-1/l I
C! -u-1l (2)
It is calculated. Here, cl is the first learning parameter vector, U is the m-th learning speech centroid, and II is the calculation of the norm (size) of the vector. The learning parameter vectors subjected to the adaptation process are stored in the learning vector storage section 8 and then supplied to the hierarchical clustering circuit 9 again.

階層的クラスタ化回路9では、再度クラスタ数が2倍に
増加され、階層的クチスタ化回路9、区分移動ベクトル
真出回路10、及び学習音声適応化回路11によって、
上記と同様に学習用パラメータベクトルの適応化処理が
行われる。この一連の処理は、クラスタ数がある事前に
定めた数に達するか、適応化ベクトルの大きさがある事
前に定めた値よりも小さくなるまで繰り返される。二の
繰り返し処理が終了すると、学習ベクトル蓄積部8に蓄
えられていた学習用パラメータベクトルの初ItlI(
i!Iが、区分移動ベクトル算出回路10に供給され、
最終状態の適応化された学習用パラメータベクトルとそ
れぞれの初期値の差として、区分移動ベクトルが算出さ
れる0区分移動ベクトルは、区分移動ベクトル算出回路
10の出力段に接続された入力適応化回路12に供給さ
れる。
In the hierarchical clustering circuit 9, the number of clusters is doubled again, and the hierarchical clustering circuit 9, the segmented movement vector extraction circuit 10, and the learning speech adaptation circuit 11,
Adaptation processing of the learning parameter vector is performed in the same manner as above. This series of processing is repeated until the number of clusters reaches a certain predetermined number or the magnitude of the adaptation vector becomes smaller than a certain predetermined value. When the second iterative process is completed, the first ItlI(
i! I is supplied to the segmented movement vector calculation circuit 10,
The 0-segment movement vector is calculated as the difference between the adapted learning parameter vector of the final state and each initial value. 12.

次に、認識すべき未知の音声信号が、音声入力端子1に
入力される。この音声信号は、音声区間検出回路2に供
給され、実際の音声の区間が判別される。音声区間と判
別された区間の音声信号は、パラメータベクトル算出回
路3に供給され、パラメータベクトル系列に変換処理さ
れる。パラメータベクトル系列は、入力適応化回路12
に供給される。人力適応化回路12では、パラメータベ
クトル系列中の各ベクトルに適応化ベクトルを加算する
ことによって、パラメータベクトルを符号帳ベクトルに
近付ける適応化処理が行われる。適応化ベクトルは、す
でに同回路に供給されている区分移動ベクトルを用いて
、次のようにして算出される。
Next, an unknown audio signal to be recognized is input to the audio input terminal 1. This audio signal is supplied to the audio section detection circuit 2, and the actual audio section is determined. The audio signal of the section determined to be a speech section is supplied to the parameter vector calculation circuit 3, where it is converted into a parameter vector series. The parameter vector series is input to the input adaptation circuit 12.
is supplied to The human adaptation circuit 12 performs an adaptation process that brings the parameter vector closer to the codebook vector by adding the adaptation vector to each vector in the parameter vector series. The adaptation vector is calculated as follows using the segmented movement vector already supplied to the circuit.

ここで、bjはパラメータベクトル系列中のj番目のベ
クトルに加算する適応化ベクトル、q、、はn番目の区
分移動ベクトル、Nは区分移動ベクトルの総数(学習用
パラメータベクトルの敗)、Vj++は重み係数で、 vjm−1/l l yj  t、11 l     
 (4)により計算される。ここで、yjはj番目のパ
ラメータベクトル、tlはn番目の学習用パラメータベ
クトルである。適応化処理されたパラメータベクトル系
列は、距離行列計算回路13に供給される。
Here, bj is the adaptation vector to be added to the j-th vector in the parameter vector series, q, is the n-th segmental movement vector, N is the total number of segmental movement vectors (loss of learning parameter vectors), and Vj++ is the With the weighting factor, vjm-1/l l yj t, 11 l
Calculated by (4). Here, yj is the j-th parameter vector, and tl is the n-th learning parameter vector. The parameter vector sequence subjected to the adaptation process is supplied to the distance matrix calculation circuit 13.

距離行列計算回路13には、並行して符号帳蓄積部5か
らすべての符号帳要素が供給され、各パラメータベクト
ルと全符号帳要素との距離が計算される。これらの距離
値は、離散化された時間軸と符号帳要素番号をそれぞれ
行と列とする行列の形に並べられて、DP演算回路14
に送られる。
All the codebook elements are supplied from the codebook storage section 5 in parallel to the distance matrix calculation circuit 13, and the distances between each parameter vector and all the codebook elements are calculated. These distance values are arranged in the form of a matrix whose rows and columns are the discretized time axis and codebook element number, respectively, and are sent to the DP calculation circuit 14.
sent to.

DP演算回路14には、同時に、単語辞書蓄積部7に蓄
えられているすべての認識対象語紮の単語辞書すなわち
符号帳系列が供給される。DP演算回路14は、入力音
声信号のスペクトル系列(以下、「入カバターン」と呼
ぶ)と、各認識対象語索の符号帳系列で表現されるスペ
クトル系列(以下、「標準パターン」と呼ぶ)との類似
の度合(距II)を計算するものである。音声の発声速
度は、同じ話者が同じ単語を繰り返し発声しても、その
度に部分的及び全体的に変化するので、両者を比較する
には、共通の音(音韻)が対応するように、一方の時間
軸を適当に非線形に伸縮して他方の時間軸に合わせ、対
応する時点のパラメータベクトルどうしを比較する必要
がある。この演算は、距離行列と単語辞書を用いた動的
計画法(DP)演算によって行うことができることがす
でに知られているので(文献、管材・古井:擬音韻標準
パタンによる大語い単語音声認識、電子通信学会論文誌
、65−d、 8. P、1041 、1982参照)
、これを用いる。
At the same time, the DP calculation circuit 14 is supplied with the word dictionary of all recognition target words stored in the word dictionary storage section 7, that is, the codebook series. The DP calculation circuit 14 calculates a spectral sequence of an input audio signal (hereinafter referred to as an "input cover pattern") and a spectral sequence expressed by a codebook sequence of each recognition target word search (hereinafter referred to as a "standard pattern"). The degree of similarity (distance II) is calculated. Even if the same speaker repeats the same word, the rate of speech changes both partially and completely each time, so in order to compare the two, it is necessary to make sure that the common sounds (phonemes) correspond. , it is necessary to appropriately non-linearly expand or contract one time axis to match the other time axis and compare the parameter vectors at corresponding points in time. It is already known that this operation can be performed by dynamic programming (DP) operation using a distance matrix and a word dictionary (Reference, Kanzai Furui: Large word speech recognition using onomatopoeic standard patterns). , Journal of the Institute of Electronics and Communication Engineers, 65-d, 8. P, 1041, 1982)
, use this.

動的計画法の演算によって標準パターンと入カバターン
の11偵度が最も大きくなるように時間軸を対応付けた
ときの、対応する時点どうしの標準パターンと入カバタ
ーンの距離を全音声区間について平均した値(以下、「
総合的距離」と呼ぶ)を計算する。このようにして得ら
れた総合的距離は、DP演算回路14の出力段に接続さ
れた認識判定回路15に出力される。
The distance between the standard pattern and the input pattern at corresponding points in time was averaged over the entire speech interval when the time axes were associated using dynamic programming calculations so that the reconnaissance of the standard pattern and the input pattern was maximized. value (hereinafter referred to as “
(referred to as "total distance"). The total distance thus obtained is output to the recognition determination circuit 15 connected to the output stage of the DP calculation circuit 14.

!!識判定回路15は、供給された総合的距離のうち、
最も値の小さい、すなわち最も類似の度合が高い標準パ
ターンを判別し、この標準パターンの示す単語を、音声
入力端子1から入力された単語であると判定し、その結
果を出力段に接続されている認識結果出力端子16を介
して出力する。
! ! The identification determination circuit 15 determines which of the supplied total distances are:
The standard pattern with the smallest value, that is, the highest degree of similarity, is determined, the word indicated by this standard pattern is determined to be the word input from the audio input terminal 1, and the result is transmitted to the output terminal connected to the output stage. The recognition result is outputted via the recognition result output terminal 16.

従来においては、適応化に用いる単語をあらかじめ決め
ておいて、その単語を発声した音声から抽出した入カバ
ターンと標準パターンとの動的計画法による時間軸整合
を用い、対応付けられた入カバターンと標準パターンの
スペクトルの差としての移動ベクトルを用いて適応化を
行っていたが、この実施例においては、入力音声と符号
帳とクラスタ化して得られたセントロイドの差に基づい
て算出した移動ベクトルを用いて適応化を行っている。
Conventionally, a word to be used for adaptation is determined in advance, and the matched input pattern and standard pattern are matched on the time axis using dynamic programming. Adaptation was performed using the movement vector as the difference between the spectra of standard patterns, but in this example, the movement vector was calculated based on the difference between the centroids obtained by clustering the input speech and the codebook. Adaptation is performed using .

その結果として、任意の少数の単語音声あるいは短い文
章音声を用いて適応化が行えるようになり、不特定多数
の発声者に対して、従来技術よりも極めて容易に認識精
度の大きな向上を達成することができる。
As a result, adaptation can be performed using an arbitrary number of word speech or short sentence speech, and it is much easier to achieve a large improvement in recognition accuracy than conventional techniques for an unspecified number of speakers. be able to.

この実施例によれば、都市名100単語を認識対象語當
として、男性4名の標準話者の音声から作成した符号帳
と単語辞書(各単語について411類の符号帳系列)を
蓄積しておき、その話者と異なる男性20名の音声に対
して、任意の10単語音声による適応化の後に認識を行
った場合、96.6%の認識精度を得るに至った。適応
化を行わなかつた場合の認識精度は95.1%であった
ことと比較すると、極めて少数の任意の単語による適応
化処理でありながら、明確な改善効果が得られることが
わかる。
According to this embodiment, a codebook and a word dictionary (codebook series of 411 classes for each word) created from the voices of four standard male speakers are accumulated, with 100 city names as recognition target words. When recognition was performed on the voices of 20 male speakers different from the speaker after adaptation using 10 arbitrary words, a recognition accuracy of 96.6% was obtained. When compared with the recognition accuracy of 95.1% without adaptation, it can be seen that a clear improvement effect can be obtained even though the adaptation process is performed using a very small number of arbitrary words.

この実施例においては符号帳全体をクラスタ化したが、
その符号帳を構成する要素がどの標準話者の音声に属す
るかをあらかじめ明示しておき、入力音声信号に最も近
い標準話者に属する符号帳要素だけを取り出して、クラ
スタ化に用いてもよい、この標準話者の選択は、入力音
声信号から算出されたパラメータベクトルと、各標準話
者に属する符号帳要素との距離を算出して、これら算出
された距離に基づいて行うことができる。この方法によ
って、上記と同様の条件で単語音声認識を行った場合、
97,2%のv2va精度を得るに至った。
In this example, the entire codebook was clustered, but
It is also possible to specify in advance to which standard speaker's speech the elements composing the codebook belong, and then extract only the codebook elements belonging to the standard speaker closest to the input speech signal and use them for clustering. This standard speaker selection can be performed by calculating the distance between the parameter vector calculated from the input audio signal and the codebook element belonging to each standard speaker, and based on these calculated distances. When word speech recognition is performed using this method under the same conditions as above,
We achieved a v2va accuracy of 97.2%.

適応化を行わなかった場合に比べて、極めて大きな改善
効果が得られることがわかる。
It can be seen that an extremely large improvement effect can be obtained compared to the case where no adaptation is performed.

また、複数の標準話者の音声を用いて符号帳作成を行う
前に、あらかじめ、この発明の技術を用いて、一人の標
準話者の音声から作成した符号帳に、他の標準話者の音
声信号を適応化しておいてもよい、この方法によってあ
らかじめ標準話者間の適応化を行ってから作成した符号
帳を用い、標準話者の選択は行わずにこの符号帳を入力
音声に適応化させた場合、上記と同様の条件で、97.
4%の認識精度を得るに至った。この方法による改善効
果も極めて大きい。
Furthermore, before creating a codebook using the voices of multiple standard speakers, the technique of this invention can be used to add other standard speakers' voices to the codebook created from the voices of one standard speaker. The speech signal may be adapted in advance. Using this method, a codebook created after adaptation between standard speakers is used, and this codebook is adapted to the input speech without selecting a standard speaker. 97. under the same conditions as above.
A recognition accuracy of 4% was achieved. The improvement effect of this method is also extremely large.

上記の実施例では、認識すべき音声と異なる単語音声を
適応化処理に用いているが、この発明によれば適応化ベ
クトルを算出するための音声として任意の音声が使える
ので、認識すべき未知の音声を一旦蓄積してそれを適応
化処理に用いてもよい、この場合には、話者は適応化の
ための音声を特別に発声する必要がなくなり、−層話者
の負担が少なくなる。
In the above embodiment, word speech different from the speech to be recognized is used for the adaptation process, but according to the present invention, any speech can be used as the speech for calculating the adaptation vector. It is also possible to temporarily store the voices of the speakers and use them for the adaptation process. In this case, the speaker does not need to specially utter the voices for adaptation, which reduces the burden on the speaker. .

上記実施例では入力パラメータベクトルを符号帳に適応
化したが、符号帳を入力パラメータに適応して、その適
応化された符号帳要素と入力パラメータベクトルとの距
離を算出してもよい。
In the above embodiment, the input parameter vector is adapted to the codebook, but the codebook may be adapted to the input parameter and the distance between the adapted codebook element and the input parameter vector may be calculated.

なお、この実施例では、音声の周波数スペクトルを示す
パラメータとして線形予測ケプストラムを用いたが、線
形予測係数、ホルマント周波数、パーコール係数、対数
断面積比、零交差数などを用いてもよい、また、入カバ
ターンと標準パターンの時間軸整合に動的計画法を用い
たが、隠れマルコフモデルなどを用いてもよい。
In this example, the linear prediction cepstrum was used as a parameter indicating the frequency spectrum of the voice, but linear prediction coefficients, formant frequencies, Percoll coefficients, log cross-sectional area ratios, number of zero crossings, etc. may also be used. Although dynamic programming was used to time-axis match the input pattern and the standard pattern, a hidden Markov model or the like may also be used.

〔発明の効果〕〔Effect of the invention〕

以上説明したように、この発明の単語音声認識装置によ
れば、話者が発声した任意の音声を用いで、蓄積されて
いる符号帳に合うように話者の音声信号を適応化するこ
とができ、あるいは話者の音声信号に合うように蓄積さ
れている符号帳を適応化することができるので、少数の
単語音声又は認識すべき音声をそのまま適応化に用いて
、不特定話者に関して高性能な音声認識を行うことがで
きる利点がある。
As explained above, according to the word speech recognition device of the present invention, it is possible to adapt the speaker's speech signal to match the stored codebook using arbitrary speech uttered by the speaker. or the stored codebook can be adapted to match the speaker's speech signal, so a small number of word speech or the speech to be recognized can be used as is for adaptation, and high It has the advantage of being able to perform high-performance speech recognition.

【図面の簡単な説明】[Brief explanation of the drawing]

第1図は、この発明の実施例を示す単語音声認識装置の
ブロック図である。 特許出願人 日本電信電話株式会社
FIG. 1 is a block diagram of a word speech recognition device showing an embodiment of the present invention. Patent applicant Nippon Telegraph and Telephone Corporation

Claims (4)

【特許請求の範囲】[Claims] (1)単語の音声信号の周波数スペクトルおよびパワー
の時間的変化を示すパラメータベクトルを算出するパラ
メータベクトル算出手段と、 一人または複数の標準話者が発生した多数の単語の音声
信号から算出した多数のパラメータベクトルをクラスタ
化して、複数の代表的なベクトル値を符号帳に蓄える符
号帳作成手段と、 各認識対象語彙を、上記符号帳に蓄えられているベクト
ル値の、一つまたは複数の時系列で表現して単語辞書に
蓄える単語辞書作成手段と、入力音声信号から算出され
た多数の入力パラメータベクトルと、上記符号帳の要素
ベクトルを、階層的にクラスタに分割化する手段と、 その分割化された各クラスタごとに上記入力パラメータ
ベクトルを上記符号帳に適応化するための、区分移動方
向ベクトルを決定する手段と、これらの区分移動方向ベ
クトルの加重平均としてすべての入力パラメータベクト
ルを適応化する手段と、 これらの適応化された入力パラメータベクトルと、上記
符号帳の各要素との距離を算出する手段と、 これらの算出された距離と上記単語辞書を用いて、動的
計画法あるいは隠れマルコフモデルによって入力音声と
各認識対象語彙との距離を算出する手段とからなる単語
音声認識装置。
(1) A parameter vector calculation means for calculating a parameter vector indicating temporal changes in the frequency spectrum and power of a word speech signal; codebook creation means for clustering parameter vectors and storing a plurality of representative vector values in a codebook; and one or more time series of vector values stored in the codebook for each recognition target vocabulary. means for creating a word dictionary and storing the expression in a word dictionary; means for hierarchically dividing a large number of input parameter vectors calculated from an input speech signal and the element vectors of the codebook into clusters; means for determining a piecewise movement direction vector for adapting said input parameter vector to said codebook for each cluster, and adapting all input parameter vectors as a weighted average of these piecewise movement direction vectors; a means for calculating distances between these adapted input parameter vectors and each element of the codebook; and a means for calculating distances between these adapted input parameter vectors and each element of the codebook; A word speech recognition device comprising means for calculating the distance between input speech and each recognition target vocabulary using a model.
(2)単語の音声信号の周波数スペクトルおよびパワー
の時間的変化を示すパラメータベクトルを算出するパラ
メータベクトル算出手段と、 一人または複数の標準話者が発声した多数の単語の音声
信号から算出した多数のパラメータベクトルをクラスタ
化して、複数の代表的なベクトル値を符号帳に蓄える符
号帳作成手段と、 各認識対象語彙を、上記符号帳に蓄えられているベクト
ル値の、一つまたは複数の時系列で表現して単語辞書に
蓄える単語辞書作成手段と、入力音声信号から算出され
た多数の入力パラメータベクトルと、上記符号帳の要素
ベクトルを、階層的にクラスタに分割化する手段と、 その分割化された各クラスタごとに上記符号帳を上記入
力パラメータベクトルに適応化するための、区分移動方
向ベクトルを決定する手段と、これらの区分移動方向ベ
クトルの加重平均としてすべての上記符号帳の要素を適
応化する手段と、これらの適応化された符号帳要素と、
入力パラメータベクトルとの距離を算出する手段と、こ
れらの算出された距離と上記単語辞書を用いて、動的計
画法あるいは隠れマルコフモデルによって入力音声と各
認識対象語彙との距離を算出する手段とからなる単語音
声認識装置。
(2) a parameter vector calculation means for calculating a parameter vector indicating temporal changes in the frequency spectrum and power of a word speech signal; codebook creation means for clustering parameter vectors and storing a plurality of representative vector values in a codebook; and one or more time series of vector values stored in the codebook for each recognition target vocabulary. means for creating a word dictionary and storing the expression in a word dictionary; means for hierarchically dividing a large number of input parameter vectors calculated from an input speech signal and the element vectors of the codebook into clusters; means for determining piecewise movement direction vectors for adapting said codebook to said input parameter vector for each cluster, and adapting all said codebook elements as a weighted average of these piecewise movement direction vectors; means for adapting these adapted codebook elements;
means for calculating distances from input parameter vectors; means for calculating distances between input speech and each recognition target vocabulary by dynamic programming or hidden Markov models using these calculated distances and the word dictionary; A word speech recognition device consisting of.
(3)上記符号帳のどの要素がどの標準話者に属するか
を表示する手段と、 入力音声信号から算出された入力パラメータベクトルと
、各標準話者の符号帳要素との距離を算出する手段と、 これら算出された距離に基づいて、上記入力音声信号に
最も近い標準話者の符号帳を選択する手段を有し、 このようにして選択された符号帳を用いることを特徴と
する請求項1記載の単語音声認識装置。
(3) means for displaying which element of the codebook belongs to which standard speaker; and means for calculating the distance between the input parameter vector calculated from the input audio signal and the codebook element of each standard speaker. and means for selecting a codebook of a standard speaker closest to the input audio signal based on these calculated distances, and the codebook selected in this way is used. 1. The word speech recognition device according to 1.
(4)上記複数標準話者による符号帳作成を行う前に、
あらかじめ、一人の標準話者から作成した符号帳を用い
て、他の標準話者の音声信号から算出されたパラメータ
ベクトルをその一人の符号帳に適応化しておくことを特
徴とする請求項1、2又は3に記載の単語音声認識装置
(4) Before creating a codebook using multiple standard speakers,
Claim 1, characterized in that a codebook created from one standard speaker is used in advance to adapt parameter vectors calculated from speech signals of other standard speakers to the codebook of that one standard speaker. 3. The word speech recognition device according to 2 or 3.
JP1042191A 1989-02-21 1989-02-21 Word voice recognition device Pending JPH02220099A (en)

Priority Applications (1)

Application Number Priority Date Filing Date Title
JP1042191A JPH02220099A (en) 1989-02-21 1989-02-21 Word voice recognition device

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
JP1042191A JPH02220099A (en) 1989-02-21 1989-02-21 Word voice recognition device

Publications (1)

Publication Number Publication Date
JPH02220099A true JPH02220099A (en) 1990-09-03

Family

ID=12629117

Family Applications (1)

Application Number Title Priority Date Filing Date
JP1042191A Pending JPH02220099A (en) 1989-02-21 1989-02-21 Word voice recognition device

Country Status (1)

Country Link
JP (1) JPH02220099A (en)

Similar Documents

Publication Publication Date Title
US5745873A (en) Speech recognition using final decision based on tentative decisions
Loizou et al. High-performance alphabet recognition
US5865626A (en) Multi-dialect speech recognition method and apparatus
US6009391A (en) Line spectral frequencies and energy features in a robust signal recognition system
JP3114975B2 (en) Speech recognition circuit using phoneme estimation
Yu et al. Speaker recognition using hidden Markov models, dynamic time warping and vector quantisation
JP4141495B2 (en) Method and apparatus for speech recognition using optimized partial probability mixture sharing
US5459815A (en) Speech recognition method using time-frequency masking mechanism
US6256607B1 (en) Method and apparatus for automatic recognition using features encoded with product-space vector quantization
CN117043857A (en) Methods, devices and computer program products for English pronunciation assessment
US6003003A (en) Speech recognition system having a quantizer using a single robust codebook designed at multiple signal to noise ratios
KR20010102549A (en) Speaker recognition
JPH08123484A (en) Signal synthesizing method and signal synthesizing apparatus
US5832181A (en) Speech-recognition system utilizing neural networks and method of using same
Paliwal Lexicon-building methods for an acoustic sub-word based speech recognizer
CN112750445B (en) Voice conversion method, device and system and storage medium
JP2898568B2 (en) Voice conversion speech synthesizer
Devi et al. A novel approach for speech feature extraction by cubic-log compression in MFCC
Syfullah et al. Efficient vector code-book generation using K-means and Linde-Buzo-Gray (LBG) algorithm for Bengali voice recognition
JPH10254473A (en) Voice conversion method and voice conversion device
Shaikh Naziya et al. Speech recognition system—a review
JP2912579B2 (en) Voice conversion speech synthesizer
Nijhawan et al. Real time speaker recognition system for hindi words
Li Speech recognition of mandarin monosyllables
EP1505572A1 (en) Voice recognition method