JPH0764590A - Voice recognizer - Google Patents
Voice recognizerInfo
- Publication number
- JPH0764590A JPH0764590A JP5209719A JP20971993A JPH0764590A JP H0764590 A JPH0764590 A JP H0764590A JP 5209719 A JP5209719 A JP 5209719A JP 20971993 A JP20971993 A JP 20971993A JP H0764590 A JPH0764590 A JP H0764590A
- Authority
- JP
- Japan
- Prior art keywords
- vector quantization
- probability
- quantization code
- arithmetic processor
- hidden markov
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Abstract
(57)【要約】
【目的】離散型隠れマルコフモデルを用いた音声認識装
置において確率計算の効率化を図る。
【構成】 複数の状態に対応するパラメータを有する離
散型隠れマルコフモデルを、認識対象の単語(あるいは
音節)毎に複数個用意し、入力音声に基づいて得られた
ベクトル量子化コード時系列に対して前記複数個の離散
型隠れマルコフモデルを用いて確率計算を行い、該計算
された確率値に基づいて認識結果を求める音声認識装置
であって、前記確率計算を行う演算プロセッサ4と、前
記複数個の離散型隠れマルコフモデルのパラメータを、
同一の前記ベクトル量子化コードに関するパラメータ毎
に一連のアドレスにまとめて格納するメモリ5(62)
とを備え、演算プロセッサ4は、前記ベクトル量子化コ
ード時系列に対して、各ベクトル量子化コードに対応す
るパラメータをメモリ5からアドレス順に読みだして前
記確率演算を行う。
(57) [Abstract] [Purpose] To improve the efficiency of probability calculation in a speech recognition system using a discrete hidden Markov model. [Structure] A plurality of discrete Hidden Markov Models having parameters corresponding to a plurality of states are prepared for each word (or syllable) to be recognized, and a vector quantization code time series obtained based on an input speech is prepared. And a plurality of discrete Hidden Markov Models for calculating probability and obtaining a recognition result based on the calculated probability value. The parameters of the discrete Hidden Markov Models are
Memory 5 (62) for collectively storing a series of addresses for each parameter relating to the same vector quantization code
The arithmetic processor 4 reads the parameters corresponding to each vector quantization code from the memory 5 in the order of addresses, and performs the probability calculation on the vector quantization code time series.
Description
【0001】[0001]
【産業上の利用分野】本発明は、音声認識装置に係り、
特に大語彙の音声認識処理を既存のハードウェアを用い
て高速に行なうのに好適な音声認識装置に関する。BACKGROUND OF THE INVENTION 1. Field of the Invention The present invention relates to a voice recognition device,
In particular, the present invention relates to a voice recognition device suitable for high-speed voice recognition processing of large vocabulary using existing hardware.
【0002】[0002]
【従来の技術】一般に音声認識には非常に大きな処理量
を要し、音声認識装置の実現には処理の効率化が要求さ
れる。特に最近広く利用されている離散型の隠れマルコ
フモデル(Hidden Markov Model、
以後HMMと記す)では各単語あるいは各音節のモデル
は膨大な量のパラメータを持ち、入力音声の認識に際し
ては全ての単語(あるいは音節)について確率計算を行
なう必要があるため膨大なパラメータの全体にアクセス
する必要が生じる。また、確率計算においては各モデル
が持つ各状態について計算を行なう必要があり処理量は
膨大なものになる。2. Description of the Related Art Generally, speech recognition requires a very large amount of processing, and it is required to improve the processing efficiency in order to realize a speech recognition apparatus. In particular, the Hidden Markov Model (discrete Hidden Markov Model), which has been widely used recently,
In the following, referred to as HMM), the model of each word or each syllable has a huge amount of parameters, and when recognizing the input speech, it is necessary to calculate the probability for all words (or syllables). Need to access. Also, in the probability calculation, it is necessary to calculate for each state of each model, and the amount of processing becomes enormous.
【0003】なお、隠れマルコフモデルの詳細について
は、例えば、“An Introduction toHidden Markov Mode
ls"、IEEE ASSP MAGAZINE、January 1
986、pp4-16に記載されている。For details of the hidden Markov model, see, for example, "An Introduction to Hidden Markov Mode.
ls ", IEEE ASSP MAGAZINE, January 1
986, pp4-16.
【0004】従来から、音声認識における処理量削減手
法として、途中まで計算して可能性が低いとみなされた
候補に関する計算処理を打ち切るビームサーチ方式、計
算量の少ない方法を使って予め認識対象の全候補の中か
ら有望な候補を選択し、選択された候補についてのみ認
識処理を行なう予備選択方式などが試みられている。Conventionally, as a method of reducing the amount of processing in speech recognition, a beam search method that terminates the calculation processing for a candidate that has been calculated halfway and considered to have a low possibility, and a method with a small amount of calculation are used in advance to recognize the object to be recognized. Preliminary selection methods have been attempted, in which a promising candidate is selected from all candidates and recognition processing is performed only on the selected candidate.
【0005】ビームサーチ方式の例としては、電子情報
通信学会論文誌、D、Vol.J71−D No.9
pp.1650−1659、(1988−9)“フレー
ム同期化、ビームサーチ、ベクトル量子化の結合による
DPマッチングの高速化”あるいは電子情報通信学会論
文誌、D−2、Vol.J72−D−2 No.8p
p.1248−1255、(1989−8)“DPビー
ムサーチのしきい値関数の検討”に記載のようなものが
ある。処理量削減の基本的考え方は、途中まで計算して
可能性が低いとみなされた候補に関する計算処理を打ち
切り、計算対象を減らすことである。これらの文献に記
載された例は認識方式としてDPマッチングを対象にし
たものであるが、削減手法自体はHMMにも適用でき
る。As an example of the beam search method, the Institute of Electronics, Information and Communication Engineers, D, Vol. J71-D No. 9
pp. 1650-1659, (1988-9) "High-speed DP matching by combining frame synchronization, beam search, and vector quantization", IEICE Transactions, D-2, Vol. J72-D-2 No. 8p
p. 1248-1255, (1989-8) "Discussion of threshold function of DP beam search". The basic idea of reducing the amount of processing is to cut off the calculation processing for candidates that have been calculated halfway and considered to be unlikely, and reduce the number of calculation targets. Although the examples described in these documents are targeted at DP matching as a recognition method, the reduction method itself can also be applied to HMMs.
【0006】一方、予備選択方式の例としては、日本音
響学会講演論文集、1−3−17、(1986−10)
“大語彙単語音声認識のためのスペクトル動特性を用い
た予備選択法”に記載のようなものがある。上記従来例
では、予め認識対象の単語毎にベクトル量子化のコード
ブックを用意しておき、入力音声の終端が検出された後
に入力音声全体を上記各コードブックを用いてそれぞれ
ベクトル量子化を行ない、このときの量子化歪みを各コ
ードブック毎に累積し、その累積値がある閾値より小さ
いものに対してのみ照合を行なう。On the other hand, as an example of the preselection method, a collection of proceedings of the Acoustical Society of Japan, 1-3-17, (1986-10).
There is a method as described in "Preliminary selection method using spectral dynamics for large vocabulary word speech recognition". In the above conventional example, a vector quantization codebook is prepared in advance for each word to be recognized, and after the end of the input speech is detected, the entire input speech is vector-quantized using each of the above codebooks. , The quantization distortion at this time is accumulated for each codebook, and only the accumulated values smaller than a certain threshold value are compared.
【0007】[0007]
【発明が解決しようとする課題】上記両従来技術は部分
的な処理結果や大まかな計算結果に基づいて認識対象の
中の可能性の低い部分を求め、その部分の処理を省くこ
とにより全体の処理量を減らすものである。もちろんこ
の両手法は処理量の削減に有効な手法であるが、HMM
における確率計算自体の効率化を図るものではない。ま
た、ハードウェア構成の観点から考えた処理速度の高速
化については言及されていない。Both of the above-mentioned prior arts find a part having a low possibility in a recognition object based on a partial processing result or a rough calculation result, and omit the processing of the part so that the whole processing is performed. It reduces the amount of processing. Of course, both of these methods are effective in reducing the amount of processing, but HMM
It does not attempt to improve the efficiency of the probability calculation itself in. Further, there is no mention of speeding up the processing speed from the viewpoint of the hardware configuration.
【0008】本発明の目的は、HMMにおける確率計算
自体の効率化を図り、また、ハードウェア構成の観点か
ら考えた処理の高速化を図り、既存のハードウェアを用
いて、コンパクトでかつ安価な大語彙の音声認識装置を
実現することにある。The object of the present invention is to improve the efficiency of the probability calculation itself in the HMM, to speed up the processing considered from the viewpoint of the hardware configuration, and to use existing hardware to make it compact and inexpensive. It is to realize a large vocabulary voice recognition device.
【0009】[0009]
【課題を解決するための手段】上記本発明の目的を達成
するために、本発明による音声認識装置は、複数の状態
に対応するパラメータを有する離散型隠れマルコフモデ
ルを、認識対象の単語(あるいは音節)毎に複数個用意
し、入力音声に基づいて得られたベクトル量子化コード
時系列に対して前記複数個の離散型隠れマルコフモデル
を用いて確率計算を行い、該計算された確率値に基づい
て認識結果を求める音声認識装置であって、前記確率計
算を行う演算プロセッサと、前記複数個の離散型隠れマ
ルコフモデルのパラメータを、同一の前記ベクトル量子
化コードに関するパラメータ毎に一連のアドレスにまと
めて格納するメモリとを備え、前記演算プロセッサは、
前記ベクトル量子化コード時系列に対して、各ベクトル
量子化コードに対応するパラメータを前記メモリからア
ドレス順に読みだして前記確率演算を行うようにしたも
のである。In order to achieve the above object of the present invention, a speech recognition apparatus according to the present invention uses a discrete Hidden Markov Model having parameters corresponding to a plurality of states as a recognition target word (or For each syllabic), a probability calculation is performed on the vector quantization code time series obtained based on the input speech using the plurality of discrete hidden Markov models, and the calculated probability value is calculated. A speech recognition apparatus for obtaining a recognition result based on a calculation processor for performing the probability calculation, and parameters of the plurality of discrete Hidden Markov Models to a series of addresses for each parameter related to the same vector quantization code. And a memory for collectively storing the arithmetic processor,
With respect to the vector quantization code time series, the parameters corresponding to each vector quantization code are read from the memory in the order of addresses, and the probability operation is performed.
【0010】[0010]
【作用】本発明の音声認識装置では、HMMを用いた確
率計算を実行する上で、効率良くHMMのデータにアク
セスできる様にパラメータ(データ)をメモリ内に配置
し、かつ、連続したアドレスを効率良くアクセスできる
メモリを用いるので、高速なHMM確率計算が実行でき
る。In the speech recognition apparatus of the present invention, in executing the probability calculation using the HMM, the parameters (data) are arranged in the memory so that the data of the HMM can be efficiently accessed, and consecutive addresses are set. Since the memory that can be efficiently accessed is used, high-speed HMM probability calculation can be executed.
【0011】また、HMMの構造上事前に計算できる遷
移確率と出現確率の掛け合わせ計算(対数領域で行なえ
ば足し合わせ計算)は予め全て行なった上でその結果を
モデルのパラメータの中に収めておき、認識時にはこれ
を用いて確率計算することにより、認識時の計算量を軽
減し、一層の高速化が図れる。Further, the multiplication calculation of the transition probabilities and the appearance probabilities that can be calculated in advance due to the structure of the HMM (addition calculation if performed in the logarithmic domain) is performed in advance, and the result is stored in the parameters of the model. Every time, the probability is calculated using this at the time of recognition, so that the amount of calculation at the time of recognition can be reduced and the speed can be further increased.
【0012】さらに、各HMM毎のベクトル量子化コー
ドの出現頻度情報を使って対応するHMMの可能性を判
定し、可能性の低いHMMについては確率計算を省略す
ることで、より一層の高速化が図れる。Further, the probability of the corresponding HMM is judged using the appearance frequency information of the vector quantization code for each HMM, and the probability calculation is omitted for the HMM with a low probability, so that the speed is further increased. Can be achieved.
【0013】さらに、複数の演算プロセッサで認識対象
となる単語(あるいは音節)を分担することにより、大
語彙に対処することができる。Furthermore, a large vocabulary can be dealt with by sharing words (or syllables) to be recognized by a plurality of arithmetic processors.
【0014】以上、各種の高速化手法を総合した本発明
によれば、最新の安価なハードウェアを用いて、コンパ
クトで安価でかつ高速な大語彙の音声認識装置を実現す
ることができる。As described above, according to the present invention in which various speed-up techniques are integrated, it is possible to realize a compact, inexpensive and high-speed large-vocabulary speech recognition device by using the latest inexpensive hardware.
【0015】[0015]
【実施例】以下、本発明の実施例を説明する。本発明は
単音節認識、単語認識、文章認識など各種の音声認識に
適用できるが、ここでは簡単のため単語認識を取り上げ
て説明する。EXAMPLES Examples of the present invention will be described below. The present invention can be applied to various types of speech recognition such as single syllable recognition, word recognition, and sentence recognition, but here, for the sake of simplicity, word recognition will be taken up and described.
【0016】図1は本発明の音声認識装置のハードウェ
ア構成を示すブロック図である。マイク1から入力され
た音声はオーディオアンプ2において増幅される。増幅
された音声信号はAD変換器3において一定時間間隔
(例えば8kHzサンプリングでは125μs)毎に取
り込まれディジタル化される。ディジタル化された音声
信号は演算プロセッサ部4において外部メモリ5の内容
を参照しながら各種処理が施され最終的に認識結果が得
られる。FIG. 1 is a block diagram showing the hardware structure of the speech recognition apparatus of the present invention. The audio input from the microphone 1 is amplified by the audio amplifier 2. The amplified audio signal is taken in the AD converter 3 at regular time intervals (for example, 125 μs in 8 kHz sampling) and digitized. The digitized voice signal is subjected to various kinds of processing in the arithmetic processor unit 4 while referring to the contents of the external memory 5, and finally a recognition result is obtained.
【0017】演算プロセッサ部4における処理を図3の
フローチャートを用いて説明する。The processing in the arithmetic processor unit 4 will be described with reference to the flowchart of FIG.
【0018】説明の簡単化のため「0」から「9」まで
の10個の数字の音声認識を例に挙げて説明する。For simplification of the description, voice recognition of ten numbers from "0" to "9" will be described as an example.
【0019】マイクに向かって数字音声、例えば、「1
(イチ)」と発声されると、AD変換器3においては音
声信号が一定時間間隔毎(例えば12kHzサンプリン
グの場合には83.3μs毎)に取り込まれデジタル化
される(S31、S32)。演算プロセッサ部4では、
音声データがデジタル化されたサンプルデータが得られ
る毎に自己相関関数の計算を行う(S33)。自己相関
関数の計算は、1サンプルデータが得られる毎に部分的
な計算を行う。式で表現すると、次のようになる。ここ
で、riは第i次の自己相関係数の部分的な結果を格納
する変数、xtは時刻tのサンプルデータを表わす。な
お、ここでは自己相関の次数を14次とする。A numerical voice, for example, "1" is input into the microphone.
("), A voice signal is taken in the AD converter 3 at regular time intervals (for example, every 83.3 μs in the case of 12 kHz sampling) and digitized (S31, S32). In the arithmetic processor unit 4,
The autocorrelation function is calculated every time sample data obtained by digitizing voice data is obtained (S33). The calculation of the autocorrelation function is performed partially every time one sample data is obtained. Expressed as an expression, it is as follows. Here, ri represents a variable for storing a partial result of the i-th-order autocorrelation coefficient, and xt represents sample data at time t. Note that the order of autocorrelation is 14th here.
【0020】ri = ri + xt × xt-i 予め決められたデータポイント数分(例えば、分析窓長
を20msとし、12kHzサンプリングとすると、2
40点)だけ上記の計算が行われると、1フレーム分の
自己相関関数ri(i=0〜14)が確定し、フレーム
単位の処理に進む(S35)。周波数分析の周期(フレ
ーム周期)を分析窓長と同じ20msとすれば、20m
sに1度ずつ自己相関係数が確定し、フレーム単位の処
理が行われることになる。Ri = ri + xt × xt-i For a predetermined number of data points (for example, when the analysis window length is 20 ms and 12 kHz sampling is performed, 2
When the above calculation is performed for 40 points), the autocorrelation function ri (i = 0 to 14) for one frame is determined, and the process proceeds in frame units (S35). If the period of the frequency analysis (frame period) is 20 ms, which is the same as the analysis window length, 20 m
The autocorrelation coefficient is determined once every s, and the processing is performed in frame units.
【0021】フレーム単位の処理では、まず、自己相関
関数から線形予測係数を計算し(S351)、さらに線
形予測係数からケプストラム係数を求める(S35
2)。求まったケプストラム係数は多次元のベクトルと
みなされ、予め用意したベクトル量子化コードブックを
用いてベクトル量子化し、ベクトル量子化コードを得る
(S353)。ベクトル量子化のレベル(コードブック
のサイズ)としては任意の値を取ることができるが、本
実施例では256とする。すなわち、ベクトル量子化後
には、量子化コードk(1から256までのいずれかの
正数値)が得られる。すなわち、入力された単語音声
「1(イチ)」の音声長がLフレーム分あったとする
と、長さLのコード系列が得られることになる。In the frame unit processing, first, a linear prediction coefficient is calculated from the autocorrelation function (S351), and a cepstrum coefficient is obtained from the linear prediction coefficient (S35).
2). The obtained cepstrum coefficient is regarded as a multidimensional vector, and vector quantization is performed using a vector quantization codebook prepared in advance to obtain a vector quantization code (S353). The vector quantization level (codebook size) can take any value, but is 256 in this embodiment. That is, after vector quantization, the quantization code k (any positive value from 1 to 256) is obtained. That is, assuming that the voice length of the input word voice "1" is L frames, a code sequence of length L is obtained.
【0022】単語音声認識は、予め用意された認識対象
すべて(今の例では10数字のすべて)の離散型隠れマ
ルコフモデル(HMM)について、上記の長さLのコー
ド系列を出力する確率を計算し、最も確率の高いHMM
を認識結果とする。In word speech recognition, the probability of outputting the above-mentioned code sequence of length L is calculated for discrete Hidden Markov Models (HMMs) of all recognition targets (all 10 numbers in this example) prepared in advance. The most probable HMM
Is the recognition result.
【0023】実際のHMMの確率計算処理は、長さLの
コード系列が求まってから行うわけではなく、図3のフ
ローチャートに示すように、ベクトル量子化コードkが
一つ求まる毎に実施し、音声終端が検出されたか否かを
判定し(S36)、検出された場合にはソーティング/
候補出力の処理に進む。音声終端が検出されず、入力音
声が継続している間、フローチャートの先頭に戻り、自
己相関の計算、フレーム単位の処理を継続する。The actual HMM probability calculation process is not performed after the length L code sequence is obtained, but is performed each time one vector quantization code k is obtained, as shown in the flowchart of FIG. It is determined whether or not the voice end is detected (S36), and if detected, sorting /
Proceed to the candidate output process. While the voice end is not detected and the input voice continues, the flow returns to the beginning of the flowchart to continue the autocorrelation calculation and the frame unit processing.
【0024】以上の処理が行われ、音声の終端が検出さ
れると確率計算を終了し、各単語(「0」から「9」の
10数字)の確率値を各単語のスコアとし、このスコア
に基づいて各単語をソーティングする(S37)。ソー
ティングされた上位L(例えばL=3)候補を認識結果
として出力する(S38)。例えば、今の例では、単語
「1」に対する確率値が高くなり、そのHMMが1位と
して出力されれば正解認識となる。When the above processing is performed and the end of the voice is detected, the probability calculation ends, and the probability value of each word (10 numbers from "0" to "9") is set as the score of each word. Each word is sorted based on (S37). The sorted upper L (eg L = 3) candidates are output as the recognition result (S38). For example, in the present example, if the probability value for the word “1” becomes high and the HMM is output as the first place, the correct answer is recognized.
【0025】なお、線形予測係数を求める処理およびケ
プストラム係数を求める処理については、例えば、古井
「ディジタル音声処理」東海大学出版などに記載されて
いる手法を使えばよい。また、本実施例の図3のフロー
チャートでは、HMMを用いた確率計算を他のフレーム
単位の処理と同期して行うような構成としているが、H
MMを用いた確率計算部分を別のプロセスとして独立さ
せ、マルチプロセスで実行することももちろん可能であ
る。For the processing for obtaining the linear prediction coefficient and the processing for obtaining the cepstrum coefficient, for example, the method described in Furui "Digital Speech Processing" Tokai University Press, etc. may be used. Further, in the flowchart of FIG. 3 of the present embodiment, the probability calculation using the HMM is configured to be performed in synchronization with other frame unit processing.
It is of course possible to make the probability calculation part using MM independent as another process and execute it in multiple processes.
【0026】つぎに、図3のフローチャートの中のHM
Mを用いた確率計算部分(S354)について詳細に説
明する。まず、HMMについて図4のHMMの説明図を
用いて説明する。HMMはいくつかの状態(状態数をN
とする。)を持った状態遷移モデルであり、各状態遷移
に対してその状態遷移が生じる確率(遷移確率)、およ
びその状態遷移が生じた際に各ベクトル量子化コードが
出現する確率(出現確率)が定義されている。状態数N
は、例えば、単語の場合には20程度、音節の場合には
5程度である。HMMは音声を表現するモデルである
が、単語を単位としてモデル化する場合(単語HMM)
と音節のような小さい単位毎にモデルを持ち(音節HM
M)これら小さなモデルの結合により単語を表す場合が
ある。本実施例では単語HMMの場合を考える。認識対
象の語彙がM(例えばM=10)個の場合、M個のHM
Mを用意する。Next, the HM in the flowchart of FIG.
The probability calculation part (S354) using M will be described in detail. First, the HMM will be described with reference to the HMM explanatory diagram of FIG. HMM has several states (number of states N
And ), The probability that the state transition will occur for each state transition (transition probability), and the probability that each vector quantization code will appear (occurrence probability) when the state transition occurs It is defined. Number of states N
Is, for example, about 20 for words and about 5 for syllables. HMM is a model that expresses speech, but when modeling with words as units (word HMM)
And a model for each small unit such as a syllable (syllable HM
M) A word may be represented by the combination of these small models. In this embodiment, the case of the word HMM will be considered. If the vocabulary to be recognized is M (for example, M = 10), M HMs
Prepare M.
【0027】HMMはN個の状態を持つ状態遷移モデル
であるが、ここでは状態数Nを5として説明する。図4
に示すのは、ある単語w(1〜Mのいずれか)に対応し
た、5状態を持つ一つのHMMである。図4で丸で示し
たのが状態であり丸の中の数字が状態番号に対応する。
状態と状態の間で遷移が許されている部分は矢印(アー
ク)で結ばれている。一般に、HMMは任意の状態から
任意の状態への状態遷移を許すが、ここでは音声認識で
良く用いられるleft−to−rightのモデルを
取り上げる。left−to−rightのモデルでは
自分自身への状態遷移と一つ先の状態(一つ番号の大き
い状態)への状態遷移のみを許す。各状態遷移には、状
態遷移確率(図中記号aで表示)と、その時の各ベクト
ル量子化コードの出現確率(図中記号bで表示)が付随
する。Although the HMM is a state transition model having N states, the number N of states will be described as 5 here. Figure 4
Shown in (1) is one HMM having five states, which corresponds to a certain word w (any one of 1 to M). The state is indicated by a circle in FIG. 4, and the number in the circle corresponds to the state number.
Portions that allow transitions between states are connected by arrows (arcs). Generally, the HMM allows a state transition from an arbitrary state to an arbitrary state, but here, a left-to-right model often used in speech recognition is taken up. The left-to-right model allows only the state transition to itself and the state transition to the next state (state with a larger number). Each state transition is accompanied by a state transition probability (indicated by symbol a in the figure) and an appearance probability of each vector quantization code (indicated by symbol b in the figure) at that time.
【0028】ベクトル量子化コード時系列k(1)、k
(2)、k(3)・・・k(t)(k(t)は1〜25
6の間の整数値)を観測して、単語wのHMMの状態i
にいる確率をP(w、i、t)と表わすことにする。M
単語の音声認識の問題は、P(w、N、T)(w=1〜
M)が最大値を与えるwを求める問題と考えることがで
きる。したがって、P(w、i、t)の計算が直接音声
認識処理につながる。P(w、i、t)の計算にはいく
つかの方法があるが、ここではビタビアルゴリズムと呼
ばれる手法を使うことにする。計算に先だって次式にし
たがって初期設定を行なう。Vector quantization code time series k (1), k
(2), k (3) ... k (t) (k (t) is 1 to 25)
6) and the HMM state i of the word w
Let us denote the probability of being at P (w, i, t). M
The problem of word voice recognition is P (w, N, T) (w = 1 to
It can be considered as a problem in which M) finds w that gives the maximum value. Therefore, the calculation of P (w, i, t) directly leads to the voice recognition process. There are several methods for calculating P (w, i, t), but here we will use a method called the Viterbi algorithm. Prior to the calculation, the initialization is performed according to the following formula.
【0029】 P(w,i,0) = 1 (i=1,w=1〜M) ・・・(1) P(w,i,0) = 0 (i≠1,w=1〜M) ・・・(2) 以後、ベクトル量子化コードk(t)(k(t)=1〜
256)が得られる毎に各単語の各状態について次式に
したがって確率値更新を行なう。P (w, i, 0) = 1 (i = 1, w = 1 to M) (1) P (w, i, 0) = 0 (i ≠ 1, w = 1 to M) ) (2) After that, vector quantization code k (t) (k (t) = 1 to
256) is obtained, the probability value is updated for each state of each word according to the following equation.
【0030】 wk1 = P(w,i-1,t-1)×a(w,i-1,i)×b(w,i-1,i,k(t)) ・・・(3) wk2 = P(w,i ,t-1)×a(w,i ,i)×b(w,i ,i,k(t)) ・・・(4) P(w,i,t) = max(wk1、wk2) ・・・(5) ここで、a(w、i、j)は単語wの状態iから状態j
への遷移確率、b(w、i、j、k(t))は単語wの
状態iから状態jへの遷移においてベクトル量子化コー
ドk(t)が出現する確率である。以上の計算フローは
フローチャートで示すと図5の様になる。なお、全ての
確率値を対数領域で表わすようにすれば、上記式(3)
(4)の確率計算中の乗算は全て加算に置き換えること
ができる。Wk1 = P (w, i-1, t-1) × a (w, i-1, i) × b (w, i-1, i, k (t)) (3) wk2 = P (w, i, t-1) × a (w, i, i) × b (w, i, i, k (t)) (4) P (w, i, t) = max (wk1, wk2) (5) where a (w, i, j) is from state i to state j of word w
, B (w, i, j, k (t)) is the probability that the vector quantized code k (t) appears in the transition from the state i to the state j of the word w. The above calculation flow is shown in a flowchart of FIG. If all probability values are expressed in the logarithmic domain, the above equation (3)
All multiplications during the probability calculation of (4) can be replaced with additions.
【0031】以上が1フレーム間のHMMの確率計算で
あるが、これを音声終端が検出されるまで繰り返し、最
終的にP(w、N、T)が全M単語について求まり、こ
れの上位のものを選ぶことで認識結果が得られる。The above is the HMM probability calculation for one frame, and this is repeated until the voice end is detected, and finally P (w, N, T) is obtained for all M words, and the higher order of these is obtained. The recognition result can be obtained by selecting one.
【0032】次に、上記HMMのデータのメモリ内での
配置について説明する。Next, the arrangement of the HMM data in the memory will be described.
【0033】上記HMMを用いた確率計算の説明におい
て示したように、1つの単語のHMMあたり、状態遷移
確率a(w、i、j)(i=1〜5、j=i、i+1)
が10ワード、ベクトル量子化コード出現確率b(w、
i、j、k)(i=1〜5、j=i、i+1、k=1〜
256)が2560ワードの計2570ワードのデータ
からなる。このデータをメモリ内でどの様に配置するか
には様々なバラエティが考えられる。As shown in the description of the probability calculation using the HMM, the state transition probability a (w, i, j) (i = 1 to 5, j = i, i + 1) per HMM of one word.
Is 10 words, vector quantization code appearance probability b (w,
i, j, k) (i = 1 to 5, j = i, i + 1, k = 1 to
256) consists of 2560 words of data, totaling 2570 words. Various varieties are conceivable as to how this data is arranged in the memory.
【0034】最も単純には、図6のa)に示す様に、ま
ず各単語毎にまとめて格納し、各単語内では各状態毎に
まとめ、各状態内では自状態への遷移と次状態への遷移
の2つの部分に分け、各部分内ではまず遷移確率aを格
納しこれに続いて256ワード分のベクトル量子化コー
ドの出現確率bをアドレス順に格納するという方法が考
えられる。しかしながら、このようにデータを配置する
と上記(3)(4)式の確率更新計算においてメモリ内
の飛び飛びのアドレスにアクセスする必要が生じ効率が
良くない。In the simplest case, as shown in FIG. 6A, the words are first stored together and stored in each word, and each word is grouped into each state, and in each state, transition to the own state and the next state are stored. It is conceivable that the method is divided into two parts, that is, the transition probability a is first stored in each part, and subsequently, the probability b of appearance of the vector quantization code of 256 words is stored in the order of address. However, if the data is arranged in this way, it is necessary to access the discrete addresses in the memory in the probability update calculation of the above formulas (3) and (4), which is not efficient.
【0035】そこで、HMMのデータを大幅に並び替
え、図6のb)に示すようにする。すなわちHMMのデ
ータを単語毎に整理するのではなく、ベクトル量子化コ
ード毎に整理する。特定のベクトル量子化コードkにつ
いてのHMMの情報は全て局所的なアドレス領域にまと
めて格納される。図6のb)に示す並びであると、ベク
トル量子化コードkが定まるとそのフレームにおける確
率計算に必要なHMMのデータは局所的な領域にまとめ
て置かれることになり、かつ、式(3)(4)の計算順
序に合わせた形でアドレス順にデータが格納されるの
で、確率計算の最中のHMMデータの参照はアドレスの
インクリメントだけで実行される。図6のb)に示す並
びでは状態遷移確率a(w、i、j)を各ベクトル量子
化コードの出現確率b(w,i,j,k)と対で格納す
るため、データ量は図6のa)に示すような格納の仕方
の場合のほぼ2倍になってしまうが、計算効率は高くな
る。Therefore, the HMM data is rearranged to a large extent, as shown in FIG. 6B). That is, the HMM data is not sorted for each word, but for each vector quantization code. All HMM information about a specific vector quantization code k is stored collectively in a local address area. With the arrangement shown in b) of FIG. 6, when the vector quantization code k is determined, the HMM data necessary for the probability calculation in that frame are put together in a local area, and the equation (3 ) Since the data is stored in the address order in a form that matches the calculation order of (4), the reference of the HMM data during the probability calculation is executed only by incrementing the address. In the arrangement shown in b) of FIG. 6, the state transition probability a (w, i, j) is stored as a pair with the appearance probability b (w, i, j, k) of each vector quantization code. Although it is almost twice as large as the case of the storage method shown in 6 a), the calculation efficiency is high.
【0036】なお、図6のb)では予めHMMのデータ
を計算効率を高めるようにメモリ内で並べ替えておいた
が、図2に示すように演算プロセッサ6内部にデータ転
送制御部63を設け、該データ転送制御部63が、ベク
トル量子化結果kが得られる毎にベクトル量子化コード
kに関するHMMのデータのみを外部メモリ5から取り
だし、これを演算プロセッサ部61の内部メモリ62の
一連のアドレス領域に収めるようにし、内部メモリ62
を用いて前記HMMの確率計算をするようにすれば同様
の計算の効率化が図れる。In FIG. 6 b), the HMM data is rearranged in the memory in advance so as to improve the calculation efficiency. However, as shown in FIG. 2, the data transfer control unit 63 is provided inside the arithmetic processor 6. , The data transfer control unit 63 fetches only the HMM data relating to the vector quantization code k from the external memory 5 every time the vector quantization result k is obtained, and outputs this to a series of addresses in the internal memory 62 of the arithmetic processor unit 61. Internal memory 62
If the HMM probability calculation is performed using, the same calculation efficiency can be achieved.
【0037】なお、実施例において、内部メモリまたは
外部メモリとして、RAMbusや同期型DRAM等の
連続したアドレスを演算プロセッサから効率よくアクセ
スできるメモリを用いてもよい。このようなメモリにつ
いては、例えば、日経エレクトロニクス1992.3.
16(no549)第95〜97頁、日経エレクトロニ
クス1992.5.11(no553)第143〜14
7頁に開示されている。In the embodiment, as the internal memory or the external memory, a memory capable of efficiently accessing continuous addresses such as a RAMbus or a synchronous DRAM may be used. For such a memory, for example, Nikkei Electronics 1992.3.
16 (no549), pp. 95-97, Nikkei Electronics 1992.5.11 (no553), 143-14.
It is disclosed on page 7.
【0038】次に、上記HMMを用いた確率計算におい
て、HMM計算の構造上事前に実行できる計算の事前実
行について説明する。式(3)(4)を見ると、同一の
式の中で現れる配列要素a(w、i、j)とb(w、
i、j、k)の添字はkを除いては全て同じであること
がわかる。すなわち、式(3)(4)式中、 a(w、i、j)×b(w、i、j、k) ・・・(6) の乗算は事前に実行できる性格のものであり、式(6)
の計算を事前に行ない、その結果をb’(w、i、j、
k)として、 b’(w,i,j,k) = a(w,i,j)×b(w,i,j,k) ・・・(7) を新たなパラメータとして格納し、これを用いてHMM
の確率計算を行なうことができる。この様にすれば式
(3)(4)と式(8)(9)の比較から明らかなよう
に計算量をほぼ半減することができる。HMMを用いた
確率計算では、式(3)(4))に代わって次式(8)
(9)を用いることになる。Next, in the probability calculation using the HMM, the pre-execution of the calculation which can be pre-executed due to the structure of the HMM calculation will be described. Looking at equations (3) and (4), array elements a (w, i, j) and b (w, which appear in the same equation)
It can be seen that the subscripts of i, j, k) are all the same except for k. That is, in the formulas (3) and (4), the multiplication of a (w, i, j) × b (w, i, j, k) ... (6) has a character that can be executed in advance, Formula (6)
Is calculated in advance, and the result is b ′ (w, i, j,
k), b ′ (w, i, j, k) = a (w, i, j) × b (w, i, j, k) (7) is stored as a new parameter, and Using HMM
The probability of can be calculated. This makes it possible to reduce the amount of calculation by almost half, as is clear from the comparison between the equations (3) and (4) and the equations (8) and (9). In probability calculation using HMM, the following equation (8) is used instead of equations (3) and (4)).
(9) will be used.
【0039】 wk1 = P(w,i-1,t-1)× b’(w,i-1,i,k(t)) ・・・(8) wk2 = P(w,i ,t-1)× b’(w,i ,i,k(t)) ・・・(9) P(w、i、t) = max(wk1、wk2) ・・・(10) このときのメモリ内のHMMのデータの配置は図7に示
すようになる。すなわち、遷移確率のデータと出現確率
のデータが事前に掛け合わされ、従来2ワード必要とし
ていた情報が1ワードに収められる。従ってメモリ量も
図6のb)の場合の配置と比べると半減する。Wk1 = P (w, i-1, t-1) × b ′ (w, i-1, i, k (t)) (8) wk2 = P (w, i, t- 1) × b ′ (w, i, i, k (t)) ・ ・ ・ (9) P (w, i, t) = max (wk1, wk2) ・ ・ ・ (10) The arrangement of HMM data is as shown in FIG. That is, the transition probability data and the appearance probability data are preliminarily multiplied, and the information that conventionally requires two words is stored in one word. Therefore, the memory amount is also halved as compared with the arrangement in the case of FIG.
【0040】つぎに複数の演算プロセッサを用いて認識
対象の語彙数を増やす場合の実施例を図8を用いて説明
する。Next, an embodiment in which the number of words to be recognized is increased by using a plurality of arithmetic processors will be described with reference to FIG.
【0041】単一の演算プロセッサでは処理能力に限界
があり認識できる語彙数も自ずと限られてしまう。演算
プロセッサを複数化するのが一つの解である。図8の実
施例では演算プロセッサの数を3個としているが、特に
演算プロセッサの数に制限がある訳ではない。図8中、
第1の演算プロセッサ71では図1に示した実施例にお
ける演算プロセッサ4とほぼ同じ処理を行なうが、ベク
トル量子化結果を他の全ての演算プロセッサ73,75
に送出する点、他の演算プロセッサから他の演算プロセ
ッサが担当している単語のHMMの確率計算結果を受け
とる点が異なる。他の演算プロセッサ73,75は、第
1の演算プロセッサ71からベクトル量子化結果kを受
けとり、これを用いてHMMの確率計算を行なう。音声
終端検出後に、HMM確率計算の最終結果を第1の演算
プロセッサ71に返す。第1の演算プロセッサ71で
は、自分が担当した単語のHMMの確率計算結果および
他の演算プロセッサから受けとった他の単語に関するH
MMの確率計算結果の全てを総合して認識結果を求め
る。以上により、語彙数の増加に対して容易に対処でき
る。本実施例の場合、図8に示すように各演算プロセッ
サ毎に一定数の単語を担当するようになるが担当する単
語のHMMのデータは各演算プロセッサ毎に個別に設け
られた外部メモリ72,74,76に格納される。した
がって、この場合には各外部メモリ72,74,76に
担当する単語毎のHMMのデータを格納する必要があ
り、図6のb)に示した様なメモリ配置にする訳にはい
かない。メモリアクセスの効率および単語毎の独立性を
両立させることを考えると、本実施例におけるHMMの
データのメモリ配置は図9に示すようなものとなる。す
なわち、HMMのデータは単語毎に分割して保持し、単
語内ではベクトル量子化コード毎に整理する形となる。With a single arithmetic processor, the processing capacity is limited and the number of vocabularies that can be recognized is naturally limited. One solution is to use multiple arithmetic processors. Although the number of arithmetic processors is three in the embodiment of FIG. 8, the number of arithmetic processors is not particularly limited. In FIG.
The first arithmetic processor 71 performs almost the same processing as that of the arithmetic processor 4 in the embodiment shown in FIG. 1, but the vector quantization result is applied to all the other arithmetic processors 73 and 75.
The difference is that the HMM probability calculation result of the word for which another arithmetic processor is in charge is received from another arithmetic processor. The other arithmetic processors 73 and 75 receive the vector quantization result k from the first arithmetic processor 71 and use it to calculate the HMM probability. After the voice end is detected, the final result of the HMM probability calculation is returned to the first arithmetic processor 71. In the first arithmetic processor 71, the HMM probability calculation result of the word for which it is in charge and the H related to another word received from another arithmetic processor.
A recognition result is obtained by integrating all the MM probability calculation results. From the above, it is possible to easily deal with the increase in the number of vocabularies. In the case of the present embodiment, as shown in FIG. 8, each arithmetic processor is in charge of a certain number of words, but the HMM data of the in-charge word is stored in the external memory 72 provided individually for each arithmetic processor. 74 and 76. Therefore, in this case, it is necessary to store the data of the HMM for each word in charge in the external memories 72, 74, and 76, and the memory arrangement as shown in FIG. 6B) cannot be achieved. Considering both the efficiency of memory access and the independence of each word, the memory layout of the HMM data in this embodiment is as shown in FIG. That is, the HMM data is divided and held for each word, and is arranged for each vector quantization code within the word.
【0042】つぎに、出現確率がある基準より低いHM
Mについて確率計算を省略することにより、全体の計算
量を削減し認識処理を高速化する手法について説明す
る。Next, HM whose appearance probability is lower than a certain standard
A method of reducing the overall calculation amount and speeding up the recognition processing by omitting the probability calculation for M will be described.
【0043】式(3)(4)から判るように、特定のベ
クトル量子化コードについての出現確率が非常に小さな
値をとるとき、そのHMMの確率は非常に小さな値とな
り、このHMMが最終的に認識結果として残る可能性は
低くなる。そこで、予め各HMM中の各ベクトル量子化
コードの出現確率を調べておき、この出現確率が非常に
低いベクトル量子化結果が得られたときにはそのHMM
の確率計算を省略することができる。各ベクトル量子化
コード毎にその出現確率が予め決められた基準より低い
遷移の存在するHMMをリストアップして図10に示す
ようなテーブルを作成する。認識時には、図10のテー
ブルを引き、ベクトル量子化結果から出現確率の低いH
MMを求め、このHMMについては確率計算を省略する
ようにする。以上により確率計算を大幅に省略できより
高速な音声認識ができる。As can be seen from the equations (3) and (4), when the appearance probability for a particular vector quantization code has a very small value, the probability of the HMM becomes a very small value, and this HMM is the final value. It is less likely to remain as a recognition result. Therefore, the appearance probability of each vector quantization code in each HMM is checked in advance, and when a vector quantization result with a very low appearance probability is obtained, the HMM
The probability calculation of can be omitted. For each vector quantization code, HMMs having transitions whose appearance probability is lower than a predetermined reference are listed to create a table as shown in FIG. At the time of recognition, the table of FIG.
MM is obtained, and probability calculation is omitted for this HMM. As a result, the probability calculation can be largely omitted and faster voice recognition can be performed.
【0044】なお、本手法を導入した場合の音声認識の
流れを図11のフローチャートに示す。本フローチャー
トは、図3に示したフローチャートに対して出現確率に
よる計算省略のステップS111、S112を挿入した
ものとなっている。The flow of voice recognition when this method is introduced is shown in the flowchart of FIG. This flowchart is obtained by inserting steps S111 and S112 for omitting calculation based on the occurrence probability into the flowchart shown in FIG.
【0045】次に、一定の時間長の区間のベクトル量子
化コードの統計情報を用いて、HMMの確率計算を省略
する手法について説明する。Next, a method of omitting the probability calculation of the HMM by using the statistical information of the vector quantization code in the section of constant time length will be described.
【0046】一定の時間長として例えば図12のa)に
示すように処理対象のフレームの前後10フレームずつ
計21フレーム(フレーム周期を20msとすれば約4
00msの区間)を考える。統計情報として図12の
b)に示すようなヒストグラムを算出する。このヒスト
グラムは21個の量子化コードについて、どのコードが
いくつあるかをカウントするだけで得られる。一方、各
HMMは各遷移毎に各ベクトル量子化コードの出現確率
を持っているが、これはそのままヒストグラムに対応す
る。そこで、この各遷移毎のヒストグラムを全遷移で平
均すればやはりヒストグラムを得ることができ、これを
このHMMのヒストグラムとして考えることができる。
両ヒストグラムを比較し、類似性が予め決められた基準
より低いときには、そのHMMについての確率計算を省
略する。As a fixed time length, for example, as shown in FIG. 12A, 10 frames before and 10 frames after the frame to be processed, a total of 21 frames (about 4 when the frame period is 20 ms).
Consider the 00 ms interval). A histogram as shown in b) of FIG. 12 is calculated as statistical information. This histogram can be obtained by simply counting the number of each of the 21 quantized codes. On the other hand, each HMM has an appearance probability of each vector quantization code for each transition, which corresponds to the histogram as it is. Therefore, if the histogram for each transition is averaged over all transitions, a histogram can be obtained, and this can be considered as a histogram of this HMM.
The two histograms are compared, and when the similarity is lower than a predetermined criterion, the probability calculation for that HMM is omitted.
【0047】類似性尺度としては、例えば、前記ヒスト
グラムを多次元ベクトルとみなし、内積をとるといった
方法が考えられる。こうして算出した類似性尺度が予め
設定した基準値(例えば0.1)より小さい場合にはH
MMの確率計算を省略する。As a similarity measure, for example, a method of considering the histogram as a multidimensional vector and taking an inner product can be considered. If the similarity measure calculated in this way is smaller than a preset reference value (for example, 0.1), H
The calculation of the probability of MM is omitted.
【0048】以上により、確率計算を大幅に省略でき、
より高速な音声認識ができる。なお、本手法を導入した
場合の音声認識の流れを図13のフローチャートに示
す。本フローチャートは、図3に示したフローチャート
にベクトル量子化コードの時系列の統計情報を使った計
算省略のステップS131、S132を挿入したものと
なっている。From the above, the probability calculation can be largely omitted,
Higher speed voice recognition is possible. The flow of speech recognition when this method is introduced is shown in the flowchart of FIG. This flowchart is obtained by inserting steps S131 and S132, which omit the calculation using the time-series statistical information of the vector quantization code, into the flowchart shown in FIG.
【0049】[0049]
【発明の効果】本発明によれば、HMMを用いた確率計
算を実行する上で、効率良くHMMのデータにアクセス
できる様にデータをメモリ内に配置し、かつ、連続した
アドレスを効率良くアクセスできるメモリを用いて構成
しているので、高速なHMM確率計算が実行できる。ま
た、HMMの構造上事前にできる計算は全て事前に済ま
せるようにしておくことにより、認識時の計算量を削減
できる。さらに各HMM毎のベクトル量子化コードの出
現頻度情報を使うことにより、可能性の低いHMMにつ
いては確率計算を省略できる。As described above, according to the present invention, in executing the probability calculation using the HMM, the data is arranged in the memory so that the data of the HMM can be accessed efficiently, and the continuous addresses are efficiently accessed. Since it is configured by using a memory that can be used, high-speed HMM probability calculation can be executed. In addition, the amount of calculation at the time of recognition can be reduced by performing all the calculations that can be performed in advance due to the structure of the HMM. Further, by using the appearance frequency information of the vector quantization code for each HMM, the probability calculation can be omitted for the HMMs with low possibility.
【0050】以上、各種の高速化手法を総合した本発明
によれば、最新の安価なハードウェアを用いて、コンパ
クトで安価でかつ高速な大語彙の音声認識装置を実現す
ることができる。As described above, according to the present invention in which various speed-up techniques are integrated, it is possible to realize a compact, inexpensive and high-speed large-vocabulary speech recognition device by using the latest inexpensive hardware.
【図1】本発明の音声認識装置の一実施例のハードウェ
ア構成を示すブロック図FIG. 1 is a block diagram showing a hardware configuration of an embodiment of a voice recognition device of the present invention.
【図2】本発明の音声認識装置の別の実施例のハードウ
ェア構成を示すブロック図FIG. 2 is a block diagram showing a hardware configuration of another embodiment of the voice recognition device of the present invention.
【図3】本発明の音声認識装置の一実施例の処理の概要
フローを示すフローチャートFIG. 3 is a flowchart showing an outline flow of processing of one embodiment of the voice recognition device of the present invention.
【図4】本発明の音声認識装置で用いる離散型隠れマル
コフモデルを説明する説明図FIG. 4 is an explanatory diagram illustrating a discrete hidden Markov model used in the speech recognition apparatus of the present invention.
【図5】本発明の音声認識装置で用いる離散型隠れマル
コフモデルによる確率計算処理の詳細な手順を示すフロ
ーチャートFIG. 5 is a flowchart showing a detailed procedure of probability calculation processing by a discrete hidden Markov model used in the speech recognition apparatus of the present invention.
【図6】本発明の音声認識装置で用いる離散型隠れマル
コフモデルのデータのメモリ内での並び方を説明する説
明図FIG. 6 is an explanatory view for explaining how the data of the discrete hidden Markov model used in the speech recognition apparatus of the present invention is arranged in the memory.
【図7】本発明の音声認識装置で用いる離散型隠れマル
コフモデルの状態遷移確率と出現確率の事前計算を説明
する説明図FIG. 7 is an explanatory diagram illustrating pre-calculation of state transition probabilities and appearance probabilities of discrete Hidden Markov Models used in the speech recognition apparatus of the present invention.
【図8】本発明の音声認識装置の複数の演算プロセッサ
による実施例を説明するブロック図FIG. 8 is a block diagram illustrating an embodiment of a plurality of arithmetic processors of a voice recognition device of the present invention.
【図9】本発明の音声認識装置の複数の演算プロセッサ
による実施例における離散型隠れマルコフモデルのデー
タのメモリ内での並び方を説明する説明図FIG. 9 is an explanatory view for explaining how the data of the discrete Hidden Markov Model is arranged in the memory in the embodiment by the plural arithmetic processors of the speech recognition apparatus of the present invention.
【図10】ベクトル量子化コードの出現確率が低い離散
型隠れマルコフモデルをリストアップしたテーブルの説
明図FIG. 10 is an explanatory diagram of a table listing discrete Hidden Markov Models with low occurrence probability of vector quantization code.
【図11】ベクトル量子化コードの出現確率に基づいて
離散型隠れマルコフモデルによる確率計算の一部を省略
する手順を説明するフローチャートFIG. 11 is a flowchart illustrating a procedure for omitting a part of probability calculation by a discrete hidden Markov model based on the appearance probability of a vector quantization code.
【図12】ベクトル量子化コード時系列の統計情報に基
づいて離散型隠れマルコフモデルによる確率計算の一部
を省略する手法を説明する説明図FIG. 12 is an explanatory diagram illustrating a method of omitting a part of the probability calculation by the discrete hidden Markov model based on the statistical information of the vector quantization code time series.
【図13】ベクトル量子化コード時系列の統計情報に基
づいて離散型隠れマルコフモデルによる確率計算の一部
を省略する手順を説明するフローチャートFIG. 13 is a flowchart illustrating a procedure of omitting a part of probability calculation by a discrete hidden Markov model based on statistical information of vector quantization code time series.
1・・・マイク、2・・・オーディオアンプ、3・・・
AD変換器 4・・・演算プロセッサ、5・・・外部メモリ、61・
・・演算プロセッサ 62・・・内部メモリ、63・・・データ転送制御部 71・・・第1の演算プロセッサ、72・・・第1の外
部メモリ 73・・・第2の演算プロセッサ、74・・・第2の外
部メモリ 75・・・第3の演算プロセッサ、76・・・第3の外
部メモリ1 ... Microphone, 2 ... Audio amplifier, 3 ...
AD converter 4 ... Arithmetic processor, 5 ... External memory, 61 ...
..Arithmetic processor 62 ... internal memory, 63 ... data transfer control unit 71 ... first arithmetic processor, 72 ... first external memory 73 ... second arithmetic processor, 74 ... ..Second external memory 75 ... Third arithmetic processor, 76 ... Third external memory
───────────────────────────────────────────────────── フロントページの続き (72)発明者 池田 宏 東京都国分寺市東恋ケ窪一丁目280番地 株式会社日立製作所中央研究所内 (72)発明者 在塚 俊之 東京都国分寺市東恋ケ窪一丁目280番地 株式会社日立製作所中央研究所内 (72)発明者 児玉 和行 東京都国分寺市東恋ケ窪一丁目280番地 株式会社日立製作所中央研究所内 (72)発明者 野口 孝樹 東京都国分寺市東恋ケ窪一丁目280番地 株式会社日立製作所中央研究所内 ─────────────────────────────────────────────────── ─── Continuation of the front page (72) Inventor Hiroshi Ikeda 1-280 Higashi Koikeku, Kokubunji, Tokyo Inside Central Research Laboratory, Hitachi, Ltd. (72) Toshiyuki Arizuka 1-280 Higashi Koikeku, Kokubunji, Tokyo Hitachi Ltd. (72) Inventor, Kazuyuki Kodama, Kazuyuki Kodama, 1-280, Higashi Koikeku, Kokubunji, Tokyo Hitachi, Ltd., Central Research Institute (72) Takaki Noguchi, 1-280, Higashi Koikeku, Kokubunji, Tokyo Hitachi, Ltd. Central Research Co., Ltd. In-house
Claims (11)
離散型隠れマルコフモデルを、認識対象の単語(あるい
は音節)毎に複数個用意し、入力音声に基づいて得られ
たベクトル量子化コード時系列に対して前記複数個の離
散型隠れマルコフモデルを用いて確率計算を行い、該計
算された確率値に基づいて認識結果を求める音声認識装
置であって、 前記確率計算を行う演算プロセッサと、 前記複数個の離散型隠れマルコフモデルのパラメータ
を、同一の前記ベクトル量子化コードに関するパラメー
タ毎に一連のアドレスにまとめて格納するメモリとを備
え、 前記演算プロセッサは、前記ベクトル量子化コード時系
列に対して、各ベクトル量子化コードに対応するパラメ
ータを前記メモリからアドレス順に読みだして前記確率
演算を行うことを特徴とする音声認識装置。1. A vector-quantized code time series obtained based on an input speech by preparing a plurality of discrete Hidden Markov Models for each recognition target word (or syllable) having parameters corresponding to a plurality of states. A plurality of discrete Hidden Markov models are used for probability calculation, a speech recognition apparatus for obtaining a recognition result based on the calculated probability value, an arithmetic processor for performing the probability calculation, A plurality of discrete Hidden Markov Model parameters, a memory for collectively storing in a series of addresses for each parameter related to the same vector quantization code, the arithmetic processor, for the vector quantization code time series And read the parameters corresponding to each vector quantization code from the memory in the order of addresses and perform the probability operation. Speech recognition apparatus according to claim.
メモリであることを特徴とする請求項1記載の音声認識
装置。2. The voice recognition device according to claim 1, wherein the memory is an external memory of the arithmetic processor.
メモリであることを特徴とする請求項1記載の音声認識
装置。3. The voice recognition device according to claim 1, wherein the memory is an internal memory of the arithmetic processor.
パラメータを格納した外部メモリと、前記演算プロセッ
サによる確率計算の対象となるパラメータを格納する内
部メモリと、前記外部メモリの分散したアドレスに存在
する、特定のベクトル量子化コードに対応するパラメー
タを取りだして前記内部メモリの一連のアドレスに転送
するデータ転送制御部とを備えたことを特徴とする請求
項3記載の音声認識装置。4. An external memory that stores parameters of the plurality of discrete Hidden Markov Models, an internal memory that stores parameters that are targets of probability calculation by the arithmetic processor, and exists at distributed addresses of the external memory. 4. The voice recognition apparatus according to claim 3, further comprising a data transfer control unit that takes out a parameter corresponding to a specific vector quantization code and transfers the parameter to a series of addresses in the internal memory.
における状態遷移に関する状態遷移確率と、当該状態遷
移における各ベクトル量子化コードの出現確率とを前記
パラメータとして有することを特徴とする請求項1記載
の音声認識装置。5. The discrete Hidden Markov Model has a state transition probability regarding a state transition in each state and an appearance probability of each vector quantization code in the state transition as the parameter. The voice recognition device described.
節)毎の離散型隠れマルコフモデルの各状態遷移に関す
る状態遷移確率と、当該状態遷移における各ベクトル量
子化コードの出現確率とを事前に掛け合わせた値を、両
パラメータに代わる新たなパラメータとして格納し、こ
のパラメータを用いて前記確率計算を行うようにしたこ
とを特徴とする請求項5記載の音声認識装置。6. The state transition probability for each state transition of the discrete Hidden Markov Model for each word (or each syllable) and the appearance probability of each vector quantization code in the state transition are stored in the memory in advance. The speech recognition apparatus according to claim 5, wherein the multiplied value is stored as a new parameter in place of both parameters, and the probability calculation is performed using this parameter.
め決められた値より小さい離散型隠れマルコフモデルを
ベクトル量子化コード毎に整理して格納したテーブルを
作成しておき、前記演算プロセッサは、ベクトル量子化
コードが得られる毎に前記テーブルを参照し、該テーブ
ル中の対応するベクトル量子化コードの欄に存在する離
散型隠れマルコフモデルに対しては確率計算を省略する
ようにしたことを特徴とする請求項5または6記載の音
声認識装置。7. A table is prepared in which discrete Hidden Markov Models whose appearance probability of the vector quantization code is smaller than a predetermined value are arranged and stored for each vector quantization code, and the arithmetic processor is configured to: Each time a vector quantization code is obtained, the table is referred to, and the probability calculation is omitted for the discrete hidden Markov model existing in the corresponding vector quantization code column in the table. The voice recognition device according to claim 5 or 6.
ル量子化コード時系列のある時間長分の統計情報を計算
し、該統計情報が各離散型隠れマルコフモデル毎に予め
決められた基準を満たした場合にはそのモデルに関する
確率計算を省略するようにしたことを特徴とする請求項
5または6記載の音声認識装置。8. The arithmetic processor calculates statistical information for a certain time length of the obtained vector quantization code time series, and the statistical information satisfies a predetermined criterion for each discrete hidden Markov model. The speech recognition apparatus according to claim 5 or 6, characterized in that the probability calculation for the model is omitted in the case.
された音声信号を増幅するオーディオアンプと、増幅さ
れた音声信号を一定時間間隔毎に取り込みディジタル化
するAD変換器と、取り込まれてディジタル化された音
声信号に対して演算を施す演算プロセッサと、演算プロ
セッサが利用する各種データを格納するメモリとから構
成される音声認識装置であって、 演算プロセッサは前記AD変換器によって取り込まれた
音声信号に対して短時間周波数分析を施して周波数スペ
クトルの時系列を求め、さらに該周波数スペクトルの時
系列に対して演算プロセッサ内部の内蔵メモリに予め格
納されたコードブックを用いてベクトル量子化を施して
ベクトル量子化コード時系列を求め、さらに該ベクトル
量子化コード時系列に対してメモリに予め格納された単
語(あるいは音節)毎の離散型隠れマルコフモデルを用
いて確率計算を行ない、計算された確率値に基づいて認
識結果を求めるようにし、 前記メモリには、各単語(あるいは各音節)の離散型隠
れマルコフモデルのパラメータを、特定のベクトル量子
化コードに関するパラメータ毎に一連のアドレスにまと
めて格納するようにしたことを特徴とする音声認識装
置。9. A microphone for inputting voice, an audio amplifier for amplifying a voice signal input from the microphone, an AD converter for fetching and digitizing the amplified voice signal at regular time intervals, and a digital signal for fetching. What is claimed is: 1. A voice recognition device, comprising: an arithmetic processor for performing an arithmetic operation on a converted voice signal; and a memory for storing various data used by the arithmetic processor, wherein the arithmetic processor is the voice captured by the AD converter. The signal is subjected to short-time frequency analysis to obtain the time series of the frequency spectrum, and the time series of the frequency spectrum is subjected to vector quantization using a codebook stored in advance in the internal memory of the arithmetic processor. To obtain the vector quantization code time series, and further to the memory for the vector quantization code time series. For each word (or syllable) stored in the memory, probability calculation is performed using a discrete Hidden Markov Model, and the recognition result is obtained based on the calculated probability value. The speech recognition device characterized in that the parameters of the discrete Hidden Markov Model of (1) are stored together in a series of addresses for each parameter related to a specific vector quantization code.
力された音声信号を増幅するオーディオアンプと、増幅
された音声信号を一定時間間隔毎に取り込みディジタル
化するAD変換器と、取り込まれてディジタル化された
音声信号に対して演算を施す複数の演算プロセッサと、
各演算プロセッサが利用する各種データを各演算プロセ
ッサ毎に個別に格納する複数のメモリとから構成される
音声認識装置であって、 第1の演算プロセッサは前記AD変換器によって取り込
まれた音声信号に対して短時間周波数分析を施して周波
数スペクトルの時系列を求め、さらに該周波数スペクト
ルの時系列に対して演算プロセッサ内部の内蔵メモリに
予め格納されたコードブックを用いてベクトル量子化を
施してベクトル量子化コード時系列を求め、さらに該ベ
クトル量子化コード時系列に対して前記メモリに予め格
納された単語(あるいは音節)毎の離散型隠れマルコフ
モデルのパラメータを用いて確率計算を行ない、さら
に、前記求められたベクトル量子化コードを他の全ての
演算プロセッサに送出し、 該他の全ての演算プロセッサは、前記第1の演算プロセ
ッサからベクトル量子化コードを受けとり、各演算プロ
セッサ毎に当該メモリに予め格納された単語(あるいは
音節)毎の離散型隠れマルコフモデルのパラメータを用
いて確率計算を行ない、最終的に得られた確率値は前記
第1の演算プロセッサに送出し、 該第1の演算プロセッサでは自身で計算した確率値およ
び他の演算プロセッサから受けとった確率値の全てを総
合して認識結果を求めるようにしたことを特徴とする音
声認識装置。10. A microphone for inputting a voice, an audio amplifier for amplifying a voice signal input from the microphone, an AD converter for fetching and digitizing the amplified voice signal at regular time intervals, and a digital signal for fetching. A plurality of arithmetic processors for performing arithmetic on the converted audio signal,
A voice recognition device comprising a plurality of memories for individually storing various data used by each arithmetic processor, wherein the first arithmetic processor converts a voice signal taken in by the AD converter. On the other hand, a short-time frequency analysis is performed to obtain the time series of the frequency spectrum, and the time series of the frequency spectrum is vector-quantized using a codebook stored in advance in the internal memory of the arithmetic processor to obtain the vector. Obtaining the quantization code time series, further performing probability calculation using the parameters of the discrete hidden Markov model for each word (or syllable) stored in advance in the memory for the vector quantization code time series, The obtained vector quantization code is sent to all other arithmetic processors, and all other arithmetic processors The processor receives the vector quantization code from the first arithmetic processor, and performs probability calculation using the parameters of the discrete hidden Markov model for each word (or syllable) stored in advance in the memory for each arithmetic processor. The probability value finally obtained is sent to the first arithmetic processor, and the first arithmetic processor combines all the probability values calculated by itself and the probability values received from other arithmetic processors. A voice recognition device characterized in that a recognition result is obtained.
ッサが担当する単語(あるいは音節)の離散型隠れマル
コフモデルのパラメータを格納し、その際、各単語(あ
るいは音節)についてベクトル量子化コードを同一とす
るパラメータを一連のアドレスにまとめて格納するよう
にしたことを特徴とする請求項10記載の音声認識装
置。11. Each of the memories stores parameters of a discrete hidden Markov model of a word (or syllable) which the corresponding arithmetic processor is in charge of, and at this time, a vector quantization code for each word (or syllable) is stored. 11. The voice recognition device according to claim 10, wherein the same parameters are collectively stored in a series of addresses.
Priority Applications (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP5209719A JPH0764590A (en) | 1993-08-24 | 1993-08-24 | Voice recognizer |
Applications Claiming Priority (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP5209719A JPH0764590A (en) | 1993-08-24 | 1993-08-24 | Voice recognizer |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| JPH0764590A true JPH0764590A (en) | 1995-03-10 |
Family
ID=16577518
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| JP5209719A Pending JPH0764590A (en) | 1993-08-24 | 1993-08-24 | Voice recognizer |
Country Status (1)
| Country | Link |
|---|---|
| JP (1) | JPH0764590A (en) |
Cited By (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| KR100464428B1 (en) * | 2002-08-12 | 2005-01-03 | 삼성전자주식회사 | Apparatus for recognizing a voice |
-
1993
- 1993-08-24 JP JP5209719A patent/JPH0764590A/en active Pending
Cited By (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| KR100464428B1 (en) * | 2002-08-12 | 2005-01-03 | 삼성전자주식회사 | Apparatus for recognizing a voice |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| EP1012827B1 (en) | Speech recognition system for recognizing continuous and isolated speech | |
| US4837831A (en) | Method for creating and using multiple-word sound models in speech recognition | |
| US5937384A (en) | Method and system for speech recognition using continuous density hidden Markov models | |
| US5822729A (en) | Feature-based speech recognizer having probabilistic linguistic processor providing word matching based on the entire space of feature vectors | |
| US6542866B1 (en) | Speech recognition method and apparatus utilizing multiple feature streams | |
| US6195634B1 (en) | Selection of decoys for non-vocabulary utterances rejection | |
| US6141641A (en) | Dynamically configurable acoustic model for speech recognition system | |
| US7054810B2 (en) | Feature vector-based apparatus and method for robust pattern recognition | |
| US5572624A (en) | Speech recognition system accommodating different sources | |
| EP0847041A2 (en) | Method and apparatus for speech recognition performing noise adaptation | |
| EP0755046B1 (en) | Speech recogniser using a hierarchically structured dictionary | |
| JP4224250B2 (en) | Speech recognition apparatus, speech recognition method, and speech recognition program | |
| EP0706171A1 (en) | Speech recognition method and apparatus | |
| EP0771461A1 (en) | Method and apparatus for speech recognition using optimised partial probability mixture tying | |
| US6502072B2 (en) | Two-tier noise rejection in speech recognition | |
| US8639510B1 (en) | Acoustic scoring unit implemented on a single FPGA or ASIC | |
| JPH09319392A (en) | Voice recognition device | |
| US20020133343A1 (en) | Method for speech recognition, apparatus for the same, and voice controller | |
| EP1369847B1 (en) | Speech recognition method and system | |
| JP4716125B2 (en) | Pronunciation rating device and program | |
| JP2000075885A (en) | Voice recognition device | |
| JPH09305195A (en) | Voice recognition device and voice recognition method | |
| JP4678464B2 (en) | Voice recognition apparatus, voice recognition method, program, and recording medium | |
| JPH09106295A (en) | Speech recognition method and apparatus thereof | |
| KR100486307B1 (en) | Apparatus for calculating an Observation Probability of Hidden Markov model algorithm |