JPH0254560B2 - - Google Patents
Info
- Publication number
- JPH0254560B2 JPH0254560B2 JP57107767A JP10776782A JPH0254560B2 JP H0254560 B2 JPH0254560 B2 JP H0254560B2 JP 57107767 A JP57107767 A JP 57107767A JP 10776782 A JP10776782 A JP 10776782A JP H0254560 B2 JPH0254560 B2 JP H0254560B2
- Authority
- JP
- Japan
- Prior art keywords
- word
- power
- speech recognition
- time
- verification
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Expired
Links
Description
【発明の詳細な説明】
(1) 発明の技術分野
本発明は多数単語音声認識方式を用いた音声の
実時間認識処理において、候補単語を選択する前
照合処理部を具え、高速かつ高い認識精度を有す
る音声認識装置に関するものである。[Detailed Description of the Invention] (1) Technical Field of the Invention The present invention provides high-speed and high recognition accuracy in real-time speech recognition processing using a multi-word speech recognition method, which includes a pre-verification processing unit that selects candidate words. The present invention relates to a speech recognition device having the following.
(2) 従来技術と問題点
従来、多数単語音声認識装置における前照合処
理方式としては、スペクトルの大域的特徴を抽出
するもの、正規化パワーの時間正規化照合するも
の、または語頭、語尾の詳細パターン照合による
もの等がある。これらには単語の基本的差異であ
る単語発音長、または母音、子音のパワー情報が
積極的に適用されていない。このため、前照合に
おけるパラメータは本照合におけるパラメータに
対し本質的に変らないことになり、分類機能を上
げるためにはかなり細かな情報を用いなければな
らず照合に時間がかかるという欠点がある。(2) Prior art and problems Conventionally, pre-matching processing methods in multi-word speech recognition devices include those that extract global features of the spectrum, those that perform time normalized matching of normalized power, or methods that extract details of the beginning and end of words. There are methods based on pattern matching. Word pronunciation length, which is a basic difference between words, or power information of vowels and consonants is not actively applied to these words. Therefore, the parameters in the pre-verification are essentially the same as those in the main verification, and there is a drawback that quite detailed information must be used in order to improve the classification function, and the verification takes time.
(3) 発明の目的
本発明の目的は多数単語音声認識方式において
単語の基本的差異を示す単語発声長または母音、
子音のパワー量情報を適用することにより、高速
かつ高い認識精度を有する音声認識装置を提供す
ることである。(3) Purpose of the Invention The purpose of the present invention is to identify word utterance lengths or vowels that indicate basic differences between words in a multi-word speech recognition system.
It is an object of the present invention to provide a speech recognition device having high speed and high recognition accuracy by applying power amount information of consonants.
(4) 発明の構成
前記目的を達成するため、本発明の音声認識装
置は多数単語音声認識方式を用いた音声の実時間
認識処理を行ない候補単語を選択する前照合処理
部を有する音声認識装置において、前記前照合処
理部に、単語発声長、正規化パワーから求めた単
語発声の全パワー量、および母音性、子音性を示
す比較的パワーの大きい時間長と比較的パワーの
小さい時間長との比の3つの特徴パラメータを用
い候補単語を選択する手段を設けたことを特徴と
するものである。(4) Structure of the Invention In order to achieve the above object, the speech recognition device of the present invention is a speech recognition device having a pre-verification processing unit that performs real-time recognition processing of speech using a multi-word speech recognition method and selects candidate words. In the pre-verification processing unit, the word utterance length, the total power amount of the word utterance obtained from the normalized power, and the time length with relatively high power and the time length with relatively low power indicating vowel nature and consonant nature are provided. The present invention is characterized by providing means for selecting candidate words using three feature parameters having a ratio of .
(5) 発明の実施例
第1図、第2図a,bは本発明の要部の説明図
である。(5) Embodiments of the invention FIG. 1 and FIGS. 2a and 2b are explanatory diagrams of the main parts of the present invention.
本発明の多数単語音声認識装置における候補単
語を選択する方式として、前照合処理のパラメー
タは、本照合処理の識別用パラメータとは別に、
3つの量、すなわち単語発声長と、正規化パワー
から求めた単語発声の全パワー量と、母音性、子
音性を示す量として正規化パワーの平均値に標準
偏差値を加えたものと、これから差引いたものの
比の3つの特徴パラメータを用いることにより候
補単語が選択される。 As a method for selecting candidate words in the multi-word speech recognition device of the present invention, the parameters of the pre-matching process are set separately from the identification parameters of the main matching process.
Three quantities: the word utterance length, the total power of the word utterance calculated from the normalized power, the average value of the normalized power plus the standard deviation value as a quantity indicating vowelness and consonance, and from this Candidate words are selected by using three feature parameters: subtract and ratio.
第1図はこの場合に使用される時間正規化に関
する説明図である。 FIG. 1 is an explanatory diagram regarding time normalization used in this case.
単語音声の発声時間長は異なる単語は勿論のこ
と、同図の波形11〜1mに示すように、同一の
単語でも発声ごとに異なつている。 As shown in waveforms 1 1 to 1m in the figure, the utterance time length of word sounds differs not only for different words but also for each utterance of the same word.
そこで、同図の波形2に示すように基準時間長
に正規化し、辞書との照合にはこの時間正規化照
合波形が用いられる。この時、照合対象の辞書と
しては極端に長さの異なるもの、すなわち長さが
2倍以上または1/2以下は除外される。従来方式
ではこの単語発声の固有量である時間長正規化が
積極的に適用されていなかつたのに対し、本発明
ではこれを設定したものである。 Therefore, as shown in waveform 2 in the figure, it is normalized to a reference time length, and this time-normalized verification waveform is used for verification with the dictionary. At this time, dictionaries to be checked that have extremely different lengths, that is, dictionaries whose length is twice or more or half or less, are excluded. In the conventional system, this time length normalization, which is an inherent quantity of word utterance, was not actively applied, but this is set in the present invention.
第2図a,bは本発明の前照合処理部で用いら
れる単語発声の全パワー量の説明図である。 FIGS. 2a and 2b are explanatory diagrams of the total power amount of word utterance used in the pre-verification processing section of the present invention.
同図a,bは横軸に時間長、縦軸にパワーをと
つた場合の単語発声の時間方向のパワー変化31,
32を例示し、第1図に示した時間正規化された
同一単語に対応している。 Figures a and b show power changes in the time direction of word utterances when the horizontal axis is the time length and the vertical axis is the power.
3 2 is shown as an example, and corresponds to the same time-normalized word shown in FIG.
通常、同一単語でもその発声の仕方により時間
成分のみならずパワーの大きさも異なつてくる。
この単語波形31,32に対し、最大パワーと最小
パワーの間で線形に正規化して単語波形41,42
が得られる。なお、発声単語の時間長は発声ごと
に変動するが、大略の値としては単語固有の長さ
が存在する。従つて、同図のように、単語の比較
的単純な固有情報量として、パワーを時間長とと
もに正規化し単語波形41,42の斜線部分より単
語の全パワー量が得られる。 Normally, even if the word is the same, not only the time component but also the magnitude of the power will differ depending on how it is uttered.
These word waveforms 3 1 , 3 2 are linearly normalized between the maximum power and the minimum power to form word waveforms 4 1 , 4 2 .
is obtained. Note that although the time length of the uttered word varies depending on the utterance, there is a rough value that is unique to the word. Therefore, as shown in the figure, the total power of the word can be obtained from the shaded portions of the word waveforms 4 1 and 4 2 by normalizing the power along with the time length as a relatively simple amount of specific information of the word.
以上の方法により、3つの特徴パラメータのう
ち第1番目のパワーは正規化された音声発声長で
あり、第2番目のパラメータは発声の正規化パラ
メータの全時間長にわたる総和、すなわち単語の
全パワー量である。この両パラメータを組合せた
第2図の全パワー量が単語発声の固有情報量とし
て安定なパラメータが設定される。 With the above method, the first power of the three feature parameters is the normalized vocal utterance length, and the second parameter is the sum of the normalized utterance parameters over the entire time length, that is, the total power of the word. It's the amount. The total amount of power shown in FIG. 2, which is a combination of these two parameters, is set as a stable parameter as the amount of unique information of word utterance.
次の第3番目のパワーは、単語中の母音らし
さ、子音らしさを示す指標として、単語発声中の
母音量/子音量という値である。母音量としては
正規化パワーが、(その平均値)+(標準偏差)を
越えた時間長が使われ、また子音量としては正規
化パワーが(その平均値)−(標準偏差)以下の時
間長が使われる。すなわち、
(母音/子音比)=(平均値)+(標準偏差
)/(平均値)−(標準偏差)〔正規化パワー〕×100
が指標となる。これは単語分類に有効なパラメー
タとなる。 The next third power is a value of vowel volume/consonant volume during word pronunciation as an index indicating vowel-likeness and consonant-likeness in a word. The length of time during which the normalized power exceeds (its average value) + (standard deviation) is used as the vowel volume, and the length of time during which the normalized power exceeds (its average value) - (standard deviation) is used as the consonant volume. long is used. In other words, the index is (vowel/consonant ratio) = (average value) + (standard deviation) / (average value) - (standard deviation) [normalized power] x 100. This becomes an effective parameter for word classification.
第3図は本発明の実施例の構成説明図である。
認識に先立ち、後述の前照合辞書16と本照合辞
書17を用意しておく。前照合辞書16は3つの
特徴パラメータに関して、各値の大きさ順に単語
が類別されており、前照合では入力音声の1つの
パラメータ値を求め、その値の±30%以内に入る
辞書項目を選択する。これを3つのパラメータに
ついて行ない、3者の論理積をとり前照合結果と
して候補単語が選択される。 FIG. 3 is an explanatory diagram of the configuration of an embodiment of the present invention.
Prior to recognition, a pre-verification dictionary 16 and a main verification dictionary 17, which will be described later, are prepared. In the pre-verification dictionary 16, words are classified in the order of the magnitude of each value regarding three feature parameters, and in the pre-verification, one parameter value of the input voice is determined, and dictionary items that fall within ±30% of that value are selected. do. This is done for the three parameters, and the logical product of the three is taken to select a candidate word as the pre-verification result.
本照合辞書17は本照合で用いるスペクトルパ
ターンのような通常の特徴パラメータが格納され
る。 The main matching dictionary 17 stores normal feature parameters such as spectral patterns used in the main matching.
同図において、認識時にマイクロホーン10か
ら音声を入力し、その電気信号は増幅器11を通
して分析回路12に送られ、音声認識のために必
要な各種パラメータの分析を行なう。 In the figure, during recognition, voice is input from a microphone 10, and the electrical signal is sent to an analysis circuit 12 through an amplifier 11, where it analyzes various parameters necessary for voice recognition.
まず、候補単語を選択する前照合処理のため
に、音声パワーがパワー正規化回路14により前
述の第1図、第2図の手法で正規化され、単語の
発声時間長、全パワー量、母音/子音比の3特徴
パラメータに変換される。これらが前照合回路1
5に送られ、前照合辞書16と照合され前述のよ
うにして候補単語が選択される。これが本照合辞
書17に送られ、対応するパラメータのたとえば
スペクトルパターンが選定される。一方、本照合
処理では前照合処理で用いるパラメータと相補的
な量を抽出するため、分析回路12の出力を特徴
抽出回路13に送り、たとえばスペクトルパター
ンの特徴パラメータが抽出され、本照合回路18
において本照合辞書17からのパラメータとの距
離計算を行ない、その結果を判定回路19に送り
判定し認識結果を出力する。 First, for pre-matching processing to select candidate words, the power normalization circuit 14 normalizes the speech power using the method shown in FIGS. / consonant ratio. These are the pre-verification circuit 1
5, the word is checked against the pre-check dictionary 16, and candidate words are selected as described above. This is sent to the main reference dictionary 17, and a corresponding parameter such as a spectrum pattern is selected. On the other hand, in the main matching process, in order to extract complementary quantities to the parameters used in the pre-matching process, the output of the analysis circuit 12 is sent to the feature extraction circuit 13, and, for example, feature parameters of the spectral pattern are extracted, and the main matching circuit 18
In this step, the distance between the parameter and the parameter from the reference dictionary 17 is calculated, and the result is sent to the determination circuit 19, which makes a determination and outputs the recognition result.
(6) 発明の効果
以上説明したように、本発明によれば、前照合
処理部において、単語発声固有の情報量として単
語の正規化パワーを基にした前述の3つの特徴パ
ラメータにより、また本照合処理部では相補的な
パラメータにより照合し、併せて2段の照合処理
を行なうので、安定にかつ高速に候補単語の選択
ができ、高い認識率で単語音声認識が可能とな
る。(6) Effects of the Invention As explained above, according to the present invention, in the pre-verification processing section, the above-mentioned three characteristic parameters based on the normalized power of the word as the amount of information specific to the word utterance are used to Since the matching processing section performs matching using complementary parameters and also performs two-stage matching processing, candidate words can be selected stably and quickly, and word speech recognition can be performed with a high recognition rate.
第1図、第2図a,bは本発明の要部の説明
図、第3図は本発明の実施例の構成説明図であ
り、41,42は正規化された全パワー量、10は
マイクロホーン、11は増幅器、12は分析回
路、13は特徴抽出回路、14はパワー正規化回
路、15は前照合回路、16は前照合辞書、17
は本照合辞書、18は本照合回路、19は判定回
路を示す。
1 and 2 a and b are explanatory diagrams of the main parts of the present invention, and FIG. 3 is an explanatory diagram of the configuration of an embodiment of the present invention, 4 1 and 4 2 are normalized total power amounts, 10 is a microphone, 11 is an amplifier, 12 is an analysis circuit, 13 is a feature extraction circuit, 14 is a power normalization circuit, 15 is a pre-verification circuit, 16 is a pre-verification dictionary, 17
Reference numeral 18 indicates a book checking dictionary, 18 a book checking circuit, and 19 a determination circuit.
Claims (1)
認識処理を行ない候補単語を選択する前照合処理
部を有する音声認識装置において、前記前照合処
理部に、単語発声長、正規化パワーから求めた単
語発声の全パワー量、および母音性、子音性を示
す比較的パワーの大きい時間長と比較的パワーの
小さい時間長との比の3つの特徴パラメータを用
い候補単語を選択する手段を設けたことを特徴と
する音声認識装置。1. In a speech recognition device that has a pre-verification processing unit that performs real-time recognition processing of speech using a multi-word speech recognition method and selects candidate words, the pre-match processing unit is configured to perform real-time speech recognition processing using a multi-word speech recognition method to select candidate words. A means is provided for selecting candidate words using three characteristic parameters: the total power of word utterance, and the ratio of the time length with relatively high power and the time length with relatively low power indicating voweliness and consonance. A voice recognition device featuring:
Priority Applications (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP57107767A JPS58224396A (en) | 1982-06-23 | 1982-06-23 | Voice recognition equipment |
Applications Claiming Priority (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP57107767A JPS58224396A (en) | 1982-06-23 | 1982-06-23 | Voice recognition equipment |
Publications (2)
| Publication Number | Publication Date |
|---|---|
| JPS58224396A JPS58224396A (en) | 1983-12-26 |
| JPH0254560B2 true JPH0254560B2 (en) | 1990-11-21 |
Family
ID=14467481
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| JP57107767A Granted JPS58224396A (en) | 1982-06-23 | 1982-06-23 | Voice recognition equipment |
Country Status (1)
| Country | Link |
|---|---|
| JP (1) | JPS58224396A (en) |
Families Citing this family (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP4604424B2 (en) * | 2001-08-07 | 2011-01-05 | カシオ計算機株式会社 | Speech recognition apparatus and method, and program |
| WO2012150658A1 (en) * | 2011-05-02 | 2012-11-08 | 旭化成株式会社 | Voice recognition device and voice recognition method |
-
1982
- 1982-06-23 JP JP57107767A patent/JPS58224396A/en active Granted
Also Published As
| Publication number | Publication date |
|---|---|
| JPS58224396A (en) | 1983-12-26 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| Vergin et al. | Generalized mel frequency cepstral coefficients for large-vocabulary speaker-independent continuous-speech recognition | |
| Bezoui et al. | Feature extraction of some Quranic recitation using mel-frequency cepstral coeficients (MFCC) | |
| JPH02195400A (en) | voice recognition device | |
| JPS6336676B2 (en) | ||
| JP2001166789A (en) | Chinese speech recognition method and apparatus using initial / final phoneme similarity vector | |
| Elenius et al. | Effects of emphasizing transitional or stationary parts of the speech signal in a discrete utterance recognition system | |
| Zhao et al. | A new hybrid approach for automatic speech signal segmentation using silence signal detection, energy convex hull, and spectral variation | |
| Blomberg et al. | Auditory models in isolated word recognition | |
| JPS6138479B2 (en) | ||
| JPS63158596A (en) | Phoneme analogy calculator | |
| JPH0558553B2 (en) | ||
| Villing et al. | Performance limits for envelope based automatic syllable segmentation | |
| Ramdinmawii et al. | Discriminating between High-Arousal and Low-Arousal Emotional States of Mind using Acoustic Analysis. | |
| Wang et al. | Automatic language recognition with tonal and non-tonal language pre-classification | |
| Kopec | Voiceless stop consonant identification using LPC spectra | |
| Wang et al. | Automatic Tonal and Non-Tonal Language Classification and Language Identification Using Prosodic Information. | |
| Sharma et al. | Speech recognition of Punjabi numerals using synergic HMM and DTW approach | |
| JPS58224396A (en) | Voice recognition equipment | |
| Ozaydin | An isolated word speaker recognition system | |
| Mengistu et al. | Text independent amharic language dialect recognition using neuro-fuzzy gaussian membership function | |
| JPH0887292A (en) | Word voice recognition device | |
| JPH0619497A (en) | Speech recognition method | |
| Sharma et al. | Automatic Segmentation of Punjabi Speech into Syllable-Like Units using Group Delay: A Review | |
| JPS63161499A (en) | Voice recognition equipment | |
| JPH0469800B2 (en) |