JPH02105200A - Speech recognizing device - Google Patents
Speech recognizing deviceInfo
- Publication number
- JPH02105200A JPH02105200A JP63257328A JP25732888A JPH02105200A JP H02105200 A JPH02105200 A JP H02105200A JP 63257328 A JP63257328 A JP 63257328A JP 25732888 A JP25732888 A JP 25732888A JP H02105200 A JPH02105200 A JP H02105200A
- Authority
- JP
- Japan
- Prior art keywords
- identification
- label
- speech
- identification label
- similarity
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Abstract
Description
【発明の詳細な説明】
(産業上の利用分野〕
本発明は、入力音声をその音声の発声を示す識別ラベル
に変換する音声認識装置に関する。DETAILED DESCRIPTION OF THE INVENTION (Field of Industrial Application) The present invention relates to a speech recognition device that converts input speech into an identification label indicating the utterance of the speech.
(従°来の技術)
音声認識装置は、一般に、予め音声から抽出した特徴パ
ラメータにその音声の発声を示す識別ラベルを付加して
記憶しておき、入力音声と、特徴パラメータが最も類似
する識別ラベルを抽出し、音声認識結果として出力して
いる。従来のこの種の音声認識装置の中で、一定時間毎
に音声認識を行い、その音声認識結果を用いて音韻を認
識する装置が知られている。ところが、音声の中には、
ある1つの音韻から次の音韻に発声が8行する場合にそ
の境界部分(遷移領域と称する)では音声の特徴が変化
する音韻があり、従来のこの種の音声認識装置はこの遷
移領域での識別ラベルの誤りこ識が音韻の認識に対して
悪影響を与えるという問題点があった。(Conventional technology) Generally, a speech recognition device stores feature parameters extracted from speech in advance with an identification label indicating the utterance of the speech, and selects an identification label whose feature parameters are most similar to the input speech. Labels are extracted and output as speech recognition results. Among conventional speech recognition devices of this type, there is a known device that performs speech recognition at regular intervals and uses the speech recognition results to recognize phonemes. However, in the audio,
When there are 8 lines of utterance from one phoneme to the next, there is a phoneme in which the characteristics of the voice change at the boundary part (referred to as a transition area), and conventional speech recognition devices of this type are difficult to recognize in this transition area. There was a problem in that incorrect knowledge of identification labels had a negative effect on phonological recognition.
この点について詳しく説明する。This point will be explained in detail.
(発明が解決しようとする課題) 第7図は従来装置の一般的な音声認識処理手順を示す。(Problem to be solved by the invention) FIG. 7 shows a general speech recognition processing procedure of a conventional device.
第7図において、例えば°°アオ°′とマイクロホンに
発声された音声は、マイクロホンからの音声信号の周波
数分析により、非常に短い一定時間毎に、入力音声の周
波数帯毎の強さ(特徴パラメータと称する)に変換され
る。この特徴パラメータと上述したように予め記憶しで
ある標準の特徴パラメータ各々との距離が計算され、こ
の特徴パラメータと各標準の特徴パラメータとの類似度
を比較することによって、その時点において最もよく類
似している識別記号(ラベルと称する)が抽出される。In Fig. 7, for example, the voice uttered into the microphone, ``°°Ao°'', is determined by frequency analysis of the audio signal from the microphone. ). The distance between this feature parameter and each standard feature parameter stored in advance as described above is calculated, and by comparing the degree of similarity between this feature parameter and each standard feature parameter, the most similar one at that point is determined. The identification symbol (referred to as a label) is extracted.
上述の”アオ′°という音声は、記号で表わすと、 “
へへへUへULIA11UOOOOOO・・・”のラベ
ル列か得られる。ところが、音韻°ア°°および°゛オ
°°遷移部分では、特徴パラメータが変化するため、”
A”を°“IJ ”と誤認識する場合が生しる。The above-mentioned sound “ao′°” can be expressed symbolically as “
A label string of "Hehehe U to ULIA11UOOOOOO..." is obtained. However, in the phoneme °A °° and °゛o °° transition parts, the feature parameters change, so "
A” may be mistakenly recognized as “IJ”.
このような誤認識を考慮して、同一の2個のラベルでは
さまれるラベルをそのラベルに変換する補正処理が行な
われた後、一定時間内に識別された連続するラベル列の
個数が最も多いラベルを音韻結果として抽出する。この
結果、第7図に示す例では、遷移部分の誤認識結果の影
響を受け“アオ”という発声に対し“UO”という誤認
識となる。In consideration of such misrecognition, after a correction process is performed to convert the label sandwiched between two identical labels into that label, the number of consecutive label strings identified within a certain period of time is the highest. Extract the label as a phonological result. As a result, in the example shown in FIG. 7, the utterance "Ao" is incorrectly recognized as "UO" due to the influence of the erroneous recognition result of the transition part.
このように連続音声を認識する従来の音声認識装置では
、1つの音韻と次の音韻の遷移部分の特徴パターンが変
化することにより音韻の誤認識を行うことが多いという
問題点があった。Conventional speech recognition devices that recognize continuous speech as described above have a problem in that they often misrecognize phonemes due to changes in the characteristic pattern of the transition portion between one phoneme and the next phoneme.
そこで、本発明の目的は、このような問題点を解決し、
より確実に連続音声を認識することが可能な音声認識装
置を提供することにある。Therefore, the purpose of the present invention is to solve such problems,
An object of the present invention is to provide a speech recognition device capable of recognizing continuous speech more reliably.
(課題を解決するための手段)
このような目的を達成するために、本発明は、連続音声
を入力する入力手段と、入力手段から入力された連続音
声を一定時間毎に特徴パラメータ化し、特徴パラメータ
と、音声の識別ラベルが予め判明している複数の標準特
徴パラメータの各々との距離計算を行い、計算結果に基
いて連続音声に対する複数の識別ラベルの類似度をそれ
ぞれ決定する第1演算手段と、第1演算手段により定め
られた現時点から、過去まで連続する一定個数の類似度
を識別ラベル毎に合計する第2演算手段と、第2演算手
段により合計された識別ラベル毎の類似度の合計結果に
基いて、最も類似度の高い合計結果と対応する識別ラベ
ルを抽出し、抽出された識別ラベルを、現時点における
音声認識結果として出力する抽出手段とを具えたことを
特徴とする。(Means for Solving the Problem) In order to achieve such an object, the present invention includes an input means for inputting continuous speech, and a feature parameter for the continuous speech inputted from the input means at fixed time intervals. a first calculating means for calculating the distance between the parameter and each of the plurality of standard feature parameters for which the identification label of the speech is known in advance, and determining the degree of similarity of the plurality of identification labels for the continuous speech based on the calculation result; and a second calculation means for summing up a certain number of consecutive similarities from the current time determined by the first calculation means to the past for each identification label; The present invention is characterized by comprising an extraction means for extracting an identification label corresponding to the summation result having the highest degree of similarity based on the summation result, and outputting the extracted identification label as the current speech recognition result.
本発明では、各時点毎に得られる類似度から始まる過去
の所定個数の類似度を識別ラベル毎に第2演算手段によ
り合計し、その合計結果の中の最も類似度の高い識別ラ
ベルを抽出手段により順次に抽出する。この結果、各時
点で得られる類似度が時系列的に平滑化されるので、例
えば入力音声に雑音が部分的に混入したり、入力音声の
一部の発声が変化したり、入力音声の音韻が変化する場
合においても、その音声の変化部分がこれまでに認識さ
れたラベルと同じラベルと認識され、以て音声の認識確
率が高くなる。In the present invention, a predetermined number of past similarities starting from the similarities obtained at each point in time are summed for each identification label by a second calculation means, and an identification label having the highest similarity among the total results is extracted. Extract sequentially by As a result, the similarity obtained at each point in time is smoothed over time, so for example, noise may be partially mixed into the input speech, the utterance of a part of the input speech may change, or the phonology of the input speech may be smoothed over time. Even when the voice changes, the changed part of the voice is recognized as the same label as the label recognized so far, and the probability of voice recognition increases.
以下、図面を参照して本発明実施例を詳細に説明する。 Embodiments of the present invention will be described in detail below with reference to the drawings.
本願出願人は、連続音声の中の特に音声の遷移部分の特
徴パターンが、前時点までにサンプリングした音声の特
徴パターンから少しずつ変化して行くという連続音声の
性質に着目し、各時点毎に計算した各標準特徴パラメー
タの類似度を時系列に一定時間(窓と称する)の範囲で
集計し、その集計結果に基き、最も類似度の高い識別ラ
ベルを各時点の認識結果として定めるようにしたもので
ある。The applicant of this application focused on the property of continuous speech in which the characteristic pattern of the transition part of continuous speech changes little by little from the characteristic pattern of the speech sampled up to the previous point in time. The calculated similarity of each standard feature parameter is aggregated over a certain period of time (referred to as a window) in chronological order, and based on the aggregated results, the identification label with the highest degree of similarity is determined as the recognition result at each point in time. It is something.
第1図は本発明実施例の基本構成を示す。FIG. 1 shows the basic configuration of an embodiment of the present invention.
第1図において、100は連続音声を入力する入力手段
である。In FIG. 1, 100 is an input means for inputting continuous speech.
200は該入力手段から入力された前記連続音声を一定
時間毎に特徴パラメータ化し、該特徴パラメータと音声
の識別ラベルが予め判明している複数の標準特徴パラメ
ータの各々との距離計算を行い、該計算結果に基いて前
記連続音声に対する複数の前記識別ラベルの類似度をそ
れぞれ訣定する第1演算手段である。200 converts the continuous speech inputted from the input means into feature parameters at regular intervals, calculates the distance between the feature parameters and each of a plurality of standard feature parameters for which the identification label of the speech is known in advance, and It is a first calculation means that determines the degree of similarity of the plurality of identification labels with respect to the continuous speech based on the calculation results.
300は、該第1演算手段により定められた現時点から
過去まで連続する一定個数の前記類似度を前記識別ラベ
ル毎に合計する第2演算手段である。Reference numeral 300 denotes a second calculation means that totals a certain number of consecutive similarities from the present time to the past determined by the first calculation means for each of the identification labels.
400は、該第2演算手段により合計された前記識別ラ
ベル毎の類似度の合計結果に基いて、最も類似度の高い
合計結果と対応する識別ラベルを抽出し、抽出された当
該識別ラベルを、前記現時点における音声認識結果とし
て出力する抽出手段である。400 extracts an identification label corresponding to the total result with the highest degree of similarity based on the total result of the similarity for each identification label summed by the second calculation means, and uses the extracted identification label as This is an extraction means that outputs the voice recognition result at the current time.
第2図は本発明実施例の具体的な構成を示す。FIG. 2 shows a specific configuration of an embodiment of the present invention.
第2図において、10は音声を入力する入力手段として
のマイクロフォンである。11はアンプであり、マイク
ロフォンの出力を増幅する。12はA/D変換器であり
、アンプ11の増幅出力をA/D変換する。13は第1
演算手段の一部を構成するフーリエ変換器であり、A/
D変換器12の出力をフーリエ変換し、周波数帯域毎の
音声の強さ(パワースペクトラム)を音声の特徴パラメ
ータとして出力する。フーリエ変換器13はLSI (
大規模集積回路)になっているものが知られているが、
コンピュータによりフーリエ変換を実行してもよい。フ
ーリエ変換器に代わりバンドパスフィルタを用いること
も可能である。In FIG. 2, 10 is a microphone serving as an input means for inputting voice. An amplifier 11 amplifies the output of the microphone. 12 is an A/D converter, which converts the amplified output of the amplifier 11 into A/D. 13 is the first
It is a Fourier transformer that constitutes a part of the calculation means, and the A/
The output of the D converter 12 is Fourier-transformed, and the sound strength (power spectrum) for each frequency band is output as a sound characteristic parameter. The Fourier transformer 13 is an LSI (
Large-scale integrated circuits) are known.
A Fourier transform may be performed by a computer. It is also possible to use a bandpass filter instead of a Fourier transformer.
14は第1.第2演算手段および抽出手段に相当するコ
ンピュータシステムであり、コンピュータシステム14
はキーボード15、フロッピディスク等を用いた外部記
憶装置(FDD)16 、プリンタ17および陰極管表
示装置(CRT) 18に接続している。15は情報を
入力可能なキーボードであり、キーボード15からは、
音声認識モードの指示や標準特徴パラメータ作・成のた
めの情報の入力認識結果の修正等を行う。FDD16に
は、作成された標準特徴パラメータと、この標準特徴パ
ラメータと対応する識別ラベルのコード番号がテーブル
(音韻マツプと称する)として記憶されている。また、
FDD16には後述の本発明に関わる確率テーブル16
−1が設けられている。14 is the first. It is a computer system corresponding to the second calculation means and the extraction means, and the computer system 14
is connected to a keyboard 15, an external storage device (FDD) 16 using a floppy disk or the like, a printer 17, and a cathode ray tube display (CRT) 18. 15 is a keyboard that can input information; from the keyboard 15,
Input information for voice recognition mode instructions and standard feature parameter creation, modify recognition results, etc. The FDD 16 stores the created standard feature parameters and the code numbers of identification labels corresponding to the standard feature parameters as a table (referred to as a phoneme map). Also,
The FDD 16 includes a probability table 16 related to the present invention, which will be described later.
-1 is provided.
第3図は本発明実施例のコンピュータシステム14の構
成の一例を示す。FIG. 3 shows an example of the configuration of the computer system 14 according to the embodiment of the present invention.
本実施例は高速演算処理を行うために、演算処理を複数
個の中央演算処理装置で行うようにしている。第3図に
おいて、 14−1はマイクロプロセッサ(MPU)で
あり、MPU14−1は入力音声のラベル付け(音声認
識)を行う、 14−2はメモリであり、メモリ14−
2にはMPU14−1が実行する制御手順が格納されて
いる。14−3は高速デジタルシグナルプロセッサ(o
sp)でありDSP14−3は入力音声の特徴パラメー
タと、音韻マツプ中の標準特徴パラメータとの類似度の
計算(距離計算)を行う、 14−4はメモリであり、
DSP14−3が実行する制御手順を格納する。14−
5はパーソナルコンピュータであり、認識モードの設定
や認識結果の表示を含む全システムの動作を続開する役
割を果たす。In this embodiment, in order to perform high-speed arithmetic processing, arithmetic processing is performed by a plurality of central processing units. In FIG. 3, 14-1 is a microprocessor (MPU), the MPU 14-1 labels input speech (speech recognition), and 14-2 is a memory.
2 stores control procedures executed by the MPU 14-1. 14-3 is a high-speed digital signal processor (o
14-4 is a memory;
Stores control procedures executed by the DSP 14-3. 14-
A personal computer 5 plays the role of continuing the operation of the entire system, including setting the recognition mode and displaying the recognition results.
なお、コンピュータシステム14には大型コンピュータ
を用いてもよく、装置の大きさ、演算処理程度に応じて
構成すればよい。Note that a large-sized computer may be used as the computer system 14, and the configuration may be configured according to the size of the device and the degree of arithmetic processing.
第4図は、第2図に示す、確率テーブル16−1のメモ
リ構成を示す。FIG. 4 shows the memory configuration of the probability table 16-1 shown in FIG. 2.
第4図において、確率テーブル16−1は各識別ラベル
毎に、音声入力の開始時点TOから一定時間間隔で算出
される確率情報を記憶する。In FIG. 4, a probability table 16-1 stores probability information calculated at fixed time intervals from the voice input start time TO for each identification label.
この確率情報は、入力音声の特徴パラメータと標準特徴
パラメータとの距離計算結果を各識別ラベル毎に百分率
で表わしたものである。なお、この確率情報に、距離計
算結果そのものを用いてもよい。This probability information is the result of calculating the distance between the input voice feature parameter and the standard feature parameter, expressed as a percentage for each identification label. Note that the distance calculation result itself may be used as this probability information.
この確率情報の示す値が高いほど入力音声の特徴パラメ
ータとそのラベルの特徴パラメータが類似していること
を示す。The higher the value of this probability information, the more similar the feature parameters of the input voice and the feature parameters of its label are.
第5図は第3図に示すMPU14−1が実行する制御手
順を示す。本発明実施例の動作を第5図のフローチャー
トを参照しながら説明する。FIG. 5 shows a control procedure executed by the MPU 14-1 shown in FIG. The operation of the embodiment of the present invention will be explained with reference to the flowchart of FIG.
第5図において、マイクロホン10から入力された音声
は入力開始時点TOからTI、T2・・・と一定時間間
隔でフーリエ変換器13により各時点毎の特徴パラメー
タに変換される。In FIG. 5, the sound input from the microphone 10 is converted into characteristic parameters for each time point by the Fourier transformer 13 at constant time intervals from input start time TO to TI, T2, . . . .
MPU14−1はこの入力特徴パラメータを入力すると
、DSP14−3により各標準パラメータとの距離計算
を指示する (ステップSl〜3)。MPU14−1は
DSP14−3から識別ラベルおよびこのラベルに対す
る距離計算結果を受は取ると、確率情報に変換し、確率
テーブル16−1 (第4図参照)に書き込む(ステッ
プS4)。従来例の記述で説明した発声゛アオ″と同じ
例を考えた場合、識別ラベル” A ” として例えば
0.7”、“i”として°゛02”・・・というように
確率情報が得られる。When the MPU 14-1 inputs this input feature parameter, it instructs the DSP 14-3 to calculate distances from each standard parameter (steps Sl to 3). When the MPU 14-1 receives the identification label and the distance calculation result for this label from the DSP 14-3, it converts it into probability information and writes it into the probability table 16-1 (see FIG. 4) (step S4). If we consider the same example as the utterance "Ao" explained in the description of the conventional example, probability information can be obtained such as, for example, 0.7" for the identification label "A", °゛02" for "i", etc. .
本実施例ではMPU114−1が入力音声の特徴パラメ
ータを入力した時点Tから過去4つの時点までの確率情
報の値を識別ラベル毎に合計し、最も犬ぎい合計結果を
有する識別ラベルを時点Tの音素識別結果として定める
。すなわち、計算式で表わすと下記の通りとなる。In this embodiment, the MPU 114-1 sums up the values of probability information for each identification label from the time T when the characteristic parameters of the input voice are input to the past four time points, and selects the identification label with the highest total result at the time T. Defined as the phoneme identification result. That is, the calculation formula is as follows.
ここで、
Pj(t):識別ラベルjの時刻tにおける確率合計、
(:t、j :時刻tにおける音素Jの確率、Bk:確
率の合計個数に対する重み係数、at:時刻tにおける
正規化係数、
! =合計すべき°確率の個数−1(時間幅の長さ)
、
本例においてはI−4、Bk−1、八t−1と設定した
例を説明している。Here, Pj(t): total probability of identification label j at time t,
(: t, j: probability of phoneme J at time t, Bk: weighting coefficient for the total number of probabilities, at: normalization coefficient at time t, ! = number of probabilities to be summed - 1 (length of time span) )
, In this example, an example in which I-4, Bk-1, and 8t-1 are set is explained.
したがって、この合計結果も第6図に示すように、識別
ラベル“A”、”I”、“0“の順に“0.7 ”、”
0.2 ” ”0.0 ” とFDD16に、MPl
+14−1により書き込む(ステップS5)。Therefore, as shown in FIG. 6, this total result is also "0.7" in the order of identification labels "A", "I", and "0".
0.2 ” “0.0 ” and FDD16, MPl
+14-1 is written (step S5).
このようにして全ての識別ラベルに対して、時刻TOで
の確率情報の値の合計が終了すると、MPU14−1は
合計値が最も高い識別ラベルを抽出し、FD016 に
記憶する(ステップ56〜57)。When the sum of probability information values at time TO is completed for all identification labels in this manner, the MPU 14-1 extracts the identification label with the highest total value and stores it in the FD016 (steps 56 to 57). ).
本例においては確率の°″0.7”を有する識別ラベル
“A′″が識別結果として記憶される。In this example, the identification label "A'" having a probability of 0.7 is stored as the identification result.
MPI+14−1はこのような手順を繰り返し実行し、
確率情報および確率の情報合計値をFD016に記憶し
て行くが、時刻T5においては、時刻T1〜T5までの
5個の確率値を合計する。以下、第6図に示すように合
計対象となる範囲(窓と称する)を順次移動させる。MPI+14-1 repeats these steps,
Probability information and probability information total values are stored in FD016, and at time T5, five probability values from times T1 to T5 are summed. Hereinafter, as shown in FIG. 6, the range to be totaled (referred to as a window) is sequentially moved.
また、このように得られる識別ラベルの中で最も多いラ
ベルを音韻認識結果として出力する。このための制御手
順は従来から周知のものを使用することができる(ステ
ップS8)。Furthermore, the label with the highest number of identification labels among the identification labels obtained in this way is output as the phoneme recognition result. A conventionally known control procedure can be used for this purpose (step S8).
連続音声では発声(音M)が変化すると、音韻毎の特徴
パターンも少しずつ変化するので、従来では発声と対応
する真の標準特徴パラメータの音韻ラベルは第2番目や
第3番目の識別候補として現われることが多い。In continuous speech, when the utterance (sound M) changes, the feature pattern for each phoneme also changes little by little, so conventionally, the phonological label of the true standard feature parameter corresponding to the utterance was used as the second or third identification candidate. often appear.
本発明は、音声の特徴パラメータ入力時点より過去の所
定時間内での距離計算結果をも音声認識処理に用いるの
で、たとえ、ある時点Tにおいて、識別結果が従来の方
式で第2番目候補となっても、連続する過去において複
数回識別結果が第1番目の候補となっていれば、その時
点Tにおいては第1番目の候補として決定される。Since the present invention also uses the distance calculation results within a predetermined period of time in the past from the time when the voice feature parameters are input, for voice recognition processing, even if at a certain time point T, the recognition result becomes the second candidate in the conventional method. However, if the identification results have been the first candidate multiple times in the past, the candidate is determined to be the first candidate at that time T.
例えば、従来例で説明した“アオ”という発声に対して
第1番目候補だけを選択する従来装置では“^AAtl
AtllJAUII00・・・”というように部分的に
異なる音素ラベルの識別結果が得られるが、本発明では
このラベル列“AAAAAA八八AUUへへ・・°°と
し)うように同一ラベルが複数個毎につながる、すなわ
ち、平滑化されたラベル列として得られる。For example, in the conventional device that selects only the first candidate for the utterance “ao” explained in the conventional example, “^AAtl
The identification result of partially different phoneme labels such as "AtllJAUII00..." is obtained, but in the present invention, the same label is identified for every plural number of phoneme labels as shown in the label string "AAAAAAA 88 AUU to...°°". In other words, it is obtained as a smoothed label sequence.
この結果、音韻の認識処理に対して従来例で説明した、
識別ラベルの補正処理を行う必要もなくなり、音声認識
確率が高まることは明らかである。As a result, as explained in the conventional example for phoneme recognition processing,
It is clear that there is no need to perform identification label correction processing, and the speech recognition probability increases.
本発明の応用形態としては次のことが考えられる。The following can be considered as an application form of the present invention.
1)本実施例では標準特徴パラメータと入力音声の特徴
パラメータとの類似度として距離計算結果から変換した
認識確率を用いているが、確率以外にも類似の順位を類
似度の指標として用いてもよい。したがって、この場合
は、識別ラベル毎に各時点での順位の値を合計し、最も
値の小さい識別ラベルが識別結果となる。1) In this example, the recognition probability converted from the distance calculation result is used as the degree of similarity between the standard feature parameter and the characteristic parameter of the input speech. good. Therefore, in this case, the ranking values at each time point are summed for each identification label, and the identification label with the smallest value becomes the identification result.
2)本実施例においては、音声認識に用いる類似度を全
てFDD16に記憶する例を示しているが、加算すべき
類似度の個数のみのレジスタやメモリを用意し、新しい
類似度が得られる毎にメモリ内容を更新十れば、音声認
識装置のメモリ容量を減じることが可能となる。2) In this embodiment, an example is shown in which all the similarities used for speech recognition are stored in the FDD 16, but registers and memories are prepared for the number of similarities to be added, and each time a new similarity is obtained, If the memory contents are updated frequently, it becomes possible to reduce the memory capacity of the speech recognition device.
3)本実施例においては過去の複数の類似度合計結果を
指標として識別ラベルを抽出しているが合計結果の他、
複数の類似度の平均値や統計的に平滑化した値を指標と
することができる。3) In this example, identification labels are extracted using the past multiple similarity total results as an index, but in addition to the total results,
An average value of multiple degrees of similarity or a statistically smoothed value can be used as an index.
(発明の効果〕
以上説明したように、本発明では、各時点毎に得られる
類似度から始まる過去の所定個数の類似度を識別ラベル
毎に第2演算手段により合計し、その合計結果の中の最
も類似度の高い識別ラベルを抽出手段により順次に抽出
する。この結果、各時点で得られる類似度が時系列的に
平滑化されるので、例えは入力音声に雑音が部分的に混
入したり、入力音声の一部の発生が変化したり、入力音
声の音韻が変化する場合においても、その音声の変化部
分が生じる悪影響が除去されて全体として正しいラベル
が認識され、以て音声の認識確率が高くなる。(Effects of the Invention) As explained above, in the present invention, a predetermined number of past similarities starting from the similarities obtained at each point in time are summed by the second calculation means for each identification label, and the total result is The identification label with the highest degree of similarity is sequentially extracted by the extraction means.As a result, the degree of similarity obtained at each point in time is smoothed over time. Even if the occurrence of a part of the input speech changes, or the phonology of the input speech changes, the negative effects caused by the changed part of the speech are removed and the correct label is recognized as a whole, thereby improving speech recognition. The probability increases.
第1図は本発明実施例の基本構成を示すブロック図、
第2図は本発明実施例の具体的な構成を示すブロック図
、
第3図は第2図に示すコンピュータシステムの回路構成
を示す回路図、
第4図は第2図に示す確率テーブル16−1のメモリ構
成を示すメモリマツプ、
第5図は第3図に示すMPUI4−1が実行する制御手
順を示すフローチャート、
第6図は本発明実り色例の類似度の計算過程を示す説明
図、
第7図は従来例の音声認識過程を示す説明図である。
!0 ・・・ マイクロホン、
14 ・・・ コンピュータシステム、14−1
・・・ マイクロプロセッサ(MPIJ)、16
・・・ フロッピディスク装置(FDD)、16−1
・・・ 確率テーブル。
不ン警甲ト芙方セイ列のコンじ°、−クシステム)4の
困−が、乞ホ11日本麿ゾ旧藁庁使1の石!畢シブル;
6−;乞示T〆tリマップフ公1し To 丁O〜
T1 TO〜T2 To−T3 丁O〜工4A
O,7152,32,43,2+ 0.2 0.
3 0.4 0.4 0.4U O,0
0,00,00,60,8未発明芙片渭1)め計算色程
を7
第6図
丁1〜T5 T2ダ6−−一
2.6 2.1
0.2 0. +
1.42・O
;す祝明詔Figure 1 is a block diagram showing the basic configuration of an embodiment of the present invention, Figure 2 is a block diagram showing a specific configuration of an embodiment of the invention, and Figure 3 is a circuit diagram of the computer system shown in Figure 2. 4 is a memory map showing the memory configuration of the probability table 16-1 shown in FIG. 2, FIG. 5 is a flow chart showing the control procedure executed by the MPUI 4-1 shown in FIG. 3, and FIG. FIG. 7 is an explanatory diagram showing the process of calculating the similarity of the invention fruit color example. FIG. 7 is an explanatory diagram showing the speech recognition process of the conventional example. ! 0...Microphone, 14...Computer system, 14-1
... Microprocessor (MPIJ), 16
... Floppy disk device (FDD), 16-1
... Probability table. The trouble of 4 is the stone of 11 Nippon Marozo former Wara Agency envoy 1, which is the combination of the non-guardian and fugata series, -ku system)! Bisible;
6-; Please show T〆t Remap Duke 1.
T1 TO~T2 To-T3 Ding O~Eng 4A
O, 7152, 32, 43, 2+ 0.2 0.
3 0.4 0.4 0.4U O,0
0, 00, 00, 60, 8 uninvented piece 1) Calculate the color degree 7 Figure 6 D1~T5 T2 da 6--1 2.6 2.1 0.2 0. + 1.42・O ;
Claims (1)
特徴パラメータ化し、該特徴パラメータと、音声の識別
ラベルが予め判明している複数の標準特徴パラメータの
各々との距離計算を行い、該計算結果に基いて前記連続
音声に対する複数の前記識別ラベルの類似度をそれぞれ
決定する第1演算手段と、 該第1演算手段により定められた、現時点から過去まで
連続する一定個数の前記類似度を前記識別ラベル毎に合
計する第2演算手段と、 該第2演算手段により合計された前記識別ラベル毎の類
似度の合計結果に基いて、最も類似度の高い合計結果と
対応する識別ラベルを抽出し、当該抽出された識別ラベ
ルを、前記現時点における音声認識結果として出力する
抽出手段と を具えたことを特徴とする音声認識装置。[Claims] 1) An input means for inputting continuous speech, and a method for converting the continuous speech inputted from the input means into feature parameters at fixed time intervals, the feature parameters and the identification label of the speech being known in advance. a first calculation means for calculating the distance between each of the plurality of standard feature parameters, and determining the degree of similarity of the plurality of identification labels for the continuous speech based on the calculation result; a second calculation means for summing up a certain number of consecutive similarities from the present time to the past for each of the identification labels; 1. A speech recognition device, comprising: extracting means for extracting an identification label corresponding to the total result with the highest degree of similarity, and outputting the extracted identification label as the speech recognition result at the current point in time.
Priority Applications (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP63257328A JPH02105200A (en) | 1988-10-14 | 1988-10-14 | Speech recognizing device |
Applications Claiming Priority (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP63257328A JPH02105200A (en) | 1988-10-14 | 1988-10-14 | Speech recognizing device |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| JPH02105200A true JPH02105200A (en) | 1990-04-17 |
Family
ID=17304836
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| JP63257328A Pending JPH02105200A (en) | 1988-10-14 | 1988-10-14 | Speech recognizing device |
Country Status (1)
| Country | Link |
|---|---|
| JP (1) | JPH02105200A (en) |
Cited By (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| EP0742546A3 (en) * | 1995-05-12 | 1998-03-25 | Nec Corporation | Speech recognizer |
-
1988
- 1988-10-14 JP JP63257328A patent/JPH02105200A/en active Pending
Cited By (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| EP0742546A3 (en) * | 1995-05-12 | 1998-03-25 | Nec Corporation | Speech recognizer |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| TWI396184B (en) | A method for speech recognition on all languages and for inputing words using speech recognition | |
| US8831942B1 (en) | System and method for pitch based gender identification with suspicious speaker detection | |
| WO2019037205A1 (en) | Voice fraud identifying method and apparatus, terminal device, and storage medium | |
| US8271282B2 (en) | Voice recognition apparatus, voice recognition method and recording medium | |
| US20080082323A1 (en) | Intelligent classification system of sound signals and method thereof | |
| US8271280B2 (en) | Voice recognition apparatus and memory product | |
| Yan et al. | Optimizing MFCC parameters for the automatic detection of respiratory diseases | |
| US8942977B2 (en) | System and method for speech recognition using pitch-synchronous spectral parameters | |
| CN108847251B (en) | Voice duplicate removal method, device, server and storage medium | |
| CN111785302A (en) | Speaker separation method, device and electronic device | |
| JP5091202B2 (en) | Identification method that can identify any language without using samples | |
| JPH02105200A (en) | Speech recognizing device | |
| JP2871120B2 (en) | Automatic transcription device | |
| JP3493849B2 (en) | Voice recognition device | |
| JP3102089B2 (en) | Automatic transcription device | |
| JP2019132948A (en) | Voice conversion model learning device, voice conversion device, method, and program | |
| JPH02141800A (en) | Speech recognition device | |
| CN107657962B (en) | A method and system for identifying and separating throat sounds and air sounds of speech signals | |
| Li | SPEech Feature Toolbox (SPEFT) design and emotional speech feature extraction | |
| JP2560785B2 (en) | Autoregressive model automatic order determination method | |
| JP2001109491A (en) | Continuous speech recognition apparatus and method | |
| JPH03223799A (en) | Method and apparatus for recognizing separated words, especially very large words | |
| JP2003022091A (en) | Speech recognition method, speech recognition device, and speech recognition program | |
| JP4219543B2 (en) | Acoustic model generation apparatus and recording medium for speech recognition | |
| CN118506791A (en) | Whale sound signal classification method and device |