JPH03228100A - Voice recognition device - Google Patents

Voice recognition device

Info

Publication number
JPH03228100A
JPH03228100A JP2023205A JP2320590A JPH03228100A JP H03228100 A JPH03228100 A JP H03228100A JP 2023205 A JP2023205 A JP 2023205A JP 2320590 A JP2320590 A JP 2320590A JP H03228100 A JPH03228100 A JP H03228100A
Authority
JP
Japan
Prior art keywords
word
phoneme
speech
standard pattern
standard
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Granted
Application number
JP2023205A
Other languages
Japanese (ja)
Other versions
JP2862306B2 (en
Inventor
Junichi Tamura
純一 田村
Tetsuo Kosaka
哲夫 小坂
Atsushi Sakurai
櫻井 穆
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Canon Inc
Original Assignee
Canon Inc
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Canon Inc filed Critical Canon Inc
Priority to JP2023205A priority Critical patent/JP2862306B2/en
Publication of JPH03228100A publication Critical patent/JPH03228100A/en
Priority to US08/194,807 priority patent/US6236964B1/en
Application granted granted Critical
Publication of JP2862306B2 publication Critical patent/JP2862306B2/en
Anticipated expiration legal-status Critical
Expired - Fee Related legal-status Critical Current

Links

Abstract

PURPOSE:To select a candidate word and segment a voice section at the same time by performing word spotting by continuous Maharanobis' DP, word by word, in a 1st stage. CONSTITUTION:A voice input part 100 inputs a voice signal from a microphone and a voice analysis part 101 transfers the input waveform. A word distance calculation part 102 by the continuous Maharanobis' DP matches a time series of feature vectors which are inputted one after another by the continuous Maharanobis' DP with the standard patterns of all words stored in a word standard pattern storage part 103 by the continuous Maharanobis' DP to calculate distances. A candidate word decision part 104 decides candidate words among word standard patterns according to the distances between respective frames and the word standard patterns which are found by the continuous Maharanobis' DP. The parameters of feature vectors in >=1 candidate section are stored in a parameter storage part 105.

Description

【発明の詳細な説明】 〔産業上の利用分野〕 本発明は音声認識装置、特に任意の話者が連続して発声
した単語等の音声を、高い認識率で認識←−−目一≠す
る音声認識装置に関するものである。
[Detailed Description of the Invention] [Industrial Application Field] The present invention is a speech recognition device, and in particular, a speech recognition device that recognizes speech such as words continuously uttered by any speaker with a high recognition rate. The present invention relates to a speech recognition device.

〔従来の技術〕[Conventional technology]

不特定話者認識に関する認識手法は、いくつか考案され
ているが、現状で最も一般的かつ本提案に比較的近い構
成を持つ不特定話者認識システムの従来例について述べ
る。
Although several recognition methods for speaker-independent recognition have been devised, we will discuss a conventional example of a speaker-independent recognition system that is currently the most common and has a configuration relatively similar to the present proposal.

従来、不特定大語量を目指した認識システムは第3図の
ような構成になっている。音声入力部1から入力された
音声は音声分析部2により音声のパワー項等を含むフィ
ルタバンク出力、LPCケプストラム等の特徴パラメー
タが求められ、ここでパラメータの圧縮等(フィルタバ
ンク出力の場合、K−L変換等による次元圧縮)も行わ
れる。
Conventionally, a recognition system aiming at an unspecified large amount of words has a configuration as shown in FIG. The voice input from the voice input unit 1 is processed by the voice analysis unit 2 to obtain characteristic parameters such as a filter bank output including the voice power term, LPC cepstrum, etc. -dimensional compression by L transformation etc.) is also performed.

(分析はフレーム単位で行われるので、以下、圧縮後の
特徴パラメータを特徴ベクトルと呼ぶ)次に連続音声中
から音素境界を決定するための処理が音素境界検出部3
により行われる。音素識別部4では、統計的な手法によ
り音素が決定される。5は多数の音素サンプルから作成
した音素標準パタンを格納する音素標準パタン格納部。
(Since the analysis is performed frame by frame, the compressed feature parameters are hereinafter referred to as feature vectors.) Next, processing for determining phoneme boundaries from continuous speech is carried out by the phoneme boundary detection unit 3.
This is done by The phoneme identification unit 4 determines phonemes using a statistical method. 5 is a phoneme standard pattern storage unit that stores phoneme standard patterns created from a large number of phoneme samples.

6は音素識別部4の出力結果から単語辞書7あるいは出
力された候補音素の中から修正規則部8により修正を行
って、最終的な認識結果を出力する単語識別部、9は認
識結果を表示する認識結果表示部である。
Reference numeral 6 denotes a word identification unit that corrects the output result of the phoneme identification unit 4 using a word dictionary 7 or the correction rule unit 8 from among the output candidate phonemes and outputs the final recognition result. 9 displays the recognition result. This is the recognition result display section.

通常、音素境界検出部3では、判別関数等を用いており
、音素識別部4でも同様に判別される。
Usually, the phoneme boundary detection section 3 uses a discriminant function or the like, and the phoneme identification section 4 also performs the same discrimination.

これら各構成要素の出力は一般的にある一定の閾値を満
足した候補が出力される。それぞれの候補について更に
複数の候補が出力されるが、7.8の様なTop  d
own的な情報等が用いられ最終的な単語に絞られる。
Each of these components generally outputs candidates that satisfy a certain threshold. Multiple candidates are output for each candidate, but Top d like 7.8
Own information is used to narrow down the words to the final word.

〔発明が解決しようとしている課題〕[Problem that the invention is trying to solve]

しかしながら、上記従来例の認識装置は基本的な構成が
ボトム・アップ型であるので、認識・過程のある箇所で
誤りが生じた場合、後の過程に悪影響を及ぼし易い形に
なっている。(例えば、音素境界検出部3において、音
素境界を誤った場合、その誤り方によっては音素識別部
4、単語識別部6に与える影響は大きい)つまり、最終
的な音声の認識率は各過程の誤り率の積に比例して下が
るので、高い認識率が得られなかった。
However, since the basic configuration of the conventional recognition apparatus described above is of a bottom-up type, if an error occurs at a certain point in the recognition process, it is likely to have an adverse effect on subsequent processes. (For example, if the phoneme boundary detection unit 3 makes a mistake in identifying a phoneme boundary, the effect on the phoneme identification unit 4 and word identification unit 6 will be large depending on how the error occurs.) In other words, the final speech recognition rate depends on each process. A high recognition rate could not be obtained because it decreased in proportion to the product of error rates.

又、特に、不特定話者を対象とする認識装置を構成する
場合各過程での判定の為の閾値の設定が非常に難しい。
In addition, especially when constructing a recognition device targeted at unspecified speakers, it is extremely difficult to set threshold values for determination in each process.

少な(とも候補の中に目的とするものが存在する様に閾
値を設定すると、各過程における候補群の数が多くなり
、複数候補単語の中から目的とする単語を正確に絞り込
む方法が非常に難しくなっていた。また、実環境下で認
識装置を使用しようとした場合、非定常ノイズ等がかな
り多く、少数単語の認識装置であっても認識率が低(、
実際、使いに(いものとなっていた。
If you set a threshold so that the target word exists among the candidates, the number of candidate groups in each process will increase, making it extremely difficult to accurately narrow down the target word from among multiple candidate words. In addition, when trying to use a recognition device in a real environment, there was a considerable amount of non-stationary noise, and the recognition rate was low even with a recognition device for a small number of words (,
In fact, it had become a useful item.

〔課題を解決するための手段〕[Means to solve the problem]

本発明によれば、上記従来の課題を解決するために、ス
ポツティング法により単語単位の音声区間の切り出し、
候補単語の選出を行い、次に音素単位でマツチングを行
うという2段階を設けることにより、候補単語の選出と
音声区間の切り出しが一気にでき、また、候補単語の絞
り込みを容易にしたものである。
According to the present invention, in order to solve the above-mentioned conventional problems, the speech section of each word is cut out by the spotting method,
By providing two steps: selecting candidate words and then performing matching on a phoneme basis, candidate words can be selected and speech sections can be extracted all at once, and candidate words can be narrowed down easily.

また、本発明によれば複数の環境下における音素の標準
パタンを用意することにより、単語の標準パタンを複数
の環境について用意するよりも少ない情報量で多(の状
況における入力音声を認識することが可能となる。
Furthermore, according to the present invention, by preparing standard patterns of phonemes under a plurality of environments, it is possible to recognize input speech in a large number of situations with a smaller amount of information than when standard patterns of words are prepared for a plurality of environments. becomes possible.

〔実施例1〕 第1図は本発明による音声認識システムの基本構成図で
、100は音声入力部、101は入力された音声を分析
、圧縮し、特徴ベクトルの時系列に変換する音声分析部
、103は多数の話者が発声した単語データから求めた
標準パタンを格納する単語標準パタン格納部、102は
音声分析部101の特徴ベクトル系列と単語標準パタン
格納部103に格納されている各々の標準パタンを入力
データのフレームごとに連続マハラノビスDPを用いて
距離を算出する連続マハラノビスアDPによる単語距離
計算部、104は連続マハラノビスDPより求めた各フ
レームと単語標準パタンとの距離の値により単語標準パ
タンの中から候補となる単語を判別する候補単語判別部
、105は候補になった1つ以上の単語区間の特徴ベク
トルのパラメータを格納するパラメータ格納部、106
は多数話者の発声した音声の中から音素単位で作成され
た標準パタンを格納する音素標準パタン格納部、107
は候補となった単語の特徴ベクトル系列について音素単
位で連続マハラノビスDPにより入力データと音素標準
パタンの距離計算を行う連続マハラノビスDPによる音
素距離計算部、108は1つ以上の候補単語のそれぞれ
についてマツチングされた各音素列から最も適当な単語
を識別して出力する音素単位の認識結果による識別部。
[Embodiment 1] FIG. 1 is a basic configuration diagram of a speech recognition system according to the present invention, in which 100 is a speech input section, and 101 is a speech analysis section that analyzes and compresses input speech and converts it into a time series of feature vectors. , 103 is a word standard pattern storage unit that stores standard patterns obtained from word data uttered by a large number of speakers, and 102 is a word standard pattern storage unit that stores the standard patterns obtained from word data uttered by many speakers. A word distance calculation unit using continuous Mahalanobis DP calculates the distance of a standard pattern using continuous Mahalanobis DP for each frame of input data; 104 is a word distance calculation unit using continuous Mahalanobis DP, which calculates a distance between each frame and the word standard pattern calculated by continuous Mahalanobis DP; 105 is a candidate word discrimination unit that discriminates candidate words from standard patterns; 105 is a parameter storage unit that stores parameters of feature vectors of one or more word sections that have become candidates; 106;
107 is a phoneme standard pattern storage unit that stores standard patterns created in units of phonemes from voices uttered by multiple speakers;
108 is a phoneme distance calculation unit using continuous Mahalanobis DP that calculates the distance between the input data and the phoneme standard pattern using continuous Mahalanobis DP on a phoneme basis for the feature vector series of candidate words, and 108 performs matching for each of one or more candidate words. An identification unit based on recognition results of phoneme units that identifies and outputs the most appropriate word from each phoneme string.

109は例えば音声応答等の手段により音声認識結果を
出力する結果出力部である。図中、第1部は音声区間の
切り出しと供に単語の候補の絞り込み、第2部は候補単
語内での音素単位の認識部を示す。
Reference numeral 109 denotes a result output unit that outputs a voice recognition result by means such as voice response. In the figure, the first part shows the extraction of speech sections and the narrowing down of word candidates, and the second part shows the recognition part of phoneme units within the candidate words.

次に動作の流れを説明する。まず、音声入力部100は
、マイクから音声信号を入力し、音声分析部101に入
力波形を転送する。音声入力部100は音声入力の受付
時間中は常に音声又は周囲のノイズ信号等を取り込み、
音声入力波形をディジタル値に変換した波形として音声
分析部101へ転送する。音声分析部101では、常に
入力されて来る波形を10m5ec〜30m5ec程度
の窓幅で分析を行い、2m5ec〜10m5ecの長さ
を持つフレームごとに、特徴パラメータを求める特徴パ
ラメータの種類としては比較的高速に分析可能なLPC
ケプストラム、LPCメルケブストラム、高精度にパラ
メータを抽出したい場合はFFTケプストラム、FFT
メルケブストラム等が一般的で、他にフィルタバンク出
力値もある。
Next, the flow of operation will be explained. First, the audio input section 100 inputs an audio signal from a microphone and transfers the input waveform to the audio analysis section 101 . The audio input unit 100 always takes in audio or ambient noise signals, etc. during the audio input reception time, and
The audio input waveform is converted into a digital value and transferred to the audio analysis unit 101 as a waveform. The speech analysis unit 101 analyzes the constantly input waveform with a window width of about 10 m5 ec to 30 m5 ec, and calculates the feature parameter for each frame with a length of 2 m5 ec to 10 m5 ec, which is relatively fast considering the type of feature parameter. LPC that can be analyzed in
Cepstrum, LPC melkebstrum, FFT cepstrum, FFT if you want to extract parameters with high precision
Melkebstrum etc. are common, and there are also filter bank output values.

また、正規化されたパワー情報を用いたり、パラメータ
の各次元ごとに重み係数を掛けたりして、システムの使
用状況に最も適したパラメータで、フレームごとに分析
される。次に、分析された特徴パラメータの次元につい
て圧縮を行う。ケプストラムパラメータは、通常係数の
1次の項〜12次の項の中から必要な次元数(例えば6
次元)だけ抜き出し、これを特徴ベクトルとする。
In addition, each frame is analyzed using the parameters most suitable for the usage status of the system by using normalized power information or by multiplying each parameter dimension by a weighting coefficient. Next, compression is performed on the dimensions of the analyzed feature parameters. Cepstral parameters are usually determined by the required number of dimensions (for example, 6
dimension) and use this as a feature vector.

フィルタバンク出力を特徴パラメータとした場合、例え
ばに−L変換、フーリエ変換等の直交変換により次元圧
縮し、低次項を用いる。これら圧縮された17レム分の
パラメータを特徴ベクトル、次元圧縮された後の特徴ベ
クトルの時系列を特徴ベクトルの系列(或は、単にパラ
メータ)と呼ぶことにする。
When the filter bank output is used as a feature parameter, the dimension is compressed by orthogonal transformation such as -L transformation or Fourier transformation, and low-order terms are used. These compressed parameters for 17 rems will be referred to as feature vectors, and the time series of feature vectors after dimension compression will be referred to as feature vector series (or simply parameters).

本実施例では分析窓長を25.6m5ecで分析し、フ
レーム周期10m5ec、FFTスペクトルのピークを
通るスペクトル包絡から、メルケプストラム係数を求め
た後、係数の2次〜6次を用い、これを1フレ一ム分の
特徴ベクトルとする。ここでメルケブストラムの0次項
はパワーを表わす。
In this example, the analysis window length was 25.6 m5 ec, the frame period was 10 m5 ec, and the mel cepstral coefficients were obtained from the spectrum envelope passing through the peak of the FFT spectrum. Let it be a feature vector for one frame. Here, the zero-order term of Melkebstrum represents power.

次に、単語標準パタン格納部103に格納する標準パタ
ンの作成方法について述べる。本システムでは例として
発声変形を含めた10数字“ゼロ、サン、二、レイ、ナ
ナ、ヨン、ゴ、マル、シ、ロク、夕、ハチ、シチ、キュ
ウ、イチ”と“ハイ、イイエ”の計17単語の認識につ
いて述べる。標準パタンは多数話者の発声した単語音声
から作成する。本実施例では1単語の標準パタンを作成
するのに50人分の音声サンプルを用いる。(音声サン
プル数は多ければ多い程良い)第2図(a)に、標準パ
タンの作成手順を表わすフローチャートを示す。
Next, a method for creating standard patterns to be stored in the word standard pattern storage section 103 will be described. In this system, for example, the 10 numbers "Zero, San, Two, Rei, Nana, Yon, Go, Maru, Shi, Roku, Yu, Hachi, Shichi, Kyu, Ichi" and "Hai, Iie" including vocalization variations are used. We will describe the recognition of a total of 17 words. Standard patterns are created from word sounds uttered by multiple speakers. In this embodiment, voice samples from 50 people are used to create a standard pattern for one word. (The larger the number of audio samples, the better.) FIG. 2(a) shows a flowchart showing the procedure for creating a standard pattern.

まず、音声サンプルから標準パタンを作成する際の仮の
比較対象となるコアパタン(核パタン)を選択する(S
 200)。選択方法は50単語の中で発声時間長と発
声パタンか最も平均的な単語を用いる。次に、サンプル
の単語を入力しく5201)、入力単語とコアパタンと
の時間軸伸縮マツチングを行い、時間正規化距離が最小
となるマツチング経路に沿って、各フレームごとに平均
ベクトル、及び分散共分散行列を作成する(S202)
。ここで時間軸伸縮マツチングの方法としてDPマツチ
ングを用いる。次に入力単語の話者番号を次々変えてゆ
き(S204)50名分の単語Si (i=1〜50)
について、各フレームごとに特徴ベクトルの平均値及び
、分散共分散行列を求める(S203.5205)。こ
の様にして計17単語についてそれぞれ上記過程と同様
にして単語標準パタンを作成し単語標準パタン格納部1
03に格納しておく。
First, select a core pattern that will be a temporary comparison target when creating a standard pattern from a voice sample (S
200). The selection method uses the word with the most average utterance duration and utterance pattern among the 50 words. Next, input a sample word (5201), perform time axis expansion/contraction matching between the input word and the core pattern, and calculate the average vector and variance/covariance for each frame along the matching path that minimizes the time normalized distance. Create a matrix (S202)
. Here, DP matching is used as a time axis expansion/contraction matching method. Next, change the speaker numbers of the input words one after another (S204) Words Si for 50 people (i=1 to 50)
For each frame, the average value of the feature vector and the variance-covariance matrix are determined (S203.5205). In this way, word standard patterns are created for a total of 17 words in the same manner as above, and the word standard pattern storage unit 1
Store it in 03.

連続マハラノビスDPによる単語距離計算部102では
、連続マハラノビスDPにより次々と入力される特徴ベ
クトルの時系列について単語標準パタン格納部103に
格納されている全ての単語の標準パタンとの連続マハラ
ノビスDPによるマツチングを行い、距離を計算する。
The word distance calculation unit 102 using continuous Mahalanobis DP matches the time series of feature vectors input one after another using continuous Mahalanobis DP with the standard patterns of all words stored in the word standard pattern storage unit 103. and calculate the distance.

ここで、連続マハラノビスDPについて説明する。連続
DPの手法は一般的で、特定話者が連続に発声した文章
の中から目的とする単語、或は、音節等の単位を探し出
す方法である。これはワードスポツティングと呼ばれ、
目的とする音声区間の切り出しと同時に認識も行ってし
まうという画期的な方法である。本実施例では連続DP
法の各々のフレーム内における距離にマハラノビス距離
を用いる事により、不特定性を吸収している。
Here, continuous Mahalanobis DP will be explained. The continuous DP method is common and is a method of searching for a target word, syllable, or other unit from a sentence continuously uttered by a specific speaker. This is called word spotting.
This is an innovative method that performs recognition at the same time as cutting out the desired audio section. In this example, continuous DP
Unspecificity is absorbed by using the Mahalanobis distance for the distance within each frame of the method.

第2図(b)は、“ゼロ という単語の標準パタンと“
ゼロ”という単語を発声した時の入力音声を無声区間も
含めて特徴ベクトルの時系列に分析したものとを連続マ
ハラノビスDPによりマツチングした結果を示したもの
である。図中、黒が濃(出ている所は標準パタンと入力
パタンの距離が大きい所、黒が薄く、白に近い所は標準
パタンと入力パタンの距離が小さい所である。マツチン
グを行った結果の下には累積距離の時間変化を示す。こ
の累積距離はその時点が終端となるDPパスの距離を示
すもので、DPパスを求めてその値をメモリに保存する
。このメモリに保存したDPパスを、音声区間の始端を
求める為につかう。例えばこの図においては距離が最小
となった時のDPパスを示したが、標準パタンと入力パ
タンが似ていた場合、累積距離が任意に定めた閾値より
小さ(なり、その標準パタンの単語を候補単語と認める
。そして、入力パタンから音声区間を切り出すために、
累積距離が闇値より小さく、更に最小である時点からD
Pパスをメモリから呼び出してバックトラックすること
により、音声区間の始端が求められる。こうして求めら
れた音声区間の特徴ベクトルの時系列をパラメータ格納
部105に格納する。
Figure 2 (b) shows the standard pattern of the word “zero” and “
This figure shows the results of matching the input speech when the word "zero" was uttered, including silent sections, analyzed into a time series of feature vectors using continuous Mahalanobis DP. The areas where the distance between the standard pattern and the input pattern is large, and the areas where black is thin and close to white are where the distance between the standard pattern and the input pattern is small. Below the matching result is the cumulative distance time. This cumulative distance indicates the distance of the DP path that ends at that point in time.The DP path is calculated and its value is stored in memory.The DP path stored in this memory is For example, this figure shows the DP path when the distance is the minimum, but if the standard pattern and the input pattern are similar, the cumulative distance is smaller than an arbitrarily determined threshold. Recognize words in the standard pattern as candidate words.Then, in order to extract speech sections from the input pattern,
From the point where the cumulative distance is smaller than the darkness value and is also the minimum, D
By recalling the P path from memory and backtracking, the start of the voice section is determined. The time series of feature vectors of the voice section thus obtained is stored in the parameter storage unit 105.

今まで説明してきた処理系により、まず候補単語と、そ
の音声区間を分析した特徴ベクトルの系列と、連続マハ
ラノビスDPによる累積距離の結果が得られる。ここで
、候補単語の中で“ンチ”と“シ”の様に音声区間が重
なっているものが複数選択された時、この場合“シチ”
の方を選択し“シ”は切り捨てる。  ロク”と”り“
も同様に、“り”の音声区間の大部分が(ここでは80
%以上とする) ロク”に含まれている時は、“り”は
切り捨てて“ロク”のみについて検証を行う。
The processing system described so far first obtains a candidate word, a series of feature vectors obtained by analyzing its speech interval, and cumulative distance results using continuous Mahalanobis DP. Here, when multiple candidate words with overlapping phonetic intervals, such as "nchi" and "shi", are selected, in this case "shichi"
Select the one and discard “shi”. “Roku” and “Ri”
Similarly, most of the vocal section of “ri” (here, 80
% or more) When it is included in "Roku", "Ri" is discarded and only "Roku" is verified.

本実施例では音素標準パタン格納部106に母音(a;
 i、u、e、o)と子音(z、s、n。
In this embodiment, the vowel (a;
i, u, e, o) and consonants (z, s, n.

rSg、m)sh i、に、h、c i)につし1て音
素の標準パタンを作成しておく、作成方法は単語標準パ
タン格納部103と同様の方法であらかじめ作成してお
く。連続マハラノビスDPによる音素距離計算部107
ではパラメータ格納部105に格納されている候補単語
として切り出された音声区間について各音素とのマツチ
ングを行う。
rSg, m) sh i, ni, h, c i) A standard phoneme pattern is created in advance using the same method as the word standard pattern storage section 103. Phoneme distance calculation unit 107 using continuous Mahalanobis DP
Then, matching of each phoneme is performed for the voice section cut out as a candidate word stored in the parameter storage unit 105.

連続マハラノビスDPによる単語距離計算部102と同
様に累積距離が最小となった位置からその音素の区間を
計算する。(候補単語判別部104と同様、累積距離が
最小となった時点をその音素の終端とし、始端は連続D
Pパスをバックトラックにより求める) 本実施例では例えば“ゼロ”#“zero“が候補単語
の場合その音声区間について“Ze”r”  0”の4
種類の音素についてのみマツチングを行う。4種の音素
と上記“zero”と判別され、候補となった音声区間
のマツチングの結果、各音素の累積距離が最小となる点
についてその位置関係と、最小距離の平均値を求めるこ
の様子を第2図(C)に示す。
Similar to the word distance calculation unit 102 using continuous Mahalanobis DP, the section of the phoneme is calculated from the position where the cumulative distance is the minimum. (Similar to the candidate word discriminator 104, the point where the cumulative distance is the minimum is the end of the phoneme, and the start is the continuous D
In this embodiment, for example, if "zero"#"zero" is a candidate word, "Ze"r"0" 4 for that speech section.
Matching is performed only for different types of phonemes. As a result of matching the four types of phonemes with the above-mentioned "zero" and the candidate speech interval, we will calculate the positional relationship of the point where the cumulative distance of each phoneme is the minimum and the average value of the minimum distance. It is shown in FIG. 2(C).

各々の音素についてマツチングの結果の距離の最小値と
、その位置をフレームで表わし音素単位の認ぷ結果によ
る認識部108に送る。この例では“Z”について最小
値は“J 1フレ一ム位置は“Z、”である。音素単位
の認識結果による認識部108では、連続マハラノビス
DPによる音素距離計算部107から送られてきたデー
タを基に最終的な単語の識別を行う。まず、候補単語の
音素列の順番(フレームの位置)がz、<6.<r+<
O+であるか否かを調べる。もしこの順番であれば認識
単語は“ゼロ (zero)”平均Hよりも小さいなら
ば、認識結果として“ゼロを出力する。
The minimum value of the distance as a result of matching for each phoneme and its position are expressed in a frame and sent to the recognition unit 108 based on the recognition results for each phoneme. In this example, the minimum value for "Z" is "J" and the 1st frame position is "Z,".The recognition unit 108 based on the recognition result of each phoneme receives the information sent from the phoneme distance calculation unit 107 using continuous Mahalanobis DP. The final word is identified based on the data.First, the order (frame position) of the phoneme string of the candidate word is z, <6.<r+<
Check whether it is O+. In this order, the recognized word is "zero". If it is smaller than the average H, "zero" is output as the recognition result.

第2図(d)は単語候補の出力結果(候補単語判別部1
04の出力結果)を示したものである。
FIG. 2(d) shows the output result of word candidates (candidate word discriminator 1
04) is shown.

■は単語“ハチ”、■は単語“シチ”、■は単語“ン”
が候補として出力される。が、ここで前に述べたように
■は■の区間に80%以上含まれており、かつ同一のシ
”が■中に存在するので音素レベルでの識別は■■につ
いて行なう。
■ is the word “Hachi”, ■ is the word “Shichi”, ■ is the word “N”
is output as a candidate. However, as mentioned above, 80% or more of ``■'' is included in the interval ``■'', and the same ``'' exists in ``■'', so identification at the phoneme level is performed for ``■■''.

ケース■ 単語S1の音素列“1hlalc”と単語S
2の音素列“1shli clil”についてマツチングした結 果、どちらも音素の順番が、候補単語と等しい場合、か
つ、個々の音素の距離がH(閾値)より小さい場合中平
均累積距離Xの小さい方、を出力する。
Case ■ Phoneme sequence “1hlalc” of word S1 and word S
As a result of matching for the phoneme sequence "1shli clil" of 2, if the order of the phonemes is the same as that of the candidate word, and the distance of each phoneme is smaller than H (threshold), the one with the smaller average cumulative distance X, Output.

ケース■ どちらも順番が異なるが個々の音素の距離が
閾値(H)より小さい場合中単語と音素列の文字列によ
るDPマツチングを行い。その距離の閾値(1)により
決定する。
Case ■ If the order is different in both cases, but the distance between individual phonemes is smaller than the threshold (H), DP matching is performed using the character string of the middle word and the phoneme string. It is determined based on the distance threshold (1).

ケース■ 順番が合っているか、個々の音素の閾値が(
H)をクリアしていない場合中リジェクト ケース■ 順番が異なり、音素の閾値もクリアしていな
い場合弁リジェクト 音素単位の認識結果による単語の識別方法は前記の方法
に限らない。後に他の実施例でも述べるが音素の単位を
どの様な形で定義し、標準パタンを作成しておくか、或
は同一の音素でも複数用意する事によって音素判別に用
いる閾値Hの値、或は識別アルゴリズムは異なる。よっ
て、平均累積距離と音素順位のどちらを優先させるか等
の識別アルゴリズムは一意に決まらない。
Case ■ Is the order correct or the threshold of each phoneme (
If H) is not cleared, medium reject case ■ If the order is different and the phoneme threshold is not cleared, the word identification method based on the recognition result of the phoneme unit is not limited to the above method. As will be described later in other embodiments, the value of the threshold H used for phoneme discrimination can be determined by defining the unit of phoneme and creating a standard pattern, or by preparing a plurality of the same phoneme. The identification algorithm is different. Therefore, the identification algorithm, such as which one to give priority to, the average cumulative distance or the phoneme order, cannot be uniquely determined.

)素中位の認忠結果による認識部108て最終結果とし
て出力した例えば音声(単語)を結果出力部109で出
力する。電話等の音声情報のみで認識さ−せる場合、認
識結果を「“ゼロ”ですね?」と、例えば音声合成手段
を用いて確認する。単語の識別の結果、距離が十分小さ
ければ認識結果を確認をせずに、それに対応した次の処
理へと移行する。
) The result output unit 109 outputs, for example, speech (word) output as the final result by the recognition unit 108 based on the recognition result of the intermediate level. When recognizing only voice information from a telephone or the like, the recognition result is confirmed using, for example, voice synthesis means, such as, "Isn't it 'zero'?" As a result of word identification, if the distance is sufficiently small, the process moves to the next corresponding process without checking the recognition result.

〔実施例2〕 前記実施例1では、後半の音素単位の認識結果を、認識
対象とする単語に含まれる全ての音素について標準パタ
ンを作成しておいた。しかし、音素はその種類によって
は、周囲の音韻環境、話者等の相異により、変形も激し
い。よって同一の音素でもパタンの異なる音素はパタン
に応じ複数用意しておくと、より確度の高い認識結果が
得られる、例えば母音11についてみると“イチ“ハチ
”シチ”に見られる様に話者によって無声化する事がか
なりある。音素レベルでの認識は候補となった単語と、
その音声区間において厳密に検定して結果を出さなけれ
ばならないので、母音filでも、有声の111、無声
化の11それぞれについて、数種類の標準パタンを作っ
ておく、他の音素についても同様で、例えば1gなどバ
ス部が存在するものとしないものがある。
[Example 2] In the above-mentioned Example 1, standard patterns were created for all the phonemes included in the word to be recognized as the recognition result of the second half of the phoneme unit. However, depending on the type of phoneme, it is subject to severe deformation due to differences in the surrounding phonetic environment, speakers, etc. Therefore, if you prepare multiple phonemes with different patterns for the same phoneme, you can obtain more accurate recognition results.For example, for vowel 11, as seen in "ichi"hachi", it is possible to obtain a more accurate recognition result. There are many cases where the voice is muted. Recognition at the phoneme level uses candidate words and
Since it is necessary to strictly test and produce results in that speech interval, several types of standard patterns are created for the vowel fil, 111 for voiced and 11 for unvoiced.The same is true for other phonemes, for example. Some, such as 1g, have a bus part and others do not.

但しこれらの音素について標準パタンを作成する場合、
少なくとも1つの標準パタンを作成する為に、各フレー
ムの特徴ベクトルの次元数をnとするとn2+α個程度
の音声データを必要とする。
However, when creating standard patterns for these phonemes,
In order to create at least one standard pattern, approximately n2+α audio data are required, assuming that the number of dimensions of the feature vector of each frame is n.

〔実施例3〕 また、音素単位で識別する別の例として、音素の単位を
変えると更に良い結果となる。前記実施例1では、la
l  lit  ・・・ (mnl  lrlに示す様
に、音声の単位としてはかなり小さい母音、子音、を別
々に扱っていた。
[Example 3] Further, as another example of identifying by phoneme unit, better results can be obtained by changing the unit of phoneme. In Example 1, la
l lit... (As shown in mnl lrl, vowels and consonants, which are quite small units of speech, were treated separately.

実際、人間が発声する連続した単語音声はアナウンサー
等を別にして日常生活においては、個々の音素の特徴を
明確に発声している事は少ない。
In fact, apart from announcers and the like, in daily life, continuous word sounds uttered by humans are rarely uttered clearly with the characteristics of individual phonemes.

データを見てもここがlalでここが1mlであると判
定出来る部分は時間的にもかなり短く、大部分は調音結
合部である。(調音結合部とは、例えば“イア”と発声
した場合“イ”の定常部がら“ア”の定常部に遷移する
(中途半端な)部分である。) よって、音素の単位を調音結合部を含むVCV型とし、
語頭に関してはCVを用いると、前記実施例1で述べた
複数候補の単語が出現した時も、順番が異なって来る場
合の割合が減少するため、最終出力単語の判別がしやす
い。(■・・・母音VowelSC・・・子音Con5
onantでvCvは、母音−子音−母音、連鎖の事)
もちろん、vCvの標準パタンは、連続音声中から切り
出したサンプルから作成する。
Looking at the data, the part where it can be determined that this is lal and this is 1ml is quite short in terms of time, and most of it is the articulatory junction. (The articulatory junction is the part where, for example, when you say "ia", the steady part of "i" transitions to the steady part of "a" (halfway).) Therefore, the unit of phoneme is the articulatory junction. A VCV type including
When CV is used for word beginnings, even when multiple candidate words described in the first embodiment appear, the proportion of cases in which they appear in different orders is reduced, making it easier to determine the final output word. (■...Vowel SC...Consonant Con5
In onant, vCv is a vowel-consonant-vowel chain)
Of course, the standard vCv pattern is created from samples extracted from continuous audio.

[実施例4〕 前記実施例では音素標準パタン格納部106に格納する
音素のパタンのマルチ化、音素単位の定義、方法につい
て述べた。
[Embodiment 4] In the above embodiment, the multiplication of phoneme patterns stored in the phoneme standard pattern storage unit 106, the definition of phoneme units, and the method were described.

単語標準パタン格納部103についても同様の事が言え
る。しかし、単語標準パタンについては、厳密にパタン
をカテゴライズしようとするとパタンの数が多くなり過
ぎる場合がある。また、個々の単語について多数話者の
発声サンプルを集め、分析する事は容易でないので、こ
こでは、個々の単語の発声時間長によりカテゴライズを
行う。本認識システムの第1段階では、候補単語の中に
、目的とする単語が100%入っている事が前提条件で
ある。本方式は基本的に時間伸縮マツチングを行ってい
るので、標準パタンから極端に外れた発声時間長の単語
だし、リジェクトされてしまう可能性が高いからである
The same can be said about the word standard pattern storage section 103. However, when it comes to word standard patterns, attempting to strictly categorize them may result in too many patterns. Furthermore, since it is not easy to collect and analyze utterance samples of many speakers for each word, here, the categorization is performed based on the length of utterance time of each word. In the first stage of this recognition system, it is a prerequisite that 100% of the candidate words contain the target word. This is because this method basically performs time expansion/contraction matching, so there is a high possibility that the word will be rejected if the length of the utterance is extremely different from the standard pattern.

よって、少なくとも認識装置に対し、協力的な話者が発
声する音声の時間長を調べ、その全時間長をカバーする
様、標準パタンをマルチ化する。
Therefore, at least for the recognition device, the time length of the voice uttered by the cooperative speaker is investigated, and the standard pattern is multiplied so as to cover the entire time length.

マルチ化する際、極端に長い発声のサンプルは得うレに
(いので、平均的な特徴ベクトルのフレーム数を第2図
(e)に示す様に2倍、3倍に増やしても良い。
When performing multiplexing, it is difficult to obtain samples of extremely long utterances (because it is difficult to obtain samples of extremely long utterances), the number of frames of the average feature vector may be doubled or tripled as shown in FIG. 2(e).

第2図(e)では、音素la1mlul  ”アム”を
単位とした基準パタンの発声時間長を2倍にした例を示
す。
FIG. 2(e) shows an example in which the utterance time length of the standard pattern in which the unit is the phoneme la1mlul "am" is doubled.

発富時闇長を拡大する際、気をつけなければならない点
は、例えばlpl、Ml、lkl等の破裂子音等を含む
場合である。この例に示す様に子音によっては発声時間
長が長くなっても、子音部の発声時間長はそれほど変わ
らない。よって、子音によって拡大の方法をテーブル等
により、個々に変える手段を持つと、簡易に正確かつ、
時間長の異なる標準パタンか作成できる。
When expanding Hatsufujiyamaga, care must be taken when, for example, plosive consonants such as lpl, Ml, and lkl are included. As shown in this example, even if the utterance time length of some consonants becomes longer, the utterance time length of the consonant part does not change much. Therefore, having a means to individually change the method of expansion depending on the consonant using a table or the like would make it easier and more accurate.
Standard patterns with different time lengths can be created.

実際に発声時間長の長い音声サンプルを集め、これらの
データから標準パタンを作成する方法がより良い標準パ
タンを作成できる。
A better standard pattern can be created by actually collecting voice samples with a long utterance time and creating a standard pattern from this data.

第2図(f)は、母音の1フレームを2倍、3倍、4倍
と重複させて標準パタン長を拡大した時、子音部のフレ
ームの重複倍率を示したテーブルである。第2図(g)
に“ログの(母音の)倍率を“3倍”にした時の様子を
示す。
FIG. 2(f) is a table showing the overlapping magnification of consonant frames when the standard pattern length is expanded by overlapping one frame of a vowel twice, three times, and four times. Figure 2 (g)
Figure 3 shows the situation when the log (vowel) magnification is set to 3x.

また、第1図の単語標準パタン格納部103は単語単位
に限らない。文節単位でも良いし、無意味音節の連鎖で
も良い。この場合単語標準パタン格納部103の単位を
(VCV、VCVCV、cv、vv、cvcv、 ・−
・等)とし、音素標準パタン格納部106の単位(CV
、VC,V・・・等)にする事も可能である。
Furthermore, the word standard pattern storage section 103 shown in FIG. 1 is not limited to word units. It can be a phrase unit or a chain of meaningless syllables. In this case, the unit of word standard pattern storage 103 is (VCV, VCVCV, cv, vv, cvcv, -
・etc.), and the unit (CV
, VC, V, etc.).

〔実施例5〕 前記実施例1では、第1図に示す処理系基本構成の第2
部において第1部の出力として得た候補単語について更
に細かい音素単位(例えばC1V、CV、cvcSvc
v等)で連続DP等のスポツティング処理を行い、結果
を出力する方法について述べた。しかし、本実施例にお
いては第1部の出力する候補単語を音素単位で認識する
方法として、スポツティング以外の方法を述べる。それ
は、複数の音声サンプルから得た音素標準パタンを候補
単語の音素系列に合わせて接続して作った単語と、音声
区間として切り出された入力音声の特徴ベクトルとのマ
ツチングを行うという方法である。この方法によっても
高い認識率が得られる。
[Example 5] In Example 1, the second example of the basic configuration of the processing system shown in FIG.
In the first part, the candidate words obtained as the output of the first part are further divided into finer phoneme units (for example, C1V, CV, cvcSvc
We have described a method for performing spotting processing such as continuous DP using a software (e.g., v) and outputting the results. However, in this embodiment, a method other than spotting will be described as a method for recognizing candidate words output in the first part in units of phonemes. This method involves matching a word created by connecting standard phoneme patterns obtained from multiple speech samples according to the phoneme sequence of a candidate word with a feature vector of the input speech extracted as a speech interval. This method also provides a high recognition rate.

本実施例における音素単位の認識処理系の基本構成を第
4図に示す。
FIG. 4 shows the basic configuration of the phoneme-based recognition processing system in this embodiment.

第1図候補単語判別部10.4において判別された候補
単語と音声区間として切り出された入力音声の特徴ベク
トルは以後第4図に示す構成において処理される。まず
、入力音声の特徴ベクトルはパラメータ格納部105に
、候補単語は標準パタン生成規則部110に送られる。
The candidate words discriminated by the candidate word discriminator 10.4 in FIG. 1 and the feature vectors of the input speech cut out as speech sections are thereafter processed in the configuration shown in FIG. 4. First, the feature vector of the input speech is sent to the parameter storage section 105, and the candidate words are sent to the standard pattern generation rule section 110.

標準パタン生成規則部110では音素標準パタン格納部
106中の音素標準パタンを候補単語の音素系列に従っ
て接続し、これとパラメータ格納部105に格納してお
いた入力音声の特徴ベクトルのパタンマツチングをパタ
ンマツチング部111において行う。
The standard pattern generation rule unit 110 connects the phoneme standard patterns in the phoneme standard pattern storage unit 106 according to the phoneme sequence of the candidate word, and performs pattern matching between this and the feature vector of the input speech stored in the parameter storage unit 105. This is performed in the pattern matching section 111.

パタンマツチングで得た音声の認識結果を結果出力部1
09より出力する。
Result output unit 1 outputs the speech recognition results obtained by pattern matching.
Output from 09.

標準パタン生成規則部110の詳細な構成図を第5図に
示す。まず、第1部の結果として出力される候補単語の
音素系列と、音声区間として切り出された入力音声の特
徴ベクトルが出力される。
A detailed configuration diagram of the standard pattern generation rule section 110 is shown in FIG. First, the phoneme sequence of the candidate word output as a result of the first part and the feature vector of the input speech cut out as a speech section are output.

ここでは、例えば“tokus imasi (徳島布
)”と入力した時に、候補単語として“tOkusim
asi    f ukus imas 1(福島布)
”  ”hirosimasi (広島布)”の3単語
が選出された場合の処理について述べる。まず、これら
の候補単語は標準パタン生成規則部110において、連
続音声認識に最適な音素に分割される。本実施例では、
語頭の音素とCV(子音中母音)、語中、語尾の音素を
VCV(母音半子音+母音)としている。
Here, for example, when you input "tokus imasi (Tokushima cloth)", "tOkusim" is selected as a candidate word.
asi f ukus imas 1 (Fukushima cloth)
We will describe the process when the three words ``hiroshimasi (Hiroshima cloth)'' are selected. First, these candidate words are divided into phonemes that are optimal for continuous speech recognition in the standard pattern generation rule unit 110. In the example,
The phoneme at the beginning of a word and the CV (vowel in a consonant), and the phoneme at the middle and end of the word are VCV (vowel semi-consonant + vowel).

次に、入力音声の特徴パラメータの長さを音素の数で割
り、1モーラ当たりの平均継続時間長を平均継続時間長
検出部152において求め、時間長の違い等により複数
種ある音素標準パタンの中から適した音素標準パタンを
選択する際に用いる。
Next, the length of the characteristic parameter of the input speech is divided by the number of phonemes, and the average duration length per mora is determined by the average duration detection unit 152. Used when selecting a suitable phoneme standard pattern from among them.

第6図(a)は候補単語として出力された単語を音素分
割処理部150において音素記号列に分割した例である
。第6図(C)は各音素との標準パタンか格納されてい
るメモリのアドレスとの対応表である。音素位置ラベル
付加部151は候補単語の音素位置に対応させて複数の
音素標準パタンの中から選択するところであるが、アド
レスの表にを二り、−Dよ、D、]とすると、D1は音
素の種類、D2は音素標準パタンの時間長、D。
FIG. 6(a) is an example in which a word output as a candidate word is divided into phoneme symbol strings by the phoneme division processing unit 150. FIG. 6(C) is a correspondence table between each phoneme and the address of the memory where the standard pattern is stored. The phoneme position label addition unit 151 selects from among a plurality of phoneme standard patterns in correspondence with the phoneme position of the candidate word. The type of phoneme, D2, is the duration of the standard phoneme pattern, D.

は音素標準パタンの複数の状況における種別であり、例
えば音素lalの標準パタンは、アドレス001−1か
ら入っている。また、アドレス001−1.1は、無声
化したlalの標準パタンか入っている。1aSa1の
ようなVCV型の音素は、アドレス931−1に入って
いる標準ものの他に、■CV全体が無声化した音(VC
V)が931−1.1に、VCVの中、CV音が無声化
した音(VCV)が931−1.2に、VCVの中、V
C音が無声化した音(VCV)が931−1.3に入っ
ている。また、これだけでなく1つの音素単位につき、
複数の標準パタンを持っている。
are types of phoneme standard patterns in a plurality of situations; for example, the standard pattern for the phoneme lal is entered from address 001-1. Further, the address 001-1.1 contains a standard pattern of devoiced lal. VCV-type phonemes such as 1aSa1 include the standard ones included in address 931-1, as well as ■VCV-type phonemes with the entire CV devoiced (VC
V) is in 931-1.1, in VCV, the sound (VCV) where CV sound is devoiced is in 931-1.2, in VCV, V
The sound (VCV) in which the C sound is devoiced is included in 931-1.3. In addition to this, for each phoneme unit,
It has multiple standard patterns.

第6図(b)は3つの候補単語の音素標準ノくタンの時
間長(D2)が1の時の音素を選択し、そのアドレスを
対応づけたものである。ここでは、「語頭・語尾は母音
部が無声化するパタンも含めて考える」という規則から
“tokusimasi”という単語は、第6図(b)
に示した音素のアドレスを使って第6図(d)に示す4
通りのパタンの組み合わせができる。
FIG. 6(b) shows the phoneme when the time length (D2) of the phoneme standard notation of three candidate words is 1, and its address is associated with the selected phoneme. Here, based on the rule that ``the beginning and end of a word should be considered including the pattern in which the vowel becomes devoiced,'' the word ``tokushimasi'' is chosen as shown in Figure 6 (b).
4 shown in Figure 6(d) using the address of the phoneme shown in Figure 6(d).
You can create combinations of street patterns.

ていないと接続できない。音素の標準パタンの種別、D
、により接続が可能な組み合わせを第6図(e)に示す
。この第6図(e)には、ある音素の標準パタンの時間
長D2と種別り、だけを示しである。例えば一番上の段
のb/bは、ある音素の標準パタンの、ある時間長(b
=とお()であり有声であるもの、b同志の接続を示す
。次の段のb/b、2はある音素の標準パタンの、ある
時間長(=bとおく)の有声であるものbと、ある音素
の標準パタンの、ある時間長(b−とおく)の前半が有
声音、後半が無声音のもの、b、2との音素の前半が等
しければ良い訳だから、第6図(e)にり、を示す必要
はなく、音素の標準パタンの時間長D2は1モ一ラ発声
時間長検出部152において1モーラ当たりの平均継続
時間長が求めであるので、これがbとなり、その単語内
では一定である。
You can't connect if you don't. Type of standard phoneme pattern, D
FIG. 6(e) shows combinations that can be connected by . FIG. 6(e) shows only the time length D2 and type of a standard pattern of a certain phoneme. For example, b/b in the top row is a certain time length (b
=to(), which is voiced, indicates the connection between b and comrades. The next row, b/b, 2, is a voiced standard pattern of a certain phoneme of a certain length of time (set as b), and a standard pattern of a certain phoneme of a certain length of time (set as b-). Since the first half of is a voiced sound and the second half is an unvoiced sound, and the first half of the phoneme b and 2 are equal, there is no need to show the time length D2 of the standard pattern of phonemes. Since the average duration per mora is determined by the mora utterance time length detection unit 152, this becomes b, which is constant within the word.

しかし、第6図(e)に示したのは音素結合規則の一部
であり、他に音声を発声する際の音響的な音素結合規則
も多くある。第6図(d)には、“t oku s i
ma s i”の組み合わせのみを示したが、同様にし
て他の候補単語についても組み合わせを作成する。音素
標準パタンの組み合わせができたら、音素標準パタン接
続部153において音素標準パタンを接続し、単語標準
パタンを作成する。接続の方法は、直接接続、線形補間
等があるが、音素0.P、Q、Rを接続する例を第5図
に示し、以下に説明する。
However, what is shown in FIG. 6(e) is only a part of the phoneme combination rules, and there are many other acoustic phoneme combination rules used when producing speech. In FIG. 6(d), “toku s i
Although only the combination of "ma s i" is shown, combinations are created for other candidate words in the same way.Once the combination of phoneme standard patterns is created, the phoneme standard patterns are connected in the phoneme standard pattern connection section 153, and the word A standard pattern is created.Connection methods include direct connection, linear interpolation, etc., and an example of connecting phonemes 0.P, Q, and R is shown in FIG. 5 and will be described below.

第7図の(a)は直接接続し、単語0PQRを生成する
例であり、(b)は音素0.P、Q、Rから補間部分と
して母音部分を数フレーム切り取ったものをQ′、P′
、Q′、R′とし、これの空白の部分を各次元のパラメ
ータの要素について線形補間しながら埋めていき、連続
した単語標準パタンを生成する例である。音素の補間方
法は、パラメータの性質によって適・不適があるので、
ここではパラメータに最適な補間法を用いる事にする。
(a) of FIG. 7 is an example of direct connection to generate the word 0PQR, and (b) is an example of the phoneme 0. The vowel parts are cut out from P, Q, and R by a few frames as interpolated parts, and are then used as Q' and P'.
, Q', and R', and the blank parts are filled in while linearly interpolating the parameter elements of each dimension to generate a continuous word standard pattern. Phoneme interpolation methods are suitable or unsuitable depending on the nature of the parameters, so
Here, we will use the interpolation method that is most suitable for the parameters.

最後に、音素標準パタン接続部153から出力された複
数の単語標準パタンと入力パタンをパタンマツチング部
111においてマツチングし、距離が最小となる単語を
結果出力部109より例えば音声として出力する。
Finally, the plurality of word standard patterns output from the phoneme standard pattern connection section 153 and the input pattern are matched in the pattern matching section 111, and the word with the minimum distance is outputted from the result output section 109 as, for example, speech.

パタンマツチング方式は、線形伸縮、DPマツチング法
法要多数るが、DPPマツチング良い結果が得られる。
Although there are many pattern matching methods such as linear expansion/contraction and DP matching, good results can be obtained with DPP matching.

ここで、距離尺度はマハラノビス距離等を代表とする統
計的な距離尺度を用いる。
Here, as the distance measure, a statistical distance measure such as Mahalanobis distance is used.

〔発明の効果〕〔Effect of the invention〕

以上説明した様に、第1段階において単語単位で連続マ
ハラノビスDPによるワードスポツティングを行うこと
により、候補単語の選出と音声区間の切り出しを同時に
行うことが可能となる。
As explained above, by performing word spotting by continuous Mahalanobis DP on a word-by-word basis in the first step, it is possible to simultaneously select candidate words and cut out speech sections.

第2段階として音素単位でマツチングを行うことにより
、2段階で認識を行う為に高い認識率が得られる。
By performing matching on a phoneme basis in the second stage, a high recognition rate can be obtained because recognition is performed in two stages.

また、複数の環境下における標準パタンを単語単位では
な(音素単位にしているため、情報量が小さくしてすむ
という効果がある。
In addition, since the standard patterns under multiple environments are made in units of phonemes (not words), the amount of information can be reduced.

また第2段階においては候補単語に対応する音素のみを
マツチングする為、時間がかからなくてすむという効果
がある。
Furthermore, in the second stage, only the phonemes corresponding to the candidate words are matched, so there is an advantage that it does not take much time.

【図面の簡単な説明】[Brief explanation of drawings]

第1図は本発明の第1の実施例の処理系の基本構成図、 第2図(a)は標準パタン作成の動作の流れを示すフロ
ーチャート、 第2図(b)は連続マハラノビスDPの様子を示す図、 第2図(C)は音素マツチングの様子を示す図、 第2図(d)は複数の候補単語と入力信号との関係を示
す図、 第2図(e)は発声時間長を2倍にした標準パタンの様
子を示す図、 第2図(f)は発声時間長の倍率変化による音素に対応
した倍率を示す図、 第2図(g)は第1図(f)の倍率に従って発声時間長
を3倍にした時の様子を示す図、第3図は従来の不特定
話者音声認識システムの構成図、 第4図は本発明の第2の音素認識処理の構成図、 第5図は標準パタン生成規則部の構成図、第6図(a)
は候補単語の音素分解の様子を示す図、 第6図(b)は候補単語の各音素の標準パタンのアドレ
スを示す図、 第6図(C)は音素標準パタンの種類によるアドレス例
を示す図、 第6図(d)は生成された標準パタンの組み合わせを示
す図、 第6図(e)は接続可能な標準パタンの組み合わせ例を
示す図、 第7図は補間方法を示す図である。 図中、1は音声入力装置、2は音声分析部、3は音素境
界検出部、4は音素識別部、5は音素標準パタン格納部
、6は単語識別部、7は単語辞書、8は修正規則部、9
は認識結果表示部、100は音声入力部、101は音声
分析部、102は連続マハラノビスDPによる距離計算
部、103は単語標準パタン格納部、104は候補単語
判別部、105はパラメータ格納部、106は音素標準
パタン格納部、107は連続マハラノビスDPによる距
離計算部、108は音素単位の認識結果による識別部、
109は結果出力部、110は標準パタン生成規則部、
111はパタンマツチング部、150は音素分割処理部
、151は音素ラベル付加部、152は1モ一ラ発声時
間長検出部、153は音素標準パタン接続部である。 第1図 処理系の基本構成 「】エコ]+09 第2図(a) 標準パターンの作成フロー 第2図(c) 音素マツチングの様子 第2図(d) 複数の候補単語と入力信号との関係 第2図(e) 発声時間長を2倍にした標準パタンの様子第2図(f) 発声時間長の倍率変化による音素に対応した倍率第2図
(9) 発声時間長を3倍にした時の様子 第4図 本発明第二の音素認識処理の構成図 第6図(a) 候補単語の音素分解の様子 第6図(b) 候補単語の各音素の標準パタンのアドレス第6図(d) 生成された標準パタンの組み合わせ 第6図(e) 接続可能な標準パタンの組み合わせ例 良−−−へへ−−レ (a) 補間力 −P −Q −R
Figure 1 is a basic configuration diagram of the processing system of the first embodiment of the present invention. Figure 2 (a) is a flowchart showing the flow of standard pattern creation operations. Figure 2 (b) is a continuous Mahalanobis DP. Figure 2 (C) is a diagram showing the state of phoneme matching, Figure 2 (d) is a diagram showing the relationship between multiple candidate words and input signals, and Figure 2 (e) is a diagram showing the utterance time length. Figure 2 (f) is a diagram showing the magnification corresponding to the phoneme due to the change in the magnification of the utterance duration, and Figure 2 (g) is the standard pattern that is doubled. A diagram showing what happens when the utterance time length is tripled according to the multiplication factor. Figure 3 is a configuration diagram of a conventional speaker-independent speech recognition system. Figure 4 is a configuration diagram of the second phoneme recognition process of the present invention. , Figure 5 is a configuration diagram of the standard pattern generation rule section, Figure 6 (a)
Figure 6(b) is a diagram showing the phoneme decomposition of a candidate word, Figure 6(b) is a diagram showing the addresses of standard patterns for each phoneme of a candidate word, and Figure 6(C) is an example of addresses depending on the type of phoneme standard pattern. Figure 6(d) is a diagram showing a combination of generated standard patterns, Figure 6(e) is a diagram showing an example of a combination of standard patterns that can be connected, and Figure 7 is a diagram showing an interpolation method. . In the figure, 1 is a speech input device, 2 is a speech analysis section, 3 is a phoneme boundary detection section, 4 is a phoneme identification section, 5 is a phoneme standard pattern storage section, 6 is a word identification section, 7 is a word dictionary, and 8 is a correction section. Rules section, 9
100 is a recognition result display section, 100 is a speech input section, 101 is a speech analysis section, 102 is a distance calculation section using continuous Mahalanobis DP, 103 is a word standard pattern storage section, 104 is a candidate word discrimination section, 105 is a parameter storage section, 106 is a phoneme standard pattern storage unit, 107 is a distance calculation unit using continuous Mahalanobis DP, 108 is an identification unit based on the recognition result of each phoneme,
109 is a result output section, 110 is a standard pattern generation rule section,
111 is a pattern matching section, 150 is a phoneme division processing section, 151 is a phoneme label adding section, 152 is a monomolar utterance time length detection section, and 153 is a phoneme standard pattern connection section. Figure 1 Basic configuration of the processing system ``]Eco]+09 Figure 2 (a) Standard pattern creation flow Figure 2 (c) Phoneme matching Figure 2 (d) Relationship between multiple candidate words and input signals Figure 2 (e) Standard pattern with twice the utterance time length Figure 2 (f) Magnification corresponding to phoneme due to change in utterance time length Figure 2 (9) utterance time length tripled Fig. 4 Configuration diagram of the second phoneme recognition process of the present invention Fig. 6 (a) Situation of phoneme decomposition of a candidate word Fig. 6 (b) Address of the standard pattern of each phoneme of the candidate word Fig. 6 ( d) Combination of generated standard patterns Figure 6 (e) Example of combination of standard patterns that can be connected Good ---Hehe-- (a) Interpolation force -P -Q -R

Claims (6)

【特許請求の範囲】[Claims] (1)入力音声を分析して特徴ベクトルの時系列を求め
る音声分析手段、 複数の音声サンプルから得た単語標準パタンを格納する
単語標準パタン格納手段、 前記入力音声特徴ベクトル時系列にスポツ テイング法を用いることにより音声区間を検出し、前記
単語標準パタンの中から候補単語を選出する候補単語識
別手段、 複数の音声サンプルから得た音素標準パタンを格納する
音素標準パタン格納手段、 前記音声区間において前記入力音声の特徴ベクトルの時
系列と前記候補単語の前記音素標準パタンとのマッチン
グを行うことにより前記入力音声を認識する認識手段、 前記認識手段により認識した結果を出力する出力手段を
有することを特徴とする音声認識装置。
(1) Speech analysis means for analyzing input speech to obtain a time series of feature vectors, word standard pattern storage means for storing word standard patterns obtained from a plurality of speech samples, and applying a spotting method to the input speech feature vector time series. candidate word identification means for detecting a speech interval and selecting a candidate word from the word standard patterns by using a phoneme standard pattern storage means for storing a phoneme standard pattern obtained from a plurality of speech samples; A recognition means for recognizing the input speech by matching a time series of feature vectors of the input speech with the phoneme standard pattern of the candidate word, and an output means for outputting a result recognized by the recognition means. speech recognition device.
(2)前記候補単語識別手段は更に統計的な距離尺度、
マハラノビス距離を用いて連続DPを行い、DPパスの
累積距離を計算する距離計算手段、 前記DPパスを記憶する記憶手段、 前記累積距離が予め設定した閾値より小さ く、かつ最小である時点を終端とする前記DPパスを前
記記憶手段より呼び出し、 該DPパスの始端を求め、音声区間を認識する音声区間
認識手段を含むことを特徴とする請求項(1)に記載の
音声認識装置。
(2) The candidate word identification means further includes a statistical distance measure;
Distance calculation means that performs continuous DP using Mahalanobis distance and calculates the cumulative distance of DP paths; Storage means that stores the DP path; A point in time when the cumulative distance is smaller than a preset threshold and is minimum is determined as the end point. 2. The speech recognition apparatus according to claim 1, further comprising speech segment recognition means for reading the DP path from the storage means, determining the starting end of the DP path, and recognizing the speech segment.
(3)前記入力音声とのマッチングは、前記候補単語に
対応する前記音素標準パタンを前記音声区間においてス
ポツテイング法を用いて行うことを特徴とする請求項(
1)に記載の音声認識装置。
(3) The matching with the input speech is performed using a spotting method of the phoneme standard pattern corresponding to the candidate word in the speech section (
1) The speech recognition device according to item 1).
(4)前記入力音声とのマッチングは、標準パタン生成
規則手段によって前記候補単語の音素列に従って前記音
素標準パタンを接続して生成した標準パタンと行うこと
を特徴とする請求項(1)に記載の音声認識装置。
(4) The matching with the input speech is performed with a standard pattern generated by connecting the phoneme standard pattern according to the phoneme string of the candidate word by a standard pattern generation rule means. speech recognition device.
(5)前記音素標準パタン格納手段に格納する音素の単
位は、CV(子音−母音)、VCV(母音−子音−母音
)、VV(母音−母音)を用いることを特徴とする請求
項(1)に記載の音声認識装置。
(5) The phoneme units stored in the phoneme standard pattern storage means are CV (consonant-vowel), VCV (vowel-consonant-vowel), and VV (vowel-vowel). ).The speech recognition device described in ).
(6)前記音素標準パタンは、話者、発声時間、発声環
境の要因による複数の標準パタンを持つことを特徴とす
る請求項(1)に記載の音声認識装置。
(6) The speech recognition device according to claim 1, wherein the phoneme standard pattern includes a plurality of standard patterns depending on factors such as a speaker, a speaking time, and a speaking environment.
JP2023205A 1990-02-01 1990-02-01 Voice recognition device Expired - Fee Related JP2862306B2 (en)

Priority Applications (2)

Application Number Priority Date Filing Date Title
JP2023205A JP2862306B2 (en) 1990-02-01 1990-02-01 Voice recognition device
US08/194,807 US6236964B1 (en) 1990-02-01 1994-02-14 Speech recognition apparatus and method for matching inputted speech and a word generated from stored referenced phoneme data

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
JP2023205A JP2862306B2 (en) 1990-02-01 1990-02-01 Voice recognition device

Publications (2)

Publication Number Publication Date
JPH03228100A true JPH03228100A (en) 1991-10-09
JP2862306B2 JP2862306B2 (en) 1999-03-03

Family

ID=12104167

Family Applications (1)

Application Number Title Priority Date Filing Date
JP2023205A Expired - Fee Related JP2862306B2 (en) 1990-02-01 1990-02-01 Voice recognition device

Country Status (1)

Country Link
JP (1) JP2862306B2 (en)

Cited By (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2016177046A (en) * 2015-03-19 2016-10-06 株式会社レイトロン Speech recognition apparatus and speech recognition program

Citations (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPS60121499A (en) * 1983-12-05 1985-06-28 富士通株式会社 Voice collation system
JPS63165900A (en) * 1986-12-27 1988-07-09 沖電気工業株式会社 Conversation voice recognition system

Patent Citations (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPS60121499A (en) * 1983-12-05 1985-06-28 富士通株式会社 Voice collation system
JPS63165900A (en) * 1986-12-27 1988-07-09 沖電気工業株式会社 Conversation voice recognition system

Cited By (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2016177046A (en) * 2015-03-19 2016-10-06 株式会社レイトロン Speech recognition apparatus and speech recognition program

Also Published As

Publication number Publication date
JP2862306B2 (en) 1999-03-03

Similar Documents

Publication Publication Date Title
Loizou et al. High-performance alphabet recognition
EP2048655B1 (en) Context sensitive multi-stage speech recognition
JPH0772840B2 (en) Speech model configuration method, speech recognition method, speech recognition device, and speech model training method
Hasija et al. Recognition of children Punjabi speech using tonal non-tonal classifier
JP5315976B2 (en) Speech recognition apparatus, speech recognition method, and program
AU2019202146B2 (en) System and method for outlier identification to remove poor alignments in speech synthesis
Yavuz et al. A phoneme-based approach for eliminating out-of-vocabulary problem of Turkish speech recognition using Hidden Markov Model.
Unnibhavi et al. LPC based speech recognition for Kannada vowels
JP2001312293A (en) Voice recognition method and apparatus, and computer-readable storage medium
Manjunath et al. Articulatory and excitation source features for speech recognition in read, extempore and conversation modes
JP5300000B2 (en) Articulation feature extraction device, articulation feature extraction method, and articulation feature extraction program
JP2862306B2 (en) Voice recognition device
Mebarkia et al. Maghrebian accent recognition using svm classifier and mfcc features
JP2943445B2 (en) Voice recognition method
JP3277522B2 (en) Voice recognition method
Ganesh et al. Syllable based continuous speech recognizer with varied length maximum likelihood character segmentation
JP2001005483A (en) Word voice recognizing method and word voice recognition device
Manjunath et al. Two-stage phone recognition system using articulatory and spectral features
JPH05303391A (en) Speech recognition device
Pangsatabam et al. Refining Tokenization Methods to Advance Low-Resource Manipuri Speech Recognition
Manjunath et al. Improvement of phone recognition accuracy using source and system features
Resch et al. Time synchronization of speech.
JP2766393B2 (en) Voice recognition method
JPH04233599A (en) Method and device speech recognition
JPH08166798A (en) Phoneme dictionary creating apparatus and method

Legal Events

Date Code Title Description
LAPS Cancellation because of no payment of annual fees