JPH0194397A - Voice recognition system - Google Patents

Voice recognition system

Info

Publication number
JPH0194397A
JPH0194397A JP62252109A JP25210987A JPH0194397A JP H0194397 A JPH0194397 A JP H0194397A JP 62252109 A JP62252109 A JP 62252109A JP 25210987 A JP25210987 A JP 25210987A JP H0194397 A JPH0194397 A JP H0194397A
Authority
JP
Japan
Prior art keywords
dictionary
pattern
axis direction
orthogonalized
orthogonalization
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Granted
Application number
JP62252109A
Other languages
Japanese (ja)
Other versions
JP2514986B2 (en
Inventor
Tsuneo Nitta
恒雄 新田
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Toshiba Corp
Original Assignee
Toshiba Corp
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Toshiba Corp filed Critical Toshiba Corp
Priority to JP62252109A priority Critical patent/JP2514986B2/en
Priority to DE3888777T priority patent/DE3888777T2/en
Priority to EP88116414A priority patent/EP0311022B1/en
Priority to US07/254,110 priority patent/US5001760A/en
Priority to KR1019880013005A priority patent/KR910007530B1/en
Publication of JPH0194397A publication Critical patent/JPH0194397A/en
Priority to SG123594A priority patent/SG123594G/en
Priority claimed from SG123594A external-priority patent/SG123594G/en
Priority to HK110794A priority patent/HK110794A/en
Application granted granted Critical
Publication of JP2514986B2 publication Critical patent/JP2514986B2/en
Anticipated expiration legal-status Critical
Expired - Lifetime legal-status Critical Current

Links

Classifications

    • C—CHEMISTRY; METALLURGY
    • C12—BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12N—MICROORGANISMS OR ENZYMES; COMPOSITIONS THEREOF; PROPAGATING, PRESERVING, OR MAINTAINING MICROORGANISMS; MUTATION OR GENETIC ENGINEERING; CULTURE MEDIA
    • C12N9/00—Enzymes; Proenzymes; Compositions thereof; Processes for preparing, activating, inhibiting, separating or purifying enzymes
    • C12N9/10—Transferases (2.)
    • C12N9/1048—Glycosyltransferases (2.4)
    • C12N9/1051—Hexosyltransferases (2.4.1)
    • C12N9/1055—Levansucrase (2.4.1.10)
    • C—CHEMISTRY; METALLURGY
    • C12—BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12N—MICROORGANISMS OR ENZYMES; COMPOSITIONS THEREOF; PROPAGATING, PRESERVING, OR MAINTAINING MICROORGANISMS; MUTATION OR GENETIC ENGINEERING; CULTURE MEDIA
    • C12N15/00—Mutation or genetic engineering; DNA or RNA concerning genetic engineering, vectors, e.g. plasmids, or their isolation, preparation or purification; Use of hosts therefor
    • C12N15/09—Recombinant DNA-technology
    • C12N15/63—Introduction of foreign genetic material using vectors; Vectors; Use of hosts therefor; Regulation of expression
    • C12N15/74—Vectors or expression systems specially adapted for prokaryotic hosts other than E. coli, e.g. Lactobacillus, Micromonospora
    • C—CHEMISTRY; METALLURGY
    • C12—BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12N—MICROORGANISMS OR ENZYMES; COMPOSITIONS THEREOF; PROPAGATING, PRESERVING, OR MAINTAINING MICROORGANISMS; MUTATION OR GENETIC ENGINEERING; CULTURE MEDIA
    • C12N15/00—Mutation or genetic engineering; DNA or RNA concerning genetic engineering, vectors, e.g. plasmids, or their isolation, preparation or purification; Use of hosts therefor
    • C12N15/09—Recombinant DNA-technology
    • C12N15/87—Introduction of foreign genetic material using processes not otherwise provided for, e.g. co-transformation
    • C12N15/90—Stable introduction of foreign DNA into chromosome

Landscapes

  • Health & Medical Sciences (AREA)
  • Genetics & Genomics (AREA)
  • Life Sciences & Earth Sciences (AREA)
  • Engineering & Computer Science (AREA)
  • Chemical & Material Sciences (AREA)
  • Organic Chemistry (AREA)
  • Bioinformatics & Cheminformatics (AREA)
  • Wood Science & Technology (AREA)
  • Zoology (AREA)
  • Biomedical Technology (AREA)
  • Biotechnology (AREA)
  • General Engineering & Computer Science (AREA)
  • Molecular Biology (AREA)
  • Microbiology (AREA)
  • Biochemistry (AREA)
  • General Health & Medical Sciences (AREA)
  • Physics & Mathematics (AREA)
  • Biophysics (AREA)
  • Plant Pathology (AREA)
  • Medicinal Chemistry (AREA)
  • Mycology (AREA)

Abstract

(57)【要約】本公報は電子出願前の出願データであるた
め要約のデータは記録されません。
(57) [Summary] This bulletin contains application data before electronic filing, so abstract data is not recorded.

Description

【発明の詳細な説明】 [発明の目的] (産業上の利用分野) 本発明は少ない学習パターンで高い認識性能を得ること
のできる音声認識方式に関する。
DETAILED DESCRIPTION OF THE INVENTION [Object of the Invention] (Industrial Application Field) The present invention relates to a speech recognition method that can obtain high recognition performance with a small number of learning patterns.

(従来の技術) 音声により情報の入出力は人間にとって自然性が高く、
マン・マシン・インターフェースとして優れており、従
来より種々研究されている。現在、実用化されている音
声認識装置の殆んどは単語音声を認識する方式のもので
、−膜内には第2図に示すように構成されている。
(Conventional technology) Inputting and outputting information through voice is highly natural for humans;
It is an excellent man-machine interface and has been studied in various ways. Most of the speech recognition devices currently in practical use are of the type that recognizes word speech, and the membrane is constructed as shown in FIG.

この装置は、発声入力された音声を電気信号に変換して
取込み、バンド・パス・フィルタ等からなる音響分析部
1にて音響分析し、始端・終端検出部2にてその単語音
声区間を検出する。そして入力音声の上記単語音声区間
の音響分析データ(特徴情報;音声パターン)と、標準
パターン辞書3に予め登録されている認識対象単語の各
標準パターンとの類似度や距離等をパターン・マツチン
グ部4にて計算し、その計算結果を判定部5にて判定し
て、例えば類似度値の最も高い標準パターンのカテゴリ
名を前記入力音声に対する認識結果として求めるものと
なっている。
This device converts the input voice into an electrical signal and captures it, acoustically analyzes it in an acoustic analysis section 1 consisting of a band pass filter, etc., and detects the word speech section in a start/end detection section 2. do. Then, a pattern matching unit calculates the similarity and distance between the acoustic analysis data (feature information; speech pattern) of the word speech section of the input speech and each standard pattern of recognition target words registered in advance in the standard pattern dictionary 3. 4, and the result of the calculation is determined by the determining unit 5, to obtain, for example, the category name of the standard pattern with the highest similarity value as the recognition result for the input voice.

しかしこのようにパターン・マツチング法による音声認
識では入力音声パターンと予め登録されている標準パタ
ーンとの時間軸方向のずれ(パターン変形)が問題とな
る。そこで従来では、専ら線形伸縮や、動的計画法(D
P)に代表される非線形伸縮等により、上述した時間軸
方向のずれに対する課題を解消している。
However, in speech recognition using the pattern matching method, a problem arises in that the input speech pattern and the standard pattern registered in advance are misaligned in the time axis direction (pattern deformation). Therefore, in the past, only linear expansion and contraction and dynamic programming (D
The problem with the shift in the time axis direction mentioned above is solved by nonlinear expansion and contraction represented by P).

一方、このようなパターン・マツチング法とは別に、予
め収集された学習パター、ンから直交化辞書を作成し、
この直交化辞書を用いて音声認識する方式(部分空間法
)が提唱されている。この方式は第3図にその構成例を
示すように、音響分析されて音声区間検出された音声パ
ターンから、標本点抽出部6にて上記音声区間を等分割
した所定点数の標本点を抽出し、(特徴ベクトルの数×
標本点数)で示される標本パターンを求める。このよう
な標本パターンを認識対象とするカテゴリ毎に所定数ず
つ収集してパターン蓄積部7に格納する。そしてグラム
・シュミット(GS)直交化部8において、上記パター
ン蓄積部7に収集された所定数(3個以上)の標本パタ
ーンを用いて以下に示す手順で直交化辞書9を作成する
。
On the other hand, apart from such a pattern matching method, an orthogonalized dictionary is created from learning patterns collected in advance.
A speech recognition method (subspace method) using this orthogonalized dictionary has been proposed. As shown in FIG. 3, an example of the configuration of this method is such that a sample point extraction unit 6 extracts a predetermined number of sample points obtained by equally dividing the speech section from a speech pattern that has been acoustically analyzed and detected a speech section. , (number of feature vectors ×
Find the sample pattern indicated by the number of sample points). A predetermined number of such sample patterns are collected for each category to be recognized and stored in the pattern storage section 7. Then, in the Gram-Schmidt (GS) orthogonalization unit 8, an orthogonalization dictionary 9 is created using the predetermined number (three or more) of sample patterns collected in the pattern storage unit 7 in the following procedure.

即ち、上記直交化辞書9の作成は、各カテゴリ毎にその
カテゴリのm回目の学習パターンをaoとし、3回発声
された学習パターンを用いる場合には、 ■ 1回目の学習データa1を第1軸の辞書b1とし、 b 1 ”” a t               
 ・・・(1)これを直交化辞書9に登録する。
That is, in creating the orthogonalized dictionary 9, for each category, the m-th learning pattern of that category is set as ao, and when a learning pattern uttered three times is used, ■ the first learning data a1 is set as the first learning pattern. Let the axis dictionary b1 be b 1 ”” a t
(1) Register this in the orthogonalization dictionary 9.

■ 2回目の学習データa2からグラム・シュミットの
直交化式を用い、 なる計算を行い、l1b211が一定値より大きい場合
、これを第2軸の辞書b2として前記直交化辞書9に登
録する。但し、(・)は内積、Tは転置、1111はノ
ルムを示す。
(2) Using the Gram-Schmidt orthogonalization formula from the second learning data a2, perform the following calculation, and if l1b211 is larger than a certain value, register this in the orthogonalization dictionary 9 as the second axis dictionary b2. However, (.) indicates the inner product, T indicates the transposition, and 1111 indicates the norm.

■ そして3回目の学習データa3から、なる計算を行
い、11b311が一定値より大きい場合、これを第3
軸の辞書b3として前記直交化辞書9に登録する。但し
、第2軸の辞書が求められていない場合には、上記(2
)式の計算を行う。
■ Then, from the third learning data a3, perform the following calculation, and if 11b311 is larger than a certain value, use this as the third
It is registered in the orthogonalization dictionary 9 as the axis dictionary b3. However, if the second axis dictionary is not required, the above (2)
) calculates the formula.

以上の■〜■の処理を各カテゴリについて繰返し実行し
て直交化辞書9を予め形成しておく。
The orthogonalized dictionary 9 is formed in advance by repeatedly performing the above processes (1) to (2) for each category.

類似度計算部10は上述した如く作成された直交化辞書
9と、入力音声パターンXとの間でとして、カテゴリi
の直交化辞書b  との間の1、r 類似度を計算するもので、この類似度値に従って上記入
力音声パターンXが認識される。尚、上記カテゴリiの
直交化辞書b  は予め正規化され1、r たちのであり、K1はカテゴリiの辞書の個数(軸数)
を示している。
The similarity calculation unit 10 calculates the category i between the orthogonalized dictionary 9 created as described above and the input speech pattern
, and the input speech pattern X is recognized according to this similarity value. Note that the above orthogonalized dictionary b of category i is normalized in advance to 1, r, and K1 is the number of dictionaries (number of axes) of category i.
It shows.

ところがこのようなGS直交化を用いる方式にあっては
、上述した各直交軸が担うパターン変動量が明確でない
と云う問題がある。この為、上述した如くして計算され
た直交化辞書9のカテゴリiの標本パターン(a   
r  a  +  a  lが、i、l   i、2 
 1.3 そのカテゴリiの本来の標準的なパターンを良く表現し
ているとは同等保障されないと云う不具合がある。
However, in a method using such GS orthogonalization, there is a problem that the amount of pattern variation carried by each of the above-mentioned orthogonal axes is not clear. For this reason, the sample pattern (a
r a + a l is i, l i, 2
1.3 There is a problem that it is not guaranteed that the original standard pattern of the category i is expressed well.

(発明が解決しようとする問題点) このように従来のGS直交化を用いた部分空間法による
音声認識にあっては、直交化された辞書自体に、例えば
収集した学習パターンの時間軸方向や周波数軸方向の変
動に起因する問題があり、その標準パターンを良く表現
しているか否かと云う点で課題が残されている。またこ
のような問題を解消するには、相当大量の学習パターン
を収集する必要がある等の不具合がある。
(Problems to be Solved by the Invention) In speech recognition using the conventional subspace method using GS orthogonalization, the orthogonalized dictionary itself contains, for example, the time axis direction of the collected learning patterns. There are problems caused by fluctuations in the direction of the frequency axis, and an issue remains as to whether the standard pattern is well represented. In addition, in order to solve such problems, there is a problem that it is necessary to collect a considerable amount of learning patterns.

本発明はこのような事情を考慮してなされたもので、そ
の目的とするところは、少ない学習パターンにてその標
準パターンを良く表現した、パターン変動に十分対処す
ることのできる直交化辞書を作成し、認識性能の向上を
図ることのできる音声認識方式を提供することにある。
The present invention was made in consideration of these circumstances, and its purpose is to create an orthogonalized dictionary that can adequately represent the standard pattern with a small number of learning patterns and can adequately cope with pattern variations. The object of the present invention is to provide a speech recognition method that can improve recognition performance.

[発明の構成] (問題点を解決するための手段) 本発明は入力音声を分析処理して求められる入力音声パ
ターンと予め収集された学習パターンに基いて作成され
ている直交化辞書との間で類似度を計算して上記入力音
声を認識する音声認識方式において、 予め収集された学習パターンに対して少なくとも平滑処
理と微分処理とを施す3種以上のフィルタを用い、例え
ば収集された学習パターンの平均パターンを求め、この
平均パターンを時間軸方向および周波数軸方向にそれぞ
れ平滑化して第1軸の辞書を求め、更に上記平均パター
ンを時間軸方向に微分して第2軸の辞書を求めると共に
、上記平均パターンを周波数軸方向に微分して第3軸の
辞書を求める等して前記直交化辞書を作成し、更に、例
えばグラムシュミットの直交化等によって上記直交化辞
書に直交する付加辞書を作成し、この付加辞書を上記直
交化辞書に付加することを特徴とするものである。
[Structure of the Invention] (Means for Solving the Problems) The present invention provides a method for solving problems between an input speech pattern obtained by analyzing input speech and an orthogonalized dictionary created based on learning patterns collected in advance. In the speech recognition method that recognizes the input speech by calculating the similarity with Find the average pattern of , smooth this average pattern in the time axis direction and the frequency axis direction to find the first axis dictionary, further differentiate the above average pattern in the time axis direction to find the second axis dictionary, and , create the orthogonalized dictionary by differentiating the average pattern in the frequency axis direction to obtain a third axis dictionary, and further create an additional dictionary orthogonal to the orthogonalized dictionary by, for example, Gram-Schmidt orthogonalization. The additional dictionary is created and added to the orthogonalized dictionary.

(作用) 3種以上のフィルタを用いて収集された学習パターンの
平均パターンを求め、この平均パターンを時間軸方向お
よび周波数軸方向にそれぞれ平滑化して第1軸の辞書を
求めるので音声パターンの時間軸方向の変動を、および
周波数軸方向の変動を効果的に吸収することができる。
(Operation) The average pattern of the learning patterns collected using three or more types of filters is obtained, and this average pattern is smoothed in the time axis direction and the frequency axis direction to obtain the first axis dictionary. It is possible to effectively absorb axial fluctuations and frequency axial fluctuations.

更には上記平均パターンを時間軸方向に微分して第2軸
の辞書を求めるので時間軸方向に対する音声パターンの
位置ずれを効果的に吸収することができ、また上記平均
パターンを周波数軸方向に微分して第3軸の辞書を求め
るので周波数軸方向に対する音声パターンの位置ずれを
効果的に吸収することができる。
Furthermore, since the average pattern is differentiated in the time axis direction to obtain a dictionary on the second axis, it is possible to effectively absorb the positional shift of the audio pattern in the time axis direction, and the above average pattern is differentiated in the frequency axis direction. Since the dictionary for the third axis is obtained by using the above method, it is possible to effectively absorb the positional deviation of the voice pattern in the direction of the frequency axis.

このようにして時間軸方向および周波数軸方向に対する
パターン変動をそれぞれ吸収した直交化辞書が作成され
るので、直交化辞書の各辞書パターンをその変動による
位置ずれに対応し得るものとすることができ、認識性能
の向上に大きく寄与する。しかも時間軸方向および周波
数軸方向のパターン変動を吸収した平均パターンから生
成される辞書パターン(第1軸)をベースとして第2軸
および第3軸の辞書を求めてその直交化辞書が生成され
ていくので、従来のように直交化辞書自体の各直交軸が
担うパターン変動量が不明確になることがなく、少ない
学習パターンを有効に用いて性能の高い直交化辞書を効
果的に作成することが可能となる。
In this way, an orthogonalized dictionary is created that absorbs pattern fluctuations in the time axis direction and frequency axis direction, so each dictionary pattern in the orthogonalized dictionary can be made to be able to cope with positional shifts due to the fluctuations. , greatly contributes to improving recognition performance. Furthermore, dictionaries for the second and third axes are obtained based on a dictionary pattern (first axis) generated from an average pattern that absorbs pattern fluctuations in the time axis direction and frequency axis direction, and the orthogonalized dictionary is generated. Therefore, unlike in the past, the amount of pattern variation carried by each orthogonal axis of the orthogonal dictionary itself does not become unclear, and it is possible to effectively create a high-performance orthogonal dictionary by effectively using a small number of learning patterns. becomes possible.

更には上記直交化辞書に直交する付加辞書が作成されて
上記直交化辞書に付加されているので、この付加辞書に
て上述した時間軸方向および周波数軸方向以外のパター
ン変動をも効果的に吸収して認識処理を行わせることが
可能となり、その認識性能の向上に大きく寄与する。
Furthermore, since an additional dictionary orthogonal to the orthogonalized dictionary is created and added to the orthogonalized dictionary, this additional dictionary can effectively absorb pattern fluctuations other than those in the time axis direction and frequency axis direction. This makes it possible to perform recognition processing using the same method, which greatly contributes to improving recognition performance.

(実施例) 以下、図面を参照して本発明の一実施例につき説明する
。
(Example) Hereinafter, an example of the present invention will be described with reference to the drawings.

第1図は本発明に係る一実施例方式を適用して構成され
る音声認識装置の概略構成図で、第3図に示した従来装
置と同一部分には同一符号を付して示しである。
FIG. 1 is a schematic configuration diagram of a speech recognition device constructed by applying an embodiment method according to the present invention, and the same parts as those of the conventional device shown in FIG. 3 are denoted by the same reference numerals. .

この実施例装置が特徴とするところは、パターン蓄積部
7に蓄積された学習パターンを用いて直交化辞書9を作
成する手段として、従来のGS直交化部8に代えて少な
くとも平滑処理と微分処理とを実行する3種以上のフィ
ルタ、例えば直交化時間・周波数フィルタからなる直交
化時間・周波数フィルタ部11を用いた点にある。そし
て更には、例えば上記GS直交化部8を用いて、上記直
交化時間・周波数フィルタ部11にて作成された直交化
辞書に直交する辞書を付加辞書として作成し、この付加
辞書を上記直交化辞書9に付加するようにしたことを特
徴としている。
The feature of this embodiment device is that, as a means for creating an orthogonalization dictionary 9 using the learning patterns accumulated in the pattern storage section 7, at least smoothing processing and differential processing are performed instead of the conventional GS orthogonalization section 8. The present invention uses an orthogonal time/frequency filter section 11 that is composed of three or more types of filters, for example, orthogonal time/frequency filters. Furthermore, for example, using the GS orthogonalization section 8, a dictionary orthogonal to the orthogonalized dictionary created in the orthogonalization time/frequency filter section 11 is created as an additional dictionary, and this additional dictionary is used for the orthogonalization. The feature is that it is added to the dictionary 9.

尚、ここではパターン蓄積部7に収集される学習パター
ンとしては、例えばj  (−1,2,〜1B)で示さ
れる16点の音響分析されな特徴ベクトルからなり、そ
の音声区間をk (−0,4,2,〜17)として17
等分する18個の標本点に亙って採取したデータ系列と
して与えられるものとして説明する。
Here, the learning pattern collected in the pattern storage section 7 is composed of a feature vector of 16 points indicated by j (-1, 2, ~1B), which has not been subjected to acoustic analysis, and its speech interval is defined as k (- 0,4,2,~17) as 17
The explanation will be given assuming that it is given as a data series collected over 18 equally divided sampling points.

しかして前記直交化時間・周波数フィルタ部11は、カ
テゴリiについて、例えば3個ずつ収集されたm番目の
学習パターンをaa+(j、k)としたとき、灰のよう
にして直交化辞書9を作成している。
Therefore, the orthogonalization time/frequency filter unit 11 converts the orthogonalization dictionary 9 into gray, for example, when the m-th learning pattern collected in groups of three is aa+(j, k) for the category i. Creating.

■ 先ず、カテゴリiの学習パターンam(j、k)か
ら、その平均パターンAU、k)を N−0,1,2,〜15.  k−0,1,2,〜1フ
]として求める。
■ First, from the learning pattern am(j,k) of category i, its average pattern AU,k) is N-0,1,2,~15. k-0,1,2,~1f].

■ しかる後、上述した如くして求めた平均パターンA
   を用いて、 (Lk) bl(j、k) +A     +A ” A(j−1,に−1)    U−1,k)   
 (j−1,に+1)+ 2*A    +A ” A(j、に−1)     (j、k)    (
j、に+1)+A     +A +A(j+1.に−1)    (j+1.k)   
 (j+1.に+1)[j−1,2,〜 14.   
k−1,2,〜 1θコ             ・
・・(6)なる演算にて第1軸の辞書b1N、k)を求
め、これを直交化辞書9に登録する。この辞書bl(j
、k)は前記平均パターンA(j、k)を時間軸方向お
よび周波数軸方向にそれぞれ平滑化したものとして求め
られ、直交化辞書9の基準となる第1軸の辞書データと
して登録される。
■ After that, the average pattern A obtained as described above
Using (Lk) bl(j,k) +A +A ” A(j-1, to-1) U-1,k)
(j-1, +1) + 2*A +A ” A (j, -1) (j, k) (
+1 to j) +A +A +A (-1 to j+1.) (j+1.k)
(+1 to j+1.) [j-1, 2, ~ 14.
k-1, 2, ~ 1θko・
The dictionary b1N,k) of the first axis is obtained by the calculation (6) and is registered in the orthogonalization dictionary 9. This dictionary bl(j
, k) are obtained by smoothing the average pattern A(j, k) in the time axis direction and the frequency axis direction, respectively, and are registered as first axis dictionary data serving as a reference for the orthogonalized dictionary 9.

■ しかる後、前記平均パターンA(j、k)を用い、
−−A           +A b2(j、k)    (j−1,に−1)   (j
−1,に+1)” =A(j、k()   (j、に+
1) ’十 A ” ” U+1.に−1)   (j+1.に+1) 
’+ A [j=1.2.〜14.  k−1,2,〜1B]  
      ・・・(7)なる演算にて第2軸の辞書b
2(j、k)を求め、これを正′規化した後、前記直交
化辞書9に登録する。
■ After that, using the average pattern A(j, k),
−-A +A b2(j, k) (j-1, to-1) (j
−1, +1)” = A(j, k() (j, +1)” = A(j, k() (j, +
1) 'ten A ” ” U+1. −1) (+1 to j+1.)
'+ A [j=1.2. ~14. k-1, 2, ~1B]
...(7) With the calculation, the dictionary b of the second axis
2(j,k) is obtained, normalized, and then registered in the orthogonalization dictionary 9.

この辞書b2(j、k)は前記平均パターンA(j、k
)を時間軸方向に微分したものとして求められる。
This dictionary b2(j,k) is the average pattern A(j,k
) is obtained by differentiating it along the time axis.

尚、このようにして計算される第2軸の辞書b2(j、
k)は、前記第1軸の辞書b1(j、k)に対して完全
には直交していないことから、必要に応じてB2(j、
k)″b2(j、k) ’ ” 2(j、k)   1(j、k))51(j、
k)・ b なる再直交化処理を施し、この再直交化された辞書デー
タB2(j、k)を正規化した後、新たな第2軸の辞書
b2(j、k)として前記直交化辞書9に登録するよう
にしても良い。しかし、このような再直交化を行わなく
ても、上述した如く求められる第2軸の辞書b2(j、
k)にて十分なる認識性能を得ることが可能である。
Note that the second axis dictionary b2(j,
k) is not completely orthogonal to the first axis dictionary b1(j, k), so B2(j,
k)″b2(j,k)′ ” 2(j,k) 1(j,k))51(j,
After performing re-orthogonalization processing such as k) and b and normalizing this re-orthogonalized dictionary data B2 (j, k), the orthogonalized dictionary is used as a new second axis dictionary b2 (j, k). 9 may be registered. However, even without such re-orthogonalization, the second axis dictionary b2(j,
It is possible to obtain sufficient recognition performance with k).

■ また前記平均パターンAU、k)を用い、卿−A 
         −A b3(j、k)    (j−1,に−1)   (j
−1,k)+A −A(j−1,に+1)   (j+1.に−1)+ 
A +A(j+1.k)      (j+1.に+1)[
j−1,2,〜14.  k−1,2,〜16]   
     ・・・(8)なる演算にて第3軸の辞書b3
(j、k)を求め、これを正規化した後、前記直交化辞
書9に登録する。
■ Also, using the average pattern AU, k),
-A b3 (j, k) (j-1, to -1) (j
-1, k) + A -A (+1 to j-1) (-1 to j+1.) +
A +A(j+1.k) (+1 to j+1.) [
j-1, 2, ~14. k-1, 2, ~16]
... (8) Dictionary b3 of the third axis with the calculation
After finding (j, k) and normalizing it, it is registered in the orthogonalization dictionary 9.

この辞書b3(j、k)は前記平均パターンA(j、k
)を周波数軸方向に微分したものとして求められる。
This dictionary b3(j,k) is the average pattern A(j,k
) is obtained by differentiating it in the frequency axis direction.

以上の■〜■の処理を各カテゴリ毎に繰返し実行するこ
とによって前記直交化辞書9が作成される。
The orthogonalized dictionary 9 is created by repeatedly performing the above-mentioned processes (1) to (2) for each category.

尚、上述した説明では直交辞書9として3軸までを求め
る例について示したが、更に2次微分を行う等して4軸
以降の辞書を作成するようにしても良い。この場合には
、学習パターンとして前述した18点ではなく、例えば
20点以上の標本点を抽出したものを用いるようにすれ
ば良い。
In the above explanation, an example was shown in which up to three axes are obtained as the orthogonal dictionary 9, but dictionaries for four axes or later may be created by further performing second-order differentiation or the like. In this case, instead of the above-mentioned 18 points as a learning pattern, for example, a pattern obtained by extracting 20 or more sample points may be used.

一方、GS直交化部8は前記パターン蓄積部7に収集さ
れた学習パターンから、上記直交辞書に直交する付加辞
書を次のようにして作成している。
On the other hand, the GS orthogonalization unit 8 creates an additional dictionary orthogonal to the orthogonal dictionary from the learning patterns collected in the pattern storage unit 7 in the following manner.

即ち、GS直交化部8は、パターン蓄積部7に収集され
た学習パターンan(j、k)について、既に求められ
ている直交化辞書の軸数をPとしたとき、なるグラムシ
ュミットの直交化式を演算している。
That is, the GS orthogonalization unit 8 performs Gram-Schmidt orthogonalization of the learning pattern an(j, k) collected in the pattern storage unit 7, where P is the number of axes of the orthogonalization dictionary that has already been obtained. Calculating an expression.

そして上記11b   11が所定値よりも大きい場合
、pm これを付加辞書として前記直交化辞書9に登録している
。この付加辞書の作成は、パターン蓄積部7に格納され
た学習パターンam(j、k)について順に行われる。
If 11b 11 is larger than a predetermined value, pm is registered in the orthogonal dictionary 9 as an additional dictionary. The creation of this additional dictionary is performed sequentially for the learning patterns am(j, k) stored in the pattern storage section 7.

このようにして直交化時間・周波数フィルタによる平滑
・微分により作成された直交化辞書、およびこの直交化
辞書をベースとしてグラムシュミットの直交化より求め
られた付加辞書とからなる直交化辞書セットを作成して
入力音声パターンを認識処理する本装置によれば、その
直交化辞書9が音声パターンの時間軸方向および周波数
軸方向への変動を吸収したものとなっており、更にはそ
の他のパターン変動をも吸収したものとなっているので
、入力音声パターンの時間軸方向および周波数軸方向の
変動に左右されることなく音声認識することが可能とな
り、その認識性能を高めることが可能となる。また直交
化時間・周波数フィルタを用いて直交化辞書9を作成し
ている、少ない学習パターンにて性能の高い直交化辞書
を効率的に構築することが可能となり、実用的効果が多
大である。
In this way, an orthogonalized dictionary set consisting of an orthogonalized dictionary created by smoothing and differentiation using an orthogonalized time/frequency filter, and an additional dictionary obtained by Gram-Schmidt orthogonalization based on this orthogonalized dictionary is created. According to this device, which recognizes and processes an input speech pattern, the orthogonalization dictionary 9 absorbs variations in the speech pattern in the time axis direction and frequency axis direction, and further absorbs other pattern variations. Since the input speech pattern has also been absorbed, speech recognition can be performed without being affected by fluctuations in the time axis direction and frequency axis direction of the input speech pattern, and the recognition performance can be improved. In addition, the orthogonalized dictionary 9 is created using the orthogonalized time/frequency filter, and it becomes possible to efficiently construct a high-performance orthogonalized dictionary with a small number of learning patterns, which has a great practical effect.

このように時間軸方向および周波数軸方向の位置ずれを
補償する微分フィルタと、2次元パターンの変動を吸収
する直交化フィルタとを用いて直交化辞書を作成して音
声認識を行う本方式によれば、少ない学習パターンによ
って高い認識性能が得られることがわかる。しかも付加
辞書によって上述したパターン変動以外のパターン変動
をも効果的に吸収して音声認識することができる。故に
、本方式は音声認識性能の向上を図る上で多大な効果を
奏すると云える。
According to this method, speech recognition is performed by creating an orthogonalized dictionary using a differential filter that compensates for positional deviations in the time axis direction and frequency axis direction and an orthogonalization filter that absorbs fluctuations in two-dimensional patterns. For example, it can be seen that high recognition performance can be obtained with a small number of learning patterns. Furthermore, the additional dictionary allows pattern variations other than those described above to be effectively absorbed for speech recognition. Therefore, it can be said that this method has a great effect on improving speech recognition performance.

尚、本発明は上述した実施例に限定されるものではない
。ここでは3軸の直交化辞書を作成する例について説明
したが、更に多くの軸数の直交化辞書を作成することも
可能である。この場合、直交化時間・周波数フィルタの
係数としては幾つかのバリエーションが考えられるが、
要は学習パターンを時間軸方向および周波数軸方向に平
滑、1次微分、2次微分、・・・すれば良いものであり
、種々変形して実施することができる。更には上記直交
化辞書に付加する付加辞書の数(軸数)も特に制限され
るものではない。また辞書の作成に供される学習パター
ンの次元数等も特に限定されるものでもない。更にはグ
ラムシュミットの直交化以外の直交化法を用いて付加辞
書を作成することも可能である。その他、本発明はその
要旨を逸脱しない範囲で変形して実施可能である。
Note that the present invention is not limited to the embodiments described above. Although an example of creating a three-axis orthogonal dictionary has been described here, it is also possible to create an orthogonal dictionary with a larger number of axes. In this case, several variations can be considered for the coefficients of the orthogonalized time/frequency filter, but
In short, the learning pattern can be smoothed, first differentiated, second differentiated, etc. in the time axis direction and the frequency axis direction, and can be implemented with various modifications. Furthermore, the number of additional dictionaries (number of axes) added to the orthogonalized dictionary is not particularly limited. Furthermore, the number of dimensions of the learning patterns used to create the dictionary is not particularly limited. Furthermore, it is also possible to create an additional dictionary using an orthogonalization method other than Gram-Schmidt orthogonalization. In addition, the present invention can be modified and implemented without departing from the gist thereof.

[発明の効果コ 以上説明したように本発明によれば3種以上のフィルタ
を用いて時間軸方向のパターン変動および周波数軸方向
のパターン変動を吸収した直交化辞書を作成し、更にこ
の直交化辞書に直交する付加辞書を作成するので、少な
い学習パターンでそのパターンの変動を効果的に表現し
た辞書を得ることができ、その認識性能の向上を図り得
る等の実用上多大なる効果を奏する。
[Effects of the Invention] As explained above, according to the present invention, an orthogonalized dictionary that absorbs pattern fluctuations in the time axis direction and pattern fluctuations in the frequency axis direction is created using three or more types of filters, and the orthogonalized dictionary is Since an additional dictionary that is orthogonal to the dictionary is created, it is possible to obtain a dictionary that effectively expresses variations in the patterns with a small number of learning patterns, and has great practical effects such as improving recognition performance.

【図面の簡単な説明】[Brief explanation of the drawing]

第1図は本発明の一実施例方式を適用して構成される音
声認識装置の概略構成図、第2図および第3図はそれぞ
れ従来の音声認識装置の概略構成を示す図である。 ■・・・音響分析部、2・・・始端・終端検出部、5・
・・判定部、6・・・標本点抽出部、7・・・パターン
蓄積部、訃・・GS直交化部、9・・・直交化辞書、1
0・・・類似度演算部、11・・・直交化時間・周波数
フィルタ。 出願人代理人 弁理士 鈴江武彦
FIG. 1 is a schematic configuration diagram of a speech recognition device constructed by applying an embodiment of the present invention, and FIGS. 2 and 3 are diagrams each showing a schematic configuration of a conventional speech recognition device. ■...Acoustic analysis section, 2...Start/end detection section, 5.
... Judgment unit, 6... Sample point extraction unit, 7... Pattern storage unit, GS orthogonalization unit, 9... Orthogonalization dictionary, 1
0... Similarity calculation unit, 11... Orthogonal time/frequency filter. Applicant's agent Patent attorney Takehiko Suzue

Claims (3)

【特許請求の範囲】[Claims] (1)入力音声を分析処理して求められる入力音声パタ
ーンと予め収集された学習パターンに基いて作成されて
いる直交化辞書との間で類似度を計算して上記入力音声
を認識する音声認識方式において、 予め収集された学習パターンに対して少なくとも平滑処
理と微分処理とを施す3種以上のフィルタを用いて前記
直交化辞書を作成する手段と、上記直交化辞書と直交す
る付加辞書を作成する手段とを具備したことを特徴とす
る音声認識方式。
(1) Speech recognition that recognizes the input speech by calculating the similarity between the input speech pattern obtained by analyzing the input speech and an orthogonal dictionary created based on the learning patterns collected in advance In the method, means for creating the orthogonalized dictionary using three or more types of filters that perform at least smoothing processing and differentiation processing on learning patterns collected in advance, and creating an additional dictionary that is orthogonal to the orthogonalized dictionary. A voice recognition method characterized by comprising means for.
(2)フィルタは、収集された学習パターンの平均パタ
ーンを求め、この平均パターンを時間軸方向および周波
数軸方向に平滑化して第1軸の辞書を求める手段と、上
記平均パターンを時間軸方向に微分して第2軸の辞書を
求める手段と、上記平均パターンを周波数軸方向に微分
して第3軸の辞書を求める手段とを備えたものである特
許請求の範囲第1項記載の音声認識方式。
(2) The filter includes means for obtaining an average pattern of the collected learning patterns, smoothing this average pattern in the time axis direction and frequency axis direction to obtain a first axis dictionary, and smoothing the average pattern in the time axis direction. Speech recognition according to claim 1, comprising means for differentiating the average pattern to obtain a second axis dictionary, and means for differentiating the average pattern in the frequency axis direction to obtain a third axis dictionary. method.
(3)付加辞書を作成する手段は、グラムシュミットの
直交化により直交化地所に直交する付加辞書を作成する
ものである特許請求の範囲第1項記載の音声認識方式。
(3) The speech recognition system according to claim 1, wherein the means for creating the additional dictionary creates an additional dictionary that is orthogonal to the orthogonalized area by Gram-Schmidt orthogonalization.
JP62252109A 1987-10-06 1987-10-06 Voice recognition system Expired - Lifetime JP2514986B2 (en)

Priority Applications (7)

Application Number Priority Date Filing Date Title
JP62252109A JP2514986B2 (en) 1987-10-06 1987-10-06 Voice recognition system
EP88116414A EP0311022B1 (en) 1987-10-06 1988-10-04 Speech recognition apparatus and method thereof
DE3888777T DE3888777T2 (en) 1987-10-06 1988-10-04 Method and device for speech recognition.
KR1019880013005A KR910007530B1 (en) 1987-10-06 1988-10-06 Voice recognition device and there method
US07/254,110 US5001760A (en) 1987-10-06 1988-10-06 Speech recognition apparatus and method utilizing an orthogonalized dictionary
SG123594A SG123594G (en) 1987-10-06 1994-08-25 Speech recognition apparatus and method thereof
HK110794A HK110794A (en) 1987-10-06 1994-10-12 Speech recognition apparatus and method thereof

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
JP62252109A JP2514986B2 (en) 1987-10-06 1987-10-06 Voice recognition system
SG123594A SG123594G (en) 1987-10-06 1994-08-25 Speech recognition apparatus and method thereof

Publications (2)

Publication Number Publication Date
JPH0194397A true JPH0194397A (en) 1989-04-13
JP2514986B2 JP2514986B2 (en) 1996-07-10

Family

ID=26540555

Family Applications (1)

Application Number Title Priority Date Filing Date
JP62252109A Expired - Lifetime JP2514986B2 (en) 1987-10-06 1987-10-06 Voice recognition system

Country Status (1)

Country Link
JP (1) JP2514986B2 (en)

Families Citing this family (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP3015477B2 (en) 1991-02-20 2000-03-06 株式会社東芝 Voice recognition method

Also Published As

Publication number Publication date
JP2514986B2 (en) 1996-07-10

Similar Documents

Publication Publication Date Title
US20090177466A1 (en) Detection of speech spectral peaks and speech recognition method and system
JPH02238495A (en) Time series signal recognizing device
Rabiner et al. Some performance benchmarks for isolated work speech recognition systems
JPS6273391A (en) Pattern recognition learning device
Kamble et al. Emotion recognition for instantaneous Marathi spoken words
Zealouk et al. Amazigh digits speech recognition system under noise car environment
JPH0225898A (en) Voice recognizing device
JP2514986B2 (en) Voice recognition system
JP2514984B2 (en) Voice recognition system
Walid Speech recognition system based on discrete wave atoms transform partial noisy environment
CN117059073A (en) Voice recognition information acquisition method
JP2514985B2 (en) Voice recognition system
JP2502880B2 (en) Speech recognition method
Chiba et al. A speaker-independent word-recognition system using multiple classification functions
JP2514983B2 (en) Voice recognition system
EP0311022B1 (en) Speech recognition apparatus and method thereof
Ganoun et al. Performance analysis of spoken arabic digits recognition techniques
JP2856429B2 (en) Voice recognition method
Bhagath et al. Multi Model Telugu Speech Signal Analysis Towards Controlling Home Appliances.
Wu Speaker recognition based on i-vector and improved local preserving projection
CN106157949A (en) A kind of modularization robot speech recognition algorithm and sound identification module thereof
KR102862872B1 (en) System and method for analysising audio
JPH01277297A (en) Sound recognizing device
JPH0225899A (en) Voice recognizing device
KR910007530B1 (en) Voice recognition device and there method