JPH0342480B2 - - Google Patents
Info
- Publication number
- JPH0342480B2 JPH0342480B2 JP58060337A JP6033783A JPH0342480B2 JP H0342480 B2 JPH0342480 B2 JP H0342480B2 JP 58060337 A JP58060337 A JP 58060337A JP 6033783 A JP6033783 A JP 6033783A JP H0342480 B2 JPH0342480 B2 JP H0342480B2
- Authority
- JP
- Japan
- Prior art keywords
- pattern
- width
- patterns
- time
- voice
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Expired - Lifetime
Links
Landscapes
- Character Discrimination (AREA)
- Image Analysis (AREA)
Description
【発明の詳細な説明】
技術分野
本発明は2次元面上の特徴量として表現される
音声パターンを比較し、パターン間の類似度を求
めるパターン比較方法に関する。DETAILED DESCRIPTION OF THE INVENTION Technical Field The present invention relates to a pattern comparison method for comparing voice patterns expressed as feature amounts on a two-dimensional plane and determining the degree of similarity between the patterns.
従来技術
近年、音声認識装置のようにパターンの類似度
又はそれに準ずるものを計算し、それによつて認
識結果を選出する装置が種々考えられている。と
ころで音声を取り扱う場合、このようなパターン
の類似度を計算する上で二つの問題点がある。一
つは発音速度の相違から同じ単語音声パターンで
も時間長が異なり、そのままパターンの比較をし
て類似度の計算ができないこと、他は話者が変る
とホルマント周波数が変化するため話者間の差が
大きくなつてしまうことである。現在前者に対し
て最も広く使われている方法として動的計画法
(DP:Dynamic Programming)によるパター
ンマツチングがある。第1図によりDPマツチン
グについて簡単に説明する。パターンf(t)と
g(t)の始端、終端を一致させ、かつ非線形な
時間軸方向の伸縮をゆるしてマツチングを行ない
類似度を求める場合である。f(t),g(t)は
一定間隔でサンプリングされたデイスクリートな
量となつており、各々のサンプル点をm1,m2,
…mN,n1,n2,…nNとすると、二つのパターン
はf(m1),f(m2),…f(mN),g(n1),g
(n2),…g(nN)で表わされる。パターンの始端
f(m1)とg(n1)、及び終端f(mN)とg(nN)
が対応づけられるものとし、他の点は両パターン
間の距離が最小になるように対応づける。そのた
めにはf(m)の中の一点(mi)はg(ni)近傍
の全ての点に対応づけてみてその中から距離を最
小にするような点を選んで対応づける。その結果
第1図にAにて示すような傾斜が求まり、これに
従つてf(t)をg(t)に写影して類似度が計算
できる。ところがこの方法は、演算量が非常に多
く、またパターンの時間長の変動は吸収すること
ができるが周波数上の変動を吸収することができ
ないという欠点がある。BACKGROUND ART In recent years, various devices have been considered, such as speech recognition devices, that calculate pattern similarity or something similar thereto and select recognition results accordingly. However, when dealing with speech, there are two problems in calculating the similarity of such patterns. One is that the same word phonetic patterns have different durations due to differences in pronunciation speed, making it impossible to directly compare the patterns and calculate the degree of similarity.The other is that the formant frequency changes when the speaker changes, so This means that the difference becomes larger. Currently, the most widely used method for the former is pattern matching using dynamic programming (DP). DP matching will be briefly explained with reference to FIG. This is a case where the similarity is determined by matching the starting and ending ends of patterns f(t) and g(t) and allowing non-linear expansion/contraction in the time axis direction. f(t) and g(t) are discrete quantities sampled at regular intervals, and each sample point is expressed as m 1 , m 2 ,
...m N , n 1 , n 2 , ...n N , the two patterns are f (m 1 ), f (m 2 ), ... f (m N ), g (n 1 ), g
(n 2 ),...g(n N ). Starting points f(m 1 ) and g(n 1 ) of the pattern, and ending points f(m N ) and g(n N )
are associated with each other, and other points are associated so that the distance between both patterns is minimized. To do this, one point (mi) in f(m) is associated with all points in the vicinity of g(ni), and the point that minimizes the distance is selected from among them and associated. As a result, a slope as shown by A in FIG. 1 is obtained, and the degree of similarity can be calculated by mapping f(t) to g(t) according to this slope. However, this method has the disadvantage that it requires a very large amount of calculations, and although it can absorb variations in the time length of the pattern, it cannot absorb variations in frequency.
このように、周波数軸と時間軸が形成する2次
元面上の音声パターンが両軸に対する変動を有す
るような場合、従来少ない計算量でこれを吸収で
きる方法がない。また、パターンのわずかな歪を
吸収して比較する方法として特開昭56−116185号
公報がある。この方法では、プリント基板等のパ
ターンチエツクを行なう事を目的に、基準となる
パターンとテストパターンの一致度を比較するも
のである。まず、基準となるパターンを一定の距
離だけ移動させ、移動前後のパターン間の論理和
と論理積をとることで、それぞれ大きなパターン
と小さなパターンを作つておく。このときの移動
距離はチエツクすべきパターンの形とパターンの
論理積および論理和をとり、基板のパターンの正
常、異常をきめるものである。しかしながら、こ
の方法では基板パターンのようなわずかな歪にた
いしては有効であるが、音声認識のパターン照合
には、次のような理由から不適当である。 In this way, when the audio pattern on the two-dimensional plane formed by the frequency axis and the time axis has fluctuations with respect to both axes, there is no conventional method that can absorb this with a small amount of calculation. Furthermore, Japanese Patent Application Laid-open No. 116185/1983 discloses a method of absorbing and comparing slight distortions in patterns. In this method, the degree of coincidence between a reference pattern and a test pattern is compared for the purpose of pattern checking on printed circuit boards and the like. First, a reference pattern is moved by a certain distance, and a large pattern and a small pattern are created by calculating the logical sum and logical product between the patterns before and after the movement. The moving distance at this time is determined by calculating the AND and OR of the shape of the pattern to be checked and the pattern to determine whether the pattern on the substrate is normal or abnormal. However, although this method is effective for slight distortions such as substrate patterns, it is inappropriate for pattern matching in voice recognition for the following reasons.
(1) 音声認識では基板のパターン等に比べ変動範
囲が大きい。したがつて、論理和、積でパター
ンの内部の形が違つてしまう。上記公報に述べ
られている方法でパターン内部の形が変ると歪
等の吸収効果が無いだけでなくパターンの比較
ができなくなつてしまう。(1) In voice recognition, the range of variation is larger than in circuit board patterns. Therefore, the internal shape of the pattern differs depending on the logical sum and product. If the internal shape of the pattern changes using the method described in the above-mentioned publication, not only will there be no effect of absorbing distortion, but it will also become impossible to compare the patterns.
(2) 上記公報に述べられている方法は、基板のよ
うなパターンがあるか無かの2値の図形処理に
向いているため、音声認識のパターンも2値化
処理しなければつかえない。したがつて、2値
にした音声パターンは情報量が減るため認識語
数の制限が小語彙となつてしまう。(2) The method described in the above-mentioned publication is suitable for binary graphic processing with or without a pattern such as a board, so it cannot be used for speech recognition patterns unless they are binarized. Therefore, since the amount of information in the binary speech pattern decreases, the number of words to be recognized is limited to a small vocabulary.
(3) 仮に、上記公報に述べられている方法で、多
値のパターンが入力された時には決められた値
以上と以下で2値として扱うという規則をつけ
て補正したとしても、第13図のような場合、
(ア)のパターンは(イ)、(ウ)どちらにも同じだけの類
似度を持つ事になつて区別できなくなる。つま
り、述べられている方法は正常か異常かのチエ
ツクはできても、いくつかある標準パターンの
中のどれと一番よく似ているかというような細
かい判定には向いていない。(3) Even if a correction is made using the method described in the above publication with a rule that when a multi-value pattern is input, it will be treated as a binary value above and below a predetermined value, the result shown in Figure 13 would be In such a case,
Pattern (a) will have the same degree of similarity in both (b) and (c), making them indistinguishable. In other words, although the described method can check whether a pattern is normal or abnormal, it is not suitable for detailed judgments such as determining which of several standard patterns it most closely resembles.
(4) 論理和、積によつてパターンに変形を加える
と、音声パターンの時間方向だけでなくそれに
直角の方向まで、多くの場合は周波数軸のデー
タの数まで変つてしまい取扱いができない。(4) When a pattern is modified by logical sum or product, not only the time direction of the audio pattern but also the direction perpendicular to it, and in many cases, the number of data on the frequency axis also changes, making it impossible to handle.
このように、上記公報に述べられた従来の方法
は、音声パターンの照合には適用できないもので
あつた。 As described above, the conventional method described in the above publication cannot be applied to voice pattern matching.
目 的
本発明は斯かる事情に鑑みてなされたもので、
周波数軸と時間軸が形成する2次元面上の音声パ
ターンの両軸に対する変動を吸収して、少ない計
算量で音声パターンの比較を行なうことのできる
音声パターン比較方法を提供しようとするもので
ある。Purpose The present invention was made in view of the above circumstances, and
The present invention aims to provide a speech pattern comparison method that can absorb variations in speech patterns on both axes on a two-dimensional plane formed by the frequency axis and the time axis, and can compare speech patterns with a small amount of calculation. .
構 成
本発明の構成について、以下、実施例に基づい
て説明する。Configuration The configuration of the present invention will be described below based on examples.
ある話者が発声した単語“size”のパターンを
第2図に示す。この図は横軸に周波数、縦軸に時
間をとつて“size”と発生した時のスペクトル分
布を濃淡で表わしたものであり黒く見える程レベ
ルが大きい。周波数は左側から右へ高くなり、
250Hz〜6.3kHzを対数等間隔で15等分してある。
同じ話者が同じ単語を別の機会に発声した例を第
3図に示す。図から明らかなように両者は時間軸
方向への長さが異なつている。 Figure 2 shows the pattern of the word "size" uttered by a certain speaker. In this figure, the horizontal axis represents frequency, and the vertical axis represents time, and the spectral distribution at the time of occurrence is expressed as "size" in shading; the blacker it appears, the higher the level. The frequency increases from left to right,
250Hz to 6.3kHz is divided into 15 equal logarithmic intervals.
FIG. 3 shows an example in which the same speaker utters the same word on different occasions. As is clear from the figure, both have different lengths in the time axis direction.
我々が発する音声を特徴づけるものにホルマン
トがある。或いはスペクトルのローカルピークと
いう概念〔音響学会誌第32巻1号(1976)第12〜
23頁〕を用いても良いが、いずれにしても言語を
発声するために我々は音道の形態を変化させ、そ
の影響が音声スペクトル上にローカルピークとし
て現われる。従つて、このようなローカルピーク
の時間変化には発せられた言語の特徴が現われて
いる。そこでローカルピークの時間変化を表わす
時間−周波数パターン(以下time−spectrum
pattern、略してT.S.Pと称する)の比較によつ
て発せられた言語を認識することを考える。第2
図、第3図に示したどちらのT.S.Pも冒険の10〜
15msが/s/、次の100ms位が/a/、続く
10ms弱が/i/でその後の数msが/Z/、最
後が短く/u/を表わすパターンである。ところ
で図に示されたような時間長の変化の他に発生者
の差がピークの周波数変化として現われるが、そ
のどちらも極端なものではない。そこで二つのパ
ターンを照合する場合に、周波数変動と時間変動
の幅を考慮して、一方のパターンの幅は広くとつ
ておき、他方のパターンは、幅のある線図形から
線の特徴を取り出す手法の一つである細線化法に
よつて幅のほぼ中央近傍の点又は中心線を取り出
してから照合を行なう。この際、時間軸方向も幅
を狭めておくことが望ましい。この場合のパター
ン幅という意味は、パターン中の図形、若くはパ
ターン中の模様のことを意味するもので、パター
ン全体の大きさを変えることを意味するものでは
ない。こうすることによつて、一方のパターンの
時間、周波数の両端が変動しても細線化した細い
線パターンは幅の広いパターンからはみ出すこと
なくマツチングがとれる。すなわち、音声の標準
パターンと入力音声パターンとを、時系列とその
時系列に対応する音響的特徴列(音声波の周波数
列)との2次元平面で形成し、前記パターンに生
成される言語的情報を担う言語情報の列(音声波
の短時間スペクトル分析に基づくスペクトル包絡
に現われるローカルピーク列)に沿つて形成され
る、かつ、各時系列に対応して前記特徴列の方向
に前記言語情報を含んで形成される特徴幅及び/
又は各特徴列に対応して前記時系列の方向に形成
される時間幅を前記標準パターンと音声パターン
とで互に異ならしめ、これら両パターンが重畳さ
れて生成されるパターンの重なり部分の大きさよ
り音声パターンの類似度を求めるものである。 Formants are what characterizes the sounds we make. Or the concept of local peaks in the spectrum [Journal of the Acoustical Society of Japan, Vol. 32, No. 1 (1976) No. 12~
[Page 23] may be used, but in any case, in order to produce language, we change the shape of the sound path, and this effect appears as local peaks on the speech spectrum. Therefore, the characteristics of the spoken language appear in the temporal change of such local peaks. Therefore, the time-frequency pattern (hereinafter referred to as time-spectrum) that represents the temporal change of local peaks is
Consider recognizing uttered language by comparing patterns (abbreviated as TSP). Second
Both TSPs shown in Figures and Figure 3 are 10~
15ms is /s/, next 100ms is /a/, etc.
The pattern is /i/ for a little less than 10 ms, followed by /Z/ for several ms, and short at the end, /u/. Incidentally, in addition to changes in time length as shown in the figure, differences in the number of occurrences appear as changes in peak frequency, but neither of these is extreme. Therefore, when comparing two patterns, one pattern is kept wide, taking into account the width of frequency fluctuation and time fluctuation, and the other pattern is created by extracting line features from a wide line figure. Verification is performed after extracting a point or center line near the center of the width using a thinning method, which is one of the methods. At this time, it is desirable to narrow the width in the time axis direction as well. The term "pattern width" in this case refers to a figure in a pattern, or rather a pattern in a pattern, and does not mean changing the size of the entire pattern. By doing this, even if both ends of time and frequency of one pattern fluctuate, the thin line pattern can be matched without protruding from the wider pattern. That is, a standard speech pattern and an input speech pattern are formed on a two-dimensional plane of a time series and an acoustic feature sequence (frequency sequence of audio waves) corresponding to the time series, and the linguistic information generated in the pattern is The linguistic information is formed along a string of linguistic information (a string of local peaks appearing in a spectrum envelope based on short-time spectrum analysis of speech waves) that carries Feature width and/or
Alternatively, the time width formed in the time series direction corresponding to each feature sequence is made different between the standard pattern and the audio pattern, and the size of the overlapping part of the pattern generated by superimposing these two patterns is This method determines the similarity of voice patterns.
第4図は、本発明のハード構成図、第5図は、
第4図に示した構成図の動作説明をするためのフ
ローチヤートで、図示のように、2次元パターン
生成処理部11に取り込まれたデータ(step1)
は該2次元パターン生成処理部11によつて2次
元パターンに生成される。2次元パターン記憶部
13に記憶されている2次元パターンは、データ
の2次元パターンとの類似度を比較するためのパ
ターンで、これを標準パターンということにす
る。この標準パターンはあらかじめ用意して記憶
しておいたものでもよいし、データの2次元パタ
ーンを取り込んで、これを標準パターンに形成し
てもよい。これらデータの2次元パターンと標準
の2次元パターンはこれらを照合するに先立つ
て、パターン幅変換処理部12において、必要に
応じたパターン幅変換処理が行なわれる。なお、
パターン幅変換処理とはデータの2次元パターン
と標準の2次元パターンとにおいて、両2次元パ
ターンと対応する各次元軸方向のパターン幅を異
なるしめる(Step2)ことである。このような処
理を行なうことによつて、一方のパターンにおい
てその各次元軸上に現われるパターン幅と、他方
のパターンにおいてその各次元軸上に現われるパ
ターン幅とに幅の差を生ずることになる。照合処
理部14は、上述のようにして形成されるパター
ン幅の互いに異なる両パターンを重ね合わせて
(step3)、その重なりの大きさを求め(step4)、
これを類似度として出力する。 FIG. 4 is a hardware configuration diagram of the present invention, and FIG. 5 is a hardware configuration diagram of the present invention.
This is a flowchart for explaining the operation of the configuration diagram shown in FIG.
is generated into a two-dimensional pattern by the two-dimensional pattern generation processing section 11. The two-dimensional pattern stored in the two-dimensional pattern storage unit 13 is a pattern for comparing the degree of similarity with the two-dimensional pattern of data, and will be referred to as a standard pattern. This standard pattern may be prepared and stored in advance, or a two-dimensional pattern of data may be imported and formed into a standard pattern. Prior to comparing the two-dimensional pattern of these data and the standard two-dimensional pattern, pattern width conversion processing is performed as necessary in a pattern width conversion processing section 12. In addition,
The pattern width conversion process is to make the pattern widths of the data two-dimensional pattern and the standard two-dimensional pattern different in each dimensional axis direction corresponding to both the two-dimensional patterns (Step 2). By performing such processing, a difference in width is generated between the pattern width appearing on each dimensional axis in one pattern and the pattern width appearing on each dimensional axis in the other pattern. The matching processing unit 14 overlaps both patterns having different pattern widths formed as described above (step 3), calculates the size of the overlap (step 4),
This is output as a degree of similarity.
次に、本発明のパターン比較装置の一実施例を
第6図に示す。 Next, an embodiment of the pattern comparison device of the present invention is shown in FIG.
第6図において、マイク1から入力された音声
信号はフイルターバンク2を通り、周波数−時間
パターンとなる。その中から音声区間切り出し部
3で音声部を切り出し、ある閾値を設定すること
により2値化部4で2値化する。この2値化は情
報量低減のためであつて、勿論2値化をしなくて
も良い。これを細線化部5によつてほぼ中央らし
い点又は中心線として辞書部6に格納しておく。
次に、スイツチ7を照合部8側にし、入力音声の
周波数−時間パターンを2値化した後、辞書部6
に格納してある各単語と照合した時すなわち二つ
のパターンを重ねた時、細線化パターンがどの程
度重なるかを求め類似度を計算する。この照合を
辞書部に格納された各パターンに対し行ない、最
も類似度の大きい単語を認識結果9とする。な
お、例として第3図に示すパターンを細線化した
ものを第7図に、第2図に示すパターンを2値化
したものを第8図に示す。ここでの細線化処理は
いろいろ考えられるが(例えば電子通信学会研究
会資料PRL−75−66、第49〜56頁参照)基本的
には2値図形の境界に接している点を図形の連結
性を保つたまま1点ずつ消していつてほぼ中央近
傍の点、又は中心線を取り出す。 In FIG. 6, an audio signal input from a microphone 1 passes through a filter bank 2 and becomes a frequency-time pattern. A voice section cutout section 3 cuts out a voice portion from the voice section, and a binarization section 4 binarizes it by setting a certain threshold value. This binarization is for reducing the amount of information, and of course binarization is not necessary. This is stored in the dictionary section 6 by the line thinning section 5 as a point or center line that appears to be approximately at the center.
Next, the switch 7 is set to the matching unit 8 side, and after the frequency-time pattern of the input voice is binarized, the dictionary unit 6
When compared with each word stored in , that is, when the two patterns are overlapped, the extent to which the thinning patterns overlap is determined and the degree of similarity is calculated. This comparison is performed for each pattern stored in the dictionary section, and the word with the highest degree of similarity is set as recognition result 9. As an example, FIG. 7 shows a thinned version of the pattern shown in FIG. 3, and FIG. 8 shows a binarized version of the pattern shown in FIG. Various thinning processes can be considered here (for example, see Institute of Electronics and Communication Engineers study group material PRL-75-66, pp. 49-56), but basically the points touching the boundaries of binary figures are connected to each other. Delete the points one by one while preserving the properties, and take out the points near the center or the center line.
なお、以上の説明において一方のパターンを細
線化したが、これは二つのパターンの比較におい
てはみ出すことなくマツチングをとるためである
から、一方を線図形の特徴を保持して太線化して
もよく、或いは一方を細線化し、他方を太線化し
ても良い。 Note that in the above explanation, one of the patterns has been made into a thin line, but this is to ensure matching without overlapping when comparing the two patterns, so one may be made into a thick line while retaining the characteristics of the line shape. Alternatively, one may be made into a thin line and the other may be made into a thick line.
第9図は、二つのパターンの一方を太線化して
他方に重ねる本発明の他の実施例を示す図で、第
6図の場合とは逆にここでは辞書登録するパター
ンには処理を加えず、認識すべきパターンに太線
化部10で太線切処理を行なつている。 FIG. 9 is a diagram showing another embodiment of the present invention in which one of two patterns is made into a thick line and superimposed on the other; contrary to the case of FIG. 6, no processing is applied to the pattern to be registered in the dictionary. , the pattern to be recognized is subjected to thick line cutting processing by the thick line making section 10.
第10図は、細線化と太線化の両方を行なう本
発明の他の実施例を示す図で、辞書登録用パター
ンを細線化し、認識用パターンを太線化している
が、勿論これを逆にしても良い。なお、以上に音
声パターンを例にとつて説明したが、本発明は音
声パターンに限定されるものでなく、他の2次元
パターンでも良いことは明らかである。。 FIG. 10 is a diagram showing another embodiment of the present invention that performs both thinning and thickening, in which the dictionary registration pattern is thinned and the recognition pattern is thickened, but of course this can be reversed. Also good. Note that although the above description has been made using a voice pattern as an example, it is clear that the present invention is not limited to voice patterns and may be applied to other two-dimensional patterns. .
第11図は、本発明の一実施例を説明するため
のフローチヤートで、この実施例は、図示のよう
にデータの2次元パターンを取り込んで(step1
→2→3)これを標準の2次元パターンとして形
成する(step4)の場合のものである。一般には、
データの2次元パターンは形成された標準の2次
パターンと比較されながら(step6)、前者の各次
元軸方向のパターン幅が後者の各対応次元軸方向
のパターンに対して異なるように決められてゆく
(step1→2→3→5)。しかし、データパターン
の各次元軸方向の幅の変換の傾向が、標準パター
ンの対応幅より常に細線化又は太線化するいずれ
か一方向にあるような変換の仕上をすれば、両パ
ターンの当該幅が異なるかどうかの比較判断
(step6)は不要となる。 FIG. 11 is a flowchart for explaining one embodiment of the present invention, and this embodiment involves importing a two-dimensional pattern of data as shown (step 1).
→2→3) This is the case of forming this as a standard two-dimensional pattern (step 4). In general,
The two-dimensional pattern of the data is compared with the formed standard two-dimensional pattern (step 6), and the pattern width of the former in each dimension axis direction is determined to be different from the latter pattern in the corresponding dimension axis direction. Go (step 1 → 2 → 3 → 5). However, if the conversion tendency of the width of the data pattern in each dimensional axis direction is always in one direction, either thinner or thicker than the corresponding width of the standard pattern, then the corresponding width of both patterns There is no need to make a comparative judgment (step 6) as to whether or not they are different.
第12図は、本発明の他の実施例を説明するた
めのフローチヤートで、この実施例も、図示のよ
うに、データの2次元パターンを取り込み、
(step1→2→3)これを標準パターンとして形成
する(step6)場合であり、この点は、第11図
に示したフローチヤートと同じである第11図に
示したフローチヤートと異なるのは、標準2次元
パターンを形成する際に、これとデータとして取
り込んだ2次元パターン(自己自身)との比較を
しながら(step5)前者の各次元軸方向のパター
ン幅が後者の各対応次元軸方向のパターンとに対
して異なるように決められてゆく(step1→2→
3→4→5→6→3)点である。しかし、この場
合にも形成される標準パターンの各次元軸方向の
幅の変換の傾向がデータパターンの対応幅より常
に細線化又は太線化するいずれか一方向にあるよ
うな変換の仕方をすれば両パターンの当該幅から
異なるかどうかの比較判断(step5)は不要とな
る。 FIG. 12 is a flowchart for explaining another embodiment of the present invention. This embodiment also captures a two-dimensional pattern of data as shown in the figure.
(Step 1 → 2 → 3) This is the case where this is formed as a standard pattern (Step 6). This point is the same as the flowchart shown in FIG. 11. What is different from the flowchart shown in FIG. When forming a standard two-dimensional pattern, while comparing it with the two-dimensional pattern (self) imported as data (step 5), the pattern width in each dimension axis direction of the former is compared with that of each corresponding dimension axis of the latter. The patterns are determined differently (step 1 → 2 →
3→4→5→6→3) points. However, even in this case, if the conversion method is such that the width of the standard pattern formed in each dimensional axis direction is always thinner or thicker than the corresponding width of the data pattern, then There is no need to compare and judge whether the widths of both patterns are different (step 5).
効 果
以上のように本発明によれば、例えば発声時に
おける単語長の時間軸の変動並びに発声者による
周波数変動のような2次元パターンの各軸に対す
る変動を吸収してパターンの比較を行なうことが
可能である。Effects As described above, according to the present invention, patterns can be compared by absorbing fluctuations in each axis of a two-dimensional pattern, such as fluctuations in the time axis of word length during utterance and frequency fluctuations due to the speaker. is possible.
第1図はDPマツチングの説明図、第2図、第
3図は時間−周波数パターンを示す図、第4図
は、本発明のハード構成図、第5図は、第4図の
構成図の動作説明をするためのフローチヤート、
第6図は、本発明によるパターン比較装置の一実
施例を示す図、第7図は、第3図のパターンを細
線化した図、第8図は、第2図のパターンを2値
化した図、第9図、第10図は、本発明によるパ
ターン比較装置の他の実施例を示す図、第11図
及び第12図は、それぞれ本発明の実施例を説明
するためのフローチヤート、第13図は、従来例
を説明するための図である。
1……マイク、2……フイルターバンク、3…
…音声区間切り出し部、4……2値化部、5……
細線化部、6……辞書部、7……スイツチ、8…
…照合部、10……太線化部、11……2次元パ
ターン生成処理部、12……パターン幅変換処理
部、13……標準の2次元パターン記憶部、14
……照合処理部。
Figure 1 is an explanatory diagram of DP matching, Figures 2 and 3 are diagrams showing time-frequency patterns, Figure 4 is a hardware configuration diagram of the present invention, and Figure 5 is a configuration diagram of the configuration diagram in Figure 4. Flowchart to explain the operation,
FIG. 6 is a diagram showing an embodiment of the pattern comparison device according to the present invention, FIG. 7 is a diagram in which the pattern in FIG. 3 is thinned, and FIG. 8 is a diagram in which the pattern in FIG. 2 is binarized. 9 and 10 are diagrams showing other embodiments of the pattern comparison device according to the present invention, and FIGS. 11 and 12 are flowcharts for explaining the embodiment of the present invention, respectively. FIG. 13 is a diagram for explaining a conventional example. 1...Microphone, 2...Filter bank, 3...
...Voice section extraction unit, 4...Binarization unit, 5...
Thinning section, 6... Dictionary section, 7... Switch, 8...
... Collation section, 10... Thick line forming section, 11... Two-dimensional pattern generation processing section, 12... Pattern width conversion processing section, 13... Standard two-dimensional pattern storage section, 14
...Verification processing section.
Claims (1)
を、時系列とその時系列に対応する音響的特徴列
との2次元平面で形成し、前記パターンに生成さ
れる言語的情報を担う言語情報の列に沿つて形成
され、かつ、各時系列に対応して前記特徴列の方
向に前記言語情報を含んで形成される特徴幅及
び/又は各特徴列に対応して前記時系列の方向に
形成される時間幅について前記標準パターンと音
声パターンのいずれか一方のものを細線化又は太
線化して他方のパターン上へ重畳させて生成され
るパターンの重なり部分の大きさより音声パター
ンの類似度を求めることを特徴とする音声パター
ン比較方法。 2 前記標準パターンと音声パターンのいずれか
一方の前記特徴幅及び/又は時間幅を細線化し、
他方のパターンの特徴幅及び/又は時間幅を太線
化し、これら両パターンを重畳させることを特徴
とする特許請求の範囲第1項に記載の音声パター
ン比較方法。 3 前記音響的特徴列は、音声波の周波数列であ
り、前記言語的情報を担う言語情報の列は音声波
の短時間スペクトル分析に基づくスペクトル包絡
に現われるローカルピーク列であることを特徴と
する特許請求の範囲第1項又は第2項に記載の音
声パターン比較方法。 4 前記短時間スペクトル分析は、帯域フイルタ
群による分析であることを特徴とする特許請求の
範囲第3項に記載の音声パターン比較方法。[Claims] 1. A standard speech pattern and an input speech pattern are formed on a two-dimensional plane of a time series and an acoustic feature sequence corresponding to the time series, and the linguistic information generated in the pattern is carried. A feature width formed along a column of linguistic information and including the linguistic information in the direction of the feature column corresponding to each time series and/or a feature width of the time series corresponding to each feature column. Regarding the time width formed in the direction, the similarity of the voice patterns is determined from the size of the overlapping part of the pattern generated by making either the standard pattern or the voice pattern thinner or thicker and superimposing it on the other pattern. A voice pattern comparison method characterized by determining. 2. Thinning the characteristic width and/or time width of either the standard pattern or the audio pattern;
2. The voice pattern comparison method according to claim 1, wherein the feature width and/or time width of the other pattern is made thicker and the two patterns are superimposed. 3. The acoustic feature sequence is a frequency sequence of a speech wave, and the linguistic information sequence carrying the linguistic information is a local peak sequence appearing in a spectrum envelope based on short-time spectrum analysis of the speech wave. A voice pattern comparison method according to claim 1 or 2. 4. The voice pattern comparison method according to claim 3, wherein the short-time spectrum analysis is an analysis using a group of band filters.
Priority Applications (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP58060337A JPS59186073A (en) | 1983-04-06 | 1983-04-06 | Pattern comparing device |
Applications Claiming Priority (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP58060337A JPS59186073A (en) | 1983-04-06 | 1983-04-06 | Pattern comparing device |
Publications (2)
| Publication Number | Publication Date |
|---|---|
| JPS59186073A JPS59186073A (en) | 1984-10-22 |
| JPH0342480B2 true JPH0342480B2 (en) | 1991-06-27 |
Family
ID=13139245
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| JP58060337A Granted JPS59186073A (en) | 1983-04-06 | 1983-04-06 | Pattern comparing device |
Country Status (1)
| Country | Link |
|---|---|
| JP (1) | JPS59186073A (en) |
Family Cites Families (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JPS56116185A (en) * | 1980-02-20 | 1981-09-11 | Hitachi Denshi Ltd | Pattern comparing method |
-
1983
- 1983-04-06 JP JP58060337A patent/JPS59186073A/en active Granted
Also Published As
| Publication number | Publication date |
|---|---|
| JPS59186073A (en) | 1984-10-22 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US7016833B2 (en) | Speaker verification system using acoustic data and non-acoustic data | |
| JPH0352640B2 (en) | ||
| JPS62217295A (en) | Voice recognition system | |
| CN112489692B (en) | Voice endpoint detection method and device | |
| Ravinder | Comparison of hmm and dtw for isolated word recognition system of punjabi language | |
| JP3069531B2 (en) | Voice recognition method | |
| JPH0222960B2 (en) | ||
| JPH0342480B2 (en) | ||
| JPH0449952B2 (en) | ||
| CN118155632A (en) | Voiceprint feature extraction algorithm based on dynamic segmentation of context-dependent spectral coefficients | |
| JPH0554118B2 (en) | ||
| JPH0283595A (en) | Voice recognition method | |
| JP2557497B2 (en) | How to identify male and female voices | |
| JPS59204897A (en) | Voice recognition dictionary registration method | |
| Shikano | Acoustic processing in the conversational speech recognition system | |
| JPS59195295A (en) | Voice recognition dictionary registration system | |
| CN106875935A (en) | Speech-sound intelligent recognizes cleaning method | |
| JPS59195293A (en) | pattern comparison device | |
| JP2886879B2 (en) | Voice recognition method | |
| JPS59195294A (en) | Voice pattern comparison device | |
| JPS61260299A (en) | Voice recognition equipment | |
| JPS59204899A (en) | Voice pattern matching device | |
| JPH0534679B2 (en) | ||
| JPS59170894A (en) | Voice section starting system | |
| Rakhi et al. | Weighted Multi-band Summary Correlogram (MBSC)-based Pitch Estimation and Voice Activity Detection for Noisy Speech |