JPS6358400A - Continuous word voice recognition equipment - Google Patents
Continuous word voice recognition equipmentInfo
- Publication number
- JPS6358400A JPS6358400A JP61203019A JP20301986A JPS6358400A JP S6358400 A JPS6358400 A JP S6358400A JP 61203019 A JP61203019 A JP 61203019A JP 20301986 A JP20301986 A JP 20301986A JP S6358400 A JPS6358400 A JP S6358400A
- Authority
- JP
- Japan
- Prior art keywords
- word
- pattern
- continuous
- isolated
- partial
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
- 238000000034 method Methods 0.000 description 11
- 238000004458 analytical method Methods 0.000 description 8
- 230000000694 effects Effects 0.000 description 7
- 230000011218 segmentation Effects 0.000 description 7
- 239000013598 vector Substances 0.000 description 5
- 238000004364 calculation method Methods 0.000 description 4
- 238000010586 diagram Methods 0.000 description 4
- 241000257465 Echinoidea Species 0.000 description 1
- 238000000354 decomposition reaction Methods 0.000 description 1
- 238000005516 engineering process Methods 0.000 description 1
- 230000005236 sound signal Effects 0.000 description 1
Abstract
(57)【要約】本公報は電子出願前の出願データであるた
め要約のデータは記録されません。(57) [Abstract] This bulletin contains application data before electronic filing, so abstract data is not recorded.
Description
【発明の詳細な説明】
〔産業上の利用分野〕
本発明は、特定話者が連続的に発声した単語列の認識を
実現する連続単語音声認識装置に関する。DETAILED DESCRIPTION OF THE INVENTION [Field of Industrial Application] The present invention relates to a continuous word speech recognition device that realizes recognition of a string of words continuously uttered by a specific speaker.
(従来の技術〕
従来、この種の連続単語音声認識装置(以下、認識装置
と略す)は、まず利用者にあらかじめ認識させる単語を
ひと通り単独に区切って発声させ(以下、孤立単語と呼
ぶ)、単語毎の音声パタンを標準パタンとして認識装置
内に記憶させ(ト記操作を標準パタンの登録と呼ぶ)、
次に、人力される連続41語音声(以下、人力パタンと
呼ぶ)に対して、各標準パタンとの間で比較操作(パタ
ンマツチング)を行い、両者の一致の度合(類似度)を
調べ、最大一致の得られる標準パタンの組合せを決定し
、これと同じ単語に属すると判定する方法を用いていた
。この方法を能率よく、かつ粒度よく実現する方法とし
て、動的計画法(ダイナミックプログラミング、以下、
DPと略ず)を利用した認識技術(特願昭50−1:1
2003およびl:]22004号以下、引用文献と称
す)が知られている。本引用文献には上記パタンマツチ
ング法による認識装置の動作原理が記載されている。こ
の原理の概要は次の通りである。何個かの単語が連続し
ている入力パタンに対し、何個かの標準パタンをあらゆ
る順列で接続することによって得られるパタンを人力パ
タンの標準パタンと考えて、大カパタン全体とのマツチ
ングを行なう。この結果得られた類似度が最大となるよ
うな標準パタンの個数と順列組合せを定めることによフ
て認識を行なう。実際には上記最大化を単語単位での最
大化処理と全体レベルでの最大化処理に分割し、各最大
化処理をDPを利用して実行することにより、処理量を
低減し実用的な処理速度を達成している。以ト述へた引
用文献記載の方法が、従来では最も有効な認識法である
と考えられる。(Prior Art) Conventionally, this type of continuous word speech recognition device (hereinafter abbreviated as a recognition device) first requires a user to separate the words to be recognized in advance and utter them individually (hereinafter referred to as isolated words). , the speech pattern for each word is stored in the recognition device as a standard pattern (the recording operation is called standard pattern registration),
Next, a comparison operation (pattern matching) is performed between the human-generated continuous 41-word speech (hereinafter referred to as a human-generated pattern) with each standard pattern, and the degree of agreement (similarity) between the two is examined. , a method was used in which the combination of standard patterns that yielded the maximum match was determined, and the combination was determined to belong to the same word. Dynamic programming (hereinafter referred to as dynamic programming) is a method to implement this method efficiently and with good granularity.
Recognition technology using DP (abbreviated as DP) (patent application 1984-1:1)
2003 and l: ] 22004 (hereinafter referred to as cited documents) are known. This cited document describes the operating principle of a recognition device using the above pattern matching method. The outline of this principle is as follows. For an input pattern in which several consecutive words are connected, the pattern obtained by connecting several standard patterns in any permutation is considered to be the standard human pattern, and is matched with the entire large Kapatan. . Recognition is performed by determining the number and permutation combination of standard patterns that maximizes the similarity obtained as a result. In reality, the above-mentioned maximization is divided into word-by-word maximization processing and overall-level maximization processing, and each maximization process is executed using DP, thereby reducing the amount of processing and making it more practical. Achieving speed. The method of citing cited documents described above is considered to be the most effective recognition method to date.
〔発明が解決しようとする問題点〕
たとえば数字“3”、rSanJを発声した場合、次に
続く数字によフてrSanJのrnJの周波数構造か大
きく変化したり、「n」のエネルギーが極端に低く(無
声化)なる場合かある。ところかrsanJのr Sa
Jは「n」に比較して周波数の変動も少なくエネルギー
値も安定して高い。しかし、従来の装置では’SaJも
rnJも同じ重みでマツチングを行なっていたため、マ
ツチングの効果が十分でなく誤認識や認識不能(マツチ
ング値が非富に悪い)となる。[Problem to be solved by the invention] For example, when the number "3", rSanJ, is uttered, the frequency structure of rnJ of rSanJ changes greatly depending on the number that follows, or the energy of "n" becomes extremely large. Sometimes it becomes low (devoiced). However, rsanJ's r Sa
Compared to "n", J has less frequency fluctuation and a stable high energy value. However, in the conventional apparatus, since 'SaJ and rnJ are matched with the same weight, the matching effect is not sufficient, resulting in erroneous recognition or inability to recognize (the matching value is extremely poor).
上述した従来の認識装置は、標準パタンの特徴と大カパ
タンの特徴を利用しないで、標準パタンと大カパタンの
比較操作を行っていたので、実用に供する場合に、種々
の要因により誤認識が生ずるという欠点がある。The conventional recognition device described above performs a comparison operation between the standard pattern and the large pattern without using the characteristics of the standard pattern and the characteristics of the large pattern, so when put into practical use, misrecognition may occur due to various factors. There is a drawback.
本発明の連続Qi語音声認識装置は、単3Δ毎に区切っ
て発声された音声パタンを孤立単語パタンとして保持し
、連続して発声された連続Q’−語バツパタンし、孤立
単語パタンをあらゆる順列で接続し、この接続された孤
立歌語パタンと連続単語パタンとの間で比較操作を行な
い、両者の一致の度合を調へ、最大の一致の得られる孤
立単語パタンの組合せを決定して、連続!…語音声を認
識する連続四ツ音声認識装置において、組合わされる孤
立単語パタンそれぞれが持つ時間−特徴情報の特徴を、
面記比較操作の際、強調する重みを記憶する重み関数記
憶部と、認識対象となる連続単語パタンが持つ時間−特
徴情報に従って、前記の比較操作の際、連続単語パタン
か持つ時間−特徴情報の特徴を強調する重みを発生する
重み関数発生部とを有することを特徴とする。The continuous Qi-word speech recognition device of the present invention retains speech patterns uttered by dividing into single 3Δ as isolated word patterns, converts consecutive Q'-word cross patterns uttered continuously, and rearranges the isolated word patterns in all permutations. , perform a comparison operation between the connected isolated song word pattern and the continuous word pattern, check the degree of matching between the two, and determine the combination of isolated word patterns that yields the maximum match. continuous! ...In a continuous four-speech recognition device that recognizes word speech, the characteristics of time-feature information possessed by each isolated word pattern to be combined are
During the orthography comparison operation, the weight function storage unit stores weights to be emphasized and time-feature information held by the continuous word pattern to be recognized. and a weighting function generating section that generates weights that emphasize the characteristics of.
すなわち、本発明は、認識対象とする各々の単語の特徴
的周波数部分を前もって調べて、その結果を重み関数と
して標準パタンと同様に各単語毎に用意し、また人力パ
タンの特徴的周波数部分も同様に調べて重み関数として
用意し、大カパタンと任意の標準パタンのマツチングの
際にト記重み関数に従ってマツチング結果を得ることに
より、単語単位でのマツチング精度を向トさせると共に
連続Q1語認識レベルにおける全体マツチングの性能を
向トさせたものである。That is, in the present invention, the characteristic frequency part of each word to be recognized is checked in advance, and the result is prepared as a weighting function for each word in the same way as a standard pattern. By similarly investigating and preparing a weight function, and obtaining a matching result according to the weight function described above when matching a large Kapatan with an arbitrary standard pattern, it is possible to improve the matching accuracy on a word-by-word basis and improve the continuous Q1 word recognition level. This improves the performance of the overall matching.
次に、本発明の実施例について図面を参照して説明をす
る。Next, embodiments of the present invention will be described with reference to the drawings.
第1図は本発明の連続9語&声認識装置の一実施例を示
すブロック図である。FIG. 1 is a block diagram showing an embodiment of the continuous nine-word and voice recognition device of the present invention.
本実施例は、マイクロホン10より人力した音声111
号を周波数分析する分析部11と、分析部11の出力を
特徴ベクトルの時系列として一時保持する大カパタンバ
ッファ12と、孤立A語パタンbjを標準パタンとして
保持する標準パタン記憶部13と、標準パタン記憶部1
3の孤立rIi語バツパタンの特徴を強調する重みW
(j)を記憶する重み関数記憶部22と、人力パタンバ
ッファ12に保持された連続単語パタンの特徴を強調す
るIRみW(i)を発生する重み関数発生部23と、重
み関数記憶部22の重みW(j)を標準パタン記憶部1
3の孤立単語パタンbjに付加し、重み関数発生部23
からの重みW(i)を入力パタンバッファ12からの連
続rlL ’5f)パタンに付加し、工1みを付加され
た孤立単語パタンと連続単語パタンとを比較操作して連
続単語パタンの部分パタンに対する各孤立単語パタンの
部分類似度と部分パタンをどの孤立単語パタンと判定し
たかの部分判定を行う第1のマツチング部14と、部分
パタンの始端と終端とに対応させて部分類似度を記憶す
る部分類似度記憶部15と、部分判定を記憶する部分判
定結果記憶部16と、漸化式値記憶部18と、部分類似
度記憶部15と部分判定結果記憶部16とのデータを入
力し、単語数設定端子1fkより入力される値に基づき
、連続単語パタンの第1の時間点までの部分類似度の総
和の最大値を第1の漸化式値として漸化式値記憶部18
に記憶し、漸化式値と第1の時間点から第2の時間点ま
での部分類似度との和の最大値を′f、2の漸化式値と
して漸化式値記憶部18に記憶し、第1の時間点を順次
変化させて館記操作を経返し、第1の時間点を仮区分点
として出力し、第1の時間点から第2の時間点までの部
分判定結果を仮判定結果として出力する第2のマツチン
グ部17と、第2のマツチング部17から出力される仮
区分点を記憶する仮区分点記憶部19と、仮判定結果を
記憶する仮判定結果記憶部20と、仮区分点記憶部j9
内の仮区分点と仮f’lJ定結果記憶部20内の仮判定
結果とを参照して各単語の区分点と単語名を決定する判
定部21とから構成されている。In this embodiment, a human voice 111 from the microphone 10 is used.
an analysis unit 11 that performs frequency analysis of the code; a large pattern buffer 12 that temporarily stores the output of the analysis unit 11 as a time series of feature vectors; and a standard pattern storage unit 13 that stores the isolated A word pattern bj as a standard pattern. Standard pattern storage section 1
Weight W that emphasizes the characteristics of the isolated rIi word cross pattern in 3
(j); a weighting function generating unit 23 that generates an IR model W(i) that emphasizes the features of the continuous word pattern held in the manual pattern buffer 12; The weight W(j) of the standard pattern storage unit 1
The weighting function generating unit 23
The weight W(i) from W(i) is added to the continuous rlL '5f) pattern from the input pattern buffer 12, and the added isolated word pattern and the continuous word pattern are compared to obtain a partial pattern of the continuous word pattern. A first matching unit 14 performs a partial determination of the partial similarity of each isolated word pattern to which isolated word pattern the partial pattern is determined to be, and stores the partial similarity in correspondence with the start and end of the partial pattern. The data of the partial similarity storage section 15 that stores partial judgments, the partial judgment result storage section 16 that stores partial judgments, the recurrence formula value storage section 18, the partial similarity storage section 15, and the partial judgment result storage section 16 is inputted. , based on the value input from the number of words setting terminal 1fk, the recurrence formula value storage unit 18 sets the maximum value of the sum of the partial similarities of the continuous word pattern up to the first time point as the first recurrence formula value.
The maximum value of the sum of the recurrence formula value and the partial similarity from the first time point to the second time point is stored in the recurrence formula value storage unit 18 as the recurrence formula value of 'f,2. The first time point is memorized, the first time point is changed sequentially, the record operation is repeated, the first time point is output as a temporary dividing point, and the partial judgment result from the first time point to the second time point is obtained. A second matching unit 17 that outputs a provisional determination result, a provisional division point storage unit 19 that stores the provisional division point output from the second matching unit 17, and a provisional determination result storage unit 20 that stores the provisional determination result. and temporary segmentation point storage section j9
The judgment unit 21 determines the segmentation point and word name of each word by referring to the tentative segmentation point in the table and the tentative judgment result in the tentative f'lJ determination result storage unit 20.
次に、本実施例の動作原理について説明する。Next, the operating principle of this embodiment will be explained.
本実施例の装置が実行する動作原理を数式的に表現する
と次のようになる。マイクロホンlOにより人力される
音声信号は分析部1】により分析処理され、周波数構造
等を表わす多次元特徴ベクトルa1の時系列パタンAと
して表わすことかできる。The operating principle executed by the apparatus of this embodiment can be expressed mathematically as follows. The audio signal manually input by the microphone IO is analyzed by the analysis unit 1 and can be expressed as a time series pattern A of a multidimensional feature vector a1 representing the frequency structure and the like.
k= al+ a2.・−、aH、・−、aj −
−−−(])一方、単独に発声された各単語(孤立1…
語)パタンも同様に分析され、時系列パタンBとして表
わすことができる。k=al+a2.・-, aH, ・-, aj −
---(]) On the other hand, each word uttered singly (isolated 1...
The pattern B is also analyzed in the same way and can be expressed as a time series pattern B.
B” = br、 b2.−−−、b’7 、 ・=
−−−−(2)nは単語を識別するための添字
、
kを連続単語に含まれる単語数として最大問題T=
(m(k) (S(^9口(1) ■ a n (
2) ■・・・■Bn(” ) ) )
−−−−(3)を計算し、最適なパラメータ(単語名
)n(k)=n (k) (k=1.2.−、 K)を
求め同時に区分点1(k)点を求める。ここで■はパタ
ンの接続を表わす演算fである。例えばBn■amは
B” Ctl ”” = br+ br、””” 1)
Ul、b7’、 b’;’。B” = br, b2.---, b'7, ・=
-----(2) where n is a subscript for identifying words, and k is the number of words included in consecutive words, the maximum problem T =
(m(k) (S(^9口(1) ■ a n (
2) ■・・・■Bn(” ) ) )
----Calculate (3), find the optimal parameter (word name) n (k) = n (k) (k = 1.2.-, K) and simultaneously find the division point 1 (k) . Here, ■ is an operation f representing the connection of patterns. For example, Bn■am is B" Ctl "" = br+ br, """ 1)
Ul, b7', b';'.
b″J1 −−−−(4)
(3)式の最大化をkおよびn(k)に関する総当り法
で計算すると膨大な計算量が必要となるが、引用文献と
同様に (3)式の最大化計算をlli語単位での処理
と全体としての処理の2段階に分割することで実用的な
処理速度を可能とする。すなわち、(+)式で表わされ
る人力パタンAの1=11よりi=mまでの部分区間と
して部分パタンA(1,m)を定義する。b″J1 ----(4) Calculating the maximization of equation (3) using the brute force method regarding k and n(k) requires a huge amount of calculation, but as in the cited document, equation (3) Practical processing speed is made possible by dividing the maximization calculation into two stages: processing in units of lli words and processing as a whole.In other words, 1 = 11 of the human pattern A expressed by the formula (+) Therefore, a partial pattern A(1, m) is defined as a partial interval up to i=m.
A(fl、 m)= a、、1.x+2. 、、、、
B、。A(fl, m)=a,,1. x+2. ,,,,
B.
以下では、lを始点、mを終点と称する。いま人カパタ
ンAニ(K−1)個の区分点fi (1) 、 l (
2) 、””1(k)・・・、ff1(K−1)を設け
、1<fi(+)<ρ(2)〈・・・・・・< 1 (
K−1) < ll(K) =■を仮定して、大カパタ
ンAをに個の部分パタンに分割する。Hereinafter, l will be referred to as the starting point and m will be referred to as the ending point. Now there are K-1 segmentation points fi (1), l (
2) , "1(k)..., ff1(K-1) are provided, and 1<fi(+)<ρ(2)<...<1 (
Assuming that K-1) < ll(K) =■, the large Kapatan A is divided into partial patterns.
A =A(+、f (1))ΦA(j2 (+) 、j
2 (2))■・・・■A(,12(k−1)!! (
k) ) ■−・・−0+八(R(K−1)、 I
)、一方、パタン間の時間軸正規化類似度を定義する
と、類似度S (八、B)はパタンの接続分解に関して
次の性質を存する。A = A (+, f (1)) ΦA (j2 (+), j
2 (2))■・・・■A(,12(k-1)!! (
k) ) ■-...-0+8(R(K-1), I
), On the other hand, if we define the time axis normalized similarity between patterns, the similarity S (8, B) has the following properties regarding the connection decomposition of patterns.
S(A・8”■8パ)“T°゛(耶)す・、x、3ニド
(T(3)式に (5)式を代入し、さらに (6)式
の関係を繰返し適用し整理すると、
となり、 (7)式の最大化問題は次のように分解して
計算することができる。S(A・8"■8pa)"T°゛(耶)su・,x,3nido(T Substitute equation (5) into equation (3), and then repeatedly apply the relationship of equation (6). When rearranged, it becomes, and the maximization problem of equation (7) can be decomposed and calculated as follows.
[11類似Jff S (八(42、m)、 [1
” ) −−−−(8)をすべてのQ<mなる部分区
間
A(n、+n)と孤立、Qt詰パタンBτ1に関して算
出する。[11 Similar Jff S (8 (42, m), [1
(8) is calculated for all partial intervals A(n, +n) where Q<m and the isolated, Qt-packed pattern Bτ1.
[21部分類似度
S (Q、 m)=max (S (八(Qlm
)、 Rn ) )部分判定結果
N (11,m)=arg max (S(八<l
、 m)、B” ) 3を計算し、テーブルに
記憶する。ここで、arg−max [・1なる記号は
[]の最大を僕える変数nを算出することを、a、味
する。[21 Partial similarity S (Q, m) = max (S (8(Qlm
), Rn )) Partial judgment result N (11, m)=arg max (S(8<l
, m), B'') 3 is calculated and stored in the table.Here, the symbol arg-max [·1 means that a, calculate the variable n that serves the maximum of [].
−−−−(+1)
なる最大問題を計算し、最適なパラメータ(+7゜分点
) fl(k) = l (k)、 k= 1.2.・
・・、Kを求める。−−−(+1) Calculate the maximum problem and find the optimal parameters (+7° point) fl(k) = l(k), k= 1.2.・
..., find K.
(11)式の最大問題は次の漸化式により計算できる。The maximum problem of equation (11) can be calculated using the following recurrence formula.
初期値 To(fl、) =O,R=1,2.・・・、
■。Initial value To(fl,)=O,R=1,2. ...,
■.
k= 1.2.・・・、に
漸化式 m=l、2.−.1 、 k=I、2.・・、
に仮置分点
L’(m)=arg max (Tk−’ (M)
+S (、Q、 m))!
−−−−(+3)
仮判定結果
N ’(111) = N (Lしくm)、 m)
−−−−(14)(+2)、 (+3)、 (目
)式の計算はに、mに関して増加する方向に計算する。k=1.2. ..., the recurrence formula m=l, 2. −. 1, k=I, 2. ...,
The hypothetical equinox L' (m) = arg max (Tk-' (M)
+S (,Q, m))! −−−−(+3) Temporary judgment result N'(111) = N (L is m), m)
----(14) (+2), (+3), Calculation of the (th) formula is performed in the direction of increasing m.
以−4二のJA埋が終了すると、(13)式のし−(m
)から区分点fl (x)か次のように決定される。After completing the above-42 JA filling, the formula (13) is
), the segmentation point fl (x) is determined as follows.
fl(に−1) −LX(+) より順次通合って仮
置分点ffi (k)を
ffi (k) = L”’ Q (k+1) )、
(k= 1.2.・・・、に−1)として、仮置分点L
’(m)のテーブルを参照して求め、それに従って、判
定結果n (k)が、(14)式の仮判定結果より
n (k) = N ’(fl (k))、 (k=
1.2.− 、K) −−−−(15)として参照する
ことで得られる。ffi (k) = L”' Q (k+1) ),
(k = 1.2..., -1), the hypothetical equinox L
Accordingly, the judgment result n (k) is calculated as n (k) = N '(fl (k)), (k =
1.2. -, K) ----(15).
以トの操作により、連続単1悟を構成する各単語の区分
点とC1i語名が旦(k)、 (k=+、2.・・・、
に−1) 。By the following operations, the segmentation point and C1i word name of each word constituting the continuous single word are dan(k), (k=+, 2...,
ni-1).
n (k) 、 (k = I、2.・・、k)として
決定される。n (k), (k = I, 2..., k).
次に1本実施例の動作について説明する。単11Δ毎に
区切って発声された音声がマイクロホン1oがら孤立中
−1iΔパタンとして分析部11に入力される。Next, the operation of this embodiment will be explained. The voices uttered in units of 11Δ are input to the analysis unit 11 as an isolated -1iΔ pattern from the microphone 1o.
分析部11で周波数分析された孤ダL昨語パタンは(2
)式で示される特徴ベクトルb」の時系列を仔するに〒
準パタンllnとして人カバターンバッファ12を介し
て標i+(パタン記憶部13に記憶される。連続的に発
声された7g声はマイクロホン1o、分析部11を軒て
(1)式で示される特徴ベクトルaiの時系列を44
−する連続’l’ +EΔパタンAとして入力パタンバ
ッファ12に記憶される。また、(瓜ウニ単語の標準パ
タン11 Tlの特徴を表現する重み関数wnU)(j
・1.2.:II、・・・1.J)か市み関2友8己憶
部22に11己憶されている。連続m1:Δパタン(人
力パタン)Aの特徴(各?11−語か持つイ〕意な特徴
成分に着[」シた)を強調1− ルrl’i、 ミ閏?
fi Wn(i) (i=1.2.3. ・・・、
I ) ’ti +lミ関数発生部23で作られて第1
のマツチング部14へ送られる。The solitary L last word pattern frequency-analyzed by the analysis unit 11 is (2
) To generate the time series of the feature vector b shown by the formula 〒
The symbol i+ (is stored in the pattern storage unit 13 as a quasi-pattern lln through the human cover pattern buffer 12.The continuously uttered 7g voice passes through the microphone 1o and the analysis unit 11, and has the characteristic shown by equation (1). The time series of vector ai is 44
- is stored in the input pattern buffer 12 as a continuous 'l' +EΔ pattern A. Also, (weight function wnU expressing the characteristics of standard pattern 11 Tl of urchin words) (j
・1.2. :II,...1. J) Kaichi Miseki 2 friends 8 members 22 members remember 11 members. Continuous m1: Δ pattern (man-powered pattern) Emphasize the features (each ?11-word) of A that have significant feature components 1-ru rl'i, mi ?
fi Wn(i) (i=1.2.3....,
I) 'ti + l The first function generated by the function generator 23
The data is sent to the matching section 14.
第2図は屯み関数Wn(、i)の−例を示す図である。FIG. 2 is a diagram showing an example of the slope function Wn(,i).
この図はrO5八Kへ Jの標準パタン[lI+に対す
る!nみ関数の一例であり、Iυff−1子音の結合部
のマツチング効果を下げ、r−a部(キ、rに摩擦音、
破’Q :’r )のマツチング効果を上げる様に重み
関数W(j)が決められている。本例ではこのようにし
てl[み関数WnU)か時間’pHl j方向にのみ決
めら打ているが、時111卜kb方向jたけてなく周波
数軸方向fに関しても特徴を統計的に調へて、市み関数
Wn(j、f)を決める方法などか考えらね、第2図の
例に限定されるものではない。This figure shows the standard pattern of J to rO58K [for lI+! This is an example of an n-shape function, which lowers the matching effect of the Iυff-1 consonant junction, and reduces the matching effect of the Iυff-1 consonant junction,
The weighting function W(j) is determined so as to enhance the matching effect of the break'Q:'r). In this example, in this way, the characteristics are determined only in the j direction of l[mifunction WnU) or time'pHl, but the characteristics are statistically adjusted not only in the direction of time 111 kb but also in the frequency axis direction f. However, the method for determining the market function Wn(j, f) is not limited to the example shown in FIG. 2.
゛)
以下余白 J
表 1
表1は重み関数発生部23で重み関数Wn(i)を発生
する際に参照する特徴−重み値テーブルの一例である。゛) Margin below J Table 1 Table 1 is an example of a feature-weight value table that is referred to when the weighting function generation unit 23 generates the weighting function Wn(i).
たとえば人力パタンのiフレーム目の特徴がSの摩擦音
であった場合、マツチングの対象が大阪の場合は重み関
数W(i) =1.9が選ばれ、東京の場合用み関数W
(i) −1,0か選ばハて各々(111)式の漸化式
に従って類似度が計算される。また屯み関数Wn(i)
も重み関数Wn(j)と同様に周波数軸方向fに関して
の特徴を統計的に調べて、重み関数Wn(i、f)を決
めることも可能である。第1のマツチング部14では次
式で定義される漸化式を各孤立単語パタンBnとパタン
Aの部分パタンA(ffi、m)に関し大カパタンベク
トルal11が人力される毎に (8)式の類似度Sを
算出する。即ち初期条件 g(!、 j” ) = S
(am 、bp )i= m=Oiミm
−−−−(+7)
漸化式
%式%(18)
なる漸化式計算をj = j” 、 j” −I、
jn−2゜・・・、1の順序で実行し、類似度
S (A(Q、 m)、 B” ) =gC(1+1
.1)−(20)を m−J”−r≦2≦m−J”
+r −−−−(21)なる範囲で算
出する。For example, if the feature of the i-th frame of the human pattern is the fricative of S, the weighting function W(i) = 1.9 is selected when the matching target is Osaka, and the weighting function W(i) = 1.9 is selected when the matching target is Osaka.
(i) Select either -1 or 0 and calculate the similarity according to the recurrence formula of equation (111). Also, the gradient function Wn(i)
Similarly to the weighting function Wn(j), it is also possible to determine the weighting function Wn(i, f) by statistically examining the characteristics in the frequency axis direction f. The first matching unit 14 calculates the recurrence formula defined by the following equation for each isolated word pattern Bn and partial pattern A(ffi, m) of pattern A, each time a large Kapatan vector al11 is manually generated using equation (8). Calculate the similarity S of . That is, the initial condition g(!, j”) = S
(am, bp) i = m = Oi mim −−−− (+7) Recurrence formula % formula % (18) Calculate the recurrence formula as j = j”, j” −I,
jn-2゜..., 1, and the similarity S (A(Q, m), B'') = gC(1+1
.. 1)-(20) m-J”-r≦2≦m-J”
+r --- (21) Calculated in the range.
上述の方法により結果として (9)式で示される部分
類似度S(1,+n)および(10)式で示される部分
判定結果N(fi、m)をそれぞれ部分類似度記憶部1
51部分判定結果記憶部16に出力する。第2のマツチ
ング部17では、部分類似度記憶部15より上記部分類
似度S(f、m)を読み出し、同時に漸化式値記憶部1
8から、l<mなる(12)式の漸化式値Tkl(2)
を、kを一定として、読み出しながら漸化式値T’(m
)を算出し、漸化式値記憶部18に出力する。同様に仮
置分点L’(m)を(13)式で算出して、仮置分点記
憶部19に出力する。仮判定結果N ’(m)は(14
)式にもとづいて部分判定結果N(42,m)と、仮置
分点L ’(m)を参照して算出され、仮判定結果記憶
部20に出力される。As a result of the above method, the partial similarity S(1,+n) shown by equation (9) and the partial judgment result N(fi, m) shown by equation (10) are respectively stored in the partial similarity storage unit 1.
51 is output to the partial determination result storage section 16. The second matching unit 17 reads out the partial similarity S(f, m) from the partial similarity storage unit 15, and at the same time reads out the partial similarity S(f, m) from the recurrence formula value storage unit 15.
8, the recurrence formula value Tkl(2) of equation (12) where l<m
, with k being constant, the recurrence formula value T'(m
) is calculated and output to the recurrence formula value storage section 18. Similarly, the temporary equinox L'(m) is calculated using equation (13) and output to the temporary equinox storage section 19. The provisional judgment result N'(m) is (14
) is calculated by referring to the partial determination result N(42, m) and the temporary equinox L′(m), and is output to the temporary determination result storage unit 20.
7J、2のマツチング部17では上記操作を単語数設定
端%ilkより入力される値を基にに=1から始め、k
=Kまで順次kを増加させながら実行する。In the matching unit 17 of 7J, 2, the above operation is started from =1 based on the value input from the word count setting terminal %ilk, and k
Execute while increasing k sequentially until =K.
かくのごとく構成された装置において、単語系列の既知
なる連続単語パタンAの始点a1がら終点a、までを順
次人力させて上述の動作を実行させることで、区分点に
関する値L’(m)と単語名を決定する値Nしくm)が
すべてのm= (+、 2.−、 I )+’に=(
1,2,・・・、K)について得られる。判定部21で
は、それぞれ仮置分点記憶部19内の仮置分点L’(m
)と仮判定結果記憶部20内の仮判定結果N ’(m)
とを参照して、(15)式に従ってkを1つづつデクリ
メントしながら順次1 (k−]) 、 j2 (k−
2) 。In the device configured as described above, the value L'(m) regarding the segmentation point can be obtained by manually performing the above-mentioned operations sequentially from the start point a1 to the end point a of the known continuous word pattern A of the word series. The value N which determines the word name (m) is for all m = (+, 2.-, I) +' = (
1, 2, ..., K). In the determination unit 21, the temporary equinox L′(m
) and the provisional determination result N′(m) in the provisional determination result storage unit 20
1 (k-]), j2 (k-
2).
・・・、+2(1)を決定する。同様にして(16)式
に従って各単語名n (k−1) 、 n (k−2)
、・” 、 n (])を決定する。..., +2(1) is determined. Similarly, each word name n (k-1), n (k-2) according to equation (16)
,・”, n (]) is determined.
以上本発明の実力’fr例を説明したが、これらの記載
は本発明の範囲を限定するものではない。例えば本明細
書では類似度を基にして動作を説明したが、距離のよう
に大小関係が逆の尺度によっても同様な処理が可能であ
る。また、抽出する部分を単語として説明したが複数の
a節からなる語句でも同様に処理することができる。さ
らに、入力音声パタンと標準パタンとの類似度を動的計
画法で説明したが、動的計画法に限定するものではない
。Although practical examples of the present invention have been described above, these descriptions do not limit the scope of the present invention. For example, in this specification, the operation has been described based on the degree of similarity, but similar processing is possible using a measure in which the magnitude relationship is reversed, such as distance. Furthermore, although the portion to be extracted has been described as a word, a word or phrase consisting of a plurality of a-clauses can also be processed in the same way. Furthermore, although the degree of similarity between the input speech pattern and the standard pattern has been explained using dynamic programming, the present invention is not limited to dynamic programming.
以北連続竿語を認識する方法を説明したが、(19)式
の制約下で表1や第2図に示す様な値の重み関数Wn(
i)、 Wn(j)で(20)式の類似度計算を実行す
ることにより、重み関数Wn(i)、 Wn(j)の極
大値近傍、つまりその孤立単語を特徴づける周波数区間
に重みか付けられてパタンマツチングが行なわれ、かつ
母音部と子(母)音部のわたり部分で比較的不安定な部
分のマツチング効果が軽減できるため、孤立単語レベル
でのマツチングの性能か向上することにより、高鯖度な
連続単語認識が実現できる。We have explained the method for recognizing continuous words from north to north, but under the constraint of equation (19), the weighting function Wn(
i), by executing the similarity calculation of equation (20) with Wn(j), weights are applied to the vicinity of the maximum values of the weighting functions Wn(i) and Wn(j), that is, to the frequency interval that characterizes the isolated word. The performance of matching at the level of isolated words can be improved because pattern matching is performed by attaching words to words, and the matching effect of relatively unstable parts at the intersection of vowel parts and consonants (vowel parts) can be reduced. This makes it possible to achieve continuous word recognition with high accuracy.
以上説明したように本発明は、連続単語パタンを接続さ
れた孤立単語パタンとして認識する際、孤ずL中1語パ
タン、連続単語パタンの仔する時間−特徴+l11報に
特徴を強調する重みを付加して比較操作することにより
、+1!−語?林位でのマツチング粒度を向トさせると
共に連続単語認識レベルにおける全体マツチングの性能
を向上させる効果がある。As explained above, when recognizing a continuous word pattern as a connected isolated word pattern, the present invention applies weights to emphasize features to the time-feature + l11 information of the single-word pattern and the continuous word pattern. +1 by adding and comparing! -Word? This has the effect of increasing the matching granularity at the forest level and improving the overall matching performance at the continuous word recognition level.
第1図は本発明の連続単語音声認識装置の一実施例を示
すブロック図、第2図は重み関数記憶部22に記憶され
ている任意の単語の標準パタンに対する重み関数W (
j)の−例を示す図である。
10・・・マイクロホン、 11・・・分析部、1
2・・・人力パタンバッファ、
13・・・標準パタン記憶部、
14・・・第1のマツチング部、
15・・・部分類似度記憶部、
16・・・部分判定結果記憶部、
17・・・第2のマツチング部、18・・・漸化式値記
憶部、19・・・仮置分点記憶部、 20・・・仮判定
結果記憶部、21・・・判定部、 22・・・
重み関数記憶部、23・・・重み関数発生部、 Wk・
・・単語数設定端子。FIG. 1 is a block diagram showing an embodiment of the continuous word speech recognition device of the present invention, and FIG. 2 is a weighting function W (
FIG. 6 is a diagram showing an example of j). 10...Microphone, 11...Analysis department, 1
2... Human pattern buffer, 13... Standard pattern storage section, 14... First matching section, 15... Partial similarity storage section, 16... Partial judgment result storage section, 17... - Second matching section, 18... Recurrence equation value storage section, 19... Temporary equinox storage section, 20... Temporary judgment result storage section, 21... Judgment section, 22...
Weighting function storage section, 23...Weighting function generation section, Wk.
・Word count setting terminal.
Claims (1)
ンとして保持し、連続して発声された連続単語パタンに
対し、孤立単語パタンをあらゆる順列で接続し、この接
続された孤立単語パタンと連続単語パタンとの間で比較
操作を行ない、両者の一致の度合を調べ、最大の一致の
得られる孤立単語パタンの組合せを決定して、連続単語
音声を認識する連続単語音声認識装置において、 組合わされる孤立単語パタンそれぞれが持つ時間−特徴
情報の特徴を、前記比較操作の際、強調する重みを記憶
する重み関数記憶部と、 認識対象となる連続単語パタンが持つ時間−特徴情報に
従って、前記比較操作の際、連続単語パタンが持つ時間
−特徴情報の特徴を強調する重みを発生する重み関数発
生部とを有することを特徴とする連続単語音声認識装置
。[Claims] A speech pattern uttered by dividing each word is held as an isolated word pattern, and the isolated word pattern is connected in any permutation to a continuous word pattern uttered continuously. Continuous word speech recognition that performs a comparison operation between isolated word patterns and continuous word patterns, checks the degree of agreement between the two, determines the combination of isolated word patterns that yields the maximum match, and recognizes continuous word speech. The apparatus includes: a weighting function storage unit that stores weights that emphasize, during the comparison operation, the characteristics of the time-feature information of each of the isolated word patterns to be combined; and the time-features of the continuous word patterns to be recognized. A continuous word speech recognition device comprising: a weighting function generator that generates a weight that emphasizes the characteristics of time-feature information of the continuous word pattern during the comparison operation according to the information.
Priority Applications (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP61203019A JPS6358400A (en) | 1986-08-28 | 1986-08-28 | Continuous word voice recognition equipment |
Applications Claiming Priority (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP61203019A JPS6358400A (en) | 1986-08-28 | 1986-08-28 | Continuous word voice recognition equipment |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| JPS6358400A true JPS6358400A (en) | 1988-03-14 |
Family
ID=16466999
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| JP61203019A Pending JPS6358400A (en) | 1986-08-28 | 1986-08-28 | Continuous word voice recognition equipment |
Country Status (1)
| Country | Link |
|---|---|
| JP (1) | JPS6358400A (en) |
Citations (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JPS58224394A (en) * | 1982-06-22 | 1983-12-26 | 日本電気株式会社 | Continuous word vice recognition equipment |
| JPS59198A (en) * | 1982-06-25 | 1984-01-05 | 中川 聖一 | Pattern comparator |
| JPS5972498A (en) * | 1982-10-19 | 1984-04-24 | 松下電器産業株式会社 | Pattern comparator |
| JPS59173883A (en) * | 1983-03-22 | 1984-10-02 | Matsushita Electric Ind Co Ltd | Pattern comparator |
-
1986
- 1986-08-28 JP JP61203019A patent/JPS6358400A/en active Pending
Patent Citations (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JPS58224394A (en) * | 1982-06-22 | 1983-12-26 | 日本電気株式会社 | Continuous word vice recognition equipment |
| JPS59198A (en) * | 1982-06-25 | 1984-01-05 | 中川 聖一 | Pattern comparator |
| JPS5972498A (en) * | 1982-10-19 | 1984-04-24 | 松下電器産業株式会社 | Pattern comparator |
| JPS59173883A (en) * | 1983-03-22 | 1984-10-02 | Matsushita Electric Ind Co Ltd | Pattern comparator |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| JP3683177B2 (en) | A method for creating context-dependent models for speech recognition | |
| JP3114975B2 (en) | Speech recognition circuit using phoneme estimation | |
| JP3412496B2 (en) | Speaker adaptation device and speech recognition device | |
| JPS58130393A (en) | Voice recognition equipment | |
| JPH05216490A (en) | Apparatus and method for speech coding and apparatus and method for speech recognition | |
| Wu et al. | Locally Linear Embedding for Exemplar-Based Spectral Conversion. | |
| JPH0535299A (en) | Method and device for coding voice | |
| JPS634200B2 (en) | ||
| CN110931045A (en) | Audio feature generation method based on convolutional neural network | |
| CN110085254A (en) | Many-to-many speech conversion method based on beta-VAE and i-vector | |
| CN110047501A (en) | Multi-to-multi phonetics transfer method based on beta-VAE | |
| Ali et al. | Gender recognition system using speech signal | |
| JPH0638199B2 (en) | Voice recognizer | |
| JP2955297B2 (en) | Speech recognition system | |
| JP2017134321A (en) | Signal processing method, signal processing device, and signal processing program | |
| CN107785030B (en) | Voice conversion method | |
| CN112967734B (en) | Multi-voice based music data recognition method, device, equipment and storage medium | |
| JP6827004B2 (en) | Speech conversion model learning device, speech converter, method, and program | |
| CA1270568A (en) | Formant pattern matching vocoder | |
| JP2980382B2 (en) | Speaker adaptive speech recognition method and apparatus | |
| JP2923243B2 (en) | Word model generation device for speech recognition and speech recognition device | |
| JPH0823758B2 (en) | Speaker-adaptive speech recognizer | |
| JP3098157B2 (en) | Speaker verification method and apparatus | |
| Das | Some dimensionality reduction studies in continuous speech recognition | |
| JP2989231B2 (en) | Voice recognition device |