JPH0566599B2 - - Google Patents
Info
- Publication number
- JPH0566599B2 JPH0566599B2 JP59269955A JP26995584A JPH0566599B2 JP H0566599 B2 JPH0566599 B2 JP H0566599B2 JP 59269955 A JP59269955 A JP 59269955A JP 26995584 A JP26995584 A JP 26995584A JP H0566599 B2 JPH0566599 B2 JP H0566599B2
- Authority
- JP
- Japan
- Prior art keywords
- word
- matching
- segment
- distance
- vowel
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Expired - Fee Related
Links
Description
産業上の利用分野
本発明は連続発声された単語や文節を音節等の
音声素片単位で認識する音声認識装置に関する。
従来の技術
人間にとつて最も自然な情報発生手段である音
声が、人間−機械系の入力手段として使用できれ
ば、その効果は非常に大きい。
従来、音声認識装置としては特定話者登録方式
によるものが実用化されている。即ち、認識装置
を使用しようとする話者が、予め、認識すべきす
べての単語を自分の声で特徴ベクトルの系列に変
換し単語辞書に標準パターンとして登録してお
き、認識時に発声された音声を、同様に特徴ベク
トルの系列に変換し、前記単語辞書中のどの単語
に最も近いかを予め定められた規即によつて計算
し、最も類似している単語を認識結果とするもの
である。
ところが、この方法によると、認識単語数が少
いときには良いが、数百、数千単語といつたよう
に増加してくると、主として次の三つの問題が無
視し得なくなる。
(1) 登録時における話者の負担が著しく増大す
る。
(2) 認識時に発声された音声と標準パターンとの
類似度あるいは距離を計算するのに要する時間
が著しく増大し、認識装置の応答速度が遅くな
る。
(3) 前記単語辞書のために要するメモリが非常に
大きくなる。
以上の欠点を回避するための方法として認識の
単位を子音+母音および母音の単音節(以後それ
ぞれCV,Vで表す。Cは子音、Vは母音を意味
する。)とする方法がある。即ち、標準パターン
として単音節を特徴ベクトルの系列として登録し
ておき、認識時に特徴ベクトルの系列に変換され
た入力音声を、前記単音節の標準パターンとマツ
チングすることにより、単音節の系列に変換する
ものである。日本語の場合、単音節はたかだか
101種類であり、単音節は仮名文字に対応してい
るから、この方法によれば、日本語の任意の単語
あるいは文章を単音節列に変換する(認識する)
ことができ、前記(1)〜(3)の問題はすべて解決され
ることになる。しかし、この場合の問題として調
音結合とセグメンテーシヨンがある。調音結合
は、音節を連続して発声すると各音節は前後の音
節の影響を受け、スペクトル構造が前後に接続さ
れる音節によつて変化する現象である。セグメン
テーシヨンは、連続して発声された音声を単音節
単位に区切ることであるが、これを確実に行う決
定的な方法は未だ見出されていない。この2つの
問題を解決するために、現在のところ各単音節を
区切つて、発声することが行われており、実用化
されている装置もある。
しかし、単音節を離散的に発声するのは不自然
であり、話者に緊張を強いるものである。
発明が解決しようとする問題点
本発明は前記連続発声された音声に対するセグ
メンテーシヨンの不確実さを回避し、あわせて、
連続発声された単語または文節を認識することが
できる連続音声認識装置を提供することを目的と
する。
問題点を解決するための手段
本発明は、単語・文節等を連続発声して得られ
る入力音声信号を特徴ベクトルの系列に変換する
特徴抽出手段と母音、子音あるいはそれらの結合
したもの等として定義される音声素片のそれぞれ
に対応した特徴ベクトルの系列を前記音声素片名
に対応づけて記憶する標準パターン記憶手段と、
入力パターンに対して素片の境界を検出する素片
境界候補検出手段と標準パターンのそれぞれと前
記入力パターンから検出された前記素片境界候補
の任意または定められた種々の組合せによつて決
定される部分区間(第1の部分区間)とのマツチ
ングを行つて両者の距離(類似度)を計算する素
片マツチング手段と、認識さるべき各単語・文節
等を前記音声素片名の系列として表現した単語・
文節等を記憶する単語・文節辞書と、この認識さ
るべき各単語・文節と前記入力パターンの任意ま
たは定められた前記素片境界候補の種々の部分区
間(第2の部分区間)との距離(類似度)を、前
記単語・文節辞書によつて指定される素片名の系
列に対応するように、前記第2の部分区間に含ま
れる前記第1の部分区間群を隣り合う区間が連続
するように最適に定めることにより、前記第1の
各部分区間の始点と終点およびその部分区間の前
記素片名に対応する距離(類似度)の総和を最小
(最大)とし、得られる最小値(最大値)を各単
語・文節に対する前記第2の部分区間の距離とし
て出力する機能を有する単語・文節マツチング手
段と、前記第2の部分区間群を隣り合う区間が連
続するように最適に定めることにより、前記第2
の各部分区間の始点と終点およびその部分区間の
前記単語・文節名に対応する距離(類似度)の総
和を最小(最大)となし、そのときの単語・文節
列を認識結果として判定する連続単語・文節判定
手段とを備えた連続音声認識装置である。
作 用
本発明は前記した構成により、単語・文節等を
連続発声して得られる入力音声信号を特徴ベクト
ルの系列に変換し、母音、子音あるいはそれらの
結合したもの等として定義される音声素片のそれ
ぞれに対応した特徴ベクトルの系列を前記音声素
片名に対応づけて記憶された標準パターンと、こ
の標準パターンのそれぞれと前記入力パターンか
ら検出された素片境界候補の任意または予め定め
られた種々の組合せによつて決定される部分区間
(第1の部分区間)とのマツチングを行つて両者
の距離(類似度)を計算し、認識さるべき各単
語・文節等を前記音声素片名の系列として表現し
た単語・文節等の認識さるべき各単語・文節と前
記入力パターンの任意または定められた前記素片
境界候補の種々の部分区間(第2の部分区間)と
の距離(類似度)を、前記単語・文節によつて指
定される素片名の系列に対応するように、前記第
2の部分区間に含まれる前記第1の部分区間群を
隣り合う区間が連続するように最適に定めること
により、前記第1の各部分区間の始点と終点およ
びその部分区間の前記素片名に対応する距離(類
似度)の総和を最小(最大)とし、得られる最小
値(最大値)を各単語・文節に対する前記第2の
部分区間の距離とし、前記第2の部分区間群を隣
り合う区間が連続するように最適に定めることに
より、前記第2の各部分区間の始点と終点および
その部分区間の前記単語・文節名に対応する距離
(類似度)の総和を最小(最大)となし、そのと
きの単語・文節列を認識結果として判定する。
実施例
以後、「単語」という言葉は「文節」という言
葉も代表するものとする。また、「類似度」は
「距離」で代表して説明する。即ち、距離が小さ
いとは類似度が大きいということである。先ず本
発明の基本であるDPマツチングについて述べる。
第2図は離散単語の認識を行う場合のDPマツ
チングを説明する格子グラフである。即ち、入力
パターンA=a1,a2…ai…aIと標準パターンBn=
bn 1,bn 2…bn j…bn jnとの距離を求める場合を示して
いる。横軸は入力パターン、縦軸は標準パターン
を示し、1は両者の特徴ベクトルの対応関係を示
す曲線である。DPマツチングは、この径路を最
適に定めることにより、その径路によつて対応づ
けられるaiとbn jとの距離dn(i,j)のこの径路
に沿う荷重平均を最小化し、その最小値を以つて
両者の距離とするものであつて、この計算を効率
的に行うものである。dn(i,j)は例えば、dn
(i,j)=|ai−bn j|等で表すことが可能であ
る。この場合、径路1を求めるについては、径路
選択のための拘束条件が設けられる。同図bは、
その径路拘束条件の一例である。即ち、点(i,
j)に至る前の点は、点(i+1,j+2)、点
(i+1,j+1)、点(i+2,j+1)であ
り、点(i,j)に至る径路は同図に示す径路に
限定される。
同図の径路上に示した数字は、その径路が選ば
れたときの重み係数を示す。本例のような、径路
の拘束を行う場合は、図aの格子グラフ上におい
て、任意の格子点間を結ぶ径路は、その選び方の
如何によらず荷重和は一定で、両点の間の入力パ
ターンの長さに等しくなる。従つて、この場合は
径路に沿う前記dn(i,j)の総和を荷重和で平
均する必要はなく総和そのものを入力パターンと
標準パターンとの距離とすることができる。具体
的な計算は次の漸化式を解くことによつて実行さ
れる。即ち
Dn(i,j)=minDn(i+1,j+2)+dn(i,j
)
Dn(i+1,j+1)+dn(i,j)
Dn(i+1,j+1)+dn(i,j)
Dn(i+2,j+1)+dn(i+1,j)+dn(i,j
)…(1)
をi=I,I−1,…,2,1,i=Jn,Jn−
1,…,2,1について初期値Dn(I,Jn)=dn
(I,Jn)のもとで解き、Dn(1,1)を両者の
距離とするものである。
径路の拘束条件を同図bのように選ぶことによ
り実際に選択可能な径路は同図aの斜線の内部に
制限される。このことは、パターンAとパターン
Bnは、同じ単語に対するものであるときは、そ
れ程ずれるはずはなく、異つた単語に対するもの
であるときは、無理な対応付をして両パターンの
距離値を不当に小さくする虞れのないようにする
という目的に合致したものである。
第3図、第4図は、DPマツチングによつて、
連続単語認識を行う場合の本発明の原理を説明す
る図である。第3図はk番目の音節境界を終点と
し、後述の範囲を始点とする入力パターンの部分
パターンと、V,CV,VV,VCV(Vは母音、C
は子音)等の音節(音声素片)標準パターンとの
DPマツチングの様子を説明する図であつて、横
軸を入力パターン、縦軸を標準パターンとする格
子グラフである。4はj=1の直線、n1,n2はそ
れぞれ音声素片標準パターンの1例を示すもので
あり、素片nのフレーム数をJnとしている。い
ま、前記入力の部分パターンと素片n1の標準パタ
ーンとマツチングする場合を考える。このとき、
第1図bの径路の拘束条件を適用すると第k番の
素片境界候補をSt(k)(k=0は語頭)とすれば、
点(St(k),Jn1)のマツチングの開始点に対して、
マツチングの範囲は直線5,6,4で囲まれる範
囲となり、点9〜点10の間の素片境界候補点を
k′とすれば、漸化式(1)の計算に従つて、k′〜kの
入力パターンの部分パターンと、n1の標準パター
ンRn1との素片累積照合距離Dn1(k′:k)はDn1
(k′:k)=Dn1(St(k′),1)で与えられる。こ
こ
に、点9は直線5と直線4との交点、点10は直
線6と直線4との交点であつて、直線5は傾き1/
2、直線6は傾き2である。この場合、第k素片
境界候補点を終点とする入力パターンの部分パタ
ーンと、標準パターンRn1とのマツチングにおい
て、始点k′の範囲は、点9〜点10の間というこ
とになり、漸化式1によりj=Jn,Jn−1,…,
1のそれぞれに対して、i′=max{St(k)−2(jn−
j),1},{St(k)−2(Jn−j)+1,1},…,
max{St(k)−〔(Jn−j+1)/2〕,1}について
Dn1(i′,j)を順次計算してゆくことにより、点
9、点10の間のk′に対するDn1(k′:k)=Dn1
(St(k′),1)は同時に求まる。ここで、max
{x,y}はx,yのうち大きい方の値を意味し、
〔x〕はxを越えない最大の整数を示す。またj
=jにおけるi′の範囲は、Dn(i′,j)はi′0に
おいては定義されていないので、上記の如くな
る。同様に、標準パターンRn2に対しては、点
(St(k),Jn2)を通る傾き1/2の直線2と直線4と
の交点7と、点(St(k),Jn2)を通る傾き2の直
線3と直線4との交点8の範囲のk′に対し、Dn2
(k′:k)=Dn2(St(k′),1)が求まる。入力の
各
境界候補フレームSt(k)において、n=1,2,
…,Nに対してこのようにして、Dn(k′:k)を
求める。
第4図は、標準パターンを単語としたとき、前
記素片標準パターンに対するのと同様な計算を行
う方法を説明している。即ち、入力の第St(k)フレ
ームを終点とし、後述の範囲を始点とする入力パ
ターンの部分パターンと単語wに対する標準パタ
ーンwとのDPマツチングの様子を説明してお
り、横軸を入力パターン、縦軸を標準パターンと
する格子グラフである。11はj′=1の直線、
w=Rs(w,1)、Rs(w,2),Rs(w,3)は単語標準パターンw
の一例を示している。ここでj′は単語標準パター
ンwの第1フレームから最終フレームまで通し
て付されたフレーム番号とし、s(w,)は単
語wの第番目の音声素片名を表す番号で、本例
では単語wは3つの素片名の系列からなり、単語
wの標準パターンwはこれに対応する3つの素
片標準パターンRs(w,1)、Rs(w,2),Rs(w,3)の結合した
ものとして表わされている。この場合も第2図b
の拘束条件を適用すると、単語wのフレーム数は
Jw=Js(w,1)+Js(w,2),+Js(w,3)であつて、点(St(k
),
Jw)のマツチングの開始点に対して、マツチン
グの範囲は、直線12,13,11で囲まれる範
囲となり、点14〜点15の間の境界候補番号を
k′とすれば、漸化式(1)と同様な計算に従つて、
k′〜kの入力パターンの部分パターンと、wと
の単語累積照合距離w(k′:k)が求まる。即
ち、この場合の漸化式はw
(i′,j′)=minw(i′+1,j′+2)+w(
i′,j′)w
(i′+1,j′+1)+w(i′,j′)w
(i′+1,j′+1)+w(i′,j′)w
(i′+2,j′+1)+w(i′+1,j′)+w
(i′,j′)…(2)
初期値 w(St(k),w)=w(St(k),w)
となり、w(i′,j′)をj′=w,w−1,…,
2,1の各々に対しi′を直線12〜13の範囲で
逐次計算してゆくことにより、w(k′:k)=
w(St(k′),1)として求めることができる。こ
こで、w(i′,j′)は入力の第i′フレームの入力
パターンの特徴ベクトルai′と単語wの標準パタ
ーンwの第j′フレームの特徴ベクトルw j′とのベ
クトル間距離であり前記dn(i,j)と同様の定
義w(i′,j′)=|ai′−bn j′|が用いられる。ま
た、
直線12は点(St(k),w)を通る傾き1/2の、直
線13は点(St(k),w)を通る傾き2の直線で
あり、点14は直線12と11との、点15は直
線13と11との交点である。次いでw∧=arg
minw〔w(k′:k)〕を計算する。argminx
in〔f(x)〕はf(x)を最小にするxを意味する。連
続発声された単語を認識するには、k=1,2,
…,Kについて以上の計算を行い、入力パターン
を個数、位置等に関して最適に分割し、分割され
たそれぞれの部分区間に対する前記最小の単語累
積照合距離を最小となし、そのときのそれぞれの
部分区間に対して求められた前記単語をそれぞれ
の区間に対する認識結果とすれば良いのである
が、単語数が厖大になつてくると前記方法で単語
累積照合距離w(k′:k)を求めるのは計算量
が厖大となる。そこで、本発明では、この単語累
積照合距離を求めるのに前記素片累積照合距離を
用いることによりこの計算量を大幅に削減してい
る点に特徴がある。即ち、本例においては、w
の最後の素片標準パターンRs(w,3)と入力パターン
のマツチングは直線12,13,16で囲まれる
領域について行われ、その結果w(St(k′),3)
は直線16上、直線12,13で挾まれる部分に
既に素片累積照合距離Ds(w,3)(k′:k)として求め
られている。k′は前記部分に含まれる素片境界候
補番号である。直線16はj′=w−Js(w,3)+1で
ある。単語標準パターンwの最終フレームから、
最後から2番目の素片Rs(w,2)までのマツチング
は、直線12,13,17で囲まれる領域につい
て行われ、その結果w(k′,2)は、直線17
の、直線12と13で挾まれる部分に求められる
ものであるが、これは、動的計画法の原理に従つ
てw
(k′,2)=min
k″
〔w(k″,3)+Ds(w,2)(k′:k″)〕
として求められる。素片累積照合距離Ds(w,2)(k′:
k″)および途中累積照合距離w(k′,3)は既
に求められているものである。ここで、直線17
はj′=w−Js(w,3)−Js(w,2)+1であつて、k′は直
線
17の直線12,13に挾まれる部分に含まれる
入力パターンの素片境界候補番号である。また
k″は例えばk′が直線17上の点20のときは、直
線16上の点であつて、点20を通り直線13に
平行な直線18と直線12に平行は直線19に挾
まれる部分と直線12と13に挾まれる部分の共
通部分の素片境界候補番号である。k′が点23の
ときも同様に、k″は、直線16の点であつて、
点23を通りそれぞれ直線13,12に平行な直
線21,22に挾まれる部分と、直線12,13
に挾まれる部分の共通部分の素片境界候補番号で
ある。これは、径路の拘束条件を図2bのように
したときは、点(St(k),w)から点20へ至る
マツチングの径路は直線12,13,18,19
で囲まれる平行四辺形の内部に限定され、点(St
(k),w)から点23へ至るマツチング径路は直
線12,13,21,22で囲まれる平行四辺形
の内部に限定されることを意味する。同様に、単
語標準パターンの最終フレームから、最後から3
番目までの素片(本例では単語wの最初の素片)
Rs(w,1)までのマツチングは直線11,12,13
で囲まれる領域について行われ、その結果w
(k′,1)は点14〜点15の部分について求め
られるものであるが、これも、動的計画法の原理
に従つて、w
(k′,1)=min
k″
〔w(k″,2)+Ds(w,1)(k′:k″)〕
として求められ、素片累積照合距離Ds(w,1)(k′:
k″)、単語途中累積照合距離w(k″,2)は既
に求められているものである。以上のようにし
て、素片累積照合距離を予め求めておき、これか
ら単語累積照合距離w(k′:k)をw(k′:
k)=w(k′,1)として求めることができる。
それぞれの単語wをそれを構成する素片標準パタ
ーンの結合で表わし、それと入力パターンと直接
マツチングする場合は各フレームにおいて単語数
だけのマツチング計算が必要であるのに比べて、
本発明の方法によれば入力の各フレームにおいて
はたかだか素片数のマツチングのみすれば良いか
ら数千語にも及ぶような大語彙単語に対する認識
の場合ははるかに少い計算量で、等価な結果が得
られるものである。
単語累積照合距離w(k′:k″)が求まると、
第St(k)フレームを最終フレームと仮定したとき、
第1フレームから第St(k)フレーム迄の最適の単語
列は動的計画法の原理により次の漸化式により求
めることができる。即ち、D〜(k)を入力の第1フレ
ームから第St(k)フレーム迄の部分パターンとそれ
に対する最適の単語列に対する特徴ベクトルの系
列との累積照合距離、B〜((k)を最後尾単語から1
つ手前の単語の最終境界候補番号、N〜(k)を最後尾
単語名とすれば、初期条件D〜(o)=0,B〜(o)=0と
して
D〜(k)=min
k′,w〔D〜(k′)+D〜w(k′:k)〕
B〜(k)=k^′
B〜(k)=k^′
N〜(k)=w^ (k^′,w^は上式を満足するk′,
w)…(3)
で与えられる。k=1,2,…,Kについて上記
計算を行えば、認識結果は次のように求まる。
最後の単語:N〜(K)
最後から2番目の単語:N〜(B(K))
最後から3番目の単語:N〜(B(B(K)))
…
最初の単語:N〜(B〜(B〜(…(B〜(K))…
)))
でB〜(B〜(…(B〜(K))…)))=0となつた
とき
終了する。
第5図はN〜(K),B〜(K)から上の単語列を
求めるフローチヤートである。
以上は単語数未知の場合の最適解を求める例で
あるが、単語数が既知の場合、オートマトン制御
による場合も式(3)の変更により簡単に本発明方法
を用いることができる。
単語数既知の場合は、Xを単語数、D〜x(k)を入
力パターンの第1フレームから第St(k)フレームま
での部分パターンと、x個の単語標準パターンを
最適に連結した標準パターンとの累積照合距離、
B〜x(k)を前記D〜x(k)に対するバツクポインタ、N〜
x
(k)を前記D〜x(k)に対する最後尾単語とすれば、式
(3)の漸化式は、初期条件D〜p(O)=0,B〜p(O
)
=0として
D〜x(k)=min
i′,w〔D〜x-1(k′)+w(k′:k)〕
B〜x(k)=k^′
B〜x(k)=k^′
N〜x(k)=w^ (k^′,w^は上式を満足するk′
,w)…(4)
によつて与えられる。
k=1,2,…,Kについて式(4)の計算を行え
ば、認識結果は次のように求まる。
最後の単語:N〜x(K)
最後から2番目の単語:N〜x-1(B〜x(K))
最後から3番目の単語:N〜x-2(B〜x-1(Bx
(K)))
…
最初の単語:N〜1(B〜2(B〜3(…(B〜x(K)
)
…)))
でB〜1(B〜2(B〜3(…(B〜x(K))…)))=
0とな
つて終了する。
第6図はN〜x(k),B〜x(k)から上の単語列を求める
フローチヤートである。
オートマトン制御の場合は次のようになる。
通常のオートマトンの認識問題と異なる点は、
時間を表わすフレーム番号も変数として入つてい
る点であり、しかも単に受理、拒否の出力でな
く、受理可能な度合(累積距離)が出力される点
である。
D〜q(k)を状態qで入力のSt(k)フレームで終端す
ると仮定したあらゆる単語列のうちの最小累積距
離、N〜q(k)をD〜q(k)に対応する単語列の最後尾単語
名、B〜q(k)をN〜q(k)の始点位置マイナス1(N〜q(
k)の
一つ前の単語の長終フレーム、即ちバツクポイン
タ)、Q〜q(k)をqへの状態遷移によつてD〜q(k)を満
たした状態名即ちΔを状態遷移規則とするときΔ
(Q〜q(k),N〜q(k))=qとするとき、次の漸化式を
解
くことで、オートマトン制御による解が得られ
る。即ち式(4)のxを状態qと読み代えることによ
つて、D〜q(k)を求める漸化式は次のようになる。
初期条件D〜p(O)=0,B〜p(O)=0として
D〜q(k)=min
k′,w,p
〔D〜p(k′)+w(k′:k)〕,q=Δ(p,n)
…(5)
をq=1,2,…,|s|−1について求め(s
は状態qの有限集合)、この式を満たすk′,w,
pをk^′,w^,p^とするとき、
N〜q(k)=w^,B〜q(k)=k^′,Q〜q(k)=p^
とする。k=Kまでこの計算を行えば、次のよう
にして最後尾の単語から逆順に単語が求まる。即
ち、
k=K,q=minqf D〜qf∈F(Fは最終状
態の集合)として
w^=N〜q(k)
B〜q(k)≠0なら、k=B〜q(k),q=Q〜q(k)と
し
てへ、B〜q(k)=0なら終了する。
第7図はフローチヤートである。
第1図は本発明の一実施例である。本実施例は
単語数未知の場合の例である。音声素片として
は、VCV音節、CV音節等を用いる場合について
説明する。この場合、音節の境界は母音定常部の
中心であるとする。100は音声信号端子であ
る。101は特徴抽出部で、フイルタバンク等で
構成されており、入力音声信号を特徴ベクトルの
系列a1,a2,…,aIに変換する。
116は母音標準パターン記憶部であつて、各
母音の標準パターンを記憶している。117は母
音認識部であつて、入力パターンの各フレームに
ついて母音標準パターン記憶部116の各母音標
準パターンと比較を行い、各フレームを母音とみ
なして母音認識を行う。これは例えば入力の各フ
レームと各母音標準パターンの距離を求めること
によつてできる。118は母音中心検出部であつ
て、母音認識部117の出力母音系列から、入力
パターンの各母音部の中心を検出する。例えば、
同一母音が連続する場合、その中心部をその母音
の母音中心とする等である。119は入力パター
ンから無音区間の検出、子音の大まかな分類を行
うものである。無音区間の検出は、入力パターン
から電力を求め、その値が予め定めた閾値より下
にあれば無音、上にあれば有音として判定でき
る。子音の大分類は、スペクトルの偏より等の周
知の方法を用いることにより、子音部の検出と摩
擦性、破裂性等の大まかな識別を行う。120は
特徴系列記憶部であつて、母音中心検出部118
で得られる母音中心の母音系列と、無音区間検
出・子音大分類部119で得られる無音、子音等
の系列を記憶するものである。102は素片標準
パターン記憶部であつて、CV,VCVのそれぞれ
に対応する特徴ベクトルの系列を標準パターンと
して記憶している。103は素片マツチング部で
あつて、入力パターンと、素片標準パターンとの
DPマツチングを行う。このとき例えば、k番目
の母音中心における処理をする場合を考えるとk
番目の母音中心部の母音の認識結果をV(k)とすれ
ば、k′k−1に対して入力パターンのフレーム
St(k′)からSt(k)までの部分パターンと、先行
母音が、V(k′)、後続母音がV(k)、子音が特徴系
列記憶部120で記憶されている第k′番の母音中
心と第k番の母音中心の間の子音大分粒結果を満
たすVCV音節標準パターンRnとのDPマツチング
を行い、素片累積照合距離Dn(k′:k)を計算す
る。ここにk′は前記第3図において説明し各音節
標準パターンに対して決定される三角形の底辺の
上に存在するもののみを考慮すれば良い。ただ
し、k=1,2,…,kに対し、max{St(k)−2
(Jn−1),1}St(k′)max{St(k)−〔(Jn−
j
+1)/2〕,1},
nは前記条件を満たすnである。
また、Dn(k′:k)はSt(k′)0においては定
義されていないので、St(k′)の範囲は第2図b
の径路の拘束条件を用いるときはここに示した範
囲となる。104は素片マツチング部で計算され
た素片累積照合距離Dn(k′:k)を記憶する部分
である。105は単語辞書であつて、認識すべき
各単語wが音声素片名の系列として表わされたも
のが記憶されている。
121は候補単語判定部であつて、単語辞書1
05から読み出される単語がマツチングすべき単
語か否かを特徴系列記憶部120の記憶内容と比
較し、予め候補単語を予備選択するものである。
今、母音中心の検出は挿入はあつても脱落はない
ものとし、挿入は2つ続けては生じないものとす
れば、第8図aに示すマツチング径路を用いて、
特徴系列同志のマツチングをとり、候補単語の選
出を行うものである。即ち、素片をVCV音節と
すれば単語辞書の単語の第+1音節の特徴が、
入力パターンの第k′番の母音中心から第k′+1あ
るいは第k′+2番の母音中心までの特徴に含まれ
れば、両者の距離dd(k′,)=0とし、含まれ
なければdd(k′,)=1とし、漸化式
DD(k′,)=minDD(k′+1,+1)+dd(k′,
)
DD(k′+2,+1)+dd(k′,) …(6)
初期値DD(k,Lw)=0
を=Lw−1,Lw−2,…,1,0について繰
り返して計算し、k−2Lwk′k−Lwの範囲で
DD(k′:o)の値が0であるか否かを判定し、
DD(k′,o)=0であればその単語は候補単語で
あり、単語マツチングの対象として採用しDD
(k′,o)≠0であれば候補外の単語であるとし
て単語マツチングの対象から省くものである。即
ち、漸化式(6)は、DPマツチングの径路の正規化
係数が標準パターンの音節数に等しくなるもので
あつて、入力側の端点自由のマツチングを行つて
いることになる。これを図的に説明すると、3音
節の単語に対する例としてマツチングの範囲は第
8図bの傾き1/2の直線122と傾き2の直線1
23で挾まれる領域となる。但し、同図におい
て、横軸は入力パターンの母音中心番号列、縦軸
はマツチングすべき単語の音節標準パターン列で
あつて、時間軸を伸縮することによつて、これら
は全て同じ間隔になるように画いてある。この図
においては、始端が、k−6〜k−3の何かから
終端が、k迄の特徴系列の中に、単語wの可能性
のある特徴系列があればDD(k′,o)=0なる
k′が存在し、その可能性がない場合は、DD(k′,
o)=0なるk′は存在しないことになる。ここで、
Lwは単語wの素片数である。
106は単語マツチング部であつて、候補単語
判定部121で選ばれた候補単語の各wに対し
て、入力パターンの第k′母音中心から第k母音中
心までの部分パターンと単語標準パターンw=
Rs(w,1),Rs(w,2),…Rs(w,Lw)とのマツチングを、前
記素片累積照合距離を基に行い、単語累積照合距
離w(k′:k)を計算する部分である。
本実施例の場合は、第9図にその例が図解され
る。これは第8図と同様に、入力パターンの母音
中心間の長さ、標準パターンの長さを同じ長さに
なるように、それぞれの軸を伸縮して画いてあ
り、3音節の単語とマツチングする場合である。
入力パターンの母音中心番号kと標準パターンの
音節番号の対応は第9図aで表わされるから、
入力パターンと標準パターンのマツチングパスは
点p=(k,3)を通り、傾き1/2と傾き1の直線
125,126で挾まれる範囲に限定される。こ
の場合、線分Aを=1なる直線の直線125と
直線126で挾まれる部分、線分Bを直線=2
の直線125と直線126で挾まれる部分とし、
rを線分A上の点、qを線分B上の点とすれば、
点pから点rまでの最小累積照合距離は点pから
線分B上の点qまでの最小累積照合距離と点qか
ら点rまでの最小累積照合距離の和を点qに関し
て最小にしたときの最小値とすることができる。
この場合、前記の説明から点pから点qまで、点
qから点rまでのそれぞれ最小累積照合距離が既
に求まつているから、点pから点rまでの最小累
積照合距離は、
w(k−3,1)=
min
min
k″=k−2,k−1〔w(k″,2)+Ds(w,2)(k−
3:k)〕
として求めることができる。w(k,3)=0と
して、=Lw−1,Lw−2,…,1,0につい
てこの操作を順次繰返すことにより入力パターン
の母音中心k−6〜k−3を始点、kを終点とす
る部分パターンと単語wのマツチング距離はw
(k′:k)=w(k′,1)として求まる。但し、
k−6k′k−3である。一般に、=にお
いてw(k′,)を計算するk′の範囲は、max
{k−2(Lw−),o}k′max{k−(Lw−
,o}となる。ここで、w(k′:k)はk′
−1では定義されていないので、k′の範囲はここ
に示したようになる。
107は単語マツチング結果記憶部であつて単
語累積照合距離w(k′:k)を記憶する部分で
ある。108は終端累積距離計算部であつて、単
語マツチング結果記憶部107の内容と終端累積
距離記憶部108の内容から漸化式3に従つて、
D〜(k),N〜(k),B〜(k)を計算する。終端累積距離
記憶
部109は、終端累積距離計算部108で計算さ
れた終端累積距離D〜(k)を必要がなくなるまで記憶
する。このD〜(k)は終端累積距離計算部108にお
ける漸化式3の計算に用いられる。110はバツ
クポインタ記憶部であつて、終端累積距離計算部
108で計算されたバツクポインタB〜(k)を記憶す
る。111は最後尾単語記憶部で、終端累積距離
記憶部109で求められた第k母音中心における
最後尾単語を記憶する。112は音声区間検出部
であつて、入力信号の大きさ等から音声区間を判
定するもので、この音声区間検出部112が音声
入力が開始されたことを検出すると、母音中心計
数部113は母音中心毎に計数を始める。前記の
処理はk母音中心についての処理であつたが、こ
の母音中心計数部113の計数値がこのkを設定
している。従つて、前記と同様の処理が母音中心
が1進む毎に行われることになる。母音中心計数
部113は音声区間が検出されると計数を始め、
音声区間が終了するとリセツトされる。最後尾単
語記憶部111、バツクポインタ記憶部100に
は、N〜(k),B〜(k)がk=1,2,…,Kについて記
憶されることになる。セグメンテーシヨン部11
4はバツクポインタ記憶部110に対し、所定の
バツクポインタを読出すべき命令を発するもので
ある。即ち、セグメンテーシヨン部、114がk
なる値をバツクポインタ記憶部110に発する
と、バツクポインタ記憶部110からはバツクポ
インタB〜(k)が読出される。セグメンテーシヨン部
114はバツクポインタ記憶部100からB〜(k)な
る値を受け取ると、その同じ値をバツクポインタ
記憶部110に発する。従つて、音声区間検出部
112が音声入力の終了が検知すると、母音中心
計数部113の最終値Kがセグメンテーシヨン部
114に供給され、セグメンテーシヨン部114
は先ずKなる値をバツクポインタ記憶部110に
発する。以後、前記、説明の動作に従つて、バツ
クポインタ記憶部110B(K),B(B(K)),
…,Oなる出力が順次得られることになる。これ
らの値は、最後から2番目の単語の終りのフレー
ム、同3番目の終りのフレーム、同4番目のフレ
ーム、…というものであり、N〜(k)はkフレームで
終る単語であつたから、この値をそのまま最後尾
単語記憶部111に与えると、最後の単語から逆
の順序で認識結果が得られることになる。この順
序を逆に(あたりまえの順序に)するには、順序
の変換をバツクポインタ記憶部110の出力か、
最後尾単語記憶部111の出力に対して行えばよ
い。
第10図は、以上の実施例の動作をプログラム
で表現したものであり、ソフトウエアで実現する
場合もこれに従えばよい。なお第10図におい
て、
INDUSTRIAL APPLICATION FIELD The present invention relates to a speech recognition device that recognizes continuously uttered words and phrases in units of speech segments such as syllables. 2. Description of the Related Art If voice, which is the most natural means of generating information for humans, could be used as an input means for a human-machine system, the effect would be enormous. Conventionally, speech recognition devices based on a specific speaker registration method have been put into practical use. That is, a speaker who intends to use a recognition device converts all the words to be recognized into a series of feature vectors using his/her own voice and registers them as standard patterns in a word dictionary, and then uses the voice uttered during recognition. is similarly converted into a series of feature vectors, which word in the word dictionary is closest is calculated according to predetermined criteria, and the most similar word is taken as the recognition result. . However, this method is good when the number of recognized words is small, but as the number of words increases to hundreds or thousands of words, the following three problems become impossible to ignore. (1) The burden on speakers during registration increases significantly. (2) The time required to calculate the similarity or distance between the voice uttered and the standard pattern during recognition increases significantly, and the response speed of the recognition device becomes slow. (3) The memory required for the word dictionary becomes very large. As a method to avoid the above-mentioned drawbacks, there is a method in which the unit of recognition is a consonant + vowel or a monosyllable of a vowel (hereinafter expressed as CV and V, respectively. C means a consonant and V means a vowel). That is, monosyllables are registered as a series of feature vectors as a standard pattern, and input speech converted into a series of feature vectors during recognition is converted into a series of monosyllables by matching with the monosyllable standard pattern. It is something to do. In Japanese, the number of monosyllables is at most
There are 101 types, and monosyllables correspond to kana characters, so this method converts (recognizes) any Japanese word or sentence into a monosyllable string.
This will solve all of the problems (1) to (3) above. However, problems in this case include articulatory combination and segmentation. Articulatory coupling is a phenomenon in which when syllables are uttered in succession, each syllable is influenced by the syllables before and after it, and the spectral structure changes depending on the syllables connected before and after it. Segmentation is the process of dividing continuously uttered speech into monosyllable units, but a definitive method for doing this reliably has not yet been found. In order to solve these two problems, the current practice is to separate each single syllable and pronounce it, and some devices are in practical use. However, uttering monosyllables discretely is unnatural and puts stress on the speaker. Problems to be Solved by the Invention The present invention avoids the uncertainty of segmentation for the continuously uttered speech, and also
It is an object of the present invention to provide a continuous speech recognition device capable of recognizing continuously uttered words or phrases. Means for Solving the Problems The present invention is defined as a feature extraction means for converting an input audio signal obtained by continuously uttering words, phrases, etc. into a series of feature vectors, and vowels, consonants, or combinations thereof. standard pattern storage means for storing a series of feature vectors corresponding to each of the speech segments in association with the speech segment name;
determined by an arbitrary or predetermined various combinations of a unit boundary candidate detection means for detecting unit boundaries with respect to an input pattern, each of the standard patterns, and the unit boundary candidates detected from the input pattern. a segment matching means for calculating the distance (similarity) between the two by matching with a subinterval (first subinterval), and representing each word, phrase, etc. to be recognized as a sequence of the phonetic segment names; words
A word/phrase dictionary that stores phrases, etc., and distances ( similarity), the first sub-interval group included in the second sub-interval is such that adjacent sections are continuous so as to correspond to the sequence of segment names specified by the word/phrase dictionary. By optimally determining as follows, the sum of the distances (similarities) corresponding to the start point and end point of each first subinterval and the segment name of that subinterval is set to the minimum (maximum), and the obtained minimum value ( a word/phrase matching means having a function of outputting a maximum value (maximum value) as the distance of the second partial interval to each word/phrase, and optimally determining the second partial interval group so that adjacent intervals are continuous. According to the second
A sequence in which the sum of the distances (similarities) corresponding to the start and end points of each subinterval and the word/phrase name in that subinterval is the minimum (maximum), and the word/phrase string at that time is determined as the recognition result. This is a continuous speech recognition device equipped with word/phrase determination means. Effect The present invention has the above-described configuration, converts an input speech signal obtained by continuously uttering words, phrases, etc. into a series of feature vectors, and converts the input speech signal into a series of feature vectors into speech elements defined as vowels, consonants, or combinations thereof. A standard pattern in which a series of feature vectors corresponding to each of The distance (similarity) between the two is calculated by matching with the subinterval (first subinterval) determined by various combinations, and each word, phrase, etc. to be recognized is identified by the phoneme name. Distance (similarity) between each word or phrase to be recognized, such as a word or phrase expressed as a sequence, and various partial intervals (second partial intervals) of the arbitrary or predetermined segment boundary candidates of the input pattern. Optimize the first subinterval group included in the second subinterval so that adjacent sections are continuous, so as to correspond to the sequence of segment names specified by the word/clause. By determining, the sum of the distances (similarities) corresponding to the start point and end point of each first subinterval and the segment name of that subinterval is the minimum (maximum), and the obtained minimum value (maximum value) is By determining the distance of the second sub-interval to each word/phrase and optimally determining the second sub-interval group so that adjacent sections are continuous, the starting point and end point of each second sub-interval and their The sum of the distances (similarities) corresponding to the word/phrase names in the partial section is set as the minimum (maximum), and the word/phrase string at that time is determined as the recognition result. Examples Hereinafter, the word "word" will also represent the word "bunsetsu." Further, "similarity" will be explained using "distance" as a representative. That is, a small distance means a large degree of similarity. First, DP matching, which is the basis of the present invention, will be described. FIG. 2 is a lattice graph explaining DP matching when recognizing discrete words. That is, input pattern A = a 1 , a 2 ... a i ... a I and standard pattern B n =
A case is shown in which the distance to b n 1 , b n 2 ... b n j ... b n j n is calculated. The horizontal axis shows the input pattern, the vertical axis shows the standard pattern, and 1 is a curve showing the correspondence between the feature vectors of both. DP matching minimizes the weight average along this path of the distance d n (i, j ) between a i and b n j that are matched by this path by optimally determining this path, and This value is used as the distance between the two, and this calculation is performed efficiently. d n (i, j) is, for example, d n
It can be expressed as (i, j)=|a i −b n j |, etc. In this case, when determining route 1, constraint conditions for route selection are provided. Figure b is
This is an example of the route constraint conditions. That is, the point (i,
The points before reaching point (i, j) are point (i+1, j+2), point (i+1, j+1), and point (i+2, j+1), and the route to point (i, j) is limited to the route shown in the figure. Ru. The numbers shown on the routes in the figure indicate the weighting coefficients when that route is selected. When constraining paths as in this example, on the lattice graph in Figure A, the sum of weights for a path connecting arbitrary lattice points is constant regardless of how it is selected, and the sum of weights between both points is constant. Equal to the length of the input pattern. Therefore, in this case, there is no need to average the sum of d n (i, j) along the path using a weighted sum, and the sum itself can be used as the distance between the input pattern and the standard pattern. The specific calculation is performed by solving the following recurrence formula. That is, D n (i, j)=minD n (i+1, j+2)+d n (i, j
) D n (i+1, j+1)+d n (i, j) D n (i+1, j+1)+d n (i, j) D n (i+2, j+1)+d n (i+1, j)+d n (i, j
)...(1) as i=I, I-1,..., 2, 1, i=J n , J n −
Initial value D n (I, J n )=d n for 1,...,2,1
(I, J n ) and let D n (1, 1) be the distance between the two. By selecting the route constraint conditions as shown in FIG. 2B, the routes that can actually be selected are limited to those within the hatched area in FIG. This means that pattern A and pattern
When B n is for the same word, it should not deviate that much, and when it is for different words, there is no risk of making an unreasonable correspondence and making the distance value of both patterns unduly small. This is consistent with the purpose of doing so. Figures 3 and 4 show that by DP matching,
FIG. 3 is a diagram illustrating the principle of the present invention when performing continuous word recognition. Figure 3 shows partial patterns of the input pattern whose ending point is the k-th syllable boundary and whose starting point is the range described below, and V, CV, VV, VCV (V is a vowel, C
is a consonant) and other syllables (speech elements) standard patterns.
It is a diagram illustrating the state of DP matching, and is a grid graph in which the horizontal axis is the input pattern and the vertical axis is the standard pattern. 4 is a straight line with j=1, n 1 and n 2 each represent an example of a standard speech unit pattern, and the number of frames of unit n is J n . Now, let us consider the case of matching the partial pattern of the input with the standard pattern of segment n1 . At this time,
Applying the path constraint shown in Figure 1b, if the k-th segment boundary candidate is St(k) (k = 0 is the beginning of a word), then
For the starting point of matching of point (St(k), J n1 ),
The matching range is the range surrounded by straight lines 5, 6, and 4, and the segment boundary candidate points between points 9 and 10 are
If k′, then according to the calculation of recurrence formula (1), the segment cumulative matching distance D n1 ( k ′ : k) is D n1
It is given by (k′:k)=D n1 (St(k′), 1). Here, point 9 is the intersection of straight line 5 and straight line 4, point 10 is the intersection of straight line 6 and straight line 4, and straight line 5 has a slope of 1/
2. The straight line 6 has a slope of 2. In this case, in matching the partial pattern of the input pattern whose end point is the k-th unit boundary candidate point and the standard pattern R n1 , the range of the starting point k' is between points 9 and 10, and gradually According to Formula 1, j=J n , J n −1,…,
1, i′=max{St(k)−2(j n −
j), 1}, {St(k)−2(J n −j)+1,1},…,
Regarding max{St(k)−[(J n −j+1)/2], 1}
By sequentially calculating D n1 (i′, j), D n1 (k′:k)=D n1 for k′ between points 9 and 10
(St(k′), 1) can be found at the same time. Here, max
{x, y} means the larger value of x, y,
[x] indicates the largest integer not exceeding x. Also j
The range of i' at =j is as above because D n (i', j) is not defined at i'0. Similarly, for the standard pattern R n2 , the intersection point 7 of the straight line 2 with a slope of 1/2 passing through the point (St(k), J n2 ) and the straight line 4, and the point (St(k), J n2 ) For k ′ in the range of intersection 8 between straight line 3 with slope 2 and straight line 4,
(k′:k)=D n2 (St(k′), 1) is found. In each input boundary candidate frame St(k), n=1, 2,
..., N in this manner to find D n (k':k). FIG. 4 explains a method for performing calculations similar to those for the segment standard pattern when the standard pattern is a word. In other words, it explains the state of DP matching between a standard pattern w for a word w and a partial pattern of an input pattern whose end point is the input St(k)th frame and whose start point is the range described below, and the horizontal axis represents the input pattern. , is a grid graph with a standard pattern on the vertical axis. 11 is the straight line of j′=1,
w = R s(w,1) , R s(w,2) , R s(w,3) are word standard patterns w
An example is shown. Here, j' is the frame number assigned from the first frame to the last frame of the word standard pattern w , and s(w,) is the number representing the name of the th phoneme segment of the word w. In this example, The word w consists of a series of three segment names, and the standard pattern w of the word w is the corresponding three segment standard patterns R s(w,1) , R s(w,2) , R s(w ,3) . In this case too, Figure 2b
Applying the constraint, the number of frames for word w is J w = J s(w,1) + J s(w,2) , +J s(w,3) , and the point (St(k
),
With respect to the matching starting point of J w ), the matching range is the range surrounded by straight lines 12, 13, and 11, and the boundary candidate number between points 14 and 15 is
If k′, then according to calculation similar to recurrence formula (1),
The cumulative word matching distance w (k':k) between the subpatterns of the input pattern k' to k and w is determined. That is, the recurrence formula in this case is w (i', j') = min w (i'+1, j'+2) + w (
i', j') w (i'+1, j'+1)+ w (i', j') w (i'+1, j'+1)+ w (i', j') w (i'+2, j′+1)+ w (i′+1,j′)+ w
(i′, j′)…(2) Initial value w (St(k), w )= w (St(k), w )
Then, w (i′, j′) is j′= w , w −1,…,
By sequentially calculating i′ for each of 2 and 1 within the range of straight line 12 to 13, w (k′:k)=
It can be obtained as w (St(k'), 1). Here, w (i′, j′) is the intervector distance between the feature vector a i ′ of the input pattern of the i′th frame of input and the feature vector w j ′ of the j′th frame of the standard pattern w of word w. The same definition as d n (i, j) above is used: w (i', j')=|a i '−b n j '|. Also,
Straight line 12 passes through the point (St(k), w ) with a slope of 1/2, straight line 13 passes through the point (St(k), w ) with a slope of 2, and the point 14 is the line between straight lines 12 and 11. , point 15 is the intersection of straight lines 13 and 11. Then w∧=arg
Calculate minw [ w (k′:k)]. argminx
in[f(x)] means x that minimizes f(x). To recognize continuously uttered words, k=1, 2,
..., K, perform the above calculations, optimally divide the input pattern in terms of number, position, etc., minimize the minimum cumulative word matching distance for each divided partial interval, and then It is sufficient to use the words obtained for each section as the recognition results for each section, but as the number of words increases, it becomes difficult to obtain the cumulative word matching distance w (k':k) using the above method. The amount of calculation becomes enormous. Therefore, the present invention is characterized in that the amount of calculation is significantly reduced by using the segment cumulative matching distance to obtain the word cumulative matching distance. That is, in this example, w
The matching of the last segment standard pattern R s (w, 3) and the input pattern is performed on the area surrounded by straight lines 12, 13, and 16, and as a result w (St (k'), 3)
has already been determined as the segment cumulative matching distance D s(w,3) (k':k) on the straight line 16 in the portion sandwiched by the straight lines 12 and 13. k' is a segment boundary candidate number included in the portion. Straight line 16 is j′= w −J s(w,3) +1. From the final frame of word standard pattern w ,
Matching up to the penultimate segment R s(w,2) is performed on the area surrounded by straight lines 12, 13, and 17, and as a result, w (k′,2) is
This is required for the part sandwiched by straight lines 12 and 13, and according to the principles of dynamic programming, w (k', 2) = min k'' [ w (k'', 3) +D s(w,2) (k′:k″)]. The segment cumulative matching distance D s(w,2) (k′:
k'') and the intermediate cumulative matching distance w (k', 3) have already been found.Here, the straight line 17
is j′= w −J s(w,3) −J s(w,2) +1, and k′ is the segment boundary of the input pattern included in the part of straight line 17 sandwiched by straight lines 12 and 13. This is the candidate number. Also
For example, when k' is a point 20 on the straight line 17, k'' is a point on the straight line 16, and is the part that passes through the point 20 and is sandwiched by the straight line 18 that is parallel to the straight line 13 and the straight line 19 that is parallel to the straight line 12. is the segment boundary candidate number of the common part of the part sandwiched by straight lines 12 and 13. Similarly, when k' is point 23, k'' is a point on straight line 16, and
The part that passes through point 23 and is sandwiched by straight lines 21 and 22 that are parallel to straight lines 13 and 12, respectively, and the straight lines 12 and 13
This is the segment boundary candidate number of the common part of the parts sandwiched by. This means that when the path constraint conditions are set as shown in Figure 2b, the matching path from point (St(k), w ) to point 20 is a straight line 12, 13, 18, 19.
The point (St
This means that the matching path from (k), w ) to point 23 is limited to the inside of the parallelogram surrounded by straight lines 12, 13, 21, and 22. Similarly, from the last frame of the word standard pattern, 3 from the end
Fragment up to the th elemental fragment (in this example, the first elemental fragment of word w)
Matching up to R s(w,1) is straight lines 11, 12, 13
The result is w
(k', 1) is found for the part from point 14 to point 15, and this also follows the principle of dynamic programming as w (k', 1)=min k'' [ w (k ″, 2) + D s(w,1) (k′:k″)], and the segment cumulative matching distance D s(w,1) (k′:
k'') and the cumulative matching distance w (k'', 2) in the middle of a word have already been determined. In the above manner, the unitary cumulative matching distance is calculated in advance, and from this, the word cumulative matching distance w (k′:k) is calculated by w (k′:
k) = w (k', 1).
If each word w is represented by a combination of its component standard patterns and then directly matched with the input pattern, matching calculations for the number of words are required in each frame.
According to the method of the present invention, it is only necessary to match the number of elements in each frame of input, so when recognizing a large vocabulary of several thousand words, the amount of calculation is much smaller, and the equivalent It's something that gets results. Once the word cumulative matching distance w (k′:k″) is found,
Assuming that the St(k)th frame is the final frame,
The optimal word string from the first frame to the St(k)th frame can be found using the following recurrence formula based on the principle of dynamic programming. That is, D~(k) is the cumulative matching distance between the input partial pattern from the first frame to the St(k)th frame and the feature vector series for the optimal word string for it, and B~((k) is the last 1 from tail word
If the final boundary candidate number of the previous word, N~(k), is the last word name, initial conditions D~(o)=0, B~(o)=0, and D~(k)=min k ′, w [D~(k')+D~ w (k':k)] B~(k)=k^'B~(k)=k^' N~(k)=w^ (k^' , w^ is k′ that satisfies the above formula,
w)…(3) is given. If the above calculation is performed for k=1, 2, . . . , K, the recognition result is obtained as follows. Last word: N~(K) Second to last word: N~(B(K)) Third to last word: N~(B(B(K)))... First word: N~( B~(B~(...(B~(K))...
))) When B~(B~(...(B~(K))...)))=0, the process ends. FIG. 5 is a flowchart for determining the upper word string from N~(K) and B~(K). The above is an example of finding the optimal solution when the number of words is unknown. However, when the number of words is known, the method of the present invention can be easily used by changing equation (3) even when using automaton control. When the number of words is known, X is the number of words, and D ~ x (k) is the standard that optimally connects the subpatterns from the first frame to the St(k)th frame of the input pattern and x word standard patterns. Cumulative matching distance with the pattern,
B~ x (k) is the back pointer for the D~ x (k), N~
x
If (k) is the last word for the above D~ x (k), then the formula
The recurrence formula (3) is given by the initial conditions D~ p (O)=0, B~ p (O
)
= 0, D~ x (k)=min i', w [D~ x-1 (k') + w (k':k)] B~ x (k)=k^' B~ x (k) =k^' N~ x (k)=w^ (k^', w^ are k' that satisfy the above formula
, w)…(4). If equation (4) is calculated for k=1, 2, . . . , K, the recognition result is obtained as follows. Last word: N~ x (K) Second to last word: N~ x-1 (B~ x (K)) Third last word: N~ x-2 (B~ x-1 (B x
(K))) … First word: N~ 1 (B~ 2 (B~ 3 ) (…(B~ x (K)
)
…))) B~ 1 (B~ 2 (B~ 3 (…(B~ x (K))…))) =
It becomes 0 and ends. FIG. 6 is a flowchart for determining the upper word string from N~ x (k) and B~ x (k). In case of automaton control, it is as follows. The difference from normal automaton recognition problems is that
The frame number representing time is also included as a variable, and the output is not just acceptance or rejection, but the degree of acceptability (cumulative distance). Let D~ q (k) be the minimum cumulative distance among all word strings assuming that it ends at the input St(k) frame in state q, and N~ q (k) be the word string corresponding to D~ q (k). The last word name of B~ q (k) is the starting point position of N~ q (k) minus 1 (N~ q (
k), the long end frame of the previous word (i.e., the back pointer), and the state name that satisfies D~ q (k), that is, Δ, by the state transition of Q~ q (k) to q, is the state transition rule. When Δ
When (Q~ q (k), N~ q (k)) = q, a solution by automaton control can be obtained by solving the following recurrence formula. That is, by replacing x in equation (4) with state q, the recurrence equation for finding D~ q (k) becomes as follows. Initial conditions D~ p (O) = 0, B~ p (O) = 0, D~ q (k) = min k', w, p [D~ p (k') + w (k': k) ], q=Δ(p,n)
...(5) for q = 1, 2, ..., |s|-1 (s
is a finite set of states q), k′, w,
When p is k^', w^, p^, N~ q (k)=w^, B~ q (k)=k^', Q~ q (k)=p^. If this calculation is performed until k=K, words can be found in reverse order starting from the last word as follows. That is, k=K, q=minq f D~ qf ∈F (F is the set of final states) w^=N~ q (k) B~ q If (k)≠0, then k=B~ q (k ), q=Q~ q (k), and if B~ q (k)=0, the process ends. FIG. 7 is a flowchart. FIG. 1 shows an embodiment of the present invention. This embodiment is an example in which the number of words is unknown. A case will be explained in which VCV syllables, CV syllables, etc. are used as speech segments. In this case, the syllable boundary is assumed to be the center of the vowel stationary part. 100 is an audio signal terminal. Reference numeral 101 denotes a feature extraction unit, which is composed of a filter bank and the like, and converts an input audio signal into a series of feature vectors a 1 , a 2 , . . . , a I . Reference numeral 116 is a vowel standard pattern storage unit that stores standard patterns for each vowel. A vowel recognition unit 117 compares each frame of the input pattern with each vowel standard pattern in the vowel standard pattern storage unit 116, and performs vowel recognition by regarding each frame as a vowel. This can be done, for example, by determining the distance between each frame of the input and each vowel standard pattern. A vowel center detection unit 118 detects the center of each vowel part of the input pattern from the output vowel sequence of the vowel recognition unit 117. for example,
When the same vowel is consecutive, the center of the same vowel is set as the vowel center of that vowel. Reference numeral 119 detects silent intervals from the input pattern and roughly classifies consonants. To detect a silent section, the power is determined from the input pattern, and if the value is below a predetermined threshold, it is determined that there is no sound, and if it is above a predetermined threshold, it is determined that there is sound. For consonant classification, consonant parts are detected and roughly identified as fricative, plosive, etc., using well-known methods such as spectral polarization. 120 is a feature series storage unit, which includes a vowel center detection unit 118
It stores the vowel series centered on vowels obtained in the above, and the series of silences, consonants, etc. obtained by the silent section detection/consonant classification section 119. Reference numeral 102 denotes a segment standard pattern storage unit which stores a series of feature vectors corresponding to each of CV and VCV as a standard pattern. Reference numeral 103 denotes a segment matching unit, which matches the input pattern and the segment standard pattern.
Perform DP matching. In this case, for example, if we consider the case where processing is performed at the center of the kth vowel, k
If the recognition result of the vowel in the center of the th vowel is V(k), then the frame of the input pattern for k′k−1 is
The partial pattern from St(k') to St(k), the preceding vowel is V(k'), the following vowel is V(k), and the k'th consonant is stored in the feature sequence storage unit 120. DP matching is performed with the VCV syllable standard pattern R n that satisfies the consonant size division result between the vowel center of and the kth vowel center, and the unitary cumulative matching distance D n (k′:k) is calculated. Here, it is sufficient to consider only k' existing on the base of the triangle explained in FIG. 3 and determined for each syllable standard pattern. However, for k=1, 2,...,k, max{St(k)−2
(J n −1), 1}St(k′)max{St(k)−[(J n −
j
+1)/2], 1}, n is n that satisfies the above condition. Also, since D n (k':k) is not defined in St(k')0, the range of St(k') is
When using the constraint conditions for the path, the range is as shown here. Reference numeral 104 denotes a part that stores the segment cumulative matching distance D n (k':k) calculated by the segment matching section. Reference numeral 105 is a word dictionary in which each word w to be recognized is expressed as a series of phoneme names. Reference numeral 121 is a candidate word determination unit, which is a word dictionary 1.
05 is a word to be matched or not, it is compared with the memory contents of the feature series storage unit 120, and candidate words are preliminarily selected.
Now, when detecting the center of a vowel, assuming that there is an insertion but no omission, and that two insertions do not occur in succession, using the matching path shown in Figure 8a,
This method selects candidate words by matching feature series. In other words, if a segment is a VCV syllable, the characteristics of the +1st syllable of a word in the word dictionary are
If it is included in the features from the k'th vowel center to the k'+1 or k'+2th vowel center of the input pattern, the distance between the two is dd(k',) = 0, and if it is not included, dd (k′,)=1, recurrence formula DD(k′,)=minDD(k′+1,+1)+dd(k′,
) DD (k'+2, +1) + dd (k',) ...(6) Repeat the initial value DD (k, L w ) = 0 for = L w -1, L w -2, ..., 1, 0. Calculate and in the range k−2L w k′k−L w
Determine whether the value of DD(k′:o) is 0 or not,
If DD (k′, o) = 0, the word is a candidate word, and DD is selected as a target for word matching.
If (k', o)≠0, the word is considered to be a non-candidate word and is omitted from the word matching target. That is, in recurrence formula (6), the normalization coefficient of the path of DP matching is equal to the number of syllables of the standard pattern, and matching is performed with free endpoints on the input side. To explain this graphically, as an example for a three-syllable word, the matching range is the straight line 122 with a slope of 1/2 and the straight line 1 with a slope of 2 in Figure 8b.
This is the area sandwiched by 23. However, in the figure, the horizontal axis is the vowel center number sequence of the input pattern, and the vertical axis is the syllable standard pattern sequence of the word to be matched, and by expanding or contracting the time axis, they all become the same interval. It is depicted as follows. In this figure, if there is a feature sequence with a possibility of word w among the feature sequences whose starting point is something between k-6 and k-3 and ending with k, then DD(k', o) =0
If k′ exists and there is no possibility, then DD(k′,
k′ with o)=0 does not exist. here,
L w is the number of prime pieces of word w. 106 is a word matching unit which, for each candidate word w selected by the candidate word determination unit 121, calculates a partial pattern from the k'th vowel center to the kth vowel center of the input pattern and a word standard pattern w =
Matching with R s(w,1) , R s(w,2) , ...R s(w,Lw) is performed based on the segment cumulative matching distance, and word cumulative matching distance w (k′:k ) is the part that calculates. In the case of this embodiment, an example is illustrated in FIG. Similar to Figure 8, this is drawn by expanding and contracting each axis so that the length between the vowel centers of the input pattern and the length of the standard pattern are the same, and matching with the three-syllable word. This is the case.
The correspondence between the vowel center number k of the input pattern and the syllable number of the standard pattern is shown in Figure 9a, so
The matching path between the input pattern and the standard pattern passes through the point p=(k, 3) and is limited to the range sandwiched by straight lines 125 and 126 with a slope of 1/2 and a slope of 1. In this case, line segment A is the part sandwiched by straight line 125 and straight line 126 where = 1, and line segment B is the straight line = 2.
The part sandwiched between the straight line 125 and the straight line 126,
If r is a point on line segment A and q is a point on line segment B, then
The minimum cumulative matching distance from point p to point r is when the sum of the minimum cumulative matching distance from point p to point q on line segment B and the minimum cumulative matching distance from point q to point r is minimized with respect to point q. can be the minimum value of
In this case, the minimum cumulative matching distance from point p to point q and from point q to point r have already been found from the above explanation, so the minimum cumulative matching distance from point p to point r is w (k −3,1)=min min k″=k−2,k−1 [ w (k″,2)+D s(w,2) (k−
3:k)]. By sequentially repeating this operation for = L w -1, L w -2, ..., 1, 0 with w (k, 3) = 0, the vowel centers k-6 to k-3 of the input pattern are the starting point, and k The matching distance between the partial pattern and word w with end point is w
It is found as (k':k)= w (k', 1). however,
k-6k'k-3. In general, the range of k′ for calculating w (k′,) at = is max
{k−2(L w −), o}k′max{k−(L w −
, o}. Here, w (k′:k) is k′
-1 is not defined, so the range of k' is as shown here. Reference numeral 107 is a word matching result storage section that stores the word cumulative matching distance w (k':k). Reference numeral 108 denotes a terminal cumulative distance calculating section which calculates, according to recurrence formula 3, from the contents of the word matching result storage section 107 and the contents of the terminal cumulative distance storage section 108.
Calculate D~(k), N~(k), and B~(k). The terminal cumulative distance storage section 109 stores the terminal cumulative distance D~(k) calculated by the terminal cumulative distance calculation section 108 until it is no longer needed. This D~(k) is used in the calculation of recurrence formula 3 in the terminal cumulative distance calculation section 108. Reference numeral 110 is a back pointer storage unit that stores the back pointers B to (k) calculated by the end cumulative distance calculation unit 108. Reference numeral 111 denotes a last word storage unit that stores the last word at the k-th vowel center determined by the end cumulative distance storage unit 109. Reference numeral 112 denotes a speech interval detection unit, which determines the speech interval based on the magnitude of the input signal, etc. When the speech interval detection unit 112 detects that voice input has started, the vowel center counting unit 113 detects the vowel. Start counting for each center. The above processing was for the k vowel center, and the count value of the vowel center counting section 113 sets this k. Therefore, the same processing as described above is performed every time the vowel center advances by one. The vowel center counting unit 113 starts counting when a speech interval is detected,
It is reset when the voice section ends. In the last word storage section 111 and the back pointer storage section 100, N~(k) and B~(k) are stored for k=1, 2, . . . , K. Segmentation part 11
4 issues an instruction to the back pointer storage section 110 to read a predetermined back pointer. That is, the segmentation section 114 is k
When a value of B to (k) is issued to the back pointer storage section 110, the back pointers B to (k) are read out from the back pointer storage section 110. When the segmentation unit 114 receives the values B to (k) from the back pointer storage unit 100, it issues the same values to the back pointer storage unit 110. Therefore, when the speech interval detection section 112 detects the end of speech input, the final value K of the vowel center counting section 113 is supplied to the segmentation section 114.
first issues a value K to the back pointer storage unit 110. Thereafter, according to the operations described above, the back pointer storage units 110B(K), B(B(K)),
..., O outputs are obtained sequentially. These values are the end frame of the second to last word, the third end frame of the last word, the fourth frame of the same word, etc. Since N~(k) is a word that ends in frame k, , if this value is directly supplied to the last word storage unit 111, recognition results will be obtained in the reverse order starting from the last word. To reverse this order (to the normal order), convert the order by using the output of the back pointer storage unit 110,
It is sufficient to perform this on the output of the last word storage section 111. FIG. 10 expresses the operation of the above embodiment as a program, and this may be followed when implementing it in software. In addition, in Figure 10,
【表】
なる記法は、条件Aが成立する間Bを行うという
ことを意味する。また、[Table] The notation means that B is performed while condition A is satisfied. Also,
【表】
なる記法は、条件Aが成立するまでBを行うとい
うことを意味する。
ステツプ200は累積距離D〜(k)、バツクポインタ
B〜(k),Dn(k−1:k),Dn(k−2:k)の初期
化を行う部分である。
ステツプ201は第k母音中心における処理を示
しており、大きくわけて素片累積照合距離Dn
(k′:k)を求める部分202と単語累積照合距
離w(k′:k)を求める部分203と終端累積
距離D〜(k)、終端バツクポインタB〜(k)、最後尾単語
N〜(k)を求める部分219に分かれる。
ステツプ202はn=1,2,…,Nについて素
片累積照合距離Dn(k′:k)を求める部分であつ
て、第1図103で行う動作に対応する。ステツ
プ204,205はステツプ206における計算の初期値
を与える部分、ステツプ209はステツプ211におけ
る計算の初期値を与える部分、ステツプ210はベ
クトル間距離dn(i′,j)を計算する部分、ステ
ツプ211は格子点(i,j)における素片累製照
合距離の途中結果Dn(i′,j)を求めている。本
実施例では第2図bの径路の拘束条件の場合を示
している。ステツプ207はDn(k′,1)を素片累
積照合距離としてDn(k′:k)に置き換えてい
る。このDn(k′:k)が素片マツチング結果記憶
部109に記憶される。
ステツプ204〜ステツプ207はnがマツチングの
条件を満たす場合に限つて実行される。即ち、n
の先行母音をVf(n),nの後続母音をVr(n),第k
母音中心の母音認識結果をV(k)とするとき、V
(k−1)orV(k−2)=Vf(n),V(k)=Vr(n)かつ
k−1〜kあるいはk−2〜kの間の子音、無音
等の特徴が標準パターンRnの特徴と一致してい
る可能性があるときのみ実行される。
ステツプ207′は、
max(St(k)−2(Jn−1),1}St(k′)max
{St(k)−〔Jn/2〕,1}
の場合にのみ実行される。
ステツプ203はw=1,2,…,Wについて単
語累積照合距離w(k′:k),wを最後尾単語と
するときの累積距離D〜w(k)、バツクポインタB〜w(k)
を計算する部分であつて、第1図の候補単語判定
部121、単語マツチング部106で行う動作に
対応する。
ステツプ203′は前記説明に従つて、DD(k′,
o)を求める部分であり、ステツプ203″は、DD
(k′,o)=0のときはw(k′:k)=(k′,
o),DD(k′,o)≠0のときはw(k′:k)=
∞とするものである。
ステツプ213はステツプ214の計算を行うに当つ
て初期化を行う部分である。ステツプ214は単語
wに対応する素片系列の最終素片s(w,Lw)か
ら、最初の素片から番目の素片までに対応す
る標準パターンの系列Rs(w,Lw),Rs(w,Lw-1),…
Rs(w, )と入力パターンの部分パターンaSt(k),aSt
(k)−1,aSt(k)−2,…,aSt(k′)との累積照合距
離を既に求めた素片累積照合距離から求める部分
である。ただし、ステツプ216の漸化式において、
s(w,)は先行母音V(k″)後続CVはwの第
1音節に等しいVCV音節である。ステツプ217,
217′は累積照合距離w(k′,1)あるいは∞を
単語累積照合距離w(k′:k)に代入する部分
である。このw(k′:k)は第1図単語マツチ
ング結果記憶部107に記憶される。ステツプ
218はwを最後尾単語とするときの累積距離D〜w
(k)、バツクポインタB〜w(k)を求める部分であある。
ステツプ219は第7図の終端累積距離計算部1
08で行う動作に対応しており、漸化式3を解い
て、D〜(k),N〜(k),B〜(k)を求める部分である。
ステツプ220,221は、ステツプ201で得られた
k=1,2,…,KについてのB〜(k),N〜(k)から認
識単語列を得る判定処理であつて、第7図のバツ
クポインタ記憶部110、セグメンテーシヨン部
114最後尾単語記憶部110で行う動作に対応
した処理を行つている。
発明の効果
本発明は、以上のように、CVやCCV音節のよ
うな音声素片を認識の単位としているので、標準
パターンの登録はいくら単語が増加してもこの音
声素片のみで済み、単語辞書はこれら素片名の系
列として表わされるので特徴ベクトルの系列とし
て記憶するのに比べて格段に少い記憶量で済み、
マツチングは前記各素片とのマツチングに費やさ
れるのがほとんどで、単語数がいくら増加しても
計算量の増加は僅かである。またDPマツチング
を行うに先立つて、母音中心、およびその認識結
果、子音、無音等に関して得られる情報のうち確
かなものを用いて、前記各素片のうちマツチング
すべき素片標準パターンを限定すること、マツチ
ングすべき単語を限定することができ、計算量は
非常に少くなる。さらに、セグメンテーシヨン
は、少々間違つていても、DPマツチングにより
最適化された結果として正しいセグメンテーシヨ
ンおよび認識が行われ、セグメンテーシヨンの不
完全さに基づく誤認識を避けることができる。
以上のことから、本発明によれば、連続発声さ
れた単語を高精度に認識することが可能となり、
実用性の高い装置である。[Table] The notation means that B is performed until condition A is satisfied. Step 200 is a part for initializing the cumulative distance D~(k), the back pointer B~(k), D n (k-1:k), and D n (k-2:k). Step 201 shows processing at the center of the k-th vowel, which can be broadly divided into unitary cumulative matching distance D n
(k':k) part 202, word cumulative matching distance w (k':k) part 203, end cumulative distance D~(k), end back pointer B~(k), last word N~ It is divided into a part 219 for calculating (k). Step 202 is a part for calculating the segment cumulative matching distance D n (k':k) for n=1, 2, . . . , N, and corresponds to the operation performed in FIG. 1 103. Steps 204 and 205 are parts that give initial values for the calculation in step 206, Step 209 is a part that gives initial values for the calculation in step 211, Step 210 is a part that calculates the distance between vectors d n (i', j), 211 calculates the intermediate result D n (i', j) of the segment cumulative matching distance at the grid point (i, j). This embodiment shows the case of the path constraint condition shown in FIG. 2b. Step 207 replaces D n (k', 1) with D n (k':k) as the segment cumulative matching distance. This D n (k':k) is stored in the segment matching result storage section 109. Steps 204 to 207 are executed only when n satisfies the matching condition. That is, n
The preceding vowel of is Vf(n), the following vowel of n is Vr(n),
When the vowel recognition result centered on the vowel is V(k), V
(k-1)orV(k-2) = Vf(n), V(k) = Vr(n) and features such as consonants and silence between k-1 and k or k-2 and k are the standard pattern R Executed only when there is a possibility of matching n features. Step 207′ is max(St(k)−2(J n −1), 1}St(k′)max
It is executed only when {St(k)−[J n /2], 1}. Step 203 calculates the cumulative word matching distance w (k':k) for w = 1, 2, ..., W, the cumulative distance D ~ w (k) when w is the last word, and the back pointer B ~ w (k). )
This is the part that calculates , and corresponds to the operation performed by the candidate word determination section 121 and word matching section 106 in FIG. Step 203′ calculates DD(k′,
o), and step 203″ is the part where DD
When (k′, o)=0, w (k′:k)=(k′,
o), when DD (k′, o)≠0, w (k′:k)=
∞. Step 213 is a part that performs initialization when performing the calculation in step 214. Step 214 starts from the final segment s(w, L w ) of the segment series corresponding to the word w and calculates a series of standard patterns R s(w,Lw) , R corresponding to the first segment to the th segment. s(w,Lw-1) ,…
R s(w, ) and subpatterns a St(k), a St of the input pattern
(k)-1, a St(k)-2, . . . , a St(k') is the part where the cumulative matching distance is calculated from the segment cumulative matching distance that has already been determined. However, in the recurrence formula of step 216,
s(w,) is the preceding vowel V(k″) and the following CV is the VCV syllable equal to the first syllable of w. Step 217.
217' is a part that substitutes the cumulative matching distance w (k', 1) or ∞ into the word cumulative matching distance w (k':k). This w (k':k) is stored in the word matching result storage section 107 in FIG. step
218 is the cumulative distance D~ w when w is the last word
(k) is the part for calculating the back pointer B~ w (k). Step 219 is the terminal cumulative distance calculation section 1 in FIG.
This corresponds to the operation performed in step 08, and is the part that solves recurrence formula 3 to obtain D~(k), N~(k), and B~(k). Steps 220 and 221 are determination processes for obtaining recognized word strings from B~(k) and N~(k) for k=1, 2, ..., K obtained in step 201, and are as shown in FIG. The back pointer storage section 110, the segmentation section 114, and the last word storage section 110 perform processing corresponding to the operations performed in the last word storage section 110. Effects of the Invention As described above, the present invention uses speech segments such as CV and CCV syllables as the unit of recognition, so no matter how many words there are, the standard pattern only needs to be registered with this speech segment. Since the word dictionary is represented as a series of these segment names, it requires much less memory than storing it as a series of feature vectors.
Most of the matching is spent on matching each of the above-mentioned fragments, and no matter how much the number of words increases, the amount of calculation increases only slightly. In addition, before performing DP matching, the standard elemental patterns to be matched among the above-mentioned elements are limited by using reliable information obtained regarding the vowel center, its recognition results, consonants, silence, etc. In addition, the words to be matched can be limited, and the amount of calculation can be extremely reduced. Furthermore, even if segmentation is slightly incorrect, DP matching will result in correct segmentation and recognition as a result of optimization, avoiding false recognition due to incomplete segmentation. . From the above, according to the present invention, it is possible to recognize continuously uttered words with high accuracy,
This is a highly practical device.
第1図は本発明の一実施例を示す図、第2図は
DPマツチングの原理を説明する図、第3図、第
4図は本発明の原理を説明する図、第5図、第6
図、第7図はそれぞれ、単語数未知の場合、単語
数既知の場合、オートマトン制御の場合に本発明
を適用した場合の認識方法の一部の動作を説明す
る図、第8図、第9図は本発明の実施例の要部の
原理を説明する図、第10図は本発明の原理をソ
フトウエア的に表現した図である。
100……音声信号入力端子、101……特徴
抽出部、102……音声素片標準パターン記憶
部、103……素片マツチング部、104……素
片マツチング結果記憶部105……単語辞書、1
06……単語マツチング部、107……単語マツ
チング結果記憶部、108……終端累積距離計算
部、109……終端累積距離記憶部、110……
バツクポインタ記憶部、111……最後尾単語記
憶部、112……音声区間検出部、113……フ
レーム数計数部、114……セグメンテーシヨン
部、115……認識結果出力端子、116……母
音標準パターン記憶部、117……母音認識部、
118……母音中心検出部、119……無音区間
検出・子音大分類部、120……特徴系列記憶
部、121……候補単語判定部。
FIG. 1 is a diagram showing an embodiment of the present invention, and FIG. 2 is a diagram showing an embodiment of the present invention.
Figures 3 and 4 are diagrams explaining the principle of DP matching, Figures 5 and 6 are diagrams explaining the principle of the present invention.
7 and 7 are diagrams illustrating a part of the operation of the recognition method when the present invention is applied when the number of words is unknown, when the number of words is known, and when automaton control is used, respectively. The figure is a diagram explaining the principle of the main part of the embodiment of the present invention, and FIG. 10 is a diagram expressing the principle of the present invention in terms of software. 100...Speech signal input terminal, 101...Feature extraction section, 102...Speech segment standard pattern storage section, 103...Segment matching section, 104...Segment matching result storage section 105...Word dictionary, 1
06...Word matching unit, 107...Word matching result storage unit, 108...Terminal cumulative distance calculation unit, 109...Terminal cumulative distance storage unit, 110...
Back pointer storage section, 111...Last word storage section, 112...Speech section detection section, 113...Frame number counting section, 114...Segmentation section, 115...Recognition result output terminal, 116...Vowel Standard pattern storage section, 117...Vowel recognition section,
118... Vowel center detection section, 119... Silent interval detection/consonant classification section, 120... Feature series storage section, 121... Candidate word determination section.
Claims (1)
声信号を特徴ベクトルの系列に変換する特徴抽出
手段と、母音、子音あるいはそれらの結合したも
の等として定義される音声素片のそれぞれに対応
した特徴ベクトルの系列を前記音声素片名に対応
づけて記憶する標準パターン記憶手段と、入力パ
ターンに対して素片の境界を検出する素片境界候
補検出手段と、標準パターンのそれぞれと前記入
力パターンから検出された前記素片境界候補の任
意または定められた種々の組合せによつて決定さ
れる部分区間(第1の部分区間)とのマツチング
を行つて両者の距離(類似度)を計算する素片マ
ツチング手段と、認識されるべき各単語・文節等
を前記音声素片名の系列として表現した単語・文
節等を記憶する単語・文節辞書と、前記認識され
るべき各単語・文節と前記入力パターンの任意ま
たは定められた前記素片境界候補の種々の部分区
間(第2の部分区間)との距離(類似度)を、前
記単語・文節辞書によつて指定される素片名の系
列に対応するように、前記第2の部分区間に含ま
れる前記第1の部分区間群を隣り合う区間が連続
するように最適に定めることにより、前記第1の
各部分区間の始点と終点およびその部分区間の前
記素片名に対応する距離(類似度)の総和を最小
(最大)とし、得られる最小値(最大値)を各単
語・文節に対する前記第2の部分区間の距離とし
て出力する機能を有する単語・文節マツチング手
段と、前記第2の部分区間群を隣り合う区間が連
続するように最適に定めることにより、前記第2
の各部分区間の始点と終点およびその部分区間の
前記単語・文節名に対応する距離(類似度)の総
和を最小(最大)となし、そのときの単語・文節
列を認識結果として判定する連続単語・文節判定
手段とを備えたことを特徴とする連続音声認識装
置。1. Feature extraction means that converts an input speech signal obtained by continuously uttering words, phrases, etc. into a series of feature vectors, and a means for extracting features corresponding to each speech element defined as a vowel, consonant, or a combination thereof. standard pattern storage means for storing a series of feature vectors in association with the phoneme segment names; segment boundary candidate detection means for detecting segment boundaries with respect to the input pattern; and each of the standard patterns and the input pattern. A element that calculates the distance (similarity) between the two by matching with a partial interval (first partial interval) determined by various arbitrary or predetermined combinations of the element boundary candidates detected from the element boundary candidate. a one-sided matching means; a word/phrase dictionary that stores words/phrases expressing each word/phrase, etc. to be recognized as a series of phoneme segment names; and each word/phrase to be recognized and the input. The distance (similarity) between the arbitrary or predetermined segment boundary candidates of the pattern and various subintervals (second subintervals) is calculated based on the sequence of segment names specified by the word/phrase dictionary. Correspondingly, by optimally determining the first sub-interval group included in the second sub-interval so that adjacent sections are continuous, the starting point and end point of each of the first sub-intervals and their portions can be determined. A function that minimizes (maximizes) the sum of distances (similarities) corresponding to the segment names of an interval and outputs the obtained minimum value (maximum value) as the distance of the second subinterval to each word/clause. By optimally determining the second partial interval group so that adjacent intervals are continuous, the second partial interval group is
A sequence in which the sum of the distances (similarities) corresponding to the start and end points of each subinterval and the word/phrase name in that subinterval is the minimum (maximum), and the word/phrase string at that time is determined as the recognition result. 1. A continuous speech recognition device comprising a word/phrase determination means.
Priority Applications (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP59269955A JPS61148498A (en) | 1984-12-21 | 1984-12-21 | Continuous voice recognition equipment |
Applications Claiming Priority (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP59269955A JPS61148498A (en) | 1984-12-21 | 1984-12-21 | Continuous voice recognition equipment |
Publications (2)
| Publication Number | Publication Date |
|---|---|
| JPS61148498A JPS61148498A (en) | 1986-07-07 |
| JPH0566599B2 true JPH0566599B2 (en) | 1993-09-22 |
Family
ID=17479541
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| JP59269955A Granted JPS61148498A (en) | 1984-12-21 | 1984-12-21 | Continuous voice recognition equipment |
Country Status (1)
| Country | Link |
|---|---|
| JP (1) | JPS61148498A (en) |
-
1984
- 1984-12-21 JP JP59269955A patent/JPS61148498A/en active Granted
Also Published As
| Publication number | Publication date |
|---|---|
| JPS61148498A (en) | 1986-07-07 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US5581655A (en) | Method for recognizing speech using linguistically-motivated hidden Markov models | |
| US6236965B1 (en) | Method for automatically generating pronunciation dictionary in speech recognition system | |
| CN117043857A (en) | Methods, devices and computer program products for English pronunciation assessment | |
| JPWO2015118645A1 (en) | Voice search apparatus and voice search method | |
| JPH07306691A (en) | Apparatus and method for speaker-independent speech recognition | |
| JP2955297B2 (en) | Speech recognition system | |
| JPH0247760B2 (en) | ||
| EP0103258B1 (en) | Pattern matching apparatus | |
| JP2001312293A (en) | Voice recognition method and apparatus, and computer-readable storage medium | |
| JP2001005483A (en) | Word voice recognizing method and word voice recognition device | |
| JPH0566598B2 (en) | ||
| KR100560916B1 (en) | Speech recognition method using posterior distance | |
| JPH0464077B2 (en) | ||
| JP3231365B2 (en) | Voice recognition device | |
| JPS60164800A (en) | Voice recognition equipment | |
| JPH0534680B2 (en) | ||
| KR100316776B1 (en) | Continuous digits recognition device and method thereof | |
| Stephenson | Speech recognition using phonetically featured syllables | |
| JP2721341B2 (en) | Voice recognition method | |
| JPS61148498A (en) | Continuous voice recognition equipment | |
| JPS6180298A (en) | voice recognition device | |
| JPS59173884A (en) | pattern comparison device | |
| JPH0554678B2 (en) | ||
| JPH01302295A (en) | Word position detecting method and generating method for phoneme standard pattern | |
| JPS60150098A (en) | voice recognition device |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| LAPS | Cancellation because of no payment of annual fees |