JPH086590A - Speech recognition method and apparatus for speech dialogue - Google Patents

Speech recognition method and apparatus for speech dialogue

Info

Publication number
JPH086590A
JPH086590A JP6134005A JP13400594A JPH086590A JP H086590 A JPH086590 A JP H086590A JP 6134005 A JP6134005 A JP 6134005A JP 13400594 A JP13400594 A JP 13400594A JP H086590 A JPH086590 A JP H086590A
Authority
JP
Japan
Prior art keywords
utterance
user
voice recognition
voice
likelihood
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Granted
Application number
JP6134005A
Other languages
Japanese (ja)
Other versions
JP3285704B2 (en
Inventor
Shingo Kuroiwa
眞吾 黒岩
Kazuya Takeda
一哉 武田
Masaki Naito
正樹 内藤
Seiichi Yamamoto
誠一 山本
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
KDDI Corp
Original Assignee
Kokusai Denshin Denwa KK
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Kokusai Denshin Denwa KK filed Critical Kokusai Denshin Denwa KK
Priority to JP13400594A priority Critical patent/JP3285704B2/en
Publication of JPH086590A publication Critical patent/JPH086590A/en
Application granted granted Critical
Publication of JP3285704B2 publication Critical patent/JP3285704B2/en
Anticipated expiration legal-status Critical
Expired - Fee Related legal-status Critical Current

Links

Abstract

(57)【要約】 【目的】 システムアナウンスとユーザの発話開始時刻
との因果関係を利用し、高精度な発話開始時刻の検出を
行って高精度な音声認識を行うこと、並びに発話開始時
刻を一点に決定することなく高精度な音声認識を行うこ
と。 【構成】 アナウンス発声装置2においてシステムアナ
ウンス中に無音区間を設けてユーザの発話開始時刻を制
御し、これに対応して用意した発話開始時刻の予測分布
から演算部12で第1の発話開始点らしさa(t)を算
出し、一方発話検出用の特徴パラメータ13aから演算
部14で第2の発話開始点らしさb(t)を算出し、こ
れらから演算部15で第3の発話開始点らしさα(t)
を算出し、α(t)と基準値Refとの比較により発話開
始時刻を決定してスイッチ17をオンさせることにより
音声認識を開始する。
(57) [Abstract] [Purpose] Utilizing the causal relationship between the system announcement and the user's utterance start time, it is possible to detect the utterance start time with high accuracy and perform high-accuracy speech recognition. Perform highly accurate voice recognition without making a single decision. [Structure] In the announcement utterance device 2, a silent section is provided during a system announcement to control the utterance start time of the user, and the calculation unit 12 uses the predicted distribution of the utterance start times corresponding to this to generate a first utterance start point. Likelihood a (t) is calculated, and on the other hand, the arithmetic unit 14 calculates the second utterance start point likelihood b (t) from the utterance detection characteristic parameter 13a, and from these, the arithmetic unit 15 calculates the third utterance start point likelihood. α (t)
Is calculated, the speech start time is determined by comparing α (t) with the reference value Ref, and the switch 17 is turned on to start voice recognition.

Description

【発明の詳細な説明】Detailed Description of the Invention

【0001】[0001]

【産業上の利用分野】本発明は音声を用いてユーザ(利
用者)との対話を行う音声対話装置に関し、特には、ユ
ーザの発話開始時刻に対する検出精度の向上、並びにユ
ーザの発話に対する音声認識精度の向上に有用なもので
ある。
BACKGROUND OF THE INVENTION 1. Field of the Invention The present invention relates to a voice dialog device for performing a dialogue with a user (user) by using voice, and more particularly, to improve detection accuracy with respect to a user's utterance start time and voice recognition for the user's utterance. It is useful for improving accuracy.

【0002】[0002]

【従来の技術】音声対話装置では音声を用いて装置側か
らユーザに話しかけることによりシステムアナウンスを
行い、ユーザの発話即ちユーザが発する音声を認識する
ことによりユーザの意思を理解して、ユーザと装置間で
対話を行う。従って、音声認識精度が重要である。
2. Description of the Related Art In a voice dialog device, a system announcement is made by talking to the user from the device side using voice, and by recognizing the user's utterance, that is, the voice uttered by the user, the user's intention is understood, and the user and the device are recognized. Have a dialogue between. Therefore, the accuracy of voice recognition is important.

【0003】図7を参照して、従来の音声対話装置にお
ける音声認識方法及び音声認識装置を説明する。図7に
おいて、音声対話装置は対話管理装置1と、アナウンス
発声装置2と、音声認識装置50とを具備している。音
声出力回路3及び音声入力回路5は音声対話装置に内蔵
されることもあり、あるいは音声対話装置とは別物で適
宜接続されることもある。後者の例としては電話機の送
受話器があり、電話回線と電話交換機を通して音声対話
装置に接続される。音声認識装置50は発話検出用の音
声信号通過スイッチ51と、発話検出用の音響分析部1
3と、発話検出部52と、音声認識用の音声信号通過ス
イッチ17と、音声認識用の音響分析部18と、音声認
識部19とを具備している。
With reference to FIG. 7, a voice recognition method and a voice recognition device in a conventional voice dialogue system will be described. In FIG. 7, the voice dialogue device includes a dialogue management device 1, an announcement utterance device 2, and a voice recognition device 50. The voice output circuit 3 and the voice input circuit 5 may be built in the voice dialog device, or may be separately connected to the voice dialog device as appropriate. An example of the latter is a telephone handset, which is connected to a voice interaction device through a telephone line and telephone exchange. The voice recognition device 50 includes a voice signal passage switch 51 for utterance detection and an acoustic analysis unit 1 for utterance detection.
3, a speech detection unit 52, a voice signal passage switch 17 for voice recognition, a sound analysis unit 18 for voice recognition, and a voice recognition unit 19.

【0004】以下、図7に示した音声対話装置の動作と
各部の機能を説明する。
The operation of the voice interactive apparatus shown in FIG. 7 and the function of each section will be described below.

【0005】(i)アナウンス発声装置2では、対話管
理装置1がコード名等により指定したシステムアナウン
スのテキスト1aに基づいて、発声すべき音声の電気信
号2aを作成し、音声出力回路3に送る。また、システ
ムアナウンスの開始を表わすアナウンス開始信号2b、
あるいはシステムアナウンスの終了を表わすアナウンス
終了信号2cを音声認識装置50の発話検出用音声信号
通過スイッチ51に送る。音声出力回路3は電気信号2
aを音声に変換して、システムアナウンス3aをユーザ
に聞かせる。このシステムアナウンス3aに対するユー
ザの発話4を音声入力回路5が受け取り、電気的音声信
号5aに変換して音声認識装置50の発話検出用及び音
声認識用の各音声信号通過スイッチ51,17に送る。
(I) In the announcement utterance device 2, the dialogue management device 1 creates an electric signal 2a of a voice to be uttered based on the text 1a of the system announcement designated by the code name or the like, and sends it to the voice output circuit 3. . Also, an announcement start signal 2b indicating the start of the system announcement,
Alternatively, an announcement end signal 2c indicating the end of the system announcement is sent to the speech detection voice signal passing switch 51 of the voice recognition device 50. The voice output circuit 3 is an electric signal 2
A is converted into voice, and the system announcement 3a is heard by the user. The voice input circuit 5 receives the user's utterance 4 to the system announcement 3a, converts it into an electric voice signal 5a, and sends it to the voice signal passing switches 51 and 17 of the voice recognition device 50 for utterance detection and voice recognition.

【0006】(ii)音声認識装置50では、システムア
ナウンス中のユーザの割り込み発話を受け付ける場合は
アナウンス開始信号2bを与えられた時からアナウンス
終了信号2cを与えられた後の一定時間まで音声信号通
過スイッチ51が閉(オン)となり、またシステムアナ
ウンス中のユーザの割込み発話を受け付けない場合はア
ナウンス終了信号2cを与えられた時から一定時間だけ
音声信号通過スイッチ51が閉(オン)となる。この音
声信号通過スイッチ51が閉じている間に送られた音声
信号を発話検出対象の信号51aとして発話検出用の音
響分析部13に送る。
(Ii) In the voice recognition device 50, in the case of accepting the interruption utterance of the user who is making the system announcement, the voice signal is passed from a time when the announcement start signal 2b is given to a certain time after the announcement end signal 2c is given. When the switch 51 is closed (ON) and the interruption utterance of the user who is making the system announcement is not accepted, the voice signal passing switch 51 is closed (ON) for a fixed time from when the announcement end signal 2c is given. The voice signal sent while the voice signal passing switch 51 is closed is sent to the utterance detection acoustic analysis unit 13 as the utterance detection target signal 51a.

【0007】(iii)この音響分析部13では、音声信号
通過スイッチ51を通過した音声信号51aから、パワ
ースペクトラムなどユーザの発話検出に適した特徴パラ
メータ13aを算出して発話検出部52に送る。発話検
出部52では、特徴パラメータ13aに基づき、ユーザ
の発話開始時刻と発話終了時刻とを各一点決定し、その
間を指定する信号52aを音声認識用の音声信号通過ス
イッチ17に送る。
(Iii) The acoustic analysis unit 13 calculates a characteristic parameter 13a suitable for detecting a user's utterance, such as a power spectrum, from the voice signal 51a that has passed through the voice signal passing switch 51 and sends it to the utterance detection unit 52. The utterance detection unit 52 determines one point each of the utterance start time and the utterance end time of the user based on the characteristic parameter 13a and sends a signal 52a designating the time point to the voice signal passing switch 17 for voice recognition.

【0008】(iv)音声信号通過スイッチ17は発話検
出部52からの信号52aにより指定された間のみ閉
(オン)となり、閉じている間に送られてきた音声信号
を音声認識対象の信号17aとして音声認識用の音響分
析部18に送る。この音響分析部18では、音声信号通
過スイッチ17を通過した音声信号17aから、音声認
識に適した特徴パラメータ18aを算出し、音声認識部
19に送る。音声認識部19では、特徴パラメータ18
aに基づいて音声認識を行い、その認識結果19aを対
話管理装置1に送る。
(Iv) The voice signal passing switch 17 is closed (ON) only during a period designated by the signal 52a from the utterance detecting section 52, and the voice signal transmitted while the voice signal passing switch 17 is closed is the signal 17a for voice recognition. To the acoustic analysis unit 18 for voice recognition. The acoustic analysis unit 18 calculates a characteristic parameter 18a suitable for voice recognition from the voice signal 17a that has passed through the voice signal passing switch 17, and sends it to the voice recognition unit 19. In the voice recognition unit 19, the characteristic parameter 18
Speech recognition is performed based on a, and the recognition result 19a is sent to the dialogue management device 1.

【0009】(v)対話管理装置1では、音声認識部1
9から与えられる認識結果19aに基づいて、次に発声
すべきシステムアナウンスのテキスト1aを決定してア
ナウンス発声装置2にコード名等を送る。
(V) In the dialogue management device 1, the voice recognition unit 1
Based on the recognition result 19a given from 9, the system announcement text 1a to be uttered next is determined, and the code name or the like is sent to the announcement utterance device 2.

【0010】以上の動作を繰り返すことにより、人間と
装置間で音声を用いた対話が行われる。なお対話管理装
置1は、必要があれば、対話内容からユーザの意思を認
識してその情報1bを外部に出力する。
By repeating the above operation, a dialogue using voice is performed between the person and the device. If necessary, the dialogue management device 1 recognizes the user's intention from the dialogue contents and outputs the information 1b to the outside.

【0011】[0011]

【発明が解決しようとする課題】音声対話装置では音声
認識の精度が重要であるが、上述した従来技術をユーザ
の割り込み発話を受け付けるように利用した場合には、
下記(a),(b)のような改善すべき点がある。
Although the accuracy of voice recognition is important in the voice interactive apparatus, when the above-mentioned conventional technique is used to receive the interrupt utterance of the user,
There are points to be improved such as (a) and (b) below.

【0012】(a)発話検出部52ではパワースペクト
ラムなどの特徴パラメータ13aのみを用いてユーザの
発話検出を行っているため、発話開始時刻の検出精度が
良くない。更に、システムアナウンス中にユーザが意味
のない発声(冗長語)や咳をしてしまうと、その時点を
ユーザの発話開始時刻として誤って検出する可能性が高
い。その結果、意味のない発声や咳をも認識対象に含ん
で音声認識を行うことになり、音声認識精度が低下す
る。
(A) Since the utterance detection unit 52 detects the utterance of the user using only the characteristic parameter 13a such as the power spectrum, the utterance start time detection accuracy is not good. Further, if the user makes a meaningless utterance (redundant word) or coughs during the system announcement, there is a high possibility that the time point is erroneously detected as the utterance start time of the user. As a result, the speech recognition is performed by including the meaningless utterance and cough in the recognition target, and the speech recognition accuracy decreases.

【0013】(b)更に、発話検出部52ではユーザの
発話開始時刻を一点のみに決定しているため、発話開始
時刻の検出に誤りが生じた場合には、音声認識部19で
は回復できない誤りとなって音声認識の精度が低下す
る、という決定的な誤りの伝搬が生じる。
(B) Furthermore, since the utterance detection unit 52 determines the utterance start time of the user to be only one point, if an error occurs in the detection of the utterance start time, an error that the voice recognition unit 19 cannot recover. Therefore, definite error propagation occurs that the accuracy of voice recognition is reduced.

【0014】そこで本発明は、ユーザの発話開始時刻の
検出精度を向上させることにより高精度な音声認識を行
うことができる音声認識方法及び装置を提供することを
目的とし、更に、ユーザの発話開始時刻の検出に誤りが
あってもこれの影響を減らして高精度な音声認識を行う
ことができる音声認識方法及び装置を提供することを他
の目的とする。
Therefore, an object of the present invention is to provide a voice recognition method and apparatus capable of performing high-accuracy voice recognition by improving the detection accuracy of the user's utterance start time, and further, to start the user's utterance start. Another object of the present invention is to provide a voice recognition method and device capable of performing highly accurate voice recognition by reducing the influence of time detection error even if there is an error.

【0015】[0015]

【課題を解決するための手段】上記目的を達成する第1
の発明は、音声を用いてユーザとの対話を行う音声対話
装置に適用される音声認識方法において:前記音声対話
装置のシステムアナウンスに対するユーザの発話開始時
刻の極大点を有する予測分布を予め用意しておき、この
予測分布に基づき、ユーザの発話が開始される期待値を
第1の発話開始点らしさとしてシステムアナウンス開始
後の時刻に応じて算出すること;電気信号に変換された
ユーザの発話を音響分析して発話検出用の特徴パラメー
タを算出し、この特徴パラメータに基づき、ユーザの発
話が開始されたであろう尤度を第2の発話開始点らしさ
として時刻に応じて算出すること;第2の発話開始点ら
しさに対して第1の発話開始点らしさにより重み付けを
行い、この重み付けで得た第3の発話開始点らしさを基
準値と比較し、基準値より大きくなった時点をユーザの
発話開始時刻であると決定すること;電気信号に変換さ
れたユーザの発話を音響分析して音声認識用の特徴パラ
メータを算出し、この特徴パラメータに基づき音声認識
を行う処理を、前記ユーザの発話開始時刻の決定に従っ
て行うこと;を特徴とする音声認識方法である。
[Means for Solving the Problems] First to achieve the above object
In a voice recognition method applied to a voice interaction device for interacting with a user using voice, a prediction distribution having a maximum point of a user's utterance start time for a system announcement of the voice interaction device is prepared in advance. Based on this prediction distribution, the expected value at which the user's utterance starts is calculated as the first utterance start point according to the time after the system announcement starts; the user's utterance converted into an electric signal is calculated. Acoustic analysis is performed to calculate a characteristic parameter for utterance detection, and based on the characteristic parameter, a likelihood that the utterance of the user is likely to be started is calculated as second utterance starting point likelihood according to time; The likelihood of the utterance start point of 2 is weighted by the likelihood of the first utterance start point, and the likelihood of the third utterance start point obtained by this weighting is compared with a reference value, The time when it becomes larger than the value is determined to be the user's utterance start time; the user's utterance converted into an electric signal is acoustically analyzed to calculate a characteristic parameter for speech recognition, and the speech recognition is performed based on this characteristic parameter. Performing the process according to the determination of the utterance start time of the user;

【0016】また第2の発明は音声を用いてユーザとの
対話を行う音声対話装置に適用される音声認識方法にお
いて:前記音声対話装置のシステムアナウンスに対する
ユーザの発話開始時刻の極大点を有する予測分布を予め
用意しておき、この予測分布に基づき、ユーザの発話が
開始される期待値を第1の発話開始点らしさとしてシス
テムアナウンス開始後の時刻に応じて算出すること;電
気信号に変換されたユーザの発話を音響分析して発話検
出用の特徴パラメータを算出し、この特徴パラメータに
基づき、ユーザの発話が開始されたであろう尤度を第2
の発話開始点らしさとして時刻に応じて算出すること;
第1の発話開始点らしさにより第1の基準値を重み付け
して時間に応じて変化する第2の基準値を算出し、第2
の発話開始点らしさをこの第2の基準値と比較し、第2
の基準値より大きくなった時点をユーザの発話開始時刻
であると決定すること;電気信号に変換されたユーザの
発話を音響分析して音声認識用の特徴パラメータを算出
し、この特徴パラメータに基づき音声認識を行う処理
を、前記ユーザの発話開始時刻の決定に従って行うこ
と;を特徴とする音声認識方法である。
A second aspect of the present invention is a voice recognition method applied to a voice dialogue device for dialogue with a user by using voice: Prediction having a maximum point of a user's utterance start time with respect to a system announcement of the voice dialogue device. Prepare a distribution in advance, and calculate the expected value at which the user's utterance starts based on this predicted distribution as the first utterance start point according to the time after the system announcement starts; converted to an electrical signal The utterance of the user is acoustically analyzed to calculate a characteristic parameter for utterance detection, and based on the characteristic parameter, the likelihood that the utterance of the user may have started
Of the utterance start point of the user according to the time;
The first reference value is weighted according to the likelihood of the first utterance to calculate the second reference value that changes with time, and the second reference value is calculated.
The utterance starting point of the person is compared with the second reference value, and the second
Is determined to be the user's utterance start time; the time when the utterance of the user is converted into an electric signal is acoustically analyzed to calculate a characteristic parameter for voice recognition, and based on the characteristic parameter Performing a process of performing voice recognition according to the determination of the utterance start time of the user;

【0017】第3の発明は、音声を用いてユーザとの対
話を行う音声対話装置に適用される音声認識方法におい
て:前記音声対話装置のシステムアナウンスに対するユ
ーザの発話開始時刻の極大点を有する予測分布を予め用
意しておき、この予測分布に基づき、ユーザの発話が開
始される期待値を第1の発話開始点らしさとしてシステ
ムアナウンス開始後の時刻に応じて算出すること;電気
信号に変換されたユーザの発話を音響分析して発話検出
用の特徴パラメータを算出し、この特徴パラメータに基
づき、ユーザの発話が開始されたであろう尤度を第2の
発話開始点らしさとして時刻に応じて算出すること;第
2の発話開始点らしさに対して第1の発話開始点らしさ
により重み付けを行い、第3の発話開始点らしさを算出
すること;電気信号に変換されたユーザの発話を音響分
析して音声認識用の特徴パラメータを算出し、この特徴
パラメータに基づき、認識開始時刻を次々にずらして音
声認識を行い、且つ、各認識開始時刻に対応した音声認
識結果毎の尤度を算出すること;各認識開始時刻毎の音
声認識結果の尤度と、前記重み付けで得た第3の発話開
始点らしさとの和または積を時刻を合わせて算出し、こ
の算出した値が最大となる認識開始時刻に対応した音声
認識結果を、ユーザの発話に対する音声認識結果と判定
すること;を特徴とする音声認識方法である。
In a third aspect of the present invention, there is provided a voice recognition method applied to a voice dialogue device for dialogue with a user by using a voice: Prediction having a maximum point of a user's utterance start time with respect to a system announcement of the voice dialogue device. Prepare a distribution in advance, and calculate the expected value at which the user's utterance starts based on this predicted distribution as the first utterance start point according to the time after the system announcement starts; converted to an electrical signal The utterance of the user is acoustically analyzed to calculate a characteristic parameter for utterance detection, and based on the characteristic parameter, the likelihood that the utterance of the user is likely to be started is determined as the second utterance start point according to the time. Calculating; weighting the likelihood of the second utterance start point with the likelihood of the first utterance start point to calculate the likelihood of the third utterance start point; The acoustic parameters of the user's utterance converted to are calculated to calculate the characteristic parameters for voice recognition, and based on the characteristic parameters, the recognition start times are shifted one after another to perform the voice recognition, and each recognition start time is corresponded. Calculating the likelihood for each voice recognition result; calculating the sum or product of the likelihood of the voice recognition result for each recognition start time and the third utterance starting point likelihood obtained by the weighting together with the time. Determining the voice recognition result corresponding to the recognition start time at which the calculated value becomes maximum as the voice recognition result for the user's utterance.

【0018】第4の発明は、音声を用いてユーザとの対
話を行う音声対話装置に適用される音声認識方法におい
て:前記音声対話装置のシステムアナウンスに対するユ
ーザの発話開始時刻の極大点を有する予測分布を予め用
意しておき、この予測分布に基づき、ユーザの発話が開
始される期待値を第1の発話開始点らしさとしてシステ
ムアナウンス開始後の時刻に応じて算出すること;電気
信号に変換されたユーザの発話を音響分析して発話検出
用の特徴パラメータを算出し、この特徴パラメータに基
づき、ユーザの発話が開始されたであろう尤度を第2の
発話開始点らしさとして時刻に応じて算出すること;第
2の発話開始点らしさに対して第1の発話開始点らしさ
により重み付けを行い、第3の発話開始点らしさを算出
すること;電気信号に変換されたユーザの発話を音響分
析して音声認識用の特徴パラメータを算出し、この特徴
パラメータに基づき、先頭に無音状態を有する確率付き
有限状態ネットワークを探索して音声認識を行うこと;
前記確率付き有限状態ネットワークの先頭の無音状態か
ら文の先頭状態へ遷移する確率を、前記重み付けで得た
第3の発話開始点らしさを用いて時刻に応じて更新する
こと;を特徴とする音声認識方法である。
In a fourth aspect of the present invention, there is provided a voice recognition method applied to a voice dialogue device for dialogue with a user by using a voice: a prediction having a maximum point of a user's utterance start time with respect to a system announcement of the voice dialogue device. Prepare a distribution in advance, and calculate the expected value at which the user's utterance starts based on this predicted distribution as the first utterance start point according to the time after the system announcement starts; converted to an electrical signal The utterance of the user is acoustically analyzed to calculate a characteristic parameter for utterance detection, and based on the characteristic parameter, the likelihood that the utterance of the user is likely to be started is determined as the second utterance start point according to the time. Calculating; weighting the likelihood of the second utterance start point with the likelihood of the first utterance start point to calculate the likelihood of the third utterance start point; Utterances converted user by acoustic analysis to calculate the characteristic parameters for voice recognition, on the basis of the characteristic parameter, leading to perform speech recognition by searching the probability finite-state network with a silent state in the;
Updating the probability of transition from the silent state at the head of the finite state network with probability to the head state of the sentence according to the time using the third utterance starting point likelihood obtained by the weighting. It is a recognition method.

【0019】そして第5の発明は、第1ないし第4の発
明において、前記ユーザの発話開始時刻の予測分布の極
大点がシステムアナウンスの無音区間に存在することを
特徴とし、第6の発明は更に前記無音区間はその長さが
0.2秒以上3秒以下であり、システムアナウンスの文
と文の間及び文節と文節との間のうち少なくとも一方に
存在することを特徴とする。
A fifth aspect of the invention is characterized in that, in the first to fourth aspects, the maximum point of the predicted distribution of the utterance start time of the user exists in the silent section of the system announcement, and the sixth aspect of the invention. Further, the silent section has a length of 0.2 seconds or more and 3 seconds or less, and is present in at least one of a sentence of the system announcement and a sentence.

【0020】次に、第7の発明は、音声を用いてユーザ
との対話を行う音声対話装置に適用される音声認識装置
において;前記音声対話装置のシステムアナウンスに対
するユーザの発話開始時刻の極大点を有する予測分布を
格納する第1手段と;前記格納された予測分布に基づ
き、ユーザの発話が開始される期待値を第1の発話開始
点らしさとしてシステムアナウンス開始後の時刻に応じ
て算出する第2手段と;電気信号に変換されたユーザの
発話を音響分析し、発話検出用の特徴パラメータを算出
する第3手段と;前記発話検出用の特徴パラメータに基
づき、ユーザの発話が開始されたであろう尤度を第2の
発話開始点らしさとして時刻に応じて算出する第4手段
と;第2の発話開始点らしさに対して第1の発話開始点
らしさにより重み付けを行い、この重み付けされた値を
第3の発話開始点らしさとして時刻に応じて算出する第
5手段と;第3の発話開始点らしさを基準値と比較し、
この基準値より大きくなった時点をユーザの発話開始時
刻であると決定する第6手段と;電気信号に変換された
ユーザの発話を音響分析して音声認識用の特徴パラメー
タを算出し、この特徴パラメータに基づき音声認識を行
う処理を、前記ユーザの発話開始時刻の決定に従って行
う第7手段と;を具備することを特徴とする音声認識装
置である。
Next, a seventh aspect of the present invention is a voice recognition device applied to a voice dialogue device for dialogue with a user using voice; the maximum point of a user's utterance start time with respect to a system announcement of the voice dialogue device. A first means for storing a prediction distribution having the following; based on the stored prediction distribution, an expected value at which the user's utterance is started is calculated as the first utterance start point according to the time after the system announcement is started. Second means; acoustic analysis of the user's utterance converted into an electric signal, and third means for calculating characteristic parameters for utterance detection; and utterance of the user based on the characteristic parameter for utterance detection And a fourth means for calculating the likelihood as the second utterance start point likelihood according to the time; and weighting the second utterance start point likelihood by the first utterance start point likelihood. Is compared with a reference value of the third utterance start point likelihood of; was carried out, the weighted values fifth means and for calculating in response to time as a third utterance start point ness
Sixth means for determining a time point when the value becomes larger than the reference value as a user's utterance start time; acoustic analysis of the user's utterance converted into an electric signal to calculate a characteristic parameter for voice recognition, and this characteristic And a seventh means for performing a process of performing voice recognition based on a parameter according to the determination of the utterance start time of the user.

【0021】第8の発明は、音声を用いてユーザとの対
話を行う音声対話装置に適用される音声認識装置におい
て;前記音声対話装置のシステムアナウンスに対するユ
ーザの発話開始時刻の極大点を有する予測分布を格納す
る第1手段と;前記格納された予測分布に基づき、ユー
ザの発話が開始される期待値を第1の発話開始点らしさ
としてシステムアナウンス開始後の時刻に応じて算出す
る第2手段と;電気信号に変換されたユーザの発話を音
響分析し、発話検出用の特徴パラメータを算出する第3
手段と;前記発話検出用の特徴パラメータに基づき、ユ
ーザの発話が開始されたであろう尤度を第2の発話開始
点らしさとして時刻に応じて算出する第4手段と;第1
の基準値に対して第1の発話開始点らしさにより重み付
けを行い、この重み付けされた値を第2の基準値として
時刻に応じて算出する第5手段と;第2の発話開始点ら
しさを前記重み付けで得た第2の基準値と比較し、この
第2の基準値より大きくなった時点をユーザの発話開始
時刻であると決定する第6手段と;電気信号に変換され
たユーザの発話を音響分析して音声認識用の特徴パラメ
ータを算出し、この特徴パラメータに基づき音声認識を
行う処理を、前記ユーザの発話開始時刻の決定に従って
行う第7手段と;を具備することを特徴とする音声認識
装置である。
An eighth aspect of the present invention is a voice recognition device applied to a voice dialogue device for dialogue with a user using voice; a prediction having a maximum point of a user's utterance start time with respect to a system announcement of the voice dialogue device. First means for storing a distribution; second means for calculating an expected value at which a user's utterance starts based on the stored predicted distribution as a first utterance start point in accordance with a time after the system announcement is started And; Thirdly, acoustic analysis is performed on the user's utterance converted into an electric signal, and characteristic parameters for utterance detection are calculated.
Fourth means for calculating the likelihood that the user's utterance has started based on the characteristic parameter for utterance detection as the second utterance start point according to the time;
Means for weighting the reference value of the first utterance starting point likeness and calculating the weighted value as the second reference value according to the time; Sixth means for comparing with a second reference value obtained by weighting, and determining a time point when the second reference value is larger than the second reference value as the utterance start time of the user; And a seventh means for performing a process of performing characteristic analysis for voice recognition by acoustic analysis and performing voice recognition based on the characteristic parameter according to the determination of the utterance start time of the user; It is a recognition device.

【0022】そして第9の発明は、第7または第8の発
明における第7手段が、電気信号に変換されたユーザの
発話を、ユーザの発話開始時刻であると決定された時点
から通過させるスイッチ手段と;このスイッチ手段を通
過したユーザの発話を音響分析して音声認識用の特徴パ
ラメータを算出する音声認識用の音響分析手段と;この
音響分析手段により算出された音声認識用の特徴パラメ
ータに基づいて音声認識を行う音声認識手段と;を具備
することを特徴とする。また第10の発明は第7または
第8の発明における第7手段が、ユーザの発話開始時刻
であると決定された時点から、電気信号に変換されたユ
ーザの発話の音響分析を開始して音声認識用の特徴パラ
メータを算出する音声認識用の音響分析手段と;この音
響分析手段により算出された音声認識用の特徴パラメー
タに基づいて音声認識を行う音声認識手段と;を具備す
ることを特徴とする。更に第11の発明は第7または第
8の発明における第7手段が、電気信号に変換されたユ
ーザの発話を音響分析して音声認識用の特徴パラメータ
を算出する音声認識用の音響分析手段と;この音響分析
手段で算出された音声認識用の特徴パラメータのうち、
ユーザの発話開始時刻であると決定された時点以降の特
徴パラメータに基づいて音声認識を行う音声認識手段
と;を具備することを特徴とする。
A ninth aspect of the present invention is a switch, wherein the seventh means of the seventh or eighth aspect of the invention passes the user's utterance converted into an electric signal from the time when it is determined to be the user's utterance start time. Means; acoustic analysis means for voice recognition for acoustically analyzing the utterance of the user passing through the switch means to calculate a characteristic parameter for voice recognition; and a characteristic parameter for voice recognition calculated by the acoustic analysis means. Voice recognition means for performing voice recognition based on the above; A tenth invention is that the seventh means in the seventh or eighth invention starts the acoustic analysis of the user's utterance converted into an electric signal from the time when it is determined to be the user's utterance start time, and outputs the voice. A voice recognition acoustic analysis unit for calculating a recognition feature parameter; and a voice recognition unit for performing voice recognition based on the voice recognition feature parameter calculated by the sound analysis unit; To do. An eleventh aspect of the invention is that the seventh means of the seventh or eighth aspect of the invention is an acoustic analysis means for voice recognition, which acoustically analyzes the user's utterance converted into an electric signal to calculate a characteristic parameter for voice recognition. Of the characteristic parameters for voice recognition calculated by this acoustic analysis means,
And a voice recognition means for performing voice recognition based on the characteristic parameter after the time when it is determined to be the utterance start time of the user.

【0023】次に第12の発明は、音声を用いてユーザ
との対話を行う音声対話装置に適用される音声認識装置
において;前記音声対話装置のシステムアナウンスに対
するユーザの発話開始時刻の極大点を有する予測分布を
格納する予測分布格納手段と;前記格納された予測分布
に基づき、ユーザの発話が開始される期待値を第1の発
話開始点らしさとしてシステムアナウンス開始後の時刻
に応じて算出する第1の演算手段と;電気信号に変換さ
れたユーザの発話を音響分析し、発話検出用の特徴パラ
メータを算出する発話検出用の音声分析手段と;前記発
話検出用の特徴パラメータに基づき、ユーザの発話が開
始されたであろう尤度を第2の発話開始点らしさとして
時刻に応じて算出する第2の演算手段と;第2の発話開
始点らしさに対して第1の発話開始点らしさにより重み
付けを行い、この重み付けされた値を第3の発話開始点
らしさとして時刻に応じて算出する第3の演算手段と;
前記電気信号に変換されたユーザの発話を音響分析し、
音声認識用の特徴パラメータを算出する音声認識用の音
響分析手段と;前記音声認識用の特徴パラメータに基づ
き、認識開始時刻を次々にずらして音声認識を行い、且
つ、各認識開始時刻に対応した音声認識結果毎の尤度を
算出する音声認識手段と;各認識開始時刻毎の音声認識
結果の尤度と、第3の発話開始点らしさとの和または積
を時刻に合せて算出し、この算出した値が最大となる認
識開始時刻に対応した音声認識結果を、ユーザの発話に
対する音声認識結果と判定する音声認識結果判定手段
と;を具備することを特徴とする音声認識装置である。
Next, a twelfth aspect of the invention is a voice recognition device applied to a voice dialogue device for dialogue with a user using voice; the maximum point of a user's utterance start time with respect to a system announcement of the voice dialogue device. A predictive distribution storage unit for storing the predictive distribution having; an expected value at which the user's utterance starts is calculated based on the stored predictive distribution as a first utterance start point in accordance with a time after the system announcement is started. First computing means; speech analysis means for utterance detection for acoustically analyzing the utterance of the user converted into an electric signal, and calculating characteristic parameters for utterance detection; user based on the characteristic parameters for utterance detection Second computing means for calculating the likelihood that the utterance of ”is likely to have started as the second utterance start point likelihood according to the time; for the second utterance start point likelihood Performs weighting by the first utterance start point ness, and third arithmetic means for calculating according to the time the weighted value as the third utterance start point ness;
Acoustic analysis of the user's speech converted into the electrical signal,
An acoustic analysis unit for voice recognition for calculating a feature parameter for voice recognition; based on the feature parameter for voice recognition, the recognition start time is shifted one after another to perform voice recognition, and corresponding to each recognition start time. A voice recognition means for calculating a likelihood for each voice recognition result; a sum or product of the likelihood of the voice recognition result for each recognition start time and the likelihood of the third utterance start point is calculated according to the time, and A voice recognition device, comprising: a voice recognition result determining means for determining a voice recognition result corresponding to a recognition start time at which the calculated value becomes maximum as a voice recognition result for a user's utterance.

【0024】第13の発明は、音声を用いてユーザとの
対話を行う音声対話装置に適用される音声認識装置にお
いて;前記音声対話装置のシステムアナウンスに対する
ユーザの発話開始時刻の極大点を有する予測分布を格納
する予測分布格納手段と;前記格納された予測分布に基
づき、ユーザの発話が開始される期待値を第1の発話開
始点らしさとしてシステムアナウンス開始後の時刻に応
じて算出する第1の演算手段と;電気信号に変換された
ユーザの発話を音響分析し、発話検出用の特徴パラメー
タを算出する発話検出用の音響分析手段と;前記発話検
出用の特徴パラメータに基づき、ユーザの発話が開始さ
れたであろう尤度を第2の発話開始点らしさとして時刻
に応じて算出する第2の演算手段と;第2の発話開始点
らしさに対して第1の発話開始点らしさにより重み付け
を行い、この重み付けされた値を第3の発話開始点らし
さとして時刻に応じて算出する第3の演算手段と;前記
電気信号に変換されたユーザの発話を音響分析し、音声
認識用の特徴パラメータを算出する音声認識用の音響分
析手段と;前記音声認識用の特徴パラメータに基づき、
先頭に無音状態を有する確率付き有限状態ネットワーク
を探索して音声認識を行う音声認識手段と;前記確率付
き有限状態ネットワークの先頭の無音状態から文の先頭
状態へ遷移する確率を、第3の発話開始点らしさを用い
て時刻に応じて更新する遷移確率更新手段と;を具備す
ることを特徴とする音声認識装置である。
A thirteenth aspect of the present invention is a voice recognition device applied to a voice dialogue device for dialogue with a user using voice; prediction having a maximum point of a user's utterance start time with respect to a system announcement of the voice dialogue device. A predictive distribution storage means for storing a distribution; a first calculating an expected value at which a user's utterance starts based on the stored predictive distribution as a first utterance start point in accordance with a time after a system announcement is started An utterance detecting acoustic analysis means for acoustically analyzing a user's utterance converted into an electric signal to calculate a utterance detecting characteristic parameter; and a user utterance based on the utterance detecting characteristic parameter. Second computing means for calculating the likelihood that the utterance has started as the second utterance starting point likelihood according to the time; Third computing means for performing weighting according to the utterance start point likelihood of the user and calculating the weighted value as the third utterance start point likelihood in accordance with time; acoustic analysis of the user's utterance converted into the electric signal. And an acoustic analysis unit for voice recognition for calculating a feature parameter for voice recognition;
A speech recognition means for performing speech recognition by searching a finite state network with a probability having a silent state at the beginning; a probability of transition from a silent state at the beginning of the finite state network with probability to a beginning state of a sentence, a third utterance And a transition probability updating means for updating according to the time using the likelihood of the starting point;

【0025】そして第14の発明は、第7ないし第13
の発明において、前記ユーザの発話開始時刻の予測分布
の極大点がシステムアナウンスの無音区間に存在するこ
とを特徴とし、第15の発明は更に前記無音区間はその
長さが0.2秒以上3秒以下であり、システムアナウン
スの文と文の間及び文節と文節との間のうち少なくとも
一方に存在することを特徴とする。
The fourteenth invention is the seventh to thirteenth inventions.
In the invention, the maximum point of the predicted distribution of the utterance start time of the user exists in the silent section of the system announcement, and the fifteenth invention further has a length of 0.2 seconds or more in the silent section. It is less than a second, and is characterized by being present in at least one of the sentence of the system announcement and the sentence and the phrase.

【0026】次に第16の発明は、第7ないし第15の
発明の音声認識装置と、システムアナウンスの指定され
たテキストを電気的音声信号に変換すると共にシステム
アナウンスの開始を前記音声認識装置に通知するアナウ
ンス発声装置と、このアナウンス発声装置に対するシス
テムアナウンスのテキストの指定及び前記音声認識装置
からの音声認識結果の入力により音声を用いたユーザと
の対話を管理する対話管理装置とを具備することを特徴
とする音声対話装置である。
A sixteenth aspect of the present invention is the speech recognition apparatus according to the seventh to fifteenth aspects of the invention, which converts the designated text of the system announcement into an electrical voice signal and starts the system announcement by the speech recognition apparatus. An announcement voicing device for notifying, and a dialogue management device for managing a dialogue with a user by using a voice by designating a text of a system announcement for the announcement voicing device and inputting a voice recognition result from the voice recognition device. Is a voice interaction device.

【0027】[0027]

【作用】第1,第2及び第7〜第11の発明では、音響
分析で得た発話検出用の特徴パラメータからユーザの発
話が開始されたであろう尤度(第2の発話開始点らし
さ)を求めて発話開始時刻を決定する際に、予測分布か
ら得た第1の発話開始点らしさで第2の発話開始点らし
さ又は基準値に対して重み付けを行う。これにより、ユ
ーザの発話開始時刻を高精度に一点決定することがで
き、音声認識の精度が向上する。またユーザの発話開始
時刻を高精度に一点決定することができることから、シ
ステムアナウンス中のユーザの割り込み発話を高精度に
音声認識することができ、音声対話装置の利用時間の短
縮が可能となる。
In the first, second and seventh to eleventh inventions, the likelihood that the user's utterance may have started from the characteristic parameter for utterance detection obtained by the acoustic analysis (the second utterance starting point ) Is determined to determine the utterance start time, the second utterance start point likelihood or the reference value is weighted by the first utterance start point likelihood obtained from the prediction distribution. As a result, it is possible to highly accurately determine one point of the speech start time of the user, and the accuracy of voice recognition is improved. Further, since the user's utterance start time can be determined with high accuracy, the interrupt utterance of the user during the system announcement can be recognized with high accuracy, and the use time of the voice dialog device can be shortened.

【0028】第3,第4,第12及び第13の発明では
ユーザの発話開始時刻を一点に決定することなく、高精
度な音声認識を可能とする。
In the third, fourth, twelfth and thirteenth inventions, highly accurate voice recognition is possible without deciding the user's utterance start time at one point.

【0029】まず第3及び第12の発明では、音声認識
をその開始時刻を次々にずらして多数行い、各認識開始
時刻に対応した音声認識結果毎の尤度を求め、この尤度
と第1の発話開始点らしさで第2の発話開始点らしさに
重み付けして得た第3の発話開始点らしさとから、最適
な音声認識結果を判定する。これにより、高精度な音声
認識を行うことができる。なお、第3の発話開始点らし
さが所定レベルを超えた時刻から音声認識を開始するこ
とも可能であり、これにより音声認識の処理量が低減す
る。第3及び第12の発明ではユーザの発話開始時刻を
高精度に一点決定することができなくても、結果的に音
声認識の精度が向上する。
First, in the third and twelfth inventions, a large number of speech recognitions are performed with their start times being shifted one after another, and the likelihood for each speech recognition result corresponding to each recognition start time is obtained. The optimal speech recognition result is determined from the third utterance start point likelihood obtained by weighting the second utterance start point likelihood by the utterance start point likelihood. As a result, highly accurate voice recognition can be performed. It is also possible to start the voice recognition from the time when the likelihood of the third utterance start point exceeds a predetermined level, which reduces the processing amount of the voice recognition. In the third and twelfth inventions, even if it is not possible to highly accurately determine one point of the utterance start time of the user, the accuracy of voice recognition is improved as a result.

【0030】次に第4及び第13の発明では、先頭に無
音状態を有する確率付き有限状態ネットワークを探索す
ることにより音声認識を行うものとする。その際に、先
頭の無音状態から文の先頭状態へ遷移する確率を、第1
の発話開始点らしさで第2の発話開始点らしさを重み付
けして得た第3の発話開始点らしさを用いて変化させ
る。従って、実質的な音声認識は発話開始が不確かな間
は行われず、最も確からしい発話開始時刻になってから
開始されることになり、高精度な音声認識を行うことが
できる。第4及び第13の発明では、ユーザの発話開始
時刻を高精度に一点決定することができなくても結果的
に音声認識の精度が向上し、更に第3及び第12の発明
に比べると、音声認識を開始時間を次々にずらして並列
的に行う必要がないから、高速な処理が可能となり、ま
たメモリ容量を削減することができる。
Next, in the fourth and thirteenth inventions, it is assumed that speech recognition is performed by searching a probability-added finite state network having a silent state at the head. At that time, the probability of transition from the silent state at the head to the head state of the sentence is
The third utterance start point likelihood obtained by weighting the second utterance start point likelihood is changed using the utterance start point likelihood. Therefore, substantial voice recognition is not performed while the utterance start is uncertain, and is started at the most probable utterance start time, and highly accurate voice recognition can be performed. In the fourth and thirteenth inventions, even if the user's utterance start time cannot be determined with high accuracy, the accuracy of the voice recognition is improved as a result, and compared with the third and twelfth inventions, Since it is not necessary to perform voice recognition in parallel by shifting start times one after another, high-speed processing is possible and the memory capacity can be reduced.

【0031】第5,第6,第14及び第15の発明で
は、より信頼性が高いユーザの発話が開始される期待値
を求めるための予測分布を得る。発明者等は、システム
アナウンスとユーザの発話開始時刻との間にどのような
因果関係があるかを調べた。これは、特徴的な因果関係
があれば、これを利用することによりユーザの発話開始
時刻を精度良く検出することができると考えたからであ
る。
In the fifth, sixth, fourteenth and fifteenth inventions, a prediction distribution for obtaining an expected value at which the utterance of a user with higher reliability is started is obtained. The inventors investigated what kind of causal relationship exists between the system announcement and the user's utterance start time. This is because, if there is a characteristic causal relationship, it is considered that the user's utterance start time can be accurately detected by using this.

【0032】具体的には、多数のユーザに音声対話装置
を利用してもらい、システムアナウンスの開始後にユー
ザが発話を開始する場合のその時刻と頻度とを調べると
いう実験を行った。その結果、ユーザの発話開始時刻が
極大点を持つ分布をすることが判った。特に、システム
アナウンスに割り込んでユーザが発話する場合は、第5
及び第14の発明のように発話開始時刻がシステムアナ
ウンスの無音区間を中心に分布することが判り、更に第
6及び第15の発明のように文と文の間あるいは文節と
文節との間に積極的に一定の無音区間を設けると、分布
の山が急峻になり、この傾向は無音区間を好ましくは
0.2秒〜3秒(より好ましくは0.4〜1.5秒)と
すると顕著であることが判った。また、システムアナウ
ンス終了後にユーザが発話を開始する場合も、システム
アナウンス終了直後を中心に発話開始時刻が特定の分布
をすることが判った。なお、無音区間とは音が全く存在
しない場合だけでなく、例えばチャイムやバックグラン
ドミュージックが流れている場合などでも、システムア
ナウンスにとって実質的に無音状態といえる場合は無音
区間である。無音区間はユーザの発話開始を促すように
制御する。
More specifically, an experiment was conducted in which a large number of users were made to use the voice dialog device and the time and frequency at which the user started speaking after the system announcement was started were examined. As a result, it was found that the utterance start time of the user has a distribution having a maximum point. In particular, if the user speaks after interrupting the system announcement, the fifth
It can be seen that the utterance start time is distributed mainly in the silent section of the system announcement as in the fourteenth invention, and further between the sentences or between the phrases as in the sixth and fifteenth inventions. When a certain silent section is positively provided, the distribution peak becomes steep, and this tendency is remarkable when the silent section is preferably 0.2 seconds to 3 seconds (more preferably 0.4 to 1.5 seconds). Was found. It was also found that when the user starts uttering after the system announcement ends, the utterance start time has a specific distribution around the time immediately after the system announcement ends. It should be noted that the silent section is a silent section not only when there is no sound at all but also when a chime or background music is played and when it can be said to be a substantially silent state for the system announcement. The silent section is controlled so as to prompt the user to start speaking.

【0033】そこで、このような実験に基づき図2に示
すようなシステムアナウンスに対するユーザの発話開始
時刻の極大点を有する予測分布100を予め作成して用
意するか、或いは、実験によらずとも無音区間もしくは
その前後に極大点を持つように正規分布、ポアソン分
布、カイ2乗分布等の確率分布を用いてシステムアナウ
ンスに対するユーザの発話開始時刻の極大点を有する予
測分布を予め用意しておくことより、システムアナウン
ス開始後の時に応じてユーザの発話が開始されるであろ
う期待値(第1の発話開始点らしさ)を求めることがで
きる。
Therefore, based on such an experiment, a prediction distribution 100 having the maximum point of the user's utterance start time for the system announcement as shown in FIG. 2 is prepared in advance and prepared, or no sound is produced regardless of the experiment. Prepare a prediction distribution that has the maximum point of the user's utterance start time for the system announcement by using a probability distribution such as a normal distribution, Poisson distribution, and chi-square distribution so that it has a maximum point before or after the interval. Thus, it is possible to obtain the expected value (likely as the first utterance start point) at which the user's utterance will start depending on the time after the system announcement is started.

【0034】第16の発明では、高精度な音声認識の下
で、ユーザと装置間で対話を行うことができる。
In the sixteenth invention, the user and the device can talk with each other under highly accurate voice recognition.

【0035】[0035]

【実施例】以下、図面を参照して発明の実施例を説明す
る。図面中、図1には第1実施例に係る音声対話装置の
ブロック構成が示されている。図2にはシステムアナウ
ンスに対するユーザの発話開始時刻の予測分布を実験に
より観測して得た例が示されている。また、図3には第
2実施例に係る音声対話装置のブロック構成が示され、
図4には第3実施例に係る音声対話装置のブロック構成
が示され、図5には第4実施例に係る音声対話装置のブ
ロック構成が示されている。図6には先頭に無音状態を
有する確率付き有限状態ネットワークの一例が示されて
いる。
Embodiments of the invention will be described below with reference to the drawings. In the drawings, FIG. 1 shows a block configuration of a voice dialogue apparatus according to the first embodiment. FIG. 2 shows an example obtained by experimentally observing the predicted distribution of the user's utterance start time for the system announcement. Further, FIG. 3 shows a block configuration of a voice dialog device according to the second embodiment.
FIG. 4 shows a block configuration of a voice dialogue apparatus according to the third embodiment, and FIG. 5 shows a block configuration of a voice dialogue apparatus according to the fourth embodiment. FIG. 6 shows an example of a finite state network with probability having a silent state at the beginning.

【0036】<第1実施例>図1に示されるように、第
1実施例に係る音声対話装置は、対話管理装置1と、ア
ナウンス発声装置2と、音声認識装置10とを具備した
ものであり、音声出力装置3及び音声入力装置5は必要
に応じて音声対話装置に内蔵されたり、あるいは音声対
話装置とは離れた別物で適宜接続されたりする。音声対
話装置が内線電話受付システムに用いられる場合は、電
話機の送受話器が音声出力回路3と音声入力回路5に相
当し、電話回線及び電話交換機を通して音声対話装置に
接続される。音声認識装置10は予測分布格納部11
と、第1の発話開始点らしさの演算部12と、発話検出
用の音響分析部13と、第2の発話開始点らしさの演算
部14と、第3の発話開始点らしさの演算部15と、発
話開始時刻決定部16と、音声認識用の音声信号通過ス
イッチ17と、音声認識用の音響分析部18と、音声認
識部19とを具備している。
<First Embodiment> As shown in FIG. 1, a voice dialogue apparatus according to the first embodiment comprises a dialogue management apparatus 1, an announcement utterance apparatus 2, and a voice recognition apparatus 10. Therefore, the voice output device 3 and the voice input device 5 may be built in the voice dialogue device as necessary, or may be appropriately connected as a separate object apart from the voice dialogue device. When the voice dialogue device is used in the extension telephone reception system, the handset of the telephone corresponds to the voice output circuit 3 and the voice input circuit 5, and is connected to the voice dialogue device through the telephone line and the telephone exchange. The speech recognition device 10 includes a prediction distribution storage unit 11
A first utterance starting point likelihood calculating section 12, an utterance detecting acoustic analysis section 13, a second utterance starting point likelihood calculating section 14, and a third utterance starting point likelihood calculating section 15. An utterance start time determination unit 16, a voice signal passage switch 17 for voice recognition, a sound analysis unit 18 for voice recognition, and a voice recognition unit 19 are provided.

【0037】アナウンス発声装置2は、対話管理装置1
がコード名等により指定したシステムアナウンスのテキ
スト1aに基づいて、発声すべき音声の電気信号2aを
作成し、音声出力回路3に送る。この時、アナウンス発
声装置2は図2に示すように、システムアナウンスの文
と文の間、または文節と文節との間に一定の無音区間2
00を設けて、音声の電気信号2aを作成する。本実施
例においては無音区間200の長さを0.5秒程度とし
てあるが、一般には0.2秒以上3秒以下が妥当であ
り、より好ましくは0.4秒以上1.5秒以下とする。
無音区間が長すぎると、ユーザに不安感を与える。無音
区間とは信号が全く存在しない場合だけでなく、例えば
チャイムやバックグラウンドミュージックが流れている
場合などでもシステムアナウンスにとって実質的な無音
状態であれば無音区間となる。また、アナウンス発声装
置2はシステムアナウンスの開始を表わすアナウンス開
始信号2bを音声認識装置10に送る。なお、システム
アナウンスの開始とはユーザに対して音声が出始める時
点そのものだけを言うのではなく、音声の出始めよりも
一定時間前をもってシステムアナウンスの開始としても
良い。
The announcement utterance device 2 is the dialogue management device 1.
Generates an electric signal 2a of a voice to be uttered based on the system announcement text 1a designated by the code name or the like and sends it to the voice output circuit 3. At this time, as shown in FIG. 2, the announcement voicing device 2 has a certain silent section 2 between sentences of the system announcement or between phrases.
00 is provided to create a voice electric signal 2a. In the present embodiment, the length of the silent section 200 is set to about 0.5 seconds, but generally 0.2 seconds or more and 3 seconds or less is appropriate, and more preferably 0.4 seconds or more and 1.5 seconds or less. To do.
If the silent section is too long, the user feels uneasy. The silent section is a silent section not only when there is no signal at all but also when there is substantially no sound for the system announcement, for example, when a chime or background music is playing. Further, the announcement utterance device 2 sends an announcement start signal 2b indicating the start of the system announcement to the voice recognition device 10. It should be noted that the start of the system announcement does not mean only the time itself at which the voice starts to be output to the user, but the system announcement may be started at a certain time before the start of the voice output.

【0038】音声出力回路3はアナウンス発声装置2か
ら送られてきた電気信号2aを音声に変換して、システ
ムアナウンス3aとしてユーザに聞かせる。このシステ
ムアナウンス3aに対してユーザの発話4があるので、
この発話4を音声入力回路5が電気信号5aに変換して
音声認識装置10に送る。
The voice output circuit 3 converts the electric signal 2a sent from the announcement utterance device 2 into a voice and makes the user hear it as a system announcement 3a. Since there is a user utterance 4 for this system announcement 3a,
The speech input circuit 5 converts the utterance 4 into an electric signal 5a and sends it to the speech recognition apparatus 10.

【0039】音声認識装置10では、予測分布格納部1
1に図2に示すようなシステムアナウンスに対するユー
ザの発話開始時刻の予測分布100を格納してある。こ
の予測分布100は、予め500名程度のユーザに内線
電話受付システムの音声対話装置を利用させて同装置か
ら文と文の間に0.5秒程度の無音区間200を設けた
システムアナウンスを発声させた場合の各ユーザの発話
開始時刻の分布を観測した実験結果から作成したもので
ある。図2中で、横軸はシステムアナウンスの開始を時
刻0とした場合の時刻tをとり、縦軸は各時刻tでユー
ザの発話が開始される期待値を表わしており、各無音区
間200に分布の極大点がある。なお、実験によらずと
も、無音区間もしくはその前後に極大点を持つ正規分
布、ポアソン分布あるいはカイ2重分布などの確率分布
を用いることにより、システムアナウンスの開始を時刻
0とした場合の各時刻tにおいてユーザの発話が開始さ
れる期待値の分布を作成して、システムアナウンスに対
するユーザの発話開始時刻の極大点を有する予測分布と
しても良い。
In the voice recognition device 10, the prediction distribution storage unit 1
1 stores a predicted distribution 100 of user's utterance start times for system announcements as shown in FIG. This predicted distribution 100 is made by a system announcement in which about 500 users have previously used the voice conversation device of the extension telephone reception system and provided a silent interval 200 of about 0.5 seconds between sentences from the device. It is created from the experimental result of observing the distribution of the utterance start time of each user in the case of making it. In FIG. 2, the horizontal axis represents time t when the start of the system announcement is time 0, and the vertical axis represents the expected value at which the user's utterance starts at each time t. There is a maximum in the distribution. It should be noted that, regardless of the experiment, by using a probability distribution such as a normal distribution, a Poisson distribution, or a chi-double distribution having a maximum point before or after the silent section, each time when the system announcement is started at time 0 It is also possible to create a distribution of expected values at which the user's utterance starts at t and use it as a predictive distribution having a maximum point of the user's utterance start time with respect to the system announcement.

【0040】演算部12はアナウンス発声装置2よりア
ナウンス開始信号2bを受けた時点から時間tに応じ
て、予測分布格納部11の予測分布に基づいて、第1の
発話開始点らしさとして、時刻tでユーザの発話が開始
されるであろう期待値a(t)を算出し、演算部15に
送る。
Based on the predicted distribution in the predicted distribution storage unit 11, the arithmetic unit 12 determines the time t from the time when the announce start signal 2b is received from the announcer 2 based on the predicted distribution in the predicted distribution storage unit 11. An expected value a (t) at which the user's utterance will start is calculated and sent to the calculation unit 15.

【0041】発話検出用の音響分析部13は音声入力回
路5から与えられる電気的音声信号5aを入力して常時
音響分布を行い、発話検出用の特徴パラメータ13aを
次々に算出して演算部14に送る。
The speech detection acoustic analysis unit 13 inputs the electrical voice signal 5a provided from the voice input circuit 5 to constantly perform acoustic distribution, successively calculates the characteristic parameters 13a for speech detection, and the arithmetic unit 14 Send to.

【0042】演算部14は発話検出用の特徴パラメータ
13aに基づいて、第2の発話開始点らしさとして、ユ
ーザの発話が開始されたであろう尤度b(t)を時間t
に応じて算出し、演算部15に送る。但し、システムア
ナウンスの開始を時刻0とする。
Based on the utterance detection characteristic parameter 13a, the calculation unit 14 determines the likelihood b (t) at which the utterance of the user is likely to be started at time t as the second utterance starting point.
According to the above, and sends it to the calculation unit 15. However, the start of the system announcement is time 0.

【0043】演算部15は第1の発話開始点らしさa
(t)により第2の発話開始点らしさb(t)に重み付
けを行い、第3の発話開始点らしさα(t)を算出し、
発話開始時刻決定部16に送る。ここで、重み付けの例
として式(1)〜式(3)をあげておく。但し、式
(2)中、0<k1 <1である。
The arithmetic unit 15 determines whether the first utterance starting point is a.
The second speech start point likelihood b (t) is weighted by (t) to calculate the third speech start point likelihood α (t),
It is sent to the utterance start time determination unit 16. Equations (1) to (3) are given here as examples of weighting. However, in the formula (2), 0 <k 1 <1.

【数1】 α(t)=a(t)+b(t) …式(1) α(t)=k1 ・a(t)+(1−k1 )・b(t) …式(2) α(t)=a(t)・b(t) …式(3)Α (t) = a (t) + b (t) Equation (1) α (t) = k 1 · a (t) + (1-k 1 ) · b (t) Equation (2) ) Α (t) = a (t) · b (t) (3)

【0044】発話開始時刻決定部16は第3の発話開始
点らしさα(t)と予め固定した基準値Refとを比較
し、最初にα(t)>Refとなった時点、もしくはα
(t)>Refが或る一定時間続いたら最初にα(t)>
efとなった時点をユーザの発話開始時刻と決定して、
その旨を表わす発話開始信号16aを音声認識用の音声
信号通過スイッチ17に送る。
The utterance start time determination unit 16 compares the third utterance start point likelihood α (t) with a reference value R ef fixed in advance, and when α (t)> R ef first, or α
When (t)> R ef continues for a certain period of time, first α (t)>
The time when it becomes Ref is determined as the utterance start time of the user,
An utterance start signal 16a indicating that is sent to the voice signal passing switch 17 for voice recognition.

【0045】このスイッチ17は発話開始時刻16aを
与えられた時点からオンとなり、音声信号5aを通過さ
せ、音声認識対象の信号17aとして音声認識用の音響
分析部18に送る。
The switch 17 is turned on from the time when the utterance start time 16a is given, passes the voice signal 5a, and sends it as the voice recognition target signal 17a to the voice analysis acoustic analysis section 18.

【0046】音響分析部18ではスイッチ17を通過し
た音声信号17aを音響分析して音声認識用の特徴パラ
メータ18aを次々に算出し、音声認識部19に送る。
The acoustic analysis unit 18 acoustically analyzes the voice signal 17a that has passed through the switch 17, successively calculates the characteristic parameters 18a for voice recognition, and sends them to the voice recognition unit 19.

【0047】音声認識部19では音声認識用の特徴パラ
メータ18aに基づいて音声認識を行う。その認識結果
19aは対話管理装置1に送られる。
The voice recognition unit 19 performs voice recognition based on the feature parameter 18a for voice recognition. The recognition result 19a is sent to the dialogue management device 1.

【0048】対話管理装置1では認識結果19aに基づ
いて、次に発声すべきシステムアナウンスのテキスト1
aを決定し、アナウンス発声装置2にコード名等を送
る。また、ユーザとの対話内容からユーザの意思を認識
して、例えば内線電話受付システムであれば内線番号の
情報1bを外部に出力する。各装置1,2,10が上述
した動作を繰り返すことにより対話が行われる。
In the dialogue management device 1, the text 1 of the system announcement to be spoken next is based on the recognition result 19a.
A is determined, and the code name or the like is sent to the announcement voicing device 2. Further, the user's intention is recognized from the content of the dialogue with the user, and, for example, in the case of an extension telephone reception system, the extension number information 1b is output to the outside. A dialogue is carried out by repeating the above-described operation of each of the devices 1, 2 and 10.

【0049】上述した第1実施例の説明ではスイッチ1
7を用いて音声認識対象の信号17aのみを音響分析部
18に与えているが、スイッチ17を用いずに次のよう
に変更しても良い。 (1)音声入力回路5からの音声信号5aを常時音響分
析部18に送り、且つ発話開始時刻決定部16から発話
開始信号16aを音響分析部18に送るものとし、音響
分析部18は発話開始信号16aを与えられた時点から
音響分析を開始する。 (2)あるいは、音声入力回路5からの音声信号5aを
常時音響分析部18に送り、且つ音響分析部18は常時
音響分析を行って特徴パラメータ18aを音声認識部1
9に送り、更に発話開始時刻決定部16から発話開始信
号16aを音声認識部19に送るものとし、音声認識部
19は発話開始信号16aを与えられた時点からの特徴
パラメータ18aを用いて音声認識を開始する。
In the above description of the first embodiment, the switch 1
Although only the signal 17a of the voice recognition target is given to the acoustic analysis unit 18 using the switch 7, it may be changed as follows without using the switch 17. (1) It is assumed that the voice signal 5a from the voice input circuit 5 is constantly sent to the acoustic analysis unit 18, and the utterance start time determination unit 16 sends the utterance start signal 16a to the acoustic analysis unit 18, which starts the utterance. The acoustic analysis is started from the time when the signal 16a is given. (2) Alternatively, the voice signal 5a from the voice input circuit 5 is constantly sent to the acoustic analysis unit 18, and the acoustic analysis unit 18 always performs acoustic analysis to determine the characteristic parameter 18a as the voice recognition unit 1.
9, and the speech start time determination unit 16 further transmits the speech start signal 16a to the speech recognition unit 19. The speech recognition unit 19 performs speech recognition using the characteristic parameter 18a from the time when the speech start signal 16a is given. To start.

【0050】<第2実施例>図3に示されるように、第
2実施例に係る音声対話装置は、対話管理装置1と、ア
ナウンス発声装置2と、音声認識装置20とを具備した
ものであり、音声出力装置3及び音声入力装置5は必要
に応じて音声対話装置に内蔵されたり、あるいは音声対
話装置とは離れた別物で適宜接続されたりする。音声認
識装置20は予測分布格納部11と、第1の発話開始点
らしさの演算部12と、発話検出用の音響分析部13
と、第2の発話開始点らしさの演算部14と、基準値演
算部21と、発話開始時刻決定部22と、音声認識用の
音声信号通過スイッチ17と、音声認識用の音響分析部
18と、音声認識部19とを具備している。これら各装
置のうち、演算部12及び14と、基準値演算部21及
び発話開始時刻決定部22とが図1に示した第1実施例
と異なり、他のもの1,2,3,5,11,13及び1
7〜19は第1実施例における同符号のものと同機能で
あるから説明を簡単にする。
<Second Embodiment> As shown in FIG. 3, the voice dialogue apparatus according to the second embodiment comprises a dialogue management apparatus 1, an announcement utterance apparatus 2, and a voice recognition apparatus 20. Therefore, the voice output device 3 and the voice input device 5 may be built in the voice dialogue device as necessary, or may be appropriately connected as a separate object apart from the voice dialogue device. The voice recognition device 20 includes a prediction distribution storage unit 11, a first utterance start point likelihood calculation unit 12, and an utterance detection acoustic analysis unit 13.
A second utterance starting point likelihood calculating unit 14, a reference value calculating unit 21, a utterance start time determining unit 22, a voice signal passing switch 17 for voice recognition, and a sound analyzing unit 18 for voice recognition. , And a voice recognition unit 19. Of these devices, the calculation units 12 and 14, the reference value calculation unit 21, and the utterance start time determination unit 22 are different from the first embodiment shown in FIG. 11, 13 and 1
Since 7 to 19 have the same functions as those of the same reference numerals in the first embodiment, the description will be simplified.

【0051】演算部12は予測分布格納部11に格納さ
れている図2に示すような予測分布に基づいて時刻tに
応じて算出した第1の発話開始点らしさa(t)を、基
準値演算部21に送る。演算部14は音響分析部13か
らの発話検出用の特徴パラメータ13aに基づいて時刻
tに応じて算出した第2の発話開始点らしさb(t)
を、発話開始時刻決定部22に送る。
The calculation unit 12 calculates the first utterance start point likelihood a (t) calculated at the time t based on the prediction distribution stored in the prediction distribution storage unit 11 as shown in FIG. It is sent to the calculation unit 21. The calculation unit 14 calculates the second utterance start point likelihood b (t) calculated according to the time t based on the utterance detection feature parameter 13a from the acoustic analysis unit 13.
Is sent to the utterance start time determination unit 22.

【0052】基準値演算部21は第1の基準値Refo
第1の発話開始点らしさa(t)により重み付けして、
時間tに応じて変化する第2の基準値Ref(t)を算出
し、発話開始時刻決定部22に送る。ここで重み付けの
例として式(4)〜式(5)をあげておく。但し、式
(5)中で、0<k2 とする。
The reference value calculation unit 21 weights the first reference value Refo by the first speech start point likelihood a (t),
A second reference value Ref (t) that changes according to the time t is calculated and sent to the utterance start time determination unit 22. Equations (4) to (5) are given here as examples of weighting. However, in the formula (5), 0 <k 2 .

【数2】 Ref(t)=S/α(t) …式(4) Ref(t)=S−k2 ・α(t) …式(5)[Equation 2] R ef (t) = S / α (t) Equation (4) R ef (t) = S−k 2 · α (t) Equation (5)

【0053】発話開始決定部22は第2の発話開始点ら
しさb(t)と重み付けされた第2の基準値Ref(t)
とを比較し、最初にb(t)>Ref(t)となった時
点、もしくはb(t)>Ref(t)が或る一定時間続い
たら最初にb(t)>Ref(t)となった時点をユーザ
の発話開始時刻と決定し、その旨を表わす発話開始信号
22aを音声信号通過スイッチ17に送る。
The utterance start determination unit 22 weights the second utterance start point likelihood b (t) and the second reference value R ef (t).
And b (t)> R ef (t) for the first time, or b (t)> R ef (t) continues for a certain period of time, then b (t)> R ef (t) for the first time. The time point t) is determined as the utterance start time of the user, and the utterance start signal 22a indicating that is sent to the voice signal passing switch 17.

【0054】このスイッチ17は発話開始信号22aを
与えられた時点からオンとなり、オンの間に送られてき
た音声信号17aのみを音声認識対象として音響分析部
18に送る。音響分析部18では、音声信号通過スイッ
チ17を通過した音声信号17aから、音声認識に適し
た特徴パラメータ18aを算出し、音声認識部19に送
る。音声認識部19では、特徴パラメータ18aに基づ
いて音声認識を行い、その認識結果19aを対話管理装
置1に送る。対話管理装置1では、音声認識部19から
与えられる認識結果19aに基づいて、次に発声すべき
システムアナウンスのテキスト1aを決定してアナウン
ス発声装置2にコード名等を送る。
The switch 17 is turned on from the point of time when the utterance start signal 22a is given, and sends only the voice signal 17a sent while it is on to the acoustic analysis unit 18 as a voice recognition target. The acoustic analysis unit 18 calculates a characteristic parameter 18a suitable for voice recognition from the voice signal 17a that has passed through the voice signal passing switch 17, and sends it to the voice recognition unit 19. The voice recognition unit 19 performs voice recognition based on the characteristic parameter 18a and sends the recognition result 19a to the dialogue management device 1. The dialogue management device 1 determines the text 1a of the system announcement to be uttered next based on the recognition result 19a provided from the voice recognition unit 19 and sends the code name or the like to the announcement utterance device 2.

【0055】上述した第2実施例の説明でもスイッチ1
7を用いて音声認識対象の信号17aのみを音響分析部
18に与えているが、スイッチ17を用いずに次のよう
に変更しても良い。 (1)音声入力回路5からの音声信号5aを常時音響分
析部18に送り、且つ発話開始時刻決定部22から発話
開始信号22aを音響分析部18に送るものとし、音響
分析部18は発話開始信号22aを与えられた時点から
音響分析を開始する。 (2)あるいは、音声入力回路5からの音声信号5aを
常時音響分析部18に送り、且つ音響分析部18は常時
音響分析を行って特徴パラメータ18aを音声認識部1
9に送り、更に発話開始時刻決定部22から発話開始信
号22aを音声認識部19に送るものとし、音声認識部
19は発話開始信号22aを与えられた時点からの特徴
パラメータ18aを用いて音声認識を開始する。
The switch 1 is also used in the description of the second embodiment.
Although only the signal 17a of the voice recognition target is given to the acoustic analysis unit 18 using the switch 7, it may be changed as follows without using the switch 17. (1) It is assumed that the voice signal 5a from the voice input circuit 5 is always sent to the acoustic analysis unit 18, and the utterance start time determination unit 22 sends the utterance start signal 22a to the acoustic analysis unit 18, which starts the utterance. The acoustic analysis is started from the time when the signal 22a is given. (2) Alternatively, the voice signal 5a from the voice input circuit 5 is constantly sent to the acoustic analysis unit 18, and the acoustic analysis unit 18 always performs acoustic analysis to determine the characteristic parameter 18a as the voice recognition unit 1.
9, and the utterance start time determination unit 22 further sends the utterance start signal 22a to the voice recognition unit 19. The voice recognition unit 19 performs the voice recognition using the characteristic parameter 18a from the time when the utterance start signal 22a is given. To start.

【0056】<第3実施例>図4に示されるように、第
3実施例に係る音声対話装置は、対話管理装置1と、ア
ナウンス発声装置2と、音声認識装置30とを具備した
ものであり、音声出力装置3及び音声入力装置5は必要
に応じて音声対話装置に内蔵されたり、あるいは音声対
話装置とは離れた別物で適宜接続されたりする。音声認
識装置30は予測分布格納部11と、第1の発話開始点
らしさの演算部12と、発話検出用の音響分析部13
と、第2の発話開始点らしさの演算部14と、第3の発
話開始点らしさの演算部15と、音声認識用の音響分析
部18と、音声認識部31と、音声認識結果判定部32
とを具備している。
<Third Embodiment> As shown in FIG. 4, the voice dialogue apparatus according to the third embodiment comprises a dialogue management apparatus 1, an announcement utterance apparatus 2, and a voice recognition apparatus 30. Therefore, the voice output device 3 and the voice input device 5 may be built in the voice dialogue device as necessary, or may be appropriately connected as a separate object apart from the voice dialogue device. The speech recognition device 30 includes a prediction distribution storage unit 11, a first utterance start point likelihood calculation unit 12, and an utterance detection acoustic analysis unit 13.
The second utterance starting point likelihood calculating unit 14, the third utterance starting point likelihood calculating unit 15, the voice analysis acoustic analysis unit 18, the voice recognizing unit 31, and the voice recognition result determining unit 32.
Is provided.

【0057】第3実施例の各装置構成要素のうち、演算
部15と、音声認識部31及び音声認識結果判定部32
とが図1に示した第1実施例と異なり、また第1実施例
における発話開始時刻決定部16及びスイッチ17が存
在しないが、他のもの1,2,3,5,11〜14及び
18は第1実施例の同符号のものと同機能であるから説
明を簡単にする。
Among the components of each device of the third embodiment, the arithmetic unit 15, the voice recognition unit 31, and the voice recognition result determination unit 32.
Is different from the first embodiment shown in FIG. 1, and the utterance start time determining unit 16 and the switch 17 in the first embodiment are not present, but the other ones 1, 2, 3, 5, 11 to 14 and 18 Has the same function as that of the same reference numeral in the first embodiment, the description will be simplified.

【0058】演算部15は前述した式(1)〜式(3)
を用いて、第1の発話開始点らしさa(t)により第2
の発話開始点らしさb(t)に対して重み付けを行い、
第3の発話開始点らしさα(t)を時間tに応じて算出
するが、これは音声認識結果判定部32に送る。なお第
1の発話開始点らしさa(t)は、予測分布格納部11
に格納されている図2に示したような予測分布に基づい
て、時刻tでユーザの発話が開始されるであろう期待値
を演算部12が算出することにより求まる。また第2の
発話開始点らしさb(t)は、音響分析部13が常時音
響分析して得られる発話検出用の特徴パラメータ13a
に基づいて、時刻tでユーザの発話が開始されたであろ
う尤度を演算部14が算出することにより求まる。但
し、アナウンス発声装置2からアナウンス開始信号2a
が与えられた時を時刻0としている。
The calculation unit 15 uses the equations (1) to (3) described above.
By using the first utterance starting point likelihood a (t)
The utterance starting point like b (t) is weighted,
The third utterance start point likelihood α (t) is calculated according to the time t, and this is sent to the voice recognition result determination unit 32. It should be noted that the first utterance start point likelihood a (t) is calculated by the prediction distribution storage unit 11
2 is calculated by the calculating unit 12 calculating an expected value at which the user's utterance will start at time t, based on the predicted distribution shown in FIG. The second utterance start point likelihood b (t) is the characteristic parameter 13a for utterance detection, which is obtained by the acoustic analysis unit 13 constantly performing acoustic analysis.
Based on, the calculation unit 14 calculates the likelihood that the user's utterance will have started at time t. However, the announcement start signal 2a from the announcement vocalization device 2
Is given as time 0.

【0059】音声認識用の音響分析部18は音声入力回
路5から与えられる音声信号5aを常時音響分析して音
声認識用の特徴パラメータ18aを次々に算出し、音声
認識部31に送る。
The voice-recognition acoustic analysis unit 18 constantly acoustically analyzes the voice signal 5a supplied from the voice input circuit 5 to successively calculate the voice-recognition characteristic parameters 18a, and sends the feature parameters 18a to the voice recognition unit 31.

【0060】音声認識部31では例えば10ミリ秒おき
の各時刻t毎にその時刻tをユーザの発話開始時刻と仮
定することにより、音声認識開始時刻を次々にずらして
複数の音声認識を行い、各時刻tから開始した場合の各
音声認識結果w(t)を音声認識結果判定部32に送る
と共に、各音声認識結果w(t)毎の尤度p(t)を算
出して音声認識結果判定部32に送る。
The voice recognition unit 31 performs a plurality of voice recognitions by shifting the voice recognition start time one after another by assuming that the time t is the user's utterance start time at each time t of every 10 milliseconds, for example. Each voice recognition result w (t) when starting from each time t is sent to the voice recognition result determination unit 32, and the likelihood p (t) is calculated for each voice recognition result w (t) to calculate the voice recognition result. It is sent to the determination unit 32.

【0061】音声認識結果判定部32は次式(6)また
は式(7)または式(8)を用いて、各認識開始時刻t
毎の音声認識結果の尤度p(t)と第3の発話開始点ら
しさα(t)とを統合した値q(t)を算出し、この値
q(t)が最大となるような時刻tmax を見い出して、
全ての音声認識結果w(t)のうちで、時刻tmax に対
応した音声認識結果w(tmax )をユーザの発話に対す
る認識結果32aと判定する。対話管理装置1にはこの
音声認識結果32aのみを送る。但し、式(7)中で、
例えば0<k3 <1とする。これにより、ユーザの発話
開始時刻を高精度に一点決定することができなくても、
結果的にユーザの発話を高精度に音声認識することがで
きる。
The voice recognition result judging section 32 uses the following equation (6), equation (7) or equation (8) to recognize each recognition start time t.
A value q (t) is calculated by integrating the likelihood p (t) of each speech recognition result and the third utterance start point likelihood α (t), and the time at which this value q (t) is maximized is calculated. find t max ,
Of all the speech recognition results w (t), the speech recognition result w (t max ) corresponding to the time t max is determined as the recognition result 32a for the user's utterance. Only the voice recognition result 32a is sent to the dialogue management device 1. However, in equation (7),
For example, 0 <k 3 <1. As a result, even if the user's utterance start time cannot be determined with high accuracy,
As a result, the speech of the user can be recognized with high accuracy.

【数3】 q(t)=α(t)+p(t) …式(6) q(t)=(1−k3 )・α(t)+k3 ・p(t) …式(7) q(t)=α(t)・p(t) …式(8)[Number 3] q (t) = α (t ) + p (t) ... formula (6) q (t) = (1-k 3) · α (t) + k 3 · p (t) ... (7) q (t) = α (t) · p (t) (8)

【0062】対話管理装置1では、音声認識結果判定部
32から与えられる認識結果32aに基づいて、次に発
声すべきシステムアナウンスのテキスト1aを決定して
アナウンス発声装置2にコード名等を送る。
In the dialogue management device 1, the text 1a of the system announcement to be uttered next is determined based on the recognition result 32a given from the voice recognition result judging section 32, and the code name or the like is sent to the announcement utterance device 2.

【0063】<第4実施例>図5に示されるように、第
4実施例に係る音声対話装置は、対話管理装置1と、ア
ナウンス発声装置2と、音声認識装置40とを具備した
ものであり、音声出力装置3及び音声入力装置5は必要
に応じて音声対話装置に内蔵されたり、あるいは音声対
話装置とは離れた別物で適宜接続されたりする。音声認
識装置40は予測分布格納部11と、第1の発話開始点
らしさの演算部12と、発話検出用の音響分析部13
と、第2の発話開始点らしさの演算部14と、第3の発
話開始点らしさの演算部15と、音声認識用の音響分析
部18と、音声認識部41と、遷移確率更新部42とを
具備している。
<Fourth Embodiment> As shown in FIG. 5, the voice dialogue apparatus according to the fourth embodiment comprises a dialogue management apparatus 1, an announcement utterance apparatus 2, and a voice recognition apparatus 40. Therefore, the voice output device 3 and the voice input device 5 may be built in the voice dialogue device as necessary, or may be appropriately connected as a separate object apart from the voice dialogue device. The speech recognition device 40 includes a prediction distribution storage unit 11, a first utterance start point likelihood calculation unit 12, and an utterance detection acoustic analysis unit 13.
A second utterance start point likelihood calculating unit 14, a third utterance start point likelihood calculating unit 15, a voice recognition acoustic analysis unit 18, a voice recognizing unit 41, and a transition probability updating unit 42. It is equipped with.

【0064】第4実施例の各装置構成要素のうち、演算
部15と、音声認識部41及び遷移確率更新部42が図
1に示した第1実施例と異なり、また第1実施例におけ
る発話開始時刻決定部16及びスイッチ17が存在しな
いが、他のもの1,2,3,5,11〜14及び18は
第1実施例の同符号のものと同機能であるから説明を簡
単にする。
Among the components of the apparatus of the fourth embodiment, the arithmetic unit 15, the voice recognition unit 41 and the transition probability updating unit 42 are different from those of the first embodiment shown in FIG. 1, and the utterance in the first embodiment is different. Although the start time determination unit 16 and the switch 17 are not present, the other units 1, 2, 3, 5, 11 to 14 and 18 have the same functions as those of the same reference numerals in the first embodiment, and therefore the description will be simplified. .

【0065】演算部15は前述した式(1)〜式(3)
を用いて、第1の発話開始点らしさa(t)により第2
の発話開始点らしさb(t)に対して重み付けを行い、
第3の発話開始点らしさα(t)を時間tに応じて算出
するが、これは遷移確率更新部42に送る。なお第1の
発話開始点らしさa(t)は、予測分布格納部11に格
納されている図2に示したような予測分布に基づいて、
時刻tでユーザの発話が開始されるであろう期待値を演
算部12が算出することにより求まる。また第2の発話
開始点らしさb(t)は、音響分析部13が常時音響分
析して得られる発話検出用の特徴パラメータ13aに基
づいて、時刻tでユーザの発話が開始されたであろう尤
度を演算部14が算出することにより求まる。但し、ア
ナウンス発声装置2からアナウンス開始信号2aが与え
られた時を時刻0としている。
The calculation section 15 uses the equations (1) to (3) described above.
By using the first utterance starting point likelihood a (t)
The utterance starting point like b (t) is weighted,
The third utterance starting point likelihood α (t) is calculated according to the time t, and this is sent to the transition probability updating unit 42. The first utterance starting point likelihood a (t) is based on the prediction distribution stored in the prediction distribution storage unit 11 as shown in FIG.
The calculation unit 12 calculates the expected value at which the user's utterance will start at time t. Further, the second utterance starting point likelihood b (t) would have started the utterance of the user at time t based on the utterance detection characteristic parameter 13a obtained by the acoustic analysis unit 13 performing acoustic analysis at all times. It is obtained by calculating the likelihood by the calculation unit 14. However, the time when the announcement start signal 2a is given from the announcement voicing device 2 is time 0.

【0066】音声認識用の音響分析部18は音声入力回
路5から与えられる音声信号5aを常時音響分析して音
声認識用の特徴パラメータ18aを次々に算出し、音声
認識部41に送る。
The voice recognition acoustic analysis unit 18 constantly performs acoustic analysis on the voice signal 5a supplied from the voice input circuit 5 to successively calculate the voice recognition characteristic parameters 18a, and sends the feature parameters 18a to the voice recognition unit 41.

【0067】音声認識部41では、音響分析部18から
与えられる特徴パラメータ18aの列に対し、常時、図
6に示すような先頭に無音状態300を有する確率付き
有限状態ネットワークを探索して、最大の尤度が得られ
る経路を音声認識結果41aとして出力し、対話管理装
置1に送る。
The voice recognition unit 41 always searches the sequence of the characteristic parameters 18a given from the acoustic analysis unit 18 for a probability-limited finite state network having a silent state 300 at the head as shown in FIG. The route from which the likelihood is obtained is output as the voice recognition result 41a and sent to the dialogue management device 1.

【0068】一般に、確率付き有限状態ネットワークは
音素や単語のHMM(隠れマルコフモデル:Hidden Mar
kov Model)によって構成されるものであり、HMMの各
状態には特徴パラメータに応じた尤度が保持されたり、
あるいは特徴パラメータに応じた尤度を計算するための
確率分布が保持されている。
In general, a finite state network with probability is an HMM (Hidden Markov Model: Hidden Marv) of phonemes and words.
kov Model), and each state of the HMM holds the likelihood according to the feature parameter,
Alternatively, the probability distribution for calculating the likelihood according to the characteristic parameter is held.

【0069】この確率付き有限ネットワークを構成する
場合に、図6に示すように、文頭に無音モデル300を
設けてある。無音モデルは音声のない区間に対応するモ
デルであるが、学習の際、背影雑音や回線雑音を用いる
ことでそれらの雑音に対応することができる。また、咳
や息などの非音声も学習しておくことにより、それらの
非音声を音声と誤認することを防ぐことができる。ま
た、雑音や非音声のモデルを別々に学習し、無音モデル
300と並列に配置することも可能である。これらによ
り、音響分析部18からの音声認識用の特徴パラメータ
18aの入力をユーザの発話開始前から常時受け付ける
ことが可能となる。
When constructing this finite network with probability, a silent model 300 is provided at the beginning of a sentence, as shown in FIG. The silent model is a model corresponding to a section without voice, but it is possible to deal with such noise by using background noise or line noise during learning. In addition, by learning non-voices such as cough and breath, it is possible to prevent the non-voices from being mistaken as voices. It is also possible to separately learn noise and non-voice models and arrange them in parallel with the silence model 300. As a result, it becomes possible to always receive the input of the characteristic parameter 18a for voice recognition from the acoustic analysis unit 18 before the user starts speaking.

【0070】遷移確率更新部42は音声認識部41で用
いられる確率付き有限状態ネットワークの先頭の無音モ
デル300から文先頭状態303へ遷移する確率を、演
算部15から与えられる第3の発話開始点らしさα
(t)を用いて、時刻tに応じて変化させる。即ち、図
6に示すように、先頭の無音モデル300には自己状態
への遷移301と、文先頭状態への遷移302とがあ
り、それぞれのアーク(弧)には状態遷移確率が付えら
れているから、第3の発話開始点らしさα(t)が大き
い時刻tでは文先頭状態303へ遷移する状態遷移確率
をα(t)に応じて大きくする。これにより、ユーザの
発話開始時刻に先頭の無音モデル300から文先頭状態
303へ遷移し易くなり、音声認識の精度が向上する。
この場合、α(t)が大きい時刻tでは同時に、自己状
態300に遷移する状態遷移確率をα(t)に応じて小
さくすると良い。
The transition probability updating unit 42 gives the probability of transition from the head silence model 300 of the finite state network with probabilities used in the speech recognition unit 41 to the sentence head state 303 to the third utterance starting point given from the arithmetic unit 15. Likeness α
Using (t), it is changed according to time t. That is, as shown in FIG. 6, the beginning silence model 300 has a transition 301 to a self-state and a transition 302 to a sentence start state, and each arc has an associated state transition probability. Therefore, at time t at which the third utterance start point likelihood α (t) is large, the state transition probability of transition to the sentence head state 303 is increased according to α (t). This facilitates the transition from the head silent model 300 to the sentence head state 303 at the user's utterance start time, and improves the accuracy of voice recognition.
In this case, at the time t when α (t) is large, the state transition probability of transitioning to the self-state 300 may be decreased at the same time according to α (t).

【0071】逆に、第3の発話開始点らしさα(t)が
小さい時刻tでは文先頭状態303へ遷移する状態遷移
確率をα(t)に応じて小さくする。これにより、ユー
ザの発話開始時刻前では先頭の無音モデル300から文
先頭状態303へは遷移し難くなり、誤った音声認識を
行い難くなるから、音声認識の精度が向上する。この場
合、α(t)が小さい時刻tでは同時に、自己状態30
0に遷移する状態遷移確率をα(t)に応じて大きくす
ると良い。このように、先頭の無音状態300から文先
頭状態303への状態遷移確率を第3の発話開始点らし
さα(t)で変化させることにより、ユーザの発話開始
時刻を高精度に一点決定することができなくても、結果
的にユーザの発話を高精度に音声認識することができ
る。また、音声認識は実質的に1回であるから、第3実
施例に比べて、処理が高速化し、メモリ容量も削減する
ことができる。
On the other hand, at time t when the third utterance start point likelihood α (t) is small, the state transition probability of transition to the sentence head state 303 is reduced according to α (t). As a result, before the user's utterance start time, it becomes difficult to transition from the head silent model 300 to the sentence head state 303, and it becomes difficult to perform erroneous voice recognition, so that the accuracy of voice recognition is improved. In this case, at time t when α (t) is small, the self-state 30
The state transition probability of transition to 0 may be increased according to α (t). Thus, by changing the state transition probability from the silent state 300 at the head to the sentence head state 303 by the third utterance start point likelihood α (t), a user's utterance start time is determined with high accuracy. Even if it is not possible, the user's utterance can be recognized with high accuracy as a result. Further, since the voice recognition is substantially performed once, the processing speed is increased and the memory capacity can be reduced as compared with the third embodiment.

【0072】なお、音声認識用の特徴パラメータ18a
は発話開始の検出には最適ではないため、無音状態30
0から文先頭状態303への状態遷移確率を固定してお
くと、先頭の無音状態300から文先頭状態303への
遷移の精度が低くなり、音声認識の精度が低下する。
The feature parameter 18a for voice recognition is used.
Is not optimal for detecting the start of speech, so the silent state 30
If the state transition probability from 0 to the sentence head state 303 is fixed, the accuracy of the transition from the head silent state 300 to the sentence head state 303 becomes low, and the accuracy of speech recognition deteriorates.

【0073】対話管理装置1では、音声認識部41から
与えられる音声認識結果41aに基づいて、次に発声す
べきシステムアナウンスのテキスト1aを決定してアナ
ウンス発声装置2にコード名等を送る。
The dialogue management device 1 determines the text 1a of the system announcement to be uttered next based on the voice recognition result 41a given from the voice recognition unit 41, and sends the code name and the like to the announcement utterance device 2.

【0074】[0074]

【発明の効果】第1,第2及び第7〜第11の発明によ
れば、システムアナウンスとユーザの発話開始時刻との
因果関係に着目して、予め用意した予測分布からユーザ
の発話が開始されるであろう期待値(第1の発話開始点
らしさ)を算出し、発話検出用の特徴パラメータから求
めたユーザの発話が開始されたであろう尤度(第2の発
話開始点らしさ)と併用してユーザの発話開始時刻を決
定するので、発話開始時刻を一点高精度に検出すること
ができ、従って高精度な音声認識を実現することができ
る。
According to the first, second and seventh to eleventh aspects of the invention, the utterance of the user is started from the prediction distribution prepared in advance, paying attention to the causal relationship between the system announcement and the utterance start time of the user. The expected value (likely as the first utterance starting point) that is likely to be calculated, and the likelihood that the utterance of the user obtained from the characteristic parameter for utterance detection is likely to start (likely as the second utterance starting point) Since the user's utterance start time is determined in combination with this, the utterance start time can be detected with one point and high accuracy, and therefore highly accurate voice recognition can be realized.

【0075】また第3〜第4及び第12〜第13の発明
によれば、音声認識を常に行うことにより、音声認識結
果の尤度が発話開始点を決定するのにも用いられること
になり、例えば無意味な発声や咳を発話開始点と決定し
てしまう等の誤りを回避することができ、結果的に高精
度な音声認識を行うことができる。
According to the third to fourth and twelfth to thirteenth inventions, the likelihood of the voice recognition result is also used to determine the utterance starting point by always performing the voice recognition. For example, it is possible to avoid an error such as meaningless utterance or determining a cough as the utterance starting point, and as a result, highly accurate voice recognition can be performed.

【0076】特に第5及び第14の発明によればシステ
ムアナウンスの無音区間に予測分布の極大点があり、更
に第6及び第15の発明によれば無音区間を故意あるい
は積極的に設けることにより、システムアナウンスとユ
ーザの発話開始時刻との因果関係が一層明確化し、発話
開始時刻の検出精度及び音声認識精度が更に向上する。
また、システムアナウンス中に無音区間を故意あるいは
積極的に設けることにより、無音区間でユーザが発話を
開始するようにユーザを制御することができるから、音
声対話装置の利用時間の短縮が可能となる。つまり、対
話における音声認識結果確認時に例えば「山本で良けれ
ばはい、さもなければいいえとお答え下さい」とシステ
ムアナウンスをする場合に比べ、「山本でよろしいでし
ょうか(1秒無音)はい、またはいいえでお答え下さ
い」とアナウンスすることにより、装置に慣れたユーザ
は無音区間に発話するようになり、「はい」以降のシス
テムアナウンスは無用となるから、システムアナウンス
を聞く時間は半分以下に短縮され、ユーザにとっての利
便性を高めると共に装置の効率的な利用が可能となる。
また、必要に応じて、発話開始時刻が決定されたならば
システムアナウンスを停止し、ユーザの発声を妨げない
ようにすることも可能となる。また、無音区間の設定に
より、初心者には十分なシステムアナウンスを聞かせ、
熟練者には短いシステムアナウンスを聞くだけで利用で
きる音声対話装置が実現する。更に、発話開始時刻を高
精度に決定できる場合には、このような利用時間の短縮
が可能な装置が一層有効に働くことができる。
Particularly, according to the fifth and fourteenth inventions, there is a maximum point of the predicted distribution in the silent section of the system announcement, and according to the sixth and fifteenth inventions, the silent section is intentionally or positively provided. , The causal relationship between the system announcement and the utterance start time of the user is further clarified, and the utterance start time detection accuracy and the voice recognition accuracy are further improved.
Further, by intentionally or positively providing the silent section during the system announcement, the user can be controlled so that the user starts the utterance in the silent section, so that the use time of the voice interaction device can be shortened. . In other words, when confirming the voice recognition result in the dialogue, for example, "Is it OK in Yamamoto, otherwise answer No?" Compared to when making a system announcement, "Is it okay in Yamamoto (1 second silence) Yes or No? By answering `` Please answer '', the user who is accustomed to the device will speak in the silent section, and the system announcement after "Yes" will be unnecessary, so the time to listen to the system announcement will be reduced to less than half, It is possible to enhance the convenience for the user and efficiently use the device.
In addition, if necessary, it is possible to stop the system announcement when the utterance start time is determined so as not to disturb the user's utterance. Also, by setting the silent section, let the beginner hear a sufficient system announcement,
A spoken dialogue system that can be used by experienced users will be realized only by listening to a short system announcement. Furthermore, when the utterance start time can be determined with high accuracy, such a device capable of shortening the use time can work more effectively.

【0077】第16の発明によれば高精度な音声認識の
下でユーザと装置間で音声を用いた対話を行うので、ス
ムーズな対話が実現する。
According to the sixteenth invention, since a dialogue using voice is performed between the user and the device under high-accuracy voice recognition, smooth dialogue can be realized.

【図面の簡単な説明】[Brief description of drawings]

【図1】第1実施例に係る音声対話装置のブロック構成
図。
FIG. 1 is a block configuration diagram of a voice interaction device according to a first embodiment.

【図2】予測分布の一例を示す図。FIG. 2 is a diagram showing an example of a predicted distribution.

【図3】第2実施例に係る音声対話装置のブロック構成
図。
FIG. 3 is a block configuration diagram of a voice dialog device according to a second embodiment.

【図4】第3実施例に係る音声対話装置のブロック構成
図。
FIG. 4 is a block configuration diagram of a voice dialog device according to a third embodiment.

【図5】第4実施例に係る音声対話装置のブロック構成
図。
FIG. 5 is a block configuration diagram of a voice dialog device according to a fourth embodiment.

【図6】先頭に無音状態を有する確率付有限状態ネット
ワークの一例を示す図。
FIG. 6 is a diagram showing an example of a finite state network with probability having a silent state at the head.

【図7】従来例を示す図。FIG. 7 is a diagram showing a conventional example.

【符号の説明】 1 対話管理装置 1a テキスト 2 アナウンス発声装置 2a,5a 音声信号 2b アナウンス開始信号 2c アナウンス終了信号 3 音声出力回路 3a システムアナウンス 4 ユーザの発話 5 音声入力回路 10,20,30,40 音声認識装置 11 予測分布格納部 12 第1の発話開始点らしさの演算部 13 発話検出用の音響分析部 13a 発話検出用の特徴パラメータ 14 第2の発話開始点らしさの演算部 15 第3の発話開始点らしさの演算部 16,22 発話開始時刻決定部 17 音声認識用の音声信号通過スイッチ 18 音声認識用の音響分析部 18a 音声認識用の特徴パラメータ 19,31,41 音声認識部 19a,32a,41a 認識結果 21 基準値演算部 32 音声認識結果判定部 42 遷移確率更新部 100 予測分布 200 無音区間 300 無音状態 303 文先頭状態 a(t) 第1の発話開始点らしさ b(t) 第2の発話開始点らしさ α(t) 第3の発話開始点らしさ p(t) 音声認識結果の尤度 Ref 基準値 Refo 第1の基準値 Ref(t) 第2の基準値[Explanation of Codes] 1 Dialog management device 1a Text 2 Announcement speech device 2a, 5a Voice signal 2b Announcement start signal 2c Announcement end signal 3 Voice output circuit 3a System announcement 4 User's speech 5 Voice input circuit 10, 20, 30, 40 Speech recognition device 11 Prediction distribution storage unit 12 First utterance start point likelihood calculation unit 13 Speech detection acoustic analysis unit 13a Speech detection characteristic parameter 14 Second utterance start point likelihood calculation unit 15 Third utterance Start point likelihood calculation unit 16, 22 Speech start time determination unit 17 Voice signal passage switch for voice recognition 18 Acoustic analysis unit for voice recognition 18a Characteristic parameter for voice recognition 19, 31, 41 Voice recognition unit 19a, 32a, 41a Recognition result 21 Reference value calculation unit 32 Speech recognition result determination unit 42 Transition probability updating unit 100 Predicted distribution 200 Silence interval 300 Silence state 303 Sentence state a (t) First utterance start point likelihood b (t) Second utterance start point likelihood α (t) Third utterance start point likelihood p (t) speech recognition result likelihood R ef reference value R EFO first reference value R ef (t) second reference value

───────────────────────────────────────────────────── フロントページの続き (72)発明者 山本 誠一 東京都新宿区西新宿二丁目3番2号 国際 電信電話株式会社内 ─────────────────────────────────────────────────── ─── Continuation of the front page (72) Inventor Seiichi Yamamoto 2-3-2 Nishishinjuku, Shinjuku-ku, Tokyo International Telegraph and Telephone Corporation

Claims (16)

【特許請求の範囲】[Claims] 【請求項1】 音声を用いてユーザとの対話を行う音声
対話装置に適用される音声認識方法において:前記音声
対話装置のシステムアナウンスに対するユーザの発話開
始時刻の極大点を有する予測分布を予め用意しておき、
この予測分布に基づき、ユーザの発話が開始される期待
値を第1の発話開始点らしさとしてシステムアナウンス
開始後の時刻に応じて算出すること;電気信号に変換さ
れたユーザの発話を音響分析して発話検出用の特徴パラ
メータを算出し、この特徴パラメータに基づき、ユーザ
の発話が開始されたであろう尤度を第2の発話開始点ら
しさとして時刻に応じて算出すること;第2の発話開始
点らしさに対して第1の発話開始点らしさにより重み付
けを行い、この重み付けで得た第3の発話開始点らしさ
を基準値と比較し、基準値より大きくなった時点をユー
ザの発話開始時刻であると決定すること;電気信号に変
換されたユーザの発話を音響分析して音声認識用の特徴
パラメータを算出し、この特徴パラメータに基づき音声
認識を行う処理を、前記ユーザの発話開始時刻の決定に
従って行うこと;を特徴とする音声認識方法。
1. A voice recognition method applied to a voice interaction device for interacting with a user using voice, wherein: a prediction distribution having a maximum point of a user's utterance start time with respect to a system announcement of the voice interaction device is prepared in advance. Well,
Based on this predicted distribution, calculate the expected value at which the user's utterance starts as the first utterance start point according to the time after the system announcement starts; perform acoustic analysis of the user's utterance converted into an electrical signal. Calculating a feature parameter for utterance detection based on the feature parameter, and calculating the likelihood that the user's utterance has started as the second utterance starting point likelihood according to the time based on the feature parameter; The likelihood of the start point is weighted by the likelihood of the first utterance start point, the likelihood of the third utterance start point obtained by this weighting is compared with a reference value, and the time point at which the likelihood becomes larger than the reference value is the utterance start time of the user. A process of performing a voice recognition based on this characteristic parameter by acoustically analyzing the user's utterance converted into an electric signal to calculate a characteristic parameter for voice recognition. Speech recognition method according to claim; be carried out according to the decision of the speech start time of the user.
【請求項2】 音声を用いてユーザとの対話を行う音声
対話装置に適用される音声認識方法において:前記音声
対話装置のシステムアナウンスに対するユーザの発話開
始時刻の極大点を有する予測分布を予め用意しておき、
この予測分布に基づき、ユーザの発話が開始される期待
値を第1の発話開始点らしさとしてシステムアナウンス
開始後の時刻に応じて算出すること;電気信号に変換さ
れたユーザの発話を音響分析して発話検出用の特徴パラ
メータを算出し、この特徴パラメータに基づき、ユーザ
の発話が開始されたであろう尤度を第2の発話開始点ら
しさとして時刻に応じて算出すること;第1の発話開始
点らしさにより第1の基準値を重み付けして時間に応じ
て変化する第2の基準値を算出し、第2の発話開始点ら
しさをこの第2の基準値と比較し、第2の基準値より大
きくなった時点をユーザの発話開始時刻であると決定す
ること;電気信号に変換されたユーザの発話を音響分析
して音声認識用の特徴パラメータを算出し、この特徴パ
ラメータに基づき音声認識を行う処理を、前記ユーザの
発話開始時刻の決定に従って行うこと;を特徴とする音
声認識方法。
2. A voice recognition method applied to a voice interaction device for interacting with a user by using a voice: A prediction distribution having a maximum point of a user's utterance start time for a system announcement of the voice interaction device is prepared in advance. Well,
Based on this predicted distribution, calculate the expected value at which the user's utterance starts as the first utterance start point according to the time after the system announcement starts; perform acoustic analysis of the user's utterance converted into an electrical signal. Calculating a feature parameter for utterance detection based on the feature parameter, and calculating a likelihood that the user's utterance has started as a second utterance starting point likelihood according to the time; the first utterance. The first reference value is weighted according to the starting point likelihood to calculate a second reference value that changes with time, and the second utterance starting point likelihood is compared with the second reference value to obtain the second reference value. The time when the value becomes larger than the value is determined to be the user's utterance start time; the user's utterance converted into an electric signal is acoustically analyzed to calculate a characteristic parameter for voice recognition, and based on this characteristic parameter Speech recognition method comprising; a process for performing voice recognition, be carried out as determined by the speech start time of the user.
【請求項3】 音声を用いてユーザとの対話を行う音声
対話装置に適用される音声認識方法において:前記音声
対話装置のシステムアナウンスに対するユーザの発話開
始時刻の極大点を有する予測分布を予め用意しておき、
この予測分布に基づき、ユーザの発話が開始される期待
値を第1の発話開始点らしさとしてシステムアナウンス
開始後の時刻に応じて算出すること;電気信号に変換さ
れたユーザの発話を音響分析して発話検出用の特徴パラ
メータを算出し、この特徴パラメータに基づき、ユーザ
の発話が開始されたであろう尤度を第2の発話開始点ら
しさとして時刻に応じて算出すること;第2の発話開始
点らしさに対して第1の発話開始点らしさにより重み付
けを行い、第3の発話開始点らしさを算出すること;電
気信号に変換されたユーザの発話を音響分析して音声認
識用の特徴パラメータを算出し、この特徴パラメータに
基づき、認識開始時刻を次々にずらして音声認識を行
い、且つ、各認識開始時刻に対応した音声認識結果毎の
尤度を算出すること;各認識開始時刻毎の音声認識結果
の尤度と、前記重み付けで得た第3の発話開始点らしさ
との和または積を時刻を合わせて算出し、この算出した
値が最大となる認識開始時刻に対応した音声認識結果
を、ユーザの発話に対する音声認識結果と判定するこ
と;を特徴とする音声認識方法。
3. A voice recognition method applied to a voice interaction device for interacting with a user by using a voice: A prediction distribution having a maximum point of a user's utterance start time with respect to a system announcement of the voice interaction device is prepared in advance. Well,
Based on this predicted distribution, calculate the expected value at which the user's utterance starts as the first utterance start point according to the time after the system announcement starts; perform acoustic analysis of the user's utterance converted into an electrical signal. Calculating a feature parameter for utterance detection based on the feature parameter, and calculating the likelihood that the user's utterance has started as the second utterance starting point likelihood according to the time based on the feature parameter; The likelihood of the start point is weighted by the likelihood of the first utterance start point to calculate the likelihood of the third utterance start point; a characteristic parameter for voice recognition by acoustically analyzing the utterance of the user converted into an electric signal. Based on this characteristic parameter, the recognition start time is shifted one after another to perform voice recognition, and the likelihood for each voice recognition result corresponding to each recognition start time is calculated. The sum or product of the likelihood of the speech recognition result for each recognition start time and the likelihood of the third utterance start point obtained by the weighting is calculated by combining the times, and the calculated start time is the maximum. A voice recognition result corresponding to the above is determined as a voice recognition result for a user's utterance;
【請求項4】 音声を用いてユーザとの対話を行う音声
対話装置に適用される音声認識方法において:前記音声
対話装置のシステムアナウンスに対するユーザの発話開
始時刻の極大点を有する予測分布を予め用意しておき、
この予測分布に基づき、ユーザの発話が開始される期待
値を第1の発話開始点らしさとしてシステムアナウンス
開始後の時刻に応じて算出すること;電気信号に変換さ
れたユーザの発話を音響分析して発話検出用の特徴パラ
メータを算出し、この特徴パラメータに基づき、ユーザ
の発話が開始されたであろう尤度を第2の発話開始点ら
しさとして時刻に応じて算出すること;第2の発話開始
点らしさに対して第1の発話開始点らしさにより重み付
けを行い、第3の発話開始点らしさを算出すること;電
気信号に変換されたユーザの発話を音響分析して音声認
識用の特徴パラメータを算出し、この特徴パラメータに
基づき、先頭に無音状態を有する確率付き有限状態ネッ
トワークを探索して音声認識を行うこと;前記確率付き
有限状態ネットワークの先頭の無音状態から文の先頭状
態へ遷移する確率を、前記重み付けで得た第3の発話開
始点らしさを用いて時刻に応じて更新すること;を特徴
とする音声認識方法。
4. A voice recognition method applied to a voice interaction device for interacting with a user by using a voice: A prediction distribution having a maximum point of a user's utterance start time for a system announcement of the voice interaction device is prepared in advance. Well,
Based on this predicted distribution, calculate the expected value at which the user's utterance starts as the first utterance start point according to the time after the system announcement starts; perform acoustic analysis of the user's utterance converted into an electrical signal. Calculating a feature parameter for utterance detection based on the feature parameter, and calculating the likelihood that the user's utterance has started as the second utterance starting point likelihood according to the time based on the feature parameter; The likelihood of the start point is weighted by the likelihood of the first utterance start point to calculate the likelihood of the third utterance start point; a characteristic parameter for voice recognition by acoustically analyzing the utterance of the user converted into an electric signal. And probing a finite-state network with a probability having a silent state at the head for speech recognition based on this characteristic parameter; Speech recognition method comprising; the probability of transition from the beginning of the silent state of the click to the top state of the sentence, be updated according to the time by using the third utterance start point likelihood of that obtained in the weighting.
【請求項5】 請求項1ないし4いずれか一つに記載の
音声認識方法において、前記ユーザの発話開始時刻の予
測分布の極大点がシステムアナウンスの無音区間に存在
することを特徴とする音声認識方法。
5. The voice recognition method according to claim 1, wherein the maximum point of the predicted distribution of the utterance start times of the users exists in a silent section of the system announcement. Method.
【請求項6】 請求項5記載の音声認識方法において、
前記無音区間はその長さが0.2秒以上3秒以下であ
り、システムアナウンスの文と文の間及び文節と文節と
の間のうち少なくとも一方に存在することを特徴とする
音声認識方法。
6. The voice recognition method according to claim 5, wherein
The voice recognition method, wherein the silent section has a length of 0.2 seconds or more and 3 seconds or less and is present in at least one of a sentence of the system announcement and a sentence or a sentence.
【請求項7】 音声を用いてユーザとの対話を行う音声
対話装置に適用される音声認識装置において;前記音声
対話装置のシステムアナウンスに対するユーザの発話開
始時刻の極大点を有する予測分布を格納する第1手段
と;前記格納された予測分布に基づき、ユーザの発話が
開始される期待値を第1の発話開始点らしさとしてシス
テムアナウンス開始後の時刻に応じて算出する第2手段
と;電気信号に変換されたユーザの発話を音響分析し、
発話検出用の特徴パラメータを算出する第3手段と;前
記発話検出用の特徴パラメータに基づき、ユーザの発話
が開始されたであろう尤度を第2の発話開始点らしさと
して時刻に応じて算出する第4手段と;第2の発話開始
点らしさに対して第1の発話開始点らしさにより重み付
けを行い、この重み付けされた値を第3の発話開始点ら
しさとして時刻に応じて算出する第5手段と;第3の発
話開始点らしさを基準値と比較し、この基準値より大き
くなった時点をユーザの発話開始時刻であると決定する
第6手段と;電気信号に変換されたユーザの発話を音響
分析して音声認識用の特徴パラメータを算出し、この特
徴パラメータに基づき音声認識を行う処理を、前記ユー
ザの発話開始時刻の決定に従って行う第7手段と;を具
備することを特徴とする音声認識装置。
7. A voice recognition device applied to a voice interaction device for interacting with a user using voice; storing a prediction distribution having a maximum point of a user's utterance start time with respect to a system announcement of the voice interaction device. First means; second means for calculating an expected value at which a user's utterance starts based on the stored predicted distribution as a first utterance start point in accordance with a time after the system announcement is started; and an electric signal Acoustic analysis of the user's utterance converted to
Third means for calculating a characteristic parameter for utterance detection; calculating a likelihood that a user's utterance has started based on the characteristic parameter for utterance detection as a second utterance start point according to time Fifth means for calculating the second utterance starting point likelihood by weighting the first utterance starting point likelihood, and calculating the weighted value as the third utterance starting point likelihood according to time. Means: A third utterance start point likelihood is compared with a reference value, and a sixth means for determining a time point when the utterance start point is greater than the reference value to be the utterance start time of the user; Acoustic analysis is performed to calculate a feature parameter for voice recognition, and voice recognition is performed based on the feature parameter according to a determination of the utterance start time of the user. Voice recognition device that.
【請求項8】 音声を用いてユーザとの対話を行う音声
対話装置に適用される音声認識装置において;前記音声
対話装置のシステムアナウンスに対するユーザの発話開
始時刻の極大点を有する予測分布を格納する第1手段
と;前記格納された予測分布に基づき、ユーザの発話が
開始される期待値を第1の発話開始点らしさとしてシス
テムアナウンス開始後の時刻に応じて算出する第2手段
と;電気信号に変換されたユーザの発話を音響分析し、
発話検出用の特徴パラメータを算出する第3手段と;前
記発話検出用の特徴パラメータに基づき、ユーザの発話
が開始されたであろう尤度を第2の発話開始点らしさと
して時刻に応じて算出する第4手段と;第1の基準値に
対して第1の発話開始点らしさにより重み付けを行い、
この重み付けで得た値を第2の基準値として時刻に応じ
て算出する第5手段と;第2の発話開始点らしさを前記
重み付けで得た第2の基準値と比較し、この第2の基準
値より大きくなった時点をユーザの発話開始時刻である
と決定する第6手段と;電気信号に変換されたユーザの
発話を音響分析して音声認識用の特徴パラメータを算出
し、この特徴パラメータに基づき音声認識を行う処理
を、前記ユーザの発話開始時刻の決定に従って行う第7
手段と;を具備することを特徴とする音声認識装置。
8. A voice recognition device applied to a voice interaction device for interacting with a user by using a voice; storing a prediction distribution having a maximum point of a user's utterance start time with respect to a system announcement of the voice interaction device. First means; second means for calculating an expected value at which a user's utterance starts based on the stored predicted distribution as a first utterance start point in accordance with a time after the system announcement is started; and an electric signal Acoustic analysis of the user's utterance converted to
Third means for calculating a characteristic parameter for utterance detection; calculating a likelihood that a user's utterance has started based on the characteristic parameter for utterance detection as a second utterance start point according to time Fourth means for performing; weighting according to the likelihood of the first utterance starting point with respect to the first reference value,
Fifth means for calculating the value obtained by this weighting as the second reference value according to the time; and comparing the second utterance start point likelihood with the second reference value obtained by the weighting, and Sixth means for determining a time point when it becomes larger than a reference value as a user's utterance start time; acoustic analysis of the user's utterance converted into an electric signal to calculate a characteristic parameter for voice recognition, and the characteristic parameter Seventh, a process of performing voice recognition based on is determined according to the determination of the utterance start time of the user.
A speech recognition apparatus comprising: means.
【請求項9】 請求項7または8に記載の音声認識装置
において:第7手段は、 電気信号に変換されたユーザの発話を、ユーザの発話開
始時刻であると決定された時点から通過させるスイッチ
手段と;このスイッチ手段を通過したユーザの発話を音
響分析して音声認識用の特徴パラメータを算出する音声
認識用の音響分析手段と;この音響分析手段により算出
された音声認識用の特徴パラメータに基づいて音声認識
を行う音声認識手段と;を具備することを特徴とする音
声認識装置。
9. The voice recognition device according to claim 7 or 8, wherein the seventh means is a switch for passing the user's utterance converted into an electric signal from a time point determined to be the user's utterance start time. Means; acoustic analysis means for voice recognition for acoustically analyzing the utterance of the user passing through the switch means to calculate a characteristic parameter for voice recognition; and a characteristic parameter for voice recognition calculated by the acoustic analysis means. A voice recognition means for performing voice recognition based on the voice recognition device;
【請求項10】 請求項7または8に記載の音声認識装
置において:第7手段は、 ユーザの発話開始時刻であると決定された時点から、電
気信号に変換されたユーザの発話の音響分析を開始して
音声認識用の特徴パラメータを算出する音声認識用の音
響分析手段と;この音響分析手段により算出された音声
認識用の特徴パラメータに基づいて音声認識を行う音声
認識手段と;を具備することを特徴とする音声認識装
置。
10. The voice recognition apparatus according to claim 7 or 8, wherein the seventh means performs an acoustic analysis of the user's utterance converted into an electric signal from the time when it is determined to be the user's utterance start time. A voice recognition acoustic analysis unit that starts to calculate a voice recognition characteristic parameter; and a voice recognition unit that performs voice recognition based on the voice recognition characteristic parameter calculated by the acoustic analysis unit; A voice recognition device characterized by the above.
【請求項11】 請求項7または8に記載の音声認識装
置において:第7手段は、 電気信号に変換されたユーザの発話を音響分析して音声
認識用の特徴パラメータを算出する音声認識用の音響分
析手段と;この音響分析手段で算出された音声認識用の
特徴パラメータのうち、ユーザの発話開始時刻であると
決定された時点以降の特徴パラメータに基づいて音声認
識を行う音声認識手段と;を具備することを特徴とする
音声認識装置。
11. The voice recognition apparatus according to claim 7 or 8, wherein the seventh means is for voice recognition for acoustically analyzing a user's utterance converted into an electric signal to calculate a feature parameter for voice recognition. Acoustic analysis means; speech recognition means for performing speech recognition based on the characteristic parameters after the time point when it is determined that the speech start time of the user is among the characteristic parameters for speech recognition calculated by the acoustic analysis means; A voice recognition device comprising:
【請求項12】 音声を用いてユーザとの対話を行う音
声対話装置に適用される音声認識装置において;前記音
声対話装置のシステムアナウンスに対するユーザの発話
開始時刻の極大点を有する予測分布を格納する予測分布
格納手段と;前記格納された予測分布に基づき、ユーザ
の発話が開始される期待値を第1の発話開始点らしさと
してシステムアナウンス開始後の時刻に応じて算出する
第1の演算手段と;電気信号に変換されたユーザの発話
を音響分析し、発話検出用の特徴パラメータを算出する
発話検出用の音声分析手段と;前記発話検出用の特徴パ
ラメータに基づき、ユーザの発話が開始されたであろう
尤度を第2の発話開始点らしさとして時刻に応じて算出
する第2の演算手段と;第2の発話開始点らしさに対し
て第1の発話開始点らしさにより重み付けを行い、この
重み付けされた値を第3の発話開始点らしさとして時刻
に応じて算出する第3の演算手段と;前記電気信号に変
換されたユーザの発話を音響分析し、音声認識用の特徴
パラメータを算出する音声認識用の音響分析手段と;前
記音声認識用の特徴パラメータに基づき、認識開始時刻
を次々にずらして音声認識を行い、且つ、各認識開始時
刻に対応した音声認識結果毎の尤度を算出する音声認識
手段と;各認識開始時刻毎の音声認識結果の尤度と、第
3の発話開始点らしさとの和または積を時刻に合せて算
出し、この算出した値が最大となる認識開始時刻に対応
した音声認識結果を、ユーザの発話に対する音声認識結
果と判定する音声認識結果判定手段と;を具備すること
を特徴とする音声認識装置。
12. A voice recognition device applied to a voice interaction device for interacting with a user using voice; storing a prediction distribution having a maximum point of a user's utterance start time with respect to a system announcement of the voice interaction device. Predictive distribution storing means; first calculating means for calculating an expected value at which the user's utterance starts based on the stored predictive distribution as a first utterance start point according to the time after the system announcement is started. An utterance detecting voice analysis means for acoustically analyzing the user's utterance converted into an electric signal and calculating a utterance detecting characteristic parameter; and an utterance of the user started based on the utterance detecting characteristic parameter. Second computing means for calculating the likelihood as the second utterance starting point likelihood according to the time; the first utterance starting point with respect to the second utterance starting point likelihood Third computing means for performing weighting according to the likelihood, and calculating the weighted value as the third speech start point likelihood according to the time; acoustic analysis of the user's speech converted into the electric signal, and voice recognition Acoustic analysis means for voice recognition for calculating feature parameters for voice recognition; voice recognition is performed by sequentially shifting recognition start times based on the voice recognition feature parameters, and voice recognition corresponding to each recognition start time. A voice recognition means for calculating the likelihood for each result; a sum or product of the likelihood of the voice recognition result for each recognition start time and the likelihood of the third utterance start point is calculated according to the time, and this calculation is performed. A voice recognition device, comprising: a voice recognition result determining means for determining a voice recognition result corresponding to a recognition start time having a maximum value as a voice recognition result for a user's utterance.
【請求項13】 音声を用いてユーザとの対話を行う音
声対話装置に適用される音声認識装置において;前記音
声対話装置のシステムアナウンスに対するユーザの発話
開始時刻の極大点を有する予測分布を格納する予測分布
格納手段と;前記格納された予測分布に基づき、ユーザ
の発話が開始される期待値を第1の発話開始点らしさと
してシステムアナウンス開始後の時刻に応じて算出する
第1の演算手段と;電気信号に変換されたユーザの発話
を音響分析し、発話検出用の特徴パラメータを算出する
発話検出用の音響分析手段と;前記発話検出用の特徴パ
ラメータに基づき、ユーザの発話が開始されたであろう
尤度を第2の発話開始点らしさとして時刻に応じて算出
する第2の演算手段と;第2の発話開始点らしさに対し
て第1の発話開始点らしさにより重み付けを行い、この
重み付けされた値を第3の発話開始点らしさとして時刻
に応じて算出する第3の演算手段と;前記電気信号に変
換されたユーザの発話を音響分析し、音声認識用の特徴
パラメータを算出する音声認識用の音響分析手段と;前
記音声認識用の特徴パラメータに基づき、先頭に無音状
態を有する確率付き有限状態ネットワークを探索して音
声認識を行う音声認識手段と;前記確率付き有限状態ネ
ットワークの先頭の無音状態から文の先頭状態へ遷移す
る確率を、第3の発話開始点らしさを用いて時刻に応じ
て更新する遷移確率更新手段と;を具備することを特徴
とする音声認識装置。
13. A voice recognition device applied to a voice interaction device for interacting with a user using voice; storing a prediction distribution having a maximum point of a user's utterance start time with respect to a system announcement of the voice interaction device. Predictive distribution storing means; first calculating means for calculating an expected value at which the user's utterance starts based on the stored predictive distribution as a first utterance start point according to the time after the system announcement is started. Utterance detection acoustic analysis means for acoustically analyzing the utterance of the user converted into an electrical signal and calculating a utterance detection characteristic parameter; utterance of the user started based on the utterance detection characteristic parameter Second computing means for calculating the likelihood as the second utterance starting point likelihood according to the time; the first utterance starting point with respect to the second utterance starting point likelihood Third computing means for performing weighting according to the likelihood, and calculating the weighted value as the third speech start point likelihood according to the time; acoustic analysis of the user's speech converted into the electric signal, and voice recognition Acoustic analysis means for voice recognition for calculating a feature parameter for voice recognition; voice recognition means for performing voice recognition by searching a finite state network with probability having a silent state at the head based on the feature parameter for voice recognition; Transition probability updating means for updating the probability of transition from the silent state at the head of the probability-limited finite state network to the head state of the sentence according to the time using the third utterance start point likelihood. And a voice recognition device.
【請求項14】 請求項7ないし13いずれか一つに記
載の音声認識装置において、前記ユーザの発話開始時刻
の予測分布の極大点がシステムアナウンスの無音区間に
存在することを特徴とする音声認識装置。
14. The voice recognition apparatus according to claim 7, wherein the maximum point of the predicted distribution of the utterance start time of the user exists in the silent section of the system announcement. apparatus.
【請求項15】 請求項14記載の音声認識装置におい
て、前記無音区間はその長さが0.2秒以上3秒以下で
あり、システムアナウンスの文と文の間及び文節と文節
との間のうち少なくとも一方に存在することを特徴とす
る音声認識装置。
15. The voice recognition apparatus according to claim 14, wherein the silent section has a length of 0.2 seconds or more and 3 seconds or less, and is between the sentences of the system announcement and between the phrases and the phrases. A voice recognition device characterized by being present in at least one of them.
【請求項16】 請求項7ないし15いずれか一つに記
載の音声認識装置と、システムアナウンスの指定された
テキストを電気的音声信号に変換すると共にシステムア
ナウンスの開始を前記音声認識装置に通知するアナウン
ス発声装置と、このアナウンス発声装置に対するシステ
ムアナウンスのテキストの指定及び前記音声認識装置か
らの音声認識結果の入力により音声を用いたユーザとの
対話を管理する対話管理装置とを具備することを特徴と
する音声対話装置。
16. The voice recognition device according to claim 7, wherein the designated text of the system announcement is converted into an electrical voice signal, and the start of the system announcement is notified to the voice recognition device. An announcement utterance device, and a dialogue management device for managing a dialogue with a user using voice by designating a text of a system announcement for the announcement utterance device and inputting a voice recognition result from the voice recognition device. Spoken dialogue device.
JP13400594A 1994-06-16 1994-06-16 Speech recognition method and apparatus for spoken dialogue Expired - Fee Related JP3285704B2 (en)

Priority Applications (1)

Application Number Priority Date Filing Date Title
JP13400594A JP3285704B2 (en) 1994-06-16 1994-06-16 Speech recognition method and apparatus for spoken dialogue

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
JP13400594A JP3285704B2 (en) 1994-06-16 1994-06-16 Speech recognition method and apparatus for spoken dialogue

Publications (2)

Publication Number Publication Date
JPH086590A true JPH086590A (en) 1996-01-12
JP3285704B2 JP3285704B2 (en) 2002-05-27

Family

ID=15118158

Family Applications (1)

Application Number Title Priority Date Filing Date
JP13400594A Expired - Fee Related JP3285704B2 (en) 1994-06-16 1994-06-16 Speech recognition method and apparatus for spoken dialogue

Country Status (1)

Country Link
JP (1) JP3285704B2 (en)

Cited By (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2001324992A (en) * 2000-05-15 2001-11-22 Fujitsu Ten Ltd Voice synthesizer and voice data storage medium
JP2006215418A (en) * 2005-02-07 2006-08-17 Nissan Motor Co Ltd Voice input device and voice input method
JP2010515292A (en) * 2006-12-25 2010-05-06 トムソン ライセンシング Method and apparatus for automatic gain control
JP2012073364A (en) * 2010-09-28 2012-04-12 Toshiba Corp Voice interactive device, method, program
JP2018017776A (en) * 2016-07-25 2018-02-01 トヨタ自動車株式会社 Voice interactive device

Cited By (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2001324992A (en) * 2000-05-15 2001-11-22 Fujitsu Ten Ltd Voice synthesizer and voice data storage medium
JP2006215418A (en) * 2005-02-07 2006-08-17 Nissan Motor Co Ltd Voice input device and voice input method
JP2010515292A (en) * 2006-12-25 2010-05-06 トムソン ライセンシング Method and apparatus for automatic gain control
JP2012073364A (en) * 2010-09-28 2012-04-12 Toshiba Corp Voice interactive device, method, program
JP2018017776A (en) * 2016-07-25 2018-02-01 トヨタ自動車株式会社 Voice interactive device

Also Published As

Publication number Publication date
JP3285704B2 (en) 2002-05-27

Similar Documents

Publication Publication Date Title
US7401017B2 (en) Adaptive multi-pass speech recognition system
EP3433855B1 (en) Speaker verification method and system
JP3004883B2 (en) End call detection method and apparatus and continuous speech recognition method and apparatus
JP5381988B2 (en) Dialogue speech recognition system, dialogue speech recognition method, and dialogue speech recognition program
CN1639768B (en) Automatic speech recognition method and device
US12424223B2 (en) Voice-controlled communication requests and responses
GB2380381A (en) Speech synthesis method and apparatus
JP2004333543A (en) Voice interaction system and voice interaction method
US10143027B1 (en) Device selection for routing of communications
US7177806B2 (en) Sound signal recognition system and sound signal recognition method, and dialog control system and dialog control method using sound signal recognition system
US10854196B1 (en) Functional prerequisites and acknowledgments
US20070136060A1 (en) Recognizing entries in lexical lists
JP3285704B2 (en) Speech recognition method and apparatus for spoken dialogue
JPH08263092A (en) Response voice generation method and voice dialogue system
US11172527B2 (en) Routing of communications to a device
JP2004251998A (en) Dialogue understanding device
JP2871420B2 (en) Spoken dialogue system
Kuroiwa et al. Robust speech detection method for telephone speech recognition system
JP4094255B2 (en) Dictation device with command input function
US11735178B1 (en) Speech-processing system
JP2003330487A (en) Conversation agent
JPH07230293A (en) Voice recognizer
US20070129945A1 (en) Voice quality control for high quality speech reconstruction
JP4636695B2 (en) voice recognition
JP2005062398A (en) Speech recognition apparatus for speech recognition, speech data collection method for speech recognition, and computer program

Legal Events

Date Code Title Description
A01 Written decision to grant a patent or to grant a registration (utility model)

Free format text: JAPANESE INTERMEDIATE CODE: A01

Effective date: 20020212

LAPS Cancellation because of no payment of annual fees