JPH1055196A - Speech recognition device and method, information storage medium - Google Patents
Speech recognition device and method, information storage mediumInfo
- Publication number
- JPH1055196A JPH1055196A JP8211078A JP21107896A JPH1055196A JP H1055196 A JPH1055196 A JP H1055196A JP 8211078 A JP8211078 A JP 8211078A JP 21107896 A JP21107896 A JP 21107896A JP H1055196 A JPH1055196 A JP H1055196A
- Authority
- JP
- Japan
- Prior art keywords
- word
- recognition
- words
- central
- speech
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Abstract
Description
【0001】[0001]
【発明の属する技術分野】本発明は、音声を認識する音
声認識装置および方法と、そのプログラム等のソフトウ
ェアが書き込まれた情報記憶媒体に関する。BACKGROUND OF THE INVENTION 1. Field of the Invention The present invention relates to a speech recognition apparatus and method for recognizing speech, and an information storage medium on which software such as a program is written.
【0002】[0002]
【従来の技術】現在、人間が発声した音声を認識する音
声認識装置が要望されており、各種の音声認識方法が考
えられている。人間が語句である単語を一つだけ発声す
る場合、これを音声認識装置が認識することは困難では
ないが、人間の自然な会話では音声は連続しており、そ
こには多数の単語が助詞等を介して含まれている。この
ように連続的な音声から必要な単語を認識する手法とし
てはワードスポッティングが考案されており、これは予
め設定された認識候補の単語を連続的な会話音声から抽
出して認識する。2. Description of the Related Art At present, there is a demand for a voice recognition device for recognizing a voice uttered by a human, and various voice recognition methods are being considered. When a human utters only one word, which is a phrase, it is not difficult for the speech recognizer to recognize it, but in a natural human conversation, speech is continuous, and many words contain particle And so on. As a method for recognizing necessary words from continuous speech, word spotting has been devised. In this method, words of preset recognition candidates are extracted from continuous speech and recognized.
【0003】このようなワードスポッティングを実行す
る音声認識装置は、認識候補の単語毎に読みが格納され
た単語辞書を有しており、連続入力の会話音声と単語辞
書の全部の単語の読みとをマッチングさせ、このマッチ
ングのスコアが基準値を超過した最大の単語を会話音声
から認識する。しかし、単純にワードスポッティングを
実行しても良好な結果は期待できないため、現在では言
語的な制約により認識精度を向上させることが一般的で
ある。A speech recognition apparatus that performs such word spotting has a word dictionary in which readings are stored for each recognition candidate word. And the largest word whose score of the matching exceeds the reference value is recognized from the conversation voice. However, good results cannot be expected even if the word spotting is simply executed, so that it is now general to improve recognition accuracy due to linguistic restrictions.
【0004】このような言語的な制約には、文法等の構
文的な性質に基づくものと、格パターン等の意味的な制
約に基づくものとがある。前者は制約として強力である
が、書き言葉と違い規則化しにくい話し言葉には向いて
おらず、その認識制御は文節内程度にしか利用できな
い。後者は語順が自由な日本語の性質、話し言葉に向い
ており、上述したワードスポッティングに利用すると認
識精度を良好に向上させることができる。Such linguistic constraints include those based on syntactic properties such as grammar, and those based on semantic constraints such as case patterns. The former is powerful as a constraint, but it is not suitable for spoken language that is difficult to regularize unlike written language, and its recognition control can be used only within phrases. The latter is suitable for the characteristics of Japanese and the spoken language in which the word order is free, and when used for the word spotting described above, the recognition accuracy can be improved satisfactorily.
【0005】[0005]
【発明が解決しようとする課題】上述のようにワードス
ポッティングでは、連続音声から必要な語句のみ認識す
ることができ、特に意味的な制約を利用すると認識精度
を向上させることができる。As described above, in word spotting, only necessary words and phrases can be recognized from continuous speech. In particular, recognition accuracy can be improved by using semantic constraints.
【0006】例えば、特開平6-102897号公報に開示され
た音声認識装置の音声認識方法では、格関係を利用して
認識する文節を予測し、絞り込みを行なっている。しか
し、これでは格関係以外の関係に対処することができ
ず、複数の関係を取り扱うこともできない。つまり、連
続音声に含まれる語句の関係は格関係だけではなく、一
つの連続音声に複数の関係が含まれることも一般的であ
る。For example, in the speech recognition method of the speech recognition apparatus disclosed in Japanese Patent Application Laid-Open No. 6-102897, a phrase to be recognized is predicted using case relations, and narrowing is performed. However, this cannot deal with relationships other than case relationships, nor can it handle multiple relationships. In other words, the relationship between words included in the continuous speech is not limited to the case relationship, but it is also common that one continuous speech includes a plurality of relationships.
【0007】例えば、会話音声が「カップラーメンのカ
レー味を一箱ください」である場合、格関係を利用して
も語句を認識することは困難である。また、会話音声が
「泣く子も黙る」である場合、「泣く・子」「子・黙
る」なる複数の関係が存在している。[0007] For example, if the conversation voice is "Please give a box of curry flavor of cup ramen", it is difficult to recognize words even by using case relations. In addition, when the conversation voice is “crying child is also silent”, there are a plurality of relationships of “crying / child” and “child / silent”.
【0008】[0008]
【課題を解決するための手段】請求項1記載の発明の音
声認識装置は、認識対象の音声の連続的な入力を受け付
ける音声入力手段と、共起関係にある複数の語句が組み
合わされて格納された認識語句辞書と、連続的な入力音
声から共起関係で組み合わされた複数の語句を認識する
語句認識手段とを有する。従って、認識語句辞書には共
起関係にある複数の語句が組み合わされて格納されてい
るので、音声入力手段に連続的に入力された認識対象の
音声から共起関係で組み合わされた複数の語句が語句認
識手段により認識される。つまり、所定の語句を共起関
係で組み合わせて設定しておけば、この共起関係にある
複数の語句は、一つの連続的な入力音声から個々に単独
で認識されず、共起関係の組み合わせに基づいて認識さ
れる。According to a first aspect of the present invention, there is provided a voice recognition apparatus for storing a combination of voice input means for receiving a continuous input of a voice to be recognized and a plurality of co-occurring words. And a phrase recognition unit for recognizing a plurality of phrases combined in a co-occurrence relationship from continuous input speech. Therefore, since a plurality of words having a co-occurrence relation are stored in combination in the recognition word dictionary, a plurality of words combined in a co-occurrence relation from the speech to be recognized continuously input to the voice input means are stored. Is recognized by the phrase recognition means. In other words, if predetermined words are combined and set in a co-occurrence relationship, a plurality of words in this co-occurrence relationship cannot be individually recognized from one continuous input voice, Is recognized based on
【0009】請求項2記載の発明では、請求項1記載の
音声認識装置において、語句認識手段は、共起関係で組
み合わされた一対の語句の一方である中心語を認識語句
辞書から読み出して入力音声から抽出してから、この抽
出された中心語と共起関係にある他方の語句である付属
語を前記認識語句辞書から読み出して入力音声から抽出
する。従って、一つの連続的な入力音声から複数の語句
が語句認識手段により認識される場合、この語句の一方
である中心語が最初に入力音声から抽出されてから、こ
の中心語と共起関係にある他方の語句である付属語が次
に入力音声から抽出される。つまり、中心語は従来のワ
ードスポッティングと同様に多数を入力音声に照合させ
ることになるが、付属語は共起関係に基づいて絞り込ま
れてから入力音声に照合させることになる。According to a second aspect of the present invention, in the speech recognition apparatus according to the first aspect, the word / phrase recognizing means reads out and inputs a central word which is one of a pair of words combined in a co-occurrence relation from the recognized word / phrase dictionary. After being extracted from the voice, an additional word, which is the other word having a co-occurrence relationship with the extracted central word, is read from the recognized phrase dictionary and extracted from the input voice. Therefore, when a plurality of words are recognized from one continuous input voice by the phrase recognition means, a central word which is one of the words is first extracted from the input voice, and then a co-occurrence relation with the central word is obtained. One other phrase, the adjunct, is then extracted from the input speech. In other words, as in the case of the conventional word spotting, a large number of central words are collated with the input voice, but the attached words are narrowed down based on the co-occurrence relation and then collated with the input voice.
【0010】請求項3記載の発明では、請求項2記載の
音声認識装置において、語句認識手段は、入力音声の中
心語を抽出した区間を排除した区間から付属語を抽出す
る。従って、一つの連続的な入力音声から共起関係にあ
る中心語と付属語とが語句認識手段により認識される場
合、最初に入力音声の全域から中心語が抽出され、この
中心語が抽出された区間以外の区間から付属語が抽出さ
れる。つまり、中心語は従来のワードスポッティングと
同様に多数を入力音声に照合させることになるが、付属
語は共起関係に基づいて絞り込まれてから中心語と重複
しない音声区間に照合させることになる。According to a third aspect of the present invention, in the speech recognition apparatus according to the second aspect, the phrase recognizing means extracts an auxiliary word from a section excluding a section from which a central word of the input speech is extracted. Therefore, when a central word and an auxiliary word having a co-occurrence relation are recognized from one continuous input voice by the phrase recognition means, first, a central word is extracted from the entire input voice, and this central word is extracted. Adjectives are extracted from sections other than the section that has been set. In other words, the central word will be matched against the input voice as in the conventional word spotting, but the attached words will be narrowed down based on the co-occurrence relationship and then matched against the voice section that does not overlap with the central word .
【0011】請求項4記載の発明では、請求項2または
3記載の音声認識装置において、認識語句辞書は、中心
語と付属語との組み合わせに順番の情報も付与されてお
り、語句認識手段は、中心語と付属語とを順番に対応し
て入力音声から認識する。従って、中心語と付属語とが
連続音声に発生する順番の情報も認識語句辞書に格納さ
れており、中心語と付属語とは入力音声に所定の順番で
発生すると語句認識手段により認識されるので、中心語
と付属語とが入力音声から個別に認識されるような場合
でも順番が適正でないと認識されない。According to a fourth aspect of the present invention, in the speech recognition apparatus according to the second or third aspect, the recognition word dictionary is further provided with information on the order of the combination of the central word and the adjunct word, and the word recognition means is , The central word and the auxiliary word are recognized from the input speech in order. Accordingly, information on the order in which the central word and the auxiliary word occur in the continuous speech is also stored in the recognition phrase dictionary, and when the central word and the auxiliary word occur in the input voice in a predetermined order, they are recognized by the phrase recognition means. Therefore, even when the central word and the auxiliary word are individually recognized from the input voice, they are not recognized unless the order is proper.
【0012】請求項5記載の発明では、請求項4記載の
音声認識装置において、認識語句辞書は、中心語と付属
語との組み合わせに中間に位置する介在語も格納されて
おり、語句認識手段は、中心語と介在語と付属語とを入
力音声から認識する。従って、中心語と付属語との組み
合わせに中間に位置する介在語も認識語句辞書に格納さ
れており、中心語と介在語と付属語とが入力音声から語
句認識手段により認識されるので、ある中心語と介在語
との付属語との読みが、他の中心語と付属語との組み合
わせの読みと同一の場合でも、これらが各々別個に認識
される。According to a fifth aspect of the present invention, in the speech recognition apparatus according to the fourth aspect, the recognition word dictionary also stores an intervening word located in the middle of the combination of the central word and the adjunct word. Recognizes a central word, intervening words, and adjunct words from input speech. Therefore, the intervening word located in the middle of the combination of the central word and the adjunct word is also stored in the recognition phrase dictionary, and the central word, the intervening word and the adjunct word are recognized from the input speech by the phrase recognition means. Even if the reading of the adjunct of the central word and the intervening word is the same as the reading of the combination of the other central words and adjuncts, they are each recognized separately.
【0013】請求項6記載の発明では、請求項2または
3記載の音声認識装置において、認識語句辞書は、中心
語と付属語との組み合わせに時間間隔の情報も付与され
ており、語句認識手段は、中心語と付属語とを時間間隔
に対応して入力音声から認識する。従って、中心語と付
属語とが連続的な入力音声に発生する時間間隔の情報も
認識語句辞書に格納されており、中心語と付属語とが入
力音声から時間間隔に対応して語句認識手段により認識
されるので、中心語と付属語とは入力音声に適正な時間
間隔で発生した場合のみ認識され、中心語と付属語とが
入力音声から個別に認識されるような場合でも時間間隔
が適正でないと認識されない。According to a sixth aspect of the present invention, in the speech recognition apparatus according to the second or third aspect, the recognition word dictionary is provided with time interval information for a combination of a central word and an adjunct word. Recognizes a central word and an adjunct word from an input voice corresponding to a time interval. Therefore, information on the time interval in which the central word and the adjunct word occur in the continuous input speech is also stored in the recognition phrase dictionary, and the central word and the adjunct word are recognized from the input speech in correspondence with the time interval. Therefore, the central word and the adjunct word are recognized only when they occur in the input speech at appropriate time intervals. Even when the central word and the adjunct word are individually recognized from the input speech, the time interval is Not recognized as improper.
【0014】請求項7記載の発明では、請求項2または
3記載の音声認識装置において、認識語句辞書は、中心
語と付属語との組み合わせが複数段階の階層構造として
格納されており、語句認識手段は、一つの中心語と複数
の付属語とを階層構造に対応して入力音声から段階的に
認識する。従って、中心語と付属語との組み合わせが複
数段階の階層構造として認識語句辞書に格納されてお
り、一つの中心語と複数の付属語とが入力音声から階層
構造に対応して語句認識手段により段階的に認識され
る。つまり、ある入力音声から一つの中心語と一つの付
属語とが認識されると、この付属語を中心語とする他の
付属語も入力音声から検索され、このような処理動作が
順次繰り返されるので、複数段階の共起関係にある一つ
の中心語と複数の付属語とが順次認識される。According to a seventh aspect of the present invention, in the speech recognition apparatus according to the second or third aspect, the recognition phrase dictionary stores a combination of a central word and an adjunct word in a hierarchical structure having a plurality of stages, and performs phrase recognition. The means recognizes one central word and a plurality of attached words stepwise from the input speech in accordance with the hierarchical structure. Therefore, the combination of the central word and the adjunct word is stored in the recognition phrase dictionary as a multi-stage hierarchical structure, and one central word and a plurality of adjunct words are recognized by the phrase recognizing means according to the hierarchical structure from the input speech. Recognized step by step. That is, when one central word and one auxiliary word are recognized from a certain input voice, other auxiliary words having this auxiliary word as the central word are also searched from the input voice, and such processing operations are sequentially repeated. Therefore, one central word and a plurality of attached words that are in a co-occurrence relationship in a plurality of stages are sequentially recognized.
【0015】請求項8記載の発明では、請求項7記載の
音声認識装置において、認識語句辞書は、一つの中心語
と複数の付属語との組み合わせに順番の情報も付与され
ており、語句認識手段は、一つの中心語と複数の付属語
とを順番に対応して入力音声から認識する。従って、一
つの中心語と複数の付属語との組み合わせの順番の情報
も認識語句辞書に格納されており、一つの中心語と複数
の付属語とが入力音声から順番に対応して語句認識手段
により認識されるので、一つの中心語と複数の付属語と
が入力音声から個別に認識されるような場合でも各々の
順番が適正でないと認識されない。According to an eighth aspect of the present invention, in the speech recognition apparatus according to the seventh aspect, the recognition word dictionary is provided with order information for a combination of one central word and a plurality of attached words. The means recognizes one central word and a plurality of attached words in order from the input speech. Therefore, the information on the order of the combination of one central word and a plurality of attached words is also stored in the recognition phrase dictionary, and the one central word and the plurality of attached words correspond in order from the input voice to the phrase recognition means. Therefore, even when one central word and a plurality of attached words are individually recognized from the input speech, they are not recognized unless the respective orders are not proper.
【0016】請求項9記載の発明では、請求項7または
8記載の音声認識装置において、認識語句辞書は、一つ
の中心語と複数の付属語との組み合わせに階層構造の深
度の情報も付与されており、語句認識手段は、一つの中
心語と複数の付属語とを深度に対応して入力音声から認
識する。従って、一つの中心語と複数の付属語との組み
合わせの階層構造の深度の情報も認識語句辞書に格納さ
れており、一つの中心語と複数の付属語とが入力音声か
ら深度に対応して語句認識手段により認識されるので、
一つの中心語を規定とした複数の付属語の段階的な認識
が所定の深度まで実行される。According to a ninth aspect of the present invention, in the speech recognition apparatus according to the seventh or eighth aspect, the recognition phrase dictionary is also provided with information on a depth of a hierarchical structure to a combination of one central word and a plurality of attached words. The word recognition means recognizes one central word and a plurality of attached words from the input speech in accordance with the depth. Therefore, the information of the depth of the hierarchical structure of the combination of one central word and a plurality of attached words is also stored in the recognition phrase dictionary, and one central word and the plurality of attached words correspond to the depth from the input speech. Since it is recognized by the phrase recognition means,
Stepwise recognition of a plurality of adjunct words defining one central word is executed to a predetermined depth.
【0017】請求項10記載の音声認識方法は、共起関
係にある複数の語句を組み合わせて設定しておき、認識
対象の音声の連続的な入力を受け付け、この連続的な入
力音声から共起関係で組み合わされた複数の語句を認識
するようにした。従って、所定の語句を共起関係で組み
合わせて設定しておけば、この共起関係にある複数の語
句は、一つの連続的な入力音声から個々に単独で認識さ
れず、共起関係の組み合わせに基づいて認識される。According to a tenth aspect of the present invention, a plurality of words having a co-occurrence relation are set in combination, a continuous input of a speech to be recognized is received, and co-occurrence is performed from the continuous input speech. Recognize multiple words combined in relationships. Therefore, if predetermined words are combined and set in a co-occurrence relationship, a plurality of words in this co-occurrence relationship cannot be individually recognized from one continuous input voice, and Is recognized based on
【0018】請求項11記載の音声認識方法は、共起関
係にある中心語と付属語とを組み合わせて設定してお
き、認識対象の音声の連続的な入力を受け付け、この連
続的な入力音声から用意された中心語を抽出し、この中
心語と共起関係にある付属語を入力音声から抽出するよ
うにした。従って、一つの連続的な入力音声から共起関
係の中心語と付属語とが認識される場合、最初に入力音
声から中心語が抽出されてから、この中心語と共起関係
にある付属語が次に入力音声から抽出される。つまり、
中心語は従来のワードスポッティングと同様に多数を入
力音声に照合させることになるが、付属語は共起関係に
基づいて絞り込まれてから入力音声に照合させることに
なる。In the speech recognition method according to the eleventh aspect, a co-occurrence central word and an adjunct word are set in combination, a continuous input of a speech to be recognized is received, and the continuous input speech is received. A central word prepared from is extracted, and an auxiliary word having a co-occurrence relation with the central word is extracted from the input speech. Therefore, when the central word and the adjunct of the co-occurrence relation are recognized from one continuous input voice, the central word is first extracted from the input speech, and then the auxiliary word having the co-occurrence relation with the central word is recognized. Is then extracted from the input speech. That is,
As with the word word spotting in the related art, a large number of words are collated with the input voice, but the attached words are narrowed down based on the co-occurrence relation and then collated with the input voice.
【0019】請求項12記載の情報記憶媒体は、コンピ
ュータが読取自在なソフトウェアが予め書き込まれた情
報記憶媒体において、共起関係にある複数の語句が組み
合わされて格納される認識語句辞書のソフトウェアと、
連続的な入力音声から共起関係で組み合わされた複数の
語句を認識するためのプログラムと、が書き込まれてい
る。従って、この情報記憶媒体のソフトウェアをコンピ
ュータに読み取らせて動作させれば、このコンピュータ
は、認識語句辞書に格納されている語句を連続的な入力
音声から認識する音声認識装置として機能する。このと
き、認識語句辞書には共起関係にある複数の語句が組み
合わされて格納されているので、連続的な入力音声から
共起関係で組み合わされた複数の語句が認識される。つ
まり、所定の語句を共起関係で組み合わせて設定してお
けば、この共起関係にある複数の語句は、一つの連続的
な入力音声から個々に単独で認識されず、共起関係の組
み合わせに基づいて認識される。According to a twelfth aspect of the present invention, there is provided an information storage medium in which a plurality of words having a co-occurrence relation are combined and stored in an information storage medium in which computer-readable software is written in advance. ,
And a program for recognizing a plurality of phrases combined in a co-occurrence relationship from continuous input speech. Therefore, if the computer reads and operates the software of the information storage medium, the computer functions as a speech recognition device that recognizes words stored in the recognized word dictionary from continuous input speech. At this time, since a plurality of words having a co-occurrence relationship are stored in combination in the recognition word dictionary, a plurality of words combined in a co-occurrence relationship are recognized from continuous input speech. In other words, if predetermined words are combined and set in a co-occurrence relationship, a plurality of words in this co-occurrence relationship cannot be individually recognized from one continuous input voice, Is recognized based on
【0020】請求項13記載の情報記憶媒体は、コンピ
ュータが読取自在なソフトウェアが予め書き込まれた情
報記憶媒体において、共起関係にある中心語と付属語と
が組み合わされて格納される認識語句辞書のソフトウェ
アと、中心語を前記認識語句辞書から読み出して連続的
な入力音声から抽出するためのプログラムと、この抽出
された中心語と共起関係にある付属語を前記認識語句辞
書から読み出して入力音声から抽出するためのプログラ
ムと、が書き込まれている。従って、この情報記憶媒体
のソフトウェアをコンピュータに読み取らせて動作させ
れば、このコンピュータは、認識語句辞書に格納されて
いる語句を連続的な入力音声から認識する音声認識装置
として機能する。このとき、一つの連続的な入力音声か
ら複数の語句が語句認識手段により認識される場合、こ
の語句の一方である中心語が最初に入力音声から抽出さ
れてから、この中心語と共起関係にある他方の語句であ
る付属語が次に入力音声から抽出される。つまり、中心
語は従来のワードスポッティングと同様に多数を入力音
声に照合させることになるが、付属語は共起関係に基づ
いて絞り込まれてから入力音声に照合させることにな
る。According to a thirteenth aspect of the present invention, there is provided an information storage medium in which a computer-readable software is written in advance. And a program for reading a central word from the recognized phrase dictionary and extracting it from continuous input speech, and reading and inputting an auxiliary word having a co-occurrence relationship with the extracted central word from the recognized phrase dictionary And a program for extracting from the voice. Therefore, if the computer reads and operates the software of the information storage medium, the computer functions as a speech recognition device that recognizes words stored in the recognized word dictionary from continuous input speech. At this time, when a plurality of words are recognized from one continuous input voice by the phrase recognition means, a central word which is one of the words is first extracted from the input voice, and then a co-occurrence relation with the central word is obtained. Is then extracted from the input speech. In other words, as in the case of the conventional word spotting, a large number of central words are collated with the input voice, but the attached words are narrowed down based on the co-occurrence relation and then collated with the input voice.
【0021】[0021]
【発明の実施の形態】本発明の実施の一形態を図面に基
づいて以下に説明する。まず、本実施の形態の音声認識
装置1は、図2および図3に示すように、そのハードウ
ェアとしてデータ処理装置であるコンピュータシステム
を有している。このコンピュータシステムからなる音声
認識装置1は、コンピュータの主体としてCPU(Centr
al Processing Unit)2を有しており、このCPU2に
は、バスライン3により、ROM(Read Only Memory)
4、RAM(Random Access Memory)5、HD(Hard Disk
…図示せず)を内蔵したHDD(HD Drive)6、FD(Flo
ppy Disk)7が装填されるFDD(FD Drive)8、CD(C
ompact Disk)−ROM9が装填されるCD−ROMドラ
イブ10、マウス11が接続されたキーボード12、デ
ィスプレイ13、入力デバイスであるマイクロフォン1
4、通信I/F(Interface)15、等が接続されてい
る。DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS One embodiment of the present invention will be described below with reference to the drawings. First, as shown in FIGS. 2 and 3, the voice recognition device 1 of the present embodiment has a computer system as a data processing device as hardware. The speech recognition device 1 including the computer system includes a CPU (Centr
al Processing Unit) 2, and the CPU 2 is connected to a ROM (Read Only Memory) by a bus line 3.
4, RAM (Random Access Memory) 5, HD (Hard Disk)
... HDD (HD Drive) 6 with built-in
ppy Disk) 7, FDD (FD Drive) 8, CD (C
ompact Disk) -CD-ROM drive 10 loaded with ROM 9, keyboard 12 connected to mouse 11, display 13, microphone 1 as input device
4, communication I / F (Interface) 15, etc. are connected.
【0022】この音声認識装置1は、前記CPU2に各
種の処理動作を実行させるプログラムが予め設定されて
おり、このプログラム等のソフトウェアは、例えば、情
報記憶媒体である前記FD7や前記CD−ROM9に予
め書き込まれている。そして、このソフトウェアが情報
記憶媒体である前記HDD6にインストールされてお
り、これが起動時に情報記憶媒体である前記RAM5に
複写されて前記CPU2に読み取られる。In the speech recognition apparatus 1, a program for causing the CPU 2 to execute various processing operations is set in advance, and software such as this program is stored in, for example, the FD 7 or the CD-ROM 9 which is an information storage medium. It has been written in advance. This software is installed in the HDD 6 as an information storage medium, and is copied to the RAM 5 as an information storage medium and read by the CPU 2 at startup.
【0023】このようにソフトウェアを前記CPU2が
読み取って各種の処理動作を実行することにより、各種
機能が各種手段として実現されている。このような各種
手段として、本実施の形態の音声認識装置1は、図1に
示すように、音声入力手段21、認識語句辞書22、語
句認識手段23、結果出力手段24、等を備えている。
前記認識語句辞書22は、単語辞書25、共起辞書2
6、語順辞書27、からなり、前記語句認識手段23
は、特徴抽出手段28、候補認識手段29、候補読出手
段30、語句探索手段31、結果確定手段32、等から
なる。As described above, various functions are realized as various means by the CPU 2 reading the software and executing various processing operations. As such various means, the voice recognition device 1 of the present embodiment includes a voice input unit 21, a recognized phrase dictionary 22, a phrase recognition unit 23, a result output unit 24, and the like, as shown in FIG. .
The recognition phrase dictionary 22 includes a word dictionary 25, a co-occurrence dictionary 2
6, word order dictionary 27, said word recognition means 23
Consists of a feature extracting unit 28, a candidate recognizing unit 29, a candidate reading unit 30, a phrase searching unit 31, a result determining unit 32, and the like.
【0024】このような音声認識装置1の各種手段は、
必要により前記ディスプレイ13や前記マイクロフォン
14等のハードウェアも利用して実現されるが、その主
体は前記RAM5等に書き込まれたソフトウェアに対応
して前記CPU2が動作することにより実現されてい
る。このため、前記RAM5には、図4に示すように、
前記認識語句辞書22のソフトウェアである辞書記憶部
41と、連続的な入力音声から共起関係で組み合わされ
た複数の語句を認識するための制御プログラム42と、
が書き込まれている。Various means of such a speech recognition device 1 include:
If necessary, hardware such as the display 13 and the microphone 14 is used, and the main component is realized by the operation of the CPU 2 corresponding to software written in the RAM 5 or the like. For this reason, as shown in FIG.
A dictionary storage unit 41 which is software of the recognition phrase dictionary 22, a control program 42 for recognizing a plurality of phrases combined in a co-occurrence relationship from continuous input speech,
Is written.
【0025】より詳細には、前記辞書記憶部41は、単
語辞書記憶部43、共起辞書記憶部44、語順辞書記憶
部45からなり、これらの記憶部43〜45に前記辞書
25〜27の各種情報が格納されている。前記制御プロ
グラム42は、モジュール構造で形成されており、特徴
抽出モジュール46、単語照合モジュール47、認識制
御モジュール48等を有している。More specifically, the dictionary storage unit 41 comprises a word dictionary storage unit 43, a co-occurrence dictionary storage unit 44, and a word order dictionary storage unit 45. These storage units 43 to 45 store the dictionaries 25 to 27 Various information is stored. The control program 42 is formed in a module structure and includes a feature extraction module 46, a word matching module 47, a recognition control module 48, and the like.
【0026】前記特徴抽出モジュール46は、連続的な
入力音声のデジタル信号をフレーム毎に分析して特徴量
を抽出するためのプログラムからなり、前記単語照合モ
ジュール47は、入力音声の特徴量に認識候補の単語の
読みである特徴量を照合させて累積類似度をスコアとし
て算出するためのプログラムからなる。前記認識制御モ
ジュール48は、前記辞書25〜27から各種情報を読
み出して前記単語照合モジュール47に伝送し、この単
語照合モジュール47が出力するスコアに基づいて認識
結果の単語を確定するためのプログラムからなる。The feature extraction module 46 comprises a program for analyzing a digital signal of a continuous input speech for each frame and extracting a feature, and the word matching module 47 recognizes the feature of the input speech. The program comprises a program for collating feature amounts which are readings of candidate words and calculating a cumulative similarity as a score. The recognition control module 48 reads various types of information from the dictionaries 25 to 27 and transmits the information to the word matching module 47. From the program for determining the word of the recognition result based on the score output by the word matching module 47, Become.
【0027】前記認識語句辞書22は、上述のように前
記RAM5の辞書記憶部41にデータファイルとして格
納されており、図5に示すように、その一部である前記
単語辞書25には、語句である複数の単語毎に意味と読
みとが登録されている。前記共起辞書26には、共起関
係にある複数の語句が、ここでは一つの中心語と複数の
付属語との組み合わせで格納されている。前記語順辞書
27には、中心語と付属語の意味の情報とが順番に設定
されているので、中心語と付属語との順番の情報が格納
されている。The recognition phrase dictionary 22 is stored as a data file in the dictionary storage unit 41 of the RAM 5 as described above, and as shown in FIG. Are registered for each of a plurality of words. In the co-occurrence dictionary 26, a plurality of words having a co-occurrence relation are stored here in a combination of one central word and a plurality of attached words. In the word order dictionary 27, information on the meaning of the central word and the auxiliary word is set in order, so that information on the order of the central word and the auxiliary word is stored.
【0028】なお、ここでは説明を簡略化するために前
記辞書25〜27の各種情報を日本語の文字として表現
しているが、実際のソフトウェアでは単語等は識別コー
ドからなり、読みは単語の音声の特徴量として設定され
ている。この読みの音声の特徴量は、例えば、音素単位
の状態遷移モデルと単語単位の音素ネットワークとして
設定されており、各状態には平均特徴量と継続時間長と
の情報が設定されている。Although the various information in the dictionaries 25 to 27 are expressed as Japanese characters for the sake of simplicity of description, words and the like are composed of identification codes in actual software, and the reading is performed in words. It is set as a feature value of the voice. The feature amount of the reading voice is set, for example, as a state transition model in phoneme units and a phoneme network in word units, and information on the average feature amount and duration is set in each state.
【0029】前記音声入力手段21は、前記マイクロフ
ォン14等により音声の連続的な入力を受け付け、この
入力音声をデジタルの電気信号にA/D(Analog/Digi
tal)変換する。前記語句認識手段23は、前記CPU2
が前記RAM5に格納された前記制御プログラム42を
読み取って対応する処理動作を実行することにより、連
続的な入力音声のデジタル信号から共起関係にある中心
語と付属語とを認識する。The voice input means 21 receives continuous voice input from the microphone 14 or the like, and converts the input voice into a digital electric signal by A / D (Analog / Digilog).
tal) The phrase recognizing means 23 includes the CPU 2
Reads the control program 42 stored in the RAM 5 and executes the corresponding processing operation, thereby recognizing the co-occurrence central word and adjunct word from the continuous input voice digital signal.
【0030】より詳細には、前記語句認識手段23の特
徴抽出手段28は、前記CPU2が前記制御プログラム
42の前記特徴抽出モジュール46を読み取って対応す
る演算処理を実行することにより、連続的な入力音声の
デジタル信号を単位時間であるフレーム毎に分析し、例
えば、LPC(Linear Predictive Coding)メルケプスト
ラムの算出により特徴量を抽出する。More specifically, the feature extracting means 28 of the phrase recognizing means 23 reads the feature extracting module 46 of the control program 42 and executes the corresponding arithmetic processing to execute the continuous input. The audio digital signal is analyzed for each frame, which is a unit time, and a feature amount is extracted by calculating, for example, an LPC (Linear Predictive Coding) mel-cepstral.
【0031】前記候補認識手段29は、前記認識制御モ
ジュール48に対応した前記CPU2の演算処理によ
り、前記認識語句辞書22の共起辞書26から全部の中
心語を読み出してから、その各々の読みを前記単語辞書
25から読み出す。さらに、前記単語照合モジュール4
7に対応した前記CPU2の演算処理により、全部の中
心語の読みを連続的な入力音声の特徴量に照合させ、そ
の各々の類似度をフレーム単位で算出して順次累積し、
この累積類似度であるスコアが最大で基準値を超過した
一つの中心語を認識候補として抽出する。The candidate recognizing means 29 reads out all the central words from the co-occurrence dictionary 26 of the recognized phrase dictionary 22 by the arithmetic processing of the CPU 2 corresponding to the recognition control module 48, and then reads each of them. Read from the word dictionary 25. Further, the word matching module 4
7, the reading of all the central words is compared with the feature amount of the continuous input voice, the similarity of each of them is calculated for each frame, and sequentially accumulated,
One central word whose score as the accumulated similarity exceeds the reference value at the maximum is extracted as a recognition candidate.
【0032】前記候補読出手段30は、前記認識制御モ
ジュール48に対応した前記CPU2の演算処理によ
り、認識候補の中心語と共起関係にある全部の付属語を
認識候補として前記認識語句辞書22の共起辞書26か
ら検出し、前記語句探索手段31は、前記単語照合モジ
ュール47に対応した前記CPU2の演算処理により、
認識候補の全部の付属語を入力音声に照合させる。The candidate reading means 30 uses the CPU 2 corresponding to the recognition control module 48 to perform all the additional words co-occurring with the central word of the recognition candidate as recognition candidates by the CPU 2. Detected from the co-occurrence dictionary 26, the word / phrase search means 31 performs an arithmetic process of the CPU 2 corresponding to the word matching module 47,
Match all attached words of the recognition candidate to the input speech.
【0033】この処理動作も中心語の場合と同様に、認
識候補の付属語の読みが前記認識語句辞書22の単語辞
書25から読み出され、これと入力音声との照合のスコ
アが最大で基準値を超過した一つの付属語が抽出され
る。ただし、このように付属語の読みを入力音声の特徴
量に照合させる際、前記認識制御モジュール48に対応
した前記CPU2の演算処理により、前記認識語句辞書
22の語順辞書27から中心語と付属語との順番の情報
が読み出され、その順番に対応して入力音声の中心語が
抽出された区間以外の区間のみに付属語の読みが照合さ
れる。In this processing operation, similarly to the case of the central word, the reading of the adjunct word of the recognition candidate is read from the word dictionary 25 of the recognized phrase dictionary 22, and the score of the collation between the word and the input speech is the maximum. One adjunct that exceeds the value is extracted. However, when the reading of the attached word is compared with the feature amount of the input voice in this way, the central word and the attached word are obtained from the word order dictionary 27 of the recognized phrase dictionary 22 by the arithmetic processing of the CPU 2 corresponding to the recognition control module 48. Is read out, and the reading of the attached word is collated only in a section other than the section in which the central word of the input voice is extracted corresponding to the order.
【0034】前記結果確定手段32は、上述のように認
識候補の付属語が入力音声から探索されると、前記認識
制御モジュール48に対応した前記CPU2の演算処理
により、認識候補の中心語と付属語とを認識結果として
確定する。前記結果出力手段24は、上述のように確定
された認識結果の中心語と付属語とを、例えば、前記デ
ィスプレイ13の文字表示により出力する。When the attached word of the recognition candidate is searched for from the input speech as described above, the result determining means 32 calculates the central word of the recognized candidate and the attached word by the arithmetic processing of the CPU 2 corresponding to the recognition control module 48. A word is determined as a recognition result. The result output means 24 outputs the central word and the auxiliary word of the recognition result determined as described above, for example, by character display on the display 13.
【0035】このような構成において、本実施の形態の
音声認識装置1は、多数の単語が連続する会話の音声が
マイクロフォン14に入力されると、この連続的な入力
音声から認識語句辞書22に格納されている単語を認識
し、この認識結果をディスプレイ13に表示出力する。
このような音声認識装置1の音声認識方法を、図6を参
照して以下に順次詳述する。In such a configuration, the speech recognition apparatus 1 according to the present embodiment, when a speech of a conversation in which a large number of words are continuous is input to the microphone 14, converts the continuous input speech into the recognized phrase dictionary 22. The stored words are recognized, and the recognition results are displayed on the display 13.
Such a voice recognition method of the voice recognition device 1 will be sequentially described in detail below with reference to FIG.
【0036】まず、マイクロフォン14に連続的に入力
された音声は、その全域がデジタル信号にA/D変換さ
れ、フレーム毎に特徴量が抽出される。つぎに、認識語
句辞書22から全部の中心語の読みが読み出され、この
全部の読みが一つの入力音声の全域の特徴量と照合され
る。このように全部の中心語に対して照合の累積類似度
がスコアとして算出されると、基準値を超過したスコア
が最大の中心語が認識候補として一つだけ選出される。First, the whole area of the voice continuously input to the microphone 14 is A / D converted into a digital signal, and the characteristic amount is extracted for each frame. Next, the readings of all the central words are read from the recognized phrase dictionary 22, and all the readings are collated with the feature values of the whole input voice. When the cumulative similarity of collation is calculated as a score for all the central words in this way, only one central word having the largest score exceeding the reference value is selected as a recognition candidate.
【0037】このように認識候補の中心語が選出される
と、これと共起関係にある全部の付属語が認識候補とし
て認識語句辞書22から読み出され、上述した中心語の
場合と同様に、全部の付属語の読みと入力音声とが照合
されて基準値を超過したスコアが最大の一つの付属語が
選出される。このとき、認識語句辞書22の語順辞書2
7から中心語と付属語との順番の情報が読み出され、そ
の順番に対応して入力音声の中心語が抽出された区間以
外の区間のみに付属語の読みが照合される。When the central word of the recognition candidate is selected as described above, all the auxiliary words co-occurring therewith are read out from the recognition phrase dictionary 22 as the recognition candidate, and the same as in the case of the central word described above. Then, the reading of all the adjuncts is compared with the input voice, and one adjunct having the maximum score exceeding the reference value is selected. At this time, the word order dictionary 2 of the recognition phrase dictionary 22
7, information on the order of the central word and the auxiliary word is read out, and the reading of the auxiliary word is collated only in a section other than the section in which the central word of the input voice is extracted in accordance with the order.
【0038】このように一つの連続的な入力音声から共
起関係にある中心語と付属語とが抽出されると、これが
認識結果として確定されてディスプレイ13の文字表示
により出力される。なお、最初に全部の中心語の照合の
スコアが基準値を超過しない場合には認識結果は無しと
され、中心語の認識候補が検出された状態で付属語が認
識語句辞書22に格納されていない場合や、全部の付属
語の照合のスコアが基準値を超過しない場合には、認識
結果は中心語のみとされる。When the co-occurrence of the central word and the adjunct word is extracted from one continuous input voice, they are determined as a recognition result and output by the character display on the display 13. If the scores of the collation of all the central words do not exceed the reference value at first, no recognition result is found, and the attached word is stored in the recognized phrase dictionary 22 in a state where the candidate for the central word is detected. If there is no matching word, or if the scores of collation of all attached words do not exceed the reference value, the recognition result is only the central word.
【0039】上述した一連の処理動作を図5を参考に具
体的に説明すると、一つの連続的な入力音声が「ご注文
はカップラーメンのカレー味ですね」の場合、これに対
して中心語である“カップラーメン,手焼き煎餅”が照
合され、スコアが高い“カップラーメン”が中心語の認
識候補として抽出される。この中心語“カップラーメ
ン”に共起する付属語として“カレー味,しょう油味”
が認識語句辞書22から検出され、この認識語句辞書2
2には中心語である“カップラーメン”より後方に意味
が“味”である付属語が位置することが規定されている
ので、連続的な入力音声から“カップラーメン”より以
後の区間「のカレー味ですね」が切り出される。この音
声区間のみに対して付属語である“カレー味,しょう油
味”が照合されるので、付属語である“カレー味”が認
識される。The above-described series of processing operations will be described in detail with reference to FIG. 5. If one continuous input voice is "order is curry taste of cup ramen""Cup ramen, hand-baked rice cracker" is collated, and "cup ramen" with a high score is extracted as a candidate for recognition of the central word. The curriculum flavor, soy sauce flavor is an auxiliary word that co-occurs with this central word “cup ramen”
Is detected from the recognition phrase dictionary 22, and the recognition phrase dictionary 2
2 specifies that an adjunct with the meaning “taste” is located behind the central word “cup ramen”, and therefore, from continuous input voice, the section “no” after “cup ramen” It has curry flavor. " Since the attached words “curry taste, soy sauce taste” are collated only for this voice section, the attached word “curry taste” is recognized.
【0040】本実施の形態の音声認識装置1の音声認識
方法では、上述のように一つの連続的な入力音声から共
起関係で組み合わされた中心語と付属語とが認識される
ので、これらの単語を個々に認識する場合より精度が良
好である。特に、中心語は多数を入力音声に照合させる
必要があるが、付属語は認識候補の中心語と共起関係に
あるもののみ入力音声に照合させれば良いので、この処
理負担が軽減されて処理速度が向上している。In the speech recognition method of the speech recognition apparatus 1 according to the present embodiment, the central word and the adjunct words combined in a co-occurrence relation are recognized from one continuous input speech as described above. Is more accurate than when recognizing the words of In particular, many central words need to be collated with the input speech, but only auxiliary words that have a co-occurrence with the central word of the recognition candidate need be collated with the input speech. Processing speed has been improved.
【0041】しかも、付属語は入力音声から中心語の区
間を排除した区間のみに照合させれば良く、この付属語
を照合させる区間も中心語との順番に基づいて一方に制
限されるので、さらに処理負担が軽減されて処理速度が
向上している。さらに、中心語に対する付属語の順番の
情報は、付属語の種類の情報により設定されており、複
数の付属語を個々に設定していないので、語順辞書27
の記憶容量も軽減されている。In addition, it is only necessary to collate the attached words only in the section where the section of the central word is excluded from the input voice, and the section in which the attached word is collated is limited to one based on the order with the central word. Further, the processing load is reduced and the processing speed is improved. Further, the information on the order of the attached words with respect to the central word is set based on the information on the type of the attached words, and a plurality of attached words are not individually set.
Storage capacity has also been reduced.
【0042】なお、本発明は上記形態に限定されるもの
ではなく、各種の変形を許容する。例えば、本実施の形
態では、最初に入力音声から照合のスコアが最大の中心
語を一つの認識候補として選出し、これと共起関係にあ
る付属語を入力音声に照合させてスコアが最大の一つを
選出することを例示したが、最初にスコアが基準値を超
過した複数の中心語を認識候補として抽出し、これらの
中心語と共起関係にある全部の付属語を入力音声に照合
させて各々のスコアを算出し、中心語と付属語とのスコ
アの合計が最大の組み合わせを認識結果とするようなこ
とも可能である。The present invention is not limited to the above-described embodiment, but allows various modifications. For example, in the present embodiment, first, a central word having the largest matching score is selected from the input speech as one recognition candidate, and an auxiliary word having a co-occurrence relationship with the selected central word is matched with the input speech to obtain the largest score. As an example of selecting one, multiple core words whose scores exceed the reference value are first extracted as recognition candidates, and all attached words co-occurring with these central words are collated with the input speech Then, each score is calculated, and the combination having the largest total of the scores of the central word and the auxiliary word can be used as the recognition result.
【0043】また、本実施の形態では、共起関係にある
複数の語句を中心語と付属語とに分類しておき、最初に
中心語を入力音声から抽出してから、この結果に基づい
て付属語を入力音声から抽出することを例示した。しか
し、このように共起関係にある複数の語句を中心語や付
属語として分類せず、全部の語句を同時に入力音声に照
合させ、合計のスコアが最大となる共起関係の組み合わ
せの語句を認識結果とするようなことも可能である。In the present embodiment, a plurality of words having a co-occurrence relation are classified into a central word and an adjunct word, and the central word is first extracted from the input speech, and based on this result, An example of extracting an accessory word from an input voice has been described. However, instead of classifying multiple words that are co-occurring in this way as central words or adjuncts, all words are matched against the input speech at the same time, and the words of the co-occurrence combination that maximizes the total score are It is also possible to use a recognition result.
【0044】さらに、本実施の形態では、認識語句辞書
22に各種辞書25〜27を用途別に個別に形成するこ
とにより、そのメンテナンスや情報登録を容易とするこ
とを想定したが、このような辞書25〜27を一つに組
み合わせた形態として認識語句辞書22を形成すること
も可能である。Further, in the present embodiment, it is assumed that maintenance and information registration are facilitated by separately forming various dictionaries 25 to 27 in the recognition phrase dictionary 22 for each application. It is also possible to form the recognition phrase dictionary 22 as a form in which 25 to 27 are combined into one.
【0045】また、本実施の形態では、音声認識装置1
をコンピュータシステムによる実験装置として想定し、
入力音声から認識した単語をディスプレイ13に表示す
ることを例示した。しかし、上述のような音声認識装置
1の各部をASIC(Application Specific Integrated
Circuit)として製作し、これを各種製品に組み込んで
音声制御に利用することも可能である。In this embodiment, the speech recognition device 1
As an experimental device using a computer system,
The display of the word recognized from the input voice on the display 13 has been exemplified. However, each part of the speech recognition device 1 as described above is integrated with an ASIC (Application Specific Integrated
It is also possible to manufacture it as a circuit and incorporate it into various products and use it for voice control.
【0046】さらに、本実施の形態では、RAM5等に
ソフトウェアとして格納されている制御プログラムに従
ってCPU2が動作することにより、音声認識装置1の
各部が実現されることを例示した。しかし、このような
各部の各々を固有のハードウェアとして製作することも
可能であり、一部をソフトウェアとしてRAM5等に格
納するとともに一部をハードウェアとして製作すること
も可能である。また、所定のソフトウェアが格納された
RAM5等や各部のハードウェアを、例えば、ファーム
ウェアとして製作することも可能である。Further, in the present embodiment, it has been exemplified that each unit of the speech recognition apparatus 1 is realized by the operation of the CPU 2 according to a control program stored as software in the RAM 5 or the like. However, it is also possible to manufacture each of these units as unique hardware, and it is also possible to store a part of the unit as software in the RAM 5 or the like and manufacture a part of the unit as hardware. Further, the RAM 5 or the like in which predetermined software is stored and hardware of each unit can be manufactured as firmware, for example.
【0047】また、本実施の形態では、音声認識装置1
の起動時に、HDD6に格納されているソフトウェアが
RAM5に複写され、このようにRAM5に格納された
ソフトウェアをCPU2が読み取ることを想定したが、
このようなソフトウェアをHDD6に格納したままCP
U2に利用させることや、RAM5に予め書き込んでお
くことも可能である。In the present embodiment, the speech recognition device 1
It is assumed that the software stored in the HDD 6 is copied to the RAM 5 at the time of startup, and the software stored in the RAM 5 is read by the CPU 2 as described above.
With such software stored in the HDD 6, the CP
It is also possible for U2 to use it or to write it in RAM5 in advance.
【0048】さらに、前述のように単体で取り扱える情
報記憶媒体であるFD7やCD−ROM9にソフトウェ
アを書き込んでおき、このFD7等からRAM5等にソ
フトウェアをインストールすることも可能であるが、こ
のようなインストールを実行することなくFD7等に書
き込まれたソフトウェアをCPU2が適宜読み取ってデ
ータ処理を実行することも可能である。Further, as described above, software can be written in the FD 7 or CD-ROM 9 which is an information storage medium that can be handled alone, and the software can be installed in the RAM 5 or the like from the FD 7 or the like. It is also possible for the CPU 2 to appropriately read software written in the FD 7 or the like without executing the installation and execute data processing.
【0049】また、このような音声認識装置1の各部を
実現する制御プログラムを、複数のソフトウェアの組み
合わせにより実現することも可能であり、その場合、単
体の製品となる情報記憶媒体には必要最小限のソフトウ
ェアのみを格納しておけば良い。例えば、オペレーティ
ングシステムが実装されている音声認識装置1に、CD
−ROM9等の情報記憶媒体によりアプリケーションソ
フトを提供するような場合、音声認識装置1の各部を実
現するソフトウェアは、アプリケーションソフトとオペ
レーティングシステムとの組み合わせで実現されるの
で、オペレーティングシステムに依存する部分のソフト
ウェアはアプリケーションソフトの情報記憶媒体から省
略することができる。It is also possible to realize a control program for realizing each part of the voice recognition apparatus 1 by a combination of a plurality of softwares. In this case, the information storage medium as a single product has a minimum required size. It is only necessary to store the limited software. For example, the voice recognition device 1 on which an operating system is mounted has a CD
In a case where application software is provided by an information storage medium such as the ROM 9, software that realizes each unit of the voice recognition device 1 is realized by a combination of the application software and the operating system. The software can be omitted from the information storage medium of the application software.
【0050】特に、本発明の音声認識装置1を、認識す
る単語が特定された業務用の装置等として製作する場合
は、その製造工程で認識語句辞書22の内容も固定的に
書き込めば良い。しかし、上述のように音声認識装置1
のアプリケーションソフトを一般ユーザに販売するよう
な場合には、認識語句辞書22の内容をユーザが自由に
登録できることが好ましい。In particular, when the speech recognition device 1 of the present invention is manufactured as a business device or the like in which words to be recognized are specified, the contents of the recognition phrase dictionary 22 may be fixedly written in the manufacturing process. However, as described above, the speech recognition device 1
It is preferable that the user can freely register the contents of the recognized phrase dictionary 22 when the application software is sold to general users.
【0051】このような製品としてCD−ROM9等の
情報記憶媒体を製造する場合には、前述した制御プログ
ラム42の他、認識語句辞書22をRAM5等に所定の
フォーマットで形成するためのプログラムと、認識語句
辞書22に各種情報を登録させるためのプログラムと
を、情報記憶媒体に書き込んでおくことになる。この場
合、これらのプログラムが情報記憶媒体における認識語
句辞書22のソフトウェアとなり、各種情報の設定澄み
の認識語句辞書22のソフトウェアは情報記憶媒体には
書き込まない。When an information storage medium such as the CD-ROM 9 is manufactured as such a product, in addition to the control program 42 described above, a program for forming the recognition phrase dictionary 22 in the RAM 5 or the like in a predetermined format includes: A program for registering various kinds of information in the recognition phrase dictionary 22 is written in the information storage medium. In this case, these programs become the software of the recognition phrase dictionary 22 in the information storage medium, and the software of the recognition phrase dictionary 22 with the setting of various information is not written in the information storage medium.
【0052】同様に、完成した製品として音声認識装置
1を製造する場合も、単語を認識する各種手段21,2
3,24等の部分は固定的に製作しておき、その認識語
句辞書22の設定内容を空白としてユーザに登録させる
ことも可能である。さらに、このような音声認識装置1
に交換自在に装着するオプション部品として、業務毎に
適正な単語を登録した認識語句辞書22を情報記憶媒体
として製作するようなことも可能である。Similarly, when manufacturing the speech recognition apparatus 1 as a completed product, various means 21 and 21 for recognizing words are used.
It is also possible to make the parts such as 3, 24, etc. fixedly, and let the user register the setting contents of the recognized phrase dictionary 22 as blank. Furthermore, such a speech recognition device 1
It is also possible to produce a recognition phrase dictionary 22 in which appropriate words are registered for each job as an information storage medium, as an optional component that can be exchangeably mounted on a computer.
【0053】なお、上述のように情報記憶媒体に書き込
んだソフトウェアをコンピュータに供給する手法は、そ
の情報記憶媒体をコンピュータに直接に装填することに
限定されない。例えば、上述のようなソフトウェアをホ
ストコンピュータの情報記憶媒体に書き込み、このホス
トコンピュータを通信ネットワークにより端末コンピュ
ータに接続し、ホストコンピュータからデータ通信によ
り端末コンピュータにソフトウェアを供給することも可
能である。The method for supplying the software written in the information storage medium to the computer as described above is not limited to loading the information storage medium directly into the computer. For example, it is also possible to write the above-mentioned software on an information storage medium of a host computer, connect the host computer to a terminal computer via a communication network, and supply the software to the terminal computer by data communication from the host computer.
【0054】この場合、端末コンピュータが自身の情報
記憶媒体にソフトウェアをダウンロードした状態でスタ
ンドアロンのデータ処理を実行することも可能である
が、ソフトウェアをダウンロードすることなくホストコ
ンピュータとのリアルタイムのデータ通信によりデータ
処理を実行することも可能である。この場合、ホストコ
ンピュータと端末コンピュータとを通信ネットワークに
より接続したシステム全体が、本発明の音声認識装置1
に相当することになる。In this case, it is possible for the terminal computer to execute stand-alone data processing in a state where the software has been downloaded to its own information storage medium, but it is possible to perform real-time data communication with the host computer without downloading the software. It is also possible to perform data processing. In this case, the entire system in which the host computer and the terminal computer are connected by the communication network is the voice recognition device 1 of the present invention.
Would be equivalent to
【0055】また、本実施の形態では、単語辞書25に
単語の意味を格納しておき、語順辞書27には付属語を
意味の情報として格納しておくことを例示したが、図7
に示すように、中心語に対する付属語の共起関係の種類
の情報を共起辞書26と語順辞書27とに格納しておく
ことも可能である。この場合、一つの中心語に複数の共
起関係で複数の付属語が対応しても、中心語に対する複
数の付属語の位置を共起関係の種類の情報で設定できる
ので、複数種類の共起関係の付属語を良好な精度で容易
に認識することができる。In the present embodiment, the word dictionary 25 stores the meaning of a word, and the word order dictionary 27 stores auxiliary words as meaning information.
As shown in (1), it is also possible to store information on the type of co-occurrence relation of an accessory word with respect to a central word in the co-occurrence dictionary 26 and the word order dictionary 27. In this case, even if a plurality of adjuncts correspond to one central word in a plurality of co-occurrence relations, the positions of a plurality of adjunct words with respect to the central word can be set by the information of the co-occurrence relation type. It is possible to easily recognize the auxiliary word of the starting relation with good accuracy.
【0056】例えば、「ご注文はカレー味のカップラー
メンですね」なる入力音声から“カップラーメン”が中
心語の認識候補として抽出された場合、この中心語“カ
ップラーメン”に共起する付属語としては“カレー味,
ミニ”が認識語句辞書22から検出される。しかし、こ
こでは種類が“味”の付属語は中心語より前方に位置す
ることが規定されており、種類が“サイズ”の付属語は
中心語より後方に位置することが規定されているので、
連続的な入力音声から「ご注文はカレー味の」の区間が
切り出されて“カレー味”の付属語が照合され、「です
ね」の音声区間が切り出されて“ミニ”の付属語が照合
される。For example, if “cup ramen” is extracted as a candidate for recognition of a central word from an input voice “order is curry flavored cup ramen,” an ancillary word co-occurring with this central word “cup ramen” As for "curry taste,
"Mini" is detected from the recognition phrase dictionary 22. However, here, it is specified that the adjunct of the type "taste" is located ahead of the central word, and the adjunct of the type "size" is the central word. Since it is specified that it is located further behind,
From the continuous input voice, the section of "Order is curry taste" is cut out and the adjunct word of "curry taste" is collated, and the voice section of "Issued" is cut out and the adjunct word of "mini" is collated Is done.
【0057】また、図8に示すように、認識語句辞書2
2の共起辞書26に、中心語と付属語との組み合わせに
中間に位置する介在語も格納しておき、語句認識手段2
3が、中心語と介在語と付属語とを入力音声から認識す
ることも可能である。この場合、中心語と介在語と付属
語とが一つの入力音声から認識されるので、ある中心語
と介在語との付属語との読みが、他の中心語と付属語と
の組み合わせの読みと同一の場合でも、これらを各々別
個に認識することができる。Also, as shown in FIG.
In the co-occurrence dictionary 26 of the second embodiment, an intervening word located in the middle of the combination of the central word and the auxiliary word is also stored.
3 can also recognize the central word, the intervening word, and the attached word from the input speech. In this case, since the central word, the intervening word and the adjunct word are recognized from one input voice, the reading of the adjunct word of a certain central word and the intervening word becomes the reading of the combination of another central word and the adjunct word. Even in the same case, these can be recognized separately from each other.
【0058】例えば、入力音声が「ご注文は手焼き煎餅
の緑茶風味ですね」の場合、中心語である“手焼き煎
餅”に対して付属語である“のり,緑茶風味”の両方が
「の緑茶風味」の音声区間から同等のスコアで認識され
ることになる。しかし、上述のように介在語として
“の”が規定されていれば、付属語として“緑茶風味”
のみを認識することができる。For example, if the input voice is “order is hand-baked rice cracker with green tea flavor”, both “nori and green tea flavor” as ancillary words are “a hand-roasted rice cracker” which is the central word. From the voice section of “green tea flavor” with the same score. However, if "no" is specified as an intervening word as described above, the adjunct "green tea flavor"
Can only recognize.
【0059】また、図9に示すように、認識語句辞書2
2の共起辞書26に、中心語と付属語との組み合わせと
ともに時間間隔の情報も格納しておき、語句認識手段2
3が、中心語と付属語とを時間間隔に対応して入力音声
から認識することも可能である。この場合、付属語を照
合させる入力音声の区間を時間間隔に対応して制限でき
るので、より高速に付属語を認識することができ、中心
語から極度に離反した付属語は認識されないので、不適
な付属語の認識を防止することもできる。Further, as shown in FIG.
In the co-occurrence dictionary 26, information on time intervals is stored together with the combination of the central word and the adjunct word.
3 can also recognize the central word and the adjunct word from the input speech corresponding to the time interval. In this case, the section of the input voice for matching the adjunct can be limited according to the time interval, so that the adjunct can be recognized more quickly and the adjunct that is extremely deviated from the central term is not recognized. It is also possible to prevent the recognition of extra adjuncts.
【0060】例えば、「ご注文はカップラーメンのカレ
ー味を二箱ですね」なる入力音声から“カップラーメ
ン”が中心語の認識候補として抽出された場合、その音
声区間から20フレーム以内の音声区間のみに意味が
“味”の付属語が照合されて“カレー味”が認識され、
100フレーム以内の音声区間のみ意味が“数量”の付属
語が照合されて“二箱”が認識される。For example, if “cup ramen” is extracted as a candidate for recognition of a central word from an input voice “Your order is two curry flavors of cup ramen,” a voice section within 20 frames from that voice section Only the adjuncts with the meaning “taste” are compared and “curry taste” is recognized.
Only the speech section within 100 frames is collated with the auxiliary word having the meaning of “quantity”, and “two boxes” is recognized.
【0061】また、図10に示すように、認識語句辞書
22の共起辞書26に、中心語と付属語との組み合わせ
を複数段階の階層構造として格納しておき、語句認識手
段23が、一つの中心語と複数の付属語とを階層構造に
対応して入力音声から段階的に認識することも可能であ
る。この場合、ある入力音声から一つの中心語と一つの
付属語とが認識されると、この付属語を中心語とする他
の付属語も入力音声から検索され、このような処理動作
が順次繰り返されるので、複数段階の共起関係にある一
つの中心語と複数の付属語とを順次認識することができ
る。As shown in FIG. 10, a combination of a central word and an adjunct word is stored in the co-occurrence dictionary 26 of the recognized phrase dictionary 22 as a hierarchical structure having a plurality of stages. It is also possible to recognize one central word and a plurality of attached words stepwise from the input speech in a hierarchical structure. In this case, when one central word and one auxiliary word are recognized from a certain input voice, other auxiliary words having this auxiliary word as the central word are also searched from the input voice, and such processing operations are sequentially repeated. Therefore, it is possible to sequentially recognize one central word and a plurality of attached words that have a co-occurrence relationship in a plurality of stages.
【0062】例えば、入力音声が「350の缶のビール
を下さい」の場合、最初に中心語として“ビール”が抽
出されて対応する付属語としては“缶”が抽出される。
次に、この“缶”を中心語として“350”なる付属語
が抽出されるので、一つの入力音声から最終的に三つの
単語が認識されることになる。For example, when the input voice is "Please give me 350 cans of beer", "beer" is first extracted as the central word, and "can" is extracted as the corresponding auxiliary word.
Next, an auxiliary word "350" is extracted with this "can" as the central word, so that three words are finally recognized from one input voice.
【0063】さらに、上述のように中心語と付属語との
組み合わせを複数段階の階層構造とした場合に、図11
に示すように、認識語句辞書22に、一つの中心語と複
数の付属語との組み合わせの順番の情報も格納してお
き、語句認識手段23が、一つの中心語と複数の付属語
とを順番に対応して入力音声から認識することも可能で
ある。この場合、一つの中心語と複数の付属語とが入力
音声から順番に対応して認識されるので、複数の付属語
を良好な精度で高速に認識することができる。Further, when the combination of the central word and the adjunct word has a hierarchical structure of a plurality of stages as described above, FIG.
As shown in (1), information on the order of combinations of one central word and a plurality of attached words is also stored in the recognized phrase dictionary 22, and the phrase recognition means 23 stores one central word and a plurality of attached words in the recognition phrase dictionary 22. It is also possible to recognize from the input voice corresponding to the order. In this case, since one central word and a plurality of attached words are recognized in order from the input speech, a plurality of attached words can be quickly and accurately recognized.
【0064】例えば、入力音声が「350の缶のビール
を下さい」の場合、最初に中心語として“ビール”が抽
出され、これより前方の音声区間である「350の缶
の」から意味が“形態”の付属語である“缶”が抽出さ
れ、これより前方の音声区間である「350の」から意
味が“サイズ”の付属語である“350”が抽出され
る。For example, when the input voice is “Please give me 350 cans of beer”, “beer” is first extracted as a central word, and the meaning is obtained from “350 cans of beer” which is a voice section ahead of this. "Can" as an adjunct word of "form" is extracted, and "350" which is an adjunct word of "size" is extracted from "350 no" which is a preceding voice section.
【0065】さらに、上述のように中心語と付属語との
組み合わせを複数段階の階層構造とした場合に、図12
に示すように、認識語句辞書22に、一つの中心語と複
数の付属語との組み合わせの階層構造の深度の情報も格
納しておき、語句認識手段23が、一つの中心語と複数
の付属語とを深度に対応して入力音声から認識すること
も可能である。この場合、一つの中心語から複数の付属
語を段階的に探索する処理動作が所定の深度まで実行さ
れるので、複数の付属語を必要な段階まで高速に認識す
ることができる。Further, when the combination of the central word and the adjunct word has a hierarchical structure of a plurality of stages as described above, FIG.
As shown in (1), information on the depth of the hierarchical structure of a combination of one central word and a plurality of attached words is also stored in the recognized phrase dictionary 22, and the phrase recognition means 23 stores one central word and a plurality of attached words. It is also possible to recognize words from input speech corresponding to depth. In this case, since the processing operation of searching for a plurality of attached words stepwise from one central word is executed to a predetermined depth, the plurality of attached words can be quickly recognized to a necessary stage.
【0066】例えば、必要な階層構造が“2”として設
定されており、入力音声が「350の缶のビールを下さ
い」の場合、最初に中心語として“ビール”が抽出され
てから第一の付属語として“缶”が抽出された時点で、
階層構造の深度は“1”となる。そこで、この“缶”を
中心語として第二の付属語として“350”が抽出され
ると、階層構造の深度は“2”となるので、この時点で
段階的な音声認識の処理動作を終了する。For example, when the required hierarchical structure is set as “2” and the input voice is “Please give me 350 cans of beer”, the first word “beer” is extracted as the central word and then the first When "can" is extracted as an appendix,
The depth of the hierarchical structure is “1”. Then, when "350" is extracted as a second adjunct word with "can" as the central word, the depth of the hierarchical structure becomes "2", and the step-by-step speech recognition processing operation ends at this point. I do.
【0067】[0067]
【発明の効果】請求項1記載の発明の音声認識装置は、
認識対象の音声の連続的な入力を受け付ける音声入力手
段と、共起関係にある複数の語句が組み合わされて格納
された認識語句辞書と、連続的な入力音声から共起関係
で組み合わされた複数の語句を認識する語句認識手段と
を有することにより、複数の語句を一つの連続的な入力
音声から共起関係の組み合わせに基づいて認識すること
ができるので、複数の語句を良好な精度で高速に認識す
ることができる。According to the first aspect of the present invention, there is provided a speech recognition apparatus.
A voice input means for receiving a continuous input of a speech to be recognized, a recognition phrase dictionary in which a plurality of co-occurring phrases are stored in combination, and a plurality of co-occurrence combinations of continuous input voices And the phrase recognition means for recognizing the multiple words can be recognized from one continuous input voice based on a combination of co-occurrence relations. Can be recognized.
【0068】請求項2記載の発明では、語句認識手段
は、共起関係で組み合わされた一対の語句の一方である
中心語を認識語句辞書から読み出して入力音声から抽出
してから、この抽出された中心語と共起関係にある他方
の語句である付属語を認識語句辞書から読み出して入力
音声から抽出することにより、中心語の抽出結果に基づ
いて入力音声に照合させる付属語を絞り込むことができ
るので、入力音声から付属語を認識する処理動作の負担
を軽減して速度を向上させることができ、共起関係にあ
る中心語と付属語とを良好な精度で高速に認識すること
ができる。According to the second aspect of the present invention, the word recognizing means reads out the central word, which is one of a pair of words combined in co-occurrence relation, from the recognized word dictionary and extracts it from the input speech, and then extracts the central word. By reading an auxiliary word, which is the other word co-occurring with the central word, from the recognized phrase dictionary and extracting it from the input speech, it is possible to narrow down the auxiliary words to be matched with the input voice based on the extraction result of the central word. Since it is possible, the burden of the processing operation of recognizing an adjunct word from the input speech can be reduced and the speed can be improved, and the co-occurring central word and the adjunct word can be quickly and accurately recognized with good accuracy. .
【0069】請求項3記載の発明では、語句認識手段
は、入力音声の中心語を抽出した区間を排除した区間か
ら付属語を抽出することにより、中心語の抽出結果に基
づいて付属語を照合させる入力音声の区間を制限するこ
とができるので、入力音声から付属語を認識する処理動
作の負担を軽減して速度を向上させることができ、共起
関係にある中心語と付属語とを良好な精度で高速に認識
することができる。According to the third aspect of the present invention, the phrase recognizing means extracts an adjunct word from a section excluding the section from which the central word of the input speech has been extracted, thereby collating the adjunct word based on the result of extracting the central word. Since the section of the input speech to be restricted can be limited, the load of the processing operation of recognizing the adjunct from the input speech can be reduced and the speed can be improved, and the co-occurrence relation between the central term and the adjunct can be improved. High-speed recognition with high accuracy.
【0070】請求項4記載の発明では、認識語句辞書
は、中心語と付属語との組み合わせに順番の情報も付与
されており、語句認識手段は、中心語と付属語とを順番
に対応して入力音声から認識することにより、付属語を
照合させる入力音声の区間を中心語の抽出区間より前方
か後方に制限することができるので、入力音声から付属
語を認識する処理動作の負担を軽減して速度を向上させ
ることができ、共起関係にある中心語と付属語とを良好
な精度で高速に認識することができる。According to the fourth aspect of the present invention, the recognition word dictionary is also provided with information on the order of the combination of the central word and the adjunct word, and the word / phrase recognition means associates the central word and the adjunct word in order. By recognizing from the input speech, the section of the input speech for collating the adjunct can be limited to the front or back from the extraction section of the central word, reducing the burden of the processing operation of recognizing the adjunct from the input speech As a result, the central word and the co-occurrence word which are co-occurring can be recognized at high speed with good accuracy.
【0071】請求項5記載の発明では、認識語句辞書
は、中心語と付属語との組み合わせに中間に位置する介
在語も格納されており、語句認識手段は、中心語と介在
語と付属語とを入力音声から認識することにより、例え
ば、ある中心語と介在語との付属語との読みが、他の中
心語と付属語との組み合わせの読みと同一の場合でも、
これらを各々別個に認識することができるので、共起関
係にある中心語と介在語と付属語とを良好な精度で認識
することができる。According to the fifth aspect of the present invention, the recognition word dictionary also stores an intervening word located in the middle of the combination of the central word and the adjunct word. By recognizing from the input voice, for example, even if the reading of the adjunct of a certain central word and the intervening word is the same as the reading of the combination of another central word and the adjunct,
Since these can be recognized separately, the co-occurrence of the central word, intervening word and adjunct word can be recognized with good accuracy.
【0072】請求項6記載の発明では、認識語句辞書
は、中心語と付属語との組み合わせに時間間隔の情報も
付与されており、語句認識手段は、中心語と付属語とを
時間間隔に対応して入力音声から認識することにより、
付属語を照合させる入力音声の区間を中心語の抽出区間
から所定の時間間隔の範囲に制限することができるの
で、入力音声から付属語を認識する処理動作の負担を軽
減して速度を向上させることができ、共起関係にある中
心語と付属語とを良好な精度で高速に認識することがで
きる。According to the sixth aspect of the present invention, the recognition phrase dictionary is also provided with information on a time interval for a combination of a central word and an adjunct word, and the phrase recognizing means converts the central word and the adjunct word into a time interval. By correspondingly recognizing from the input voice,
Since the section of the input speech for matching the adjunct can be limited to a range of a predetermined time interval from the central term extraction section, the load of the processing for recognizing the adjunct from the input speech is reduced and the speed is improved. Thus, the co-occurrence of the central word and the auxiliary word can be quickly recognized with good accuracy.
【0073】請求項7記載の発明では、認識語句辞書
は、中心語と付属語との組み合わせが複数段階の階層構
造として格納されており、語句認識手段は、一つの中心
語と複数の付属語とを階層構造に対応して入力音声から
段階的に認識することにより、複数段階の共起関係にあ
る一つの中心語と複数の付属語とを段階的に順次認識す
ることができ、一つの語句の抽出結果に基づいて入力音
声に照合させる次の語句を絞り込むことができるので、
入力音声から複数の語句を段階的に順次認識する処理動
作の負担を軽減して速度を向上させることができ、一つ
の入力音声から多数の語句を良好な精度で高速に認識す
ることができる。According to the seventh aspect of the present invention, the recognition phrase dictionary stores a combination of a central word and an adjunct word in a hierarchical structure having a plurality of stages, and the phrase recognizing means comprises one central word and a plurality of adjunct words. Are recognized stepwise from the input speech in accordance with the hierarchical structure, so that one central word and a plurality of adjuncts in a co-occurrence relation of a plurality of steps can be sequentially recognized step by step. Based on the phrase extraction results, you can narrow down the next phrase to match with the input voice,
The load on the processing operation of sequentially recognizing a plurality of words in a step-by-step manner from the input voice can be reduced and the speed can be improved, and many words can be recognized from one input voice at high speed with good accuracy.
【0074】請求項8記載の発明では、認識語句辞書
は、一つの中心語と複数の付属語との組み合わせに順番
の情報も付与されており、語句認識手段は、一つの中心
語と複数の付属語とを順番に対応して入力音声から認識
することにより、一つの語句の抽出結果に基づいて次の
語句を入力音声に照合させる場合に、この照合区間を直
前の語句の抽出区間より前方か後方に制限することがで
きるので、入力音声から複数の語句を段階的に順次認識
する処理動作の負担を軽減して速度を向上させることが
でき、一つの入力音声から多数の語句を良好な精度で高
速に認識することができる。According to the eighth aspect of the present invention, the recognition word dictionary is also provided with order information for a combination of one central word and a plurality of attached words. By recognizing adjunct words from the input speech in correspondence with each other, if the next phrase is to be collated with the input speech based on the extraction result of one phrase, this collation section is located ahead of the preceding phrase extraction section. Can be restricted to the rear, so that the load of the processing operation of recognizing a plurality of words in a step-by-step manner from the input voice can be reduced and the speed can be improved. It can be recognized at high speed with high accuracy.
【0075】請求項9記載の発明では、認識語句辞書
は、一つの中心語と複数の付属語との組み合わせに階層
構造の深度の情報も付与されており、語句認識手段は、
一つの中心語と複数の付属語とを深度に対応して入力音
声から認識することにより、複数段階の共起関係にある
一つの中心語と複数の付属語とを段階的に順次認識する
処理動作を所定の深度まで実行することができるので、
一つの入力音声から多数の語句を必要な段階まで認識す
ることができる。According to the ninth aspect of the present invention, in the recognition word dictionary, information on the depth of the hierarchical structure is also given to a combination of one central word and a plurality of attached words.
A process of sequentially recognizing one central word and a plurality of attached words in a multi-stage co-occurrence relationship by recognizing one central word and a plurality of attached words from the input speech in accordance with the depth Since the operation can be performed to a predetermined depth,
Many words and phrases can be recognized from one input voice to a necessary stage.
【0076】請求項10記載の音声認識方法は、共起関
係にある複数の語句を組み合わせて設定しておき、認識
対象の音声の連続的な入力を受け付け、この連続的な入
力音声から共起関係で組み合わされた複数の語句を認識
するようにしたことにより、複数の語句が一つの連続的
な入力音声から共起関係の組み合わせに基づいて認識さ
れるので、複数の語句を良好な精度で高速に認識するこ
とができる。According to a tenth aspect of the present invention, a plurality of words having a co-occurrence relationship are set in combination, a continuous input of a speech to be recognized is received, and a co-occurrence is obtained from the continuous input speech. By recognizing multiple words combined in a relationship, multiple words are recognized based on a combination of co-occurrence relationships from one continuous input voice, so multiple words can be recognized with good accuracy. Can be recognized at high speed.
【0077】請求項11記載の音声認識方法は、共起関
係にある中心語と付属語とを組み合わせて設定してお
き、認識対象の音声の連続的な入力を受け付け、この連
続的な入力音声から用意された中心語を抽出し、この中
心語と共起関係にある付属語を入力音声から抽出するよ
うにしたことにより、共起関係で組み合わされた中心語
と付属語とが一つの連続的な入力音声から認識され、中
心語の抽出結果に基づいて入力音声に照合させる付属語
を絞り込むことができるので、入力音声から付属語を認
識する処理動作の負担を軽減して速度を向上させること
ができ、共起関係にある中心語と付属語とを良好な精度
で高速に認識することができる。In the speech recognition method according to the eleventh aspect, a co-occurrence central word and an adjunct word are set in combination, a continuous input of a speech to be recognized is received, and the continuous input speech is received. From the input speech, the central word and the auxiliary word combined in the co-occurrence relation are one continuous word. It is possible to narrow the attached words that are recognized from the typical input speech and to match with the input speech based on the extraction result of the central word, thereby reducing the burden of processing for recognizing the attached words from the input speech and improving the speed. Thus, the co-occurrence of the central word and the auxiliary word can be quickly recognized with good accuracy.
【0078】請求項12記載の情報記憶媒体は、共起関
係にある複数の語句が組み合わされて格納される認識語
句辞書のソフトウェアと、連続的な入力音声から共起関
係で組み合わされた複数の語句を認識するためのプログ
ラムと、が書き込まれているので、この情報記憶媒体の
ソフトウェアをコンピュータに読み取らせて動作させれ
ば、このコンピュータは、複数の語句を一つの連続的な
入力音声から共起関係の組み合わせに基づいて認識する
ことができるので、複数の語句を良好な精度で高速に認
識することができる。According to a twelfth aspect of the present invention, there is provided an information storage medium, comprising: a recognition phrase dictionary software in which a plurality of words having a co-occurrence relation are stored in combination; Since a program for recognizing words and phrases has been written, if the computer reads and operates the software of the information storage medium, the computer can share a plurality of words from one continuous input voice. Since recognition can be performed based on a combination of occurrence relationships, a plurality of phrases can be recognized at high speed with good accuracy.
【0079】請求項13記載の情報記憶媒体は、共起関
係にある中心語と付属語とが組み合わされて格納される
認識語句辞書のソフトウェアと、中心語を認識語句辞書
から読み出して連続的な入力音声から抽出するためのプ
ログラムと、この抽出された中心語と共起関係にある付
属語を認識語句辞書から読み出して入力音声から抽出す
るためのプログラムと、が書き込まれていることによ
り、この情報記憶媒体のソフトウェアをコンピュータに
読み取らせて動作させれば、このコンピュータは、共起
関係で組み合わされた中心語と付属語とを一つの連続的
な入力音声から認識することができ、中心語の抽出結果
に基づいて入力音声に照合させる付属語を絞り込むこと
ができるので、入力音声から付属語を認識する処理動作
の負担を軽減して速度を向上させることができ、共起関
係にある中心語と付属語とを良好な精度で高速に認識す
ることができる。According to a thirteenth aspect of the present invention, there is provided an information storage medium, comprising: a recognition word dictionary software in which a co-occurring central word and an adjunct word are stored in combination; By writing a program for extracting from the input speech and a program for extracting an auxiliary word having a co-occurrence relation with the extracted central word from the recognition phrase dictionary and extracting from the input speech, If the computer reads and operates the software of the information storage medium, the computer can recognize the central word and the adjunct words combined in a co-occurrence relationship from one continuous input voice, Since the auxiliary words to be matched with the input voice can be narrowed down based on the extraction result of the Can be improved, the center word and the accessory words can be recognized at high speed with good accuracy in the co-occurrence relation.
【図1】本発明の実施の一形態の音声認識装置の論理的
構造を示す模式図である。FIG. 1 is a schematic diagram showing a logical structure of a speech recognition device according to an embodiment of the present invention.
【図2】音声認識装置の物理的構造を示すブロック図で
ある。FIG. 2 is a block diagram showing a physical structure of the speech recognition device.
【図3】音声認識装置の外観を示す斜視図である。FIG. 3 is a perspective view showing an external appearance of the voice recognition device.
【図4】情報記憶媒体であるRAMに書き込まれたソフ
トウェアの論理的構造を示す模式図である。FIG. 4 is a schematic diagram showing a logical structure of software written in a RAM serving as an information storage medium.
【図5】認識語句辞書の記憶内容を示し、(a)は単語
辞書、(b)は共起辞書、(c)は語順辞書、を示す模
式図である。5A and 5B are schematic diagrams showing storage contents of a recognized phrase dictionary, wherein FIG. 5A is a word dictionary, FIG. 5B is a co-occurrence dictionary, and FIG. 5C is a word order dictionary.
【図6】音声認識装置の音声認識方法を示すフローチャ
ートである。FIG. 6 is a flowchart illustrating a speech recognition method of the speech recognition device.
【図7】第一の変形例の認識語句辞書の共起辞書と語順
辞書との記憶内容を示す模式図である。FIG. 7 is a schematic diagram showing storage contents of a co-occurrence dictionary and a word order dictionary of a recognized phrase dictionary according to a first modified example.
【図8】第二の変形例の認識語句辞書の共起辞書の記憶
内容を示す模式図である。FIG. 8 is a schematic diagram showing storage contents of a co-occurrence dictionary of a recognized phrase dictionary according to a second modified example.
【図9】第三の変形例の認識語句辞書の共起辞書の記憶
内容を示す模式図である。FIG. 9 is a schematic diagram showing storage contents of a co-occurrence dictionary of a recognized phrase dictionary of a third modified example.
【図10】第四の変形例の認識語句辞書の共起辞書の記
憶内容を示す模式図である。FIG. 10 is a schematic diagram showing storage contents of a co-occurrence dictionary of a recognized phrase dictionary according to a fourth modified example.
【図11】第五の変形例の認識語句辞書の単語辞書と語
順辞書との記憶内容を示す模式図である。FIG. 11 is a schematic diagram showing storage contents of a word dictionary and a word order dictionary of a recognized phrase dictionary of a fifth modified example.
【図12】第六の変形例の認識語句辞書の語順辞書の記
憶内容を示す模式図である。FIG. 12 is a schematic diagram showing storage contents of a word order dictionary of a recognized phrase dictionary according to a sixth modified example.
1 音声認識装置 2 コンピュータ 4〜7,9 情報記憶媒体 21 音声入力手段 22 認識語句辞書 23 語句認識手段 41,42 ソフトウェア 42 プログラム DESCRIPTION OF SYMBOLS 1 Speech recognition apparatus 2 Computer 4-7, 9 Information storage medium 21 Speech input means 22 Recognized phrase dictionary 23 Phrase recognition means 41, 42 Software 42 Program
─────────────────────────────────────────────────────
────────────────────────────────────────────────── ───
【手続補正書】[Procedure amendment]
【提出日】平成8年9月20日[Submission date] September 20, 1996
【手続補正1】[Procedure amendment 1]
【補正対象書類名】明細書[Document name to be amended] Statement
【補正対象項目名】全文[Correction target item name] Full text
【補正方法】変更[Correction method] Change
【補正内容】[Correction contents]
【書類名】 明細書[Document Name] Statement
【発明の名称】 音声認識装置および方法、情報記憶媒
体Patent application title: Speech recognition apparatus and method, information storage medium
【特許請求の範囲】[Claims]
【発明の詳細な説明】DETAILED DESCRIPTION OF THE INVENTION
【0001】[0001]
【発明の属する技術分野】本発明は、音声を認識する音
声認識装置および方法と、そのプログラム等のソフトウ
ェアが書き込まれた情報記憶媒体に関する。BACKGROUND OF THE INVENTION 1. Field of the Invention The present invention relates to a speech recognition apparatus and method for recognizing speech, and an information storage medium on which software such as a program is written.
【0002】[0002]
【従来の技術】現在、人間が発声した音声を認識する音
声認識装置が要望されており、各種の音声認識方法が考
えられている。人間が語句である単語を一つだけ発声す
る場合、これを音声認識装置が認識することは困難では
ないが、人間の自然な会話では音声は連続しており、そ
こには多数の単語が助詞等を介して含まれている。この
ように連続的な音声から必要な単語を認識する手法とし
てはワードスポッティングが考案されており、これは予
め設定された認識候補の単語を連続的な会話音声から抽
出して認識する。2. Description of the Related Art At present, there is a demand for a voice recognition device for recognizing a voice uttered by a human, and various voice recognition methods are being considered. When a human utters only one word, which is a phrase, it is not difficult for the speech recognizer to recognize it, but in a natural human conversation, speech is continuous, and many words contain particle And so on. As a method for recognizing necessary words from continuous speech, word spotting has been devised. In this method, words of preset recognition candidates are extracted from continuous speech and recognized.
【0003】このようなワードスポッティングを実行す
る音声認識装置は、認識候補の単語毎に読みが格納され
た単語辞書を有しており、連続入力の会話音声と単語辞
書の全部の単語の読みとをマッチングさせ、このマッチ
ングのスコアが基準値を超過した最大の単語を会話音声
から認識する。しかし、単純にワードスポッティングを
実行しても良好な結果は期待できないため、現在では言
語的な制約により認識精度を向上させることが一般的で
ある。A speech recognition apparatus that performs such word spotting has a word dictionary in which readings are stored for each recognition candidate word. And the largest word whose score of the matching exceeds the reference value is recognized from the conversation voice. However, good results cannot be expected even if the word spotting is simply executed, so that it is now general to improve recognition accuracy due to linguistic restrictions.
【0004】このような言語的な制約には、文法等の構
文的な性質に基づくものと、格パターン等の意味的な制
約に基づくものとがある。前者は制約として強力である
が、書き言葉と違い規則化しにくい話し言葉には向いて
おらず、その認識制御は文節内程度にしか利用できな
い。後者は語順が自由な日本語の性質、話し言葉に向い
ており、上述したワードスポッティングに利用すると認
識精度を良好に向上させることができる。Such linguistic constraints include those based on syntactic properties such as grammar, and those based on semantic constraints such as case patterns. The former is powerful as a constraint, but it is not suitable for spoken language that is difficult to regularize unlike written language, and its recognition control can be used only within phrases. The latter is suitable for the characteristics of Japanese and the spoken language in which the word order is free, and when used for the word spotting described above, the recognition accuracy can be improved satisfactorily.
【0005】[0005]
【発明が解決しようとする課題】上述のようにワードス
ポッティングでは、連続音声から必要な語句のみ認識す
ることができ、特に意味的な制約を利用すると認識精度
を向上させることができる。As described above, in word spotting, only necessary words and phrases can be recognized from continuous speech. In particular, recognition accuracy can be improved by using semantic constraints.
【0006】例えば、特開平6-102897号公報に開示され
た音声認識装置の音声認識方法では、格関係を利用して
認識する文節を予測し、絞り込みを行なっている。しか
し、これでは格関係以外の関係に対処することができ
ず、複数の関係を取り扱うこともできない。つまり、連
続音声に含まれる語句の関係は格関係だけではなく、一
つの連続音声に複数の関係が含まれることも一般的であ
る。For example, in the speech recognition method of the speech recognition apparatus disclosed in Japanese Patent Application Laid-Open No. 6-102897, a phrase to be recognized is predicted using case relations, and narrowing is performed. However, this cannot deal with relationships other than case relationships, nor can it handle multiple relationships. In other words, the relationship between words included in the continuous speech is not limited to the case relationship, but it is also common that one continuous speech includes a plurality of relationships.
【0007】例えば、会話音声が「カップラーメンのカ
レー味を一箱ください」である場合、格関係を利用して
も語句を認識することは困難である。また、会話音声が
「泣く子も黙る」である場合、「泣く・子」「子・黙
る」なる複数の関係が存在している。[0007] For example, if the conversation voice is "Please give a box of curry flavor of cup ramen", it is difficult to recognize words even by using case relations. In addition, when the conversation voice is “crying child is also silent”, there are a plurality of relationships of “crying / child” and “child / silent”.
【0008】[0008]
【課題を解決するための手段】請求項1記載の発明の音
声認識装置は、認識対象の音声の連続的な入力を受け付
ける音声入力手段と、共起関係にある複数の語句が組み
合わされて格納された認識語句辞書と、連続的な入力音
声から共起関係で組み合わされた複数の語句を認識する
語句認識手段とを有する。従って、認識語句辞書には共
起関係にある複数の語句が組み合わされて格納されてい
るので、音声入力手段に連続的に入力された認識対象の
音声から共起関係で組み合わされた複数の語句が語句認
識手段により認識される。つまり、所定の語句を共起関
係で組み合わせて設定しておけば、この共起関係にある
複数の語句は、一つの連続的な入力音声から個々に単独
で認識されず、共起関係の組み合わせに基づいて認識さ
れる。According to a first aspect of the present invention, there is provided a voice recognition apparatus for storing a combination of voice input means for receiving a continuous input of a voice to be recognized and a plurality of co-occurring words. And a phrase recognition unit for recognizing a plurality of phrases combined in a co-occurrence relationship from continuous input speech. Therefore, since a plurality of words having a co-occurrence relation are stored in combination in the recognition word dictionary, a plurality of words combined in a co-occurrence relation from the speech to be recognized continuously input to the voice input means are stored. Is recognized by the phrase recognition means. In other words, if predetermined words are combined and set in a co-occurrence relationship, a plurality of words in this co-occurrence relationship cannot be individually recognized from one continuous input voice, Is recognized based on
【0009】請求項2記載の発明では、請求項1記載の
音声認識装置において、語句認識手段は、共起関係で組
み合わされた一対の語句の一方である中心語を認識語句
辞書から読み出して入力音声から抽出してから、この抽
出された中心語と共起関係にある他方の語句である共起
語を前記認識語句辞書から読み出して入力音声から抽出
する。従って、一つの連続的な入力音声から複数の語句
が語句認識手段により認識される場合、この語句の一方
である中心語が最初に入力音声から抽出されてから、こ
の中心語と共起関係にある他方の語句である共起語が次
に入力音声から抽出される。つまり、中心語は従来のワ
ードスポッティングと同様に多数を入力音声に照合させ
ることになるが、共起語は共起関係に基づいて絞り込ま
れてから入力音声に照合させることになる。According to a second aspect of the present invention, in the speech recognition apparatus according to the first aspect, the word / phrase recognizing means reads out and inputs a central word which is one of a pair of words combined in a co-occurrence relation from the recognized word / phrase dictionary. After being extracted from the voice, a co-occurrence word, which is the other word having a co- occurrence relationship with the extracted central word, is read from the recognized phrase dictionary and extracted from the input voice. Therefore, when a plurality of words are recognized from one continuous input voice by the phrase recognition means, a central word which is one of the words is first extracted from the input voice, and then a co-occurrence relation with the central word is obtained. One other phrase, a co-occurrence word, is then extracted from the input speech. In other words, as in the case of conventional word spotting, a large number of central words are collated with the input speech, but co-occurring words are collated with the input speech after being narrowed down based on the co-occurrence relationship.
【0010】請求項3記載の発明では、請求項2記載の
音声認識装置において、語句認識手段は、入力音声の中
心語を抽出した区間を排除した区間から共起語を抽出す
る。従って、一つの連続的な入力音声から共起関係にあ
る中心語と共起語とが語句認識手段により認識される場
合、最初に入力音声の全域から中心語が抽出され、この
中心語が抽出された区間以外の区間から共起語が抽出さ
れる。つまり、中心語は従来のワードスポッティングと
同様に多数を入力音声に照合させることになるが、共起
語は共起関係に基づいて絞り込まれてから中心語と重複
しない音声区間に照合させることになる。According to a third aspect of the present invention, in the speech recognition apparatus according to the second aspect, the word / phrase recognizing means extracts a co-occurrence word from a section excluding a section from which a central word of the input speech is extracted. Therefore, when a co-occurrence central word and a co-occurrence word are recognized from one continuous input voice by the phrase recognition means, the central word is first extracted from the entire input voice, and this central word is extracted. A co-occurrence word is extracted from a section other than the section performed. That is, although the center word will be collated with the input speech a number similar to the conventional word spotting, the co-occurrence <br/> words in speech interval that does not overlap the center Language narrowed down based on the co-occurrence relation It will be collated.
【0011】請求項4記載の発明では、請求項2または
3記載の音声認識装置において、認識語句辞書は、中心
語と共起語との組み合わせに順番の情報も付与されてお
り、語句認識手段は、中心語と共起語とを順番に対応し
て入力音声から認識する。従って、中心語と共起語とが
連続音声に発生する順番の情報も認識語句辞書に格納さ
れており、中心語と共起語とは入力音声に所定の順番で
発生すると語句認識手段により認識されるので、中心語
と共起語とが入力音声から個別に認識されるような場合
でも順番が適正でないと認識されない。According to a fourth aspect of the present invention, in the speech recognition apparatus according to the second or third aspect, the recognition word dictionary is further provided with information on the order of the combination of the central word and the co-occurrence word. Recognizes a central word and a co-occurrence word from the input speech in order. Therefore, information on the order in which the central word and the co-occurring word occur in the continuous speech is also stored in the recognition phrase dictionary. When the central word and the co-occurring word occur in the input speech in a predetermined order, the phrase recognition unit recognizes the central word and the co-occurring word. Therefore, even when the central word and the co-occurring word are individually recognized from the input speech, they are not recognized unless the order is proper.
【0012】請求項5記載の発明では、請求項4記載の
音声認識装置において、認識語句辞書は、中心語と共起
語との組み合わせに中間に位置する付属語も格納されて
おり、語句認識手段は、中心語と付属語と共起語とを入
力音声から認識する。従って、中心語と共起語との組み
合わせに中間に位置する付属語も認識語句辞書に格納さ
れており、中心語と付属語と共起語とが入力音声から語
句認識手段により認識されるので、ある中心語と付属語
との共起語との読みが、他の中心語と共起語との組み合
わせの読みと同一の場合でも、これらが各々別個に認識
される。According to a fifth aspect of the present invention, in the speech recognition apparatus according to the fourth aspect, the recognition word dictionary also stores an auxiliary word located in the middle of the combination of the central word and the co-occurrence word. The word recognition means recognizes a central word, an adjunct word, and a co-occurrence word from the input speech. Therefore, the auxiliary word located in the middle of the combination of the central word and the co-occurrence word is also stored in the recognition phrase dictionary, and the central word, the auxiliary word, and the co-occurrence word are recognized from the input speech by the phrase recognition means. Even if the reading of a co-occurring word of a certain central word and an adjunct word is the same as the reading of a combination of another central word and a co-occurring word, these are recognized separately.
【0013】請求項6記載の発明では、請求項2または
3記載の音声認識装置において、認識語句辞書は、中心
語と共起語との組み合わせに時間間隔の情報も付与され
ており、語句認識手段は、中心語と共起語とを時間間隔
に対応して入力音声から認識する。従って、中心語と共
起語とが連続的な入力音声に発生する時間間隔の情報も
認識語句辞書に格納されており、中心語と共起語とが入
力音声から時間間隔に対応して語句認識手段により認識
されるので、中心語と共起語とは入力音声に適正な時間
間隔で発生した場合のみ認識され、中心語と共起語とが
入力音声から個別に認識されるような場合でも時間間隔
が適正でないと認識されない。According to a sixth aspect of the present invention, in the speech recognition apparatus according to the second or third aspect, the recognition word dictionary is provided with time interval information for a combination of a central word and a co-occurrence word. The means recognizes the central word and the co-occurrence word from the input voice corresponding to the time interval. Therefore, the center language and co
Recognized by electromotive word and information time interval that occurs a continuous input speech is also stored in the recognition word dictionary, the center word and occurrence word and is compatible from the input speech to the time interval phrase recognition means Therefore, the central word and the co-occurring word are recognized only when they occur at an appropriate time interval in the input voice, and the time interval is not appropriate even when the central word and the co-occurring word are individually recognized from the input voice. Is not recognized.
【0014】請求項7記載の発明では、請求項2または
3記載の音声認識装置において、認識語句辞書は、中心
語と共起語との組み合わせが複数段階の階層構造として
格納されており、語句認識手段は、一つの中心語と複数
の共起語とを階層構造に対応して入力音声から段階的に
認識する。従って、中心語と共起語との組み合わせが複
数段階の階層構造として認識語句辞書に格納されてお
り、一つの中心語と複数の共起語とが入力音声から階層
構造に対応して語句認識手段により段階的に認識され
る。つまり、ある入力音声から一つの中心語と一つの共
起語とが認識されると、この共起語を中心語とする他の
共起語も入力音声から検索され、このような処理動作が
順次繰り返されるので、複数段階の共起関係にある一つ
の中心語と複数の共起語とが順次認識される。According to a seventh aspect of the present invention, in the speech recognition apparatus according to the second or third aspect, the recognition word dictionary stores a combination of a central word and a co-occurrence word as a hierarchical structure having a plurality of stages. The recognizing means recognizes one central word and a plurality of co-occurring words stepwise from the input speech in accordance with the hierarchical structure. Therefore, a combination of a central word and a co-occurrence word is stored in the recognition phrase dictionary as a hierarchical structure having a plurality of stages, and one central word and a plurality of co-occurrence words are recognized from the input speech in a hierarchical structure corresponding to the hierarchical structure. Recognized step by step. In other words, one central word and one shared word from a certain input voice
When the electromotive word is recognized, the other centered words the occurrence word
Co-occurrence words are also retrieved from the input speech, and such processing operations are sequentially repeated, so that one central word and a plurality of co-occurrence words having a co- occurrence relationship in a plurality of stages are sequentially recognized.
【0015】請求項8記載の発明では、請求項7記載の
音声認識装置において、認識語句辞書は、一つの中心語
と複数の共起語との組み合わせに順番の情報も付与され
ており、語句認識手段は、一つの中心語と複数の共起語
とを順番に対応して入力音声から認識する。従って、一
つの中心語と複数の共起語との組み合わせの順番の情報
も認識語句辞書に格納されており、一つの中心語と複数
の共起語とが入力音声から順番に対応して語句認識手段
により認識されるので、一つの中心語と複数の共起語と
が入力音声から個別に認識されるような場合でも各々の
順番が適正でないと認識されない。According to the invention described in claim 8, in the speech recognition apparatus according to claim 7, the recognition word dictionary is further provided with order information for a combination of one central word and a plurality of co-occurring words. The recognizing means recognizes one central word and a plurality of co-occurring words in order from the input speech. Therefore, the information on the order of the combination of one central word and a plurality of co-occurring words is also stored in the recognition phrase dictionary, and the one central word and the plurality of co-occurring words correspond to the phrases in order from the input speech. Since it is recognized by the recognition means, even when one central word and a plurality of co-occurring words are individually recognized from the input voice, it is not recognized that their respective orders are not proper.
【0016】請求項9記載の発明では、請求項7または
8記載の音声認識装置において、認識語句辞書は、一つ
の中心語と複数の共起語との組み合わせに階層構造の深
度の情報も付与されており、語句認識手段は、一つの中
心語と複数の共起語とを深度に対応して入力音声から認
識する。従って、一つの中心語と複数の共起語との組み
合わせの階層構造の深度の情報も認識語句辞書に格納さ
れており、一つの中心語と複数の共起語とが入力音声か
ら深度に対応して語句認識手段により認識されるので、
一つの中心語を規定とした複数の共起語の段階的な認識
が所定の深度まで実行される。According to a ninth aspect of the present invention, in the speech recognition apparatus according to the seventh or eighth aspect, the recognition phrase dictionary also adds information of a hierarchical structure depth to a combination of one central word and a plurality of co-occurring words. The word recognition means recognizes one central word and a plurality of co-occurring words from the input speech in accordance with the depth. Therefore, the information on the depth of the hierarchical structure of the combination of one central word and multiple co-occurring words is also stored in the recognition phrase dictionary, and one central word and multiple co-occurring words correspond to the depth from the input speech. And is recognized by the phrase recognition means,
Stepwise recognition of a plurality of co-occurring words defining one central word is executed to a predetermined depth.
【0017】請求項10記載の音声認識方法は、共起関
係にある複数の語句を組み合わせて設定しておき、認識
対象の音声の連続的な入力を受け付け、この連続的な入
力音声から共起関係で組み合わされた複数の語句を認識
するようにした。従って、所定の語句を共起関係で組み
合わせて設定しておけば、この共起関係にある複数の語
句は、一つの連続的な入力音声から個々に単独で認識さ
れず、共起関係の組み合わせに基づいて認識される。According to a tenth aspect of the present invention, a plurality of words having a co-occurrence relation are set in combination, a continuous input of a speech to be recognized is received, and co-occurrence is performed from the continuous input speech. Recognize multiple words combined in relationships. Therefore, if predetermined words are combined and set in a co-occurrence relationship, a plurality of words in this co-occurrence relationship cannot be individually recognized from one continuous input voice, and Is recognized based on
【0018】請求項11記載の音声認識方法は、共起関
係にある中心語と共起語とを組み合わせて設定してお
き、認識対象の音声の連続的な入力を受け付け、この連
続的な入力音声から用意された中心語を抽出し、この中
心語と共起関係にある共起語を入力音声から抽出するよ
うにした。従って、一つの連続的な入力音声から共起関
係の中心語と共起語とが認識される場合、最初に入力音
声から中心語が抽出されてから、この中心語と共起関係
にある共起語が次に入力音声から抽出される。つまり、
中心語は従来のワードスポッティングと同様に多数を入
力音声に照合させることになるが、共起語は共起関係に
基づいて絞り込まれてから入力音声に照合させることに
なる。According to the speech recognition method of the present invention, a co-occurrence central word and a co-occurrence word are set in combination, and a continuous input of a speech to be recognized is received. A prepared central word is extracted from the voice, and co- occurring words having a co- occurrence relationship with the central word are extracted from the input voice. Therefore, co-located if one from a continuous input speech centered word cooccurrence and occurrence word is recognized, the center word is extracted from the first input speech, the co-occurrence relationship with the central word The spoken word is then extracted from the input speech. That is,
As with the central word, a large number of words are collated with the input speech, as in the conventional word spotting, but the co-occurring words are collated with the input speech after being narrowed down based on the co-occurrence relationship.
【0019】請求項12記載の情報記憶媒体は、コンピ
ュータが読取自在なソフトウェアが予め書き込まれた情
報記憶媒体において、共起関係にある複数の語句が組み
合わされて格納される認識語句辞書のソフトウェアと、
連続的な入力音声から共起関係で組み合わされた複数の
語句を認識するためのプログラムと、が書き込まれてい
る。従って、この情報記憶媒体のソフトウェアをコンピ
ュータに読み取らせて動作させれば、このコンピュータ
は、認識語句辞書に格納されている語句を連続的な入力
音声から認識する音声認識装置として機能する。このと
き、認識語句辞書には共起関係にある複数の語句が組み
合わされて格納されているので、連続的な入力音声から
共起関係で組み合わされた複数の語句が認識される。つ
まり、所定の語句を共起関係で組み合わせて設定してお
けば、この共起関係にある複数の語句は、一つの連続的
な入力音声から個々に単独で認識されず、共起関係の組
み合わせに基づいて認識される。According to a twelfth aspect of the present invention, there is provided an information storage medium in which a plurality of words having a co-occurrence relation are combined and stored in an information storage medium in which computer-readable software is written in advance. ,
And a program for recognizing a plurality of phrases combined in a co-occurrence relationship from continuous input speech. Therefore, if the computer reads and operates the software of the information storage medium, the computer functions as a speech recognition device that recognizes words stored in the recognized word dictionary from continuous input speech. At this time, since a plurality of words having a co-occurrence relationship are stored in combination in the recognition word dictionary, a plurality of words combined in a co-occurrence relationship are recognized from continuous input speech. In other words, if predetermined words are combined and set in a co-occurrence relationship, a plurality of words in this co-occurrence relationship cannot be individually recognized from one continuous input voice, Is recognized based on
【0020】請求項13記載の情報記憶媒体は、コンピ
ュータが読取自在なソフトウェアが予め書き込まれた情
報記憶媒体において、共起関係にある中心語と共起語と
が組み合わされて格納される認識語句辞書のソフトウェ
アと、中心語を前記認識語句辞書から読み出して連続的
な入力音声から抽出するためのプログラムと、この抽出
された中心語と共起関係にある共起語を前記認識語句辞
書から読み出して入力音声から抽出するためのプログラ
ムと、が書き込まれている。従って、この情報記憶媒体
のソフトウェアをコンピュータに読み取らせて動作させ
れば、このコンピュータは、認識語句辞書に格納されて
いる語句を連続的な入力音声から認識する音声認識装置
として機能する。このとき、一つの連続的な入力音声か
ら複数の語句が語句認識手段により認識される場合、こ
の語句の一方である中心語が最初に入力音声から抽出さ
れてから、この中心語と共起関係にある他方の語句であ
る共起語が次に入力音声から抽出される。つまり、中心
語は従来のワードスポッティングと同様に多数を入力音
声に照合させることになるが、共起語は共起関係に基づ
いて絞り込まれてから入力音声に照合させることにな
る。According to a thirteenth aspect of the present invention, there is provided an information storage medium in which a computer readable software is written in advance, and a recognition word stored in combination with a co-occurring central word and a co-occurrence word. Dictionary software, a program for reading a central word from the recognized phrase dictionary and extracting it from continuous input speech, and reading a co-occurring word having a co- occurrence relationship with the extracted central word from the recognized phrase dictionary And a program for extracting from the input voice. Therefore, if the computer reads and operates the software of the information storage medium, the computer functions as a speech recognition device that recognizes words stored in the recognized word dictionary from continuous input speech. At this time, when a plurality of words are recognized from one continuous input voice by the phrase recognition means, a central word which is one of the words is first extracted from the input voice, and then a co-occurrence relation with the central word is obtained. occurrence word is the other word in is extracted from the next input speech. In other words, as in the case of conventional word spotting, a large number of central words are collated with the input speech, but co-occurring words are collated with the input speech after being narrowed down based on the co-occurrence relationship.
【0021】[0021]
【発明の実施の形態】本発明の実施の一形態を図面に基
づいて以下に説明する。まず、本実施の形態の音声認識
装置1は、図2および図3に示すように、そのハードウ
ェアとしてデータ処理装置であるコンピュータシステム
を有している。このコンピュータシステムからなる音声
認識装置1は、コンピュータの主体としてCPU(Centr
al Processing Unit)2を有しており、このCPU2に
は、バスライン3により、ROM(Read Only Memory)
4、RAM(Random Access Memory)5、HD(Hard Disk
…図示せず)を内蔵したHDD(HD Drive)6、FD(Flo
ppy Disk)7が装填されるFDD(FD Drive)8、CD(C
ompact Disk)−ROM9が装填されるCD−ROMドラ
イブ10、マウス11が接続されたキーボード12、デ
ィスプレイ13、入力デバイスであるマイクロフォン1
4、通信I/F(Interface)15、等が接続されてい
る。DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS One embodiment of the present invention will be described below with reference to the drawings. First, as shown in FIGS. 2 and 3, the voice recognition device 1 of the present embodiment has a computer system as a data processing device as hardware. The speech recognition device 1 including the computer system includes a CPU (Centr
al Processing Unit) 2, and the CPU 2 is connected to a ROM (Read Only Memory) by a bus line 3.
4, RAM (Random Access Memory) 5, HD (Hard Disk)
... HDD (HD Drive) 6 with built-in
ppy Disk) 7, FDD (FD Drive) 8, CD (C
ompact Disk) -CD-ROM drive 10 loaded with ROM 9, keyboard 12 connected to mouse 11, display 13, microphone 1 as input device
4, communication I / F (Interface) 15, etc. are connected.
【0022】この音声認識装置1は、前記CPU2に各
種の処理動作を実行させるプログラムが予め設定されて
おり、このプログラム等のソフトウェアは、例えば、情
報記憶媒体である前記FD7や前記CD−ROM9に予
め書き込まれている。そして、このソフトウェアが情報
記憶媒体である前記HDD6にインストールされてお
り、これが起動時に情報記憶媒体である前記RAM5に
複写されて前記CPU2に読み取られる。In the speech recognition apparatus 1, a program for causing the CPU 2 to execute various processing operations is set in advance, and software such as this program is stored in, for example, the FD 7 or the CD-ROM 9 which is an information storage medium. It has been written in advance. This software is installed in the HDD 6 as an information storage medium, and is copied to the RAM 5 as an information storage medium and read by the CPU 2 at startup.
【0023】このようにソフトウェアを前記CPU2が
読み取って各種の処理動作を実行することにより、各種
機能が各種手段として実現されている。このような各種
手段として、本実施の形態の音声認識装置1は、図1に
示すように、音声入力手段21、認識語句辞書22、語
句認識手段23、結果出力手段24、等を備えている。
前記認識語句辞書22は、単語辞書25、共起辞書2
6、語順辞書27、からなり、前記語句認識手段23
は、特徴抽出手段28、候補認識手段29、候補読出手
段30、語句探索手段31、結果確定手段32、等から
なる。As described above, various functions are realized as various means by the CPU 2 reading the software and executing various processing operations. As such various means, the voice recognition device 1 of the present embodiment includes a voice input unit 21, a recognized phrase dictionary 22, a phrase recognition unit 23, a result output unit 24, and the like, as shown in FIG. .
The recognition phrase dictionary 22 includes a word dictionary 25, a co-occurrence dictionary 2
6, word order dictionary 27, said word recognition means 23
Consists of a feature extracting unit 28, a candidate recognizing unit 29, a candidate reading unit 30, a phrase searching unit 31, a result determining unit 32, and the like.
【0024】このような音声認識装置1の各種手段は、
必要により前記ディスプレイ13や前記マイクロフォン
14等のハードウェアも利用して実現されるが、その主
体は前記RAM5等に書き込まれたソフトウェアに対応
して前記CPU2が動作することにより実現されてい
る。このため、前記RAM5には、図4に示すように、
前記認識語句辞書22のソフトウェアである辞書記憶部
41と、連続的な入力音声から共起関係で組み合わされ
た複数の語句を認識するための制御プログラム42と、
が書き込まれている。Various means of such a speech recognition device 1 include:
If necessary, hardware such as the display 13 and the microphone 14 is used, and the main component is realized by the operation of the CPU 2 corresponding to software written in the RAM 5 or the like. For this reason, as shown in FIG.
A dictionary storage unit 41 which is software of the recognition phrase dictionary 22, a control program 42 for recognizing a plurality of phrases combined in a co-occurrence relationship from continuous input speech,
Is written.
【0025】より詳細には、前記辞書記憶部41は、単
語辞書記憶部43、共起辞書記憶部44、語順辞書記憶
部45からなり、これらの記憶部43〜45に前記辞書
25〜27の各種情報が格納されている。前記制御プロ
グラム42は、モジュール構造で形成されており、特徴
抽出モジュール46、単語照合モジュール47、認識制
御モジュール48等を有している。More specifically, the dictionary storage unit 41 comprises a word dictionary storage unit 43, a co-occurrence dictionary storage unit 44, and a word order dictionary storage unit 45. These storage units 43 to 45 store the dictionaries 25 to 27 Various information is stored. The control program 42 is formed in a module structure and includes a feature extraction module 46, a word matching module 47, a recognition control module 48, and the like.
【0026】前記特徴抽出モジュール46は、連続的な
入力音声のデジタル信号をフレーム毎に分析して特徴量
を抽出するためのプログラムからなり、前記単語照合モ
ジュール47は、入力音声の特徴量に認識候補の単語の
読みである特徴量を照合させて累積類似度をスコアとし
て算出するためのプログラムからなる。前記認識制御モ
ジュール48は、前記辞書25〜27から各種情報を読
み出して前記単語照合モジュール47に伝送し、この単
語照合モジュール47が出力するスコアに基づいて認識
結果の単語を確定するためのプログラムからなる。The feature extraction module 46 comprises a program for analyzing a digital signal of a continuous input speech for each frame and extracting a feature, and the word matching module 47 recognizes the feature of the input speech. The program comprises a program for collating feature amounts which are readings of candidate words and calculating a cumulative similarity as a score. The recognition control module 48 reads various types of information from the dictionaries 25 to 27 and transmits the information to the word matching module 47. From the program for determining the word of the recognition result based on the score output by the word matching module 47, Become.
【0027】前記認識語句辞書22は、上述のように前
記RAM5の辞書記憶部41にデータファイルとして格
納されており、図5に示すように、その一部である前記
単語辞書25には、語句である複数の単語毎に意味と読
みとが登録されている。前記共起辞書26には、共起関
係にある複数の語句が、ここでは一つの中心語と複数の
共起語との組み合わせで格納されている。前記語順辞書
27には、中心語と共起語の意味の情報とが順番に設定
されているので、中心語と共起語との順番の情報が格納
されている。The recognition phrase dictionary 22 is stored as a data file in the dictionary storage unit 41 of the RAM 5 as described above, and as shown in FIG. Are registered for each of a plurality of words. In the co-occurrence dictionary 26, a plurality of words having a co-occurrence relationship are stored in the co-occurrence dictionary.
Stored in combination with co-occurrence words. The word order dictionary 27 stores information on the order of the central word and the co-occurring words because the central word and the information on the meaning of the co-occurring words are set in order.
【0028】なお、ここでは説明を簡略化するために前
記辞書25〜27の各種情報を日本語の文字として表現
しているが、実際のソフトウェアでは単語等は識別コー
ドからなり、読みは単語の音声の特徴量として設定され
ている。この読みの音声の特徴量は、例えば、音素単位
の状態遷移モデルと単語単位の音素ネットワークとして
設定されており、各状態には平均特徴量と継続時間長と
の情報が設定されている。Although the various information in the dictionaries 25 to 27 are expressed as Japanese characters for the sake of simplicity of description, words and the like are composed of identification codes in actual software, and the reading is performed in words. It is set as a feature value of the voice. The feature amount of the reading voice is set, for example, as a state transition model in phoneme units and a phoneme network in word units, and information on the average feature amount and duration is set in each state.
【0029】前記音声入力手段21は、前記マイクロフ
ォン14等により音声の連続的な入力を受け付け、この
入力音声をデジタルの電気信号にA/D(Analog/Digi
tal)変換する。前記語句認識手段23は、前記CPU2
が前記RAM5に格納された前記制御プログラム42を
読み取って対応する処理動作を実行することにより、連
続的な入力音声のデジタル信号から共起関係にある中心
語と共起語とを認識する。The voice input means 21 receives continuous voice input from the microphone 14 or the like, and converts the input voice into a digital electric signal by A / D (Analog / Digilog).
tal) The phrase recognizing means 23 includes the CPU 2
Reads the control program 42 stored in the RAM 5 and executes a corresponding processing operation, thereby recognizing a co-occurrence center word and a co-occurrence word from a continuous input voice digital signal.
【0030】より詳細には、前記語句認識手段23の特
徴抽出手段28は、前記CPU2が前記制御プログラム
42の前記特徴抽出モジュール46を読み取って対応す
る演算処理を実行することにより、連続的な入力音声の
デジタル信号を単位時間であるフレーム毎に分析し、例
えば、LPC(Linear Predictive Coding)メルケプスト
ラムの算出により特徴量を抽出する。More specifically, the feature extracting means 28 of the phrase recognizing means 23 reads the feature extracting module 46 of the control program 42 and executes the corresponding arithmetic processing to execute the continuous input. The audio digital signal is analyzed for each frame, which is a unit time, and a feature amount is extracted by calculating, for example, an LPC (Linear Predictive Coding) mel-cepstral.
【0031】前記候補認識手段29は、前記認識制御モ
ジュール48に対応した前記CPU2の演算処理によ
り、前記認識語句辞書22の共起辞書26から全部の中
心語を読み出してから、その各々の読みを前記単語辞書
25から読み出す。さらに、前記単語照合モジュール4
7に対応した前記CPU2の演算処理により、全部の中
心語の読みを連続的な入力音声の特徴量に照合させ、そ
の各々の類似度をフレーム単位で算出して順次累積し、
この累積類似度であるスコアが最大で基準値を超過した
一つの中心語を認識候補として抽出する。The candidate recognizing means 29 reads out all the central words from the co-occurrence dictionary 26 of the recognized phrase dictionary 22 by the arithmetic processing of the CPU 2 corresponding to the recognition control module 48, and then reads each of them. Read from the word dictionary 25. Further, the word matching module 4
7, the reading of all the central words is compared with the feature amount of the continuous input voice, the similarity of each of them is calculated for each frame, and sequentially accumulated,
One central word whose score as the accumulated similarity exceeds the reference value at the maximum is extracted as a recognition candidate.
【0032】前記候補読出手段30は、前記認識制御モ
ジュール48に対応した前記CPU2の演算処理によ
り、認識候補の中心語と共起関係にある全部の共起語を
認識候補として前記認識語句辞書22の共起辞書26か
ら検出し、前記語句探索手段31は、前記単語照合モジ
ュール47に対応した前記CPU2の演算処理により、
認識候補の全部の共起語を入力音声に照合させる。The candidate reading means 30 uses the CPU 2 corresponding to the recognition control module 48 to perform all the co-occurrence words having a co- occurrence relationship with the central word of the recognition candidate as recognition candidates by the CPU 2 corresponding to the recognition control module 48. From the co-occurrence dictionary 26, the word / phrase search means 31 performs a calculation process of the CPU 2 corresponding to the word matching module 47,
All the co-occurring words of the recognition candidates are collated with the input speech.
【0033】この処理動作も中心語の場合と同様に、認
識候補の共起語の読みが前記認識語句辞書22の単語辞
書25から読み出され、これと入力音声との照合のスコ
アが最大で基準値を超過した一つの共起語が抽出され
る。ただし、このように共起語の読みを入力音声の特徴
量に照合させる際、前記認識制御モジュール48に対応
した前記CPU2の演算処理により、前記認識語句辞書
22の語順辞書27から中心語と共起語との順番の情報
が読み出され、その順番に対応して入力音声の中心語が
抽出された区間以外の区間のみに共起語の読みが照合さ
れる。In this processing operation, similarly to the case of the central word, the reading of the co-occurrence word of the recognition candidate is read from the word dictionary 25 of the recognized phrase dictionary 22, and the score of the collation between the word and the input speech is maximum. One co-occurrence word exceeding the reference value is extracted. However, when the reading of the co-occurrence word is compared with the feature amount of the input voice in this way, the CPU 2 corresponding to the recognition control module 48 performs the co- processing with the central word from the word order dictionary 27 of the recognition word dictionary 22. information order with raised words are read out, occurrence word reading is matched only in a section other than the central language of the input speech corresponding to the order is extracted section.
【0034】前記結果確定手段32は、上述のように認
識候補の共起語が入力音声から探索されると、前記認識
制御モジュール48に対応した前記CPU2の演算処理
により、認識候補の中心語と共起語とを認識結果として
確定する。前記結果出力手段24は、上述のように確定
された認識結果の中心語と共起語とを、例えば、前記デ
ィスプレイ13の文字表示により出力する。When the co-occurrence word of the recognition candidate is searched from the input voice as described above, the result determination means 32 calculates the central word of the recognition candidate by the arithmetic processing of the CPU 2 corresponding to the recognition control module 48. A co-occurrence word is determined as a recognition result. The result output unit 24 outputs the central word and the co-occurrence word of the recognition result determined as described above, for example, by character display on the display 13.
【0035】このような構成において、本実施の形態の
音声認識装置1は、多数の単語が連続する会話の音声が
マイクロフォン14に入力されると、この連続的な入力
音声から認識語句辞書22に格納されている単語を認識
し、この認識結果をディスプレイ13に表示出力する。
このような音声認識装置1の音声認識方法を、図6を参
照して以下に順次詳述する。In such a configuration, the speech recognition apparatus 1 according to the present embodiment, when a speech of a conversation in which a large number of words are continuous is input to the microphone 14, converts the continuous input speech into the recognized phrase dictionary 22. The stored words are recognized, and the recognition results are displayed on the display 13.
Such a voice recognition method of the voice recognition device 1 will be sequentially described in detail below with reference to FIG.
【0036】まず、マイクロフォン14に連続的に入力
された音声は、その全域がデジタル信号にA/D変換さ
れ、フレーム毎に特徴量が抽出される。つぎに、認識語
句辞書22から全部の中心語の読みが読み出され、この
全部の読みが一つの入力音声の全域の特徴量と照合され
る。このように全部の中心語に対して照合の累積類似度
がスコアとして算出されると、基準値を超過したスコア
が最大の中心語が認識候補として一つだけ選出される。First, the whole area of the voice continuously input to the microphone 14 is A / D converted into a digital signal, and the characteristic amount is extracted for each frame. Next, the readings of all the central words are read from the recognized phrase dictionary 22, and all the readings are collated with the feature values of the whole input voice. When the cumulative similarity of collation is calculated as a score for all the central words in this way, only one central word having the largest score exceeding the reference value is selected as a recognition candidate.
【0037】このように認識候補の中心語が選出される
と、これと共起関係にある全部の共起語が認識候補とし
て認識語句辞書22から読み出され、上述した中心語の
場合と同様に、全部の共起語の読みと入力音声とが照合
されて基準値を超過したスコアが最大の一つの共起語が
選出される。このとき、認識語句辞書22の語順辞書2
7から中心語と共起語との順番の情報が読み出され、そ
の順番に対応して入力音声の中心語が抽出された区間以
外の区間のみに共起語の読みが照合される。When the central word of the recognition candidate is selected as described above, all co-occurring words having a co- occurrence relation with the selected central word are read out from the recognized phrase dictionary 22 as recognition candidates, and are the same as in the case of the central word described above. Then, the reading of all the co-occurring words and the input voice are collated, and one co-occurring word having the maximum score exceeding the reference value is selected. At this time, the word order dictionary 2 of the recognition phrase dictionary 22
7, information on the order of the central word and the co-occurring word is read, and the reading of the co-occurring word is collated only in a section other than the section in which the central word of the input voice is extracted in accordance with the order.
【0038】このように一つの連続的な入力音声から共
起関係にある中心語と共起語とが抽出されると、これが
認識結果として確定されてディスプレイ13の文字表示
により出力される。なお、最初に全部の中心語の照合の
スコアが基準値を超過しない場合には認識結果は無しと
され、中心語の認識候補が検出された状態で共起語が認
識語句辞書22に格納されていない場合や、全部の共起
語の照合のスコアが基準値を超過しない場合には、認識
結果は中心語のみとされる。When a central word and a co-occurrence word having a co-occurrence relation are extracted from one continuous input voice, they are determined as a recognition result and output by character display on the display 13. If the scores of the collation of all the central words do not exceed the reference value at first, no recognition result is found, and the co-occurrence word is stored in the recognized phrase dictionary 22 in a state where the recognition candidate of the central word is detected. If not, or if the scores of collation of all co-occurring words do not exceed the reference value, the recognition result is only the central word.
【0039】上述した一連の処理動作を図5を参考に具
体的に説明すると、一つの連続的な入力音声が「ご注文
はカップラーメンのカレー味ですね」の場合、これに対
して中心語である“カップラーメン,手焼き煎餅”が照
合され、スコアが高い“カップラーメン”が中心語の認
識候補として抽出される。この中心語“カップラーメ
ン”に共起する共起語として“カレー味,しょう油味”
が認識語句辞書22から検出され、この認識語句辞書2
2には中心語である“カップラーメン”より後方に意味
が“味”である共起語が位置することが規定されている
ので、連続的な入力音声から“カップラーメン”より以
後の区間「のカレー味ですね」が切り出される。この音
声区間のみに対して共起語である“カレー味,しょう油
味”が照合されるので、共起語である“カレー味”が認
識される。The above-described series of processing operations will be described in detail with reference to FIG. 5. If one continuous input voice is "order is curry taste of cup ramen""Cup ramen, hand-baked rice cracker" is collated, and "cup ramen" with a high score is extracted as a candidate for recognition of the central word. "Curry flavor, soy sauce flavor" is a co-occurring word that co- occurs with this central word "cup ramen"
Is detected from the recognition phrase dictionary 22, and the recognition phrase dictionary 2
2 specifies that a co-occurrence word whose meaning is “taste” is located behind the central word “cup ramen”, so that a continuous input voice and a subsequent section “coup ramen” It's curry flavor. " Since the co-occurrence word “curry taste, soy sauce taste” is collated only for this voice section, the co-occurrence word “curry taste” is recognized.
【0040】本実施の形態の音声認識装置1の音声認識
方法では、上述のように一つの連続的な入力音声から共
起関係で組み合わされた中心語と共起語とが認識される
ので、これらの単語を個々に認識する場合より精度が良
好である。特に、中心語は多数を入力音声に照合させる
必要があるが、共起語は認識候補の中心語と共起関係に
あるもののみ入力音声に照合させれば良いので、この処
理負担が軽減されて処理速度が向上している。In the speech recognition method of the speech recognition apparatus 1 according to the present embodiment, as described above, the central word and the co-occurrence word combined in a co-occurrence relation are recognized from one continuous input speech. The accuracy is better than when these words are individually recognized. In particular, the center word, it is necessary to match the number on the input speech, occurrence word since only it is sufficient to match the input speech that the center word co-occurrence relationships recognition candidate, the processing load is reduced Processing speed has been improved.
【0041】しかも、共起語は入力音声から中心語の区
間を排除した区間のみに照合させれば良く、この共起語
を照合させる区間も中心語との順番に基づいて一方に制
限されるので、さらに処理負担が軽減されて処理速度が
向上している。さらに、中心語に対する共起語の順番の
情報は、共起語の種類の情報により設定されており、複
数の共起語を個々に設定していないので、語順辞書27
の記憶容量も軽減されている。[0041] Moreover, occurrence word is limited to one based on it is sufficient only matched to the section which eliminated the central word section, the order of the section is also central word for matching the occurrence word from the input speech Therefore, the processing load is further reduced, and the processing speed is improved. Further, the information on the order of co-occurring words with respect to the central word is set by the information on the type of co-occurring words, and a plurality of co-occurring words are not individually set.
Storage capacity has also been reduced.
【0042】なお、本発明は上記形態に限定されるもの
ではなく、各種の変形を許容する。例えば、本実施の形
態では、最初に入力音声から照合のスコアが最大の中心
語を一つの認識候補として選出し、これと共起関係にあ
る共起語を入力音声に照合させてスコアが最大の一つを
選出することを例示したが、最初にスコアが基準値を超
過した複数の中心語を認識候補として抽出し、これらの
中心語と共起関係にある全部の共起語を入力音声に照合
させて各々のスコアを算出し、中心語と共起語とのスコ
アの合計が最大の組み合わせを認識結果とするようなこ
とも可能である。The present invention is not limited to the above-described embodiment, but allows various modifications. For example, in the present embodiment, first, a central word having the largest matching score is selected from the input speech as one recognition candidate, and a co- occurring word having a co- occurrence relationship with the selected central word is compared with the input speech to obtain the maximum score. Was selected as an example. First, a plurality of central words whose scores exceeded the reference value were extracted as recognition candidates, and all co-occurring words having a co- occurrence relationship with these central words were input speech. , The respective scores are calculated, and the combination in which the sum of the scores of the central word and the co-occurrence word is the largest can be used as the recognition result.
【0043】また、本実施の形態では、共起関係にある
複数の語句を中心語と共起語とに分類しておき、最初に
中心語を入力音声から抽出してから、この結果に基づい
て共起語を入力音声から抽出することを例示した。しか
し、このように共起関係にある複数の語句を中心語や共
起語として分類せず、全部の語句を同時に入力音声に照
合させ、合計のスコアが最大となる共起関係の組み合わ
せの語句を認識結果とするようなことも可能である。In the present embodiment, a plurality of co-occurring terms are classified into a central word and a co-occurring word, and the central word is first extracted from the input speech, and then based on the result. Extracting co-occurring words from input speech has been exemplified. However, the center word and co multiple words in this manner cooccurrence
Not classified as causing word, is collated simultaneously input speech all the words, the score of the sum is also possible, as a recognition result a word or phrase in a combination of co-occurrence relation with a maximum.
【0044】さらに、本実施の形態では、認識語句辞書
22に各種辞書25〜27を用途別に個別に形成するこ
とにより、そのメンテナンスや情報登録を容易とするこ
とを想定したが、このような辞書25〜27を一つに組
み合わせた形態として認識語句辞書22を形成すること
も可能である。Further, in the present embodiment, it is assumed that maintenance and information registration are facilitated by separately forming various dictionaries 25 to 27 in the recognition phrase dictionary 22 for each application. It is also possible to form the recognition phrase dictionary 22 as a form in which 25 to 27 are combined into one.
【0045】また、本実施の形態では、音声認識装置1
をコンピュータシステムによる実験装置として想定し、
入力音声から認識した単語をディスプレイ13に表示す
ることを例示した。しかし、上述のような音声認識装置
1の各部をASIC(Application Specific Integrated
Circuit)として製作し、これを各種製品に組み込んで
音声制御に利用することも可能である。In this embodiment, the speech recognition device 1
As an experimental device using a computer system,
The display of the word recognized from the input voice on the display 13 has been exemplified. However, each part of the speech recognition device 1 as described above is integrated with an ASIC (Application Specific Integrated
It is also possible to manufacture it as a circuit and incorporate it into various products and use it for voice control.
【0046】さらに、本実施の形態では、RAM5等に
ソフトウェアとして格納されている制御プログラムに従
ってCPU2が動作することにより、音声認識装置1の
各部が実現されることを例示した。しかし、このような
各部の各々を固有のハードウェアとして製作することも
可能であり、一部をソフトウェアとしてRAM5等に格
納するとともに一部をハードウェアとして製作すること
も可能である。また、所定のソフトウェアが格納された
RAM5等や各部のハードウェアを、例えば、ファーム
ウェアとして製作することも可能である。Further, in the present embodiment, it has been exemplified that each unit of the speech recognition apparatus 1 is realized by the operation of the CPU 2 according to a control program stored as software in the RAM 5 or the like. However, it is also possible to manufacture each of these units as unique hardware, and it is also possible to store a part of the unit as software in the RAM 5 or the like and manufacture a part of the unit as hardware. Further, the RAM 5 or the like in which predetermined software is stored and hardware of each unit can be manufactured as firmware, for example.
【0047】また、本実施の形態では、音声認識装置1
の起動時に、HDD6に格納されているソフトウェアが
RAM5に複写され、このようにRAM5に格納された
ソフトウェアをCPU2が読み取ることを想定したが、
このようなソフトウェアをHDD6に格納したままCP
U2に利用させることや、RAM5に予め書き込んでお
くことも可能である。In the present embodiment, the speech recognition device 1
It is assumed that the software stored in the HDD 6 is copied to the RAM 5 at the time of startup, and the software stored in the RAM 5 is read by the CPU 2 as described above.
With such software stored in the HDD 6, the CP
It is also possible for U2 to use it or to write it in RAM5 in advance.
【0048】さらに、前述のように単体で取り扱える情
報記憶媒体であるFD7やCD−ROM9にソフトウェ
アを書き込んでおき、このFD7等からRAM5等にソ
フトウェアをインストールすることも可能であるが、こ
のようなインストールを実行することなくFD7等に書
き込まれたソフトウェアをCPU2が適宜読み取ってデ
ータ処理を実行することも可能である。Further, as described above, software can be written in the FD 7 or CD-ROM 9 which is an information storage medium that can be handled alone, and the software can be installed in the RAM 5 or the like from the FD 7 or the like. It is also possible for the CPU 2 to appropriately read software written in the FD 7 or the like without executing the installation and execute data processing.
【0049】また、このような音声認識装置1の各部を
実現する制御プログラムを、複数のソフトウェアの組み
合わせにより実現することも可能であり、その場合、単
体の製品となる情報記憶媒体には必要最小限のソフトウ
ェアのみを格納しておけば良い。例えば、オペレーティ
ングシステムが実装されている音声認識装置1に、CD
−ROM9等の情報記憶媒体によりアプリケーションソ
フトを提供するような場合、音声認識装置1の各部を実
現するソフトウェアは、アプリケーションソフトとオペ
レーティングシステムとの組み合わせで実現されるの
で、オペレーティングシステムに依存する部分のソフト
ウェアはアプリケーションソフトの情報記憶媒体から省
略することができる。It is also possible to realize a control program for realizing each part of the voice recognition apparatus 1 by a combination of a plurality of softwares. In this case, the information storage medium as a single product has a minimum required size. It is only necessary to store the limited software. For example, the voice recognition device 1 on which an operating system is mounted has a CD
In a case where application software is provided by an information storage medium such as the ROM 9, software that realizes each unit of the voice recognition device 1 is realized by a combination of the application software and the operating system. The software can be omitted from the information storage medium of the application software.
【0050】特に、本発明の音声認識装置1を、認識す
る単語が特定された業務用の装置等として製作する場合
は、その製造工程で認識語句辞書22の内容も固定的に
書き込めば良い。しかし、上述のように音声認識装置1
のアプリケーションソフトを一般ユーザに販売するよう
な場合には、認識語句辞書22の内容をユーザが自由に
登録できることが好ましい。In particular, when the speech recognition device 1 of the present invention is manufactured as a business device or the like in which words to be recognized are specified, the contents of the recognition phrase dictionary 22 may be fixedly written in the manufacturing process. However, as described above, the speech recognition device 1
It is preferable that the user can freely register the contents of the recognized phrase dictionary 22 when the application software is sold to general users.
【0051】このような製品としてCD−ROM9等の
情報記憶媒体を製造する場合には、前述した制御プログ
ラム42の他、認識語句辞書22をRAM5等に所定の
フォーマットで形成するためのプログラムと、認識語句
辞書22に各種情報を登録させるためのプログラムと
を、情報記憶媒体に書き込んでおくことになる。この場
合、これらのプログラムが情報記憶媒体における認識語
句辞書22のソフトウェアとなり、各種情報の設定澄み
の認識語句辞書22のソフトウェアは情報記憶媒体には
書き込まない。When an information storage medium such as the CD-ROM 9 is manufactured as such a product, in addition to the control program 42 described above, a program for forming the recognition phrase dictionary 22 in the RAM 5 or the like in a predetermined format includes: A program for registering various kinds of information in the recognition phrase dictionary 22 is written in the information storage medium. In this case, these programs become the software of the recognition phrase dictionary 22 in the information storage medium, and the software of the recognition phrase dictionary 22 with the setting of various information is not written in the information storage medium.
【0052】同様に、完成した製品として音声認識装置
1を製造する場合も、単語を認識する各種手段21,2
3,24等の部分は固定的に製作しておき、その認識語
句辞書22の設定内容を空白としてユーザに登録させる
ことも可能である。さらに、このような音声認識装置1
に交換自在に装着するオプション部品として、業務毎に
適正な単語を登録した認識語句辞書22を情報記憶媒体
として製作するようなことも可能である。Similarly, when manufacturing the speech recognition apparatus 1 as a completed product, various means 21 and 21 for recognizing words are used.
It is also possible to make the parts such as 3, 24, etc. fixedly, and let the user register the setting contents of the recognized phrase dictionary 22 as blank. Furthermore, such a speech recognition device 1
It is also possible to produce a recognition phrase dictionary 22 in which appropriate words are registered for each job as an information storage medium, as an optional component that can be exchangeably mounted on a computer.
【0053】なお、上述のように情報記憶媒体に書き込
んだソフトウェアをコンピュータに供給する手法は、そ
の情報記憶媒体をコンピュータに直接に装填することに
限定されない。例えば、上述のようなソフトウェアをホ
ストコンピュータの情報記憶媒体に書き込み、このホス
トコンピュータを通信ネットワークにより端末コンピュ
ータに接続し、ホストコンピュータからデータ通信によ
り端末コンピュータにソフトウェアを供給することも可
能である。The method for supplying the software written in the information storage medium to the computer as described above is not limited to loading the information storage medium directly into the computer. For example, it is also possible to write the above-mentioned software on an information storage medium of a host computer, connect the host computer to a terminal computer via a communication network, and supply the software to the terminal computer by data communication from the host computer.
【0054】この場合、端末コンピュータが自身の情報
記憶媒体にソフトウェアをダウンロードした状態でスタ
ンドアロンのデータ処理を実行することも可能である
が、ソフトウェアをダウンロードすることなくホストコ
ンピュータとのリアルタイムのデータ通信によりデータ
処理を実行することも可能である。この場合、ホストコ
ンピュータと端末コンピュータとを通信ネットワークに
より接続したシステム全体が、本発明の音声認識装置1
に相当することになる。In this case, it is possible for the terminal computer to execute stand-alone data processing in a state where the software has been downloaded to its own information storage medium, but it is possible to perform real-time data communication with the host computer without downloading the software. It is also possible to perform data processing. In this case, the entire system in which the host computer and the terminal computer are connected by the communication network is the voice recognition device 1 of the present invention.
Would be equivalent to
【0055】また、本実施の形態では、単語辞書25に
単語の意味を格納しておき、語順辞書27には共起語を
意味の情報として格納しておくことを例示したが、図7
に示すように、中心語に対する共起語の共起関係の種類
の情報を共起辞書26と語順辞書27とに格納しておく
ことも可能である。この場合、一つの中心語に複数の共
起関係で複数の共起語が対応しても、中心語に対する複
数の共起語の位置を共起関係の種類の情報で設定できる
ので、複数種類の共起関係の共起語を良好な精度で容易
に認識することができる。In this embodiment, the word dictionary 25 stores the meanings of words, and the word order dictionary 27 stores co-occurrence words as meaning information.
As shown in (1), it is also possible to store information on the type of co-occurrence relation of a co- occurrence word with the central word in the co- occurrence dictionary 26 and the word order dictionary 27. In this case, even if a plurality of co-occurring words correspond to one central word in a plurality of co-occurring relations, the positions of the plurality of co-occurring words with respect to the central word can be set by the information of the type of the co-occurring relation. the co-occurrence relationship cooccurrence words can be easily recognized with good accuracy.
【0056】例えば、「ご注文はカレー味のカップラー
メンですね」なる入力音声から“カップラーメン”が中
心語の認識候補として抽出された場合、この中心語“カ
ップラーメン”に共起する共起語としては“カレー味,
ミニ”が認識語句辞書22から検出される。しかし、こ
こでは種類が“味”の共起語は中心語より前方に位置す
ることが規定されており、種類が“サイズ”の共起語は
中心語より後方に位置することが規定されているので、
連続的な入力音声から「ご注文はカレー味の」の区間が
切り出されて“カレー味”の共起語が照合され、「です
ね」の音声区間が切り出されて“ミニ”の共起語が照合
される。For example, if “cup ramen” is extracted as a candidate for recognition of a central word from an input voice “Your order is curry flavored cup ramen,” a co- occurrence co- occurring with this central word “cup ramen” The word is "curry taste,
Mini "is detected from the recognition word dictionary 22. However, where the kind of" taste "co-occurrence word is defined to be located forward of the central words, the type is" occurrence word size "is Since it is specified that it is located after the central word,
From the continuous input voice, the section "Order is curry taste" is cut out and the co-occurrence word "curry taste" is collated, and the voice section of "is" is cut out and co-occurrence word "mini" Are collated.
【0057】また、図8に示すように、認識語句辞書2
2の共起辞書26に、中心語と共起語との組み合わせに
中間に位置する付属語も格納しておき、語句認識手段2
3が、中心語と付属語と共起語とを入力音声から認識す
ることも可能である。この場合、中心語と付属語と共起
語とが一つの入力音声から認識されるので、ある中心語
と付属語との共起語との読みが、他の中心語と共起語と
の組み合わせの読みと同一の場合でも、これらを各々別
個に認識することができる。Also, as shown in FIG.
In the co-occurrence dictionary 26 of FIG. 2, an auxiliary word located in the middle of the combination of the central word and the co-occurrence word is also stored.
3 can also recognize the central word, the adjunct word, and the co-occurrence word from the input speech. In this case, since the central word, the adjunct word and the co-occurrence word are recognized from one input voice, the reading of the co-occurrence word of a certain central word and the adjunct word is not co- occurred with other central words . even if the same reading in combination with electromotive words, it is possible to recognize them each separately.
【0058】例えば、入力音声が「ご注文は手焼き煎餅
の緑茶風味ですね」の場合、中心語である“手焼き煎
餅”に対して共起語である“のり,緑茶風味”の両方が
「の緑茶風味」の音声区間から同等のスコアで認識され
ることになる。しかし、上述のように付属語として
“の”が規定されていれば、共起語として“緑茶風味”
のみを認識することができる。For example, if the input voice is “order is hand-baked rice cracker green tea flavor”, both of the co-occurrence words “Nori and green tea flavor” are used for the central word “hand-baked rice cracker”. Recognition is performed with the same score from the voice section of “no green tea flavor”. However, if it is stipulated "in" as an accessory words, as described above, "green tea flavor" as the co-occurrence word
Can only recognize.
【0059】また、図9に示すように、認識語句辞書2
2の共起辞書26に、中心語と共起語との組み合わせと
ともに時間間隔の情報も格納しておき、語句認識手段2
3が、中心語と共起語とを時間間隔に対応して入力音声
から認識することも可能である。この場合、共起語を照
合させる入力音声の区間を時間間隔に対応して制限でき
るので、より高速に共起語を認識することができ、中心
語から極度に離反した共起語は認識されないので、不適
な共起語の認識を防止することもできる。Further, as shown in FIG.
In the co-occurrence dictionary 26, information on the time interval is stored together with the combination of the central word and the co-occurrence word.
3 can also recognize the central word and the co-occurrence word from the input voice corresponding to the time interval. In this case, since the section of the input speech to match the occurrence word can be limited to correspond to the time intervals, it can recognize the occurrence word faster, occurrence word was extremely away from the center word is not recognized Therefore, it is possible to prevent recognition of inappropriate co-occurrence words.
【0060】例えば、「ご注文はカップラーメンのカレ
ー味を二箱ですね」なる入力音声から“カップラーメ
ン”が中心語の認識候補として抽出された場合、その音
声区間から20フレーム以内の音声区間のみに意味が
“味”の共起語が照合されて“カレー味”が認識され、
100フレーム以内の音声区間のみ意味が“数量”の共起
語が照合されて“二箱”が認識される。For example, if “cup ramen” is extracted as a candidate for recognition of a central word from an input voice “Your order is two curry flavors of cup ramen,” a voice section within 20 frames from that voice section Only the co-occurrence word with the meaning “taste” is compared and “curry taste” is recognized,
Only co-occurrence words having a meaning of “quantity” are collated only in a voice section within 100 frames, and “two boxes” are recognized.
【0061】また、図10に示すように、認識語句辞書
22の共起辞書26に、中心語と共起語との組み合わせ
を複数段階の階層構造として格納しておき、語句認識手
段23が、一つの中心語と複数の共起語とを階層構造に
対応して入力音声から段階的に認識することも可能であ
る。この場合、ある入力音声から一つの中心語と一つの
共起語とが認識されると、この共起語を中心語とする他
の共起語も入力音声から検索され、このような処理動作
が順次繰り返されるので、複数段階の共起関係にある一
つの中心語と複数の共起語とを順次認識することができ
る。As shown in FIG. 10, a combination of a central word and a co-occurrence word is stored in the co-occurrence dictionary 26 of the recognized word dictionary 22 as a hierarchical structure having a plurality of stages. It is also possible to recognize one central word and a plurality of co-occurring words stepwise from the input speech in a hierarchical structure. In this case, one central word and one
When occurrence word and is recognized, the other co-occurrence word centered words the occurrence word is also retrieved from the input speech, such a process operation is repeated sequentially, in co-occurrence relationship a plurality of stages One central word and a plurality of co-occurring words can be sequentially recognized.
【0062】例えば、入力音声が「350の缶のビール
を下さい」の場合、最初に中心語として“ビール”が抽
出されて対応する共起語としては“缶”が抽出される。
次に、この“缶”を中心語として“350”なる共起語
が抽出されるので、一つの入力音声から最終的に三つの
単語が認識されることになる。For example, when the input voice is “Please give me 350 cans of beer”, “beer” is first extracted as the central word, and “can” is extracted as the corresponding co-occurrence word.
Next, since the co-occurrence word "350" is extracted with this "can" as the central word, three words are finally recognized from one input voice.
【0063】さらに、上述のように中心語と共起語との
組み合わせを複数段階の階層構造とした場合に、図11
に示すように、認識語句辞書22に、一つの中心語と複
数の共起語との組み合わせの順番の情報も格納してお
き、語句認識手段23が、一つの中心語と複数の共起語
とを順番に対応して入力音声から認識することも可能で
ある。この場合、一つの中心語と複数の共起語とが入力
音声から順番に対応して認識されるので、複数の共起語
を良好な精度で高速に認識することができる。Further, when the combination of the central word and the co-occurring word has a hierarchical structure of a plurality of stages as described above, FIG.
As shown in (1), information on the order of combinations of one central word and a plurality of co-occurring words is also stored in the recognized phrase dictionary 22, and the phrase recognizing means 23 stores one central word and a plurality of co-occurring words. Can be recognized from the input voice in order. In this case, since one central word and a plurality of co-occurring words are recognized in order from the input speech, a plurality of co-occurring words can be recognized with good accuracy and at high speed.
【0064】例えば、入力音声が「350の缶のビール
を下さい」の場合、最初に中心語として“ビール”が抽
出され、これより前方の音声区間である「350の缶
の」から意味が“形態”の共起語である“缶”が抽出さ
れ、これより前方の音声区間である「350の」から意
味が“サイズ”の共起語である“350”が抽出され
る。For example, when the input voice is “Please give me 350 cans of beer”, “beer” is first extracted as a central word, and the meaning is obtained from “350 cans of beer” which is a voice section ahead of this. "Can" which is a co-occurrence word of "form" is extracted, and "350" which is a co-occurrence word of "size" is extracted from "350 of" which is a preceding voice section.
【0065】さらに、上述のように中心語と共起語との
組み合わせを複数段階の階層構造とした場合に、図12
に示すように、認識語句辞書22に、一つの中心語と複
数の共起語との組み合わせの階層構造の深度の情報も格
納しておき、語句認識手段23が、一つの中心語と複数
の共起語とを深度に対応して入力音声から認識すること
も可能である。この場合、一つの中心語から複数の共起
語を段階的に探索する処理動作が所定の深度まで実行さ
れるので、複数の共起語を必要な段階まで高速に認識す
ることができる。Further, when the combination of the central word and the co-occurring word has a hierarchical structure of a plurality of stages as described above, FIG.
As shown in (1), information on the depth of the hierarchical structure of a combination of one central word and a plurality of co-occurring words is also stored in the recognized phrase dictionary 22, and the phrase recognition means 23 It is also possible to recognize a co-occurrence word from an input voice in accordance with the depth. In this case, since a processing operation of searching for a plurality of co-occurring words stepwise from one central word is executed to a predetermined depth, it is possible to quickly recognize a plurality of co-occurring words to a necessary stage. Can be.
【0066】例えば、必要な階層構造が“2”として設
定されており、入力音声が「350の缶のビールを下さ
い」の場合、最初に中心語として“ビール”が抽出され
てから第一の共起語として“缶”が抽出された時点で、
階層構造の深度は“1”となる。そこで、この“缶”を
中心語として第二の共起語として“350”が抽出され
ると、階層構造の深度は“2”となるので、この時点で
段階的な音声認識の処理動作を終了する。For example, when the required hierarchical structure is set as “2” and the input voice is “Please give me 350 cans of beer”, the first word “beer” is extracted as the central word and then the first When “can” is extracted as a co-occurrence word,
The depth of the hierarchical structure is “1”. Then, when "350" is extracted as the second co-occurrence word with this "can" as the central word, the depth of the hierarchical structure becomes "2". finish.
【0067】[0067]
【発明の効果】請求項1記載の発明の音声認識装置は、
認識対象の音声の連続的な入力を受け付ける音声入力手
段と、共起関係にある複数の語句が組み合わされて格納
された認識語句辞書と、連続的な入力音声から共起関係
で組み合わされた複数の語句を認識する語句認識手段と
を有することにより、複数の語句を一つの連続的な入力
音声から共起関係の組み合わせに基づいて認識すること
ができるので、複数の語句を良好な精度で高速に認識す
ることができる。According to the first aspect of the present invention, there is provided a speech recognition apparatus.
A voice input means for receiving a continuous input of a speech to be recognized, a recognition phrase dictionary in which a plurality of co-occurring phrases are stored in combination, and a plurality of co-occurrence combinations of continuous input voices And the phrase recognition means for recognizing the multiple words can be recognized from one continuous input voice based on a combination of co-occurrence relations. Can be recognized.
【0068】請求項2記載の発明では、語句認識手段
は、共起関係で組み合わされた一対の語句の一方である
中心語を認識語句辞書から読み出して入力音声から抽出
してから、この抽出された中心語と共起関係にある他方
の語句である共起語を認識語句辞書から読み出して入力
音声から抽出することにより、中心語の抽出結果に基づ
いて入力音声に照合させる共起語を絞り込むことができ
るので、入力音声から共起語を認識する処理動作の負担
を軽減して速度を向上させることができ、共起関係にあ
る中心語と共起語とを良好な精度で高速に認識すること
ができる。According to the second aspect of the present invention, the word recognizing means reads out the central word, which is one of a pair of words combined in co-occurrence relation, from the recognized word dictionary and extracts it from the input speech, and then extracts the central word. central word and by extraction reads from the input speech from the other recognition word dictionary the occurrence word is a word in the co-occurrence relation, Filter occurrence word to be collated with the input speech based on the extracted result of the central word Can reduce the burden of processing operations for recognizing co-occurring words from input speech and improve the speed, and can quickly and quickly recognize co-occurring central words and co-occurring words with good accuracy. can do.
【0069】請求項3記載の発明では、語句認識手段
は、入力音声の中心語を抽出した区間を排除した区間か
ら共起語を抽出することにより、中心語の抽出結果に基
づいて共起語を照合させる入力音声の区間を制限するこ
とができるので、入力音声から共起語を認識する処理動
作の負担を軽減して速度を向上させることができ、共起
関係にある中心語と共起語とを良好な精度で高速に認識
することができる。[0069] In the present invention of claim 3, wherein, the phrase recognition means, by extracting the occurrence word from eliminating the section a section of extracting center language of the input speech, occurrence word based on the extraction result of the central word it is possible to limit the period of the input speech to be collated, the burden of recognizing processing operation occurrence word from the input speech speed can be improved by co-occurrence a central word in the co-occurrence relation Words can be quickly recognized with good accuracy.
【0070】請求項4記載の発明では、認識語句辞書
は、中心語と共起語との組み合わせに順番の情報も付与
されており、語句認識手段は、中心語と共起語とを順番
に対応して入力音声から認識することにより、共起語を
照合させる入力音声の区間を中心語の抽出区間より前方
か後方に制限することができるので、入力音声から共起
語を認識する処理動作の負担を軽減して速度を向上させ
ることができ、共起関係にある中心語と共起語とを良好
な精度で高速に認識することができる。[0070] In the present invention of claim 4, wherein the recognition word dictionary, also the order of the information to the combination of the central word and occurrence word are applied, the phrase recognition means, in turn a central word and occurrence word by recognizing the input speech corresponding to, it is possible to limit the forward or backward from the extraction section of the central words the section of the input speech to match the co-occurrence word, the co-occurrence <br/> word from the input speech The speed of the recognition operation can be reduced by reducing the load of the recognition operation, and the co-occurrence center word and the co-occurrence word can be quickly and accurately recognized with good accuracy.
【0071】請求項5記載の発明では、認識語句辞書
は、中心語と共起語との組み合わせに中間に位置する付
属語も格納されており、語句認識手段は、中心語と付属
語と共起語とを入力音声から認識することにより、例え
ば、ある中心語と付属語との共起語との読みが、他の中
心語と共起語との組み合わせの読みと同一の場合でも、
これらを各々別個に認識することができるので、共起関
係にある中心語と付属語と共起語とを良好な精度で認識
することができる。According to the fifth aspect of the present invention, the recognition phrase dictionary is a combination of a central word and a co-occurrence word .
Generic words are also stored, and the phrase recognition means recognizes a central word, an adjunct word, and a co-occurrence word from the input voice, for example, a co-occurrence word of a certain central word and an adjunct word. Is the same as the reading of the combination of other central words and co-occurring words,
Since these can be recognized separately, the central word, the auxiliary word, and the co-occurrence word which are in a co-occurrence relationship can be recognized with good accuracy.
【0072】請求項6記載の発明では、認識語句辞書
は、中心語と共起語との組み合わせに時間間隔の情報も
付与されており、語句認識手段は、中心語と共起語とを
時間間隔に対応して入力音声から認識することにより、
共起語を照合させる入力音声の区間を中心語の抽出区間
から所定の時間間隔の範囲に制限することができるの
で、入力音声から共起語を認識する処理動作の負担を軽
減して速度を向上させることができ、共起関係にある中
心語と共起語とを良好な精度で高速に認識することがで
きる。[0072] In the present invention of claim 6, wherein the recognition word dictionary, information of the time interval to a combination of the central word and occurrence word also been granted, the phrase recognition means, the center word and occurrence word and the time By recognizing from the input voice corresponding to the interval,
Since the section of the input speech for collating co- occurred words can be limited to a range of a predetermined time interval from the extraction section of the central word, the load of the processing operation of recognizing co- occurred words from the input speech is reduced and the speed is reduced. Thus, the co-occurrence word and the co-occurrence word can be quickly recognized with good accuracy.
【0073】請求項7記載の発明では、認識語句辞書
は、中心語と共起語との組み合わせが複数段階の階層構
造として格納されており、語句認識手段は、一つの中心
語と複数の共起語とを階層構造に対応して入力音声から
段階的に認識することにより、複数段階の共起関係にあ
る一つの中心語と複数の共起語とを段階的に順次認識す
ることができ、一つの語句の抽出結果に基づいて入力音
声に照合させる次の語句を絞り込むことができるので、
入力音声から複数の語句を段階的に順次認識する処理動
作の負担を軽減して速度を向上させることができ、一つ
の入力音声から多数の語句を良好な精度で高速に認識す
ることができる。[0073] In the present invention of claim 7, wherein the recognition word dictionary, a combination of a central word and occurrence word is stored as a hierarchical structure of a plurality of stages, the phrase recognition means, one of the central words and more co by stepwise recognized from the input speech corresponding to the hierarchical structure and cause words, it is possible to sequentially recognize stepwise and one central word and a plurality of co-occurrence word in the co-occurrence relationship a plurality of stages , You can narrow down the next phrase to match against the input voice based on the result of extracting one phrase,
The load on the processing operation of sequentially recognizing a plurality of words in a step-by-step manner from the input voice can be reduced and the speed can be improved, and many words can be recognized from one input voice at high speed with good accuracy.
【0074】請求項8記載の発明では、認識語句辞書
は、一つの中心語と複数の共起語との組み合わせに順番
の情報も付与されており、語句認識手段は、一つの中心
語と複数の共起語とを順番に対応して入力音声から認識
することにより、一つの語句の抽出結果に基づいて次の
語句を入力音声に照合させる場合に、この照合区間を直
前の語句の抽出区間より前方か後方に制限することがで
きるので、入力音声から複数の語句を段階的に順次認識
する処理動作の負担を軽減して速度を向上させることが
でき、一つの入力音声から多数の語句を良好な精度で高
速に認識することができる。According to the eighth aspect of the present invention, the recognition word dictionary is also provided with order information for a combination of one central word and a plurality of co-occurring words. When the next phrase is collated with the input voice based on the result of extracting one phrase by recognizing the co-occurrence words of the input speech in order, Since it is possible to restrict the forward or backward direction, it is possible to reduce the burden of the processing operation of sequentially recognizing a plurality of phrases from the input voice and improve the speed, and it is possible to improve a plurality of phrases from one input voice. High-speed recognition can be performed with good accuracy.
【0075】請求項9記載の発明では、認識語句辞書
は、一つの中心語と複数の共起語との組み合わせに階層
構造の深度の情報も付与されており、語句認識手段は、
一つの中心語と複数の共起語とを深度に対応して入力音
声から認識することにより、複数段階の共起関係にある
一つの中心語と複数の共起語とを段階的に順次認識する
処理動作を所定の深度まで実行することができるので、
一つの入力音声から多数の語句を必要な段階まで認識す
ることができる。According to the ninth aspect of the present invention, in the recognition word dictionary, information on the depth of the hierarchical structure is also given to a combination of one central word and a plurality of co-occurrence words.
Recognizing one central word and multiple co-occurring words from the input speech corresponding to depth, and sequentially recognizing one central word and multiple co-occurring words in a multi-stage co-occurrence relationship Can be performed to a predetermined depth,
Many words and phrases can be recognized from one input voice to a necessary stage.
【0076】請求項10記載の音声認識方法は、共起関
係にある複数の語句を組み合わせて設定しておき、認識
対象の音声の連続的な入力を受け付け、この連続的な入
力音声から共起関係で組み合わされた複数の語句を認識
するようにしたことにより、複数の語句が一つの連続的
な入力音声から共起関係の組み合わせに基づいて認識さ
れるので、複数の語句を良好な精度で高速に認識するこ
とができる。According to a tenth aspect of the present invention, a plurality of words having a co-occurrence relationship are set in combination, a continuous input of a speech to be recognized is received, and a co-occurrence is obtained from the continuous input speech. By recognizing multiple words combined in a relationship, multiple words are recognized based on a combination of co-occurrence relationships from one continuous input voice, so multiple words can be recognized with good accuracy. Can be recognized at high speed.
【0077】請求項11記載の音声認識方法は、共起関
係にある中心語と共起語とを組み合わせて設定してお
き、認識対象の音声の連続的な入力を受け付け、この連
続的な入力音声から用意された中心語を抽出し、この中
心語と共起関係にある共起語を入力音声から抽出するよ
うにしたことにより、共起関係で組み合わされた中心語
と共起語とが一つの連続的な入力音声から認識され、中
心語の抽出結果に基づいて入力音声に照合させる共起語
を絞り込むことができるので、入力音声から共起語を認
識する処理動作の負担を軽減して速度を向上させること
ができ、共起関係にある中心語と共起語とを良好な精度
で高速に認識することができる。In the speech recognition method according to the present invention, a central word and a co-occurrence word having a co-occurrence relation are set in combination, and a continuous input of a speech to be recognized is received. By extracting a central word prepared from the voice and extracting a co-occurrence word having a co- occurrence relation with the central word from the input voice, the central word and the co-occurrence word combined in the co-occurrence relation are extracted. Since co-occurring words that are recognized from one continuous input voice and are matched with the input speech based on the extraction result of the central word can be narrowed down, the burden of processing operations for recognizing co-occurring words from the input voice can be reduced. Speed can be improved, and a central word and a co-occurring word in a co-occurring relationship can be quickly recognized with good accuracy.
【0078】請求項12記載の情報記憶媒体は、共起関
係にある複数の語句が組み合わされて格納される認識語
句辞書のソフトウェアと、連続的な入力音声から共起関
係で組み合わされた複数の語句を認識するためのプログ
ラムと、が書き込まれているので、この情報記憶媒体の
ソフトウェアをコンピュータに読み取らせて動作させれ
ば、このコンピュータは、複数の語句を一つの連続的な
入力音声から共起関係の組み合わせに基づいて認識する
ことができるので、複数の語句を良好な精度で高速に認
識することができる。According to a twelfth aspect of the present invention, there is provided an information storage medium, comprising: a recognition phrase dictionary software in which a plurality of words having a co-occurrence relation are stored in combination; Since a program for recognizing words and phrases has been written, if the computer reads and operates the software of the information storage medium, the computer can share a plurality of words from one continuous input voice. Since recognition can be performed based on a combination of occurrence relationships, a plurality of phrases can be recognized at high speed with good accuracy.
【0079】請求項13記載の情報記憶媒体は、共起関
係にある中心語と共起語とが組み合わされて格納される
認識語句辞書のソフトウェアと、中心語を認識語句辞書
から読み出して連続的な入力音声から抽出するためのプ
ログラムと、この抽出された中心語と共起関係にある共
起語を認識語句辞書から読み出して入力音声から抽出す
るためのプログラムと、が書き込まれていることによ
り、この情報記憶媒体のソフトウェアをコンピュータに
読み取らせて動作させれば、このコンピュータは、共起
関係で組み合わされた中心語と共起語とを一つの連続的
な入力音声から認識することができ、中心語の抽出結果
に基づいて入力音声に照合させる共起語を絞り込むこと
ができるので、入力音声から共起語を認識する処理動作
の負担を軽減して速度を向上させることができ、共起関
係にある中心語と共起語とを良好な精度で高速に認識す
ることができる。According to a thirteenth aspect of the present invention, there is provided an information storage medium, comprising: a recognition word dictionary software for storing a co-occurring central word and a co-occurring word in combination; a program for extracting the Do input speech, co-located on the extracted central word co-occurrence relationships
A program for extracting from the input speech by reading an electromotive Language recognition word dictionary, by is written, be operated is read by the software of the information storage medium into the computer, the computer cooccurrence Since the central word and co-occurrence word combined in the relation can be recognized from one continuous input voice, and the co-occurrence word to be collated with the input voice based on the extraction result of the central word can be narrowed down, The load on the processing operation of recognizing a co-occurring word from the input speech can be reduced and the speed can be improved, and the central word and the co-occurring word that have a co-occurring relationship can be quickly and accurately recognized.
【図面の簡単な説明】[Brief description of the drawings]
【図1】本発明の実施の一形態の音声認識装置の論理的
構造を示す模式図である。FIG. 1 is a schematic diagram showing a logical structure of a speech recognition device according to an embodiment of the present invention.
【図2】音声認識装置の物理的構造を示すブロック図で
ある。FIG. 2 is a block diagram showing a physical structure of the speech recognition device.
【図3】音声認識装置の外観を示す斜視図である。FIG. 3 is a perspective view showing an external appearance of the voice recognition device.
【図4】情報記憶媒体であるRAMに書き込まれたソフ
トウェアの論理的構造を示す模式図である。FIG. 4 is a schematic diagram showing a logical structure of software written in a RAM serving as an information storage medium.
【図5】認識語句辞書の記憶内容を示し、(a)は単語
辞書、(b)は共起辞書、(c)は語順辞書、を示す模
式図である。5A and 5B are schematic diagrams showing storage contents of a recognized phrase dictionary, wherein FIG. 5A is a word dictionary, FIG. 5B is a co-occurrence dictionary, and FIG. 5C is a word order dictionary.
【図6】音声認識装置の音声認識方法を示すフローチャ
ートである。FIG. 6 is a flowchart illustrating a speech recognition method of the speech recognition device.
【図7】第一の変形例の認識語句辞書の共起辞書と語順
辞書との記憶内容を示す模式図である。FIG. 7 is a schematic diagram showing storage contents of a co-occurrence dictionary and a word order dictionary of a recognized phrase dictionary according to a first modified example.
【図8】第二の変形例の認識語句辞書の共起辞書の記憶
内容を示す模式図である。FIG. 8 is a schematic diagram showing storage contents of a co-occurrence dictionary of a recognized phrase dictionary according to a second modified example.
【図9】第三の変形例の認識語句辞書の共起辞書の記憶
内容を示す模式図である。FIG. 9 is a schematic diagram showing storage contents of a co-occurrence dictionary of a recognized phrase dictionary of a third modified example.
【図10】第四の変形例の認識語句辞書の共起辞書の記
憶内容を示す模式図である。FIG. 10 is a schematic diagram showing storage contents of a co-occurrence dictionary of a recognized phrase dictionary according to a fourth modified example.
【図11】第五の変形例の認識語句辞書の単語辞書と語
順辞書との記憶内容を示す模式図である。FIG. 11 is a schematic diagram showing storage contents of a word dictionary and a word order dictionary of a recognized phrase dictionary of a fifth modified example.
【図12】第六の変形例の認識語句辞書の語順辞書の記
憶内容を示す模式図である。FIG. 12 is a schematic diagram showing storage contents of a word order dictionary of a recognized phrase dictionary according to a sixth modified example.
【符号の説明】 1 音声認識装置 2 コンピュータ 4〜7,9 情報記憶媒体 21 音声入力手段 22 認識語句辞書 23 語句認識手段 41,42 ソフトウェア 42 プログラム[Description of Signs] 1 Voice Recognition Device 2 Computer 4-7, 9 Information Storage Medium 21 Voice Input Means 22 Recognized Word Dictionary 23 Word Recognition Means 41, 42 Software 42 Program
【手続補正2】[Procedure amendment 2]
【補正対象書類名】図面[Document name to be amended] Drawing
【補正対象項目名】図5[Correction target item name] Fig. 5
【補正方法】変更[Correction method] Change
【補正内容】[Correction contents]
【図5】 FIG. 5
【手続補正3】[Procedure amendment 3]
【補正対象書類名】図面[Document name to be amended] Drawing
【補正対象項目名】図6[Correction target item name] Fig. 6
【補正方法】変更[Correction method] Change
【補正内容】[Correction contents]
【図6】 FIG. 6
【手続補正4】[Procedure amendment 4]
【補正対象書類名】図面[Document name to be amended] Drawing
【補正対象項目名】図7[Correction target item name] Fig. 7
【補正方法】変更[Correction method] Change
【補正内容】[Correction contents]
【図7】 FIG. 7
【手続補正5】[Procedure amendment 5]
【補正対象書類名】図面[Document name to be amended] Drawing
【補正対象項目名】図8[Correction target item name] Fig. 8
【補正方法】変更[Correction method] Change
【補正内容】[Correction contents]
【図8】 FIG. 8
【手続補正6】[Procedure amendment 6]
【補正対象書類名】図面[Document name to be amended] Drawing
【補正対象項目名】図9[Correction target item name] Fig. 9
【補正方法】変更[Correction method] Change
【補正内容】[Correction contents]
【図9】 FIG. 9
【手続補正7】[Procedure amendment 7]
【補正対象書類名】図面[Document name to be amended] Drawing
【補正対象項目名】図10[Correction target item name] FIG.
【補正方法】変更[Correction method] Change
【補正内容】[Correction contents]
【図10】 FIG. 10
Claims (13)
ける音声入力手段と、共起関係にある複数の語句が組み
合わされて格納された認識語句辞書と、連続的な入力音
声から共起関係で組み合わされた複数の語句を認識する
語句認識手段と、を有することを特徴とする音声認識装
置。1. A speech input means for receiving a continuous input of a speech to be recognized, a recognition phrase dictionary in which a plurality of co-occurring phrases are stored in combination, and a co-occurrence relationship from the continuous input speech And a phrase recognizing means for recognizing a plurality of phrases combined with each other.
れた一対の語句の一方である中心語を認識語句辞書から
読み出して入力音声から抽出してから、この抽出された
中心語と共起関係にある他方の語句である付属語を前記
認識語句辞書から読み出して入力音声から抽出すること
を特徴とする請求項1記載の音声認識装置。2. The phrase recognition means reads a central word, which is one of a pair of phrases combined in co-occurrence relation, from a recognized phrase dictionary and extracts it from an input voice, and then co-occurs with the extracted central word. 2. The speech recognition apparatus according to claim 1, wherein an attached word, which is the other related phrase, is read from the recognized phrase dictionary and extracted from the input speech.
出した区間を排除した区間から付属語を抽出することを
特徴とする請求項2記載の音声認識装置。3. The speech recognition apparatus according to claim 2, wherein the word recognition means extracts an auxiliary word from a section excluding a section from which a central word of the input speech has been extracted.
み合わせに順番の情報も付与されており、語句認識手段
は、中心語と付属語とを順番に対応して入力音声から認
識することを特徴とする請求項2または3記載の音声認
識装置。4. The recognition phrase dictionary is also provided with order information for a combination of a central word and an adjunct word, and the phrase recognizing means recognizes the central word and the adjunct word from the input speech in correspondence with the order. The speech recognition device according to claim 2 or 3, wherein:
み合わせに中間に位置する介在語も格納されており、語
句認識手段は、中心語と介在語と付属語とを入力音声か
ら認識することを特徴とする請求項4記載の音声認識装
置。5. The recognition phrase dictionary also stores intervening words located in the middle of a combination of a central word and an adjunct word, and the phrase recognizing means recognizes the central word, the intervening word and the adjunct word from the input speech. The voice recognition device according to claim 4, wherein the voice recognition is performed.
み合わせに時間間隔の情報も付与されており、語句認識
手段は、中心語と付属語とを時間間隔に対応して入力音
声から認識することを特徴とする請求項2または3記載
の音声認識装置。6. The recognition phrase dictionary is also provided with information on a time interval for a combination of a central word and an adjunct word, and the phrase recognizing means converts the central word and the adjunct word from an input voice in accordance with the time interval. 4. The voice recognition device according to claim 2, wherein the voice recognition is performed.
み合わせが複数段階の階層構造として格納されており、
語句認識手段は、一つの中心語と複数の付属語とを階層
構造に対応して入力音声から段階的に認識することを特
徴とする請求項2または3記載の音声認識装置。7. The recognition phrase dictionary stores a combination of a central word and an adjunct word in a hierarchical structure having a plurality of stages.
4. The speech recognition apparatus according to claim 2, wherein the phrase recognition means recognizes one central word and a plurality of attached words stepwise from the input speech corresponding to a hierarchical structure.
付属語との組み合わせに順番の情報も付与されており、
語句認識手段は、一つの中心語と複数の付属語とを順番
に対応して入力音声から認識することを特徴とする請求
項7記載の音声認識装置。8. The recognition phrase dictionary is also provided with order information for a combination of one central word and a plurality of attached words,
8. The speech recognition apparatus according to claim 7, wherein the phrase recognition unit recognizes one central word and a plurality of attached words in order from the input speech.
付属語との組み合わせに階層構造の深度の情報も付与さ
れており、語句認識手段は、一つの中心語と複数の付属
語とを深度に対応して入力音声から認識することを特徴
とする請求項7または8記載の音声認識装置。9. The recognition phrase dictionary is provided with information on the depth of a hierarchical structure for a combination of one central word and a plurality of attached words, and the phrase recognition means includes one central word and a plurality of attached words. 9. The speech recognition apparatus according to claim 7, wherein the speech recognition unit recognizes the input speech from the input speech in accordance with the depth.
せて設定しておき、認識対象の音声の連続的な入力を受
け付け、この連続的な入力音声から共起関係で組み合わ
された複数の語句を認識するようにしたことを特徴とす
る音声認識方法。10. A plurality of words having a co-occurrence relationship are set in combination, a continuous input of a speech to be recognized is received, and a plurality of words combined in a co-occurrence relationship from the continuous input speech. A voice recognition method characterized by recognizing a voice.
み合わせて設定しておき、認識対象の音声の連続的な入
力を受け付け、この連続的な入力音声から用意された中
心語を抽出し、この中心語と共起関係にある付属語を入
力音声から抽出するようにしたことを特徴とする音声認
識方法。11. A co-occurrence central word and an adjunct word are set in combination, a continuous input of a speech to be recognized is received, and a prepared central word is extracted from the continuous input voice. And a supplementary word having a co-occurrence relationship with the central word is extracted from the input speech.
アが予め書き込まれた情報記憶媒体において、共起関係
にある複数の語句が組み合わされて格納される認識語句
辞書のソフトウェアと、連続的な入力音声から共起関係
で組み合わされた複数の語句を認識するためのプログラ
ムと、が書き込まれていることを特徴とする情報記憶媒
体。12. An information storage medium in which software readable by a computer is written in advance, and a software for a recognition phrase dictionary in which a plurality of words having a co-occurrence relation are combined and stored, and a continuous input voice. An information storage medium, in which a program for recognizing a plurality of phrases combined in a starting relationship is written.
アが予め書き込まれた情報記憶媒体において、共起関係
にある中心語と付属語とが組み合わされて格納される認
識語句辞書のソフトウェアと、中心語を前記認識語句辞
書から読み出して連続的な入力音声から抽出するための
プログラムと、この抽出された中心語と共起関係にある
付属語を前記認識語句辞書から読み出して入力音声から
抽出するためのプログラムと、が書き込まれていること
を特徴とする情報記憶媒体。13. An information storage medium in which software that is readable by a computer is pre-written, wherein a software for a recognition phrase dictionary in which co-occurring central words and adjunct words are stored in combination, A program for reading from the recognition phrase dictionary and extracting it from the continuous input speech, and a program for reading from the recognition phrase dictionary an auxiliary word having a co-occurrence relationship with the extracted central word from the input speech. , An information storage medium characterized by being written therein.
Priority Applications (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP8211078A JPH1055196A (en) | 1996-08-09 | 1996-08-09 | Speech recognition device and method, information storage medium |
Applications Claiming Priority (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP8211078A JPH1055196A (en) | 1996-08-09 | 1996-08-09 | Speech recognition device and method, information storage medium |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| JPH1055196A true JPH1055196A (en) | 1998-02-24 |
Family
ID=16600051
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| JP8211078A Pending JPH1055196A (en) | 1996-08-09 | 1996-08-09 | Speech recognition device and method, information storage medium |
Country Status (1)
| Country | Link |
|---|---|
| JP (1) | JPH1055196A (en) |
Cited By (9)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP2000293189A (en) * | 1999-04-02 | 2000-10-20 | Toshiba Corp | Speech recognition apparatus and method |
| JP2001005488A (en) * | 1999-06-18 | 2001-01-12 | Mitsubishi Electric Corp | Spoken dialogue system |
| JP2006209022A (en) * | 2005-01-31 | 2006-08-10 | Toshiba Corp | Information retrieval system, method and program |
| JP2009139862A (en) * | 2007-12-10 | 2009-06-25 | Fujitsu Ltd | Speech recognition apparatus and computer program |
| JP2009295101A (en) * | 2008-06-09 | 2009-12-17 | Hitachi Ltd | Speech data retrieval system |
| JP2011169960A (en) * | 2010-02-16 | 2011-09-01 | Nec Corp | Apparatus for estimation of speech content, language model forming device, and method and program used therefor |
| JP2012189829A (en) * | 2011-03-10 | 2012-10-04 | Fujitsu Ltd | Voice recognition device, voice recognition method, and voice recognition program |
| JP2013200362A (en) * | 2012-03-23 | 2013-10-03 | Dowango:Kk | Voice recognition device, voice recognition program and voice recognition method |
| JP2017151665A (en) * | 2016-02-24 | 2017-08-31 | 日本電気株式会社 | Information processing apparatus, information processing method, and program |
-
1996
- 1996-08-09 JP JP8211078A patent/JPH1055196A/en active Pending
Cited By (10)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP2000293189A (en) * | 1999-04-02 | 2000-10-20 | Toshiba Corp | Speech recognition apparatus and method |
| JP2001005488A (en) * | 1999-06-18 | 2001-01-12 | Mitsubishi Electric Corp | Spoken dialogue system |
| JP2006209022A (en) * | 2005-01-31 | 2006-08-10 | Toshiba Corp | Information retrieval system, method and program |
| JP2009139862A (en) * | 2007-12-10 | 2009-06-25 | Fujitsu Ltd | Speech recognition apparatus and computer program |
| US8271280B2 (en) | 2007-12-10 | 2012-09-18 | Fujitsu Limited | Voice recognition apparatus and memory product |
| JP2009295101A (en) * | 2008-06-09 | 2009-12-17 | Hitachi Ltd | Speech data retrieval system |
| JP2011169960A (en) * | 2010-02-16 | 2011-09-01 | Nec Corp | Apparatus for estimation of speech content, language model forming device, and method and program used therefor |
| JP2012189829A (en) * | 2011-03-10 | 2012-10-04 | Fujitsu Ltd | Voice recognition device, voice recognition method, and voice recognition program |
| JP2013200362A (en) * | 2012-03-23 | 2013-10-03 | Dowango:Kk | Voice recognition device, voice recognition program and voice recognition method |
| JP2017151665A (en) * | 2016-02-24 | 2017-08-31 | 日本電気株式会社 | Information processing apparatus, information processing method, and program |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| JP5040909B2 (en) | Speech recognition dictionary creation support system, speech recognition dictionary creation support method, and speech recognition dictionary creation support program | |
| JP3848319B2 (en) | Information processing method and information processing apparatus | |
| US6334102B1 (en) | Method of adding vocabulary to a speech recognition system | |
| US6910012B2 (en) | Method and system for speech recognition using phonetically similar word alternatives | |
| US20060020473A1 (en) | Method, apparatus, and program for dialogue, and storage medium including a program stored therein | |
| US20070198245A1 (en) | Apparatus, method, and computer program product for supporting in communication through translation between different languages | |
| US20020095289A1 (en) | Method and apparatus for identifying prosodic word boundaries | |
| US20090138266A1 (en) | Apparatus, method, and computer program product for recognizing speech | |
| EP2317507B1 (en) | Corpus compilation for language model generation | |
| US20020091520A1 (en) | Method and apparatus for text input utilizing speech recognition | |
| KR101424193B1 (en) | Non-direct data-based pronunciation variation modeling system and method for improving performance of speech recognition system for non-native speaker speech | |
| JPH08248971A (en) | Text aloud reading device | |
| JP2020134719A (en) | Translation equipment, translation methods, and translation programs | |
| US20060241936A1 (en) | Pronunciation specifying apparatus, pronunciation specifying method and recording medium | |
| JPH1055196A (en) | Speech recognition device and method, information storage medium | |
| US7103533B2 (en) | Method for preserving contextual accuracy in an extendible speech recognition language model | |
| JP5611270B2 (en) | Word dividing device and word dividing method | |
| CN112966491A (en) | Character tone recognition method based on electronic book, electronic equipment and storage medium | |
| JP3441400B2 (en) | Language conversion rule creation device and program recording medium | |
| JPH08248980A (en) | Voice recognition device | |
| JPH0962286A (en) | Speech synthesizer and speech synthesis method | |
| JP3029403B2 (en) | Sentence data speech conversion system | |
| JP3865149B2 (en) | Speech recognition apparatus and method, dictionary creation apparatus, and information storage medium | |
| JP2006040150A (en) | Voice data retrieval device | |
| JP2009086911A (en) | Unique expression extraction device, its method, program, and recording medium |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| A02 | Decision of refusal |
Free format text: JAPANESE INTERMEDIATE CODE: A02 Effective date: 20041005 |