JPH0827639B2 - Voice recognition device - Google Patents
Voice recognition deviceInfo
- Publication number
- JPH0827639B2 JPH0827639B2 JP60144744A JP14474485A JPH0827639B2 JP H0827639 B2 JPH0827639 B2 JP H0827639B2 JP 60144744 A JP60144744 A JP 60144744A JP 14474485 A JP14474485 A JP 14474485A JP H0827639 B2 JPH0827639 B2 JP H0827639B2
- Authority
- JP
- Japan
- Prior art keywords
- independent
- sequence
- series
- word
- word dictionary
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Expired - Lifetime
Links
Description
【発明の詳細な説明】 〔発明の利用分野〕 本発明は、単語・文節単位あるいは複数文節で発声さ
れた音声を認識し出力する音声認識装置の改良に関す
る。Description: FIELD OF THE INVENTION The present invention relates to an improvement of a voice recognition device for recognizing and outputting a voice uttered in a unit of word / phrase or in a plurality of phrases.
従来の装置は、特開昭59−116837に記載のように、文
節単位で発声された音声を認識し、複数の候補系列よ
り、音声認識結果の確からしさ以外の自立語長、頻度を
含む条件により評価して認識候補の出力順序変更するよ
うになつていた。しかし、該装置では、語長の短かい自
立語を含有する文節を入力した場合には認識性能が低下
することがあり、自立語長が影響する場合の認識性能に
対しては配慮されていなかつた。As described in Japanese Patent Laid-Open No. 59-116837, a conventional device recognizes a voice uttered in a phrase unit, and a condition including an independent word length and frequency other than the certainty of a voice recognition result from a plurality of candidate sequences. Then, the output order of the recognition candidates is evaluated and changed. However, in the device, the recognition performance may be deteriorated when a bunsetsu containing an independent word having a short word length is input, and consideration is not given to the recognition performance when the independent word length is affected. It was
本発明の目的は、単語・文節・複数文節単位で発声さ
れた音声を正しく認識する音声認識装置を提供すること
にある。It is an object of the present invention to provide a voice recognition device that correctly recognizes a voice uttered in units of words, phrases and plural phrases.
上記の目的を達成するために、本発明は以下の構成を
採る。In order to achieve the above object, the present invention has the following configurations.
発声された音声を認識する認識手段と、認識手段によ
り認識された音声から複数の候補系列を作成する候補系
列作成手段と、自立語に関する情報を登録している自立
語辞書と、付属語に関する情報を登録している付属語辞
書と、自立語辞書および付属語辞書を用いて、上記複数
の候補系列より正しい系列を選択する選択手段とを備え
た音声認識装置において、選択手段は、候補系列の自立
語数、自立語頻度および音響的類似度より正しい系列を
選択することを特徴とする。Recognition means for recognizing spoken voice, candidate series creation means for creating a plurality of candidate series from the speech recognized by the recognition means, independent word dictionary in which information on independent words is registered, and information on adjunct words In the speech recognition device having an adjunct word dictionary that has registered, and an independent word dictionary and an adjunct word dictionary, and a selection means that selects a correct series from the plurality of candidate series, the selection means is The feature is that the correct sequence is selected based on the number of independent words, the frequency of independent words, and the acoustic similarity.
以下、本発明の一実施例を、図を用いて説明する。第
1図は、本発明の概要を説明するための音声認識装置の
一例のブロツク図である。入力された単語/文節/複数
文節音声は、音響認識部1にて音韻/音節/単語等の標
準パターン2との照合が行なわれ、複数候補系列を音響
的類似度とともに認識結果バツフアメモリ3に格納す
る。認識結果バツフアメモリ3の内容の一例を第2図
(a)に示す。候補系列選択部4は、自立語辞書6およ
び付属語辞書7を有し、自立語辞書6には第2図(b)
のように自立語および該自立語の表記・出現頻度が登録
されており、付属語辞書7には第2図(c)のように付
属語および該付属語の表記が登録されている。An embodiment of the present invention will be described below with reference to the drawings. FIG. 1 is a block diagram of an example of a voice recognition device for explaining the outline of the present invention. The input word / syllable / plural syllable voice is collated with the standard pattern 2 such as phoneme / syllable / word in the acoustic recognition unit 1 and the plural candidate sequences are stored in the recognition result buffer memory 3 together with the acoustic similarity. To do. An example of the contents of the recognition result buffer memory 3 is shown in FIG. The candidate sequence selection unit 4 has an independent word dictionary 6 and an auxiliary word dictionary 7, and the independent word dictionary 6 is shown in FIG.
The independent word and the notation / appearance frequency of the independent word are registered, and the adjunct word and the notation of the adjunct word are registered in the adjunct word dictionary 7 as shown in FIG.
認識結果バツフアメモリ3中の候補系列は、音響的類
似度とともに1系列ずつ候補系列選択部4に送られ、自
立語辞書6および付属語辞書7の中の各項目の文字列と
の比較照合により、自立語および付属語に分解する。該
分解文字列は、自立語数,自立語頻度より求めた系列頻
度,未知語フラグ,音響的類似度とともに、系列解折バ
ツフアメモリ5に記憶される。例えば、第2図(a)の
第1系列の“ヒツヨウテアル”場合は、自立語辞書6中
の“ヒツ”、“ウ”、“アル”および付属語辞書7中の
“ヨ”、“テ”と照合がとれて、 ヒツ+ヨ+ウ+テ+アル と分解される。この時、自立語数は3、系列頻度は、
(“ヒツ”の頻度)+(“ウ”の頻度)+(“アル”の
頻度)を自立語数で除したもの、すなわち、(27+34+
451)/3=171となる。未知語フラグは、辞書項目と照合
がとれない文字列が存在する場合に値1をもつ。第1系
列“ヒツヨウテアル”の未知語フラグの値は0である。
第2図(a)の5系列について、系列解折バツフアメモ
リ5に記憶される自立語数,系列頻度,未知語フラグ,
音響的類似度の内容を第3図に示す。The candidate sequences in the recognition result buffer memory 3 are sent to the candidate sequence selection unit 4 one by one together with the acoustic similarity, and are compared and collated with the character strings of the respective items in the independent word dictionary 6 and the auxiliary word dictionary 7. Decompose into independent words and attached words. The decomposed character string is stored in the sequence solving buffer memory 5 together with the number of independent words, the sequence frequency obtained from the independent word frequency, the unknown word flag, and the acoustic similarity. For example, in the case of the first series “Hitsuyoutearu” in FIG. 2 (a), “Hits”, “U”, and “Al” in the independent word dictionary 6 and “Yo” and “Te” in the auxiliary word dictionary 7 are used. ", And it is disassembled as Hits + Yo + U + Te + Al. At this time, the number of independent words is 3, and the sequence frequency is
(The frequency of "hits") + (the frequency of "c") + (the frequency of "al") divided by the number of independent words, that is, (27 + 34 +
451) / 3 = 171. The unknown word flag has a value of 1 when there is a character string that cannot be matched with the dictionary item. The value of the unknown word flag of the first series “Hitoyo Teal” is 0.
Regarding the 5 sequences in FIG. 2 (a), the number of independent words stored in the sequence solution buffer memory 5, the sequence frequency, the unknown word flag,
The content of the acoustic similarity is shown in FIG.
次に、系列解折バツフアメモリ5に記憶されている系
列の中から、未知語フラグの値が0である系列の番号を
選択系列メモリ8に記憶する。例では、1,2,3,4の系列
の未知語フラグの値が0であるため番号1,2,3,4が選択
系列メモリ8に記憶される。次に、選択系列メモリ8に
記憶された番号の系列のうち、自立語数が最小のものの
系列の番号を、選択系列メモリ8に記憶し直す。例で
は、3,4の系列の自立語数両者ともに2で最小であるの
で、選択系メモリ8に番号3,4が記憶される。次に、選
択系列メモリ8に記憶された番号の系列のうち、音響的
類似度の値が最大である系列の番号を、選択系列メモリ
8に記憶し直す。例では、3,4の系列の音響的類似度の
値が両者とも71であるので、選択系列メモリ8に番号3,
4が記憶される。次に、選択系列メモリ8に記憶された
番号の系列のうち、系列頻度の値が最大の系列の番号
を、選択系列メモリ8に記憶し直す。例では、3の系列
の系列頻度の値が182、4の系列の系列頻度の値が161で
あるので、選択系列メモリ8に番号3が記憶される。Next, among the sequences stored in the sequence analysis buffer memory 5, the number of the sequence in which the value of the unknown word flag is 0 is stored in the selected sequence memory 8. In the example, the numbers 1, 2, 3, 4 are stored in the selected sequence memory 8 because the value of the unknown word flag of the sequence 1, 2, 3, 4 is 0. Next, of the series of numbers stored in the selected series memory 8, the number of the series having the smallest number of independent words is stored again in the selected series memory 8. In the example, both the numbers of independent words in the series of 3 and 4 are 2 and the minimum, so that the numbers 3 and 4 are stored in the selection system memory 8. Next, of the series of numbers stored in the selected series memory 8, the number of the series having the maximum acoustic similarity is stored again in the selected series memory 8. In the example, since the acoustic similarity values of the sequences 3 and 4 are both 71, the selected sequence memory 8 has the number 3,
4 is memorized. Next, of the series of numbers stored in the selected series memory 8, the series number having the largest series frequency value is stored in the selected series memory 8 again. In the example, the value of the sequence frequency of the sequence of 3 is 182, and the value of the sequence frequency of the sequence of 4 is 161. Therefore, the number 3 is stored in the selected sequence memory 8.
最後に、選択系列メモリ8に記憶されている番号の系
列をデイスプレイ用バツフアメモリ9に記憶する。自立
語辞書6,付属語辞書7の各項目と比較照合を行なう際
に、照合のとれた項目の「表記」を記憶しておいて、最
終的に選択系列メモリ8に記憶されている番号の系列の
「表記」列を、デイスプレイ用バツフアメモリ9に記憶
してもよい。Finally, the series of numbers stored in the selected series memory 8 is stored in the display buffer memory 9. When performing comparison and collation with each item of the independent word dictionary 6 and the adjunct word dictionary 7, the “notation” of the collated item is stored and finally the number stored in the selected series memory 8 is stored. The "notation" column of the sequence may be stored in the display buffer memory 9.
そして、デイスプレイ用バツフアメモリ9の内容を、
デイスプレイ10に表示する。Then, the contents of the display buffer memory 9 for display are
Display on display 10.
なお、音響認識部1の直後に、日本語において出現し
得る音節の組合せ等の情報を用いて候補系列を少数に絞
ることも可能である。Immediately after the acoustic recognition unit 1, it is possible to narrow down the candidate sequence to a small number by using information such as syllable combinations that can appear in Japanese.
以上説明したように、本発明では、音声認識の結果の
候補系列に対して、系列中の自立語数,自立語頻度,音
響的類似度を用いて、正しい系列を選択するので、日本
語として妥当な系列が高い精度で選ばれ出力される。As described above, according to the present invention, the correct sequence is selected for the candidate sequence resulting from the speech recognition by using the number of independent words in the sequence, the independent word frequency, and the acoustic similarity. Series are selected and output with high accuracy.
第1図は本発明の一実施例の全体構成図、第2図(a)
は文節「ヒツヨウデアル」を認識した時の音響認識部の
出力候補系列を示す図、同図(b)は本発明で使用する
自立語辞書を示す図、同図(c)は本発明で使用する付
属語辞書を示す図、第3図は「ヒツヨウデアル」を認識
した時の候補系列についての自立語数,系列頻度,未知
語フラグ,音響的類似度の一例を示す図である。 1…音響認識部、2…音韻/音節/単語標準パターン、
3…認識結果バツフアメモリ、4…候補系列選択部、5
…系列解折バツフアメモリ、6…自立語辞書、7…付属
語辞書、8…選択系列バツフアメモリ、9…デイスプレ
イ用バツフアメモリ、10…デイスプレイ。FIG. 1 is an overall configuration diagram of an embodiment of the present invention, and FIG. 2 (a).
Is a diagram showing an output candidate sequence of the acoustic recognition unit when recognizing the phrase "Hitoyodeal", FIG. 7B is a diagram showing an independent word dictionary used in the present invention, and FIG. 7C is used in the present invention. FIG. 3 is a diagram showing an adjunct word dictionary to be used, and FIG. 3 is a diagram showing an example of the number of independent words, the sequence frequency, the unknown word flag, and the acoustic similarity regarding the candidate sequence when "Hitoyodeal" is recognized. 1 ... acoustic recognition unit, 2 ... phoneme / syllable / word standard pattern,
3 ... Recognition result buffer memory, 4 ... Candidate sequence selection unit, 5
... Sequence analysis buffer memory, 6 ... Independent word dictionary, 7 ... Adjunct word dictionary, 8 ... Selected series buffer memory, 9 ... Display memory for display, 10 ... Display.
───────────────────────────────────────────────────── フロントページの続き (72)発明者 阿部 正博 東京都国分寺市東恋ヶ窪1丁目280番地 株式会社日立製作所中央研究所内 (72)発明者 武市 宜之 東京都国分寺市東恋ヶ窪1丁目280番地 株式会社日立製作所中央研究所内 (72)発明者 遠藤 裕英 東京都国分寺市東恋ヶ窪1丁目280番地 株式会社日立製作所中央研究所内 (56)参考文献 特開 昭59−180629(JP,A) ─────────────────────────────────────────────────── ─── Continued Front Page (72) Inventor Masahiro Abe 1-280 Higashi Koigakubo, Kokubunji City, Tokyo Inside Hitachi Central Research Laboratory (72) Inventor Yoshiyuki Takeichi 1-280 Higashi Koigakubo, Kokubunji City, Tokyo Hitachi Ltd. (72) Inventor Hirohide Endo 1-280, Higashikoigakubo, Kokubunji, Tokyo (56) References, Hitachi, Ltd. Central Research Laboratory (56) Reference JP-A-59-180629 (JP, A)
Claims (1)
を作成する候補系列作成手段と、 自立語に関する情報を登録している自立語辞書と、 付属語に関する情報を登録している付属語辞書と、 上記自立語辞書および上記付属語辞書を用いて、上記複
数の候補系列より正しい系列を選択する選択手段とを備
えた音声認識装置において、 上記選択手段は、上記複数の候補系列の自立語数、自立
語頻度および音響的類似度より正しい系列を選択するこ
とを特徴とする音声認識装置。1. A recognition means for recognizing a spoken voice, a candidate series creation means for creating a plurality of candidate series from the sounds recognized by the recognition means, and an independent word dictionary in which information about an independent word is registered. In a speech recognition device comprising: an adjunct word dictionary in which information about adjunct words is registered; and selecting means for selecting a correct sequence from the plurality of candidate sequences using the independent word dictionary and the adjunct word dictionary. The speech recognition device, wherein the selection means selects a correct sequence from the number of independent words, the independent word frequency, and the acoustic similarity of the plurality of candidate sequences.
Priority Applications (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP60144744A JPH0827639B2 (en) | 1985-07-03 | 1985-07-03 | Voice recognition device |
Applications Claiming Priority (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP60144744A JPH0827639B2 (en) | 1985-07-03 | 1985-07-03 | Voice recognition device |
Publications (2)
| Publication Number | Publication Date |
|---|---|
| JPS627095A JPS627095A (en) | 1987-01-14 |
| JPH0827639B2 true JPH0827639B2 (en) | 1996-03-21 |
Family
ID=15369350
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| JP60144744A Expired - Lifetime JPH0827639B2 (en) | 1985-07-03 | 1985-07-03 | Voice recognition device |
Country Status (1)
| Country | Link |
|---|---|
| JP (1) | JPH0827639B2 (en) |
Family Cites Families (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JPS59180629A (en) * | 1983-03-30 | 1984-10-13 | Comput Basic Mach Technol Res Assoc | Voice inputting device of japanese |
-
1985
- 1985-07-03 JP JP60144744A patent/JPH0827639B2/en not_active Expired - Lifetime
Also Published As
| Publication number | Publication date |
|---|---|
| JPS627095A (en) | 1987-01-14 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US5875426A (en) | Recognizing speech having word liaisons by adding a phoneme to reference word models | |
| KR20040076035A (en) | Method and apparatus for speech recognition using phone connection information | |
| JPH10503033A (en) | Speech recognition method and device based on new word modeling | |
| US6230128B1 (en) | Path link passing speech recognition with vocabulary node being capable of simultaneously processing plural path links | |
| El Méliani et al. | Accurate keyword spotting using strictly lexical fillers | |
| JP2002278579A (en) | Voice data search device | |
| JP2820093B2 (en) | Monosyllable recognition device | |
| JPH05119793A (en) | Method and device for speech recognition | |
| JPH0210957B2 (en) | ||
| JP3240691B2 (en) | Voice recognition method | |
| JP3299170B2 (en) | Voice registration recognition device | |
| JP2001147698A (en) | Pseudo-word generation method for speech recognition and speech recognition device | |
| JPH0736481A (en) | Interpolation speech recognition device | |
| JPS627095A (en) | Voice recognition equipment | |
| JPS63165925A (en) | Sentence read-out system | |
| JPH049320B2 (en) | ||
| JPS58186836A (en) | Voice input data processor | |
| Hwang et al. | Efficient speech recognition techniques for the Finals of Mandarin syllables | |
| JPH03179498A (en) | Voice japanese conversion system | |
| JPH04291399A (en) | Voice recognizing method | |
| JPS6073595A (en) | Voice input unit | |
| JPH0632021B2 (en) | Japanese speech recognizer | |
| JP3084864B2 (en) | Text input device | |
| Okawa et al. | Phrase recognition in conversational speech using prosodic and phonemic information | |
| JPH0473694A (en) | Japanese language speech recognizing method |