JPH01197796A - voice recognition device - Google Patents

voice recognition device

Info

Publication number
JPH01197796A
JPH01197796A JP63022443A JP2244388A JPH01197796A JP H01197796 A JPH01197796 A JP H01197796A JP 63022443 A JP63022443 A JP 63022443A JP 2244388 A JP2244388 A JP 2244388A JP H01197796 A JPH01197796 A JP H01197796A
Authority
JP
Japan
Prior art keywords
section
recognition device
switch
storage
compared
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
JP63022443A
Other languages
Japanese (ja)
Inventor
Junichiro Fujimoto
潤一郎 藤本
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Ricoh Co Ltd
Original Assignee
Ricoh Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Ricoh Co Ltd filed Critical Ricoh Co Ltd
Priority to JP63022443A priority Critical patent/JPH01197796A/en
Publication of JPH01197796A publication Critical patent/JPH01197796A/en
Pending legal-status Critical Current

Links

Abstract

(57)【要約】本公報は電子出願前の出願データであるた
め要約のデータは記録されません。
(57) [Summary] This bulletin contains application data before electronic filing, so abstract data is not recorded.

Description

【発明の詳細な説明】 挟権宛」 本発明は、音声認識装置に係る。[Detailed description of the invention] To the rightful person” The present invention relates to a speech recognition device.

更欺挟帆 音声認識の研究は盛んであるが、音声認識には使用者が
自分の音声を登録してから使用する特定話者方式と、何
の準備もなく使用できる不特定話者方式があり、前者の
方が後者より認識率が高いという特徴がある。一方、こ
れらを別々の機械とせず、一つの中に両者の機能を持た
せ、必要に応じて、これらを使い分けることが考えられ
る。この場合、多勢が共通に使う言葉は不特定話者方式
で認識し、そうでないものは特定話者方式で認識するよ
うにしておくと便利である。或は、特定の指定をするた
めのコマンドを不特定にしておき、例えば、自分の名前
を言うと、その人の音声辞書がディスクの中から検索さ
れてロードされるような使い方が便利である。このよう
な場合、一つの装置を何人かで使用する訳であるが、あ
る人が使用中に席を立ったような場合、次の人が使用す
る時には前の人の音声辞書がロードされたままであり、
認識しないという欠点がある。
Research on speech recognition is active, but there are two types of speech recognition: the speaker-specific method, in which the user registers his or her own voice, and the speaker-independent method, which can be used without any preparation. The former has a higher recognition rate than the latter. On the other hand, instead of using these as separate machines, it is conceivable to have both functions in one and use them properly as needed. In this case, it is convenient to recognize words that are commonly used by a large number of people using the speaker-independent method, and to recognize words that are not commonly used using the speaker-specific method. Alternatively, it may be convenient to leave commands for specific specifications unspecified, so that when you say your name, the person's voice dictionary is searched from the disk and loaded. . In such cases, one device is used by several people, and if one person gets up while using it, the voice dictionary of the previous person will be loaded when the next person uses it. There are up to
It has the disadvantage of not being recognized.

目     的 本発明は、上述のごとき実情に鑑みてなされたもので、
特に、一つの認識装置を複数の人で使う場合、いつでも
どの人の声でも認識できるような装置を提供することを
目的としてなされたものである。
Purpose The present invention was made in view of the above-mentioned circumstances.
In particular, when a single recognition device is used by multiple people, the purpose of this invention is to provide a device that can recognize any person's voice at any time.

構成 本発明は、上記目的を達成するために、音声を電気信号
に変換する手段と、この信号から特徴をとり出す特徴量
変換部と、その特徴量を格納する第1の記憶部と、これ
とは別の第2の記憶部と。
Configuration In order to achieve the above object, the present invention provides means for converting audio into an electrical signal, a feature converter for extracting features from this signal, a first storage for storing the features, and the like. and a second storage unit separate from the .

特徴量を比較すべき比較部と、比較した類似度合を保持
する保持部を有する音声認識装置において。
In a speech recognition device having a comparison unit for comparing feature amounts and a holding unit for holding the compared similarity degree.

時間計測部を有し、該時間計測部において音声が入力さ
れない時間を計測し、その値が定められた値を超えた時
1次に入力された音声を前記第2の記憶部に記憶された
特徴量とのみ比較するようにしたことを特徴としたこと
、或いは、音声を電気信号に変換する機器を配置する場
所を備えると共にスイッチを具備し、該機器を配置する
ことにより眞記スイッチが開閉され、それによって前記
比較部で比較する対象を前記第2の記憶部に記憶された
ものに限定するようにしたことを特徴としたものである
。以下、本発明の実施例に基いて説明する。
It has a time measuring section, and the time measuring section measures the time during which no voice is input, and when the value exceeds a predetermined value, the first inputted voice is stored in the second storage section. The feature is that the comparison is made only with the feature quantity, or it is equipped with a place for placing a device that converts audio into an electrical signal, and is also equipped with a switch, and by placing the device, the true key switch can be opened and closed. Accordingly, the objects to be compared in the comparison section are limited to those stored in the second storage section. Hereinafter, the present invention will be explained based on examples.

第1図は、本発明の一実施例を説明するための構成図で
、図中、1はマイクロフォン、2はバンドパスフィルタ
群、3は切換スイッチ、4.5はレジスタ、6は切換ス
イッチ、7は比較部、8は合算部、9は比較器、10は
閾値発生部、11はタイムカウンタ、12は認識結果出
力部で、この実施例は、音声を電気信号に変換する手段
と、この信号から特徴をとり出す特徴社変換部と、その
特徴量を格納する記憶部と、これとは別の第2の記憶部
と特徴量を比較すべき比較部と、比較した類似度合を保
持する保持部を有する音声認識装置において、時間計測
部を設け、音声が入力されない時間を計測し、その値が
定められた値を超えた時1次に入力された音声は第2の
記憶部内の特徴量とのみ比較するようにしたものである
。第1図において、音響、電気信号変換器1としてマイ
クロフォンを用い、その出力を特徴量変換部であるバン
ドパスフィルタ群へ印加せしめる。−バンドパスフィル
タ群2とマイクロフォン1の間にマイクアンプや音声区
間検出部を入れても良いことは言うまでもない。その後
、スイッチ3によって音声登録と認識とを分ける。スイ
ッチ3をC側へ倒すと登録で、D側へ倒すと認識である
。レジスタ4には不特定話者用の辞書が、レジスタ5に
は各使用者が登録した音声が格納されている。例えば、
レジスタ4は「スタート」 「エンドJ  rlJ  
r2Jr3J  r4J・・・「9」等の音声パターン
が登録されたROMだと考えれば良い。又、レジスタ5
の内容はハードディスク等に格納したり、ロードしたり
出来るようにするのが望ましい(ハードディスクは図示
せず)、比較部7はスイッチ6で指定されたレジスタ4
又は5のパターンと入力されたパターンを比較し、類似
度性を求めるようになっている。この際のパターンの比
較はDPマツチングとして知られている動的計画法を用
いる照合法や2値化したデータを用いる方法などの方法
を用いても差支えない。
FIG. 1 is a configuration diagram for explaining one embodiment of the present invention, in which 1 is a microphone, 2 is a group of band-pass filters, 3 is a changeover switch, 4.5 is a register, 6 is a changeover switch, 7 is a comparison section, 8 is a summation section, 9 is a comparator, 10 is a threshold generation section, 11 is a time counter, and 12 is a recognition result output section. A feature conversion unit that extracts features from a signal, a storage unit that stores the feature values, a comparison unit that compares the feature values with a separate second storage unit, and a storage unit that stores the compared similarity degrees. In a speech recognition device having a holding section, a time measuring section is provided to measure the time during which no voice is input, and when the value exceeds a predetermined value, the first input voice is recorded as a characteristic in the second storage section. It is designed to compare only the quantity. In FIG. 1, a microphone is used as the acoustic/electrical signal converter 1, and its output is applied to a group of band-pass filters, which is a feature converter. - It goes without saying that a microphone amplifier or a voice section detection section may be inserted between the bandpass filter group 2 and the microphone 1. Thereafter, the switch 3 separates voice registration and recognition. When switch 3 is turned to the C side, it is registered, and when it is turned to the D side, it is recognized. The register 4 stores a dictionary for unspecified speakers, and the register 5 stores voices registered by each user. for example,
Register 4 is “Start” “End J rlJ
r2Jr3J r4J... It can be thought of as a ROM in which voice patterns such as "9" are registered. Also, register 5
It is desirable to be able to store or load the contents into a hard disk or the like (the hard disk is not shown).
Or, the pattern No. 5 is compared with the input pattern to determine the degree of similarity. In this case, the patterns may be compared using a matching method using dynamic programming known as DP matching or a method using binarized data.

比較部7で比較した結果、最も類似していると判定され
た音声に対応する信号を認識結果12として出力する。
As a result of the comparison by the comparison unit 7, a signal corresponding to the voice determined to be most similar is output as a recognition result 12.

又、スイイチ3をD側に倒して使用している時は、常に
その出力をバンドパスフィルタの数分だけ合算部8で合
計し、つまり時間ごとのエネルギーを求めてその値が閾
値部1oに定められた値より大きいか小さいかを比較す
る。閾値の設定のしかたはマイクロフォンから音声を入
力しない時のE点での出力をある時間平均してその値と
して決めれば良い。この閾値より大きな入力が比較部9
に対してあった時、タイ11カウンタ11をリセットす
る。タイムカウンタ11は一定の時間が計測できるよう
なものであればどのようなものでも良く1例えば、リセ
ットと共に大きな値を設定し、減算をくり返し、Oにな
るまでの時間を測定すれば良い。0になった時に、信号
によってスイッチ6をA側へ倒すようにする。タイムカ
ウンタ11の設定値は1分〜5分程度が望ましい。まず
、使用者によって番号が割りあてられ。
Also, when the switch 3 is turned to the D side and used, the output is always summed up by the number of bandpass filters in the summation section 8, that is, the energy for each time is calculated and the value is added to the threshold section 1o. Compare whether it is greater or less than a specified value. The threshold value can be set by averaging the output at point E over a certain period of time when no voice is input from the microphone and determining that value. If the input is larger than this threshold, the comparator 9
11, the tie 11 counter 11 is reset. The time counter 11 may be of any type as long as it can measure a certain amount of time. For example, it may be reset and set to a large value, and the time taken until it reaches O by repeating subtraction is measured. When the value becomes 0, the switch 6 is turned to the A side by a signal. The setting value of the time counter 11 is preferably about 1 minute to 5 minutes. First, a number is assigned by the user.

3番の使用者が使うものとすると、スイッチ3をD側に
倒し、マイクロフォン1から「3」と発声すると、レジ
スタ4の不特定話者用の辞書と照合され、結果「3」が
得られる。この認識結果はどのような使い方をしても良
いが、ここではディスクの中から3番の話者の音声デー
タをレジスタ5ヘロードするようにするとする(図示せ
ず)。スイッチ6をB側へ倒し、3番の話者の辞書によ
って認識させる。途中、この使用者は、席を立ったとす
ると、タイムカウンタ11が働き、スイッチ6を自動的
にA側へもどす0次に、この装置を2番の話者が使う時
はマイクロフォン1に向かって「2」と言えば良く誰の
辞書がロードされているかを確認する必要はない。
Assuming that user number 3 is using it, when he flips switch 3 to the D side and utters "3" from microphone 1, it is checked against the dictionary for unspecified speakers in register 4, and the result is "3". . This recognition result may be used in any way, but here it is assumed that the voice data of the third speaker from the disk is loaded into the register 5 (not shown). Turn the switch 6 to the B side, and the speaker will be recognized using the dictionary of the third speaker. If this user gets up from his or her seat midway, the time counter 11 will operate and the switch 6 will automatically return to the A side.Next, when the second speaker uses this device, he or she should face the microphone 1. You can just say "2" and there is no need to check whose dictionary is loaded.

第2図は、本発明の他の実施例を説明するための構成図
で1図中、13はマイクスタンドを示し、その他、第1
図に実施例と同様の作用をする部分には第1図の場合と
同一の参照番号が付しである。
FIG. 2 is a configuration diagram for explaining another embodiment of the present invention. In FIG. 1, numeral 13 indicates a microphone stand;
In the figure, parts having the same function as in the embodiment are given the same reference numerals as in FIG. 1.

而して、この実施例は、音声を電気信号に変換する手段
と、この信号から特徴をとり出す特徴量変換部と、その
特徴量を格納する記憶部と、これとは別の第2の記憶部
と、特徴量を比較すべき比較部と、比較した類似度合を
保持する保持部を有する音声認識装置において、音声を
電気信号に変換する機器を配置する場所を備えると共に
スイッチを具備し、該機器を配置することで前記スイッ
チが開閉され、それによって、比較部で比較する対象を
第2の記憶部内のものに限定するようにしたものである
。この第2図に示した実施例の場合、第1図に示した実
施例において使用したタイムカウンタ11や比較器9が
なく、これに代ってマイクスタンド13がついており、
このマイクスタンド13に荷重がかかればスイッチ6を
A側へ倒すようになっている。而して、第1図に示した
実施例の場合は使用者が使用しない時間を計測すること
によってスイッチ6をA側つまり不特定話者用のレジス
タ4側へ倒すようになっていたが、第2図に示した実施
例では、使用者が席を立つ時にマイクロフォン1をマイ
クスタンド13に乗せるとスイッチ6がA側へ倒れるよ
うになっている。゛なお、以上の例では、不特定話者用
のレジスタ4はROMとして説明したが、この必要はな
く、これも登録可能にしておいて使用者各自が自分の声
で自分の名前を登録しておいても良い。
Therefore, this embodiment includes a means for converting audio into an electrical signal, a feature amount converting section that extracts features from this signal, a storage section that stores the feature amounts, and a second separate device. A speech recognition device having a storage section, a comparison section for comparing feature amounts, and a holding section for holding the compared degree of similarity, the speech recognition device including a place for arranging a device for converting speech into an electrical signal and a switch, By arranging the device, the switch is opened and closed, thereby limiting the objects to be compared in the comparison section to those in the second storage section. In the case of the embodiment shown in FIG. 2, there is no time counter 11 or comparator 9 used in the embodiment shown in FIG. 1, and a microphone stand 13 is provided instead.
When a load is applied to the microphone stand 13, the switch 6 is turned to the A side. In the case of the embodiment shown in FIG. 1, the switch 6 is moved to the A side, that is, the register 4 side for unspecified speakers, by measuring the time when the user does not use the device. In the embodiment shown in FIG. 2, when the user leaves his/her seat and places the microphone 1 on the microphone stand 13, the switch 6 is tilted toward the A side.゛Although in the above example, register 4 for unspecified speakers was explained as a ROM, this is not necessary, and it can also be made registrable so that each user can register his or her name using his/her own voice. You can leave it there.

紘−一來 以上の説明から明らかなように、本発明によると、使用
者が変った時にも、最初から音声認識装置が利用でき、
わずられしい手続きが不要となる。
Kazuki HiroAs is clear from the above explanation, according to the present invention, even when the user changes, the voice recognition device can be used from the beginning.
There is no need for complicated procedures.

【図面の簡単な説明】[Brief explanation of the drawing]

第1図及び第2図は、それぞれ本発明の詳細な説明する
ための構成図である。 1・・・マイクロフォン、2・・・バンドパスフィルタ
群、3・・・切換スイッチ、4.5・・・レジスタ、6
・・・切換スイッチ、7・・・比較部、8・・・合算器
、9・・・比較器、10・・・閾値発生部、11・・・
タイムカウンタ、12・・・認識結果出力部、13・・
・マイクスタンド。
FIG. 1 and FIG. 2 are configuration diagrams for explaining the present invention in detail, respectively. DESCRIPTION OF SYMBOLS 1... Microphone, 2... Bandpass filter group, 3... Changeover switch, 4.5... Register, 6
... Selector switch, 7... Comparison section, 8... Adder, 9... Comparator, 10... Threshold value generation section, 11...
Time counter, 12... Recognition result output unit, 13...
·Mike stand.

Claims (1)

【特許請求の範囲】 1、音声を電気信号に変換する手段と、この信号から特
徴をとり出す特徴量変換部と、その特徴量を格納する第
1の記憶部と、これとは別の第2の記憶部と、特徴量を
比較すべき比較部と、比較した類似度合を保持する保持
部を有する音声認識装置において、時間計測部を有し、
該時間計測部において音声が入力されない時間を計測し
、その値が定められた値を超えた時、次に入力された音
声を前記第2の記憶部に記憶された特徴量とのみ比較す
るようにしたことを特徴とする音声認識装置。 2、音声を電気信号に変換する手段と、この信号から特
徴をとり出す特徴量変換部と、その特徴量を格納する第
1の記憶部と、これとは別の第2の記憶部と、特徴量を
比較すべき比較部と、比較した類似度合を保持する保持
部とを有する音声認識装置において、音声を電気信号に
変換する機器を配置する場所を備えると共にスイッチを
具備し、該機器を配置することにより前記スイッチが開
閉され、それによって前記比較部で比較する対象を前記
第2の記憶部に記憶されたものに限定するようにしたこ
とを特徴とする音声認識装置。
[Claims] 1. A means for converting audio into an electrical signal, a feature amount converting section for extracting features from this signal, a first storage section for storing the feature amounts, and a separate second storage section. 2, a comparison unit for comparing feature amounts, and a storage unit for holding the compared similarity degree, the speech recognition device having a time measurement unit,
The time measurement unit measures the time during which no audio is input, and when the value exceeds a predetermined value, the next input audio is compared only with the feature amount stored in the second storage unit. A voice recognition device characterized by: 2. means for converting audio into an electrical signal, a feature converter for extracting features from this signal, a first storage for storing the features, and a second storage separate from this; A speech recognition device having a comparison section for comparing feature amounts and a holding section for holding the compared similarity degree, the speech recognition device having a place for arranging a device for converting speech into an electrical signal, and a switch for arranging the device. The voice recognition device is characterized in that the switch is opened and closed by arranging the switch, thereby limiting the objects to be compared in the comparison section to those stored in the second storage section.
JP63022443A 1988-02-02 1988-02-02 voice recognition device Pending JPH01197796A (en)

Priority Applications (1)

Application Number Priority Date Filing Date Title
JP63022443A JPH01197796A (en) 1988-02-02 1988-02-02 voice recognition device

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
JP63022443A JPH01197796A (en) 1988-02-02 1988-02-02 voice recognition device

Publications (1)

Publication Number Publication Date
JPH01197796A true JPH01197796A (en) 1989-08-09

Family

ID=12082851

Family Applications (1)

Application Number Title Priority Date Filing Date
JP63022443A Pending JPH01197796A (en) 1988-02-02 1988-02-02 voice recognition device

Country Status (1)

Country Link
JP (1) JPH01197796A (en)

Similar Documents

Publication Publication Date Title
JPS6217240B2 (en)
JPS5876893A (en) Voice recognition equipment
JPH01197796A (en) voice recognition device
JP4440414B2 (en) Speaker verification apparatus and method
JPH02178698A (en) Voice recognition device and telephone set using same
JPH02210500A (en) Standard pattern registering system
JPH0222699A (en) voice recognition device
JPS6075898A (en) Word voice recognition equipment
JPS6193499A (en) Audio pattern matching method
JPH06318099A (en) Talker recognition device
JPS5876892A (en) Voice recognition equipment
JPH0233199A (en) voice recognition device
JPS58105200A (en) Voice section detection device
JPS58136097A (en) Recognition pattern collation system
JPH09244684A (en) Personal authentication device
JPS60498A (en) Voice detector
JP2002133417A (en) Fingerprint collation device
JPS61278896A (en) Speaker collator
JPS6019884A (en) Introductory management device
JPS5923400A (en) Voice recognition equipment
JPS61246800A (en) Voice response switch
JPH02287398A (en) Voice recognizing system and voice recognizing device
JPS58205199A (en) Voice recognition equipment
JPH0536499U (en) Voice recognizer
JPS6063900U (en) voice recognition device