JPH04181297A - voice recognition device - Google Patents
voice recognition deviceInfo
- Publication number
- JPH04181297A JPH04181297A JP2310472A JP31047290A JPH04181297A JP H04181297 A JPH04181297 A JP H04181297A JP 2310472 A JP2310472 A JP 2310472A JP 31047290 A JP31047290 A JP 31047290A JP H04181297 A JPH04181297 A JP H04181297A
- Authority
- JP
- Japan
- Prior art keywords
- output
- section
- signal
- speech
- outputs
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
- 238000001514 detection method Methods 0.000 claims abstract description 34
- 238000006243 chemical reaction Methods 0.000 abstract description 6
- 230000011218 segmentation Effects 0.000 abstract 3
- 230000029058 respiratory gaseous exchange Effects 0.000 abstract 1
- 239000000872 buffer Substances 0.000 description 31
- 238000010586 diagram Methods 0.000 description 7
- 238000000605 extraction Methods 0.000 description 4
- 238000003384 imaging method Methods 0.000 description 4
- 239000003795 chemical substances by application Substances 0.000 description 1
- 238000007796 conventional method Methods 0.000 description 1
Abstract
Description
【発明の詳細な説明】
[産業上の利用分野]
本発明は、音声認識装置に関する6
[従来の技術]
従来1発声者が発話した音声をマイクロフォンから入力
して、その音声信号から音声区間を検出して音声認識を
行う、音声認識装置においては。[Detailed Description of the Invention] [Industrial Application Field] The present invention relates to a speech recognition device. In a speech recognition device that detects and performs speech recognition.
第2図に示すように、マイクロフォンから入力された信
号の、比較的レベルが低い部分を検出するか、あるいは
2発話直前の周囲の雑音を取り込んで、その信号レベル
と比較することによって、音声区間を検出していた。As shown in Figure 2, by detecting a relatively low-level part of the signal input from the microphone, or by capturing the surrounding noise immediately before two utterances and comparing it with the signal level, the speech interval can be detected. was detected.
[発明が解決しようとする課題及び目的]しかし、従来
の方法では2周囲の騒音による影響を受は易く、騒音レ
ベルの大きいところでは使用できなかった1本発明は、
かかる問題点に鑑みて成されたもので、その目的とする
ところは9周囲の雑音のレベルによらず、容易に音声区
間の検出を可能し、なおかつ、多様な発話に対する。音
声認識装置の適応範囲を広げることである。[Problems and objectives to be solved by the invention] However, conventional methods are easily affected by surrounding noise and cannot be used in places with high noise levels.
This system has been developed in view of these problems, and its purpose is to easily detect voice sections regardless of the level of surrounding noise, and to be able to handle a wide variety of utterances. The goal is to expand the range of applicability of speech recognition devices.
[課題を解決する手段]
本発明の音声認識装置は2発声者が発話した音声をマイ
クロフォンから入力して、その音声信号から音声区間を
検出して音声認識を行う、音声認識装置において2発声
者の唇付近の動画像を撮影する撮像手段と、前記撮像手
段の出力をデジタル化して、デジタルデータを出力する
。データ変換手段と、前記データ変換手段の出力を記憶
し、直前のデータと比較をして、その結果を出力する。[Means for Solving the Problems] The speech recognition device of the present invention inputs the voices uttered by two speakers from a microphone, and performs speech recognition by detecting a speech section from the voice signal. an imaging means for photographing a moving image of the vicinity of the lips of the person, and digitizing the output of the imaging means to output digital data. A data conversion means stores the output of the data conversion means, compares it with the immediately previous data, and outputs the result.
画像記憶部と、前記記憶部の出力を監視し、前記比較の
結果、差異がある程度以下になった状態が一定時間続い
たら信号を出力する。音声区間検出部と9発声者が発話
する際に生じる。鼻からの呼気の気流を検出し、信号を
出力する。気流検出手段と、前記気流検出手段の出力を
、基準値と比較し、比較結果を出力する。気流検出部と
、前記音声区間検出部と前記気流検出部の出力を受けて
。The outputs of the image storage section and the storage section are monitored, and if the difference remains below a certain level as a result of the comparison for a certain period of time, a signal is output. This occurs when the voice section detection unit and the speaker 9 speak. It detects the airflow of exhaled air from the nose and outputs a signal. The airflow detection means and the output of the airflow detection means are compared with a reference value, and a comparison result is output. receiving outputs from the airflow detection section, the voice section detection section, and the airflow detection section;
前記マイクロフォンから入力された音声データに。to the audio data input from the microphone.
音声区間を示すデータを刊加して出力する。音声区間切
り出し部と、前記音声区間切り出し部の出力を、音声と
して認識し、認識結果を出力する。Data indicating the voice section is added and output. A speech section cutting section and the output of the speech section cutting section are recognized as speech, and a recognition result is output.
認識部を有することを特徴とする。It is characterized by having a recognition section.
[作用]
本発明の音声認識装置の作用について説明すると、第1
図の機能ブロック図に示すように、1の撮像手段である
カメラで2発声者の唇付近を撮影し、その信号を2のデ
ータ変換部へ送る。ここで。[Operation] To explain the operation of the speech recognition device of the present invention, the first
As shown in the functional block diagram in the figure, a camera serving as an imaging means (1) photographs the vicinity of the lips of a speaker (2), and the signal is sent to a data conversion unit (2). here.
発声者の唇付近の画像がサンプリングされてデジタルデ
ータになり、3の画像記憶部に送られる、ここで、デジ
タル化した画像を記憶すると同時に。The image near the speaker's lips is sampled into digital data and sent to the image storage section 3, where the digitized image is simultaneously stored.
直前の画像データとの違いを計り、結果を4の音声区間
検出部へ送る。4の音声区間検出部では。The difference from the immediately previous image data is measured and the result is sent to the voice section detection section 4. In the voice section detection section 4.
3の出力を監視し、直前のデータとの差異がある程度以
下になった状態が一定時間続くまで、信号を6の音声区
間切り出し部に送る。一方、5の気流検出手段である圧
力センサは9発声者が発話する際に、鼻からの呼気の圧
力を検出し、6の気流検出部へ送る。6の気流検出部で
は、5の圧力センサの出力を、基準値と比較し、基準値
より大きい間、信号を8の音声区間切り出し部へ送る。The output of 3 is monitored, and the signal is sent to the audio section cutout section 6 until the difference with the previous data remains below a certain level for a certain period of time. On the other hand, the pressure sensor 5, which is the airflow detection means, detects the pressure of exhalation from the nose when the speaker 9 speaks, and sends it to the airflow detection section 6. The airflow detection section 6 compares the output of the pressure sensor 5 with a reference value, and sends a signal to the voice section extraction section 8 while the output is greater than the reference value.
8の音声区間切り出し部では、4の音声区間検出部と6
の気流検出部の出力を受けて、7のマイクロフォンから
入力された音声を2区間毎に分けて。The voice section extraction section 8 includes the voice section detection section 4 and the voice section detection section 6.
In response to the output of the airflow detection section, the audio input from the 7 microphones is divided into two sections.
9の認識部へ渡す、9の認識部では、8の音声区間切り
出し部の出力を認識して、認識結果を出力する。The recognition unit 9 passes the output to the recognition unit 9, which recognizes the output of the speech segment extraction unit 8 and outputs a recognition result.
[実施例1
以下に2本発明の音声認識装置の詳細を図示した実施例
に基づいて説明する。[Embodiment 1] The details of the speech recognition device of the present invention will be described below based on the illustrated embodiment.
第3図は9本発明の実施例を示す構成図である。FIG. 3 is a configuration diagram showing nine embodiments of the present invention.
図中符号31は、カメラであり、ここでは、 CCDカ
メラを使用している。32は、31のカメラの出力をデ
ジタル化するデータ変換手段である。 A/D変換器1
である。33の破線で囲まれた部分は2画像記憶部であ
り、331のフレームバッファ制御[R:、332のフ
レームバッファ1と、333のフレームバッファ2と、
334,335の排他的論理和(以下、 XORと云う
、)ゲートと、336のオアゲートで構成されている。Reference numeral 31 in the figure is a camera, and here a CCD camera is used. 32 is a data conversion means for digitizing the output of the camera 31. A/D converter 1
It is. The part surrounded by the broken line 33 is two image storage units, and the frame buffer control 331 [R:, the frame buffer 1 of 332, the frame buffer 2 of 333,
It consists of 334 and 335 exclusive OR (hereinafter referred to as XOR) gates and 336 OR gates.
34は、音声区間検出部であり、341のカウンタと、
342のコンパレータ1で構成されている。35は圧力
センサである6 36の点線で囲まれた部分は、気流検
出部であり、361の増幅器と、その出力と。34 is a voice section detection unit, which includes a counter 341;
It is composed of 342 comparators 1. 35 is a pressure sensor 6 The part surrounded by a dotted line 36 is an airflow detection section, and 361 is an amplifier and its output.
363の可変抵抗器によって決定される基準値とを比較
する362のコンパレータ2から構成されている。37
は、マイクロフォンである。38は。It is composed of 362 comparators 2 that compare the reference value determined by 363 variable resistors. 37
is a microphone. 38 is.
音声区間切り出し部で、381のA/D変換器2と。381 A/D converter 2 in the audio section cutting section.
382のバッファ制御回路と、383のバッファ1と、
384のバッファ2と、386のオアゲートで構成され
ている。39は、音声を認識して結果を表示する。認識
部である。382 buffer control circuit, 383 buffer 1,
It consists of 384 buffers 2 and 386 OR gates. 39 recognizes the voice and displays the result. This is the recognition part.
次に、このように構成した装置の動作を、第3図に示し
た構成図に基づいて説明する。Next, the operation of the apparatus configured as described above will be explained based on the configuration diagram shown in FIG.
31のカメラは9発声者の唇に焦点を合わせ。Camera 31 focuses on the lips of speaker 9.
また1口元以外は、撮影しないように調整されている。Also, the camera is adjusted so that it does not photograph anything other than one mouth.
31のカメラで撮影された映像は、32のA/D変換器
1で、デジタルデータに変換されて。The video taken by the camera 31 is converted into digital data by the A/D converter 1 32.
33の画像記憶部に送られる。 このデータは。The image data is sent to the image storage unit No. 33. This data is.
332のフレームバッファ1か、333のフレームバッ
ファ2に記憶される。33の画像記憶部について説明す
る0本装置が動作を開始した後の。332 frame buffer 1 or 333 frame buffer 2. After the device starts operating, the image storage unit of 33 will be described.
最初の画像データは、332のフレームバッファ1に記
憶される。その次の画像データは、333のフレームバ
ッファ2に記憶される。このとき。Initial image data is stored in frame buffer 1 at 332. The next image data is stored in the frame buffer 2 of 333. At this time.
332のフレームバッファlは読みだし状態になり、8
売みだされたデータは、334のXORゲートに送られ
る。フレームバッファのデータの読みだしは、32のA
/D変換器1の出力に同期して行われる。また、32の
A/D変換器1の出力も334のXORゲートに入力さ
れているため、334のXORゲートで、現在の画像デ
ータと、その直前の画像データとのXORがとられる6
そして、336のオアゲートを通って、34の音声区間
検出部に送られる。以後の画像データに対しては、フレ
ームバッファlとフレームバッファ2の役割を、交互に
切り替えて行う6以上のように、現在の画像と。The frame buffer l of 332 is in the reading state, and the frame buffer l of 8
The marketed data is sent to 334 XOR gates. Reading frame buffer data is 32 A.
This is done in synchronization with the output of the /D converter 1. In addition, since the output of the A/D converter 1 of 32 is also input to the XOR gate of 334, the current image data and the immediately previous image data are XORed by the XOR gate of 334.
Then, it passes through the OR gate 336 and is sent to the voice section detection section 34. For subsequent image data, the roles of frame buffer 1 and frame buffer 2 are alternately switched between the current image and the above 6.
その直前の画像とのXORをとると2画像に変化がない
部分、即ち、動きがない部分はOで表され。When XORed with the image just before that, the portion where there is no change in the two images, that is, the portion where there is no movement, is represented by O.
動きがある部分は1で表される。従って、■の個数をカ
ウントすることによって1画像に動きがあるかどうかを
判別することができる0以上のようにして、XQRがと
られたデータは、34の音声区間検出部に送られる。こ
こでは、341のカウンタで、入力されたデータの1の
個数を数え、数えた結果を342のコンパレータ1に送
って、予め設定された値(以下、スレショルドレベルと
云う)と比較する。その結果lの個数が、スレショルド
レベルを下回るまで、信号を出力する。なお。Parts with movement are represented by 1. Therefore, it is possible to determine whether or not there is movement in one image by counting the number of ■. Here, the counter 341 counts the number of 1's in the input data, and the counted result is sent to the comparator 1 342, where it is compared with a preset value (hereinafter referred to as a threshold level). As a result, a signal is output until the number l falls below the threshold level. In addition.
このスレショルドレベルは、カメラ、あるいは。This threshold level is the camera or.
A/D変換器lのノイズレベルに依存する。35の圧力
センサは1発声者の鼻からの呼気のみが十分に当たる位
置に取り付けられている。この圧力センサの出力は、3
6の気流検出部に入力される。It depends on the noise level of A/D converter l. No. 35 pressure sensors are installed at positions that are sufficiently exposed to only the exhaled air from the nose of one speaker. The output of this pressure sensor is 3
It is input to the airflow detection section 6.
その後、361の増幅器で増幅され、362のコンパレ
ータ2で、363の可変抵抗器によって設定された基準
値と比較される。比較した結果、圧力センサの出力が、
基準値より大きければ、信号が出力される。38の音声
区間切り出し部では。Thereafter, it is amplified by an amplifier 361, and compared with a reference value set by a variable resistor 363 by a comparator 2 362. As a result of the comparison, the output of the pressure sensor is
If it is larger than the reference value, a signal is output. 38 in the audio section extraction section.
37のマイクロフォンで集音した音声を、381の、A
/D変換器2でデジタルデータに変換し。The sound collected by microphone 37 is transferred to 381 and A.
/D converter 2 converts it into digital data.
383のバッファ1.あるいは、384のバッファ2に
記憶する。382のバッファ制御回路は。383 buffers 1. Alternatively, it is stored in buffer 2 of 384. 382 buffer control circuit.
34の音声区間検出部の信号、または、36の気流検出
部の信号を受けると、バッファ1.あるいは、バッファ
2に記憶されている音声データを1つの音声区間のデー
タとして、386のオアゲートを経て、39の認識部へ
出力する0通常、34の音声区間検出部の信号と、36
の気流検出部の信号は、は−ぼ同時に、38の音声区間
切り出し部に入力されるが2例えば、相づちのように、
鼻音のみで構成された音声が人力された場合2発声者の
唇はほとんど動かない、従って、34の音声区間検出部
からの信号はないが、36の気流検出部は、鼻音を発声
したことによる気流を検出することができ、信号を出力
できる。ゆえに2発声された音声が鼻音だけであっても
音声区間を検出できる。また2本装置が動作を開始した
直後の、34の音声区間検出部、あるいは、36の気流
検出部の信号は、音声入力開始の信号ともなる2本装置
が動作を開始し、342のコンパレータl、あるいは、
362のコンパレータ2の最初の信号が入力されると、
その信号が人力された時点からの音声が、最初の音声と
なり、バッファlに記憶される、382のバッファ制御
回路が、342のコンパレータ1.あるいは、362の
コンパレータ2の次の信号を受は取ると、バッファlに
記憶された音声データを読みだし、39の認識部へ送る
。When receiving the signal from the voice section detecting section 34 or the signal from the airflow detecting section 36, the buffer 1. Alternatively, the audio data stored in the buffer 2 is output as one audio section data through the OR gate 386 to the recognition section 39. Normally, the signal from the audio section detection section 34 and the signal from the audio section detection section 36 are output.
The signals from the airflow detecting section are input to the voice section cutting section 38 almost at the same time.
When a voice consisting only of nasal sounds is produced manually, the lips of the speaker hardly move.Therefore, there is no signal from the voice section detection section 34, but the airflow detection section 36 detects the result of the nasal sound being uttered. It can detect airflow and output a signal. Therefore, even if the two uttered voices are only nasal sounds, the voice section can be detected. Immediately after the two devices start operating, the signal from the voice section detecting section 34 or the airflow detecting section 36 serves as a signal to start voice input. ,or,
When the first signal of 362 comparator 2 is input,
The sound from the time when the signal is manually input becomes the first sound and is stored in the buffer l.The buffer control circuit 382 controls the comparator 1.342. Alternatively, when the next signal from the comparator 2 of 362 is received, the voice data stored in the buffer l is read out and sent to the recognition section of 39.
同時に、384のバッファ2を書き込み状態に設定し2
次に、342のコンパレータl、あるいは。At the same time, set buffer 2 of 384 to write state.
Next, the comparator l of 342, or.
362のコンパレータ2から信号が入力されるまで、3
84のバッファ2に音声データを書き込む。3 until the signal is input from comparator 2 of 362.
Audio data is written to buffer 2 of 84.
これらの動作を、342のコンパレータ1の出力がある
度に、切り替える。These operations are switched every time there is an output from the comparator 1 of 342.
尚2本発明に於て、カメラは、唇の動きのみを撮影する
ようにする必要がある6そのため2本実施例では、第4
図に示すように、マイクロフォン装置の43の集音ユニ
ットが取り付けられている42のアーム部分に44の小
型CCDカメラを取り付けているが、集音ユニットと一
体化することもできるし9発話者の頭部と一体になって
動く部分であれば、どこでも取り付けることができる。2. In the present invention, it is necessary for the camera to photograph only the movement of the lips. 6 Therefore, in the present embodiment, the fourth
As shown in the figure, 44 small CCD cameras are attached to arm 42, to which 43 sound collection units of the microphone device are attached, but they can also be integrated with the sound collection unit, or It can be attached anywhere as long as it moves in unison with the head.
また、45の圧力センサは2発声者の鼻からの呼気のみ
を検出するために、第4図の42のアームから、さらに
アームを延ばし2発話者の鼻の真下に取り付けているが
、42のアームと分離することもできるし9発話者の鼻
からの呼気のみを検出することができる部分なら、どこ
にでも取り付けることができる。In addition, in order to detect only the exhaled air from the nose of the second speaker, the pressure sensor 45 is further extended from the arm 42 in Fig. 4 and attached directly below the nose of the second speaker. It can be separated from the arm, or it can be attached to any part that can detect only the exhalation from the speaker's nose.
[発明の効果1
以上説明したように2本発明によれば2周囲の騒音レベ
ルに全く影響されることなく、正確に。[Effects of the Invention 1 As explained above, 2 According to the present invention, 2 accurate noise can be achieved without being affected by the surrounding noise level at all.
音声区間を検出し、さらに、音声認識装置が扱える音声
の範囲をより拡大することができ、その効果は顕著であ
る。It is possible to detect voice sections and further expand the range of voices that can be handled by the voice recognition device, and the effect is remarkable.
第1図は2本発明の音声認識装置の構成を示す機能ブロ
ック図。
第2図は、従来例を示す機能ブロック図、第3図は2本
発明の実施例を示す構成図。
第4図は2本発明の、カメラと圧力センサとマイクロフ
ォンの取り付けの一例を示す概念図。
1 ・・・撮像手段
2 ・・・データ変換部
3 ・・・画像記憶部
4・・・音声区間検出部
5・・・気流検出手段
6・・・気流検出部
7・・・マイクロフォン
8・・・音声区間切り出し部
9・・・認識部
21・・・マイクロフォン
22・・・音声レベル検出部
23・・・音声区間切り出し部
24・・・認識部
31・・・カメラ
32・・・A/D変換器1
33・・・画像記憶部
331・・・フレームバッファ制御回路332・・・フ
レームバッファ1
333・・・フレームバッファ2
334、335・・−XORゲート
336、366・・・オアゲート
34・・・音声区間検出部
341・・・カウンタ
342・・・コンパレータ1
35・・・圧力センサ
36・・・気流検出部
361・・・増幅器
362・・・コンパレータ2
363・・・可変抵抗器
37・・・マイクロフォン
38・・・音声区間切り出し部
381・・・A/D変換器2
382・・・バッファ制御回路
383・・・バッファ1
384・・・バッファ2
39・・・認識部
41・・・発話者の口
42・・・マイクロフォン装置のアーム43・・・集音
ユニット
44・・・CCDカメラ
45・・・圧力センサ
以 上
出願人 セイコーエプソン株式会社
代理人 弁理士 鈴木喜三部 他1名
第1図FIG. 1 is a functional block diagram showing the configuration of a speech recognition device according to the present invention. FIG. 2 is a functional block diagram showing a conventional example, and FIG. 3 is a configuration diagram showing two embodiments of the present invention. FIG. 4 is a conceptual diagram showing an example of how to attach a camera, a pressure sensor, and a microphone according to the present invention. 1...Imaging means 2...Data conversion section 3...Image storage section 4...Audio section detection section 5...Airflow detection means 6...Airflow detection section 7...Microphone 8...・Voice section cutting section 9...Recognition section 21...Microphone 22...Speech level detection section 23...Voice section cutting section 24...Recognition section 31...Camera 32...A/D Converter 1 33... Image storage unit 331... Frame buffer control circuit 332... Frame buffer 1 333... Frame buffer 2 334, 335...-XOR gates 336, 366... OR gate 34...・Voice section detection unit 341...Counter 342...Comparator 1 35...Pressure sensor 36...Air flow detection unit 361...Amplifier 362...Comparator 2 363...Variable resistor 37...・Microphone 38...Voice section cutting unit 381...A/D converter 2 382...Buffer control circuit 383...Buffer 1 384...Buffer 2 39...Recognition unit 41...Speech Person's mouth 42...Arm 43 of the microphone device...Sound collection unit 44...CCD camera 45...Pressure sensor or more Applicant: Seiko Epson Co., Ltd. Agent Patent attorney Kizobe Suzuki and 1 other person Figure 1
Claims (1)
て、その音声信号から音声区間を検出して音声認識を行
う、音声認識装置において、b)発声者の唇付近の動画
像を撮影する撮像手段と、 c)前記撮像手段の出力をデジタル化して、デジタルデ
ータを出力する、データ変換手段と、d)前記データ変
換手段の出力を記憶し、直前のデータと比較をして、そ
の結果を出力する、画像記憶部と、 e)前記記憶部の出力を監視し、前記比較の結果、差異
がある程度以下になった状態が一定時間続いたら信号を
出力する、音声区間検出部と、f)発声者が発話する際
に生じる、鼻からの呼気の気流を検出し、信号を出力す
る、気流検出手段と、 g)前記気流検出手段の出力を、基準値と比較し、比較
結果を出力する、気流検出部と、 h)前記音声区間検出部と前記気流検出部の出力を受け
て、前記マイクロフォンから入力された音声データに、
音声区間を示すデータを付加して出力する、音声区間切
り出し部と、 i)前記音声区間切り出し部の出力を、音声として認識
し、認識結果を出力する、認識部を有することを特徴と
する、音声認識装置。[Scope of Claims] a) A voice recognition device that performs voice recognition by inputting the voice uttered by a speaker from a microphone and detecting a voice section from the voice signal, b) A moving image of the vicinity of the speaker's lips. c) data converting means for digitizing the output of the image capturing means and outputting digital data; d) storing the output of the data converting means and comparing it with the immediately preceding data. e) an image storage unit that monitors the output of the storage unit and outputs a signal if the difference remains below a certain level as a result of the comparison for a certain period of time; f) airflow detection means for detecting the airflow of exhaled air from the nose that occurs when the speaker speaks and outputting a signal; g) comparing the output of the airflow detection means with a reference value; an airflow detection section that outputs a comparison result;
A speech segment cutting unit that adds and outputs data indicating a speech segment; and i) a recognition unit that recognizes the output of the speech segment cutting unit as speech and outputs a recognition result. Speech recognition device.
Priority Applications (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP2310472A JPH04181297A (en) | 1990-11-16 | 1990-11-16 | voice recognition device |
Applications Claiming Priority (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP2310472A JPH04181297A (en) | 1990-11-16 | 1990-11-16 | voice recognition device |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| JPH04181297A true JPH04181297A (en) | 1992-06-29 |
Family
ID=18005657
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| JP2310472A Pending JPH04181297A (en) | 1990-11-16 | 1990-11-16 | voice recognition device |
Country Status (1)
| Country | Link |
|---|---|
| JP (1) | JPH04181297A (en) |
Cited By (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| EP0765084A3 (en) * | 1995-09-21 | 1997-10-15 | At & T Corp | Automatic video tracking system |
| JP2002358089A (en) * | 2001-06-01 | 2002-12-13 | Denso Corp | Audio processing device and audio processing method |
| JP2007171637A (en) * | 2005-12-22 | 2007-07-05 | Toshiba Tec Corp | Audio processing device |
-
1990
- 1990-11-16 JP JP2310472A patent/JPH04181297A/en active Pending
Cited By (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| EP0765084A3 (en) * | 1995-09-21 | 1997-10-15 | At & T Corp | Automatic video tracking system |
| JP2002358089A (en) * | 2001-06-01 | 2002-12-13 | Denso Corp | Audio processing device and audio processing method |
| JP2007171637A (en) * | 2005-12-22 | 2007-07-05 | Toshiba Tec Corp | Audio processing device |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| JP2687712B2 (en) | Integrated video camera | |
| TWI383377B (en) | Multi-sensory speech recognition system and method | |
| JP6651989B2 (en) | Video processing apparatus, video processing method, and video processing system | |
| US20210158828A1 (en) | Audio processing device, image processing device, microphone array system, and audio processing method | |
| US11405584B1 (en) | Smart audio muting in a videoconferencing system | |
| TW201801069A (en) | Method, system and device for receiving voice information | |
| Yoshinaga et al. | Audio-visual speech recognition using lip movement extracted from side-face images. | |
| JP4715738B2 (en) | Utterance detection device and utterance detection method | |
| JP2015175983A (en) | Voice recognition device, voice recognition method, and program | |
| GB2375276A (en) | Method and system of sound processing | |
| JP2019192092A (en) | Conference support device, conference support system, conference support method, and program | |
| JP3838159B2 (en) | Speech recognition dialogue apparatus and program | |
| US8315865B2 (en) | Method and apparatus for adaptive conversation detection employing minimal computation | |
| JP5645393B2 (en) | Audio signal processing device | |
| JPH04180096A (en) | Voice recognition device | |
| JP4127155B2 (en) | Hearing aids | |
| Yoshinaga et al. | Audio-visual speech recognition using new lip features extracted from side-face images | |
| JPH04184495A (en) | Voice recognition device | |
| JP2000276191A (en) | Voice recognizing method | |
| JPH04186400A (en) | Speech recognition device | |
| TWI687917B (en) | Voice system and voice detection method | |
| JPH04181300A (en) | voice recognition device | |
| JP6112913B2 (en) | Surveillance camera system and method | |
| JPH01310399A (en) | Speech recognition device | |
| JPS63152277A (en) | Portable video camera |