JP2000276191A - Voice recognizing method - Google Patents

Voice recognizing method

Info

Publication number
JP2000276191A
JP2000276191A JP11077381A JP7738199A JP2000276191A JP 2000276191 A JP2000276191 A JP 2000276191A JP 11077381 A JP11077381 A JP 11077381A JP 7738199 A JP7738199 A JP 7738199A JP 2000276191 A JP2000276191 A JP 2000276191A
Authority
JP
Japan
Prior art keywords
microphone
sound
voice
voice recognition
signal
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Granted
Application number
JP11077381A
Other languages
Japanese (ja)
Other versions
JP3649032B2 (en
Inventor
Takashi Miki
敬 三木
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Oki Electric Industry Co Ltd
Original Assignee
Oki Electric Industry Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Oki Electric Industry Co Ltd filed Critical Oki Electric Industry Co Ltd
Priority to JP07738199A priority Critical patent/JP3649032B2/en
Publication of JP2000276191A publication Critical patent/JP2000276191A/en
Application granted granted Critical
Publication of JP3649032B2 publication Critical patent/JP3649032B2/en
Anticipated expiration legal-status Critical
Expired - Fee Related legal-status Critical Current

Links

Abstract

PROBLEM TO BE SOLVED: To correctly recognize a voice by removing the breath and snort sound of a speaker that a clitic type microphone detects. SOLUTION: A main microphone 101 is provided nearby a mouth of a speaker and a submicrophone 102 is provided between the mouth and ear of the speaker respectively; and the difference in sound power between an acoustic signal obtained from the main microphone 101 and the acoustic signal obtained from the submicrophone 102 is used to discriminate the breath and snort sound of the speaker from the acoustic signal obtained from the main microphone 101, thereby performing voice recognition.

Description

【発明の詳細な説明】DETAILED DESCRIPTION OF THE INVENTION

【0001】[0001]

【発明の属する技術分野】本発明は音声認識用マイクロ
フォンならびにそのマイクロフォンを用いた音声認識方
法に関する。
BACKGROUND OF THE INVENTION 1. Field of the Invention The present invention relates to a microphone for speech recognition and a speech recognition method using the microphone.

【0002】[0002]

【従来の技術】従来の音声の認識においては、音声信号
を感度よく収録するために、接話型マイクを使用するこ
とがあった。この接話型マイクは、口とマイクを接近さ
せることができるために、比較的周囲の雑音の影響を受
けることなく音声だけの信号が得られる利点があった。
2. Description of the Related Art In conventional voice recognition, a close-talking microphone has been used to record voice signals with high sensitivity. The close-talking microphone has an advantage that a signal of only voice can be obtained relatively without being affected by ambient noise since the microphone can be brought close to the mouth.

【0003】[0003]

【発明が解決しようとする課題】しかしながら、接話型
マイクは話者がマイクロフォンを吹いてしまうと、これ
は大音量の音声信号となりエラーの原因となる。また、
呼吸音や鼻息なども音声信号として検出してしまい、誤
認識の原因となっていた。これらの課題を解決する為に
特開平10−69293号公報の「音声認識装置および
方法、情報記憶媒体」に示されるようなマイク波形の絶
対値から呼吸音や鼻息音を判定する方法があった。しか
し波形のレベルは発話者の声の大きさに依存するため、
音の大きな話者や小さな話者などの違いや、口とマイク
の位置関係によって左右されるため、呼吸音や鼻息音の
正確な判定ができない場合があった。
However, when the speaker blows the microphone, the close-talking microphone becomes a loud sound signal, which causes an error. Also,
Respiratory sounds and nasal breaths are also detected as voice signals, causing misrecognition. In order to solve these problems, there has been a method of determining a breathing sound or a nasal breathing sound from an absolute value of a microphone waveform as disclosed in “Speech recognition device and method, information storage medium” in JP-A-10-69293. . However, the level of the waveform depends on the loudness of the speaker,
Because it depends on the difference between loud or small speakers and the positional relationship between the mouth and the microphone, accurate determination of respiratory sounds and nasal breath sounds may not be possible.

【0004】[0004]

【課題を解決するための手段】本発明に係る音声認識方
法は、話者の口に近接した位置に主マイクと、前記話者
の口と耳との間の位置に副マイクとをそれぞれ設け、前
記主マイクから得られた音響信号と副マイクから得られ
た音響信号との音響パワーの相違を利用して前記主マイ
クから得られた音響信号から話者の呼吸音または鼻息音
を弁別して音声認識を行う工程を有するものである。そ
の結果、主マイクが検出してしまう話者の吸呼音や鼻息
音を除去して正しく音声を認識することができる。
According to the speech recognition method of the present invention, a main microphone is provided at a position close to a mouth of a speaker, and a sub microphone is provided at a position between a mouth and an ear of the speaker. By using the difference in sound power between the sound signal obtained from the main microphone and the sound signal obtained from the sub microphone, the breathing sound or nasal breath sound of the speaker is distinguished from the sound signal obtained from the main microphone. It has a step of performing voice recognition. As a result, it is possible to correctly recognize the voice by removing the breath sound and the nasal breath sound of the speaker detected by the main microphone.

【0005】[0005]

【発明の実施の形態】実施形態1 実施形態1では、ヘッドセットマイクの主マイクのほか
に、その耳受けと主マイクの間のムーブに副マイクを取
りつけ、2本のマイク信号の音響パワーの違いを利用し
て、呼吸音や鼻息音などを安定かつ正確に峻別する。図
1は本発明の実施形態1〜3に係るヘッドセットマイク
の構成図であり、図2は図1のヘッドセットマイクを使
用した本実施形態1に係る音声認識処理装置の構成図で
ある。
BEST MODE FOR CARRYING OUT THE INVENTION Embodiment 1 In Embodiment 1, in addition to a main microphone of a headset microphone, a sub microphone is attached to a move between the earrest and the main microphone, and the acoustic power of two microphone signals is reduced. Taking advantage of the difference, it stably and accurately distinguishes breath sounds and nasal breath sounds. FIG. 1 is a configuration diagram of a headset microphone according to Embodiments 1 to 3 of the present invention, and FIG. 2 is a configuration diagram of a voice recognition processing device according to Embodiment 1 using the headset microphone of FIG.

【0006】図1において、101は主マイク、102
は副マイク、103は主マイク信号出力端子、104は
副マイク信号出力端子、105,106は耳当て、10
7はムーブであり、主マイク101の位置を調整できる
ように耳当て105に摺動自在に結合されている。なお
101〜107をすべて含んだ構成をヘッドセットマイ
クという。図1においては、通常のヘッドセットマイク
における音声を感度よく収集する主マイク101に加え
て、耳当て105と主マイク101を結ぶムーブ107
上でかつ耳当て105に近い部分に副マイク102を設
置する。また主マイク101と副マイク102の信号は
それぞれ主マイク信号出力端子103と副マイク信号出
力端子104から取り出す。106は耳当て105の反
対側の耳当てである。
In FIG. 1, reference numeral 101 denotes a main microphone;
, A sub microphone signal output terminal; 104, a sub microphone signal output terminal;
A move 7 is slidably coupled to the earpiece 105 so that the position of the main microphone 101 can be adjusted. A configuration including all of 101 to 107 is called a headset microphone. In FIG. 1, in addition to a main microphone 101 that collects sound from a normal headset microphone with high sensitivity, a move 107 that connects the earpiece 105 and the main microphone 101.
The sub-microphone 102 is placed above and near the earpiece 105. The signals of the main microphone 101 and the sub microphone 102 are taken out from the main microphone signal output terminal 103 and the sub microphone signal output terminal 104, respectively. Reference numeral 106 denotes an earpiece opposite to the earpiece 105.

【0007】図2において、201は音声判定音や、2
02は音声認識部、203は総合判定部、204は副マ
イク信号、205は主マイク信号である。図2の音声認
識処理装置は、副マイク信号204と主マイク信号20
5から呼吸音や鼻息音の判定するための判定信号Zfを
算出する音声判定部201と、主マイク信号205から
音声認識を行い、認識候補Cを求める音声認識部202
と、判定信号Zfと認識候補Cから音声認識結果Rを出
力する総合判定部203からなる。
In FIG. 2, reference numeral 201 denotes a voice judgment sound, 2
02 is a voice recognition unit, 203 is a comprehensive judgment unit, 204 is a sub microphone signal, and 205 is a main microphone signal. The voice recognition processing device of FIG.
5, a speech determination unit 201 that calculates a determination signal Zf for determining a breathing sound or a nasal breathing sound, and a voice recognition unit 202 that performs voice recognition from the main microphone signal 205 to obtain a recognition candidate C.
And a comprehensive determination unit 203 that outputs a speech recognition result R from the determination signal Zf and the recognition candidate C.

【0008】図1、2により実施形態1の動作を説明す
る。図1の主マイク101は口と近接しているために、
微少な音声の変化も捉える高感度な音声情報を収集でき
る。その一方で微少な呼吸音や鼻息なども確実に拾って
しまう。逆に副マイク102はそれほどの高感度ではな
いものの、呼吸や鼻息の流れからは離れているために呼
吸音や鼻息にはほとんど感知しない。この2つのマイク
の特性を生かすために、図2の音声判定部201では、
主マイク信号205の短時間音響パワーMfと副マイク
信号204の短時間音響パワーSfを計算する。
The operation of the first embodiment will be described with reference to FIGS. Since the main microphone 101 in FIG. 1 is close to the mouth,
Highly sensitive voice information that captures even small voice changes can be collected. On the other hand, minute breath sounds and sniffing are surely picked up. Conversely, although the secondary microphone 102 is not so sensitive, it is hardly perceived by breath sounds and nasal breaths because it is far from the flow of breathing and nasal breathing. In order to make use of the characteristics of these two microphones, the sound determination unit 201 in FIG.
The short-time sound power Mf of the main microphone signal 205 and the short-time sound power Sf of the sub microphone signal 204 are calculated.

【0009】音声判定部201では、まず入力される主
マイク信号205と副マイク信号204をそれぞれA/
D変換してデジタルの電気信号に変換する。このデジタ
ル信号に変換後の主マイク信号をMt、副マイク信号を
Stとすると、フレームと呼ばれる時間間隔T毎に、ま
ず下記式(1)、(2)により、主マイク信号205と
副マイク信号204の短時間音響パワーMfとSfを計
算し、次に下記式(3)によりMfとSfの比を判定信
号Zfとして求める。なお、式(1)〜(3)における
fはフレーム番号である。
[0009] First, an audio determination unit 201 converts an input main microphone signal 205 and sub microphone signal 204 into A / A signals.
The signal is D-converted to a digital electric signal. Assuming that the main microphone signal after conversion into the digital signal is Mt and the sub microphone signal is St, the main microphone signal 205 and the sub microphone signal are first determined at the time intervals T called frames by the following equations (1) and (2). Then, the short-time sound power Mf and Sf at 204 are calculated, and then the ratio between Mf and Sf is determined as a determination signal Zf by the following equation (3). Note that f in Expressions (1) to (3) is a frame number.

【0010】[0010]

【数1】 (Equation 1)

【0011】判定信号Zfは、通常の音声ではある一定
値V以下になるが、呼吸音や鼻息では、主マイクの短時
間音響パワーMfは大きいが、副マイクの短時間音響パ
ワーSfは非常に小さいままであるためZfはVを大幅
に超える値をとる。ただし、破裂音(日本語では例えば
「パ行音」等)などでは、破裂時点で、主マイク101
に息がかかることがあるために、瞬間的にZfが大きく
なることがある。しかし、長時間にわたってZfが大き
くなることは通常の音声ではありえない。
The judgment signal Zf is lower than a certain fixed value V for normal voices. For breathing sounds and nasal breaths, the short-time sound power Mf of the main microphone is large, but the short-time sound power Sf of the sub microphone is very low. Since it remains small, Zf takes a value that greatly exceeds V. However, in the case of a plosive sound (for example, in Japanese, for example, “pa-line sound”), the main microphone 101
In some cases, Zf may be instantaneously increased due to breathing. However, increasing Zf for a long time cannot be a normal voice.

【0012】音声認識部202の動作を説明する。先ず
音声認識部202は、主マイク信号205を音響特徴分
析し、フレーム周期f毎にi次元の音響特徴パラメータ
Xfiを算出する。この音響特徴パラメータXfiとし
ては、フーリエ解析や自己相関関数から算出されるLP
Cケプストラムが用いられるのが一般的である。LPC
ケプストラムの算出方法については、例えば“古井貞
煕、「ディジタル音声処理」、1985年9月、東海大
学出版会(以下、文献[1]と称す)、PP44〜4
8”に示されている。次に音声認識部202は、Xfi
の時間変化から音声が発せられている区間情報fs,f
eを見出す。fsは音声の始端時刻を示すフレーム番号
であり、feは認識候補の音声の終了時刻を示すフレー
ム番号である。この音声が発せられている区間を決定す
ることは、ほとんどの音声認識システムでは必須であ
り、その手法は同業者ならば周知の事項である。具体的
には、前記文献[1]のPP153〜154、などに記
載されている手法がある。
The operation of the speech recognition unit 202 will be described. First, the voice recognition unit 202 analyzes the acoustic characteristics of the main microphone signal 205 and calculates an i-dimensional acoustic characteristic parameter Xfi for each frame period f. As the acoustic feature parameter Xfi, LP calculated from Fourier analysis or autocorrelation function
Generally, C cepstrum is used. LPC
For the method of calculating the cepstrum, see, for example, “Sadahiro Furui,“ Digital Audio Processing ”, September 1985, Tokai University Press (hereinafter referred to as reference [1]), PP44-4.
8 ″. Next, the voice recognition unit 202
Section information fs, f where the voice is emitted from the time change of
Find e. fs is a frame number indicating the start time of the voice, and fe is a frame number indicating the end time of the voice of the recognition candidate. Determining the section in which the voice is being emitted is essential for most voice recognition systems, and the method is well known to those skilled in the art. Specifically, there is a method described in PP153 to 154 of the above document [1].

【0013】次に音声認識部202は、音声認識部20
2内部に記憶されている音声の音響特徴パラメータのA
fi(w)と主マイク信号205からの音響特徴パラメ
ータXfiを照合し、その類似性を検証する。ここで、
wは音声の特徴パタンAfi(w)の内容を示す番号
(例えばw=1は「東京」、w=2は「大阪」など予め
取り決めておく)である。この類似性の検証方法も、ほ
とんどの音声認識システムでは必須であり、その手法は
同業者ならば周知の事項である。例えば前記文献[1]
のPP167〜169、などに記載されている。最後に
認識候補w1を決定する。認識候補w1は、もっとも類
似している音声の音響特徴パラメータがAfi(w1)
とした場合、その内容を示す番号w1とする。
Next, the voice recognition unit 202
2 A of acoustic feature parameters of speech stored inside
fi (w) is collated with the acoustic feature parameter Xfi from the main microphone signal 205, and the similarity is verified. here,
w is a number indicating the contents of the feature pattern Afi (w) of the voice (for example, w = 1 is determined in advance such as “Tokyo”, w = 2 is determined in advance such as “Osaka”). This similarity verification method is also essential for most speech recognition systems, and the method is well known to those skilled in the art. For example, the above document [1]
PP167-169, etc. Finally, the recognition candidate w1 is determined. For the recognition candidate w1, the acoustic feature parameter of the most similar voice is Afi (w1).
In this case, a number w1 indicating the content is set.

【0014】総合判定部203の動作を説明する。先ず
総合判定部203は、音声認識部202で決定された音
声始端フレームfsから音声終端フレームfeまでの区
間に対して、音声判定部201で算出された判定信号Z
fの平均値ZZを次式(4)により計算する。
The operation of the comprehensive judgment section 203 will be described. First, the comprehensive determination unit 203 determines the determination signal Z calculated by the voice determination unit 201 for the section from the voice start frame fs to the voice end frame fe determined by the voice recognition unit 202.
The average value ZZ of f is calculated by the following equation (4).

【0015】[0015]

【数2】 (Equation 2)

【0016】このZZが定められた値Vより大きい場
合、この区間の主マイク信号205は呼吸音や鼻息によ
るものと判断され、最終的な認識結果Rは認識候補w1
でなく、エラーが発生したと見なす。逆に、ZZがVよ
り小さい場合、この区間の主マイク信号205は音声で
あるため、最終的な認識結果Rは認識候補w1とする。
If ZZ is larger than a predetermined value V, it is determined that the main microphone signal 205 in this section is due to a breathing sound or a nose breath, and the final recognition result R is a recognition candidate w1.
Instead, assume that an error has occurred. Conversely, if ZZ is smaller than V, the main microphone signal 205 in this section is a voice, so the final recognition result R is a recognition candidate w1.

【0017】以上説明したように、本実施形態1によれ
ば、ヘッドセットマイクに主マイク101と副マイク1
02の2つのマイクを設け、一方の主マイク101は口
と近接しているために、微少な音声の変化も捉える高感
度な音声情報を収集でき、他方の副マイク102はそれ
ほどの高感度ではないものの、呼吸や鼻息の流れからは
離れているために呼吸音や鼻息にはほとんど感知しない
よう構成される。そしてこの2つのマイク信号を組み合
わせて音声認識を行うことで、周囲の雑音の影響を受け
ることなく音声認識が行え、かつ呼吸音や鼻息による誤
認識をなくすことができる音声認識方式が実現できる。
As described above, according to the first embodiment, the main microphone 101 and the sub microphone 1 are connected to the headset microphone.
02, and one main microphone 101 is close to the mouth, so that it is possible to collect high-sensitivity voice information that captures even a small change in voice, and the other sub-microphone 102 is not so sensitive. Although it is not, it is configured so that it is hardly perceived by breath sounds and nasal breaths because it is far from the flow of breath and nasal breath. Then, by performing voice recognition by combining the two microphone signals, a voice recognition method can be realized in which voice recognition can be performed without being affected by ambient noise and erroneous recognition due to breathing sounds and nasal breaths can be eliminated.

【0018】実施形態2 実施形態2では、ヘッドセットマイクの主マイクのほか
に、その耳受けと主マイクの間のムーブに副マイクを取
りつけ、2本のマイク信号の音響パワーの違いを利用し
て、呼吸音や鼻息音などを安定かつ正確に峻別するとと
もに、雑音除去と音声特徴強調も行う。 図3は図1の
ヘッドセットマイクを用いた本実施形態2に係る音声認
識処理装置の構成図であり、図の301は音声判定部、
302は音声認識部、303は総合判定部、304は副
マイク信号、305は主マイク信号、306は雑音除去
部である。
Embodiment 2 In Embodiment 2, in addition to the main microphone of the headset microphone, a sub microphone is attached to the move between the earrest and the main microphone, and the difference in acoustic power between the two microphone signals is used. In addition to stably and accurately distinguishing breath sounds and nasal breath sounds, noise removal and voice feature enhancement are also performed. FIG. 3 is a configuration diagram of a voice recognition processing device according to the second embodiment using the headset microphone of FIG. 1.
Reference numeral 302 denotes a voice recognition unit, 303 denotes a general determination unit, 304 denotes a sub microphone signal, 305 denotes a main microphone signal, and 306 denotes a noise removal unit.

【0019】図3の音声認識処理装置は、副マイク信号
304と主マイク信号305から呼吸音や鼻息音の判定
するための判定信号Zfを算出する音声判定部301
と、副マイク信号304と主マイク信号305とから雑
音除去を行い、音声信号を取り出す雑音除去部306
と、この雑音除去後の音声信号を入力し、音声始端フレ
ームfsと音声終端フレームfeならびに認識候補w1
を求める音声認識部302と、音声認識部302の出力
する認識候補と前記判別信号Zfから音声認識結果Rを
出力する総合判定部303からなる。
The voice recognition processing apparatus shown in FIG. 3 is a voice determination unit 301 for calculating a determination signal Zf for determining a breathing sound or a nasal breathing sound from the sub microphone signal 304 and the main microphone signal 305.
And a noise removing unit 306 that removes noise from the sub microphone signal 304 and the main microphone signal 305 to extract an audio signal.
And the speech signal after the noise removal, and inputs the speech start frame fs, speech end frame fe, and the recognition candidate w1.
And a comprehensive determination unit 303 that outputs a voice recognition result R from the recognition candidate output by the voice recognition unit 302 and the determination signal Zf.

【0020】図3により実施形態2の動作を説明する。
図3の音声判定部301では、実施形態1の音声判定部
201と同様に主マイク信号305の短時間音響パワー
Mfと副マイク信号304の短時間音響パワーSfなら
びにその比である判定信号Zfを計算する。雑音除去部
306は主マイク信号305と副マイク信号304の性
質の違いから、背景ノイズ分をキャンセルする処理を行
い、主マイク信号305に混入した推定雑音成分を取り
除いた音声信号成分を取り出す。
The operation of the second embodiment will be described with reference to FIG.
3, the short-term sound power Mf of the main microphone signal 305 and the short-time sound power Sf of the sub microphone signal 304 and the judgment signal Zf, which is the ratio thereof, are determined in the same manner as the sound judgment unit 201 of the first embodiment. calculate. The noise removing unit 306 performs a process of canceling the background noise component due to the difference in properties between the main microphone signal 305 and the sub microphone signal 304, and extracts an audio signal component from which the estimated noise component mixed in the main microphone signal 305 has been removed.

【0021】この雑音除去処理は一般にアクティブ・ノ
イズ・キャンセラ(ANC)と呼ばれ、入力された雑音
信号(本実施形態2では副マイク信号304)と同振
幅、逆位相の逆位相のキャンセル信号を生成し、雑音信
号が混入した音声信号(本実施形態2では主マイク信号
305)にキャンセル信号を加えることで、雑音信号を
消しさる技術である。音声認識にこのANCを適用した
効果については、例えば“日本音響学会講演論文集、平
成5年10月、中山 昭ほか、3−8−10「適応ノイ
ズキャンセラを用いた騒音下音声認識」、PP−143
〜144”に示されている。
This noise elimination processing is generally called an active noise canceller (ANC), which cancels a cancel signal having the same amplitude and opposite phase as the input noise signal (the sub microphone signal 304 in the second embodiment). This is a technique for eliminating a noise signal by adding a cancel signal to a generated audio signal mixed with a noise signal (the main microphone signal 305 in the second embodiment). Regarding the effect of applying this ANC to speech recognition, see, for example, “Transactions of the Acoustical Society of Japan, October 1993, Akira Nakayama et al. 143
~ 144 ".

【0022】音声認識部302は、雑音除去部306で
生成された音声信号成分から実施形態1の音声認識部2
02と同様な音声認識処理を行い、音声始端フレームf
sと音声終端フレームfeならびに認識候補w1を求め
る。総合判定部303は、判定信号Zfと音声始端フレ
ームfsと音声終端フレームfeならびに認識候補w1
から最終的な認識結果Rを出力する。ここでの判定ロジ
ックは実施形態1と同様に行えば良い。
The speech recognition unit 302 uses the speech signal component generated by the noise removal unit 306 to
02 performs the same speech recognition processing as in the case of
s, speech end frame fe and recognition candidate w1 are obtained. The comprehensive determination unit 303 determines the determination signal Zf, the voice start frame fs, the voice end frame fe, and the recognition candidate w1.
Output the final recognition result R. The determination logic here may be performed in the same manner as in the first embodiment.

【0023】以上説明したように、本実施形態2によれ
ば、ヘッドセットマイクには、口に近接した主マイク1
01と口から少し離れた副マイク102の2つのマイク
を設け、主マイク信号305と副マイク信号304の両
マイク信号から呼吸音や鼻息の判定を行うと同時に、主
マイク信号305と副マイク信号304から背景雑音成
分を除去した音声成分を抽出し、この音声成分に対し
て、音声認識を行うことで、周囲の雑音の高度に除去し
た音声認識が行え、かつ呼吸音や鼻息による誤認識をな
くすことができる音声認識方式が実現できる。
As described above, according to the second embodiment, the headset microphone includes the main microphone 1 close to the mouth.
01 and a sub-microphone 102 slightly away from the mouth, a breathing sound and a nasal breath are determined from both the main microphone signal 305 and the sub-microphone signal 304, and at the same time, the main microphone signal 305 and the sub-microphone signal are determined. By extracting a speech component from which background noise components have been removed from 304 and performing speech recognition on the speech components, speech recognition with a high degree of elimination of surrounding noise can be performed, and erroneous recognition due to breathing sounds and nasal breaths can be prevented. A speech recognition method that can be eliminated can be realized.

【0024】実施形態3 実施形態3では、ヘッドセットマイクの主マイクのほか
に、その耳受けと主マイクの間のムーブに副マイクを取
りつけ、2本のマイク信号の音響パワーの違いを利用し
て、呼吸音や鼻息音などを安定かつ正確に峻別する判定
信号を音声認識類似度計算に反映させることで、音声の
一部分に呼吸音や鼻息がかかった場合でも正しく音声認
識を行うことができるようになる。図4は図1のヘッド
セットマイクを用いた本実施形態3に係る音声認識装置
の構成図であり、図の401は音声判定部、402は音
声認識部、403は総合判定部、404は副マイク信
号、405は主マイク信号である。
Embodiment 3 In Embodiment 3, in addition to the main microphone of the headset microphone, a sub microphone is attached to the move between the earrest and the main microphone, and the difference in acoustic power between the two microphone signals is used. By reflecting the determination signal for stably and accurately distinguishing breath sounds and nasal breath sounds in the speech recognition similarity calculation, correct speech recognition can be performed even when breath sounds or nasal breaths are applied to a part of the speech. Become like FIG. 4 is a configuration diagram of a voice recognition device according to the third embodiment using the headset microphone of FIG. 1. In FIG. 4, reference numeral 401 denotes a voice determination unit; 402, a voice recognition unit; A microphone signal 405 is a main microphone signal.

【0025】図4の音声認識処理装置は、副マイク信号
404と主マイク信号405から呼吸音や鼻息音の判定
するための判定信号Zfを算出する音声判定部401
と、主マイク信号405と判定信号Zfを入力し、音声
始端フレームfsと音声終端フレームfeならびに認識
候補w1を求める音声認識部402と、音声認識部40
2の出力する認識候補と前記判別信号Zfから音声認識
結果Rを出力する総合判定部403からなる。
The speech recognition processing apparatus shown in FIG. 4 is a speech judgment unit 401 for calculating a judgment signal Zf for judging a breathing sound or a nasal breathing sound from the sub microphone signal 404 and the main microphone signal 405.
And a main microphone signal 405 and a determination signal Zf, and a voice recognition unit 402 for obtaining a voice start frame fs and a voice end frame fe and a recognition candidate w1, and a voice recognition unit 40
2 and a general determination unit 403 that outputs a speech recognition result R from the recognition candidate output from the control unit 2 and the determination signal Zf.

【0026】図4により実施形態3の動作を説明する。
図4の音声判定部401では、実施形態1の音声判定部
201と同様に主マイク信号405の短時間音響パワー
Mfと副マイク信号404の短時間音響パワーSfなら
びにその比である判定信号Zfを計算する。音声認識部
402は、主マイク信号405と判定信号Zfから実施
形態1の音声認識部202と同様な音声認識処理を行
う。ただし、判定信号Zfが予め定められたVより大き
いフレームでは、その部分の主マイク信号が呼吸音や鼻
息である可能性が高いため、そのフレームでの認識処理
は認識候補w1の判定に与える影響を軽減する。
The operation of the third embodiment will be described with reference to FIG.
4, the short-time sound power Mf of the main microphone signal 405 and the short-time sound power Sf of the sub microphone signal 404 and the judgment signal Zf, which is the ratio thereof, are similar to the sound judgment unit 201 of the first embodiment. calculate. The voice recognition unit 402 performs the same voice recognition processing as the voice recognition unit 202 of the first embodiment from the main microphone signal 405 and the determination signal Zf. However, in a frame in which the determination signal Zf is larger than the predetermined V, the main microphone signal in that portion is likely to be a breathing sound or a nose breath, so that the recognition processing in that frame affects the determination of the recognition candidate w1. To reduce

【0027】次に音声認識部402は、音声認識部40
2内部に記憶されている音声の音響特徴パラメータAf
i(w)と主マイク信号405からの音響特徴パラメー
タXfiを照合し、その類似性を検証する。一般的な類
似性の検証方法では、音響特徴パラメータXfiと音声
の音響特徴パラメータAfi(w)とフレーム毎の局所
類似度Xfを計算し、その局所類似度Xfをフレーム毎
に算出して順次累積させ、この累積類似度が最大となる
w1を見出す。
Next, the voice recognition unit 402
2 Acoustic feature parameter Af of the voice stored inside
The sound characteristic parameter Xfi from the main microphone signal 405 is compared with i (w), and the similarity is verified. In a general similarity verification method, an acoustic feature parameter Xfi, an acoustic feature parameter Afi (w) of speech, and a local similarity Xf for each frame are calculated, and the local similarity Xf is calculated for each frame and sequentially accumulated. Then, w1 that maximizes the cumulative similarity is found.

【0028】しかし、本実施形態3では、この局所類似
度Xfに対して判定信号Zfに応じて重みを掛ける。重
みのかけ方としては、例えば、予め定められたVより大
きいフレームfではXfの値に0.3を乗じた値とする
ことで、そのフレームfの累積類似度に与える影響を軽
減できる。この処理により、呼吸音や鼻息がかかったと
想定されるフレームの影響を考慮した累積類似度が算出
でき、認識候補w1の判定精度が著しく向上する。総合
判定部403は、判定信号Zfと音声始端フレームfs
と音声終端フレームfeならびに認識候補w1から最終
的な認識結果Rを出力する。ここでの判定ロジックは実
施形態1と同様に行えば良い。
However, in the third embodiment, the local similarity Xf is weighted according to the determination signal Zf. As a method of weighting, for example, in the case of a frame f larger than a predetermined V, a value obtained by multiplying the value of Xf by 0.3 can reduce the influence on the cumulative similarity of the frame f. By this processing, the cumulative similarity can be calculated in consideration of the influence of the frame in which the breath sound or the nose breath is assumed to be applied, and the determination accuracy of the recognition candidate w1 is significantly improved. The overall determination unit 403 determines whether the determination signal Zf and the voice start frame fs
Then, a final recognition result R is output from the voice termination frame fe and the recognition candidate w1. The determination logic here may be performed in the same manner as in the first embodiment.

【0029】以上説明したように、本実施形態3によれ
ば、ヘッドセットマイクには、口に近接した主マイク1
01と口から少し離れた副マイク102の2つのマイク
を設け、主マイク信号405と副マイク信号404の両
マイク信号から呼吸音や鼻息の判定を行うと同時に、主
マイク信号405と副マイク信号404から求まる判定
信号Zfを認識処理の類似度計算に反映させることで、
音声の一部分に呼吸音や鼻息がかかった場合でも正しく
音声認識を行うことができるようになる。
As described above, according to the third embodiment, the headset microphone includes the main microphone 1 close to the mouth.
01 and a sub-microphone 102 slightly away from the mouth, a breathing sound and a nasal breath are determined from both the main microphone signal 405 and the sub-microphone signal 404, and at the same time, the main microphone signal 405 and the sub-microphone signal are determined. By reflecting the determination signal Zf obtained from 404 in the similarity calculation of the recognition process,
Even when a part of the voice is breathing or nose-breathing, the voice can be correctly recognized.

【0030】上記本発明に係る音声認識方法は、あらゆ
る音声認識装置や音声通話装置に適用して利用すること
ができる。なお、上記各実施形態では、副マイクは、ヘ
ッドセットマイクの耳受けと主マイクとの間のムーブに
取り付ける例を説明したが、本発明はこれに限定される
ものではなく、例えばイヤホーンマイク(イヤホーンと
一体化された小型マイク)を副マイクとして用いるよう
にしてもよい。
The voice recognition method according to the present invention can be applied to any voice recognition device and voice communication device. In each of the above embodiments, an example has been described in which the sub microphone is attached to the move between the earphone of the headset microphone and the main microphone. However, the present invention is not limited to this. For example, an earphone microphone ( A small microphone integrated with the earphone) may be used as the sub microphone.

【0031】[0031]

【発明の効果】以上のように本発明によれば、話者の口
に近接した位置に主マイクと、前記話者の口と耳との間
の位置に副マイクとをそれぞれ設け、前記主マイクから
得られた音響信号と副マイクから得られた音響信号との
音響パワーの相違を利用して前記主マイクから得られた
音響信号から話者の呼吸音または鼻息音を弁別して音声
認識を行う工程を有するようにしたので、主マイクが検
出してしまう話者の呼吸音や鼻息音を除去して正しく音
声を認識することができる。
As described above, according to the present invention, the main microphone is provided at a position close to the mouth of the speaker and the auxiliary microphone is provided at a position between the mouth and the ear of the speaker. Utilizing the difference in sound power between the sound signal obtained from the microphone and the sound signal obtained from the sub-microphone, the sound signal obtained from the main microphone is used to discriminate a breathing sound or a nasal breath sound of a speaker to perform voice recognition. Since the method includes the step of performing, it is possible to remove the breathing sound and the nasal breathing sound of the speaker which is detected by the main microphone, and to correctly recognize the voice.

【図面の簡単な説明】[Brief description of the drawings]

【図1】本発明の実施形態1〜3に係るヘッドセットマ
イクの構成図である。
FIG. 1 is a configuration diagram of a headset microphone according to Embodiments 1 to 3 of the present invention.

【図2】本発明の実施形態1に係る音声認識装置の構成
図である。
FIG. 2 is a configuration diagram of a speech recognition device according to the first embodiment of the present invention.

【図3】本発明の実施形態2に係る音声認識装置の構成
図である。
FIG. 3 is a configuration diagram of a voice recognition device according to a second embodiment of the present invention.

【図4】本発明の実施形態3に係る音声認識装置の構成
図である。
FIG. 4 is a configuration diagram of a speech recognition device according to a third embodiment of the present invention.

【符号の説明】[Explanation of symbols]

101 主マイク 102 副マイク 103 主マイク出力端子 104 副マイク出力端子 105,106 耳当て 107 ムーブ 201,301,401 音声判定部 202,302,402 音声認識部 203,303,403 総合判定部 204,304,404 副マイク信号 205,305,405 主マイク信号 306 雑音除去部 Reference Signs List 101 Main microphone 102 Sub microphone 103 Main microphone output terminal 104 Sub microphone output terminal 105, 106 Earpiece 107 Move 201, 301, 401 Voice determination unit 202, 302, 402 Voice recognition unit 203, 303, 403 Total determination unit 204, 304 , 404 Sub microphone signal 205, 305, 405 Main microphone signal 306 Noise removing unit

Claims (6)

【特許請求の範囲】[Claims] 【請求項1】 話者の口に近接した位置に主マイクと、
前記話者の口と耳との間の位置に副マイクとをそれぞれ
設け、 前記主マイクから得られた音響信号と副マイクから得ら
れた音響信号との音響パワーの相違を利用して前記主マ
イクから得られた音響信号から話者の呼吸音または鼻息
音を弁別して音声認識を行う工程を有することを特徴と
する音声認識方法。
1. A main microphone close to a mouth of a speaker;
A sub microphone is provided at a position between the mouth and the ear of the speaker, and the main microphone is provided by utilizing a difference in sound power between an audio signal obtained from the main microphone and an audio signal obtained from the sub microphone. A speech recognition method comprising a step of performing speech recognition by distinguishing a breathing sound or a nasal breathing sound of a speaker from an acoustic signal obtained from a microphone.
【請求項2】 前記音声認識を行う工程は、 前記主マイクと副マイクから得られた2つの音響信号の
短時間音響パワーの比を呼吸音または鼻息音の判定信号
として算出する判定信号算出工程と、 前記主マイクから得られた音響信号の音響特徴分析、音
声発生区間の決定並びに音響特徴パラメータとの照合及
び類似性の検証を行い、最も類似している音声の音響特
徴パラメータを音声認識候補として得る音声認識候補取
得工程と、 前記音声認識候補取得工程が決定した各音声発生区間毎
に、前記判定信号算出工程の算出した判定信号の平均値
を求め、該平均値が所定の値を越えたか否かによって、
該当音声発生区間の音声認識候補を呼吸音または鼻息音
として識別するか、または音声として識別する音声識別
工程とを有することを特徴とする請求項1記載の音声認
識方法。
2. The step of performing voice recognition includes a step of calculating a ratio of short-time sound power of two sound signals obtained from the main microphone and the sub microphone as a judgment signal of a breath sound or a nasal breath sound. And performing acoustic feature analysis of the acoustic signal obtained from the main microphone, determining a sound generation section, collating and verifying similarity with the acoustic feature parameter, and determining the acoustic feature parameter of the most similar speech as a speech recognition candidate. Obtaining a voice recognition candidate obtaining step, and for each voice generation section determined by the voice recognition candidate obtaining step, obtains an average value of the determination signals calculated in the determination signal calculating step, and the average value exceeds a predetermined value. Depending on whether
2. The voice recognition method according to claim 1, further comprising: a voice recognition step of identifying a voice recognition candidate in the voice generation section as a breathing sound or a nasal breathing sound, or identifying it as a voice.
【請求項3】 前記音声認識を行う工程は、 前記主マイクと副マイクから得られた2つの音響信号の
短時間音響パワーの比を呼吸音または鼻息音の判定信号
として算出する判定信号算出工程と、 前記主マイクより得られた音響信号から前記副マイクよ
り得られた音響信号を減算して混入雑音を除去する雑音
除去工程と、 前記雑音除去工程によって主マイクより得られた音響信
号から混入雑音の除去された音響信号の音響特徴分析、
音声発生区間の決定並びに音響特徴パラメータとの照合
及び類似性の検証を行い、最も類似している音声の音響
特徴パラメータを音声認識候補として得る音声認識候補
取得工程と、 前記音声認識候補取得工程が決定した各音声発生区間毎
に、前記判定信号算出工程の算出した判定信号の平均値
を求め、該平均値が所定の値を越えたか否かによって、
該当音声発生区間の音声認識候補を呼吸音または鼻息音
として識別するか、または音声として識別する音声識別
工程とを有することを特徴とする請求項1記載の音声認
識方法。
3. The step of performing voice recognition includes a step of calculating a ratio of short-time sound power of two sound signals obtained from the main microphone and the sub microphone as a judgment signal of a breath sound or a nasal breath sound. A noise removing step of subtracting an acoustic signal obtained from the sub microphone from an acoustic signal obtained from the main microphone to remove mixed noise; and mixing a sound signal obtained from the main microphone by the noise removing step. Acoustic feature analysis of noise-free acoustic signals,
Determining a voice generation section and performing verification and similarity verification with an audio feature parameter to obtain a voice feature parameter of the most similar voice as a voice recognition candidate; For each determined sound generation section, determine the average value of the determination signal calculated in the determination signal calculation step, depending on whether the average value exceeds a predetermined value,
2. The voice recognition method according to claim 1, further comprising: a voice recognition step of identifying a voice recognition candidate in the voice generation section as a breathing sound or a nasal breathing sound, or identifying it as a voice.
【請求項4】 前記音声認識を行う工程は、 前記主マイクと副マイクから得られた2つの音響信号の
短時間音響パワーの比を呼吸音または鼻息音の判定信号
として算出する判定信号算出工程と、 前記主マイクから得られた音響信号の音響特徴分析、音
声発生区間の決定及び音響パラメータとの照合を行い、
次に前記判定信号算出工程の算出した判定信号を重み係
数とした類似度計算により前記音響パラメータとの類似
性の検証を行い、最も類似している音声の音響特徴パラ
メータを音声認識候補として得る音声認識候補取得工程
と、 前記音声認識候補取得工程が決定した各音声発生区間毎
に、前記判定信号算出工程の算出した判定信号の平均値
を求め、該平均値が所定の値を越えたか否かによって、
該当音声発生区間の音声認識候補を呼吸音または鼻息音
として識別するか、または音声として識別する音声識別
工程とを有することを特徴とする請求項1記載の音声認
識方法。
4. The step of performing voice recognition includes a step of calculating a ratio of short-time sound power of two sound signals obtained from the main microphone and the sub microphone as a judgment signal of a breath sound or a nasal breath sound. And, performing an acoustic feature analysis of the acoustic signal obtained from the main microphone, determining a sound generation section and collating with acoustic parameters,
Next, a similarity calculation using the determination signal calculated in the determination signal calculation step as a weighting factor to verify the similarity with the acoustic parameter, and obtain a voice characteristic parameter of the most similar voice as a voice recognition candidate Recognition candidate acquisition step, For each voice generation section determined by the speech recognition candidate acquisition step, determine the average value of the determination signal calculated in the determination signal calculation step, whether or not the average value exceeds a predetermined value By
2. The voice recognition method according to claim 1, further comprising: a voice recognition step of identifying a voice recognition candidate in the voice generation section as a breathing sound or a nasal breathing sound, or identifying it as a voice.
【請求項5】 前記副マイクは、ヘッドセットマイクの
耳受けと前記主マイクとの間のムーブに取り付けられた
ことを特徴とする請求項1から4までのいずれかの請求
項に記載の音声認識方法。
5. The audio according to claim 1, wherein the auxiliary microphone is attached to a move between an ear receiver of a headset microphone and the main microphone. Recognition method.
【請求項6】 前記副マイクは、イヤホーンマイクを用
いることを特徴とする請求項1から4までのいずれかの
請求項に記載の音声認識方法。
6. The voice recognition method according to claim 1, wherein an earphone microphone is used as the sub microphone.
JP07738199A 1999-03-23 1999-03-23 Speech recognition method Expired - Fee Related JP3649032B2 (en)

Priority Applications (1)

Application Number Priority Date Filing Date Title
JP07738199A JP3649032B2 (en) 1999-03-23 1999-03-23 Speech recognition method

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
JP07738199A JP3649032B2 (en) 1999-03-23 1999-03-23 Speech recognition method

Publications (2)

Publication Number Publication Date
JP2000276191A true JP2000276191A (en) 2000-10-06
JP3649032B2 JP3649032B2 (en) 2005-05-18

Family

ID=13632325

Family Applications (1)

Application Number Title Priority Date Filing Date
JP07738199A Expired - Fee Related JP3649032B2 (en) 1999-03-23 1999-03-23 Speech recognition method

Country Status (1)

Country Link
JP (1) JP3649032B2 (en)

Cited By (7)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2005303574A (en) * 2004-04-09 2005-10-27 Toshiba Corp Voice recognition headset
CN102368793A (en) * 2011-10-12 2012-03-07 惠州Tcl移动通信有限公司 Cell phone and conversation signal processing method thereof
JP2014063018A (en) * 2012-09-21 2014-04-10 Systec:Kk Ultra-small voice input device
JP2016042162A (en) * 2014-08-19 2016-03-31 大学共同利用機関法人情報・システム研究機構 Living body detection device, living body detection method, and program
WO2018038381A1 (en) * 2016-08-26 2018-03-01 삼성전자 주식회사 Portable device for controlling external device, and audio signal processing method therefor
CN109087648A (en) * 2018-08-21 2018-12-25 平安科技(深圳)有限公司 Sales counter voice monitoring method, device, computer equipment and storage medium
JP7100746B1 (en) 2021-06-21 2022-07-13 アルインコ株式会社 Wireless relay device and wireless communication system

Cited By (13)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2005303574A (en) * 2004-04-09 2005-10-27 Toshiba Corp Voice recognition headset
CN102368793A (en) * 2011-10-12 2012-03-07 惠州Tcl移动通信有限公司 Cell phone and conversation signal processing method thereof
CN102368793B (en) * 2011-10-12 2014-03-19 惠州Tcl移动通信有限公司 Cell phone and conversation signal processing method thereof
JP2014063018A (en) * 2012-09-21 2014-04-10 Systec:Kk Ultra-small voice input device
JP2016042162A (en) * 2014-08-19 2016-03-31 大学共同利用機関法人情報・システム研究機構 Living body detection device, living body detection method, and program
KR20180023617A (en) * 2016-08-26 2018-03-07 삼성전자주식회사 Portable device for controlling external device and audio signal processing method thereof
WO2018038381A1 (en) * 2016-08-26 2018-03-01 삼성전자 주식회사 Portable device for controlling external device, and audio signal processing method therefor
US11170767B2 (en) 2016-08-26 2021-11-09 Samsung Electronics Co., Ltd. Portable device for controlling external device, and audio signal processing method therefor
KR102814684B1 (en) * 2016-08-26 2025-05-29 삼성전자주식회사 Portable device for controlling external device and audio signal processing method thereof
CN109087648A (en) * 2018-08-21 2018-12-25 平安科技(深圳)有限公司 Sales counter voice monitoring method, device, computer equipment and storage medium
CN109087648B (en) * 2018-08-21 2023-10-20 平安科技(深圳)有限公司 Counter voice monitoring method, device, computer equipment and storage medium
JP7100746B1 (en) 2021-06-21 2022-07-13 アルインコ株式会社 Wireless relay device and wireless communication system
JP2023001751A (en) * 2021-06-21 2023-01-06 アルインコ株式会社 Radio relay device and radio communication system

Also Published As

Publication number Publication date
JP3649032B2 (en) 2005-05-18

Similar Documents

Publication Publication Date Title
US9959886B2 (en) Spectral comb voice activity detection
CN103262577B (en) The method of hearing aids and enhancing voice reproduction
KR100636317B1 (en) Distributed speech recognition system and method
CN109195042B (en) Low-power-consumption efficient noise reduction earphone and noise reduction system
KR101340520B1 (en) Apparatus and method for removing noise
JP4940414B2 (en) Audio processing method, audio processing program, and audio processing apparatus
JP2008299221A (en) Speech detection device
US20130231932A1 (en) Voice Activity Detection and Pitch Estimation
US20220180886A1 (en) Methods for clear call under noisy conditions
JP3649032B2 (en) Speech recognition method
CN113707156A (en) Vehicle-mounted voice recognition method and system
CN116935900A (en) Voice detection method
Liu et al. Leakage model and teeth clack removal for air-and bone-conductive integrated microphones
JP2005338454A (en) Spoken dialogue device
JP2007288242A (en) Operator evaluation method, apparatus, operator evaluation program, recording medium
Maganti et al. A perceptual masking approach for noise robust speech recognition
CN111226278B (en) Low-complexity voiced speech detection and pitch estimation
KR100574883B1 (en) Speech Extraction Method by Non-Voice Rejection
JP3520430B2 (en) Left and right sound image direction extraction method
JP4752028B2 (en) Discrimination processing method for non-speech speech in speech
JP2010164992A (en) Speech interaction device
CN105810198A (en) Channel robust speaker identification method and device based on characteristic domain compensation
JP2882792B2 (en) Standard pattern creation method
Pfau et al. Hidden markov model based speech activity detection for the ICSI meeting project
JP2007264132A (en) Voice detection apparatus and method

Legal Events

Date Code Title Description
A977 Report on retrieval

Free format text: JAPANESE INTERMEDIATE CODE: A971007

Effective date: 20041109

A131 Notification of reasons for refusal

Free format text: JAPANESE INTERMEDIATE CODE: A131

Effective date: 20041116

A521 Request for written amendment filed

Free format text: JAPANESE INTERMEDIATE CODE: A523

Effective date: 20050114

TRDD Decision of grant or rejection written
A01 Written decision to grant a patent or to grant a registration (utility model)

Free format text: JAPANESE INTERMEDIATE CODE: A01

Effective date: 20050201

A61 First payment of annual fees (during grant procedure)

Free format text: JAPANESE INTERMEDIATE CODE: A61

Effective date: 20050207

R150 Certificate of patent or registration of utility model

Free format text: JAPANESE INTERMEDIATE CODE: R150

FPAY Renewal fee payment (event date is renewal date of database)

Free format text: PAYMENT UNTIL: 20090225

Year of fee payment: 4

FPAY Renewal fee payment (event date is renewal date of database)

Free format text: PAYMENT UNTIL: 20090225

Year of fee payment: 4

FPAY Renewal fee payment (event date is renewal date of database)

Free format text: PAYMENT UNTIL: 20100225

Year of fee payment: 5

FPAY Renewal fee payment (event date is renewal date of database)

Free format text: PAYMENT UNTIL: 20110225

Year of fee payment: 6

FPAY Renewal fee payment (event date is renewal date of database)

Free format text: PAYMENT UNTIL: 20110225

Year of fee payment: 6

FPAY Renewal fee payment (event date is renewal date of database)

Free format text: PAYMENT UNTIL: 20120225

Year of fee payment: 7

LAPS Cancellation because of no payment of annual fees