JPH04343399A - Speech feature extraction method and speech recognition device - Google Patents

Speech feature extraction method and speech recognition device

Info

Publication number
JPH04343399A
JPH04343399A JP3143760A JP14376091A JPH04343399A JP H04343399 A JPH04343399 A JP H04343399A JP 3143760 A JP3143760 A JP 3143760A JP 14376091 A JP14376091 A JP 14376091A JP H04343399 A JPH04343399 A JP H04343399A
Authority
JP
Japan
Prior art keywords
level
voice
input signal
section
frequency analysis
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
JP3143760A
Other languages
Japanese (ja)
Inventor
Takashi Ariyoshi
有吉 敬
Mitsugi Matsushita
貢 松下
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Ricoh Co Ltd
Original Assignee
Ricoh Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Ricoh Co Ltd filed Critical Ricoh Co Ltd
Priority to JP3143760A priority Critical patent/JPH04343399A/en
Publication of JPH04343399A publication Critical patent/JPH04343399A/en
Pending legal-status Critical Current

Links

Abstract

PURPOSE:To improve the accuracy of voice recognition by correcting an utterance fluctuation under noise and extracting the feature quantity of arm utterer. CONSTITUTION:This device is provided with a level measuring means 11 for measuring a level of an input signal a frequency analyzing means 12 for executing a frequency analysis of the input signal, a voice section detecting means 13 for detecting a voice section in the input signal, and a coefficient setting means 14 for setting a coefficient of the frequency analyzing means, based on a level of the input signal measured by the level measuring means 11, at the time point when the voice section is not detected by the voice section detecting means. In such a state, a level of a noise for exerting the influence on an utterer is measured, and in accordance with its level, a coefficient of a frequency analysis filter is changed, and the feature quantity of a voice which corrects a spectrum deformation of utterance by a Lombard effect is obtained.

Description

【発明の詳細な説明】[Detailed description of the invention]

【0001】0001

【技術分野】本発明は、音声特徴量抽出方式及び該音声
特徴量抽出方式を用いた音声認識装置、より詳細には、
騒音下での発声変換を補正する技術に関し、例えば、事
務所内、自動車内、工場内、家庭内で使用される音声認
識装置に応用して好適なものである。
TECHNICAL FIELD The present invention relates to a voice feature extraction method and a speech recognition device using the voice feature extraction method, and more particularly, to
The present invention relates to a technology for correcting speech conversion under noisy conditions, and is suitable for application to speech recognition devices used in offices, cars, factories, and homes, for example.

【0002】0002

【従来技術】騒音下での音声は、無騒音下の音声と比べ
て、音量の増大やスペクトル変形が起こることが知られ
ている。Lombard効果と呼ばれてるこの効果は、
発声者自身に対する音声のフィードバックが騒音によっ
て妨げられ、結果として、発声者が大きな声を出すこと
によって起こる。このことは、騒音下での音声認識を困
難にする原因の一つとなっている。この対策としての従
来技術は、例えば、特開平3−2793号公報に記載の
音声認識用前処理装置があるが、この装置では、発声者
に、自分自身の音声を適切な音量でフィードバックさせ
るようにすることにより、Lombard効果自体を防
ぐようにしている。しかしながら、この方法では、再生
アンプ、ヘッドホンなどの再生装置が必要なことが一つ
の欠点であり、また、マイクロホンとして接話マイクが
使えないなど、音声と周囲の騒音の比(SN比)がある
程度悪い場合には、音量の増大もなくなるため、この方
法を用いない通常の発声よりSN比が悪化し、認識が困
難になることが別の欠点である。
2. Description of the Related Art It is known that the volume of speech in noise increases and the spectrum is distorted compared to speech in no-noise conditions. This effect is called the Lombard effect.
This occurs when the voice feedback to the speaker himself is blocked by the noise, resulting in the speaker speaking loudly. This is one of the reasons why speech recognition under noisy conditions is difficult. As a countermeasure against this, for example, there is a speech recognition preprocessing device described in Japanese Patent Application Laid-Open No. 3-2793, but this device has a system that allows the speaker to feed back his or her own voice at an appropriate volume. By doing so, the Lombard effect itself is prevented. However, one drawback of this method is that it requires a playback device such as a playback amplifier and headphones, and the ratio of the sound to surrounding noise (SN ratio) is limited to a certain extent, such as the inability to use a close-talking microphone. In the worst case, there is no increase in volume, so the S/N ratio becomes worse than normal speech without using this method, making recognition difficult, which is another drawback.

【0003】Lombard効果を受けた音声は、特に
低域において、ホルマント周波数が上昇することが知ら
れている。例えば、「雑音下での発声変形を考慮した認
識方式の検討」音講論、平1−10月、2−1−5によ
れば、86dBSPL騒音時では、無騒音時に対して、
1,5KHz以下のホルマント周波数が平均で約120
Hz上昇している。このような、発声変形を持った音声
の特徴量をそのまま用いて認識を行なうことは、困難で
ある。
It is known that the formant frequency of speech subjected to the Lombard effect increases, especially in the low range. For example, according to "Study of recognition method considering vocalization deformation under noise", Onkoron, January-October 1999, 2-1-5, when there is 86 dBSPL noise, compared to when there is no noise,
The average formant frequency below 1.5 KHz is approximately 120
Hz is rising. It is difficult to perform recognition by directly using the feature values of voices with such vocalization distortions.

【0004】0004

【目的】本発明は、上述のごとき実情に鑑みてなされた
もので、SN比を良くしようとする発声者の自然な努力
をそのまま用いて、その際に受けるLombard効果
を排除するために、発声者に影響を与える騒音のレベル
を計測し、そのレベルに応じて、Lombard効果に
よって移動したホルマント周波数を補正した音声の特徴
量を得ること、更には、騒音下の音声認識においても、
Lombard効果の影響を受けずに精度の良い認識を
行なうことを目的としてなされたものである。
[Purpose] The present invention has been made in view of the above-mentioned actual circumstances, and is a method for improving the SN ratio by directly using the natural efforts of the speaker to improve the SN ratio while eliminating the Lombard effect. In addition, it is possible to measure the level of noise that affects a person, and obtain voice features that correct formant frequencies shifted by the Lombard effect according to the level.Furthermore, in speech recognition in noise,
This was done with the aim of performing highly accurate recognition without being affected by the Lombard effect.

【0005】[0005]

【構成】本発明は、上記目的を達成するために、(1)
入力信号のレベルを計測するレベル計測手段と、上記入
力信号の周波数分析を行なう周波数分析手段と、上記入
力信号のうちの音声区間を検出する音声区間検出手段と
、上記音声区間検出手段で音声区間が検出されていない
時点の、上記レベル計測手段で計測された入力信号のレ
ベルに基づいて、上記周波数分析手段の係数を設定する
係数設定手段とを具備して成る音声特徴量抽出方式、更
には、(2)音声を入力するためのマイクロホンと、上
記マイクロホンから得られる入力信号から、前記(1)
記載の音声特徴量抽出方式によって音声の特徴量を得る
特徴量抽出部と、上記特徴量抽出部で得られた音声の特
徴量から、入力パターンを得る入力パターン生成部と、
予め登録された音声の標準パターンを記憶する標準パタ
ーンメモリと、上記入力パターン生成部で得られた入力
パターンと上記標準パターンメモリに記憶された標準パ
ターンとで認識を行ない、結果を出力する認識部とを具
備して成る音声認識装置を特徴としたものである。以下
、本発明の実施例に基いて説明する。
[Structure] In order to achieve the above objects, the present invention provides (1)
a level measuring means for measuring the level of an input signal; a frequency analyzing means for performing frequency analysis of the input signal; a voice section detecting means for detecting a voice section of the input signal; a coefficient setting means for setting a coefficient of the frequency analysis means based on the level of the input signal measured by the level measurement means at a time when the frequency analysis means is not detected; , (2) from a microphone for inputting audio and an input signal obtained from the microphone, the above (1)
a feature amount extraction unit that obtains a voice feature amount using the voice feature amount extraction method described above; an input pattern generation unit that obtains an input pattern from the voice feature amount obtained by the feature amount extraction unit;
a standard pattern memory that stores standard patterns of speech registered in advance; and a recognition unit that performs recognition using the input pattern obtained by the input pattern generation unit and the standard pattern stored in the standard pattern memory and outputs the result. The present invention is characterized by a speech recognition device comprising: Hereinafter, the present invention will be explained based on examples.

【0006】図1は、請求項1に記載の発明の一実施例
を説明するための構成図で、レベル計測手段11は、絶
対値演算、ローパスフィルタリング(LPF)を行なっ
て入力信号のレベルを求め、10ms毎に出力する。周
波数分析手段12は、15チャンネルのバンドパスフィ
ルタリング(BPF)、絶対値演算、ローパスフィルタ
リング(LPF)によって、入力信号のスペクトルを求
め、10ms毎に音声の特徴量として出力する。レベル
計測手段11と周波数分析手段12は、デジタルフィル
タで構成され、周波数分析手段12の15チャンネルの
BPFの中心周波数は、係数設定手段14によって設定
される。音声区間検出手段13は、レベル計測手段11
で10ms毎に得られた入力信号のレベルの時系列デー
タから、公知の2しきい値法を用いて音声区間を検出す
る。ここで用いる音声区間検出法は本発明の本質ではな
く、他の方法を用いても良い。
FIG. 1 is a configuration diagram for explaining an embodiment of the invention as claimed in claim 1, in which a level measuring means 11 performs absolute value calculation and low-pass filtering (LPF) to measure the level of an input signal. It is calculated and output every 10ms. The frequency analysis means 12 obtains the spectrum of the input signal by 15 channels of band pass filtering (BPF), absolute value calculation, and low pass filtering (LPF), and outputs it as a voice feature every 10 ms. The level measuring means 11 and the frequency analyzing means 12 are composed of digital filters, and the center frequency of the 15-channel BPF of the frequency analyzing means 12 is set by the coefficient setting means 14. The voice section detection means 13 is the level measurement means 11.
From the time-series data of the level of the input signal obtained every 10 ms, a voice section is detected using a known two-threshold method. The speech interval detection method used here is not the essence of the present invention, and other methods may be used.

【0007】係数設定手段14は、周波数分析手段12
の各BPF毎に3個の中心周波数に対応するそれぞれの
3組の係数を予め用意していて、騒音レベルに応じて、
その係数を変更する。騒音レベルは、音声区間検出手段
13で、音声区間でないとされた区間の騒音レベルの複
数個のデータを平滑化したものを用いる。騒音レベルを
3段階に区別するための騒音の2つのしきい値は、使用
されるマイクロホン位置(より好ましくは、発声者の耳
の位置)で絶対音圧がそれぞれ70dBSPL,80d
BSPLに相当するように予め設定されている。このし
きい値は、当然、使用するシステムのマイクロホンの感
度やマイクアンプのゲインによって異なってくる。ここ
で、各BPFの中心周波数は、(a)騒音レベルが70
dBSPL以下に相当する場合、250Hzから1/3
オクターブ間隔で設定され、(b)騒音レベルが70〜
80dBSPLに相当する場合、(a)の各中心周波数
にそれぞれ50Hzを加えた周波数とし、(c)騒音レ
ベルが80dBSPL以上に相当する場合、(a)の各
中心周波数にそれぞれ100Hzを加えた周波数とする
[0007] The coefficient setting means 14 includes the frequency analysis means 12
Three sets of coefficients corresponding to three center frequencies are prepared in advance for each BPF, and depending on the noise level,
Change its coefficients. The noise level is obtained by smoothing a plurality of pieces of noise level data in a section determined by the voice section detecting means 13 to be not a voice section. The two noise thresholds for distinguishing the noise level into three levels are absolute sound pressures of 70 dBSPL and 80 dBSPL, respectively, at the microphone position used (more preferably at the speaker's ear position).
It is set in advance to correspond to BSPL. This threshold value naturally varies depending on the microphone sensitivity and microphone amplifier gain of the system used. Here, the center frequency of each BPF is (a) when the noise level is 70
If equivalent to dBSPL or less, 1/3 from 250Hz
(b) The noise level is set at octave intervals, and (b) the noise level is 70~
If the noise level corresponds to 80 dBSPL, the frequency shall be the sum of each center frequency in (a) plus 50 Hz, and (c) If the noise level corresponds to 80 dBSPL or more, the frequency shall be the sum of each center frequency in (a) plus 100 Hz. do.

【0008】従って、周囲の騒音レベルが変化し、Lo
mbard効果によりあるホルマントの周波数が変化し
ても、騒音レベルにかかわらず、同じチャンネルでその
ホルマントを捕らえることができるので、周波数分析手
段12で得られるスペクトル、即ち、音声の特徴量は、
Lombard効果の影響を補正されることになる。こ
の実施例の(b)、(c)の設定では、全てのBPFの
中心周波数を一律に上昇させているが、基本となる(a
)の設定が、LOGスケールで等間隔になっているので
、高域においては、設定変更の影響は少ないし、また、
各BPFを独立に設定変更可能なので一律に上昇させる
必要も特にない。この実施例では、各BPFの中心周波
数を3組ずつ用意して設定し直しているが、計算のため
の処理量が問題にならないシステムでは、一定期間毎に
予め定められた式から最適な中心周波数を計算して設定
すれば、より精度の良い音声の特徴量が得られる。上記
本発明の特徴量抽出方式によれば、中心周波数の変更の
みならず、一部のBPFのゲインを変えることになり、
周波数特性をも同時に変更することもできる。Lomb
ard効果は、ホルマント周波数の移動のみならず、特
定の帯域のエネルギの上昇という現象があることも知ら
れており、本発明は、この補正に利用することができ、
この補正方法も本発明の範囲に含まれるものである。
[0008] Therefore, the ambient noise level changes and Lo
Even if the frequency of a certain formant changes due to the mbard effect, that formant can be captured in the same channel regardless of the noise level, so the spectrum obtained by the frequency analysis means 12, that is, the feature amount of the voice, is
The influence of the Lombard effect will be corrected. In the settings (b) and (c) of this example, the center frequencies of all BPFs are raised uniformly, but the basic (a)
) settings are at equal intervals on the LOG scale, so changing settings has little effect on high frequencies, and
Since the settings of each BPF can be changed independently, there is no particular need to uniformly increase the BPF. In this example, three sets of center frequencies for each BPF are prepared and reset. However, in a system where the amount of calculation processing is not an issue, the optimum center frequency is By calculating and setting the frequency, more accurate voice features can be obtained. According to the above feature extraction method of the present invention, not only the center frequency is changed, but also the gain of some BPFs is changed.
It is also possible to change the frequency characteristics at the same time. Lomb
It is known that the ard effect involves not only a shift in formant frequency but also an increase in energy in a specific band, and the present invention can be used for this correction.
This correction method is also included within the scope of the present invention.

【0009】図2は、請求項2の発明の一実施例の構成
を示す図で、図中、マイクロホン1は、音声を入力し、
電気信号に変換し、前処理部2は、信号の増幅と高域遮
断を行なう。A/D変換部3は、アナログの信号を、1
6kHzのサンプリングで12ビットのデジタル値に変
換する。特徴量抽出部10は、前述の特徴量抽出を実行
するもので、この中のレベル計測手段11、周波数分析
手段12、音声区間検出手段13、係数設定手段14は
、図1の例と同様である。入力パターン生成部4は、音
声区間検出手段13で音声区間とされた区間の、特徴量
抽出部10で得られた音声の特徴量から、入力パターン
を得る。標準パターンメモリ5は、予め登録された音声
の標準パターンを記憶する。認識部6は、入力パターン
生成部4で得られた入力パターンと標準パターンメモリ
5に記憶された標準パターンとで認識を行ない、結果を
出力する。ここで、入力パターン生成部4、標準パター
ンメモリ5、認識部6の動作は、「2値のTSPを用い
た単語音声認識システム」電学論C、108巻、昭63
−10月、pp.858−865などで公知のBTST
方式音声認識技術が用いられるが、他の方法でも良い。
FIG. 2 is a diagram showing the configuration of an embodiment of the invention according to claim 2. In the figure, a microphone 1 inputs audio,
The preprocessor 2 converts the signal into an electrical signal, and amplifies the signal and cuts off the high frequency range. The A/D converter 3 converts the analog signal into one
Convert to 12-bit digital value with 6kHz sampling. The feature quantity extraction unit 10 executes the above-mentioned feature quantity extraction, and the level measurement means 11, frequency analysis means 12, voice section detection means 13, and coefficient setting means 14 are the same as in the example of FIG. be. The input pattern generation section 4 obtains an input pattern from the voice feature amount obtained by the feature amount extraction section 10 of the section determined as a voice section by the voice section detection means 13. The standard pattern memory 5 stores standard patterns of voices registered in advance. The recognition unit 6 performs recognition using the input pattern obtained by the input pattern generation unit 4 and the standard pattern stored in the standard pattern memory 5, and outputs the result. Here, the operations of the input pattern generation section 4, standard pattern memory 5, and recognition section 6 are as follows: "Word speech recognition system using binary TSP", Electrical Engineering Theory C, Vol. 108, 1983.
-October, pp. BTST known as 858-865 etc.
Although a standard voice recognition technique is used, other methods may also be used.

【0010】図3は、図2に示した実施例の構成をハー
ドウェアにした場合の一例の図で、図中、アナログ部2
1は、図2の前処理部2に対応し、A/D変換器22は
、A/D変換部2に対応する。DPS23は、特徴量抽
出部10を含み、ROM25は、標準パターンメモリ5
のデータと、入力パターン生成部4のプログラム、認識
部6のプログラムとを含む。CPU27は、入力パター
ン生成部4のプログラム、認識部6のプログラムなどを
実行する。RAM26は、プログラム実行のためのワー
クメモリとして使用される。I/Oボート28は、RS
232Cボートで、認識結果の出力など外部との通信を
行なうものである。
FIG. 3 is a diagram showing an example in which the configuration of the embodiment shown in FIG. 2 is implemented as hardware.
1 corresponds to the preprocessing section 2 in FIG. 2, and the A/D converter 22 corresponds to the A/D converting section 2. The DPS 23 includes a feature extraction unit 10, and the ROM 25 includes a standard pattern memory 5.
data, a program for the input pattern generation section 4, and a program for the recognition section 6. The CPU 27 executes the program of the input pattern generation section 4, the program of the recognition section 6, and the like. RAM 26 is used as a work memory for program execution. The I/O boat 28 is an RS
The 232C port is used to communicate with the outside, such as outputting recognition results.

【0011】[0011]

【効果】請求項1に対応する効果;発声者に影響を与え
る騒音のレベルを計測し、そのレベルに応じて、周波数
分析フィルタの係数を変更しているので、Lombar
d効果による発声のスペクトル変形を補正した音声の特
徴量を得ることができる。また、周波数分析の時点で補
正が行なわれているので、得られた特徴量に補正のため
の演算処理を行なうことを必要とせず、その特徴量から
直ちに認識処理を行なうことができる。請求項2に対応
する効果;Lombard効果による発声のスペクトル
変形を補正する音声の特徴量を求めているので、騒音下
の音声認識においても、Lombard効果の影響を受
けずに精度の良い認識を行なうことができる。
[Effect] Effect corresponding to claim 1; Since the level of noise that affects the speaker is measured and the coefficients of the frequency analysis filter are changed according to the level, Lombar
It is possible to obtain the feature amount of the voice that corrects the spectrum deformation of the voice due to the d effect. Further, since the correction is performed at the time of frequency analysis, it is not necessary to perform arithmetic processing for correction on the obtained feature amount, and recognition processing can be performed immediately from the feature amount. Effect corresponding to claim 2: Since a voice feature quantity that corrects the spectral deformation of vocalization due to the Lombard effect is obtained, accurate recognition is performed without being affected by the Lombard effect even in voice recognition in noise. be able to.

【図面の簡単な説明】[Brief explanation of the drawing]

【図1】  請求項1に記載の発明の一実施例を説明す
るための構成図である。
FIG. 1 is a configuration diagram for explaining an embodiment of the invention according to claim 1.

【図2】  請求項2に記載の発明の一実施例を説明す
るための構成図である。
FIG. 2 is a configuration diagram for explaining an embodiment of the invention according to claim 2.

【図3】  図3に示した実施例をハードウェアにした
場合の構成例を示す図である。
FIG. 3 is a diagram showing an example of a configuration in which the embodiment shown in FIG. 3 is implemented as hardware.

【符号の説明】[Explanation of symbols]

1…マイクロフォン、2…前処理部、3…A/D変換部
、4…入力パターン生成部、5…標準パターンメモリ、
6…認識部、10…特徴量抽出部、11…レベル計測手
段、12…周波数分析手段、12…音声区間検出手段。
DESCRIPTION OF SYMBOLS 1... Microphone, 2... Preprocessing part, 3... A/D conversion part, 4... Input pattern generation part, 5... Standard pattern memory,
6... Recognition section, 10... Feature amount extraction section, 11... Level measurement means, 12... Frequency analysis means, 12... Voice section detection means.

Claims (2)

【特許請求の範囲】[Claims] 【請求項1】  入力信号のレベルを計測するレベル計
測手段と、上記入力信号の周波数分析を行なう周波数分
析手段と、上記入力信号のうちの音声区間を検出する音
声区間検出手段と、上記音声区間検出手段で音声区間が
検出されていない時点の、上記レベル計測手段で計測さ
れた入力信号のレベルに基づいて、上記周波数分析手段
の係数を設定する係数設定手段とを具備して成ることを
特徴とする音声特徴量抽出方式。
1. Level measurement means for measuring the level of an input signal; frequency analysis means for frequency analysis of the input signal; voice section detection means for detecting a voice section of the input signal; and coefficient setting means for setting the coefficients of the frequency analysis means based on the level of the input signal measured by the level measurement means at a time when no voice section is detected by the detection means. A method for extracting audio features.
【請求項2】  音声を入力するためのマイクロホンと
、上記マイクロホンから得られる入力信号から、請求項
1記載の音声特徴量抽出方式によって音声の特徴量を得
る特徴量抽出部と、上記特徴量抽出部で得られた音声の
特徴量から、入力パターンを得る入力パターン生成部と
、予め登録された音声の標準パターンを記憶する標準パ
ターンメモリと、上記入力パターン生成部で得られた入
力パターンと上記標準パターンメモリに記憶された標準
パターンとで認識を行ない、結果を出力する認識部とを
具備して成ることを特徴とする音声認識装置。
2. A microphone for inputting sound; a feature extraction unit for obtaining a sound feature by the sound feature extraction method according to claim 1 from an input signal obtained from the microphone; an input pattern generation section that obtains an input pattern from the feature amount of the voice obtained in the section; a standard pattern memory that stores a standard pattern of voice registered in advance; A speech recognition device comprising: a recognition unit that performs recognition using a standard pattern stored in a standard pattern memory and outputs a result.
JP3143760A 1991-05-20 1991-05-20 Speech feature extraction method and speech recognition device Pending JPH04343399A (en)

Priority Applications (1)

Application Number Priority Date Filing Date Title
JP3143760A JPH04343399A (en) 1991-05-20 1991-05-20 Speech feature extraction method and speech recognition device

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
JP3143760A JPH04343399A (en) 1991-05-20 1991-05-20 Speech feature extraction method and speech recognition device

Publications (1)

Publication Number Publication Date
JPH04343399A true JPH04343399A (en) 1992-11-30

Family

ID=15346389

Family Applications (1)

Application Number Title Priority Date Filing Date
JP3143760A Pending JPH04343399A (en) 1991-05-20 1991-05-20 Speech feature extraction method and speech recognition device

Country Status (1)

Country Link
JP (1) JPH04343399A (en)

Cited By (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
WO2000023984A1 (en) * 1998-10-16 2000-04-27 Dragon Systems Uk Research & Development Limited Speech processing

Cited By (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
WO2000023984A1 (en) * 1998-10-16 2000-04-27 Dragon Systems Uk Research & Development Limited Speech processing

Similar Documents

Publication Publication Date Title
EP1208563B1 (en) Noisy acoustic signal enhancement
JP2962732B2 (en) Hearing aid signal processing system
US9591410B2 (en) Hearing assistance apparatus
US6718301B1 (en) System for measuring speech content in sound
JP2561850B2 (en) Voice processor
US9454976B2 (en) Efficient discrimination of voiced and unvoiced sounds
JPH0968997A (en) Audio processing method and apparatus
US7424119B2 (en) Voice matching system for audio transducers
EP1229517B1 (en) Method for recognizing speech with noise-dependent variance normalization
US5944672A (en) Digital hearing impairment simulation method and hearing aid evaluation method using the same
JP3420831B2 (en) Bone conduction voice noise elimination device
CN113963699A (en) Intelligent voice interaction method for financial equipment
CN115312071A (en) Voice data processing method and device, electronic equipment and storage medium
JPH0424692A (en) Voice section detection method
JPH0416900A (en) Speech recognition device
JPH0573090A (en) Speech recognizing method
JPH05273964A (en) Attack time detection device used for automatic music transcription device etc.
JP4856559B2 (en) Received audio playback device
JPH0675596A (en) Speech and acoustic phenomenon analysis device
JPH02178699A (en) Voice recognition device
JPH03147000A (en) Voice input device
JPH02272499A (en) Voice recognizing device
KR100531776B1 (en) How to set the gain of the amplifier according to the user
JPS6355280B2 (en)
JPH02165198A (en) Voice recognizing device