JPH0348900A - Sound detecting device - Google Patents
Sound detecting deviceInfo
- Publication number
- JPH0348900A JPH0348900A JP1183684A JP18368489A JPH0348900A JP H0348900 A JPH0348900 A JP H0348900A JP 1183684 A JP1183684 A JP 1183684A JP 18368489 A JP18368489 A JP 18368489A JP H0348900 A JPH0348900 A JP H0348900A
- Authority
- JP
- Japan
- Prior art keywords
- parameter
- frame
- noise
- feature
- parameters
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Granted
Links
Landscapes
- Time-Division Multiplex Systems (AREA)
- Transmission Systems Not Characterized By The Medium Used For Transmission (AREA)
Abstract
Description
【発明の詳細な説明】
[発明の目的]
(産業上の利用分野)
本発明は、ATM (Asynchronous T
ransfer Mode)通信、DSI(Digi
tal 5peech Interplation
)、パケット通信、音声認識の分野に適用され、音声信
号中の有音区間を精度良く検出する有音検出装置に関す
る。[Detailed Description of the Invention] [Object of the Invention] (Industrial Application Field) The present invention provides an ATM (Asynchronous T
transfer mode) communication, DSI (Digi
tal 5peech Interpretation
), it is applied to the fields of packet communication and voice recognition, and relates to a sound detection device that accurately detects sound intervals in a voice signal.
(従来の技術) 第10図は従来の有音検出装置の一構成を示している。(Conventional technology) FIG. 10 shows a configuration of a conventional sound detection device.
入力端子100に入力された音声信号中から電力、零交
差数、自己相関関数、スペクトルなどの特徴パラメータ
がフレーム単位で特徴パラメータ計算器101によって
計算される。Feature parameters such as power, number of zero crossings, autocorrelation function, spectrum, etc. are calculated by the feature parameter calculator 101 on a frame-by-frame basis from the audio signal input to the input terminal 100.
計算された特徴パラメータは、マツチング器102へ出
力され、予め設定された有音標準パターン103及び雑
音標準パターン104と比較し、それぞれの距離が算出
される。The calculated feature parameters are output to the matching device 102, and compared with a preset voice standard pattern 103 and a noise standard pattern 104, and the respective distances are calculated.
もし、特徴パラメータと有音標準パターン103の距離
が特徴パラメータと雑音パターン104との距離よりも
小さければ、入力フレームは有音に属し、反対であれば
雑音に属すると判定され、その判定結果が出力端子10
5から出力される。If the distance between the feature parameter and the sound standard pattern 103 is smaller than the distance between the feature parameter and the noise pattern 104, it is determined that the input frame belongs to the sound pattern, and if the opposite is true, it is determined that the input frame belongs to the noise, and the determination result is Output terminal 10
Output from 5.
(発明が解決しようとする課題)
しかしながら、有音であっても子音の電力は母音と異な
り背景雑音の電力を下回ることが多い。(Problems to be Solved by the Invention) However, unlike vowels, the power of consonants is often lower than the power of background noise even if they are voiced.
このため、背景雑音が大きい環境下では、子音区間の特
徴パラメータに背景雑音の特徴が大きく出てしまう。For this reason, in an environment with large background noise, the characteristics of the background noise will appear significantly in the characteristic parameters of the consonant section.
上記従来の有音検出装置によれば、背景雑音の影響を受
けた特徴パラメータをそのまま判定に用いていたので、
背景雑音が大きい場合には、子音の検出誤りが多くなっ
ていた。According to the conventional sound detection device described above, the feature parameters affected by background noise are used as they are for determination.
When the background noise was large, there were many consonant detection errors.
このことによって、通信の分野では音質の劣化の要因と
なり、また、音声認識の分野で認識率の低下を招いてい
た。This has caused a deterioration in sound quality in the field of communications, and has caused a decline in recognition rates in the field of speech recognition.
本発明は上記事情に鑑みてなされたものであり、その目
的は、背景雑音が大きい場合にあっても有音の検出精度
を向上することができる音声検出装置を提供することに
ある。The present invention has been made in view of the above circumstances, and an object of the present invention is to provide a voice detection device that can improve the accuracy of detecting presence of voice even when background noise is large.
[発明の構成]
(課題を解決するための手段)
上記課題を解決するために、第1の発明は、入力された
音声信号中からフレーム単位で音声の特徴パラメータを
演算する手段と、
演算されたフレーム単位の特徴パラメータを全であるい
は雑音区間の特徴パラメータを順次蓄積するバッファと
、
蓄積された特徴パラメータの内、入力フレームからSフ
レーム前の特徴パラメータを基準として過去のNフレー
ム分の特徴パラメータ集合を取り出し、前記入力フレー
ムの特徴パラメータとの距離ベクトル若しくはベクトル
のノルムを演算することによって入力フレームの特徴パ
ラメータの変換パラメータを生成する手段と、
前記変換パラメータと予め設定された標準パターンとを
比較して前記音声信号中の有音区間を判別する手段と、
を有することを特徴とする。[Structure of the Invention] (Means for Solving the Problems) In order to solve the above problems, the first invention provides means for calculating voice characteristic parameters on a frame-by-frame basis from an input voice signal; A buffer for sequentially accumulating all feature parameters for each frame or feature parameters for a noise interval, and a buffer for sequentially accumulating feature parameters for each frame or for a noise interval, and a buffer for sequentially accumulating feature parameters for the past N frames based on the feature parameters for S frames before the input frame from among the accumulated feature parameters. means for generating transformation parameters for the feature parameters of the input frame by extracting a set and calculating a distance vector or norm of the vector with the feature parameters of the input frame; and comparing the transformation parameters with a preset standard pattern. and means for determining a sound interval in the audio signal.
また、第2の発明は、入力された音声信号中からフレー
ム単位で音声の特徴パラメータを演算する手段と、
演算されたフレーム単位の特徴パラメータを順次蓄積す
るバッファと、
蓄積された特徴パラメータから回帰係数を演算し、入力
フレームの特徴パラメータの変換パラメータを生成する
手段、
前記変換パラメータと予め設定された標準パターンと比
較して前記音声信号中の有音区間を判断する手段と、
を有することを特徴とする。The second invention also provides a means for calculating audio feature parameters in frame units from an input audio signal, a buffer for sequentially accumulating the calculated feature parameters in frame units, and regression from the accumulated feature parameters. means for calculating a coefficient to generate a conversion parameter of a characteristic parameter of an input frame; and means for comparing the conversion parameter with a preset standard pattern to determine a sound section in the audio signal. Features.
(作用)
以上の構成において、第1の発明ではフレーム単位で演
算された特徴パラメータの全であるいは雑音領域での特
徴パターンの任意のNフレーム分の集合から入力フレー
ムの特徴パラメータとの距離ベクトル若しくはベクトル
のノルムを演算してこれらを変換パラメータとする。そ
して、この変換パラメータと標準パターンとを比較する
ことにより、雑音の影響を回避した有音判別がされる。(Operation) In the above configuration, in the first invention, the distance vector or Calculate the norm of the vector and use these as transformation parameters. Then, by comparing this conversion parameter with a standard pattern, the presence of a voice is determined while avoiding the influence of noise.
また、第2の発明では蓄積された特徴パラメータから回
帰係数を求め、この回帰係数から変換パラメータを求め
る。そして、この変換パラメータと標準パターンとを比
較することにより、同様に雑音の影響を回避した有音判
別がされる。Furthermore, in the second invention, a regression coefficient is determined from the accumulated feature parameters, and a conversion parameter is determined from this regression coefficient. Then, by comparing this conversion parameter with the standard pattern, the sound presence is determined while avoiding the influence of noise.
(実施例)
第1図は本発明に係る有音検出装置の概略的構成を示す
ブロック図であり、この装置は、特徴パラメータ計算器
1と、特徴パラメータ変換器2と、有音判定器3とから
構成される。(Embodiment) FIG. 1 is a block diagram showing a schematic configuration of a sound detection device according to the present invention, which includes a feature parameter calculator 1, a feature parameter converter 2, and a sound presence determiner 3. It consists of
なお、以下の各実施例では、音声信号をフレーム単位に
分析し有無・音声の判定を行なっていく。In each of the following embodiments, the audio signal is analyzed frame by frame to determine the presence or absence of audio.
例えば、音声信号を8KHzでサンプリングし、160
サンプルづつまとめて1フレームとする。For example, if an audio signal is sampled at 8KHz and
Each sample is combined into one frame.
ただし、フレーム長は、常に一定長である必要はない。However, the frame length does not always have to be a constant length.
特徴パラメータ計算器1では、フレーム単位にDurb
in法などを用いて線形予測係数を計算する。ここで、
線形予測係数からPARCOR係数、LPCケプストラ
ム、メルケプストラム等を計算し、特徴パラメータとし
てもよい。また、電力、自己相関関数、零交差数、等も
計算してもよい。The feature parameter calculator 1 calculates Durb for each frame.
Calculate linear prediction coefficients using the in method or the like. here,
A PARCOR coefficient, LPC cepstrum, mel cepstrum, etc. may be calculated from the linear prediction coefficients and used as feature parameters. Power, autocorrelation function, number of zero crossings, etc. may also be calculated.
現在有音か無音かを判定しようとしているフレームを以
下では入力フレームという。また、特徴パラメータ計算
器1で得られた入力フレームの特徴パラメータをX(n
)とする。nはフレームのシーケンシャルな番号である
。特徴パラメータは、p次元のベクトルで、次の(1)
の式で書き表わされる。The frame for which it is currently being determined whether there is a sound or no sound is hereinafter referred to as an input frame. In addition, the feature parameters of the input frame obtained by the feature parameter calculator 1 are set to X(n
). n is the sequential number of the frame. The feature parameter is a p-dimensional vector, as shown in (1) below.
It is written as the formula.
X(n) = (xt (n)、x2(n)s=、
XP (n))・・・(1)
特徴パラメータ変換器2では、音声と雑音の違いを強調
するために特徴パラメータを変換する。X(n) = (xt (n), x2(n)s=,
XP (n)) (1) The feature parameter converter 2 converts feature parameters in order to emphasize the difference between speech and noise.
ここで変換された特徴パラメータを、以下では変換パラ
メータと呼び、次の(2)式で書き表わせる。The feature parameters converted here are hereinafter referred to as conversion parameters, and can be expressed by the following equation (2).
変化パラメータはr(≦P)次元のベクトルである。The change parameter is an r (≦P)-dimensional vector.
Y(n) = (yt (n)、y2(n)、=−、y
r (n))・・・(2)
第2図は本発明の一実施例の特徴部分である特徴パラメ
ータ変換器2の構成を示すブロック図であり、本実施例
は背景雑音の影響を回避するため雑音区間の特徴パラメ
ータを蓄積し、その蓄積データから入力フレームの特徴
パラメータの距離ベクトルを変換パラメータY (n)
として求めた後、変換パラメータY(n)を標準パター
ンと比較して有音区間を判別するものである。Y(n) = (yt (n), y2(n), =-, y
r (n))...(2) FIG. 2 is a block diagram showing the configuration of the feature parameter converter 2, which is a characteristic part of an embodiment of the present invention, and this embodiment avoids the influence of background noise. In order to
After determining the conversion parameter Y(n), the conversion parameter Y(n) is compared with a standard pattern to determine a sound interval.
本実施例の特徴パラメータ変換器2は、電力測定器4と
、雑音判定器5と、スイッチ(SW)6と、バッファ7
と、距離ベクトル計算器8とから構成されている。The feature parameter converter 2 of this embodiment includes a power measuring device 4, a noise determiner 5, a switch (SW) 6, and a buffer 7.
and a distance vector calculator 8.
電力測定器4では、フレーム単位に次の(3)式で平均
電力Pを測定する。The power measuring device 4 measures the average power P on a frame-by-frame basis using the following equation (3).
フレーム内の音声信号のサンプルをa(1)(1−0,
1、N−1) 、1フレームのサンプル数をNとすると
、
平均電力P = ”? a(1)2/ N −
(3)雑音判定器5では、入力信号の中から、確実に雑
音であるという区間を検出するためにあらかじめ、与え
られているしきい値Tの電力測定器で測定した電力Pと
比較する。The samples of the audio signal in the frame are a(1)(1-0,
1, N-1), and the number of samples in one frame is N, then the average power P = "? a(1)2/ N -
(3) The noise determiner 5 compares the input signal with the power P measured with a power measuring device having a given threshold value T in order to detect a section that is definitely noise from the input signal.
もし、P≧Tならば雑音でないと判定し“0”SW6に
出力する。If P≧T, it is determined that it is not noise and outputs “0” to SW6.
そうでなければ雑音と判定し“1“をSW6に出力する
。If not, it is determined to be noise and "1" is output to SW6.
SW6は、雑音判定器の出力が“1”ならば、バッファ
42にそのフレームの特徴パラメータヲ記憶させる。If the output of the noise determiner is "1", SW6 causes the buffer 42 to store the characteristic parameters of that frame.
バッファ7では、特徴パラメータがバッファ7に蓄積さ
れる時間の順序関係を保存するために、特徴パラメータ
がバッファに入力された順番で、バッファのヘッドから
テイルに向かって蓄積する。In the buffer 7, the feature parameters are accumulated from the head of the buffer to the tail in the order in which they were input to the buffer, in order to preserve the order of the time in which the feature parameters are accumulated in the buffer 7.
すなわち、一番新しい特徴パラメータ(現在判定すべき
フレームの特徴パラメータ)をバッファのヘッドに、一
番過去の特徴パラメータをテイルに蓄積する。第3図に
はバッファ7の構成例が示されている。That is, the newest feature parameter (the feature parameter of the frame to be currently determined) is stored in the head of the buffer, and the oldest feature parameter is stored in the tail. FIG. 3 shows an example of the configuration of the buffer 7.
距離ベクトル計算器8では、バッファ7に蓄積された特
徴パラメータのうち、入力フレームのSフレーム前(バ
ッファのヘッドからSフレームめ)からバッファのテイ
ルに向かってNフレーム分の特徴パラメータ集合Ωを取
り出し、Ωと入力フレームの特徴パラメータX(n)と
の距離ベクトルを計算する。The distance vector calculator 8 extracts a feature parameter set Ω for N frames from S frames before the input frame (S frames from the head of the buffer) to the tail of the buffer from among the feature parameters accumulated in the buffer 7. , Ω and the feature parameter X(n) of the input frame.
なお、前記Sフレーム、Nフレームは任意の数フレーム
を取り得るが、数フレームから20フレ一ム程度が好適
である。Note that the S frame and the N frame can be any number of frames, but preferably several frames to about 20 frames.
まず、バッファ7内で、入力フレームのSフレーム前の
フレームから数えてNフレーム過去の特徴パラメータの
集合Ωを第3図に示すように、Ω: XL (0)、X
L (1)、・・・、 X、 (N−1,) +とする
。First, in the buffer 7, a set Ω of feature parameters for N frames past, counting from the frame S frames before the input frame, is expressed as Ω: XL (0), X
Let L (1), ..., X, (N-1,) +.
X(n)とΩの雑音距離ベクトルをY (n)とする。Let the noise distance vector between X(n) and Ω be Y(n).
Y(n)は、Ωの平均値MとX(n)の差をΩの標準偏
差で正規化したものであり、次の(4)式で表される。Y(n) is the difference between the average value M of Ω and X(n) normalized by the standard deviation of Ω, and is expressed by the following equation (4).
第4図にはX(n)、Ω、Y(n)の関係が図示されて
いる。FIG. 4 shows the relationship between X(n), Ω, and Y(n).
Y(n) = (y+ (n)、y2(n)、−、3’
q (n))・・・(4)
M” (mt 1m2* ・”1mq ) ・・・
(5)とすると、
Y+ (n) −(x+ (n)、mt ) /at
−(6)・・・(8)
ここで、1=I−,2,−、q、 r > 1のときq
=rで、変換パラメータはY (n)である。Y(n) = (y+ (n), y2(n), -, 3'
q (n))...(4) M" (mt 1m2* ・"1mq)...
(5), then Y+ (n) −(x+ (n), mt ) /at
-(6)...(8) Here, 1=I-, 2,-, q, when r > 1, q
= r and the transformation parameter is Y (n).
r=1のとき、qは1≦q≦pの任意の大きさを取るこ
とができる。特に、r=1、p>lの場合には、11・
11をベクトルのノルムとすると、変換パラメータをI
IY(11) +1.!1.する。When r=1, q can take any size in the range 1≦q≦p. In particular, when r=1 and p>l, 11.
11 is the norm of the vector, then the transformation parameter is I
IY(11) +1. ! 1. do.
有音判定器3では、特徴パラメータ変換器2から得られ
た変換パラメータをもとに、有音区間を判定する。この
有音判定器3は第5図に示すように、マツチング器9と
、M個の標準パターン10とから構成されている。The sound determination unit 3 determines a sound interval based on the conversion parameters obtained from the feature parameter converter 2. As shown in FIG. 5, the utterance determiner 3 is composed of a matching device 9 and M standard patterns 10.
マツチング器9では、標準パターンと変換パラメータの
距離を測定し、音声に属する標準パターンにマツチング
された場合音声、そうでない場合無音と判定する。The matching device 9 measures the distance between the standard pattern and the conversion parameter, and determines that if the standard pattern is matched with the standard pattern belonging to audio, it is audio, and if not, it is determined to be silence.
まず、次式より各標準パターンω、 (i=1゜・・
・、M)との距離を測定する。First, each standard pattern ω, (i=1°...
・, M).
DI (Y)
= (Y−μ()′ Σ+ −’ (Y−μ、)
+101 Σ 1
・・・(9)Yは
i =min D + (Y)
・・・(10)なるω1にYが属しているとする。もし
ω1が音声に属していれば、そのフレームは有音、ω、
が雑音に属していれば、そのフレームは雑音であると判
定する。DI (Y) = (Y−μ()′ Σ+ −′ (Y−μ,)
+101 Σ 1
...(9) Y is i = min D + (Y)
...(10) Suppose that Y belongs to ω1. If ω1 belongs to voice, then the frame is voiced, ω,
If the frame belongs to noise, the frame is determined to be noise.
標準パターン10は以下のように定義できる。Standard pattern 10 can be defined as follows.
標準パターン10はマツチング器9の構成から分は標準
パターンのクラスを示すiを簡易のため省略する。The standard pattern 10 is based on the configuration of the matching device 9, and i, which indicates the class of the standard pattern, is omitted for the sake of simplicity.
クラスωに属するL個のr次元変換パラメータをY、
(j)−(y−t(j)、y 、2(j)、・・・、
y、、(j))・・・(11)
j−1,2,・・・、L
とする。Let L r-dimensional transformation parameters belonging to class ω be Y,
(j)-(y-t(j), y, 2(j),...
y,, (j))...(11) Let j-1, 2,..., L.
また、μとΣの各要素をμ1、Σ1とすると、第5図の
平均・共分散行列計算器8で次式を用いて計算される。Further, assuming that the respective elements of μ and Σ are μ1 and Σ1, the calculation is performed by the mean/covariance matrix calculator 8 in FIG. 5 using the following equation.
((yWt(j) −μL) ・・・(13)標準
パターン(標準パターント・・標準パターンM)の作成
は以下のようにして行なう。((yWt(j) - μL) (13) The standard pattern (standard pattern M) is created as follows.
標準パターンを作成するため、各標準パターンのクラス
ごとに、音声データを作成する。To create standard patterns, audio data is created for each class of each standard pattern.
その作成方法は、まず、複数の被検者に各クラスに属す
る音韻を発音してもらい、それを録音する。The method for creating it is to first have multiple subjects pronounce the phonemes belonging to each class, and then record them.
このようにして得られた音声信号に対し、フレーム単位
に、子音と雑音の区別をつけるためのラベルを付してい
く。ラベル付けは、音声信号の波形やスペクトルをCR
Tに表示して、それを見ながらフレーム単位にラベルを
付けていく。A label is attached to the audio signal obtained in this manner on a frame-by-frame basis to distinguish between consonants and noise. Labeling is done by CR of the waveform and spectrum of the audio signal.
T, and label each frame while looking at it.
ラベル付けされた音声データベース11から、第6図に
示すように、変換パラメータを計算する。From the labeled speech database 11, conversion parameters are calculated as shown in FIG.
このとき、ラベルデータベース12を参照して、そのク
ラスに属するフレームの変換パラメータならば、平均・
共分散行列計算器13で、平均・共分散を計算する。At this time, referring to the label database 12, if the conversion parameter of the frame belonging to that class is the average
A covariance matrix calculator 13 calculates the average and covariance.
このようにして作成された標準パターンと前述のように
して求められた変換パラメータとが比較されて雑音の影
響を回避した有音区間が判別されるのである。The standard pattern created in this manner is compared with the conversion parameters determined as described above, and a sound section that avoids the influence of noise is determined.
なお、前記実施例では予め雑音判定をしてこの区間の特
徴パラメータと蓄積するため、電力測定機4、雑音判定
器5及びスイッチ6とを用いる構成としたが、雑音判定
をしない場合には、全ての区間の特徴パラメータをバッ
ファ7へ蓄積するようにすればよい。この場合は、これ
ら電力測定器4、雑音判定器5及びスイッチ6は不要で
ある。In the above embodiment, the power measuring device 4, the noise determining device 5, and the switch 6 are used in order to determine the noise in advance and store it as the characteristic parameter of this section. However, when the noise is not determined, The feature parameters of all sections may be stored in the buffer 7. In this case, the power measuring device 4, noise determining device 5, and switch 6 are unnecessary.
また、雑音を判定するために電力測定器4及び雑音判定
器5を用いたが、第5図に示した有音判定器と同様に、
雑音の標準パターンとマツチング器を用いて雑音区間を
判別するようにしてもよい。In addition, a power measuring device 4 and a noise determining device 5 were used to determine noise, but similar to the presence determining device shown in FIG.
The noise section may be determined using a standard pattern of noise and a matching device.
第7図は本発明の他の実施例の特徴部分である特徴パラ
メータ変換器2の構成を示すブロック図であり、本実施
例は背景雑音の影響を回避するために、雑音区間の特徴
パラメータを蓄積し、その蓄積データから回避係数と演
算し、この回帰係数から変換パラメータを求めて有音区
間を判別するものである。FIG. 7 is a block diagram showing the configuration of the feature parameter converter 2, which is a characteristic part of another embodiment of the present invention. The data is accumulated, an avoidance coefficient is calculated from the accumulated data, and a conversion parameter is determined from this regression coefficient to determine a sound interval.
バッファ14は、前記バッファ7と同一構成であり、特
徴パラメータをフレーム単位で順次蓄積する。The buffer 14 has the same configuration as the buffer 7, and sequentially accumulates feature parameters on a frame-by-frame basis.
回帰係数計算器15では、以下のようにして回帰係数及
び変換パラメータを求める。The regression coefficient calculator 15 calculates regression coefficients and conversion parameters as follows.
入力フレームの特徴パラメータをX(n) 、バッファ
内のに番目のフレームのp次元特徴パラメータベクトル
の各要素X+(k)が(n−(N−1))≦に≦nの範
囲で与えられたとき、各要素x+(k)の(n−(N−
1))≦に≦nでの近似関数が、f + (k)−a
to + a ILK” + a 12に2十・・・+
a、、に’ ・・・(14)で表されると
する。これを第8図に示す。Let the feature parameters of the input frame be X(n), and each element X+(k) of the p-dimensional feature parameter vector of the th frame in the buffer is given in the range of (n-(N-1))≦≦n. Then, each element x+(k) has (n-(N-
1)) ≦ to ≦n, the approximation function is f + (k)-a
to + a ILK” + a 12 to 20...+
Assume that a, , ni'... is expressed as (14). This is shown in FIG.
各点におけるx+(k)とf+(k)の差、すなわち残
差
e+ (k) =f+ (k) −Xi (k)
・(t5)の2乗和をJlとすると、J、は
J、= Σ e l 2(k)
となる。このJlを最小となるように各回帰係数aII
)+ ・・・、a、を定める。The difference between x+(k) and f+(k) at each point, that is, the residual e+ (k) = f+ (k) −Xi (k)
- If the sum of squares of (t5) is Jl, J becomes J, = Σ e l 2(k). Each regression coefficient aII is set so that this Jl is minimized.
) + ..., a, is determined.
入力フレームの変換パラメータは、次の(17)式%式
%
(17)
または複数のiの組合せである。例えば、Y (n)
” (az+ a21+ ”’+ arl+ a
12+a 22.−、 a r2) ・・・(
[1)となる。The conversion parameter of the input frame is the following formula (17) % formula % (17) or a combination of a plurality of i's. For example, Y (n)
” (az+ a21+ ”'+ arl+ a
12+a 22. -, a r2) ...(
[1] becomes.
このようにして求められた変換パラメータY(n)は有
音判定器3で標準パターンと比較され有音区間が判定さ
れる。なお、有音区間の判定は先の実施例と同一の過程
によりなされるので、説明は省略する。The conversion parameter Y(n) obtained in this manner is compared with a standard pattern by the sound presence determination unit 3 to determine the sound interval. It should be noted that the determination of the voiced section is performed by the same process as in the previous embodiment, so the explanation will be omitted.
以上の各実施例の効果を具体的な測定結果を基に説明す
る。The effects of each of the above embodiments will be explained based on specific measurement results.
母音と異なり、子音の電力は背景音電力を下回ることが
多い。そのため、背景雑音が大きな環境では、子音区間
でも特徴パラメータに雑音の特徴が大きく出てしまう。Unlike vowels, the power of consonants is often lower than the background sound power. Therefore, in an environment with large background noise, the characteristics of the noise will appear significantly in the feature parameters even in consonant intervals.
従来の方式では、背景雑音の影響を受けた特徴パラメー
タをそのまま判定に用いていたなめ、背景雑音が大きな
場合には、子音の検出誤りが多くなっていた。In the conventional method, feature parameters affected by background noise are used as they are for determination, so when the background noise is large, consonant detection errors increase.
本発明の各実施例では、雑音と音声の特徴を強調するた
め、S/N比が20dBから14dBはどの、背景雑音
の大きな環境でも検出率が良好な検出率が得られた。In each of the embodiments of the present invention, in order to emphasize the characteristics of noise and voice, a good detection rate was obtained even in an environment with large background noise, where the S/N ratio was between 20 dB and 14 dB.
以下に、特徴パラメータ・特徴パラメータ変換法を変え
たときの語頭子音の検出結果を示す。Below, the detection results of word-initial consonants when changing the feature parameters and feature parameter conversion methods are shown.
音声データに付けられたラベルが子音を示しているフレ
ームが子音のクラスのうちいずれかであると判定された
場合、正しく検出されたものであるとする。If it is determined that a frame whose label attached to the audio data indicates a consonant is in one of the consonant classes, it is determined that the frame has been correctly detected.
第10図に示した検出率は子音検出率と雑音検出率の平
均値である。子音検出率は、次式で定義される。The detection rate shown in FIG. 10 is the average value of the consonant detection rate and the noise detection rate. The consonant detection rate is defined by the following equation.
子音検出率−
また、雑音データのフレームが、雑音クラスのうちいず
れかであると判定された場合、正しく検出されたものと
する。これが雑音検出率であり、次式で定義される。Consonant Detection Rate - Furthermore, if it is determined that the frame of noise data belongs to any of the noise classes, it is assumed that it has been correctly detected. This is the noise detection rate and is defined by the following equation.
雑音検出率=
第10図において、縦軸は検出率である。また、横軸は
特徴パラメータの種類を示しており、LPCはLPCケ
プストラム、Pはフレーム内平均電力、P+LPCはP
とLPCの併用である。Noise detection rate= In FIG. 10, the vertical axis is the detection rate. In addition, the horizontal axis indicates the type of feature parameter, where LPC is the LPC cepstrum, P is the average power within the frame, and P+LPC is the P
and LPC.
なお、以下ではLPCケプストラム分析次元は12次、
変換パラメータ次元は特徴パラメータがLPGのとき4
次、P+LPGのとき5次とした。In addition, in the following, the LPC cepstrum analysis dimension is 12th,
The conversion parameter dimension is 4 when the feature parameter is LPG.
Next, when P+LPG, it was set to 5th order.
特徴パラメータ変換法は、プロットを変えて示した。The feature parameter conversion method is shown by changing the plot.
Cは、特徴パラメータ変換を行わない従来の方法である
。C is a conventional method that does not perform feature parameter conversion.
nは、第2図に示した実施例であり、雑音判定をしてい
るものである。n is the embodiment shown in FIG. 2, in which noise is determined.
■は、第2図に示した実施例で、雑音判定をしていない
ものである。2 is the embodiment shown in FIG. 2, in which noise determination is not performed.
rは、第7図に示した他の実施例であり、特徴パラメー
タを回帰係数に変換したものである。r is another example shown in FIG. 7, in which feature parameters are converted into regression coefficients.
[発明の効果]
以上説明したように本発明によれば、特徴パラメータ変
換により特徴パラメータから雑音の影響を除去できるの
で、背景雑音が大きい環境下にあっても精確に有音区間
を判別することができる。[Effects of the Invention] As explained above, according to the present invention, the influence of noise can be removed from the feature parameters by feature parameter conversion, so that it is possible to accurately determine a sound interval even in an environment with large background noise. I can do it.
第1図は本発明に係る有音検出装置の概略構成を示すブ
ロック図、第2図は本発明の一実施例の特徴部分を示す
ブロック図、第3図は同実施例で使用されるバッファの
構成図、第4図は同実施例の変換パラメータの説明図、
第5図は有音判定器の構成例を示すブロック図、第6図
は標準パターンを作成する装置の構成例を示すブロック
図、第7図は本発明の他の実施例の特徴部分を示すブロ
ック図、第8図は同実施例の回帰係数を求める過程の近
似関数を示す図、第9図は各実施例における特徴パラメ
ータと検出率との関係を示す特性図、第10図は従来の
有音検出装置の構成例を示すブロック図である。
1・・・特徴パラメータ計算器
2・・・特徴パラメータ変換器
3・・・有音判定器
4・・・電力判定器
5・・・雑音判定器
6・・・スイッチ
8・・・雑音ベクトル計算器
9・・・マツチング器
10・・・標準パターンFIG. 1 is a block diagram showing a schematic configuration of a sound detection device according to the present invention, FIG. 2 is a block diagram showing characteristic parts of an embodiment of the present invention, and FIG. 3 is a buffer used in the embodiment. Fig. 4 is an explanatory diagram of the conversion parameters of the same embodiment.
FIG. 5 is a block diagram showing an example of the configuration of a voice determination device, FIG. 6 is a block diagram showing an example of the configuration of a device for creating a standard pattern, and FIG. 7 is a block diagram showing a characteristic part of another embodiment of the present invention. The block diagram, FIG. 8 is a diagram showing the approximation function in the process of calculating the regression coefficient in the same embodiment, FIG. 9 is a characteristic diagram showing the relationship between the feature parameters and detection rate in each embodiment, and FIG. 10 is the conventional FIG. 2 is a block diagram showing a configuration example of a sound detection device. 1... Feature parameter calculator 2... Feature parameter converter 3... Sound determiner 4... Power determiner 5... Noise determiner 6... Switch 8... Noise vector calculation Device 9...Matching device 10...Standard pattern
Claims (2)
特徴パラメータを演算する手段と、 演算されたフレーム単位の全ての特徴パラメータをある
いは雑音区間の特徴パラメータを順次蓄積するバッファ
と、 蓄積された特徴パラメータの内、入力フレームからSフ
レーム前の特徴パラメータを基準として過去のNフレー
ム分の特徴パラメータ集合を取り出し、前記入力フレー
ムの特徴パラメータとの距離ベクトル若しくはベクトル
のノルムを演算することによって入力フレームの特徴パ
ラメータの変換パラメータを生成する手段と、 前記変換パラメータと予め設定された標準パターンとを
比較して前記音声信号中の有音区間を判別する手段と、 を有することを特徴とする有音検出装置。(1) A means for calculating voice characteristic parameters in frame units from an input voice signal, a buffer for sequentially accumulating all the calculated characteristic parameters in frame units or characteristic parameters of noise intervals, Among the feature parameters, a set of feature parameters for past N frames is extracted based on the feature parameters S frames before the input frame, and the distance vector or norm of the vector from the feature parameters of the input frame is calculated. a means for generating a conversion parameter of the characteristic parameter; and a means for comparing the conversion parameter with a preset standard pattern to determine a sound section in the audio signal. Detection device.
特徴パラメータを演算する手段と、 演算されたフレーム単位の特徴パラメータを順次蓄積す
るバッファと、 蓄積された特徴パラメータから回帰係数を演算し、入力
フレームの特徴パラメータの変換パラメータを生成する
手段、 前記変換パラメータと予め設定された標準パターンとを
比較して前記音声信号中の有音区間を判別する手段と、 を有することを特徴とする有音検出装置。(2) means for calculating audio feature parameters in frame units from an input audio signal; a buffer for sequentially accumulating the calculated feature parameters in frame units; and calculating regression coefficients from the accumulated feature parameters; An apparatus characterized by comprising: means for generating a conversion parameter of a characteristic parameter of an input frame; and means for comparing the conversion parameter with a preset standard pattern to determine a sound interval in the audio signal. Sound detection device.
Priority Applications (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP1183684A JP3032215B2 (en) | 1989-07-18 | 1989-07-18 | Sound detection device and method |
Applications Claiming Priority (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP1183684A JP3032215B2 (en) | 1989-07-18 | 1989-07-18 | Sound detection device and method |
Publications (2)
| Publication Number | Publication Date |
|---|---|
| JPH0348900A true JPH0348900A (en) | 1991-03-01 |
| JP3032215B2 JP3032215B2 (en) | 2000-04-10 |
Family
ID=16140121
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| JP1183684A Expired - Fee Related JP3032215B2 (en) | 1989-07-18 | 1989-07-18 | Sound detection device and method |
Country Status (1)
| Country | Link |
|---|---|
| JP (1) | JP3032215B2 (en) |
Cited By (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN111755029A (en) * | 2020-05-27 | 2020-10-09 | 北京大米科技有限公司 | Voice processing method, device, storage medium and electronic device |
-
1989
- 1989-07-18 JP JP1183684A patent/JP3032215B2/en not_active Expired - Fee Related
Cited By (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN111755029A (en) * | 2020-05-27 | 2020-10-09 | 北京大米科技有限公司 | Voice processing method, device, storage medium and electronic device |
| CN111755029B (en) * | 2020-05-27 | 2023-08-25 | 北京大米科技有限公司 | Voice processing method, device, storage medium and electronic equipment |
Also Published As
| Publication number | Publication date |
|---|---|
| JP3032215B2 (en) | 2000-04-10 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US5692104A (en) | Method and apparatus for detecting end points of speech activity | |
| US5596680A (en) | Method and apparatus for detecting speech activity using cepstrum vectors | |
| JP3162994B2 (en) | Method for recognizing speech words and system for recognizing speech words | |
| JPS6336676B2 (en) | ||
| US4937870A (en) | Speech recognition arrangement | |
| Thomson et al. | Use of periodicity and jitter as speech recognition features | |
| EP1511007B1 (en) | Vocal tract resonance tracking using a target-guided constraint | |
| JPH10105187A (en) | Signal segmentalization method basing cluster constitution | |
| US5806031A (en) | Method and recognizer for recognizing tonal acoustic sound signals | |
| JP2002538514A (en) | Speech detection method using stochastic reliability in frequency spectrum | |
| Wang et al. | A multi-space distribution (MSD) approach to speech recognition of tonal languages. | |
| Slaney et al. | Pitch-gesture modeling using subband autocorrelation change detection. | |
| Kalaiarasi et al. | Performance Analysis and Comparison of Speaker Independent Isolated Speech Recognition System | |
| JPH0458297A (en) | Sound detecting device | |
| JPH02205897A (en) | Sound detector | |
| Sharma et al. | Speech recognition of Punjabi numerals using synergic HMM and DTW approach | |
| CN121600911B (en) | A dynamic error correction method and system for speech recognition results | |
| JP3032215B2 (en) | Sound detection device and method | |
| Sigmund | Search for keywords and vocal elements in audio recordings | |
| Seman et al. | Hybrid methods of Brandt’s generalised likelihood ratio and short-term energy for Malay word speech segmentation | |
| JPH0772899A (en) | Voice recognizer | |
| KR100269429B1 (en) | Transient voice determining method in voice recognition | |
| JPH0398098A (en) | Voice recognition device | |
| AlDahri et al. | Detection of Voice Onset Time (VOT) for unvoiced stop sound in Modern Standard Arabic (MSA) based on power signal | |
| Akdemir et al. | HMM topology for boundary refinement in automatic speech segmentation |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| LAPS | Cancellation because of no payment of annual fees |