JPH06195093A - Automatic tone quality evaluating device - Google Patents
Automatic tone quality evaluating deviceInfo
- Publication number
- JPH06195093A JPH06195093A JP4263349A JP26334992A JPH06195093A JP H06195093 A JPH06195093 A JP H06195093A JP 4263349 A JP4263349 A JP 4263349A JP 26334992 A JP26334992 A JP 26334992A JP H06195093 A JPH06195093 A JP H06195093A
- Authority
- JP
- Japan
- Prior art keywords
- input
- speech signal
- voice signal
- reproduced
- dynamic
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Granted
Links
- 238000011156 evaluation Methods 0.000 claims abstract description 86
- 238000000605 extraction Methods 0.000 claims abstract description 37
- 239000000284 extract Substances 0.000 claims abstract description 9
- 230000005236 sound signal Effects 0.000 claims description 36
- 238000000034 method Methods 0.000 claims description 23
- 238000013441 quality evaluation Methods 0.000 claims description 19
- 239000002131 composite material Substances 0.000 claims description 6
- 238000001228 spectrum Methods 0.000 description 13
- 238000010586 diagram Methods 0.000 description 6
- 230000005284 excitation Effects 0.000 description 2
- 238000012545 processing Methods 0.000 description 2
- 230000003595 spectral effect Effects 0.000 description 2
- 238000006243 chemical reaction Methods 0.000 description 1
- 238000004891 communication Methods 0.000 description 1
- 230000002596 correlated effect Effects 0.000 description 1
- 238000002474 experimental method Methods 0.000 description 1
- 230000037433 frameshift Effects 0.000 description 1
- 229940035637 spectrum-4 Drugs 0.000 description 1
Landscapes
- Transmission Systems Not Characterized By The Medium Used For Transmission (AREA)
Abstract
Description
【0001】[0001]
【産業上の利用分野】本発明は、音質評価装置に関し、
特に入力音声信号の特徴パラメータと再生音声信号の特
徴パラメータとの距離より再生音声信号の主観評価値を
予測する装置に関する。BACKGROUND OF THE INVENTION 1. Field of the Invention The present invention relates to a sound quality evaluation device,
In particular, the present invention relates to an apparatus for predicting a subjective evaluation value of a reproduced voice signal from the distance between the characteristic parameter of the input voice signal and the characteristic parameter of the reproduced voice signal.
【0002】[0002]
【従来の技術】従来、音声の音質評価は、聴取実験で評
価音声を主観的に評価する他に、音声の特徴パラメータ
原音声と評価音声から抽出し、特徴パラメータ間の比や
距離を求めて客観的に評価する方法がある。客観評価に
は、SN比(特開昭63−273895)(文献1)や
セグメンタルSN比などの他にケプストラム距離やBS
D距離(Proc.ICASSP 91,IEEE S
peech Processing.,vol.1,p
p.493−496,1991)(文献2)などスペク
トルの歪みが使われている。また最近では、これら客観
評価に用いられている音声の特徴パラメータから、主観
評価値を予測するモデルも研究されている(電子情報通
信学会論文誌A Vol.J73−A No.6 p
p.1039−1047 1990年6月)(文献
3)。2. Description of the Related Art Conventionally, in order to evaluate the sound quality of a voice, in addition to subjectively evaluating the evaluation voice in a listening experiment, it is extracted from the original feature voice of the voice and the evaluation voice to obtain the ratio and distance between the feature parameters. There is a method of objective evaluation. For objective evaluation, in addition to the SN ratio (Japanese Patent Laid-Open No. 63-273895) (Reference 1) and the segmental SN ratio, the cepstrum distance and BS
D distance (Proc.ICASSP 91, IEEE S
Peach Processing. , Vol. 1, p
p. 493-496, 1991) (reference 2) and the like, spectral distortion is used. Recently, a model for predicting a subjective evaluation value from the characteristic parameters of the voice used for the objective evaluation has also been studied (Journal of the Institute of Electronics, Information and Communication Engineers A Vol. J73-A No. 6 p.
p. 1039-1047 June 1990) (Reference 3).
【0003】[0003]
【発明が解決しようとする課題】しかし、これら従来の
音質評価方式は、8kb/s以下の低ビットレートの符
号化音声に対して、主観評価値との対応がつきにくいと
いう問題がある。However, these conventional sound quality evaluation methods have a problem that it is difficult to correspond to a subjective evaluation value for coded speech having a low bit rate of 8 kb / s or less.
【0004】[0004]
【課題を解決するための手段】上述した問題点を解決す
るため、第1の発明による自動音質評価装置は、音声信
号を入力し、前記入力音声信号を符号化/復号化し再生
音声信号を作成する符号化/復号化部と、前記入力音声
信号から特徴パラメータを抽出する入力音声信号特徴抽
出部と、前記再生音声信号から特徴パラメータを抽出る
再生音声信号特徴抽出部と、前記入力音声信号の特徴パ
ラメータを用いて動的特徴パラメータを抽出する入力音
声信号動的特徴抽出部と、前記再生音声信号の特徴パラ
メータを用いて動的特徴パラメータを抽出する再生音声
信号動的特徴抽出部と、前記入力音声信号の動的特徴パ
ラメータと前記再生音声信号の動的特徴パラメータとの
予め定められた距離を出力する第1の客観評価部と、前
記第1の客観評価部にて求められた客観評価値を用いて
前記再生音声信号の主観評価値を予測する主観評価予測
部を有する。In order to solve the above problems, an automatic sound quality evaluation apparatus according to the first invention inputs a voice signal and encodes / decodes the input voice signal to generate a reproduced voice signal. An encoding / decoding unit for extracting the characteristic parameter from the input audio signal, a reproduced audio signal characteristic extracting unit for extracting the characteristic parameter from the reproduced audio signal, and an input audio signal characteristic extracting unit for extracting the characteristic parameter from the reproduced audio signal. An input voice signal dynamic feature extraction unit that extracts a dynamic feature parameter using a feature parameter; a playback voice signal dynamic feature extraction unit that extracts a dynamic feature parameter using a feature parameter of the playback voice signal; A first objective evaluation unit for outputting a predetermined distance between a dynamic characteristic parameter of an input audio signal and a dynamic characteristic parameter of the reproduced audio signal; and the first objective evaluation. Using an objective evaluation value obtained by having a subjective evaluation prediction section for predicting a subjective evaluation value of the reproduced audio signal.
【0005】または第2の発明による自動音質評価装置
は第1の発明において、さらに前記入力音声信号の特徴
パラメータと、前記再生音声信号の特徴パラメータとの
距離を出力する第2の客観評価部と、前記第1の客観評
価部にて求められた客観評価値と、前記第2の客観評価
部にて求められた客観評価値とを用いて主観評価値を予
測する複合主観評価予測部を有する。Alternatively, the automatic sound quality evaluation apparatus according to the second invention is the automatic sound quality evaluation apparatus according to the first invention, further comprising a second objective evaluation section for outputting a distance between the characteristic parameter of the input audio signal and the characteristic parameter of the reproduced audio signal. A composite subjective evaluation prediction unit that predicts a subjective evaluation value by using the objective evaluation value obtained by the first objective evaluation unit and the objective evaluation value obtained by the second objective evaluation unit .
【0006】あるいは、第3の発明による自動音質評価
装置は第1の発明において、前記符号化/復号化部のか
わりに、予め定められた符号化方式により符号化/復号
化した再生音声を入力する再生音声入力端子を有する。Alternatively, the automatic sound quality evaluation apparatus according to the third invention is the automatic sound quality evaluation apparatus according to the first invention, wherein the reproduced sound coded / decoded by a predetermined coding method is input instead of the coding / decoding unit. It has a reproduction voice input terminal for.
【0007】[0007]
【実施例】図1、図2、図3を参照して本発明の実施例
について説明する。図1は第1の発明の音質評価装置の
実施例を示すブロック図である。DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS An embodiment of the present invention will be described with reference to FIGS. FIG. 1 is a block diagram showing an embodiment of the sound quality evaluation apparatus of the first invention.
【0008】図1において、入力端子1からは、音声信
号が入力され、符号化/復号化部3と入力音声信号特徴
抽出部4へ送られる。符号化/復号化部3では、原信号
を用いて符号化/復号化を行ない再生音声信号を作成す
る。符号化/復号化には例えばCELP(Proc.I
nt.Conf.Acoust.,Speech,Si
gnal Processing,pp.200−20
3,1989)(文献4)などが用いられる。前記符号
化/復号化部3にて作成された再生音声信号は、再生音
声信号特徴抽出部5へ送られる。In FIG. 1, a voice signal is input from an input terminal 1 and sent to an encoding / decoding unit 3 and an input voice signal feature extraction unit 4. The encoding / decoding unit 3 performs encoding / decoding using the original signal to create a reproduced voice signal. For encoding / decoding, for example, CELP (Proc.
nt. Conf. Accout. , Speech, Si
general Processing, pp. 200-20
3, 1989) (reference 4) and the like. The reproduced audio signal generated by the encoding / decoding unit 3 is sent to the reproduced audio signal feature extracting unit 5.
【0009】入力音声信号特徴抽出部4では、入力音声
信号Sx を用いて一定時間(フレーム)毎に、特徴パラ
メータxparを求める。入力音声信号特徴抽出部4に
て求められた特徴パラメータxparは、入力音声信号
動的特徴抽出部6へ送られる。入力音声信号特徴抽出部
6で求められる特徴パラメータxparとしては、例え
ばrms、Barkスペクトル(文献2)、ピッチ、ケ
プストラムなど周知のものを使うことができる。The input voice signal feature extraction unit 4 uses the input voice signal S x to obtain the feature parameter xpar at regular time intervals (frames). The feature parameter xpar obtained by the input voice signal feature extraction unit 4 is sent to the input voice signal dynamic feature extraction unit 6. As the feature parameter xpar obtained by the input voice signal feature extraction unit 6, well-known ones such as rms, Bark spectrum (Reference 2), pitch, and cepstrum can be used.
【0010】ここでは、入力音声信号Sx から特徴パラ
メータxparを求める例として、以下でrmsとBa
rkスペクトルを求める式を示す。Here, as an example of obtaining the characteristic parameter xpar from the input voice signal S x , rms and Ba will be described below.
The formula for obtaining the rk spectrum is shown.
【0011】まず、入力音声信号から、第kフレームの
特徴パラメータxrms( k ) を求める式を示す。First, an equation for obtaining the characteristic parameter xrms (k) of the k-th frame from the input voice signal will be shown.
【0012】[0012]
【数1】 [Equation 1]
【0013】次に、入力音声信号からBarkスペクト
ルを抽出する手順を説明する。Next, the procedure for extracting the Bark spectrum from the input voice signal will be described.
【0014】1.入力音声信号Sx ( k ) に対し、FF
Tを行ない、パワースペクトルX( k ) (f)を求め
る。1. FF for the input voice signal S x (k)
T is performed to obtain the power spectrum X ( k) (f).
【0015】2.パワースペクトルX( k ) (f)をB
arkスケールY( k ) (x)に変換する。xからyへ
の変換は、以下 の関係式を用いて行なう。2. Power spectrum X (k) (f) is B
Convert to ark scale Y (k) (x). The conversion from x to y is performed using the following relational expression.
【0016】f=600sinh(i/6) 3.Barkスケール変換したパワースペクトルY
( k ) (x)に臨界帯域フィルタF( k ) (x)をか
け、excitation patternD
( k ) (x)を求める。臨界フィルタF( k ) (x)
は、以下の式で表される。F = 600 sinh (i / 6) 3. Power spectrum Y converted to Bark scale
The critical band filter F (k) (x) is applied to (k) (x) to obtain the excitation patternD.
(k) (x) is calculated. Critical filter F (k) (x)
Is represented by the following formula.
【0017】 10log1 0 F( k ) (x)=7−7.5(x−α)
−17.5[0.196+(x−α)2 ]1 / 2 ここで、α=0.215とする。 D( k ) (x)=F( k ) (x)*Y( k ) (x) (*は、畳み込み演算子である。) D( k ) (x)−Excitation patter
n F( k ) (x)−臨界帯域フィルタ Y( k ) (x)−Barkスケール変換したパワースペ
クトル 4.Excitation patternD( k ) に
聴感重み付けを行なう。10 log 10 F (k) (x) = 7-7.5 (x-α)
−17.5 [0.196+ (x−α) 2 ] 1/2 Here, α = 0.215. D (k) (x) = F (k) (x) * Y (k) (x) (* is a convolution operator.) D (k) (x) -Excitation pattern
3. N F (k) (x) -critical band filter Y (k) (x) -Bark scale converted power spectrum 4. The perceptual weighting is applied to the Excitation pattern D (k) .
【0018】1800〜3400Hzの聴感重み付けH
(f)は以下の式より求めることができる。Hearing weight H from 1800 to 3400 Hz
(F) can be obtained from the following equation.
【0019】H(f)=(2.6+e- 2 j π f )/
(1.6+e- 2 j π f) 1800Hz以下ではH(f)=1とし3400Hz以
上では3400Hzと同じ値をとる。H (f) = (2.6 + e -2 j π f ) /
(1.6 + e −2 j π f ) H (f) = 1 at 1800 Hz or lower, and the same value as 3400 Hz at 3400 Hz or higher.
【0020】[0020]
【数2】 [Equation 2]
【0021】5.第kフレームのBarkスペクトル
{xB( k ) [i],i=1,...,I}は、以下の
式によって求められる。5. Bark spectrum {xB (k) [i], i = 1 ,. . . , I} is calculated by the following equation.
【0022】[0022]
【数3】 [Equation 3]
【0023】但し、channel番号iとBark
Scale xとの間には、x=1.0×iとなる関係
がある。However, channel number i and Bark
There is a relationship of x = 1.0 × i with Scale x.
【0024】以上の手順で求めた入力信号rmsやBa
rkスペクトルの特徴パラメータxrmsやxBは、入
力音声信号動的特徴抽出部6に出力される。The input signals rms and Ba obtained by the above procedure
The feature parameters xrms and xB of the rk spectrum are output to the input voice signal dynamic feature extraction unit 6.
【0025】再生音声信号特徴抽出部5では、再生音声
信号Sy を用いて、一定時間(フレーム)毎に特徴パラ
メータyparを求める。再生音声信号特徴抽出部5に
て求められた、特徴パラメータyparは、再生音声信
号動的特徴抽出部7へ送られる。再生音声信号特徴抽出
部5で求められる特徴パラメータyparとして例えば
特徴パラメータrms Barkスペクトル(文献
2)、ピッチ、ケプストラム などがある。再生音声信
号Sy から再生音声信号の特徴パラメータyparを求
める方法は、前記入力音声信号特徴抽出部4において、
入力音声信号Sx を用いて入力音声信号の特徴パラメー
タを求める方法と同じであるため、ここでは説明を省略
する。The reproduced voice signal feature extraction unit 5 uses the reproduced voice signal S y to obtain the characteristic parameter ypar for each fixed time (frame). The characteristic parameter ypar obtained by the reproduced voice signal characteristic extraction unit 5 is sent to the reproduced voice signal dynamic characteristic extraction unit 7. The characteristic parameters ypar obtained by the reproduced voice signal characteristic extraction unit 5 include, for example, characteristic parameter rms Bark spectrum (reference 2), pitch, cepstrum, and the like. The method for obtaining the characteristic parameter ypar of the reproduced audio signal from the reproduced audio signal S y is as follows.
Since the method is the same as the method of obtaining the characteristic parameter of the input audio signal using the input audio signal S x , the description thereof is omitted here.
【0026】入力音声信号動的特徴抽出部6について説
明する。次に、入力音声信号特徴抽出部4にて求められ
た、入力音声信号Sx の特徴パラメータxparを動的
特徴パラメータδxparに変換し、第1の客観評価部
8へと送る動作を行なう、入力音声信号動的特徴抽出部
6について説明する。入力音声信号動的特徴抽出部6
で、入力音声信号の特徴パラメータxparを動的特徴
パラメータδxparに変換する方法はいくつかあるた
め、ここで例をあげておく。The input voice signal dynamic feature extraction unit 6 will be described. Next, an operation of converting the characteristic parameter xpar of the input speech signal S x obtained by the input speech signal characteristic extraction unit 4 into a dynamic characteristic parameter δxpar and sending it to the first objective evaluation unit 8 is performed. The audio signal dynamic feature extraction unit 6 will be described. Input voice signal dynamic feature extraction unit 6
Since there are several methods for converting the characteristic parameter xpar of the input speech signal into the dynamic characteristic parameter δxpar, an example will be given here.
【0027】はじめに、入力音声信号の特徴パラメータ
xparを、動的特徴パラメータδxparに変換する
方法を、入力音声信号の特徴パラメータxpar1 と、
xpar1 から変換された動的特徴パラメータδxpa
r1 を使って説明する。[0027] First, the characteristic parameters XPAR of the input speech signal, the method for converting a dynamic characteristic parameter Derutaxpar, wherein parameters XPAR 1 of the input audio signal,
Dynamic feature parameter δxpa converted from xpar 1
An explanation will be given using r 1 .
【0028】第kフレームで抽出された特徴パラメータ
xpar1 ( k ) の動的特徴パラメータδxpar1
( k ) は、lフレーム前の特徴パラメータxpar1
( k + l) からlフレーム後の特徴パラメータxpar
1 ( k - l ) の差によって求める。Dynamic feature parameter δxpar 1 of the feature parameter xpar 1 (k) extracted in the k-th frame
(k) is the feature parameter xpar 1 before the 1-frame
Feature parameter xpar after 1 frame from (k + l)
Calculated by the difference of 1 (k-l) .
【0029】δxpar1 ( k ) =xpar1
( k + l ) −xpar1 ( k - l ) k −フレーム番号 l −フレームシフト数 xpar1 ( k ) −入力音声信号の第kフレームの特
徴パラメータ δxpar1 ( k ) −入力音声信号の第kフレームの動
的特徴パラメータ 入力音声信号の特徴パラメータxpar1 がrmsであ
った場合の動的特徴パラメータの求め方を以下に示す。
ここで、入力音声信号のrmsの特徴パラメータをxr
m、動的特徴パラメータをδxrmsとすると、第kフ
レームのrmsの動的特徴パラメータδxrms( k )
は、以下の式で求めることができる。Δxpar 1 (k) = xpar 1
(K + l) -xpar 1 ( k - l) k - frame number l - frame shift number xpar 1 (k) - characteristic parameters Derutaxpar 1 of the k frame of the input audio signal (k) - k-th input speech signal Dynamic Feature Parameter of Frame A method of obtaining the dynamic feature parameter when the feature parameter xpar 1 of the input speech signal is rms is shown below.
Here, the characteristic parameter of rms of the input voice signal is xr
m and the dynamic feature parameter is δxrms, the dynamic feature parameter δxrms (k) of the rms of the k-th frame
Can be calculated by the following formula.
【0030】δxrms( k ) =xrms( k + l ) −
xrms( k - l ) k −フレーム番号 l −フレームシフト数 xrms( k ) −入力音声信号の第kフレームのrm
s δxrms( k ) −入力音声信号の第kフレームのrm
sの動的特徴パラメータ 次に、入力音声信号の特徴パラメータxparを、動的
特徴パラメータδxparに変換する方法を、入力音声
信号の特徴パラメータxpar2 が多次元である場合
に、xpar2 から変換される動的特徴パラメータδx
par2 として説明する。Δxrms (k) = xrms (k + l) -
xrms (k-l) k-frame number l-frame shift number xrms (k) -rm of kth frame of input speech signal
s δxrms (k) -rm of the kth frame of the input speech signal
dynamic characteristic parameter of s then the characteristic parameters XPAR of the input speech signal, the method for converting a dynamic characteristic parameter Derutaxpar, when the characteristic parameter XPAR 2 of the input audio signal is a multi-dimensional, is converted from XPAR 2 Dynamic feature parameter δx
This will be described as par 2 .
【0031】第kフレームで抽出されたT次元目の特徴
パラメータ{xpar2 ( k ) [t],t=
1,...,T}の動的特徴パラメータδpar2
( k ) [t]は、以下の式により求めることができる。Characteristic parameter {xpar 2 (k) [t], t = of the T-th dimension extracted in the k-th frame
1 ,. . . , T} dynamic feature parameter δ par 2
(k) [t] can be calculated by the following equation.
【0032】δxpar2 ( k ) [t]=xpar2
( k + l ) [t]−xpar2 ( k - l ) [t] t=1,...,T k −フレーム番号 t −次元番号 T −次元数 l −フレームシフト数 xpar2 ( k ) [t] −入力音声信号の第kフレー
ム、第t次元目の特徴ラメータ δxpar2 ( k ) [t]−入力音声信号の第kフレー
ム、第t次元目の動的特徴パラメータ 入力音声信号の特徴パラメータxpar2 がBarkス
ペクトルであった場合の動的特徴パラメータの求め方を
以下に示す。ここで、入力音声信号のBarkスペクト
ルの特徴パラメータをxB、動的特徴パラメータをδx
Bとすると、第kフレームで抽出されたI次元目Bar
kスペクトルの特徴パラメータ{xB(K ) [i],i
=1,...,I}の動的特徴パラメータδxB( K )
[i]は、以下の式により求めることができる。Δxpar 2 (k) [t] = xpar 2
(K + l) [t] -xpar 2 (k - l) [t] t = 1 ,. . . , T k-frame number t-dimension number T-dimension number l-frame shift number xpar 2 (k) [t] -kth frame of the input speech signal, characteristic parameter of the tth dimension δxpar 2 (k) [t ] -Dynamic feature parameter of k-th frame and t-th dimension of input voice signal The method of obtaining the dynamic feature parameter when the feature parameter xpar 2 of the input voice signal is the Bark spectrum is shown below. Here, the characteristic parameter of the Bark spectrum of the input speech signal is xB, and the dynamic characteristic parameter is δx.
Let B be the I-th dimension Bar extracted in the kth frame
Characteristic parameter of k spectrum {xB (K) [i], i
= 1 ,. . . , I} dynamic feature parameter δxB (K)
[I] can be obtained by the following formula.
【0033】δxB( k ) =xB( k + l ) [i]−x
B( k - 1 ) [i] k −フレーム番号 i −次元番号 l −フレームシフト数 xB( k ) −入力音声信号の第kフレームのBark
スペクトル δxB( k ) −入力音声信号の第kフレームのBark
スペクトルの動的特徴パラメータ また、入力音声信号の特徴パラメータxparを、動的
特徴パラメータδxparに変換する上記以外の方法
を、入力音声信号の特徴パラメータxpar3 と、xp
ar3 から変換された動的特徴パラメータδxpar3
を使って説明する。ΔxB (k) = xB (k + l) [i] -x
B (k-1) [i] k-frame number i-dimension number l-frame shift number xB (k) -bark of k-th frame of input speech signal
Spectrum δxB (k) -Bark of k-th frame of input speech signal
Dynamic Feature Parameter of Spectral A method other than the above for converting the feature parameter xpar of the input speech signal into the dynamic feature parameter δxpar is used as the feature parameters xpar 3 and xp of the input speech signal.
ar 3 dynamic characteristic parameters transformed from Derutaxpar 3
Use to explain.
【0034】第kフレームの動的特徴パラメータδxp
ar3 ( k ) は、特徴パラメータxpar3 ( k ) と特
徴パラメータxpar3 の平均特徴パラメータavg−
xparの差により求まる。Dynamic feature parameter δxp of the k-th frame
ar 3 (k) is an average feature parameter avg− of the feature parameter xpar 3 (k) and the feature parameter xpar 3.
It is obtained by the difference of xpar.
【0035】[0035]
【数4】 [Equation 4]
【0036】さらに、入力音声信号の特徴パラメータx
parを、動的特徴パラメータδxparに変換する方
法として、入力音声信号の特徴パラメータxpar
4 と、xpar4 から変換された動的特徴パラメータδ
xpar4 を使って説明する。Further, the characteristic parameter x of the input voice signal
As a method of converting the par into the dynamic feature parameter δxpar, the feature parameter xpar of the input audio signal is used.
4 and the dynamic feature parameter δ converted from xpar 4
Explain using xpar 4 .
【0037】[0037]
【数5】 [Equation 5]
【0038】また、再生音声信号特徴抽出部5にて求め
られた再生音声信号Sy の特徴パラメータyparは、
再生音声信号動的特徴抽出部7にて、動的特徴パラメー
タδyparに変換され、第1の客観評価部8へ送られ
る。再生音声信号動的特徴抽出部7での、動的特徴パラ
メータの求め方は、入力音声信号特徴抽出部6と同じで
あるため、ここでは説明を省略する。The characteristic parameter ypar of the reproduced voice signal S y obtained by the reproduced voice signal characteristic extraction unit 5 is
The reproduced voice signal dynamic feature extraction unit 7 converts the dynamic feature parameter δypar and sends the dynamic feature parameter δypar to the first objective evaluation unit 8. The method for obtaining the dynamic feature parameter in the reproduced voice signal dynamic feature extraction unit 7 is the same as that in the input voice signal feature extraction unit 6, and therefore the description is omitted here.
【0039】入力音声信号動的特徴抽出部6と、再生音
声信号動的特徴抽出部7にて求められた動的特徴パラメ
ータは、第1の客観評価部8で、入力信号の動的特徴パ
ラメータxparと再生信号の動的特徴パラメータyp
arとの距離を計算し動的特徴パラメータの客観評価値
δAVGとして、主観評価予測部9へ送られる。動的特
徴パラメータから客観評価値を求める方法を以下に示
す。The dynamic feature parameters obtained by the input voice signal dynamic feature extraction unit 6 and the reproduced voice signal dynamic feature extraction unit 7 are supplied to the first objective evaluation unit 8 as dynamic feature parameters of the input signal. xpar and the dynamic characteristic parameter yp of the reproduced signal
The distance from ar is calculated and sent as an objective evaluation value δAVG of the dynamic feature parameter to the subjective evaluation prediction unit 9. The method of obtaining the objective evaluation value from the dynamic feature parameters is shown below.
【0040】まず、動的特徴パラメータが、1次元であ
った場合の客観評価値の求め方を説明する。ここで、入
力音声信号の動的特徴パラメータをδxpar1 、再生
音声信号の動的特徴パラメータをδypar1 、客観評
価値をδAVG1 とすると、第kフレームの入力音声信
号の動的特徴パラメータδxpar1 ( k ) と、第kフ
レームの再生音声信号の動的特徴パラメータδypar
1 ( k ) との、客観評価値δAVG1 は以下の式で求め
られる。First, a method of obtaining an objective evaluation value when the dynamic feature parameter is one-dimensional will be described. Here, assuming that the dynamic characteristic parameter of the input speech signal is δxpar 1 , the dynamic characteristic parameter of the reproduced speech signal is δypar 1 , and the objective evaluation value is δAVG 1 , the dynamic characteristic parameter δxpar 1 of the input speech signal of the k-th frame. (k) and the dynamic feature parameter δypar of the reproduced audio signal of the k-th frame
The objective evaluation value δAVG 1 with 1 (k) is calculated by the following equation.
【0041】[0041]
【数6】 [Equation 6]
【0042】次に、入力音声の動的特徴パラメータδx
par1 をδxrms、再生音声の動的特徴パラメータ
δypar1 をδyrms、動的特徴パラメータの客観
評価値δAVG1 をδAVGr m s とすると、第kフレ
ームの入力音声信号のrmsの動的特徴パラメータδx
rms1 ( k ) と、第kフレームの再生音声信号のrm
sの動的特徴パラメータδyrms1 ( k ) との、客観
評価値δAVGr m sは以下の式で求められる。Next, the dynamic feature parameter δx of the input voice
The par 1 δxrms, δyrms dynamic characteristic parameter Derutaypar 1 playback sound, when an objective evaluation value DerutaAVG 1 of the dynamic characteristic parameter and DerutaAVG rms, dynamic characteristic parameter of the rms of the input speech signal of the k frame δx
rms 1 (k) and rm of the reproduced audio signal of the k-th frame
The objective evaluation value δAVG rms with the dynamic feature parameter δyrms 1 (k) of s is obtained by the following formula.
【0043】[0043]
【数7】 [Equation 7]
【0044】さらに、動的特徴パラメータが多次元であ
る場合の客観評価値の求め方を、入力音声信号の動的特
徴パラメータをδxpar2 、再生音声信号の動的特徴
パラメータをδypar2 、客観評価値をδAVG2 と
して説明する。Further, the method of obtaining the objective evaluation value when the dynamic feature parameter is multidimensional is as follows: the dynamic feature parameter of the input voice signal is δxpar 2 , the dynamic feature parameter of the reproduced voice signal is δypar 2 , and the objective assessment is performed. The value will be described as δAVG 2 .
【0045】[0045]
【数8】 [Equation 8]
【0046】ここで、入力音声の動的特徴パラメータδ
xpar2 をδxB、再生音声の動的特徴パラメータδ
ypar2 をδyB、動的特徴パラメータの客観評価値
δAVG2 をδAVGB S D として、BSDの求め方を
説明する。第kフレームの入力音声信号のBSDの動的
特徴パラメータδxB( k ) と、第kフレームの再生音
声信号のBSDの動的特徴パラメータδyB( k ) と
の、客観評価値δAVGB S D は以下の式で求められ
る。Here, the dynamic feature parameter δ of the input voice
xpar 2 is δxB, dynamic characteristic parameter δ of reproduced voice
A method of obtaining BSD will be described, where ypar 2 is δyB and the objective evaluation value δAVG 2 of the dynamic feature parameter is δAVG BSD . Dynamic characteristic parameter δxB the BSD of the input audio signal of the k-th frame (k), the dynamic characteristic parameter δyB the BSD of reproduced audio signal of the k-th frame (k), an objective evaluation value DerutaAVG BSD following formula Required by.
【0047】客観評価値δAVGB S D は以下の式で求
めることができる。The objective evaluation value δAVG BSD can be calculated by the following formula.
【0048】[0048]
【数9】 [Equation 9]
【0049】以上の方法によって第1の客観評価部8で
求めた、動的特徴パラメータの客観評価値は、主観評価
予測部8へ送られる。The objective evaluation value of the dynamic feature parameter obtained by the first objective evaluation unit 8 by the above method is sent to the subjective evaluation prediction unit 8.
【0050】主観評価予測部8では、少なくとも1つの
客観評価値と少なくとも2つの予測係数で主観評価値を
予測し、評価結果を出力端子2より出力する。予測係数
は、予め大量の音声データを用いて集めた主観評価値と
予測評価値の誤差が、最小になるように求められる。予
測係数aと主観評価値との関係を以下に示す。The subjective evaluation prediction unit 8 predicts the subjective evaluation value with at least one objective evaluation value and at least two prediction coefficients, and outputs the evaluation result from the output terminal 2. The prediction coefficient is obtained so that the error between the subjective evaluation value and the prediction evaluation value collected in advance using a large amount of voice data is minimized. The relationship between the prediction coefficient a and the subjective evaluation value is shown below.
【0051】[0051]
【数10】 [Equation 10]
【0052】予測係数と、第1の客観評価部8で求めた
客観評価値とを用いて予測主観評価値を求める式を以下
に示す。予測主観評価値はThe formula for obtaining the predicted subjective evaluation value using the prediction coefficient and the objective evaluation value obtained by the first objective evaluation unit 8 is shown below. The predicted subjective evaluation value is
【0053】[0053]
【数11】 [Equation 11]
【0054】、予測係数はa、b、c、第1の客観評価
部8にて求めた動的特徴パラメータの客観評価値はδA
VGaとδAVGbとする。The prediction coefficients are a, b, and c, and the objective evaluation value of the dynamic feature parameter obtained by the first objective evaluation unit 8 is δA.
Let VGa and δAVGb.
【0055】[0055]
【数12】 [Equation 12]
【0056】図2は第2の発明の音質評価装置の実施例
を示すブロック図である。図2において、同一の番号の
ある構成要素は、図1と同一の動作をするので説明は省
略する。FIG. 2 is a block diagram showing an embodiment of the sound quality evaluation apparatus of the second invention. In FIG. 2, components having the same numbers operate in the same manner as in FIG.
【0057】第2の客観評価部10では、入力音声信号
Sx から入力音声信号特徴抽出部4で求めた、特徴パラ
メータxparと、再生音声信号Sy から再生音声信号
特徴抽出部5で求めた、特徴パラメータyparとの、
距離を客観評価値AVGとして計算し、複合主観評価予
測部12へ出力する。In the second objective evaluation unit 10, the characteristic parameter xpar obtained by the input voice signal feature extraction unit 4 from the input voice signal S x and the reproduced voice signal feature extraction unit 5 from the reproduced voice signal S y are obtained. , With the feature parameter ypar,
The distance is calculated as the objective evaluation value AVG and output to the composite subjective evaluation prediction unit 12.
【0058】ここで、入力音声信号の特徴パラメータを
xpar、再生音声信号の特徴パラメータをypar、
客観評価値をAVGとすると、第kフレームの入力音声
信号の特徴パラメータxpar( k ) と、第kフレーム
の再生音声信号の特徴パラメータypar( k ) との、
客観評価値AVGは以下の式で求められる。Here, the characteristic parameter of the input audio signal is xpar, the characteristic parameter of the reproduced audio signal is ypar,
When AVG objective evaluation value, the characteristic parameters xpar of the input speech signal of the k-th frame (k), the characteristic parameters ypar the reproduced audio signal of the k-th frame (k),
The objective evaluation value AVG is calculated by the following formula.
【0059】[0059]
【数13】 [Equation 13]
【0060】複合主観評価予測部12では、第1の客観
評価部8にて求められた少なくとも1つの客観評価値
と、第2の客観評価部10にて求められた少なくとも1
つの客観評価値と、少なくとも3つの予測係数で主観評
価値を予測し、評価結果を出力端子2より出力する。複
合主観評価予測部12の動作は、主観評価予測部9と同
じであるため、説明は省略する。In the composite subjective evaluation prediction unit 12, at least one objective evaluation value obtained by the first objective evaluation unit 8 and at least one objective evaluation value obtained by the second objective evaluation unit 10.
The subjective evaluation value is predicted with one objective evaluation value and at least three prediction coefficients, and the evaluation result is output from the output terminal 2. The operation of the composite subjective evaluation prediction unit 12 is the same as that of the subjective evaluation prediction unit 9, and thus the description thereof will be omitted.
【0061】図3は第3の発明の音質評価装置の実施例
を示すブロック図である。図3において、同一の番号の
ある構成要素は、図1及び図2と同一の動作をするので
説明は省略する。FIG. 3 is a block diagram showing an embodiment of the sound quality evaluation apparatus of the third invention. In FIG. 3, components having the same numbers operate in the same way as in FIGS.
【0062】再生音声入力端子6からは、入力端子1か
ら入力される音声信号と同一の音声信号で符号化/復号
化を行ない求めた再生音声を入力し、再生音声信号特徴
抽出部5に出力する。From the reproduced voice input terminal 6, the reproduced voice obtained by encoding / decoding with the same voice signal as the voice signal input from the input terminal 1 is input and output to the reproduced voice signal feature extraction unit 5. To do.
【0063】[0063]
【発明の効果】以上説明したように、本発明による音質
評価方式は、従来用いられている特徴パラメータから動
的特徴パラメータに変換することで、低ビットレートの
音声でも、主観評価値との相関の高い音質評価を行なう
ことができる。As described above, the sound quality evaluation method according to the present invention converts a conventionally used feature parameter into a dynamic feature parameter, so that even a low bit rate voice is correlated with a subjective evaluation value. It is possible to perform high-quality sound evaluation.
【図1】第1の発明の音質評価装置の実施例を示すブロ
ック図。FIG. 1 is a block diagram showing an embodiment of a sound quality evaluation apparatus of the first invention.
【図2】第2の発明の音質評価装置の実施例を示すブロ
ック図。FIG. 2 is a block diagram showing an embodiment of a sound quality evaluation apparatus of the second invention.
【図3】第3の発明の音質評価装置の実施例を示すブロ
ック図。FIG. 3 is a block diagram showing an embodiment of a sound quality evaluation apparatus of the third invention.
1 入力端子 2 出力端子 3 符号化/復号化部 4 入力音声信号特徴抽出部 5 再生音声信号特徴抽出部 6 入力音声信号動的特徴抽出部 7 再生音声信号動的特徴抽出部 8 第1の客観評価部 9 主観評価予測部 10 第2の客観評価部 12 複合主観評価予測部 13 再生音声入力端子 1 Input Terminal 2 Output Terminal 3 Encoding / Decoding Section 4 Input Voice Signal Feature Extraction Section 5 Playback Voice Signal Feature Extraction Section 6 Input Voice Signal Dynamic Feature Extraction Section 7 Playback Voice Signal Dynamic Feature Extraction Section 8 First Objective Evaluation unit 9 Subjective evaluation prediction unit 10 Second objective evaluation unit 12 Composite subjective evaluation prediction unit 13 Playback audio input terminal
Claims (3)
符号化/復号化し再生音声信号を作成する符号化/復号
化部と、前記入力音声信号から特徴パラメータを抽出す
る入力音声信号特徴抽出部と、前記再生音声信号から特
徴パラメータを抽出する再生音声信号特徴抽出部と、前
記入力音声信号の特徴パラメータを用いて動的特徴パラ
メータを抽出する入力音声信号動的特徴抽出部と、前記
再生音声信号の特徴パラメータを用いて動的特徴パラメ
ータを抽出する再生音声信号動的特徴抽出部と、前記入
力音声信号の動的特徴パラメータと前記再生音声信号の
動的特徴パラメータとの予め定められた距離を出力する
第1の客観評価部と、前記第1の客観評価部にて求めら
れた客観評価値を用いて前記再生音声信号の主観評価値
を予測する主観評価予測部を有することを特徴とする音
質評価装置。1. An encoding / decoding unit for inputting an audio signal and encoding / decoding the input audio signal to create a reproduced audio signal; and an input audio signal feature extraction for extracting a characteristic parameter from the input audio signal. A playback voice signal feature extraction unit that extracts a feature parameter from the playback voice signal, an input voice signal dynamic feature extraction unit that extracts a dynamic feature parameter using the feature parameter of the input voice signal, and the playback A playback voice signal dynamic feature extraction unit that extracts a dynamic feature parameter using a feature parameter of a voice signal, and a predetermined dynamic feature parameter of the input voice signal and a dynamic feature parameter of the playback voice signal. A first objective evaluation unit that outputs a distance, and a subjective evaluation that predicts the subjective evaluation value of the reproduced audio signal using the objective evaluation value obtained by the first objective evaluation unit. A sound quality evaluation apparatus having a prediction unit.
さらに、前記入力音声信号の特徴パラメータと、前記再
生音声信号の特徴パラメータとの距離を出力する第2の
客観評価部と、前記第1の客観評価部にて求められた客
観評価値と、前記第2の客観評価部にて求められた客観
評価値とを用いて主観評価値を予測する複合主観評価予
測部を有することを特徴とする音質評価装置。2. The sound quality evaluation apparatus according to claim 1, wherein
Furthermore, a second objective evaluation unit that outputs a distance between the characteristic parameter of the input audio signal and the characteristic parameter of the reproduced audio signal, the objective evaluation value obtained by the first objective evaluation unit, and A sound quality evaluation apparatus comprising a composite subjective evaluation prediction unit that predicts a subjective evaluation value using the objective evaluation value obtained by the second objective evaluation unit.
前記符号化/復号化部のかわりに、予め定められた符号
化方式により符号化/復号化した再生音声を入力する再
生音声入力端子を備えることを特徴とする音質評価装
置。3. The sound quality evaluation apparatus according to claim 1,
A sound quality evaluation apparatus comprising, instead of the encoding / decoding unit, a reproduced sound input terminal for inputting reproduced sound encoded / decoded by a predetermined encoding method.
Priority Applications (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP4263349A JP2972459B2 (en) | 1992-10-01 | 1992-10-01 | Automatic sound quality evaluation device |
Applications Claiming Priority (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP4263349A JP2972459B2 (en) | 1992-10-01 | 1992-10-01 | Automatic sound quality evaluation device |
Publications (2)
| Publication Number | Publication Date |
|---|---|
| JPH06195093A true JPH06195093A (en) | 1994-07-15 |
| JP2972459B2 JP2972459B2 (en) | 1999-11-08 |
Family
ID=17388242
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| JP4263349A Expired - Lifetime JP2972459B2 (en) | 1992-10-01 | 1992-10-01 | Automatic sound quality evaluation device |
Country Status (1)
| Country | Link |
|---|---|
| JP (1) | JP2972459B2 (en) |
Citations (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JPS622286A (en) * | 1985-06-27 | 1987-01-08 | 松下電器産業株式会社 | Pronounciation practicing apparatus |
| JPS63273895A (en) * | 1987-05-06 | 1988-11-10 | 松下電器産業株式会社 | Sound quality evaluation device for high-efficiency speech code/decoder |
| JPH02153397A (en) * | 1988-12-06 | 1990-06-13 | Nec Corp | Voice recording device |
-
1992
- 1992-10-01 JP JP4263349A patent/JP2972459B2/en not_active Expired - Lifetime
Patent Citations (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JPS622286A (en) * | 1985-06-27 | 1987-01-08 | 松下電器産業株式会社 | Pronounciation practicing apparatus |
| JPS63273895A (en) * | 1987-05-06 | 1988-11-10 | 松下電器産業株式会社 | Sound quality evaluation device for high-efficiency speech code/decoder |
| JPH02153397A (en) * | 1988-12-06 | 1990-06-13 | Nec Corp | Voice recording device |
Also Published As
| Publication number | Publication date |
|---|---|
| JP2972459B2 (en) | 1999-11-08 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| KR100427753B1 (en) | Method and apparatus for reproducing voice signal, method and apparatus for voice decoding, method and apparatus for voice synthesis and portable wireless terminal apparatus | |
| US8315861B2 (en) | Wideband speech decoding apparatus for producing excitation signal, synthesis filter, lower-band speech signal, and higher-band speech signal, and for decoding coded narrowband speech | |
| JP5343098B2 (en) | LPC harmonic vocoder with super frame structure | |
| JP5226777B2 (en) | Recovery of hidden data embedded in audio signals | |
| JP4489960B2 (en) | Low bit rate coding of unvoiced segments of speech. | |
| JP2004508596A (en) | Output-based objective speech quality evaluation method and apparatus | |
| JP3765171B2 (en) | Speech encoding / decoding system | |
| JP3628268B2 (en) | Acoustic signal encoding method, decoding method and apparatus, program, and recording medium | |
| JPH06236198A (en) | Tone quality subjective evaluation prediction system | |
| JPH07261800A (en) | Transform coding method, decoding method | |
| JP2972459B2 (en) | Automatic sound quality evaluation device | |
| JP4373693B2 (en) | Hierarchical encoding method and hierarchical decoding method for acoustic signals | |
| JP3496618B2 (en) | Apparatus and method for speech encoding / decoding including speechless encoding operating at multiple rates | |
| JPH10111700A (en) | Audio compression encoding method and audio compression encoding device | |
| JP2900431B2 (en) | Audio signal coding device | |
| JP3350340B2 (en) | Voice coding method and voice decoding method | |
| JPH0786952A (en) | Predictive coding method for speech | |
| JPH0481199B2 (en) | ||
| JP3715417B2 (en) | Audio compression encoding apparatus, audio compression encoding method, and computer-readable recording medium storing a program for causing a computer to execute each step of the method | |
| JPH06208398A (en) | Sound source waveform generation method | |
| JPH043878B2 (en) | ||
| JPH11133999A (en) | Audio encoding / decoding device | |
| JPH0426119B2 (en) | ||
| JPH08328598A (en) | Sound coding/decoding device | |
| JPS6151200A (en) | Voice signal coding system |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| A02 | Decision of refusal |
Free format text: JAPANESE INTERMEDIATE CODE: A02 Effective date: 19970513 |