JPH0122638B2 - - Google Patents

Info

Publication number
JPH0122638B2
JPH0122638B2 JP55029987A JP2998780A JPH0122638B2 JP H0122638 B2 JPH0122638 B2 JP H0122638B2 JP 55029987 A JP55029987 A JP 55029987A JP 2998780 A JP2998780 A JP 2998780A JP H0122638 B2 JPH0122638 B2 JP H0122638B2
Authority
JP
Japan
Prior art keywords
analysis
value
circuit
window
formant
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Expired
Application number
JP55029987A
Other languages
Japanese (ja)
Other versions
JPS56126895A (en
Inventor
Katsunobu Fushikida
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
NEC Corp
Original Assignee
Nippon Electric Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Nippon Electric Co Ltd filed Critical Nippon Electric Co Ltd
Priority to JP2998780A priority Critical patent/JPS56126895A/en
Publication of JPS56126895A publication Critical patent/JPS56126895A/en
Publication of JPH0122638B2 publication Critical patent/JPH0122638B2/ja
Granted legal-status Critical Current

Links

Landscapes

  • Transmission Systems Not Characterized By The Medium Used For Transmission (AREA)

Description

【発明の詳細な説明】 音声分析合成あるいは音声認識等に用いる音声
分析装置として音声波形を10〜30msec程度の分
析周期(フレーム周期)毎に矩形窓、ハミング窓
等の分析窓を用いて自己相関値、フオルマント周
波数等の音声の特徴パラメータを抽出する音声分
析装置が知られている。しかしながら従来の音声
分析装置は音声波形が準周期的な波形を持つ等の
理由により前記分析窓の位置により抽出される前
記特徴パラメータ値が影響を受け、分析エラーに
よるゆらぎが生ずる欠点がある。特に1ピツチ周
期以内の短時間区間内で特徴パラメータ値を算出
する場合には声帯音源の励振時点により大きな影
響を受け分析エラーが生ずるおそれがある。この
欠点を防ぐため声帯音源の励振時点を抽出し、前
記励振時点に同期して分析を行なう方式も知られ
ているが、励振時点の正確な検出は非常に困難で
ある。
[Detailed Description of the Invention] As a speech analysis device used for speech analysis and synthesis or speech recognition, speech waveforms are autocorrelated using an analysis window such as a rectangular window or a Hamming window at every analysis period (frame period) of about 10 to 30 msec. 2. Description of the Related Art Speech analysis devices that extract voice characteristic parameters such as values and formant frequencies are known. However, the conventional speech analysis apparatus has a drawback that the feature parameter value extracted is affected by the position of the analysis window due to the fact that the speech waveform has a quasi-periodic waveform, and fluctuations occur due to analysis errors. In particular, when calculating characteristic parameter values within a short period of one pitch cycle or less, there is a risk that an analysis error may occur due to the large influence of the excitation point of the vocal cord sound source. In order to prevent this drawback, a method is known in which the excitation time point of the vocal cord sound source is extracted and analysis is performed in synchronization with the excitation time point, but accurate detection of the excitation time point is extremely difficult.

本発明の目的は、分析窓の時間的位置と励振時
点とのずれ等により生ずる分析誤差を減少させ比
較的ゆらぎが少なく正確な音声の特徴パラメータ
値系列を抽出することのできる音声分析装置を提
供することにある。
An object of the present invention is to provide a speech analysis device that is capable of reducing analysis errors caused by differences between the temporal position of an analysis window and the excitation point, etc., and extracting an accurate speech characteristic parameter value series with relatively little fluctuation. It's about doing.

本発明は、分析フレーム周期毎に複数の時間位
置の異なる分析窓を用いて特徴パラメータ値を算
出する手段と、時間的に隣接する前記分析区間に
おいて抽出された特徴パラメータ値間の距離を算
出する手段と、複数の前記分析区間を含む時間区
間内において、前記時間位置の異なる分析窓に対
する特徴パラメータのなかから、前記特徴パラメ
ータ値間の距離の総和を最小とする特徴パラメー
タ値の系列を選択する手段とから構成されてい
る。
The present invention provides a means for calculating feature parameter values using a plurality of analysis windows at different temporal positions for each analysis frame period, and a method for calculating a distance between feature parameter values extracted in temporally adjacent analysis sections. and selecting a series of feature parameter values that minimizes the sum of distances between the feature parameter values from among feature parameters for analysis windows at different time positions within a time interval including a plurality of said analysis intervals. It consists of means.

本発明の特徴は、まず各分析区間内において、
時間的位置の異なる複数の分析窓に対して、それ
ぞれ、フオルマント周波数等の特徴パラメータ値
を算出しておき、それらの特徴パラメータ値のな
かから最適なものを時間的に相隣る分析区間にお
いて選択される特徴パラメータ値との距離の総和
が最小となるように選択することにある。
The feature of the present invention is that, within each analysis interval, first,
Characteristic parameter values such as formant frequencies are calculated for multiple analysis windows at different temporal positions, and the optimal one is selected from among these characteristic parameter values in temporally adjacent analysis intervals. The objective is to select such that the sum of the distances from the characteristic parameter values to be selected is the minimum.

従つて、本発明によれば分析フレーム間のゆら
ぎが比較的少ない特徴パラメータ値系列を選択で
きることは明らかである。
Therefore, it is clear that according to the present invention, a feature parameter value series with relatively little fluctuation between analysis frames can be selected.

最適な特徴パラメータ値系列の選択は以下に述
べるダイナミツク・プログラミング法(DP法)
を用いて能率よく選択することができる。
The optimal feature parameter value series is selected using the dynamic programming method (DP method) described below.
can be selected efficiently using

n番目の分析フレームにおけるi番目の分析窓
に対する特徴パラメータ値をベクトル〓oiで表
わし、特徴ベクトル〓oiと〓o-1jとの距離を
R(n、i|n−1、j)で表わすと最適な各分
析区間における特徴パラメータ値系列は次の漸化
式(1)、(2)を解くことにより求めることができる。
The feature parameter value for the i-th analysis window in the n-th analysis frame is represented by a vector 〓o , i , and the distance between the feature vector 〓o , i and 〓o -1 , j is expressed as R(n, i | n-1 , j), the optimal feature parameter value series in each analysis interval can be obtained by solving the following recurrence equations (1) and (2).

S(n、i)=MinR(n、i|n‐1、1)+S(n‐1、1) R(n、i|n‐1、2)+S(n‐1、2) 〓 〓 R(n、i|n‐1、j)+S(n‐1、j) 〓 〓 R(n、i|n‐1、I)+S(n‐1、I) …(1) ここでi=1、2、…、I(但しIは分析窓の
数) n=1、2、…、N(Nは分析区間数) また、S(n、i)はn番目の分析フレームに
おけるi番目の分析窓に対する評価値であり距離
Rの積算値となつている。
S(n, i)=MinR(n, i|n‐1, 1)+S(n‐1, 1) R(n, i|n‐1, 2)+S(n‐1, 2) 〓 〓 R(n, i|n-1, j)+S(n-1, j) 〓 〓 R(n, i|n-1, I)+S(n-1, I) …(1) Here where i = 1, 2, ..., I (where I is the number of analysis windows) n = 1, 2, ..., N (N is the number of analysis sections), and S (n, i) is the number of analysis windows in the nth analysis frame. This is an evaluation value for the i-th analysis window, and is an integrated value of distance R.

Smin=MinS(N、1) S(N、2) 〓 S(N、I) ……(2) 最適な特徴パラメータ値系列は式(2)で表わされ
る最小評価値Sminに付随する(n、i)系列と
して決定される。
Smin=MinS(N, 1) S(N, 2) 〓 S(N, I) ...(2) The optimal feature parameter value series is associated with the minimum evaluation value Smin expressed by equation (2) (n, i) Determined as a series.

また、特徴パラメータ値としては例えばフオル
マント周波数を用いることができる。フオルマン
ト周波数は、例えば自己相関値を線形予測係数に
変換し、さらに線形予測係数を極周波数に変換す
る方法により求めることができる。この方法は下
記参照資料(1)に詳しいので、ここでは説明を省略
する。
Furthermore, for example, a formant frequency can be used as the characteristic parameter value. The formant frequency can be determined, for example, by converting an autocorrelation value into a linear prediction coefficient, and further converting the linear prediction coefficient into a polar frequency. This method is detailed in reference material (1) below, so the explanation will be omitted here.

資料(1):“Speech Analysis and Synthesis by
Linear Prediction of the Speech Wave”、 B、S、Atal and S.L.Hanauer、The
Journal of the Acoustical Society of
America、Vol.50、Num.2、1971 次に図面を用いて本発明を詳細に説明する。図
は本発明の一実施例を示すブロツク図である。
Material (1): “Speech Analysis and Synthesis by
Linear Prediction of the Speech Wave”, B. S. Atal and SL Hanauer, The
Journal of the Acoustical Society of
America, Vol. 50, Num. 2, 1971 Next, the present invention will be explained in detail using the drawings. The figure is a block diagram showing one embodiment of the present invention.

まず音声波形が音声波形入力端子1を介して音
声波形バツフア2に入力され一時記憶される。音
声波形バツフア2に一時記憶された音声波形は、
制御回路12より音声波形バツフア制御データ伝
送路6を介して出力される波形出力制御データに
従い窓回路3に入力される。窓回路3は制御回路
12より窓回路制御データ伝送路7を介して入力
される窓位置制御データに従つて前記音声波形に
対して時間位置の異なる窓をかけた波形を算出
し、順次フオルマントパラメータ抽出回路4に出
力する。フオルマント抽出回路4は前記、時間位
置の異なる窓をかけられた音声波形に対してそれ
ぞれフオルマント周波数を算出しフオルマントパ
ラメータ値一時記憶回路5に記憶させる。フオル
マントパラメータ値一時記憶回路5は、制御回路
12よりフオルマントパラメータ値一時記憶回路
制御データ伝送路9を介して与えられるフオルマ
ントパラメータ値出力データに従つて該分析区間
と直前の分析区間における二つのフオルマントパ
ラメータ値を距離算出回路11に出力する。距離
算出回路11は、制御回路12より距離算出回路
制御データ伝送路10を介して与えられる距離算
出回路制御データに従つて前記二つのフオルマン
トパラメータ値の距離(前記式(1)のRに対応す
る)を算出し、評価値算出回路16に出力する。
評価値算出回路16は制御回路12から評価値算
出回路制御データ伝送路13を介して与えられる
評価値算出回路制御データに従い、前記距離デー
タと、評価値一時記憶回路15より出力される前
段の評価値(式(1)におけるS(n−1、j)に対
応する)とから新たな評価値(式(1)におけるS
(n、i)に対応する)を算出し、最適パスデー
タ(式(1)において最小値を与えるjの値に対応す
る)を最適パスデータ一時記憶回路17に記憶さ
せるとともに、前記新たな評価値を評価値一時記
憶回路15に記憶させる。
First, a voice waveform is input to the voice waveform buffer 2 via the voice waveform input terminal 1 and temporarily stored. The audio waveform temporarily stored in the audio waveform buffer 2 is
The waveform output control data output from the control circuit 12 via the audio waveform buffer control data transmission line 6 is input to the window circuit 3. The window circuit 3 calculates a waveform obtained by multiplying the audio waveform by windows at different time positions according to the window position control data inputted from the control circuit 12 via the window circuit control data transmission line 7, and sequentially performs the filtering. It is output to the cloak parameter extraction circuit 4. The formant extraction circuit 4 calculates formant frequencies for each of the voice waveforms that are windowed at different time positions, and stores them in the formant parameter value temporary storage circuit 5. The formant parameter value temporary storage circuit 5 stores the analysis period and the previous analysis according to the formant parameter value output data given from the control circuit 12 via the formant parameter value temporary storage circuit control data transmission line 9. The two formant parameter values in the section are output to the distance calculation circuit 11. The distance calculation circuit 11 calculates the distance between the two formant parameter values (R in the above formula (1)) according to distance calculation circuit control data given from the control circuit 12 via the distance calculation circuit control data transmission line 10. corresponding) is calculated and output to the evaluation value calculation circuit 16.
The evaluation value calculation circuit 16 follows the evaluation value calculation circuit control data given from the control circuit 12 via the evaluation value calculation circuit control data transmission line 13, and calculates the distance data and the previous evaluation output from the evaluation value temporary storage circuit 15. value (corresponding to S(n-1, j) in equation (1)) to a new evaluation value (S in equation (1))
(corresponding to n, i)), and stores the optimal path data (corresponding to the value of j that gives the minimum value in equation (1)) in the optimal path data temporary storage circuit 17, and the new evaluation The value is stored in the evaluation value temporary storage circuit 15.

以上の処理を、あらかじめ定められたフレーム
数(前記式(1)、(2)のNに対応する)回だけ繰り返
した後、最適データアドレス生成回路18は制御
回路12より最適データアドレス生成回路制御デ
ータ伝送路14を介して与えられる制御データに
従つて、評価値一時記憶回路15より出力される
N番目の分析区間における最小値を検出し(前記
式(2)のSnioに対応する)た後、前記最小値(Snio
に付随する最適パスデータを最適パスデータ一時
記憶回路17より順次出力させることにより、各
分析区間における最適なフオルマントパラメータ
値に対するアドレスデータを生成しアドレスデー
タ伝送路19を介して、フオルマントパラメータ
値一時記憶回路5に出力する。フオルマントパラ
メータ値一時記憶回路5は前記アドレスデータに
従い該フオルマントパラメータ値を順次特徴パラ
メータ値出力端子20より出力させる。
After repeating the above process a predetermined number of frames (corresponding to N in equations (1) and (2) above), the optimal data address generation circuit 18 is controlled by the control circuit 12 to control the optimal data address generation circuit. According to the control data given via the data transmission path 14, the minimum value in the Nth analysis interval output from the evaluation value temporary storage circuit 15 is detected (corresponding to S nio in the above equation (2)). After that, the minimum value (S nio )
By sequentially outputting the optimal path data associated with the optimal path data from the optimal path data temporary storage circuit 17, address data for the optimal formant parameter value in each analysis interval is generated, and the formant is transmitted via the address data transmission line 19. It is output to the parameter value temporary storage circuit 5. The formant parameter value temporary storage circuit 5 sequentially outputs the formant parameter values from the characteristic parameter value output terminal 20 according to the address data.

以上の説明においては、特徴パラメータとし
て、フオルマント周波数を用いたが、特徴パラメ
ータとして自己相関値、フイルターバンク出力
値、線形予測係数等を用いても同様の効果の得ら
れる音声分析装置が実現できることは明らかであ
る。
In the above explanation, the formant frequency was used as the feature parameter, but it is possible to realize a speech analysis device with similar effects using autocorrelation values, filter bank output values, linear prediction coefficients, etc. as the feature parameters. it is obvious.

【図面の簡単な説明】[Brief explanation of drawings]

図は本発明の実施例を説明するためのブロツク
図である。図において、1は音声波形入力端子、
2は音声波形バツフア、3は窓回路、4はフオル
マントパラメータ抽出回路、5はフオルマントパ
ラメータ値一時記憶回路、6は音声波形バツフア
制御データ伝送路、7は窓回路制御データ伝送
路、8はフオルマントパラメータ抽出回路制御デ
ータ伝送路、9はフオルマントパラメータ一時記
憶回路制御データ伝送路、10は距離算出回路制
御データ伝送路、11は距離算出回路、12は制
御回路、13は評価値算出回路制御データ伝送
路、14は最適データアドレス生成回路制御デー
タ伝送路、15は評価値一時記憶回路、16は評
価値算出回路、17は最適パスデータ一時記憶回
路、18は最適データアドレス生成回路、19は
アドレスデータ伝送路、20は特徴パラメータ値
出力端子である。
The figure is a block diagram for explaining an embodiment of the present invention. In the figure, 1 is an audio waveform input terminal;
2 is an audio waveform buffer, 3 is a window circuit, 4 is a formant parameter extraction circuit, 5 is a formant parameter value temporary storage circuit, 6 is an audio waveform buffer control data transmission line, 7 is a window circuit control data transmission line, 8 is a formant parameter extraction circuit control data transmission line, 9 is a formant parameter temporary storage circuit control data transmission line, 10 is a distance calculation circuit control data transmission line, 11 is a distance calculation circuit, 12 is a control circuit, and 13 is a control circuit. 14 is an evaluation value calculation circuit control data transmission line, 14 is an optimum data address generation circuit control data transmission line, 15 is an evaluation value temporary storage circuit, 16 is an evaluation value calculation circuit, 17 is an optimum path data temporary storage circuit, and 18 is an optimum data address. 19 is an address data transmission path, and 20 is a characteristic parameter value output terminal.

Claims (1)

【特許請求の範囲】[Claims] 1 音声波形を分析周期毎に分析し各分析区間n
(=1、2、…、N)における音声の特徴パラメ
ータ値を抽出する音声分析装置において、各分析
区間毎に複数個の時間位置の異なる分析窓i(=
1、2、…、I)に対する特徴パラメータ値〓o,i
を算出する手段と、分析区間nにおいては、分析
窓iについて分析区間n−1においては分析窓j
(=1、2、…、I)について得られた特徴パラ
メータ値〓oi、〓o-1jとの距離R(n、i|n
−1、j)をすべてのi、j、nについて得る手
段と、分析区間n、分析窓iの各々については、
保持している分析区間n−1の分析区間n−1の
分析窓j(=1、2、…、I)による特徴パラメ
ータ値〓o-1jまでの距離累積値S(n−1、j)
の各々に、前記R(n、i|n−1、j)を加算
してR(n、i|n−1、j)+S(n−1、j);
(j=1、2、…、I)を得、このI個の値の中
で最小のものをこの分析区間nの分析窓iまでの
距離累積値S(n、i)とする処理を分析区間N
まで得ない、距離累積値S(N、i);(i=1、
2、…、I)を得、このS(N、i);(i=1、
2、…、N)の中で最小値を与える特徴パラメー
タ値の組合せを出力する手段とを有することを特
徴とする音声分析装置。
1 Analyze the audio waveform for each analysis period and analyze each analysis section n
In a speech analysis device that extracts feature parameter values of speech at (=1, 2, ..., N), a plurality of analysis windows i (=
1, 2, ..., I) Feature parameter values 〓 o,i
and means for calculating the analysis window j in the analysis interval n-1 for the analysis window i in the analysis interval n.
Distance R ( n , i | n
−1, j) for all i, j, n, and for each analysis interval n and analysis window i,
Feature parameter value by analysis window j (=1, 2, ..., I) of analysis section n-1 of analysis section n-1 held〓o -1 , cumulative distance value S(n-1, j)
The above R(n, i|n-1, j) is added to each of R(n, i|n-1, j)+S(n-1, j);
Analyze the process of obtaining (j = 1, 2, ..., I) and setting the minimum value among these I values as the cumulative distance value S (n, i) of this analysis interval n to the analysis window i. Section N
Distance cumulative value S(N, i); (i=1,
2,...,I), and this S(N,i); (i=1,
2, . . . , N) for outputting a combination of feature parameter values that gives the minimum value.
JP2998780A 1980-03-10 1980-03-10 Voice analyzer Granted JPS56126895A (en)

Priority Applications (1)

Application Number Priority Date Filing Date Title
JP2998780A JPS56126895A (en) 1980-03-10 1980-03-10 Voice analyzer

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
JP2998780A JPS56126895A (en) 1980-03-10 1980-03-10 Voice analyzer

Publications (2)

Publication Number Publication Date
JPS56126895A JPS56126895A (en) 1981-10-05
JPH0122638B2 true JPH0122638B2 (en) 1989-04-27

Family

ID=12291303

Family Applications (1)

Application Number Title Priority Date Filing Date
JP2998780A Granted JPS56126895A (en) 1980-03-10 1980-03-10 Voice analyzer

Country Status (1)

Country Link
JP (1) JPS56126895A (en)

Families Citing this family (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US4731846A (en) * 1983-04-13 1988-03-15 Texas Instruments Incorporated Voice messaging system with pitch tracking based on adaptively filtered LPC residual signal
JPS6155700A (en) * 1984-08-27 1986-03-20 富士通株式会社 Pitch extraction processing system
JPS61156182A (en) * 1984-12-28 1986-07-15 函館工業高等専門学校長 Intonation display unit

Also Published As

Publication number Publication date
JPS56126895A (en) 1981-10-05

Similar Documents

Publication Publication Date Title
JP3167787B2 (en) Digital speech coder
JPH0736475A (en) Reference pattern formation method in voice analysis
KR880700387A (en) Speech processing system and voice processing method
RU2510954C2 (en) Method of re-sounding audio materials and apparatus for realising said method
EP1995723A1 (en) Neuroevolution training system
JP3255190B2 (en) Speech coding apparatus and its analyzer and synthesizer
JP2798003B2 (en) Voice band expansion device and voice band expansion method
US4873723A (en) Method and apparatus for multi-pulse speech coding
JPS62229200A (en) Pitch detector
Hasan et al. An approach to voice conversion using feature statistical mapping
CN114974271B (en) Voice reconstruction method based on sound channel filtering and glottal excitation
JP2539351B2 (en) Speech synthesis method
JPS62102294A (en) Voice coding system
JP3288052B2 (en) Fundamental frequency extraction method
JPH0754438B2 (en) Voice processor
JPS5965895A (en) Voice synthesization
JPH0679238B2 (en) Pitch extractor
JP2515609B2 (en) Speaker recognition method
Funaki et al. A time varying ARMAX speech modeling with phase compensation using glottal source model
JP3414637B2 (en) Articulatory parameter time-series extracted speech analysis method, device thereof, and program recording medium
JPS61256400A (en) Speech analysis and synthesis method
JPH0448239B2 (en)
JP3263136B2 (en) Signal pitch synchronous position extraction method and signal synthesis method
Hacioglu et al. Pulse-by-pulse reoptimization of the synthesis filter in pulse-based coders
JPH0736119B2 (en) Piecewise optimal function approximation method