JPH0284699A - Voice analyzing and synthesizing device - Google Patents

Voice analyzing and synthesizing device

Info

Publication number
JPH0284699A
JPH0284699A JP63237102A JP23710288A JPH0284699A JP H0284699 A JPH0284699 A JP H0284699A JP 63237102 A JP63237102 A JP 63237102A JP 23710288 A JP23710288 A JP 23710288A JP H0284699 A JPH0284699 A JP H0284699A
Authority
JP
Japan
Prior art keywords
sound source
waveform
circuit
analysis
synthesis
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Granted
Application number
JP63237102A
Other languages
Japanese (ja)
Other versions
JP2650355B2 (en
Inventor
Hirohisa Tazaki
裕久 田崎
Kunio Nakajima
中島 邦男
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Mitsubishi Electric Corp
Original Assignee
Mitsubishi Electric Corp
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Mitsubishi Electric Corp filed Critical Mitsubishi Electric Corp
Priority to JP63237102A priority Critical patent/JP2650355B2/en
Publication of JPH0284699A publication Critical patent/JPH0284699A/en
Application granted granted Critical
Publication of JP2650355B2 publication Critical patent/JP2650355B2/en
Anticipated expiration legal-status Critical
Expired - Fee Related legal-status Critical Current

Links

Abstract

PURPOSE:To transmit the waveforms with a relatively small information quantity and to improve the sound quality which is like the buzzer sound specific to a VOCODER system and lacks personality and naturalness by using optimum sound source vectors to form the sound source waveforms within an analysis frame and forming synthetic waves. CONSTITUTION:A representative sound source forming circuit 8 segments the representative sound source waveforms 2b of a pitch period length from the predicted residual waveforms 2a within the analysis frame of the specified length of the vocal voices. Sound source selecting circuits 10, 11 select the optimum sound source vectors to obtain the synthesized waveform synthesized from this representative sound source by using the envelope parameter of the above-mentioned frame and the synthesized waveform of the smallest waveform distortion from the sound source vector groups in code note memories 9, 18. A sound source forming circuit 19 forms the sound source waveforms 2d by arraying the optimum sound source vectors at every pitch period within the above-mentioned analysis frame. Synthesizing filter circuits 10, 11, 22 and driven by the predicted residual signal 2a, by which the efficient transmission of the predicted residual signal is executed. The synthetic sounds of good quality are efficiently synthesized in this way with the relatively small transmission information quantity and the synthetic sounds which are natural and have the ample personality is obtd.

Description

【発明の詳細な説明】 〔産業上の利用分野〕 本発明は音声分析合成装置に関するものである。[Detailed description of the invention] [Industrial application field] The present invention relates to a speech analysis and synthesis device.

〔従来の技術〕[Conventional technology]

音声の分析合成により音声情報の圧縮を行う方法として
、分析側で入力音声波形の周波数スペクトラム包絡形状
を表す特徴パラメータ(以降包絡パラメータと呼ぶ)と
ピンチ周期を抽出し、合成側において前記包絡パラメー
タとピッチ周期を用いて合成音声波形を生成する方式(
ボコーダと呼ばれている)が知られている。このボコー
ダ方式では、分析部において有声/無声判別を行い、有
声音と判別した場合にはさらにピンチ抽出を行う。
As a method for compressing audio information by analyzing and synthesizing audio, the analysis side extracts the characteristic parameters (hereinafter referred to as envelope parameters) representing the frequency spectrum envelope shape of the input audio waveform and the pinch period, and the synthesis side extracts the pinch period and the frequency spectrum envelope shape of the input audio waveform. A method for generating synthetic speech waveforms using pitch periods (
(called a vocoder) is known. In this vocoder method, voiced/unvoiced discrimination is performed in the analysis section, and if it is determined that the sound is voiced, pinch extraction is further performed.

そして合成部において、前記有声/無声判別結果が有声
音である場合にはピッチ周期で繰り返すインパルス列、
無声音である場合には白色雑音を駆動音源として、包絡
パラメータを用いた合成フィルタを駆動することで合成
音声を得る。この方式は、比較的少ない伝送情報量で実
用上十分な明瞭性を持つ合成音声が得られる利点がある
ものの、人力音声に比べ個人性、自然性の欠落した貧弱
な音質であり、しばしば長時間の聴取に耐えないブザー
音を伴う欠点を持つ。
Then, in the synthesis section, if the voiced/unvoiced discrimination result is a voiced sound, an impulse train that repeats at a pitch period;
If the sound is unvoiced, synthesized speech is obtained by using white noise as a driving sound source and driving a synthesis filter using an envelope parameter. Although this method has the advantage of producing synthesized speech with sufficient clarity for practical use with a relatively small amount of transmitted information, the sound quality is poor and lacks individuality and naturalness compared to human-generated speech, and it is often used for long periods of time. It has the disadvantage of producing a buzzer sound that cannot be heard.

従来この音質改善法としては、例えば下記資料(1)に
示されるようなインパルス列等価音源を用いた有声音源
の改良法がある。
As a conventional method for improving sound quality, there is a method for improving a voiced sound source using an impulse train equivalent sound source as shown in the following document (1), for example.

[高品質音声合成のためのインパルス列等価音源」電子
通信学会論文誌(1985−11)Vol、J68−^
N 0 、11−−−−−− (11有声音区間の予測
残差信号は分析フレーム長全体としてはその周波数スペ
クトルはほぼ平坦であるが、ピンチ周期長以下のPi 
FJtA時間の周波数スペクトルは高域強調区間と低域
強調区間がピッチ周期毎に繰り返す構造を持つ。資料f
l) (71)方式においては、分析フレーム全体の周
波数スペクトラムの平坦性を保ちつつ、極短時間の周波
数スペクトルにピッチ周期毎に繰り返す変動を与えた音
声音源(インパルス列等価音源と呼ぶ)を生成し、これ
を用いることで合成音声の音質に大きく寄竪する有声音
区間の品質を改善するというものであった。
[Impulse train equivalent sound source for high-quality speech synthesis] Transactions of the Institute of Electronics and Communication Engineers (1985-11) Vol. J68-^
N 0 , 11 --- (11) Although the frequency spectrum of the predicted residual signal of the voiced section is almost flat for the entire analysis frame length, Pi
The frequency spectrum of the FJtA time has a structure in which a high-frequency emphasis section and a low-frequency emphasis section repeat every pitch period. Material f
l) In the (71) method, a speech sound source (called an impulse train equivalent sound source) is generated in which the frequency spectrum of an extremely short period of time is given fluctuations that repeat every pitch period while maintaining the flatness of the frequency spectrum of the entire analysis frame. However, by using this, the quality of voiced sound sections, which greatly interferes with the sound quality of synthesized speech, could be improved.

第3°図はこの資料(11に示された従来の方式を表す
ブロック図である。まず音声波形が音声波形入力端子3
を介して分析部1内の包絡パラメータ抽出回路4と有声
/無声判別回路5とピンチ抽出回路6にそれぞれ入力さ
れる。包絡パラメータ抽出回路4は前記音声波形より包
絡パラメータの算出を行い、そのパラメータを包絡パラ
メータ伝送路15を介して合成部2内の合成フィルタ回
路22に伝送する。有声/無声判別回路5は前記音声波
形が有声音区間であるか否かを判別し、判別結果をピッ
チ抽出回路6に出力し、さらに有声/無声判別結果伝送
路16を介して合成部2内の音源切換回路21へ出力す
る。ピンチ抽出回路6は前記有声/無声判別結果が有声
音区間である場合に音声波形からピッチ周期分析を行い
、抽出されたと・ノチデータをピッチデータ伝送路17
を介して合成部2内のインパルス列等価音源生成回路2
4に出力する。合成部2内のインパルス列等価音源生成
回路24は予め与えられるlピッチ周期長のインパルス
列等価音源を前記ピッチ周期毎に繰り返し生成し音源切
換回路21に出力する。無声音源生成回路20は白色雑
音の生成を行い音源切換回路21に出)Jする。音源切
換回路21は前記有声/無声判別結果が有声音である場
合にインパルス列等価音源を遺沢し、有声音以外の場合
には白色雑音を選択し、選択した音源波形を合成フィル
タ回路22に出力する。そして合成フィルタ回路22は
、分析部l内の包絡パラメータ抽出回路4より入力され
た包絡パラメータ及び前記音源波形を用いて合成波形を
生成し合成波形出力端子23より出力するというもので
ある。
Figure 3 is a block diagram showing the conventional method shown in this document (11).First, the audio waveform is input to the audio waveform input terminal 3.
The signal is inputted to the envelope parameter extraction circuit 4, the voiced/unvoiced discrimination circuit 5, and the pinch extraction circuit 6 in the analysis section 1, respectively. The envelope parameter extraction circuit 4 calculates envelope parameters from the audio waveform, and transmits the parameters to the synthesis filter circuit 22 in the synthesis section 2 via the envelope parameter transmission path 15. The voiced/unvoiced discrimination circuit 5 discriminates whether or not the speech waveform is a voiced sound section, outputs the discrimination result to the pitch extraction circuit 6, and further outputs the discrimination result to the synthesis unit 2 via the voiced/unvoiced discrimination result transmission line 16. It is output to the sound source switching circuit 21 of. The pinch extraction circuit 6 performs pitch period analysis from the voice waveform when the voiced/unvoiced discrimination result is a voiced section, and transmits the extracted pinch data to the pitch data transmission line 17.
Impulse train equivalent sound source generation circuit 2 in synthesis unit 2 via
Output to 4. The impulse train equivalent sound source generating circuit 24 in the synthesizing section 2 repeatedly generates an impulse train equivalent sound source having a predetermined l pitch period length for each pitch period and outputs it to the sound source switching circuit 21 . The unvoiced sound source generation circuit 20 generates white noise and outputs it to the sound source switching circuit 21). The sound source switching circuit 21 leaves behind the impulse train equivalent sound source when the voiced/unvoiced discrimination result is a voiced sound, selects white noise when the voiced sound is not a voiced sound, and sends the selected sound source waveform to the synthesis filter circuit 22. Output. The synthesis filter circuit 22 generates a synthesized waveform using the envelope parameter inputted from the envelope parameter extraction circuit 4 in the analysis section 1 and the sound source waveform, and outputs it from the synthesized waveform output terminal 23.

(発明が解決しようとするIKg) 以上説明したインパルス列等価音源を用いる従来の音声
分析合成装置では、合成音のブザー音的音質の若干の低
減効果が得られるが、本来、話者、音韻により変化する
音源を固定音源で表しているため、個人性、自然性の回
復はほとんどなく、十分な音質改善は得られないという
課題があった。
(IKg to be solved by the invention) In the conventional speech analysis and synthesis device using the impulse train equivalent sound source described above, it is possible to obtain a slight reduction effect on the buzzer-like sound quality of the synthesized speech, but originally it depends on the speaker and phoneme. Since a changing sound source is represented by a fixed sound source, there is little recovery of individuality and naturalness, and there is a problem in that sufficient sound quality improvement cannot be obtained.

本発明の目的は、予測残差波形に含まれる音源情報を比
較的少ない情報量で伝送し、ボコーダ方式特有のブザー
音的で、個人性、自然性の欠落した音質を改善すること
にある。
An object of the present invention is to transmit the sound source information included in the predicted residual waveform with a relatively small amount of information, and to improve the sound quality that is characteristic of the vocoder method, which has a buzzing sound and lacks individuality and naturalness.

〔課題を解決するための手段〕 本発明に係る音声分析合成装置は、一定長の分析フレー
ム毎に該当フレームが有声音区間であると判定された場
合にそのフレームの予測残差波形から1ピッチ周期長の
代表音源波形を切り出す代表音源抽出回路と、有限個の
音源ベクトルを記憶する符号帳メモリと、この音源ベク
トルの中から前記代表音源抽出回路で抽出された代表音
源に最も等価な音源ベクトルを最適音源ベクトルとして
選択する音源選択回路を備え、合成部において前記最適
音源ベクトルを用いて当該分析フレームの内の音源波形
を生成する音源生成回路と、この音源生成回路で生成さ
れた音源波形により合成波形を生成する合成フィルタを
備えるように構成したものである。
[Means for Solving the Problems] The speech analysis and synthesis device according to the present invention extracts one pitch from the predicted residual waveform of the frame when it is determined that the frame is a voiced section for each analysis frame of a certain length. A representative sound source extraction circuit that extracts a representative sound source waveform with a period length, a codebook memory that stores a finite number of sound source vectors, and a sound source vector that is most equivalent to the representative sound source extracted by the representative sound source extraction circuit from among these sound source vectors. a sound source selection circuit that selects a sound source vector as an optimal sound source vector, a sound source generation circuit that generates a sound source waveform in the analysis frame using the optimal sound source vector in a synthesis section; It is configured to include a synthesis filter that generates a synthesized waveform.

〔作用〕[Effect]

この発明における代表音源生成回路は有声音声の一定長
の分析フレーム内の予測残差波形からピンチ周期長の代
表音源波形を切り出し、音源選択回路は当該フレームの
包絡パラメータを用いてこの代表音源から合成される合
成波形と最も波形歪が小さい合成波形を得る最適音源ベ
クトルを符号帳メモリ内の音源ベクトル群より選択し、
音源生成回路はこの最適音源ベクトルを当該分析フレー
ム内でピッチ周期毎に並べることで音源波形を生成する
The representative sound source generation circuit in this invention cuts out a representative sound source waveform with a pinch period length from the predicted residual waveform in an analysis frame of a certain length of voiced speech, and the sound source selection circuit synthesizes it from this representative sound source using the envelope parameter of the frame. selects an optimal sound source vector from the group of sound source vectors in the codebook memory to obtain a combined waveform with the smallest waveform distortion and a combined waveform with the smallest waveform distortion;
The sound source generation circuit generates a sound source waveform by arranging the optimal sound source vectors for each pitch period within the analysis frame.

予測残差信号により合成フィルタ回路を駆動することで
人力音声信号が再生されることから、この予測残差信号
を効率よく伝送することができればより入力音声信号に
近い合成音が得られることは明らかである。ボコーダの
音質に大きく影響する有声音区間における予測残差波形
を視ると、概形が良(領た波形がピッチ周期で繰り返す
構造を持っており、1分析フレーム内において一つのピ
ッチ長残差信号のみを切り出して伝送し、合成部ではこ
れをピンチ周期毎とに並べることで音源波形を生成する
ことにすれば、この予測残差信号に含まれる冗長度を大
幅に削減して伝送することができる。さらに予め用意し
た符号長ベクトル内から切り出して予測残差信号に対す
る最適音源ベクトルを選択する方式を用いることにより
、1分析フレーム当り数ビットの伝送情報量の付加によ
り効率よ(この音源情報の伝送が実現できる。
Since the human voice signal is reproduced by driving the synthesis filter circuit with the prediction residual signal, it is clear that if this prediction residual signal can be efficiently transmitted, a synthesized sound that is closer to the input voice signal can be obtained. It is. Looking at the predicted residual waveform in the voiced interval, which greatly affects the sound quality of the vocoder, the approximate shape is good (the predicted waveform has a structure that repeats with a pitch period, and one pitch length residual within one analysis frame). If only the signal is cut out and transmitted, and the synthesizer generates the sound source waveform by arranging it for each pinch period, it is possible to significantly reduce the redundancy contained in this prediction residual signal and transmit it. Furthermore, by using a method that selects the optimal excitation vector for the prediction residual signal by cutting out the code length vector prepared in advance, efficiency can be improved by adding several bits of transmission information per analysis frame (this excitation information transmission can be realized.

〔発明の実施例〕[Embodiments of the invention]

以下、この発明の一実施例を第1図及び第2図について
説明する0図において、7は逆フィルタ回路、8は代表
音源抽出回路、9及び18は符号帳メモリ、10.11
及び22は合成フィルタ回路、12は波形歪算出回路、
13は比較回路、14は音源ベクトル番号伝送路、19
は1フレーム長音源生成回路である。1〜6.15〜1
7.20〜23は従来例と同じであるので説明を省略す
る。2aは予測残差波形、2bは代表音源波形、2Cは
選択音源ヘクトル、2dは1フレーム長音源波形である
Hereinafter, one embodiment of the present invention will be explained with reference to FIGS. 1 and 2. In FIG. 0, 7 is an inverse filter circuit, 8 is a representative sound source extraction circuit, 9 and 18 are codebook memories, and 10.11
and 22 is a synthesis filter circuit, 12 is a waveform distortion calculation circuit,
13 is a comparison circuit, 14 is a sound source vector number transmission line, 19
is a one frame long sound source generation circuit. 1~6.15~1
7. 20 to 23 are the same as the conventional example, so their explanation will be omitted. 2a is a predicted residual waveform, 2b is a representative sound source waveform, 2C is a selected sound source vector, and 2d is a 1-frame long sound source waveform.

分析部1内の逆フィルタ回路7は、包絡パラメータ抽出
回路4によって抽出された包絡パラメータを用いて、入
力端子3を介して入力された音声波形に逆フィルタ処理
を行い、得られた予測残差波形2aを代表音源抽出回路
8に出力する。代表音源抽出回路8は、有声/無声判別
回路5がを声音区間であると判別した場合に前記予測残
差波形2aから1ピッチ周期長の残差波形を代表音源波
形2bとして切り出し、得られた代表音源波形2bを合
成フィルタ回路lOに出力する。切り出し処理の方法と
しては例えば残差波形中振幅最大の部分の直前のゼロ交
差点から1ピンチ周期長の波形の切り出しを行う、予測
残差波形の中から振幅の大きな部分が先頭部近傍にくる
ように切り出すことで、記憶しておく音源ベクトルとし
て振幅の大きな部分が先頭部近傍にあるものだけにする
ことができ、音源ベクトル群の数を少なくできる0分析
部l内の符号長メモリ9は、記憶している有限個の音源
ベクトルを順次合成フィルタ回路11に出力する。
The inverse filter circuit 7 in the analysis unit 1 performs inverse filter processing on the speech waveform input via the input terminal 3 using the envelope parameters extracted by the envelope parameter extraction circuit 4, and calculates the obtained prediction residual. The waveform 2a is output to the representative sound source extraction circuit 8. When the voiced/unvoiced discrimination circuit 5 determines that the voiced/unvoiced segment is a voiced sound section, the representative sound source extraction circuit 8 extracts a residual waveform of one pitch period length from the predicted residual waveform 2a as a representative sound source waveform 2b. The representative sound source waveform 2b is output to the synthesis filter circuit IO. An example of a method for cutting out processing is to cut out a waveform with a length of one pinch cycle from the zero intersection immediately before the part with the maximum amplitude in the residual waveform, or to cut out a waveform with a length of one pinch cycle so that the part with the largest amplitude in the predicted residual waveform is near the beginning. The code length memory 9 in the analysis unit 1 can reduce the number of sound source vector groups by cutting out the sound source vectors so that only the sound source vectors that have large amplitude parts near the beginning are stored. The stored finite number of sound source vectors are sequentially output to the synthesis filter circuit 11.

前記合成フィルタ回路10及び11は、各々入力された
音源ベクトルと上記包絡パラメータを用いて合成波形を
生成し、その合成波形を波形歪算出回路12に出力する
。波形歪算出回路12は、代表音源波形2aを用いて得
られた合成波形と符号帳メモリ9に記憶されていた各音
源ベクトルを用いて得られた合成波形の間の波形歪を算
出し、その結果を比較回路13に出力する。比較回路1
3は前記波形歪を比較し、最小の波形歪を与えた音源ベ
クトルの番号を音源ベクトル番号伝送路14を介して合
成部2内の符号帳メモIJ18に出力する。符号帳メモ
リ18は入力された音源ベクトル番号により指定された
選択音源ベクトル2cを1フレーム長音源生成回路19
に出力する。1フレーム長音源生成回路19は前記選択
音源ベクトル2cを、ピッチデータ伝送路17を介して
入力されたピッチ周期間隔で並べ立てることで1フレー
ム長音源波形2dを生成し、これを音源切換回路21に
出力する。音源切換回路21は有声/無声判別結果伝送
路16を介して入力された有声/無声判別結果が有声音
である場合に前記1フレーム長音源波形2dを選択し、
有声音以外の場合には白色雑音を選択し、選択した音源
波形を合成フィルタ回路22に出力する。そして合成フ
ィルタ回路22は、分析部1内の包絡パラメータ抽出回
路4より入力された包絡パラメータ及び前記音源波形を
用いて合成波形を生成し合成波形出力端子23より出力
する。
The synthesis filter circuits 10 and 11 each generate a synthesized waveform using the input sound source vector and the envelope parameter, and output the synthesized waveform to the waveform distortion calculation circuit 12. The waveform distortion calculation circuit 12 calculates the waveform distortion between the composite waveform obtained using the representative sound source waveform 2a and the composite waveform obtained using each sound source vector stored in the codebook memory 9, and The result is output to the comparison circuit 13. Comparison circuit 1
3 compares the waveform distortions, and outputs the number of the excitation vector giving the minimum waveform distortion to the codebook memo IJ 18 in the synthesis section 2 via the excitation vector number transmission line 14. The codebook memory 18 converts the selected excitation vector 2c specified by the input excitation vector number into a one frame length excitation generator circuit 19.
Output to. The one-frame long sound source generation circuit 19 generates a one-frame long sound source waveform 2d by arranging the selected sound source vectors 2c at pitch period intervals inputted via the pitch data transmission line 17, and sends this to the sound source switching circuit 21. Output. The sound source switching circuit 21 selects the one frame long sound source waveform 2d when the voiced/unvoiced discrimination result inputted via the voiced/unvoiced discrimination result transmission line 16 is a voiced sound,
If the sound is not a voiced sound, white noise is selected and the selected sound source waveform is output to the synthesis filter circuit 22. Then, the synthesis filter circuit 22 generates a synthesized waveform using the envelope parameter input from the envelope parameter extraction circuit 4 in the analysis section 1 and the sound source waveform, and outputs it from the synthesized waveform output terminal 23.

上記符号帳メモリ9内の有限個の音源ベクトルは、多く
の入力音声信号より上記分析部2内の代表音源抽出回路
8を用いて抽出した代表音源波形の集合中から所望の個
数だけ選択して予め容易する。その選択の方法としては
、例えば有声音区間の平均的なスペクトル包絡形状を持
つ合成フィルタ回路を構成し、この合成フィルタ回路を
前記代表音源波形の集合で駆動して得られる合成波形に
おける歪を最小にする基準のクラスタリング手法を用い
ることができる。
A desired number of finite sound source vectors in the codebook memory 9 are selected from a set of representative sound source waveforms extracted from many input audio signals using the representative sound source extraction circuit 8 in the analysis section 2. Make it easy in advance. The selection method is, for example, to configure a synthesis filter circuit with an average spectral envelope shape of a voiced sound section, and to minimize the distortion in the synthesized waveform obtained by driving this synthesis filter circuit with the set of representative sound source waveforms. A clustering method based on the following criteria can be used.

〔他の実施例の説明、他の用途への転用例の説明〕上記
実施例では、各演算処理を回路内で実現する例について
述べたが、これを信号処理プロセッサ等の汎用演算装置
によるソフトウェア処理により実現してもよい、また、
音源選択法として合成フィルタ出力波形における歪を最
小にする音源を選択する方式を述べたが、これを包絡パ
ラメータ及び音源波形をDFT等の周波数スペクトルで
表しその周波数軸上における歪を最小にするように選択
する方式とすることもできる。
[Explanation of other embodiments, explanation of examples of diversion to other uses] In the above embodiments, an example was described in which each arithmetic process is realized within a circuit, but this can be implemented using software using a general-purpose arithmetic device such as a signal processing processor. It may be realized by processing, and
As a sound source selection method, we have described a method of selecting a sound source that minimizes the distortion in the synthesis filter output waveform, but this can be done by representing the envelope parameter and the sound source waveform as a frequency spectrum such as DFT, and minimizing the distortion on the frequency axis. It is also possible to select a method.

〔発明の効果〕〔Effect of the invention〕

以上のように、この発明によれば、比較的少ない伝送情
報量により効率よく品質のよい合成音を合成できる有声
音音源情報の伝送が可能であり、インパルス列等等価音
源等の従来の固定音源を用いた装置で問題であったボコ
ーダの合成音的な音が改善されたより自然で個人性豊か
な合成音が得られる。
As described above, according to the present invention, it is possible to transmit voiced sound source information that can efficiently synthesize a high-quality synthesized sound with a relatively small amount of transmitted information, and it is possible to transmit voiced sound source information that can efficiently synthesize high-quality synthesized sounds with a relatively small amount of transmitted information, and it is possible to transmit voiced sound source information that can efficiently synthesize high-quality synthesized sounds with a relatively small amount of transmitted information. The synthesized sound of the vocoder, which was a problem with devices using the vocoder, has been improved and a more natural and individualized synthesized sound can be obtained.

【図面の簡単な説明】[Brief explanation of drawings]

第1図はこの発明の1実施例による音声分析合成装置を
示すブロック図、第2図はその実施例における代表音源
波形抽出と1フレーム長音源波形生成の様子を示す模式
図、第3図は従来の音声分析合成装置を示すブロック図
である・ 図において1は分析部、2は合成部、3は音声波形入力
端子、4は包絡パラメータ抽出回路、51よ有声/無声
判別回路、6はピッチ抽出回路、7は逆フィルタ回路、
8は代表音源抽出回路、9及び18は符号帳メモリ、1
0.11及び22は合成フィルタ回路、12は波形歪算
出回路、13は比較回路、14は音源ヘクトル番号伝送
路、15は包絡パラメータ伝送路、16は有声/無声判
別結果伝送路、17はピッチデータ伝送路、19は1フ
レーム長音源生成回路、20は無声音源生成回路、21
は音源切換回路、23は合成波形出力端子、24はイン
パルス列等価音′a生成回路である。2aは予測残差波
形、2bは代表音源波形、2cは選択音源ベクトル、2
dは1フレーム長音源波形である。 代理人   大  岩  増  雄 第2図
FIG. 1 is a block diagram showing a speech analysis and synthesis device according to an embodiment of the present invention, FIG. 2 is a schematic diagram showing representative sound source waveform extraction and one-frame long sound source waveform generation in the embodiment, and FIG. 1 is a block diagram showing a conventional speech analysis and synthesis device. In the figure, 1 is an analysis section, 2 is a synthesis section, 3 is a speech waveform input terminal, 4 is an envelope parameter extraction circuit, 51 is a voiced/unvoiced discrimination circuit, and 6 is a pitch extraction circuit, 7 is an inverse filter circuit,
8 is a representative sound source extraction circuit, 9 and 18 are codebook memories, 1
0.11 and 22 are synthesis filter circuits, 12 is a waveform distortion calculation circuit, 13 is a comparison circuit, 14 is a sound source vector number transmission line, 15 is an envelope parameter transmission line, 16 is a voiced/unvoiced discrimination result transmission line, and 17 is a pitch a data transmission line; 19 is a 1 frame length sound source generation circuit; 20 is an unvoiced sound source generation circuit; 21
2 is a sound source switching circuit, 23 is a composite waveform output terminal, and 24 is an impulse train equivalent sound 'a generating circuit. 2a is the predicted residual waveform, 2b is the representative sound source waveform, 2c is the selected sound source vector, 2
d is a one frame long sound source waveform. Agent Masuo Oiwa Figure 2

Claims (1)

【特許請求の範囲】[Claims] 音声波形の情報量圧縮を行う音声分析合成装置において
、入力音声波形を一定長の分析フレーム単位にこの入力
音声波形のスペクトル包絡パラメータを用いて逆フィル
タリングし、当該分析フレームの予測残差波形を求める
逆フィルタ回路と、この逆フィルタ回路で求められた予
測残差波形から、この予測残差波形の持つ一ピッチ周期
長分を切り出し当該分析フレームの代表音源波形とする
代表音源抽出回路と、有限個の音源ベクトルを記憶する
符号帳メモリと、前記周波数特徴パラメータを用いて構
成された合成フィルタ回路と、この合成フィルタ回路を
前記代表音源抽出回路で切り出された代表音源波形で駆
動して得る合成波形とこの合成フィルタを前記符号帳メ
モリ内の音源ベクトルで駆動して得る合成波形との歪が
最小化するように前記符号幅メモリ内から最適音源ベク
トルを選択する音源選択回路を分析部に備え、この分析
部で求められた最適音源ベクトルを当該分析フレーム内
でピッチ周期毎に繰り返し並べることで当該分析フレー
ムの音源波形を生成する1フレーム長音源生成回路と、
この1フレーム長音源生成回路で生成された音源波形を
駆動源とし前記スペクトル包絡パラメータを用いて合成
波形を求める合成フィルタ回路を合成部に備えることを
特徴とする音声分析合成装置
In a speech analysis and synthesis device that compresses the amount of information in a speech waveform, the input speech waveform is inverse filtered in units of analysis frames of a certain length using the spectral envelope parameters of this input speech waveform, and the predicted residual waveform of the analysis frame is obtained. an inverse filter circuit, a representative sound source extraction circuit that extracts one pitch cycle length of the predicted residual waveform obtained by the inverse filter circuit and uses it as a representative sound source waveform of the analysis frame; a codebook memory that stores sound source vectors; a synthesis filter circuit configured using the frequency feature parameters; and a synthesis waveform obtained by driving this synthesis filter circuit with the representative sound source waveform extracted by the representative sound source extraction circuit. The analysis unit includes an excitation selection circuit that selects an optimal excitation vector from within the code width memory so as to minimize distortion between the synthesis filter and a synthesized waveform obtained by driving the synthesis filter with an excitation vector within the codebook memory; a 1-frame long sound source generation circuit that generates a sound source waveform for the analysis frame by repeatedly arranging the optimal sound source vectors determined by the analysis unit for each pitch cycle within the analysis frame;
A speech analysis and synthesis device characterized in that the synthesis section includes a synthesis filter circuit that uses the sound source waveform generated by the one-frame long sound source generation circuit as a driving source and obtains a synthesized waveform using the spectral envelope parameter.
JP63237102A 1988-09-21 1988-09-21 Voice analysis and synthesis device Expired - Fee Related JP2650355B2 (en)

Priority Applications (1)

Application Number Priority Date Filing Date Title
JP63237102A JP2650355B2 (en) 1988-09-21 1988-09-21 Voice analysis and synthesis device

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
JP63237102A JP2650355B2 (en) 1988-09-21 1988-09-21 Voice analysis and synthesis device

Publications (2)

Publication Number Publication Date
JPH0284699A true JPH0284699A (en) 1990-03-26
JP2650355B2 JP2650355B2 (en) 1997-09-03

Family

ID=17010443

Family Applications (1)

Application Number Title Priority Date Filing Date
JP63237102A Expired - Fee Related JP2650355B2 (en) 1988-09-21 1988-09-21 Voice analysis and synthesis device

Country Status (1)

Country Link
JP (1) JP2650355B2 (en)

Cited By (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPH06250694A (en) * 1993-02-25 1994-09-09 Idou Tsushin Syst Kaihatsu Kk Voice coding and decoding device
JP2003522965A (en) * 1998-12-21 2003-07-29 クゥアルコム・インコーポレイテッド Periodic speech coding

Citations (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPS634300A (en) * 1986-06-24 1988-01-09 日本電気株式会社 Voice encoding method and apparatus
JPS6337724A (en) * 1986-07-31 1988-02-18 Fujitsu Ltd Coding transmitter

Patent Citations (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPS634300A (en) * 1986-06-24 1988-01-09 日本電気株式会社 Voice encoding method and apparatus
JPS6337724A (en) * 1986-07-31 1988-02-18 Fujitsu Ltd Coding transmitter

Cited By (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPH06250694A (en) * 1993-02-25 1994-09-09 Idou Tsushin Syst Kaihatsu Kk Voice coding and decoding device
JP2003522965A (en) * 1998-12-21 2003-07-29 クゥアルコム・インコーポレイテッド Periodic speech coding
JP4824167B2 (en) * 1998-12-21 2011-11-30 クゥアルコム・インコーポレイテッド Periodic speech coding

Also Published As

Publication number Publication date
JP2650355B2 (en) 1997-09-03

Similar Documents

Publication Publication Date Title
US7149682B2 (en) Voice converter with extraction and modification of attribute data
KR100427753B1 (en) Method and apparatus for reproducing voice signal, method and apparatus for voice decoding, method and apparatus for voice synthesis and portable wireless terminal apparatus
JP3328080B2 (en) Code-excited linear predictive decoder
JPH10307599A (en) Waveform interpolating voice coding using spline
WO1980002211A1 (en) Residual excited predictive speech coding system
US6064955A (en) Low complexity MBE synthesizer for very low bit rate voice messaging
JPH10319996A (en) Efficient decomposition of noise and periodic signal waveform in waveform interpolation
AU6125594A (en) Method for generating a spectral noise weighting filter for use in a speech coder
JP3439307B2 (en) Speech rate converter
US6003000A (en) Method and system for speech processing with greatly reduced harmonic and intermodulation distortion
JPH0284699A (en) Voice analyzing and synthesizing device
Acero Source-filter models for time-scale pitch-scale modification of speech
JP2615548B2 (en) Highly efficient speech coding system and its device.
JP2841797B2 (en) Voice analysis and synthesis equipment
CN116110424B (en) Voice bandwidth expansion method and related device
JP2829978B2 (en) Audio encoding / decoding method, audio encoding device, and audio decoding device
JP3515216B2 (en) Audio coding device
JP2560682B2 (en) Speech signal coding / decoding method and apparatus
JP3515215B2 (en) Audio coding device
JPWO2003042648A1 (en) Speech coding apparatus, speech decoding apparatus, speech coding method, and speech decoding method
JP3218680B2 (en) Voiced sound synthesis method
JP2508002B2 (en) Speech coding method and apparatus thereof
JP2844590B2 (en) Audio coding system and its device
JPH09258796A (en) Voice synthesis method
JPH0836397A (en) Speech synthesizer

Legal Events

Date Code Title Description
LAPS Cancellation because of no payment of annual fees