JPH1091193A - Voice coding method and voice decoding method - Google Patents

Voice coding method and voice decoding method

Info

Publication number
JPH1091193A
JPH1091193A JP8246443A JP24644396A JPH1091193A JP H1091193 A JPH1091193 A JP H1091193A JP 8246443 A JP8246443 A JP 8246443A JP 24644396 A JP24644396 A JP 24644396A JP H1091193 A JPH1091193 A JP H1091193A
Authority
JP
Japan
Prior art keywords
signal
codebook
conversion pattern
pitch
conversion
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
JP8246443A
Other languages
Japanese (ja)
Inventor
Ko Amada
皇 天田
Masami Akamine
政巳 赤嶺
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Toshiba Corp
Original Assignee
Toshiba Corp
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Toshiba Corp filed Critical Toshiba Corp
Priority to JP8246443A priority Critical patent/JPH1091193A/en
Publication of JPH1091193A publication Critical patent/JPH1091193A/en
Pending legal-status Critical Current

Links

Landscapes

  • Analogue/Digital Conversion (AREA)
  • Compression, Expansion, Code Conversion, And Decoders (AREA)
  • Transmission Systems Not Characterized By The Medium Used For Transmission (AREA)

Abstract

(57)【要約】 【課題】計算量が少なくかつ高音質である音声符号化方
法を提供する。 【解決手段】適応符号帳110から適応符号ベクトルを
ピッチ励振信号として取り出して合成フィルタ102に
よりピッチ応答信号を生成し、変換パターン符号帳10
4から取り出された変換パターンでピッチ応答信号に変
換部105により変換を施して変換応答信号を生成す
る。ピッチ応答信号および変換応答信号をゲイン乗算器
106,107を介して加算器109で合成して合成音
声信号131を生成し、入力音声信号132に対する合
成音声信号131の歪が最小となる適応符号ベクトルお
よび変換パターンを符号帳101,104から探索し
て、合成フィルタ102の係数と符号帳101,104
から探索した適応符号ベクトルおよび変換パターンを示
すインデックスを符号化パラメータとして出力する。
(57) [Summary] [Problem] To provide a speech coding method with a small amount of calculation and high sound quality. An adaptive code vector is extracted from an adaptive code book as a pitch excitation signal, and a pitch response signal is generated by a synthesis filter.
The conversion unit 105 performs conversion on the pitch response signal using the conversion pattern extracted from Step 4 to generate a conversion response signal. An adder 109 synthesizes the pitch response signal and the converted response signal via gain multipliers 106 and 107 to generate a synthesized speech signal 131, and an adaptive code vector that minimizes distortion of the synthesized speech signal 131 with respect to the input speech signal 132 And a conversion pattern is searched from the codebooks 101 and 104, and the coefficients of the synthesis filter 102 and the codebooks 101 and 104 are searched.
And outputs an index indicating the adaptive code vector and the conversion pattern searched from, as encoding parameters.

Description

【発明の詳細な説明】DETAILED DESCRIPTION OF THE INVENTION

【0001】[0001]

【発明の属する技術分野】本発明は、ディジタル電話等
において音声信号を圧縮符号化するための音声符号化方
法および音声復号方法に関する。
BACKGROUND OF THE INVENTION 1. Field of the Invention The present invention relates to a voice coding method and a voice decoding method for compressing and coding a voice signal in a digital telephone or the like.

【0002】[0002]

【従来の技術】近年、電話帯域の音声を効率良く圧縮符
号化する方法として、CELP(CodeExcited Linear Pr
ediction)方式が良く用いられている。CELP方式に
関しては、M.R.Schroeder and B.S.Atal, “Code Exci
ted Linear Prediction (CELP): High Quality Speech
at Very Low Bit Rates,”Proc. ICASSP, pp.937-940,1
985(文献1)、および W.S.Kleijin, D.J.Krasinski e
t al.“Improved Speech Quality and Efficient Vecto
r Quantization in SELP,”Proc.ICASSP, pp155-158,19
88 (文献2)で詳しく述べられている。
2. Description of the Related Art In recent years, CELP (Code Excited Linear Pr
ediction) method is often used. Regarding the CELP method, see MRSchroeder and BSAtal, “Code Exci
ted Linear Prediction (CELP): High Quality Speech
at Very Low Bit Rates, ”Proc. ICASSP, pp.937-940,1
985 (Reference 1), and WSKleijin, DJKrasinskie
t al. “Improved Speech Quality and Efficient Vecto
r Quantization in SELP, ”Proc.ICASSP, pp155-158,19
88 (Ref. 2).

【0003】CELP方式における主要な符号化パラメ
ータはフィルタ係数、駆動信号、ゲインである。符号化
方式は、いわゆるAnalysis-by-synthesis 、すなわち分
析合成法に基づく方式であるため、エンコーダにデコー
ダがそのまま含まれている。エンコーダでは、前述の符
号化パラメータの値を変えながら、このデコーダを用い
て実際に復号を行い(ローカルデコード)、復号音声と
入力音声の歪みがより小さくなる符号化パラメータの組
合せを探して、これを伝送する。デコーダでは、送られ
てきた符号化パラメータの組合せで復号を行い、復号音
声を得る。
[0003] The main coding parameters in the CELP system are a filter coefficient, a drive signal, and a gain. Since the encoding method is a so-called analysis-by-synthesis, that is, a method based on an analysis-synthesis method, a decoder is directly included in an encoder. The encoder performs actual decoding by using this decoder while changing the value of the above-mentioned encoding parameter (local decoding), and searches for a combination of encoding parameters in which the distortion between the decoded speech and the input speech becomes smaller. Is transmitted. The decoder performs decoding with the combination of the received coding parameters to obtain a decoded speech.

【0004】典型的なCELP方式によるエンコーダを
図6に、デコーダを図7にそれぞれ示す。まず、図6に
示すエンコーダでは、フレーム単位に分割された音声が
入力端子1060に入力され、これを線形予測分析部1
050で分析することによって、聴感重み付き合成フィ
ルタ1030の係数と聴感重みフィルタ1040の係数
を決定する。次に、聴感重みフィルタ1040で重み付
けされた重み付き入力音声1041と復号音声1031
との誤差を誤差評価部1070で評価し、この誤差つま
り復号音声1031の歪みが最小になるように、適応符
号帳1011と雑音符号帳1012の符号ベクトルを選
択する。通常、これら二つの符号帳1011,1012
は逐次探索される。すなわち、まず最初に適応符号帳1
011の探索を行って適応符号ベクトルを決定し、次に
雑音符号帳1012の探索を行って雑音符号ベクトルを
決定する。最後に、ゲイン乗算器1021,1022で
適応符号ベクトル、雑音符号ベクトルにそれぞれ乗じる
べきゲインをゲイン符号帳1013から求める。ゲイン
乗算器1021,1022で適応符号ベクトル、雑音符
号ベクトルにゲインを乗じた後、加算器1023で足し
合わせて得られる駆動信号で聴感重み付き合成フィルタ
1030を駆動することにより、復号音声1031が得
られる。
FIG. 6 shows a typical encoder based on the CELP system, and FIG. 7 shows a decoder. First, in the encoder shown in FIG. 6, the audio divided in units of frames is input to an input terminal 1060, and this is input to a linear prediction analysis unit 1.
By analyzing at 050, the coefficients of the perceptual weighting synthesis filter 1030 and the perceptual weight filter 1040 are determined. Next, the weighted input voice 1041 and the decoded voice 1031 weighted by the perceptual weight filter 1040
Is evaluated by the error evaluator 1070, and code vectors of the adaptive codebook 1011 and the noise codebook 1012 are selected so that the error, that is, the distortion of the decoded speech 1031 is minimized. Usually, these two codebooks 1011 and 1012
Are searched sequentially. That is, first, adaptive codebook 1
011 to determine the adaptive code vector, and then search the random codebook 1012 to determine the random code vector. Finally, gains to be multiplied by the adaptive code vector and the noise code vector by the gain multipliers 1021 and 1022 are obtained from the gain codebook 1013. After multiplying the adaptive code vector and the noise code vector by the gains in the gain multipliers 1021 and 1022, the adder 1023 drives the perceptually weighted synthesis filter 1030 with a drive signal obtained by the addition, whereby the decoded speech 1031 is obtained. Can be

【0005】そして、線形予測分析部1050で生成さ
れLPC量子化部1080で量子化されたLPC係数
(聴感重み付き合成フィルタ1030の係数)と、適応
符号帳1011および雑音符号帳1012のインデック
ス(駆動信号)およびゲイン符号帳1013のインデッ
クス(ゲイン)が符号化パラメータとして出力される。
Then, the LPC coefficients (coefficients of the perceptually weighted synthesis filter 1030) generated by the linear prediction analysis unit 1050 and quantized by the LPC quantization unit 1080, and the indexes (drives) of the adaptive codebook 1011 and the noise codebook 1012 are used. The signal (signal) and the index (gain) of the gain codebook 1013 are output as coding parameters.

【0006】一方、デコーダは図7に示すように図6に
示したエンコーダの一部と同じ構造になっている。エン
コーダから送られてきた符号化パラメータを基に、適応
符号帳2011および雑音符号帳2012から符号ベク
トル、ゲイン符号帳2013からゲインをそれぞれ復号
し、さらにLPC逆量子化部2081で合成フィルタ2
035の係数を復号する。ゲイン乗算器2021,20
22でゲインを乗じた後、加算器2023で足し合わせ
て得られた駆動信号で合成フィルタ2035を駆動する
ことにより、復号音声が得られる。通常、聴覚的な品質
を向上させるために、復号音声はさらにポストフィルタ
2090を通して出力される。
On the other hand, the decoder has the same structure as a part of the encoder shown in FIG. 6, as shown in FIG. Based on the encoding parameters sent from the encoder, the adaptive codebook 2011 and the noise codebook 2012 decode the code vector and the gain codebook 2013 respectively the gain, and the LPC inverse quantization unit 2081 decodes the synthesis filter 2.
Decode the 035 coefficient. Gain multipliers 2021, 20
After multiplying the gain by 22, the adder 2023 drives the synthesis filter 2035 with a drive signal obtained by adding them, thereby obtaining a decoded voice. Normally, the decoded speech is further output through a post filter 2090 to improve the auditory quality.

【0007】CELP方式は高音質ではあるが、エンコ
ーダで多くの計算量を必要とする方式として知られてい
る。計算量の大半は、適応符号帳1011および雑音符
号帳1012からの符号ベクトルの探索に費やされてい
る。この探索では適応符号帳1011および雑音符号帳
1012に格納されている全ての符号ベクトルに対し
て、駆動信号を重み付き合成フィルタ1030を通した
後に目標ベクトルに対する歪みを計算するため、フィル
タリングに多くの演算量が必要となる。この演算量をい
かに削減するかが実現上のポイントであった。
Although the CELP system has high sound quality, it is known as a system requiring a large amount of calculation in an encoder. Most of the amount of calculation is spent searching for code vectors from adaptive codebook 1011 and noise codebook 1012. In this search, for all the code vectors stored in the adaptive codebook 1011 and the noise codebook 1012, the drive signal is passed through the weighted synthesis filter 1030, and then the distortion with respect to the target vector is calculated. The amount of calculation is required. The point in realizing was how to reduce the amount of calculation.

【0008】符号帳1011,1012のうち、特に適
応符号帳1011の探索は計算量の削減がしやすい。適
応符号帳1011は、ピッチ成分を表す符号ベクトルを
生成するため、入力信号や過去の復号音声のピッチ周期
を分析することで、おおよそどの符号ベクトルが適して
いるか見当を付けることが可能であるからである。こう
して見当を付けた符号ベクトルとその周辺を探すこと
で、ほぼ最適な符号ベクトルを探索することができる。
[0008] Of the codebooks 1011 and 1012, the search for the adaptive codebook 1011 is particularly easy to reduce the amount of calculation. Since the adaptive codebook 1011 generates a code vector representing a pitch component, it is possible to roughly determine which code vector is appropriate by analyzing a pitch period of an input signal or a past decoded speech. It is. By searching for the code vector thus provided and its surroundings, it is possible to search for an almost optimal code vector.

【0009】これに対し、雑音符号帳1012の探索は
計算量の削減がしにくい。雑音符号帳1012は、適応
符号ベクトルで表し切れなかった、目標ベクトルとのず
れを表す役割がある。したがって、格納されているベク
トルも相関の少ない雑音的な符号ベクトルであり、適応
符号帳1011の探索のように見当を付けて探すことが
難しい。
On the other hand, the search for the random codebook 1012 is difficult to reduce the amount of calculation. The noise codebook 1012 has a role of indicating a deviation from the target vector, which cannot be completely expressed by the adaptive code vector. Therefore, the stored vector is also a noise-like code vector with a small correlation, and it is difficult to search with an aim like a search of the adaptive codebook 1011.

【0010】そこで、従来より計算量が削減できるよう
に符号帳の構造に手を加えた、いわゆる構造化符号帳が
広く用いられてきている。VSELP,ACELPなど
は、この構造化符号帳を用いて符号化を行う典型的な例
である。しかし、構造化符号帳を用いることは、計算量
の削減には有効であるが、格納する符号ベクトルに制約
を課していることになるため、構造化しない場合に比べ
て音質が劣化する問題があった。
Therefore, a so-called structured codebook in which the structure of the codebook is modified so that the amount of calculation can be reduced has been widely used. VSELP, ACELP, and the like are typical examples of performing encoding using the structured codebook. However, although the use of a structured codebook is effective in reducing the amount of calculation, it imposes restrictions on the code vectors to be stored, so that the sound quality is deteriorated compared to the case where the structure is not structured. was there.

【0011】[0011]

【発明が解決しようとする課題】上述したように、CE
LP方式による従来の音声符号化方法では、特に雑音符
号帳からの符号ベクトルの探索に多くの計算を必要とす
るという問題があり、また計算量を削減するため符号帳
の構造化を行うと、音質が劣化するという問題点があっ
た。本発明は、計算量が少なくかつ高音質である音声符
号化方法および音声復号方法を提供することを目的とす
る。
As described above, the CE
In the conventional speech coding method based on the LP method, there is a problem that a large number of calculations are required particularly for searching for a code vector from a noise codebook, and when the codebook is structured to reduce the calculation amount, There was a problem that the sound quality deteriorated. An object of the present invention is to provide a speech encoding method and a speech decoding method which require a small amount of calculation and have high sound quality.

【0012】[0012]

【課題を解決するための手段】上記の課題を解決するた
め、本発明の音声符号化方法では、入力音声信号の分析
結果に基づいて係数が決定される合成フィルタを駆動す
るための過去の駆動信号に基づいて生成される適応符号
ベクトルを格納した適応符号帳から適応符号ベクトルを
取り出してピッチ励振信号とし、このピッチ励振信号を
合成フィルタに通して第1の応答信号を生成する。一
方、複数の変換パターンを格納した変換パターン符号帳
を用意し、この符号帳から取り出される変換パターンで
第1の応答信号に変換を施して第2の応答信号を生成す
る。そして、これら第1および第2の応答信号を合成し
て合成音声信号を生成し、入力音声信号に対する合成音
声信号の歪がより小さくなる適応符号ベクトルおよび変
換パターンを適応符号帳および変換パターン符号帳から
それぞれ探索して、少なくとも合成フィルタの係数の情
報と適応符号帳および変換パターン符号帳から探索した
適応符号ベクトルおよび変換パターンを示す情報を符号
化パラメータとして出力する。
In order to solve the above-mentioned problems, a speech encoding method according to the present invention employs a past drive for driving a synthesis filter whose coefficients are determined based on an analysis result of an input speech signal. An adaptive code vector is extracted from an adaptive codebook that stores an adaptive code vector generated based on the signal and used as a pitch excitation signal, and the pitch excitation signal is passed through a synthesis filter to generate a first response signal. On the other hand, a conversion pattern codebook that stores a plurality of conversion patterns is prepared, and the first response signal is converted with the conversion pattern extracted from the codebook to generate a second response signal. Then, the first and second response signals are combined to generate a synthesized speech signal, and the adaptive code vector and the conversion pattern that reduce the distortion of the synthesized speech signal with respect to the input speech signal are converted into an adaptive codebook and a conversion pattern codebook. , And outputs at least information on the coefficients of the synthesis filter and information indicating the adaptive code vector and the conversion pattern searched from the adaptive codebook and the conversion pattern codebook as coding parameters.

【0013】この音声符号化方法では、従来方法の畳み
込まれた雑音符号ベクトルに相当する第2の応答信号を
求める際に、第1の応答信号に対して変換パターン符号
帳に格納された変換パターンに従って様々な変換を適用
し、入力音声信号に対する合成音声信号の歪みが最小に
なる変換パターンを探索して、この変換パターンを示す
インデックスを出力する。
In this speech coding method, when a second response signal corresponding to a convolved noise code vector according to the conventional method is obtained, the conversion stored in the conversion pattern codebook with respect to the first response signal is performed. Various conversions are applied according to the pattern, a conversion pattern that minimizes the distortion of the synthesized voice signal with respect to the input voice signal is searched, and an index indicating the conversion pattern is output.

【0014】このようにすることで、従来の雑音符号帳
探索で必要であった雑音符号ベクトルのフィルタリング
演算が不要となる。最適な変換を探索するために増加す
る計算量はフィルタリング演算に比べ僅かなので、全体
として第2の応答信号を従来よりも大幅に少ない計算量
で求めることが可能になる。
This eliminates the need for a noise code vector filtering operation required in the conventional noise codebook search. Since the amount of calculation that is increased to search for the optimal conversion is small compared to the filtering operation, it is possible to obtain the second response signal as a whole with a significantly smaller amount of calculation than in the past.

【0015】ここで、変換パターン符号帳に格納される
変換パターンとしては、例えば行列演算が用いられる。
また、最適な変換パターンを探索するための計算を行う
際、この行列の多くの要素を零とし、一部が非零である
ように、例えば各行に非零の要素が5個以下となるよう
に構成する。このようにすると、計算量削減の効果がさ
らに大きくなる。
Here, as the conversion pattern stored in the conversion pattern codebook, for example, a matrix operation is used.
In addition, when performing a calculation for searching for an optimal conversion pattern, many elements of this matrix are set to zero and some of them are non-zero, for example, each row has five or less non-zero elements. To be configured. In this way, the effect of reducing the amount of calculation is further increased.

【0016】一方、この音声符号化方法に対応する本発
明の音声復号方法では、符号化側からの少なくとも合成
フィルタのフィルタ係数と適応符号ベクトルおよび変換
パターンを示すインデックスを符号化パラメータとして
入力し、適応符号ベクトルを格納した適応符号帳から符
号化パラメータに従って取り出される適応符号ベクトル
をピッチ励振信号として、このピッチ励振信号を前記符
号化パラメータに従って係数が決定される合成フィルタ
に通して第1の応答信号を生成する。また、複数の変換
パターンを格納した変換パターン符号帳から符号化パラ
メータに従って取り出される変換パターンで第1の応答
信号に変換を施して、第2の応答信号を生成する。そし
て、これら第1および第2の応答信号を合成して復号音
声信号を生成する。
On the other hand, in the speech decoding method of the present invention corresponding to this speech encoding method, at least a filter coefficient of a synthesis filter, an adaptive code vector, and an index indicating a conversion pattern from a coding side are input as coding parameters. An adaptive code vector extracted from an adaptive codebook storing an adaptive code vector according to a coding parameter is used as a pitch excitation signal, and the pitch excitation signal is passed through a synthesis filter whose coefficient is determined according to the coding parameter to obtain a first response signal. Generate The second response signal is generated by converting the first response signal using a conversion pattern extracted from a conversion pattern codebook storing a plurality of conversion patterns in accordance with an encoding parameter. Then, the first and second response signals are combined to generate a decoded audio signal.

【0017】本発明の音声合成方法は、合成フィルタが
ない構成でもよく、その場合は適応符号帳から適応符号
ベクトルを取り出して第1のピッチ信号とし、変換パタ
ーン符号帳から取り出された変換パターンで第1のピッ
チ信号に変換を施して第2のピッチ信号を生成し、これ
ら第1および第2のピッチ信号を合成して合成音声信号
を生成する。そして、入力音声信号に対する合成音声信
号の歪がより小さくなる適応符号ベクトルおよび変換パ
ターンを適応符号帳および変換パターン符号帳からそれ
ぞれ探索し、少なくとも適応符号帳および変換パターン
符号帳から探索した適応符号ベクトルおよび変換パター
ンを示すインデックスを符号化パラメータとして出力す
る。
The speech synthesizing method of the present invention may have a configuration without a synthesis filter. In this case, an adaptive code vector is extracted from the adaptive codebook and used as a first pitch signal, and the conversion pattern extracted from the conversion pattern codebook is used. The first pitch signal is converted to generate a second pitch signal, and the first and second pitch signals are synthesized to generate a synthesized speech signal. Then, an adaptive code vector and a conversion pattern in which the distortion of the synthesized voice signal with respect to the input voice signal is smaller are searched from the adaptive code book and the conversion pattern code book, respectively, and at least the adaptive code vector searched from the adaptive code book and the conversion pattern code book. And an index indicating the conversion pattern is output as an encoding parameter.

【0018】この音声符号化方法に対応する本発明の音
声復号方法では、符号化側からの少なくとも適応符号ベ
クトルおよび変換パターンを示すインデックスを符号化
パラメータとして入力し、適応符号帳から符号化パラメ
ータに従って取り出される適応符号ベクトルを第1のピ
ッチ信号とし、変換パターン符号帳から符号化パラメー
タに従って取り出される変換パターンで第1のピッチ信
号に変換を施して第2のピッチ信号を生成する。そし
て、これら第1および第2のピッチ信号を合成して復号
音声信号を生成する。
According to the speech decoding method of the present invention corresponding to this speech encoding method, at least an index indicating an adaptive code vector and a conversion pattern from the encoding side is input as a coding parameter, and is input from the adaptive codebook according to the coding parameter. The extracted adaptive code vector is used as a first pitch signal, and the second pitch signal is generated by converting the first pitch signal with a conversion pattern extracted from the conversion pattern codebook in accordance with an encoding parameter. Then, the first and second pitch signals are combined to generate a decoded audio signal.

【0019】[0019]

【発明の実施の形態】BEST MODE FOR CARRYING OUT THE INVENTION

(第1の実施形態)図1に、第1の実施形態に係る音声
符号化方法を適用した音声符号化装置の構成を示す。こ
の音声符号化装置は、適応符号帳101、合成フィルタ
102、この合成フィルタ102の逆フィルタである分
析フィルタ103、変換パターン符号帳104、変換部
105、ゲイン乗算器106,107、利得符号帳10
8、加算器109、線形予測分析部110、入力音声信
号132の入力端子111、減算器112、聴感重みフ
ィルタ113および評価部114から構成される。
(First Embodiment) FIG. 1 shows a configuration of a speech coding apparatus to which a speech coding method according to a first embodiment is applied. The speech coding apparatus includes an adaptive codebook 101, a synthesis filter 102, an analysis filter 103 which is an inverse filter of the synthesis filter 102, a conversion pattern codebook 104, a conversion unit 105, gain multipliers 106 and 107, and a gain codebook 10.
8, an adder 109, a linear prediction analysis unit 110, an input terminal 111 for an input audio signal 132, a subtractor 112, an audibility weighting filter 113, and an evaluation unit 114.

【0020】適応符号帳101は、合成フィルタ102
を駆動する過去の駆動信号から作られる複数の適応符号
ベクトルを格納している。ここで、駆動信号は入力音声
信号132を分析する線形予測分析部110によりフィ
ルタ係数が決定される分析フィルタ103に合成音声信
号131を通すことによって得られる。合成フィルタ1
02は、同様に線形予測分析部110によりフィルタ係
数が決定され、適応符号帳101から取り出される適応
符号ベクトルをピッチ励振信号として入力して、ピッチ
応答信号(第1の応答信号)を出力する。
The adaptive codebook 101 includes a synthesis filter 102
And a plurality of adaptive code vectors generated from the past drive signals for driving. Here, the drive signal is obtained by passing the synthesized voice signal 131 through the analysis filter 103 whose filter coefficients are determined by the linear prediction analysis unit 110 that analyzes the input voice signal 132. Synthesis filter 1
In step 02, a filter coefficient is determined by the linear prediction analysis unit 110, an adaptive code vector extracted from the adaptive codebook 101 is input as a pitch excitation signal, and a pitch response signal (first response signal) is output.

【0021】変換パターン符号帳104は、合成フィル
タ102からのピッチ応答信号に施すべき変換を表す複
数の変換パターンを格納しており、変換部105は変換
パターン符号帳104から取り出された変換パターンで
ピッチ応答信号に変換を施して、変換応答信号(第2の
応答信号)を出力する。
The conversion pattern codebook 104 stores a plurality of conversion patterns representing conversions to be performed on the pitch response signal from the synthesis filter 102. The conversion unit 105 uses the conversion pattern extracted from the conversion pattern codebook 104. It converts the pitch response signal and outputs a converted response signal (second response signal).

【0022】合成フィルタ102から出力されるピッチ
応答信号および変換部105から出力される変換応答信
号は、それぞれゲイン乗算器106,107により利得
符号帳108から与えられるゲインが乗じられた後、加
算器109で足し合わせられ、合成音声信号131が生
成される。減算器112は、この合成音声信号131と
入力音声信号132との差信号を出力する。聴感重みフ
ィルタ113は、この差信号に聴感重み付けを行う。
The pitch response signal output from the synthesis filter 102 and the conversion response signal output from the conversion section 105 are multiplied by gains given from a gain codebook 108 by gain multipliers 106 and 107, respectively, and then added by an adder. The sum is added at 109 to generate a synthesized speech signal 131. The subtractor 112 outputs a difference signal between the synthesized voice signal 131 and the input voice signal 132. The audibility weighting filter 113 weights the difference signal with audibility.

【0023】評価部114は、聴感重みフィルタ113
によって聴感重み付けられた差信号の評価を行い、この
差信号のパワが最小になるように、すなわち、入力音声
信号132に対する合成音声信号131の聴感重み付き
の歪が最小となるように、適応符号帳101、変換パタ
ーン符号帳104および利得符号帳108の探索を行
う。
The evaluation unit 114 includes an audibility weighting filter 113
The difference between the perceptually weighted difference signals is evaluated, and the adaptive code is set so that the power of the difference signal is minimized, that is, the distortion with the perceptual weight of the synthesized speech signal 131 with respect to the input speech signal 132 is minimized. The book 101, the conversion pattern codebook 104, and the gain codebook 108 are searched.

【0024】評価部114による探索の結果、利得符号
帳108のインデックス121、変換パターン符号帳1
04のインデックス122、合成フィルタ102のイン
デックス123および適応符号帳101のインデックス
124が得られ、これらのインデックスが符号化パラメ
ータとして出力される。
As a result of the search by the evaluation unit 114, the index 121 of the gain codebook 108, the conversion pattern codebook 1
An index 122, an index 123 of the synthesis filter 102, and an index 124 of the adaptive codebook 101 are obtained, and these indexes are output as coding parameters.

【0025】次に、本実施形態における音声符号化処理
の手順を図2に示すフローチャートを用いて説明する。 [ステップS1]まず、線形予測分析部110におい
て、入力端子111に所定フレーム長で入力される入力
音声信号132を線形予測分析し、合成フィルタ10
2、分析フィルタ103および聴感重みフィルタ113
の係数を求める。合成フィルタ102の係数は、通常ベ
クトル量子化され、合成フィルタ102のインデックス
123として出力される。
Next, the procedure of the speech encoding process in the present embodiment will be described with reference to the flowchart shown in FIG. [Step S1] First, the linear prediction analysis unit 110 performs linear prediction analysis on an input audio signal 132 input to the input terminal 111 with a predetermined frame length, and
2. Analysis filter 103 and auditory weight filter 113
Find the coefficient of. The coefficients of the synthesis filter 102 are usually vector-quantized and output as an index 123 of the synthesis filter 102.

【0026】[ステップS2]次に、適応符号帳101
の探索を行う。すなわち、聴感重みフィルタ113で重
み付けした差信号が最小となるピッチ励振信号が出力さ
れるように、適応符号帳101から一つの適応符号ベク
トルを探索する。適応符号帳101からピッチ励振信号
の候補となる適応符号ベクトルを探索する方法は、当該
技術分野において周知であり、本実施形態においてもそ
の方法を用いることができる。この探索結果は、適応符
号帳101のインデックスとして出力される。また、こ
うして求められた最適なピッチ励振信号を合成フィルタ
102に通して、ピッチ応答信号を求めておく。
[Step S2] Next, the adaptive codebook 101
Search for. That is, one adaptive code vector is searched from the adaptive codebook 101 so that the pitch excitation signal that minimizes the difference signal weighted by the perceptual weight filter 113 is output. A method of searching the adaptive codebook 101 for an adaptive code vector that is a candidate for a pitch excitation signal is well known in the art, and the method can be used in the present embodiment. This search result is output as an index of adaptive codebook 101. Further, the optimum pitch excitation signal thus obtained is passed through the synthesis filter 102 to obtain a pitch response signal.

【0027】[ステップS3]次に、変換部105にお
いて、変換パターン符号帳104に格納されている変換
パターンで示される変換をピッチ応答信号に施して変換
応答信号を作り、ステップS2で求められたピッチ応答
信号と合わせた時に聴感重み付きの差信号のパワが最小
になる変換パターンを探索し、探索結果を変換パターン
符号帳104のインデックス122として出力する。
[Step S3] Next, the conversion section 105 applies a conversion indicated by the conversion pattern stored in the conversion pattern codebook 104 to the pitch response signal to generate a conversion response signal, which is obtained in step S2. A search is made for a conversion pattern that minimizes the power of the perceptually weighted difference signal when combined with the pitch response signal, and the search result is output as an index 122 of the conversion pattern codebook 104.

【0028】[ステップS4]次に、目標ベクトルとの
聴感重み付き誤差が最小になるように、ゲイン乗算器1
06,107で乗じるゲインを利得符号帳108から探
索し、その探索結果を利得符号帳108のインデックス
121として出力する。
[Step S4] Next, the gain multiplier 1 is set so that the perceptual weighted error from the target vector is minimized.
The gain multiplied by 06 and 107 is searched from gain codebook 108, and the search result is output as index 121 of gain codebook 108.

【0029】[ステップS5]最後に、入力音声信号1
32の次のフレームの処理に備えるため、合成音声信号
131を分析フィルタ103に入力して残差信号を生成
し、この残差信号を用いて適応符号帳101の内容を更
新する。
[Step S5] Finally, the input audio signal 1
In order to prepare for the processing of the frame next to 32, the synthesized speech signal 131 is input to the analysis filter 103 to generate a residual signal, and the content of the adaptive codebook 101 is updated using the residual signal.

【0030】次に、本実施形態による効果について述べ
る。従来の雑音符号帳探索では、符号帳に格納されてい
る符号ベクトル全てに対してフィルタによる畳み込み演
算を行う必要があり、この畳み込み演算が計算量増加の
主要な原因になっていた。
Next, effects of the present embodiment will be described. In the conventional noise codebook search, it is necessary to perform a convolution operation using a filter on all the code vectors stored in the codebook, and this convolution operation has been a major cause of an increase in the amount of calculation.

【0031】これに対し、本実施形態ではステップS2
の適応符号帳101の探索終了時にはピッチ励振信号は
確定しており、これを合成フィルタ102で畳み込んだ
ピッチ応答信号を活用することにより、従来の雑音符号
帳探索に相当する計算量を削減できる。すなわち、ステ
ップS3に示したようにピッチ応答信号に適当な変換を
施すことで、従来の雑音符号ベクトルの応答に相当する
変換応答信号を作り出すものである。このようにするこ
とで、従来の雑音符号帳探索での畳み込み演算を不要に
し、計算量の大幅な削減が可能になる。
On the other hand, in the present embodiment, step S2
At the end of the search of the adaptive codebook 101, the pitch excitation signal is determined. By using the pitch response signal obtained by convolving the signal with the synthesis filter 102, the amount of calculation corresponding to the conventional noise codebook search can be reduced. . That is, by performing appropriate conversion on the pitch response signal as shown in step S3, a conversion response signal corresponding to the response of the conventional noise code vector is created. By doing so, the convolution operation in the conventional random codebook search becomes unnecessary, and the amount of calculation can be greatly reduced.

【0032】次に、本発明の特徴をなす変換パターン符
号帳104および変換部105について具体的に説明す
る。変換部105によりピッチ応答信号に対して施され
る変換の変換パターン、すなわち変換パターン符号帳1
04に格納されている変換パターンは、例えば行列演算
で表される。合成フィルタ102から出力されるピッチ
応答信号を表すべクトルをp、変換パターンを表す行列
をAi(i=1,…,M、ただしMは変換パターンの
数)とした場合、変換部105から出力される変換応答
信号を表すベクトルxiは xi=Ai p (1) と表すことができる。
Next, the conversion pattern codebook 104 and the conversion unit 105 which characterize the present invention will be described in detail. Conversion pattern of conversion applied to pitch response signal by conversion section 105, that is, conversion pattern codebook 1
The conversion pattern stored in 04 is represented by, for example, a matrix operation. If the vector representing the pitch response signal output from the synthesis filter 102 is p and the matrix representing the conversion pattern is Ai (i = 1,..., M, where M is the number of conversion patterns), the output from the conversion unit 105 The vector xi representing the converted response signal to be expressed can be expressed as xi = Aip (1).

【0033】評価部114により変換部パターン符号帳
104の探索を行う際には、全てのAi(i=1,…,
M)に対しxiを計算して、目標べクトルに対する歪み
を最小にするxiを求め、その時のAiを最適な変換パ
ターンとして、その変換パターンを表す変換パターン符
号帳104のインデックスiを出力する。
When the evaluation unit 114 searches the conversion unit pattern codebook 104, all Ai (i = 1,...,
Xi is calculated for M) to obtain xi that minimizes distortion with respect to the target vector, and Ai at that time is set as an optimum conversion pattern, and an index i of the conversion pattern codebook 104 representing the conversion pattern is output.

【0034】ベクトルxiとpは通常、フレーム(また
はサブフレーム)長の次元のベクトルであり、この次元
をNとすると、AiはN*Nの行列になる。従って、こ
の要素が全て非零の要素であると、xiを求める計算量
が大きくなり、従来の雑音符号帳を用い方式に比べた場
合の計算量削減のメリットが低下する。しかし、非零の
要素数を制限する、例えばどの行も非零の要素は5個以
下に制限することによって、変換に必要な計算量を畳み
込み演算の場合と比べて大幅に削減することができる。
The vectors xi and p are usually vectors of the dimension of the frame (or sub-frame) length. If this dimension is N, Ai becomes an N * N matrix. Therefore, if all of these elements are non-zero elements, the amount of calculation for obtaining xi increases, and the merit of reducing the amount of calculation as compared with the conventional method using a random codebook decreases. However, by limiting the number of non-zero elements, for example, by limiting the number of non-zero elements in each row to 5 or less, the amount of computation required for conversion can be significantly reduced as compared to the case of convolution operation. .

【0035】このことは、変換部105から出力される
変換応答信号のバリエーションを制限しているようにも
見えるが、もともとN次元ベクトルの変換にN*Nの冗
長な行列を用いているため、制限していることにはなら
ない。実際に、非零の要素がN個あれば、任意の応答ベ
クトルを生成することが可能である。各行に1つだけ非
零の要素を持ち、その他を零とした場合、xiを求める
変換をN回の演算で行うことができ、従来法に比べ計算
量を大幅に削減することができる。
Although this seems to limit the variation of the conversion response signal output from the conversion unit 105, since an N * N redundant matrix is originally used for the conversion of the N-dimensional vector, It is not a restriction. In fact, if there are N non-zero elements, it is possible to generate an arbitrary response vector. When each row has only one non-zero element and the others are zero, the conversion for obtaining xi can be performed by N operations, and the amount of calculation can be greatly reduced as compared with the conventional method.

【0036】数式を用いて説明すると、本実施形態にお
ける変換パターン符号帳104の探索は、入力音声信号
132から作られる目標ベクトルをrとした場合、次式
で示される評価値 E=<r,xi>2 /|xi|2 (2) を最大にするxiを探索する。この評価値Eの値を一回
計算するのに必要な演算量は、(1)式の計算でN回、
(2)式の分子でN+1回、分母でN回なので、約3N
回である。
Explaining using mathematical expressions, the search of the conversion pattern codebook 104 in the present embodiment is performed when the target vector generated from the input speech signal 132 is r, and the evaluation value E = <r, Search for xi that maximizes xi> 2 // xi | 2 (2). The amount of operation required to calculate the value of the evaluation value E once is N times in the calculation of the expression (1),
Since N + 1 times in the numerator of the equation (2) and N times in the denominator, about 3N
Times.

【0037】これに対し、従来の雑音符号帳探索は合成
フィルタによる畳み込みを行列H、雑音符号ベクトルを
ciで表すと、 E=<r,Hci>2 /|Hci|2 (3) を計算する必要がある。この方法では、Hciの畳み込
み演算にN(N+1)/2回の演算が必要であり、分母
と分子の内積でそれぞれN回の演算が必要になるため、
合計2N+N(N+1)/2回の計算量が必要になる。
On the other hand, in the conventional noise codebook search, when convolution by a synthesis filter is represented by a matrix H and a noise code vector is represented by ci, E = <r, Hci> 2 / | Hci | 2 (3) There is a need. In this method, the convolution operation of Hci requires N (N + 1) / 2 operations, and the inner product of the denominator and the numerator requires N operations, respectively.
A total of 2N + N (N + 1) / 2 calculations are required.

【0038】今、サブフレーム長N=40(5mse
c)で8ビットの符号帳を用いたと仮定すると、雑音符
号帳を用いる従来方式は(2*40+40*41/2)
*256/5msec=46.1MOPSとなるのに対
し、本実施形態では3*40*256/5msec=
6.1MOPSで変換パターン符号帳104の探索を行
うことができ、探索に必要な計算量は従来方式の1/7
以下に低減される。
Now, the subframe length N = 40 (5 mse
Assuming that an 8-bit codebook is used in c), the conventional method using the noise codebook is (2 * 40 + 40 * 41/2)
* 256/5 msec = 46.1 MOPS, whereas in the present embodiment, 3 * 40 * 256/5 msec =
The conversion pattern codebook 104 can be searched at 6.1 MOPS, and the amount of calculation required for the search is 1/7 that of the conventional method.
It is reduced below.

【0039】(2)(3)式の評価式は、探索の方法に
よって変わる。例えば直交化探索法を用いた場合は、x
iの代わりにこれをピッチ応答ベクトルに対し直交化し
たベクトルxi′を用いればよい。
The evaluation expressions (2) and (3) vary depending on the search method. For example, when the orthogonal search method is used, x
Instead of i, a vector xi 'obtained by making this orthogonal to the pitch response vector may be used.

【0040】さらに、本実施形態におけるピッチ応答信
号を変換して得られた変換応答信号は、入力音声信号1
32の性質に合った応答信号となる利点もある。ピッチ
励振信号となる適応符号帳101に格納された適応符号
ベクトルは、過去の駆動信号から作られるため、既に符
号化した入力音声信号の性質を反映しているからであ
る。従って、これを変形して得られる変換応答信号には
その性質が残るため、入力音声信号に合った応答信号が
得られ、音質の向上につながるのである。
Further, the converted response signal obtained by converting the pitch response signal in the present embodiment is the input voice signal 1
There is also an advantage that the response signal conforms to the characteristics of the T.32. This is because the adaptive code vector stored in the adaptive codebook 101 serving as the pitch excitation signal is created from the past drive signal, and thus reflects the properties of the already encoded input speech signal. Therefore, since the converted response signal obtained by transforming it has its properties, a response signal suitable for the input audio signal is obtained, which leads to improvement in sound quality.

【0041】なお、従来の雑音符号帳探索では、複数の
符号帳を備えておき、入力音声信号に応じて切替えて用
いる方法や、入力音声信号のピッチ周期で符号ベクトル
を周期化する方法など、入力音声信号に適応化しようと
する試みはあるが、本実施形態のように自動的に入力音
声信号に合った変換応答信号を生成することは極めて困
難であった。
In the conventional noise codebook search, a method of providing a plurality of codebooks and switching them according to the input speech signal, a method of making the code vector periodic with the pitch cycle of the input speech signal, and the like are available. Although attempts have been made to adapt to an input audio signal, it has been extremely difficult to automatically generate a conversion response signal that matches the input audio signal as in the present embodiment.

【0042】CELP方式では通常、計算量を削減する
ため、合成音声信号と入力音声信号との差をとる前に、
それぞれの信号に聴感重みフィルタによるフィルタリン
グを施しておくことが多い。本実施形態に関しても、こ
のような工夫を行うことは可能である。図1のブロック
図は、探索の仕組みを理解しやすくするために理論に即
した表記をしたものであり、聴感重みフィルタ113の
位置を限定するものではない。
In the CELP system, usually, in order to reduce the amount of calculation, before taking the difference between the synthesized speech signal and the input speech signal,
In many cases, each signal is subjected to filtering by an audibility weighting filter. Also in the present embodiment, it is possible to make such a contrivance. The block diagram in FIG. 1 is based on the theory in order to facilitate understanding of the search mechanism, and does not limit the position of the audibility weighting filter 113.

【0043】ステップS4のゲイン探索法も、他に様々
な方法がある。例えば、ステップS2でピッチ励振信号
が確定した直後に、ゲイン乗算器107で乗じるゲイン
の値を確定し、その後にステップS3の処理を行い、ス
テップS4ではゲイン乗算器106で乗じるゲインの値
を決めるといった具合に、逐次ゲインを確定してゆく方
法をとってもよい。すなわち、ステップS4でのゲイン
の確定方法は一例であり、これに限定されるものではな
い。
There are various other methods for the gain search in step S4. For example, immediately after the pitch excitation signal is determined in step S2, the value of the gain to be multiplied by the gain multiplier 107 is determined, and then the process of step S3 is performed. In step S4, the value of the gain to be multiplied by the gain multiplier 106 is determined. For example, a method of determining the gain sequentially may be used. That is, the method of determining the gain in step S4 is an example, and the method is not limited to this.

【0044】次に、図3を用いて本実施形態に係る音声
復号装置の構成を説明する。本発明は分析合成法に基づ
く音声符号化方法であるため、復号装置は符号化装置に
組み込まれている復号装置(ローカルデコーダ)と同様
である。すなわち、この音声復号装置は適応符号帳20
1、合成フィルタ202、合成フィルタ202の逆フィ
ルタ203、変換パターン符号帳204、変換部20
5、ゲイン乗算器206,207、利得符号帳208、
加算器209、逆量子化部210、符号化パラメータの
入力端子221,222,223および224により構
成される。
Next, the configuration of the speech decoding apparatus according to this embodiment will be described with reference to FIG. Since the present invention is a speech coding method based on the analysis and synthesis method, the decoding device is the same as a decoding device (local decoder) incorporated in the coding device. That is, this speech decoding apparatus is adapted to the adaptive codebook 20.
1, synthesis filter 202, inverse filter 203 of synthesis filter 202, conversion pattern codebook 204, conversion unit 20
5, gain multipliers 206 and 207, gain codebook 208,
It comprises an adder 209, an inverse quantization unit 210, and input terminals 221, 222, 223 and 224 for coding parameters.

【0045】図3において、入力端子221,222,
223および224には、符号化パラメータとして、図
1の音声符号化装置から出力された利得符号帳108の
インデックス121、変換パターン符号帳104のイン
デックス122、合成フィルタ102のインデックス1
23および適応符号帳101のインデックス124がそ
れぞれ入力される。
In FIG. 3, input terminals 221, 222,
223 and 224 include, as coding parameters, an index 121 of the gain codebook 108, an index 122 of the conversion pattern codebook 104, and an index 1 of the synthesis filter 102 output from the speech coding apparatus of FIG.
23 and the index 124 of the adaptive codebook 101 are input.

【0046】インデックス121は利得符号帳208に
与えられ、これに基づき利得符号帳208からゲイン乗
算器206,207にゲインの値が読み出される。イン
デックス122は変換パターン符号帳204に与えら
れ、これに基づき変換パターンが変換部205に与えら
れる。インデックス123は逆量子化部210を介して
合成フィルタ202および逆フィルタ203に与えら
れ、これらのフィルタ202,203の係数が決定され
る。インデックス124は適応符号帳201に与えら
れ、合成フィルタ202に入力されるピッチ励振信号と
なる適応符号ベクトルが選択される。
The index 121 is given to the gain codebook 208, and the gain value is read out from the gain codebook 208 to the gain multipliers 206 and 207 based on the index 121. The index 122 is provided to the conversion pattern codebook 204, and a conversion pattern is provided to the conversion unit 205 based on the index. The index 123 is provided to the synthesis filter 202 and the inverse filter 203 via the inverse quantization unit 210, and the coefficients of these filters 202 and 203 are determined. The index 124 is given to the adaptive codebook 201, and an adaptive code vector serving as a pitch excitation signal input to the synthesis filter 202 is selected.

【0047】この結果、合成フィルタ202から出力さ
れるピッチ応答信号と変換部205から出力される変換
応答信号がゲイン乗算器206,207でゲインを乗じ
られた後、加算器209で加算されることによって、復
号音声信号211が生成される。この復号音声信号21
1は、必要に応じて図示しないポストフィルタで聴覚的
に品質が向上するように処理されることがある。
As a result, the pitch response signal output from the synthesis filter 202 and the conversion response signal output from the conversion unit 205 are multiplied by gains in the gain multipliers 206 and 207 and then added by the adder 209. As a result, a decoded audio signal 211 is generated. This decoded audio signal 21
1 may be processed by a post-filter (not shown) so that the quality is improved audibly as needed.

【0048】(第2の実施形態)図4に、第2の実施形
態に係る音声符号化方法を適用した音声符号化装置を示
す。本実施形態は、第1の実施形態と比較して聴感重み
フィルタの位置が異なっている。
(Second Embodiment) FIG. 4 shows a speech coding apparatus to which a speech coding method according to a second embodiment is applied. This embodiment is different from the first embodiment in the position of the audibility weighting filter.

【0049】第1の実施形態の説明において聴感重みフ
ィルタの位置はどこでも良いと述べたが、本実施形態の
構成をとった場合、第1の実施形態と等価にはならな
い。なぜなら、本実施形態では合成フィルタ102から
出力されるピッチ応答信号を聴感重みフィルタ141に
通した後に変換を行うのに対し、第1の実施形態ではピ
ッチ応答信号を変換した後に聴感重みフィルタ113に
通すからである。一般には、変換とフィルタリングの順
番は入れ換えられないため、等価ではない。等価である
か否かは本発明の効果とは関係がなく、どちらの構成で
も変換パターン符号帳104をその構成に合うように設
計すれば良い。なお、本実施形態では合成フィルタ10
2の出力側に聴感重みフィルタ141を設けたことに伴
い、逆フィルタ103の入力側に聴感重み逆フィルタ1
40を設け、さらに入力音声信号も聴感重みフィルタ1
42を通して減算器112に入力している。
In the description of the first embodiment, the position of the audibility weighting filter may be anywhere. However, when the configuration of the present embodiment is adopted, it is not equivalent to the first embodiment. This is because in the present embodiment, the pitch response signal output from the synthesis filter 102 is passed through the perceptual weight filter 141 and then converted, whereas in the first embodiment the pitch response signal is converted and then transmitted to the perceptual weight filter 113. Because it passes. In general, the order of conversion and filtering is not equivalent because they are not interchangeable. Whether they are equivalent or not is not related to the effect of the present invention, and the conversion pattern codebook 104 may be designed so as to conform to either configuration. In the present embodiment, the synthesis filter 10
2 is provided with the audibility weight filter 141 on the output side, and the audibility weight inverse filter 1 is provided on the input side of the inverse filter 103.
40, and the input audio signal is also applied to the audibility weighting filter 1
It is input to the subtractor 112 through.

【0050】(第3の実施形態)第1の実施形態では計
算量は削減できるが、変換を表す行列Aiのためのメモ
リ量が比較的大きくなる。つまり、各行に非零が1要素
だけとしても256候補用意するには、40*256=
10kワードのメモリが必要になる。
(Third Embodiment) In the first embodiment, the amount of calculation can be reduced, but the amount of memory for the matrix Ai representing the transformation becomes relatively large. That is, to prepare 256 candidates even if each row has only one non-zero element, 40 * 256 =
10k words of memory are required.

【0051】そこで、本実施形態では、変換パターン符
号帳104をオーバラップ符号帳、特に非零の要素をオ
ーバラップ構造とすることと、各Aiは対角方向に非零
の要素を持つようにすることで、メモリ量を大幅に削減
している。
Therefore, in the present embodiment, the conversion pattern codebook 104 has an overlap codebook, particularly non-zero elements have an overlap structure, and each Ai has a nonzero element in a diagonal direction. By doing so, the amount of memory has been significantly reduced.

【0052】図5に、その様子を摸式的に示す。対角成
分を格納した−本のオーバラップ符号帳のi番目の位置
からサブフレーム長のべクトルを切り出し、Aiの対角
要素とする。従って、AiとAi+1はN−1個の共通
した要素、すなわち重複した成分を持つことになる。こ
のときのメモリ量は、N+M−1(Mは変換パターンの
数)となる。これは第1の実施形態で説明したと同様
に、N=40,M=256では、0.295kワードと
なり、より低メモリでの実装が容易となる。
FIG. 5 schematically shows the state. A vector having a subframe length is cut out from the i-th position of the -overlapping codebook in which the diagonal components are stored, and used as a diagonal element of Ai. Therefore, Ai and Ai + 1 have N-1 common elements, that is, overlapping components. The amount of memory at this time is N + M-1 (M is the number of conversion patterns). As described in the first embodiment, this is 0.295 k words when N = 40 and M = 256, which facilitates mounting with a lower memory.

【0053】なお、ここではオーバラップ符号帳のシフ
ト数を1としたが、シフト数を2以上の値にすることも
可能である。この場合、隣り合う変換行列がN−2以下
の重複した要素を持つことになり、これだけ自由度が高
くなるが、メモリ量は若干増加する。
Although the number of shifts of the overlap codebook is set to 1 here, the number of shifts can be set to a value of 2 or more. In this case, adjacent transformation matrices have overlapping elements equal to or less than N-2, and the degree of freedom increases accordingly, but the amount of memory slightly increases.

【0054】また、ここでは変換行列として対角成分の
み値を持ち、この対角成分をオーバラップさせたが、値
を持つ成分を対角方向だけに限定する必要はなく、例え
ば帯状行列としてオーバラップ化することも可能であ
る。
Although only the diagonal components have values as the transformation matrix and the diagonal components are overlapped here, it is not necessary to limit the components having the values only in the diagonal direction. Wrapping is also possible.

【0055】なお、CELP方式には図1および図3に
示したように合成フィルタ102,202が通常存在す
るが、本発明は合成フィルタを用いない、言い換えれば
合成フィルタの重みが1である場合にも適用することが
できる。
Although the CELP system normally has the synthesis filters 102 and 202 as shown in FIGS. 1 and 3, the present invention does not use a synthesis filter. In other words, when the weight of the synthesis filter is 1, Can also be applied.

【0056】合成フィルタを用いない場合、符号化側で
は適応符号帳101から取り出した適応符号ベクトルを
第1のピッチ信号とし、変換パターン符号帳104から
取り出した変換パターンで第1のピッチ信号を変換した
第2のピッチ信号と合成して合成音声信号を生成すれば
よい。この場合、当然のことながら符号化パラメータに
は合成フィルタのフィルタ係数は含まれない。一方、復
号側においては少なくとも適応符号ベクトルおよび変換
パターンを示すインデックスを符号化パラメータとして
入力し、符号化側と同様に適応符号帳から取り出した適
応符号ベクトルを第1のピッチ信号とし、変換パターン
符号帳から取り出した変換パターンで第1のピッチ信号
を変換した第2のピッチ信号と合成して復号音声信号を
生成すればよい。
When the synthesis filter is not used, the encoding side uses the adaptive code vector extracted from the adaptive codebook 101 as the first pitch signal, and converts the first pitch signal with the conversion pattern extracted from the conversion pattern codebook 104. What is necessary is just to synthesize | combine with the 2nd pitch signal and to generate | occur | produce a synthetic | combination audio signal. In this case, the encoding parameters do not include the filter coefficients of the synthesis filter. On the decoding side, on the other hand, at least an index indicating an adaptive code vector and a conversion pattern is input as an encoding parameter, and the adaptive code vector extracted from the adaptive codebook is used as a first pitch signal in the same manner as the encoding side, and the conversion pattern code The decoded voice signal may be generated by combining the first pitch signal with the second pitch signal obtained by converting the first pitch signal using the conversion pattern extracted from the book.

【0057】[0057]

【発明の効果】以上説明したように、本発明によれば従
来の雑音符号帳探索に必要であったフィルタリングの演
算が不要になり、低演算量での符号化が可能になるとと
もに、入力音声信号の性質に合った合成音声信号を生成
しやすく、音質が向上するという効果がある。
As described above, according to the present invention, the filtering operation required for searching for a noise codebook in the related art is not required, the encoding can be performed with a small amount of operation, and the input speech can be obtained. There is an effect that it is easy to generate a synthesized speech signal that matches the signal properties, and the sound quality is improved.

【図面の簡単な説明】[Brief description of the drawings]

【図1】本発明の第1の実施形態に係る音声符号化装置
の構成を示すブロック図
FIG. 1 is a block diagram illustrating a configuration of a speech encoding device according to a first embodiment of the present invention.

【図2】同実施形態の処理手順を示すフローチャートFIG. 2 is a flowchart showing a processing procedure according to the embodiment;

【図3】同実施形態に係る音声復号装置の構成を示すブ
ロック図
FIG. 3 is a block diagram showing the configuration of the speech decoding apparatus according to the embodiment;

【図4】本発明の第2の実施形態に係る音声符号化装置
の構成を示すブロック図
FIG. 4 is a block diagram illustrating a configuration of a speech encoding device according to a second embodiment of the present invention.

【図5】本発明における第1の応答信号に施す変換の様
子を模式的に示す図
FIG. 5 is a diagram schematically showing a state of conversion applied to a first response signal in the present invention.

【図6】従来のCELP方式の音声符号化装置の構成を
示すブロック図
FIG. 6 is a block diagram showing a configuration of a conventional CELP-type speech coding apparatus.

【図7】従来のCELP方式の音声復号化装置の構成を
示すブロック図
FIG. 7 is a block diagram showing a configuration of a conventional CELP-based speech decoding apparatus.

【符号の説明】[Explanation of symbols]

101…適応符号帳 102…合成フィルタ 103…分析フィルタ 104…変換パターン符号帳 105…変換部 106,107…ゲイン乗算器 108…利得符号帳 109…加算器 110…線形予測分析部 111…入力端子 112…減算器 113,141,142…聴感重みフィルタ 140…聴感重み逆フィルタ 114…評価部 121…利得符号帳のインデックス 122…変換パターン符号帳のインデックス 123…合成フィルタのインデックス 124…適応符号帳のインデックス 131…合成音声信号 132…入力音声信号 210…逆量子化部 201〜204…入力端子 201…適応符号帳 202…合成フィルタ 203…分析フィルタ 204…変換パターン符号帳 2105…変換部 206,207…ゲイン乗算器 208…利得符号帳 209…加算器 210…逆量子化部 211…復号音声信号信号 1011…適応符号帳 1012…雑音符号帳 1013…ゲイン符号帳 1021,1022,…ゲイン乗算器 1023…加算器 1030…聴感重み付き合成フィルタ 1031…復号音声信号 2035…合成フィルタ 1040…聴感重みフィルタ 1041…聴感重み付き入力音声信号 1050…線形予測分析部 1060…入力端子 1070…誤差評価部 1071…減算器 1080…LPC量子化部 2011…適応符号帳 2012…雑音符号帳 2013…ゲイン符号帳 2021,2022,…ゲイン乗算器 2023…加算器 2035…合成フィルタ 2081…LPC逆量子化部 2090…ポストフィルタ Reference Signs List 101 adaptive codebook 102 synthesis filter 103 analysis filter 104 conversion pattern codebook 105 conversion units 106 and 107 gain multiplier 108 gain codebook 109 adder 110 linear prediction analysis unit 111 input terminal 112 ... subtractors 113, 141, 142 ... perceptual weight filter 140 ... perceptual weight inverse filter 114 ... evaluator 121 ... gain codebook index 122 ... conversion pattern codebook index 123 ... synthesis filter index 124 ... adaptive codebook index 131 ... Synthesized speech signal 132 ... Input speech signal 210 ... Dequantizer 201-204 ... Input terminal 201 ... Adaptive codebook 202 ... Synthesis filter 203 ... Analysis filter 204 ... Conversion pattern codebook 2105 ... Conversion units 206 and 207 ... Gain Multiplier 208 ... Acquisition codebook 209 Adder 210 Dequantizer 211 Decoded speech signal signal 1011 Adaptive codebook 1012 Noise codebook 1013 Gain codebook 1021, 1022, Gain multiplier 1023 Adder 1030 Perceptual weight Attached synthesis filter 1031 decoded speech signal 2035 synthesis filter 1040 perceptual weight filter 1041 perceptual weighted input voice signal 1050 linear predictive analysis unit 1060 input terminal 1070 error evaluator 1071 subtractor 1080 LPC quantizer 2011 Adaptive codebook 2012 Noise codebook 2013 Gain codebook 2021, 2022 Gain multiplier 2023 Adder 2035 Synthesis filter 2081 LPC dequantizer 2090 Post filter

Claims (7)

【特許請求の範囲】[Claims] 【請求項1】入力音声信号の分析結果に基づいて係数が
決定される合成フィルタを駆動するための過去の駆動信
号に基づいて生成される適応符号ベクトルを格納した適
応符号帳から適応符号ベクトルを取り出してピッチ励振
信号とし、 このピッチ励振信号を前記合成フィルタに通して第1の
応答信号を生成し、 複数の変換パターンを格納した変換パターン符号帳から
取り出された変換パターンで前記第1の応答信号に変換
を施して第2の応答信号を生成し、 前記第1および第2の応答信号を合成して合成音声信号
を生成し、 前記入力音声信号に対する合成音声信号の歪がより小さ
くなる適応符号ベクトルおよび変換パターンを前記適応
符号帳および変換パターン符号帳からそれぞれ探索し、 少なくとも前記合成フィルタの係数と前記適応符号帳お
よび変換パターン符号帳から探索した適応符号ベクトル
および変換パターンを示すインデックスを符号化パラメ
ータとして出力することを特徴とする音声符号化方法。
An adaptive code vector is generated from an adaptive code book storing an adaptive code vector generated based on a past drive signal for driving a synthesis filter whose coefficient is determined based on an analysis result of an input voice signal. The pitch excitation signal is taken out as a pitch excitation signal. The pitch excitation signal is passed through the synthesis filter to generate a first response signal. The first response signal is obtained by using a conversion pattern extracted from a conversion pattern codebook storing a plurality of conversion patterns. Converting the signal to generate a second response signal; synthesizing the first and second response signals to generate a synthesized audio signal; and adapting the input audio signal to reduce the distortion of the synthesized audio signal. A code vector and a conversion pattern are respectively searched from the adaptive codebook and the conversion pattern codebook, and at least the coefficients of the synthesis filter and the adaptive Speech encoding method and outputting an index indicating an adaptive code vector and transformation pattern is searched from the issue book and conversion pattern codebook as an encoding parameter.
【請求項2】過去の合成音声信号に基づいて生成される
適応符号ベクトルを格納した適応符号帳から適応符号ベ
クトルを取り出して第1のピッチ信号とし、 複数の変換パターンを格納した変換パターン符号帳から
取り出された変換パターンで前記第1のピッチ信号に変
換を施して第2のピッチ信号を生成し、 前記第1および第2のピッチ信号を合成して合成音声信
号を生成し、 入力音声信号に対する合成音声信号の歪がより小さくな
る適応符号ベクトルおよび変換パターンを前記適応符号
帳および変換パターン符号帳からそれぞれ探索し、 少なくとも前記適応符号帳および変換パターン符号帳か
ら探索した適応符号ベクトルおよび変換パターンを示す
インデックスを符号化パラメータとして出力することを
特徴とする音声符号化方法。
2. A conversion pattern codebook that extracts an adaptive code vector from an adaptive codebook that stores an adaptive code vector generated based on a past synthesized speech signal and that is used as a first pitch signal, and that stores a plurality of conversion patterns. Converting the first pitch signal with the conversion pattern extracted from the first pitch signal to generate a second pitch signal; synthesizing the first and second pitch signals to generate a synthesized voice signal; The adaptive code vector and the conversion pattern searched from the adaptive code book and the conversion pattern code book, respectively, for the adaptive code vector and the conversion pattern which reduce the distortion of the synthesized speech signal with respect to the adaptive code book and the conversion pattern code book. Is output as an encoding parameter.
【請求項3】前記変換パターン符号帳に格納された複数
の変換パターンは、行列演算で表されることを特徴とす
る請求項1または2に記載の音声符号化方法。
3. The speech encoding method according to claim 1, wherein the plurality of conversion patterns stored in the conversion pattern codebook are represented by a matrix operation.
【請求項4】前記行列に非零の成分が5個以下である行
が存在することを特徴とする請求項3に記載の音声符号
化方法。
4. The speech coding method according to claim 3, wherein the matrix has rows in which the number of non-zero components is 5 or less.
【請求項5】前記行列が対角行列であり、前記変換パタ
ーン符号帳は、隣り合う行列の対角成分が重複する成分
を持つように構成されていることを特徴とする請求項3
に記載の音声符号化方法。
5. The conversion pattern codebook according to claim 3, wherein the matrix is a diagonal matrix, and the conversion pattern codebook is configured such that diagonal components of adjacent matrices have overlapping components.
3. The speech encoding method according to claim 1.
【請求項6】少なくとも合成フィルタのフィルタ係数と
適応符号ベクトルおよび変換パターンを示すインデック
スを符号化パラメータとして入力し、 適応符号ベクトルを格納した適応符号帳から前記符号化
パラメータに従って取り出される適応符号ベクトルをピ
ッチ励振信号として、このピッチ励振信号を前記符号化
パラメータに従って係数が決定される合成フィルタに通
して第1の応答信号を生成し、 複数の変換パターンを格納した変換パターン符号帳から
前記符号化パラメータに従って取り出される変換パター
ンで前記第1の応答信号に変換を施して第2の応答信号
を生成し、 前記第1および第2の応答信号を合成して復号音声信号
を生成することを特徴とする音声復号化方法。
6. An adaptive code vector extracted at least according to the coding parameter from an adaptive code book storing an adaptive code vector, at least an index indicating a filter coefficient of a synthesis filter, an adaptive code vector, and a conversion pattern is input. As a pitch excitation signal, the pitch excitation signal is passed through a synthesis filter whose coefficient is determined according to the encoding parameter to generate a first response signal, and the encoding parameter is obtained from a conversion pattern codebook storing a plurality of conversion patterns. Converting the first response signal with a conversion pattern extracted according to the following formula to generate a second response signal, and combining the first and second response signals to generate a decoded speech signal. Audio decoding method.
【請求項7】少なくとも適応符号ベクトルおよび変換パ
ターンを示すインデックスを符号化パラメータとして入
力し、 適応符号ベクトルを格納した適応符号帳から前記符号化
パラメータに従って取り出される適応符号ベクトルを第
1のピッチ信号とし、 複数の変換パターンを格納した変換パターン符号帳から
前記符号化パラメータに従って取り出される変換パター
ンで前記第1のピッチ信号に変換を施して第2のピッチ
信号を生成し、 前記第1および第2のピッチ信号を合成して復号音声信
号を生成することを特徴とする音声復号化方法。
7. An adaptive code vector extracted from an adaptive codebook storing an adaptive code vector in accordance with the encoding parameter is inputted as at least an index indicating an adaptive code vector and a conversion pattern as a first pitch signal. Converting the first pitch signal with a conversion pattern extracted according to the encoding parameter from a conversion pattern codebook storing a plurality of conversion patterns to generate a second pitch signal; A speech decoding method comprising: synthesizing a pitch signal to generate a decoded speech signal.
JP8246443A 1996-09-18 1996-09-18 Voice coding method and voice decoding method Pending JPH1091193A (en)

Priority Applications (1)

Application Number Priority Date Filing Date Title
JP8246443A JPH1091193A (en) 1996-09-18 1996-09-18 Voice coding method and voice decoding method

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
JP8246443A JPH1091193A (en) 1996-09-18 1996-09-18 Voice coding method and voice decoding method

Publications (1)

Publication Number Publication Date
JPH1091193A true JPH1091193A (en) 1998-04-10

Family

ID=17148535

Family Applications (1)

Application Number Title Priority Date Filing Date
JP8246443A Pending JPH1091193A (en) 1996-09-18 1996-09-18 Voice coding method and voice decoding method

Country Status (1)

Country Link
JP (1) JPH1091193A (en)

Cited By (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
WO2002063610A1 (en) * 2001-02-02 2002-08-15 Nec Corporation Voice code sequence converting device and method

Cited By (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
WO2002063610A1 (en) * 2001-02-02 2002-08-15 Nec Corporation Voice code sequence converting device and method
US7505899B2 (en) 2001-02-02 2009-03-17 Nec Corporation Speech code sequence converting device and method in which coding is performed by two types of speech coding systems

Similar Documents

Publication Publication Date Title
US5142584A (en) Speech coding/decoding method having an excitation signal
US20050261897A1 (en) Method and device for robust predictive vector quantization of linear prediction parameters in variable bit rate speech coding
JP3180762B2 (en) Audio encoding device and audio decoding device
JPH0990995A (en) Speech coding device
JP3628268B2 (en) Acoustic signal encoding method, decoding method and apparatus, program, and recording medium
JP3357795B2 (en) Voice coding method and apparatus
WO2002071394A1 (en) Sound encoding apparatus and method, and sound decoding apparatus and method
JP3275247B2 (en) Audio encoding / decoding method
JP3095133B2 (en) Acoustic signal coding method
JP3531780B2 (en) Voice encoding method and decoding method
JP6644848B2 (en) Vector quantization device, speech encoding device, vector quantization method, and speech encoding method
JP3268750B2 (en) Speech synthesis method and system
JPH1063300A (en) Audio decoding device and audio encoding device
JPH1091193A (en) Voice coding method and voice decoding method
JPH0519795A (en) Speech excitation signal encoding / decoding method
JP3088204B2 (en) Code-excited linear prediction encoding device and decoding device
JPH06282298A (en) Voice coding method
JP2002221998A (en) Acoustic parameter encoding / decoding method, apparatus and program, audio encoding / decoding method, apparatus and program
JP3276977B2 (en) Audio coding device
JP3192051B2 (en) Audio coding device
JPH0519796A (en) Speech excitation signal encoding / decoding method
JP2808841B2 (en) Audio coding method
JPH09179593A (en) Speech encoding device
JPH03243999A (en) Voice encoding system
JP2002527777A (en) Method for encoding or decoding audio signal samples and encoder or decoder