JPH0364800A - Audio encoding and decoding method - Google Patents

Audio encoding and decoding method

Info

Publication number
JPH0364800A
JPH0364800A JP1201741A JP20174189A JPH0364800A JP H0364800 A JPH0364800 A JP H0364800A JP 1201741 A JP1201741 A JP 1201741A JP 20174189 A JP20174189 A JP 20174189A JP H0364800 A JPH0364800 A JP H0364800A
Authority
JP
Japan
Prior art keywords
intersection
level
frame
encoding
crossing
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
JP1201741A
Other languages
Japanese (ja)
Inventor
Masashi Kuwabara
昌史 桑原
Hitoshi Takanashi
高梨 斉
Shinsaku Mori
森 真作
Wasaku Yamada
山田 和作
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Ricoh Co Ltd
Original Assignee
Ricoh Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Ricoh Co Ltd filed Critical Ricoh Co Ltd
Priority to JP1201741A priority Critical patent/JPH0364800A/en
Publication of JPH0364800A publication Critical patent/JPH0364800A/en
Pending legal-status Critical Current

Links

Landscapes

  • Transmission Systems Not Characterized By The Medium Used For Transmission (AREA)

Abstract

PURPOSE:To improve voice quality and shorten a processing time by encoding the value of an intersection level of each frame, the number of intersections, data showing which of positive and negative levels is crossed, and a time interval between an intersection and a last intersection. CONSTITUTION:This system has a means 10 which puts each specific number of samples in a frame and a means 20 which calculates a positive-negative intersection level (threshold value) from the values of the respective samples in the frame, and the value of the intersection level of each frame, the number of intersections, data indicating which of the negative and positive levels is crossed at each intersection, and the time interval (sample interval) between one intersection and its last intersection are encoded. Consequently, the voice quality is improved and the processing time is shortened.

Description

【発明の詳細な説明】 枝先立夏 本発明は、音声符号化及び復号化方式に関する。[Detailed description of the invention] branch tip summer The present invention relates to audio encoding and decoding systems.

丈米挟先 電話をはじめとする各種の音声の伝送技術はこれまで深
く研究され、進歩してきた。さらに、1938年に発明
されたPCM方式に端を発するデジタル通信方式は現在
でもピットレートを低くするという方向で研究が続けら
れている。この通信方式で音声等の伝送を行なう際に欠
かせないのが、デジタル音声処理である。代表的な音声
処理法は、波形符号化方式としてはA D P CM 
(N、S、JayantP、Cumm1nskey  
and  J、L、Flanagan、Adaptiv
e  quan−tization in diffe
rential pen carding of 5p
e−ech、Be1l 5yst、Tech、J、、5
2,7.1105−1118.1973.)、聴覚的符
号化方式としてはA T C(R,Zelinskia
nd P、No11.Adaptive transf
orm cording ofspeech sign
als、IEEE  Trans、^5SP−25,2
99−309゜1977、) 、声生成モデルに基づが
れた分析合成方式としてはPARCOR(北馬、板倉、
斉藤。
Various voice transmission technologies, including long-distance telephones, have been deeply researched and have made progress. Furthermore, research into digital communication systems, which originated from the PCM system invented in 1938, is still ongoing in the direction of lowering the pit rate. Digital audio processing is indispensable when transmitting audio, etc. using this communication method. A typical audio processing method is ADPCM as a waveform encoding method.
(N, S, JayantP, Cumm1nskey
and J. L., Flanagan, Adaptiv.
e quan-tization in diffe
rental pen carding of 5p
e-ech, Be1l 5yst, Tech, J,, 5
2, 7.1105-1118.1973. ), and the auditory encoding method is ATC (R, Zelinskia
nd P, No11. Adaptive transf
orm coding of speech sign
als, IEEE Trans, ^5SP-25,2
99-309゜1977,), PARCOR (Hokuma, Itakura,
Saito.

Parcor形音声分析合成系における最適符号構成。Optimal code structure in Parcor type speech analysis and synthesis system.

信学論(A) 、J61−A、2.119−126.1
978.) 、L S P〔管材、板倉、線スペクトル
対(LSP)音声分析合成方式による音声情報圧縮、信
学論(A)、J64−A、8,599−606.198
1.)などが提案されている。
IEICE Theory (A), J61-A, 2.119-126.1
978. ), L S P [Pipe Material, Itakura, Speech Information Compression Using Line Spectral Pair (LSP) Speech Analysis and Synthesis Method, IEICE Theory (A), J64-A, 8,599-606.198
1. ) have been proposed.

これらは、いづれの方法も音質の劣化を最小限にとどめ
つつ情報伝送量を下げることを目的とした符号化法であ
るが、いづれの符号化法も伝送量の低下に伴う音声品質
の劣化、処理時間の増大といった問題がある。しかし、
近年LSI技術の進展により高速な演算処理が可能とな
ったことを背景として、デジタル通信網における伝送コ
ストの低減や移動・衛星通信における周波数帯域の有効
利用を目的とした16kbps以下の低ビツトレート高
能率音声符号化方式の研究が活発に進められている。
All of these encoding methods aim to reduce the amount of information transmitted while minimizing the deterioration of sound quality. There is a problem of increased processing time. but,
In recent years, advances in LSI technology have made it possible to perform high-speed arithmetic processing, and as a result, low bit rate high efficiency of 16 kbps or less is aimed at reducing transmission costs in digital communication networks and effectively utilizing frequency bands in mobile and satellite communications. Research on speech coding methods is actively underway.

且−−」0 本発明は、上述のごとき実情に鑑みてなされたもので、
特に、移動通信などで将来需要が予想されている16k
bps以下の符号量における音声符号化の一方式として
、比較的処理時間の少ない、時間領域と周波数領域を組
み合わせた符号化方式を堤供することを目的としてなさ
れたものである。
And--"0 The present invention was made in view of the above-mentioned circumstances,
In particular, 16k is expected to be in demand in the future in mobile communications, etc.
This was developed with the aim of providing an encoding method that combines time domain and frequency domain and requires relatively little processing time as a method for audio encoding with a code amount of less than bps.

盪−一基 本発明は、上記目的を達成するために、(1)PCM符
号化された音声信号の圧縮方式において、所定のサンプ
ル数毎にフレーム化する手段と、フレーム内の各サンプ
ルの値を基に正負の交差レベル(しきい値)を演算する
手段を有し、各フレーム毎の交差レベルの値と、交差点
の数と、各交差点ごとに正負のどちらのレベルと交差し
たかを示す符号と、該交差点とその前の交差点との時間
間隔(サンプル間隔)を符号化すること、或いは、(2
)前記(1)において符号化された信号を復号する方式
において、前記交差点が前の交差点と異なる符号のレベ
ルで交差し、かつ、その間隔が所定の値よりも大きいと
きには該2点間を3次曲線で補間し、その間隔が所定の
値よりも小さいときには該2点間を直線で補間し、該交
差点と同じ符号のレベルと交差したときには、該2点間
を2次曲線で補間すること、或いは、(3)PCM符号
化された音声信号の圧縮方式において、入力されたPG
M音声信号を所定の区間でフレーム化し、その区間でピ
ッチ抽出し、ピッチ抽出が可能なときにはそのピッチ区
間を新たなフレームとしてフレーム内の各サンプルの値
を基に正負の交差レベル(しきい値)を演算する手段を
もち、各フレーム毎の交差レベルの値と、交差点の数と
、各交差点ごとに正負のどちらのレベルと交差したかを
示す符号と、該交差点とその前の交差点との時間間隔(
サンプル間隔)を符号化するとともに、これらの符号化
結果からローカル復号をする際に該交差点が前の交差点
と異なる符号のレベルで交差し、かつ、その間隔が所定
の値よりも大きいときには該2点間を3次曲線で補間し
、その間隔が所定の値よりも小さいときには該2点間を
直線で補間し、該交差点と同じ符号のレベルと交差した
ときには、該2点間を2次曲線で補間し、このローカル
復号の結果と元信号との差分を取り、その信号をフーリ
エ変換し、各調波のスペクトルのピーク点の周波数、振
幅と、前記フレーム毎の交差レベルの値と、交差点の数
と、各交差点ごとに正負のどちらのレベルと交差したか
を示す符号と、該交差点とその前の交差点との時間間隔
(サンプル間隔)を符号化すること、或いは、(4)前
記(3)に示した音声符号化データを復号する方式にお
いて、前記フレーム毎の交差レベルの値と、交差点の数
と、各交差点ごとに正負のどちらのレベルと交差したか
を示す符号と、該交差点とその前の交差点との時間間隔
(サンプル間隔)から、フレーム内の交差点が前の交差
点と異なる符号のレベルで交差し、かつ、その間隔が所
定の値よりも大きいときにはこれら2点間を3次曲線で
補間し、その間隔が所定の値よりも小さいときには該2
点間を直線で補間し、該交差点と同じ符号のレベルと交
差したときには、該2点間を2次曲線で補間することで
復号し、さらに、各調波のスペクトルのピーク点の周波
数、振幅を示す符号から、逆フーリエ変換を行って復号
した音声信号とを加算して復号することを特徴としたも
のである。以下本発明の実施例に基づいて説明する。
(2) In order to achieve the above object, the basic invention provides: (1) In a compression method for a PCM encoded audio signal, means for forming a frame every predetermined number of samples; It has a means for calculating positive and negative crossing levels (thresholds) based on the data, and a code indicating the crossing level value for each frame, the number of crossing points, and which level (positive or negative) is crossed for each crossing point. and encoding the time interval (sample interval) between the intersection and the previous intersection, or (2
) In the method for decoding the encoded signal in (1) above, when the intersection intersects with the previous intersection at a different code level and the interval is larger than a predetermined value, the distance between the two points is 3. Interpolate with a quadratic curve, and when the interval is smaller than a predetermined value, interpolate between the two points with a straight line, and when it intersects a level with the same sign as the intersection, interpolate between the two points with a quadratic curve. , or (3) In the compression method of a PCM encoded audio signal, the input PG
The M audio signal is framed in a predetermined section, the pitch is extracted in that section, and when pitch extraction is possible, the pitch section is used as a new frame and a positive/negative crossing level (threshold value) is set based on the value of each sample in the frame. ), and calculates the intersection level value for each frame, the number of intersections, a sign indicating which level (positive or negative) each intersection has crossed, and the relationship between the intersection and the previous intersection. Time interval(
When performing local decoding from these encoding results, if this intersection intersects at a different code level from the previous intersection and the interval is larger than a predetermined value, A cubic curve is used to interpolate between points, and when the interval is smaller than a predetermined value, a straight line is interpolated between the two points, and when the intersection intersects a level with the same sign as the intersection, a quadratic curve is used to interpolate between the two points. The difference between this local decoding result and the original signal is taken, and the signal is Fourier transformed, and the frequency and amplitude of the peak point of the spectrum of each harmonic, the value of the cross level for each frame, and the intersection point are calculated. , a code indicating which level (positive or negative) has been crossed for each intersection, and a time interval (sample interval) between the intersection and the previous intersection, or (4) the above ( In the method for decoding audio encoded data shown in 3), the value of the crossing level for each frame, the number of crossing points, a code indicating which level (positive or negative) is crossed for each crossing point, and the crossing point. Based on the time interval (sample interval) between the intersection and the previous intersection, if the intersection in the frame intersects at a different sign level from the previous intersection, and the interval is larger than a predetermined value, the distance between these two points is 3. Interpolate with the following curve, and if the interval is smaller than a predetermined value, the 2
A straight line is interpolated between the points, and when the intersection crosses a level with the same sign as the intersection point, decoding is performed by interpolating a quadratic curve between the two points, and the frequency and amplitude of the peak point of the spectrum of each harmonic are This method is characterized in that decoding is performed by adding a voice signal decoded by performing inverse Fourier transform from a code indicating . The present invention will be explained below based on examples.

本発明による符号化方式は、時間領域の符号化である適
応レベル交差法〔高架、秋田、森、2レベル交差法によ
る音声波形の符号化、信学全大、1367.1986.
)、(高架、森、適応レベル交差法による音声波形の符
号化、信学全大、1325、↓987.)と周波数領域
の符号化であるHarmonic Cording (
R,J、McAulay and T、F、Qua−t
iefi、5peech analysis/5ynt
hesis based on 5i−nusoida
l representation、IEEE Tra
ns、ASSP−34゜774−754.1986.)
、 (L、B、A11nsida and J、M、T
ribolet、Non5tationary 5pe
ctral modeling of voteeds
peech、IEEE Trans、ASSP−31,
664−678,1983,) 。
The encoding method according to the present invention is based on the adaptive level crossing method, which is time-domain encoding [Koukashi, Akita, Mori, Coding of speech waveforms by two-level crossing method, IEICE National University, 1367.1986.
), (Kakashi, Mori, Speech Waveform Coding Using Adaptive Level Crossing Method, IEICE University, 1325, ↓987.) and Harmonic Coding, which is frequency domain coding (
R, J, McAulay and T, F, Qua-t.
iefi, 5peech analysis/5ynt
hesis based on 5i-nusoida
l representation, IEEE Tra.
ns, ASSP-34゜774-754.1986. )
, (L, B, A11nsida and J, M, T
ribolet, Non5tationary 5pe
ctral modeling of votes
peach, IEEE Trans, ASSP-31,
664-678, 1983).

〔犬山、音声及び残差信号の調波構造を利用した2つの
音声分析合成系、音声研資、584−50゜1984、
)、(松雪、多田、音声の調波構造を利用した中帯域符
号化方式の検討、信学技報、5P88−57.1988
.3を組み合わせ、適応レベル交差法を用いる際にピッ
チ抽出を行い、1フレーム中の1ピッチ周期分の符号化
を行なうことによりピットレートの低減をはかったもの
である。
[Inuyama, Two speech analysis and synthesis systems using the harmonic structure of speech and residual signals, Speech Research Fund, 584-50゜1984,
), (Matsuyuki, Tada, Study of medium band coding method using harmonic structure of voice, IEICE Technical Report, 5P88-57.1988
.. 3 is combined, pitch extraction is performed when using the adaptive level crossing method, and the pit rate is reduced by encoding one pitch period in one frame.

最初に適応レベル交差法の概要について説明する。隣接
するサンプルが正負の異なる符号の値になったときに零
と交差をするが、その零交差を用いると、短時間の平均
周波数を推定することができるということが知られてい
る。この零交差だけで符号化をおこなうのでは、良好な
合成音を再生することは難しいと考えられる。波形符号
化において零交差を符号化するのではなく極点の符号化
を行なうことによって低ビツトレートにする符号化方法
として極点抽出法が提案されている〔松材。
First, an overview of the adaptive level crossing method will be explained. It is known that when adjacent samples have values of different signs (positive or negative), they cross zero, and by using these zero crossings, it is possible to estimate the short-term average frequency. It is thought that it is difficult to reproduce good synthesized speech if encoding is performed using only these zero crossings. A pole point extraction method has been proposed as an encoding method that lowers the bit rate by encoding the pole points rather than the zero crossings in waveform encoding [Matsuzai.

片動、吉田、菅田、極点抽出法(PVP法)による音声
符号化、信学技報、C385−32,1985、〕。こ
の極点抽出法の符号化方法は、極点間の時間間隔と極点
間の振幅の差を符号化し、それらのパラメータより極点
間の補間をおこなう。
Single motion, Yoshida, Suda, Speech coding using the polar point extraction method (PVP method), IEICE Technical Report, C385-32, 1985]. The encoding method of this pole point extraction method encodes the time interval between pole points and the difference in amplitude between pole points, and performs interpolation between pole points using these parameters.

極点抽出法と同様に、波形符号化による低ビツトレート
にする方法として零交差を応用した正負のレベルでの交
差による符号化を行なった適応しベル交差法が提案され
ている〔高梨、秋UJ1.森、2レベル交差法による音
声波形の符号化、信学全大、工367.1986.)、
(高梨、森、適応レベル交差法による音声波形の符号化
、信学全大、1325.1987.1゜この符号化方式
では、正または負の設定されたレベルに交差したところ
の時間間隔の符号化をおこなう。この符号化では、交差
した点を補間する方法なので、音声波形の概形を符号化
するにすぎない。そのため音声波形の低域成分を符号化
するために用いる。
Similar to the pole extraction method, an adaptive bell crossing method has been proposed in which coding is performed using crossings at positive and negative levels by applying zero crossings as a method of reducing the bit rate by waveform encoding [Takanashi, Aki UJ1. Mori, Coding of speech waveforms using two-level crossing method, IEICE, Engineering 367. 1986. ),
(Takanashi, Mori, Coding of speech waveforms by adaptive level crossing method, IEICE, 1325.1987.1゜In this coding method, the sign of the time interval at which the positive or negative set level is crossed is This encoding is a method of interpolating crossing points, so it only encodes the outline of the audio waveform.Therefore, it is used to encode the low-frequency components of the audio waveform.

次に、音声の符号化について説明する。Next, audio encoding will be explained.

第6図は、音声符号化の一例を説明するためのフローチ
ャートで、まず、lフレームの音声波形より平均振幅a
mpを(1)式により計算する。
FIG. 6 is a flowchart for explaining an example of audio encoding. First, from the audio waveform of 1 frame, the average amplitude a
mp is calculated using equation (1).

この平均振幅に係数を掛けることによって正負の交差レ
ベルを決定する。決定された正負のレベルに交差する音
声波形の交差点を符号化する。もし、隣接する2点間の
間に交差レベルが来た場合は交差レベルとの距離が近い
方を交差点とする。
The positive and negative crossing levels are determined by multiplying this average amplitude by a coefficient. The intersection of the audio waveform that intersects the determined positive and negative levels is encoded. If a crossing level comes between two adjacent points, the one that is closer to the crossing level is determined as the crossing point.

符号化するパラメータは、 (1)フレームごとに 1、交差レベル 2、交差点の数 (II)交差点ごとに 1、正負どちらのレベルと交差したかの区別 2、交差点と前の交差点との時間間隔Tである。The parameters to be encoded are (1) For each frame 1.Cross level 2. Number of intersections (II) At each intersection 1. Distinguishing whether the level is positive or negative 2. Time interval T between the intersection and the previous intersection.

次に、復号化について説明する。Next, decoding will be explained.

符号器から伝送されたパラメータより復号化を行なうが
、第7図はその方法を説明するためのフローチャートで (I)前の交差点と異符号のレベルの交差点のとき 1、時間間隔Tが大きいとき この区間を3次曲線で補間する 2、時間間隔が小さいとき この区間を直線で補間する (II)前の交差点と同符号のレベルの交差点のとき この区間を2次曲線で補間する。
Decoding is performed based on the parameters transmitted from the encoder, and FIG. 7 is a flowchart to explain the method. This section is interpolated with a cubic curve. 2. When the time interval is small, this section is interpolated with a straight line. (II) When the intersection has the same sign level as the previous intersection, this section is interpolated with a quadratic curve.

第8図は、適応レベル符号化の基本構成を示す図で1図
中、10は符号化部、20は復号化部で。
FIG. 8 is a diagram showing the basic configuration of adaptive level encoding. In the figure, 10 is an encoding section and 20 is a decoding section.

以下、交差点間を補間する時の係数の決定の仕方につい
て説明するが、以下、(I)前の交差点と異符号のレベ
ルの交差点のときと、(■)前の交差点と同符号のレベ
ルの交差点のときとは分けて説明する。
How to determine coefficients when interpolating between intersections will be explained below. This will be explained separately from the intersection.

最初に、前の交差点と異符号レベルの交差点のときにつ
いて説明すると、このときの交差点の取り方は、第9図
(a)、(b)に示すように時間間隔Tが大きいとき(
(a)図)と、時間間隔Tが小さいとき((b)図)の
2種類ある。
First, we will explain when the intersection is at a different sign level from the previous intersection.The way to take the intersection at this time is when the time interval T is large (
There are two types: (Figure (a)) and when the time interval T is small (Figure (b)).

今、現在の交差点をPL (T、−cd)、1つ前の交
差点をPO(0,cd)とする。時間間隔Tが大きいと
き、たとえば、(T>10)のときは、3次曲線によっ
て補間を行うが、その3次曲線による補間式を f  (t)=a t’+b t”+c t+d   
 (2)で定義する。このときの係数at bt at
 dの決定法について説明すると。
Now, assume that the current intersection is PL (T, -cd) and the previous intersection is PO (0, cd). When the time interval T is large, for example (T>10), interpolation is performed using a cubic curve.
Defined in (2). Coefficient at this time bt at
Let me explain how to determine d.

PO,PLがなりたつようにするために、補間式%式%
(3) (4) また、復号した波形がなめらかにつながるように、1つ
前の交差点POで、その1つ前の補間した曲線または直
線の傾きに等しいようにする。前の補間した曲線または
直線での交差点POの傾きをmとすると、補間式を微分
した が等しくなるようにすればよいので。
In order to make PO and PL equal, use the interpolation formula % formula %
(3) (4) Also, so that the decoded waveforms are smoothly connected, the slope of the previous intersection PO is made equal to the slope of the previous interpolated curve or straight line. Assuming that the slope of the intersection point PO on the previous interpolated curve or straight line is m, the interpolation formula can be differentiated to make it equal.

そして、T/2でOを通るように すると、係数at b、c、dは次のようになる。Then, pass through O at T/2 Then, the coefficients at b, c, and d are as follows.

C= m d=cd しかし、補間する3次曲線は、POとPlの間で交差点
をもたないはずなので、補間式はT/4でd = c 
d また、時間間隔Tが小さいとき、たとえば(T≦10)
のときは、直線によって補間を行なうが、その直線によ
る補間式を f(t)=at+b     (12)で定義すると、
POとPLを通るので、係数a。
C= m d=cd However, since the cubic curve to be interpolated should not have an intersection between PO and Pl, the interpolation formula is T/4 and d = c
d Also, when the time interval T is small, for example (T≦10)
When , interpolation is performed using a straight line, but if the interpolation formula using the straight line is defined as f(t)=at+b (12),
Since it passes through PO and PL, the coefficient a.

bは次のようになる。b becomes as follows.

が成り立つはずである。もし上の式が成り立たなかった
らT/4で−c d / 2となるようにとし、交差点
POでの前の曲線または直線との傾きが等しいというの
を条件から除くと、係数a。
should hold true. If the above equation does not hold, then set it to -c d / 2 at T/4, and remove from the condition that the slope is equal to the previous curve or straight line at the intersection PO, then the coefficient a.

b、c、dは次のようになる。b, c, and d are as follows.

b=cd 次に、前の交差点と同符号のレベルの交差点のときの取
り方について説明するが、この時の交差点の取り方は、
第9図(Q)に示すように1種類で、2次曲線で補間す
る。前記と同様に、現在の交差点をPL (T、cd)
、1つ前の交差点をPO(0,cd)とする。2つの交
差点が同符号のときは、2次曲線によって補間を行なう
が、その2次曲線による補間式を f(t)=a t2+b t+c  (14)で、定義
する。PO,PLがなりたつようにするために、補間式
は、 f(0)=cd=c    (15) f(T)=cd=aT”+bT+c  (16)となる
、また、復号した波形がなめらかにつながるように、1
つ前の交差点POで、その1つ前の補間した曲線または
直線の傾きに等しいようにするために、1つ前の交差点
POでの1つ前の補間した曲線または直線の傾きをmと
すると、補間式を微分した が等しくなるようにすればいいので。
b=cd Next, we will explain how to take an intersection when it has the same sign and level as the previous intersection.
As shown in FIG. 9(Q), one type of interpolation is performed using a quadratic curve. As before, let the current intersection be PL (T, cd)
, the previous intersection is PO(0, cd). When two intersections have the same sign, interpolation is performed using a quadratic curve, and the interpolation formula using the quadratic curve is defined as f(t)=a t2+b t+c (14). In order to make PO and PL equal, the interpolation formula is f(0)=cd=c (15) f(T)=cd=aT''+bT+c (16) Also, the decoded waveform is smooth. To connect, 1
In order to make the slope equal to the slope of the previous interpolated curve or straight line at the previous intersection PO, let m be the slope of the previous interpolated curve or straight line at the previous intersection PO. , we differentiated the interpolation formula, but we just need to make it equal.

ここで、補間する2次曲線は、POとPlの間で交差点
をもたないはずなので、補間式はT/2でf (T/2
)がcdと異符号のとき、が成り立つはずである。もし
、上の式(20)が%式% とし、交差点POでの前の曲線または直線との傾きが等
しいというのを条件から除くと、係数a。
Here, since the quadratic curve to be interpolated should not have an intersection between PO and Pl, the interpolation formula is T/2 and f (T/2
) should hold true when cd has a different sign. If the above equation (20) is expressed as %, and if we exclude from the condition that the slope of the previous curve or straight line at the intersection PO is equal, then the coefficient a.

b、cは次のようになる。b and c are as follows.

となる、よって、係数a、b、cは次のようになる。Therefore, the coefficients a, b, and c are as follows.

a = −一 b=m (19) c=cd c=cd また、T/2でf(T/2)が同符号で、3cdよりも
大きくなる場合は、 f(T/2)があまりにも大きく
なりすぎるので、その場合は、T/2で2cdになるよ
うに とし、交差点POでの前の曲線または直線との傾きが等
しいというのを条件から除くと、係数a。
a = -1b=m (19) c=cd c=cd Also, if f(T/2) has the same sign at T/2 and becomes larger than 3cd, then f(T/2) is too large. In that case, T/2 is set to 2 cd, and if we remove from the condition that the slope is equal to the previous curve or straight line at the intersection PO, the coefficient a.

b、cは次のようになる。b and c are as follows.

c=cd 次に、ハーモニックコーディング(HarmonicC
ording)について説明する。
c=cd Next, harmonic coding (HarmonicC
(ordering) will be explained.

音声の有声音区間では、信号の周期性により、その短時
間フーリエスペクトルは、一般に調波構造(基本周波数
とその高調波による周期的なスペクトル構造)を有する
。この調波構造に着目した符号化方式として、Harm
onic Cording (以下HCと書<)方式が
提案されている(R,J、McAulay andT、
F、(luatiefi、5peech analys
is/5ynthesis bas−ed on 5i
nusoidal representation、I
EEE Trans。
In a voiced period of speech, the short-time Fourier spectrum generally has a harmonic structure (a periodic spectral structure consisting of a fundamental frequency and its harmonics) due to the periodicity of the signal. As an encoding method that focuses on this harmonic structure, Harm
The onic coding (hereinafter referred to as HC) method has been proposed (R, J, McAulay and T,
F, (luatiefi, 5peech analyzes
is/5 synthesis bas-ed on 5i
Nusoidal representation, I
EEE Trans.

ASSP−34,774−754,1986) 、 (
L、B、A11neida and J。
ASSP-34, 774-754, 1986), (
L, B, A11neida and J.

M、Tribolet、Non5tationary 
 5pectral  modelingof voi
ced 5peech、IEEE Trans、ASS
P−31,664−678゜1983、) 、  (犬
山、音声及び残差信号の調波構造を利用した2つの音声
分析合成系、音声研資、584−50.1984.)、
(松雪、多田、音声の調波構造を利用した中′41F域
符号住方式の検討、信学技報、5P88−57、工98
8〕。これは周波数領域における調波構造を線スペクト
ルで表現し、音声波形を正弦波の重量により復号化する
ため、モデルが非常に単純でありパラメータの量子化ビ
ット数を大幅に低減することができる。
M, Tribolet, Non5tationary
5 pectral modeling of voi
ced 5peech, IEEE Trans, ASS
P-31, 664-678゜1983,), (Inuyama, Two speech analysis and synthesis systems using the harmonic structure of speech and residual signals, Voice Research Fund, 584-50.1984.),
(Matsuyuki, Tada, Study of the middle '41F range code system using the harmonic structure of speech, IEICE Technical Report, 5P88-57, Eng. 98
8]. This expresses the harmonic structure in the frequency domain as a line spectrum and decodes the audio waveform using the weight of a sine wave, so the model is very simple and the number of parameter quantization bits can be significantly reduced.

第10図は、HCの代表的な構成を示す図で、図中、1
0は符号器、20は復号器を示し、符号器10では、ま
ず、フレーム単位(20〜32m sec程度)で入力
音声にハミング窓(window)を掛けた後、フーリ
エ変換(DFT)により周波数領域に変換する。次に、
各周波の振幅スペクトルにおける各調波のピーク点を抽
出しくpeakPicking)、抽出したピーク点の
周波数、振幅、位相を符号化パラメータとして伝送する
。復号器20では伝送されたパラメータをもとにして音
声信号を再生する。復号化されたパラメータの周波数を
fl、振幅をml、位相をφ、としたとき、符号化音声
s (t)は1次式(25)となる。
FIG. 10 is a diagram showing a typical configuration of HC, and in the figure, 1
0 indicates an encoder, and 20 indicates a decoder. In the encoder 10, first, a Hamming window is applied to the input audio in frame units (approximately 20 to 32 msec), and then the frequency domain is processed by Fourier transform (DFT). Convert to next,
The peak point of each harmonic in the amplitude spectrum of each frequency is extracted (peakPicking), and the frequency, amplitude, and phase of the extracted peak point are transmitted as encoding parameters. The decoder 20 reproduces the audio signal based on the transmitted parameters. When the frequency of the decoded parameter is fl, the amplitude is ml, and the phase is φ, the encoded speech s (t) is expressed by the linear equation (25).

調波成分の抽出は、変換したスペクトルの振幅成分から
調波構造の各ピークを抽出して行うが。
The harmonic components are extracted by extracting each peak of the harmonic structure from the amplitude component of the converted spectrum.

ここで、ピークとは、FFT点のうち、隣り合う5点の
振幅スペクトル値の中心が最大で両側の点は順に小さく
なっている場合の中心の点をいう。
Here, the peak refers to the center point when the center of the amplitude spectrum values of five adjacent points among the FFT points is the maximum, and the points on both sides become smaller in order.

次に高周波補償を考慮した適応レベル交差法による符号
化について説明するが、最初に、適応レベル交差法によ
る音声劣化の要因について説明する。
Next, encoding using the adaptive level crossing method that takes high frequency compensation into consideration will be explained, but first, the causes of audio deterioration caused by the adaptive level crossing method will be explained.

第11図は、音声の有声音区間の振幅スペクトルを示す
図で、同図をみるとはっきりとしたホルマントと呼ばれ
るスペクトルのピークがみられる。
FIG. 11 is a diagram showing the amplitude spectrum of a voiced sound section of speech. Looking at the diagram, a clear peak of the spectrum called a formant can be seen.

普通、有声音、特に母音には3個程度の特徴的なホルマ
ントがあり、同じ音素でも発声者によりかなり大幅に変
動する。そのために、合成した音声に対しても入力音声
と同じようなスペクトル包絡をもつ必要がある。
Usually, voiced sounds, especially vowels, have about three characteristic formants, and even the same phoneme can vary considerably depending on the speaker. Therefore, it is necessary for the synthesized speech to have the same spectral envelope as the input speech.

第12図に、適応レベル交差法により符号化を行なった
ときの合成した音声の振幅スペクトルを示す、この合成
した音声の振幅スペクトルを見てみると、低周波の成分
に対しては、入力音声とほとんど変わりはないが、中・
高周波を見てみると、入力音声とはかなりずれてしまう
。そのため、第1ホルマント、第2ホルマントまでは再
生されているが、それより周波数の高いところでの振幅
のスペクトル包絡は再生することができず、第3ホルマ
ント以降のホルマントを再生することができない。適応
レベル交差法だけで符号化を行なうと高周波成分を再生
することができないので、このことが合成音声の劣化に
つながる要因と考えられる。
Figure 12 shows the amplitude spectrum of the synthesized speech when encoding is performed using the adaptive level crossing method. Looking at the amplitude spectrum of the synthesized speech, we see that for low frequency components, the input speech There is almost no difference between
If you look at the high frequencies, they will deviate considerably from the input audio. Therefore, although the first and second formants are reproduced, the spectral envelope of the amplitude at higher frequencies cannot be reproduced, and the formants after the third formant cannot be reproduced. If encoding is performed using only the adaptive level crossing method, high frequency components cannot be reproduced, and this is considered to be a factor that leads to deterioration of synthesized speech.

第工図は、本発明による符号化方式の一実施例を説明す
るための基本構成図、第2図(a)。
FIG. 2(a) is a basic configuration diagram for explaining an embodiment of the encoding method according to the present invention.

(b)は、符号器の動作を説明するためのフローチャー
ト、第3図は、復号器の動作を説明するためのフローチ
ャートで、第1図において、10は符号器、20は復号
器である。
(b) is a flowchart for explaining the operation of the encoder, and FIG. 3 is a flowchart for explaining the operation of the decoder. In FIG. 1, 10 is an encoder and 20 is a decoder.

前述のように、適応レベル交差法によって符号化を行っ
ただけでは、高周波成分を再生することができない。そ
のため、高周波酸を再生することが必要である。高周波
成分を再生するために、前述のHCを用いる。ここで用
いるHCは高周波成分のみを再生することを目的として
いるので、正弦波を重量するときの周波数は、高周波の
みである。普通、HCによって符号化を行うときは周波
数、振幅、位相をパラメータとして伝送している。
As mentioned above, high frequency components cannot be reproduced simply by encoding using the adaptive level crossing method. Therefore, it is necessary to regenerate the high frequency acid. The aforementioned HC is used to reproduce high frequency components. Since the HC used here is intended to reproduce only high frequency components, the frequencies used when weighing a sine wave are only high frequencies. Normally, when encoding is performed using HC, frequency, amplitude, and phase are transmitted as parameters.

しかし、低周波成分に対しては、適応レベル交差法によ
って符号化を行うことによって再生していると考えられ
るので、ここで用いるH Cでは、高周波のみに対して
符号化を行えばよいと考えられる。また、人間の聴覚は
高周波の位相に対しては鈍感であると言われていて、そ
の性質に着目した符号化方式も知られている〔守谷、誉
田、位相等化とベクトル量子化に基づく中帯域符号化、
信学論(A) 、 J 70. A、 S、 445−
452゜1987、)。ここではその性質を利用して、
位相を伝送せず周波数と振幅をパラメータとして伝送す
る。
However, since low frequency components are considered to be reproduced by encoding them using the adaptive level crossing method, it is thought that with the HC used here, it is sufficient to encode only high frequencies. It will be done. Furthermore, human hearing is said to be insensitive to the phase of high frequencies, and there are encoding methods that focus on this property [Moriya, Honda, band encoding,
Theory of Faith (A), J 70. A, S, 445-
452°1987,). Here, we will take advantage of this property,
Transmits frequency and amplitude as parameters without transmitting phase.

適応レベル交差法を行ったあとに、tI Cを用いて符
号化を行うことになるが、その方法は、入力音声と、適
応レベル交差法によって符号化したものの部分復号との
残差に対してHCによって符号化を行う。そのときHC
で符号化するときに伝送するパラメータは高周波の周波
数と振幅である。
After performing the adaptive level crossing method, encoding is performed using tIC, but this method uses the residual difference between the input speech and the partial decoding of the coded by the adaptive level crossing method. Encoding is performed by HC. At that time H.C.
The parameters transmitted when encoding are the frequency and amplitude of the high frequency.

第1図は、このHCによって高域を強調した適応レベル
符号化の基構成を示す図で、この符号化方式による有声
音区間での合成した音声の振幅スペクトルを第4図に示
す。第12図に示したHCを用いていないものに比べて
第11図に示した原音声の振幅スペクトルに近いものが
再生されていることがわかる。また、第5図に原音声(
(a)図)と本発明による符号化方式による合成した音
声の波形((b)図)を示す。
FIG. 1 is a diagram showing the basic configuration of adaptive level encoding in which high frequencies are emphasized by this HC, and FIG. 4 shows the amplitude spectrum of synthesized speech in a voiced section using this encoding method. It can be seen that the amplitude spectrum closer to the original sound shown in FIG. 11 is reproduced compared to the one not using HC shown in FIG. 12. In addition, the original audio (
FIG. 3(a) shows the waveform of synthesized speech using the encoding method according to the present invention (see FIG. 3(b)).

次に、ピッチを考慮した適応レベルの符号化について説
明する。
Next, adaptive level encoding that takes pitch into consideration will be explained.

有声音においては、はぼ周期的な波形を示すので、適応
レベル符号化を行うときに、1フレームの中で、工周期
分の波形のみを適応レベル符号化を行うことによって、
ビットレートを低くすることが可能である。
Since voiced sounds exhibit periodic waveforms, when performing adaptive level encoding, by performing adaptive level encoding only on the waveform for the period of time in one frame,
It is possible to lower the bit rate.

周期を抽出する方法は、始めに高いレベルで適応レベル
符号化を行って、その交差点を求める。
The method for extracting the period is to first perform adaptive level encoding at a high level and find the intersection point.

次に、その中の1周期分の波形に対して適応レベル符号
化を行い、交差点を求める。これらのパラメータを復号
器に伝送する。復号器では、始めに1周期分の波形を復
号し、その波形を高いレベルで行なった交差点に合わせ
るようにlフレーム分の復号を行なう。
Next, adaptive level encoding is performed on one period of the waveform, and an intersection point is determined. Transmit these parameters to the decoder. The decoder first decodes one period of the waveform, and then decodes one frame so that the waveform matches the intersection point made at a high level.

また、無声音やわたりの部分などにおいて、うまくピッ
チを抽出できないフレームに対しては、規定値の時間で
フレームとし、フレーム全体での適応レベル符号化を行
う。
Furthermore, for frames in which the pitch cannot be successfully extracted, such as in unvoiced sounds or transition parts, the frame is set at a specified time, and adaptive level encoding is performed on the entire frame.

次に、符号器及び復号器の手順について具体的に説明す
る。
Next, the procedures of the encoder and decoder will be specifically explained.

まず、第2図を参照しながら符号器について説明すると
、符号器の手順は、 1、フレーム長256サンプル、フレームオーバーラツ
プ15サンプルでフレーム切り出しを行なう。
First, the encoder will be explained with reference to FIG. 2. The encoder's steps are as follows: 1. Frame extraction is performed using a frame length of 256 samples and a frame overlap of 15 samples.

2、ピッチを抽出する。2. Extract the pitch.

ピッチを抽出できなかった場合は、フレーム全体で適応
レベル符号化を行なう。
If the pitch cannot be extracted, adaptive level encoding is performed on the entire frame.

ピッチを抽出できた場合は、フレーム中の1ピツチに対
して適応レベル符号化を行なう。
If the pitch can be extracted, adaptive level encoding is performed on one pitch in the frame.

ピッチ抽出の有無とそれぞれに対しての交差点数と交差
点間の間隔を伝送する。ピッチを抽出を行なったときは
、抽出したピッチの長さも伝送する。
The presence or absence of pitch extraction, the number of intersections for each, and the interval between intersections are transmitted. When a pitch is extracted, the length of the extracted pitch is also transmitted.

3、部分復号を行い、入力音声との残差をとる。3. Perform partial decoding and take the residual difference from the input voice.

4、残差に対してFFTを行い、高周波成分のところの
FFTのピークの周波数と振幅を求め、その結果を伝送
する。
4. Perform FFT on the residual, find the frequency and amplitude of the FFT peak at the high frequency component, and transmit the results.

となる。becomes.

次に、第3図を参照しなから復号器の手順について説明
すると復号器の手順は、 1.平均振幅と交差点数と交差点間の間隔から適応レベ
ル交差法による復号を行なう。
Next, the decoder procedure will be explained with reference to FIG. 3. The decoder procedure is as follows: 1. Decoding is performed using the adaptive level crossing method from the average amplitude, number of intersections, and interval between intersections.

2、周波数と振幅によりHCによる復号を行なう。2. Perform HC decoding using frequency and amplitude.

3、適応レベル符号化とHCのそれぞれによる復号化し
たものを足し合わせて全体の復号化を行なう。
3. The decoded data obtained by adaptive level coding and HC are added together to perform overall decoding.

となる。becomes.

勲二−二最 以上の説明から明らかなように、本発明によると、適応
レベル交差を行ったあとに、HCを用いて符号化を行っ
ているため、HCを用いないものに比して、原音声の振
幅音声に近いものを再生することができる。また、適応
レベル符号化を行うときに、1フレームの中で、■周期
分の波形のみを適応レベル符号化を行うことによって、
ビットレートを低くすることができる。
As is clear from the above explanation, according to the present invention, encoding is performed using HC after adaptive level crossing, so the original It is possible to reproduce audio amplitudes close to audio. Also, when performing adaptive level encoding, by performing adaptive level encoding only on the waveform for ■cycles in one frame,
Bitrate can be lowered.

【図面の簡単な説明】[Brief explanation of drawings]

第1図は、本発明による符号化・方式の基本構成を示す
図、第2図は、音声符号化のフローチャート、第3図は
、復号化のフローチャート、第4図は、本発明による符
号化方式によって得られた有声音の振幅スペクトルを示
す図、第5図は、原音声(、)及び音声音声(b)の波
形を示す図、第6図は、音声符号化の一例を説明するた
めのフローチャート、第7図は、復号化の一例を説明す
るためのフローチャート、第8図は適応レベル符号化の
従来技術の一例を説明するための基本構成図、第9図は
、交差点間の補間の方法を説明するための図、第10図
は、HCの基本構成図、第11図は、有声音原音声の振
幅スペクトルを示す図、第12図は、適応レベル交差法
を用いた場合の有声音の振幅スペクトルを示す図である
。 10・・・符号器、20・・・復号器 第1図 第2図 (a) ORDER DECORDER 第 2 図 第 6 図 第 図 第 8 図 第 図 (a) (b) a P。 (C) Pす 第 1 図 frequenc) 第 2 図 frequency
Figure 1 is a diagram showing the basic configuration of the encoding method according to the present invention, Figure 2 is a flowchart of speech encoding, Figure 3 is a flowchart of decoding, and Figure 4 is a diagram showing the encoding method according to the present invention. Figure 5 is a diagram showing the amplitude spectrum of the voiced sound obtained by the method. Figure 5 is a diagram showing the waveforms of the original voice (,) and voiced voice (b). Figure 6 is for explaining an example of voice encoding. FIG. 7 is a flowchart for explaining an example of decoding, FIG. 8 is a basic configuration diagram for explaining an example of a conventional technique of adaptive level encoding, and FIG. 9 is a flowchart for explaining an example of decoding. Figure 10 is a basic configuration diagram of HC, Figure 11 is a diagram showing the amplitude spectrum of the voiced original voice, and Figure 12 is a diagram for explaining the method when using the adaptive level crossing method. FIG. 3 is a diagram showing an amplitude spectrum of a voiced sound. 10... Encoder, 20... Decoder Figure 1 Figure 2 (a) ORDER DECORDER Figure 2 Figure 6 Figure 8 Figure 8 (a) (b) a P. (C) Figure 1 Frequency) Figure 2 Frequency

Claims (1)

【特許請求の範囲】 1、PCM符号化された音声信号の圧縮方式において、
所定のサンプル数毎にフレーム化する手段と、フレーム
内の各サンプルの値を基に正負の交差レベルを演算する
手段とを有し、各フレーム毎の交差レベルの値と、交差
点の数と、各交差点ごとに正負のどちらのレベルと交差
したかを示す符号と、該交差点とその前の交差点との時
間間隔を符号化することを特徴とする音声符号化方式。 2、請求項第1項において符号化された信号を復号する
方式において、前記交差点が前の交差点と異なる符号の
レベルで交差し、かつ、その間隔が所定の値よりも大き
いときには該2点間を3次曲線で補間し、その間隔が所
定の値よりも小さいときには該2点間を直線で補間し、
該交差点と同じ符号のレベルと交差したときには、該2
点間を2次曲線で補間することを特徴とする音声復号化
方式。 3、PCM符号化された音声信号の圧縮方式において、
入力されたPCM音声信号を所定の区間でフレーム化し
、その区間でピッチ抽出し、ピッチ抽出が可能なときに
はそのピッチ区間を新たなフレームとしてフレーム内の
各サンプルの値を基に正負の交差レベルを演算する手段
を有し、各フレーム毎の交差レベルの値と、交差点の数
と、各交差点ごとに正負のどちらのレベルと交差したか
を示す符号と、該交差点とその前の交差点との時間間隔
を符号化するとともに、これらの符号化結果からローカ
ル復号をする際に該交差点が前の交差点と異なる符号の
レベルで交差し、かつ、その間隔が所定の値よりも大き
いときには該2点間を3次曲線で補間し、その間隔が所
定の値よりも小さいときには該2点間を直線で補間し、
該交差点と同じ符号のレベルと交差したときには、該2
点間を2次曲線で補間し、このローカル復号の結果と元
信号との差分を取り、その信号をフーリエ変換し、各調
波のスペクトルのピーク点の周波数、振幅と、前記フレ
ーム毎の交差レベルの値と、交差点の数と、各交差点ご
とに正負のどちらのレベルと交差したかを示す符号と、
該交差点とその前の交差点との時間間隔を符号化するこ
とを特徴とする音声符号化方式。 4、請求項第3項に示した音声符号化データを復号する
方式において、前記フレーム毎の交差レベルの値と、交
差点の数と、各交差点ごとに正負のどちらのレベルと交
差したかを示す符号と、該交差点とその前の交差点との
時間間隔から、フレーム内の交差点が前の交差点と異な
る符号のレベルで交差し、かつ、その間隔が所定の値よ
りも大きいときには該2点間を3次曲線で補間し、その
間隔が所定の値よりも小さいときには該2点間を直線で
補間し、該交差点と同じ符号のレベルと交差したときに
は、該2点間を2次曲線で補間することで復号し、さら
に、各調波のスペクトルのピーク点の周波数、振幅を示
す符号から、逆フーリエ変換を行って復号した音声信号
とを加算して復号することを特徴とする音声復号化方式
[Claims] 1. In a compression method for a PCM encoded audio signal,
It has means for creating a frame for each predetermined number of samples, and means for calculating positive and negative crossing levels based on the values of each sample in the frame, and calculating the crossing level value for each frame, the number of crossing points, A voice encoding method characterized by encoding a code indicating whether a positive or negative level has been crossed for each intersection, and a time interval between the intersection and the previous intersection. 2. In the method for decoding encoded signals in claim 1, when the intersection intersects with the previous intersection at a different code level and the interval between the two points is larger than a predetermined value, interpolate with a cubic curve, and when the interval is smaller than a predetermined value, interpolate with a straight line between the two points,
When crossing the level with the same code as the intersection, the 2
An audio decoding method characterized by interpolating between points using a quadratic curve. 3. In the compression method for PCM encoded audio signals,
The input PCM audio signal is framed in a predetermined section, the pitch is extracted in that section, and when pitch extraction is possible, the pitch section is used as a new frame and the positive and negative crossing levels are calculated based on the values of each sample in the frame. It has means for calculating, and calculates the value of the intersection level for each frame, the number of intersections, a sign indicating whether the level is positive or negative for each intersection, and the time between the intersection and the previous intersection. When encoding the interval and performing local decoding from these encoding results, if the intersection intersects with the previous intersection at a different code level and the interval is larger than a predetermined value, the interpolate with a cubic curve, and when the interval is smaller than a predetermined value, interpolate with a straight line between the two points,
When crossing the level with the same code as the intersection, the 2
Interpolate between points using a quadratic curve, take the difference between this local decoding result and the original signal, perform Fourier transform on that signal, and calculate the frequency and amplitude of the peak point of the spectrum of each harmonic and the intersection for each frame. The value of the level, the number of intersections, and a sign indicating which level (positive or negative) was crossed for each intersection,
A voice encoding method characterized by encoding the time interval between the intersection and the intersection before it. 4. In the method for decoding encoded audio data set forth in claim 3, the value of the crossing level for each frame, the number of crossing points, and whether positive or negative levels are crossed for each crossing point are indicated. Based on the code and the time interval between the intersection and the previous intersection, if the intersection in the frame intersects at a different code level from the previous intersection, and the interval is larger than a predetermined value, the intersection between the two points is determined. Interpolates with a cubic curve, and when the interval is smaller than a predetermined value, interpolates with a straight line between the two points, and when it intersects a level with the same sign as the intersection, interpolates with a quadratic curve. An audio decoding method is characterized in that the audio signal is decoded by performing inverse Fourier transform from the code indicating the frequency and amplitude of the peak point of the spectrum of each harmonic, and the decoded audio signal is added to perform decoding. .
JP1201741A 1989-08-03 1989-08-03 Audio encoding and decoding method Pending JPH0364800A (en)

Priority Applications (1)

Application Number Priority Date Filing Date Title
JP1201741A JPH0364800A (en) 1989-08-03 1989-08-03 Audio encoding and decoding method

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
JP1201741A JPH0364800A (en) 1989-08-03 1989-08-03 Audio encoding and decoding method

Publications (1)

Publication Number Publication Date
JPH0364800A true JPH0364800A (en) 1991-03-20

Family

ID=16446171

Family Applications (1)

Application Number Title Priority Date Filing Date
JP1201741A Pending JPH0364800A (en) 1989-08-03 1989-08-03 Audio encoding and decoding method

Country Status (1)

Country Link
JP (1) JPH0364800A (en)

Cited By (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2003527622A (en) * 1999-07-19 2003-09-16 クゥアルコム・インコーポレイテッド Method and apparatus for identifying frequency bands to calculate a linear phase shift between frame prototypes in a speech coder
JP2005531014A (en) * 2002-06-27 2005-10-13 サムスン エレクトロニクス カンパニー リミテッド Audio coding method and apparatus using harmonic components
JP2020187377A (en) * 2013-12-27 2020-11-19 ソニー株式会社 Decoding device and method, and program

Cited By (7)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2003527622A (en) * 1999-07-19 2003-09-16 クゥアルコム・インコーポレイテッド Method and apparatus for identifying frequency bands to calculate a linear phase shift between frame prototypes in a speech coder
JP4860860B2 (en) * 1999-07-19 2012-01-25 クゥアルコム・インコーポレイテッド Method and apparatus for identifying frequency bands to calculate a linear phase shift between frame prototypes in a speech coder
JP2005531014A (en) * 2002-06-27 2005-10-13 サムスン エレクトロニクス カンパニー リミテッド Audio coding method and apparatus using harmonic components
JP2020187377A (en) * 2013-12-27 2020-11-19 ソニー株式会社 Decoding device and method, and program
JP2021177260A (en) * 2013-12-27 2021-11-11 ソニーグループ株式会社 Decoding device and method, and program
US11705140B2 (en) 2013-12-27 2023-07-18 Sony Corporation Decoding apparatus and method, and program
US12183353B2 (en) 2013-12-27 2024-12-31 Sony Group Corporation Decoding apparatus and method, and program

Similar Documents

Publication Publication Date Title
CN100370517C (en) A method for decoding encoded signals
CN100369112C (en) Variable Rate Speech Coding
KR100427753B1 (en) Method and apparatus for reproducing voice signal, method and apparatus for voice decoding, method and apparatus for voice synthesis and portable wireless terminal apparatus
US11721349B2 (en) Methods, encoder and decoder for linear predictive encoding and decoding of sound signals upon transition between frames having different sampling rates
Kondoz Digital speech: coding for low bit rate communication systems
CN102664003B (en) Residual excitation signal synthesis and voice conversion method based on harmonic plus noise model (HNM)
CN101359978B (en) Method for control of rate variant multi-mode wideband encoding rate
Tachibana et al. An investigation of noise shaping with perceptual weighting for WaveNet-based speech generation
US6138092A (en) CELP speech synthesizer with epoch-adaptive harmonic generator for pitch harmonics below voicing cutoff frequency
JP2004101720A (en) Acoustic encoding apparatus and acoustic encoding method
EP0843302B1 (en) Voice coder using sinusoidal analysis and pitch control
TWI281657B (en) Method and system for speech coding
JPH10124094A (en) Voice analysis method and method and device for voice coding
JPS61134000A (en) Speech analysis and synthesis method
TW463143B (en) Low-bit rate speech encoding method
CN105765653A (en) Adaptive high-pass post-filter
JP2003526123A (en) Audio decoder and method for decoding audio
KR20060059297A (en) Codevector Generation Method with Bitrate Elasticity and Wideband Vocoder Using the Same
JPH0364800A (en) Audio encoding and decoding method
US8719012B2 (en) Methods and apparatus for coding digital audio signals using a filtered quantizing noise
CN1650156A (en) Method and device for speech coding in an analysis-by-synthesis speech coder
Alku et al. Linear predictive method for improved spectral modeling of lower frequencies of speech with small prediction orders
CN101533639B (en) Voice signal processing method and device
Bhatia et al. Matrix quantization and LPC vocoder based linear predictive for low-resource speech recognition system
JP4826580B2 (en) Audio signal reproduction method and apparatus