JPH0378639B2 - - Google Patents

Info

Publication number
JPH0378639B2
JPH0378639B2 JP57040735A JP4073582A JPH0378639B2 JP H0378639 B2 JPH0378639 B2 JP H0378639B2 JP 57040735 A JP57040735 A JP 57040735A JP 4073582 A JP4073582 A JP 4073582A JP H0378639 B2 JPH0378639 B2 JP H0378639B2
Authority
JP
Japan
Prior art keywords
encoded data
pcm
adpcm
speech
audio
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Expired - Lifetime
Application number
JP57040735A
Other languages
Japanese (ja)
Other versions
JPS58158693A (en
Inventor
Masayoshi Yurugi
Shizuo Nagata
Naoji Akutsu
Takanori Murata
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Oki Electric Industry Co Ltd
Original Assignee
Oki Electric Industry Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Oki Electric Industry Co Ltd filed Critical Oki Electric Industry Co Ltd
Priority to JP57040735A priority Critical patent/JPS58158693A/en
Publication of JPS58158693A publication Critical patent/JPS58158693A/en
Publication of JPH0378639B2 publication Critical patent/JPH0378639B2/ja
Granted legal-status Critical Current

Links

Landscapes

  • Compression, Expansion, Code Conversion, And Decoders (AREA)
  • Transmission Systems Not Characterized By The Medium Used For Transmission (AREA)

Description

【発明の詳細な説明】 (従来技術) この発明は語単位の原音声を、適応差分パルス
符号変調方式(Adaptive Differential Pulse
Code Modulation、以下ADPCMと略す)を用
いて符号化する音声符号化方法に関する。
DETAILED DESCRIPTION OF THE INVENTION (Prior Art) This invention modulates word-by-word original speech using adaptive differential pulse code modulation.
This invention relates to a speech encoding method that uses code modulation (hereinafter abbreviated as ADPCM).

(背景技術) 従来の音声編集システムのブロツク図を第1図
に示す。第1図において、1は音声の入力端子、
2は音声を分析しADPCM符号化を行なう分析
部、3は分析部2で符号化された符号化データを
一担格納するサブメモリ部、4はサブメモリ部か
ら必要とする符号化データを格納するメインメモ
リ部、5は符号化データを原音声に変換する合成
部、6は分析部2、サブメモリ部3、メインメモ
リ部4及び合成部5の制御及び分析、合成の機能
を備えた中央処理装置CPUである。
(Background Art) A block diagram of a conventional audio editing system is shown in FIG. In Fig. 1, 1 is an audio input terminal;
2 is an analysis section that analyzes audio and performs ADPCM encoding; 3 is a sub-memory section that stores the encoded data encoded by the analysis section 2; and 4 stores the necessary encoded data from the sub-memory section. 5 is a synthesis unit that converts encoded data into original speech; 6 is a central unit that has functions for controlling, analyzing, and synthesizing the analysis unit 2, submemory unit 3, main memory unit 4, and synthesis unit 5; Processing unit CPU.

従来の音声編集方式は第2図に示すフローチヤ
ートのように、語単位の音声(以下音声片とい
う)を入力端子1に入力し(ステツプ8)、分析
部2により分析してADPCM符号化し(ステツプ
9)、該結果を音声の符号化データとして一旦サ
ブメモリ部3に格納しておき(ステツプ10)、必
要とする語単位(例えば“A”)の符号化データ
をサブメモリ部3から読出し(ステツプ11)、該
データをメインメモリ4に格納し(ステツプ12)、
次の音声片(例えば“B”)を含んだ入力音声を
同様に分析してADPCM符号化し、必要とする音
声片“B”の符号化データを読出し、メインメモ
リ4に格納する。以下同様にして編集に必要な音
声片の読出し及び格納を行なう。
In the conventional audio editing method, as shown in the flowchart shown in Figure 2, a word-by-word audio (hereinafter referred to as a audio piece) is input to the input terminal 1 (step 8), analyzed by the analysis unit 2, and encoded into ADPCM ( Step 9), the result is temporarily stored in the sub-memory section 3 as audio encoded data (step 10), and the encoded data of the required word unit (for example, "A") is read out from the sub-memory section 3. (Step 11), stores the data in the main memory 4 (Step 12),
The input speech including the next speech piece (for example, "B") is similarly analyzed and ADPCM encoded, and the encoded data of the necessary speech piece "B" is read out and stored in the main memory 4. Thereafter, the audio pieces necessary for editing are read out and stored in the same manner.

次に編集時に希望する順序(例えば“B”,
“A”,“C”の順)で各音声片の符号化データを
メインメモリ部4から読み出し(ステツプ14)、
音声合成器5により原音声に再生し(ステツプ
15)、希望の音声出力“BAC”を得る(ステツプ
16)。
Next, the desired order when editing (e.g. “B”,
The encoded data of each voice piece is read out from the main memory unit 4 in the order of “A” and “C” (step 14),
The voice synthesizer 5 reproduces the original voice (step
15), obtain the desired audio output “BAC” (step
16).

しかしながら、ADPCM方式は周知のように、
入力音声の変化量(すなわち差分値)に応じてデ
ルタパラメータΔpが最小値“0”から最大値
“MAX(Δp)”の範囲で変わるため、各音声片の
デルタパラメータΔpが“0”で終らないことが
ある。例えば連続した音声から音声片を切り出す
場合、音声片(例えば“B”)の終端が第3図a
に示すように、変化量が大きい状態で終ることが
ある。この場合前述したように音声片終端のデル
タパラメータΔp(B)は第3図bに示すように“0”
以上のある値“α”となつている。しかしなが
ら、第4図aに示す次の音声片(例えば“A”)
の先端は第4図bに示すようにデルタパラメータ
Δp(A)=αが初期化され分析され符号化されてい
るため、音声片“B,A”をこのまま結合し、合
成すると、本来は先端のΔp(A)が“0”であれば
原音声に再生されるが実際は“α”となつている
ので、第5図に示すように波形が変化してしまい
(a)、またゼロレベルがシフトする(b)というような
不都合が生ずる。
However, as is well known, the ADPCM method
Since the delta parameter Δp changes in the range from the minimum value "0" to the maximum value "MAX (Δp)" depending on the amount of change in the input audio (that is, the difference value), the delta parameter Δp of each voice piece does not end at "0". Sometimes there isn't. For example, when cutting out a speech segment from continuous speech, the end of the speech segment (for example, "B") is
As shown in Figure 2, the amount of change may end up being large. In this case, as mentioned above, the delta parameter Δp(B) at the end of the speech segment is "0" as shown in Figure 3b.
The value "α" is set to a certain value above. However, the next speech piece shown in Figure 4a (e.g. "A")
As shown in Figure 4b, the delta parameter Δp(A)=α has been initialized, analyzed, and encoded at the tip of , so if the speech pieces "B, A" are combined and synthesized as they are, the tip of If Δp(A) is “0”, the original audio will be reproduced, but since it is actually “α”, the waveform will change as shown in Figure 5.
Inconveniences such as (a) and (b) that the zero level shifts occur.

(発明の目的) この発明は、このような従来の問題点に着目し
てなされたもので、原音声を一担PCM符号化デ
ータとし、該PCM符号化データの冗長部分(連
続音声では語と語をつなぐ音声若しくはその符号
化データであり、語単位の音声では語と語の間の
無音部若しくはその符号化データである)を切断
した各音声片の終端に強制的に無音データを付加
しその後ADPCM符号化することにより上記問題
点を解決することを目的としている。
(Purpose of the Invention) The present invention has been made by focusing on the above-mentioned conventional problems. Silence data is forcibly added to the end of each piece of speech that has been cut off (speech that connects words or its encoded data; in the case of word-based speech, it is the silent part between words or its encoded data). The aim is to solve the above problem by subsequently performing ADPCM encoding.

(発明の構成及び作用) この発明を第6図に示すフローチヤートを用い
て説明する。
(Structure and operation of the invention) This invention will be explained using the flowchart shown in FIG.

この発明を実施するための装置は、第1図に示
す従来の装置と同様な構成である。まず連続音声
を入力端子1に入力し(ステツプ17)、分析部2
により分析してADPCM符号化し(ステツプ18)、
該結果を符号化データとして一旦サブメモリ部3
に格納し(ステツプ19)、該符号化データを中央
処理装置6によりPCM符号化データに再生し、
かつ該PCM符号化データから必要とする音声片
の符号化データを切り出し、該データの冗長部を
切断、除去し(ステツプ20)、残存する音声片の
符号化データの終端に無音データ(例えば分析部
内A/Dコンバータを12ビツトバイポーラ方式で
使用した場合、7FF(H))をあるサンプル数個以上
付加し(ステツプ21)、再び中央処理装置により
分析してADPCM符号化し(ステツプ22)、該
ADPCM符号化データを主メモリ部に格納する
(ステツプ23)。以下同様にして語単位で切り出し
及び格納を行なう。
The apparatus for carrying out this invention has a similar structure to the conventional apparatus shown in FIG. First, input continuous audio into input terminal 1 (step 17), and then
is analyzed and ADPCM encoded (step 18),
The result is temporarily stored in the sub memory unit 3 as encoded data.
(step 19), and the central processing unit 6 reproduces the encoded data into PCM encoded data.
Then, the encoded data of the necessary voice piece is cut out from the PCM encoded data, the redundant part of the data is cut and removed (step 20), and silent data (for example, analysis is added to the end of the encoded data of the remaining voice piece). When the internal A/D converter is used in a 12-bit bipolar system, a certain number or more samples of 7FF(H) are added (step 21), analyzed again by the central processing unit, and ADPCM encoded (step 22).
The ADPCM encoded data is stored in the main memory section (step 23). Thereafter, extraction and storage are performed word by word in the same manner.

第7図に音声片“B”の符号化データの終端に
無音データ7FF(H)を付加し(ステツプ21)
ADPCM符号化した時(ステツプ22)のデルタパ
ラメータΔpの波形29を示す。同図において、
音声片“B”28の符号化データの終端がデルタパ
ラメータΔp(B)=αであつた場合、デルタパラメ
ータΔp29の減少はADPCM方式の特徴として1
づつであるため、デルタパラメータΔpが最小値
かつ初期値である“0”になるまでにはαサンプ
ル分を要する。従つてサンプル周期をtsとすると
付加すべき無音データの長さTは T=α・ts であり、付加する無音データはα個要する。また
音声片の終端でデルタパラメータΔpが最大値
MAX(Δp)となる場合もありうるため、付加無
音区間Tを一定値とするためには、 TMAX(Δp)・ts とする必要がある。
In Figure 7, silent data 7FF(H) is added to the end of the encoded data of voice piece "B" (step 21).
A waveform 29 of the delta parameter Δp when ADPCM encoding is performed (step 22) is shown. In the same figure,
When the end of the encoded data of speech piece “B”28 is delta parameter Δp(B)=α, the decrease in delta parameter Δp29 is 1 as a feature of the ADPCM method.
Therefore, it takes α samples for the delta parameter Δp to reach the minimum value and initial value “0”. Therefore, if the sampling period is ts , the length T of the silent data to be added is T=α· ts , and α pieces of silent data are required. Also, the delta parameter Δp has its maximum value at the end of the voice piece.
MAX(Δp). Therefore, in order to keep the additional silent period T at a constant value, it is necessary to set TMAX(Δp)· ts .

第8図に、付加無音区間T=MAX(Δp)・tsと
し、音声片“B”及び“A”を合成した場合(ス
テツプ25〜27)の合成音声波形30及びデルタパ
ラメータΔpの波形31を示す。同図に示すよう
に、音声片“A”の始端32までにデルタパラメ
ータΔpを“0”とすることができるので、次の
音声素片“A”は完全に原音声波形に再生され
(a)、またゼロシフトの現象も発生しない(b)。
FIG. 8 shows a synthesized speech waveform 30 and a waveform 31 of delta parameter Δp when speech pieces “B” and “A” are synthesized (steps 25 to 27) with additional silent interval T=MAX(Δp)· ts . shows. As shown in the figure, the delta parameter Δp can be set to "0" by the start point 32 of the speech segment "A", so the next speech segment "A" is completely reproduced into the original speech waveform.
(a), and no zero shift phenomenon occurs (b).

具体例として、MAX(Δp)=48とし、サンプリ
ング周期tsを周波数f=6KHzとして ts=1/6=167(μs)とした場合、付加無音区間
Tは TMAX(Δp)・ts=48/6=8(ms) となるが、8ms程度の無音は無音として関知で
きないものであるから、このような処理を行つて
も不自然な音声に再生されることはない。また切
断、除去される冗長部は、通常100〜200ms程度
であり、従つて付加される無音部ははるかに少な
い。
As a specific example, if MAX (Δp) = 48, sampling period t s is frequency f = 6KHz, and t s = 1/6 = 167 (μs), the additional silent interval T is TMAX (Δp)・t s = 48/6=8 (ms) However, since approximately 8 ms of silence cannot be perceived as silence, even if such processing is performed, unnatural sound will not be reproduced. Further, the length of the redundant section to be cut and removed is usually about 100 to 200 ms, and therefore the number of silent sections to be added is much smaller.

尚、上記の実施例は原音声信号を一担ADPCM
方式によりADPCM符号化し、該ADPCM符号化
データをPCM符号化データに再生して、該PCM
符号化データの終端部に前述した切断、除去等の
一連の処理を行なう場合であるが、第2の実施例
として原音声をいきなりPCM符号化し、該PCM
符号化データに一連の処理を行なうこともでき
る。また上記第1の実施例は、原音声が連続音で
ある場合であるが、原音声が語単位の音声である
場合にも同様に実施できる。
In addition, in the above embodiment, the original audio signal is
The ADPCM encoded data is regenerated into PCM encoded data, and the PCM encoded data is encoded using the ADPCM encoding method.
This is a case where a series of processes such as cutting and removal described above are performed on the end of encoded data, but as a second embodiment, the original audio is suddenly PCM encoded, and the PCM
A series of processes can also be performed on encoded data. Furthermore, although the above first embodiment deals with the case where the original speech is a continuous sound, it can be similarly implemented when the original speech is a speech on a word-by-word basis.

(発明の効果) 以上説明したように、この発明によればその構
成をADPCM方式によりADPCM符号化データと
した後該ADPCM符号化データをPCM符号化デ
ータとし、又は原音声をPCM方式によりPCM符
号化データとし、これらのPCM符号化データの
終端部の冗長部分を切断、除去し、残存する
PCM符号化データの終端に所定数の無音符号を
付加し、該PCM符号をADPCM方式により符号
化することとしたため、音声片をどの時点で切り
出してもその符号化データ終端のデルタパラメー
タを円滑に初期化することができ、編集時におけ
る音声の合成部における波形の変形等の欠点が除
去され、原音声を完全に再生することができる。
(Effects of the Invention) As explained above, according to the present invention, the structure is converted into ADPCM encoded data using the ADPCM method, and then the ADPCM encoded data is converted into PCM encoded data, or original audio is converted into PCM encoded data using the PCM method. The redundant parts at the end of these PCM encoded data are cut and removed to remain.
Because we decided to add a predetermined number of silence codes to the end of PCM encoded data and encode the PCM code using the ADPCM method, the delta parameter at the end of the encoded data can be smoothly adjusted no matter what point a voice piece is cut out. It can be initialized, defects such as waveform deformation in the audio synthesis section during editing can be removed, and the original audio can be perfectly reproduced.

【図面の簡単な説明】[Brief explanation of drawings]

第1図は従来の又は本発明を実施するための音
声編集システムのブロツク図、第2図は従来の音
声編集を示すフローチヤートを示す図、第3図a
及びbはそれぞれ連続音声から切出した音声片終
端の波形及びデルタパラメータΔpの波形を示す
図、第4図a及びbはそれぞれ連続音声から切出
した音声片先端の波形及びデルタパラメータΔp
の波形を示す図、第5図a及びbはそれぞれ第3
図及び第4図に示す音声片を合成した場合の合成
波形及びデルタパラメータΔpの波形を示す図、
第6図はこの発明による音声編集を示すフローチ
ヤートを示す図、第7図a及びbはそれぞれこの
発明により音声片の終端に無音データを付加した
場合の音声波形及びデルタパラメータΔpの波形
を示す図、第8図a,bはそれぞれ第6図に示す
編集を行つた場合の編集合成波形及びデルタパラ
メータΔpの波形を示す図である。 1……音声入力端子、2……分析部、3……サ
ブメモリ部、4……メインメモリ部、5……合成
部、6……中央処理装置。
FIG. 1 is a block diagram of a conventional audio editing system or for implementing the present invention, FIG. 2 is a flowchart showing conventional audio editing, and FIG. 3 a
4 and b are diagrams showing the waveform at the end of a speech piece cut out from continuous speech and the waveform of delta parameter Δp, respectively, and FIGS. 4a and b show the waveform and delta parameter Δp at the tip of a speech piece cut out from continuous speech, respectively.
Figures 5a and 5b show the waveforms of the third
A diagram showing the synthesized waveform and the waveform of the delta parameter Δp when the speech pieces shown in FIGS.
FIG. 6 is a flowchart illustrating audio editing according to the present invention, and FIGS. 7a and 7b respectively show the audio waveform and the waveform of the delta parameter Δp when silence data is added to the end of a voice piece according to the present invention. FIGS. 8a and 8b are diagrams respectively showing the edited composite waveform and the waveform of the delta parameter Δp when the editing shown in FIG. 6 is performed. DESCRIPTION OF SYMBOLS 1...Audio input terminal, 2...Analysis section, 3...Sub memory section, 4...Main memory section, 5...Synthesizing section, 6...Central processing unit.

Claims (1)

【特許請求の範囲】[Claims] 1 原音声を適用差分パルス符号変調方式
(ADPCM方式)により符号化する音声符号化方
法において、原音声をADPCM方式により
ADPCM符号化データとした後該ADPCM符号化
データをPCM符号化データとし、又は原音声を
PCM方式によりPCM符号化データとし、上記い
ずれかの方法により得られたPCM符号化データ
の終端部の冗長部分を切断、除去し、残存する
PCM符号化データの終端にデルタパラメータΔp
の最大値Max(Δp)とサンプル同期周期tsとの積
Max(Δp)・tsで定まる長さの無音符号を付加し、
該PCM符号をADPCM方式により符号化するこ
とを特徴とする音声符号化方法。
1 In an audio encoding method that encodes the original audio using the applied differential pulse code modulation method (ADPCM method), the original audio is encoded using the ADPCM method.
After converting it into ADPCM encoded data, convert the ADPCM encoded data into PCM encoded data, or convert the original audio into PCM encoded data.
PCM coded data is generated using the PCM method, and the redundant portion at the end of the PCM coded data obtained by any of the above methods is cut off and removed to remain.
Delta parameter Δp at the end of PCM encoded data
The product of the maximum value Max (Δp) and the sample synchronization period ts
Add a silence code with a length determined by Max(Δp)・ts,
A voice encoding method characterized in that the PCM code is encoded using an ADPCM method.
JP57040735A 1982-03-17 1982-03-17 Voice coding Granted JPS58158693A (en)

Priority Applications (1)

Application Number Priority Date Filing Date Title
JP57040735A JPS58158693A (en) 1982-03-17 1982-03-17 Voice coding

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
JP57040735A JPS58158693A (en) 1982-03-17 1982-03-17 Voice coding

Publications (2)

Publication Number Publication Date
JPS58158693A JPS58158693A (en) 1983-09-20
JPH0378639B2 true JPH0378639B2 (en) 1991-12-16

Family

ID=12588886

Family Applications (1)

Application Number Title Priority Date Filing Date
JP57040735A Granted JPS58158693A (en) 1982-03-17 1982-03-17 Voice coding

Country Status (1)

Country Link
JP (1) JPS58158693A (en)

Families Citing this family (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPH0833742B2 (en) * 1986-07-17 1996-03-29 日本電気株式会社 Speech synthesis method
JPH0833743B2 (en) * 1986-11-07 1996-03-29 日本電気株式会社 Waveform synthesis method
JPH0833758B2 (en) * 1988-02-24 1996-03-29 日本電気株式会社 Speech synthesis method

Family Cites Families (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPS5173418A (en) * 1974-12-20 1976-06-25 Sanyo Electric Co Jikanjikuhenkansochino zatsuonyokuatsukairo
JPS5736943Y2 (en) * 1979-09-07 1982-08-14

Also Published As

Publication number Publication date
JPS58158693A (en) 1983-09-20

Similar Documents

Publication Publication Date Title
EP0380572B1 (en) Generating speech from digitally stored coarticulated speech segments
US5682502A (en) Syllable-beat-point synchronized rule-based speech synthesis from coded utterance-speed-independent phoneme combination parameters
JP2612868B2 (en) Voice utterance speed conversion method
JP4867076B2 (en) Compression unit creation apparatus for speech synthesis, speech rule synthesis apparatus, and method used therefor
JP2005018037A (en) Device and method for speech synthesis and program
JPS58158693A (en) Voice coding
JPH05303399A (en) Audio time axis companding device
JP3342310B2 (en) Audio decoding device
JP2001282246A (en) Waveform data time expansion / compression device
JP3241582B2 (en) Prosody control device and method
JP2648138B2 (en) How to compress audio patterns
US6418406B1 (en) Synthesis of high-pitched sounds
JPS6295595A (en) Voice response method
JPS5948399B2 (en) Speech synthesis method
JP2861005B2 (en) Audio storage and playback device
JPH01197793A (en) Speech synthesizer
JPH0756589A (en) Speech synthesis method
JP2527393Y2 (en) Speech synthesizer
JP2007108450A (en) Voice reproducing device, voice distributing device, voice distribution system, voice reproducing method, voice distributing method, and program
JPS58196598A (en) Rule type voice synthesizer
JPH01191900A (en) Voice response device
JP2007256303A (en) Voice compression system
JPS63178300A (en) Voice encoder
JPS60256987A (en) Time axis converter of acoustic signal
JPH0355840B2 (en)