JPS60500A - Speech analysis and synthesis method - Google Patents

Speech analysis and synthesis method

Info

Publication number
JPS60500A
JPS60500A JP58108766A JP10876683A JPS60500A JP S60500 A JPS60500 A JP S60500A JP 58108766 A JP58108766 A JP 58108766A JP 10876683 A JP10876683 A JP 10876683A JP S60500 A JPS60500 A JP S60500A
Authority
JP
Japan
Prior art keywords
speech
sound source
voiced
analysis
data
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Granted
Application number
JP58108766A
Other languages
Japanese (ja)
Other versions
JPH0344319B2 (en
Inventor
平岡 省二
謙二 加賀
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Panasonic Holdings Corp
Original Assignee
Matsushita Electric Industrial Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Matsushita Electric Industrial Co Ltd filed Critical Matsushita Electric Industrial Co Ltd
Priority to JP58108766A priority Critical patent/JPS60500A/en
Publication of JPS60500A publication Critical patent/JPS60500A/en
Publication of JPH0344319B2 publication Critical patent/JPH0344319B2/ja
Granted legal-status Critical Current

Links

Abstract

(57)【要約】本公報は電子出願前の出願データであるた
め要約のデータは記録されません。
(57) [Summary] This bulletin contains application data before electronic filing, so abstract data is not recorded.

Description

【発明の詳細な説明】 産業上の利用分野 本発明は音声信号をディジタル化した後、分析し、分析
して得られたパラメー′夕を低データレートで伝送また
は記憶し、再び音声信号に復元する音声分析合成装置に
関するものである。−従来例の構成とその問題点 通常、音声分析合成装置においては入力された音声から
分析器で声道パラメータと音源ノ(ラメータを抽出し、
各パラメータにコード化等のデータレート低減のための
処理を施し、伝送路または記憶素子へ送出し、これを合
成器で音声に再合成する。この場合の方式として従来、
音源)くラメータの違いにより、(1)分析時に抽出さ
れる分析残差波形をそのまま或いは差分等の処理でデー
タ圧縮して伝送または記憶する方式、(2り音声の大き
さを表わす振幅パラメータ、声の高さを表わすピッチ・
(ラメータおよび有声無声切換情報を抽出して伝送また
は記憶する方式、(3)音声データを記憶する例では前
記2の方式の各ノ(ラメータと話者の分析残差波形の一
部を記憶しておき再合成する方式がある0 (1)の方式では話者の声質をよく再合成できる反面伝
送または記憶時のデータレートが高いという欠点がある
0 し)の方式では(1)の方式と反対にデータレートは低
いが話者の違いに関係なく一定の有声音源データを使用
するため声のつや等の個性的特徴が失なわれた合成音と
なる欠点がある。
DETAILED DESCRIPTION OF THE INVENTION Field of Industrial Application The present invention digitizes an audio signal, analyzes it, transmits or stores the parameters obtained by the analysis at a low data rate, and restores the audio signal again. This invention relates to a speech analysis and synthesis device. - Conventional configuration and its problems Normally, in a speech analysis and synthesis device, vocal tract parameters and sound source parameters are extracted from input speech using an analyzer.
Each parameter is subjected to processing to reduce the data rate, such as encoding, and sent to a transmission line or storage element, and then resynthesized into speech by a synthesizer. Conventionally, the method in this case is
(1) A method in which the analysis residual waveform extracted at the time of analysis is transmitted or stored as it is, or compressed by processing such as a difference, and (2) an amplitude parameter representing the loudness of the sound; Pitch, which indicates the pitch of the voice
(3) In the example of storing voice data, each method of the above 2 (method of extracting and transmitting or storing the parameter and voiced/unvoiced switching information) There is a method that resynthesizes the speaker's voice quality well, but the disadvantage is that the data rate during transmission or storage is high. On the other hand, although the data rate is low, it uses voiced sound source data that is constant regardless of the speaker, so it has the disadvantage of resulting in a synthesized sound that loses individual characteristics such as the luster of the voice.

(′4の方式は(1)および(ロ)の方式の中間的特徴
をもつが話者が一定でない実時間分析合成の例では適さ
ない。
(Method '4' has intermediate characteristics between methods (1) and (b), but is not suitable for real-time analysis and synthesis in which the speaker is not constant.

発明の目的 本発明は従来の技術の上記欠点を改善するもので、その
目的は音声の実時間伝送における情報量を極端に増大す
ることなく、話者の個性的な特徴を含んだ音声を再合成
するだめの伝送方式を提供するものである。
OBJECT OF THE INVENTION The present invention aims to improve the above-mentioned drawbacks of the prior art, and its purpose is to reproduce speech that includes the individual characteristics of the speaker without significantly increasing the amount of information in real-time transmission of speech. This provides a transmission method that does not involve combining.

発明の構成 本発明は音声をディジタル信号に変換し線形予測法など
で声道、振幅、ピッチ、有声無声判定の各パラメータを
抽出するパラメータ分析部の他に分析残差波形の一部を
有声駆動音源データとして切出す回路を分析器に有し、
パラメータと有声駆動音源データを低データレートで伝
送する伝送路を有し、有声駆動音源データが伝送されて
くる以前には予め定めた波形を使用し有声駆動音源デー
タが伝送されてから以降は新しい駆動音源データを使用
して音声合成する合成器を有する音声分析合成装置であ
る0 実施例の説明 以下、本発明の実施例を詳細に説明する0第1図は本発
明による音声分析合成装置の構成を示すブロック画であ
る0第1図において、1はマイクロフォン等の収音器で
伝送する音声を収音しアナログ信号に変換して、音声分
析器2に与える。音声分析器2はアナログ信号を8に〜
10KHz程度でサンプリングしディジタル信号に変換
した後6〜20 ms程度の区間(フレームと呼ぶ)毎
に線形予測分析等によシ声道パラメータと音源ノくラメ
ータをめ、このパラメータを符号化等によシさらに帯域
圧縮し、伝送路3に送出、するO伝送路3は通常の電話
回線のように実時間で伝送される系のほか、書込可能な
メモリ素子(RAM)等のような記憶媒体であってもよ
い0圧縮ノくラメータを受信した音声合成器4では音声
分析器2で行なった帯域圧縮の逆の操作を行ない音声信
号を復元する。この復元した音声信垂をスピーカ6に与
え音声再生する。
Structure of the Invention The present invention includes a parameter analysis section that converts speech into a digital signal and extracts each parameter of the vocal tract, amplitude, pitch, and voiced/unvoiced judgment using a linear prediction method, etc., as well as a parameter analysis section that converts speech into a digital signal and extracts each parameter of vocal tract, amplitude, pitch, and voiced/unvoiced judgment. The analyzer has a circuit that extracts it as sound source data,
It has a transmission line that transmits parameters and voiced driving sound source data at a low data rate, and uses a predetermined waveform before the voiced driving sound source data is transmitted, and a new waveform after the voiced driving sound source data is transmitted. This is a speech analysis and synthesis device having a synthesizer that synthesizes speech using driving sound source data.Description of EmbodimentsHereinafter, embodiments of the present invention will be described in detail.0Figure 1 shows a speech analysis and synthesis device according to the present invention. In FIG. 1, which is a block diagram showing the configuration, numeral 1 collects the transmitted sound with a sound collector such as a microphone, converts it into an analog signal, and supplies it to the sound analyzer 2. Speech analyzer 2 converts the analog signal into 8~
After sampling at approximately 10 KHz and converting to a digital signal, vocal tract parameters and sound source parameters are determined by linear predictive analysis for each interval of approximately 6 to 20 ms (called a frame), and these parameters are used for encoding, etc. The band is further compressed and sent to the transmission line 3.The transmission line 3 is not only a system that transmits in real time like a normal telephone line, but also a storage system such as a writable memory element (RAM). The speech synthesizer 4 that receives the zero compression parameter, which may be a medium, performs the inverse operation of the band compression performed by the speech analyzer 2 to restore the speech signal. The restored audio signal is given to the speaker 6 for audio reproduction.

帯域圧縮技術として本実施例では線形予測分析法の一つ
であるPAftCOR法を用いている。
As a band compression technique, this embodiment uses the PAftCOR method, which is one of the linear predictive analysis methods.

PARCOR法を用いた音声分析器については後述する
A speech analyzer using the PARCOR method will be described later.

第2図に示した音声信号をPARCOR分析、パラメー
タ伝送、PARCOR合成する際、伝送路の容量はパラ
メータの単位時間当りの最大データレートで定まるが、
実際の伝送では第2図において区間(→、(4)に比し
て区間(1) 9 (3) 9 (5)のような無音区
間では転送データレートは極端に低い。そこで、本発明
では区間(2)や(4)で分析して得られた残差波形の
うち定常的な母音区間の一部を区間0)や(句で伝送し
合成器で有声駆動音源として使用する。この残差波形に
はパラメータで表わされない話者の個人性が含まれてい
るので個人性豊かな音声が合成できる。有声駆動音源デ
ータは通常1ピッチ周期以下のデータ列であるが本実施
例では8ビット×31点で構成しているため248ビツ
トを無音区間に転送する必要があり、今、2400ビッ
ト/秒の伝送路を使用すれば、この残差データの伝送に
約1o○ミリ秒所要するが、通常の発声では数百ミ+7
秒程度の無音区間はよく存在するので十分伝送できる。
When performing PARCOR analysis, parameter transmission, and PARCOR synthesis of the audio signal shown in Figure 2, the capacity of the transmission path is determined by the maximum data rate per unit time of the parameters.
In actual transmission, the transfer data rate is extremely low in silent sections such as sections (1) 9 (3) 9 (5) compared to sections (→, (4)) in FIG. Of the residual waveforms obtained by analyzing sections (2) and (4), a part of the stationary vowel section is transmitted in sections 0) and (phrases) and used as a voiced driving sound source in the synthesizer. Since the difference waveform includes the individuality of the speaker that is not expressed by parameters, speech with rich individuality can be synthesized.Voiced driving sound source data is normally a data string of one pitch period or less, but in this example, it is 8 pitch periods or less. Since it is composed of bits x 31 points, it is necessary to transfer 248 bits to the silent section, and if a 2400 bit/second transmission line is used, it will take approximately 100 milliseconds to transmit this residual data. However, in normal vocalization, it is several hundred mi+7
There are often silent intervals of about seconds, so it can be transmitted sufficiently.

なお有声駆動音源データは差分法等でデータ圧縮し短か
い無音区間で伝送することもできる〇一方、合成器は区
間(2)のような発声開始時点ではまだ駆動音源データ
が伝送されていないので予め定めたインノ<ルス波形等
を有声駆動音源データとして使用する0 第3図は第1図中2に相当する音声分析器の構成を示す
ブロック甲である。21は音声信号をサンプリンブレデ
ィジタル信号に変換するムD変換器でディジタル信号は
PA、ROOR分析器22、ピッチ抽出器23、有声無
声判定器24、無音区間検出器26に送られるoP、A
RCOf’1分析器22で得られた残差信号は残差切出
回路26で残差信号の一部を切出され一時蓄わ見られる
Oまた振幅決定回路27で振幅パラメータがめられる。
Note that voiced driving sound source data can be compressed using a differential method or the like and transmitted in a short silent section.On the other hand, the synthesizer does not transmit driving sound source data yet at the start of vocalization, such as in section (2). Therefore, a predetermined inno<lus waveform or the like is used as the voiced drive sound source data.0 FIG. 3 is a block A showing the configuration of a speech analyzer corresponding to 2 in FIG. 1. 21 is a mu-D converter that converts the audio signal into a sampled digital signal; the digital signal is sent to the PA, the ROOR analyzer 22, the pitch extractor 23, the voiced/unvoiced determiner 24, and the silent section detector 26;
A part of the residual signal obtained by the RCOf'1 analyzer 22 is extracted in a residual extraction circuit 26 and temporarily stored for viewing.An amplitude parameter is determined in an amplitude determination circuit 27.

PARCOFt分析器22、ピッチ抽出器23、有声無
声判定器24、無音区間検出器26および振幅決定回路
27でめられたパラメータは符号器28で符号化され、
切換器29を経である時間区間(フレーム)を代表する
パラメータ値として伝送路3に送出される。無音区間検
出器25で無音区間が検出されると切換器29は反転し
残差切出 −回路26で切出された残差波形の一部が合
成器の有声音源データとして伝送される。
The parameters determined by the PARCOFT analyzer 22, pitch extractor 23, voiced/unvoiced determiner 24, silent section detector 26, and amplitude determination circuit 27 are encoded by an encoder 28,
It is sent to the transmission line 3 via the switch 29 as a parameter value representative of a certain time interval (frame). When the silent interval detector 25 detects a silent interval, the switch 29 is inverted, and a portion of the residual waveform extracted by the residual extraction circuit 26 is transmitted as voiced sound source data to the synthesizer.

第4図は伝送されてくるパラメータおよび有声音源デー
タを受けて音声信号を合成するPARCOR方式音声合
成器の構成を示すブロック図であり、第1図の4に対応
する。伝送されてくるデータは選択器41で2種類に分
離され、パラメータはパラメータメモリ42に蓄わえら
れ、有声音源データは音源メモリ43に蓄わえられる。
FIG. 4 is a block diagram showing the configuration of a PARCOR type speech synthesizer that receives transmitted parameters and voiced sound source data and synthesizes a speech signal, and corresponds to 4 in FIG. 1. The transmitted data is separated into two types by a selector 41, parameters are stored in a parameter memory 42, and voiced sound source data is stored in a sound source memory 43.

電源投入直後および長時間の無音区間を検出した時は前
記の予め定めたインパルス波形等の有声音源データが音
源メモリ43に自動的にセットされ、選択器41より新
しい有声音源データがセットされるまで保持される。4
4は無声音源発生器で、有声。
Immediately after the power is turned on or when a long period of silence is detected, voiced sound source data such as the predetermined impulse waveform is automatically set in the sound source memory 43 until new voiced sound source data is set by the selector 41. Retained. 4
4 is a voiceless sound source generator and is voiced.

無声選択器46で音源メモリ43または無声音源発生器
44のいずれかのデータが選択され、ツク゛ラメータメ
モリ42内のノ(ラメータとともにディジタルフィルタ
46で演算され、その結果力より人変換器47でアナロ
グ信号に変換されて音声信号となシ、増幅器48で増幅
されてスピーカ6へ供給される。
The data of either the sound source memory 43 or the unvoiced sound source generator 44 is selected by the silent selector 46, and the data in the parameter memory 42 is calculated by the digital filter 46. The signal is converted into an audio signal, amplified by an amplifier 48, and supplied to the speaker 6.

発明の効果 以上のように、本発明は実時間で音声波形を分析、伝送
、合成する際に定常の有声音区間を分析して得られた残
差波形の一部を有声駆動音源データとして無音中の低デ
ータレートの区間に伝送するようにした音声分析合成装
置で、音声分析時に抽出される残差波形の一部を、)く
ラメータを伝送しない無音区間に伝送することにより、
伝送路の最大転送データ容量を増大させることなく、声
のつやや丸やかさ等といった話者特有の声質層75為な
音声を合成することができる。
Effects of the Invention As described above, the present invention analyzes a stationary voiced sound section when analyzing, transmitting, and synthesizing speech waveforms in real time, and uses a portion of the residual waveform obtained as voiced driving sound source data to generate silent sound. By transmitting a part of the residual waveform extracted during speech analysis to the silent section where the parameter is not transmitted, the speech analysis and synthesis device transmits data to the middle low data rate section.
It is possible to synthesize speech based on the voice quality layer 75 unique to the speaker, such as the luster and roundness of the voice, without increasing the maximum transfer data capacity of the transmission path.

【図面の簡単な説明】[Brief explanation of drawings]

第1図は本発明による音声分析合成装置の構成を示すブ
ロック図、第2図は音声波形と伝送するデータの時間関
係を示した波形図、第3図は本発明による音声分析合成
装置の音声分析器部の構成を示すブロック図、第4図は
本発明による音声合成分析装置の音声合成器部の構成を
示すブロック図である。 1・・・・・・収音器、2・・・・・・音声分析器、3
・・・・・・伝送路、4・・・・・・音声合成器、6・
・パ・・・スピーカ、21・・・・・・AD変換器、2
2・・・・・・PAR(30R分析器、23・・・・・
・ピッチ抽出器、24・・・・・・有声無声判定器、2
5・・・・・・無声区間検出器、26・・・・・・残差
切出回路、27・・・・・・振幅決定回路、28・・・
・・・符号器、29・・・・・・切換器、41・・・・
・・選択器、42・・・・・・)・フメータメモリ、4
3・・・・・・音源メモリ、44・・・・・・無声音源
発生器、46・・・・・・ディジタルフィルタ、47・
・・・・・Dム変換器。 代理人の氏名 弁理士 中 尾 敏 男 ほか1名68
FIG. 1 is a block diagram showing the configuration of the speech analysis and synthesis device according to the present invention, FIG. 2 is a waveform diagram showing the time relationship between speech waveforms and transmitted data, and FIG. 3 is the speech analysis and synthesis device according to the invention. A block diagram showing the structure of the analyzer section. FIG. 4 is a block diagram showing the structure of the speech synthesizer section of the speech synthesis and analysis apparatus according to the present invention. 1...Sound collector, 2...Speech analyzer, 3
...Transmission line, 4...Speech synthesizer, 6.
・Pa...Speaker, 21...AD converter, 2
2...PAR (30R analyzer, 23...
・Pitch extractor, 24... Voiced/unvoiced determiner, 2
5... Silent section detector, 26... Residual extraction circuit, 27... Amplitude determining circuit, 28...
...Encoder, 29...Switcher, 41...
・・Selector, 42・・・・)・Fumemeter memory, 4
3... Sound source memory, 44... Silent sound source generator, 46... Digital filter, 47...
...DM converter. Name of agent: Patent attorney Toshio Nakao and 1 other person68
1

Claims (2)

【特許請求の範囲】[Claims] (1)定常の有声音区間を分析して得られた残差波形の
一部を有声駆動音源データとして無音区間に伝送し音声
合成することを特徴とする音声分析合成装置。
(1) A speech analysis and synthesis device characterized in that a part of the residual waveform obtained by analyzing a stationary voiced sound section is transmitted as voiced driving sound source data to a silent section and synthesized into speech.
(2) 残差波形の一部が抽出される以前゛は予め定め
た波形を有声駆動音源デ゛−夕として音声合成する特許
請求の範囲第1項記載の音声分析合成装置0
(2) The speech analysis and synthesis device 0 according to claim 1, which synthesizes speech using a predetermined waveform as a voiced driving sound source data before a part of the residual waveform is extracted.
JP58108766A 1983-06-16 1983-06-16 Speech analysis and synthesis method Granted JPS60500A (en)

Priority Applications (1)

Application Number Priority Date Filing Date Title
JP58108766A JPS60500A (en) 1983-06-16 1983-06-16 Speech analysis and synthesis method

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
JP58108766A JPS60500A (en) 1983-06-16 1983-06-16 Speech analysis and synthesis method

Publications (2)

Publication Number Publication Date
JPS60500A true JPS60500A (en) 1985-01-05
JPH0344319B2 JPH0344319B2 (en) 1991-07-05

Family

ID=14492944

Family Applications (1)

Application Number Title Priority Date Filing Date
JP58108766A Granted JPS60500A (en) 1983-06-16 1983-06-16 Speech analysis and synthesis method

Country Status (1)

Country Link
JP (1) JPS60500A (en)

Cited By (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPS61166600A (en) * 1985-01-19 1986-07-28 三洋電機株式会社 Voice snthesizer

Cited By (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPS61166600A (en) * 1985-01-19 1986-07-28 三洋電機株式会社 Voice snthesizer

Also Published As

Publication number Publication date
JPH0344319B2 (en) 1991-07-05

Similar Documents

Publication Publication Date Title
EP0380572B1 (en) Generating speech from digitally stored coarticulated speech segments
JP3607450B2 (en) Audio information classification device
JPH0344319B2 (en)
JPH10133678A (en) Audio playback device
JPH0376480B2 (en)
JPH0854895A (en) Playback device
KR100359988B1 (en) real-time speaking rate conversion system
JP2860991B2 (en) Audio storage and playback device
JPH0772896A (en) Device for compressing/expanding sound
JPS62102294A (en) Voice coding system
JPS5912479A (en) Pronuntiation practicing apparatus
JPH09146587A (en) Speech speed changer
KR100194659B1 (en) Voice recording method of digital recorder
JP2535809B2 (en) Linear predictive speech analysis and synthesis device
JPS58113992A (en) Voice signal compression system
JPS635398A (en) Voice analysis system
JPH0376479B2 (en)
JPH03160500A (en) Speech synthesizer
JPH01261700A (en) Voice coding system
JPH02245800A (en) Voice reproduction system
Linggard Neural networks for speech processing: An introduction
JPS61128299A (en) Voice analysis/analytic synthesization system
JPS5926795A (en) Speech unit
JPS58171095A (en) Noise suppression system
JPS5915300A (en) Audio recording and playback device