JPH01120599A - Voice synthesization system - Google Patents
Voice synthesization systemInfo
- Publication number
- JPH01120599A JPH01120599A JP62278541A JP27854187A JPH01120599A JP H01120599 A JPH01120599 A JP H01120599A JP 62278541 A JP62278541 A JP 62278541A JP 27854187 A JP27854187 A JP 27854187A JP H01120599 A JPH01120599 A JP H01120599A
- Authority
- JP
- Japan
- Prior art keywords
- syllable
- vowel
- long
- long vowel
- voice
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Granted
Links
Abstract
Description
【発明の詳細な説明】
[概 要コ
本発明は人間が発声した音節単位の音声を短い時間間隔
ごとに分析して、これを該音声のパラメータ時系列デー
タとして、音節ごとに蓄積−しておいて、これらのパラ
メータ時系列データから成る音声を結合することにより
、任意の音声を合成する音声合成方式に関し、
合成すべき音声が長音を含む場合の自然らしさと了解度
の高い合成音声を得る手段を提供することを目的とし、
長音を合成する場合に、長音の直前の音節のCI:音部
を伸張して長音部として用いることにより構成する。[Detailed Description of the Invention] [Overview] The present invention analyzes the syllable-based sounds uttered by humans at short time intervals, and stores this as parameter time-series data for each syllable. Regarding a speech synthesis method that synthesizes arbitrary speech by combining speech composed of these parameter time series data, we aim to obtain synthesized speech with high naturalness and intelligibility when the speech to be synthesized includes long sounds. The purpose of this method is to provide a means for synthesizing long sounds by stretching the CI: syllable part of the syllable immediately before the long sound and using it as the long sound part.
[産業上の利用分野]
本発明は音声の合成方式の内、実際に人間が発声した音
節単位の音声を短い時間間隔ごとに分析して、これを該
音声のパラメータ時系列データとして、音節ごとに蓄積
しておいて、これらのパラメータ時系列データから成る
音声を結合することにより、任意の音声を合成する音声
合成方式に関し、特に、長音を伴う場合の音声の自然性
を向上せしめ得る音声合成方式に係る。[Industrial Application Field] Among the speech synthesis methods, the present invention analyzes syllable-based speech actually uttered by humans at short time intervals, and uses this as parameter time-series data of the speech for each syllable. This method relates to a speech synthesis method that synthesizes arbitrary speech by combining speech composed of these parameter time-series data stored in Regarding the method.
[従来の技術]
人工的に音声を合成する方式を大別すると、■録音編集
方式、■分析合成方式、■純粋合成方式の3方式に分け
られる。[Prior Art] Methods for artificially synthesizing speech can be roughly divided into three methods: ■ Recording/editing method, ■ Analysis/synthesis method, and ■ Pure synthesis method.
これらの内、■の録音編集方式は予め録音した人間の音
声波形をつなぎ合わせて合成するもので装置や制御が比
較的簡単であり、音質も良好であるという長所もあるが
情報量が非常に多いため、大容量の記憶装置を必要とす
る欠点がある。Of these, the recording/editing method (■) combines pre-recorded human voice waveforms and synthesizes them, and has the advantage of relatively simple equipment and control, and good sound quality, but the amount of information is very large. Since there are a lot of data, there is a drawback that a large capacity storage device is required.
また、■の純粋合成方式は文字等から一定の合成規則を
用いて音声を合成する方式で、人間の音声を予め分析す
る等の必要はないが、自然性の良好な音声を作り出すた
めには、前記合成規則が非常に複雑なものとなり、その
実現は容易ではない。In addition, the pure synthesis method (■) synthesizes speech from characters, etc. using certain synthesis rules, and there is no need to analyze human speech in advance. , the above-mentioned synthesis rule becomes very complicated, and its implementation is not easy.
これらに対して、■の分析合成方式は予め人の音声を分
析して、該音声をその特徴を表すパラメータに変換して
記憶しておいて、これを元に音声を合成する方式であっ
て、この方式にお−いては、音声を波形としてではなく
、音声をその調音によるスベクタルの形状を表すパラメ
ータと音源の状態を表すパラメータとに変換して、それ
らの情報を圧縮して記憶することができるので、記憶容
量が少なくて済む上、音質も比較的良好なものが得られ
るから、近年、この方式を用いた音声合成方式が背反し
つつある。On the other hand, the analysis-synthesis method (2) analyzes human speech in advance, converts the speech into parameters representing its characteristics, stores them, and synthesizes speech based on these parameters. In this method, instead of converting the audio into waveforms, the audio is converted into parameters that represent the shape of the subectals resulting from its articulation and parameters that represent the state of the sound source, and these information is compressed and stored. In recent years, speech synthesis methods using this method have been losing ground, as it requires less storage capacity and provides relatively good sound quality.
このような方式にもとづいて人間の発声した音節単位の
音声(例えば“ア”イ“つ”・・・・・・等)を分析し
て置き、それらの音節を結合して声の高さをM御するこ
とで任意の音声を合成する方式が広く用いられている。Based on this method, human utterances are analyzed in syllable units (for example, "a", "i", "tsu", etc.), and these syllables are combined to determine the pitch of the voice. A method of synthesizing arbitrary speech by controlling M is widely used.
[発明が解決しようとする問題点]
上述したような従来の分析合成方式にもとづく任意語の
音声合成方式において、長音は同種の母音連鎖と同様の
方法で合成されていた。[Problems to be Solved by the Invention] In the speech synthesis method for arbitrary words based on the conventional analysis and synthesis method as described above, long sounds are synthesized in the same manner as vowel chains of the same type.
例えば、“キー”(Key)も“キイ″(奇異)も全く
同様に音節“キ”の後に音節“イ°°を接続(同種母音
連鎖)することにより合成していた。For example, "Key" and "Kii" (strange) were synthesized in exactly the same way by connecting the syllable "i" after the syllable "ki" (homogeneous vowel chain).
そのなめ、長音部の音声が不自然であるばかりでなく、
長音と同種母音連鎖との区別が付かないという問題点が
あった。。Not only is the sound of the long note unnatural,
There was a problem in that it was not possible to distinguish between long sounds and homologous vowel chains. .
本発明はこのような従来の問題点に鑑み、長音の発声が
自然であると共に、長音と同種母音連鎖との区別を明確
にすることのできる任意語の音声合成方式を提供するこ
とを目的としている。In view of these conventional problems, it is an object of the present invention to provide a speech synthesis method for arbitrary words that allows the pronunciation of long sounds to be natural and that also makes it possible to clearly distinguish between long sounds and homologous vowel chains. There is.
[問題点を解決するための手段]
本発明によれば、上述の目的は前記特許請求の範囲に記
載した手段により達成される。すなわち、本発明は、人
間が発声した音節単位の音声を短い時間間隔ごとに分析
して、これを該音声のパラメータ時系列データとして、
音節ごとに蓄積しておいて、これらのパラメータ時系列
データから成る音声を結合することにより、任意の音声
を合成する音声合成方式はおいて、長音を合成する場合
に、長音の直前の音節の母音部を伸張して長音部として
用いる音声合成方式[]
第1図は本発明の詳細な説明する図であって、(a)は
“奇異”の音声合成について示しており、(b)は“キ
ー”の音声合成について示したもので、1は音節“キ”
、2は音節“イ”、3は伸張部を表している。[Means for Solving the Problems] According to the present invention, the above objects are achieved by the means described in the claims. That is, the present invention analyzes syllable-based sounds uttered by humans at short time intervals, and uses this as parameter time-series data of the sounds.
In addition to the speech synthesis method that synthesizes arbitrary speech by combining the sounds made up of these parameter time series data that are accumulated for each syllable, when synthesizing a long sound, the vowel of the syllable immediately before the long sound is used. Speech synthesis method that expands the part and uses it as a long part [ ] Figure 1 is a diagram explaining the present invention in detail, (a) shows "unusual" speech synthesis, and (b) shows " 1 is the syllable “ki”.
, 2 represents the syllable "i", and 3 represents the extended part.
従来は、“奇異”、“キー”共に(a)で示す方法で谷
成していたが、本発明によれば、“キーパの場合に(b
)に示す方法によって音声合成を行なうことにより、こ
のような長音と (a)の場合のような同種母音連鎖と
の区別が明確に間き分けられる合成音声を得ることがで
きる。Conventionally, both "oddness" and "key" were determined by the method shown in (a), but according to the present invention, in the case of "keeper", (b)
By performing speech synthesis using the method shown in (a), it is possible to obtain synthesized speech that clearly distinguishes between such long sounds and homogeneous vowel chains as in case (a).
[実施例]
第2図は本発明の一実施例の機能ブロック図であって、
4は音節読出部、5゛は音節格納部、6は音節結合部、
7は時間長設定部、8はピッチパターン設定部、9は波
形合成部、10は長音検出部を表している。[Embodiment] FIG. 2 is a functional block diagram of an embodiment of the present invention,
4 is a syllable reading part, 5 is a syllable storing part, 6 is a syllable combining part,
Reference numeral 7 represents a time length setting section, 8 a pitch pattern setting section, 9 a waveform synthesis section, and 10 a long tone detection section.
同図において、音声合成すべき入力文字列が入力される
と、音節読出し部4は音節格納部5から必要な音節の音
声パラメータを読み出して音節結合部6へ送る。In the figure, when an input character string to be synthesized into speech is input, a syllable reading section 4 reads out the speech parameters of the necessary syllables from a syllable storage section 5 and sends them to a syllable combining section 6.
一方、時間長設定部7は前記入力文字列から各音節の時
間長を設定する
また、長音検出部10は入力文字列を監視していて、長
音が検出された場合にはその直前の音節の母音を伸張す
るように音節結合部6に連絡する。On the other hand, a time length setting unit 7 sets the time length of each syllable from the input character string.A long sound detection unit 10 monitors the input character string, and when a long sound is detected, the syllable immediately before it. The syllable combining unit 6 is contacted to stretch the vowel.
該音節結合部6は、前記時間長設定部7が設定した音節
の時間長に従って、音声パラメータを結合する。The syllable combining unit 6 combines voice parameters according to the syllable time length set by the time length setting unit 7.
このとき、読み出されたパラメータと設定された時間長
に不一致があれば母音部時間長を伸縮することによって
調整を行なう、そして、同時に長音部に係る母音伸張を
行ない、結合した音声パラメータを波形合成部9へ送る
。At this time, if there is a discrepancy between the read parameters and the set time length, adjustments are made by expanding or contracting the vowel part time length. At the same time, the vowel expansion related to the long part is performed, and the combined voice parameters are converted into a waveform. Send it to the synthesis section 9.
波形合成部9は該音声パラメータとピッチパターン設定
部8で設定されたピッチパターンとによって音声波形を
合成する。The waveform synthesis section 9 synthesizes a speech waveform using the speech parameters and the pitch pattern set by the pitch pattern setting section 8.
[発明の効果1
以上、説明したように、本発明によれば、長音を有する
合成音の自然性を増して音質を向上せしめ得る利点を有
する。[Effect of the Invention 1] As described above, the present invention has the advantage of increasing the naturalness of synthesized sounds having long sounds and improving the sound quality.
そして従来区別できなかった長音く例えば“ア一″)と
同一母音連鎖(“アア“)あるいは同種母音連鎖との区
別が明確になる効果があるIt also has the effect of making it clearer to distinguish between long sounds (for example, “a-ichi”), which could not be distinguished in the past, and the same vowel chain (“aa”) or the same vowel chain.
第1図は本発明の詳細な説明する図、第2図は本発明の
一実施例の機能ブロック図である。
1・・・・・・音節“キ”、2・・・・・・音節“イ”
、3・・・・・・伸張部、4・・・・・・音節読出部、
5・・・・・・音節格納部、6・・・・・・音節結合部
、7・・・・・・時間長設定部、8・・・・・・ピッチ
パターン設定部、9・・・・・・波形合成部、10・・
・・・・長音検出部
へFIG. 1 is a diagram explaining the present invention in detail, and FIG. 2 is a functional block diagram of an embodiment of the present invention. 1... syllable "ki", 2... syllable "i"
, 3...Extension unit, 4...Syllable reading unit,
5...Syllable storage section, 6...Syllable combining section, 7...Time length setting section, 8...Pitch pattern setting section, 9... ...Waveform synthesis section, 10...
...to long sound detection section
Claims (1)
析して、これを該音声のパラメータ時系列データとして
、音節ごとに蓄積しておいて、これらのパラメータ時系
列データから成る音声を結合することにより、任意の音
声を合成する音声合成方式において、 長音を合成する場合に、長音の直前の音節の母音部を伸
張して長音部として用いることを特徴とする音声合成方
式。[Claims] The syllable-based sounds uttered by humans are analyzed at short time intervals, and this is stored as parameter time-series data for each syllable, and these parameter time-series data are A speech synthesis method that synthesizes arbitrary speech by combining sounds consisting of , and is characterized in that when synthesizing a long sound, the vowel part of the syllable immediately before the long sound is stretched and used as the long sound part. method.
Priority Applications (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP62278541A JPH0738113B2 (en) | 1987-11-04 | 1987-11-04 | Speech synthesizer |
Applications Claiming Priority (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP62278541A JPH0738113B2 (en) | 1987-11-04 | 1987-11-04 | Speech synthesizer |
Publications (2)
| Publication Number | Publication Date |
|---|---|
| JPH01120599A true JPH01120599A (en) | 1989-05-12 |
| JPH0738113B2 JPH0738113B2 (en) | 1995-04-26 |
Family
ID=17598700
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| JP62278541A Expired - Lifetime JPH0738113B2 (en) | 1987-11-04 | 1987-11-04 | Speech synthesizer |
Country Status (1)
| Country | Link |
|---|---|
| JP (1) | JPH0738113B2 (en) |
Citations (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JPS59177598A (en) * | 1983-03-28 | 1984-10-08 | 三洋電機株式会社 | Voice synthesizer |
-
1987
- 1987-11-04 JP JP62278541A patent/JPH0738113B2/en not_active Expired - Lifetime
Patent Citations (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JPS59177598A (en) * | 1983-03-28 | 1984-10-08 | 三洋電機株式会社 | Voice synthesizer |
Also Published As
| Publication number | Publication date |
|---|---|
| JPH0738113B2 (en) | 1995-04-26 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| JPH031200A (en) | Regulation type voice synthesizing device | |
| JPS62160495A (en) | Voice synthesization system | |
| JP5360489B2 (en) | Phoneme code converter and speech synthesizer | |
| JPH01120599A (en) | Voice synthesization system | |
| JPH01118200A (en) | Voice synthesization system | |
| JP2956069B2 (en) | Data processing method of speech synthesizer | |
| JP2900454B2 (en) | Syllable data creation method for speech synthesizer | |
| JPH0642158B2 (en) | Speech synthesizer | |
| JPS5914752B2 (en) | Speech synthesis method | |
| JP3241582B2 (en) | Prosody control device and method | |
| JPH0358100A (en) | Rule type voice synthesizer | |
| JPS5842099A (en) | Voice synthsizing system | |
| JP2586040B2 (en) | Voice editing and synthesis device | |
| JP2995774B2 (en) | Voice synthesis method | |
| JP3284634B2 (en) | Rule speech synthesizer | |
| JPH0553595A (en) | Speech synthesizing device | |
| JPH03160500A (en) | Speech synthesizer | |
| JPH0756598B2 (en) | Speech synthesis method of speech synthesizer | |
| JPH0258640B2 (en) | ||
| JPS61278900A (en) | Voice synthesizer | |
| KR19980065482A (en) | Speech synthesis method to change the speaking style | |
| JPS6325700A (en) | Long note combination method | |
| JPS63262699A (en) | Voice analyzer/synthesizer | |
| JPS62262100A (en) | Ruled speech synthesizer | |
| JPH038000A (en) | Voice rule synthesizing device |