JPH0594196A - Speech synthesizing device - Google Patents

Speech synthesizing device

Info

Publication number
JPH0594196A
JPH0594196A JP3278806A JP27880691A JPH0594196A JP H0594196 A JPH0594196 A JP H0594196A JP 3278806 A JP3278806 A JP 3278806A JP 27880691 A JP27880691 A JP 27880691A JP H0594196 A JPH0594196 A JP H0594196A
Authority
JP
Japan
Prior art keywords
voice
unit
speech
analysis
synthesis
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
JP3278806A
Other languages
Japanese (ja)
Inventor
Keiichi Yamada
敬一 山田
Yoshiaki Oikawa
芳明 及川
Naoto Iwahashi
直人 岩橋
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Sony Corp
Original Assignee
Sony Corp
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Sony Corp filed Critical Sony Corp
Priority to JP3278806A priority Critical patent/JPH0594196A/en
Publication of JPH0594196A publication Critical patent/JPH0594196A/en
Pending legal-status Critical Current

Links

Abstract

(57)【要約】 【目的】本発明は、実際の人間の音声に比して品質の劣
化が少なく違和感のない合成音を発声し得る音声合成装
置を実現しようとするものである。 【構成】実音声の分析合成において高品質なピツチ変換
法である複素ケプストラム分析を用いると共に、音声の
合成に用いる音声単位データそれぞれ1つにおいて、そ
の有声部分においては音源情報としてのインパルスと声
道特性としての単位応答波形の必要複数組を持ち、その
無声部分においては実音声の切り出し波形を持つように
したことにより、ピツチパターンの変化によるスペクト
ル包絡の歪みを生じることなく、人間の音声に近い高品
質な合成音声を任意に生成し得る。
(57) [Summary] [Object] The present invention is intended to realize a speech synthesizer capable of synthesizing a synthetic speech with less deterioration in quality compared to an actual human speech and without discomfort. [Structure] A complex cepstrum analysis, which is a high-quality Pitch transform method, is used in the analysis and synthesis of real speech, and impulses and vocal tracts are used as sound source information in the voiced part of each voice unit data used in the synthesis of speech. By having multiple sets of required unit response waveforms as characteristics, and having a cutout waveform of the actual voice in the unvoiced part, it is close to human voice without causing distortion of the spectrum envelope due to changes in pitch pattern. It is possible to arbitrarily generate high-quality synthetic speech.

Description

【発明の詳細な説明】Detailed Description of the Invention

【0001】[0001]

【目次】以下の順序で本発明を説明する。 産業上の利用分野 従来の技術 発明が解決しようとする課題 課題を解決するための手段(図1及び図2) 作用(図1及び図2) 実施例 (1)実施例の原理 (2)実施例の音声合成装置(図1及び図2) (3)他の実施例 発明の効果[Table of Contents] The present invention will be described in the following order. Field of Industrial Application Conventional Technology Problem to be Solved by the Invention Means for Solving the Problem (FIGS. 1 and 2) Action (FIGS. 1 and 2) Example (1) Principle of Example (2) Implementation Example speech synthesizer (FIGS. 1 and 2) (3) Other embodiments Effect of the invention

【0002】[0002]

【産業上の利用分野】本発明は音声合成装置に関し、特
に規則合成方式による音声合成装置に適用して好適なも
のである。
BACKGROUND OF THE INVENTION 1. Field of the Invention The present invention relates to a speech synthesizer, and is particularly suitable for application to a speech synthesizer using a rule synthesis method.

【0003】[0003]

【従来の技術】従来、規則合成方式による音声合成装置
においては、入力された文字の系列を解析した後、所定
の規則に従つてパラメータを合成することにより、いか
なる言葉でも音声合成し得るようになされている(特開
平3-119396号公報)。
2. Description of the Related Art Conventionally, in a speech synthesizing apparatus based on a rule synthesizing method, after analyzing a series of input characters and synthesizing parameters according to a predetermined rule, it is possible to synthesize speech with any words. (Japanese Patent Laid-Open No. 3-119396).

【0004】すなわち規則合成方式による音声合成装置
は、入力された文字の系列を解析した後、所定の規則に
従つて、各文節ごとにアクセントを検出し、各文節の並
びから、文字系列全体としての抑揚、ポース等を表現す
るピツチパラメータを合成する。
That is, the speech synthesizer based on the rule synthesizing method analyzes an input character sequence, detects an accent for each phrase according to a predetermined rule, and detects the accent of each phrase as a whole character sequence. Pitch parameters that express intonation, porcelain, etc. are synthesized.

【0005】さらに音声合成装置は、同様に所定の規則
に従つて各文節を例えばCV単位のような音声単位に分
割した後、そのスペクトラムを表現する合成パラメータ
を生成する。これにより、ピツチパラメータ及び合成パ
ラメータに基づいて合成音を発声するようになされてい
る。
Further, the voice synthesizer similarly divides each clause into voice units such as CV units according to a predetermined rule, and then generates a synthesis parameter expressing the spectrum. As a result, a synthesized sound is produced based on the pitch parameter and the synthesis parameter.

【0006】[0006]

【発明が解決しようとする課題】ところで従来の合成パ
ラメータに基づく一般的な合成法として線形予測分析を
用いた残差駆動による合成方式では、合成音声のピツチ
を変更する場合に音源情報である予測残差波形に対して
処理を施している。
By the way, in a conventional residual-based synthesis method using linear prediction analysis as a general synthesis method based on synthesis parameters, when the pitch of synthesized speech is changed, the prediction of the sound source information is performed. The residual waveform is processed.

【0007】すなわち合成音声のピツチ周波数を低くす
る場合には前記予測残差波形の終端から所望のピツチ周
期となるように一定値0を挿入し、また合成音声のピツ
チ周波数を高くする場合には前記予測残差波形を途中で
打ち切ることによつて、所望のピツチ周期にするという
処理である。
That is, when the pitch frequency of the synthesized speech is lowered, a constant value 0 is inserted from the end of the prediction residual waveform so that the desired pitch period is obtained, and when the pitch frequency of the synthesized speech is raised. It is a process of setting the desired pitch cycle by cutting off the prediction residual waveform on the way.

【0008】ところが線形予測分析においては実音声か
らの音源情報(ピツチ情報)と声道特性(スペクトル包
絡)の分離が不完全なため、この波形処理によつて前記
予測残差波形に含まれているスペクトル包絡情報に歪み
を与えることになり、この歪みによつて合成音声の品質
が劣化しやすい問題があつた。
However, in the linear predictive analysis, since the source information (pitch information) and the vocal tract characteristic (spectral envelope) are not completely separated from the actual voice, the waveform is included in the predictive residual waveform by this waveform processing. This results in distortion of the existing spectrum envelope information, and this distortion tends to deteriorate the quality of synthesized speech.

【0009】本発明は以上の点を考慮してなされたもの
で、実際の人間の音声に比して品質の劣化が少なく違和
感のない合成音を発声することができる音声合成装置を
提案しようとするものである。
The present invention has been made in consideration of the above points, and it is an object of the present invention to propose a voice synthesizing device capable of uttering a synthetic voice with less deterioration in quality than an actual human voice and having no discomfort. To do.

【0010】[0010]

【課題を解決するための手段】かかる課題を解決するた
めに第1の発明においては、入力された文字の系列を解
析して得られた単語、文節の境界及び基本アクセントを
蓄積する解析情報蓄積部3と、音声単位内において周期
性を有する有声部分に関しては、実音声の分析処理によ
つて得られた各1ピツチ周期分に対応する音声波形デー
タを音声単位として蓄積し、音声単位内において周期性
のない無声部分に関しては、実音声をそのまま音声波形
データとして蓄積するメモリ部2と、解析情報蓄積部3
の解析情報に基づき、所定の音韻規則及び韻律規則に従
つてピツチパターンを生成する音声合成規則部4と、メ
モリ部2の音声単位及び音声合成規則部4のピツチパタ
ーンに基づいて、音声を合成する音声合成部5とを設け
るようにした。
In order to solve such a problem, in the first invention, analysis information storage for accumulating a word, a bunsetsu boundary, and a basic accent obtained by analyzing a series of input characters With respect to the unit 3 and the voiced portion having the periodicity in the voice unit, the voice waveform data corresponding to each one-pitch cycle obtained by the analysis process of the real voice is accumulated as the voice unit, and the voice waveform data is stored in the voice unit. As for the unvoiced portion having no periodicity, the memory unit 2 that stores the actual voice as it is as voice waveform data and the analysis information storage unit 3
A voice synthesis rule unit 4 for generating a pitch pattern in accordance with a predetermined phonological rule and a prosodic rule based on the analysis information of 1. and a voice unit of the memory unit 2 and a voice synthesis based on the pitch pattern of the voice synthesis rule unit 4. The voice synthesizing unit 5 is provided.

【0011】また第2の発明においては、解析情報蓄積
部3に入力された文章の文字の系列を単語、文節の境界
及び基本アクセントに解析する文章解析部3を設けるよ
うにした。
Further, in the second aspect of the invention, the sentence analysis unit 3 is provided for analyzing the series of characters of the sentence input to the analysis information storage unit 3 into words, bunsetsu boundaries and basic accents.

【0012】[0012]

【作用】実音声の分析合成において高品質なピツチ変換
法である複素ケプストラム分析を用いると共に、音声の
合成に用いる音声単位データそれぞれ1つにおいて、そ
の有声部分においては音源情報としてのインパルスと声
道特性としての単位応答波形の必要複数組を持ち、その
無声部分においては実音声の切り出し波形を持つことに
よつて、ピツチパターンの変化によるスペクトル包絡の
歪みを生じることなく、人間の音声に近い高品質な合成
音声を任意に生成し得る音声合成装置を実現できる。
The complex cepstrum analysis, which is a high-quality Pitch transform method, is used in the analysis and synthesis of the real voice, and the impulse and vocal tract as the sound source information in the voiced part of each voice unit data used for the voice synthesis. By having the required multiple sets of unit response waveforms as characteristics, and having the cutout waveform of the actual voice in the unvoiced part, it is possible to obtain a high-level sound close to that of human voice without causing distortion of the spectrum envelope due to the change of pitch pattern. It is possible to realize a voice synthesizing device that can arbitrarily generate high quality synthetic voice.

【0013】[0013]

【実施例】以下図面について、本発明の一実施例を詳述
する。
DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS An embodiment of the present invention will be described in detail with reference to the drawings.

【0014】(1)実施例の原理 この実施例の場合、合成に使用する音声単位の分析処理
で、実音声の有声部分における音源情報と声道特性の分
離に複素ケプストラム分析を用い、音源情報をインパル
スとして抽出し、また声道特性は音源情報であるインパ
ルスの単位応答として抽出する。
(1) Principle of the embodiment In the case of this embodiment, the complex cepstrum analysis is used to separate the sound source information and the vocal tract characteristics in the voiced part of the real speech in the analysis processing for each voice unit used for synthesis. Is extracted as an impulse, and the vocal tract characteristics are extracted as a unit response of impulse which is sound source information.

【0015】この複素ケプストラム分析は、実音声の分
析合成において高品質なピツチ変換法、発話速度変換法
として既知の分析手法であり、この音声の分析合成にお
いて有益な分析手法を任意文発声の規則合成に用いるよ
うになされている。
This complex cepstrum analysis is an analysis method known as a high-quality pitch conversion method and a speech rate conversion method in the analysis and synthesis of real speech, and the analysis method useful in this analysis and synthesis of speech is a rule of arbitrary sentence utterance. It is designed to be used for synthesis.

【0016】合成に使用する音声単位の有声部分には、
複素ケプストラム分析手法によつて抽出されたインパル
スと単位応答の両者を1つの組合せとして、音声単位有
声部分に必要なフレーム数だけの組合せを有声部分のデ
ータとして貯えておく。また、音声単位の無声部分にお
いては、実音声の無声部分をそのまま切り出してデータ
として貯えておく。
The voiced part of the voice unit used for synthesis is
Both the impulse and the unit response extracted by the complex cepstrum analysis method are set as one combination, and as many combinations as the number of frames required for the voice unit voiced portion are stored as voiced portion data. Further, in the unvoiced portion of each voice unit, the unvoiced portion of the actual voice is cut out as it is and stored as data.

【0017】これにより音声単位はインパルスとその単
位応答からなる複数フレーム分の組合せか、無声部分で
ある実音声の切り出し波形か、あるいはその両者から構
成されることとなる。
As a result, the voice unit is composed of a combination of a plurality of frames consisting of an impulse and its unit response, a cutout waveform of a real voice which is an unvoiced portion, or both.

【0018】このためまずこのような内容で構成される
音声単位データをメモリに貯えた音声単位記憶部と共
に、入力された文字の系列を解析して、単語、文節の境
界及び基本アクセントを検出する文章解析部とを設け
る。
Therefore, first, together with the voice unit storage unit in which voice unit data having such contents is stored in the memory, the sequence of input characters is analyzed to detect the boundaries of words and phrases and basic accents. A sentence analysis unit is provided.

【0019】さらにこれに加えてこの文章解析部の検出
結果に基づいて、所定の韻律規則に従つて、合成音声の
ピツチパターンを生成し、また音韻規則に従つて合成音
声に必要な合成波形データを前記音声単位記憶部から読
み出しを行なう音声合成規則部と、合成波形データ及び
ピツチパターンに基づいて、合成音を生成する音声合成
部とを設ける。
In addition to this, based on the detection result of the sentence analysis unit, a pitch pattern of synthetic speech is generated according to a predetermined prosody rule, and synthetic waveform data necessary for synthetic speech is also generated according to the phonological rule. Is provided from the voice unit storage unit, and a voice synthesis unit that generates a synthetic voice based on the synthetic waveform data and the pitch pattern.

【0020】このようにすれば音声単位データ内のイン
パルスを所望のピツチパターンに対応するピツチ周期の
間隔に順次配置して、それぞれのインパルスと組合せに
なつている単位応答を1組ずつ重畳することによつて音
声を合成するのであるが、音源情報がインパルスである
ためピツチ周期が伸縮してもそれによる音源情報への影
響はほとんどなく、ピツチパターンが大きく変化するよ
うな場合でもスペクトル包絡に歪みが生じない。
In this way, the impulses in the voice unit data are sequentially arranged at intervals of the pitch cycle corresponding to the desired pitch pattern, and the unit responses in combination with the respective impulses are superposed one by one. However, even if the pitch period expands or contracts, it has almost no effect on the sound source information because the sound source information is an impulse, and even if the pitch pattern changes significantly, the spectrum envelope is distorted. Does not occur.

【0021】このように音声のピツチ変換に適した複素
ケプストラム分析を規則合成に用いることによつて、人
間の音声に近い高品質な任意合成音が得られる。また合
成パラメータによる合成方式のように複雑な演算処理を
必要としないため、音声合成部における処理を高速化し
得るようになされている。
As described above, by using the complex cepstrum analysis suitable for the pitch conversion of speech for the rule synthesis, a high quality arbitrary synthesized speech close to human speech can be obtained. Further, since it does not require a complicated calculation process unlike the synthesizing method using the synthesizing parameter, it is possible to speed up the process in the voice synthesizing unit.

【0022】(2)実施例の音声合成装置 図1において、1は全体として演算処理装置構成の音声
合成装置の概略構成を示し、音声単位記憶部2、文章解
析部3、音声合成規則部4及び音声合成部5に分割さ
れ、まず文章解析部3は、所定の入力装置から入力され
たテキスト入力(文字の系列で表された文章等でなる)
を所定の辞書を基準にして解析し、仮名文字列に変換し
た後、単語、文節毎に分解する。
(2) Speech synthesizer of the embodiment In FIG. 1, reference numeral 1 shows a schematic configuration of a speech synthesizer having an arithmetic processing unit as a whole, a voice unit storage unit 2, a sentence analysis unit 3, and a speech synthesis rule unit 4. And the speech synthesis unit 5, and the sentence analysis unit 3 first inputs a text input from a predetermined input device (consists of a sentence represented by a sequence of characters).
Is analyzed based on a predetermined dictionary, converted into a kana character string, and then decomposed into words and phrases.

【0023】すなわち日本語においては、英語のように
単語が分かち書きされていないことから、例えば「米国
産業界」のような言葉は、「米国/産業・界」、「米/
国産/業界」のように2種類区分化し得る。
In other words, in Japanese, words are not separated into words like English, so words such as "US industry" are "US / industry / world" and "US /
There are two types, such as "domestic / industry".

【0024】このため文章解析部3は辞書を参考にしな
がら、言葉の連続関係及び単語の統計的性質を利用し
て、テキスト入力を単語、文節毎に分解するようになさ
れ、これにより単語、文節の境界を検出する。さらに文
章解析部3は、各単語毎に基本アクセントを検出した
後、これらを音声合成規則部4に出力する。
For this reason, the sentence analysis unit 3 is designed to decompose the text input into words and phrases by using the continuity of words and the statistical property of words while referring to the dictionary. Detect the boundaries of. Further, the sentence analysis unit 3 detects basic accents for each word and then outputs them to the speech synthesis rule unit 4.

【0025】音声合成規則部4は日本語の特徴に基づい
て設定された所定の音韻規則に従つて、文章解析部3の
検出結果及びテキスト入力を処理するようになされてい
る。すなわち日本語の自然な音声は、言語学的特性に基
づいて区別すると、約100程度の発声の単位に区分し
得ることが知られており、例えば「さくら」という単語
を発声の単位に区分すると、「sa」+「ku」+「ra」の
3つのCV単位に分割することができる。
The voice synthesis rule unit 4 is adapted to process the detection result of the sentence analysis unit 3 and the text input according to a predetermined phonological rule set based on the characteristics of Japanese. That is, it is known that natural Japanese speech can be divided into about 100 voicing units when distinguished based on linguistic characteristics. For example, when the word "Sakura" is divided into voicing units. , “Sa” + “ku” + “ra” can be divided into three CV units.

【0026】また日本語は単語が連続する場合、連なつ
た後ろの語の語頭音節が濁音化したり(すなわち続濁で
なる)、語頭以外のガ行音が鼻音化したりして、単語単
体の場合と発声が変化する特徴がある。
Further, in Japanese, when words are continuous, the initial syllable of the succeeding words becomes dull (that is, it becomes continuous), and the ga-sound other than the beginning becomes nasal. There is a feature that the utterance changes with the case.

【0027】従つて音声合成規則部4はこれら日本語の
特徴に従つて音韻規則が設定されるようになされ、この
音韻規則に従つてテキスト入力を音韻記号列(すなわち
上述の「sa」+「ku」+「ra」等の連続する列でなる)
に変換するようになされている。さらに音声合成規則部
4は、当該音韻記号列に基づいて、音声単位記憶部2か
ら各音声単位のデータをロードする。
Accordingly, the speech synthesis rule unit 4 is configured to set the phonological rules according to the characteristics of Japanese, and the text input according to the phonological rules is a phonological symbol string (that is, "sa" + "mentioned above). (Consecutive columns such as "ku" + "ra")
It is designed to be converted into. Further, the voice synthesis rule unit 4 loads data of each voice unit from the voice unit storage unit 2 based on the phoneme symbol string.

【0028】ここでこの音声合成装置1においては、波
形編集の手法を用いて合成音を発声するようになされ、
音声単位記憶部2からロードされるデータは、各CV単
位で表される合成音を生成する際に用いられる波形デー
タでなる。この波形合成に用いられる音声単位データは
次のような構成からなる。
Here, in the voice synthesizer 1, a synthesized voice is produced by using a waveform editing method.
The data loaded from the voice unit storage unit 2 is waveform data used when generating a synthetic sound represented by each CV unit. The voice unit data used for this waveform synthesis has the following configuration.

【0029】音声単位データの有声部に関しては、実音
声の有声部分において前記複素ケプストラム分析を用い
て抽出された、1ピツチに対応するインパルスと単位応
答波形を一組として、この組を1つの音声単位データと
して必要なピツチ分だけ貯えたものからなり、また音声
単位データの無声部に関しては、実音声の無声部分の波
形を切り出してそのまま貯えたものからなる。
With respect to the voiced part of the voice unit data, the impulse corresponding to one pitch and the unit response waveform extracted by using the complex cepstrum analysis in the voiced part of the real voice are set as one set, and this set is set as one voice. The unvoiced part of the voice unit data is cut out from the waveform of the unvoiced part of the actual voice and stored as it is.

【0030】従つて音声単位データがCV単位である場
合には、1つの音声単位CVの子音部Cが無声子音であ
る時には無声部分の切り出し波形と、インパルスと単位
応答波形からなる複数組によつて1つの音声単位データ
が構成され、また1つの音声単位CVの子音部Cが有声
子音である時にはインパルスと単位応答波形からなる複
数組のみによつて1つの音声単位データが構成されるこ
ととなる。
Therefore, when the voice unit data is in CV units, when the consonant part C of one voice unit CV is an unvoiced consonant, the unvoiced part cutout waveform and a plurality of sets of impulses and unit response waveforms are used. One voice unit data is formed, and when the consonant part C of one voice unit CV is a voiced consonant, one voice unit data is formed by only a plurality of pairs of impulses and unit response waveforms. Become.

【0031】音声合成規則部4は音声単位記憶部2から
ロードされた音声単位データを、テキスト入力に応じた
順序(以下このデータを合成波形データと呼ぶ)で合成
し、かくして抑揚のない状態で、テキスト入力を読み上
げた合成音声波形を得ることができる。
The voice synthesis rule unit 4 synthesizes the voice unit data loaded from the voice unit storage unit 2 in the order corresponding to the text input (hereinafter, this data will be referred to as synthesized waveform data), and thus, without inflection. , It is possible to obtain a synthetic speech waveform that reads the text input.

【0032】さらに音声合成規則部4は所定の韻律規則
に基づいて、テキスト入力を適当な長さで分割して、切
れ目(すなわちポーズでなる)を検出する。このように
して、図2に示すように、例えばテキスト入力として文
章「きれいな花を山田さんからもらいました」が入力さ
れた場合は(図2(A))、当該テキスト入力は、「き
れいな」、「はな」、「やまださんから」、「もらいま
した」に分解された後、「はな」及び「やまださんか
ら」間にポーズが検出される(図2(B))。
Further, the voice synthesis rule unit 4 divides the text input into appropriate lengths based on a predetermined prosody rule to detect a break (that is, a pause). In this way, as shown in FIG. 2, for example, when the text “A beautiful flower was received from Mr. Yamada” is input as the text input (FIG. 2 (A)), the text input is “clean”. , "Hana", "From Yamada-san", and "I got it", then a pose is detected between "Hana" and "From Yamada-san" (Fig. 2 (B)).

【0033】さらに音声合成規則部4は韻律規則及び各
単語の基本アクセントに基づいて、各文節のアクセント
を検出する。すなわち日本語の文節単体のアクセント
は、感覚的に仮名文字を単位として(以下モーラと呼
ぶ)高低の2レベルで表現することができる。このとき
文節の内容等に応じて、文節のアクセント位置を区別す
ることができる。
Further, the speech synthesis rule unit 4 detects the accent of each phrase based on the prosody rule and the basic accent of each word. That is, the accent of a Japanese phrase alone can be expressed sensuously in two levels, high and low, in units of kana characters (hereinafter referred to as mora). At this time, the accent position of the phrase can be distinguished according to the content of the phrase.

【0034】例えば端、箸、橋は2モーラの単語で、そ
れぞれアクセントのない0型、アクセントの位置が先頭
のモーラにある1型、アクセントの位置が2モーラ目に
ある2型に分類することができる。かくしてこの実施例
において音声合成規則部4は、テキスト入力の各文節
を、1型、2型、0型、4型と分類し(図2(C))こ
れにより文節単位でアクセント及びポーズを検出する。
For example, edges, chopsticks, and bridges are 2-mora words, and are classified into 0 type with no accent, 1 type with accent position in the first mora, and 2 type with accent position in 2nd mora. You can Thus, in this embodiment, the speech synthesis rule unit 4 classifies each phrase of the text input into 1 type, 2 type, 0 type and 4 type (FIG. 2 (C)). To do.

【0035】さらに音声合成規則部4はアクセント及び
ポーズの検出結果に基づいて、テキスト入力全体の抑揚
を表す基本ピツチパターンを生成する。すなわち日本語
においては、文節のアクセントは、感覚的に2レベルで
表し得るのに対し、実際の抑揚は、アクセントの位置か
ら徐々に低下する特徴がある(図2(D))。
Furthermore, the voice synthesis rule unit 4 generates a basic pitch pattern representing the intonation of the entire text input, based on the accent and pause detection results. That is, in Japanese, the accent of a bunsetsu can be sensuously expressed in two levels, while the actual intonation is characterized by gradually decreasing from the position of the accent (FIG. 2 (D)).

【0036】さらに日本語においては、文節が連続して
1つの文章になると、ポーズから続くポーズに向かつ
て、抑揚が徐々に低下する特徴がある(図2(E))。
従つて音声合成規則部4は、かかる日本語の特徴に基づ
いて、テキスト入力全体の抑揚を表すパラメータを各モ
ーラ毎に生成した後、人間が発声した場合と同様に抑揚
が滑らかに変化するように、モーラ間に補間によりパラ
メータを設定する。
Further, in Japanese, when the bunsetsu becomes one sentence continuously, the intonation gradually decreases from one pose to the next (FIG. 2 (E)).
Therefore, the speech synthesis rule unit 4 generates a parameter representing the intonation of the entire text input for each mora based on the characteristics of the Japanese language, and then the intonation changes smoothly as if a human uttered. Then, parameters are set by interpolation between mora.

【0037】かくして音声合成規則部4は、テキスト入
力に応じた順序で、各モーラのパラメータ及び補間した
パラメータを合成し(以下ピツチパターンと呼ぶ)、か
くしてテキスト入力を読み上げた音声の抑揚を表すピツ
チパターン(図2(F))を得ることができる。
Thus, the speech synthesis rule unit 4 synthesizes the parameters of each mora and the interpolated parameters in the order according to the text input (hereinafter referred to as a pitch pattern), and thus the pitch representing the intonation of the speech read out from the text input. A pattern (FIG. 2 (F)) can be obtained.

【0038】次に音声合成部5は合成波形データ及びピ
ツチパターンに基づいて波形合成処理を行ない、合成音
を生成する。この波形合成処理は、次のようなことを行
なつている。合成音声の有声部分においては、合成波形
データ内のインパルスをピツチパターンに基づいて並
べ、その並べられたインパルスそれぞれに対応する単位
応答波形を各インパルスに重畳する。
Next, the voice synthesizing unit 5 performs a waveform synthesizing process based on the synthetic waveform data and the pitch pattern to generate a synthetic sound. This waveform synthesis processing is performed as follows. In the voiced part of the synthetic speech, impulses in the synthetic waveform data are arranged based on a pitch pattern, and a unit response waveform corresponding to each of the arranged impulses is superimposed on each impulse.

【0039】また合成音声の無声部分においては、合成
波形データ内の切り出し波形をそのまま所望の合成音声
の波形とする。これにより、ピツチパターンの変化に追
従して抑揚の変化する合成音を得ることができる。
In the unvoiced part of the synthetic voice, the cut-out waveform in the synthetic waveform data is used as it is as the waveform of the desired synthetic voice. As a result, it is possible to obtain a synthetic sound in which the intonation changes according to the change in the pitch pattern.

【0040】従つて合成音において音源情報にインパル
スを用いているため、合成音のピツチ周期が伸縮しても
それによる音源情報への影響はほとんどなく、ピツチパ
ターンが大きく変化するような場合でもスペクトル包絡
に歪みが生じることなく、人間の音声に近い高品質な任
意合成音が得られる。
Therefore, since the impulse is used for the sound source information in the synthesized sound, even if the pitch period of the synthesized sound expands or contracts, there is almost no effect on the sound source information, and even if the pitch pattern greatly changes, the spectrum is changed. It is possible to obtain a high-quality arbitrary synthesized voice that is similar to human voice without causing distortion in the envelope.

【0041】以上の構成において、所定の入力装置から
入力されたテキスト入力は、文章解析部2で、所定の辞
書を基準にして解析され、単語、文節の境界及び基本ア
クセントが検出される。この単語、文節の境界及び基本
アクセントの検出結果は、音声合成規則部4で、所定の
音韻規則に従つて処理され、抑揚のない状態でテキスト
入力を読み上げた音声を表す合成波形データが生成され
る。
In the above configuration, the text input input from the predetermined input device is analyzed by the sentence analysis unit 2 with reference to the predetermined dictionary, and the word, the boundary of the phrase and the basic accent are detected. The results of detecting the words, the boundaries of the clauses, and the basic accents are processed by the speech synthesis rule unit 4 in accordance with a predetermined phonological rule to generate synthetic waveform data representing a speech in which the text input is read aloud without intonation. It

【0042】さらに単語、文節の境界及び基本アクセン
トの検出結果は、音声合成規則部4で、所定の韻律規則
に従つて処理され、テキスト入力全体の抑揚を表すピツ
チパターンが生成される。ピツチパターンは、合成波形
データと共に音声合成部5に出力され、ここでピツチパ
ターン及び合成波形データに基づいて合成音が生成され
る。
Further, the detection result of the word and phrase boundaries and the basic accent is processed by the voice synthesis rule section 4 in accordance with a predetermined prosody rule to generate a pitch pattern representing the intonation of the entire text input. The pitch pattern is output to the voice synthesizing unit 5 together with the synthetic waveform data, and a synthetic sound is generated based on the pitch pattern and the synthetic waveform data.

【0043】以上の構成によれば、音声の合成に用いる
音声単位データそれぞれ1つにおいて、その有声部分に
おいては音源情報としてのインパルスと声道特性として
の単位応答波形の必要複数組を持ち、その無声部分にお
いては実音声の切り出し波形を持つことによつて、ピツ
チパターンの変化によるスペクトル包絡の歪みを生じる
ことなく、人間の音声に近い高品質な合成音声を任意に
生成し得る音声合成装置を実現できる。
According to the above configuration, in each voice unit data used for synthesizing a voice, the voiced portion has a plurality of required sets of impulses as sound source information and unit response waveforms as vocal tract characteristics. Since the unvoiced part has the cut-out waveform of the actual voice, a voice synthesizer capable of arbitrarily generating high-quality synthesized voice close to human voice without causing distortion of the spectrum envelope due to change in pitch pattern is provided. realizable.

【0044】(3)他の実施例 なお上述の実施例においては、音源情報としてインパル
スを用いて単位応答波形と重畳することによつて波形合
成を行なつているが、このインパルスを理想的なインパ
ルスと見なすことによつて、インパルスをピツチ周期間
隔に並べて重畳することなく、直接単位応答波形をピツ
チパターンに対応するように並べることで、所望の合成
音を生成するようにしてもよい。
(3) Other Embodiments In the above-mentioned embodiment, the waveform is synthesized by superimposing the unit response waveform by using the impulse as the sound source information, but the impulse is ideal. By considering it as an impulse, a desired synthesized sound may be generated by arranging the unit response waveforms directly so as to correspond to the pitch pattern without arranging and superimposing the impulses at pitch period intervals.

【0045】さらに上述の実施例においては、音声単位
記憶部2において音声単位データをCV単位で保持して
いるが、これはCV単位のみではなく、CVC単位など
別の音声単位でデータを保持してもよい。
Further, in the above-described embodiment, the voice unit data is held in the voice unit storage unit 2 in the unit of CV. However, the data is held not only in the unit of CV but also in another unit of voice such as CVC unit. May be.

【0046】[0046]

【発明の効果】上述のように本発明によれば、実音声の
分析合成において高品質なピツチ変換法である複素ケプ
ストラム分析を用いるようにしたことにより、人間の音
声に近い高品質な合成音を任意に合成し得る音声合成装
置を実現できる。
As described above, according to the present invention, a complex cepstrum analysis, which is a high-quality Pitch transform method, is used in the analysis and synthesis of real voice, so that a high-quality synthesized voice close to human voice is obtained. It is possible to realize a voice synthesizing device capable of synthesizing any voice.

【図面の簡単な説明】[Brief description of drawings]

【図1】本発明の一実施例による音声合成装置を示すブ
ロツク図である。
FIG. 1 is a block diagram showing a speech synthesizer according to an embodiment of the present invention.

【図2】その動作の説明に供する略線図である。FIG. 2 is a schematic diagram for explaining the operation.

【符号の説明】[Explanation of symbols]

1……音声合成装置、2……音声単位記憶部、3……文
章解析部、4……音声合成規則部、5……音声合成部。
1 ... Voice synthesis device, 2 ... Voice unit storage unit, 3 ... Text analysis unit, 4 ... Voice synthesis rule unit, 5 ... Voice synthesis unit.

Claims (2)

【特許請求の範囲】[Claims] 【請求項1】入力された文字の系列を解析して得られた
単語、文節の境界及び基本アクセントを蓄積する解析情
報蓄積部と、 音声単位内において周期性を有する有声部分に関して
は、実音声の分析処理によつて得られた各1ピツチ周期
分に対応する音声波形データを上記音声単位として蓄積
し、上記音声単位内において周期性のない無声部分に関
しては、上記実音声をそのまま上記音声波形データとし
て蓄積するメモリ部と、 上記解析情報蓄積部の解析情報に基づき、所定の音韻規
則及び韻律規則に従つてピツチパターンを生成する音声
合成規則部と、 上記メモリ手段の上記音声単位及び上記音声合成規則部
の上記ピツチパターンに基づいて、音声を合成する音声
合成部とを具えることを特徴とする音声合成装置。
1. An analysis information storage unit for storing a word, a segment boundary, and a basic accent obtained by analyzing a sequence of input characters, and a voiced portion having periodicity in a voice unit, a real voice. The voice waveform data corresponding to each one-pitch cycle obtained by the analysis processing of step 1 is stored as the voice unit, and for the unvoiced portion having no periodicity in the voice unit, the actual voice is directly used as the voice waveform. A memory unit for storing as data, a voice synthesis rule unit for generating a pitch pattern according to a predetermined phonological rule and a prosodic rule based on the analysis information of the analysis information storage unit, the voice unit and the voice of the memory means. A voice synthesizing device comprising: a voice synthesizing unit for synthesizing a voice based on the pitch pattern of the synthesis rule unit.
【請求項2】上記解析情報蓄積部には入力された文章の
文字の系列を単語、文節の境界及び基本アクセントに解
析する文章解析部を具えることを特徴とする請求項1に
記載の音声合成装置。
2. The speech analysis system according to claim 1, wherein the analysis information storage unit includes a sentence analysis unit that analyzes a sequence of characters of an input sentence into words, boundaries between phrases, and basic accents. Synthesizer.
JP3278806A 1991-09-30 1991-09-30 Speech synthesizing device Pending JPH0594196A (en)

Priority Applications (1)

Application Number Priority Date Filing Date Title
JP3278806A JPH0594196A (en) 1991-09-30 1991-09-30 Speech synthesizing device

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
JP3278806A JPH0594196A (en) 1991-09-30 1991-09-30 Speech synthesizing device

Publications (1)

Publication Number Publication Date
JPH0594196A true JPH0594196A (en) 1993-04-16

Family

ID=17602433

Family Applications (1)

Application Number Title Priority Date Filing Date
JP3278806A Pending JPH0594196A (en) 1991-09-30 1991-09-30 Speech synthesizing device

Country Status (1)

Country Link
JP (1) JPH0594196A (en)

Similar Documents

Publication Publication Date Title
Isewon et al. Design and implementation of text to speech conversion for visually impaired people
US20230058658A1 (en) Text-to-speech (tts) processing
JP2000206982A (en) Speech synthesizer and machine-readable recording medium recording sentence-to-speech conversion program
JPH0833744B2 (en) Speech synthesizer
JPH086591A (en) Audio output device
JPH0632020B2 (en) Speech synthesis method and apparatus
JP2761552B2 (en) Voice synthesis method
JPH0887297A (en) Speech synthesis system
US6829577B1 (en) Generating non-stationary additive noise for addition to synthesized speech
EP1543503B1 (en) Method for controlling duration in speech synthesis
Rama et al. Thirukkural: a text-to-speech synthesis system
van Rijnsoever A multilingual text-to-speech system
JPH08335096A (en) Text voice synthesizer
JPH0580791A (en) Device and method for speech rule synthesis
JP2001034284A (en) Speech synthesis method and apparatus, and recording medium recording sentence / speech conversion program
JPH06318094A (en) Speech rule synthesizer
JP3235747B2 (en) Voice synthesis device and voice synthesis method
JPH0756590A (en) Speech synthesizer, speech synthesis method and recording medium
Ng Survey of data-driven approaches to Speech Synthesis
JP3614874B2 (en) Speech synthesis apparatus and method
JPH0764586A (en) Speech synthesizer
JPH09292897A (en) Voice synthesizing device
JPH01321496A (en) Speech synthesizing device
JPH07129192A (en) Speech synthesizer
JP2000172286A (en) Simultaneous articulator for Chinese speech synthesis