JPS593496A - Fundamental frequency control system for rule synthesization system - Google Patents

Fundamental frequency control system for rule synthesization system

Info

Publication number
JPS593496A
JPS593496A JP57112864A JP11286482A JPS593496A JP S593496 A JPS593496 A JP S593496A JP 57112864 A JP57112864 A JP 57112864A JP 11286482 A JP11286482 A JP 11286482A JP S593496 A JPS593496 A JP S593496A
Authority
JP
Japan
Prior art keywords
information
fundamental frequency
rule
frequency control
amplitude
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
JP57112864A
Other languages
Japanese (ja)
Inventor
金盛 亨
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Fujitsu Ltd
Original Assignee
Fujitsu Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Fujitsu Ltd filed Critical Fujitsu Ltd
Priority to JP57112864A priority Critical patent/JPS593496A/en
Publication of JPS593496A publication Critical patent/JPS593496A/en
Pending legal-status Critical Current

Links

Abstract

(57)【要約】本公報は電子出願前の出願データであるた
め要約のデータは記録されません。
(57) [Summary] This bulletin contains application data before electronic filing, so abstract data is not recorded.

Description

【発明の詳細な説明】 〔発明の技術分野〕 本発明は規則合成方式の音声合成装置に関し、簡単な制
御により十分な性能が得られるようにしたものでるる。
DETAILED DESCRIPTION OF THE INVENTION [Technical Field of the Invention] The present invention relates to a speech synthesis device using a regular synthesis method, and is capable of obtaining sufficient performance through simple control.

〔発明の従来技術〕[Prior art to the invention]

音声出力装置には大別して2通りの出力方式がある。1
つは出力すべき単語ないし文章をすべて用意しておき、
それらを所望の組合わせで順次出力するものであるが、
出力すべき文章の種類が多数の場合には膨大々記憶容量
を必要とする欠点があるつ これに対して本願の対象とする出力方式は、自然音声の
音韻単位を必要な種類だけパラメータ化して用意してお
き、任意の単語ないし文章をそれら音韻単位から合成す
る規則合成方式である この場合、各音韻単位を並べる
だけでなく、それらのピッチや振幅を制御して、より自
然な音声とすることができる。。
There are roughly two types of output methods for audio output devices. 1
One is to prepare all the words or sentences to be output,
It outputs them sequentially in a desired combination,
When there are many types of sentences to be output, there is a drawback that a huge amount of storage capacity is required.In contrast, the output method that is the subject of this application parameterizes the phonological units of natural speech into only the necessary types. This is a rule synthesis method that synthesizes arbitrary words or sentences from those phonological units.In this case, not only do you arrange each phonological unit, but you also control their pitch and amplitude to make more natural speech. be able to. .

〔従来技術の問題点〕[Problems with conventional technology]

従来方式において、人間の発声に近い滑らかな基本周波
数の時間変化パターンを得るためには、10m5ec程
度の短い時間間隔にて、関数近似等の複雑な手段を用い
て基本周波数を設定す周波数の対数値を用いて行なわれ
る1、これは人間の聴感上は周波数の対数に応じて操作
する方が優れた近似が得られることによる。しかし、こ
の基本周波数情報に応じて音声信号を作成する、いわゆ
る音声合成LSI1liすべ″′C基本周期(周波数の
逆数)を入力するようになっており従って周波数の対数
をm期に変換する必要があシ、そのための複雑な回路及
び変換処理時間を要するという問題がある。
In the conventional method, in order to obtain a smooth time-varying pattern of the fundamental frequency close to that of human speech, the fundamental frequency is set using complex means such as function approximation at short time intervals of about 10 m5ec. 1, which is performed using numerical values, because from the perspective of human hearing, a better approximation can be obtained by operating according to the logarithm of the frequency. However, the so-called speech synthesis LSI 1li, which creates an audio signal according to this fundamental frequency information, requires input of the C fundamental period (reciprocal of the frequency), and therefore it is necessary to convert the logarithm of the frequency into an m period. However, there is a problem in that it requires a complicated circuit and conversion processing time.

さらに、2つの音韻単位を接続する場合、第1の音韻単
位の最終振幅と第2の音韻単位の初期振幅とが異なる場
合、その間を滑らかになるよう補関する必要がある。こ
の振幅情報は一般に対数表示をさらに情報圧縮した形式
のパラメータで与えられるが、これを補関するには一但
整数表現にしてから補関し、再度元のバフメータ形式に
戻して音声合成LSIに入力しており、そのための回路
及び処理時間も大きくなるという問題がある。
Furthermore, when connecting two phoneme units, if the final amplitude of the first phoneme unit and the initial amplitude of the second phoneme unit are different, it is necessary to interpolate the difference so that the difference between them is smooth. This amplitude information is generally given as a parameter in a logarithmic representation with further information compression, but in order to interpolate it, it must be expressed as a simple integer, then interpolated, and then returned to the original buff meter format and input to the speech synthesis LSI. However, there is a problem in that the circuit and processing time required for this increase are large.

〔発明の目的〕[Purpose of the invention]

←発明今月弁→ 本発明はこれらの問題を解決し、簡単な回路で高速処理
ができ、かつ十分な性能を得られる音声合成制御方式を
提供することにある。
←Invention this month valve→ The purpose of the present invention is to solve these problems and provide a speech synthesis control method that can perform high-speed processing with a simple circuit and obtain sufficient performance.

〔発明の構成〕 本発明は上記欠点を解決するため、先ず、基本周波数を
制御する情報は大まかな近似で与え、その情報(デジタ
ル情報)をデジタルフィルタにより平滑化することで、
細かい近似で与えた場合に近い清らがさを得るようにし
ている。
[Structure of the Invention] In order to solve the above-mentioned drawbacks, the present invention first provides information for controlling the fundamental frequency as a rough approximation, and then smoothes the information (digital information) using a digital filter.
I am trying to obtain a clarity that is close to that obtained by giving a detailed approximation.

また、この基本周波数情報をピッチ周期に基づくパラメ
ータに変換するのにテーブル索引方式を採用し、ハード
ウェアの単純化、4?性変更の容易化を達成する。
In addition, a table lookup method is adopted to convert this fundamental frequency information into parameters based on the pitch period, which simplifies the hardware. Achieve ease of gender change.

また、振幅情報の補間処理においては、−但整数値に変
換することをせず、振幅パラメータをそのまま純2進数
と見なして補間することにより、高速かつ高性能の補間
処理を行なうようにしている。
In addition, in the interpolation process of amplitude information, the amplitude parameter is treated as a pure binary number and interpolated without converting it to a negative integer value, thereby achieving high-speed and high-performance interpolation process. .

〔発明の実施例〕[Embodiments of the invention]

第1図は規則合虜方式の説明図であり、「ヤマガタ」と
いう単語を合成する場合を例にしている。音韻単位の分
は形にはいくつかの方式があるが、ここではいわゆるV
CV (母音・子音・母音)の組合わせを用いている。
FIG. 1 is an explanatory diagram of the rule-based method, taking as an example the case where the word "Yamagata" is synthesized. There are several ways to form phonological units, but here we use the so-called V
It uses a combination of CV (vowels, consonants, vowels).

「YAMAGATAJは5つの音韻単位「YAJ 。``YAMAGATAJ is made up of five phonetic units ``YAJ''.

FAMA」、「AGA」、EATA」、「A」を接続し
て得られる(同図(a))。各単位間は同一母音で接続
すればよいので、自然な接続が比較的容易に得られる。
It is obtained by connecting "FAMA", "AGA", "EATA", and "A" ((a) in the same figure). Since each unit can be connected using the same vowel, natural connections can be obtained relatively easily.

同図(b)は基本周波数情報(対数表現)を示しており
、「マ」にアクセントがあることを示している。
FIG. 6B shows fundamental frequency information (logarithmic expression), indicating that "ma" has an accent.

同図(c)は基本周波数のピンチ(周期)表現を示して
おり、対数表現では直線のものが周期では曲線になって
いる。
Figure (c) shows a pinch (periodic) expression of the fundamental frequency, where the logarithmic expression is a straight line, but the period is a curved line.

第2図は本発明の一実施例ブロック図であり1は合成す
べき単語/文章を文字コードで人力する手段、2は文字
列を音韻単位(vcv)の系列情報へ変換する手段、3
は音韻単位の系列をV、Cの系列情報へ分解整列する手
段、4はアクセント、イントネーションのパターンを指
定する手段、5は韻律テーブルでイントネーション等に
応じたパラメータを与えるもの、6は基本周波数や各音
韻間の接続時間の時系列的なパターンを作成する手段、
7はデジタルフィルタ。
FIG. 2 is a block diagram of one embodiment of the present invention, in which 1 is means for manually inputting words/sentences to be synthesized using character codes, 2 is means for converting character strings into sequence information of phoneme units (VCV), and 3
4 is a means for specifying accent and intonation patterns, 5 is a prosodic table that gives parameters according to intonation, etc. 6 is a means for disassembling and arranging the phoneme unit sequence into V and C sequence information, 6 is a means for specifying accent and intonation patterns, 6 is a means for giving parameters according to intonation, etc. means for creating a chronological pattern of connection times between each phoneme;
7 is a digital filter.

8は変換テーブル、9は各音韻単位(VCV)が例えば
10m5ecのサンプリング周期でパラメータ化されて
格納されたファイル、10はVCVパラメータを結合す
る手段、11がパラメータに応じて音声信号を合成する
手段、12はスピーカである。
8 is a conversion table; 9 is a file in which each phoneme unit (VCV) is parameterized and stored at a sampling period of, for example, 10 m5ec; 10 is a means for combining VCV parameters; and 11 is a means for synthesizing an audio signal according to the parameters. , 12 are speakers.

イントネーション情報等によって音韻テーブル5を索引
して基本周波数情報(logf)を求める場合、従来で
は充分滑らかなI ogf曲線を得るには、2の音韻テ
ーブル5の情報を上記サンプリング周期と同程度の周期
で詳細に用意するか、又はテーブル5の情報は徂っぽく
してその代わり比較的複雑な所定の関数を用いてその間
を補間するかしていた。
When finding the fundamental frequency information (logf) by indexing the phoneme table 5 using intonation information, etc., conventionally, in order to obtain a sufficiently smooth Iogf curve, the information in the phoneme table 5 of 2 is indexed at a period comparable to the above-mentioned sampling period. Either the information in Table 5 was prepared in detail, or the information in Table 5 was changed and instead a relatively complicated predetermined function was used to interpolate between them.

これに対して本発明ではテーブル5の情報は例えば10
0m5 e c毎程度に粗っぽくし、かつ補間は直線補
間等の単純な関数で行なう。その代わり、その情報はデ
ジタルフィルタ7にて平滑化を施される。
On the other hand, in the present invention, the information in table 5 is, for example, 10
It is made coarser every 0 m5 e c, and interpolation is performed using a simple function such as linear interpolation. Instead, the information is smoothed by the digital filter 7.

第3図はデジタルフィルタ7の一実施例であり、人力又
は減算器31にて出力Yの遅延(Z−’ ) したもの
yz−iと差をとられ、それをアンプ32で8倍(1>
a>0 ) L、それにYZ 1を加算器33で加算し
て出力される。出力YVi】サンプリング周期(例えば
10m5ec )遅延素子34で遅延されて減算器31
.加算器33にフィードバックされる。この回路の伝達
特性はであり、Z=1−aのときに極であり、S平面に
写像するとZ=e  ”πfTとなる。ここでfは極周
波数、Tはサンプリング周期である。今、a=174.
T=0.01秒とすると、f−−ユA’n(1−a)÷
4.6Hy、、即ちカットオノ2πT 周波数4.611z 、−6db10ctのローパス、
フィルタとなる。
FIG. 3 shows an embodiment of the digital filter 7, in which the difference is taken manually or by a subtracter 31 from the delayed output Y (Z-'), yz-i, and the amplifier 32 multiplies it by eight times (1 >
a>0) L, and YZ 1 are added to it by an adder 33 and output. Output YVi] Sampling period (for example, 10 m5ec) is delayed by the delay element 34 and sent to the subtracter 31
.. It is fed back to the adder 33. The transfer characteristic of this circuit is , which is a pole when Z = 1-a, and when mapped to the S plane, Z = e "πfT. Here, f is the polar frequency and T is the sampling period. Now, a=174.
If T=0.01 seconds, then f--UA'n(1-a)÷
4.6Hy, i.e. cut ono 2πT frequency 4.611z, -6db10ct low pass,
It becomes a filter.

このようなりィルクを用いれば、第1図(h)の実線の
ような直線近似でも同図点線のような滑らかなIogf
情報となる。
If such a dirk is used, even a linear approximation like the solid line in Fig. 1(h) will result in a smooth Iogf like the dotted line in Fig. 1(h).
It becomes information.

さらに、このようにして得た基本周波数の対数表現情報
を周期情報に変換するには変換テーブル8を用いる。第
4図はテーブルの一実施例を示し、ROM(リードオン
リーメモリ)を用い、人力には例えば8ビツト2進数で (0〜255)10が与えられ、これに対して出力には
7ビツト2進数で(111〜21)1oが出力される。
Furthermore, a conversion table 8 is used to convert the logarithmic expression information of the fundamental frequency obtained in this way into period information. FIG. 4 shows an example of a table, using a ROM (read only memory). For example, 8-bit binary numbers (0 to 255) 10 are given to the human input, whereas the output is 7-bit binary numbers. (111 to 21) 1o is output in base number.

このようなテーブル変換方式によれば、ROMを又換又
は切換えるのみで、任意の特性の装置に変換することが
できる。例えば男声から女声への変更等が容易に行える
According to such a table conversion method, it is possible to convert the device into a device with arbitrary characteristics simply by converting or switching the ROM. For example, changing from a male voice to a female voice can be easily performed.

また、各音韻単位間の振幅補間は以下のとおりに行う。Further, amplitude interpolation between each phoneme unit is performed as follows.

尚、VCvファイル9には各音韻単位が10m5ec毎
の各種パラメータの時系列集合として記憶されている。
Note that each phoneme unit is stored in the VCv file 9 as a time-series set of various parameters every 10 m5ec.

振幅情報もそのノくラメータの1つである。第1図(a
)の[YAJの最後と1’−A M AJ の最初のよ
うに同一母音「A」でもその振幅は一般に異なっており
、それらを接続するときには滑らかにつながるように補
間を施す必要がある。パラメータ中の振幅情報は、一般
に次表のように表わされる。
Amplitude information is also one of the parameters. Figure 1 (a
) The amplitudes of the same vowel "A" are generally different, such as at the end of [YAJ and at the beginning of 1'-A M AJ , and when connecting them, it is necessary to perform interpolation to connect them smoothly. The amplitude information in the parameters is generally expressed as shown in the table below.

即ち、10進整数は浮動小数点の2進表示にされ、その
小数点以下第1桁は常に”1°゛であるので(正視化し
である故)それを省いてビット数を圧縮している。例え
ばある音(H(の最終振幅、<ラメータが1′1000
1 ”で、次の音韻の最初の振幅パラメータが°’10
101”の場合、従来は各〕(ラメータを整数値″10
”と°°20“とに−但変換し、その間の直線補間をと
って古び);ラメータ形式に戻している。
In other words, a decimal integer is expressed as a floating point binary number, and the first digit below the decimal point is always "1°", so it is omitted to compress the number of bits.For example: The final amplitude of a certain sound (H), < rammeter is 1'1000
1”, the first amplitude parameter of the next phoneme is °’10
101", conventionally each] (parameter was set to an integer value of "10")
” and °°20” (However, linear interpolation between them is taken and the old); it is returned to the parameter format.

これに対して本発明では、ノくラメータ形式をその′ま
ま純2進数と見なして直接に直線補間を行う。即ち上記
の例では、” 10001 ”と” 10101 ”の
差は00100 ”であり、仮に4点間に直線補間する
とすれば、各点に’00001″の差をつければよい。
In contrast, in the present invention, the parameter format is treated as a pure binary number and linear interpolation is directly performed. That is, in the above example, the difference between "10001" and "10101" is 00100", and if linear interpolation is performed between four points, it is sufficient to add a difference of "00001" to each point.

従って前夫のとおり、ノくラメータ形式は’10001
”から1′10101 ” tで”00001”きざみ
で補間される。これを10進、表示に直してみると、+
110”、121“、′14“l、16″。
Therefore, as my ex-husband said, the parameter format is '10001'.
``1'10101'' t is interpolated in steps of ``00001''. If you convert this to decimal and display, +
110", 121",'14"l, 16".

1′20°′となり、指数関数的な補間になっているの
が判る。この方法は無駄な変換が不要なだけでなく、む
しろ聴感特性上も優れているという効果がある。という
のは、人間の聴感上は周波数、振幅等が対数的に変化し
ても、それが直線的な変化にしか感じないという%性が
あり、補間を施すにも対数関数的な補間の方が優れてい
るのである。
It becomes 1'20°', and it can be seen that it is an exponential interpolation. This method not only eliminates the need for unnecessary conversion, but also has the advantage of superior auditory characteristics. This is because human hearing senses that even if frequency, amplitude, etc. change logarithmically, it only feels like a linear change. is superior.

第5図に補間用回路の例を示すが、この回路はVCV結
合部10の中にあると考えてよい。
An example of an interpolation circuit is shown in FIG. 5, and this circuit may be considered to be within the VCV coupling unit 10.

前回の音韻単位の最終振幅パラメータはレジスタ41に
、また次音韻単位の先頭振幅パラメータはレジスタ42
にセットされ、夫々を純2進数として減算器43にて差
をとる。一方、音韻単位間の接続時間を決定する手段6
からの接続時間(サンプル点数)でその差を除算器44
で割り、その商を加算器45にて前回最終振幅値に加算
して行けばよい。
The final amplitude parameter of the previous phonetic unit is stored in the register 41, and the first amplitude parameter of the next phonetic unit is stored in the register 42.
, and the subtracter 43 calculates the difference by treating each as a pure binary number. On the other hand, means 6 for determining the connection time between phonological units
Divider 44 divides the difference by the connection time (number of sample points) from
The adder 45 adds the quotient to the previous final amplitude value.

〔発明の効果〕〔Effect of the invention〕

以上の如く本発明によれば、単純な回路で単純な処理を
行うことで従来と同等、或いはより優れた特性を得るこ
とができ、音声出力装置のコストダウン化に有効である
As described above, according to the present invention, by performing simple processing with a simple circuit, it is possible to obtain characteristics equivalent to or better than the conventional ones, and it is effective in reducing the cost of the audio output device.

【図面の簡単な説明】[Brief explanation of the drawing]

第1図は規則合成方式の説明図、第2図は本発明の一実
施例概略ブロック図、第3図はデジタルフィルタの一実
施例プロック図、第4図は変換テーブルの一実施例ブロ
ック図、第5図は補間回路の一実施例ブロック図である
。 第 1 図 晃 λ 口
Fig. 1 is an explanatory diagram of a rule synthesis method, Fig. 2 is a schematic block diagram of an embodiment of the present invention, Fig. 3 is a block diagram of an embodiment of a digital filter, and Fig. 4 is a block diagram of an embodiment of a conversion table. , FIG. 5 is a block diagram of an embodiment of the interpolation circuit. Figure 1 Akira λ Mouth

Claims (1)

【特許請求の範囲】[Claims] 複数の音韻単位を所定のサンプリング周期でパラメータ
化して蓄積し、所望の音韻単位を接続し、かつ基本周波
数情報の時間変化パターンを与えて音声合成する規則合
成方式において、上記基本周波数情報をアドレスとして
索引されるテーブルVcjって該情報を他のパラメータ
形式に変換することを特徴とする規則合成方式における
基本周波数制御方式。
In a rule synthesis method that parameterizes and stores multiple phoneme units at a predetermined sampling period, connects the desired phoneme units, and synthesizes speech by giving a time-varying pattern of fundamental frequency information, the fundamental frequency information is used as an address. A basic frequency control method in a rule synthesis method, characterized in that the indexed table Vcj converts the information into another parameter format.
JP57112864A 1982-06-30 1982-06-30 Fundamental frequency control system for rule synthesization system Pending JPS593496A (en)

Priority Applications (1)

Application Number Priority Date Filing Date Title
JP57112864A JPS593496A (en) 1982-06-30 1982-06-30 Fundamental frequency control system for rule synthesization system

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
JP57112864A JPS593496A (en) 1982-06-30 1982-06-30 Fundamental frequency control system for rule synthesization system

Publications (1)

Publication Number Publication Date
JPS593496A true JPS593496A (en) 1984-01-10

Family

ID=14597433

Family Applications (1)

Application Number Title Priority Date Filing Date
JP57112864A Pending JPS593496A (en) 1982-06-30 1982-06-30 Fundamental frequency control system for rule synthesization system

Country Status (1)

Country Link
JP (1) JPS593496A (en)

Cited By (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPS60225198A (en) * 1984-04-20 1985-11-09 三洋電機株式会社 Voice synthesizer by rule
JPS6170597A (en) * 1984-09-14 1986-04-11 株式会社日立製作所 speech synthesizer
JPS6375799A (en) * 1986-09-19 1988-04-06 富士通株式会社 Voice rule synthesizer

Citations (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPS56158553A (en) * 1980-05-12 1981-12-07 Nec Corp Communication controller providing code converting function
JPS5792959A (en) * 1980-12-01 1982-06-09 Fujitsu Ltd System for terminal test

Patent Citations (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPS56158553A (en) * 1980-05-12 1981-12-07 Nec Corp Communication controller providing code converting function
JPS5792959A (en) * 1980-12-01 1982-06-09 Fujitsu Ltd System for terminal test

Cited By (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPS60225198A (en) * 1984-04-20 1985-11-09 三洋電機株式会社 Voice synthesizer by rule
JPS6170597A (en) * 1984-09-14 1986-04-11 株式会社日立製作所 speech synthesizer
JPS6375799A (en) * 1986-09-19 1988-04-06 富士通株式会社 Voice rule synthesizer

Similar Documents

Publication Publication Date Title
CA1065490A (en) Emphasis controlled speech synthesizer
US5490234A (en) Waveform blending technique for text-to-speech system
US7831420B2 (en) Voice modifier for speech processing systems
JPS59192295A (en) Multiplication/addition circuit
CN110599998B (en) Method and device for generating voice data
JPWO2011004579A1 (en) Voice quality conversion device, pitch conversion device, and voice quality conversion method
CN113160849A (en) Singing voice synthesis method and device, electronic equipment and computer readable storage medium
WO1982002109A1 (en) Method and system for modelling a sound channel and speech synthesizer using the same
US20240274120A1 (en) Speech synthesis method and apparatus, electronic device, and readable storage medium
JPS593496A (en) Fundamental frequency control system for rule synthesization system
CN119132314B (en) Variable bit rate de-redundant speech semantic coding method and device
JPS593495A (en) Fundamental frequency control system for rule synthesization system
JPS593497A (en) Fundamental frequency control system for rule synthesization system
CN114242034A (en) A kind of speech synthesis method, apparatus, terminal equipment and storage medium
JPS58161000A (en) Voice synthesizer
JP2987089B2 (en) Speech unit creation method, speech synthesis method and apparatus therefor
JP2650480B2 (en) Speech synthesizer
JP2003066983A (en) Speech synthesis apparatus, speech synthesis method, and program recording medium
JPS5968793A (en) Voice synthesizer
JPS58168095A (en) Voice synthesizer
JPS6295595A (en) Voice response method
JPH0632037B2 (en) Speech synthesizer
JPS63285596A (en) Speech speed altering system for voice synthesization
JP2861005B2 (en) Audio storage and playback device
JPS6143797A (en) Voice editing output system