WO2004109660A1 - 音声データを選択するための装置、方法およびプログラム - Google Patents
音声データを選択するための装置、方法およびプログラム Download PDFInfo
- Publication number
- WO2004109660A1 WO2004109660A1 PCT/JP2004/008088 JP2004008088W WO2004109660A1 WO 2004109660 A1 WO2004109660 A1 WO 2004109660A1 JP 2004008088 W JP2004008088 W JP 2004008088W WO 2004109660 A1 WO2004109660 A1 WO 2004109660A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- data
- speech
- unit
- voice
- representing
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Classifications
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L13/00—Speech synthesis; Text to speech systems
- G10L13/02—Methods for producing synthetic speech; Speech synthesisers
- G10L13/027—Concept to speech synthesisers; Generation of natural phrases from machine-based concepts
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L13/00—Speech synthesis; Text to speech systems
- G10L13/06—Elementary speech units used in speech synthesisers; Concatenation rules
Definitions
- the present invention relates to an audio data selection device, an audio data selection method, and a program.
- the recording and editing method is used for voice guidance systems at stations and navigation devices for vehicles.
- a word is associated with voice data representing a voice that reads out the word, a sentence to be subjected to voice synthesis is divided into words, and voice data associated with these words is acquired. It is a technique of connecting together.
- prosody prediction is an extremely complicated process
- it is necessary to use a processor with high processing power or to perform the processing over a long time. Therefore, this method is not suitable for applications that require high-speed processing using a device with a simple configuration.
- the present invention has been made in view of the above situation, and has as its object to provide an audio data selection device, an audio data selection method, and a program for obtaining a natural synthesized voice at high speed with a simple configuration. .
- the audio data selection device basically includes a storage unit that stores a plurality of audio data representing a waveform of an audio; Search means for inputting sentence information representing a sentence, and searching for audio data representing a waveform of a sound unit having a common reading with a sound unit constituting the sentence from among the sound data; Value of the speech data corresponding to each of the speech units constituting the sentence, and the pitch difference at the boundary between adjacent speech units in the entire sentence. And selecting means for selecting so that is minimized.
- the voice data selecting device may further include voice synthesizing means for generating data representing a synthesized voice by combining the selected voice data with each other.
- the voice data selection method of the present invention basically stores a plurality of voice data representing a voice waveform, inputs text information representing a text, and forms the text from the voice data. Speech data representing the waveform of a speech unit having the same reading as the speech unit is found, and one of the found speech data is speech data corresponding to each of the speech units constituting the sentence. Each time, a pitch difference at a boundary between adjacent sound pieces is selected so that a value obtained by accumulating the entire text is minimized.
- the computer program according to the present invention further comprises: a storage unit for storing a plurality of voice data representing a waveform of a voice; and text information representing a text.
- a search unit that searches for a speech data representing a waveform of a speech unit having the same reading as the constituent speech units, and a search unit that searches the speech data for each of the speech units constituting the sentence from the retrieved speech data.
- Selecting means for selecting the corresponding voice data one by one, and selecting the pitch difference at the boundary between adjacent speech pieces so that the value obtained by accumulating the total of the entire sentence is minimized. It has become something.
- the voice selection device basically inputs storage information for storing a plurality of voice data representing a voice waveform, text information representing a text, and reads the text.
- a prediction unit that predicts a temporal change in the pitch of the speech unit by performing prosodic prediction on the speech unit that constitutes the speech unit; and a speech unit that has the same reading as the speech unit that constitutes the sentence among the speech data.
- selecting means for selecting audio data which represents a waveform and whose time change in pitch has the highest correlation with the result of prediction by the predicting means.
- the selection means includes a regression calculation for performing a first-order regression between a time change of a pitch of a sound piece represented by voice data and a time change of a pitch of a sound piece in the sentence having the same reading as the sound piece. Based on the result, the strength of the correlation between the time change of the pitch of the audio data and the result of the prediction by the prediction means may be specified.
- the selecting means is configured to determine, based on a correlation coefficient between a temporal change of a pitch of a voice unit represented by voice data and a temporal change of a pitch of a voice unit in the text that is commonly read with the voice unit, the voice Any correlation strength may be specified as a result of the time change of the pitch over time and the result of prediction by the prediction means.
- Another voice selecting device of the present invention is a storage means for storing a plurality of voice data representing a voice waveform, and inputting text information representing a text, and predicting a prosody of a speech unit in the text.
- a prediction unit for predicting the time length of the speech unit and the temporal change of the pitch of the speech unit, and each audio data representing the waveform of the speech unit having the same reading as the speech unit in the text.
- Selecting means for specifying an evaluation value for evening and selecting audio data having the highest evaluation value, wherein the evaluation value is a temporal change in pitch of a sound piece represented by the audio data.
- the numerical value representing the correlation is obtained by a first-order regression between the time change of the pitch of the sound piece represented by the voice data and the time change of the pitch of the sound piece in the sentence having the same reading as the sound piece.
- the gradient of the linear function It may be.
- the numerical value representing the correlation is obtained by a first-order regression between the time change of the pitch of the sound piece represented by the voice data and the time change of the pitch of the sound piece in the sentence having the same reading as the sound piece. It may consist of the intercept of the obtained linear function.
- the numerical value representing the correlation is a correlation coefficient between the time change of the pitch of the sound piece represented by the voice data and the predicted result of the time change of the pitch of the sound piece in the sentence having the same reading as the sound piece. It may consist of.
- the numerical value representing the correlation is a function represented by the data representing the time change of the pitch of the sound piece represented by the voice data, which is obtained by shifting the number of bits cyclically.
- the sound in the sentence having the same reading as the sound piece is read. It may consist of the maximum value of the correlation coefficient with a function representing the prediction result of the temporal change of the pitch of the piece.
- the storage means may store phonogram data representing the reading of the voice data in association with the voice data, and the selecting means may store a reading matching the reading of the speech unit in the text. Speech data to which the represented phonogram data is associated may be handled as speech data representing the waveform of the vocal piece having the same reading as the relevant vocal piece.
- the voice selecting device may further include a voice synthesizing unit that generates data representing a synthesized voice by combining the selected voice data with each other.
- the voice selecting device may be configured to represent a waveform of the voice unit of the voice unit in which the voice data is not selected by the storage unit, without using the voice data stored in the storage unit.
- the apparatus may further comprise a missing part synthesizing means for synthesizing voice data, wherein the voice synthesizing means comprises: The data synthesized by the means may be combined with each other to generate a data representing the synthesized voice.
- the voice selection method of the present invention stores a plurality of voice data representing the waveform of a voice, inputs text information representing a text, and performs prosody prediction on a voice unit constituting the text, thereby obtaining a prosody of the voice unit.
- a time change of pitch is predicted, and a waveform of a voice unit having a common reading with a voice unit constituting the sentence is represented from among the voice data, and the time change of the pitch is determined by the prediction unit. And selecting the speech data that shows the highest correlation with the result of the prediction by.
- another voice selection method of the present invention stores a plurality of voice data representing a voice waveform, inputs text information representing a text, and performs prosody prediction on a speech unit in the text.
- Estimate the time length of the voice unit and the time change of the pitch of the voice unit specify the evaluation value of each voice data representing the waveform of the voice unit having the same reading as the voice unit in the sentence, The voice data having the highest evaluation value is selected, and the evaluation value is determined based on the time change of the pitch of the sound piece represented by the voice data and the text in which the reading is common to the sound piece.
- the computer program of the present invention is a computer program for storing a plurality of voice data representing voice waveforms, inputting text information representing a text, and performing prosodic prediction on a speech unit constituting the text.
- Prediction means for predicting a temporal change in the pitch of the speech piece, and a speech piece constituting the text from each of the speech data Represents the waveform of a common speech unit, and a selection means for selecting voice data whose temporal change in pitch has the highest correlation with the result of the prediction by the prediction means. It has become.
- another computer program is a computer program, comprising: a storage means for storing a plurality of voice data representing a waveform of a voice; and text information representing a text, and a prosody prediction for a speech unit in the text.
- a prediction unit that predicts the time length of the speech unit and the time change of the pitch of the speech unit, and each voice data representing the waveform of the speech unit having the same reading as the speech unit in the text
- a program for specifying an evaluation value of the speech data and causing it to function as selection means for selecting audio data representing the highest evaluation, wherein the evaluation value is a speech data represented by a speech unit represented by evening.
- a function of a numerical value representing the correlation between the time change of the pitch of the sound piece and the prediction result of the time change of the pitch of the sound piece in the sentence that is common to the sound piece and the time of the sound piece represented by the sound data The chief, It is obtained from the function of the difference between the speech unit and the prediction result of the time length of the speech unit in the sentence having the same reading.
- the speech data selection device basically comprises a storage means for storing a plurality of speech data representing a speech waveform, and a sentence for inputting sentence information representing a sentence.
- An information input unit a search unit that searches for voice data having a portion that is common to a voice unit in the text represented by the text information, and a search unit that connects the searched voice data in accordance with the text represented by the text information
- Selecting means for obtaining an evaluation value according to a predetermined evaluation criterion based on a relationship between mutually adjacent voice data and selecting a combination of output voice data based on the evaluation value.
- the evaluation criterion is a criterion for determining a correlation between a voice represented by voice data and a prosody prediction result and an evaluation value indicating a relationship between voice data adjacent to each other, wherein the evaluation value is a value of a voice represented by the voice data.
- An evaluation expression including at least one of a parameter indicating a feature, a parameter indicating a feature of a voice obtained by combining the voices represented by the voice data with each other, and a parameter indicating a feature related to the speech time length. It may be one that is obtained based on this.
- the evaluation criterion is a criterion for determining a correlation between a voice represented by voice data and a prosody prediction result and an evaluation value indicating a relationship between voice data adjacent to each other.
- the parameters indicating the characteristics of the voice obtained by combining the voices represented by the voice data with each other, and the parameters indicating the characteristics of the voice represented by the voice data and the parameters indicating the characteristics related to the speech duration may be obtained based on an evaluation formula including at least one of them.
- the parameter indicating the characteristic of the voice obtained by combining the voices represented by the voice data with each other is selected from among voice data representing a waveform of a voice having a portion common to a speech unit in a text represented by the text information and a reading. This is obtained based on the pitch difference at the boundary between adjacent audio data when one audio data corresponding to each sound piece constituting the sentence is selected one by one. There may be.
- the speech unit data selection device inputs sentence information representing a sentence and predicts the prosody of the speech unit in the sentence, thereby predicting the time length of the speech unit and the time change of the pitch of the speech unit.
- the evaluation criterion defines an evaluation value indicating a correlation or a difference between a voice represented by voice data and a prosody prediction result of the prosody prediction means.
- the evaluation value is a criterion, and the evaluation value is a correlation between a time change of a pitch of a sound piece represented by voice data and a prediction result of a time change of a pitch of a sound piece in the sentence having the same reading as the sound piece.
- the numerical value representing the correlation is obtained by a first-order regression between the time change of the pitch of the sound piece represented by the voice data and the time change of the pitch of the sound piece in the sentence having the same reading as the relevant sound piece. It may consist of the slope and / or intercept of a linear function.
- the numerical value representing the correlation is a correlation coefficient between the time change of the pitch of the sound piece represented by the voice data and the predicted result of the time change of the pitch of the sound piece in the sentence having the same reading as the sound piece. It may consist of.
- the numerical value representing the correlation may be a function represented by data obtained by shifting the pitch of the speech piece represented by the voice data over time by various numbers of bits, and a value represented in the sentence having the same reading as the speech piece. It may be made up of the maximum value of the correlation coefficient with the function representing the prediction result of the time change of the pitch of the sound piece.
- the storage unit may store phonogram data representing a reading of voice data in association with the voice data, and the selecting unit may represent a reading matching a reading of a speech unit in the text.
- the voice data associated with the phonetic data may be handled as voice data representing the waveform of a voice unit having the same reading as the relevant voice unit.
- the speech unit data selection device includes a speech synthesis unit that generates data representing a synthesized speech by combining the selected speech data with each other. 'It may have more.
- the sound piece de-night selection device for the sound pieces in which the selecting means cannot select the sound data among the sound pieces in the text, without using the sound data stored in the storage means.
- Missing voice synthesizing means for synthesizing voice data representing the waveform of the voice signal, wherein the voice synthesizing means combines the voice data selected by the selecting means and the voice data synthesized by the missing voice portion with each other. By combining the data, data representing the synthesized voice may be generated.
- the voice data selection method of the present invention stores a plurality of voice data representing a voice waveform, inputs text information representing a text, and has a portion that is commonly read with a speech unit in the text represented by the text information.
- voice data is searched, and the searched voice data is connected in accordance with a text represented by text information, an evaluation value is obtained and output according to a predetermined evaluation criterion based on a relationship between adjacent voice data. It includes a first-speed processing step of selecting a combination of audio data based on the evaluation value.
- the computer program of the present invention comprises: a computer, a storage unit for storing a plurality of voice data representing a voice waveform, a text information input unit for inputting text information representing text, and a text information represented by the text information.
- a search unit that searches for voice data having a portion that is common to the voice unit of the voice unit, and voice data that are adjacent to each other when the searched voice data are connected in accordance with the text represented by text information.
- the evaluation means obtains an evaluation value in accordance with a predetermined evaluation criterion based on the above relationship, and functions as a selecting means for selecting a combination of output audio data based on the evaluation value.
- FIG. 1 is a block diagram showing a configuration of a speech synthesis system according to each embodiment of the present invention.
- FIG. 2 is a diagram schematically showing a data structure of a sound piece database according to the first embodiment of the present invention.
- FIG. 4B is a graph for explaining a process of linearly regressing the change
- FIG. 4B is a graph showing an example of a prediction result data and a value of pitch component data used for obtaining a correlation coefficient. is there.
- FIG. 4 is a diagram schematically showing a data structure of a speech piece database according to a second embodiment of the present invention.
- Fig. 5 (a) is a diagram showing the reading of a fixed message
- Fig. 5 (b) is a list of the speech unit data supplied to the speech unit editing unit
- Fig. 5 (c) is
- FIG. 9D is a diagram showing the absolute value of the difference between the frequency of the pitch component at the end of the preceding speech unit and the frequency of the pitch component at the beginning of the succeeding speech unit.
- FIG. It is a figure which shows whether one day is selected.
- FIG. 6 is a flowchart showing processing when a personal computer performing the function of the speech synthesis system according to each embodiment of the present invention has acquired free text data.
- FIG. 7 is a flowchart showing processing when a personal computer that performs the function of the speech synthesis system according to each embodiment of the present invention has acquired distribution character string data.
- FIG. 8 is a diagram showing a speech synthesis according to the first embodiment of the present invention.
- FIG. 5 is a flowchart showing a process performed when a personal computer performing the above function obtains fixed message data and utterance speed data.
- FIG. 9 is a flowchart showing a process when a personal computer performing the function of the speech synthesis system according to the second embodiment of the present invention acquires fixed message data and utterance speed data.
- FIG. 10 is a flowchart showing a process performed when a personal computer performing the function of the speech synthesis system according to the third embodiment of the present invention acquires fixed message data and utterance speed data. is there.
- FIG. 1 is a diagram showing a configuration of a speech synthesis system according to a first embodiment of the present invention.
- the speech synthesis system includes a main unit M and a speech unit registration unit R.
- the main unit M consists of a language processing unit 1, a general word dictionary 2, a user word dictionary 3, a sound processing unit 4, a search unit 5, a decompression unit 6, a waveform data base 7, and a speech unit. It comprises an editing unit 8, a search unit 9, a speech unit database 10, and a speech speed conversion unit 11.
- the language processing unit 1, the sound processing unit 4, the search unit 5, the decompression unit 6, the speech unit editing unit 8, the search unit 9, and the speech speed conversion unit 11 are all CPUs (Central ⁇
- a processor such as a processing unit (DP) and a digital signal processor (DP), and a memory for storing a program to be executed by the processor, and performs processing described later.
- DP processing unit
- DP digital signal processor
- a single processor performs part or all of the functions of the language processing unit 1, the sound processing unit 4, the search unit 5, the decompression unit 6, the speech unit editing unit 8, the search unit 9, and the speech speed conversion unit 11. You may do so.
- the general word dictionary 2 is composed of a non-volatile memory such as a PROM (Programmable Read Only Memory) and a hard disk device.
- the general word dictionary 2 contains words including ideographic characters (for example, kanji) and phonograms (for example, kana and phonograms) representing the reading of the words and the like. They are stored in advance in association with each other by a manufacturer or the like.
- the user word dictionary 3 is a data rewritable non-volatile memory such as an EEPROM (Electrically Erasable / Programmable Read Only Memory) and a hard disk device, and a control for controlling data writing to the non-volatile memory. It consists of a circuit. Note that the processor may perform the function of this control circuit, and the language processing unit 1, the sound processing unit 4, the search unit 5, the decompression unit 6, the speech unit editing unit 8, the search unit 9, and the speech speed conversion unit 11 A processor that performs a part or all of the functions may perform the function of the control circuit of the user word dictionary 3.
- EEPROM Electrically Erasable / Programmable Read Only Memory
- the user word dictionary 3 acquires words including ideographic characters and phonograms indicating reading of the words and the like from outside according to user operations, and stores them in association with each other.
- the user word dictionary 3 stores words and the like not stored in the general word dictionary 2 and phonograms representing their readings. Is enough.
- the waveform database 7 is composed of a nonvolatile memory such as a PR ⁇ M or a hard disk device.
- the waveform data base 7 contains the phonograms and the compressed waveform data obtained by subjecting the speech data representing the waveform of the unit voice represented by the phonograms to the entrance-to-peak coding. They are stored in advance in association with each other by a stem manufacturer or the like.
- the unit speech is a speech that is short enough to be used in the rule-based synthesis method, and specifically, is speech that is separated by units such as phonemes and V CV (Vowel-Consonant-Vowel) syllables.
- the waveform data before being subjected to the entropy encoding may be composed of, for example, digital data that has been subjected to pulse code modulation (PCM).
- PCM pulse code modulation
- the voice element data base 10 is composed of a nonvolatile memory such as a PROM and a hard disk device.
- the speech unit database 10 stores, for example, data having a data structure shown in FIG. That is, as shown in the figure, the data stored in the speech unit database 10 is divided into four types: a header section HDR, an index section IDX, a directory section DIR, and a data section DAT. ing.
- the storage of the data in the voice unit database 10 is performed in advance by, for example, the manufacturer of the voice synthesis system, and is performed by the Z or the voice unit registration unit R performing an operation described later.
- the header HDR contains data identifying the speech unit database 10 and data indicating the data amount, data format, copyright, etc. of the index part IDX, directory part DIR and data part DAT. But Will be delivered.
- the data section DAT stores the compressed speech unit data obtained by entropy-encoding the speech unit data representing the speech unit waveform.
- a speech unit is a continuous section containing one or more phonemes in a voice, and usually consists of one or more words.
- the speech piece data before entropy encoding is the same format as the waveform data before entropy encoding for generating the above-described compressed waveform data (for example, digital Format).
- the directory section DIR contains individual compressed audio data
- FIG. 2 shows that the data included in the data part DAT is a compressed speech piece data having a data amount of 141 h bytes, which represents the waveform of a speech piece whose reading is “Saitama”.
- a 3 6 A 6 'h first The case where it is stored in a logical position is illustrated. (In this specification and the drawings, the number suffixed with “h” indicates a hexadecimal number.)
- the pitch component data is, for example, as shown in the figure, the frequency of the pitch component of the sound piece. It is assumed that the data represents a sample Y (i) obtained by sampling (where n is a positive integer equal to or less than n, where n is the total number of samples).
- At least the data (A) (that is, the speech unit reading data) of the above set of data items (A) to (E) is sorted according to the order determined based on the phonetic characters represented by the speech unit reading data. In a single state (for example, if the phonetic characters are kana, they are arranged in descending address order according to the Japanese syllabary order) and stored in the storage area of the speech unit database 10.
- the index section IDX stores data for specifying the approximate logical position of the data in the directory section DIR based on the speech unit reading data. Specifically, for example, assuming that the speech unit reading data represents power, the kana character and the ', and the range of addresses of the speech unit reading data whose first character is this kana character are Is stored in association with each other.
- a single nonvolatile memory may perform some or all of the functions of the general word dictionary 2, the user word dictionary 3, the waveform database 7, and the speech unit database 10.
- the storage of the data in the speech unit database 10 is performed by the speech unit registration unit R shown in FIG.
- the speech unit registration unit R includes a recorded speech unit data set storage unit 12, a speech unit database creation unit 13, and a compression unit 14.
- the sound registration The nit R may be removably connected to the speech unit database 10 in this case.
- the speech unit With the registration unit R separated from the main unit M, the main unit M may perform the operation described below.
- the recorded sound piece data set storage unit 12 is composed of a non-volatile memory that can be rewritten overnight, such as a hard disk device.
- the stored sound data storage unit 12 contains phonograms that represent the reading of a sound piece, and sounds that represent the waveform obtained by collecting the actual sound of this sound piece.
- the piece data is stored in association with each other in advance by the manufacturer of the speech synthesis system or the like. It should be noted that the sound piece data may be composed of, for example, digital data converted into PCM.
- the voice unit data base creation unit 13 and the compression unit 14 are composed of a processor such as a CPU, a memory for storing a program to be executed by the processor, and the like. Do.
- a single processor may perform some or all of the functions of the speech unit database creation unit 13 and the compression unit 14. Also, the language processing unit 1, the sound processing unit 4, and the search unit 5 The processor that performs part or all of the functions of the decompression unit 6, the speech unit editing unit 8, the search unit 9, and the speech speed conversion unit 11 further performs the functions of the speech unit database creation unit 13 and compression unit 14. You may. Further, a processor that performs the functions of the speech unit database creation unit 13 and the compression unit 14 may also serve as the control circuit of the recorded speech unit data set storage unit 12.
- the sound piece data set creation section 13 From step 2, the phonetic character and the speech unit data that are associated with each other are read out, and the time change of the frequency of the pitch component of the voice represented by the speech unit data and the utterance speed are specified.
- the utterance speed may be specified, for example, by counting the number of samples in the sound piece.
- the time change of the frequency of the pitch component may be specified by performing cepstrum analysis on the sound piece data, for example.
- the waveform represented by the speech unit data is divided into a number of small parts on the time axis, and the strength of each obtained small part is calculated as the logarithm of the original value (the base of the logarithm is arbitrary) Transforms this small portion of the spectrum (that is, the cepstrum) into a value that is substantially equal to, and generates data that represents the result of the Fourier transform of a discrete variable (or a discrete variable). Any other method).
- the minimum value of the frequencies giving the maximum value of the cepstrum is specified as the frequency of the pitch component in this small portion.
- the time change of the frequency of the pitch component can be calculated, for example, by converting the sound piece data into a pitch waveform data according to the method disclosed in Japanese Patent Application Laid-Open No. 2003-108172. After that, good results can be expected if identification is performed based on this pitch waveform data.
- the pitch signal is extracted by filtering the speech unit data, and the waveform represented by the speech unit data is divided into sections of unit pitch length based on the extracted pitch signal. It is sufficient to specify the phase shift based on the correlation with and to convert the speech unit data into a pitch waveform signal by aligning the phases of the respective sections. Then, the obtained pitch waveform signal is treated as sound piece data, and cepstrum analysis is performed.
- the time change of the frequency of the pitch component may be specified.
- the speech unit database creation unit 13 supplies the speech unit data read from the recorded speech unit data set storage unit 12 to the compression unit 14.
- the compression unit 14 creates compressed speech unit data by entropy-encoding the speech unit data supplied from the speech unit database creation unit 13, and returns it to the speech unit data base creation unit 13.
- the time change of the utterance speed and the frequency of the pitch component of the speech unit data is identified, and this speech unit data is subjected to entropy coding and returned as a compressed speech unit data from the compression unit 14.
- the voice unit data base creation unit 13 writes the compressed voice unit data into the storage area of the voice unit database 10 as data constituting the data unit DAT.
- the speech unit data base creation unit 13 also converts the phonogram read out from the recorded speech unit data set storage unit 12 as an indication of the reading of the speech unit represented by the written compressed speech unit data, and Write the read data to the storage area of the base unit 10 as the read data.
- the head address of the written compressed speech piece data in the storage area of the speech piece database 10 is specified, and this address is written to the storage area of the speech piece database 10 as the above-mentioned (B) data. .
- the data length of the compressed speech piece data is specified, and the specified data length is written to the storage area of the speech piece database 10 as the data (C).
- a data indicating the time change of the utterance speed and the frequency of the pitch component of the voice unit represented by the compressed voice data is generated, and the data is generated as speed initial value data and pitch component data.
- the language processing unit 1 obtains free text data describing a sentence (free text) including ideographic characters prepared by a user as an object for synthesizing a speech in the speech synthesis system from the outside. I do.
- the language processing unit 1 may obtain the free text data by any method.
- the language processing unit 1 may obtain the free text data from an external device network via an interface circuit (not shown), or a recording device (not shown).
- the medium may be read from a recording medium (for example, a floppy (registered trademark) disk or a CD_ROM) set in the drive device via the recording medium drive device.
- the processor performing the function of the language processing unit 1 may transfer text data used in other processing being executed by itself to the processing of the language processing unit 1 as free text data. .
- the language processing unit 1 searches the phonetic character representing the reading of each ideographic character included in the free text by searching the general word dictionary 2 and the user word dictionary 3. Identify. Then, this ideogram is replaced with the specified phonogram. Then, the language processing unit 1 supplies a phonogram string obtained as a result of replacing all ideograms in the free text with phonograms to the sound processing unit 4.
- the sound processing section 4 searches for the waveform of the unit speech represented by the phonogram for each phonogram included in the phonogram string. Instruct the search unit 5.
- the search unit 5 searches the waveform database 7 in response to this instruction, The compressed waveform data representing the waveform of the unit voice represented by each phonetic character included in the phonetic character string is searched for. Then, the retrieved compressed waveform data is supplied to the decompression unit 6.
- the decompression unit 6 restores the compressed waveform data supplied from the search unit 5 to the waveform data before being compressed, and returns it to the search unit 5.
- the search unit 5 supplies the waveform data returned from the decompression unit 6 to the sound processing unit 4 as a search result.
- the sound processing unit 4 converts the waveform data supplied from the search unit 5 into a speech unit editing unit according to the order of each phonogram in the phonogram string supplied from the language processing unit 1.
- Supply to 8. -Upon receiving the waveform data from the sound processing unit 4, the speech unit editing unit 8 combines the waveform data with each other in the order of supply and outputs the combined data as data representing synthesized speech (synthesized speech data). I do.
- This synthesized speech synthesized based on the free text data is equivalent to the speech synthesized by the rule synthesis method.
- the method by which the sound piece editing unit 8 outputs synthesized speech data is arbitrary.
- a D / A (Digital-to-Analog) converter (not shown)
- the synthesized voice represented by the synthesized voice data may be reproduced.
- the data may be sent to an external device network via an interface circuit (not shown), or may be written to a recording medium set in a recording medium drive device (not shown) via the recording medium drive device. May be.
- the processor performing the function of the sound piece editing unit 8 may transfer the synthesized speech data to another process executed by itself.
- the sound processing unit 4 represents the phonetic character string distributed from the outside. It is assumed that data (delivery character string data) has been obtained. (Note that the method by which the sound processing unit 4 acquires the distribution character string data is also arbitrary. For example, the language processing unit 1 may acquire the distribution character string data by the same method as the method of acquiring the free text data. )
- the sound processing unit 4 treats the phonetic character string represented by the distribution character string data in the same manner as the phonetic character string supplied from the language processing unit 1.
- the compressed waveform data corresponding to the phonetic characters included in the phonetic character string represented by the distribution character string data is retrieved by the search unit 5, and the waveform data before compression is retrieved by the decompression unit 6. Will be restored.
- Each of the restored waveform data is supplied to the sound piece editing unit 8 via the sound processing unit 4, and the sound unit editing unit 8 converts this waveform data into each of the phonetic character strings represented by the distribution character string data. Combine them in the order of phonetic characters and output them as synthesized speech data.
- the synthesized speech data synthesized based on the distribution character string data also indicates the speech synthesized by the rule synthesis method.
- the speech piece editing unit 8 has acquired the fixed message data and the utterance speed data.
- the fixed message data is data representing a fixed message as a phonetic character string
- the utterance speed data is a fixed message message—the specified value of the utterance speed of the fixed message represented by the evening (this fixed message is (The specified value of the utterance time length).
- the method by which the sound piece editing unit 8 obtains the fixed message data and the utterance speed data is arbitrary.
- the fixed message data and the utterance are obtained in the same manner as the method by which the language processing unit 1 obtains the free text data. What is necessary is just to acquire speed data.
- the standard message data and utterance speed data are sent to the sound piece editing unit 8.
- the speech unit editing unit 8 searches for all the compressed speech unit data associated with the phonetic characters that match the phonetic characters representing the reading of the speech units included in the fixed message.
- the search unit 9 is instructed.
- the search unit 9 searches the speech unit database 10 in response to the instruction of the speech unit editing unit 8, and finds the corresponding compressed speech unit data and the above-described speech unit reading associated with the corresponding compressed speech unit data. Data, speed initial value data and pitch component data are retrieved, and the retrieved compressed waveform data is supplied to the decompression unit 6. Even when a plurality of compressed speech piece data correspond to one speech piece, all of the corresponding compressed speech piece data are searched for as candidates for data used for speech synthesis. On the other hand, when there is a speech unit for which compressed speech unit data could not be found, the search unit 9 generates data for identifying the corresponding speech unit (hereinafter, referred to as missing portion identification data).
- the decompression unit 6 restores the compressed speech unit data supplied from the retrieval unit 9 to the speech unit data before being compressed, and returns the data to the retrieval unit 9.
- the search unit 9 uses the speech unit data returned from the decompression unit 6 and the retrieved speech unit reading data, the speed initial value data and the pitch component data as search results as speech speed conversion units 1 1 To supply. When the missing part identification data is generated, the missing part identification data is also supplied to the speech speed conversion unit 11.
- the speech unit editing unit 8 converts the speech unit data supplied to the speech speed conversion unit 11 into the speech speed conversion unit 11, and determines the time length of the speech unit represented by the speech unit data. Instruct the user to match the speed indicated by the utterance speed data.
- the speech speed conversion unit 11 responds to the instruction of the speech unit editing unit 8, converts the speech unit data supplied from the search unit 9 so as to match the instruction, and edits the speech unit. Supply to Part 8. Specifically, for example, the original time length of the speech unit data supplied from the search unit 9 is specified based on the searched speed initial value data, and the speech unit data is resampled. Then, the number of samples of the speech piece data may be set to a time length that matches the speed specified by the speech piece editing unit 8.
- the speech speed conversion unit 11 also supplies the speech unit reading data, speed initial value data and pitch component data supplied from the retrieval unit 9 to the speech unit editing unit 8, and retrieves the missing part identification data. If supplied from 9, the missing part identification data is also supplied to the sound piece editing unit 8.
- the speech unit editing unit 8 When the utterance speed data is not supplied to the speech unit editing unit 8, the speech unit editing unit 8 outputs the speech unit supplied to the speech speed conversion unit 11 to the speech speed conversion unit 11. What is necessary is just to instruct the speech unit editing unit 8 to supply the data without converting the data, and the speech speed conversion unit 11 responds to this instruction and outputs the speech unit data supplied from the search unit 9 as it is. It may be supplied to the one-side editing unit 8.
- the speech unit editing unit 8 When the speech unit editing unit 8 is supplied with the speech unit data, the speech unit reading data, the speed initial value data and the pitch component data from the speech speed conversion unit 11, the speech unit editing unit 8 From the above, select one voice unit for each voice unit that represents the waveform that best approximates the waveform of the voice unit that composes the fixed message.
- the speech unit editing unit 8 adds a prosody prediction such as a “Fujisaki model” or “To BI (To neand Break Indices)” to the fixed message represented by the fixed message data.
- a prosody prediction such as a “Fujisaki model” or “To BI (To neand Break Indices)”
- To BI To neand Break Indices
- the speech unit editing unit 8 determines, for each speech unit in the fixed message, that the prediction result data representing the prediction result of the time change of the frequency of the pitch component of the speech unit matches the reading of the speech unit.
- the correlation with pitch component data representing the time change of the frequency of the pitch component of the speech unit data representing the waveform of the speech unit to be performed is obtained.
- the speech piece editing unit 8 calculates, for each of the pitch component data supplied from the speech speed conversion unit 11, for example, a value shown on the right side of Equation 1 and a value shown on the right side of Equation 2] 3 Ask. n '
- Fig. 3 (a) As shown in Fig. 3 (a), as a linear function of the value X (i) (i is an integer) of the i-th sample of the prediction result data (the total number of samples is n) for a certain speech unit First-order regression is performed on the value of the i-th sample Y (i) of pitch component data (the total number of samples is assumed to be n) for the speech unit data representing the waveform of the speech unit whose reading matches that of this speech unit.
- the gradient of this linear function has an intercept of ⁇ .
- the unit of the slope may be, for example, [Hertz / second]
- the unit of intercept / 3 may be, for example, [Hertz].
- the total number of samples differs between the prediction result data and the pitch component data for the same reading speech unit, one (or both) of the two is replaced by linear interpolation, Lagrange interpolation, or any other method. It is sufficient to resample after interpolating by the method, and to obtain the correlation after aligning the total number of both samples.
- the speech unit editing unit 8 uses the speed initial value data supplied from the speech speed conversion unit 11 and the fixed message data and utterance speed data supplied to the speech unit editing unit 8. Then, the value dt on the right side of Equation 3 is obtained.
- This value dt is a coefficient representing the time difference between the utterance speed of the speech unit represented by the speech unit and the utterance speed of the speech unit in the fixed message whose reading matches the reading.
- the evaluation value cost 1 is set to be larger as the correlation between the prediction result of the pitch of the voice unit and the pitch of the voice unit is higher, so that the reciprocal of the linear function of the value i 1 - ⁇ I Therefore, the evaluation value cost 1 becomes larger as the value I 1 - ⁇ I approaches 0.
- the inflection of speech is characterized by the temporal change of the frequency of the pitch component of a speech unit. Therefore, the value of the slope H has the property of reflecting the difference in the intonation of the voice to the sensitivity.
- the value of the intercept i3 is close to 0. Therefore, the value of [intercept] 3 has the property of sensitively reflecting the difference in the base pitch frequency of voice.
- the evaluation value c ost 1 since the evaluation value c ost 1 has a form that can be regarded as the reciprocal of the linear function of the value I iS I, the evaluation value c ost 1 becomes larger as the value I ⁇ I becomes closer to 0.
- the base pitch frequency of a voice is a factor that governs the voice quality of a voice speaker, and the gender difference of the speaker is remarkable.
- the coefficient W Get a value of 2 It is desirable to make it larger.
- the speech unit editing unit 8 selects the speech unit data representing a waveform close to the waveform of the speech unit in the fixed message, while identifying the missing part from the speech speed conversion unit 11. If data is also supplied, a phonetic character string representing the reading of the speech piece indicated by the missing part identification data is extracted from the fixed message data and supplied to the acoustic processing unit 4, and the waveform of this speech piece is obtained. Instruct them to combine.
- the sound processing unit 4 treats the phonetic character string supplied from the speech unit editing unit 8 in the same manner as the phonetic character string represented by the distribution character string data.
- compressed waveform data representing the waveform of the speech indicated by the phonetic characters included in the phonetic character string is retrieved by the search unit 5, and the compressed waveform data is restored by the decompression unit 6 to the original waveform data. It is restored and supplied to the sound processing unit 4 via the search unit 5.
- the sound processing section 4 supplies the waveform data to the sound piece editing section 8.
- the voice data editing unit 8 When the voice processing unit 4 returns the waveform data from the sound processing unit 4, the voice data editing unit 8 outputs the waveform data and the voice data editing unit 8 out of the voice data supplied from the speech speed conversion unit 11. Are combined with each other in the order according to the sequence of the sound pieces in the fixed message indicated by the fixed message data, and output as data representing the synthesized speech.
- the speech unit editing unit 8 immediately specified it without instructing the sound processing unit 4 to synthesize a waveform.
- the speech unit data may be combined with each other in the order according to the sequence of the speech units in the fixed message indicated by the fixed message data, and output as data representing the synthesized speech.
- units larger than phonemes The speech unit data representing the waveform of the speech unit can be naturally spliced by the recording and editing method based on the prediction result of the prosody, and the speech that reads out the fixed message is synthesized.
- the storage capacity of the speech unit database 10 can be made smaller than in the case where a waveform is stored for each phoneme, and a search can be performed at high speed. Therefore, this speech synthesis system can be configured to be small and lightweight, and can follow high-speed processing.
- the correlation between the predicted result of the speech unit waveform and the speech unit data is evaluated using multiple evaluation criteria (for example, evaluation based on the gradient or intercept in the case of first-order regression, evaluation based on the time difference of the speech unit, etc.).
- evaluation criteria for example, evaluation based on the gradient or intercept in the case of first-order regression, evaluation based on the time difference of the speech unit, etc.
- discrepancies in the results of these evaluations can often occur.
- the results of evaluations based on multiple evaluation criteria are integrated based on one evaluation value, and appropriate evaluation is performed.
- waveform data ⁇ speech piece data does not need to be in PCM format data, and the data format is arbitrary.
- the waveform database 7 ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ speech data base 10 does not necessarily need to store the waveform data ⁇ speech data in a compressed state. Waveform data base 7 ⁇ Speech data base 10 If waveform data ⁇ Speech data is stored in an uncompressed state, it is not necessary for main unit M to have decompression unit 6. Absent.
- the speech unit database creation unit 13 performs a new compression from the recording medium set in the recording medium drive unit (not shown) to be added to the speech unit data base 10 via the recording medium drive unit. It is also possible to read the sound piece data or the phonetic character string which is the material of the sound piece data.
- the speech unit registration unit R is not necessarily the recorded speech unit data set. It is not necessary to have a memory 12.
- the speech unit editing unit 8 stores in advance a prosody registration data representing the prosody of a specific speech unit, and if the specific message unit is included in the fixed message, the prosody registration data is stored.
- the prosody represented may be treated as the result of prosody prediction.
- the sound piece editing unit 8 may newly store the result of the past prosody prediction as prosody registration data.
- the speech piece editing unit 8 calculates, for each pitch component data supplied from the speech speed conversion unit 11, for example, the value RX y shown on the right side of Expression 5 (j) is determined by taking the value of j as an integer from 0 to less than n to obtain a total of n, and among the n correlation coefficients from R xy (0) to R y (n-1) obtained, The maximum value may be specified. [ ⁇ X (i) -mx ⁇ ⁇ ⁇ Y j (i) -my ⁇ ]
- R xy (j) is the prediction result for a certain sound piece (the total number of samples is n.
- X (i) in Equation 5 is the same as that in Equation 1).
- a sequence of samples obtained by cyclically shifting the pitch component data (total number of samples n) of the speech unit data representing the waveform of the matching speech unit by j in a fixed direction (note that Y j (i ) Is the value of the i-th sample in this sample column.)
- FIG. 3 (b) is a graph showing an example of values of prediction result data and pitch component data used for obtaining values of R xy (0) and R xy (j).
- the speech unit editing unit 8 performs a speech unit decoding process that represents a speech unit that matches the reading of the speech unit in the fixed message.
- the value on the right-hand side of Equation 6 (evaluation value) that has the largest cost 2 may be selected.
- Rm ax is the maximum value among the R xy (0) ⁇ R xy ( n one 1)
- the sound piece editing unit 8 does not necessarily need to obtain the above-described correlation coefficient for the pitch component data obtained by performing various cyclic shifts.
- the value of R xy (0) is used as the maximum value of the correlation coefficient as it is. It may be handled.
- the evaluation values co s t 1 and c os t 2 may not include the term of the coefficient d t, and in this case, the speech piece editing unit 8 does not need to find the coefficient d t.
- the speech unit editing unit 8 may use the value of the coefficient dt as it is as the evaluation value.
- the speech unit editing unit may use the gradient ⁇ , the intercept] 3, There is no need to find the value of R xy (j).
- the pitch component data may be data representing a temporal change of the pitch length of the sound piece represented by the sound piece data.
- the speech unit editing unit 8 creates, as the prediction result data, data representing the prediction result of the time change of the pitch length of the speech unit, and generates the waveform of the speech unit whose reading matches the reading of the speech unit.
- the correlation with the pitch component data representing the temporal change of the pitch length of the sound element data may be obtained.
- the sound piece database creation unit 13 may include a microphone, an amplifier, a sampling circuit, an A / D (Analog-to-Digital) converter, a PCM encoder, and the like.
- the sound piece data creation base section 13 instead of acquiring the sound piece data from the recorded sound piece data set storage section 12, the sound piece data creation base section 13 generates a sound representing the sound collected by its own microphone. After amplifying the signal, sampling and A / D converting it, PCM modulation may be applied to the sampled audio signal to create a sound unit data.
- the speech unit editing unit 8 supplies the waveform data returned from the sound processing unit 4 to the speech speed conversion unit 11 so that the speech speed data indicates the time length of the waveform represented by the waveform data. It may be made to match with the one.
- the speech unit editing unit 8 acquires free text data together with the language processing unit 1, for example, and obtains a speech unit data representing a waveform closest to the waveform of the speech unit included in the free text represented by the free text data. May be selected by performing substantially the same processing as the processing of selecting the speech piece data representing the waveform closest to the speech piece waveform included in the fixed message, and used for speech synthesis. ' In this case, the sound processing unit 4 does not need to cause the search unit 5 to search for waveform data representing the waveform of the speech unit represented by the speech unit selected by the speech unit editing unit 8. .
- the sound piece editing unit 8 notifies the sound processing unit 4 of a sound unit that does not need to be synthesized by the sound processing unit 4, and the sound processing unit 4 responds to this notification to respond to the notification by rewriting the unit speech constituting the sound unit. What is necessary is to stop the search of the waveform of.
- the speech unit editing unit 8 acquires distribution character string data together with, for example, the acoustic processing unit 4, and generates a speech unit representing a waveform closest to the waveform of the speech unit included in the distribution character string represented by the distribution character string data. It is also possible to select data by performing processing that is substantially the same as the processing of selecting a speech unit that represents the waveform closest to the waveform of the speech unit included in the fixed message, and use it for speech synthesis. Good. In this case, the sound processing unit 4 does not need to cause the search unit 5 to search for waveform data representing the waveform of the sound unit represented by the sound unit data selected by the sound unit editing unit 8. .
- the physical configuration of the speech synthesis system according to the second embodiment of the present invention is substantially the same as the configuration in the above-described first embodiment.
- the directory section DIR of the speech unit database 10 in the speech synthesis system includes, for example, as shown in FIG.
- the data of A) to (D) are stored in association with each other, and instead of the data of (E) described above, (F) the compressed sound piece data is stored as pitch component data.
- the data representing the pitch component frequencies at the beginning and end of the sound piece to be expressed are stored in a form associated with these (A) to (D) data. Have been.
- Fig. 4 shows the data included in the data section DAT, which represents the waveform of a sound piece whose reading is "Saitama," as in Fig. 2.
- the one-side data is stored at a logical position starting from the address 01 A36A6h.
- at least the data of (A) in the above set of data of (A) to (D) and (F) are sorted according to the order determined based on the phonetic characters represented by the phoneme reading data. It is assumed that it is stored in the storage area of the sound piece database 10 in a single state.
- the speech unit database creation unit 13 of the speech unit registration unit R reads out the phonograms and the speech unit data that are associated with each other from the recorded speech unit data set storage unit 12, and The utterance speed of the voice represented by the piece of data and the frequency of the pitch component at the beginning and end shall be specified.
- the read speech piece data is supplied to the compression section 14, and when the compressed speech piece data is returned, the compressed speech piece data, the phonogram read out from the recorded speech piece data set storage section 12,
- the first address in the storage area of the speech unit database 10 of the compressed speech unit data, the data length of the compressed speech unit data, and the speed initial value data indicating the specified utterance speed are stored in the first
- the data is written in the storage area of the speech unit database 10, and data indicating the result of specifying the frequency of the pitch component at the beginning and end of the sound is stored. It is generated and written in the storage area of the speech piece database 10 as pitch component data.
- the utterance speed and the frequency of the pitch component are specified, for example, as follows:
- the method may be performed by substantially the same method as the method performed by the sound piece database creating unit 13 of the first embodiment.
- the operation of the speech synthesis system of the first embodiment is as follows.
- the operation is substantially the same as the operation to be performed.
- the method by which the language processing unit 1 obtains free text data and the method by which the sound processing unit 4 obtains distribution character string data are arbitrary.
- both methods are used in the language processing unit in the first embodiment.
- Free text data or distribution character string data may be obtained by the same method as that performed by 1 and the sound processing unit 4.
- the speech piece editing unit 8 has acquired the fixed message data and the utterance speed data.
- the method by which the sound piece editing unit 8 acquires the fixed message data and the utterance speed data is also arbitrary.
- the fixed message data may be obtained by the same method as the method performed by the sound unit editing unit 8 in the first embodiment. Or utterance speed data.
- the speech unit editing unit 8 When the standard message data and the utterance speed data are supplied to the speech unit editing unit 8, the speech unit editing unit 8 is included in the standard message, similarly to the speech unit editing unit 8 in the first embodiment.
- the search unit 9 is instructed to search all the compressed speech unit data associated with the phonogram corresponding to the phonogram representing the reading of the speech unit.
- the speech unit conversion unit 11 converts the speech unit data supplied to the speech speed conversion unit 11 in the same manner as the speech unit editing unit 8 in the first embodiment, and converts the speech unit.
- the time length of the speech unit represented by the data matches the speed indicated by the utterance speed data. Instructs them to do so.
- the search unit 9, the expansion unit 6, and the speech speed conversion unit 11 perform substantially the same operation as the operation of the search unit 9, the expansion unit 6, and the speech speed conversion unit 11 in the first embodiment.
- the speech unit data, the speech unit reading data, and the pitch component data are supplied from the speech speed conversion unit 11 to the speech unit editing unit 8.
- the missing part identification data is supplied from the search unit 9 to the speech speed conversion unit 11, the missing part identification data is also supplied to the speech piece editing unit 8.
- the speech unit editing unit 8 When the speech unit editing unit 8 is supplied with the speech unit data, the speech unit reading data, and the pitch component data from the speech speed conversion unit ⁇ 1, the speech unit editing unit 8 performs a fixed form from among the supplied speech unit data according to the procedure described below. Select one piece of voice data for each voice piece that represents the waveform that can be regarded as the waveform of the voice piece that composes the message.
- the speech unit editing unit 8 first and last of the speech unit data supplied from the speech speed conversion unit 11 Specify the frequency of the pitch component at each point in time. Then, from the speech unit data supplied from the speech speed conversion unit 11, the absolute value of the frequency difference of the pitch component at the boundary between adjacent speech units in the fixed message is accumulated for the entire fixed message. Select the voice unit to satisfy the condition that the resulting value is minimized.
- FIGS. 5 (a) to 5 (d) The conditions for selecting the voice unit will be described with reference to FIGS. 5 (a) to 5 (d).
- a fixed message data representing a fixed message reading "This Saki Mikika-bu-da” is supplied to the sound piece editing unit 8, and
- the fixed message consists of three sound pieces, "Konosaki”, “Migikaichi” and "is” Shall be.
- the speech unit de-even base 10 has three compressed speech unit data whose readings are "Konosaki"("A" in Fig. 5 (b)).
- Fig. 5 (c) shows, for example, the difference between the frequency of the pitch component at the end of the speech unit represented by the speech unit A1 and the frequency of the pitch component at the beginning of the speech unit represented by the speech unit data B1. This indicates that the absolute value is “1 2 3.”
- the unit of this absolute value is, for example, “Hertz”.
- the speech piece editing unit 8 selects the speech piece data A3, B2, and C2 as shown in FIG. 5 (d).
- the speech unit editing unit 8 sets, for example, the absolute value of the frequency difference between the pitch components at the boundary between adjacent speech units in the fixed message as the distance. What is necessary is just to define and select the sound piece by the DP (Dynamic Programming) matching method.
- the speech unit editing unit 8 converts the phonetic character string representing the reading of the speech unit indicated by the missing portion identification data into a fixed message.
- the data is extracted from the data and supplied to the acoustic processing unit 4 to instruct to synthesize the waveform of the sound piece.
- the sound processing unit 4 treats the phonetic character string supplied from the speech unit editing unit 8 in the same manner as the phonetic character string represented by the distribution character string data.
- compressed waveform data representing the waveform of the voice indicated by the phonogram contained in the phonogram string is retrieved by the search unit 5, and the compressed waveform data is converted into the original waveform data by the decompression unit 6. Is restored and supplied to the sound processing unit 4 via the search unit 5.
- the sound processing section 4 supplies the waveform data to the sound piece editing section 8.
- the voice unit editing unit 8 selects the waveform data and the voice unit editing unit 8 out of the voice unit data supplied from the speech speed conversion unit 11. These are combined with each other in the order according to the sequence of the sound pieces in the fixed message indicated by the fixed message data, and output as data representing the synthesized speech.
- the sound processing unit is used as in the first embodiment.
- the speech unit data selected by the speech unit editing unit 8 is immediately combined with the sequence of each of the speech units in the standard message indicated by the standard message without immediately instructing the synthesis of the waveform in 4. Then, it may be output as a data representing synthesized speech.
- the cumulative total of the amount of the discontinuous change in the frequency of the pitch component at the boundary between the speech units is minimized in the entire fixed message.
- the sound unit is selected so that it can be connected naturally by the recording and editing method, so that the synthesized speech becomes natural.
- this speech synthesis system does not perform prosody prediction with complicated processing, and can follow high-speed processing with a simple configuration.
- the configuration of the speech synthesis system according to the second embodiment is not limited to the configuration described above.
- the pitch component data may be data representing a pitch length at the beginning and end of the sound piece represented by the sound piece data.
- the speech unit editing unit 8 determines the pitch length at the beginning and end of each speech unit data supplied from the speech speed conversion unit 11 based on the pitch component data supplied from the speech speed conversion unit 11. If the speech unit data is selected so as to satisfy the condition that the absolute value of the pitch length difference at the boundary between adjacent speech units in the fixed message is minimized over the entire fixed message, Good.
- the speech unit editing unit 8 acquires free text data together with the language processing unit 1, and determines the speech unit data representing a waveform that can be regarded as a waveform of a speech unit included in the free text represented by the free text data. By extracting the speech unit data representing the waveform that can be regarded as the waveform of the speech unit included in the type message. And may be used for speech synthesis.
- the sound processing unit 4 does not need to cause the search unit 5 to search for waveform data representing the waveform of the speech unit represented by the speech unit data extracted by the speech unit editing unit 8. .
- the sound piece editing unit 8 notifies the sound processing unit 4 of a sound unit that does not need to be synthesized by the sound processing unit 4, and the sound processing unit 4 responds to this notification to respond to the notification by rewriting the unit speech constituting the sound unit. What is necessary is to stop the search of the waveform of.
- the sound piece editing unit 8 acquires distribution character string data together with, for example, the sound processing unit 4, and generates a sound representing a waveform that can be regarded as a waveform of a sound unit included in the distribution character string represented by the distribution character string data.
- the segment data is extracted by performing substantially the same processing as that for extracting the speech unit data representing the waveform that can be regarded as the waveform of the speech unit included in the fixed message, and can be used for speech synthesis. Good.
- the sound processing unit 4 does not need to cause the search unit 5 to search for waveform data representing the waveform of the sound unit represented by the sound unit data extracted by the sound unit editing unit 8. .
- the physical configuration of the speech synthesis system according to the third embodiment of the present invention is substantially the same as the configuration in the above-described first embodiment.
- the operation when the language processing unit 1 of the speech synthesis system obtains the free text data from outside and the sound processing unit 4 obtains the distribution character string data from the outside are described in the first or second implementation.
- the operation is substantially the same as that performed by the speech synthesis system of the first embodiment.
- the language processing unit 1 acquires free text data and the sound processing unit 4 uses The method of acquiring the evening is arbitrary.
- the method of obtaining free text data or the delivery character by the same method as the method performed by the language processing unit 1 or the sound processing unit 4 in the first or second embodiment is used. All you have to do is get the column data.
- the speech piece editing unit 8 has acquired the fixed message data and the utterance speed data.
- the method by which the sound piece editing unit 8 acquires the fixed message data—evening and utterance speed data is also optional.
- the sound pattern editing unit 8 uses the same method as the method performed by the sound unit editing unit 8 in the first embodiment. What is necessary is just to acquire message data and utterance speed data.
- the speech synthesis system is a part of an in-vehicle system such as a force navigation system, and other devices constituting the in-vehicle system (for example, performing speech recognition and performing speech recognition).
- the Device that performs agent processing based on the information obtained as a result of recognition), determines the content and speed of speech to the user, and generates data representing the result of the determination.
- the speech synthesis system may receive (acquire) the generated data and handle the data as fixed message data and utterance speed data.
- the sound unit editing unit 8 is included in the fixed message, similarly to the sound unit editing unit 8 in the first embodiment.
- the search unit 9 is instructed to search for all the compressed speech unit data associated with the phonetic characters that match the phonetic characters representing the reading of the speech unit.
- the speech rate conversion section 11 converts the speech piece data supplied to the speech rate conversion section 11 in the same manner as the speech piece editing section 8 in the first embodiment, and converts the speech The time length of the sound piece represented by the piece data matches the speed indicated by the utterance speed data. Instructs them to do so.
- the search unit 9, the expansion unit 6, and the speech speed conversion unit 11 perform substantially the same operation as the operation of the search unit 9, the expansion unit 6, and the speech speed conversion unit 11 in the first embodiment.
- the speech speed conversion unit 11 sends the speech unit editing unit 8 to the speech unit data, the speech unit reading data, the speed initial value data representing the speech speed of the speech unit represented by the speech unit data, and the pitch. Ingredient data is provided. Further, when the missing portion identification data is supplied from the search unit 9 to the speech speed conversion unit 11, the missing portion identification data is also supplied to the speech unit editing unit 8.
- the speech unit editing unit 8 When the speech unit editing unit 8 receives the speech unit data, the speech unit reading data, and the pitch component data from the speech speed conversion unit 11, the speech unit editing unit 8 performs the above-described processing on each pitch component data supplied from the speech speed conversion unit 11. Using the initial value of the speed and the fixed message data and the utterance speed data supplied to the sound piece editing unit 8, Find the value dt.
- the speech unit editing unit 8 calculates, for each of the speech unit data supplied from the speech speed conversion unit 11, the ⁇ of the speech unit data obtained by itself (hereinafter, referred to as the speech unit data X). , ⁇ , Rmax, and dt, and the sound data (hereinafter referred to as sound data Y) that represents the sound data adjacent to the sound data represented by the sound data in the fixed message.
- the evaluation value ⁇ ⁇ shown in Expression 7 is specified based on the frequency of the pitch component.
- H XY (W A -cost-A) + (W B -cost-B) + (W c -cost-C)
- the value cost—A included in the right-hand side of Equation 7 is the pitch component of the pitch component at the boundary between the speech unit represented by speech unit data X and the speech unit represented by speech unit data Y, which are adjacent to each other in the fixed message.
- the speech unit editing unit 8 determines the cost-A value based on the pitch component data supplied from the speech speed conversion unit 11 in order to identify the value of cost-A. 11.
- the frequency of the pitch component at each of the beginning and end of each piece of sound piece data supplied from 1 may be specified.
- the value c os t—B included in the right-hand side of Expression 7 is a value obtained when the evaluation value c os t—B is obtained for the voice unit X according to Expression 8.
- cost _B 1 / (W B 1 I 1-a I + W B 2 I ⁇ I + W B 3
- the value co st — C included in the right-hand side of Equation 7 is a value obtained when the evaluation value co st — C of the voice unit X is calculated according to Equation 9.
- cost _C 1 / (W c 1 I Rm ax I + W C 2 -dt)
- the speech piece editing section 8 in place of Equation 7 to Equation 9, may be specified evaluation value Eta chi gamma according to Equation 1 0) and (1 1.
- any value of the coefficient W B 3 and W c 3 above is 0.
- the terms (W B 3 ⁇ dt) and (W C 2 ⁇ dt) in Equations 8 and 9 need not be provided.
- ⁇ ⁇ ( ⁇ -cost— A) + (W B -cosf _B) + (W c- cost—C) + (W D -cost—D)
- W D is a predetermined coefficient that is not 0
- W d is a predetermined coefficient that is not 0
- the speech unit editing unit 8 generates, from among the speech unit data supplied from the speech speed conversion unit 11, the speech unit 1 forming the standard message represented by the standard message data supplied to the speech unit editing unit 8.
- the speech unit editing unit 8 synthesize the speech that reads out the fixed message with the largest sum of the evaluation value ⁇ ⁇ of each piece of speech piece data belonging to the combination To select the best combination of sound pieces for the night.
- a fixed message representing a fixed message message is composed of speech pieces A, B, and C, and is used as a candidate for speech piece data representing speech piece A.
- A1, A2, and A3 are found, and speech unit data B1 and B2 are found as candidates for speech unit data representing speech unit B, and speech units are found as candidates for speech unit data representing speech unit C.
- the combination includes a speech unit data P representing the speech unit p and a speech unit data Q representing the speech unit Q.
- the speech unit P precedes the speech unit q.
- the evaluation value HpQ when the speech unit p is adjacent to the speech unit is used.
- the sound piece editing unit 8 treats the value of (W A ⁇ cost—A) as being 0, while The values of the coefficients W B , W c, and W D are each treated as a predetermined value different from the case of calculating the evaluation value ⁇ ⁇ ⁇ of other sound piece data.
- the speech unit editing unit 8 calculates the evaluation value indicating the relationship between the speech unit data X and the adjacent speech unit data Y before the speech unit represented by the speech unit data X using Expression 7 or Expression 11.
- the evaluation value H XY may be specified as including XY . In this case, the value of cost- ⁇ cannot be determined because there is no preceding speech unit for the first speech unit of the fixed message.
- the sound piece editing unit 8 treats the value of (W ⁇ ⁇ cost- ⁇ ) as being 0, while The values of the coefficients W B , W c, and W D may be respectively treated as predetermined values different from the case of calculating the evaluation value ⁇ ⁇ of other sound piece data.
- the speech piece editing unit 8 outputs the missing part identification data from the speech speed conversion unit 11. Is also supplied, the phonetic character string representing the reading of the speech piece indicated by the missing part identification data is extracted from the fixed message data and supplied to the acoustic processing unit 4, and the waveform of this speech piece is synthesized. To do so.
- the sound processing unit 4 treats the phonetic character string supplied from the speech unit editing unit 8 in the same manner as the phonetic character string represented by the distribution character string data.
- compressed waveform data representing the voice waveform indicated by the phonetic characters included in this phonetic character string is retrieved by the search unit 5, and the compressed waveform data is restored to the original waveform data by the decompression unit 6.
- the data is supplied to the sound processing unit 4 via the search unit 5.
- the sound processing section 4 supplies the waveform data to the sound piece editing section 8.
- Speech piece editing section 8 when it is sent back to the waveform data from the acoustic processing unit 4, and the waveform data, speech speed converting section 1 1 supplied speech piece data sac Chi than, the sum of the evaluation values Eta chi gamma
- the data that represents the synthesized speech is combined with the one that belongs to the combination selected by the speech unit editing unit 8 as the largest combination in the order of each speech unit in the standard message indicated by the standard message Is output as
- the sound processing unit 4 must be instructed to synthesize a waveform, as in the first embodiment.
- the speech unit selected by the speech unit editing unit 8 is combined with each other in the order according to the sequence of each speech unit in the standard message indicated by the standard message data, and output as data representing the synthesized voice do it.
- the speech units are spliced together naturally by the recording and editing method, and the speech for reading the fixed message is synthesized.
- One piece of sound The storage capacity of 10 can be smaller than that of storing a waveform for each phoneme, and a high-speed search can be performed. Therefore, this speech synthesis system can be configured to be small and lightweight, and can follow high-speed processing.
- various evaluation criteria for evaluating the appropriateness of a combination of speech piece data selected for synthesizing speech for reading a fixed message are provided.
- the configuration of the speech synthesis system according to the third embodiment is not limited to the configuration described above.
- the evaluation values used by the speech unit editing unit 8 to select the optimum combination of speech unit data are not limited to those shown in Expressions 7 to 13, and are obtained by combining the speech units represented by the speech unit data with each other. It may be any value that represents an evaluation of how similar or dissimilar the speech that is made to the speech uttered by a person.
- evaluation expression representing the evaluation value necessarily an expression
- the evaluation expression is not limited to those included in ⁇ 13, and the evaluation expression can be obtained by arbitrarily setting parameters representing the characteristics of the sound unit represented by the sound unit data, or by combining the sound units together.
- a mathematical expression including an arbitrary parameter indicating a feature of the voice or an optional parameter indicating a feature expected to be included in the voice when a person utters the voice is used. May be.
- the criterion for selecting the optimal combination of speech unit data does not necessarily need to be one that can be expressed in the form of an evaluation value, and is obtained by combining the speech units represented by the speech unit data with each other. Any criteria are possible as long as the criteria lead to the determination of the optimal combination of speech piece data based on an evaluation of how similar or different the speech is to human speech.
- the speech unit editing unit 8 acquires free text data together with the language processing unit 1, for example, and generates speech unit data representing a waveform that can be regarded as a waveform of a speech unit included in the free text represented by the free text data. Alternatively, it may be extracted by performing substantially the same processing as the processing of extracting speech piece data representing a waveform that can be regarded as a speech piece waveform included in a fixed message, and used for speech synthesis. In this case, the sound processing unit 4 does not need to cause the search unit 5 to search for waveform data representing the waveform of the sound unit represented by the sound unit data extracted by the sound unit editing unit 8.
- the sound piece editing unit 8 notifies the sound processing unit 4 of a sound unit that does not need to be synthesized by the sound processing unit 4, and the sound processing unit 4 responds to the notification to generate a unit constituting the sound unit.
- the search for the audio waveform may be stopped.
- the sound piece editing unit 8 acquires the distribution character string data together with the sound processing unit 4, for example, and generates the sound unit data representing the waveform that can be regarded as the waveform of the sound unit included in the distribution character string represented by the distribution character string data. May be extracted by performing substantially the same processing as the processing of extracting speech piece data representing a waveform that can be regarded as the waveform of the speech piece included in the fixed message, and used for speech synthesis.
- the sound processing unit 4 performs, for the sound unit represented by the sound unit data extracted by the sound unit editing unit 8, a waveform representing the waveform of the sound unit. It is not necessary for the search unit 5 to search for shape data.
- the audio data selecting device can be realized using a normal computer system, not a dedicated system.
- a language processing unit 1 For example, in a personal computer, a language processing unit 1, a general word dictionary 2, a user word dictionary 3, a sound processing unit 4, a search unit 5, an expansion unit 6, a waveform database 7, a speech unit in the first embodiment described above.
- a medium CD-ROM, MII, floppy (registered trademark) disk, etc.
- a program for causing a personal computer to execute the operations of the recorded sound piece data set storage section 12, the sound piece data base creation section 13 and the compression section 14 in the first embodiment described above By installing the program from the medium storing the ram, the personal computer can perform the function of the sound piece registration unit R of the above-described first embodiment.
- a personal computer that executes these programs and functions as the main unit unit M and the speech unit registration unit R in the first embodiment is executed as a process corresponding to the operation of the speech synthesis system in FIG.
- the processing shown in FIGS. 6 to 8 is to be performed.
- FIG. 6 is a flowchart showing processing when the personal computer acquires free text data.
- Fig. 7 shows that this personal computer 9 is a flowchart showing a process when the information is obtained.
- FIG. 8 is a flowchart showing a process when the personal computer acquires the fixed message data and the utterance speed data.
- step S101 when the personal computer obtains the above-described free text data from outside (FIG. 6, step S101), for each ideographic character included in the free text represented by the free text data, The phonogram representing the reading is specified by searching the general word dictionary 2 and the user word dictionary 3, and the ideogram is replaced with the specified phonogram (step S102).
- the method by which the personal computer acquires the free text data is arbitrary.
- each phonogram included in the phonogram string is obtained.
- the waveform of the unit speech represented by the phonetic character is searched from the waveform database 7, and the compressed waveform data representing the waveform of the unit speech represented by each phonetic character included in the phonetic character string is retrieved ( Step S103).
- the personal computer restores the extracted compressed wave data to the waveform data before compression (step S104), and converts the restored waveform data into a phonetic character string.
- the phonograms in the sequence are combined with each other in the same order and output as synthesized speech data (step S105).
- the method by which this personal computer outputs synthesized speech data is arbitrary.
- this personal computer receives the above-mentioned distribution statement from outside.
- the character string data is obtained by an arbitrary method (FIG. 7, step S201)
- the unit represented by the phonetic character The voice waveform is searched from the waveform database 7, and compressed waveform data representing the unit voice waveform represented by each phonogram included in the phonogram string is retrieved (step S202).
- the personal computer restores the extracted compressed waveform data to the waveform data before compression (step S203), and converts the restored waveform data into a phonetic character string.
- the phonograms in the sequence are combined with each other in the same order, and are output as synthesized speech data by the same processing as the processing in step S105 (step S204).
- step S301 when the personal computer obtains the above-mentioned fixed message data and the utterance speed data from an external device by any method (FIG. 8, step S301), first, the fixed message data is obtained. All compressed speech piece data associated with phonograms that match the phonetic readings contained in the fixed message included in the evening message are retrieved (step S302).
- step S302 the above-mentioned speech piece reading data, speed initial value data, and pitch component data associated with the corresponding compressed speech piece data are also retrieved. If more than one piece of compressed speech data corresponds to a single speech piece, search for the entire compressed speech piece data. On the other hand, when there is a speech unit for which compressed speech unit data cannot be found, the above-described missing portion identification data is generated.
- the personal computer restores the retrieved compressed speech piece data to the speech piece data before being compressed (step S303). Then, the reconstructed speech unit data is converted by the same processing as that performed by the speech unit editing unit 8 described above, and the time length of the speech unit represented by the speech unit data matches the speed indicated by the utterance speed data. (Step S304). When the utterance speed data is not supplied, the restored speech piece data need not be converted.
- the personal computer converts the speech unit data representing the waveform closest to the waveform of the speech unit constituting the fixed message from the speech unit data in which the time length of the speech unit has been converted to the above-described speech unit.
- one sound piece is selected one by one (steps S305 to S308).
- the personal computer predicts the prosody of the fixed message by adding an analysis based on the prosody prediction method to the fixed message represented by the fixed message data (step S305). Then, for each of the sound pieces in the fixed message, a prediction result of the time change of the frequency of the pitch component of this sound piece and a sound piece data representing the waveform of the sound piece whose reading matches that of this sound piece.
- the correlation with the pitch component data representing the time change of the frequency of the pitch component of is obtained (step S306). More specifically, for each of the retrieved pitch component data, for example, the values of the above-described gradient a and intercept] 3 are obtained.
- the personal computer obtains the above value d t using the retrieved initial speed value data, the fixed message data and the utterance speed data obtained from outside (step S 30).
- the personal computer calculates the value of ⁇ obtained in step S306 and the value of 'dt' obtained in step S307. Then, a speech piece data representing a speech piece that matches the reading of the speech piece in the fixed message is selected from those having the largest evaluation value c st1 (step S308).
- this personal computer may determine the above-described maximum value of R Xy (j) in step S306 instead of obtaining the values of ⁇ and 3 described above.
- step S308 based on the maximum value of RXy (j) and the coefficient dt obtained in step S307, the speech unit that matches the reading of the speech unit in the fixed message is used.
- the one that maximizes the above-mentioned evaluation value c 0 st 2 may be selected from the sound elements that represent
- the personal computer extracts a phonetic character string representing the reading of the sound piece indicated by the missing part identifying data from the fixed message data, and generates a phoneme for this phonetic character string.
- Each phonetic character in this phonetic character string is treated in the same way as the phonetic character string represented by the distribution character string data and processed in steps S202 to S203 described above.
- the waveform data representing the waveform of the indicated voice is restored (step S309).
- the personal computer compares the restored waveform data and the sound piece data selected in step S308 in the order according to the order of each sound piece in the fixed message indicated by the fixed message message. These are combined with each other and output as data representing synthesized speech (step S310).
- the personal computer has a language processing unit 1, a general word dictionary 2, a user word dictionary 3, a sound processing unit 4, a search unit 5, a decompression unit 6, and a waveform data base in the second embodiment described above. 7, the operation of the speech unit editing unit 8, the search unit 9, the speech unit data base 10 and the speech speed conversion unit 11 'is executed.
- a program for causing a personal computer to execute the operations of the recorded speech unit data set storage unit 12, the speech unit database creation unit 13 and the compression unit 14 in the second embodiment described above By installing the program from the medium storing the sound unit, the personal computer can perform the function of the sound piece registration unit R in the above-described second embodiment.
- a personal computer that executes these programs and functions as the main unit unit M and the speech unit registration unit R in the second embodiment is executed as processing corresponding to the operation of the speech synthesis system in FIG.
- the above-described processing shown in FIGS. 6 and 7 is performed, and the processing shown in FIG. 9 is performed.
- FIG. 9 is a flowchart showing a process when the personal computer acquires the fixed message data and the utterance speed data.
- step S402 when the personal computer obtains the above-mentioned fixed message data and utterance speed data from an external device by an arbitrary method (FIG. 9, step S401), first, the above-described step S302 is performed. Similarly to the processing, the compressed speech piece data in which the phonogram matching the phonogram representing the reading of the speech piece included in the fixed message represented by the fixed message data is associated with the corresponding compressed speech piece data. The above-mentioned speech unit reading data and speed initial value data And all the pitch component data are retrieved (step S402). Note that, even in step S402, if one compressed sound piece data is extracted from one sound piece, the entire compressed sound piece data is searched for. If there is a speech piece that could not be found, the above-described missing portion identification data is generated.
- the personal computer restores the retrieved compressed speech data to the original speech data before compression (step S403), and restores the restored speech data. Then, conversion is performed by the same processing as that performed by the speech unit editing unit 8 described above, and the time length of the speech unit represented by the speech unit data is matched with the speed indicated by the utterance speed data (step S404) ). If the utterance speed data is not supplied, the restored speech piece data need not be converted.
- the personal computer converts the speech unit data representing the waveform that can be regarded as the waveform of the speech unit constituting the fixed message from the speech unit data obtained by converting the time length of the speech unit into the second unit described above.
- the personal computer converts the speech unit data representing the waveform that can be regarded as the waveform of the speech unit constituting the fixed message from the speech unit data obtained by converting the time length of the speech unit into the second unit described above.
- this personal computer calculates the pitch component frequency at each of the beginning and end of each piece of sound piece data in which the time length of the sound piece has been converted. It is specified based on the evening (step S405). Then, the condition that the sum of the absolute values of the frequency differences of the pitch components at the boundaries between adjacent sound units in the fixed message in the fixed message among these sound unit data is minimized.
- the speech piece data is selected so as to satisfy the condition (step S406).
- the absolute value of the difference between the frequencies of the pitch components at the boundary between adjacent speech units in the fixed message is defined as the distance, and the DP unit is used to select the speech unit by the DP matching method. Just fine.
- the personal computer when the personal computer generates the missing part identification data, the personal computer extracts a phonetic character string representing the reading of the sound piece indicated by the missing part identifying data from the standard message and reads the phonetic character string.
- Each phoneme is treated in the same manner as the phonetic character string represented by the delivery character string data, and the processing in steps S202 to S203 described above is performed, whereby each table in the phonetic character string is processed.
- the waveform data representing the waveform of the voice indicated by the phonetic character is restored (step S407).
- the personal computer compares the restored waveform data and the sound piece data selected in step S406 with the order of each sound piece in the fixed message indicated by the fixed message message. And output as data representing the synthesized speech (step S408).
- a personal computer is provided with a language processing unit 1, a general word dictionary 2, a user word dictionary 3, a sound processing unit 4, a search unit 5, a decompression unit 6, a waveform data base 7 according to the third embodiment.
- a medium storing a program for causing a personal computer to execute the operations of the recorded speech unit data set storage unit 12, the speech unit database creation unit 13 and the compression unit 14 in the third embodiment described above.
- the personal computer can perform the function of the sound piece registration unit R in the third embodiment described above.
- a personal computer that executes these programs and functions as the main unit unit M and the speech unit registration unit R in the third embodiment performs processing corresponding to the operation of the speech synthesis system in FIG.
- the above-described processing shown in FIGS. 6 and 7 is performed, and the processing shown in FIG. 10 is performed. .
- FIG. 10 is a flowchart showing a process when the personal computer obtains the fixed message data and the utterance speed data.
- step S501 when the personal computer obtains the above-mentioned fixed message data and the utterance speed data from the outside by any method (FIG. 10, step S501), first, the above-mentioned step S501 is executed. Similarly to the process of 302, a compressed speech unit decoder in which a phonogram matching a phonogram representing a reading of a speech unit included in the fixed message represented by the fixed message data is associated with the compressed message unit. All the above-mentioned speech piece reading data, speed initial value data and pitch component data associated with the corresponding compressed speech piece data are retrieved (step S502).
- step S502 if a plurality of compressed sound piece data are included in one sound piece, the corresponding compressed sound piece data is searched for, and one of the compressed sound piece data is searched for. If there is a voice piece that could not be found overnight, the above-mentioned missing part identification data is generated.
- the personal computer restores the extracted compressed speech piece data to the speech piece data before being compressed (step S503),
- the restored speech piece data is converted by the same processing as that performed by the speech piece editing unit 8 described above, and the time length of the speech piece represented by the speech piece data is made to match the speed indicated by the utterance speed data (Ste S504). If the utterance speed data is not supplied, the restored speech unit may not be converted.
- the personal computer determines the optimal combination of the speech unit data for synthesizing the voice to read the fixed message from the speech unit data in which the time length of the speech unit is converted, according to the third embodiment described above.
- the selection is performed by performing the same processing as the processing performed by the sound piece editing unit 8 in the mode (steps S505 to S507).
- this personal computer obtains the above-mentioned value, the set of ⁇ and / or Rmax for each pitch component data searched out in step S502, and the speed initial value data,
- the above-mentioned value dt is obtained using the fixed message data and the utterance speed data obtained in step S501 (step S501).
- the personal computer calculates the values of a, ⁇ , Rmax, and dt obtained in step S505 for each of the speech piece data converted in step S504, and outputs the values in the fixed message.
- the above-mentioned evaluation value ⁇ ⁇ is specified based on the frequency of the pitch component of the sound piece data representing the sound piece adjacent to the sound piece represented by the sound piece data (step S506 ) o
- the personal computer computes the sound unit 1 ⁇ which constitutes the fixed message represented by the fixed message data acquired in step S501 from the sound unit data converted in step S504.
- the speech with the largest total sum of the evaluation values ⁇ ⁇ ⁇ ⁇ of each piece of speech data belonging to the combination is synthesized as a voice reading a fixed message (Step S507).
- the evaluation value ⁇ ⁇ ⁇ used to calculate the sum a value that correctly reflects the connection relation of the sound pieces in the combination is selected.
- the personal computer when the personal computer generates the missing part identification data, the personal computer extracts a phonetic character string representing the reading of the speech piece indicated by the missing part identifying data from the standard message data, and By processing the above steps S202 to S203 for each phoneme in the same way as the phonetic sequence represented by the distribution character string data, each phoneme in this phonetic string is Restores waveform data representing the waveform of the voice indicated by the phonetic characters
- the personal computer compares the restored waveform data and the sound piece data belonging to the combination selected in step S507 with the order of each sound piece in the fixed message indicated by the fixed message data. And output as data representing synthesized speech
- a program that causes a personal computer to perform the functions of the main unit unit M and the speech unit registration unit R is, for example, a communication board bulletin board.
- BTS Backbone System
- a device that modulates a carrier with signals representing these programs transmits the resulting modulated wave, and receives this modulated wave May restore the programs by demodulating the modulated wave. Then, by starting these programs and executing them in the same manner as other application programs under the control of the OS, the above-described processing can be executed.
- the program excluding the part is stored in the recording medium. May be. Also in this case, in the present invention, it is assumed that the recording medium stores a program for executing each function or step executed by the computer.
- a voice selecting device a voice selecting method, and a program for obtaining a natural synthesized voice at high speed with a simple configuration are realized.
Landscapes
- Engineering & Computer Science (AREA)
- Computational Linguistics (AREA)
- Health & Medical Sciences (AREA)
- Audiology, Speech & Language Pathology (AREA)
- Human Computer Interaction (AREA)
- Physics & Mathematics (AREA)
- Acoustics & Sound (AREA)
- Multimedia (AREA)
- Machine Translation (AREA)
- Information Retrieval, Db Structures And Fs Structures Therefor (AREA)
- Compression, Expansion, Code Conversion, And Decoders (AREA)
Abstract
Description
Claims
Priority Applications (4)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US10/559,573 US20070100627A1 (en) | 2003-06-04 | 2004-06-03 | Device, method, and program for selecting voice data |
| CN2004800187934A CN1816846B (zh) | 2003-06-04 | 2004-06-03 | 用于选择话音数据的设备和方法 |
| EP04735989A EP1632933A4 (en) | 2003-06-04 | 2004-06-03 | DEVICE, METHOD AND PROGRAM FOR SELECTING VOICE DATA |
| DE04735989T DE04735989T1 (de) | 2003-06-04 | 2004-06-03 | Einrichtung, verfahren und programm zur auswahl von voice-daten |
Applications Claiming Priority (6)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP2003159880 | 2003-06-04 | ||
| JP2003-159880 | 2003-06-04 | ||
| JP2003165582 | 2003-06-10 | ||
| JP2003-165582 | 2003-06-10 | ||
| JP2004155306A JP4264030B2 (ja) | 2003-06-04 | 2004-05-25 | 音声データ選択装置、音声データ選択方法及びプログラム |
| JP2004-155306 | 2004-05-25 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2004109660A1 true WO2004109660A1 (ja) | 2004-12-16 |
Family
ID=33514559
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/JP2004/008088 Ceased WO2004109660A1 (ja) | 2003-06-04 | 2004-06-03 | 音声データを選択するための装置、方法およびプログラム |
Country Status (7)
| Country | Link |
|---|---|
| US (1) | US20070100627A1 (ja) |
| EP (1) | EP1632933A4 (ja) |
| JP (1) | JP4264030B2 (ja) |
| KR (1) | KR20060015744A (ja) |
| CN (1) | CN1816846B (ja) |
| DE (1) | DE04735989T1 (ja) |
| WO (1) | WO2004109660A1 (ja) |
Families Citing this family (9)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US7109208B2 (en) | 2001-04-11 | 2006-09-19 | Senju Pharmaceutical Co., Ltd. | Visual function disorder improving agents |
| CN1813285B (zh) * | 2003-06-05 | 2010-06-16 | 株式会社建伍 | 语音合成设备和方法 |
| JP4516863B2 (ja) * | 2005-03-11 | 2010-08-04 | 株式会社ケンウッド | 音声合成装置、音声合成方法及びプログラム |
| JP2008185805A (ja) * | 2007-01-30 | 2008-08-14 | Internatl Business Mach Corp <Ibm> | 高品質の合成音声を生成する技術 |
| JP5387410B2 (ja) * | 2007-10-05 | 2014-01-15 | 日本電気株式会社 | 音声合成装置、音声合成方法および音声合成プログラム |
| JP5093387B2 (ja) * | 2011-07-19 | 2012-12-12 | ヤマハ株式会社 | 音声特徴量算出装置 |
| CN111506736B (zh) * | 2020-04-08 | 2023-08-08 | 北京百度网讯科技有限公司 | 文本发音获取方法、装置和电子设备 |
| CN112669810B (zh) * | 2020-12-16 | 2023-08-01 | 平安科技(深圳)有限公司 | 语音合成的效果评估方法、装置、计算机设备及存储介质 |
| CN114495902B (zh) * | 2022-02-25 | 2025-10-17 | 北京有竹居网络技术有限公司 | 语音合成方法、装置、计算机可读介质及电子设备 |
Citations (11)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JPH01284898A (ja) * | 1988-05-11 | 1989-11-16 | Nippon Telegr & Teleph Corp <Ntt> | 音声合成方法 |
| JPH07319497A (ja) * | 1994-05-23 | 1995-12-08 | N T T Data Tsushin Kk | 音声合成装置 |
| JPH0944191A (ja) * | 1995-05-25 | 1997-02-14 | Sanyo Electric Co Ltd | 音声合成装置 |
| JPH09230893A (ja) * | 1996-02-22 | 1997-09-05 | N T T Data Tsushin Kk | 規則音声合成方法及び音声合成装置 |
| JPH1097268A (ja) * | 1996-09-24 | 1998-04-14 | Sanyo Electric Co Ltd | 音声合成装置 |
| JPH11249679A (ja) * | 1998-03-04 | 1999-09-17 | Ricoh Co Ltd | 音声合成装置 |
| JPH11259083A (ja) * | 1998-03-09 | 1999-09-24 | Canon Inc | 音声合成装置および方法 |
| JP2001013982A (ja) * | 1999-04-28 | 2001-01-19 | Victor Co Of Japan Ltd | 音声合成装置 |
| JP2001034284A (ja) * | 1999-07-23 | 2001-02-09 | Toshiba Corp | 音声合成方法及び装置、並びに文音声変換プログラムを記録した記録媒体 |
| JP2001092481A (ja) * | 1999-09-24 | 2001-04-06 | Sanyo Electric Co Ltd | 規則音声合成方法 |
| JP2003513311A (ja) * | 1999-10-28 | 2003-04-08 | シーメンス アクチエンゲゼルシヤフト | 合成すべき音声応答の基本周波数の時間特性を定めるための方法 |
Family Cites Families (8)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US5636325A (en) * | 1992-11-13 | 1997-06-03 | International Business Machines Corporation | Speech synthesis and analysis of dialects |
| JP3587048B2 (ja) * | 1998-03-02 | 2004-11-10 | 株式会社日立製作所 | 韻律制御方法及び音声合成装置 |
| JP3180764B2 (ja) * | 1998-06-05 | 2001-06-25 | 日本電気株式会社 | 音声合成装置 |
| US6505152B1 (en) * | 1999-09-03 | 2003-01-07 | Microsoft Corporation | Method and apparatus for using formant models in speech systems |
| US6496801B1 (en) * | 1999-11-02 | 2002-12-17 | Matsushita Electric Industrial Co., Ltd. | Speech synthesis employing concatenated prosodic and acoustic templates for phrases of multiple words |
| US6865533B2 (en) * | 2000-04-21 | 2005-03-08 | Lessac Technology Inc. | Text to speech |
| CA2359771A1 (en) * | 2001-10-22 | 2003-04-22 | Dspfactory Ltd. | Low-resource real-time audio synthesis system and method |
| US20040030555A1 (en) * | 2002-08-12 | 2004-02-12 | Oregon Health & Science University | System and method for concatenating acoustic contours for speech synthesis |
-
2004
- 2004-05-25 JP JP2004155306A patent/JP4264030B2/ja not_active Expired - Fee Related
- 2004-06-03 US US10/559,573 patent/US20070100627A1/en not_active Abandoned
- 2004-06-03 KR KR1020057023078A patent/KR20060015744A/ko not_active Ceased
- 2004-06-03 CN CN2004800187934A patent/CN1816846B/zh not_active Expired - Lifetime
- 2004-06-03 DE DE04735989T patent/DE04735989T1/de active Pending
- 2004-06-03 WO PCT/JP2004/008088 patent/WO2004109660A1/ja not_active Ceased
- 2004-06-03 EP EP04735989A patent/EP1632933A4/en not_active Withdrawn
Patent Citations (11)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JPH01284898A (ja) * | 1988-05-11 | 1989-11-16 | Nippon Telegr & Teleph Corp <Ntt> | 音声合成方法 |
| JPH07319497A (ja) * | 1994-05-23 | 1995-12-08 | N T T Data Tsushin Kk | 音声合成装置 |
| JPH0944191A (ja) * | 1995-05-25 | 1997-02-14 | Sanyo Electric Co Ltd | 音声合成装置 |
| JPH09230893A (ja) * | 1996-02-22 | 1997-09-05 | N T T Data Tsushin Kk | 規則音声合成方法及び音声合成装置 |
| JPH1097268A (ja) * | 1996-09-24 | 1998-04-14 | Sanyo Electric Co Ltd | 音声合成装置 |
| JPH11249679A (ja) * | 1998-03-04 | 1999-09-17 | Ricoh Co Ltd | 音声合成装置 |
| JPH11259083A (ja) * | 1998-03-09 | 1999-09-24 | Canon Inc | 音声合成装置および方法 |
| JP2001013982A (ja) * | 1999-04-28 | 2001-01-19 | Victor Co Of Japan Ltd | 音声合成装置 |
| JP2001034284A (ja) * | 1999-07-23 | 2001-02-09 | Toshiba Corp | 音声合成方法及び装置、並びに文音声変換プログラムを記録した記録媒体 |
| JP2001092481A (ja) * | 1999-09-24 | 2001-04-06 | Sanyo Electric Co Ltd | 規則音声合成方法 |
| JP2003513311A (ja) * | 1999-10-28 | 2003-04-08 | シーメンス アクチエンゲゼルシヤフト | 合成すべき音声応答の基本周波数の時間特性を定めるための方法 |
Non-Patent Citations (1)
| Title |
|---|
| See also references of EP1632933A4 * |
Also Published As
| Publication number | Publication date |
|---|---|
| KR20060015744A (ko) | 2006-02-20 |
| EP1632933A4 (en) | 2007-11-14 |
| JP4264030B2 (ja) | 2009-05-13 |
| EP1632933A1 (en) | 2006-03-08 |
| JP2005025173A (ja) | 2005-01-27 |
| US20070100627A1 (en) | 2007-05-03 |
| CN1816846B (zh) | 2010-06-09 |
| DE04735989T1 (de) | 2006-10-12 |
| CN1816846A (zh) | 2006-08-09 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| CN1813285B (zh) | 语音合成设备和方法 | |
| JP4516863B2 (ja) | 音声合成装置、音声合成方法及びプログラム | |
| US20070011009A1 (en) | Supporting a concatenative text-to-speech synthesis | |
| JP4264030B2 (ja) | 音声データ選択装置、音声データ選択方法及びプログラム | |
| US7089187B2 (en) | Voice synthesizing system, segment generation apparatus for generating segments for voice synthesis, voice synthesizing method and storage medium storing program therefor | |
| JP4287785B2 (ja) | 音声合成装置、音声合成方法及びプログラム | |
| JP2005018036A (ja) | 音声合成装置、音声合成方法及びプログラム | |
| JP2010224418A (ja) | 音声合成装置、方法およびプログラム | |
| JP2004272236A (ja) | ピッチ波形信号分割装置、音声信号圧縮装置、データベース、音声信号復元装置、音声合成装置、ピッチ波形信号分割方法、音声信号圧縮方法、音声信号復元方法、音声合成方法、記録媒体及びプログラム | |
| JP4209811B2 (ja) | 音声選択装置、音声選択方法及びプログラム | |
| JP4780188B2 (ja) | 音声データ選択装置、音声データ選択方法及びプログラム | |
| JP7183556B2 (ja) | 合成音生成装置、方法、及びプログラム | |
| JP2010224419A (ja) | 音声合成装置、方法およびプログラム | |
| JP4574333B2 (ja) | 音声合成装置、音声合成方法及びプログラム | |
| JP4184157B2 (ja) | 音声データ管理装置、音声データ管理方法及びプログラム | |
| KR20100003574A (ko) | 음성음원정보 생성 장치 및 시스템, 그리고 이를 이용한음성음원정보 생성 방법 | |
| JP2006145848A (ja) | 音声合成装置、音片記憶装置、音片記憶装置製造装置、音声合成方法、音片記憶装置製造方法及びプログラム | |
| JP2006145690A (ja) | 音声合成装置、音声合成方法及びプログラム | |
| JP2007240988A (ja) | 音声合成装置、データベース、音声合成方法及びプログラム | |
| JP2007240989A (ja) | 音声合成装置、音声合成方法及びプログラム | |
| JP2006195207A (ja) | 音声合成装置、音声合成方法及びプログラム | |
| JP2007240987A (ja) | 音声合成装置、音声合成方法及びプログラム | |
| JP2007240990A (ja) | 音声合成装置、音声合成方法及びプログラム |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| AK | Designated states |
Kind code of ref document: A1 Designated state(s): AE AG AL AM AT AU AZ BA BB BG BR BW BY BZ CA CH CN CO CR CU CZ DE DK DM DZ EC EE EG ES FI GB GD GE GH GM HR HU ID IL IN IS KE KG KP KR KZ LC LK LR LS LT LU LV MA MD MG MK MN MW MX MZ NA NI NO NZ OM PG PH PL PT RO RU SC SD SE SG SK SL SY TJ TM TN TR TT TZ UA UG US UZ VC VN YU ZA ZM ZW |
|
| AL | Designated countries for regional patents |
Kind code of ref document: A1 Designated state(s): GM KE LS MW MZ NA SD SL SZ TZ UG ZM ZW AM AZ BY KG KZ MD RU TJ TM AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HU IE IT LU MC NL PL PT RO SE SI SK TR BF BJ CF CG CI CM GA GN GQ GW ML MR NE SN TD TG |
|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application | ||
| WWE | Wipo information: entry into national phase |
Ref document number: 2004735989 Country of ref document: EP |
|
| WWE | Wipo information: entry into national phase |
Ref document number: 1020057023078 Country of ref document: KR |
|
| WWE | Wipo information: entry into national phase |
Ref document number: 20048187934 Country of ref document: CN |
|
| WWP | Wipo information: published in national office |
Ref document number: 1020057023078 Country of ref document: KR |
|
| WWP | Wipo information: published in national office |
Ref document number: 2004735989 Country of ref document: EP |
|
| WWE | Wipo information: entry into national phase |
Ref document number: 2007100627 Country of ref document: US Ref document number: 10559573 Country of ref document: US |
|
| WWP | Wipo information: published in national office |
Ref document number: 10559573 Country of ref document: US |
