WO2004109660A1 - 音声データを選択するための装置、方法およびプログラム - Google Patents

音声データを選択するための装置、方法およびプログラム Download PDF

Info

Publication number
WO2004109660A1
WO2004109660A1 PCT/JP2004/008088 JP2004008088W WO2004109660A1 WO 2004109660 A1 WO2004109660 A1 WO 2004109660A1 JP 2004008088 W JP2004008088 W JP 2004008088W WO 2004109660 A1 WO2004109660 A1 WO 2004109660A1
Authority
WO
WIPO (PCT)
Prior art keywords
data
speech
unit
voice
representing
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/JP2004/008088
Other languages
English (en)
French (fr)
Inventor
Yasushi Sato
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Kenwood KK
Original Assignee
Kenwood KK
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Kenwood KK filed Critical Kenwood KK
Priority to US10/559,573 priority Critical patent/US20070100627A1/en
Priority to CN2004800187934A priority patent/CN1816846B/zh
Priority to EP04735989A priority patent/EP1632933A4/en
Priority to DE04735989T priority patent/DE04735989T1/de
Publication of WO2004109660A1 publication Critical patent/WO2004109660A1/ja
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Classifications

    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L13/00Speech synthesis; Text to speech systems
    • G10L13/02Methods for producing synthetic speech; Speech synthesisers
    • G10L13/027Concept to speech synthesisers; Generation of natural phrases from machine-based concepts
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L13/00Speech synthesis; Text to speech systems
    • G10L13/06Elementary speech units used in speech synthesisers; Concatenation rules

Definitions

  • the present invention relates to an audio data selection device, an audio data selection method, and a program.
  • the recording and editing method is used for voice guidance systems at stations and navigation devices for vehicles.
  • a word is associated with voice data representing a voice that reads out the word, a sentence to be subjected to voice synthesis is divided into words, and voice data associated with these words is acquired. It is a technique of connecting together.
  • prosody prediction is an extremely complicated process
  • it is necessary to use a processor with high processing power or to perform the processing over a long time. Therefore, this method is not suitable for applications that require high-speed processing using a device with a simple configuration.
  • the present invention has been made in view of the above situation, and has as its object to provide an audio data selection device, an audio data selection method, and a program for obtaining a natural synthesized voice at high speed with a simple configuration. .
  • the audio data selection device basically includes a storage unit that stores a plurality of audio data representing a waveform of an audio; Search means for inputting sentence information representing a sentence, and searching for audio data representing a waveform of a sound unit having a common reading with a sound unit constituting the sentence from among the sound data; Value of the speech data corresponding to each of the speech units constituting the sentence, and the pitch difference at the boundary between adjacent speech units in the entire sentence. And selecting means for selecting so that is minimized.
  • the voice data selecting device may further include voice synthesizing means for generating data representing a synthesized voice by combining the selected voice data with each other.
  • the voice data selection method of the present invention basically stores a plurality of voice data representing a voice waveform, inputs text information representing a text, and forms the text from the voice data. Speech data representing the waveform of a speech unit having the same reading as the speech unit is found, and one of the found speech data is speech data corresponding to each of the speech units constituting the sentence. Each time, a pitch difference at a boundary between adjacent sound pieces is selected so that a value obtained by accumulating the entire text is minimized.
  • the computer program according to the present invention further comprises: a storage unit for storing a plurality of voice data representing a waveform of a voice; and text information representing a text.
  • a search unit that searches for a speech data representing a waveform of a speech unit having the same reading as the constituent speech units, and a search unit that searches the speech data for each of the speech units constituting the sentence from the retrieved speech data.
  • Selecting means for selecting the corresponding voice data one by one, and selecting the pitch difference at the boundary between adjacent speech pieces so that the value obtained by accumulating the total of the entire sentence is minimized. It has become something.
  • the voice selection device basically inputs storage information for storing a plurality of voice data representing a voice waveform, text information representing a text, and reads the text.
  • a prediction unit that predicts a temporal change in the pitch of the speech unit by performing prosodic prediction on the speech unit that constitutes the speech unit; and a speech unit that has the same reading as the speech unit that constitutes the sentence among the speech data.
  • selecting means for selecting audio data which represents a waveform and whose time change in pitch has the highest correlation with the result of prediction by the predicting means.
  • the selection means includes a regression calculation for performing a first-order regression between a time change of a pitch of a sound piece represented by voice data and a time change of a pitch of a sound piece in the sentence having the same reading as the sound piece. Based on the result, the strength of the correlation between the time change of the pitch of the audio data and the result of the prediction by the prediction means may be specified.
  • the selecting means is configured to determine, based on a correlation coefficient between a temporal change of a pitch of a voice unit represented by voice data and a temporal change of a pitch of a voice unit in the text that is commonly read with the voice unit, the voice Any correlation strength may be specified as a result of the time change of the pitch over time and the result of prediction by the prediction means.
  • Another voice selecting device of the present invention is a storage means for storing a plurality of voice data representing a voice waveform, and inputting text information representing a text, and predicting a prosody of a speech unit in the text.
  • a prediction unit for predicting the time length of the speech unit and the temporal change of the pitch of the speech unit, and each audio data representing the waveform of the speech unit having the same reading as the speech unit in the text.
  • Selecting means for specifying an evaluation value for evening and selecting audio data having the highest evaluation value, wherein the evaluation value is a temporal change in pitch of a sound piece represented by the audio data.
  • the numerical value representing the correlation is obtained by a first-order regression between the time change of the pitch of the sound piece represented by the voice data and the time change of the pitch of the sound piece in the sentence having the same reading as the sound piece.
  • the gradient of the linear function It may be.
  • the numerical value representing the correlation is obtained by a first-order regression between the time change of the pitch of the sound piece represented by the voice data and the time change of the pitch of the sound piece in the sentence having the same reading as the sound piece. It may consist of the intercept of the obtained linear function.
  • the numerical value representing the correlation is a correlation coefficient between the time change of the pitch of the sound piece represented by the voice data and the predicted result of the time change of the pitch of the sound piece in the sentence having the same reading as the sound piece. It may consist of.
  • the numerical value representing the correlation is a function represented by the data representing the time change of the pitch of the sound piece represented by the voice data, which is obtained by shifting the number of bits cyclically.
  • the sound in the sentence having the same reading as the sound piece is read. It may consist of the maximum value of the correlation coefficient with a function representing the prediction result of the temporal change of the pitch of the piece.
  • the storage means may store phonogram data representing the reading of the voice data in association with the voice data, and the selecting means may store a reading matching the reading of the speech unit in the text. Speech data to which the represented phonogram data is associated may be handled as speech data representing the waveform of the vocal piece having the same reading as the relevant vocal piece.
  • the voice selecting device may further include a voice synthesizing unit that generates data representing a synthesized voice by combining the selected voice data with each other.
  • the voice selecting device may be configured to represent a waveform of the voice unit of the voice unit in which the voice data is not selected by the storage unit, without using the voice data stored in the storage unit.
  • the apparatus may further comprise a missing part synthesizing means for synthesizing voice data, wherein the voice synthesizing means comprises: The data synthesized by the means may be combined with each other to generate a data representing the synthesized voice.
  • the voice selection method of the present invention stores a plurality of voice data representing the waveform of a voice, inputs text information representing a text, and performs prosody prediction on a voice unit constituting the text, thereby obtaining a prosody of the voice unit.
  • a time change of pitch is predicted, and a waveform of a voice unit having a common reading with a voice unit constituting the sentence is represented from among the voice data, and the time change of the pitch is determined by the prediction unit. And selecting the speech data that shows the highest correlation with the result of the prediction by.
  • another voice selection method of the present invention stores a plurality of voice data representing a voice waveform, inputs text information representing a text, and performs prosody prediction on a speech unit in the text.
  • Estimate the time length of the voice unit and the time change of the pitch of the voice unit specify the evaluation value of each voice data representing the waveform of the voice unit having the same reading as the voice unit in the sentence, The voice data having the highest evaluation value is selected, and the evaluation value is determined based on the time change of the pitch of the sound piece represented by the voice data and the text in which the reading is common to the sound piece.
  • the computer program of the present invention is a computer program for storing a plurality of voice data representing voice waveforms, inputting text information representing a text, and performing prosodic prediction on a speech unit constituting the text.
  • Prediction means for predicting a temporal change in the pitch of the speech piece, and a speech piece constituting the text from each of the speech data Represents the waveform of a common speech unit, and a selection means for selecting voice data whose temporal change in pitch has the highest correlation with the result of the prediction by the prediction means. It has become.
  • another computer program is a computer program, comprising: a storage means for storing a plurality of voice data representing a waveform of a voice; and text information representing a text, and a prosody prediction for a speech unit in the text.
  • a prediction unit that predicts the time length of the speech unit and the time change of the pitch of the speech unit, and each voice data representing the waveform of the speech unit having the same reading as the speech unit in the text
  • a program for specifying an evaluation value of the speech data and causing it to function as selection means for selecting audio data representing the highest evaluation, wherein the evaluation value is a speech data represented by a speech unit represented by evening.
  • a function of a numerical value representing the correlation between the time change of the pitch of the sound piece and the prediction result of the time change of the pitch of the sound piece in the sentence that is common to the sound piece and the time of the sound piece represented by the sound data The chief, It is obtained from the function of the difference between the speech unit and the prediction result of the time length of the speech unit in the sentence having the same reading.
  • the speech data selection device basically comprises a storage means for storing a plurality of speech data representing a speech waveform, and a sentence for inputting sentence information representing a sentence.
  • An information input unit a search unit that searches for voice data having a portion that is common to a voice unit in the text represented by the text information, and a search unit that connects the searched voice data in accordance with the text represented by the text information
  • Selecting means for obtaining an evaluation value according to a predetermined evaluation criterion based on a relationship between mutually adjacent voice data and selecting a combination of output voice data based on the evaluation value.
  • the evaluation criterion is a criterion for determining a correlation between a voice represented by voice data and a prosody prediction result and an evaluation value indicating a relationship between voice data adjacent to each other, wherein the evaluation value is a value of a voice represented by the voice data.
  • An evaluation expression including at least one of a parameter indicating a feature, a parameter indicating a feature of a voice obtained by combining the voices represented by the voice data with each other, and a parameter indicating a feature related to the speech time length. It may be one that is obtained based on this.
  • the evaluation criterion is a criterion for determining a correlation between a voice represented by voice data and a prosody prediction result and an evaluation value indicating a relationship between voice data adjacent to each other.
  • the parameters indicating the characteristics of the voice obtained by combining the voices represented by the voice data with each other, and the parameters indicating the characteristics of the voice represented by the voice data and the parameters indicating the characteristics related to the speech duration may be obtained based on an evaluation formula including at least one of them.
  • the parameter indicating the characteristic of the voice obtained by combining the voices represented by the voice data with each other is selected from among voice data representing a waveform of a voice having a portion common to a speech unit in a text represented by the text information and a reading. This is obtained based on the pitch difference at the boundary between adjacent audio data when one audio data corresponding to each sound piece constituting the sentence is selected one by one. There may be.
  • the speech unit data selection device inputs sentence information representing a sentence and predicts the prosody of the speech unit in the sentence, thereby predicting the time length of the speech unit and the time change of the pitch of the speech unit.
  • the evaluation criterion defines an evaluation value indicating a correlation or a difference between a voice represented by voice data and a prosody prediction result of the prosody prediction means.
  • the evaluation value is a criterion, and the evaluation value is a correlation between a time change of a pitch of a sound piece represented by voice data and a prediction result of a time change of a pitch of a sound piece in the sentence having the same reading as the sound piece.
  • the numerical value representing the correlation is obtained by a first-order regression between the time change of the pitch of the sound piece represented by the voice data and the time change of the pitch of the sound piece in the sentence having the same reading as the relevant sound piece. It may consist of the slope and / or intercept of a linear function.
  • the numerical value representing the correlation is a correlation coefficient between the time change of the pitch of the sound piece represented by the voice data and the predicted result of the time change of the pitch of the sound piece in the sentence having the same reading as the sound piece. It may consist of.
  • the numerical value representing the correlation may be a function represented by data obtained by shifting the pitch of the speech piece represented by the voice data over time by various numbers of bits, and a value represented in the sentence having the same reading as the speech piece. It may be made up of the maximum value of the correlation coefficient with the function representing the prediction result of the time change of the pitch of the sound piece.
  • the storage unit may store phonogram data representing a reading of voice data in association with the voice data, and the selecting unit may represent a reading matching a reading of a speech unit in the text.
  • the voice data associated with the phonetic data may be handled as voice data representing the waveform of a voice unit having the same reading as the relevant voice unit.
  • the speech unit data selection device includes a speech synthesis unit that generates data representing a synthesized speech by combining the selected speech data with each other. 'It may have more.
  • the sound piece de-night selection device for the sound pieces in which the selecting means cannot select the sound data among the sound pieces in the text, without using the sound data stored in the storage means.
  • Missing voice synthesizing means for synthesizing voice data representing the waveform of the voice signal, wherein the voice synthesizing means combines the voice data selected by the selecting means and the voice data synthesized by the missing voice portion with each other. By combining the data, data representing the synthesized voice may be generated.
  • the voice data selection method of the present invention stores a plurality of voice data representing a voice waveform, inputs text information representing a text, and has a portion that is commonly read with a speech unit in the text represented by the text information.
  • voice data is searched, and the searched voice data is connected in accordance with a text represented by text information, an evaluation value is obtained and output according to a predetermined evaluation criterion based on a relationship between adjacent voice data. It includes a first-speed processing step of selecting a combination of audio data based on the evaluation value.
  • the computer program of the present invention comprises: a computer, a storage unit for storing a plurality of voice data representing a voice waveform, a text information input unit for inputting text information representing text, and a text information represented by the text information.
  • a search unit that searches for voice data having a portion that is common to the voice unit of the voice unit, and voice data that are adjacent to each other when the searched voice data are connected in accordance with the text represented by text information.
  • the evaluation means obtains an evaluation value in accordance with a predetermined evaluation criterion based on the above relationship, and functions as a selecting means for selecting a combination of output audio data based on the evaluation value.
  • FIG. 1 is a block diagram showing a configuration of a speech synthesis system according to each embodiment of the present invention.
  • FIG. 2 is a diagram schematically showing a data structure of a sound piece database according to the first embodiment of the present invention.
  • FIG. 4B is a graph for explaining a process of linearly regressing the change
  • FIG. 4B is a graph showing an example of a prediction result data and a value of pitch component data used for obtaining a correlation coefficient. is there.
  • FIG. 4 is a diagram schematically showing a data structure of a speech piece database according to a second embodiment of the present invention.
  • Fig. 5 (a) is a diagram showing the reading of a fixed message
  • Fig. 5 (b) is a list of the speech unit data supplied to the speech unit editing unit
  • Fig. 5 (c) is
  • FIG. 9D is a diagram showing the absolute value of the difference between the frequency of the pitch component at the end of the preceding speech unit and the frequency of the pitch component at the beginning of the succeeding speech unit.
  • FIG. It is a figure which shows whether one day is selected.
  • FIG. 6 is a flowchart showing processing when a personal computer performing the function of the speech synthesis system according to each embodiment of the present invention has acquired free text data.
  • FIG. 7 is a flowchart showing processing when a personal computer that performs the function of the speech synthesis system according to each embodiment of the present invention has acquired distribution character string data.
  • FIG. 8 is a diagram showing a speech synthesis according to the first embodiment of the present invention.
  • FIG. 5 is a flowchart showing a process performed when a personal computer performing the above function obtains fixed message data and utterance speed data.
  • FIG. 9 is a flowchart showing a process when a personal computer performing the function of the speech synthesis system according to the second embodiment of the present invention acquires fixed message data and utterance speed data.
  • FIG. 10 is a flowchart showing a process performed when a personal computer performing the function of the speech synthesis system according to the third embodiment of the present invention acquires fixed message data and utterance speed data. is there.
  • FIG. 1 is a diagram showing a configuration of a speech synthesis system according to a first embodiment of the present invention.
  • the speech synthesis system includes a main unit M and a speech unit registration unit R.
  • the main unit M consists of a language processing unit 1, a general word dictionary 2, a user word dictionary 3, a sound processing unit 4, a search unit 5, a decompression unit 6, a waveform data base 7, and a speech unit. It comprises an editing unit 8, a search unit 9, a speech unit database 10, and a speech speed conversion unit 11.
  • the language processing unit 1, the sound processing unit 4, the search unit 5, the decompression unit 6, the speech unit editing unit 8, the search unit 9, and the speech speed conversion unit 11 are all CPUs (Central ⁇
  • a processor such as a processing unit (DP) and a digital signal processor (DP), and a memory for storing a program to be executed by the processor, and performs processing described later.
  • DP processing unit
  • DP digital signal processor
  • a single processor performs part or all of the functions of the language processing unit 1, the sound processing unit 4, the search unit 5, the decompression unit 6, the speech unit editing unit 8, the search unit 9, and the speech speed conversion unit 11. You may do so.
  • the general word dictionary 2 is composed of a non-volatile memory such as a PROM (Programmable Read Only Memory) and a hard disk device.
  • the general word dictionary 2 contains words including ideographic characters (for example, kanji) and phonograms (for example, kana and phonograms) representing the reading of the words and the like. They are stored in advance in association with each other by a manufacturer or the like.
  • the user word dictionary 3 is a data rewritable non-volatile memory such as an EEPROM (Electrically Erasable / Programmable Read Only Memory) and a hard disk device, and a control for controlling data writing to the non-volatile memory. It consists of a circuit. Note that the processor may perform the function of this control circuit, and the language processing unit 1, the sound processing unit 4, the search unit 5, the decompression unit 6, the speech unit editing unit 8, the search unit 9, and the speech speed conversion unit 11 A processor that performs a part or all of the functions may perform the function of the control circuit of the user word dictionary 3.
  • EEPROM Electrically Erasable / Programmable Read Only Memory
  • the user word dictionary 3 acquires words including ideographic characters and phonograms indicating reading of the words and the like from outside according to user operations, and stores them in association with each other.
  • the user word dictionary 3 stores words and the like not stored in the general word dictionary 2 and phonograms representing their readings. Is enough.
  • the waveform database 7 is composed of a nonvolatile memory such as a PR ⁇ M or a hard disk device.
  • the waveform data base 7 contains the phonograms and the compressed waveform data obtained by subjecting the speech data representing the waveform of the unit voice represented by the phonograms to the entrance-to-peak coding. They are stored in advance in association with each other by a stem manufacturer or the like.
  • the unit speech is a speech that is short enough to be used in the rule-based synthesis method, and specifically, is speech that is separated by units such as phonemes and V CV (Vowel-Consonant-Vowel) syllables.
  • the waveform data before being subjected to the entropy encoding may be composed of, for example, digital data that has been subjected to pulse code modulation (PCM).
  • PCM pulse code modulation
  • the voice element data base 10 is composed of a nonvolatile memory such as a PROM and a hard disk device.
  • the speech unit database 10 stores, for example, data having a data structure shown in FIG. That is, as shown in the figure, the data stored in the speech unit database 10 is divided into four types: a header section HDR, an index section IDX, a directory section DIR, and a data section DAT. ing.
  • the storage of the data in the voice unit database 10 is performed in advance by, for example, the manufacturer of the voice synthesis system, and is performed by the Z or the voice unit registration unit R performing an operation described later.
  • the header HDR contains data identifying the speech unit database 10 and data indicating the data amount, data format, copyright, etc. of the index part IDX, directory part DIR and data part DAT. But Will be delivered.
  • the data section DAT stores the compressed speech unit data obtained by entropy-encoding the speech unit data representing the speech unit waveform.
  • a speech unit is a continuous section containing one or more phonemes in a voice, and usually consists of one or more words.
  • the speech piece data before entropy encoding is the same format as the waveform data before entropy encoding for generating the above-described compressed waveform data (for example, digital Format).
  • the directory section DIR contains individual compressed audio data
  • FIG. 2 shows that the data included in the data part DAT is a compressed speech piece data having a data amount of 141 h bytes, which represents the waveform of a speech piece whose reading is “Saitama”.
  • a 3 6 A 6 'h first The case where it is stored in a logical position is illustrated. (In this specification and the drawings, the number suffixed with “h” indicates a hexadecimal number.)
  • the pitch component data is, for example, as shown in the figure, the frequency of the pitch component of the sound piece. It is assumed that the data represents a sample Y (i) obtained by sampling (where n is a positive integer equal to or less than n, where n is the total number of samples).
  • At least the data (A) (that is, the speech unit reading data) of the above set of data items (A) to (E) is sorted according to the order determined based on the phonetic characters represented by the speech unit reading data. In a single state (for example, if the phonetic characters are kana, they are arranged in descending address order according to the Japanese syllabary order) and stored in the storage area of the speech unit database 10.
  • the index section IDX stores data for specifying the approximate logical position of the data in the directory section DIR based on the speech unit reading data. Specifically, for example, assuming that the speech unit reading data represents power, the kana character and the ', and the range of addresses of the speech unit reading data whose first character is this kana character are Is stored in association with each other.
  • a single nonvolatile memory may perform some or all of the functions of the general word dictionary 2, the user word dictionary 3, the waveform database 7, and the speech unit database 10.
  • the storage of the data in the speech unit database 10 is performed by the speech unit registration unit R shown in FIG.
  • the speech unit registration unit R includes a recorded speech unit data set storage unit 12, a speech unit database creation unit 13, and a compression unit 14.
  • the sound registration The nit R may be removably connected to the speech unit database 10 in this case.
  • the speech unit With the registration unit R separated from the main unit M, the main unit M may perform the operation described below.
  • the recorded sound piece data set storage unit 12 is composed of a non-volatile memory that can be rewritten overnight, such as a hard disk device.
  • the stored sound data storage unit 12 contains phonograms that represent the reading of a sound piece, and sounds that represent the waveform obtained by collecting the actual sound of this sound piece.
  • the piece data is stored in association with each other in advance by the manufacturer of the speech synthesis system or the like. It should be noted that the sound piece data may be composed of, for example, digital data converted into PCM.
  • the voice unit data base creation unit 13 and the compression unit 14 are composed of a processor such as a CPU, a memory for storing a program to be executed by the processor, and the like. Do.
  • a single processor may perform some or all of the functions of the speech unit database creation unit 13 and the compression unit 14. Also, the language processing unit 1, the sound processing unit 4, and the search unit 5 The processor that performs part or all of the functions of the decompression unit 6, the speech unit editing unit 8, the search unit 9, and the speech speed conversion unit 11 further performs the functions of the speech unit database creation unit 13 and compression unit 14. You may. Further, a processor that performs the functions of the speech unit database creation unit 13 and the compression unit 14 may also serve as the control circuit of the recorded speech unit data set storage unit 12.
  • the sound piece data set creation section 13 From step 2, the phonetic character and the speech unit data that are associated with each other are read out, and the time change of the frequency of the pitch component of the voice represented by the speech unit data and the utterance speed are specified.
  • the utterance speed may be specified, for example, by counting the number of samples in the sound piece.
  • the time change of the frequency of the pitch component may be specified by performing cepstrum analysis on the sound piece data, for example.
  • the waveform represented by the speech unit data is divided into a number of small parts on the time axis, and the strength of each obtained small part is calculated as the logarithm of the original value (the base of the logarithm is arbitrary) Transforms this small portion of the spectrum (that is, the cepstrum) into a value that is substantially equal to, and generates data that represents the result of the Fourier transform of a discrete variable (or a discrete variable). Any other method).
  • the minimum value of the frequencies giving the maximum value of the cepstrum is specified as the frequency of the pitch component in this small portion.
  • the time change of the frequency of the pitch component can be calculated, for example, by converting the sound piece data into a pitch waveform data according to the method disclosed in Japanese Patent Application Laid-Open No. 2003-108172. After that, good results can be expected if identification is performed based on this pitch waveform data.
  • the pitch signal is extracted by filtering the speech unit data, and the waveform represented by the speech unit data is divided into sections of unit pitch length based on the extracted pitch signal. It is sufficient to specify the phase shift based on the correlation with and to convert the speech unit data into a pitch waveform signal by aligning the phases of the respective sections. Then, the obtained pitch waveform signal is treated as sound piece data, and cepstrum analysis is performed.
  • the time change of the frequency of the pitch component may be specified.
  • the speech unit database creation unit 13 supplies the speech unit data read from the recorded speech unit data set storage unit 12 to the compression unit 14.
  • the compression unit 14 creates compressed speech unit data by entropy-encoding the speech unit data supplied from the speech unit database creation unit 13, and returns it to the speech unit data base creation unit 13.
  • the time change of the utterance speed and the frequency of the pitch component of the speech unit data is identified, and this speech unit data is subjected to entropy coding and returned as a compressed speech unit data from the compression unit 14.
  • the voice unit data base creation unit 13 writes the compressed voice unit data into the storage area of the voice unit database 10 as data constituting the data unit DAT.
  • the speech unit data base creation unit 13 also converts the phonogram read out from the recorded speech unit data set storage unit 12 as an indication of the reading of the speech unit represented by the written compressed speech unit data, and Write the read data to the storage area of the base unit 10 as the read data.
  • the head address of the written compressed speech piece data in the storage area of the speech piece database 10 is specified, and this address is written to the storage area of the speech piece database 10 as the above-mentioned (B) data. .
  • the data length of the compressed speech piece data is specified, and the specified data length is written to the storage area of the speech piece database 10 as the data (C).
  • a data indicating the time change of the utterance speed and the frequency of the pitch component of the voice unit represented by the compressed voice data is generated, and the data is generated as speed initial value data and pitch component data.
  • the language processing unit 1 obtains free text data describing a sentence (free text) including ideographic characters prepared by a user as an object for synthesizing a speech in the speech synthesis system from the outside. I do.
  • the language processing unit 1 may obtain the free text data by any method.
  • the language processing unit 1 may obtain the free text data from an external device network via an interface circuit (not shown), or a recording device (not shown).
  • the medium may be read from a recording medium (for example, a floppy (registered trademark) disk or a CD_ROM) set in the drive device via the recording medium drive device.
  • the processor performing the function of the language processing unit 1 may transfer text data used in other processing being executed by itself to the processing of the language processing unit 1 as free text data. .
  • the language processing unit 1 searches the phonetic character representing the reading of each ideographic character included in the free text by searching the general word dictionary 2 and the user word dictionary 3. Identify. Then, this ideogram is replaced with the specified phonogram. Then, the language processing unit 1 supplies a phonogram string obtained as a result of replacing all ideograms in the free text with phonograms to the sound processing unit 4.
  • the sound processing section 4 searches for the waveform of the unit speech represented by the phonogram for each phonogram included in the phonogram string. Instruct the search unit 5.
  • the search unit 5 searches the waveform database 7 in response to this instruction, The compressed waveform data representing the waveform of the unit voice represented by each phonetic character included in the phonetic character string is searched for. Then, the retrieved compressed waveform data is supplied to the decompression unit 6.
  • the decompression unit 6 restores the compressed waveform data supplied from the search unit 5 to the waveform data before being compressed, and returns it to the search unit 5.
  • the search unit 5 supplies the waveform data returned from the decompression unit 6 to the sound processing unit 4 as a search result.
  • the sound processing unit 4 converts the waveform data supplied from the search unit 5 into a speech unit editing unit according to the order of each phonogram in the phonogram string supplied from the language processing unit 1.
  • Supply to 8. -Upon receiving the waveform data from the sound processing unit 4, the speech unit editing unit 8 combines the waveform data with each other in the order of supply and outputs the combined data as data representing synthesized speech (synthesized speech data). I do.
  • This synthesized speech synthesized based on the free text data is equivalent to the speech synthesized by the rule synthesis method.
  • the method by which the sound piece editing unit 8 outputs synthesized speech data is arbitrary.
  • a D / A (Digital-to-Analog) converter (not shown)
  • the synthesized voice represented by the synthesized voice data may be reproduced.
  • the data may be sent to an external device network via an interface circuit (not shown), or may be written to a recording medium set in a recording medium drive device (not shown) via the recording medium drive device. May be.
  • the processor performing the function of the sound piece editing unit 8 may transfer the synthesized speech data to another process executed by itself.
  • the sound processing unit 4 represents the phonetic character string distributed from the outside. It is assumed that data (delivery character string data) has been obtained. (Note that the method by which the sound processing unit 4 acquires the distribution character string data is also arbitrary. For example, the language processing unit 1 may acquire the distribution character string data by the same method as the method of acquiring the free text data. )
  • the sound processing unit 4 treats the phonetic character string represented by the distribution character string data in the same manner as the phonetic character string supplied from the language processing unit 1.
  • the compressed waveform data corresponding to the phonetic characters included in the phonetic character string represented by the distribution character string data is retrieved by the search unit 5, and the waveform data before compression is retrieved by the decompression unit 6. Will be restored.
  • Each of the restored waveform data is supplied to the sound piece editing unit 8 via the sound processing unit 4, and the sound unit editing unit 8 converts this waveform data into each of the phonetic character strings represented by the distribution character string data. Combine them in the order of phonetic characters and output them as synthesized speech data.
  • the synthesized speech data synthesized based on the distribution character string data also indicates the speech synthesized by the rule synthesis method.
  • the speech piece editing unit 8 has acquired the fixed message data and the utterance speed data.
  • the fixed message data is data representing a fixed message as a phonetic character string
  • the utterance speed data is a fixed message message—the specified value of the utterance speed of the fixed message represented by the evening (this fixed message is (The specified value of the utterance time length).
  • the method by which the sound piece editing unit 8 obtains the fixed message data and the utterance speed data is arbitrary.
  • the fixed message data and the utterance are obtained in the same manner as the method by which the language processing unit 1 obtains the free text data. What is necessary is just to acquire speed data.
  • the standard message data and utterance speed data are sent to the sound piece editing unit 8.
  • the speech unit editing unit 8 searches for all the compressed speech unit data associated with the phonetic characters that match the phonetic characters representing the reading of the speech units included in the fixed message.
  • the search unit 9 is instructed.
  • the search unit 9 searches the speech unit database 10 in response to the instruction of the speech unit editing unit 8, and finds the corresponding compressed speech unit data and the above-described speech unit reading associated with the corresponding compressed speech unit data. Data, speed initial value data and pitch component data are retrieved, and the retrieved compressed waveform data is supplied to the decompression unit 6. Even when a plurality of compressed speech piece data correspond to one speech piece, all of the corresponding compressed speech piece data are searched for as candidates for data used for speech synthesis. On the other hand, when there is a speech unit for which compressed speech unit data could not be found, the search unit 9 generates data for identifying the corresponding speech unit (hereinafter, referred to as missing portion identification data).
  • the decompression unit 6 restores the compressed speech unit data supplied from the retrieval unit 9 to the speech unit data before being compressed, and returns the data to the retrieval unit 9.
  • the search unit 9 uses the speech unit data returned from the decompression unit 6 and the retrieved speech unit reading data, the speed initial value data and the pitch component data as search results as speech speed conversion units 1 1 To supply. When the missing part identification data is generated, the missing part identification data is also supplied to the speech speed conversion unit 11.
  • the speech unit editing unit 8 converts the speech unit data supplied to the speech speed conversion unit 11 into the speech speed conversion unit 11, and determines the time length of the speech unit represented by the speech unit data. Instruct the user to match the speed indicated by the utterance speed data.
  • the speech speed conversion unit 11 responds to the instruction of the speech unit editing unit 8, converts the speech unit data supplied from the search unit 9 so as to match the instruction, and edits the speech unit. Supply to Part 8. Specifically, for example, the original time length of the speech unit data supplied from the search unit 9 is specified based on the searched speed initial value data, and the speech unit data is resampled. Then, the number of samples of the speech piece data may be set to a time length that matches the speed specified by the speech piece editing unit 8.
  • the speech speed conversion unit 11 also supplies the speech unit reading data, speed initial value data and pitch component data supplied from the retrieval unit 9 to the speech unit editing unit 8, and retrieves the missing part identification data. If supplied from 9, the missing part identification data is also supplied to the sound piece editing unit 8.
  • the speech unit editing unit 8 When the utterance speed data is not supplied to the speech unit editing unit 8, the speech unit editing unit 8 outputs the speech unit supplied to the speech speed conversion unit 11 to the speech speed conversion unit 11. What is necessary is just to instruct the speech unit editing unit 8 to supply the data without converting the data, and the speech speed conversion unit 11 responds to this instruction and outputs the speech unit data supplied from the search unit 9 as it is. It may be supplied to the one-side editing unit 8.
  • the speech unit editing unit 8 When the speech unit editing unit 8 is supplied with the speech unit data, the speech unit reading data, the speed initial value data and the pitch component data from the speech speed conversion unit 11, the speech unit editing unit 8 From the above, select one voice unit for each voice unit that represents the waveform that best approximates the waveform of the voice unit that composes the fixed message.
  • the speech unit editing unit 8 adds a prosody prediction such as a “Fujisaki model” or “To BI (To neand Break Indices)” to the fixed message represented by the fixed message data.
  • a prosody prediction such as a “Fujisaki model” or “To BI (To neand Break Indices)”
  • To BI To neand Break Indices
  • the speech unit editing unit 8 determines, for each speech unit in the fixed message, that the prediction result data representing the prediction result of the time change of the frequency of the pitch component of the speech unit matches the reading of the speech unit.
  • the correlation with pitch component data representing the time change of the frequency of the pitch component of the speech unit data representing the waveform of the speech unit to be performed is obtained.
  • the speech piece editing unit 8 calculates, for each of the pitch component data supplied from the speech speed conversion unit 11, for example, a value shown on the right side of Equation 1 and a value shown on the right side of Equation 2] 3 Ask. n '
  • Fig. 3 (a) As shown in Fig. 3 (a), as a linear function of the value X (i) (i is an integer) of the i-th sample of the prediction result data (the total number of samples is n) for a certain speech unit First-order regression is performed on the value of the i-th sample Y (i) of pitch component data (the total number of samples is assumed to be n) for the speech unit data representing the waveform of the speech unit whose reading matches that of this speech unit.
  • the gradient of this linear function has an intercept of ⁇ .
  • the unit of the slope may be, for example, [Hertz / second]
  • the unit of intercept / 3 may be, for example, [Hertz].
  • the total number of samples differs between the prediction result data and the pitch component data for the same reading speech unit, one (or both) of the two is replaced by linear interpolation, Lagrange interpolation, or any other method. It is sufficient to resample after interpolating by the method, and to obtain the correlation after aligning the total number of both samples.
  • the speech unit editing unit 8 uses the speed initial value data supplied from the speech speed conversion unit 11 and the fixed message data and utterance speed data supplied to the speech unit editing unit 8. Then, the value dt on the right side of Equation 3 is obtained.
  • This value dt is a coefficient representing the time difference between the utterance speed of the speech unit represented by the speech unit and the utterance speed of the speech unit in the fixed message whose reading matches the reading.
  • the evaluation value cost 1 is set to be larger as the correlation between the prediction result of the pitch of the voice unit and the pitch of the voice unit is higher, so that the reciprocal of the linear function of the value i 1 - ⁇ I Therefore, the evaluation value cost 1 becomes larger as the value I 1 - ⁇ I approaches 0.
  • the inflection of speech is characterized by the temporal change of the frequency of the pitch component of a speech unit. Therefore, the value of the slope H has the property of reflecting the difference in the intonation of the voice to the sensitivity.
  • the value of the intercept i3 is close to 0. Therefore, the value of [intercept] 3 has the property of sensitively reflecting the difference in the base pitch frequency of voice.
  • the evaluation value c ost 1 since the evaluation value c ost 1 has a form that can be regarded as the reciprocal of the linear function of the value I iS I, the evaluation value c ost 1 becomes larger as the value I ⁇ I becomes closer to 0.
  • the base pitch frequency of a voice is a factor that governs the voice quality of a voice speaker, and the gender difference of the speaker is remarkable.
  • the coefficient W Get a value of 2 It is desirable to make it larger.
  • the speech unit editing unit 8 selects the speech unit data representing a waveform close to the waveform of the speech unit in the fixed message, while identifying the missing part from the speech speed conversion unit 11. If data is also supplied, a phonetic character string representing the reading of the speech piece indicated by the missing part identification data is extracted from the fixed message data and supplied to the acoustic processing unit 4, and the waveform of this speech piece is obtained. Instruct them to combine.
  • the sound processing unit 4 treats the phonetic character string supplied from the speech unit editing unit 8 in the same manner as the phonetic character string represented by the distribution character string data.
  • compressed waveform data representing the waveform of the speech indicated by the phonetic characters included in the phonetic character string is retrieved by the search unit 5, and the compressed waveform data is restored by the decompression unit 6 to the original waveform data. It is restored and supplied to the sound processing unit 4 via the search unit 5.
  • the sound processing section 4 supplies the waveform data to the sound piece editing section 8.
  • the voice data editing unit 8 When the voice processing unit 4 returns the waveform data from the sound processing unit 4, the voice data editing unit 8 outputs the waveform data and the voice data editing unit 8 out of the voice data supplied from the speech speed conversion unit 11. Are combined with each other in the order according to the sequence of the sound pieces in the fixed message indicated by the fixed message data, and output as data representing the synthesized speech.
  • the speech unit editing unit 8 immediately specified it without instructing the sound processing unit 4 to synthesize a waveform.
  • the speech unit data may be combined with each other in the order according to the sequence of the speech units in the fixed message indicated by the fixed message data, and output as data representing the synthesized speech.
  • units larger than phonemes The speech unit data representing the waveform of the speech unit can be naturally spliced by the recording and editing method based on the prediction result of the prosody, and the speech that reads out the fixed message is synthesized.
  • the storage capacity of the speech unit database 10 can be made smaller than in the case where a waveform is stored for each phoneme, and a search can be performed at high speed. Therefore, this speech synthesis system can be configured to be small and lightweight, and can follow high-speed processing.
  • the correlation between the predicted result of the speech unit waveform and the speech unit data is evaluated using multiple evaluation criteria (for example, evaluation based on the gradient or intercept in the case of first-order regression, evaluation based on the time difference of the speech unit, etc.).
  • evaluation criteria for example, evaluation based on the gradient or intercept in the case of first-order regression, evaluation based on the time difference of the speech unit, etc.
  • discrepancies in the results of these evaluations can often occur.
  • the results of evaluations based on multiple evaluation criteria are integrated based on one evaluation value, and appropriate evaluation is performed.
  • waveform data ⁇ speech piece data does not need to be in PCM format data, and the data format is arbitrary.
  • the waveform database 7 ⁇ ⁇ ⁇ ⁇ ⁇ ⁇ speech data base 10 does not necessarily need to store the waveform data ⁇ speech data in a compressed state. Waveform data base 7 ⁇ Speech data base 10 If waveform data ⁇ Speech data is stored in an uncompressed state, it is not necessary for main unit M to have decompression unit 6. Absent.
  • the speech unit database creation unit 13 performs a new compression from the recording medium set in the recording medium drive unit (not shown) to be added to the speech unit data base 10 via the recording medium drive unit. It is also possible to read the sound piece data or the phonetic character string which is the material of the sound piece data.
  • the speech unit registration unit R is not necessarily the recorded speech unit data set. It is not necessary to have a memory 12.
  • the speech unit editing unit 8 stores in advance a prosody registration data representing the prosody of a specific speech unit, and if the specific message unit is included in the fixed message, the prosody registration data is stored.
  • the prosody represented may be treated as the result of prosody prediction.
  • the sound piece editing unit 8 may newly store the result of the past prosody prediction as prosody registration data.
  • the speech piece editing unit 8 calculates, for each pitch component data supplied from the speech speed conversion unit 11, for example, the value RX y shown on the right side of Expression 5 (j) is determined by taking the value of j as an integer from 0 to less than n to obtain a total of n, and among the n correlation coefficients from R xy (0) to R y (n-1) obtained, The maximum value may be specified. [ ⁇ X (i) -mx ⁇ ⁇ ⁇ Y j (i) -my ⁇ ]
  • R xy (j) is the prediction result for a certain sound piece (the total number of samples is n.
  • X (i) in Equation 5 is the same as that in Equation 1).
  • a sequence of samples obtained by cyclically shifting the pitch component data (total number of samples n) of the speech unit data representing the waveform of the matching speech unit by j in a fixed direction (note that Y j (i ) Is the value of the i-th sample in this sample column.)
  • FIG. 3 (b) is a graph showing an example of values of prediction result data and pitch component data used for obtaining values of R xy (0) and R xy (j).
  • the speech unit editing unit 8 performs a speech unit decoding process that represents a speech unit that matches the reading of the speech unit in the fixed message.
  • the value on the right-hand side of Equation 6 (evaluation value) that has the largest cost 2 may be selected.
  • Rm ax is the maximum value among the R xy (0) ⁇ R xy ( n one 1)
  • the sound piece editing unit 8 does not necessarily need to obtain the above-described correlation coefficient for the pitch component data obtained by performing various cyclic shifts.
  • the value of R xy (0) is used as the maximum value of the correlation coefficient as it is. It may be handled.
  • the evaluation values co s t 1 and c os t 2 may not include the term of the coefficient d t, and in this case, the speech piece editing unit 8 does not need to find the coefficient d t.
  • the speech unit editing unit 8 may use the value of the coefficient dt as it is as the evaluation value.
  • the speech unit editing unit may use the gradient ⁇ , the intercept] 3, There is no need to find the value of R xy (j).
  • the pitch component data may be data representing a temporal change of the pitch length of the sound piece represented by the sound piece data.
  • the speech unit editing unit 8 creates, as the prediction result data, data representing the prediction result of the time change of the pitch length of the speech unit, and generates the waveform of the speech unit whose reading matches the reading of the speech unit.
  • the correlation with the pitch component data representing the temporal change of the pitch length of the sound element data may be obtained.
  • the sound piece database creation unit 13 may include a microphone, an amplifier, a sampling circuit, an A / D (Analog-to-Digital) converter, a PCM encoder, and the like.
  • the sound piece data creation base section 13 instead of acquiring the sound piece data from the recorded sound piece data set storage section 12, the sound piece data creation base section 13 generates a sound representing the sound collected by its own microphone. After amplifying the signal, sampling and A / D converting it, PCM modulation may be applied to the sampled audio signal to create a sound unit data.
  • the speech unit editing unit 8 supplies the waveform data returned from the sound processing unit 4 to the speech speed conversion unit 11 so that the speech speed data indicates the time length of the waveform represented by the waveform data. It may be made to match with the one.
  • the speech unit editing unit 8 acquires free text data together with the language processing unit 1, for example, and obtains a speech unit data representing a waveform closest to the waveform of the speech unit included in the free text represented by the free text data. May be selected by performing substantially the same processing as the processing of selecting the speech piece data representing the waveform closest to the speech piece waveform included in the fixed message, and used for speech synthesis. ' In this case, the sound processing unit 4 does not need to cause the search unit 5 to search for waveform data representing the waveform of the speech unit represented by the speech unit selected by the speech unit editing unit 8. .
  • the sound piece editing unit 8 notifies the sound processing unit 4 of a sound unit that does not need to be synthesized by the sound processing unit 4, and the sound processing unit 4 responds to this notification to respond to the notification by rewriting the unit speech constituting the sound unit. What is necessary is to stop the search of the waveform of.
  • the speech unit editing unit 8 acquires distribution character string data together with, for example, the acoustic processing unit 4, and generates a speech unit representing a waveform closest to the waveform of the speech unit included in the distribution character string represented by the distribution character string data. It is also possible to select data by performing processing that is substantially the same as the processing of selecting a speech unit that represents the waveform closest to the waveform of the speech unit included in the fixed message, and use it for speech synthesis. Good. In this case, the sound processing unit 4 does not need to cause the search unit 5 to search for waveform data representing the waveform of the sound unit represented by the sound unit data selected by the sound unit editing unit 8. .
  • the physical configuration of the speech synthesis system according to the second embodiment of the present invention is substantially the same as the configuration in the above-described first embodiment.
  • the directory section DIR of the speech unit database 10 in the speech synthesis system includes, for example, as shown in FIG.
  • the data of A) to (D) are stored in association with each other, and instead of the data of (E) described above, (F) the compressed sound piece data is stored as pitch component data.
  • the data representing the pitch component frequencies at the beginning and end of the sound piece to be expressed are stored in a form associated with these (A) to (D) data. Have been.
  • Fig. 4 shows the data included in the data section DAT, which represents the waveform of a sound piece whose reading is "Saitama," as in Fig. 2.
  • the one-side data is stored at a logical position starting from the address 01 A36A6h.
  • at least the data of (A) in the above set of data of (A) to (D) and (F) are sorted according to the order determined based on the phonetic characters represented by the phoneme reading data. It is assumed that it is stored in the storage area of the sound piece database 10 in a single state.
  • the speech unit database creation unit 13 of the speech unit registration unit R reads out the phonograms and the speech unit data that are associated with each other from the recorded speech unit data set storage unit 12, and The utterance speed of the voice represented by the piece of data and the frequency of the pitch component at the beginning and end shall be specified.
  • the read speech piece data is supplied to the compression section 14, and when the compressed speech piece data is returned, the compressed speech piece data, the phonogram read out from the recorded speech piece data set storage section 12,
  • the first address in the storage area of the speech unit database 10 of the compressed speech unit data, the data length of the compressed speech unit data, and the speed initial value data indicating the specified utterance speed are stored in the first
  • the data is written in the storage area of the speech unit database 10, and data indicating the result of specifying the frequency of the pitch component at the beginning and end of the sound is stored. It is generated and written in the storage area of the speech piece database 10 as pitch component data.
  • the utterance speed and the frequency of the pitch component are specified, for example, as follows:
  • the method may be performed by substantially the same method as the method performed by the sound piece database creating unit 13 of the first embodiment.
  • the operation of the speech synthesis system of the first embodiment is as follows.
  • the operation is substantially the same as the operation to be performed.
  • the method by which the language processing unit 1 obtains free text data and the method by which the sound processing unit 4 obtains distribution character string data are arbitrary.
  • both methods are used in the language processing unit in the first embodiment.
  • Free text data or distribution character string data may be obtained by the same method as that performed by 1 and the sound processing unit 4.
  • the speech piece editing unit 8 has acquired the fixed message data and the utterance speed data.
  • the method by which the sound piece editing unit 8 acquires the fixed message data and the utterance speed data is also arbitrary.
  • the fixed message data may be obtained by the same method as the method performed by the sound unit editing unit 8 in the first embodiment. Or utterance speed data.
  • the speech unit editing unit 8 When the standard message data and the utterance speed data are supplied to the speech unit editing unit 8, the speech unit editing unit 8 is included in the standard message, similarly to the speech unit editing unit 8 in the first embodiment.
  • the search unit 9 is instructed to search all the compressed speech unit data associated with the phonogram corresponding to the phonogram representing the reading of the speech unit.
  • the speech unit conversion unit 11 converts the speech unit data supplied to the speech speed conversion unit 11 in the same manner as the speech unit editing unit 8 in the first embodiment, and converts the speech unit.
  • the time length of the speech unit represented by the data matches the speed indicated by the utterance speed data. Instructs them to do so.
  • the search unit 9, the expansion unit 6, and the speech speed conversion unit 11 perform substantially the same operation as the operation of the search unit 9, the expansion unit 6, and the speech speed conversion unit 11 in the first embodiment.
  • the speech unit data, the speech unit reading data, and the pitch component data are supplied from the speech speed conversion unit 11 to the speech unit editing unit 8.
  • the missing part identification data is supplied from the search unit 9 to the speech speed conversion unit 11, the missing part identification data is also supplied to the speech piece editing unit 8.
  • the speech unit editing unit 8 When the speech unit editing unit 8 is supplied with the speech unit data, the speech unit reading data, and the pitch component data from the speech speed conversion unit ⁇ 1, the speech unit editing unit 8 performs a fixed form from among the supplied speech unit data according to the procedure described below. Select one piece of voice data for each voice piece that represents the waveform that can be regarded as the waveform of the voice piece that composes the message.
  • the speech unit editing unit 8 first and last of the speech unit data supplied from the speech speed conversion unit 11 Specify the frequency of the pitch component at each point in time. Then, from the speech unit data supplied from the speech speed conversion unit 11, the absolute value of the frequency difference of the pitch component at the boundary between adjacent speech units in the fixed message is accumulated for the entire fixed message. Select the voice unit to satisfy the condition that the resulting value is minimized.
  • FIGS. 5 (a) to 5 (d) The conditions for selecting the voice unit will be described with reference to FIGS. 5 (a) to 5 (d).
  • a fixed message data representing a fixed message reading "This Saki Mikika-bu-da” is supplied to the sound piece editing unit 8, and
  • the fixed message consists of three sound pieces, "Konosaki”, “Migikaichi” and "is” Shall be.
  • the speech unit de-even base 10 has three compressed speech unit data whose readings are "Konosaki"("A" in Fig. 5 (b)).
  • Fig. 5 (c) shows, for example, the difference between the frequency of the pitch component at the end of the speech unit represented by the speech unit A1 and the frequency of the pitch component at the beginning of the speech unit represented by the speech unit data B1. This indicates that the absolute value is “1 2 3.”
  • the unit of this absolute value is, for example, “Hertz”.
  • the speech piece editing unit 8 selects the speech piece data A3, B2, and C2 as shown in FIG. 5 (d).
  • the speech unit editing unit 8 sets, for example, the absolute value of the frequency difference between the pitch components at the boundary between adjacent speech units in the fixed message as the distance. What is necessary is just to define and select the sound piece by the DP (Dynamic Programming) matching method.
  • the speech unit editing unit 8 converts the phonetic character string representing the reading of the speech unit indicated by the missing portion identification data into a fixed message.
  • the data is extracted from the data and supplied to the acoustic processing unit 4 to instruct to synthesize the waveform of the sound piece.
  • the sound processing unit 4 treats the phonetic character string supplied from the speech unit editing unit 8 in the same manner as the phonetic character string represented by the distribution character string data.
  • compressed waveform data representing the waveform of the voice indicated by the phonogram contained in the phonogram string is retrieved by the search unit 5, and the compressed waveform data is converted into the original waveform data by the decompression unit 6. Is restored and supplied to the sound processing unit 4 via the search unit 5.
  • the sound processing section 4 supplies the waveform data to the sound piece editing section 8.
  • the voice unit editing unit 8 selects the waveform data and the voice unit editing unit 8 out of the voice unit data supplied from the speech speed conversion unit 11. These are combined with each other in the order according to the sequence of the sound pieces in the fixed message indicated by the fixed message data, and output as data representing the synthesized speech.
  • the sound processing unit is used as in the first embodiment.
  • the speech unit data selected by the speech unit editing unit 8 is immediately combined with the sequence of each of the speech units in the standard message indicated by the standard message without immediately instructing the synthesis of the waveform in 4. Then, it may be output as a data representing synthesized speech.
  • the cumulative total of the amount of the discontinuous change in the frequency of the pitch component at the boundary between the speech units is minimized in the entire fixed message.
  • the sound unit is selected so that it can be connected naturally by the recording and editing method, so that the synthesized speech becomes natural.
  • this speech synthesis system does not perform prosody prediction with complicated processing, and can follow high-speed processing with a simple configuration.
  • the configuration of the speech synthesis system according to the second embodiment is not limited to the configuration described above.
  • the pitch component data may be data representing a pitch length at the beginning and end of the sound piece represented by the sound piece data.
  • the speech unit editing unit 8 determines the pitch length at the beginning and end of each speech unit data supplied from the speech speed conversion unit 11 based on the pitch component data supplied from the speech speed conversion unit 11. If the speech unit data is selected so as to satisfy the condition that the absolute value of the pitch length difference at the boundary between adjacent speech units in the fixed message is minimized over the entire fixed message, Good.
  • the speech unit editing unit 8 acquires free text data together with the language processing unit 1, and determines the speech unit data representing a waveform that can be regarded as a waveform of a speech unit included in the free text represented by the free text data. By extracting the speech unit data representing the waveform that can be regarded as the waveform of the speech unit included in the type message. And may be used for speech synthesis.
  • the sound processing unit 4 does not need to cause the search unit 5 to search for waveform data representing the waveform of the speech unit represented by the speech unit data extracted by the speech unit editing unit 8. .
  • the sound piece editing unit 8 notifies the sound processing unit 4 of a sound unit that does not need to be synthesized by the sound processing unit 4, and the sound processing unit 4 responds to this notification to respond to the notification by rewriting the unit speech constituting the sound unit. What is necessary is to stop the search of the waveform of.
  • the sound piece editing unit 8 acquires distribution character string data together with, for example, the sound processing unit 4, and generates a sound representing a waveform that can be regarded as a waveform of a sound unit included in the distribution character string represented by the distribution character string data.
  • the segment data is extracted by performing substantially the same processing as that for extracting the speech unit data representing the waveform that can be regarded as the waveform of the speech unit included in the fixed message, and can be used for speech synthesis. Good.
  • the sound processing unit 4 does not need to cause the search unit 5 to search for waveform data representing the waveform of the sound unit represented by the sound unit data extracted by the sound unit editing unit 8. .
  • the physical configuration of the speech synthesis system according to the third embodiment of the present invention is substantially the same as the configuration in the above-described first embodiment.
  • the operation when the language processing unit 1 of the speech synthesis system obtains the free text data from outside and the sound processing unit 4 obtains the distribution character string data from the outside are described in the first or second implementation.
  • the operation is substantially the same as that performed by the speech synthesis system of the first embodiment.
  • the language processing unit 1 acquires free text data and the sound processing unit 4 uses The method of acquiring the evening is arbitrary.
  • the method of obtaining free text data or the delivery character by the same method as the method performed by the language processing unit 1 or the sound processing unit 4 in the first or second embodiment is used. All you have to do is get the column data.
  • the speech piece editing unit 8 has acquired the fixed message data and the utterance speed data.
  • the method by which the sound piece editing unit 8 acquires the fixed message data—evening and utterance speed data is also optional.
  • the sound pattern editing unit 8 uses the same method as the method performed by the sound unit editing unit 8 in the first embodiment. What is necessary is just to acquire message data and utterance speed data.
  • the speech synthesis system is a part of an in-vehicle system such as a force navigation system, and other devices constituting the in-vehicle system (for example, performing speech recognition and performing speech recognition).
  • the Device that performs agent processing based on the information obtained as a result of recognition), determines the content and speed of speech to the user, and generates data representing the result of the determination.
  • the speech synthesis system may receive (acquire) the generated data and handle the data as fixed message data and utterance speed data.
  • the sound unit editing unit 8 is included in the fixed message, similarly to the sound unit editing unit 8 in the first embodiment.
  • the search unit 9 is instructed to search for all the compressed speech unit data associated with the phonetic characters that match the phonetic characters representing the reading of the speech unit.
  • the speech rate conversion section 11 converts the speech piece data supplied to the speech rate conversion section 11 in the same manner as the speech piece editing section 8 in the first embodiment, and converts the speech The time length of the sound piece represented by the piece data matches the speed indicated by the utterance speed data. Instructs them to do so.
  • the search unit 9, the expansion unit 6, and the speech speed conversion unit 11 perform substantially the same operation as the operation of the search unit 9, the expansion unit 6, and the speech speed conversion unit 11 in the first embodiment.
  • the speech speed conversion unit 11 sends the speech unit editing unit 8 to the speech unit data, the speech unit reading data, the speed initial value data representing the speech speed of the speech unit represented by the speech unit data, and the pitch. Ingredient data is provided. Further, when the missing portion identification data is supplied from the search unit 9 to the speech speed conversion unit 11, the missing portion identification data is also supplied to the speech unit editing unit 8.
  • the speech unit editing unit 8 When the speech unit editing unit 8 receives the speech unit data, the speech unit reading data, and the pitch component data from the speech speed conversion unit 11, the speech unit editing unit 8 performs the above-described processing on each pitch component data supplied from the speech speed conversion unit 11. Using the initial value of the speed and the fixed message data and the utterance speed data supplied to the sound piece editing unit 8, Find the value dt.
  • the speech unit editing unit 8 calculates, for each of the speech unit data supplied from the speech speed conversion unit 11, the ⁇ of the speech unit data obtained by itself (hereinafter, referred to as the speech unit data X). , ⁇ , Rmax, and dt, and the sound data (hereinafter referred to as sound data Y) that represents the sound data adjacent to the sound data represented by the sound data in the fixed message.
  • the evaluation value ⁇ ⁇ shown in Expression 7 is specified based on the frequency of the pitch component.
  • H XY (W A -cost-A) + (W B -cost-B) + (W c -cost-C)
  • the value cost—A included in the right-hand side of Equation 7 is the pitch component of the pitch component at the boundary between the speech unit represented by speech unit data X and the speech unit represented by speech unit data Y, which are adjacent to each other in the fixed message.
  • the speech unit editing unit 8 determines the cost-A value based on the pitch component data supplied from the speech speed conversion unit 11 in order to identify the value of cost-A. 11.
  • the frequency of the pitch component at each of the beginning and end of each piece of sound piece data supplied from 1 may be specified.
  • the value c os t—B included in the right-hand side of Expression 7 is a value obtained when the evaluation value c os t—B is obtained for the voice unit X according to Expression 8.
  • cost _B 1 / (W B 1 I 1-a I + W B 2 I ⁇ I + W B 3
  • the value co st — C included in the right-hand side of Equation 7 is a value obtained when the evaluation value co st — C of the voice unit X is calculated according to Equation 9.
  • cost _C 1 / (W c 1 I Rm ax I + W C 2 -dt)
  • the speech piece editing section 8 in place of Equation 7 to Equation 9, may be specified evaluation value Eta chi gamma according to Equation 1 0) and (1 1.
  • any value of the coefficient W B 3 and W c 3 above is 0.
  • the terms (W B 3 ⁇ dt) and (W C 2 ⁇ dt) in Equations 8 and 9 need not be provided.
  • ⁇ ⁇ ( ⁇ -cost— A) + (W B -cosf _B) + (W c- cost—C) + (W D -cost—D)
  • W D is a predetermined coefficient that is not 0
  • W d is a predetermined coefficient that is not 0
  • the speech unit editing unit 8 generates, from among the speech unit data supplied from the speech speed conversion unit 11, the speech unit 1 forming the standard message represented by the standard message data supplied to the speech unit editing unit 8.
  • the speech unit editing unit 8 synthesize the speech that reads out the fixed message with the largest sum of the evaluation value ⁇ ⁇ of each piece of speech piece data belonging to the combination To select the best combination of sound pieces for the night.
  • a fixed message representing a fixed message message is composed of speech pieces A, B, and C, and is used as a candidate for speech piece data representing speech piece A.
  • A1, A2, and A3 are found, and speech unit data B1 and B2 are found as candidates for speech unit data representing speech unit B, and speech units are found as candidates for speech unit data representing speech unit C.
  • the combination includes a speech unit data P representing the speech unit p and a speech unit data Q representing the speech unit Q.
  • the speech unit P precedes the speech unit q.
  • the evaluation value HpQ when the speech unit p is adjacent to the speech unit is used.
  • the sound piece editing unit 8 treats the value of (W A ⁇ cost—A) as being 0, while The values of the coefficients W B , W c, and W D are each treated as a predetermined value different from the case of calculating the evaluation value ⁇ ⁇ ⁇ of other sound piece data.
  • the speech unit editing unit 8 calculates the evaluation value indicating the relationship between the speech unit data X and the adjacent speech unit data Y before the speech unit represented by the speech unit data X using Expression 7 or Expression 11.
  • the evaluation value H XY may be specified as including XY . In this case, the value of cost- ⁇ cannot be determined because there is no preceding speech unit for the first speech unit of the fixed message.
  • the sound piece editing unit 8 treats the value of (W ⁇ ⁇ cost- ⁇ ) as being 0, while The values of the coefficients W B , W c, and W D may be respectively treated as predetermined values different from the case of calculating the evaluation value ⁇ ⁇ of other sound piece data.
  • the speech piece editing unit 8 outputs the missing part identification data from the speech speed conversion unit 11. Is also supplied, the phonetic character string representing the reading of the speech piece indicated by the missing part identification data is extracted from the fixed message data and supplied to the acoustic processing unit 4, and the waveform of this speech piece is synthesized. To do so.
  • the sound processing unit 4 treats the phonetic character string supplied from the speech unit editing unit 8 in the same manner as the phonetic character string represented by the distribution character string data.
  • compressed waveform data representing the voice waveform indicated by the phonetic characters included in this phonetic character string is retrieved by the search unit 5, and the compressed waveform data is restored to the original waveform data by the decompression unit 6.
  • the data is supplied to the sound processing unit 4 via the search unit 5.
  • the sound processing section 4 supplies the waveform data to the sound piece editing section 8.
  • Speech piece editing section 8 when it is sent back to the waveform data from the acoustic processing unit 4, and the waveform data, speech speed converting section 1 1 supplied speech piece data sac Chi than, the sum of the evaluation values Eta chi gamma
  • the data that represents the synthesized speech is combined with the one that belongs to the combination selected by the speech unit editing unit 8 as the largest combination in the order of each speech unit in the standard message indicated by the standard message Is output as
  • the sound processing unit 4 must be instructed to synthesize a waveform, as in the first embodiment.
  • the speech unit selected by the speech unit editing unit 8 is combined with each other in the order according to the sequence of each speech unit in the standard message indicated by the standard message data, and output as data representing the synthesized voice do it.
  • the speech units are spliced together naturally by the recording and editing method, and the speech for reading the fixed message is synthesized.
  • One piece of sound The storage capacity of 10 can be smaller than that of storing a waveform for each phoneme, and a high-speed search can be performed. Therefore, this speech synthesis system can be configured to be small and lightweight, and can follow high-speed processing.
  • various evaluation criteria for evaluating the appropriateness of a combination of speech piece data selected for synthesizing speech for reading a fixed message are provided.
  • the configuration of the speech synthesis system according to the third embodiment is not limited to the configuration described above.
  • the evaluation values used by the speech unit editing unit 8 to select the optimum combination of speech unit data are not limited to those shown in Expressions 7 to 13, and are obtained by combining the speech units represented by the speech unit data with each other. It may be any value that represents an evaluation of how similar or dissimilar the speech that is made to the speech uttered by a person.
  • evaluation expression representing the evaluation value necessarily an expression
  • the evaluation expression is not limited to those included in ⁇ 13, and the evaluation expression can be obtained by arbitrarily setting parameters representing the characteristics of the sound unit represented by the sound unit data, or by combining the sound units together.
  • a mathematical expression including an arbitrary parameter indicating a feature of the voice or an optional parameter indicating a feature expected to be included in the voice when a person utters the voice is used. May be.
  • the criterion for selecting the optimal combination of speech unit data does not necessarily need to be one that can be expressed in the form of an evaluation value, and is obtained by combining the speech units represented by the speech unit data with each other. Any criteria are possible as long as the criteria lead to the determination of the optimal combination of speech piece data based on an evaluation of how similar or different the speech is to human speech.
  • the speech unit editing unit 8 acquires free text data together with the language processing unit 1, for example, and generates speech unit data representing a waveform that can be regarded as a waveform of a speech unit included in the free text represented by the free text data. Alternatively, it may be extracted by performing substantially the same processing as the processing of extracting speech piece data representing a waveform that can be regarded as a speech piece waveform included in a fixed message, and used for speech synthesis. In this case, the sound processing unit 4 does not need to cause the search unit 5 to search for waveform data representing the waveform of the sound unit represented by the sound unit data extracted by the sound unit editing unit 8.
  • the sound piece editing unit 8 notifies the sound processing unit 4 of a sound unit that does not need to be synthesized by the sound processing unit 4, and the sound processing unit 4 responds to the notification to generate a unit constituting the sound unit.
  • the search for the audio waveform may be stopped.
  • the sound piece editing unit 8 acquires the distribution character string data together with the sound processing unit 4, for example, and generates the sound unit data representing the waveform that can be regarded as the waveform of the sound unit included in the distribution character string represented by the distribution character string data. May be extracted by performing substantially the same processing as the processing of extracting speech piece data representing a waveform that can be regarded as the waveform of the speech piece included in the fixed message, and used for speech synthesis.
  • the sound processing unit 4 performs, for the sound unit represented by the sound unit data extracted by the sound unit editing unit 8, a waveform representing the waveform of the sound unit. It is not necessary for the search unit 5 to search for shape data.
  • the audio data selecting device can be realized using a normal computer system, not a dedicated system.
  • a language processing unit 1 For example, in a personal computer, a language processing unit 1, a general word dictionary 2, a user word dictionary 3, a sound processing unit 4, a search unit 5, an expansion unit 6, a waveform database 7, a speech unit in the first embodiment described above.
  • a medium CD-ROM, MII, floppy (registered trademark) disk, etc.
  • a program for causing a personal computer to execute the operations of the recorded sound piece data set storage section 12, the sound piece data base creation section 13 and the compression section 14 in the first embodiment described above By installing the program from the medium storing the ram, the personal computer can perform the function of the sound piece registration unit R of the above-described first embodiment.
  • a personal computer that executes these programs and functions as the main unit unit M and the speech unit registration unit R in the first embodiment is executed as a process corresponding to the operation of the speech synthesis system in FIG.
  • the processing shown in FIGS. 6 to 8 is to be performed.
  • FIG. 6 is a flowchart showing processing when the personal computer acquires free text data.
  • Fig. 7 shows that this personal computer 9 is a flowchart showing a process when the information is obtained.
  • FIG. 8 is a flowchart showing a process when the personal computer acquires the fixed message data and the utterance speed data.
  • step S101 when the personal computer obtains the above-described free text data from outside (FIG. 6, step S101), for each ideographic character included in the free text represented by the free text data, The phonogram representing the reading is specified by searching the general word dictionary 2 and the user word dictionary 3, and the ideogram is replaced with the specified phonogram (step S102).
  • the method by which the personal computer acquires the free text data is arbitrary.
  • each phonogram included in the phonogram string is obtained.
  • the waveform of the unit speech represented by the phonetic character is searched from the waveform database 7, and the compressed waveform data representing the waveform of the unit speech represented by each phonetic character included in the phonetic character string is retrieved ( Step S103).
  • the personal computer restores the extracted compressed wave data to the waveform data before compression (step S104), and converts the restored waveform data into a phonetic character string.
  • the phonograms in the sequence are combined with each other in the same order and output as synthesized speech data (step S105).
  • the method by which this personal computer outputs synthesized speech data is arbitrary.
  • this personal computer receives the above-mentioned distribution statement from outside.
  • the character string data is obtained by an arbitrary method (FIG. 7, step S201)
  • the unit represented by the phonetic character The voice waveform is searched from the waveform database 7, and compressed waveform data representing the unit voice waveform represented by each phonogram included in the phonogram string is retrieved (step S202).
  • the personal computer restores the extracted compressed waveform data to the waveform data before compression (step S203), and converts the restored waveform data into a phonetic character string.
  • the phonograms in the sequence are combined with each other in the same order, and are output as synthesized speech data by the same processing as the processing in step S105 (step S204).
  • step S301 when the personal computer obtains the above-mentioned fixed message data and the utterance speed data from an external device by any method (FIG. 8, step S301), first, the fixed message data is obtained. All compressed speech piece data associated with phonograms that match the phonetic readings contained in the fixed message included in the evening message are retrieved (step S302).
  • step S302 the above-mentioned speech piece reading data, speed initial value data, and pitch component data associated with the corresponding compressed speech piece data are also retrieved. If more than one piece of compressed speech data corresponds to a single speech piece, search for the entire compressed speech piece data. On the other hand, when there is a speech unit for which compressed speech unit data cannot be found, the above-described missing portion identification data is generated.
  • the personal computer restores the retrieved compressed speech piece data to the speech piece data before being compressed (step S303). Then, the reconstructed speech unit data is converted by the same processing as that performed by the speech unit editing unit 8 described above, and the time length of the speech unit represented by the speech unit data matches the speed indicated by the utterance speed data. (Step S304). When the utterance speed data is not supplied, the restored speech piece data need not be converted.
  • the personal computer converts the speech unit data representing the waveform closest to the waveform of the speech unit constituting the fixed message from the speech unit data in which the time length of the speech unit has been converted to the above-described speech unit.
  • one sound piece is selected one by one (steps S305 to S308).
  • the personal computer predicts the prosody of the fixed message by adding an analysis based on the prosody prediction method to the fixed message represented by the fixed message data (step S305). Then, for each of the sound pieces in the fixed message, a prediction result of the time change of the frequency of the pitch component of this sound piece and a sound piece data representing the waveform of the sound piece whose reading matches that of this sound piece.
  • the correlation with the pitch component data representing the time change of the frequency of the pitch component of is obtained (step S306). More specifically, for each of the retrieved pitch component data, for example, the values of the above-described gradient a and intercept] 3 are obtained.
  • the personal computer obtains the above value d t using the retrieved initial speed value data, the fixed message data and the utterance speed data obtained from outside (step S 30).
  • the personal computer calculates the value of ⁇ obtained in step S306 and the value of 'dt' obtained in step S307. Then, a speech piece data representing a speech piece that matches the reading of the speech piece in the fixed message is selected from those having the largest evaluation value c st1 (step S308).
  • this personal computer may determine the above-described maximum value of R Xy (j) in step S306 instead of obtaining the values of ⁇ and 3 described above.
  • step S308 based on the maximum value of RXy (j) and the coefficient dt obtained in step S307, the speech unit that matches the reading of the speech unit in the fixed message is used.
  • the one that maximizes the above-mentioned evaluation value c 0 st 2 may be selected from the sound elements that represent
  • the personal computer extracts a phonetic character string representing the reading of the sound piece indicated by the missing part identifying data from the fixed message data, and generates a phoneme for this phonetic character string.
  • Each phonetic character in this phonetic character string is treated in the same way as the phonetic character string represented by the distribution character string data and processed in steps S202 to S203 described above.
  • the waveform data representing the waveform of the indicated voice is restored (step S309).
  • the personal computer compares the restored waveform data and the sound piece data selected in step S308 in the order according to the order of each sound piece in the fixed message indicated by the fixed message message. These are combined with each other and output as data representing synthesized speech (step S310).
  • the personal computer has a language processing unit 1, a general word dictionary 2, a user word dictionary 3, a sound processing unit 4, a search unit 5, a decompression unit 6, and a waveform data base in the second embodiment described above. 7, the operation of the speech unit editing unit 8, the search unit 9, the speech unit data base 10 and the speech speed conversion unit 11 'is executed.
  • a program for causing a personal computer to execute the operations of the recorded speech unit data set storage unit 12, the speech unit database creation unit 13 and the compression unit 14 in the second embodiment described above By installing the program from the medium storing the sound unit, the personal computer can perform the function of the sound piece registration unit R in the above-described second embodiment.
  • a personal computer that executes these programs and functions as the main unit unit M and the speech unit registration unit R in the second embodiment is executed as processing corresponding to the operation of the speech synthesis system in FIG.
  • the above-described processing shown in FIGS. 6 and 7 is performed, and the processing shown in FIG. 9 is performed.
  • FIG. 9 is a flowchart showing a process when the personal computer acquires the fixed message data and the utterance speed data.
  • step S402 when the personal computer obtains the above-mentioned fixed message data and utterance speed data from an external device by an arbitrary method (FIG. 9, step S401), first, the above-described step S302 is performed. Similarly to the processing, the compressed speech piece data in which the phonogram matching the phonogram representing the reading of the speech piece included in the fixed message represented by the fixed message data is associated with the corresponding compressed speech piece data. The above-mentioned speech unit reading data and speed initial value data And all the pitch component data are retrieved (step S402). Note that, even in step S402, if one compressed sound piece data is extracted from one sound piece, the entire compressed sound piece data is searched for. If there is a speech piece that could not be found, the above-described missing portion identification data is generated.
  • the personal computer restores the retrieved compressed speech data to the original speech data before compression (step S403), and restores the restored speech data. Then, conversion is performed by the same processing as that performed by the speech unit editing unit 8 described above, and the time length of the speech unit represented by the speech unit data is matched with the speed indicated by the utterance speed data (step S404) ). If the utterance speed data is not supplied, the restored speech piece data need not be converted.
  • the personal computer converts the speech unit data representing the waveform that can be regarded as the waveform of the speech unit constituting the fixed message from the speech unit data obtained by converting the time length of the speech unit into the second unit described above.
  • the personal computer converts the speech unit data representing the waveform that can be regarded as the waveform of the speech unit constituting the fixed message from the speech unit data obtained by converting the time length of the speech unit into the second unit described above.
  • this personal computer calculates the pitch component frequency at each of the beginning and end of each piece of sound piece data in which the time length of the sound piece has been converted. It is specified based on the evening (step S405). Then, the condition that the sum of the absolute values of the frequency differences of the pitch components at the boundaries between adjacent sound units in the fixed message in the fixed message among these sound unit data is minimized.
  • the speech piece data is selected so as to satisfy the condition (step S406).
  • the absolute value of the difference between the frequencies of the pitch components at the boundary between adjacent speech units in the fixed message is defined as the distance, and the DP unit is used to select the speech unit by the DP matching method. Just fine.
  • the personal computer when the personal computer generates the missing part identification data, the personal computer extracts a phonetic character string representing the reading of the sound piece indicated by the missing part identifying data from the standard message and reads the phonetic character string.
  • Each phoneme is treated in the same manner as the phonetic character string represented by the delivery character string data, and the processing in steps S202 to S203 described above is performed, whereby each table in the phonetic character string is processed.
  • the waveform data representing the waveform of the voice indicated by the phonetic character is restored (step S407).
  • the personal computer compares the restored waveform data and the sound piece data selected in step S406 with the order of each sound piece in the fixed message indicated by the fixed message message. And output as data representing the synthesized speech (step S408).
  • a personal computer is provided with a language processing unit 1, a general word dictionary 2, a user word dictionary 3, a sound processing unit 4, a search unit 5, a decompression unit 6, a waveform data base 7 according to the third embodiment.
  • a medium storing a program for causing a personal computer to execute the operations of the recorded speech unit data set storage unit 12, the speech unit database creation unit 13 and the compression unit 14 in the third embodiment described above.
  • the personal computer can perform the function of the sound piece registration unit R in the third embodiment described above.
  • a personal computer that executes these programs and functions as the main unit unit M and the speech unit registration unit R in the third embodiment performs processing corresponding to the operation of the speech synthesis system in FIG.
  • the above-described processing shown in FIGS. 6 and 7 is performed, and the processing shown in FIG. 10 is performed. .
  • FIG. 10 is a flowchart showing a process when the personal computer obtains the fixed message data and the utterance speed data.
  • step S501 when the personal computer obtains the above-mentioned fixed message data and the utterance speed data from the outside by any method (FIG. 10, step S501), first, the above-mentioned step S501 is executed. Similarly to the process of 302, a compressed speech unit decoder in which a phonogram matching a phonogram representing a reading of a speech unit included in the fixed message represented by the fixed message data is associated with the compressed message unit. All the above-mentioned speech piece reading data, speed initial value data and pitch component data associated with the corresponding compressed speech piece data are retrieved (step S502).
  • step S502 if a plurality of compressed sound piece data are included in one sound piece, the corresponding compressed sound piece data is searched for, and one of the compressed sound piece data is searched for. If there is a voice piece that could not be found overnight, the above-mentioned missing part identification data is generated.
  • the personal computer restores the extracted compressed speech piece data to the speech piece data before being compressed (step S503),
  • the restored speech piece data is converted by the same processing as that performed by the speech piece editing unit 8 described above, and the time length of the speech piece represented by the speech piece data is made to match the speed indicated by the utterance speed data (Ste S504). If the utterance speed data is not supplied, the restored speech unit may not be converted.
  • the personal computer determines the optimal combination of the speech unit data for synthesizing the voice to read the fixed message from the speech unit data in which the time length of the speech unit is converted, according to the third embodiment described above.
  • the selection is performed by performing the same processing as the processing performed by the sound piece editing unit 8 in the mode (steps S505 to S507).
  • this personal computer obtains the above-mentioned value, the set of ⁇ and / or Rmax for each pitch component data searched out in step S502, and the speed initial value data,
  • the above-mentioned value dt is obtained using the fixed message data and the utterance speed data obtained in step S501 (step S501).
  • the personal computer calculates the values of a, ⁇ , Rmax, and dt obtained in step S505 for each of the speech piece data converted in step S504, and outputs the values in the fixed message.
  • the above-mentioned evaluation value ⁇ ⁇ is specified based on the frequency of the pitch component of the sound piece data representing the sound piece adjacent to the sound piece represented by the sound piece data (step S506 ) o
  • the personal computer computes the sound unit 1 ⁇ which constitutes the fixed message represented by the fixed message data acquired in step S501 from the sound unit data converted in step S504.
  • the speech with the largest total sum of the evaluation values ⁇ ⁇ ⁇ ⁇ of each piece of speech data belonging to the combination is synthesized as a voice reading a fixed message (Step S507).
  • the evaluation value ⁇ ⁇ ⁇ used to calculate the sum a value that correctly reflects the connection relation of the sound pieces in the combination is selected.
  • the personal computer when the personal computer generates the missing part identification data, the personal computer extracts a phonetic character string representing the reading of the speech piece indicated by the missing part identifying data from the standard message data, and By processing the above steps S202 to S203 for each phoneme in the same way as the phonetic sequence represented by the distribution character string data, each phoneme in this phonetic string is Restores waveform data representing the waveform of the voice indicated by the phonetic characters
  • the personal computer compares the restored waveform data and the sound piece data belonging to the combination selected in step S507 with the order of each sound piece in the fixed message indicated by the fixed message data. And output as data representing synthesized speech
  • a program that causes a personal computer to perform the functions of the main unit unit M and the speech unit registration unit R is, for example, a communication board bulletin board.
  • BTS Backbone System
  • a device that modulates a carrier with signals representing these programs transmits the resulting modulated wave, and receives this modulated wave May restore the programs by demodulating the modulated wave. Then, by starting these programs and executing them in the same manner as other application programs under the control of the OS, the above-described processing can be executed.
  • the program excluding the part is stored in the recording medium. May be. Also in this case, in the present invention, it is assumed that the recording medium stores a program for executing each function or step executed by the computer.
  • a voice selecting device a voice selecting method, and a program for obtaining a natural synthesized voice at high speed with a simple configuration are realized.

Landscapes

  • Engineering & Computer Science (AREA)
  • Computational Linguistics (AREA)
  • Health & Medical Sciences (AREA)
  • Audiology, Speech & Language Pathology (AREA)
  • Human Computer Interaction (AREA)
  • Physics & Mathematics (AREA)
  • Acoustics & Sound (AREA)
  • Multimedia (AREA)
  • Machine Translation (AREA)
  • Information Retrieval, Db Structures And Fs Structures Therefor (AREA)
  • Compression, Expansion, Code Conversion, And Decoders (AREA)

Abstract

 本発明は、簡単な構成で高速に自然な合成音声を得るための音声データ選択装置等を提供するものである。本発明の音声データ選択装置においては、定型メッセージを表すデータが供給されると、音片編集部は、定型メッセージ内の音片と読みが合致する音片の音片データを音片データベースから索出させる。一方で音片編集部は定型メッセージの韻律予測を行い、索出された音片データのうちから定型メッセージ内の各音片に最もよく合致するものを1個ずつ、評価式に基づいて特定する。評価式は、韻律予測結果−音片データ間におけるピッチ成分の周波数の1次回帰の結果や発声スピードの時間差を変数とするものである。そして、特定した音片データや、特定ができないため代わりに音響処理部に供給させた波形データを互いに結合して、合成音声を表すデータを生成する。

Description

明 細 書 音声データを選択するための装置、 方法およびプログラム
技術分野
本発明は、 音声データ選択装置、 音声データ選択方法及びプログ ラムに関する。
背景技術
• 音声を合成する手法として、録音編集方式と呼ばれる手法がある。 録音編集方式は、 駅の音声案内システムや、 車載用のナビゲーショ ン装置などに用いられている。
録音編集方式は、 単語と、 この単語を読み上げる音声を表す音声 データとを対応付けておき、 音声合成する対象の文章を単語に区切 つてから、 これらの単語に対応付けられた音声データを取得してつ なぎ合わせる、 という手法である。
この録音編集方式については、 例えば、 特開平 1 0— 4 9 1 9 3 号公報 (以降、 文献 1 と呼ぶ) において詳細に説明されている。
しかし、 音声データを単につなぎ合わせた場合、 音声データ同士 の境界では通常、音声のピッチ成分の周波数が不連続的に変化する、 等の理由で、 合成音声が不自然なものとなる。
この問題を解決する手法としては、 同一の音素を互いに異なった 韻律で読み上げる音声を表す複数の音声データを用意し、 一方で音 声合成する対象の文章に韻律予測を施して、 予測結果に合致する音 声デ—夕を選び出してつなぎ合わせる、 という手法が考えられる。 しかし、 音声データを音素毎に用意して録音編集方式により自然 な合成音声を得ようとすると、 音声データを記憶する記憶装置には 膨大な記憶容量が必要となり、 小型軽量な装置を用いる必要がある 用途には適さない。 また、 検索する対象のデータの量も膨大なもの となるから、 高速な処理が要求される用途にも適さない。
また、 韻律予測は極めて複雑な処理であるので、 韻律予測を用い たこの手法を実現するには、処理能力が高いプロセッサなどを用い、 あるいは長時間をかけて処理を行わせる必要がある。 従ってこの手 法は、 構成が簡単な装置を用いた高速な処理が要求される用途には 適さない。
この発明は、 上記実状に鑑みてなされたものであり、 簡単な構成 で高速に自然な合成音声を得るための音声データ選択装置、 音声デ ―夕選択方法及びプログラムを提供することを目的とする。
発明の開示
( 1 ) 上記発明目的を達成するために、 本発明の音声データ選択装 置は、 第 1の局面においては、 基本的に、 音声の波形を表す音声デ 一夕を複数記憶する記憶手段と、 文章を表す文章情報を入力し、 各 前記音声データのうちから、 前記文章を構成する音片と読みが共通 する音片の波形を表している音声データを索出する検索手段と、 索 出された音声デ一夕のうちから、 前記文章を構成するそれぞれの音 片に相当する音声データを 1個ずつ、 互いに隣接する音片同士の境 界でのピッチの差を前記文章全体で累計した値が最小となるように 選択する選択手段と、 から構成される。
前記音声データ選択装置は、 選択された音声データを互いに結合 することにより、 合成音声を表すデ一夕を生成する音声合成手段を 更に備えていてもよい。 また、 本発明の音声データ選択方法は、 基本的に、 音声の波形を 表す音声データを複数記憶し、 文章を表す文章情報を入力し、 各前 記音声データのうちから、 前記文章を構成する音片と読みが共通す る音片の波形を表している音声データを索出し、 および索出された 音声データのうちから、 前記文章を構成するそれぞれの音片に相当 する音声データを 1個ずつ、 互いに隣接する音片同士の境界でのピ ツチの差を前記文章全体で累計した値が最小となるように選択する. という一連の処理ステップを含む。
また、 この発明のコンピュータプログラムは、 コンピュ一夕を、 音声の波形を表す音声データを複数記憶する記憶手段と、 文章を表 す文章情報を入力し、 各前記音声データのうちから、 前記文章を構 成する音片と読みが共通する音片の波形を表している音声デ一夕を 索出する検索手段と、 索出された音声データのうちから、 前記文章 を構成するそれぞれの音片に相当する音声データを 1個ずつ、 互い に隣接する音片同士の境界でのピツチの差を前記文章全体で累計し た値が最小となるように選択する選択手段と、 して機能させるため のものとなっている。
( 2 )本発明の第 2の局面においては、音声選択装置は、基本的に、 音声の波形を表す音声データを複数記憶する記憶手段と、 文章を表 す文章情報を入力し、 当該文章を構成する音片について韻律予測を 行うことにより、 当該音片のピッチの時間変化を予測する予測手段 と、 各前記音声データのうちから、 前記文章を構成する音片と読み が共通する音片の波形を表していて、 且つ、 ピッチの時間変化が前 記予測手段による予測の結果と最も高い相関を示す音声データを選 択する選択手段と、 から構成される。 · 前記選択手段は、 音声データが表す音片のピツチの時間変化と、 当該音片と読みが共通する前記文章内の音片のピッチの時間変化と の間での 1次回帰を行う回帰計算の結果に基づいて、 当該音声デー 夕のピッチの時間変化と前記予測手段による予測の結果との相関の 強さを特定するものであってもよい。
前記選択手段は、 音声データが表す音片のピツチの時間変化と、 当該音片と読みが共通する前記文章内の音片のピッチの時間変化と の間の相関係数に基づいて、 当該音声デ一夕のピッチの時間変化と 前記予測手段による予測の結果どの相関の強さを特定するものであ つてもよい。
また、 この発明の別の音声選択装置は、 音声の波形を表す音声デ —タを複数記憶する記憶手段と、 文章を表す文章情報を入力し、 当 該文章内の音片について韻律予測を行うことにより、 当該音片の時 間長、 及び、 当該音片のピッチの時間変化を予測する予測手段と、 前記文章内の音片と読みが共通する音片の波形を表す各々の音声デ —夕についての評価値を特定し、 評価値が最も高い評価を表してい る音声データを選択する選択手段と、 より構成されており、 前記評 価値は、 音声データが表す音片のピッチの時間変化と、 当該音片と 読みが共通する前記文章内の音片のピッチの時間変化の予測結果と の相関を表す数値の関数、 及び、 当該音声データが表す音片の時間 長と、 当該音片と読みが共通する前記文章内の音片の時間長の予測 結果との差の関数より得られるようになつている。
前記相関を表す数値は、 音声データが表す音片のピッチの時間変 化と、 当該音片と読みが共通する前記文章内の音片のピッチの時間 変化との間での 1次回帰により得られる 1次関数の勾配からなって いてもよい。
また、 前記相関を表す数値は、 音声データが表す音片のピッチの 時間変化と、 当該音片と読みが共通する前記文章内の音片のピツチ の時間変化との間での 1次回帰により得られる 1次関数の切片から なっていてもよい。
前記相関を表す数値は、 音声データが表す音片のピツチの時間変 化と、 当該音片と読みが共通する前記文章内の音片のピツチの時間 変化の予測結果との間の相関係数からなっていてもよい。
前記相関を表す数値は、 音声データが表す音片のピツチの時間変 化を表すデータを種々のビッ ト数循環シフトしたものが表す関数と. 当該音片と読みが共通する前記文章内の音片のピッチの時間変化の 予測結果を表す関数との相関係数の最大値からなっていてもよい。 前記記憶手段は、 音声データの読みを表す表音データを、 当該音 声データに対応付けて記憶していてもよく、 また前記選択手段は、 前記文章内の音片の読みに合致する読みを表す表音データが対応付 けられている音声データを、 当該音片と読みが共通する音片の波形 を表す音声データとして扱うものであってもよい。
前記音声選択装置は、 選択された音声デ一夕を互いに結合するこ とにより、 合成音声を表すデータを生成する音声合成手段を更に備 えていてもよい。
前記音声選択装置は、 前記文章内の音片のうち、 前記選択手段が 音声データを選択できなかった音片について、 前記記憶手段が記憶 する音声データを用いることなく、 当該音片の波形を表す音声デー 夕を合成する欠落部分合成手段を備えていてもよく、 前記音声合成 手段は、 前記選択手段が選択した音声データ及び前記欠落部分合成 手段が合成した音声デ一夕を互いに結合することにより、 合成音声 を表すデ一夕を生成するものであってもよい。
また、 この発明の音声選択方法は、 音声の波形を表す音声データ を複数記憶し、 文章を表す文章情報を入力し、 当該文章を構成する 音片について韻律予測を行うことにより、 当該音片のピッチの時間 変化を予測し、 および各前記音声データのうちから、 前記文章を構 成する音片と読みが共通する音片の波形を表していて、 且つ、 ピッ チの時間変化が前記予測手段による予測の結果と最も高い相関を示 す音声デ一夕を選択する、 という一連の処理ステップを含む。
また、 この発明の別の音声選択方法は、 音声の波形を表す音声デ 一夕を複数記憶し、 文章を表す文章情報を入力し、 当該文章内の音 片について韻律予測を行うことにより、 当該音片の時間長、 及び、 当該音片のピッチの時間変化を予測し、 前記文章内の音片と読みが 共通する音片の波形を表す各々の音声データについての評価値を特 定し、 評価値が最も高い評価を表している音声データを選択するよ うになつており、 前記評価値は、 音声データが表す音片のピッチの 時間変化と、 当該音片と読みが共通する前記文章内の音片のピッチ の時間変化の予測結果との相関を表す数値の関数、 及び、 当該音声 データが表す音片の時間長と、 当該音片と読みが共通する前記文章 内の音片の時間長の予測結果との差の関数より得られるものである。 また、 この発明のコンピュータプログラムは、 コンピュータを、 音声の波形を表す音声データを複数記憶する記憶手段と、 文章を表 す文章情報を入力し、 当該文章を構成する音片について韻律予測を 行うことにより、 当該音片のピッチの時間変化を予測する予測手段 と、 各前記音声データのうちから、 前記文章を構成する音片と読み が共通する音片の波形を表していて、 且つ、 ピッチの時間変化が前 記予測手段による予測の結果と最も高い相関を示す音声データを選 択する選択手段と、 して機能させるためのものとなっている。
さらに、 この発明の別のコンピュータプログラムは、 コンビユー タを、 音声の波形を表す音声データを複数記憶する記憶手段と、 文 章を表す文章情報を入力し、 当該文章内の音片について韻律予測を 行うことにより、 当該音片の時間長、 及び、 当該音片のピッチの時 間変化を予測する予測手段と、 前記文章内の音片と読みが共通する 音片の波形を表す各々の音声データについての評価値を特定し、 評 価値が最も高い評価を表している音声データを選択する選択手段と して機能させるためのプログラムであって、 前記評価値は、 音声デ —夕が表す音片のピツチの時間変化と、 当該音片と読みが共通する 前記文章内の音片のピツチの時間変化の予測結果との相関を表す数 値の関数、 及び、 当該音声データが表す音片の時間長と、 当該音片 と読みが共通する前記文章内の音片の時間長の予測結果との差の関 数より得られるものである、 となっている。
( 3 ) 本発明の第 3の局面においては、 音声デ一夕選択装置は、 基 本的に、 音声の波形を表す音声データを複数記憶する記憶手段と、 文章を表す文章情報を入力する文章情報入力手段と、 前記文章情報 が表す文章内の音片と読みが共通する部分を有する音声データを索 出する検索部と、 前記索出されたそれぞれの音声データを文章情報 が表す文章に従って接続した際に互いに隣接する音声データ同士の 関係に基づいた所定の評価基準に従って評価値を求め、 出力する音 声データの組み合わせを当該評価値に基づいて選択する選択手段と, から構成される。 前記評価基準は、 音声データが表す音声と韻律予測結果との相関 及び互いに隣接する音声データ同士の関係を示す評価値を定める基 準であって、 前記評価値は、 前記音声データが表す音声の特徴を示 すパラメ一夕、 前記音声データが表す音声を互いに結合して得られ る音声の特徴を示すパラメータ、 及び、 発話時間長に関する特徴を 示すパラメータのうち、 少なくともいずれかを含む評価式に基づい て得られるものであるものであってもよい。
あるいは、 前記評価基準は、 音声データが表す音声と韻律予測結 果との相関及び互いに隣接する音声データ同士の関係を示す評価値 を定める基準であって、 前記評価値は、 前記音声デ一夕が表す音声 を互いに結合して得られる音声の特徴を示すパラメ一夕を含み、 ま た、 前記音声データが表す音声の特徴を示すパラメ一夕と発話時間 長に関する特徴を示すパラメ一夕のうち、 少なく ともいずれかを含 む評価式に基づいて得られるものであってもよい。
前記音声データが表す音声を互いに結合して得られる音声の特徴 を示すパラメータは、 前記文章情報が表す文章内の音片と読みが共 通する部分を有する音声の波形を表す音声データのうちから、 前記 文章を構成するそれぞれの音片に相当する音声デ一夕を 1個ずつ選 択した場合における、 互いに隣接する音声デ一夕同士の境界でのピ ツチの差に基づいて得られるものであってもよい。
前記音片データ選択装置は、 文章を表す文章情報を入力し、 当該 文章内の音片について韻律予測を行うことにより、 当該音片の時間 長、 及び、 当該音片のピッチの時間変化を予測する予測手段を備え ていてもよく、 前記評価基準は、 音声データが表す音声と前記韻律 予測手段の韻律予測結果との相関ないし差異を示す評価値を定める 基準であって、 前記評価値は、 音声データが表す音片のピッチの時 間変化と、 当該音片と読みが共通する前記文章内の音片のピツチの 時間変化の予測結果との相関を表す数値の関数、 及び/又は、 当該 音声データが表す音片の時間長と、 当該音片と読みが共通する前記 文章内の音片の時間長の予測結果との差の関数に基づいて得られる ものであってもよい。
前記相関を表す数値は、 音声データが表す音片のピツチの時間変 化と、 当該音片と読みが共通する前記文章内の音片のピツチの時間 変化との間での 1次回帰により得られる 1次関数の勾配及び/又は 切片からなっていてもよい。
前記相関を表す数値は、 音声データが表す音片のピッチの時間変 化と、 当該音片と読みが共通する前記文章内の音片のピツチの時間 変化の予測結果との間の相関係数からなっていてもよい。
あるいは、 前記相関を表す数値は、 音声データが表す音片のピッ チの時間変化を表すデータを種々のビッ ト数循環シフトしたものが 表す関数と、 当該音片と読みが共通する前記文章内の音片のピッチ の時間変化の予測結果を表す関数との相関係数の最大値からなって いてもよい。
前記記憶手段は、 音声データの読みを表す表音データを、 当該音 声データに対応付けて記憶していてもよく、 前記選択手段は、 前記 文章内の音片の読みに合致する読みを表す表音データが対応付けら れている音声データを、 当該音片と読みが共通する音片の波形を表 す音声データとして扱うものであってもよい。
前記音片データ選択装置は、 選択された音声データを互いに結合 することにより、 合成音声を表すデータを生成する'音声合成手段を '更に備えていてもよい。
前記音片デ一夕選択装置は、 前記文章内の音片のうち、 前記選択 手段が音声データを選択できなかった音片について、 前記記憶手段 が記憶する音声データを用いることなく、 当該音片の波形を表す音 声データを合成する欠落部分合成手段を備えていてもよく、 前記音 声合成手段は、 前記選択手段が選択した音声データ及び前記欠落部 分合成手段が合成した音声データを互いに結合することにより、 合 成音声を表すデータを生成するものであってもよい。
また、 この発明の音声データ選択方法は、 音声の波形を表す音声 データを複数記憶し、 文章を表す文章情報を入力し、 前記文章情報 が表す文章内の音片と読みが共通する部分を有する音声データを索 出し、 前記索出されたそれぞれの音声データを文章情報が表す文章 に従って接続した際に互いに隣接する音声データ同士の関係に基づ いた所定の評価基準に従って評価値を求め、 出力する音声データの 組み合わせを当該評価値に基づいて選択する、 一速の処理ステップ を含む。
また、 この発明のコンピュータプログラムは、 コンピュータを、 音声の波形を表す音声データを複数記憶する記憶手段と、 文章を表 す文章情報を入力する文章情報入力手段と、 前記文章情報が表す文 章内の音片と読みが共通する部分を有する音声データを索出する検 索部と、 前記'索出されたそれぞれの音声データを文章情報が表す文 章に従って接続した際に互いに隣接する音声データ同士の関係に基 づいた所定の評価基準に従って評価値を求め、 出力する音声データ の組み合わせを当該評価値に基づいて選択する選択手段と、 して機 能させるためのものとなっている。 , 図面の簡単な説明
第 1図は、 この発明の各実施の形態に係る音声合成システムの構 成を示すブロック図である。
第 2図は、 この発明の第 1の実施の形態における音片デ一夕べ一 スのデ一夕構造を模式的に示す図である。
第 3図の ( a ) は、 音片についてのピッチ成分の周波数の予測結 果と、 この音片と読みが合致する音片の波形を表す音片デ一夕のピ ツチ成分の周波数の時間変化とを 1次回帰させる処理を説明するた めのグラフであり、 同図 ( b ) は、 相関係数を求めるために用いる 予測結果デ一夕及びピッチ成分データの値の一例を示すグラフであ る。
第 4図は、 この発明の第 2の実施の形態における音片デ一夕べ一 スのデ一夕構造を模式的に示す図である。
第 5図の ( a ) は、 定型メッセージの読みを示す図であり、 同図 ( b ) は、 音片編集部に供給された音片データのリス トであり、 同 図 ( c ) は、 先行する音片の末尾におけるピッチ成分の周波数と後 続の音片の先頭におけるピッチ成分の周波数との差の絶対値を示す 図であり、 同図 ( d ) は、 音片編集部がどの音片デ一夕を選択する かを示す図である。
第 6図は、 この発明の各実施の形態に係る音声合成システムの機 能を行うパーソナルコンピュータがフリーテキス トデータを取得し た場合の処理を示すフローチヤ一トである。
第 7図は、 この発明の各実施の形態に係る音声合成システムの機 能を行うパーソナルコンピュータが配信文字列データを取得した場 合の処理を示すフローチャートである。 ' 第 8図は、 この発明の第 1の実施の形態に係る音声合成
の機能を行うパーソナルコンピュー夕が定型メッセージデータ及び 発声スピ一ドデ一夕を取得した場合の処理を示すフローチヤ一トで ある。
第 9図は、 この発明の第 2の実施の形態に係る音声合成システム の機能を行うパーソナルコンピュー夕が定型メッセージデータ及び 発声スピ一ドデータを取得した場合の処理を示すフローチヤ一トで ある。
第 1 0図は、 この発明の第 3の実施の形態に係る音声合成システ ムの機能を行うパーソナルコンピュータが定型メッセージデ一タ及 び発声スピードデータを取得した場合の処理を示すフローチヤ一ト である。
発明を実施するための最良の形態
以下、 この発明の実施の形態を、 音声合成システムを例とし、 図 面を参照して説明する。
(第 1の実施の形態) . 第 1図は、 この発明の第 1の実施の形態に係る音声合成システム の構成を示す図である。図示するように、この音声合成システムは、 本体ュニッ ト Mと、 音片登録ュニッ ト Rとにより構成されている。 本体ユニッ ト Mは、 言語処理部 1 と、 一般単語辞書 2と、 ユーザ 単語辞書 3 と、 音響処理部 4と、 検索部 5と、 伸長部 6 と、 波形デ 一夕ベース 7と、 音片編集部 8 と、 検索部 9と、 音片データベース 1 0 と、 話速変換部 1 1 とにより構成されている。
言語処理部 1、 音響処理部 4、 検索部 5、 伸長部 6、 音片編集部 8、 検索部 9及び話速変換部 1 1は、 いずれも、 C P U ( C e ntr al ί
Processing Unit) や D P (Digital Signal Processor) 等 のプロセッサゃ、 このプロセッサが実行するためのプログラムを記 憶するメモリなどより構成されており、 それぞれ後述する処理を行 Ό。
なお、 言語処理部 1、 音響処理部 4、 検索部 5、 伸長部 6、 音片 編集部 8、 検索部 9及び話速変換部 1 1の一部又は全部の機能を単 一のプロセッサが行うようにしてもよい。
一般単語辞書 2 は、 P R O M (Programmable Read Only Memory) やハードディスク装置等の不揮発性メモリより構成され ている。 一般単語辞書 2には、 表意文字 (例えば、 漢字など) を含 む単語等と、 この単語等の読みを表す表音文字 (例えば、 カナや発 音記号など) とが、 この音声合成システムの製造者等によって、 あ らかじめ互いに対応付けて記憶されている。
ユ ー ザ 単語辞 書 3 は 、 E E P R O M ( Electrically Erasable/Programmable Read Oniv Memo ry) やハードディ スク装置等のデータ書き換え可能な不揮発性メモリと、 この不揮発 性メモリへのデ一夕の書き込みを制御する制御回路とにより構成さ れている。なお、プロセッサがこの制御回路の機能を行ってもよく、 言語処理部 1、 音響処理部 4、 検索部 5、 伸長部 6、 音片編集部 8、 検索部 9及び話速変換部 1 1の一部又は全部の機能を行うプロセッ サがユーザ単語辞書 3の制御回路の機能を行うようにしてもよい。
ユーザ単語辞書 3は、 表意文字を含む単語等と、 この単語等の読 みを表す表音文字とを、 ユーザの操作に従って外部より取得し、 互 いに対応付けて記憶する。 ユーザ単語辞書 3には、 一般単語辞書 2 に記憶されていない単語等とその読みを表す表音文字とが格納され ていれば十分である。
波形データベース 7は、 P R〇Mやハードディスク装置等の不揮 発性メモリより構成されている。 波形データべ一ス 7には、 表音文 字と、 この表音文字が表す単位音声の波形を表す波形データをェン ト口ピー符号化して得られる圧縮波形データとが、 この音声合成シ ステムの製造者等によって、 あらかじめ互いに対応付けて記憶され ている。 単位音声は、 規則合成方式の手法で用いられる程度の短い 音声であり、 具体的には、 音素や、 V CV (Vowel-Consonant- Vowel) 音節などの単位で区切られる音声である。 なお、 ェントロ ピー符号化される前の波形データは、 例えば、 P C M (Pulse Code Modulation) 化されたデジタル形式のデータからなっていれぱょ い。
音片デ一夕ベース 1 0は、 P ROMやハードディスク装置等の不 揮発性メモリより構成されている。
音片データべ一ス 1 0には、 例えば、 第 2図に示すデ一夕構造を 有するデータが記憶されている。 すなわち、 図示するように、 音片 デ—夕べ—ス 1 0に格納されているデータは、 ヘッダ部 HD R、 ィ ンデックス部 I D X、 ディ レク トリ部 D I R及びデ一夕部 DATの 4種に分かれている。
なお、 音片データベース 1 0へのデータの格納は、 例えば、 この 音声合成システムの製造者によりあらかじめ行われ、 及び Z又は、 音片登録ュニッ ト Rが後述する動作を行うことにより行われる。 へッダ部 HD Rには、音片データベース 1 0を識別するデータや、 インデックス部 I D X、 ディ レク トリ部 D I R及びデータ部 D A T のデータ量、 データの形式、 著作権等の帰属などを示すデータが格 納される。
データ部 DATには、 音片の波形を表す音片データをェントロピ 一符号化して得られる圧縮音片デ一夕が格納されている。
なお、 音片とは、 音声のうち音素 1個以上を含む連続した 1区間 をいい、 通常は単語 1個分又は複数個分の区間からなる。
また、 エントロピー符号化される前の音片データは、 上述の圧縮 波形データの生成のためェント口ピ一符号化される前の波形デ一夕 と同じ形式のデータ (例えば、 P CMされたデジタル形式のデータ) からなっていればよい。
ディ レク トリ部 D I Rには、 個々の圧縮音声データについて、
(A) この圧縮音片デ一夕が表す音片の読みを示す表音文字を表 すデータ (音片読みデータ) 、
(B) この圧縮音片データが格納されている記憶位置の先頭のァ ドレスを表すデ一夕、
(C) この圧縮音片デ一夕のデータ長を表すデータ、
(D) この圧縮音片デ一夕が表す音片の発声スピード (再生した 場合の時間長) を表すデータ (スピード初期値データ) 、
(E) この音片のピッチ成分の周波数の時間変化を表すデータ(ピ ツチ成分 "5 "—タ) 、
が、 互いに対応付けられた形で格納されている。 (なお、 音片デ 一夕ベース 1 0の記憶領域にはアドレスが付されているものとす る。 )
なお、 第 2図は、 データ部 DATに含まれるデータとして、 読み が 「サイタマ」 である音片の波形を表す、 データ量 1 4 1 0 hバイ 卜の圧縮音片データが、 アドレス 0 0 1 A 3 6 A 6' hを先頭とする 論理的位置に格納されている場合を例示している。 (なお、 本明細 書及び図面において、末尾に" h "を付した数字は 1 6進数を表す。) また、 ピッチ成分データは、 例えば、 図示するように、 音片のピ ツチ成分の周波数をサンプリングして得られたサンプル Y ( i ) (サ ンプルの総数を nとして、 i は n以下の正の整数) を表すデ一夕で あるものとする。
なお、 上述の (A ) 〜 (E ) のデータの集合のうち少なくとも (A ) のデータ (すなわち音片読みデータ) は、 音片読みデータが表す表 音文字に基づいて決められた順位に従ってソ一トされた状態で (例 えば、 表音文字がカナであれば、 五十音順に従って、 アドレス降順 に並んだ状態で) 、 音片データベース 1 0の記憶領域に格納されて いる。
インデックス部 I D Xには、 ディ レク トリ部 D I Rのデータのお およその論理的位置を音片読みデータに基づいて特定するためのデ —夕が格納されている。 具体的には、 例えば、 音片読みデータが力 ナを表すものであるとして、 カナ文字と'、 先頭 1字がこのカナ文字 であるような音片読みデータがどのような範囲のアドレスにあるか を示すデータとが、 互いに対応付けて格納されている。
なお、 一般単語辞書 2、 ユーザ単語辞書 3、 波形データベース 7 及び音片データベース 1 0の一部又は全部の機能を単一の不揮発性 メモリが行うようにしてもよい。
音片データべ一ス 1 0へのデ一夕の格納は、 第 1図に示す音片登 録ユニッ ト Rにより行われる。 音片登録ュニッ ト Rは、 図示するよ うに、 収録音片データセッ ト記憶部 1 2と、 音片デ一夕べ一ス作成 部 1 3 と、 圧縮部 1 4とにより構成されている。 なお、 音片登録ュ ニッ ト Rは音片デ一夕べ一ス 1 0とは着脱可能に接続されていても よく、 この場合は、 音片データベース 1 0に新たにデータを書き込 むときを除いては、 音片登録ュニッ ト Rを本体ュニッ ト Mから切り 離した状態で本体ュニッ ト Mに後述の動作を行わせてよい。
収録音片データセッ ト記憶部 1 2は、 ハードディスク装置等のデ 一夕書き換え可能な不揮発性メモリより構成されている。
収録音片デ一夕セッ ト記憶部 1 2には、 音片の読みを表す表音文 字と、 この音片を人が実際に発声したものを集音して得た波形を表 す音片データとが、 この音声合成システムの製造者等によって、 あ らかじめ互いに対応付けて記憶されている。 なお、 この音片デ一夕 は、 例えば、 P C M化されたデジタル形式のデータからなっていれ ばよい。
音片デ一夕ベース作成部 1 3及び圧縮部 1 4は、 C P U等のプロ セッサゃ、 このプロセッサが実行するためのプロダラムを記憶する メモリなどより構成されており、 このプログラムに従って後述する 処理を行う。
なお、 音片データベース作成部 1 3及び圧縮部 1 4の一部又は全 部の機能を単一のプロセッサが行うようにしてもよく、 また、 言語 処理部 1、 音響処理部 4、 検索部 5、 伸長部 6、 音片編集部 8、 検 索部 9及び話速変換部 1 1の一部又は全部の機能を行うプロセッサ が音片データベース作成部 1 3や圧縮部 1 4の機能を更に行っても よい。 また、 音片データベース作成部 1 3や圧縮部 1 4の機能を行 うプロセッサが、 収録音片データセッ ト記憶部 1 2の制御回路の機 能を兼ねてもよい。
音片デ一夕ベース作成部 1 3は、 収録音片デ一夕セッ ト記愴部 1 2より、 互いに対応付けられている表音文字及び音片デ一夕を読み 出し、 この音片データが表す音声のピッチ成分の周波数の時間変化 と、 発声スピードとを特定する。
発声スピードの特定は、 例えば、 この音片デ一夕のサンプル数を 数えることにより特定すればよい。
一方、 ピッチ成分の周波数の時間変化は、 例えば、 この音片デー 夕にケプストラム解析を施すことにより特定すればよい。 具体的に は、 例えば、 音片データが表す波形を時間軸上で多数の小部分へと 区切り、 得られたそれぞれの小'部分の強度を、 元の値の対数 (対数 の底は任意) に実質的に等しい値へと変換し、 値が変換されたこの 小部分のスペク トル (すなわち、 ケプストラム) を、 高速フーリエ 変換の手法 (あるいは、 離散的変数をフーリエ変換した結果を表す データを生成する他の任意の手法) により求める。 そして、 このケ プストラムの極大値を与える周波数のうちの最小値を、 この小部分 におけるピッチ成分の周波数として特定する。
なお、 ピッチ成分の周波数の時間変化は、 例えば、 特開 2 0 0 3 - 1 0 8 1 7 2号公報に開示された手法に従って音片デ一夕をピッ チ波形デ一夕へと変換してから、 このピッチ波形データに基づいて 特定するようにすると良好な結果が期待できる。 具体的には、 音片 データをフィルタリングしてピッチ信号を抽出し、 抽出されたピッ チ信号に基づいて、 音片データが表す波形を単位ピッチ長の区間へ と区切り、 各区間について、 ピッチ信号との相関関係に基づいて位 相のずれを特定して各区間の位相を揃えることにより、 音片データ をピッチ波形信号へと変換すればよい。 そして、 得られたピッチ波 形信号を音片データとして扱い、 ケプス トラム解析を行う等するこ とにより、 ピッチ成分の周波数の時間変化を特定すればよい。
一方、 音片データベース作成部 1 3は、 収録音片データセッ ト記 憶部 1 2より読み出した音片データを圧縮部 1 4に供給する。
圧縮部 1 4は、 音片データベース作成部 1 3より供給された音片 データをェントロピー符号化して圧縮音片データを作成し、 音片デ 一夕ベース作成部 1 3に返送する。
音片デ一夕の発声スピード及びピッチ成分の周波数の時間変化を 特定し、 この音片デ一夕がェントロピー符号化され圧縮音片デ一夕 となって圧縮部 1 4より返送されると、 音片デ一夕ベース作成部 1 3は、 この圧縮音片データを、 データ部 D A Tを構成するデータと して、 音片データベース 1 0の記憶領域に書き込む。
また、 音片デ一夕ベース作成部 1 3は、 書き込んだ圧縮音片デー 夕が表す音片の読みを示すものとして収録音片データセッ ト記憶部 1 2より読み出した表音文字を、 音片読みデータとして音片デ一夕 ベース 1 0の記憶領域に書き込む。
また、 書き込んだ圧縮音片データの、 音片データベース 1 0の記 憶領域内での先頭のァドレスを特定し、 このアドレスを上述の (B ) のデータとして音片データベース 1 0の記憶領域に書き込む。
また、 この圧縮音片データのデ一夕長を特定し、 特定したデ一夕 長を、 (C ) のデータとして音片データベース 1 0の記憶領域に書 き込む。
また、 この圧縮音片デ一夕が表す音片の発声スピード及ぴピッチ 成分の周波数の時間変化を特定した結果を示すデ一夕を生成し、 ス ピ一ド初期値データ及びピッチ成分データとして音片デ一夕ベース 1 0の記憶領域に書き込む。 次に、 この音声合成システムの動作を説明する。
まず、 言語処理部 1が、 この音声合成システムに音声を合成させ る対象としてュ一ザが用意した、 表意文字を含む文章 (フリーテキ ス ト) を記述したフリーテキス トデータを外部から取得したとして 説明する。
なお、 言語処理部 1がフリ一テキストデータを取得する手法は任 意であり、 例えば、 図示しないインターフェース回路を介して外部 の装置ゃネッ トヮ一クから取得してもよいし、 図示しない記録媒体 ドライブ装置にセッ トされた記録媒体 (例えば、 フロッピー (登録 商標) ディスクや C D _ R O Mなど) から、 この記録媒体ドライブ 装置を介して読み取ってもよい。 また、 言語処理部 1の機能を行つ ているプロセッサが、 自ら実行している他の処理で用いたテキスト データを、 フリーテキストデータとして、 言語処理部 1の処理へと 引き渡すようにしてもよい。
フリーテキス トデータを取得すると、 言語処理部 1は、 このフリ 一テキストに含まれるそれぞれの表意文字について、 その読みを表 す表音文字を、 一般単語辞書 2やユーザ単語辞書 3を検索すること により特定する。 そして、 この表意文字を、 特定した表音文字へと 置換する。 そして、 言語処理部 1は、 フリーテキスト内の表意文字 がすべて表音文字へと置換した結果得られる表音文字列を、 音響処 理部 4へと供給する。
音響処理部 4は、 言語処理部 1より表音文字列を供給されると、 この表音文字列に含まれるそれぞれの表音文字について、 当該表音 文字が表す単位音声の波形を検索するよう、 検索部 5に指示する。 検索部 5は、 この指示に応答して波形データべ ス 7を検索し、 表音文字列に含まれるそれぞれの表音文字が表す単位音声の波形を 表す圧縮波形データを索出する。 そして、 索出された圧縮波形デー 夕を伸長部 6へと供給する。
伸長部 6は、 検索部 5より供給された圧縮波形デ一夕を、 圧縮さ れる前の波形データへと復元し、 検索部 5へと返送する。 検索部 5 は、'伸長部 6より返送された波形データを、 検索結果として音響処 理部 4へと供給する。
音響処理部 4は、 検索部 5より供給された波形デ一夕を、 言語処 理部 1より供給された表音文字列内での各表音文字の並びに従った 順序で、 音片編集部 8へと供給する。 - 音片編集部 8は、 音響処理部 4より波形デ一タを供給されると、 この波形データを、 供給された順序で互いに結合し、 合成音声を表 すデータ (合成音声データ) として出力する。 フリーテキストデー タに基づいて合成されたこの合成音声は、 規則合成方式の手法によ り合成された音声に相当する。
なお、 音片編集部 8が合成音声データを出力する手法は任意であ り、 例えば、 図示しない D / A ( D i git al - t o - An al o g ) 変換器ゃス ピー力を介して、 この合成音声デ一夕が表す合成音声を再生するよ うにしてもよい。 また、 図示しないインタ一フェース回路を介して 外部の装置ゃネッ トワークに送出してもよいし、 図示しない記録媒 体ドライブ装置にセッ トされた記録媒体へ、 この記録媒体ドライブ 装置を介して書き込んでもよい。 また、 音片編集部 8の機能を行つ ているプロセッサが、 自ら実行している他の処理へと、 合成音声デ 一夕を引き渡すようにしてもよい。
次に、 音響処理部 4が、 外部より配信された、 表音文字列を表す データ (配信文字列データ) を取得したとする。 (なお、 音響処理 部 4が配信文字列データを取得する手法も任意であり、 例えば、 言 語処理部 1がフリーテキストデータを取得する手法と同様の手法で 配信文字列データを取得すればよい。 )
この場合、 音響処理部 4は、 配信文字列データが表す表音文字列 を、 言語処理部 1より供給された表音文字列と同様に扱う。 この結 果、 配信文字列データが表す表音文字列に含まれる表音文字に対応 する圧縮波形データが検索部 5により索出され、 圧縮される前の波 形デ一夕が伸長部 6により復元される。 復元された各波形データは 音響処理部 4を介して音片編集部 8へと供給され、音片編集部 8が、 この波形データを、 配信文字列データが表す表音文字列内での各表 音文字の並びに従った順序で互いに結合し、 合成音声デ一夕として 出力する。 配信文字列データに基づいて合成されたこの合成音声デ —タも、 規則合成方式の手法により合成された音声を表す。
次に、 音片編集部 8が、 定型メッセージデータ及び発声スピード データを取得したとする。
なお、 定型メッセ一ジデータは、 定型メッセージを表音文字列と して表すデータであり、 発声スピードデータは、 定型メッセ一ジデ —夕が表す定型メッセージの発声スピードの指定値 (この定型メッ セージを発声する時間長の指定値) を示すデータである。
また、 音片編集部 8が定型メッセージデータや発声スピードデ一 夕を取得する手法は任意であり、 例えば、 言語処理部 1がフリーテ キストデータを取得する手法と同様の手法で定型メッセージデータ や発声スピードデータを取得すればよい。
定型メッセージデータ及び発声スピードデータが音片編集部 8に 供給されると、 音片編集部 8は、 定型メッセージに含まれる音片の 読みを表す表音文字に合致する表音文字が対応付けられている圧縮 音片データをすベて索出するよう、 検索部 9に指示する。
検索部 9は、 音片編集部 8の指示に応答して音片データベース 1 0を検索し、 該当する圧縮音片データと、 該当する圧縮音片データ に対応付けられている上述の音片読みデータ、 スピード初期値デー 夕及びピッチ成分データとを索出し、 索出された圧縮波形データを 伸長部 6へと供給する。 1個の音片にっき複数の圧縮音片データが 該当する場合も、 該当する圧縮音片データすべてが、 音声合成に用 いられるデータの候補として索出される。 一方、 圧縮音片データを 索出できなかった音片があった場合、 検索部 9は、 該当する音片を 識別するデータ (以下、 欠落部分識別データと呼ぶ) を生成する。 伸長部 6は、 検索部 9より供給された圧縮音片デ一夕を、 圧縮さ れる前の音片データへと復元し、 検索部 9へと返送する。 検索部 9 は、 伸長部 6より返送された音片デ一夕と、 索出された音片読みデ 一夕、 スピード初期値データ及びピッチ成分データとを、 検索結果 として話速変換部 1 1へと供給する。 また、 欠落部分識別データを 生成した場合は、 この欠落部分識別データも話速変換部 1 1へと供 給する。
一方、 音片編集部 8は、 話速変換部 1 1に対し、 話速変換部 1 1 に供給された音片デ一夕を変換して、 当該音片データが表す音片の 時間長を、 発声スピ一ドデータが示すスピードに合致するようにす ることを指示する。
話速変換部 1 1は、 音片編集部 8の指示に応答し、 検索部 9より 供給された音片データを指示に合致するように変換して、 音片編集 部 8に供給する。 具体的には、 例えば、 検索部 9より供給された音 片データの元の時間長を、 索出されたスピード初期値データに基づ いて特定した上、 この音片デ一タをリサンプリングして、 この音片 データのサンプル数を、 音片編集部 8の指示したスピードに合致す る時間長にすればよい。
また、 話速変換部 1 1は、 検索部 9より供給された音片読みデー タ、 スピ一ド初期値データ及びピッチ成分データも音片編集部 8に 供給し、 欠落部分識別データを検索部 9より供給された場合は、 更 にこの欠落部分識別データも音片編集部 8に供給する。
なお、 発声スピードデ一夕が音片編集部 8に供給されていない場 合、 音片編集部 8は、 話速変換部 1 1 に対し、 話速変換部 1 1に供 給された音片デ一夕を変換せずに音片編集部 8に供給するよう指示 すればよく、 話速変換部 1 1は、 この指示に応答し、 検索部 9より 供給された'音片データをそのまま音片編集部 8に供給すればよい。 音片編集部 8は、 話速変換部 1 1より音片デ一夕、 音片読みデー 夕、 スピード初期値データ及びピッチ成分データを供給されると、 供給された音片デ一夕のうちから、 定型メッセージを構成する音片 の波形に最もよく近似できる波形を表す音片デ一夕を、 音片 1個に つき 1個ずつ選択する。
具体的には、 まず、 音片編集部 8は、 定型メッセージデータが表 す定型メッセージに、例えば「藤崎モデル」や「T o B I ( To n e a n d B r e a k I n dic e s ) 」 等の韻律予測の手法に基づいた解析を加えるこ とにより、 この定型メッセージ内の各音片のピッチ成分の周波数の 時間変化を予測する。 そして、 音片毎に、 ピッチ成分の周波数の時 間変化の予測結果をサンプリングしたものを表すデジタル形式のデ 一夕 (以下、 予測結果データと呼ぶ) を生成する。
次に、 音片編集部 8は、 定型メッセージ内のそれぞれの音片につ いて、 この音片のピッチ成分の周波数の時間変化の予測結果を表す 予測結果データと、 この音片と読みが合致する音片の波形を表す音 片データのピッチ成分の周波数の時間変化を表すピッチ成分データ との相関を求める。
より具体的には、 音片編集部 8は、 話速変換部 1 1より供給され た各々のピッチ成分データについて、 例えば、 数式 1の右辺に示す 値 及び数式 2の右辺に示す値 ]3を求める。 n '
[は ( i ) -mx}' · {Y ( i ) - my} ]
Figure imgf000027_0001
た 7こし、 m x =
Figure imgf000027_0002
i = 1 , i = 1 β =m y— ( a · m x )
第 3図 ( a ) に示すように、 ある音片についての予測結果データ (サンプルの総数は n個とする) の i番目のサンプルの値 X ( i ) ( i は整数) の 1次関数として、 この音片と読みが合致する音片の 波形を表す音片データについてのピツチ成分データ (サンプルの総 数は n個とする) の i番目のサンプル Y ( i ) の値を 1次回帰させ た場合、 この 1次関数の勾配は 、 切片は^となる。 (勾配ひ の単 位は例えば [ヘルツ/秒] であればよく、 切片 /3の単位は例えば [へ ルツ] であればよい。 )
なお、 同一の読みの音片について、 予測結果データとピッチ成分 データとでサンプルの総数が互いに異なる場合は、 両者のうち一方 (または両方) を、 1次補間やラグランジェ補間あるいはその他任 意の手法により補間した上でリサンプリングし、 両者のサンプルの 総数を揃えてから相関を求めるようにすればよい。
一方、 音片編集部 8は、 話速変換部 1 1より供給されたスピード 初期値データと、 音片編集部 8に供給された定型メッセージデ一夕 及び発声スピ一ドデ一夕とを用いて、 数式 3の右辺の値 d t を求め る。 この値 d tは、 音片デ一夕が表す音片の発声スピードと、 この 音片と読みが合致する定型メッセージ内の音片の発声スピードとの 時間差を表す係数である。
d t = I ( X t - Y t ) / Ύ t I
(ただし、 Y tは音片データが表す音片の発声スピード、 X tはこ の音片と読みが合致する定型メッセージ内の音片の発声スピード) そして、 音片編集部 8は、 1次回帰により得られた上述の α及び /3の値と、 上述の係数 d t とに基づいて、 定型メッセージ内の音片 の読みと一致する音片を表す音片デ一夕のうち、 数式 4の右辺の値 (評価値) c 0 s t 1が最大となるものを選択する。
c o s t 1 = 1 / ( Ψ 1 I 1 - α I + W 2 I β l + d t ) (ただし、 及び W 2は所定の正の係数)
音片のピッチ成分の周波数の時間変化の予測結果と、 この音片と 読みが合致する音片の波形を表す音片データのピヅチ成分の周波数 の時間変化とが互いに近いほど、 勾配 の値は 1に近くなり、 従つ て、 値 I 1 — ひ I は 0に近くなる。 そして、 評価値 c o s t 1は、 音片のピッチの予測結果と音片デ一夕のピッチとの相関が高いほど 大きな値となるようにするため、 値 i 1 一 α I の 1次関数の逆数の 形をとつているので、 評価値 c o s t 1は、 値 I 1 一 α I が 0に近 くなるほど大きな値となる。
一方、 音声の抑揚は、 音片のピッチ成分の周波数の時間変化によ り特徴付けられる。 従って、 勾配ひ の値は、 音声の抑揚の差異を敏 感に反映する性質を有する。
このため、 合成されるべき音声について抑揚の正確さが重視され る場合 (例えば、 電子メール等のテキストを読み上げる音声を合成 する場合等) は、 上述の係数 の値をなるベく大きくすることが 望ましい。
これに対し、 音片のピッチ成分の基本周波数 (ベースピッチ周波 数) の予測結果と、 この音片と読みが合致する音片の波形を表す音 片データのベースピッチ周波数とが互いに近いほど、 切片 i3の値は 0に近くなる。 従って、 切片 ]3の値は、 音声のベースピッチ周波数 の差異を敏感に反映する性質を有する。一方、評価値 c o s t 1は、 値 I iS I の 1次関数の逆数とみることもできる形をとつているので、 評価値 c o s t 1は、値 I β I が 0に近くなるほど大きな値となる。
一方、 音声のベースピッチ周波数は、 音声の話者の声質を支配す る要因であり、 話者の性別による差異も顕著である。
このため、 合成されるべき音声についてベースピッチ周波数の正 確さが重視される場合 (例えば、 合成音声の話者の性別や声質を明 確にする必要がある場合など) は、 上述の係数 W 2の値をなるベく 大きくすることが望ましい。
動作の説明に戻ると、 音片編集部 8は、 定型メッセ一ジ内の音片 の波形に近い波形を表す音片デ一夕を選択する一方で、 話速変換部 1 1より欠落部分識別データも供給されている場合には、 欠落部分 識別データが示す音片の読みを表す表音文字列を定型メッセージデ 一夕より抽出して音響処理部 4に供給し、 この音片の波形を合成す るよう指示する。
指示を受けた音響処理部 4は、 音片編集部 8より供給された表音 文字列を、 配信文字列データが表す表音文字列と同様に扱う。 この 結果、 この表音文字列に含まれる表音文字が示す音声の波形を表す 圧縮波形データが検索部 5により索出され、 この圧縮波形データが 伸長部 6により元の波形デ一夕へと復元され、 検索部 5を介して音 響処理部 4へと供給される。 音響処理部 4は、 この波形データを音 片編集部 8へと供給する。
音片編集部 8は、 音響処理部 4より波形デ一夕を返送されると、 この波形データと、 話速変換部 1 1 より供給された音片デ一夕のう ち音片編集部 8が特定したものとを、 定型メッセージデータが示す 定型メッセージ内での各音片の並びに従った順序で互いに結合し、 合成音声を表すデータとして出力する。
なお、 話速変換部 1 1より供給されたデータに欠落部分識別デー 夕が含まれていない場合は、 音響処理部 4に波形の合成を指示する ことなく直ちに、 音片編集部 8が特定した音片データを、 定型メッ セージデータが示す定型メッセージ内での各音片の並びに従った順 序で互いに結合し、 合成音声を表すデータとして出力すればよい。 以上説明した、 この音声合成システムでは、 音素より大きな単位 であり得る音片の波形を表す音片データが、 韻律の予測結果に基づ いて、 録音編集方式により自然につなぎ合わせられ、 定型メッセ一 ジを読み上げる音声が合成される。 音片データベース 1 0の記憶容 量は、 音素毎に波形を記憶する場合に比べて小さくでき、 また、 高 速に検索できる。 このため、 この音声合成システムは小型軽量に構 成することができ、 また高速な処理にも追随できる。
また、 音片の波形の予測結果と音片デ一夕との相関を複数の評価 基準 (例えば、 1次回帰させた場合の勾配や切片による評価と、 音 片の時間差による評価、 など) で評価した場合は、 これらの評価の 結果に食い違いが生じる場合が多々あり得る。 しかし、 この音声合 成システムでは、 複数の評価基準で評価した結果が 1個の評価値に 基づいて総合され、 適正な評価が行われる。
なお、 この音声合成システムの構成は上述のものに限られない。 例えば、 波形データゃ音片データは P C M形式のデータである必 要はなく、 デ一タ形式は任意である。
また、 波形データベース 7ゃ音片データべ一ス 1 0は波形データ ゃ音片データを必ずしもデータ圧縮された状態で記憶している必要 はない。 波形デ一夕ベース 7ゃ音片データべ一ス 1 0が波形データ ゃ音片データをデータ圧縮されていない状態で記憶している場合、 本体ュニッ ト Mは伸長部 6を備えている必要はない。
また、 音片データベース作成部 1 3は、 図示しない記録媒体ドラ イブ装置にセッ 卜された記録媒体から、 この記録媒体ドライブ装置 を介して、 音片デ一夕ベース 1 0に追加する新たな圧縮音片データ の材料となる音片デ一夕や表音文字列を読み取ってもよい。
また、 音片登録ユニッ ト Rは、 必ずしも収録音片データセッ ト記 憶部 1 2を備えている必要はない。
また、 音片編集部 8は、 特定の音片の韻律を表す韻律登録デ一夕 をあらかじめ記憶し、 定型メッセージにこの特定の音片が含まれて いる場合は、 この韻律登録デ一夕が表す韻律を、 韻律予測の結果と して扱うようにしてもよい。
また、 音片編集部 8は、 過去の韻律予測の結果を韻律登録データ として新たに記憶するようにしてもよい。
また、 音片編集部 8は、 上述の α及び ]3の値を求める代わりに、 話速変換部 1 1より供給された各々のピッチ成分データについて、 例えば、 数式 5の右辺に示す値 R X y ( j ) を、 j の値を 0以上 n 未満の各整数として、 合計 n個求め、 得られた R x y ( 0 ) から R y ( n - 1 ) までの n個の,相関係数のうちの最大値を特定するよ うにしてもよい。 [ {X (i) -mx} · {Y j ( i) -my} ]
Rxy ( j )
2 {X ( i ) -mx} 2 ( i ) I 2
R x y ( j ) は、 ある音片についての予測結果デ一夕 (サンプル 総数 n個。 なお、 数式 5における X ( i ) は数式 1 におけるものと 同一である) と、 この音片と読みが合致する音片の波形を表す音片 データについてのピッチ成分データ (サンプルの総数 n個) を一定 の方向へ j 個循環シフトして得られたサンプルの列 (なお、 数式 5 において Y j ( i ) は、 このサンプルの列の i番目のサンプルの値 である) との相関係数の値である。 ' なお、 第 3図 ( b ) は、 R x y ( 0 ) 及び R x y ( j ) の値を求 めるために用いる予測結果データ及びピッチ成分データの値の一例 を示すグラフである。 ただし、 Y ( p ) の値 (ただし、 pは 1以上 n以下の整数) は、 循環シフ トを行う前のピッチ成分データの p番 目のサンプルの値である。 従って、 例えば、 音片データのサンプル が時刻の早い順に並んでおり、 循環シフ トが下位方向 (つまり時刻 が遅い方) へと行われるものとすれば、 j く pの場合は Y j ( p ) = Y (p— j ) であり、 一方、 l≤ p≤ j の場合は Y j ( p ) = Y ( n - j + p ) である。
そして、 音片編集部 8は、 上述の R x y ( j ) の最大値と、 上述 の係数 d tとに基づいて、 定型メッセージ内の音片の読みと一致す る音片を表す音片デ一夕のうち、 数式 6の右辺の値 (評価値) c o s t 2が最大となるものを選択すればよい。
c o s t 2 = 1 / (W3 I Rm a x I + d t )
(ただし、 W3は所定の係数、 Rm a xは R x y ( 0 ) 〜R x y (n 一 1 ) のうちの最大値)
なお、 音片編集部 8は、 必ずしもピッチ成分データを種々循環シ フトしたものについて上述の相関係数を求める必要はなく、例えば、 R x y ( 0 ) の値をそのまま相関係数の最大値として扱うようにし てもよい。
また、 評価値 c o s t 1や c o s t 2は、 係数 d tの項を含まな くてもよく、 この場合、 音片編集部 8は、 係数 d tを求める必要が ない。
あるいは、 音片編集部 8は、 係数 d tの値をそのまま評価値とし て用いてもよく、 この場合、 音片編集部は、 勾配 άや、 切片 ]3や、 R x y ( j ) の値を求める必要がない。
また、 ピッチ成分データは音片データが表す音片のピツチ長の時 間変化を表すデータであってもよい。 この場合、 音片編集部 8は、 予測結果データとして、 音片のピッチ長の時間変化の予測結果を表 すデータを作成するものとし、 この音片と読みが合致する音片の波 形を表す音片デ一夕のピッチ長の時間変化を表すピッチ成分データ との相関を求めるようにすればよい。
また、音片データベース作成部 1 3は、マイクロフォン、増幅器、 サンプリング回路、 A / D (Analo g - to - D i git al) コンバータ及び P C Mエンコーダなどを備えていてもよい。 この場合、 音片デ一夕 ベース作成部 1 3は、 収録音片デ一夕セッ ト記憶部 1 2より音片デ —夕を取得する代わりに、 自己のマイクロフォンが集音した音声を 表す音声信号を増幅し、 サンプリングして A / D変換した後、 サン プリングされた音声信号に P C M変調を施すことにより、 音片デー 夕を作成してもよい。
また、 音片編集部 8は、 音響処理部 4より返送された波形データ を話速変換部 1 1に供給することにより、 当該波形データが表す波 形の時間長を、 発声スピードデータが示すスピ一ドに合致させるよ うにしてもよい。
また、 音片編集部 8は、 例えば、 言語処理部 1 と共にフリーテキ ストデータを取得し、 このフリ一テキストデータが表すフリーテキ ストに含まれる音片の波形に最も近い波形を表す音片デ一夕を、 定 型メッセージに含まれる音片の波形に最も近い波形を表す音片デー 夕を選択する処理と実質的に同一の処理を行うことによって選択し て、 音声の合成に用いてもよい。 ' この場合、 音響処理部 4は、 音片編集部 8が選択した音片デ一夕 が表す音片については、 この音片の波形を表す波形データを検索部 5に索出させなくてもよい。 なお、 音片編集部 8は、 音響処理部 4 が合成しなくてよい音片を音響処理部 4に通知し、 音響処理部 4は この通知に応答して、 この音片を構成する単位音声の波形の検索を 中止するようにすればよい。
また、 音片編集部 8は、 例えば、 音響処理部 4と共に配信文字列 データを取得し、 この配信文字列データが表す配信文字列に含まれ る音片の波形に最も近い波形を表す音片データを、 定型メッセージ に含まれる音片の波形に最も近い波形を表す音片デ一夕を選択する 処理と実質的に同一の処理を行うことによって選択して、 音声の合 成に用いてもよい。 この場合、 音響処理部 4は、 音片編集部 8が選 択した音片データが表す音片については、 この音片の波形を表す波 形データを検索部 5に索出させなくてもよい。
(第 2の実施の形態)
次に、 この発明の第 2の実施の形態を説明する。 この発明の第 2 の実施の形態に係る音声合成システムの物理的構成は、 上述した第 1の実施の形態における構成と実質的に同一である。
ただし、 第 2の実施の形態の音声合成システムにおける音片デー 夕ベース 1 0のディ レク トリ部 D I Rには、 例えば第 4図に示すよ うに、 個々の圧縮音声デ一夕について、 上述の (A ) 〜 (D ) のデ 一夕が互いに対応づけられた形で格納されているほか、 上述の (E ) のデータに代え、 ピッチ成分データとして、 (F ) この圧縮音片デ 一夕が表す音片の先頭と末尾におけるピツチ成分の周波数を表すデ —夕が、 これら (A ) 〜 (D ) のデータに対応付けられた形で格納 されている。
なお、 第 4図は、 第 2図と同様、 データ部 D A Tに含まれるデ一 タとして、 読みが 「サイタマ」 である音片の波形を表す、 データ量 1 4 1 0 hバイ トの圧縮音片デ一夕が、 アドレス 0 0 1 A 3 6 A 6 hを先頭とする論理的位置に格納されている場合を例示している。) また、 上述の (A ) 〜 (D ) 及び (F ) のデータの集合のうち少な く とも (A ) のデータは、 音片読みデータが表す表音文字に基づい て決められた順位に従ってソ一トされた状態で音片データベース 1 0の記憶領域に格納されているものとする。
そして、 音片登録ュニッ ト Rの音片データベース作成部 1 3は、 収録音片データセッ ト記憶部 1 2より、 互いに対応付けられている 表音文字及び音片デ一夕を読み出すと、 この音片データが表す音声 の発声スピードと、 先頭及び末尾でのピッチ成分の周波数とを特定 するものとする。
そして、 読み出した音片データを圧縮部 1 4に供給し、 圧縮音片 データの返送を受けると、 この圧縮音片データ、 収録音片データセ ッ 卜記憶部 1 2より読み出した表音文字、 この圧縮音片デ一夕の音 片データベース 1 0の記憶領域内での先頭のァドレス、 この圧縮音 片デ一夕のデータ長、 及び、 特定した発声スピードを示すスピード 初期値データを、 第 1の実施の形態の音片データベース作成部 1 3 と同様の動作を行うことにより音片データベース 1 0の記憶領域に 書き込み、 また、 音声の先頭及び末尾におけるピッチ成分の周波数 を特定した結果を示すデータを生成して、 ピツチ成分デ一夕として 音片データベース 1 0の記憶領域に書き込むものとする。
なお、 発声スピード及びピッチ成分の周波数の特定は、 例えば、 第 1の実施の形態の音片データベース作成部 1 3が行う手法と実質 的に同一の手法により行えばよい。
次に、 この音声合成システムの動作を説明する。
この音声合成システムの言語処理部 1がフリーテキス トデータを 外部から取得した場合、 及び、 音響処理部 4が配信文字列データを 取得した場合の動作は、 第 1の実施の形態の音声合成システムが行 う動作と実質的に同一である。 (なお、 言語処理部 1がフリーテキ ストデータを取得する手法や音響処理部 4が配信文字列データを取 得する手法はいずれも任意であり、 例えば、 いずれも第 1の実施の 形態における言語処理部 1や音響処理部 4が行う手法と同様の手法 によりフリーテキストデータあるいは配信文字列データを取得すれ ばよい。 )·
次に、 音片編集部 8が、 定型メッセージデータ及び発声スピード データを取得したとする。 (なお、 音片編集部 8が定型メッセージ データや発声スピードデータを取得する手法も任意であり、例えば、 第 1の実施の形態の音片編集部 8が行う手法と同様の手法で定型メ ッセージデータや発声スピードデータを取得すればよい。 )
定型メッセージデ一夕及び発声スピードデータが音片編集部 8に 供給されると、 音片編集部 8は、 第 1の実施の形態における音片編 集部 8 と同様に、 定型メッセージに含まれる音片の読みを表す表音 文字に合致する表音文字が対応付けられている圧縮音片デ一夕をす ベて索出するよう、 検索部 9に指示する。 また、 話速変換部 1 1に 対しても、 第 1の実施の形態における音片編集部 8 と同様に、 話速 変換部 1 1に供給される音片データを変換して、 当該音片データが 表す音片の時間長を、 発声スピードデータが示すスピードに合致す るようにすることを指示する。
すると、 検索部 9、 伸張部 6及び話速変換部 1 1が、 第 1の実施 の形態における検索部 9、 伸張部 6及び話速変換部 1 1の動作と実 質的に同一の動作を行い、 この結果、 話速変換部 1 1から音片編集 部 8へと、 音片データ、 音片読みデータ及びピッチ成分データが供 給される。 また、 欠落部分識別データが検索部 9より話速変換部 1 1へと供給された場合は、 更にこの欠落部分識別データも音片編集 部 8へと供給される。
音片編集部 8は、 話速変換部 ί 1より音片データ、 音片読みデー 夕及びピツチ成分データを供給されると、以下説明する手順に従い、 供給された音片データのうちから、 定型メッセ一ジを構成する音片 の波形とみなせる波形を表す音片データを、 音片 1個につき 1個ず つ選択する。
具体的には、 まず、 音片編集部 8は、 話速変換部 1 1より供給さ れたピッチ成分データに基づき、 話速変換部 1 1より供給された各 音片データの先頭及び末尾の各時点でのピッチ成分の周波数を特定 する。 そして、 話速変換部 1 1より供給された音片デ一夕のうちか ら、 定型メッセージ内で隣接する音片同士の境界でのピッチ成分の 周波数の差の絶対値を定型メッセージ全体で累計した値が最小にな る、 という条件を満たすように、 音片デ一夕を選択する。
音片デ一夕を選択する条件を、 第 5図 ( a ) 〜 ( d ) を参照して 説明する。 例えば、 第 5図 ( a ) に示すような、 「このさきみぎか —ぶです」 という読みの定型メッセ一ジを表す定型メッセ一ジデー 夕が音片編集部 8に供給されたものとし、この定型メッセージが「こ のさき」 、 「みぎか一ぶ」 及び 「です」 という 3個の音片からなる ものとする。 そして、 第 5図 (b) にリストを示すように、 音片デ —夕ベース 1 0が、 読みが 「このさき」 である圧縮音片データが 3 個 (第 5図 ( b) において 「A 1」 「A 2」 あるいは 「A 3」 とし て表したもの) 、 読みが 「みぎか一ぶ」 である圧縮音片デ一夕が 2 個 (第 5図 (b) において 「B 1」 あるいは 「B 2」 として表した もの) 、 読みが 「です」 である圧縮音片デ一夕が 3個 (第 5図 (b) において 「C 1」 「C 2」 'あるいは 「C 3」 として表したもの) 、 それぞれ索出され、 伸長され、 音片データとして音片編集部 8へと 供給されたとする。
一方、 読みが 「このさき」 である各音片データが表す各音片の末 尾におけるピッチ成分の周波数と読みが 「みぎか一ぶ」 である各音 片データが表す各音片の先頭におけるピツチ成分の周波数との差の 絶対値は第 5図 ( c ) に示す通りであったとする。 (第 5図 ( c ) は、 例えば、 音片デ一夕 A 1が表す音片の末尾におけるピッチ成分 の周波数と音片データ B 1が表す音片の先頭におけるピッチ成分の 周波数との差の絶対値は「 1 2 3」であることを示している。なお、 この絶対値の単位は、 例えば 「ヘルツ」 である。 )
また、 読みが 「みぎか一ぶ」 である各音片デ一夕が表す各音片の 末尾におけるピッチ成分の周波数と読みが 「です」 である各音片デ 一夕が表す各音片の先頭におけるピツチ成分の周波数との差の絶対 値は第 5図 ( c ) に示す通りであったとする。
この場合において、 「このさきみぎか一ぶです」 という定型メッ セージを読み上げる音声の波形を音片データを用いて生成した場合、 隣接する音片同士の境界でのピッチ成分の周波数の差の絶対値の累 計が最小になる組み合わせは、 A 3、 B 2及び C 2 という組み合わ せである。 従ってこの場合、 音片編集部 8は、 第 5図 (d) に示す ように、 音片デ一夕 A 3、 B 2及び C 2を選択する。
この条件を満たす音片デ一夕を選択するために、音片編集部 8は、 例えば、 定型メッセージ内で隣接する音片同士の境界でのピッチ成 分の周波数の差の絶対値を距離として定義し、 D P (Dynamic Programming) マツチングの手法により音片デ一夕を選ぶように すればよい。
一方、 音片編集部 8は、 話速変換部 1 1より欠落部分識別デ一夕 も供給されている場合には、 欠落部分識別データが示す音片の読み を表す表音文字列を定型メッセージデータより抽出して音響処理部 4に供給し、 この音片の波形を合成するよう指示する。
指示を受けた音響処理部 4は、 音片編集部 8より供給された表音 文字列を、 配信文字列データが表す表音文字列と同様に扱う。 この 結果、 この表音文字列に含まれる表音文字が示す音声の波形を表す 圧縮波形データが検索部 5によ,り索出され、 この圧縮波形データが 伸長部 6により元の波形データへと復元され、 検索部 5を介して音 響処理部 4へと供給される。 音響処理部 4は、 この波形データを音 片編集部 8へと供給する。
音片編集部 8は、 音響処理部 4より波形データを返送されると、 この波形データと、 話速変換部 1 1より供給された音片デ一夕のう ち音片編集部 8が選択したものとを、 定型メッセージデータが示す 定型メッセージ内での各音片の並びに従った順序で互いに結合し、 合成音声を表すデータとして出力する。
なお、 話速変換部 1 1より供給されたデータに欠落部分識別デー 夕が含まれていない場合は、 第 1の実施の形態と同様、 音響処理部 4に波形の合成を指示することなく直ちに、 音片編集部 8が選択し た音片データを、 定型メッセ一ジデ一夕が示す定型メッセージ内で の各音片の並びに従った順序で互いに結合し、 合成音声を表すデ一 夕として出力すればよい。
以上説明したように、 この第 2の実施の形態の音声合成システム では、 音片デ一夕同士の境界でのピツチ成分の周波数の不連続的な 変化の量の累計が定型メッセージ全体で最小となるように音片デ一 夕が選ばれ、 録音編集方式により自然につなぎ合わせられるため、 合成音声が自然なものとなる。 また、 この音声合成システムでは、 処理が複雑な韻律予測は行われないので、 簡単な構成で高速な処理 にも追随できる。
なお、 この第 2の実施の形態の音声合成システムの構成も、 上述 のものに限られない。
例えば、 ピッチ成分データは音片データが表す音.片の先頭及び末 尾でのピッチ長を表すデータであってもよい。 この場合、 音片編集 部 8は、 話速変換部 1 1より供給された各音片データの先頭及び末 尾でのピッチ長を話速変換部 1 1より供給されたピツチ成分データ に基づいて特定し、 定型メッセージ内で隣接する音片同士の境界で のピッチ長の差の絶対値を定型メッセージ全体で累計した値が最小 になる、という条件を満たすように、音片データを選択すればよい。 また、 音片編集部 8は、,例えば、 言語処理部 1 と共にフリーテキ ストデータを取得し、 このフリーテキストデータが表すフリーテキ ストに含まれる音片の波形とみなせる波形を表す音片データを、 定 型メッセージに含まれる音片の波形とみなせる波形を表す音片デー 夕を抽出する処理と実質的に同一の処理を行うことによって抽出し て、 音声の合成に用いてもよい。
この場合、 音響処理部 4は、 音片編集部 8が抽出した音片デ一夕 が表す音片については、 この音片の波形を表す波形データを検索部 5に索出させなくてもよい。 なお、 音片編集部 8は、 音響処理部 4 が合成しなくてよい音片を音響処理部 4に通知し、 音響処理部 4は この通知に応答して、 この音片を構成する単位音声の波形の検索を 中止するようにすればよい。
また、 音片編集部 8は、 例えば、 音響処理部 4と共に配信文字列 データを取得し、 この配信文字列デ一夕が表す配信文字列に含まれ る音片の波形とみなせる波形を表す音片データを、 定型メッセージ に含まれる音片の波形とみなせる波形を表す音片デ一夕を抽出する 処理と実質的に同一の処理を行うことによって抽出して、 音声の合 成に用いてもよい。 この場合、 音響処理部 4は、 音片編集部 8が抽 出した音片データが表す音片については、 この音片の波形を表す波 形データを検索部 5に索出させなくてもよい。
(第 3の実施の形態)
次に、 この発明の第 3の実施の形態を説明する。 この発明の第 3 の実施の形態に係る音声合成システムの物理的構成は、 上述した第 1の実施の形態における構成と実質的に同一である。
次に、 この音声合成システムの動作を説明する。
この音声合成システムの言語処理部 1がフリーテキストデ一夕を 外部から取得した場合、 及び、 音響処理部 4が配信文字列デ一夕を 取得した場合の動作は、 第 1又は第 2の実施の形態の音声合成シス テムが行う動作と実質的に同一である。 (なお、 言語処理部 1がフ リーテキス トデータを取得する手法や音響処理部 4が配信文字列デ —夕を取得する手法はいずれも任意であり、 例えば、 いずれも第 1 又は第 2の実施の形態における言語処理部 1や音響処理部 4が行う 手法と同様の手法によりフリーテキストデータあるいは配信文字列 データを取得すればよい。 )
次に、 音片編集部 8が、 定型メッセージデ一タ及び発声スピード データを取得したとする。 なお、 音片編集部 8が定型メッセージデ —夕や発声スピードデータを取得する手法も任意であり、 例えば、 第 1の実施の形態の音片編集部 8が行う手法と同様の手法で定型メ ッセージデータや発声スピードデータを取得すればよい。あるいは、 例えばこの音声合成システムが力一ナビゲーシヨンシステム等の車 両内システムの一部をなすものであって、 この車両内システムを構 成するの他の装置 (例えば、 音声認識を行い、 音声認識の結果得ら れた情報に基づいてエージェント処理を実行する装置など) が、 ュ 一ザ一に対して発話する内容や発話スピードを決定し、 決定結果を 表すデータを生成するものである場合、 この音声合成システムは、 生成されたこのデータを受信 (取得) し、 定型メッセージデ一夕及 び発声スピードデータとして扱うようにしてもよい。
定型メッセージデータ及び発声スピードデ一夕が音片編集部 8に 供給されると、 音片編集部 8は、 第 1の実施の形態における音片編 集部 8と同様に、 定型メッセージに含まれる音片の読みを表す表音 文字に合致する表音文字が対応付けられている圧縮音片データをす ベて索出するよう、 検索部 9に指示する。 また、 話速変換部 1 1に 対しても、 第 1の実施の形態における音片編集部 8 と崗様に、 話速 変換部 1 1に供給される音片データを変換して、 当該音片データが 表す音片の時間長を、 発声スピ一ドデータが示すスピードに合致す るようにすることを指示する。
すると、 検索部 9、 伸張部 6及び話速変換部 1 1が、 第 1の実施 の形態における検索部 9、 伸張部 6及び話速変換部 1 1の動作と実 質的に同一の動作を行い、 この結果、 話速変換部 1 1から音片編集 部 8へと、 音片データ、 音片読みデータ、 この音片データが表す音 片の発声スピードを表すスピ一ド初期値データ及びピッチ成分デー 夕が供給される。 また、 欠落部分識別データが検索部 9より話速変 換部 1 1へと供給された場合は、 更にこの欠落部分識別デ一タも音 片編集部 8へと供給される。
音片編集部 8は、 話速変換部 1 1より音片データ、 音片読みデー タ及びピッチ成分データを供給されると、 話速変換部 1 1より供給 された各々のピッチ成分データについて上述の値《、 βの組及び Ζ 又は Rma xを求め、 また、 このスピード初期値デ一夕と、 音片編 集部 8に供給された定型メッセージデータ及び発声スピードデータ とを用いて、 上述の値 d tを求める。
そして、 音片編集部 8は、 話速変換部 1 1より供給されたそれぞ れの音片データにつき、 自ら求めた当該音片データ (以下、 音片デ 一夕 Xと記す) についての α、 β、 Rma x及び d tの値と、 定型 メッセージ内で当該音片デ一夕が表す音片の後に隣接する音片を表 す音片データ (以下、 音片デ一夕 Yと記す) のピッチ成分の周波数 とに基づいて、 数式 7に示す評価値 Ηχγを特定する。
HXY= (WA - c o s t— A) + (WB - c o s t— B) + (Wc - c o s t― C )
(ただし、 WA、 WB及び Wcはいずれも所定の係数であり、 WAは 0ではないものとする) 数式 7の右辺に含まれる値 c o s t— Aは、 当該定型メッセ一ジ 内で互いに隣接する、 音片データ Xが表す音片と音片データ Yが表 す音片との境界でのピツチ成分の周波数の差の絶対値の逆数である, なお、 音片編集部 8は、 c o s t—Aの値を特定するため、 話速 変換部 1 1より供給されたピッチ成分データに基づき、 話速変換部 1 1より供給された各音片データの先頭及び末尾の各時点でのピッ チ成分の周波数を特定するようにすればよい。
また、 数式 7の右辺に含まれる値 c o s t— Bは、 音片デ一夕 X について数式 8に従って評価値 c o s t—Bを求めた場合の値であ る。
c o s t _B = 1 / (WB 1 I 1 - a I + WB 2 I β I + WB 3 · d t )
(ただし、 WB 1、 WB 2及び WB 3は所定の正の係数)
また、 数式 7の右辺に含まれる値 c o s t— Cは、 音片デ一夕 X について数式 9に従って評価値 c o s t— Cを求めた場合における 値である。
c o s t _C = 1 / (Wc 1 I Rm a x I +WC 2 - d t )
(ただし、 WC 1及び WC 2は所定の係数)
あるいは、 音片編集部 8は、 数式 7〜数式 9に代えて、 数式 1 0 及び数式 1 1に従って評価値 Ηχ γを特定するようにしてもよい。た だし、 数式 1 0に含まれる c o s t—B及び c 0 s t— Cについて は、 上述の係数 WB 3及び Wc 3の値はいずれも 0 とする。 また、 数 式 8及び数式 9における (WB 3 · d t ) 及び (WC 2 · d t ) の項 を備えなく ともよい。
Ηχγ= ( Λ - c o s t— A) + (WB - c o s f _B ) + (Wc - c o s t— C) + (WD - c o s t— D)
(ただし、 WDは 0でない所定の係数)
c o s t— D = 1 / (Wd ! · d t )
(ただし、 Wdェは 0でない所定の係数)
そして、 音片編集部 8は、 話速変換部 1 1より供給された各音片 データのうちから、 音片編集部 8に供給された定型メッセージデー 夕が表す定型メッセージを構成する音片 1個につき 1個ずつの音片 データを選ぶことにより得られる各組み合わせのうち、 組み合わせ に属する各音片データの評価値 Ηχγの総和が最大となるものを、定 型メッセージを読み上げる音声を合成するための最適な音片デ一夕 の組み合わせとして選択する。
つまり、 例えば第 5図に示すように、 定型メッセ一ジデ一夕が表 す定型メッセージが音片 A, B及び Cより構成され、 音片 Aを表す 音片データの候補として音片デ一夕 A 1,A 2及び A 3が索出され、 音片 Bを表す音片データの候補として音片データ B 1及び B 2が索 出され、 音片 Cを表す音片データの候補として音片データ C 1, C 2及び C 3が索出された場合、 音片デ一夕 Α 1·, Α 2及び A 3のう ちから 1個、 音片デ一夕 Β 1及び Β 2のうちから 1個、 音片データ C 1 , C 2及び C 3のうちから 1個、 計 3個選ぶことにより得られ る組み合わせ計 1 8通りのうち、 組み合わせに属する各音片データ' の評価値 Ηχγの総和が最大となるものを、定型メッセージを読み上 げる音声を合成するための最適な音片データの組み合わせとして選 択する。
ただし、 総和を求めるために用いられる評価値 Η χ γとしては、 組 み合わせ内での音片の接続関係を正しく反映したものが選ばれるも のとする。 つまり、 例えば組み合わせ内に、 音片 pを表す音片デー 夕 P及び音片 Qを表す音片デ一夕 Qが含まれており、 定型メッセー ジ内では音片 Pが音片 qに先行する形で互いに隣接するという場合 音片データ Pの評価値としては、 音片 pが音片 (1に先行する形で互 いに隣接する場合における評価値 H p Qが甩いられるものとする。
また、 定型メッセージの末尾の音片 (例えば、 第 5図を参照して 前述した例でいえば、 音片 C 1 , 。 2及びじ 3 ) については、 後続 する音片が存在しないため、 c o s t—Aの値を定めることができ ない。 このため、 これら末尾の音片を表す音片データの評価値 Ηχγ を算定するにあたって、 音片編集部 8は、 (WA · c o s t— A) の値を 0であるものとして扱い、 一方、 係数 WB, Wc及び WDの値 は、 それぞれ、 他の音片データの評価値 Ηχγを算定する場合とは異 なる所定の値であるものとして扱う。
なお、 音片編集部 8は、 数式 7あるいは数式 1 1を用いて、 音片 データ Xについて、 当該音片データ Xが表す音片の前に隣接する音 片データ Yとの関係を表す評価値を含むものとして評価値 HXYを 特定してもよい。 この場合は、 定型メッセージの先頭の音片につい て、 先行する音片が存在しないため、 c o s t— Αの値を定めるこ とができないこととなる。 このため、 これら先頭の音片を表す音片 データの評価値 Ηχγを算定するにあたって、 音片編集部 8は、 (W Α · c o s t— Α)の値を 0であるものとして扱い、一方、係数 WB, Wc及び WDの値は、 それぞれ、 他の音片データの評価値 Ηχγを算 定する場合とは異なる所定の値であるものとして扱うようにすれば よい。
一方、 音片編集部 8は、 話速変換部 1 1より欠落部分識別データ も供給されている場合には、 欠落部分識別デ一夕が示す音片の読み を表す表音文字列を定型メッセージデータより抽出して音響処理部 4に供給し、 この音片の波形を合成するよう指示する。
指示を受けた音響処理部 4は、 音片編集部 8より供給された表音 文字列を、 配信文字列データが表す表音文字列と同様に扱う。 この 結果、 この表音文字列に含まれる表音文字が示す音声の波形を表す 圧縮波形データが検索部 5により索出され、 この圧縮波形データが 伸長部 6により元の波形データへと復元され、 検索部 5を介して音 響処理部 4へと供給される。 音響処理部 4は、 この波形データを音 片編集部 8へと供給する。
音片編集部 8は、 音響処理部 4より波形データを返送されると、 この波形データと、 話速変換部 1 1より供給された音片データのう ち、評価値 Η χ γの総和が最大となる組み合わせとして音片編集部 8 が選択した組み合わせに属するものとを、 定型メッセージデ一夕が 示す定型メッセージ内での各音片の並びに従った順序で互いに結合 し、 合成音声を表すデータとして出力する。
なお、 話速変換部 1 1より供給されたデ一夕に欠落部分識別デー 夕が含まれていない場合は、 第 1の実施の形態と同様、 音響処理部 4に波形の合成を指示することなく直ちに、 音片編集部 8が選択し た音片デ一夕を、 定型メッセージデータが示す定型メッセージ内で の各音片の並びに従った順序で互いに結合し、 合成音声を表すデー 夕として出力すればよい。
以上説明したように、 この第 3の実施の形態の音声合成システム でも、 音片デ一夕が録音編集方式により自然につなぎ合わせられ、 定型メッセージを読み上げる音声が合成される。 音片デ一夕べ一ス 1 0の記憶容量は、 音素毎に波形を記憶する場合に比べて小さくで き、 また、 高速に検索できる。 このため、 この音声合成システムは 小型軽量に構成することができ、 また高速な処理にも追随できる。 そして、 第 3の実施の形態の音声合成システムによれば、 定型メ ッセ一ジを読み上げる音声を合成するために選択される音片データ の組み合わせの適切さを評価するための様々な評価基準 (例えば、 音片の波形の予測結果と音片デ一夕との相関を 1次回帰させた場合 の勾配や切片による評価や、 音片の時間差による評価や、 音片デー 夕同士の境界でのピッチ成分の周波数の不連続的な変化の量の累計, など) が、 1個の評価値に影響を及ぼす形で総合的に反映され、 こ の結果、 最も自然な合成音声を合成するために選択すべき最適な音 片デ一夕の組み合わせが、 適正に決定される。
なお、 この第 3の実施の形態の音声合成システムの構成も、 上述 のものに限られない。
例えば、 最適な音片データの組み合わせを選択するために音片編 集部 8が用いる評価値は数式 7〜 1 3に示すものに限られず、 音片 データが表す音片を互いに結合して得られる音声が、 人の発する音 声にどの程度類似又は相違しているかについての評価を表す任意の 値であってよい。
また、 評価値を表す数式 (評価式) に含まれる変数ないし定数も 必ずしも数式?〜 1 3に含まれているものに限られず、 評価式とし ては、 音片データが表す音片の特徴を示す任意のパラメ一夕や、 あ るいは当該音片を互いに結合して得られる音声の特徴を示す任意の パラメ一夕や、 あるいは当該音声を人が発した場合に当該音声に備 わると予測される特徴を示す任意のパラメ一夕を含んだ数式が用い られてよい。
また、 最適な音片デ一夕の組み合わせを選択するための基準は必 ずしも評価値の形で表現可能なものである必要はなく、 音片データ が表す音片を互いに結合して得られる音声が人の発する音声にどの 程度類似又は相違しているかについての評価に基づいて音片データ の最適な組み合わせを特定するに至るような基準である限り任意で ある。
また、 音片編集部 8は、 例えば、 言語処理部 1 と共にフリーテキ ストデータを取得し、 このフリーテキストデ一夕が表すフリーテキ ス 卜に含まれる音片の波形とみなせる波形を表す音片データを、 定 型メッセージに含まれる音片の波形とみなせる波形を表す音片デー 夕を抽出する処理と実質的に同一の処理を行うことによって抽出し て、 音声の合成に用いてもよい。 この場合、 音響処理部 4は、 音片 編集部 8が抽出した音片データが表す音片については、 この音片の 波形を表す波形データを検索部 5に索出させなくてもよい。 なお、 音片編集部 8は、 音響処理部 4が合成しなくてよい音片を音響処理 部 4に通知し、 音響処理部 4はこの通知に応答して、 この音片を構 成する単位音声の波形の検索を中止するようにすればよい。
また、 音片編集部 8は、 例えば、 音響処理部 4と共に配信文字列 データを取得し、 この配信文字列データが表す配信文字列に含まれ る音片の波形とみなせる波形を表す音片データを、 定型メッセージ に含まれる音片の波形とみなせる波形を表す音片データを抽出する 処理と実質的に同一の処理を行うことによって抽出して、 音声の合 成に用いてもよい。 この場合、 音響処理部 4は、 音片編集部 8が抽 出した音片データが表す音片については、 この音片の波形を表す波 形データを検索部 5に索出させなくてもよい。
以上、 この発明の実施の形態を説明したが、 この発明にかかる音 声データ選択装置は、 専用のシステムによらず、 通常のコンピュー 夕システムを用いて実現可能である。
例えば、 パーソナルコンピュータに上述の第 1の実施の形態にお ける言語処理部 1、 一般単語辞書 2、 ユーザ単語辞書 3、 音響処理 部 4、 検索部 5、 伸長部 6、 波形データベース 7、 音片編集部 8 、 検索部 9、 音片データベース 1 0及び話速変換部 1 1の動作を実行 させるためのプログラムを格納した媒体 (C D— R O M、 M〇、 フ ロッピー (登録商標) ディスク等) から該プログラムをインスト一 ルすることにより、 当該パーソナルコンピュータに、 上述の第 1の 実施の形態の本体ュニッ ト Mの機能を行わせることができる。
また、 パーソナルコンピュータに、 上述の第 1の実施の形態にお ける収録音片データセッ 卜記憶部 1 2、 音片デ一夕ベース作成部 1 3及び圧縮部 1 4の動作を実行させるためのプログ,ラムを格納した 媒体から該プログラムをインスト一ルすることにより、 当該パーソ ナルコンピュータに、 上述の第 1の実施の形態の音片登録ュニッ ト Rの機能を行わせることができる。
そして、 これらのプログラムを実行し、 第 1の実施の形態におけ る本体ュニッ ト Mゃ音片登録ュニッ ト Rとして機能するパーソナル コンピュータが、 第 1図の音声合成システムの動作に相当する処理 として、 第 6図〜第 8図に示す処理を行うものとする。
第 6図は、 このパーソナルコンピュータがフリーテキストデータ を取得した場合の処理を示すフローチヤ一トである。
第 7図は、 このパーソナルコンピュータが配信文字列データを取 得した場合の処理を示すフローチヤ一トである。
第 8図は、 このパーソナルコンピュータが定型メッセージデータ 及び発声スピードデ一夕を取得した場合の処理を示すフローチヤ一 卜である。
すなわち、 まず、 このパーソナルコンピュータが、 外部より、 上 述のフリーテキストデータを取得すると (第 6図、 ステップ S 1 0 1 ) 、 このフリーテキストデータが表すフリーテキストに含まれる それぞれの表意文字について、 その読みを表す表音文字を、 一般単 語辞書 2やユーザ単語辞書 3を検索することにより特定し、 この表 意文字を、 特定した表音文字へと置換する (ステップ S 1 0 2 ) 。 なお、 このパーソナルコンピュータがフリ一テキストデ一夕を取得 する手法は任意である。
そして、 このパーソナルコンピュータは、 フリーテキスト内の表 意文字をすベて表音文字へと置換した結果を表す表音文字列が得ら れると、 この表音文字列に含まれるそれぞれの表音文字について、 当該表音文字が表す単位音声の波形を波形データベース 7より検索 し、 表音文字列に含まれるそれぞれの表音文字が表す単位音声の波 形を表す圧縮波形データを索出する (ステップ S 1 0 3 ) 。
次に、 このパーソナルコンピュータは、 索出された圧縮波 デー 夕を、圧縮される前の波形デ一夕へと復元し(ステップ S 1 0 4 )、 復元された波形データを、 表音文字列内での各表音文字の並びに従 つた順序で互いに結合し、 合成音声データとして出力する (ステツ プ S 1 0 5 ) 。 なお、 このパーソナルコンピュータが合成音声デー 夕を出力する手法は任意である。
また、 このパーソナルコンピュータが、 外部より、 上述の配信文 字列データを任意の手法で取得すると(第 7図、ステップ S 2 0 1 )、 この配信文字列データが表す表音文字列に含まれるそれぞれの表音 文字について、 当該表音文字が表す単位音声の波形を波形データべ —ス 7より検索し、 表音文字列に含まれるそれぞれの表音文字が表 す単位音声の波形を表す圧縮波形データを索出する (ステップ S 2 0 2 ) 。
次に、 このパーソナルコンピュータは、 索出された圧縮波形デー 夕を、圧縮される前の波形デ一夕へと復元し(ステップ S 2 0 3 ) 、 復元された波形データを、 表音文字列内での各表音文字の並びに従 つた順序で互いに結合し、 合成音声データとしてステップ S 1 0 5 の処理と同様の処理により出力する (ステップ S 2 0 4 ) 。
一方、 このパーソナルコンピュータが、 外部より、 上述の定型メ ッセージデータ及び発声スピードデ一夕を任意の手法により取得す ると (第 8図、 ステップ S 3 0 1 ) 、 まず、 この定型メッセ一ジデ —夕が表す定型メッセージに含まれる音片の読みを表す表音文字に 合致する表音文字が対応付けられている圧縮音片データをすベて索 出する (ステップ S 3 0 2 ) 。
また、 ステップ S 3 0 2では、 該当する圧縮音片デ一夕に対応付 けられている上述の音片読みデータ、 スピ一ド初期値デ一夕及びピ ツチ成分データも索出する。 なお、 1個の音片にっき複数の圧縮音 片データが該当する場合は、 該当する圧縮音片デ一夕すベてを索出 する。 一方、 圧縮音片データを索出できなかった音片があった場合 は、 上述の欠落部分識別データを生成する。
次に、 このパーソナルコンピュータは、 索出された圧縮音片デー 夕を、圧縮される前の音片データへと復元する(ステップ S 3 0 3 )。 そして、 復元された音片データを、 上述の音片編集部 8が行う処理 と同様の処理により変換して、 当該音片データが表す音片の時間長 を、 発声スピードデータが示すスピードに合致させる (ステップ S 3 0 4 ) 。 なお、 発声スピードデータが供給されていない場合は、 復元された音片データを変換しなくてもよい。
次に、 このパーソナルコンピュータは、 音片の時間長が変換され た音片デ一夕のうちから、 定型メッセージを構成する音片の波形に 最も近い波形を表す音片データを、 上述の音片編集部 8が行う処理 と同様の処理を行うことにより、 音片 1個につき 1個ずつ選択する (ステップ S 3 0 5 〜 S 3 0 8 ) 。
すなわち、 このパーソナルコンピュータは、 定型メッセ一ジデー 夕が表す定型メッセージに韻律予測の手法に基づいた解析を加える ことにより、 この定型メッセージの韻律を予測する (ステップ S 3 0 5 ) 。 そして、 定型メッセ一ジ内のそれぞれの音片について、 こ の音片のピツチ成分の周波数の時間変化の予測結果と、 この音片と 読みが合致する音片の波形を表す音片デ一夕のピッチ成分の周波数 の時間変化を表すピッチ成分データとの相関を求める (ステップ S 3 0 6 ) 。 より具体的には、 索出された各々のピッチ成分データに ついて、 例えば、 上述した勾配 a及び切片 ]3の値を求める。
一方で、 このパーソナルコンピュータは、 索出されたスピード初 期値データと、 外部より取得した定型メッセージデータ及び発声ス ピードデータとを用いて、 上述の値 d t を求める (ステップ S 3 0
7 ) 。
そして、 このパーソナルコンピュータは、 ステップ S 3 0 6で求 めたひ、 βの値、 及び、 ステップ S 3 0 7で求めた' d tの値に基づ いて、 定型メッセージ内の音片の読みと一致する音片を表す音片デ —夕のうち、 上述の評価値 c ο s t 1が最大となるものを選択する (ステップ S 3 0 8 ) 。
なお、 このパーソナルコンピュータは、 ステップ S 3 0 6で、 上 述の α及び 3の値を求める代わりに、 上述の R X y ( j ) の最大値 を求めるようにしてもよい。 この場合は、 ステップ S 3 0 8で、 R X y ( j ) の最大値と、 ステップ S 3 0 7で求めた係数 d t とに基 づいて、 定型メッセージ内の音片の読みと一致する音片を表す音片 デ一夕のうち、 上述の評価値 c 0 s t 2が最大となるものを選択す ればよい。
一方、 このパーソナルコンピュータは、 欠落部分識別データを生 成した場合、 欠落部分識別データが示す音片の読みを表す表音文字 列を定型メッセージデータより抽出し、 この表音文字列につき、 音 素毎に、 配信文字列データが表す表音文字列と同様に扱って上述の ステップ S 2 0 2 〜 S 2 0 3の処理を行うことにより、 この表音文 字列内の各表音文字が示す音声の波形を表す波形データを復元する (ステップ S 3 0 9 ) 。 '
そして、このパーソナルコンピュ一夕は、復元した波形データと、 ステップ S 3 0 8で選択した音片データとを、 定型メッセ一ジデー 夕が示す定型メッセージ内での各音片の並びに従った順序で互いに 結合し、合成音声を表すデ一夕として出力する(ステップ S 3 1 0 )。
また、 パーソナルコンピュータに上述の第 2の実施の形態におけ る言語処理部 1、 一般単語辞書 2、 ユーザ単語辞書 3、 音響処理部 4、 検索部 5、 伸長部.6、 波形デ一夕ベース 7、 音片編集部 8、 検 索部 9、 音片デ一夕ベース 1 0及び話速変換部 1 1'の動作を実行さ せるためのプログラムを格納した媒体から該プログラムをインスト —ルすることにより、 当該パーソナルコンピュータに、 上述の第 2 の実施の形態における本体ュニッ ト Mの機能を行わせることができ る。
また、 パニソナルコンピュータに上述の第 2の実施の形態におけ る収録音片デ一夕セッ ト記憶部 1 2、 音片データベース作成部 1 3 及び圧縮部 1 4の動作を実行させるためのプログラムを格納した媒 体から該プログラムをインスト一ルすることにより、 当該パーソナ ルコンピュータに、 上述の第 2の実施の形態における音片登録ュニ ッ ト Rの機能を行わせることができる。
そして、 これらのプログラムを実行し、 第 2の実施の形態におけ る本体ュニッ ト Mゃ音片登録ュニッ ト Rとして機能するパーソナル コンピュータが、 第 1図の音声合成システムの動作に相当する処理 として、 第 6図及び第 7図に示す上述の処理を行い、 また、 第 9図 に示す処理を行うものとする。
第 9図は、 このパーソナルコンピュータが定型メッセ一ジデータ 及び発声スピードデ一夕を取得した場合の処理を示すフロ一チヤ一 トである。
すなわち、 このパーソナルコンピュータが、 外部より、 上述の定 型メッセージデータ及び発声スピードデータを任意の手法により取 得すると (第 9図、 ステップ S 4 0 1 ) 、 まず、 上述のステツプ S 3 0 2の処理と同様に、 この定型メッセージデータが表す定型メッ セージに含まれる音片の読みを表す表音文字に合致する表音文字が 対応付けられている圧縮音片データと、 該当する圧縮音片データに 対応付けられている上述の音片読みデータ、 スピ ド初期値デ一夕 及びピツチ成分データとを、すべて索出する(ステップ S 4 0 2 ) 。 なお、 ステップ S 4 0 2でも、 1個の音片にっき複数の圧縮音片デ —夕が該当する場合は該当する圧縮音片デ一夕すベてを索出し、 一 方で圧縮音片データを索出できなかった音片があった場合は、 上述 の欠落部分識別データを生成する。
次に、 このパーソナルコンピュータは、 索出された圧縮音片デー 夕を、圧縮される前の音片デ一夕へと復元し(ステップ S 4 0 3 )、 復元された音片デ一夕を、 上述の音片編集部 8が行う処理と同様の 処理により変換して、 当該音片デ一夕が表す音片の時間長を、 発声 スピードデータが示すスピードに合致させる(ステップ S 4 0 4 )。 なお、 発声スピードデータが供給されていない場合は、 復元された 音片データを変換しなくてもよい。
次に、 このパーソナルコンピュータは、 音片の時間長が変換され た音片データのうちから、 定型メッセージを構成する音片の波形と みなせる波形を表す音片デ一タを、 上述の第 2の実施の形態におけ る音片編集部 8が行う処理と同様の処理を行う ことにより、 音片 1 個につき 1個ずつ選択する (ステップ S 4 0 5 〜 S 4 0 6 ) 。
具体的には、 まず、 このパーソナルコンピュータは、 音片の時間 長が変換された各音片デ一夕の先頭及び末尾の各時点でのピツチ成 分の周波数を、索出されたピツチ成分デ一夕に基づいて特定する(ス テツプ S 4 0 5 ) 。 そして、 これらの音片データのうちから、 定型 メッセージ内で隣接する音片同士の境界でのピッチ成分の周波数の 差の絶対値を定型メッセージ全体で累計した値が最小になる、 とい う条件を満たすように、音片データを選択する(ステツプ S 4 0 6 )。 この条件を満たす音片データを選択するために、 このパーソナルコ ンピュー夕は、 例えば、 定型メッセージ内で隣接する音片同士の境 界でのピッチ成分の周波数の差の絶対値を距離として定義し、 D P マッチングの手法により音片デ一夕を選ぶようにすればよい。
一方、 このパーソナルコンピュータは、 欠落部分識別データを生 成した場合、 欠落部分識別データが示す音片の読みを表す表音文字 列を定型メッセ一ジデ一夕より抽出し、 この表音文字列につき、 音 素毎に、 配信文字列データが表す表音文字列と同様に扱って上述の ステップ S 2 0 2 〜 S 2 0 3の処理を行うことにより、 この表音文 字列内の各表音文字が示す音声の波形を表す波形データを復元する (ステップ S 4 0 7 ) 。
そして、このパーソナルコンピュータは、復元した波形データと、 ステップ S 4 0 6で選択した音片デ一夕とを、 定型メッセ一ジデ一 夕が示す定型メッセージ内での各音片の並びに従った順序で互いに 結合し、合成音声を表すデータとして出力する(ステップ S 4 0 8 )。
また、 パーソナルコンピュータに上述の第 3の実施の形態におけ る言語処理部 1、 一般単語辞書 2、 ユーザ単語辞書 3、 音響処理部 4、 検索部 5、 伸長部 6、 波形データべ一ス 7、 音片編集部 8、 検 索部 9、 音片データベース 1 0及び話速変換部 1 1 の動作を実行さ せるためのプログラムを格納した媒体から該プログラムをインスト ールすることにより、 当該パーソナルコンピュータに、 上述の第 3 の実施の形態における本体ュニッ ト Mの機能を行わせることができ る。
また、 パーソナルコンピュータに上述の第 3の実施の形態におけ る収録音片データセッ ト記憶部 1 2、 音片データベース作成部 1 3 及び圧縮部 1 4の動作を実行させるためのプログラムを格納した媒 体から該プログラムをインストールすることにより、 当該パ一ソナ ルコンピュータに、 上述の第 3の実施の形態における音片登録ュニ ッ ト Rの機能を行わせることができる。
そして、 これらのプログラムを実行し、 第 3の実施の形態におけ る本体ュニッ ト Mゃ音片登録ュニッ ト Rとして機能するパーソナル コンピュータが、 第 1図の音声合成システムの動作に相当する処理 として、 第 6図及び第 7図に示す上述の処理 ¾行い、 また、 第 1 0 図に示す処理を行うものとする。.
第 1 0図は、 このパーソナルコンピュ一夕が定型メッセージデ一 タ及び発声スピードデータを取得した場合の処理を示すフローチヤ 一卜である。
すなわち、 このパーソナルコンピュータが、 外部より、 上述の定 型メッセージデータ及ぴ発声スピードデ一夕を任意の手法により取 得すると (第 1 0図、 ステップ S 5 0 1 ) 、 まず、 上述のステップ S 3 0 2の処理と同様に、 この定型メッセージデータが表す定型メ ッセージに含まれる音片の読みを表す表音文字に合致する表音文字 が対応付けられている圧縮音片デ一ダと、 該当する圧縮音片データ に対応付けられている上述の音片読みデータ、 スピ一ド初期値デ一 夕及びピッチ成分データとを、すべて索出する(ステップ S 5 0 2 )。 なお、 ステップ S 5 0 2でも、 1個の音片にっき複数の圧縮音片デ 一夕が該当する場合は該当する圧縮音片デ一夕すベてを索出し、 一 方で圧縮音片デ一夕を索出できなかった音片があった場合は、 上述 の欠落部分識別データを生成する。
次に、 このパーソナルコンピュータは、 索出された圧縮音片デー 夕を、圧縮される前の音片データへと復元し(ステヅプ S 5 0 3 )、 復元された音片データを、 上述の音片編集部 8が行う処理と同様の 処理により変換して、 当該音片データが表す音片の時間長を、 発声 スピードデータが示すスピードに合致させる(ステップ S 5 0 4 )。 なお、 発声スピードデータが供給されていない場合は、 復元された 音片デ一夕を変換しなくてもよい。
次に、 このパーソナルコンピュータは、 音片の時間長が変換され た音片データのうちから、 定型メッセージを読み上げる音声を合成 するための最適な音片データの組み合わせを、 上述の第 3の実施の 形態における音片編集部 8が行う処理と同様の処理を行うことによ り選択する (ステップ S 5 0 5〜S 5 0 7 ) 。
すなわち、 まず、 このパーソナルコンピュータは、 ステップ S 5 0 2で索出された各々のピッチ成分データについて上述の値 、 β の組及び/又は Rm a Xを求め、 また、 このスピード初期値データ と、 ステップ S 5 0 1で取得した定型メッセージデータ及び発声ス ピ―ドデータとを用いて、 上述の値 d tを求める (ステップ S 5 0
5 ) 。
次に、 このパーソナルコンピュータは、 ステップ S 5 0 4で変換 されたそれぞれの音片デ一夕につき、ステツプ S 5 0 5で求めた a、 β、 Rm a x及び d tの値と、 定型メッセージ内で当該音片データ が表す音片の後に隣接する音片を表す音片データのピッチ成分の周 波数とに基づいて、 上述した評価値 Ηχγを特定する (ステップ S 5 0 6 ) o
そして、 このパーソナルンピュータは、 ステップ S 5 0 4で変換 された各音片データのうちから、 ステップ S 5 0 1で取得した定型 メッセージデータが表す定型メッセージを構成する音片 1偭にっき 1個ずつの音片デ一夕を選ぶことにより得られる各組み合わせのう ち、組み合わせに属する各音片データの評価値 Η χ γの総和が最大と なるものを、 定型メッセージを読み上げる音声を合成するための最 適な音片データの組み合わせとして選択する(ステップ S 5 0 7 )。 ただし、 総和を求めるために用いられる評価値 Η χ γとしては、 組み 合わせ内での音片の接続関係を正しく反映したものが選ばれるもの とする。
一方、 このパーソナルコンピュータは、 欠落部分識別デ一タを生 成した場合、 欠落部分識別データが示す音片の読みを表す表音文字 列を定型メッセ一ジデータより抽出し、 この表音文字列につき、 音 素毎に、 配信文字列データが表す表音文宇列と同様に扱って上述の ズテツプ S 2 0 2〜 S 2 0 3の処理を行うことにより、 この表音文 字列内の各表音文字が示す音声の波形を表す波形データを復元する
(ステップ S 5 0 8 ) 。
そして、このパーソナルコンピュータは、復元した波形データと、 ステップ S 5 0 7で選択した組み合わせに属する音片デ一夕とを、 定型メッセージデータが示す定型メッセージ内での各音片の並びに 従った順序で互いに結合し、 合成音声を表すデータとして出力する
(ステップ S 5 0 9 ) 。
なお、 パーソナルコンピュータに本体ュニッ ト Mゃ音片登録ュニ ッ ト Rの機能を行わせるプログラムは、 例えば、 通信回線の掲示板
( B B S ) にアップロードし、 これを通信回線を介して配信しても よく、また、これらのプログラムを表す信号により搬送波を変調し、 得られた変調波を伝送し、 この変調波を受信した装置が変調波を復 調してこれらのプログラムを復元するようにしても'よい。 そして、 これらのプログラムを起動し、 O Sの制御下に、 他のァ プリケーションプログラムと同様に実行することにより、 上述の処 理を実行することができる。
なお、 〇 Sが処理の一部を分担する場合、 あるいは、 O Sが本願 発明の 1つの構成要素の一部を構成するような場合には、 記録媒体 には、その部分を除いたプログラムを格納してもよい。この場合も、 この発明では、 その記録媒体には、 コンピュータが実行する各機能 又はステツプを実行するためのプログラムが格納されているものと する。
産業上利用可能性
本発明によれば、 簡単な構成で高速に自然な合成音声を得るため の音声選択装置、 音声選択方法及びプログラムが実現される。

Claims

請求の範囲
1 . 音声の波形を表す音声データを複数記憶する記憶手段と、 文章を表す文章情報を入力し、 各前記音声データのうちから、 前 記文章を構成する音片と読みが共通する音片の波形を表している音 声データを索出する検索手段と、
索出された音声データのうちから、 前記文章を構成するそれぞれ の音片に相当する音声データを 1個ずつ、 互いに隣接する音片同士 の境界でのピッチの差を前記文章全体で累計した値が最小となるよ うに選択する選択手段と、
より構成されることを特徴とする音声データ選択装置。
2 . 選択された音声データを互いに結合することにより、 合成音声 を表すデータを生成する音声合成手段を更に備える、
ことを特徴とする請求項 1に記載の音声データ選択装置。
3 . 音声の波形を表す音声データを複数記憶し、
文章を表す文章情報を入力し、 各前記音声データのうちから、 前 記文章を構成する音片と読みが共通する音片の波形を表している音 声データを索出し、
索出された音声データのうちから、 前記文章を構成するそれぞれ の音片に相当する音声デ一夕を 1個ずつ、 互いに隣接する音片同士 の境界でのピッチの差を前記文章全体で累計した値が最小となるよ うに選択する、
ことを特徴とする音声データ選択方法。
4 . コンピュータを、
音声の波形を表す音声データを複数記憶する記憶手段と、 文章を表す文章情報を入力し、 各前記音声データのうちから、 前 記文章を構成する音片と読みが共通する音片の波形を表している音 声データを索出する検索手段と、
索出された音声デ一夕のうちから、 前記文章を構成するそれぞれ の音片に相当する音声データを 1個ずつ、 互いに隣接する音片同士 の境界でのピツチの差を前記文章全体で累計した値が最小となるよ うに選択する選択手段と、
して機能させるためのプログラム。
5 . 音声の波形を表す音声データを複数記憶する記憶手段と、 文章を表す文章情報を入力し、 当該文章を構成する音片について 韻律予測を行うことにより、 当該音片のピッチの時間変化を予測す る予測手段と、
各前記音声データのうちから、 前記文章を構成する音片と読みが 共通する音片の波形を表していて、 且つ、 ピッチの時間変化が前記 予測手段による予測の結果と最も高い相関を示す音声データを選択 する選択手段と、
より構成されることを特徴とする音声選択装置。
6 . 前記選択手段は、 音声データが表す音片のピッチの時間変化 と、 当該音片と読みが共通する前記文章内の音片のピツチの時間変 化との間での 1次回帰を行う回帰計算の結果に基づいて、 当該音声 データのピッチの時間変化と前記予測手段による予測の結果との相 関の強さを特定するものである、
ことを特徴とする請求項 5に記載の音声選択装置。
7 . 前記選択手段は、 音声データが表す音片のピッチの時間変化 と、 当該音片と読みが共通する前記文章内の音片のピッチの時間変 化との間の相関係数に基づいて、 当該音声データのピツチの時間変 化と前記予測手段による予測の結果との相関の強さを特定するもの である、
ことを特徴とする請求項 5に記載の音声選択装置。
8 . 音声の波形を表す音声データを複数記憶する記憶手段と、 文章を表す文章情報を入力し、 当該文章内の音片について韻律予 測を行うことにより、 -当該音片の時間長、 及び、 当該音片のピッチ の時間変化を予測する予測手段と、
前記文章内の音片と読みが共通する音片の波形を表す各々の音声 デ一夕についての評価値を特定し、 評価値が最も高い評価を表して いる音声データを選択する選択手段と、 より構成されており、 前記評価値は、 音声データが表す音片のピッチの時間変化と、 当 該音片と読みが共通する前記文章内の音片のピツチの時間変化の予 測結果との相関を表す数値の関数、 及び、 当該音声データが表す音 片の時間長と、 当該音片と読みが共通する前記文章内の音片の時間 長の予測結果との差の関数より得られるものである、
ことを特徴とする音声選択装置。
9 . 前記相関を表す数値は、 音声データが表す音片のピッチの時 間変化と、 当該音片と読みが共通する前記文章内の音片のピツチの 時間変化との間での 1次回帰により得られる 1次関数の勾配からな る、
ことを特徴とする請求項 8に記載の音声選択装置。
1 0 . 前記相関を表す数値は、 音声データが表す音片のピッチの時 間変化と、 当該音片と読みが共通する前記文章内の音片のピッチの 時間変化との間での 1次回帰により得られる 1次関数の切片からな る、
ことを特徴とする請求項 8に記載の音声選択装置。
1 1 . 前記相関を表す数値は、 音声データが表す音片のピッチの時 間変化と、 当該音片と読みが共通する前記文章内の音片のピッチの 時間変化の予測結果との間の相関係数からなる、
ことを特徴とする請求項 8に記載の音声選択装置。
1 2 . 前記相関を表す数値は、 音声データが表す音片のピッチの時 間変化を表すデ一夕を種々のビッ ト数循環シフ トしたものが表す関 数と、 当該音片と読みが共通する前記文章内の音片のピツチの時間 変化の予測結果を表す関数との相関係数の最大値からなる、
ことを特徴とする請求項 8に記載の音声選択装置。
1 3 . 前記記憶手段は、 音声デ一夕の読みを表す表音デ一夕を、 当 該音声データに対応付けて記憶しており、
前記選択手段は、 前記文章内の音片の読みに合致する読みを表す 表音データが対応付けられている音声データを、 当該音片と読みが 共通する音片の波形を表す音声データとして扱う、 ' ことを特徴とする請求項 5乃至 1 2のいずれか 1項に記載の音声 選択装置。
1 4 . 選択された音声データを互いに結合することにより、 合成音 声を表すデータを生成する音声合成手段を更に備える、
ことを特徴とする請求項 5乃至 1 3のいずれか 1項に記載の音声 選択装置。
1 5 .. 前記文章内の音片のうち、 前記選択手段が音声データを選択 できなかった音片について、 前記記憶手段が記憶する音声データを V, 用いることなく、 当該音片の波形を表す音声データを合成する欠落 部分合成手段を備え、
前記音声合成手段は、 前記選択手段が選択した音声データ及び前 記欠落部分合成手段が合成した音声データを互いに結合することに より、 合成音声を表すデータを生成する、
ことを特徴とする請求項 1 4に記載の音声選択装置。
1 6 . 音声の波形を表す音声データを複数記憶し、
文章を表す文章情報を入力し、 当該文章を構成する音片について 韻律予測を行うことにより、当該音片のピツチの時間変化を予測し、 各前記音声データのうちから、 前記文章を構成する音片と読みが 共通する音片の波形を表していて、 且つ、 ピッチの時間変化が前記 予測手段による予測の結果と最も高い相関を示す音声デ一夕を選択 する、
ことを特徴とする音声選択方法。
1 7 . 音声の波形を表す音声データを複数記憶し、
文章を表す文章情報を入力し、 当該文章内の音片について韻律予 測を行うことにより、 当該音片の時間長、 及び、 当該音片のピッチ の時間変化を予測し、
前記文章内の音片と読みが共通する音片の波形を表す各々の音声 データについての評価値を特定し、 評価値が最も高い評価を表して いる音声データを選択する、 ことを特徴とする音声選択方法であつ て、
前記評価値は、 音声データが表す音片のピッチの時間変化と、 当 該音片と読みが共通する前記文章内の音片のピッチの時間変化の予 測結果との相関を表す数値の関数、 及び、 当該音声データが表す音 片の時間長と、 当該音片と読みが共通する前記文章内の音片の時間 長の予測結果との差の関数より得られるものである、 ことを特徴とする音声選択方法。
1 8 . コンピュータを、
音声の波形を表す音声データを複数記憶する記憶手段と、 文章を表す文章情報を入力し、 当該文章を構成する音片について 韻律予測を行うことにより、 当該音片のピッチの時間変化を予測す る予測手段と、
各前記音声データのうちから、 前記文章を構成する音片と読みが 共通する音片の波形を表していて、 且つ、 ピッチの時間変化が前記 予測手段による予測の結果と最も高い相関を示す音声データを選択 する選択手段と、
して機能させるためのプログラム。
1 9 . コンピュータを、
音声の波形を表す音声データを複数記憶する記憶手段と、 文章を表す文章情報を入力し、 当該文章内の音片について韻律予 測を行うことにより、 当該音片の時間長、 及び、 当該音片のピッチ の時間変化を予測する予測手段と、
前記文章内の音片と読みが共通する音片の波形を表す各々の音声 データについての評価値を特定し、 評価値が最も高い評価を表して いる音声データを選択する選択手段として機能させるためのプログ ラムであって、
前記評価値は、 音声データが表す音片のピッチの時間変化と、 当 該音片と読みが共通する前記文章内の音片のピッチの時間変化の予 測結果との相関を表す数値の関数、 及び、 当該音声データが表す音 片の時間長と、 当該音片と読みが共通する前記文章内の音片の時間 長の予測結果との差の関数より得られるものである、 . ことを特徴とするプログラム。
2 0 . 音声の波形を表す音声データを複数記憶する記憶手段と、 文章を表す文章情報を入力する文章情報入力手段と、
前記文章情報が表す文章内の音片と読みが共通する部分を有する 音声データを索出する検索部と、 前記索出されたそれぞれの音声デ 一夕を文章情報が表す文章に従って接続した際に互いに隣接する音 声データ同士の関係に基づいた所定の評価基準に従って評価値を求 め、 出力する音声データの組み合わせを当該評価値に基づいて選択 する選択手段と、 を備える、
ことを特徴とする音声データ選択装置。
2 1 . 前記評価基準は、 互いに隣接する音声データ同士の関係を示 す評価値を定める基準であって、 前記評価値は、 前記音声データが 表す音声の特徴を示すパラメ一夕、 前記音声データが表す音声を互 いに結合して得られる音声の特徴を示すパラメ一タ、 及び、 発話時 間長に関する特徴を示すパラメ一夕のうち、 少なく ともいずれかを 含む評価式に基づいて得られるものである、
ことを特徴とする請求項 2 0に記載の音声データ選択装置。
2 2 . 前記評価基準は、 互いに隣接する音声データ同士の関係を示 す評価値を定める基準であって、 前記評価値は、 前記音声データが 表す音声を互いに結合して得られる音声の特徴を示すパラメ一夕を 含み、 また、 前記音声データが表す音声の特徴を示すパラメ一夕と 発話時間長に関する特徴を示すパラメータのうち、 少なく ともいず れかを含む評価式に基づいて得られるものである、
ことを特徵とする請求項 2 0に記載の音声データ選択装置。
2 3 . 前記音声データが表す音声を互いに結合して得られる音声の 特徴を示すパラメータは、 前記文章情報が表す文章内の音片と読み が共通する部分を有する音声の波形を表す音声データのうちから、 前記文章を構成するそれぞれの音片に相当する音声データを 1個ず つ選択した場合における、 互いに隣接する音声データ同士の境界で のピッチの差に基づいて得られるものである、
ことを特徴とする請求項 2 1又は 2 2に記載の音声デ一夕選択装
2 4 . 前記評価基準は、 更に音声データが表す音声との韻律予測結 果との相関ないし差異を示す評価値を定める基準を含み、 前記評価 値は、 音声データが表す音片のピッチの時間変化と、 当該音片と読 みが共通する前記文章内の音片のピツチの時間変化の予測結果との 相関を表す数値の関数、 及び/又は、 当該音声データが表す音片の 時間長と、 当該音片と読みが共通する前記文章内の音片の時間長の 予測結果との差の関数に基づいて得られるものである、
ことを特徴とする請求項 2 0乃至 2 3のいずれか 1項に記載の音 声データ選択装置。
2 5 . 前記相関を表す数値は、 音声データが表す音片のピッチの時 間変化と、 当該音片と読みが共通する前記文章内の音片のピッチの 時間変化との間での 1次回帰により得られる 1次関数の勾配及び Z 又は切片からなる、
ことを特徴とする請求項 2 に記載の音声データ選択装置。
2 6 . 前記相関を表す数値は、 音声データが表す音片のピッチの時 間変化と、 当該音片と読みが共通する前記文章内の音片のピッチの 時間変化の予測結果との間の相関係数からなる、 ' ことを特徴とする請求項 2 4又は 2 5に記載の音声データ選択装'
2 7 . 前記相関を表す数値は、 音声データが表す音片のピッチの時 間変化を表すデータを種々のビッ ト数循環シフ 卜したものが表す関 数と、 当該音片と読みが共通する前記文章内の音片のピツチの時間 変化の予測結果を表す関数との相関係数の最大値からなる、
ことを特徴とする請求項 2 4又は 2 5に記載の音声データ選択装 置。
2 8 . 前記記憶手段は、 音声データの読みを表す表音データを、 当 該音声デ一夕に対応付けて記憶しており、
前記選択手段は、 前記文章内の音片の読みに合致する読みを表す 表音データが対応付けられている音声データを、 当該音片と読みが 共通する音片の波形を表す音声デ一夕として扱う、
ことを特徴とする請求項 2 0乃至 2 7のいずれか 1項に記載の音 声データ選択装置。
2 9 . 選択された音声データを互いに結合することにより、 合成音 声を表すデータを生成する音声合成手段を更に備える、
ことを特徴とする請求項 2 0乃至 2 8のいずれか 1項に記載の音 声データ選択装置。
3 0 . 前記文章内の音片のうち、 前記選択手段が音声デ一夕を選択 できなかった音片について、 前記記憶手段が記憶する音声データを 用いることなく、 当該音片の波形を表す音声データを合成する欠落 部分合成手段を備え、
前記音声合成手段は、 前記選択手段が選択した音声データ及び前 記欠落部分合成手段が合成した音声データを互いに結合することに より、 合成音声を表すデ一夕を生成する、
ことを特徴とする請求項 2 9に記載の音声データ選択装置。
3 1 . 音声の波形を表す音声データを複数記憶し、
文章を表す文章情報を入力し、
前記文章情報が表す文章内の音片と読みが共通する部分を有する 音声データを索出し、
前記索出されたそれぞれの音声データを文章情報が表す文章に従 つて接続した際に互いに隣接する音声デ一夕同士の関係に基づいた 所定の評価基準に従って評価値を求め、 出力する音声データの組み 合わせを当該評価値に基づいて選択する、
ことを特徴とする音声データ選択方法。
3 2 . コンピュータを、
音声の波形を表す音声データを複数記憶する記憶手段と、 文章を表す文章情報を入力する文章情報入力手段と、
前記文章情報が表す文章内の音片と読みが共通する部分を有する 音声データを索出する検索部と、
前記索出されたそれぞれの音声デ一夕を文章情報が表す文章に従 つて接続した際に互いに隣接する音声データ同士の関係に基づいた 所定の評価基準に従って評価値を求め、 出力する音声データの組み 合わせを当該評価値に基づいて選択する選択手段と、
して機能させるためのプログラム。
PCT/JP2004/008088 2003-06-04 2004-06-03 音声データを選択するための装置、方法およびプログラム Ceased WO2004109660A1 (ja)

Priority Applications (4)

Application Number Priority Date Filing Date Title
US10/559,573 US20070100627A1 (en) 2003-06-04 2004-06-03 Device, method, and program for selecting voice data
CN2004800187934A CN1816846B (zh) 2003-06-04 2004-06-03 用于选择话音数据的设备和方法
EP04735989A EP1632933A4 (en) 2003-06-04 2004-06-03 DEVICE, METHOD AND PROGRAM FOR SELECTING VOICE DATA
DE04735989T DE04735989T1 (de) 2003-06-04 2004-06-03 Einrichtung, verfahren und programm zur auswahl von voice-daten

Applications Claiming Priority (6)

Application Number Priority Date Filing Date Title
JP2003159880 2003-06-04
JP2003-159880 2003-06-04
JP2003165582 2003-06-10
JP2003-165582 2003-06-10
JP2004155306A JP4264030B2 (ja) 2003-06-04 2004-05-25 音声データ選択装置、音声データ選択方法及びプログラム
JP2004-155306 2004-05-25

Publications (1)

Publication Number Publication Date
WO2004109660A1 true WO2004109660A1 (ja) 2004-12-16

Family

ID=33514559

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/JP2004/008088 Ceased WO2004109660A1 (ja) 2003-06-04 2004-06-03 音声データを選択するための装置、方法およびプログラム

Country Status (7)

Country Link
US (1) US20070100627A1 (ja)
EP (1) EP1632933A4 (ja)
JP (1) JP4264030B2 (ja)
KR (1) KR20060015744A (ja)
CN (1) CN1816846B (ja)
DE (1) DE04735989T1 (ja)
WO (1) WO2004109660A1 (ja)

Families Citing this family (9)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US7109208B2 (en) 2001-04-11 2006-09-19 Senju Pharmaceutical Co., Ltd. Visual function disorder improving agents
CN1813285B (zh) * 2003-06-05 2010-06-16 株式会社建伍 语音合成设备和方法
JP4516863B2 (ja) * 2005-03-11 2010-08-04 株式会社ケンウッド 音声合成装置、音声合成方法及びプログラム
JP2008185805A (ja) * 2007-01-30 2008-08-14 Internatl Business Mach Corp <Ibm> 高品質の合成音声を生成する技術
JP5387410B2 (ja) * 2007-10-05 2014-01-15 日本電気株式会社 音声合成装置、音声合成方法および音声合成プログラム
JP5093387B2 (ja) * 2011-07-19 2012-12-12 ヤマハ株式会社 音声特徴量算出装置
CN111506736B (zh) * 2020-04-08 2023-08-08 北京百度网讯科技有限公司 文本发音获取方法、装置和电子设备
CN112669810B (zh) * 2020-12-16 2023-08-01 平安科技(深圳)有限公司 语音合成的效果评估方法、装置、计算机设备及存储介质
CN114495902B (zh) * 2022-02-25 2025-10-17 北京有竹居网络技术有限公司 语音合成方法、装置、计算机可读介质及电子设备

Citations (11)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPH01284898A (ja) * 1988-05-11 1989-11-16 Nippon Telegr & Teleph Corp <Ntt> 音声合成方法
JPH07319497A (ja) * 1994-05-23 1995-12-08 N T T Data Tsushin Kk 音声合成装置
JPH0944191A (ja) * 1995-05-25 1997-02-14 Sanyo Electric Co Ltd 音声合成装置
JPH09230893A (ja) * 1996-02-22 1997-09-05 N T T Data Tsushin Kk 規則音声合成方法及び音声合成装置
JPH1097268A (ja) * 1996-09-24 1998-04-14 Sanyo Electric Co Ltd 音声合成装置
JPH11249679A (ja) * 1998-03-04 1999-09-17 Ricoh Co Ltd 音声合成装置
JPH11259083A (ja) * 1998-03-09 1999-09-24 Canon Inc 音声合成装置および方法
JP2001013982A (ja) * 1999-04-28 2001-01-19 Victor Co Of Japan Ltd 音声合成装置
JP2001034284A (ja) * 1999-07-23 2001-02-09 Toshiba Corp 音声合成方法及び装置、並びに文音声変換プログラムを記録した記録媒体
JP2001092481A (ja) * 1999-09-24 2001-04-06 Sanyo Electric Co Ltd 規則音声合成方法
JP2003513311A (ja) * 1999-10-28 2003-04-08 シーメンス アクチエンゲゼルシヤフト 合成すべき音声応答の基本周波数の時間特性を定めるための方法

Family Cites Families (8)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US5636325A (en) * 1992-11-13 1997-06-03 International Business Machines Corporation Speech synthesis and analysis of dialects
JP3587048B2 (ja) * 1998-03-02 2004-11-10 株式会社日立製作所 韻律制御方法及び音声合成装置
JP3180764B2 (ja) * 1998-06-05 2001-06-25 日本電気株式会社 音声合成装置
US6505152B1 (en) * 1999-09-03 2003-01-07 Microsoft Corporation Method and apparatus for using formant models in speech systems
US6496801B1 (en) * 1999-11-02 2002-12-17 Matsushita Electric Industrial Co., Ltd. Speech synthesis employing concatenated prosodic and acoustic templates for phrases of multiple words
US6865533B2 (en) * 2000-04-21 2005-03-08 Lessac Technology Inc. Text to speech
CA2359771A1 (en) * 2001-10-22 2003-04-22 Dspfactory Ltd. Low-resource real-time audio synthesis system and method
US20040030555A1 (en) * 2002-08-12 2004-02-12 Oregon Health & Science University System and method for concatenating acoustic contours for speech synthesis

Patent Citations (11)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPH01284898A (ja) * 1988-05-11 1989-11-16 Nippon Telegr & Teleph Corp <Ntt> 音声合成方法
JPH07319497A (ja) * 1994-05-23 1995-12-08 N T T Data Tsushin Kk 音声合成装置
JPH0944191A (ja) * 1995-05-25 1997-02-14 Sanyo Electric Co Ltd 音声合成装置
JPH09230893A (ja) * 1996-02-22 1997-09-05 N T T Data Tsushin Kk 規則音声合成方法及び音声合成装置
JPH1097268A (ja) * 1996-09-24 1998-04-14 Sanyo Electric Co Ltd 音声合成装置
JPH11249679A (ja) * 1998-03-04 1999-09-17 Ricoh Co Ltd 音声合成装置
JPH11259083A (ja) * 1998-03-09 1999-09-24 Canon Inc 音声合成装置および方法
JP2001013982A (ja) * 1999-04-28 2001-01-19 Victor Co Of Japan Ltd 音声合成装置
JP2001034284A (ja) * 1999-07-23 2001-02-09 Toshiba Corp 音声合成方法及び装置、並びに文音声変換プログラムを記録した記録媒体
JP2001092481A (ja) * 1999-09-24 2001-04-06 Sanyo Electric Co Ltd 規則音声合成方法
JP2003513311A (ja) * 1999-10-28 2003-04-08 シーメンス アクチエンゲゼルシヤフト 合成すべき音声応答の基本周波数の時間特性を定めるための方法

Non-Patent Citations (1)

* Cited by examiner, † Cited by third party
Title
See also references of EP1632933A4 *

Also Published As

Publication number Publication date
KR20060015744A (ko) 2006-02-20
EP1632933A4 (en) 2007-11-14
JP4264030B2 (ja) 2009-05-13
EP1632933A1 (en) 2006-03-08
JP2005025173A (ja) 2005-01-27
US20070100627A1 (en) 2007-05-03
CN1816846B (zh) 2010-06-09
DE04735989T1 (de) 2006-10-12
CN1816846A (zh) 2006-08-09

Similar Documents

Publication Publication Date Title
CN1813285B (zh) 语音合成设备和方法
JP4516863B2 (ja) 音声合成装置、音声合成方法及びプログラム
US20070011009A1 (en) Supporting a concatenative text-to-speech synthesis
JP4264030B2 (ja) 音声データ選択装置、音声データ選択方法及びプログラム
US7089187B2 (en) Voice synthesizing system, segment generation apparatus for generating segments for voice synthesis, voice synthesizing method and storage medium storing program therefor
JP4287785B2 (ja) 音声合成装置、音声合成方法及びプログラム
JP2005018036A (ja) 音声合成装置、音声合成方法及びプログラム
JP2010224418A (ja) 音声合成装置、方法およびプログラム
JP2004272236A (ja) ピッチ波形信号分割装置、音声信号圧縮装置、データベース、音声信号復元装置、音声合成装置、ピッチ波形信号分割方法、音声信号圧縮方法、音声信号復元方法、音声合成方法、記録媒体及びプログラム
JP4209811B2 (ja) 音声選択装置、音声選択方法及びプログラム
JP4780188B2 (ja) 音声データ選択装置、音声データ選択方法及びプログラム
JP7183556B2 (ja) 合成音生成装置、方法、及びプログラム
JP2010224419A (ja) 音声合成装置、方法およびプログラム
JP4574333B2 (ja) 音声合成装置、音声合成方法及びプログラム
JP4184157B2 (ja) 音声データ管理装置、音声データ管理方法及びプログラム
KR20100003574A (ko) 음성음원정보 생성 장치 및 시스템, 그리고 이를 이용한음성음원정보 생성 방법
JP2006145848A (ja) 音声合成装置、音片記憶装置、音片記憶装置製造装置、音声合成方法、音片記憶装置製造方法及びプログラム
JP2006145690A (ja) 音声合成装置、音声合成方法及びプログラム
JP2007240988A (ja) 音声合成装置、データベース、音声合成方法及びプログラム
JP2007240989A (ja) 音声合成装置、音声合成方法及びプログラム
JP2006195207A (ja) 音声合成装置、音声合成方法及びプログラム
JP2007240987A (ja) 音声合成装置、音声合成方法及びプログラム
JP2007240990A (ja) 音声合成装置、音声合成方法及びプログラム

Legal Events

Date Code Title Description
AK Designated states

Kind code of ref document: A1

Designated state(s): AE AG AL AM AT AU AZ BA BB BG BR BW BY BZ CA CH CN CO CR CU CZ DE DK DM DZ EC EE EG ES FI GB GD GE GH GM HR HU ID IL IN IS KE KG KP KR KZ LC LK LR LS LT LU LV MA MD MG MK MN MW MX MZ NA NI NO NZ OM PG PH PL PT RO RU SC SD SE SG SK SL SY TJ TM TN TR TT TZ UA UG US UZ VC VN YU ZA ZM ZW

AL Designated countries for regional patents

Kind code of ref document: A1

Designated state(s): GM KE LS MW MZ NA SD SL SZ TZ UG ZM ZW AM AZ BY KG KZ MD RU TJ TM AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HU IE IT LU MC NL PL PT RO SE SI SK TR BF BJ CF CG CI CM GA GN GQ GW ML MR NE SN TD TG

121 Ep: the epo has been informed by wipo that ep was designated in this application
WWE Wipo information: entry into national phase

Ref document number: 2004735989

Country of ref document: EP

WWE Wipo information: entry into national phase

Ref document number: 1020057023078

Country of ref document: KR

WWE Wipo information: entry into national phase

Ref document number: 20048187934

Country of ref document: CN

WWP Wipo information: published in national office

Ref document number: 1020057023078

Country of ref document: KR

WWP Wipo information: published in national office

Ref document number: 2004735989

Country of ref document: EP

WWE Wipo information: entry into national phase

Ref document number: 2007100627

Country of ref document: US

Ref document number: 10559573

Country of ref document: US

WWP Wipo information: published in national office

Ref document number: 10559573

Country of ref document: US