EP0058130A2 - Procédé pour la synthèse de la parole avec un vocabulaire illimité et dispositif pour la mise en oeuvre dudit procédé - Google Patents

Procédé pour la synthèse de la parole avec un vocabulaire illimité et dispositif pour la mise en oeuvre dudit procédé Download PDF

Info

Publication number
EP0058130A2
EP0058130A2 EP82730011A EP82730011A EP0058130A2 EP 0058130 A2 EP0058130 A2 EP 0058130A2 EP 82730011 A EP82730011 A EP 82730011A EP 82730011 A EP82730011 A EP 82730011A EP 0058130 A2 EP0058130 A2 EP 0058130A2
Authority
EP
European Patent Office
Prior art keywords
sound
elements
sounds
samples
synthesis
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Granted
Application number
EP82730011A
Other languages
German (de)
English (en)
Other versions
EP0058130B1 (fr
EP0058130A3 (en
Inventor
Eberhard Dr.-Ing. Grossmann
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Grossmann Eberhard Dr-Ing
Original Assignee
Fraunhofer Institut fuer Nachrichtentechnik Heinrich Hertz Institute HHI
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Fraunhofer Institut fuer Nachrichtentechnik Heinrich Hertz Institute HHI filed Critical Fraunhofer Institut fuer Nachrichtentechnik Heinrich Hertz Institute HHI
Priority to AT82730011T priority Critical patent/ATE20784T1/de
Publication of EP0058130A2 publication Critical patent/EP0058130A2/fr
Publication of EP0058130A3 publication Critical patent/EP0058130A3/de
Application granted granted Critical
Publication of EP0058130B1 publication Critical patent/EP0058130B1/fr
Expired legal-status Critical Current

Links

Images

Classifications

    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L13/00Speech synthesis; Text to speech systems
    • G10L13/02Methods for producing synthetic speech; Speech synthesisers
    • G10L13/04Details of speech synthesis systems, e.g. synthesiser structure or memory management

Definitions

  • the invention relates to a method for the synthesis of speech with unlimited vocabulary in the time domain from sound elements, which are obtained from natural speech samples and coded in digital form, with low redundancy, stored and also in view of the required storage space in length in each case on the significant range of relevant sound-typical time signal and the number are reduced by utilizing mutually convertible related sounds, for speech synthesis these sound elements are chained in the required form, number and order to form digital signal sequences based on input commands and predetermined linking rules, from which by means of digital-analog Conversion and controllable amplification as speech-perceptible sound waves are generated, as well as on a circuit arrangement for performing the method.
  • Speech synthesis is understood to mean the conversion of a text present as a symbol sequence into the equivalent acoustic signal by means of technical equipment. It is of fundamental importance that between the input of the symbol sequence in the apparatus and the output of the equivalent acoustic signal, all processes take place immediately, without the interposition of human mind powers. The precisely determined individual technical measures follow the planned use of predictable and controllable natural forces.
  • the evaluation criteria for synthetic language are intelligibility and naturalness.
  • the standards for this are, even if e.g. in terms of intelligibility ascertainable from an objective point of view, subjective nature. Nevertheless, there are circumstances that everyone can immediately use for the assessment. These are the course of the basic pitch (pitch frequency), the speaking rhythm and the course of the intensity.
  • the individual sounds merge into one another in the course of the natural language signal. They are characterized by several sound generation frequencies (formants). These sound generation frequencies are independent of the fundamental pitch, i.e. regardless of the speech height.
  • the speech synthesis system known from DE-OS 20 16 572 takes into account the problems at the transitions between successive phonemes in particular with regard to intelligibility. Since the formant frequencies - taking into account the three main formants is sufficient - to the Increase, decrease or remain transitions, there are nine versions for every phoneme to be stored. In order not to have to increase the storage capacity by practically a further power of ten, the solution in this known prior art aims to get by with a stored version and to modify this representation according to the requirements during the synthesis process. In addition, only the significant range of the individual sounds is saved, which, for example with a / s / sound, only has to be 10% of the total duration and can therefore be reproduced exactly enough and understandably by repeating ten times.
  • the stored sections should start with an oscillation zero crossing.
  • the suitability at the transition to other phonemes must also be selected in a special way - a subjective test. With this compromise, abrupt transitions can be avoided or at least reduced to a small extent, but on the other hand, completely bumpless transitions must be avoided.
  • the language generator known from DE-OS 23 06 816 is based on the task in the preparation of phonetic segments to create a comprehensive pitch period control range of the synthesized sounds, which should benefit the improvement of naturalness and intelligibility.
  • the solution given is to pick out sound waveforms from natural language for voiced sounds with a defined periodicity of each pitch length and to add a waveform to each such waveform at the end area that was obtained by a rough calculation for the waveform of the respective sound. Sound waveforms of unvoiced sounds and the transitions between consonants and vowels that have an undefined periodicity should be divided into fixed lengths.
  • the pitch period can be changed, lengthened or shortened, and thus the basic pitch can be lowered or raised accordingly, if samples are used or omitted at these discrete locations, without the sound character changing thereby.
  • samples approximately 30 within such a significant area, special “samples” are used, the marker words, which allow these locations to be found at any time.
  • the marker words themselves are omitted when the elements are linked to form the digital signal sequences.
  • 60 samples for example those adjacent to a marker word, depending on whether they are used or not, permit a practically continuous variation in the pitch, that is to say a large number of melody lines. In particular, this also enables the fundamental speech frequency curves at the transitions to the following sounds to be designed continuously, that is, to avoid bumps.
  • transitions - with the exception of combinations of plosives - can be inverted in time; by lengthening or shortening the length of time, vowel conversions take place; by shortening the length of time, there are also consonant conversions.
  • the required sound elements are made up of almost 60 elements for transitional sounds, 27 elements for voiced individual sounds and 13 elements for unvoiced individual sounds. Further details follow in connection with the description of the figures.
  • Particularly preferred embodiments of the invention consist in providing additional samples in the digitally stored elements for the voiced individual sounds for the purpose of pitch variation. This measure leads to a slight increase of approx. 1000 bytes of the required storage space, but allows more extensive variations in the melody.
  • an additional sample value has an interpolated value lying between the adjacent true sample values. In this way, any discontinuities that would occur between the true samples that are definitely needed and used can be reduced or avoided.
  • Marking words should preferably be provided at points with a slight slope in the time signal.
  • An associated error signal has very small deflections at such locations and thus allows the desired discrete locations to be determined, localized and marked in a simple manner.
  • marker words are preferred in the embodiments of the invention.
  • a marker word and a true or additional sample can have digital patterns of the same stock.
  • marker words are to be reserved for digital words which do not occur in the sample values.
  • embodiments of the invention it is of particular importance for embodiments of the invention to be able to determine the shape of the sound elements required for the concatenation of the next word following the pauses on the basis of the input commands. This avoids discontinuities in the output of the individual words.
  • the duration for determining the shape of the required synthesis building blocks, even for very long words, is in the range of a few milliseconds. Determining the shape is to be understood here: searching for the relevant sound element, possibly inverting in time, lengthening or shortening the duration of the sound and specifying the number of repetitions of the stored sound element.
  • a further essential advantage of the invention is that sequences of conventional characters entered via an alphanumeric keyboard can be automatically transcribed into a sequence of phonetic characters suitable as input commands in a method step preceding the actual synthesis process. As a result, even inexperienced or untrained users will find it much easier to use, or even opened up. Of course, there is also the option of entering phonetic characters or the appropriate input commands directly.
  • a circuit arrangement for carrying out the method according to the invention can be constructed with a microprocessor, to which the read-only memory with a total storage capacity of 32 kbytes and a working memory for 1 kbyte are to be connected, and also has a decomposing digital-to-analog converter and a volume-controllable low-frequency amplifier and a loudspeaker as an electro-acoustic transducer device.
  • Such circuit elements and components are common on the market. The concept but also enables extensive integration. Decomparing before the digital-to-analog conversion naturally means that previously the stored data has been subjected to a coding which reduces the data rate.
  • the logarithmic PCM and the adaptive delta PCM are common and increasingly reducing methods in the order given. Relevant components are known from common voice transmission systems and can also be used without further ado in embodiments of the invention.
  • the data input i.e. the writing or phonetic symbol sequences
  • the output of the acoustic signals can take place both directly on the device and at remote locations.
  • a V24 interface or a low-frequency socket can be provided at the output.
  • a speech synthesis system in embodiments according to the invention essentially consists of two units, the one for the transcription and the one for the synthesis itself. Either a character string is to be entered, which is done via an alphanumeric keyboard or via a V24 interface can, or a sound string. Although experienced or trained users can also enter the sound strings directly using suitable keyboards, in most applications, if the transcription is not used, the synthesis unit will then receive the corresponding input signals from a remote location via a data line and the V24 interface. Of course, other interface conditions can also be complied with and implemented within the scope of professional skills.
  • the transcription unit uses prepared rules, summarized under the term grammar, the synthesis unit essentially uses the stored sound elements.
  • the synthesized Sampling value sequences arrive via a digital-to-analog converter D / A and a controllable amplifier either directly via a loudspeaker or via a low-frequency socket and a voice transmission line, not shown, and at a remote location via a loudspeaker as sound waves for reproduction, better output,
  • the block diagram shown in FIG. 2 shows, in particular in the size comparison of the individual blocks, the storage space requirement with the proportions that are required for the synthesis and the transcription as a whole.
  • the system is designed on the basis of a microprocessor pP.
  • An alphanumeric keyboard is provided for entering the character strings, and a conventional electro-acoustic converter is provided for outputting the sound waves perceptible as speech.
  • the microprocessor pP works with the transcription program TP and the transcription grammar TG, for speech synthesis with the synthesis program SP and the synthesis matrix SM, the required sound elements being taken from the sound element memory SE as required and into the RAM stored in the working memory from which Derived shape of the relevant sound string, chained in the relevant number and order and passed to the digital-to-analog converter (see FIG. 1, D / A).
  • a volume control within the synthesized words and sentences takes place, also controlled by the microprocessor pP and according to commands entered therefor, in the controllable low-frequency amplifier (see FIG. 1) before the radiation of the sound waves or the transmission of the low-frequency signal.
  • the position of the first three formants for nine different sounds shown in FIG. 3 shows that the first and the second formants in particular are of considerable importance for the formation of sounds. Due to the linear division of the frequency scale, it should not be overlooked that in the third formant, the range is about half an octave.
  • FIGS. 6a, 6b and 6c for temporal inversion of transitions (FIG. 6a), for vowel conversion (FIG. 6b) and for consonant conversion (FIG. 6c) speak for themselves and therefore do not require any further explanation here.
  • shortening or lengthening the duration of the sound not only brings about a shift in the pitch, but in particular causes a sound conversion.
  • the 16 sounds indicated in FIG. 6 c only those given in the first place in each line need to be stored. Although these are the sounds with the most required sample values in each case, this saves storage space of a good 60% compared to storing all of these sounds.
  • the change in the auditory impression shown in FIG. 7 indicates that 20 test persons should find a consonant conversion (in brackets) which - apart from two persons when the starting point was shifted to 160 ms - confirmed the stated auditory impression in the individual conversion forms.
  • FIG. 8a, 8b and 8c show an example of the manner in which the pitch variation which is essential in the invention is made possible.
  • a basic frequency period of the sound / a / is plotted in FIG. 8a.
  • the associated error signal (FIG. 8b) is first generated by a prediction error filter. From this it can be seen that discrete places can be specified where modifications have to be made without changing the sound character but its pitch.
  • 8c shows the period of the sound / a / shortened by approximately 20% compared to FIG. 8a. It can be seen in the comparison of the curves of Figures 8a and 8c that a shortening of the period, i.e. an increase in pitch, the actual characteristic image does not change, the sound / a / is therefore retained as such and - as desired - sounds higher.
  • the block shown in FIG. 10 is intended to illustrate the ratio of the storage space requirement which is required for the synthesis building blocks, the elements of the individual and the transition sounds. These are primarily the true sample values WAW of the elements, but also the marking words MAW and the mathematically determined additional sample values ZAW for the voiced individual sounds or the voiced areas of transition sounds. The dashed line between the areas for the individual sound and the transition sound elements shows a division roughly in a ratio of 4: 6.
  • the lexical processing shows that it is no exception.
  • the word analysis is done according to:

Landscapes

  • Engineering & Computer Science (AREA)
  • Computational Linguistics (AREA)
  • Health & Medical Sciences (AREA)
  • Audiology, Speech & Language Pathology (AREA)
  • Human Computer Interaction (AREA)
  • Physics & Mathematics (AREA)
  • Acoustics & Sound (AREA)
  • Multimedia (AREA)
  • Machine Translation (AREA)
  • Electrophonic Musical Instruments (AREA)
  • Compression, Expansion, Code Conversion, And Decoders (AREA)
EP82730011A 1981-02-11 1982-02-11 Procédé pour la synthèse de la parole avec un vocabulaire illimité et dispositif pour la mise en oeuvre dudit procédé Expired EP0058130B1 (fr)

Priority Applications (1)

Application Number Priority Date Filing Date Title
AT82730011T ATE20784T1 (de) 1981-02-11 1982-02-11 Verfahren zur synthese von sprache mit unbegrenztem wortschatz und schaltungsanordnung zur durchfuehrung des verfahrens.

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
DE19813105518 DE3105518A1 (de) 1981-02-11 1981-02-11 Verfahren zur synthese von sprache mit unbegrenztem wortschatz und schaltungsanordnung zur durchfuehrung des verfahrens
DE3105518 1981-02-11

Publications (3)

Publication Number Publication Date
EP0058130A2 true EP0058130A2 (fr) 1982-08-18
EP0058130A3 EP0058130A3 (en) 1982-09-08
EP0058130B1 EP0058130B1 (fr) 1986-07-16

Family

ID=6124949

Family Applications (1)

Application Number Title Priority Date Filing Date
EP82730011A Expired EP0058130B1 (fr) 1981-02-11 1982-02-11 Procédé pour la synthèse de la parole avec un vocabulaire illimité et dispositif pour la mise en oeuvre dudit procédé

Country Status (4)

Country Link
EP (1) EP0058130B1 (fr)
AT (1) ATE20784T1 (fr)
CA (1) CA1172365A (fr)
DE (2) DE3105518A1 (fr)

Cited By (7)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
EP0107945A1 (fr) * 1982-10-19 1984-05-09 Kabushiki Kaisha Toshiba Dispositif pour la synthèse de la parole
EP0144731A3 (en) * 1983-11-01 1985-07-03 Nec Corporation Speech synthesizer
EP0181339A4 (fr) * 1984-04-10 1986-12-08 First Byte Systeme de conversion texte-parole en temps reel.
DE3530856A1 (de) * 1985-08-29 1987-03-05 Telefonbau & Normalzeit Gmbh Schaltungsanordnung zur akustischen bedienungshilfe bei fernmelde-, insbesondere fernsprechendgeraeten
US4862504A (en) * 1986-01-09 1989-08-29 Kabushiki Kaisha Toshiba Speech synthesis system of rule-synthesis type
FR2655759A1 (fr) * 1989-12-13 1991-06-14 Hodys Edgar Dispositif electronique de synthese vocale sans vocabulaire de reference.
CN118016079A (zh) * 2024-04-07 2024-05-10 广州市艾索技术有限公司 一种智能语音转写方法及系统

Families Citing this family (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
DE3513243A1 (de) * 1985-04-13 1986-10-16 Telefonbau Und Normalzeit Gmbh, 6000 Frankfurt Verfahren zur sprachuebertragung und sprachspeicherung
DE19860133C2 (de) * 1998-12-17 2001-11-22 Cortologic Ag Verfahren und Vorrichtung zur Sprachkompression

Family Cites Families (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPS5331323B2 (fr) * 1972-11-13 1978-09-01

Cited By (8)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
EP0107945A1 (fr) * 1982-10-19 1984-05-09 Kabushiki Kaisha Toshiba Dispositif pour la synthèse de la parole
EP0144731A3 (en) * 1983-11-01 1985-07-03 Nec Corporation Speech synthesizer
EP0181339A4 (fr) * 1984-04-10 1986-12-08 First Byte Systeme de conversion texte-parole en temps reel.
DE3530856A1 (de) * 1985-08-29 1987-03-05 Telefonbau & Normalzeit Gmbh Schaltungsanordnung zur akustischen bedienungshilfe bei fernmelde-, insbesondere fernsprechendgeraeten
US4862504A (en) * 1986-01-09 1989-08-29 Kabushiki Kaisha Toshiba Speech synthesis system of rule-synthesis type
FR2655759A1 (fr) * 1989-12-13 1991-06-14 Hodys Edgar Dispositif electronique de synthese vocale sans vocabulaire de reference.
CN118016079A (zh) * 2024-04-07 2024-05-10 广州市艾索技术有限公司 一种智能语音转写方法及系统
CN118016079B (zh) * 2024-04-07 2024-06-07 广州市艾索技术有限公司 一种智能语音转写方法及系统

Also Published As

Publication number Publication date
EP0058130B1 (fr) 1986-07-16
CA1172365A (fr) 1984-08-07
DE3105518A1 (de) 1982-08-19
DE3271965D1 (en) 1986-08-21
ATE20784T1 (de) 1986-08-15
EP0058130A3 (en) 1982-09-08

Similar Documents

Publication Publication Date Title
DE69031165T2 (de) System und methode zur text-sprache-umsetzung mit hilfe von kontextabhängigen vokalallophonen
EP0886853B1 (fr) Procede de synthese vocale a base de microsegments
DE60035001T2 (de) Sprachsynthese mit Prosodie-Mustern
DE69821673T2 (de) Verfahren und Vorrichtung zum Editieren synthetischer Sprachnachrichten, sowie Speichermittel mit dem Verfahren
DE69718284T2 (de) Sprachsynthesesystem und Wellenform-Datenbank mit verringerter Redundanz
DE60118874T2 (de) Prosodiemustervergleich für Text-zu-Sprache Systeme
DE69719270T2 (de) Sprachsynthese unter Verwendung von Hilfsinformationen
DE60126564T2 (de) Verfahren und Anordnung zur Sprachsysnthese
DE60112512T2 (de) Kodierung von Ausdruck in Sprachsynthese
DE69620399T2 (de) Sprachsynthese
DE69506037T2 (de) Audioausgabeeinheit und Methode
EP1184839A2 (fr) Conversion graphème-phonème
DD143970A1 (de) Verfahren und anordnung zur synthese von sprache
EP1105867B1 (fr) Procede et dispositif permettant de concatener des segments audio en tenant compte de la coarticulation
DE4237563A1 (fr)
DE112004000187T5 (de) Verfahren und Vorrichtung der prosodischen Simulations-Synthese
EP1058235B1 (fr) Procédé de reproduction pour systèmes contrôlés par la voix avec synthèse de la parole basée sur texte
EP0058130B1 (fr) Procédé pour la synthèse de la parole avec un vocabulaire illimité et dispositif pour la mise en oeuvre dudit procédé
DE60205421T2 (de) Verfahren und Vorrichtung zur Sprachsynthese
DE2519483A1 (de) Verfahren und anordnung zur sprachsynthese
EP1110203B1 (fr) Procede et dispositif de traitement numerique de la voix
DE69607928T2 (de) Verfahren und vorrichtung zur bereitstellung und verwendung von diphonen für mehrsprachige text-nach-sprache systeme
EP1344211B1 (fr) Vorrichtung und verfahren zur differenzierten sprachausgabe
EP1554715B1 (fr) Procede de synthese de la parole assistee par ordinateur sous forme de signal vocal analogique a partir d'un texte electronique, dispositif de synthese de la parole et appareil de telecommunication
DE3232835C2 (fr)

Legal Events

Date Code Title Description
PUAI Public reference made under article 153(3) epc to a published international application that has entered the european phase

Free format text: ORIGINAL CODE: 0009012

PUAL Search report despatched

Free format text: ORIGINAL CODE: 0009013

AK Designated contracting states

Designated state(s): AT CH DE FR GB

AK Designated contracting states

Designated state(s): AT CH DE FR GB

17P Request for examination filed

Effective date: 19830210

RAP1 Party data changed (applicant data changed or rights of an application transferred)

Owner name: GROSSMANN, EBERHARD, DR.-ING.

RIN1 Information on inventor provided before grant (corrected)

Inventor name: GROSSMANN, EBERHARD, DR.-ING.

GRAA (expected) grant

Free format text: ORIGINAL CODE: 0009210

AK Designated contracting states

Kind code of ref document: B1

Designated state(s): AT CH DE FR GB LI

PG25 Lapsed in a contracting state [announced via postgrant information from national office to epo]

Ref country code: FR

Free format text: THE PATENT HAS BEEN ANNULLED BY A DECISION OF A NATIONAL AUTHORITY

Effective date: 19860716

REF Corresponds to:

Ref document number: 20784

Country of ref document: AT

Date of ref document: 19860815

Kind code of ref document: T

REF Corresponds to:

Ref document number: 3271965

Country of ref document: DE

Date of ref document: 19860821

EN Fr: translation not filed
PG25 Lapsed in a contracting state [announced via postgrant information from national office to epo]

Ref country code: AT

Effective date: 19870211

PLBE No opposition filed within time limit

Free format text: ORIGINAL CODE: 0009261

26N No opposition filed
GBPC Gb: european patent ceased through non-payment of renewal fee
PG25 Lapsed in a contracting state [announced via postgrant information from national office to epo]

Ref country code: GB

Effective date: 19881121

PGFP Annual fee paid to national office [announced via postgrant information from national office to epo]

Ref country code: CH

Payment date: 19900323

Year of fee payment: 9

PGFP Annual fee paid to national office [announced via postgrant information from national office to epo]

Ref country code: DE

Payment date: 19900430

Year of fee payment: 9

PG25 Lapsed in a contracting state [announced via postgrant information from national office to epo]

Ref country code: LI

Effective date: 19910228

Ref country code: CH

Effective date: 19910228

REG Reference to a national code

Ref country code: CH

Ref legal event code: PL

PG25 Lapsed in a contracting state [announced via postgrant information from national office to epo]

Ref country code: DE

Effective date: 19911101