CN1991981A - Method for voice data classification - Google Patents
Method for voice data classification Download PDFInfo
- Publication number
- CN1991981A CN1991981A CNA2005101217187A CN200510121718A CN1991981A CN 1991981 A CN1991981 A CN 1991981A CN A2005101217187 A CNA2005101217187 A CN A2005101217187A CN 200510121718 A CN200510121718 A CN 200510121718A CN 1991981 A CN1991981 A CN 1991981A
- Authority
- CN
- China
- Prior art keywords
- phoneme
- high amplitude
- data
- sound
- speech
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
- 238000000034 method Methods 0.000 title claims abstract description 40
- 238000001228 spectrum Methods 0.000 claims abstract description 29
- 238000010183 spectrum analysis Methods 0.000 claims abstract description 7
- 238000004891 communication Methods 0.000 claims description 13
- 238000001914 filtration Methods 0.000 claims description 13
- 238000012797 qualification Methods 0.000 claims description 2
- 238000004220 aggregation Methods 0.000 abstract 1
- 230000002776 aggregation Effects 0.000 abstract 1
- 230000006870 function Effects 0.000 description 13
- 238000012545 processing Methods 0.000 description 8
- 238000005516 engineering process Methods 0.000 description 7
- 238000013507 mapping Methods 0.000 description 7
- 230000008901 benefit Effects 0.000 description 6
- 238000005070 sampling Methods 0.000 description 6
- 230000009471 action Effects 0.000 description 5
- 238000001514 detection method Methods 0.000 description 3
- 238000012986 modification Methods 0.000 description 3
- 230000004048 modification Effects 0.000 description 3
- 230000003068 static effect Effects 0.000 description 3
- 206010038743 Restlessness Diseases 0.000 description 2
- 230000015572 biosynthetic process Effects 0.000 description 2
- 230000008878 coupling Effects 0.000 description 2
- 238000010168 coupling process Methods 0.000 description 2
- 238000005859 coupling reaction Methods 0.000 description 2
- 238000010586 diagram Methods 0.000 description 2
- 230000008859 change Effects 0.000 description 1
- 238000006243 chemical reaction Methods 0.000 description 1
- 238000013461 design Methods 0.000 description 1
- 230000000694 effects Effects 0.000 description 1
- 230000002650 habitual effect Effects 0.000 description 1
- 239000000463 material Substances 0.000 description 1
- 238000005259 measurement Methods 0.000 description 1
- 238000010295 mobile communication Methods 0.000 description 1
- 230000001360 synchronised effect Effects 0.000 description 1
- 238000003786 synthesis reaction Methods 0.000 description 1
- 238000012360 testing method Methods 0.000 description 1
- 230000000007 visual effect Effects 0.000 description 1
- 230000001755 vocal effect Effects 0.000 description 1
Images
Classifications
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L15/00—Speech recognition
- G10L15/02—Feature extraction for speech recognition; Selection of recognition unit
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L25/00—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00
- G10L25/93—Discriminating between voiced and unvoiced parts of speech signals
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L15/00—Speech recognition
- G10L15/02—Feature extraction for speech recognition; Selection of recognition unit
- G10L2015/025—Phonemes, fenemes or fenones being the recognition units
Landscapes
- Engineering & Computer Science (AREA)
- Computational Linguistics (AREA)
- Health & Medical Sciences (AREA)
- Audiology, Speech & Language Pathology (AREA)
- Human Computer Interaction (AREA)
- Physics & Mathematics (AREA)
- Acoustics & Sound (AREA)
- Multimedia (AREA)
- Signal Processing (AREA)
- Computer Vision & Pattern Recognition (AREA)
- Processing Or Creating Images (AREA)
- Telephonic Communication Services (AREA)
Abstract
A method which classifies the real voice data and not dense can be used to improve the cartoon of virtual role. The method includes acoustic voice section identified with voice data (step 410). Then the high amplitude frequency spectrum can be ensured by executing spectrum analysis for high amplitude component (step 415). The high amplitude frequency spectrum is classified to vowel phoneme that is selected from the vowel aggregation (step 440).
Description
Technical field
The present invention relates generally to speech recognition.Especially, though and non-exclusively, the present invention relates to speech data is analyzed and classified to help to make virtual role (avatars) animation.
Background technology
Speech recognition is that a kind of acoustical signal that will for example receive in the loudspeaker place is converted to for example processing of the language element of phoneme, word and sentence and so on.Speech recognition is useful to many functions, and these functions comprise and convert spoken language the dictation of penman text to and utilize verbal order to come the computer control that software application is controlled.
The application of the speech recognition technology that further emerges is the virtual role (avatars) that control computer generates.According to Hindoo mythology, virtual role (avatars) is the incarnation that plays with the god of the such effect of human intermediary.In the virtual world of electronic communication, virtual role is " two dimension " as the cartoon or " three-dimensional " diagrammatic representation of people or all kinds of biologies.As one " speaking head ", virtual role can allow the electronic communication of voice call for example or Email and so on become lively by represent the virtual image of communicating by letter to send to the recipient.For example, utilize speech synthesis technique and by virtual role with the text of Email " say to " recipient.In addition, utilize the virtual role of speaking, can will only the sound data be transformed the conference call that is as the criterion from the traditional telephony call that the calling party is sent to the callee.For the participator, this accurate conference call still need be transmitted the bandwidth of much less than actual video data than traditional more interesting and more abundant information of the Conference Calling that audio frequency is only arranged.
Use the accurate video conference of virtual role to adopt speech recognition technology to identify the language element in the received voice data.For example, shown virtual role can make calling party's speech animation in real time on the mobile phone screen.When calling party's speech spreads by telelecture, the language element in speech recognition software in phone sign calling party's the speech, and this language element is mapped to the figured variation of the mouth of virtual role.Utilize calling party's speech in real time, virtual role (avatars) therefore seems similarly to be the telephone subscriber who is talking.
Summary of the invention
According to an aspect, the present invention is a kind of method that is used for voice data classification.This method comprises sound (voiced) voice segments of logos sound data.Then, by being carried out spectrum analysis, the high amplitude composition of this speech sound section determines the high amplitude frequency spectrum.Afterwards this high amplitude frequency spectrum is categorized as vowel phoneme, wherein this vowel phoneme is to choose from the vowel set of simplifying.
Therefore, use the present invention, utilize the real-time voice data can improve animation virtual role.Method of the present invention is compared the computational intensity that has still less with most conventional speech recognition methods, and this can make that method of the present invention is carried out quickly, uses less processor resource simultaneously.
Description of drawings
In order can easily to understand the present invention and the present invention to be dropped into actual use, with reference now to by illustrated with reference to the accompanying drawings exemplary embodiment, wherein in each figure similar reference number represent identical or function on similar element.According to the present invention, accompanying drawing is included in the instructions with following detailed description and has constituted the part of instructions, and is used for further specifying embodiment and explains various principle and advantages, wherein:
Fig. 1 has provided the synoptic diagram with the mobile device of wireless telephone form that is used to carry out method of the present invention;
Fig. 2 is the chart and relevant spectrogram that has illustrated according to the speech data that for example receives in the mobile device place and handle of the embodiment of the invention;
Fig. 3 has provided the module map according to the functional part of the phonetic classification of the embodiment of the invention and the processing of mouth Motion mapping; And
Fig. 4 is the general flow figure to the method for voice data classification that has illustrated according to the embodiment of the invention.
Those skilled in the art it should be understood that for simple and clear for the purpose of, the element in the accompanying drawing is illustrated, and it is not necessarily drawn in proportion.For example, some size of component have been exaggerated for other elements in the accompanying drawing, to help to improve the understanding to embodiments of the invention.
Embodiment
Before describing in detail, should be noted that this embodiment mainly is and is used for the relevant method step of the method for voice data classification and the combination of device feature according to embodiments of the invention.Therefore, in the accompanying drawings in suitable place with habitual symbolic representation device feature and method step, wherein only show those specific detail relevant, so that can be because of not blured the disclosure for benefiting from the details that this those of ordinary skills of instructions, is readily understood that with understanding embodiments of the invention.
In this article, only be used to distinguish an entity or action and another entity or action, and not necessarily require or hint the relation of this any reality between this entity or the action or in proper order such as a left side and right, such relational terms such as first and second.Term " comprises ", " comprise " or its any other variation is used to cover comprising of nonexcludability, thereby make processing, method, goods or the device comprise a series of key elements not only comprise these key elements, but also comprise that other are not clearly listed or for this processing, method, goods or install intrinsic key element.The key element that is limited by " comprising one ... " is not precluded within the situation that is not subjected to more restrictions in processing, method, goods or the device and also has other identical key elements.
Referring to Fig. 1, this schematic view illustrating be used to carry out method of the present invention, with the mobile device of the form of wireless telephone 100.This phone 100 comprises the radio frequency communications unit 102 that is coupled into row communication with processor 103.This phone 100 also has keypad 106 and the display screen 105 that is coupled into row communication with processor 103.It will be apparent to one skilled in the art that screen 105 can be that thereby to make keypad 106 be selectable to touch-screen.
This processor 103 comprises the code ROM (read-only memory) (ROM) 112 that following data are stored in encoder/decoder 111 and relevant being used to, and can carry out Code And Decode by speech or other signal that wireless telephone 100 sends or receives when described data are used for.Processor 103 also comprise by common data and address bus 117 and with microprocessor 113, character ROM (read-only memory) (ROM) 114, random-access memory (ram) 104, static programmable memory 116 and the SIM interface 118 of encoder/decoder 111 couplings.Static programmable memory 116 and operationally with the SIM of SIM interface 118 coupling in each all especially can store selected arrival text message and telephone number database TND (telephone directory), the name field that this telephone number database TND comprises the number field that is used for telephone number and is used for the identifier that one of the number with name field is associated.For example, clauses and subclauses among the telephone number database TND may be in name field, have relevant identifier " Steven C! At work " 91999111111 (being input in the number field).
Microprocessor 113 has and is used for the port that is coupled with keypad 106 and screen 105 and warning horn 115, and wherein warning horn 115 generally comprises alert speaker, vibrating motor and relevant driver.In addition, microprocessor 113 has and is used for the port that is coupled with microphone 135 and communications speaker 140.Character ROM (read-only memory) 114 storage is used for the code to being decoded or be encoded by the text message that communication unit 102 is received.In this embodiment, this character ROM (read-only memory) 114 is gone back the operation code (OC) of storage microprocessor 113 and the code that is used to carry out the function that is associated with wireless telephone 100.
Radio frequency communications unit 102 is receiver and the reflection machines with combination of community antenna 107.This communication unit 102 has the transceiver 108 that is coupled with antenna 107 by radio frequency amplifier 109.This transceiver 108 also is coupled with the modulator/demodulator 110 that makes up, and this modulator/demodulator 110 is coupled communication unit 102 and processor 103.
Referring to Fig. 2, according to embodiments of the invention, chart 200 has been described the speech data that for example receives and handle in wireless telephone 100 places with relevant spectrogram 205-n.Chart 200 is depicted as sound amplitude with respect to the time with speech data.Those skilled in the art can be identified as sound (voiced) voice with three kinds of main peak value waveform envelope 210-n, and the interval of the relative short arc between the peak value waveform envelope 210-n is identified as noiseless (unvoiced) voice.
Traditional voice recognition processing has solved the such complicated technology problem of sign phoneme, and phoneme is the minimum vowel sound unit that is used to make speech.Speech recognition normally needs speech data is calculated the statistical treatment of intensive analysis.This analysis comprises: the sound variation the noise that identification such as ground unrest and sensor cause; And the voice changeability the sound difference of identification in single phoneme.
According to an embodiment, the present invention is that a kind of computational intensity is wanted significantly the method less than conventional speech recognition methods, this method be used for voice data classification so that the animation of the mouth feature of virtual role (avatars) more credible just look at truer.For example, virtual role may be displayed on the screen 105 of phone 100, and looks similarly to be those language that received and amplified on communications speaker 140 by transceiver 108 of saying the calling party in real time.To describe a kind of like this method below in detail.
At first, by the speech sound section of logos sound data for example speech data described in Figure 200 is carried out filtering.The sign of speech sound section when carrying out by the known in the art various technology of use such as energy spectrometer and zero-crossing rate analysis.The high-energy component of speech data is relevant with sound sound usually, and low-yield speech data is relevant with noiseless sound usually to middle energy speech data simultaneously.The extremely low-yield composition of speech data is common and silent or ground unrest is relevant.
Zero-crossing rate is the simple measurement that the frequency content to speech data carries out.The low-frequency component of speech data is relevant with sound voice usually, and the radio-frequency component of speech data is relevant with noiseless voice usually.
After the sound voice segments of sign, for each section is determined a high amplitude frequency spectrum.Therefore, for each section, the fast Fourier transform (FFT) standardization of the high amplitude composition by making each speech sound section according to amplitude, the fast Fourier transform (FFT) data of settling the standard.For example, in chart 200, be " key frame " with the high amplitude component identification of each section, should " key frame " comprise the peak amplitude within each section.This key frame generally has regular time window (approximately 30ms), and the number of times of sampling can change according to the sampling rate of this speech data.For example, typical key frame can comprise the sampling with length L=256 of 8kHz sampling rate, or with the sampling of the L=512 of 16kHz sampling rate.
After this standardized FFT data are carried out filtering, so that make the peak value in these data more obvious.For example, can use that to have threshold setting be 0.1 Hi-pass filter, the value that all in its FFT data are lower than threshold setting is set to zero.
Then by one or more peak detctors handle by standardization and filtering the FFT data.This peak detctor is to detecting such as the so various peak value attributes of peak value number, peak Distribution and peak energy.After this be used to data from peak detctor, may represent main vowel sound the high amplitude frequency spectrum by standardization and filtering the FFT data be divided into sub-band.For example, according to one embodiment of present invention, use to be indexed as four sub-bands of from 0 to 3.If the concentration of energy of high amplitude frequency spectrum is in sub-band 1 or 2, this frequency spectrum probably is classified as corresponding with pivot sound phoneme/a/ so.If the concentration of energy of this high amplitude frequency spectrum is in sub-band 0 and 2, then this frequency spectrum probably is classified as corresponding with pivot sound phoneme/i/.At last, if the concentration of energy of this high amplitude frequency spectrum in sub-band 0, then this frequency spectrum probably is classified as corresponding with pivot sound phoneme/u/.Fig. 2 has described and peak value waveform envelope 210-1 and the corresponding standardization frequency spectrum of pivot sound phoneme/a/ 205-1; With peak value waveform envelope 210-2 and the corresponding standardization frequency spectrum of pivot sound phoneme/i/ 205-2; And with peak value waveform envelope 210-3 and the corresponding standardization frequency spectrum of pivot sound phoneme/u/ 205-3.
According to one embodiment of present invention, utilize classified frequency spectrum to make the feature animation of virtual role (avatars), in fact " saying " the such impression of speech data so that create this virtual role.This animation moves and carries out by classified frequency spectrum being mapped to discontinuous mouth.As last field is known, use a series of pronunciation mouth shapes (viseme) that discontinuous mouth motion is repeated by virtual role, described a series of pronunciation mouth shapes (viseme) come down to be mapped to the basic phonetic unit in the visual city.Each pronunciation mouth shape (viseme) expression mouth shape static state, formation contrast visually, it is corresponding with employed mouth shape when a people sends particular phoneme usually.
The present invention can carry out the mapping of this phoneme to pronunciation mouth shape effectively by using the following fact, and the described fact is exactly that phoneme number in the language is much larger than the number of corresponding pronunciation mouth shape.Further, with above-mentioned pivot sound phoneme/a/ ,/i/ and/each of u/ is mapped to one of three diverse pronunciation mouth shapes.---its with open and and then relevant to the picture frame of the mouth of make-position motion from being closed into---as cartoon, can create believable mouth motion by only using these three different pronunciation mouth shapes.Owing in speech data, only discern three main vowel phonemes, so the speech recognition in the embodiments of the invention obviously has the less processor closeness than speech recognition of the prior art.For example, according to one embodiment of present invention, utilize as following table 1 as shown in/a/ ,/i/ and/three main vowel phonemes of u/, the vowel of various vowel phonemes grouping the becoming simplification in the English is gathered.
The vowel set of the simplification in the table 1-English
| /a/ | ax,aa,ae,ao,aw,er,ay,eh,ey |
| /i/ | ih,iy |
| /u/ | ow,oy,uh,uw |
Utilization is such as carrying out mouth width mapping or carry out technology the mapping of mouth shape according to the spectrum structure of speech data according to speech energy, the speech data that religious doctrine according to the present invention is classified can be used to control the action of the mouth and the lip figure of virtual role (avatars).For example, the mapping of mouth width is relevant with the opening and closing of mouth during peak value waveform envelope 210-n.Consideration will be numbered as from 0 to i-1 i image and be used to describe peak value waveform envelope 210-n.The mouth width mapping at first unvoiced segments of the beginning of this peak value waveform envelope 210-n is set to zero, thus the closed mouth of expression.According to the speech energy in each respective frame, will in peak value waveform envelope 210-n, be mapped to this image 1 to i-1 by remaining Frame then.At last, the action of the mouth of the virtual role that perceives in order to make and lip looks more natural, carries out aftertreatment to mouth and lip figure so that the level and smooth conversion between the image to be provided.
Referring to Fig. 3, schematic block diagram 300 has been described the functional part according to the phonetic classification of the embodiment of the invention and the processing of mouth Motion mapping.Processing can be divided into three major function pieces: the synthetic piece 315 of key frame home block 305, the vowel sets classification piece of simplifying 310 and animation.
In key frame home block 305, in energy spectrometer piece 320 and zero-crossing rate piece 325, receive and handle the speech data of input concurrently.To offer from the data of energy spectrometer piece 320 and zero-crossing rate piece 325 and be used to separate sound and sound/no sound detection pieces 330 unvoiced speech the data.Then, to handling from the data of sound/no sound detection piece 330, this sound envelope generator piece 335 is used for the speech sound section of logos sound data in sound envelope generator piece 335.As shown in the figure, in above-mentioned key frame home block 305, in sound envelope generator piece 335, also use raw data from energy spectrometer piece 320.
In the vowel sets classification piece of simplifying 310, provide sound voice segments to key frame spectrum analysis piece 340, the high amplitude composition of 340 pairs of speech sound sections of this key frame spectrum analysis piece is carried out spectrum analysis and is determined the high amplitude frequency spectrum.Next, in classification block 345, with the high amplitude frequency spectrum be categorized as pivot sound phoneme/a/ ,/i/ or/u/.
At last, in the synthetic piece 315 of animation, will be mapped to the pronunciation mouth shape (visemes) that is used to make the virtual role animation from the pivot sound phoneme of classification block 350 outputs.The synthetic piece 315 of this animation retrieve such information from cartoon material database 355, this information comprise that the mouth shape of for example pronouncing defines and with the relevant information in employed traditional mouth opening and closing position when between phoneme, changing.Therefore the last output of the synthetic piece 315 of this animation be the animation with voice synchronous.
Referring to Fig. 4, general flow figure has described the method 400 to voice data classification of being used for according to the embodiment of the invention.At first, in step 405, receive speech data at the mobile radio communication apparatus place of for example wireless telephone 100 and so on.In step 410, identify the speech sound section of this speech data.Next, in step 415,, the high amplitude spectrum component of speech sound section determines the high amplitude frequency spectrum by being carried out spectrum analysis.Step 415 can comprise following substep: in step 420, standardized FFT data are determined in the FFT standardization of the high amplitude composition by making the speech sound section according to amplitude.In step 425, standardized FFT data are carried out filtering, with create standardization and filtering the FFT data.Afterwards, in step 430, detect standardization and filtering data in peak value.In step 435, based on the attribute of the institute's detection peak such as peak value sum, peak Distribution and peak energy with standardization and filtering the FFT data qualification be sub-band.Then, in step 440, with the high amplitude frequency spectrum be categorized as for example above-mentioned pivot phoneme/a/ ,/i/ or/the such vowel phoneme of one of u/.At last, in step 445, main vowel phoneme is mapped to the pronunciation mouth shape that for example is associated with the mouth position of virtual role.
Therefore advantage of the present invention comprises the animation of utilizing real-time speech data and having improved virtual role.In addition, method of the present invention is compared the computational intensity that has still less with most conventional speech recognition methods, and this makes method of the present invention to be carried out quickly, uses less processor resource simultaneously.Therefore, embodiments of the invention are particularly suitable for having the limited processor and the mobile communication equipment of memory resource.
Above detailed description only provides exemplary embodiment, and does not mean that and define scope of the present invention, usability and structure.On the contrary, above-mentioned detailed description of illustrative embodiments provides the description that allows to realize exemplary embodiment of the present invention to those skilled in the art.Should be understood that, under the situation that does not break away from the spirit and scope of the present invention illustrated in the claims, can carry out various changes the function and the setting of element and step.It should be understood that, embodiments of the invention as described herein can comprise one or more traditional processors and unique institute's program stored instruction, described programmed instruction is used for one or more processors are controlled, to realize as described herein to some functions of voice data classification, most of function or all functions in conjunction with some non-processor circuit.This non-processor circuit can be including but not limited to radio receiver, radio reflections machine, signal driver, clock circuit, power circuit and user input apparatus.Similarly, can be the method step that is used for voice data classification with these functional interpretations.Perhaps, part or all function can be realized by the state machine that does not have the program stored instruction, perhaps use one or more special ICs (ASIC) to realize that some combination of each function or some function can be used as the logic realization of customization in described special IC.Certainly, can use the combination of two kinds of methods.Therefore, the method and apparatus that is used for these functions is described here.In addition, what can reckon with is, though owing to pot life for example, current technology and the consideration of economic aspect have inspired possible remarkable result and many design alternatives, but those of ordinary skill is when being subjected to the instructing of notion disclosed herein and principle, and can utilize minimum test and produces this software instruction and program and IC at an easy rate.
In above-mentioned instructions, specific embodiments of the invention have been described.Yet those of ordinary skills should be appreciated that under the situation that does not break away from the spirit and scope of the present invention illustrated in the claim can make various modifications and variations.Therefore, instructions and accompanying drawing are considered to illustrative and not restrictive, and intention is included within the scope of the present invention all this modifications.The solution of benefit, advantage, problem and can produce any benefit, advantage or solution or make its significant more any key element that becomes should not be considered to key, essential or basic feature or the key element that any one claim or all authority require.The present invention is only defined by claims, comprising any modification made during the pending trial of this application and the equivalent of these claims.
Claims (10)
1, a kind of method to voice data classification comprises:
The speech sound section of logos sound data;
Carry out spectrum analysis by high amplitude composition, determine the high amplitude frequency spectrum this speech sound section; With
This high amplitude frequency spectrum is categorized as vowel phoneme, and wherein this vowel phoneme is to choose from the vowel set of simplifying.
2, according to the method in the claim 1, wherein the vowel of this simplification set only comprise pivot sound phoneme/a/ ,/i/ and/u/.
3, according to the method in the claim 2, wherein pivot sound phoneme/a/ comprise from by/ax/ ,/aa/ ,/ae/ ,/ao/ ,/aw/ ,/er/ ,/ay/ ,/eh/ and/english phoneme selected the group that ey/ constitutes; Pivot sound phoneme/i/ comprise from by/ih/ and/english phoneme selected the group that iy/ constitutes; And pivot sound phoneme/u/ comprise from by/ow/ ,/oy/ ,/uh/ and/english phoneme selected the group that uw/ constitutes.
4, according to the method in the claim 1, further comprise: vowel phoneme is mapped to pronunciation mouth shape so that the virtual role animation.
5, according to the method in the claim 1, further comprise: receive speech data at the mobile radio communication apparatus place.
6,, wherein when receiving speech data, in real time the high amplitude composition of speech sound section is classified according to the method in the claim 5.
7,, determine that wherein the step of high amplitude frequency spectrum comprises according to the method in the claim 1:
The fast fourier transform of the high amplitude composition by making the speech sound section according to amplitude, be the FFT standardization, determine normalized fast fourier transform, be the FFT data;
Normalized FFT data are carried out filtering, with create standardization and filtering the FFT data;
Detect standardization and filtering the FFT data in peak value; With
Based on detected peak value and will be standardization and filtering the FFT data qualification be sub-band.
8, according to the method in the claim 7, wherein detect standardization and filtering the FFT data in the step of peak value comprise: the number to peak value is counted, is measured peak Distribution and measures peak energy.
9, according to the method in the claim 7, wherein sub-band is indexed as from 0 to 3, and the vowel of simplifying set comprises following pivot sound phoneme:
/ a/, wherein the concentration of energy of high amplitude frequency spectrum is in sub-band 1 or 2;
/ i/, wherein the concentration of energy of high amplitude frequency spectrum is in sub-band 0 and 2; With
/ u/, wherein the concentration of energy of high amplitude frequency spectrum is in sub-band 0.
10, according to the method in the claim 1, wherein the high amplitude composition of speech sound section comprises the key frame of speech data.
Priority Applications (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CNA2005101217187A CN1991981A (en) | 2005-12-29 | 2005-12-29 | Method for voice data classification |
| PCT/US2006/062032 WO2007076279A2 (en) | 2005-12-29 | 2006-12-13 | Method for classifying speech data |
Applications Claiming Priority (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CNA2005101217187A CN1991981A (en) | 2005-12-29 | 2005-12-29 | Method for voice data classification |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| CN1991981A true CN1991981A (en) | 2007-07-04 |
Family
ID=38214193
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| CNA2005101217187A Pending CN1991981A (en) | 2005-12-29 | 2005-12-29 | Method for voice data classification |
Country Status (2)
| Country | Link |
|---|---|
| CN (1) | CN1991981A (en) |
| WO (1) | WO2007076279A2 (en) |
Cited By (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN109087629A (en) * | 2018-08-24 | 2018-12-25 | 苏州玩友时代科技股份有限公司 | A kind of mouth shape cartoon implementation method and device based on speech recognition |
| CN111326143A (en) * | 2020-02-28 | 2020-06-23 | 科大讯飞股份有限公司 | Voice processing method, device, equipment and storage medium |
Families Citing this family (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| GB2468140A (en) * | 2009-02-26 | 2010-09-01 | Dublin Inst Of Technology | A character animation tool which associates stress values with the locations of vowels |
| WO2016154800A1 (en) * | 2015-03-27 | 2016-10-06 | Intel Corporation | Avatar facial expression and/or speech driven animations |
| US11176960B2 (en) * | 2018-06-18 | 2021-11-16 | University Of Florida Research Foundation, Incorporated | Method and apparatus for differentiating between human and electronic speaker for voice interface security |
Family Cites Families (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US5884267A (en) * | 1997-02-24 | 1999-03-16 | Digital Equipment Corporation | Automated speech alignment for image synthesis |
| US6909453B2 (en) * | 2001-12-20 | 2005-06-21 | Matsushita Electric Industrial Co., Ltd. | Virtual television phone apparatus |
-
2005
- 2005-12-29 CN CNA2005101217187A patent/CN1991981A/en active Pending
-
2006
- 2006-12-13 WO PCT/US2006/062032 patent/WO2007076279A2/en not_active Ceased
Cited By (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN109087629A (en) * | 2018-08-24 | 2018-12-25 | 苏州玩友时代科技股份有限公司 | A kind of mouth shape cartoon implementation method and device based on speech recognition |
| CN111326143A (en) * | 2020-02-28 | 2020-06-23 | 科大讯飞股份有限公司 | Voice processing method, device, equipment and storage medium |
| CN111326143B (en) * | 2020-02-28 | 2022-09-06 | 科大讯飞股份有限公司 | Voice processing method, device, equipment and storage medium |
Also Published As
| Publication number | Publication date |
|---|---|
| WO2007076279A2 (en) | 2007-07-05 |
| WO2007076279A3 (en) | 2008-04-24 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| CN1991982A (en) | Method of activating image by using voice data | |
| Gabbay et al. | Visual speech enhancement | |
| US8560307B2 (en) | Systems, methods, and apparatus for context suppression using receivers | |
| EP4531037A2 (en) | End-to-end speech conversion | |
| CN110634507A (en) | Speech classification of audio for voice wakeup | |
| US20110282650A1 (en) | Automatic normalization of spoken syllable duration | |
| CN1742321B (en) | Rhythmic imitation synthesis method and device | |
| JP2007501444A (en) | Speech recognition method using signal-to-noise ratio | |
| JP2005202854A (en) | Image processor, image processing method and image processing program | |
| CN106157957A (en) | Audio recognition method, device and subscriber equipment | |
| CN114495907B (en) | Adaptive voice activity detection method, device, equipment and storage medium | |
| CN111696580A (en) | Voice detection method and device, electronic equipment and storage medium | |
| JPH10293860A (en) | Method and apparatus for displaying human image using voice drive | |
| EP1908053B1 (en) | Speech analysis system | |
| CN1991981A (en) | Method for voice data classification | |
| CN102857650B (en) | Method for dynamically regulating voice | |
| KR20080008432A (en) | Lip sync synchronization method and apparatus for voice signal | |
| KR100399057B1 (en) | Apparatus for Voice Activity Detection in Mobile Communication System and Method Thereof | |
| CN107785020B (en) | Voice recognition processing method and device | |
| CN106899625A (en) | A kind of method and device according to user mood state adjusting device environment configuration information | |
| KR101095867B1 (en) | Speech Synthesis Device and Method | |
| AU2021107566A4 (en) | Mobile device with whisper function | |
| CN114783406B (en) | Speech synthesis method, apparatus and computer-readable storage medium | |
| JP2011158515A (en) | Device and method for recognizing speech | |
| CN120600018A (en) | VAD adaptive terminal wake-up method and system based on user voice feature recognition |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| C06 | Publication | ||
| PB01 | Publication | ||
| C10 | Entry into substantive examination | ||
| SE01 | Entry into force of request for substantive examination | ||
| C02 | Deemed withdrawal of patent application after publication (patent law 2001) | ||
| WD01 | Invention patent application deemed withdrawn after publication |