WO2023210149A1 - 情報処理装置及び情報処理方法、並びにコンピュータプログラム - Google Patents
情報処理装置及び情報処理方法、並びにコンピュータプログラム Download PDFInfo
- Publication number
- WO2023210149A1 WO2023210149A1 PCT/JP2023/007479 JP2023007479W WO2023210149A1 WO 2023210149 A1 WO2023210149 A1 WO 2023210149A1 JP 2023007479 W JP2023007479 W JP 2023007479W WO 2023210149 A1 WO2023210149 A1 WO 2023210149A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- voice
- normal
- whisper
- unit
- recognition
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/04—Architecture, e.g. interconnection topology
- G06N3/045—Combinations of networks
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/08—Learning methods
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L25/00—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00
- G10L25/48—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 specially adapted for particular use
- G10L25/51—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 specially adapted for particular use for comparison or discrimination
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L15/00—Speech recognition
- G10L15/28—Constructional details of speech recognition systems
- G10L15/32—Multiple recognisers used in sequence or in parallel; Score combination systems therefor, e.g. voting systems
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L17/00—Speaker identification or verification techniques
- G10L17/26—Recognition of special voice characteristics, e.g. for use in lie detectors; Recognition of animal voices
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L25/00—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00
- G10L25/27—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 characterised by the analysis technique
- G10L25/30—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 characterised by the analysis technique using neural networks
Definitions
- this disclosure relates to an information processing device and an information processing method that process audio, and a computer program.
- Speech recognition is widely used. Speech recognition is also used for transcription, which converts utterances into text, but the original utterances do not always consist of only characters; for example, they can include symbols such as punctuation marks and quotation marks, and editing commands for input text. (or when the speaker wants to instruct symbol input or editing commands).
- speech recognition technology it is difficult to distinguish between text input, symbol input, and editing commands. For example, if you say ⁇ mountain'' and input it as ⁇ Kagikkko Yama Kakikakko Jiru,'' it may be input as is. Commands such as "delete one line" are also entered as text.
- a sentence-final symbol assignment norm corresponding to the situation in which the dialogue took place is selected from among multiple sentence-final symbol assignment norms based on the dialogue situation characteristics, and the selected sentence-final symbol assignment norm, the acoustic characteristics of the dialogue, and the language are selected.
- a device has been proposed that uses features to estimate the end-of-sentence symbol for the text resulting from speech recognition of the dialogue (see Patent Document 1); however, this device can only add punctuation marks to the speech-recognized text. You cannot enter symbols such as square brackets into the text.
- An object of the present disclosure is to provide an information processing device, an information processing method, and a computer program that perform processing related to voice input.
- the present disclosure has been made in consideration of the above problems, and the first aspect thereof is: a classification unit that classifies the spoken voice into normal voice and whispered voice based on voice features; a recognition unit that recognizes the whisper classified by the classification unit; a control unit that controls processing based on the recognition result of the recognition unit;
- This is an information processing device comprising:
- the classification unit is configured to classify normal speech and whispering using a first trained neural network
- the recognition unit is configured to recognize whispering using a second trained neural network.
- the second trained neural network includes a feature extraction layer and a transformer layer
- the first trained neural network is configured to share the feature extraction layer with the second trained neural network.
- the information processing device further includes a normal speech recognition unit that recognizes the normal speech classified by the classification unit, and the control unit is configured to respond to the recognition result of the normal speech recognition unit.
- the device is configured to perform processing corresponding to the recognition result of the whispering voice by the recognition unit.
- the control section executes processing of the whisper command recognized by the recognition section on the text into which the normal speech recognition section has converted normal speech.
- a second aspect of the present disclosure is: a classification step of classifying the spoken voice into normal voice and whispered voice based on voice features; a recognition step of recognizing the whisper classified in the classification step; a control step for controlling processing based on the recognition result in the recognition step;
- This is an information processing method having the following.
- a third aspect of the present disclosure is: a classification unit that classifies spoken voice into normal voice and whispered voice based on voice features; a recognition unit that recognizes the whisper classified by the classification unit; a control unit that controls processing based on the recognition result of the recognition unit; A computer program written in computer-readable form to cause a computer to function as a computer program.
- a computer program according to the third aspect of the present disclosure defines a computer program written in a computer readable format so as to implement predetermined processing on a computer.
- a cooperative effect is exerted on the computer, and the same effect as that of the information processing device according to the first aspect of the present disclosure is achieved. effect can be obtained.
- FIG. 1 is a diagram showing the basic configuration of a voice input system 100 to which the present disclosure is applied.
- FIG. 2 is a diagram showing spectrograms for each frequency band of normal voice and whispering voice.
- FIG. 3 is a diagram showing the operation of the voice input system 100.
- FIG. 4 is a diagram showing the overall structure of a system that performs whisper voice recognition using a neural network.
- FIG. 5 is a diagram showing the results of classification samples of whispering voice and normal voice.
- FIG. 6 is a diagram showing a two-dimensional representation of the situation in which feature vectors are processed within the classification unit 101, compressed in dimension based on the t-SNE algorithm.
- FIG. 7 is a diagram showing another example of the neural network configuration of the classification unit 101.
- FIG. 1 is a diagram showing the basic configuration of a voice input system 100 to which the present disclosure is applied.
- FIG. 2 is a diagram showing spectrograms for each frequency band of normal voice and whispering voice.
- FIG. 3 is
- FIG. 8 is a diagram illustrating a specific operational example of a voice input process using the voice input system 100 according to the present disclosure.
- FIG. 9 is a diagram illustrating a specific operational example of a voice input process using the voice input system 100 according to the present disclosure.
- FIG. 10 is a diagram illustrating a specific operational example of a voice input process using the voice input system 100 according to the present disclosure.
- FIG. 11 is a diagram illustrating a specific operational example of a voice input process using the voice input system 100 according to the present disclosure.
- FIG. 12 is a diagram illustrating a specific operational example of a voice input process using the voice input system 100 according to the present disclosure.
- FIG. 13 is a diagram illustrating a specific operational example of a voice input process using the voice input system 100 according to the present disclosure.
- FIG. 14 is a diagram illustrating a specific operational example of a voice input process using the voice input system 100 according to the present disclosure.
- FIG. 15 is a diagram illustrating a specific operational example of a voice input process using the voice input system 100 according to the present disclosure.
- FIG. 16 is a diagram illustrating an operation example when the present disclosure is applied to a remote conference.
- FIG. 17 is a diagram showing an example of the functional configuration of conference terminal 1700.
- FIG. 18 is a diagram showing an example of the functional configuration of the avatar control system 1800.
- FIG. 19 is a diagram showing an example of the functional configuration of a mask-type wearable interface 1900.
- FIG. 20 is a diagram showing a configuration example of the information processing device 2000.
- ASR Automatic speech recognition
- voice recognition has the problem of recognition errors. Misrecognitions caused by voice are difficult to correct using spoken commands, and recognition errors must be corrected using a keyboard or mouse, thereby eliminating the benefits of hands-free operation.
- voice when entering text using voice, problems arise when entering special characters or commands by voice, such as symbols outside of regular text. For example, when entering "?", you can say “question mark”, but the text “question mark” itself may be entered instead of the symbol "?”. If you want to start a new paragraph and say “new paragraph,” it may be recognized literally and entered as the text "new paragraph.” Similar problems can occur when entering commands commonly used in text editing, such as "delete word” or "new paragraph.”
- the voice spoken by the user is a mixture of text input and spoken commands, but in conventional automatic speech recognition, there is no easy way to distinguish between commands and text input from voice.
- text can be input, and commands can also be input by pressing modifier keys such as the Shift key, Control key, and Alt key at the same time as alphanumeric keys.
- modifier keys such as the Shift key, Control key, and Alt key at the same time as alphanumeric keys.
- the modality during voice input may be switched by inputting a modifier key, but the advantage of hands-free operation is lost.
- the uttered voice is classified into normal voice and whispered voice based on voice features, the classified whispered voice is recognized, and based on the recognition result, the normal voice of the original uttered voice is classified into normal voice and whispered voice.
- This paper proposes an information processing device that controls processing related to the audio portion of a computer.
- the information processing device according to the present disclosure can realize voice input through completely hands-free operation without requiring any special hardware other than a normal microphone.
- text input is performed using the normal speech part of the input speech, and whispering is classified from the input speech, and the recognition results of the whispering are divided into periods, commas, It can be used for inputting symbols such as quotation marks or special characters, and for text editing commands such as deleting text and line breaks. That is, by applying the present disclosure to voice recognition, text can be input using normal voice, while non-text information such as commands can be input using whispered voice.
- FIG. 1 schematically shows the basic configuration of a voice input system 100 to which the present disclosure is applied.
- the voice input system 100 is a system that gives multiple meanings to utterances using whispered voice and normal voice.
- the voice input system 100 can be configured using a general information processing device such as a personal computer (PC), and does not require any special input equipment other than a microphone for voice input.
- PC personal computer
- a typical information processing device is equipped with an input device such as a keyboard or a mouse for a user to perform an input operation, but in performing voice input according to the present disclosure, there is no need to use an input device other than a microphone.
- the voice input system 100 shown in FIG. 1 includes a classification section 101, a whispered voice recognition section 102, a normal voice recognition section 103, and a control section 104.
- the voice spoken by the user is input into the voice input system using a microphone (or headset).
- the classification unit 101 classifies input speech into normal speech and whispering based on the speech feature amount.
- the input audio includes a normal voice section, a whisper section, and a silent section. Therefore, the classification unit 101 attaches marks for identifying the normal voice section and the whispering voice section to the input audio signal, and sends the marks to each of the whisper voice recognition section 102 and the normal voice recognition section 103 in the subsequent stage. Output.
- the normal speech recognition unit 103 performs normal speech recognition processing using only the normal speech sections of the time-series audio signal input from the classification unit 101, and converts it into text.
- the whisper recognition unit 102 executes whisper recognition processing using only the whisper section of the audio signal input from the classification unit 101, and recognizes symbols, special characters, and editing commands for text input by normal voice. Convert to non-text information (such as text deletion or line breaks).
- control unit 104 performs processing on the text output from the normal speech recognition unit 102 based on the non-text information output from the whispered voice recognition unit 103.
- the control unit 104 can respond to text input as normal voice by adding symbols or special characters input in a whispered voice to corresponding locations in the text. Perform input, select text conversion candidates specified in a whisper, and execute commands entered in a whisper, such as deletion or line break.
- voice input system 100 can also be applied to potential applications other than text input, but the details will be discussed later.
- Two new neural networks are used to realize the voice input system 100 to which the present disclosure is applied.
- One is a neural network used in the classification unit 101 to distinguish between normal voice and whispering voice.
- the other is a neural network used in the whisper recognition unit 102 to recognize whispers. Details of each neural network will be given in the subsequent section D.
- the present disclosure can be broadly utilized in a variety of situations where speech recognition is already available.
- One example implementation of the present disclosure uses existing speech recognition systems (e.g., Google's Cloud Speech-to-Text) to perform regular speech recognition, and uses a customized neural network to perform whispered voice recognition. Recognize.
- the former conventional speech recognition, is trained on large corpora, is speaker independent, and can be used without special training.
- the whispered voice recognition part has no choice but to be trained on a smaller corpus, as there is no existing large-scale corpus, and it is difficult for individual users to whisper example sentences for speaker adaptation (i.e. customization). Requires required training stages.
- Whispered voice is a mode of speech that many people can pronounce without special training, and can be used freely between normal voice and whispered voice while speaking. This is one of the reasons why we use whisper voice as an alternative voice input mode to normal voice in this disclosure.
- Speech modes in which a person can produce pronunciation include normal voice and whispering, as well as voice pitch switching.
- normal voice is the voice used when the user speaks normally
- whisper voice is the voice used when the user talks in private.
- the pronunciation methods are different between normal voice and whisper voice.
- the human vocal cords exhibit regular and periodic vibrations, and this vibration frequency is called the fundamental frequency.
- the vibrations of the vocal cords are not noticeable and exhibit irregular and random vibrations, that is, it can be said that there is no fundamental frequency. Therefore, even if you forcefully raise the volume of the whispering voice, it will not become the same as normal voice.
- FIGS. 2(A) and 2(B) show spectrograms for each frequency band of normal voice and whispering voice, respectively.
- This is a case in which the speaker utters the same phrase "A quick brown fox jumps over the lazy black dog.” in a normal voice and in a whisper.
- the spectrogram patterns of the two are different due to the difference in the biological phenomenon that the vocal cords vibrate regularly and periodically when speaking in a normal voice, but there is almost no vibration when speaking in a whisper. The difference is obvious.
- Pattern recognition is one of the main applications of neural networks, which can be used to accurately distinguish between the frequency spectrograms of normal speech and whispered voices in real time.
- One is a neural network used in the classification unit 101 to distinguish between normal voice and whispering voice.
- the other is a neural network used in the whisper recognition unit 102 to recognize whispers.
- FIG. 3 illustrates the operation of the voice input system 100 shown in FIG. 1 from the perspective of recognizing a frequency spectrogram using a neural network.
- the voice input system 100 receives the speaker's voice, which is a mixture of normal voice and whispered voice, and is indicated by reference number 301.
- the classification unit 101 using the first neural network classifies whether the input voice is a normal voice or a whisper in a predetermined time unit (for example, every 100 milliseconds).
- the whisper recognition unit 102 uses the second neural network to perform recognition processing on the section of the input voice classified as a whisper, indicated by reference number 302. Further, the normal speech recognition unit 103 performs normal speech recognition processing on a section of the input speech classified as normal speech, indicated by reference number 303, using the third neural network.
- the normal speech recognition unit 103 converts speech in a normal speech section into text.
- the whisper recognition unit 102 recognizes the whisper voice section and converts it into non-text information such as symbols, special characters, and editing commands (text deletion, line break, etc.) for text input in normal voice.
- the control unit 104 performs processing on the text output from the normal speech recognition unit 102 based on the non-text information output from the whispered voice recognition unit 103.
- the voice input system 100 can give multiple meanings to utterances using whispered voices and normal voices. For example, saying “line break” in a normal voice means entering the text, but saying “line break” in a whisper means invoking the line break command. Similarly, to correct a recognition error, the user can display alternative recognition candidates by whispering "candidate” and select them by whispering the number attached to the candidate ("1", "2", etc.). You can choose one.
- Conventional voice recognition methods can be compared to a keyboard that does not include command keys or symbol keys.
- Keyboards used to operate computers are provided with function keys, command keys, symbol keys, and other means for inputting commands as well as character input. You can call these functions without explicitly changing the interaction mode or taking your hands off the keyboard. That is, a computer keyboard allows text input and command input to coexist without any special mode conversion operation.
- the voice input system 100 the relationship between normal voice and whispering can be compared to the relationship between text input and command input.
- the user can use normal voice and whispered voice to coexist text input and command input with hands-free operation.
- This section D describes the configuration of a neural network that enables whisper recognition and classification in the voice input system 100 to which the present disclosure is applied.
- whisper recognition i.e., the whisper recognition unit 102 uses wav2vec2.0 (see Non-Patent Documents 1 and 2) or HuBERT (Non-Patent Document 2). 3). Both wav3ves2.0 and HuBERT are self-supervised neural networks designed for speech processing systems.
- Both wav2vec2.0 and HuBERT assume a combination of pre-training and self-supervised representation learning using unlabeled audio data and fine tuning using labeled audio data. These systems are primarily targeted at speech recognition applications, but have also been applied to speaker recognition, language recognition, and emotion recognition.
- FIG. 4 shows the overall structure of a system (whisper recognition neural network 400) that performs whisper recognition using a neural network.
- a neural network 400 configured using wav2vec2.0 or HuBERT can be roughly divided into a feature extraction layer 410 and a transformer layer 420.
- This method of pre-training the whispered voice recognition neural network 400 is similar to BERT's masked language model in natural language processing. It is designed to mask part of the input and estimate the corresponding representational features (feature vectors) from the remaining input. This pre-training should allow it to learn the acoustic properties and voice features of the input data.
- the whispered voice recognition neural network 400 shown in FIG. 4 can achieve speech recognition accuracy comparable to conventional state-of-the-art ASR by fine-tuning using only a small amount of labeled voice data set (Non-Patent Document 1 and Non-Patent Document 2). Therefore, this architecture is considered suitable for recognizing whispers under a limited whisper corpus.
- the feature extraction layer 410 converts the raw audio waveform into a latent feature vector.
- the feature extraction layer 410 is composed of a plurality of (for example, five) convolutional neural networks (CNNs). Section data obtained by dividing a voice waveform signal produced by a user into time sections of a predetermined length is input to each CNN. However, the data is divided into each section data so as to include overlapping areas in adjacent time sections.
- CNNs convolutional neural networks
- Each CNN is composed of a seven-block one-dimensional convolutional layer (Conv1D) and a temporal convolutional layer (GroupNorm), similar to the original wav2vec2 and HuBERT.
- Each block has 512 channels consisting of stride (5, 2, 2, 2, 2, 2, 2) and kernel width (10, 3, 3, 3, 2, 2).
- the feature extraction layer 410 is designed to output 512 latent dimension feature vectors every 20 milliseconds.
- the classification unit 101 distinguishes between whispering voices and normal voice input using an audio signal of a fixed length (eg, 100 milliseconds).
- the right side of FIG. 4 shows in detail the configuration of the feature extraction layer 410 of the whispered voice recognition neural network 400 made of wav2vec2.0 or HuBERT.
- the feature extraction layer 410 converts the acoustic signal into a 512-dimensional feature vector every 20 milliseconds.
- the classification unit 101 partially shares the neural network (that is, the feature extraction layer 410) with the whispered voice recognition neural network 400, so that the network size of the voice input system 100 as a whole is reduced.
- the classification unit 101 shares the portion surrounded by a broken line with the whisper recognition neural network 400 of the feature extraction layer 410 in the whisper recognition neural network 400 shown in the center of FIG.
- the right side of FIG. 4 shows an enlarged view of the inside of the classification section 101.
- the classification unit 101 applies a normalization layer (Layer Norm) 411, an average pooling layer (Avg Pool) 412, and two subsequent FC (fully connected) layers 413 and 414 to the feature vector extracted from the audio waveform signal.
- Layer Norm Layer Norm
- Avg Pool average pooling layer
- FC fully connected
- FIG. 5 shows the results of the classification samples of whispering and normal voices obtained by the classification unit 101 shown on the right side of FIG. 4 using a frequency spectrum.
- the audio waveform signal input to the audio input system 100 includes a whisper section, a normal voice section, and a silent section.
- the classification unit 101 identifies the whisper section from the frequency spectrogram of the input audio waveform signal and marks it as “Classified as whisper”, and also marks the normal speech section as "Classified as whisper”. It is identified and marked as "Classified as normal.”
- the classification unit 101 Based on the "Classified as normal” mark, the classification unit 101 generates an audio stream from which the normal voice is removed, as shown in FIG. . Furthermore, the classification unit 101 generates an audio stream from which whispers have been removed, as shown in FIG. .
- FIG. 6A shows the 512-dimensional feature vector classified as input to the second FC (fully connected) layer 414, together with the feature vector input to the first normalization layer, and t-SNE ( It shows the results of two-dimensional representation based on a visualization method using a dimension reduction algorithm called t-Distributed Stochastic Neighbor Embedding. Comparing FIG. 6A with the input to the normalization layer 411 and the second FC (fully connected) layer 414 shown in FIG. It can be seen that the feature vectors of utterances and whispers are well distinguished.
- FIG. 7 shows another configuration example of the classification unit 101 that is shared with the feature extraction layer 410 in the whispered voice recognition neural network 400.
- the illustrated classification unit 101 includes a normalization layer gMLP 701, an average pooling layer (Avg Pool) 702, an FC (fully connected) layer 703, and a multi-class classification output layer for the feature vector extracted from the audio waveform signal.
- gMLP multi-layer perceptron with gating
- gMLP multi-layer perceptron with gating
- Ru is a deep learning model announced by Google Brain, and is said to have performance comparable to a transformer by simply combining a multi-layer perceptron with a gate mechanism without using an attention mechanism. Ru.
- the whisper recognition neural network 400 is further fine-tuned using two types of whisper data sets, wTIMIT and per-user database.
- wTIMIT whispering voices
- Non-Patent Document 4 Each speaker follows the TIMIT prompt and says 450 phonetically balanced sentences in both normal speech and whispers. There were 29 speakers, with a good gender balance. The total number of utterances used was 11,324 (1,011 minutes).
- the corpus of wTIMIT includes normal speech, and can be used for training the classification unit 101 to classify whispering and normal speech (described later). The data is provided in two parts: training and testing. In this embodiment, the division between training and testing is used as is.
- Per-user database is audio data dubbed by each user in a whisper.
- the selected sequence of voice commands will be used as a script.
- These phrases are primarily intended to be commands used during text entry.
- Each user repeats each phrase five times, resulting in a total of 110 phrases being recorded.
- These phrases are further randomly concatenated and used as a data set.
- the total number of utterances after this connection is 936 (approximately 82 minutes).
- Fine-tuning is a learning method well known in the industry that trains a trained model to be tailored to each task. Specifically, in fine-tuning, a model is trained using unlabeled data, and then the model parameters are tuned using supervised data for a specific task to be solved.
- fine tuning of the whisper recognition neural network 400 is performed in two stages. In the first stage, we train on wTIMIT (whisper), and in the second stage we use the whisper command set (a dataset of whispers for each user) blown by the user.
- the audio data of normal speech and whispering included in wTIMIT is used (as described above, The wTIMIT corpus comes with regular voices as well as whispers).
- the length of the audio supplied to the classification unit 101 is set to 1,600 samples (100 millisecond size with 16 Kps audio sampling). This matches the length of audio chunks used in speech recognition cloud services at later stages.
- NAMELINES which means creating a new line
- the voice input system 100 treats it as a command and creates a new line for the text being input by voice.
- the voice input system 100 treats it as a command and deletes the last word, You can change the words you are interested in.
- the voice input system 100 can treat it as a command and present another recognition candidate.
- the candidates on the menu are labeled 1, 2, and 3, and users can select their favorite candidate from the menu by whispering ⁇ ONE'', ⁇ TWO'', ⁇ THREE'', etc.
- the voice input system 100 it is also possible to combine normal voice input and whispered voice commands. For example, if you whisper “SPELL” immediately after inputting the spelling "w a v 2 v e c" in normal voice, the word “ wav2vec2", which is quite difficult to input with normal ASR, will be generated. Similarly, it is also possible to input emoticons by whispering "EMOTION” immediately after saying "smile” in normal voice.
- the voice input system 100 will not treat it as a symbol or command. , convert it to text in the same way as traditional voice input text creation.
- FIGS. 8 to 15. 8 to 15 are text editing screens based on voice input.
- FIGS. 8 to 10 show an example of the operation when a command is used in a whispered voice while inputting text in a normal voice.
- the finalized part of the text input by normal voice is displayed in black text, and the unconfirmed part immediately after the voice input is displayed in gray text.
- the user whispers "MENU".
- the voice input system 100 treats the whispered voice "MENU" as a command.
- the classification unit 101 classifies the audio waveform corresponding to “MENU” as a whisper, and the whisper recognition unit 102 recognizes that the audio waveform is a command “MENU”.
- the voice input system 100 pops up a menu window listing voice recognition candidates for the text displayed in gray on the screen.
- the illustrated menu window five text conversion candidates for unconfirmed input speech are displayed.
- the control unit 104 activates a normal voice recognition unit for the target (or most recently input) normal voice.
- a menu window listing these conversion candidates is generated, and the menu window is displayed as a pop-up on the text editing screen.
- the candidates on the menu are labeled 1, 2, 3, etc., and the user can enter the menu by whispering "ONE”, “TWO”, “THREE”, etc. You can select your favorite candidate. Here, the user whispers "FOUR” and selects the fourth candidate "of sitting by her sister.”
- the classification unit 101 classifies the voice waveform corresponding to “FOUR” as a whisper, and the whisper recognition unit 102 recognizes that the voice waveform is a command to select the fourth candidate “FOUR”. do. Then, when the control unit 104 confirms the selection of the fourth text conversion candidate in response to the recognition result of "FOUR", that is, selection of the fourth candidate, by the whisper recognition unit 102, as shown in FIG. The menu window is closed, and the undetermined text displayed in gray on the text editing screen shown in FIG. 8 is replaced with the selected text candidate "of sitting by her sister" and displayed.
- the classification unit 101 classifies the speech waveform up to “...conversations in it” as normal speech, the normal speech recognition unit 103 converts it into text, and the control unit 104 converts it into text on the text editing screen. The text up to "...conversations in it” is displayed. Next, the classification unit 101 classifies the voice waveform corresponding to "COMMA DOUBLE QUOTE" as a whisper, and the whisper recognition unit 102 recognizes that the voice waveform is a symbol input of "COMMA” and "DOUBLE QUOTE".
- the control unit 104 displays the input confirmed text on the text editing screen as shown in FIG.
- Each symbol ",” and “" is successively connected to the end of the symbol.
- special characters can also be input in a whisper using the same method as symbol input.
- Figures 12 and 13 show a voice input method for inputting words, such as abbreviations, that are quite difficult to input using normal ASR (or words that are not yet registered in the dictionary) by combining normal voice input and whispered voice commands.
- An example of the operation of the input system 100 is shown.
- the user inputs the spelling "w a v 2 v e c 2" in normal voice, and immediately after that, utters "SPELL” in a whisper.
- the classification unit 101 classifies the voice waveform corresponding to “w a v 2 v e c 2” as normal voice
- the normal voice recognition unit 103 classifies the voice waveform into each of the alphabets. Convert to characters “w”, “a”, “v”, “2”, “v”, “e”, “c”, and “2”.
- the control unit 104 adds an alphabetical character to the end of the input text on the text editing screen, as shown in FIG.
- the characters "w”, “a”, “v”, “2”, “v”, “e”, “c”, and “2" are successively connected.
- the classification unit 101 classifies the audio waveform corresponding to "SPELL" as a whisper, and when the whisper recognition unit 102 recognizes that the audio waveform is a command to input the spelling of the word "SPELL", the control starts.
- the unit 104 generates the word “wav2vec2” by combining the letters of the alphabet input just before in the order in which they were uttered.
- the control unit 104 adds “wav2vec2” to the end of the input text on the text editing screen, as shown in FIG. Connect words.
- the classification unit 101 classifies the speech waveforms up to "what a good day today" and "smile” as normal speech, and the normal speech recognition section 103 classifies the speech waveforms as "what a good day”. Convert it to the text "today smile”. Then, as shown in FIG. 14, the control unit 104 displays the text up to "what a good day today” and "smile” on the text editing screen. Next, the classification unit 101 classifies the voice waveform corresponding to the word "EMOTION" as a whisper, and the whisper recognition unit 102 recognizes that the voice waveform is a command instructing input of pictographs. Then, in response to the recognition result of the input of a pictogram by the whisper recognition unit 102, the control unit 104 converts the last text "smile” into a pictogram on the text editing screen, as shown in FIG. indicate.
- Section F describes some applications other than the voice interaction input described in Section E above that can utilize the present disclosure.
- each participant is connected via a network such as TCP/IP, and communicates images of each other's facial images and uttered audio in real time. View real-time facial images and speech audio of other participants, and share meeting materials among participants as needed.
- a network such as TCP/IP
- participants in the conference can give commands using whispers, and voice commands can be prevented from becoming part of the utterances during the conference. That is, when voices uttered by conference participants are input to the voice input system 100, the classification unit 101 distinguishes between whispered voices and normal voices. The whisper part is then sent to the whisper recognition unit 102, and the command is processed based on the recognition result. Additionally, to prevent whispers from becoming part of the speech during a meeting, the whispers are removed from the conference participant's audio stream and then multiplexed with the video of the participant's facial image to allow other participants' Send in real time to the conference terminal.
- FIG. 16A shows an example of the original video of a participant speaking alternately in a normal voice and a whisper.
- the section in which the whispering voice is uttered is a silent section in which other participants do not speak, the participant's mouth is moving in the same way as in the section where normal voices can be heard. It feels unnatural to other participants watching such videos.
- an image of the lips can be generated from the utterance using deep learning, for example using Wav2Lip (see Non-Patent Document 6).
- Wav2Lip Wav2Lip
- both the facial image stream and the audio stream can be adjusted so that the parts of the whispered commands are no longer visible or audible to other participants.
- FIG. 17 shows an example of a functional configuration of a conference terminal 1700 to which the voice input system 100 according to the present disclosure is applied.
- FIG. 17 only the input-side functional blocks of the conference terminal 1700 that capture video and audio of participants and transmit them to other conference terminals are illustrated, and for convenience of explanation and simplification of the drawing, The functional blocks on the output side that output video and audio received from other conference terminals are not shown.
- the same functional blocks as those included in the voice input system 100 shown in FIG. 1 are shown with the same names and symbols.
- the conference terminal 1700 further includes a microphone 1701, a camera 1702, a data processing section 1703, and a data transmission section 1704.
- the microphone 1701 inputs the voices spoken by the conference participants, and the camera 1702 images the conference participants.
- the classification unit 101 classifies the voices of conference participants input from the microphone 1701 into normal voices and whispers.
- the whisper recognition unit 102 recognizes and processes the voice signal classified as a whisper and converts it into a command.
- the control unit 104 executes processing of the command recognized by the whisper recognition unit 102.
- the content of the command is not particularly limited. Furthermore, if text is not input on the conference terminal 1700, the normal speech recognition unit 103 is not necessary.
- the data processing unit 1703 multiplexes the audio stream classified as normal audio by the classification unit 101 and the video stream captured by the camera 1702 into a predetermined format (for example, MPEG (Moving Picture Experts Group), etc.). convert and encode.
- a predetermined format for example, MPEG (Moving Picture Experts Group), etc.
- the audio stream from which the whispering portion has been deleted is input from the classification unit 101 to the data processing unit 1703.
- the conference participants move their mouths when speaking either normal voice or whispering, and in the section where the whispering voice is deleted, the conference participants move their mouths. and audio are not consistent. Therefore, as described with reference to FIG. 16(B), the data processing unit 1703 uses Wav2Lip, for example, in the section where the participant is speaking in a whisper to make it appear as if the participant is not speaking.
- the video is processed by replacing the lips of the participant's video, and then multiplexed with the audio stream and encoded.
- the data communication unit 1704 transmits the data (video and audio stream) processed by the data processing unit 1703 to other conference terminals via the network in accordance with a predetermined communication protocol such as TCP/IP.
- the first avatar can be made to speak using the user's normal voice
- the second avatar can be operated using the user's whispering voice.
- the ⁇ operation'' of the avatar here includes various actions of the avatar, including physical movements and speech of the avatar.
- Another example of simultaneous control of multiple avatars is to have the first avatar speak using the user's normal voice, and the second avatar speak using the user's whispered voice. can.
- the voice conversion technology disclosed in Non-Patent Document 7 and Non-Patent Document 8 is used to convert the whisper into normal voice and use it as the voice of the avatar. Good too.
- FIG. 18 schematically shows an example of a functional configuration of an avatar control system 1800 that simultaneously controls a plurality of avatars and is configured by incorporating the functions of the voice input system 100 according to the present disclosure.
- the same functional blocks as those included in the voice input system 100 shown in FIG. 1 are shown with the same names and symbols.
- the operation of the avatar control system 1800 when the first avatar speaks using the user's normal voice and the second avatar speaks using the user's whispered voice will be described.
- the classification unit 101 classifies the user's voice input from a microphone or headset into normal voice and whisper voice.
- the first avatar voice generation unit 1801 converts the voice signal classified into normal voice by the classification unit 101 to generate the voice of the first avatar.
- the algorithm for converting normal voice into other voices is arbitrary and may utilize currently available voice changers.
- the second avatar voice generation unit 1802 converts the voice signal classified as a whisper by the classification unit 101 into normal voice, and generates the voice of the second avatar.
- the voice conversion technology disclosed in Non-Patent Document 7 and Non-Patent Document 8 may be used to convert the whispered voice into other normal voice and use it as the utterance of the avatar.
- Silent Speech is an interface that allows voice input without being noticed by those around you, either silently or in a low voice, ensuring that the voice commands do not become noise to those around you, and that confidential information is not disclosed.
- the main purpose is to protect privacy.
- the sound pressure level of a typical conversation is about 60 dB, while the sound pressure level of a whisper is 30 to 40 dB. In this way, the purpose of silent speech can be largely achieved by using whispers as spoken commands.
- a mask that enables powered ventilation for breathing has been proposed to protect against air pollution and infectious diseases (see Patent Document 2).
- a mask-type wearable interface can be realized.
- a microphone can be placed near the mouth, making it possible to pick up whispers even at low sound pressure levels.
- FIG. 19 shows an example of a functional configuration of a mask-type wearable interface 1900 incorporating the functions of the voice input system 100 according to the present disclosure.
- the same functional blocks as those included in voice input system 100 shown in FIG. 1 are shown with the same names and symbols.
- the wearable interface 1900 has at least a microphone 1901 and a speaker 1902 mounted on a mask-type main body. At least some of the components of the voice input system 100, such as the classification unit 101 and the whispered voice recognition unit 102, may also be installed in the wearable interface 1900, or the functions of the voice input system 100 may be placed outside the wearable interface 1900. You can leave it there.
- the classification unit 101 classifies the user's voice input from the microphone 1901 into normal voice and whisper voice.
- the whisper recognition unit 102 recognizes a voice command from the voice signal classified as a whisper by the classification unit 101.
- the control unit 104 then processes the whisper command.
- the content of the whisper command is not particularly limited.
- the whisper command may be a process for normal voice input from the microphone 1901, or may be a command for a personal computer (PC) or other information terminal to which the wearable interface 1900 is connected.
- PC personal computer
- the microphone 1901 placed near the mouth can be used to pick up whispering voices with a low sound pressure level, so it is possible to prevent voice commands from being missed.
- the audio signal classified as normal audio by the classification unit 101 is amplified by an amplifier 1903 and then output as audio from a speaker 1902 attached to a mask, that is, a wearable interface 1900.
- Wearing the mask-type wearable interface 1900 covers the user's mouth, making it difficult to hear normal voices, but by amplifying the sound with the speakers 1902, it is possible to compensate for the attenuation of the voices caused by the mask.
- the mask-type wearable interface 1900 by picking up whispered voices with the microphone inside the mask, it is possible to achieve an effect almost equivalent to silent speech. That is, by introducing the voice input system 100 according to the present disclosure into a mask that is always worn, it is possible to construct a voice interface that can be used at all times without interfering with normal conversation.
- the mask-type wearable interface 1900 there is a possibility of using three modalities: silent speech, normal voice, and whispered voice. For example, if we can recognize lip reading and whispering, we can obtain three types of speech modalities in addition to normal speech.
- Non-Patent Document 9 discloses "voice shift" which specifies the mode of voice input by intentionally controlling the pitch of voice.
- voice shift a different mode is determined when the fundamental frequency (F0) of speech exceeds a specified threshold.
- F0 fundamental frequency
- Voice shifting requires the user to speak at an unnaturally high pitch in order to stably recognize the two voice input modes. In the first place, it is difficult for the user to understand which frequencies should be used for different pitches.
- switching between whispering and normal speech is more natural and can be performed more clearly without setting a threshold (unknown to the user).
- Non-Patent Document 10 discloses a method for automatically detecting a filled (uttered) pause, which is one of the hesitation phenomena in uttering a voice command, and suggesting candidates that can fill the command. has been done. According to this method, for example, if a user says “play, Beeee" and stops, the system will detect “eeee” as a pause filled with hesitation, and the system will detect the "eeee” as a pause filled with hesitation, and the Filled candidates can be suggested.
- this method shows the possibility of indicating nonverbal intent in speech, it only utilizes hesitation and cannot utilize arbitrary commands as in the present disclosure. Moreover, making a vocalized pause in this way is only possible after the vowel of the utterance, and not after the consonant.
- PrivateTalk A technique called "PrivateTalk" disclosed in Non-Patent Document 11 uses a hand that partially covers the mouth from one side to activate voice commands.
- the primary purpose of PrivateTalk is to protect privacy, it can also be used to differentiate between normal speech (without hand covering) and commands (hand covering).
- this disclosure differs in that it is no longer a "hands-free" interaction as explicit hand gestures are required.
- PrivateTalk requires two microphones (connected to the left and right earphones) to recognize the effect of the hand cover. In contrast, in the present disclosure, only one standard microphone is sufficient. Further, according to the present disclosure, privacy can be protected by a natural and effective method of covering the mouth and speaking in a whisper.
- DualBreath uses breathing as a command and distinguishes between inhaling and exhaling air through the nose and mouth at the same time, thereby distinguishing it from normal exhalation. DualBreath can express a trigger by simply pressing a button, but it cannot express a command as rich as a whisper as in the present disclosure.
- ProxiMic disclosed in Non-Patent Document 13 is a sensing technology that detects the user's utterances using a microphone device placed near the mouth, and is a sensing technology that detects the user's utterances using a microphone device placed near the mouth. It is intended to be used as an utterance of "wake-up-free,” such as “wake-up-free.”
- ProxiMic requires physical movement such as moving the microphone close to the mouth, it is not easy to mix normal speech and speech near the mouth.
- Non-Patent Document 14 discloses a technology called "SilentVoice” that allows voice input using "inhalation sounds” that are uttered while inhaling.
- SilentVoice is primarily designed for silent speech, it can also distinguish between normal speech and intrusive sounds. However, it requires a special microphone placed very close to the mouth, and training is required for the user to speak correctly in inhalation mode. Furthermore, it is difficult for a person to frequently switch between normal speech and inspired speech.
- Alexa smart speaker supports whisper mode. When this mode is set, if you speak to Alexa in a whisper, Alexa also responds in a whisper, but unlike the present disclosure, voice commands are not entered in a whisper.
- Section H describes the information processing device used to realize the voice input system 100 according to the present disclosure, and also to realize the various applications of the present disclosure introduced in Section F above. do.
- FIG. 20 shows a configuration example of an information processing device 2000 that performs classification between normal voice and whisper voice, recognizes whisper voice, or realizes various applications that utilize the whisper voice recognition results. Further, the information processing device 2000 can also be used as the conference terminal 1700 described in the above section F-1.
- the information processing device 2000 shown in FIG. 20 includes a CPU (Central Processing Unit) 2001, a ROM (Read Only Memory) 2002, a RAM (Random Access Memory) 2003, and a host bus 20. 04, bridge 2005, and expansion bus 2006. , an interface section 2007, an input section 2008, an output section 2009, a storage section 2010, a drive 2011, and a communication section 2013.
- a CPU Central Processing Unit
- ROM Read Only Memory
- RAM Random Access Memory
- the CPU 2001 functions as an arithmetic processing device and a control device, and controls the overall operation of the information processing device 2000 according to various programs.
- the ROM 2002 non-volatilely stores programs used by the CPU 2001 (such as a basic input/output system) and calculation parameters.
- the RAM 2003 is used to load programs used in the execution of the CPU 2001, and to temporarily store parameters such as work data that change as appropriate during program execution. Programs loaded into the RAM 2003 and executed by the CPU 2001 include, for example, various application programs and an operating system (OS).
- OS operating system
- the CPU 2001, ROM 2002, and RAM 2003 are interconnected by a host bus 2004 composed of a CPU bus and the like. Through the cooperative operation of the ROM 2002 and the RAM 2003, the CPU 2001 can execute various application programs in an execution environment provided by the OS to realize various functions and services.
- the OS is, for example, Microsoft Windows or Unix.
- the information processing device 2000 is an information terminal such as a smartphone or a tablet
- the OS is, for example, iOS from Apple Inc. or Android from Google Inc.
- the application programs include applications that classify normal voices and whispered voices and recognize whispered voices, and various applications that utilize the results of whispering voice recognition.
- the host bus 2004 is connected to an expansion bus 2006 via a bridge 2005.
- the expansion bus 2006 is, for example, a PCI (Peripheral Component Interconnect) bus or PCI Express, and the bridge 2005 is based on the PCI standard.
- PCI Peripheral Component Interconnect
- the bridge 2005 is based on the PCI standard.
- the interface unit 2007 connects peripheral devices such as an input unit 2008, an output unit 2009, a storage unit 2010, a drive 2011, and a communication unit 2013 in accordance with the standard of the expansion bus 2006.
- peripheral devices such as an input unit 2008, an output unit 2009, a storage unit 2010, a drive 2011, and a communication unit 2013 in accordance with the standard of the expansion bus 2006.
- the information processing apparatus 2000 may further include peripheral devices that are not shown.
- the peripheral devices may be built into the main body of the information processing device 2000, or some peripheral devices may be externally connected to the main body of the information processing device 2000.
- the input unit 2008 includes an input control circuit that generates an input signal based on input from the user and outputs it to the CPU 2001.
- the input unit 2008 may include a keyboard, a mouse, and a touch panel, and may also include a camera and a microphone.
- the input unit 2008 is, for example, a touch panel, a camera, or a microphone, but may further include other mechanical operators such as buttons.
- input devices other than the microphone are almost unnecessary.
- the output unit 2009 includes, for example, a display device such as a liquid crystal display (LCD) device, an organic EL (Electro-Luminescence) display device, and an LED (Light Emitting Diode).
- a display device such as a liquid crystal display (LCD) device, an organic EL (Electro-Luminescence) display device, and an LED (Light Emitting Diode).
- voice dialogue on the information processing device 2000 as in this embodiment, text input using normal voice, special characters such as symbols input using a whispered voice, and editing commands such as deletion and line feed are used.
- the execution results are presented using a display device.
- the output unit 2009 may include an audio output device such as a speaker and headphones, and output at least a part of the message to the user displayed on the UI screen as an audio message.
- the storage unit 2010 stores files such as programs (applications, OS, etc.) executed by the CPU 2001 and various data.
- the data stored in the storage unit 2010 may include a corpus of ordinary voices and whispers (described above) for training a neural network.
- the storage unit 2010 is configured with a large-capacity storage device such as an SSD (Solid State Drive) or an HDD (Hard Disk Drive), but may also include an external storage device.
- the removable storage medium 2012 is a cartridge-type storage medium such as a microSD card, for example.
- the drive 2011 performs read and write operations on the loaded removable storage medium 113.
- the drive 2011 outputs data read from the removable recording medium 2012 to the RAM 2003 or the storage unit 2010, or writes data on the RAM 2003 or the storage unit 2010 to the removable recording medium 2012.
- the communication unit 2013 is a device that performs wireless communication such as Wi-Fi (registered trademark), Bluetooth (registered trademark), and cellular communication networks such as 4G and 5G.
- the communication unit 2013 also includes terminals such as USB (Universal Serial Bus) and HDMI (registered trademark) (High-Definition Multimedia Interface), and enables HDMI (registered trademark) communication with USB devices such as scanners and printers, displays, etc. It may further include a function to perform the following.
- the information processing device 2000 is not limited to one device, and may be distributed over two or more devices to realize the voice input system 100 shown in FIG. It is also possible to execute the processing of the various applications introduced in Section F.
- PC personal computer
- a voice input system that can input non-text commands using a whispered voice and input text using normal voice.
- ordinary voice input can be used for text input, and various commands can be input simply by whispering.
- No special hardware other than a regular microphone is required to implement the present disclosure, and it can be used in a wide range of situations where voice recognition is already available.
- two useful neural networks can be provided.
- One is a neural network that can be used to distinguish between whispers and normal speech, and the other is a neural network that can be used to recognize whispers.
- excellent whisper recognition accuracy can be achieved by fine-tuning a model pre-trained on normal speech using whispered speech.
- HuBERT Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units. (June 2 021). arXiv:2106.07447 [ cs.CL] Boon Pang Lim. 2010. Computational differences between whispered and non-whispered speech. Ph.D. Dissertation. University of Illinois Urbana-Champaign. Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur. 2015. Librispeech: An ASR corpus based on public domain audio books. In 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP).
- ICASSP Acoustics, Speech and Signal Processing
- ProxiMic Convenient Voice Activation via Close-to-Mic SpeechDetect ed by a Single Microphone. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems. Association for Computing Machinery, New York, NY, USA, 1-12. Masaaki Fukumoto. 2018. SilentVoice: Unnoticeable Voice Input by Ingressive Speech. In Proceedings of the 31st Annual ACM Symposium on User Interface Software and Technology (Berlin, Germany) (UIST '18). Association for Computing Machinery, New York, NY, USA , 237-246.
- ASR automatic speech recognition
- commands using whispered voices in remote conferences voice switching of multiple avatars, and combination with silent speech.
- An information processing device comprising:
- the classification unit uses a first trained neural network to classify normal speech and whispering;
- the recognition unit recognizes a whisper using a second trained neural network.
- the second trained neural network is made of wave2vec2.0 or HuBERT;
- the second trained neural network is pre-trained on a normal speech corpus, and then fine-tuned using whispered voices.
- the information processing device according to any one of (2) or (3) above.
- the fine-tuning includes a first-stage fine-tuning using a general-purpose whispering voice corpus and a second-stage fine-tuning using a whispering voice database for each user.
- the information processing device according to (4) above.
- the second trained neural network includes a feature extraction layer and a transformer layer,
- the first trained neural network is configured to share the feature extraction layer with the second trained neural network,
- the information processing device according to any one of (2) to (5) above.
- the control unit performs processing corresponding to the recognition result of the whispering voice by the recognition unit on the recognition result of the normal voice recognition unit.
- the control unit executes processing of the whisper command recognized by the recognition unit on the text into which the normal voice recognition unit has converted the normal voice.
- the information processing device according to (7) above.
- the control unit executes at least one of inputting a symbol or special character to the text, selecting a text conversion candidate, deleting the text, and starting a line in the text based on the whisper command.
- the control unit In response to the recognition unit recognizing a whisper command instructing input of pictograms, the control unit converts the normal voice immediately before the whisper into a pictogram.
- the information processing device according to any one of (8) to (10) above.
- the control unit removes the voice classified as a whisper from the original speech voice and transmits it to an external device.
- the information processing device according to any one of (1) to (11) above.
- the control unit performs a process of replacing the lips of the image of the speaker in the section classified as a whisper by the classification unit with an image of the person not speaking.
- the information processing device according to (12) above.
- a first voice generation unit that generates the voice of the first avatar based on the voice classified as normal voice by the recognition unit;
- a second voice generation unit that generates the voice of a second avatar based on the voice classified as a whisper by the recognition unit;
- the information processing device according to any one of (1) to (14) above.
- a classification unit that classifies spoken voice into normal voice and whispered voice based on voice features; a recognition unit that recognizes the whisper classified by the classification unit; a control unit that controls processing based on the recognition result of the recognition unit;
- Data processing section, 1704... Data transmission section 1800... Avatar control system 1801... First avatar voice generation section 1802... Second avatar voice generation section 1900... Wearable interface , 1901...Microphone 1902...Speaker, 1903...Amplification section 2000...Information processing device, 2001...CPU, 2002...ROM 2003...RAM, 2004...Host bus, 2005...Bridge 2006...Expansion bus, 2007...Interface section 2008...Input section, 2009...Output section, 2010...Storage section 2011...Drive, 2012...Removable recording medium 2013...Communication section
Landscapes
- Engineering & Computer Science (AREA)
- Physics & Mathematics (AREA)
- Theoretical Computer Science (AREA)
- Computational Linguistics (AREA)
- Health & Medical Sciences (AREA)
- Artificial Intelligence (AREA)
- General Physics & Mathematics (AREA)
- Evolutionary Computation (AREA)
- Biophysics (AREA)
- Computing Systems (AREA)
- Multimedia (AREA)
- Biomedical Technology (AREA)
- Software Systems (AREA)
- Data Mining & Analysis (AREA)
- Mathematical Physics (AREA)
- General Health & Medical Sciences (AREA)
- Molecular Biology (AREA)
- Life Sciences & Earth Sciences (AREA)
- General Engineering & Computer Science (AREA)
- Audiology, Speech & Language Pathology (AREA)
- Human Computer Interaction (AREA)
- Acoustics & Sound (AREA)
- Signal Processing (AREA)
- Machine Translation (AREA)
- Quality & Reliability (AREA)
- Computer Vision & Pattern Recognition (AREA)
- Facsimiles In General (AREA)
Abstract
Description
発話音声を音声特徴量に基づいて通常の音声とささやき声に分類する分類部と、
前記分類部によって分類されたささやき声を認識する認識部と、
前記認識部の認識結果に基づく処理を制御する制御部と、
を具備する情報処理装置である。
発話音声を音声特徴量に基づいて通常の音声とささやき声に分類する分類ステップと、
前記分類ステップにおいて分類されたささやき声を認識する認識ステップと、
前記認識ステップにおける認識結果に基づく処理を制御する制御ステップと、
を有する情報処理方法である。
発話音声を音声特徴量に基づいて通常の音声とささやき声に分類する分類部、
前記分類部によって分類されたささやき声を認識する認識部、
前記認識部の認識結果に基づく処理を制御する制御部、
としてコンピュータを機能させるようにコンピュータ可読形式で記述されたコンピュータプログラムである。
B.基本構成
C.ささやき声の利用
D.認識アーキテクチャ
D-1.ささやき声の認識
D-2.ささやき声の分類
D-3.トレーニング用データセット
E.対話システムへの適用
F.他のアプリケーション
F-1.遠隔会議でのささやきコマンドの使用
F-2.複数のアバターの制御
F-3.サイレントスピーチとの組み合わせ
G.先行技術との対比
H.情報処理装置の構成
I.結論
自動音声認識(ASR)は、スマートスピーカーやカーナビゲーションシステムなどのデバイスの操作からテキスト入力の手法に至るまで、さまざまなアプリケーションで既に使用されている。音声入力はマイク以外の特別な入力機器を必要としないので、ハンズフリーで操作可能である。さらに、音声に基づくテキスト入力はタイプ入力よりもはるかに高速であることから、音声入力はアイデアを書き留めたり、原稿の下書きをすばやく入力したりするためにも使用される。
図1には、本開示を適用した音声入力システム100の基本構成を模式的に示している。音声入力システム100は、ささやき声と通常の音声で、発話に複数の意味を与えるシステムである。音声入力システム100は、例えばパーソナルコンピュータ(PC)のような一般的な情報処理装置を用いて構成することができ、音声入力にはマイク以外の特別な入力機器を必要としない。一般的な情報処理装置はキーボードやマウスといったユーザが入力操作を行う入力装置を備えているが、本開示により音声入力を行う上では、マイク以外の入力装置を使用する必要はない。
ささやき声は、多くの人が特別なトレーニングなしで発音できる発話のモードであり、発話中に通常の音声とささやき声に自在に使い分けることができる。これが、本開示において、通常の音声の別の音声入力モードとしてささやき声を使用する理由の1つである。
このD項では、本開示を適用した音声入力システム100において、ささやき声の認識と分類を可能にするニューラルネットワークの構成について説明する。
本開示の最良の実施形態では、ささやき声の認識(すなわち、ささやき声認識部102)には、wav2vec2.0(非特許文献1及び非特許文献2を参照のこと)又はHuBERT(非特許文献3を参照のこと)を使用する。wav3ves2.0及びHuBERTはいずれも音声処理システム用に設計された自己教師ありニューラルネットワークである。
分類部101は、固定長(100ミリ秒など)の音声信号により、ささやき声と通常の音声入力を区別する。図4右には、wav2vec2.0又はHuBERTからなるささやき声認識ニューラルネットワーク400のうち、特徴抽出層410の構成を詳細に示している。特徴抽出層410は、音響信号を20ミリ秒毎に512次元の特徴ベクトルに変換する。分類部101は、ささやき声認識ニューラルネットワーク400とニューラルネットワーク(すなわち、特徴抽出層410)を部分的に共有することによって、音声入力システム100全体としてネットワークサイズが縮小される。
ささやき声認識ニューラルネットワーク400のトレーニングに関しては、まず、通常の音声データを用いて事前トレーニング及びファインチューニングされたニューラルネットワークから開始する。具体的には、通常の音声データとしてLibrispeechデータセット(非特許文献5を参照のこと)を使用し、通常の音声を用いた960時間のトレーニングを実施する。事前トレーニングには教師データ(音声に対応するテキスト)は必要でなく、音声データのみが必要であるという点には留意されたい。
本開示に係る音声入力システム100では、通常の音声入力は従来の音声入力テキスト作成と同じようにテキストに変換される。対照的に、「COMMA」、「PERIOD」、「QUOTE」などの記号をささやき声で話すと、音声入力システム100は、これらを記号として扱う。
このF項では、上記E項で説明した音声対話入力以外の、本開示を利用できるいくつかのアプリケーションについて説明する。
本開示の他のアプリケーションとして、遠隔会議中の音声コマンドの使用が挙げられる。遠隔会議中に音声コマンドを使用することを想定する。コマンドが会議中の発話の一部になるのを防ぐために、会議の参加者はささやき声を使ってコマンドを与える。
本開示のさらに他のアプリケーションとして、1人のユーザが仮想空間における複数のアバターの制御を同時に行うことができる。
サイレントスピーチは、無声音又は小声で周囲に気づかれずに音声入力を可能にするインターフェースであり、音声コマンドの音声が周囲にノイズにならないようにすることと、機密情報が開示されないことによってプライバシーを保護することを主な目的とする。一般的な会話の音圧レベルは約60dBであるが、ささやき声の音圧レベルは30~40dBである。このように、ささやき声を発話コマンドとして使用することで、サイレントスピーチの目的を大幅に達成することができる。
このG項では、本出願人が調査した先行技術と本開示との比較結果及び本開示の優位な点について説明する。
このH項では、本開示に係る音声入力システム100を実現するために、さらには上記F項で紹介した本開示の各種アプリケーションを実現するために利用される情報処理装置について説明する。
最後に、本開示の利点及び本開示によってもたらされる効果についてまとめておく。
前記分類部によって分類されたささやき声を認識する認識部と、
前記認識部の認識結果に基づく処理を制御する制御部と、
を具備する情報処理装置。
前記認識部は、第2の学習済みニューラルネットワークを用いてささやき声の認識を行う、
上記(1)に記載の情報処理装置。
上記(2)に記載の情報処理装置。
上記(2)又は(3)のいずれか1つに記載の情報処理装置。
上記(4)に記載の情報処理装置。
前記第1の学習済みニューラルネットワークは、前記第2の学習済みニューラルネットワークと前記特徴抽出層を共有して構成される、
上記(2)乃至(5)のいずれか1つに記載の情報処理装置。
前記制御部は、前記通常の音声認識部の認識結果に対して前記認識部によるささやき声の認識結果に対応する処理を実施する、
上記(1)乃至(6)のいずれか1つに記載の情報処理装置。
上記(7)に記載の情報処理装置。
上記(8)に記載の情報処理装置。
上記(8)又は(9)のいずれか1つに記載の情報処理装置。
上記(8)乃至(10)のいずれか1つに記載の情報処理装置。
上記(1)乃至(11)のいずれか1つに記載の情報処理装置。
上記(12)に記載の情報処理装置。
前記認識部がささやき声に分類した音声に基づいて第2のアバターの音声を生成する第2の音声生成部と、
をさらに備える上記(1)乃至(13)のいずれか1つに記載の情報処理装置。
前記分類部によって通常の音声に分類された音声信号を増幅する増幅部と、
をさらに備え、
前記増幅部で増幅した音声信号を前記スピーカーから音声出力する、
上記(1)乃至(14)のいずれか1つに記載の情報処理装置。
前記分類ステップにおいて分類されたささやき声を認識する認識ステップと、
前記認識ステップにおける認識結果に基づく処理を制御する制御ステップと、
を有する情報処理方法。
前記分類部によって分類されたささやき声を認識する認識部、
前記認識部の認識結果に基づく処理を制御する制御部、
としてコンピュータを機能させるようにコンピュータ可読形式で記述されたコンピュータプログラム。
102…ささやき声認識部、103…通常の音声認識部
104…制御部
400…ささやき声認識ニューラルネットワーク
410…特徴抽出層、411…正規化層(Layer Norm)
412…平均プーリング層(Avg Pool)
413、414…FC(全結合)層
415…出力層(Logsoftmax)、420…トランスフォーマ層
430…Projection層、440…CTC層
701…正規化層(gMLP)
702平均プーリング層(Avg Pool)
703…FC(全結合)層、704…出力層(Logsoftmax)
1700…会議端末、1701…マイク、1702…カメラ
1703…データ処理部、1704…データ送信部
1800…アバター制御システム
1801…第1のアバター音声生成部
1802…第2のアバター音声生成部
1900…ウェアラブルインターフェース、1901…マイク
1902…スピーカー、1903…増幅部
2000…情報処理装置、2001…CPU、2002…ROM
2003…RAM、2004…ホストバス、2005…ブリッジ
2006…拡張バス、2007…インターフェース部
2008…入力部、、2009…出力部、2010…ストレージ部
2011…ドライブ、2012…リムーバブル記録媒体
2013…通信部
Claims (17)
- 発話音声を音声特徴量に基づいて通常の音声とささやき声に分類する分類部と、
前記分類部によって分類されたささやき声を認識する認識部と、
前記認識部の認識結果に基づく処理を制御する制御部と、
を具備する情報処理装置。 - 前記分類部は、第1の学習済みニューラルネットワークを用いて通常の音声とささやき声の分類を行い、
前記認識部は、第2の学習済みニューラルネットワークを用いてささやき声の認識を行う、
請求項1に記載の情報処理装置。 - 前記第2の学習済みニューラルネットワークは、wave2vec2.0又はHuBERTからなる、
請求項2に記載の情報処理装置。 - 前記第2の学習済みニューラルネットワークは、通常の音声コーパスで事前トレーニングされた後にささやき声によるファインチューニングが実施されている、
請求項2に記載の情報処理装置。 - 前記ファインチューニングは、汎用のささやき声コーパスを用いた第1段階のファインチューニングと、ユーザ毎のささやき声のデータベースを用いた第2段階のファインチューニングを含む、
請求項4に記載の情報処理装置。 - 前記第2の学習済みニューラルネットワークは、特徴抽出層及びトランスフォーマ層を含み、
前記第1の学習済みニューラルネットワークは、前記第2の学習済みニューラルネットワークと前記特徴抽出層を共有して構成される、
請求項2に記載の情報処理装置。 - 前記分類部によって分類された通常の音声を認識する通常音声認識部をさらに備え、
前記制御部は、前記通常音声認識部の認識結果に対して前記認識部によるささやき声の認識結果に対応する処理を実施する、
請求項1に記載の情報処理装置。 - 前記制御部は、前記通常音声認識部が通常の音声を変換したテキストに対して、前記認識部が認識したささやき声コマンドの処理を実行する、
請求項7に記載の情報処理装置。 - 前記制御部は、前記ささやき声コマンドに基づいて、前記テキストに対する記号又は特殊文字の入力、テキスト変換候補の選択、テキストの削除、テキストの改行のうち少なくとも1つを実行する、
請求項8に記載の情報処理装置。 - 通常の音声で複数の文字が発話された後にささやき声で「SPELL」と発話されたときに、前記認識部は単語の綴りの入力を指示するコマンドであることを認識し。前記制御部は前記複数の文字を発話された順に連結した単語を生成する、
請求項8に記載の情報処理装置。 - 前記認識部が絵文字入力を指示するささやき声コマンドを認識したことに応答して、前記制御部はそのささやき声の直前の通常の音声を絵文字に変換する、
請求項8に記載の情報処理装置。 - 前記制御部は、ささやき声に分類された音声を元の発話音声から除去して、外部の装置に送信する、
請求項1に記載の情報処理装置。 - 前記制御部は、前記分類部がささやき声に分類した区間の話者の映像の唇の部分を発話していないときの映像に置き換える処理を行う、
請求項12に記載の情報処理装置。 - 前記認識部が通常の音声に分類した音声に基づいて第1のアバターの音声を生成する第1の音声生成部と、
前記認識部がささやき声に分類した音声に基づいて第2のアバターの音声を生成する第2の音声生成部と、
をさらに備える請求項1に記載の情報処理装置。 - 話者が装着するマスクに搭載されたマイク及びスピーカーと、
前記分類部によって通常の音声に分類された音声信号を増幅する増幅部と、
をさらに備え、
前記増幅部で増幅した音声信号を前記スピーカーから音声出力する、
請求項1に記載の情報処理装置。 - 発話音声を音声特徴量に基づいて通常の音声とささやき声に分類する分類ステップと、
前記分類ステップにおいて分類されたささやき声を認識する認識ステップと、
前記認識ステップにおける認識結果に基づく処理を制御する制御ステップと、
を有する情報処理方法。 - 発話音声を音声特徴量に基づいて通常の音声とささやき声に分類する分類部、
前記分類部によって分類されたささやき声を認識する認識部、
前記認識部の認識結果に基づく処理を制御する制御部、
としてコンピュータを機能させるようにコンピュータ可読形式で記述されたコンピュータプログラム。
Priority Applications (4)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US18/858,123 US20250273203A1 (en) | 2022-04-26 | 2023-03-01 | Information processing device, information processing method, and computer program |
| JP2024517869A JPWO2023210149A1 (ja) | 2022-04-26 | 2023-03-01 | |
| EP23795897.0A EP4517746A4 (en) | 2022-04-26 | 2023-03-01 | INFORMATION PROCESSING DEVICE, INFORMATION PROCESSING METHOD, AND COMPUTER PROGRAM |
| CN202380034669.XA CN119054015A (zh) | 2022-04-26 | 2023-03-01 | 信息处理装置、信息处理方法和计算机程序 |
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP2022072123 | 2022-04-26 | ||
| JP2022-072123 | 2022-04-26 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2023210149A1 true WO2023210149A1 (ja) | 2023-11-02 |
Family
ID=88518435
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/JP2023/007479 Ceased WO2023210149A1 (ja) | 2022-04-26 | 2023-03-01 | 情報処理装置及び情報処理方法、並びにコンピュータプログラム |
Country Status (5)
| Country | Link |
|---|---|
| US (1) | US20250273203A1 (ja) |
| EP (1) | EP4517746A4 (ja) |
| JP (1) | JPWO2023210149A1 (ja) |
| CN (1) | CN119054015A (ja) |
| WO (1) | WO2023210149A1 (ja) |
Cited By (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN118298855A (zh) * | 2024-06-05 | 2024-07-05 | 山东第一医科大学附属省立医院(山东省立医院) | 一种婴儿哭声识别护理方法、系统及存储介质 |
Citations (8)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP2000338986A (ja) * | 1999-05-28 | 2000-12-08 | Canon Inc | 音声入力装置及びその制御方法及び記憶媒体 |
| JP2001079265A (ja) * | 1999-09-14 | 2001-03-27 | Sega Corp | ゲーム装置 |
| JP2005140859A (ja) * | 2003-11-04 | 2005-06-02 | Canon Inc | 音声認識装置および方法 |
| JP2015219480A (ja) | 2014-05-21 | 2015-12-07 | 日本電信電話株式会社 | 対話状況特徴計算装置、文末記号推定装置、これらの方法及びプログラム |
| JP2016186515A (ja) * | 2015-03-27 | 2016-10-27 | 日本電信電話株式会社 | 音響特徴量変換装置、音響モデル適応装置、音響特徴量変換方法、およびプログラム |
| KR20190133325A (ko) * | 2018-05-23 | 2019-12-03 | 카페24 주식회사 | 음성인식 방법 및 장치 |
| JP2020515877A (ja) * | 2018-04-12 | 2020-05-28 | アイフライテック カンパニー,リミテッド | ささやき声変換方法、装置、デバイス及び可読記憶媒体 |
| US20220054870A1 (en) * | 2020-08-23 | 2022-02-24 | Joseph LaCombe | Face Mask Communication System |
Family Cites Families (11)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US7577564B2 (en) * | 2003-03-03 | 2009-08-18 | The United States Of America As Represented By The Secretary Of The Air Force | Method and apparatus for detecting illicit activity by classifying whispered speech and normally phonated speech according to the relative energy content of formants and fricatives |
| US7728866B2 (en) * | 2005-11-03 | 2010-06-01 | Broadcom Corp. | Video telephony image processing |
| WO2008015800A1 (en) * | 2006-08-02 | 2008-02-07 | National University Corporation NARA Institute of Science and Technology | Speech processing method, speech processing program, and speech processing device |
| US20120004910A1 (en) * | 2009-05-07 | 2012-01-05 | Romulo De Guzman Quidilig | System and method for speech processing and speech to text |
| WO2011025462A1 (en) * | 2009-08-25 | 2011-03-03 | Nanyang Technological University | A method and system for reconstructing speech from an input signal comprising whispers |
| US9867012B2 (en) * | 2015-06-03 | 2018-01-09 | Dsp Group Ltd. | Whispered speech detection |
| US10255907B2 (en) * | 2015-06-07 | 2019-04-09 | Apple Inc. | Automatic accent detection using acoustic models |
| US11089396B2 (en) * | 2017-06-09 | 2021-08-10 | Microsoft Technology Licensing, Llc | Silent voice input |
| JP7000924B2 (ja) * | 2018-03-06 | 2022-01-19 | 株式会社Jvcケンウッド | 音声内容制御装置、音声内容制御方法、及び音声内容制御プログラム |
| US20210027802A1 (en) * | 2020-10-09 | 2021-01-28 | Himanshu Bhalla | Whisper conversion for private conversations |
| US11848019B2 (en) * | 2021-06-16 | 2023-12-19 | Hewlett-Packard Development Company, L.P. | Private speech filterings |
-
2023
- 2023-03-01 WO PCT/JP2023/007479 patent/WO2023210149A1/ja not_active Ceased
- 2023-03-01 JP JP2024517869A patent/JPWO2023210149A1/ja active Pending
- 2023-03-01 US US18/858,123 patent/US20250273203A1/en active Pending
- 2023-03-01 EP EP23795897.0A patent/EP4517746A4/en active Pending
- 2023-03-01 CN CN202380034669.XA patent/CN119054015A/zh active Pending
Patent Citations (8)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP2000338986A (ja) * | 1999-05-28 | 2000-12-08 | Canon Inc | 音声入力装置及びその制御方法及び記憶媒体 |
| JP2001079265A (ja) * | 1999-09-14 | 2001-03-27 | Sega Corp | ゲーム装置 |
| JP2005140859A (ja) * | 2003-11-04 | 2005-06-02 | Canon Inc | 音声認識装置および方法 |
| JP2015219480A (ja) | 2014-05-21 | 2015-12-07 | 日本電信電話株式会社 | 対話状況特徴計算装置、文末記号推定装置、これらの方法及びプログラム |
| JP2016186515A (ja) * | 2015-03-27 | 2016-10-27 | 日本電信電話株式会社 | 音響特徴量変換装置、音響モデル適応装置、音響特徴量変換方法、およびプログラム |
| JP2020515877A (ja) * | 2018-04-12 | 2020-05-28 | アイフライテック カンパニー,リミテッド | ささやき声変換方法、装置、デバイス及び可読記憶媒体 |
| KR20190133325A (ko) * | 2018-05-23 | 2019-12-03 | 카페24 주식회사 | 음성인식 방법 및 장치 |
| US20220054870A1 (en) * | 2020-08-23 | 2022-02-24 | Joseph LaCombe | Face Mask Communication System |
Non-Patent Citations (16)
| Title |
|---|
| ABHISHEK NIRANJANMUKESH SHARMASAI BHARATH CHANDRA GUTHAM ALI BASHA SHAIK, END-TO-END WHISPER TO NATURAL SPEECH CONVERSION USING MODIFIED TRANSFORMER NETWORK., 2020 |
| ALEXEI BAEVSKIHENRY ZHOUABDELRAHMAN MOHAMEDMICHAEL AULI, ARXIV [CS.CL, vol. A Framework for Self-Supervised Learning of Speech, June 2020 (2020-06-01) |
| BOON PANG LIM.: "Ph.D. Dissertation", 2010, UNIVERSITY OF ILLINOIS URBANA-CHAMPAIGN, article "Computational differences between whispered and non-whispered speech" |
| JI WON YOON; BEOM JUN WOO; NAM SOO KIM: "HuBERT-EE: Early Exiting HuBERT for Efficient Speech Recognition", ARXIV.ORG, 13 April 2022 (2022-04-13), XP091203726 * |
| K R PRAJWALRUDRABHA MUKHOPADHYAYVINAY P. NAMBOODIRIC.V. JAWAHAR: "Proceedings of the 28th ACM International Conference on Multimedia", 2020, ASSOCIATION FOR COMPUTING MACHINERY, article "A Lip Sync Expert Is All You Need for Speech to Lip Generation In the Wild", pages: 484 - 492 |
| MASAAKI FUKUMOTO: "Proceedings of the 31st Annual ACM Symposium on User Interface Software and Technology", 2018, ASSOCIATION FOR COMPUTING MACHINERY, article "SilentVoice: Unnoticeable Voice Input by Ingressive Speech", pages: 237 - 246 |
| MASATAKA GOTO.: "A Real-time Filled Pause Detection System for Spontaneous Speech Recognition", PROC. OF EUROSPEECH, 1999, pages 99 |
| MASATAKA GOTOYUKIHIRO OMOTOKATUNOBU ITOUTETSUNORI KOBAYASHI.: "Speech shift: direct speech-input-mode switching through intentional control of voice pitch", EUROPEAN CONFERENCE ON SPEECH COMMUNICATION AND TECHNOLOGY, EUROSPEECH 2003 - INTERSPEECH, 2003 |
| REKIMOTO JUN REKIMOTO@ACM.ORG: "DualVoice: Speech Interaction that Discriminates between Normal and Whispered Voice Input", THE ADJUNCT PUBLICATION OF THE 35TH ANNUAL ACM SYMPOSIUM ON USER INTERFACE SOFTWARE AND TECHNOLOGY, ACMPUB27, NEW YORK, NY, USA, 29 October 2022 (2022-10-29) - 23 September 2022 (2022-09-23), New York, NY, USA, pages 1 - 10, XP058912221, ISBN: 978-1-4503-9427-7, DOI: 10.1145/3526113.3545685 * |
| RYOYA ONISHITAO MORISAKISHUN SUZUKISAYA MIZUTANITAKAAKI KAMIGAKIMASAHIRO FUJIWARAYASUTOSHI MAKINOHIROYUKI SHINODA: "Augmented Humans Conference 2021", 2021, ASSOCIATION FOR COMPUTING MACHINERY, article "DualBreath: Input Method Using Nasal and Mouth Breathing", pages: 283 - 285 |
| SANTIAGO PASCUALANTONIO BONAFONTEJOAN SERRAJOSE A. GONZALEZ, WHISPERED-TO-VOICED ALARYNGEAL SPEECH CONVERSION WITH GENERATIVE ADVERSARIAL NETWORKS., 2018 |
| See also references of EP4517746A4 |
| VASSIL PANAYOTOVGUOGUO CHENDANIEL POVEYSANJEEV KHUDANPUR: "Librispeech: An ASR corpus based on public domain audio books", 2015 IEEE INTERNATIONAL CONFERENCE ON ACOUSTICS, SPEECH AND SIGNAL PROCESSING (ICASSP, 2015, pages 5206 - 5210, XP055270589, Retrieved from the Internet <URL:https://doi.org/10.1109/ICASSP.2015.7178964> DOI: 10.1109/ICASSP.2015.7178964 |
| WEI-NING HSUBENJAMIN BOLTEYAO- HUNG HUBERT TSAIKUSHAL LAKHOTIARUSLAN SALAKHUTDINOVABDELRAHMAN MOHAMED.: "HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units.", ARXIV:2106.07447 [CS.CL, June 2021 (2021-06-01) |
| YUE QINCHUN YUZHAOHENG LIMINGYUAN ZHONGYUKANG YANYUANCHUN SHI: "Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems", 2021, ASSOCIATION FOR COMPUTING MACHINERY, article "ProxiMic: Convenient Voice Activation via Close-to-Mic Speech Detected by a Single Microphone.", pages: 1 - 12 |
| YUKANG YANCHUN YUYINGTIAN SHIMINXING XIE: "Proceedings of the 32nd Annual ACM Symposium on User Interface Software and Technology", 2019, ASSOCIATION FOR COMPUTING MACHINERY, article "PrivateTalk: Activating Voice Input with Hand-On-Mouth Gesture Detected by Bluetooth Earphones", pages: 1013 - 1020 |
Cited By (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN118298855A (zh) * | 2024-06-05 | 2024-07-05 | 山东第一医科大学附属省立医院(山东省立医院) | 一种婴儿哭声识别护理方法、系统及存储介质 |
Also Published As
| Publication number | Publication date |
|---|---|
| CN119054015A (zh) | 2024-11-29 |
| EP4517746A1 (en) | 2025-03-05 |
| JPWO2023210149A1 (ja) | 2023-11-02 |
| EP4517746A4 (en) | 2025-07-02 |
| US20250273203A1 (en) | 2025-08-28 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US11727914B2 (en) | Intent recognition and emotional text-to-speech learning | |
| JP6714607B2 (ja) | 音声を要約するための方法、コンピュータ・プログラムおよびコンピュータ・システム | |
| US10089974B2 (en) | Speech recognition and text-to-speech learning system | |
| Dhanjal et al. | An automatic machine translation system for multi-lingual speech to Indian sign language | |
| Rekimoto | WESPER: Zero-shot and realtime whisper to normal voice conversion for whisper-based speech interactions | |
| US12387711B2 (en) | Speech synthesis device and speech synthesis method | |
| KR102628211B1 (ko) | 전자 장치 및 그 제어 방법 | |
| KR101819459B1 (ko) | 음성 인식 오류 수정을 지원하는 음성 인식 시스템 및 장치 | |
| WO2020105349A1 (ja) | 情報処理装置および情報処理方法 | |
| US9028255B2 (en) | Method and system for acquisition of literacy | |
| JP2019208138A (ja) | 発話認識装置、及びコンピュータプログラム | |
| JP2020181022A (ja) | 会議支援装置、会議支援システム、および会議支援プログラム | |
| Fellbaum et al. | Principles of electronic speech processing with applications for people with disabilities | |
| Rekimoto | DualVoice: speech interaction that discriminates between normal and whispered voice input | |
| Pandey et al. | MELDER: The design and evaluation of a real-time silent speech recognizer for mobile devices | |
| JP2002344915A (ja) | コミュニケーション把握装置、および、その方法 | |
| KR102114365B1 (ko) | 음성인식 방법 및 장치 | |
| US20250273203A1 (en) | Information processing device, information processing method, and computer program | |
| Lin et al. | Predicting Turn-Taking and Backchannel in Human-Machine Conversations Using Linguistic, Acoustic, and Visual Signals | |
| EP3477634A1 (en) | Information processing device and information processing method | |
| WO2018079294A1 (ja) | 情報処理装置及び情報処理方法 | |
| Rekimoto | DualVoice: A speech interaction method using whisper-voice as commands | |
| CN113178187B (zh) | 一种语音处理方法、装置、设备及介质、程序产品 | |
| Venkatagiri | Speech recognition technology applications in communication disorders | |
| US20210082427A1 (en) | Information processing apparatus and information processing method |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 23795897 Country of ref document: EP Kind code of ref document: A1 |
|
| ENP | Entry into the national phase |
Ref document number: 2024517869 Country of ref document: JP Kind code of ref document: A |
|
| WWE | Wipo information: entry into national phase |
Ref document number: 202380034669.X Country of ref document: CN |
|
| WWE | Wipo information: entry into national phase |
Ref document number: 18858123 Country of ref document: US |
|
| WWE | Wipo information: entry into national phase |
Ref document number: 2023795897 Country of ref document: EP |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| ENP | Entry into the national phase |
Ref document number: 2023795897 Country of ref document: EP Effective date: 20241126 |
|
| WWP | Wipo information: published in national office |
Ref document number: 18858123 Country of ref document: US |