WO2023035397A1 - 一种语音识别方法、装置、设备及存储介质 - Google Patents
一种语音识别方法、装置、设备及存储介质 Download PDFInfo
- Publication number
- WO2023035397A1 WO2023035397A1 PCT/CN2021/129733 CN2021129733W WO2023035397A1 WO 2023035397 A1 WO2023035397 A1 WO 2023035397A1 CN 2021129733 W CN2021129733 W CN 2021129733W WO 2023035397 A1 WO2023035397 A1 WO 2023035397A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- speech
- speaker
- features
- target
- voice
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L15/00—Speech recognition
- G10L15/02—Feature extraction for speech recognition; Selection of recognition unit
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L15/00—Speech recognition
- G10L15/06—Creation of reference templates; Training of speech recognition systems, e.g. adaptation to the characteristics of the speaker's voice
- G10L15/063—Training
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L25/00—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00
- G10L25/48—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 specially adapted for particular use
- G10L25/51—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 specially adapted for particular use for comparison or discrimination
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L15/00—Speech recognition
- G10L15/06—Creation of reference templates; Training of speech recognition systems, e.g. adaptation to the characteristics of the speaker's voice
- G10L15/065—Adaptation
- G10L15/07—Adaptation to the speaker
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L17/00—Speaker identification or verification techniques
- G10L17/02—Preprocessing operations, e.g. segment selection; Pattern representation or modelling, e.g. based on linear discriminant analysis [LDA] or principal components; Feature selection or extraction
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L17/00—Speaker identification or verification techniques
- G10L17/04—Training, enrolment or model building
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L17/00—Speaker identification or verification techniques
- G10L17/18—Artificial neural networks; Connectionist approaches
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L25/00—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00
- G10L25/27—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 characterised by the analysis technique
- G10L25/30—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 characterised by the analysis technique using neural networks
Definitions
- the present application relates to the technical field of speech recognition, and in particular to a speech recognition method, device, equipment and storage medium.
- the voice collected by the smart device is a mixed voice.
- voice interaction in order to obtain a better user experience, it is necessary to identify the voice content of the target speaker from the mixed voice, and how to identify the voice content of the target speaker from the mixed voice is an urgent need to be solved question.
- the present application provides a speech recognition method, device, device and storage medium for more accurately recognizing the speech content of the target speaker from the mixed speech.
- the technical solution is as follows:
- a speech recognition method comprising:
- the target speech feature is a speech feature used to obtain a speech recognition result consistent with the target speaker's real speech content
- acquiring speaker characteristics of the target speaker includes:
- Short-term voiceprint features and long-term voiceprint features are extracted from the registered voice of the target speaker to obtain multi-scale voiceprint features as the speaker features of the target speaker.
- the extraction direction is toward the target speech feature
- the target is extracted from the speech feature of the target mixed speech according to the speech feature of the target mixed speech and the speaker feature of the target speaker.
- the speech characteristics of the speaker including:
- the feature extraction model adopts the voice features of the training mixed voice including the voice of the specified speaker and the speaker features of the specified speaker, and optimizes the speech recognition result obtained based on the extracted voice features of the specified speaker.
- the target training results in that the extracted speech features of the designated speaker are the speech features of the designated speaker extracted from the speech features of the training mixed speech.
- the feature extraction model is trained with the extracted speech features of the designated speaker and the speech recognition results obtained based on the extracted speech features of the designated speaker as optimization targets.
- the speech characteristics of the speaker including:
- the acquiring the speech recognition result of the target speaker according to the extracted speech features of the target speaker includes:
- the registered speech feature of the target speaker is the speech feature of the registered speech of the target speaker.
- the acquiring the speech recognition result of the target speaker according to the extracted speech features of the target speaker includes:
- the speech recognition model is jointly trained with the feature extraction model, the speech recognition model uses the extracted speech features of the specified speaker, and the speech recognition result obtained based on the extracted speech features of the specified speaker is the optimization target Get trained.
- inputting the speech recognition input features into the speech recognition model to obtain the speech recognition result of the target speaker including:
- the audio-related feature vector required for decoding at the time of decoding is extracted from the encoding result
- the decoder module based on the speech recognition model decodes the audio-related feature vector extracted from the encoding result to obtain the recognition result at the decoding moment.
- the joint training process of the speech recognition model and the feature extraction model includes:
- the training mixed voice corresponds to the voice of the designated speaker
- the parameter updating of the feature extraction model according to the extracted speech features of the designated speaker and the speech recognition result of the designated speaker, and the parameter updating of the speech recognition model according to the speech recognition result of the designated speaker include: :
- the training mixed voice and the voice of the specified speaker corresponding to the training mixed voice are obtained from a pre-built training data set;
- the construction process of the training data set includes:
- each voice is the voice of a single speaker, and each voice has annotated text
- the training data set is composed of all training data obtained.
- a speech recognition device comprising: a feature acquisition module, a feature extraction module and a speech recognition module;
- the feature acquisition module is used to acquire the speech features of the target mixed voice and the speaker features of the target speaker;
- the feature extraction module is used to extract the target voice feature as the extraction direction, and extract from the voice feature of the target mixed voice according to the voice feature of the target mixed voice and the speaker feature of the target speaker.
- the speech features of the target speaker to obtain the extracted speech features of the target speaker, wherein the target speech feature is a speech feature used to obtain a speech recognition result consistent with the real speech content of the target speaker ;
- the speech recognition module is configured to obtain a speech recognition result of the target speaker according to the extracted speech features of the target speaker.
- the feature acquisition module includes: a speaker feature acquisition module
- the speaker feature acquisition module is configured to acquire the registered voice of the target speaker, and extract short-term voiceprint features and long-term voiceprint features from the registered voice of the target speaker to obtain multi-scale voiceprint features, is the speaker feature of the target speaker.
- the feature extraction module is specifically configured to use a pre-established feature extraction model, based on the speech features of the target mixed voice and the speaker features of the target speaker, from the voice of the target mixed voice Extract the voice features of the target speaker from the features;
- the feature extraction model adopts the voice features of the training mixed voice including the voice of the specified speaker and the speaker features of the specified speaker, and uses the speech recognition results obtained based on the extracted voice features of the specified speaker as the optimization target training It is obtained that the extracted speech features of the specified speaker are the speech features of the specified speaker extracted from the speech features of the training mixed speech.
- the speech recognition module is specifically configured to obtain the speech recognition result of the target speaker according to the extracted speech features of the target speaker and the registered speech features of the target speaker;
- the registered speech feature of the target speaker is the speech feature of the registered speech of the target speaker.
- the speech recognition module is specifically configured to input speech recognition input features including at least the extracted speech features of the target speaker into a pre-established speech recognition model to obtain a speech recognition result of the target speaker;
- the speech recognition model is jointly trained with the feature extraction model, the speech recognition model uses the extracted speech features of the specified speaker, and the speech recognition result obtained based on the extracted speech features of the specified speaker is The optimization objective is trained to get.
- a speech recognition device comprising: a memory and a processor
- the memory is used to store programs
- the processor is configured to execute the program to implement each step of the speech recognition method described in any one of the above.
- a readable storage medium on which a computer program is stored, and when the computer program is executed by a processor, each step of the speech recognition method described in any one of the above is realized.
- the speech recognition method, device, device and storage medium provided by the present application can extract the target speaker from the speech characteristics of the target mixed speech according to the speech characteristics of the target mixed speech and the speaker characteristics of the target speaker.
- the voice features of the target speaker can be obtained according to the extracted voice features of the target speaker, and the speech recognition result of the target speaker can be obtained. Since the application extracts the voice features of the target speaker from the voice features of the target mixed voice, it tends to the target speaker.
- Voice feature for obtaining the voice feature of the voice recognition result consistent with the real voice content of the target target speaker
- the extracted voice feature is the target voice feature or the voice feature approaching the target voice feature
- Fig. 1 is a schematic flow chart of the voice recognition method provided by the embodiment of the present application.
- Fig. 2 is a schematic flow chart of the joint training of the feature extraction model and the speech recognition model provided by the embodiment of the present application;
- FIG. 3 is a schematic diagram of the joint training process of the feature extraction model and the speech recognition model provided by the embodiment of the present application;
- FIG. 4 is a schematic structural diagram of a voice recognition device provided in an embodiment of the present application.
- FIG. 5 is a schematic structural diagram of a speech recognition device provided by an embodiment of the present application.
- the initial idea is: first train a feature extraction model, and then train a speech recognition model; obtain the registered voice of the target speaker, and The registered voice of the target speaker is extracted from the d-vector as the speaker feature of the target speaker; based on the pre-trained feature extraction model, based on the speaker features of the target speaker and the voice features of the target mixed voice, the target mixed voice.
- the speech features of the target speaker are extracted from the speech features of the target speaker; the speech of the target speaker is obtained by performing a series of transformation processes on the speech features of the extracted target speaker; the speech of the target speaker is input into the pre-trained speech
- the recognition model performs speech recognition to obtain the speech recognition result of the target speaker.
- the applicant further conducted research, and finally proposed a speech recognition method that can perfectly overcome the above-mentioned defects.
- the speech content of the speaker can be applied to a terminal with data processing capabilities, the terminal can recognize the speech content of the target speaker from the target mixed speech according to the speech recognition method provided by this application, and the terminal can include a processing component , a memory, an input/output interface and a power supply component.
- the terminal may also include a multimedia component, an audio component, a sensor component, a communication component, and the like. in:
- the processing component is used for data processing, and it can perform speech synthesis processing in this case.
- the processing component may include one or more processors, and the processing component may also include one or more modules to facilitate interaction with other components.
- the memory is configured to store various types of data, and the memory can be implemented with any type of volatile or non-volatile memory device or a combination of them, such as static random access memory (SRAM), electrically erasable programmable memory One of read memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, optical disk, etc. or Various combinations.
- SRAM static random access memory
- EEPROM electrically erasable programmable memory
- EPROM erasable programmable read-only memory
- PROM programmable read-only memory
- ROM read-only memory
- magnetic memory flash memory
- flash memory magnetic disk
- optical disk etc.
- the power supply component provides power for various components of the terminal, and the power supply component may include a power management system, one or more power supplies, and the like.
- the multimedia component may include a screen.
- the screen may be a touch display screen, and the touch display screen may receive input signals from a user.
- the multimedia component may also include a front camera and/or a rear camera.
- the audio component is configured to output and/or input audio signals
- the audio component may include a microphone configured to receive an external audio signal
- the audio component may further include a speaker configured to output an audio signal
- the voice synthesized by the terminal may pass through Speaker output.
- the input/output interface is the interface between the processing component and the peripheral interface module.
- the peripheral interface module can be a keyboard, a button, etc., wherein the button can include but is not limited to a home button, a volume button, a start button, a lock button, etc.
- the sensor component may include one or more sensors for providing status assessment of various aspects of the terminal, for example, the sensor component may detect the open/closed state of the terminal, whether the user is in contact with the terminal, the orientation, speed, temperature, etc. of the device.
- the sensor component may include, but is not limited to, one or a combination of image sensors, acceleration sensors, gyroscope sensors, pressure sensors, temperature sensors, and the like.
- the communication component is configured to facilitate wired or wireless communication between the terminal and other devices.
- the terminal can access wireless networks based on communication standards, such as one or a combination of WiFi, 2G, 3G, 4G, and 5G.
- the terminal can be implemented by one or more Application Specific Integrated Circuits (ASICs), Digital Signal Processors (ASP), Digital Signal Processing Devices (DSPDs), Programmable Logic Devices (PLDs), Field Programmable Gate Arrays (FPGAs) ), a controller, a microcontroller, a microprocessor or other electronic components for implementing the simultaneous interpretation method provided in this application.
- ASICs Application Specific Integrated Circuits
- ASP Digital Signal Processor
- DSPDs Digital Signal Processing Devices
- PLDs Programmable Logic Devices
- FPGAs Field Programmable Gate Arrays
- the voice recognition method provided by this application can also be applied to the server, and the server can recognize the voice content of the target speaker from the target mixed voice according to the voice recognition method provided by this application.
- the server can be connected to the terminal through the network , the terminal acquires the target mixed voice, transmits the target mixed voice to the server through the network connected to the server, and the server recognizes the voice content of the target speaker from the target mixed voice according to the voice recognition method provided by this application, and then transmits the target speaker through the network
- the human voice content is transmitted to the terminal.
- the server may include one or more than one central processing unit and memory, wherein the memory is configured to store various types of data, and the memory may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Read Only Memory (PROM), Read Only Memory (ROM), One or more combinations of magnetic memory, flash memory, magnetic disk, optical disk, etc.
- the server may also include one or more power supplies, one or more wired network interfaces and/or one or more wireless network interfaces, one or more operating systems.
- FIG. 1 shows a schematic flow chart of a speech recognition method provided in an embodiment of the present application.
- the method may include:
- Step S101 Obtain the speech features of the target mixed speech and the speaker features of the target speaker.
- the target mixed voice is the voice of multiple speakers, which includes the voice of other speakers in addition to the voice of the target speaker. recognize the speech content of the target speaker.
- the process of obtaining the speech features of the target mixed speech includes: obtaining the feature vector (such as spectral features) of each speech frame in the target mixed speech to obtain the feature vector sequence, and using the obtained feature vector sequence as the speech feature of the target mixed speech .
- the target mixed speech includes K speech frames, and the feature vector of the kth speech frame is expressed as x k , then the speech features of the target mixed speech can be expressed as [x 1 ,x 2 ,...,x k ,...,x K ] .
- this embodiment provides the following two optional ways of realization:
- the registered voice of the target speaker can be obtained, and the target speaker can speak to the target speaker
- the d-vector is extracted from the registered voice of the person, and the extracted d-vector is used as the speaker feature of the target speaker; considering that the voiceprint information contained in the d-vector is relatively simple and not rich enough, in order to improve the effect of subsequent feature extraction, this embodiment
- Another preferred implementation method is provided, that is, to obtain the registered voice of the target speaker, and extract short-term voiceprint features and long-term voiceprint features from the registered voice of the target speaker to obtain multi-scale voiceprint features. Scale voiceprint features are used as speaker features of the target speaker.
- the speaker characteristics obtained through the second implementation above contain richer voiceprint information, which makes subsequent use of the speaker features obtained through the second implementation above Speaker features are used for feature extraction to obtain better feature extraction results.
- the process of extracting short-term voiceprint features and long-term voiceprint features from the registered voice of the target speaker may include: using a pre-established speaker representation extraction model to extract short-term voiceprint features and long-term voiceprint features from the registered voice of the target speaker.
- Temporal voiceprint features Specifically, the speech feature sequence of the registered speech of the target speaker is obtained, and the speech feature sequence of the registered speech of the target speaker is input into the pre-established speaker representation extraction model, and the short-term voiceprint feature and long-term voiceprint feature of the target speaker are obtained. striation feature.
- the speaker representation extraction model can use a convolutional neural network, and input the voice feature sequence of the registered speech of the target speaker into the convolutional neural network for feature extraction, so as to obtain shallow features and deep features, wherein the shallow features Because the smaller receptive field can better represent the short-term voiceprint, therefore, the shallow features are used as the short-term voiceprint features, and the deep features are better able to represent the long-term voiceprint due to the larger receptive field. Therefore, the deep features are used as the long-term voiceprint features. Voiceprint features.
- the speaker characterization extraction model in this embodiment is trained by using a large number of training voices with real speaker labels (the training voice here is preferably the voice of a single speaker), where the real speaker label of the training voice represents the training voice Voice corresponding to the speaker.
- the speaker representation extraction model may be trained using a Cross Entropy (Cross Entropy, CE) criterion or a Metric Learning (ML) criterion.
- Step S102 Take the target speech feature as the extraction direction, according to the speech feature of the target mixed speech and the speaker feature of the target speaker, extract the speech feature of the target speaker from the speech feature of the target mixed speech to obtain the target speaker Extract speech features.
- the target speech feature is a speech feature used to obtain a speech recognition result consistent with the real speech content of the target speaker.
- the target speech feature or the speech feature approaching the target speech feature can be extracted from the speech feature of the target mixed speech, that is, taking the approaching target speech feature as the extraction direction can extract the target speech feature from the target mixture
- Speech features that are beneficial to subsequent speech recognition are extracted from the speech features of the speech, and speech recognition is performed based on the speech features that are conducive to speech recognition, so that a better speech recognition effect can be obtained.
- the process of extracting speech features of people may include: using a pre-established feature extraction model, based on the target mixed speech features and the target speaker features, extracting the speech features of the specified speaker from the target mixed speech features to obtain the target speaker Extract speech features.
- the feature extraction model adopts the voice features of the training mixed voice including the voice of the specified speaker and the speaker features of the specified speaker, and uses the speech recognition results obtained based on the extracted voice features of the specified speaker as the optimization target training.
- the input of the feature extraction model is the speech features of the above-mentioned training mixed speech and the speaker features of the designated speaker
- the output is the speech features of the designated speaker extracted from the training mixed speech features.
- the speech recognition result obtained based on the extracted speech features of the specified speaker is used as the optimization goal.
- the speech recognition results obtained based on the extracted speech features of the specified speaker are used as the optimization goal to train the feature extraction model, so that the feature extraction model can extract the speech features that are beneficial to speech recognition from the mixed speech features.
- the extracted speech features of the specified speaker and the speech recognition results obtained based on the extracted speech features of the specified speaker are used as optimization goals.
- the extracted speech features of the specified speaker and the speech recognition results obtained based on the extracted speech features of the specified speaker are the optimization goals, so that the feature extraction model can extract speech features that are conducive to speech recognition and approach the specified speech from the mixed speech features.
- the speech characteristics of a person's standard speech characteristics It should be noted that the standard speech features of the target speaker refer to the speech features obtained from the speech (clean speech) of the specified speaker.
- Step S103 Obtain the speech recognition result of the target speaker according to the extracted speech features of the target speaker.
- the speech of the target speaker can be obtained only according to the extracted speech features of the target speaker Recognition results; in order to improve the speech recognition effect, in another possible implementation, it can be based on the target speaker's extracted speech features and the target speaker's registered speech features (the target speaker's registered speech features refer to the target speech The voice features of the registered voice of the person) to obtain the voice recognition result of the target speaker, wherein the registered voice features of the target speaker are used as recognition auxiliary information to improve the voice recognition effect.
- the pre-established speech recognition model can be used to obtain the speech recognition result of the target speaker. More specifically, the extracted speech features of the target speaker are used as speech recognition input features, or the extracted speech features of the target speaker and The registered speech features of the target speaker are used as speech recognition input features, and the speech recognition input features are input into the pre-established speech recognition model to obtain the speech recognition results of the target speaker.
- the extracted speech features of the target speaker and the registered speech features of the target speaker are input into the speech recognition model as speech recognition input features
- the registered speech features of the target speaker can be compared with the extracted speech features of the target speaker. If it is accurate, the auxiliary speech recognition model performs speech recognition, thereby improving the effect of speech recognition.
- the speech recognition model can be jointly trained with the feature extraction model, and the speech recognition model uses the above-mentioned "extracted speech features of the specified speaker" as a training sample, and optimizes the speech recognition result obtained based on the extracted speech features of the specified speaker. target training.
- the feature extraction model is jointly trained with the speech recognition model, so that the feature extraction model can be optimized in a direction that is conducive to speech recognition.
- the speech recognition method provided by the embodiment of the present application can extract the speech features of the target speaker from the speech features of the target mixed speech, and then can obtain the speech recognition result of the target speaker according to the extracted speech features of the target speaker.
- the speech features tending to the target speech feature (the speech feature used to obtain the speech recognition result consistent with the real speech content of the target target speaker) is Therefore, the extracted speech feature is the target speech feature or a speech feature close to the target speech feature, and speech recognition based on the speech feature can obtain a better speech recognition effect, that is, a more accurate speech recognition can be obtained As a result, the user experience is better.
- the above-mentioned embodiment mentions that the feature extraction model used to extract the speech features of the target speaker from the speech features of the target mixed speech, and the speech recognition model used to obtain the speech recognition result of the target speaker according to the features extracted by the feature extraction model , which can be obtained through joint training.
- This embodiment focuses on the joint training process of the feature extraction model and the speech recognition model.
- the joint training process of the feature extraction model and the speech recognition model may include:
- Step S201 Obtain training mixed speech s m from the pre-built training data set S.
- the training data set S includes multiple pieces of training data, each piece of training data includes the voice (clean voice) of the specified speaker, and the training mixed voice comprising the voice of the specified speaker, wherein the voice of the specified speaker has Annotated text (annotated text is the voice content of the specified speaker's voice).
- the construction process of the training data set S includes:
- Step a1 acquiring multiple voices from multiple speakers.
- Each of the multiple voices obtained in this step is the voice of a single speaker, and each voice has a marked text.
- the marked text of the voice is " ⁇ s>, today, fill, day, gas, no, wrong, ⁇ /s>", where " ⁇ s>” is the sentence start character, and " ⁇ /s>” is the sentence end character.
- Step a2 using part of the voices or each voice in all the voices as the voice of the designated speaker: mixing one or more voices of other speakers in other voices with the voice of the designated speaker to A training mixed voice is obtained, and the training mixed voice and the voice of the designated speaker are used as a piece of training data.
- the acquired multiple voices include a voice of speaker a, a voice of speaker b, a voice of speaker c and a voice of speaker d, where each voice is of a single speaker Clean speech
- the speech of speaker a can be used as the speech of the specified speaker
- the speech of other speakers one or more speakers
- the speech of person b is mixed with the speech of speaker a, or, the speech of speaker b, the speech of speaker c is mixed with the speech of speaker a
- the speech of speaker a is mixed with the speech of speaker a and the speech of other speakers
- the training mixed speech obtained by mixing human speech is used as a piece of training data.
- the speech of speaker b can be used as the speech of the designated speaker, and the speech of other speakers (one or more speakers) can be combined with the speech of speaker b.
- Speech mixing to obtain a training mixed voice using speaker b's voice and the training mixed voice obtained by mixing speaker b's voice with other speakers' voices as a piece of training data, and multiple pieces of training data can be obtained in this way .
- the speech of the designated speaker includes K speech frames, that is, the length of the speech of the designated speaker is K, if the length of the speech of other speakers is greater than K, then the K+1th speech in the speech of other speakers can be Frame and the following speech frames are deleted, that is, only the first K speech frames are kept. If the length of other speakers' speech is less than K, assuming it is L, then K-L speech frames are copied from the front to supplement.
- Step a3 forming a training data set from all the obtained training data.
- Step S202 Obtain the speech features of the training mixed speech s m as the training mixed speech features X m , and obtain the speaker features of the specified speaker as the training speaker features.
- a speaker representation extraction model can be established in advance, and speaker features can be extracted from the registered speech of a specified speaker by using the pre-established speaker representation extraction model, and the extracted speaker features can be used as training speaker features.
- the speaker representation extraction model 300 is used to extract short-term voiceprint features and long-term voiceprint features from the registered voice of the specified speaker, and the extracted short-term voiceprint features and long-term voiceprint features are used as the designated speaker speaker characteristics.
- the speaker representation extraction model is pre-trained before the joint training of the feature extraction model and the speech recognition model. During the joint training stage of the feature extraction model and speech recognition model, its parameters are fixed and do not change Parameter update with speech recognition model.
- Step S203 Using the feature extraction model, based on the training mixed speech feature X m and the training speaker feature, extract the speech feature of the designated speaker from the training mixed speech feature X m , as the extracted speech feature of the designated speaker
- the training mixed speech feature X m and the training speaker feature are input into the feature extraction model to obtain the feature mask M corresponding to the specified speaker, and then according to the feature mask M corresponding to the specified speaker, from the training mixed speech feature X Extract the speech features of the specified speaker in m as the extracted speech features of the specified speaker
- the training mixed speech feature X m and the training speaker feature are input into the feature extraction model 301, and the feature extraction model 301 determines the feature mask corresponding to the specified speaker according to the input training mixed speech feature X m and the training speaker feature Code M and output.
- the feature extraction model 301 in this embodiment may be, but not limited to, a recurrent neural network (Recurrent Neural Network, RNN), a convolutional neural network (Convolution Neural Network, CNN), a deep neural network (Deep Neural Network, DNN) and the like.
- the training mixed speech feature X m is the feature vector sequence [x m1 ,x m2 ,...,x mk ,...,x mK ] (K is the training mixed speech total number of speech frames), when training the mixed speech feature X m and the training speaker feature input feature extraction model 301, the training speaker feature can be spliced with the feature vector of each speech frame in the training mixed speech, after splicing Input feature extraction model 301 .
- the feature vector of each voice frame in the training mixed voice is 40 dimensions, and the short-term voiceprint feature and the long-term voiceprint feature in the training speaker feature are both 40 dimensions, then each voice in the training mixed voice
- a 120-dimensional spliced feature vector can be obtained.
- the combination of short-term voiceprint features and long-term voiceprint features increases the richness of the input information, which makes the feature extraction model better extract the speech features of the specified speaker.
- the feature mask M corresponding to the specified speaker can represent the proportion of the voice features of the specified speaker in the training mixed voice features X m .
- the training mixed speech feature X m is expressed as [x m1 , x m2 ,...,x mk ,...,x mK ]
- the feature mask M corresponding to the specified speaker is expressed as [m 1 ,m 2 ,..., m k ,...,m K ]
- m 1 represents the proportion of the speech features of the specified speaker in x m1
- m 2 represents the proportion of the speech features of the specified speaker in x m2
- m K represents x Proportion of speech features of the specified speaker in mK
- m 1 ⁇ m K are all values between [0,1].
- the training mixed speech feature X m is multiplied frame by frame by the feature mask M corresponding to the specified speaker, and the specified feature extracted from the training mixed speech feature X m can be obtained speaker's phonetic features
- Step S204 Extract the speech features of the specified speaker Input the speech recognition model to obtain the speech recognition result of the specified speaker
- the registered voice features of the specified speaker (the registered voice features of the specified speaker refer to the voice features of the registered voice of the specified speaker)
- Xe [x e1 , x e2 ,...,x ek ,...,x eK ], except that the extracted speech features of the specified speaker
- the registered speech feature X e of the designated speaker is also input into the speech recognition model, and the speech recognition model is assisted by the registered speech feature X e of the designated speaker.
- the speech recognition model in this embodiment may include: an encoder module, an attention module, and a decoder module. in:
- the input to the encoder module consists of the extracted speech features of the specified speaker
- two encoding modules can be set in the encoder module, as shown in Figure 3, the first encoding module 3021 is set in the encoder module and the second encoding module 3022, wherein the first encoding module 3021 is used to extract speech features of the designated speaker Encoding, the second encoding module is used to encode the registered speech features X e of the designated speaker; in another possible implementation, an encoding module is set in the encoder module to extract the speech features of the designated speaker
- the encoding operation of and the encoding operation of the registered speech feature X e of the designated speaker are both performed by this encoding module, that is, the two encoding processes share one encoding module.
- each encoding module can include one or more encoding layers, and the encoding layer can adopt a long-short-term memory layer in a one-way or two-way long-short-term memory neural network , or use the convolutional layer of a convolutional neural network.
- Attention module for extracting speech features from specified speakers respectively
- the audio-related feature vector required for decoding at the time of decoding is extracted from the encoding result H x of the specified speaker and the encoding result of the registered speech feature X e of the specified speaker.
- the decoding module is used to decode the audio-related feature vector extracted by the attention module, so as to obtain the recognition result at the decoding moment.
- the attention module 3023 is based on the attention mechanism, and at each decoding moment, respectively, from Extracting the current _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ Audio-related feature vectors required at the decoding instant.
- the extracted audio-related feature vector represents the audio content of the character to be decoded at the t-th decoding moment.
- the attention mechanism refers to using a vector as a query item (query), performing an attention mechanism operation on a set of feature vector sequences, and selecting the feature vector that best matches the query item as an output, specifically, Calculate a matching coefficient between the query item and each feature vector in the feature vector sequence, and then multiply and sum these matching coefficients with the corresponding feature vectors to obtain a new feature vector that is the feature vector that best matches the query item.
- the state feature vector d t of the decoder module 3024 is determined according to the recognition result y t-1 at the t-1th decoding moment and c t-1 x and c t-1 e output by the attention module.
- the decoder module 3024 may include multiple neural network layers, for example, two layers of unidirectional long-short-term memory layers.
- the recognition result y t-1 at one decoding moment and the c t-1 x and c t-1 e output by the attention module 3023 are used as input to calculate the state feature vector d t of the decoder, and d t is input to the attention module 3023 , used to calculate c t x and c t e at the tth decoding moment, and then c t x and c t e are spliced, and the spliced vector is used as the input of the second layer of long short-term memory layer of the decoder module 3024 (for example, Both c t x and c t e are 128-dimensional vectors, splicing c t x and c t e can obtain a 256-dimensional spliced vector, and the 256-dimensional spliced vector is input to the second layer of the long short-term memory layer of the decoder module 3024) , calculate the output h t
- Step S205 Extract speech features according to the specified speaker and speech recognition results for the specified speaker Update the parameters of the feature extraction model, and based on the speech recognition results of the specified speaker Update the parameters of the speech recognition model.
- step S205 may include:
- Step S2051 obtain the labeled text T t of the specified speaker's voice st (voice of the specified speaker) corresponding to the training mixed voice s m , and obtain the voice feature of the specified speaker's voice st as the standard voice feature X of the specified speaker t .
- Step S2052 extracting speech features according to the specified speaker and the standard speech feature X t of the specified speaker to determine the first prediction loss Loss1, and according to the speech recognition result of the specified speaker and the labeled text T t of the specified speaker's voice st to determine the second prediction loss Loss2.
- the extracted speech features for a given speaker can be computed The minimum mean square error with the standard speech feature Xt of the specified speaker, as the first prediction loss Loss1, according to the speech recognition result of the specified speaker Compute the cross-entropy loss with the annotated text T t of the specified speaker's voice st as the second prediction loss.
- Step S2053 update the parameters of the feature extraction model according to the first prediction loss Loss1 and the second prediction loss Loss2, and update the parameters of the speech recognition model according to the second prediction loss Loss2.
- the parameters of the feature extraction model are updated so that the feature extraction model can extract the speech that is close to the standard speech features of the specified speaker and is conducive to speech recognition from the training mixed speech features feature, the speech feature is input into the speech recognition model for speech recognition, and better speech recognition effect can be obtained.
- this embodiment is specific to the first embodiment of "using the pre-established feature extraction model, based on the target mixed voice features and the target speaker features, to extract the specified The speech features of the speaker, and the process of extracting the speech features of the target speaker" is introduced.
- extracting the voice features of the specified speaker from the target mixed voice features, so as to obtain the process of extracting voice features of the target speaker may include:
- Step b1 Input the speech features of the target mixed speech and the speaker features of the target speaker into the feature extraction model to obtain a feature mask corresponding to the target speaker.
- the feature mask corresponding to the target speaker can represent the proportion of the target speaker's voice features in the voice features of the target mixed voice.
- Step b2 Extract the speech features of the target speaker from the speech features of the target mixed speech according to the feature mask corresponding to the target speaker, so as to obtain the extracted speech features of the target speaker.
- the speech features of the target mixed speech are multiplied frame by frame by the feature mask corresponding to the target speaker, so as to obtain the extracted speech features of the target speaker.
- the process of obtaining the speech recognition result of the target speaker may include:
- Step c1 The encoder module based on the speech recognition model encodes the extracted speech features of the target speaker and the registered speech features of the target speaker respectively to obtain two encoding results.
- Step c2 Based on the attention module of the speech recognition model, the audio-related feature vectors required for decoding at the time of decoding are respectively extracted from the two encoding results.
- the decoder module based on the speech recognition model decodes the audio-related feature vectors extracted from the two encoding results respectively to obtain the recognition result at the decoding moment.
- the process of inputting the extracted speech features of the target speaker into the speech recognition model to obtain the speech recognition result of the target speaker is the same as the process of combining the extracted speech features of the designated speaker and the registered speech features of the designated speaker in the training stage.
- the implementation process of inputting the speech recognition model and obtaining the speech recognition result of the designated speaker is similar.
- the specific implementation process of steps c1 to c3 can be found in the introduction of the encoder module, attention module and decoder module in the second embodiment. The embodiment will not be repeated here.
- the speech recognition method provided by the present application has the following advantages: First, the present application extracts multi-scale voiceprint features from the registered speech of the target speaker and inputs the feature extraction model, Increase the richness of the input information of the feature extraction model, and improve the feature extraction effect of the feature extraction model; second, the joint training of the feature extraction model and the speech recognition model enables the prediction loss of the speech recognition model to act on the feature extraction model, thereby making the feature
- the extraction model can extract speech features that are beneficial to speech recognition, thereby improving the accuracy of speech recognition results; third, the speech features of the target speaker's registered speech are used as additional input to the speech recognition model, and the features extracted by the feature extraction model When the speech characteristics are not good, it can assist the speech recognition model to perform speech recognition, so as to obtain more accurate speech recognition results. To sum up, the speech recognition method provided by the present application can accurately recognize the speech content of the target speaker in the case of complex human voice interference.
- the embodiment of the present application also provides a speech recognition device.
- the speech recognition device provided in the embodiment of the present application is described below.
- the speech recognition device described below and the speech recognition method described above can be referred to in correspondence.
- FIG. 4 shows a schematic structural diagram of a speech recognition device provided by an embodiment of the present application, which may include: a feature acquisition module 401 , a feature extraction module 402 and a speech recognition module 403 .
- the feature acquisition module 401 is configured to acquire speech features of the target mixed voice and speaker features of the target speaker.
- the feature extraction module 402 is used to extract the target speech features from the speech features of the target mixed speech according to the speech features of the target mixed speech and the speaker features of the target speaker. Speech features of the target speaker to obtain the extracted speech features of the target speaker, wherein the target speech feature is a speech feature used to obtain a speech recognition result consistent with the real speech content of the target speaker.
- the speech recognition module 403 is configured to obtain a speech recognition result of the target speaker according to the extracted speech features of the target speaker.
- the feature acquisition module 401 includes: a voice feature acquisition module and a speaker feature acquisition module.
- the voice feature acquisition module is used to acquire the voice features of the target mixed voice.
- the speaker feature acquisition module is used to acquire speaker features of the target speaker.
- the speaker feature acquisition module acquires the speaker features of the target speaker
- it is specifically configured to acquire the registered voice of the target speaker, and extract short-term voiceprint features from the registered voice of the target speaker and long-term voiceprint features to obtain multi-scale voiceprint features as the speaker features of the target speaker.
- the feature extraction module 402 is specifically configured to use a pre-established feature extraction model, based on the speech features of the target mixed speech and the speaker features of the target speaker, from the speech features of the target mixed speech Extract the speech features of the target speaker.
- the feature extraction model adopts the voice features of the training mixed voice including the voice of the specified speaker and the speaker features of the specified speaker, and uses the speech recognition results obtained based on the extracted voice features of the specified speaker as the optimization target training It is obtained that the extracted speech features of the specified speaker are the speech features of the specified speaker extracted from the speech features of the training mixed speech.
- the feature extraction model is trained with the extracted speech features of the designated speaker and the speech recognition results obtained based on the extracted speech features of the designated speaker as optimization targets.
- the feature extraction module 402 may include: a feature mask determination submodule and a speech feature extraction submodule.
- the feature mask determination submodule is configured to input the speech features of the target mixed speech and the speaker features of the target speaker into the feature extraction model to obtain a feature mask corresponding to the target speaker, wherein , the feature mask can represent the proportion of the speech features of the corresponding speaker in the speech features of the target mixed speech.
- the speech feature extraction submodule is used to extract the speech features of the target speaker from the speech features of the target mixed speech according to the speech features of the target mixed speech and the feature mask corresponding to the target speaker .
- the speech recognition module 403 is specifically configured to obtain the speech recognition result of the target speaker according to the extracted speech features of the target speaker and the registered speech features of the target speaker; wherein, the target speaker The registered speech feature of the person is the speech feature of the registered speech of the target speaker.
- the speech recognition module 403 is specifically configured to input speech recognition input features including at least the extracted speech features of the target speaker into a pre-established speech recognition model to obtain a speech recognition result of the target speaker.
- the speech recognition model is jointly trained with the feature extraction model, the speech recognition model uses the extracted speech features of the specified speaker, and the speech recognition result obtained based on the extracted speech features of the specified speaker is The optimization objective is trained to get.
- the speech recognition module 403 is specifically used to:
- the speech recognition input features are encoded to obtain a coding result; based on the attention module of the speech recognition model, the information required for decoding at the time of decoding is extracted from the coding result
- An audio-related feature vector the decoder module based on the speech recognition model decodes the audio-related feature vector extracted from the encoding result to obtain a recognition result at the decoding moment.
- the speech recognition device may further include: a model training module.
- the model training module may include: an acquisition module for extracting speech features, an acquisition module for speech recognition results, and a parameter update module.
- the extracted speech feature acquisition module is configured to use a feature extraction model to extract the speech features of the specified speaker from the speech features of the training mixed speech, so as to obtain the extracted speech features of the specified speaker.
- the speech recognition result obtaining module is used to obtain the speech recognition result of the designated speaker by using the speech recognition model and the extracted speech features of the designated speaker.
- the model update module is used to update the parameters of the feature extraction model according to the extracted speech features of the designated speaker and the speech recognition results of the designated speaker, and perform speech recognition according to the speech recognition results of the designated speaker.
- the model is updated with parameters.
- the model update module may include: an annotation text acquisition module, a standard speech feature acquisition module, a prediction loss determination module, and a parameter update module.
- the training mixed speech corresponds to the speech of the designated speaker.
- the standard speech feature acquisition module is used to obtain the speech feature of the voice of the designated speaker as the standard speech feature of the designated speaker.
- the annotation text obtaining module is used to obtain the annotation text of the voice of the designated speaker.
- the prediction loss determining module is configured to determine a first prediction loss according to the extracted speech features of the specified speaker and the standard speech features of the specified speaker, and to determine the first prediction loss according to the speech recognition result of the specified speaker and the specified Annotated text of the speaker's speech, determining a second prediction loss.
- the parameter update module is configured to update the parameters of the feature extraction model according to the first prediction loss and the second prediction loss, and update the parameters of the speech recognition model according to the second prediction loss.
- the training mixed voice and the voice of the designated speaker corresponding to the training mixed voice are obtained from a pre-built training data set, and the voice recognition device provided in the embodiment of the present application may further include: building a training data set module.
- the training dataset building blocks are used to:
- each voice is the voice of a single speaker, and each voice has annotated text; taking part of the multiple voices or each voice in all the voices as a specified utterance Human voice: Mix one or more voices of other speakers in other voices with the voice of the specified speaker to obtain a training mixed voice, and use the voice of the specified speaker and the training mixed voice obtained by mixing as A piece of training data; the training data set is composed of all obtained training data.
- the speech recognition device can extract the speech features of the target speaker from the speech features of the target mixed speech, and then can obtain the speech recognition result of the target speaker according to the extracted speech features of the target speaker.
- the speech features tending to the target speech feature (the speech feature used to obtain the speech recognition result consistent with the real speech content of the target target speaker) is Therefore, the extracted speech feature is the target speech feature or a speech feature close to the target speech feature, and speech recognition based on the speech feature can obtain a better speech recognition effect, that is, a more accurate speech recognition can be obtained As a result, the user experience is better.
- the embodiment of the present application also provides a speech recognition device.
- FIG. 5, shows a schematic structural diagram of the speech recognition device.
- the speech recognition device may include: at least one processor 501, at least one communication interface 502, at least one memory 503 and at least one communication bus 504;
- the number of processor 501, communication interface 502, memory 503, and communication bus 504 is at least one, and the processor 501, communication interface 502, and memory 503 complete mutual communication through the communication bus 504;
- Processor 501 may be a central processing unit CPU, or a specific integrated circuit ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiments of the present invention, etc.;
- ASIC Application Specific Integrated Circuit
- the memory 503 may include a high-speed RAM memory, and may also include a non-volatile memory (non-volatile memory), such as at least one disk memory;
- the memory stores a program
- the processor can call the program stored in the memory, and the program is used for:
- the speech feature of the target speaker is extracted from the speech feature of the target mixed speech to obtain the extracted speech of the target speaker feature, wherein the target speech feature is a speech feature used to obtain a speech recognition result consistent with the target speaker's real speech content;
- the embodiment of the present application also provides a readable storage medium, which can store a program suitable for execution by a processor, and the program is used for:
- the speech feature of the target speaker is extracted from the speech feature of the target mixed speech to obtain the extracted speech of the target speaker feature, wherein the target speech feature is a speech feature used to obtain a speech recognition result consistent with the target speaker's real speech content;
Landscapes
- Engineering & Computer Science (AREA)
- Multimedia (AREA)
- Health & Medical Sciences (AREA)
- Audiology, Speech & Language Pathology (AREA)
- Human Computer Interaction (AREA)
- Physics & Mathematics (AREA)
- Acoustics & Sound (AREA)
- Computational Linguistics (AREA)
- Computer Vision & Pattern Recognition (AREA)
- Artificial Intelligence (AREA)
- Signal Processing (AREA)
- Evolutionary Computation (AREA)
- Image Analysis (AREA)
- Information Retrieval, Db Structures And Fs Structures Therefor (AREA)
- Electrically Operated Instructional Devices (AREA)
- Machine Translation (AREA)
- Telephonic Communication Services (AREA)
Abstract
Description
Claims (18)
- 一种语音识别方法,其特征在于,包括:获取目标混合语音的语音特征以及目标说话人的说话人特征;以趋于目标语音特征为提取方向,根据所述目标混合语音的语音特征以及所述目标说话人的说话人特征,从所述目标混合语音的语音特征中提取所述目标说话人的语音特征,以得到所述目标说话人的提取语音特征,其中,所述目标语音特征为用于获得与所述目标说话人的真实语音内容一致的语音识别结果的语音特征;根据所述目标说话人的提取语音特征,获取所述目标说话人的语音识别结果。
- 根据权利要求1所述的语音识别方法,其特征在于,获取所述目标说话人的说话人特征,包括:获取所述目标说话人的注册语音;对所述目标说话人的注册语音提取短时声纹特征和长时声纹特征,以得到多尺度声纹特征,作为所述目标说话人的说话人特征。
- 根据权利要求1所述的语音识别方法,其特征在于,所述以趋于目标语音特征为提取方向,根据所述目标混合语音的语音特征以及所述目标说话人的说话人特征,从所述目标混合语音的语音特征中提取所述目标说话人的语音特征,包括:利用预先建立的特征提取模型,以所述目标混合语音的语音特征以及所述目标说话人的说话人特征为依据,从所述目标混合语音的语音特征中提取所述目标说话人的语音特征;其中,所述特征提取模型采用包含指定说话人的语音的训练混合语音的语音特征和所述指定说话人的说话人特征,以基于所述指定说话人的提取语音特征获取的语音识别结果为优化目标训练得到,所述指定说话人的提取语音特征为从所述训练混合语音的语音特征中提取的所述指定说话人的语音特征。
- 根据权利要求3所述的语音识别方法,其特征在于,所述特征提取模型同时以所述指定说话人的提取语音特征和基于所述指定说话人的提取语音特征获取的语音识别结果为优化目标训练得到。
- 根据权利要求3或4所述的语音识别方法,其特征在于,所述利用预 先建立的特征提取模型,以所述目标混合语音的语音特征以及所述目标说话人的说话人特征为依据,从所述目标混合语音的语音特征中提取所述目标说话人的语音特征,包括:将所述目标混合语音的语音特征以及所述目标说话人的说话人特征输入所述特征提取模型,得到所述目标说话人对应的特征掩码;根据所述目标混合语音的语音特征和所述目标说话人对应的特征掩码,从所述目标混合语音的语音特征中提取所述目标说话人的语音特征。
- 根据权利要求1所述的语音识别方法,其特征在于,所述根据所述目标说话人的提取语音特征,获取所述目标说话人的语音识别结果,包括:根据所述目标说话人的提取语音特征以及所述目标说话人的注册语音特征,获取所述目标说话人的语音识别结果;其中,所述目标说话人的注册语音特征为所述目标说话人的注册语音的语音特征。
- 根据权利要求3或4所述的语音识别方法,其特征在于,所述根据所述目标说话人的提取语音特征,获取所述目标说话人的语音识别结果,包括:将至少包括所述目标说话人的提取语音特征的语音识别输入特征输入预先建立的语音识别模型,得到所述目标说话人的语音识别结果;所述语音识别模型与所述特征提取模型联合训练得到,所述语音识别模型采用所述指定说话人的提取语音特征,以基于所述指定说话人的提取语音特征获得的语音识别结果为优化目标训练得到。
- 根据权利要求7所述的语音识别方法,其特征在于,将所述语音识别输入特征输入所述语音识别模型,得到所述目标说话人的语音识别结果,包括:基于所述语音识别模型的编码器模块,对所述语音识别输入特征进行编码,以得到编码结果;基于所述语音识别模型的注意力模块,从所述编码结果中提取解码时刻解码所需的音频相关特征向量;基于所述语音识别模型的解码器模块,对从所述编码结果中提取的所述音频相关特征向量进行解码,得到所述解码时刻的识别结果。
- 根据权利要求7所述的语音识别方法,其特征在于,所述语音识别模型与所述特征提取模型联合训练的过程包括:利用特征提取模型,从所述训练混合语音的语音特征中提取所述指定说话人的语音特征,以得到所述指定说话人的提取语音特征;利用语音识别模型和所述指定说话人的提取语音特征,获取所述指定说话人的语音识别结果;根据所述指定说话人的提取语音特征和所述指定说话人的语音识别结果对特征提取模型进行参数更新,并根据所述指定说话人的语音识别结果对语音识别模型进行参数更新。
- 根据权利要求9所述的语音识别方法,其特征在于,所述训练混合语音对应有所述指定说话人的语音;所述根据所述指定说话人的提取语音特征和所述指定说话人的语音识别结果对特征提取模型进行参数更新,并根据所述指定说话人的语音识别结果对语音识别模型进行参数更新,包括:获取所述指定说话人的语音的标注文本,并获取所述指定说话人的语音的语音特征作为所述指定说话人的标准语音特征;根据所述指定说话人的提取语音特征和所述指定说话人的标准语音特征确定第一预测损失,并根据所述指定说话人的语音识别结果和所述指定说话人的语音的标注文本,确定第二预测损失;根据所述第一预测损失和所述第二预测损失对特征提取模型进行参数更新,并根据所述第二预测损失对语音识别模型进行参数更新。
- 根据权利要求10所述的语音识别方法,其特征在于,所述训练混合语音以及所述训练混合语音对应的所述指定说话人的语音从预先构建的训练数据集中获取;所述训练数据集的构建过程包括:获取多个说话人的多条语音,其中,每条语音为单一说话人的语音,每条语音具有标注文本;将所述多条语音中的部分语音或全部语音中的每条语音作为指定说话人的语音:将其它语音中其他说话人的一条或多条语音与该指定说话人的语音进行混合,以得到一条训练混合语音,将该指定说话人的语音与通过混合得到的训练混合语音作为一条训练数据;由获得的所有训练数据组成所述训练数据集。
- 一种语音识别装置,其特征在于,包括:特征获取模块、特征提取模块和语音识别模块;所述特征获取模块,用于获取目标混合语音的语音特征以及目标说话人的说话人特征;所述特征提取模块,用于以趋于目标语音特征为提取方向,根据所述目标混合语音的语音特征以及所述目标说话人的说话人特征,从所述目标混合语音的语音特征中提取所述目标说话人的语音特征,以得到所述目标说话人的提取语音特征,其中,所述目标语音特征为用于获得与所述目标说话人的真实语音内容一致的语音识别结果的语音特征;所述语音识别模块,用于根据所述目标说话人的提取语音特征,获取所述目标说话人的语音识别结果。
- 根据权利要求12所述的语音识别装置,其特征在于,所述特征获取模块包括:说话人特征获取模块;所述说话人特征获取模块,用于获取所述目标说话人的注册语音,对所述目标说话人的注册语音提取短时声纹特征和长时声纹特征,以得到多尺度声纹特征,作为所述目标说话人的说话人特征。
- 根据权利要求12所述的语音识别装置,其特征在于,所述特征提取模块具体用于利用预先建立的特征提取模型,以所述目标混合语音的语音特征以及所述目标说话人的说话人特征为依据,从所述目标混合语音的语音特征中提取所述目标说话人的语音特征;其中,所述特征提取模型采用包含指定说话人的语音的训练混合语音的语音特征和指定说话人的说话人特征,以基于所述指定说话人的提取语音特征获取的语音识别结果为优化目标训练得到,所述指定说话人的提取语音特征为从所述训练混合语音的语音特征中提取的所述指定说话人的语音特征。
- 根据权利要求12所述的语音识别装置,其特征在于,所述语音识别模块,具体用于根据所述目标说话人的提取语音特征以及所述目标说话人的注册语音特征,获取所述目标说话人的语音识别结果;其中,所述目标说话人的注册语音特征为所述目标说话人的注册语音的语音特征。
- 根据权利要求14所述的语音识别装置,其特征在于,所述语音识别 模块,用于将至少包括所述目标说话人的提取语音特征的语音识别输入特征输入预先建立的语音识别模型,得到所述目标说话人的语音识别结果;其中,所述语音识别模型与所述特征提取模型联合训练得到,所述语音识别模型采用所述指定说话人的提取语音特征,以基于所述指定说话人的提取语音特征获得的语音识别结果为优化目标训练得到。
- 一种语音识别设备,其特征在于,包括:存储器和处理器;所述存储器,用于存储程序;所述处理器,用于执行所述程序,实现如权利要求1~11中任一项所述的语音识别方法的各个步骤。
- 一种计算机可读存储介质,其上存储有计算机程序,其特征在于,所述计算机程序被处理器执行时,实现如权利要求1~11中任一项所述的语音识别方法的各个步骤。
Priority Applications (4)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US18/689,668 US12626687B2 (en) | 2021-09-07 | 2021-11-10 | Speech recognition method, apparatus and device, and storage medium |
| EP21956569.4A EP4401074A4 (en) | 2021-09-07 | 2021-11-10 | SPEECH RECOGNITION METHOD, APPARATUS AND DEVICE, AND STORAGE MEDIUM |
| JP2024514680A JP7786691B2 (ja) | 2021-09-07 | 2021-11-10 | 音声認識方法、装置、設備及び記憶媒体 |
| KR1020247011082A KR20240050447A (ko) | 2021-09-07 | 2021-11-10 | 음성 인식 방법, 장치, 디바이스 및 저장매체 |
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN202111042821.8 | 2021-09-07 | ||
| CN202111042821.8A CN113724713B (zh) | 2021-09-07 | 2021-09-07 | 一种语音识别方法、装置、设备及存储介质 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2023035397A1 true WO2023035397A1 (zh) | 2023-03-16 |
Family
ID=78682155
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2021/129733 Ceased WO2023035397A1 (zh) | 2021-09-07 | 2021-11-10 | 一种语音识别方法、装置、设备及存储介质 |
Country Status (6)
| Country | Link |
|---|---|
| US (1) | US12626687B2 (zh) |
| EP (1) | EP4401074A4 (zh) |
| JP (1) | JP7786691B2 (zh) |
| KR (1) | KR20240050447A (zh) |
| CN (1) | CN113724713B (zh) |
| WO (1) | WO2023035397A1 (zh) |
Cited By (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN117198272A (zh) * | 2023-11-07 | 2023-12-08 | 浙江同花顺智能科技有限公司 | 一种语音处理方法、装置、电子设备及存储介质 |
| CN119229875A (zh) * | 2024-09-04 | 2024-12-31 | 武汉大学 | 一种基于多参考线索融合的目标语音提取方法及装置 |
Families Citing this family (7)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN114898756B (zh) * | 2022-04-18 | 2025-08-15 | 北京荣耀终端有限公司 | 语音分离方法及装置 |
| CN116486818A (zh) * | 2022-08-30 | 2023-07-25 | 重庆蚂蚁消费金融有限公司 | 基于语音的身份识别方法、装置以及电子设备 |
| CN116978359A (zh) * | 2022-11-30 | 2023-10-31 | 腾讯科技(深圳)有限公司 | 音素识别方法、装置、电子设备及存储介质 |
| CN115713939B (zh) * | 2023-01-06 | 2023-04-21 | 阿里巴巴达摩院(杭州)科技有限公司 | 语音识别方法、装置及电子设备 |
| CN116312503A (zh) * | 2023-02-22 | 2023-06-23 | 哲库科技(上海)有限公司 | 语音数据的识别方法、装置、芯片及电子设备 |
| CN116403603B (zh) * | 2023-04-28 | 2025-09-05 | 科大讯飞股份有限公司 | 一种假音检测方法、假音检测模型获取方法及相关设备 |
| CN119851670B (zh) * | 2025-01-10 | 2025-09-23 | 中国科学技术大学 | 一种基于内容相关的帧级说话人声纹建模的语音匿名化方法 |
Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN109874096A (zh) * | 2019-01-17 | 2019-06-11 | 天津大学 | 一种基于智能终端选择输出的双耳麦克风助听器降噪算法 |
| CN110288989A (zh) * | 2019-06-03 | 2019-09-27 | 安徽兴博远实信息科技有限公司 | 语音交互方法及系统 |
| CN110827853A (zh) * | 2019-11-11 | 2020-02-21 | 广州国音智能科技有限公司 | 语音特征信息提取方法、终端及可读存储介质 |
| CN111128197A (zh) * | 2019-12-25 | 2020-05-08 | 北京邮电大学 | 基于声纹特征与生成对抗学习的多说话人语音分离方法 |
| CN111433847A (zh) * | 2019-12-31 | 2020-07-17 | 深圳市优必选科技股份有限公司 | 语音转换的方法及训练方法、智能装置和存储介质 |
Family Cites Families (29)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN111081231B (zh) * | 2016-03-23 | 2023-09-05 | 谷歌有限责任公司 | 用于多声道语音识别的自适应音频增强 |
| KR102596430B1 (ko) * | 2016-08-31 | 2023-10-31 | 삼성전자주식회사 | 화자 인식에 기초한 음성 인식 방법 및 장치 |
| US11133011B2 (en) * | 2017-03-13 | 2021-09-28 | Mitsubishi Electric Research Laboratories, Inc. | System and method for multichannel end-to-end speech recognition |
| US10699698B2 (en) | 2018-03-29 | 2020-06-30 | Tencent Technology (Shenzhen) Company Limited | Adaptive permutation invariant training with auxiliary information for monaural multi-talker speech recognition |
| US10957337B2 (en) * | 2018-04-11 | 2021-03-23 | Microsoft Technology Licensing, Llc | Multi-microphone speech separation |
| US10811000B2 (en) * | 2018-04-13 | 2020-10-20 | Mitsubishi Electric Research Laboratories, Inc. | Methods and systems for recognizing simultaneous speech by multiple speakers |
| CN109166586B (zh) * | 2018-08-02 | 2023-07-07 | 平安科技(深圳)有限公司 | 一种识别说话人的方法及终端 |
| US11475898B2 (en) * | 2018-10-26 | 2022-10-18 | Apple Inc. | Low-latency multi-speaker speech recognition |
| CN109326302B (zh) * | 2018-11-14 | 2022-11-08 | 桂林电子科技大学 | 一种基于声纹比对和生成对抗网络的语音增强方法 |
| US10923111B1 (en) * | 2019-03-28 | 2021-02-16 | Amazon Technologies, Inc. | Speech detection and speech recognition |
| EP4047596B1 (en) * | 2019-06-04 | 2025-02-19 | Google LLC | Two-pass end to end speech recognition |
| CN112331181B (zh) * | 2019-07-30 | 2024-07-05 | 中国科学院声学研究所 | 一种基于多说话人条件下目标说话人语音提取方法 |
| US20210065712A1 (en) * | 2019-08-31 | 2021-03-04 | Soundhound, Inc. | Automotive visual speech recognition |
| JP7329393B2 (ja) * | 2019-09-02 | 2023-08-18 | 日本電信電話株式会社 | 音声信号処理装置、音声信号処理方法、音声信号処理プログラム、学習装置、学習方法及び学習プログラム |
| CN110517698B (zh) * | 2019-09-05 | 2022-02-01 | 科大讯飞股份有限公司 | 一种声纹模型的确定方法、装置、设备及存储介质 |
| CN111145736B (zh) * | 2019-12-09 | 2022-10-04 | 华为技术有限公司 | 语音识别方法及相关设备 |
| CN111009237B (zh) * | 2019-12-12 | 2022-07-01 | 北京达佳互联信息技术有限公司 | 语音识别方法、装置、电子设备及存储介质 |
| CN111261146B (zh) | 2020-01-16 | 2022-09-09 | 腾讯科技(深圳)有限公司 | 语音识别及模型训练方法、装置和计算机可读存储介质 |
| CN111326143B (zh) * | 2020-02-28 | 2022-09-06 | 科大讯飞股份有限公司 | 语音处理方法、装置、设备及存储介质 |
| CN111508505B (zh) * | 2020-04-28 | 2023-11-03 | 讯飞智元信息科技有限公司 | 一种说话人识别方法、装置、设备及存储介质 |
| CN111583916B (zh) * | 2020-05-19 | 2023-07-25 | 科大讯飞股份有限公司 | 一种语音识别方法、装置、设备及存储介质 |
| CN111899727B (zh) * | 2020-07-15 | 2022-05-06 | 思必驰科技股份有限公司 | 用于多说话人的语音识别模型的训练方法及系统 |
| CN111833886B (zh) * | 2020-07-27 | 2021-03-23 | 中国科学院声学研究所 | 全连接多尺度的残差网络及其进行声纹识别的方法 |
| CN111899758B (zh) * | 2020-09-07 | 2024-01-30 | 腾讯科技(深圳)有限公司 | 语音处理方法、装置、设备和存储介质 |
| CN112735390B (zh) * | 2020-12-25 | 2023-02-28 | 江西台德智慧科技有限公司 | 一种具有语音识别功能的智能语音终端设备 |
| CN112599118B (zh) * | 2020-12-30 | 2024-02-13 | 中国科学技术大学 | 语音识别方法、装置、电子设备和存储介质 |
| CN112786057B (zh) * | 2021-02-23 | 2023-06-02 | 厦门熵基科技有限公司 | 一种声纹识别方法、装置、电子设备及存储介质 |
| CN113077795B (zh) * | 2021-04-06 | 2022-07-15 | 重庆邮电大学 | 一种通道注意力传播与聚合下的声纹识别方法 |
| CN113221673B (zh) * | 2021-04-25 | 2024-03-19 | 华南理工大学 | 基于多尺度特征聚集的说话人认证方法及系统 |
-
2021
- 2021-09-07 CN CN202111042821.8A patent/CN113724713B/zh active Active
- 2021-11-10 JP JP2024514680A patent/JP7786691B2/ja active Active
- 2021-11-10 EP EP21956569.4A patent/EP4401074A4/en active Pending
- 2021-11-10 KR KR1020247011082A patent/KR20240050447A/ko active Pending
- 2021-11-10 WO PCT/CN2021/129733 patent/WO2023035397A1/zh not_active Ceased
- 2021-11-10 US US18/689,668 patent/US12626687B2/en active Active
Patent Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN109874096A (zh) * | 2019-01-17 | 2019-06-11 | 天津大学 | 一种基于智能终端选择输出的双耳麦克风助听器降噪算法 |
| CN110288989A (zh) * | 2019-06-03 | 2019-09-27 | 安徽兴博远实信息科技有限公司 | 语音交互方法及系统 |
| CN110827853A (zh) * | 2019-11-11 | 2020-02-21 | 广州国音智能科技有限公司 | 语音特征信息提取方法、终端及可读存储介质 |
| CN111128197A (zh) * | 2019-12-25 | 2020-05-08 | 北京邮电大学 | 基于声纹特征与生成对抗学习的多说话人语音分离方法 |
| CN111433847A (zh) * | 2019-12-31 | 2020-07-17 | 深圳市优必选科技股份有限公司 | 语音转换的方法及训练方法、智能装置和存储介质 |
Non-Patent Citations (1)
| Title |
|---|
| See also references of EP4401074A4 |
Cited By (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN117198272A (zh) * | 2023-11-07 | 2023-12-08 | 浙江同花顺智能科技有限公司 | 一种语音处理方法、装置、电子设备及存储介质 |
| CN117198272B (zh) * | 2023-11-07 | 2024-01-30 | 浙江同花顺智能科技有限公司 | 一种语音处理方法、装置、电子设备及存储介质 |
| CN119229875A (zh) * | 2024-09-04 | 2024-12-31 | 武汉大学 | 一种基于多参考线索融合的目标语音提取方法及装置 |
Also Published As
| Publication number | Publication date |
|---|---|
| US20240395242A1 (en) | 2024-11-28 |
| CN113724713A (zh) | 2021-11-30 |
| JP2024530353A (ja) | 2024-08-16 |
| KR20240050447A (ko) | 2024-04-18 |
| EP4401074A4 (en) | 2025-07-23 |
| US12626687B2 (en) | 2026-05-12 |
| JP7786691B2 (ja) | 2025-12-16 |
| EP4401074A1 (en) | 2024-07-17 |
| CN113724713B (zh) | 2024-07-05 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| JP7786691B2 (ja) | 音声認識方法、装置、設備及び記憶媒体 | |
| US20230127787A1 (en) | Method and apparatus for converting voice timbre, method and apparatus for training model, device and medium | |
| CN110956959B (zh) | 语音识别纠错方法、相关设备及可读存储介质 | |
| CN113362813B (zh) | 一种语音识别方法、装置和电子设备 | |
| WO2020107878A1 (zh) | 文本摘要生成方法、装置、计算机设备及存储介质 | |
| CN108305617B (zh) | 语音关键词的识别方法和装置 | |
| CN110516253B (zh) | 中文口语语义理解方法及系统 | |
| CN108170686B (zh) | 文本翻译方法及装置 | |
| CN108899013B (zh) | 语音搜索方法、装置和语音识别系统 | |
| CN109003601A (zh) | 一种针对低资源土家语的跨语言端到端语音识别方法 | |
| WO2019196196A1 (zh) | 一种耳语音恢复方法、装置、设备及可读存储介质 | |
| CN114694255B (zh) | 基于通道注意力与时间卷积网络的句子级唇语识别方法 | |
| CN110428820A (zh) | 一种中英文混合语音识别方法及装置 | |
| CN108630199A (zh) | 一种声学模型的数据处理方法 | |
| CN111161724B (zh) | 中文视听结合语音识别方法、系统、设备及介质 | |
| CN107221330A (zh) | 标点添加方法和装置、用于标点添加的装置 | |
| CN115394287A (zh) | 混合语种语音识别方法、装置、系统及存储介质 | |
| CN115240712A (zh) | 一种基于多模态的情感分类方法、装置、设备及存储介质 | |
| WO2020238045A1 (zh) | 智能语音识别方法、装置及计算机可读存储介质 | |
| CN113889087B (zh) | 语音识别及模型建立方法、装置、设备和存储介质 | |
| CN112017643A (zh) | 语音识别模型训练方法、语音识别方法及相关装置 | |
| CN114373443A (zh) | 语音合成方法和装置、计算设备、存储介质及程序产品 | |
| CN116564330A (zh) | 弱监督语音预训练方法、电子设备和存储介质 | |
| CN109979461A (zh) | 一种语音翻译方法及装置 | |
| CN116416966A (zh) | 文本到语音合成方法、装置、设备和存储介质 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 21956569 Country of ref document: EP Kind code of ref document: A1 |
|
| WWE | Wipo information: entry into national phase |
Ref document number: 2024514680 Country of ref document: JP |
|
| WWE | Wipo information: entry into national phase |
Ref document number: 18689668 Country of ref document: US |
|
| ENP | Entry into the national phase |
Ref document number: 20247011082 Country of ref document: KR Kind code of ref document: A |
|
| WWE | Wipo information: entry into national phase |
Ref document number: 2021956569 Country of ref document: EP |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| ENP | Entry into the national phase |
Ref document number: 2021956569 Country of ref document: EP Effective date: 20240408 |