WO2023035397A1 - 一种语音识别方法、装置、设备及存储介质 - Google Patents

一种语音识别方法、装置、设备及存储介质 Download PDF

Info

Publication number
WO2023035397A1
WO2023035397A1 PCT/CN2021/129733 CN2021129733W WO2023035397A1 WO 2023035397 A1 WO2023035397 A1 WO 2023035397A1 CN 2021129733 W CN2021129733 W CN 2021129733W WO 2023035397 A1 WO2023035397 A1 WO 2023035397A1
Authority
WO
WIPO (PCT)
Prior art keywords
speech
speaker
features
target
voice
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/CN2021/129733
Other languages
English (en)
French (fr)
Inventor
方昕
刘俊华
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
iFlytek Co Ltd
Original Assignee
iFlytek Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by iFlytek Co Ltd filed Critical iFlytek Co Ltd
Priority to US18/689,668 priority Critical patent/US12626687B2/en
Priority to EP21956569.4A priority patent/EP4401074A4/en
Priority to JP2024514680A priority patent/JP7786691B2/ja
Priority to KR1020247011082A priority patent/KR20240050447A/ko
Publication of WO2023035397A1 publication Critical patent/WO2023035397A1/zh
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Images

Classifications

    • G—PHYSICS
    • G10—MUSICAL INSTRUMENTS; ACOUSTICS
    • G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L15/00—Speech recognition
    • G10L15/02—Feature extraction for speech recognition; Selection of recognition unit
    • G—PHYSICS
    • G10—MUSICAL INSTRUMENTS; ACOUSTICS
    • G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L15/00—Speech recognition
    • G10L15/06—Creation of reference templates; Training of speech recognition systems, e.g. adaptation to the characteristics of the speaker's voice
    • G10L15/063—Training
    • G—PHYSICS
    • G10—MUSICAL INSTRUMENTS; ACOUSTICS
    • G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L25/00—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00
    • G10L25/48—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 specially adapted for particular use
    • G10L25/51—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 specially adapted for particular use for comparison or discrimination
    • G—PHYSICS
    • G10—MUSICAL INSTRUMENTS; ACOUSTICS
    • G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L15/00—Speech recognition
    • G10L15/06—Creation of reference templates; Training of speech recognition systems, e.g. adaptation to the characteristics of the speaker's voice
    • G10L15/065—Adaptation
    • G10L15/07—Adaptation to the speaker
    • G—PHYSICS
    • G10—MUSICAL INSTRUMENTS; ACOUSTICS
    • G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L17/00—Speaker identification or verification techniques
    • G10L17/02—Preprocessing operations, e.g. segment selection; Pattern representation or modelling, e.g. based on linear discriminant analysis [LDA] or principal components; Feature selection or extraction
    • G—PHYSICS
    • G10—MUSICAL INSTRUMENTS; ACOUSTICS
    • G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L17/00—Speaker identification or verification techniques
    • G10L17/04—Training, enrolment or model building
    • G—PHYSICS
    • G10—MUSICAL INSTRUMENTS; ACOUSTICS
    • G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L17/00—Speaker identification or verification techniques
    • G10L17/18—Artificial neural networks; Connectionist approaches
    • G—PHYSICS
    • G10—MUSICAL INSTRUMENTS; ACOUSTICS
    • G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L25/00—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00
    • G10L25/27—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 characterised by the analysis technique
    • G10L25/30—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 characterised by the analysis technique using neural networks

Definitions

  • the present application relates to the technical field of speech recognition, and in particular to a speech recognition method, device, equipment and storage medium.
  • the voice collected by the smart device is a mixed voice.
  • voice interaction in order to obtain a better user experience, it is necessary to identify the voice content of the target speaker from the mixed voice, and how to identify the voice content of the target speaker from the mixed voice is an urgent need to be solved question.
  • the present application provides a speech recognition method, device, device and storage medium for more accurately recognizing the speech content of the target speaker from the mixed speech.
  • the technical solution is as follows:
  • a speech recognition method comprising:
  • the target speech feature is a speech feature used to obtain a speech recognition result consistent with the target speaker's real speech content
  • acquiring speaker characteristics of the target speaker includes:
  • Short-term voiceprint features and long-term voiceprint features are extracted from the registered voice of the target speaker to obtain multi-scale voiceprint features as the speaker features of the target speaker.
  • the extraction direction is toward the target speech feature
  • the target is extracted from the speech feature of the target mixed speech according to the speech feature of the target mixed speech and the speaker feature of the target speaker.
  • the speech characteristics of the speaker including:
  • the feature extraction model adopts the voice features of the training mixed voice including the voice of the specified speaker and the speaker features of the specified speaker, and optimizes the speech recognition result obtained based on the extracted voice features of the specified speaker.
  • the target training results in that the extracted speech features of the designated speaker are the speech features of the designated speaker extracted from the speech features of the training mixed speech.
  • the feature extraction model is trained with the extracted speech features of the designated speaker and the speech recognition results obtained based on the extracted speech features of the designated speaker as optimization targets.
  • the speech characteristics of the speaker including:
  • the acquiring the speech recognition result of the target speaker according to the extracted speech features of the target speaker includes:
  • the registered speech feature of the target speaker is the speech feature of the registered speech of the target speaker.
  • the acquiring the speech recognition result of the target speaker according to the extracted speech features of the target speaker includes:
  • the speech recognition model is jointly trained with the feature extraction model, the speech recognition model uses the extracted speech features of the specified speaker, and the speech recognition result obtained based on the extracted speech features of the specified speaker is the optimization target Get trained.
  • inputting the speech recognition input features into the speech recognition model to obtain the speech recognition result of the target speaker including:
  • the audio-related feature vector required for decoding at the time of decoding is extracted from the encoding result
  • the decoder module based on the speech recognition model decodes the audio-related feature vector extracted from the encoding result to obtain the recognition result at the decoding moment.
  • the joint training process of the speech recognition model and the feature extraction model includes:
  • the training mixed voice corresponds to the voice of the designated speaker
  • the parameter updating of the feature extraction model according to the extracted speech features of the designated speaker and the speech recognition result of the designated speaker, and the parameter updating of the speech recognition model according to the speech recognition result of the designated speaker include: :
  • the training mixed voice and the voice of the specified speaker corresponding to the training mixed voice are obtained from a pre-built training data set;
  • the construction process of the training data set includes:
  • each voice is the voice of a single speaker, and each voice has annotated text
  • the training data set is composed of all training data obtained.
  • a speech recognition device comprising: a feature acquisition module, a feature extraction module and a speech recognition module;
  • the feature acquisition module is used to acquire the speech features of the target mixed voice and the speaker features of the target speaker;
  • the feature extraction module is used to extract the target voice feature as the extraction direction, and extract from the voice feature of the target mixed voice according to the voice feature of the target mixed voice and the speaker feature of the target speaker.
  • the speech features of the target speaker to obtain the extracted speech features of the target speaker, wherein the target speech feature is a speech feature used to obtain a speech recognition result consistent with the real speech content of the target speaker ;
  • the speech recognition module is configured to obtain a speech recognition result of the target speaker according to the extracted speech features of the target speaker.
  • the feature acquisition module includes: a speaker feature acquisition module
  • the speaker feature acquisition module is configured to acquire the registered voice of the target speaker, and extract short-term voiceprint features and long-term voiceprint features from the registered voice of the target speaker to obtain multi-scale voiceprint features, is the speaker feature of the target speaker.
  • the feature extraction module is specifically configured to use a pre-established feature extraction model, based on the speech features of the target mixed voice and the speaker features of the target speaker, from the voice of the target mixed voice Extract the voice features of the target speaker from the features;
  • the feature extraction model adopts the voice features of the training mixed voice including the voice of the specified speaker and the speaker features of the specified speaker, and uses the speech recognition results obtained based on the extracted voice features of the specified speaker as the optimization target training It is obtained that the extracted speech features of the specified speaker are the speech features of the specified speaker extracted from the speech features of the training mixed speech.
  • the speech recognition module is specifically configured to obtain the speech recognition result of the target speaker according to the extracted speech features of the target speaker and the registered speech features of the target speaker;
  • the registered speech feature of the target speaker is the speech feature of the registered speech of the target speaker.
  • the speech recognition module is specifically configured to input speech recognition input features including at least the extracted speech features of the target speaker into a pre-established speech recognition model to obtain a speech recognition result of the target speaker;
  • the speech recognition model is jointly trained with the feature extraction model, the speech recognition model uses the extracted speech features of the specified speaker, and the speech recognition result obtained based on the extracted speech features of the specified speaker is The optimization objective is trained to get.
  • a speech recognition device comprising: a memory and a processor
  • the memory is used to store programs
  • the processor is configured to execute the program to implement each step of the speech recognition method described in any one of the above.
  • a readable storage medium on which a computer program is stored, and when the computer program is executed by a processor, each step of the speech recognition method described in any one of the above is realized.
  • the speech recognition method, device, device and storage medium provided by the present application can extract the target speaker from the speech characteristics of the target mixed speech according to the speech characteristics of the target mixed speech and the speaker characteristics of the target speaker.
  • the voice features of the target speaker can be obtained according to the extracted voice features of the target speaker, and the speech recognition result of the target speaker can be obtained. Since the application extracts the voice features of the target speaker from the voice features of the target mixed voice, it tends to the target speaker.
  • Voice feature for obtaining the voice feature of the voice recognition result consistent with the real voice content of the target target speaker
  • the extracted voice feature is the target voice feature or the voice feature approaching the target voice feature
  • Fig. 1 is a schematic flow chart of the voice recognition method provided by the embodiment of the present application.
  • Fig. 2 is a schematic flow chart of the joint training of the feature extraction model and the speech recognition model provided by the embodiment of the present application;
  • FIG. 3 is a schematic diagram of the joint training process of the feature extraction model and the speech recognition model provided by the embodiment of the present application;
  • FIG. 4 is a schematic structural diagram of a voice recognition device provided in an embodiment of the present application.
  • FIG. 5 is a schematic structural diagram of a speech recognition device provided by an embodiment of the present application.
  • the initial idea is: first train a feature extraction model, and then train a speech recognition model; obtain the registered voice of the target speaker, and The registered voice of the target speaker is extracted from the d-vector as the speaker feature of the target speaker; based on the pre-trained feature extraction model, based on the speaker features of the target speaker and the voice features of the target mixed voice, the target mixed voice.
  • the speech features of the target speaker are extracted from the speech features of the target speaker; the speech of the target speaker is obtained by performing a series of transformation processes on the speech features of the extracted target speaker; the speech of the target speaker is input into the pre-trained speech
  • the recognition model performs speech recognition to obtain the speech recognition result of the target speaker.
  • the applicant further conducted research, and finally proposed a speech recognition method that can perfectly overcome the above-mentioned defects.
  • the speech content of the speaker can be applied to a terminal with data processing capabilities, the terminal can recognize the speech content of the target speaker from the target mixed speech according to the speech recognition method provided by this application, and the terminal can include a processing component , a memory, an input/output interface and a power supply component.
  • the terminal may also include a multimedia component, an audio component, a sensor component, a communication component, and the like. in:
  • the processing component is used for data processing, and it can perform speech synthesis processing in this case.
  • the processing component may include one or more processors, and the processing component may also include one or more modules to facilitate interaction with other components.
  • the memory is configured to store various types of data, and the memory can be implemented with any type of volatile or non-volatile memory device or a combination of them, such as static random access memory (SRAM), electrically erasable programmable memory One of read memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, optical disk, etc. or Various combinations.
  • SRAM static random access memory
  • EEPROM electrically erasable programmable memory
  • EPROM erasable programmable read-only memory
  • PROM programmable read-only memory
  • ROM read-only memory
  • magnetic memory flash memory
  • flash memory magnetic disk
  • optical disk etc.
  • the power supply component provides power for various components of the terminal, and the power supply component may include a power management system, one or more power supplies, and the like.
  • the multimedia component may include a screen.
  • the screen may be a touch display screen, and the touch display screen may receive input signals from a user.
  • the multimedia component may also include a front camera and/or a rear camera.
  • the audio component is configured to output and/or input audio signals
  • the audio component may include a microphone configured to receive an external audio signal
  • the audio component may further include a speaker configured to output an audio signal
  • the voice synthesized by the terminal may pass through Speaker output.
  • the input/output interface is the interface between the processing component and the peripheral interface module.
  • the peripheral interface module can be a keyboard, a button, etc., wherein the button can include but is not limited to a home button, a volume button, a start button, a lock button, etc.
  • the sensor component may include one or more sensors for providing status assessment of various aspects of the terminal, for example, the sensor component may detect the open/closed state of the terminal, whether the user is in contact with the terminal, the orientation, speed, temperature, etc. of the device.
  • the sensor component may include, but is not limited to, one or a combination of image sensors, acceleration sensors, gyroscope sensors, pressure sensors, temperature sensors, and the like.
  • the communication component is configured to facilitate wired or wireless communication between the terminal and other devices.
  • the terminal can access wireless networks based on communication standards, such as one or a combination of WiFi, 2G, 3G, 4G, and 5G.
  • the terminal can be implemented by one or more Application Specific Integrated Circuits (ASICs), Digital Signal Processors (ASP), Digital Signal Processing Devices (DSPDs), Programmable Logic Devices (PLDs), Field Programmable Gate Arrays (FPGAs) ), a controller, a microcontroller, a microprocessor or other electronic components for implementing the simultaneous interpretation method provided in this application.
  • ASICs Application Specific Integrated Circuits
  • ASP Digital Signal Processor
  • DSPDs Digital Signal Processing Devices
  • PLDs Programmable Logic Devices
  • FPGAs Field Programmable Gate Arrays
  • the voice recognition method provided by this application can also be applied to the server, and the server can recognize the voice content of the target speaker from the target mixed voice according to the voice recognition method provided by this application.
  • the server can be connected to the terminal through the network , the terminal acquires the target mixed voice, transmits the target mixed voice to the server through the network connected to the server, and the server recognizes the voice content of the target speaker from the target mixed voice according to the voice recognition method provided by this application, and then transmits the target speaker through the network
  • the human voice content is transmitted to the terminal.
  • the server may include one or more than one central processing unit and memory, wherein the memory is configured to store various types of data, and the memory may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Read Only Memory (PROM), Read Only Memory (ROM), One or more combinations of magnetic memory, flash memory, magnetic disk, optical disk, etc.
  • the server may also include one or more power supplies, one or more wired network interfaces and/or one or more wireless network interfaces, one or more operating systems.
  • FIG. 1 shows a schematic flow chart of a speech recognition method provided in an embodiment of the present application.
  • the method may include:
  • Step S101 Obtain the speech features of the target mixed speech and the speaker features of the target speaker.
  • the target mixed voice is the voice of multiple speakers, which includes the voice of other speakers in addition to the voice of the target speaker. recognize the speech content of the target speaker.
  • the process of obtaining the speech features of the target mixed speech includes: obtaining the feature vector (such as spectral features) of each speech frame in the target mixed speech to obtain the feature vector sequence, and using the obtained feature vector sequence as the speech feature of the target mixed speech .
  • the target mixed speech includes K speech frames, and the feature vector of the kth speech frame is expressed as x k , then the speech features of the target mixed speech can be expressed as [x 1 ,x 2 ,...,x k ,...,x K ] .
  • this embodiment provides the following two optional ways of realization:
  • the registered voice of the target speaker can be obtained, and the target speaker can speak to the target speaker
  • the d-vector is extracted from the registered voice of the person, and the extracted d-vector is used as the speaker feature of the target speaker; considering that the voiceprint information contained in the d-vector is relatively simple and not rich enough, in order to improve the effect of subsequent feature extraction, this embodiment
  • Another preferred implementation method is provided, that is, to obtain the registered voice of the target speaker, and extract short-term voiceprint features and long-term voiceprint features from the registered voice of the target speaker to obtain multi-scale voiceprint features. Scale voiceprint features are used as speaker features of the target speaker.
  • the speaker characteristics obtained through the second implementation above contain richer voiceprint information, which makes subsequent use of the speaker features obtained through the second implementation above Speaker features are used for feature extraction to obtain better feature extraction results.
  • the process of extracting short-term voiceprint features and long-term voiceprint features from the registered voice of the target speaker may include: using a pre-established speaker representation extraction model to extract short-term voiceprint features and long-term voiceprint features from the registered voice of the target speaker.
  • Temporal voiceprint features Specifically, the speech feature sequence of the registered speech of the target speaker is obtained, and the speech feature sequence of the registered speech of the target speaker is input into the pre-established speaker representation extraction model, and the short-term voiceprint feature and long-term voiceprint feature of the target speaker are obtained. striation feature.
  • the speaker representation extraction model can use a convolutional neural network, and input the voice feature sequence of the registered speech of the target speaker into the convolutional neural network for feature extraction, so as to obtain shallow features and deep features, wherein the shallow features Because the smaller receptive field can better represent the short-term voiceprint, therefore, the shallow features are used as the short-term voiceprint features, and the deep features are better able to represent the long-term voiceprint due to the larger receptive field. Therefore, the deep features are used as the long-term voiceprint features. Voiceprint features.
  • the speaker characterization extraction model in this embodiment is trained by using a large number of training voices with real speaker labels (the training voice here is preferably the voice of a single speaker), where the real speaker label of the training voice represents the training voice Voice corresponding to the speaker.
  • the speaker representation extraction model may be trained using a Cross Entropy (Cross Entropy, CE) criterion or a Metric Learning (ML) criterion.
  • Step S102 Take the target speech feature as the extraction direction, according to the speech feature of the target mixed speech and the speaker feature of the target speaker, extract the speech feature of the target speaker from the speech feature of the target mixed speech to obtain the target speaker Extract speech features.
  • the target speech feature is a speech feature used to obtain a speech recognition result consistent with the real speech content of the target speaker.
  • the target speech feature or the speech feature approaching the target speech feature can be extracted from the speech feature of the target mixed speech, that is, taking the approaching target speech feature as the extraction direction can extract the target speech feature from the target mixture
  • Speech features that are beneficial to subsequent speech recognition are extracted from the speech features of the speech, and speech recognition is performed based on the speech features that are conducive to speech recognition, so that a better speech recognition effect can be obtained.
  • the process of extracting speech features of people may include: using a pre-established feature extraction model, based on the target mixed speech features and the target speaker features, extracting the speech features of the specified speaker from the target mixed speech features to obtain the target speaker Extract speech features.
  • the feature extraction model adopts the voice features of the training mixed voice including the voice of the specified speaker and the speaker features of the specified speaker, and uses the speech recognition results obtained based on the extracted voice features of the specified speaker as the optimization target training.
  • the input of the feature extraction model is the speech features of the above-mentioned training mixed speech and the speaker features of the designated speaker
  • the output is the speech features of the designated speaker extracted from the training mixed speech features.
  • the speech recognition result obtained based on the extracted speech features of the specified speaker is used as the optimization goal.
  • the speech recognition results obtained based on the extracted speech features of the specified speaker are used as the optimization goal to train the feature extraction model, so that the feature extraction model can extract the speech features that are beneficial to speech recognition from the mixed speech features.
  • the extracted speech features of the specified speaker and the speech recognition results obtained based on the extracted speech features of the specified speaker are used as optimization goals.
  • the extracted speech features of the specified speaker and the speech recognition results obtained based on the extracted speech features of the specified speaker are the optimization goals, so that the feature extraction model can extract speech features that are conducive to speech recognition and approach the specified speech from the mixed speech features.
  • the speech characteristics of a person's standard speech characteristics It should be noted that the standard speech features of the target speaker refer to the speech features obtained from the speech (clean speech) of the specified speaker.
  • Step S103 Obtain the speech recognition result of the target speaker according to the extracted speech features of the target speaker.
  • the speech of the target speaker can be obtained only according to the extracted speech features of the target speaker Recognition results; in order to improve the speech recognition effect, in another possible implementation, it can be based on the target speaker's extracted speech features and the target speaker's registered speech features (the target speaker's registered speech features refer to the target speech The voice features of the registered voice of the person) to obtain the voice recognition result of the target speaker, wherein the registered voice features of the target speaker are used as recognition auxiliary information to improve the voice recognition effect.
  • the pre-established speech recognition model can be used to obtain the speech recognition result of the target speaker. More specifically, the extracted speech features of the target speaker are used as speech recognition input features, or the extracted speech features of the target speaker and The registered speech features of the target speaker are used as speech recognition input features, and the speech recognition input features are input into the pre-established speech recognition model to obtain the speech recognition results of the target speaker.
  • the extracted speech features of the target speaker and the registered speech features of the target speaker are input into the speech recognition model as speech recognition input features
  • the registered speech features of the target speaker can be compared with the extracted speech features of the target speaker. If it is accurate, the auxiliary speech recognition model performs speech recognition, thereby improving the effect of speech recognition.
  • the speech recognition model can be jointly trained with the feature extraction model, and the speech recognition model uses the above-mentioned "extracted speech features of the specified speaker" as a training sample, and optimizes the speech recognition result obtained based on the extracted speech features of the specified speaker. target training.
  • the feature extraction model is jointly trained with the speech recognition model, so that the feature extraction model can be optimized in a direction that is conducive to speech recognition.
  • the speech recognition method provided by the embodiment of the present application can extract the speech features of the target speaker from the speech features of the target mixed speech, and then can obtain the speech recognition result of the target speaker according to the extracted speech features of the target speaker.
  • the speech features tending to the target speech feature (the speech feature used to obtain the speech recognition result consistent with the real speech content of the target target speaker) is Therefore, the extracted speech feature is the target speech feature or a speech feature close to the target speech feature, and speech recognition based on the speech feature can obtain a better speech recognition effect, that is, a more accurate speech recognition can be obtained As a result, the user experience is better.
  • the above-mentioned embodiment mentions that the feature extraction model used to extract the speech features of the target speaker from the speech features of the target mixed speech, and the speech recognition model used to obtain the speech recognition result of the target speaker according to the features extracted by the feature extraction model , which can be obtained through joint training.
  • This embodiment focuses on the joint training process of the feature extraction model and the speech recognition model.
  • the joint training process of the feature extraction model and the speech recognition model may include:
  • Step S201 Obtain training mixed speech s m from the pre-built training data set S.
  • the training data set S includes multiple pieces of training data, each piece of training data includes the voice (clean voice) of the specified speaker, and the training mixed voice comprising the voice of the specified speaker, wherein the voice of the specified speaker has Annotated text (annotated text is the voice content of the specified speaker's voice).
  • the construction process of the training data set S includes:
  • Step a1 acquiring multiple voices from multiple speakers.
  • Each of the multiple voices obtained in this step is the voice of a single speaker, and each voice has a marked text.
  • the marked text of the voice is " ⁇ s>, today, fill, day, gas, no, wrong, ⁇ /s>", where " ⁇ s>” is the sentence start character, and " ⁇ /s>” is the sentence end character.
  • Step a2 using part of the voices or each voice in all the voices as the voice of the designated speaker: mixing one or more voices of other speakers in other voices with the voice of the designated speaker to A training mixed voice is obtained, and the training mixed voice and the voice of the designated speaker are used as a piece of training data.
  • the acquired multiple voices include a voice of speaker a, a voice of speaker b, a voice of speaker c and a voice of speaker d, where each voice is of a single speaker Clean speech
  • the speech of speaker a can be used as the speech of the specified speaker
  • the speech of other speakers one or more speakers
  • the speech of person b is mixed with the speech of speaker a, or, the speech of speaker b, the speech of speaker c is mixed with the speech of speaker a
  • the speech of speaker a is mixed with the speech of speaker a and the speech of other speakers
  • the training mixed speech obtained by mixing human speech is used as a piece of training data.
  • the speech of speaker b can be used as the speech of the designated speaker, and the speech of other speakers (one or more speakers) can be combined with the speech of speaker b.
  • Speech mixing to obtain a training mixed voice using speaker b's voice and the training mixed voice obtained by mixing speaker b's voice with other speakers' voices as a piece of training data, and multiple pieces of training data can be obtained in this way .
  • the speech of the designated speaker includes K speech frames, that is, the length of the speech of the designated speaker is K, if the length of the speech of other speakers is greater than K, then the K+1th speech in the speech of other speakers can be Frame and the following speech frames are deleted, that is, only the first K speech frames are kept. If the length of other speakers' speech is less than K, assuming it is L, then K-L speech frames are copied from the front to supplement.
  • Step a3 forming a training data set from all the obtained training data.
  • Step S202 Obtain the speech features of the training mixed speech s m as the training mixed speech features X m , and obtain the speaker features of the specified speaker as the training speaker features.
  • a speaker representation extraction model can be established in advance, and speaker features can be extracted from the registered speech of a specified speaker by using the pre-established speaker representation extraction model, and the extracted speaker features can be used as training speaker features.
  • the speaker representation extraction model 300 is used to extract short-term voiceprint features and long-term voiceprint features from the registered voice of the specified speaker, and the extracted short-term voiceprint features and long-term voiceprint features are used as the designated speaker speaker characteristics.
  • the speaker representation extraction model is pre-trained before the joint training of the feature extraction model and the speech recognition model. During the joint training stage of the feature extraction model and speech recognition model, its parameters are fixed and do not change Parameter update with speech recognition model.
  • Step S203 Using the feature extraction model, based on the training mixed speech feature X m and the training speaker feature, extract the speech feature of the designated speaker from the training mixed speech feature X m , as the extracted speech feature of the designated speaker
  • the training mixed speech feature X m and the training speaker feature are input into the feature extraction model to obtain the feature mask M corresponding to the specified speaker, and then according to the feature mask M corresponding to the specified speaker, from the training mixed speech feature X Extract the speech features of the specified speaker in m as the extracted speech features of the specified speaker
  • the training mixed speech feature X m and the training speaker feature are input into the feature extraction model 301, and the feature extraction model 301 determines the feature mask corresponding to the specified speaker according to the input training mixed speech feature X m and the training speaker feature Code M and output.
  • the feature extraction model 301 in this embodiment may be, but not limited to, a recurrent neural network (Recurrent Neural Network, RNN), a convolutional neural network (Convolution Neural Network, CNN), a deep neural network (Deep Neural Network, DNN) and the like.
  • the training mixed speech feature X m is the feature vector sequence [x m1 ,x m2 ,...,x mk ,...,x mK ] (K is the training mixed speech total number of speech frames), when training the mixed speech feature X m and the training speaker feature input feature extraction model 301, the training speaker feature can be spliced with the feature vector of each speech frame in the training mixed speech, after splicing Input feature extraction model 301 .
  • the feature vector of each voice frame in the training mixed voice is 40 dimensions, and the short-term voiceprint feature and the long-term voiceprint feature in the training speaker feature are both 40 dimensions, then each voice in the training mixed voice
  • a 120-dimensional spliced feature vector can be obtained.
  • the combination of short-term voiceprint features and long-term voiceprint features increases the richness of the input information, which makes the feature extraction model better extract the speech features of the specified speaker.
  • the feature mask M corresponding to the specified speaker can represent the proportion of the voice features of the specified speaker in the training mixed voice features X m .
  • the training mixed speech feature X m is expressed as [x m1 , x m2 ,...,x mk ,...,x mK ]
  • the feature mask M corresponding to the specified speaker is expressed as [m 1 ,m 2 ,..., m k ,...,m K ]
  • m 1 represents the proportion of the speech features of the specified speaker in x m1
  • m 2 represents the proportion of the speech features of the specified speaker in x m2
  • m K represents x Proportion of speech features of the specified speaker in mK
  • m 1 ⁇ m K are all values between [0,1].
  • the training mixed speech feature X m is multiplied frame by frame by the feature mask M corresponding to the specified speaker, and the specified feature extracted from the training mixed speech feature X m can be obtained speaker's phonetic features
  • Step S204 Extract the speech features of the specified speaker Input the speech recognition model to obtain the speech recognition result of the specified speaker
  • the registered voice features of the specified speaker (the registered voice features of the specified speaker refer to the voice features of the registered voice of the specified speaker)
  • Xe [x e1 , x e2 ,...,x ek ,...,x eK ], except that the extracted speech features of the specified speaker
  • the registered speech feature X e of the designated speaker is also input into the speech recognition model, and the speech recognition model is assisted by the registered speech feature X e of the designated speaker.
  • the speech recognition model in this embodiment may include: an encoder module, an attention module, and a decoder module. in:
  • the input to the encoder module consists of the extracted speech features of the specified speaker
  • two encoding modules can be set in the encoder module, as shown in Figure 3, the first encoding module 3021 is set in the encoder module and the second encoding module 3022, wherein the first encoding module 3021 is used to extract speech features of the designated speaker Encoding, the second encoding module is used to encode the registered speech features X e of the designated speaker; in another possible implementation, an encoding module is set in the encoder module to extract the speech features of the designated speaker
  • the encoding operation of and the encoding operation of the registered speech feature X e of the designated speaker are both performed by this encoding module, that is, the two encoding processes share one encoding module.
  • each encoding module can include one or more encoding layers, and the encoding layer can adopt a long-short-term memory layer in a one-way or two-way long-short-term memory neural network , or use the convolutional layer of a convolutional neural network.
  • Attention module for extracting speech features from specified speakers respectively
  • the audio-related feature vector required for decoding at the time of decoding is extracted from the encoding result H x of the specified speaker and the encoding result of the registered speech feature X e of the specified speaker.
  • the decoding module is used to decode the audio-related feature vector extracted by the attention module, so as to obtain the recognition result at the decoding moment.
  • the attention module 3023 is based on the attention mechanism, and at each decoding moment, respectively, from Extracting the current _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ Audio-related feature vectors required at the decoding instant.
  • the extracted audio-related feature vector represents the audio content of the character to be decoded at the t-th decoding moment.
  • the attention mechanism refers to using a vector as a query item (query), performing an attention mechanism operation on a set of feature vector sequences, and selecting the feature vector that best matches the query item as an output, specifically, Calculate a matching coefficient between the query item and each feature vector in the feature vector sequence, and then multiply and sum these matching coefficients with the corresponding feature vectors to obtain a new feature vector that is the feature vector that best matches the query item.
  • the state feature vector d t of the decoder module 3024 is determined according to the recognition result y t-1 at the t-1th decoding moment and c t-1 x and c t-1 e output by the attention module.
  • the decoder module 3024 may include multiple neural network layers, for example, two layers of unidirectional long-short-term memory layers.
  • the recognition result y t-1 at one decoding moment and the c t-1 x and c t-1 e output by the attention module 3023 are used as input to calculate the state feature vector d t of the decoder, and d t is input to the attention module 3023 , used to calculate c t x and c t e at the tth decoding moment, and then c t x and c t e are spliced, and the spliced vector is used as the input of the second layer of long short-term memory layer of the decoder module 3024 (for example, Both c t x and c t e are 128-dimensional vectors, splicing c t x and c t e can obtain a 256-dimensional spliced vector, and the 256-dimensional spliced vector is input to the second layer of the long short-term memory layer of the decoder module 3024) , calculate the output h t
  • Step S205 Extract speech features according to the specified speaker and speech recognition results for the specified speaker Update the parameters of the feature extraction model, and based on the speech recognition results of the specified speaker Update the parameters of the speech recognition model.
  • step S205 may include:
  • Step S2051 obtain the labeled text T t of the specified speaker's voice st (voice of the specified speaker) corresponding to the training mixed voice s m , and obtain the voice feature of the specified speaker's voice st as the standard voice feature X of the specified speaker t .
  • Step S2052 extracting speech features according to the specified speaker and the standard speech feature X t of the specified speaker to determine the first prediction loss Loss1, and according to the speech recognition result of the specified speaker and the labeled text T t of the specified speaker's voice st to determine the second prediction loss Loss2.
  • the extracted speech features for a given speaker can be computed The minimum mean square error with the standard speech feature Xt of the specified speaker, as the first prediction loss Loss1, according to the speech recognition result of the specified speaker Compute the cross-entropy loss with the annotated text T t of the specified speaker's voice st as the second prediction loss.
  • Step S2053 update the parameters of the feature extraction model according to the first prediction loss Loss1 and the second prediction loss Loss2, and update the parameters of the speech recognition model according to the second prediction loss Loss2.
  • the parameters of the feature extraction model are updated so that the feature extraction model can extract the speech that is close to the standard speech features of the specified speaker and is conducive to speech recognition from the training mixed speech features feature, the speech feature is input into the speech recognition model for speech recognition, and better speech recognition effect can be obtained.
  • this embodiment is specific to the first embodiment of "using the pre-established feature extraction model, based on the target mixed voice features and the target speaker features, to extract the specified The speech features of the speaker, and the process of extracting the speech features of the target speaker" is introduced.
  • extracting the voice features of the specified speaker from the target mixed voice features, so as to obtain the process of extracting voice features of the target speaker may include:
  • Step b1 Input the speech features of the target mixed speech and the speaker features of the target speaker into the feature extraction model to obtain a feature mask corresponding to the target speaker.
  • the feature mask corresponding to the target speaker can represent the proportion of the target speaker's voice features in the voice features of the target mixed voice.
  • Step b2 Extract the speech features of the target speaker from the speech features of the target mixed speech according to the feature mask corresponding to the target speaker, so as to obtain the extracted speech features of the target speaker.
  • the speech features of the target mixed speech are multiplied frame by frame by the feature mask corresponding to the target speaker, so as to obtain the extracted speech features of the target speaker.
  • the process of obtaining the speech recognition result of the target speaker may include:
  • Step c1 The encoder module based on the speech recognition model encodes the extracted speech features of the target speaker and the registered speech features of the target speaker respectively to obtain two encoding results.
  • Step c2 Based on the attention module of the speech recognition model, the audio-related feature vectors required for decoding at the time of decoding are respectively extracted from the two encoding results.
  • the decoder module based on the speech recognition model decodes the audio-related feature vectors extracted from the two encoding results respectively to obtain the recognition result at the decoding moment.
  • the process of inputting the extracted speech features of the target speaker into the speech recognition model to obtain the speech recognition result of the target speaker is the same as the process of combining the extracted speech features of the designated speaker and the registered speech features of the designated speaker in the training stage.
  • the implementation process of inputting the speech recognition model and obtaining the speech recognition result of the designated speaker is similar.
  • the specific implementation process of steps c1 to c3 can be found in the introduction of the encoder module, attention module and decoder module in the second embodiment. The embodiment will not be repeated here.
  • the speech recognition method provided by the present application has the following advantages: First, the present application extracts multi-scale voiceprint features from the registered speech of the target speaker and inputs the feature extraction model, Increase the richness of the input information of the feature extraction model, and improve the feature extraction effect of the feature extraction model; second, the joint training of the feature extraction model and the speech recognition model enables the prediction loss of the speech recognition model to act on the feature extraction model, thereby making the feature
  • the extraction model can extract speech features that are beneficial to speech recognition, thereby improving the accuracy of speech recognition results; third, the speech features of the target speaker's registered speech are used as additional input to the speech recognition model, and the features extracted by the feature extraction model When the speech characteristics are not good, it can assist the speech recognition model to perform speech recognition, so as to obtain more accurate speech recognition results. To sum up, the speech recognition method provided by the present application can accurately recognize the speech content of the target speaker in the case of complex human voice interference.
  • the embodiment of the present application also provides a speech recognition device.
  • the speech recognition device provided in the embodiment of the present application is described below.
  • the speech recognition device described below and the speech recognition method described above can be referred to in correspondence.
  • FIG. 4 shows a schematic structural diagram of a speech recognition device provided by an embodiment of the present application, which may include: a feature acquisition module 401 , a feature extraction module 402 and a speech recognition module 403 .
  • the feature acquisition module 401 is configured to acquire speech features of the target mixed voice and speaker features of the target speaker.
  • the feature extraction module 402 is used to extract the target speech features from the speech features of the target mixed speech according to the speech features of the target mixed speech and the speaker features of the target speaker. Speech features of the target speaker to obtain the extracted speech features of the target speaker, wherein the target speech feature is a speech feature used to obtain a speech recognition result consistent with the real speech content of the target speaker.
  • the speech recognition module 403 is configured to obtain a speech recognition result of the target speaker according to the extracted speech features of the target speaker.
  • the feature acquisition module 401 includes: a voice feature acquisition module and a speaker feature acquisition module.
  • the voice feature acquisition module is used to acquire the voice features of the target mixed voice.
  • the speaker feature acquisition module is used to acquire speaker features of the target speaker.
  • the speaker feature acquisition module acquires the speaker features of the target speaker
  • it is specifically configured to acquire the registered voice of the target speaker, and extract short-term voiceprint features from the registered voice of the target speaker and long-term voiceprint features to obtain multi-scale voiceprint features as the speaker features of the target speaker.
  • the feature extraction module 402 is specifically configured to use a pre-established feature extraction model, based on the speech features of the target mixed speech and the speaker features of the target speaker, from the speech features of the target mixed speech Extract the speech features of the target speaker.
  • the feature extraction model adopts the voice features of the training mixed voice including the voice of the specified speaker and the speaker features of the specified speaker, and uses the speech recognition results obtained based on the extracted voice features of the specified speaker as the optimization target training It is obtained that the extracted speech features of the specified speaker are the speech features of the specified speaker extracted from the speech features of the training mixed speech.
  • the feature extraction model is trained with the extracted speech features of the designated speaker and the speech recognition results obtained based on the extracted speech features of the designated speaker as optimization targets.
  • the feature extraction module 402 may include: a feature mask determination submodule and a speech feature extraction submodule.
  • the feature mask determination submodule is configured to input the speech features of the target mixed speech and the speaker features of the target speaker into the feature extraction model to obtain a feature mask corresponding to the target speaker, wherein , the feature mask can represent the proportion of the speech features of the corresponding speaker in the speech features of the target mixed speech.
  • the speech feature extraction submodule is used to extract the speech features of the target speaker from the speech features of the target mixed speech according to the speech features of the target mixed speech and the feature mask corresponding to the target speaker .
  • the speech recognition module 403 is specifically configured to obtain the speech recognition result of the target speaker according to the extracted speech features of the target speaker and the registered speech features of the target speaker; wherein, the target speaker The registered speech feature of the person is the speech feature of the registered speech of the target speaker.
  • the speech recognition module 403 is specifically configured to input speech recognition input features including at least the extracted speech features of the target speaker into a pre-established speech recognition model to obtain a speech recognition result of the target speaker.
  • the speech recognition model is jointly trained with the feature extraction model, the speech recognition model uses the extracted speech features of the specified speaker, and the speech recognition result obtained based on the extracted speech features of the specified speaker is The optimization objective is trained to get.
  • the speech recognition module 403 is specifically used to:
  • the speech recognition input features are encoded to obtain a coding result; based on the attention module of the speech recognition model, the information required for decoding at the time of decoding is extracted from the coding result
  • An audio-related feature vector the decoder module based on the speech recognition model decodes the audio-related feature vector extracted from the encoding result to obtain a recognition result at the decoding moment.
  • the speech recognition device may further include: a model training module.
  • the model training module may include: an acquisition module for extracting speech features, an acquisition module for speech recognition results, and a parameter update module.
  • the extracted speech feature acquisition module is configured to use a feature extraction model to extract the speech features of the specified speaker from the speech features of the training mixed speech, so as to obtain the extracted speech features of the specified speaker.
  • the speech recognition result obtaining module is used to obtain the speech recognition result of the designated speaker by using the speech recognition model and the extracted speech features of the designated speaker.
  • the model update module is used to update the parameters of the feature extraction model according to the extracted speech features of the designated speaker and the speech recognition results of the designated speaker, and perform speech recognition according to the speech recognition results of the designated speaker.
  • the model is updated with parameters.
  • the model update module may include: an annotation text acquisition module, a standard speech feature acquisition module, a prediction loss determination module, and a parameter update module.
  • the training mixed speech corresponds to the speech of the designated speaker.
  • the standard speech feature acquisition module is used to obtain the speech feature of the voice of the designated speaker as the standard speech feature of the designated speaker.
  • the annotation text obtaining module is used to obtain the annotation text of the voice of the designated speaker.
  • the prediction loss determining module is configured to determine a first prediction loss according to the extracted speech features of the specified speaker and the standard speech features of the specified speaker, and to determine the first prediction loss according to the speech recognition result of the specified speaker and the specified Annotated text of the speaker's speech, determining a second prediction loss.
  • the parameter update module is configured to update the parameters of the feature extraction model according to the first prediction loss and the second prediction loss, and update the parameters of the speech recognition model according to the second prediction loss.
  • the training mixed voice and the voice of the designated speaker corresponding to the training mixed voice are obtained from a pre-built training data set, and the voice recognition device provided in the embodiment of the present application may further include: building a training data set module.
  • the training dataset building blocks are used to:
  • each voice is the voice of a single speaker, and each voice has annotated text; taking part of the multiple voices or each voice in all the voices as a specified utterance Human voice: Mix one or more voices of other speakers in other voices with the voice of the specified speaker to obtain a training mixed voice, and use the voice of the specified speaker and the training mixed voice obtained by mixing as A piece of training data; the training data set is composed of all obtained training data.
  • the speech recognition device can extract the speech features of the target speaker from the speech features of the target mixed speech, and then can obtain the speech recognition result of the target speaker according to the extracted speech features of the target speaker.
  • the speech features tending to the target speech feature (the speech feature used to obtain the speech recognition result consistent with the real speech content of the target target speaker) is Therefore, the extracted speech feature is the target speech feature or a speech feature close to the target speech feature, and speech recognition based on the speech feature can obtain a better speech recognition effect, that is, a more accurate speech recognition can be obtained As a result, the user experience is better.
  • the embodiment of the present application also provides a speech recognition device.
  • FIG. 5, shows a schematic structural diagram of the speech recognition device.
  • the speech recognition device may include: at least one processor 501, at least one communication interface 502, at least one memory 503 and at least one communication bus 504;
  • the number of processor 501, communication interface 502, memory 503, and communication bus 504 is at least one, and the processor 501, communication interface 502, and memory 503 complete mutual communication through the communication bus 504;
  • Processor 501 may be a central processing unit CPU, or a specific integrated circuit ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiments of the present invention, etc.;
  • ASIC Application Specific Integrated Circuit
  • the memory 503 may include a high-speed RAM memory, and may also include a non-volatile memory (non-volatile memory), such as at least one disk memory;
  • the memory stores a program
  • the processor can call the program stored in the memory, and the program is used for:
  • the speech feature of the target speaker is extracted from the speech feature of the target mixed speech to obtain the extracted speech of the target speaker feature, wherein the target speech feature is a speech feature used to obtain a speech recognition result consistent with the target speaker's real speech content;
  • the embodiment of the present application also provides a readable storage medium, which can store a program suitable for execution by a processor, and the program is used for:
  • the speech feature of the target speaker is extracted from the speech feature of the target mixed speech to obtain the extracted speech of the target speaker feature, wherein the target speech feature is a speech feature used to obtain a speech recognition result consistent with the target speaker's real speech content;

Landscapes

  • Engineering & Computer Science (AREA)
  • Multimedia (AREA)
  • Health & Medical Sciences (AREA)
  • Audiology, Speech & Language Pathology (AREA)
  • Human Computer Interaction (AREA)
  • Physics & Mathematics (AREA)
  • Acoustics & Sound (AREA)
  • Computational Linguistics (AREA)
  • Computer Vision & Pattern Recognition (AREA)
  • Artificial Intelligence (AREA)
  • Signal Processing (AREA)
  • Evolutionary Computation (AREA)
  • Image Analysis (AREA)
  • Information Retrieval, Db Structures And Fs Structures Therefor (AREA)
  • Electrically Operated Instructional Devices (AREA)
  • Machine Translation (AREA)
  • Telephonic Communication Services (AREA)

Abstract

本申请提供了一种语音识别方法、装置、设备及存储介质,其中,方法包括:获取目标混合语音的语音特征以及指定说话人的说话人特征;以趋于目标语音特征为提取方向,根据目标混合语音的语音特征以及目标说话人的说话人特征,从目标混合语音的语音特征中提取目标说话人的语音特征,以得到目标说话人的提取语音特征,其中,目标语音特征为用于获得与目标说话人的真实语音内容一致的语音识别结果的语音特征;根据指定说话人的提取语音特征,获取指定说话人的语音识别结果。经由本申请提供的语音识别方法可从包含指定说话人语音的混合语音中较为准确的识别出指定说话人的语音内容,用户体验较好。

Description

一种语音识别方法、装置、设备及存储介质
本申请要求于2021年09月07日提交中国专利局、申请号为CN202111042821.8、发明名称为“一种语音识别方法、装置、设备及存储介质”的中国专利申请的优先权,其全部内容通过引用结合在本申请中。
技术领域
本申请涉及语音识别技术领域,尤其涉及一种语音识别方法、装置、设备及存储介质。
背景技术
随着人工智能技术的飞速发展,智能设备在人们的生活中扮演着越来越重要的角色,语音交互作为最方便自然的人机交互方式深受用户喜爱。
在用户使用智能设备时,其可能处在一个存在其他人声的复杂环境中,在这种情况下,智能设备采集的语音为混合语音。在进行语音交互时,为了能够获得较好的用户体验,就需要从混合语音中识别出目标说话人的语音内容,而如何从混合语音中识别出目标说话人的语音内容是目前亟需解决的问题。
发明内容
有鉴于此,本申请提供了一种语音识别方法、装置、设备及存储介质,用以从混合语音中较为准确地识别出目标说话人的语音内容,其技术方案如下:
一种语音识别方法,包括:
获取目标混合语音的语音特征以及目标说话人的说话人特征;
以趋于目标语音特征为提取方向,根据所述目标混合语音的语音特征以及所述目标说话人的说话人特征,从所述目标混合语音的语音特征中提取所述目标说话人的语音特征,以得到所述目标说话人的提取语音特征,其中,所述目标语音特征为用于获得与所述目标说话人的真实语音内容一致的语音识别结果的语音特征;
根据所述目标说话人的提取语音特征,获取所述目标说话人的语音识别结果。
可选的,获取所述目标说话人的说话人特征,包括:
获取所述目标说话人的注册语音;
对所述目标说话人的注册语音提取短时声纹特征和长时声纹特征,以得到 多尺度声纹特征,作为所述目标说话人的说话人特征。
可选的,所述以趋于目标语音特征为提取方向,根据所述目标混合语音的语音特征以及所述目标说话人的说话人特征,从所述目标混合语音的语音特征中提取所述目标说话人的语音特征,包括:
利用预先建立的特征提取模型,以所述目标混合语音的语音特征以及所述目标说话人的说话人特征为依据,从所述目标混合语音的语音特征中提取所述目标说话人的语音特征;
其中,所述特征提取模型采用包含指定说话人的语音的训练混合语音的语音特征和所述指定说话人的说话人特征,以基于所述指定说话人的提取语音特征获取的语音识别结果为优化目标训练得到,所述指定说话人的提取语音特征为从所述训练混合语音的语音特征中提取的所述指定说话人的语音特征。
可选的,所述特征提取模型同时以所述指定说话人的提取语音特征和基于所述指定说话人的提取语音特征获取的语音识别结果为优化目标训练得到。
可选的,所述利用预先建立的特征提取模型,以所述目标混合语音的语音特征以及所述目标说话人的说话人特征为依据,从所述目标混合语音的语音特征中提取所述目标说话人的语音特征,包括:
将所述目标混合语音的语音特征以及所述目标说话人的说话人特征输入所述特征提取模型,得到所述目标说话人对应的特征掩码;
根据所述目标混合语音的语音特征和所述目标说话人对应的特征掩码,从所述目标混合语音的语音特征中提取所述目标说话人的语音特征。
可选的,所述根据所述目标说话人的提取语音特征,获取所述目标说话人的语音识别结果,包括:
根据所述目标说话人的提取语音特征以及所述目标说话人的注册语音特征,获取所述目标说话人的语音识别结果;
其中,所述目标说话人的注册语音特征为所述目标说话人的注册语音的语音特征。
可选的,所述根据所述目标说话人的提取语音特征,获取所述目标说话人的语音识别结果,包括:
将至少包括所述目标说话人的提取语音特征的语音识别输入特征输入预先建立的语音识别模型,得到所述目标说话人的语音识别结果;
所述语音识别模型与所述特征提取模型联合训练得到,所述语音识别模型采用所述指定说话人的提取语音特征,以基于所述指定说话人的提取语音特征获得的语音识别结果为优化目标训练得到。
可选的,将所述语音识别输入特征输入所述语音识别模型,得到所述目标说话人的语音识别结果,包括:
基于所述语音识别模型的编码器模块,对所述语音识别输入特征进行编码,以得到编码结果;
基于所述语音识别模型的注意力模块,从所述编码结果中提取解码时刻解码所需的音频相关特征向量;
基于所述语音识别模型的解码器模块,对从所述编码结果中提取的所述音频相关特征向量进行解码,得到所述解码时刻的识别结果。
可选的,所述语音识别模型与所述特征提取模型联合训练的过程包括:
利用特征提取模型,从所述训练混合语音的语音特征中提取所述指定说话人的语音特征,以得到所述指定说话人的提取语音特征;
利用语音识别模型和所述指定说话人的提取语音特征,获取所述指定说话人的语音识别结果;
根据所述指定说话人的提取语音特征和所述指定说话人的语音识别结果对特征提取模型进行参数更新,并根据所述指定说话人的语音识别结果对语音识别模型进行参数更新。
可选的,所述训练混合语音对应有所述指定说话人的语音;
所述根据所述指定说话人的提取语音特征和所述指定说话人的语音识别结果对特征提取模型进行参数更新,并根据所述指定说话人的语音识别结果对语音识别模型进行参数更新,包括:
获取所述指定说话人的语音的标注文本,并获取所述指定说话人的语音的语音特征作为所述指定说话人的标准语音特征;
根据所述指定说话人的提取语音特征和所述指定说话人的标准语音特征确定第一预测损失,并根据所述指定说话人的语音识别结果和所述指定说话人的语音的标注文本,确定第二预测损失;
根据所述第一预测损失和所述第二预测损失对特征提取模型进行参数更新,并根据所述第二预测损失对语音识别模型进行参数更新。
可选的,所述训练混合语音以及所述训练混合语音对应的所述指定说话人的语音从预先构建的训练数据集中获取;
所述训练数据集的构建过程包括:
获取多个说话人的多条语音,其中,每条语音为单一说话人的语音,每条语音具有标注文本;
将所述多条语音中的部分语音或全部语音中的每条语音作为指定说话人的语音:将其它语音中其他说话人的一条或多条语音与该指定说话人的语音进行混合,以得到一条训练混合语音,将该指定说话人的语音与通过混合得到的训练混合语音作为一条训练数据;
由获得的所有训练数据组成所述训练数据集。
一种语音识别装置,包括:特征获取模块、特征提取模块和语音识别模块;
所述特征获取模块,用于获取目标混合语音的语音特征以及目标说话人的说话人特征;
所述特征提取模块,用于以提取趋于目标语音特征为提取方向,根据所述目标混合语音的语音特征以及所述目标说话人的说话人特征,从所述目标混合语音的语音特征中提取所述目标说话人的语音特征,以得到所述目标说话人的提取语音特征,其中,所述目标语音特征为用于获得与所述目标说话人的真实语音内容一致的语音识别结果的语音特征;
所述语音识别模块,用于根据所述目标说话人的提取语音特征,获取所述目标说话人的语音识别结果。
可选的,所述特征获取模块包括:说话人特征获取模块;
所述说话人特征获取模块,用于获取所述目标说话人的注册语音,对所述目标说话人的注册语音提取短时声纹特征和长时声纹特征,以得到多尺度声纹特征,作为所述目标说话人的说话人特征。
可选的,所述特征提取模块具体用于利用预先建立的特征提取模型,以所述目标混合语音的语音特征以及所述目标说话人的说话人特征为依据,从所述目标混合语音的语音特征中提取所述目标说话人的语音特征;
其中,所述特征提取模型采用包含指定说话人的语音的训练混合语音的语音特征和指定说话人的说话人特征,以基于所述指定说话人的提取语音特征获取的语音识别结果为优化目标训练得到,所述指定说话人的提取语音特征为从 所述训练混合语音的语音特征中提取的所述指定说话人的语音特征。
可选的,所述语音识别模块,具体用于根据所述目标说话人的提取语音特征以及所述目标说话人的注册语音特征,获取所述目标说话人的语音识别结果;
其中,所述目标说话人的注册语音特征为所述目标说话人的注册语音的语音特征。
可选的,所述语音识别模块,具体用于将至少包括所述目标说话人的提取语音特征的语音识别输入特征输入预先建立的语音识别模型,得到所述目标说话人的语音识别结果;
其中,所述语音识别模型与所述特征提取模型联合训练得到,所述语音识别模型采用所述指定说话人的提取语音特征,以基于所述指定说话人的提取语音特征获得的语音识别结果为优化目标训练得到。
一种语音识别设备,包括:存储器和处理器;
所述存储器,用于存储程序;
所述处理器,用于执行所述程序,实现上述任一项所述的语音识别方法的各个步骤。
一种可读存储介质,其上存储有计算机程序,所述计算机程序被处理器执行时,实现上述任一项所述的语音识别方法的各个步骤。
经由上述方案可知,本申请提供的语音识别方法、装置、设备及存储介质,能够根据目标混合语音的语音特征以及目标说话人的说话人特征,从目标混合语音的语音特征中提取出目标说话人的语音特征,进而能够根据提取出的目标说话人的语音特征,获得目标说话人的语音识别结果,由于本申请在从目标混合语音的语音特征提取目标说话人的语音特征时,以趋于目标语音特征(用于获得与目标目标说话人的真实语音内容一致的语音识别结果的语音特征)为提取方向,因此,提取出的语音特征为目标语音特征或者趋近于目标语音特征的语音特征,可见,经由上述方式提取出的语音特征为有利于语音识别的特征,基于提取出的语音特征进行语音识别,能够获得较好的语音识别效果,即能够获得较为准确的语音识别结果,用户体验较好。
附图说明
为了更清楚地说明本发明实施例或现有技术中的技术方案,下面将对实施 例或现有技术描述中所需要使用的附图作简单地介绍,显而易见地,下面描述中的附图仅仅是本发明的实施例,对于本领域普通技术人员来讲,在不付出创造性劳动的前提下,还可以根据提供的附图获得其他的附图。
图1为本申请实施例提供的语音识别方法的流程示意图;
图2为本申请实施例提供的特征提取模型与语音识别模型联合训练的流程示意图;
图3为本申请实施例提供的特征提取模型与语音识别模型联合训练的过程示意图;
图4为本申请实施例提供的语音识别装置的结构示意图;
图5为本申请实施例提供的语音识别设备的结构示意图。
具体实施方式
下面将结合本发明实施例中的附图,对本发明实施例中的技术方案进行清楚、完整地描述,显然,所描述的实施例仅仅是本发明一部分实施例,而不是全部的实施例。基于本发明中的实施例,本领域普通技术人员在没有做出创造性劳动前提下所获得的所有其他实施例,都属于本发明保护的范围。
在外界环境中,人们通常被许多不同的声源围绕着,比如,多个人同时说话的声音、交通噪声、自然噪声等,经过研究者的不懈努力,以上描述的背景噪声的分离问题即通常意义上的语音增强问题已经得到了较好的解决,而相比之下,多个人同时说话的情况下,如何识别目标说话人的语音内容,即如何从混合语音中识别目标说话人的语音内容为难度更大的问题,其更具备研究意义。
为了能够从混合语音中识别出目标说话人的语音内容,申请人进行了研究,起初的思路是:先训练一个特征提取模型,再训练一个语音识别模型;获取目标说话人的注册语音,并对目标说话人的注册语音提取d-vector作为目标说话人的说话人特征;基于预先训练得到的特征提取模型,以目标说话人的说话人特征和目标混合语音的语音特征为依据,从目标混合语音的语音特征中提取出目标说话人的语音特征;通过对提取出的目标说话人的语音特征进行一系列的变换处理来获得目标说话人的语音;将目标说话人的语音输入预先训练得到的语音识别模型进行语音识别,从而获得目标说话人的语音识别结果。
申请人通过对上述思路进行研究发现,上述思路存在诸多缺陷,主要包括 如下几个方面:其一,对目标说话人的注册语音提取的d-vector所包含的声纹信息不足,影响后续特征提取的效果;其二,特征提取模型与语音识别模型是单独训练的,二者是完全割裂的,不能有效联合优化,级联两个独立训练得到的模型进行语音识别会存在级联误差,进而影响语音识别效果;其三,在前端的特征提取部分提取的特征不佳时,后端的语音识别部分没有任何补救措施,可能导致语音识别效果较差。
申请人在上述思路以及上述思路所存在的缺陷的基础上,进一步进行研究,最终提出了一种能够完美克服上述缺陷的语音识别方法,该语音识别方法能够从混合语音中较为准确地识别出目标说话人的语音内容,该语音识别方法可应用于具有数据处理能力的终端,终端可按本申请提供的语音识别方法从目标混合语音中识别出目标说话人的语音内容,该终端可以包括处理组件、存储器、输入/输出接口和电源组件,可选的,该终端还可以包括多媒体组件、音频组件、传感器组件和通信组件等。其中:
处理组件用于进行数据处理,其可以进行本案的语音合成处理,处理组件可以包括一个或多个处理器,处理组件还可以包括一个或多个模块,便于与其它组件之间的交互。
存储器被配置为存储各种类型的数据,存储器可以有任何类型的易失性或非易失性存储设备或者它们的组合实现,如静态随机存取存储器(SRAM)、电可擦除可编程只读存储器(EEPROM)、可擦除可编程只读存储器(EPROM)、可编程只读存储器(PROM)、只读存储器(ROM)、磁存储器、快闪存储器、磁盘、光盘等中的一种或多种的组合。
电源组件为终端的各种组件提供电力,电源组件可以包括电源管理系统、一个或多个电源等。
多媒体组件可以包括屏幕,优选的,屏幕可以为触摸显示屏,触摸显示屏可接收来自用户的输入信号。多媒体组件还可以包括前置摄像头和/或后置摄像头。
音频组件被配置为输出和/或输入音频信号,如音频组件可以包括麦克风,麦克风被配置为接收外部音频信号,音频组件还可以包括扬声器,扬声器被配置为输出音频信号,终端合成的语音可通过扬声器输出。
输入/输出接口为处理组件与外围接口模块之间的接口,外围接口模块可 以为键盘、按钮等,其中,按钮可包括但不限定于主页按钮、音量按钮、启动按钮、锁定按钮等。
传感器组件可以包括一个或多个传感器,用于为终端提供各个方面的状态评估,例如,传感器组件可以检测终端的打开/关闭状态、用户与终端是否接触、装置的方位、速度、温度等。传感器组件可以包括但不限定于图像传感器、加速度传感器、陀螺仪传感器、压力传感器、温度传感器等中的一种或多种的组合。
通信组件被配置为便于终端和其它设备进行有线或无线通信。终端可接入基于通信标准的无线网络,如WiFi、2G、3G、4G、5G中的一种或多种的组合。
可选的,终端可被一个或多个应用专用集成电路(ASIC)、数字信号处理器(ASP)、数字信号处理设备(DSPD)、可编程逻辑器件(PLD)、现场可编程门阵列(FPGA)、控制器、微控制器、微处理器或其他电子元件实现,用于执行本申请提供的同传翻译方法。
本申请提供的语音识别方法还可应用于服务器,服务器可按本申请提供的语音识别方法从目标混合语音中识别出目标说话人的语音内容,在一种场景中,服务器可通过网络与终端连接,终端获取目标混合语音,将目标混合语音通过与服务器连接的网络传输至服务器,服务器按本申请提供的语音识别方法从目标混合语音中识别出目标说话人的语音内容,再通过网络将目标说话人的语音内容传输至终端。服务器可以包括一个或一个以上的中央处理器和存储器,其中,存储器被配置为存储各种类型的数据,存储器可以有任何类型的易失性或非易失性存储设备或者它们的组合实现,如静态随机存取存储器(SRAM)、电可擦除可编程只读存储器(EEPROM)、可擦除可编程只读存储器(EPROM)、可编程只读存储器(PROM)、只读存储器(ROM)、磁存储器、快闪存储器、磁盘、光盘等中的一种或多种的组合。服务器还可以包括一个或一个以上电源、一个或一个以上有线网络接口和/或一个或一个以上无线网络接口、一个或一个以上操作系统。
接下来通过下述实施例对本申请提供的语音识别方法进行介绍。
第一实施例
请参阅图1,示出了本申请实施例提供的语音识别方法的流程示意图,该方法可以包括:
步骤S101:获取目标混合语音的语音特征以及目标说话人的说话人特征。
其中,目标混合语音为多个说话人的语音,其除了包括目标说话人的语音外,还包括其他说话人的语音,本申请意在实现,在存在其他说话人的语音的情况下,较为准确地识别出目标说话人的语音内容。
其中,获取目标混合语音的语音特征的过程包括:获取目标混合语音中每个语音帧的特征向量(比如频谱特征),以得到特征向量序列,将获得的特征向量序列作为目标混合语音的语音特征。假设目标混合语音包括K个语音帧,第k个语音帧的特征向量表示为x k,则目标混合语音的语音特征可表示为[x 1,x 2,…,x k,…,x K]。
其中,获取目标说话人的说话人特征实现方式有多种,本实施例提供如下两种可选的实现方式:在一种可能的实现方式中,可获取目标说话人的注册语音,对目标说话人的注册语音提取d-vector,提取的d-vector作为目标说话人的说话人特征;考虑到d-vector包含的声纹信息较为单一,不够丰富,为了提升后续特征提取的效果,本实施例提供另一种较为优选的实现方式,即,获取目标说话人的注册语音,对目标说话人的注册语音提取短时声纹特征和长时声纹特征,以得到多尺度声纹特征,将多尺度声纹特征作为目标说话人的说话人特征。
相比于经由上述第一种实现方式获得的说话人特征,经由上述第二种实现方式获得的说话人特征含有更为丰富的声纹信息,这使得后续利用经由上述第二种实现方式获得的说话人特征进行特征提取,能够获得更佳的特征提取效果。
接下来对上述的第二种实现方式中,“对指定说话人的注册语音提取短时声纹特征和长时声纹特征”的具体实现过程进行介绍。
对目标说话人的注册语音提取短时声纹特征和长时声纹特征的过程可以包括:利用预先建立的说话人表征提取模型,从目标说话人的注册语音中提取短时声纹特征和长时声纹特征。具体的,获取目标说话人的注册语音的语音特征序列,将目标说话人的注册语音的语音特征序列输入预先建立的说话人表征提取模型,获得目标说话人的短时声纹特征和长时声纹特征。
可选的,说话人表征提取模型可以采用卷积神经网络,将目标说话人的注册语音的语音特征序列输入卷积神经网络进行特征提取,以获得浅层特征和深 层特征,其中,浅层特征因感受野较小更能表征短时声纹,因此,将浅层特征作为短时声纹特征,而深层特征因感受野较大更能表征长时声纹,因此,将深层特征作为长时声纹特征。
本实施例中的说话人表征提取模型采用大量带真实说话人标签的训练语音(此处的训练语音优选为单一说话人的语音)训练得到,其中,训练语音的真实说话人标签代表的是训练语音对应的说话人。可选的,可采用交叉熵(Cross Entropy,CE)准则或者度量学习(Metric Learning,ML)准则训练说话人表征提取模型。
步骤S102:以趋于目标语音特征为提取方向,根据目标混合语音的语音特征以及目标说话人的说话人特征,从目标混合语音的语音特征中提取目标说话人的语音特征,以得到目标说话人的提取语音特征。
其中,目标语音特征为用于获得与目标说话人的真实语音内容一致的语音识别结果的语音特征。
以趋于目标语音特征为提取方向,能够从目标混合语音的语音特征中提取出目标语音特征或者趋近于目标语音特征的语音特征,即,以趋于目标语音特征为提取方向能够从目标混合语音的语音特征中提取出有利于后续语音识别的语音特征,根据有利于语音识别的语音特征进行语音识别,能够获得较好的语音识别效果。
可选的,以趋于目标语音特征为提取方向,根据目标混合语音的语音特征以及目标说话人的说话人特征,从目标混合语音的语音特征中提取目标说话人的语音特征,以得到目标说话人的提取语音特征的过程可以包括:利用预先建立的特征提取模型,以目标混合语音特征和目标说话人特征为依据,从目标混合语音特征中提取指定说话人的语音特征,以得到目标说话人的提取语音特征。
其中,特征提取模型采用包含指定说话人的语音的训练混合语音的语音特征和指定说话人的说话人特征,以基于指定说话人的提取语音特征获取的语音识别结果为优化目标训练得到。需要说明的是,在训练阶段,特征提取模型的输入为上述的训练混合语音的语音特征和指定说话人的说话人特征,输出为从训练混合语音特征中提取的指定说话人的语音特征。
在一种可能的实现方式中,特征提取模型训练时,以基于指定说话人的提 取语音特征获取的语音识别结果为优化目标。以基于指定说话人的提取语音特征获取的语音识别结果为优化目标训练特征提取模型,使得基于特征提取模型能够从混合语音特征中提取出有利于语音识别的语音特征。
为了提升特征提取效果,在另一种可能的实现方式中,特征提取模型训练时,以指定说话人的提取语音特征以及基于指定说话人的提取语音特征获取的语音识别结果为优化目标。同时以指定说话人的提取语音特征以及基于指定说话人的提取语音特征获取的语音识别结果为优化目标,使得基于特征提取模型能够从混合语音特征中提取出有利于语音识别且趋近于指定说话人的标准语音特征的语音特征。需要说明的是,目标说话人的标准语音特征指的是根据指定说话人的语音(干净语音)获得的语音特征。
步骤S103:根据目标说话人的提取语音特征,获取目标说话人的语音识别结果。
根据目标说话人的提取语音特征,获取目标说话人的语音识别结果的实现方式有多种:在一种可能的实现方式中,可只根据目标说话人的提取语音特征,获取目标说话人的语音识别结果;为了提升语音识别效果,在另一种可能的实现方式中,可根据目标说话人的提取语音特征以及目标说话人的注册语音特征(目标说话人的注册语音特征指的是,目标说话人的注册语音的语音特征),获取目标说话人的语音识别结果,其中,目标说话人的注册语音特征作为识别辅助信息,用以提升语音识别效果。
具体的,可利用预先建立的语音识别模型,获取目标说话人的语音识别结果,更为具体的,将目标说话人的提取语音特征作为语音识别输入特征,或者将目标说话人的提取语音特征以及目标说话人的注册语音特征作为语音识别输入特征,将语音识别输入特征输入预先建立的语音识别模型,以得到目标说话人的语音识别结果。
需要说明的是,在将目标说话人的提取语音特征以及目标说话人的注册语音特征作为语音识别输入特征输入语音识别模型时,目标说话人的注册语音特征能够在目标说话人的提取语音特征不准确的情况下,辅助语音识别模型进行语音识别,从而提升语音识别效果。
优选的,语音识别模型可与特征提取模型联合训练得到,语音识别模型采用上述的“指定说话人的提取语音特征”为训练样本,以基于指定说话人的提 取语音特征获得的语音识别结果为优化目标训练得到。将特征提取模型与语音识别模型联合训练,使得特征提取模型能够朝着利于语音识别的方向去优化。
本申请实施例提供的语音识别方法能够从目标混合语音的语音特征中提取出目标说话人的语音特征,进而能够根据提取出的目标说话人的语音特征,获得目标说话人的语音识别结果,由于本申请实施例在从目标混合语音的语音特征提取目标说话人的语音特征时,以趋于目标语音特征(用于获得与目标目标说话人的真实语音内容一致的语音识别结果的语音特征)为提取方向,因此,提取出的语音特征为目标语音特征或者趋近于目标语音特征的语音特征,基于该语音特征进行语音识别,能够获得较好的语音识别效果,即能够获得较为准确的语音识别结果,用户体验较好。
第二实施例
上述实施例提到,用于从目标混合语音的语音特征中提取目标说话人的语音特征的特征提取模型,以及用于根据特征提取模型提取的特征获取目标说话人的语音识别结果的语音识别模型,可通过联合训练方式训练得到。本实施例重点对特征提取模型与语音识别模型的联合训练过程进行介绍。
下面在图2的基础上结合图3对特征提取模型与语音识别模型的联合训练过程进行介绍,特征提取模型与语音识别模型的联合训练过程可以包括:
步骤S201:从预先构建的训练数据集S中获取训练混合语音s m。
其中,训练数据集S中包括多条训练数据,每条训练数据均包括指定说话人的语音(干净语音),以及包含该指定说话人的语音的训练混合语音,其中,指定说话人的语音具有标注文本(标注文本为指定说话人的语音的语音内容)。
训练数据集S的构建过程包括:
步骤a1、获取多个说话人的多条语音。
本步骤获取的多条语音中的每条语音为单一说话人的语音,每条语音具有标注文本,假设一说话人的语音为内容为“今天天气不错”的语音,则该语音的标注文本为“<s>,今,填,天,气,不,错,</s>”,其中,“<s>”为句子开始符,“</s>”为句子结束符。
需要说明的是,多条语音的数量与多条语音对应的说话人的数量可以相同,也可以不同,假设步骤a1获取了P个说话人的Q多条语音,则P与Q的关系可以为P=Q(比如,获取说话人a的一条语音、说话人b的一条语音、说 话人c的一条语音),也可以为P<Q(比如,获取说话人a的两条语音、说话人b的一条语音、说话人c的三条语音),也就是说,针对每个说话人,可获取一条语音,也可以获取多条语音。
步骤a2、将多条语音中的部分语音或全部语音中的每条语音作为指定说话人的语音:将其它语音中其他说话人的一条或多条语音与该指定说话人的语音进行混合,以得到一条训练混合语音,将该条训练混合语音与该指定说话人的语音作为一条训练数据。
示例性的,获取的多条语音包括说话人a的一条语音、说话人b的一条语音、说话人c的一条语音和说话人d的一条语音,此处的每条语音都是单一说话人的干净语音,可将说话人a的语音作为指定说话人的语音,将其他说话人(一个或多个说话人)的语音与说话人a的语音混合,以得到一条训练混合语音,比如,将说话人b的语音与说话人a语音混合,或者,将说话人b的语音、说话人c的语音与说话人a的语音混合,将说话人a的语音以及通过将说话人a的语音与其他说话人的语音混合得到的训练混合语音作为一条训练数据,同样的,可将说话人b的语音作为指定说话人的语音,将其他说话人(一个或多个说话人)的语音与说话人b的语音混合,以得到一条训练混合语音,将说话人b的语音以及通过将说话人b的语音与其他说话人的语音混合得到的训练混合语音作为一条训练数据,按该方式可获得多条训练数据。
需要说明的是,在将指定说话人的语音与其他说话人的语音混合时,若其他说话人的语音的长度与指定说话人的语音的长度不同,则需要将其他说话人的语音处理成与指定说话人的语音长度相同。假设指定说话人的语音包括K个语音帧,即指定说话人的语音的长度为K,若其他说话人的语音的长度大于K,则可将其他说话人的语音中的第K+1个语音帧以及后面的语音帧删除,即只保留前K个语音帧,若其他说话人的语音的长度小于K,假设为L,则从前面复制K-L个语音帧进行补充。
步骤a3、由获得的所有训练数据组成训练数据集。
步骤S202:获取训练混合语音s m的语音特征作为训练混合语音特征X m,,并获取指定说话人的说话人特征作为训练说话人特征。
如第一实施例所述,可预先建立说话人表征提取模型,利用预先建立的说话人表征提取模型对指定说话人的注册语音提取说话人特征,提取的说话人 特征作为训练说话人特征。如图3所示,利用说话人表征提取模型300对指定说话人的注册语音提取短时声纹特征和长时声纹特征,提取的短时声纹特征和长时声纹特征作为指定说话人的说话人特征。
需要说明的是,说话人表征提取模型是在对特征提取模型与语音识别模型进行联合训练之前预先训练好的,在特征提取模型与语音识别模型的联合训练阶段,其参数固定,不随特征提取模型与语音识别模型进行参数更新。
步骤S203:利用特征提取模型,以训练混合语音特征X m和训练说话人特征为依据,从训练混合语音特征X m中提取指定说话人的语音特征,作为指定说话人的提取语音特征
Figure PCTCN2021129733-appb-000001
具体的,首先将训练混合语音特征X m和训练说话人特征输入特征提取模型,得到指定说话人对应的特征掩码M,然后根据指定说话人对应的特征掩码M,从训练混合语音特征X m中提取指定说话人的语音特征,作为指定说话人的提取语音特征
Figure PCTCN2021129733-appb-000002
如图3所示,将训练混合语音特征X m和训练说话人特征输入特征提取模型301,特征提取模型301根据输入的训练混合语音特征X m和训练说话人特征确定指定说话人对应的特征掩码M并输出。本实施例中的特征提取模型301可以但不限为循环神经网络(Recurrent Neural Network,RNN)、卷积神经网络(Convolution Neural Network,CNN)、深度神经网络(Deep Neural Network,DNN)等。
需要说明的是,训练混合语音特征X m为训练混合语音中各语音帧的特征向量组成的特征向量序列[x m1,x m2,…,x mk,…,x mK](K为训练混合语音的语音帧的总数量),在将训练混合语音特征X m和训练说话人特征输入特征提取模型301时,可将训练说话人特征与训练混合语音中每一语音帧的特征向量拼接,拼接后输入特征提取模型301。示例性的,训练混合语音中每一语音帧的特征向量为40维,训练说话人特征中的短时声纹特征和长时声纹特征均为40维,则在训练混合语音中每一语音帧的特征向量拼接上短时声纹特征和长时声纹特征后,可得到120维的拼接特征向量。在提取指定说话人的语音特征时,结合短时声纹特征和长时声纹特征,增加了输入信息的丰富程度,这使得特征提取模型更好的提取到指定说话人的语音特征。
在本实施例中,指定说话人对应的特征掩码M能够表征指定说话人的语 音特征在训练混合语音特征X m中的占比。若将训练混合语音特征X m表示为[x m1,x m2,…,x mk,…,x mK],将指定说话人对应的特征掩码M表示为[m 1,m 2,……,m k,……,m K],则m 1表示x m1中指定说话人的语音特征的占比,m 2表示x m2中指定说话人的语音特征的占比,以此类推,m K表示x mK中指定说话人的语音特征的占比,m 1~m K均为[0,1]之间的值。在获得指定说话人对应的特征掩码M后,将训练混合语音特征X m与指定说话人对应的特征掩码M逐帧相乘,便可得到从训练混合语音特征X m中提取出的指定说话人的语音特征
Figure PCTCN2021129733-appb-000003
步骤S204:将指定说话人的提取语音特征
Figure PCTCN2021129733-appb-000004
输入语音识别模型,获得指定说话人的语音识别结果
Figure PCTCN2021129733-appb-000005
优选的,为了提升语音识别模型的识别效果,可获取指定说话人的注册语音特征(指定说话人的注册语音特征指的是,指定说话人的注册语音的语音特征)Xe=[x e1,x e2,……,x ek,……,x eK],除了将指定说话人的提取语音特征
Figure PCTCN2021129733-appb-000006
输入语音识别模型外,将指定说话人的注册语音特征X e也输入语音识别模型,用指定说话人的注册语音特征X e辅助语音识别模型进行语音识别。
可选的,本实施例中的语音识别模型可以包括:编码器模块、注意力模块和解码器模块。其中:
编码器模块,用于对指定说话人的提取语音特征
Figure PCTCN2021129733-appb-000007
进行编码,以得到
Figure PCTCN2021129733-appb-000008
的编码结果H x=[h 1 x,h 2 x,……,h K x],以及对指定说话人的注册语音特征X e进行编码,以得到X e的编码结果H e=[h 1 e,h 2 e,……,h K e]。需要说明的是,若只将指定说话人的提取语音特征
Figure PCTCN2021129733-appb-000009
输入语音识别模型,则编码器模块只需要对指定说话人的提取语音特征
Figure PCTCN2021129733-appb-000010
进行编码即可。
在编码器模块的输入包括指定说话人的提取语音特征
Figure PCTCN2021129733-appb-000011
和指定说话人的注册语音特征X e的情况下:在一种可能的实现方式中,编码器模块中可设置两个编码模块,如图3所示,编码器模块中设置第一编码模块3021和第二编码模块3022,其中,第一编码模块3021用于对指定说话人的提取语音特征
Figure PCTCN2021129733-appb-000012
进行编码,第二编码模块用于对指定说话人的注册语音特征X e进行编码;在另一种可能的实现方式中,编码器模块中设置一个编码模块,对指定说话人的提取语音特征
Figure PCTCN2021129733-appb-000013
的编码操作和对指定说话人的注册语音特征X e的编码操作均由这一个编码模块执行,即两个编码过程共用一个编码模块。在编码器模块的 输入只包括指定说话人的提取语音特征
Figure PCTCN2021129733-appb-000014
的情况下,编码器模块中只需要设置一个编码模块即可。不管编码器模块中设置一个编码模块还是设置两个编码模块,每个编码模块均可以包括一层或多层编码层,编码层可以采用单向或双向长短时记忆神经网络中的长短时记忆层,或者采用卷积神经网络的卷积层。
注意力模块,用于分别从指定说话人的提取语音特征
Figure PCTCN2021129733-appb-000015
的编码结果H x和指定说话人的注册语音特征X e的编码结果中提取解码时刻解码所需的音频相关特征向量。
解码模块,用于对注意力模块提取出的音频相关特征向量进行解码,以得到解码时刻的识别结果。
如图3所示,注意力模块3023基于注意力机制,在每个解码时刻分别从
Figure PCTCN2021129733-appb-000016
的编码结果H x=[h 1 x,h 2 x,……,h K x]和X e的编码结果H e=[h 1 e,h 2 e,……,h K e]中提取当前解码时刻所需的音频相关特征向量。对于第t个解码时刻,提取的音频相关特征向量表征的是第t个解码时刻待解码字符的音频内容。
需要说明的是,注意力机制指的是,使用一个向量作为查询项(query),对一组特征向量序列进行注意力机制操作,选出与查询项最匹配的特征向量作为输出,具体为,将查询项与特征向量序列中每个特征向量计算一个匹配系数,然后将这些匹配系数与对应的特征向量相乘并求和,得到一个新的特征向量即为与查询项最匹配的特征向量。
对于第t个解码时刻:注意力模块3023将解码器模块3024的状态特征向量d t作为查询项,计算d t与H x=[h 1 x,h 2 x,……,h K x]中每个特征向量的匹配系数w 1 x、w 2 x、……、w K x,然后将匹配系数w 1 x、w 2 x、……、w K x与H x=[h 1 x,h 2 x,……,h K x]中对应的特征向量相乘后求和,求和得到的特征向量作为音频相关特征向量c t x,同样的,注意力模块3023计算d t与H e=[h 1 e,h 2 e,……,h K e]中每个特征向量的匹配系数w 1 e、w 2 e、……、w K e,然后将匹配系数w 1 e、w 2 e、……、w K e与H e=[h 1 e,h 2 e,……,h K e]中对应的特征向量相乘后求和,求和得到的特征向量作为音频相关特征向量c t e,在获得音频相关特征向量c t x和c t e后,将音频相关特征向量c t x和c t e输入解码器模块3024进行解码,以得到第t个解码时刻的识别结果。
其中,解码器模块3024的状态特征向量d t根据第t-1个解码时刻的识别结果y t-1和注意力模块输出的c t-1 x和c t-1 e确定。可选的,解码器模块3024可以 包括多个神经网络层,比如,两层单向长短时记忆层,在第t个解码时刻,解码器模块3024的第一层长短时记忆层以第t-1个解码时刻的识别结果y t-1以及注意力模块3023输出的c t-1 x和c t-1 e作为输入,计算得到解码器的状态特征向量d t,d t输入注意力模块3023,用于计算第t个解码时刻的c t x和c t e,然后将c t x和c t e拼接,拼接后向量作为解码器模块3024的第二层长短时记忆层的输入(比如,c t x和c t e均为128维向量,将c t x与c t e拼接,可获得256维的拼接向量,256维的拼接向量输入解码器模块3024的第二层长短时记忆层),计算得到解码器的输出h t d,最终根据h t d计算输出字符的后验概率,从而根据输出字符的后验概率确定第t个解码时刻的识别结果。
步骤S205:根据指定说话人的提取语音特征
Figure PCTCN2021129733-appb-000017
和指定说话人的语音识别结果
Figure PCTCN2021129733-appb-000018
对特征提取模型进行参数更新,并根据指定说话人的语音识别结果
Figure PCTCN2021129733-appb-000019
对语音识别模型进行参数更新。
具体的,步骤S205的实现过程可以包括:
步骤S2051、获取训练混合语音s m对应的指定说话人语音s t(指定说话人的语音)的标注文本T t,并获取指定说话人语音s t的语音特征作为指定说话人的标准语音特征X t。
需要说明的是,此处的指定说话人语音s t与上述指定说话人的注册语音为指定说话人的不同语音。
步骤S2052、根据指定说话人的提取语音特征
Figure PCTCN2021129733-appb-000020
和指定说话人的标准语音特征X t确定第一预测损失Loss1,并根据指定说话人的语音识别结果
Figure PCTCN2021129733-appb-000021
和指定说话人语音s t的标注文本T t,确定第二预测损失Loss2。
可选的,可计算指定说话人的提取语音特征
Figure PCTCN2021129733-appb-000022
与指定说话人的标准语音特征X t的最小均方误差,作为第一预测损失Loss1,根据指定说话人的语音识别结果
Figure PCTCN2021129733-appb-000023
和指定说话人语音s t的标注文本T t计算交叉熵损失,作为第二预测损失。
步骤S2053、根据第一预测损失Loss1和第二预测损失Loss2对特征提取模型进行参数更新,并根据第二预测损失Loss2对语音识别模型进行参数更新。
根据第一预测损失Loss1和第二预测损失Loss2对特征提取模型进行参数更新使得基于特征提取模型能够从训练混合语音特征中提取出趋近于指定说 话人的标准语音特征且有利于语音识别的语音特征,将该语音特征输入语音识别模型进行语音识别,能够获得较好的语音识别效果。
第三实施例
在上述第三实施例的基础上,本实施例对第一实施例中的“利用预先建立的特征提取模型,以目标混合语音特征和目标说话人特征为依据,从目标混合语音特征中提取指定说话人的语音特征,以得到目标说话人的提取语音特征”的过程进行介绍。
利用预先建立的特征提取模型,以目标混合语音特征和目标说话人特征为依据,从目标混合语音特征中提取指定说话人的语音特征,以得到目标说话人的提取语音特征的过程可以包括:
步骤b1、将目标混合语音的语音特征和目标说话人的说话人特征输入特征提取模型,得到目标说话人对应的特征掩码。
其中,目标说话人对应的特征掩码能够表征目标混合语音的语音特征中目标说话人的语音特征的占比。
步骤b2、根据目标说话人对应的特征掩码,从目标混合语音的语音特征中提取目标说话人的语音特征,以得到目标说话人的提取语音特征。
具体的,将目标混合语音的语音特征与目标说话人对应的特征掩码逐帧相乘,从而得到目标说话人的提取语音特征。
在获得目标说话人的提取语音特征后,将目标说话人的提取语音特征和目标说话人的注册语音特征输入语音识别模型,得到目标说话人的语音识别结果,具体的,将目标说话人的提取语音特征和目标说话人的注册语音特征输入语音识别模型,得到目标说话人的语音识别结果的过程可以包括:
步骤c1、基于语音识别模型的编码器模块,分别对目标说话人的提取语音特征和目标说话人的注册语音特征进行编码,以得到两个编码结果。
步骤c2、基于语音识别模型的注意力模块,分别从两个编码结果中提取解码时刻解码所需的音频相关特征向量。
步骤c3、基于语音识别模型的解码器模块,对分别从两个编码结果中提取的音频相关特征向量进行解码,得到解码时刻的识别结果。
需要说明的是,将目标说话人的提取语音特征输入语音识别模型,得到所述目标说话人的语音识别结果的过程,与训练阶段将指定说话人的提取语音特 征和指定说话人的注册语音特征输入语音识别模型,得到指定说话人的语音识别结果的实现过程类似,步骤c1~步骤c3的具体实现过程可参见第二实施例中关于编码器模块、注意力模块和解码器模块的介绍,本实施例在此不做赘述。
经由上述第一实施例至第三实施例可知,本申请提供的语音识别方法具有如下几方面的优势:其一,本申请对目标说话人的注册语音提取多尺度声纹特征输入特征提取模型,增加特征提取模型输入信息的丰富程度,提升了特征提取模型的特征提取效果;其二,特征提取模型与语音识别模型进行联合训练,使得语音识别模型预测损失能够作用于特征提取模型,进而使得特征提取模型能够提取有利于语音识别的语音特征,从而能够提升语音识别结果的准确度;其三,将目标说话人的注册语音的语音特征作为语音识别模型的额外输入,以在特征提取模型提取的语音特征不佳时,能够辅助语音识别模型进行语音识别,从而得到较为准确的语音识别结果。综上,本申请提供的语音识别方法能够在复杂人声干扰的情况下,准确地识别出目标说话人的语音内容。
第四实施例
本申请实施例还提供了一种语音识别装置,下面对本申请实施例提供的语音识别装置进行描述,下文描述的语音识别装置与上文描述的语音识别方法可相互对应参照。
请参阅图4,示出了本申请实施例提供的语音识别装置的结构示意图,可以包括:特征获取模块401、特征提取模块402和语音识别模块403。
特征获取模块401,用于获取目标混合语音的语音特征以及目标说话人的说话人特征。
特征提取模块402,用于以趋于目标语音特征为提取方向,根据所述目标混合语音的语音特征以及所述目标说话人的说话人特征,从所述目标混合语音的语音特征中提取所述目标说话人的语音特征,以得到所述目标说话人的提取语音特征,其中,所述目标语音特征为用于获得与所述目标说话人的真实语音内容一致的语音识别结果的语音特征。
语音识别模块403,用于根据所述目标说话人的提取语音特征,获取所述目标说话人的语音识别结果。
可选的,特征获取模块401包括:语音特征获取模块和说话人特征获取模 块。
所述语音特征获取模块,用于获取目标混合语音的语音特征。
所述说话人特征获取模块,用于获取目标说话人的说话人特征。
可选的,所述说话人特征获取模块在获取目标说话人的说话人特征时,具体用于获取所述目标说话人的注册语音,对所述目标说话人的注册语音提取短时声纹特征和长时声纹特征,以得到多尺度声纹特征,作为所述目标说话人的说话人特征。
可选的,特征提取模块402具体用于利用预先建立的特征提取模型,以所述目标混合语音的语音特征以及所述目标说话人的说话人特征为依据,从所述目标混合语音的语音特征中提取所述目标说话人的语音特征。
其中,所述特征提取模型采用包含指定说话人的语音的训练混合语音的语音特征和指定说话人的说话人特征,以基于所述指定说话人的提取语音特征获取的语音识别结果为优化目标训练得到,所述指定说话人的提取语音特征为从所述训练混合语音的语音特征中提取的所述指定说话人的语音特征。
可选的,所述特征提取模型同时以所述指定说话人的提取语音特征和基于所述指定说话人的提取语音特征获取的语音识别结果为优化目标训练得到。
可选的,特征提取模块402可以包括:特征掩码确定子模块和语音特征提取子模块。
所述特征掩码确定子模块,用于将所述目标混合语音的语音特征以及所述目标说话人的说话人特征输入所述特征提取模型,得到所述目标说话人对应的特征掩码,其中,所述特征掩码能够表征对应说话人的语音特征在所述目标混合语音的语音特征中的占比。
所述语音特征提取子模块,用于根据所述目标混合语音的语音特征和所述目标说话人对应的特征掩码,从所述目标混合语音的语音特征中提取所述目标说话人的语音特征。
可选的,语音识别模块403,具体用于根据所述目标说话人的提取语音特征以及所述目标说话人的注册语音特征,获取所述目标说话人的语音识别结果;其中,所述目标说话人的注册语音特征为所述目标说话人的注册语音的语音特征。
可选的,语音识别模块403,具体用于将至少包括所述目标说话人的提取语音特征的语音识别输入特征输入预先建立的语音识别模型,得到所述目标说话人的语音识别结果。
其中,所述语音识别模型与所述特征提取模型联合训练得到,所述语音识别模型采用所述指定说话人的提取语音特征,以基于所述指定说话人的提取语音特征获得的语音识别结果为优化目标训练得到。
可选的,语音识别模块403在将至少包括所述目标说话人的提取语音特征的语音识别输入特征输入预先建立的语音识别模型,得到所述目标说话人的语音识别结果时,具体用于:
基于所述语音识别模型的编码器模块,对所述语音识别输入特征进行编码,以得到编码结果;基于所述语音识别模型的注意力模块,从所述编码结果中提取解码时刻解码所需的音频相关特征向量;基于所述语音识别模型的解码器模块,对从所述编码结果中提取的所述音频相关特征向量进行解码,得到所述解码时刻的识别结果。
可选的,本申请实施例提供的语音识别装置还可以包括:模型训练模块。模型训练模块可以包括:提取语音特征获取模块、语音识别结果获取模块、参数更新模块。
所述提取语音特征获取模块,用于利用特征提取模型,从所述训练混合语音的语音特征中提取所述指定说话人的语音特征,以得到所述指定说话人的提取语音特征。
所述语音识别结果获取模块,用于利用语音识别模型和所述指定说话人的提取语音特征,获取所述指定说话人的语音识别结果。
所述模型更新模块,用于根据所述指定说话人的提取语音特征和所述指定说话人的语音识别结果对特征提取模型进行参数更新,并根据所述指定说话人的语音识别结果对语音识别模型进行参数更新。
可选的,模型更新模块可以包括:标注文本获取模块、标准语音特征获取模块、预测损失确定模块和参数更新模块。
所述训练混合语音对应有所述指定说话人的语音。
所述标准语音特征获取模块,用于获取获取所述指定说话人的语音的语音 特征作为所述指定说话人的标准语音特征。
所述标注文本获取模块,用于获取所述指定说话人的语音的标注文本。
所述预测损失确定模块,用于根据所述指定说话人的提取语音特征和所述指定说话人的标准语音特征确定第一预测损失,并根据所述指定说话人的语音识别结果和所述指定说话人的语音的标注文本,确定第二预测损失。
所述参数更新模块,用于根据所述第一预测损失和所述第二预测损失对特征提取模型进行参数更新,并根据所述第二预测损失对语音识别模型进行参数更新。
可选的,所述训练混合语音以及所述训练混合语音对应的所述指定说话人的语音从预先构建的训练数据集中获取,本申请实施例提供的语音识别装置还可以包括:训练数据集构建模块。
所述训练数据集构建模块用于:
获取多个说话人的多条语音,其中,每条语音为单一说话人的语音,每条语音具有标注文本;将所述多条语音中的部分语音或全部语音中的每条语音作为指定说话人的语音:将其它语音中其他说话人的一条或多条语音与该指定说话人的语音进行混合,以得到一条训练混合语音,将该指定说话人的语音与通过混合得到的训练混合语音作为一条训练数据;由获得的所有训练数据组成所述训练数据集。
本申请实施例提供的语音识别装置能够从目标混合语音的语音特征中提取出目标说话人的语音特征,进而能够根据提取出的目标说话人的语音特征,获得目标说话人的语音识别结果,由于本申请实施例在从目标混合语音的语音特征提取目标说话人的语音特征时,以趋于目标语音特征(用于获得与目标目标说话人的真实语音内容一致的语音识别结果的语音特征)为提取方向,因此,提取出的语音特征为目标语音特征或者趋近于目标语音特征的语音特征,基于该语音特征进行语音识别,能够获得较好的语音识别效果,即能够获得较为准确的语音识别结果,用户体验较好。
第五实施例
本申请实施例还提供了一种语音识别设备,请参阅图5,示出了该语音识别设备的结构示意图,该语音识别设备可以包括:至少一个处理器501,至少 一个通信接口502,至少一个存储器503和至少一个通信总线504;
在本申请实施例中,处理器501、通信接口502、存储器503、通信总线504的数量为至少一个,且处理器501、通信接口502、存储器503通过通信总线504完成相互间的通信;
处理器501可能是一个中央处理器CPU,或者是特定集成电路ASIC(Application Specific Integrated Circuit),或者是被配置成实施本发明实施例的一个或多个集成电路等;
存储器503可能包含高速RAM存储器,也可能还包括非易失性存储器(non-volatile memory)等,例如至少一个磁盘存储器;
其中,存储器存储有程序,处理器可调用存储器存储的程序,所述程序用于:
获取目标混合语音的语音特征以及目标说话人的说话人特征;
以趋于目标语音特征为提取方向,根据目标混合语音的语音特征以及目标说话人的说话人特征,从目标混合语音的语音特征中提取目标说话人的语音特征,以得到目标说话人的提取语音特征,其中,所述目标语音特征为用于获得与目标说话人的真实语音内容一致的语音识别结果的语音特征;
根据所述目标说话人的提取语音特征,获取所述目标说话人的语音识别结果。
可选的,所述程序的细化功能和扩展功能可参照上文描述。
第六实施例
本申请实施例还提供一种可读存储介质,该可读存储介质可存储有适于处理器执行的程序,所述程序用于:
获取目标混合语音的语音特征以及目标说话人的说话人特征;
以趋于目标语音特征为提取方向,根据目标混合语音的语音特征以及目标说话人的说话人特征,从目标混合语音的语音特征中提取目标说话人的语音特征,以得到目标说话人的提取语音特征,其中,所述目标语音特征为用于获得与目标说话人的真实语音内容一致的语音识别结果的语音特征;
根据所述目标说话人的提取语音特征,获取所述目标说话人的语音识别结果。
可选的,所述程序的细化功能和扩展功能可参照上文描述。
最后,还需要说明的是,在本文中,诸如第一和第二等之类的关系术语仅仅用来将一个实体或者操作与另一个实体或操作区分开来,而不一定要求或者暗示这些实体或操作之间存在任何这种实际的关系或者顺序。而且,术语“包括”、“包含”或者其任何其他变体意在涵盖非排他性的包含,从而使得包括一系列要素的过程、方法、物品或者设备不仅包括那些要素,而且还包括没有明确列出的其他要素,或者是还包括为这种过程、方法、物品或者设备所固有的要素。在没有更多限制的情况下,由语句“包括一个……”限定的要素,并不排除在包括所述要素的过程、方法、物品或者设备中还存在另外的相同要素。
本说明书中各个实施例采用递进的方式描述,每个实施例重点说明的都是与其他实施例的不同之处,各个实施例之间相同相似部分互相参见即可。
对所公开的实施例的上述说明,使本领域专业技术人员能够实现或使用本发明。对这些实施例的多种修改对本领域的专业技术人员来说将是显而易见的,本文中所定义的一般原理可以在不脱离本发明的精神或范围的情况下,在其它实施例中实现。因此,本发明将不会被限制于本文所示的这些实施例,而是要符合与本文所公开的原理和新颖特点相一致的最宽的范围。

Claims (18)

  1. 一种语音识别方法,其特征在于,包括:
    获取目标混合语音的语音特征以及目标说话人的说话人特征;
    以趋于目标语音特征为提取方向,根据所述目标混合语音的语音特征以及所述目标说话人的说话人特征,从所述目标混合语音的语音特征中提取所述目标说话人的语音特征,以得到所述目标说话人的提取语音特征,其中,所述目标语音特征为用于获得与所述目标说话人的真实语音内容一致的语音识别结果的语音特征;
    根据所述目标说话人的提取语音特征,获取所述目标说话人的语音识别结果。
  2. 根据权利要求1所述的语音识别方法,其特征在于,获取所述目标说话人的说话人特征,包括:
    获取所述目标说话人的注册语音;
    对所述目标说话人的注册语音提取短时声纹特征和长时声纹特征,以得到多尺度声纹特征,作为所述目标说话人的说话人特征。
  3. 根据权利要求1所述的语音识别方法,其特征在于,所述以趋于目标语音特征为提取方向,根据所述目标混合语音的语音特征以及所述目标说话人的说话人特征,从所述目标混合语音的语音特征中提取所述目标说话人的语音特征,包括:
    利用预先建立的特征提取模型,以所述目标混合语音的语音特征以及所述目标说话人的说话人特征为依据,从所述目标混合语音的语音特征中提取所述目标说话人的语音特征;
    其中,所述特征提取模型采用包含指定说话人的语音的训练混合语音的语音特征和所述指定说话人的说话人特征,以基于所述指定说话人的提取语音特征获取的语音识别结果为优化目标训练得到,所述指定说话人的提取语音特征为从所述训练混合语音的语音特征中提取的所述指定说话人的语音特征。
  4. 根据权利要求3所述的语音识别方法,其特征在于,所述特征提取模型同时以所述指定说话人的提取语音特征和基于所述指定说话人的提取语音特征获取的语音识别结果为优化目标训练得到。
  5. 根据权利要求3或4所述的语音识别方法,其特征在于,所述利用预 先建立的特征提取模型,以所述目标混合语音的语音特征以及所述目标说话人的说话人特征为依据,从所述目标混合语音的语音特征中提取所述目标说话人的语音特征,包括:
    将所述目标混合语音的语音特征以及所述目标说话人的说话人特征输入所述特征提取模型,得到所述目标说话人对应的特征掩码;
    根据所述目标混合语音的语音特征和所述目标说话人对应的特征掩码,从所述目标混合语音的语音特征中提取所述目标说话人的语音特征。
  6. 根据权利要求1所述的语音识别方法,其特征在于,所述根据所述目标说话人的提取语音特征,获取所述目标说话人的语音识别结果,包括:
    根据所述目标说话人的提取语音特征以及所述目标说话人的注册语音特征,获取所述目标说话人的语音识别结果;
    其中,所述目标说话人的注册语音特征为所述目标说话人的注册语音的语音特征。
  7. 根据权利要求3或4所述的语音识别方法,其特征在于,所述根据所述目标说话人的提取语音特征,获取所述目标说话人的语音识别结果,包括:
    将至少包括所述目标说话人的提取语音特征的语音识别输入特征输入预先建立的语音识别模型,得到所述目标说话人的语音识别结果;
    所述语音识别模型与所述特征提取模型联合训练得到,所述语音识别模型采用所述指定说话人的提取语音特征,以基于所述指定说话人的提取语音特征获得的语音识别结果为优化目标训练得到。
  8. 根据权利要求7所述的语音识别方法,其特征在于,将所述语音识别输入特征输入所述语音识别模型,得到所述目标说话人的语音识别结果,包括:
    基于所述语音识别模型的编码器模块,对所述语音识别输入特征进行编码,以得到编码结果;
    基于所述语音识别模型的注意力模块,从所述编码结果中提取解码时刻解码所需的音频相关特征向量;
    基于所述语音识别模型的解码器模块,对从所述编码结果中提取的所述音频相关特征向量进行解码,得到所述解码时刻的识别结果。
  9. 根据权利要求7所述的语音识别方法,其特征在于,所述语音识别模型与所述特征提取模型联合训练的过程包括:
    利用特征提取模型,从所述训练混合语音的语音特征中提取所述指定说话人的语音特征,以得到所述指定说话人的提取语音特征;
    利用语音识别模型和所述指定说话人的提取语音特征,获取所述指定说话人的语音识别结果;
    根据所述指定说话人的提取语音特征和所述指定说话人的语音识别结果对特征提取模型进行参数更新,并根据所述指定说话人的语音识别结果对语音识别模型进行参数更新。
  10. 根据权利要求9所述的语音识别方法,其特征在于,所述训练混合语音对应有所述指定说话人的语音;
    所述根据所述指定说话人的提取语音特征和所述指定说话人的语音识别结果对特征提取模型进行参数更新,并根据所述指定说话人的语音识别结果对语音识别模型进行参数更新,包括:
    获取所述指定说话人的语音的标注文本,并获取所述指定说话人的语音的语音特征作为所述指定说话人的标准语音特征;
    根据所述指定说话人的提取语音特征和所述指定说话人的标准语音特征确定第一预测损失,并根据所述指定说话人的语音识别结果和所述指定说话人的语音的标注文本,确定第二预测损失;
    根据所述第一预测损失和所述第二预测损失对特征提取模型进行参数更新,并根据所述第二预测损失对语音识别模型进行参数更新。
  11. 根据权利要求10所述的语音识别方法,其特征在于,所述训练混合语音以及所述训练混合语音对应的所述指定说话人的语音从预先构建的训练数据集中获取;
    所述训练数据集的构建过程包括:
    获取多个说话人的多条语音,其中,每条语音为单一说话人的语音,每条语音具有标注文本;
    将所述多条语音中的部分语音或全部语音中的每条语音作为指定说话人的语音:将其它语音中其他说话人的一条或多条语音与该指定说话人的语音进行混合,以得到一条训练混合语音,将该指定说话人的语音与通过混合得到的训练混合语音作为一条训练数据;
    由获得的所有训练数据组成所述训练数据集。
  12. 一种语音识别装置,其特征在于,包括:特征获取模块、特征提取模块和语音识别模块;
    所述特征获取模块,用于获取目标混合语音的语音特征以及目标说话人的说话人特征;
    所述特征提取模块,用于以趋于目标语音特征为提取方向,根据所述目标混合语音的语音特征以及所述目标说话人的说话人特征,从所述目标混合语音的语音特征中提取所述目标说话人的语音特征,以得到所述目标说话人的提取语音特征,其中,所述目标语音特征为用于获得与所述目标说话人的真实语音内容一致的语音识别结果的语音特征;
    所述语音识别模块,用于根据所述目标说话人的提取语音特征,获取所述目标说话人的语音识别结果。
  13. 根据权利要求12所述的语音识别装置,其特征在于,所述特征获取模块包括:说话人特征获取模块;
    所述说话人特征获取模块,用于获取所述目标说话人的注册语音,对所述目标说话人的注册语音提取短时声纹特征和长时声纹特征,以得到多尺度声纹特征,作为所述目标说话人的说话人特征。
  14. 根据权利要求12所述的语音识别装置,其特征在于,所述特征提取模块具体用于利用预先建立的特征提取模型,以所述目标混合语音的语音特征以及所述目标说话人的说话人特征为依据,从所述目标混合语音的语音特征中提取所述目标说话人的语音特征;
    其中,所述特征提取模型采用包含指定说话人的语音的训练混合语音的语音特征和指定说话人的说话人特征,以基于所述指定说话人的提取语音特征获取的语音识别结果为优化目标训练得到,所述指定说话人的提取语音特征为从所述训练混合语音的语音特征中提取的所述指定说话人的语音特征。
  15. 根据权利要求12所述的语音识别装置,其特征在于,所述语音识别模块,具体用于根据所述目标说话人的提取语音特征以及所述目标说话人的注册语音特征,获取所述目标说话人的语音识别结果;
    其中,所述目标说话人的注册语音特征为所述目标说话人的注册语音的语音特征。
  16. 根据权利要求14所述的语音识别装置,其特征在于,所述语音识别 模块,用于将至少包括所述目标说话人的提取语音特征的语音识别输入特征输入预先建立的语音识别模型,得到所述目标说话人的语音识别结果;
    其中,所述语音识别模型与所述特征提取模型联合训练得到,所述语音识别模型采用所述指定说话人的提取语音特征,以基于所述指定说话人的提取语音特征获得的语音识别结果为优化目标训练得到。
  17. 一种语音识别设备,其特征在于,包括:存储器和处理器;
    所述存储器,用于存储程序;
    所述处理器,用于执行所述程序,实现如权利要求1~11中任一项所述的语音识别方法的各个步骤。
  18. 一种计算机可读存储介质,其上存储有计算机程序,其特征在于,所述计算机程序被处理器执行时,实现如权利要求1~11中任一项所述的语音识别方法的各个步骤。
PCT/CN2021/129733 2021-09-07 2021-11-10 一种语音识别方法、装置、设备及存储介质 Ceased WO2023035397A1 (zh)

Priority Applications (4)

Application Number Priority Date Filing Date Title
US18/689,668 US12626687B2 (en) 2021-09-07 2021-11-10 Speech recognition method, apparatus and device, and storage medium
EP21956569.4A EP4401074A4 (en) 2021-09-07 2021-11-10 SPEECH RECOGNITION METHOD, APPARATUS AND DEVICE, AND STORAGE MEDIUM
JP2024514680A JP7786691B2 (ja) 2021-09-07 2021-11-10 音声認識方法、装置、設備及び記憶媒体
KR1020247011082A KR20240050447A (ko) 2021-09-07 2021-11-10 음성 인식 방법, 장치, 디바이스 및 저장매체

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
CN202111042821.8 2021-09-07
CN202111042821.8A CN113724713B (zh) 2021-09-07 2021-09-07 一种语音识别方法、装置、设备及存储介质

Publications (1)

Publication Number Publication Date
WO2023035397A1 true WO2023035397A1 (zh) 2023-03-16

Family

ID=78682155

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2021/129733 Ceased WO2023035397A1 (zh) 2021-09-07 2021-11-10 一种语音识别方法、装置、设备及存储介质

Country Status (6)

Country Link
US (1) US12626687B2 (zh)
EP (1) EP4401074A4 (zh)
JP (1) JP7786691B2 (zh)
KR (1) KR20240050447A (zh)
CN (1) CN113724713B (zh)
WO (1) WO2023035397A1 (zh)

Cited By (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN117198272A (zh) * 2023-11-07 2023-12-08 浙江同花顺智能科技有限公司 一种语音处理方法、装置、电子设备及存储介质
CN119229875A (zh) * 2024-09-04 2024-12-31 武汉大学 一种基于多参考线索融合的目标语音提取方法及装置

Families Citing this family (7)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN114898756B (zh) * 2022-04-18 2025-08-15 北京荣耀终端有限公司 语音分离方法及装置
CN116486818A (zh) * 2022-08-30 2023-07-25 重庆蚂蚁消费金融有限公司 基于语音的身份识别方法、装置以及电子设备
CN116978359A (zh) * 2022-11-30 2023-10-31 腾讯科技(深圳)有限公司 音素识别方法、装置、电子设备及存储介质
CN115713939B (zh) * 2023-01-06 2023-04-21 阿里巴巴达摩院(杭州)科技有限公司 语音识别方法、装置及电子设备
CN116312503A (zh) * 2023-02-22 2023-06-23 哲库科技(上海)有限公司 语音数据的识别方法、装置、芯片及电子设备
CN116403603B (zh) * 2023-04-28 2025-09-05 科大讯飞股份有限公司 一种假音检测方法、假音检测模型获取方法及相关设备
CN119851670B (zh) * 2025-01-10 2025-09-23 中国科学技术大学 一种基于内容相关的帧级说话人声纹建模的语音匿名化方法

Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN109874096A (zh) * 2019-01-17 2019-06-11 天津大学 一种基于智能终端选择输出的双耳麦克风助听器降噪算法
CN110288989A (zh) * 2019-06-03 2019-09-27 安徽兴博远实信息科技有限公司 语音交互方法及系统
CN110827853A (zh) * 2019-11-11 2020-02-21 广州国音智能科技有限公司 语音特征信息提取方法、终端及可读存储介质
CN111128197A (zh) * 2019-12-25 2020-05-08 北京邮电大学 基于声纹特征与生成对抗学习的多说话人语音分离方法
CN111433847A (zh) * 2019-12-31 2020-07-17 深圳市优必选科技股份有限公司 语音转换的方法及训练方法、智能装置和存储介质

Family Cites Families (29)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN111081231B (zh) * 2016-03-23 2023-09-05 谷歌有限责任公司 用于多声道语音识别的自适应音频增强
KR102596430B1 (ko) * 2016-08-31 2023-10-31 삼성전자주식회사 화자 인식에 기초한 음성 인식 방법 및 장치
US11133011B2 (en) * 2017-03-13 2021-09-28 Mitsubishi Electric Research Laboratories, Inc. System and method for multichannel end-to-end speech recognition
US10699698B2 (en) 2018-03-29 2020-06-30 Tencent Technology (Shenzhen) Company Limited Adaptive permutation invariant training with auxiliary information for monaural multi-talker speech recognition
US10957337B2 (en) * 2018-04-11 2021-03-23 Microsoft Technology Licensing, Llc Multi-microphone speech separation
US10811000B2 (en) * 2018-04-13 2020-10-20 Mitsubishi Electric Research Laboratories, Inc. Methods and systems for recognizing simultaneous speech by multiple speakers
CN109166586B (zh) * 2018-08-02 2023-07-07 平安科技(深圳)有限公司 一种识别说话人的方法及终端
US11475898B2 (en) * 2018-10-26 2022-10-18 Apple Inc. Low-latency multi-speaker speech recognition
CN109326302B (zh) * 2018-11-14 2022-11-08 桂林电子科技大学 一种基于声纹比对和生成对抗网络的语音增强方法
US10923111B1 (en) * 2019-03-28 2021-02-16 Amazon Technologies, Inc. Speech detection and speech recognition
EP4047596B1 (en) * 2019-06-04 2025-02-19 Google LLC Two-pass end to end speech recognition
CN112331181B (zh) * 2019-07-30 2024-07-05 中国科学院声学研究所 一种基于多说话人条件下目标说话人语音提取方法
US20210065712A1 (en) * 2019-08-31 2021-03-04 Soundhound, Inc. Automotive visual speech recognition
JP7329393B2 (ja) * 2019-09-02 2023-08-18 日本電信電話株式会社 音声信号処理装置、音声信号処理方法、音声信号処理プログラム、学習装置、学習方法及び学習プログラム
CN110517698B (zh) * 2019-09-05 2022-02-01 科大讯飞股份有限公司 一种声纹模型的确定方法、装置、设备及存储介质
CN111145736B (zh) * 2019-12-09 2022-10-04 华为技术有限公司 语音识别方法及相关设备
CN111009237B (zh) * 2019-12-12 2022-07-01 北京达佳互联信息技术有限公司 语音识别方法、装置、电子设备及存储介质
CN111261146B (zh) 2020-01-16 2022-09-09 腾讯科技(深圳)有限公司 语音识别及模型训练方法、装置和计算机可读存储介质
CN111326143B (zh) * 2020-02-28 2022-09-06 科大讯飞股份有限公司 语音处理方法、装置、设备及存储介质
CN111508505B (zh) * 2020-04-28 2023-11-03 讯飞智元信息科技有限公司 一种说话人识别方法、装置、设备及存储介质
CN111583916B (zh) * 2020-05-19 2023-07-25 科大讯飞股份有限公司 一种语音识别方法、装置、设备及存储介质
CN111899727B (zh) * 2020-07-15 2022-05-06 思必驰科技股份有限公司 用于多说话人的语音识别模型的训练方法及系统
CN111833886B (zh) * 2020-07-27 2021-03-23 中国科学院声学研究所 全连接多尺度的残差网络及其进行声纹识别的方法
CN111899758B (zh) * 2020-09-07 2024-01-30 腾讯科技(深圳)有限公司 语音处理方法、装置、设备和存储介质
CN112735390B (zh) * 2020-12-25 2023-02-28 江西台德智慧科技有限公司 一种具有语音识别功能的智能语音终端设备
CN112599118B (zh) * 2020-12-30 2024-02-13 中国科学技术大学 语音识别方法、装置、电子设备和存储介质
CN112786057B (zh) * 2021-02-23 2023-06-02 厦门熵基科技有限公司 一种声纹识别方法、装置、电子设备及存储介质
CN113077795B (zh) * 2021-04-06 2022-07-15 重庆邮电大学 一种通道注意力传播与聚合下的声纹识别方法
CN113221673B (zh) * 2021-04-25 2024-03-19 华南理工大学 基于多尺度特征聚集的说话人认证方法及系统

Patent Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN109874096A (zh) * 2019-01-17 2019-06-11 天津大学 一种基于智能终端选择输出的双耳麦克风助听器降噪算法
CN110288989A (zh) * 2019-06-03 2019-09-27 安徽兴博远实信息科技有限公司 语音交互方法及系统
CN110827853A (zh) * 2019-11-11 2020-02-21 广州国音智能科技有限公司 语音特征信息提取方法、终端及可读存储介质
CN111128197A (zh) * 2019-12-25 2020-05-08 北京邮电大学 基于声纹特征与生成对抗学习的多说话人语音分离方法
CN111433847A (zh) * 2019-12-31 2020-07-17 深圳市优必选科技股份有限公司 语音转换的方法及训练方法、智能装置和存储介质

Non-Patent Citations (1)

* Cited by examiner, † Cited by third party
Title
See also references of EP4401074A4

Cited By (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN117198272A (zh) * 2023-11-07 2023-12-08 浙江同花顺智能科技有限公司 一种语音处理方法、装置、电子设备及存储介质
CN117198272B (zh) * 2023-11-07 2024-01-30 浙江同花顺智能科技有限公司 一种语音处理方法、装置、电子设备及存储介质
CN119229875A (zh) * 2024-09-04 2024-12-31 武汉大学 一种基于多参考线索融合的目标语音提取方法及装置

Also Published As

Publication number Publication date
US20240395242A1 (en) 2024-11-28
CN113724713A (zh) 2021-11-30
JP2024530353A (ja) 2024-08-16
KR20240050447A (ko) 2024-04-18
EP4401074A4 (en) 2025-07-23
US12626687B2 (en) 2026-05-12
JP7786691B2 (ja) 2025-12-16
EP4401074A1 (en) 2024-07-17
CN113724713B (zh) 2024-07-05

Similar Documents

Publication Publication Date Title
JP7786691B2 (ja) 音声認識方法、装置、設備及び記憶媒体
US20230127787A1 (en) Method and apparatus for converting voice timbre, method and apparatus for training model, device and medium
CN110956959B (zh) 语音识别纠错方法、相关设备及可读存储介质
CN113362813B (zh) 一种语音识别方法、装置和电子设备
WO2020107878A1 (zh) 文本摘要生成方法、装置、计算机设备及存储介质
CN108305617B (zh) 语音关键词的识别方法和装置
CN110516253B (zh) 中文口语语义理解方法及系统
CN108170686B (zh) 文本翻译方法及装置
CN108899013B (zh) 语音搜索方法、装置和语音识别系统
CN109003601A (zh) 一种针对低资源土家语的跨语言端到端语音识别方法
WO2019196196A1 (zh) 一种耳语音恢复方法、装置、设备及可读存储介质
CN114694255B (zh) 基于通道注意力与时间卷积网络的句子级唇语识别方法
CN110428820A (zh) 一种中英文混合语音识别方法及装置
CN108630199A (zh) 一种声学模型的数据处理方法
CN111161724B (zh) 中文视听结合语音识别方法、系统、设备及介质
CN107221330A (zh) 标点添加方法和装置、用于标点添加的装置
CN115394287A (zh) 混合语种语音识别方法、装置、系统及存储介质
CN115240712A (zh) 一种基于多模态的情感分类方法、装置、设备及存储介质
WO2020238045A1 (zh) 智能语音识别方法、装置及计算机可读存储介质
CN113889087B (zh) 语音识别及模型建立方法、装置、设备和存储介质
CN112017643A (zh) 语音识别模型训练方法、语音识别方法及相关装置
CN114373443A (zh) 语音合成方法和装置、计算设备、存储介质及程序产品
CN116564330A (zh) 弱监督语音预训练方法、电子设备和存储介质
CN109979461A (zh) 一种语音翻译方法及装置
CN116416966A (zh) 文本到语音合成方法、装置、设备和存储介质

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 21956569

Country of ref document: EP

Kind code of ref document: A1

WWE Wipo information: entry into national phase

Ref document number: 2024514680

Country of ref document: JP

WWE Wipo information: entry into national phase

Ref document number: 18689668

Country of ref document: US

ENP Entry into the national phase

Ref document number: 20247011082

Country of ref document: KR

Kind code of ref document: A

WWE Wipo information: entry into national phase

Ref document number: 2021956569

Country of ref document: EP

NENP Non-entry into the national phase

Ref country code: DE

ENP Entry into the national phase

Ref document number: 2021956569

Country of ref document: EP

Effective date: 20240408