WO2020029332A1 - 一种基于rnn的实时会议降噪方法及装置 - Google Patents
一种基于rnn的实时会议降噪方法及装置 Download PDFInfo
- Publication number
- WO2020029332A1 WO2020029332A1 PCT/CN2018/101820 CN2018101820W WO2020029332A1 WO 2020029332 A1 WO2020029332 A1 WO 2020029332A1 CN 2018101820 W CN2018101820 W CN 2018101820W WO 2020029332 A1 WO2020029332 A1 WO 2020029332A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- noise reduction
- rnn
- real
- signal
- time
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L21/00—Speech or voice signal processing techniques to produce another audible or non-audible signal, e.g. visual or tactile, in order to modify its quality or its intelligibility
- G10L21/02—Speech enhancement, e.g. noise reduction or echo cancellation
- G10L21/0208—Noise filtering
- G10L21/0216—Noise filtering characterised by the method used for estimating noise
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L21/00—Speech or voice signal processing techniques to produce another audible or non-audible signal, e.g. visual or tactile, in order to modify its quality or its intelligibility
- G10L21/02—Speech enhancement, e.g. noise reduction or echo cancellation
- G10L21/0208—Noise filtering
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/04—Architecture, e.g. interconnection topology
- G06N3/044—Recurrent networks, e.g. Hopfield networks
- G06N3/0442—Recurrent networks, e.g. Hopfield networks characterised by memory or gating, e.g. long short-term memory [LSTM] or gated recurrent units [GRU]
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/08—Learning methods
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/08—Learning methods
- G06N3/09—Supervised learning
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L21/00—Speech or voice signal processing techniques to produce another audible or non-audible signal, e.g. visual or tactile, in order to modify its quality or its intelligibility
- G10L21/02—Speech enhancement, e.g. noise reduction or echo cancellation
- G10L21/0208—Noise filtering
- G10L21/0216—Noise filtering characterised by the method used for estimating noise
- G10L21/0224—Processing in the time domain
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L21/00—Speech or voice signal processing techniques to produce another audible or non-audible signal, e.g. visual or tactile, in order to modify its quality or its intelligibility
- G10L21/02—Speech enhancement, e.g. noise reduction or echo cancellation
- G10L21/0208—Noise filtering
- G10L21/0216—Noise filtering characterised by the method used for estimating noise
- G10L21/0232—Processing in the frequency domain
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L25/00—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00
- G10L25/03—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 characterised by the type of extracted parameters
- G10L25/18—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 characterised by the type of extracted parameters the extracted parameters being spectral information of each sub-band
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L25/00—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00
- G10L25/27—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 characterised by the analysis technique
- G10L25/30—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 characterised by the analysis technique using neural networks
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L25/00—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00
- G10L25/45—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 characterised by the type of analysis window
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L25/00—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00
- G10L25/48—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 specially adapted for particular use
- G10L25/51—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 specially adapted for particular use for comparison or discrimination
- G10L25/60—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 specially adapted for particular use for comparison or discrimination for measuring the quality of voice signals
Definitions
- the invention relates to a method and a system for real-time conference noise reduction, in particular to a method and a device for real-time conference noise reduction based on RNN.
- Real-time conference systems have been solving a problem for decades, that is, the separation between noise and speech.
- the main problem is that during the real-time conference, when the microphone picks up the speaker, it will also pick up the surrounding environment noise (air conditioning noise, keyboard noise, environmental noise, etc.).
- the application direction is mainly divided into two aspects.
- the disadvantages are the high cost of the device and the troublesome deployment. And it has no effect on noise in the same direction as the speaker.
- One is to use a single microphone for traditional noise suppression, perform noise estimation based on the characteristics of the noise, and then perform noise suppression.
- the advantage is that it is easy to deploy, and the disadvantage is that it can only have a significant effect on the stable noise floor of the environment.
- the use of ambient noise floor is based on a stationary Gaussian distribution. However, the sudden noise cannot be accurately estimated.
- Deep learning is based on machine learning theory. It builds and simulates the human brain for analysis and learning of neural networks.
- Geoffery Hinton proposed a deep belief network (DNN) method (DNN) for deep belief networks (DBN, Deep Belief Neworks) superimposed by multiple restricted Boltzmann machines (RBM, Restricted Boltzmann machines). Belief Nets).
- DNN deep belief network
- BBM Restricted Boltzmann machines
- the invention provides a method and device for real-time conference noise reduction based on RNN, which aims to solve the problems that the current method of noise reduction for RNN has the disadvantages of poor real-time performance and large amount of calculation, and cannot be used in a real-time conference system.
- the present invention provides a method for real-time conference noise reduction based on RNN, including the following steps:
- the RNN-based real-time conference noise reduction method disclosed in the present invention obtains the logarithmic spectrum of the speech signal by framing and windowing the speech signal, and puts the logarithmic spectrum in the RNN model to determine the reduction.
- the noise suppression coefficient is obtained by reducing the noise reduction coefficient to the logarithmic spectrum of the original signal, and the application of the RNN noise reduction method in a real-time conference is realized.
- this general RNN model cannot handle long-range dependencies well, it is not suitable for noise reduction in real-time conferences, and the existence of LSTMs that can handle long-range dependencies well has a complicated structure and difficult design. The large amount of calculations and the difficulty of meeting real-time requirements cannot be used in real-time conference noise reduction.
- the structure of GRU is relatively simple and suitable for real-time conference noise reduction.
- the invention chose the RNN model using GRU as the RNN model.
- the present invention uses a GRU model, and by updating and resetting the gate, the information at the previous moments is retained to a certain extent to ensure the reliability of the training model.
- the present invention utilizes the characteristics that the GRU model can retain the information of the previous moments to a certain extent, and selects a suitable window length for the framed windowing of the speech signal, when inputting the RNN model for estimation, only the current frame needs to be input
- the logarithmic spectrum of the RNN model according to the present invention has low requirements for input information, and does not need to do a large amount of preprocessing on the received voice signals, which further reduces the amount of calculations, speeds up the response speed, and improves real-time performance.
- the present invention adopts the logarithmic spectrum in the frequency domain for the signal pairing. The logarithmic spectrum can extremely obviously highlight the signal change. When the signal comparison is performed, the logarithmic spectrum can be used to compare the signals.
- the present invention uses the frequency domain.
- the logarithmic spectrum is also used to further reduce the signal comparison time and improve the real-time performance of the present invention.
- the processing result is subjected to window overlap and inverse Fourier transform. Since the signal after noise reduction still undergoes window overlap processing, the signal can be further guaranteed. Processing effect to avoid excessive noise reduction and ensure speech integrity.
- the present invention reduces the computational load of the general RNN noise reduction method by performing framed windowing on the speech signal, adopting an appropriate window length, using an RNN model using GRU, and using a logarithmic spectrum in the frequency domain for signal comparison.
- the invention provides a real-time conference noise reduction method based on RNN, which solves the shortcomings of the current RNN noise reduction method, such as poor real-time performance, large calculation amount, and cannot be used in a real-time conference system, and provides a simple structure and signal processing. Simple, small amount of data calculation, high real-time, can be used in real-time meetings based on RNN noise reduction method.
- step S1 includes the following steps:
- the log spectrum of the noisy speech signal passes through a fully connected layer and two GRU layers to generate a corresponding estimated log spectrum, and the expected suppression parameters are obtained according to the estimated log spectrum and the log spectrum of the noisy speech signal.
- the present invention provides a method for real-time conference noise reduction based on RNN.
- RNN model By training an RNN model using GRU, appropriate parameters of the RNN model are determined.
- the training signal is composed of a pure speech signal and a noisy speech signal to synthesize a noisy speech signal, calculate the mean square error of the accurate suppression parameter and the expected suppression parameter, and use the mean square error derivative to update the parameters of the RNN model using GRU.
- the expected suppression parameter in the present invention is calculated from the estimated log spectrum and the log spectrum of the noisy speech signal, and the estimated log spectrum is obtained from the log spectrum of the noisy speech signal after passing through a fully connected layer and two GRU layers. generate.
- the one fully connected layer and two GRU layers are the RNN model structure used in the present invention.
- the RNN model is composed of one fully connected layer and two GRU layers.
- the model structure is simple, and the structure of the GRU layer itself is also relatively simple.
- the invention adopts a simple structure of the RNN model, which improves the real-time performance of signal noise reduction to a certain extent.
- this RNN model has a simple structure, it can also achieve very good noise reduction with the signal processing steps and the characteristics of the GRU itself. effect. Therefore, the RNN model provided by the present invention realizes the use of RNN noise reduction in real-time conferences, and solves the shortcomings of the current RNN noise reduction methods that are poor in real-time performance and large in computation, and cannot be used in real-time conference systems.
- An RNN-based noise reduction method with simple structure, simple signal processing, small amount of data calculation, and high real-time performance can be used in real-time conferences.
- the signal is framed and windowed, the window length is set to 256 samples, and the framed signal is overlapped by 50%.
- the optimal selection of the appropriate window length according to the present invention is 256 samples, which can improve the operation efficiency on the premise that information is not lost. Specifically, the information of the previous moments that the GRU model can retain is limited.
- the present invention also makes a design for the processing of voice signals. Since the minimum unit of a piece of speech can be divided into a basic syllable, the basic syllables of ordinary people are between 80ms and 160ms, and the fast Fourier transform is based on the data of 2 ⁇ N length to improve the calculation speed. Therefore, the window length is set in the present invention. It is 256 samples.
- the frame is divided into 256sample frames, the previous frame information will be lost.
- the GRU model can retain the information of previous moments to a certain extent, such a window length is set to ensure that the information is not lost.
- Calculation speed, and in the present invention the framed signals are overlapped by 50%, which can avoid inter-frame abrupt changes.
- a suitable window length is selected for the frame and window of the speech signal, so when inputting the RNN model for estimation, only the input of the current frame is required.
- Digital spectrum the RNN model of the present invention has low requirements on input information, and does not need to do a lot of preprocessing on the received voice signal, which further reduces the amount of calculation, speeds up the response speed, and improves the real-time performance.
- the present invention provides a real-time conference noise reduction device based on RNN, which includes a collecting device, a computing device, and a playback device; the collecting device collects a noisy voice signal and sends the signal to a computing device, and the computing device processes a band.
- the noisy speech signal is obtained and sent to the playback device; the computing device is a computing device using the RNN-based real-time conference noise reduction method according to any one of claims 1-3.
- the RNN-based real-time conference noise reduction device disclosed by the present invention collects voice signals through a collection device, performs voice noise reduction processing through a computing device, and plays through a playback device, where the computing device uses the present
- the invention describes a method for real-time meeting noise reduction based on RNN. Due to the RNN-based real-time meeting noise reduction method described in the present invention, the speech signal is framed and windowed to obtain the logarithmic spectrum of the speech signal. The logarithmic spectrum is placed in the RNN model to determine the noise reduction suppression coefficient.
- the noise reduction coefficient is converted to the logarithmic spectrum of the original signal to obtain the noise-reduced speech signal, and the application of the RNN noise reduction method in a real-time conference is realized.
- this general RNN model cannot handle long-range dependencies well, it is not suitable for noise reduction in real-time conferences, and the existence of LSTMs that can handle long-range dependencies well has a complicated structure and difficult design. The large amount of calculations and the difficulty of meeting real-time requirements cannot be used in real-time conference noise reduction.
- the structure of GRU is relatively simple and suitable for real-time conference noise reduction.
- the invention chose the RNN model using GRU as the RNN model.
- the present invention uses a GRU model, and by updating and resetting the gate, the information at the previous moments is retained to a certain extent to ensure the reliability of the training model.
- the information of the previous moments that the GRU model can retain is limited.
- the present invention also designs a speech signal processing. Since the minimum unit of a piece of speech can be divided into a basic syllable, the basic syllables of ordinary people are between 80ms and 160ms, and the fast Fourier transform is based on the data of 2 ⁇ N length to improve the calculation speed. Therefore, the window length is set in the present invention. It is 256 samples.
- the frame is divided into 256sample frames, the previous frame information will be lost.
- the GRU model can retain the information of previous moments to a certain extent, such a window length is set to ensure that the information is not lost.
- the framed signals are overlapped by 50%, which can avoid inter-frame abrupt changes.
- a suitable window length is selected for the frame and window of the speech signal, so when inputting the RNN model for estimation, only the input of the current frame is required.
- the RNN model of the present invention has low requirements on input information, and does not need to do a lot of preprocessing on the received voice signal, which further reduces the amount of calculation, speeds up the response speed, and improves the real-time performance.
- the present invention adopts the logarithmic spectrum in the frequency domain for the signal pairing.
- the logarithmic spectrum can extremely obviously highlight the signal change.
- the logarithmic spectrum can be used to compare the signals.
- the present invention uses the frequency domain.
- the logarithmic spectrum is also used to further reduce the signal comparison time and improve the real-time performance of the present invention.
- the present invention reduces the calculation amount of the general RNN noise reduction method by performing frame windowing on the speech signal, adopting an appropriate window length, using an RNN model using GRU, and using a logarithmic spectrum in the frequency domain for signal comparison. , Simplified model structure and signal processing process, improve real-time performance, and through the window overlap processing on the signal after noise reduction, to avoid excessive noise reduction and ensure speech integrity. Therefore, the RNN-based real-time conference noise reduction device disclosed in the present invention realizes the application of the RNN noise reduction method in a real-time conference system, and solves the problem that there is currently no RNN-based real-time conference noise reduction device.
- the acquisition device further includes a remote receiving unit, and the remote receiving unit is connected to the computing device.
- the invention discloses a RNN-based real-time conference noise reduction device.
- the acquisition device is used to collect voice information and send it to a computing device.
- the RNN-based real-time conference noise reduction device is used in a real-time conference system to reduce noise in a real-time conference.
- the current real-time conference noise reduction is often to perform noise reduction processing on the voice signal collected by the microphone, and the noise-reduced voice signal is sent to the playback device for playback through the network.
- the playback device here is often another real-time conference system.
- current real-time conferences often use two real-time conference systems to communicate with each other. One conference system sends voice information to the other conference system.
- the current real-time conference noise reduction is after the conference system receives the voice.
- Noise reduction is performed, and the information after noise reduction is sent to another conference service system, because the other conference system plays the noise reduction voice.
- the present invention provides a RNN-based real-time conference noise reduction device.
- a remote receiving unit is set in the acquisition device, that is, the acquisition device can accept the noise-reduced voice sent by other conference service systems, and send the received voice.
- the computing unit performs noise reduction once again, and the voice information after the noise reduction is sent by the computing device to the playback device for playback, that is, the conference system accepts the voice information of the noise reduction sent by other conference systems, and performs the noise reduction on the voice information. Noise reduction is performed again, and the conference service system plays the voice information of noise reduction again.
- the RNN-based real-time conference noise reduction device supports two communication parties in a real-time conference to reduce noise twice. Repeated noise reduction guarantees voice information.
- the noise reduction effect of the real-time conference is maximized, and because of the RNN-based real-time conference noise reduction device, the RNN model used is simple in structure and the voice signal processing method is simple, which makes the RNN-based
- the structure of the real-time conference noise reduction device is simple, the signal processing is simple, the amount of data calculation is small, and the real-time performance is high.
- the two noise reductions of the two communication parties also meet the real-time requirements of the real-time conference.
- a device for applying an RNN noise reduction method in a real-time conference system is provided. The meeting's requirements for real-time performance have solved the problem that there is no real-time meeting noise reduction device based on RNN.
- FIG. 1 is a flowchart of a method for real-time conference noise reduction based on RNN of the present invention
- FIG. 3 is a schematic structural diagram of a GRU layer of a RNN-based real-time conference noise reduction method according to the present invention.
- FIG. 4 is a schematic structural diagram of a fully connected layer of a RNN-based real-time conference noise reduction method according to the present invention.
- FIG. 5 is an activation function diagram of a fully connected layer of a real-time conference noise reduction method based on RNN of the present invention
- FIG. 6 is a structural block diagram 1 of an RNN-based real-time conference noise reduction device according to the present invention.
- FIG. 7 is a structural block diagram 2 of a real-time conference noise reduction device based on RNN of the present invention.
- the RNN-based real-time conference noise reduction method includes the following steps: S1. Training an RNN model using a GRU to determine appropriate parameters of the RNN model to obtain a trained RNN model; S2 , Frame-by-frame and window calculation of the speech signal transmitted by the acquisition unit to obtain the log spectrum of each frame of the speech signal in the frequency domain; S3, put the log spectrum of the current frame into the trained RNN model to estimate the current speech S4, estimate the logarithmic spectrum of the estimated log spectrum and the original signal, calculate the signal-to-noise ratio of the current frame, and calculate the current noise reduction suppression coefficient based on the signal-to-noise ratio; S5, apply noise reduction suppression Coefficient to logarithmic spectrum of the original signal, window overlap and inverse Fourier transform on the result, send it to the corresponding playback device through the network, and play the processed signal.
- the RNN-based real-time conference noise reduction method disclosed in the present invention obtains the logarithmic spectrum of the speech signal by framing and windowing the speech signal, and puts the logarithmic spectrum in the RNN model to determine the reduction.
- the noise suppression coefficient is obtained by reducing the noise reduction coefficient to the logarithmic spectrum of the original signal, and the application of the RNN noise reduction method in a real-time conference is realized.
- this general RNN model cannot handle long-range dependencies well, it is not suitable for noise reduction in real-time conferences, and the existence of LSTMs that can handle long-range dependencies well has a complicated structure and difficult design. The large amount of calculations and the difficulty of meeting real-time requirements cannot be used in real-time conference noise reduction.
- the structure of GRU is relatively simple and suitable for real-time conference noise reduction. The invention chose the RNN model using GRU as the RNN model.
- the operating principle of the GRU used in the present invention is as follows:
- the Sigmoid function restricts the result of passing two gates to [0,1].
- Is the candidate implicit state Use a reset gate to control the inflow of the last hidden state that contains information about past moments. The smaller the reset gate, the more the previous implicit state is discarded. Therefore, the reset gate provides a mechanism to discard past implicit states that are not related to the future.
- the hidden state h t uses an update gate z t to update the previous hidden state h t-1 and the candidate hidden state.
- Update gates can control the importance of past implicit states at the current moment. If the update gate is always close to 1, the past hidden state will be saved through time and passed to the current moment. This design can cope with the gradient attenuation problem in the recurrent neural network and better capture the larger interval dependencies in the time series data.
- the original input passes a link layer and two GRUs to generate the corresponding estimated signal, and then the MSE is obtained from the log spectrum of the clean speech signal before we synthesize it. Continuous activation and iteration through activation functions. Update the Dense and GRU parameters to obtain the minimum error.
- This method can be implemented by any hardware device with the function of calculating instructions.
- the invention uses multiple GPUs for training and speeds up the training process.
- the present invention uses a GRU model, and by updating and resetting the gate, the information at the previous moments is retained to a certain extent to ensure the reliability of the training model. Because the present invention utilizes the characteristics that the GRU model can retain the information of the previous moments to a certain extent, and selects a suitable window length for the framed windowing of the speech signal, when inputting the RNN model for estimation, only the current frame needs to be input.
- the logarithmic spectrum of the RNN model according to the present invention has low requirements for input information, and does not need to do a large amount of preprocessing on the received voice signals, which further reduces the amount of calculations, speeds up the response speed, and improves real-time performance. .
- the present invention adopts the logarithmic spectrum in the frequency domain for the signal pairing.
- the logarithmic spectrum can extremely obviously highlight the signal change.
- the logarithmic spectrum can be used to compare the signals.
- the present invention uses the frequency domain.
- the logarithmic spectrum is also used to further reduce the signal comparison time and improve the real-time performance of the present invention.
- the processing result is subjected to window overlap and inverse Fourier transform. Since the signal after noise reduction still undergoes window overlap processing, the signal can be further guaranteed. Processing effect to avoid excessive noise reduction and ensure speech integrity.
- the present invention reduces the computational load of the general RNN noise reduction method by performing framed windowing on the speech signal, adopting an appropriate window length, using an RNN model using GRU, and using a logarithmic spectrum in the frequency domain for signal comparison. , Simplified model structure and signal processing process, improve real-time performance, and through the window overlap processing on the signal after noise reduction, to avoid excessive noise reduction and ensure speech integrity.
- the invention provides a real-time conference noise reduction method based on RNN, which solves the shortcomings of the current RNN noise reduction method, such as poor real-time performance, large calculation amount, and cannot be used in a real-time conference system, and provides a simple structure and signal processing. Simple, small amount of data calculation, high real-time, can be used in real-time meetings based on RNN noise reduction method.
- the step S1 includes the following steps: S11, collecting pure voice signals and noisy voice signals, superimposing the pure voice signals and the noisy voice signals on the time domain to generate a noisy voice signal;
- the noisy speech signal and the pure speech signal are framed and windowed separately, and the log spectrum of each frame in the frequency domain is calculated.
- the log spectrum of the noisy speech signal and the log spectrum of the pure speech signal are compared to obtain the corresponding accurate suppression.
- S22 using the logarithmic spectrum of the noisy speech signal obtained after framed windowing as the input of the RNN model using GRU
- S23 the logarithmic spectrum of the noisy speech signal passing through a fully connected layer and two GRU layers Generate the corresponding estimated log spectrum, and obtain the expected suppression parameters based on the estimated log spectrum and the log spectrum of the noisy speech signal
- S24 calculate the mean square error using the expected suppression parameter and the accurate suppression parameter to determine whether the mean square error is less than a threshold value, If yes, end the step, if not, use the mean square error to perform differentiation, update the parameters of the RNN model using the GRU, and return to step S11.
- the step S1 of the present invention is the training step of the RNN model using the GRU. Specifically, in the conference room scene, we collect enough clean and clear speech data, and then collect enough data with only noise. We assume that the noise that needs to be processed is additive noise, so the two signals are superimposed in the time domain to generate noisy speech data.
- Feature extraction technology We perform feature extraction on noisy speech signals and pure speech signals, respectively.
- Feature extraction technology We use windowed short-time Fourier transform to avoid the problem of inter-frame mutation and overlap the framed signals by 50%.
- Fast Fourier transform based on 2 N length data can improve the calculation speed, so the window length is set to 256smaple.
- the window is defined as follows: After performing short-time Fourier transform according to the frame-length windowed frame length data, the log spectrum of each frame in the frequency domain is calculated and used as the input of the neural network.
- the present invention provides a method for real-time conference noise reduction based on RNN.
- RNN model By training an RNN model using GRU, appropriate parameters of the RNN model are determined.
- the training signal is composed of a pure speech signal and a noisy speech signal to synthesize a noisy speech signal, calculate the mean square error of the accurate suppression parameter and the expected suppression parameter, and use the mean square error derivative to update the parameters of the RNN model using GRU.
- the expected suppression parameter in the present invention is calculated from the estimated log spectrum and the log spectrum of the noisy speech signal, and the estimated log spectrum is obtained from the log spectrum of the noisy speech signal after passing through a fully connected layer and two GRU layers. generate.
- the one fully connected layer and two GRU layers are the RNN model structure used in the present invention.
- the RNN model is composed of one fully connected layer and two GRU layers.
- the model structure is simple, and the structure of the GRU layer itself is also relatively simple.
- the invention adopts a simple structure of the RNN model, which improves the real-time performance of signal noise reduction to a certain extent.
- this RNN model has a simple structure, it can also achieve very good noise reduction with the signal processing steps and the characteristics of the GRU itself. effect. Therefore, the RNN model provided by the present invention realizes the use of RNN noise reduction in real-time conferences, and solves the shortcomings of the current RNN noise reduction methods, which are poor in real-time performance and large in computation, and cannot be used in real-time conference systems.
- An RNN-based noise reduction method with simple structure, simple signal processing, small amount of data calculation, and high real-time performance can be used in real-time conferences.
- the activation function of the fully connected layer uses a tanh function, and the average value of the tanh function is 0.
- the structure of the fully connected layer is shown in Figure 4.
- the activation function uses tanh. As shown in Figure 5, the average value of tanh is 0, which can have better results in practical applications.
- the noise reduction suppression coefficient is obtained by smoothing a desired suppression parameter, and the expected suppression parameter is obtained by estimating a log spectrum and a log spectrum of an original signal. Since the noise reduction suppression coefficient is a smoothing process of the desired suppression parameter, the speech coherence is guaranteed.
- the framed windowing is performed on the signal, the window length is set to 256 samples, and the framed signal is overlapped by 50%.
- the optimal selection of the appropriate window length according to the present invention is 256 samples, which can improve the operation efficiency on the premise that information is not lost. Specifically, the information of the previous moments that the GRU model can retain is limited.
- the present invention also makes a design for the processing of voice signals. Since the minimum unit of a piece of speech can be divided into a basic syllable, the basic syllables of ordinary people are between 80ms and 160ms, and the fast Fourier transform is based on the data of 2 ⁇ N length to improve the calculation speed.
- the window length is set in the present invention. It is 256 samples. Although the frame is divided into 256sample frames, the previous frame information will be lost. However, because the GRU model can retain the information of previous moments to a certain extent, such a window length is set to ensure that the information is not lost. Calculation speed, and in the present invention, the framed signals are overlapped by 50%, which can avoid inter-frame abrupt changes. Because the present invention utilizes the feature that the GRU model can retain the information of the previous moments to a certain extent, a suitable window length is selected for the frame and window of the speech signal, so when inputting the RNN model for estimation, only the input of the current frame is required. Digital spectrum, the RNN model of the present invention has low requirements on input information, and does not need to do a lot of preprocessing on the received voice signal, which further reduces the amount of calculation, speeds up the response speed, and improves the real-time performance.
- the present invention provides a real-time conference noise reduction device based on RNN, which includes a collecting device, a computing device, and a playback device; the collecting device collects a noisy voice signal and sends it to the computing device, and the computing device processes The noisy speech signal is obtained and sent to the playback device; the computing device is a computing device using the RNN-based real-time conference noise reduction method according to any one of claims 1-3.
- the RNN-based real-time conference noise reduction device disclosed by the present invention collects voice signals through a collection device, performs voice noise reduction processing through a computing device, and plays through a playback device, where the computing device uses the present
- the invention describes a method for real-time meeting noise reduction based on RNN. Due to the RNN-based real-time meeting noise reduction method described in the present invention, the speech signal is framed and windowed to obtain the logarithmic spectrum of the speech signal. The logarithmic spectrum is placed in the RNN model to determine the noise reduction suppression coefficient. The noise reduction coefficient is converted to the logarithmic spectrum of the original signal to obtain the noise-reduced speech signal, and the application of the RNN noise reduction method in a real-time conference is realized.
- this general RNN model cannot handle long-range dependencies well, it is not suitable for noise reduction in real-time conferences, and the existence of LSTMs that can handle long-range dependencies well has a complicated structure and difficult design. The large amount of calculations and the difficulty of meeting real-time requirements cannot be used in real-time conference noise reduction.
- the structure of GRU is relatively simple and suitable for real-time conference noise reduction.
- the invention chose the RNN model using GRU as the RNN model.
- the present invention uses a GRU model, and by updating and resetting the gate, the information at the previous moments is retained to a certain extent to ensure the reliability of the training model.
- the information of the previous moments that the GRU model can retain is limited.
- the present invention also designs a speech signal processing. Since the minimum unit of a piece of speech can be divided into a basic syllable, the basic syllables of ordinary people are between 80ms and 160ms, and the fast Fourier transform is based on the data of 2 ⁇ N length to improve the calculation speed. Therefore, the window length is set in the present invention. It is 256 samples. Although the frame is divided into 256sample frames, the previous frame information will be lost. However, because the GRU model can retain the information of previous moments to a certain extent, such a window length is set to ensure that the information is not lost. Calculation speed, and in the present invention, the framed signals are overlapped by 50%, which can avoid inter-frame abrupt changes.
- the present invention utilizes the feature that the GRU model can retain the information of the previous moments to a certain extent, a suitable window length is selected for the frame and window of the speech signal, so when inputting the RNN model for estimation, only the input of the current frame is required.
- Digital spectrum the RNN model of the present invention has low requirements on input information, and does not need to do a lot of preprocessing on the received voice signal, which further reduces the amount of calculation, speeds up the response speed, and improves the real-time performance.
- the present invention adopts the logarithmic spectrum in the frequency domain for the signal pairing. The logarithmic spectrum can extremely obviously highlight the signal change. When the signal comparison is performed, the logarithmic spectrum can be used to compare the signals.
- the present invention uses the frequency domain.
- the logarithmic spectrum is also used to further reduce the signal comparison time and improve the real-time performance of the present invention.
- the processing result is subjected to window overlap and inverse Fourier transform. Since the signal after noise reduction still undergoes window overlap processing, the signal can be further guaranteed. Processing effect to avoid excessive noise reduction and ensure speech integrity.
- the present invention reduces the computational load of the general RNN noise reduction method by performing framed windowing on the speech signal, adopting an appropriate window length, using an RNN model using GRU, and using a logarithmic spectrum in the frequency domain for signal comparison.
- Simplified model structure and signal processing process improve real-time performance, and through the window overlap processing on the signal after noise reduction, to avoid excessive noise reduction and ensure speech integrity. Therefore, the RNN-based real-time conference noise reduction device disclosed in the present invention realizes the application of the RNN noise reduction method in a real-time conference system, and solves the problem that there is currently no RNN-based real-time conference noise reduction device.
- the acquisition device includes a microphone and an AD converter
- the microphone is connected to the computing device through the AD converter
- the playback device is connected to the computing device through a network.
- the microphone collects the ambient sound
- the AD converter converts the ambient sound into the digital signal required for calculation.
- the acquisition device further includes a remote receiving unit, and the remote receiving unit is connected to the computing device.
- the invention discloses a RNN-based real-time conference noise reduction device.
- the acquisition device is used to collect voice information and send it to a computing device.
- the RNN-based real-time conference noise reduction device is used in a real-time conference system to reduce noise in a real-time conference.
- the current real-time conference noise reduction is often to perform noise reduction processing on the voice signal collected by the microphone, and the noise-reduced voice signal is sent to the playback device for playback through the network.
- the playback device here is often another real-time conference system. Specifically, current real-time conferences often use two real-time conference systems to communicate with each other.
- One conference system sends voice information to the other conference system.
- the current real-time conference noise reduction is after the conference system receives the voice. Noise reduction is performed, and the information after noise reduction is sent to another conference service system, because the other conference system plays the noise reduction voice.
- the present invention provides a RNN-based real-time conference noise reduction device.
- a remote receiving unit is set in the acquisition device, that is, the acquisition device can accept the noise-reduced voice sent by other conference service systems, and send the received voice.
- the computing unit performs noise reduction once again, and the voice information after the noise reduction is sent by the computing device to the playback device for playback, that is, the conference system accepts the voice information of the noise reduction sent by other conference systems, and performs the noise reduction on the voice information.
- the RNN-based real-time conference noise reduction device provided by the present invention supports two communication parties in a real-time conference to reduce noise twice. Repeated noise reduction guarantees voice information.
- the noise reduction effect of the real-time conference is maximized, and because of the RNN-based real-time conference noise reduction device, the RNN model used is simple in structure and the voice signal processing method is simple, which makes the RNN-based
- the structure of the real-time conference noise reduction device is simple, the signal processing is simple, the amount of data calculation is small, and the real-time performance is high.
- the two noise reductions of the two communication parties also meet the real-time requirements of real-time conferences.
- a device for applying an RNN noise reduction method in a real-time conference system is provided, retaining the advantages of the better noise reduction effect of the RNN noise reduction method, while meeting the The real-time requirements of real-time conferences solve the problem that there is currently no real-time conference noise reduction device based on RNN.
- the computing device is a multi-CPU hardware device with a computing instruction function. Because this kind of real-time meeting noise reduction method based on RNN can be implemented by any hardware device with calculation instruction function.
- the invention adopts multiple CPUs for processing, which can speed up the processing process, improve operation efficiency, and improve real-time performance.
Landscapes
- Engineering & Computer Science (AREA)
- Physics & Mathematics (AREA)
- Health & Medical Sciences (AREA)
- Computational Linguistics (AREA)
- Human Computer Interaction (AREA)
- Audiology, Speech & Language Pathology (AREA)
- Signal Processing (AREA)
- Acoustics & Sound (AREA)
- Multimedia (AREA)
- Quality & Reliability (AREA)
- Theoretical Computer Science (AREA)
- Artificial Intelligence (AREA)
- Evolutionary Computation (AREA)
- Biomedical Technology (AREA)
- General Engineering & Computer Science (AREA)
- Biophysics (AREA)
- Data Mining & Analysis (AREA)
- General Health & Medical Sciences (AREA)
- Molecular Biology (AREA)
- Computing Systems (AREA)
- Life Sciences & Earth Sciences (AREA)
- General Physics & Mathematics (AREA)
- Mathematical Physics (AREA)
- Software Systems (AREA)
- Spectroscopy & Molecular Physics (AREA)
- Telephonic Communication Services (AREA)
- Telephone Function (AREA)
- Two-Way Televisions, Distribution Of Moving Picture Or The Like (AREA)
Abstract
本发明公开了一种基于RNN的实时会议降噪方法,对语音信号进行分帧加窗处理得到语音信号的对数谱,将对数谱放入RNN模型中确定降噪抑制系数,通过降噪抑制系数到原始信号的对数谱得到降噪后的语音信号,实现了将RNN降噪方法在实时会议中的应用。由于本发明利用了GRU模型可以一定程度上保留前些时刻的信息的特点为语音信号的分帧加窗选择了合适的窗长,所以在输入RNN模型进行估算时,仅需输入当前帧的对数谱,本发明所述的RNN模型对输入信息的要求低,无需对接收到的语音信号做大量的预处理,这也进一步的减少了运算量,加快了响应速度,提高了实时性,提供了一种可以在实时会议中使用的基于RNN的降噪方法。
Description
本发明涉及实时会议降噪方法及系统,尤其涉及一种基于RNN的实时会议降噪方法及装置。
实时会议系统数十年来都在解决一个问题,就是噪声与语音的是分离。主要解决的问题是实时会议过程中,当麦克风拾到发言人的同时也会同时拾取到周围环境的噪声(空调噪声,键盘噪声,环境底噪等)。现在应用方向主要分成两个方面。
一种是依靠多个麦克风装置的麦克风阵列降噪,这种是依赖多个麦克风同时拾音,计算多个信号之间的相位差,获取发声源的空间信息。通过MVDR等技术消去旁边的声源,提高信噪比。但是缺陷是需要装置成本高,部署麻烦。而且对与发言人同一个方向的噪声没有效果。一种是使用单个麦克风进行传统噪声抑制,通过噪声的特性进行噪声估计,然后再进行噪声抑制。优点是部署简易,缺点是只能针对环境平稳底噪有明显效果。利用环境底噪是基于平稳高斯分布。但是对突发性噪声无法进行正确的估计。
近些年来有开始利用深度学习技术进行噪声抑制的技术。深度学习基于机器学习理论,通过建立和模拟人脑进行分析学习的神经网络。Geoffery Hinton提出一种由多个受限玻尔兹曼机(RBM,Restricted Boltzmann machines)叠加而成的深度信念网络(DBN,Deep Belief Neworks)的深度学习(DNN)方法(A Fast Learning Algorithm for Deep Belief Nets)。近两年开始有通过大量标定的噪声数据和语音进行学习语音的特性,来进行语音降噪的功能。这种利用大量标定的噪声数据和语音学习语音的特性,来进行语音降噪的方法可以通过回归神经网 络(RNN)来实现,但是目前利用RNN实现的语音降噪依旧存在许多问题限制了基于RNN的降噪方法在实时会议中的应用,首当其冲的一点是目前的RNN语音降噪方法难以满足实时会议提出的实时性要求,其次是基于RNN降噪方法的数据处理量大难以集成与实时会务系统中使用。由于RNN降噪方法存在实时性差,运算量大的缺点,致使这种具有较好效果的降噪方法不能使用在实时会议系统中,对此,人们需要只用可以在实时会议中使用的RNN降噪方法。
发明内容
本发明提供了一种基于RNN的实时会议降噪方法及装置,旨在解决目前RNN降噪方法存在实时性差,运算量大的缺点,不能在实时会议系统中使用的问题。
为实现上述目的,本发明提供了一种基于RNN的实时会议降噪方法,包括以下步骤:
S1,对使用GRU的RNN模型进行训练确定RNN模型的合适参数,得到训练完成的RNN模型;
S2,对采集单元传输的语音信号进行分帧加窗,计算得到语音信号每帧在频域上的对数谱;
S3,将当前帧的对数谱放入训练完成的RNN模型进行估算,得到当前语音的估计对数谱;
S4,根据估计对数谱与原始信号的对数谱进行估计,算出当前帧的信噪比,根据信噪比计算出当前的降噪抑制系数;
S5,应用降噪抑制系数到原始信号的对数谱,对结果进行窗重叠和傅里叶逆变换,通过网络发送到对应的播放设备上,对处理后的信号进行播放。
与现有技术相比,本发明公开的一种基于RNN的实时会议降噪方法,对语音信号进行分帧加窗处理得到语音信号的对数谱,将对数谱放入RNN模型中确定降噪抑制系数,通过降噪抑制系数到原始信号的对数谱得到降噪后的语音信号,实 现了将RNN降噪方法在实时会议中的应用。具体而言,由于本通常的RNN模型由于无法很好处理远距离依赖而不适用于实时会议中的降噪,而对于可以很好处理远距离依赖的LSTM又存在则结构复杂,设计困难大,运算量大,实时性要求难以满足等问题,而不能使用在实时会议降噪中,而GRU的结构相对简单适合使用在实时会议降噪中,发明在RNN模型选择了采用GRU的RNN模型。本发明使用GRU的模型,通过更新门和重置门,一定程度上保留前些时刻的信息,保证训练模型的可靠性。由于本发明利用了GRU模型可以一定程度上保留前些时刻的信息的特点,并为语音信号的分帧加窗选择了合适的窗长,所以在输入RNN模型进行估算时,仅需输入当前帧的对数谱,本发明所述的RNN模型对输入信息的要求低,无需对接收到的语音信号做大量的预处理,这也进一步的减少了运算量,加快了响应速度,提高了实时性。同时,本发明在进行信号对于是采用的是频域上的对数谱,对数谱可以极其明显的突出信号变化,在进行信号对比时采用对数谱可以方面信号对比,本发明采用频域上的对数谱也是为了进一步减少信号对比时间,提高本发明的实时性。最后,在通过降噪抑制系数对原始信号的对数谱进行处理后,对处理结果进行窗重叠和傅里叶逆变换,由于降噪后的信号依然经过窗重叠处理,可以进一步的保证信号的处理效果,避免过分降噪,保证语音完整。对此,本发明通过对语音信号进行分帧加窗,采取合适的窗长度,采用使用GRU的RNN模型,采用频域上的对数谱进行信号对比,降低了一般RNN降噪方法的运算量,简化的模型结构和信号处理过程,提高了实时性,并且通过对降噪后的信号依然经过窗重叠处理,避免过分降噪,保证语音完整。本发明提供的一种基于RNN的实时会议降噪方法,解决目前RNN降噪方法存在实时性差,运算量大的缺点,不能在实时会议系统中使用的问题,提供了一种结构简单,信号处理简单,数据运算量小,实时性高的,可以在实时会议中使用的基于RNN的降噪方法。
进一步,所述步骤S1包括以下步骤:
S11,采集纯净语音信号和噪音语音信号,对纯净语音信号和噪音语音信号进行 时域上的叠加,产生带噪语音信号;
S12,对带噪语音信号和纯净语音信号分别进行分帧加窗,计算每帧在频域上的对数谱,将带噪语音信号的对数谱和纯净语音信号的对数谱进行对比得到对应的准确抑制参数;
S22,将分帧加窗后得到的带噪语音信号的对数谱作为使用GRU的RNN模型的输入;
S23,带噪语音信号的对数谱经过一个全连接层和两个GRU层后生成对应的估计对数谱,根据估计对数谱和带噪语音信号的对数谱得到期望抑制参数;
S24,使用期望抑制参数和准确抑制参数计算均方误差,判断均方误差是否小于阈值,是则结束步骤,不是则利用均方误差进行求导,更新使用GRU的RNN模型的参数并返回步骤S11。
本发明提供的一种基于RNN的实时会议降噪方法,通过对采用GRU的RNN模型进行训练,确定RNN模型的合适参数。所述训练信号通过纯净语音信号和噪音语音信号合成带噪语音信号,计算准确抑制参数和期望抑制参数的均方差,利用均方误差求导更新使用GRU的RNN模型的参数。本发明中的期望抑制参数由估计对数谱和带噪语音信号的对数谱计算得到,所述估计对数谱由带噪语音信号的对数谱经过一个全连接层和两个GRU层后生成。所述的一个全连接层和两个GRU层即为本发明采用的RNN模型结构,该RNN模型由一个全连接层和两个GRU层构成,模型结构简单,并且GRU层本身的结构也较为简单,本发明采用RNN模型结构简单,一定程度上提高了信号降噪的实时性,并且,这种RNN模型虽然结构简单,但是配合信号处理步骤和GRU本身的特点,也可以实现非常好的降噪效果。故而,本发明提供的这种RNN模型,实现了RNN降噪在实时会议中的使用,解决目前RNN降噪方法存在实时性差,运算量大的缺点,不能在实时会议系统中使用的问题,提供了一种结构简单,信号处理简单,数据运算量小,实时性高的,可以在实时会议中使用的基于RNN的降噪方法。
进一步,所述对信号进行分帧加窗,设置窗长为256样本,对分帧信号进行50%重叠。
本发明所述的适当窗长的最优选择为256样本,可以在保证信息不丢失的前提下,提升运算效率。具体而言,GRU模型可以保留的前些时刻信息有限,为了满足实时会议对降噪效果和实时性的要求,本发明还在语音信号的处理上作出了设计。由于一段语音的最小单位可以分割成一个基本音节,普通人讲话的基本音节都在80ms~160ms之间,而快速傅立叶变换基于2^N长度的数据进行可以提高计算速度,故而本发明设置窗长为256样本,虽然分帧成256sample一帧的话,就会丢失前面帧信息,但由于GRU模型可以一定程度上保留前些时刻的信息,这样的窗长设置在保证不丢失信息的前提下,提高的计算速度,而本发明中对分帧信号进行50%重叠,可以避免帧间突变。由于本发明利用了GRU模型可以一定程度上保留前些时刻的信息的特点为语音信号的分帧加窗选择了合适的窗长,所以在输入RNN模型进行估算时,仅需输入当前帧的对数谱,本发明所述的RNN模型对输入信息的要求低,无需对接收到的语音信号做大量的预处理,这也进一步的减少了运算量,加快了响应速度,提高了实时性。
为实现上述目的,本发明提供了一种基于RNN的实时会议降噪装置,包括采集装置,计算装置和播放装置;所述采集装置采集带噪语音信号发送至计算装置,所述计算装置处理带噪语音信号得到降噪语音信号发送至播放装置;所述计算装置为采用权利要求1-3任一项所述的一种基于RNN的实时会议降噪方法的计算装置。
与现有技术相比,本发明公开的一种基于RNN的实时会议降噪装置,通过采集装置采集语音信号,通过计算装置进行语音降噪处理,通过播放装置进行播放,其中计算装置采用了本发明所述的一种基于RNN的实时会议降噪方法。由于本发明所述的一种基于RNN的实时会议降噪方法,对语音信号进行分帧加窗处理得到语音信号的对数谱,将对数谱放入RNN模型中确定降噪抑制系数,通过降噪抑制 系数到原始信号的对数谱得到降噪后的语音信号,实现了将RNN降噪方法在实时会议中的应用。具体而言,由于本通常的RNN模型由于无法很好处理远距离依赖而不适用于实时会议中的降噪,而对于可以很好处理远距离依赖的LSTM又存在则结构复杂,设计困难大,运算量大,实时性要求难以满足等问题,而不能使用在实时会议降噪中,而GRU的结构相对简单适合使用在实时会议降噪中,发明在RNN模型选择了采用GRU的RNN模型。本发明使用GRU的模型,通过更新门和重置门,一定程度上保留前些时刻的信息,保证训练模型的可靠性。但与LSTM相比,GRU模型可以保留的前些时刻信息有限,为了满足实时会议对降噪效果和实时性的要求,本发明还在语音信号的处理上作出了设计。由于一段语音的最小单位可以分割成一个基本音节,普通人讲话的基本音节都在80ms~160ms之间,而快速傅立叶变换基于2^N长度的数据进行可以提高计算速度,故而本发明设置窗长为256样本,虽然分帧成256sample一帧的话,就会丢失前面帧信息,但由于GRU模型可以一定程度上保留前些时刻的信息,这样的窗长设置在保证不丢失信息的前提下,提高的计算速度,而本发明中对分帧信号进行50%重叠,可以避免帧间突变。由于本发明利用了GRU模型可以一定程度上保留前些时刻的信息的特点为语音信号的分帧加窗选择了合适的窗长,所以在输入RNN模型进行估算时,仅需输入当前帧的对数谱,本发明所述的RNN模型对输入信息的要求低,无需对接收到的语音信号做大量的预处理,这也进一步的减少了运算量,加快了响应速度,提高了实时性。同时,本发明在进行信号对于是采用的是频域上的对数谱,对数谱可以极其明显的突出信号变化,在进行信号对比时采用对数谱可以方面信号对比,本发明采用频域上的对数谱也是为了进一步减少信号对比时间,提高本发明的实时性。最后,在通过降噪抑制系数对原始信号的对数谱进行处理后,对处理结果进行窗重叠和傅里叶逆变换,由于降噪后的信号依然经过窗重叠处理,可以进一步的保证信号的处理效果,避免过分降噪,保证语音完整。对此,本发明通过对语音信号进行分帧加窗,采取合适的窗长度,采用使用GRU的RNN模型, 采用频域上的对数谱进行信号对比,降低了一般RNN降噪方法的运算量,简化的模型结构和信号处理过程,提高了实时性,并且通过对降噪后的信号依然经过窗重叠处理,避免过分降噪,保证语音完整。故而,本发明公开的一种基于RNN的实时会议降噪装置,实现了RNN降噪方法在实时会务系统中的应用,解决了目前没有基于RNN的实时会议降噪装置的问题。
进一步,所述采集装置还包括远程接收单元,所述远程接收单元与计算装置连接。
本发明公开的一种基于RNN的实时会议降噪装置,采集装置用于采集语音信息发送至计算装置,由于该种基于RNN的实时会议降噪装置是应用在实时会务系统对实时会议进行降噪的,目前的实时会议降噪往往是对麦克风采集的语音信号进行降噪处理,降噪后的语音信号通过网络发送至播放装置进行播放,此处的播放装置往往是另一台实时会务系统。具体而言,目前的实时会议往往使用两个实时会务系统相互通信,由一台会务系统向另一台会务系统发送语音信息,目前的实时会议降噪则是在本会务系统接收到语音后即进行降噪,将降噪后的信息发送至另一台会务系统,由于另一台会务系统播放降噪后的语音。但是,本发明提供的一种基于RNN的实时会议降噪装置,在采集装置中设置远程接收单元,即采集装置可以接受其他会务系统发送的已经经过降噪的语音,而对接收到的语音发送的计算单元再一次进行降噪,而再次降噪后的语音信息有计算装置发送至播放装置进行播放,即本会务系统接受其他会务系统发送的已经降噪的语音信息,对已降噪语音信息进行再次降噪,本会务系统播放再次降噪的语音信息,本发明提供的一种基于RNN的实时会议降噪装置,支持实时会议中的通信双方两次降噪,重复降噪保障了语音信息的降噪效果,最大程度的出去了实时会议中的语音噪音,并且,由于该种基于RNN的实时会议降噪装置,采用的RNN模型结构简单,语音信号处理方法简单,使得该种基于RNN的实时会议降噪装置的结构简单,信号处理简单,数据运算量小,实时性高,通信双方两次降噪同样满足实时会议的实时 性要求,提供了一种RNN降噪方法在实时会务系统中应用的装置,保留了RNN降噪方法较好降噪效果的优点,同时满足了实时会议对实时性的要求,解决了目前没有基于RNN的实时会议降噪装置的问题。
图1是本发明一种基于RNN的实时会议降噪方法的流程图;
图2是本发明一种基于RNN的实时会议降噪方法步骤S1的流程图;
图3是本发明一种基于RNN的实时会议降噪方法的GRU层的结构示意图;
图4是本发明一种基于RNN的实时会议降噪方法的全连接层的结构示意图;
图5是本发明一种基于RNN的实时会议降噪方法的全连接层的激活函数图;
图6是本发明一种基于RNN的实时会议降噪装置的结构框图1;
图7是本发明一种基于RNN的实时会议降噪装置的结构框图2。
如图1所示,本发明所述一种基于RNN的实时会议降噪方法,包括以下步骤:S1,对使用GRU的RNN模型进行训练确定RNN模型的合适参数,得到训练完成的RNN模型;S2,对采集单元传输的语音信号进行分帧加窗计算得到语音信号每帧在频域上的对数谱;S3,将当前帧的对数谱放入训练完成的RNN模型进行估算,得到当前语音的估计对数谱;S4,根据估计对数谱与原始信号的对数谱进行估计,算出当前帧的信噪比,根据信噪比计算出当前的降噪抑制系数;S5,应用降噪抑制系数到原始信号的对数谱,对结果进行窗重叠和傅里叶逆变换,通过网络发送到对应的播放设备上,对处理后的信号进行播放。
与现有技术相比,本发明公开的一种基于RNN的实时会议降噪方法,对语音信号进行分帧加窗处理得到语音信号的对数谱,将对数谱放入RNN模型中确定降噪抑制系数,通过降噪抑制系数到原始信号的对数谱得到降噪后的语音信号,实 现了将RNN降噪方法在实时会议中的应用。
具体而言,由于本通常的RNN模型由于无法很好处理远距离依赖而不适用于实时会议中的降噪,而对于可以很好处理远距离依赖的LSTM又存在则结构复杂,设计困难大,运算量大,实时性要求难以满足等问题,而不能使用在实时会议降噪中,而GRU的结构相对简单适合使用在实时会议降噪中,发明在RNN模型选择了采用GRU的RNN模型。
如图3所示,本发明使用的GRU的运行原理如下:
重置门r
t定义如下所示:r
t=σ(W
t×[h
t-1,x
t]);
更新门z
t定义如下所示:z
t=σ(W
z×[h
t-1,x
t]);
Sigmoid函数将通过两个门的结果限制在[0,1]。
隐含状态h
t使用更新门z
t来对上一个隐含状态h
t-1和候选隐含状态进行更新。更新门可以控制过去的隐含状态在当前时刻的重要性。如果更新门一直近似1,过去的隐含状态将一直通过时间保存并传递至当前时刻。这个设计可以应对循环神经网络中的梯度衰减问题,并更好地捕捉时序数据中间隔较大的依赖关系。
y
t=σ(W
0×h
t)
原始输入经过一个链接层和两个GRU之后生成对应估计信号,再与我们合成之前的干净语音信号的对数谱求得MSE。通过激活函数进行不断收敛和迭代。更新Dense和GRU的参数获取最小误差。该方法可以有任何具有计算指令功能的硬件设备实施。本发明采用多个GPU进行训练,加速训练过程。
本发明使用GRU的模型,通过更新门和重置门,一定程度上保留前些时刻的信息,保证训练模型的可靠性。由于本发明利用了GRU模型可以一定程度上保留前些时刻的信息的特点,并为语音信号的分帧加窗选择了合适的窗长,所以在输入RNN模型进行估算时,仅需输入当前帧的对数谱,本发明所述的RNN模型对输入信息的要求低,无需对接收到的语音信号做大量的预处理,这也进一步的减少了运算量,加快了响应速度,提高了实时性。同时,本发明在进行信号对于是采用的是频域上的对数谱,对数谱可以极其明显的突出信号变化,在进行信号对比时采用对数谱可以方面信号对比,本发明采用频域上的对数谱也是为了进一步减少信号对比时间,提高本发明的实时性。最后,在通过降噪抑制系数对原始信号的对数谱进行处理后,对处理结果进行窗重叠和傅里叶逆变换,由于降噪后的信号依然经过窗重叠处理,可以进一步的保证信号的处理效果,避免过分降噪,保证语音完整。
对此,本发明通过对语音信号进行分帧加窗,采取合适的窗长度,采用使用GRU的RNN模型,采用频域上的对数谱进行信号对比,降低了一般RNN降噪方法的运算量,简化的模型结构和信号处理过程,提高了实时性,并且通过对降噪后的信号依然经过窗重叠处理,避免过分降噪,保证语音完整。本发明提供的一种基于RNN的实时会议降噪方法,解决目前RNN降噪方法存在实时性差,运算量大的缺点,不能在实时会议系统中使用的问题,提供了一种结构简单,信号处理简单,数据运算量小,实时性高的,可以在实时会议中使用的基于RNN的降噪方法。
如图2所示,所述步骤S1包括以下步骤:S11,采集纯净语音信号和噪音语音信号,对纯净语音信号和噪音语音信号进行时域上的叠加,产生带噪语音信号;S12,对带噪语音信号和纯净语音信号分别进行分帧加窗,计算每帧在频域上的对数谱,将带噪语音信号的对数谱和纯净语音信号的对数谱进行对比得到对应的准确抑制参数;S22,将分帧加窗后得到的带噪语音信号的对数谱作为使用GRU的RNN模型的输入;S23,带噪语音信号的对数谱经过一个全连接层和两个GRU层后 生成对应的估计对数谱,根据估计对数谱和带噪语音信号的对数谱得到期望抑制参数;S24,使用期望抑制参数和准确抑制参数计算均方误差,判断均方误差是否小于阈值,是则结束步骤,不是则利用均方误差进行求导,更新使用GRU的RNN模型的参数并返回步骤S11。
本发明所述步骤S1即采用GRU的RNN模型的训练步骤,具体而言,在会议室场景中我们收集足够多干净清晰的语音部分数据,再收集足够的只有噪声部分的数据。我们假定需要处理的噪声都属于加性噪声,所以在时域上对两个信号进行叠加,产生带噪声的语音数据。如下公式所示:SN
t=Speech
t+Noise
t。
我们分别对带噪语音信号和纯净语音信号进行特征提取。特征提取技术我们采用加窗后的短时傅立叶变换,避免帧间突变问题,对分帧的信号进行50%重叠。基于2
N长度的数据进行快速傅立叶变换可以提高计算速度,所以设置窗长256smaple。窗如下定义:
根据分帧加窗的帧长数据进行短时傅立叶变换后,计算每帧在频域的上的对数谱,作为神经网络的输入。
本发明提供的一种基于RNN的实时会议降噪方法,通过对采用GRU的RNN模型进行训练,确定RNN模型的合适参数。所述训练信号通过纯净语音信号和噪音语音信号合成带噪语音信号,计算准确抑制参数和期望抑制参数的均方差,利用均方误差求导更新使用GRU的RNN模型的参数。本发明中的期望抑制参数由估计对数谱和带噪语音信号的对数谱计算得到,所述估计对数谱由带噪语音信号的对数谱经过一个全连接层和两个GRU层后生成。所述的一个全连接层和两个GRU层即为本发明采用的RNN模型结构,该RNN模型由一个全连接层和两个GRU层构成,模型结构简单,并且GRU层本身的结构也较为简单,本发明采用RNN模型结构简单,一定程度上提高了信号降噪的实时性,并且,这种RNN模型虽然结构简单,但是配合信号处理步骤和GRU本身的特点,也可以实现非常好的降噪效果。故而,本发明提供的这种RNN模型,实现了RNN降噪在实时会议中的使用,解决目前RNN降噪方法存在实时性差,运算量大的缺点,不能在实时会议系统中使用的问题, 提供了一种结构简单,信号处理简单,数据运算量小,实时性高的,可以在实时会议中使用的基于RNN的降噪方法。
所述全连接层的激活函数采用tanh函数,所述tanh函数的均值为0。全连接层的结构如图4,激活函数采用tanh。如图5所示,tanh的均值为0,实际应用中可以有较好的效果。
所述降噪抑制系数为期望抑制参数进行平滑处理得到,所述期望抑制参数为计对数谱与原始信号的对数谱进行估计得到。由于降噪抑制系数为期望抑制参数进行平滑处理得到保证语音的连贯性。
所述对信号进行分帧加窗,设置窗长为256样本,对分帧信号进行50%重叠。本发明所述的适当窗长的最优选择为256样本,可以在保证信息不丢失的前提下,提升运算效率。具体而言,GRU模型可以保留的前些时刻信息有限,为了满足实时会议对降噪效果和实时性的要求,本发明还在语音信号的处理上作出了设计。由于一段语音的最小单位可以分割成一个基本音节,普通人讲话的基本音节都在80ms~160ms之间,而快速傅立叶变换基于2^N长度的数据进行可以提高计算速度,故而本发明设置窗长为256样本,虽然分帧成256sample一帧的话,就会丢失前面帧信息,但由于GRU模型可以一定程度上保留前些时刻的信息,这样的窗长设置在保证不丢失信息的前提下,提高的计算速度,而本发明中对分帧信号进行50%重叠,可以避免帧间突变。由于本发明利用了GRU模型可以一定程度上保留前些时刻的信息的特点为语音信号的分帧加窗选择了合适的窗长,所以在输入RNN模型进行估算时,仅需输入当前帧的对数谱,本发明所述的RNN模型对输入信息的要求低,无需对接收到的语音信号做大量的预处理,这也进一步的减少了运算量,加快了响应速度,提高了实时性。
如图6所示,本发明提供了一种基于RNN的实时会议降噪装置,包括采集装置,计算装置和播放装置;所述采集装置采集带噪语音信号发送至计算装置,所述计算装置处理带噪语音信号得到降噪语音信号发送至播放装置;所述计算装置 为采用权利要求1-3任一项所述的一种基于RNN的实时会议降噪方法的计算装置。
与现有技术相比,本发明公开的一种基于RNN的实时会议降噪装置,通过采集装置采集语音信号,通过计算装置进行语音降噪处理,通过播放装置进行播放,其中计算装置采用了本发明所述的一种基于RNN的实时会议降噪方法。由于本发明所述的一种基于RNN的实时会议降噪方法,对语音信号进行分帧加窗处理得到语音信号的对数谱,将对数谱放入RNN模型中确定降噪抑制系数,通过降噪抑制系数到原始信号的对数谱得到降噪后的语音信号,实现了将RNN降噪方法在实时会议中的应用。
具体而言,由于本通常的RNN模型由于无法很好处理远距离依赖而不适用于实时会议中的降噪,而对于可以很好处理远距离依赖的LSTM又存在则结构复杂,设计困难大,运算量大,实时性要求难以满足等问题,而不能使用在实时会议降噪中,而GRU的结构相对简单适合使用在实时会议降噪中,发明在RNN模型选择了采用GRU的RNN模型。本发明使用GRU的模型,通过更新门和重置门,一定程度上保留前些时刻的信息,保证训练模型的可靠性。但与LSTM相比,GRU模型可以保留的前些时刻信息有限,为了满足实时会议对降噪效果和实时性的要求,本发明还在语音信号的处理上作出了设计。由于一段语音的最小单位可以分割成一个基本音节,普通人讲话的基本音节都在80ms~160ms之间,而快速傅立叶变换基于2^N长度的数据进行可以提高计算速度,故而本发明设置窗长为256样本,虽然分帧成256sample一帧的话,就会丢失前面帧信息,但由于GRU模型可以一定程度上保留前些时刻的信息,这样的窗长设置在保证不丢失信息的前提下,提高的计算速度,而本发明中对分帧信号进行50%重叠,可以避免帧间突变。
由于本发明利用了GRU模型可以一定程度上保留前些时刻的信息的特点为语音信号的分帧加窗选择了合适的窗长,所以在输入RNN模型进行估算时,仅需输入当前帧的对数谱,本发明所述的RNN模型对输入信息的要求低,无需对接收到 的语音信号做大量的预处理,这也进一步的减少了运算量,加快了响应速度,提高了实时性。同时,本发明在进行信号对于是采用的是频域上的对数谱,对数谱可以极其明显的突出信号变化,在进行信号对比时采用对数谱可以方面信号对比,本发明采用频域上的对数谱也是为了进一步减少信号对比时间,提高本发明的实时性。最后,在通过降噪抑制系数对原始信号的对数谱进行处理后,对处理结果进行窗重叠和傅里叶逆变换,由于降噪后的信号依然经过窗重叠处理,可以进一步的保证信号的处理效果,避免过分降噪,保证语音完整。
对此,本发明通过对语音信号进行分帧加窗,采取合适的窗长度,采用使用GRU的RNN模型,采用频域上的对数谱进行信号对比,降低了一般RNN降噪方法的运算量,简化的模型结构和信号处理过程,提高了实时性,并且通过对降噪后的信号依然经过窗重叠处理,避免过分降噪,保证语音完整。故而,本发明公开的一种基于RNN的实时会议降噪装置,实现了RNN降噪方法在实时会务系统中的应用,解决了目前没有基于RNN的实时会议降噪装置的问题。
如图6所示,所述采集装置包括麦克风和AD转换器,所述麦克风通过AD转换器与计算装置连接;所述播放装置通过网络与计算装置连接。麦克风采集环境声音,AD转换器将环境中的声音转换成计算需要的数字信号。
如图7所示,所述采集装置还包括远程接收单元,所述远程接收单元与计算装置连接。本发明公开的一种基于RNN的实时会议降噪装置,采集装置用于采集语音信息发送至计算装置,由于该种基于RNN的实时会议降噪装置是应用在实时会务系统对实时会议进行降噪的,目前的实时会议降噪往往是对麦克风采集的语音信号进行降噪处理,降噪后的语音信号通过网络发送至播放装置进行播放,此处的播放装置往往是另一台实时会务系统。具体而言,目前的实时会议往往使用两个实时会务系统相互通信,由一台会务系统向另一台会务系统发送语音信息,目前的实时会议降噪则是在本会务系统接收到语音后即进行降噪,将降噪后的信息发送至另一台会务系统,由于另一台会务系统播放降噪后的语音。但是,本发 明提供的一种基于RNN的实时会议降噪装置,在采集装置中设置远程接收单元,即采集装置可以接受其他会务系统发送的已经经过降噪的语音,而对接收到的语音发送的计算单元再一次进行降噪,而再次降噪后的语音信息有计算装置发送至播放装置进行播放,即本会务系统接受其他会务系统发送的已经降噪的语音信息,对已降噪语音信息进行再次降噪,本会务系统播放再次降噪的语音信息,本发明提供的一种基于RNN的实时会议降噪装置,支持实时会议中的通信双方两次降噪,重复降噪保障了语音信息的降噪效果,最大程度的出去了实时会议中的语音噪音,并且,由于该种基于RNN的实时会议降噪装置,采用的RNN模型结构简单,语音信号处理方法简单,使得该种基于RNN的实时会议降噪装置的结构简单,信号处理简单,数据运算量小,实时性高,通信双方两次降噪同样满足实时会议的实时性要求,提供了一种RNN降噪方法在实时会务系统中应用的装置,保留了RNN降噪方法较好降噪效果的优点,同时满足了实时会议对实时性的要求,解决了目前没有基于RNN的实时会议降噪装置的问题。
所述计算装置为具有计算指令功能的多CPU硬件设备。由于该种基于RNN的实时会议降噪方法可以由任何具有计算指令功能的硬件设备实施。本发明采用多个CPU进行处理,可以加速处理过程,提高运行效率,提高实时性。
以上所述是本发明的优选实施方式,应当指出,对于本技术领域的普通技术人员来说,在不脱离本发明原理的前提下,还可以做出若干改进和润饰,这些改进和润饰也视为本发明的保护范围。
Claims (10)
- 一种基于RNN的实时会议降噪方法,其特征在于,包括以下步骤:S1,对使用GRU的RNN模型进行训练确定RNN模型的合适参数,得到训练完成的RNN模型;S2,对采集单元传输的语音信号进行分帧加窗,计算得到语音信号每帧在频域上的对数谱;S3,将当前帧的对数谱放入训练完成的RNN模型进行估算,得到当前语音的估计对数谱;S4,根据估计对数谱与原始信号的对数谱进行估计,算出当前帧的信噪比,根据信噪比计算出当前的降噪抑制系数;S5,应用降噪抑制系数到原始信号的对数谱,对结果进行窗重叠和傅里叶逆变换,通过网络发送到对应的播放设备上,对处理后的信号进行播放。
- 根据权利要求1所述的一种基于RNN的实时会议降噪方法,其特征在于,所述步骤S1包括以下步骤:S11,采集纯净语音信号和噪音语音信号,对纯净语音信号和噪音语音信号进行时域上的叠加,产生带噪语音信号;S12,对带噪语音信号和纯净语音信号分别进行分帧加窗,计算每帧在频域上的对数谱,将带噪语音信号的对数谱和纯净语音信号的对数谱进行对比得到对应的准确抑制参数;S22,将分帧加窗后得到的带噪语音信号的对数谱作为使用GRU的RNN模型的输入;S23,带噪语音信号的对数谱经过一个全连接层和两个GRU层后生成对应的估计对数谱,根据估计对数谱和带噪语音信号的对数谱得到期望抑制参数;S24,使用期望抑制参数和准确抑制参数计算均方误差,判断均方误差是否小于阈 值,是则结束步骤,不是则利用均方误差进行求导,更新使用GRU的RNN模型的参数并返回步骤S11。
- 根据权利要求2所述的一种基于RNN的实时会议降噪方法,其特征在于,所述全连接层的激活函数采用tanh函数,所述tanh函数的均值为0。
- 根据权利要求1任一项所述的一种基于RNN的实时会议降噪方法,其特征在于,所述降噪抑制系数为期望抑制参数进行平滑处理得到,所述期望抑制参数为计对数谱与原始信号的对数谱进行估计得到。
- 根据权利要求1-4任一项所述的一种基于RNN的实时会议降噪方法,其特征在于,所述对信号进行分帧加窗,设置窗长为256样本,对分帧信号进行50%重叠。
- 一种基于RNN的实时会议降噪装置,其特征在于,包括采集装置,计算装置和播放装置;所述采集装置采集带噪语音信号发送至计算装置,所述计算装置处理带噪语音信号得到降噪语音信号发送至播放装置;所述计算装置为采用权利要求1-4任一项所述的一种基于RNN的实时会议降噪方法的计算装置。
- 一种基于RNN的实时会议降噪装置,其特征在于,包括采集装置,计算装置和播放装置;所述采集装置采集带噪语音信号发送至计算装置,所述计算装置处理带噪语音信号得到降噪语音信号发送至播放装置;所述计算装置为采用权利要求5所述的一种基于RNN的实时会议降噪方法的计算装置。
- 根据权利要求7所述的一种基于RNN的实时会议降噪装置,其特征在于,所述采集装置包括麦克风和AD转换器,所述麦克风通过AD转换器与计算装置连接; 所述播放装置通过网络与计算装置连接。
- 根据权利要求8所述的一种基于RNN的实时会议降噪装置,其特征在于,所述采集装置还包括远程接收单元,所述远程接收单元与计算装置连接。
- 根据权利要求7-9所述的一种基于RNN的实时会议降噪装置,其特征在于,所述计算装置为具有计算指令功能的多CPU硬件设备。
Priority Applications (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| EP18923762.1A EP3633676A4 (en) | 2018-08-09 | 2018-08-22 | RNN-BASED NOISE REDUCTION METHOD AND DEVICE FOR REAL-TIME CONFERENCE |
| US16/628,679 US11024324B2 (en) | 2018-08-09 | 2018-08-22 | Methods and devices for RNN-based noise reduction in real-time conferences |
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN201810904699.2 | 2018-08-09 | ||
| CN201810904699.2A CN109273021B (zh) | 2018-08-09 | 2018-08-09 | 一种基于rnn的实时会议降噪方法及装置 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2020029332A1 true WO2020029332A1 (zh) | 2020-02-13 |
Family
ID=65153280
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2018/101820 Ceased WO2020029332A1 (zh) | 2018-08-09 | 2018-08-22 | 一种基于rnn的实时会议降噪方法及装置 |
Country Status (4)
| Country | Link |
|---|---|
| US (1) | US11024324B2 (zh) |
| EP (1) | EP3633676A4 (zh) |
| CN (1) | CN109273021B (zh) |
| WO (1) | WO2020029332A1 (zh) |
Families Citing this family (22)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN109712628B (zh) * | 2019-03-15 | 2020-06-19 | 哈尔滨理工大学 | 一种基于rnn建立的drnn降噪模型的语音降噪方法及语音识别方法 |
| CN110189260B (zh) * | 2019-04-15 | 2021-01-26 | 浙江大学 | 一种基于多尺度并行门控神经网络的图像降噪方法 |
| CN110010144A (zh) * | 2019-04-24 | 2019-07-12 | 厦门亿联网络技术股份有限公司 | 语音信号增强方法及装置 |
| CN110689901B (zh) * | 2019-09-09 | 2022-06-28 | 苏州臻迪智能科技有限公司 | 语音降噪的方法、装置、电子设备及可读存储介质 |
| CN110648680B (zh) * | 2019-09-23 | 2024-05-14 | 腾讯科技(深圳)有限公司 | 语音数据的处理方法、装置、电子设备及可读存储介质 |
| CN111223493B (zh) * | 2020-01-08 | 2022-08-02 | 北京声加科技有限公司 | 语音信号降噪处理方法、传声器和电子设备 |
| CN111477239B (zh) * | 2020-03-31 | 2023-05-09 | 厦门快商通科技股份有限公司 | 一种基于gru神经网络的去除噪声方法及系统 |
| CN111508519B (zh) * | 2020-04-03 | 2022-04-26 | 北京达佳互联信息技术有限公司 | 一种音频信号人声增强的方法及装置 |
| CN111653285B (zh) * | 2020-06-01 | 2023-06-30 | 北京猿力未来科技有限公司 | 丢包补偿方法及装置 |
| CN111710344B (zh) * | 2020-06-28 | 2025-06-27 | 腾讯科技(深圳)有限公司 | 一种信号处理方法、装置、设备及计算机可读存储介质 |
| CN111739555B (zh) * | 2020-07-23 | 2020-11-24 | 深圳市友杰智新科技有限公司 | 基于端到端深度神经网络的音频信号处理方法及装置 |
| SE545513C2 (en) * | 2021-05-12 | 2023-10-03 | Audiodo Ab Publ | Voice optimization in noisy environments |
| CN113611292B (zh) * | 2021-08-06 | 2023-11-10 | 思必驰科技股份有限公司 | 用于语音分离、识别的短时傅里叶变化的优化方法及系统 |
| GB2612621A (en) * | 2021-11-05 | 2023-05-10 | Iris Audio Tech Limited | Audio processing device and method for suppressing noise |
| CN115527547B (zh) * | 2022-04-29 | 2023-06-16 | 荣耀终端有限公司 | 噪声处理方法及电子设备 |
| CN115171642B (zh) * | 2022-07-07 | 2024-10-29 | 山东省计算中心(国家超级计算济南中心) | 一种基于改进的循环神经网络的主动降噪方法及系统 |
| CN115565542A (zh) * | 2022-09-27 | 2023-01-03 | 睿云联(厦门)网络通讯技术有限公司 | 一种基于纯时域信息的实时语音去噪方法和装置以及设备 |
| CN115762540B (zh) * | 2022-10-27 | 2025-12-02 | 深圳市龙芯威半导体科技有限公司 | 多维度的rnn的语音降噪方法、装置、设备及介质 |
| US12469510B2 (en) * | 2022-11-16 | 2025-11-11 | Cisco Technology, Inc. | Transforming speech signals to attenuate speech of competing individuals and other noise |
| CN116597852B (zh) * | 2022-12-27 | 2026-04-03 | 厦门亿联网络技术股份有限公司 | 一种语音降噪模型训练方法及装置 |
| US20240304186A1 (en) * | 2023-03-08 | 2024-09-12 | Google Llc | Audio signal synthesis from a network of devices |
| CN119360873B (zh) * | 2024-12-26 | 2025-04-22 | 深圳市海威恒泰智能科技有限公司 | 基于ai的会议音频流智能降噪方法、装置、设备及介质 |
Citations (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN103310789A (zh) * | 2013-05-08 | 2013-09-18 | 北京大学深圳研究生院 | 一种基于改进的并行模型组合的声音事件识别方法 |
| US9008329B1 (en) * | 2010-01-26 | 2015-04-14 | Audience, Inc. | Noise reduction using multi-feature cluster tracker |
| CN107452389A (zh) * | 2017-07-20 | 2017-12-08 | 大象声科(深圳)科技有限公司 | 一种通用的单声道实时降噪方法 |
| CN108197327A (zh) * | 2018-02-07 | 2018-06-22 | 腾讯音乐娱乐(深圳)有限公司 | 歌曲推荐方法、装置及存储介质 |
Family Cites Families (9)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| FI100840B (fi) * | 1995-12-12 | 1998-02-27 | Nokia Mobile Phones Ltd | Kohinanvaimennin ja menetelmä taustakohinan vaimentamiseksi kohinaises ta puheesta sekä matkaviestin |
| US6097820A (en) * | 1996-12-23 | 2000-08-01 | Lucent Technologies Inc. | System and method for suppressing noise in digitally represented voice signals |
| JP5870476B2 (ja) * | 2010-08-04 | 2016-03-01 | 富士通株式会社 | 雑音推定装置、雑音推定方法および雑音推定プログラム |
| US8239196B1 (en) * | 2011-07-28 | 2012-08-07 | Google Inc. | System and method for multi-channel multi-feature speech/noise classification for noise suppression |
| US10013975B2 (en) * | 2014-02-27 | 2018-07-03 | Qualcomm Incorporated | Systems and methods for speaker dictionary based speech modeling |
| US9881631B2 (en) * | 2014-10-21 | 2018-01-30 | Mitsubishi Electric Research Laboratories, Inc. | Method for enhancing audio signal using phase information |
| CN104952448A (zh) * | 2015-05-04 | 2015-09-30 | 张爱英 | 一种双向长短时记忆递归神经网络的特征增强方法及系统 |
| WO2017164954A1 (en) * | 2016-03-23 | 2017-09-28 | Google Inc. | Adaptive audio enhancement for multichannel speech recognition |
| CN107886967B (zh) * | 2017-11-18 | 2018-11-13 | 中国人民解放军陆军工程大学 | 一种深度双向门递归神经网络的骨导语音增强方法 |
-
2018
- 2018-08-09 CN CN201810904699.2A patent/CN109273021B/zh active Active
- 2018-08-22 EP EP18923762.1A patent/EP3633676A4/en not_active Ceased
- 2018-08-22 WO PCT/CN2018/101820 patent/WO2020029332A1/zh not_active Ceased
- 2018-08-22 US US16/628,679 patent/US11024324B2/en active Active
Patent Citations (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US9008329B1 (en) * | 2010-01-26 | 2015-04-14 | Audience, Inc. | Noise reduction using multi-feature cluster tracker |
| CN103310789A (zh) * | 2013-05-08 | 2013-09-18 | 北京大学深圳研究生院 | 一种基于改进的并行模型组合的声音事件识别方法 |
| CN107452389A (zh) * | 2017-07-20 | 2017-12-08 | 大象声科(深圳)科技有限公司 | 一种通用的单声道实时降噪方法 |
| CN108197327A (zh) * | 2018-02-07 | 2018-06-22 | 腾讯音乐娱乐(深圳)有限公司 | 歌曲推荐方法、装置及存储介质 |
Non-Patent Citations (2)
| Title |
|---|
| See also references of EP3633676A4 * |
| ZHANG, JUNYANG ET AL.: "Review of Deep Learning", APPLICATION RESEARCH OF COMPUTERS, vol. 35, no. 7, 31 July 2018 (2018-07-31), pages 1921 - 1936, XP009519636 * |
Also Published As
| Publication number | Publication date |
|---|---|
| US11024324B2 (en) | 2021-06-01 |
| CN109273021A (zh) | 2019-01-25 |
| EP3633676A4 (en) | 2020-05-06 |
| US20210035594A1 (en) | 2021-02-04 |
| EP3633676A1 (en) | 2020-04-08 |
| CN109273021B (zh) | 2021-11-30 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| WO2020029332A1 (zh) | 一种基于rnn的实时会议降噪方法及装置 | |
| CN112735456B (zh) | 一种基于dnn-clstm网络的语音增强方法 | |
| JP5608678B2 (ja) | パーティクルフィルタリングを利用した音源位置の推定 | |
| CN111445919B (zh) | 结合ai模型的语音增强方法、系统、电子设备和介质 | |
| CN110070880B (zh) | 用于分类的联合统计模型的建立方法及应用方法 | |
| CN110600050A (zh) | 基于深度神经网络的麦克风阵列语音增强方法及系统 | |
| CN118486318B (zh) | 一种户外直播环境杂音消除方法、介质及系统 | |
| WO2019080551A1 (zh) | 目标语音检测方法及装置 | |
| CN113707136B (zh) | 服务型机器人语音交互的音视频混合语音前端处理方法 | |
| CN111292762A (zh) | 一种基于深度学习的单通道语音分离方法 | |
| CN108109617A (zh) | 一种远距离拾音方法 | |
| CN113284504B (zh) | 姿态检测方法、装置、电子设备及计算机可读存储介质 | |
| CN116013344A (zh) | 一种多种噪声环境下的语音增强方法 | |
| CN115620739A (zh) | 指定方向的语音增强方法及电子设备和存储介质 | |
| CN117746874A (zh) | 一种音频数据处理方法、装置以及可读存储介质 | |
| CN115862632A (zh) | 语音识别方法、装置、电子设备和存储介质 | |
| JP6265903B2 (ja) | 信号雑音減衰 | |
| CN107346658B (zh) | 混响抑制方法及装置 | |
| CN120808810B (zh) | 多模态感知的智能麦克风阵列信号处理方法与系统 | |
| CN115376526A (zh) | 一种基于声纹识别的电力设备故障检测方法及系统 | |
| CN119049506A (zh) | 一种拾音性能优化方法及装置 | |
| Lei et al. | A study on a two-stage UAV noise removal method based on deep residual neural networks | |
| CN107393559B (zh) | 检校语音检测结果的方法及装置 | |
| CN107393558B (zh) | 语音活动检测方法及装置 | |
| Yan et al. | A Frogman Speech Enhancement Algorithm with Pre-Enhanced DNN Joint Model |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| ENP | Entry into the national phase |
Ref document number: 2018923762 Country of ref document: EP Effective date: 20200103 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |