WO2021068120A1 - 一种融合骨振动传感器和麦克风信号的深度学习语音提取和降噪方法 - Google Patents

一种融合骨振动传感器和麦克风信号的深度学习语音提取和降噪方法 Download PDF

Info

Publication number
WO2021068120A1
WO2021068120A1 PCT/CN2019/110080 CN2019110080W WO2021068120A1 WO 2021068120 A1 WO2021068120 A1 WO 2021068120A1 CN 2019110080 W CN2019110080 W CN 2019110080W WO 2021068120 A1 WO2021068120 A1 WO 2021068120A1
Authority
WO
WIPO (PCT)
Prior art keywords
vibration sensor
bone vibration
microphone
noise reduction
neural network
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/CN2019/110080
Other languages
English (en)
French (fr)
Inventor
闫永杰
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Elevoc Technology Co Ltd
Original Assignee
Elevoc Technology Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Elevoc Technology Co Ltd filed Critical Elevoc Technology Co Ltd
Priority to EP19920643.4A priority Critical patent/EP4044181A4/en
Priority to KR1020207028217A priority patent/KR102429152B1/ko
Priority to US17/042,973 priority patent/US20220392475A1/en
Priority to JP2020563485A priority patent/JP2022505997A/ja
Priority to PCT/CN2019/110080 priority patent/WO2021068120A1/zh
Publication of WO2021068120A1 publication Critical patent/WO2021068120A1/zh
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Images

Classifications

    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L21/00Speech or voice signal processing techniques to produce another audible or non-audible signal, e.g. visual or tactile, in order to modify its quality or its intelligibility
    • G10L21/02Speech enhancement, e.g. noise reduction or echo cancellation
    • G10L21/0208Noise filtering
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L21/00Speech or voice signal processing techniques to produce another audible or non-audible signal, e.g. visual or tactile, in order to modify its quality or its intelligibility
    • G10L21/02Speech enhancement, e.g. noise reduction or echo cancellation
    • G10L21/0208Noise filtering
    • G10L21/0216Noise filtering characterised by the method used for estimating noise
    • G10L21/0232Processing in the frequency domain
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L19/00Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis
    • G10L19/04Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis using predictive techniques
    • G10L19/26Pre-filtering or post-filtering
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L21/00Speech or voice signal processing techniques to produce another audible or non-audible signal, e.g. visual or tactile, in order to modify its quality or its intelligibility
    • G10L21/02Speech enhancement, e.g. noise reduction or echo cancellation
    • G10L21/0208Noise filtering
    • G10L21/0216Noise filtering characterised by the method used for estimating noise
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L21/00Speech or voice signal processing techniques to produce another audible or non-audible signal, e.g. visual or tactile, in order to modify its quality or its intelligibility
    • G10L21/02Speech enhancement, e.g. noise reduction or echo cancellation
    • G10L21/038Speech enhancement, e.g. noise reduction or echo cancellation using band spreading techniques
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L25/00Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00
    • G10L25/03Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 characterised by the type of extracted parameters
    • G10L25/18Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 characterised by the type of extracted parameters the extracted parameters being spectral information of each sub-band
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L25/00Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00
    • G10L25/27Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 characterised by the analysis technique
    • G10L25/30Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 characterised by the analysis technique using neural networks
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04RLOUDSPEAKERS, MICROPHONES, GRAMOPHONE PICK-UPS OR LIKE ACOUSTIC ELECTROMECHANICAL TRANSDUCERS; ELECTRIC HEARING AIDS; PUBLIC ADDRESS SYSTEMS
    • H04R1/00Details of transducers, loudspeakers or microphones
    • H04R1/08Mouthpieces; Microphones; Attachments therefor
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04RLOUDSPEAKERS, MICROPHONES, GRAMOPHONE PICK-UPS OR LIKE ACOUSTIC ELECTROMECHANICAL TRANSDUCERS; ELECTRIC HEARING AIDS; PUBLIC ADDRESS SYSTEMS
    • H04R11/00Transducers of moving-armature or moving-core type
    • H04R11/04Microphones
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04RLOUDSPEAKERS, MICROPHONES, GRAMOPHONE PICK-UPS OR LIKE ACOUSTIC ELECTROMECHANICAL TRANSDUCERS; ELECTRIC HEARING AIDS; PUBLIC ADDRESS SYSTEMS
    • H04R3/00Circuits for transducers
    • H04R3/005Circuits for transducers for combining the signals of two or more microphones
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L21/00Speech or voice signal processing techniques to produce another audible or non-audible signal, e.g. visual or tactile, in order to modify its quality or its intelligibility
    • G10L21/02Speech enhancement, e.g. noise reduction or echo cancellation
    • G10L21/0208Noise filtering
    • G10L21/0216Noise filtering characterised by the method used for estimating noise
    • G10L2021/02161Number of inputs available containing the signal or the noise to be suppressed
    • G10L2021/02165Two microphones, one receiving mainly the noise signal and the other one mainly the speech signal
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04RLOUDSPEAKERS, MICROPHONES, GRAMOPHONE PICK-UPS OR LIKE ACOUSTIC ELECTROMECHANICAL TRANSDUCERS; ELECTRIC HEARING AIDS; PUBLIC ADDRESS SYSTEMS
    • H04R2460/00Details of hearing devices, i.e. of ear- or headphones covered by H04R1/10 or H04R5/033 but not provided for in any of their subgroups, or of hearing aids covered by H04R25/00 but not provided for in any of its subgroups
    • H04R2460/13Hearing devices using bone conduction transducers

Definitions

  • the invention relates to the technical field of electronic equipment voice noise reduction, and more specifically, to a deep learning noise reduction method that integrates a bone vibration sensor and a microphone signal.
  • Voice noise reduction technology refers to the separation of speech signals from noisy speech signals. This technology has a wide range of applications. There are usually single-microphone noise reduction technology and multi-microphone noise reduction technology. However, traditional noise reduction technology has some shortcomings. The single-microphone noise reduction technology presupposes that the noise is smooth noise, which is not highly adaptable and has large limitations; while the traditional multi-microphone noise reduction technology requires two or more microphones, which increases the cost, and the multi-microphone structure is for the structural design of the product The requirements are higher, which limits the structural design of the product. Moreover, the multi-microphone noise reduction technology relies on directional information for noise reduction, and cannot suppress the noise from the target human voice direction. The above defects are worthy of improvement.
  • Multi-microphones have higher requirements for product structure design, which restricts product structure design
  • Multi-mic noise reduction technology relies on directional information to reduce noise, and cannot suppress noise coming from the direction of the human voice approaching the target;
  • the single-microphone noise reduction technology relies on noise estimation, and its pre-installed noise is a steady sound, which has limitations.
  • the invention combines the signals of the bone vibration sensor and the traditional microphone, adopts deep learning for fusion to realize noise reduction, and realizes the extraction of target human voice and reduces interference noise under various noise environments.
  • This technology can be applied to earphones, mobile phones, and other call scenarios that fit the ears (or other body parts). Compared with the technology that only uses one or more microphones to reduce noise, combined with the bone vibration sensor can still maintain a good call experience in environments with extremely low signal-to-noise ratios, such as subways and wind noise.
  • this technology does not make any assumptions about noise (traditional single-mic wind noise reduction technology presupposes that the noise is stable noise), and uses the powerful modeling capabilities of deep neural networks to have a good degree of vocal reproduction And extremely strong noise suppression ability, can solve the problem of human voice extraction in complex noise scenes.
  • traditional multi-microphone noise reduction technology that requires 2 or more microphones for beamforming, we use a single microphone.
  • the bone vibration sensor signal sampling is mainly in the low frequency range, but it is not interfered by air conduction noise.
  • this technology uses bone conduction signals as low-frequency input signals, and after high-frequency reconstruction (optional), and The microphone signals are sent to the deep neural network for overall fusion to achieve noise reduction.
  • the present invention introduces the signal of the bone vibration sensor, and uses the characteristic that the bone vibration sensor is not interfered by air noise. It uses deep neural network fusion with the air conduction microphone signal to achieve a high-quality noise reduction effect even at a very low signal-to-noise ratio.
  • the bone vibration sensor signal is used as a sign of voice activity detection.
  • We combine the bone vibration sensor signal with the microphone signal As the input of the deep neural network, organic fusion of the signal layer is carried out to achieve high-quality noise reduction effect.
  • the technical problem to be solved by the present invention is how to use a deep learning noise reduction method that integrates bone vibration sensors and microphone signals to solve the multi-microphone limitation product structure, high cost, and traditional single-microphone noise reduction in the prior art.
  • Technology has limitations and other issues. Unlike other technologies that combine bone vibration sensors and air conduction microphones, which only use bone vibration sensor signals as activation detection signs, this technology uses the characteristics of bone vibration sensor signals that are not interfered by air conduction noise, and uses bone vibration signals as direct input signals. After frequency reconstruction (optional), it is sent to the deep neural network together with the microphone signal for overall fusion and noise reduction. With the help of bone vibration sensors, we can obtain high-quality low-frequency signals, and on this basis, greatly improve the accuracy of deep neural network predictions and make noise reduction effects better.
  • the technical solution adopted by the present invention to solve its technical problems is to construct a deep learning noise reduction method that integrates bone vibration sensors and microphone signals, combines the respective advantages of bone vibration sensors and traditional microphone signals, and adopts deep learning human voice extraction and Noise reduction technology can extract target human voice and reduce interference noise in various noise environments.
  • This technology can be applied to earphones, mobile phones, and other earphones (or other body parts) call scenes, and it is low cost and easy to implement.
  • the deep learning noise reduction method of fusion bone vibration sensor and microphone signal includes the following steps:
  • S1 bone vibration sensor and microphone collect audio signals, and obtain bone vibration sensor audio signals and microphone audio signals respectively;
  • S2 inputs the audio signal of the bone vibration sensor to the high-pass filtering module for high-pass filtering
  • S3 inputs the high-pass filtered audio signal of the bone vibration sensor and the audio signal of the microphone into the deep neural network module;
  • the S4 deep neural network module predicts the noise-reduced speech after fusion.
  • the high-pass filter module corrects the DC offset of the audio signal of the bone vibration sensor and filters out low-frequency clutter signals.
  • the audio signal of the bone vibration sensor is processed by high-pass filtering, and more preferably, the frequency is further widened through high-frequency reconstruction, that is, the method of widening the frequency band. Range, broaden the audio signal of the bone vibration sensor to more than two kilohertz, and then input it into the deep neural network module.
  • the deep neural network module further includes a fusion module, and the fusion module fuses the microphone audio signal and the bone vibration sensor audio signal and reduces noise.
  • one implementation method of the deep neural network module is to implement a convolutional recurrent neural network, and obtain a pure speech amplitude spectrum through prediction.
  • the deep neural network module is composed of several layers of convolutional networks, several layers of long and short-term memory networks, and three corresponding layers of deconvolutional networks. .
  • the training target of the deep neural network module is a pure speech amplitude spectrum.
  • the pure speech is subjected to a short-time Fourier transform, and then the pure speech amplitude spectrum is obtained as the training target, that is, the target amplitude spectrum.
  • the input signal of the deep neural network module is composed of the amplitude spectrum of the audio signal of the bone vibration sensor (or the amplitude spectrum after the frequency band is widened) and the microphone The amplitude spectrum of the audio signal is stacked;
  • the audio signal of the bone vibration sensor and the audio signal of the microphone are respectively subjected to short-time Fourier transform, and then two amplitude spectra are obtained respectively and stacked.
  • the stacked amplitude spectrum is passed through the deep neural network module to obtain the predicted amplitude spectrum and output.
  • the target amplitude spectrum and the predicted amplitude spectrum are taken as the mean square error.
  • the beneficial effect is that the present invention provides a deep learning voice extraction and noise reduction method that integrates bone vibration sensors and microphone signals, and uses the powerful modeling capabilities of deep neural networks to have good people. Sound reproduction and strong noise suppression capabilities can solve the problem of human voice extraction in complex noise scenes.
  • the invention utilizes the characteristic that the bone vibration sensor is not interfered by air conduction noise, and can still maintain a good call experience in an environment with extremely low signal-to-noise ratio, such as subway, wind noise and other scenes. And the use of a single microphone significantly simplifies implementation and reduces costs.
  • this technology uses the characteristics of bone vibration sensor signals that are not interfered by air conduction noise, and uses bone transmission signals as low-frequency input signals. After high-frequency reconstruction (optional), it is sent to the deep neural network together with the microphone signal for overall fusion to obtain the human voice. With the help of bone vibration sensors, we can obtain high-quality low-frequency signals, and on this basis, greatly improve the accuracy of deep neural network prediction of human voice, so that the noise reduction effect is better.
  • Fig. 1 is a flow chart of a deep learning noise reduction method fusion of bone vibration sensor and microphone signal according to the present invention
  • Figure 2 is a block diagram of a method of high-frequency reconstruction
  • FIG. 3 is a block diagram of a deep neural network fusion module structure of a deep learning noise reduction method that integrates bone vibration sensors and microphone signals according to the present invention
  • FIG. 4 is a schematic diagram of the frequency spectrum of the audio signal collected by the bone vibration sensor of the deep learning noise reduction method of the fusion of the bone vibration sensor and the microphone signal of the present invention
  • FIG. 5 is a schematic diagram of the frequency spectrum of the audio signal collected by the microphone of the deep learning noise reduction method that integrates the bone vibration sensor and the microphone signal of the present invention
  • FIG. 6 is a schematic diagram of an audio signal frequency spectrum after processing by a deep learning noise reduction method that integrates a bone vibration sensor and a microphone signal according to the present invention
  • Fig. 7 is a comparison diagram of the noise reduction effect of a method of fusion bone vibration sensor and microphone signal of the present invention and a deep learning real-time noise reduction method corresponding to a single channel without a bone vibration sensor.
  • the present invention is a deep learning voice extraction and noise reduction method fusing bone vibration sensor and microphone signal, including the following steps:
  • S1 bone vibration sensor and microphone collect audio signals, and obtain bone vibration sensor audio signals and microphone audio signals respectively;
  • S2 inputs the audio signal of the bone vibration sensor to the high-pass filter module and performs high-pass filtering
  • S3 inputs the high-pass filtered audio signal of the bone vibration sensor and the audio signal of the microphone into the deep neural network module;
  • the S4 deep neural network module predicts the voice after fusion and noise reduction.
  • the present invention introduces a bone vibration sensor, and utilizes its characteristic of not being interfered by air noise, uses a deep neural network to fuse the bone vibration sensor signal and the air conduction microphone signal, and achieves ideal noise reduction even at a very low signal-to-noise ratio. effect.
  • the most advanced practical speech noise reduction scheme was a feedforward deep neural network (DNN) trained with a large amount of data, although this scheme can separate specific human voices from untrained noisy human voices.
  • this model has a poor noise reduction effect on non-specific human voices.
  • the most effective method is to add the voices of multiple speakers in the training set. However, this will cause the DNN to confuse the speech and background noise, and tend to misclassify the noise as speech.
  • the published application number is 201710594168.3.
  • the patent (named as a universal mono real-time noise reduction method) relates to a universal mono real-time noise reduction method, including the following steps: Receive noisy speech in electronic format, which contains speech And non-human voice interference noise; extract the short-time Fourier amplitude spectrum from the received sound frame by frame as the acoustic feature; use the deep regression neural network with long and short-term memory to generate the ratio film frame by frame; use the generated ratio film to band The amplitude spectrum of the noisy speech is masked; the masked amplitude spectrum and the original phase of the noisy speech are used to synthesize the speech waveform again after inverse Fourier transform.
  • This invention uses a supervised learning method for speech noise reduction, and estimates the ideal ratio film by using a recurrent neural network with long and short-term memory; the recurrent neural network proposed by this invention uses a large number of noisy speech for training, which contains a variety of realities Acoustic scene and microphone impulse response finally realize universal speech noise reduction independent of background noise, speaker and transmission channel.
  • mono noise reduction refers to the processing of signals collected by a single microphone. Compared with the noise reduction method of a microphone array of beamforming, mono noise reduction has a wider range of practicability and low cost.
  • the invention uses a supervised learning method for speech noise reduction, and estimates the ideal ratio film by using a regression neural network with long and short-term memory.
  • This invention introduces the technology to eliminate dependence on future time frames, and realizes the efficient calculation of the regression neural network model in the noise reduction process. Under the premise of not affecting the noise reduction performance, by further simplifying the calculation, a very small Regression neural network model to achieve real-time speech noise reduction.
  • the bone vibration sensor can collect low-frequency speech without being disturbed by air noise.
  • the bone vibration sensor signal and the air conduction microphone signal are fused with a deep neural network to achieve an ideal full-band noise reduction effect even at a very low signal-to-noise ratio.
  • the bone vibration sensor in this embodiment is the prior art.
  • the speech signal has a strong correlation in the time dimension, and this correlation is very helpful for speech separation.
  • a deep neural network-based method splices the current frame and several consecutive frames before and after into a vector with a larger dimension as the input feature.
  • the method is executed by a computer program to extract acoustic features from noisy speech, estimate the ideal time-frequency ratio film, and re-synthesize the noise-reduced speech waveform.
  • the method includes one or more program modules, and any system or hardware device with executable computer programming instructions is used to execute one or more modules above.
  • the high-pass filtering module corrects the DC offset of the audio signal of the bone vibration sensor, and filters out low-frequency clutter signals.
  • the high-pass filter module can be implemented by digital filter filtering.
  • the audio signal of the bone vibration sensor is processed by high-pass filtering, more preferably, it is reconstructed by high frequency. That is, the frequency band widening method is used to further widen the frequency range, and the audio signal of the bone vibration sensor is widened to more than two kilohertz, and then it is input into the deep neural network module.
  • the function of the high-frequency reconstruction module is to further broaden the bandwidth of the bone vibration signal and is an optional module.
  • a deep neural network is currently the most effective method.
  • a structure of a deep neural network is given as an example.
  • High-pass filtering the audio signal of the bone vibration sensor to correct the DC offset of the bone conduction signal to filter out low-frequency noise; broaden the bone vibration signal to more than 2kHz through the method of frequency band widening (high-frequency reconstruction).
  • This step is optional.
  • the original bone vibration signal in step S1 can be used directly; the output of step S2 and the microphone signal are sent to the deep neural network module; the deep neural network module predicts the voice after fusion and noise reduction.
  • Deep neural networks can be used for reconstruction. Deep neural networks can be implemented in many ways. Figure 2 shows one of them (but Not limited to this network), a high-frequency reconstruction method of deep regression neural network based on long and short-term memory.
  • the published application number is 201811199154.2 patent (named a system that recognizes user voice through human body vibration to control electronic equipment) including a human body vibration sensor for sensing the user's human body vibration; a processing circuit coupled with the human body vibration sensor, When it is determined that the output signal of the human body vibration sensor includes the user's voice signal, control the sound pickup device to start sound pickup; a communication module, coupled with the processing circuit and the sound pickup device, is used for the processing circuit and the sound pickup device. Communication between pickup devices.
  • this patent which uses the bone vibration sensor signal as a sign of voice activity detection, we use the bone vibration sensor signal and the microphone signal as the input of the deep neural network to perform deep fusion of the signal layer, so as to achieve an excellent noise reduction effect.
  • the deep neural network module also includes a fusion module.
  • the function of the fusion module based on the deep neural network is to complete the fusion and noise reduction of the microphone audio signal and the bone vibration sensor audio signal.
  • one implementation method of the deep neural network module is to implement a convolutional recurrent neural network, and obtain a pure speech amplitude spectrum (Speech Magnitude Spectrum) through prediction.
  • the network structure in the deep neural network-based fusion module takes the convolutional recurrent neural network as an example, and it can also replace the growth short-term neural network, deep full convolutional network and other structures.
  • the deep neural network module can be composed of a three-layer convolutional network, a three-layer long short-term memory network, and a three-layer deconvolution network.
  • Fig. 3 shows a block diagram of the deep neural network fusion module structure of the deep learning noise reduction method of fusion bone vibration sensor and microphone signal of the present invention, and shows the implementation of the convolutional recurrent neural network of the deep neural network module, that is, deep neural network
  • the training target of the network module is the pure speech amplitude spectrum (Speech Magnitude Spectrum).
  • the pure speech (Clean Speech) is subjected to the short-time Fourier transform (STFT), and then the pure speech amplitude spectrum (Speech Magnitude Spectrum) is obtained.
  • STFT short-time Fourier transform
  • Traffic Magnitude Spectrum As a training target (Training Target), that is, Target Magnitude Spectrum.
  • the input signal of the deep neural network module is formed by stacking the amplitude spectrum of the audio signal of the bone vibration sensor and the amplitude spectrum of the microphone audio signal;
  • the audio signal of the bone vibration sensor and the audio signal of the microphone are respectively subjected to short-time Fourier transform (STFT), and then two amplitude spectra (Magnitude Spectrum) are obtained respectively, and stacked (Stacking).
  • STFT short-time Fourier transform
  • Magnitude Spectrum two amplitude spectra
  • Stacking two amplitude spectra
  • the stacked amplitude spectrum is passed through the deep neural network module to obtain an estimated amplitude spectrum (Estimated Magnitude Spectrum) and output.
  • the target amplitude spectrum and the estimated amplitude spectrum are taken as mean-square error (MSE), and the mean-square error (MSE) is a kind of difference between the estimator and the estimator. measure.
  • the training process uses the back propagation-gradient descent method to update the network parameters, and continuously feeds the network training data and updates the network parameters until the network converges.
  • the inference process uses the combination of the phase of the result of the short-time Fourier transform (STFT) of the microphone data and the predicted amplitude spectrum (Estimated Magnitude Spectrum) to recover the predicted clean speech (Clean Speech).
  • STFT short-time Fourier transform
  • Estimated Magnitude Spectrum predicted amplitude spectrum
  • this patent uses a single microphone as the input. Therefore, it has the characteristics of strong robustness, controllable cost, and low requirements for product structure design.
  • robustness means that the noise reduction performance of the noise reduction system is interfered by microphone consistency, etc.
  • strong robustness means that there is no requirement for microphone consistency and placement, and it can adapt to various microphones.
  • FIG. 7 there is shown a comparison diagram of the noise reduction effect of a deep learning noise reduction method fused with a bone vibration sensor and a microphone signal and a corresponding mono deep learning noise reduction method without a bone vibration sensor. It specifically compared the processing results of the method (Only-Mic) and the method (Sensor-Mic) described in this technology in the "A Universal Mono Real-time Noise Reduction Method" (Application No.: 201710594168.3) in 8 noise scenarios. The objective test results in Figure 7 are obtained.
  • the eight types of noise are: bar noise, highway noise, crossroad noise, railway station noise, car noise at 130km/h, coffee shop noise, table noise, and office noise.
  • the test standard is subjective speech quality assessment (PESQ), and its value range is [-0.5, 4.5]. From the table, we can see that in each scenario, the PESQ score has been greatly improved after processing by this technology, and the average improvement of the eight scenarios is 0.26. This shows that this technology has a higher degree of speech restoration and a stronger noise suppression capability.
  • This method utilizes the characteristics of the bone vibration sensor that is not interfered by air noise, and uses a deep neural network to fuse the bone vibration sensor signal and the air conduction microphone signal to achieve an ideal noise reduction effect even at a very low signal-to-noise ratio.
  • the present invention does not make any assumptions about noise (traditional single-mic wind noise reduction technology generally presupposes that the noise is stationary noise), and uses the powerful modeling ability of the deep neural network, which is very good
  • the human voice reproduction and strong noise suppression ability can solve the problem of human voice extraction in complex noise scenes.
  • This technology can be applied to earphones, mobile phones and other ear (or other body parts) call scenes.
  • this technology uses the characteristic that the bone vibration sensor signal is not interfered by air conduction noise, and uses the bone transmission signal as a low-frequency input signal . After high-frequency reconstruction (optional), it is sent to the deep neural network together with the microphone signal for overall noise reduction and fusion.
  • high-frequency reconstruction (optional)
  • it is sent to the deep neural network together with the microphone signal for overall noise reduction and fusion.
  • we can obtain high-quality low-frequency signals, and on this basis, greatly improve the accuracy of deep neural network predictions and make noise reduction effects better. It is also possible to directly output the result of the bone vibration sensor signal after the frequency band is widened.
  • the function of the high-frequency reconstruction module is to further broaden the bandwidth of the bone vibration signal, and is an optional module.
  • There are many methods for high-frequency reconstruction. Deep neural network is one of the most effective recent methods. In the specific embodiment, only a structure of deep neural network is given as an example.
  • the network structure of the deep neural network-based fusion module uses a convolutional recurrent neural network as an example, and it can also replace structures such as a long-term neural network and a deep full convolutional network.
  • the present invention provides a deep learning voice extraction and noise reduction method that integrates bone vibration sensors and microphone signals, combines the respective advantages of bone vibration sensors and traditional microphone signals, and uses the powerful modeling capabilities of deep neural networks to achieve high levels of people
  • the sound reproduction and strong noise suppression ability can solve the problem of human voice extraction in complex noise scenes, realize the extraction of target human voice, reduce interference noise, and adopt a single microphone structure to reduce the complexity and cost of implementation.

Landscapes

  • Engineering & Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • Signal Processing (AREA)
  • Acoustics & Sound (AREA)
  • Health & Medical Sciences (AREA)
  • Audiology, Speech & Language Pathology (AREA)
  • Computational Linguistics (AREA)
  • Human Computer Interaction (AREA)
  • Multimedia (AREA)
  • Quality & Reliability (AREA)
  • General Health & Medical Sciences (AREA)
  • Otolaryngology (AREA)
  • Spectroscopy & Molecular Physics (AREA)
  • Artificial Intelligence (AREA)
  • Evolutionary Computation (AREA)
  • Electromagnetism (AREA)
  • Circuit For Audible Band Transducer (AREA)
  • Details Of Audible-Bandwidth Transducers (AREA)

Abstract

一种融合骨振动传感器和麦克风信号的深度学习降噪方法,包括如下步骤:S1骨振动传感器和麦克风采集音频信号,分别得到骨振动传感器音频信号和麦克风音频信号;S2将骨振动传感器音频信号输入高通滤波模块,并进行高通滤波;S3将经过高通滤波后的骨振动传感器音频信号或经过频带拓宽后的信号,与麦克风音频信号输入深度神经网络模块;S4深度神经网络模块经过预测得出融合降噪后的语音。该方法结合了骨震动传感器以及传统麦克风的信号,利用深度神经网络强大的建模能力实现了很高的人声还原度及极强的噪声抑制能力,可以解决复杂噪声场景下的人声提取问题,实现提取目标人声,降低干扰噪声,并可采用单麦克风结构减少成本。还可将骨振动传感器音频信号经过频带拓宽后的信号直接作为输出。

Description

一种融合骨振动传感器和麦克风信号的深度学习语音提取和降噪方法 技术领域
本发明涉及电子设备语音降噪技术领域,更具体地说,涉及一种融合骨振动传感器和麦克风信号的深度学习降噪方法。
背景技术
语音降噪技术是指从带噪语音信号中分离出语音信号,该技术拥有广泛的应用,通常有单麦克风降噪技术和多麦克风降噪技术,然而传统的降噪技术中存在一些缺陷,传统的单麦克风降噪技术预先假设噪声为平稳噪声,适应性不高,局限较大;而传统的多麦克风降噪技术需要两个及以上的麦克风,增加了成本,多麦克风结构对于产品的结构设计要求更高,限制了产品的结构设计,而且,多麦克风降噪技术依靠方向信息进行降噪,无法抑制来自目标人声方向的噪音,以上缺陷值得改进。
传统多麦克风和单麦克风通话降噪技术存在以下缺陷:
1.麦克风数量与成本呈线性关系,麦克数量越多,成本越高;
2.多麦克风对产品结构设计要求更高,限制产品的结构设计;
3.多麦克降噪技术依靠方向信息进行降噪,无法抑制来自于接近目标人声方向的噪音;
4.单麦克风降噪技术依赖噪声估计,其预先架设噪声为平稳声,具有局限性。
本发明结合了骨震动传感器及传统麦克风的信号,采用深度学习进行融合从而实现降噪,在各种噪声环境下,实现提取目标人声,降低干扰噪声。该技术可应用于耳机、手机等贴合耳部(或其它身体部位)的通话场景。相比于仅采用一个或多个麦克风降噪的技术,结合骨振动传感器可在信噪比极低的环境下,诸如:地铁、风噪等场景,依然可以保持良好的通话体验。相比传统单麦克风降噪技术,本技术不对噪声做任何假设(传统单麦风降噪技术预先假设噪声为平稳噪声),利用深度神经网络强大的建模能力,有很好的人声还原度及极强的噪声抑制能力,可以在解决复杂噪声场景下人声提取问题。相比于传统多麦克风降噪技术需要2个及以上麦克风进行波束形成的降噪方案,我们采用单麦克风。
相对于气导麦克风,骨振动传感器信号采样主要在低频范围,但不受气导噪声干扰。不同于其他结合骨震动传感器及气导麦克风降噪方式仅利用骨震动传感器信号作为人声激活检测的标志,本技术将骨传导信号作为低频输入信号,通过高频重建(可选)后,与麦克风信号一同送入深度神经网络进行整体融合后实现降噪。借助骨振动传感器,我们能够得到优质的低频信号,并以此为基础,极大地提高深度神经网络预测的准确性,使得降噪效果更佳。
相比申请号为201710594168.3的专利(名称为一种通用的单声道实时降噪方法),本发明引入了骨振动传感器信号,利用骨振动传感器不受空气噪音 干扰的特性,将骨振动传感器信号与气导麦克风信号使用深度神经网络融合,达到了在极低信噪比下也能有优质的降噪效果。
相比申请号为201811199154.2的专利(名称为一种通过人体振动识别用户语音以控制电子设备的系统)中将骨振动传感器信号作为语音活动检测的标志不同,我们将骨振动传感器信号与麦克风信号一起作为深度神经网络的输入,进行信号层的有机融合,从而达到优质的降噪效果。
发明内容
本发明要解决的技术问题在于如何通过采用一种融合骨振动传感器和麦克风信号的深度学习降噪方法,以解决现有技术中多麦克风限制产品结构、成本过高、而且传统的单麦克风降噪技术有局限性等问题。不同于其他结合骨震动传感器和气导麦克风技术中仅利用骨震动传感器信号作为激活检测的标志,本技术利用骨振动传感器信号不受气导噪声干扰的特性,将骨传信号作为直接输入信号,通过高频重建(可选)后,与麦克风信号一同送入深度神经网络进行整体融合及降噪。借助骨振动传感器,我们能够得到优质的低频信号,并以此为基础,极大地提高深度神经网络预测的准确性,使得降噪效果更佳。
本发明解决其技术问题所采用的技术方案是:构造一种融合骨振动传感器和麦克风信号的深度学习降噪方法,结合了骨震动传感器及传统麦克风的信号各自优势,采用深度学习人声提取及降噪技术,在各种噪声环境下,实现提取 目标人声,降低干扰噪声。该技术可应用于耳机、手机等贴合耳部(或其它身体部位)的通话场景,且成本低易实现。
在本发明所述的一种融合骨振动传感器和麦克风信号的深度学习降噪方法中,所述一种融合骨振动传感器和麦克风信号的深度学习降噪方法,包括如下步骤:
S1骨振动传感器和麦克风采集音频信号,分别得到骨振动传感器音频信号和麦克风音频信号;
S2将骨振动传感器音频信号输入高通滤波模块,进行高通滤波;
S3将经过高通滤波后的骨振动传感器音频信号与麦克风音频信号输入深度神经网络模块;
S4深度神经网络模块经过融合后预测得出降噪语音。
在本发明的一种融合骨振动传感器和麦克风信号的深度学习降噪方法中,高通滤波模块修正骨振动传感器音频信号直流偏移,并滤除低频杂波信号。
在本发明的一种融合骨振动传感器和麦克风信号的深度学习降噪方法中,骨振动传感器音频信号经过高通滤波处理后,更优选的,通过高频重建,即频带拓宽的方法,进一步拓宽频率范围,将骨振动传感器音频信号拓宽至两千赫兹以上,随后将其输入深度神经网络模块。
进一步,亦可仅使用频带拓宽后的骨振动信号作为最终的输出信号,从而无需依赖麦克风信号。
在本发明的一种融合骨振动传感器和麦克风信号的深度学习降噪方法中,深度神经网络模块还包括融合模块,融合模块将麦克风音频信号和骨振动传感器音频信号融合及降噪。
在本发明的一种融合骨振动传感器和麦克风信号的深度学习降噪方法中,深度神经网络模块的一种实现方法是通过卷积循环神经网络实现,并通过预测得到纯净语音幅度谱。
在本发明的一种融合骨振动传感器和麦克风信号的深度学习降噪方法中,深度神经网络模块由数层卷积网络、数层长短期记忆网络和三相对应的数层反卷积网络构成。
在本发明的一种融合骨振动传感器和麦克风信号的深度学习降噪方法中,深度神经网络模块的训练目标是纯净语音幅度谱。首先将纯净语音经过短时傅里叶变换后,再获得纯净语音幅度谱作为训练目标,即目标幅度谱。
在本发明的一种融合骨振动传感器和麦克风信号的深度学习降噪方法中,深度神经网络模块的输入信号是由骨振动传感器音频信号的幅度谱(或经过频带拓宽后的幅度谱)和麦克风音频信号的幅度谱堆叠而成;
首先将骨振动传感器音频信号和麦克风音频信号分别经过短时傅里叶变换,再分别取得两路幅度谱,并进行堆叠。
在本发明的一种融合骨振动传感器和麦克风信号的深度学习降噪方法中,将堆叠后的幅度谱经过深度神经网络模块,得到预测幅度谱,并输出。
在本发明的一种融合骨振动传感器和麦克风信号的深度学习降噪方法中,将目标幅度谱与预测幅度谱做均方误差。
根据上述方案的本发明,其有益效果在于,本发明提供了一种融合骨振动传感器和麦克风信号的深度学习语音提取和降噪方法,利用深度神经网络强大的建模能力,有很好的人声还原度及极强的噪声抑制能力,可以解决复杂噪声场景下的人声提取问题。本发明利用骨振动传感器不受气导噪声干扰的特性,可在信噪比极低的环境下,诸如:地铁、风噪等场景,依然保持良好的通话体验。且采用单麦克风显著地简化实现和降低成本。不同于其他结合骨震动传感器和气导麦克风降噪方式中仅利用骨震动传感器信号作为激活检测的标志,本技术利用骨振动传感器信号不受气导噪声干扰的特性,将骨传信号作为低频输入信号,通过高频重建(可选)后,与麦克风信号一同送入深度神经网络进行整体融合而获取人声。借助骨振动传感器,我们能够得到优质的低频信号,并以此为基础,极大地提高深度神经网络预测人声的准确性,使得降噪效果更佳。
附图说明
下面将结合附图及实施例对本发明作进一步说明。附图中:
图1是本发明的一种融合骨振动传感器和麦克风信号的深度学习降噪方法的流程框图;
图2是高频重建的一种方法原理框图;
图3是本发明的一种融合骨振动传感器和麦克风信号的深度学习降噪方法的深度神经网络融合模块结构框图;
图4是本发明的一种融合骨振动传感器和麦克风信号的深度学习降噪方法的骨震动传感器采集到的音频信号频谱图示意;
图5是本发明的一种融合骨振动传感器和麦克风信号的深度学习降噪方法的麦克风采集到的音频信号频谱图示意;
图6是本发明的一种融合骨振动传感器和麦克风信号的深度学习降噪方法处理后的音频信号频谱图示意;
图7是本发明的一种融合骨振动传感器和麦克风信号的降噪方法和一种无骨震动传感器的单声道对应的深度学习实时降噪方法的降噪效果对比图。
具体实施方式
为了使本发明的目的、技术方案及优点更加清楚明白,以下结合附图及实施例,对本发明进行进一步详细说明。应当理解,此处所描述的具体实施例仅用以解释本发明,并不用于限定本发明。
如图1所示,本发明是一种融合骨振动传感器和麦克风信号的深度学习语音提取和降噪方法,包括如下步骤:
S1骨振动传感器和麦克风采集音频信号,分别得到骨振动传感器音频信号和麦克风音频信号;
S2将骨振动传感器音频信号输入高通滤波模块,并进行高通滤波;
S3将经过高通滤波后的骨振动传感器音频信号与麦克风音频信号输入深度神经网络模块;
S4深度神经网络模块经过预测得出融合降噪后的语音。本发明引入了骨振动传感器,利用其不受空气噪音干扰的特性,将骨振动传感器信号与气导麦克风信号使用深度神经网络融合,达到了在极低信噪比下也能有理想的降噪效果。
之前最先进的实用语音降噪方案是使用大量数据训练的前馈型深度神经网络(Deep neural network,DNN),尽管该方案可以实现从未经训练的带噪人声中分离出特定人声,但该模型对非特定人声的降噪效果并不好。为了提升非特定人声的降噪效果,最有效的方法是在训练集中加入多个说话人的语音,然而这样会使得DNN对语音和背景噪声出现混淆,并且倾向于将噪声错分为语音。
公开的申请号为201710594168.3专利(名称为一种通用的单声道实时降噪方法)涉及一种通用的单声道实时降噪方法,包括以下步骤:接收电子格式的带噪语音,其中包含语音和非人声干扰噪声;从接收到的声音中逐帧提取短时傅里叶幅度谱作为声学特征;使用具有长短期记忆的深度回归神经网络逐帧产生比值膜;利用产生的比值膜对带噪语音的幅度谱进行掩蔽;使用掩蔽后的幅度谱和带噪语音的原始相位,经过逆傅里叶变换,再次合成语音波形。该发明采用有监督学习方法进行语音降噪,通过使用带有长短期记忆的回归神经网络 来估计理想比值膜;该发明提出的回归神经网络使用大量带噪语音进行训练,其中包含了各种现实声学场景和麦克风脉冲响应,最终实现了独立于背景噪声、说话人和传输信道的通用语音降噪。其中,单声道降噪是指对单个麦克风采集的信号进行处理,相比波束形成的麦克风阵列降噪方法,单声道降噪具有更广泛的实用性及低成本。该发明采用有监督学习方法进行语音降噪,通过使用带有长短期记忆的回归神经网络来估计理想比值膜。该发明引入了消除对未来时间帧依赖的技术,并实现了降噪过程中回归神经网络模型的高效计算,在不影响降噪性能的前提下,通过进一步的简化计算,构造了一个非常小的回归神经网络模型,从而实现了实时语音降噪。
进一步地,引入了骨振动传感器。骨振动传感器能采集低频语音、不受空气噪音干扰。将骨振动传感器信号与气导麦克风信号使用深度神经网络融合,达到了在极低信噪比下也能有理想的全频段降噪效果。本实施例中的骨振动传感器为现有技术。
语音信号在时间维度上具有较强的相关性,而且这种相关性对语音分离有很大帮助。为了利用这一上下文信息提高分离性能,基于深度神经网络的方法将当前帧和前后连续几帧拼接成一个维度较大的向量作为输入特征。该方法由计算机程序执行,从带噪语音中提取声学特征,估计理想时频比值膜,并重新合成降噪后的语音波形。该方法包含一个或多个程序模块,任何系统或带有可执行计算机编程指令的硬件设备用来执行上述的一个或多个模块。
进一步地,高通滤波模块修正骨振动传感器音频信号直流偏移,并滤除低频杂波信号。
更进一步地,高通滤波模块可通过数字滤波器滤波实现。
进一步地,骨振动传感器音频信号经过高通滤波处理后,更优选的,通过高频重建。即利用频带拓宽方法进一步拓宽频率范围,将骨振动传感器音频信号拓宽至两千赫兹以上,随后将其输入深度神经网络模块。
进一步地,高频重建模块的作用是进一步拓宽骨振动信号的带宽,是可选模块。
更进一步地,高频重建的方法有很多,深度神经网络是目前最有效的方法,本实施例中仅示例给出了一种深度神经网络的结构作为示例。
将骨振动传感器音频信号进行高通滤波,修正骨传导信号直流偏移,滤除低频噪音;通过频带拓宽(高频重建)的方法,将骨振动信号拓宽至2kHz以上,此步骤可选,此步可直接使用步骤S1中原始的骨振动信号;将步骤S2的输出与麦克风的信号送入深度神经网络模块;深度神经网络模块预测出融合降噪后的语音。
如图2所示,高频重建的作用是进一步拓宽骨振动信号的频率范围,可以采用深度神经网络进行重建,其中深度神经网络可以有多种实现方式,图2给出了其中一种(但不限于该网络),基于长短期记忆的深度回归神经网络的高频重建方式。
公开的申请号为201811199154.2专利(名称为一种通过人体振动识别用户语音以控制电子设备的系统)包括人体振动传感器,用于感应用户的人体振动;处理电路,与所述人体振动传感器相耦合,用于当确定所述人体振动传感器的输出信号包括用户语音信号时,控制拾音设备开始拾音;通信模块,与处理电路和所述拾音设备相耦合,用于所述处理电路和所述拾音设备之间的通信。与该专利将骨振动传感器信号作为语音活动检测的标志不同,我们将骨振动传感器信号与麦克风信号一起作为深度神经网络的输入,进行信号层的深度融合,从而达到优良的降噪效果。
进一步地,深度神经网络模块还包括融合模块,基于深度神经网络的融合模块作用是完成麦克风音频信号和骨振动传感器音频信号融合及降噪。
进一步地,深度神经网络模块的一种实现方法是通过卷积循环神经网络实现,并通过预测得到纯净语音幅度谱(Speech Magnitude Spectrum)。
更进一步地,基于深度神经网络的融合模块中网络结构以卷积循环神经网络作为示例,也可替换成长短期神经网络、深度全卷积网络等结构。
作为示例,深度神经网络模块可由三层卷积网络、三层长短期记忆网络和三层反卷积网络构成。
图3示出了本发明的一种融合骨振动传感器和麦克风信号的深度学习降噪方法的深度神经网络融合模块结构框图,给出了深度神经网络模块的卷积循环神经网络实现,即深度神经网络模块的训练目标(Training Target)是纯净 语音幅度谱(Speech Magnitude Spectrum),首先将纯净语音(Clean Speech)经过短时傅里叶变换(STFT)后,再获得纯净语音幅度谱(Speech Magnitude Spectrum)作为训练目标(Training Target),即目标幅度谱(Target Magnitude Spectrum)。
进一步地,深度神经网络模块的输入信号是由骨振动传感器音频信号的幅度谱和麦克风音频信号的幅度谱堆叠(Stacking)而成;
首先将骨振动传感器音频信号和麦克风音频信号分别经过短时傅里叶变换(STFT),再分别取得两路幅度谱(Magnitude Spectrum),并进行堆叠(Stacking)。
进一步地,将堆叠(Stacking)后的幅度谱经过深度神经网络模块,得到预测幅度谱(Estimated Magnitude Spectrum),并输出。
进一步地,将目标幅度谱与预测幅度谱(Estimated Magnitude Spectrum)做均方误差(mean-square error,MSE),均方误差(MSE)是反映估计量与被估计量之间差异程度的一种度量。更进一步地,训练过程(Training)采用反向传播-梯度下降的方式更新网络参数,不断地送入网络训练数据、更新网络参数,直至网络收敛。
进一步地,推理过程(Inference)使用麦克风数据短时傅里叶变换(STFT)后结果的相位和预测的幅度谱(Estimated Magnitude Spectrum)结合,恢复出预测后的纯净语音(Clean Speech)。
相对传统多麦降噪技术,本专利采用单麦克风作为输入。因此具有鲁棒性强、成本可控、对产品结构设计要求低等特点。在本实施例中,鲁棒性是指降噪系统的降噪性能受麦克风一致性等干扰,鲁棒性强指的是对麦克风一致性及放置等没有要求,能适应各种麦克风。
如图7所示,示出了一种融合骨振动传感器和麦克风信号的深度学习降噪方法和相对应一种无骨震动传感器的单声道深度学习降噪方法的降噪效果对比图。具体对比了8种噪音场景下分别使用《一种通用的单声道实时降噪方法》(申请号:201710594168.3)中方法(Only-Mic)与本技术所述方法(Sensor-Mic)处理结果,得出了图7中的客观测试结果。八种噪声分别为:酒吧噪声、公路噪声、十字路口噪声、火车站噪声、130km/h速度行驶的汽车噪声、咖啡厅噪声、餐桌上的噪声以及办公室噪声。测试标准为主观语音质量评估(PESQ),其值范围为[-0.5,4.5]。从表中我们可以看到,在各场景下,经本技术处理后PESQ得分都有很大提升,八个场景平均提升在0.26。这表明本技术对语音还原度更高、噪声抑制能力更强。本方法利用骨振动传感器不受空气噪音干扰的特性,将骨振动传感器信号与气导麦克风信号使用深度神经网络融合,达到了在极低信噪比下也能有理想的降噪效果。
更进一步地,相比传统单麦克风降噪技术,本发明不对噪声做任何假设(传统单麦风降噪技术一般预先假设噪声为平稳噪声),利用深度神经网络强大的建模能力,有很好的人声还原度及极强的噪声抑制能力,可以解决复杂噪声场 景下的人声提取问题,该技术可应用于耳机、手机等贴合耳部(或其它身体部位)的通话场景。不同于其他结合骨震动传感器及气导麦克风降噪方式中仅利用骨震动传感器信号作为激活检测的标志,本技术利用骨振动传感器信号不受气导噪声干扰的特性,将骨传信号作为低频输入信号,通过高频重建(可选)后,与麦克风信号一同送入深度神经网络进行整体降噪、融合。借助骨振动传感器,我们能够得到优质的低频信号,并以此为基础,极大地提高深度神经网络预测的准确性,使得降噪效果更佳。亦可单独将骨振动传感器信号经过频带拓宽后的结果直接作为输出。
在本实施例中,高频重建模块的作用是进一步拓宽骨振动信号的带宽,是一种可选模块。高频重建的方法有很多,深度神经网络是一种效果最优秀的近期方法,具体实施例中仅示例给出了一种深度神经网络的结构作为示例。实施例中基于深度神经网络的融合模块中网络结构以卷积循环神经网络作为示例,也可替换成长短期神经网络、深度全卷积网络等结构。
本发明提供一种融合骨振动传感器和麦克风信号的深度学习语音提取和降噪方法,结合了骨震动传感器及传统麦克风信号的各自优势,利用深度神经网络强大的建模能力实现了很高的人声还原度及极强的噪声抑制能力,可以解决复杂噪声场景下的人声提取问题,实现提取目标人声,降低干扰噪声,并采用单麦克风结构,减少了实现复杂度及成本。
尽管通过以上实施例对本发明进行了揭示,但本发明的保护范围并不局限 于此,在不偏离本发明构思的条件下,对以上各构件所做的变形、替换等均将落入本发明的权利要求范围内。

Claims (10)

  1. 一种融合骨振动传感器和麦克风信号的深度学习降噪方法,其特征在于,包括如下步骤:
    S1骨振动传感器和麦克风采集音频信号,分别得到骨振动传感器音频信号和麦克风音频信号;
    S2将所述骨振动传感器音频信号输入高通滤波模块,并进行高通滤波;
    S3将经过高通滤波后的所述骨振动传感器音频信号与所述麦克风音频信号输入深度神经网络模块;
    S4所述深度神经网络模块经过预测得出融合降噪后的语音。
  2. 根据权利要求1所述的一种融合骨振动传感器和麦克风信号的深度学习降噪方法,其特征在于,所述高通滤波模块修正所述骨振动传感器音频信号直流偏移,并滤除低频杂波信号。
  3. 根据权利要求2所述的一种融合骨振动传感器和麦克风信号的深度学习降噪方法,其特征在于,所述骨振动传感器音频信号经过高通滤波处理后,更优选的,通过高频重建,即频带拓宽的方法,进一步拓宽频率范围,将所述骨振动传感器音频信号拓宽至两千赫兹以上,随后将其输入所述深度神经网络模块。
  4. 根据权利要求3所述的将骨振动传感器信号经过高频重建(频带拓宽)后的结果亦可直接作为本发明输出。5、根据权利要求1所述的一种融合骨振 动传感器和麦克风信号的深度学习降噪方法,其特征在于,所述深度神经网络模块还包括融合模块,所述融合模块将所述麦克风音频信号和所述骨振动传感器音频信号融合及降噪。
  5. 根据权利要求5所述的一种融合骨振动传感器和麦克风信号的深度学习降噪方法,其特征在于,所述深度神经网络模块的一种实现方法是通过卷积循环神经网络实现,并通过预测得到纯净语音幅度谱。
  6. 根据权利要求1所述的一种融合骨振动传感器和麦克风信号的深度学习降噪方法,其特征在于,所述深度神经网络模块由数层卷积网络、数层长短期记忆网络和相对应的数层反卷积网络构成。
  7. 根据权利要求6所述的一种融合骨振动传感器和麦克风信号的深度学习降噪方法,其特征在于,所述深度神经网络模块的训练目标是所述纯净语音幅度谱,首先将所述纯净语音经过短时傅里叶变换后,再获得所述纯净语音幅度谱作为训练目标,即目标幅度谱。
  8. 根据权利要求6所述的一种融合骨振动传感器和麦克风信号的深度学习降噪方法,其特征在于,所述深度神经网络模块的输入信号是由所述骨振动传感器音频信号的幅度谱和所述麦克风音频信号的幅度谱堆叠而成;
    首先将所述骨振动传感器音频信号和所述麦克风音频信号分别经过短时傅里叶变换,再分别取得两路幅度谱,并进行堆叠。
  9. 根据权利要求9所述的一种融合骨振动传感器和麦克风信号的深度学 习降噪方法,其特征在于,将堆叠后的幅度谱经过所述深度神经网络模块,得到预测幅度谱,并输出。
  10. 根据权利要求8或10所述的一种融合骨振动传感器和麦克风信号的深度学习降噪方法,其特征在于,将所述目标幅度谱与所述预测幅度谱做均方误差。
PCT/CN2019/110080 2019-10-09 2019-10-09 一种融合骨振动传感器和麦克风信号的深度学习语音提取和降噪方法 Ceased WO2021068120A1 (zh)

Priority Applications (5)

Application Number Priority Date Filing Date Title
EP19920643.4A EP4044181A4 (en) 2019-10-09 2019-10-09 Deep learning speech extraction and noise reduction method fusing signals of bone vibration sensor and microphone
KR1020207028217A KR102429152B1 (ko) 2019-10-09 2019-10-09 골진동 센서 및 마이크로폰 신호를 융합한 딥 러닝 음성 추출 및 노이즈 저감 방법
US17/042,973 US20220392475A1 (en) 2019-10-09 2019-10-09 Deep learning based noise reduction method using both bone-conduction sensor and microphone signals
JP2020563485A JP2022505997A (ja) 2019-10-09 2019-10-09 骨振動センサーとマイクの信号を融合するディープラーニング音声抽出及びノイズ低減方法
PCT/CN2019/110080 WO2021068120A1 (zh) 2019-10-09 2019-10-09 一种融合骨振动传感器和麦克风信号的深度学习语音提取和降噪方法

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
PCT/CN2019/110080 WO2021068120A1 (zh) 2019-10-09 2019-10-09 一种融合骨振动传感器和麦克风信号的深度学习语音提取和降噪方法

Publications (1)

Publication Number Publication Date
WO2021068120A1 true WO2021068120A1 (zh) 2021-04-15

Family

ID=75436918

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2019/110080 Ceased WO2021068120A1 (zh) 2019-10-09 2019-10-09 一种融合骨振动传感器和麦克风信号的深度学习语音提取和降噪方法

Country Status (5)

Country Link
US (1) US20220392475A1 (zh)
EP (1) EP4044181A4 (zh)
JP (1) JP2022505997A (zh)
KR (1) KR102429152B1 (zh)
WO (1) WO2021068120A1 (zh)

Cited By (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN115171713A (zh) * 2022-06-30 2022-10-11 歌尔科技有限公司 语音降噪方法、装置、设备及计算机可读存储介质
WO2023056280A1 (en) * 2021-09-30 2023-04-06 Sonos, Inc. Noise reduction using synthetic audio
WO2024002896A1 (en) * 2022-06-29 2024-01-04 Analog Devices International Unlimited Company Audio signal processing method and system for enhancing a bone-conducted audio signal using a machine learning model
JP2024528596A (ja) * 2021-07-15 2024-07-30 ドルビー ラボラトリーズ ライセンシング コーポレイション 発話向上
US12567428B1 (en) * 2021-10-11 2026-03-03 Meta Platforms Technologies, Llc Contact transducer based audio enhancement

Families Citing this family (10)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP7814723B2 (ja) * 2021-08-26 2026-02-17 国立大学法人九州工業大学 個人認証方法、個人認証装置及び個人認証用プログラム
KR102790372B1 (ko) * 2022-07-22 2025-04-01 재단법인대구경북과학기술원 신경망 모델에 기반하여 고주파 생체 신호를 복원하는 방법 및 장치
JP2024044550A (ja) 2022-09-21 2024-04-02 株式会社メタキューブ デジタルフィルタ回路、方法、および、プログラム
CN116030823B (zh) * 2023-03-30 2023-06-16 北京探境科技有限公司 一种语音信号处理方法、装置、计算机设备及存储介质
WO2024232876A1 (en) * 2023-05-09 2024-11-14 Google Llc Machine learning based robust voice communication via head-worn device
CN119339734A (zh) * 2023-07-21 2025-01-21 北京三星通信技术研究有限公司 由电子设备执行的方法、电子设备及存储介质
CN116687379B (zh) * 2023-07-28 2026-02-17 南京理工大学 一种基于加速度计的生理信息处理方法及系统
KR102922169B1 (ko) * 2023-09-12 2026-02-03 주식회사 인투스 고품질 음성 획득 장치 및 그 방법
WO2025165461A1 (en) * 2024-01-31 2025-08-07 Qualcomm Incorporated Generative speech restoration using vibration sensor data
CN118465305B (zh) * 2024-07-10 2024-11-05 南京大学 基于监控相机音频数据的风速测量深度学习方法及系统

Citations (9)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN102027536A (zh) * 2008-05-14 2011-04-20 索尼爱立信移动通讯有限公司 响应于说话时在用户面部中感测到的振动对麦克风信号进行自适应滤波
CN102761643A (zh) * 2011-04-26 2012-10-31 鹦鹉股份有限公司 组合话筒和耳机的音频头戴式耳机
US20130070935A1 (en) * 2011-09-19 2013-03-21 Bitwave Pte Ltd Multi-sensor signal optimization for speech communication
CN103229238A (zh) * 2010-11-24 2013-07-31 皇家飞利浦电子股份有限公司 用于产生音频信号的系统和方法
CN107452389A (zh) 2017-07-20 2017-12-08 大象声科(深圳)科技有限公司 一种通用的单声道实时降噪方法
CN108231086A (zh) * 2017-12-24 2018-06-29 航天恒星科技有限公司 一种基于fpga的深度学习语音增强器及方法
CN108986834A (zh) * 2018-08-22 2018-12-11 中国人民解放军陆军工程大学 基于编解码器架构与递归神经网络的骨导语音盲增强方法
CN109346075A (zh) 2018-10-15 2019-02-15 华为技术有限公司 通过人体振动识别用户语音以控制电子设备的方法和系统
CN109767783A (zh) * 2019-02-15 2019-05-17 深圳市汇顶科技股份有限公司 语音增强方法、装置、设备及存储介质

Family Cites Families (7)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPH08223677A (ja) * 1995-02-15 1996-08-30 Nippon Telegr & Teleph Corp <Ntt> 送話器
JP2003264883A (ja) * 2002-03-08 2003-09-19 Denso Corp 音声処理装置および音声処理方法
JP2008042740A (ja) * 2006-08-09 2008-02-21 Nara Institute Of Science & Technology 非可聴つぶやき音声採取用マイクロホン
US9418675B2 (en) * 2010-10-04 2016-08-16 LI Creative Technologies, Inc. Wearable communication system with noise cancellation
US10090001B2 (en) * 2016-08-01 2018-10-02 Apple Inc. System and method for performing speech enhancement using a neural network-based combined symbol
US10535364B1 (en) * 2016-09-08 2020-01-14 Amazon Technologies, Inc. Voice activity detection using air conduction and bone conduction microphones
US10847173B2 (en) * 2018-02-13 2020-11-24 Intel Corporation Selection between signal sources based upon calculated signal to noise ratio

Patent Citations (9)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN102027536A (zh) * 2008-05-14 2011-04-20 索尼爱立信移动通讯有限公司 响应于说话时在用户面部中感测到的振动对麦克风信号进行自适应滤波
CN103229238A (zh) * 2010-11-24 2013-07-31 皇家飞利浦电子股份有限公司 用于产生音频信号的系统和方法
CN102761643A (zh) * 2011-04-26 2012-10-31 鹦鹉股份有限公司 组合话筒和耳机的音频头戴式耳机
US20130070935A1 (en) * 2011-09-19 2013-03-21 Bitwave Pte Ltd Multi-sensor signal optimization for speech communication
CN107452389A (zh) 2017-07-20 2017-12-08 大象声科(深圳)科技有限公司 一种通用的单声道实时降噪方法
CN108231086A (zh) * 2017-12-24 2018-06-29 航天恒星科技有限公司 一种基于fpga的深度学习语音增强器及方法
CN108986834A (zh) * 2018-08-22 2018-12-11 中国人民解放军陆军工程大学 基于编解码器架构与递归神经网络的骨导语音盲增强方法
CN109346075A (zh) 2018-10-15 2019-02-15 华为技术有限公司 通过人体振动识别用户语音以控制电子设备的方法和系统
CN109767783A (zh) * 2019-02-15 2019-05-17 深圳市汇顶科技股份有限公司 语音增强方法、装置、设备及存储介质

Non-Patent Citations (1)

* Cited by examiner, † Cited by third party
Title
See also references of EP4044181A4

Cited By (10)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2024528596A (ja) * 2021-07-15 2024-07-30 ドルビー ラボラトリーズ ライセンシング コーポレイション 発話向上
JP7739583B2 (ja) 2021-07-15 2025-09-16 ドルビー ラボラトリーズ ライセンシング コーポレイション 発話向上
WO2023056280A1 (en) * 2021-09-30 2023-04-06 Sonos, Inc. Noise reduction using synthetic audio
US12499901B2 (en) 2021-09-30 2025-12-16 Sonos, Inc. Noise reduction using synthetic audio
US12567428B1 (en) * 2021-10-11 2026-03-03 Meta Platforms Technologies, Llc Contact transducer based audio enhancement
WO2024002896A1 (en) * 2022-06-29 2024-01-04 Analog Devices International Unlimited Company Audio signal processing method and system for enhancing a bone-conducted audio signal using a machine learning model
US12080313B2 (en) 2022-06-29 2024-09-03 Analog Devices International Unlimited Company Audio signal processing method and system for enhancing a bone-conducted audio signal using a machine learning model
CN115171713A (zh) * 2022-06-30 2022-10-11 歌尔科技有限公司 语音降噪方法、装置、设备及计算机可读存储介质
WO2024000854A1 (zh) * 2022-06-30 2024-01-04 歌尔科技有限公司 语音降噪方法、装置、设备及计算机可读存储介质
CN115171713B (zh) * 2022-06-30 2025-08-01 歌尔科技有限公司 语音降噪方法、装置、设备及计算机可读存储介质

Also Published As

Publication number Publication date
KR102429152B1 (ko) 2022-08-03
KR20210043485A (ko) 2021-04-21
US20220392475A1 (en) 2022-12-08
EP4044181A1 (en) 2022-08-17
JP2022505997A (ja) 2022-01-17
EP4044181A4 (en) 2023-10-18

Similar Documents

Publication Publication Date Title
TWI763073B (zh) 融合骨振動感測器信號及麥克風信號的深度學習降噪方法
KR102429152B1 (ko) 골진동 센서 및 마이크로폰 신호를 융합한 딥 러닝 음성 추출 및 노이즈 저감 방법
CN109065067B (zh) 一种基于神经网络模型的会议终端语音降噪方法
Li et al. ICASSP 2021 deep noise suppression challenge: Decoupling magnitude and phase optimization with a two-stage deep network
CN111916101B (zh) 一种融合骨振动传感器和双麦克风信号的深度学习降噪方法及系统
JP6703525B2 (ja) 音源を強調するための方法及び機器
WO2022027423A1 (zh) 一种融合骨振动传感器和双麦克风信号的深度学习降噪方法及系统
CN102164328B (zh) 一种用于家庭环境的基于传声器阵列的音频输入系统
CN111292759A (zh) 一种基于神经网络的立体声回声消除方法及系统
Doclo et al. Multi-microphone noise reduction and dereverberation techniques for speech applications
US11832072B2 (en) Audio processing using distributed machine learning model
JP2009522942A (ja) 発話改善のためにマイク間レベル差を用いるシステム及び方法
US20240096343A1 (en) Voice quality enhancement method and related device
US20240331716A1 (en) Low-latency noise suppression
CN107564538A (zh) 一种实时语音通信的清晰度增强方法及系统
CN110364175B (zh) 语音增强方法及系统、通话设备
CN116030823A (zh) 一种语音信号处理方法、装置、计算机设备及存储介质
CN106328160A (zh) 一种基于双麦克的降噪方法
CN116129930B (zh) 无参考回路的回声消除装置及方法
Ohlenbusch et al. Low-complexity own voice reconstruction for hearables with an in-ear microphone
EP4571740A1 (en) Audio-visual speech enhancement
US20240363133A1 (en) Noise suppression model using gated linear units
CN113990337B (zh) 音频优化方法及相关装置、电子设备、存储介质
Balasubrahmanyam et al. A Comprehensive Review of Conventional to Modern Algorithms of Speech Enhancement
Sunohara et al. Low-latency real-time blind source separation with binaural directional hearing aids

Legal Events

Date Code Title Description
ENP Entry into the national phase

Ref document number: 2020563485

Country of ref document: JP

Kind code of ref document: A

NENP Non-entry into the national phase

Ref country code: DE

ENP Entry into the national phase

Ref document number: 2019920643

Country of ref document: EP

Effective date: 20220509