EP2363853A1 - Verfahren zur Schätzung des rauschfreien Spektrums eines Signals - Google Patents
Verfahren zur Schätzung des rauschfreien Spektrums eines Signals Download PDFInfo
- Publication number
- EP2363853A1 EP2363853A1 EP10450036A EP10450036A EP2363853A1 EP 2363853 A1 EP2363853 A1 EP 2363853A1 EP 10450036 A EP10450036 A EP 10450036A EP 10450036 A EP10450036 A EP 10450036A EP 2363853 A1 EP2363853 A1 EP 2363853A1
- Authority
- EP
- European Patent Office
- Prior art keywords
- signal
- spectrum
- coefficients
- noise
- model
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Withdrawn
Links
- 238000001228 spectrum Methods 0.000 title claims abstract description 48
- 238000000034 method Methods 0.000 title claims abstract description 35
- 239000000654 additive Substances 0.000 claims abstract description 7
- 230000000996 additive effect Effects 0.000 claims abstract description 7
- 238000012546 transfer Methods 0.000 claims abstract description 5
- 230000003595 spectral effect Effects 0.000 claims description 39
- 230000002708 enhancing effect Effects 0.000 claims description 4
- 230000006870 function Effects 0.000 description 11
- 238000012545 processing Methods 0.000 description 6
- 238000007476 Maximum Likelihood Methods 0.000 description 3
- 238000013459 approach Methods 0.000 description 3
- 230000000694 effects Effects 0.000 description 3
- 230000009466 transformation Effects 0.000 description 3
- 238000010420 art technique Methods 0.000 description 2
- 238000010586 diagram Methods 0.000 description 2
- 238000001914 filtration Methods 0.000 description 2
- 230000005236 sound signal Effects 0.000 description 2
- 238000004891 communication Methods 0.000 description 1
- 238000009472 formulation Methods 0.000 description 1
- 239000000203 mixture Substances 0.000 description 1
- 230000011218 segmentation Effects 0.000 description 1
- 230000000087 stabilizing effect Effects 0.000 description 1
- 238000007619 statistical method Methods 0.000 description 1
- 238000011410 subtraction method Methods 0.000 description 1
- 230000001629 suppression Effects 0.000 description 1
Images
Classifications
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L21/00—Speech or voice signal processing techniques to produce another audible or non-audible signal, e.g. visual or tactile, in order to modify its quality or its intelligibility
- G10L21/02—Speech enhancement, e.g. noise reduction or echo cancellation
- G10L21/0208—Noise filtering
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L25/00—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00
- G10L25/03—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 characterised by the type of extracted parameters
- G10L25/12—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 characterised by the type of extracted parameters the extracted parameters being prediction coefficients
Definitions
- the present invention relates to a method for estimating the clean spectrum of a signal degraded by additive noise, in particular a speech signal, by determining the coefficients of a predictive model of said clean spectrum.
- the invention further relates to a method for enhancing a signal based on this clean spectrum estimation.
- the enhancement of speech by digital signal processing means improves the quality and intelligibility of voice communication for a wide fan of applications, such as mobile telephony, hearing aids, teleconference systems, dictation systems, voice coders and automatic speech recognition systems.
- minimizing is intended to comprise both, making the cost function minimal as well as making the cost function at least a sufficiently low value, i.e. a value within a given or acceptable tolerance interval from that minimum.
- the biological hearing sense responds to the logarithm of the sound intensity.
- the invention is based on the insight that this bio-acoustic principle of logarithmic sense can be introduced into a novel cost function as stated above which takes into account the actual signal-to-noise ratio in each portion of the signal spectrum.
- the proposed cost function fits the model to the data for those regions with high SNR, and - as will be detailed later on - in low-SNR areas the fitting process is driven by the mentioned good fitting performance taking place on adjacent high-SNR areas.
- the inventive method thus leads to an interpolation effect from high-SNR to low-SNR spectral regions.
- said equation can be solved by holding E ( ⁇ ) and M ( ⁇ ) constant, solving the remaining linear problem, using the solution to re-evaluate the previous constant terms, and proceeding further iteratively.
- the method of the invention is suited for any predictive model known in the art.
- a parametric all-pole filter model an autoregressive coefficients filter (ARC) model, a reflection coefficients filter (RC) model, and/or a line spectral frequencies (LSF) model is used.
- ARC autoregressive coefficients filter
- RC reflection coefficients filter
- LSF line spectral frequencies
- a method for enhancing a digital signal, in particular a speech signal, with increased quality comprises the further steps of
- the signal is enhanced by means of a Wiener filter, a MMSE-based enhancement, or variants thereof, using said spectral signal-to-noise ratio.
- the inventor has found out analytically that the IWF method is equivalent to a method that results from the following minimization problem arg min a ⁇ ⁇ 2 ⁇ ⁇ ⁇
- the inventor has found out that a functional built on the ratio between the samples and the model, such as in (1), does not possess the desirable property of frequency selectivity while such a property would be desirable when not all spectral samples are available:
- the spectral samples at which the a priori SNR is low or very low do not represent a trustful reference for the estimation of the autoregressive model.
- the method of the present invention for estimating the clean speech spectrum is related to the minimization of the maximum likelihood (ML) of the ratio between the input noisy spectrum X ( ⁇ ) and the model of clean speech corrupted by additive noise.
- ML maximum likelihood
- X( ⁇ ) is modelled by a Gaussian distribution
- said maximum likelihood estimation turns out arg min a ⁇ ⁇ 2 ⁇ ⁇ ⁇
- the clean speech follows the autoregressive model defined in (2)
- a is the vector containing the autoregressive coefficients
- S v ( ⁇ ) is the power spectral density of the noise which is available a priori.
- the spectral mask is defined in terms of the a-priori signal-to-noise ratio for each frequency ("spectral" signal-to-noise ratio), SNR( ⁇ ).
- equation (4) Since equation (4) is nonlinear with respect to the autoregressive coefficients, its solution must and can be obtained by means of an iterative procedure, in which at each iteration a positive-definite Toeplitz linear system must be solved.
- Several techniques are available to solve Toeplitz systems, such as the well-known Levinson algorithm.
- the spectral mask (7) weights the importance of the spectral error between the noisy samples and the model of clean speech plus additive noise. This weight at each frequency depends on the respective signal-to-noise ratio.
- the spectral mask is close to 1 at that frequency, and the information at that frequency is valuable in the estimation.
- the spectral mask tends to zero, which implies that the relevance of the information at the frequency is low.
- the spectral mask, the signal-to-noise ratio, and therewith the clean speech model are estimated in an iterative fashion.
- the final solution is obtained either after several iterations or when successive partial solutions do not differ from each other substantially.
- the noise-substracted power spectrum can be
- the notation in the integrals refers to ⁇ ⁇ M ⁇ ⁇ ⁇ ⁇ ⁇ - ⁇ ⁇ M ⁇ ⁇ where M ⁇ ⁇ is the spectral weight (mask) M ( ⁇ ) at the K iteration. Since the spectral weight is present in all terms of the inverse problem (8f), its effect is that of weighting the relevance of the spectral samples. The magnitude of the weight depends on the local SNR ⁇ ⁇ , such that in areas with high SNR >> 1) the spectral weight tends to one, while in low-SNR areas ( ⁇ ⁇ ⁇ 1) it tends to zero. Note as comparison that in the noiseless case the spectral weight turns one for all frequencies, this meaning that the noiseless case need not require spectral selectivity.
- step (8f) is a linear inverse problem involving a positive-semidefinite symmetric Toeplitz system.
- it can be efficiently solved with the Levinson algorithm or any other algorithm to solve Toeplitz systems.
- Fig. 1 shows in a simplified fashion the processing-block diagram of a speech enhancement front-end (apparatus 100) that uses the method of the present invention.
- Fig. 2 shows the function of the clean speech estimation step (block 40) of Fig. 1 in detail.
- Block 10 performs the usual segmentation of the input digital signal into segments.
- Block 20 performs the spectral transformation of said segment.
- Said spectral transformation corresponds to the "Discrete Fourier Transform", “Discrete Sinus Transform” and/or to the “Fan-Chirp Transform”, among other popular choices.
- Block 30 carries out the estimation of the power spectrum of the noise according to known ad-hoc techniques. It is assumed that this block has memory facilities in such a way that the spectrum of the previous segments are stored therein. Therefore, if required, the estimation of the noise power spectrum can be performed by statistical methods over spectral data stretching within a reasonably long time span.
- Block 40 carries out the estimation of the clean speech model from the spectrum of the segment and the estimation of the noise power spectrum.
- the estimation of the clean speech model is based on the numerical implementation of the minimization problem (3), which represents the core method of the present invention.
- Block 50 computes numerically the signal-to-noise ratio for each frequency (spectral signal-to-noise ratio) from the estimated clean speech model and noise model.
- Block 60 enhances the spectrum of the input signal by means of state-of-art techniques that require the signal-to-noise ratio for each frequency.
- Wiener filter and its variants e.g. the root-square of the Wiener filter
- MMSE minimum-mean-square-error
- its variants e.g. the log-MMSE, et cet. (see P. J. Wolfe and S. J. Godsill, loc.cit.).
- Block 70 performs the inverse spectral transformation to block 20.
- the output of block 70 is the enhanced segment of the audio signal.
Landscapes
- Engineering & Computer Science (AREA)
- Computational Linguistics (AREA)
- Quality & Reliability (AREA)
- Signal Processing (AREA)
- Health & Medical Sciences (AREA)
- Audiology, Speech & Language Pathology (AREA)
- Human Computer Interaction (AREA)
- Physics & Mathematics (AREA)
- Acoustics & Sound (AREA)
- Multimedia (AREA)
- Circuit For Audible Band Transducer (AREA)
Priority Applications (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| EP10450036A EP2363853A1 (de) | 2010-03-04 | 2010-03-04 | Verfahren zur Schätzung des rauschfreien Spektrums eines Signals |
Applications Claiming Priority (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| EP10450036A EP2363853A1 (de) | 2010-03-04 | 2010-03-04 | Verfahren zur Schätzung des rauschfreien Spektrums eines Signals |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| EP2363853A1 true EP2363853A1 (de) | 2011-09-07 |
Family
ID=42316009
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| EP10450036A Withdrawn EP2363853A1 (de) | 2010-03-04 | 2010-03-04 | Verfahren zur Schätzung des rauschfreien Spektrums eines Signals |
Country Status (1)
| Country | Link |
|---|---|
| EP (1) | EP2363853A1 (de) |
Cited By (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| WO2013061232A1 (en) * | 2011-10-24 | 2013-05-02 | Koninklijke Philips Electronics N.V. | Audio signal noise attenuation |
| CN112562701A (zh) * | 2020-11-16 | 2021-03-26 | 华南理工大学 | 心音信号双通道自适应降噪算法、装置、介质及设备 |
| CN115238233A (zh) * | 2022-07-29 | 2022-10-25 | 中国科学院声学研究所 | 一种多测量矢量线谱估计方法及计算机设备和存储介质 |
| US20220358904A1 (en) * | 2019-03-20 | 2022-11-10 | Research Foundation Of The City University Of New York | Method for extracting speech from degraded signals by predicting the inputs to a speech vocoder |
Citations (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| WO2008031124A1 (de) * | 2006-09-15 | 2008-03-20 | Technische Universität Graz | Vorrichtung zur geräuschunterdrückung bei einem audiosignal |
| EP1970893A1 (de) * | 2007-03-13 | 2008-09-17 | Österreichische Akademie der Wissenschaften | Verfahren zur Schätzung von Signalkodierungsparametern |
-
2010
- 2010-03-04 EP EP10450036A patent/EP2363853A1/de not_active Withdrawn
Patent Citations (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| WO2008031124A1 (de) * | 2006-09-15 | 2008-03-20 | Technische Universität Graz | Vorrichtung zur geräuschunterdrückung bei einem audiosignal |
| EP1970893A1 (de) * | 2007-03-13 | 2008-09-17 | Österreichische Akademie der Wissenschaften | Verfahren zur Schätzung von Signalkodierungsparametern |
| WO2008109904A1 (en) | 2007-03-13 | 2008-09-18 | Österreichische Akademie der Wissenschaften | A method for estimating signal coding parameters |
Non-Patent Citations (8)
| Title |
|---|
| B. SIM; Y. TONG; J. CHANG; C. TAN: "A parametric formulation of the generalized spectral subtraction method", IEEE TRANSACTIONS ON SPEECH AND AUDIO PROCESSING, vol. 6, no. 4, July 1998 (1998-07-01), pages 328 - 337 |
| E. ZA- VAREHEI; S. VASEGHI; Q. YAN: "Speech enhancement using Kalman filters for restoration of short-time DFT trajectories", IEEE WORKSHOP AUTOMATIC SPEECH RECOGNITION AND UNDERSTANDING, 2005, pages 313 - 318 |
| J. H. L. HANSEN; M. A. CLEMENTS: "Constrained iterative speech enhancement with application to speech, recognition", IEEE TRANS. SIGNAL PROCESSING, vol. 39, no. 4, April 1991 (1991-04-01), pages 795 - 805 |
| K. FUNAKI: "Speech enhancement based on iterative Wiener filter using complex speech analysis", PROC. EUSIPCO 2008, 29 August 2008 (2008-08-29), Lausanne, Switzerland, pages 1 - 5, XP002593133, Retrieved from the Internet <URL:http://www.eurasip.org/Proceedings/Eusipco/Eusipco2008/papers/1569105040.pdf> [retrieved on 20100722] * |
| P. C. LOIZOU: "Speech enhancement: Theory and practice", 2007, CRC PRESS |
| P. J. WOLFE; S. J. GODSILL: "Efficient alternatives to the Ephraim and Malah suppression rule for au dio signal enhancement", EURASIP J. APPLIED SIGNAL PROCESSING, vol. 2003, no. 10, 2003, pages 1043 - 1051 |
| T. V. SREENIVAS; P. KIRNAPURE: "Codebook constrained Wiener filtering for speech enhancement", IEEE TRANS. SPEECH, AUDIO PROCESSING, vol. 4, no. 5, September 1996 (1996-09-01), pages 383 - 389 |
| Y. EPHRAIM; D. MALAH: "Speech enhancement using a minimum mean-square error short-time spectral amplitude estimator", IEEE TRANS. ACOUST., SPEECH, SIGNAL PROCESSING, vol. 32, no. 6, 1984, pages 1109 - 1121 |
Cited By (7)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| WO2013061232A1 (en) * | 2011-10-24 | 2013-05-02 | Koninklijke Philips Electronics N.V. | Audio signal noise attenuation |
| US9875748B2 (en) | 2011-10-24 | 2018-01-23 | Koninklijke Philips N.V. | Audio signal noise attenuation |
| US20220358904A1 (en) * | 2019-03-20 | 2022-11-10 | Research Foundation Of The City University Of New York | Method for extracting speech from degraded signals by predicting the inputs to a speech vocoder |
| US12020682B2 (en) * | 2019-03-20 | 2024-06-25 | Research Foundation Of The City University Of New York | Method for extracting speech from degraded signals by predicting the inputs to a speech vocoder |
| AU2020242078B2 (en) * | 2019-03-20 | 2026-01-29 | Research Foundation Of The City University Of New York | Method for extracting speech from degraded signals by predicting the inputs to a speech vocoder |
| CN112562701A (zh) * | 2020-11-16 | 2021-03-26 | 华南理工大学 | 心音信号双通道自适应降噪算法、装置、介质及设备 |
| CN115238233A (zh) * | 2022-07-29 | 2022-10-25 | 中国科学院声学研究所 | 一种多测量矢量线谱估计方法及计算机设备和存储介质 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US7313518B2 (en) | Noise reduction method and device using two pass filtering | |
| JP5068653B2 (ja) | 雑音のある音声信号を処理する方法および該方法を実行する装置 | |
| TWI420509B (zh) | 語音增強用雜訊變異量估計器 | |
| EP0807305B1 (de) | Verfahren zur rauschunterdrückung mittels spektraler subtraktion | |
| Soon et al. | Speech enhancement using 2-D Fourier transform | |
| CN100543842C (zh) | 基于多统计模型和最小均方误差实现背景噪声抑制的方法 | |
| US20100023327A1 (en) | Method for improving speech signal non-linear overweighting gain in wavelet packet transform domain | |
| CN109308904A (zh) | 一种阵列语音增强算法 | |
| WO2000017855A1 (en) | Noise suppression for low bitrate speech coder | |
| US7016839B2 (en) | MVDR based feature extraction for speech recognition | |
| US20130138437A1 (en) | Speech recognition apparatus based on cepstrum feature vector and method thereof | |
| CN103578477A (zh) | 基于噪声估计的去噪方法和装置 | |
| CN108962275A (zh) | 一种音乐噪声抑制方法及装置 | |
| CN116913308B (zh) | 一种平衡降噪量和语音音质的单通道语音增强方法 | |
| Poovarasan et al. | Speech enhancement using sliding window empirical mode decomposition and hurst-based technique | |
| Lei et al. | Speech enhancement for nonstationary noises by wavelet packet transform and adaptive noise estimation | |
| Daqrouq et al. | An investigation of speech enhancement using wavelet filtering method | |
| KR20110061781A (ko) | 실시간 잡음 추정에 기반하여 잡음을 제거하는 음성 처리 장치 및 방법 | |
| Batina et al. | Noise power spectrum estimation for speech enhancement using an autoregressive model for speech power spectrum dynamics | |
| EP1635331A1 (de) | Verfahren zur Abschätzung eines Signal-Rauschverhältnisses | |
| Bolisetty et al. | Speech enhancement using modified wiener filter based MMSE and speech presence probability estimation | |
| Gupta et al. | Speech enhancement using MMSE estimation and spectral subtraction methods | |
| Funaki | Speech enhancement based on iterative wiener filter using complex speech analysis | |
| Erkelens et al. | Speech enhancement based on Rayleigh mixture modeling of speech spectral amplitude distributions | |
| Gui et al. | Adaptive subband Wiener filtering for speech enhancement using critical-band gammatone filterbank |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| PUAI | Public reference made under article 153(3) epc to a published international application that has entered the european phase |
Free format text: ORIGINAL CODE: 0009012 |
|
| AK | Designated contracting states |
Kind code of ref document: A1 Designated state(s): AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC MK MT NL NO PL PT RO SE SI SK SM TR |
|
| AX | Request for extension of the european patent |
Extension state: AL BA ME RS |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: THE APPLICATION IS DEEMED TO BE WITHDRAWN |
|
| 18D | Application deemed to be withdrawn |
Effective date: 20120308 |