EP2363853A1 - Verfahren zur Schätzung des rauschfreien Spektrums eines Signals - Google Patents

Verfahren zur Schätzung des rauschfreien Spektrums eines Signals Download PDF

Info

Publication number
EP2363853A1
EP2363853A1 EP10450036A EP10450036A EP2363853A1 EP 2363853 A1 EP2363853 A1 EP 2363853A1 EP 10450036 A EP10450036 A EP 10450036A EP 10450036 A EP10450036 A EP 10450036A EP 2363853 A1 EP2363853 A1 EP 2363853A1
Authority
EP
European Patent Office
Prior art keywords
signal
spectrum
coefficients
noise
model
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Withdrawn
Application number
EP10450036A
Other languages
English (en)
French (fr)
Inventor
Luis Weruaga
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Innovationsagentur GmbH
Osterreichische Akademie der Wissenschaften
Original Assignee
Innovationsagentur GmbH
Osterreichische Akademie der Wissenschaften
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Innovationsagentur GmbH, Osterreichische Akademie der Wissenschaften filed Critical Innovationsagentur GmbH
Priority to EP10450036A priority Critical patent/EP2363853A1/de
Publication of EP2363853A1 publication Critical patent/EP2363853A1/de
Withdrawn legal-status Critical Current

Links

Images

Classifications

    • G—PHYSICS
    • G10—MUSICAL INSTRUMENTS; ACOUSTICS
    • G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L21/00—Speech or voice signal processing techniques to produce another audible or non-audible signal, e.g. visual or tactile, in order to modify its quality or its intelligibility
    • G10L21/02—Speech enhancement, e.g. noise reduction or echo cancellation
    • G10L21/0208—Noise filtering
    • G—PHYSICS
    • G10—MUSICAL INSTRUMENTS; ACOUSTICS
    • G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L25/00—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00
    • G10L25/03—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 characterised by the type of extracted parameters
    • G10L25/12—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 characterised by the type of extracted parameters the extracted parameters being prediction coefficients

Definitions

  • the present invention relates to a method for estimating the clean spectrum of a signal degraded by additive noise, in particular a speech signal, by determining the coefficients of a predictive model of said clean spectrum.
  • the invention further relates to a method for enhancing a signal based on this clean spectrum estimation.
  • the enhancement of speech by digital signal processing means improves the quality and intelligibility of voice communication for a wide fan of applications, such as mobile telephony, hearing aids, teleconference systems, dictation systems, voice coders and automatic speech recognition systems.
  • minimizing is intended to comprise both, making the cost function minimal as well as making the cost function at least a sufficiently low value, i.e. a value within a given or acceptable tolerance interval from that minimum.
  • the biological hearing sense responds to the logarithm of the sound intensity.
  • the invention is based on the insight that this bio-acoustic principle of logarithmic sense can be introduced into a novel cost function as stated above which takes into account the actual signal-to-noise ratio in each portion of the signal spectrum.
  • the proposed cost function fits the model to the data for those regions with high SNR, and - as will be detailed later on - in low-SNR areas the fitting process is driven by the mentioned good fitting performance taking place on adjacent high-SNR areas.
  • the inventive method thus leads to an interpolation effect from high-SNR to low-SNR spectral regions.
  • said equation can be solved by holding E ( ⁇ ) and M ( ⁇ ) constant, solving the remaining linear problem, using the solution to re-evaluate the previous constant terms, and proceeding further iteratively.
  • the method of the invention is suited for any predictive model known in the art.
  • a parametric all-pole filter model an autoregressive coefficients filter (ARC) model, a reflection coefficients filter (RC) model, and/or a line spectral frequencies (LSF) model is used.
  • ARC autoregressive coefficients filter
  • RC reflection coefficients filter
  • LSF line spectral frequencies
  • a method for enhancing a digital signal, in particular a speech signal, with increased quality comprises the further steps of
  • the signal is enhanced by means of a Wiener filter, a MMSE-based enhancement, or variants thereof, using said spectral signal-to-noise ratio.
  • the inventor has found out analytically that the IWF method is equivalent to a method that results from the following minimization problem arg min a ⁇ ⁇ 2 ⁇ ⁇ ⁇
  • the inventor has found out that a functional built on the ratio between the samples and the model, such as in (1), does not possess the desirable property of frequency selectivity while such a property would be desirable when not all spectral samples are available:
  • the spectral samples at which the a priori SNR is low or very low do not represent a trustful reference for the estimation of the autoregressive model.
  • the method of the present invention for estimating the clean speech spectrum is related to the minimization of the maximum likelihood (ML) of the ratio between the input noisy spectrum X ( ⁇ ) and the model of clean speech corrupted by additive noise.
  • ML maximum likelihood
  • X( ⁇ ) is modelled by a Gaussian distribution
  • said maximum likelihood estimation turns out arg min a ⁇ ⁇ 2 ⁇ ⁇ ⁇
  • the clean speech follows the autoregressive model defined in (2)
  • a is the vector containing the autoregressive coefficients
  • S v ( ⁇ ) is the power spectral density of the noise which is available a priori.
  • the spectral mask is defined in terms of the a-priori signal-to-noise ratio for each frequency ("spectral" signal-to-noise ratio), SNR( ⁇ ).
  • equation (4) Since equation (4) is nonlinear with respect to the autoregressive coefficients, its solution must and can be obtained by means of an iterative procedure, in which at each iteration a positive-definite Toeplitz linear system must be solved.
  • Several techniques are available to solve Toeplitz systems, such as the well-known Levinson algorithm.
  • the spectral mask (7) weights the importance of the spectral error between the noisy samples and the model of clean speech plus additive noise. This weight at each frequency depends on the respective signal-to-noise ratio.
  • the spectral mask is close to 1 at that frequency, and the information at that frequency is valuable in the estimation.
  • the spectral mask tends to zero, which implies that the relevance of the information at the frequency is low.
  • the spectral mask, the signal-to-noise ratio, and therewith the clean speech model are estimated in an iterative fashion.
  • the final solution is obtained either after several iterations or when successive partial solutions do not differ from each other substantially.
  • the noise-substracted power spectrum can be
  • the notation in the integrals refers to ⁇ ⁇ M ⁇ ⁇ ⁇ ⁇ ⁇ - ⁇ ⁇ M ⁇ ⁇ where M ⁇ ⁇ is the spectral weight (mask) M ( ⁇ ) at the K iteration. Since the spectral weight is present in all terms of the inverse problem (8f), its effect is that of weighting the relevance of the spectral samples. The magnitude of the weight depends on the local SNR ⁇ ⁇ , such that in areas with high SNR >> 1) the spectral weight tends to one, while in low-SNR areas ( ⁇ ⁇ ⁇ 1) it tends to zero. Note as comparison that in the noiseless case the spectral weight turns one for all frequencies, this meaning that the noiseless case need not require spectral selectivity.
  • step (8f) is a linear inverse problem involving a positive-semidefinite symmetric Toeplitz system.
  • it can be efficiently solved with the Levinson algorithm or any other algorithm to solve Toeplitz systems.
  • Fig. 1 shows in a simplified fashion the processing-block diagram of a speech enhancement front-end (apparatus 100) that uses the method of the present invention.
  • Fig. 2 shows the function of the clean speech estimation step (block 40) of Fig. 1 in detail.
  • Block 10 performs the usual segmentation of the input digital signal into segments.
  • Block 20 performs the spectral transformation of said segment.
  • Said spectral transformation corresponds to the "Discrete Fourier Transform", “Discrete Sinus Transform” and/or to the “Fan-Chirp Transform”, among other popular choices.
  • Block 30 carries out the estimation of the power spectrum of the noise according to known ad-hoc techniques. It is assumed that this block has memory facilities in such a way that the spectrum of the previous segments are stored therein. Therefore, if required, the estimation of the noise power spectrum can be performed by statistical methods over spectral data stretching within a reasonably long time span.
  • Block 40 carries out the estimation of the clean speech model from the spectrum of the segment and the estimation of the noise power spectrum.
  • the estimation of the clean speech model is based on the numerical implementation of the minimization problem (3), which represents the core method of the present invention.
  • Block 50 computes numerically the signal-to-noise ratio for each frequency (spectral signal-to-noise ratio) from the estimated clean speech model and noise model.
  • Block 60 enhances the spectrum of the input signal by means of state-of-art techniques that require the signal-to-noise ratio for each frequency.
  • Wiener filter and its variants e.g. the root-square of the Wiener filter
  • MMSE minimum-mean-square-error
  • its variants e.g. the log-MMSE, et cet. (see P. J. Wolfe and S. J. Godsill, loc.cit.).
  • Block 70 performs the inverse spectral transformation to block 20.
  • the output of block 70 is the enhanced segment of the audio signal.

Landscapes

  • Engineering & Computer Science (AREA)
  • Computational Linguistics (AREA)
  • Quality & Reliability (AREA)
  • Signal Processing (AREA)
  • Health & Medical Sciences (AREA)
  • Audiology, Speech & Language Pathology (AREA)
  • Human Computer Interaction (AREA)
  • Physics & Mathematics (AREA)
  • Acoustics & Sound (AREA)
  • Multimedia (AREA)
  • Circuit For Audible Band Transducer (AREA)
EP10450036A 2010-03-04 2010-03-04 Verfahren zur Schätzung des rauschfreien Spektrums eines Signals Withdrawn EP2363853A1 (de)

Priority Applications (1)

Application Number Priority Date Filing Date Title
EP10450036A EP2363853A1 (de) 2010-03-04 2010-03-04 Verfahren zur Schätzung des rauschfreien Spektrums eines Signals

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
EP10450036A EP2363853A1 (de) 2010-03-04 2010-03-04 Verfahren zur Schätzung des rauschfreien Spektrums eines Signals

Publications (1)

Publication Number Publication Date
EP2363853A1 true EP2363853A1 (de) 2011-09-07

Family

ID=42316009

Family Applications (1)

Application Number Title Priority Date Filing Date
EP10450036A Withdrawn EP2363853A1 (de) 2010-03-04 2010-03-04 Verfahren zur Schätzung des rauschfreien Spektrums eines Signals

Country Status (1)

Country Link
EP (1) EP2363853A1 (de)

Cited By (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
WO2013061232A1 (en) * 2011-10-24 2013-05-02 Koninklijke Philips Electronics N.V. Audio signal noise attenuation
CN112562701A (zh) * 2020-11-16 2021-03-26 华南理工大学 心音信号双通道自适应降噪算法、装置、介质及设备
CN115238233A (zh) * 2022-07-29 2022-10-25 中国科学院声学研究所 一种多测量矢量线谱估计方法及计算机设备和存储介质
US20220358904A1 (en) * 2019-03-20 2022-11-10 Research Foundation Of The City University Of New York Method for extracting speech from degraded signals by predicting the inputs to a speech vocoder

Citations (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
WO2008031124A1 (de) * 2006-09-15 2008-03-20 Technische Universität Graz Vorrichtung zur geräuschunterdrückung bei einem audiosignal
EP1970893A1 (de) * 2007-03-13 2008-09-17 Österreichische Akademie der Wissenschaften Verfahren zur Schätzung von Signalkodierungsparametern

Patent Citations (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
WO2008031124A1 (de) * 2006-09-15 2008-03-20 Technische Universität Graz Vorrichtung zur geräuschunterdrückung bei einem audiosignal
EP1970893A1 (de) * 2007-03-13 2008-09-17 Österreichische Akademie der Wissenschaften Verfahren zur Schätzung von Signalkodierungsparametern
WO2008109904A1 (en) 2007-03-13 2008-09-18 Österreichische Akademie der Wissenschaften A method for estimating signal coding parameters

Non-Patent Citations (8)

* Cited by examiner, † Cited by third party
Title
B. SIM; Y. TONG; J. CHANG; C. TAN: "A parametric formulation of the generalized spectral subtraction method", IEEE TRANSACTIONS ON SPEECH AND AUDIO PROCESSING, vol. 6, no. 4, July 1998 (1998-07-01), pages 328 - 337
E. ZA- VAREHEI; S. VASEGHI; Q. YAN: "Speech enhancement using Kalman filters for restoration of short-time DFT trajectories", IEEE WORKSHOP AUTOMATIC SPEECH RECOGNITION AND UNDERSTANDING, 2005, pages 313 - 318
J. H. L. HANSEN; M. A. CLEMENTS: "Constrained iterative speech enhancement with application to speech, recognition", IEEE TRANS. SIGNAL PROCESSING, vol. 39, no. 4, April 1991 (1991-04-01), pages 795 - 805
K. FUNAKI: "Speech enhancement based on iterative Wiener filter using complex speech analysis", PROC. EUSIPCO 2008, 29 August 2008 (2008-08-29), Lausanne, Switzerland, pages 1 - 5, XP002593133, Retrieved from the Internet <URL:http://www.eurasip.org/Proceedings/Eusipco/Eusipco2008/papers/1569105040.pdf> [retrieved on 20100722] *
P. C. LOIZOU: "Speech enhancement: Theory and practice", 2007, CRC PRESS
P. J. WOLFE; S. J. GODSILL: "Efficient alternatives to the Ephraim and Malah suppression rule for au dio signal enhancement", EURASIP J. APPLIED SIGNAL PROCESSING, vol. 2003, no. 10, 2003, pages 1043 - 1051
T. V. SREENIVAS; P. KIRNAPURE: "Codebook constrained Wiener filtering for speech enhancement", IEEE TRANS. SPEECH, AUDIO PROCESSING, vol. 4, no. 5, September 1996 (1996-09-01), pages 383 - 389
Y. EPHRAIM; D. MALAH: "Speech enhancement using a minimum mean-square error short-time spectral amplitude estimator", IEEE TRANS. ACOUST., SPEECH, SIGNAL PROCESSING, vol. 32, no. 6, 1984, pages 1109 - 1121

Cited By (7)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
WO2013061232A1 (en) * 2011-10-24 2013-05-02 Koninklijke Philips Electronics N.V. Audio signal noise attenuation
US9875748B2 (en) 2011-10-24 2018-01-23 Koninklijke Philips N.V. Audio signal noise attenuation
US20220358904A1 (en) * 2019-03-20 2022-11-10 Research Foundation Of The City University Of New York Method for extracting speech from degraded signals by predicting the inputs to a speech vocoder
US12020682B2 (en) * 2019-03-20 2024-06-25 Research Foundation Of The City University Of New York Method for extracting speech from degraded signals by predicting the inputs to a speech vocoder
AU2020242078B2 (en) * 2019-03-20 2026-01-29 Research Foundation Of The City University Of New York Method for extracting speech from degraded signals by predicting the inputs to a speech vocoder
CN112562701A (zh) * 2020-11-16 2021-03-26 华南理工大学 心音信号双通道自适应降噪算法、装置、介质及设备
CN115238233A (zh) * 2022-07-29 2022-10-25 中国科学院声学研究所 一种多测量矢量线谱估计方法及计算机设备和存储介质

Similar Documents

Publication Publication Date Title
US7313518B2 (en) Noise reduction method and device using two pass filtering
JP5068653B2 (ja) 雑音のある音声信号を処理する方法および該方法を実行する装置
TWI420509B (zh) 語音增強用雜訊變異量估計器
EP0807305B1 (de) Verfahren zur rauschunterdrückung mittels spektraler subtraktion
Soon et al. Speech enhancement using 2-D Fourier transform
CN100543842C (zh) 基于多统计模型和最小均方误差实现背景噪声抑制的方法
US20100023327A1 (en) Method for improving speech signal non-linear overweighting gain in wavelet packet transform domain
CN109308904A (zh) 一种阵列语音增强算法
WO2000017855A1 (en) Noise suppression for low bitrate speech coder
US7016839B2 (en) MVDR based feature extraction for speech recognition
US20130138437A1 (en) Speech recognition apparatus based on cepstrum feature vector and method thereof
CN103578477A (zh) 基于噪声估计的去噪方法和装置
CN108962275A (zh) 一种音乐噪声抑制方法及装置
CN116913308B (zh) 一种平衡降噪量和语音音质的单通道语音增强方法
Poovarasan et al. Speech enhancement using sliding window empirical mode decomposition and hurst-based technique
Lei et al. Speech enhancement for nonstationary noises by wavelet packet transform and adaptive noise estimation
Daqrouq et al. An investigation of speech enhancement using wavelet filtering method
KR20110061781A (ko) 실시간 잡음 추정에 기반하여 잡음을 제거하는 음성 처리 장치 및 방법
Batina et al. Noise power spectrum estimation for speech enhancement using an autoregressive model for speech power spectrum dynamics
EP1635331A1 (de) Verfahren zur Abschätzung eines Signal-Rauschverhältnisses
Bolisetty et al. Speech enhancement using modified wiener filter based MMSE and speech presence probability estimation
Gupta et al. Speech enhancement using MMSE estimation and spectral subtraction methods
Funaki Speech enhancement based on iterative wiener filter using complex speech analysis
Erkelens et al. Speech enhancement based on Rayleigh mixture modeling of speech spectral amplitude distributions
Gui et al. Adaptive subband Wiener filtering for speech enhancement using critical-band gammatone filterbank

Legal Events

Date Code Title Description
PUAI Public reference made under article 153(3) epc to a published international application that has entered the european phase

Free format text: ORIGINAL CODE: 0009012

AK Designated contracting states

Kind code of ref document: A1

Designated state(s): AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC MK MT NL NO PL PT RO SE SI SK SM TR

AX Request for extension of the european patent

Extension state: AL BA ME RS

STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: THE APPLICATION IS DEEMED TO BE WITHDRAWN

18D Application deemed to be withdrawn

Effective date: 20120308