WO2015188627A1 - 一种音频信号的时域包络处理方法及装置、编码器 - Google Patents

一种音频信号的时域包络处理方法及装置、编码器 Download PDF

Info

Publication number
WO2015188627A1
WO2015188627A1 PCT/CN2015/071727 CN2015071727W WO2015188627A1 WO 2015188627 A1 WO2015188627 A1 WO 2015188627A1 CN 2015071727 W CN2015071727 W CN 2015071727W WO 2015188627 A1 WO2015188627 A1 WO 2015188627A1
Authority
WO
WIPO (PCT)
Prior art keywords
signal
current frame
subframe
subframes
band signal
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/CN2015/071727
Other languages
English (en)
French (fr)
Inventor
刘泽新
苗磊
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Huawei Technologies Co Ltd
Original Assignee
Huawei Technologies Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Huawei Technologies Co Ltd filed Critical Huawei Technologies Co Ltd
Priority to JP2016572398A priority Critical patent/JP6510566B2/ja
Priority to KR1020167033851A priority patent/KR101896486B1/ko
Priority to EP19169470.2A priority patent/EP3579229B1/en
Priority to EP15806700.9A priority patent/EP3133599B1/en
Publication of WO2015188627A1 publication Critical patent/WO2015188627A1/zh
Priority to US15/372,130 priority patent/US9799343B2/en
Anticipated expiration legal-status Critical
Priority to US15/708,617 priority patent/US10170128B2/en
Priority to US16/201,647 priority patent/US10580423B2/en
Ceased legal-status Critical Current

Links

Images

Classifications

    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L19/00Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis
    • G10L19/02Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis using spectral analysis, e.g. transform vocoders or subband vocoders
    • G10L19/022Blocking, i.e. grouping of samples in time; Choice of analysis windows; Overlap factoring
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L19/00Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis
    • G10L19/02Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis using spectral analysis, e.g. transform vocoders or subband vocoders
    • G10L19/032Quantisation or dequantisation of spectral components
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L19/00Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis
    • G10L19/04Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis using predictive techniques
    • G10L19/08Determination or coding of the excitation function; Determination or coding of the long-term prediction parameters
    • G10L19/12Determination or coding of the excitation function; Determination or coding of the long-term prediction parameters the excitation function being a code excitation, e.g. in code excited linear prediction [CELP] vocoders
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L19/00Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis
    • G10L19/04Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis using predictive techniques
    • G10L19/08Determination or coding of the excitation function; Determination or coding of the long-term prediction parameters
    • G10L19/12Determination or coding of the excitation function; Determination or coding of the long-term prediction parameters the excitation function being a code excitation, e.g. in code excited linear prediction [CELP] vocoders
    • G10L19/135Vector sum excited linear prediction [VSELP]
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L19/00Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis
    • G10L19/04Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis using predictive techniques
    • G10L19/16Vocoder architecture
    • G10L19/18Vocoders using multiple modes
    • G10L19/20Vocoders using multiple modes using sound class specific coding, hybrid encoders or object based coding
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L21/00Speech or voice signal processing techniques to produce another audible or non-audible signal, e.g. visual or tactile, in order to modify its quality or its intelligibility
    • G10L21/02Speech enhancement, e.g. noise reduction or echo cancellation
    • G10L21/038Speech enhancement, e.g. noise reduction or echo cancellation using band spreading techniques
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L25/00Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00
    • G10L25/45Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 characterised by the type of analysis window

Definitions

  • the embodiments of the present invention relate to the field of communications technologies, and in particular, to a time domain envelope processing method and apparatus for an audio signal, and an encoder.
  • the existing process of calculating and quantizing the time domain envelope is: according to the number of calculation time domain envelopes M, M is positive An integer, which divides the pre-processed original high-band signal and the predicted high-band signal into M subframes, adds a window to the subframe, and then calculates the pre-processed original high-band signal and the predicted high-band signal in each subframe. Energy or amplitude ratio.
  • the number M of the calculated time domain envelopes set in advance is determined according to the length of the lookahead buffer.
  • the forward buffer is the current frame in order to calculate some parameters.
  • the last sample of the input signal is not used for buffering. It is used when calculating parameters in the next frame.
  • the current frame uses the sample buffer of the previous frame.
  • the cached samples are forward buffers, and the number of cached samples is the length of the forward buffer.
  • the above problem with the processing of the time domain envelope is: when solving the time domain envelope, all the symmetric windows are used, and in order to ensure the inter-subframe and inter-frame aliasing, according to the forward cache (lookahead)
  • the length calculates multiple time domain envelopes.
  • the time domain resolution of the signal is too high, it will cause discontinuity of the energy in the frame, thus introducing a poor hearing experience.
  • Embodiments of the present invention provide a time domain envelope processing method and apparatus for an audio signal, and an encoder, which can solve the problem of discontinuity of intra-frame energy caused by calculating a time domain envelope.
  • an embodiment of the present invention provides a time domain envelope processing method for an audio signal, including:
  • the calculating the time domain envelope of each of the subframes includes:
  • a subframe other than the topmost subframe and the lastmost subframe in the M subframes is windowed.
  • the time domain envelope is solved by using different window lengths and/or window shapes under different conditions, and the introduction of the time domain envelope difference is introduced.
  • the effect of energy discontinuity can improve the performance of the output signal.
  • the method further includes:
  • the asymmetric window is determined according to a length of a forward buffer of the high band signal of the current frame signal and the number M of time domain envelopes.
  • the subframe and the subframe except the front end of the M subframes are windowed, including:
  • a subframe other than the topmost subframe and the lastmost subframe among the M subframes is windowed by an asymmetric window.
  • the window length of the asymmetric window is different from the subframe of the M subframe except the frontmost end and the last end
  • the window used for windowing outside the frame is the same as the window length of the window.
  • the determining, by the length of the forward buffer of the high-band signal of the current frame audio signal, the asymmetric window comprises:
  • the length of the forward buffer of the high band signal of the current frame signal is less than the first threshold
  • the length of the forward buffer according to the high band signal of the previous frame signal of the current frame and the high band signal of the current frame signal Determining the asymmetric window, wherein an asymmetric window of the highest-end signal of the high-band signal of the previous frame signal of the current frame and an asymmetry window of the high-band signal of the current frame signal are asymmetric
  • the aliasing portion of the window is equal to the length of the forward buffer of the high band signal of the current frame signal, the first threshold being equal to the frame length of the high band signal of the current frame divided by M.
  • Determining the length of the forward buffer of the high-band signal of the current frame signal determines an asymmetric window, including:
  • the high-band signal according to the previous frame signal of the current frame and the forward buffer of the high-band signal of the current frame signal Determining the asymmetric window, wherein the asymmetric window of the highest-end signal of the high-band signal of the previous frame signal of the current frame and the foremost terminal frame of the high-band signal of the current frame signal are used
  • the aliasing portion of the asymmetric window is equal to the first threshold, the first threshold being equal to the frame length of the high band signal of the current frame divided by M.
  • the time domain envelope is determined according to one of the following manners Number M:
  • M1 and M2 are positive integers, and M2>M1.
  • the method further includes:
  • the time domain of each of the subframes The envelope is smoothed.
  • an embodiment of the present invention provides a time domain envelope processing apparatus for an audio signal, including:
  • a high-band signal acquisition module configured to obtain a high-band signal of the current frame signal according to the received current frame signal
  • a subframe acquiring module configured to divide the high-band signal of the current frame into M subframes according to a predetermined number of time-domain envelopes M, where M is an integer greater than or equal to 2;
  • a time domain envelope acquisition module configured to calculate a time domain envelope of each of the subframes
  • the time domain envelope obtaining module is specifically configured to:
  • a subframe other than the topmost subframe and the lastmost subframe in the M subframes is windowed.
  • the time domain envelope is solved by using different window lengths and/or window shapes under different conditions, thereby reducing the introduction of the time domain envelope difference.
  • the effect of energy discontinuity can improve the performance of the output signal.
  • the time domain envelope obtaining module is further configured to:
  • the asymmetric window is determined according to a length of a forward buffer of the high band signal of the current frame signal and the number M of time domain envelopes.
  • the time domain envelope obtaining module is specifically configured to:
  • Windowing the foremost subframe of the M subframes and the last subframe of the M subframes by using an asymmetric window, and selecting the subframe of the M subframes except the frontmost end and the Subframes other than the endmost subframe are windowed with an asymmetric window.
  • a window length of the asymmetric window and a subframe that is the most front end and the maximum of the M subframes is the same as the window length.
  • M1 and M2 are positive integers, and M2>M1.
  • An embodiment of the third aspect of the present invention discloses an encoder, and the encoder is specifically configured to:
  • the calculating a time domain envelope of the predicted highband signal includes:
  • the quantized time domain envelope is encoded.
  • the encoder provided by the embodiment of the present invention solves the time domain envelope by using different window lengths and/or window shapes under different conditions, thereby reducing the influence of energy discontinuity introduced due to too much difference in the time domain envelope, and improving The performance of the output signal.
  • 1 is a schematic diagram of a process of encoding an audio signal
  • Embodiment 2 is a flowchart of Embodiment 1 of a time domain envelope processing method for an audio signal according to the present invention
  • FIG. 3 is a schematic diagram of processing an audio signal in an embodiment of the present invention.
  • FIG. 4 is a schematic diagram of processing an audio signal according to another embodiment of the present invention.
  • FIG. 5 is a schematic diagram of processing an audio signal according to another embodiment of the present invention.
  • Embodiment 6 is a flowchart of Embodiment 2 of a time domain envelope processing method for an audio signal according to the present invention
  • FIG. 7 is a schematic structural diagram of a time domain envelope processing apparatus according to an embodiment of the present invention.
  • FIG. 8 is a schematic structural diagram of an encoder according to an embodiment of the present invention.
  • FIG. 1 is a schematic diagram of a process of encoding a speech audio signal.
  • the original audio signal is first decomposed to obtain a low-band signal of the original audio signal.
  • High-band signal followed by encoding the low-band signal through an existing algorithm
  • ACELP Algebraic Code Excited Linear Prediction
  • CELP Code Excited Linear Prediction
  • the low-band excitation signal is obtained, and the low-band excitation signal is preprocessed; for the high-band signal of the original audio signal, the pre-processing is first performed, and then linear prediction is performed.
  • the LP coefficient is analyzed to obtain the LP coefficient, and the LP coefficient is quantized.
  • the preprocessed low-band excitation signal is passed through the LP synthesis filter (the filter coefficient is the quantized LP coefficient) to obtain a predicted high-band signal.
  • the processed high-band signal and the predicted high-band signal calculate and quantize the time-domain envelope of the high-band signal, and finally output the coded code stream (MUX).
  • the process of calculating and quantizing the time-domain envelope of the high-band signal is: The number N of the time domain envelopes set in advance is divided into N subframes by the preprocessed highband signal and the predicted highband signal, and each subframe is entered.
  • N of the time domain envelopes set in advance is determined according to the length of the lookahead cache, and N is a positive integer.
  • Embodiments of the present invention provide a time domain envelope processing method for an audio signal, which is mainly used for the steps of calculating and quantizing a time domain envelope shown in FIG. 1, and can also be used for solving a time domain envelope using the same principle.
  • the time domain envelope processing method of the audio signal provided by the embodiment of the present invention is described in detail below with reference to the accompanying drawings.
  • Embodiment 1 of a time domain envelope processing method for an audio signal according to the present invention. As shown in FIG. 2, the method in this embodiment includes:
  • the current frame signal may be a voice signal, a music signal, or a noise signal, and no specific limitation is imposed herein.
  • the number of time domain envelopes M to be determined may be determined according to the overall algorithm requirements and empirical values.
  • the number of time domain envelopes M is determined, for example, by the encoder according to the overall algorithm or empirical value, and will not change after the determination. For example, for an input signal of 20 ms one frame, if the input signal is relatively stable, four or two time domain envelopes are solved, but for some non-stationary signals, more than eight time domain envelopes need to be solved.
  • calculating the time domain envelope of each subframe includes:
  • the topmost subframe of the M subframes and the last subframe of the M subframes are windowed by using an asymmetric window.
  • a sub-frame other than the topmost sub-frame and the last-most sub-frame among the M sub-frames is windowed.
  • the method of this embodiment may further include: before the windowing of the topmost subframe of the M subframes and the last subframe of the M subframes by using the asymmetric window, the method in this embodiment may further include:
  • the asymmetric window is determined based on the length of the forward buffer of the high band signal of the current frame signal and the number M of time domain envelopes.
  • the windowing of the subframes other than the frontmost subframe and the last subframe of the M subframes may include:
  • Sub-frames other than the foremost sub-frame and the last-most sub-frame of the M sub-frames are windowed by a symmetric window; or,
  • a subframe other than the foremost subframe and the last subframe of the M subframes is windowed by an asymmetric window.
  • the window length of the asymmetric window used for windowing the front terminal frame and the endmost subframe is different from the subframe of the M subframe except the frontmost end and the endmost subframe.
  • the window used for windowing is the same as the window length of the window.
  • determining an asymmetric window according to a length of a forward buffer of a high-band signal of a current frame audio signal includes:
  • the length of the forward buffer of the high band signal of the current frame signal is less than the first threshold, determining the asymmetric window according to the high band signal of the previous frame signal of the current frame and the forward buffer length of the high band signal of the current frame signal , wherein the asymmetric window of the highest-end subframe of the high-band signal of the previous frame signal of the current frame and the asymmetrical window of the front-end terminal frame of the high-band signal of the current frame signal are equal to the current frame signal
  • the length of the forward buffer of the high band signal, the first threshold being equal to the frame length of the high band signal of the current frame divided by M.
  • the forward buffer of the high band signal according to the current frame signal The length determines the asymmetric window, including:
  • the first threshold is equal to the frame length of the high band signal of the current frame divided by M.
  • the number of time domain envelopes M is determined according to one of the following ways:
  • the method of this embodiment may further include:
  • the time domain envelope of each subframe is smoothed.
  • the time domain envelope is smoothed, and the time domain envelopes of the two adjacent subframes are weighted, and the weighted time domain envelope is used as the time domain envelope of the two subframes.
  • the pitch period of the low band signal is greater than a given threshold (greater than 70 samples, at this time, the low band signal
  • the sampling rate is 12.8 kHz sampling
  • the decoded high-band signal time domain envelope is smoothed, otherwise the time domain envelope is kept unchanged. Smoothing can be:
  • Env[N] 0.5*(env[N-1]+env[N]).
  • env[] is the time domain envelope.
  • step numbers are merely examples for facilitating understanding of the embodiments of the present invention, and are not intended to limit the embodiments of the present invention.
  • a sub-frame other than the front-end and the end-most sub-frames may be windowed, and then the front-end and end-most sub-frames are windowed.
  • FIG. 3 is a schematic diagram of processing an audio signal in an embodiment of the present invention.
  • the original audio signal is first decomposed to obtain a low-band signal and a high-band signal of the original audio signal, and then the low-band signal is encoded by an existing algorithm. Obtaining a low-band code stream.
  • a low-band excitation signal is obtained, and the low-band excitation signal is preprocessed; for the original high-band signal of the original audio signal, the pre-processing is performed first, and then The LP analysis yields an LP coefficient that is quantized.
  • the preprocessed low band excitation signal is then passed through an LP synthesis filter (the filter coefficients are quantized LP coefficients) to obtain a predicted high band signal.
  • the time domain envelope of the high band signal is calculated and quantized according to the preprocessed high band signal and the predicted high band signal, and finally the coded code stream is output.
  • the N+1th frame is divided into M subframes according to the number of time domain envelopes to be calculated, and M is a positive integer.
  • M can be 3, 4, 5, 8, and the like. There are no restrictions here.
  • the window of the foremost subframe among the M subframes and the last subframe of the M subframes are windowed by an asymmetric window.
  • the topmost subframe among the M subframes of the N+1 frame is a subframe that overlaps with the signal of the previous frame (N frame); the last subframe is the next frame (N+2 frame, in the figure)
  • the signal not shown) has a sub-frame of overlapping portions.
  • the frontmost subframe is the leftmost subframe of the N+1 frame
  • the last subframe is the rightmost subframe of the N+1 frame. It can be understood that the leftmost and rightmost are only a specific example in conjunction with FIG. 3, and are not intended to limit the embodiments of the present invention.
  • the division of the actual neutron frame is such that there is no directional restriction of the leftmost and rightmost.
  • the asymmetric window used for windowing the most front-end and end-most sub-frames can be completely Same, it can be different. There are no restrictions here.
  • the window length of the asymmetric window used by the frontmost terminal frame is the same as the window length of the asymmetric window used by the endmost subframe.
  • a subframe other than the topmost subframe and the last subframe of the M subframes of the N+1 frame is windowed by a symmetric window.
  • the window length of the asymmetric window used for windowing the topmost and lastmost subframes is equal to the window length of the symmetric window employed for the other subframes. It can be understood that in another possible manner, the window length of the asymmetric window and the window length of the symmetric window may also be unequal.
  • the frame length of the N+1th frame is 80 samples and the sampling rate is 4 kHz, eight time domain envelopes can be solved.
  • the number N of time domain envelopes may be predetermined based on other information of the N+1 frame.
  • the following is an example of an implementation that determines the number N of time domain envelopes:
  • the second threshold can be 70 samples.
  • the low-band signal of the (N+1)th frame can be obtained when the signal of the (N+1)th frame is decomposed, and the method used for signal decomposition and the method of solving the pitch period of the low-band signal can be used.
  • the method used for signal decomposition and the method of solving the pitch period of the low-band signal can be used. There is any one of the techniques, and no specific limitation is imposed here.
  • the asymmetric window when windowing the foremost sub-frame and the endmost sub-frame with an asymmetric window, is determined according to the length of the forward buffer.
  • the window length of the asymmetric window used for windowing is The window length of the symmetrical window can be 20 samples.
  • the first threshold is obtained by dividing the frame length by the number of envelopes, in this example the first threshold is equal to 10. Then, when the length of the current buffer is less than 10 samples, the 8th subframe (ie, the endmost child) The frame used and the alias used by the first subframe (ie, the topmost subframe) are equal to the length of the forward buffer.
  • the length of the left side of the window adopted by the 8th subframe and the left side of the window adopted by the 1st subframe may be equal to the other side (for example, the first subframe is adopted)
  • the window length (10 samples) on the right side of the window or the left side of the window used in the eighth sub-frame can also be set according to experience (for example, the same as when the forward buffer is less than 10 samples) length).
  • the window length of the asymmetric window used for windowing is The window length of the symmetrical window can be 40 samples.
  • the first threshold is obtained by dividing the frame length by the number of envelopes, in this example the first threshold is equal to 20.
  • the time domain energy of the pre-processed original high-band signal and the predicted high-band signal in each sub-frame or the average of the amplitude of each sample in the sub-frame is calculated.
  • the method for signal processing provided by the embodiment of the present invention is different from the prior art in determining the shape of the window used in windowing and the number of required windowing. .
  • Other calculation methods can refer to the manner provided in the prior art.
  • the time domain envelope is solved by using different window lengths and/or window shapes under different conditions, and the introduction of the time domain envelope difference is introduced.
  • the effect of energy discontinuity can improve the performance of the output signal.
  • FIG. 4 is a schematic diagram of processing an audio signal according to another embodiment of the present invention.
  • the number of time domain envelopes to be calculated according to the N+1 frame is divided into M subframes, M is a positive integer.
  • the value of M can be 3, 4, 5, 8, and the like. There are no restrictions here.
  • the window of the foremost subframe among the M subframes and the last subframe of the M subframes are windowed by an asymmetric window.
  • the asymmetric window used for windowing the foremost end frame and the endmost sub-frame is different.
  • the window length of the asymmetric window used by the frontmost terminal frame and the window length of the asymmetric window used by the endmost subframe may be the same or different.
  • subframes other than the foremost subframe and the last subframe of the M subframes of the N+1 frame are performed by using asymmetric windows of the same shape. Add window.
  • the frame length of the N+1th frame is 80 samples and the sampling rate is 4 kHz, eight time domain envelopes can be solved.
  • the number N of time domain envelopes may be predetermined based on other information of the N+1 frame.
  • the following is an example of an implementation that determines the number N of time domain envelopes:
  • the second threshold can be 70 samples.
  • the low-band signal of the (N+1)th frame can be obtained when the signal of the (N+1)th frame is decomposed, and the method used for signal decomposition and the method of solving the pitch period of the low-band signal can be used.
  • the method used for signal decomposition and the method of solving the pitch period of the low-band signal can be used. There is any one of the techniques, and no specific limitation is imposed here.
  • the asymmetric window when windowing the foremost sub-frame and the endmost sub-frame with an asymmetric window, is determined according to the length of the forward buffer.
  • the window length of the asymmetric window used for windowing is The window length of the symmetrical window can be 20 samples.
  • the first threshold is obtained by dividing the frame length by the number of envelopes, in this example the first threshold is equal to 10. If the length of the current buffer is less than 10 samples, the alias of the window used by the 8th subframe (ie, the last subframe) and the 1st subframe (ie, the frontmost subframe) Equal to the length of the forward buffer.
  • the length of the left side of the window used by the 8th subframe and the left side of the window used by the 1st subframe may be equal to the other side (for example, the window used by the 1st subframe)
  • the length of the window (10 samples) on the right side or the left side of the window used in the 8th sub-frame can also be set according to experience (for example, the same length as when the forward buffer is less than 10 samples) ).
  • the window length of the asymmetric window used for windowing is The window length of the symmetrical window can be Take 40 samples.
  • the first threshold is obtained by dividing the frame length by the number of envelopes, in this example the first threshold is equal to 20.
  • the time domain energy of the pre-processed original high-band signal and the predicted high-band signal in each sub-frame or the average of the amplitude of each sample in the sub-frame is calculated.
  • the method for signal processing provided by the embodiment of the present invention is different from the prior art in determining the shape of the window used in windowing and the number of required windowing. .
  • Other calculation methods can refer to the manner provided in the prior art.
  • FIG. 5 is a schematic diagram of processing an audio signal according to another embodiment of the present invention.
  • FIG. 5 after obtaining an original audio signal at the encoding end, first decomposing the original audio signal to obtain a low original audio signal. With a signal and a high-band signal, the low-band signal is then encoded by an existing algorithm to obtain a low-band code stream. At the same time, in the low-band coding process, a low-band excitation signal is obtained, and the low-band excitation signal is pre-processed. Processing; for the high-band signal of the original audio signal, first pre-processing, and then LP analysis to obtain the LP coefficient, and quantize the LP coefficient.
  • the preprocessed low band excitation signal is then passed through an LP synthesis filter (the filter coefficients are quantized LP coefficients) to obtain a predicted high band signal.
  • the time domain envelope of the high band signal is calculated and quantized according to the preprocessed high band signal and the predicted high band signal, and finally the coded code stream is output.
  • the N+1th frame is divided into M subframes according to the number of time domain envelopes to be calculated, and M is a positive integer.
  • M can be 3, 4, 5, 8, and the like. There are no restrictions here.
  • the window of the foremost subframe among the M subframes and the last subframe of the M subframes are windowed by an asymmetric window.
  • the topmost subframe among the M subframes of the N+1 frame is a subframe that overlaps with the signal of the previous frame (N frame); the last subframe is the next frame (N+2 frame, in the figure)
  • the signal not shown) has a sub-frame of overlapping portions.
  • the frontmost subframe is the leftmost subframe of the N+1 frame
  • the last subframe is the rightmost subframe of the N+1 frame. Understandable Yes, the leftmost and rightmost are only a specific example in conjunction with FIG. 3, and are not intended to limit the embodiments of the present invention.
  • the division of the actual neutron frame is such that there is no directional restriction of the leftmost and rightmost.
  • the asymmetric window used for windowing the most front-end subframe and the last-end subframe may be identical or different. There are no restrictions here. In a possible implementation, the window length of the asymmetric window used by the frontmost terminal frame is the same as the window length of the asymmetric window used by the endmost subframe.
  • the foremost subframe of the M subframes and the last subframe of the M subframes are windowed by an asymmetric window, where the front end of the M subframes
  • the asymmetric window used in the subframe is different from the shape of the asymmetric window used in the last subframe of the M subframes, and one asymmetric window is rotated 180 degrees in the horizontal direction to coincide with the other asymmetric window.
  • the window length of the asymmetric window used by the frontmost terminal frame is the same as the window length of the asymmetric window used by the endmost subframe.
  • subframes other than the topmost subframe and the last subframe of the M subframes of the N+1 frame are windowed by using a symmetric window.
  • the window length of the symmetrical window is different from the window length of the asymmetric window. For example, for a signal with a frame length of 20 ms (80 samples) and a sampling rate of 4 kHz: if the forward buffer is 5 samples, and 4 time domain envelopes are solved, the window of this embodiment is used, and the window lengths at both ends are For 30 samples, the number of samples for the two consecutive frames is 5 samples, the middle two windows are 50 samples, and the 25 samples are aliased.
  • subframes other than the topmost subframe and the last subframe of the M subframes of the N+1 frame are windowed by using a symmetric window.
  • the window length of the asymmetric window used for windowing the topmost and lastmost subframes is equal to the window length of the symmetric window employed for the other subframes. It can be understood that in another possible manner, the window length of the asymmetric window and the window length of the symmetric window may also be unequal.
  • the frame length of the N+1th frame is 80 samples and the sampling rate is 4 kHz, eight time domain envelopes can be solved.
  • the number N of time domain envelopes may be predetermined based on other information of the N+1 frame.
  • the following is an example of an implementation that determines the number N of time domain envelopes:
  • the second threshold can be 70 samples.
  • the low-band signal of the (N+1)th frame can be obtained when the signal of the (N+1)th frame is decomposed, and the method used for signal decomposition and the method of solving the pitch period of the low-band signal can be used.
  • the method used for signal decomposition and the method of solving the pitch period of the low-band signal can be used. There is any one of the techniques, and no specific limitation is imposed here.
  • the asymmetric window when windowing the foremost sub-frame and the endmost sub-frame with an asymmetric window, is determined according to the length of the forward buffer.
  • the window length of the asymmetric window used for windowing is The window length of the symmetrical window can be 20 samples.
  • the first threshold is obtained by dividing the frame length by the number of envelopes, in this example the first threshold is equal to 10. If the length of the current buffer is less than 10 samples, the window used by the 8th subframe (ie, the last subframe) and the alias of the window used by the 1st subframe (ie, the topmost subframe) Equal to the length of the forward buffer.
  • the length of the left side of the window adopted by the 8th subframe and the left side of the window adopted by the 1st subframe may be equal to the other side (for example, the first subframe is adopted)
  • the window length (10 samples) on the right side of the window or the left side of the window used in the eighth sub-frame can also be set according to experience (for example, the same as when the forward buffer is less than 10 samples) length).
  • the window length of the asymmetric window used for windowing is The window length of the symmetrical window can be 40 samples.
  • the first threshold is obtained by dividing the frame length by the number of envelopes, in this example the first threshold is equal to 20.
  • the time domain energy of the pre-processed original high-band signal and the predicted high-band signal in each sub-frame or the average of the amplitude of each sample in the sub-frame is calculated.
  • the method for signal processing provided by the embodiment of the present invention is different from the prior art in determining the shape of the window used in windowing and the number of required windowing. .
  • Other calculation methods can refer to the manner provided in the prior art.
  • a method for processing a time domain envelope of an audio signal according to an embodiment of the present invention in different strips
  • the time domain envelope is solved by using different window lengths and/or window shapes to reduce the influence of energy discontinuity introduced due to too much difference in the time domain envelope, and the performance of the output signal can be improved.
  • the time domain envelope processing method for the audio signal obtained by this embodiment obtains a high band signal of the audio frame according to the received audio frame signal, and then the high band signal of the audio frame according to the predetermined number of time envelopes M Divided into M subframes, and finally calculates the time domain envelope of each subframe. Therefore, the problem that the lookahead is very short and the problem of solving the excessive time domain envelope caused by the good aliasing between the sub-frames is effectively avoided, thereby avoiding the energy introduced by the excessive solution of the time domain envelope for some signals. Discontinuous problems while reducing computational complexity.
  • FIG. 6 is a flowchart of Embodiment 2 of a time domain envelope processing method for an audio signal according to the present invention. As shown in FIG. 6, the method in this embodiment may include:
  • the determining the number of time domain envelopes M to be processed specifically includes:
  • M is equal to M1
  • M1 is greater than M2
  • M1 and M2 are positive integers
  • the preset threshold is based on The sampling rate is determined.
  • the stationary state means that the mean value of the energy or amplitude of the time domain signal does not change much within a certain period of time, or the deviation of the time domain signal within a certain time is less than a given threshold.
  • window processing when window processing is performed on each subframe, it is not limited to which windowing method is used for windowing processing.
  • the time domain envelope processing method of the audio signal provided by the embodiment can solve the energy caused by the excessive time domain envelope of the signal under certain conditions by solving different time domain envelopes according to different conditions. Discontinuity, and thus the resulting auditory quality is degraded, and at the same time, the average complexity of the algorithm can be effectively reduced.
  • the embodiment of the present invention further provides a time domain envelope processing device for an audio signal, which can be used to execute some of the methods shown in FIG. 1 to FIG. 5, and can also be used for other processing processes for solving a time domain envelope using the same principle. in.
  • a time domain envelope processing apparatus for the audio signal provided by the embodiment of the present invention is described in detail below with reference to the accompanying drawings.
  • FIG. 7 is a schematic structural diagram of a time domain envelope processing apparatus according to an embodiment of the present invention.
  • the time domain envelope processing apparatus 70 of the present embodiment includes: a highband signal acquisition module 71, configured to receive according to The current frame signal is obtained as a high-band signal of the current frame signal; the subframe acquisition module 72 is configured to divide the high-band signal of the current frame into M subframes according to the predetermined number of time-domain envelopes M, where M is greater than or equal to An integer of 2; a time domain envelope obtaining module 73, configured to calculate a time domain envelope of each subframe; wherein the time domain envelope obtaining module 73 is specifically configured to: use an asymmetric window to the forefront of the M subframes The subframe and the last subframe of the M subframes are windowed; and the subframes other than the foremost subframe and the last subframe of the M subframes are windowed.
  • the time domain envelope obtaining module 73 is further configured to:
  • the asymmetric window is determined based on the length of the forward buffer of the high band signal of the current frame signal and the number M of time domain envelopes.
  • the time domain envelope obtaining module 73 is specifically configured to:
  • An asymmetrical window is used to window the topmost subframe of the M subframes and the last subframe of the M subframes, except for the subframes of the M subframes except the frontmost subframe and the last subframe.
  • Subframes are windowed with symmetric windows; or,
  • the topmost subframe of the M subframes and the last of the M subframes The sub-frame is windowed, and the sub-frames other than the foremost sub-frame and the last-most sub-frame of the M sub-frames are windowed by an asymmetric window.
  • a window length of an asymmetric window and a window used for windowing a subframe other than the frontmost subframe and the last subframe of the M subframes The window length is the same.
  • the time domain envelope obtaining module 73 is further configured to: obtain a pitch period of the low band signal of the current frame signal according to the current frame signal;
  • the time domain envelope of each subframe is smoothed.
  • the time domain envelope is smoothed, and the time domain envelopes of the two adjacent subframes are weighted, and the weighted time domain envelope is used as the time domain envelope of the two subframes.
  • the pitch period of the low band signal is greater than a given threshold (greater than 70 samples, at this time, the low band signal
  • the sampling rate is 12.8 kHz sampling
  • the decoded high-band signal time domain envelope is smoothed, otherwise the time domain envelope is kept unchanged. Smoothing can be:
  • Env[N] 0.5*(env[N-1]+env[N]).
  • env[] is the time domain envelope.
  • the time domain envelope processing apparatus 70 further includes: a determining module 74, configured to determine the number of time domain envelopes M according to one of the following ways:
  • M1 and M2 are positive integers, and M2>M1.
  • the number M of time domain envelopes to be predetermined may be determined based on overall algorithm requirements and empirical values.
  • the number of time domain envelopes M is, for example, the encoder according to the whole
  • the algorithm or empirical value is determined and will not change after the determination. For example, for an input signal of 20 ms one frame, if the input signal is relatively stable, four or two time domain envelopes are solved, but for some non-stationary signals, more than eight time domain envelopes need to be solved.
  • the original audio signal is first decomposed to obtain a low-band signal and a high-band signal of the original audio signal, and then the low-band signal is encoded by an existing algorithm. Obtaining a low-band code stream.
  • a low-band excitation signal is obtained, and the low-band excitation signal is preprocessed; for the original high-band signal of the original audio signal, the pre-processing is performed first, and then The LP analysis yields an LP coefficient that is quantized.
  • the preprocessed low band excitation signal is then passed through an LP synthesis filter (the filter coefficients are quantized LP coefficients) to obtain a predicted high band signal.
  • the time domain envelope of the high band signal is calculated and quantized according to the preprocessed high band signal and the predicted high band signal, and finally the coded code stream is output.
  • the device of this embodiment can be used to implement the technical solution of the method embodiment shown in FIG. 2 to FIG. 5, and the implementation principle is similar.
  • the original audio signal is first decomposed to obtain a low-band signal and a high-band signal of the original audio signal, and then the low-band signal is passed through an existing algorithm.
  • the coded low-band code stream is obtained.
  • the low-band excitation signal is obtained, and the low-band excitation signal is pre-processed; for the original high-band signal of the original audio signal, the pre-processing is performed first, and then LP analysis is performed to obtain LP coefficients, and the LP coefficients are quantized.
  • the preprocessed low band excitation signal is then passed through an LP synthesis filter (the filter coefficients are quantized LP coefficients) to obtain a predicted high band signal.
  • the time domain envelope of the high band signal is calculated and quantized according to the preprocessed high band signal and the predicted high band signal, and finally the coded code stream is output.
  • the N+1th frame is divided into M subframes according to the number of time domain envelopes to be calculated, and M is a positive integer.
  • M can be 3, 4, 5, 8, and the like. There are no restrictions here.
  • the window of the foremost subframe among the M subframes and the last subframe of the M subframes are windowed by an asymmetric window.
  • the topmost subframe among the M subframes of the N+1 frame is heavier than the signal of the previous frame (N frame)
  • the sub-frame of the overlap portion; the last-most sub-frame is a sub-frame having an overlapping portion with the signal of the next frame (N + 2 frames, not shown).
  • the frontmost subframe is the leftmost subframe of the N+1 frame
  • the last subframe is the rightmost subframe of the N+1 frame. It is to be understood that the leftmost and rightmost are only a specific example, and are not intended to limit the embodiments of the present invention.
  • the division of the actual neutron frame is such that there is no directional restriction of the leftmost and rightmost.
  • the asymmetric window used for windowing the most front-end subframe and the last-end subframe may be identical or different. There are no restrictions here. In a possible implementation, the window length of the asymmetric window used by the frontmost terminal frame is the same as the window length of the asymmetric window used by the endmost subframe.
  • a subframe other than the topmost subframe and the last subframe of the M subframes of the N+1 frame is windowed by a symmetric window.
  • the window length of the asymmetric window used for windowing the topmost and lastmost subframes is equal to the window length of the symmetric window employed for the other subframes. It can be understood that in another possible manner, the window length of the asymmetric window and the window length of the symmetric window may also be unequal.
  • the frame length of the N+1th frame is 80 samples and the sampling rate is 4 kHz, eight time domain envelopes can be solved.
  • the number N of time domain envelopes may be predetermined based on other information of the N+1 frame.
  • the following is an example of an implementation that determines the number N of time domain envelopes:
  • the second threshold can be 70 samples.
  • the low-band signal of the (N+1)th frame can be obtained when the signal of the (N+1)th frame is decomposed, and the method used for signal decomposition and the pitch period for solving the low-band signal can be any one of the prior art. In this way, no specific restrictions are imposed here.
  • the asymmetrical window is used to match the topmost subframe and the endmost When the sub-frame is windowed, the asymmetric window is determined according to the length of the forward buffer.
  • the window length of the asymmetric window used for windowing is The window length of the symmetrical window can be 20 samples.
  • the first threshold is obtained by dividing the frame length by the number of envelopes, in this example the first threshold is equal to 10. If the length of the current buffer is less than 10 samples, the window used by the 8th subframe (ie, the last subframe) and the alias of the window used by the 1st subframe (ie, the topmost subframe) Equal to the length of the forward buffer.
  • the length of the left side of the window adopted by the 8th subframe and the left side of the window adopted by the 1st subframe may be equal to the other side (for example, the first subframe is adopted)
  • the window length (10 samples) on the right side of the window or the left side of the window used in the eighth sub-frame can also be set according to experience (for example, the same as when the forward buffer is less than 10 samples) length).
  • the window length of the asymmetric window used for windowing is The window length of the symmetrical window can be 40 samples.
  • the first threshold is obtained by dividing the frame length by the number of envelopes, in this example the first threshold is equal to 20.
  • the time domain energy of the pre-processed original high-band signal and the predicted high-band signal in each sub-frame or the average of the amplitude of each sample in the sub-frame is calculated.
  • the method for signal processing provided by the embodiment of the present invention is different from the prior art in determining the shape of the window used in windowing and the number of required windowing. .
  • Other calculation methods can refer to the manner provided in the prior art.
  • the time domain envelope processing device of the audio signal provided by the embodiment effectively solves the energy caused by the excessive time domain envelope of the signal under certain conditions by solving different time domain envelopes according to different conditions. Discontinuity, and thus the resulting auditory quality is degraded, and at the same time, the average complexity of the algorithm can be effectively reduced.
  • FIG. 8 is a schematic structural diagram of an encoder according to an embodiment of the present invention. As shown in FIG. 8, the encoder 80 is specifically configured to:
  • the time domain envelope for calculating the predicted highband signal includes:
  • the quantized time domain envelope is encoded.
  • encoder 80 can be used to perform any of the method embodiments described above.
  • the time domain envelope processing device 70 of any of the embodiments may also be included.
  • the specific encoder 80 reference may be made to the foregoing method and device embodiments, and details are not described herein.
  • the aforementioned program can be stored in a computer readable storage medium.
  • the program when executed, performs the steps including the foregoing method embodiments; and the foregoing storage medium includes various media that can store program codes, such as a ROM, a RAM, a magnetic disk, or an optical disk.

Landscapes

  • Engineering & Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • Computational Linguistics (AREA)
  • Signal Processing (AREA)
  • Health & Medical Sciences (AREA)
  • Audiology, Speech & Language Pathology (AREA)
  • Human Computer Interaction (AREA)
  • Acoustics & Sound (AREA)
  • Multimedia (AREA)
  • Spectroscopy & Molecular Physics (AREA)
  • Quality & Reliability (AREA)
  • Compression, Expansion, Code Conversion, And Decoders (AREA)

Abstract

一种音频信号的时域包络处理方法及装置、编码器,在求解多个时域包络能够很好的保持信号能量的连续,同时降低了计算时域包络的复杂度。该方法包括:根据接收到的当前帧音频信号,得到当前帧音频信号的高带信号(S21);根据预先确定的时域包络个数M将当前帧音频信号的高带信号分成M个子帧,其中,M为大于等于2的整数(S22);计算每一个子帧的时域包络(S23);采用非对称窗对M个子帧中的最前端的子帧和M个子帧中的最末端的子帧进行加窗;对M个子帧中除最前端的子帧和最末端的子帧之外的子帧进行加窗。

Description

一种音频信号的时域包络处理方法及装置、编码器 技术领域
本发明实施例涉及通信技术领域,尤其涉及一种音频信号的时域包络处理方法及装置、编码器。
背景技术
随着语音频压缩技术的高速发展,各种语音频编码算法也相继出现。在语音频编码算法的处理过程中,需要计算时域包络,现有的计算并量化时域包络的过程为:根据事先设定好的计算时域包络的个数M,M为正整数,将预处理后的原始高带信号和预测的高带信号分别分成M个子帧,对子帧进行加窗,然后计算各个子帧内预处理后的原始高带信号和预测的高带信号的能量或幅度比。其中,事先设定好的计算时域包络的个数M是根据前向缓存(lookahead buffer)的长度来确定。前向缓存是当前帧为了计算一些参数的需要,将输入信号的最后某些样点缓存不用,在下一帧计算参数时使用,当前帧使用的是前一帧缓存的样点。缓存的这些样点即为前向缓存,缓存的样点的个数即为前向缓存的长度。
上述对时域包络的处理过程存在的问题是:在求解时域包络时,利用的都是对称窗,同时为了保证子帧间和帧间的混叠,根据前向缓存(lookahead)的长度计算了多个时域包络。但在计算时域包络时,如果信号的时域分辨率太高,会造成帧内能量的不连续,从而引入很差的听觉感受。
发明内容
本发明实施例提供一种音频信号的时域包络处理方法及装置、编码器,可解决在计算时域包络时造成的帧内能量的不连续的问题。
第一方面,本发明实施例提供一种音频信号的时域包络处理方法,包括:
根据接收到的当前帧信号,得到所述当前帧信号的高带信号;
根据预先确定的时域包络个数M将所述当前帧的高带信号分成M个子帧,其中,M为大于等于2的整数;
计算每一个所述子帧的时域包络;
其中,所述计算每一个所述子帧的时域包络包括:
采用非对称窗对所述M个子帧中的最前端的子帧和所述M个子帧中的最末端的子帧进行加窗;
对所述M个子帧中除所述最前端的子帧和所述最末端的子帧之外的子帧进行加窗。
根据本发明实施例提供的音频信号的时域包络的处理方法,在不同的条件下采用不同的窗长度和/或窗形状求解时域包络,减少因为时域包络差别太大引入的能量不连续的影响,能够提升输出信号的性能。
在第一方面的第一种可能的实施方式中,在采用非对称窗对所述M个子帧中的最前端的子帧和所述M个子帧中的最末端的子帧进行加窗之前,所述方法还包括:
根据所述当前帧信号的高带信号的前向缓存的长度确定所述非对称窗;或者,
根据所述当前帧信号的高带信号的前向缓存的长度和所述时域包络个数M确定所述非对称窗。
结合第一方面或第一方面的第一种可能的实施方式,在第一方面的第二种可能的实施方式中,所述对所述M个子帧中除所述最前端的子帧和所述最末端的子帧之外的子帧进行加窗,包括:
对所述M个子帧中除最前端的子帧和所述最末端的子帧之外的子帧采用对称窗进行加窗;或者,
对所述M个子帧中除最前端的子帧和所述最末端的子帧之外的子帧采用非对称窗进行加窗。
结合第一方面,在第一方面的第三种可能的实施方式中,所述非对称窗的窗长与对所述M个子帧中除所述最前端的子帧和所述最末端的子帧之外的子帧进行加窗所采用的窗的窗长相同。
结合第一方面的第一种可能的实施方式至第一方面的第三种可能的 实施方式任意之一所述的方法,在第一方面的第四种可能的实施方式中,所述根据所述当前帧音频信号的高带信号的前向缓存的长度确定非对称窗,包括:
当所述当前帧信号的高带信号的前向缓存的长度小于第一阈值时,根据当前帧的前一帧信号的高带信号和所述当前帧信号的高带信号的前向缓存的长度确定所述非对称窗,其中,所述当前帧的前一帧信号的高带信号的最末端子帧采用的非对称窗和所述当前帧信号的高带信号的最前端子帧采用的非对称窗的混叠部分等于所述当前帧信号的高带信号的前向缓存的长度,所述第一阈值等于所述当前帧的高带信号的帧长除以M。
结合第一方面的第一种可能的实施方式至第一方面的第三种可能的实施方式任意之一所述的方法,在第一方面的第五种可能的实施方式中,所述根据所述当前帧信号的高带信号的前向缓存的长度确定非对称窗,包括:
当所述当前帧信号的高带信号的前向缓存的长度大于第一阈值时,根据所述当前帧的前一帧信号的高带信号和所述当前帧信号的高带信号的前向缓存的长度确定所述非对称窗,其中,所述当前帧的前一帧信号的高带信号的最末端子帧采用的非对称窗和所述当前帧信号的高带信号的最前端子帧采用的非对称窗的混叠部分等于所述第一阈值,所述第一阈值等于所述当前帧的高带信号的帧长除以M。
结合第一方面至第一方面的第五种可能的实施方式任意之一所述的方法,在第一方面的第六种可能的实施方式中,根据下列之一方式确定所述时域包络个数M:
根据所述当前帧信号得到所述当前帧信号的低带信号,当所述当前帧信号的低带信号的基音周期大于第二阈值时,M=M1;或者,
根据所述当前帧信号得到所述当前帧信号的低带信号,当所述当前帧信号的低带信号的基音周期不大于第二阈值时,M=M2;
其中,M1,M2均为正整数,且M2>M1。
结合第一方面至第一方面的第五种可能的实施方式任意之一所述的方法,在第一方面的第七种可能的实施方式中,所述方法还包括:
根据所述当前帧信号得到所述当前帧信号的低带信号的基音周期;
当所述当前帧信号的类型与所述当前帧的前一帧信号的类型相同,且所述当前帧的低带信号的基音周期大于第三阈值时,对每一个所述子帧的时域包络进行平滑处理。
第二方面,本发明实施例提供一种音频信号的时域包络处理装置,包括:
高带信号获取模块,用于根据接收到的当前帧信号,得到所述当前帧信号的高带信号;
子帧获取模块,用于根据预先确定的时域包络个数M将所述当前帧的高带信号分成M个子帧,其中,M为大于等于2的整数;
时域包络获取模块,用于计算每一个所述子帧的时域包络;
其中,所述时域包络获取模块具体用于:
采用非对称窗对所述M个子帧中的最前端的子帧和所述M个子帧中的最末端的子帧进行加窗;
对所述M个子帧中除所述最前端的子帧和所述最末端的子帧之外的子帧进行加窗。
根据本发明实施例提供的音频信号的时域包络的处理装置,在不同的条件下采用不同的窗长度和/或窗形状求解时域包络,减少因为时域包络差别太大引入的能量不连续的影响,能够提升输出信号的性能。
在第二方面的第一种可能的实施方式中,所述时域包络获取模块还用于:
根据所述当前帧信号的高带信号的前向缓存的长度确定所述非对称窗;或者,
根据所述当前帧信号的高带信号的前向缓存的长度和所述时域包络个数M确定所述非对称窗。
结合第二方面的实施方式,在第二方面的第二种可能的实施方式中,所述时域包络获取模块具体用于:
采用非对称窗对所述M个子帧中的最前端的子帧和所述M个子帧中的最末端的子帧进行加窗,对所述M个子帧中除最前端的子帧和所述最末端的子帧之外的子帧采用对称窗进行加窗;或者,
采用非对称窗对所述M个子帧中的最前端的子帧和所述M个子帧中的最末端的子帧进行加窗,对所述M个子帧中除最前端的子帧和所述最末端的子帧之外的子帧采用非对称窗进行加窗。
结合第二方面的实施方式,在第二方面的第三种可能的实施方式中,所述非对称窗的窗长与对所述M个子帧中除所述最前端的子帧和所述最末端的子帧之外的子帧进行加窗所采用的窗的窗长相同。
结合第二方面至第二方面的第三种可能的实施方式任意之一所述的装置,在第二方面的第四种可能的实施方式中,还包括:确定模块,用于根据下列之一方式确定所述时域包络个数M:
根据所述当前帧信号得到所述当前帧信号的低带信号,当所述当前帧信号的低带信号的基音周期大于第二阈值时,M=M1;或者,
根据所述当前帧信号得到所述当前帧信号的低带信号,当所述当前帧信号的低带信号的基音周期不大于第二阈值时,M=M2;
其中,M1,M2均为正整数,且M2>M1。
本发明第三方面的实施例公开了一种编码器,所述编码器具体用于:
用于根据接收到的当前帧信号,得到所述当前帧信号的低带信号和所述当前帧信号的高带信号;
对所述当前帧信号的低带信号进行编码,得到低带编码的激励信号;
对所述当前帧信号的高带信号进行线性预测,得到线性预测系数;
量化所述线性预测系数,得到量化后的线性预测系数;
根据所述低带编码的激励信号和所述量化后的线性预测系数得到预测的高带信号;
计算及量化所述预测的高带信号的时域包络;
其中,所述计算所述预测的高带信号的时域包络包括:
根据预先确定的时域包络个数M将所述预测的高带信号分成M个子帧,其中,M为大于等于2的整数,
采用非对称窗对所述M个子帧中的最前端的子帧和所述M个子帧中的最末端的子帧进行加窗,
对所述M个子帧中除所述最前端的子帧和所述最末端的子帧之外的子帧进行加窗;
对量化后的时域包络进行编码。
根据本发明实施例提供的编码器,在不同的条件下采用不同的窗长度和/或窗形状求解时域包络,减少因为时域包络差别太大引入的能量不连续的影响,能够提升输出信号的性能。
附图说明
为了更清楚地说明本发明实施例中的技术方案,下面将对实施例描述中所需要使用的附图作一简单地介绍,显而易见地,下面描述中的附图是本发明的一些实施例,对于本领域普通技术人员来讲,在不付出创造性劳动性的前提下,还可以根据这些附图获得其他的附图。
图1为一种对音频信号进行编码的过程示意图;
图2为本发明音频信号的时域包络处理方法实施例一的流程图;
图3为本发明实施例中对音频信号进行处理的示意图;
图4为本发明另一实施例的对音频信号进行处理的示意图;
图5为本发明另一实施例的对音频信号进行处理的示意图;
图6为本发明音频信号的时域包络处理方法实施例二的流程图;
图7为本发明实施例的时域包络处理装置的结构示意图;
图8为本发明实施例的编码器的结构示意图。
具体实施方式
为使本发明实施例的目的、技术方案和优点更加清楚,下面将结合本发明实施例中的附图,对本发明实施例中的技术方案进行清楚、完整地描述,显然,所描述的实施例是本发明一部分实施例,而不是全部的实施例。基于本发明中的实施例,本领域普通技术人员在没有作出创造性劳动前提下所获得的所有其他实施例,都属于本发明保护的范围。
图1为一种对语音频信号进行编码的过程示意图,如图1所示,在编码端,在获得原始音频信号后,首先对原始音频信号进行信号分解,得到原始音频信号的低带信号和高带信号,接着对低带信号通过已有算法进行编码得 到低带的码流,已有算法(例如代数码本激励线性预测编码(Algebraic Code Excited Linear Prediction,简称:ACELP),或码本激励线性预测编码(Code Excited Linear Prediction,简称:CELP等算法),同时,在进行低带编码过程中,得到低带的激励信号,并对低带激励信号进行预处理;对于原始音频信号的高带信号,首先进行预处理,然后做线性预测(Linear prediction,以下简称:LP)分析得到LP系数,量化该LP系数。接着将预处理后的低带激励信号通过LP合成滤波器(滤波器系数为量化后的LP系数)得到预测的高带信号。根据预处理后的高带信号和预测的高带信号,计算及量化高带信号的时域包络,最后输出编码码流(MUX)。计算并量化高带信号的时域包络的过程为:根据事先设定好的时域包络的个数N,将预处理后的高带信号和预测的高带信号分别分成N个子帧,对每一个子帧进行加窗,然后计算预处理后的原始高带信号每一个子帧和预测的高带信号的相对应的每一个子帧的时域能量或子帧内每个样点幅度的平均值。其中,事先设定好的时域包络的个数N是根据前向缓存(lookahead)的长度来确定的,N为正整数。
本发明实施例提供一种音频信号的时域包络处理方法,主要用于图1中所示的计算及量化时域包络的步骤,还可以用于其它采用同样原理的求解时域包络的处理流程中。下面结合附图详细说明本发明实施例提供的音频信号的时域包络处理方法。
图2为本发明音频信号的时域包络处理方法实施例一的流程图,如图2所示,本实施例的方法包括:
S21、根据接收到的当前帧信号,得到当前帧信号的高带信号。
当前帧信号即可以是语音信号,也可以是音乐信号,还可能是噪音信号,在此不做具体的限制。
S22、根据预先确定的时域包络个数M将当前帧的高带信号分成M个子帧,其中,M为大于等于2的整数。
其中,具体来说,要预先确定的时域包络个数M可以是根据整体算法要求和经验值确定。时域包络个数M例如是编码器事先根据整体算法或经验值确定,确定后不会改变。例如一般对20ms一帧的输入信号,如果输入信号相对平稳,求解4个或者2个时域包络,但对一些非平稳信号,需要求解更多如8个时域包络。
S23、计算每一个子帧的时域包络。
其中,计算每一个子帧的时域包络包括:
采用非对称窗对M个子帧中的最前端的子帧和M个子帧中的最末端的子帧进行加窗。
对M个子帧中除最前端的子帧和最末端的子帧之外的子帧进行加窗。
进一步地,在采用非对称窗对M个子帧中的最前端的子帧和M个子帧中的最末端的子帧进行加窗之前,本实施例的方法还可以包括:
根据当前帧信号的高带信号的前向缓存的长度确定非对称窗;或者,
根据当前帧信号的高带信号的前向缓存的长度和时域包络个数M确定非对称窗。
其中,对M个子帧中除最前端的子帧和最末端的子帧之外的子帧进行加窗,具体可以包括:
对M个子帧中除最前端的子帧和最末端的子帧之外的子帧采用对称窗进行加窗;或者,
对M个子帧中除最前端的子帧和最末端的子帧之外的子帧采用非对称窗进行加窗。
其中,在一种可能的实施方式中,对最前端子帧和最末端子帧加窗使用的非对称窗的窗长与对M个子帧中除最前端的子帧和最末端的子帧之外的子帧进行加窗所采用的窗的窗长相同。
在上述实施例中,作为一种可实施的方式,根据当前帧音频信号的高带信号的前向缓存的长度确定非对称窗,包括:
当当前帧信号的高带信号的前向缓存的长度小于第一阈值时,根据当前帧的前一帧信号的高带信号和当前帧信号的高带信号的前向缓存的长度确定非对称窗,其中,当前帧的前一帧信号的高带信号的最末端子帧采用的非对称窗和当前帧信号的高带信号的最前端子帧采用的非对称窗的混叠部分等于当前帧信号的高带信号的前向缓存的长度,第一阈值等于当前帧的高带信号的帧长除以M。
在一种可能的实施方式中,根据当前帧信号的高带信号的前向缓存的 长度确定非对称窗,包括:
当当前帧信号的高带信号的前向缓存的长度大于第一阈值时,根据当前帧的前一帧信号的高带信号和当前帧信号的高带信号的前向缓存的长度确定非对称窗,其中,当前帧的前一帧信号的高带信号的最末端子帧采用的非对称窗和当前帧信号的高带信号的最前端子帧采用的非对称窗的混叠部分等于第一阈值,第一阈值等于当前帧的高带信号的帧长除以M。
在本发明的一种实施例中,根据下列之一方式确定时域包络个数M:
根据当前帧信号得到当前帧信号的低带信号,当当前帧信号的低带信号的基音周期大于第二阈值时,M=M1;或者,
根据当前帧信号得到当前帧信号的低带信号,当当前帧信号的低带信号的基音周期不大于第二阈值时,M=M2;
其中,M1,M2均为正整数,且M2>M1。在一种可能的方式中,M1=4,M2=8。
在上述实施例中,进一步地,本实施例的方法还可以包括:
根据当前帧信号得到当前帧信号的低带信号的基音周期;
当当前帧信号的类型与当前帧的前一帧信号的类型相同,且当前帧的低带信号的基音周期大于第三阈值时,对每一个子帧的时域包络进行平滑处理。
对时域包络做平滑处理,具体可以是:将相邻的两个子帧的时域包络加权,加权后的时域包络作为这两个子帧的时域包络。例如,当解码端连续两帧信号都是浊音信号,或者一帧是浊音信号一帧是普通信号,且低带信号的基音周期大于给定阈值(大于70个样点,此时低带信号的采样率为12.8kHz采样)时,则对解码的高带信号时域包络做平滑处理,否则保持时域包络不变。平滑处理可以为:
env[0]=0.5*(env[0]+env[1]);
env[1]=0.5*(env[0]+env[1]);
env[N-1]=0.5*(env[N-1]+env[N]);
env[N]=0.5*(env[N-1]+env[N])。
其中,env[]为时域包络。
可以理解的是,上述步骤序号只是为了帮助理解本发明实施例而做出的一种示例,而不是对本发明实施例的具体限制。在实际的处理过程中,并不需要严格的按照上述顺序的限制。例如,可以先对除最前端和最末端的子帧之外的子帧进行加窗,再对最前端和最末端的子帧进行加窗。
图3为本发明实施例中对音频信号进行处理的示意图。
如图3所示,在编码端,在获得原始音频信号后,首先对原始音频信号进行信号分解,得到原始音频信号的低带信号和高带信号,接着对低带信号通过已有算法进行编码得到低带的码流,同时,在进行低带编码过程中,得到低带的激励信号,并对低带激励信号进行预处理;对于原始音频信号的高带信号,首先进行预处理,然后做LP分析得到LP系数,量化该LP系数。接着将预处理后的低带激励信号通过LP合成滤波器(滤波器系数为量化后的LP系数)得到预测的高带信号。根据预处理后的高带信号和预测的高带信号,计算及量化高带信号的时域包络,最后输出编码码流。
除了计算及量化高带信号的时域包络的步骤之外,对于音频信号的其它步骤的处理可以参考现有技术中所采用的方法,在此不再赘述。
下面以具体对图3中所示的N+1帧的处理来描述本发明实施例中计算及量化时域包络的步骤。
如图3所示,将第N+1帧按照需要计算的时域包络的个数划分为M个子帧,M为正整数。在一种可能的实施方式中,M的值可以是3、4、5、8等。在此不做限制。
对M个子帧中的最前端的子帧和M个子帧中最末端的子帧采用非对称窗进行加窗。N+1帧的M个子帧中最前端的子帧为与前一帧(N帧)的信号有重叠部分的子帧;最末端的子帧为与后一帧(N+2帧,图中未示出)的信号有重叠部分的子帧。在一种可能的方式中,如图3所示,最前端的子帧即为N+1帧中最左端的子帧,最末端的子帧即为N+1帧中最右端的子帧。可以理解的是,最左和最右只是结合图3的一种具体示例,而不是对本发明实施例的限制。实际中子帧的划分是不存在最左、最右这种方向性限制的。
对于最前端的子帧和最末端的子帧加窗所使用的非对称窗可以完全相 同,也可以不同。在此不做限制。在一种可能的实现方式中,最前端子帧使用的非对称窗的窗长和最末端子帧所使用的非对称窗的窗长相同。
在本发明的一个实施例中,如图3所示,对N+1帧的M个子帧中除最前端的子帧和最末端的子帧之外的子帧采用对称窗进行加窗。
在本发明的一个实施例中,对于最前端的子帧和最末端的子帧加窗所采用的非对称窗的窗长与对其它子帧采用的对称窗的窗长相等。可以理解的是,在另一种可能的方式中,非对称窗的窗长和对称窗的窗长也可以不等。
在本发明的一个实施例中,当第N+1帧的帧长为80个样点,采样率为4kHz时,可以求解8个时域包络。
在一种可能的实现方式中,当第N+1帧的帧长为80个样点,采样率为4kHz时,也可以求解4个时域包络。
在本发明的一个实施例中,除了预先设定之外,还可以根据N+1帧的其它信息预先确定时域包络的个数N。下面是确定时域包络的个数N的实现方式的示例:
在一种可能实现的方式中,当第N+1帧的低带信号的基音周期大于第二阈值时,N=4;或者,当第N+1帧的低带信号的基音周期不大于第二阈值时,N=8。对于采用率为12.8kHz的低带信号,第二阈值可以为70个样点。可以理解的是,上述数值只是为了帮助理解本发明实施例而做出的一种具体举例,而不是对本发明实施例的具体限制。如图3所示,在对第N+1帧的信号进行信号分解时可以得到第N+1帧的低带信号,信号分解所采用的方法和求解低带信号的基音周期的方式可以采用现有技术中的任意一种方式,在此不做具体的限制。
可以理解的是,除了利用低带信号的基音周期以外,还可以利用信号的能量等其它参数。
在本发明的一个实施例中,在利用非对称窗对最前端的子帧和最末端的子帧进行加窗时,根据前向缓存的长度确定非对称窗。
在一种可能的实现方式中,当第N+1帧的帧长为80个样点,采样率为4kHz,求解8个时域包络时,加窗所采用的非对称窗的窗长和对称窗的窗长可以都为20个样点。利用帧长除以包络个数得到第一阈值,此示例中第一阈值等于10。则当前向缓存的长度小于10个样点时,第8个子帧(即,最末端的子 帧)采用的窗和第1个子帧(即,最前端的子帧)采用的窗的混叠部分等于前向缓存的长度。当前向缓存的长度大于等于10个样点时,第8个子帧采用的窗的右侧和第1个子帧采用的窗的左侧的长度可以等于另一侧(例如第一个子帧采用的窗的右侧或第八个子帧采用的窗的左侧)的窗长(10个样点),也可以根据经验设定一个长度(如,保持和前向缓存小于10个样点时相同的长度)。
在一种可能的实现方式中,当第N+1帧的帧长为80个样点,采样率为4kHz,求解4个时域包络时,加窗所采用的非对称窗的窗长和对称窗的窗长可以都为40个样点。利用帧长除以包络个数得到第一阈值,此示例中第一阈值等于20。
在加窗后,计算各个子帧内预处理后的原始高带信号和预测的高带信号的时域能量或子帧内每个样点幅度的平均值。具体的计算方式可参考现有技术中提供的方式,本发明实施例提供的信号处理的方法在加窗时所采用的窗的形状和所需要加窗的个数的确定方式与现有技术不同。其它的计算方式均可参考现有技术中提供的方式。
根据本发明实施例提供的音频信号的时域包络的处理方法,在不同的条件下采用不同的窗长度和/或窗形状求解时域包络,减少因为时域包络差别太大引入的能量不连续的影响,能够提升输出信号的性能。
下面以具体对图4中所示的N+1帧的处理来描述本发明另一实施例中计算及量化时域包络的步骤。
图4为本发明另一实施例的对音频信号进行处理的示意图,如图4所示,和图3所示类似,将第N+1帧按照需要计算的时域包络的个数划分为M个子帧,M为正整数。在一种可能的实施方式中,M的值可以是3、4、5、8等。在此不做限制。
对M个子帧中的最前端的子帧和M个子帧中最末端的子帧采用非对称窗进行加窗。如图4所示,对于最前端的子帧和最末端的子帧加窗所使用的非对称窗不同。在一种可能的实现方式中,最前端子帧使用的非对称窗的窗长和最末端子帧所使用的非对称窗的窗长相同,也可以不同。
在本发明的一个实施例中,如图4所示,对N+1帧的M个子帧中除最前端的子帧和最末端的子帧之外的子帧采用形状相同的非对称窗进行加窗。
在本发明的一个实施例中,当第N+1帧的帧长为80个样点,采样率为4kHz时,可以求解8个时域包络。
在一种可能的实现方式中,当第N+1帧的帧长为80个样点,采样率为4kHz时,也可以求解4个时域包络。
在本发明的一个实施例中,除了预先设定之外,还可以根据N+1帧的其它信息预先确定时域包络的个数N。下面是确定时域包络的个数N的实现方式的示例:
在一种可能实现的方式中,当第N+1帧的低带信号的基音周期大于第二阈值时,N=4;或者,当第N+1帧的低带信号的基音周期不大于第二阈值时,N=8。对于采用率为12.8kHz的低带信号,第二阈值可以为70个样点。可以理解的是,上述数值只是为了帮助理解本发明实施例而做出的一种具体举例,而不是对本发明实施例的具体限制。如图4所示,在对第N+1帧的信号进行信号分解时可以得到第N+1帧的低带信号,信号分解所采用的方法和求解低带信号的基音周期的方式可以采用现有技术中的任意一种方式,在此不做具体的限制。
可以理解的是,除了利用低带信号的基音周期以外,还可以利用信号的能量等其它参数。
在本发明的一个实施例中,在利用非对称窗对最前端的子帧和最末端的子帧进行加窗时,根据前向缓存的长度确定非对称窗。
在一种可能的实现方式中,当第N+1帧的帧长为80个样点,采样率为4kHz,求解8个时域包络时,加窗所采用的非对称窗的窗长和对称窗的窗长可以都为20个样点。利用帧长除以包络个数得到第一阈值,此示例中第一阈值等于10。则当前向缓存的长度小于10个样点时,第8个子帧采用的窗(即,最末端的子帧)和第1个子帧(即,最前端的子帧)采用的窗的混叠部分等于前向缓存的长度。当前向缓存的长度大于等于10个样点时,第8个子帧采用的窗的右侧和第1个子帧采用的窗的左侧的长度可以等于另一侧(例如第1个子帧采用的窗的右侧或第8个子帧采用的窗的左侧)的窗长(10个样点),也可以根据经验设定一个长度(如,保持和前向缓存小于10个样点时相同的长度)。
在一种可能的实现方式中,当第N+1帧的帧长为80个样点,采样率为4kHz,求解4个时域包络时,加窗所采用的非对称窗的窗长和对称窗的窗长可 以都为40个样点。利用帧长除以包络个数得到第一阈值,此示例中第一阈值等于20。
在加窗后,计算各个子帧内预处理后的原始高带信号和预测的高带信号的时域能量或子帧内每个样点幅度的平均值。具体的计算方式可参考现有技术中提供的方式,本发明实施例提供的信号处理的方法在加窗时所采用的窗的形状和所需要加窗的个数的确定方式与现有技术不同。其它的计算方式均可参考现有技术中提供的方式。
下面以具体对图5中所示的N+1帧的处理来描述本发明另一实施例中计算及量化时域包络的步骤。
图5为本发明另一实施例的对音频信号进行处理的示意图,如图5所示,在编码端,在获得原始音频信号后,首先对原始音频信号进行信号分解,得到原始音频信号的低带信号和高带信号,接着对低带信号通过已有算法进行编码得到低带的码流,同时,在进行低带编码过程中,得到低带的激励信号,并对低带激励信号进行预处理;对于原始音频信号的高带信号,首先进行预处理,然后做LP分析得到LP系数,量化该LP系数。接着将预处理后的低带激励信号通过LP合成滤波器(滤波器系数为量化后的LP系数)得到预测的高带信号。根据预处理后的高带信号和预测的高带信号,计算及量化高带信号的时域包络,最后输出编码码流。
除了计算及量化高带信号的时域包络的步骤之外,对于音频信号的其它步骤的处理可以参考现有技术中所采用的方法,在此不再赘述。
下面以具体对图5中所示的N+1帧的处理来描述本发明实施例中计算及量化时域包络的步骤。
如图5所示,将第N+1帧按照需要计算的时域包络的个数划分为M个子帧,M为正整数。在一种可能的实施方式中,M的值可以是3、4、5、8等。在此不做限制。
对M个子帧中的最前端的子帧和M个子帧中最末端的子帧采用非对称窗进行加窗。N+1帧的M个子帧中最前端的子帧为与前一帧(N帧)的信号有重叠部分的子帧;最末端的子帧为与后一帧(N+2帧,图中未示出)的信号有重叠部分的子帧。在一种可能的方式中,如图3所示,最前端的子帧即为N+1帧中最左端的子帧,最末端的子帧即为N+1帧中最右端的子帧。可以理解的 是,最左和最右只是结合图3的一种具体示例,而不是对本发明实施例的限制。实际中子帧的划分是不存在最左、最右这种方向性限制的。
对于最前端的子帧和最末端的子帧加窗所使用的非对称窗可以完全相同,也可以不同。在此不做限制。在一种可能的实现方式中,最前端子帧使用的非对称窗的窗长和最末端子帧所使用的非对称窗的窗长相同。
在本发明的一种可能实现的方式中,对M个子帧中的最前端的子帧和M个子帧中最末端的子帧采用非对称窗进行加窗,其中对M个子帧中的最前端的子帧采用的非对称窗与对M个子帧中最末端的子帧采用的非对称窗的形状不同,其中一个非对称窗以水平方向旋转180度可以与另一个非对称窗重合。在一种可能的实现方式中,最前端子帧使用的非对称窗的窗长和最末端子帧所使用的非对称窗的窗长相同。在本发明的一个实施例中,如图5所示,对N+1帧的M个子帧中除最前端的子帧和最末端的子帧之外的子帧采用对称窗进行加窗。对称窗的窗长与非对称窗的窗长不同。例如,对帧长为20ms(80个样点)采样率为4kHz的信号:如果前向缓存为5个样点,求解4个时域包络,采用本实施例的窗,两端的窗长为30个样点,连续两帧混叠时的样点数为5个样点,中间的两个窗长为50个样点,混叠25个样点。
在本发明的一个实施例中,如图5所示,对N+1帧的M个子帧中除最前端的子帧和最末端的子帧之外的子帧采用对称窗进行加窗。
在本发明的一个实施例中,对于最前端的子帧和最末端的子帧加窗所采用的非对称窗的窗长与对其它子帧采用的对称窗的窗长相等。可以理解的是,在另一种可能的方式中,非对称窗的窗长和对称窗的窗长也可以不等。
在本发明的一个实施例中,当第N+1帧的帧长为80个样点,采样率为4kHz时,可以求解8个时域包络。
在一种可能的实现方式中,当第N+1帧的帧长为80个样点,采样率为4kHz时,也可以求解4个时域包络。
在本发明的一个实施例中,除了预先设定之外,还可以根据N+1帧的其它信息预先确定时域包络的个数N。下面是确定时域包络的个数N的实现方式的示例:
在一种可能实现的方式中,当第N+1帧的低带信号的基音周期大于第二 阈值时,N=4;或者,当第N+1帧的低带信号的基音周期不大于第二阈值时,N=8。对于采用率为12.8kHz的低带信号,第二阈值可以为70个样点。可以理解的是,上述数值只是为了帮助理解本发明实施例而做出的一种具体举例,而不是对本发明实施例的具体限制。如图3所示,在对第N+1帧的信号进行信号分解时可以得到第N+1帧的低带信号,信号分解所采用的方法和求解低带信号的基音周期的方式可以采用现有技术中的任意一种方式,在此不做具体的限制。
可以理解的是,除了利用低带信号的基音周期以外,还可以利用信号的能量等其它参数。
在本发明的一个实施例中,在利用非对称窗对最前端的子帧和最末端的子帧进行加窗时,根据前向缓存的长度确定非对称窗。
在一种可能的实现方式中,当第N+1帧的帧长为80个样点,采样率为4kHz,求解8个时域包络时,加窗所采用的非对称窗的窗长和对称窗的窗长可以都为20个样点。利用帧长除以包络个数得到第一阈值,此示例中第一阈值等于10。则当前向缓存的长度小于10个样点时,第8个子帧(即,最末端的子帧)采用的窗和第1个子帧(即,最前端的子帧)采用的窗的混叠部分等于前向缓存的长度。当前向缓存的长度大于等于10个样点时,第8个子帧采用的窗的右侧和第1个子帧采用的窗的左侧的长度可以等于另一侧(例如第一个子帧采用的窗的右侧或第八个子帧采用的窗的左侧)的窗长(10个样点),也可以根据经验设定一个长度(如,保持和前向缓存小于10个样点时相同的长度)。
在一种可能的实现方式中,当第N+1帧的帧长为80个样点,采样率为4kHz,求解4个时域包络时,加窗所采用的非对称窗的窗长和对称窗的窗长可以都为40个样点。利用帧长除以包络个数得到第一阈值,此示例中第一阈值等于20。
在加窗后,计算各个子帧内预处理后的原始高带信号和预测的高带信号的时域能量或子帧内每个样点幅度的平均值。具体的计算方式可参考现有技术中提供的方式,本发明实施例提供的信号处理的方法在加窗时所采用的窗的形状和所需要加窗的个数的确定方式与现有技术不同。其它的计算方式均可参考现有技术中提供的方式。
根据本发明实施例提供的音频信号的时域包络的处理方法,在不同的条 件下采用不同的窗长度和/或窗形状求解时域包络,减少因为时域包络差别太大引入的能量不连续的影响,能够提升输出信号的性能。
本实施例提供的音频信号的时域包络处理方法,通过根据接收到的音频帧信号得到音频帧的高带信号,然后根据预先确定的时域包络个数M将音频帧的高带信号分成M个子帧,最后计算每一个子帧的时域包络。从而有效避免了在lookahead很短,同时要保证子帧间很好的混叠引起的求解过多时域包络的问题,进而避免了对一些信号,因过多求解时域包络而引入的能量不连续的问题,同时降低了计算复杂度。
图6为本发明音频信号的时域包络处理方法实施例二的流程图,如图6所示,本实施例的方法可以包括:
S60、接收到待处理信号后,根据第一频带内时域信号的平稳状态或第二频带信号的基音周期大小,确定对待处理信号计算的时域包络个数M,第一频带为待处理信号的时域信号的频带或整个输入信号的频带,第二频带为低于给定阈值的频带或整个输入信号的频带。
其中,确定对待处理信号计算的时域包络个数M,具体包括:
当第一频带内时域信号处于平稳状态或第二频带信号的基音周期大于预设阈值时,M等于M1,否则M等于M2,M1大于M2,M1、M2都为正整数,预设阈值根据采样率确定。
平稳状态是指时域信号在一定时间内的能量或幅度的均值变化不大,或时域信号在一定时间内的偏差小于给定阈值。
例如,对帧长为20ms(80个样点)采样率为4kHz的高带信号,如果高带时域信号子帧间的能量的比值小于给定阈值(小于0.5),或低带信号的基音周期大于给定阈值(大于70个样点,此时低带信号的采样率为12.8kHz采样),则在对高带信号求解时域包络时,求解4个时域包络;否则,求解8个时域包络。
例如,对帧长为20ms(320个样点)采样率为16kHz的高带信号,如果高带时域信号子帧间的能量的比值小于给定阈值(小于0.5),或低带信号的基音周期大于给定阈值(大于70个样点,此时低带信号的采样率为12.8kHz采样),则在对高带信号求解时域包络时,求解2个时域包络;否则,求解4个时域包络。
S61、将待处理信号分成M个子帧,计算每一个子帧的时域包络。
其中,本实施例对每一个子帧进行加窗处理时,不限定采用何种加窗方式进行加窗处理。
本实施例提供的音频信号的时域包络处理方法,通过根据不同的条件求解不同个数的时域包络,有效避免了对一定条件下的信号求解过多的时域包络造成的能量不连续,进而引起的听觉质量下降,同时,可以有效降低算法的平均复杂度。
本发明实施例还提供一种音频信号的时域包络处理装置,可以用于执行图1-图5中所示部分方法,还可以用于其它采用同样原理的求解时域包络的处理流程中。下面结合附图详细说明本发明实施例提供的音频信号的时域包络处理装置的结构。
图7为本发明实施例的时域包络处理装置的结构示意图,如图7所示,本实施例的时域包络处理装置70包括:高带信号获取模块71,用于根据接收到的当前帧信号,得到当前帧信号的高带信号;子帧获取模块72,用于根据预先确定的时域包络个数M将当前帧的高带信号分成M个子帧,其中,M为大于等于2的整数;时域包络获取模块73,用于计算每一个子帧的时域包络;其中,时域包络获取模块73具体用于:采用非对称窗对M个子帧中的最前端的子帧和M个子帧中的最末端的子帧进行加窗;对M个子帧中除最前端的子帧和最末端的子帧之外的子帧进行加窗。
在本发明实施例一种可能的方式中,时域包络获取模块73还用于:
根据当前帧信号的高带信号的前向缓存的长度确定非对称窗;或者,
根据当前帧信号的高带信号的前向缓存的长度和时域包络个数M确定非对称窗。
在本发明一个实施例中,时域包络获取模块73具体用于:
采用非对称窗对M个子帧中的最前端的子帧和M个子帧中的最末端的子帧进行加窗,对M个子帧中除最前端的子帧和最末端的子帧之外的子帧采用对称窗进行加窗;或者,
采用非对称窗对M个子帧中的最前端的子帧和M个子帧中的最末端 的子帧进行加窗,对M个子帧中除最前端的子帧和最末端的子帧之外的子帧采用非对称窗进行加窗。
在本发明实施例一种可能的实现方式中,非对称窗的窗长与对M个子帧中除最前端的子帧和最末端的子帧之外的子帧进行加窗所采用的窗的窗长相同。在本发明的一个实施例中,时域包络获取模块73还用于:根据当前帧信号得到当前帧信号的低带信号的基音周期;
当当前帧信号的类型与当前帧的前一帧信号的类型相同,且当前帧的低带信号的基音周期大于第三阈值时,对每一个子帧的时域包络进行平滑处理。
对时域包络做平滑处理,具体可以是:将相邻的两个子帧的时域包络加权,加权后的时域包络作为这两个子帧的时域包络。例如,当解码端连续两帧信号都是浊音信号,或者一帧是浊音信号一帧是普通信号,且低带信号的基音周期大于给定阈值(大于70个样点,此时低带信号的采样率为12.8kHz采样)时,则对解码的高带信号时域包络做平滑处理,否则保持时域包络不变。平滑处理可以为:
env[0]=0.5*(env[0]+env[1]);
env[1]=0.5*(env[0]+env[1]);
env[N-1]=0.5*(env[N-1]+env[N]);
env[N]=0.5*(env[N-1]+env[N])。
其中,env[]为时域包络。
在本发明的一个实施例中,时域包络处理装置70还包括:确定模块74,用于根据下列之一方式确定时域包络个数M:
根据当前帧信号得到当前帧信号的低带信号,当当前帧信号的低带信号的基音周期大于第二阈值时,M=M1;或者,
根据当前帧信号得到当前帧信号的低带信号,当当前帧信号的低带信号的基音周期不大于第二阈值时,M=M2;
其中,M1,M2均为正整数,且M2>M1。
在本发明的实施例中,要预先确定的时域包络个数M可以是根据整体算法要求和经验值确定。时域包络个数M例如是编码器事先根据整体 算法或经验值确定,确定后不会改变。例如一般对20ms一帧的输入信号,如果输入信号相对平稳,求解4个或者2个时域包络,但对一些非平稳信号,需要求解更多如8个时域包络。
具体来说,首先,在编码端,在获得原始音频信号后,首先对原始音频信号进行信号分解,得到原始音频信号的低带信号和高带信号,接着对低带信号通过已有算法进行编码得到低带的码流,同时,在进行低带编码过程中,得到低带的激励信号,并对低带激励信号进行预处理;对于原始音频信号的高带信号,首先进行预处理,然后做LP分析得到LP系数,量化该LP系数。接着将预处理后的低带激励信号通过LP合成滤波器(滤波器系数为量化后的LP系数)得到预测的高带信号。根据预处理后的高带信号和预测的高带信号,计算及量化高带信号的时域包络,最后输出编码码流。
除了计算及量化高带信号的时域包络的步骤之外,对于音频信号的其它步骤的处理可以参考现有技术中所采用的方法,在此不再赘述。
本实施例的装置,可以用于执行图2-图5所示方法实施例的技术方案,其实现原理类似。
在一个具体的示例中,在编码端,在获得原始音频信号后,首先对原始音频信号进行信号分解,得到原始音频信号的低带信号和高带信号,接着对低带信号通过已有算法进行编码得到低带的码流,同时,在进行低带编码过程中,得到低带的激励信号,并对低带激励信号进行预处理;对于原始音频信号的高带信号,首先进行预处理,然后做LP分析得到LP系数,量化该LP系数。接着将预处理后的低带激励信号通过LP合成滤波器(滤波器系数为量化后的LP系数)得到预测的高带信号。根据预处理后的高带信号和预测的高带信号,计算及量化高带信号的时域包络,最后输出编码码流。
除了计算及量化高带信号的时域包络的步骤之外,对于音频信号的其它步骤的处理可以参考现有技术中所采用的方法,在此不再赘述。
将第N+1帧按照需要计算的时域包络的个数划分为M个子帧,M为正整数。在一种可能的实施方式中,M的值可以是3、4、5、8等。在此不做限制。
对M个子帧中的最前端的子帧和M个子帧中最末端的子帧采用非对称窗进行加窗。N+1帧的M个子帧中最前端的子帧为与前一帧(N帧)的信号有重 叠部分的子帧;最末端的子帧为与后一帧(N+2帧,图中未示出)的信号有重叠部分的子帧。在一种可能的方式中,最前端的子帧即为N+1帧中最左端的子帧,最末端的子帧即为N+1帧中最右端的子帧。可以理解的是,最左和最右只是一种具体示例,而不是对本发明实施例的限制。实际中子帧的划分是不存在最左、最右这种方向性限制的。
对于最前端的子帧和最末端的子帧加窗所使用的非对称窗可以完全相同,也可以不同。在此不做限制。在一种可能的实现方式中,最前端子帧使用的非对称窗的窗长和最末端子帧所使用的非对称窗的窗长相同。
在本发明的一个实施例中,对N+1帧的M个子帧中除最前端的子帧和最末端的子帧之外的子帧采用对称窗进行加窗。
在本发明的一个实施例中,对于最前端的子帧和最末端的子帧加窗所采用的非对称窗的窗长与对其它子帧采用的对称窗的窗长相等。可以理解的是,在另一种可能的方式中,非对称窗的窗长和对称窗的窗长也可以不等。
在本发明的一个实施例中,当第N+1帧的帧长为80个样点,采样率为4kHz时,可以求解8个时域包络。
在一种可能的实现方式中,当第N+1帧的帧长为80个样点,采样率为4kHz时,也可以求解4个时域包络。
在本发明的一个实施例中,除了预先设定之外,还可以根据N+1帧的其它信息预先确定时域包络的个数N。下面是确定时域包络的个数N的实现方式的示例:
在一种可能实现的方式中,当第N+1帧的低带信号的基音周期大于第二阈值时,N=4;或者,当第N+1帧的低带信号的基音周期不大于第二阈值时,N=8。对于采用率为12.8kHz的低带信号,第二阈值可以为70个样点。可以理解的是,上述数值只是为了帮助理解本发明实施例而做出的一种具体举例,而不是对本发明实施例的具体限制。在对第N+1帧的信号进行信号分解时可以得到第N+1帧的低带信号,信号分解所采用的方法和求解低带信号的基音周期的方式可以采用现有技术中的任意一种方式,在此不做具体的限制。
可以理解的是,除了利用低带信号的基音周期以外,还可以利用信号的能量等其它参数。
在本发明的一个实施例中,在利用非对称窗对最前端的子帧和最末端的 子帧进行加窗时,根据前向缓存的长度确定非对称窗。
在一种可能的实现方式中,当第N+1帧的帧长为80个样点,采样率为4kHz,求解8个时域包络时,加窗所采用的非对称窗的窗长和对称窗的窗长可以都为20个样点。利用帧长除以包络个数得到第一阈值,此示例中第一阈值等于10。则当前向缓存的长度小于10个样点时,第8个子帧(即,最末端的子帧)采用的窗和第1个子帧(即,最前端的子帧)采用的窗的混叠部分等于前向缓存的长度。当前向缓存的长度大于等于10个样点时,第8个子帧采用的窗的右侧和第1个子帧采用的窗的左侧的长度可以等于另一侧(例如第一个子帧采用的窗的右侧或第八个子帧采用的窗的左侧)的窗长(10个样点),也可以根据经验设定一个长度(如,保持和前向缓存小于10个样点时相同的长度)。
在一种可能的实现方式中,当第N+1帧的帧长为80个样点,采样率为4kHz,求解4个时域包络时,加窗所采用的非对称窗的窗长和对称窗的窗长可以都为40个样点。利用帧长除以包络个数得到第一阈值,此示例中第一阈值等于20。
在加窗后,计算各个子帧内预处理后的原始高带信号和预测的高带信号的时域能量或子帧内每个样点幅度的平均值。具体的计算方式可参考现有技术中提供的方式,本发明实施例提供的信号处理的方法在加窗时所采用的窗的形状和所需要加窗的个数的确定方式与现有技术不同。其它的计算方式均可参考现有技术中提供的方式。
本实施例提供的音频信号的时域包络处理装置,通过根据不同的条件求解不同个数的时域包络,有效避免了对一定条件下的信号求解过多的时域包络造成的能量不连续,进而引起的听觉质量下降,同时,可以有效降低算法的平均复杂度。
下面结合图8描述本发明实施例的一种编码器80,图8为本发明实施例的编码器的结构示意图,如图8所示,编码器80具体用于:
用于根据接收到的当前帧信号,得到当前帧信号的低带信号和当前帧信号的高带信号;
对当前帧信号的低带信号进行编码,得到低带编码的激励信号;
对当前帧信号的高带信号进行线性预测,得到线性预测系数;
量化线性预测系数,得到量化后的线性预测系数;
根据低带编码的激励信号和量化后的线性预测系数得到预测的高带信号;
计算及量化预测的高带信号的时域包络;
其中,计算所述预测的高带信号的时域包络包括:
根据预先确定的时域包络个数M将预测的高带信号分成M个子帧,其中,M为大于等于2的整数,
采用非对称窗对M个子帧中的最前端的子帧和M个子帧中的最末端的子帧进行加窗,
对M个子帧中除所述最前端的子帧和最末端的子帧之外的子帧进行加窗;
对量化后的时域包络进行编码。
可以理解的是,编码器80可以用于执行上述任意的方法实施例。也可以包括任意实施例的时域包络处理装置70。具体的编码器80所执行的功能可参考前述方法和装置实施例,在此不再赘述。
本领域普通技术人员可以理解:实现上述各方法实施例的全部或部分步骤可以通过程序指令相关的硬件来完成。前述的程序可以存储于一计算机可读取存储介质中。该程序在执行时,执行包括上述各方法实施例的步骤;而前述的存储介质包括:ROM、RAM、磁碟或者光盘等各种可以存储程序代码的介质。
最后应说明的是:以上各实施例仅用以说明本发明的技术方案,而非对其限制;尽管参照前述各实施例对本发明进行了详细的说明,本领域的普通技术人员应当理解:其依然可以对前述各实施例所记载的技术方案进行修改,或者对其中部分或者全部技术特征进行等同替换;而这些修改或者替换,并不使相应技术方案的本质脱离本发明各实施例技术方案的范围。

Claims (15)

  1. 一种音频信号的时域包络处理方法,其特征在于,包括:
    根据接收到的当前帧信号,得到所述当前帧信号的高带信号;
    根据预先确定的时域包络个数M将所述当前帧的高带信号分成M个子帧,其中,M为大于等于2的整数;
    计算每一个所述子帧的时域包络;
    其中,所述计算每一个所述子帧的时域包络包括:
    采用非对称窗对所述M个子帧中的最前端的子帧和所述M个子帧中的最末端的子帧进行加窗;
    对所述M个子帧中除所述最前端的子帧和所述最末端的子帧之外的子帧进行加窗。
  2. 根据权利要求1所述的方法,其特征在于,在采用非对称窗对所述M个子帧中的最前端的子帧和所述M个子帧中的最末端的子帧进行加窗之前,所述方法还包括:
    根据所述当前帧信号的高带信号的前向缓存的长度确定所述非对称窗;或者,
    根据所述当前帧信号的高带信号的前向缓存的长度和所述时域包络个数M确定所述非对称窗。
  3. 根据权利要求1或2所述的方法,其特征在于,所述对所述M个子帧中除所述最前端的子帧和所述最末端的子帧之外的子帧进行加窗,包括:
    对所述M个子帧中除最前端的子帧和所述最末端的子帧之外的子帧采用对称窗进行加窗;或者,
    对所述M个子帧中除最前端的子帧和所述最末端的子帧之外的子帧采用非对称窗进行加窗。
  4. 根据权利要求1所述的方法,其特征在于,所述非对称窗的窗长与对所述M个子帧中除所述最前端的子帧和所述最末端的子帧之外的子帧进行加窗所采用的窗的窗长相同。
  5. 根据权利要求2-4任意之一所述的方法,其特征在于,所述根据所述当前帧音频信号的高带信号的前向缓存的长度确定非对称窗,包 括:
    当所述当前帧信号的高带信号的前向缓存的长度小于第一阈值时,根据当前帧的前一帧信号的高带信号和所述当前帧信号的高带信号的前向缓存的长度确定所述非对称窗,其中,所述当前帧的前一帧信号的高带信号的最末端子帧采用的非对称窗和所述当前帧信号的高带信号的最前端子帧采用的非对称窗的混叠部分等于所述当前帧信号的高带信号的前向缓存的长度,所述第一阈值等于所述当前帧的高带信号的帧长除以M。
  6. 根据权利要求2-4任意之一所述的方法,其特征在于,所述根据所述当前帧信号的高带信号的前向缓存的长度确定非对称窗,包括:
    当所述当前帧信号的高带信号的前向缓存的长度大于第一阈值时,根据所述当前帧的前一帧信号的高带信号和所述当前帧信号的高带信号的前向缓存的长度确定所述非对称窗,其中,所述当前帧的前一帧信号的高带信号的最末端子帧采用的非对称窗和所述当前帧信号的高带信号的最前端子帧采用的非对称窗的混叠部分等于所述第一阈值,所述第一阈值等于所述当前帧的高带信号的帧长除以M。
  7. 根据权利要求1-6任意之一所述的方法,其特征在于,根据下列之一方式确定所述时域包络个数M:
    根据所述当前帧信号得到所述当前帧信号的低带信号,当所述当前帧信号的低带信号的基音周期大于第二阈值时,M=M1;或者,
    根据所述当前帧信号得到所述当前帧信号的低带信号,当所述当前帧信号的低带信号的基音周期不大于第二阈值时,M=M2;
    其中,M1,M2均为正整数,且M2>M1。
  8. 根据权利要求1-6任意之一所述的方法,其特征在于,所述方法还包括:
    根据所述当前帧信号得到所述当前帧信号的低带信号的基音周期;
    当所述当前帧信号的类型与所述当前帧的前一帧信号的类型相同,且所述当前帧的低带信号的基音周期大于第三阈值时,对每一个所述子帧的时域包络进行平滑处理。
  9. 一种音频信号的时域包络处理装置,其特征在于,包括:
    高带信号获取模块,用于根据接收到的当前帧信号,得到所述当前帧信号的高带信号;
    子帧获取模块,用于根据预先确定的时域包络个数M将所述当前帧的高带信号分成M个子帧,其中,M为大于等于2的整数;
    时域包络获取模块,用于计算每一个所述子帧的时域包络;
    其中,所述时域包络获取模块具体用于:
    采用非对称窗对所述M个子帧中的最前端的子帧和所述M个子帧中的最末端的子帧进行加窗;
    对所述M个子帧中除所述最前端的子帧和所述最末端的子帧之外的子帧进行加窗。
  10. 根据权利要求9所述的装置,其特征在于,所述时域包络获取模块还用于:
    根据所述当前帧信号的高带信号的前向缓存的长度确定所述非对称窗;或者,
    根据所述当前帧信号的高带信号的前向缓存的长度和所述时域包络个数M确定所述非对称窗。
  11. 根据权利要求9所述的装置,其特征在于,所述时域包络获取模块具体用于:
    采用非对称窗对所述M个子帧中的最前端的子帧和所述M个子帧中的最末端的子帧进行加窗,对所述M个子帧中除最前端的子帧和所述最末端的子帧之外的子帧采用对称窗进行加窗;或者,
    采用非对称窗对所述M个子帧中的最前端的子帧和所述M个子帧中的最末端的子帧进行加窗,对所述M个子帧中除最前端的子帧和所述最末端的子帧之外的子帧采用非对称窗进行加窗。
  12. 根据权利要求9所述的装置,其特征在于,所述非对称窗的窗长与对所述M个子帧中除所述最前端的子帧和所述最末端的子帧之外的子帧进行加窗所采用的窗的窗长相同。
  13. 根据权利要求9-12任意之一所述的装置,其特征在于,还包括:确定模块,用于根据下列之一方式确定所述时域包络个数M:
    根据所述当前帧信号得到所述当前帧信号的低带信号,当所述当前帧 信号的低带信号的基音周期大于第二阈值时,M=M1;或者,
    根据所述当前帧信号得到所述当前帧信号的低带信号,当所述当前帧信号的低带信号的基音周期不大于第二阈值时,M=M2;
    其中,M1,M2均为正整数,且M2>M1。
  14. 根据权利要求9-13任意之一所述的装置,其特征在于,所述时域包络获取模块还用于:
    根据所述当前帧信号得到所述当前帧信号的低带信号的基音周期;
    当所述当前帧信号的类型与所述当前帧的前一帧信号的类型相同,且所述当前帧的低带信号的基音周期大于第三阈值时,对每一个所述子帧的时域包络进行平滑处理。
  15. 一种编码器,其特征在于,所述编码器具体用于:
    用于根据接收到的当前帧信号,得到所述当前帧信号的低带信号和所述当前帧信号的高带信号;
    对所述当前帧信号的低带信号进行编码,得到低带编码的激励信号;
    对所述当前帧信号的高带信号进行线性预测,得到线性预测系数;
    量化所述线性预测系数,得到量化后的线性预测系数;
    根据所述低带编码的激励信号和所述量化后的线性预测系数得到预测的高带信号;
    计算及量化所述预测的高带信号的时域包络;
    其中,所述计算所述预测的高带信号的时域包络包括:
    根据预先确定的时域包络个数M将所述预测的高带信号分成M个子帧,其中,M为大于等于2的整数,
    采用非对称窗对所述M个子帧中的最前端的子帧和所述M个子帧中的最末端的子帧进行加窗,
    对所述M个子帧中除所述最前端的子帧和所述最末端的子帧之外的子帧进行加窗;
    对量化后的时域包络进行编码。
PCT/CN2015/071727 2014-06-12 2015-01-28 一种音频信号的时域包络处理方法及装置、编码器 Ceased WO2015188627A1 (zh)

Priority Applications (7)

Application Number Priority Date Filing Date Title
JP2016572398A JP6510566B2 (ja) 2014-06-12 2015-01-28 オーディオ信号の時間包絡線を処理するための方法および装置、ならびにエンコーダ
KR1020167033851A KR101896486B1 (ko) 2014-06-12 2015-01-28 오디오 신호의 시간 엔벨로프를 처리하기 위한 방법과 장치, 및 인코더
EP19169470.2A EP3579229B1 (en) 2014-06-12 2015-01-28 Method and encoder for processing temporal envelope of audio signal
EP15806700.9A EP3133599B1 (en) 2014-06-12 2015-01-28 Method and encoder of processing temporal envelope of audio signal
US15/372,130 US9799343B2 (en) 2014-06-12 2016-12-07 Method and apparatus for processing temporal envelope of audio signal, and encoder
US15/708,617 US10170128B2 (en) 2014-06-12 2017-09-19 Method and apparatus for processing temporal envelope of audio signal, and encoder
US16/201,647 US10580423B2 (en) 2014-06-12 2018-11-27 Method and apparatus for processing temporal envelope of audio signal, and encoder

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
CN201410260730.5A CN105336336B (zh) 2014-06-12 2014-06-12 一种音频信号的时域包络处理方法及装置、编码器
CN201410260730.5 2014-06-12

Related Child Applications (1)

Application Number Title Priority Date Filing Date
US15/372,130 Continuation US9799343B2 (en) 2014-06-12 2016-12-07 Method and apparatus for processing temporal envelope of audio signal, and encoder

Publications (1)

Publication Number Publication Date
WO2015188627A1 true WO2015188627A1 (zh) 2015-12-17

Family

ID=54832857

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2015/071727 Ceased WO2015188627A1 (zh) 2014-06-12 2015-01-28 一种音频信号的时域包络处理方法及装置、编码器

Country Status (8)

Country Link
US (3) US9799343B2 (zh)
EP (2) EP3579229B1 (zh)
JP (2) JP6510566B2 (zh)
KR (1) KR101896486B1 (zh)
CN (2) CN105336336B (zh)
ES (1) ES2895495T3 (zh)
PT (1) PT3579229T (zh)
WO (1) WO2015188627A1 (zh)

Cited By (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
WO2017125840A1 (en) * 2016-01-19 2017-07-27 Hua Kanru Method for analysis and synthesis of aperiodic signals

Families Citing this family (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN105336336B (zh) * 2014-06-12 2016-12-28 华为技术有限公司 一种音频信号的时域包络处理方法及装置、编码器
JP6501259B2 (ja) * 2015-08-04 2019-04-17 本田技研工業株式会社 音声処理装置及び音声処理方法
CN108109629A (zh) * 2016-11-18 2018-06-01 南京大学 一种基于线性预测残差分类量化的多描述语音编解码方法和系统
CN111402917B (zh) * 2020-03-13 2023-08-04 北京小米松果电子有限公司 音频信号处理方法及装置、存储介质

Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPH10222194A (ja) * 1997-02-03 1998-08-21 Gotai Handotai Kofun Yugenkoshi 音声符号化における有声音と無声音の識別方法
JP2001166800A (ja) * 1999-12-09 2001-06-22 Nippon Telegr & Teleph Corp <Ntt> 音声符号化方法及び音声復号化方法
CN1424712A (zh) * 2002-12-19 2003-06-18 北京工业大学 2.3kb/s谐波激励线性预测语音编码方法
US20120016668A1 (en) * 2010-07-19 2012-01-19 Futurewei Technologies, Inc. Energy Envelope Perceptual Correction for High Band Coding
WO2013066238A2 (en) * 2011-11-02 2013-05-10 Telefonaktiebolaget L M Ericsson (Publ) Generation of a high band extension of a bandwidth extended audio signal

Family Cites Families (27)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN1062963C (zh) * 1990-04-12 2001-03-07 多尔拜实验特许公司 用于产生高质量声音信号的解码器和编码器
US5754534A (en) * 1996-05-06 1998-05-19 Nahumi; Dror Delay synchronization in compressed audio systems
JP3518737B2 (ja) * 1999-10-25 2004-04-12 日本ビクター株式会社 オーディオ符号化装置、オーディオ符号化方法、及びオーディオ符号化信号記録媒体
EP1199711A1 (en) * 2000-10-20 2002-04-24 Telefonaktiebolaget Lm Ericsson Encoding of audio signal using bandwidth expansion
US7424434B2 (en) * 2002-09-04 2008-09-09 Microsoft Corporation Unified lossy and lossless audio compression
US7630902B2 (en) * 2004-09-17 2009-12-08 Digital Rise Technology Co., Ltd. Apparatus and methods for digital audio coding using codebook application ranges
RU2386179C2 (ru) * 2005-04-01 2010-04-10 Квэлкомм Инкорпорейтед Способ и устройство для кодирования речевых сигналов с расщеплением полосы
TWI324336B (en) 2005-04-22 2010-05-01 Qualcomm Inc Method of signal processing and apparatus for gain factor smoothing
KR101390188B1 (ko) * 2006-06-21 2014-04-30 삼성전자주식회사 적응적 고주파수영역 부호화 및 복호화 방법 및 장치
US9159333B2 (en) 2006-06-21 2015-10-13 Samsung Electronics Co., Ltd. Method and apparatus for adaptively encoding and decoding high frequency band
US8532984B2 (en) * 2006-07-31 2013-09-10 Qualcomm Incorporated Systems, methods, and apparatus for wideband encoding and decoding of active frames
US8260609B2 (en) * 2006-07-31 2012-09-04 Qualcomm Incorporated Systems, methods, and apparatus for wideband encoding and decoding of inactive frames
US9454974B2 (en) * 2006-07-31 2016-09-27 Qualcomm Incorporated Systems, methods, and apparatus for gain factor limiting
PT3550564T (pt) * 2007-08-27 2020-08-18 Ericsson Telefon Ab L M Análise/síntese espectral de baixa complexidade utilizando resolução temporal selecionável
CN101615394B (zh) * 2008-12-31 2011-02-16 华为技术有限公司 分配子帧的方法和装置
WO2010084756A1 (ja) * 2009-01-22 2010-07-29 パナソニック株式会社 ステレオ音響信号符号化装置、ステレオ音響信号復号装置およびそれらの方法
US8457975B2 (en) * 2009-01-28 2013-06-04 Fraunhofer-Gesellschaft Zur Foerderung Der Angewandten Forschung E.V. Audio decoder, audio encoder, methods for decoding and encoding an audio signal and computer program
US8718804B2 (en) * 2009-05-05 2014-05-06 Huawei Technologies Co., Ltd. System and method for correcting for lost data in a digital audio signal
PL2471061T3 (pl) * 2009-10-08 2014-03-31 Fraunhofer Ges Forschung Działający w wielu trybach dekoder sygnału audio, działający w wielu trybach koder sygnału audio, sposoby i program komputerowy stosujące kształtowanie szumu oparte o kodowanie z wykorzystaniem predykcji liniowej
BR112012009032B1 (pt) 2009-10-20 2021-09-21 Fraunhofer-Gesellschaft zur Förderung der angewandten Forschung e. V. Codificador de sinal de áudio, decodificador de sinal de áudio, método para prover uma representação codificada de um conteúdo de áudio, método para prover uma representação decodificada de um conteúdo de áudio para uso em aplicações de baixo retardamento
US9047875B2 (en) * 2010-07-19 2015-06-02 Futurewei Technologies, Inc. Spectrum flatness control for bandwidth extension
CN102436820B (zh) * 2010-09-29 2013-08-28 华为技术有限公司 高频带信号编码方法及装置、高频带信号解码方法及装置
BR112013020239B1 (pt) * 2011-02-14 2021-12-21 Fraunhofer-Gellschaft Zur Förderung Der Angewandten Forschung E.V. Geração de ruído em codecs de áudio
CN106128473B (zh) * 2011-06-30 2019-12-10 三星电子株式会社 用于产生带宽扩展信号的设备和方法
US9275644B2 (en) * 2012-01-20 2016-03-01 Qualcomm Incorporated Devices for redundant frame coding and decoding
US9384746B2 (en) * 2013-10-14 2016-07-05 Qualcomm Incorporated Systems and methods of energy-scaled signal processing
CN105336336B (zh) * 2014-06-12 2016-12-28 华为技术有限公司 一种音频信号的时域包络处理方法及装置、编码器

Patent Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPH10222194A (ja) * 1997-02-03 1998-08-21 Gotai Handotai Kofun Yugenkoshi 音声符号化における有声音と無声音の識別方法
JP2001166800A (ja) * 1999-12-09 2001-06-22 Nippon Telegr & Teleph Corp <Ntt> 音声符号化方法及び音声復号化方法
CN1424712A (zh) * 2002-12-19 2003-06-18 北京工业大学 2.3kb/s谐波激励线性预测语音编码方法
US20120016668A1 (en) * 2010-07-19 2012-01-19 Futurewei Technologies, Inc. Energy Envelope Perceptual Correction for High Band Coding
WO2013066238A2 (en) * 2011-11-02 2013-05-10 Telefonaktiebolaget L M Ericsson (Publ) Generation of a high band extension of a bandwidth extended audio signal

Cited By (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
WO2017125840A1 (en) * 2016-01-19 2017-07-27 Hua Kanru Method for analysis and synthesis of aperiodic signals

Also Published As

Publication number Publication date
EP3133599A4 (en) 2017-07-12
US20180005638A1 (en) 2018-01-04
US20190096415A1 (en) 2019-03-28
JP2017523448A (ja) 2017-08-17
US10170128B2 (en) 2019-01-01
EP3579229A1 (en) 2019-12-11
US9799343B2 (en) 2017-10-24
JP6765471B2 (ja) 2020-10-07
ES2895495T3 (es) 2022-02-21
US10580423B2 (en) 2020-03-03
CN105336336B (zh) 2016-12-28
PT3579229T (pt) 2021-08-20
JP2019135551A (ja) 2019-08-15
CN106409304A (zh) 2017-02-15
EP3133599B1 (en) 2019-07-10
EP3579229B1 (en) 2021-07-28
CN106409304B (zh) 2020-08-25
KR20160147048A (ko) 2016-12-21
EP3133599A1 (en) 2017-02-22
US20170098451A1 (en) 2017-04-06
JP6510566B2 (ja) 2019-05-08
KR101896486B1 (ko) 2018-09-07
CN105336336A (zh) 2016-02-17

Similar Documents

Publication Publication Date Title
CN103493129B (zh) 用于使用瞬态检测及质量结果将音频信号的部分编码的装置与方法
CN102834863B (zh) 用于包括通用音频和语音帧的音频信号的解码器
CN103109318B (zh) 利用前向混迭消除技术的编码器
JP6765471B2 (ja) オーディオ信号の時間包絡線を処理するための方法および装置、ならびにエンコーダ
CN110444219B (zh) 选择第一编码演算法或第二编码演算法的装置与方法
RU2618848C2 (ru) Устройство и способ для выбора одного из первого алгоритма кодирования аудио и второго алгоритма кодирования аудио
CN101197134A (zh) 消除编码模式切换影响的方法和装置以及解码方法和装置
US20130096913A1 (en) Method and apparatus for adaptive multi rate codec
Li et al. A 1.8 kbps vocoder based on Mixed Excitation Linear Prediction
HK1192049B (zh) 用於使用瞬态检测及品质结果将音讯信号的部分编码
HK1192049A (zh) 用於使用瞬态检测及品质结果将音讯信号的部分编码

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 15806700

Country of ref document: EP

Kind code of ref document: A1

REEP Request for entry into the european phase

Ref document number: 2015806700

Country of ref document: EP

WWE Wipo information: entry into national phase

Ref document number: 2015806700

Country of ref document: EP

ENP Entry into the national phase

Ref document number: 20167033851

Country of ref document: KR

Kind code of ref document: A

ENP Entry into the national phase

Ref document number: 2016572398

Country of ref document: JP

Kind code of ref document: A

NENP Non-entry into the national phase

Ref country code: DE