TWI882003B - Low-latency, low-frequency effects codec - Google Patents

Low-latency, low-frequency effects codec Download PDF

Info

Publication number
TWI882003B
TWI882003B TW109130176A TW109130176A TWI882003B TW I882003 B TWI882003 B TW I882003B TW 109130176 A TW109130176 A TW 109130176A TW 109130176 A TW109130176 A TW 109130176A TW I882003 B TWI882003 B TW I882003B
Authority
TW
Taiwan
Prior art keywords
coefficients
channel signal
lfe channel
lfe
processors
Prior art date
Application number
TW109130176A
Other languages
Chinese (zh)
Other versions
TW202211206A (en
Inventor
里沙普 塔吉
大衛 S 麥格拉斯
Original Assignee
美商杜拜研究特許公司
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by 美商杜拜研究特許公司 filed Critical 美商杜拜研究特許公司
Priority to TW109130176A priority Critical patent/TWI882003B/en
Publication of TW202211206A publication Critical patent/TW202211206A/en
Application granted granted Critical
Publication of TWI882003B publication Critical patent/TWI882003B/en

Links

Landscapes

  • Compression, Expansion, Code Conversion, And Decoders (AREA)

Abstract

In some implementations, a method of encoding a low-frequency effect (LFE) channel comprises: receiving a time-domain LFE channel signal; filtering, using a low-pass filter, the time-domain LFE channel signal; converting the filtered time-domain LFE channel signal into a frequency-domain representation of the LFE channel signal that includes a number of coefficients representing a frequency spectrum of the LFE channel signal; arranging coefficients into a number of subband groups corresponding to different frequency bands of the LFE channel signal; quantizing coefficients in each subband group according to a frequency response curve of the low-pass filter; encoding the quantized coefficients in each subband group using an entropy coder tuned for the subband group; and generating a bitstream including the encoded quantized coefficients; and storing the bitstream on a storage device or streaming the bitstream to a downstream device.

Description

低延遲、低頻率效應之編碼解碼器Low latency, low frequency effect codec

本發明大體而言係關於音訊信號處理,且確切而言,係關於處理低頻率效應(LFE)聲道。 The present invention relates generally to audio signal processing, and more particularly to processing low frequency effects (LFE) channels.

舉例而言,沉浸式服務之標準化努力包含針對聲音、多串流電傳會議、虛擬實境(VR)、使用者產生之現場及非現場內容串流傳輸開發一沉浸式聲音與音訊服務(IVAS)編碼解碼器。IVAS標準之一目標係開發音訊品質出色、延遲低、支援空間音訊寫碼、具有一恰當位元速率範圍、高品質錯誤恢復及一實際實施複雜性的一單個編碼解碼器。為達成此目標,期望開發可基於能夠進行IVAS之裝置或能夠處理LFE信號之任何其他裝置來處置低延遲LFE操作的一IVAS編碼解碼器。LFE聲道用於範圍為20Hz至120Hz之深度低音調聲響,且通常發送至經設計以再生低頻率音訊內容之一揚聲器。 For example, standardization efforts for immersive services include the development of an Immersive Sound and Audio Services (IVAS) codec for sound, multi-stream teleconferencing, virtual reality (VR), user-generated live and non-live content streaming. One of the goals of the IVAS standard is to develop a single codec with excellent audio quality, low latency, support for spatial audio coding, an appropriate bit rate range, high-quality error resilience, and a practical implementation complexity. To achieve this goal, it is desirable to develop an IVAS codec that can handle low-latency LFE operations based on IVAS-capable devices or any other device capable of processing LFE signals. The LFE channel is used for deep bass-pitched sounds in the 20Hz to 120Hz range and is typically sent to a speaker designed to reproduce low-frequency audio content.

揭示一可配置低延遲LFE編碼解碼器之實施方案。 An implementation scheme of a configurable low-latency LFE codec is disclosed.

在一些實施方案中,一種對一低頻率效應(LFE)聲道進行編碼之方法包括:使用一或多個處理器接收一時域LFE聲道信號;使用一低通濾波器對該時域LFE聲道信號進行濾波;使用該一或多個處理器將經濾 波之該時域LFE聲道信號轉換成該LFE聲道信號的包含表示該LFE聲道信號之一頻譜之一定數目個係數之一頻域表示;使用該一或多個處理器將係數配置至與該LFE聲道信號之不同頻帶對應之一定數目個次頻帶群組中;使用該一或多個處理器根據該低通濾波器之一頻率回應曲線將每一次頻帶群組中之係數量化;使用該一或多個處理器使用針對每一次頻帶群組調諧之一熵寫碼器對該次頻帶群組中之該等經量化係數進行編碼;及使用該一或多個處理器產生包含經編碼之該等經量化係數之一位元串流;以及使用該一或多個處理器將該位元串流儲存於一儲存裝置上或將該位元串流串流傳輸至一下游裝置。 In some implementations, a method for encoding a low frequency effects (LFE) channel includes: receiving a time domain LFE channel signal using one or more processors; filtering the time domain LFE channel signal using a low pass filter; converting the filtered time domain LFE channel signal into a frequency domain representation of the LFE channel signal including a certain number of coefficients representing a spectrum of the LFE channel signal using the one or more processors; arranging the coefficients to correspond to different frequency bands of the LFE channel signal using the one or more processors; quantize the coefficients in each sub-band group according to a frequency response curve of the low-pass filter using the one or more processors; encode the quantized coefficients in the sub-band group using an entropy encoder tuned for each sub-band group using the one or more processors; and generate a bit stream including the encoded quantized coefficients using the one or more processors; and store the bit stream on a storage device or stream the bit stream to a downstream device using the one or more processors.

在一些實施方案中,將每一次頻帶群組中之該等係數量化進一步包括基於可用量化點之一最大數目及該等係數之絕對值之一和來產生一縮放移位因數;及使用該縮放移位因數將該等係數量化。 In some implementations, quantizing the coefficients in each sub-band grouping further includes generating a scale shift factor based on a maximum number of available quantization points and a sum of the absolute values of the coefficients; and quantizing the coefficients using the scale shift factor.

在一些實施方案中,若一經量化係數超過量化點之該最大數目,則將該縮放移位因數減小且再次將該等係數量化。 In some implementations, if a quantized coefficient exceeds the maximum number of quantization points, the scaling factor is reduced and the coefficients are quantized again.

在一些實施方案中,對於每一次頻帶群組而言,該等量化點係不同的。 In some implementations, the quantization points are different for each band grouping.

在一些實施方案中,根據一精細量化方案或一粗略量化方案將每一次頻帶群組中之該等係數量化,其中與根據該粗略量化方案指派給一或多個次頻帶群組之量化點相比,利用該精細量化方案將更多量化點分配給該等各別次頻帶群組。 In some implementations, the coefficients in each sub-band grouping are quantized according to a fine quantization scheme or a coarse quantization scheme, wherein more quantization points are allocated to the respective sub-band groupings using the fine quantization scheme than quantization points assigned to the one or more sub-band groupings according to the coarse quantization scheme.

在一些實施方案中,該等係數之正負號位元與該等係數被分開寫碼。 In some implementations, the sign bits of the coefficients are coded separately from the coefficients.

在一些實施方案中,存在四個次頻帶群組,且一第一次頻 帶群組對應於0Hz至100Hz之一第一頻率範圍,一第二次頻帶群組對應於100Hz至200Hz之一第二頻率範圍,一第三次頻帶群組對應於200Hz至300Hz之一第三頻率範圍,且一第四次頻帶群組對應於300Hz至400Hz之一第四頻率範圍。 In some embodiments, there are four sub-band groups, and a first sub-band group corresponds to a first frequency range of 0 Hz to 100 Hz, a second sub-band group corresponds to a second frequency range of 100 Hz to 200 Hz, a third sub-band group corresponds to a third frequency range of 200 Hz to 300 Hz, and a fourth sub-band group corresponds to a fourth frequency range of 300 Hz to 400 Hz.

在一些實施方案中,該熵寫碼器係一算術熵寫碼器。 In some implementations, the entropy coder is an arithmetic entropy coder.

在一些實施方案中,將經濾波之該時域LFE聲道信號轉換成該LFE聲道信號的包括表示該LFE聲道信號之一頻譜之一定數目個係數的一頻域表示進一步包括:判定該LFE聲道信號之一第一步長;基於該第一步長指定一窗函數之一第一窗口大小;將該第一窗口大小應用於該時域LFE聲道信號之一或多個訊框;及將一修改型離散餘弦變換(MDCT)應用於經窗口化之該訊框以產生該等係數。 In some implementations, converting the filtered time-domain LFE channel signal into a frequency-domain representation of the LFE channel signal including a certain number of coefficients representing a spectrum of the LFE channel signal further includes: determining a first step length of the LFE channel signal; specifying a first window size of a window function based on the first step length; applying the first window size to one or more frames of the time-domain LFE channel signal; and applying a modified discrete cosine transform (MDCT) to the windowed frames to generate the coefficients.

在一些實施方案中,該方法進一步包括:判定該LFE聲道信號之一第二步長;基於該第二步長指定該窗函數之一第二窗口大小;及將該第二窗口大小應用於該時域LFE聲道信號之該一或多個訊框。 In some implementations, the method further includes: determining a second step size of the LFE channel signal; specifying a second window size of the window function based on the second step size; and applying the second window size to the one or more frames of the time-domain LFE channel signal.

在一些實施方案中,該第一步長係N毫秒(ms),N大於或等於5ms且小於或等於60ms,該第一窗口大小高於或等於10ms,該第二步長係5ms且該第二窗口大小係10ms。 In some implementations, the first step length is N milliseconds (ms), N is greater than or equal to 5ms and less than or equal to 60ms, the first window size is greater than or equal to 10ms, the second step length is 5ms and the second window size is 10ms.

在一些實施方案中,該第一步長係20毫秒(ms),該第一窗口大小係10ms或20ms或40ms,該第二步長係10ms且該第二窗口大小係10ms或20ms。 In some implementations, the first step length is 20 milliseconds (ms), the first window size is 10ms or 20ms or 40ms, the second step length is 10ms and the second window size is 10ms or 20ms.

在一些實施方案中,該第一步長係10毫秒(ms),該第一窗口大小係10ms或20ms,該第二步長係5ms,且該第二窗口大小係10ms。 In some implementations, the first step size is 10 milliseconds (ms), the first window size is 10ms or 20ms, the second step size is 5ms, and the second window size is 10ms.

在一些實施方案中,該第一步長係20毫秒(ms),該第一窗口大小係10ms、20ms或40ms,該第二步長係5ms且該第二窗口大小係10ms。 In some implementations, the first step size is 20 milliseconds (ms), the first window size is 10ms, 20ms, or 40ms, the second step size is 5ms, and the second window size is 10ms.

在一些實施方案中,該窗函數係具有一可配置漸隱長度之一凱撒-貝索導出(KBD)窗函數。 In some embodiments, the window function is a Kaiser-Besso derived (KBD) window function with a configurable gradient length.

在一些實施方案中,該低通濾波器係一截止頻率為約130Hz或低於130Hz之一四階巴特沃斯低通濾波器。 In some embodiments, the low-pass filter is a fourth-order Butterworth low-pass filter having a cutoff frequency of about 130 Hz or less.

在一些實施方案中,該方法進一步包括:使用該一或多個處理器判定該LFE聲道信號之一訊框之一能量位準是否低於一臨限值;根據該能量位準低於一臨限位準,產生一靜寂訊框指示符以指示該解碼器;將該靜寂訊框指示符插入至該LFE聲道位元串流之後設資料中;及在偵測到靜寂訊框時減小一LFE聲道位元速率。 In some implementations, the method further includes: using the one or more processors to determine whether an energy level of a frame of the LFE channel signal is lower than a threshold value; generating a silence frame indicator to indicate the decoder according to the energy level being lower than a threshold level; inserting the silence frame indicator into the metadata of the LFE channel bit stream; and reducing an LFE channel bit rate when a silence frame is detected.

在一些實施方案中,一種對一低頻率效應(LFE)聲道位元串流進行解碼之方法包括:使用一或多個處理器接收一LFE聲道位元串流,該LFE聲道位元串流包含表示一時域LFE聲道信號之一頻譜之熵寫碼係數;使用該一或多個處理器使用一熵解碼器將經量化係數解碼;使用該一或多個處理器將經量化係數逆量化,其中該等係數已根據用於在一編碼器中對該時域LFE聲道信號進行濾波之一低通濾波器之一頻率回應曲線而在與頻帶對應之次頻帶群組中被量化;使用該一或多個處理器將經逆量化之該等係數轉換成一時域LFE聲道信號;使用該一或多個處理器調整該時域LFE聲道信號之一延時;及使用一低通濾波器對經延時調整之該LFE聲道信號進行濾波。 In some implementations, a method for decoding a low frequency effect (LFE) channel bit stream includes: receiving, using one or more processors, an LFE channel bit stream, the LFE channel bit stream comprising entropy-coded coefficients representing a spectrum of a time-domain LFE channel signal; decoding, using the one or more processors, the quantized coefficients using an entropy decoder; and dequantizing, using the one or more processors, the quantized coefficients, wherein the coefficients quantized in a sub-band group corresponding to a frequency band according to a frequency response curve of a low-pass filter used to filter the time-domain LFE channel signal in an encoder; converting the inversely quantized coefficients into a time-domain LFE channel signal using the one or more processors; adjusting a delay of the time-domain LFE channel signal using the one or more processors; and filtering the delay-adjusted LFE channel signal using a low-pass filter.

在一些實施方案中,低通濾波器之一階數經組態以確保由 於對包含該LFE聲道信號之一多聲道音訊信號中之該LFE聲道進行編碼及解碼所致之一第一總演算法延時小於或等於由於對其他聲道進行編碼及解碼所致之一第二總演算法延時。 In some implementations, an order of the low pass filter is configured to ensure that a first total algorithm delay due to encoding and decoding the LFE channel in a multi-channel audio signal including the LFE channel signal is less than or equal to a second total algorithm delay due to encoding and decoding other channels.

在一些實施方案中,該方法進一步包括:判定該第二總演算法延時是否超過一臨限值;及根據該第二總演算法延時超過該臨限值,將該低通濾波器組態為一N階低通濾波器,其中N係大於或等於2之一整數;及根據該第二總演算法延時不超過該臨限值,將該低通濾波器之該階數組態為小於N。 In some implementations, the method further includes: determining whether the second total algorithm delay exceeds a critical value; and configuring the low-pass filter as an N-order low-pass filter according to the second total algorithm delay exceeding the critical value, wherein N is an integer greater than or equal to 2; and configuring the order of the low-pass filter to be less than N according to the second total algorithm delay not exceeding the critical value.

本文中所揭示之其他實施方案係關於一系統、設備及電腦可讀媒體。在隨附圖式及下文說明中陳述一或多個所揭示實施方案的細節。依據說明、圖式及申請專利範圍明瞭其他特徵、目標及優點。 Other embodiments disclosed herein relate to a system, apparatus, and computer-readable medium. Details of one or more disclosed embodiments are set forth in the accompanying drawings and the following description. Other features, objects, and advantages will be apparent from the description, drawings, and claims.

本文中所揭示之特定實施例提供以下優點中之一或多者。所揭示之低延遲LFE編碼解碼器:1)主要針對LFE聲道;2)主要針對20Hz至120Hz之一頻率範圍,但在低/中等位元速率情景中攜載達300Hz之音訊且在高位元速率情景中攜載達400Hz之音訊;3)藉由根據一輸入低通濾波器之一頻率回應曲線應用一量化方案來達成一低位元速率;4)具有一低演算法延遲且經設計而以20毫秒(ms)之一步幅操作並且具有33毫秒之一總演算法延遲(包含成框);5)可經組態至更小之步幅及更低之演算法延遲以支援其他情景,包含組態成低至5毫秒之步幅及13毫秒之總演算法延遲(包含成框);6)基於LFE編碼解碼器可用之延遲在解碼器輸出處自動地選擇一低通濾波器;7)具有在靜寂期間具有50位元/秒(bps)之一低位元速率之一靜寂模式;及8)在有效訊框期間,位元速率基於所使用之量化位準而在2千位元/秒(kbps)至4kbps之間波動,且在靜寂訊框期間位元速率係50 bps。 Specific embodiments disclosed herein provide one or more of the following advantages. The disclosed low-latency LFE codec: 1) is primarily targeted at the LFE channel; 2) is primarily targeted at a frequency range of 20 Hz to 120 Hz, but carries audio up to 300 Hz in low/medium bit rate scenarios and up to 400 Hz in high bit rate scenarios; 3) achieves a low bit rate by applying a quantization scheme based on a frequency response curve of an input low-pass filter; 4) has a low algorithmic delay and is designed to operate in a step size of 20 milliseconds (ms) and has a total algorithmic delay of 33 ms (including framing); 5) can be configured to smaller steps and lower algorithm delays to support other scenarios, including configuration to steps as low as 5 ms and 13 ms total algorithm delay (including framing); 6) automatically select a low pass filter at the decoder output based on the available delay of the LFE codec; 7) have a silent mode with a low bit rate of 50 bits per second (bps) during silent periods; and 8) during active frames, the bit rate fluctuates between 2 kilobits per second (kbps) and 4 kbps based on the quantization level used, and during silent frames the bit rate is 50 bps.

100:沉浸式聲音與音訊服務編碼解碼器 100: Immersive sound and audio service codec

101:音訊資料 101: Audio data

102:空間分析與降混單元 102: Spatial analysis and downmixing unit

103:主音訊聲道編碼單元/主音訊聲道編碼器 103: Main audio channel encoding unit/main audio channel encoder

104:空間後設資料編碼單元 104: Spatial metadata coding unit

105:低頻率效應聲道編碼單元 105: Low frequency effect channel coding unit

106:空間後設資料解碼單元 106: Spatial metadata decoding unit

107:主音訊聲道解碼單元 107: Main audio channel decoding unit

108:低頻率效應聲道解碼單元 108: Low frequency effect channel decoding unit

109:空間合成/升混/呈現單元 109: Spatial synthesis/upmixing/presentation unit

201:輸入低通濾波器/四階輸入低通濾波器/低通濾波器 201: Input low-pass filter/fourth-order input low-pass filter/low-pass filter

202:修改型離散餘弦變換單元 202: Modified Discrete Cosine Transformation Unit

203:量化與熵寫碼單元 203: Quantization and entropy coding unit

204:熵解碼與逆量化單元 204: Entropy decoding and inverse quantization unit

205:逆修改型離散餘弦變換與窗口化單元 205: Inverse modified discrete cosine transform and windowing unit

206:延時調整單元 206: Delay adjustment unit

207:輸出低通濾波器/低通濾波器/四階輸出低通濾波器/二階輸出低通濾波器 207: Output low-pass filter/Low-pass filter/Fourth-order output low-pass filter/Second-order output low-pass filter

900:程序 900:Procedure

901:步驟 901: Steps

902:步驟 902: Steps

903:步驟 903: Steps

904:步驟 904: Steps

905:步驟 905: Steps

906:步驟 906: Steps

907:步驟 907: Steps

908:步驟 908: Steps

1000:程序 1000:Program

1001:步驟 1001: Steps

1002:步驟 1002: Steps

1003:步驟 1003: Steps

1004:步驟 1004: Steps

1005:步驟 1005: Steps

1100:系統 1100:System

1101:中央處理單元 1101: Central processing unit

1102:唯讀記憶體 1102: Read-only memory

1103:隨機存取記憶體 1103: Random Access Memory

1104:匯流排 1104: Bus

1105:輸入/輸出介面 1105: Input/output interface

1109:通信單元 1109: Communication unit

在圖式中,為易於說明起見展示示意性元件(例如,表示裝置、單元、指令區塊及資料元素之元件)之具體配置或次序。然而,熟習此項技術者應理解,圖式中之示意性元件之具體次序或配置並不意在暗示需要一特定處理次序或順序或程序分離。此外,一圖式中包含一示意性元件並不意在暗示在所有實施例中皆需要此等元件,或並不意在暗示由此等元件表示之特徵可不包含於一些實施方案中或與在一些實施方案中之其他元件組合。 In the drawings, the specific configuration or order of schematic elements (e.g., elements representing devices, units, instruction blocks, and data elements) is shown for ease of illustration. However, those skilled in the art will understand that the specific order or configuration of schematic elements in the drawings is not intended to imply the need for a particular processing order or sequence or process separation. Furthermore, the inclusion of a schematic element in a drawing is not intended to imply that such element is required in all embodiments, or that the features represented by such elements may not be included in some embodiments or combined with other elements in some embodiments.

此外,在圖式中,使用連接元素(例如實線或虛線或箭頭)圖解說明兩個或兩個以上其他示意性元件之間或當中的一連接、關係或關聯性,不存在任何此等連接元素並不意在暗示連接、關係或關聯性不可存在。換言之,圖式中未展示元件之間的一些連接、關係或關聯性以免使本發明模糊。另外,為易於圖解說明,使用一單個連接元件來表示元件之間的多個連接、關係或關聯性。舉例而言,在一連接元素表示信號、資料或指令之一通信,熟習此項技術者應理解此等元素視需要表示影響通信之一或多個信號路徑。 In addition, in the drawings, connecting elements (such as solid or dashed lines or arrows) are used to illustrate a connection, relationship or association between or among two or more other schematic elements, and the absence of any such connecting elements is not intended to imply that the connection, relationship or association cannot exist. In other words, some connections, relationships or associations between elements are not shown in the drawings to avoid obscuring the present invention. In addition, for ease of illustration, a single connecting element is used to represent multiple connections, relationships or associations between elements. For example, when a connecting element represents a communication of a signal, data or instruction, those skilled in the art should understand that such elements represent one or more signal paths that affect the communication as needed.

圖1圖解說明根據一或多個實施方案的用於對IVAS及LFE位元串流進行編碼及解碼之一IVAS編碼解碼器。 FIG. 1 illustrates an IVAS codec for encoding and decoding IVAS and LFE bit streams according to one or more implementations.

圖2A係圖解說明根據一或多個實施方案的LFE編碼之一方塊圖。 FIG. 2A is a block diagram illustrating LFE encoding according to one or more implementations.

圖2B係圖解說明根據一或多個實施方案之LFE解碼之一方塊圖。 FIG. 2B is a block diagram illustrating LFE decoding according to one or more implementations.

圖3係圖解說明根據一或多個實施方案的具有130Hz之一拐角一截止點之四階巴特沃斯低通濾波器之一頻率回應之一曲線圖。 FIG. 3 is a graph illustrating a frequency response of a fourth-order Butterworth low-pass filter having a corner cutoff point of 130 Hz according to one or more embodiments.

圖4係圖解說明根據一或多個實施方案之一菲爾德(Fielder)窗口之一曲線圖。 FIG. 4 is a graph illustrating a Fielder window according to one or more implementations.

圖5圖解說明根據一或多個實施方案的精細量化點隨頻率之變化。 FIG5 illustrates the variation of fine quantization points with frequency according to one or more implementations.

圖6圖解說明根據一或多個實施方案的粗略量化點隨頻率之變化。 FIG6 illustrates the variation of coarse quantization points with frequency according to one or more implementations.

圖7圖解說明根據一或多個實施方案的經量化MDCT係數隨精細量化之一概率分佈。 FIG. 7 illustrates a probability distribution of quantized MDCT coefficients with fine quantization according to one or more implementations.

圖8圖解說明根據一或多個實施方案的經量化MDCT係數隨粗略量化之一概率分佈。 FIG8 illustrates a probability distribution of quantized MDCT coefficients with coarse quantization according to one or more implementations.

圖9係根據一或多個實施方案的對修改型離散餘弦變換(MDCT)係數進行編碼之一程序之一流程圖。 FIG. 9 is a flow chart of a process for encoding modified discrete cosine transform (MDCT) coefficients according to one or more implementations.

圖10係根據一或多個實施方案的對修改型離散餘弦變換(MDCT)係數進行解碼之一程序之一流程圖。 FIG. 10 is a flow chart of a process for decoding modified discrete cosine transform (MDCT) coefficients according to one or more implementation schemes.

圖11係根據一或多個實施方案的用於實施參考圖1至圖10所闡述之特徵及程序之一系統之一方塊圖。 FIG. 11 is a block diagram of a system for implementing the features and procedures described with reference to FIGS. 1 to 10 according to one or more implementation schemes.

各個種圖式中使用相同參考符號來指示相似元件。 The same reference symbols are used in the various drawings to indicate similar elements.

在以下詳細說明中,陳述眾多具體細節以提供對各種所闡述實施例之一透徹理解。熟習此項技術者應明瞭,可在不具有此等具體細節之情況下實踐各種所闡述之實施方案。在其他例項中,未詳細闡述眾所 周知之方法、過程、組件及電路以免使實施例之態樣發生不必要模糊。下文闡述可各自彼此獨立地使用或與其他特徵之任何組合使用之數個特徵。 In the following detailed description, numerous specific details are set forth to provide a thorough understanding of one of the various described embodiments. Those skilled in the art will appreciate that the various described embodiments may be practiced without these specific details. In other instances, well-known methods, processes, components, and circuits are not described in detail to avoid unnecessarily obscuring the aspects of the embodiments. Several features are described below that may be used independently of each other or in any combination with other features.

命名法 Nomenclature

如本文中所使用,將術語「包含」及其變化形式解讀為開放式術語,意指「包含但不限於」。將術語「或」解讀為「及/或」,除非內容脈絡另有明確指示。將術語「基於」解讀為「至少部分地基於」。將術語「一項實例性實施方案」及「實例性實施方案」解讀為「至少一項實例性實施方案」。將術語「另一實施方案」解讀為「至少一項其他實施方案」。將術語「經判定」或「判定」解讀為獲得、接收、運算、計算、估計、預測或導出。除非另有定義,否則本文所使用之所有技術術語及科學術語皆具有與熟習本發明所屬領域技術者通常所理解的相同的意義。 As used herein, the term "including" and its variations are interpreted as open-ended terms, meaning "including but not limited to". The term "or" is interpreted as "and/or", unless the context clearly indicates otherwise. The term "based on" is interpreted as "based at least in part on". The terms "an exemplary embodiment" and "exemplary embodiment" are interpreted as "at least one exemplary embodiment". The term "another embodiment" is interpreted as "at least one other embodiment". The term "determined" or "determined" is interpreted as obtaining, receiving, calculating, computing, estimating, predicting or deriving. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which the invention belongs.

系統概述 System Overview

圖1圖解說明根據一或多個實施方案的用於對IVAS位元串流(包含一LFE聲道位元串流)進行編碼及解碼之一IVAS編碼解碼器100。在編碼時,IVAS編碼解碼器100接收音訊資料101之N+1個聲道,其中音訊資料101之N個聲道被輸入至空間分析與降混單元102中,且一個LFE聲道被輸入至LFE聲道編碼單元105中。音訊資料101包括但不限於:單聲道信號、立體聲信號、雙耳信號、空間音訊信號(例如,多聲道空間音訊對象)、一階高保真度立體聲響複製(FoA)、高階高保真度立體聲響複製(HoA)及任何其他音訊資料。 FIG. 1 illustrates an IVAS codec 100 for encoding and decoding an IVAS bitstream (including an LFE channel bitstream) according to one or more embodiments. During encoding, the IVAS codec 100 receives N+1 channels of audio data 101, wherein the N channels of audio data 101 are input to a spatial analysis and downmix unit 102, and one LFE channel is input to an LFE channel encoding unit 105. The audio data 101 includes, but is not limited to: a mono signal, a stereo signal, a binaural signal, a spatial audio signal (e.g., a multi-channel spatial audio object), first-order Ambisonics (FoA), higher-order Ambisonics (HoA), and any other audio data.

在一些實施方案中,空間分析與降混單元102經組態以實施複雜先進耦合(CACPL)以用於對立體聲音訊資料進行分析/降混,及/或實施空間重構(SPAR)以用於對FoA音訊資料進行分析/降混。在其他實施方 案中,空間分析與降混單元102實施其他格式。空間分析與降混單元102之輸出包括空間後設資料及音訊資料之1至N個聲道。將空間後設資料輸入至空間後設資料編碼單元104中,空間後設資料編碼單元104經組態以對空間後設資料進行量化及熵寫碼。在一些實施方案中,量化可包含精細、適度、粗略及額外粗略量化策略,且熵寫碼可包含霍夫曼或算術寫碼。 In some implementations, the spatial analysis and downmix unit 102 is configured to implement complex advanced coupling (CACPL) for analysis/downmixing of stereo audio data, and/or spatial reconstruction (SPAR) for analysis/downmixing of FoA audio data. In other implementations, the spatial analysis and downmix unit 102 implements other formats. The output of the spatial analysis and downmix unit 102 includes spatial metadata and 1 to N channels of audio data. The spatial metadata is input to the spatial metadata encoding unit 104, which is configured to quantize and entropy code the spatial metadata. In some implementations, quantization may include fine, moderate, coarse, and extra-coarse quantization strategies, and entropy coding may include Huffman or arithmetic coding.

將音訊資料之1至N個聲道輸入至主音訊聲道編碼單元103中,主音訊聲道編碼單元103經組態以將音訊資料之1至N個聲道編碼成一或多個經增強聲音服務(EVS)位元串流。在一些實施方案中,主音訊聲道編碼單元103遵照3GPP TS 26.445且提供各種各樣之功能性,例如為窄頻(EVS-NB)及寬頻(EVS-WB)語音服務增強品質與寫碼效率、使用超寬頻(EVS-SWB)語音增強品質、在對話應用中增強混合內容與音樂之品質、封包遺失及延時抖動的穩健性以及與AMR-WB編碼解碼器之反向相容性。 1 to N channels of audio data are input to the main audio channel coding unit 103, which is configured to encode 1 to N channels of audio data into one or more enhanced voice service (EVS) bit streams. In some implementations, the main audio channel coding unit 103 complies with 3GPP TS 26.445 and provides a variety of functionalities, such as enhancing quality and coding efficiency for narrowband (EVS-NB) and wideband (EVS-WB) voice services, enhancing quality of voice using ultra-wideband (EVS-SWB), enhancing the quality of mixed content and music in conversational applications, robustness to packet loss and delay jitter, and backward compatibility with AMR-WB codecs.

在一些實施方案中,主音訊聲道編碼單元103包括一預處理與模式選擇單元,該預處理與模式選擇單元基於模式/位元速率控制以一規定的位元速率在用於對語音信號進行編碼之一語音寫碼器及用於對音訊信號進行編碼之一感知寫碼器之間做出選擇。在一些實施方案中,語音編碼器係代數碼激式線性預測(ACELP)之一改良變化形式,其針對不同語音類別而擴展出專門的基於LP之模式。 In some embodiments, the main audio channel coding unit 103 includes a pre-processing and mode selection unit that selects between a speech codec for encoding speech signals and a perceptual codec for encoding audio signals at a specified bit rate based on mode/bit rate control. In some embodiments, the speech codec is a modified variation of Algebraic Coded Excitation Linear Prediction (ACELP) that is extended with specialized LP-based modes for different speech categories.

在一些實施方案中,音訊編碼器係一修改型離散餘弦變換(MDCT)編碼器,其在低延時/低位元速率下效率得到提高且經設計以在語音編碼器與音訊編碼器之間實行無縫且可靠切換。 In some implementations, the audio codec is a modified discrete cosine transform (MDCT) codec that has improved efficiency at low latency/low bit rates and is designed to perform seamless and reliable switching between the speech codec and the audio codec.

如先前所闡述,LFE聲道信號用於範圍為20Hz至120Hz之深度低音調聲響,且通常發送至經設計以再生低頻率音訊內容之一揚聲器(例如,一低音揚聲器)。將LFE聲道信號輸入至LFE聲道信號編碼單元105中,LFE聲道信號編碼單元105經組態以按照參考圖2A之闡述對LFE聲道信號進行編碼。 As previously explained, the LFE channel signal is used for deep bass sound in the range of 20 Hz to 120 Hz and is typically sent to a speaker (e.g., a subwoofer) designed to reproduce low-frequency audio content. The LFE channel signal is input to the LFE channel signal encoding unit 105, which is configured to encode the LFE channel signal as explained with reference to FIG. 2A.

在一些實施方案中,一IVAS解碼器包括:空間後設資料解碼單元106,其經組態以恢復空間後設資料;及主音訊聲道解碼單元107,其經組態以恢復1至N個聲道音訊信號。將所恢復之空間後設資料及所恢復之1至N個聲道音訊信號輸入至空間合成/升混/呈現單元109中,空間合成/升混/呈現單元109經組態以使用空間後設資料將1至N個聲道音訊信號合成並呈現為N個或N個以上聲道輸出音訊信號以供在各種音訊系統之揚聲器上播放,包含但不限於:家庭影院系統、視訊會議室系統、虛擬實境(VR)裝備及能夠呈現音訊之任何其他音訊系統。LFE聲道解碼單元108接收LFE位元串流且經組態以對LFE位元串流進行解碼,如參考圖2B所闡述。 In some embodiments, an IVAS decoder includes: a spatial metadata decoding unit 106 configured to recover spatial metadata; and a main audio channel decoding unit 107 configured to recover 1 to N channel audio signals. The recovered spatial metadata and the recovered 1 to N channel audio signals are input into a spatial synthesis/upmixing/presentation unit 109, which is configured to use the spatial metadata to synthesize and present the 1 to N channel audio signals into N or more channel output audio signals for playback on speakers of various audio systems, including but not limited to: home theater systems, video conference room systems, virtual reality (VR) equipment, and any other audio system capable of presenting audio. The LFE channel decoding unit 108 receives the LFE bit stream and is configured to decode the LFE bit stream, as described with reference to FIG. 2B .

儘管上文所闡述之LFE編碼/解碼之實例性實施方案係藉由一IVAS編碼解碼器來實行,但下文所闡述之低延遲LFE編碼解碼器可係一獨立LFE編碼解碼器,或其可包含於在需要或期望低延遲及可組態性之音訊應用中對低頻率信號進行編碼及解碼之任何專用或標準化音訊編碼解碼器。 Although the exemplary implementation of LFE encoding/decoding described above is implemented by an IVAS codec, the low-latency LFE codec described below may be a standalone LFE codec, or it may be included in any dedicated or standardized audio codec that encodes and decodes low frequency signals in audio applications where low latency and configurability are required or desired.

圖2A係圖解說明根據一或多個實施例的圖1中所展示之LFE聲道編碼單元105之功能組件之一方塊圖。圖2B係圖解說明根據一或多個實施例的圖1中所展示之LFE聲道解碼器108之功能組件之一方塊圖。LFE 聲道解碼器108包括熵解碼與逆量化單元204、逆MDCT與窗口化單元205、延時調整單元206及輸出LPF 207。延時調整單元206可位於LPF 207之前或之後,且實行延時調整(例如,藉由緩衝經解碼LFE聲道信號)以使經解碼LFE聲道信號與主編碼解碼器經解碼輸出相匹配。在後文中,參考圖2B所闡述之LFE聲道編碼單元105及LFE聲道解碼單元108統稱為一LFE編碼解碼器。 FIG. 2A is a block diagram illustrating functional components of the LFE channel encoding unit 105 shown in FIG. 1 according to one or more embodiments. FIG. 2B is a block diagram illustrating functional components of the LFE channel decoder 108 shown in FIG. 1 according to one or more embodiments. The LFE channel decoder 108 includes an entropy decoding and inverse quantization unit 204, an inverse MDCT and windowing unit 205, a delay adjustment unit 206, and an output LPF 207. The delay adjustment unit 206 may be located before or after the LPF 207 and performs delay adjustment (e.g., by buffering the decoded LFE channel signal) to match the decoded LFE channel signal with the decoded output of the main encoder/decoder. In the following text, the LFE channel encoding unit 105 and the LFE channel decoding unit 108 described in reference to FIG. 2B are collectively referred to as an LFE codec.

LFE聲道編碼單元105包括輸入低通濾波器(LPF)201、窗口化與MDCT單元202以及量化與熵寫碼單元203。在一實施例中,輸入音訊信號係一脈衝碼調變(PCM)音訊信號,且LFE聲道編碼單元105預期一步幅為5毫秒、10毫秒或20毫秒之一輸入音訊信號。內在地,LFE聲道編碼單元105對5毫秒或10毫秒子訊框進行操作,且對此等子訊框之一組合實行窗口化及MDCT。在一實施例中,LFE聲道編碼單元105以一20毫秒輸入步幅運行且內在地將此輸入劃分成相等長度之兩個子訊框。去往LFE之先前輸入訊框之最後子訊框與去往LFE之當前輸入訊框之第一子訊框級聯且被窗口化。去往LFE之當前輸入訊框之第一子訊框與去往LFE之當前輸入訊框之第二子訊框級聯且被窗口化。實行MDCT兩次,每一經窗口化區塊上各一次。 The LFE channel coding unit 105 includes an input low pass filter (LPF) 201, a windowing and MDCT unit 202, and a quantization and entropy coding unit 203. In one embodiment, the input audio signal is a pulse code modulation (PCM) audio signal, and the LFE channel coding unit 105 expects an input audio signal with a step size of 5 ms, 10 ms, or 20 ms. Internally, the LFE channel coding unit 105 operates on 5 ms or 10 ms subframes, and performs windowing and MDCT on a combination of these subframes. In one embodiment, the LFE channel coding unit 105 operates with a 20 ms input step size and internally divides the input into two subframes of equal length. The last subframe of the previous input frame to the LFE is concatenated with the first subframe of the current input frame to the LFE and is windowed. The first subframe of the current input frame to the LFE is concatenated with the second subframe of the current input frame to the LFE and is windowed. The MDCT is performed twice, once on each windowed block.

在一實施例中,演算法延時(不包含成框延時)等於8毫秒+由輸入LPF 103引致之延時+由輸出LPF 207引致之延時。在四階輸入LPF 201及四階輸出LPF 207之情況下,總系統延遲係大約15毫秒。在四階輸入LPF 201及二階輸出LPF 207之情況下,總LFE編碼解碼器延遲大約13毫秒。 In one embodiment, the algorithmic delay (excluding the framing delay) is equal to 8 ms + the delay caused by the input LPF 103 + the delay caused by the output LPF 207. In the case of a four-stage input LPF 201 and a four-stage output LPF 207, the total system delay is about 15 ms. In the case of a four-stage input LPF 201 and a two-stage output LPF 207, the total LFE codec delay is about 13 ms.

圖3係圖解說明根據一或多個實施例的一實例性輸入LPF 201之一頻率回應之一曲線圖。在所展示之實例中,LPF 201係截止頻率為130Hz之四階巴特沃斯濾波器。其他實施例可使用具有相同或不同階數以及相同或不同截止頻率的一不同類型之LPF(例如,契比雪夫(Chebyshev)、貝索(Bessel))。 FIG. 3 is a graph illustrating a frequency response of an exemplary input LPF 201 according to one or more embodiments. In the example shown, LPF 201 is a fourth-order Butterworth filter with a cutoff frequency of 130 Hz. Other embodiments may use a different type of LPF (e.g., Chebyshev, Bessel) with the same or different order and the same or different cutoff frequency.

圖4係圖解說明,根據一或多個實施例之一菲爾德窗口之一曲線圖。在一實施例中,由窗口化與MDCT單元202應用之窗函數係漸隱長度為8毫秒之一菲爾德窗函數。菲爾德窗口係α=5之一凱澤-貝索導出(KBD)窗,其係藉由建構來滿足MDCT之Princen-Bradley條件,且因此以先進音訊寫碼(AAC)數位音訊格式使用的一窗口。亦可使用其他窗函數。 FIG. 4 is a graph illustrating a plot of a Field window according to one or more embodiments. In one embodiment, the window function applied by the windowing and MDCT unit 202 is a Field window function with a gradual concealment length of 8 milliseconds. The Field window is a Kaiser-Besso derived (KBD) window with α=5, which is constructed to satisfy the Princen-Bradley condition of the MDCT and is therefore a window used in the Advanced Audio Coding (AAC) digital audio format. Other window functions may also be used.

量化及熵寫碼 Quantization and entropy coding

在一實施例中,量化與熵寫碼單元203實施符合輸入LPF 201頻率回應曲線之一量化策略以更高效地將MDCT係數量化。在一實施例中,將頻率範圍劃分為表示4個頻帶之4個次頻帶群組:0Hz至100Hz、100Hz至200Hz、200Hz至300Hz及300Hz至400Hz。此等頻帶是實例,且更多或更少的頻帶可與相同或不同頻率範圍一起使用。更確切而言,MDCT係數係使用基於一特定訊框中之MDCT係數值動態地運算之一縮放移位因數加以量化,且依據LPF頻率回應曲線選擇量化點,如圖5至圖8中所展示。此量化策略有助於減少MDCT係數的屬100Hz至200Hz、200Hz至300Hz及300Hz至400Hz頻帶之量化點,而為主LFE頻帶0Hz至100Hz保留最佳量化點,0Hz至100Hz中將存在最低頻率效應能量(例如,隆隆聲)。 In one embodiment, the quantization and entropy coding unit 203 implements a quantization strategy that conforms to the frequency response curve of the input LPF 201 to more efficiently quantize the MDCT coefficients. In one embodiment, the frequency range is divided into 4 sub-band groups representing 4 frequency bands: 0Hz to 100Hz, 100Hz to 200Hz, 200Hz to 300Hz, and 300Hz to 400Hz. These frequency bands are examples, and more or fewer frequency bands can be used with the same or different frequency ranges. More specifically, the MDCT coefficients are quantized using a scaling factor that is dynamically calculated based on the MDCT coefficient values in a particular frame, and the quantization points are selected according to the LPF frequency response curve, as shown in Figures 5 to 8. This quantization strategy helps reduce the quantization points of the MDCT coefficients belonging to the 100Hz to 200Hz, 200Hz to 300Hz, and 300Hz to 400Hz bands, while retaining the best quantization points for the main LFE band 0Hz to 100Hz, where the lowest frequency effect energy (e.g., rumble) will exist.

在一實施例中,下文闡述去往LFE聲道編碼單元105之一F len 毫秒(ms)輸入PCM步幅(輸入訊框長度)之一量化策略,其中訊框長度 F len 可取由5*f ms給定之任何值,在此1<=f<=12。 In one embodiment, a quantization strategy for an F len millisecond (ms) input PCM step (input frame length) to the LFE channel coding unit 105 is described below, where the frame length F len can take any value given by 5* f ms, where 1<= f <=12.

首先,將輸入PCM步幅劃分成相等長度的N個子訊框,每一子訊框寬度(Sw)=F len /N ms。N應經選擇以使得每一S w 係5ms之倍數(舉例而言,若F len =20ms,則N可係1、2或4;若F len =10ms,則N可係1或2;且若F len =5ms,則N等於1)。使Si係任何給定訊框中的第i子訊框,在此i係範圍係0<=i<=N之一整數,其中S 0 對應於去往LFE編碼單元105之先前輸入訊框中之最後子訊框,且S 1 S N 係當前訊框中之N個子訊框。 First, divide the input PCM stride into N subframes of equal length, each subframe width (S w ) = F len /N ms. N should be selected so that each Sw is a multiple of 5 ms (for example, if F len =20 ms, then N can be 1, 2 or 4; if F len =10 ms, then N can be 1 or 2; and if F len =5 ms, then N is equal to 1). Let S i be the i- th subframe in any given frame, where i is an integer in the range 0 <= i <= N , where S 0 corresponds to the last subframe in the previous input frame to the LFE coding unit 105, and S 1 to S N are the N subframes in the current frame.

接下來,將每一S i S i+1 子訊框級聯且利用一菲爾德窗口(參見圖4)來窗口化,且然後對此等經窗口化樣本實行MDCT。此使得對每一訊框進行總共N次MDCT。來自每一MDCT之MDCT係數之數目(num_coeffs)=取樣頻率*S w /1000。每一MDCT之頻率解析度(每一MDCT係數之寬度)(W mdct )係大約1000/(2*S w )Hz。鑒於低音揚聲器通常具有大約100Hz至120Hz之一LPF截止點,且在400Hz之後的後LPF能量通常非常低,將高達400Hz的MDCT係數量化且發送至LFE解碼單元108,而將MDCT係數之其餘部分量化為0。發送高達400Hz的MDCT係數確保在LFE解碼單元108處高達120Hz的高品質重構。因此,用於量化及碼(N quant )之MDCT係數之總數目等於N*400/W mdct Next, each Si and Si + 1 subframe is concatenated and windowed using a Field window (see FIG. 4 ), and then MDCT is performed on these windowed samples. This results in a total of N MDCTs being performed on each frame. The number of MDCT coefficients from each MDCT ( num_coeffs ) = sampling frequency * Sw /1000. The frequency resolution of each MDCT (width of each MDCT coefficient) (Wmdct ) is approximately 1000/(2* Sw ) Hz . Given that subwoofers typically have an LPF cutoff point of approximately 100 Hz to 120 Hz, and the post-LPF energy after 400 Hz is typically very low, the MDCT coefficients up to 400 Hz are quantized and sent to the LFE decoding unit 108, while the remainder of the MDCT coefficients are quantized to zero. Sending MDCT coefficients up to 400 Hz ensures high quality reconstruction up to 120 Hz at the LFE decoding unit 108. Therefore, the total number of MDCT coefficients used for quantization and coding ( Nquant ) is equal to N*400/Wmdct .

接下來,將MDCT係數配置於M個次頻帶群組中,其中每一次頻帶群組之寬度係W mdct 之一倍數且所有次頻帶群組之寬度之和等於400Hz。使每一次頻帶之寬度係SBWm Hz,其中m係範圍為1<=m<=M之一整數。在此寬度下,第m次頻帶群組中之係數之數目=SN quant =N*SBW m /W mdct (即,來自每一MDCT之SBW m /W mdct 係數)。然後,根據下 文所闡述之一移位縮放因數(shift)來縮放每一次頻帶群組中之MDCT係數,該移位縮放因數係由所有N quant MDCT係數之絕對值之和或最大值確定。然後,在編碼器輸入處使用符合LPF曲線之一量化方案單獨地將每一次頻帶群組中之縮放MDCT係數量化及寫碼。利用一熵寫碼器(例如,一算術或霍夫曼寫碼器)對經量化MDCT係數進行寫碼。利用一不同熵寫碼器來對每一次頻帶群組進行寫碼,且每一熵寫碼器使用一恰當概率分佈模型來高效地對各別次頻帶群組進行寫碼。 Next, the MDCT coefficients are arranged in M sub-band groups, where the width of each sub-band group is a multiple of W mdct and the sum of the widths of all sub-band groups is equal to 400 Hz. Let the width of each sub-band be SBW m Hz, where m is an integer in the range 1 <= m <= M. At this width, the number of coefficients in the mth sub-band group = SN quant = N * SBW m /W mdct (i.e., the SBW m /W mdct coefficients from each MDCT). The MDCT coefficients in each sub-band group are then scaled according to a shift scaling factor (shift) as explained below, which is determined by the sum or maximum of the absolute values of all N quant MDCT coefficients. The scaled MDCT coefficients in each sub-band group are then quantized and coded individually at the encoder input using a quantization scheme that conforms to the LPF curve. The quantized MDCT coefficients are coded using an entropy coder (e.g., an arithmetic or Huffman coder). Each sub-band group is coded using a different entropy coder, and each entropy coder uses an appropriate probability distribution model to efficiently code the respective sub-band group.

現在將闡述具有一20毫秒(ms)之步幅(F len =20ms)、2個子訊框(N=2)且取樣頻率=48000之一實例性量化策略。在此實例性輸入組態下,子訊框寬度S w =10ms且MDCT之數目=N=2。對一20ms區塊實行第一MDCT。藉由將先前20ms輸入中之一10ms至20ms子訊框與當前20ms輸入中之一0ms至10ms子訊框級聯來形成此區塊,且然後對20ms長之菲爾德窗口進行窗口化(參見圖4)。在N=1且N=4之情況下,相應地對菲爾德窗口進行縮放且漸隱長度改變為16/N ms。對藉由利用一20ms長菲爾德窗口給當前20ms輸入訊框窗口化所形成之一20ms區塊實行第二MDCT。每一MDCT的MDCT係數數目(num_coeffs)=480,每一MDCT係數之寬度W mdct =50Hz,量化及寫碼之係數之總數目N quant =16,且量化及寫碼之係數之總數目/MDCT=16/N=8。 An exemplary quantization strategy with a 20 ms stride ( F len = 20 ms), 2 subframes ( N = 2) and sampling frequency = 48000 will now be described. In this exemplary input configuration, the subframe width Sw = 10 ms and the number of MDCTs = N = 2. The first MDCT is performed on a 20 ms block. This block is formed by concatenating a 10 ms to 20 ms subframe in the previous 20 ms input with a 0 ms to 10 ms subframe in the current 20 ms input, and then windowing a 20 ms long Field window (see FIG. 4 ). In the case of N = 1 and N = 4, the Field window is scaled accordingly and the gradient length is changed to 16/N ms. The second MDCT is performed on a 20ms block formed by windowing the current 20ms input frame using a 20ms long Field window. The number of MDCT coefficients per MDCT ( num_coeffs ) = 480, the width of each MDCT coefficient Wmdct = 50Hz , the total number of quantized and coded coefficients Nquant = 16, and the total number of quantized and coded coefficients/MDCT = 16/N = 8.

接下來,將MDCT係數配置於4個次頻帶群組(M=4)中,其中每一次頻帶群組對應於一100Hz頻帶(0至100、100至200、200至300、300至400、SBW m =100Hz,每一次頻帶群組中之係數之數目=SN quant =N*SBW m /W mdct =4)。使a1、a2、a3、a4、a5、a6、a7、a8作為將自第一MDCT量化之前8個MDCT係數,且b1、b2、b3、b4、b5、b6、b7、b8 作為將自第二MDCT量化之前8個MDCT係數。4個次頻帶群組經配置以具有以下係數:次頻帶群組1={a1,a2,b1,b2},次頻帶群組2={a3,a4,b3,b4},次頻帶群組3={a5,a6,b5,b6},次頻帶群組4={a7,a8,b7,b8},其中每一次頻帶群組對應於一100Hz頻帶。 Next, the MDCT coefficients are arranged in 4 sub-band groups ( M =4), where each sub-band group corresponds to a 100 Hz band (0 to 100, 100 to 200, 200 to 300, 300 to 400, SBWm = 100Hz , the number of coefficients in each sub- band group = SNquant = N * SBWm / Wmdct = 4). Let a1 , a2 , a3 , a4 , a5 , a6 , a7 , a8 be the first 8 MDCT coefficients to be quantized from the first MDCT, and b1 , b2 , b3 , b4 , b5 , b6 , b7 , b8 be the first 8 MDCT coefficients to be quantized from the second MDCT. The four sub-band groups are configured to have the following coefficients: sub-band group 1 = {a 1 , a 2 , b 1 , b 2 }, sub-band group 2 = {a 3 , a 4 , b 3 , b 4 }, sub-band group 3 = {a 5 , a 6 , b 5 , b 6 }, sub-band group 4 = {a 7 , a 8 , b 7 , b 8 }, where each sub-band group corresponds to a 100 Hz frequency band.

一增益為大約-30dB(或小於-30dB)之訊框可具有值為大於10-2或10-1或更低之MDCT係數,而具有滿量程增益之一訊框可具有值為20或高於20之MDCT係數。為滿足此寬的值範圍,基於可用之最大量化點(max_value)及MDCT係數(lfe_dct_new)之絕對值之一和來運算一縮放移位因數(shift),如下:shift=floor(shifts_per_double*log2(max_value/sum(abs(lfe_dct_new))))。 A frame with a gain of approximately -30 dB (or less than -30 dB) may have MDCT coefficients with values greater than 10-2 or 10-1 or less, while a frame with full -scale gain may have MDCT coefficients with values of 20 or more. To accommodate this wide range of values, a scaling shift factor (shift ) is calculated based on the maximum available quantization point ( max_value ) and a sum of the absolute values of the MDCT coefficients ( lfe_dct_new ), as follows : shift = floor ( shift s_per_double *log2( max_value / sum(abs (lfe_dct_new ) ))).

在一實施方案中,lfe_dct_new係16個MDCT係數之一陣列,shifts_per_double係一常數(例如4),max_value係為精細量化(例如,63個量化值)及粗略量化(例如,31個量化值)而選擇之一整數,且在精細量化時移位僅限於自4至35之5位元值,且在粗略量化時移位僅限於2至33之5位元值。 In one embodiment, lfe_dct_new is an array of 16 MDCT coefficients, shifts_per_double is a constant (e.g., 4 ) , max_value is an integer selected for fine quantization (e.g., 63 quantization values) and coarse quantization ( e.g., 31 quantization values), and shifts are limited to 5-bit values from 4 to 35 in fine quantization and 2 to 33 in coarse quantization.

然後,如下運算經量化MDCT係數:vals=round(lfe_dct_new*(2^(shift/shifts_per_double))),其中round( )運算將結果四捨五入至最接近整數值。 The quantized MDCT coefficients are then calculated as follows: vals = round( lfe _ dct _ new *(2^( shift/shifts _ per _ double ))), where the round() operation rounds the result to the nearest integer value.

若經量化值(vals)超過最大允許的可用量化點數目 (max_val),則減小縮放移位因數(shift)且再次計算經量化值(vals)。在其他實施方案中,代替和函數sum(abs(lfe_dct_new))),可使用最大值函數max(abs(lfe_dct_new)))來運算縮放移位因數(shift),但使用max( )函數將將量化值更分散,而使得設計一高效的熵寫碼器更困難。 If the quantized value ( vals ) exceeds the maximum allowed number of available quantization points ( max_val ), the scaling shift factor ( shift ) is reduced and the quantized value ( vals ) is calculated again. In other implementation schemes, instead of the sum function sum ( abs ( lfe_dct_new ))), the maximum function max(abs( lfe_dct_new ) )) can be used to calculate the scaling shift factor ( shift ), but using the max() function will make the quantized values more dispersed, making it more difficult to design an efficient entropy coder.

在上文所闡述之量化步驟中,在一個迴路中一起計算每一次頻帶群組之經量化值,但每一次頻帶群組之量化點係不同的。若第一次頻帶群組超過允許的範圍,則減小縮放移位因數。若其他次頻帶群組中之任一者超過允許的範圍,則將該次頻帶群組刪減為max_value。針對每一次頻帶群組單獨地對所有MDCT係數之正負號位元及經量化MDCT係數之絕對值進行寫碼。 In the quantization step described above, the quantized values of each sub-band grouping are calculated together in one loop, but the quantization point is different for each sub-band grouping. If the first sub-band grouping exceeds the allowed range, the scaling factor is reduced. If any of the other sub-band groups exceeds the allowed range, the sub-band grouping is deleted to max_value . The sign bits of all MDCT coefficients and the absolute value of the quantized MDCT coefficients are coded separately for each sub-band grouping.

圖5圖解說明根據一或多個實施方案精細量化點隨頻率之變化。在精細量化時,次頻帶群組1(0Hz至100Hz)具有64個量化點,次頻帶群組2(100Hz至200Hz)具有32個量化點,次頻帶群組3(200Hz至300Hz)具有8個量化點且次頻帶群組4(300Hz至400Hz)具有2個量化點。在一實施例中,利用熵寫碼器(例如,一算術或霍夫曼熵寫碼器)對每一次頻帶群組進行熵寫碼,其中每一熵寫碼器使用一不同概率分佈。因此,主0Hz至100Hz範圍被分配的量化點最多。 FIG. 5 illustrates the variation of fine quantization points with frequency according to one or more embodiments. In fine quantization, subband group 1 (0 Hz to 100 Hz) has 64 quantization points, subband group 2 (100 Hz to 200 Hz) has 32 quantization points, subband group 3 (200 Hz to 300 Hz) has 8 quantization points, and subband group 4 (300 Hz to 400 Hz) has 2 quantization points. In one embodiment, each subband group is entropy-coded using an entropy coder (e.g., an arithmetic or Huffman entropy coder), where each entropy coder uses a different probability distribution. Therefore, the main 0 Hz to 100 Hz range is assigned the most quantization points.

注意,為次頻帶群組1至次頻帶群組4分配量化點係遵循LPF頻率回應曲線之形狀,該LPF頻率回應曲線在較低頻率中所具有之資訊多於在較高頻率中之資訊,且在截止頻率之外無資訊。為正確地重構高達130Hz之頻率,亦對與高於130Hz之頻率對應之MDCT係數進行編碼以避免或最小化頻疊。在一些實施方案中,對高達400Hz之MDCT係數進行編碼以使得可在解碼單元處恰當地重構高達130Hz之頻率。 Note that the assignment of quantization points to subband group 1 to subband group 4 follows the shape of the LPF frequency response curve, which has more information in lower frequencies than in higher frequencies and no information outside the cutoff frequency. To properly reconstruct frequencies up to 130 Hz, the MDCT coefficients corresponding to frequencies above 130 Hz are also encoded to avoid or minimize frequency overlap. In some implementations, MDCT coefficients up to 400 Hz are encoded so that frequencies up to 130 Hz can be properly reconstructed at the decoding unit.

圖6圖解說明根據一或多個實施方案粗略量化點隨頻率之變化。在粗略量化時,次頻帶群組1(0Hz至100Hz)具有32個量化點,次頻帶群組2(100Hz至200Hz)具有16個量化點,次頻帶群組3(200Hz至300Hz)具有4個量化點且次頻帶群組4(300Hz至400Hz)未經量化及熵寫碼。在一實施例中,利用使用一不同概率分佈之一單獨熵寫碼器來對每一次頻帶群組進行熵寫碼。 FIG6 illustrates the variation of coarse quantization points with frequency according to one or more implementations. In coarse quantization, subband group 1 (0 Hz to 100 Hz) has 32 quantization points, subband group 2 (100 Hz to 200 Hz) has 16 quantization points, subband group 3 (200 Hz to 300 Hz) has 4 quantization points, and subband group 4 (300 Hz to 400 Hz) is not quantized and entropy coded. In one embodiment, entropy coding is performed on each subband group using a separate entropy coder using a different probability distribution.

圖7圖解說明根據一或多個實施方案在精細量化時經量化MDCT係數之一概率分佈。y軸係出現頻率且x軸係量化點數目。Sg1係與0Hz至100Hz頻帶中之經量化MDCT係數對應之次頻帶群組1,Sg2係與100Hz至200Hz頻帶中之經量化MDCT係數對應之次頻帶群組2。Sg3係與200Hz至300Hz頻帶中之經量化MDCT係數對應之次頻帶群組3。Sg4係與頻帶300Hz至400Hz中之經量化MDCT係數對應之次頻帶群組4。 FIG. 7 illustrates a probability distribution of quantized MDCT coefficients in fine quantization according to one or more implementations. The y-axis is the frequency of occurrence and the x-axis is the number of quantization points. Sg1 is subband group 1 corresponding to quantized MDCT coefficients in the 0 Hz to 100 Hz band, Sg2 is subband group 2 corresponding to quantized MDCT coefficients in the 100 Hz to 200 Hz band. Sg3 is subband group 3 corresponding to quantized MDCT coefficients in the 200 Hz to 300 Hz band. Sg4 is subband group 4 corresponding to quantized MDCT coefficients in the 300 Hz to 400 Hz band.

圖8圖解說明根據一或多個實施方案在粗略量化時經量化MDCT係數之一概率分佈。y軸係出現頻率且x軸係量化點數目。Sg1係與0Hz至100Hz頻帶中之經量化MDCT係數對應之次頻帶群組1,Sg2係與100Hz至200Hz頻帶中之經量化MDCT係數對應之次頻帶群組2。Sg3係與200Hz至300Hz頻帶中之經量化MDCT係數對應之次頻帶群組3。Sg4係與頻帶300Hz至400Hz中之經量化MDCT係數對應之次頻帶群組4。 FIG8 illustrates a probability distribution of quantized MDCT coefficients during coarse quantization according to one or more implementations. The y-axis is the frequency of occurrence and the x-axis is the number of quantization points. Sg1 is subband group 1 corresponding to quantized MDCT coefficients in the 0 Hz to 100 Hz band, Sg2 is subband group 2 corresponding to quantized MDCT coefficients in the 100 Hz to 200 Hz band. Sg3 is subband group 3 corresponding to quantized MDCT coefficients in the 200 Hz to 300 Hz band. Sg4 is subband group 4 corresponding to quantized MDCT coefficients in the 300 Hz to 400 Hz band.

注意,主頻帶(0Hz至100Hz)係發現LFE效應最多的頻帶且因此分配更多量化點以達到更大解析度。然而,在粗略量化中分配給主頻帶之位元比精細量化少。在一實施例中,針對MDCT係數之一訊框是使用精細量化還是粗略量化取決於由主音訊聲道編碼器103設定之所期望目標位元速率。主音訊聲道編碼器103在初始化期間一次性設定此值,或基 於對每一訊框中之主音訊聲道進行編碼所需或所使用之位元而逐訊框地動態設定此值。 Note that the main band (0 Hz to 100 Hz) is the band where the LFE effect is found most and therefore more quantization points are allocated to achieve greater resolution. However, fewer bits are allocated to the main band in coarse quantization than fine quantization. In one embodiment, whether fine quantization or coarse quantization is used for a frame of MDCT coefficients depends on the desired target bit rate set by the main audio channel encoder 103. The main audio channel encoder 103 sets this value once during initialization or dynamically on a frame-by-frame basis based on the bits required or used to encode the main audio channel in each frame.

靜寂訊框 Silent message frame

在一些實施方案中,在LFE聲道位元串流中添加一信號以指示靜寂訊框。一靜寂訊框係具有低於一所規定臨限值之能量之一訊框。在一些實施方案中,將1位元包含於傳輸至解碼器之LFE聲道位元串流中(例如,插入於訊框標頭中)以指示一靜寂訊框,且將LFE聲道位元串流中之所有MDCT係數設定為0。在靜寂訊框期間此技術可將位元速率減小至50bps。 In some implementations, a signal is added to the LFE channel bitstream to indicate a silent frame. A silent frame is a frame with energy below a specified threshold. In some implementations, 1 bit is included in the LFE channel bitstream transmitted to the decoder (e.g., inserted in a frame header) to indicate a silent frame, and all MDCT coefficients in the LFE channel bitstream are set to 0. This technique can reduce the bit rate to 50bps during silent frames.

解碼器LPF Decoder LPF

在LFE聲道解碼單元108之輸出處提供實施LPF 207(參見圖2B)之兩個選項。基於可用延時(其他音訊聲道之總延時減去LFE漸隱延時減去輸入LPF延時)選擇LPF 207。注意,預期由主音訊聲道編碼單元103/主音訊聲道解碼單元107對其他聲道進行編碼/解碼,且該等聲道之延時取決於主音訊聲道編碼單元103/主音訊聲道解碼單元107之演算法延時。 Two options for implementing LPF 207 (see FIG. 2B ) are provided at the output of the LFE channel decoding unit 108. LPF 207 is selected based on the available delay (total delay of other audio channels minus LFE fade delay minus input LPF delay). Note that other channels are expected to be encoded/decoded by the main audio channel encoding unit 103/main audio channel decoding unit 107 and their delays depend on the algorithmic delay of the main audio channel encoding unit 103/main audio channel decoding unit 107.

在一實施方案中,若可用延時小於3.5ms,則使用截止點為130Hz之二階巴特沃斯LPF;否則使用截止點為130Hz之四階巴特沃斯LPF。因此,在LFE聲道解碼單元108處,需在移除在截止頻率以外的疊頻能量與演算法延時之間做出折衷。在一些實施方案中,可完全移除LPF 207,此乃因低音揚聲器通常具有一LPF。LPF 20有助於減小LFE解碼器輸出自身的截止點以外的疊頻能量,且可有助於高效後處理。 In one embodiment, if the available delay is less than 3.5ms, a second-order Butterworth LPF with a cutoff point of 130Hz is used; otherwise, a fourth-order Butterworth LPF with a cutoff point of 130Hz is used. Therefore, at the LFE channel decoding unit 108, a trade-off needs to be made between removing the folded frequency energy outside the cutoff frequency and the algorithm delay. In some embodiments, LPF 207 can be completely removed because the subwoofer usually has an LPF. LPF 20 helps to reduce the folded frequency energy outside the cutoff point of the LFE decoder output itself, and can help efficient post-processing.

實例性程序 Example procedures

圖9係根據一或多個實施方案的對MDCT係數進行編碼之一程序900之一流程圖。可使用例如參考圖11所闡述之系統1100實施程序900。 FIG. 9 is a flow chart of a process 900 for encoding MDCT coefficients according to one or more implementations. Process 900 may be implemented using, for example, system 1100 as described with reference to FIG. 11 .

程序900包含以下步驟:接收一時域LFE聲道信號(901);使用一低通濾波器對該時域LFE聲道信號進行濾波(902);將經濾波時域LFE聲道信號轉換成LFE聲道信號的包括表示LFE聲道信號之一頻譜之一定數目個係數之一頻域表示(903);將係數配置至與LFE聲道信號之不同頻帶對應之一定數目個次頻帶群組中(904);根據低通濾波器之一頻率回應曲線使用縮放移位因數將每一次頻帶群組中之係數量化(905);使用針對次頻帶群組組態之一熵寫碼器對每一次頻帶群組中之經量化係數進行編碼(906);產生包含經編碼之經量化係數之一位元串流(907);及將位元串流儲存於一儲存裝置上或將位元串流串流傳輸至一下游裝置(908)。 The process 900 includes the following steps: receiving a time domain LFE channel signal (901); filtering the time domain LFE channel signal using a low pass filter (902); converting the filtered time domain LFE channel signal into a frequency domain representation of the LFE channel signal including a certain number of coefficients representing a spectrum of the LFE channel signal (903); allocating the coefficients to a certain number of sub-band groups corresponding to different frequency bands of the LFE channel signal (904); 904); quantizing the coefficients in each sub-band group using a scaling factor according to a frequency response curve of a low pass filter (905); encoding the quantized coefficients in each sub-band group using an entropy encoder configured for the sub-band group (906); generating a bit stream including the encoded quantized coefficients (907); and storing the bit stream on a storage device or transmitting the bit stream to a downstream device (908).

圖10係根據一或多個實施方案的對MDCT係數進行解碼之一程序1000之一流程圖。可使用例如參考圖11所闡述之系統1100來實施程序1000。 FIG. 10 is a flow chart of a process 1000 for decoding MDCT coefficients according to one or more implementations. Process 1000 may be implemented using, for example, system 1100 as described with reference to FIG. 11 .

程序1000包含以下步驟:接收一LFE聲道位元串流(1001),其中LFE聲道位元串流包括表示一時域LFE聲道信號之一頻譜之經熵寫碼係數;對係數進行解碼及逆量化(1002),其中係數係使用一縮放移位因數根據一低通濾波器之一頻率回應曲線在與不同頻帶對應之次頻帶群組中被量化;將經解碼且經逆量化係數轉換成一時域LFE聲道信號(1003);調整時域LFE聲道信號之一延時(1004);及使用一低通濾波器對經延時調整之LFE聲道信號進行濾波(1005)。在一實施例中,可基於可自用於對包括時域LFE聲道信號之一多聲道音訊信號之全頻帶寬度聲道進行 編碼/解碼之一主編碼解碼器得到之一總演算法延時來對低通濾波器之階數進行組態。在一些實施方案中,解碼單元僅需要獲悉編碼單元是利用精細量化還是粗略量化對MDCT係數進行編碼即可。可使用LFE位元串流標頭中之一位元或任何其他適合的傳信機制來指示量化類型。 Process 1000 includes the following steps: receiving an LFE channel bit stream (1001), wherein the LFE channel bit stream includes entropy coded coefficients representing a spectrum of a time-domain LFE channel signal; decoding and inverse quantizing the coefficients (1002), wherein the coefficients are quantized in sub-band groups corresponding to different frequency bands using a scale shift factor according to a frequency response curve of a low-pass filter; converting the decoded and inverse quantized coefficients into a time-domain LFE channel signal (1003); adjusting a delay of the time-domain LFE channel signal (1004); and filtering the delay-adjusted LFE channel signal using a low-pass filter (1005). In one embodiment, the order of the low pass filter may be configured based on a total algorithmic delay available from a main codec for encoding/decoding the full bandwidth channels of a multi-channel audio signal including a time domain LFE channel signal. In some embodiments, the decoding unit only needs to know whether the coding unit encodes the MDCT coefficients using fine quantization or coarse quantization. The quantization type may be indicated using a bit in the LFE bit stream header or any other suitable signaling mechanism.

在一些實施方案中,按照如下方式實行經逆量化係數至時域PCM樣本之解碼。將每一次頻帶群組中之經逆量化係數重新配置至N個群組中(N是在編碼單元處運算之MDCT之數目),其中每一群組具有與各別MDCT對應之係數。根據上文所闡述之實例性實施方案,編碼單元對以下4個次頻帶群組進行編碼:次頻帶群組1={a1,a2,b1,b2},次頻帶群組2={a3,a4,b3,b4},次頻帶群組3={a5,a6,b5,b6},次頻帶群組4={a7,a8,b7,b8}。 In some implementations, decoding of the inverse quantized coefficients into time domain PCM samples is performed as follows: The inverse quantized coefficients in each sub-band grouping are re-arranged into N groups (N being the number of MDCTs operated at the coding unit), where each group has coefficients corresponding to a respective MDCT. According to the exemplary implementation scheme described above, the encoding unit encodes the following four sub-band groups: sub-band group 1 = {a 1 , a 2 , b 1 , b 2 }, sub-band group 2 = {a 3 , a 4 , b 3 , b 4 }, sub-band group 3 = {a 5 , a 6 , b 5 , b 6 }, sub-band group 4 = {a 7 , a 8 , b 7 , b 8 }.

解碼單元對4個次頻帶群組進行解碼且將其重新配置回至{a1,a2,a3,a4,a5,a6,a7,a8}及{b1,b2,b3,b4,b5,b6,b7,b8},且然後將零填補到群組以得到所期望之逆MDCT(iMDCT)輸入長度。實行N次iMDCT以將每一群組中之MDCT係數逆變換成時域區塊。在此實例中,每一區塊係2*Sw ms寬,其中Sw係上文所界定之子訊框寬度。接下來,由圖4中所展示之LFE編碼單元使用之同一菲爾德窗口來給此區塊窗口化。藉由恰當地疊加先前iMDCT輸出之經窗口化資料與當前iMDCT輸出來重新建構每一子訊框Si(i係1<=i<=N之間的一整數)。最後,藉由級聯所有N個子訊框來重新建構(1003)之輸出。 The decoding unit decodes and reconfigures the 4 subband groups back to {a 1 ,a 2 ,a 3 ,a 4 ,a 5 ,a 6 ,a 7 ,a 8 } and {b 1 ,b 2 ,b 3 ,b 4 ,b 5 ,b 6 ,b 7 ,b 8 } and then pads the groups with zeros to get the desired inverse MDCT (iMDCT) input length. The iMDCT is performed N times to inverse transform the MDCT coefficients in each group into time domain blocks. In this example, each block is 2*Sw ms wide, where Sw is the subframe width defined above. Next, this block is windowed by the same Field window used by the LFE encoding unit shown in Figure 4. Each subframe S i (i is an integer between 1 <= i <= N ) is reconstructed by appropriately superimposing the windowed data of the previous iMDCT output with the current iMDCT output. Finally, the output of (1003) is reconstructed by concatenating all N subframes.

實例性系統架構 Example system architecture

圖11係根據一或多個實施方案的用於實施參考圖1至圖10所闡述之特徵及程序之一系統1100之一方塊圖。系統1100包括一或多個伺服器電腦或任何用戶端裝置,包含但不限於:叫用伺服器、使用者裝備、會議室系統、家庭影院系統、虛擬實境(VR)裝備及沉浸式內容攝取裝置。系統1100包括任何消費型裝置,包含但不限於:智慧型電話、平板電腦、可穿戴電腦、車輛電腦、遊戲主控台、環繞式系統、資訊站等。 FIG. 11 is a block diagram of a system 1100 for implementing the features and procedures described in reference to FIG. 1 to FIG. 10 according to one or more implementation schemes. System 1100 includes one or more server computers or any client devices, including but not limited to: call servers, user equipment, conference room systems, home theater systems, virtual reality (VR) equipment, and immersive content capture devices. System 1100 includes any consumer device, including but not limited to: smart phones, tablet computers, wearable computers, car computers, game consoles, surround systems, information stations, etc.

如所展示,系統1100包括一中央處理單元(CPU)1101,中央處理單元1101能夠根據儲存於例如一唯讀記憶體(ROM)1102中之一程式或自例如一儲存單元1108載入至一隨機存取記憶體(RAM)1103之一程式實行各種程序。在RAM 1103中亦視需要儲存當CPU 1101實行各種程序時所需之資料。CPU 1101、ROM 1102及RAM 1103經由一匯流排1104彼此連接。一輸入/輸出(I/O)介面1105亦連接至匯流排1104。 As shown, the system 1100 includes a central processing unit (CPU) 1101, which can execute various programs according to a program stored in, for example, a read-only memory (ROM) 1102 or a program loaded from, for example, a storage unit 1108 to a random access memory (RAM) 1103. Data required when the CPU 1101 executes various programs is also stored in RAM 1103 as needed. CPU 1101, ROM 1102, and RAM 1103 are connected to each other via a bus 1104. An input/output (I/O) interface 1105 is also connected to bus 1104.

以下組件連接至I/O介面1105:一輸入單元1106,其可包含一鍵盤、一滑鼠等;一輸出單元1107,其可包含一顯示器,例如一液晶顯示器(LCD)及一或多個揚聲器;儲存單元1108,其包含一硬碟或另一適合的儲存裝置;及一通信單元1109,其包含一網路介面卡,例如一網路卡(例如,有線或無線)。 The following components are connected to the I/O interface 1105: an input unit 1106, which may include a keyboard, a mouse, etc.; an output unit 1107, which may include a display, such as a liquid crystal display (LCD) and one or more speakers; a storage unit 1108, which includes a hard disk or another suitable storage device; and a communication unit 1109, which includes a network interface card, such as a network card (e.g., wired or wireless).

在一些實施方案中,輸入單元1106包括處於不同位置中(根據主機裝置)且能夠以各種格式擷取音訊信號(例如,單聲道、立體聲、空間、沉浸式及其他適合的格式)之一或多個麥克風。 In some implementations, input unit 1106 includes one or more microphones located in different locations (depending on the host device) and capable of capturing audio signals in various formats (e.g., mono, stereo, spatial, immersive, and other suitable formats).

在一些實施方案中,輸出單元1107包含具有各種數目個揚聲器之系統。輸出單元1107(根據主機裝置之能力)可以各種格式(例如,單聲道、立體聲、沉浸式、雙聲道及其他適合的格式)呈現音訊信號。 In some implementations, output unit 1107 includes a system having a varying number of speakers. Output unit 1107 (depending on the capabilities of the host device) can present audio signals in a variety of formats (e.g., mono, stereo, immersive, binaural, and other suitable formats).

通信單元1109經組態以與其他裝置進行通信(例如,經由一網路)。一磁碟機1110亦視需要連接至I/O介面1105。一可抽換式媒體1111(例如一磁碟、一光碟、一磁光碟、一快閃磁碟機或另一適合可抽換式媒體)安裝於磁碟機1110上,使得視需要將自可抽換式媒體1111讀取之一電腦程式安裝至儲存單元1108中。熟習此項技術者應理解儘管系統1100被闡述為包含上文所闡述之組件,但在實際應用中,可添加、移除及/或替換此等組件中之一些組件且所有此等修改或更改全部處於本發明之範疇內。 The communication unit 1109 is configured to communicate with other devices (e.g., via a network). A disk drive 1110 is also connected to the I/O interface 1105 as needed. A removable medium 1111 (e.g., a disk, an optical disk, a magneto-optical disk, a flash disk drive, or another suitable removable medium) is installed on the disk drive 1110 so that a computer program read from the removable medium 1111 can be installed into the storage unit 1108 as needed. Those skilled in the art should understand that although the system 1100 is described as including the components described above, in actual applications, some of these components may be added, removed, and/or replaced and all such modifications or changes are within the scope of the present invention.

根據本發明之實例性實施例,上文所闡述之程序可被實施為電腦軟體程式或實施於一電腦可讀儲存媒體上。舉例而言,本發明之實施例包含一電腦程式產品,該電腦程式產品包含有形地體現於一機器可讀媒體上之一電腦程式,該電腦程式包含用於實行方法之程式碼。在此等實施例中,電腦程式可經由通信單元1109自網路下載且安裝,及/或自可抽換式媒體1111安裝。 According to exemplary embodiments of the present invention, the procedures described above may be implemented as a computer software program or implemented on a computer-readable storage medium. For example, an embodiment of the present invention includes a computer program product, which includes a computer program tangibly embodied on a machine-readable medium, and the computer program includes a program code for implementing the method. In these embodiments, the computer program can be downloaded and installed from the network via the communication unit 1109, and/or installed from the removable medium 1111.

通常,本發明之各種實例性實施例可被實施為硬體或特殊用途電路(例如,控制電路系統)、軟體、邏輯或其任何組合。舉例而言,上文所論述之單元可由控制電路系統(例如,與圖11之其他組件組合之一CPU)執行,因此控制電路系統可正在實行本發明中所闡述之動作。一些態樣可被實施為硬體,而其他態樣可被實施為可由一控制器、微處理器或其他運算裝置(例如,控制電路系統)執行之韌體或軟體。雖然本發明之實例性實施例之各項態樣可被圖解說明且闡述為方塊圖、流程圖或使用某一其他圖形表示,但應瞭解,本文中所闡述之這些方塊、設備、系統、技術或方法可在(作為非限制性實例)硬體、軟體、韌體、特殊用途電路或邏 輯、一般用途硬體或控制器或者其他計算裝置或其某一組合中實施。 In general, various exemplary embodiments of the present invention may be implemented as hardware or special purpose circuits (e.g., control circuit systems), software, logic, or any combination thereof. For example, the units discussed above may be executed by a control circuit system (e.g., a CPU in combination with other components of FIG. 11 ), so that the control circuit system may be performing the actions described in the present invention. Some aspects may be implemented as hardware, while other aspects may be implemented as firmware or software that may be executed by a controller, microprocessor, or other computing device (e.g., control circuit system). Although various aspects of exemplary embodiments of the present invention may be illustrated and described as block diagrams, flow charts, or using some other graphical representation, it should be understood that these blocks, devices, systems, techniques, or methods described herein may be implemented in (as non-limiting examples) hardware, software, firmware, special purpose circuits or logic, general purpose hardware or controllers, or other computing devices, or some combination thereof.

另外,流程圖中所展示之各種區塊可被視為方法步驟及/或由電腦程式碼之操作達成之操作,及/或被視為經構造以施行相關聯功能之複數個耦合邏輯電路元件。舉例而言,本發明之實施例包含一電腦程式產品,該電腦程式產品包含有形地體現於一機器可讀媒體上之一電腦程式,該電腦程式含有經組態以施行上文所闡述之方法之程式碼。 In addition, the various blocks shown in the flowchart may be viewed as method steps and/or operations achieved by the operation of computer program code, and/or as a plurality of coupled logic circuit elements constructed to perform related functions. For example, an embodiment of the present invention includes a computer program product, which includes a computer program tangibly embodied on a machine-readable medium, the computer program containing program code configured to implement the method described above.

在本發明之內容脈絡中,一機器/電腦可讀媒體可係可含有或儲存由一指令執行系統、設備或裝置使用之一程式或與一指令執行系統、設備或裝置結合之任何有形媒體。機器/電腦可讀媒體可係一機器/電腦可讀信號媒體或一機器/電腦可讀儲存媒體。一機器/電腦可讀媒體可係非暫時性的且可包含但不限於一電子、磁性、光學、電磁、紅外線或半導體系統、設備或裝置或前述各項之任何適合的組合。機器/電腦可讀儲存媒體之更具體實例將包含具有一或多個配線之一電連接、一可攜式電腦磁碟、一硬碟、RAM、ROM、一可抹除可程式化唯讀記憶體(EPROM或快閃記憶體)、一光纖、一可攜式光碟唯讀記憶體(CD-ROM)、一光學儲存裝置、一磁性儲存裝置或前述各項之任何適合組合。 In the context of the present invention, a machine/computer readable medium may be any tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine/computer readable medium may be a machine/computer readable signal medium or a machine/computer readable storage medium. A machine/computer readable medium may be non-transitory and may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared or semiconductor system, apparatus or device, or any suitable combination of the foregoing. More specific examples of machine/computer readable storage media would include an electrical connection having one or more wiring, a portable computer disk, a hard drive, RAM, ROM, an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

可以一或多種程式設計語言之任何組合撰寫施行本發明方法之電腦程式碼。可將此等電腦程式碼提供給一般用途電腦、特殊用途電腦之一處理器或具有控制電路系統之其他可程式化資料處理設備,以使得程式碼在由電腦或其他可程式化資料處理設備之處理器執行時使得將實施在流程圖及/或方塊圖中所規定之功能/操作。程式碼可作為一獨立軟體封裝完全地在一電腦上、部分地在電腦上執行,部分地在電腦上且部分地在一遠端電腦上執行,或者完全地在遠端電腦或伺服器上執行,亦或者散佈 於一或多個遠端電腦及/或伺服器上。 Computer program code for implementing the method of the present invention may be written in any combination of one or more programming languages. Such computer program code may be provided to a general purpose computer, a processor of a special purpose computer, or other programmable data processing device having a control circuit system so that when the program code is executed by the processor of the computer or other programmable data processing device, the functions/operations specified in the flowchart and/or block diagram will be implemented. The program code may be executed entirely on a computer, partially on a computer, partially on a computer and partially on a remote computer, or entirely on a remote computer or server as a stand-alone software package, or distributed on one or more remote computers and/or servers.

雖然本文件含有諸多具體實施方案細節,但此等細節不應被解釋為對可主張內容之範疇之限制,而是應被解釋為可係對特定實施例特有之特徵之說明。本說明書中在單獨實施例之內容脈絡中所闡述之特定特徵亦可以組合方式實施於一單個實施例中。相反地,在一單個實施例之內容脈絡中闡述之各種特徵亦可單獨地或以任何適合子組合方式實施於多個實施例中。此外,儘管上文可將特徵闡述為以特定組合方式起作用且甚至最初如此主張,但來自一所主張組合之一或多個特徵在一些情形中可自該組合去除,且該所主張組合可針對於一子組合或一子組合之變化形式。圖中所描繪之邏輯流程不需要所展示之特定次序或順序次序來達成所期望結果。另外,可提供其他步驟,或可自所闡述之流程清除步驟,且可為所闡述之系統添加或自所闡述之系統移除其他組件。因此,其他實施方案處於所附申請專利範圍之範疇內。 Although this document contains many specific implementation details, these details should not be construed as limitations on the scope of what may be claimed, but rather as descriptions of features that may be unique to particular embodiments. Specific features described in this specification in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented in multiple embodiments individually or in any suitable subcombination. Furthermore, although features may be described above as functioning in a particular combination and even initially claimed as such, one or more features from a claimed combination may be removed from that combination in some cases, and the claimed combination may be directed to a subcombination or variations of a subcombination. The logical flow depicted in the figure does not require the particular order or sequential order shown to achieve the desired results. Additionally, other steps may be provided, or steps may be eliminated from the described flow, and other components may be added to or removed from the described system. Thus, other implementations are within the scope of the appended claims.

100:沉浸式聲音與音訊服務編碼解碼器 100: Immersive sound and audio service codec

101:音訊資料 101: Audio data

102:空間分析與降混單元 102: Spatial analysis and downmixing unit

103:主音訊聲道編碼單元/主音訊聲道編碼器 103: Main audio channel encoding unit/main audio channel encoder

104:空間後設資料編碼單元 104: Spatial metadata coding unit

105:低頻率效應聲道編碼單元 105: Low frequency effect channel coding unit

106:空間後設資料解碼單元 106: Spatial metadata decoding unit

107:主音訊聲道解碼單元 107: Main audio channel decoding unit

108:低頻率效應聲道解碼單元 108: Low frequency effect channel decoding unit

109:空間合成/升混/呈現單元 109: Spatial synthesis/upmixing/presentation unit

Claims (22)

一種對一低頻率效應(LFE)聲道進行編碼之方法,其包括: 使用一或多個處理器接收一時域LFE聲道信號; 使用一低通濾波器對該時域LFE聲道信號進行濾波; 使用該一或多個處理器將經濾波之該時域LFE聲道信號轉換成該LFE聲道信號的包含表示該LFE聲道信號之一頻譜之一定數目個係數之一頻域表示; 使用該一或多個處理器將係數配置至與該LFE聲道信號之不同頻帶對應之一定數目個次頻帶群組中; 使用該一或多個處理器根據該低通濾波器之一頻率回應曲線將每一次頻帶群組中之係數量化; 使用該一或多個處理器使用針對每一次頻帶群組調諧之一熵寫碼器對該次頻帶群組中之該等經量化係數進行編碼;及 使用該一或多個處理器產生包含經編碼之該等經量化係數之一位元串流;及 使用該一或多個處理器將該位元串流儲存於一儲存裝置上或將該位元串流串流傳輸至一下游裝置。A method for encoding a low frequency effect (LFE) channel, comprising: Using one or more processors to receive a time domain LFE channel signal; Using a low pass filter to filter the time domain LFE channel signal; Using the one or more processors to convert the filtered time domain LFE channel signal into a frequency domain representation of the LFE channel signal including a certain number of coefficients representing a spectrum of the LFE channel signal; Using the one or more processors to allocate the coefficients to a certain number of coefficients corresponding to different frequency bands of the LFE channel signal; number of sub-band groups; quantizing the coefficients in each sub-band group according to a frequency response curve of the low-pass filter using the one or more processors; encoding the quantized coefficients in the sub-band group using an entropy encoder tuned for each sub-band group using the one or more processors; and generating a bit stream including the encoded quantized coefficients using the one or more processors; and storing the bit stream on a storage device or streaming the bit stream to a downstream device using the one or more processors. 如請求項1之方法,其中將每一次頻帶群組中之該等係數量化進一步包括: 基於可用量化點之一最大數目及該等係數之絕對值之一和來產生一縮放移位因數;及 使用該縮放移位因數將該等係數量化。The method of claim 1, wherein quantizing the coefficients in each frequency band group further comprises: generating a scaling factor based on a maximum number of available quantization points and a sum of the absolute values of the coefficients; and quantizing the coefficients using the scaling factor. 如請求項2之方法,若一經量化係數超過量化點之該最大數目,則將該縮放移位因數減小且再次將該等係數量化。As in the method of claim 2, if a quantized coefficient exceeds the maximum number of quantization points, the scaling factor is reduced and the coefficients are quantized again. 如前述請求項1至3中任一項之方法,其中對於每一次頻帶群組而言,該等量化點係不同的。The method of any one of claims 1 to 3 above, wherein the quantization points are different for each frequency band grouping. 如前述請求項1至3中任一項之方法,其中根據一精細量化方案或一粗略量化方案將每一次頻帶群組中之該等係數量化,其中與根據該粗略量化方案指派給一或多個次頻帶群組之量化點相比,利用該精細量化方案將更多量化點分配給該等各別次頻帶群組。A method as in any of claims 1 to 3 above, wherein the coefficients in each sub-band group are quantized according to a fine quantization scheme or a coarse quantization scheme, wherein the fine quantization scheme is used to allocate more quantization points to the respective sub-band groups than the quantization points assigned to one or more sub-band groups according to the coarse quantization scheme. 如前述請求項1至3中任一項之方法,其中該等係數之正負號位元與該等係數被分開寫碼。A method as in any one of claims 1 to 3 above, wherein the sign bits of the coefficients are coded separately from the coefficients. 如前述請求項1至3中任一項之方法,其中存在四個次頻帶群組,且一第一次頻帶群組對應於0 Hz至100 Hz之一第一頻率範圍,一第二次頻帶群組對應於100 Hz至200 Hz之一第二頻率範圍,一第三次頻帶群組對應於200 Hz至300 Hz之一第三頻率範圍,且一第四次頻帶群組對應於300 Hz至400 Hz之一第四頻率範圍。A method as in any of claims 1 to 3 above, wherein there are four sub-band groups, and a first sub-band group corresponds to a first frequency range of 0 Hz to 100 Hz, a second sub-band group corresponds to a second frequency range of 100 Hz to 200 Hz, a third sub-band group corresponds to a third frequency range of 200 Hz to 300 Hz, and a fourth sub-band group corresponds to a fourth frequency range of 300 Hz to 400 Hz. 如前述請求項1至3中任一項之方法,其中該熵寫碼器係一算術熵寫碼器。A method as in any one of claims 1 to 3 above, wherein the entropy coder is an arithmetic entropy coder. 如前述請求項1至3中任一項之方法,其中將經濾波之該時域LFE聲道信號轉換成該LFE聲道信號的包含表示該LFE聲道信號之一頻譜之一定數目個係數的一頻域表示進一步包括: 判定該LFE聲道信號之一第一步長; 基於該第一步長指定一窗函數之一第一窗口大小; 將該第一窗口大小應用於該時域LFE聲道信號之一或多個訊框;及 將一修改型離散餘弦變換(MDCT)應用於經窗口化之該等訊框以產生該等係數。A method as in any one of claims 1 to 3 above, wherein converting the filtered time-domain LFE channel signal into a frequency-domain representation of the LFE channel signal comprising a certain number of coefficients representing a spectrum of the LFE channel signal further comprises: Determining a first step length of the LFE channel signal; Specifying a first window size of a window function based on the first step length; Applying the first window size to one or more frames of the time-domain LFE channel signal; and Applying a modified discrete cosine transform (MDCT) to the windowed frames to generate the coefficients. 如請求項9之方法,其進一步包括: 判定該LFE聲道信號之一第二步長; 基於該第二步長指定該窗函數之一第二窗口大小;及 將該第二窗口大小應用於該時域LFE聲道信號之該一或多個訊框。The method of claim 9 further comprises: determining a second step size of the LFE channel signal; specifying a second window size of the window function based on the second step size; and applying the second window size to the one or more frames of the time domain LFE channel signal. 如請求項10之方法,其中: 該第一步長係N毫秒(ms); N大於或等於5 ms且小於或等於60 ms; 該第一窗口大小高於或等於10 ms; 該第二步長係5 ms;且 該第二窗口大小係10 ms。The method of claim 10, wherein: the first step size is N milliseconds (ms); N is greater than or equal to 5 ms and less than or equal to 60 ms; the first window size is greater than or equal to 10 ms; the second step size is 5 ms; and the second window size is 10 ms. 如請求項10之方法,其中: 該第一步長係20毫秒(ms); 該第一窗口大小係10 ms、20 ms或40 ms; 該第二步長係10 ms;且 該第二窗口大小係10 ms或20 ms。The method of claim 10, wherein: the first step size is 20 milliseconds (ms); the first window size is 10 ms, 20 ms, or 40 ms; the second step size is 10 ms; and the second window size is 10 ms or 20 ms. 如請求項10之方法,其中: 該第一步長係10毫秒(ms); 該第一窗口大小係10 ms或20 ms; 該第二步長係5 ms;且 該第二窗口大小係10 ms。A method as claimed in claim 10, wherein: the first step size is 10 milliseconds (ms); the first window size is 10 ms or 20 ms; the second step size is 5 ms; and the second window size is 10 ms. 如請求項10之方法,其中: 該第一步長係20毫秒(ms); 該第一窗口大小係10 ms、20 ms或40 ms; 該第二步長係5 ms;且 該第二窗口大小係10 ms。The method of claim 10, wherein: the first step size is 20 milliseconds (ms); the first window size is 10 ms, 20 ms, or 40 ms; the second step size is 5 ms; and the second window size is 10 ms. 如請求項9之方法,其中該窗函數係具有一可配置漸隱長度之一凱撒-貝索導出(KBD)窗函數。The method of claim 9, wherein the window function is a Kaiser-Besso derived (KBD) window function having a configurable gradual hidden length. 如前述請求項1至3中任一項之方法,其中該低通濾波器係一截止頻率為約130 Hz或低於130 Hz之一四階巴特沃斯低通濾波器。A method as in any of claims 1 to 3 above, wherein the low pass filter is a fourth order Butterworth low pass filter having a cutoff frequency of about 130 Hz or less. 如請求項1至3中任一項之方法,其進一步包括: 使用該一或多個處理器判定該LFE聲道信號之一訊框之一能量位準是否低於一臨限值; 根據該能量位準低於一臨限位準, 產生一靜寂訊框指示符以指示該解碼器; 將該靜寂訊框指示符插入至該LFE聲道位元串流之後設資料中;及 在偵測到靜寂訊框時減小一LFE聲道位元速率。The method of any one of claims 1 to 3 further comprises: Using the one or more processors to determine whether an energy level of a frame of the LFE channel signal is lower than a critical value; Based on the energy level being lower than a critical level, Generate a silence frame indicator to indicate the decoder; Insert the silence frame indicator into the metadata of the LFE channel bit stream; and Reduce an LFE channel bit rate when a silence frame is detected. 一種對一低頻率效應(LFE)聲道位元串流進行解碼之方法,其包括: 使用一或多個處理器接收一LFE聲道位元串流,該LFE聲道位元串流包含表示一時域LFE聲道信號之一頻譜之熵寫碼係數; 使用該一或多個處理器使用一熵解碼器將經量化係數解碼; 使用該一或多個處理器將經量化係數逆量化,其中該等係數已根據用於在一編碼器中對該時域LFE聲道信號進行濾波之一低通濾波器之一頻率回應曲線而在與頻帶對應之次頻帶群組中被量化; 使用該一或多個處理器將經逆量化之該等係數轉換成一時域LFE聲道信號; 使用該一或多個處理器調整該時域LFE聲道信號之一延時;及 使用一低通濾波器對經延時調整之該LFE聲道信號進行濾波。A method for decoding a low frequency effect (LFE) channel bit stream, comprising: Using one or more processors to receive an LFE channel bit stream, the LFE channel bit stream comprising entropy coded coefficients representing a spectrum of a time domain LFE channel signal; Using the one or more processors to decode the quantized coefficients using an entropy decoder; Using the one or more processors to inverse quantize the quantized coefficients, wherein the coefficients have been quantized according to the entropy decoder; A frequency response curve of a low pass filter for filtering the time domain LFE channel signal in an encoder is quantized in a sub-band group corresponding to the frequency band; The inverse quantized coefficients are converted into a time domain LFE channel signal using the one or more processors; A delay of the time domain LFE channel signal is adjusted using the one or more processors; and The delay-adjusted LFE channel signal is filtered using a low pass filter. 如請求項18之方法,其中低通濾波器之一階數經組態以確保由於對該LFE聲道進行編碼及解碼所致之一第一總演算法延時小於或等於由於對包含該LFE聲道信號之一多聲道音訊信號中之其他聲道進行編碼及解碼所致之一第二總演算法延時。A method as in claim 18, wherein an order of the low pass filter is configured to ensure that a first total algorithm delay due to encoding and decoding the LFE channel is less than or equal to a second total algorithm delay due to encoding and decoding other channels in a multi-channel audio signal including the LFE channel signal. 如請求項19之方法,其進一步包括: 判定該第二總演算法延時是否超過一臨限值;及 根據該第二總演算法延時超過該臨限值, 將該低通濾波器組態為一N階低通濾波器,其中N係大於或等於2之一整數;及 根據該第二總演算法延時不超過該臨限值, 將該低通濾波器之該階數組態為小於N。The method of claim 19 further includes: Determining whether the second total algorithm delay exceeds a critical value; and Based on the second total algorithm delay exceeding the critical value, Configuring the low-pass filter as an N-order low-pass filter, where N is an integer greater than or equal to 2; and Based on the second total algorithm delay not exceeding the critical value, Configuring the order of the low-pass filter to be less than N. 一種低延遲低頻率效應(LFE)之編碼器/解碼器,其包括: 一或多個處理器;及 一儲存指令之非暫時性電腦可讀媒體,該等指令在由該一或多個處理器執行時使得該一或多個處理器實行如方法請求項1至20中任一項之操作。A low latency low frequency effect (LFE) encoder/decoder comprising: one or more processors; and a non-transitory computer-readable medium storing instructions that, when executed by the one or more processors, cause the one or more processors to perform any of method requests 1 to 20. 一種儲存指令之非暫時性電腦可讀媒體,該等指令在由一或多個處理器執行時使得該一或多個處理器實行如方法請求項1至20中任一項之操作。A non-transitory computer-readable medium storing instructions which, when executed by one or more processors, cause the one or more processors to perform the operations of any one of method claims 1 to 20.
TW109130176A 2020-09-03 2020-09-03 Low-latency, low-frequency effects codec TWI882003B (en)

Priority Applications (1)

Application Number Priority Date Filing Date Title
TW109130176A TWI882003B (en) 2020-09-03 2020-09-03 Low-latency, low-frequency effects codec

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
TW109130176A TWI882003B (en) 2020-09-03 2020-09-03 Low-latency, low-frequency effects codec

Publications (2)

Publication Number Publication Date
TW202211206A TW202211206A (en) 2022-03-16
TWI882003B true TWI882003B (en) 2025-05-01

Family

ID=81731788

Family Applications (1)

Application Number Title Priority Date Filing Date
TW109130176A TWI882003B (en) 2020-09-03 2020-09-03 Low-latency, low-frequency effects codec

Country Status (1)

Country Link
TW (1) TWI882003B (en)

Citations (8)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
EP0563832A1 (en) * 1992-03-30 1993-10-06 Matsushita Electric Industrial Co., Ltd. Stereo audio encoding apparatus and method
EP0875999A2 (en) * 1997-03-31 1998-11-04 Sony Corporation Encoding method and apparatus, decoding method and apparatus and recording medium
US6016473A (en) * 1998-04-07 2000-01-18 Dolby; Ray M. Low bit-rate spatial coding method and system
US6226616B1 (en) * 1999-06-21 2001-05-01 Digital Theater Systems, Inc. Sound quality of established low bit-rate audio coding systems without loss of decoder compatibility
EP0864146B1 (en) * 1995-12-01 2004-10-13 Digital Theater Systems, Inc. Multi-channel predictive subband coder using psychoacoustic adaptive bit allocation
US20040267543A1 (en) * 2003-04-30 2004-12-30 Nokia Corporation Support of a multichannel audio extension
US20160012825A1 (en) * 2013-04-05 2016-01-14 Dolby International Ab Audio encoder and decoder
US20160267914A1 (en) * 2013-11-29 2016-09-15 Dolby Laboratories Licensing Corporation Audio object extraction

Patent Citations (9)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
EP0563832A1 (en) * 1992-03-30 1993-10-06 Matsushita Electric Industrial Co., Ltd. Stereo audio encoding apparatus and method
EP0864146B1 (en) * 1995-12-01 2004-10-13 Digital Theater Systems, Inc. Multi-channel predictive subband coder using psychoacoustic adaptive bit allocation
EP0875999A2 (en) * 1997-03-31 1998-11-04 Sony Corporation Encoding method and apparatus, decoding method and apparatus and recording medium
US6016473A (en) * 1998-04-07 2000-01-18 Dolby; Ray M. Low bit-rate spatial coding method and system
US6226616B1 (en) * 1999-06-21 2001-05-01 Digital Theater Systems, Inc. Sound quality of established low bit-rate audio coding systems without loss of decoder compatibility
US20040267543A1 (en) * 2003-04-30 2004-12-30 Nokia Corporation Support of a multichannel audio extension
US20160012825A1 (en) * 2013-04-05 2016-01-14 Dolby International Ab Audio encoder and decoder
US20170301362A1 (en) * 2013-04-05 2017-10-19 Dolby International Ab Audio decoder for interleaving signals
US20160267914A1 (en) * 2013-11-29 2016-09-15 Dolby Laboratories Licensing Corporation Audio object extraction

Also Published As

Publication number Publication date
TW202211206A (en) 2022-03-16

Similar Documents

Publication Publication Date Title
JP7829668B2 (en) Encoding and decoding of IVAS bitstreams
JP2026062726A (en) Low-latency, bass-enhancing codec
JP2023551732A (en) Immersive voice and audio services (IVAS) with adaptive downmix strategy
TW202230335A (en) Apparatus, method, or computer program for processing an encoded audio scene using a parameter smoothing
TW202242852A (en) Adaptive gain control
TWI882003B (en) Low-latency, low-frequency effects codec
TWI921149B (en) Low-latency, low-frequency effects codec
RU2809977C1 (en) Low latency codec with low frequency effects
CN120226077A (en) Method, apparatus and medium for audio bitstream encoding and decoding
TW202534670A (en) Low-latency, low-frequency effects codec
HK40073280A (en) Low-latency, low-frequency effects codec
TW202230334A (en) Apparatus, method, or computer program for processing an encoded audio scene using a parameter conversion
TWI897027B (en) Decoder and decoding method for discontinuous transmission of parametrically coded independent streams with metadata
TWI897026B (en) Encoder and encoding method for discontinuous transmission of parametrically coded independent streams with metadata
HK40097526A (en) Spatial noise filling in multi-channel codec
CN120077434A (en) Methods, apparatus and media for encoding and decoding of audio bitstreams and associated echo reference signals
CN120226074A (en) Method and apparatus for discontinuous transmission in object-based audio codec
CN119998875A (en) Method, device and medium for decoding an audio signal having skippable blocks
CN116547748A (en) Spatial noise filling in multi-channel codecs
HK40128273A (en) Bitrate distribution in immersive voice and audio services
TW202219942A (en) Apparatus, method, or computer program for processing an encoded audio scene using a bandwidth extension
CN119998873A (en) Method, apparatus and medium for encoding and decoding audio bitstreams using flexible block-based syntax
HK40071164A (en) Encoding and decoding ivas bitstreams
HK40076195B (en) Bitrate distribution in immersive voice and audio services