WO2022012553A1 - 多声道音频信号的编解码方法和装置 - Google Patents
多声道音频信号的编解码方法和装置 Download PDFInfo
- Publication number
- WO2022012553A1 WO2022012553A1 PCT/CN2021/106101 CN2021106101W WO2022012553A1 WO 2022012553 A1 WO2022012553 A1 WO 2022012553A1 CN 2021106101 W CN2021106101 W CN 2021106101W WO 2022012553 A1 WO2022012553 A1 WO 2022012553A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- channel
- pair
- audio frame
- correlation value
- channel pair
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L19/00—Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis
- G10L19/008—Multichannel audio signal coding or decoding using interchannel correlation to reduce redundancy, e.g. joint-stereo, intensity-coding or matrixing
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L25/00—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00
- G10L25/03—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 characterised by the type of extracted parameters
- G10L25/06—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 characterised by the type of extracted parameters the extracted parameters being correlation coefficients
Definitions
- the present application relates to audio processing technologies, and in particular, to a method and device for encoding and decoding multi-channel audio signals.
- Encoding and decoding of multi-channel audio is a technique for encoding or decoding audio that contains more than two channels.
- Common multi-channel audios include 5.1-channel audio, 7.1-channel audio, 7.1.4-channel audio, and 22.2-channel audio.
- MPS MPEG Surround
- the present application provides a method and device for encoding and decoding multi-channel audio signals, so as to reduce redundancy between channel signals and improve audio encoding efficiency.
- the present application provides a method for encoding a multi-channel audio signal, including: acquiring a first audio frame to be encoded, where the first audio frame includes at least five channel signals; acquiring a correlation value set, the The correlation value set includes respective correlation values of a plurality of channel pairs, one channel pair includes two channel signals in the at least five channel signals, and the correlation value of the one channel pair is used to represent the one channel pair.
- M is a positive integer less than or equal to the set value
- M is a positive integer less than or equal to the set value
- the first audio frame in this embodiment may be any frame in the multi-channel audio signal to be encoded, and the first audio frame includes five or more channel signals. Coding two channel signals with higher correlation together can reduce redundancy and improve coding efficiency. Therefore, in this embodiment, the pairing is determined according to the correlation value between the two channel signals. In order to find the channel pair set with the highest correlation as much as possible, the correlation value between at least five channel signals in the first audio frame can be calculated to obtain the correlation value set of the first audio frame. For example, five channel signals may form 10 channel pairs in total, and correspondingly, the correlation value set may include 10 correlation values.
- all the correlation values included in the correlation value set can be sorted in descending order, and the top M correlation values in the front row are selected from them, and the M correlation values must be greater than or equal to the group pair threshold, This is because a correlation value smaller than the group pair threshold indicates that the correlation between the two channel signals in the corresponding channel pair is low, and there is no need for group pair coding.
- an upper limit N of M is set, that is, at most N correlation values can be selected.
- the target channel pair set contains The sum of the correlation values of all channel pairs is the largest, and the number of channel pairs in a group pair is increased as much as possible to reduce the redundancy between channel signals and improve the audio coding efficiency.
- the M channel pair sets include a first channel pair set, and the acquiring the M channel pair sets acquires the first channel pair set; the acquiring the first channel pair set A set of channel pairs, comprising: adding a first channel pair in the M channel pairs to the first channel pair set, where the first channel pair is the first channel pair in the M channel pairs Any one; when the other channel pairs other than the associated channel in the multiple channel pairs include channel pairs whose correlation value is greater than the set pair threshold, select the maximum correlation value from the other channel pairs One channel pair of the channel pair is added to the first channel pair set, and the associated channel pair includes any one of the channel signals included in the channel pair that has been added to the first channel pair set.
- the multiple channel pairs with larger correlation values are used as the first channel pair added to the channel pair set, and then the channel pair corresponding to the largest correlation value in the remaining channel pairs is selected.
- Add the corresponding channel pair set obtain the sum of the correlation values of multiple channel pair sets as much as possible, and then determine the channel pair set corresponding to the maximum correlation value sum as the target channel pair set, so that the target sound can be achieved.
- the sum of the correlation values of all channel pairs included in the channel pair set is the largest, and the number of channel pairs in a group pair is increased as much as possible, reducing redundancy between channel signals and improving audio coding efficiency.
- the selecting M correlation values from the correlation value set includes: selecting N correlation values from the correlation value set, where the N correlation values are all greater than the correlation values other correlation values in the value set except the N correlation values, where N is the set value; select correlation values greater than or equal to the pair threshold from the N correlation values, and the correlation values greater than or equal to The number of correlation values of the group to the threshold is M.
- the M correlation values are greater than or equal to the pair threshold, and M is a positive integer less than or equal to a set value (eg, N).
- a set value eg, N
- all the correlation values included in the correlation value set may be sorted in descending order, and the top N correlation values in the front row may be selected from the correlation values, and the N correlation values may have correlation values smaller than the group pair threshold. Therefore, M correlation values greater than or equal to the group pair threshold are selected from the N correlation values, because the correlation values smaller than the group pair threshold represent the correlation between the two channel signals in the corresponding channel pair. Low sex, no group pair coding necessary.
- the correlation value is a normalized value.
- the normalization process can incorporate the related values with large differences in the value range into a unified range for comparison and processing, thereby improving the operation efficiency.
- the correlation value of the one channel pair is set to 0.
- a small correlation value indicates that the correlation between the corresponding two channel signals is small, and there is no need for group pairs. Therefore, the correlation value of the two channel signals in this case is set to 0, which is convenient for subsequent calculations and improves. Operational efficiency.
- the present application provides a method for encoding a multi-channel audio signal, including: acquiring a first audio frame to be encoded, where the first audio frame includes at least five channel signals; acquiring a correlation value set, the The correlation value set includes respective correlation values of a plurality of channel pairs, one channel pair includes two channel signals in the at least five channel signals, and the correlation value of the one channel pair is used to represent the one channel pair.
- Correlation between two channel signals of a channel pair multiple channel pair sets are obtained according to the multiple channel pairs, when the channel pair set includes more than two channel pairs, the two More than one channel pair does not contain the same channel signal; obtain the sum of the correlation values of all channel pairs included in each channel pair set in the multiple channel pair sets according to the correlation value set; determine the target sound A channel pair set, the sum of the correlation values of all channel pairs in the target channel pair set is the largest among the multiple channel pair sets; according to the target channel pair set, the first audio frame to encode.
- the obtaining multiple channel pair sets according to the multiple channel pairs includes: obtaining the obtained channel pairs according to other channel pairs in the multiple channel pairs other than unrelated channels The plurality of channel pairs are set, and the correlation value of the uncorrelated channel pairs is less than the group pair threshold.
- a small correlation value indicates that the correlation between the corresponding two channel signals is small, and there is no need for group pairs. Therefore, the correlation value of the two channel signals in this case and the sound of the two channel signals
- the deletion of the track pair can reduce the amount of subsequent calculations and improve the operation efficiency.
- the correlation value is a normalized value.
- the normalization process can incorporate the related values with large differences in the value range into a unified range for comparison and processing, thereby improving the operation efficiency.
- the correlation value of the one channel pair is set to 0.
- a small correlation value indicates that the correlation between the corresponding two channel signals is small, and there is no need for group pairs. Therefore, the correlation value of the two channel signals in this case is set to 0, which is convenient for subsequent calculations and improves. Operational efficiency.
- the present application provides a method for encoding a multi-channel audio signal, comprising: acquiring a first audio frame to be encoded, where the first audio frame includes at least five channel signals; acquiring the first audio frame
- the correlation value set of the first audio frame includes the respective correlation values of multiple channel pairs, one channel pair includes two channel signals in the at least five channel signals, and the one channel pair includes two channel signals in the at least five channel signals.
- the correlation value of the channel pair is used to represent the correlation between the two channel signals of the one channel pair; the correlation value set of the second audio frame is obtained, and the correlation value set of the second audio frame includes the The respective correlation values of a plurality of channel pairs of the second audio frame, one channel pair includes two channel signals among the at least five channel signals of the second audio frame, and the correlation value of the one channel pair Used to represent the correlation between the two channel signals of the one channel pair, the second audio frame is the previous frame of the first audio frame; according to the set of correlation values of the first audio frame and the correlation value set of the second audio frame to judge whether it is necessary to re-acquire the target channel pair set of the first audio frame; if it is necessary to re-acquire the target channel pair set of the first audio frame, the above
- the method according to any one of the first to second aspects acquires a target channel pair set of the first audio frame, and encodes the first audio frame according to the target channel pair set; Acquiring the target channel pair set of the first audio frame
- the sum of the difference between the correlation value set of the current audio frame and the correlation value set of the previous audio frame it is determined whether it is necessary to re-acquire the target channel pair set of the current frame. Reduce the amount of calculation and improve the coding efficiency. Even if the audio changes greatly and the target channel pair set needs to be re-acquired, the sum of the correlation values of multiple channel pair sets can still be obtained as much as possible, and then the sum of the maximum correlation value can be obtained.
- the corresponding channel pair set is determined as the target channel pair set, which can maximize the sum of the correlation values of all channel pairs included in the target channel pair set, and increase the number of channel pairs in the group pair as much as possible, reducing the The redundancy between channel signals improves the coding efficiency of audio.
- determining whether the target channel pair set of the first audio frame needs to be re-acquired according to the correlation value set of the first audio frame and the correlation value set of the second audio frame including: calculating the absolute value of the difference between the correlation value set of the first audio frame and the correlation value set of the second audio frame corresponding to the same channel pair; The corresponding sum of the absolute values; when the sum of the absolute values is less than the change threshold, it is determined that it is not necessary to re-acquire the target channel pair set of the first audio frame; when the sum of the absolute values is greater than or equal to the When the threshold value is changed, it is determined that the target channel pair set of the first audio frame needs to be re-acquired.
- the change threshold can be, for example, ⁇ the number of channel pairs, where the value of ⁇ can be 0.14 or 0.15, and the number of channel pairs refers to the set of correlation values of the first audio frame (or the correlation value of the second audio frame). The number of channel pairs included in the value set).
- the present application provides a method for encoding a multi-channel audio signal, comprising: acquiring a first audio frame to be encoded, where the first audio frame includes K channel signals, where K is an integer greater than or equal to 5
- K is an integer greater than or equal to 5
- K is an integer greater than or equal to 5
- K is an integer greater than or equal to 5
- K is an integer greater than or equal to 5
- K is an integer greater than or equal to 5
- K is an integer greater than or equal to 5
- K is an integer greater than or equal to 5
- K is an integer greater than or equal to 5
- K is an integer greater than or equal to 5
- K is an integer greater than or equal to 5
- K is an integer greater than or equal to 5
- K is an integer greater than or equal to 5
- K is an integer greater than or equal to 5
- K is an integer greater than or equal to 5
- K is an integer greater than or equal to 5
- K is an integer greater than or equal to 5
- K is an integer greater
- the difference from the method of the first aspect or the second aspect is that the method of the first aspect and the second aspect is fused, that is, according to the number of channel signals contained in the first audio frame, it is determined which audio frame to use for the first audio frame.
- a method gets its target channel pair collection.
- the method of the second aspect is adopted, it is necessary to exhaustively enumerate all target channel pairs, which will increase the amount of calculation. Therefore, the method of the first aspect will reduce the amount of calculation. A lot of computation.
- the method of the second aspect can be used to obtain the sum of the correlation values of all channel pair sets, so as to ensure that the final selected target channel pair set must be the highest Optimal results that match the characteristics of the first audio frame.
- the present application provides an encoding device, comprising: an acquisition module configured to acquire a first audio frame to be encoded, where the first audio frame includes at least five channel signals; acquire a correlation value set, the correlation The value set includes respective correlation values of a plurality of channel pairs, one channel pair includes two channel signals in the at least five channel signals, and the correlation value of the one channel pair is used to represent the one channel pair.
- the M correlation values are all greater than or equal to the group pair threshold, and M is a positive integer less than or equal to the set value; acquire M channel pair sets, each of which includes at least the One of the M channel pairs corresponding to the M correlation values, and when the channel pair set includes more than two channel pairs, the two or more channel pairs do not contain the same channel signal; determine A module for determining a target channel pair set from the M channel pair sets, where the sum of the correlation values of all channel pairs in the target channel pair set is the largest in the M channel pair sets an encoding module, configured to encode the first audio frame according to the target channel pair set.
- the M channel pair sets include a first channel pair set; and the acquiring module is specifically configured to add the first channel pair in the M channel pairs to all channel pairs.
- the first channel pair set, the first channel pair is any one of the M channel pairs;
- the obtaining module is specifically configured to select N correlation values from the correlation value set, and the N correlation values are all greater than the N correlation values in the correlation value set divided by the N correlation values.
- Other correlation values other than the value, N is the set value; from the N correlation values, select the correlation value greater than or equal to the threshold value of the group pair, the correlation value greater than or equal to the threshold value of the group pair The number is M.
- the correlation value is a normalized value.
- the correlation value of the one channel pair is set to 0.
- the present application provides an encoding device, comprising: an acquisition module for acquiring a first audio frame to be encoded, the first audio frame including at least five channel signals; acquiring a correlation value set, the correlation The value set includes respective correlation values of a plurality of channel pairs, one channel pair includes two channel signals in the at least five channel signals, and the correlation value of the one channel pair is used to represent the one channel pair.
- Correlation between two channel signals of a channel pair multiple channel pair sets are obtained according to the multiple channel pairs, and when the channel pair set includes more than two channel pairs, the two The above channel pairs do not contain the same channel signal; obtain the sum of the correlation values of all channel pairs included in each channel pair set in the multiple channel pair sets according to the correlation value set; determine the module, use In determining a target channel pair set, the sum of the correlation values of all channel pairs in the target channel pair set is the largest in the multiple channel pair sets; an encoding module is used for according to the target channel pair set.
- the set encodes the first audio frame.
- the obtaining module is specifically configured to obtain the set of multiple channel pairs according to other channel pairs in the multiple channel pairs other than the non-correlated channels to the outside, and the non-correlated channels The correlation value of the channel pair is less than the group pair threshold.
- the correlation value is a normalized value.
- the correlation value of the one channel pair is set to 0.
- the present application provides an encoding device, comprising: an acquisition module configured to acquire a first audio frame to be encoded, the first audio frame including at least five channel signals; A set of correlation values, the set of correlation values of the first audio frame includes respective correlation values of a plurality of channel pairs, one channel pair includes two channel signals in the at least five channel signals, and the one channel pair includes two channel signals in the at least five channel signals.
- the correlation value of the channel pair is used to represent the correlation between the two channel signals of the one channel pair; the correlation value set of the second audio frame is obtained, and the correlation value set of the second audio frame includes the first audio frame.
- one channel pair includes two channel signals out of at least five channel signals of the second audio frame, and the correlation value of the one channel pair is defined by
- the second audio frame is the previous frame of the first audio frame; the encoding module is used for according to the first audio frame.
- the correlation value set of the second audio frame and the correlation value set of the second audio frame judge whether it is necessary to re-acquire the target channel pair set of the first audio frame; if it is necessary to re-acquire the target channel pair set of the first audio frame, Then execute the method according to any one of claims 1-9 to obtain the target channel pair set of the first audio frame, and encode the first audio frame according to the target channel pair set; if It is not necessary to re-acquire the target channel pair set of the first audio frame, then determine the target channel pair set of the second audio frame as the target channel pair set of the first audio frame, and determine the target channel pair set of the second audio frame as the target channel pair set of the first audio frame.
- the set of target channel pairs encodes the first audio frame.
- the encoding module is specifically configured to calculate the difference between the correlation value set of the first audio frame and the correlation value set corresponding to the same channel pair in the correlation value set of the second audio frame The absolute value of the difference; calculate the sum of the absolute values corresponding to the plurality of the channel pairs respectively; when the sum of the absolute values is less than the change threshold, it is determined that the target channel of the first audio frame does not need to be re-acquired Pair set; when the sum of the absolute values is greater than or equal to the change threshold, it is determined that the target channel pair set of the first audio frame needs to be re-acquired.
- the present application provides an encoding device, comprising: an acquisition module configured to acquire a first audio frame to be encoded, where the first audio frame includes K channel signals, and K is an integer greater than or equal to 5; an encoding module, configured to perform the method according to any one of the above-mentioned first aspects to encode the first audio frame when K is greater than the threshold of the number of channel signals; when K is less than or equal to the threshold of the number of channel signals , performing the method according to any one of the above second aspects to encode the first audio frame.
- the present application provides a device, comprising: one or more processors; a memory for storing one or more programs; when the one or more programs are executed by the one or more processors, The one or more processors are caused to implement the method of any one of the first to fourth aspects above.
- the present application provides a computer-readable storage medium, comprising a computer program, which, when executed on a computer, causes the computer to execute the method according to any one of the first to fourth aspects.
- the present application provides a computer-readable storage medium, characterized by comprising an encoded code stream obtained according to the encoding method for a multi-channel audio signal according to any one of the first to fourth aspects above.
- FIG. 1 exemplarily presents a schematic block diagram of an audio decoding system 10 applied in the present application
- FIG. 2 exemplarily presents a schematic block diagram of an audio decoding device 200 to which the present application is applied;
- FIG. 3 is a flowchart of an exemplary embodiment of a method for encoding a multi-channel audio signal provided by the present application
- FIG. 4 is an exemplary structural diagram of an encoding device to which the multi-channel audio signal encoding method provided by the present application is applied;
- FIG. 5 is a flowchart of an exemplary embodiment of a method for encoding a multi-channel audio signal provided by the present application
- FIG. 6 is a flowchart of an exemplary embodiment of a method for encoding a multi-channel audio signal provided by the present application
- FIG. 7 is a flowchart of an exemplary embodiment of a method for encoding a multi-channel audio signal provided by the present application
- FIG. 8 is an exemplary structural diagram of a decoding device to which the decoding method for a multi-channel audio signal provided by the present application is applied;
- FIG. 9 is a schematic structural diagram of an embodiment of an encoding device of the present application.
- FIG. 10 is a schematic structural diagram of an embodiment of a device of the present application.
- At least one (item) refers to one or more, and "a plurality” refers to two or more.
- “And/or” is used to describe the relationship between related objects, indicating that there can be three kinds of relationships, for example, “A and/or B” can mean: only A, only B, and both A and B exist , where A and B can be singular or plural.
- the character “/” generally indicates that the associated objects are an “or” relationship.
- At least one item(s) below” or similar expressions thereof refer to any combination of these items, including any combination of single item(s) or plural items(s).
- At least one (a) of a, b or c can mean: a, b, c, "a and b", “a and c", “b and c", or "a and b and c" ", where a, b, c can be single or multiple.
- Audio frame Audio data is streaming.
- the amount of audio data within a period of time is usually taken as a frame of audio. This period is called “sampling time", which can be determined according to the codec. Determine its value according to the requirements of the device and specific applications, for example, the duration is 2.5ms to 60ms, and ms is milliseconds.
- Audio signal is the information carrier of frequency and amplitude variation of regular sound waves with speech, music and sound effects. Audio is a continuously changing analog signal that can be represented by a continuous curve called a sound wave. Audio is a digital signal generated by analog-to-digital conversion or by a computer. Sound waves have three important parameters: frequency, amplitude and phase, which determine the characteristics of the audio signal.
- Channel signal refers to the independent audio signals that are collected or played back at different spatial positions during recording or playback. Therefore, the number of channels is the number of sound sources during sound recording or the number of speakers during playback.
- FIG. 1 exemplarily shows a schematic block diagram of an audio decoding system 10 applied in the present application.
- the audio coding system 10 may include a source device 12 and a destination device 14.
- the source device 12 generates an encoded code stream, and thus, the source device 12 may be referred to as an audio encoding device.
- the destination device 14 may decode the encoded codestream generated by the source device 12, and thus, the destination device 14 may be referred to as an audio decoding device.
- the source device 12 includes an encoder 20 and, optionally, an audio source 16 , an audio preprocessor 18 , and a communication interface 22 .
- Audio source 16 may include or be any type of audio capture device for capturing real world speech, music, sound effects, etc., and/or any type of audio generation device, such as an audio processor for generating speech, music, and sound effects or equipment.
- the audio source may be any type of memory or storage that stores the above audio.
- the audio preprocessor 18 is used to receive (raw) audio data 17 and to preprocess the audio data 17 to obtain preprocessed audio data 19 .
- the preprocessing performed by the audio preprocessor 18 may include trimming or denoising. It is understood that the audio preprocessing unit 18 may be an optional component.
- An encoder 20 is used to receive preprocessed audio data 19 and provide encoded audio data 21 .
- a communication interface 22 in source device 12 may be used to receive encoded audio data 21 and send encoded audio data 21 over communication channel 13 to destination device 14 for storage or direct reconstruction.
- the destination device 14 includes a decoder 30 and, optionally, a communication interface 28 , an audio post-processor 32 and a playback device 34 .
- the communication interface 28 in the destination device 14 is used to receive the encoded audio data 21 directly from the source device 12 and to provide the encoded audio data 21 to the decoder 30 .
- Communication interface 22 and communication interface 28 may be used through a direct communication link between source device 12 and destination device 14, such as a direct wired or wireless connection, etc., or through any type of network, such as a wired network, a wireless network, or any A combination, any type of private network and public network, or any type of combination, transmits or receives encoded audio data 21 .
- the communication interface 22 may be used to encapsulate the encoded audio data 21 into a suitable format such as a message, and/or to process the encoded audio data 21 using any type of transfer encoding or processing for transmission over a communication link or communication network .
- the communication interface 28 corresponds to the communication interface 22 and may be used, for example, to receive transmission data and process the transmission data to obtain encoded audio data 21 using any type of corresponding transmission decoding or processing and/or decapsulation.
- Both the communication interface 22 and the communication interface 28 can be configured as a one-way communication interface as indicated by the arrow in FIG. 1 from the corresponding communication channel 13 of the source device 12 to the destination device 14, or a two-way communication interface, and can be used to send and receive messages etc. to establish a connection, acknowledge and exchange any other information related to a communication link and/or data transfer such as encoded audio data, etc.
- a decoder 30 is used to receive encoded audio data 21 and provide decoded audio data 31 .
- the audio post-processor 32 is used for post-processing the decoded audio data 31 to obtain post-processed post-processed audio data 33 .
- the post-processing performed by the audio post-processor 32 may include, for example, trimming or resampling, and the like.
- Playback device 34 is used to receive post-processed audio data 33 to play audio to a user or listener.
- Playback device 34 may be or include any type of player for playing reconstructed audio, eg, integrated or external speakers.
- speakers may include speakers, speakers, and the like.
- FIG. 2 exemplarily shows a schematic block diagram of an audio decoding device 200 applied in the present application.
- the audio coding apparatus 200 may be an audio decoder (eg, decoder 30 of FIG. 1 ) or an audio encoder (eg, encoder 20 of FIG. 1 ).
- the audio decoding device 200 includes: an input port 210 and a receiving unit (Rx) 220 for receiving data, a processor, logic unit or central processing unit 230 for processing data, and a transmitting unit (Tx) 240 for transmitting data and egress port 250, and a memory 260 for storing data.
- the audio decoding device 200 may also include photoelectric conversion components and electro-optical (EO) components coupled with the input port 210, the receiving unit 220, the transmitting unit 240, and the output port 250 for the exit or entrance of optical or electrical signals.
- EO electro-optical
- the processor 230 is implemented by hardware and software.
- the processor 230 may be implemented as one or more CPU chips, cores (eg, multi-core processors), FPGAs, ASICs, and DSPs.
- the processor 230 communicates with the ingress port 210 , the receiving unit 220 , the transmitting unit 240 , the egress port 250 and the memory 260 .
- the processor 230 includes a decoding module 270 (eg, an encoding module or a decoding module).
- the decoding module 270 implements the embodiments disclosed in this application, so as to implement the encoding and decoding methods for multi-channel audio signals provided in this application.
- the transcoding module 270 implements, processes, or provides various encoding operations.
- decoding module 270 is implemented as instructions stored in memory 260 and executed by processor 230 .
- Memory 260 includes one or more magnetic disks, tape drives, and solid-state drives, and may serve as an overflow data storage device for storing programs as they are selectively executed, and for storing instructions and data read during program execution.
- Memory 260 may be volatile and/or non-volatile, and may be read only memory (ROM), random access memory (RAM), random access memory (ternary content-addressable memory, TCAM) and/or static Random Access Memory (SRAM).
- ROM read only memory
- RAM random access memory
- TCAM ternary content-addressable memory
- SRAM static Random Access Memory
- the present application provides a method for encoding and decoding a multi-channel audio signal.
- FIG. 3 is a flowchart of an exemplary embodiment of a method for encoding a multi-channel audio signal provided by the present application.
- the process 300 may be performed by the source device 12 or the audio coding device 200 in the audio coding system 10 .
- Process 300 is described as a series of steps or operations, and it should be understood that process 300 may be performed in various orders and/or concurrently, and is not limited to the order of execution shown in FIG. 3 .
- the method includes:
- Step 301 Obtain a first audio frame to be encoded.
- the first audio frame in this embodiment may be any frame in the multi-channel audio signal to be encoded, and the first audio frame includes five or more channel signals.
- a 5.1 channel includes a center channel (C), a front left channel (left, L), a front right channel (right, R), a left surround channel (LS), a back The right surround channel (right surround, RS) and the 0.1 channel low frequency effects (low frequency effects, LFE) a total of six channel signals.
- the 7.1 channel includes C, L, R, LS, RS, LB, RB and LFE a total of eight channel signals, where LFE is the audio channel from 3-120Hz, which is usually sent to a channel specially designed for low tones. designed speakers.
- Step 302 Obtain a set of related values.
- the correlation value set includes the respective correlation values of a plurality of channel pairs, wherein one channel pair includes two channel signals in at least five channel signals, and the correlation value of one channel pair is used to represent the two channels of the channel pair. correlation between channel signals.
- the multiple channel pairs may include all channel pairs corresponding to at least five channel signals, or, the multiple channel pairs may also include some channel pairs corresponding to at least five channel signals. Make specific restrictions.
- the pairing is determined according to the correlation value between the two channel signals.
- the correlation value set of the first audio frame may be obtained by first calculating the correlation value between at least five channel signals in the first audio frame. For example, five channel signals may form 10 channel pairs in total, and correspondingly, the correlation value set may include 10 correlation values.
- the correlation values can be normalized, so that the correlation values of all channel pairs are limited to a specific range, so as to set a unified judgment standard for the correlation values, such as a group pair threshold, the group pair threshold can be set. It is a value greater than or equal to 0.2 and less than or equal to 1, such as 0.3, 0.4, or 0.35, etc., so that as long as the normalized correlation value of the two channel signals is less than the group pair threshold, the two channels are considered to be The correlation of the channel signals is poor, and group pair coding is not required.
- the correlation value between two channel signals can be calculated using the following formula:
- corr_norm(ch1, ch2) represents the normalized correlation value between the channel signal ch1 and the channel signal ch2
- spec_ch1(i) represents the frequency domain coefficient of the ith frequency point of the channel signal ch1
- spec_ch2(i ) is the frequency domain coefficient of the ith frequency point of the channel signal ch2
- N represents the total number of frequency points of an audio frame.
- the correlation value calculated by using the above algorithm or formula can be used as the initial correlation value, and then it is determined whether the initial correlation value needs to be modified according to a preset condition.
- the restriction condition may include: calculating whether the amplitude ratio between the two channel signals of the initial correlation value is greater than a preset group pair threshold. When the amplitude ratio is greater than the pair threshold, the initial correlation value is modified; when the amplitude ratio is less than or equal to the pair threshold, the initial correlation value is kept unchanged.
- the modification may be to reduce the initial correlation value, for example, the initial correlation value may be directly modified to 0, so as to prevent the two channel signals from being processed in groups.
- the amplitude level(ch) of the current frame of the channel signal ch can be calculated by using the following formula:
- i represents the ith sample of the current frame of the channel signal ch
- N represents the total number of samples of the current frame
- sepc_coeff(ch, i) is the frequency domain coefficient of the ith sample of the current frame.
- Step 303 Select M correlation values from the correlation value set.
- the M correlation values are all greater than other correlation values in the correlation value set except the M correlation values, the M correlation values are all greater than or equal to the group pair threshold, and M is a positive value less than or equal to a set value (for example, N). Integer.
- all the correlation values included in the correlation value set can be sorted in descending order, and the top M correlation values in the front row are selected from them, and the M correlation values must be greater than or equal to the group pair threshold, This is because a correlation value smaller than the group pair threshold indicates that the correlation between the two channel signals in the corresponding channel pair is low, and there is no need for group pair coding.
- an upper limit N of M is set, that is, at most N correlation values can be selected.
- N can be an integer greater than or equal to 2, and the maximum value of N cannot exceed the number of all channel pairs corresponding to all channel signals of the first audio frame.
- the larger the value of N the larger the amount of computation involved, while the smaller the value of N is, the channel pair set may be lost, thereby reducing the coding efficiency.
- the subsequent steps do not need to be performed, and it is sufficient to perform mono encoding on each channel signal of the first audio frame. If M correlation values are selected from the set of correlation values, the following steps may be performed.
- Step 304 Acquire M channel pair sets.
- Each channel pair set includes at least one of the M channel pairs corresponding to the M correlation values, and when the channel pair set includes more than two channel pairs, the two or more channel pairs do not contain the same sound channel signal.
- the 3 channel pairs corresponding to the largest correlation value selected according to the correlation value set are (L, R), (R, C) and (LS, RS), where the correlation of (LS, RS) The value is less than the group pair threshold, so it is excluded, then the remaining two channel pairs (L, R) and (R, C) can obtain two channel pair sets, one of which includes (L ,R) and the other includes (R,C).
- the method for acquiring the M channel pair sets in this embodiment may include: adding the first channel pair to the first channel pair.
- a channel pair set, the M channel pair sets include the first channel pair set, when the other channel pairs except the associated channel in the multiple channel pairs include channels whose correlation value is greater than the group pair threshold
- a channel pair whose correlation value is greater than the group pair threshold is included, select a channel pair with the largest correlation value from other channel pairs and add it to the first channel pair set.
- step b may be iteratively executed.
- the correlation value smaller than the pair pair threshold may be deleted from the correlation value set, so that the number of channel pairs can be reduced, thereby reducing the number of iterations.
- Step 305 Determine a target channel pair set from the M channel pair sets.
- the sum of the correlation values of all channel pairs in the target channel pair set is the largest among the M channel pair sets. After the above-mentioned M channel pair sets are obtained, the sum of the correlation values of all channel pairs included in each channel pair set can be calculated, and finally the channel pair set with the largest sum of the correlation values is determined as the target channel pair. gather.
- Step 306 Encode the first audio frame according to the target channel pair set.
- At least five audio channels in the first audio frame may be acquired.
- the channel signals are separately processed for energy equalization to obtain at least five equalized channel signals, and then stereo processing is performed on the at least five equalized channel signals.
- the encoded object is related to the equalized channel signals.
- the energy equalization mode may include a first energy equalization mode and/or a second energy equalization mode, wherein the first energy equalization mode only uses two channel signals in one channel pair to obtain two equalized channels corresponding to one channel pair Signal.
- the second energy equalization mode uses two channel signals in one channel pair and at least one channel signal outside one channel to obtain two equalized channel signals corresponding to one channel pair.
- the average value of the energy or amplitude values of the two channel signals included in the current channel pair may be calculated, and according to the average value Perform energy equalization processing on the two channel signals respectively to obtain two corresponding equalized channel signals.
- the fluctuation interval value of at least five channel signals is large, energy balance can be performed only between the two related channel signals, so that the allocation of bits during stereo processing is more in line with the energy characteristics of channel signals, avoiding In a low bit rate coding environment, the coding noise of the channel pair with high energy may be much larger than the coding noise of the channel pair with low energy due to insufficient bits, and the bits of the channel pair with low energy may be redundant.
- the energy equalization mode is the second energy equalization mode
- the average value of the energy or amplitude values of the at least five channel signals can be calculated, and the energy equalization processing is performed on the at least five channel signals according to the average value to obtain at least five equalized sound signals. channel signal.
- the target channel pair set contains The sum of the correlation values of all channel pairs is the largest, and the number of channel pairs in a group pair is increased as much as possible to reduce the redundancy between channel signals and improve the audio coding efficiency.
- FIG. 4 is an exemplary structural diagram of an encoding apparatus to which the encoding method for a multi-channel audio signal provided by the present application is applied.
- the encoding apparatus may be the encoder 20 of the source device 12 in the audio decoding system 10, or may be is the decoding module 270 in the audio decoding apparatus 200 .
- the encoding device may include a channel pair set generation module, a multi-channel processing module, a channel encoding module and a code stream multiplexing interface, wherein,
- the input of the channel pair set generation module is n channel signals (CH1-CHn) of multi-channel audio, where n is an integer greater than or equal to 5, and all the n channel signals can be processed in stereo.
- the channel pair set generation module calculates the correlation value between any two channel signals in the n channel signals, so as to obtain the target channel pair set by using the method of the embodiment shown in FIG. 3 according to these correlation values, for example, (CH1 , CH2), (CH3, CH4), ..., (CHi-1, CHi).
- the multi-channel processing module includes a plurality of stereo processing units, and the stereo processing unit can use prediction-based or Karhunen-Loeve Transform (Karhunen-Loeve Transform, KLT)-based processing, that is, the input two channel signals are rotated (for example, via 2 ⁇ 2 rotation matrix) to maximize energy compression, thereby concentrating the signal energy into one channel.
- KLT Karhunen-Loeve Transform
- Each channel pair in the target channel pair set output by the channel pair set generation module is respectively input to a stereo processing unit, for example, (CH1, CH2) input stereo processing unit 1, (CH3, CH4) input stereo processing unit 2 , ..., (CHi-1, CHi) are input to the stereo processing unit m.
- the stereo processing unit processes the input two channel signals, it outputs the processed channel signal (P) corresponding to the two channel signals and the multi-channel parameter (SIDE_PAIR).
- the multi-channel parameters include channel pair index, energy Equalization side information, stereo processing side information.
- the stereo processing unit 1 processes CH1 and CH2 to obtain P1 and P2 and SIDE_PAIR1
- the stereo processing unit 2 processes CH3 and CH4 to obtain P3 and P4, and SIDE_PAIR2, ...
- the stereo processing unit m pairs CHi-1 and CHi Process, get Pi-1 and Pi, and SIDE_PAIRm.
- the channel encoding module encodes the processed channel signal output by the multi-channel processing module by using a mono encoding unit (or a mono box or a mono tool) to output a corresponding encoded channel signal (E).
- a mono encoding unit or a mono box or a mono tool
- the channel signal with higher energy (or higher amplitude) is allocated more bits
- the channel signal with less energy (or lower amplitude) is allocated more bits. Allocate fewer bits.
- the channel encoding module may also use a stereo encoding unit, such as a parametric stereo encoder or a lossy stereo encoder, to encode the processed channel signal output by the multi-channel processing module.
- P1, P2, P3, P4, ..., Pi1, Pi are respectively encoded by a mono coding unit to obtain E1, E2, E3, E4, ..., Ei1, Ei.
- the unpaired channel signals (such as CHj) in the channel pair set generation module do not need to be processed by the stereo processing unit in the multi-channel processing module, and can be directly input into a single channel in the channel encoding module.
- the channel coding unit obtains Ej.
- the code stream multiplexing interface generates an encoded multi-channel signal, and the encoded multi-channel signal includes the encoded channel signal output by the channel encoding module and the multi-channel parameters output by the multi-channel processing module.
- the encoded multi-channel signal includes E1, E2, E3, E4, ..., Ei1, Ei, and SIDE_PAIR1, SIDE_PAIR2, ..., SIDE_PAIRm.
- the code stream multiplexing interface can process the encoded multi-channel signal into a serial signal or a serial bit stream.
- the processing flow for obtaining the target channel pair set provided by the present application may be implemented by the channel pair set generation module in the encoding apparatus shown in FIG. 4 .
- the 5.1 channel includes a center channel (C), a front left channel (left, L), a front right channel (right, R), and a left surround channel (left surround). , LS), right surround back channel (right surround, RS) and 0.1 channel low frequency effects (low frequency effects, LFE).
- the channel pair set generation module can use a multi-channel mask to remove the channels that do not need multi-channel processing to improve the coding efficiency.
- the LFE channel can be removed from the 5.1 channel, so the input sound
- the channel signals of the channel pair set generation module include C, L, R, LS and RS.
- the method for obtaining the target channel pair set may include the following steps:
- the present application can use the following formula to calculate the correlation value between two channel signals (for example, the channel signal ch1 and the channel signal ch2):
- corr_norm(ch1, ch2) represents the normalized correlation value between the channel signal ch1 and the channel signal ch2
- spec_ch1(i) represents the frequency domain coefficient of the ith frequency point of the channel signal ch1
- spec_ch2(i ) is the frequency domain coefficient of the ith frequency point of the channel signal ch2
- N represents the total number of frequency points of an audio frame.
- the obtained correlation value set may include at most Correlation values for each channel pair.
- Table 1 shows an example of a set of correlation values for 5.1 channels.
- the pairing threshold is set to 0.3, and only two channel signals with a correlation value greater than 0.3 can be paired. Therefore, delete the correlation values in Table 1 that are less than the pairing threshold to obtain Table 1a, which can be used in the iterative process. Channel signals with less correlation are not considered, thereby reducing the amount of computation.
- R, C is the first channel pair added to the first channel pair set, and the correlation value of the channel pair including R and/or C is deleted from Table 1a to obtain Table 1b.
- the maximum correlation value in Table 1b is 0.42 (LS, RS), so LS and RS are formed into a second channel pair and added to the first channel pair set. At this time, there is only one channel signal L left in the five channel signals, and the grouping cannot be continued. Therefore, the final first channel pairing set includes two channel pairs (R, C) and (LS, RS).
- (L, C) is the first channel pair added to the second channel pair set, and the correlation value of the channel pair including L and/or C is deleted from Table 1a to obtain Table 1c.
- the maximum correlation value in Table 1c is 0.42 (LS, RS), so LS and RS are formed into a second channel pair and added to the second channel pair set. At this time, there is only one channel signal R left in the five channel signals, and the grouping cannot be continued. Therefore, the final second channel pairing set includes two channel pairs (L, C) and (LS, RS).
- (LS, RS) is the first channel pair added to the third channel pair set, and the correlation value of the channel pair including LS and/or RS is deleted from Table 1a to obtain Table 1d.
- the maximum correlation value in Table 1d is 0.57(R, C), so R and C are combined into the second channel pair and added to the third channel pair set. At this time, there is only one channel signal L left in the five channel signals, and the grouping cannot be continued. Therefore, the final third channel pairing set includes two channel pairs (LS, RS) and (R, C).
- the channel pair set corresponding to (1) (or S(3)) is used as the target channel pair set, that is, the channel pairs available for 5.1 channels in this embodiment include (L, C) and (LS, RS).
- the target channel pair set can be represented by an index, and index values can be set for the channel pairs corresponding to all the correlation values in Table 1. After the target channel pair set is determined, the channel pairs in the target channel pair set can be used. The corresponding index value is indicated to save the number of bits in the code stream.
- the 7.1 channel includes C, L, R, LS, RS, a left back channel (LB), a right back channel (RB) and an LFE.
- the channel pair set generation module can use the multi-channel mask to remove the channels that do not need multi-channel processing to improve the coding efficiency.
- the LFE channel can be removed from the 7.1 channel, so the input sound
- the channel signals of the channel pair set generation module include C, L, R, LS, RS, LB and RB.
- the method for obtaining the target channel pair set may include the following steps:
- the formula of the above-mentioned first embodiment can also be used to calculate the correlation value between the two channel signals.
- the obtained correlation value set may include at most Correlation values for each channel pair.
- Table 2 shows an example of a set of correlation values for 7.1 channels.
- the pairing threshold is set to 0.3, that is, only two channel signals with a correlation value greater than 0.3 can be paired. Therefore, by deleting the correlation value in Table 2 that is less than the pairing threshold, Table 2a can be obtained. In this way, in the process of iterative processing Channel signals with less correlation can be ignored, thereby reducing the amount of computation.
- (LS, LB) is the first channel pair added to the first channel pair set, and the correlation value of the channel pair including LS and/or LB is deleted from Table 2a to obtain Table 2b.
- the maximum correlation value in Table 2b is 0.57(R, C), so R and C are combined to form a second channel pair and added to the first channel pair set. Correlation values for channel pairs containing R and/or C are removed from Table 2b, resulting in Table 2c.
- the final first channel pair set includes two channel pairs (LS,LB) and (R,C).
- (RS, LB) is the first channel pair added to the second channel pair set, and the correlation value of the channel pair including RS and/or LB is deleted from Table 2a to obtain Table 2d.
- the maximum correlation value in Table 2d is 0.57(R, C), so R and C are combined into a second channel pair and added to the second channel pair set. Correlation values for channel pairs containing R and/or C are removed from Table 2d, resulting in Table 2e.
- Table 2e The maximum correlation value in Table 2e is 0.39(L, LS), so L and LS are combined into a third channel pair and added to the second channel pair set. Correlation values for channel pairs containing L and/or LS are removed from Table 2e, resulting in Table 2f.
- the final first channel pair set includes three channel pairs (RS, LB), (R, C) and (L, LS).
- R, C is the first channel pair added to the third channel pair set, and the correlation value of the channel pair including R and/or C is deleted from Table 2a to obtain Table 2g.
- Table 2g The maximum correlation value in Table 2g is 0.67 (LS, LB), so LS and LB are formed into a second channel pair and added to the third channel pair set. Correlation values for channel pairs containing LS and/or LB are removed from Table 2g, resulting in Table 2h.
- the final first channel pair set includes two channel pairs (R,C) and (LS,LB).
- (L, C) is the first channel pair added to the fourth channel pair set, and the correlation value of the channel pair including L and/or C is deleted from Table 2a to obtain Table 2i.
- Table 2i The maximum correlation value in Table 2i is 0.67(LS, LB), so LS and LB are formed into the second channel pair and added to the fourth channel pair set.
- Table 2j is obtained by deleting the correlation values of channel pairs containing LS and/or LB from Table 2i.
- the final first channel pair set includes two channel pairs (L,C) and (LS,LB).
- S(2) is the largest among S(1), S(2), S(3) and S(4), so the channel pair set corresponding to S(2) is used as the target channel pair set, that is, this implementation
- the channel pairs available for the 7.1 channel in the example include (RS, LB), (R, C) and (L, LS).
- the second embodiment has one more iterative processing process, and the number of channel pairs included in the target channel pair set is also one more, which is related to the number of channel signals participating in the group pair.
- FIG. 5 is a flowchart of an exemplary embodiment of a method for encoding a multi-channel audio signal provided by the present application.
- the process 500 may be performed by the source device 12 or the audio coding device 200 in the audio coding system 10 .
- Process 500 is described as a series of steps or operations, and it should be understood that process 500 may be performed in various orders and/or concurrently, and is not limited to the order of execution shown in FIG. 5 .
- the method includes:
- Step 501 Obtain a first audio frame to be encoded.
- Step 502 Obtain a set of related values.
- steps 501 and 502 in this embodiment reference may be made to the above-mentioned steps 301 and 302, and details are not repeated here.
- Step 503 Acquire multiple channel pair sets according to the multiple channel pairs.
- the correlation value set includes correlation values of multiple channel pairs of at least five channel signals of the first audio frame, and the multiple channel pairs are regularly combined (that is, multiple sound channels in the same channel pair set are combined. Channel pairs cannot contain the same channel signal), and multiple channel pair sets corresponding to the at least five channel signals can be obtained.
- the following formula can be used to calculate the number of all channel pair sets:
- Pair_num represents the number of all channel pair sets
- CH represents the number of channel signals involved in multi-channel processing in the first audio frame, which is the result of filtering by the multi-channel mask.
- multiple channel pair sets can be obtained according to other channel pairs in the multiple channel pairs that are not related to the outside, and the correlation value of the uncorrelated channel pair can be obtained. is smaller than the group pair threshold, in this way, the number of channel pairs participating in the calculation can be reduced when obtaining the channel pair set, thereby reducing the number of channel pair sets, and the calculation amount of the sum of correlation values can also be reduced in subsequent steps.
- the channel signals whose correlation values with other channel signals are all smaller than the group pair threshold can be deleted, that is, such channel signals do not consider the group pair, and when the sound channel signal is obtained, the channel signal can be deleted.
- the channel pairs are set, the number of the channel pairs involved in the calculation can be reduced, thereby reducing the number of the channel pair sets, and the calculation amount of the sum of the correlation values can also be reduced in the subsequent steps.
- Step 504 Obtain the sum of the correlation values of all channel pairs included in each channel pair set in the multiple channel pair sets according to the correlation value set.
- the sum of the correlation values of all channel pairs contained in the channel pair set is calculated.
- Step 505 Determine the target channel pair set.
- Step 506 Encode the first audio frame according to the target channel pair set.
- steps 505 and 506 in this embodiment reference may be made to the above-mentioned steps 305 and 306, and details are not repeated here.
- the 5.1 channel includes C, L, R, LS, RS, and LFE.
- the channel pair set generation module can use the multi-channel mask to remove the channels that do not need multi-channel processing to improve the coding efficiency.
- the LFE channel can be removed from the 5.1 channel, so the input sound
- the channel signals of the channel pair set generation module include C, L, R, LS and RS.
- the method for obtaining the target channel pair set may include the following steps:
- the formula of the above-mentioned first embodiment can also be used to calculate the correlation value between the two channel signals.
- the obtained correlation value set may include at most The correlation values of each channel pair are shown in Table 1.
- 10 correlation values can be obtained from five channel signals, and correspondingly, 10 channel pairs can be obtained, and then the 10 channel pairs can be obtained at most A collection of channel pairs. For example, ⁇ (L,R),(LS,RS) ⁇ , ⁇ (L,R),(C,RS) ⁇ , ⁇ (L,R),(LS,C) ⁇ , .
- the correlation value of the channel pair may be set to 0.
- the channel pairs whose correlation value is less than the group pair threshold can be excluded, so that the number of channel pairs can be reduced when obtaining the channel pair set, thereby reducing the number of channel pairs.
- the number of channel pair sets can be excluded, so that the number of channel pairs can be reduced when obtaining the channel pair set, thereby reducing the number of channel pairs.
- FIG. 6 is a flowchart of an exemplary embodiment of a method for encoding a multi-channel audio signal provided by the present application.
- the process 600 may be performed by the source device 12 or the audio coding device 200 in the audio coding system 10 .
- Process 600 is described as a series of steps or operations, and it should be understood that process 600 may be performed in various orders and/or concurrently, and is not limited to the order of execution shown in FIG. 6 .
- the method includes:
- Step 601 Obtain a first audio frame to be encoded.
- step 601 reference may be made to the above-mentioned step 301, which will not be repeated here.
- Step 602 Obtain a correlation value set of the first audio frame.
- the correlation value set of the first audio frame includes respective correlation values of a plurality of channel pairs, one channel pair includes two channel signals in at least five channel signals, and the correlation value of one channel pair is used to represent an audio channel. Correlation between the two channel signals of a channel pair.
- Step 603 Obtain a correlation value set of the second audio frame.
- the set of correlation values of the second audio frame includes respective correlation values of a plurality of channel pairs of the second audio frame, one channel pair includes two channel signals among the at least five channel signals of the second audio frame, one channel pair
- the correlation value of the channel pair is used to represent the correlation between the two channel signals of one channel pair, and the second audio frame is the previous frame of the first audio frame.
- this embodiment in addition to acquiring the correlation value set of the first audio frame, this embodiment also needs to acquire the correlation value set of the previous frame (ie, the second audio frame) of the first audio frame.
- the encoding device Since the encoding of the second audio frame is prior to the encoding of the first audio frame, when the first audio frame is processed, the encoding device has already acquired relevant information when encoding the second audio frame, including the correlation value of the second audio frame Therefore, in this embodiment, the correlation value set of the second audio frame may be directly read from the cache or the memory, and the correlation value set of the second audio frame does not need to be calculated again.
- Step 604 Determine whether the target channel pair set of the first audio frame needs to be re-acquired according to the correlation value set of the first audio frame and the correlation value set of the second audio frame.
- the difference between the correlation value set of the first audio frame and the correlation value set of the second audio frame can be calculated as the judgment basis, that is, the correlation value set of the first audio frame and the correlation value of the second audio frame can be calculated.
- the absolute value of the difference between the correlation values corresponding to the same channel pair in the set is calculated, and the sum of the absolute values corresponding to the plurality of channel pairs is calculated.
- the sum of the absolute values is less than the change threshold, it is determined that the target channel pair set of the first audio frame does not need to be re-acquired; when the sum of the absolute values is greater than or equal to the change threshold, it is determined that the target channel of the first audio frame needs to be re-acquired pair collection.
- the correlation value difference corresponds to the same channel pair, calculate the correlation value difference respectively, and then calculate the sum of the absolute values of the difference values of all channel pairs, so that the first audio frame relative to the second audio frame, the signal of each channel can be obtained.
- the change of the correlation value between the two exceeds the change threshold, if not, it means that the change from the second audio frame to the first audio frame is not large, and it is not necessary to rebuild the target channel pair set for the first audio frame, which reduces the calculation If it exceeds, it means that the change from the second audio frame to the first audio frame is large, and the target channel pair set of the first audio frame needs to be re-acquired.
- Step 605 if it is necessary to re-acquire the target channel pair set of the first audio frame, then adopt the method of the embodiment shown in FIG. 3 or FIG. 5 to obtain the target channel pair set of the first audio frame, and according to the target channel pair set The first audio frame is encoded.
- the method in the embodiment shown in FIG. 3 or FIG. 5 may be used to obtain the correlation value set of the first audio frame, which will not be repeated here.
- Step 606 if it is not necessary to re-acquire the target channel pair set of the first audio frame, then determine the target channel pair set of the second audio frame as the target channel pair set of the first audio frame, and according to the target channel pair set.
- the set encodes the first audio frame.
- the target channel pair set of the second audio frame when it is determined that the target channel pair set of the first audio frame does not need to be re-acquired, the target channel pair set of the second audio frame can be directly used as the target channel pair set of the first audio frame, thereby reducing the amount of calculation. Improve coding efficiency.
- the channel pair set corresponding to the sum of the values is determined as the target channel pair set, which can maximize the sum of the correlation values of all channel pairs included in the target channel pair set, and increase the number of channel pairs in the group pair as much as possible. It reduces the redundancy between channel signals and improves the coding efficiency of audio.
- the 5.1 channel includes C, L, R, LS, RS, and LFE.
- the channel pair set generation module can use the multi-channel mask to remove the channels that do not need multi-channel processing to improve the coding efficiency.
- the LFE channel can be removed from the 5.1 channel, so the input sound
- the channel signals of the channel pair set generation module include C, L, R, LS and RS.
- the method for obtaining the target channel pair set may include the following steps:
- the formula of the above-mentioned first embodiment can also be used to calculate the correlation value between the two channel signals.
- the obtained correlation value set may include at most The correlation values of each channel pair are shown in Table 1.
- the correlation value set of the first audio frame and the correlation value set of the second audio frame are both represented in the form of matrices, and matrices Matrix1 and Matrix2 are obtained respectively, and the value of each element in the matrix corresponds to the value of the correlation value set in the correlation value set.
- a correlation value, the sum of the differences can be calculated by the following formula:
- D represents the sum of the difference between the correlation value set of the first audio frame and the correlation value set of the second audio frame
- Matrix1(i) represents the i-th element value in the matrix corresponding to the correlation value set of the first audio frame
- Matrix2(i) represents the i-th element value in the matrix corresponding to the correlation value set of the second audio frame.
- a change threshold is set, and whether the target channel pair set of the first audio frame needs to be re-acquired is defined by the threshold.
- FIG. 7 is a flowchart of an exemplary embodiment of a method for encoding a multi-channel audio signal provided by the present application.
- the process 700 may be performed by the source device 12 or the audio coding device 200 in the audio coding system 10 .
- Process 700 is described as a series of steps or operations, and it should be understood that process 700 may be performed in various orders and/or concurrently, and is not limited to the order of execution shown in FIG. 7 .
- the method includes:
- Step 701 Acquire a first audio frame to be encoded, where the first audio frame includes K channel signals.
- step 701 reference may be made to the above-mentioned step 301, which will not be repeated here.
- Step 702 When K is greater than the threshold of the number of channel signals, use the method of the embodiment shown in FIG. 3 to encode the first audio frame.
- Step 703 When K is less than or equal to the threshold of the number of channel signals, use the method of the embodiment shown in FIG. 5 to encode the first audio frame.
- the difference between this embodiment and the embodiment shown in FIG. 3 or FIG. 5 is that the method in FIG. 3 and FIG. 5 is combined in this embodiment, that is, according to the number of channel signals contained in the first audio frame Which method an audio frame uses to obtain its target channel pair set.
- the method of the second aspect is adopted, it is necessary to exhaustively enumerate all target channel pairs, which will increase the amount of calculation. Therefore, the method of the first aspect will reduce the amount of calculation. A lot of computation.
- the method of the second aspect can be used to obtain the sum of the correlation values of all channel pair sets, so as to ensure that the final selected target channel pair set must be the highest Optimal results that match the characteristics of the first audio frame.
- FIG. 8 is an exemplary structural diagram of a decoding apparatus to which the decoding method for a multi-channel audio signal provided by the present application is applied.
- the decoding apparatus may be the decoder 30 of the destination device 14 in the audio decoding system 10, or may be is the decoding module 270 in the audio decoding apparatus 200 .
- the decoding device may include a code stream demultiplexing interface, a channel decoding module and a multi-channel processing module, wherein,
- the code stream demultiplexing interface receives the encoded multi-channel signal (eg serial bit stream bitstream) from the encoding device, and obtains the encoded channel signal (E) and multi-channel parameters (SIDE_PAIR) after demultiplexing.
- E encoded channel signal
- SIDE_PAIR multi-channel parameters
- the channel decoding module uses a monaural decoding unit (or a monaural box, a monaural tool) to decode the coded channel signal output by the code stream demultiplexing interface and output the decoded channel signal (D).
- a monaural decoding unit or a monaural box, a monaural tool
- D decoded channel signal
- the multi-channel processing module includes a plurality of stereo processing units.
- the stereo processing unit can adopt prediction-based or KLT-based processing, that is, the input two channel signals are inversely rotated (for example, via a 2 ⁇ 2 rotation matrix), so that the signal Transform to the original signal direction.
- the decoded channel signal output by the channel decoding module can identify which two decoded channel signal groups are paired by the multi-channel parameters, and input the decoded channel signal of the pair into the stereo processing unit, and the stereo processing unit decodes the two input channel signals.
- the channel signal (CH) corresponding to the two decoded channel signals is output.
- stereo processing unit 1 processes D1 and D2 according to SIDE_PAIR1 to obtain CH1 and CH2
- stereo processing unit 2 processes D3 and D4 according to SIDE_PAIR2 to obtain CH3 and CH4, ...
- the unpaired channel signal (such as CHj) does not need to be processed by the stereo processing unit in the multi-channel processing module, and can be directly output after decoding.
- FIG. 9 is a schematic structural diagram of an embodiment of an encoding apparatus of the present application. As shown in FIG. 9 , the apparatus may be applied to the source device 12 or the audio decoding device 200 in the above-mentioned embodiment.
- the encoding apparatus in this embodiment may include: an acquisition module 901 , an encoding module 902 and a determination module 903 .
- the obtaining module 901 is configured to obtain a first audio frame to be encoded, where the first audio frame includes at least five channel signals; obtain a correlation value set, where the correlation value set includes multiple Correlation values of each channel pair, where one channel pair includes two channel signals in the at least five channel signals, and the correlation value of the one channel pair is used to represent the two channel signals of the one channel pair.
- M correlation values are selected from the correlation value set, and the M correlation values are all greater than other correlation values in the correlation value set except the M correlation values,
- the M correlation values are all greater than or equal to the group pair threshold, and M is a positive integer less than or equal to a set value; acquire M channel pair sets, each of which includes at least the M correlation values one of the corresponding M channel pairs, and when the channel pair set includes more than two channel pairs, the two or more channel pairs do not contain the same channel signal; the determining module 903, using In determining a target channel pair set from the M channel pair sets, the sum of the correlation values of all channel pairs in the target channel pair set is the largest in the M channel pair sets; encoding Module 902, configured to encode the first audio frame according to the target channel pair set.
- the M channel pair sets include a first channel pair set; the acquiring module 901 is specifically configured to add the first channel pair in the M channel pairs The first channel pair set, the first channel pair is any one of the M channel pairs; when the other channel pairs except the associated channel in the multiple channel pairs include: When the channel pair whose correlation value is greater than the set pair threshold, select a channel pair with the largest correlation value from the other channel pairs and add it to the first channel pair set, where the associated channel pair includes the added channel pair. Any one of the channel signals included in the channel pair of the first channel pair set.
- the obtaining module 901 is specifically configured to select N correlation values from the correlation value set, and the N correlation values are all greater than the N correlation values in the correlation value set divided by the N correlation values.
- N is the set value; from the N correlation values, select the correlation value greater than or equal to the threshold of the group pair, and the correlation value greater than or equal to the threshold value of the group pair The number is M.
- the correlation value is a normalized value.
- the correlation value of the one channel pair is set to 0.
- the obtaining module 901 is configured to obtain a first audio frame to be encoded, where the first audio frame includes at least five channel signals; obtain a correlation value set, where the correlation value set includes multiple Correlation values of the respective channel pairs, one channel pair includes two channel signals in the at least five channel signals, and the correlation value of the one channel pair is used to represent the two channel signals of the one channel pair.
- multiple channel pair sets are obtained according to the multiple channel pairs, when the channel pair set includes more than two channel pairs, the two or more channel pairs does not contain the same channel signal; obtains the sum of the correlation values of all channel pairs included in each channel pair set in the multiple channel pair sets according to the correlation value set; the determining module 903 is used to determine the target A channel pair set, where the sum of the correlation values of all channel pairs in the target channel pair set is the largest among the multiple channel pair sets; the encoding module 902 is configured to, according to the target channel pair set The first audio frame is encoded.
- the obtaining module 901 is specifically configured to obtain the set of multiple channel pairs according to other channel pairs in the multiple channel pairs that are not related to the outside Correlation values for related channel pairs are less than the group pair threshold.
- the obtaining module 901 is configured to obtain a first audio frame to be encoded, where the first audio frame includes at least five channel signals; obtain a correlation value set of the first audio frame, The correlation value set of the first audio frame includes respective correlation values of a plurality of channel pairs, one channel pair includes two channel signals in the at least five channel signals, and the correlation value of the one channel pair is The value is used to represent the correlation between the two channel signals of the one channel pair; the correlation value set of the second audio frame is obtained, and the correlation value set of the second audio frame includes the correlation value set of the second audio frame.
- Correlation values for each of a plurality of channel pairs where one channel pair includes two channel signals in at least five channel signals of the second audio frame, and the correlation value of the one channel pair is used to represent the Correlation between two channel signals of a channel pair, the second audio frame is the previous frame of the first audio frame; the encoding module 902 is configured to use the correlation value of the first audio frame according to The set and the correlation value set of the second audio frame determine whether it is necessary to re-acquire the target channel pair set of the first audio frame; if it is necessary to re-acquire the target channel pair set of the first audio frame, execute the 3 or the method of the embodiment shown in FIG.
- the 5 acquires the target channel pair set of the first audio frame, and encodes the first audio frame according to the target channel pair set;
- the target channel pair set of the first audio frame, then the target channel pair set of the second audio frame is determined as the target channel pair set of the first audio frame, and according to the target channel pair set
- the first audio frame is encoded.
- the encoding module 902 is specifically configured to calculate the correlation value corresponding to the same channel pair in the correlation value set of the first audio frame and the correlation value set of the second audio frame The absolute value of the difference; calculate the sum of the absolute values corresponding to a plurality of the channel pairs respectively; when the sum of the absolute values is less than the change threshold, it is determined that the target sound of the first audio frame does not need to be re-acquired A channel pair set; when the sum of the absolute values is greater than or equal to the change threshold, it is determined that the target channel pair set of the first audio frame needs to be re-acquired.
- the acquisition module is used to acquire the first audio frame to be encoded, the first audio frame includes K channel signals, and K is an integer greater than or equal to 5; the encoding module is used for When K is greater than the threshold of the number of channel signals, the method of the embodiment shown in FIG. 3 is performed to encode the first audio frame; when K is less than or equal to the threshold of the number of channel signals, the method of the embodiment shown in FIG. 5 is performed The first audio frame is encoded.
- the apparatus of this embodiment can be used to implement the technical solutions of the method embodiments shown in FIG. 3 , FIG. 5 , FIG. 6 or FIG. 7 , and the implementation principles and technical effects thereof are similar, and are not repeated here.
- FIG. 10 is a schematic structural diagram of an embodiment of a device of the present application.
- the device may be the encoding device in the above-mentioned embodiment.
- the device in this embodiment may include: a processor 1001 and a memory 1002, where the memory 1002 is used to store one or more programs; when the one or more programs are executed by the processor 1001, the processor 1001 realizes the The technical solution of the method embodiment shown in FIG. 3 , FIG. 5 , FIG. 6 or FIG. 7 .
- each step of the above method embodiments may be completed by a hardware integrated logic circuit in a processor or an instruction in the form of software.
- the processor can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other Programming logic devices, discrete gate or transistor logic devices, discrete hardware components.
- DSP digital signal processor
- ASIC application-specific integrated circuit
- FPGA field programmable gate array
- a general purpose processor may be a microprocessor or the processor may be any conventional processor or the like.
- the steps of the method disclosed in the present application can be directly embodied as executed by a hardware encoding processor, or executed by a combination of hardware and software modules in the encoding processor.
- the software modules may be located in random access memory, flash memory, read-only memory, programmable read-only memory or electrically erasable programmable memory, registers and other storage media mature in the art.
- the storage medium is located in the memory, and the processor reads the information in the memory, and completes the steps of the above method in combination with its hardware.
- the memory mentioned in the above embodiments may be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory.
- the non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically programmable Erase programmable read-only memory (electrically EPROM, EEPROM) or flash memory.
- Volatile memory may be random access memory (RAM), which acts as an external cache.
- RAM random access memory
- DRAM dynamic random access memory
- SDRAM synchronous DRAM
- SDRAM double data rate synchronous dynamic random access memory
- ESDRAM enhanced synchronous dynamic random access memory
- SLDRAM synchronous link dynamic random access memory
- direct rambus RAM direct rambus RAM
- the disclosed system, apparatus and method may be implemented in other manners.
- the apparatus embodiments described above are only illustrative.
- the division of the units is only a logical function division. In actual implementation, there may be other division methods.
- multiple units or components may be combined or Can be integrated into another system, or some features can be ignored, or not implemented.
- the shown or discussed mutual coupling or direct coupling or communication connection may be through some interfaces, indirect coupling or communication connection of devices or units, and may be in electrical, mechanical or other forms.
- the units described as separate components may or may not be physically separated, and components displayed as units may or may not be physical units, that is, may be located in one place, or may be distributed to multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution in this embodiment.
- each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically alone, or two or more units may be integrated into one unit.
- the functions, if implemented in the form of software functional units and sold or used as independent products, may be stored in a computer-readable storage medium.
- the technical solution of the present application can be embodied in the form of a software product in essence, or the part that contributes to the prior art or the part of the technical solution.
- the computer software product is stored in a storage medium, including Several instructions are used to cause a computer device (personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application.
- the aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), magnetic disk or optical disk and other media that can store program codes .
Landscapes
- Engineering & Computer Science (AREA)
- Physics & Mathematics (AREA)
- Computational Linguistics (AREA)
- Signal Processing (AREA)
- Health & Medical Sciences (AREA)
- Audiology, Speech & Language Pathology (AREA)
- Human Computer Interaction (AREA)
- Acoustics & Sound (AREA)
- Multimedia (AREA)
- Mathematical Physics (AREA)
- Compression, Expansion, Code Conversion, And Decoders (AREA)
Abstract
Description
| 声道信号\相关值 | R | C | LS | RS |
| L | 0.36 | 0.47 | 0.39 | 0.27 |
| R | 0.57 | 0.22 | 0.08 | |
| C | 0.31 | 0.26 | ||
| LS | 0.42 |
| 声道信号\相关值 | R | C | LS | RS | LB | RB |
| L | 0.36 | 0.47 | 0.39 | 0.27 | 0.43 | 0.24 |
| R | 0.57 | 0.22 | 0.08 | 0.19 | 0.21 | |
| C | 0.31 | 0.26 | 0.36 | 0.07 | ||
| LS | 0.42 | 0.67 | 0.03 | |||
| RS | 0.64 | 0.07 | ||||
| LB | 0.19 |
Claims (27)
- 一种多声道音频信号的编码方法,其特征在于,包括:获取待编码的第一音频帧,所述第一音频帧包括至少五个声道信号;获取相关值集合,所述相关值集合包括多个声道对各自的相关值,一个声道对包括所述至少五个声道信号中的两个声道信号,所述一个声道对的相关值用于表示所述一个声道对的两个声道信号之间的相关性;从所述相关值集合中选取M个相关值,所述M个相关值均大于所述相关值集合中除所述M个相关值外的其他相关值,所述M个相关值均大于或等于组对阈值,M为小于或等于设定值的正整数;获取M个声道对集合,每个所述声道对集合包括与所述M个相关值对应的一个或多个声道对,且当所述声道对集合包括两个以上声道对时,所述两个以上声道对不包含相同的声道信号;从所述M个声道对集合中确定目标声道对集合,所述目标声道对集合中的所有声道对的相关值之和是所述M个声道对集合中最大的;根据所述目标声道对集合对所述第一音频帧进行编码。
- 根据权利要求1所述的方法,其特征在于,所述M个声道对集合包括第一声道对集合,所述获取M个声道对集合包括获取所述第一声道对集合;所述获取所述第一声道对集合,包括:将所述M个声道对中的第一声道对加入所述第一声道对集合,所述第一声道对为所述M个声道对中的任意一个;当所述多个声道对中除关联声道对外的其他声道对中包括相关值大于所述组对阈值的声道对时,从所述其他声道对中选取相关值最大的一个声道对加入所述第一声道对集合,所述关联声道对包括已加入所述第一声道对集合的声道对所包括的声道信号中的任意一个。
- 根据权利要求1或2所述的方法,其特征在于,所述从所述相关值集合中选取M个相关值,包括:从所述相关值集合中选取N个相关值,所述N个相关值均大于所述相关值集合中除所述N个相关值外的其他相关值,N为所述设定值;从所述N个相关值中选取大于或等于所述组对阈值的相关值,所述大于或等于所述组对阈值的相关值的个数为M。
- 根据权利要求1-3中任一项所述的方法,其特征在于,所述相关值为经归一化处理的值。
- 根据权利要求1-4中任一项所述的方法,其特征在于,当所述一个声道对的相关值小于所述组对阈值时,所述一个声道对的相关值设置为0。
- 一种多声道音频信号的编码方法,其特征在于,包括:获取待编码的第一音频帧,所述第一音频帧包括至少五个声道信号;获取相关值集合,所述相关值集合包括多个声道对各自的相关值,一个声道对包括所述至少五个声道信号中的两个声道信号,所述一个声道对的相关值用于表示所述一个声道 对的两个声道信号之间的相关性;根据所述多个声道对获取多个声道对集合,当所述声道对集合包括两个以上声道对时,所述两个以上声道对不包含相同的声道信号;根据所述相关值集合获取所述多个声道对集合中每一个声道对集合包含的所有声道对的相关值之和;确定目标声道对集合,所述目标声道对集合中的所有声道对的相关值之和是所述多个声道对集合中最大的;根据所述目标声道对集合对所述第一音频帧进行编码。
- 根据权利要求6所述的方法,其特征在于,所述根据所述多个声道对获取多个声道对集合,包括:根据所述多个声道对中除非相关声道对外的其他声道对获取所述多个声道对集合,所述非相关声道对的相关值小于组对阈值。
- 根据权利要求6或5所述的方法,其特征在于,所述相关值为经归一化处理的值。
- 根据权利要求6-8中任一项所述的方法,其特征在于,当所述一个声道对的相关值小于组对阈值时,所述一个声道对的相关值设置为0。
- 一种多声道音频信号的编码方法,其特征在于,包括:获取待编码的第一音频帧,所述第一音频帧包括至少五个声道信号;获取所述第一音频帧的相关值集合,所述第一音频帧的相关值集合包括多个声道对各自的相关值,一个声道对包括所述至少五个声道信号中的两个声道信号,所述一个声道对的相关值用于表示所述一个声道对的两个声道信号之间的相关性;获取第二音频帧的相关值集合,所述第二音频帧的相关值集合包括所述第二音频帧的多个声道对各自的相关值,一个声道对包括所述第二音频帧的至少五个声道信号中的两个声道信号,所述一个声道对的相关值用于表示所述一个声道对的两个声道信号之间的相关性,所述第二音频帧是所述第一音频帧的上一帧;根据所述第一音频帧的相关值集合和所述第二音频帧的相关值集合判断是否需要重新获取所述第一音频帧的目标声道对集合;若需要重新获取所述第一音频帧的目标声道对集合,则采用如权利要求1-9中任一项所述的方法获取所述第一音频帧的目标声道对集合,并根据所述目标声道对集合对所述第一音频帧进行编码;若不需要重新获取所述第一音频帧的目标声道对集合,则将所述第二音频帧的目标声道对集合确定为所述第一音频帧的目标声道对集合,并根据所述目标声道对集合对所述第一音频帧进行编码。
- 根据权利要求10所述的方法,其特征在于,所述根据所述第一音频帧的相关值集合和所述第二音频帧的相关值集合判断是否需要重新获取所述第一音频帧的目标声道对集合,包括:计算所述第一音频帧的相关值集合和所述第二音频帧的相关值集合中对应于同一声道对的相关值之差的绝对值;计算多个所述声道对分别对应的所述绝对值之和;当所述绝对值之和小于变更阈值时,确定不需要重新获取所述第一音频帧的目标声道 对集合;当所述绝对值之和大于或等于所述变更阈值时,确定需要重新获取所述第一音频帧的目标声道对集合。
- 一种多声道音频信号的编码方法,其特征在于,包括:获取待编码的第一音频帧,所述第一音频帧包括K个声道信号,K为大于或等于5的整数;当K大于声道信号数量阈值时,采用权利要求1-5中任一项所述的方法对所述第一音频帧进行编码;当K小于或等于声道信号数量阈值时,采用权利要求6-9中任一项所述的方法对所述第一音频帧进行编码。
- 一种编码装置,其特征在于,包括:获取模块,用于获取待编码的第一音频帧,所述第一音频帧包括至少五个声道信号;获取相关值集合,所述相关值集合包括多个声道对各自的相关值,一个声道对包括所述至少五个声道信号中的两个声道信号,所述一个声道对的相关值用于表示所述一个声道对的两个声道信号之间的相关性;从所述相关值集合中选取M个相关值,所述M个相关值均大于所述相关值集合中除所述M个相关值外的其他相关值,所述M个相关值均大于或等于组对阈值,M为小于或等于设定值的正整数;获取M个声道对集合,每个所述声道对集合至少包括所述M个相关值对应的M个声道对的其中之一,且当所述声道对集合包括两个以上声道对时,所述两个以上声道对不包含相同的声道信号;确定模块,用于从所述M个声道对集合中确定目标声道对集合,所述目标声道对集合中的所有声道对的相关值之和是所述M个声道对集合中最大的;编码模块,用于根据所述目标声道对集合对所述第一音频帧进行编码。
- 根据权利要求13所述的装置,其特征在于,所述M个声道对集合包括第一声道对集合;所述获取模块,具体用于将所述M个声道对中的第一声道对加入所述第一声道对集合,所述第一声道对为所述M个声道对中的任意一个;当所述多个声道对中除关联声道对外的其他声道对中包括相关值大于所述组对阈值的声道对时,从所述其他声道对中选取相关值最大的一个声道对加入所述第一声道对集合,所述关联声道对包括已加入所述第一声道对集合的声道对所包括的声道信号中的任意一个。
- 根据权利要求13或14所述的装置,其特征在于,所述获取模块,具体用于从所述相关值集合中选取N个相关值,所述N个相关值均大于所述相关值集合中除所述N个相关值外的其他相关值,N为所述设定值;从所述N个相关值中选取大于或等于所述组对阈值的相关值,所述大于或等于所述组对阈值的相关值的个数为M。
- 根据权利要求13-15中任一项所述的装置,其特征在于,所述相关值为经归一化处理的值。
- 根据权利要求13-16中任一项所述的装置,其特征在于,当所述一个声道对的相关值小于所述组对阈值时,所述一个声道对的相关值设置为0。
- 一种编码装置,其特征在于,包括:获取模块,用于获取待编码的第一音频帧,所述第一音频帧包括至少五个声道信号;获取相关值集合,所述相关值集合包括多个声道对各自的相关值,一个声道对包括所述至 少五个声道信号中的两个声道信号,所述一个声道对的相关值用于表示所述一个声道对的两个声道信号之间的相关性;根据所述多个声道对获取多个声道对集合,当所述声道对集合包括两个以上声道对时,所述两个以上声道对不包含相同的声道信号;根据所述相关值集合获取所述多个声道对集合中每一个声道对集合包含的所有声道对的相关值之和;确定模块,用于确定目标声道对集合,所述目标声道对集合中的所有声道对的相关值之和是所述多个声道对集合中最大的;编码模块,用于根据所述目标声道对集合对所述第一音频帧进行编码。
- 根据权利要求18所述的装置,其特征在于,所述获取模块,具体用于根据所述多个声道对中除非相关声道对外的其他声道对获取所述多个声道对集合,所述非相关声道对的相关值小于组对阈值。
- 根据权利要求18或19所述的装置,其特征在于,所述相关值为经归一化处理的值。
- 根据权利要求18-20中任一项所述的装置,其特征在于,当所述一个声道对的相关值小于组对阈值时,所述一个声道对的相关值设置为0。
- 一种编码装置,其特征在于,包括:获取模块,用于获取待编码的第一音频帧,所述第一音频帧包括至少五个声道信号;获取所述第一音频帧的相关值集合,所述第一音频帧的相关值集合包括多个声道对各自的相关值,一个声道对包括所述至少五个声道信号中的两个声道信号,所述一个声道对的相关值用于表示所述一个声道对的两个声道信号之间的相关性;获取第二音频帧的相关值集合,所述第二音频帧的相关值集合包括所述第二音频帧的多个声道对各自的相关值,一个声道对包括所述第二音频帧的至少五个声道信号中的两个声道信号,所述一个声道对的相关值用于表示所述一个声道对的两个声道信号之间的相关性,所述第二音频帧是所述第一音频帧的上一帧;编码模块,用于根据所述第一音频帧的相关值集合和所述第二音频帧的相关值集合判断是否需要重新获取所述第一音频帧的目标声道对集合;若需要重新获取所述第一音频帧的目标声道对集合,则执行如权利要求1-9中任一项所述的方法获取所述第一音频帧的目标声道对集合,并根据所述目标声道对集合对所述第一音频帧进行编码;若不需要重新获取所述第一音频帧的目标声道对集合,则将所述第二音频帧的目标声道对集合确定为所述第一音频帧的目标声道对集合,并根据所述目标声道对集合对所述第一音频帧进行编码。
- 根据权利要求22所述的装置,其特征在于,所述编码模块,具体用于计算所述第一音频帧的相关值集合和所述第二音频帧的相关值集合中对应于同一声道对的相关值之差的绝对值;计算多个所述声道对分别对应的所述绝对值之和;当所述绝对值之和小于变更阈值时,确定不需要重新获取所述第一音频帧的目标声道对集合;当所述绝对值之和大于或等于所述变更阈值时,确定需要重新获取所述第一音频帧的目标声道对集合。
- 一种编码装置,其特征在于,包括:获取模块,用于获取待编码的第一音频帧,所述第一音频帧包括K个声道信号,K为大于或等于5的整数;编码模块,用于当K大于声道信号数量阈值时,执行如权利要求1-5中任一项所述的方法对所述第一音频帧进行编码;当K小于或等于声道信号数量阈值时,执行如权利要求 6-9中任一项所述的方法对所述第一音频帧进行编码。
- 一种设备,其特征在于,包括:一个或多个处理器;存储器,用于存储一个或多个程序;当所述一个或多个程序被所述一个或多个处理器执行,使得所述一个或多个处理器实现如权利要求1-11中任一项所述的方法。
- 一种计算机可读存储介质,其特征在于,包括计算机程序,所述计算机程序在计算机上被执行时,使得所述计算机执行权利要求1-11中任一项所述的方法。
- 一种计算机可读存储介质,其特征在于,包括根据如权利要求1-11中任一项所述的多声道音频信号的编码方法获得的编码码流。
Priority Applications (4)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| EP21843116.1A EP4174855A4 (en) | 2020-07-17 | 2021-07-13 | METHOD AND DEVICE FOR ENCODING/DECODING A MULTI-CHANNEL AUDIO SIGNAL |
| JP2023502888A JP7519531B2 (ja) | 2020-07-17 | 2021-07-13 | マルチチャネルオーディオ信号符号化および復号方法および装置 |
| KR1020237004819A KR102948552B1 (ko) | 2020-07-17 | 2021-07-13 | 다중 채널 오디오 신호 인코딩 및 디코딩 방법 및 장치 |
| US18/153,128 US12437767B2 (en) | 2020-07-17 | 2023-01-11 | Multi-channel audio signal encoding and decoding method and apparatus |
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN202010699706.7A CN113948095B (zh) | 2020-07-17 | 2020-07-17 | 多声道音频信号的编解码方法和装置 |
| CN202010699706.7 | 2020-07-17 |
Related Child Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| US18/153,128 Continuation US12437767B2 (en) | 2020-07-17 | 2023-01-11 | Multi-channel audio signal encoding and decoding method and apparatus |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2022012553A1 true WO2022012553A1 (zh) | 2022-01-20 |
Family
ID=79326898
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2021/106101 Ceased WO2022012553A1 (zh) | 2020-07-17 | 2021-07-13 | 多声道音频信号的编解码方法和装置 |
Country Status (6)
| Country | Link |
|---|---|
| US (1) | US12437767B2 (zh) |
| EP (1) | EP4174855A4 (zh) |
| JP (1) | JP7519531B2 (zh) |
| KR (1) | KR102948552B1 (zh) |
| CN (1) | CN113948095B (zh) |
| WO (1) | WO2022012553A1 (zh) |
Families Citing this family (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN116434760A (zh) * | 2023-04-14 | 2023-07-14 | 北京小米移动软件有限公司 | 一种音频编码方法、装置、电子设备及存储介质 |
| CN116564319B (zh) * | 2023-05-10 | 2025-12-09 | 北京达佳互联信息技术有限公司 | 音频处理方法、装置、电子设备及存储介质 |
| CN117730367A (zh) * | 2023-10-31 | 2024-03-19 | 北京小米移动软件有限公司 | 分组方法、编码器、解码器以及存储介质 |
Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20090112606A1 (en) * | 2007-10-26 | 2009-04-30 | Microsoft Corporation | Channel extension coding for multi-channel source |
| CN101695150A (zh) * | 2009-10-12 | 2010-04-14 | 清华大学 | 多声道音频编码方法、编码器、解码方法和解码器 |
| US20110123031A1 (en) * | 2009-05-08 | 2011-05-26 | Nokia Corporation | Multi channel audio processing |
| CN104364842A (zh) * | 2012-04-18 | 2015-02-18 | 诺基亚公司 | 立体声音频信号编码器 |
| CN109416912A (zh) * | 2016-06-30 | 2019-03-01 | 杜塞尔多夫华为技术有限公司 | 一种对多声道音频信号进行编码和解码的装置和方法 |
Family Cites Families (14)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US8359196B2 (en) * | 2007-12-28 | 2013-01-22 | Panasonic Corporation | Stereo sound decoding apparatus, stereo sound encoding apparatus and lost-frame compensating method |
| AU2015246158B2 (en) | 2009-03-17 | 2017-10-26 | Dolby International Ab | Advanced stereo coding based on a combination of adaptively selectable left/right or mid/side stereo coding and of parametric stereo coding. |
| US9197978B2 (en) * | 2009-03-31 | 2015-11-24 | Panasonic Intellectual Property Management Co., Ltd. | Sound reproduction apparatus and sound reproduction method |
| US9100766B2 (en) * | 2009-10-05 | 2015-08-04 | Harman International Industries, Inc. | Multichannel audio system having audio channel compensation |
| JP2015011076A (ja) | 2013-06-26 | 2015-01-19 | 日本放送協会 | 音響信号符号化装置、音響信号符号化方法、および音響信号復号化装置 |
| TWI774136B (zh) | 2013-09-12 | 2022-08-11 | 瑞典商杜比國際公司 | 多聲道音訊系統中之解碼方法、解碼裝置、包含用於執行解碼方法的指令之非暫態電腦可讀取的媒體之電腦程式產品、包含解碼裝置的音訊系統 |
| CN105898667A (zh) * | 2014-12-22 | 2016-08-24 | 杜比实验室特许公司 | 从音频内容基于投影提取音频对象 |
| EP3067885A1 (en) * | 2015-03-09 | 2016-09-14 | Fraunhofer-Gesellschaft zur Förderung der angewandten Forschung e.V. | Apparatus and method for encoding or decoding a multi-channel signal |
| CN109389984B (zh) * | 2017-08-10 | 2021-09-14 | 华为技术有限公司 | 时域立体声编解码方法和相关产品 |
| CN109389987B (zh) * | 2017-08-10 | 2022-05-10 | 华为技术有限公司 | 音频编解码模式确定方法和相关产品 |
| CN109389985B (zh) * | 2017-08-10 | 2021-09-14 | 华为技术有限公司 | 时域立体声编解码方法和相关产品 |
| ES3059239T3 (en) | 2018-07-04 | 2026-03-19 | Fraunhofer Ges Forschung | Multisignal encoder, multisignal decoder, and related methods using signal whitening or signal post processing |
| WO2020164751A1 (en) * | 2019-02-13 | 2020-08-20 | Fraunhofer-Gesellschaft zur Förderung der angewandten Forschung e.V. | Decoder and decoding method for lc3 concealment including full frame loss concealment and partial frame loss concealment |
| EP3719799A1 (en) * | 2019-04-04 | 2020-10-07 | FRAUNHOFER-GESELLSCHAFT zur Förderung der angewandten Forschung e.V. | A multi-channel audio encoder, decoder, methods and computer program for switching between a parametric multi-channel operation and an individual channel operation |
-
2020
- 2020-07-17 CN CN202010699706.7A patent/CN113948095B/zh active Active
-
2021
- 2021-07-13 EP EP21843116.1A patent/EP4174855A4/en active Pending
- 2021-07-13 KR KR1020237004819A patent/KR102948552B1/ko active Active
- 2021-07-13 JP JP2023502888A patent/JP7519531B2/ja active Active
- 2021-07-13 WO PCT/CN2021/106101 patent/WO2022012553A1/zh not_active Ceased
-
2023
- 2023-01-11 US US18/153,128 patent/US12437767B2/en active Active
Patent Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20090112606A1 (en) * | 2007-10-26 | 2009-04-30 | Microsoft Corporation | Channel extension coding for multi-channel source |
| US20110123031A1 (en) * | 2009-05-08 | 2011-05-26 | Nokia Corporation | Multi channel audio processing |
| CN101695150A (zh) * | 2009-10-12 | 2010-04-14 | 清华大学 | 多声道音频编码方法、编码器、解码方法和解码器 |
| CN104364842A (zh) * | 2012-04-18 | 2015-02-18 | 诺基亚公司 | 立体声音频信号编码器 |
| CN109416912A (zh) * | 2016-06-30 | 2019-03-01 | 杜塞尔多夫华为技术有限公司 | 一种对多声道音频信号进行编码和解码的装置和方法 |
Non-Patent Citations (1)
| Title |
|---|
| See also references of EP4174855A4 |
Also Published As
| Publication number | Publication date |
|---|---|
| US20230154471A1 (en) | 2023-05-18 |
| CN113948095B (zh) | 2025-02-25 |
| KR20230036146A (ko) | 2023-03-14 |
| EP4174855A4 (en) | 2023-12-06 |
| JP2023533366A (ja) | 2023-08-02 |
| EP4174855A1 (en) | 2023-05-03 |
| CN113948095A (zh) | 2022-01-18 |
| JP7519531B2 (ja) | 2024-07-19 |
| US12437767B2 (en) | 2025-10-07 |
| KR102948552B1 (ko) | 2026-04-03 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| WO2022012553A1 (zh) | 多声道音频信号的编解码方法和装置 | |
| CN113593586B (zh) | 音频信号编码方法、解码方法、编码设备以及解码设备 | |
| CN115691514B (zh) | 一种多声道信号的编解码方法和装置 | |
| WO2023005414A1 (zh) | 一种音频信号的编解码方法和装置 | |
| JP7636516B2 (ja) | マルチ・チャネル・オーディオ信号符号化方法及び装置 | |
| CN116913293A (zh) | 一种多声道音频的混合模式编码方法、装置、设备及介质 | |
| WO2022247651A1 (zh) | 多声道音频信号的编码方法和装置 | |
| WO2022012675A1 (zh) | 多声道音频信号的编码方法和装置 | |
| US12431144B2 (en) | Multi-channel audio signal encoding and decoding method and apparatus | |
| CN118782053A (zh) | 非负矩阵分解的蓝牙接收端单声道上混方法、装置、介质 | |
| CN120266204A (zh) | 参数空间音频编码 | |
| KR102869250B1 (ko) | 선형 예측 코딩 파라미터 코딩 방법 및 코딩 장치 | |
| RU2811412C1 (ru) | СПОСОБ КОДИРОВАНИЯ ПАРАМЕТРОВ КОДИРОВАНИЯ С ЛИНЕЙНЫМ ПРОГНОЗИРОВАНИЕМ и УСТРОЙСТВО КОДИРОВАНИЯ | |
| WO2023173941A1 (zh) | 一种多声道信号的编解码方法和编解码设备以及终端设备 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 21843116 Country of ref document: EP Kind code of ref document: A1 |
|
| ENP | Entry into the national phase |
Ref document number: 2023502888 Country of ref document: JP Kind code of ref document: A |
|
| WWE | Wipo information: entry into national phase |
Ref document number: 202317004150 Country of ref document: IN |
|
| ENP | Entry into the national phase |
Ref document number: 20237004819 Country of ref document: KR Kind code of ref document: A |
|
| ENP | Entry into the national phase |
Ref document number: 2021843116 Country of ref document: EP Effective date: 20230124 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| WWG | Wipo information: grant in national office |
Ref document number: 202317004150 Country of ref document: IN |























