WO2022012553A1 - 多声道音频信号的编解码方法和装置 - Google Patents

多声道音频信号的编解码方法和装置 Download PDF

Info

Publication number
WO2022012553A1
WO2022012553A1 PCT/CN2021/106101 CN2021106101W WO2022012553A1 WO 2022012553 A1 WO2022012553 A1 WO 2022012553A1 CN 2021106101 W CN2021106101 W CN 2021106101W WO 2022012553 A1 WO2022012553 A1 WO 2022012553A1
Authority
WO
WIPO (PCT)
Prior art keywords
channel
pair
audio frame
correlation value
channel pair
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/CN2021/106101
Other languages
English (en)
French (fr)
Inventor
王智
丁建策
夏丙寅
王宾
王喆
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Huawei Technologies Co Ltd
Original Assignee
Huawei Technologies Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Huawei Technologies Co Ltd filed Critical Huawei Technologies Co Ltd
Priority to EP21843116.1A priority Critical patent/EP4174855A4/en
Priority to JP2023502888A priority patent/JP7519531B2/ja
Priority to KR1020237004819A priority patent/KR102948552B1/ko
Publication of WO2022012553A1 publication Critical patent/WO2022012553A1/zh
Priority to US18/153,128 priority patent/US12437767B2/en
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Images

Classifications

    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L19/00Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis
    • G10L19/008Multichannel audio signal coding or decoding using interchannel correlation to reduce redundancy, e.g. joint-stereo, intensity-coding or matrixing
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L25/00Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00
    • G10L25/03Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 characterised by the type of extracted parameters
    • G10L25/06Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 characterised by the type of extracted parameters the extracted parameters being correlation coefficients

Definitions

  • the present application relates to audio processing technologies, and in particular, to a method and device for encoding and decoding multi-channel audio signals.
  • Encoding and decoding of multi-channel audio is a technique for encoding or decoding audio that contains more than two channels.
  • Common multi-channel audios include 5.1-channel audio, 7.1-channel audio, 7.1.4-channel audio, and 22.2-channel audio.
  • MPS MPEG Surround
  • the present application provides a method and device for encoding and decoding multi-channel audio signals, so as to reduce redundancy between channel signals and improve audio encoding efficiency.
  • the present application provides a method for encoding a multi-channel audio signal, including: acquiring a first audio frame to be encoded, where the first audio frame includes at least five channel signals; acquiring a correlation value set, the The correlation value set includes respective correlation values of a plurality of channel pairs, one channel pair includes two channel signals in the at least five channel signals, and the correlation value of the one channel pair is used to represent the one channel pair.
  • M is a positive integer less than or equal to the set value
  • M is a positive integer less than or equal to the set value
  • the first audio frame in this embodiment may be any frame in the multi-channel audio signal to be encoded, and the first audio frame includes five or more channel signals. Coding two channel signals with higher correlation together can reduce redundancy and improve coding efficiency. Therefore, in this embodiment, the pairing is determined according to the correlation value between the two channel signals. In order to find the channel pair set with the highest correlation as much as possible, the correlation value between at least five channel signals in the first audio frame can be calculated to obtain the correlation value set of the first audio frame. For example, five channel signals may form 10 channel pairs in total, and correspondingly, the correlation value set may include 10 correlation values.
  • all the correlation values included in the correlation value set can be sorted in descending order, and the top M correlation values in the front row are selected from them, and the M correlation values must be greater than or equal to the group pair threshold, This is because a correlation value smaller than the group pair threshold indicates that the correlation between the two channel signals in the corresponding channel pair is low, and there is no need for group pair coding.
  • an upper limit N of M is set, that is, at most N correlation values can be selected.
  • the target channel pair set contains The sum of the correlation values of all channel pairs is the largest, and the number of channel pairs in a group pair is increased as much as possible to reduce the redundancy between channel signals and improve the audio coding efficiency.
  • the M channel pair sets include a first channel pair set, and the acquiring the M channel pair sets acquires the first channel pair set; the acquiring the first channel pair set A set of channel pairs, comprising: adding a first channel pair in the M channel pairs to the first channel pair set, where the first channel pair is the first channel pair in the M channel pairs Any one; when the other channel pairs other than the associated channel in the multiple channel pairs include channel pairs whose correlation value is greater than the set pair threshold, select the maximum correlation value from the other channel pairs One channel pair of the channel pair is added to the first channel pair set, and the associated channel pair includes any one of the channel signals included in the channel pair that has been added to the first channel pair set.
  • the multiple channel pairs with larger correlation values are used as the first channel pair added to the channel pair set, and then the channel pair corresponding to the largest correlation value in the remaining channel pairs is selected.
  • Add the corresponding channel pair set obtain the sum of the correlation values of multiple channel pair sets as much as possible, and then determine the channel pair set corresponding to the maximum correlation value sum as the target channel pair set, so that the target sound can be achieved.
  • the sum of the correlation values of all channel pairs included in the channel pair set is the largest, and the number of channel pairs in a group pair is increased as much as possible, reducing redundancy between channel signals and improving audio coding efficiency.
  • the selecting M correlation values from the correlation value set includes: selecting N correlation values from the correlation value set, where the N correlation values are all greater than the correlation values other correlation values in the value set except the N correlation values, where N is the set value; select correlation values greater than or equal to the pair threshold from the N correlation values, and the correlation values greater than or equal to The number of correlation values of the group to the threshold is M.
  • the M correlation values are greater than or equal to the pair threshold, and M is a positive integer less than or equal to a set value (eg, N).
  • a set value eg, N
  • all the correlation values included in the correlation value set may be sorted in descending order, and the top N correlation values in the front row may be selected from the correlation values, and the N correlation values may have correlation values smaller than the group pair threshold. Therefore, M correlation values greater than or equal to the group pair threshold are selected from the N correlation values, because the correlation values smaller than the group pair threshold represent the correlation between the two channel signals in the corresponding channel pair. Low sex, no group pair coding necessary.
  • the correlation value is a normalized value.
  • the normalization process can incorporate the related values with large differences in the value range into a unified range for comparison and processing, thereby improving the operation efficiency.
  • the correlation value of the one channel pair is set to 0.
  • a small correlation value indicates that the correlation between the corresponding two channel signals is small, and there is no need for group pairs. Therefore, the correlation value of the two channel signals in this case is set to 0, which is convenient for subsequent calculations and improves. Operational efficiency.
  • the present application provides a method for encoding a multi-channel audio signal, including: acquiring a first audio frame to be encoded, where the first audio frame includes at least five channel signals; acquiring a correlation value set, the The correlation value set includes respective correlation values of a plurality of channel pairs, one channel pair includes two channel signals in the at least five channel signals, and the correlation value of the one channel pair is used to represent the one channel pair.
  • Correlation between two channel signals of a channel pair multiple channel pair sets are obtained according to the multiple channel pairs, when the channel pair set includes more than two channel pairs, the two More than one channel pair does not contain the same channel signal; obtain the sum of the correlation values of all channel pairs included in each channel pair set in the multiple channel pair sets according to the correlation value set; determine the target sound A channel pair set, the sum of the correlation values of all channel pairs in the target channel pair set is the largest among the multiple channel pair sets; according to the target channel pair set, the first audio frame to encode.
  • the obtaining multiple channel pair sets according to the multiple channel pairs includes: obtaining the obtained channel pairs according to other channel pairs in the multiple channel pairs other than unrelated channels The plurality of channel pairs are set, and the correlation value of the uncorrelated channel pairs is less than the group pair threshold.
  • a small correlation value indicates that the correlation between the corresponding two channel signals is small, and there is no need for group pairs. Therefore, the correlation value of the two channel signals in this case and the sound of the two channel signals
  • the deletion of the track pair can reduce the amount of subsequent calculations and improve the operation efficiency.
  • the correlation value is a normalized value.
  • the normalization process can incorporate the related values with large differences in the value range into a unified range for comparison and processing, thereby improving the operation efficiency.
  • the correlation value of the one channel pair is set to 0.
  • a small correlation value indicates that the correlation between the corresponding two channel signals is small, and there is no need for group pairs. Therefore, the correlation value of the two channel signals in this case is set to 0, which is convenient for subsequent calculations and improves. Operational efficiency.
  • the present application provides a method for encoding a multi-channel audio signal, comprising: acquiring a first audio frame to be encoded, where the first audio frame includes at least five channel signals; acquiring the first audio frame
  • the correlation value set of the first audio frame includes the respective correlation values of multiple channel pairs, one channel pair includes two channel signals in the at least five channel signals, and the one channel pair includes two channel signals in the at least five channel signals.
  • the correlation value of the channel pair is used to represent the correlation between the two channel signals of the one channel pair; the correlation value set of the second audio frame is obtained, and the correlation value set of the second audio frame includes the The respective correlation values of a plurality of channel pairs of the second audio frame, one channel pair includes two channel signals among the at least five channel signals of the second audio frame, and the correlation value of the one channel pair Used to represent the correlation between the two channel signals of the one channel pair, the second audio frame is the previous frame of the first audio frame; according to the set of correlation values of the first audio frame and the correlation value set of the second audio frame to judge whether it is necessary to re-acquire the target channel pair set of the first audio frame; if it is necessary to re-acquire the target channel pair set of the first audio frame, the above
  • the method according to any one of the first to second aspects acquires a target channel pair set of the first audio frame, and encodes the first audio frame according to the target channel pair set; Acquiring the target channel pair set of the first audio frame
  • the sum of the difference between the correlation value set of the current audio frame and the correlation value set of the previous audio frame it is determined whether it is necessary to re-acquire the target channel pair set of the current frame. Reduce the amount of calculation and improve the coding efficiency. Even if the audio changes greatly and the target channel pair set needs to be re-acquired, the sum of the correlation values of multiple channel pair sets can still be obtained as much as possible, and then the sum of the maximum correlation value can be obtained.
  • the corresponding channel pair set is determined as the target channel pair set, which can maximize the sum of the correlation values of all channel pairs included in the target channel pair set, and increase the number of channel pairs in the group pair as much as possible, reducing the The redundancy between channel signals improves the coding efficiency of audio.
  • determining whether the target channel pair set of the first audio frame needs to be re-acquired according to the correlation value set of the first audio frame and the correlation value set of the second audio frame including: calculating the absolute value of the difference between the correlation value set of the first audio frame and the correlation value set of the second audio frame corresponding to the same channel pair; The corresponding sum of the absolute values; when the sum of the absolute values is less than the change threshold, it is determined that it is not necessary to re-acquire the target channel pair set of the first audio frame; when the sum of the absolute values is greater than or equal to the When the threshold value is changed, it is determined that the target channel pair set of the first audio frame needs to be re-acquired.
  • the change threshold can be, for example, ⁇ the number of channel pairs, where the value of ⁇ can be 0.14 or 0.15, and the number of channel pairs refers to the set of correlation values of the first audio frame (or the correlation value of the second audio frame). The number of channel pairs included in the value set).
  • the present application provides a method for encoding a multi-channel audio signal, comprising: acquiring a first audio frame to be encoded, where the first audio frame includes K channel signals, where K is an integer greater than or equal to 5
  • K is an integer greater than or equal to 5
  • K is an integer greater than or equal to 5
  • K is an integer greater than or equal to 5
  • K is an integer greater than or equal to 5
  • K is an integer greater than or equal to 5
  • K is an integer greater than or equal to 5
  • K is an integer greater than or equal to 5
  • K is an integer greater than or equal to 5
  • K is an integer greater than or equal to 5
  • K is an integer greater than or equal to 5
  • K is an integer greater than or equal to 5
  • K is an integer greater than or equal to 5
  • K is an integer greater than or equal to 5
  • K is an integer greater than or equal to 5
  • K is an integer greater than or equal to 5
  • K is an integer greater than or equal to 5
  • K is an integer greater
  • the difference from the method of the first aspect or the second aspect is that the method of the first aspect and the second aspect is fused, that is, according to the number of channel signals contained in the first audio frame, it is determined which audio frame to use for the first audio frame.
  • a method gets its target channel pair collection.
  • the method of the second aspect is adopted, it is necessary to exhaustively enumerate all target channel pairs, which will increase the amount of calculation. Therefore, the method of the first aspect will reduce the amount of calculation. A lot of computation.
  • the method of the second aspect can be used to obtain the sum of the correlation values of all channel pair sets, so as to ensure that the final selected target channel pair set must be the highest Optimal results that match the characteristics of the first audio frame.
  • the present application provides an encoding device, comprising: an acquisition module configured to acquire a first audio frame to be encoded, where the first audio frame includes at least five channel signals; acquire a correlation value set, the correlation The value set includes respective correlation values of a plurality of channel pairs, one channel pair includes two channel signals in the at least five channel signals, and the correlation value of the one channel pair is used to represent the one channel pair.
  • the M correlation values are all greater than or equal to the group pair threshold, and M is a positive integer less than or equal to the set value; acquire M channel pair sets, each of which includes at least the One of the M channel pairs corresponding to the M correlation values, and when the channel pair set includes more than two channel pairs, the two or more channel pairs do not contain the same channel signal; determine A module for determining a target channel pair set from the M channel pair sets, where the sum of the correlation values of all channel pairs in the target channel pair set is the largest in the M channel pair sets an encoding module, configured to encode the first audio frame according to the target channel pair set.
  • the M channel pair sets include a first channel pair set; and the acquiring module is specifically configured to add the first channel pair in the M channel pairs to all channel pairs.
  • the first channel pair set, the first channel pair is any one of the M channel pairs;
  • the obtaining module is specifically configured to select N correlation values from the correlation value set, and the N correlation values are all greater than the N correlation values in the correlation value set divided by the N correlation values.
  • Other correlation values other than the value, N is the set value; from the N correlation values, select the correlation value greater than or equal to the threshold value of the group pair, the correlation value greater than or equal to the threshold value of the group pair The number is M.
  • the correlation value is a normalized value.
  • the correlation value of the one channel pair is set to 0.
  • the present application provides an encoding device, comprising: an acquisition module for acquiring a first audio frame to be encoded, the first audio frame including at least five channel signals; acquiring a correlation value set, the correlation The value set includes respective correlation values of a plurality of channel pairs, one channel pair includes two channel signals in the at least five channel signals, and the correlation value of the one channel pair is used to represent the one channel pair.
  • Correlation between two channel signals of a channel pair multiple channel pair sets are obtained according to the multiple channel pairs, and when the channel pair set includes more than two channel pairs, the two The above channel pairs do not contain the same channel signal; obtain the sum of the correlation values of all channel pairs included in each channel pair set in the multiple channel pair sets according to the correlation value set; determine the module, use In determining a target channel pair set, the sum of the correlation values of all channel pairs in the target channel pair set is the largest in the multiple channel pair sets; an encoding module is used for according to the target channel pair set.
  • the set encodes the first audio frame.
  • the obtaining module is specifically configured to obtain the set of multiple channel pairs according to other channel pairs in the multiple channel pairs other than the non-correlated channels to the outside, and the non-correlated channels The correlation value of the channel pair is less than the group pair threshold.
  • the correlation value is a normalized value.
  • the correlation value of the one channel pair is set to 0.
  • the present application provides an encoding device, comprising: an acquisition module configured to acquire a first audio frame to be encoded, the first audio frame including at least five channel signals; A set of correlation values, the set of correlation values of the first audio frame includes respective correlation values of a plurality of channel pairs, one channel pair includes two channel signals in the at least five channel signals, and the one channel pair includes two channel signals in the at least five channel signals.
  • the correlation value of the channel pair is used to represent the correlation between the two channel signals of the one channel pair; the correlation value set of the second audio frame is obtained, and the correlation value set of the second audio frame includes the first audio frame.
  • one channel pair includes two channel signals out of at least five channel signals of the second audio frame, and the correlation value of the one channel pair is defined by
  • the second audio frame is the previous frame of the first audio frame; the encoding module is used for according to the first audio frame.
  • the correlation value set of the second audio frame and the correlation value set of the second audio frame judge whether it is necessary to re-acquire the target channel pair set of the first audio frame; if it is necessary to re-acquire the target channel pair set of the first audio frame, Then execute the method according to any one of claims 1-9 to obtain the target channel pair set of the first audio frame, and encode the first audio frame according to the target channel pair set; if It is not necessary to re-acquire the target channel pair set of the first audio frame, then determine the target channel pair set of the second audio frame as the target channel pair set of the first audio frame, and determine the target channel pair set of the second audio frame as the target channel pair set of the first audio frame.
  • the set of target channel pairs encodes the first audio frame.
  • the encoding module is specifically configured to calculate the difference between the correlation value set of the first audio frame and the correlation value set corresponding to the same channel pair in the correlation value set of the second audio frame The absolute value of the difference; calculate the sum of the absolute values corresponding to the plurality of the channel pairs respectively; when the sum of the absolute values is less than the change threshold, it is determined that the target channel of the first audio frame does not need to be re-acquired Pair set; when the sum of the absolute values is greater than or equal to the change threshold, it is determined that the target channel pair set of the first audio frame needs to be re-acquired.
  • the present application provides an encoding device, comprising: an acquisition module configured to acquire a first audio frame to be encoded, where the first audio frame includes K channel signals, and K is an integer greater than or equal to 5; an encoding module, configured to perform the method according to any one of the above-mentioned first aspects to encode the first audio frame when K is greater than the threshold of the number of channel signals; when K is less than or equal to the threshold of the number of channel signals , performing the method according to any one of the above second aspects to encode the first audio frame.
  • the present application provides a device, comprising: one or more processors; a memory for storing one or more programs; when the one or more programs are executed by the one or more processors, The one or more processors are caused to implement the method of any one of the first to fourth aspects above.
  • the present application provides a computer-readable storage medium, comprising a computer program, which, when executed on a computer, causes the computer to execute the method according to any one of the first to fourth aspects.
  • the present application provides a computer-readable storage medium, characterized by comprising an encoded code stream obtained according to the encoding method for a multi-channel audio signal according to any one of the first to fourth aspects above.
  • FIG. 1 exemplarily presents a schematic block diagram of an audio decoding system 10 applied in the present application
  • FIG. 2 exemplarily presents a schematic block diagram of an audio decoding device 200 to which the present application is applied;
  • FIG. 3 is a flowchart of an exemplary embodiment of a method for encoding a multi-channel audio signal provided by the present application
  • FIG. 4 is an exemplary structural diagram of an encoding device to which the multi-channel audio signal encoding method provided by the present application is applied;
  • FIG. 5 is a flowchart of an exemplary embodiment of a method for encoding a multi-channel audio signal provided by the present application
  • FIG. 6 is a flowchart of an exemplary embodiment of a method for encoding a multi-channel audio signal provided by the present application
  • FIG. 7 is a flowchart of an exemplary embodiment of a method for encoding a multi-channel audio signal provided by the present application
  • FIG. 8 is an exemplary structural diagram of a decoding device to which the decoding method for a multi-channel audio signal provided by the present application is applied;
  • FIG. 9 is a schematic structural diagram of an embodiment of an encoding device of the present application.
  • FIG. 10 is a schematic structural diagram of an embodiment of a device of the present application.
  • At least one (item) refers to one or more, and "a plurality” refers to two or more.
  • “And/or” is used to describe the relationship between related objects, indicating that there can be three kinds of relationships, for example, “A and/or B” can mean: only A, only B, and both A and B exist , where A and B can be singular or plural.
  • the character “/” generally indicates that the associated objects are an “or” relationship.
  • At least one item(s) below” or similar expressions thereof refer to any combination of these items, including any combination of single item(s) or plural items(s).
  • At least one (a) of a, b or c can mean: a, b, c, "a and b", “a and c", “b and c", or "a and b and c" ", where a, b, c can be single or multiple.
  • Audio frame Audio data is streaming.
  • the amount of audio data within a period of time is usually taken as a frame of audio. This period is called “sampling time", which can be determined according to the codec. Determine its value according to the requirements of the device and specific applications, for example, the duration is 2.5ms to 60ms, and ms is milliseconds.
  • Audio signal is the information carrier of frequency and amplitude variation of regular sound waves with speech, music and sound effects. Audio is a continuously changing analog signal that can be represented by a continuous curve called a sound wave. Audio is a digital signal generated by analog-to-digital conversion or by a computer. Sound waves have three important parameters: frequency, amplitude and phase, which determine the characteristics of the audio signal.
  • Channel signal refers to the independent audio signals that are collected or played back at different spatial positions during recording or playback. Therefore, the number of channels is the number of sound sources during sound recording or the number of speakers during playback.
  • FIG. 1 exemplarily shows a schematic block diagram of an audio decoding system 10 applied in the present application.
  • the audio coding system 10 may include a source device 12 and a destination device 14.
  • the source device 12 generates an encoded code stream, and thus, the source device 12 may be referred to as an audio encoding device.
  • the destination device 14 may decode the encoded codestream generated by the source device 12, and thus, the destination device 14 may be referred to as an audio decoding device.
  • the source device 12 includes an encoder 20 and, optionally, an audio source 16 , an audio preprocessor 18 , and a communication interface 22 .
  • Audio source 16 may include or be any type of audio capture device for capturing real world speech, music, sound effects, etc., and/or any type of audio generation device, such as an audio processor for generating speech, music, and sound effects or equipment.
  • the audio source may be any type of memory or storage that stores the above audio.
  • the audio preprocessor 18 is used to receive (raw) audio data 17 and to preprocess the audio data 17 to obtain preprocessed audio data 19 .
  • the preprocessing performed by the audio preprocessor 18 may include trimming or denoising. It is understood that the audio preprocessing unit 18 may be an optional component.
  • An encoder 20 is used to receive preprocessed audio data 19 and provide encoded audio data 21 .
  • a communication interface 22 in source device 12 may be used to receive encoded audio data 21 and send encoded audio data 21 over communication channel 13 to destination device 14 for storage or direct reconstruction.
  • the destination device 14 includes a decoder 30 and, optionally, a communication interface 28 , an audio post-processor 32 and a playback device 34 .
  • the communication interface 28 in the destination device 14 is used to receive the encoded audio data 21 directly from the source device 12 and to provide the encoded audio data 21 to the decoder 30 .
  • Communication interface 22 and communication interface 28 may be used through a direct communication link between source device 12 and destination device 14, such as a direct wired or wireless connection, etc., or through any type of network, such as a wired network, a wireless network, or any A combination, any type of private network and public network, or any type of combination, transmits or receives encoded audio data 21 .
  • the communication interface 22 may be used to encapsulate the encoded audio data 21 into a suitable format such as a message, and/or to process the encoded audio data 21 using any type of transfer encoding or processing for transmission over a communication link or communication network .
  • the communication interface 28 corresponds to the communication interface 22 and may be used, for example, to receive transmission data and process the transmission data to obtain encoded audio data 21 using any type of corresponding transmission decoding or processing and/or decapsulation.
  • Both the communication interface 22 and the communication interface 28 can be configured as a one-way communication interface as indicated by the arrow in FIG. 1 from the corresponding communication channel 13 of the source device 12 to the destination device 14, or a two-way communication interface, and can be used to send and receive messages etc. to establish a connection, acknowledge and exchange any other information related to a communication link and/or data transfer such as encoded audio data, etc.
  • a decoder 30 is used to receive encoded audio data 21 and provide decoded audio data 31 .
  • the audio post-processor 32 is used for post-processing the decoded audio data 31 to obtain post-processed post-processed audio data 33 .
  • the post-processing performed by the audio post-processor 32 may include, for example, trimming or resampling, and the like.
  • Playback device 34 is used to receive post-processed audio data 33 to play audio to a user or listener.
  • Playback device 34 may be or include any type of player for playing reconstructed audio, eg, integrated or external speakers.
  • speakers may include speakers, speakers, and the like.
  • FIG. 2 exemplarily shows a schematic block diagram of an audio decoding device 200 applied in the present application.
  • the audio coding apparatus 200 may be an audio decoder (eg, decoder 30 of FIG. 1 ) or an audio encoder (eg, encoder 20 of FIG. 1 ).
  • the audio decoding device 200 includes: an input port 210 and a receiving unit (Rx) 220 for receiving data, a processor, logic unit or central processing unit 230 for processing data, and a transmitting unit (Tx) 240 for transmitting data and egress port 250, and a memory 260 for storing data.
  • the audio decoding device 200 may also include photoelectric conversion components and electro-optical (EO) components coupled with the input port 210, the receiving unit 220, the transmitting unit 240, and the output port 250 for the exit or entrance of optical or electrical signals.
  • EO electro-optical
  • the processor 230 is implemented by hardware and software.
  • the processor 230 may be implemented as one or more CPU chips, cores (eg, multi-core processors), FPGAs, ASICs, and DSPs.
  • the processor 230 communicates with the ingress port 210 , the receiving unit 220 , the transmitting unit 240 , the egress port 250 and the memory 260 .
  • the processor 230 includes a decoding module 270 (eg, an encoding module or a decoding module).
  • the decoding module 270 implements the embodiments disclosed in this application, so as to implement the encoding and decoding methods for multi-channel audio signals provided in this application.
  • the transcoding module 270 implements, processes, or provides various encoding operations.
  • decoding module 270 is implemented as instructions stored in memory 260 and executed by processor 230 .
  • Memory 260 includes one or more magnetic disks, tape drives, and solid-state drives, and may serve as an overflow data storage device for storing programs as they are selectively executed, and for storing instructions and data read during program execution.
  • Memory 260 may be volatile and/or non-volatile, and may be read only memory (ROM), random access memory (RAM), random access memory (ternary content-addressable memory, TCAM) and/or static Random Access Memory (SRAM).
  • ROM read only memory
  • RAM random access memory
  • TCAM ternary content-addressable memory
  • SRAM static Random Access Memory
  • the present application provides a method for encoding and decoding a multi-channel audio signal.
  • FIG. 3 is a flowchart of an exemplary embodiment of a method for encoding a multi-channel audio signal provided by the present application.
  • the process 300 may be performed by the source device 12 or the audio coding device 200 in the audio coding system 10 .
  • Process 300 is described as a series of steps or operations, and it should be understood that process 300 may be performed in various orders and/or concurrently, and is not limited to the order of execution shown in FIG. 3 .
  • the method includes:
  • Step 301 Obtain a first audio frame to be encoded.
  • the first audio frame in this embodiment may be any frame in the multi-channel audio signal to be encoded, and the first audio frame includes five or more channel signals.
  • a 5.1 channel includes a center channel (C), a front left channel (left, L), a front right channel (right, R), a left surround channel (LS), a back The right surround channel (right surround, RS) and the 0.1 channel low frequency effects (low frequency effects, LFE) a total of six channel signals.
  • the 7.1 channel includes C, L, R, LS, RS, LB, RB and LFE a total of eight channel signals, where LFE is the audio channel from 3-120Hz, which is usually sent to a channel specially designed for low tones. designed speakers.
  • Step 302 Obtain a set of related values.
  • the correlation value set includes the respective correlation values of a plurality of channel pairs, wherein one channel pair includes two channel signals in at least five channel signals, and the correlation value of one channel pair is used to represent the two channels of the channel pair. correlation between channel signals.
  • the multiple channel pairs may include all channel pairs corresponding to at least five channel signals, or, the multiple channel pairs may also include some channel pairs corresponding to at least five channel signals. Make specific restrictions.
  • the pairing is determined according to the correlation value between the two channel signals.
  • the correlation value set of the first audio frame may be obtained by first calculating the correlation value between at least five channel signals in the first audio frame. For example, five channel signals may form 10 channel pairs in total, and correspondingly, the correlation value set may include 10 correlation values.
  • the correlation values can be normalized, so that the correlation values of all channel pairs are limited to a specific range, so as to set a unified judgment standard for the correlation values, such as a group pair threshold, the group pair threshold can be set. It is a value greater than or equal to 0.2 and less than or equal to 1, such as 0.3, 0.4, or 0.35, etc., so that as long as the normalized correlation value of the two channel signals is less than the group pair threshold, the two channels are considered to be The correlation of the channel signals is poor, and group pair coding is not required.
  • the correlation value between two channel signals can be calculated using the following formula:
  • corr_norm(ch1, ch2) represents the normalized correlation value between the channel signal ch1 and the channel signal ch2
  • spec_ch1(i) represents the frequency domain coefficient of the ith frequency point of the channel signal ch1
  • spec_ch2(i ) is the frequency domain coefficient of the ith frequency point of the channel signal ch2
  • N represents the total number of frequency points of an audio frame.
  • the correlation value calculated by using the above algorithm or formula can be used as the initial correlation value, and then it is determined whether the initial correlation value needs to be modified according to a preset condition.
  • the restriction condition may include: calculating whether the amplitude ratio between the two channel signals of the initial correlation value is greater than a preset group pair threshold. When the amplitude ratio is greater than the pair threshold, the initial correlation value is modified; when the amplitude ratio is less than or equal to the pair threshold, the initial correlation value is kept unchanged.
  • the modification may be to reduce the initial correlation value, for example, the initial correlation value may be directly modified to 0, so as to prevent the two channel signals from being processed in groups.
  • the amplitude level(ch) of the current frame of the channel signal ch can be calculated by using the following formula:
  • i represents the ith sample of the current frame of the channel signal ch
  • N represents the total number of samples of the current frame
  • sepc_coeff(ch, i) is the frequency domain coefficient of the ith sample of the current frame.
  • Step 303 Select M correlation values from the correlation value set.
  • the M correlation values are all greater than other correlation values in the correlation value set except the M correlation values, the M correlation values are all greater than or equal to the group pair threshold, and M is a positive value less than or equal to a set value (for example, N). Integer.
  • all the correlation values included in the correlation value set can be sorted in descending order, and the top M correlation values in the front row are selected from them, and the M correlation values must be greater than or equal to the group pair threshold, This is because a correlation value smaller than the group pair threshold indicates that the correlation between the two channel signals in the corresponding channel pair is low, and there is no need for group pair coding.
  • an upper limit N of M is set, that is, at most N correlation values can be selected.
  • N can be an integer greater than or equal to 2, and the maximum value of N cannot exceed the number of all channel pairs corresponding to all channel signals of the first audio frame.
  • the larger the value of N the larger the amount of computation involved, while the smaller the value of N is, the channel pair set may be lost, thereby reducing the coding efficiency.
  • the subsequent steps do not need to be performed, and it is sufficient to perform mono encoding on each channel signal of the first audio frame. If M correlation values are selected from the set of correlation values, the following steps may be performed.
  • Step 304 Acquire M channel pair sets.
  • Each channel pair set includes at least one of the M channel pairs corresponding to the M correlation values, and when the channel pair set includes more than two channel pairs, the two or more channel pairs do not contain the same sound channel signal.
  • the 3 channel pairs corresponding to the largest correlation value selected according to the correlation value set are (L, R), (R, C) and (LS, RS), where the correlation of (LS, RS) The value is less than the group pair threshold, so it is excluded, then the remaining two channel pairs (L, R) and (R, C) can obtain two channel pair sets, one of which includes (L ,R) and the other includes (R,C).
  • the method for acquiring the M channel pair sets in this embodiment may include: adding the first channel pair to the first channel pair.
  • a channel pair set, the M channel pair sets include the first channel pair set, when the other channel pairs except the associated channel in the multiple channel pairs include channels whose correlation value is greater than the group pair threshold
  • a channel pair whose correlation value is greater than the group pair threshold is included, select a channel pair with the largest correlation value from other channel pairs and add it to the first channel pair set.
  • step b may be iteratively executed.
  • the correlation value smaller than the pair pair threshold may be deleted from the correlation value set, so that the number of channel pairs can be reduced, thereby reducing the number of iterations.
  • Step 305 Determine a target channel pair set from the M channel pair sets.
  • the sum of the correlation values of all channel pairs in the target channel pair set is the largest among the M channel pair sets. After the above-mentioned M channel pair sets are obtained, the sum of the correlation values of all channel pairs included in each channel pair set can be calculated, and finally the channel pair set with the largest sum of the correlation values is determined as the target channel pair. gather.
  • Step 306 Encode the first audio frame according to the target channel pair set.
  • At least five audio channels in the first audio frame may be acquired.
  • the channel signals are separately processed for energy equalization to obtain at least five equalized channel signals, and then stereo processing is performed on the at least five equalized channel signals.
  • the encoded object is related to the equalized channel signals.
  • the energy equalization mode may include a first energy equalization mode and/or a second energy equalization mode, wherein the first energy equalization mode only uses two channel signals in one channel pair to obtain two equalized channels corresponding to one channel pair Signal.
  • the second energy equalization mode uses two channel signals in one channel pair and at least one channel signal outside one channel to obtain two equalized channel signals corresponding to one channel pair.
  • the average value of the energy or amplitude values of the two channel signals included in the current channel pair may be calculated, and according to the average value Perform energy equalization processing on the two channel signals respectively to obtain two corresponding equalized channel signals.
  • the fluctuation interval value of at least five channel signals is large, energy balance can be performed only between the two related channel signals, so that the allocation of bits during stereo processing is more in line with the energy characteristics of channel signals, avoiding In a low bit rate coding environment, the coding noise of the channel pair with high energy may be much larger than the coding noise of the channel pair with low energy due to insufficient bits, and the bits of the channel pair with low energy may be redundant.
  • the energy equalization mode is the second energy equalization mode
  • the average value of the energy or amplitude values of the at least five channel signals can be calculated, and the energy equalization processing is performed on the at least five channel signals according to the average value to obtain at least five equalized sound signals. channel signal.
  • the target channel pair set contains The sum of the correlation values of all channel pairs is the largest, and the number of channel pairs in a group pair is increased as much as possible to reduce the redundancy between channel signals and improve the audio coding efficiency.
  • FIG. 4 is an exemplary structural diagram of an encoding apparatus to which the encoding method for a multi-channel audio signal provided by the present application is applied.
  • the encoding apparatus may be the encoder 20 of the source device 12 in the audio decoding system 10, or may be is the decoding module 270 in the audio decoding apparatus 200 .
  • the encoding device may include a channel pair set generation module, a multi-channel processing module, a channel encoding module and a code stream multiplexing interface, wherein,
  • the input of the channel pair set generation module is n channel signals (CH1-CHn) of multi-channel audio, where n is an integer greater than or equal to 5, and all the n channel signals can be processed in stereo.
  • the channel pair set generation module calculates the correlation value between any two channel signals in the n channel signals, so as to obtain the target channel pair set by using the method of the embodiment shown in FIG. 3 according to these correlation values, for example, (CH1 , CH2), (CH3, CH4), ..., (CHi-1, CHi).
  • the multi-channel processing module includes a plurality of stereo processing units, and the stereo processing unit can use prediction-based or Karhunen-Loeve Transform (Karhunen-Loeve Transform, KLT)-based processing, that is, the input two channel signals are rotated (for example, via 2 ⁇ 2 rotation matrix) to maximize energy compression, thereby concentrating the signal energy into one channel.
  • KLT Karhunen-Loeve Transform
  • Each channel pair in the target channel pair set output by the channel pair set generation module is respectively input to a stereo processing unit, for example, (CH1, CH2) input stereo processing unit 1, (CH3, CH4) input stereo processing unit 2 , ..., (CHi-1, CHi) are input to the stereo processing unit m.
  • the stereo processing unit processes the input two channel signals, it outputs the processed channel signal (P) corresponding to the two channel signals and the multi-channel parameter (SIDE_PAIR).
  • the multi-channel parameters include channel pair index, energy Equalization side information, stereo processing side information.
  • the stereo processing unit 1 processes CH1 and CH2 to obtain P1 and P2 and SIDE_PAIR1
  • the stereo processing unit 2 processes CH3 and CH4 to obtain P3 and P4, and SIDE_PAIR2, ...
  • the stereo processing unit m pairs CHi-1 and CHi Process, get Pi-1 and Pi, and SIDE_PAIRm.
  • the channel encoding module encodes the processed channel signal output by the multi-channel processing module by using a mono encoding unit (or a mono box or a mono tool) to output a corresponding encoded channel signal (E).
  • a mono encoding unit or a mono box or a mono tool
  • the channel signal with higher energy (or higher amplitude) is allocated more bits
  • the channel signal with less energy (or lower amplitude) is allocated more bits. Allocate fewer bits.
  • the channel encoding module may also use a stereo encoding unit, such as a parametric stereo encoder or a lossy stereo encoder, to encode the processed channel signal output by the multi-channel processing module.
  • P1, P2, P3, P4, ..., Pi1, Pi are respectively encoded by a mono coding unit to obtain E1, E2, E3, E4, ..., Ei1, Ei.
  • the unpaired channel signals (such as CHj) in the channel pair set generation module do not need to be processed by the stereo processing unit in the multi-channel processing module, and can be directly input into a single channel in the channel encoding module.
  • the channel coding unit obtains Ej.
  • the code stream multiplexing interface generates an encoded multi-channel signal, and the encoded multi-channel signal includes the encoded channel signal output by the channel encoding module and the multi-channel parameters output by the multi-channel processing module.
  • the encoded multi-channel signal includes E1, E2, E3, E4, ..., Ei1, Ei, and SIDE_PAIR1, SIDE_PAIR2, ..., SIDE_PAIRm.
  • the code stream multiplexing interface can process the encoded multi-channel signal into a serial signal or a serial bit stream.
  • the processing flow for obtaining the target channel pair set provided by the present application may be implemented by the channel pair set generation module in the encoding apparatus shown in FIG. 4 .
  • the 5.1 channel includes a center channel (C), a front left channel (left, L), a front right channel (right, R), and a left surround channel (left surround). , LS), right surround back channel (right surround, RS) and 0.1 channel low frequency effects (low frequency effects, LFE).
  • the channel pair set generation module can use a multi-channel mask to remove the channels that do not need multi-channel processing to improve the coding efficiency.
  • the LFE channel can be removed from the 5.1 channel, so the input sound
  • the channel signals of the channel pair set generation module include C, L, R, LS and RS.
  • the method for obtaining the target channel pair set may include the following steps:
  • the present application can use the following formula to calculate the correlation value between two channel signals (for example, the channel signal ch1 and the channel signal ch2):
  • corr_norm(ch1, ch2) represents the normalized correlation value between the channel signal ch1 and the channel signal ch2
  • spec_ch1(i) represents the frequency domain coefficient of the ith frequency point of the channel signal ch1
  • spec_ch2(i ) is the frequency domain coefficient of the ith frequency point of the channel signal ch2
  • N represents the total number of frequency points of an audio frame.
  • the obtained correlation value set may include at most Correlation values for each channel pair.
  • Table 1 shows an example of a set of correlation values for 5.1 channels.
  • the pairing threshold is set to 0.3, and only two channel signals with a correlation value greater than 0.3 can be paired. Therefore, delete the correlation values in Table 1 that are less than the pairing threshold to obtain Table 1a, which can be used in the iterative process. Channel signals with less correlation are not considered, thereby reducing the amount of computation.
  • R, C is the first channel pair added to the first channel pair set, and the correlation value of the channel pair including R and/or C is deleted from Table 1a to obtain Table 1b.
  • the maximum correlation value in Table 1b is 0.42 (LS, RS), so LS and RS are formed into a second channel pair and added to the first channel pair set. At this time, there is only one channel signal L left in the five channel signals, and the grouping cannot be continued. Therefore, the final first channel pairing set includes two channel pairs (R, C) and (LS, RS).
  • (L, C) is the first channel pair added to the second channel pair set, and the correlation value of the channel pair including L and/or C is deleted from Table 1a to obtain Table 1c.
  • the maximum correlation value in Table 1c is 0.42 (LS, RS), so LS and RS are formed into a second channel pair and added to the second channel pair set. At this time, there is only one channel signal R left in the five channel signals, and the grouping cannot be continued. Therefore, the final second channel pairing set includes two channel pairs (L, C) and (LS, RS).
  • (LS, RS) is the first channel pair added to the third channel pair set, and the correlation value of the channel pair including LS and/or RS is deleted from Table 1a to obtain Table 1d.
  • the maximum correlation value in Table 1d is 0.57(R, C), so R and C are combined into the second channel pair and added to the third channel pair set. At this time, there is only one channel signal L left in the five channel signals, and the grouping cannot be continued. Therefore, the final third channel pairing set includes two channel pairs (LS, RS) and (R, C).
  • the channel pair set corresponding to (1) (or S(3)) is used as the target channel pair set, that is, the channel pairs available for 5.1 channels in this embodiment include (L, C) and (LS, RS).
  • the target channel pair set can be represented by an index, and index values can be set for the channel pairs corresponding to all the correlation values in Table 1. After the target channel pair set is determined, the channel pairs in the target channel pair set can be used. The corresponding index value is indicated to save the number of bits in the code stream.
  • the 7.1 channel includes C, L, R, LS, RS, a left back channel (LB), a right back channel (RB) and an LFE.
  • the channel pair set generation module can use the multi-channel mask to remove the channels that do not need multi-channel processing to improve the coding efficiency.
  • the LFE channel can be removed from the 7.1 channel, so the input sound
  • the channel signals of the channel pair set generation module include C, L, R, LS, RS, LB and RB.
  • the method for obtaining the target channel pair set may include the following steps:
  • the formula of the above-mentioned first embodiment can also be used to calculate the correlation value between the two channel signals.
  • the obtained correlation value set may include at most Correlation values for each channel pair.
  • Table 2 shows an example of a set of correlation values for 7.1 channels.
  • the pairing threshold is set to 0.3, that is, only two channel signals with a correlation value greater than 0.3 can be paired. Therefore, by deleting the correlation value in Table 2 that is less than the pairing threshold, Table 2a can be obtained. In this way, in the process of iterative processing Channel signals with less correlation can be ignored, thereby reducing the amount of computation.
  • (LS, LB) is the first channel pair added to the first channel pair set, and the correlation value of the channel pair including LS and/or LB is deleted from Table 2a to obtain Table 2b.
  • the maximum correlation value in Table 2b is 0.57(R, C), so R and C are combined to form a second channel pair and added to the first channel pair set. Correlation values for channel pairs containing R and/or C are removed from Table 2b, resulting in Table 2c.
  • the final first channel pair set includes two channel pairs (LS,LB) and (R,C).
  • (RS, LB) is the first channel pair added to the second channel pair set, and the correlation value of the channel pair including RS and/or LB is deleted from Table 2a to obtain Table 2d.
  • the maximum correlation value in Table 2d is 0.57(R, C), so R and C are combined into a second channel pair and added to the second channel pair set. Correlation values for channel pairs containing R and/or C are removed from Table 2d, resulting in Table 2e.
  • Table 2e The maximum correlation value in Table 2e is 0.39(L, LS), so L and LS are combined into a third channel pair and added to the second channel pair set. Correlation values for channel pairs containing L and/or LS are removed from Table 2e, resulting in Table 2f.
  • the final first channel pair set includes three channel pairs (RS, LB), (R, C) and (L, LS).
  • R, C is the first channel pair added to the third channel pair set, and the correlation value of the channel pair including R and/or C is deleted from Table 2a to obtain Table 2g.
  • Table 2g The maximum correlation value in Table 2g is 0.67 (LS, LB), so LS and LB are formed into a second channel pair and added to the third channel pair set. Correlation values for channel pairs containing LS and/or LB are removed from Table 2g, resulting in Table 2h.
  • the final first channel pair set includes two channel pairs (R,C) and (LS,LB).
  • (L, C) is the first channel pair added to the fourth channel pair set, and the correlation value of the channel pair including L and/or C is deleted from Table 2a to obtain Table 2i.
  • Table 2i The maximum correlation value in Table 2i is 0.67(LS, LB), so LS and LB are formed into the second channel pair and added to the fourth channel pair set.
  • Table 2j is obtained by deleting the correlation values of channel pairs containing LS and/or LB from Table 2i.
  • the final first channel pair set includes two channel pairs (L,C) and (LS,LB).
  • S(2) is the largest among S(1), S(2), S(3) and S(4), so the channel pair set corresponding to S(2) is used as the target channel pair set, that is, this implementation
  • the channel pairs available for the 7.1 channel in the example include (RS, LB), (R, C) and (L, LS).
  • the second embodiment has one more iterative processing process, and the number of channel pairs included in the target channel pair set is also one more, which is related to the number of channel signals participating in the group pair.
  • FIG. 5 is a flowchart of an exemplary embodiment of a method for encoding a multi-channel audio signal provided by the present application.
  • the process 500 may be performed by the source device 12 or the audio coding device 200 in the audio coding system 10 .
  • Process 500 is described as a series of steps or operations, and it should be understood that process 500 may be performed in various orders and/or concurrently, and is not limited to the order of execution shown in FIG. 5 .
  • the method includes:
  • Step 501 Obtain a first audio frame to be encoded.
  • Step 502 Obtain a set of related values.
  • steps 501 and 502 in this embodiment reference may be made to the above-mentioned steps 301 and 302, and details are not repeated here.
  • Step 503 Acquire multiple channel pair sets according to the multiple channel pairs.
  • the correlation value set includes correlation values of multiple channel pairs of at least five channel signals of the first audio frame, and the multiple channel pairs are regularly combined (that is, multiple sound channels in the same channel pair set are combined. Channel pairs cannot contain the same channel signal), and multiple channel pair sets corresponding to the at least five channel signals can be obtained.
  • the following formula can be used to calculate the number of all channel pair sets:
  • Pair_num represents the number of all channel pair sets
  • CH represents the number of channel signals involved in multi-channel processing in the first audio frame, which is the result of filtering by the multi-channel mask.
  • multiple channel pair sets can be obtained according to other channel pairs in the multiple channel pairs that are not related to the outside, and the correlation value of the uncorrelated channel pair can be obtained. is smaller than the group pair threshold, in this way, the number of channel pairs participating in the calculation can be reduced when obtaining the channel pair set, thereby reducing the number of channel pair sets, and the calculation amount of the sum of correlation values can also be reduced in subsequent steps.
  • the channel signals whose correlation values with other channel signals are all smaller than the group pair threshold can be deleted, that is, such channel signals do not consider the group pair, and when the sound channel signal is obtained, the channel signal can be deleted.
  • the channel pairs are set, the number of the channel pairs involved in the calculation can be reduced, thereby reducing the number of the channel pair sets, and the calculation amount of the sum of the correlation values can also be reduced in the subsequent steps.
  • Step 504 Obtain the sum of the correlation values of all channel pairs included in each channel pair set in the multiple channel pair sets according to the correlation value set.
  • the sum of the correlation values of all channel pairs contained in the channel pair set is calculated.
  • Step 505 Determine the target channel pair set.
  • Step 506 Encode the first audio frame according to the target channel pair set.
  • steps 505 and 506 in this embodiment reference may be made to the above-mentioned steps 305 and 306, and details are not repeated here.
  • the 5.1 channel includes C, L, R, LS, RS, and LFE.
  • the channel pair set generation module can use the multi-channel mask to remove the channels that do not need multi-channel processing to improve the coding efficiency.
  • the LFE channel can be removed from the 5.1 channel, so the input sound
  • the channel signals of the channel pair set generation module include C, L, R, LS and RS.
  • the method for obtaining the target channel pair set may include the following steps:
  • the formula of the above-mentioned first embodiment can also be used to calculate the correlation value between the two channel signals.
  • the obtained correlation value set may include at most The correlation values of each channel pair are shown in Table 1.
  • 10 correlation values can be obtained from five channel signals, and correspondingly, 10 channel pairs can be obtained, and then the 10 channel pairs can be obtained at most A collection of channel pairs. For example, ⁇ (L,R),(LS,RS) ⁇ , ⁇ (L,R),(C,RS) ⁇ , ⁇ (L,R),(LS,C) ⁇ , .
  • the correlation value of the channel pair may be set to 0.
  • the channel pairs whose correlation value is less than the group pair threshold can be excluded, so that the number of channel pairs can be reduced when obtaining the channel pair set, thereby reducing the number of channel pairs.
  • the number of channel pair sets can be excluded, so that the number of channel pairs can be reduced when obtaining the channel pair set, thereby reducing the number of channel pairs.
  • FIG. 6 is a flowchart of an exemplary embodiment of a method for encoding a multi-channel audio signal provided by the present application.
  • the process 600 may be performed by the source device 12 or the audio coding device 200 in the audio coding system 10 .
  • Process 600 is described as a series of steps or operations, and it should be understood that process 600 may be performed in various orders and/or concurrently, and is not limited to the order of execution shown in FIG. 6 .
  • the method includes:
  • Step 601 Obtain a first audio frame to be encoded.
  • step 601 reference may be made to the above-mentioned step 301, which will not be repeated here.
  • Step 602 Obtain a correlation value set of the first audio frame.
  • the correlation value set of the first audio frame includes respective correlation values of a plurality of channel pairs, one channel pair includes two channel signals in at least five channel signals, and the correlation value of one channel pair is used to represent an audio channel. Correlation between the two channel signals of a channel pair.
  • Step 603 Obtain a correlation value set of the second audio frame.
  • the set of correlation values of the second audio frame includes respective correlation values of a plurality of channel pairs of the second audio frame, one channel pair includes two channel signals among the at least five channel signals of the second audio frame, one channel pair
  • the correlation value of the channel pair is used to represent the correlation between the two channel signals of one channel pair, and the second audio frame is the previous frame of the first audio frame.
  • this embodiment in addition to acquiring the correlation value set of the first audio frame, this embodiment also needs to acquire the correlation value set of the previous frame (ie, the second audio frame) of the first audio frame.
  • the encoding device Since the encoding of the second audio frame is prior to the encoding of the first audio frame, when the first audio frame is processed, the encoding device has already acquired relevant information when encoding the second audio frame, including the correlation value of the second audio frame Therefore, in this embodiment, the correlation value set of the second audio frame may be directly read from the cache or the memory, and the correlation value set of the second audio frame does not need to be calculated again.
  • Step 604 Determine whether the target channel pair set of the first audio frame needs to be re-acquired according to the correlation value set of the first audio frame and the correlation value set of the second audio frame.
  • the difference between the correlation value set of the first audio frame and the correlation value set of the second audio frame can be calculated as the judgment basis, that is, the correlation value set of the first audio frame and the correlation value of the second audio frame can be calculated.
  • the absolute value of the difference between the correlation values corresponding to the same channel pair in the set is calculated, and the sum of the absolute values corresponding to the plurality of channel pairs is calculated.
  • the sum of the absolute values is less than the change threshold, it is determined that the target channel pair set of the first audio frame does not need to be re-acquired; when the sum of the absolute values is greater than or equal to the change threshold, it is determined that the target channel of the first audio frame needs to be re-acquired pair collection.
  • the correlation value difference corresponds to the same channel pair, calculate the correlation value difference respectively, and then calculate the sum of the absolute values of the difference values of all channel pairs, so that the first audio frame relative to the second audio frame, the signal of each channel can be obtained.
  • the change of the correlation value between the two exceeds the change threshold, if not, it means that the change from the second audio frame to the first audio frame is not large, and it is not necessary to rebuild the target channel pair set for the first audio frame, which reduces the calculation If it exceeds, it means that the change from the second audio frame to the first audio frame is large, and the target channel pair set of the first audio frame needs to be re-acquired.
  • Step 605 if it is necessary to re-acquire the target channel pair set of the first audio frame, then adopt the method of the embodiment shown in FIG. 3 or FIG. 5 to obtain the target channel pair set of the first audio frame, and according to the target channel pair set The first audio frame is encoded.
  • the method in the embodiment shown in FIG. 3 or FIG. 5 may be used to obtain the correlation value set of the first audio frame, which will not be repeated here.
  • Step 606 if it is not necessary to re-acquire the target channel pair set of the first audio frame, then determine the target channel pair set of the second audio frame as the target channel pair set of the first audio frame, and according to the target channel pair set.
  • the set encodes the first audio frame.
  • the target channel pair set of the second audio frame when it is determined that the target channel pair set of the first audio frame does not need to be re-acquired, the target channel pair set of the second audio frame can be directly used as the target channel pair set of the first audio frame, thereby reducing the amount of calculation. Improve coding efficiency.
  • the channel pair set corresponding to the sum of the values is determined as the target channel pair set, which can maximize the sum of the correlation values of all channel pairs included in the target channel pair set, and increase the number of channel pairs in the group pair as much as possible. It reduces the redundancy between channel signals and improves the coding efficiency of audio.
  • the 5.1 channel includes C, L, R, LS, RS, and LFE.
  • the channel pair set generation module can use the multi-channel mask to remove the channels that do not need multi-channel processing to improve the coding efficiency.
  • the LFE channel can be removed from the 5.1 channel, so the input sound
  • the channel signals of the channel pair set generation module include C, L, R, LS and RS.
  • the method for obtaining the target channel pair set may include the following steps:
  • the formula of the above-mentioned first embodiment can also be used to calculate the correlation value between the two channel signals.
  • the obtained correlation value set may include at most The correlation values of each channel pair are shown in Table 1.
  • the correlation value set of the first audio frame and the correlation value set of the second audio frame are both represented in the form of matrices, and matrices Matrix1 and Matrix2 are obtained respectively, and the value of each element in the matrix corresponds to the value of the correlation value set in the correlation value set.
  • a correlation value, the sum of the differences can be calculated by the following formula:
  • D represents the sum of the difference between the correlation value set of the first audio frame and the correlation value set of the second audio frame
  • Matrix1(i) represents the i-th element value in the matrix corresponding to the correlation value set of the first audio frame
  • Matrix2(i) represents the i-th element value in the matrix corresponding to the correlation value set of the second audio frame.
  • a change threshold is set, and whether the target channel pair set of the first audio frame needs to be re-acquired is defined by the threshold.
  • FIG. 7 is a flowchart of an exemplary embodiment of a method for encoding a multi-channel audio signal provided by the present application.
  • the process 700 may be performed by the source device 12 or the audio coding device 200 in the audio coding system 10 .
  • Process 700 is described as a series of steps or operations, and it should be understood that process 700 may be performed in various orders and/or concurrently, and is not limited to the order of execution shown in FIG. 7 .
  • the method includes:
  • Step 701 Acquire a first audio frame to be encoded, where the first audio frame includes K channel signals.
  • step 701 reference may be made to the above-mentioned step 301, which will not be repeated here.
  • Step 702 When K is greater than the threshold of the number of channel signals, use the method of the embodiment shown in FIG. 3 to encode the first audio frame.
  • Step 703 When K is less than or equal to the threshold of the number of channel signals, use the method of the embodiment shown in FIG. 5 to encode the first audio frame.
  • the difference between this embodiment and the embodiment shown in FIG. 3 or FIG. 5 is that the method in FIG. 3 and FIG. 5 is combined in this embodiment, that is, according to the number of channel signals contained in the first audio frame Which method an audio frame uses to obtain its target channel pair set.
  • the method of the second aspect is adopted, it is necessary to exhaustively enumerate all target channel pairs, which will increase the amount of calculation. Therefore, the method of the first aspect will reduce the amount of calculation. A lot of computation.
  • the method of the second aspect can be used to obtain the sum of the correlation values of all channel pair sets, so as to ensure that the final selected target channel pair set must be the highest Optimal results that match the characteristics of the first audio frame.
  • FIG. 8 is an exemplary structural diagram of a decoding apparatus to which the decoding method for a multi-channel audio signal provided by the present application is applied.
  • the decoding apparatus may be the decoder 30 of the destination device 14 in the audio decoding system 10, or may be is the decoding module 270 in the audio decoding apparatus 200 .
  • the decoding device may include a code stream demultiplexing interface, a channel decoding module and a multi-channel processing module, wherein,
  • the code stream demultiplexing interface receives the encoded multi-channel signal (eg serial bit stream bitstream) from the encoding device, and obtains the encoded channel signal (E) and multi-channel parameters (SIDE_PAIR) after demultiplexing.
  • E encoded channel signal
  • SIDE_PAIR multi-channel parameters
  • the channel decoding module uses a monaural decoding unit (or a monaural box, a monaural tool) to decode the coded channel signal output by the code stream demultiplexing interface and output the decoded channel signal (D).
  • a monaural decoding unit or a monaural box, a monaural tool
  • D decoded channel signal
  • the multi-channel processing module includes a plurality of stereo processing units.
  • the stereo processing unit can adopt prediction-based or KLT-based processing, that is, the input two channel signals are inversely rotated (for example, via a 2 ⁇ 2 rotation matrix), so that the signal Transform to the original signal direction.
  • the decoded channel signal output by the channel decoding module can identify which two decoded channel signal groups are paired by the multi-channel parameters, and input the decoded channel signal of the pair into the stereo processing unit, and the stereo processing unit decodes the two input channel signals.
  • the channel signal (CH) corresponding to the two decoded channel signals is output.
  • stereo processing unit 1 processes D1 and D2 according to SIDE_PAIR1 to obtain CH1 and CH2
  • stereo processing unit 2 processes D3 and D4 according to SIDE_PAIR2 to obtain CH3 and CH4, ...
  • the unpaired channel signal (such as CHj) does not need to be processed by the stereo processing unit in the multi-channel processing module, and can be directly output after decoding.
  • FIG. 9 is a schematic structural diagram of an embodiment of an encoding apparatus of the present application. As shown in FIG. 9 , the apparatus may be applied to the source device 12 or the audio decoding device 200 in the above-mentioned embodiment.
  • the encoding apparatus in this embodiment may include: an acquisition module 901 , an encoding module 902 and a determination module 903 .
  • the obtaining module 901 is configured to obtain a first audio frame to be encoded, where the first audio frame includes at least five channel signals; obtain a correlation value set, where the correlation value set includes multiple Correlation values of each channel pair, where one channel pair includes two channel signals in the at least five channel signals, and the correlation value of the one channel pair is used to represent the two channel signals of the one channel pair.
  • M correlation values are selected from the correlation value set, and the M correlation values are all greater than other correlation values in the correlation value set except the M correlation values,
  • the M correlation values are all greater than or equal to the group pair threshold, and M is a positive integer less than or equal to a set value; acquire M channel pair sets, each of which includes at least the M correlation values one of the corresponding M channel pairs, and when the channel pair set includes more than two channel pairs, the two or more channel pairs do not contain the same channel signal; the determining module 903, using In determining a target channel pair set from the M channel pair sets, the sum of the correlation values of all channel pairs in the target channel pair set is the largest in the M channel pair sets; encoding Module 902, configured to encode the first audio frame according to the target channel pair set.
  • the M channel pair sets include a first channel pair set; the acquiring module 901 is specifically configured to add the first channel pair in the M channel pairs The first channel pair set, the first channel pair is any one of the M channel pairs; when the other channel pairs except the associated channel in the multiple channel pairs include: When the channel pair whose correlation value is greater than the set pair threshold, select a channel pair with the largest correlation value from the other channel pairs and add it to the first channel pair set, where the associated channel pair includes the added channel pair. Any one of the channel signals included in the channel pair of the first channel pair set.
  • the obtaining module 901 is specifically configured to select N correlation values from the correlation value set, and the N correlation values are all greater than the N correlation values in the correlation value set divided by the N correlation values.
  • N is the set value; from the N correlation values, select the correlation value greater than or equal to the threshold of the group pair, and the correlation value greater than or equal to the threshold value of the group pair The number is M.
  • the correlation value is a normalized value.
  • the correlation value of the one channel pair is set to 0.
  • the obtaining module 901 is configured to obtain a first audio frame to be encoded, where the first audio frame includes at least five channel signals; obtain a correlation value set, where the correlation value set includes multiple Correlation values of the respective channel pairs, one channel pair includes two channel signals in the at least five channel signals, and the correlation value of the one channel pair is used to represent the two channel signals of the one channel pair.
  • multiple channel pair sets are obtained according to the multiple channel pairs, when the channel pair set includes more than two channel pairs, the two or more channel pairs does not contain the same channel signal; obtains the sum of the correlation values of all channel pairs included in each channel pair set in the multiple channel pair sets according to the correlation value set; the determining module 903 is used to determine the target A channel pair set, where the sum of the correlation values of all channel pairs in the target channel pair set is the largest among the multiple channel pair sets; the encoding module 902 is configured to, according to the target channel pair set The first audio frame is encoded.
  • the obtaining module 901 is specifically configured to obtain the set of multiple channel pairs according to other channel pairs in the multiple channel pairs that are not related to the outside Correlation values for related channel pairs are less than the group pair threshold.
  • the obtaining module 901 is configured to obtain a first audio frame to be encoded, where the first audio frame includes at least five channel signals; obtain a correlation value set of the first audio frame, The correlation value set of the first audio frame includes respective correlation values of a plurality of channel pairs, one channel pair includes two channel signals in the at least five channel signals, and the correlation value of the one channel pair is The value is used to represent the correlation between the two channel signals of the one channel pair; the correlation value set of the second audio frame is obtained, and the correlation value set of the second audio frame includes the correlation value set of the second audio frame.
  • Correlation values for each of a plurality of channel pairs where one channel pair includes two channel signals in at least five channel signals of the second audio frame, and the correlation value of the one channel pair is used to represent the Correlation between two channel signals of a channel pair, the second audio frame is the previous frame of the first audio frame; the encoding module 902 is configured to use the correlation value of the first audio frame according to The set and the correlation value set of the second audio frame determine whether it is necessary to re-acquire the target channel pair set of the first audio frame; if it is necessary to re-acquire the target channel pair set of the first audio frame, execute the 3 or the method of the embodiment shown in FIG.
  • the 5 acquires the target channel pair set of the first audio frame, and encodes the first audio frame according to the target channel pair set;
  • the target channel pair set of the first audio frame, then the target channel pair set of the second audio frame is determined as the target channel pair set of the first audio frame, and according to the target channel pair set
  • the first audio frame is encoded.
  • the encoding module 902 is specifically configured to calculate the correlation value corresponding to the same channel pair in the correlation value set of the first audio frame and the correlation value set of the second audio frame The absolute value of the difference; calculate the sum of the absolute values corresponding to a plurality of the channel pairs respectively; when the sum of the absolute values is less than the change threshold, it is determined that the target sound of the first audio frame does not need to be re-acquired A channel pair set; when the sum of the absolute values is greater than or equal to the change threshold, it is determined that the target channel pair set of the first audio frame needs to be re-acquired.
  • the acquisition module is used to acquire the first audio frame to be encoded, the first audio frame includes K channel signals, and K is an integer greater than or equal to 5; the encoding module is used for When K is greater than the threshold of the number of channel signals, the method of the embodiment shown in FIG. 3 is performed to encode the first audio frame; when K is less than or equal to the threshold of the number of channel signals, the method of the embodiment shown in FIG. 5 is performed The first audio frame is encoded.
  • the apparatus of this embodiment can be used to implement the technical solutions of the method embodiments shown in FIG. 3 , FIG. 5 , FIG. 6 or FIG. 7 , and the implementation principles and technical effects thereof are similar, and are not repeated here.
  • FIG. 10 is a schematic structural diagram of an embodiment of a device of the present application.
  • the device may be the encoding device in the above-mentioned embodiment.
  • the device in this embodiment may include: a processor 1001 and a memory 1002, where the memory 1002 is used to store one or more programs; when the one or more programs are executed by the processor 1001, the processor 1001 realizes the The technical solution of the method embodiment shown in FIG. 3 , FIG. 5 , FIG. 6 or FIG. 7 .
  • each step of the above method embodiments may be completed by a hardware integrated logic circuit in a processor or an instruction in the form of software.
  • the processor can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other Programming logic devices, discrete gate or transistor logic devices, discrete hardware components.
  • DSP digital signal processor
  • ASIC application-specific integrated circuit
  • FPGA field programmable gate array
  • a general purpose processor may be a microprocessor or the processor may be any conventional processor or the like.
  • the steps of the method disclosed in the present application can be directly embodied as executed by a hardware encoding processor, or executed by a combination of hardware and software modules in the encoding processor.
  • the software modules may be located in random access memory, flash memory, read-only memory, programmable read-only memory or electrically erasable programmable memory, registers and other storage media mature in the art.
  • the storage medium is located in the memory, and the processor reads the information in the memory, and completes the steps of the above method in combination with its hardware.
  • the memory mentioned in the above embodiments may be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory.
  • the non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically programmable Erase programmable read-only memory (electrically EPROM, EEPROM) or flash memory.
  • Volatile memory may be random access memory (RAM), which acts as an external cache.
  • RAM random access memory
  • DRAM dynamic random access memory
  • SDRAM synchronous DRAM
  • SDRAM double data rate synchronous dynamic random access memory
  • ESDRAM enhanced synchronous dynamic random access memory
  • SLDRAM synchronous link dynamic random access memory
  • direct rambus RAM direct rambus RAM
  • the disclosed system, apparatus and method may be implemented in other manners.
  • the apparatus embodiments described above are only illustrative.
  • the division of the units is only a logical function division. In actual implementation, there may be other division methods.
  • multiple units or components may be combined or Can be integrated into another system, or some features can be ignored, or not implemented.
  • the shown or discussed mutual coupling or direct coupling or communication connection may be through some interfaces, indirect coupling or communication connection of devices or units, and may be in electrical, mechanical or other forms.
  • the units described as separate components may or may not be physically separated, and components displayed as units may or may not be physical units, that is, may be located in one place, or may be distributed to multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution in this embodiment.
  • each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically alone, or two or more units may be integrated into one unit.
  • the functions, if implemented in the form of software functional units and sold or used as independent products, may be stored in a computer-readable storage medium.
  • the technical solution of the present application can be embodied in the form of a software product in essence, or the part that contributes to the prior art or the part of the technical solution.
  • the computer software product is stored in a storage medium, including Several instructions are used to cause a computer device (personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application.
  • the aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), magnetic disk or optical disk and other media that can store program codes .

Landscapes

  • Engineering & Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • Computational Linguistics (AREA)
  • Signal Processing (AREA)
  • Health & Medical Sciences (AREA)
  • Audiology, Speech & Language Pathology (AREA)
  • Human Computer Interaction (AREA)
  • Acoustics & Sound (AREA)
  • Multimedia (AREA)
  • Mathematical Physics (AREA)
  • Compression, Expansion, Code Conversion, And Decoders (AREA)

Abstract

一种多声道音频信号的编解码方法和装置。该多声道音频信号的编码方法包括:获取待编码的第一音频帧(S301);获取相关值集合(S302),相关值集合包括多个声道对各自的相关值,一个声道对包括至少五个声道信号中的两个声道信号;从相关值集合中选取M个相关值(S303),该M个相关值均大于相关值集合中除M个相关值外的其他相关值,M个相关值均大于或等于组对阈值;获取M个声道对集合(S304),每个声道对集合至少包括M个相关值对应的M个声道对的其中之一;从M个声道对集合中确定目标声道对集合(S305),目标声道对集合中的所有声道对的相关值之和是M个声道对集合中最大的;根据目标声道对集合对第一音频帧进行编码(S306)。减少声道信号之间的冗余,提升音频的编码效率。

Description

多声道音频信号的编解码方法和装置
本申请要求于2020年7月17日提交中国专利局、申请号为202010699706.7、申请名称为“多声道音频信号的编解码方法和装置”的中国专利申请的优先权,其全部内容通过引用结合在本申请中。
技术领域
本申请涉及音频处理技术,尤其涉及一种多声道音频信号的编解码方法和装置。
背景技术
多声道音频的编解码是对包含两个以上声道的音频进行编码或解码的技术。常见的多声道音频有5.1声道音频、7.1声道音频、7.1.4声道音频以及22.2声道音频等。
MPEG环绕声(MPEG Surround,MPS)标准规定了针对四个声道的联合编码,但仍需有可以针对上述各种多声道音频信号的编解码方法。
发明内容
本申请提供一种多声道音频信号的编解码方法和装置,以减少声道信号之间的冗余,提升音频的编码效率。
第一方面,本申请提供一种多声道音频信号的编码方法,包括:获取待编码的第一音频帧,所述第一音频帧包括至少五个声道信号;获取相关值集合,所述相关值集合包括多个声道对各自的相关值,一个声道对包括所述至少五个声道信号中的两个声道信号,所述一个声道对的相关值用于表示所述一个声道对的两个声道信号之间的相关性;从所述相关值集合中选取M个相关值,所述M个相关值均大于所述相关值集合中除所述M个相关值外的其他相关值,所述M个相关值均大于或等于组对阈值,M为小于或等于设定值的正整数;获取M个声道对集合,每个所述声道对集合至少包括所述M个相关值对应的M个声道对的其中之一,且当所述声道对集合包括两个以上声道对时,所述两个以上声道对不包含相同的声道信号;从所述M个声道对集合中确定目标声道对集合,所述目标声道对集合中的所有声道对的相关值之和是所述M个声道对集合中最大的;根据所述目标声道对集合对所述第一音频帧进行编码。
本实施例的第一音频帧可以是待编码的多声道音频信号中的任意一个帧,该第一音频帧包括了五个或五个以上的声道信号。将相关性越高的两个声道信号放在一起编码可以减少冗余,提升编码效率,因此本实施例在组对时,是依据两个声道信号之间的相关值来确定的。为了尽可能找寻相关性最高的声道对集合,可以计算第一音频帧中的至少五个声道信号中两两之间的相关值得到第一音频帧的相关值集合。例如五个声道信号一共可以组成10个声道对,相对应的,相关值集合中可以包括10个相关值。本实施例可以将相关值集合中包括的所有相关值按照从大到小的顺序排序,从中选取排在前面的前M个相关值,该M个相关值必须是大于或等于组对阈值的,这是因为小于组对阈值的相关值,表示其所对应的声道对中的两个声道信号之间的相关性较低,没有组对编码的必要。而为了提高编码效率,无需把所有大于或等于组对阈值的相关值全都选出来,因此设定了一个M的 上限N,即最多选取N个相关值即可。
本实施例通过尽量多的获取多个声道对集合的相关值之和,进而将最大相关值之和对应的声道对集合确定为目标声道对集合,可以实现目标声道对集合所包含的所有声道对的相关值之和最大,并尽可能增加组对的声道对的个数,减少声道信号之间的冗余,提升音频的编码效率。
在一种可能的实现方式中,所述M个声道对集合包括第一声道对集合,所述获取M个声道对集合获取所述第一声道对集合;所述获取所述第一声道对集合,包括:将所述M个声道对中的第一声道对加入所述第一声道对集合,所述第一声道对为所述M个声道对中的任意一个;当所述多个声道对中除关联声道对外的其他声道对中包括相关值大于所述组对阈值的声道对时,从所述其他声道对中选取相关值最大的一个声道对加入所述第一声道对集合,所述关联声道对包括已加入所述第一声道对集合的声道对所包括的声道信号中的任意一个。
将多个声道对中,相关值大小较大的多个声道对分别作为声道对集合中加入的第一个声道对,然后选取剩余声道对中最大相关值对应的声道对加入对应的声道对集合,通过尽量多的获取多个声道对集合的相关值之和,进而将最大相关值之和对应的声道对集合确定为目标声道对集合,可以实现目标声道对集合所包含的所有声道对的相关值之和最大,并尽可能增加组对的声道对的个数,减少声道信号之间的冗余,提升音频的编码效率。
在一种可能的实现方式中,所述从所述相关值集合中选取M个相关值,包括:从所述相关值集合中选取N个相关值,所述N个相关值均大于所述相关值集合中除所述N个相关值外的其他相关值,N为所述设定值;从所述N个相关值中选取大于或等于所述组对阈值的相关值,所述大于或等于所述组对阈值的相关值的个数为M。
M个相关值大于或等于组对阈值,M为小于或等于设定值(例如N)的正整数。本实施例可以将相关值集合中包括的所有相关值按照从大到小的顺序排序,从中选取排在前面的前N个相关值,该N个相关值可能存在小于组对阈值的相关值,因此从N个相关值中选取大于或等于组对阈值的M个相关值,这是因为小于组对阈值的相关值,表示其所对应的声道对中的两个声道信号之间的相关性较低,没有组对编码的必要。
在一种可能的实现方式中,所述相关值为经归一化处理的值。
归一化处理可以将取值范围差别较大的相关值纳入一个统一的范围内进行比较和处理,提高运算效率。
在一种可能的实现方式中,当所述一个声道对的相关值小于所述组对阈值时,所述一个声道对的相关值设置为0。
较小的相关值说明对应的两个声道信号之间的相关性较小,没有组对的必要,因此将这种情况的两个声道信号的相关值设置为0,便于后续计算,提高运算效率。
第二方面,本申请提供一种多声道音频信号的编码方法,包括:获取待编码的第一音频帧,所述第一音频帧包括至少五个声道信号;获取相关值集合,所述相关值集合包括多个声道对各自的相关值,一个声道对包括所述至少五个声道信号中的两个声道信号,所述一个声道对的相关值用于表示所述一个声道对的两个声道信号之间的相关性;根据所述多个声道对获取多个声道对集合,当所述声道对集合包括两个以上声道对时,所述两个以上声道对不包含相同的声道信号;根据所述相关值集合获取所述多个声道对集合中每一个声 道对集合包含的所有声道对的相关值之和;确定目标声道对集合,所述目标声道对集合中的所有声道对的相关值之和是所述多个声道对集合中最大的;根据所述目标声道对集合对所述第一音频帧进行编码。
通过尽可能多的获取多个声道对集合的相关值之和,进而将最大相关值之和对应的声道对集合确定为目标声道对集合,可以实现目标声道对集合所包含的所有声道对的相关值之和最大,并尽可能增加组对的声道对的个数,减少声道信号之间的冗余,提升音频的编码效率。
在一种可能的实现方式中,所述根据所述多个声道对获取多个声道对集合,包括:根据所述多个声道对中除非相关声道对外的其他声道对获取所述多个声道对集合,所述非相关声道对的相关值小于组对阈值。
较小的相关值说明对应的两个声道信号之间的相关性较小,没有组对的必要,因此将这种情况的两个声道信号的相关值及该两个声道信号的声道对删除,可以减少后续计算量,提高运算效率。
在一种可能的实现方式中,所述相关值为经归一化处理的值。
归一化处理可以将取值范围差别较大的相关值纳入一个统一的范围内进行比较和处理,提高运算效率。
在一种可能的实现方式中,当所述一个声道对的相关值小于组对阈值时,所述一个声道对的相关值设置为0。
较小的相关值说明对应的两个声道信号之间的相关性较小,没有组对的必要,因此将这种情况的两个声道信号的相关值设置为0,便于后续计算,提高运算效率。
第三方面,本申请提供一种多声道音频信号的编码方法,包括:获取待编码的第一音频帧,所述第一音频帧包括至少五个声道信号;获取所述第一音频帧的相关值集合,所述第一音频帧的相关值集合包括多个声道对各自的相关值,一个声道对包括所述至少五个声道信号中的两个声道信号,所述一个声道对的相关值用于表示所述一个声道对的两个声道信号之间的相关性;获取第二音频帧的相关值集合,所述第二音频帧的相关值集合包括所述第二音频帧的多个声道对各自的相关值,一个声道对包括所述第二音频帧的至少五个声道信号中的两个声道信号,所述一个声道对的相关值用于表示所述一个声道对的两个声道信号之间的相关性,所述第二音频帧是所述第一音频帧的上一帧;根据所述第一音频帧的相关值集合和所述第二音频帧的相关值集合判断是否需要重新获取所述第一音频帧的目标声道对集合;若需要重新获取所述第一音频帧的目标声道对集合,则采用如上述第一至二方面中任一项所述的方法获取所述第一音频帧的目标声道对集合,并根据所述目标声道对集合对所述第一音频帧进行编码;若不需要重新获取所述第一音频帧的目标声道对集合,则将所述第二音频帧的目标声道对集合确定为所述第一音频帧的目标声道对集合,并根据所述目标声道对集合对所述第一音频帧进行编码。
通过获取当前音频帧的相关值集合和上一音频帧的相关值集合的差值之和,从而确定是否需要重新获取当前帧的目标声道对集合,可以在音频变化较小的情况下,大大减少计算量,提高编码效率,而即使音频变化较大,需要重新获取目标声道对集合,仍可以尽可能多的获取多个声道对集合的相关值之和,进而将最大相关值之和对应的声道对集合确定为目标声道对集合,可以实现目标声道对集合所包含的所有声道对的相关值之和最大,并 尽可能增加组对的声道对的个数,减少声道信号之间的冗余,提升音频的编码效率。
在一种可能的实现方式中,所述根据所述第一音频帧的相关值集合和所述第二音频帧的相关值集合判断是否需要重新获取所述第一音频帧的目标声道对集合,包括:计算所述第一音频帧的相关值集合和所述第二音频帧的相关值集合中对应于同一声道对的相关值之差的绝对值;计算多个所述声道对分别对应的所述绝对值之和;当所述绝对值之和小于变更阈值时,确定不需要重新获取所述第一音频帧的目标声道对集合;当所述绝对值之和大于或等于所述变更阈值时,确定需要重新获取所述第一音频帧的目标声道对集合。变更阈值例如可以是α×声道对的个数,其中,α的取值可以是0.14或者0.15,声道对的个数是指第一音频帧的相关值集合(或者第二音频帧的相关值集合)中包括的声道对的个数。
第四方面,本申请提供一种多声道音频信号的编码方法,包括:获取待编码的第一音频帧,所述第一音频帧包括K个声道信号,K为大于或等于5的整数;当K大于声道信号数量阈值时,采用上述第一方面中任一项所述的方法对所述第一音频帧进行编码;当K小于或等于声道信号数量阈值时,采用上述第二方面中任一项所述的方法对所述第一音频帧进行编码。声道信号数量阈值例如可以是5、6或者7等。
与第一方面或第二方面的方法的区别在于,将第一方面和第二方面的方法进行融合,即根据第一音频帧包含的声道信号的个数来确定对第一音频帧采用哪一种方法获取其目标声道对集合。当第一音频帧包含的声道信号的个数较多时,如果采用第二方面的方法,需要穷举所有目标声道对集合,会增加计算量,因此此时采用第一方面的方法会减少很多的计算量。而当第一音频帧包含的声道信号的个数较少时,采用第二方面的方法可以获取到所有声道对集合的相关值之和,确保最终选取的目标声道对集合一定是最符合第一音频帧的特性的最优结果。
第五方面,本申请提供一种编码装置,包括:获取模块,用于获取待编码的第一音频帧,所述第一音频帧包括至少五个声道信号;获取相关值集合,所述相关值集合包括多个声道对各自的相关值,一个声道对包括所述至少五个声道信号中的两个声道信号,所述一个声道对的相关值用于表示所述一个声道对的两个声道信号之间的相关性;从所述相关值集合中选取M个相关值,所述M个相关值均大于所述相关值集合中除所述M个相关值外的其他相关值,所述M个相关值均大于或等于组对阈值,M为小于或等于设定值的正整数;获取M个声道对集合,每个所述声道对集合至少包括所述M个相关值对应的M个声道对的其中之一,且当所述声道对集合包括两个以上声道对时,所述两个以上声道对不包含相同的声道信号;确定模块,用于从所述M个声道对集合中确定目标声道对集合,所述目标声道对集合中的所有声道对的相关值之和是所述M个声道对集合中最大的;编码模块,用于根据所述目标声道对集合对所述第一音频帧进行编码。
在一种可能的实现方式中,所述M个声道对集合包括第一声道对集合;所述获取模块,具体用于将所述M个声道对中的第一声道对加入所述第一声道对集合,所述第一声道对为所述M个声道对中的任意一个;当所述多个声道对中除关联声道对外的其他声道对中包括相关值大于所述组对阈值的声道对时,从所述其他声道对中选取相关值最大的一个声道对加入所述第一声道对集合,所述关联声道对包括已加入所述第一声道对集合的声道对所包括的声道信号中的任意一个。
在一种可能的实现方式中,所述获取模块,具体用于从所述相关值集合中选取N个相 关值,所述N个相关值均大于所述相关值集合中除所述N个相关值外的其他相关值,N为所述设定值;从所述N个相关值中选取大于或等于所述组对阈值的相关值,所述大于或等于所述组对阈值的相关值的个数为M。
在一种可能的实现方式中,所述相关值为经归一化处理的值。
在一种可能的实现方式中,当所述一个声道对的相关值小于所述组对阈值时,所述一个声道对的相关值设置为0。
第六方面,本申请提供一种编码装置,包括:获取模块,用于获取待编码的第一音频帧,所述第一音频帧包括至少五个声道信号;获取相关值集合,所述相关值集合包括多个声道对各自的相关值,一个声道对包括所述至少五个声道信号中的两个声道信号,所述一个声道对的相关值用于表示所述一个声道对的两个声道信号之间的相关性;根据所述多个声道对获取多个声道对集合,当所述声道对集合包括两个以上声道对时,所述两个以上声道对不包含相同的声道信号;根据所述相关值集合获取所述多个声道对集合中每一个声道对集合包含的所有声道对的相关值之和;确定模块,用于确定目标声道对集合,所述目标声道对集合中的所有声道对的相关值之和是所述多个声道对集合中最大的;编码模块,用于根据所述目标声道对集合对所述第一音频帧进行编码。
在一种可能的实现方式中,所述获取模块,具体用于根据所述多个声道对中除非相关声道对外的其他声道对获取所述多个声道对集合,所述非相关声道对的相关值小于组对阈值。
在一种可能的实现方式中,所述相关值为经归一化处理的值。
在一种可能的实现方式中,当所述一个声道对的相关值小于组对阈值时,所述一个声道对的相关值设置为0。
第七方面,本申请提供一种编码装置,包括:获取模块,用于获取待编码的第一音频帧,所述第一音频帧包括至少五个声道信号;获取所述第一音频帧的相关值集合,所述第一音频帧的相关值集合包括多个声道对各自的相关值,一个声道对包括所述至少五个声道信号中的两个声道信号,所述一个声道对的相关值用于表示所述一个声道对的两个声道信号之间的相关性;获取第二音频帧的相关值集合,所述第二音频帧的相关值集合包括所述第二音频帧的多个声道对各自的相关值,一个声道对包括所述第二音频帧的至少五个声道信号中的两个声道信号,所述一个声道对的相关值用于表示所述一个声道对的两个声道信号之间的相关性,所述第二音频帧是所述第一音频帧的上一帧;编码模块,用于根据所述第一音频帧的相关值集合和所述第二音频帧的相关值集合判断是否需要重新获取所述第一音频帧的目标声道对集合;若需要重新获取所述第一音频帧的目标声道对集合,则执行如权利要求1-9中任一项所述的方法获取所述第一音频帧的目标声道对集合,并根据所述目标声道对集合对所述第一音频帧进行编码;若不需要重新获取所述第一音频帧的目标声道对集合,则将所述第二音频帧的目标声道对集合确定为所述第一音频帧的目标声道对集合,并根据所述目标声道对集合对所述第一音频帧进行编码。
在一种可能的实现方式中,所述编码模块,具体用于计算所述第一音频帧的相关值集合和所述第二音频帧的相关值集合中对应于同一声道对的相关值之差的绝对值;计算多个所述声道对分别对应的所述绝对值之和;当所述绝对值之和小于变更阈值时,确定不需要重新获取所述第一音频帧的目标声道对集合;当所述绝对值之和大于或等于所述变更阈值 时,确定需要重新获取所述第一音频帧的目标声道对集合。
第八方面,本申请提供一种编码装置,包括:获取模块,用于获取待编码的第一音频帧,所述第一音频帧包括K个声道信号,K为大于或等于5的整数;编码模块,用于当K大于声道信号数量阈值时,执行如上述第一方面中任一项所述的方法对所述第一音频帧进行编码;当K小于或等于声道信号数量阈值时,执行如上述第二方面中任一项所述的方法对所述第一音频帧进行编码。
第九方面,本申请提供一种设备,包括:一个或多个处理器;存储器,用于存储一个或多个程序;当所述一个或多个程序被所述一个或多个处理器执行,使得所述一个或多个处理器实现如上述第一至四方面中任一项所述的方法。
第十方面,本申请提供一种计算机可读存储介质,包括计算机程序,所述计算机程序在计算机上被执行时,使得所述计算机执行上述第一至四方面中任一项所述的方法。
第十一方面,本申请提供一种计算机可读存储介质,其特征在于,包括根据如上述第一至四方面中任一项所述的多声道音频信号的编码方法获得的编码码流。
附图说明
图1示例性地给出了本申请所应用的音频译码系统10的示意性框图;
图2示例性地给出了本申请所应用的音频译码设备200的示意性框图;
图3是本申请提供的多声道音频信号的编码方法的一个示例性的实施例的流程图;
图4是本申请提供的多声道音频信号的编方法所应用的编码装置的一个示例性的结构图;
图5是本申请提供的多声道音频信号的编码方法的一个示例性的实施例的流程图;
图6是本申请提供的多声道音频信号的编码方法的一个示例性的实施例的流程图;
图7是本申请提供的多声道音频信号的编码方法的一个示例性的实施例的流程图;
图8是本申请提供的多声道音频信号的解码方法所应用的解码装置的一个示例性的结构图;
图9为本申请编码装置实施例的结构示意图;
图10为本申请设备实施例的结构示意图。
具体实施方式
为使本申请的目的、技术方案和优点更加清楚,下面将结合本申请中的附图,对本申请中的技术方案进行清楚、完整地描述,显然,所描述的实施例是本申请一部分实施例,而不是全部的实施例。基于本申请中的实施例,本领域普通技术人员在没有作出创造性劳动前提下所获取的所有其他实施例,都属于本申请保护的范围。
本申请的说明书实施例和权利要求书及附图中的术语“第一”、“第二”等仅用于区分描述的目的,而不能理解为指示或暗示相对重要性,也不能理解为指示或暗示顺序。此外,术语“包括”和“具有”以及他们的任何变形,意图在于覆盖不排他的包含,例如,包含了一系列步骤或单元。方法、系统、产品或设备不必限于清楚地列出的那些步骤或单元,而是可包括没有清楚地列出的或对于这些过程、方法、产品或设备固有的其它步骤或单元。
应当理解,在本申请中,“至少一个(项)”是指一个或者多个,“多个”是指两个或两 个以上。“和/或”,用于描述关联对象的关联关系,表示可以存在三种关系,例如,“A和/或B”可以表示:只存在A,只存在B以及同时存在A和B三种情况,其中A,B可以是单数或者复数。字符“/”一般表示前后关联对象是一种“或”的关系。“以下至少一项(个)”或其类似表达,是指这些项中的任意组合,包括单项(个)或复数项(个)的任意组合。例如,a,b或c中的至少一项(个),可以表示:a,b,c,“a和b”,“a和c”,“b和c”,或“a和b和c”,其中a,b,c可以是单个,也可以是多个。
本申请涉及到的相关名词解释:
音频帧:音频数据是流式的,在实际应用中,为了便于音频处理和传输,通常取一时长内的音频数据量作为一帧音频,该时长被称为“采样时间”,可以根据编解码器和具体应用的需求确定其值,例如该时长为2.5ms~60ms,ms为毫秒。
音频信号:音频信号是带有语音、音乐和音效的有规律的声波的频率、幅度变化信息载体。音频是一种连续变化的模拟信号,可用一条连续的曲线来表示,称为声波。音频通过模数转换或计算机生成的数字信号即为音频信号。声波有三个重要参数:频率、幅度和相位,这也就决定了音频信号的特征。
声道信号:是指声音在录制或播放时在不同空间位置采集或回放的相互独立的音频信号。因此声道数也就是声音录制时的音源数量或回放时的扬声器数量。
以下是本申请所应用的系统架构。
图1示例性地给出了本申请所应用的音频译码系统10的示意性框图。如图1所示,音频译码系统10可包括源设备12和目的设备14,源设备12产生经编码的码流,因此,源设备12可被称为音频编码装置。目的设备14可对由源设备12所产生的经编码的码流进行解码,因此,目的设备14可被称为音频解码装置。
源设备12包括编码器20,可选地,可包括音频源16、音频预处理器18、通信接口22。
音频源16可包括或可以为任意类型的用于捕获现实世界语音、音乐和音效等的音频捕获设备,和/或任意类型的音频生成设备,例如用于生成语音、音乐和音效的音频处理器或设备。所述音频源可以为存储上述音频的任意类型的内存或存储器。
音频预处理器18用于接收(原始)音频数据17,并对音频数据17进行预处理,得到预处理音频数据19。例如,音频预处理器18执行的预处理可包括修剪或去噪。可以理解的是,音频预处理单元18可以为可选组件。
编码器20用于接收预处理音频数据19并提供编码音频数据21。
源设备12中的通信接口22可用于接收编码音频数据21并通过通信信道13向目的设备14发送编码音频数据21,以便存储或直接重建。
目的设备14包括解码器30,可选地,可包括通信接口28、音频后处理器32和播放设备34。
目的设备14中的通信接口28用于直接从源设备12接收编码音频数据21,并将编码音频数据21提供给解码器30。
通信接口22和通信接口28可用于通过源设备12与目的设备14之间的直连通信链路,例如直接有线或无线连接等,或者通过任意类型的网络,例如有线网络、无线网络或其任意组合、任意类型的私网和公网或其任意类型的组合,发送或接收编码音频数据21。
例如,通信接口22可用于将编码音频数据21封装为报文等合适的格式,和/或使用任意类型的传输编码或处理来处理编码音频数据21,以便在通信链路或通信网络上进行传输。
通信接口28与通信接口22对应,例如,可用于接收传输数据,并使用任意类型的对应传输解码或处理和/或解封装,对传输数据进行处理,得到编码音频数据21。
通信接口22和通信接口28均可配置为如图1中从源设备12指向目的设备14的对应通信信道13的箭头所指示的单向通信接口,或双向通信接口,并且可用于发送和接收消息等,以建立连接,确认并交换与通信链路和/或编码音频数据等数据传输相关的任何其它信息,等等。
解码器30用于接收编码音频数据21并提供解码音频数据31。
音频后处理器32用于对解码音频数据31进行后处理,得到后处理后的后处理音频数据33。音频后处理器32执行的后处理可以包括例如修剪或重采样等。
播放设备34用于接收后处理音频数据33,以向用户或收听者播放音频。播放设备34可以为或包括任意类型的用于播放重建后音频的播放器,例如,集成或外部扬声器。例如,扬声器可包括喇叭、音响等。
图2示例性地给出了本申请所应用的音频译码设备200的示意性框图。在一个实施例中,音频译码设备200可以是音频解码器(例如图1的解码器30)或音频编码器(例如图1的编码器20)。
音频译码设备200包括:用于接收数据的入端口210和接收单元(Rx)220,用于处理数据的处理器、逻辑单元或中央处理器230,用于传输数据的发射单元(Tx)240和出端口250,以及,用于存储数据的存储器260。音频译码设备200还可以包括与入端口210、接收单元220、发射单元240和出端口250耦合的光电转换组件和电光(EO)组件,用于光信号或电信号的出口或入口。
处理器230通过硬件和软件实现。处理器230可以实现为一个或多个CPU芯片、核(例如,多核处理器)、FPGA、ASIC和DSP。处理器230与入端口210、接收单元220、发射单元240、出端口250和存储器260通信。处理器230包括译码模块270(例如编码模块或解码模块)。译码模块270实现本申请中所公开的实施例,以实现本申请所提供的多声道音频信号的编解码方法。例如,译码模块270实现、处理或提供各种编码操作。因此,通过译码模块270为音频译码设备200的功能提供了实质性的改进,并影响了音频译码设备200到不同状态的转换。或者,以存储在存储器260中并由处理器230执行的指令来实现译码模块270。
存储器260包括一个或多个磁盘、磁带机和固态硬盘,可以用作溢出数据存储设备,用于在选择性地执行这些程序时存储程序,并存储在程序执行过程中读取的指令和数据。存储器260可以是易失性和/或非易失性的,可以是只读存储器(ROM)、随机存取存储器(RAM)、随机存取存储器(ternary content-addressable memory,TCAM)和/或静态随机存取存储器(SRAM)。
基于上述实施例的描述,本申请提供了一种多声道音频信号的编解码方法。
图3是本申请提供的多声道音频信号的编码方法的一个示例性的实施例的流程图。该过程300可由音频译码系统10中的源设备12或音频译码设备200执行。过程300描述为 一系列的步骤或操作,应当理解的是,过程300可以以各种顺序执行和/或同时发生,不限于图3所示的执行顺序。如图3所示,该方法包括:
步骤301、获取待编码的第一音频帧。
本实施例的第一音频帧可以是待编码的多声道音频信号中的任意一个帧,该第一音频帧包括了五个或五个以上的声道信号。例如,5.1声道包括中央声道(C)、前置左声道(left,L)、前置右声道(right,R)、后置左环绕声道(left surround,LS)、后置右环绕声道(right surround,RS)以及0.1声道低频效果(low frequency effects,LFE)共六个声道信号。7.1声道包括C、L、R、LS、RS、LB、RB和LFE共八个声道信号,其中,LFE是从3-120Hz的音频声道,该声道通常发送到专门为低音调而设计的扬声器。
步骤302、获取相关值集合。
相关值集合包括多个声道对各自的相关值,其中一个声道对包括至少五个声道信号中的两个声道信号,一个声道对的相关值用于表示该声道对的两个声道信号之间的相关性。可选的,多个声道对可以包括至少五个声道信号对应的所有声道对,或者,多个声道对也可以包括至少五个声道信号对应的部分声道对,对此不做具体限定。
将相关性越高的两个声道信号放在一起编码可以减少冗余,提升编码效率,因此本实施例在组对时,是依据两个声道信号之间的相关值来确定的。为了尽可能找寻相关性最高的声道对集合,可以先计算第一音频帧中的至少五个声道信号中两两之间的相关值得到第一音频帧的相关值集合。例如五个声道信号一共可以组成10个声道对,相对应的,相关值集合中可以包括10个相关值。
可选的,可以对相关值做归一化处理,这样所有声道对的相关值都限定在特定范围内,以便于设置相关值的统一判断标准,例如组对阈值,该组对阈值可以设置为大于或等于0.2、且小于或等于1的值,例如可以是0.3,0.4,或0.35等等,这样只要两个声道信号的归一化相关值小于组对阈值,就认为该两个声道信号的相关性较差,不需要组对编码。
在一种可能的实现方式中,可以采用以下公式计算两个声道信号(例如ch1和ch2)之间的相关值:
Figure PCTCN2021106101-appb-000001
其中,corr_norm(ch1,ch2)表示声道信号ch1和声道信号ch2之间归一化的相关值,spec_ch1(i)表示声道信号ch1的第i个频点的频域系数,spec_ch2(i)是声道信号ch2的第i个频点的频域系数,N表示一个音频帧的总频点数。
需要说明的是,还可以采用其他的算法或公式计算两个声道信号之间的相关值,本申请对此不做具体限定。
在一些实施方式中,采用上述算法或公式计算的相关值可以作为初始相关值,然后根据预设条件确定是否需要对所述初始相关值进行修改。例如,所述限制条件可以包括:计算所述初始相关值的两个声道信号之间的幅度比值是否大于预设组对阈值。当所述幅度比值大于所述组对阈值时,对所述初始相关值进行修改;当所述幅度比值小于或等于所述组对阈值时,保持所述初始相关值不变。其中,所述修改可以是减小所述初始相关值,例如 可以直接将所述初始相关值修改为0,从而避免所述两个声道信号被组对进行处理。
示例性的,声道信号ch的当前帧的幅度level(ch)可以采用如下公式计算得到:
Figure PCTCN2021106101-appb-000002
其中,i表示声道信号ch的当前帧的第i个样点,N表示当前帧的样点总数,sepc_coeff(ch,i)为当前帧的第i个样点的频域系数。
假设组对幅度阈值ThreholdCoupling=2,当
Figure PCTCN2021106101-appb-000003
或者
Figure PCTCN2021106101-appb-000004
时,将corr_norm(ch1,ch2)置为0,使得ch1和ch2不会被组对。
步骤303、从相关值集合中选取M个相关值。
该M个相关值均大于相关值集合中除该M个相关值外的其他相关值,该M个相关值均大于或等于组对阈值,M为小于或等于设定值(例如N)的正整数。本实施例可以将相关值集合中包括的所有相关值按照从大到小的顺序排序,从中选取排在前面的前M个相关值,该M个相关值必须是大于或等于组对阈值的,这是因为小于组对阈值的相关值,表示其所对应的声道对中的两个声道信号之间的相关性较低,没有组对编码的必要。而为了提高编码效率,无需把所有大于或等于组对阈值的相关值全都选出来,因此设定了一个M的上限N,即最多选取N个相关值即可。
N可以选取大于或等于2的整数,N的最大值也不能超过第一音频帧的所有声道信号对应的所有声道对的个数。N的值越大,伴随的计算量会增加,而N的值越小,可能会出现声道对集合丢失的情况,从而降低编码效率。
可选的,可以将N设置为最大声道对数加一,即
Figure PCTCN2021106101-appb-000005
CH表示第一音频帧包含的声道信号的个数。例如,5.1声道包含五个声道信号(不考虑LFE声道),则N=3;7.1声道包含七个声道信号(不考虑LFE声道),则N=4。
如果相关值集合中不包括大于或等于组对阈值的相关值,则不需要执行后续的步骤,对第一音频帧的各个声道信号分别进行单声道编码即可。如果从相关值集合中选出了M个相关值,则可以执行以下步骤。
步骤304、获取M个声道对集合。
每个声道对集合至少包括M个相关值对应的M个声道对的其中之一,且当声道对集合包括两个以上声道对时,两个以上声道对不包含相同的声道信号。例如,5.1声道,根据相关值集合选出来的最大相关值对应的3个声道对是(L,R)、(R,C)和(LS,RS),其中(LS,RS)的相关值小于组对阈值,因此排除,那么剩余的两个声道对(L,R)和(R,C)可以得到两个声道对集合,这两个声道对集合的其中一个包括(L,R),另一个包括(R,C)。
以M个相关值对应的M个声道对中的任意一个(例如第一声道对)为例,本实施例获取M个声道对集合的方法可以包括:将第一声道对加入第一声道对集合,M个声道对集合包括该第一声道对集合该,当多个声道对中除关联声道对外的其他声道对中包括相关值大于组对阈值的声道对时,从其他声道对中选取相关值最大的一个声道对加入第一声道 对集合,关联声道对包括已加入第一声道对集合的声道对所包括的声道信号中的任意一个。
上述过程除将第一声道对加入第一声道对集合的步骤外,均为迭代处理步骤。即
a、判断多个声道对中除关联声道对外的其他声道对中是否包括相关值大于组对阈值的声道对。
b、若包括相关值大于组对阈值的声道对,则从其他声道对中选取相关值最大的一个声道对加入第一声道对集合。
此时只要其他声道对中包括相关值大于组对阈值的声道对,就可以迭代执行上述步骤b。
可选的,为了减少计算量,可以从相关值集合中将小于组对阈值的相关值删除,这样可以减少声道对的个数,进而减少迭代的次数。
步骤305、从M个声道对集合中确定目标声道对集合。
目标声道对集合中的所有声道对的相关值之和是M个声道对集合中最大的。得到上述M个声道对集合后,可以计算各个声道对集合中包含的所有声道对的相关值之和,最后将相关值之和之和最大的声道对集合确定为目标声道对集合。
步骤306、根据目标声道对集合对第一音频帧进行编码。
根据目标声道对集合对第一音频帧进行编码的过程可参考下文图4所示实施例,此处不再赘述。
可选的,本实施例可以在获取在对第一音频帧编码之前,尤其是在对第一音频帧的至少五个声道信号进行立体声处理之前,先对第一音频帧中的至少五个声道信号分别进行能量均衡处理,得到至少五个均衡声道信号,再对该至少五个均衡声道信号进行立体声处理,此时编码的对象是与均衡声道信号相关的。
能量均衡模式可以包括第一能量均衡模式和/或第二能量均衡模式,其中,第一能量均衡模式仅使用一个声道对中两个声道信号获取一个声道对对应的两个均衡声道信号。第二能量均衡模式使用一个声道对中两个声道信号以及一个声道对外至少一个声道信号来获取一个声道对对应的两个均衡声道信号。
当能量均衡模式为第一能量均衡模式时,可以针对目标声道对集合中的当前声道对,计算当前声道对包含的两个声道信号的能量或幅度值的平均值,根据平均值分别对两个声道信号进行能量均衡处理以得到对应的两个均衡声道信号。这样当至少五个声道信号的波动区间值较大时,可以只在相关的两个声道信号之间进行能量均衡,使得立体声处理时对于比特的分配更符合声道信号的能量特性,避免在低码率的编码环境中能量大的声道对因比特不足导致编码噪声可能会远大于能量小的声道对的编码噪声,而能量小的声道对的比特会有冗余的问题。
当能量均衡模式为第二能量均衡模式时,可以计算至少五个声道信号的能量或幅度值的平均值,根据平均值分别对至少五个声道信号进行能量均衡处理得到至少五个均衡声道信号。
本实施例通过尽量多的获取多个声道对集合的相关值之和,进而将最大相关值之和对应的声道对集合确定为目标声道对集合,可以实现目标声道对集合所包含的所有声道对的相关值之和最大,并尽可能增加组对的声道对的个数,减少声道信号之间的冗余,提升音频的编码效率。
以下通过两个具体的实施例对图3所示方法实施例中如何获取目标声道对集合的过程进行描述。
图4是本申请提供的多声道音频信号的编码方法所应用的编码装置的一个示例性的结构图,该编码装置可以是音频译码系统10中的源设备12的编码器20,也可以是音频译码设备200中的译码模块270。该编码装置可以包括声道对集合生成模块、多声道处理模块、声道编码模块和码流复用接口,其中,
声道对集合生成模块的输入是多声道音频的n个声道信号(CH1-CHn),n是大于或等于5的整数,该n个声道信号均可以进行立体声处理的。声道对集合生成模块计算n个声道信号中任意两个声道信号之间的相关值,从而根据这些相关值采用图3所示实施例的方法得到目标声道对集合,例如,(CH1,CH2),(CH3,CH4),…,(CHi-1,CHi)。
多声道处理模块包括多个立体声处理单元,立体声处理单元可以采用基于预测的或者基于Karhunen-Loeve变换(Karhunen-Loeve Transform,KLT)的处理,即输入的两个声道信号被旋转(例如经由2×2旋转矩阵)以最大化能量压缩,从而将信号能量集中于一个声道内。
声道对集合生成模块输出的目标声道对集合中的各个声道对分别被输入一个立体声处理单元,例如,(CH1,CH2)输入立体声处理单元1,(CH3,CH4)输入立体声处理单元2,…,(CHi-1,CHi)输入立体声处理单元m。立体声处理单元对输入的两个声道信号处理后,输出该两个声道信号对应的处理声道信号(P)以及多声道参数(SIDE_PAIR),多声道参数包括声道对索引、能量均衡边信息、立体声处理边信息。例如,立体声处理单元1对CH1和CH2处理,得到P1和P2、以及SIDE_PAIR1,立体声处理单元2对CH3和CH4处理,得到P3和P4、以及SIDE_PAIR2,…,立体声处理单元m对CHi-1和CHi处理,得到Pi-1和Pi、以及SIDE_PAIRm。
声道编码模块使用单声道编码单元(或者单声道声道盒、单声道工具)对多声道处理模块输出的处理声道信号进行编码输出对应的编码声道信号(E)。单声道编码单元对声道信号编码过程中,对具有较高能量(或较高振幅)的声道信号分配较多的比特数,对具有较少能量(或较少振幅)的声道信号分配较少的比特数。可选的,声道编码模块也可以采用立体声编码单元,例如参数立体声编码器或损耗立体声编码器对多声道处理模块输出的处理声道信号进行编码。例如,P1、P2、P3、P4、…、Pi1、Pi分别通过一个单声道编码单元进行编码得到E1、E2、E3、E4、…、Ei1、Ei。
需要说明的是,在声道对集合生成模块中未组对的声道信号(例如CHj)不需要经过多声道处理模块中的立体声处理单元处理,可以直接输入声道编码模块中的一个单声道编码单元得到Ej。
码流复用接口产生编码多声道信号,该编码多声道信号包括声道编码模块输出的编码声道信号和多声道处理模块输出的多声道参数。例如,编码多声道信号包括E1、E2、E3、E4、…、Ei1、Ei,以及SIDE_PAIR1,SIDE_PAIR2,…,SIDE_PAIRm。可选的,码流复用接口可以将编码多声道信号处理成串行信号或串行比特流。
如上所述,本申请提供的获取目标声道对集合的处理流程,可以由图4所示的编码装置中的声道对集合生成模块实现。
实施例一
以5.1声道为例,该5.1声道包括中央声道(C)、前置左声道(left,L)、前置右声道(right,R)、后置左环绕声道(left surround,LS)、后置右环绕声道(right surround,RS)以及0.1声道低频效果(low frequency effects,LFE)。针对这几个声道,声道对集合生成模块可以使用多声道掩码去掉不需要经过多声道处理的声道,以提升编码效率,5.1声道中可以去掉LFE声道,因此输入声道对集合生成模块的声道信号包括C、L、R、LS和RS。获取目标声道对集合的方法可以包括以下步骤:
(1)计算五个声道信号中任意两个之间的相关值。
本申请可以采用以下公式计算两个声道信号(例如声道信号ch1和声道信号ch2)之间的相关值:
Figure PCTCN2021106101-appb-000006
其中,corr_norm(ch1,ch2)表示声道信号ch1和声道信号ch2之间归一化的相关值,spec_ch1(i)表示声道信号ch1的第i个频点的频域系数,spec_ch2(i)是声道信号ch2的第i个频点的频域系数,N表示一个音频帧的总频点数。
本实施例中5.1声道参与组对的声道信号有五个,因此得到的相关值集合可以最多包括
Figure PCTCN2021106101-appb-000007
个声道对的相关值。表1示出了5.1声道的相关值集合的一个示例。
表1
声道信号\相关值 R C LS RS
L 0.36 0.47 0.39 0.27
R   0.57 0.22 0.08
C     0.31 0.26
LS       0.42
组对阈值设置为0.3,只有相关值大于0.3的两个声道信号才可以组对,因此将表1中小于组对阈值的相关值删除,可以得到表1a,这样在迭代处理的过程中可以不考虑相关性较小的声道信号,进而减少计算量。
表1a
Figure PCTCN2021106101-appb-000008
N设置为最大声道对数加一,即
Figure PCTCN2021106101-appb-000009
从表1a中选取N=3个最大的相关值,例如从大到小依次为0.57(R,C)、0.47(L,C)、0.42(LS,RS),这三个相关值均大于组对阈值0.3。
(2)第一个迭代处理流程
(R,C)是加入第一声道对集合的第一个声道对,从表1a中将包含了R和/或C的声道对的相关值删除,得到表1b。
表1b
Figure PCTCN2021106101-appb-000010
表1b中最大的相关值为0.42(LS,RS),因此将LS和RS组成第二个声道对加入第一声道对集合。此时五个声道信号只剩下一个声道信号L,无法继续组对,因此最终的第一声道对集合包括两个声道对(R,C)和(LS,RS)。
计算第一声道对集合的相关值之和S(1)=0.57+0.42=0.99。
(3)第二个迭代处理流程
(L,C)是加入第二声道对集合的第一个声道对,从表1a中将包含了L和/或C的声道对的相关值删除,得到表1c。
表1c
Figure PCTCN2021106101-appb-000011
表1c中最大的相关值为0.42(LS,RS),因此将LS和RS组成第二个声道对加入第二声道对集合。此时五个声道信号只剩下一个声道信号R,无法继续组对,因此最终的第二声道对集合包括两个声道对(L,C)和(LS,RS)。
计算第一声道对集合的相关值之和S(2)=0.47+0.42=0.89。
(4)第三个迭代处理流程
(LS,RS)是加入第三声道对集合的第一个声道对,从表1a中将包含了LS和/或RS的声道对的相关值删除,得到表1d。
表1d
Figure PCTCN2021106101-appb-000012
表1d中最大的相关值为0.57(R,C),因此将R和C组成第二个声道对加入第三声道对集合。此时五个声道信号只剩下一个声道信号L,无法继续组对,因此最终的第三声道对集合包括两个声道对(LS,RS)和(R,C)。
计算第一声道对集合的相关值之和S(3)=0.42+0.57=0.99。
(5)获取目标声道对集合
S(1)、S(2)和S(3)中最大的是S(1)和S(3),其所对应的两个声道对集合包含的声道对是相同的,因此将S(1)(或S(3))对应的声道对集合作为目标声道对集合,即本实施例中5.1声道可以得到的声道对包括(L,C)和(LS,RS)。目标声道对集合可以用索引表示,可以对表1中的所有相关值对应的声道对设置索引值,当确定目标声道对集合后,可以将目标声道对集合中的声道对用对应的索引值表示,以节省码流中的比特数。
实施例二
以7.1声道为例,该7.1声道包括C、L、R、LS、RS,左后置声道(left back,LB)、右后置声道(right back,RB)以及LFE。针对这几个声道,声道对集合生成模块可以使用多声道掩码去掉不需要经过多声道处理的声道,以提升编码效率,7.1声道中可以去掉LFE声道,因此输入声道对集合生成模块的声道信号包括C、L、R、LS、RS、LB和RB。获取目标声道对集合的方法可以包括以下步骤:
(1)计算七个声道信号中任意两个之间的相关值。
本实施例也可以采用上述实施例一的公式计算两个声道信号之间的相关值。
本实施例中7.1声道参与组对的声道信号有七个,因此得到的相关值集合可以最多包括
Figure PCTCN2021106101-appb-000013
个声道对的相关值。表2示出了7.1声道的相关值集合的一个示例。
表2
声道信号\相关值 R C LS RS LB RB
L 0.36 0.47 0.39 0.27 0.43 0.24
R   0.57 0.22 0.08 0.19 0.21
C     0.31 0.26 0.36 0.07
LS       0.42 0.67 0.03
RS         0.64 0.07
LB           0.19
组对阈值设置为0.3,即只有相关值大于0.3的两个声道信号才可以组对,因此将表2中小于组对阈值的相关值删除,可以得到表2a,这样在迭代处理的过程中可以不考虑相关性较小的声道信号,进而减少计算量。
表2a
Figure PCTCN2021106101-appb-000014
Figure PCTCN2021106101-appb-000015
N设置为最大声道对数加一,即
Figure PCTCN2021106101-appb-000016
从表2a中选取N=4个最大的相关值,例如从大到小依次为0.67(LS,LB)、0.64(RS,LB)、0.57(R,C)、0.47(L,C),这四个相关值均大于组对阈值0.3。
(2)第一个迭代处理流程
(LS,LB)是加入第一声道对集合的第一个声道对,从表2a中将包含了LS和/或LB的声道对的相关值删除,得到表2b。
表2b
Figure PCTCN2021106101-appb-000017
表2b中最大的相关值为0.57(R,C),因此将R和C组成第二个声道对加入第一声道对集合。从表2b中将包含了R和/或C的声道对的相关值删除,得到表2c。
表2c
Figure PCTCN2021106101-appb-000018
表2c中已无可用的相关值,因此最终的第一声道对集合包括两个声道对(LS,LB)和(R,C)。
计算第一声道对集合的相关值之和S(1)=0.67+0.57=1.24。
(3)第二个迭代处理流程
(RS,LB)是加入第二声道对集合的第一个声道对,从表2a中将包含了RS和/或LB的声道对的相关值删除,得到表2d。
表2d
Figure PCTCN2021106101-appb-000019
Figure PCTCN2021106101-appb-000020
表2d中最大的相关值为0.57(R,C),因此将R和C组成第二个声道对加入第二声道对集合。从表2d中将包含了R和/或C的声道对的相关值删除,得到表2e。
表2e
Figure PCTCN2021106101-appb-000021
表2e中最大的相关值为0.39(L,LS),因此将L和LS组成第三个声道对加入第二声道对集合。从表2e中将包含了L和/或LS的声道对的相关值删除,得到表2f。
表2f
Figure PCTCN2021106101-appb-000022
表2f中已无可用的相关值,因此最终的第一声道对集合包括三个声道对(RS,LB)、(R,C)和(L,LS)。
计算第二声道对集合的相关值之和S(2)=0.64+0.57+0.39=1.6。
(4)第三个迭代处理流程
(R,C)是加入第三声道对集合的第一个声道对,从表2a中将包含了R和/或C的声道对的相关值删除,得到表2g。
表2g
Figure PCTCN2021106101-appb-000023
Figure PCTCN2021106101-appb-000024
表2g中最大的相关值为0.67(LS,LB),因此将LS和LB组成第二个声道对加入第三声道对集合。从表2g中将包含了LS和/或LB的声道对的相关值删除,得到表2h。
表2h
Figure PCTCN2021106101-appb-000025
表2h中已无可用的相关值,因此最终的第一声道对集合包括两个声道对(R,C)和(LS,LB)。
计算第二声道对集合的相关值之和S(3)=0.57+0.67=1.24。
(5)第四个迭代处理流程
(L,C)是加入第四声道对集合的第一个声道对,从表2a中将包含了L和/或C的声道对的相关值删除,得到表2i。
表2i
Figure PCTCN2021106101-appb-000026
表2i中最大的相关值为0.67(LS,LB),因此将LS和LB组成第二个声道对加入第四声道对集合。从表2i中将包含了LS和/或LB的声道对的相关值删除,得到表2j。
表2j
Figure PCTCN2021106101-appb-000027
Figure PCTCN2021106101-appb-000028
表2j中已无可用的相关值,因此最终的第一声道对集合包括两个声道对(L,C)和(LS,LB)。
计算第二声道对集合的相关值之和S(4)=0.47+0.67=1.14。
(6)获取目标声道对集合
S(1)、S(2)、S(3)和S(4)中最大的是S(2),因此将S(2)对应的声道对集合作为目标声道对集合,即本实施例中7.1声道可以得到的声道对包括(RS,LB)、(R,C)和(L,LS)。
实施例二相较于实施例一,多了一次迭代处理过程,目标声道对集合中包括的声道对的个数也多一个,这均与参与组对的声道信号的数量有关。
图5是本申请提供的多声道音频信号的编码方法的一个示例性的实施例的流程图。该过程500可由音频译码系统10中的源设备12或音频译码设备200执行。过程500描述为一系列的步骤或操作,应当理解的是,过程500可以以各种顺序执行和/或同时发生,不限于图5所示的执行顺序。如图5所示,该方法包括:
步骤501、获取待编码的第一音频帧。
步骤502、获取相关值集合。
本实施例的步骤501和502可参考上述步骤301和302,此处不再赘述。
步骤503、根据多个声道对获取多个声道对集合。
相关值集合包括了第一音频帧的至少五个声道信号的多个声道对的相关值,将该多个声道对进行有规则的组合(即同一声道对集合中的多个声道对之间不能包含相同的声道信号),可以得到该至少五个声道信号对应的多个声道对集合。
在一种可能的实现方式中,当声道信号的个数为奇数时,可以采用以下公式计算所有声道对集合的个数:
Figure PCTCN2021106101-appb-000029
在一种可能的实现方式中,当声道信号的个数为偶数时,可以采用以下公式计算所有声道对集合的个数:
Figure PCTCN2021106101-appb-000030
其中,Pair_num表示所有声道对集合的个数,CH表示第一音频帧里参与多声道处理的声道信号的个数,是经过多声道掩码筛选后的结果。
可选的,为了减少计算量,得到相关值集合之后,可以根据多个声道对中除非相关声道对外的其他声道对获取多个声道对集合,该非相关声道对的相关值小于组对阈值,这样在获取声道对集合时可以减少参与计算的声道对的个数,进而减少声道对集合的个数,在后续步骤也可以减少相关值之和的计算量。
可选的,为了减少计算量,得到相关值集合之后,可以将与其他声道信号的相关值均小于组对阈值的声道信号删除,即这样的声道信号不考虑组对,在获取声道对集合时可以 减少参与计算的声道对的个数,进而减少声道对集合的个数,在后续步骤也可以减少相关值之和的计算量。
步骤504、根据相关值集合获取多个声道对集合中每一个声道对集合包含的所有声道对的相关值之和。
针对每一个声道对集合,计算该声道对集合中包含的所有声道对的相关值之和。
步骤505、确定目标声道对集合。
步骤506、根据目标声道对集合对第一音频帧进行编码。
本实施例的步骤505和506可参考上述步骤305和306,此处不再赘述。
本实施例通过尽可能多的获取多个声道对集合的相关值之和,进而将最大相关值之和对应的声道对集合确定为目标声道对集合,可以实现目标声道对集合所包含的所有声道对的相关值之和最大,并尽可能增加组对的声道对的个数,减少声道信号之间的冗余,提升音频的编码效率。
以下通过一个具体的实施例对图5所示方法实施例中如何获取目标声道对集合的过程进行描述。该过程仍然由图4所示的编码装置中的声道对集合生成模块实现。
实施例三
以5.1声道为例,该5.1声道包括C、L、R、LS、RS以及LFE。针对这几个声道,声道对集合生成模块可以使用多声道掩码去掉不需要经过多声道处理的声道,以提升编码效率,5.1声道中可以去掉LFE声道,因此输入声道对集合生成模块的声道信号包括C、L、R、LS和RS。获取目标声道对集合的方法可以包括以下步骤:
(1)计算五个声道信号中任意两个之间的相关值。
本实施例也可以采用上述实施例一的公式计算两个声道信号之间的相关值。
本实施例中5.1声道参与组对的声道信号有五个,因此得到的相关值集合可以最多包括
Figure PCTCN2021106101-appb-000031
个声道对的相关值,如表1所示。
(2)计算五个声道信号对应的所有声道对集合的相关值之和。
如表1所示,五个声道信号可以得到10个相关值,相应的,也就可以得到10个声道对,进而该10个声道对可以得到最多
Figure PCTCN2021106101-appb-000032
个声道对集合。例如,{(L,R),(LS,RS)},{(L,R),(C,RS)},{(L,R),(LS,C)},……。
针对声道对集合S(i),计算S(i)中包括的所有声道对的相关值之和,1≤i≤15。例如,S(1)=corr(L,R)+corr(LS,RS),S(2)=corr(L,R)+corr(C,RS),S(3)=corr(L,R)+corr(LS,C),……。
可选的,当计算相关值之和时,若某一声道对的相关值小于组对阈值,可以将该声道对的相关值设置为0。
可选的,为了减少计算量,在获取声道对集合之前,可以将相关值小于组对阈值的声道对排除掉,这样在获取声道对集合时可以减少声道对的数量,进而减少声道对集合的数量。
图6是本申请提供的多声道音频信号的编码方法的一个示例性的实施例的流程图。该过程600可由音频译码系统10中的源设备12或音频译码设备200执行。过程600描述为一系列的步骤或操作,应当理解的是,过程600可以以各种顺序执行和/或同时发生,不限于图6所示的执行顺序。如图6所示,该方法包括:
步骤601、获取待编码的第一音频帧。
步骤601可参考上述步骤301,此处不再赘述。
步骤602、获取第一音频帧的相关值集合。
第一音频帧的相关值集合包括多个声道对各自的相关值,一个声道对包括至少五个声道信号中的两个声道信号,一个声道对的相关值用于表示一个声道对的两个声道信号之间的相关性。
步骤603、获取第二音频帧的相关值集合。
第二音频帧的相关值集合包括第二音频帧的多个声道对各自的相关值,一个声道对包括第二音频帧的至少五个声道信号中的两个声道信号,一个声道对的相关值用于表示一个声道对的两个声道信号之间的相关性,第二音频帧是第一音频帧的上一帧。
本实施例与上述步骤302的区别在于,本实施例除了获取第一音频帧的相关值集合,还需要获取第一音频帧的上一帧(即第二音频帧)的相关值集合。
获取第一音频帧的相关值集合的方法可参考上述步骤302,此处不再赘述。
由于第二音频帧的编码在第一音频帧的编码之前,因此当处理到第一音频帧时,编码装置已经获取了对第二音频帧编码时的相关信息,包括第二音频帧的相关值集合,因此本实施例获取第二音频帧的相关值集合可以是直接从缓存或内存中读取即可,不需要再次计算获取第二音频帧的相关值集合。
步骤604、根据第一音频帧的相关值集合和第二音频帧的相关值集合判断是否需要重新获取第一音频帧的目标声道对集合。
本实施例可以通过计算第一音频帧的相关值集合和第二音频帧的相关值集合的差值之和作为判断依据,即计算第一音频帧的相关值集合和第二音频帧的相关值集合中对应于同一声道对的相关值之差的绝对值,计算多个声道对分别对应的绝对值之和。当绝对值之和小于变更阈值时,确定不需要重新获取第一音频帧的目标声道对集合;当绝对值之和大于或等于变更阈值时,确定需要重新获取第一音频帧的目标声道对集合。
对应于相同的声道对,分别计算其相关值差值,然后计算所有声道对的差值的绝对值之和,这样可以得到第一音频帧相对于第二音频帧,各声道信号之间的相关值的变化是否超过了变更阈值,如果没有超过,说明第二音频帧到第一音频帧的变化不大,可以不需要对第一音频帧重新组建目标声道对集合,减少了计算量,提高编码效率;如果超过,说明第二音频帧到第一音频帧的变化较大,需要重新获取第一音频帧的目标声道对集合。
步骤605、若需要重新获取第一音频帧的目标声道对集合,则采用图3或图5所示实施例的方法获取第一音频帧的目标声道对集合,并根据目标声道对集合对第一音频帧进行编码。
本实施例在确定需要重新获取第一音频帧的目标声道对集合,可以采用图3或图5所示实施例中的方法获取第一音频帧的相关值集合,此处不再赘述。
步骤606、若不需要重新获取第一音频帧的目标声道对集合,则将第二音频帧的目标声道对集合确定为第一音频帧的目标声道对集合,并根据目标声道对集合对第一音频帧进行编码。
本实施例在确定不需要重新获取第一音频帧的目标声道对集合,可以直接将第二音频帧的目标声道对集合作为第一音频帧的目标声道对集合,从而减少计算量,提高编码效率。
本实施例通过获取当前音频帧的相关值集合和上一音频帧的相关值集合的差值之和,从而确定是否需要重新获取当前帧的目标声道对集合,可以在音频变化较小的情况下,大大减少计算量,提高编码效率,而即使音频变化较大,需要重新获取目标声道对集合,仍可以尽可能多的获取多个声道对集合的相关值之和,进而将最大相关值之和对应的声道对集合确定为目标声道对集合,可以实现目标声道对集合所包含的所有声道对的相关值之和最大,并尽可能增加组对的声道对的个数,减少声道信号之间的冗余,提升音频的编码效率。
以下通过一个具体的实施例对图6所示方法实施例中如何获取目标声道对集合的过程进行描述。该过程仍然由图4所示的编码装置中的声道对集合生成模块实现。
实施例四
以5.1声道为例,该5.1声道包括C、L、R、LS、RS以及LFE。针对这几个声道,声道对集合生成模块可以使用多声道掩码去掉不需要经过多声道处理的声道,以提升编码效率,5.1声道中可以去掉LFE声道,因此输入声道对集合生成模块的声道信号包括C、L、R、LS和RS。获取目标声道对集合的方法可以包括以下步骤:
(1)计算五个声道信号中任意两个之间的相关值。
本实施例也可以采用上述实施例一的公式计算两个声道信号之间的相关值。
本实施例中5.1声道参与组对的声道信号有五个,因此得到的相关值集合可以最多包括
Figure PCTCN2021106101-appb-000033
个声道对的相关值,如表1所示。
(2)计算第一音频帧的相关值集合和第二音频帧的相关值集合的差值之和。
本实施例将第一音频帧的相关值集合和第二音频帧的相关值集合均以矩阵的方式表示,分别得到矩阵Matrix1和Matrix2,矩阵中的每个元素的取值对应相关值集合中的一个相关值,可以通过以下公式计算差值之和:
Figure PCTCN2021106101-appb-000034
其中,D表示第一音频帧的相关值集合和第二音频帧的相关值集合的差值之和,Matrix1(i)表示第一音频帧的相关值集合对应的矩阵中的第i个元素值,Matrix2(i)表示第二音频帧的相关值集合对应的矩阵中的第i个元素值。
(3)根据相关值之和D确定是否需要重新获取第一音频帧的目标声道对集合。
本实施例设置一个变更阈值,通过该阈值界定是否需要重新获取第一音频帧的目标声道对集合。可选的,本实施例还可以设置一个标识keepFlag,当keepFlag=1时,表示第一音频帧可以保留上一帧的目标声道对集合,即不需要重新获取第一音频帧的目标声道对集合;当keepFlag=0时,表示第一音频帧不能保留上一帧的目标声道对集合,即需要重新获取第一音频帧的目标声道对集合。
基于上述设置,当D<变更阈值时,keepFlag=1;当D≥变更阈值时,keepFlag=0。
(4)获取第一音频帧的目标声道对集合
根据上述标识keepFlag的取值,编码装置可以获取第一音频帧的目标声道对集合,即当keepFlag=1时,编码装置直接将第二音频帧的目标声道对集合作为第一音频帧的目标声道对集合;当keepFlag=0时,编码装置可以采用图3或图5所示实施例的方法获取第 一音频帧的目标声道对集合,此处不再赘述。
图7是本申请提供的多声道音频信号的编码方法的一个示例性的实施例的流程图。该过程700可由音频译码系统10中的源设备12或音频译码设备200执行。过程700描述为一系列的步骤或操作,应当理解的是,过程700可以以各种顺序执行和/或同时发生,不限于图7所示的执行顺序。如图7所示,该方法包括:
步骤701、获取待编码的第一音频帧,第一音频帧包括K个声道信号。
步骤701可参考上述步骤301,此处不再赘述。
步骤702、当K大于声道信号数量阈值时,采用图3所示实施例的方法对第一音频帧进行编码。
步骤703、当K小于或等于声道信号数量阈值时,采用图5所示实施例的方法对第一音频帧进行编码。
本实施例与上述图3或图5所示实施例的区别在于,本实施例将图3和图5的方法进行融合,即根据第一音频帧包含的声道信号的个数来确定对第一音频帧采用哪一种方法获取其目标声道对集合。当第一音频帧包含的声道信号的个数较多时,如果采用第二方面的方法,需要穷举所有目标声道对集合,会增加计算量,因此此时采用第一方面的方法会减少很多的计算量。而当第一音频帧包含的声道信号的个数较少时,采用第二方面的方法可以获取到所有声道对集合的相关值之和,确保最终选取的目标声道对集合一定是最符合第一音频帧的特性的最优结果。
图8是本申请提供的多声道音频信号的解码方法所应用的解码装置的一个示例性的结构图,该解码装置可以是音频译码系统10中的目的设备14的解码器30,也可以是音频译码设备200中的译码模块270。该解码装置可以包括码流解复用接口、声道解码模块和多声道处理模块,其中,
码流解复用接口接收来自编码装置的编码多声道信号(例如串行比特流bitstream),解复用后得到编码声道信号(E)和多声道参数(SIDE_PAIR)。例如,E1、E2、E3、E4、…、Ei1、Ei,以及SIDE_PAIR1,SIDE_PAIR2,…,SIDE_PAIRm。
声道解码模块使用单声道解码单元(或者单声道声道盒、单声道工具)对码流解复用接口输出的编码声道信号进行解码输出解码声道信号(D)。例如,E1、E2、E3、E4、…、Ei1、Ei分别通过一个单声道解码单元进行解码得到E1解码得D1、D2、D3、D4、…、Di1、Di。
多声道处理模块包括多个立体声处理单元,立体声处理单元可以采用基于预测的或者基于KLT的处理,即输入的两个声道信号被反旋转(例如经由2×2旋转矩阵),从而将信号变换到原始信号方向。
声道解码模块输出的解码声道信号藉由多声道参数可以识别哪两个解码声道信号组对,将组对的解码声道信号输入立体声处理单元,立体声处理单元对输入的两个解码声道信号处理后,输出该两个解码声道信号对应的声道信号(CH)。例如,立体声处理单元1根据SIDE_PAIR1对D1和D2处理,得到CH1和CH2,立体声处理单元2根据SIDE_PAIR2对D3和D4处理,得到CH3和CH4,…,立体声处理单元m根据SIDE_PAIRm对Di-1和Di处理,得到CHi-1和CHi。
需要说明的是,针对未组对的声道信号(例如CHj)不需要经过多声道处理模块中的 立体声处理单元处理,可以解码后直接输出。
图9为本申请编码装置实施例的结构示意图,如图9所示,该装置可以应用于上述实施例中的源设备12或音频译码设备200。本实施例的编码装置可以包括:获取模块901、编码模块902和确定模块903。
在一种可能的实现方式中,获取模块901,用于获取待编码的第一音频帧,所述第一音频帧包括至少五个声道信号;获取相关值集合,所述相关值集合包括多个声道对各自的相关值,一个声道对包括所述至少五个声道信号中的两个声道信号,所述一个声道对的相关值用于表示所述一个声道对的两个声道信号之间的相关性;从所述相关值集合中选取M个相关值,所述M个相关值均大于所述相关值集合中除所述M个相关值外的其他相关值,所述M个相关值均大于或等于组对阈值,M为小于或等于设定值的正整数;获取M个声道对集合,每个所述声道对集合至少包括所述M个相关值对应的M个声道对的其中之一,且当所述声道对集合包括两个以上声道对时,所述两个以上声道对不包含相同的声道信号;确定模块903,用于从所述M个声道对集合中确定目标声道对集合,所述目标声道对集合中的所有声道对的相关值之和是所述M个声道对集合中最大的;编码模块902,用于根据所述目标声道对集合对所述第一音频帧进行编码。
在一种可能的实现方式中,所述M个声道对集合包括第一声道对集合;所述获取模块901,具体用于将所述M个声道对中的第一声道对加入所述第一声道对集合,所述第一声道对为所述M个声道对中的任意一个;当所述多个声道对中除关联声道对外的其他声道对中包括相关值大于所述组对阈值的声道对时,从所述其他声道对中选取相关值最大的一个声道对加入所述第一声道对集合,所述关联声道对包括已加入所述第一声道对集合的声道对所包括的声道信号中的任意一个。
在一种可能的实现方式中,所述获取模块901,具体用于从所述相关值集合中选取N个相关值,所述N个相关值均大于所述相关值集合中除所述N个相关值外的其他相关值,N为所述设定值;从所述N个相关值中选取大于或等于所述组对阈值的相关值,所述大于或等于所述组对阈值的相关值的个数为M。
在一种可能的实现方式中,所述相关值为经归一化处理的值。
在一种可能的实现方式中,当所述一个声道对的相关值小于所述组对阈值时,所述一个声道对的相关值设置为0。
在一种可能的实现方式中,获取模块901,用于获取待编码的第一音频帧,所述第一音频帧包括至少五个声道信号;获取相关值集合,所述相关值集合包括多个声道对各自的相关值,一个声道对包括所述至少五个声道信号中的两个声道信号,所述一个声道对的相关值用于表示所述一个声道对的两个声道信号之间的相关性;根据所述多个声道对获取多个声道对集合,当所述声道对集合包括两个以上声道对时,所述两个以上声道对不包含相同的声道信号;根据所述相关值集合获取所述多个声道对集合中每一个声道对集合包含的所有声道对的相关值之和;确定模块903,用于确定目标声道对集合,所述目标声道对集合中的所有声道对的相关值之和是所述多个声道对集合中最大的;编码模块902,用于根据所述目标声道对集合对所述第一音频帧进行编码。
在一种可能的实现方式中,所述获取模块901,具体用于根据所述多个声道对中除非相关声道对外的其他声道对获取所述多个声道对集合,所述非相关声道对的相关值小于组 对阈值。
在一种可能的实现方式中,获取模块901,用于获取待编码的第一音频帧,所述第一音频帧包括至少五个声道信号;获取所述第一音频帧的相关值集合,所述第一音频帧的相关值集合包括多个声道对各自的相关值,一个声道对包括所述至少五个声道信号中的两个声道信号,所述一个声道对的相关值用于表示所述一个声道对的两个声道信号之间的相关性;获取第二音频帧的相关值集合,所述第二音频帧的相关值集合包括所述第二音频帧的多个声道对各自的相关值,一个声道对包括所述第二音频帧的至少五个声道信号中的两个声道信号,所述一个声道对的相关值用于表示所述一个声道对的两个声道信号之间的相关性,所述第二音频帧是所述第一音频帧的上一帧;编码模块902,用于根据所述第一音频帧的相关值集合和所述第二音频帧的相关值集合判断是否需要重新获取所述第一音频帧的目标声道对集合;若需要重新获取所述第一音频帧的目标声道对集合,则执行图3或图5所示实施例的方法获取所述第一音频帧的目标声道对集合,并根据所述目标声道对集合对所述第一音频帧进行编码;若不需要重新获取所述第一音频帧的目标声道对集合,则将所述第二音频帧的目标声道对集合确定为所述第一音频帧的目标声道对集合,并根据所述目标声道对集合对所述第一音频帧进行编码。
在一种可能的实现方式中,所述编码模块902,具体用于计算所述第一音频帧的相关值集合和所述第二音频帧的相关值集合中对应于同一声道对的相关值之差的绝对值;计算多个所述声道对分别对应的所述绝对值之和;当所述绝对值之和小于变更阈值时,确定不需要重新获取所述第一音频帧的目标声道对集合;当所述绝对值之和大于或等于所述变更阈值时,确定需要重新获取所述第一音频帧的目标声道对集合。
在一种可能的实现方式中,获取模块,用于获取待编码的第一音频帧,所述第一音频帧包括K个声道信号,K为大于或等于5的整数;编码模块,用于当K大于声道信号数量阈值时,执行图3所示实施例的方法对所述第一音频帧进行编码;当K小于或等于声道信号数量阈值时,执行图5所示实施例的方法对所述第一音频帧进行编码。
本实施例的装置,可以用于执行图3、图5、图6或图7所示方法实施例的技术方案,其实现原理和技术效果类似,此处不再赘述。
图10为本申请设备实施例的结构示意图,如图10所示,该设备可以是上述实施例中的编码设备。本实施例的设备可以包括:处理器1001和存储器1002,存储器1002,用于存储一个或多个程序;当所述一个或多个程序被所述处理器1001执行,使得所述处理器1001实现如图3、图5、图6或图7所示方法实施例的技术方案。
在实现过程中,上述方法实施例的各步骤可以通过处理器中的硬件的集成逻辑电路或者软件形式的指令完成。处理器可以是通用处理器、数字信号处理器(digital signal processor,DSP)、特定应用集成电路(application-specific integrated circuit,ASIC)、现场可编程门阵列(field programmable gate array,FPGA)或其他可编程逻辑器件、分立门或者晶体管逻辑器件、分立硬件组件。通用处理器可以是微处理器或者该处理器也可以是任何常规的处理器等。本申请公开的方法的步骤可以直接体现为硬件编码处理器执行完成,或者用编码处理器中的硬件及软件模块组合执行完成。软件模块可以位于随机存储器,闪存、只读存储器,可编程只读存储器或者电可擦写可编程存储器、寄存器等本领域成熟的存储介质中。该存储介质位于存储器,处理器读取存储器中的信息,结合其硬件完成上述方法的步骤。
上述各实施例中提及的存储器可以是易失性存储器或非易失性存储器,或可包括易失性和非易失性存储器两者。其中,非易失性存储器可以是只读存储器(read-only memory,ROM)、可编程只读存储器(programmable ROM,PROM)、可擦除可编程只读存储器(erasable PROM,EPROM)、电可擦除可编程只读存储器(electrically EPROM,EEPROM)或闪存。易失性存储器可以是随机存取存储器(random access memory,RAM),其用作外部高速缓存。通过示例性但不是限制性说明,许多形式的RAM可用,例如静态随机存取存储器(static RAM,SRAM)、动态随机存取存储器(dynamic RAM,DRAM)、同步动态随机存取存储器(synchronous DRAM,SDRAM)、双倍数据速率同步动态随机存取存储器(double data rate SDRAM,DDR SDRAM)、增强型同步动态随机存取存储器(enhanced SDRAM,ESDRAM)、同步连接动态随机存取存储器(synchlink DRAM,SLDRAM)和直接内存总线随机存取存储器(direct rambus RAM,DR RAM)。应注意,本文描述的系统和方法的存储器旨在包括但不限于这些和任意其它适合类型的存储器。
本领域普通技术人员可以意识到,结合本文中所公开的实施例描述的各示例的单元及算法步骤,能够以电子硬件、或者计算机软件和电子硬件的结合来实现。这些功能究竟以硬件还是软件方式来执行,取决于技术方案的特定应用和设计约束条件。专业技术人员可以对每个特定的应用来使用不同方法来实现所描述的功能,但是这种实现不应认为超出本申请的范围。
所属领域的技术人员可以清楚地了解到,为描述的方便和简洁,上述描述的系统、装置和单元的具体工作过程,可以参考前述方法实施例中的对应过程,在此不再赘述。
在本申请所提供的几个实施例中,应该理解到,所揭露的系统、装置和方法,可以通过其它的方式实现。例如,以上所描述的装置实施例仅仅是示意性的,例如,所述单元的划分,仅仅为一种逻辑功能划分,实际实现时可以有另外的划分方式,例如多个单元或组件可以结合或者可以集成到另一个系统,或一些特征可以忽略,或不执行。另一点,所显示或讨论的相互之间的耦合或直接耦合或通信连接可以是通过一些接口,装置或单元的间接耦合或通信连接,可以是电性,机械或其它的形式。
所述作为分离部件说明的单元可以是或者也可以不是物理上分开的,作为单元显示的部件可以是或者也可以不是物理单元,即可以位于一个地方,或者也可以分布到多个网络单元上。可以根据实际的需要选择其中的部分或者全部单元来实现本实施例方案的目的。
另外,在本申请各个实施例中的各功能单元可以集成在一个处理单元中,也可以是各个单元单独物理存在,也可以两个或两个以上单元集成在一个单元中。
所述功能如果以软件功能单元的形式实现并作为独立的产品销售或使用时,可以存储在一个计算机可读取存储介质中。基于这样的理解,本申请的技术方案本质上或者说对现有技术做出贡献的部分或者该技术方案的部分可以以软件产品的形式体现出来,该计算机软件产品存储在一个存储介质中,包括若干指令用以使得一台计算机设备(个人计算机,服务器,或者网络设备等)执行本申请各个实施例所述方法的全部或部分步骤。而前述的存储介质包括:U盘、移动硬盘、只读存储器(read-only memory,ROM)、随机存取存储器(random access memory,RAM)、磁碟或者光盘等各种可以存储程序代码的介质。
以上所述,仅为本申请的具体实施方式,但本申请的保护范围并不局限于此,任何熟悉本技术领域的技术人员在本申请揭露的技术范围内,可轻易想到变化或替换,都应涵盖 在本申请的保护范围之内。因此,本申请的保护范围应以所述权利要求的保护范围为准。

Claims (27)

  1. 一种多声道音频信号的编码方法,其特征在于,包括:
    获取待编码的第一音频帧,所述第一音频帧包括至少五个声道信号;
    获取相关值集合,所述相关值集合包括多个声道对各自的相关值,一个声道对包括所述至少五个声道信号中的两个声道信号,所述一个声道对的相关值用于表示所述一个声道对的两个声道信号之间的相关性;
    从所述相关值集合中选取M个相关值,所述M个相关值均大于所述相关值集合中除所述M个相关值外的其他相关值,所述M个相关值均大于或等于组对阈值,M为小于或等于设定值的正整数;
    获取M个声道对集合,每个所述声道对集合包括与所述M个相关值对应的一个或多个声道对,且当所述声道对集合包括两个以上声道对时,所述两个以上声道对不包含相同的声道信号;
    从所述M个声道对集合中确定目标声道对集合,所述目标声道对集合中的所有声道对的相关值之和是所述M个声道对集合中最大的;
    根据所述目标声道对集合对所述第一音频帧进行编码。
  2. 根据权利要求1所述的方法,其特征在于,所述M个声道对集合包括第一声道对集合,所述获取M个声道对集合包括获取所述第一声道对集合;
    所述获取所述第一声道对集合,包括:
    将所述M个声道对中的第一声道对加入所述第一声道对集合,所述第一声道对为所述M个声道对中的任意一个;
    当所述多个声道对中除关联声道对外的其他声道对中包括相关值大于所述组对阈值的声道对时,从所述其他声道对中选取相关值最大的一个声道对加入所述第一声道对集合,所述关联声道对包括已加入所述第一声道对集合的声道对所包括的声道信号中的任意一个。
  3. 根据权利要求1或2所述的方法,其特征在于,所述从所述相关值集合中选取M个相关值,包括:
    从所述相关值集合中选取N个相关值,所述N个相关值均大于所述相关值集合中除所述N个相关值外的其他相关值,N为所述设定值;
    从所述N个相关值中选取大于或等于所述组对阈值的相关值,所述大于或等于所述组对阈值的相关值的个数为M。
  4. 根据权利要求1-3中任一项所述的方法,其特征在于,所述相关值为经归一化处理的值。
  5. 根据权利要求1-4中任一项所述的方法,其特征在于,当所述一个声道对的相关值小于所述组对阈值时,所述一个声道对的相关值设置为0。
  6. 一种多声道音频信号的编码方法,其特征在于,包括:
    获取待编码的第一音频帧,所述第一音频帧包括至少五个声道信号;
    获取相关值集合,所述相关值集合包括多个声道对各自的相关值,一个声道对包括所述至少五个声道信号中的两个声道信号,所述一个声道对的相关值用于表示所述一个声道 对的两个声道信号之间的相关性;
    根据所述多个声道对获取多个声道对集合,当所述声道对集合包括两个以上声道对时,所述两个以上声道对不包含相同的声道信号;
    根据所述相关值集合获取所述多个声道对集合中每一个声道对集合包含的所有声道对的相关值之和;
    确定目标声道对集合,所述目标声道对集合中的所有声道对的相关值之和是所述多个声道对集合中最大的;
    根据所述目标声道对集合对所述第一音频帧进行编码。
  7. 根据权利要求6所述的方法,其特征在于,所述根据所述多个声道对获取多个声道对集合,包括:
    根据所述多个声道对中除非相关声道对外的其他声道对获取所述多个声道对集合,所述非相关声道对的相关值小于组对阈值。
  8. 根据权利要求6或5所述的方法,其特征在于,所述相关值为经归一化处理的值。
  9. 根据权利要求6-8中任一项所述的方法,其特征在于,当所述一个声道对的相关值小于组对阈值时,所述一个声道对的相关值设置为0。
  10. 一种多声道音频信号的编码方法,其特征在于,包括:
    获取待编码的第一音频帧,所述第一音频帧包括至少五个声道信号;
    获取所述第一音频帧的相关值集合,所述第一音频帧的相关值集合包括多个声道对各自的相关值,一个声道对包括所述至少五个声道信号中的两个声道信号,所述一个声道对的相关值用于表示所述一个声道对的两个声道信号之间的相关性;
    获取第二音频帧的相关值集合,所述第二音频帧的相关值集合包括所述第二音频帧的多个声道对各自的相关值,一个声道对包括所述第二音频帧的至少五个声道信号中的两个声道信号,所述一个声道对的相关值用于表示所述一个声道对的两个声道信号之间的相关性,所述第二音频帧是所述第一音频帧的上一帧;
    根据所述第一音频帧的相关值集合和所述第二音频帧的相关值集合判断是否需要重新获取所述第一音频帧的目标声道对集合;
    若需要重新获取所述第一音频帧的目标声道对集合,则采用如权利要求1-9中任一项所述的方法获取所述第一音频帧的目标声道对集合,并根据所述目标声道对集合对所述第一音频帧进行编码;
    若不需要重新获取所述第一音频帧的目标声道对集合,则将所述第二音频帧的目标声道对集合确定为所述第一音频帧的目标声道对集合,并根据所述目标声道对集合对所述第一音频帧进行编码。
  11. 根据权利要求10所述的方法,其特征在于,所述根据所述第一音频帧的相关值集合和所述第二音频帧的相关值集合判断是否需要重新获取所述第一音频帧的目标声道对集合,包括:
    计算所述第一音频帧的相关值集合和所述第二音频帧的相关值集合中对应于同一声道对的相关值之差的绝对值;
    计算多个所述声道对分别对应的所述绝对值之和;
    当所述绝对值之和小于变更阈值时,确定不需要重新获取所述第一音频帧的目标声道 对集合;
    当所述绝对值之和大于或等于所述变更阈值时,确定需要重新获取所述第一音频帧的目标声道对集合。
  12. 一种多声道音频信号的编码方法,其特征在于,包括:
    获取待编码的第一音频帧,所述第一音频帧包括K个声道信号,K为大于或等于5的整数;
    当K大于声道信号数量阈值时,采用权利要求1-5中任一项所述的方法对所述第一音频帧进行编码;
    当K小于或等于声道信号数量阈值时,采用权利要求6-9中任一项所述的方法对所述第一音频帧进行编码。
  13. 一种编码装置,其特征在于,包括:
    获取模块,用于获取待编码的第一音频帧,所述第一音频帧包括至少五个声道信号;获取相关值集合,所述相关值集合包括多个声道对各自的相关值,一个声道对包括所述至少五个声道信号中的两个声道信号,所述一个声道对的相关值用于表示所述一个声道对的两个声道信号之间的相关性;从所述相关值集合中选取M个相关值,所述M个相关值均大于所述相关值集合中除所述M个相关值外的其他相关值,所述M个相关值均大于或等于组对阈值,M为小于或等于设定值的正整数;获取M个声道对集合,每个所述声道对集合至少包括所述M个相关值对应的M个声道对的其中之一,且当所述声道对集合包括两个以上声道对时,所述两个以上声道对不包含相同的声道信号;
    确定模块,用于从所述M个声道对集合中确定目标声道对集合,所述目标声道对集合中的所有声道对的相关值之和是所述M个声道对集合中最大的;
    编码模块,用于根据所述目标声道对集合对所述第一音频帧进行编码。
  14. 根据权利要求13所述的装置,其特征在于,所述M个声道对集合包括第一声道对集合;所述获取模块,具体用于将所述M个声道对中的第一声道对加入所述第一声道对集合,所述第一声道对为所述M个声道对中的任意一个;当所述多个声道对中除关联声道对外的其他声道对中包括相关值大于所述组对阈值的声道对时,从所述其他声道对中选取相关值最大的一个声道对加入所述第一声道对集合,所述关联声道对包括已加入所述第一声道对集合的声道对所包括的声道信号中的任意一个。
  15. 根据权利要求13或14所述的装置,其特征在于,所述获取模块,具体用于从所述相关值集合中选取N个相关值,所述N个相关值均大于所述相关值集合中除所述N个相关值外的其他相关值,N为所述设定值;从所述N个相关值中选取大于或等于所述组对阈值的相关值,所述大于或等于所述组对阈值的相关值的个数为M。
  16. 根据权利要求13-15中任一项所述的装置,其特征在于,所述相关值为经归一化处理的值。
  17. 根据权利要求13-16中任一项所述的装置,其特征在于,当所述一个声道对的相关值小于所述组对阈值时,所述一个声道对的相关值设置为0。
  18. 一种编码装置,其特征在于,包括:
    获取模块,用于获取待编码的第一音频帧,所述第一音频帧包括至少五个声道信号;获取相关值集合,所述相关值集合包括多个声道对各自的相关值,一个声道对包括所述至 少五个声道信号中的两个声道信号,所述一个声道对的相关值用于表示所述一个声道对的两个声道信号之间的相关性;根据所述多个声道对获取多个声道对集合,当所述声道对集合包括两个以上声道对时,所述两个以上声道对不包含相同的声道信号;根据所述相关值集合获取所述多个声道对集合中每一个声道对集合包含的所有声道对的相关值之和;
    确定模块,用于确定目标声道对集合,所述目标声道对集合中的所有声道对的相关值之和是所述多个声道对集合中最大的;
    编码模块,用于根据所述目标声道对集合对所述第一音频帧进行编码。
  19. 根据权利要求18所述的装置,其特征在于,所述获取模块,具体用于根据所述多个声道对中除非相关声道对外的其他声道对获取所述多个声道对集合,所述非相关声道对的相关值小于组对阈值。
  20. 根据权利要求18或19所述的装置,其特征在于,所述相关值为经归一化处理的值。
  21. 根据权利要求18-20中任一项所述的装置,其特征在于,当所述一个声道对的相关值小于组对阈值时,所述一个声道对的相关值设置为0。
  22. 一种编码装置,其特征在于,包括:
    获取模块,用于获取待编码的第一音频帧,所述第一音频帧包括至少五个声道信号;获取所述第一音频帧的相关值集合,所述第一音频帧的相关值集合包括多个声道对各自的相关值,一个声道对包括所述至少五个声道信号中的两个声道信号,所述一个声道对的相关值用于表示所述一个声道对的两个声道信号之间的相关性;获取第二音频帧的相关值集合,所述第二音频帧的相关值集合包括所述第二音频帧的多个声道对各自的相关值,一个声道对包括所述第二音频帧的至少五个声道信号中的两个声道信号,所述一个声道对的相关值用于表示所述一个声道对的两个声道信号之间的相关性,所述第二音频帧是所述第一音频帧的上一帧;
    编码模块,用于根据所述第一音频帧的相关值集合和所述第二音频帧的相关值集合判断是否需要重新获取所述第一音频帧的目标声道对集合;若需要重新获取所述第一音频帧的目标声道对集合,则执行如权利要求1-9中任一项所述的方法获取所述第一音频帧的目标声道对集合,并根据所述目标声道对集合对所述第一音频帧进行编码;若不需要重新获取所述第一音频帧的目标声道对集合,则将所述第二音频帧的目标声道对集合确定为所述第一音频帧的目标声道对集合,并根据所述目标声道对集合对所述第一音频帧进行编码。
  23. 根据权利要求22所述的装置,其特征在于,所述编码模块,具体用于计算所述第一音频帧的相关值集合和所述第二音频帧的相关值集合中对应于同一声道对的相关值之差的绝对值;计算多个所述声道对分别对应的所述绝对值之和;当所述绝对值之和小于变更阈值时,确定不需要重新获取所述第一音频帧的目标声道对集合;当所述绝对值之和大于或等于所述变更阈值时,确定需要重新获取所述第一音频帧的目标声道对集合。
  24. 一种编码装置,其特征在于,包括:
    获取模块,用于获取待编码的第一音频帧,所述第一音频帧包括K个声道信号,K为大于或等于5的整数;
    编码模块,用于当K大于声道信号数量阈值时,执行如权利要求1-5中任一项所述的方法对所述第一音频帧进行编码;当K小于或等于声道信号数量阈值时,执行如权利要求 6-9中任一项所述的方法对所述第一音频帧进行编码。
  25. 一种设备,其特征在于,包括:
    一个或多个处理器;
    存储器,用于存储一个或多个程序;
    当所述一个或多个程序被所述一个或多个处理器执行,使得所述一个或多个处理器实现如权利要求1-11中任一项所述的方法。
  26. 一种计算机可读存储介质,其特征在于,包括计算机程序,所述计算机程序在计算机上被执行时,使得所述计算机执行权利要求1-11中任一项所述的方法。
  27. 一种计算机可读存储介质,其特征在于,包括根据如权利要求1-11中任一项所述的多声道音频信号的编码方法获得的编码码流。
PCT/CN2021/106101 2020-07-17 2021-07-13 多声道音频信号的编解码方法和装置 Ceased WO2022012553A1 (zh)

Priority Applications (4)

Application Number Priority Date Filing Date Title
EP21843116.1A EP4174855A4 (en) 2020-07-17 2021-07-13 METHOD AND DEVICE FOR ENCODING/DECODING A MULTI-CHANNEL AUDIO SIGNAL
JP2023502888A JP7519531B2 (ja) 2020-07-17 2021-07-13 マルチチャネルオーディオ信号符号化および復号方法および装置
KR1020237004819A KR102948552B1 (ko) 2020-07-17 2021-07-13 다중 채널 오디오 신호 인코딩 및 디코딩 방법 및 장치
US18/153,128 US12437767B2 (en) 2020-07-17 2023-01-11 Multi-channel audio signal encoding and decoding method and apparatus

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
CN202010699706.7A CN113948095B (zh) 2020-07-17 2020-07-17 多声道音频信号的编解码方法和装置
CN202010699706.7 2020-07-17

Related Child Applications (1)

Application Number Title Priority Date Filing Date
US18/153,128 Continuation US12437767B2 (en) 2020-07-17 2023-01-11 Multi-channel audio signal encoding and decoding method and apparatus

Publications (1)

Publication Number Publication Date
WO2022012553A1 true WO2022012553A1 (zh) 2022-01-20

Family

ID=79326898

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2021/106101 Ceased WO2022012553A1 (zh) 2020-07-17 2021-07-13 多声道音频信号的编解码方法和装置

Country Status (6)

Country Link
US (1) US12437767B2 (zh)
EP (1) EP4174855A4 (zh)
JP (1) JP7519531B2 (zh)
KR (1) KR102948552B1 (zh)
CN (1) CN113948095B (zh)
WO (1) WO2022012553A1 (zh)

Families Citing this family (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN116434760A (zh) * 2023-04-14 2023-07-14 北京小米移动软件有限公司 一种音频编码方法、装置、电子设备及存储介质
CN116564319B (zh) * 2023-05-10 2025-12-09 北京达佳互联信息技术有限公司 音频处理方法、装置、电子设备及存储介质
CN117730367A (zh) * 2023-10-31 2024-03-19 北京小米移动软件有限公司 分组方法、编码器、解码器以及存储介质

Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20090112606A1 (en) * 2007-10-26 2009-04-30 Microsoft Corporation Channel extension coding for multi-channel source
CN101695150A (zh) * 2009-10-12 2010-04-14 清华大学 多声道音频编码方法、编码器、解码方法和解码器
US20110123031A1 (en) * 2009-05-08 2011-05-26 Nokia Corporation Multi channel audio processing
CN104364842A (zh) * 2012-04-18 2015-02-18 诺基亚公司 立体声音频信号编码器
CN109416912A (zh) * 2016-06-30 2019-03-01 杜塞尔多夫华为技术有限公司 一种对多声道音频信号进行编码和解码的装置和方法

Family Cites Families (14)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US8359196B2 (en) * 2007-12-28 2013-01-22 Panasonic Corporation Stereo sound decoding apparatus, stereo sound encoding apparatus and lost-frame compensating method
AU2015246158B2 (en) 2009-03-17 2017-10-26 Dolby International Ab Advanced stereo coding based on a combination of adaptively selectable left/right or mid/side stereo coding and of parametric stereo coding.
US9197978B2 (en) * 2009-03-31 2015-11-24 Panasonic Intellectual Property Management Co., Ltd. Sound reproduction apparatus and sound reproduction method
US9100766B2 (en) * 2009-10-05 2015-08-04 Harman International Industries, Inc. Multichannel audio system having audio channel compensation
JP2015011076A (ja) 2013-06-26 2015-01-19 日本放送協会 音響信号符号化装置、音響信号符号化方法、および音響信号復号化装置
TWI774136B (zh) 2013-09-12 2022-08-11 瑞典商杜比國際公司 多聲道音訊系統中之解碼方法、解碼裝置、包含用於執行解碼方法的指令之非暫態電腦可讀取的媒體之電腦程式產品、包含解碼裝置的音訊系統
CN105898667A (zh) * 2014-12-22 2016-08-24 杜比实验室特许公司 从音频内容基于投影提取音频对象
EP3067885A1 (en) * 2015-03-09 2016-09-14 Fraunhofer-Gesellschaft zur Förderung der angewandten Forschung e.V. Apparatus and method for encoding or decoding a multi-channel signal
CN109389984B (zh) * 2017-08-10 2021-09-14 华为技术有限公司 时域立体声编解码方法和相关产品
CN109389987B (zh) * 2017-08-10 2022-05-10 华为技术有限公司 音频编解码模式确定方法和相关产品
CN109389985B (zh) * 2017-08-10 2021-09-14 华为技术有限公司 时域立体声编解码方法和相关产品
ES3059239T3 (en) 2018-07-04 2026-03-19 Fraunhofer Ges Forschung Multisignal encoder, multisignal decoder, and related methods using signal whitening or signal post processing
WO2020164751A1 (en) * 2019-02-13 2020-08-20 Fraunhofer-Gesellschaft zur Förderung der angewandten Forschung e.V. Decoder and decoding method for lc3 concealment including full frame loss concealment and partial frame loss concealment
EP3719799A1 (en) * 2019-04-04 2020-10-07 FRAUNHOFER-GESELLSCHAFT zur Förderung der angewandten Forschung e.V. A multi-channel audio encoder, decoder, methods and computer program for switching between a parametric multi-channel operation and an individual channel operation

Patent Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20090112606A1 (en) * 2007-10-26 2009-04-30 Microsoft Corporation Channel extension coding for multi-channel source
US20110123031A1 (en) * 2009-05-08 2011-05-26 Nokia Corporation Multi channel audio processing
CN101695150A (zh) * 2009-10-12 2010-04-14 清华大学 多声道音频编码方法、编码器、解码方法和解码器
CN104364842A (zh) * 2012-04-18 2015-02-18 诺基亚公司 立体声音频信号编码器
CN109416912A (zh) * 2016-06-30 2019-03-01 杜塞尔多夫华为技术有限公司 一种对多声道音频信号进行编码和解码的装置和方法

Non-Patent Citations (1)

* Cited by examiner, † Cited by third party
Title
See also references of EP4174855A4

Also Published As

Publication number Publication date
US20230154471A1 (en) 2023-05-18
CN113948095B (zh) 2025-02-25
KR20230036146A (ko) 2023-03-14
EP4174855A4 (en) 2023-12-06
JP2023533366A (ja) 2023-08-02
EP4174855A1 (en) 2023-05-03
CN113948095A (zh) 2022-01-18
JP7519531B2 (ja) 2024-07-19
US12437767B2 (en) 2025-10-07
KR102948552B1 (ko) 2026-04-03

Similar Documents

Publication Publication Date Title
WO2022012553A1 (zh) 多声道音频信号的编解码方法和装置
CN113593586B (zh) 音频信号编码方法、解码方法、编码设备以及解码设备
CN115691514B (zh) 一种多声道信号的编解码方法和装置
WO2023005414A1 (zh) 一种音频信号的编解码方法和装置
JP7636516B2 (ja) マルチ・チャネル・オーディオ信号符号化方法及び装置
CN116913293A (zh) 一种多声道音频的混合模式编码方法、装置、设备及介质
WO2022247651A1 (zh) 多声道音频信号的编码方法和装置
WO2022012675A1 (zh) 多声道音频信号的编码方法和装置
US12431144B2 (en) Multi-channel audio signal encoding and decoding method and apparatus
CN118782053A (zh) 非负矩阵分解的蓝牙接收端单声道上混方法、装置、介质
CN120266204A (zh) 参数空间音频编码
KR102869250B1 (ko) 선형 예측 코딩 파라미터 코딩 방법 및 코딩 장치
RU2811412C1 (ru) СПОСОБ КОДИРОВАНИЯ ПАРАМЕТРОВ КОДИРОВАНИЯ С ЛИНЕЙНЫМ ПРОГНОЗИРОВАНИЕМ и УСТРОЙСТВО КОДИРОВАНИЯ
WO2023173941A1 (zh) 一种多声道信号的编解码方法和编解码设备以及终端设备

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 21843116

Country of ref document: EP

Kind code of ref document: A1

ENP Entry into the national phase

Ref document number: 2023502888

Country of ref document: JP

Kind code of ref document: A

WWE Wipo information: entry into national phase

Ref document number: 202317004150

Country of ref document: IN

ENP Entry into the national phase

Ref document number: 20237004819

Country of ref document: KR

Kind code of ref document: A

ENP Entry into the national phase

Ref document number: 2021843116

Country of ref document: EP

Effective date: 20230124

NENP Non-entry into the national phase

Ref country code: DE

WWG Wipo information: grant in national office

Ref document number: 202317004150

Country of ref document: IN