WO2022242534A1 - 编解码方法、装置、设备、存储介质及计算机程序 - Google Patents

编解码方法、装置、设备、存储介质及计算机程序 Download PDF

Info

Publication number
WO2022242534A1
WO2022242534A1 PCT/CN2022/092385 CN2022092385W WO2022242534A1 WO 2022242534 A1 WO2022242534 A1 WO 2022242534A1 CN 2022092385 W CN2022092385 W CN 2022092385W WO 2022242534 A1 WO2022242534 A1 WO 2022242534A1
Authority
WO
WIPO (PCT)
Prior art keywords
adjustment factor
variable
bits
encoding
latent variable
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/CN2022/092385
Other languages
English (en)
French (fr)
Inventor
夏丙寅
李佳蔚
王喆
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Huawei Technologies Co Ltd
Original Assignee
Huawei Technologies Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Huawei Technologies Co Ltd filed Critical Huawei Technologies Co Ltd
Priority to MX2023013828A priority Critical patent/MX2023013828A/es
Priority to BR112023024341A priority patent/BR112023024341A2/pt
Priority to KR1020237044093A priority patent/KR20240011767A/ko
Priority to JP2023572207A priority patent/JP7656094B2/ja
Priority to EP22803856.8A priority patent/EP4333432A4/en
Publication of WO2022242534A1 publication Critical patent/WO2022242534A1/zh
Priority to US18/515,612 priority patent/US12412587B2/en
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Images

Classifications

    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N19/00Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
    • H04N19/10Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
    • H04N19/102Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the element, parameter or selection affected or controlled by the adaptive coding
    • H04N19/13Adaptive entropy coding, e.g. adaptive variable length coding [AVLC] or context adaptive binary arithmetic coding [CABAC]
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L19/00Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis
    • G10L19/04Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis using predictive techniques
    • G10L19/16Vocoder architecture
    • G10L19/18Vocoders using multiple modes
    • G10L19/24Variable rate codecs, e.g. for generating different qualities using a scalable representation such as hierarchical encoding or layered encoding
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L19/00Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis
    • G10L19/0017Lossless audio signal coding; Perfect reconstruction of coded audio signal by transmission of coding error
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04LTRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
    • H04L65/00Network arrangements, protocols or services for supporting real-time applications in data packet communication
    • H04L65/60Network streaming of media packets
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L19/00Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis
    • G10L19/002Dynamic bit allocation
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L19/00Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis
    • G10L19/04Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis using predictive techniques
    • G10L19/16Vocoder architecture
    • G10L19/167Audio streaming, i.e. formatting and decoding of an encoded audio signal representation into a data stream for transmission or storage purposes
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L25/00Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00
    • G10L25/27Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 characterised by the analysis technique
    • G10L25/30Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 characterised by the analysis technique using neural networks
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04LTRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
    • H04L65/00Network arrangements, protocols or services for supporting real-time applications in data packet communication
    • H04L65/60Network streaming of media packets
    • H04L65/75Media network packet handling
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04LTRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
    • H04L65/00Network arrangements, protocols or services for supporting real-time applications in data packet communication
    • H04L65/80Responding to QoS
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N19/00Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
    • H04N19/10Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
    • H04N19/134Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the element, parameter or criterion affecting or controlling the adaptive coding
    • H04N19/146Data rate or code amount at the encoder output
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N19/00Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
    • H04N19/10Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
    • H04N19/169Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding
    • H04N19/184Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding the unit being bits, e.g. of the compressed video stream
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N19/00Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
    • H04N19/90Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using coding techniques not provided for in groups H04N19/10-H04N19/85, e.g. fractals
    • H04N19/91Entropy coding, e.g. variable length coding [VLC] or arithmetic coding
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/40Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
    • H04N21/43Processing of content or additional data, e.g. demultiplexing additional data from a digital video stream; Elementary client operations, e.g. monitoring of home network or synchronising decoder's clock; Client middleware
    • H04N21/439Processing of audio elementary streams
    • H04N21/4398Processing of audio elementary streams involving reformatting operations of audio signals
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/40Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
    • H04N21/45Management operations performed by the client for facilitating the reception of or the interaction with the content or administrating data related to the end-user or to the client device itself, e.g. learning user preferences for recommending movies, resolving scheduling conflicts
    • H04N21/466Learning process for intelligent management, e.g. learning user preferences for recommending movies
    • H04N21/4662Learning process for intelligent management, e.g. learning user preferences for recommending movies characterized by learning algorithms
    • H04N21/4666Learning process for intelligent management, e.g. learning user preferences for recommending movies characterized by learning algorithms using neural networks, e.g. processing the feedback provided by the user
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/80Generation or processing of content or additional data by content creator independently of the distribution process; Content per se
    • H04N21/81Monomedia components thereof
    • H04N21/8106Monomedia components thereof involving special audio data, e.g. different tracks for different languages
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L19/00Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis
    • G10L19/02Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis using spectral analysis, e.g. transform vocoders or subband vocoders
    • G10L19/0212Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis using spectral analysis, e.g. transform vocoders or subband vocoders using orthogonal transformation
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L19/00Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis
    • G10L19/02Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis using spectral analysis, e.g. transform vocoders or subband vocoders
    • G10L19/032Quantisation or dequantisation of spectral components
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L19/00Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis
    • G10L19/04Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis using predictive techniques
    • G10L19/16Vocoder architecture
    • G10L19/173Transcoding, i.e. converting between two coded representations avoiding cascaded coding-decoding
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04LTRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
    • H04L65/00Network arrangements, protocols or services for supporting real-time applications in data packet communication
    • H04L65/60Network streaming of media packets
    • H04L65/61Network streaming of media packets for supporting one-way streaming services, e.g. Internet radio
    • H04L65/611Network streaming of media packets for supporting one-way streaming services, e.g. Internet radio for multicast or broadcast

Definitions

  • the embodiments of the present application relate to the technical field of encoding and decoding, and in particular, relate to an encoding and decoding method, device, equipment, storage medium, and computer program.
  • Codec technology is an indispensable link in media applications such as media communication and media broadcasting. Therefore, how to encode and decode has become one of the concerns of the industry.
  • the related art proposes a method for encoding and decoding an audio signal.
  • the audio signal in the time domain is processed by a modified discrete cosine transform (modified discrete cosine transform, MDCT) to obtain audio signal in the frequency domain.
  • the audio signal in the frequency domain is processed by encoding the neural network model to obtain a latent variable, and the latent variable is used to indicate the feature of the audio signal in the frequency domain.
  • the latent variable is quantized to obtain the quantized latent variable, entropy coding is performed on the quantized latent variable, and the entropy coding result is written into the code stream.
  • the quantized latent variable is determined based on the code stream, and the quantized latent variable is dequantized to obtain the latent variable.
  • the latent variable is processed by decoding the neural network model to obtain the audio signal in the frequency domain, and the audio signal in the frequency domain is processed by inverse discrete cosine transform (Inverse modified discrete cosine transform, IMDCT) to obtain the reconstructed time domain audio signal.
  • IMDCT inverse discrete cosine transform
  • Embodiments of the present application provide an encoding and decoding method, device, equipment, storage medium, and computer program, which can meet the requirements of an encoder for a stable encoding rate. Described technical scheme is as follows:
  • an encoding method is provided, and the method can be applied to a codec not including a context model, and can also be applied to a codec including a context model. Moreover, not only the latent variables generated by the media data to be encoded can be adjusted by the adjustment factor, but also the latent variables determined by the context model can be adjusted by the adjustment factor. Therefore, the method will be explained in detail in several cases.
  • the media data to be encoded is processed through the first encoding neural network model to obtain the first latent variable, the first latent variable is used to indicate the characteristics of the media data to be encoded; the first latent variable is determined based on the first latent variable A variable adjustment factor, the first variable adjustment factor is used to make the number of encoded bits of the entropy encoding result of the second latent variable meet the preset encoding rate condition, and the second latent variable is adjusted by the first variable adjustment factor to the first latent Obtaining after variable adjustment; obtaining an entropy coding result of the second latent variable; writing the entropy coding result of the second latent variable and the coding result of the first variable adjustment factor into the code stream.
  • the number of encoded bits of the entropy encoding result of the second latent variable satisfies the preset encoding rate condition, it can be ensured that the encoding bit number of the entropy encoding result of the latent variable corresponding to each frame of media data can meet the preset encoding rate condition, that is Yes, it can ensure that the number of encoded bits of the entropy encoding result of the latent variable corresponding to each frame of media data is basically consistent, rather than changing dynamically, thus meeting the encoder's requirement for a stable encoding rate.
  • the need to transmit side information (such as window type, time-domain noise shaping (TNS: Temporal Noise Shaping) parameters, frequency-domain noise shaping (FDNS: Frequency-domain noise shaping) parameters, and/or bandwidth extension (BWE : bandwidth extension) parameter, etc.), it can ensure that the number of encoded bits of the entropy encoding result of the latent variable corresponding to each frame of media data and the encoded bit number of the side information are basically consistent, so as to meet the encoder's need for a stable encoding rate .
  • TMS Temporal Noise Shaping
  • FDNS Frequency-domain noise shaping
  • BWE bandwidth extension
  • the media data to be encoded is an audio signal, a video signal, or an image.
  • the form of the media data to be encoded may be in any form, which is not limited in this embodiment of the present application.
  • the implementation process of processing the media data to be encoded by the first encoding neural network model is: input the media data to be encoded into the first encoding neural network model, and obtain the first latent variable output by the first encoding neural network model.
  • the media data to be encoded is preprocessed, and the preprocessed media data is input into the first encoding neural network model to obtain the first latent variable output by the first encoding neural network model.
  • the media data to be encoded can be used as the input of the first encoding neural network model to determine the first latent variable, and the media data to be encoded can also be preprocessed and then used as the input of the first encoding neural network model Determine the first latent variable.
  • satisfying the preset encoding rate condition includes that the number of encoding bits is less than or equal to the target encoding bit number; or, satisfying the preset encoding rate condition includes encoding the number of bits Less than or equal to the target number of coding bits, and the difference between the number of coding bits and the target number of coding bits is less than the threshold value of the number of bits; or, meeting the preset coding rate conditions includes the number of coding bits being the maximum number of coding bits less than or equal to the target number of coding bits .
  • satisfying the preset encoding rate condition includes that the absolute value of the difference between the number of encoded bits and the target number of encoded bits is smaller than a bit number threshold. That is, satisfying the preset encoding rate condition includes that the number of encoded bits is less than or equal to the target encoded bit number, and the difference between the target encoded bit number and the encoded bit number is less than the bit number threshold; or, meeting the preset encoding rate condition includes encoding bits The number is greater than or equal to the target encoding bit number, and the difference between the encoding bit number and the target encoding bit number is less than the bit number threshold.
  • the target number of coding bits may be set in advance.
  • the target number of coding bits may also be determined based on the coding rate, and different coding rates correspond to different target coding bits.
  • the media data to be encoded may be encoded using a fixed code rate, or the media data to be encoded may be encoded using a variable code rate.
  • the number of bits of the media data to be encoded in the current frame can be determined based on the fixed bit rate, and then the number of used bits in the current frame can be subtracted to obtain the target of the current frame
  • the number of encoded bits may be the number of bits for encoding side information, etc., and usually, the side information of each frame of media data is different, so the target number of encoding bits of each frame of media data is usually different.
  • the number of bits of media data to be encoded in the current frame can be determined based on the specified code rate, and then the number of used bits in the current frame can be subtracted to obtain the target number of encoded bits in the current frame.
  • the number of used bits can be the number of bits encoded by side information, and in some cases, the side information of media data of different frames can be different, so the target number of encoded bits of media data of different frames is usually is different.
  • the initial number of coding bits may be determined based on the first latent variable, and the first variable adjustment factor may be determined based on the initial number of coding bits and the target number of coding bits.
  • the initial coding bit number is the coding bit number of the entropy coding result of the first latent variable; or, the initial coding bit number is the coding bit number of the entropy coding result of the first latent variable adjusted by the first initial adjustment factor.
  • the first initial adjustment factor may be a first preset adjustment factor.
  • the implementation process of adjusting the first latent variable based on the first initial adjustment factor is: multiply each element in the first latent variable by the corresponding element in the first initial adjustment factor to obtain the adjusted first latent variable .
  • each element in the first latent variable may be divided by the corresponding element in the first initial adjustment factor to obtain the adjusted first latent variable.
  • the embodiment of the present application does not limit the adjustment method.
  • the first initial adjustment factor is determined as the first variable adjustment factor. If the initial number of encoded bits is not equal to the target number of encoded bits, the first variable adjustment factor is determined in a first round-robin manner based on the initial number of encoded bits and the target number of encoded bits.
  • the i-th cycle processing of the first cycle method includes the following steps: determining the adjustment factor of the i-th cycle processing, i is a positive integer, and adjusting the first latent variable based on the adjustment factor of the i-th cycle processing to obtain the first
  • the first latent variable after i adjustments Determine the number of coded bits of the entropy coding result of the first latent variable after the i-th adjustment, so as to obtain the number of coded bits of the i-th time.
  • the (i+1)th loop processing of the first loop mode is executed.
  • the execution of the first loop mode is terminated, and the first variable adjustment factor is determined based on the adjustment factor of the i-th loop process.
  • the realization process of determining the adjustment factor of the ith loop processing is: based on the adjustment factor of the i-1th loop processing in the first loop mode, the number of coded bits of the i-1th time, and the target number of coded bits, determine the Adjustment factor for i-cycle processing.
  • the adjustment factor of the i-1 th round of processing is the first initial adjustment factor
  • the number of encoded bits of the i-1 th time is the initial number of encoded bits.
  • the continuous adjustment condition includes that the number of coded bits of the i-1th time and the number of coded bits of the ith time are both smaller than the target number of coded bits, or the continued adjustment condition includes the number of coded bits of the i-1th time and the number of coded bits of the ith time
  • the number of encoded bits is greater than the target number of encoded bits.
  • the condition for continuing to adjust includes that the number of coded bits for the ith time does not exceed the target number of coded bits.
  • the meaning of not crossing here means: the number of encoded bits for the first i-1 times is always smaller than the target number of encoded bits, and the number of encoded bits for the ith time is still smaller than the target number of encoded bits.
  • the number of encoded bits for the first i-1 times is always greater than the target number of encoded bits, and the number of encoded bits for the ith time is still greater than the target number of encoded bits.
  • overstepping means that the number of encoded bits for the first i-1 times is always smaller than the target number of encoded bits, and the number of encoded bits for the ith time is greater than the target number of encoded bits.
  • the number of encoded bits for the first i-1 times is always greater than the target number of encoded bits, and the number of encoded bits for the ith time is smaller than the target number of encoded bits.
  • the implementation process of determining the first variable adjustment factor based on the adjustment factor of the i-th cyclic processing includes: in the case that the i-th coding bit number is equal to the target coding bit number, determining the i-th cyclic processing adjustment factor as First variable adjustment factor. Or, in the case that the number of coded bits of the ith time is not equal to the target number of coded bits, the first variable is determined based on the adjustment factor of the i-th round-robin processing and the adjustment factor of the i-1-th round-robin processing of the first round-robin mode adjustment factor.
  • the adjustment factor of the i-th loop processing is the adjustment factor obtained last time through the above-mentioned first round-robin manner, and the number of coded bits of the i-th time is the number of coded bits obtained last time.
  • the last obtained adjustment factor is determined as the first variable adjustment factor.
  • the first variable adjustment factor is determined based on the adjustment factors obtained twice later.
  • the implementation process of determining the first variable adjustment factor includes: determining the adjustment factor of the i th cyclic processing and the first An average value of the adjustment factors of i-1 cycles of processing, and the first variable adjustment factor is determined based on the average value.
  • the average value may be directly determined as the first variable adjustment factor, or the average value may be multiplied by a preset constant to obtain the first variable adjustment factor.
  • this constant can be less than 1.
  • the implementation process of determining the first variable adjustment factor can also be: based on the adjustment factor of the i-th cyclic processing and the adjustment factor of the i-1th cyclic processing, the first variable adjustment factor is determined through the second cyclic manner.
  • the j-th loop processing of the second loop manner includes the following steps: determining the third adjustment factor of the j-th loop processing based on the first adjustment factor of the j-th loop processing and the second adjustment factor of the j-th loop processing.
  • Adjustment factor wherein, in the case that j is equal to 1, the first adjustment factor of the jth cycle processing is one of the adjustment factor of the i-th cycle processing and the adjustment factor of the i-1th cycle processing, and the jth cycle processing
  • the second adjustment factor of the second cycle processing is the other one of the adjustment factor of the i-th cycle processing and the adjustment factor of the i-1th cycle processing, and the first adjustment factor of the j-th cycle processing corresponds to the j-th cycle processing.
  • the number of coded bits, the second adjustment factor of the jth cyclic process corresponds to the second coded bit number of the jth time, and the first coded bit number of the jth time refers to the adjusted first adjustment factor of the jth cyclic process
  • the number of coded bits of the entropy coding result of the first latent variable, the second coded bit number of the jth time refers to the coded bits of the entropy coding result of the first latent variable adjusted by the second adjustment factor of the jth cycle processing
  • the number of first coded bits of the jth time is less than the number of second coded bits of the jth time.
  • the third number of encoded bits of the jth time refers to the number of encoded bits of the entropy coding result of the first latent variable adjusted by the third adjustment factor of the jth cycle processing. If the number of third encoded bits of the jth time does not satisfy the condition for continuing the loop, the execution of the second loop mode is terminated, and the third adjustment factor of the jth loop process is determined as the first variable adjustment factor.
  • the third coding bit number of the jth time satisfies the continuation loop condition, the third coding bit number of the jth time is greater than the target coding bit number and less than the second coding bit number of the jth time, the third coded bit number of the jth time loop processing
  • the adjustment factor is used as the second adjustment factor of the j+1th round of processing, the first adjustment factor of the jth round of processing is used as the first adjustment factor of the j+1th round of processing, and the j+1th round of the second round is executed. 1 cycle processing.
  • the third encoding bit number of the jth time satisfies the continuation loop condition, the third encoding bit number of the jth time is less than the target encoding bit number and greater than the first encoding bit number of the jth time, the third encoding bit number of the jth loop processing
  • the adjustment factor is used as the first adjustment factor of the j+1th round of processing, and the second adjustment factor of the jth round of processing is used as the second adjustment factor of the j+1th round of processing, and the j+1th round of the second round is executed. 1 cycle processing.
  • the j-th loop processing in the second loop manner includes the following steps: determining the j-th loop processing based on the first adjustment factor of the j-th loop processing and the second adjustment factor of the j-th loop processing.
  • Three adjustment factors wherein, in the case that j is equal to 1, the first adjustment factor of the jth cyclic processing is one of the adjustment factor of the i th cyclic processing and the adjustment factor of the i-1 th cyclic processing, the first The second adjustment factor of the j-th cycle processing is the other one of the adjustment factor of the i-th cycle processing and the adjustment factor of the i-1-th cycle processing, and the first adjustment factor of the j-th cycle processing corresponds to the j-th cycle processing.
  • the number of first coded bits, the second adjustment factor of the jth cyclic processing corresponds to the second number of coded bits of the jth time, the number of first coded bits of the jth time is less than the number of second coded bits of the jth time, and j is positive integer.
  • the execution of the second loop mode is terminated, and the third adjustment factor of the jth loop process is determined as the first variable adjustment factor. If j reaches the maximum number of cycles and the number of third coded bits of the jth time satisfies the continuation cycle condition, the execution of the second cycle mode is terminated, and the first variable adjustment factor is determined based on the first adjustment factor of the jth cycle processing.
  • the number of third coded bits of the jth time satisfies the condition for continuing the cycle, and the number of third coded bits of the jth time is greater than the number of target coded bits and less than the number of second coded bits of the jth time
  • the The third adjustment factor of the j cycle processing is used as the second adjustment factor of the j+1 cycle processing
  • the first adjustment factor of the j cycle processing is used as the first adjustment factor of the j+1 cycle processing
  • the first adjustment factor is executed.
  • the number of third coded bits of the jth time satisfies the condition for continuing the cycle, and the number of third coded bits of the jth time is less than the target number of coded bits and greater than the number of first coded bits of the jth time
  • the The third adjustment factor of the j loop processing is used as the first adjustment factor of the j+1 loop processing
  • the second adjustment factor of the j loop processing is used as the second adjustment factor of the j+1 loop processing
  • the first adjustment factor is executed.
  • the implementation process of determining the third adjustment factor of the j-th cyclic processing based on the first adjustment factor of the j-th cyclic processing and the second adjustment factor of the j-th cyclic processing includes: determining the first adjustment factor of the j-th cyclic processing Factor and the average value of the second adjustment factor of the j-th cycle processing, and determine the third adjustment factor of the j-th cycle processing based on the average value.
  • the average value may be directly determined as the third adjustment factor for the j-th cycle processing, or the average value may be multiplied by a preset constant to obtain the third adjustment factor for the j-th cycle processing.
  • this constant can be less than 1.
  • the implementation process of obtaining the third coded bit number of the jth time includes: adjusting the first latent variable based on the third adjustment factor of the jth cyclic process to obtain the adjusted first latent variable, and adjusting the adjusted first latent variable A latent variable is quantified to obtain a quantized first latent variable. Entropy encoding is performed on the quantized first latent variable, and the number of encoded bits of the entropy encoding result is counted to obtain the third encoded bit number of the jth time.
  • the implementation process of determining the first variable adjustment factor based on the first adjustment factor of the j-th cycle processing includes: the first adjustment factor of the j-th cycle processing The factor is determined as the first variable adjustment factor.
  • the implementation process of determining the first variable adjustment factor based on the first adjustment factor of the jth cycle processing includes: determining the target encoding bit number and the jth A first difference between the number of coded bits, and a second difference between the second number of coded bits determined for the jth time and the target number of coded bits.
  • the first adjustment factor of the j-th loop processing is determined as the first variable adjustment factor. If the second difference is smaller than the first difference, the second adjustment factor of the j-th loop processing is determined as the first variable adjustment factor. If the first difference is equal to the second difference, determine the first adjustment factor of the jth cyclic process as the first variable adjustment factor, or determine the second adjustment factor of the j th cyclic process as the first variable adjustment factor.
  • the continuation loop condition includes that the jth third encoding bit number is greater than the target encoding bit number, or the continuation loop condition includes the jth third encoding bit number
  • the number of bits is smaller than the target number of encoded bits, and the difference between the target number of encoded bits and the j-th third encoded number of bits is greater than the threshold value of the number of bits.
  • the continuation loop condition includes that the absolute value of the difference between the target encoding bit number and the j-th third encoding bit number is greater than the bit number threshold.
  • condition for continuing the loop includes that the third coded bit number of the jth time is greater than the target coded bit number, and the difference between the jth third coded bit number and the target coded bit number is greater than the bit number threshold, or the loop continues
  • the conditions include that the number of third encoded bits of the jth time is smaller than the target number of encoded bits, and the difference between the target number of encoded bits and the third number of encoded bits of the jth time is greater than a threshold value of the number of bits.
  • the first initial adjustment factor is determined as the first variable adjustment factor.
  • the first variable adjustment factor is determined according to the above first implementation manner.
  • the adjustment factor of the i-1th loop processing of the first loop mode is adjusted according to the length of the first step to obtain The adjustment factor of the ith loop processing, at this time, the condition for continuing to adjust includes that the number of coded bits in the ith time is less than the target number of coded bits.
  • the adjustment condition includes that the i-th number of encoded bits is greater than the target number of encoded bits.
  • adjusting the adjustment factor of the i-1th loop processing of the first loop method according to the first step length can refer to increasing the adjustment factor of the i-1th loop process according to the first step length, and adjusting the adjustment factor according to the second step size
  • the adjustment factor of the i-1 th loop processing in the first loop manner may refer to reducing the adjustment factor of the i-1 th loop processing according to the second step size.
  • the above-mentioned increase processing and decrease processing may be linear or non-linear.
  • the sum of the adjustment factor of the i-1th loop processing and the first step length can be determined as the adjustment factor of the i-th loop processing, and the adjustment factor of the i-1th loop processing can be combined with the second step size The difference is determined as the adjustment factor for the ith cycle processing.
  • first step size and the second step size may be set in advance, and the first step size and the second step size may be adjusted based on different requirements.
  • first step length and the second step length may be equal or not.
  • the first initial adjustment factor is determined as the first variable adjustment factor.
  • the first variable adjustment factor may be determined according to the above-mentioned first implementation manner.
  • the first variable adjustment factor may also be determined according to the situation that the initial number of coding bits is greater than the target number of coding bits in the above-mentioned second implementation manner.
  • the entropy coding result of the third latent variable is also obtained.
  • the third latent variable is determined based on the second latent variable through the context model, and the third latent variable is used to indicate the second a probability distribution of the latent variable; writing the entropy encoding result of the third latent variable into the code stream.
  • the total number of encoded bits of the entropy encoding result of the second latent variable and the entropy encoding result of the third latent variable satisfies the preset encoding rate condition.
  • determining the first variable adjustment factor based on the first latent variable includes: determining the initial number of coding bits based on the first latent variable; and determining the first variable adjustment factor based on the initial number of coding bits and the target number of coding bits.
  • the manner of determining the number of initial coding bits based on the first latent variable may include two manners, which will be introduced respectively next.
  • the corresponding context initial coding bit number and the initial entropy coding model parameter are determined through the context model, and the coding bit number of the entropy coding result of the first latent variable is determined based on the initial entropy coding model parameter , to get the number of basic initial coded bits.
  • the initial encoding bit number is determined based on the context initial encoding bit number and the basic initial encoding bit number.
  • the context model includes a context encoding neural network model and a context decoding neural network model.
  • the implementation process of determining the corresponding context initial encoding bits and initial entropy encoding model parameters through the context model includes: processing the first latent variable through the context encoding neural network model to obtain the fifth latent variable, the fifth The latent variable is used to indicate the probability distribution of the first latent variable.
  • the entropy coding result of the fifth latent variable is determined, and the coding bit number of the entropy coding result of the fifth latent variable is used as the initial coding bit number of the context.
  • the fifth latent variable is reconstructed based on the entropy coding result of the fifth latent variable, and the reconstructed fifth latent variable is processed through the context decoding neural network model to obtain initial entropy coding model parameters.
  • the first preset adjustment factor is used as the first initial adjustment factor
  • the first latent variable is adjusted based on the first initial adjustment factor
  • the corresponding context is determined through the context model based on the adjusted first latent variable
  • the initial encoding bit number and the initial entropy encoding model parameter determine the encoding bit number of the entropy encoding result of the adjusted first latent variable, and obtain the basic initial encoding bit number.
  • the initial encoding bit number is determined based on the context initial encoding bit number and the basic initial encoding bit number.
  • the context model includes a context encoding neural network model and a context decoding neural network model.
  • the implementation process of determining the corresponding context initial encoding bits and initial entropy encoding model parameters through the context model includes: processing the adjusted first latent variable through the context encoding neural network model to obtain the first Of the six latent variables, the sixth latent variable is used to indicate the adjusted probability distribution of the first latent variable.
  • the entropy coding result of the sixth latent variable is determined, and the coding bit number of the entropy coding result of the sixth latent variable is used as the initial coding bit number of the context.
  • the sixth latent variable is reconstructed based on the entropy coding result of the sixth latent variable, and the reconstructed sixth latent variable is processed through the context decoding neural network model to obtain initial entropy coding model parameters.
  • the first initial adjustment factor is determined as the first variable adjustment factor. If the initial number of encoded bits is not equal to the target number of encoded bits, the first variable adjustment factor is determined in a first round-robin manner based on the initial number of encoded bits and the target number of encoded bits.
  • the i-th cycle processing of the first cycle method includes the following steps: determining the adjustment factor of the i-th cycle processing, i is a positive integer, and adjusting the first latent variable based on the adjustment factor of the i-th cycle processing to obtain the first The first latent variable after i adjustments. Based on the i-th adjusted first latent variable, determine the corresponding i-th context coding bit number and i-th entropy coding model parameters through the context model, and determine the i-th time based on the i-th entropy coding model parameters.
  • the number of coded bits of the entropy coding result of the adjusted first latent variable is obtained to obtain the number of basic coded bits of the ith time, and the code of the ith time is determined based on the number of coded bits of the context code of the ith time and the number of basic coded bits of the ith time number of bits.
  • the (i+1)th loop processing of the first loop mode is executed.
  • the execution of the first loop mode is terminated, and the first variable adjustment factor is determined based on the adjustment factor of the i-th loop process.
  • the first initial adjustment factor is determined as the first variable adjustment factor.
  • the first variable adjustment factor is determined according to the above first implementation manner.
  • the adjustment factor of the i-1th loop processing of the first loop mode is adjusted according to the length of the first step to obtain The adjustment factor of the ith loop processing, at this time, the condition for continuing to adjust includes that the number of coded bits in the ith time is less than the target number of coded bits.
  • the adjustment condition includes that the i-th number of encoded bits is greater than the target number of encoded bits.
  • adjusting the adjustment factor of the i-1th loop processing of the first loop method according to the first step length can refer to increasing the adjustment factor of the i-1th loop process according to the first step length, and adjusting the adjustment factor according to the second step size
  • the adjustment factor of the i-1 th loop processing in the first loop manner may refer to reducing the adjustment factor of the i-1 th loop processing according to the second step size.
  • the above-mentioned increase processing and decrease processing may be linear or non-linear.
  • the sum of the adjustment factor of the i-1th loop processing and the first step length can be determined as the adjustment factor of the i-th loop processing, and the adjustment factor of the i-1th loop processing can be combined with the second step size The difference is determined as the adjustment factor for the ith cycle processing.
  • first step size and the second step size may be set in advance, and the first step size and the second step size may be adjusted based on different requirements.
  • first step length and the second step length may be equal or not.
  • the first initial adjustment factor is determined as the first variable adjustment factor.
  • the first variable adjustment factor may be determined according to the above-mentioned first implementation manner.
  • the first variable adjustment factor may also be determined according to the situation that the initial number of coding bits is greater than the target number of coding bits in the above-mentioned second implementation manner.
  • the second variable adjustment factor is also determined based on the first latent variable; the entropy encoding result of the fourth latent variable is obtained, and the fourth latent variable is based on the second variable adjustment factor to the third
  • the latent variable is adjusted, the third latent variable is determined based on the second latent variable through the context model, and the third latent variable is used to indicate the probability distribution of the second latent variable; the entropy encoding result of the fourth latent variable and the second
  • the encoding result of the variable adjustment factor is written into the code stream; wherein, the total number of encoded bits of the entropy encoding result of the second latent variable and the entropy encoding result of the fourth latent variable satisfies the preset encoding rate condition.
  • the manner of determining the first variable adjustment factor and the second variable adjustment factor based on the first latent variable may include two ways, respectively:
  • the corresponding context initial coding bit number and the initial entropy coding model parameter are determined through the context model; based on the initial entropy coding model parameter, the coding bit number of the entropy coding result of the first latent variable is determined , to obtain the basic initial encoding bit number; based on the context initial encoding bit number, the basic initial encoding bit number and the target encoding bit number, determine a first variable adjustment factor and a second variable adjustment factor.
  • the first preset adjustment factor is used as the first initial adjustment factor
  • the second preset adjustment factor is used as the second initial adjustment factor.
  • the first latent variable is adjusted based on the first initial adjustment factor.
  • the corresponding context initial coding bit number and initial entropy coding model parameters are determined through the context model.
  • An encoding model parameter is used to determine the encoding bit number of the entropy encoding result of the first latent variable, so as to obtain the basic initial encoding bit number.
  • a first variable adjustment factor and a second variable adjustment factor are determined based on the context initial encoding bit number, the basic initial encoding bit number and the target encoding bit number.
  • the context model includes a context encoding neural network model and a context decoding neural network model.
  • the implementation process of determining the corresponding context initial coding bits and initial entropy coding model parameters through the context model includes: using the context coding neural network model to adjust the adjusted first latent Variables are processed to obtain a sixth latent variable, which is used to indicate the adjusted probability distribution of the first latent variable.
  • the sixth latent variable is adjusted based on the second initial adjustment factor to obtain a seventh latent variable.
  • the entropy coding result of the seventh latent variable is determined, and the coding bit number of the entropy coding result of the seventh latent variable is used as the initial coding bit number of the context.
  • the seventh latent variable is reconstructed based on the entropy coding result of the seventh latent variable, and the reconstructed seventh latent variable is processed through a context decoding neural network model to obtain initial entropy coding model parameters.
  • the implementation process of adjusting the sixth latent variable based on the second initial adjustment factor is: multiplying each element in the sixth latent variable by the corresponding element in the second initial adjustment factor to obtain the seventh latent variable.
  • each element in the sixth latent variable may be divided by the corresponding element in the second initial adjustment factor to obtain the seventh latent variable.
  • the embodiment of the present application does not limit the adjustment method.
  • the implementation process of determining the first variable adjustment factor and the second variable adjustment factor may include two implementations, which will be introduced respectively next.
  • the second variable adjustment factor is set as the second initial adjustment factor, and based on at least one of the basic initial encoding bit number and the context initial encoding bit number, and the target encoding bit number, the basic target encoding bit number is determined .
  • the first variable adjustment factor and the basic actual number of encoded bits are determined.
  • the basic actual number of encoded bits refers to the first variable adjusted by the first variable adjustment factor.
  • the number of encoded bits of the entropy encoding result of the latent variable.
  • the context target number of coding bits is determined.
  • the second variable adjustment factor is determined based on the target number of coding bits of the context and the number of initial coding bits of the context.
  • the implementation process of determining the basic target coding bit number includes: subtracting the context initial coding bit number from the target coding bit number, Get the number of base target coded bits. Alternatively, determine the ratio between the basic initial coding bit number and the context initial coding bit number, and determine the basic target coding bit number based on the ratio and the target coding bit number. Alternatively, the basic target number of coding bits is determined based on the ratio of the target number of coding bits to the basic initial number of coding bits. Of course, it can also be determined through other implementation processes.
  • Mode 11 when the basic initial number of coding bits is equal to the basic target number of coding bits, determine the first initial adjustment factor as the first variable adjustment factor. In the case that the basic initial coding bit number is not equal to the basic target coding bit number, the first variable adjustment factor is determined in a first round-robin manner based on the basic initial coding bit number and the basic target coding bit number.
  • the i-th cycle processing of the first cycle method includes the following steps: determining the adjustment factor of the i-th cycle processing, i is a positive integer, and adjusting the first latent variable based on the adjustment factor of the i-th cycle processing to obtain the first The first latent variable after i adjustments.
  • the i-th context coding bit number and the i-th entropy coding model parameters are determined through the context model, and based on the i-th entropy coding model parameters, Determine the number of coding bits of the entropy coding result of the first latent variable after the i-th adjustment, and obtain the basic coding bits of the i-th time, and determine the first The number of encoded bits for i times.
  • the (i+1)th loop processing of the first loop mode is executed.
  • the execution of the first loop mode is terminated, and the first variable adjustment factor is determined based on the adjustment factor of the i-th loop process.
  • Mode 12 when the basic initial number of coding bits is equal to the basic target number of coding bits, determine the first initial adjustment factor as the first variable adjustment factor.
  • the first variable adjustment factor is determined according to the above-mentioned manner 11.
  • the adjustment factor of the i-1th cycle processing of the first cycle method is adjusted according to the length of the first step to obtain the first The adjustment factor for the i-time loop processing, at this time, the condition for continuing the adjustment includes that the number of coded bits for the ith time is less than the basic target number of coded bits.
  • the condition for continuing to adjust includes that the number of coded bits for the ith time is greater than the number of basic target coded bits.
  • Mode 13 when the basic initial coding bit number is less than or equal to the basic target coding bit number, determine the first initial adjustment factor as the first variable adjustment factor.
  • the first variable adjustment factor can be determined according to the above method 11.
  • the first variable adjustment factor may also be determined according to the condition that the basic initial coded bit number is greater than the basic target coded bit number in the above manner 12.
  • the implementation process of determining the second variable adjustment factor based on the context target number of coding bits and the context initial number of coding bits is similar to the implementation process of determining the first variable adjustment factor based on the basic initial number of coding bits and the basic target number of coding bits.
  • the following are also divided into three ways to introduce respectively.
  • the second initial adjustment factor is determined as the second variable adjustment factor.
  • the first latent variable is adjusted based on the first variable adjustment factor to obtain the second latent variable.
  • the second variable adjustment factor is determined in a first loop manner.
  • the i-th cycle processing of the first cycle method includes the following steps: determining the adjustment factor of the i-th cycle processing, i is a positive integer, based on the second latent variable and the adjustment factor of the i-th cycle processing, determined through the context model
  • the i-th context coding bit number and the i-th entropy coding model parameters based on the i-th entropy coding model parameters, determine the number of coding bits of the entropy coding result of the second latent variable, and obtain the i-th basic coding bits
  • the number of coded bits of the i-th time is determined based on the number of coded bits of the i-th time context and the number of basic coded bits of the ith time.
  • the (i+1)th loop processing of the first loop mode is executed.
  • the execution of the first loop mode is terminated, and the second variable adjustment factor is determined based on the adjustment factor of the i-th loop process.
  • the implementation process of determining the i-th context encoding bit number and the i-th entropy encoding model parameters through the context model includes: using the context encoding neural network model to The second latent variable is processed to obtain a third latent variable, and the third latent variable is used to indicate the probability distribution of the second latent variable.
  • the third latent variable is adjusted based on the adjustment factor of the i-th cyclic process to obtain the i-th adjusted third latent variable.
  • the i-th adjusted third latent variable is reconstructed based on the entropy encoding result of the i-th adjusted third latent variable.
  • the adjusted third latent variable reconstructed for the ith time is adjusted by the adjustment factor of the ith cycle processing to obtain the reconstructed third latent variable.
  • the reconstructed third latent variable is processed through the context decoding neural network model to obtain the i-th entropy encoding model parameters.
  • the second initial adjustment factor is determined as the second variable adjustment factor.
  • the second variable adjustment factor is determined according to the above-mentioned manner 21.
  • the adjustment factor of the i-1th round of the first round-robin method is adjusted according to the length of the first step, and the first The adjustment factor for the i-time cyclic processing.
  • the continuous adjustment condition includes that the number of coded bits for the ith time is less than the target number of coded bits for the context.
  • adjust the adjustment factor of the i-1th loop processing of the first loop mode according to the second step length, and obtain the adjustment factor of the i-th loop process, at this time continue to adjust the condition that the number of coded bits for the ith time is greater than the target number of coded bits for the context.
  • the second initial adjustment factor is determined as the second variable adjustment factor.
  • the second variable adjustment factor may be determined according to the above manner 21.
  • the second variable adjustment factor can also be determined according to the situation that the initial number of coded bits of the context is greater than the number of target coded bits of the context in the above manner 22.
  • the target coding bit number is divided into the basic target coding bit number and the context target coding bit number, and the first variable adjustment factor is determined based on the basic target coding bit number and the basic initial coding bit number.
  • the second variable adjustment factor is determined based on the target number of coding bits of the context and the number of initial coding bits of the context.
  • a decoding method is provided, and in this method, the introduction is also divided into multiple situations.
  • the reconstructed second latent variable and the reconstructed first variable adjustment factor are determined based on the code stream; based on the reconstructed first variable adjustment factor, the reconstructed second latent variable is adjusted to obtain the reconstructed
  • the reconstructed first latent variable is used to indicate the characteristics of the media data to be decoded; the reconstructed first latent variable is processed by the first decoding neural network model to obtain the reconstructed media data.
  • the reconstructed third latent variable and the reconstructed first variable adjustment factor are determined based on the code stream; and the reconstructed second latent variable is determined based on the code stream and the reconstructed third latent variable.
  • the reconstructed second latent variable is adjusted to obtain the reconstructed first latent variable, and the reconstructed first latent variable is used to indicate the characteristics of the media data to be decoded; through the second A decoding neural network model processes the reconstructed first latent variable to obtain reconstructed media data.
  • determining the reconstructed second latent variable includes: processing the reconstructed third latent variable through the context decoding neural network model to obtain the reconstructed first entropy Coding model parameters; based on the code stream and the reconstructed first entropy coding model parameters, determine the reconstructed second latent variable.
  • the reconstructed fourth latent variable, the reconstructed second variable adjustment factor and the reconstructed first variable adjustment factor are determined based on the code stream; based on the code stream, the reconstructed fourth latent variable and the reconstructed The second variable adjustment factor, which determines the reconstructed second latent variable.
  • a reconstructed second latent variable is determined.
  • the reconstructed second latent variable is adjusted to obtain the reconstructed first latent variable, and the reconstructed first latent variable is used to indicate the characteristics of the media data to be decoded; through the second A decoding neural network model processes the reconstructed first latent variable to obtain reconstructed media data.
  • determining the reconstructed second latent variable includes: based on the reconstructed second variable adjustment factor, adjusting the reconstructed fourth The latent variable is adjusted to obtain the reconstructed third latent variable; the reconstructed third latent variable is processed through the context decoding neural network model to obtain the reconstructed second entropy coding model parameters; based on the code stream and the reconstructed A second entropy encodes the model parameters to determine a reconstructed second latent variable.
  • the media data is an audio signal, a video signal or an image.
  • an encoding device in a third aspect, is provided, and the encoding device has a function of implementing the behavior of the encoding method in the first aspect above.
  • the encoding device includes at least one module, and the at least one module is used to implement the encoding method provided in the first aspect above.
  • a decoding device in a fourth aspect, has the function of realizing the behavior of the decoding method in the second aspect above.
  • the decoding device includes at least one module, and the at least one module is used to implement the decoding method provided by the second aspect above.
  • an encoding device includes a processor and a memory, and the memory is used to store a program for executing the encoding method provided in the first aspect above.
  • the processor is configured to execute the program stored in the memory, so as to implement the encoding method provided in the first aspect above.
  • the encoding end device may further include a communication bus, which is used to establish a connection between the processor and the memory.
  • a decoding end device in a sixth aspect, includes a processor and a memory, and the memory is used to store a program for executing the decoding method provided in the second aspect above.
  • the processor is configured to execute the program stored in the memory, so as to implement the decoding method provided by the second aspect above.
  • the decoding device may further include a communication bus, which is used to establish a connection between the processor and the memory.
  • a computer-readable storage medium and instructions are stored in the storage medium, and when the instructions are run on a computer, the computer is made to execute the steps of the encoding method described in the first aspect above, or execute The steps of the decoding method described in the second aspect above.
  • a computer program product containing instructions, which, when the instructions are run on a computer, cause the computer to execute the steps of the encoding method described in the above-mentioned first aspect, or perform the decoding described in the above-mentioned second aspect method steps.
  • a computer program is provided, and when the computer program is executed, the steps of the encoding method described in the above-mentioned first aspect are realized, or the steps of the decoding method described in the above-mentioned second aspect are realized.
  • a ninth aspect provides a computer-readable storage medium, where the computer-readable storage medium includes the code stream obtained by the encoding method described in the first aspect.
  • the first latent variable is adjusted by the first variable adjustment factor to obtain the second latent variable, and the number of encoded bits of the entropy coding result of the second latent variable satisfies the preset coding rate condition, which can ensure that the potential of each frame of media data corresponds to
  • the number of encoded bits of the entropy encoding result of the variable can meet the preset encoding rate conditions, that is, it can be ensured that the number of encoded bits of the entropy encoding result of the latent variable corresponding to each frame of media data is basically consistent, rather than dynamically changing. Therefore, the encoder's requirement for a stable encoding rate is met.
  • the need to transmit side information (such as window type, time-domain noise shaping (TNS: Temporal Noise Shaping) parameters, frequency-domain noise shaping (FDNS: Frequency-domain noise shaping) parameters, and/or bandwidth extension (BWE : bandwidth extension) parameter, etc.), it can ensure that the number of encoded bits of the entropy encoding result of the latent variable corresponding to each frame of media data and the encoded bit number of the side information are basically consistent, so as to meet the encoder's need for a stable encoding rate .
  • TMS Temporal Noise Shaping
  • FDNS Frequency-domain noise shaping
  • BWE bandwidth extension
  • FIG. 1 is a schematic diagram of an implementation environment provided by an embodiment of the present application.
  • FIG. 2 is a schematic diagram of an implementation environment of a terminal scenario provided by an embodiment of the present application.
  • FIG. 3 is a schematic diagram of an implementation environment of a transcoding scenario of a wireless or core network device provided in an embodiment of the present application;
  • FIG. 4 is a schematic diagram of an implementation environment of a broadcast television scene provided by an embodiment of the present application.
  • FIG. 5 is a schematic diagram of an implementation environment of a virtual reality streaming scene provided by an embodiment of the present application.
  • Fig. 6 is a flow chart of the first encoding method provided by the embodiment of the present application.
  • Fig. 7 is a schematic diagram of the form of the first latent variable provided by the embodiment of the present application.
  • Fig. 8 is a schematic diagram of the form of the second latent variable provided by the embodiment of the present application.
  • FIG. 9 is a flow chart of the first decoding method provided by the embodiment of the present application.
  • Fig. 10 is a flow chart of the second encoding method provided by the embodiment of the present application.
  • Fig. 11 is a flow chart of the second decoding method provided by the embodiment of the present application.
  • Fig. 12 is an exemplary block diagram of an encoding method shown in Fig. 10 provided by an embodiment of the present application;
  • FIG. 13 is an exemplary block diagram of a decoding method shown in FIG. 11 provided by an embodiment of the present application.
  • Fig. 14 is a flow chart of the third encoding method provided by the embodiment of the present application.
  • FIG. 15 is a flow chart of a third decoding method provided by an embodiment of the present application.
  • Fig. 16 is an exemplary block diagram of an encoding method shown in Fig. 14 provided by an embodiment of the present application;
  • FIG. 17 is an exemplary block diagram of a decoding method shown in FIG. 15 provided by an embodiment of the present application.
  • FIG. 18 is a schematic structural diagram of an encoding device provided by an embodiment of the present application.
  • FIG. 19 is a schematic structural diagram of a decoding device provided by an embodiment of the present application.
  • Fig. 20 is a schematic block diagram of a codec device provided by an embodiment of the present application.
  • Encoding refers to the process of compressing the media data to be encoded into a code stream.
  • the media data to be encoded mainly includes audio signals, video signals and images.
  • the encoding of the audio signal is the process of compressing the audio frame sequence included in the audio signal to be encoded into a code stream.
  • the encoding of the video signal is the process of compressing the image sequence included in the video to be encoded into a code stream.
  • the encoding of the image is The process of compressing the image to be encoded into a code stream.
  • the media data after the media data is compressed into a code stream, it may be referred to as encoded media data or compressed media data.
  • encoded media data For example, for an audio signal, after the audio signal is compressed into a bit stream, it can be called a coded audio signal or a compressed audio signal, and after a video signal is compressed into a bit stream, it can also be called a coded video signal or a compressed audio signal.
  • a compressed video signal after an image is compressed into a code stream, may also be referred to as a coded image or a compressed image.
  • Decoding refers to the process of restoring the coded stream into reconstructed media data according to specific grammatical rules and processing methods. Among them, the decoding of the audio code stream refers to the processing process of restoring the audio code stream into a reconstructed audio signal, the decoding of the video code stream refers to the processing process of restoring the video code stream into a reconstructed video signal, and the decoding of the image code stream refers to converting The image code stream is restored to the process of reconstructing the image.
  • Entropy coding refers to the coding that does not lose any information according to the principle of entropy during the coding process. That is, a lossless data compression method. Entropy coding is coded based on the occurrence probability of an element, that is, for the same element, when the occurrence probability of the element is different, the number of coded bits of the entropy coding result of the element is different. Entropy coding usually includes arithmetic coding (arithmetic coding), range coding (range coding, RC), Huffman (huffman) coding and so on.
  • Constant bit rate refers to the encoding bit rate is a fixed value, for example, the fixed value is the target encoding bit rate.
  • Variable bit rate means that the encoding rate can exceed the target encoding bit rate, or can be less than the target encoding bit rate, but the difference between the target encoding bit rate and the target encoding bit rate is small.
  • FIG. 1 is a schematic diagram of an implementation environment provided by an embodiment of the present application.
  • the implementation environment includes source device 10 , destination device 20 , link 30 and storage device 40 .
  • the source device 10 may generate encoded media data. Therefore, the source device 10 may also be called a media data encoding device.
  • Destination device 20 may decode the encoded media data generated by source device 10 . Accordingly, destination device 20 may also be referred to as a media data decoding device.
  • Link 30 may receive encoded media data generated by source device 10 and may transmit the encoded media data to destination device 20 .
  • the storage device 40 can receive the encoded media data generated by the source device 10, and can store the encoded media data.
  • the destination device 20 can directly obtain the encoded media from the storage device 40.
  • the storage device 40 may correspond to a file server or another intermediate storage device that may save encoded media data generated by the source device 10, in which case the destination device 20 may transmit or download the media data from the storage device 40 via streaming or downloading. Stored encoded media data.
  • Both the source device 10 and the destination device 20 may include one or more processors and a memory coupled to the one or more processors, and the memory may include random access memory (random access memory, RAM), read-only memory ( read-only memory, ROM), charged erasable programmable read-only memory (electrically erasable programmable read-only memory, EEPROM), flash memory, can be used to store the desired program in the form of instructions or data structures that can be accessed by the computer Any other media etc. of the code.
  • RAM random access memory
  • ROM read-only memory
  • EEPROM electrically erasable programmable read-only memory
  • flash memory can be used to store the desired program in the form of instructions or data structures that can be accessed by the computer Any other media etc. of the code.
  • both source device 10 and destination device 20 may include desktop computers, mobile computing devices, notebook (e.g., laptop) computers, tablet computers, set-top boxes, telephone handsets such as so-called “smart" phones, Televisions, cameras, display devices, digital media players, video game consoles, in-vehicle computers, or the like.
  • Link 30 may include one or more media or devices capable of transmitting encoded media data from source device 10 to destination device 20 .
  • link 30 may include one or more communication media that enable source device 10 to transmit encoded media data directly to destination device 20 in real-time.
  • the source device 10 may modulate the encoded media data based on a communication standard, such as a wireless communication protocol, etc., and may send the modulated media data to the destination device 20 .
  • the one or more communication media may include wireless and/or wired communication media, for example, the one or more communication media may include radio frequency (radio frequency, RF) spectrum or one or more physical transmission lines.
  • the one or more communication media may form part of a packet-based network, such as a local area network, a wide area network, or a global network (eg, the Internet), among others.
  • the one or more communication media may include routers, switches, base stations, or other devices that facilitate communication from the source device 10 to the destination device 20, etc., which are not specifically limited in this embodiment of the present application.
  • the storage device 40 may store the received encoded media data sent by the source device 10 , and the destination device 20 may directly obtain the encoded media data from the storage device 40 .
  • the storage device 40 may include any one of a variety of distributed or locally accessed data storage media, for example, any one of the various distributed or locally accessed data storage media may be Hard disk drive, Blu-ray Disc, digital versatile disc (DVD), compact disc read-only memory (CD-ROM), flash memory, volatile or nonvolatile memory, or Any other suitable digital storage medium for storing encoded media data, etc.
  • the storage device 40 may correspond to a file server or another intermediate storage device that may save the encoded media data generated by the source device 10, and the destination device 20 may transmit or download the storage device via streaming or downloading. 40 stored media data.
  • the file server may be any type of server capable of storing encoded media data and sending the encoded media data to destination device 20 .
  • the file server may include a network server, a file transfer protocol (file transfer protocol, FTP) server, a network attached storage (network attached storage, NAS) device, or a local disk drive.
  • Destination device 20 may obtain encoded media data over any standard data connection, including an Internet connection.
  • Any standard data connection may include a wireless channel (e.g., a Wi-Fi connection), a wired connection (e.g., a digital subscriber line (DSL), cable modem, etc.), or is suitable for obtaining encoded data stored on a file server.
  • a wireless channel e.g., a Wi-Fi connection
  • a wired connection e.g., a digital subscriber line (DSL), cable modem, etc.
  • DSL digital subscriber line
  • cable modem etc.
  • the transmission of encoded media data from storage device 40 may be a streaming transmission, a download transmission, or a combination of both.
  • the implementation environment shown in Figure 1 is only a possible implementation, and the technology of the embodiment of the present application is not only applicable to the source device 10 shown in Figure 1 that can encode media data, but also can encode the encoded media
  • the destination device 20 for decoding data may also be applicable to other devices capable of encoding media data and decoding encoded media data, which is not specifically limited in this embodiment of the present application.
  • the source device 10 includes a data source 120 , an encoder 100 and an output interface 140 .
  • output interface 140 may include a conditioner/demodulator (modem) and/or a transmitter, where a transmitter may also be referred to as a transmitter.
  • Data source 120 may include an image capture device (e.g., video camera, etc.), an archive containing previously captured media data, a feed interface for receiving media data from a media data content provider, and/or a computer for generating media data graphics system, or a combination of these sources of media data.
  • the data source 120 may send media data to the encoder 100, and the encoder 100 may encode the received media data sent by the data source 120 to obtain encoded media data.
  • An encoder may send encoded media data to an output interface.
  • source device 10 sends the encoded media data directly to destination device 20 via output interface 140 .
  • encoded media data may also be stored on storage device 40 for later retrieval by destination device 20 for decoding and/or display.
  • the destination device 20 includes an input interface 240 , a decoder 200 and a display device 220 .
  • input interface 240 includes a receiver and/or a modem.
  • the input interface 240 can receive the encoded media data via the link 30 and/or from the storage device 40, and then send it to the decoder 200, and the decoder 200 can decode the received encoded media data to obtain the decoded media data. media data.
  • the decoder may transmit the decoded media data to the display device 220 .
  • the display device 220 may be integrated with the destination device 20 or may be external to the destination device 20 . In general, the display device 220 displays the decoded media data.
  • the display device 220 can be any type of display device in various types, for example, the display device 220 can be a liquid crystal display (liquid crystal display, LCD), a plasma display, an organic light-emitting diode (organic light-emitting diode, OLED) monitor or other type of display device.
  • the display device 220 can be a liquid crystal display (liquid crystal display, LCD), a plasma display, an organic light-emitting diode (organic light-emitting diode, OLED) monitor or other type of display device.
  • encoder 100 and decoder 200 may be individually integrated with the encoder and decoder, and may include appropriate multiplexer-demultiplexer (multiplexer-demultiplexer) , MUX-DEMUX) unit or other hardware and software for encoding both audio and video in a common data stream or in separate data streams.
  • the MUX-DEMUX unit may conform to the ITU H.223 multiplexer protocol, or other protocols such as user datagram protocol (UDP), if applicable.
  • Each of the encoder 100 and the decoder 200 can be any one of the following circuits: one or more microprocessors, digital signal processing (digital signal processing, DSP), application specific integrated circuit (application specific integrated circuit, ASIC) ), field-programmable gate array (FPGA), discrete logic, hardware, or any combination thereof. If the techniques of the embodiments of the present application are implemented partially in software, the device may store instructions for the software in a suitable non-transitory computer-readable storage medium, and may use one or more processors in hardware The instructions are executed to implement the technology of the embodiments of the present application. Any of the foregoing (including hardware, software, a combination of hardware and software, etc.) may be considered to be one or more processors. Each of encoder 100 and decoder 200 may be included in one or more encoders or decoders, either of which may be integrated into a combined encoding in a corresponding device Part of a codec/decoder (codec).
  • codec codec/decoder
  • Embodiments of the present application may generally refer to the encoder 100 as “signaling” or “sending” certain information to another device such as the decoder 200 .
  • the term “signaling” or “sending” may generally refer to the transmission of syntax elements and/or other data for decoding compressed media data. This transfer can occur in real time or near real time. Alternatively, this communication may occur after a period of time, such as upon encoding when storing syntax elements in an encoded bitstream to a computer-readable storage medium, which the decoding device may then perform after the syntax elements are stored on this medium The syntax element is retrieved at any time.
  • the encoding and decoding methods provided in the embodiments of the present application can be applied to various scenarios. Next, taking the media data to be encoded as an audio signal as an example, several scenarios will be introduced respectively.
  • FIG. 2 is a schematic diagram of an implementation environment in which a codec method provided by an embodiment of the present application is applied to a terminal scenario.
  • the implementation environment includes a first terminal 101 and a second terminal 201 , and the first terminal 101 and the second terminal 201 are connected in communication.
  • the communication connection may be a wireless connection or a wired connection, which is not limited in this embodiment of the present application.
  • the first terminal 101 may be a sending end device or a receiving end device.
  • the second terminal 201 may be a receiving end device or a sending end device.
  • the first terminal 101 is a sending end device
  • the second terminal 201 is a receiving end device
  • the second terminal 201 is a sending end device.
  • the first terminal 101 may be the source device 10 in the implementation environment shown in FIG. 1 above.
  • the second terminal 201 may be the destination device 20 in the implementation environment shown in FIG. 1 above.
  • both the first terminal 101 and the second terminal 201 include an audio collection module, an audio playback module, an encoder, a decoder, a channel encoding module and a channel decoding module.
  • the audio collection module in the first terminal 101 collects the audio signal and transmits it to the encoder, and the encoder encodes the audio signal using the encoding method provided in the embodiment of the present application, and the encoding may be referred to as information source encoding.
  • the channel coding module needs to perform channel coding again, and then transmit the coded stream through the wireless or wired network communication device in the digital channel.
  • the second terminal 201 receives the code stream transmitted in the digital channel through a wireless or wired network communication device, the channel decoding module performs channel decoding on the code stream, and then the decoder uses the decoding method provided by the embodiment of the application to decode the audio signal, and then passes the audio Playback module to play.
  • the first terminal 101 and the second terminal 201 can be any electronic product that can interact with the user through one or more ways such as keyboard, touch pad, touch screen, remote control, voice interaction or handwriting equipment, etc.,
  • personal computer personal computer, PC
  • mobile phone smart phone
  • personal digital assistant Personal Digital Assistant, PDA
  • wearable device PPC (pocket PC)
  • tablet computer smart car machine, smart TV, smart speaker Wait.
  • FIG. 3 is a schematic diagram of an implementation environment in which a codec method provided by an embodiment of the present application is applied to a transcoding scenario of a wireless or core network device.
  • the implementation environment includes a channel decoding module, an audio decoder, an audio encoder and a channel encoding module.
  • the audio decoder may be a decoder using the decoding method provided in the embodiment of the present application, or may be a decoder using other decoding methods.
  • the audio encoder may be an encoder using the encoding method provided by the embodiment of the present application, or may be an encoder using other encoding methods.
  • the audio encoder is a coder using other encoding methods
  • the audio The encoder is an encoder using the encoding method provided by the embodiment of the present application.
  • the audio decoder is a decoder using the decoding method provided by the embodiment of the present application, and the audio encoder is an encoder using other encoding methods.
  • the channel decoding module is used to perform channel decoding on the received code stream, and then the audio decoder is used to use the decoding method provided by the embodiment of the application to perform source decoding, and then the audio encoder is used to encode according to other encoding methods to achieve a
  • the conversion from one format to another is known as transcoding. After that, it is sent after channel coding.
  • the audio decoder is a decoder using other decoding methods
  • the audio encoder is an encoder using the encoding method provided by the embodiment of the present application.
  • the channel decoding module is used to perform channel decoding on the received code stream, and then the audio decoder is used to use other decoding methods to perform source decoding, and then the audio encoder uses the encoding method provided by the embodiment of the application to perform encoding to realize a
  • the conversion from one format to another is known as transcoding. After that, it is sent after channel coding.
  • the wireless device may be a wireless access point, a wireless router, a wireless connector, and the like.
  • a core network device may be a mobility management entity, a gateway, and the like.
  • FIG. 4 is a schematic diagram of an implementation environment in which a codec method provided by an embodiment of the present application is applied to a broadcast television scene.
  • the broadcast TV scene is divided into a live scene and a post-production scene.
  • the implementation environment includes a live program 3D sound production module, a 3D sound encoding module, a set-top box and a speaker group, and the set-top box includes a 3D sound decoding module.
  • the implementation environment includes post-program 3D sound production modules, 3D sound coding modules, network receivers, mobile terminals, earphones, etc.
  • the three-dimensional sound production module of the live program produces a three-dimensional sound signal
  • the three-dimensional sound signal is encoded by applying the coding method of the embodiment of the application to obtain a code stream
  • the code stream is transmitted to the user side through the radio and television network, and is transmitted by the set-top box
  • the 3D sound decoder uses the decoding method provided by the embodiment of the present application to perform decoding, so as to reconstruct the 3D sound signal, which is played back by the speaker group.
  • the code stream is transmitted to the user side through the Internet, and the 3D sound decoder in the network receiver decodes it using the decoding method provided by the embodiment of the present application, so as to reconstruct the 3D sound signal and play it back by the speaker group.
  • the code stream is transmitted to the user side through the Internet, and the 3D sound decoder in the mobile terminal decodes it using the decoding method provided by the embodiment of the present application, thereby reconstructing the 3D sound signal, and playing it back by the earphone.
  • the post-program 3D sound production module produces a 3D sound signal
  • the 3D sound signal is encoded by applying the coding method of the embodiment of the application to obtain a code stream
  • the code stream is transmitted to the user side through the radio and television network, and is transmitted by the set-top box
  • the 3D sound decoder uses the decoding method provided by the embodiment of the present application to decode, thereby reconstructing the 3D sound signal, which is played back by the speaker group.
  • the code stream is transmitted to the user side through the Internet, and the 3D sound decoder in the network receiver decodes it using the decoding method provided by the embodiment of the present application, so as to reconstruct the 3D sound signal and play it back by the speaker group.
  • the code stream is transmitted to the user side through the Internet, and the 3D sound decoder in the mobile terminal decodes it using the decoding method provided by the embodiment of the present application, thereby reconstructing the 3D sound signal, and playing it back by the earphone.
  • FIG. 5 is a schematic diagram of an implementation environment in which a codec method provided by an embodiment of the present application is applied to a virtual reality streaming scene.
  • the implementation environment includes an encoding end and a decoding end.
  • the encoding end includes an acquisition module, a preprocessing module, an encoding module, a packaging module and a sending module
  • the decoding end includes an unpacking module, a decoding module, a rendering module and earphones.
  • the acquisition module collects audio signals, and then performs preprocessing operations through the preprocessing module.
  • the preprocessing operations include filtering out the low-frequency part of the signal, usually with 20Hz or 50Hz as the cut-off point, and extracting the orientation information in the signal.
  • use the encoding module to perform encoding processing using the encoding method provided by the embodiment of the present application.
  • After encoding use the packing module to pack and send to the decoding end through the sending module.
  • the unpacking module at the decoding end first unpacks, and then uses the decoding method provided by the embodiment of the application to decode through the decoding module, and then performs binaural rendering processing on the decoded signal through the rendering module, and the rendered signal is mapped to the listener's earphones superior.
  • the earphone can be an independent earphone, or an earphone on a virtual reality glasses device.
  • any of the following encoding methods may be executed by the encoder 100 in the source device 10 .
  • Any of the following decoding methods may be performed by the decoder 200 in the destination device 20 .
  • the embodiments of the present application may be applied to a codec that does not include a context model, and may also be applied to a codec that includes a context model.
  • the adjustment factor and the variable adjustment factor involved in the embodiment of the present application may be a value before quantization or a value after quantization, which is not limited in the embodiment of the present application.
  • FIG. 6 is a flowchart of the first encoding method provided by the embodiment of the present application.
  • the method does not include a context model, only the latent variables generated by the media data to be encoded are adjusted by an adjustment factor.
  • the encoding method is applied to an encoding end device, and includes the following steps.
  • Step 601 Process the media data to be encoded through the first encoding neural network model to obtain a first latent variable, and the first latent variable is used to indicate the characteristics of the media data to be encoded.
  • the media data to be encoded is an audio signal, a video signal, or an image.
  • the form of the media data to be encoded may be in any form, which is not limited in this embodiment of the present application.
  • the media data to be encoded may be media data in the time domain, or media data in the frequency domain obtained after time-frequency transformation of the media data in the time domain, for example, it may be media data in the time domain after MDCT transformation
  • the obtained media data in the frequency domain, or the media data in the frequency domain obtained after the media data in the time domain undergoes fast Fourier transformation (FFT).
  • FFT fast Fourier transformation
  • the media data to be encoded can also be the media data in the complex frequency domain obtained after the media data in the time domain is filtered by an orthogonal mirror filter (quandrature mirror filter, QMF), or the media data to be encoded is media data in the time domain.
  • the characteristic signal obtained by data extraction, such as the Mel cepstral coefficient, or the media data to be encoded can also be a residual signal, such as the residual signal of other codes or the residual after filtering by linear predictive coding (LPC) Signal.
  • LPC linear predictive coding
  • the implementation process of processing the media data to be encoded by the first encoding neural network model is: input the media data to be encoded into the first encoding neural network model, and obtain the first latent variable output by the first encoding neural network model.
  • the media data to be encoded is preprocessed, and the preprocessed media data is input into the first encoding neural network model to obtain the first latent variable output by the first encoding neural network model.
  • the media data to be encoded can be used as the input of the first encoding neural network model to determine the first latent variable, and the media data to be encoded can also be preprocessed and then used as the input of the first encoding neural network model Determine the first latent variable.
  • the preprocessing operation may be temporal noise shaping (temporal noise shaping, TNS) processing, frequency domain noise shaping (frequency domain noise shaping, FDNS) processing, channel downmixing processing, and the like.
  • the first encoding neural network model is pre-trained, and the embodiment of the present application does not limit the network structure and training method of the first encoding neural network model.
  • the network structure of the first encoding neural network model may be a fully connected network or a convolutional neural network (convolutional neural network, CNN) network.
  • CNN convolutional neural network
  • the embodiment of the present application does not limit the number of layers included in the network structure of the first coding neural network model and the number of nodes in each layer.
  • the form of latent variables output by encoding neural network models with different network structures may be different.
  • the first latent variable is a vector
  • the dimension M of the vector is the size of the latent variable (latent size), as shown in FIG. 7 .
  • the first latent variable is an N*M dimensional matrix, where N is the number of channels (channels) of the CNN network, and M is the potential of each channel of the CNN network
  • the size of the variable (latent size) as shown in Figure 8. It should be noted that Figure 7 and Figure 8 only give an illustration of the latent variables of the fully connected network and the latent variables of the CNN network. The channel number can be counted from 1 or 0, and the latent variables in each channel The same is true for the element numbers.
  • Step 602 Determine the first variable adjustment factor based on the first latent variable.
  • the first variable adjustment factor is used to make the number of encoded bits of the entropy coding result of the second latent variable meet the preset coding rate condition.
  • the second latent variable is obtained through the first
  • the variable adjustment factor is obtained after adjusting the first latent variable.
  • the initial number of coding bits may be determined based on the first latent variable, and the first variable adjustment factor may be determined based on the initial number of coding bits and the target number of coding bits.
  • the target number of coding bits may be set in advance.
  • the target number of coding bits may also be determined based on the coding rate, and different coding rates correspond to different target coding bits.
  • the media data to be encoded may be encoded using a fixed code rate, or the media data to be encoded may be encoded using a variable code rate.
  • the number of bits of the media data to be encoded in the current frame can be determined based on the fixed bit rate, and then the number of used bits in the current frame can be subtracted to obtain the target of the current frame
  • the number of encoded bits may be the number of bits for encoding side information, etc., and usually, the side information of each frame of media data is different, so the target number of encoding bits of each frame of media data is usually different.
  • the number of bits of media data to be encoded in the current frame can be determined based on the specified code rate, and then the number of used bits in the current frame can be subtracted to obtain the target number of encoded bits in the current frame.
  • the number of used bits can be the number of bits encoded by side information, and in some cases, the side information of media data of different frames can be different, so the target number of encoded bits of media data of different frames is usually is different.
  • the manner of determining the number of initial coding bits based on the first latent variable may include two manners, which will be introduced respectively next.
  • the number of coded bits of the entropy coding result of the first latent variable is determined to obtain the initial number of coded bits. That is, the initial number of coding bits is the number of coding bits of the entropy coding result of the first latent variable.
  • the quantization process is performed on the first latent variable to obtain the quantized first latent variable.
  • Entropy coding is performed on the quantized first latent variable to obtain an initial coding result of the first latent variable.
  • the number of coding bits of the initial coding result of the first latent variable is counted to obtain the number of coding bits of the initial coding.
  • the manner of performing quantization processing on the first latent variable may include multiple manners, for example, performing scalar quantization on each element in the first latent variable.
  • the quantization step size of scalar quantization can be determined based on different coding rates, that is, the corresponding relationship between the coding rate and the quantization step size is stored in advance, and the corresponding quantization can be obtained from the corresponding relationship based on the coding rate adopted in the embodiment of the present application. step size.
  • scalar quantization may also have an offset, that is, the first latent variable is biased by the offset and then scalar quantization is performed according to the quantization step size.
  • entropy coding When performing entropy coding on the quantized first latent variable, entropy coding based on an adjustable entropy coding model may be used, or an entropy coding model with a preset probability distribution may be used to perform entropy coding, which is not limited in this embodiment of the present application.
  • the entropy coding may adopt one of arithmetic coding (arithmetic coding), range coding (range coding, RC) or Huffman (huffman) coding, which is not limited in this embodiment of the present application.
  • quantization processing method and entropy coding method are similar to those here, and the following quantization processing method and entropy coding method can refer to the method here, and the embodiments of the present application will not repeat them hereafter.
  • the first latent variable is adjusted based on the first initial adjustment factor, and the number of encoded bits of the entropy encoding result of the adjusted first latent variable is determined to obtain the initial number of encoded bits. That is, the initial number of coding bits is the number of coding bits of the entropy coding result of the first latent variable adjusted by the first initial adjustment factor.
  • the first initial adjustment factor may be a first preset adjustment factor.
  • the first latent variable is adjusted based on the first initial adjustment factor to obtain an adjusted first latent variable, and the adjusted first latent variable is quantized to obtain a quantized first latent variable.
  • Entropy coding is performed on the quantized first latent variable to obtain an initial coding result of the first latent variable.
  • the number of coding bits of the initial coding result of the first latent variable is counted to obtain the number of coding bits of the initial coding.
  • the implementation process of adjusting the first latent variable based on the first initial adjustment factor is: multiply each element in the first latent variable by the corresponding element in the first initial adjustment factor to obtain the adjusted first latent variable .
  • each element in the first latent variable may be divided by the corresponding element in the first initial adjustment factor to obtain the adjusted first latent variable.
  • the embodiment of the present application does not limit the adjustment method.
  • an initial value of the adjustment factor is set for the first latent variable, and the initial value of the adjustment factor is usually equal to 1.
  • the first preset adjustment factor may be greater than or equal to the initial value of the adjustment factor, and may also be smaller than the initial value of the adjustment factor.
  • the first preset adjustment factor is a constant such as 1 or 2.
  • the first initial adjustment factor is the initial value of the adjustment factor; in the case of determining the initial number of encoded bits through the above-mentioned second implementation, the first initial The adjustment factor is a first preset adjustment factor.
  • the first initial adjustment factor may be a scalar or a vector.
  • the first latent variable output by it is a vector
  • the dimension M of the vector is the size of the first latent variable (latent size). If the first initial adjustment factor is a scalar, in this case, the adjustment factor value corresponding to each element in the first latent variable whose dimension is M and is a vector is the same, that is, the first initial adjustment factor includes one element.
  • the adjustment factor value corresponding to each element in the first latent variable vector whose dimension is M and is a vector is not exactly the same, and multiple elements can share one adjustment factor
  • the value, that is, the first initial adjustment factor includes a plurality of elements, and each element corresponds to one or more elements in the first latent variable.
  • the first potential variable output by it is an N*M dimensional matrix, where N is the channel number (channel) of the CNN network, and M is each channel of the CNN network.
  • the adjustment factor values corresponding to each element in the N*M-dimensional first latent variable matrix are not exactly the same, and the elements of the latent variables that can belong to the same channel correspond to
  • the same adjustment factor value, that is, the first initial adjustment factor includes N elements, and each element corresponds to M elements with the same channel number in the first latent variable.
  • the above-mentioned first variable adjustment factor is the final value of the adjustment factor of the first latent variable.
  • the second latent variable is obtained after the first latent variable is adjusted by the first variable adjustment factor, and the number of coded bits of the entropy coding result of the second latent variable satisfies a preset coding rate condition.
  • satisfying the preset encoding rate condition includes that the number of encoding bits is less than or equal to the target number of encoding bits.
  • satisfying the preset coding rate condition includes that the number of coding bits is less than or equal to the target number of coding bits, and the difference between the number of coding bits and the target number of coding bits is smaller than a threshold value of the number of bits.
  • satisfying the preset encoding rate condition includes that the absolute value of the difference between the number of encoded bits and the target number of encoded bits is smaller than the bit number threshold. That is, satisfying the preset encoding rate condition includes that the number of encoded bits is less than or equal to the target number of encoded bits, and the difference between the target number of encoded bits and the number of encoded bits is smaller than the bit number threshold.
  • satisfying the preset encoding rate condition includes that the number of encoded bits is greater than or equal to the target number of encoded bits, and the difference between the number of encoded bits and the target number of encoded bits is smaller than a threshold value of the number of bits.
  • bit number threshold can be set in advance, and the bit number threshold can be adjusted based on different requirements.
  • the first initial adjustment factor is determined as the first variable adjustment factor. If the initial number of encoded bits is not equal to the target number of encoded bits, the first variable adjustment factor is determined in a first round-robin manner based on the initial number of encoded bits and the target number of encoded bits.
  • the i-th cycle processing of the first cycle method includes the following steps: determine the adjustment factor of the i-th cycle process, i is a positive integer, and adjust the first latent variable based on the adjustment factor of the i-th cycle process to obtain The first latent variable after the i-th adjustment. Determine the number of coded bits of the entropy coding result of the first latent variable after the i-th adjustment, so as to obtain the number of coded bits of the i-th time. In the case that the number of coded bits of the ith time satisfies the continuous adjustment condition, the (i+1)th loop processing of the first loop mode is executed. In the case that the number of encoded bits of the i-th time does not satisfy the condition for continuing adjustment, the execution of the first loop mode is terminated, and the first variable adjustment factor is determined based on the adjustment factor of the i-th loop process.
  • the realization process of determining the adjustment factor of the ith loop processing is: based on the adjustment factor of the i-1th loop processing in the first loop mode, the number of coded bits of the i-1th time, and the target number of coded bits, determine the Adjustment factor for i-cycle processing.
  • the adjustment factor of the i-1 th round of processing is the first initial adjustment factor
  • the number of encoded bits of the i-1 th time is the initial number of encoded bits.
  • the continuous adjustment condition includes that the number of coded bits of the i-1th time and the number of coded bits of the ith time are both smaller than the target number of coded bits, or the continued adjustment condition includes the number of coded bits of the i-1th time and the number of coded bits of the ith time
  • the number of encoded bits is greater than the target number of encoded bits.
  • the condition for continuing to adjust includes that the number of coded bits for the ith time does not exceed the target number of coded bits.
  • the meaning of not crossing here means: the number of encoded bits for the first i-1 times is always smaller than the target number of encoded bits, and the number of encoded bits for the ith time is still smaller than the target number of encoded bits.
  • the number of encoded bits for the first i-1 times is always greater than the target number of encoded bits, and the number of encoded bits for the ith time is still greater than the target number of encoded bits.
  • overstepping means that the number of encoded bits for the first i-1 times is always smaller than the target number of encoded bits, and the number of encoded bits for the ith time is greater than the target number of encoded bits.
  • the number of encoded bits for the first i-1 times is always greater than the target number of encoded bits, and the number of encoded bits for the ith time is smaller than the target number of encoded bits.
  • the adjustment factor is a quantized value
  • the adjustment factor, the number of coded bits of the i-1th time, and the target number of coded bits may be based on the adjustment factor of the i-1th round of processing in the first round-robin manner
  • the adjustment factor for the ith cycle processing is determined according to the following formula (1).
  • scale(i) refers to the adjustment factor of the ith cycle processing
  • scale(i-1) refers to the adjustment factor of the i-1th cycle processing
  • target refers to the target coding bit curr(i-1) refers to the number of coded bits of the i-1th time.
  • i is a positive integer greater than 0.
  • Q ⁇ x ⁇ refers to obtaining the quantized value of x.
  • the adjustment factor is the value before quantization
  • the adjustment factor of the i-th cycle is determined by the above formula (1)
  • the right side of the equation of the above formula (1) can be used without Q ⁇ x ⁇ to deal with.
  • the implementation process of determining the first variable adjustment factor based on the adjustment factor of the i-th cyclic processing includes: in the case that the i-th coding bit number is equal to the target coding bit number, determining the i-th cyclic processing adjustment factor as First variable adjustment factor. Or, in the case that the number of coded bits of the i-th time is not equal to the target number of coded bits, the first variable adjustment factor is determined based on the adjustment factor of the i-th cyclic process and the adjustment factor of the i-1 th cyclic process.
  • the adjustment factor of the i-th loop processing is the adjustment factor obtained last time through the above-mentioned first round-robin manner, and the number of coded bits of the i-th time is the number of coded bits obtained last time.
  • the last obtained adjustment factor is determined as the first variable adjustment factor.
  • the first variable adjustment factor is determined based on the adjustment factors obtained twice later.
  • the implementation process of determining the first variable adjustment factor includes: determining the adjustment factor of the i-th cycle processing and the i-th cycle processing - an average value of the adjustment factors of 1 cycle treatment, based on which average value the first variable adjustment factor is determined.
  • the average value may be directly determined as the first variable adjustment factor, or the average value may be multiplied by a preset constant to obtain the first variable adjustment factor.
  • this constant can be less than 1.
  • the realization process of determining the first variable adjustment factor includes: based on the adjustment factor of the i th cyclic processing and the first The adjustment factor for i-1 times of cyclic processing, the first variable adjustment factor is determined through the second cyclic manner.
  • the j-th loop processing of the second loop manner includes the following steps: determining the third adjustment factor of the j-th loop processing based on the first adjustment factor of the j-th loop processing and the second adjustment factor of the j-th loop processing.
  • Adjustment factor wherein, in the case that j is equal to 1, the first adjustment factor of the jth cycle processing is one of the adjustment factor of the i-th cycle processing and the adjustment factor of the i-1th cycle processing, and the jth cycle processing
  • the second adjustment factor of the second cycle processing is the other one of the adjustment factor of the i-th cycle processing and the adjustment factor of the i-1th cycle processing, and the first adjustment factor of the j-th cycle processing corresponds to the j-th cycle processing.
  • the number of coded bits, the second adjustment factor of the jth cyclic process corresponds to the second coded bit number of the jth time, and the first coded bit number of the jth time refers to the adjusted first adjustment factor of the jth cyclic process
  • the number of coded bits of the entropy coding result of the first latent variable, the second coded bit number of the jth time refers to the coded bits of the entropy coding result of the first latent variable adjusted by the second adjustment factor of the jth cycle processing
  • the number of first coded bits of the jth time is less than the number of second coded bits of the jth time.
  • the third number of encoded bits of the jth time refers to the number of encoded bits of the entropy coding result of the first latent variable adjusted by the third adjustment factor of the jth cycle processing. If the number of third encoded bits of the jth time does not satisfy the condition for continuing the loop, the execution of the second loop mode is terminated, and the third adjustment factor of the jth loop process is determined as the first variable adjustment factor.
  • the third coding bit number of the jth time satisfies the continuation loop condition, the third coding bit number of the jth time is greater than the target coding bit number and less than the second coding bit number of the jth time, the third coded bit number of the jth time loop processing
  • the adjustment factor is used as the second adjustment factor of the j+1th round of processing, the first adjustment factor of the jth round of processing is used as the first adjustment factor of the j+1th round of processing, and the j+1th round of the second round is executed. 1 cycle processing.
  • the third encoding bit number of the jth time satisfies the continuation loop condition, the third encoding bit number of the jth time is less than the target encoding bit number and greater than the first encoding bit number of the jth time, the third encoding bit number of the jth loop processing
  • the adjustment factor is used as the first adjustment factor of the j+1th round of processing, and the second adjustment factor of the jth round of processing is used as the second adjustment factor of the j+1th round of processing, and the j+1th round of the second round is executed. 1 cycle processing.
  • the j-th loop processing in the second loop manner includes the following steps: determining the j-th loop processing based on the first adjustment factor of the j-th loop processing and the second adjustment factor of the j-th loop processing.
  • Three adjustment factors wherein, in the case that j is equal to 1, the first adjustment factor of the jth cyclic processing is one of the adjustment factor of the i th cyclic processing and the adjustment factor of the i-1 th cyclic processing, the first The second adjustment factor of the j-th cycle processing is the other one of the adjustment factor of the i-th cycle processing and the adjustment factor of the i-1-th cycle processing, and the first adjustment factor of the j-th cycle processing corresponds to the j-th cycle processing.
  • the number of first coded bits, the second adjustment factor of the jth cyclic processing corresponds to the second number of coded bits of the jth time, the number of first coded bits of the jth time is less than the number of second coded bits of the jth time, and j is positive integer.
  • the execution of the second loop mode is terminated, and the third adjustment factor of the jth loop process is determined as the first variable adjustment factor. If j reaches the maximum number of cycles and the number of third encoded bits of the jth time satisfies the continuation cycle condition, the execution of the second cycle mode is terminated, and the first variable adjustment factor is determined based on the first adjustment factor of the jth cycle processing.
  • the number of third coded bits of the jth time satisfies the condition for continuing the cycle, and the number of third coded bits of the jth time is greater than the number of target coded bits and less than the number of second coded bits of the jth time
  • the The third adjustment factor of the j cycle processing is used as the second adjustment factor of the j+1 cycle processing
  • the first adjustment factor of the j cycle processing is used as the first adjustment factor of the j+1 cycle processing
  • the first adjustment factor is executed.
  • the number of third coded bits of the jth time satisfies the condition for continuing the cycle, and the number of third coded bits of the jth time is less than the target number of coded bits and greater than the number of first coded bits of the jth time
  • the The third adjustment factor of the j loop processing is used as the first adjustment factor of the j+1 loop processing
  • the second adjustment factor of the j loop processing is used as the second adjustment factor of the j+1 loop processing
  • the first adjustment factor is executed.
  • the implementation process of determining the third adjustment factor of the j-th cyclic processing based on the first adjustment factor of the j-th cyclic processing and the second adjustment factor of the j-th cyclic processing includes: determining the first adjustment factor of the j-th cyclic processing Factor and the average value of the second adjustment factor of the j-th cycle processing, and determine the third adjustment factor of the j-th cycle processing based on the average value.
  • the average value may be directly determined as the third adjustment factor for the j-th cycle processing, or the average value may be multiplied by a preset constant to obtain the third adjustment factor for the j-th cycle processing.
  • this constant can be less than 1.
  • the implementation process of obtaining the third coded bit number of the jth time includes: adjusting the first latent variable based on the third adjustment factor of the jth cyclic process to obtain the adjusted first latent variable, and adjusting the adjusted first latent variable A latent variable is quantified to obtain a quantized first latent variable. Entropy encoding is performed on the quantized first latent variable, and the number of encoded bits of the entropy encoding result is counted to obtain the third encoded bit number of the jth time.
  • the implementation process of determining the first variable adjustment factor based on the first adjustment factor of the j-th cycle processing includes: the first adjustment factor of the j-th cycle processing The factor is determined as the first variable adjustment factor.
  • the implementation process of determining the first variable adjustment factor based on the first adjustment factor of the jth cycle processing includes: determining the target encoding bit number and the jth A first difference between the number of coded bits, and a second difference between the second number of coded bits determined for the jth time and the target number of coded bits.
  • the first adjustment factor of the j-th loop processing is determined as the first variable adjustment factor. If the second difference is smaller than the first difference, the second adjustment factor of the j-th loop processing is determined as the first variable adjustment factor. If the first difference is equal to the second difference, determine the first adjustment factor of the jth cyclic process as the first variable adjustment factor, or determine the second adjustment factor of the j th cyclic process as the first variable adjustment factor.
  • the continuation loop condition includes that the jth third encoding bit number is greater than the target encoding bit number, or the continuation loop condition includes the jth third encoding bit number
  • the number of bits is smaller than the target number of encoded bits, and the difference between the target number of encoded bits and the j-th third encoded number of bits is greater than the threshold value of the number of bits.
  • the continuation loop condition includes that the absolute value of the difference between the target encoding bit number and the j-th third encoding bit number is greater than the bit number threshold.
  • condition for continuing the loop includes that the third coded bit number of the jth time is greater than the target coded bit number, and the difference between the jth third coded bit number and the target coded bit number is greater than the bit number threshold, or the loop continues
  • the conditions include that the number of third encoded bits of the jth time is smaller than the target number of encoded bits, and the difference between the target number of encoded bits and the third number of encoded bits of the jth time is greater than a threshold value of the number of bits.
  • the continuation loop condition includes: bits_curr>target
  • the continuation loop condition includes: bits_curr>target&&(bits_curr-target)>TH
  • bits_curr refers to the number of third coded bits of the jth time
  • target refers to the number of target coded bits
  • TH is a threshold value of the number of bits.
  • the first adjustment factor of the jth cyclic processing can be recorded as scale_lower
  • the jth first encoding bit number can be recorded as bits_lower
  • the second adjustment factor of the j th cyclic processing can be recorded as scale_upper
  • the jth cyclic processing can be recorded as scale_upper.
  • the number of second coded bits of the second time is recorded as bits_upper
  • the third adjustment factor of the jth cycle processing is recorded as scale_curr
  • the third adjustment factor of the jth cycle processing is used as the first adjustment factor of the j+1th cycle processing
  • the pseudocode of the second adjustment factor for the j+1th cycle processing is as follows:
  • the first initial adjustment factor is determined as the first variable adjustment factor.
  • the first variable adjustment factor is determined according to the above first implementation manner.
  • the adjustment factor of the i-1th round of processing is adjusted according to the length of the first step to obtain the i-th round of processing
  • the condition for continuing to adjust includes that the number of coded bits for the ith time is less than the target number of coded bits.
  • the initial number of encoded bits is greater than the target number of encoded bits
  • the continuing adjustment conditions include the i
  • the second number of encoded bits is greater than the target number of encoded bits.
  • adjusting the adjustment factor of the i-1th loop processing according to the first step length may refer to increasing the adjustment factor of the i-1th loop processing according to the first step length, and adjusting the i-1th loop processing according to the second step length
  • the adjustment factor of the cyclic processing may refer to reducing the adjustment factor of the i-1 th cyclic processing according to the second step size.
  • the above-mentioned increase processing and decrease processing may be linear or non-linear.
  • the sum of the adjustment factor of the i-1th loop processing and the first step length can be determined as the adjustment factor of the i-th loop processing, and the adjustment factor of the i-1th loop processing can be combined with the second step size The difference is determined as the adjustment factor for the ith cycle processing.
  • first step size and the second step size may be set in advance, and the first step size and the second step size may be adjusted based on different requirements.
  • first step length and the second step length may be equal or not.
  • the first initial adjustment factor is determined as the first variable adjustment factor.
  • the first variable adjustment factor may be determined according to the above-mentioned first implementation manner.
  • the first variable adjustment factor may also be determined according to the situation that the initial number of coding bits is greater than the target number of coding bits in the above-mentioned second implementation manner.
  • Step 603 Obtain an entropy encoding result of the second latent variable.
  • the first variable adjustment factor may be the first initial adjustment factor, may be the adjustment factor of the i-th cycle processing, or may be the third adjustment factor of the j-th cycle processing , may also be determined based on the first adjustment factor of the j-th cycle processing.
  • the corresponding entropy encoding result is determined in the above cyclic process, so the entropy encoding result corresponding to the first variable adjustment factor can be obtained directly from the entropy encoding result determined in the above cyclic process, that is, the second Entropy encoding results of latent variables.
  • the first latent variable is directly adjusted based on the first variable adjustment factor to obtain the second latent variable.
  • Entropy coding is performed on the quantized second latent variable to obtain an entropy coding result of the second latent variable.
  • the first initial adjustment factor may be a scalar or a vector
  • the first variable adjustment factor is determined based on the first initial adjustment factor, so the first variable adjustment factor may be a scalar or a vector.
  • the first variable adjustment factor includes one element or multiple elements, and when the first variable adjustment factor includes multiple elements, one element in the first variable adjustment factor corresponds to one or more elements in the first latent variable .
  • Step 604 Write the entropy coding result of the second latent variable and the coding result of the first variable adjustment factor into the code stream.
  • the first variable adjustment factor may be a value before quantization, or a value after quantization. If the first variable adjustment factor is the value before quantization, at this time, the implementation process of determining the encoding result of the first variable adjustment factor includes: performing quantization and encoding processing on the first variable adjustment factor to obtain the encoding result of the first variable adjustment factor . If the first variable adjustment factor is a quantized value, at this time, the implementation process of determining the encoding result of the first variable adjustment factor includes: encoding the first variable adjustment factor to obtain the encoding result of the first variable adjustment factor.
  • the first variable adjustment factor may be encoded in any encoding manner, which is not limited in this embodiment of the present application.
  • the quantization process and the encoding process are performed in one process, that is, the quantization result and the encoding result can be obtained through one process. Therefore, for the first variable adjustment factor, the encoding result of the first variable adjustment factor may also be obtained during the above-mentioned process of determining the first variable adjustment factor, so the encoding result of the first variable adjustment factor can be obtained directly. That is, the encoding result of the first variable adjustment factor can also be directly obtained during the process of determining the first variable adjustment factor.
  • the embodiment of the present application may also determine a quantization index corresponding to a quantization step of the first variable adjustment factor to obtain the first quantization index.
  • a quantization index corresponding to a quantization step of the second latent variable may be determined to obtain a second quantization index. Encode the first quantization index and the second quantization index into the code stream.
  • the quantization index is used to indicate the corresponding quantization step size, that is, the first quantization index is used to indicate the quantization step size of the first variable adjustment factor, and the second quantization index is used to indicate the quantization step size of the second latent variable.
  • the first latent variable is adjusted by the first variable adjustment factor to obtain the second latent variable, and the number of encoded bits of the entropy coding result of the second latent variable satisfies the preset coding rate condition, which can ensure
  • the number of encoded bits of the entropy encoding result of the latent variable corresponding to each frame of media data can meet the preset encoding rate conditions, that is, it can be guaranteed that the number of encoded bits of the entropy encoding result of the latent variable corresponding to each frame of media data is basically consistent , rather than dynamically changing, thus meeting the encoder's need for a stable encoding rate.
  • the need to transmit side information (such as window type, time-domain noise shaping (TNS: Temporal Noise Shaping) parameters, frequency-domain noise shaping (FDNS: Frequency-domain noise shaping) parameters, and/or bandwidth extension (BWE : bandwidth extension) parameter, etc.), it can ensure that the number of encoded bits of the entropy encoding result of the latent variable corresponding to each frame of media data and the encoded bit number of the side information are basically consistent, so as to meet the encoder's need for a stable encoding rate .
  • TMS Temporal Noise Shaping
  • FDNS Frequency-domain noise shaping
  • BWE bandwidth extension
  • FIG. 9 is a flowchart of a first decoding method provided by an embodiment of the present application, and the method is applied to a decoding end. This method corresponds to the encoding method shown in FIG. 6 . The method includes the following steps.
  • Step 901 Determine the reconstructed second latent variable and the reconstructed first variable adjustment factor based on the code stream.
  • entropy decoding may be performed on the entropy encoding result of the second latent variable in the code stream, and the encoding result of the first variable adjustment factor in the code stream may be decoded to obtain the quantized second latent variable and The quantized adjustment factor for the first variable.
  • the quantized second latent variable and the quantized first variable adjustment factor are dequantized to obtain the reconstructed second latent variable and the reconstructed first variable adjustment factor.
  • the decoding method in this step corresponds to the encoding method at the encoding end
  • the dequantization processing in this step corresponds to the quantization processing at the encoding end. That is, the decoding method is the inverse process of the encoding method, and the dequantization process is the inverse process of the quantization process.
  • the first quantization index and the second quantization index can be parsed from the code stream, the first quantization index is used to indicate the quantization step size of the first variable adjustment factor, and the second quantization index is used to indicate the quantization of the second latent variable step size.
  • the quantized first variable adjustment factor is dequantized to obtain the reconstructed first variable adjustment factor.
  • the quantized second latent variable is dequantized to obtain a reconstructed second latent variable.
  • Step 902 Adjust the reconstructed second latent variable based on the reconstructed first variable adjustment factor to obtain the reconstructed first latent variable, and the reconstructed first latent variable is used to indicate the characteristics of the media data to be decoded .
  • the reconstructed second latent variable can be adjusted based on the reconstructed first variable adjustment factor to obtain the reconstructed first variable a latent variable.
  • the process of adjusting the reconstructed second latent variable based on the reconstructed first variable adjustment factor is an inverse process of adjusting the first latent variable at the encoding side. For example, in the case where the encoding side multiplies each element in the first latent variable with the corresponding element in the first variable adjustment factor, here each element in the reconstructed second latent variable can be divided by the reconstructed The first variable adjusts the corresponding elements in the factor to obtain the reconstructed first latent variable.
  • each element in the reconstructed second latent variable can be multiplied by the reconstructed first variable adjustment
  • the corresponding elements in the factor get the reconstructed first latent variable.
  • the first initial adjustment factor may be a scalar or a vector
  • the finally obtained first variable adjustment factor may be a scalar or a vector.
  • the network structure of the first encoding neural network model is a fully connected network
  • the second latent variable is a vector
  • the dimension M of the vector is the size of the second latent variable (latent size). If the first variable adjustment factor is a scalar, in this case, the value of the adjustment factor corresponding to each element in the second latent variable whose dimension is M and is a vector is the same, that is, the first variable adjustment factor includes one element.
  • the adjustment factor value corresponding to each element in the second latent variable vector whose dimension is M and is a vector is not exactly the same, and multiple elements can share one adjustment factor
  • the value, that is, the first variable adjustment factor includes a plurality of elements, each element corresponding to one or more elements of the second latent variable.
  • the network structure of the first coded neural network model is a CNN network
  • the second latent variable is an N*M dimensional matrix, where N is the number of channels of the CNN network, and M is the potential of each channel of the CNN network.
  • the size of the variable (latent size). If the first variable adjustment factor is a scalar, in this case, the value of the adjustment factor corresponding to each element in the N*M-dimensional second latent variable matrix is the same, that is, the first variable adjustment factor includes one element.
  • the adjustment factor value corresponding to each element in the N*M-dimensional second latent variable matrix is not exactly the same, and the elements of the latent variable that can belong to the same channel correspond to
  • the same adjustment factor value, that is, the first variable adjustment factor includes N elements, and each element corresponds to M elements with the same channel number in the second latent variable.
  • Step 903 Process the reconstructed first latent variable through the first decoding neural network model to obtain reconstructed media data.
  • the reconstructed first latent variable may be input into the first decoding neural network model to obtain reconstructed media data output by the first decoding neural network model.
  • perform post-processing on the reconstructed first latent variable input the post-processed first latent variable into the first decoding neural network model, and obtain reconstructed media data output by the first decoding neural network model.
  • the reconstructed first latent variable can be used as the input of the first decoding neural network model to determine the reconstructed media data, or the reconstructed first latent variable can be post-processed and then used as the first decoding The input of the neural network model to determine the reconstructed media data.
  • the first decoding neural network model corresponds to the first encoding neural network model, and both are pre-trained.
  • the embodiment of the present application does not limit the network structure and training method of the first decoding neural network model.
  • the network structure of the first decoding neural network model may be a fully connected network or a CNN network.
  • the embodiment of the present application does not limit the number of layers included in the network structure of the first decoding neural network model and the number of nodes in each layer.
  • the output of the first decoding neural network model may be reconstructed media data in the time domain or reconstructed media data in the frequency domain. If it is media data in the frequency domain, it needs to be converted from the frequency domain to the time domain to obtain the media data in the time domain.
  • the output of the first decoding neural network model may also be a residual signal. In this case, other corresponding processing is required to obtain an audio signal or a video signal.
  • the encoding rate condition is set, that is, it can ensure that the number of encoded bits of the entropy encoding result of the latent variable corresponding to each frame of media data is basically consistent, rather than dynamically changing, thereby meeting the encoder's need for a stable encoding rate.
  • the need to transmit side information (such as window type, time-domain noise shaping (TNS: Temporal Noise Shaping) parameters, frequency-domain noise shaping (FDNS: Frequency-domain noise shaping) parameters, and/or bandwidth extension (BWE : bandwidth extension) parameter, etc.), it can ensure that the number of encoded bits of the entropy encoding result of the latent variable corresponding to each frame of media data and the encoded bit number of the side information are basically consistent, so as to meet the encoder's need for a stable encoding rate .
  • TMS Temporal Noise Shaping
  • FDNS Frequency-domain noise shaping
  • BWE bandwidth extension
  • FIG. 10 is a flowchart of a second encoding method provided by an embodiment of the present application.
  • the method includes a context model, but only latent variables generated by the media data to be encoded are adjusted by adjustment factors.
  • the encoding method is applied to an encoding end device, and includes the following steps.
  • Step 1001 Process the media data to be encoded through the first encoding neural network model to obtain a first latent variable, and the first latent variable is used to indicate the characteristics of the media data to be encoded.
  • step 100 for the implementation process of step 1001, reference may be made to the implementation process of the above-mentioned step 601, which will not be repeated here.
  • Step 1002 Determine a first variable adjustment factor based on the first latent variable.
  • the initial number of coding bits may be determined based on the first latent variable
  • the first variable adjustment factor may be determined based on the initial number of coding bits and the target number of coding bits.
  • the target number of coding bits reference may be made to the content in step 602, which will not be repeated here.
  • the manner of determining the number of initial coding bits based on the first latent variable may include two manners, which will be introduced respectively next.
  • the corresponding context initial coding bit number and the initial entropy coding model parameter are determined through the context model, and the coding bit number of the entropy coding result of the first latent variable is determined based on the initial entropy coding model parameter , to get the number of basic initial coded bits.
  • the initial encoding bit number is determined based on the context initial encoding bit number and the basic initial encoding bit number.
  • the context model includes a context encoding neural network model and a context decoding neural network model.
  • the implementation process of determining the corresponding context initial encoding bits and initial entropy encoding model parameters through the context model includes: processing the first latent variable through the context encoding neural network model to obtain the fifth latent variable, the fifth The latent variable is used to indicate the probability distribution of the first latent variable.
  • the entropy coding result of the fifth latent variable is determined, and the coding bit number of the entropy coding result of the fifth latent variable is used as the initial coding bit number of the context.
  • the fifth latent variable is reconstructed based on the entropy coding result of the fifth latent variable, and the reconstructed fifth latent variable is processed through the context decoding neural network model to obtain initial entropy coding model parameters.
  • the fifth latent variable is obtained by processing the first latent variable through the context encoding neural network model.
  • the fifth latent variable is quantified to obtain the quantized fifth latent variable.
  • Entropy encoding is performed on the quantized fifth latent variable, and the number of encoded bits of the entropy encoding result is counted to obtain the initial number of encoded bits of the context.
  • entropy decoding is performed on the entropy encoding result of the fifth latent variable to obtain a quantized fifth latent variable, and dequantization processing is performed on the quantized fifth latent variable to obtain a reconstructed fifth latent variable.
  • the reconstructed fifth latent variable is input into the context decoding neural network model to obtain initial entropy encoding model parameters output by the context decoding neural network model.
  • the implementation process of processing the first latent variable through the context coding neural network model is: input the first latent variable into the context coding neural network model, and obtain the fifth latent variable output by the context coding neural network model.
  • the implementation process of processing the first latent variable through the context coding neural network model is: input the first latent variable into the context coding neural network model, and obtain the fifth latent variable output by the context coding neural network model.
  • take the absolute value of each element in the first latent variable and then input the context coding neural network model to obtain the fifth latent variable output by the context coding neural network model.
  • the number of coding bits of the entropy coding result of the first latent variable is determined, and the realization process of obtaining the basic initial coding bit number is as follows: from the entropy coding model with adjustable coding model parameters, the initial entropy The entropy encoding model corresponding to the encoding model parameters. Perform quantization processing on the first latent variable to obtain the quantized first latent variable. Based on the entropy coding model corresponding to the initial entropy coding model parameters, entropy coding is performed on the quantized first latent variable to obtain an initial coding result of the first latent variable. The number of coding bits of the initial coding result of the first latent variable is counted to obtain the basic number of coding bits of the initial coding.
  • step 1002 For the quantization processing method and entropy coding method in step 1002, reference may be made to the quantization processing method and entropy coding method in step 602, which will not be repeated here.
  • the implementation process of determining the initial encoding bit number based on the context initial encoding bit number and the basic initial encoding bit number includes: determining the sum of the context initial encoding bit number and the basic initial encoding bit number as the initial encoding bit number.
  • it can also be determined through other implementation manners.
  • the first latent variable is adjusted based on the first initial adjustment factor, and based on the adjusted first latent variable, the corresponding context initial coding bit number and initial entropy coding model parameters are determined through the context model, and based on the initial entropy
  • the coding model parameters determine the number of coding bits of the entropy coding result of the adjusted first latent variable, and obtain the basic initial coding bit number.
  • the initial encoding bit number is determined based on the context initial encoding bit number and the basic initial encoding bit number.
  • the first initial adjustment factor may be a first preset adjustment factor.
  • the context model includes a context encoding neural network model and a context decoding neural network model.
  • the implementation process of determining the corresponding context initial encoding bits and initial entropy encoding model parameters through the context model includes: processing the adjusted first latent variable through the context encoding neural network model to obtain the first Of the six latent variables, the sixth latent variable is used to indicate the adjusted probability distribution of the first latent variable.
  • the entropy coding result of the sixth latent variable is determined, and the coding bit number of the entropy coding result of the sixth latent variable is used as the initial coding bit number of the context.
  • the sixth latent variable is reconstructed based on the entropy coding result of the sixth latent variable, and the reconstructed sixth latent variable is processed through the context decoding neural network model to obtain initial entropy coding model parameters.
  • the adjusted first latent variable is processed through the context coding neural network model to obtain the sixth latent variable.
  • the sixth latent variable is quantified to obtain the quantized sixth latent variable.
  • Entropy encoding is performed on the quantized sixth latent variable, and the number of encoded bits of the entropy encoding result is counted to obtain the initial number of encoded bits of the context.
  • entropy decoding is performed on the entropy encoding result of the sixth latent variable to obtain a quantized sixth latent variable, and dequantization processing is performed on the quantized sixth latent variable to obtain a reconstructed sixth latent variable.
  • the reconstructed sixth latent variable is input into the context decoding neural network model to obtain initial entropy encoding model parameters output by the context decoding neural network model.
  • the above-mentioned first variable adjustment factor is the final value of the adjustment factor of the first latent variable.
  • the second latent variable is obtained after the first latent variable is adjusted by the first variable adjustment factor, and the second latent variable is processed by the context encoding neural network model to obtain the third latent variable, and the entropy coding result of the second latent variable and the third latent variable
  • the total number of encoded bits of the entropy encoding result of the latent variable satisfies the preset encoding rate condition.
  • satisfying the preset encoding rate condition includes that the number of encoding bits is less than or equal to the target number of encoding bits.
  • satisfying the preset coding rate condition includes that the number of coding bits is less than or equal to the target number of coding bits, and the difference between the number of coding bits and the target number of coding bits is smaller than a threshold value of the number of bits.
  • satisfying the preset encoding rate condition includes that the absolute value of the difference between the number of encoded bits and the target number of encoded bits is smaller than the bit number threshold. That is, satisfying the preset encoding rate condition includes that the number of encoded bits is less than or equal to the target number of encoded bits, and the difference between the target number of encoded bits and the number of encoded bits is smaller than the bit number threshold.
  • satisfying the preset encoding rate condition includes that the number of encoded bits is greater than or equal to the target number of encoded bits, and the difference between the number of encoded bits and the target number of encoded bits is smaller than a threshold value of the number of bits.
  • bit number threshold can be set in advance, and the bit number threshold can be adjusted based on different requirements.
  • the first initial adjustment factor is determined as the first variable adjustment factor. If the initial number of encoded bits is not equal to the target number of encoded bits, the first variable adjustment factor is determined in a first round-robin manner based on the initial number of encoded bits and the target number of encoded bits.
  • the i-th cycle processing of the first cycle method includes the following steps: determining the adjustment factor of the i-th cycle processing, i is a positive integer, and adjusting the first latent variable based on the adjustment factor of the i-th cycle processing to obtain the first The first latent variable after i adjustments. Based on the i-th adjusted first latent variable, determine the corresponding i-th context coding bit number and i-th entropy coding model parameters through the context model, and determine the i-th time based on the i-th entropy coding model parameters.
  • the number of coded bits of the entropy coding result of the adjusted first latent variable is obtained to obtain the number of basic coded bits of the ith time, and the code of the ith time is determined based on the number of coded bits of the context code of the ith time and the number of basic coded bits of the ith time number of bits.
  • the (i+1)th loop processing of the first loop mode is executed.
  • the execution of the first loop mode is terminated, and the first variable adjustment factor is determined based on the adjustment factor of the i-th loop process.
  • the implementation process of determining the adjustment factor of the i-th loop processing can refer to the relevant description in the above step 602, which will not be repeated here.
  • the implementation process of determining the i-th context coding bit number and the i-th entropy coding model parameters through the context model can refer to the above-mentioned determination of the context initial coding bit number and initial entropy coding model The process of parameters will not be repeated here.
  • the implementation process of determining the number of encoded bits of the entropy encoding result of the first latent variable after the i-th adjustment can refer to the above-mentioned process of determining the number of basic initial encoding bits based on the initial entropy encoding model parameters. I won't repeat them here.
  • the implementation process of determining the first variable adjustment factor based on the adjustment factor of the i-th loop processing can refer to the description in the above step 602, which will not be repeated here.
  • the first initial adjustment factor is determined as the first variable adjustment factor.
  • the first variable adjustment factor is determined according to the above first implementation manner.
  • the adjustment factor of the i-1th loop processing of the first loop mode is adjusted according to the length of the first step to obtain The adjustment factor of the ith loop processing, at this time, the condition for continuing to adjust includes that the number of coded bits in the ith time is less than the target number of coded bits.
  • the adjustment condition includes that the i-th number of encoded bits is greater than the target number of encoded bits.
  • adjusting the adjustment factor of the i-1th loop processing according to the first step length may refer to increasing the adjustment factor of the i-1th loop processing according to the first step length, and adjusting the i-1th loop processing according to the second step length
  • the adjustment factor of the cyclic processing may refer to reducing the adjustment factor of the i-1 th cyclic processing according to the second step size.
  • the above-mentioned increase processing and decrease processing may be linear or non-linear.
  • the sum of the adjustment factor of the i-1th loop processing and the first step length can be determined as the adjustment factor of the i-th loop processing, and the adjustment factor of the i-1th loop processing can be combined with the second step size The difference is determined as the adjustment factor for the ith cycle processing.
  • first step size and the second step size may be set in advance, and the first step size and the second step size may be adjusted based on different requirements.
  • first step length and the second step length may be equal or not.
  • the first initial adjustment factor is determined as the first variable adjustment factor.
  • the first variable adjustment factor may be determined according to the above-mentioned first implementation manner.
  • the first variable adjustment factor may also be determined according to the situation that the initial number of coding bits is greater than the target number of coding bits in the above-mentioned second implementation manner.
  • Step 1003 Obtain the entropy coding result of the second latent variable and the entropy coding result of the third latent variable, the second latent variable is obtained by adjusting the first latent variable by the first variable adjustment factor, and the third latent variable is used to indicate the The probability distribution of the two latent variables, the total number of encoded bits of the entropy coding result of the second latent variable and the entropy coding result of the third latent variable satisfy the preset coding rate condition.
  • the first variable adjustment factor may be the first initial adjustment factor, the adjustment factor for the i-th cycle processing, or the third adjustment factor for the j-th cycle processing , may also be determined based on the first adjustment factor of the j-th loop processing.
  • the corresponding entropy coding result is determined in the above-mentioned cyclic process, so the entropy coding result of the second latent variable and the entropy coding result of the third latent variable can be obtained directly from the entropy coding result determined in the above-mentioned cyclic process. Entropy encoding results.
  • the first latent variable is directly adjusted based on the first variable adjustment factor to obtain the second latent variable.
  • the entropy coding result of the third latent coding and the parameters of the first entropy coding model are determined through the context model.
  • the quantized entropy encoding result of the second latent variable is determined.
  • the context model includes a context coding neural network model and a context decoding neural network model
  • the implementation process of determining the entropy coding result of the third latent coding and the parameters of the first entropy coding model through the context model includes:
  • the neural network model processes the second latent variable to obtain the third latent variable.
  • the third latent variable is quantified to obtain the quantified third latent variable.
  • Entropy encoding is performed on the quantized third latent variable to obtain an entropy encoding result of the third latent variable.
  • Entropy decoding is performed on the entropy encoding result of the third latent variable to obtain a quantized third latent variable, and dequantization processing is performed on the quantized third latent variable to obtain a reconstructed third latent variable.
  • the reconstructed third latent variable is processed through the context decoding neural network model to obtain the parameters of the first entropy encoding model.
  • the implementation process of determining the entropy coding result of the quantized second latent variable includes: determining the entropy coding corresponding to the first entropy coding model parameters from the entropy coding model with adjustable coding model parameters Model. Based on the entropy coding model corresponding to the parameters of the first entropy coding model, entropy coding is performed on the quantized second latent variable to obtain an entropy coding result of the second latent variable.
  • Step 1004 Write the entropy coding result of the second latent variable, the entropy coding result of the third latent variable, and the coding result of the first variable adjustment factor into the code stream.
  • step 604 For details about the encoding result of the first variable adjustment factor, reference may be made to the description in step 604 , which will not be repeated here.
  • the embodiment of the present application may also determine a quantization index corresponding to a quantization step of the first variable adjustment factor to obtain the first quantization index.
  • the embodiment of the present application may also determine the quantization index corresponding to the quantization step of the second latent variable to obtain the second quantization index, and determine the quantization index corresponding to the quantization step of the third latent variable to obtain the third quantization index . Encode the first quantization index, the second quantization index and the third quantization index into the code stream.
  • the quantization index is used to indicate the corresponding quantization step size, that is, the first quantization index is used to indicate the quantization step size of the first variable adjustment factor, the second quantization index is used to indicate the quantization step size of the second latent variable, and the third The quantization index is used to indicate the quantization step size of the third latent variable.
  • the first latent variable is adjusted by the first variable adjustment factor to obtain the second latent variable
  • the second latent variable is processed by the context coding neural network model to obtain the third latent variable
  • the second The total number of encoded bits of the entropy encoding result of the latent variable and the entropy encoding result of the third latent variable satisfies the preset encoding rate condition, so that it can be ensured that the encoding bits of the entropy encoding result of the latent variable corresponding to each frame of media data can meet the preset encoding rate condition.
  • the encoding rate condition is set, that is, it can ensure that the number of encoded bits of the entropy encoding result of the latent variable corresponding to each frame of media data is basically consistent, rather than dynamically changing, thereby meeting the encoder's need for a stable encoding rate.
  • the need to transmit side information (such as window type, time-domain noise shaping (TNS: Temporal Noise Shaping) parameters, frequency-domain noise shaping (FDNS: Frequency-domain noise shaping) parameters, and/or bandwidth extension (BWE : bandwidth extension) parameter, etc.), it can ensure that the number of encoded bits of the entropy encoding result of the latent variable corresponding to each frame of media data and the encoded bit number of the side information are basically consistent, so as to meet the encoder's need for a stable encoding rate .
  • TMS Temporal Noise Shaping
  • FDNS Frequency-domain noise shaping
  • BWE bandwidth extension
  • FIG. 11 is a flowchart of a second decoding method provided by an embodiment of the present application. This method is applied to a decoding end, and this method corresponds to the encoding method shown in FIG. 10 . The method includes the following steps.
  • Step 1101 Determine the reconstructed third latent variable and the reconstructed first variable adjustment factor based on the code stream.
  • entropy decoding may be performed on the entropy encoding result of the third latent variable in the code stream, and the encoding result of the first variable adjustment factor in the code stream may be decoded to obtain the quantized third latent variable and The quantized adjustment factor for the first variable.
  • the quantized third latent variable and the quantized first variable adjustment factor are dequantized to obtain the reconstructed third latent variable and the reconstructed first variable adjustment factor.
  • the decoding method in this step corresponds to the encoding method at the encoding end
  • the dequantization processing in this step corresponds to the quantization processing at the encoding end. That is, the decoding method is the inverse process of the encoding method, and the dequantization process is the inverse process of the quantization process.
  • the first quantization index and the third quantization index can be parsed from the code stream, the first quantization index is used to indicate the quantization step size of the first variable adjustment factor, and the third quantization index is used to indicate the quantization of the third latent variable step size.
  • the quantized first variable adjustment factor is dequantized to obtain the reconstructed first variable adjustment factor.
  • the quantized third latent variable is dequantized to obtain a reconstructed third latent variable.
  • Step 1102 Determine a reconstructed second latent variable based on the code stream and the reconstructed third latent variable.
  • the reconstructed third latent variable is processed by the context decoding neural network model to obtain the reconstructed first entropy coding model parameters, and the reconstruction is determined based on the code stream and the reconstructed first entropy coding model parameters.
  • the second latent variable of the structure is processed by the context decoding neural network model to obtain the reconstructed first entropy coding model parameters, and the reconstruction is determined based on the code stream and the reconstructed first entropy coding model parameters.
  • an entropy decoding model corresponding to the reconstructed first entropy coding model parameters may be determined. Based on the entropy decoding model corresponding to the reconstructed first entropy coding model parameters, entropy decoding is performed on the entropy coding result of the second latent variable in the code stream to obtain the quantized second latent variable. The quantized second latent variable is dequantized to obtain a reconstructed second latent variable.
  • the decoding method in this step corresponds to the encoding method at the encoding end
  • the dequantization processing in this step corresponds to the quantization processing at the encoding end. That is, the decoding method is the inverse process of the encoding method, and the dequantization process is the inverse process of the quantization process.
  • the second quantization index may be parsed from the code stream, and the second quantization index is used to indicate the quantization step size of the second latent variable. According to the quantization step size indicated by the second quantization index, the quantized second latent variable is dequantized to obtain a reconstructed second latent variable.
  • Step 1103 Adjust the reconstructed second latent variable based on the reconstructed first variable adjustment factor to obtain the reconstructed first latent variable, which is used to indicate the characteristics of the media data to be decoded .
  • step 1103 for the implementation process of step 1103, reference may be made to the implementation process of the above-mentioned step 902, which will not be repeated here.
  • Step 1104 Process the reconstructed first latent variable through the first decoding neural network model to obtain reconstructed media data.
  • step 1104 for the implementation process of step 1104, reference may be made to the implementation process of the above-mentioned step 903, which will not be repeated here.
  • the entropy of the latent variable corresponding to each frame of media data can be guaranteed
  • the number of encoded bits of the encoding result can meet the preset encoding rate conditions, that is, it can be guaranteed that the number of encoded bits of the entropy encoding result of the latent variable corresponding to each frame of media data is basically consistent, rather than dynamically changing, thus satisfying The encoder's need for a stable encoding rate.
  • the need to transmit side information (such as window type, time-domain noise shaping (TNS: Temporal Noise Shaping) parameters, frequency-domain noise shaping (FDNS: Frequency-domain noise shaping) parameters, and/or bandwidth extension (BWE : bandwidth extension) parameter, etc.), it can ensure that the number of encoded bits of the entropy encoding result of the latent variable corresponding to each frame of media data and the encoded bit number of the side information are basically consistent, so as to meet the encoder's need for a stable encoding rate .
  • TMS Temporal Noise Shaping
  • FDNS Frequency-domain noise shaping
  • BWE bandwidth extension
  • FIG. 12 is a block diagram of an exemplary encoding method provided by an embodiment of the present application.
  • FIG. 12 is mainly an exemplary explanation of the encoding method shown in FIG. 10 .
  • an audio signal is taken as an example.
  • the audio signal can be processed by windowing to obtain the audio signal of the current frame.
  • the frequency domain signal of the current frame is obtained.
  • the first latent variable is output after being processed by the first encoding neural network model.
  • the first latent variable is adjusted based on the first latent variable adjustment factor to obtain the second latent variable.
  • the second latent variable is processed through the context coding neural network model to obtain the third latent variable.
  • Quantization processing and entropy coding are performed on the third latent variable to obtain an entropy coding result of the third latent variable, and the entropy coding result of the third latent variable is written into a code stream.
  • entropy decoding is performed on the entropy encoding result of the third latent variable to obtain a quantized third latent variable.
  • the quantified third latent variable is dequantized to obtain a reconstructed third latent variable.
  • the reconstructed third latent variable is processed through the context decoding neural network model to obtain the parameters of the first entropy coding model.
  • An entropy coding model corresponding to the parameters of the first entropy coding model is selected from the parameter-adjustable entropy coding models. Quantize the second latent variable, perform entropy coding on the quantized second latent variable based on the selected entropy coding model, obtain an entropy coding result of the second latent variable, and write the entropy coding result of the second latent variable into the code stream. Then, the encoding result of the first variable adjustment factor is written into the code stream.
  • FIG. 13 is a block diagram of an exemplary decoding method provided by an embodiment of the present application.
  • FIG. 13 is mainly an exemplary explanation of the decoding method shown in FIG. 11 .
  • an audio signal is taken as an example.
  • Entropy decoding is performed on the entropy coding result of the third latent variable in the code stream by using an entropy decoding model to obtain a quantized third latent variable.
  • the quantified third latent variable is dequantized to obtain a reconstructed third latent variable.
  • the reconstructed third variable is processed through the context decoding neural network model to obtain the reconstructed first entropy encoding model parameters.
  • a corresponding entropy decoding model is selected from parameter-adjustable entropy decoding models.
  • Entropy decoding is performed on the entropy encoding result of the second latent variable in the code stream according to the selected entropy decoding model to obtain the quantized second latent variable.
  • the quantized second latent variable is dequantized to obtain a reconstructed second latent variable.
  • the encoding result of the first variable adjustment factor in the code stream is decoded to obtain the reconstructed first variable adjustment factor.
  • the reconstructed second latent variable is adjusted based on the reconstructed first variable adjustment factor to obtain the reconstructed first latent variable.
  • the reconstructed first latent variable is processed through the first decoding neural network model to obtain a reconstructed frequency domain signal of the current frame. Perform IMDCT transformation processing and de-windowing processing on the reconstructed frequency domain signal of the current frame to obtain a reconstructed audio signal.
  • Figure 14 is a flow chart of the third encoding method provided by the embodiment of the present application.
  • This method includes a context model, and not only adjusts the latent variables generated by the media data to be encoded through adjustment factors, but also adjusts the context model The identified latent variables are adjusted by adjustment factors.
  • the encoding method is applied to an encoding end device, and includes the following steps.
  • Step 1401 Process the media data to be encoded through the first encoding neural network model to obtain a first latent variable, and the first latent variable is used to indicate the characteristics of the media data to be encoded.
  • step 140 For the implementation process of step 1401, reference may be made to the implementation process of step 601 above, which will not be repeated here.
  • Step 1402 Determine a first variable adjustment factor and a second variable adjustment factor based on the first latent variable.
  • the corresponding context initial coding bit number and initial entropy coding model parameters can be determined through the context model, and the coding bits of the entropy coding result of the first latent variable can be determined based on the initial entropy coding model parameters to obtain the basic initial coding bit number, and determine the first variable adjustment factor and the second variable adjustment factor based on the context initial coding bit number, the basic initial coding bit number and the target coding bit number.
  • the implementation process of determining the corresponding context initial coding bit number and the initial entropy coding model parameter through the context model, and determining the coding bit number of the entropy coding result of the first latent variable based on the initial entropy coding model parameter the implementation process of obtaining the basic initial number of encoded bits can refer to the relevant description in the above-mentioned step 1002, which will not be repeated here.
  • the first preset adjustment factor may be used as the first initial adjustment factor
  • the second preset adjustment factor may be used as the second initial adjustment factor.
  • the first latent variable is adjusted based on the first initial adjustment factor.
  • the corresponding context initial coding bit number and initial entropy coding model parameters are determined through the context model.
  • An encoding model parameter is used to determine the encoding bit number of the entropy encoding result of the first latent variable, so as to obtain the basic initial encoding bit number.
  • a first variable adjustment factor and a second variable adjustment factor are determined based on the context initial encoding bit number, the basic initial encoding bit number and the target encoding bit number.
  • the context model includes a context encoding neural network model and a context decoding neural network model.
  • the implementation process of determining the corresponding context initial coding bits and initial entropy coding model parameters through the context model includes: using the context coding neural network model to adjust the adjusted first latent Variables are processed to obtain a sixth latent variable, which is used to indicate the adjusted probability distribution of the first latent variable.
  • the sixth latent variable is adjusted based on the second initial adjustment factor to obtain a seventh latent variable.
  • the entropy coding result of the seventh latent variable is determined, and the coding bit number of the entropy coding result of the seventh latent variable is used as the initial coding bit number of the context.
  • the seventh latent variable is reconstructed based on the entropy coding result of the seventh latent variable, and the reconstructed seventh latent variable is processed through a context decoding neural network model to obtain initial entropy coding model parameters.
  • the adjusted first latent variable is processed through the context coding neural network model to obtain the sixth latent variable.
  • the sixth latent variable is adjusted based on the second initial adjustment factor to obtain a seventh latent variable.
  • the seventh latent variable is quantified to obtain the quantized seventh latent variable.
  • Entropy encoding is performed on the quantized seventh latent variable, and the number of encoded bits of the entropy encoding result is counted to obtain the initial number of encoded bits of the context.
  • entropy decoding is performed on the entropy encoding result of the seventh latent variable to obtain a quantized seventh latent variable, and dequantization processing is performed on the quantized seventh latent variable to obtain a reconstructed seventh latent variable. Inputting the reconstructed seventh latent variable into the context decoding neural network model to obtain initial entropy coding model parameters output by the context decoding neural network model.
  • the implementation process of adjusting the sixth latent variable based on the second initial adjustment factor is: multiplying each element in the sixth latent variable by the corresponding element in the second initial adjustment factor to obtain the seventh latent variable.
  • each element in the sixth latent variable may be divided by the corresponding element in the second initial adjustment factor to obtain the seventh latent variable.
  • the embodiment of the present application does not limit the adjustment method.
  • an initial value of the adjustment factor is set for the third latent variable, and the initial value of the adjustment factor is usually equal to 1.
  • the second preset adjustment factor may be greater than or equal to the initial value of the adjustment factor, and may also be smaller than the initial value of the adjustment factor.
  • the second preset adjustment factor is a constant such as 1 or 2.
  • the second initial adjustment factor is the initial value of the adjustment factor
  • the first variable adjustment factor is determined through the above-mentioned second implementation method
  • the second initial adjustment factor is the second preset adjustment factor.
  • the second initial adjustment factor can be a scalar or a vector.
  • the related introduction of the first initial adjustment factor please refer to the related introduction of the first initial adjustment factor.
  • variable adjustment factor and the second variable adjustment factor may include two types, which will be introduced respectively next.
  • the second variable adjustment factor is set as the second initial adjustment factor, and based on at least one of the basic initial encoding bit number and the context initial encoding bit number, and the target encoding bit number, the basic target encoding bit number is determined .
  • the first variable adjustment factor and the basic actual number of encoded bits are determined.
  • the basic actual number of encoded bits refers to the first variable adjusted by the first variable adjustment factor.
  • the number of encoded bits of the entropy encoding result of the latent variable.
  • the context target number of coding bits is determined.
  • the second variable adjustment factor is determined based on the target number of coding bits of the context and the number of initial coding bits of the context.
  • the implementation process of determining the basic target coding bit number includes: subtracting the context initial coding bit number from the target coding bit number, Get the number of base target coded bits. Alternatively, determine the ratio between the basic initial coding bit number and the context initial coding bit number, and determine the basic target coding bit number based on the ratio and the target coding bit number. Alternatively, the basic target number of coding bits is determined based on the ratio of the target number of coding bits to the basic initial number of coding bits. Of course, it can also be determined through other implementation processes.
  • the target encoding bits may be multiplied by 5/8 to obtain the basic target encoding bits.
  • the ratio of the target coded bit number to the basic initial coded bit number determines the ratio of the target coded bit number to the basic initial coded bit number, if the ratio of the target coded bit number to the basic initial coded bit number is greater than the first ratio threshold, then the basic initial coded bit number to the preset first ratio The sum of the adjustment steps is determined as the basic target coding bit number. If the ratio of the target number of coding bits to the basic initial number of coding bits is smaller than the second proportional threshold, the difference between the basic initial coding number of bits and the preset second proportional adjustment step is determined as the basic target number of coding bits. Wherein, the first ratio threshold is greater than the second ratio threshold. If the ratio of the target number of encoded bits to the basic initial number of encoded bits is greater than or equal to the second ratio threshold and less than or equal to the first ratio threshold, then the basic initial number of encoded bits is determined as the basic target number of encoded bits.
  • first proportional adjustment step can be larger than the second proportional adjustment step, or smaller than the second proportional adjustment step, of course, it can also be equal to the second proportional adjustment step.
  • the size relationship between the adjustment step and the second proportional adjustment step is not limited.
  • Mode 11 when the basic initial number of coding bits is equal to the basic target number of coding bits, determine the first initial adjustment factor as the first variable adjustment factor. In the case that the basic initial coding bit number is not equal to the basic target coding bit number, the first variable adjustment factor is determined in a first round-robin manner based on the basic initial coding bit number and the basic target coding bit number.
  • the i-th cycle processing of the first cycle method includes the following steps: determining the adjustment factor of the i-th cycle processing, i is a positive integer, and adjusting the first latent variable based on the adjustment factor of the i-th cycle processing to obtain the first The first latent variable after i adjustments.
  • the i-th context coding bit number and the i-th entropy coding model parameters are determined through the context model, and based on the i-th entropy coding model parameters, Determine the number of coding bits of the entropy coding result of the first latent variable after the i-th adjustment, and obtain the basic coding bits of the i-th time, and determine the first The number of encoded bits for i times.
  • the (i+1)th loop processing of the first loop mode is executed.
  • the execution of the first loop mode is terminated, and the first variable adjustment factor is determined based on the adjustment factor of the i-th loop process.
  • the content in the manner 11 of determining the first variable adjustment factor is similar to the content in the first implementation manner of determining the first variable adjustment factor in step 1002 .
  • the difference lies in that this step is based on the i-th adjusted first latent variable and the second initial adjustment factor, and the i-th context coding bit number and the i-th entropy coding model parameters are determined through the context model.
  • the implementation process of determining the i-th context coding bit number and the i-th entropy coding model parameters through the context model can refer to the above adjustment based on After the first latent variable and the second initial adjustment factor, the implementation process of determining the corresponding context initial encoding bit number and initial entropy encoding model parameters through the context model will not be repeated here.
  • Mode 12 when the basic initial number of coding bits is equal to the basic target number of coding bits, determine the first initial adjustment factor as the first variable adjustment factor.
  • the first variable adjustment factor is determined according to the above-mentioned manner 11.
  • the adjustment factor of the i-1th cycle processing of the first cycle method is adjusted according to the length of the first step to obtain the first The adjustment factor for the i-time loop processing, at this time, the condition for continuing the adjustment includes that the number of coded bits for the ith time is less than the basic target number of coded bits.
  • the condition for continuing to adjust includes that the number of coded bits for the ith time is greater than the number of basic target coded bits.
  • Mode 13 when the basic initial coding bit number is less than or equal to the basic target coding bit number, determine the first initial adjustment factor as the first variable adjustment factor.
  • the first variable adjustment factor may be determined according to the above manner 11.
  • the first variable adjustment factor may also be determined according to the condition that the basic initial coded bit number is greater than the basic target coded bit number in the above manner 12.
  • the realization process of determining the context target coding bit number includes: subtracting the target coding bit number from the basic actual coding bit number to obtain the context target coding bit number.
  • the implementation process of determining the second variable adjustment factor based on the context target number of coding bits and the context initial number of coding bits is similar to the implementation process of determining the first variable adjustment factor based on the basic initial number of coding bits and the basic target number of coding bits.
  • the following are also divided into three ways to introduce respectively.
  • the second initial adjustment factor is determined as the second variable adjustment factor.
  • the first latent variable is adjusted based on the first variable adjustment factor to obtain the second latent variable.
  • the second variable adjustment factor is determined in a first loop manner.
  • the i-th cycle processing of the first cycle method includes the following steps: determining the adjustment factor of the i-th cycle processing, i is a positive integer, based on the second latent variable and the adjustment factor of the i-th cycle processing, determined through the context model
  • the i-th context coding bit number and the i-th entropy coding model parameters based on the i-th entropy coding model parameters, determine the number of coding bits of the entropy coding result of the second latent variable, and obtain the i-th basic coding bits
  • the number of coded bits of the i-th time is determined based on the number of coded bits of the i-th time context and the number of basic coded bits of the ith time.
  • the (i+1)th loop processing of the first loop mode is executed.
  • the execution of the first loop mode is terminated, and the second variable adjustment factor is determined based on the adjustment factor of the i-th loop process.
  • the implementation process of determining the i-th context encoding bit number and the i-th entropy encoding model parameters through the context model includes: using the context encoding neural network model to The second latent variable is processed to obtain a third latent variable, and the third latent variable is used to indicate the probability distribution of the second latent variable.
  • the third latent variable is adjusted based on the adjustment factor of the i-th cyclic process to obtain the i-th adjusted third latent variable.
  • the i-th adjusted third latent variable is reconstructed based on the entropy encoding result of the i-th adjusted third latent variable.
  • the adjusted third latent variable reconstructed for the ith time is adjusted by the adjustment factor of the ith cycle processing to obtain the reconstructed third latent variable.
  • the reconstructed third latent variable is processed through the context decoding neural network model to obtain the i-th entropy encoding model parameters.
  • the second initial adjustment factor is determined as the second variable adjustment factor.
  • the second variable adjustment factor is determined according to the above-mentioned manner 21.
  • the adjustment factor of the i-1th round of the first round-robin method is adjusted according to the length of the first step, and the first The adjustment factor for the i-time cyclic processing.
  • the continuous adjustment condition includes that the number of coded bits for the ith time is less than the target number of coded bits for the context.
  • adjust the adjustment factor of the i-1th loop processing of the first loop mode according to the second step length, and obtain the adjustment factor of the i-th loop process, at this time continue to adjust the condition that the number of coded bits for the ith time is greater than the target number of coded bits for the context.
  • the second initial adjustment factor is determined as the second variable adjustment factor.
  • the second variable adjustment factor may be determined according to the above manner 21.
  • the second variable adjustment factor may also be determined according to the situation that the initial number of coded bits of the context is greater than the number of target coded bits of the context in the above manner 22.
  • the target coding bit number is divided into the basic target coding bit number and the context target coding bit number, and the first variable adjustment factor is determined based on the basic target coding bit number and the basic initial coding bit number.
  • the second variable adjustment factor is determined based on the target number of coding bits of the context and the number of initial coding bits of the context.
  • the second variable adjustment factor when determining the first variable adjustment factor based on the number of basic target coding bits and the basic initial coding bit number, the second variable adjustment factor can be set as the second initial adjustment factor, based on the second initial adjustment factor, the basic target coding bit number and the number of basic initial coding bits to determine the first variable adjustment factor.
  • the specific implementation process can refer to the previous description.
  • the first variable adjustment factor has already been determined. At this time, it can be directly based on the first variable adjustment factor, the context target number of coding bits and The number of initial encoding bits of the context determines the second variable adjustment factor.
  • the specific implementation process can refer to the previous description.
  • Step 1403 Obtain the entropy encoding result of the second latent variable and the entropy encoding result of the fourth latent variable, the second latent variable is obtained by adjusting the first latent variable by the first variable adjustment factor, and the fourth latent variable is obtained based on the second
  • the variable adjustment factor is obtained after adjusting the third latent variable, the third latent variable is determined based on the second latent variable through the context model, the third latent variable is used to indicate the probability distribution of the second latent variable, and the entropy of the second latent variable
  • the total number of encoded bits of the encoding result and the entropy encoding result of the fourth latent variable satisfies a preset encoding rate condition.
  • step 1402 In the process of determining the first variable adjustment factor and the second variable adjustment factor in step 1402, there is an entropy encoding result corresponding to the adjustment factor for each cycle processing in the above-mentioned loop process, so it can be determined directly from the above-mentioned loop process.
  • the entropy encoding result of the second latent variable and the entropy encoding result of the fourth latent variable are obtained from the entropy encoding result of the entropy encoding.
  • the first latent variable is directly adjusted based on the first variable adjustment factor to obtain the second latent variable.
  • the entropy coding result of the fourth latent coding and the second entropy coding model parameters are determined through the context model. Perform quantization processing on the second latent variable to obtain the quantized second latent variable.
  • the quantized entropy encoding result of the second latent variable is determined.
  • the context model includes a context encoding neural network model and a context decoding neural network model, and based on the second latent variable, the implementation process of determining the entropy encoding result of the fourth latent encoding and the parameters of the second entropy encoding model through the context model includes: through context encoding
  • the neural network model processes the second latent variable to obtain the third latent variable.
  • the third latent variable is adjusted based on the second variable adjustment factor to obtain a fourth latent variable.
  • the fourth latent variable is quantified to obtain the quantized fourth latent variable.
  • Entropy coding is performed on the quantized fourth latent variable to obtain an entropy coding result of the fourth latent variable.
  • Entropy decoding is performed on the entropy encoding result of the fourth latent variable to obtain a quantized fourth latent variable, and dequantization processing is performed on the quantized fourth latent variable to obtain a reconstructed fourth latent variable.
  • the reconstructed fourth latent variable is adjusted based on the second variable adjustment factor to obtain the reconstructed third latent variable.
  • the reconstructed third latent variable is processed through the context decoding neural network model to obtain the parameters of the second entropy coding model.
  • the implementation process of determining the entropy coding result of the quantized second latent variable includes: determining the entropy coding corresponding to the second entropy coding model parameters from the entropy coding model with adjustable coding model parameters Model. Based on the entropy coding model corresponding to the parameters of the second entropy coding model, entropy coding is performed on the quantized second latent variable to obtain an initial coding result of the second latent variable.
  • Step 1404 Write the entropy encoding result of the second latent variable, the entropy encoding result of the fourth latent variable, the encoding result of the first variable adjustment factor, and the encoding result of the second variable adjustment factor into the code stream.
  • step 604 For details about the encoding result of the first variable adjustment factor, reference may be made to the description in step 604 , which will not be repeated here.
  • the second variable adjustment factor may be a value before quantization, or a value after quantization. If the second variable adjustment factor is the value before quantization, at this time, the realization process of determining the encoding result of the second variable adjustment factor includes: performing quantization and encoding processing on the second variable adjustment factor to obtain the encoding result of the second variable adjustment factor . If the second variable adjustment factor is a quantized value, at this time, the implementation process of determining the encoding result of the second variable adjustment factor includes: encoding the second variable adjustment factor to obtain the encoding result of the second variable adjustment factor.
  • the second variable adjustment factor may be encoded in any encoding manner, which is not limited in this embodiment of the present application.
  • the quantization process and the encoding process are performed in one process, that is, the quantization result and the encoding result can be obtained through one process. Therefore, for the second variable adjustment factor, the encoding result of the second variable adjustment factor may also be obtained in the process of determining the second variable adjustment factor, so the encoding result of the second variable adjustment factor can be obtained directly. That is, the encoding result of the second variable adjustment factor can also be directly obtained during the process of determining the second variable adjustment factor.
  • this embodiment of the present application may also determine the quantization index corresponding to the quantization step of the first variable adjustment factor to obtain the first quantization index, and A quantization step size of the second variable adjustment factor is determined to obtain a fifth quantization index.
  • the quantization index corresponding to the quantization step of the second latent variable may be determined to obtain the second quantization index
  • the quantization index corresponding to the quantization step of the fourth latent variable may be determined to obtain the fourth quantization index, Encode the first quantization index, the second quantization index, the fourth quantization index and the fifth quantization index into the code stream.
  • the quantization index is used to indicate the corresponding quantization step size, that is, the first quantization index is used to indicate the quantization step size of the first variable adjustment factor, the second quantization index is used to indicate the quantization step size of the second latent variable, and the fourth The quantization index is used to indicate the quantization step size of the fourth latent variable, and the fifth quantization index is used to indicate the quantization step size of the second variable adjustment factor.
  • the first latent variable is adjusted by the first variable adjustment factor to obtain the second latent variable
  • the second latent variable is processed by the context encoding neural network model to obtain the third latent variable.
  • the third latent variable is adjusted by the second variable adjustment factor to obtain the fourth latent variable.
  • the total number of encoded bits of the entropy encoding result of the second latent variable and the entropy encoding result of the fourth latent variable satisfies the preset encoding rate condition, so that the encoding bits of the entropy encoding result of the latent variable corresponding to each frame of media data can be guaranteed to be equal.
  • the need to transmit side information (such as window type, time-domain noise shaping (TNS:Temporal Noise Shaping) parameters, frequency-domain noise shaping (FDNS:Frequency-domain noise shaping) parameters, and/or bandwidth extension (BWE : bandwidth extension) parameters, etc.), it can ensure that the number of coded bits of the entropy coding result of the latent variable corresponding to each frame of media data and the coded bits of the side information are basically consistent as a whole, so as to meet the needs of the encoder for a stable coding rate .
  • side information such as window type, time-domain noise shaping (TNS:Temporal Noise Shaping) parameters, frequency-domain noise shaping (FDNS:Frequency-domain noise shaping) parameters, and/or bandwidth extension (BWE : bandwidth extension) parameters, etc.
  • FIG. 15 is a flowchart of a third decoding method provided by an embodiment of the present application, which is applied to a decoding end. This method corresponds to the encoding method shown in FIG. 14, and the method includes the following steps.
  • Step 1501 Determine a reconstructed fourth latent variable, a reconstructed first variable adjustment factor, and a reconstructed second variable adjustment factor based on the code stream.
  • entropy decoding may be performed on the entropy encoding result of the fourth latent variable in the code stream, and the encoding result of the first variable adjustment factor and the second variable adjustment factor in the code stream may be decoded to obtain the quantized The fourth latent variable of , the quantized adjustment factor for the first variable, and the quantized adjustment factor for the second variable. Dequantize the quantized fourth latent variable, the quantized first variable adjustment factor and the quantized second variable adjustment factor to obtain the reconstructed fourth latent variable, the reconstructed first variable adjustment factor and the reconstructed Adjustment factor for the second variable of the structure.
  • the decoding method in this step corresponds to the encoding method at the encoding end
  • the dequantization processing in this step corresponds to the quantization processing at the encoding end. That is, the decoding method is the inverse process of the encoding method, and the dequantization process is the inverse process of the quantization process.
  • the first quantization index, the fourth quantization index and the fifth quantization index can be parsed from the code stream, the first quantization index is used to indicate the quantization step size of the first variable adjustment factor, and the fourth quantization index is used to indicate the first quantization index
  • the quantization step size of the four latent variables, the fifth quantization index is used to indicate the quantization step size of the adjustment factor of the second variable.
  • the quantized first variable adjustment factor is dequantized to obtain the reconstructed first variable adjustment factor.
  • the quantized fourth latent variable is dequantized to obtain a reconstructed fourth latent variable.
  • the quantized second variable adjustment factor is dequantized to obtain the reconstructed second variable adjustment factor.
  • Step 1502 Determine a reconstructed second latent variable based on the code stream, the reconstructed fourth latent variable and the reconstructed second variable adjustment factor.
  • the reconstructed fourth latent variable is adjusted based on the reconstructed second variable adjustment factor to obtain the reconstructed third latent variable.
  • an entropy decoding model corresponding to the reconstructed second entropy coding model parameters may be determined. Based on the entropy decoding model corresponding to the parameters of the second entropy coding model, entropy decoding is performed on the entropy coding result of the second latent variable in the code stream to obtain the quantized second latent variable. The quantized second latent variable is dequantized to obtain a reconstructed second latent variable.
  • the decoding method in this step corresponds to the encoding method at the encoding end
  • the dequantization processing in this step corresponds to the quantization processing at the encoding end. That is, the decoding method is the inverse process of the encoding method, and the dequantization process is the inverse process of the quantization process.
  • the second quantization index may be parsed from the code stream, and the second quantization index is used to indicate the quantization step size of the second latent variable. According to the quantization step size indicated by the second quantization index, the quantized second latent variable is dequantized to obtain a reconstructed second latent variable.
  • Step 1503 Adjust the reconstructed second latent variable based on the reconstructed first variable adjustment factor to obtain the reconstructed first latent variable.
  • step 1503 For the implementation process of step 1503, reference may be made to the implementation process of step 902 above, which will not be repeated here.
  • Step 1504 Process the reconstructed first latent variable through the first decoding neural network model to obtain reconstructed media data.
  • step 1504 For the implementation process of step 1504, reference may be made to the implementation process of step 903 above, which will not be repeated here.
  • the entropy of the latent variable corresponding to each frame of media data can be guaranteed
  • the number of encoded bits of the encoding result can meet the preset encoding rate conditions, that is, it can be guaranteed that the number of encoded bits of the entropy encoding result of the latent variable corresponding to each frame of media data is basically consistent, rather than dynamically changing, thus satisfying The encoder's need for a stable encoding rate.
  • the need to transmit side information (such as window type, time-domain noise shaping (TNS: Temporal Noise Shaping) parameters, frequency-domain noise shaping (FDNS: Frequency-domain noise shaping) parameters, and/or bandwidth extension (BWE : bandwidth extension) parameters, etc.), it can ensure that the number of coded bits of the entropy coding result of the latent variable corresponding to each frame of media data and the coded bits of the side information are basically consistent as a whole, so as to meet the needs of the encoder for a stable coding rate .
  • TMS Temporal Noise Shaping
  • FDNS Frequency-domain noise shaping
  • BWE bandwidth extension
  • FIG. 16 is a block diagram of an exemplary encoding method provided by an embodiment of the present application.
  • FIG. 16 is mainly an exemplary explanation of the encoding method shown in FIG. 14 .
  • an audio signal is taken as an example.
  • the audio signal can be processed by windowing to obtain the audio signal of the current frame.
  • the frequency domain signal of the current frame is obtained.
  • the first latent variable is output after being processed by the first encoding neural network model.
  • the first latent variable is adjusted based on the first latent variable adjustment factor to obtain the second latent variable.
  • the second latent variable is processed through the context coding neural network model to obtain the third latent variable.
  • the third latent variable is adjusted based on the second variable adjustment factor to obtain a fourth latent variable.
  • Quantization processing and entropy coding are performed on the fourth latent variable to obtain an entropy coding result of the fourth latent variable, and the entropy coding result of the fourth latent variable is written into a code stream.
  • entropy decoding is performed on the entropy encoding result of the fourth latent variable to obtain a quantized fourth latent variable.
  • the quantized fourth latent variable is dequantized to obtain a reconstructed fourth latent variable.
  • the reconstructed fourth latent variable is adjusted based on the second variable adjustment factor to obtain the reconstructed third latent variable.
  • the reconstructed third latent variable is processed through the context decoding neural network model to obtain the parameters of the second entropy encoding model.
  • An entropy coding model corresponding to the parameters of the second entropy coding model is selected from the parameter-adjustable entropy coding models.
  • Quantize the second latent variable perform entropy coding on the quantized second latent variable based on the selected entropy coding model, obtain an entropy coding result of the second latent variable, and write the entropy coding result of the second latent variable into the code stream. Then, the encoding result of the first variable adjustment factor is written into the code stream.
  • FIG. 17 is a block diagram of an exemplary decoding method provided by an embodiment of the present application.
  • FIG. 17 is mainly an exemplary explanation of the decoding method shown in FIG. 15 .
  • an audio signal is taken as an example.
  • Entropy decoding is performed on the entropy coding result of the fourth latent variable in the code stream by using an entropy decoding model to obtain a quantized fourth latent variable.
  • the quantized fourth latent variable is dequantized to obtain a reconstructed fourth latent variable.
  • the encoding result of the second variable adjustment factor in the code stream is decoded to obtain the quantized second variable adjustment factor.
  • the quantized second variable adjustment factor is dequantized to obtain the reconstructed second variable adjustment factor.
  • the reconstructed fourth latent variable is adjusted based on the reconstructed second variable adjustment factor to obtain the reconstructed third latent variable.
  • the reconstructed third variable is processed through the context decoding neural network model to obtain the reconstructed second entropy encoding model parameters.
  • a corresponding entropy decoding model is selected from parameter-adjustable entropy decoding models.
  • Entropy decoding is performed on the entropy encoding result of the second latent variable in the code stream according to the selected entropy decoding model to obtain the quantized second latent variable.
  • the quantized second latent variable is dequantized to obtain a reconstructed second latent variable.
  • the encoding result of the first variable adjustment factor in the code stream is decoded to obtain the reconstructed first variable adjustment factor.
  • the reconstructed second latent variable is adjusted based on the reconstructed first variable adjustment factor to obtain the reconstructed first latent variable.
  • the reconstructed first latent variable is processed through the first decoding neural network model to obtain a reconstructed frequency domain signal of the current frame. Perform IMDCT transformation processing and de-windowing processing on the reconstructed frequency domain signal of the current frame to obtain a reconstructed audio signal.
  • Fig. 18 is a schematic structural diagram of an encoding device provided by an embodiment of the present application.
  • the encoding device can be implemented by software, hardware or a combination of the two to become part or all of the encoding end device.
  • the encoding end device can be as shown in Fig. 1 source device. Referring to FIG. 18 , the device includes: a data processing module 1801 , an adjustment factor determination module 1802 , a first encoding result acquisition module 1803 and a first encoding result writing module 1804 .
  • the data processing module 1801 is configured to process the media data to be encoded through the first encoding neural network model to obtain a first latent variable, and the first latent variable is used to indicate the characteristics of the media data to be encoded.
  • first latent variable is used to indicate the characteristics of the media data to be encoded.
  • the adjustment factor determination module 1802 is configured to determine a first variable adjustment factor based on the first latent variable, the first variable adjustment factor is used to make the number of encoded bits of the entropy encoding result of the second latent variable meet the preset encoding rate condition, and the second latent variable The variable is obtained by adjusting the first latent variable by the first variable adjustment factor.
  • the first encoding result acquisition module 1803 is configured to acquire the entropy encoding result of the second latent variable.
  • the detailed implementation process refer to the corresponding content in the foregoing embodiments, and details are not repeated here.
  • the first encoding result writing module 1804 is configured to write the entropy encoding result of the second latent variable and the encoding result of the first variable adjustment factor into the code stream.
  • satisfying the preset encoding rate condition includes that the number of encoding bits is less than or equal to the target encoding bit number; or, satisfying the preset encoding rate condition includes encoding the number of bits is less than or equal to the target number of encoded bits, and the difference between the number of encoded bits and the target number of encoded bits is smaller than the threshold value of the number of bits.
  • satisfying the preset encoding rate condition includes that the absolute value of the difference between the number of encoded bits and the target number of encoded bits is smaller than a bit number threshold.
  • the adjustment factor determining module 1802 includes:
  • the bit number determination submodule is used to determine the initial encoding bit number based on the first latent variable
  • the first factor determining submodule is configured to determine a first variable adjustment factor based on the initial number of coding bits and the target number of coding bits.
  • the initial number of coding bits is the number of coding bits of the entropy coding result of the first latent variable
  • the initial number of coding bits is the number of coding bits of the entropy coding result of the first latent variable adjusted by the first initial adjustment factor.
  • the first factor determination submodule is specifically used for:
  • a first variable adjustment factor is determined through a first loop
  • the ith loop processing of the first loop mode comprises the following steps:
  • the execution of the first loop mode is terminated, and the first variable adjustment factor is determined based on the adjustment factor of the i-th loop process.
  • the device also includes:
  • the second encoding result acquisition module is used to acquire the entropy encoding result of the third latent variable, and the third latent variable is obtained through the context model based on the second latent variable;
  • the second encoding result writing module is used to write the entropy encoding result of the third latent variable into the code stream;
  • the total number of encoded bits of the entropy encoding result of the second latent variable and the entropy encoding result of the third latent variable satisfies the preset encoding rate condition.
  • the adjustment factor determination module 1802 includes:
  • the bit number determination submodule is used to determine the initial encoding bit number based on the first latent variable
  • the first factor determining submodule is configured to determine a first variable adjustment factor based on the initial number of coding bits and the target number of coding bits.
  • bit number determination submodule is specifically used for:
  • the initial encoding bit number is determined based on the context initial encoding bit number and the basic initial encoding bit number.
  • the first factor determination submodule is specifically used for:
  • a first variable adjustment factor is determined through a first loop
  • the ith loop processing of the first loop mode comprises the following steps:
  • the execution of the first loop mode is terminated, and the first variable adjustment factor is determined based on the adjustment factor of the i-th loop process.
  • the first factor determining submodule is specifically used for:
  • the adjustment factor of the i-1th loop processing in the first loop mode Based on the adjustment factor of the i-1th loop processing in the first loop mode, the number of coded bits of the i-1th time and the target number of coded bits, determine the adjustment factor of the i-th loop process;
  • the adjustment factor of the i-1th cyclic processing is the first initial adjustment factor
  • the number of encoded bits of the i-1th time is the initial number of encoded bits
  • the continuous adjustment condition includes that the number of coded bits of the i-1th time and the number of coded bits of the i-th time are both less than the target number of coded bits, or the continuous adjustment condition includes the number of coded bits of the i-1th time and the number of coded bits of the ith time The numbers are greater than the target number of encoded bits.
  • the first factor determination submodule is specifically used to:
  • the adjustment factor of the i-1th loop processing is the first initial adjustment factor
  • the condition for continuing to adjust includes that the number of encoded bits for the ith time is smaller than the target number of encoded bits.
  • the first factor determination submodule is specifically used to:
  • the adjustment factor of the i-1th loop processing is the first initial adjustment factor
  • the condition for continuing to adjust includes that the i-th number of encoded bits is greater than the target number of encoded bits.
  • the first factor determining submodule is specifically used for:
  • the first initial adjustment factor is determined as the first variable adjustment factor.
  • the first factor determining submodule is specifically used for:
  • the adjustment factor of the ith cycle processing is determined as the first variable adjustment factor
  • the first variable adjustment factor is determined based on the adjustment factor of the i-th round-robin processing and the adjustment factor of the i-1th round-robin processing of the first round-robin mode .
  • the first factor determining submodule is specifically used for:
  • a first variable adjustment factor is determined based on the average.
  • the first factor determining submodule is specifically used for:
  • the first variable adjustment factor is determined through the second cycle
  • the jth loop processing of the second loop mode comprises the following steps:
  • the third adjustment factor of the j th cyclic processing is determined, wherein, when j is equal to 1, the j cyclic processing.
  • the first adjustment factor is one of the adjustment factor of the i-th cycle processing and the adjustment factor of the i-1-th cycle processing
  • the second adjustment factor of the j-th cycle processing is the adjustment factor of the i-th cycle processing and the adjustment factor of the i-1th cycle processing.
  • the number of second coded bits, the number of first coded bits of the jth time is less than the number of second coded bits of the jth time, and j is a positive integer;
  • the third coding bit number of the jth time refers to the number of coding bits of the entropy coding result of the first latent variable adjusted by the third adjustment factor of the jth cycle processing;
  • the third coding bit number of the jth time satisfies the continuation loop condition, the third coding bit number of the jth time is greater than the target coding bit number and less than the second coding bit number of the jth time, the third coded bit number of the jth time loop processing
  • the adjustment factor is used as the second adjustment factor of the j+1th round of processing, the first adjustment factor of the jth round of processing is used as the first adjustment factor of the j+1th round of processing, and the j+1th round of the second round is executed. 1 cycle processing;
  • the third encoding bit number of the jth time satisfies the continuation loop condition, the third encoding bit number of the jth time is less than the target encoding bit number and greater than the first encoding bit number of the jth time, the third encoding bit number of the jth loop processing
  • the adjustment factor is used as the first adjustment factor of the j+1th round of processing, and the second adjustment factor of the jth round of processing is used as the second adjustment factor of the j+1th round of processing, and the j+1th round of the second round is executed. 1 cycle processing.
  • the first factor determining submodule is specifically used for:
  • the first variable adjustment factor is determined through the second cycle
  • the jth loop processing of the second loop mode comprises the following steps:
  • the third adjustment factor of the j th cyclic processing is determined, wherein, when j is equal to 1, the j cyclic processing.
  • the first adjustment factor is one of the adjustment factor of the i-th cycle processing and the adjustment factor of the i-1-th cycle processing
  • the second adjustment factor of the j-th cycle processing is the adjustment factor of the i-th cycle processing and the adjustment factor of the i-1th cycle processing.
  • the number of second coded bits, the number of first coded bits of the jth time is less than the number of second coded bits of the jth time, and j is a positive integer;
  • the third coding bit number of the jth time refers to the number of coding bits of the entropy coding result of the first latent variable adjusted by the third adjustment factor of the jth cycle processing;
  • the number of third coded bits of the jth time satisfies the condition for continuing the cycle, and the number of third coded bits of the jth time is greater than the number of target coded bits and less than the number of second coded bits of the jth time
  • the The third adjustment factor of the j cycle processing is used as the second adjustment factor of the j+1 cycle processing
  • the first adjustment factor of the j cycle processing is used as the first adjustment factor of the j+1 cycle processing, and the first adjustment factor is executed.
  • the number of third coded bits of the jth time satisfies the condition for continuing the cycle, and the number of third coded bits of the jth time is less than the target number of coded bits and greater than the number of first coded bits of the jth time
  • the The third adjustment factor of the j loop processing is used as the first adjustment factor of the j+1 loop processing
  • the second adjustment factor of the j loop processing is used as the second adjustment factor of the j+1 loop processing
  • the first adjustment factor is executed.
  • the first factor determining submodule is specifically used for:
  • the first adjustment factor of the jth loop processing is determined as the first variable adjustment factor.
  • the first factor determining submodule is specifically used for:
  • first difference is equal to the second difference, determine the first adjustment factor of the j-th cyclic process as the first variable adjustment factor, or determine the second adjustment factor of the j-th cyclic process as the first variable adjustment factor .
  • the continuation loop condition includes that the third coded bit number of the jth time is greater than the target coded bit number, or the continuation loop condition includes the jth time
  • the third encoding bit number is smaller than the target encoding bit number, and the difference between the target encoding bit number and the jth third encoding bit number is greater than the bit number threshold.
  • the continuation loop condition includes that the absolute value of the difference between the target encoding bit number and the j-th third encoding bit number is greater than the bit number threshold.
  • the first factor determining submodule is specifically used for:
  • the first initial adjustment factor is determined as the first variable adjustment factor.
  • the device also includes:
  • the adjustment factor determination module 1802 is further configured to determine a second variable adjustment factor based on the first latent variable
  • the third encoding result acquisition module is used to acquire the entropy encoding result of the fourth latent variable, the fourth latent variable is obtained after adjusting the third latent variable based on the second variable adjustment factor, and the third latent variable is obtained based on the second latent variable by The context model is determined;
  • the third encoding result writing module is used to write the entropy encoding result of the fourth latent variable and the encoding result of the second variable adjustment factor into the code stream;
  • the total number of encoded bits of the entropy encoding result of the second latent variable and the entropy encoding result of the fourth latent variable satisfies the preset encoding rate condition.
  • the adjustment factor determining module 1802 includes:
  • the first determining submodule is used to determine the corresponding context initial encoding bit number and initial entropy encoding model parameters through the context model based on the first latent variable;
  • the second determination sub-module is used to determine the number of coding bits of the entropy coding result of the first latent variable based on the initial entropy coding model parameters, so as to obtain the basic initial coding bit number;
  • the second factor determining submodule is configured to determine a first variable adjustment factor and a second variable adjustment factor based on the context initial encoding bit number, the basic initial encoding bit number and the target encoding bit number.
  • the second factor determining submodule is specifically used for:
  • the basic target number of encoded bits and the basic initial number of encoded bits, the first variable adjustment factor and the basic actual number of encoded bits are determined.
  • the basic actual number of encoded bits refers to the first variable adjusted by the first variable adjustment factor.
  • the second variable adjustment factor is determined based on the target number of coding bits of the context and the number of initial coding bits of the context.
  • the second factor determining submodule is specifically used for:
  • the second variable adjustment factor is determined based on the target number of coding bits of the context and the number of initial coding bits of the context.
  • the media data is an audio signal, a video signal or an image.
  • the number of encoded bits of the entropy encoding result of the second latent variable satisfies the preset encoding rate condition, it can be ensured that the encoding bit number of the entropy encoding result of the latent variable corresponding to each frame of media data can meet the preset encoding rate condition, that is Yes, it can ensure that the number of encoded bits of the entropy encoding result of the latent variable corresponding to each frame of media data is basically consistent, rather than changing dynamically, thus meeting the encoder's requirement for a stable encoding rate.
  • the need to transmit side information (such as window type, time-domain noise shaping (TNS: Temporal Noise Shaping) parameters, frequency-domain noise shaping (FDNS: Frequency-domain noise shaping) parameters, and/or bandwidth extension (BWE : bandwidth extension) parameter, etc.), it can ensure that the number of encoded bits of the entropy encoding result of the latent variable corresponding to each frame of media data and the encoded bit number of the side information are basically consistent, so as to meet the encoder's need for a stable encoding rate .
  • TMS Temporal Noise Shaping
  • FDNS Frequency-domain noise shaping
  • BWE bandwidth extension
  • the encoding device provided in the above-mentioned embodiments performs encoding, it only uses the division of the above-mentioned functional modules as an example for illustration. In practical applications, the above-mentioned function allocation can be completed by different functional modules based on needs. The internal structure of the system is divided into different functional modules to complete all or part of the functions described above.
  • the encoding device and the encoding method embodiments provided in the above embodiments belong to the same idea, and the specific implementation process thereof is detailed in the method embodiments, and will not be repeated here.
  • Figure 19 is a schematic structural diagram of a decoding device provided by an embodiment of the present application.
  • the decoding device can be implemented by software, hardware or a combination of the two to become part or all of the decoding end device.
  • the decoding end device can be as shown in Figure 1 destination device. Referring to FIG. 19 , the device includes: a first determination module 1901 , a variable adjustment module 1902 and a variable processing module 1903 .
  • the first determining module 1901 is configured to determine the reconstructed second latent variable and the reconstructed first variable adjustment factor based on the code stream. For the detailed implementation process, refer to the corresponding content in the foregoing embodiments, and details are not repeated here.
  • the variable adjustment module 1902 is configured to adjust the reconstructed second latent variable based on the reconstructed first variable adjustment factor to obtain the reconstructed first latent variable, and the reconstructed first latent variable is used to indicate the to-be-decoded Characteristics of media data.
  • the reconstructed first latent variable is used to indicate the to-be-decoded Characteristics of media data.
  • the variable processing module 1903 is configured to process the reconstructed first latent variables through the first decoding neural network model to obtain reconstructed media data.
  • the variable processing module 1903 is configured to process the reconstructed first latent variables through the first decoding neural network model to obtain reconstructed media data.
  • the first determining module 1901 includes:
  • a first determining submodule configured to determine a reconstructed third latent variable based on the code stream
  • the second determining submodule is configured to determine a reconstructed second latent variable based on the code stream and the reconstructed third latent variable.
  • the second determining submodule is specifically used for:
  • a reconstructed second latent variable is determined.
  • the first determining module 1901 includes:
  • the third determining submodule is used to determine the reconstructed fourth latent variable and the reconstructed second variable adjustment factor based on the code stream;
  • the fourth determining submodule is configured to determine the reconstructed second latent variable based on the code stream, the reconstructed fourth latent variable and the reconstructed second variable adjustment factor.
  • the fourth determining submodule is specifically used for:
  • a reconstructed second latent variable is determined.
  • the media data is an audio signal, a video signal or an image.
  • the encoding rate condition is set, that is, it can ensure that the number of encoded bits of the entropy encoding result of the latent variable corresponding to each frame of media data is basically consistent, rather than dynamically changing, thereby meeting the encoder's need for a stable encoding rate.
  • the need to transmit side information (such as window type, time-domain noise shaping (TNS: Temporal Noise Shaping) parameters, frequency-domain noise shaping (FDNS: Frequency-domain noise shaping) parameters, and/or bandwidth extension (BWE : bandwidth extension) parameter, etc.), it can ensure that the number of encoded bits of the entropy encoding result of the latent variable corresponding to each frame of media data and the encoded bit number of the side information are basically consistent, so as to meet the encoder's need for a stable encoding rate .
  • TMS Temporal Noise Shaping
  • FDNS Frequency-domain noise shaping
  • BWE bandwidth extension
  • the decoding device when the decoding device provided in the above-mentioned embodiments performs decoding, the division of the above-mentioned functional modules is used as an example for illustration. In practical applications, the above-mentioned function allocation can be completed by different functional modules based on needs, namely, the device The internal structure of the system is divided into different functional modules to complete all or part of the functions described above.
  • the decoding device and the decoding method embodiments provided in the above embodiments belong to the same idea, and the specific implementation process thereof is detailed in the method embodiments, and will not be repeated here.
  • FIG. 20 is a schematic block diagram of a codec device 2000 used in an embodiment of the present application.
  • the codec apparatus 2000 may include a processor 2001 , a memory 2002 and a bus system 2003 .
  • the processor 2001 and the memory 2002 are connected through the bus system 2003, the memory 2002 is used to store instructions, and the processor 2001 is used to execute the instructions stored in the memory 2002 to perform various encoding or decoding described in the embodiments of this application method. To avoid repetition, no detailed description is given here.
  • the processor 2001 can be a central processing unit (central processing unit, CPU), and the processor 2001 can also be other general-purpose processors, DSP, ASIC, FPGA or other programmable logic devices, discrete gates Or transistor logic devices, discrete hardware components, etc.
  • a general-purpose processor may be a microprocessor, or the processor may be any conventional processor, or the like.
  • the memory 2002 may include a ROM device or a RAM device. Any other suitable type of storage device may also be used as memory 2002 .
  • Memory 2002 may include code and data 20021 accessed by processor 2001 using bus 2003 .
  • the memory 2002 may further include an operating system 20023 and an application program 20022, where the application program 20022 includes at least one program that allows the processor 2001 to execute the encoding or decoding method described in the embodiment of this application.
  • the application program 20022 may include applications 1 to N, which further include an encoding or decoding application (codec application for short) that executes the encoding or decoding method described in the embodiment of this application.
  • the bus system 2003 may include not only a data bus, but also a power bus, a control bus, and a status signal bus. However, for clarity of illustration, the various buses are labeled as bus system 2003 in the figure.
  • the codec apparatus 2000 may also include one or more output devices, such as a display 2004 .
  • display 2004 may be a touch-sensitive display that incorporates a display with a haptic unit operable to sense touch input.
  • the display 2004 can be connected to the processor 2001 via the bus 2003 .
  • the codec apparatus 2000 may execute the encoding method in the embodiment of the present application, and may also execute the decoding method in the embodiment of the present application.
  • Computer-readable media may include computer-readable storage media, which correspond to tangible media, such as data storage media, or communication media including any medium that facilitates transfer of a computer program from one place to another (eg, based on a communication protocol) .
  • a computer-readable medium may generally correspond to (1) a non-transitory tangible computer-readable storage medium, or (2) a communication medium, such as a signal or carrier wave.
  • Data storage media may be any available media that can be accessed by one or more computers or one or more processors to retrieve instructions, code and/or data structures for implementation of the techniques described in this application.
  • a computer program product may include a computer readable medium.
  • such computer-readable storage media may include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage, flash memory, or any other medium that can contain the desired program code in the form of a computer and can be accessed by a computer.
  • any connection is properly termed a computer-readable medium.
  • coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave
  • coaxial cable Wire, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of media.
  • Disk and disc includes compact disc (CD), laser disc, optical disc, DVD and blu-ray disc where disks usually reproduce data magnetically, while discs reproduce data optically with lasers. Combinations of the above should also be included within the scope of computer-readable media.
  • DSPs digital signal processors
  • ASICs application specific integrated circuits
  • FPGAs field programmable logic arrays
  • DSPs digital signal processors
  • ASICs application specific integrated circuits
  • FPGAs field programmable logic arrays
  • DSPs digital signal processors
  • ASICs application specific integrated circuits
  • FPGAs field programmable logic arrays
  • the term "processor,” as used herein may refer to any of the foregoing structure or any other structure suitable for implementation of the techniques described herein.
  • the functionality described by the various illustrative logical blocks, modules, and steps described herein may be provided within dedicated hardware and/or software modules configured for encoding and decoding, or in conjunction with into the combined codec.
  • the techniques may be fully implemented in one or more circuits or logic elements.
  • various illustrative logical blocks, units, and modules in the encoder 100 and the decoder 200 may be understood as corresponding circuit devices or logic elements.
  • inventions of the present application may be implemented in a wide variety of devices or devices, including a wireless handset, an integrated circuit (IC), or a group of ICs (eg, a chipset).
  • IC integrated circuit
  • a group of ICs eg, a chipset
  • Various components, modules or units are described in the embodiments of the present application to emphasize the functional aspects of the apparatus for performing the disclosed technology, but they do not necessarily need to be realized by different hardware units. Indeed, as described above, the various units may be combined in a codec hardware unit in conjunction with suitable software and/or firmware, or by interoperating hardware units (comprising one or more processors as described above) to supply.
  • all or part of them may be implemented by software, hardware, firmware or any combination thereof.
  • software When implemented using software, it may be implemented in whole or in part in the form of a computer program product.
  • the computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on the computer, the processes or functions according to the embodiments of the present application will be generated in whole or in part.
  • the computer can be a general purpose computer, a special purpose computer, a computer network or other programmable devices.
  • the computer instructions may be stored in or transmitted from one computer-readable storage medium to another computer-readable storage medium, for example, the computer instructions may be transmitted from a website, computer, server or data center Transmission to another website site, computer, server or data center by wired (eg coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (eg infrared, wireless, microwave, etc.).
  • the computer-readable storage medium may be any available medium that can be accessed by a computer, or may be a data storage device such as a server or a data center integrated with one or more available media.
  • the available medium may be a magnetic medium (for example: floppy disk, hard disk, magnetic tape), an optical medium (for example: digital versatile disc (digital versatile disc, DVD)) or a semiconductor medium (for example: solid state disk (solid state disk, SSD)) Wait.
  • a magnetic medium for example: floppy disk, hard disk, magnetic tape
  • an optical medium for example: digital versatile disc (digital versatile disc, DVD)
  • a semiconductor medium for example: solid state disk (solid state disk, SSD)
  • the computer-readable storage medium mentioned in the embodiment of the present application may be a non-volatile storage medium, in other words, may be a non-transitory storage medium.

Landscapes

  • Engineering & Computer Science (AREA)
  • Multimedia (AREA)
  • Signal Processing (AREA)
  • Physics & Mathematics (AREA)
  • Acoustics & Sound (AREA)
  • Computational Linguistics (AREA)
  • Health & Medical Sciences (AREA)
  • Audiology, Speech & Language Pathology (AREA)
  • Human Computer Interaction (AREA)
  • Computer Networks & Wireless Communication (AREA)
  • Artificial Intelligence (AREA)
  • Evolutionary Computation (AREA)
  • Databases & Information Systems (AREA)
  • Quality & Reliability (AREA)
  • Compression, Expansion, Code Conversion, And Decoders (AREA)
  • Compression Or Coding Systems Of Tv Signals (AREA)

Abstract

本申请实施例公开了一种编解码方法、装置、设备、存储介质及计算机程序,属于编解码技术领域。在该方法中,通过第一变量调整因子对第一潜在变量进行调整,得到第二潜在变量,而且第二潜在变量的熵编码结果的编码比特数满足预设编码速率条件,这样可以保证每帧媒体数据对应的潜在变量的熵编码结果的编码比特数均能够满足预设编码速率条件,也即是,可以保证每帧媒体数据对应的潜在变量的熵编码结果的编码比特数基本保持一致,而不是动态变化的,从而满足了编码器对稳定编码速率的需求。

Description

编解码方法、装置、设备、存储介质及计算机程序
本申请要求于2021年05月21日提交的申请号为202110559102.7、发明名称为“编解码方法、装置、设备、存储介质及计算机程序”的中国专利申请的优先权,其全部内容通过引用结合在本申请中。
技术领域
本申请实施例涉及编解码技术领域,特别涉及一种编解码方法、装置、设备、存储介质及计算机程序。
背景技术
编解码技术是媒体通信和媒体广播等媒体应用中不可或缺的环节。因此如何进行编解码成为业界的关注点之一。
相关技术提出了一种音频信号的编解码方法,在该方法中,当进行音频信号的编码时,对时域的音频信号进行修正的离散余弦变换(modified discrete cosine transform,MDCT)处理,以得到频域的音频信号。通过编码神经网络模型对频域的音频信号进行处理,得到潜在变量,该潜在变量用于指示频域的音频信号的特征。对该潜在变量进行量化处理,得到量化后的潜在变量,对量化后的潜在变量进行熵编码,并将熵编码结果写入码流。当进行音频信号的解码时,基于该码流确定量化后的潜在变量,将量化后的潜在变量进行去量化处理,得到潜在变量。通过解码神经网络模型对该潜在变量进行处理,得到频域的音频信号,对频域的音频信号进行修正的离散余弦逆变换(Inverse modified discrete cosine transform,IMDCT)处理,以得到重构的时域的音频信号。
然而,由于熵编码是对不同概率的元素采用不同的比特数来编码,所以对于相邻两帧音频信号来说,这两帧音频信号对应的潜在变量中的元素出现的概率可能不同,导致这两帧音频信号的潜在变量的编码比特数不同,从而无法满足稳定编码速率的需求。
发明内容
本申请实施例提供了一种编解码方法、装置、设备、存储介质及计算机程序,可以满足编码器稳定编码速率的需求。所述技术方案如下:
第一方面,提供了一种编码方法,该方法可以应用于不包括上下文模型的编解码器中,也可以应用于包括上下文模型的编解码器中。而且,不仅可以对待编码的媒体数据生成的潜在变量通过调整因子进行调整,还可以对上下文模型确定的潜在变量通过调整因子进行调整。因此,接下来将分为多种情况,对该方法进行详细地解释说明。
第一种情况,通过第一编码神经网络模型对待编码的媒体数据进行处理,以得到第一潜在变量,第一潜在变量用于指示待编码的媒体数据的特征;基于第一潜在变量确定第一变量调整因子,第一变量调整因子用于使得第二潜在变量的熵编码结果的编码比特数满足预设编 码速率条件,第二潜在变量是通过所述第一变量调整因子对所述第一潜在变量调整后得到;获取第二潜在变量的熵编码结果;将第二潜在变量的熵编码结果以及第一变量调整因子的编码结果写入码流。
由于第二潜在变量的熵编码结果的编码比特数满足预设编码速率条件,这样可以保证每帧媒体数据对应的潜在变量的熵编码结果的编码比特数均能够满足预设编码速率条件,也即是,可以保证每帧媒体数据对应的潜在变量的熵编码结果的编码比特数基本保持一致,而不是动态变化的,从而满足了编码器对稳定编码速率的需求。进一步地,在考虑到需要传输边信息(例如窗型,时域噪声整形(TNS:Temporal Noise Shaping)参数,频域噪声整形(FDNS:Frequency-domain noise shaping)参数,和/或带宽扩展(BWE:bandwidth extension)参数等等)时,能够保证每帧媒体数据对应的潜在变量的熵编码结果的编码比特数和边信息的编码比特数整体基本保持一致,从而满足编码器对稳定编码速率的需求。
其中,待编码的媒体数据为音频信号、视频信号或者图像等。而且,待编码的媒体数据的形式可以为任何一种形式,本申请实施例对此不做限定。
通过第一编码神经网络模型对待编码的媒体数据进行处理的实现过程为:将待编码的媒体数据输入第一编码神经网络模型,得到第一编码神经网络模型输出的第一潜在变量。或者,对待编码的媒体数据进行预处理,将预处理后的媒体数据输入第一编码神经网络模型,得到第一编码神经网络模型输出的第一潜在变量。
也就是说,可以将待编码的媒体数据作为第一编码神经网络模型的输入来确定第一潜在变量,也可以对待编码的媒体数据进行预处理后,再作为第一编码神经网络模型的输入来确定第一潜在变量。
可选地,在使用固定码率对待编码的媒体数据进行编码的情况下,满足预设编码速率条件包括编码比特数小于或等于目标编码比特数;或者,满足预设编码速率条件包括编码比特数小于或等于目标编码比特数,且编码比特数与目标编码比特数的差值小于比特数阈值;或者,满足预设编码速率条件包括编码比特数为小于或等于目标编码比特数的最大编码比特数。
可选地,在使用可变码率对待编码的媒体数据进行编码的情况下,满足预设编码速率条件包括编码比特数与目标编码比特数的差值的绝对值小于比特数阈值。也即是,满足预设编码速率条件包括编码比特数小于或等于目标编码比特数,且目标编码比特数与编码比特数的差值小于比特数阈值;或者,满足预设编码速率条件包括编码比特数大于或等于目标编码比特数,且编码比特数与目标编码比特数的差值小于比特数阈值。
其中,目标编码比特数可以为事先设置的。当然,目标编码比特数也可以基于编码速率来确定,且不同的编码速率对应不同的目标编码比特数。
在本申请实施例中,可以使用固定码率对待编码的媒体数据进行编码,也可以使用可变码率对待编码的媒体数据进行编码。
在使用固定码率对待编码的媒体数据进行编码的情况下,可以基于固定码率确定当前帧的待编码的媒体数据的比特数,再减去当前帧的已使用比特数,得到当前帧的目标编码比特数。其中,已使用比特数可以为边信息等进行编码的比特数,而且通常情况下,每帧媒体数据的边信息不同,所以,每帧媒体数据的目标编码比特数通常是不同的。
在使用可变码率对待编码的媒体数据进行编码的情况下,通常会指定一个码率,实际码率会在指定的码率的上下波动。此时,可以基于指定的码率确定当前帧的待编码的媒体数据 的比特数,再减去当前帧的已使用比特数,得到当前帧的目标编码比特数。其中,已使用比特数可以为边信息等进行编码的比特数,而且在某些情况下,不同帧的媒体数据的边信息可以是不同的,所以,不同帧的媒体数据的目标编码比特数通常是不同的。
其中,可以基于第一潜在变量确定初始编码比特数,基于初始编码比特数和目标编码比特数,确定第一变量调整因子。
其中,初始编码比特数为第一潜在变量的熵编码结果的编码比特数;或者,初始编码比特数为经过第一初始调整因子调整后的第一潜在变量的熵编码结果的编码比特数。其中,第一初始调整因子可以为第一预设调整因子。
其中,基于第一初始调整因子对第一潜在变量进行调整的实现过程为:将第一潜在变量中的各个元素与第一初始调整因子中对应的元素相乘,得到调整后的第一潜在变量。
值得注意的是,上述实现过程仅仅为一种示例,实际应用中,还可以采用其他的方法来调整。比如,可以将第一潜在变量中的各个元素除以第一初始调整因子中对应的元素,得到调整后的第一潜在变量。本申请实施例对调整方法不做限定。
其中,基于初始编码比特数和目标编码比特数,确定第一变量调整因子的实现方式可以包括多种,接下来对其中的三种进行介绍。
第一种实现方式,在初始编码比特数等于目标编码比特数的情况下,将第一初始调整因子确定为第一变量调整因子。在初始编码比特数不等于目标编码比特数的情况下,基于初始编码比特数和目标编码比特数,通过第一循环方式确定第一变量调整因子。
其中,第一循环方式的第i次循环处理包括如下步骤:确定第i次循环处理的调整因子,i为正整数,基于第i次循环处理的调整因子对第一潜在变量进行调整,得到第i次调整后的第一潜在变量。确定第i次调整后的第一潜在变量的熵编码结果的编码比特数,以得到第i次的编码比特数。在第i次的编码比特数满足继续调整条件的情况下,执行第一循环方式的第i+1次循环处理。在第i次的编码比特数不满足继续调整条件的情况下,终止第一循环方式的执行,基于第i次循环处理的调整因子确定第一变量调整因子。
其中,确定第i次循环处理的调整因子的实现过程为:基于第一循环方式的第i-1次循环处理的调整因子、第i-1次的编码比特数和目标编码比特数,确定第i次循环处理的调整因子。其中,在i=1的情况下,第i-1次循环处理的调整因子为第一初始调整因子,第i-1次的编码比特数为初始编码比特数。
此时,继续调整条件包括第i-1次的编码比特数和第i次的编码比特数均小于目标编码比特数,或者,继续调整条件包括第i-1次的编码比特数和第i次的编码比特数均大于目标编码比特数。
换句话说,继续调整条件包括第i次的编码比特数没有越过目标编码比特数。这里的没有越过的意思是指:前i-1次的编码比特数一直小于目标编码比特数,第i次的编码比特数仍小于目标编码比特数。或者,前i-1次的编码比特数一直大于目标编码比特数,第i次的编码比特数仍大于目标编码比特数。相反地,越过的意思是指:前i-1次的编码比特数一直小于目标编码比特数,第i次的编码比特数大于目标编码比特数。或者,前i-1次的编码比特数一直大于目标编码比特数,第i次的编码比特数小于目标编码比特数。
其中,基于第i次循环处理的调整因子确定第一变量调整因子的实现过程包括:在第i次的编码比特数等于目标编码比特数的情况下,将第i次循环处理的调整因子确定为第一变量 调整因子。或者,在第i次的编码比特数不等于目标编码比特数的情况下,基于第i次循环处理的调整因子和第一循环方式的第i-1次循环处理的调整因子,确定第一变量调整因子。
也即是,第i次循环处理的调整因子为通过上述第一循环方式最后一次得到的调整因子,第i次的编码比特数为最后一次得到的编码比特数。在最后一次得到的编码比特数等于目标编码比特数的情况下,将最后一次得到的调整因子确定为第一变量调整因子。在最后一次得到的编码比特数不等于目标编码比特数的情况下,基于后两次得到的调整因子确定第一变量调整因子。
其中,基于第i次循环处理的调整因子和第一循环方式的第i-1次循环处理的调整因子,确定第一变量调整因子的实现过程包括:确定第i次循环处理的调整因子和第i-1次循环处理的调整因子的平均值,基于该平均值确定第一变量调整因子。
其中,可以直接将该平均值确定为第一变量调整因子,也可以将该平均值乘以一个预设的常数,得到第一变量调整因子。可选地,该常数可以小于1。
当然,基于第i次循环处理的调整因子和第一循环方式的第i-1次循环处理的调整因子,确定第一变量调整因子的实现过程还可以为:基于第i次循环处理的调整因子和第i-1次循环处理的调整因子,通过第二循环方式确定第一变量调整因子。
作为一种示例,第二循环方式的第j次循环处理包括如下步骤:基于第j次循环处理的第一调整因子和第j次循环处理的第二调整因子确定第j次循环处理的第三调整因子,其中,在j等于1的情况下,第j次循环处理的第一调整因子为第i次循环处理的调整因子和第i-1次循环处理的调整因子中的一者,第j次循环处理的第二调整因子为第i次循环处理的调整因子和第i-1次循环处理的调整因子中的另一者,第j次循环处理的第一调整因子对应第j次的第一编码比特数,第j次循环处理的第二调整因子对应第j次的第二编码比特数,第j次的第一编码比特数是指经过第j次循环处理的第一调整因子调整后的第一潜在变量的熵编码结果的编码比特数,第j次的第二编码比特数是指经过第j次循环处理的第二调整因子调整后的第一潜在变量的熵编码结果的编码比特数,第j次的第一编码比特数小于第j次的第二编码比特数。获取第j次的第三编码比特数,第j次的第三编码比特数是指经过第j次循环处理的第三调整因子调整后的第一潜在变量的熵编码结果的编码比特数。若第j次的第三编码比特数不满足继续循环条件,终止第二循环方式的执行,将第j次循环处理的第三调整因子确定为第一变量调整因子。若第j次的第三编码比特数满足继续循环条件、第j次的第三编码比特数大于目标编码比特数且小于第j次的第二编码比特数,将第j次循环处理的第三调整因子作为第j+1次循环处理的第二调整因子,将第j次循环处理的第一调整因子作为第j+1次循环处理的第一调整因子,执行第二循环方式的第j+1次循环处理。若第j次的第三编码比特数满足继续循环条件、第j次的第三编码比特数小于目标编码比特数且大于第j次的第一编码比特数,将第j次循环处理的第三调整因子作为第j+1次循环处理的第一调整因子,将第j次循环处理的第二调整因子作为第j+1次循环处理的第二调整因子,执行第二循环方式的第j+1次循环处理。
作为另一种示例,第二循环方式的第j次循环处理包括如下步骤:基于第j次循环处理的第一调整因子和第j次循环处理的第二调整因子确定第j次循环处理的第三调整因子,其中,在j等于1的情况下,第j次循环处理的第一调整因子为第i次循环处理的调整因子和第i-1次循环处理的调整因子中的一者,第j次循环处理的第二调整因子为第i次循环处理的调整 因子和第i-1次循环处理的调整因子中的另一者,第j次循环处理的第一调整因子对应第j次的第一编码比特数,第j次循环处理的第二调整因子对应第j次的第二编码比特数,第j次的第一编码比特数小于第j次的第二编码比特数,j为正整数。获取第j次的第三编码比特数,第j次的第三编码比特数是指经过第j次循环处理的第三调整因子调整后的第一潜在变量的熵编码结果的编码比特数。若第j次的第三编码比特数不满足继续循环条件,终止第二循环方式的执行,将第j次循环处理的第三调整因子确定为第一变量调整因子。若j达到最大循环次数且第j次的第三编码比特数满足继续循环条件,终止第二循环方式的执行,基于第j次循环处理的第一调整因子确定第一变量调整因子。若j未达到最大循环次数、第j次的第三编码比特数满足继续循环条件、第j次的第三编码比特数大于目标编码比特数且小于第j次的第二编码比特数,将第j次循环处理的第三调整因子作为第j+1次循环处理的第二调整因子,将第j次循环处理的第一调整因子作为第j+1次循环处理的第一调整因子,执行第二循环方式的第j+1次循环处理。若j未达到最大循环次数、第j次的第三编码比特数满足继续循环条件、第j次的第三编码比特数小于目标编码比特数且大于第j次的第一编码比特数,将第j次循环处理的第三调整因子作为第j+1次循环处理的第一调整因子,将第j次循环处理的第二调整因子作为第j+1次循环处理的第二调整因子,执行第二循环方式的第j+1次循环处理。
其中,基于第j次循环处理的第一调整因子和第j次循环处理的第二调整因子确定第j次循环处理的第三调整因子的实现过程包括:确定第j次循环处理的第一调整因子和第j次循环处理的第二调整因子的平均值,基于该平均值确定第j次循环处理的第三调整因子。作为一种示例,可以直接将该平均值确定为第j次循环处理的第三调整因子,也可以将该平均值乘以一个预设的常数,得到第j次循环处理的第三调整因子。可选地,该常数可以小于1。
另外,获取第j次的第三编码比特数的实现过程包括:基于第j次循环处理的第三调整因子对第一潜在变量进行调整,得到调整后的第一潜在变量,对调整后的第一潜在变量进行量化处理,得到量化后的第一潜在变量。对量化后的第一潜在变量进行熵编码,统计该熵编码结果的编码比特数,得到第j次的第三编码比特数。
其中,在使用固定码率对待编码的媒体数据进行编码的情况下,基于第j次循环处理的第一调整因子确定第一变量调整因子的实现过程包括:将第j次循环处理的第一调整因子确定为第一变量调整因子。在使用可变码率对待编码的媒体数据进行编码的情况下,基于第j次循环处理的第一调整因子确定第一变量调整因子的实现过程包括:确定目标编码比特数与第j次的第一编码比特数之间的第一差值,以及确定第j次的第二编码比特数与目标编码比特数之间的第二差值。若第一差值小于第二差值,将第j次循环处理的第一调整因子确定为第一变量调整因子。若第二差值小于第一差值,将第j次循环处理的第二调整因子确定为第一变量调整因子。若第一差值等于第二差值,将第j次循环处理的第一调整因子确定为第一变量调整因子,或者将第j次循环处理的第二调整因子确定为第一变量调整因子。
其中,在使用固定码率对待编码的媒体数据进行编码的情况下,继续循环条件包括第j次的第三编码比特数大于目标编码比特数,或者,继续循环条件包括第j次的第三编码比特数小于目标编码比特数,且目标编码比特数与第j次的第三编码比特数的差值大于比特数阈值。在使用可变码率对待编码的媒体数据进行编码的情况下,继续循环条件包括目标编码比特数与第j次的第三编码比特数的差值的绝对值大于比特数阈值。也即是,继续循环条件包括第j次的第三编码比特数大于目标编码比特数,且第j次的第三编码比特数与目标编码比特数的 差值大于比特数阈值,或者,继续循环条件包括第j次的第三编码比特数小于目标编码比特数,且目标编码比特数与第j次的第三编码比特数的差值大于比特数阈值。
第二种实现方式,在初始编码比特数等于目标编码比特数的情况下,将第一初始调整因子确定为第一变量调整因子。在初始编码比特数不等于目标编码比特数的情况下,按照上述第一种实现方式来确定第一变量调整因子。但是,与上述第一种实现方式不同的是,在初始编码比特数小于目标编码比特数的情况下,按照第一步长调整第一循环方式的第i-1次循环处理的调整因子,得到第i次循环处理的调整因子,此时,继续调整条件包括第i次的编码比特数小于目标编码比特数。在初始编码比特数大于目标编码比特数的情况下,按照第二步长调整第一循环方式的第i-1次循环处理的调整因子,得到第i次循环处理的调整因子,此时,继续调整条件包括第i次的编码比特数大于目标编码比特数。
其中,按照第一步长调整第一循环方式的第i-1次循环处理的调整因子可以是指按照第一步长增大第i-1次循环处理的调整因子,按照第二步长调整第一循环方式的第i-1次循环处理的调整因子可以是指按照第二步长减小第i-1次循环处理的调整因子。
上述的增大处理和减小处理可以为线性的,也可以为非线性的。示例地,可以将第i-1次循环处理的调整因子与第一步长之和确定为第i次循环处理的调整因子,可以将第i-1次循环处理的调整因子与第二步长之差确定为第i次循环处理的调整因子。
需要说明的是,第一步长和第二步长可以为事先设置的,且第一步长和第二步长可以基于不同的需求来调整。另外,第一步长与第二步长可以相等,也可以不相等。
第三种实现方式,在初始编码比特数小于或等于目标编码比特数的情况下,将第一初始调整因子确定为第一变量调整因子。在初始编码比特数大于目标编码比特数的情况下,可以按照上述第一种实现方式来确定第一变量调整因子。也可以按照上述第二种实现方式中初始编码比特数大于目标编码比特数的情况来确定第一变量调整因子。
第二种情况,在第一种情况的基础上还获取第三潜在变量的熵编码结果,第三潜在变量是基于第二潜在变量通过上下文模型确定得到,且第三潜在变量用于指示第二潜在变量的概率分布;将第三潜在变量的熵编码结果写入码流。其中,第二潜在变量的熵编码结果和第三潜在变量的熵编码结果的编码总比特数满足预设编码速率条件。
其中,基于第一潜在变量确定第一变量调整因子,包括:基于第一潜在变量确定初始编码比特数;基于初始编码比特数和目标编码比特数,确定第一变量调整因子。
其中,基于第一潜在变量确定初始编码比特数的方式可以包括两种,接下来将分别介绍。
第一种实现方式,基于第一潜在变量,通过上下文模型确定对应的上下文初始编码比特数和初始熵编码模型参数,基于初始熵编码模型参数,确定第一潜在变量的熵编码结果的编码比特数,以得到基础初始编码比特数。基于上下文初始编码比特数与基础初始编码比特数确定为初始编码比特数。
其中,上下文模型包括上下文编码神经网络模型和上下文解码神经网络模型。基于第一潜在变量,通过上下文模型确定对应的上下文初始编码比特数和初始熵编码模型参数的实现过程包括:通过上下文编码神经网络模型对第一潜在变量进行处理,得到第五潜在变量,第五潜在变量用于指示第一潜在变量的概率分布。确定第五潜在变量的熵编码结果,将第五潜在变量的熵编码结果的编码比特数作为上下文初始编码比特数。基于第五潜在变量的熵编码结果重构第五潜在变量,通过上下文解码神经网络模型对重构得到的第五潜在变量进行处理, 得到初始熵编码模型参数。
第二种实现方式,将第一预设调整因子作为第一初始调整因子,基于第一初始调整因子对第一潜在变量进行调整,基于调整后的第一潜在变量,通过上下文模型确定对应的上下文初始编码比特数和初始熵编码模型参数,基于初始熵编码模型参数,确定调整后的第一潜在变量的熵编码结果的编码比特数,得到基础初始编码比特数。基于上下文初始编码比特数与基础初始编码比特数确定初始编码比特数。
其中,上下文模型包括上下文编码神经网络模型和上下文解码神经网络模型。基于调整后的第一潜在变量,通过上下文模型确定对应的上下文初始编码比特数和初始熵编码模型参数的实现过程包括:通过上下文编码神经网络模型对调整后的第一潜在变量进行处理,得到第六潜在变量,第六潜在变量用于指示调整后的第一潜在变量的概率分布。确定第六潜在变量的熵编码结果,将第六潜在变量的熵编码结果的编码比特数作为上下文初始编码比特数。基于第六潜在变量的熵编码结果重构第六潜在变量,通过上下文解码神经网络模型对重构得到的第六潜在变量进行处理,得到初始熵编码模型参数。
其中,基于初始编码比特数和目标编码比特数,确定第一变量调整因子的实现方式可以包括多种,接下来对其中的三种进行介绍。
第一种实现方式,在初始编码比特数等于目标编码比特数的情况下,将第一初始调整因子确定为第一变量调整因子。在初始编码比特数不等于目标编码比特数的情况下,基于初始编码比特数和目标编码比特数,通过第一循环方式确定第一变量调整因子。
其中,第一循环方式的第i次循环处理包括如下步骤:确定第i次循环处理的调整因子,i为正整数,基于第i次循环处理的调整因子对第一潜在变量进行调整,得到第i次调整后的第一潜在变量。基于第i次调整后的第一潜在变量,通过上下文模型确定对应的第i次的上下文编码比特数和第i次的熵编码模型参数,基于第i次的熵编码模型参数,确定第i次调整后的第一潜在变量的熵编码结果的编码比特数,得到第i次的基础编码比特数,基于第i次的上下文编码比特数与第i次的基础编码比特数确定第i次的编码比特数。在第i次的编码比特数满足继续调整条件的情况下,执行第一循环方式的第i+1次循环处理。在第i次的编码比特数不满足继续调整条件的情况下,终止第一循环方式的执行,基于第i次循环处理的调整因子确定第一变量调整因子。
第二种实现方式,在初始编码比特数等于目标编码比特数的情况下,将第一初始调整因子确定为第一变量调整因子。在初始编码比特数不等于目标编码比特数的情况下,按照上述第一种实现方式来确定第一变量调整因子。但是,与上述第一种实现方式不同的是,在初始编码比特数小于目标编码比特数的情况下,按照第一步长调整第一循环方式的第i-1次循环处理的调整因子,得到第i次循环处理的调整因子,此时,继续调整条件包括第i次的编码比特数小于目标编码比特数。在初始编码比特数大于目标编码比特数的情况下,按照第二步长调整第一循环方式的第i-1次循环处理的调整因子,得到第i次循环处理的调整因子,此时,继续调整条件包括第i次的编码比特数大于目标编码比特数。
其中,按照第一步长调整第一循环方式的第i-1次循环处理的调整因子可以是指按照第一步长增大第i-1次循环处理的调整因子,按照第二步长调整第一循环方式的第i-1次循环处理的调整因子可以是指按照第二步长减小第i-1次循环处理的调整因子。
上述的增大处理和减小处理可以为线性的,也可以为非线性的。示例地,可以将第i-1次 循环处理的调整因子与第一步长之和确定为第i次循环处理的调整因子,可以将第i-1次循环处理的调整因子与第二步长之差确定为第i次循环处理的调整因子。
需要说明的是,第一步长和第二步长可以为事先设置的,且第一步长和第二步长可以基于不同的需求来调整。另外,第一步长与第二步长可以相等,也可以不相等。
第三种实现方式,在初始编码比特数小于或等于目标编码比特数的情况下,将第一初始调整因子确定为第一变量调整因子。在初始编码比特数大于目标编码比特数的情况下,可以按照上述第一种实现方式来确定第一变量调整因子。也可以按照上述第二种实现方式中初始编码比特数大于目标编码比特数的情况来确定第一变量调整因子。
第三种情况,在第一种情况的基础上还基于第一潜在变量确定第二变量调整因子;获取第四潜在变量的熵编码结果,第四潜在变量是基于第二变量调整因子对第三潜在变量调整后得到,第三潜在变量是基于第二潜在变量通过上下文模型确定得到,且第三潜在变量用于指示第二潜在变量的概率分布;将第四潜在变量的熵编码结果以及第二变量调整因子的编码结果写入码流;其中,第二潜在变量的熵编码结果和第四潜在变量的熵编码结果的编码总比特数满足预设编码速率条件。
其中,基于第一潜在变量确定第一变量调整因子和第二变量调整因子的方式可以包括两种,分别为:
第一种实现方式,基于第一潜在变量,通过上下文模型确定对应的上下文初始编码比特数和初始熵编码模型参数;基于初始熵编码模型参数,确定第一潜在变量的熵编码结果的编码比特数,以得到基础初始编码比特数;基于上下文初始编码比特数、基础初始编码比特数和目标编码比特数,确定第一变量调整因子和第二变量调整因子。
第二种实现方式,将第一预设调整因子作为第一初始调整因子,将第二预设调整因子作为第二初始调整因子。基于第一初始调整因子对第一潜在变量进行调整,基于调整后的第一潜在变量和第二初始调整因子,通过上下文模型确定对应的上下文初始编码比特数和初始熵编码模型参数,基于初始熵编码模型参数,确定第一潜在变量的熵编码结果的编码比特数,以得到基础初始编码比特数。基于上下文初始编码比特数、基础初始编码比特数和目标编码比特数,确定第一变量调整因子和第二变量调整因子。
其中,上下文模型包括上下文编码神经网络模型和上下文解码神经网络模型。基于调整后的第一潜在变量和第二初始调整因子,通过上下文模型确定对应的上下文初始编码比特数和初始熵编码模型参数的实现过程包括:通过上下文编码神经网络模型对调整后的第一潜在变量进行处理,得到第六潜在变量,第六潜在变量用于指示调整后的第一潜在变量的概率分布。基于第二初始调整因子对第六潜在变量进行调整,得到第七潜在变量。确定第七潜在变量的熵编码结果,将第七潜在变量的熵编码结果的编码比特数作为上下文初始编码比特数。基于第七潜在变量的熵编码结果重构第七潜在变量,通过上下文解码神经网络模型对重构得到的第七潜在变量进行处理,得到初始熵编码模型参数。
其中,基于第二初始调整因子对第六潜在变量进行调整的实现过程为:将第六潜在变量中的各个元素与第二初始调整因子中对应的元素相乘,得到第七潜在变量。
值得注意的是,上述实现过程仅仅为一种示例,实际应用中,还可以采用其他的方法来调整。比如,可以将第六潜在变量中的各个元素除以第二初始调整因子中对应的元素,得到第七潜在变量。本申请实施例对调整方法不做限定。
其中,基于上下文初始编码比特数、基础初始编码比特数和目标编码比特数,确定第一变量调整因子和第二变量调整因子的实现过程可以包括两种,接下来分别介绍。
第一种实现方式,将第二变量调整因子置为第二初始调整因子,基于基础初始编码比特数和上下文初始编码比特数中的至少一者,以及目标编码比特数,确定基础目标编码比特数。基于第二初始调整因子、基础目标编码比特数和基础初始编码比特数,确定第一变量调整因子和基础实际编码比特数,基础实际编码比特数是指经过第一变量调整因子调整后的第一潜在变量的熵编码结果的编码比特数。基于目标编码比特数和基础实际编码比特数,确定上下文目标编码比特数。基于上下文目标编码比特数和上下文初始编码比特数,确定第二变量调整因子。
其中,基于基础初始编码比特数和上下文初始编码比特数中的至少一者,以及目标编码比特数,确定基础目标编码比特数的实现过程包括:将目标编码比特数减去上下文初始编码比特数,得到基础目标编码比特数。或者,确定基础初始编码比特数和上下文初始编码比特数之间的比例,基于该比例和目标编码比特数,确定基础目标编码比特数。或者,基于目标编码比特数与基础初始编码比特数的比值,确定基础目标编码比特数。当然,还可以通过其他实现过程来确定。
其中,基于第二初始调整因子、基础目标编码比特数和基础初始编码比特数确定第一变量调整因子的实现方式也包括三种,接下来分别介绍。
方式11,在基础初始编码比特数等于基础目标编码比特数的情况下,将第一初始调整因子确定为第一变量调整因子。在基础初始编码比特数不等于基础目标编码比特数的情况下,基于基础初始编码比特数和基础目标编码比特数,通过第一循环方式确定第一变量调整因子。
其中,第一循环方式的第i次循环处理包括如下步骤:确定第i次循环处理的调整因子,i为正整数,基于第i次循环处理的调整因子对第一潜在变量进行调整,得到第i次调整后的第一潜在变量。基于第i次调整后的第一潜在变量和第二初始调整因子,通过上下文模型确定第i次的上下文编码比特数和第i次的熵编码模型参数,基于第i次的熵编码模型参数,确定第i次调整后的第一潜在变量的熵编码结果的编码比特数,得到第i次的基础编码比特数,基于第i次的上下文编码比特数与第i次的基础编码比特数确定第i次的编码比特数。在第i次的编码比特数满足继续调整条件的情况下,执行第一循环方式的第i+1次循环处理。在第i次的编码比特数不满足继续调整条件的情况下,终止第一循环方式的执行,基于第i次循环处理的调整因子确定第一变量调整因子。
方式12,在基础初始编码比特数等于基础目标编码比特数的情况下,将第一初始调整因子确定为第一变量调整因子。在基础初始编码比特数不等于基础目标编码比特数的情况下,按照上述方式11来确定第一变量调整因子。但是,与上述方式11不同的是,在基础初始编码比特数小于基础目标编码比特数的情况下,按照第一步长调整第一循环方式的第i-1次循环处理的调整因子,得到第i次循环处理的调整因子,此时,继续调整条件包括第i次的编码比特数小于基础目标编码比特数。在基础初始编码比特数大于基础目标编码比特数的情况下,按照第二步长调整第一循环方式的第i-1次循环处理的调整因子,得到第i次循环处理的调整因子,此时,继续调整条件包括第i次的编码比特数大于基础目标编码比特数。
方式13,在基础初始编码比特数小于或等于基础目标编码比特数的情况下,将第一初始调整因子确定为第一变量调整因子。在基础初始编码比特数大于基础目标编码比特数的情况 下,可以按照上述方式11来确定第一变量调整因子。也可以按照上述方式12中基础初始编码比特数大于基础目标编码比特数的情况来确定第一变量调整因子。
其中,基于上下文目标编码比特数和上下文初始编码比特数确定第二变量调整因子的实现过程与基于基础初始编码比特数和基础目标编码比特数确定第一变量调整因子的实现过程类似。接下来也分为三种方式分别介绍。
方式21,在上下文初始编码比特数等于上下文目标编码比特数的情况下,将第二初始调整因子确定为第二变量调整因子。在上下文初始编码比特数不等于上下文目标编码比特数的情况下,基于第一变量调整因子对第一潜在变量进行调整,得到第二潜在变量。基于上下文初始编码比特数、上下文目标编码比特数和第二潜在变量,通过第一循环方式确定第二变量调整因子。
其中,第一循环方式的第i次循环处理包括如下步骤:确定第i次循环处理的调整因子,i为正整数,基于第二潜在变量和第i次循环处理的调整因子,通过上下文模型确定第i次的上下文编码比特数和第i次的熵编码模型参数,基于第i次的熵编码模型参数,确定第二潜在变量的熵编码结果的编码比特数,得到第i次的基础编码比特数,基于第i次的上下文编码比特数与第i次的基础编码比特数确定第i次的编码比特数。在第i次的编码比特数满足继续调整条件的情况下,执行第一循环方式的第i+1次循环处理。在第i次的编码比特数不满足继续调整条件的情况下,终止第一循环方式的执行,基于第i次循环处理的调整因子确定第二变量调整因子。
其中,基于第二潜在变量和第i次循环处理的调整因子,通过上下文模型确定第i次的上下文编码比特数和第i次的熵编码模型参数的实现过程包括:通过上下文编码神经网络模型对第二潜在变量进行处理,得到第三潜在变量,第三潜在变量用于指示第二潜在变量的概率分布。基于第i次循环处理的调整因子对第三潜在变量进行调整,得到第i次调整后的第三潜在变量。确定第i次调整后的第三潜在变量的熵编码结果,将第i次调整后的第三潜在变量的熵编码结果的编码比特数作为第i次的上下文编码比特数。基于第i次调整后的第三潜在变量的熵编码结果重构第i次调整后的第三潜在变量。通过第i次循环处理的调整因子,对重构的第i次调整后的第三潜在变量进行调整,得到重构的第三潜在变量。通过上下文解码神经网络模型对重构的第三潜在变量进行处理,得到第i次的熵编码模型参数。
方式22,在上下文初始编码比特数等于上下文目标编码比特数的情况下,将第二初始调整因子确定为第二变量调整因子。在上下文初始编码比特数不等于上下文目标编码比特数的情况下,按照上述方式21来确定第二变量调整因子。但是,与上述方式21不同的是,在上下文初始编码比特数小于上下文目标编码比特数的情况下,按照第一步长调整第一循环方式的第i-1次循环处理的调整因子,得到第i次循环处理的调整因子,此时,继续调整条件包括第i次的编码比特数小于上下文目标编码比特数。在上下文初始编码比特数大于上下文目标编码比特数的情况下,按照第二步长调整第一循环方式的第i-1次循环处理的调整因子,得到第i次循环处理的调整因子,此时,继续调整条件包括第i次的编码比特数大于上下文目标编码比特数。
方式23,在上下文初始编码比特数小于或等于上下文目标编码比特数的情况下,将第二初始调整因子确定为第二变量调整因子。在上下文初始编码比特数大于上下文目标编码比特数的情况下,可以按照上述方式21来确定第二变量调整因子。也可以按照上述方式22中上 下文初始编码比特数大于上下文目标编码比特数的情况来确定第二变量调整因子。
第二种实现方式,将目标编码比特数划分为基础目标编码比特数和上下文目标编码比特数,基于基础目标编码比特数和基础初始编码比特数,确定第一变量调整因子。基于上下文目标编码比特数和上下文初始编码比特数,确定第二变量调整因子。
第二方面,提供了一种解码方法,在该方法中,也分为多种情况进行介绍。
第一种情况,基于码流确定重构的第二潜在变量和重构的第一变量调整因子;基于重构的第一变量调整因子,对重构的第二潜在变量进行调整,得到重构的第一潜在变量,重构的第一潜在变量用于指示待解码的媒体数据的特征;通过第一解码神经网络模型对重构的第一潜在变量进行处理,以得到重构的媒体数据。
第二种情况,基于码流确定重构的第三潜在变量和重构的第一变量调整因子;基于码流和重构的第三潜在变量,确定重构的第二潜在变量。基于重构的第一变量调整因子,对重构的第二潜在变量进行调整,得到重构的第一潜在变量,重构的第一潜在变量用于指示待解码的媒体数据的特征;通过第一解码神经网络模型对重构的第一潜在变量进行处理,以得到重构的媒体数据。
其中,基于码流和重构的第三潜在变量,确定重构的第二潜在变量,包括:通过上下文解码神经网络模型对重构的第三潜在变量进行处理,以得到重构的第一熵编码模型参数;基于码流和重构的第一熵编码模型参数,确定重构的第二潜在变量。
第三种情况,基于码流确定重构的第四潜在变量、重构的第二变量调整因子和重构的第一变量调整因子;基于码流、重构的第四潜在变量和重构的第二变量调整因子,确定重构的第二潜在变量。基于码流和重构的第三潜在变量,确定重构的第二潜在变量。基于重构的第一变量调整因子,对重构的第二潜在变量进行调整,得到重构的第一潜在变量,重构的第一潜在变量用于指示待解码的媒体数据的特征;通过第一解码神经网络模型对重构的第一潜在变量进行处理,以得到重构的媒体数据。
其中,基于码流、重构的第四潜在变量和重构的第二变量调整因子,确定重构的第二潜在变量,包括:基于重构的第二变量调整因子,对重构的第四潜在变量进行调整,得到重构的第三潜在变量;通过上下文解码神经网络模型对重构的第三潜在变量进行处理,以得到重构的第二熵编码模型参数;基于码流和重构的第二熵编码模型参数,确定重构的第二潜在变量。
其中,媒体数据为音频信号、视频信号或者图像。
第三方面,提供了一种编码装置,所述编码装置具有实现上述第一方面中编码方法行为的功能。所述编码装置包括至少一个模块,该至少一个模块用于实现上述第一方面所提供的编码方法。
第四方面,提供了一种解码装置,所述解码装置具有实现上述第二方面中解码方法行为的功能。所述解码装置包括至少一个模块,该至少一个模块用于实现上述第二方面所提供的解码方法。
第五方面,提供了一种编码端设备,所述编码端设备包括处理器和存储器,所述存储器用于存储执行上述第一方面所提供的编码方法的程序。所述处理器被配置为用于执行所述存储器中存储的程序,以实现上述第一方面提供的编码方法。
可选地,所述编码端设备还可以包括通信总线,该通信总线用于该处理器与存储器之间建立连接。
第六方面,提供了一种解码端设备,所述解码端设备包括处理器和存储器,所述存储器用于存储执行上述第二方面所提供的解码方法的程序。所述处理器被配置为用于执行所述存储器中存储的程序,以实现上述第二方面提供的解码方法。
可选地,所述解码端设备还可以包括通信总线,该通信总线用于该处理器与存储器之间建立连接。
第七方面,提供了一种计算机可读存储介质,所述存储介质内存储有指令,当所述指令在计算机上运行时,使得计算机执行上述第一方面所述的编码方法的步骤,或者执行上述第二方面所述的解码方法的步骤。
第八方面,提供了一种包含指令的计算机程序产品,当所述指令在计算机上运行时,使得计算机执行上述第一方面所述的编码方法的步骤,或者执行上述第二方面所述的解码方法的步骤。或者说,提供了一种计算机程序,所述计算机程序被执行时实现上述第一方面所述的编码方法的步骤,或者实现上述第二方面所述的解码方法的步骤。
第九方面,提供了一种计算机可读存储介质,所述计算机可读存储介质包括上述第一方面所述的编码方法所获得的码流。
上述第三方面、第四方面、第五方面、第六方面、第七方面、第八方面和第九方面所得到的技术效果与第一方面或第二方面中对应的技术手段得到的技术效果近似,在这里不再赘述。
本申请实施例提供的技术方案至少可以带来以下有益效果:
通过第一变量调整因子对第一潜在变量进行调整,得到第二潜在变量,而且第二潜在变量的熵编码结果的编码比特数满足预设编码速率条件,这样可以保证每帧媒体数据对应的潜在变量的熵编码结果的编码比特数均能够满足预设编码速率条件,也即是,可以保证每帧媒体数据对应的潜在变量的熵编码结果的编码比特数基本保持一致,而不是动态变化的,从而满足了编码器对稳定编码速率的需求。进一步地,在考虑到需要传输边信息(例如窗型,时域噪声整形(TNS:Temporal Noise Shaping)参数,频域噪声整形(FDNS:Frequency-domain noise shaping)参数,和/或带宽扩展(BWE:bandwidth extension)参数等等)时,能够保证每帧媒体数据对应的潜在变量的熵编码结果的编码比特数和边信息的编码比特数整体基本保持一致,从而满足编码器对稳定编码速率的需求。
附图说明
图1是本申请实施例提供的一种实施环境的示意图;
图2是本申请实施例提供的一种终端场景的实施环境的示意图;
图3是本申请实施例提供的一种无线或核心网设备的转码场景的实施环境的示意图;
图4是本申请实施例提供的一种广播电视场景的实施环境的示意图;
图5是本申请实施例提供的一种虚拟现实流场景的实施环境的示意图;
图6是本申请实施例提供的第一种编码方法的流程图;
图7是本申请实施例提供的第一种潜在变量的形式示意图;
图8是本申请实施例提供的第二种潜在变量的形式示意图;
图9是本申请实施例提供的第一种解码方法的流程图;
图10是本申请实施例提供的第二种编码方法的流程图;
图11是本申请实施例提供的第二种解码方法的流程图;
图12是本申请实施例提供的一种关于图10所示的编码方法的示例性框图;
图13是本申请实施例提供的一种关于图11所示的解码方法的示例性框图;
图14是本申请实施例提供的第三种编码方法的流程图;
图15是本申请实施例提供的第三种解码方法的流程图;
图16是本申请实施例提供的一种关于图14所示的编码方法的示例性框图;
图17是本申请实施例提供的一种关于图15所示的解码方法的示例性框图;
图18是本申请实施例提供的一种编码装置的结构示意图;
图19是本申请实施例提供的一种解码装置的结构示意图;
图20是本申请实施例提供的一种编解码装置的示意性框图。
具体实施方式
为使本申请实施例的目的、技术方案和优点更加清楚,下面将结合附图对本申请实施方式作进一步地详细描述。
在对本申请实施例提供的编解码方法进行详细地解释说明之前,先对本申请实施例涉及的术语和实施环境进行介绍。
为了便于理解,首先对本申请实施例涉及的术语进行解释。
编码:是指将待编码的媒体数据压缩成码流的处理过程。其中,待编码的媒体数据主要包括音频信号、视频信号和图像。音频信号的编码是将待编码的音频信号包括的音频帧序列压缩成码流的处理过程,视频信号的编码是将待编码的视频包括的图像序列压缩成码流的处理过程,图像的编码是将待编码的图像压缩成码流的处理过程。
需要说明的是,媒体数据被压缩成码流之后可以称为经编码的媒体数据或者经压缩的媒体数据。比如,对于音频信号来说,音频信号被压缩成码流之后可以称为经编码的音频信号或者经压缩的音频信号,视频信号被压缩成码流之后也可以称为经编码的视频信号或者经压缩的视频信号,图像被压缩成码流之后也可以称为经编码的图像或者经压缩的图像。
解码:是指将编码码流按照特定的语法规则和处理方法恢复成重建媒体数据的处理过程。其中,音频码流的解码是指将音频码流恢复成重建音频信号的处理过程,视频码流的解码是 指将视频码流恢复成重建视频信号的处理过程,图像码流的解码是指将图像码流恢复成重建图像的处理过程。
熵编码:是指编码过程中按熵原理不丢失任何信息的编码。也即是一种无损数据压缩方法。熵编码是基于元素的出现概率来编码的,即,对于同一元素,当该元素的出现概率不同时,该元素的熵编码结果的编码比特数不同。熵编码通常包括算术编码(arithmetic coding)、区间编码(range coding,RC)、哈夫曼(huffman)编码等等。
固定码率(constant bit rate,CBR):是指编码码率是一个固定值,比如该固定值为目标编码码率。
可变码率(variable bit rate,VBR):是指编码速率可超过目标编码码率,或者可小于目标编码码率,但是与目标编码码率之间的差值较小。
接下来对本申请实施例涉及的实施环境进行介绍。
请参考图1,图1是本申请实施例提供的一种实施环境的示意图。该实施环境包括源装置10、目的地装置20、链路30和存储装置40。其中,源装置10可以产生经编码的媒体数据。因此,源装置10也可以被称为媒体数据编码装置。目的地装置20可以对由源装置10所产生的经编码的媒体数据进行解码。因此,目的地装置20也可以被称为媒体数据解码装置。链路30可以接收源装置10所产生的经编码的媒体数据,并可以将该经编码的媒体数据传输给目的地装置20。存储装置40可以接收源装置10所产生的经编码的媒体数据,并可以将该经编码的媒体数据进行存储,这样的条件下,目的地装置20可以直接从存储装置40中获取经编码的媒体数据。或者,存储装置40可以对应于文件服务器或可以保存由源装置10产生的经编码的媒体数据的另一中间存储装置,这样的条件下,目的地装置20可以经由流式传输或下载存储装置40存储的经编码的媒体数据。
源装置10和目的地装置20均可以包括一个或多个处理器以及耦合到该一个或多个处理器的存储器,该存储器可以包括随机存取存储器(random access memory,RAM)、只读存储器(read-only memory,ROM)、带电可擦可编程只读存储器(electrically erasable programmable read-only memory,EEPROM)、快闪存储器、可用于以可由计算机存取的指令或数据结构的形式存储所要的程序代码的任何其它媒体等。例如,源装置10和目的地装置20均可以包括桌上型计算机、移动计算装置、笔记型(例如,膝上型)计算机、平板计算机、机顶盒、例如所谓的“智能”电话等电话手持机、电视机、相机、显示装置、数字媒体播放器、视频游戏控制台、车载计算机或其类似者。
链路30可以包括能够将经编码的媒体数据从源装置10传输到目的地装置20的一个或多个媒体或装置。在一种可能的实现方式中,链路30可以包括能够使源装置10实时地将经编码的媒体数据直接发送到目的地装置20的一个或多个通信媒体。在本申请实施例中,源装置10可以基于通信标准来调制经编码的媒体数据,该通信标准可以为无线通信协议等,并且可以将经调制的媒体数据发送给目的地装置20。该一个或多个通信媒体可以包括无线和/或有线通信媒体,例如该一个或多个通信媒体可以包括射频(radio frequency,RF)频谱或一个或多个物理传输线。该一个或多个通信媒体可以形成基于分组的网络的一部分,基于分组的网络可以为局域网、广域网或全球网络(例如,因特网)等。该一个或多个通信媒体可以包括路由器、交换器、基站或促进从源装置10到目的地装置20的通信的其它设备等,本申请实 施例对此不做具体限定。
在一种可能的实现方式中,存储装置40可以将接收到的由源装置10发送的经编码的媒体数据进行存储,目的地装置20可以直接从存储装置40中获取经编码的媒体数据。这样的条件下,存储装置40可以包括多种分布式或本地存取的数据存储媒体中的任一者,例如,该多种分布式或本地存取的数据存储媒体中的任一者可以为硬盘驱动器、蓝光光盘、数字多功能光盘(digital versatile disc,DVD)、只读光盘(compact disc read-only memory,CD-ROM)、快闪存储器、易失性或非易失性存储器,或用于存储经编码媒体数据的任何其它合适的数字存储媒体等。
在一种可能的实现方式中,存储装置40可以对应于文件服务器或可以保存由源装置10产生的经编码媒体数据的另一中间存储装置,目的地装置20可经由流式传输或下载存储装置40存储的媒体数据。文件服务器可以为能够存储经编码的媒体数据并且将经编码的媒体数据发送给目的地装置20的任意类型的服务器。在一种可能的实现方式中,文件服务器可以包括网络服务器、文件传输协议(file transfer protocol,FTP)服务器、网络附属存储(network attached storage,NAS)装置或本地磁盘驱动器等。目的地装置20可以通过任意标准数据连接(包括因特网连接)来获取经编码媒体数据。任意标准数据连接可以包括无线信道(例如,Wi-Fi连接)、有线连接(例如,数字用户线路(digital subscriber line,DSL)、电缆调制解调器等),或适合于获取存储在文件服务器上的经编码的媒体数据的两者的组合。经编码的媒体数据从存储装置40的传输可为流式传输、下载传输或两者的组合。
图1所示的实施环境仅为一种可能的实现方式,并且本申请实施例的技术不仅可以适用于图1所示的可以对媒体数据进行编码的源装置10,以及可以对经编码的媒体数据进行解码的目的地装置20,还可以适用于其他可以对媒体数据进行编码和对经编码的媒体数据进行解码的装置,本申请实施例对此不做具体限定。
在图1所示的实施环境中,源装置10包括数据源120、编码器100和输出接口140。在一些实施例中,输出接口140可以包括调节器/解调器(调制解调器)和/或发送器,其中发送器也可以称为发射器。数据源120可以包括图像捕获装置(例如,摄像机等)、含有先前捕获的媒体数据的存档、用于从媒体数据内容提供者接收媒体数据的馈入接口,和/或用于产生媒体数据的计算机图形系统,或媒体数据的这些来源的组合。
数据源120可以向编码器100发送媒体数据,编码器100可以对接收到由数据源120发送的媒体数据进行编码,得到经编码的媒体数据。编码器可以将经编码的媒体数据发送给输出接口。在一些实施例中,源装置10经由输出接口140将经编码的媒体数据直接发送到目的地装置20。在其它实施例中,经编码的媒体数据还可存储到存储装置40上,供目的地装置20以后获取并用于解码和/或显示。
在图1所示的实施环境中,目的地装置20包括输入接口240、解码器200和显示装置220。在一些实施例中,输入接口240包括接收器和/或调制解调器。输入接口240可经由链路30和/或从存储装置40接收经编码的媒体数据,然后再发送给解码器200,解码器200可以对接收到的经编码的媒体数据进行解码,得到经解码的媒体数据。解码器可以将经解码的媒体数据发送给显示装置220。显示装置220可与目的地装置20集成或可在目的地装置20外部。一般来说,显示装置220显示经解码的媒体数据。显示装置220可以为多种类型中的任一种类型的显示装置,例如,显示装置220可以为液晶显示器(liquid crystal display,LCD)、 等离子显示器、有机发光二极管(organic light-emitting diode,OLED)显示器或其它类型的显示装置。
尽管图1中未示出,但在一些方面,编码器100和解码器200可各自与编码器和解码器集成,且可以包括适当的多路复用器-多路分用器(multiplexer-demultiplexer,MUX-DEMUX)单元或其它硬件和软件,用于共同数据流或单独数据流中的音频和视频两者的编码。在一些实施例中,如果适用的话,那么MUX-DEMUX单元可符合ITU H.223多路复用器协议,或例如用户数据报协议(user datagram protocol,UDP)等其它协议。
编码器100和解码器200各自可为以下各项电路中的任一者:一个或多个微处理器、数字信号处理器(digital signal processing,DSP)、专用集成电路(application specific integrated circuit,ASIC)、现场可编程门阵列(field-programmable gate array,FPGA)、离散逻辑、硬件或其任何组合。如果部分地以软件来实施本申请实施例的技术,那么装置可将用于软件的指令存储在合适的非易失性计算机可读存储媒体中,且可使用一个或多个处理器在硬件中执行所述指令从而实施本申请实施例的技术。前述内容(包括硬件、软件、硬件与软件的组合等)中的任一者可被视为一个或多个处理器。编码器100和解码器200中的每一者都可以包括在一个或多个编码器或解码器中,所述编码器或所述解码器中的任一者可以集成为相应装置中的组合编码器/解码器(编码解码器)的一部分。
本申请实施例可大体上将编码器100称为将某些信息“发信号通知”或“发送”到例如解码器200的另一装置。术语“发信号通知”或“发送”可大体上指代用于对经压缩的媒体数据进行解码的语法元素和/或其它数据的传送。此传送可实时或几乎实时地发生。替代地,此通信可经过一段时间后发生,例如可在编码时在经编码位流中将语法元素存储到计算机可读存储媒体时发生,解码装置接着可在所述语法元素存储到此媒体之后的任何时间检索所述语法元素。
本申请实施例提供的编解码方法可以应用于多种场景,接下来以待编码的媒体数据为音频信号为例,对其中的几种场景分别进行介绍。
请参考图2,图2是本申请实施例提供的一种编解码方法应用于终端场景的实施环境的示意图。该实施环境包括第一终端101和第二终端201,第一终端101与第二终端201进行通信连接。该通信连接可以为无线连接,也可以为有线连接,本申请实施例对此不做限定。
其中,第一终端101可以为发送端设备,也可以为接收端设备,同理,第二终端201可以为接收端设备,也可以为发送端设备。在第一终端101为发送端设备的情况下,第二终端201为接收端设备,在第一终端101为接收端设备的情况下,第二终端201为发送端设备。
接下来以第一终端101为发送端设备,第二终端201为接收端设备为例进行介绍。
第一终端101可以为上述图1所示的实施环境中的源装置10。第二终端201可以为上述图1所示的实施环境中的目的地装置20。其中,第一终端101和第二终端201均包括音频采集模块、音频回放模块、编码器、解码器、信道编码模块和信道解码模块。
第一终端101中的音频采集模块采集音频信号并传输给编码器,编码器利用本申请实施例提供的编码方法对音频信号进行编码,该编码可以称为信源编码。之后,为了实现音频信号在信道中的传输,信道编码模块还需要再进行信道编码,然后将编码得到的码流通过无线或者有线网络通信设备在数字信道中传输。
第二终端201通过无线或者有线网络通信设备接收数字信道中传输的码流,信道解码模块对码流进行信道解码,然后解码器利用本申请实施例提供的解码方法解码得到音频信号,再通过音频回放模块进行播放。
其中,第一终端101和第二终端201可以是任何一种可与用户通过键盘、触摸板、触摸屏、遥控器、语音交互或手写设备等一种或多种方式进行人机交互的电子产品,例如个人计算机(personal computer,PC)、手机、智能手机、个人数字助手(Personal Digital Assistant,PDA)、可穿戴设备、掌上电脑PPC(pocket PC)、平板电脑、智能车机、智能电视、智能音箱等。
本领域技术人员应能理解上述终端仅为举例,其他现有的或今后可能出现的终端如可适用于本申请实施例,也应包含在本申请实施例保护范围以内,并在此以引用方式包含于此。
请参考图3,图3是本申请实施例提供的一种编解码方法应用于无线或核心网设备的转码场景的实施环境的示意图。该实施环境包括信道解码模块、音频解码器、音频编码器和信道编码模块。
其中,音频解码器可以为利用本申请实施例提供的解码方法的解码器,也可以为利用其他解码方法的解码器。音频编码器可以为利用本申请实施例提供的编码方法的编码器,也可以为利用其他编码方法的编码器。在音频解码器为利用本申请实施例提供的解码方法的解码器的情况下,音频编码器为利用其他编码方法的编码器,在音频解码器为利用其他解码方法的解码器的情况下,音频编码器为利用本申请实施例提供的编码方法的编码器。
第一种情况,音频解码器为利用本申请实施例提供的解码方法的解码器,音频编码器为利用其他编码方法的编码器。
此时,信道解码模块用于对接收的码流进行信道解码,然后音频解码器用于利用本申请实施例提供的解码方法进行信源解码,再通过音频编码器按照其他编码方法进行编码,实现一种格式到另一种格式的转换,即转码。之后,再通过信道编码后发送。
第二种情况,音频解码器为利用其他解码方法的解码器,音频编码器为利用本申请实施例提供的编码方法的编码器。
此时,信道解码模块用于对接收的码流进行信道解码,然后音频解码器用于利用其他解码方法进行信源解码,再通过音频编码器利用本申请实施例提供的编码方法进行编码,实现一种格式到另一种格式的转换,即转码。之后,再通过信道编码后发送。
其中,无线设备可以为无线接入点、无线路由器、无线连接器等等。核心网设备可以为移动性管理实体、网关等等。
本领域技术人员应能理解上述无线设备或者核心网设备仅为举例,其他现有的或今后可能出现的无线或核心网设备如可适用于本申请实施例,也应包含在本申请实施例保护范围以内,并在此以引用方式包含于此。
请参考图4,图4是本申请实施例提供的一种编解码方法应用于广播电视场景的实施环境的示意图。广播电视场景分为直播场景和后期制作场景。对于直播场景来说,该实施环境包括直播节目三维声制作模块、三维声编码模块、机顶盒和扬声器组,机顶盒包括三维声解码模块。对于后期制作场景来说,该实施环境包括后期节目三维声制作模块、三维声编码模 块、网络接收器、移动终端、耳机等。
直播场景下,直播节目三维声制作模块制作出三维声信号,该三维声信号经过应用本申请实施例的编码方法的编码得到码流,该码流经广电网络传输到用户侧,由机顶盒中的三维声解码器利用本申请实施例提供的解码方法进行解码,从而重建三维声信号,由扬声器组进行回放。或者,该码流经互联网传输到用户侧,由网络接收器中的三维声解码器利用本申请实施例提供的解码方法进行解码,从而重建三维声信号,由扬声器组进行回放。又或者,该码流经互联网传输到用户侧,由移动终端中的三维声解码器利用本申请实施例提供的解码方法进行解码,从而重建三维声信号,由耳机进行回放。
后期制作场景下,后期节目三维声制作模块制作出三维声信号,该三维声信号经过应用本申请实施例的编码方法的编码得到码流,该码流经广电网络传输到用户侧,由机顶盒中的三维声解码器利用本申请实施例提供的解码方法进行解码,从而重建三维声信号,由扬声器组进行回放。或者,该码流经互联网传输到用户侧,由网络接收器中的三维声解码器利用本申请实施例提供的解码方法进行解码,从而重建三维声信号,由扬声器组进行回放。又或者,该码流经互联网传输到用户侧,由移动终端中的三维声解码器利用本申请实施例提供的解码方法进行解码,从而重建三维声信号,由耳机进行回放。
请参考图5,图5是本申请实施例提供的一种编解码方法应用于虚拟现实流场景的实施环境的示意图。该实施环境包括编码端和解码端,编码端包括采集模块、预处理模块、编码模块、打包模块和发送模块,解码端包括解包模块、解码模块、渲染模块和耳机。
采集模块采集音频信号,然后通过预处理模块进行预处理操作,预处理操作包括滤除掉信号中的低频部分,通常是以20Hz或者50Hz为分界点,提取信号中的方位信息等。之后通过编码模块,利用本申请实施例提供的编码方法进行编码处理,编码之后通过打包模块进行打包,进而通过发送模块发送给解码端。
解码端的解包模块首先进行解包,之后通过解码模块,利用本申请实施例提供的解码方法进行解码,然后通过渲染模块对解码信号进行双耳渲染处理,渲染处理后的信号映射到收听者耳机上。该耳机可以为独立的耳机,也可以是基于虚拟现实的眼镜设备上的耳机。
接下来对本申请实施例提供的编解码方法进行详细地解释说明。需要说明的是,结合图1所示的实施环境,下文中的任一种编码方法可以是源装置10中的编码器100执行的。下文中的任一种解码方法可以是目的地装置20中的解码器200执行的。
需要说明的是,本申请实施例可以应用于不包括上下文模型的编解码器中,也可以应用于包括上下文模型的编解码器中。而且,本申请实施例不仅可以对待编码的媒体数据生成的潜在变量通过调整因子进行调整,还可以对上下文模型确定的潜在变量通过调整因子进行调整。因此,接下来将分为多个实施例,对本申请实施例提供的编解码方法进行详细地解释说明。另外,本申请实施例中涉及的调整因子和变量调整因子可以是量化前的值,也可以是量化后的值,本申请实施例对此不做限定。
请参考图6,图6是本申请实施例提供的第一种编码方法的流程图。该方法不包括上下文模型,只对待编码的媒体数据生成的潜在变量通过调整因子进行调整。该编码方法应用于编码端设备,包括如下步骤。
步骤601:通过第一编码神经网络模型对待编码的媒体数据进行处理,以得到第一潜在变量,第一潜在变量用于指示待编码的媒体数据的特征。
其中,待编码的媒体数据为音频信号、视频信号或者图像等。而且,待编码的媒体数据的形式可以为任何一种形式,本申请实施例对此不做限定。
示例地,待编码的媒体数据可以为时域的媒体数据,也可以为时域的媒体数据经过时频变换后得到的频域的媒体数据,例如,可以是时域的媒体数据经过MDCT变换后得到的频域的媒体数据,或者是时域的媒体数据经过快速傅氏变换(fast fourier transformation,FFT)后得到的频域的媒体数据。待编码的媒体数据还可以为时域的媒体数据经过正交镜象滤波器(quandrature mirror filter,QMF)滤波后得到的复频域的媒体数据,或者待编码的媒体数据为从时域的媒体数据提取得到的特征信号,例如梅尔倒谱系数,或者待编码的媒体数据还可以是残差信号,例如其他编码的残差信号或者线性预测编码(linear predictive coding,LPC)滤波后的残差信号。
通过第一编码神经网络模型对待编码的媒体数据进行处理的实现过程为:将待编码的媒体数据输入第一编码神经网络模型,得到第一编码神经网络模型输出的第一潜在变量。或者,对待编码的媒体数据进行预处理,将预处理后的媒体数据输入第一编码神经网络模型,得到第一编码神经网络模型输出的第一潜在变量。
也就是说,可以将待编码的媒体数据作为第一编码神经网络模型的输入来确定第一潜在变量,也可以对待编码的媒体数据进行预处理后,再作为第一编码神经网络模型的输入来确定第一潜在变量。
其中,该预处理操作可以为时域噪声整形(temporal noise shaping,TNS)处理、频域噪声整形(frequency domain noise shaping,FDNS)处理、声道下混处理等等。
第一编码神经网络模型是预先训练好的,本申请实施例对第一编码神经网络模型的网络结构和训练方法不做限定。例如,第一编码神经网络模型的网络结构可以为全连接网络或者卷积神经网络(convolutional neural network,CNN)网络。另外,本申请实施例对第一编码神经网络模型的网络结构所包含的层数和每一层的节点数也不做限定。
不同网络结构的编码神经网络模型输出的潜在变量的形式可能不同。例如,在第一编码神经网络模型的网络结构是全连接网络的情况下,第一潜在变量为一个矢量,矢量的维数M是潜在变量的大小(latent size),如图7所示。在第一编码神经网络模型的网络结构是CNN网络的情况下,第一潜在变量为一个N*M维矩阵,其中N为CNN网络的通道数(channel),M为CNN网络的每个通道潜在变量的大小(latent size),如图8所示。需要注意的是,图7和图8仅给出了全连接网络的潜在变量和CNN网络的潜在变量的一种示意,通道序号可以从1开始计数也可以从0开始计数,各个通道内潜在变量的元素序号也一样。
步骤602:基于第一潜在变量确定第一变量调整因子,第一变量调整因子用于使得第二潜在变量的熵编码结果的编码比特数满足预设编码速率条件,第二潜在变量是通过第一变量调整因子对第一潜在变量调整后得到。
在一些实施例中,可以基于第一潜在变量确定初始编码比特数,基于初始编码比特数和目标编码比特数,确定第一变量调整因子。
其中,目标编码比特数可以为事先设置的。当然,目标编码比特数也可以基于编码速率来确定,且不同的编码速率对应不同的目标编码比特数。
在本申请实施例中,可以使用固定码率对待编码的媒体数据进行编码,也可以使用可变码率对待编码的媒体数据进行编码。
在使用固定码率对待编码的媒体数据进行编码的情况下,可以基于固定码率确定当前帧的待编码的媒体数据的比特数,再减去当前帧的已使用比特数,得到当前帧的目标编码比特数。其中,已使用比特数可以为边信息等进行编码的比特数,而且通常情况下,每帧媒体数据的边信息不同,所以,每帧媒体数据的目标编码比特数通常是不同的。
在使用可变码率对待编码的媒体数据进行编码的情况下,通常会指定一个码率,实际码率会在指定的码率的上下波动。此时,可以基于指定的码率确定当前帧的待编码的媒体数据的比特数,再减去当前帧的已使用比特数,得到当前帧的目标编码比特数。其中,已使用比特数可以为边信息等进行编码的比特数,而且在某些情况下,不同帧的媒体数据的边信息可以是不同的,所以,不同帧的媒体数据的目标编码比特数通常是不同的。
接下来将分别对初始编码比特数和第一变量调整因子的确定过程进行详细解释说明。
确定初始编码比特数
其中,基于第一潜在变量确定初始编码比特数的方式可以包括两种,接下来将分别介绍。
第一种实现方式,确定第一潜在变量的熵编码结果的编码比特数,以得到初始编码比特数。也即是,初始编码比特数为第一潜在变量的熵编码结果的编码比特数。
作为一种示例,对第一潜在变量进行量化处理,得到量化后的第一潜在变量。对量化后的第一潜在变量进行熵编码,得到第一潜在变量的初始编码结果。统计第一潜在变量的初始编码结果的编码比特数,得到初始编码比特数。
其中,对第一潜在变量进行量化处理的方式可以包括多种,比如,对第一潜在变量中的每个元素进行标量量化。标量量化的量化步长可以基于不同的编码速率来确定,也即是,事先存储编码速率与量化步长的对应关系,可以基于本申请实施例采用的编码速率从该对应关系中获取对应的量化步长。另外,标量量化还可以存在偏置量,即,通过偏置量对第一潜在变量进行偏置处理后再按照量化步长进行标量量化。
对量化后的第一潜在变量进行熵编码时,可以采用基于可调熵编码模型进行熵编码,也可以采用预置概率分布的熵编码模型进行熵编码,本申请实施例对此不做限定。其中,熵编码可以采用算术编码(arithmetic coding)、区间编码(range coding,RC)或者哈夫曼(huffman)编码中的一种,本申请实施例不做限定。
需要说明的是,下文的量化处理方式以及熵编码方式与此处的类似,下文的量化处理方式和熵编码方式可以参考此处的方式,本申请实施例在后文不再赘述。
第二种实现方式,基于第一初始调整因子对第一潜在变量进行调整,确定调整后的第一潜在变量的熵编码结果的编码比特数,以得到初始编码比特数。也即是,初始编码比特数为经过第一初始调整因子调整后的第一潜在变量的熵编码结果的编码比特数。其中,第一初始调整因子可以为第一预设调整因子。
作为一种示例,基于第一初始调整因子对第一潜在变量进行调整,得到调整后的第一潜在变量,对调整后的第一潜在变量进行量化处理,得到量化后的第一潜在变量。对量化后的第一潜在变量进行熵编码,得到第一潜在变量的初始编码结果。统计第一潜在变量的初始编码结果的编码比特数,得到初始编码比特数。
其中,基于第一初始调整因子对第一潜在变量进行调整的实现过程为:将第一潜在变量 中的各个元素与第一初始调整因子中对应的元素相乘,得到调整后的第一潜在变量。
值得注意的是,上述实现过程仅仅为一种示例,实际应用中,还可以采用其他的方法来调整。比如,可以将第一潜在变量中的各个元素除以第一初始调整因子中对应的元素,得到调整后的第一潜在变量。本申请实施例对调整方法不做限定。
需要说明的是,本申请实施例针对第一潜在变量设置有调整因子初值,调整因子初值通常等于1。第一预设调整因子可以大于或等于调整因子初值,也可以小于调整因子初值,比如,第一预设调整因子为1或2等常数。在通过上述第一种实现方式来确定初始编码比特数的情况下,第一初始调整因子为调整因子初值,在通过上述第二种实现方式来确定初始编码比特数的情况下,第一初始调整因子为第一预设调整因子。
另外,第一初始调整因子可以是标量也可以是矢量。例如,假设第一编码神经网络模型的网络结构为全连接网络,其输出的第一潜在变量为一个矢量,矢量的维数M是第一潜在变量的大小(latent size)。如果第一初始调整因子是标量,这种情况下,维数M且为矢量的第一潜在变量中的每个元素对应的调整因子值是相同的,即第一初始调整因子包括一个元素。如果第一初始调整因子是矢量,这种情况下,维数M且为矢量的第一潜在变量矢量中的每个元素对应的调整因子值是不完全相同的,可以多个元素共用一个调整因子值,即第一初始调整因子包括多个元素,每个元素对应第一潜在变量中的一个或多个元素。
同样的,假设第一编码神经网络模型的网络结构为CNN网络,其输出的第一潜在变量为一个N*M维矩阵,其中N为CNN网络的通道数(channel),M为CNN网络的每个通道潜在变量的大小(latent size)。如果第一初始调整因子是标量,这种情况下,N*M维的第一潜在变量矩阵中的每个元素对应的调整因子值是相同的,即第一初始调整因子包括一个元素。如果第一初始调整因子是矢量,这种情况下,N*M维的第一潜在变量矩阵中的每个元素对应的调整因子值是不完全相同的,可以属于同一通道的潜在变量的元素对应相同的调整因子值,即第一初始调整因子包括N个元素,每个元素对应第一潜在变量中通道序号相同的M个元素。
确定第一变量调整因子
上述第一变量调整因子为第一潜在变量的调整因子最终值。通过第一变量调整因子对第一潜在变量进行调整后得到第二潜在变量,且第二潜在变量的熵编码结果的编码比特数满足预设编码速率条件。其中,在使用固定码率对待编码的媒体数据进行编码的情况下,满足预设编码速率条件包括编码比特数小于或等于目标编码比特数。或者,满足预设编码速率条件包括编码比特数小于或等于目标编码比特数,且编码比特数与目标编码比特数的差值小于比特数阈值。在使用可变码率对待编码的媒体数据进行编码的情况下,满足预设编码速率条件包括编码比特数与目标编码比特数的差值的绝对值小于比特数阈值。也即是,满足预设编码速率条件包括编码比特数小于或等于目标编码比特数,且目标编码比特数与编码比特数的差值小于比特数阈值。或者,满足预设编码速率条件包括编码比特数大于或等于目标编码比特数,且编码比特数与目标编码比特数的差值小于比特数阈值。
需要说明的是,比特数阈值可以事先设置,且比特数阈值可以基于不同的需求进行调整。
其中,基于初始编码比特数和目标编码比特数,确定第一变量调整因子的实现方式可以包括多种,接下来对其中的三种进行介绍。
第一种实现方式,在初始编码比特数等于目标编码比特数的情况下,将第一初始调整因 子确定为第一变量调整因子。在初始编码比特数不等于目标编码比特数的情况下,基于初始编码比特数和目标编码比特数,通过第一循环方式确定第一变量调整因子。
其中,第一循环方式的第i次循环处理包括如下步骤:确定第i次循环处理的调整因子,i为正整数,基于第i次循环处理的调整因子对第一潜在变量进行调整,以得到第i次调整后的第一潜在变量。确定第i次调整后的第一潜在变量的熵编码结果的编码比特数,以得到第i次的编码比特数。在第i次的编码比特数满足继续调整条件的情况下,执行第一循环方式的第i+1次循环处理。在第i次的编码比特数不满足继续调整条件的情况下,终止第一循环方式的执行,基于第i次循环处理的调整因子确定第一变量调整因子。
其中,确定第i次循环处理的调整因子的实现过程为:基于第一循环方式的第i-1次循环处理的调整因子、第i-1次的编码比特数和目标编码比特数,确定第i次循环处理的调整因子。其中,在i=1的情况下,第i-1次循环处理的调整因子为第一初始调整因子,第i-1次的编码比特数为初始编码比特数。
此时,继续调整条件包括第i-1次的编码比特数和第i次的编码比特数均小于目标编码比特数,或者,继续调整条件包括第i-1次的编码比特数和第i次的编码比特数均大于目标编码比特数。
换句话说,继续调整条件包括第i次的编码比特数没有越过目标编码比特数。这里的没有越过的意思是指:前i-1次的编码比特数一直小于目标编码比特数,第i次的编码比特数仍小于目标编码比特数。或者,前i-1次的编码比特数一直大于目标编码比特数,第i次的编码比特数仍大于目标编码比特数。相反地,越过的意思是指:前i-1次的编码比特数一直小于目标编码比特数,第i次的编码比特数大于目标编码比特数。或者,前i-1次的编码比特数一直大于目标编码比特数,第i次的编码比特数小于目标编码比特数。
作为一种示例,在调整因子为量化后的值的情况下,可以基于第一循环方式的第i-1次循环处理的调整因子、第i-1次的编码比特数和目标编码比特数,按照下述公式(1)确定第i次循环处理的调整因子。
scale(i)=Q{scale(i-1)*[target/curr(i-1)]}       (1)
其中,在上述公式(1)中,scale(i)是指第i次循环处理的调整因子,scale(i-1)是指第i-1次循环处理的调整因子,target是指目标编码比特数,curr(i-1)是指第i-1次的编码比特数。其中,i为大于0的正整数。Q{x}是指获取x量化后的值。
需要说明的是,在调整因子为量化前的值的情况下,通过上述公式(1)确定第i次循环处理的调整因子时,上述公式(1)的等式右边可以不用通过Q{x}来处理。
其中,基于第i次循环处理的调整因子确定第一变量调整因子的实现过程包括:在第i次的编码比特数等于目标编码比特数的情况下,将第i次循环处理的调整因子确定为第一变量调整因子。或者,在第i次的编码比特数不等于目标编码比特数的情况下,基于第i次循环处理的调整因子和第i-1次循环处理的调整因子,确定第一变量调整因子。
也即是,第i次循环处理的调整因子为通过上述第一循环方式最后一次得到的调整因子,第i次的编码比特数为最后一次得到的编码比特数。在最后一次得到的编码比特数等于目标编码比特数的情况下,将最后一次得到的调整因子确定为第一变量调整因子。在最后一次得到的编码比特数不等于目标编码比特数的情况下,基于后两次得到的调整因子确定第一变量 调整因子。
在一些实施例中,基于第i次循环处理的调整因子和第i-1次循环处理的调整因子,确定第一变量调整因子的实现过程包括:确定第i次循环处理的调整因子和第i-1次循环处理的调整因子的平均值,基于该平均值确定第一变量调整因子。
作为一种示例,可以直接将该平均值确定为第一变量调整因子,也可以将该平均值乘以一个预设的常数,得到第一变量调整因子。可选地,该常数可以小于1。
在另一些实施例中,基于第i次循环处理的调整因子和第i-1次循环处理的调整因子,确定第一变量调整因子的实现过程包括:基于第i次循环处理的调整因子和第i-1次循环处理的调整因子,通过第二循环方式确定第一变量调整因子。
作为一种示例,第二循环方式的第j次循环处理包括如下步骤:基于第j次循环处理的第一调整因子和第j次循环处理的第二调整因子确定第j次循环处理的第三调整因子,其中,在j等于1的情况下,第j次循环处理的第一调整因子为第i次循环处理的调整因子和第i-1次循环处理的调整因子中的一者,第j次循环处理的第二调整因子为第i次循环处理的调整因子和第i-1次循环处理的调整因子中的另一者,第j次循环处理的第一调整因子对应第j次的第一编码比特数,第j次循环处理的第二调整因子对应第j次的第二编码比特数,第j次的第一编码比特数是指经过第j次循环处理的第一调整因子调整后的第一潜在变量的熵编码结果的编码比特数,第j次的第二编码比特数是指经过第j次循环处理的第二调整因子调整后的第一潜在变量的熵编码结果的编码比特数,第j次的第一编码比特数小于第j次的第二编码比特数。获取第j次的第三编码比特数,第j次的第三编码比特数是指经过第j次循环处理的第三调整因子调整后的第一潜在变量的熵编码结果的编码比特数。若第j次的第三编码比特数不满足继续循环条件,终止第二循环方式的执行,将第j次循环处理的第三调整因子确定为第一变量调整因子。若第j次的第三编码比特数满足继续循环条件、第j次的第三编码比特数大于目标编码比特数且小于第j次的第二编码比特数,将第j次循环处理的第三调整因子作为第j+1次循环处理的第二调整因子,将第j次循环处理的第一调整因子作为第j+1次循环处理的第一调整因子,执行第二循环方式的第j+1次循环处理。若第j次的第三编码比特数满足继续循环条件、第j次的第三编码比特数小于目标编码比特数且大于第j次的第一编码比特数,将第j次循环处理的第三调整因子作为第j+1次循环处理的第一调整因子,将第j次循环处理的第二调整因子作为第j+1次循环处理的第二调整因子,执行第二循环方式的第j+1次循环处理。
作为另一种示例,第二循环方式的第j次循环处理包括如下步骤:基于第j次循环处理的第一调整因子和第j次循环处理的第二调整因子确定第j次循环处理的第三调整因子,其中,在j等于1的情况下,第j次循环处理的第一调整因子为第i次循环处理的调整因子和第i-1次循环处理的调整因子中的一者,第j次循环处理的第二调整因子为第i次循环处理的调整因子和第i-1次循环处理的调整因子中的另一者,第j次循环处理的第一调整因子对应第j次的第一编码比特数,第j次循环处理的第二调整因子对应第j次的第二编码比特数,第j次的第一编码比特数小于第j次的第二编码比特数,j为正整数。获取第j次的第三编码比特数,第j次的第三编码比特数是指经过第j次循环处理的第三调整因子调整后的第一潜在变量的熵编码结果的编码比特数。若第j次的第三编码比特数不满足继续循环条件,终止第二循环方式的执行,将第j次循环处理的第三调整因子确定为第一变量调整因子。若j达到最大循环 次数且第j次的第三编码比特数满足继续循环条件,终止第二循环方式的执行,基于第j次循环处理的第一调整因子确定第一变量调整因子。若j未达到最大循环次数、第j次的第三编码比特数满足继续循环条件、第j次的第三编码比特数大于目标编码比特数且小于第j次的第二编码比特数,将第j次循环处理的第三调整因子作为第j+1次循环处理的第二调整因子,将第j次循环处理的第一调整因子作为第j+1次循环处理的第一调整因子,执行第二循环方式的第j+1次循环处理。若j未达到最大循环次数、第j次的第三编码比特数满足继续循环条件、第j次的第三编码比特数小于目标编码比特数且大于第j次的第一编码比特数,将第j次循环处理的第三调整因子作为第j+1次循环处理的第一调整因子,将第j次循环处理的第二调整因子作为第j+1次循环处理的第二调整因子,执行第二循环方式的第j+1次循环处理。
其中,基于第j次循环处理的第一调整因子和第j次循环处理的第二调整因子确定第j次循环处理的第三调整因子的实现过程包括:确定第j次循环处理的第一调整因子和第j次循环处理的第二调整因子的平均值,基于该平均值确定第j次循环处理的第三调整因子。作为一种示例,可以直接将该平均值确定为第j次循环处理的第三调整因子,也可以将该平均值乘以一个预设的常数,得到第j次循环处理的第三调整因子。可选地,该常数可以小于1。
另外,获取第j次的第三编码比特数的实现过程包括:基于第j次循环处理的第三调整因子对第一潜在变量进行调整,得到调整后的第一潜在变量,对调整后的第一潜在变量进行量化处理,得到量化后的第一潜在变量。对量化后的第一潜在变量进行熵编码,统计该熵编码结果的编码比特数,得到第j次的第三编码比特数。
其中,在使用固定码率对待编码的媒体数据进行编码的情况下,基于第j次循环处理的第一调整因子确定第一变量调整因子的实现过程包括:将第j次循环处理的第一调整因子确定为第一变量调整因子。在使用可变码率对待编码的媒体数据进行编码的情况下,基于第j次循环处理的第一调整因子确定第一变量调整因子的实现过程包括:确定目标编码比特数与第j次的第一编码比特数之间的第一差值,以及确定第j次的第二编码比特数与目标编码比特数之间的第二差值。若第一差值小于第二差值,将第j次循环处理的第一调整因子确定为第一变量调整因子。若第二差值小于第一差值,将第j次循环处理的第二调整因子确定为第一变量调整因子。若第一差值等于第二差值,将第j次循环处理的第一调整因子确定为第一变量调整因子,或者将第j次循环处理的第二调整因子确定为第一变量调整因子。
其中,在使用固定码率对待编码的媒体数据进行编码的情况下,继续循环条件包括第j次的第三编码比特数大于目标编码比特数,或者,继续循环条件包括第j次的第三编码比特数小于目标编码比特数,且目标编码比特数与第j次的第三编码比特数的差值大于比特数阈值。在使用可变码率对待编码的媒体数据进行编码的情况下,继续循环条件包括目标编码比特数与第j次的第三编码比特数的差值的绝对值大于比特数阈值。也即是,继续循环条件包括第j次的第三编码比特数大于目标编码比特数,且第j次的第三编码比特数与目标编码比特数的差值大于比特数阈值,或者,继续循环条件包括第j次的第三编码比特数小于目标编码比特数,且目标编码比特数与第j次的第三编码比特数的差值大于比特数阈值。
即,在使用固定码率对待编码的媒体数据进行编码的情况下,继续循环条件包括:bits_curr>target||(bits_curr<target&&(target-bits_curr)>TH),在使用可变码率对待编码的媒体数据进行编码的情况下,继续循环条件包括:bits_curr>target&&(bits_curr-target)>TH||(bits_curr<target&&(target-bits_curr)>TH)。其中,bits_curr是指第j次的第三编码比特数, target是指目标编码比特数,TH为比特数阈值。
其中,可以将第j次循环处理的第一调整因子记为scale_lower,将第j次的第一编码比特数记为bits_lower,将第j次循环处理的第二调整因子记为scale_upper,将第j次的第二编码比特数记为bits_upper,将第j次循环处理的第三调整因子记为scale_curr,将第j次循环处理的第三调整因子作为第j+1次循环处理的第一调整因子或者第j+1次循环处理的第二调整因子的伪代码如下:
Figure PCTCN2022092385-appb-000001
第二种实现方式,在初始编码比特数等于目标编码比特数的情况下,将第一初始调整因子确定为第一变量调整因子。在初始编码比特数不等于目标编码比特数的情况下,按照上述第一种实现方式来确定第一变量调整因子。但是,与上述第一种实现方式不同的是,在初始编码比特数小于目标编码比特数的情况下,按照第一步长调整第i-1次循环处理的调整因子,得到第i次循环处理的调整因子,此时,继续调整条件包括第i次的编码比特数小于目标编码比特数。在初始编码比特数大于目标编码比特数的情况下,按照第二步长调整第i-1次循环处理的调整因子,得到第i次循环处理的调整因子,此时,继续调整条件包括第i次的编码比特数大于目标编码比特数。
其中,按照第一步长调整第i-1次循环处理的调整因子可以是指按照第一步长增大第i-1次循环处理的调整因子,按照第二步长调整第i-1次循环处理的调整因子可以是指按照第二步长减小第i-1次循环处理的调整因子。
上述的增大处理和减小处理可以为线性的,也可以为非线性的。示例地,可以将第i-1次循环处理的调整因子与第一步长之和确定为第i次循环处理的调整因子,可以将第i-1次循环处理的调整因子与第二步长之差确定为第i次循环处理的调整因子。
需要说明的是,第一步长和第二步长可以为事先设置的,且第一步长和第二步长可以基于不同的需求来调整。另外,第一步长与第二步长可以相等,也可以不相等。
第三种实现方式,在初始编码比特数小于或等于目标编码比特数的情况下,将第一初始调整因子确定为第一变量调整因子。在初始编码比特数大于目标编码比特数的情况下,可以按照上述第一种实现方式来确定第一变量调整因子。也可以按照上述第二种实现方式中初始编码比特数大于目标编码比特数的情况来确定第一变量调整因子。
步骤603:获取第二潜在变量的熵编码结果。
在步骤602确定第一变量调整因子的过程中,第一变量调整因子可能为第一初始调整因子,可能为第i次循环处理的调整因子,也可能是第j次循环处理的第三调整因子,也可能是基于第j次循环处理的第一调整因子确定得到的。不管是哪种情况,上述循环过程中都有确定对应的熵编码结果,所以,可以直接从上述循环过程中确定得到的熵编码结果中获取第一变量调整因子对应的熵编码结果,即第二潜在变量的熵编码结果。
当然,还可以重新处理一次。即,直接基于第一变量调整因子对第一潜在变量进行调整, 得到第二潜在变量。对第二潜在变量进行量化处理,得到量化后的第二潜在变量。对量化后的第二潜在变量进行熵编码,得到第二潜在变量的熵编码结果。
基于上述描述,第一初始调整因子可以为标量也可以是矢量,第一变量调整因子是基于第一初始调整因子确定的,所以第一变量调整因子可以为标量也可以是矢量。
即,第一变量调整因子包括一个元素或者多个元素,在第一变量调整因子包括多个元素的情况下,第一变量调整因子中的一个元素对应第一潜在变量中的一个或多个元素。
步骤604:将第二潜在变量的熵编码结果以及第一变量调整因子的编码结果写入码流。
基于上述描述,第一变量调整因子可以为量化前的值,也可以为量化后的值。若第一变量调整因子为量化前的值,此时,确定第一变量调整因子的编码结果的实现过程包括:对第一变量调整因子进行量化及编码处理,得到第一变量调整因子的编码结果。若第一变量调整因子为量化后的值,此时,确定第一变量调整因子的编码结果的实现过程包括:对第一变量调整因子进行编码,得到第一变量调整因子的编码结果。
其中,第一变量调整因子可以采用任何一种编码方式进行编码,本申请实施例对此不做限定。
另外,对于某些编码器来说,量化过程和编码过程是在一次处理中执行的,也即是,通过一次处理能够得到量化结果和编码结果。所以,对于第一变量调整因子来说,在上述确定第一变量调整因子的过程中可能也会得到第一变量调整因子的编码结果,所以,可以直接获取到第一变量调整因子的编码结果。也即是,第一变量调整因子的编码结果也可以在确定第一变量调整因子的过程中直接得到。
可选地,若第一变量调整因子为量化前的值,本申请实施例还可以确定第一变量调整因子的量化步长对应的量化索引,得到第一量化索引。可选地,本申请实施例还可以确定第二潜在变量的量化步长对应的量化索引,得到第二量化索引。将第一量化索引和第二量化索引编入码流。其中,量化索引用于指示对应的量化步长,即,第一量化索引用于指示第一变量调整因子的量化步长,第二量化索引用于指示第二潜在变量的量化步长。
在本申请实施例中,通过第一变量调整因子对第一潜在变量进行调整,得到第二潜在变量,而且第二潜在变量的熵编码结果的编码比特数满足预设编码速率条件,这样可以保证每帧媒体数据对应的潜在变量的熵编码结果的编码比特数均能够满足预设编码速率条件,也即是,可以保证每帧媒体数据对应的潜在变量的熵编码结果的编码比特数基本保持一致,而不是动态变化的,从而满足了编码器对稳定编码速率的需求。进一步地,在考虑到需要传输边信息(例如窗型,时域噪声整形(TNS:Temporal Noise Shaping)参数,频域噪声整形(FDNS:Frequency-domain noise shaping)参数,和/或带宽扩展(BWE:bandwidth extension)参数等等)时,能够保证每帧媒体数据对应的潜在变量的熵编码结果的编码比特数和边信息的编码比特数整体基本保持一致,从而满足编码器对稳定编码速率的需求。
请参考图9,图9是本申请实施例提供的第一种解码方法的流程图,该方法应用于解码端。该方法对应于图6所示的编码方法。该方法包括如下步骤。
步骤901:基于码流确定重构的第二潜在变量和重构的第一变量调整因子。
在一些实施例中,可以对码流中的第二潜在变量的熵编码结果进行熵解码,以及对码流中的第一变量调整因子的编码结果进行解码,得到经量化的第二潜在变量和经量化的第一变 量调整因子。对经量化的第二潜在变量和经量化的第一变量调整因子进行去量化处理,得到重构的第二潜在变量和重构的第一变量调整因子。
其中,本步骤的解码方法与编码端的编码方法相对应,本步骤的去量化处理与编码端的量化处理相对应。即,解码方法为编码方法的逆过程,去量化处理为量化处理的逆过程。
示例地,可以从码流中解析出第一量化索引和第二量化索引,第一量化索引用于指示第一变量调整因子的量化步长,第二量化索引用于指示第二潜在变量的量化步长。按照第一量化索引所指示的量化步长,对经量化的第一变量调整因子进行去量化处理,得到重构的第一变量调整因子。按照第二量化索引所指示的量化步长,对经量化的第二潜在变量进行去量化处理,得到重构的第二潜在变量。
步骤902:基于重构的第一变量调整因子,对重构的第二潜在变量进行调整,得到重构的第一潜在变量,重构的第一潜在变量用于指示待解码的媒体数据的特征。
由于第二潜在变量是第一潜在变量经过第一变量调整因子调整后得到的,所以,可以基于重构的第一变量调整因子,对重构的第二潜在变量进行调整,得到重构的第一潜在变量。
其中,基于重构的第一变量调整因子对重构的第二潜在变量进行调整的过程为编码侧对第一潜在变量进行调整的逆过程。示例地,在编码侧将第一潜在变量中的各个元素与第一变量调整因子中对应的元素相乘的情况下,此处可以将重构的第二潜在变量中各个元素除以重构的第一变量调整因子中对应的元素,得到重构的第一潜在变量。在编码侧将第一潜在变量中的各个元素除以第一变量调整因子中对应的元素的情况下,此处可以将重构的第二潜在变量中各个元素乘以重构的第一变量调整因子中对应的元素,得到重构的第一潜在变量。
基于上述描述,第一初始调整因子可以是标量也可以是矢量,所以最终得到的第一变量调整因子可以是标量也可以是矢量。例如,假设第一编码神经网络模型的网络结构为全连接网络,第二潜在变量为一个矢量,矢量的维数M是第二潜在变量的大小(latent size)。如果第一变量调整因子是标量,这种情况下,维数M且为矢量的第二潜在变量中的每个元素对应的调整因子值是相同的,即第一变量调整因子包括一个元素。如果第一变量调整因子是矢量,这种情况下,维数M且为矢量的第二潜在变量矢量中的每个元素对应的调整因子值是不完全相同的,可以多个元素共用一个调整因子值,即第一变量调整因子包括多个元素,每个元素对应第二潜在变量中的一个或多个元素。
同样的,假设第一编码神经网络模型的网络结构为CNN网络,第二潜在变量为一个N*M维矩阵,其中N为CNN网络的通道数(channel),M为CNN网络的每个通道潜在变量的大小(latent size)。如果第一变量调整因子是标量,这种情况下,N*M维的第二潜在变量矩阵中的每个元素对应的调整因子值是相同的,即第一变量调整因子包括一个元素。如果第一变量调整因子是矢量,这种情况下,N*M维的第二潜在变量矩阵中的每个元素对应的调整因子值是不完全相同的,可以属于同一通道的潜在变量的元素对应相同的调整因子值,即第一变量调整因子包括N个元素,每个元素对应第二潜在变量中通道序号相同的M个元素。
步骤903:通过第一解码神经网络模型对重构的第一潜在变量进行处理,以得到重构的媒体数据。
在一些实施例中,可以将重构的第一潜在变量输入第一解码神经网络模型,得到第一解码神经网络模型输出的重构的媒体数据。或者,对重构的第一潜在变量进行后处理,将后处理后的第一潜在变量输入第一解码神经网络模型,得到第一解码神经网络模型输出的重构的 媒体数据。
也就是说,可以将重构的第一潜在变量作为第一解码神经网络模型的输入来确定重构的媒体数据,也可以对重构的第一潜在变量进行后处理后,再作为第一解码神经网络模型的输入来确定重构的媒体数据。
其中,第一解码神经网络模型与第一编码神经网络模型对应,均是预先训练好的,本申请实施例对第一解码神经网络模型的网络结构和训练方法不做限定。例如,第一解码神经网络模型的网络结构可以为全连接网络或者CNN网络。另外,本申请实施例对第一解码神经网络模型的网络结构所包含的层数和每一层的节点数也不做限定。
在媒体数据为音频信号和视频信号的情况下,第一解码神经网络模型的输出可以是重构的时域的媒体数据,也可以是重构的频域的媒体数据。如果是频域的媒体数据,则需要经过频域到时域的变换后得到时域的媒体数据。第一解码神经网络模型的输出还可以是残差信号,此时,还需要进行其他相应处理才能得到音频信号或者视频信号。
在本申请实施例中,由于第二潜在变量的熵编码结果的编码比特数满足预设编码速率条件,这样可以保证每帧媒体数据对应的潜在变量的熵编码结果的编码比特数均能够满足预设编码速率条件,也即是,可以保证每帧媒体数据对应的潜在变量的熵编码结果的编码比特数基本保持一致,而不是动态变化的,从而满足了编码器对稳定编码速率的需求。进一步地,在考虑到需要传输边信息(例如窗型,时域噪声整形(TNS:Temporal Noise Shaping)参数,频域噪声整形(FDNS:Frequency-domain noise shaping)参数,和/或带宽扩展(BWE:bandwidth extension)参数等等)时,能够保证每帧媒体数据对应的潜在变量的熵编码结果的编码比特数和边信息的编码比特数整体基本保持一致,从而满足编码器对稳定编码速率的需求。
请参考图10,图10是本申请实施例提供的第二种编码方法的流程图,该方法包括上下文模型,但只对待编码的媒体数据生成的潜在变量通过调整因子进行调整。该编码方法应用于编码端设备,包括如下步骤。
步骤1001:通过第一编码神经网络模型对待编码的媒体数据进行处理,以得到第一潜在变量,第一潜在变量用于指示待编码的媒体数据的特征。
其中,步骤1001的实现过程可以参考上述步骤601的实现过程,此处不再赘述。
步骤1002:基于第一潜在变量确定第一变量调整因子。
在一些实施例中,可以基于第一潜在变量确定初始编码比特数,基于初始编码比特数和目标编码比特数,确定第一变量调整因子。其中,目标编码比特数的相关描述可以参考步骤602中的内容,此处不再赘述。接下来将分别对初始编码比特数和第一变量调整因子的确定过程进行详细解释说明。
确定初始编码比特数
其中,基于第一潜在变量确定初始编码比特数的方式可以包括两种,接下来将分别介绍。
第一种实现方式,基于第一潜在变量,通过上下文模型确定对应的上下文初始编码比特数和初始熵编码模型参数,基于初始熵编码模型参数,确定第一潜在变量的熵编码结果的编码比特数,以得到基础初始编码比特数。基于上下文初始编码比特数与基础初始编码比特数确定为初始编码比特数。
其中,上下文模型包括上下文编码神经网络模型和上下文解码神经网络模型。基于第一 潜在变量,通过上下文模型确定对应的上下文初始编码比特数和初始熵编码模型参数的实现过程包括:通过上下文编码神经网络模型对第一潜在变量进行处理,得到第五潜在变量,第五潜在变量用于指示第一潜在变量的概率分布。确定第五潜在变量的熵编码结果,将第五潜在变量的熵编码结果的编码比特数作为上下文初始编码比特数。基于第五潜在变量的熵编码结果重构第五潜在变量,通过上下文解码神经网络模型对重构得到的第五潜在变量进行处理,得到初始熵编码模型参数。
作为一种示例,通过上下文编码神经网络模型对第一潜在变量进行处理,得到第五潜在变量。对第五潜在变量进行量化处理,得到量化后的第五潜在变量。对量化后的第五潜在变量进行熵编码,并统计该熵编码结果的编码比特数,得到上下文初始编码比特数。之后,对第五潜在变量的熵编码结果进行熵解码,得到经量化的第五潜在变量,对经量化的第五潜在变量进行去量化处理,得到重构的第五潜在变量。将重构的第五潜在变量输入上下文解码神经网络模型,得到上下文解码神经网络模型输出的初始熵编码模型参数。
其中,通过上下文编码神经网络模型对第一潜在变量进行处理的实现过程为:将第一潜在变量输入上下文编码神经网络模型,得到上下文编码神经网络模型输出的第五潜在变量。或者,对第一潜在变量中的每个元素取绝对值,然后再输入上下文编码神经网络模型,得到上下文编码神经网络模型输出的第五潜在变量。
其中,基于初始熵编码模型参数,确定第一潜在变量的熵编码结果的编码比特数,得到基础初始编码比特数的实现过程为:从编码模型参数可调的熵编码模型中确定与该初始熵编码模型参数对应的熵编码模型。对第一潜在变量进行量化处理,得到量化后的第一潜在变量。基于与该初始熵编码模型参数对应的熵编码模型,对量化后的第一潜在变量进行熵编码,得到第一潜在变量的初始编码结果。统计第一潜在变量的初始编码结果的编码比特数,得到基础初始编码比特数。
其中,步骤1002中的量化处理方式和熵编码方式可以参考步骤602中的量化处理方式和熵编码方式,此处不再赘述。
其中,基于上下文初始编码比特数与基础初始编码比特数确定为初始编码比特数的实现过程包括:将上下文初始编码比特数与基础初始编码比特数之和确定为初始编码比特数。当然,还可以通过其他的实现方式来确定。
第二种实现方式,基于第一初始调整因子对第一潜在变量进行调整,基于调整后的第一潜在变量,通过上下文模型确定对应的上下文初始编码比特数和初始熵编码模型参数,基于初始熵编码模型参数,确定调整后的第一潜在变量的熵编码结果的编码比特数,得到基础初始编码比特数。基于上下文初始编码比特数与基础初始编码比特数确定初始编码比特数。其中,第一初始调整因子可以为第一预设调整因子。
其中,上下文模型包括上下文编码神经网络模型和上下文解码神经网络模型。基于调整后的第一潜在变量,通过上下文模型确定对应的上下文初始编码比特数和初始熵编码模型参数的实现过程包括:通过上下文编码神经网络模型对调整后的第一潜在变量进行处理,得到第六潜在变量,第六潜在变量用于指示调整后的第一潜在变量的概率分布。确定第六潜在变量的熵编码结果,将第六潜在变量的熵编码结果的编码比特数作为上下文初始编码比特数。基于第六潜在变量的熵编码结果重构第六潜在变量,通过上下文解码神经网络模型对重构得到的第六潜在变量进行处理,得到初始熵编码模型参数。
作为一种示例,通过上下文编码神经网络模型对调整后的第一潜在变量进行处理,得到第六潜在变量。对第六潜在变量进行量化处理,得到量化后的第六潜在变量。对量化后的第六潜在变量进行熵编码,并统计该熵编码结果的编码比特数,得到上下文初始编码比特数。之后,对第六潜在变量的熵编码结果进行熵解码,得到经量化的第六潜在变量,对经量化的第六潜在变量进行去量化处理,得到重构的第六潜在变量。将重构的第六潜在变量输入上下文解码神经网络模型,得到上下文解码神经网络模型输出的初始熵编码模型参数。
需要说明的是,第二种实现方式中的其他内容可以参考上文对应步骤的描述,此处不再赘述。
确定第一变量调整因子
上述第一变量调整因子为第一潜在变量的调整因子最终值。通过第一变量调整因子对第一潜在变量进行调整后得到第二潜在变量,第二潜在变量经过上下文编码神经网络模型处理后得到第三潜在变量,且第二潜在变量的熵编码结果和第三潜在变量的熵编码结果的编码总比特数满足预设编码速率条件。其中,在使用固定码率对待编码的媒体数据进行编码的情况下,满足预设编码速率条件包括编码比特数小于或等于目标编码比特数。或者,满足预设编码速率条件包括编码比特数小于或等于目标编码比特数,且编码比特数与目标编码比特数的差值小于比特数阈值。在使用可变码率对待编码的媒体数据进行编码的情况下,满足预设编码速率条件包括编码比特数与目标编码比特数的差值的绝对值小于比特数阈值。也即是,满足预设编码速率条件包括编码比特数小于或等于目标编码比特数,且目标编码比特数与编码比特数的差值小于比特数阈值。或者,满足预设编码速率条件包括编码比特数大于或等于目标编码比特数,且编码比特数与目标编码比特数的差值小于比特数阈值。
需要说明的是,比特数阈值可以事先设置,且比特数阈值可以基于不同的需求进行调整。
其中,基于初始编码比特数和目标编码比特数,确定第一变量调整因子的实现方式可以包括多种,接下来对其中的三种进行介绍。
第一种实现方式,在初始编码比特数等于目标编码比特数的情况下,将第一初始调整因子确定为第一变量调整因子。在初始编码比特数不等于目标编码比特数的情况下,基于初始编码比特数和目标编码比特数,通过第一循环方式确定第一变量调整因子。
其中,第一循环方式的第i次循环处理包括如下步骤:确定第i次循环处理的调整因子,i为正整数,基于第i次循环处理的调整因子对第一潜在变量进行调整,得到第i次调整后的第一潜在变量。基于第i次调整后的第一潜在变量,通过上下文模型确定对应的第i次的上下文编码比特数和第i次的熵编码模型参数,基于第i次的熵编码模型参数,确定第i次调整后的第一潜在变量的熵编码结果的编码比特数,得到第i次的基础编码比特数,基于第i次的上下文编码比特数与第i次的基础编码比特数确定第i次的编码比特数。在第i次的编码比特数满足继续调整条件的情况下,执行第一循环方式的第i+1次循环处理。在第i次的编码比特数不满足继续调整条件的情况下,终止第一循环方式的执行,基于第i次循环处理的调整因子确定第一变量调整因子。
其中,确定第i次循环处理的调整因子的实现过程可以参考上述步骤602中的相关描述,此处不再赘述。基于第i次调整后的第一潜在变量,通过上下文模型确定第i次的上下文编码比特数和第i次的熵编码模型参数的实现过程可以参考上述确定上下文初始编码比特数和初始熵编码模型参数的过程,此处不再赘述。基于第i次的编码模型参数,确定第i次调整后的 第一潜在变量的熵编码结果的编码比特数的实现过程可以参考上述基于初始熵编码模型参数确定基础初始编码比特数的过程,此处不再赘述。
其中,基于第i次循环处理的调整因子确定第一变量调整因子的实现过程可以参考上述步骤602中的描述,此处不再赘述。
第二种实现方式,在初始编码比特数等于目标编码比特数的情况下,将第一初始调整因子确定为第一变量调整因子。在初始编码比特数不等于目标编码比特数的情况下,按照上述第一种实现方式来确定第一变量调整因子。但是,与上述第一种实现方式不同的是,在初始编码比特数小于目标编码比特数的情况下,按照第一步长调整第一循环方式的第i-1次循环处理的调整因子,得到第i次循环处理的调整因子,此时,继续调整条件包括第i次的编码比特数小于目标编码比特数。在初始编码比特数大于目标编码比特数的情况下,按照第二步长调整第一循环方式的第i-1次循环处理的调整因子,得到第i次循环处理的调整因子,此时,继续调整条件包括第i次的编码比特数大于目标编码比特数。
其中,按照第一步长调整第i-1次循环处理的调整因子可以是指按照第一步长增大第i-1次循环处理的调整因子,按照第二步长调整第i-1次循环处理的调整因子可以是指按照第二步长减小第i-1次循环处理的调整因子。
上述的增大处理和减小处理可以为线性的,也可以为非线性的。示例地,可以将第i-1次循环处理的调整因子与第一步长之和确定为第i次循环处理的调整因子,可以将第i-1次循环处理的调整因子与第二步长之差确定为第i次循环处理的调整因子。
需要说明的是,第一步长和第二步长可以为事先设置的,且第一步长和第二步长可以基于不同的需求来调整。另外,第一步长与第二步长可以相等,也可以不相等。
第三种实现方式,在初始编码比特数小于或等于目标编码比特数的情况下,将第一初始调整因子确定为第一变量调整因子。在初始编码比特数大于目标编码比特数的情况下,可以按照上述第一种实现方式来确定第一变量调整因子。也可以按照上述第二种实现方式中初始编码比特数大于目标编码比特数的情况来确定第一变量调整因子。
步骤1003:获取第二潜在变量的熵编码结果和第三潜在变量的熵编码结果,第二潜在变量是通过第一变量调整因子对第一潜在变量调整后得到,第三潜在变量用于指示第二潜在变量的概率分布,第二潜在变量的熵编码结果和第三潜在变量的熵编码结果的编码总比特数满足预设编码速率条件。
在步骤1002确定第一变量调整因子的过程中,第一变量调整因子可能为第一初始调整因子,可能为第i次循环处理的调整因子,也可能是第j次循环处理的第三调整因子,也可能是基于第j次循环处理的第一调整因子确定的。不管是哪种情况,上述循环过程中都有确定对应的熵编码结果,所以,可以直接从上述循环过程中确定得到的熵编码结果中获取第二潜在变量的熵编码结果和第三潜在变量的熵编码结果。
当然,还可以重新处理一次。即,直接基于第一变量调整因子对第一潜在变量进行调整,得到第二潜在变量。基于第二潜在变量,通过上下文模型确定第三潜在编码的熵编码结果和第一熵编码模型参数。对第二潜在变量进行量化处理,得到量化后的第二潜在变量。基于第一熵编码模型参数,确定量化后的第二潜在变量的熵编码结果,即第二潜在变量的熵编码结果。
其中,上下文模型包括上下文编码神经网络模型和上下文解码神经网络模型,基于第二 潜在变量,通过上下文模型确定第三潜在编码的熵编码结果和第一熵编码模型参数的实现过程包括:通过上下文编码神经网络模型对第二潜在变量进行处理,得到第三潜在变量。对第三潜在变量进行量化处理,得到量化后的第三潜在变量。对量化后的第三潜在变量进行熵编码,得到第三潜在变量的熵编码结果。对第三潜在变量的熵编码结果进行熵解码,得到经量化的第三潜在变量,对经量化的第三潜在变量进行去量化处理,得到重构的第三潜在变量。通过上下文解码神经网络模型对重构的第三潜在变量进行处理,得到第一熵编码模型参数。
其中,基于第一熵编码模型参数,确定量化后的第二潜在变量的熵编码结果的实现过程包括:从编码模型参数可调的熵编码模型中确定与第一熵编码模型参数对应的熵编码模型。基于与第一熵编码模型参数对应的熵编码模型,对量化后的第二潜在变量进行熵编码,得到第二潜在变量的熵编码结果。
步骤1004:将第二潜在变量的熵编码结果、第三潜在变量的熵编码结果以及第一变量调整因子的编码结果写入码流。
其中,关于第一变量调整因子的编码结果的相关内容可以参考步骤604中的描述,此处不再赘述。
可选地,若第一变量调整因子为量化前的值,本申请实施例还可以确定第一变量调整因子的量化步长对应的量化索引,得到第一量化索引。可选地,本申请实施例还可以确定第二潜在变量的量化步长对应的量化索引,得到第二量化索引,以及确定第三潜在变量的量化步长对应的量化索引,得到第三量化索引。将第一量化索引、第二量化索引和第三量化索引编入码流。其中,量化索引用于指示对应的量化步长,即,第一量化索引用于指示第一变量调整因子的量化步长,第二量化索引用于指示第二潜在变量的量化步长,第三量化索引用于指示第三潜在变量的量化步长。
在本申请实施例中,通过第一变量调整因子对第一潜在变量进行调整,得到第二潜在变量,通过上下文编码神经网络模型对第二潜在变量进行处理,得到第三潜在变量,而且第二潜在变量的熵编码结果和第三潜在变量的熵编码结果的编码总比特数满足预设编码速率条件,这样可以保证每帧媒体数据对应的潜在变量的熵编码结果的编码比特数均能够满足预设编码速率条件,也即是,可以保证每帧媒体数据对应的潜在变量的熵编码结果的编码比特数基本保持一致,而不是动态变化的,从而满足了编码器对稳定编码速率的需求。进一步地,在考虑到需要传输边信息(例如窗型,时域噪声整形(TNS:Temporal Noise Shaping)参数,频域噪声整形(FDNS:Frequency-domain noise shaping)参数,和/或带宽扩展(BWE:bandwidth extension)参数等等)时,能够保证每帧媒体数据对应的潜在变量的熵编码结果的编码比特数和边信息的编码比特数整体基本保持一致,从而满足编码器对稳定编码速率的需求。
请参考图11,图11是本申请实施例提供的第二种解码方法的流程图,该方法应用于解码端,且该方法对应于图10所示的编码方法。该方法包括如下步骤。
步骤1101:基于码流确定重构的第三潜在变量和重构的第一变量调整因子。
在一些实施例中,可以对码流中的第三潜在变量的熵编码结果进行熵解码,以及对码流中的第一变量调整因子的编码结果进行解码,得到经量化的第三潜在变量和经量化的第一变量调整因子。对经量化的第三潜在变量和经量化的第一变量调整因子进行去量化处理,得到重构的第三潜在变量和重构的第一变量调整因子。
其中,本步骤的解码方法与编码端的编码方法相对应,本步骤的去量化处理与编码端的量化处理相对应。即,解码方法为编码方法的逆过程,去量化处理为量化处理的逆过程。
示例地,可以从码流中解析出第一量化索引和第三量化索引,第一量化索引用于指示第一变量调整因子的量化步长,第三量化索引用于指示第三潜在变量的量化步长。按照第一量化索引所指示的量化步长,对经量化的第一变量调整因子进行去量化处理,得到重构的第一变量调整因子。按照第三量化索引所指示的量化步长,对经量化的第三潜在变量进行去量化处理,得到重构的第三潜在变量。
步骤1102:基于码流和重构的第三潜在变量,确定重构的第二潜在变量。
在一些实施例中,通过上下文解码神经网络模型对重构的第三潜在变量进行处理,得到重构的第一熵编码模型参数,基于码流和重构的第一熵编码模型参数,确定重构的第二潜在变量。
在一些实施例中,可以确定与重构的第一熵编码模型参数对应的熵解码模型。基于与重构的第一熵编码模型参数对应的熵解码模型,对码流中的第二潜在变量的熵编码结果进行熵解码,得到经量化的第二潜在变量。对经量化的第二潜在变量进行去量化处理,得到重构的第二潜在变量。
其中,本步骤的解码方法与编码端的编码方法相对应,本步骤的去量化处理与编码端的量化处理相对应。即,解码方法为编码方法的逆过程,去量化处理为量化处理的逆过程。
示例地,可以从码流中解析出第二量化索引,第二量化索引用于指示第二潜在变量的量化步长。按照第二量化索引所指示的量化步长,对经量化的第二潜在变量进行去量化处理,得到重构的第二潜在变量。
步骤1103:基于重构的第一变量调整因子,对重构的第二潜在变量进行调整,得到重构的第一潜在变量,重构的第一潜在变量用于指示待解码的媒体数据的特征。
其中,步骤1103的实现过程可以参考上述步骤902的实现过程,此处不再赘述。
步骤1104:通过第一解码神经网络模型对重构的第一潜在变量进行处理,以得到重构的媒体数据。
其中,步骤1104的实现过程可以参考上述步骤903的实现过程,此处不再赘述。
在本申请实施例中,由于第二潜在变量的熵编码结果和第三潜在变量的熵编码结果的编码总比特数满足预设编码速率条件,这样可以保证每帧媒体数据对应的潜在变量的熵编码结果的编码比特数均能够满足预设编码速率条件,也即是,可以保证每帧媒体数据对应的潜在变量的熵编码结果的编码比特数基本保持一致,而不是动态变化的,从而满足了编码器对稳定编码速率的需求。进一步地,在考虑到需要传输边信息(例如窗型,时域噪声整形(TNS:Temporal Noise Shaping)参数,频域噪声整形(FDNS:Frequency-domain noise shaping)参数,和/或带宽扩展(BWE:bandwidth extension)参数等等)时,能够保证每帧媒体数据对应的潜在变量的熵编码结果的编码比特数和边信息的编码比特数整体基本保持一致,从而满足编码器对稳定编码速率的需求。
请参考图12,图12是本申请实施例提供的一种示例性编码方法的框图。图12主要是对图10所示的编码方法进行示例性解释。在图12中,以音频信号为例。可以对音频信号经过加窗处理,得到当前帧音频信号。对当前帧音频信号进行MDCT变换处理后,得到当前帧的 频域信号。基于当前帧的频域信号,经过第一编码神经网络模型处理,输出第一潜在变量。基于第一潜在变量调整因子对第一潜在变量进行调整,得到第二潜在变量。通过上下文编码神经网络模型对第二潜在变量进行处理,得到第三潜在变量。对第三潜在变量进行量化处理和熵编码,得到第三潜在变量的熵编码结果,将第三潜在变量的熵编码结果写入码流。同时,对第三潜在变量的熵编码结果进行熵解码,得到经量化的第三潜在变量。对经量化的第三潜在变量进行去量化处理,得到重构的第三潜在变量。通过上下文解码神经网络模型对重构的第三潜在变量进行处理,得到第一熵编码模型参数。从参数可调的熵编码模型中选择与第一熵编码模型参数对应的熵编码模型。对第二潜在变量进行量化,基于选择的熵编码模型对量化后的第二潜在变量进行熵编码,得到第二潜在变量的熵编码结果,将第二潜在变量的熵编码结果写入码流。然后,将第一变量调整因子的编码结果写入码流。
请参考图13,图13是本申请实施例提供的一种示例性解码方法的框图。图13主要是对图11所示的解码方法进行示例性解释。在图13中,以音频信号为例。通过熵解码模型对码流中的第三潜在变量的熵编码结果进行熵解码,得到经量化的第三潜在变量。对经量化的第三潜在变量进行去量化处理,得到重构的第三潜在变量。通过上下文解码神经网络模型对重构的第三变量进行处理,得到重构的第一熵编码模型参数。基于重构的第一熵编码模型参数,从参数可调的熵解码模型中选择对应的熵解码模型。按照选择的熵解码模型对码流中的第二潜在变量的熵编码结果进行熵解码,得到经量化的第二潜在变量。对经量化的第二潜在变量进行去量化处理,得到重构的第二潜在变量。对码流中第一变量调整因子的编码结果进行解码,得到重构的第一变量调整因子。基于重构的第一变量调整因子对重构的第二潜在变量进行调整,得到重构的第一潜在变量。通过第一解码神经网络模型对重构的第一潜在变量进行处理,得到重构的当前帧的频域信号。对重构的当前帧的频域信号进行IMDCT变换处理以及去加窗处理,得到重构的音频信号。
请参考图14,图14是本申请实施例提供的第三种编码方法的流程图,该方法包括上下文模型,而且不仅对待编码的媒体数据生成的潜在变量通过调整因子进行调整,还对上下文模型确定的潜在变量通过调整因子进行调整。该编码方法应用于编码端设备,包括如下步骤。
步骤1401:通过第一编码神经网络模型对待编码的媒体数据进行处理,以得到第一潜在变量,第一潜在变量用于指示待编码的媒体数据的特征。
其中,步骤1401的实现过程可以参考上述步骤601的实现过程,此处不再赘述。
步骤1402:基于第一潜在变量确定第一变量调整因子和第二变量调整因子。
在一些实施例中,可以基于第一潜在变量,通过上下文模型确定对应的上下文初始编码比特数和初始熵编码模型参数,基于初始熵编码模型参数,确定第一潜在变量的熵编码结果的编码比特数,以得到基础初始编码比特数,基于上下文初始编码比特数、基础初始编码比特数和目标编码比特数,确定第一变量调整因子和第二变量调整因子。
其中,基于第一潜在变量,通过上下文模型确定对应的上下文初始编码比特数和初始熵编码模型参数的实现过程,以及基于初始熵编码模型参数,确定第一潜在变量的熵编码结果的编码比特数,以得到基础初始编码比特数的实现过程可以参考上述步骤1002中的相关描述,此处不再赘述。
在另一些实施例中,可以将第一预设调整因子作为第一初始调整因子,将第二预设调整因子作为第二初始调整因子。基于第一初始调整因子对第一潜在变量进行调整,基于调整后的第一潜在变量和第二初始调整因子,通过上下文模型确定对应的上下文初始编码比特数和初始熵编码模型参数,基于初始熵编码模型参数,确定第一潜在变量的熵编码结果的编码比特数,以得到基础初始编码比特数。基于上下文初始编码比特数、基础初始编码比特数和目标编码比特数,确定第一变量调整因子和第二变量调整因子。
其中,基于第一初始调整因子对第一潜在变量进行调整的实现过程可以参考上述步骤1002中的相关描述,此处不再赘述。
其中,上下文模型包括上下文编码神经网络模型和上下文解码神经网络模型。基于调整后的第一潜在变量和第二初始调整因子,通过上下文模型确定对应的上下文初始编码比特数和初始熵编码模型参数的实现过程包括:通过上下文编码神经网络模型对调整后的第一潜在变量进行处理,得到第六潜在变量,第六潜在变量用于指示调整后的第一潜在变量的概率分布。基于第二初始调整因子对第六潜在变量进行调整,得到第七潜在变量。确定第七潜在变量的熵编码结果,将第七潜在变量的熵编码结果的编码比特数作为上下文初始编码比特数。基于第七潜在变量的熵编码结果重构第七潜在变量,通过上下文解码神经网络模型对重构得到的第七潜在变量进行处理,得到初始熵编码模型参数。
作为一种示例,通过上下文编码神经网络模型对调整后的第一潜在变量进行处理,得到第六潜在变量。基于第二初始调整因子对第六潜在变量进行调整,得到第七潜在变量。对第七潜在变量进行量化处理,得到量化后的第七潜在变量。对量化后的第七潜在变量进行熵编码,并统计该熵编码结果的编码比特数,得到上下文初始编码比特数。之后,对第七潜在变量的熵编码结果进行熵解码,得到经量化的第七潜在变量,对经量化的第七潜在变量进行去量化处理,得到重构的第七潜在变量。将重构的第七潜在变量输入上下文解码神经网络模型,得到上下文解码神经网络模型输出的初始熵编码模型参数。
其中,基于第二初始调整因子对第六潜在变量进行调整的实现过程为:将第六潜在变量中的各个元素与第二初始调整因子中对应的元素相乘,得到第七潜在变量。
值得注意的是,上述实现过程仅仅为一种示例,实际应用中,还可以采用其他的方法来调整。比如,可以将第六潜在变量中的各个元素除以第二初始调整因子中对应的元素,得到第七潜在变量。本申请实施例对调整方法不做限定。
其中,其余步骤的具体实现过程可以参考前文对应步骤的描述,此处不再赘述。
需要说明的是,本申请实施例针对第三潜在变量设置有调整因子初值,调整因子初值通常等于1。第二预设调整因子可以大于或等于调整因子初值,也可以小于调整因子初值,比如,第二预设调整因子为1或2等常数。在通过上述第一种实现方式来确定第一变量调整因子和第二变量调整因子的情况下,第二初始调整因子为调整因子初值,在通过上述第二种实现方式来确定第一变量调整因子和第二变量调整因子的情况下,第二初始调整因子为第二预设调整因子。
另外,第二初始调整因子可以是标量也可以是矢量。具体介绍请参考第一初始调整因子的相关介绍。
其中,基于上下文初始编码比特数、基础初始编码比特数和目标编码比特数,确定第一 变量调整因子和第二变量调整因子的实现过程可以包括两种,接下来分别介绍。
第一种实现方式,将第二变量调整因子置为第二初始调整因子,基于基础初始编码比特数和上下文初始编码比特数中的至少一者,以及目标编码比特数,确定基础目标编码比特数。基于第二初始调整因子、基础目标编码比特数和基础初始编码比特数,确定第一变量调整因子和基础实际编码比特数,基础实际编码比特数是指经过第一变量调整因子调整后的第一潜在变量的熵编码结果的编码比特数。基于目标编码比特数和基础实际编码比特数,确定上下文目标编码比特数。基于上下文目标编码比特数和上下文初始编码比特数,确定第二变量调整因子。
其中,基于基础初始编码比特数和上下文初始编码比特数中的至少一者,以及目标编码比特数,确定基础目标编码比特数的实现过程包括:将目标编码比特数减去上下文初始编码比特数,得到基础目标编码比特数。或者,确定基础初始编码比特数和上下文初始编码比特数之间的比例,基于该比例和目标编码比特数,确定基础目标编码比特数。或者,基于目标编码比特数与基础初始编码比特数的比值,确定基础目标编码比特数。当然,还可以通过其他实现过程来确定。
比如,确定基础初始编码比特数和上下文初始编码比特数之间的比例为5:3,此时,可以将目标编码比特数乘以5/8,得到基础目标编码比特数。
又比如,确定目标编码比特数与基础初始编码比特数的比值,如果目标编码比特数与基础初始编码比特数的比值大于第一比例门限,则将基础初始编码比特数与预设的第一比例调整步长之和确定为基础目标编码比特数。如果目标编码比特数与基础初始编码比特数的比值小于第二比例门限,则将基础初始编码比特数与预设的第二比例调整步长之差确定为基础目标编码比特数。其中,第一比例门限大于第二比例门限。如果目标编码比特数与基础初始编码比特数的比值大于或等于第二比例门限且小于或等于第一比例门限,则将基础初始编码比特数确定为基础目标编码比特数。
需要说明的是,第一比例调整步长可以大于第二比例调整步长,也可以小于第二比例调整步长,当然,也可以等于第二比例调整步长,本申请实施例对第一比例调整步长与第二比例调整步长的大小关系不做限定。
其中,基于第二初始调整因子、基础目标编码比特数和基础初始编码比特数确定第一变量调整因子的实现方式也包括三种,接下来分别介绍。
方式11,在基础初始编码比特数等于基础目标编码比特数的情况下,将第一初始调整因子确定为第一变量调整因子。在基础初始编码比特数不等于基础目标编码比特数的情况下,基于基础初始编码比特数和基础目标编码比特数,通过第一循环方式确定第一变量调整因子。
其中,第一循环方式的第i次循环处理包括如下步骤:确定第i次循环处理的调整因子,i为正整数,基于第i次循环处理的调整因子对第一潜在变量进行调整,得到第i次调整后的第一潜在变量。基于第i次调整后的第一潜在变量和第二初始调整因子,通过上下文模型确定第i次的上下文编码比特数和第i次的熵编码模型参数,基于第i次的熵编码模型参数,确定第i次调整后的第一潜在变量的熵编码结果的编码比特数,得到第i次的基础编码比特数,基于第i次的上下文编码比特数与第i次的基础编码比特数确定第i次的编码比特数。在第i次的编码比特数满足继续调整条件的情况下,执行第一循环方式的第i+1次循环处理。在第i次的编码比特数不满足继续调整条件的情况下,终止第一循环方式的执行,基于第i次循环处理的调整因子确定第一变量调整因子。
其中,确定第一变量调整因子的方式11中的内容与步骤1002中确定第一变量调整因子的第一种实现方式中的内容类似。不同的地方在于,本步骤是基于第i次调整后的第一潜在变量和第二初始调整因子,通过上下文模型确定第i次的上下文编码比特数和第i次的熵编码模型参数。其中,基于第i次调整后的第一潜在变量和第二初始调整因子,通过上下文模型确定第i次的上下文编码比特数和第i次的熵编码模型参数的实现过程可以参考上文基于调整后的第一潜在变量和第二初始调整因子,通过上下文模型确定对应的上下文初始编码比特数和初始熵编码模型参数的实现过程,此处不再赘述。
方式12,在基础初始编码比特数等于基础目标编码比特数的情况下,将第一初始调整因子确定为第一变量调整因子。在基础初始编码比特数不等于基础目标编码比特数的情况下,按照上述方式11来确定第一变量调整因子。但是,与上述方式11不同的是,在基础初始编码比特数小于基础目标编码比特数的情况下,按照第一步长调整第一循环方式的第i-1次循环处理的调整因子,得到第i次循环处理的调整因子,此时,继续调整条件包括第i次的编码比特数小于基础目标编码比特数。在基础初始编码比特数大于基础目标编码比特数的情况下,按照第二步长调整第一循环方式的第i-1次循环处理的调整因子,得到第i次循环处理的调整因子,此时,继续调整条件包括第i次的编码比特数大于基础目标编码比特数。
其中,方式12中的相关内容可以参考步骤1002中第二种实现方式的内容,此处不再赘述。
方式13,在基础初始编码比特数小于或等于基础目标编码比特数的情况下,将第一初始调整因子确定为第一变量调整因子。在基础初始编码比特数大于基础目标编码比特数的情况下,可以按照上述方式11来确定第一变量调整因子。也可以按照上述方式12中基础初始编码比特数大于基础目标编码比特数的情况来确定第一变量调整因子。
其中,基于目标编码比特数和基础实际编码比特数,确定上下文目标编码比特数的实现过程包括:将目标编码比特数减去基础实际编码比特数,得到上下文目标编码比特数。
其中,基于上下文目标编码比特数和上下文初始编码比特数确定第二变量调整因子的实现过程与基于基础初始编码比特数和基础目标编码比特数确定第一变量调整因子的实现过程类似。接下来也分为三种方式分别介绍。
方式21,在上下文初始编码比特数等于上下文目标编码比特数的情况下,将第二初始调整因子确定为第二变量调整因子。在上下文初始编码比特数不等于上下文目标编码比特数的情况下,基于第一变量调整因子对第一潜在变量进行调整,得到第二潜在变量。基于上下文初始编码比特数、上下文目标编码比特数和第二潜在变量,通过第一循环方式确定第二变量调整因子。
其中,第一循环方式的第i次循环处理包括如下步骤:确定第i次循环处理的调整因子,i为正整数,基于第二潜在变量和第i次循环处理的调整因子,通过上下文模型确定第i次的上下文编码比特数和第i次的熵编码模型参数,基于第i次的熵编码模型参数,确定第二潜在变量的熵编码结果的编码比特数,得到第i次的基础编码比特数,基于第i次的上下文编码比特数与第i次的基础编码比特数确定第i次的编码比特数。在第i次的编码比特数满足继续调整条件的情况下,执行第一循环方式的第i+1次循环处理。在第i次的编码比特数不满足继续调整条件的情况下,终止第一循环方式的执行,基于第i次循环处理的调整因子确定第二变量调整因子。
其中,基于第二潜在变量和第i次循环处理的调整因子,通过上下文模型确定第i次的上下文编码比特数和第i次的熵编码模型参数的实现过程包括:通过上下文编码神经网络模型对第二潜在变量进行处理,得到第三潜在变量,第三潜在变量用于指示第二潜在变量的概率分布。基于第i次循环处理的调整因子对第三潜在变量进行调整,得到第i次调整后的第三潜在变量。确定第i次调整后的第三潜在变量的熵编码结果,将第i次调整后的第三潜在变量的熵编码结果的编码比特数作为第i次的上下文编码比特数。基于第i次调整后的第三潜在变量的熵编码结果重构第i次调整后的第三潜在变量。通过第i次循环处理的调整因子,对重构的第i次调整后的第三潜在变量进行调整,得到重构的第三潜在变量。通过上下文解码神经网络模型对重构的第三潜在变量进行处理,得到第i次的熵编码模型参数。
其余步骤可以参考前文相应内容的描述,此处不再赘述。
方式22,在上下文初始编码比特数等于上下文目标编码比特数的情况下,将第二初始调整因子确定为第二变量调整因子。在上下文初始编码比特数不等于上下文目标编码比特数的情况下,按照上述方式21来确定第二变量调整因子。但是,与上述方式21不同的是,在上下文初始编码比特数小于上下文目标编码比特数的情况下,按照第一步长调整第一循环方式的第i-1次循环处理的调整因子,得到第i次循环处理的调整因子,此时,继续调整条件包括第i次的编码比特数小于上下文目标编码比特数。在上下文初始编码比特数大于上下文目标编码比特数的情况下,按照第二步长调整第一循环方式的第i-1次循环处理的调整因子,得到第i次循环处理的调整因子,此时,继续调整条件包括第i次的编码比特数大于上下文目标编码比特数。
方式23,在上下文初始编码比特数小于或等于上下文目标编码比特数的情况下,将第二初始调整因子确定为第二变量调整因子。在上下文初始编码比特数大于上下文目标编码比特数的情况下,可以按照上述方式21来确定第二变量调整因子。也可以按照上述方式22中上下文初始编码比特数大于上下文目标编码比特数的情况来确定第二变量调整因子。
第二种实现方式,将目标编码比特数划分为基础目标编码比特数和上下文目标编码比特数,基于基础目标编码比特数和基础初始编码比特数,确定第一变量调整因子。基于上下文目标编码比特数和上下文初始编码比特数,确定第二变量调整因子。
其中,在基于基础目标编码比特数和基础初始编码比特数,确定第一变量调整因子时,可以将第二变量调整因子置为第二初始调整因子,基于第二初始调整因子、基础目标编码比特数和基础初始编码比特数,确定第一变量调整因子。具体实现过程可以参考前文描述。
另外,在基于上下文目标编码比特数和上下文初始编码比特数,确定第二变量调整因子时,第一变量调整因子已经确定,此时,可以直接基于第一变量调整因子、上下文目标编码比特数和上下文初始编码比特数,确定第二变量调整因子。具体实现过程可以参考前文描述。
步骤1403:获取第二潜在变量的熵编码结果和第四潜在变量的熵编码结果,第二潜在变量是通过第一变量调整因子对第一潜在变量调整后得到,第四潜在变量是基于第二变量调整因子对第三潜在变量调整后得到,第三潜在变量是基于第二潜在变量通过上下文模型确定得到,第三潜在变量用于指示第二潜在变量的概率分布,且第二潜在变量的熵编码结果和第四潜在变量的熵编码结果的编码总比特数满足预设编码速率条件。
在步骤1402确定第一变量调整因子和第二变量调整因子的过程中,上述循环过程中都有确定每次循环处理的调整因子对应的熵编码结果,所以,可以直接从上述循环过程中确定得 到的熵编码结果中获取第二潜在变量的熵编码结果和第四潜在变量的熵编码结果。
当然,还可以重新处理一次。即,直接基于第一变量调整因子对第一潜在变量进行调整,得到第二潜在变量。基于第二潜在变量,通过上下文模型确定第四潜在编码的熵编码结果和第二熵编码模型参数。对第二潜在变量进行量化处理,得到量化后的第二潜在变量。基于第二熵编码模型参数,确定量化后的第二潜在变量的熵编码结果,即第二潜在变量的熵编码结果。
其中,上下文模型包括上下文编码神经网络模型和上下文解码神经网络模型,基于第二潜在变量,通过上下文模型确定第四潜在编码的熵编码结果和第二熵编码模型参数的实现过程包括:通过上下文编码神经网络模型对第二潜在变量进行处理,得到第三潜在变量。基于第二变量调整因子对第三潜在变量进行调整,得到第四潜在变量。对第四潜在变量进行量化处理,得到量化后的第四潜在变量。对量化后的第四潜在变量进行熵编码,得到第四潜在变量的熵编码结果。对第四潜在变量的熵编码结果进行熵解码,得到经量化的第四潜在变量,对经量化的第四潜在变量进行去量化处理,得到重构的第四潜在变量。基于第二变量调整因子对重构的第四潜在变量进行调整,得到重构的第三潜在变量。通过上下文解码神经网络模型对重构的第三潜在变量进行处理,得到第二熵编码模型参数。
其中,基于第二熵编码模型参数,确定量化后的第二潜在变量的熵编码结果的实现过程包括:从编码模型参数可调的熵编码模型中确定与第二熵编码模型参数对应的熵编码模型。基于与第二熵编码模型参数对应的熵编码模型,对量化后的第二潜在变量进行熵编码,得到第二潜在变量的初始编码结果。
步骤1404:将第二潜在变量的熵编码结果、第四潜在变量的熵编码结果、第一变量调整因子的编码结果以及第二变量调整因子的编码结果写入码流。
其中,关于第一变量调整因子的编码结果的相关内容可以参考步骤604中的描述,此处不再赘述。
基于上述描述,第二变量调整因子可以为量化前的值,也可以为量化后的值。若第二变量调整因子为量化前的值,此时,确定第二变量调整因子的编码结果的实现过程包括:对第二变量调整因子进行量化及编码处理,得到第二变量调整因子的编码结果。若第二变量调整因子为量化后的值,此时,确定第二变量调整因子的编码结果的实现过程包括:对第二变量调整因子进行编码,得到第二变量调整因子的编码结果。其中,第二变量调整因子可以采用任何一种编码方式进行编码,本申请实施例对此不做限定。
另外,对于某些编码器来说,量化过程和编码过程是在一次处理中执行的,也即是,通过一次处理能够得到量化结果和编码结果。所以,对于第二变量调整因子来说,在上述确定第二变量调整因子的过程中可能也会得到第二变量调整因子的编码结果,所以,可以直接获取到第二变量调整因子的编码结果。也即是,第二变量调整因子的编码结果也可以在确定第二变量调整因子的过程中直接得到。
可选地,若第一变量调整因子和第二变量调整因子为量化前的值,本申请实施例还可以确定第一变量调整因子的量化步长对应的量化索引,得到第一量化索引,以及确定第二变量调整因子的量化步长,得到第五量化索引。可选地,本申请实施例还可以确定第二潜在变量的量化步长对应的量化索引,得到第二量化索引,确定第四潜在变量的量化步长对应的量化索引,得到第四量化索引,将第一量化索引、第二量化索引、第四量化索引和第五量化索引 编入码流。其中,量化索引用于指示对应的量化步长,即,第一量化索引用于指示第一变量调整因子的量化步长,第二量化索引用于指示第二潜在变量的量化步长,第四量化索引用于指示第四潜在变量的量化步长,第五量化索引用于指示第二变量调整因子的量化步长。
在本申请实施例中,通过第一变量调整因子对第一潜在变量进行调整,得到第二潜在变量,通过上下文编码神经网络模型对第二潜在变量进行处理,得到第三潜在变量。通过第二变量调整因子对第三潜在变量进行调整,得到第四潜在变量。而且第二潜在变量的熵编码结果和第四潜在变量的熵编码结果的编码总比特数满足预设编码速率条件,这样可以保证每帧媒体数据对应的潜在变量的熵编码结果的编码比特数均能够满足预设编码速率条件,也即是,可以保证每帧媒体数据对应的潜在变量的熵编码结果的编码比特数基本保持一致,而不是动态变化的,从而满足了编码器对稳定编码速率的需求。进一步地,在考虑到需要传输边信息(例如窗型,时域噪声整形(TNS:Temporal Noise Shaping)参数,频域噪声整形(FDNS:Frequency-domain noise shaping)参数,和/或带宽扩展(BWE:bandwidth extension)参数等等)时,能够保证每帧媒体数据对应的潜在变量的熵编码结果的编码比特数和边信息的编码比特数整体基本保持一致,从而满足编码器对稳定编码速率的需求。
请参考图15,图15是本申请实施例提供的第三种解码方法的流程图,该方法应用于解码端。该方法对应于图14所示的编码方法,该方法包括如下步骤。
步骤1501:基于码流确定重构的第四潜在变量、重构的第一变量调整因子和重构的第二变量调整因子。
在一些实施例中,可以对码流中的第四潜在变量的熵编码结果进行熵解码,以及对码流中的第一变量调整因子和第二变量调整因子的编码结果进行解码,得到经量化的第四潜在变量、经量化的第一变量调整因子和经量化的第二变量调整因子。对经量化的第四潜在变量、经量化的第一变量调整因子和经量化的第二变量调整因子进行去量化处理,得到重构的第四潜在变量、重构的第一变量调整因子和重构的第二变量调整因子。
其中,本步骤的解码方法与编码端的编码方法相对应,本步骤的去量化处理与编码端的量化处理相对应。即,解码方法为编码方法的逆过程,去量化处理为量化处理的逆过程。
示例地,可以从码流中解析出第一量化索引、第四量化索引和第五量化索引,第一量化索引用于指示第一变量调整因子的量化步长,第四量化索引用于指示第四潜在变量的量化步长,第五量化索引用于指示第二变量调整因子的量化步长。按照第一量化索引所指示的量化步长,对经量化的第一变量调整因子进行去量化处理,得到重构的第一变量调整因子。按照第四量化索引所指示的量化步长,对经量化的第四潜在变量进行去量化处理,得到重构的第四潜在变量。按照第五量化索引所指示的量化步长,对经量化的第二变量调整因子进行去量化处理,得到重构的第二变量调整因子。
步骤1502:基于码流、重构的第四潜在变量和重构的第二变量调整因子,确定重构的第二潜在变量。
在一些实施例中,基于重构的第二变量调整因子对重构的第四潜在变量进行调整,得到重构的第三潜在变量。通过上下文解码神经网络模型对重构的第三潜在变量进行处理,得到重构的第二熵编码模型参数,基于码流和重构的第二熵编码模型参数,确定重构的第二潜在变量。
在一些实施例中,可以确定与重构的第二熵编码模型参数对应的熵解码模型。基于与第二熵编码模型参数对应的熵解码模型,对码流中的第二潜在变量的熵编码结果进行熵解码,得到经量化的第二潜在变量。对经量化的第二潜在变量进行去量化处理,得到重构的第二潜在变量。
其中,本步骤的解码方法与编码端的编码方法相对应,本步骤的去量化处理与编码端的量化处理相对应。即,解码方法为编码方法的逆过程,去量化处理为量化处理的逆过程。
示例地,可以从码流中解析出第二量化索引,第二量化索引用于指示第二潜在变量的量化步长。按照第二量化索引所指示的量化步长,对经量化的第二潜在变量进行去量化处理,得到重构的第二潜在变量。
步骤1503:基于重构的第一变量调整因子,对重构的第二潜在变量进行调整,得到重构的第一潜在变量。
其中,步骤1503的实现过程可以参考上述步骤902的实现过程,此处不再赘述。
步骤1504:通过第一解码神经网络模型对重构的第一潜在变量进行处理,以得到重构的媒体数据。
其中,步骤1504的实现过程可以参考上述步骤903的实现过程,此处不再赘述。
在本申请实施例中,由于第二潜在变量的熵编码结果和第四潜在变量的熵编码结果的编码总比特数满足预设编码速率条件,这样可以保证每帧媒体数据对应的潜在变量的熵编码结果的编码比特数均能够满足预设编码速率条件,也即是,可以保证每帧媒体数据对应的潜在变量的熵编码结果的编码比特数基本保持一致,而不是动态变化的,从而满足了编码器对稳定编码速率的需求。进一步地,在考虑到需要传输边信息(例如窗型,时域噪声整形(TNS:Temporal Noise Shaping)参数,频域噪声整形(FDNS:Frequency-domain noise shaping)参数,和/或带宽扩展(BWE:bandwidth extension)参数等等)时,能够保证每帧媒体数据对应的潜在变量的熵编码结果的编码比特数和边信息的编码比特数整体基本保持一致,从而满足编码器对稳定编码速率的需求。
请参考图16,图16是本申请实施例提供的一种示例性编码方法的框图。图16主要是对图14所示的编码方法进行示例性解释。在图16中,以音频信号为例。可以对音频信号经过加窗处理,得到当前帧音频信号。对当前帧音频信号进行MDCT变换处理后,得到当前帧的频域信号。基于当前帧的频域信号,经过第一编码神经网络模型处理,输出第一潜在变量。基于第一潜在变量调整因子对第一潜在变量进行调整,得到第二潜在变量。通过上下文编码神经网络模型对第二潜在变量进行处理,得到第三潜在变量。基于第二变量调整因子对第三潜在变量进行调整,得到第四潜在变量。对第四潜在变量进行量化处理和熵编码,得到第四潜在变量的熵编码结果,将第四潜在变量的熵编码结果写入码流。同时,对第四潜在变量的熵编码结果进行熵解码,得到经量化的第四潜在变量。对经量化的第四潜在变量进行去量化处理,得到重构的第四潜在变量。基于第二变量调整因子对重构的第四潜在变量进行调整,得到重构的第三潜在变量。通过上下文解码神经网络模型对重构的第三潜在变量进行处理,得到第二熵编码模型参数。从参数可调的熵编码模型中选择与第二熵编码模型参数对应的熵编码模型。对第二潜在变量进行量化,基于选择的熵编码模型对量化后的第二潜在变量进行熵编码,得到第二潜在变量的熵编码结果,将第二潜在变量的熵编码结果写入码流。然后, 将第一变量调整因子的编码结果写入码流。
请参考图17,图17是本申请实施例提供的一种示例性解码方法的框图。图17主要是对图15所示的解码方法进行示例性解释。在图17中,以音频信号为例。通过熵解码模型对码流中的第四潜在变量的熵编码结果进行熵解码,得到经量化的第四潜在变量。对经量化的第四潜在变量进行去量化处理,得到重构的第四潜在变量。对码流中第二变量调整因子的编码结果进行解码,得到经量化的第二变量调整因子。对经量化的第二变量调整因子进行去量化处理,得到重构的第二变量调整因子。基于重构的第二变量调整因子对重构的第四潜在变量进行调整,得到重构的第三潜在变量。通过上下文解码神经网络模型对重构的第三变量进行处理,得到重构的第二熵编码模型参数。基于重构的第二熵编码模型参数,从参数可调的熵解码模型中选择对应的熵解码模型。按照选择的熵解码模型对码流中的第二潜在变量的熵编码结果进行熵解码,得到经量化的第二潜在变量。对经量化的第二潜在变量进行去量化处理,得到重构的第二潜在变量。对码流中第一变量调整因子的编码结果进行解码,得到重构的第一变量调整因子。基于重构的第一变量调整因子对重构的第二潜在变量进行调整,得到重构的第一潜在变量。通过第一解码神经网络模型对重构的第一潜在变量进行处理,得到重构的当前帧的频域信号。对重构的当前帧的频域信号进行IMDCT变换处理以及去加窗处理,得到重构的音频信号。
图18是本申请实施例提供的一种编码装置的结构示意图,该编码装置可以由软件、硬件或者两者的结合实现成为编码端设备的部分或者全部,该编码端设备可以为图1所示的源装置。参见图18,该装置包括:数据处理模块1801、调整因子确定模块1802、第一编码结果获取模块1803和第一编码结果写入模块1804。
数据处理模块1801,用于通过第一编码神经网络模型对待编码的媒体数据进行处理,以得到第一潜在变量,第一潜在变量用于指示待编码的媒体数据的特征。详细实现过程参考上述各个实施例中对应的内容,此处不再赘述。
调整因子确定模块1802,用于基于第一潜在变量确定第一变量调整因子,第一变量调整因子用于使得第二潜在变量的熵编码结果的编码比特数满足预设编码速率条件,第二潜在变量是通过第一变量调整因子对第一潜在变量调整后得到。详细实现过程参考上述各个实施例中对应的内容,此处不再赘述。
第一编码结果获取模块1803,用于获取第二潜在变量的熵编码结果。详细实现过程参考上述各个实施例中对应的内容,此处不再赘述。
第一编码结果写入模块1804,用于将第二潜在变量的熵编码结果以及第一变量调整因子的编码结果写入码流。详细实现过程参考上述各个实施例中对应的内容,此处不再赘述。
可选地,在使用固定码率对待编码的媒体数据进行编码的情况下,满足预设编码速率条件包括编码比特数小于或等于目标编码比特数;或者,满足预设编码速率条件包括编码比特数小于或等于目标编码比特数,且编码比特数与目标编码比特数的差值小于比特数阈值。
可选地,在使用可变码率对待编码的媒体数据进行编码的情况下,满足预设编码速率条件包括编码比特数与目标编码比特数的差值的绝对值小于比特数阈值。
可选地,调整因子确定模块1802包括:
比特数确定子模块,用于基于第一潜在变量确定初始编码比特数;
第一因子确定子模块,用于基于初始编码比特数和目标编码比特数,确定第一变量调整因子。
可选地,初始编码比特数为第一潜在变量的熵编码结果的编码比特数;或者
初始编码比特数为经过第一初始调整因子调整后的第一潜在变量的熵编码结果的编码比特数。
可选地,在初始编码比特数不等于目标编码比特数的情况下,第一因子确定子模块具体用于:
基于初始编码比特数和目标编码比特数,通过第一循环方式确定第一变量调整因子;
第一循环方式的第i次循环处理包括如下步骤:
确定第i次循环处理的调整因子,i为正整数;
基于第i次循环处理的调整因子对第一潜在变量进行调整,以得到第i次调整后的第一潜在变量;
确定第i次调整后的第一潜在变量的熵编码结果的编码比特数,以得到第i次的编码比特数;
若第i次的编码比特数满足继续调整条件,执行第一循环方式的第i+1次循环处理;
若第i次的编码比特数不满足继续调整条件,终止第一循环方式的执行,基于第i次循环处理的调整因子确定第一变量调整因子。
可选地,该装置还包括:
第二编码结果获取模块,用于获取第三潜在变量的熵编码结果,第三潜在变量是基于第二潜在变量通过上下文模型确定得到;
第二编码结果写入模块,用于将第三潜在变量的熵编码结果写入码流;
其中,第二潜在变量的熵编码结果和第三潜在变量的熵编码结果的编码总比特数满足预设编码速率条件。
可选地,调整因子确定模块1802包括:
比特数确定子模块,用于基于第一潜在变量确定初始编码比特数;
第一因子确定子模块,用于基于初始编码比特数和目标编码比特数,确定第一变量调整因子。
可选地,比特数确定子模块具体用于:
基于第一潜在变量,通过上下文模型确定对应的上下文初始编码比特数和初始熵编码模型参数;
基于初始熵编码模型参数,确定第一潜在变量的熵编码结果的编码比特数,以得到基础初始编码比特数;
基于上下文初始编码比特数与基础初始编码比特数确定初始编码比特数。
可选地,在初始编码比特数不等于目标编码比特数的情况下,第一因子确定子模块具体用于:
基于初始编码比特数和目标编码比特数,通过第一循环方式确定第一变量调整因子;
第一循环方式的第i次循环处理包括如下步骤:
确定第i次循环处理的调整因子,i为正整数;
基于第i次循环处理的调整因子对第一潜在变量进行调整,得到第i次调整后的第一潜在变量;
基于第i次调整后的第一潜在变量,通过上下文模型确定第i次的上下文编码比特数和第i次的熵编码模型参数;
基于第i次的熵编码模型参数,确定第i次调整后的第一潜在变量的熵编码结果的编码比特数,以得到第i次的基础编码比特数;
基于第i次的上下文编码比特数与第i次的基础编码比特数确定第i次的编码比特数;
若第i次的编码比特数满足继续调整条件,执行第一循环方式的第i+1次循环处理;
若第i次的编码比特数不满足继续调整条件,终止第一循环方式的执行,基于第i次循环处理的调整因子确定第一变量调整因子。
可选地,第一因子确定子模块具体用于:
基于第一循环方式的第i-1次循环处理的调整因子、第i-1次的编码比特数和目标编码比特数,确定第i次循环处理的调整因子;
其中,在i=1的情况下,第i-1次循环处理的调整因子为第一初始调整因子,第i-1次的编码比特数为初始编码比特数;
继续调整条件包括第i-1次的编码比特数和第i次的编码比特数均小于目标编码比特数,或者,继续调整条件包括第i-1次的编码比特数和第i次的编码比特数均大于目标编码比特数。
可选地,在初始编码比特数小于目标编码比特数的情况下,第一因子确定子模块具体用于:
按照第一步长调整第一循环方式的第i-1次循环处理的调整因子,得到第i次循环处理的调整因子;
其中,在i=1的情况下,第i-1次循环处理的调整因子为第一初始调整因子;
继续调整条件包括第i次的编码比特数小于目标编码比特数。
可选地,在初始编码比特数大于目标编码比特数的情况下,第一因子确定子模块具体用于:
按照第二步长调整第一循环方式的第i-1次循环处理的调整因子,得到第i次循环处理的调整因子;
其中,在i=1的情况下,第i-1次循环处理的调整因子为第一初始调整因子;
继续调整条件包括第i次的编码比特数大于目标编码比特数。
可选地,第一因子确定子模块具体用于:
在初始编码比特数小于目标编码比特数的情况下,将第一初始调整因子确定为第一变量调整因子。
可选地,第一因子确定子模块具体用于:
在第i次的编码比特数等于目标编码比特数的情况下,将第i次循环处理的调整因子确定为第一变量调整因子;或者,
在第i次的编码比特数不等于目标编码比特数的情况下,基于第i次循环处理的调整因子和第一循环方式的第i-1次循环处理的调整因子,确定第一变量调整因子。
可选地,第一因子确定子模块具体用于:
确定第i次循环处理的调整因子和第i-1次循环处理的调整因子的平均值;
基于该平均值确定第一变量调整因子。
可选地,第一因子确定子模块具体用于:
基于第i次循环处理的调整因子和第i-1次循环处理的调整因子,通过第二循环方式确定第一变量调整因子;
第二循环方式的第j次循环处理包括如下步骤:
基于第j次循环处理的第一调整因子和第j次循环处理的第二调整因子确定第j次循环处理的第三调整因子,其中,在j等于1的情况下,第j次循环处理的第一调整因子为第i次循环处理的调整因子和第i-1次循环处理的调整因子中的一者,第j次循环处理的第二调整因子为第i次循环处理的调整因子和第i-1次循环处理的调整因子中的另一者,第j次循环处理的第一调整因子对应第j次的第一编码比特数,第j次循环处理的第二调整因子对应第j次的第二编码比特数,第j次的第一编码比特数小于第j次的第二编码比特数,j为正整数;
获取第j次的第三编码比特数,第j次的第三编码比特数是指经过第j次循环处理的第三调整因子调整后的第一潜在变量的熵编码结果的编码比特数;
若第j次的第三编码比特数不满足继续循环条件,终止第二循环方式的执行,将第j次循环处理的第三调整因子确定为第一变量调整因子;
若第j次的第三编码比特数满足继续循环条件、第j次的第三编码比特数大于目标编码比特数且小于第j次的第二编码比特数,将第j次循环处理的第三调整因子作为第j+1次循环处理的第二调整因子,将第j次循环处理的第一调整因子作为第j+1次循环处理的第一调整因子,执行第二循环方式的第j+1次循环处理;
若第j次的第三编码比特数满足继续循环条件、第j次的第三编码比特数小于目标编码比特数且大于第j次的第一编码比特数,将第j次循环处理的第三调整因子作为第j+1次循环处理的第一调整因子,将第j次循环处理的第二调整因子作为第j+1次循环处理的第二调整因子,执行第二循环方式的第j+1次循环处理。
可选地,第一因子确定子模块具体用于:
基于第i次循环处理的调整因子和第i-1次循环处理的调整因子,通过第二循环方式确定第一变量调整因子;
第二循环方式的第j次循环处理包括如下步骤:
基于第j次循环处理的第一调整因子和第j次循环处理的第二调整因子确定第j次循环处理的第三调整因子,其中,在j等于1的情况下,第j次循环处理的第一调整因子为第i次循环处理的调整因子和第i-1次循环处理的调整因子中的一者,第j次循环处理的第二调整因子为第i次循环处理的调整因子和第i-1次循环处理的调整因子中的另一者,第j次循环处理的第一调整因子对应第j次的第一编码比特数,第j次循环处理的第二调整因子对应第j次的第二编码比特数,第j次的第一编码比特数小于第j次的第二编码比特数,j为正整数;
获取第j次的第三编码比特数,第j次的第三编码比特数是指经过第j次循环处理的第三调整因子调整后的第一潜在变量的熵编码结果的编码比特数;
若第j次的第三编码比特数不满足继续循环条件,终止第二循环方式的执行,将第j次循环处理的第三调整因子确定为第一变量调整因子;
若j达到最大循环次数且第j次的第三编码比特数满足继续循环条件,终止第二循环方 式的执行,基于第j次循环处理的第一调整因子确定第一变量调整因子;
若j未达到最大循环次数、第j次的第三编码比特数满足继续循环条件、第j次的第三编码比特数大于目标编码比特数且小于第j次的第二编码比特数,将第j次循环处理的第三调整因子作为第j+1次循环处理的第二调整因子,将第j次循环处理的第一调整因子作为第j+1次循环处理的第一调整因子,执行第二循环方式的第j+1次循环处理;
若j未达到最大循环次数、第j次的第三编码比特数满足继续循环条件、第j次的第三编码比特数小于目标编码比特数且大于第j次的第一编码比特数,将第j次循环处理的第三调整因子作为第j+1次循环处理的第一调整因子,将第j次循环处理的第二调整因子作为第j+1次循环处理的第二调整因子,执行第二循环方式的第j+1次循环处理。
可选地,第一因子确定子模块具体用于:
在使用固定码率对待编码的媒体数据进行编码的情况下,将第j次循环处理的第一调整因子确定为第一变量调整因子。
可选地,第一因子确定子模块具体用于:
在使用可变码率对待编码的媒体数据进行编码的情况下,确定目标编码比特数与第j次的第一编码比特数之间的第一差值,以及确定第j次的第二编码比特数与目标编码比特数之间的第二差值;
若第一差值小于第二差值,将第j次循环处理的第一调整因子确定为第一变量调整因子;
若第二差值小于第一差值,将第j次循环处理的第二调整因子确定为第一变量调整因子;
若第一差值等于第二差值,将第j次循环处理的第一调整因子确定为第一变量调整因子,或者,将第j次循环处理的第二调整因子确定为第一变量调整因子。
可选地,在使用固定码率对待编码的媒体数据进行编码的情况下,继续循环条件包括第j次的第三编码比特数大于目标编码比特数,或者,继续循环条件包括第j次的第三编码比特数小于目标编码比特数,且目标编码比特数与第j次的第三编码比特数的差值大于比特数阈值。
可选地,在使用可变码率对待编码的媒体数据进行编码的情况下,继续循环条件包括目标编码比特数与第j次的第三编码比特数的差值的绝对值大于比特数阈值。
可选地,第一因子确定子模块具体用于:
在初始编码比特数等于目标编码比特数的情况下,将第一初始调整因子确定为第一变量调整因子。
可选地,该装置还包括:
调整因子确定模块1802,还用于基于第一潜在变量确定第二变量调整因子;
第三编码结果获取模块,用于获取第四潜在变量的熵编码结果,第四潜在变量是基于第二变量调整因子对第三潜在变量调整后得到,第三潜在变量是基于第二潜在变量通过上下文模型确定得到;
第三编码结果写入模块,用于将第四潜在变量的熵编码结果以及第二变量调整因子的编码结果写入码流;
其中,第二潜在变量的熵编码结果和第四潜在变量的熵编码结果的编码总比特数满足预设编码速率条件。
可选地,调整因子确定模块1802包括:
第一确定子模块,用于基于第一潜在变量,通过上下文模型确定对应的上下文初始编码比特数和初始熵编码模型参数;
第二确定子模块,用于基于初始熵编码模型参数,确定第一潜在变量的熵编码结果的编码比特数,以得到基础初始编码比特数;
第二因子确定子模块,用于基于上下文初始编码比特数、基础初始编码比特数和目标编码比特数,确定第一变量调整因子和第二变量调整因子。
可选地,第二因子确定子模块具体用于:
基于基础初始编码比特数和上下文初始编码比特数中的至少一者,以及目标编码比特数,确定基础目标编码比特数;
基于第二初始调整因子、基础目标编码比特数和基础初始编码比特数,确定第一变量调整因子和基础实际编码比特数,基础实际编码比特数是指经过第一变量调整因子调整后的第一潜在变量的熵编码结果的编码比特数;
基于目标编码比特数和基础实际编码比特数,确定上下文目标编码比特数;
基于上下文目标编码比特数和上下文初始编码比特数,确定第二变量调整因子。
可选地,第二因子确定子模块具体用于:
将目标编码比特数划分为基础目标编码比特数和上下文目标编码比特数;
基于基础目标编码比特数和基础初始编码比特数,确定第一变量调整因子;
基于上下文目标编码比特数和上下文初始编码比特数,确定第二变量调整因子。
可选地,媒体数据为音频信号、视频信号或者图像。
由于第二潜在变量的熵编码结果的编码比特数满足预设编码速率条件,这样可以保证每帧媒体数据对应的潜在变量的熵编码结果的编码比特数均能够满足预设编码速率条件,也即是,可以保证每帧媒体数据对应的潜在变量的熵编码结果的编码比特数基本保持一致,而不是动态变化的,从而满足了编码器对稳定编码速率的需求。进一步地,在考虑到需要传输边信息(例如窗型,时域噪声整形(TNS:Temporal Noise Shaping)参数,频域噪声整形(FDNS:Frequency-domain noise shaping)参数,和/或带宽扩展(BWE:bandwidth extension)参数等等)时,能够保证每帧媒体数据对应的潜在变量的熵编码结果的编码比特数和边信息的编码比特数整体基本保持一致,从而满足编码器对稳定编码速率的需求。
需要说明的是:上述实施例提供的编码装置在进行编码时,仅以上述各功能模块的划分进行举例说明,实际应用中,可以基于需要而将上述功能分配由不同的功能模块完成,即将装置的内部结构划分成不同的功能模块,以完成以上描述的全部或者部分功能。另外,上述实施例提供的编码装置与编码方法实施例属于同一构思,其具体实现过程详见方法实施例,这里不再赘述。
图19是本申请实施例提供的一种解码装置的结构示意图,该解码装置可以由软件、硬件或者两者的结合实现成为解码端设备的部分或者全部,该解码端设备可以为图1所示的目的地装置。参见图19,该装置包括:第一确定模块1901、变量调整模块1902和变量处理模块1903。
第一确定模块1901,用于基于码流确定重构的第二潜在变量和重构的第一变量调整因子。详细实现过程参考上述各个实施例中对应的内容,此处不再赘述。
变量调整模块1902,用于基于重构的第一变量调整因子,对重构的第二潜在变量进行调整,得到重构的第一潜在变量,重构的第一潜在变量用于指示待解码的媒体数据的特征。详细实现过程参考上述各个实施例中对应的内容,此处不再赘述。
变量处理模块1903,用于通过第一解码神经网络模型对重构的第一潜在变量进行处理,以得到重构的媒体数据。详细实现过程参考上述各个实施例中对应的内容,此处不再赘述。
可选地,第一确定模块1901包括:
第一确定子模块,用于基于码流确定重构的第三潜在变量;
第二确定子模块,用于基于码流和重构的第三潜在变量,确定重构的第二潜在变量。
可选地,第二确定子模块具体用于:
通过上下文解码神经网络模型对重构的第三潜在变量进行处理,以得到重构的第一熵编码模型参数;
基于码流和重构的第一熵编码模型参数,确定重构的第二潜在变量。
可选地,第一确定模块1901包括:
第三确定子模块,用于基于码流确定重构的第四潜在变量和重构的第二变量调整因子;
第四确定子模块,用于基于码流、重构的第四潜在变量和重构的第二变量调整因子,确定重构的第二潜在变量。
可选地,第四确定子模块具体用于:
基于重构的第二变量调整因子,对重构的第四潜在变量进行调整,得到重构的第三潜在变量;
通过上下文解码神经网络模型对重构的第三潜在变量进行处理,以得到重构的第二熵编码模型参数;
基于码流和重构的第二熵编码模型参数,确定重构的第二潜在变量。
可选地,媒体数据为音频信号、视频信号或者图像。
在本申请实施例中,由于第二潜在变量的熵编码结果的编码比特数满足预设编码速率条件,这样可以保证每帧媒体数据对应的潜在变量的熵编码结果的编码比特数均能够满足预设编码速率条件,也即是,可以保证每帧媒体数据对应的潜在变量的熵编码结果的编码比特数基本保持一致,而不是动态变化的,从而满足了编码器对稳定编码速率的需求。进一步地,在考虑到需要传输边信息(例如窗型,时域噪声整形(TNS:Temporal Noise Shaping)参数,频域噪声整形(FDNS:Frequency-domain noise shaping)参数,和/或带宽扩展(BWE:bandwidth extension)参数等等)时,能够保证每帧媒体数据对应的潜在变量的熵编码结果的编码比特数和边信息的编码比特数整体基本保持一致,从而满足编码器对稳定编码速率的需求。
需要说明的是:上述实施例提供的解码装置在进行解码时,仅以上述各功能模块的划分进行举例说明,实际应用中,可以基于需要而将上述功能分配由不同的功能模块完成,即将装置的内部结构划分成不同的功能模块,以完成以上描述的全部或者部分功能。另外,上述实施例提供的解码装置与解码方法实施例属于同一构思,其具体实现过程详见方法实施例,这里不再赘述。
图20为用于本申请实施例的一种编解码装置2000的示意性框图。其中,编解码装置2000可以包括处理器2001、存储器2002和总线系统2003。其中,处理器2001和存储器2002通 过总线系统2003相连,该存储器2002用于存储指令,该处理器2001用于执行该存储器2002存储的指令,以执行本申请实施例描述的各种的编码或解码方法。为避免重复,这里不再详细描述。
在本申请实施例中,该处理器2001可以是中央处理单元(central processing unit,CPU),该处理器2001还可以是其他通用处理器、DSP、ASIC、FPGA或者其他可编程逻辑器件、分立门或者晶体管逻辑器件、分立硬件组件等。通用处理器可以是微处理器或者该处理器也可以是任何常规的处理器等。
该存储器2002可以包括ROM设备或者RAM设备。任何其他适宜类型的存储设备也可以用作存储器2002。存储器2002可以包括由处理器2001使用总线2003访问的代码和数据20021。存储器2002可以进一步包括操作系统20023和应用程序20022,该应用程序20022包括允许处理器2001执行本申请实施例描述的的编码或解码方法的至少一个程序。例如,应用程序20022可以包括应用1至N,其进一步包括执行在本申请实施例描述的的编码或解码方法的编码或解码应用(简称编解码应用)。
该总线系统2003除包括数据总线之外,还可以包括电源总线、控制总线和状态信号总线等。但是为了清楚说明起见,在图中将各种总线都标为总线系统2003。
可选地,编解码装置2000还可以包括一个或多个输出设备,诸如显示器2004。在一个示例中,显示器2004可以是触感显示器,其将显示器与可操作地感测触摸输入的触感单元合并。显示器2004可以经由总线2003连接到处理器2001。
需要指出的是,编解码装置2000可以执行本申请实施例中的的编码方法,也可执行本申请实施例中的的解码方法。
本领域技术人员能够领会,结合本文公开描述的各种说明性逻辑框、模块和算法步骤所描述的功能可以硬件、软件、固件或其任何组合来实施。如果以软件来实施,那么各种说明性逻辑框、模块、和步骤描述的功能可作为一或多个指令或代码在计算机可读媒体上存储或传输,且由基于硬件的处理单元执行。计算机可读媒体可包含计算机可读存储媒体,其对应于有形媒体,例如数据存储媒体,或包括任何促进将计算机程序从一处传送到另一处的媒体(例如,基于通信协议)的通信媒体。以此方式,计算机可读媒体大体上可对应于(1)非暂时性的有形计算机可读存储媒体,或(2)通信媒体,例如信号或载波。数据存储媒体可为可由一或多个计算机或一或多个处理器存取以检索用于实施本申请中描述的技术的指令、代码和/或数据结构的任何可用媒体。计算机程序产品可包含计算机可读媒体。
作为实例而非限制,此类计算机可读存储媒体可包括RAM、ROM、EEPROM、CD-ROM或其它光盘存储装置、磁盘存储装置或其它磁性存储装置、快闪存储器或可用来存储指令或数据结构的形式的所要程序代码并且可由计算机存取的任何其它媒体。并且,任何连接被恰当地称作计算机可读媒体。举例来说,如果使用同轴缆线、光纤缆线、双绞线、数字订户线(DSL)或例如红外线、无线电和微波等无线技术从网站、服务器或其它远程源传输指令,那么同轴缆线、光纤缆线、双绞线、DSL或例如红外线、无线电和微波等无线技术包含在媒体的定义中。但是,应理解,所述计算机可读存储媒体和数据存储媒体并不包括连接、载波、信号或其它暂时媒体,而是实际上针对于非暂时性有形存储媒体。如本文中所使用,磁盘和光盘包含压缩光盘(CD)、激光光盘、光学光盘、DVD和蓝光光盘,其中磁盘通常以磁性方式再现数据,而光盘利用激光以光学方式再现数据。以上各项的组合也应包含在计算机可读媒体 的范围内。
可通过例如一或多个数字信号处理器(DSP)、通用微处理器、专用集成电路(ASIC)、现场可编程逻辑阵列(FPGA)或其它等效集成或离散逻辑电路等一或多个处理器来执行指令。因此,如本文中所使用的术语“处理器”可指前述结构或适合于实施本文中所描述的技术的任一其它结构中的任一者。另外,在一些方面中,本文中所描述的各种说明性逻辑框、模块、和步骤所描述的功能可以提供于经配置以用于编码和解码的专用硬件和/或软件模块内,或者并入在组合编解码器中。而且,所述技术可完全实施于一或多个电路或逻辑元件中。在一种示例下,编码器100及解码器200中的各种说明性逻辑框、单元、模块可以理解为对应的电路器件或逻辑元件。
本申请实施例的技术可在各种各样的装置或设备中实施,包含无线手持机、集成电路(IC)或一组IC(例如,芯片组)。本申请实施例中描述各种组件、模块或单元是为了强调用于执行所揭示的技术的装置的功能方面,但未必需要由不同硬件单元实现。实际上,如上文所描述,各种单元可结合合适的软件和/或固件组合在编码解码器硬件单元中,或者通过互操作硬件单元(包含如上文所描述的一或多个处理器)来提供。
也就是说,在上述实施例中,可以全部或部分地通过软件、硬件、固件或者其任意结合来实现。当使用软件实现时,可以全部或部分地以计算机程序产品的形式实现。所述计算机程序产品包括一个或多个计算机指令。在计算机上加载和执行所述计算机指令时,全部或部分地产生按照本申请实施例所述的流程或功能。所述计算机可以是通用计算机、专用计算机、计算机网络或其他可编程装置。所述计算机指令可以存储在计算机可读存储介质中,或者从一个计算机可读存储介质向另一个计算机可读存储介质传输,例如,所述计算机指令可以从一个网站站点、计算机、服务器或数据中心通过有线(例如:同轴电缆、光纤、数据用户线(digital subscriber line,DSL))或无线(例如:红外、无线、微波等)方式向另一个网站站点、计算机、服务器或数据中心进行传输。所述计算机可读存储介质可以是计算机能够存取的任何可用介质,或者是包含一个或多个可用介质集成的服务器、数据中心等数据存储设备。所述可用介质可以是磁性介质(例如:软盘、硬盘、磁带)、光介质(例如:数字通用光盘(digital versatile disc,DVD))或半导体介质(例如:固态硬盘(solid state disk,SSD))等。值得注意的是,本申请实施例提到的计算机可读存储介质可以为非易失性存储介质,换句话说,可以是非瞬时性存储介质。
应当理解的是,本文提及的“多个”是指两个或两个以上。在本申请实施例的描述中,除非另有说明,“/”表示或的意思,例如,A/B可以表示A或B;本文中的“和/或”仅仅是一种描述关联对象的关联关系,表示可以存在三种关系,例如,A和/或B,可以表示:单独存在A,同时存在A和B,单独存在B这三种情况。另外,为了便于清楚描述本申请实施例的技术方案,在本申请实施例中,采用了“第一”、“第二”等字样对功能和作用基本相同的相同项或相似项进行区分。本领域技术人员可以理解“第一”、“第二”等字样并不对数量和执行次序进行限定,并且“第一”、“第二”等字样也并不限定一定不同。
以上所述为本申请提供的实施例,并不用以限制本申请,凡在本申请的精神和原则之内,所作的任何修改、等同替换、改进等,均应包含在本申请的保护范围之内。

Claims (73)

  1. 一种编码方法,其特征在于,所述方法包括:
    通过第一编码神经网络模型对待编码的媒体数据进行处理,以得到第一潜在变量,所述第一潜在变量用于指示所述待编码的媒体数据的特征;
    基于所述第一潜在变量确定第一变量调整因子,所述第一变量调整因子用于使得第二潜在变量的熵编码结果的编码比特数满足预设编码速率条件,所述第二潜在变量是通过所述第一变量调整因子对所述第一潜在变量调整后得到;
    获取所述第二潜在变量的熵编码结果;
    将所述第二潜在变量的熵编码结果以及所述第一变量调整因子的编码结果写入码流。
  2. 如权利要求1所述的方法,其特征在于,
    在使用固定码率对所述待编码的媒体数据进行编码的情况下,所述满足预设编码速率条件包括所述编码比特数小于或等于目标编码比特数;或者,所述满足预设编码速率条件包括所述编码比特数小于或等于所述目标编码比特数,且所述编码比特数与所述目标编码比特数的差值小于比特数阈值。
  3. 如权利要求1所述的方法,其特征在于,
    在使用可变码率对所述待编码的媒体数据进行编码的情况下,所述满足预设编码速率条件包括所述编码比特数与所述目标编码比特数的差值的绝对值小于比特数阈值。
  4. 如权利要求1所述的方法,其特征在于,所述基于所述第一潜在变量确定第一变量调整因子,包括:
    基于所述第一潜在变量确定初始编码比特数;
    基于所述初始编码比特数和目标编码比特数,确定所述第一变量调整因子。
  5. 如权利要求4所述的方法,其特征在于,所述初始编码比特数为所述第一潜在变量的熵编码结果的编码比特数;或者
    所述初始编码比特数为经过第一初始调整因子调整后的第一潜在变量的熵编码结果的编码比特数。
  6. 如权利要求4所述的方法,其特征在于,在所述初始编码比特数不等于所述目标编码比特数的情况下,所述基于所述初始编码比特数和目标编码比特数,确定所述第一变量调整因子,包括:
    基于所述初始编码比特数和所述目标编码比特数,通过第一循环方式确定所述第一变量调整因子;
    所述第一循环方式的第i次循环处理包括如下步骤:
    确定所述第i次循环处理的调整因子,i为正整数;
    基于所述第i次循环处理的调整因子对所述第一潜在变量进行调整,以得到第i次调整后的第一潜在变量;
    确定所述第i次调整后的第一潜在变量的熵编码结果的编码比特数,以得到第i次的编码比特数;
    若所述第i次的编码比特数满足继续调整条件,执行所述第一循环方式的第i+1次循环处理;
    若所述第i次的编码比特数不满足所述继续调整条件,终止所述第一循环方式的执行,基于所述第i次循环处理的调整因子确定所述第一变量调整因子。
  7. 如权利要求1所述的方法,其特征在于,所述方法还包括:
    获取第三潜在变量的熵编码结果,所述第三潜在变量是基于所述第二潜在变量通过上下文模型确定得到;
    将所述第三潜在变量的熵编码结果写入所述码流;
    其中,所述第二潜在变量的熵编码结果和所述第三潜在变量的熵编码结果的编码总比特数满足所述预设编码速率条件。
  8. 如权利要求7所述的方法,其特征在于,所述基于所述第一潜在变量确定第一变量调整因子,包括:
    基于所述第一潜在变量确定初始编码比特数;
    基于所述初始编码比特数和目标编码比特数,确定所述第一变量调整因子。
  9. 如权利要求8所述的方法,其特征在于,所述基于所述第一潜在变量确定初始编码比特数,包括:
    基于所述第一潜在变量,通过所述上下文模型确定对应的上下文初始编码比特数和初始熵编码模型参数;
    基于所述初始熵编码模型参数,确定所述第一潜在变量的熵编码结果的编码比特数,以得到基础初始编码比特数;
    基于所述上下文初始编码比特数与所述基础初始编码比特数确定所述初始编码比特数。
  10. 如权利要求8所述的方法,其特征在于,在所述初始编码比特数不等于所述目标编码比特数的情况下,所述基于所述初始编码比特数和目标编码比特数,确定所述第一变量调整因子,包括:
    基于所述初始编码比特数和所述目标编码比特数,通过第一循环方式确定所述第一变量调整因子;
    所述第一循环方式的第i次循环处理包括如下步骤:
    确定所述第i次循环处理的调整因子,i为正整数;
    基于所述第i次循环处理的调整因子对所述第一潜在变量进行调整,以得到第i次调整后的第一潜在变量;
    基于所述第i次调整后的第一潜在变量,通过所述上下文模型确定第i次的上下文编码 比特数和第i次的熵编码模型参数;
    基于所述第i次的熵编码模型参数,确定所述第i次调整后的第一潜在变量的熵编码结果的编码比特数,以得到第i次的基础编码比特数;
    基于所述第i次的上下文编码比特数与所述第i次的基础编码比特数确定第i次的编码比特数;
    若所述第i次的编码比特数满足继续调整条件,执行所述第一循环方式的第i+1次循环处理;
    若所述第i次的编码比特数不满足所述继续调整条件,终止所述第一循环方式的执行,基于所述第i次循环处理的调整因子确定所述第一变量调整因子。
  11. 如权利要求6或10所述的方法,其特征在于,所述确定所述第i次循环处理的调整因子,包括:
    基于所述第一循环方式的第i-1次循环处理的调整因子、第i-1次的编码比特数和所述目标编码比特数,确定所述第i次循环处理的调整因子;
    其中,在i=1的情况下,所述第i-1次循环处理的调整因子为第一初始调整因子,所述第i-1次的编码比特数为所述初始编码比特数;
    所述继续调整条件包括所述第i-1次的编码比特数和所述第i次的编码比特数均小于所述目标编码比特数,或者,所述继续调整条件包括所述第i-1次的编码比特数和所述第i次的编码比特数均大于所述目标编码比特数。
  12. 如权利要求6或10所述的方法,其特征在于,在所述初始编码比特数小于所述目标编码比特数的情况下,所述确定所述第i次循环处理的调整因子,包括:
    按照第一步长调整所述第一循环方式的第i-1次循环处理的调整因子,以得到所述第i次循环处理的调整因子;
    其中,在i=1的情况下,所述第i-1次循环处理的调整因子为第一初始调整因子;
    所述继续调整条件包括所述第i次的编码比特数小于所述目标编码比特数。
  13. 如权利要求6或10所述的方法,其特征在于,在所述初始编码比特数大于所述目标编码比特数的情况下,所述确定所述第i次循环处理的调整因子,包括:
    按照第二步长调整所述第一循环方式的第i-1次循环处理的调整因子,以得到所述第i次循环处理的调整因子;
    其中,在i=1的情况下,所述第i-1次循环处理的调整因子为第一初始调整因子;
    所述继续调整条件包括所述第i次的编码比特数大于所述目标编码比特数。
  14. 如权利要求4或8所述的方法,其特征在于,所述基于所述初始编码比特数和目标编码比特数,确定所述第一变量调整因子包括:
    在所述初始编码比特数小于所述目标编码比特数的情况下,将第一初始调整因子确定为所述第一变量调整因子。
  15. 如权利要求6或10所述的方法,其特征在于,所述基于所述第i次循环处理的调整因子确定所述第一变量调整因子,包括:
    在所述第i次的编码比特数等于所述目标编码比特数的情况下,将所述第i次循环处理的调整因子确定为所述第一变量调整因子;或者,
    在所述第i次的编码比特数不等于所述目标编码比特数的情况下,基于所述第i次循环处理的调整因子和所述第一循环方式的第i-1次循环处理的调整因子,确定所述第一变量调整因子。
  16. 如权利要求15所述的方法,其特征在于,所述基于所述第i次循环处理的调整因子和所述第一循环方式的第i-1次循环处理的调整因子,确定所述第一变量调整因子,包括:
    确定所述第i次循环处理的调整因子和所述第i-1次循环处理的调整因子的平均值;
    基于所述平均值确定所述第一变量调整因子。
  17. 如权利要求15所述的方法,其特征在于,所述基于所述第i次循环处理的调整因子和所述第一循环方式的第i-1次循环处理的调整因子,确定所述第一变量调整因子,包括:
    基于所述第i次循环处理的调整因子和所述第i-1次循环处理的调整因子,通过第二循环方式确定所述第一变量调整因子;
    所述第二循环方式的第j次循环处理包括如下步骤:
    基于所述第j次循环处理的第一调整因子和所述第j次循环处理的第二调整因子确定所述第j次循环处理的第三调整因子;其中,在j等于1的情况下,所述第j次循环处理的第一调整因子为所述第i次循环处理的调整因子和所述第i-1次循环处理的调整因子中的一者,所述第j次循环处理的第二调整因子为所述第i次循环处理的调整因子和所述第i-1次循环处理的调整因子中的另一者,所述第j次循环处理的第一调整因子对应第j次的第一编码比特数,所述第j次循环处理的第二调整因子对应第j次的第二编码比特数,所述第j次的第一编码比特数小于所述第j次的第二编码比特数,j为正整数;
    获取第j次的第三编码比特数,所述第j次的第三编码比特数是指经过所述第j次循环处理的第三调整因子调整后的第一潜在变量的熵编码结果的编码比特数;
    若所述第j次的第三编码比特数不满足继续循环条件,终止所述第二循环方式的执行,将所述第j次循环处理的第三调整因子确定为所述第一变量调整因子;
    若所述第j次的第三编码比特数满足所述继续循环条件、所述第j次的第三编码比特数大于所述目标编码比特数且小于所述第j次的第二编码比特数,将所述第j次循环处理的第三调整因子作为第j+1次循环处理的第二调整因子,将所述第j次循环处理的第一调整因子作为第j+1次循环处理的第一调整因子,执行所述第二循环方式的第j+1次循环处理;若所述第j次的第三编码比特数满足所述继续循环条件、所述第j次的第三编码比特数小于所述目标编码比特数且大于所述第j次的第一编码比特数,将所述第j次循环处理的第三调整因子作为第j+1次循环处理的第一调整因子,将所述第j次循环处理的第二调整因子作为第j+1次循环处理的第二调整因子,执行所述第二循环方式的第j+1次循环处理。
  18. 如权利要求15所述的方法,其特征在于,所述基于所述第i次循环处理的调整因子和 所述第一循环方式的第i-1次循环处理的调整因子,确定所述第一变量调整因子,包括:
    基于所述第i次循环处理的调整因子和所述第i-1次循环处理的调整因子,通过第二循环方式确定所述第一变量调整因子;
    所述第二循环方式的第j次循环处理包括如下步骤:
    基于所述第j次循环处理的第一调整因子和所述第j次循环处理的第二调整因子确定所述第j次循环处理的第三调整因子;其中,在j等于1的情况下,所述第j次循环处理的第一调整因子为所述第i次循环处理的调整因子和所述第i-1次循环处理的调整因子中的一者,所述第j次循环处理的第二调整因子为所述第i次循环处理的调整因子和所述第i-1次循环处理的调整因子中的另一者,所述第j次循环处理的第一调整因子对应第j次的第一编码比特数,所述第j次循环处理的第二调整因子对应第j次的第二编码比特数,所述第j次的第一编码比特数小于所述第j次的第二编码比特数,j为正整数;
    获取第j次的第三编码比特数,所述第j次的第三编码比特数是指经过所述第j次循环处理的第三调整因子调整后的第一潜在变量的熵编码结果的编码比特数;
    若所述第j次的第三编码比特数不满足继续循环条件,终止所述第二循环方式的执行,将所述第j次循环处理的第三调整因子确定为所述第一变量调整因子;
    若所述j达到最大循环次数且所述第j次的第三编码比特数满足所述继续循环条件,终止所述第二循环方式的执行,基于所述第j次循环处理的第一调整因子确定所述第一变量调整因子;
    若所述j未达到所述最大循环次数、所述第j次的第三编码比特数满足所述继续循环条件、所述第j次的第三编码比特数大于所述目标编码比特数且小于所述第j次的第二编码比特数,将所述第j次循环处理的第三调整因子作为第j+1次循环处理的第二调整因子,将所述第j次循环处理的第一调整因子作为第j+1次循环处理的第一调整因子,执行所述第二循环方式的第j+1次循环处理;
    若所述j未达到所述最大循环次数、所述第j次的第三编码比特数满足所述继续循环条件、所述第j次的第三编码比特数小于所述目标编码比特数且大于所述第j次的第一编码比特数,将所述第j次循环处理的第三调整因子作为第j+1次循环处理的第一调整因子,将所述第j次循环处理的第二调整因子作为第j+1次循环处理的第二调整因子,执行所述第二循环方式的第j+1次循环处理。
  19. 如权利要求18所述的方法,其特征在于,所述基于所述第j次循环处理的第一调整因子确定所述第一变量调整因子,包括:
    在使用固定码率对所述待编码的媒体数据进行编码的情况下,将所述第j次循环处理的第一调整因子确定为所述第一变量调整因子。
  20. 如权利要求18所述的方法,其特征在于,所述基于所述第j次循环处理的第一调整因子确定所述第一变量调整因子,包括:
    在使用可变码率对所述待编码的媒体数据进行编码的情况下,确定所述目标编码比特数与所述第j次的第一编码比特数之间的第一差值,以及确定所述第j次的第二编码比特数与所述目标编码比特数之间的第二差值;
    若所述第一差值小于所述第二差值,将所述第j次循环处理的第一调整因子确定为所述第一变量调整因子;
    若所述第二差值小于所述第一差值,将所述第j次循环处理的第二调整因子确定为所述第一变量调整因子;
    若所述第一差值等于所述第二差值,将所述第j次循环处理的第一调整因子确定为所述第一变量调整因子,或者,将所述第j次循环处理的第二调整因子确定为所述第一变量调整因子。
  21. 如权利要求17或18所述的方法,其特征在于,
    在使用固定码率对所述待编码的媒体数据进行编码的情况下,所述继续循环条件包括所述第j次的第三编码比特数大于所述目标编码比特数,或者,所述继续循环条件包括所述第j次的第三编码比特数小于所述目标编码比特数,且所述目标编码比特数与所述第j次的第三编码比特数的差值大于比特数阈值。
  22. 如权利要求17或18所述的方法,其特征在于,
    在使用可变码率对所述待编码的媒体数据进行编码的情况下,所述继续循环条件包括所述目标编码比特数与所述第j次的第三编码比特数的差值的绝对值大于所述比特数阈值。
  23. 如权利要求4或8所述的方法,其特征在于,所述基于所述初始编码比特数和目标编码比特数,确定所述第一变量调整因子,包括:
    在所述初始编码比特数等于所述目标编码比特数的情况下,将第一初始调整因子确定为所述第一变量调整因子。
  24. 如权利要求1所述的方法,其特征在于,所述方法还包括:
    基于所述第一潜在变量确定第二变量调整因子;
    获取第四潜在变量的熵编码结果,所述第四潜在变量是基于所述第二变量调整因子对第三潜在变量调整后得到,所述第三潜在变量是基于所述第二潜在变量通过上下文模型确定得到;
    将所述第四潜在变量的熵编码结果以及所述第二变量调整因子的编码结果写入所述码流;
    其中,所述第二潜在变量的熵编码结果和所述第四潜在变量的熵编码结果的编码总比特数满足所述预设编码速率条件。
  25. 如权利要求24所述的方法,其特征在于,基于所述第一潜在变量确定第一变量调整因子和第二变量调整因子,包括:
    基于所述第一潜在变量,通过所述上下文模型确定对应的上下文初始编码比特数和初始熵编码模型参数;
    基于所述初始熵编码模型参数,确定所述第一潜在变量的熵编码结果的编码比特数,以得到基础初始编码比特数;
    基于所述上下文初始编码比特数、所述基础初始编码比特数和目标编码比特数,确定所 述第一变量调整因子和所述第二变量调整因子。
  26. 如权利要求25所述的方法,其特征在于,所述基于所述上下文初始编码比特数、所述基础初始编码比特数和目标编码比特数,确定所述第一变量调整因子和所述第二变量调整因子,包括:
    基于所述基础初始编码比特数和所述上下文初始编码比特数中的至少一者,以及所述目标编码比特数,确定基础目标编码比特数;
    基于第二初始调整因子、所述基础目标编码比特数和所述基础初始编码比特数,确定所述第一变量调整因子和基础实际编码比特数,所述基础实际编码比特数是指经过所述第一变量调整因子调整后的第一潜在变量的熵编码结果的编码比特数;
    基于所述目标编码比特数和所述基础实际编码比特数,确定上下文目标编码比特数;
    基于所述上下文目标编码比特数和所述上下文初始编码比特数,确定所述第二变量调整因子。
  27. 如权利要求25所述的方法,其特征在于,所述基于所述上下文初始编码比特数、所述基础初始编码比特数和目标编码比特数,确定所述第一变量调整因子和所述第二变量调整因子,包括:
    将所述目标编码比特数划分为基础目标编码比特数和上下文目标编码比特数;
    基于所述基础目标编码比特数和所述基础初始编码比特数,确定所述第一变量调整因子;
    基于所述上下文目标编码比特数和所述上下文初始编码比特数,确定所述第二变量调整因子。
  28. 如权利要求1-27任一所述的方法,其特征在于,所述媒体数据为音频信号、视频信号或者图像。
  29. 一种解码方法,其特征在于,所述方法包括:
    基于码流确定重构的第二潜在变量和重构的第一变量调整因子;
    基于所述重构的第一变量调整因子,对所述重构的第二潜在变量进行调整,得到重构的第一潜在变量,所述重构的第一潜在变量用于指示待解码的媒体数据的特征;
    通过第一解码神经网络模型对所述重构的第一潜在变量进行处理,以得到重构的媒体数据。
  30. 如权利要求29所述的方法,其特征在于,所述基于码流确定重构的第二潜在变量,包括:
    基于所述码流确定重构的第三潜在变量;
    基于所述码流和所述重构的第三潜在变量,确定所述重构的第二潜在变量。
  31. 如权利要求30所述的方法,其特征在于,所述基于所述码流和所述重构的第三潜在变量,确定所述重构的第二潜在变量,包括:
    通过上下文解码神经网络模型对所述重构的第三潜在变量进行处理,以得到重构的第一熵编码模型参数;
    基于所述码流和所述重构的第一熵编码模型参数,确定所述重构的第二潜在变量。
  32. 如权利要求29所述的方法,其特征在于,所述基于码流确定重构的第二潜在变量,包括:
    基于所述码流确定重构的第四潜在变量和重构的第二变量调整因子;
    基于所述码流、所述重构的第四潜在变量和所述重构的第二变量调整因子,确定所述重构的第二潜在变量。
  33. 如权利要求32所述的方法,其特征在于,所述基于所述码流、所述重构的第四潜在变量和所述重构的第二变量调整因子,确定所述重构的第二潜在变量,包括:
    基于所述重构的第二变量调整因子,对所述重构的第四潜在变量进行调整,得到重构的第三潜在变量;
    通过上下文解码神经网络模型对所述重构的第三潜在变量进行处理,以得到重构的第二熵编码模型参数;
    基于所述码流和所述重构的第二熵编码模型参数,确定所述重构的第二潜在变量。
  34. 如权利要求29-33任一所述的方法,其特征在于,所述媒体数据为音频信号、视频信号或者图像。
  35. 一种编码装置,其特征在于,所述装置包括:
    数据处理模块,用于通过第一编码神经网络模型对待编码的媒体数据进行处理,以得到第一潜在变量,所述第一潜在变量用于指示所述待编码的媒体数据的特征;
    调整因子确定模块,用于基于所述第一潜在变量确定第一变量调整因子,所述第一变量调整因子用于使得第二潜在变量的熵编码结果的编码比特数满足预设编码速率条件,所述第二潜在变量是通过所述第一变量调整因子对所述第一潜在变量调整后得到;
    第一编码结果获取模块,用于获取所述第二潜在变量的熵编码结果;
    第一编码结果写入模块,用于将所述第二潜在变量的熵编码结果以及所述第一变量调整因子的编码结果写入码流。
  36. 如权利要求35所述的装置,其特征在于,
    在使用固定码率对所述待编码的媒体数据进行编码的情况下,所述满足预设编码速率条件包括所述编码比特数小于或等于目标编码比特数;或者,所述满足预设编码速率条件包括所述编码比特数小于或等于所述目标编码比特数,且所述编码比特数与所述目标编码比特数的差值小于比特数阈值。
  37. 如权利要求35所述的装置,其特征在于,
    在使用可变码率对所述待编码的媒体数据进行编码的情况下,所述满足预设编码速率条 件包括所述编码比特数与所述目标编码比特数的差值的绝对值小于所述比特数阈值。
  38. 如权利要求35所述的装置,其特征在于,所述调整因子确定模块包括:
    比特数确定子模块,用于基于所述第一潜在变量确定初始编码比特数;
    第一因子确定子模块,用于基于所述初始编码比特数和目标编码比特数,确定所述第一变量调整因子。
  39. 如权利要求38所述的装置,其特征在于,所述初始编码比特数为所述第一潜在变量的熵编码结果的编码比特数;或者
    所述初始编码比特数为经过第一初始调整因子调整后的第一潜在变量的熵编码结果的编码比特数。
  40. 如权利要求38所述的装置,其特征在于,在所述初始编码比特数不等于所述目标编码比特数的情况下,所述第一因子确定子模块具体用于:
    基于所述初始编码比特数和所述目标编码比特数,通过第一循环方式确定所述第一变量调整因子;
    所述第一循环方式的第i次循环处理包括如下步骤:
    确定所述第i次循环处理的调整因子,i为正整数;
    基于所述第i次循环处理的调整因子对所述第一潜在变量进行调整,以得到第i次调整后的第一潜在变量;
    确定所述第i次调整后的第一潜在变量的熵编码结果的编码比特数,以得到第i次的编码比特数;
    若所述第i次的编码比特数满足继续调整条件,执行所述第一循环方式的第i+1次循环处理;
    若所述第i次的编码比特数不满足所述继续调整条件,终止所述第一循环方式的执行,基于所述第i次循环处理的调整因子确定所述第一变量调整因子。
  41. 如权利要求35所述的装置,其特征在于,所述装置还包括:
    第二编码结果获取模块,用于获取第三潜在变量的熵编码结果,所述第三潜在变量是基于所述第二潜在变量通过上下文模型确定得到;
    第二编码结果写入模块,用于将所述第三潜在变量的熵编码结果写入所述码流;
    其中,所述第二潜在变量的熵编码结果和所述第三潜在变量的熵编码结果的编码总比特数满足所述预设编码速率条件。
  42. 如权利要求41所述的装置,其特征在于,所述调整因子确定模块包括:
    比特数确定子模块,用于基于所述第一潜在变量确定初始编码比特数;
    第一因子确定子模块,用于基于所述初始编码比特数和目标编码比特数,确定所述第一变量调整因子。
  43. 如权利要求42所述的装置,其特征在于,所述比特数确定子模块具体用于:
    基于所述第一潜在变量,通过所述上下文模型确定对应的上下文初始编码比特数和初始熵编码模型参数;
    基于所述初始熵编码模型参数,确定所述第一潜在变量的熵编码结果的编码比特数,以得到基础初始编码比特数;
    基于所述上下文初始编码比特数与所述基础初始编码比特数确定所述初始编码比特数。
  44. 如权利要求42所述的装置,其特征在于,在所述初始编码比特数不等于所述目标编码比特数的情况下,所述第一因子确定子模块具体用于:
    基于所述初始编码比特数和所述目标编码比特数,通过第一循环方式确定所述第一变量调整因子;
    所述第一循环方式的第i次循环处理包括如下步骤:
    确定所述第i次循环处理的调整因子,i为正整数;
    基于所述第i次循环处理的调整因子对所述第一潜在变量进行调整,以得到第i次调整后的第一潜在变量;
    基于所述第i次调整后的第一潜在变量,通过所述上下文模型确定第i次的上下文编码比特数和第i次的熵编码模型参数;
    基于所述第i次的熵编码模型参数,确定所述第i次调整后的第一潜在变量的熵编码结果的编码比特数,以得到第i次的基础编码比特数;
    基于所述第i次的上下文编码比特数与所述第i次的基础编码比特数确定第i次的编码比特数;
    若所述第i次的编码比特数满足继续调整条件,执行所述第一循环方式的第i+1次循环处理;
    若所述第i次的编码比特数不满足所述继续调整条件,终止所述第一循环方式的执行,基于所述第i次循环处理的调整因子确定所述第一变量调整因子。
  45. 如权利要求40或44所述的装置,其特征在于,所述第一因子确定子模块具体用于:
    基于所述第一循环方式的第i-1次循环处理的调整因子、第i-1次的编码比特数和所述目标编码比特数,确定所述第i次循环处理的调整因子;
    其中,在i=1的情况下,所述第i-1次循环处理的调整因子为第一初始调整因子,所述第i-1次的编码比特数为所述初始编码比特数;
    所述继续调整条件包括所述第i-1次的编码比特数和所述第i次的编码比特数均小于所述目标编码比特数,或者,所述继续调整条件包括所述第i-1次的编码比特数和所述第i次的编码比特数均大于所述目标编码比特数。
  46. 如权利要求40或44所述的装置,其特征在于,在所述初始编码比特数小于所述目标编码比特数的情况下,所述第一因子确定子模块具体用于:
    按照第一步长调整所述第一循环方式的第i-1次循环处理的调整因子,以得到所述第i次循环处理的调整因子;
    其中,在i=1的情况下,所述第i-1次循环处理的调整因子为第一初始调整因子;
    所述继续调整条件包括所述第i次的编码比特数小于所述目标编码比特数。
  47. 如权利要求40或44所述的装置,其特征在于,在所述初始编码比特数大于所述目标编码比特数的情况下,所述第一因子确定子模块具体用于:
    按照第二步长调整所述第一循环方式的第i-1次循环处理的调整因子,以得到所述第i次循环处理的调整因子;
    其中,在i=1的情况下,所述第i-1次循环处理的调整因子为第一初始调整因子;
    所述继续调整条件包括所述第i次的编码比特数大于所述目标编码比特数。
  48. 如权利要求38或42所述的装置,其特征在于,所述第一因子确定子模块具体用于:
    在所述初始编码比特数小于所述目标编码比特数的情况下,将第一初始调整因子确定为所述第一变量调整因子。
  49. 如权利要求40或44所述的装置,其特征在于,所述第一因子确定子模块具体用于:
    在所述第i次的编码比特数等于所述目标编码比特数的情况下,将所述第i次循环处理的调整因子确定为所述第一变量调整因子;或者,
    在所述第i次的编码比特数不等于所述目标编码比特数的情况下,基于所述第i次循环处理的调整因子和所述第一循环方式的第i-1次循环处理的调整因子,确定所述第一变量调整因子。
  50. 如权利要求49所述的装置,其特征在于,所述第一因子确定子模块具体用于:
    确定所述第i次所述第一循环方式的调整因子和所述第i-1次所述第一循环方式的调整因子的平均值;
    基于所述平均值确定所述第一变量调整因子。
  51. 如权利要求49所述的装置,其特征在于,所述第一因子确定子模块具体用于:
    基于所述第i次所述第一循环方式的调整因子和所述第i-1次所述第一循环方式的调整因子,通过第二循环方式确定所述第一变量调整因子;
    所述第二循环方式的第j次循环处理包括如下步骤:
    基于所述第j次循环处理的第一调整因子和所述第j次循环处理的第二调整因子确定所述第j次循环处理的第三调整因子;其中,在j等于1的情况下,所述第j次循环处理的第一调整因子为所述第i次循环处理的调整因子和所述第i-1次循环处理的调整因子中的一者,所述第j次循环处理的第二调整因子为所述第i次循环处理的调整因子和所述第i-1次循环处理的调整因子中的另一者,所述第j次循环处理的第一调整因子对应第j次的第一编码比特数,所述第j次循环处理的第二调整因子对应第j次的第二编码比特数,所述第j次的第一编码比特数小于所述第j次的第二编码比特数,j为正整数;
    获取第j次的第三编码比特数,所述第j次的第三编码比特数是指经过所述第j次循环处理的第三调整因子调整后的第一潜在变量的熵编码结果的编码比特数;
    若所述第j次的第三编码比特数不满足继续循环条件,终止所述第二循环方式的执行,将所述第j次循环处理的第三调整因子确定为所述第一变量调整因子;
    若所述第j次的第三编码比特数满足所述继续循环条件、所述第j次的第三编码比特数大于所述目标编码比特数且小于所述第j次的第二编码比特数,将所述第j次循环处理的第三调整因子作为第j+1次循环处理的第二调整因子,将所述第j次循环处理的第一调整因子作为第j+1次循环处理的第一调整因子,执行所述第二循环方式的第j+1次循环处理;若所述第j次的第三编码比特数满足所述继续循环条件、所述第j次的第三编码比特数小于所述目标编码比特数且大于所述第j次的第一编码比特数,将所述第j次循环处理的第三调整因子作为第j+1次循环处理的第一调整因子,将所述第j次循环处理的第二调整因子作为第j+1次循环处理的第二调整因子,执行所述第二循环方式的第j+1次循环处理。
  52. 如权利要求49所述的装置,其特征在于,所述第一因子确定子模块具体用于:
    基于所述第i次循环处理的调整因子和所述第i-1次循环处理的调整因子,通过第二循环方式确定所述第一变量调整因子;
    所述第二循环方式的第j次循环处理包括如下步骤:
    基于所述第j次循环处理的第一调整因子和所述第j次循环处理的第二调整因子确定所述第j次循环处理的第三调整因子;其中,在j等于1的情况下,所述第j次循环处理的第一调整因子为所述第i次循环处理的调整因子和所述第i-1次循环处理的调整因子中的一者,所述第j次循环处理的第二调整因子为所述第i次循环处理的调整因子和所述第i-1次循环处理的调整因子中的另一者,所述第j次循环处理的第一调整因子对应第j次的第一编码比特数,所述第j次循环处理的第二调整因子对应第j次的第二编码比特数,所述第j次的第一编码比特数小于所述第j次的第二编码比特数,j为正整数;
    获取第j次的第三编码比特数,所述第j次的第三编码比特数是指经过所述第j次循环处理的第三调整因子调整后的第一潜在变量的熵编码结果的编码比特数;
    若所述第j次的第三编码比特数不满足继续循环条件,终止所述第二循环方式的执行,将所述第j次循环处理的第三调整因子确定为所述第一变量调整因子;
    若所述j达到最大循环次数且所述第j次的第三编码比特数满足所述继续循环条件,终止所述第二循环方式的执行,基于所述第j次循环处理的第一调整因子确定所述第一变量调整因子;
    若所述j未达到所述最大循环次数、所述第j次的第三编码比特数满足所述继续循环条件、所述第j次的第三编码比特数大于所述目标编码比特数且小于所述第j次的第二编码比特数,将所述第j次循环处理的第三调整因子作为第j+1次循环处理的第二调整因子,将所述第j次循环处理的第一调整因子作为第j+1次循环处理的第一调整因子,执行所述第二循环方式的第j+1次循环处理;
    若所述j未达到所述最大循环次数、所述第j次的第三编码比特数满足所述继续循环条件、所述第j次的第三编码比特数小于所述目标编码比特数且大于所述第j次的第一编码比特数,将所述第j次循环处理的第三调整因子作为第j+1次循环处理的第一调整因子,将所述第j次循环处理的第二调整因子作为第j+1次循环处理的第二调整因子,执行所述第二循环方式的第j+1次循环处理。
  53. 如权利要求52所述的装置,其特征在于,所述第一因子确定子模块具体用于:
    在使用固定码率对所述待编码的媒体数据进行编码的情况下,将所述第j次循环处理的第一调整因子确定为所述第一变量调整因子。
  54. 如权利要求52所述的装置,其特征在于,所述第一因子确定子模块具体用于:
    在使用可变码率对所述待编码的媒体数据进行编码的情况下,确定所述目标编码比特数与所述第j次的第一编码比特数之间的第一差值,以及确定所述第j次的第二编码比特数与所述目标编码比特数之间的第二差值;
    若所述第一差值小于所述第二差值,将所述第j次循环处理的第一调整因子确定为所述第一变量调整因子;
    若所述第二差值小于所述第一差值,将所述第j次循环处理的第二调整因子确定为所述第一变量调整因子;
    若所述第一差值等于所述第二差值,将所述第j次循环处理的第一调整因子确定为所述第一变量调整因子,或者,将所述第j次循环处理的第二调整因子确定为所述第一变量调整因子。
  55. 如权利要求51或52所述的装置,其特征在于,
    在使用固定码率对所述待编码的媒体数据进行编码的情况下,所述继续循环条件包括所述第j次的第三编码比特数大于所述目标编码比特数,或者,所述继续循环条件包括所述第j次的第三编码比特数小于所述目标编码比特数,且所述目标编码比特数与所述第j次的第三编码比特数的差值大于比特数阈值。
  56. 如权利要求51或52所述的装置,其特征在于,
    在使用可变码率对所述待编码的媒体数据进行编码的情况下,所述继续循环条件包括所述目标编码比特数与所述第j次的第三编码比特数的差值的绝对值大于所述比特数阈值。
  57. 如权利要求38或42所述的装置,其特征在于,所述第一因子确定子模块具体用于:
    在所述初始编码比特数等于所述目标编码比特数的情况下,将第一初始调整因子确定为所述第一变量调整因子。
  58. 如权利要求35所述的装置,其特征在于,所述装置还包括:
    所述调整因子确定模块,还用于基于所述第一潜在变量确定第二变量调整因子;
    第三编码结果获取模块,用于获取第四潜在变量的熵编码结果,所述第四潜在变量是基于所述第二变量调整因子对第三潜在变量调整后得到,所述第三潜在变量是基于所述第二潜在变量通过上下文模型确定得到;
    第三编码结果写入模块,用于将所述第四潜在变量的熵编码结果以及所述第二变量调整因子的编码结果写入所述码流;
    其中,所述第二潜在变量的熵编码结果和所述第四潜在变量的熵编码结果的编码总比特 数满足所述预设编码速率条件。
  59. 如权利要求58所述的装置,其特征在于,所述调整因子确定模块包括:
    第一确定子模块,用于基于所述第一潜在变量,通过所述上下文模型确定对应的上下文初始编码比特数和初始熵编码模型参数;
    第二确定子模块,用于基于所述初始熵编码模型参数,确定所述第一潜在变量的熵编码结果的编码比特数,以得到基础初始编码比特数;
    第二因子确定子模块,用于基于所述上下文初始编码比特数、所述基础初始编码比特数和目标编码比特数,确定所述第一变量调整因子和所述第二变量调整因子。
  60. 如权利要求59所述的装置,其特征在于,所述第二因子确定子模块具体用于:
    基于所述基础初始编码比特数和所述上下文初始编码比特数中的至少一者,以及所述目标编码比特数,确定基础目标编码比特数;
    基于第二初始调整因子、所述基础目标编码比特数和所述基础初始编码比特数,确定所述第一变量调整因子和基础实际编码比特数,所述基础实际编码比特数是指经过所述第一变量调整因子调整后的第一潜在变量的熵编码结果的编码比特数;
    基于所述目标编码比特数和所述基础实际编码比特数,确定上下文目标编码比特数;
    基于所述上下文目标编码比特数和所述上下文初始编码比特数,确定所述第二变量调整因子。
  61. 如权利要求59所述的装置,其特征在于,所述第二因子确定子模块具体用于:
    将所述目标编码比特数划分为基础目标编码比特数和上下文目标编码比特数;
    基于所述基础目标编码比特数和所述基础初始编码比特数,确定所述第一变量调整因子;
    基于所述上下文目标编码比特数和所述上下文初始编码比特数,确定所述第二变量调整因子。
  62. 如权利要求35-61任一所述的装置,其特征在于,所述媒体数据为音频信号、视频信号或者图像。
  63. 一种解码装置,其特征在于,所述装置包括:
    第一确定模块,用于基于码流确定重构的第二潜在变量和重构的第一变量调整因子;
    变量调整模块,用于基于所述重构的第一变量调整因子,对所述重构的第二潜在变量进行调整,得到重构的第一潜在变量,所述重构的第一潜在变量用于指示待解码的媒体数据的特征;
    变量处理模块,用于通过第一解码神经网络模型对所述重构的第一潜在变量进行处理,以得到重构的媒体数据。
  64. 如权利要求63所述的装置,其特征在于,所述第一确定模块包括:
    第一确定子模块,用于基于所述码流确定重构的第三潜在变量;
    第二确定子模块,用于基于所述码流和所述重构的第三潜在变量,确定所述重构的第二潜在变量。
  65. 如权利要求64所述的装置,其特征在于,所述第二确定子模块具体用于:
    通过上下文解码神经网络模型对所述重构的第三潜在变量进行处理,以得到重构的第一熵编码模型参数;
    基于所述码流和所述重构的第一熵编码模型参数,确定所述重构的第二潜在变量。
  66. 如权利要求63所述的装置,其特征在于,所述第一确定模块包括:
    第三确定子模块,用于基于所述码流确定重构的第四潜在变量和重构的第二变量调整因子;
    第四确定子模块,用于基于所述码流、所述重构的第四潜在变量和所述重构的第二变量调整因子,确定所述重构的第二潜在变量。
  67. 如权利要求66所述的装置,其特征在于,所述第四确定子模块具体用于:
    基于所述重构的第二变量调整因子,对所述重构的第四潜在变量进行调整,得到重构的第三潜在变量;
    通过上下文解码神经网络模型对所述重构的第三潜在变量进行处理,以得到重构的第二熵编码模型参数;
    基于所述码流和所述重构的第二熵编码模型参数,确定所述重构的第二潜在变量。
  68. 如权利要求63-67任一所述的装置,其特征在于,所述媒体数据为音频信号、视频信号或者图像。
  69. 一种编码端设备,其特征在于,所述编码端设备包括存储器和处理器;
    所述存储器用于存储计算机程序,所述处理器用于执行所述存储器中存储的计算机程序,以实现权利要求1-28任一所述的编码方法。
  70. 一种解码端设备,其特征在于,所述解码端设备包括存储器和处理器;
    所述存储器用于存储计算机程序,所述处理器用于执行所述存储器中存储的计算机程序,以实现权利要求29-34任一所述的解码方法。
  71. 一种计算机可读存储介质,其特征在于,所述存储介质内存储有指令,当所述指令在所述计算机上运行时,使得所述计算机执行权利要求1-34任一所述的方法的步骤。
  72. 一种计算机程序,其特征在于,所述计算机程序被执行时实现如权利要求1-34中任一项所述的方法。
  73. 一种计算机可读存储介质,其特征在于,包括如权利要求1-28中任一项所述的编码方 法所获得的码流。
PCT/CN2022/092385 2021-05-21 2022-05-12 编解码方法、装置、设备、存储介质及计算机程序 Ceased WO2022242534A1 (zh)

Priority Applications (6)

Application Number Priority Date Filing Date Title
MX2023013828A MX2023013828A (es) 2021-05-21 2022-05-12 Método y aparato de codificación, método y aparato de decodificación, dispositivo, medio de almacenamiento y programa informático.
BR112023024341A BR112023024341A2 (pt) 2021-05-21 2022-05-12 Método e aparelho de codificação, método e aparelho de decodificação, dispositivo, meio de armazenamento e programa de computador
KR1020237044093A KR20240011767A (ko) 2021-05-21 2022-05-12 인코딩 방법 및 장치, 디코딩 방법 및 장치, 디바이스, 저장 매체 및 컴퓨터 프로그램
JP2023572207A JP7656094B2 (ja) 2021-05-21 2022-05-12 エンコーディング方法および装置、デコーディング方法および装置、デバイス、記憶媒体、ならびにコンピュータプログラム
EP22803856.8A EP4333432A4 (en) 2021-05-21 2022-05-12 ENCODING METHOD AND DEVICE, DECODING METHOD AND DEVICE, DEVICE, STORAGE MEDIUM AND COMPUTER PROGRAM
US18/515,612 US12412587B2 (en) 2021-05-21 2023-11-21 Encoding method and apparatus, decoding method and apparatus, device, storage medium, and computer program

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
CN202110559102.7 2021-05-21
CN202110559102.7A CN115460182B (zh) 2021-05-21 2021-05-21 编解码方法、装置、设备、存储介质及计算机程序

Related Child Applications (1)

Application Number Title Priority Date Filing Date
US18/515,612 Continuation US12412587B2 (en) 2021-05-21 2023-11-21 Encoding method and apparatus, decoding method and apparatus, device, storage medium, and computer program

Publications (1)

Publication Number Publication Date
WO2022242534A1 true WO2022242534A1 (zh) 2022-11-24

Family

ID=84141105

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2022/092385 Ceased WO2022242534A1 (zh) 2021-05-21 2022-05-12 编解码方法、装置、设备、存储介质及计算机程序

Country Status (8)

Country Link
US (1) US12412587B2 (zh)
EP (1) EP4333432A4 (zh)
JP (1) JP7656094B2 (zh)
KR (1) KR20240011767A (zh)
CN (2) CN118694750A (zh)
BR (1) BR112023024341A2 (zh)
MX (1) MX2023013828A (zh)
WO (1) WO2022242534A1 (zh)

Citations (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20200210808A1 (en) * 2018-12-27 2020-07-02 Paypal, Inc. Data augmentation in transaction classification using a neural network
CN111787326A (zh) * 2020-07-31 2020-10-16 广州市百果园信息技术有限公司 一种熵编码及熵解码的方法和装置
CN112019865A (zh) * 2020-07-26 2020-12-01 杭州皮克皮克科技有限公司 一种用于深度学习编码的跨平台熵编码方法及解码方法
CN112449189A (zh) * 2019-08-27 2021-03-05 腾讯科技(深圳)有限公司 一种视频数据处理方法、装置及视频编码器、存储介质

Family Cites Families (10)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPH0993135A (ja) 1995-09-26 1997-04-04 Victor Co Of Japan Ltd 発声音データの符号化装置及び復号化装置
US20120082395A1 (en) * 2010-09-30 2012-04-05 Microsoft Corporation Entropy Coder for Image Compression
WO2014120368A1 (en) * 2013-01-30 2014-08-07 Intel Corporation Content adaptive entropy coding for next generation video
CN109479136A (zh) * 2016-08-04 2019-03-15 深圳市大疆创新科技有限公司 用于比特率控制的系统和方法
EP3704638A1 (en) * 2017-10-30 2020-09-09 Fraunhofer Gesellschaft zur Förderung der Angewand Neural network representation
WO2019155064A1 (en) * 2018-02-09 2019-08-15 Deepmind Technologies Limited Data compression using jointly trained encoder, decoder, and prior neural networks
CN111937389B (zh) * 2018-03-29 2022-01-14 华为技术有限公司 用于视频编解码的设备和方法
KR20210033781A (ko) * 2019-09-19 2021-03-29 주식회사 케이티 얼굴 분석 시스템 및 방법
CN112599139B (zh) * 2020-12-24 2023-11-24 维沃移动通信有限公司 编码方法、装置、电子设备及存储介质
FR3133265A1 (fr) * 2022-03-02 2023-09-08 Orange Codage et décodage optimisé d’un signal audio utilisant un auto-encodeur à base de réseau de neurones

Patent Citations (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20200210808A1 (en) * 2018-12-27 2020-07-02 Paypal, Inc. Data augmentation in transaction classification using a neural network
CN112449189A (zh) * 2019-08-27 2021-03-05 腾讯科技(深圳)有限公司 一种视频数据处理方法、装置及视频编码器、存储介质
CN112019865A (zh) * 2020-07-26 2020-12-01 杭州皮克皮克科技有限公司 一种用于深度学习编码的跨平台熵编码方法及解码方法
CN111787326A (zh) * 2020-07-31 2020-10-16 广州市百果园信息技术有限公司 一种熵编码及熵解码的方法和装置

Non-Patent Citations (1)

* Cited by examiner, † Cited by third party
Title
See also references of EP4333432A4

Also Published As

Publication number Publication date
JP7656094B2 (ja) 2025-04-02
CN115460182A (zh) 2022-12-09
CN115460182B (zh) 2024-07-05
KR20240011767A (ko) 2024-01-26
US20240087585A1 (en) 2024-03-14
JP2024518647A (ja) 2024-05-01
EP4333432A1 (en) 2024-03-06
US12412587B2 (en) 2025-09-09
MX2023013828A (es) 2024-02-15
BR112023024341A2 (pt) 2024-02-06
CN118694750A (zh) 2024-09-24
EP4333432A4 (en) 2024-04-24

Similar Documents

Publication Publication Date Title
CN115881138B (zh) 解码方法、装置、设备、存储介质及计算机程序产品
KR101647576B1 (ko) 스테레오 오디오 신호 인코더
JP7123910B2 (ja) インデックスコーディング及びビットスケジューリングを備えた量子化器
CN115881139B (zh) 编解码方法、装置、设备、存储介质及计算机程序
CN115881140B (zh) 编解码方法、装置、设备、存储介质及计算机程序产品
WO2021213128A1 (zh) 音频信号编码方法和装置
CN109983535B (zh) 具有子带能量平滑的基于变换的音频编解码器和方法
EP4336498B1 (en) Audio data encoding method and related apparatus, audio data decoding method and related apparatus, and computer-readable storage medium
US12412587B2 (en) Encoding method and apparatus, decoding method and apparatus, device, storage medium, and computer program
TWI854237B (zh) 音訊訊號的編解碼方法、裝置、設備、儲存介質及電腦程式
WO2023082773A1 (zh) 视频编解码方法、装置、设备、存储介质及计算机程序
CN118283485A (zh) 虚拟扬声器的确定方法及相关装置
CN120020946A (zh) 音频编解码方法、装置、设备
TW202349966A (zh) 濾波方法、濾波模型訓練方法及相關裝置
KR20240126443A (ko) Isobmff의 cmaf 스위칭 세트를 시그널링하는 방법 및 장치
CN120220703A (zh) 音频编解码方法、装置、设备及存储介质

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 22803856

Country of ref document: EP

Kind code of ref document: A1

WWE Wipo information: entry into national phase

Ref document number: 202317079057

Country of ref document: IN

Ref document number: 2023572207

Country of ref document: JP

Ref document number: MX/A/2023/013828

Country of ref document: MX

WWE Wipo information: entry into national phase

Ref document number: 2022803856

Country of ref document: EP

REG Reference to national code

Ref country code: BR

Ref legal event code: B01A

Ref document number: 112023024341

Country of ref document: BR

ENP Entry into the national phase

Ref document number: 2022803856

Country of ref document: EP

Effective date: 20231127

ENP Entry into the national phase

Ref document number: 20237044093

Country of ref document: KR

Kind code of ref document: A

WWE Wipo information: entry into national phase

Ref document number: 1020237044093

Country of ref document: KR

NENP Non-entry into the national phase

Ref country code: DE

WWE Wipo information: entry into national phase

Ref document number: 11202308669T

Country of ref document: SG

ENP Entry into the national phase

Ref document number: 112023024341

Country of ref document: BR

Kind code of ref document: A2

Effective date: 20231121