WO2023197717A1 - 一种图像解码方法、编码方法及装置 - Google Patents
一种图像解码方法、编码方法及装置 Download PDFInfo
- Publication number
- WO2023197717A1 WO2023197717A1 PCT/CN2023/071928 CN2023071928W WO2023197717A1 WO 2023197717 A1 WO2023197717 A1 WO 2023197717A1 CN 2023071928 W CN2023071928 W CN 2023071928W WO 2023197717 A1 WO2023197717 A1 WO 2023197717A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- feature map
- feature
- image
- optical flow
- map
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
- G06T9/00—Image coding
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/10—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
- H04N19/169—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding
- H04N19/17—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding the unit being an image region, e.g. an object
- H04N19/172—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding the unit being an image region, e.g. an object the region being a picture, frame or field
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
- G06T3/00—Geometric image transformations in the plane of the image
- G06T3/18—Image warping, e.g. rearranging pixels individually
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
- G06T7/00—Image analysis
- G06T7/20—Analysis of motion
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V10/00—Arrangements for image or video recognition or understanding
- G06V10/70—Arrangements for image or video recognition or understanding using pattern recognition or machine learning
- G06V10/77—Processing image or video features in feature spaces; using data integration or data reduction, e.g. principal component analysis [PCA] or independent component analysis [ICA] or self-organising maps [SOM]; Blind source separation
- G06V10/771—Feature selection, e.g. selecting representative features from a multi-dimensional feature space
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V10/00—Arrangements for image or video recognition or understanding
- G06V10/70—Arrangements for image or video recognition or understanding using pattern recognition or machine learning
- G06V10/77—Processing image or video features in feature spaces; using data integration or data reduction, e.g. principal component analysis [PCA] or independent component analysis [ICA] or self-organising maps [SOM]; Blind source separation
- G06V10/80—Fusion, i.e. combining data from various sources at the sensor level, preprocessing level, feature extraction level or classification level
- G06V10/806—Fusion, i.e. combining data from various sources at the sensor level, preprocessing level, feature extraction level or classification level of extracted features
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/10—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
- H04N19/102—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the element, parameter or selection affected or controlled by the adaptive coding
- H04N19/103—Selection of coding mode or of prediction mode
- H04N19/105—Selection of the reference unit for prediction within a chosen coding or prediction mode, e.g. adaptive choice of position and number of pixels used for prediction
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/42—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals characterised by implementation details or hardware specially adapted for video compression or decompression, e.g. dedicated software implementation
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/44—Decoders specially adapted therefor, e.g. video decoders which are asymmetric with respect to the encoder
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/50—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/50—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding
- H04N19/503—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding involving temporal prediction
- H04N19/51—Motion estimation or motion compensation
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/50—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding
- H04N19/503—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding involving temporal prediction
- H04N19/51—Motion estimation or motion compensation
- H04N19/577—Motion compensation with bidirectional frame interpolation, i.e. using B-pictures
Definitions
- the present application relates to the field of video processing technology, and in particular, to an image decoding method, encoding method and device.
- Video compression systems perform spatial (intra-image) prediction and/or temporal (inter-image) prediction to reduce or remove redundant information inherent in video sequences; video enhancement techniques are used to improve the display quality of images.
- the decoder uses the optical flow mapping (warping) method to decode the image frames contained in the video.
- Optical flow mapping means that the decoding end obtains the image domain optical flow between the image frame and the reference frame, and decodes the image frame based on this optical flow.
- Image domain optical flow is used to indicate the motion speed and motion direction between corresponding pixels in two adjacent frames of images.
- optical flow mapping is more sensitive to optical flow accuracy, and subtle changes in optical flow accuracy will affect the accuracy of warping. Since the error in optical flow prediction between two adjacent frames is large, such as an error of 5 pixels or more, the accuracy of decoding the image frame based on the image domain optical flow at the decoder is low, resulting in decoding The clarity of the images obtained is affected. Therefore, how to provide a more effective image decoding method has become an urgent problem that needs to be solved.
- the present application provides an image decoding method, encoding method and device, which solves the problem that the accuracy of decoding image frames based on image domain optical flow is low, causing the clarity of the decoded image to be affected.
- this application provides an image decoding method, which method is applied to a video coding and decoding system, or a physical device that supports the implementation of the video coding and decoding system to implement the image decoding method.
- the physical device is a decoding end or a video decoder. (decoder), in some cases, the physical device may include a chip system.
- the image decoding method performed by the decoding end is used as an example.
- the image decoding method includes: first, the decoding end parses the code stream to obtain at least one optical flow set; the at least one optical flow set includes a first optical flow set, and the first optical flow set
- the flow set corresponds to the first feature map of the reference frame of the first image frame, and the first optical flow set includes one or more feature domain optical flows, any one of the one or more feature domain optical flows.
- the stream is used to indicate motion information between the feature map of the first image frame and the aforementioned first feature map.
- the decoding end processes the first feature map based on one or more feature domain optical flows included in the first optical flow set, and obtains one or more intermediate feature maps corresponding to the first feature map; and, the decoding end fuses the aforementioned one or more A plurality of intermediate feature maps are used to obtain a first predicted feature map of the first image frame. Finally, the decoder decodes the first image frame according to the first prediction feature map to obtain the first image.
- the error of the pixels in the feature domain optical flow is smaller than the error of the pixels in the image domain optical flow. Therefore, the intermediate feature map determined by the decoding end based on the feature domain optical flow brings The resulting decoding error is lower than the decoding error caused by image domain optical flow in common technology.
- the decoding end decodes the image frame based on the feature domain optical flow, which reduces the decoding error caused by the image domain optical flow between two adjacent frames. Decoding errors improve the accuracy of image decoding.
- the decoder fuses multiple intermediate feature maps determined by the optical flow set and the feature map of the reference frame to obtain a predicted feature map of the image frame.
- This predicted feature map contains more image information than the optical flow of a single image domain.
- the decoder decodes the image frame based on the prediction feature map obtained by fusion, it avoids the difficulty of accurately representing the information indicated by the optical flow in a single image domain.
- the problem of reaching the first image is improved, and the accuracy of image decoding and image quality (such as image clarity, etc.) are improved.
- the decoder processes the first feature map based on one or more feature domain optical flows included in the first optical flow set, and obtains one or more intermediate feature maps corresponding to the first feature map, including :
- the decoding end parses the code stream to obtain the first feature map of the aforementioned reference frame; and, for one or more feature domain optical flows included in the first optical flow set, the decoding end analyzes the first feature domain optical flow based on the first feature domain optical flow.
- the feature map is subjected to optical flow mapping (warping) to obtain an intermediate feature map corresponding to the first feature domain optical flow; the first feature domain optical flow is any one of the aforementioned one or more feature domain optical flows.
- the decoder performs warping on the first feature map based on a set of feature domain optical flows. Then, after the decoder obtains the image corresponding to the reference frame, the decoder can warp the first feature map in the image corresponding to the reference frame according to a set of The feature domain optical flow interpolates the corresponding positions of the set of feature domain optical flows in the image corresponding to the reference frame to obtain the predicted value of the pixel value in the first image, thereby obtaining the first image, avoiding the need for the decoder to rely on the image domain optical flow. Predicting all pixel values of the first image reduces the amount of calculation required by the decoder for image decoding and improves the efficiency of image decoding.
- the aforementioned first optical flow set also corresponds to the second feature map of the reference frame.
- the image decoding method before the decoding end decodes the first image frame according to the first prediction feature map to obtain the first image, the image decoding method further includes: a first step, the decoding end includes based on the first optical flow set: One or more feature domain optical flows process the second feature map to obtain one or more intermediate feature maps corresponding to the second feature map.
- the decoder fuses one or more intermediate feature maps corresponding to the second feature map to obtain the second predicted feature map of the first image frame. In this way, the decoder can decode the first image frame according to the first prediction feature map and the second prediction feature map to obtain the first image.
- multiple feature maps (or a group of feature maps) of the reference frame can correspond to an optical flow set (or a set of feature domain optical flows).
- the decoding end processes the feature map according to a set of feature domain optical flows corresponding to the feature map, thereby obtaining an intermediate feature map corresponding to the feature map.
- the decoder end processes the intermediate feature map corresponding to the feature map. Fusion is performed to obtain the predicted feature map corresponding to the feature map.
- the decoder decodes the image frame according to the prediction feature maps corresponding to all feature maps of the reference frame to obtain the target image.
- the decoder can divide the multiple feature maps into one group or multiple groups.
- the feature maps belonging to the same group share an optical flow set, and
- the intermediate feature map corresponding to the feature map is fused to obtain the predicted feature map, which avoids the problem that the decoder reconstructs the image based on the feature map with low accuracy and slow speed when the reference frame or image frame has more information, and improves the image quality. Decoding accuracy. It is worth noting that the number of channels of feature maps belonging to different groups can be different.
- the decoder fuses one or more intermediate feature maps to obtain the first predicted feature map of the first image frame, including: the decoder obtains one or more of the one or more intermediate feature maps. Multiple weights, where one intermediate feature map corresponds to one weight; and, the decoder processes the intermediate feature maps corresponding to the one or more weights based on the one or more intermediate feature map weights, and all processes The subsequent intermediate feature maps are added to obtain the first predicted feature map.
- the weight is used to indicate the weight of the intermediate feature map in the first predicted feature map.
- the weight of each intermediate feature map can be different.
- the decoder can set different weights for the intermediate feature maps according to the requirements of image decoding. , For example, if the images corresponding to some intermediate feature maps are relatively blurry, the weights of these intermediate feature maps are reduced, thereby improving the clarity of the first image.
- the decoder fuses one or more intermediate feature maps to obtain the first predicted feature map of the first image frame, including: converting one or more intermediate feature maps corresponding to the first feature map Input the feature fusion model to obtain the first predicted feature map.
- the feature fusion model includes a convolutional network layer, which is used to fuse intermediate feature maps.
- the decoder obtains multiple image domain optical flows between image frames and reference frames, and obtains multiple images based on the multiple image domain optical flows and reference frames, thereby fusing multiple images to obtain image frames.
- the decoding end needs to predict the pixel values of multiple images when decoding one frame of image, and fuse multiple images to obtain the target image, resulting in large computing resources required for image decoding. Streaming decoded video is less efficient.
- the decoder uses a feature fusion model to fuse multiple intermediate feature maps, and then , the decoding end decodes the image frame based on the prediction feature map obtained by fusing multiple intermediate feature maps, that is, the decoding end only needs to predict the pixel value of the image position indicated by the prediction feature map in the image based on the prediction feature map, without prediction. All pixel values of multiple images reduce the computing resources required for image decoding and improve the efficiency of image decoding.
- the image decoding method provided by this application also includes: first, the decoding end obtains the feature map of the first image. Secondly, the decoding end obtains the enhanced feature map based on the feature map of the first image, the first feature map and the first predicted feature map. Finally, the decoder processes the first image according to the enhanced feature map to obtain the second image.
- the second image has a higher definition than the first image.
- the decoder can fuse the feature map of the first image, the first feature map and the first predicted feature map, and perform video enhancement processing on the first image based on the enhanced feature map obtained through the fusion to obtain better clarity. Excellent second image, thereby improving the decoded image clarity and image display effect.
- the decoder processes the first image according to the enhanced feature map to obtain the second image, including: the decoder obtains the enhancement layer image of the first image based on the enhanced feature map, and reconstructs the image based on the enhancement layer image. Construct the first image and obtain the second image.
- the enhancement layer image may refer to an image determined by the decoder based on the reference frame and the enhancement feature map.
- the decoder adds part or all of the information of the enhancement layer image to the first image to obtain the second image; or , the decoder uses the enhancement layer image as the reconstructed image of the first image, that is, the aforementioned second image.
- the decoder obtains multiple feature maps of the image at different stages, and obtains the enhanced feature map determined by these multiple feature maps, so as to reconstruct and enhance the first image based on the enhanced feature map. Improved image clarity and image display after decoding.
- the second aspect provides an image coding method, which method is applied to a video coding and decoding system, or a physical device that supports the implementation of the video coding and decoding system to implement the image coding method.
- the physical device is an encoding end or a video encoder (encoder). ), in some cases, the physical device may include a system-on-a-chip.
- an image coding method performed by the encoding end is used as an example.
- the image encoding method includes: in the first step, the encoding end obtains the feature map of the first image frame and the first feature map of the reference frame of the first image frame.
- the encoding end obtains at least one optical flow set based on the feature map and the first feature map of the first image frame; the at least one optical flow set includes a first optical flow set, and the first optical flow set corresponds to the aforementioned first optical flow set.
- feature map, and the first optical flow set includes one or more feature domain optical flows, any one of the aforementioned one or more feature domain optical flows is used to indicate the feature map and the first feature of the first image frame Motion information between graphs.
- the encoding end processes the first feature map based on one or more feature domain optical flows included in the first optical flow set, and obtains one or more intermediate feature maps corresponding to the first feature map.
- the encoding end fuses the aforementioned one or more intermediate feature maps to obtain the first predicted feature map of the first image frame.
- the encoding end encodes the first image frame according to the first prediction feature map to obtain a code stream.
- the code stream includes: a feature domain optical flow code stream corresponding to at least one optical flow set, and a residual code stream of the image area corresponding to the first prediction feature map.
- the error of pixels in the feature domain optical flow is smaller than that of the image domain optical flow.
- the error of pixels in the flow Therefore, the encoding error caused by the intermediate feature map determined by the encoding end based on the feature domain optical flow is lower than the encoding error caused by the image domain optical flow in the common technology.
- the encoding end determines based on the feature domain optical flow.
- Encoding image frames reduces coding errors caused by image domain optical flow between two adjacent frames and improves the accuracy of image coding.
- the encoding end processes the feature maps of the reference frame based on multiple feature domain optical flows to obtain multiple intermediate feature maps, and fuses the multiple intermediate feature maps to determine the predicted feature map of the image frame.
- the encoding end combines the optical flow set and the reference frame.
- the multiple intermediate feature maps determined by the feature map are fused to obtain the prediction feature map of the image frame.
- the prediction feature map contains more image information, so that when the encoding end encodes the image frame based on the prediction feature map obtained by the fusion, This avoids the problem that a single intermediate feature map is difficult to accurately express the first image, and improves the accuracy of image coding and image quality (such as image clarity, etc.).
- the encoding end processes the first feature map based on one or more feature domain optical flows included in the first optical flow set, and obtains one or more intermediate feature maps corresponding to the first feature map, including : For one or more feature domain optical flows included in the first optical flow set, optical flow mapping is performed on the first feature map based on the first feature domain optical flow, and an intermediate feature map corresponding to the first feature domain optical flow is obtained. .
- the first feature map is any one of one or more feature domain optical flows included in the first optical flow set.
- the encoding end uses a feature fusion model to fuse multiple intermediate feature maps, and then , the encoding end encodes the image frame based on the prediction feature map obtained by fusing multiple intermediate feature maps, that is, the encoding end only needs to predict the pixel value of the image position indicated by the prediction feature map in the image based on the prediction feature map, without prediction. All pixel values of multiple images reduce the computing resources required for image encoding and improve the efficiency of image encoding.
- the first optical flow set also corresponds to the second feature map of the reference frame.
- the image encoding method provided in this embodiment further includes: a first step, the encoding end bases one or more of the first optical flow set on The feature domain optical flow processes the second feature map to obtain one or more intermediate feature maps corresponding to the second feature map.
- the encoding end fuses one or more intermediate feature maps corresponding to the second feature map to obtain the second predicted feature map of the first image frame.
- the aforementioned coding end encoding the first image frame according to the first prediction feature map to obtain a code stream may include: the encoding end encoding the first image frame according to the first prediction feature map and the second prediction feature map to obtain a code stream.
- multiple feature maps (or a group of feature maps) of the reference frame can correspond to an optical flow set (or a set of feature domain optical flows).
- the encoding end processes the feature map according to a set of feature domain optical flows corresponding to the feature map, thereby obtaining an intermediate feature map corresponding to the feature map.
- the encoding end converts the intermediate feature map corresponding to the feature map Fusion is performed to obtain the predicted feature map corresponding to the feature map.
- the encoding end encodes the image frame according to the prediction feature maps corresponding to all feature maps of the reference frame to obtain the target code stream. In this way, during the image encoding process, if the reference frame corresponds to multiple feature maps, the encoding end can divide the multiple feature maps into one group or multiple groups.
- the feature maps belonging to the same group share an optical flow set, and
- the intermediate feature map corresponding to the feature map is fused to obtain the predicted feature map, which avoids the problem that when the reference frame or image frame has more information, the encoding end obtains a code stream based on the feature map encoding with more redundancy and lower accuracy, and improves improve the accuracy of image coding. It is worth noting that the number of channels of feature maps belonging to different groups can be different.
- the encoding end fuses one or more intermediate feature maps to obtain the first predicted feature map of the first image frame, including: the encoding end obtains one or more of the one or more intermediate feature maps. Multiple weights, wherein one intermediate feature map corresponds to one weight; and, the encoding end processes the intermediate feature maps corresponding to the one or more weights based on the aforementioned one or more weights, and converts all processed The intermediate feature maps are added to obtain the first predicted feature map.
- the weight is used to indicate the weight of the intermediate feature map in the first predicted feature map. In this embodiment, for the first For multiple intermediate feature maps corresponding to the feature map, the weight of each intermediate feature map can be different.
- the encoding end can set different weights for the intermediate feature map according to the requirements of image encoding. For example, if some intermediate features When the image corresponding to the image is relatively blurry, the weights of these intermediate feature maps are reduced to improve the clarity of the first image.
- the encoding end fuses one or more intermediate feature maps to obtain the first predicted feature map of the first image frame, including: the encoding end fuses one or more intermediate feature maps corresponding to the first feature map.
- the feature map is input into the feature fusion model to obtain the first predicted feature map.
- the feature fusion model includes a convolutional network layer, which is used to fuse intermediate feature maps.
- the encoding end uses a feature fusion model to fuse multiple intermediate feature maps.
- the encoding end uses a feature fusion model to fuse multiple intermediate feature maps.
- the predicted feature map obtained from the intermediate feature map encodes the image frame, reducing the computing resources required for image encoding and improving the efficiency of image encoding.
- an image decoding device which may be applied to a decoding end or support a video coding and decoding system that implements the foregoing image decoding method.
- the image decoding device includes various modules for executing the image decoding method in the first aspect or any possible implementation of the first aspect.
- the image decoding device includes: a code stream unit, a processing unit, a fusion unit and a decoding unit.
- the code stream unit is used to parse the code stream to obtain at least one optical flow set.
- the at least one optical flow set includes a first optical flow set corresponding to the first feature map of the reference frame of the first image frame.
- the first optical flow set includes one or more feature domain optical flows, as described above. Any one of the one or more feature domain optical flows is used to indicate motion information between the feature map of the first image frame and the first feature map.
- a processing unit configured to process the first feature map based on one or more feature domain optical flows included in the first optical flow set, and obtain one or more intermediate feature maps corresponding to the first feature map.
- the fusion unit is used to fuse one or more intermediate feature maps to obtain the first predicted feature map of the first image frame.
- a decoding unit configured to decode the first image frame according to the first prediction feature map to obtain the first image.
- the image decoding device has the function of implementing the behaviors in the method examples of any of the above first aspects.
- the functions described can be implemented by hardware, or can be implemented by hardware executing corresponding software.
- the hardware or software includes one or more modules corresponding to the above functions.
- the processing unit is specifically configured to: parse the code stream to obtain the first feature map of the reference frame, and perform optical flow mapping (warping) on the first feature map based on the first feature domain optical flow. , obtain an intermediate feature map corresponding to the first feature domain optical flow; the first feature domain optical flow is any one of one or more feature domain optical flows included in the aforementioned first optical flow set.
- the first optical flow set also corresponds to the second feature map of the reference frame.
- the processing unit is also configured to process the second feature map based on one or more feature domain optical flows included in the first optical flow set, and obtain one or more intermediate feature maps corresponding to the second feature map.
- the fusion unit is also used to fuse one or more intermediate feature maps corresponding to the second feature map to obtain the second predicted feature map of the first image frame.
- the decoding unit is specifically configured to decode the first image frame according to the first prediction feature map and the second prediction feature map to obtain the first image.
- the fusion unit is specifically used to: obtain one or more weights of the aforementioned one or more intermediate feature maps, where one intermediate feature map corresponds to a weight, and the weight is expressed as Indicates the weight of the intermediate feature map in the first predicted feature map. and, based on one or more weights, process intermediate feature maps respectively corresponding to the aforementioned one or more weights, and add all processed intermediate feature maps to obtain a first prediction feature map.
- the fusion unit is specifically configured to input one or more intermediate feature maps corresponding to the first feature map into the feature fusion model to obtain the first predicted feature map.
- the feature fusion model includes a convolutional network layer, which is used to fuse intermediate feature maps.
- the image decoding device further includes: an acquisition unit and an enhancement unit.
- An acquisition unit is used to acquire the feature map of the first image.
- the fusion unit is also configured to obtain an enhanced feature map based on the feature map of the first image, the first feature map and the first predicted feature map.
- the enhancement unit is used to process the first image according to the enhanced feature map to obtain the second image.
- the second image has a higher definition than the first image.
- the enhancement unit is specifically configured to: obtain the enhancement layer image of the first image based on the enhancement feature map. And, reconstruct the first image based on the enhancement layer image to obtain the second image.
- the fourth aspect provides an image encoding device, which can be applied to the encoding end, or supports a video encoding and decoding system that implements the aforementioned image encoding method.
- the image encoding device includes various modules for executing the image encoding method in the second aspect or any possible implementation of the second aspect.
- the image coding device includes: an acquisition unit, a processing unit, a fusion unit and a coding unit.
- the obtaining unit is configured to obtain the feature map of the first image frame and the first feature map of the reference frame of the first image frame.
- a processing unit configured to obtain at least one optical flow set according to the feature map of the first image frame and the first feature map.
- the at least one optical flow set includes a first optical flow set corresponding to the first feature map of the reference frame of the first image frame.
- the first optical flow set includes one or more feature domain optical flows, as described above. Any one of the one or more feature domain optical flows is used to indicate motion information between the feature map of the first image frame and the first feature map.
- the processing unit is also configured to process the first feature map based on one or more feature domain optical flows included in the first optical flow set, and obtain one or more intermediate feature maps corresponding to the first feature map.
- the fusion unit is used to fuse one or more intermediate feature maps to obtain the first predicted feature map of the first image frame.
- An encoding unit configured to encode the first image frame according to the first prediction feature map to obtain a code stream.
- the code stream includes: a feature domain optical flow code stream corresponding to at least one optical flow set, and a residual code stream of the image area corresponding to the first prediction feature map.
- the image encoding device has the function of implementing the behaviors in the method examples of any of the above-mentioned first aspects.
- the functions described can be implemented by hardware, or can be implemented by hardware executing corresponding software.
- the hardware or software includes one or more modules corresponding to the above functions.
- the processing unit is specifically configured to: and, for one or more feature domain optical flows included in the first optical flow set, calculate the first feature based on the first feature domain optical flow therein.
- Optical flow mapping (warping) is performed on the graph to obtain an intermediate feature map corresponding to the first feature domain optical flow; the first feature domain optical flow is any one of one or more feature domain optical flows included in the first optical flow set.
- the first optical flow set also corresponds to the second feature map of the reference frame.
- the processing unit is also configured to process the second feature map based on one or more feature domain optical flows included in the first optical flow set, and obtain one or more intermediate feature maps corresponding to the second feature map.
- the fusion unit is also used to fuse one or more intermediate feature maps corresponding to the second feature map to obtain the second predicted feature map of the first image frame.
- the encoding unit is specifically configured to encode the first image frame according to the first prediction feature map and the second prediction feature map to obtain a code stream.
- the fusion unit is specifically configured to: obtain one or more weights of the aforementioned one or more intermediate feature maps, where one intermediate feature map corresponds to a weight, and the weight is expressed as to indicate that the intermediate feature map is in The weight of the first predicted feature map. and, based on the one or more weights, process the intermediate feature maps respectively corresponding to the one or more weights, and add all the processed intermediate feature maps to obtain the first prediction feature map.
- the fusion unit is specifically configured to input one or more intermediate feature maps corresponding to the first feature map into the feature fusion model to obtain the first predicted feature map.
- the feature fusion model includes a convolutional network layer, which is used to fuse intermediate feature maps.
- an image decoding device in a fifth aspect, includes at least one processor and a memory.
- the memory is used to store program code.
- the processor calls the program code, the first aspect or The operation steps of the image decoding method in any possible implementation of the first aspect, for example, the image decoding device may be a decoding end or a video decoder included in the video encoding and decoding system.
- the image decoding device may be a coding end or a video encoder included in the video coding and decoding system.
- a computer-readable storage medium In a sixth aspect, a computer-readable storage medium is provided. Computer programs or instructions are stored in the storage medium. When the computer program or instructions are executed by an electronic device, it is possible to implement the first aspect or any one of the first aspects.
- the operation steps of the image decoding method in the manner, and/or, the operation steps of the image encoding method in the second aspect or any possible implementation manner of the second aspect are performed.
- the electronic device refers to the aforementioned image decoding device.
- another computer-readable storage medium which stores a code stream obtained according to the image encoding method in the second aspect or any of the possible implementations of the second aspect.
- the code stream may include: a feature domain optical flow code stream corresponding to at least one optical flow set in the second aspect, and a residual code stream of the image area corresponding to the first prediction feature map.
- An eighth aspect provides a video encoding and decoding system, including an encoding end and a decoding end.
- the decoding end can perform the operation steps of the image decoding method in the first aspect or any possible implementation of the first aspect.
- the encoding end can Execute the operational steps of the image encoding method in the second aspect or any possible implementation manner of the second aspect.
- beneficial effects please refer to the description of any one of the first aspects, or the description of any one of the second aspects, and will not be described again here.
- a computer program product When the computer program product is run on an electronic device, it causes the electronic device to perform the operation steps of the method described in the first aspect or any possible implementation of the first aspect, And/or, when the computer program product is run on the computer, the electronic device is caused to perform the operation steps of the method described in the first aspect or any possible implementation manner of the first aspect.
- the electronic device refers to the aforementioned image decoding device.
- a chip including a control circuit and an interface circuit.
- the interface circuit is used to receive signals from other devices outside the chip and transmit them to the processor, or to send signals from the control circuit to devices outside the chip.
- Other devices, control circuits, logic circuits or execution code instructions are used to implement: the above-mentioned first aspect or the operational steps of the method in any possible implementation of the first aspect, and/or the above-mentioned second aspect or the second aspect. Operational steps of the method in any possible implementation.
- the chip refers to the aforementioned image decoding device.
- Figure 1 is an exemplary block diagram of the video encoding and decoding system provided by this application.
- Figure 2 is a mapping relationship diagram between optical flow and color provided by this application.
- FIG. 3 is a schematic diagram of optical flow mapping provided by this application.
- FIG. 4 is a schematic diagram of the architecture of the video compression system provided by this application.
- Figure 5 is a schematic flow chart of the image encoding method provided by this application.
- Figure 6 is a schematic structural diagram of the feature extraction network provided by this application.
- Figure 7 is a schematic structural diagram of the optical flow estimation network provided by this application.
- Figure 8 is a schematic flow chart of optical flow mapping and feature fusion provided by this application.
- Figure 9 is a schematic flow chart 1 of the image decoding method provided by this application.
- Figure 10 is a schematic structural diagram of the feature encoding network and feature decoding network provided by this application.
- Figure 11 is a schematic flow chart 2 of the image decoding method provided by this application.
- Figure 12 is a schematic structural diagram of the image reconstruction network provided by this application.
- Figure 13 is a schematic diagram of the variable convolution network provided by this application.
- FIG 14 is a schematic diagram of the efficiency comparison provided by this application.
- Figure 15 is a schematic structural diagram of the image encoding device provided by this application.
- Figure 16 is a schematic structural diagram of the image decoding device provided by this application.
- Figure 17 is a schematic structural diagram of an image decoding device provided by this application.
- Embodiments of the present application provide an image decoding (encoding) method, including: the decoding end (or encoding end) obtains a set of feature domain optical flows (or an optical flow set) of the image frame and the feature map of the reference frame. multiple intermediate feature maps between them, and fuse the multiple intermediate feature maps to obtain the predicted feature map of the image frame, so that the decoding end (or encoding end) decodes (or encodes) the image frame based on the predicted feature map , obtain the target image (or code stream) corresponding to the image frame.
- the image decoding process as an example, for a single feature domain optical flow, the error of the pixels in the feature domain optical flow is smaller than the error of the pixels in the image domain optical flow.
- the decoding end determines the intermediate feature map based on the feature domain optical flow.
- the decoding error caused is lower than the decoding error caused by the image domain optical flow in common technology.
- the decoder decodes the image frame based on the feature domain optical flow, which reduces the image domain optical flow between two adjacent frames.
- the decoding error improves the accuracy of image decoding.
- the decoding end processes the feature maps of the reference frame based on multiple feature domain optical flows to obtain multiple intermediate feature maps, and fuses the multiple intermediate feature maps to determine the predicted feature map of the image frame. That is, the decoding end combines the optical flow set and the reference frame.
- the multiple intermediate feature maps determined by the feature map are fused to obtain the predicted feature map of the image frame.
- the predicted feature map contains more image information, so that when the decoder decodes the image frame based on the predicted feature map obtained by fusion, This avoids the problem that a single intermediate feature map is difficult to accurately express the first image, and improves the accuracy of image decoding and image quality (such as image clarity, etc.).
- image decoding and image quality such as image clarity, etc.
- FIG. 1 is an exemplary block diagram of a video encoding and decoding system provided by this application.
- the term "video decoder” generally refers to both a video encoder and a video decoder.
- video coding or “decoding” may generally refer to video encoding or video decoding.
- the encoding end or the decoding end may be collectively referred to as an image decoding device.
- the video encoding and decoding system includes an encoding end 10 and a decoding end 20.
- Encoding end 10 generates encoded video data. Therefore, the encoding end 10 may be called a video encoding device.
- the decoding end 20 may decode the encoded video data (such as a video including one or more image frames) generated by the encoding end 10 . Therefore, the decoding end 20 may be called a video decoding device.
- Various implementations of encoding end 10, decoding end 20, or both may include one or more processors and memory coupled to the one or more processors.
- the memory may include, but is not limited to, random access memory (random access memory, RAM), flash memory, read-only memory (ROM), programmable ROM (PROM), erasable programmable read-only memory (erasable PROM, EPROM), electronic Electrically erasable programmable read-only memory (EEPROM), flash memory, or any other medium that can be used to store desired program code in the form of instructions or data structures accessible by a computer.
- RAM random access memory
- ROM read-only memory
- PROM programmable ROM
- EPROM erasable programmable read-only memory
- EEPROM electronic Electrically erasable programmable read-only memory
- flash memory or any other medium that can be used to store desired program code in the form of instructions or data structures accessible by a computer.
- Encoding end 10 and decoding end 20 may include a variety of devices, including desktop computers, mobile computing devices, notebook (eg, laptop) computers, tablet computers, set-top boxes, telephone handsets such as so-called “smart” phones. , televisions, cameras, display devices, digital media players, video game consoles, vehicle-mounted computers, or the like.
- Decoding end 20 may receive encoded video data from encoding end 10 via link 30 .
- Link 30 may include one or more media or devices capable of moving encoded video data from encoding end 10 to decoding end 20 .
- link 30 may include one or more communication media that enables encoding end 10 to transmit encoded video data directly to decoding end 20 in real time.
- the encoding end 10 may modulate the encoded video data according to a communication standard (eg, a wireless communication protocol), and may transmit the modulated video data to the decoding end 20 .
- the one or more communication media may include wireless and/or wired communication media, such as the radio frequency (RF) spectrum or one or more physical transmission lines.
- RF radio frequency
- the one or more communication media may form part of a packet-based network, such as a local area network, a wide area network, or a global network (eg, the Internet).
- the one or more communication media may include routers, switches, base stations, or other devices that facilitate communication from the encoding end 10 to the decoding end 20.
- the encoded data may be output from output interface 140 to storage device 40.
- encoded data may be accessed from storage device 40 through input interface 240.
- Storage device 40 may include any of a variety of distributed or locally accessed data storage media, such as a hard drive, Blu-ray Disc, digital video disc (DVD), compact disc read-only memory, CD-ROM), flash memory, volatile or non-volatile memory, or any other suitable digital storage medium for storing encoded video data.
- storage device 40 may correspond to a file server or another intermediate storage device that may hold encoded video generated by encoding end 10 .
- the decoding end 20 may access the stored video data from the storage device 40 via streaming or downloading.
- the file server may be any type of server capable of storing encoded video data and transmitting the encoded video data to decoding end 20 .
- Example file servers include network servers (for example, for websites), file transfer protocol (FTP) servers, network attached storage (NAS) devices, or local disk drives. Decoder 20 may access the encoded video data through any standard data connection, including an Internet connection.
- the transmission of encoded video data from storage device 40 may be a streaming transmission, a download transmission, or a combination of both.
- the image decoding method provided by the present application can be applied to video encoding and decoding to support a variety of multimedia applications, such as over-the-air television broadcasting, cable television transmission, satellite television transmission, streaming video transmission (for example, via the Internet), for storage in data Encoding of video data on a storage medium, decoding of video data stored on a data storage medium, or other applications.
- a video codec system may be used to support one-way or two-way video transmission to support applications such as video streaming, video enhancement, video playback, video broadcasting, and/or video telephony.
- the video codec system illustrated in FIG. 1 is only an example, and the techniques of this application are applicable to video coding arrangements (eg, video encoding or video decoding) that do not necessarily involve any data communication between encoding devices and decoding devices.
- the data is retrieved from local storage, streamed over the network, and so on.
- the video encoding device can process the data
- the data is encoded and stored to memory, and/or the video decoding device may retrieve the data from memory and decode the data.
- encoding and decoding are performed by devices that do not communicate with each other but merely encode data to and/or retrieve data from memory and decode data.
- the encoding end 10 includes a video source 120 , a video encoder 100 and an output interface 140 .
- output interface 140 may include a regulator/demodulator (modem) and/or a transmitter.
- Video source 120 may include a video capture device (e.g., a video camera), a video archive containing previously captured video data, a video feed interface to receive video data from a video content provider, and/or a computer for generating video data. graphics system, or a combination of these sources of video data.
- Video encoder 100 may encode video data from video source 120 .
- encoder 10 transmits the encoded video data directly to decoder 20 via output interface 140 .
- the encoded video data may also be stored on the storage device 40 for later access by the decoder 20 for decoding and/or playback.
- the decoding terminal 20 includes an input interface 240 , a video decoder 200 and a display device 220 .
- input interface 240 includes a receiver and/or modem.
- Input interface 240 may receive encoded video data via link 30 and/or from storage device 40.
- the display device 220 may be integrated with the decoding terminal 20 or may be external to the decoding terminal 20 . Generally speaking, display device 220 displays decoded video data.
- the display device 220 may include a variety of display devices, such as a liquid crystal display (LCD), a plasma display, an organic light-emitting diode (OLED) display, or other types of display devices.
- LCD liquid crystal display
- OLED organic light-emitting diode
- video encoder 100 and video decoder 200 may each be integrated with an audio encoder and decoder, and may include appropriate multiplexer-demultiplexer units or other hardware and software to handle the encoding of both audio and video in a common data stream or in separate data streams.
- the demultiplexer (MUX-DEMUX) unit may conform to the ITU H.223 multiplexer protocol, or other protocols such as user datagram protocol (UDP), if applicable.
- Video encoder 100 and video decoder 200 may each be implemented as any of a variety of circuits, such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (application-specific integrated circuit, ASIC), field programmable gate array (field programmable gate array, FPGA), discrete logic, hardware, or any combination thereof. If the present application is implemented partly in software, the device may store instructions for the software in a suitable non-volatile computer-readable storage medium, and the instructions may be executed in hardware using one or more processors Thereby implementing the technology of this application. Any of the foregoing (including hardware, software, a combination of hardware and software, etc.) may be considered one or more processors. Each of video encoder 100 and video decoder 200 may be included in one or more encoders or decoders, either of which may be integrated as a combined encoder in the respective device. / part of the decoder (codec).
- codec part of the decoder
- Video encoder 100 may be generally referred to herein as "signaling” or “transmitting” certain information to another device, such as video decoder 200 .
- the term “signaling” or “transmission” may generally refer to the transmission of syntax elements and/or other data used to decode compressed video data. This transfer can occur in real time or near real time. Alternatively, this communication may occur over a period of time, such as when encoding to store the syntax elements in the encoded codestream to a computer-readable storage medium, and the decoding device may then occur after the syntax elements are stored to such media. retrieve the syntax element at any time.
- a video sequence usually contains a series of video frames or images.
- a group of pictures exemplarily includes a series, one or more video images.
- the GOP can be in the header information of the GOP, in one or more of the images
- the header information or elsewhere contains syntax data describing the number of images contained in the GOP.
- Each slice of an image may contain slice syntax data describing the encoding mode of the corresponding image.
- Video encoder 100 typically operates on video blocks within individual video slices in order to encode video data. Video blocks may correspond to decoding nodes within a CU. Video blocks may be of fixed or varying size, and may differ in size depending on the specified decoding standard.
- video encoder 100 may scan the quantized transform coefficients using a predefined scan order to produce serialized vectors that may be entropy encoded. In other possible implementations, video encoder 100 may perform adaptive scanning. After scanning the quantized transform coefficients to form a one-dimensional vector, video encoder 100 may perform context-based adaptive variable-length code (CAVLC), context-adaptive binary arithmetic decoding (context-based adaptive variable-length code). based adaptive binary arithmetic coding (CABAC), syntax-based adaptive binary arithmetic coding (SBAC), probability interval partitioning entropy (PIPE) decoding or other entropy decoding methods. Entropy decodes a one-dimensional vector. Video encoder 100 may also entropy encode syntax elements associated with the encoded video data for use by video decoder 200 in decoding the video data.
- CAVLC context-based adaptive variable-length code
- CABAC based adaptive binary arithmetic coding
- SBAC syntax-based adaptive binary arithmetic
- video encoder 100 may assign context within the context model to symbols to be transmitted.
- the context may relate to whether the adjacent values of the symbol are non-zero.
- video encoder 100 may select a variable length code of symbols to be transmitted. Codewords in variable-length code (VLC) can be constructed such that relatively short codes correspond to more likely symbols, while longer codes correspond to less likely symbols. In this way, the use of VLC can achieve code rate savings relative to using equal length codewords for each symbol to be transmitted. Probabilities in CABAC can be determined based on the context assigned to the symbol.
- This application may refer to the image currently being decoded by the video decoder as the current image.
- the video encoding and decoding system provided by this application can be applied to video compression scenarios, such as artificial intelligence (artificial intelligence, AI) video encoding/decoding modules.
- artificial intelligence artificial intelligence, AI
- the video encoding and decoding system provided by this application can be used to store compressed video files for different services, such as data (image or video) storage of terminal albums, video surveillance, Huawei TM Cloud, etc.
- the video encoding and decoding system provided by this application can be used to transmit compressed video files, such as Huawei TM Cloud, live video, etc.
- the process of the output interface 140 sending the data stream to the outside can also be called push streaming, and the server cluster sends the data stream to the input interface.
- the process of sending data streams can also be called distribution.
- the encoding end 10 may use optical flow mapping to encode and compress the image frames.
- the decoding end 20 may also use optical flow mapping to decode and reconstruct image frames.
- optical flow represents the movement speed and direction of pixels in two adjacent frames of images.
- image domain optical flow represents the motion information between pixels in two adjacent frames of images.
- the motion information may include the motion speed and motion direction between pixels in two adjacent frames of images.
- Feature domain optical flow It represents the motion information between the feature maps of two adjacent image frames.
- the motion information indicates: the motion speed and direction of motion between the feature maps of the image frame and the feature maps of adjacent image frames of the image frame.
- the adjacent image frame eg, image frame 2 of the aforementioned image frame (eg, image frame 1) may also be called a reference frame of the image frame 1.
- Optical flow has two directions in the time dimension: direction 1, the optical flow from the previous frame image to the next frame image; direction 2, the optical flow from the next frame image to the previous frame image.
- Optical flow in one direction is usually represented digitally, such as using a three-dimensional array [2, h, w].
- “2" means that the optical flow contains 2 channels, and the first channel represents the deflection of the image in the x direction. Shift direction and size, the second channel represents the shift direction and size of the image in the y direction.
- h is the height of the image
- w is the width of the image.
- a positive value means the object moves to the left, while a negative value means the object moves to the right; in the y direction, a positive value means the object moves upward, while a negative value means the object moves downward.
- Figure 2 is a mapping relationship between optical flow and color provided by this application.
- the arrow in (A) of Figure 2 shows the direction of the optical flow in the image, and the length of the arrow shows the size of the optical flow in the image; in Figure 2 (B) shows the color and brightness of the image determined based on the optical flow.
- the color represents the direction of the optical flow
- the brightness represents the size of the optical flow.
- the greater the brightness the greater the value corresponding to the optical flow.
- (B) in Figure 2 takes grayscale as an example to illustrate the color of the image, but this should not be understood as a limitation of the present application.
- the method of predicting the current frame through the reference frame and the optical flow between the two frames is called optical flow mapping (Warping), usually expressed as in, is the predicted frame, x t-1 is the reference frame, and v t is the optical flow from the reference frame to the predicted frame.
- Warping optical flow mapping
- the pixel points of the current frame correspond to the position of the reference frame. According to the corresponding By interpolating the position of the reference frame, an estimate of the pixel value of the current frame can be obtained.
- Optical flow mapping includes forward mapping (forward warping) and backward mapping (backward warping).
- Figure 3 is a schematic diagram of the optical flow mapping provided by this application:
- a in Figure 3 is forward mapping, which means that the decoder uses the optical flow between the previous frame image (image frame 1) and the previous frame. The flow predicts the next frame of image (Image Frame 2);
- B in Figure 3 is the backward mapping, which means that the decoder predicts the previous frame of image (Image Frame 2) based on the optical flow between the next frame of image (Image Frame 2) and the previous and next frames.
- Image frame 1 A in Figure 3 is forward mapping, which means that the decoder uses the optical flow between the previous frame image (image frame 1) and the previous frame. The flow predicts the next frame of image (Image Frame 2);
- B in Figure 3 is the backward mapping, which means that the decoder predicts the previous frame of image (Image Frame 2) based on the optical flow between the next frame of image (Image Frame 2) and the previous and next frames. Image frame 1).
- FIG. 4 is a schematic architectural diagram of the video compression system provided by this application.
- the video compression system 400 includes: a motion estimation module 410, an optical flow compression module 420, a motion compensation module 430, and a residual Compression (residual compression) module 440, entropy coding (entropy coding) module 450, and multi-frame feature fusion (multi-frame feature fusion) module 460.
- the motion estimation module 410, the optical flow compression module 420 and the motion compensation module 430 may be collectively referred to as variable motion compensation components.
- the image frame X t obtains the feature map F t of the image frame after feature extraction.
- the reference frame X t-1 obtains the feature map F t-1 of the reference frame after feature extraction.
- the reference frame X t-1 Can be stored in the reconstructed frame buffer (decoded frame buffer), which is used to provide data storage space for multiple frames.
- the motion estimation module 410 determines the feature domain optical flow O t between the image frame and the reference frame based on F t and F t-1 .
- the optical flow compression module 420 compresses O t to obtain the feature domain optical flow code stream O' t .
- the motion compensation module 430 performs feature prediction to determine the predicted feature map F t_pre corresponding to the image frame X t based on the feature map F t-1 of the reference frame and the decoded feature domain optical stream O ⁇ t .
- the residual compression module 440 outputs the compressed decoding residual R ⁇ t according to the obtained feature domain residual R t .
- the prediction feature map F t_pre and the decoding residual R ⁇ t can be used to determine the initial reconstruction feature F ⁇ t_intial of the image frame.
- the multi-frame feature fusion module 460 calculates the reconstructed feature maps corresponding to the initial reconstructed feature F ⁇ t_intial and multiple reference frames (F t-1_ref , F t-2_ref and F t-3_ref shown in Figure 4) , determine the final reconstructed feature map F ⁇ t_final of the image frame X t .
- the final reconstructed feature map F ⁇ t_final undergoes frame reconstruction to obtain the reconstructed frame X ⁇ t corresponding to the image frame X t .
- the entropy encoding module 450 is used to encode at least one of the feature domain optical flow O t and the feature domain residual R t to obtain a binary code stream.
- Figures 1 to 4 are only examples provided by this application.
- the encoding end 10, the decoding end 20, the video encoder 100, the video decoder 200 and the video encoding and decoding system may include more or fewer components or Unit, this application is not limited to this.
- Figure 5 is a schematic flow chart of an image coding method provided by this application.
- the image coding method can be applied to the video encoding and decoding system shown in Figure 1 or the video compression system shown in Figure 4.
- Example the image encoding method can be executed by the encoding end 10 or the video encoder 100.
- the encoding end 10 executes the image encoding method provided in this embodiment as an example for description.
- the image encoding method provided by this embodiment includes the following steps S510 to S560.
- S510 The encoding end obtains the feature map of image frame 1 and the first feature map of the reference frame of image frame 1.
- the image frame 1 and the reference frame may belong to the same GOP included in the video.
- the video includes one or more image frames.
- Image frame 1 can be any image frame in the video.
- the reference frame can be an image frame adjacent to image frame 1 in the video.
- the reference frame is in image frame 1 The previous adjacent image frame, or the reference frame is the adjacent image frame after image frame 1. It is worth noting that in some cases, image frame 1 may also be called the first image frame.
- the way in which the encoding end obtains the feature map may include, but is not limited to: the encoding end is implemented based on a neural network model. Assume that the size of each image frame in the video is [3, H, W], that is, the image frame has 3 channels, a height of H, and a width of W.
- Figure 6 is a schematic structural diagram of the feature extraction network provided by this application.
- the feature extraction network includes 1 convolutional layer (conv) and 3 residual blocks (residual block). , resblock) processing layer, the convolution kernel of the convolution layer is 64 ⁇ 5 ⁇ 5/2, and the convolution kernel size of the resblock processing layer is 64 ⁇ 3 ⁇ 3.
- the encoding end determines that the feature map corresponding to image frame 1 is f t , the first feature map of the reference frame is f t-1 .
- Figure 6 is only an example of a feature extraction method provided in this embodiment and should not be understood as a limitation of this application.
- the features shown in Figure 6 The parameters of each network layer in the extracted network can also be different.
- the encoding end may also use other methods to extract features of the image frame 1 and the reference frame, which is not limited in this application.
- S520 The encoding end obtains at least one optical flow set based on the feature map and the first feature map of image frame 1.
- the first optical flow set in the at least one optical flow set corresponds to the aforementioned first characteristic map.
- the first optical flow set may be the optical flow set 1 shown in FIG. 5 .
- the first optical flow set may include one or more feature domain optical flows v t .
- the feature domain optical flows v t are used to indicate the relationship between the feature map of the image frame 1 (or the first image frame) and the first feature map. movement information. This motion information can be used to indicate motion information and motion direction between the feature map of image frame 1 and the first feature map.
- the process of obtaining the optical flow set by the encoding end is actually an optical flow estimation process.
- the optical flow estimation process can be implemented by the motion estimation module 410 shown in Figure 4. This operation The motion estimation module 410 can be supported by the encoding end.
- the encoding end may utilize an optical flow estimation network to determine the aforementioned optical flow set.
- Figure 7 is a schematic structural diagram of an optical flow estimation network provided by this application.
- the optical flow estimation network includes an upsampling network and a downsampling network, where the upsampling network includes three network layers 1 , Network layer 1 includes: 3 residual block processing layers (convolution kernel is 64 ⁇ 3 ⁇ 3) and 1 convolution layer (convolution kernel is 64 ⁇ 5 ⁇ 5/2).
- the downsampling network includes 3 network layers 2, which in turn include: 1 convolution layer (convolution kernel is 64 ⁇ 5 ⁇ 5/2) and 3 residual block processing layers (convolution kernel is 64 ⁇ 3 ⁇ 3).
- a residual block processing layer includes: a convolution layer (convolution kernel is 64 ⁇ 3 ⁇ 3), an activation layer and a convolution layer (convolution kernel is 64 ⁇ 3 ⁇ 3).
- the activation layer may refer to a linear rectified layer (rectified linear units layer, ReLU), or a parametric rectified linear unit (PReLU), etc.
- Figure 7 is only an example of the optical flow estimation network provided by the embodiment of the present application and should not be understood as a limitation of the present application.
- the size of the convolution kernel of each network layer, the number of feature map channels of the input network layer, The downsampling position, number of convolutional layers, and network activation layers can all be adjusted.
- the optical flow estimation network can also use more complex network structures.
- S530 The encoding end performs feature domain optical flow coding on the optical flow set obtained in S520 to obtain a feature domain optical flow code stream.
- the feature domain optical flow code stream includes the code stream corresponding to optical flow set 1.
- the feature domain optical stream code stream may be a binary file, or the feature domain optical stream code stream may also be other types of files that comply with the multimedia transmission protocol, without limitation.
- S540 The encoding end processes the first feature map based on the feature domain optical flow included in the optical flow set 1, and obtains one or more intermediate feature maps corresponding to the first feature map.
- a feasible processing method is provided here.
- the encoding end performs optical flow mapping on the first feature map based on the first feature domain optical flow (v 1 ). (warping), obtain the intermediate feature map corresponding to the optical flow in the first feature domain. It is worth noting that warping a feature domain optical flow with the first feature map will obtain an intermediate feature map.
- the number of intermediate feature maps corresponding to the first feature map is consistent with the number of feature domain optical flows included in optical flow set 1.
- Figure 8 is a schematic process diagram of optical flow mapping and feature fusion provided by this application.
- Optical flow set 1 contains multiple feature domain optical flows, such as v 1 to v m , and m is a positive integer.
- the encoding end warps each feature domain optical flow included in the optical flow set 1 with the first feature map f t to obtain m intermediate feature maps.
- S550 The encoding end fuses one or more intermediate feature maps to obtain the first predicted feature map of the first image frame.
- the decoder obtains multiple image domain optical flows between image frames and reference frames, and obtains multiple images based on multiple image domain optical flows and reference frames, thereby fusing multiple images to obtain image frame correspondence.
- the target image Therefore, the decoder decodes one frame of image and needs to predict the pixel values of multiple images and fuse the multiple images to obtain the target image. This results in large computing resources required for image decoding.
- the decoder uses image domain optical flow Decoding video is less efficient.
- the encoding end inputs the aforementioned one or more intermediate feature maps into the feature fusion model to obtain the first predicted feature map.
- the feature fusion model includes a convolutional network layer, which is used to fuse intermediate feature maps.
- the decoder uses a feature fusion model to fuse multiple intermediate feature maps.
- the decoder uses a feature fusion model to fuse multiple intermediate feature maps.
- the prediction feature map obtained from the intermediate feature map decodes the image frame, that is, the decoder only needs to predict the pixel value of the image position indicated by the prediction feature map in the image based on the prediction feature map, and does not need to predict all pixels of multiple images. value, reducing the computing resources required for image decoding and improving the efficiency of image decoding.
- the encoding end obtains one or more weights of the aforementioned one or more intermediate feature maps, and processes the one or more weights based on the one or more weights.
- the intermediate feature maps corresponding to respective values, and the first predicted feature map is obtained by adding all processed intermediate feature maps.
- the weight is used to indicate the weight of the intermediate feature map in the first predicted feature map.
- the weight of each intermediate feature map can be different.
- the encoding end can set different weights for the intermediate feature maps according to the requirements of image encoding. , For example, if the images corresponding to some intermediate feature maps are relatively blurry, the weights of the intermediate feature maps corresponding to these blurred images are reduced, thereby improving the clarity of the first image.
- the weight may refer to the mask value corresponding to each optical flow.
- the optical flow set 1 includes 9 feature domain optical flows.
- the weights of each feature domain optical flow are as follows: For example, the size of the optical flow in the feature domain is [2, H/s, W/s], and the size of the mask is [1, H/s, W/s].
- the encoding end fuses the intermediate feature maps corresponding to the optical flows in these nine feature domains to obtain the first predicted feature map corresponding to the first feature map.
- the feature map fusion in this application can only adopt the above two methods.
- the encoding end can warping each feature domain optical flow included in the optical flow set 1 with the first feature map f t respectively.
- Four intermediate feature maps are obtained; furthermore, the encoding end performs feature fusion on the four intermediate feature maps to obtain the first predicted feature map.
- the image encoding method provided in this embodiment also includes the following step S560.
- S560 The encoding end encodes the residual corresponding to image frame 1 according to the first prediction feature map to obtain a residual code stream.
- the residual code stream of the image area corresponding to the first prediction feature map and the feature domain optical flow code stream determined in S530 can be collectively referred to as the code stream corresponding to image frame 1 (or the first image frame).
- the code stream includes: a feature domain optical flow code stream corresponding to the aforementioned optical flow set and a residual code stream of the image area corresponding to the first prediction feature map.
- the encoding end determines a set of feature domain optical flows ( As the aforementioned optical flow set 1), and process the first feature map of the reference frame to obtain one or more intermediate feature maps; secondly, obtain one or more intermediate feature maps corresponding to the first feature map at the encoding end Finally, the one or more intermediate feature maps are fused to obtain the first prediction feature map corresponding to the first image frame; finally, the encoding end encodes the first image frame according to the first prediction feature map to obtain a code stream.
- a set of feature domain optical flows As the aforementioned optical flow set 1
- the encoding end encodes the image frame according to the feature domain optical flow, which reduces the coding error caused by the image domain optical flow between two adjacent frames and improves the coding error.
- Image encoding accuracy Moreover, the encoding end processes the feature maps of the reference frame based on multiple feature domain optical flows to obtain multiple intermediate feature maps, and fuses the multiple intermediate feature maps to determine the predicted feature map of the image frame.
- the encoding end combines the optical flow set and the reference frame.
- the multiple intermediate feature maps determined by the feature map are fused to obtain the prediction feature map of the image frame.
- the prediction feature map contains more image information, so that when the encoding end encodes the image frame based on the prediction feature map obtained by fusion, This avoids the problem that a single intermediate feature map is difficult to accurately express the first image, and improves the accuracy of image coding and image quality (such as image clarity, etc.).
- the reference frame may correspond to multiple feature maps, such as the aforementioned first feature map and the second feature map.
- the first feature map and the second feature map may refer to different channels among the multiple channels included in the reference frame. Two feature maps of the same channel.
- a set of feature maps corresponding to the reference frame may correspond to an optical flow set.
- the aforementioned optical flow set 1 may also correspond to the second feature map.
- the encoding end can also process the second feature map based on the feature domain optical flow included in the optical flow set 1 to obtain the second feature map.
- One or more intermediate feature maps corresponding to the feature map (f ⁇ t-1 to f ⁇ tm as shown in Figure 8, m is a positive integer).
- the encoding end also fuses one or more intermediate feature maps corresponding to the second feature map to obtain the second predicted feature map of the first image frame.
- the encoding end encoding the first image frame may include the following content: the encoding end bases the first prediction feature map and the second feature map on The second prediction feature map is used to encode the first image frame to obtain a code stream.
- the code stream may include the residual code stream of the image area corresponding to the first feature map and the second feature map in the first image frame, and the feature domain optical flow code stream corresponding to the aforementioned optical flow set 1.
- multiple feature maps (or a set of feature maps, or a feature map group, etc.) of the reference frame may correspond to an optical flow set (or a set of feature domain optical flows). For example, for Multiple feature maps belonging to the same group share a set of feature domain flows.
- the encoding end processes the feature map according to a set of feature domain optical flows corresponding to the feature map, thereby obtaining the intermediate feature map corresponding to the feature map. Further, the encoding end will The intermediate feature map corresponding to the feature map is fused to obtain the predicted feature map corresponding to the feature map. Finally, the encoding end encodes the image frame according to the prediction feature maps corresponding to all feature maps of the reference frame to obtain the target code stream.
- the encoding end can divide the multiple feature maps into one group or multiple groups.
- the feature maps belonging to the same group share an optical flow set, and
- the intermediate feature map corresponding to the feature map is fused to obtain the predicted feature map, which avoids the problem that when the reference frame or image frame has more information, the code stream obtained by the encoding end based on the feature map encoding has more redundancy and lower accuracy, and improves improve the accuracy of image coding. It is worth noting that the number of channels of feature maps belonging to different groups can be different.
- the encoding end can use different feature fusion methods for different feature maps of the reference frame.
- the intermediate feature map corresponding to the feature map of the reference frame is fused to obtain a predicted feature map corresponding to the feature map of the reference frame.
- the reference frame corresponds to feature map 1 and feature map 2.
- the encoding end inputs multiple intermediate feature maps corresponding to feature map 1 into the feature fusion model to determine the predicted feature map corresponding to feature map 1; the encoding end obtains feature map 2.
- the corresponding weight of each intermediate feature map is processed, and the corresponding intermediate feature map is processed according to these weights to obtain the predicted feature map corresponding to feature map 2.
- the encoding end can set the confidence level for the predicted feature map output by each feature fusion method, so as to meet the different coding requirements of the user.
- FIG. 9 is a schematic flowchart 1 of the image decoding method provided by this application.
- the image decoding method can be applied to the video encoding and decoding system shown in Figure 1 or the video encoding and decoding system shown in Figure 4.
- Video compression system, exemplary image encoding The method may be executed by the decoding end 20 or the video decoder 200.
- the decoding end 20 executes the image decoding method provided in this embodiment as an example for description.
- the image decoding method provided by this embodiment includes the following steps S910 to S940.
- the decoding end parses the code stream to obtain at least one optical flow set.
- the at least one optical flow set includes a first optical flow set (optical flow set 1 as shown in Figure 9), the optical flow set 1 corresponds to the first feature map of the reference frame, and the optical flow set 1 includes one or more A feature domain optical flow, wherein any one of the one or more feature domain optical flows is used to indicate motion information between the feature map of the first image frame and the aforementioned first feature map.
- a first optical flow set optical flow set 1 as shown in Figure 9
- the optical flow set 1 corresponds to the first feature map of the reference frame
- the optical flow set 1 includes one or more A feature domain optical flow, wherein any one of the one or more feature domain optical flows is used to indicate motion information between the feature map of the first image frame and the aforementioned first feature map.
- the reference frame may refer to the image frame adjacent to the first image frame that has been decoded by the decoding end (the image frame before the first image frame, or the image frame after the first image frame), or,
- the reference frame is an image frame carried in the code stream that contains complete information of an image.
- the decoding end processes the first feature map based on the feature domain optical flow included in the optical flow set 1, and obtains one or more intermediate feature maps corresponding to the first feature map.
- the process by which the decoder processes the feature map of the reference frame based on the feature domain optical flow is also called feature alignment, feature prediction, or feature align, and this application is not limited to this.
- the decoding end can also process the first feature map in the same manner as the aforementioned S520, and obtain one or more features corresponding to the first feature map. an intermediate feature map.
- the decoder performs warping on the first feature map based on a set of feature domain optical flows.
- the decoder can interpolate based on the position of a set of feature domain optical flows corresponding to the reference frame to obtain the predicted value of the pixel value in the first image to avoid This eliminates the need for the decoder to predict all pixel values of the first image based on the image domain optical flow, reduces the amount of calculation required for image decoding at the decoder, and improves the efficiency of image decoding.
- S930 The decoder fuses one or more intermediate feature maps determined in S920 to obtain the first predicted feature map of the first image frame.
- the decoding end may input the one or more intermediate feature maps determined in S920 to the feature fusion model to obtain the first predicted feature map.
- the feature fusion model includes a convolutional network layer, which is used to fuse intermediate feature maps.
- the decoder uses a feature fusion model to fuse multiple intermediate feature maps.
- the decoder uses a feature fusion model to fuse multiple intermediate feature maps.
- the predicted feature map obtained from the intermediate feature map decodes the image frame, reducing the computing resources required for image decoding and improving the efficiency of image decoding.
- the decoder may also obtain one or more weights of the aforementioned one or more intermediate feature maps, where one intermediate feature map corresponds to one weight; and the decoder may process the intermediate feature map based on the weight of the intermediate feature map, and All processed intermediate feature maps are added to obtain the first predicted feature map.
- the weight is used to indicate the weight of the intermediate feature map in the first predicted feature map.
- the weight of each intermediate feature map can be different.
- the decoder can set different weights for the intermediate feature maps according to the requirements of image decoding.
- the weights of these intermediate feature maps are reduced, thereby improving the clarity of the first image.
- the weight of the intermediate feature map please refer to the relevant content of S550 and will not be described in detail here.
- the decoder decodes the first image frame according to the first prediction feature map determined in S930 to obtain the first image.
- the decoder can decode the first image frame according to the reconstructed feature map f_res to obtain the first image.
- the decoder determines the first feature of the reference frame based on a set of feature domain optical flows corresponding to the first image frame (such as the aforementioned first optical flow set).
- the image is processed to obtain one or more intermediate feature maps; secondly, after the decoder obtains one or more intermediate feature maps corresponding to the first feature map, the one or more intermediate feature maps are fused to obtain the first The first prediction feature map corresponding to the image frame; finally, the decoder decodes the first image frame according to the first prediction feature map to obtain the first image.
- the decoder decodes the image frame according to the feature domain optical flow, which reduces the decoding error caused by the image domain optical flow between two adjacent frames and improves improve the accuracy of image decoding.
- the decoding end processes the feature maps of the reference frame based on multiple feature domain optical flows to obtain multiple intermediate feature maps, and fuses the multiple intermediate feature maps to determine the predicted feature map of the image frame. That is, the decoding end combines the optical flow set and the reference frame.
- the multiple intermediate feature maps determined by the feature map are fused to obtain the predicted feature map of the image frame.
- the predicted feature map contains more image information, so that when the decoder decodes the image frame based on the predicted feature map obtained by fusion, This avoids the problem that a single intermediate feature map is difficult to accurately express the first image, and improves the accuracy of image decoding and image quality (such as image clarity, etc.).
- the encoding end can encode the feature map through a feature encoding network
- the decoding end can encode the code stream corresponding to the feature map through a feature decoding network.
- Figure 10 is the feature encoding network provided by this application.
- the feature encoding network includes 3 network layers 1.
- the network layer 1 includes: 3 residual block processing layers (convolution kernel is 64 ⁇ 3 ⁇ 3) and 1 convolution layer (convolution layer).
- the core is 64 ⁇ 5 ⁇ 5/2).
- the feature decoding network includes 3 network layers 2.
- Network layer 2 includes: 1 convolution layer (convolution kernel is 64 ⁇ 5 ⁇ 5/2) and 3 residual block processing layers (convolution kernel is 64 ⁇ 3 ⁇ 3).
- a residual block processing layer includes in sequence: a convolution layer (the convolution kernel is 64 ⁇ 3 ⁇ 3), an activation layer and a convolution layer (the convolution kernel is 64 ⁇ 3 ⁇ 3). 64 ⁇ 3 ⁇ 3).
- the activation layer may refer to the ReLU layer or PReLU, etc.
- Figure 10 is only an example of the feature encoding network and the feature decoding network provided by the embodiment of the present application, and should not be understood as a limitation of the present application.
- the convolution kernel size of each network layer, the feature map of the input network layer The number of channels, downsampling positions, number of convolutional layers, and network activation layers can all be adjusted. In some more complex scenes, the optical flow estimation network can also use more complex network structures.
- the reference frame may correspond to multiple feature maps, such as the aforementioned first feature map and the second feature map.
- a set of feature maps corresponding to the reference frame may correspond to an optical flow set.
- the aforementioned optical flow set 1 may also correspond to the second feature map of the reference frame.
- the decoder can also process the second feature map based on the feature domain optical flow included in the optical flow set 1 to obtain the second feature map.
- One or more intermediate feature maps corresponding to the feature map (f ⁇ t-1 to f ⁇ tm as shown in Figure 8, m is a positive integer).
- the decoder also fuses one or more intermediate feature maps corresponding to the second feature map to obtain the second predicted feature map of the first image frame.
- multiple feature maps (or a group of feature maps) of the reference frame can correspond to an optical flow set (or a set of feature domain optical flows).
- the decoding end processes the feature map according to a set of feature domain optical flows corresponding to the feature map, thereby obtaining an intermediate feature map corresponding to the feature map.
- the decoder end processes the intermediate feature map corresponding to the feature map. Fusion is performed to obtain the predicted feature map corresponding to the feature map.
- the decoder decodes the image frame according to the prediction feature maps corresponding to all feature maps of the reference frame to obtain the target image.
- the decoder can divide the multiple feature maps into one group or multiple groups.
- the feature maps belonging to the same group share an optical flow set, and
- the intermediate feature map corresponding to the feature map is fused to obtain the predicted feature map, which avoids the problem of low accuracy in reconstructing the image based on the feature map at the decoder when the reference frame or image frame has more information, and improves the accuracy of image decoding.
- FIG 11 is a schematic flow chart 2 of the image decoding method provided by this application.
- the image decoding method shown in Figure 11 can be compared with the image encoding method provided by the previous embodiment. Combined with the image decoding method, it can also be implemented separately.
- the decoder executes the image decoding method shown in Figure 11 as an example.
- the image decoding method provided by this embodiment includes the following steps S1110 to S1130.
- the decoder obtains the feature map of the first image.
- the decoder can obtain the feature map of the first image according to the feature extraction network shown in Figure 6, which will not be described again here.
- the decoder obtains an enhanced feature map based on the feature map of the first image, the first feature map of the reference frame, and the first predicted feature map.
- the decoder can use a feature fusion model to fuse multiple feature maps included in S1120 to obtain an enhanced feature map.
- the feature fusion model may refer to the feature fusion model provided by the aforementioned S550 or S930, or may refer to other models including a convolution layer.
- the convolution kernel of the convolution layer is 3 ⁇ 3, which is not limited in this application. .
- the decoder can also set different weights for the feature map of the first image, the first feature map of the reference frame, and the first prediction feature map, so as to fuse these multiple feature maps to obtain enhanced features. picture.
- the enhancement feature map may be used to determine the enhancement layer image of the first image.
- the video quality of the enhancement layer image is higher than the video quality of the first image.
- the video quality may refer to or include image signal-to-noise ratio (SNR), resolution (image resolution) and peak signal-to-noise ratio (SNR). Peak signal-to-noise ratio, PSNR) at least one.
- SNR image signal-to-noise ratio
- PSNR peak signal-to-noise ratio
- the image signal-to-noise ratio refers to the ratio of the signal mean of the image to the background standard deviation.
- the "signal mean” here generally refers to the grayscale mean of the image, and the background standard deviation can be used as the variance of the background signal value of the image.
- the variance of the background signal value refers to the noise power; for images, the greater the image signal-to-noise ratio, the better the image quality.
- Resolution is the number of pixels per unit area of a single frame image. The higher the resolution of the image, the better the quality of the image.
- PSNR is used to indicate the subjective quality of an image. The larger the PSNR, the better the quality of the image. For more information about SNR, resolution and PSNR, please refer to the relevant descriptions of the prior art, which will not be described again here.
- the decoder processes the first image according to the enhanced feature map to obtain the second image.
- the second image indicates the same content as the first image, but the second image is sharper than the first image.
- image sharpness may be indicated by peak signal-to-noise ratio (PSNR),
- PSNR peak signal-to-noise ratio
- the decoder can fuse the feature map of the first image, the first feature map and the first predicted feature map, and perform video enhancement processing on the first image based on the enhanced feature map obtained through the fusion to obtain better clarity. Excellent second image, thereby improving the decoded image clarity and image display effect.
- the decoder processes the first image according to the enhanced feature map to obtain the second image, including: the decoder obtains the enhancement layer image of the first image based on the enhanced feature map, and reconstructs the first image based on the enhancement layer image to obtain the second image.
- the enhancement layer image may refer to an image determined by the decoder based on the reference frame and the enhancement feature map.
- the decoder adds part or all of the information of the enhancement layer image to the first image to obtain the second image; or , the decoder uses the enhancement layer image as the reconstructed image of the first image, that is, the aforementioned second image.
- the decoder obtains multiple feature maps of the image at different stages, and obtains the enhanced feature map determined by these multiple feature maps, so as to reconstruct and enhance the first image based on the enhanced feature map. Improved image clarity and image display after decoding.
- the decoding end may reconstruct the first image based on an image reconstruction network, as shown in Figure 12.
- Figure 12 is a schematic structural diagram of the image reconstruction network provided by this application.
- the image The reconstruction network includes: 3 residual block processing layers (convolution kernel is 64 ⁇ 3 ⁇ 3) and 1 deconvolution layer (convolution kernel is 3 ⁇ 5 ⁇ 5/2), in which the residual block
- the processing layer can include 1 convolution layer (convolution kernel is 64 ⁇ 3 ⁇ 3), an activation layer and 1 convolution layer (convolution kernel is 64 ⁇ 3 ⁇ 3).
- the encoding end/decoding end predicts the feature map from an intermediate feature map set determined by performing optical flow mapping on the feature map of the reference frame.
- the encoding end/decoding end can also predict the feature map based on the deformable Networks (deformable convolutional networks, DCN) process feature maps.
- DCN is implemented based on convolution. The following is a mathematical expression of convolution:
- n is the convolution kernel size
- w is the convolution kernel weight
- F is the input feature map
- p is the convolution position
- p k is the enumeration value of the position relative to p in the convolution kernel.
- DCN is based on a network learning offset (offset), so that the convolution kernel is offset at the sampling point of the input feature map and concentrated on the region of interest (ROI) or target area. Its mathematical expression is:
- ⁇ p k is the offset relative to p k , so that the convolution sampling position becomes an irregular position.
- m(p k ) indicates that the mask value of position p k is the penalty term for position p k .
- the convolution operation of convolving the most central point through points in the neighborhood (a) is the common sampling method of 3x3 convolution kernel, (b) is after sampling deformable convolution plus offset changes in the sampling points, where (c) and (d) are special forms of deformable convolution.
- the multiple pixel points in (c) are scaled to obtain the pixel prediction value of the corresponding position in the target image, and as (c) The multiple pixels in d) are rotated and scaled to obtain the pixel prediction value of the corresponding position in the target image.
- the image encoding method and image decoding method provided by the embodiments of the present application can also be implemented by DCN. It can be understood that when the reference frame corresponds to multiple channel different feature maps, each feature map can use the same or different DCN processing methods, as shown in Figure 13, the four possible DCN processing methods.
- Figure 14 is a schematic diagram of the efficiency comparison provided by this application.
- Figure 14 provides two indicators to compare the technical solution provided by this application and common technology. These two The indicators are: PSNR and bits per pixel (BPP).
- BPP refers to the number of bits used to store each pixel. BPP is also used to indicate the resolution of the image.
- the BPP of the solution provided by the common technology is higher than the technical solution provided by the embodiment of the present application, that is, the storage occupied by the code stream generated by the common technology
- the space is larger and the network bandwidth required to transmit the code stream is larger; however, the PSNR of the solution provided by the general technology is lower than that of the technical solution provided by the embodiment of the present application. That is, in the technical solution provided by the embodiment of the present application, after the code stream is decoded
- the subjective quality of the video (or image) obtained is better.
- the end-to-end time of the technical solution to test the frame image is usually 0.3206 seconds (second, s).
- the technical solution test provided by the embodiment of the present application The end-to-end time of this frame image is 0.2188s. That is to say, during the video encoding and decoding process, the technical solution provided by the embodiment of the present application can not only save the bit rate (reflected by BPP), but also ensure the quality (reflected by PSNR), and also reduce the encoding and decoding time of a single frame image. delay, thereby improving the overall efficiency of video encoding and decoding.
- image encoding method and image decoding method can not only be applied to video encoding and decoding, video enhancement, video compression and other scenarios, but also can be applied to all video frame requirements such as video prediction, video frame insertion, and video analysis.
- the encoding end and the decoding end include corresponding hardware structures and/or software modules for performing each function.
- the units and method steps of each example described in conjunction with the embodiments disclosed in this application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed by hardware or computer software driving the hardware depends on the specific application scenarios and design constraints of the technical solution.
- Figure 15 is a schematic structural diagram of an image encoding device provided by the present application.
- the image encoding device can be used to implement the functions of the encoding end or video encoder in the above method embodiments, and therefore can also achieve the beneficial effects of the above image encoding method embodiments.
- the image encoding device may be the encoding end 10 or the video encoder 100 as shown in FIG. 1 , or may be a module (such as a chip) applied to the encoding end 10 or the video encoder 100 .
- the image encoding device 1500 includes: an acquisition unit 1510 , a processing unit 1520 , a fusion unit 1530 and an encoding unit 1540 .
- the image encoding device 1500 is used to implement the functions of the encoding end or video encoder in the method embodiments shown in FIG. 5 and FIG. 8 .
- the acquisition unit 1510 is used to perform S510
- the processing unit 1520 is used to perform S520 and S540
- the fusion unit 1530 is used to perform S550
- the encoding unit 1540 is used to perform S530 and S560.
- the processing unit 1520 and the fusion unit 1530 are used to implement the functions of optical flow mapping and feature fusion.
- FIG. 16 is a schematic structural diagram of the image decoding device provided by the present application.
- the image decoding device can be used to implement decoding in the above method embodiments.
- the function of the terminal or video decoder can also achieve the beneficial effects of the above image decoding method embodiments.
- the image decoding device may be the decoding terminal 20 or the video decoder 200 as shown in FIG. 1 , or may be a module (such as a chip) applied to the decoding terminal 20 or the video decoder 200 .
- the image decoding device 1600 includes: a code stream unit 1610, a processing unit 1620, a fusion unit 1630 and a decoding unit 1640.
- the image decoding device 1600 is used to implement the functions of the decoding end or video decoder in the method embodiments shown in FIG. 8 and FIG. 9 .
- the processing unit 1620 and the fusion unit 1630 are used to implement the functions of optical flow mapping and feature fusion.
- the code stream unit 1610 is used to perform S910
- the processing unit 1620 is used to perform S920
- the fusion unit 1630 is used to perform S930
- the decoding unit 1640 is used to perform S940 .
- the image decoding device 1600 may further include an enhancement unit configured to process the first image according to the enhanced feature map to obtain the second image.
- the enhancement unit is specifically configured to: obtain the enhancement layer image of the first image based on the enhancement feature map. And, reconstruct the first image based on the enhancement layer image to obtain the second image.
- the image encoding (or image decoding) device and its respective units may also be software modules.
- the software module is called by the processor to implement the above image encoding (or image decoding) method.
- the processor can be a central processing unit (CPU), an application-specific integrated circuit (ASIC) implementation, or a programmable logic device (PLD), which can be a complex program Logic device (complex programmable logical device, CPLD), field programmable gate array (field programmable gate array, FPGA), general array logic (generic array logic, GAL) or any combination thereof.
- image encoding (or image decoding) device For a more detailed description of the above image encoding (or image decoding) device, please refer to the embodiments shown in the preceding figures. The relevant descriptions are not repeated here. It can be understood that the image encoding (or image decoding) device shown in the foregoing drawings is only an example provided by this embodiment.
- the image encoding (or image decoding) device may vary depending on the image encoding (or image decoding) process or service. More or fewer units may be included, which is not limited by this application.
- the hardware may be implemented by a processor or a chip.
- the chip includes interface circuit and control circuit.
- the interface circuit is used to receive data from other devices other than the processor and transmit it to the control circuit, or to send data from the control circuit to other devices other than the processor.
- control circuit is used to implement any of the possible implementation methods in the above embodiments through logic circuits or executing code instructions.
- the beneficial effects can be found in the description of any aspect in the above embodiments and will not be described again here.
- processor in the embodiments of the present application may be a CPU, a neural processing unit (NPU) or a graphics processing unit (GPU), or other general-purpose processors, digital signal Processor (digital signal processor, DSP), ASIC, FPGA or other programmable logic devices, transistor logic devices, hardware components or any combination thereof.
- a general-purpose processor can be a microprocessor or any conventional processor.
- the method steps in the embodiments of the present application can be implemented by means of hardware.
- the hardware is an image decoding device, as shown in Figure 17.
- Figure 17 is a schematic structural diagram of the image decoding device provided by the present application.
- the image decoding device 1700 includes a memory 1710 and at least one processor 1720.
- the processor 1720 can implement the image encoding method and the image decoding method provided in the above embodiments.
- the memory 1710 is used to store the correspondence between the above image encoding method and the image decoding method. software instructions.
- the image decoding device 1700 may refer to a chip or chip system encapsulating one or more processors 1720 .
- the processor 1720 included in the image decoding device 1700 executes the steps of the above method and its possible sub-steps.
- the image decoding device 1700 may also include a communication interface 1730, which may be used to send and receive data.
- the communication interface 1730 is used to receive the user's encoding request, decoding request, or send code stream, receive code stream, etc.
- the communication interface 1730, the processor 1720, and the memory 1710 can be connected through a bus 1740.
- the bus 1740 can be divided into an address bus, a data bus, a control bus, etc.
- the image decoding device 1700 can also perform the functions of the image encoding device 1500 shown in FIG. 15 and the functions of the image decoding device 1600 shown in FIG. 16 , which will not be described again here.
- the image decoding device 1700 provided in this embodiment may be a server, a personal computer, or other image decoding device 1700 with data processing functions, which is not limited by this application.
- the image decoding device 1700 may be the aforementioned encoding end 10 (or video encoder 100), or the decoding end 20 (or video decoder 200).
- the image decoding device 1700 may also have the functions of the aforementioned encoding end 10 and decoding end 20.
- the image decoding device 1700 refers to a video codec system (or video compression system) with a video codec function. ).
- the method steps in the embodiments of the present application can also be implemented by a processor executing software instructions.
- Software instructions can be composed of corresponding software modules.
- Software modules can be stored in random access memory (random access memory, RAM), flash memory, read-only memory (read-only memory, ROM), programmable read-only memory (programmable ROM) , PROM), erasable programmable read-only memory (erasable PROM, EPROM), electrically erasable programmable read-only memory (electrically EPROM, EEPROM), register, hard disk, mobile hard disk, CD-ROM or other well-known in the art any other form of storage media.
- An exemplary storage medium is coupled to the processor such that the processor can read information from the storage medium and write information to the storage medium.
- the storage medium may also be an integral part of the processor.
- the processor and storage media may be located in an ASIC. Additionally, the ASIC can be located in network equipment or terminal equipment. Of course, the processor and the storage medium can also exist as discrete components in network equipment or terminal equipment.
- this application also provides a computer-readable storage medium, which stores a code stream obtained according to the image encoding method provided in any of the foregoing embodiments.
- the computer-readable storage medium may be, but is not limited to: RAM, flash memory, ROM, PROM, EPROM, EEPROM, register, hard disk, removable hard disk, CD-ROM or any other form of storage medium well known in the art.
- the computer program product includes one or more computer programs or instructions.
- the computer may be a general purpose computer, a special purpose computer, a computer network, a network device, a user equipment, or other programmable device.
- the computer program or instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another.
- the computer program or instructions may be transmitted from a website, computer, A server or data center transmits via wired or wireless means to another website site, computer, server, or data center.
- the computer-readable storage medium may be any available medium that can be accessed by a computer or a data storage device such as a server or data center that integrates one or more available media.
- the available media may be magnetic media, such as floppy disks, hard disks, and magnetic tapes; they may also be optical media, such as digital video discs (DVDs); they may also be semiconductor media, such as solid state drives (solid state drives). ,SSD).
- “at least one” refers to one or more, and “plurality” refers to two or more.
- “And/or” describes the association of associated objects, indicating that there can be three relationships, for example, A and/or B, which can mean: A exists alone, A and B exist simultaneously, and B exists alone, where A, B can be singular or plural.
- the character “/” generally indicates that the related objects before and after are an “or” relationship; in the formula of this application, the character “/” indicates that the related objects before and after are a kind of "division” Relationship.
Landscapes
- Engineering & Computer Science (AREA)
- Multimedia (AREA)
- Signal Processing (AREA)
- Theoretical Computer Science (AREA)
- Computer Vision & Pattern Recognition (AREA)
- Physics & Mathematics (AREA)
- General Physics & Mathematics (AREA)
- Artificial Intelligence (AREA)
- Health & Medical Sciences (AREA)
- Computing Systems (AREA)
- Databases & Information Systems (AREA)
- Evolutionary Computation (AREA)
- General Health & Medical Sciences (AREA)
- Medical Informatics (AREA)
- Software Systems (AREA)
- Compression Or Coding Systems Of Tv Signals (AREA)
Abstract
Description
Claims (30)
- 一种图像解码方法,其特征在于,所述方法包括:解析码流获得至少一个光流集;其中,所述至少一个光流集包括第一光流集,所述第一光流集包括一个或多个特征域光流,所述第一光流集对应第一图像帧的参考帧的第一特征图,所述一个或多个特征域光流中的任一个特征域光流用于指示第一图像帧的特征图与所述第一特征图之间的运动信息;基于所述一个或多个特征域光流处理所述第一特征图,获得所述第一特征图对应的一个或多个中间特征图;融合所述一个或多个中间特征图,获得所述第一图像帧的第一预测特征图;根据所述第一预测特征图解码所述第一图像帧获得第一图像。
- 根据权利要求1所述的方法,其特征在于,所述基于所述一个或多个特征域光流处理所述第一特征图,获得所述第一特征图对应的一个或多个中间特征图,包括:解析所述码流获得所述第一特征图;基于第一特征域光流对所述第一特征图进行光流映射warping,获得所述第一特征域光流对应的中间特征图,其中,所述第一特征域光流为所述一个或多个特征域光流中的任一个。
- 根据权利要求1或2所述的方法,其特征在于,所述第一光流集还对应所述参考帧的第二特征图;在所述根据所述第一预测特征图解码所述第一图像帧获得第一图像之前,所述方法还包括:基于所述一个或多个特征域光流处理所述第二特征图,获得所述第二特征图对应的一个或多个中间特征图;融合所述第二特征图对应的一个或多个中间特征图,获得所述第一图像帧的第二预测特征图;所述根据所述第一预测特征图解码所述第一图像帧获得第一图像,包括:根据所述第一预测特征图和所述第二预测特征图,解码所述第一图像帧获得第一图像。
- 根据权利要求1至3中任一项所述的方法,其特征在于,所述融合所述一个或多个中间特征图,获得所述第一图像帧的第一预测特征图,包括:获取所述一个或多个中间特征图的一个或多个权值,其中,一个中间特征图对应一个权值,所述权值用于指示所述中间特征图在所述第一预测特征图中所占的权重;基于所述一个或多个权值处理与所述一个或多个权值各自对应的中间特征图,并将所有处理后的中间特征图相加获得所述第一预测特征图。
- 根据权利要求1至3中任一项所述的方法,其特征在于,所述融合所述一个或多个中间特征图,获得所述第一图像帧的第一预测特征图,包括:将所述第一特征图对应的一个或多个中间特征图输入特征融合模型,获得第一预测特征图;所述特征融合模型包括卷积网络层,所述卷积网络层用于融合中间特征图。
- 根据权利要求1至5中任一项所述的方法,其特征在于,所述方法还包括:获取所述第一图像的特征图;根据所述第一图像的特征图、所述第一特征图和所述第一预测特征图,获得增强特征 图;根据所述增强特征图处理所述第一图像,获得第二图像;所述第二图像的清晰度高于所述第一图像的清晰度。
- 根据权利要求6所述的方法,其特征在于,所述根据所述增强特征图处理所述第一图像,获得第二图像,包括:依据所述增强特征图获得所述第一图像的增强层图像;基于所述增强层图像重构所述第一图像,获得所述第二图像。
- 一种图像编码方法,其特征在于,所述方法包括:获取第一图像帧的特征图,以及所述第一图像帧的参考帧的第一特征图;根据所述第一图像帧的特征图和所述第一特征图,获得至少一个光流集;所述至少一个光流集包括第一光流集,所述第一光流集包括一个或多个特征域光流,所述第一光流集对应所述第一特征图,所述一个或多个特征域光流中的任一个特征域光流用于指示所述第一图像帧的特征图与所述第一特征图之间的运动信息;基于所述一个或多个特征域光流处理所述第一特征图,获得所述第一特征图对应的一个或多个中间特征图;融合所述一个或多个中间特征图,获得所述第一图像帧的第一预测特征图;根据所述第一预测特征图对所述第一图像帧进行编码获得码流。
- 根据权利要求8所述的方法,其特征在于,所述基于所述一个或多个特征域光流处理所述第一特征图,获得所述第一特征图对应的一个或多个中间特征图,包括:基于第一特征域光流对所述第一特征图进行光流映射warping,获得所述第一特征域光流对应的中间特征图,其中,所述第一特征域光流为所述一个或多个特征域光流中的任一个。
- 根据权利要求8或9所述的方法,其特征在于,所述第一光流集还对应所述参考帧的第二特征图;在所述根据所述第一预测特征图对所述第一图像帧进行编码获得码流之前,所述方法还包括:基于所述一个或多个特征域光流处理所述第二特征图,获得所述第二特征图对应的一个或多个中间特征图;融合所述第二特征图对应的一个或多个中间特征图,获得所述第一图像帧的第二预测特征图;所述根据所述第一预测特征图对所述第一图像帧进行编码获得码流,包括:根据所述第一预测特征图和所述第二预测特征图,编码所述第一图像帧获得码流。
- 根据权利要求8至10中任一项所述的方法,其特征在于,所述融合所述一个或多个中间特征图,获得所述第一图像帧的第一预测特征图,包括:获取所述一个或多个中间特征图的一个或多个权值,其中,一个中间特征图对应一个权值,所述权值用于指示所述中间特征图在所述第一预测特征图中所占的权重;基于所述一个或多个权值处理与所述一个或多个权值各自对应的中间特征图,并将所有处理后的中间特征图相加获得所述第一预测特征图。
- 根据权利要求8至11中任一项所述的方法,其特征在于,所述融合所述一个或多 个中间特征图,获得所述第一图像帧的第一预测特征图,包括:将所述第一特征图对应的一个或多个中间特征图输入特征融合模型,获得第一预测特征图;所述特征融合模型包括卷积网络层,所述卷积网络层用于融合中间特征图。
- 一种图像解码装置,其特征在于,所述装置包括:码流单元,用于解析码流获得至少一个光流集;所述至少一个光流集包括第一光流集,所述第一光流集包括一个或多个特征域光流,其中的第一光流集对应第一图像帧的参考帧的第一特征图,所述一个或多个特征域光流中的任一个特征域光流用于指示所述第一图像帧的特征图与所述第一特征图之间的运动信息;处理单元,用于基于所述一个或多个特征域光流处理所述第一特征图,获得所述第一特征图对应的一个或多个中间特征图;融合单元,用于融合所述一个或多个中间特征图,获得所述第一图像帧的第一预测特征图;解码单元,用于根据所述第一预测特征图解码所述第一图像帧获得第一图像。
- 根据权利要求13所述的装置,其特征在于,所述处理单元,具体用于:解析所述码流获得所述第一特征图;以及,基于第一特征域光流对所述第一特征图进行光流映射warping,获得所述第一特征域光流对应的中间特征图;所述第一特征域光流为所述一个或多个特征域光流中的任一个。
- 根据权利要求13或14所述的装置,其特征在于,所述第一光流集还对应所述参考帧的第二特征图;所述处理单元,还用于基于所述一个或多个特征域光流处理所述第二特征图,获得所述第二特征图对应的一个或多个中间特征图;所述融合单元,还用于融合所述第二特征图对应的一个或多个中间特征图,获得所述第一图像帧的第二预测特征图;所述解码单元,具体用于根据所述第一预测特征图和所述第二预测特征图,解码所述第一图像帧获得第一图像。
- 根据权利要求13至15中任一项所述的装置,其特征在于,所述融合单元,具体用于:获取所述一个或多个中间特征图的一个或多个权值,其中,一个中间特征图对应一个权值,所述权值用于指示所述中间特征图在所述第一预测特征图中所占的权重;以及,基于所述一个或多个权值处理与所述一个或多个权值各自对应的中间特征图,并将所有处理后的中间特征图相加获得所述第一预测特征图。
- 根据权利要求13至15中任一项所述的装置,其特征在于,所述融合单元,具体用于:将所述第一特征图对应的一个或多个中间特征图输入特征融合模型,获得第一预测特征图;所述特征融合模型包括卷积网络层,所述卷积网络层用于融合中间特征图。
- 根据权利要求13至17中任一项所述的装置,其特征在于,所述装置还包括:获取单元和增强单元;所述获取单元,用于获取所述第一图像的特征图;所述融合单元,还用于根据所述第一图像的特征图、所述第一特征图和所述第一预测特征图,获得增强特征图;所述增强单元,用于根据所述增强特征图处理所述第一图像,获得第二图像;所述第 二图像的清晰度高于所述第一图像的清晰度。
- 根据权利要求18所述的装置,其特征在于,所述增强单元,具体用于:依据所述增强特征图获得所述第一图像的增强层图像;以及,基于所述增强层图像重构所述第一图像,获得所述第二图像。
- 一种图像编码装置,其特征在于,所述装置包括:获取单元,用于获取第一图像帧的特征图,以及所述第一图像帧的参考帧的第一特征图;处理单元,用于根据所述第一图像帧的特征图和所述第一特征图,获得至少一个光流集;所述至少一个光流集包括第一光流集,所述第一光流集包括一个或多个特征域光流,所述第一光流集对应所述第一特征图,所述一个或多个特征域光流中的任一个特征域光流用于指示所述第一图像帧的特征图与所述第一特征图之间的运动信息;所述处理单元,还用于基于所述一个或多个特征域光流处理所述第一特征图,获得所述第一特征图对应的一个或多个中间特征图;融合单元,用于融合所述一个或多个中间特征图,获得所述第一图像帧的第一预测特征图;编码单元,用于根据所述第一预测特征图对所述第一图像帧进行编码获得码流。
- 根据权利要求20所述的装置,其特征在于,所述处理单元,具体用于:基于第一特征域光流对所述第一特征图进行光流映射warping,获得所述第一特征域光流对应的中间特征图,所述第一特征域光流为所述一个或多个特征域光流中的任一个。
- 根据权利要求20或21所述的装置,其特征在于,所述第一光流集还对应所述参考帧的第二特征图;所述处理单元,还用于基于所述一个或多个特征域光流处理所述第二特征图,获得所述第二特征图对应的一个或多个中间特征图;所述融合单元,还用于融合所述第二特征图对应的一个或多个中间特征图,获得所述第一图像帧的第二预测特征图;所述编码单元,具体用于根据所述第一预测特征图和所述第二预测特征图,编码所述第一图像帧获得码流。
- 根据权利要求20至22中任一项所述的装置,其特征在于,所述融合单元,具体用于:获得所述一个或多个中间特征图的一个或多个权值,其中,一个中间特征图对应一个权值,所述权值用于指示所述中间特征图在所述第一预测特征图中所占的权重;以及,基于所述一个或多个权值处理与所述一个或多个权值各自对应的中间特征图,并将所有处理后的中间特征图相加获得所述第一预测特征图。
- 根据权利要求20至22中任一项所述的装置,其特征在于,所述融合单元,具体用于:将所述第一特征图对应的一个或多个中间特征图输入特征融合模型,获得第一预测特征图;所述特征融合模型包括卷积网络层,所述卷积网络层用于融合中间特征图。
- 一种图像译码的装置,其特征在于,包括:存储器和处理器;所述存储器用于存储程序代码,所述处理器用于调用所述程序代码实现权利要求1至7中任意一项所述的方法,或者实现权利要求8至12中任意一项所述的方法。
- 一种计算机可读存储介质,其特征在于,所述存储介质中存储有计算机程序或指令, 当所述计算机程序或指令被电子设备执行时,实现权利要求1至7中任意一项所述的方法,或者实现权利要求8至12中任意一项所述的方法。
- 一种计算机可读存储介质,其特征在于,所述存储介质中存储有根据权利要求8-12任意一项所述的方法获取的码流。
- 一种视频编解码系统,其特征在于,包括:编码端和解码端;所述编码端用于执行权利要求1至7中任意一项所述的方法;所述解码端用于执行权利要求8至12中任意一项所述的方法。
- 一种计算机程序产品,其特征在于,当所述计算机程序产品在电子设备上运行时,所述电子设备实现权利要求1至7中任意一项所述的方法,和/或,所述电子设备实现权利要求8至12中任意一项所述的方法。
- 一种芯片,其特征在于,包括:控制电路和接口电路;所述接口电路用于:接收来自所述芯片之外的其它设备的信号并传输至处理器,或将来自所述控制电路的信号发送给所述芯片之外的其它设备;所述控制电路用于:通过逻辑电路或执行代码指令来实现权利要求1至7中任意一项所述的方法,或者实现权利要求8至12中任意一项所述的方法。
Priority Applications (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| EP23787366.6A EP4498678A4 (en) | 2022-04-15 | 2023-01-12 | IMAGE DECODING METHOD AND APPARATUS, AND IMAGE ENCODING METHOD AND APPARATUS |
| US18/914,881 US20250037317A1 (en) | 2022-04-15 | 2024-10-14 | Image Decoding Method, Image Encoding Method, and Apparatus |
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN202210397258.4 | 2022-04-15 | ||
| CN202210397258.4A CN116962706B (zh) | 2022-04-15 | 2022-04-15 | 图像解码方法、编码方法及装置 |
Related Child Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| US18/914,881 Continuation US20250037317A1 (en) | 2022-04-15 | 2024-10-14 | Image Decoding Method, Image Encoding Method, and Apparatus |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2023197717A1 true WO2023197717A1 (zh) | 2023-10-19 |
Family
ID=88328782
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2023/071928 Ceased WO2023197717A1 (zh) | 2022-04-15 | 2023-01-12 | 一种图像解码方法、编码方法及装置 |
Country Status (4)
| Country | Link |
|---|---|
| US (1) | US20250037317A1 (zh) |
| EP (1) | EP4498678A4 (zh) |
| CN (1) | CN116962706B (zh) |
| WO (1) | WO2023197717A1 (zh) |
Citations (6)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20180098066A1 (en) * | 2016-10-05 | 2018-04-05 | Qualcomm Incorporated | Systems and methods of switching interpolation filters |
| CN110913218A (zh) * | 2019-11-29 | 2020-03-24 | 合肥图鸭信息科技有限公司 | 一种视频帧预测方法、装置及终端设备 |
| US20210314474A1 (en) * | 2020-04-01 | 2021-10-07 | Samsung Electronics Co., Ltd. | System and method for motion warping using multi-exposure frames |
| CN113542651A (zh) * | 2021-05-28 | 2021-10-22 | 北京迈格威科技有限公司 | 模型训练方法、视频插帧方法及对应装置 |
| CN113949883A (zh) * | 2020-07-16 | 2022-01-18 | 武汉Tcl集团工业研究院有限公司 | 一种双向预测帧的编码方法、解码方法以及编解码系统 |
| CN114339219A (zh) * | 2021-12-31 | 2022-04-12 | 浙江大华技术股份有限公司 | 帧间预测方法、装置、编解码方法、编解码器及电子设备 |
Family Cites Families (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN110163370B (zh) * | 2019-05-24 | 2021-09-17 | 上海肇观电子科技有限公司 | 深度神经网络的压缩方法、芯片、电子设备及介质 |
| CN110689558B (zh) * | 2019-09-30 | 2022-07-22 | 清华大学 | 多传感器图像增强方法及装置 |
| CN111083501A (zh) * | 2019-12-31 | 2020-04-28 | 合肥图鸭信息科技有限公司 | 一种视频帧重构方法、装置及终端设备 |
-
2022
- 2022-04-15 CN CN202210397258.4A patent/CN116962706B/zh active Active
-
2023
- 2023-01-12 WO PCT/CN2023/071928 patent/WO2023197717A1/zh not_active Ceased
- 2023-01-12 EP EP23787366.6A patent/EP4498678A4/en active Pending
-
2024
- 2024-10-14 US US18/914,881 patent/US20250037317A1/en active Pending
Patent Citations (6)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20180098066A1 (en) * | 2016-10-05 | 2018-04-05 | Qualcomm Incorporated | Systems and methods of switching interpolation filters |
| CN110913218A (zh) * | 2019-11-29 | 2020-03-24 | 合肥图鸭信息科技有限公司 | 一种视频帧预测方法、装置及终端设备 |
| US20210314474A1 (en) * | 2020-04-01 | 2021-10-07 | Samsung Electronics Co., Ltd. | System and method for motion warping using multi-exposure frames |
| CN113949883A (zh) * | 2020-07-16 | 2022-01-18 | 武汉Tcl集团工业研究院有限公司 | 一种双向预测帧的编码方法、解码方法以及编解码系统 |
| CN113542651A (zh) * | 2021-05-28 | 2021-10-22 | 北京迈格威科技有限公司 | 模型训练方法、视频插帧方法及对应装置 |
| CN114339219A (zh) * | 2021-12-31 | 2022-04-12 | 浙江大华技术股份有限公司 | 帧间预测方法、装置、编解码方法、编解码器及电子设备 |
Non-Patent Citations (1)
| Title |
|---|
| See also references of EP4498678A4 |
Also Published As
| Publication number | Publication date |
|---|---|
| EP4498678A1 (en) | 2025-01-29 |
| EP4498678A4 (en) | 2025-07-02 |
| US20250037317A1 (en) | 2025-01-30 |
| CN116962706A (zh) | 2023-10-27 |
| CN116962706B (zh) | 2025-12-12 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| TWI893191B (zh) | 圖像編碼方法、圖像解碼方法及相關裝置 | |
| JP2012508485A (ja) | Gpu加速を伴うソフトウエアビデオトランスコーダ | |
| TWI882138B (zh) | 圖像編碼方法、圖像解碼方法及相關裝置 | |
| KR20060043115A (ko) | 베이스 레이어를 이용하는 영상신호의 엔코딩/디코딩 방법및 장치 | |
| CN113747242B (zh) | 图像处理方法、装置、电子设备及存储介质 | |
| US20060072837A1 (en) | Mobile imaging application, device architecture, and service platform architecture | |
| CN110545433B (zh) | 视频编解码方法和装置及存储介质 | |
| CN110572673B (zh) | 视频编解码方法和装置、存储介质及电子装置 | |
| CN114125448B (zh) | 视频编码方法、解码方法及相关装置 | |
| CN110662071B (zh) | 视频解码方法和装置、存储介质及电子装置 | |
| US9681129B2 (en) | Scalable video encoding using a hierarchical epitome | |
| CN110636295B (zh) | 视频编解码方法和装置、存储介质及电子装置 | |
| CN110572677B (zh) | 视频编解码方法和装置、存储介质及电子装置 | |
| CN110572672B (zh) | 视频编解码方法和装置、存储介质及电子装置 | |
| CN116962706B (zh) | 图像解码方法、编码方法及装置 | |
| CN114051140B (zh) | 视频编码方法、装置、计算机设备及存储介质 | |
| WO2024108931A1 (zh) | 视频编解码方法及装置 | |
| CN116405665A (zh) | 编码方法、装置、设备及存储介质 | |
| CN115866297A (zh) | 视频处理方法、装置、设备及存储介质 | |
| US8982948B2 (en) | Video system with quantization matrix coding mechanism and method of operation thereof | |
| CN115334305A (zh) | 视频数据传输方法、装置、电子设备和介质 | |
| US20250234043A1 (en) | Training with corruption for quality propagation in low-delay video coders | |
| US20250233904A1 (en) | Discrete cosine hyperprior in neural image coding | |
| US20240244229A1 (en) | Systems and methods for predictive coding | |
| WO2025156771A1 (zh) | 编解码方法以及相关装置 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 23787366 Country of ref document: EP Kind code of ref document: A1 |
|
| WWE | Wipo information: entry into national phase |
Ref document number: 202417079056 Country of ref document: IN |
|
| WWE | Wipo information: entry into national phase |
Ref document number: 2023787366 Country of ref document: EP |
|
| ENP | Entry into the national phase |
Ref document number: 2023787366 Country of ref document: EP Effective date: 20241021 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |