WO2020063599A1 - 图像预测方法、装置以及相应的编码器和解码器 - Google Patents
图像预测方法、装置以及相应的编码器和解码器 Download PDFInfo
- Publication number
- WO2020063599A1 WO2020063599A1 PCT/CN2019/107614 CN2019107614W WO2020063599A1 WO 2020063599 A1 WO2020063599 A1 WO 2020063599A1 CN 2019107614 W CN2019107614 W CN 2019107614W WO 2020063599 A1 WO2020063599 A1 WO 2020063599A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- block
- motion vector
- prediction
- current
- current image
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/50—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding
- H04N19/503—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding involving temporal prediction
- H04N19/51—Motion estimation or motion compensation
- H04N19/513—Processing of motion vectors
- H04N19/521—Processing of motion vectors for estimating the reliability of the determined motion vectors or motion vector field, e.g. for smoothing the motion vector field or for correcting motion vectors
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/10—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
- H04N19/102—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the element, parameter or selection affected or controlled by the adaptive coding
- H04N19/119—Adaptive subdivision aspects, e.g. subdivision of a picture into rectangular or non-rectangular coding blocks
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/10—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
- H04N19/134—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the element, parameter or criterion affecting or controlling the adaptive coding
- H04N19/157—Assigned coding mode, i.e. the coding mode being predefined or preselected to be further used for selection of another element or parameter
- H04N19/159—Prediction type, e.g. intra-frame, inter-frame or bidirectional frame prediction
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/10—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
- H04N19/169—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding
- H04N19/17—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding the unit being an image region, e.g. an object
- H04N19/176—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding the unit being an image region, e.g. an object the region being a block, e.g. a macroblock
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/50—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding
- H04N19/503—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding involving temporal prediction
- H04N19/51—Motion estimation or motion compensation
- H04N19/513—Processing of motion vectors
- H04N19/517—Processing of motion vectors by encoding
- H04N19/52—Processing of motion vectors by encoding by predictive encoding
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/50—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding
- H04N19/503—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding involving temporal prediction
- H04N19/51—Motion estimation or motion compensation
- H04N19/537—Motion estimation other than block-based
- H04N19/54—Motion estimation other than block-based using feature points or meshes
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N23/00—Cameras or camera modules comprising electronic image sensors; Control thereof
- H04N23/10—Cameras or camera modules comprising electronic image sensors; Control thereof for generating image signals from different wavelengths
- H04N23/17—Cameras or camera modules comprising electronic image sensors; Control thereof for generating image signals from different wavelengths using opto-mechanical scanning means only
Definitions
- the present application relates to the technical field of video encoding and decoding, and in particular, to an image prediction method and device, and a corresponding encoder and decoder.
- Video encoding (video encoding and decoding) is widely used in digital video applications, such as broadcast digital TV, video transmission on the Internet and mobile networks, real-time conversation applications such as video chat and video conferencing, DVD and Blu-ray discs, video content acquisition and editing systems And security applications for camcorders.
- Video Coding AVC
- ITU-T H.265 High Efficiency Video Coding
- 3D three-dimensional
- HEVC High Efficiency Video Coding
- the embodiments of the present application provide an image prediction method and device, and a corresponding encoder and encoder, which can reduce complexity to a certain extent, thereby improving codec performance.
- an embodiment of the present application provides an image prediction method. It should be understood that the method of the embodiment of the present application may be a video decoder or an electronic device with a video decoding function, or the method of the embodiment of the present application.
- the execution subject may be a video encoder or an electronic device with a video encoding function, and the method may include:
- Obtaining a motion vector of a control point of a current image block for example, a current affine image block
- An affine transformation model is used to obtain a motion vector of each sub-block in the current image block according to a motion vector (such as a motion vector group) of a control point (such as an affine control point) of the current image block, where the size of the sub-block is Determined based on the prediction direction of the current image block;
- the pixel prediction value of each sub-block may be further modified or may not be modified.
- the pixel prediction values of multiple sub-blocks of the current image block are obtained.
- a size of a sub-block in the current image block is UxV; or,
- the size of the subblock in the current image block is MxN
- U M represents the width of the sub-block
- V represents the height of the sub-block
- U, V, M, N are all 2n
- n is a positive integer
- M is 4 and N is 4.
- U is 8 and V is 8.
- the method further includes: parsing an affine-related syntax element from a code stream (E.g. affine_inter_flag, affine_merge_flag).
- the obtaining a motion vector of a control point of a current image block includes:
- the motion vector of the control point of the current image block is determined according to the target candidate MVP group and the motion vector difference MVD parsed from the code stream. It should be understood that the present application is not limited to this.
- the decoder may obtain the target candidate MVP group through other methods instead of the index.
- the prediction direction of the current image block is obtained in the following manner:
- the prediction direction indication information is used to indicate unidirectional prediction or bidirectional prediction, wherein the prediction direction indication information is obtained by parsing or deriving from a code stream.
- the acquiring a motion vector of a control point of a current image block includes:
- target candidate motion information Determining target candidate motion information from a candidate motion information list according to the index, wherein the target candidate motion information includes at least one target candidate motion vector group (e.g., a target candidate motion vector corresponding to a first reference frame list (e.g., L0) Group and / or target candidate motion vectors corresponding to the second reference frame list (for example, L1), the target candidate motion vector group being a motion vector of a control point of the current image block.
- target candidate motion vector group e.g., a target candidate motion vector corresponding to a first reference frame list (e.g., L0) Group and / or target candidate motion vectors corresponding to the second reference frame list (for example, L1)
- the target candidate motion vector group being a motion vector of a control point of the current image block.
- the prediction direction of the current image block is obtained in the following manner:
- the prediction direction of the current image block is bidirectional prediction, wherein the target candidate motion information corresponding to an index (such as a candidate index) in the candidate motion information list includes a first target candidate motion vector group corresponding to a first reference frame list ( (For example, a target candidate motion vector group in the first direction), and a second target candidate motion vector group (for example, a second direction target candidate motion vector group) corresponding to the second reference frame list;
- the target candidate motion information corresponding to an index (such as a candidate index) in the candidate motion information list includes a first target candidate motion vector group corresponding to a first reference frame list ( (For example, a target candidate motion vector group in the first direction), and a second target candidate motion vector group (for example, a second direction target candidate motion vector group) corresponding to the second reference frame list;
- the prediction direction of the current image block is one-way prediction, wherein the target candidate motion information corresponding to an index (such as a candidate index) in the candidate motion vector prediction value MVP list includes a first target candidate corresponding to a first reference frame list Motion vector groups (such as target candidate motion vector groups in the first direction),
- the target candidate motion information corresponding to an index (for example, a candidate index) in the candidate motion vector prediction value MVP list includes: a second target candidate motion vector group (for example, a second direction target candidate) corresponding to a second reference frame list Motion vector set).
- the method is characterized in that, according to the obtained motion vector value of the control point of the current image block, an affine transformation model is used to obtain a motion vector value of each sub-block in the current image block, include:
- a motion vector of each sub-block in the current image block is obtained.
- the current decoding block in the prior art which is divided into MxN (that is, 4x4) sub-blocks, that is, each MxN (that is, 4x4) sub-blocks use corresponding motion vectors for motion compensation.
- the current The size of the sub-blocks in the decoded block is determined based on the prediction direction of the current decoded block; for example, if the current decoded block is unidirectional prediction, the size of the sub-block of the current decoded block is 4x4; if the current decoded block is bi-directionally predicted , Then the size of the sub-block of the current decoding block is 8x8.
- the size of the subblocks (or motion compensation units) of certain image blocks in the embodiments of the present application is relatively larger than the subblocks (or motion compensation units) in the prior art.
- the average number of read reference pixels required for each pixel for motion compensation is relatively small, and the computational complexity of interpolation is relatively low. Therefore, the embodiment of the present application reduces the motion compensation to some extent while taking into account the prediction efficiency. Complexity, which improves codec performance.
- an embodiment of the present application provides an image prediction method. It should be understood that the method of the embodiment of the present application may be a video decoder or an electronic device with a video decoding function, or the method of the embodiment of the present application.
- the execution subject may be a video encoder or an electronic device with a video encoding function, and the method may include:
- An affine transformation model is used to obtain the motion vector of each sub-block in the current image block according to the motion vector (motion vector group) of the control point (affine control point) of the current image block, where if the current image block (current affine) Image block) is unidirectional prediction, and the size of the subblock of the current image block is set to 4x4; or, if the current image block is bidirectional prediction, the size of the subblock of the current image block is set to 8x8.
- the prediction direction of the current image block is considered to determine the size of the subblock of the current image block.
- the size of the subblock is 4 * 4;
- the prediction mode of the current coded image block is bidirectional prediction, and the size of the sub-block is 8 * 8; in this way, a balance between the complexity of motion compensation and the prediction efficiency is achieved, that is, the complexity of motion compensation in the prior art is reduced.
- an embodiment of the present application provides an image prediction device, including several functional units for implementing any one of the methods in the first aspect.
- the image prediction apparatus may include:
- An obtaining unit configured to obtain a motion vector of a control point of a current image block
- An inter prediction processing unit configured to obtain an motion vector of each sub-block in the current image block by using an affine transformation model according to a motion vector of a control point of the current image block, wherein the size of the sub-block is based on the current image block The prediction direction is determined; motion compensation is performed according to the motion vector of each sub-block in the current image block to obtain the pixel prediction value of each sub-block.
- an embodiment of the present application provides an image prediction apparatus including several functional units for implementing any one of the methods in the first aspect or the second aspect.
- the image prediction apparatus may include:
- An obtaining unit configured to obtain a motion vector of a control point of a current image block
- An inter prediction processing unit configured to obtain an motion vector of each sub-block in the current image block by using an affine transformation model according to a motion vector of a control point of the current image block, where if the current image block is unidirectional prediction, the current image The size of the sub-block of the block is 4x4; or, if the current image block is bidirectional prediction, the size of the sub-block of the current image block is 8x8; motion compensation is performed according to the motion vector of each sub-block in the current image block to obtain each sub-block The pixel predicted value of the block.
- the present application provides a decoding method including: a video decoder determining an inter-frame direction of a current decoding block (which may be specifically a current affine decoding block); parsing a code stream to obtain an index and a motion vector difference MVD Determining a target motion vector group (also referred to as a target candidate motion vector group) from a candidate motion vector prediction value MVP list (for example, an affine transformation candidate motion vector list) according to the index, where the target motion vector group represents The motion vector prediction value of a set of control points of the current decoding block; the motion vector of the control point of the current decoding block is determined according to the target candidate motion vector group and the motion vector difference value MVD parsed from the code stream; according to the determined The motion vector value of the control point of the current decoding block uses an affine transformation model to obtain the motion vector value of each sub-block in the current decoding block, where the size of the sub-block is determined based on the prediction direction of the current decoding block,
- the current The size of a sub-block in a decoding block is determined based on the prediction direction of the current decoding block, or the size of the sub-block is determined based on the prediction direction of the current decoding block and the size of the current decoding block; for example, if If the current decoding block is unidirectional prediction in the AMVP mode, the size of the sub-block of the current decoding block is set to 4x4; if the current decoding block is the bidirectional prediction in the AMVP mode, the size of the sub-block of the current decoding block is set to 8x4 or 4x8.
- the size of the sub-block (or motion compensation unit) in the embodiment of the present application is relatively larger than the sub-block (or motion compensation unit) in the prior art.
- the average of each pixel is The number of reference pixels to be read during motion compensation is relatively small, and the computational complexity of interpolation is relatively low. Therefore, the embodiment of the present application reduces the complexity of motion compensation to a certain extent while taking into account the prediction efficiency. Improved codec performance.
- the size of the subblock in the current decoding block is UxV;
- the size of the sub-blocks in the current decoding block is MxN.
- the size of the sub-blocks in the current decoding block is MxN.
- the sub-block for unidirectional prediction of the current affine decoding block is divided into MxN
- M is an integer such as 4, 8, or 16 and N is an integer such as 4, 8, or 16.
- the parsing code stream further includes: parsing the affine-related code stream from the code stream. Syntax elements (eg, affine_inter_flag, affine_merge_flag).
- the size of the sub-block of the current affine decoding block is set to 4x4; if the current affine decoding block is bidirectional prediction, the The sub-block size of the current affine decoding block is set to 8x4.
- the inter prediction mode of the current decoding block is modified to a unidirectional prediction mode.
- the modifying the inter prediction mode of the current decoding block to a unidirectional prediction mode includes: discarding backward predicted motion information and converting it to forward prediction; or , Discard the forward predicted motion information and convert it to backward prediction.
- the flag bits of the bidirectional prediction need not be parsed.
- the size of the sub-block of the current affine decoding block is set to 4x4; if the current affine decoding block is bidirectional prediction, Then the size of the sub-block of the current affine decoding block is set to 4x8.
- the size of the sub-blocks of the current affine decoding block is set to 4x4; if the current affine decoding block is bidirectional prediction , The adaptive division is performed according to the size of the affine decoding block; the division method may be any one of the following three methods:
- the width A of the affine decoded block is greater than H, set the size of the sub-block of the current affine decoded block to 8x4, and if the width W of the affine decoded block is less than or equal to H, then the current affine decoded block The size of the child block is set to 4x8; or
- the code stream may also be limited, so that when the width W of the affine decoded block is equal to 8 and the height H is equal to 8, the flag bits for bidirectional prediction need not be parsed.
- each candidate motion vector group may be a motion vector binary group or a motion vector triplet.
- the target motion vector group can be directly determined without analyzing the bitstream to obtain an index.
- the determining an inter prediction mode of a current decoding block includes parsing a code stream to obtain an inter frame.
- One or more syntax elements related to prediction where the one or more syntax elements are used to indicate that the current decoded block adopts the AMVP mode, and are used to indicate that the current decoded block adopts unidirectional prediction or bidirectional prediction.
- the obtaining the motion vector value of each sub-block in the current decoding block by using an affine transformation model according to the determined motion vector value of the control point of the current decoding block includes:
- the affine transformation model obtains the motion vector of one or more sub-blocks of the current decoding block (for example, substituting the coordinates of the center point of the one or more sub-blocks into the affine transformation model, thereby obtaining the motion of one or more sub-blocks Vector), wherein the affine transformation model is determined based on position coordinates of a set of control points of the current decoding block and a motion vector of a set of control points of the current decoding block, in other words, the affine transformation
- the model parameters of the model are determined based on position coordinates of a set of control points of the current decoding block and motion vectors of a set of control points of the current decoding block.
- the present application provides another decoding method, including: a video decoder determining a prediction direction of a current decoded block; parsing a bitstream to obtain an index; and the video decoder determining from a candidate motion information list according to the index Target motion vector group; the video decoder uses the parameter affine transformation model to obtain the motion vector value of each sub-block in the current decoding block according to the determined motion vector value of the control point of the current decoding block, wherein the size of the sub-block is based on The inter-frame direction of the current decoding block is determined, or the size of the sub-block is determined based on the inter-frame direction of the current decoding block and the size of the current decoding block; according to the motion of each sub-block in the current decoding block The vector value is subjected to motion compensation to obtain the pixel prediction value of each sub-block.
- the determining a prediction direction of a current decoding block includes: determining, based on the index (for example, affine_merge_idx), that the current decoding block adopts Unidirectional prediction or bidirectional prediction, where the prediction direction of the current decoded block is the same as the prediction direction of the candidate motion information indicated by the index (for example, affine_merge_idx);
- the current The size of a sub-block in a decoding block is determined based on the prediction direction of the current decoding block, or the size of the sub-block is determined based on the prediction direction of the current decoding block and the size of the current decoding block; for example, if If the current decoding block is a unidirectional prediction of the merge mode, the size of the subblock of the current decoding block is set to 4x4; if the current decoding block is a bidirectional prediction of the merge mode, the size of the subblock of the current decoding block is set to 8x4 or 4x8.
- the size of the sub-blocks (or motion compensation units) of some image blocks in the embodiments of the present application is relatively larger than the size of the sub-blocks (or motion compensation units) in the prior art.
- the average number of reference pixels to be read for each pixel for motion compensation is relatively small, and the computational complexity of interpolation is relatively low. Therefore, in the embodiment of the present application, while taking into account the prediction efficiency, it is reduced to a certain extent.
- the complexity of motion compensation improves codec performance.
- the inter prediction mode of the current decoding block is a bidirectional prediction mode of the MERGE mode
- the size of the subblock in the current decoding block is UxV
- the inter prediction mode of the current decoding block is a unidirectional prediction mode of the MERGE mode
- the size of the sub-blocks in the current decoding block is MxN.
- the inter prediction mode of the current decoding block is a unidirectional prediction mode of the MERGE mode
- the size of the sub-blocks in the current decoding block is MxN.
- the sub-block for unidirectional prediction of the affine decoding block is divided into MxN
- M is an integer such as 4, 8, or 16 and N is an integer such as 4, 8, or 16.
- the parsing code stream includes: parsing the code stream to obtain an affine-related syntax Element (eg affine_inter_flag, affine_merge_flag).
- an affine-related syntax Element eg affine_inter_flag, affine_merge_flag.
- the size of the basic motion compensation unit of the current affine decoded block is 4x4, and if the current affine decoded block is unidirectional prediction , Then set the size of the motion compensation unit of the current affine decoding block to 8x4.
- the inter prediction mode of the current decoding block is modified to a unidirectional prediction mode.
- the modifying the inter prediction mode of the current decoding block to a unidirectional prediction mode includes: discarding backward predicted motion information and converting it to forward prediction; or , Discard the forward predicted motion information and convert it to backward prediction.
- the flag bits of the bidirectional prediction need not be parsed.
- the size of the motion compensation unit of the current affine decoding block is set to 4x4, and if the current affine decoding block is bidirectional prediction , Set the size of its motion compensation unit to 4x8.
- the affine decoding block is unidirectional prediction
- the size of its motion compensation unit is set to 4x4
- the affine decoding block is bidirectional prediction
- the affine decoding is performed according to the affine decoding.
- the size of the block is adaptively divided; the division can be any of the following three ways:
- the inter prediction mode of the current decoded block is modified to a unidirectional prediction mode.
- the method may further limit the code stream so that when the width W of the affine decoding block is equal to 8 and the height H is equal to 8, the flag bits for bidirectional prediction need not be parsed.
- an embodiment of the present application provides a video decoder, including several functional units for implementing any one of the foregoing methods.
- a video decoder may include:
- Entropy decoding unit used to parse the code stream to obtain the index and motion vector difference MVD;
- An inter prediction unit configured to determine a prediction direction of a current decoding block; determining a target candidate MVP group from a candidate motion vector prediction value MVP list according to the index; according to the target candidate MVP group and a motion vector parsed from a code stream
- the difference MVD determines the motion vector of the control point of the current decoding block; and according to the determined motion vector value of the control point of the current decoding block, an affine transformation model is used to obtain the motion vector value of each sub-block in the current decoding block, where The size of the sub-block is determined based on the prediction direction of the current decoding block, or the size of the sub-block is determined based on the prediction direction of the current decoding block and the size of the current decoding block; The motion vector value of each sub-block is subjected to motion compensation to obtain the pixel prediction value of each sub-block.
- an embodiment of the present application provides another video decoder, including several functional units for implementing any one of the foregoing methods.
- a video decoder may include:
- Entropy decoding unit used to parse the code stream to obtain the index
- An inter prediction unit configured to determine a prediction direction of a current decoding block; determining a target motion vector group from a candidate motion information list according to the index; and using a parameter affine according to the determined motion vector value of a control point of the current decoding block
- the transformation model obtains the motion vector value of each sub-block in the current decoding block, where the size of the sub-block is determined based on the prediction direction of the current decoding block, or the size of the sub-block is based on the prediction direction of the current decoding block and The size of the current decoding block is determined; motion compensation is performed according to a motion vector value of each sub-block in the current decoding block to obtain a pixel prediction value of each sub-block.
- the method according to the first aspect of the present application may be performed by a device according to the third aspect of the present application.
- Other features and implementations of the method according to the first aspect of the application directly depend on the functionality of the device according to the third aspect of the application and its different implementations.
- the method according to the second aspect of the present application may be performed by an apparatus according to the fourth aspect of the present application.
- Other features and implementations of the method according to the second aspect of the application directly depend on the functionality of the device according to the fourth aspect of the application and its different implementations.
- the present application relates to a device for decoding a video stream, including a processor and a memory.
- the memory stores instructions that cause the processor to execute a method according to the first aspect or the second aspect.
- the present application relates to a device for encoding a video stream, including a processor and a memory.
- the memory stores instructions that cause the processor to execute a method according to the first aspect or the second aspect.
- a computer-readable storage medium on which instructions are stored, which, when executed, cause one or more processors to encode video data.
- the instructions cause the one or more processors to perform a method according to the first or second aspect or any possible embodiment of the first or second aspect.
- the present application relates to a computer program including program code that, when run on a computer, performs a method according to the first or second aspect or any possible embodiment of the first or second aspect.
- an embodiment of the present application provides a device for decoding video data, where the device includes:
- Memory for storing video data in the form of a stream
- a video decoder for parsing a bitstream to obtain an index and a motion vector difference MVD; and determining a target motion vector group from a candidate motion vector prediction value MVP list according to the index, the target motion vector group representing a current decoded block Motion vector prediction value of a set of control points; determining the motion vector of the control point of the current decoding block according to the target motion vector group and the motion vector difference MVD parsed from the code stream; and also used to determine the current decoding block's Prediction direction; using an affine transformation model to obtain a motion vector value of each sub-block (for example, one or more sub-blocks) in the current decoding block according to the determined motion vector value of the control point of the current decoding block, wherein the sub-blocks
- the size of is determined based on the prediction direction of the current decoding block, or the size of the sub-block is determined based on the prediction direction of the current decoding block and the size of the current decoding block; according to each sub-block in the current decoding
- an embodiment of the present application provides a device for decoding video data, where the device includes:
- Memory for storing video data in the form of a stream
- a video decoder for parsing a bitstream to obtain an index; and determining a target motion vector group from a candidate motion information list according to the index, where the target motion vector group represents a motion vector of a set of control points of a current decoding block; And, it is also used to determine the prediction direction of the current decoded block; according to the determined motion vector value of the control point of the current decoded block, a parameter affine transformation model is used to obtain the A motion vector value, wherein the size of the sub-block is determined based on the prediction direction of the current decoding block, or the size of the sub-block is determined based on the prediction direction of the current decoding block and the size of the current decoding block; Perform motion compensation according to a motion vector value of each sub-block (for example, one or more sub-blocks) in the current decoding block to obtain a pixel prediction value of each sub-block; in other words, based on one or more sub-blocks of the current decoding block The motion vector of the pixel to predict the
- FIG. 1 is a block diagram of a video encoding and decoding system in an implementation manner described in an embodiment of the present application;
- FIG. 2A is a block diagram of a video encoder in an implementation described in an embodiment of the present application.
- 2B is a schematic diagram of inter prediction in an implementation manner described in an embodiment of the present application.
- FIG. 2C is a block diagram of a video decoder in an implementation described in an embodiment of the present application.
- FIG. 3 is a schematic diagram of candidate positions of motion information in the implementation described in the embodiment of the present application.
- FIG. 5A is a schematic diagram of prediction of a control point motion vector constructed in the implementation manner described in the embodiment of the present application.
- FIG. 5B is a schematic flowchart of combining control point motion information to obtain a structured control point motion information in the implementation described in the embodiment of the present application;
- 6A is a flowchart of a decoding method in an implementation manner described in an embodiment of the present application.
- FIG. 6B is a schematic diagram of constructing a candidate motion vector list in the implementation described in the embodiment of the present application.
- 6C is a schematic diagram of a sub-block (also referred to as a motion compensation unit) in an implementation manner described in an embodiment of the present application;
- FIG. 7A is a schematic flowchart of an image prediction method provided in an embodiment of the present application.
- FIG. 7B is a schematic structural diagram of a sub-block (also referred to as a motion compensation unit) according to an embodiment of the present application;
- FIG. 7C is a schematic structural diagram of still another seed block (also referred to as a motion compensation unit) according to an embodiment of the present application.
- 7D is a schematic flowchart of a decoding method provided in an embodiment of the present application.
- FIG. 8A is a schematic structural diagram of an image prediction apparatus according to an embodiment of the present invention.
- 8B is a schematic structural diagram of an encoding device or a decoding device according to an embodiment of the present invention.
- FIG. 9 is a video encoding system 1100 including the encoder 20 of FIG. 2A and / or the decoder 30 of FIG. 2C according to an exemplary embodiment.
- FIG. 1 is a schematic block diagram of a video encoding and decoding system 10 according to an embodiment of the present application.
- the system 10 includes a source device 11 and a destination device 12.
- the source device 11 generates encoded video data and sends the encoded video data to the destination device 12.
- the destination device 12 is configured to receive the encoded video data, decode and display the encoded video data. .
- the source device 11 and the destination device 12 may include any of a wide range of devices, including desktop computers, notebook computers, tablet computers, set-top boxes, telephone handsets such as so-called “smart” phones, and so-called “smart” Touchpads, TVs, cameras, display devices, digital media players, video game consoles, video streaming devices, and more.
- FIG. 1 is a schematic block diagram of a video encoding and decoding system 10 according to an embodiment of the present application.
- the system 10 includes a source device 11 and a destination device 12.
- the source device 11 generates encoded video data and sends the encoded video data to the destination device 12.
- the destination device 12 is configured to receive the encoded video data, decode and display the encoded video data. .
- the source device 11 and the destination device 12 may include any of a wide range of devices, including desktop computers, notebook computers, tablet computers, set-top boxes, telephone handsets such as so-called “smart” phones, and so-called “smart” Touchpads, TVs, cameras, display devices, digital media players, video game consoles, video streaming devices, and more.
- the destination device 12 may receive the encoded video data to be decoded via the link 16.
- the link 16 may include any type of media or device capable of passing the encoded video data from the source device 11 to the destination device 12.
- the link 16 may include a communication medium that enables the source device 11 to directly transmit the encoded video data to the destination device 12 in real time.
- the encoded video data may be modulated according to a communication standard (eg, a wireless communication protocol) and transmitted to the destination device 12.
- Communication media may include any wireless or wired communication media, such as a radio frequency spectrum or one or more physical transmission lines. Communication media may form part of a packet-based network, such as a global network of a local area network, a wide area network, or the Internet.
- the communication medium may include a router, a switch, a base station, or any other device that may be used to facilitate communication from the source device 11 to the destination device 12.
- the video encoding and decoding system 10 further includes a storage device that can output the encoded data from the output interface 14 to the storage device.
- the encoded data can be accessed from the storage device by the input interface 15.
- the storage device may include any of a variety of distributed or locally accessible data storage media, such as hard drives, Blu-ray discs, DVDs, CD-ROMs, flash memory, volatile or non-volatile memory, or Any other suitable digital storage medium for storing encoded video data.
- the storage device may correspond to a file server or another intermediate storage device that may hold the encoded video produced by the source device 11.
- the destination device 12 can access the stored video data from the storage device via streaming or downloading.
- the file server may be any type of server capable of storing encoded video data and transmitting this encoded video data to the destination device 12. Possible implementations The file server includes a web server, a file transfer protocol server, a network attached storage device, or a local disk drive.
- the destination device 12 can access the encoded video data via any standard data connection including an Internet connection. This data connection may include a wireless channel (e.g., a Wi-Fi connection), a wired connection (e.g., a cable modem, etc.), or a combination of both, suitable for accessing the encoded video data stored on the file server.
- the transmission of the encoded video data from the storage device may be a streaming transmission, a download transmission, or a combination of the two.
- the techniques of this application are not necessarily limited to wireless applications or settings.
- Technology can be applied to video decoding to support any of a variety of multimedia applications, such as over-the-air television broadcasting, cable television transmission, satellite television transmission, streaming video transmission (e.g., via the Internet), encoding digital video for storage Decoding digital video or other applications on a data storage medium.
- the system 10 may be configured to support one-way or two-way video transmission to support applications such as video streaming, video playback, video broadcasting, and / or video telephony.
- the source device 11 may include a video source 13, a video encoder 20, and an output interface 14.
- the output interface 14 may include a modulator / demodulator (modem) and / or a transmitter.
- the video source 13 may include, for example, various source devices such as a video capture device (e.g., a video camera), an archive containing previously captured video, and a video feed interface for receiving video from a video content provider , And / or a computer graphics system for generating computer graphics data as a source video, or a combination of these sources.
- the source device 11 and the destination device 12 may form a so-called camera phone or video phone.
- the techniques described in this application may be exemplarily applicable to video decoding, and may be applicable to wireless and / or wired applications.
- the captured, pre-captured, or computationally generated video may be encoded by video encoder 20.
- the encoded video data can be directly transmitted to the destination device 12 via the output interface 14 of the source device 11.
- the encoded video data may also (or alternatively) be stored on a storage device for later access by the destination device 12 or other device for decoding and / or playback.
- the destination device 12 includes an input interface 15, a video decoder 30, and a display device 17.
- the input interface 15 may include a receiver and / or a modem.
- the input interface 15 of the destination device 12 receives the encoded video data via the link 16.
- the encoded video data communicated or provided on the storage device via the link 16 may include various syntax elements generated by the video encoder 20 for use by the video decoder 30 of the video decoder 30 to decode the video data. These syntax elements may be included with the encoded video data transmitted on a communication medium, stored on a storage medium, or stored on a file server.
- the display device 17 may be integrated with or external to the destination device 12.
- the destination device 12 may include an integrated display device and is also configured to interface with an external display device.
- the destination device 12 may be a display device.
- the display device 17 displays the decoded video data to a user, and may include any of a variety of display devices, such as a liquid crystal display, a plasma display, an organic light emitting diode display, or another type of display device.
- Video encoder 20 and video decoder 30 may operate according to, for example, the next-generation video codec compression standard (H.266) currently under development and may conform to the H.266 test model (JEM).
- the video encoder 20 and the video decoder 30 may be based on, for example, the ITU-TH.265 standard, also referred to as a high-efficiency video decoding standard, or other proprietary or industrial standards of the ITU-TH.264 standard or extensions of these standards
- the ITU-TH.264 standard is alternatively referred to as MPEG-4 Part 10, also known as advanced video coding (AVC).
- AVC advanced video coding
- the techniques of this application are not limited to any particular decoding standard.
- Other possible implementations of the video compression standard include MPEG-2 and ITU-TH.263.
- video encoder 20 and video decoder 30 may each be integrated with an audio encoder and decoder, and may include an appropriate multiplexer-demultiplexer (MUX-DEMUX) unit or other hardware and software to handle encoding of both audio and video in a common or separate data stream.
- MUX-DEMUX multiplexer-demultiplexer
- the MUX-DEMUX unit may conform to the ITUH.223 multiplexer protocol or other protocols such as the User Datagram Protocol (UDP).
- UDP User Datagram Protocol
- Video encoder 20 and video decoder 30 may each be implemented as any of a variety of suitable encoder circuits, such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (application specific integrated circuit (ASIC)), field-programmable gate array (FPGA), discrete logic, software, hardware, firmware, or any combination thereof.
- DSPs digital signal processors
- ASIC application specific integrated circuit
- FPGA field-programmable gate array
- the device may store the software's instructions in a suitable non-transitory computer-readable medium and execute the instructions in hardware using one or more processors to perform the techniques of this application.
- Each of the video encoder 20 and the video decoder 30 may be included in one or more encoders or decoders, and any of them may be integrated as a combined encoder / decoder (CODEC) in a corresponding device. part.
- CDEC combined encoder / decoder
- H.265 HEVC
- HM HEVC test model
- the latest standard document of H.265 can be obtained from http://www.itu.int/rec/T-REC-H.265.
- the latest version of the standard document is H.265 (12/16).
- the standard document is in full text. The citation is incorporated herein. HM assumes that video decoding devices have several additional capabilities over existing algorithms of ITU-TH.264 / AVC.
- H.266 test model The evolution model of the video decoding device.
- the algorithm description of H.266 can be obtained from http://phenix.int-evry.fr/jvet. The latest algorithm description is included in JVET-F1001-v2.
- the algorithm description document is incorporated herein by reference in its entirety.
- the reference software for the JEM test model can be obtained from https://jvet.hhi.fraunhofer.de/svn/svn_HMJEMSoftware/, which is also incorporated herein by reference in its entirety.
- HM can divide a video frame or image into a sequence of tree blocks or maximum coding units (LCUs) containing both luminance and chrominance samples.
- LCUs are also known as CTUs.
- the tree block has a similar purpose as the macro block of the H.264 standard.
- a slice contains several consecutive tree blocks in decoding order.
- a video frame or image can be split into one or more slices.
- Each tree block can be split into coding units according to a quadtree. For example, a tree block that is a root node of a quad tree may be split into four child nodes, and each child node may be a parent node and split into another four child nodes.
- the final indivisible child nodes that are leaf nodes of the quadtree include decoding nodes, such as decoded image blocks.
- decoding nodes such as decoded image blocks.
- the syntax data associated with the decoded codestream can define the maximum number of times a tree block can be split, and can also define the minimum size of a decoding node.
- the coding unit includes a decoding node and a prediction unit (PU) and a transformation unit (TU) associated with the decoding node.
- the size of the CU corresponds to the size of the decoding node and the shape must be square.
- the size of the CU can range from 8 ⁇ 8 pixels to a maximum 64 ⁇ 64 pixels or larger tree block size.
- Each CU may contain one or more PUs and one or more TUs.
- the syntax data associated with a CU may describe a case where a CU is partitioned into one or more PUs.
- the partitioning mode may be different between cases where the CU is skipped or is encoded in direct mode, intra prediction mode, or inter prediction mode.
- the PU can be divided into non-square shapes.
- the syntax data associated with a CU may also describe a case where a CU is partitioned into one or more TUs according to a quadtree.
- the shape of the TU can be square or non-square.
- the HEVC standard allows transformation based on the TU, which can be different for different CUs.
- the TU is usually sized based on the size of the PUs within a given CU defined for the partitioned LCU, but this may not always be the case.
- the size of the TU is usually the same as or smaller than the PU.
- a quad-tree structure called "residual quad tree" (RQT) may be used to subdivide the residual samples corresponding to the CU into smaller units.
- the leaf node of RQT may be called TU.
- the pixel difference values associated with the TU may be transformed to produce a transformation coefficient, which may be quantized.
- the PU contains data related to the prediction process.
- the PU may include data describing the intra-prediction mode of the PU.
- the PU may include data defining a motion vector of the PU.
- the data defining the motion vector of the PU may describe the horizontal component of the motion vector, the vertical component of the motion vector, the resolution of the motion vector (e.g., quarter-pixel accuracy or eighth-pixel accuracy), motion vector The reference image pointed to, and / or the reference image list of the motion vector (eg, list 0, list 1 or list C).
- TU uses transform and quantization processes.
- a given CU with one or more PUs may also contain one or more TUs.
- video encoder 20 may calculate a residual value corresponding to the PU.
- the residual values include pixel differences that can be transformed into transform coefficients, quantized, and scanned using TU to generate serialized transform coefficients for entropy decoding.
- This application generally uses the term "image block" to refer to the decoding node of a CU.
- image block may also be used in this application to refer to a tree block including a decoding node and a PU and a TU, such as an LCU or a CU.
- the video encoder 20 encodes video data.
- Video data may include one or more pictures.
- the video encoder 20 may generate a code stream, which includes encoding information of video data in the form of a bit stream.
- the encoding information may include encoded picture data and associated data.
- the associated data may include a sequence parameter set (SPS), a picture parameter set (PPS), and other syntax structures.
- SPS may contain parameters applied to zero or more sequences.
- SPS describes high-level parameters of the general characteristics of a coded video sequence (CVS).
- the sequence parameter set SPS contains information required by all slices in the CVS.
- the PPS may contain parameters applied to zero or more pictures.
- a grammatical structure is a collection of zero or more grammatical elements arranged in a specified order in a code stream.
- HM supports prediction of various PU sizes. Assuming that the size of a specific CU is 2N ⁇ 2N, HM supports intra prediction of PU sizes of 2N ⁇ 2N or N ⁇ N, and symmetric PU sizes of 2N ⁇ 2N, 2N ⁇ N, N ⁇ 2N, or N ⁇ N between frames. prediction. HM also supports asymmetric partitioning of PU-sized inter predictions of 2N ⁇ nU, 2N ⁇ nD, nL ⁇ 2N, and nR ⁇ 2N. In asymmetric partitioning, one direction of the CU is not partitioned, and the other direction is partitioned into 25% and 75%.
- 2N ⁇ nU refers to a horizontally-divided 2N ⁇ 2NCU, where 2N ⁇ 0.5NPU is at the top and 2N ⁇ 1.5NPU is at the bottom.
- N ⁇ N and “N times N” are used interchangeably to refer to the pixel size of an image block according to the vertical and horizontal dimensions, for example, 16 ⁇ 16 pixels or 16 ⁇ 16 pixels.
- N ⁇ N blocks generally have N pixels in the vertical direction and N pixels in the horizontal direction, where N represents a non-negative integer value. Pixels in a block can be arranged in rows and columns.
- the block does not necessarily need to have the same number of pixels in the horizontal direction as in the vertical direction.
- a block may include N ⁇ M pixels, where M is not necessarily equal to N.
- the video encoder 20 may calculate the residual data of the TU of the CU.
- a PU may include pixel data in a spatial domain (also referred to as a pixel domain), and a TU may include transforming (e.g., discrete cosine transform (DCT), integer transform, wavelet transform, or conceptually similar transform) Coefficients in the transform domain after being applied to the residual video data.
- the residual data may correspond to a pixel difference between a pixel of an uncoded image and a prediction value corresponding to a PU.
- Video encoder 20 may form a TU containing residual data of the CU, and then transform the TU to generate transform coefficients for the CU.
- the JEM model further improves the coding structure of video images.
- a block coding structure called "Quad Tree Combined with Binary Tree” (QTBT) is introduced.
- QTBT Quality Tree Combined with Binary Tree
- a CU can be square or rectangular.
- a CTU first performs a quadtree partition, and the leaf nodes of the quadtree further perform a binary tree partition.
- there are two partitioning modes in binary tree partitioning symmetrical horizontal partitioning and symmetrical vertical partitioning.
- the leaf nodes of a binary tree are called CUs.
- JEM's CUs cannot be further divided during the prediction and transformation process, which means that JEM's CU, PU, and TU have the same block size.
- the maximum size of the CTU is 256 ⁇ 256 luminance pixels.
- FIG. 2A is a schematic block diagram of a video encoder 20 according to an embodiment of the present application.
- the video encoder 20 may include a prediction module 21, a summer 22, a transform module 23, a quantization module 24, and an entropy encoding module 25.
- the prediction module 21 may include an inter prediction module 211 and an intra prediction module 212, and the internal structure of the prediction module 21 is not limited in this embodiment of the present application.
- the video encoder 20 may also include an inverse quantization module 26, an inverse transform module 27, and a summer 28.
- the video encoder 20 may further include a storage module 29. It should be understood that the storage module 29 may also be disposed outside the video encoder 20.
- the video encoder 20 may further include a filter (not illustrated in FIG. 2A) to filter the boundary of the image block to remove artifacts from the reconstructed video image.
- the filter filters the output of the summer 28 when needed.
- the video encoder 20 may further include a segmentation unit (not shown in FIG. 2A).
- the video encoder 20 receives video data, and the dividing unit divides the video data into image blocks.
- This segmentation may also include segmentation into slices, image blocks, or other larger units, and segmentation of image blocks, for example, based on the quad-tree structure of the LCU and CU.
- Video encoder 20 exemplarily illustrates the components of an image block encoded in a video slice to be encoded.
- a slice may be divided into a plurality of image blocks (and may be divided into a set called an image block).
- the types of slices include I (mainly used for intra-frame image coding), P (for inter-frame forward reference prediction image coding), and B (for inter-bidirectional reference prediction image coding).
- the prediction module 21 is configured to perform intra or inter prediction on an image block that needs to be processed to obtain a prediction value of the current block (which may be referred to as prediction information in this application).
- a prediction value of the current block which may be referred to as prediction information in this application.
- an image block that needs to be processed currently may be simply referred to as a to-be-processed block, or may be simply referred to as a current image block, or may be simply referred to as a current block.
- the image blocks that currently need to be processed during the encoding phase may also be referred to as current coding blocks for short, and the image blocks that are currently required to be processed during the decoding phase may be referred to as current decoding blocks or current decoding blocks for short.
- the inter prediction module 211 included in the prediction module 21 performs inter prediction on the current block to obtain an inter prediction value.
- the intra prediction module 212 performs intra prediction on the current block to obtain an intra prediction value.
- the inter prediction module 211 finds a matching reference block for the current block in the current image, and uses the pixel value of the pixel point in the reference block as the prediction information or prediction value of the pixel value of the pixel point in the current block. (In the following, information and values are no longer distinguished.) This process is called motion estimation (ME) (as shown in FIG. 2B), and the motion information of the current block is transmitted.
- ME motion estimation
- the motion information of the image block includes indication information of the prediction direction (usually forward prediction, backward prediction or bidirectional prediction), one or two motion vectors (Motion vector, MV) pointing to the reference block, and The indication information (usually referred to as the reference frame index) of the picture in which the reference block is located.
- Forward prediction refers to that the current block selects a reference image from the forward reference image set to obtain a reference block.
- Backward prediction means that the current block selects a reference image from the backward reference image set to obtain a reference block.
- Bidirectional prediction refers to selecting a reference image from the forward and backward reference image sets to obtain a reference block. When the bidirectional prediction method is used, there are two reference blocks in the current block. Each reference block needs to be indicated by a motion vector and a reference frame index, and then the pixels in the current block are determined according to the pixel values of the pixels in the two reference blocks. The predicted value of the pixel value.
- the motion estimation process needs to try multiple reference blocks in the reference image for the current block. Which one or several reference blocks are ultimately used for prediction is determined using rate-distortion optimization (RDO) or other methods.
- RDO rate-distortion optimization
- the video encoder 20 forms residual information by subtracting the prediction value from the current block.
- the transform module 23 is configured to transform the residual information.
- the transform module 23 uses, for example, discrete cosine transform (DCT) or a conceptually similar transform (for example, discrete sine transform DST) to transform the residual information into residual transform coefficients.
- the transform module 23 may send the obtained residual transform coefficients to a quantization module 24.
- the quantization module 24 quantizes the residual transform coefficients to further reduce the bit rate.
- the quantization module 24 may then perform a scan of the matrix containing the quantized transform coefficients.
- the entropy encoding module 25 may perform scanning.
- the entropy encoding module 25 may entropy encode the quantized residual transform coefficients to obtain a code stream.
- the entropy encoding module 25 may perform context adaptive variable length decoding (CAVLC), context adaptive binary arithmetic decoding (CABAC), syntax-based context adaptive binary arithmetic decoding (SBAC), probability interval partitioning entropy (PIPE) decoding or another entropy coding method or technique.
- CAVLC context adaptive variable length decoding
- CABAC context adaptive binary arithmetic decoding
- SBAC syntax-based context adaptive binary arithmetic decoding
- PIPE probability interval partitioning entropy
- the encoded code stream may be transmitted to the video decoder 30 or archived for later transmission or retrieved by the video decoder 30.
- the inverse quantization module 26 and the inverse transform module 27 respectively apply inverse quantization and inverse transform to reconstruct a residual block in the pixel domain for later use as a reference block of a reference image.
- the summer 28 adds the reconstructed residual information and the prediction value generated by the prediction module 21 to generate a reconstruction block, and uses the reconstruction block as a reference block for storage in the storage module 29.
- These reference blocks can be used by the prediction module 21 as reference blocks to predict inter- or intra-blocks in subsequent video frames or images.
- the video encoder 20 may directly quantize the residual information without processing by the transform module 23 and correspondingly does not need to be processed by the inverse transform module 27; or, for some image blocks Or image frames, the video encoder 20 does not generate residual information, and accordingly does not need to be processed by the transform module 23, the quantization module 24, the inverse quantization module 26, and the inverse transform module 27; or, the video encoder 20 may convert the reconstructed image
- the blocks are stored directly as reference blocks without being processed by the filter unit; or, the quantization module 24 and the inverse quantization module 26 in the video encoder 20 may be merged together; or the transform module 23 and the inverse transform in the video encoder 20 Modules 27 may be merged together; alternatively, summer 22 and summer 28 may be merged together.
- FIG. 2C is a schematic block diagram of a video decoder 30 according to an embodiment of the present application.
- the video decoder 30 may include an entropy decoding module 31, a prediction module 32, an inverse quantization module 34, an inverse transform module 35, and a reconstruction module 36.
- the prediction module 32 may include a motion compensation module 322 and an intra prediction module 321, which is not limited in this embodiment of the present application.
- the video decoder 30 may further include a storage module 33. It should be understood that the storage module 33 may also be disposed outside the video decoder 30. In some feasible implementations, the video decoder 30 may perform an exemplary reciprocal decoding process that is inverse to the encoding process described with respect to the video encoder 20 from FIG. 2A.
- video decoder 30 receives a code stream from video encoder 20.
- the code stream received by the video decoder 30 successively performs entropy decoding, inverse quantization, and inverse transformation on the entropy decoding module 31, inverse quantization module 34, and inverse transform module 35 to obtain residual information.
- the motion compensation module 322 needs to parse out the motion information, use the parsed motion information to determine a reference block in the reconstructed image block, and use the pixel values of the pixels in the reference block as prediction information ( This process is called motion compensation (MC).
- the reconstruction module 36 can obtain the reconstruction information by using the prediction information and the residual information.
- this application exemplarily relates to inter-frame decoding.
- certain techniques of this application may be performed by the motion compensation module 322.
- one or more other units of video decoder 30 may additionally or alternatively be responsible for performing the techniques of this application.
- AMVP advanced motion vector prediction
- AMVP mode For AMVP mode, first traverse the current block spatial or time-domain adjacent coded blocks (denoted as neighboring blocks), build a candidate motion vector list (also known as motion information candidate list) based on the motion information of each neighboring block, and then pass The rate distortion cost determines the optimal motion vector from the candidate motion vector list, and the candidate motion information with the lowest rate distortion cost is used as the motion vector predictor (MVP) of the current block.
- MVP motion vector predictor
- the rate-distortion cost is calculated by formula (1), where J represents the rate-distortion cost RD Cost, SAD is the sum of the absolute error between the predicted pixel value and the original pixel value obtained after motion estimation using the candidate motion vector prediction value of absolute differences (SAD), where R is the code rate and ⁇ is the Lagrangian multiplier.
- the encoding end passes the index value of the selected motion vector prediction value in the candidate motion vector list and the reference frame index value to the decoding end. Further, a motion search is performed in a neighborhood centered on the MVP to obtain the actual motion vector of the current block, and the encoder transmits the difference between the MVP and the actual motion vector (motion vector difference) to the decoder.
- the motion information of the current block in the spatial or time-domain adjacent coded blocks is used to construct a list of candidate motion vectors, and then the optimal motion information is determined from the list of candidate motion vectors by calculating the rate-distortion cost as the current block's The motion information, and then the index value of the position of the optimal motion information in the candidate motion vector list (referred to as merge index, the same applies hereinafter) to the decoding end.
- merge index the index value of the position of the optimal motion information in the candidate motion vector list
- the spatial and temporal candidate motion information of the current block is shown in Figure 3.
- the spatial motion candidate information comes from the spatially adjacent 5 blocks (A0, A1, B0, B1, and B2).
- the motion information of the neighboring block is not added to the candidate motion vector list.
- the time-domain candidate motion information of the current block is obtained after scaling the MV of the corresponding position block in the reference frame according to the reference frame and the picture order count (POC) of the current frame.
- POC picture order count
- the positions of the neighboring blocks in the Merge mode and their traversal order are also predefined, and the positions of the neighboring blocks and their traversal order may be different in different modes.
- AMVP modes can be divided into AMVP modes based on translational models and AMVP modes based on non-translational models;
- Merge modes can be divided into Merge modes based on translational models and non-translational movement models.
- Merge mode can be divided into Merge modes based on translational models and non-translational movement models.
- Non-translational motion model prediction refers to the use of the same motion model at the codec side to derive the motion information of each sub-motion compensation unit in the current block, and perform motion compensation based on the motion information of the sub-motion compensation unit to obtain the prediction block, thereby improving the prediction. effectiveness.
- Commonly used non-translational motion models are 4-parameter affine motion models or 6-parameter affine motion models.
- the sub-motion compensation unit involved in the embodiment of the present application may be a pixel or a pixel block of size N 1 ⁇ N 2 divided according to a specific method, where N 1 and N 2 are both positive integers and N 1 It may be equal to N 2 or not equal to N 2 .
- the 4-parameter affine motion model can be represented by the motion vector of two pixels and their coordinates relative to the top left vertex pixel of the current block.
- the pixels used to represent the parameters of the motion model are called control points. If the upper left vertex (0,0) and upper right vertex (W, 0) pixels are used as control points, first determine the motion vectors (vx0, vy0) and (vx1, vy1) of the upper left vertex and upper right vertex control points of the current block. Then, the motion information of each sub motion compensation unit in the current block is obtained according to formula (3), where (x, y) is the coordinate of the sub motion compensation unit relative to the top left vertex pixel of the current block, and W is the width of the current block.
- the 6-parameter affine motion model can be represented by the motion vector of three pixels and its coordinates relative to the top left vertex pixel of the current block. If the upper left vertex (0,0), upper right vertex (W, 0), and lower left vertex (0, H) pixels are used as control points, the motion vectors of the upper left vertex, upper right vertex, and lower left vertex control point of the current block are determined respectively Is (vx0, vy0) and (vx1, vy1) and (vx2, vy2), and then the motion information of each sub motion compensation unit in the current block is obtained according to formula (5), where (x, y) is the sub motion compensation unit Relative to the coordinates of the top left pixel of the current block, W and H are the width and height of the current block, respectively.
- the coding block predicted by the affine motion model is called an affine coding block.
- an advanced motion vector prediction (AMVP) mode based on an affine motion model or a merge mode based on an affine motion model can be used to obtain the motion information of the control points of the affine coding block.
- AMVP advanced motion vector prediction
- the motion information of the control points of the current coding block can be obtained by an inherited control point motion vector prediction method or a constructed control point motion vector prediction method.
- the inherited control point motion vector prediction method refers to using a motion model of an adjacent coded affine coding block to determine a candidate control point motion vector of the current block.
- the affine coding block where the position block is located obtains the control point motion information of the affine coding block, and then uses the motion model constructed by the control point motion information of the affine coding block to derive the control point motion vector of the current block (for Merge Mode) or motion vector prediction value of control point (for AMVP mode).
- A1-> B1-> B0-> A0-> B2 is only an example, and the order of other combinations is also applicable to this application.
- the adjacent position blocks are not limited to A1, B1, B0, A0, and B2.
- the adjacent position block can be a pixel point, and a pixel block of a preset size divided according to a specific method.
- it can be a 4x4 pixel block, it can also be a 4x2 pixel block, or it can be a pixel block of other sizes. limited.
- the motion vector (vx4, vy4) and the upper right vertex (x5, y5) of the upper left vertex (x4, y4) of the affine coding block are obtained.
- the combination of the motion vector (vx0, vy0) of the upper left vertex (x0, y0) of the current block and the motion vector (vx1, vy1) of the upper right vertex (x1, y1) obtained based on the affine coding block where A1 is located as above is the current Candidate control point motion vector for the block.
- the motion vector (vx4, vy4) of the upper left vertex (x4, y4) and the motion vector (vx5, y5) of the upper right vertex (x4, y5) of the affine coding block are obtained.
- the motion vector (vx1, vy1) of the upper-right vertex (x1, y1), and the lower left of the current block is the candidate control point motion vector of the current block.
- the constructed control point motion vector prediction method refers to combining the motion vectors of neighboring coded blocks around the control points of the current block as the motion vectors of the control points of the current affine coding block, without having to consider the neighboring neighboring coded blocks. Whether the coding block is an affine coding block.
- the motion information of the upper left vertex and the upper right vertex of the current block is determined by using the motion information of the coded blocks adjacent to the current coded block.
- the method for predicting the constructed control point motion vector is described by taking FIG. 5A as an example. It should be noted that FIG. 5A is only used as an example.
- the motion vectors of the upper left vertex adjacent coded blocks A2, B2, and B3 are used as candidate motion vectors for the motion vectors of the upper left vertex of the current block; the upper right vertex is adjacent to the encoded blocks B1 and B0 blocks.
- the motion vector is a candidate motion vector of the motion vector of the top right vertex of the current block.
- the above candidate motion vectors of the upper left vertex and the upper right vertex are combined to form a plurality of two-tuples.
- the motion vectors of two coded blocks included in the tuple can be used as candidate control point motion vectors of the current block. See the following formula ( 11A) shown:
- v A2 represents the motion vector of A2
- v B1 represents the motion vector of B1
- v B0 represents the motion vector of B0
- v B2 represents the motion vector of B2
- v B3 represents the motion vector of B3.
- the motion vectors of the upper left vertex adjacent coded blocks A2, B2, and B3 are used as candidate motion vectors for the motion vectors of the upper left vertex of the current block; the upper right vertex is adjacent to the encoded blocks B1 and B0 blocks.
- the motion vector is a candidate motion vector of the motion vector of the upper right vertex of the current block, and the motion vector of the coded blocks A0, A1 adjacent to the sitting vertex is used as the motion vector of the lower left vertex of the current block.
- the candidate motion vectors of the upper left vertex, upper right vertex, and lower left vertex are combined to form a triplet.
- the motion vectors of the three coded blocks included in the triplet can be used as candidate control point motion vectors for the current block. See the following formula (11B),
- v A2 indicates the motion vector of A2
- v B1 indicates the motion vector of B1
- v B0 indicates the motion vector of B0
- v B2 indicates the motion vector of B2
- v B3 indicates the motion vector of B3
- v A0 indicates the motion vector of A0
- v A1 represents the motion vector of A1.
- Step 501 Obtain motion information of each control point of the current block.
- A0, A1, A2, B0, B1, B2, and B3 are the spatially adjacent positions of the current block and are used to predict CP1, CP2, or CP3;
- T is the temporally adjacent positions of the current block and used to predict CP4.
- the inspection order is B2-> A2-> B3. If B2 is available, the motion information of B2 is used. Otherwise, detect A2, B3. If motion information is not available at all three locations, CP1 motion information cannot be obtained.
- the checking order is B0-> B1; if B0 is available, CP2 uses the motion information of B0. Otherwise, detect B1. If motion information is not available at both locations, CP2 motion information cannot be obtained.
- X can be obtained to indicate that the block including the position of X (X is A0, A1, A2, B0, B1, B2, B3, or T) has been encoded and adopts the inter prediction mode; otherwise, the X position is not available.
- Step 502 Combine the motion information of the control points to obtain the structured control point motion information.
- the motion information of the two control points is combined to form a two-tuple, which is used to construct a 4-parameter affine motion model.
- the combination of the two control points can be ⁇ CP1, CP4 ⁇ , ⁇ CP2, CP3 ⁇ , ⁇ CP1, CP2 ⁇ , ⁇ CP2, CP4 ⁇ , ⁇ CP1, CP3 ⁇ , ⁇ CP3, CP4 ⁇ .
- Affine CP1, CP2
- the combination of the three control points can be ⁇ CP1, CP2, CP4 ⁇ , ⁇ CP1, CP2, CP3 ⁇ , ⁇ CP2, CP3, CP4 ⁇ , ⁇ CP1, CP3, CP4 ⁇ .
- a 6-parameter affine motion model constructed using a triple of CP1, CP2, and CP3 control points can be written as Affine (CP1, CP2, CP3).
- a quadruple formed by combining the motion information of the four control points is used to construct an 8-parameter bilinear model.
- An 8-parameter bilinear model constructed using a quaternion of CP1, CP2, CP3, and CP4 control points is denoted as Bilinear (CP1, CP2, CP3, CP4).
- the motion information combination of two control points is simply referred to as a tuple, and the motion information of three control points (or three coded blocks) is combined. It is abbreviated as a triple, and the combination of motion information of four control points (or four coded blocks) is abbreviated as a quad.
- CurPoc represents the POC number of the current frame
- DesPoc represents the POC number of the reference frame of the current block
- SrcPoc represents the POC number of the reference frame of the control point
- MV s represents the scaled motion vector
- MV represents the motion vector of the control point.
- control points can also be converted into a control point at the same position.
- the 4-parameter affine motion model obtained by combining ⁇ CP1, CP4 ⁇ , ⁇ CP2, CP3 ⁇ , ⁇ CP2, CP4 ⁇ , ⁇ CP1, CP3 ⁇ , ⁇ CP3, CP4 ⁇ is converted to ⁇ CP1, CP2 ⁇ or ⁇ CP1, CP2, CP3 ⁇ .
- the conversion method is to substitute the motion vector of the control point and its coordinate information into formula (2) to obtain the model parameters, and then substitute the coordinate information of ⁇ CP1, CP2 ⁇ into formula (3) to obtain its motion vector.
- the conversion can be performed according to the following formulas (13)-(21), where W represents the width of the current block, H represents the height of the current block, and in formulas (13)-(21), (vx 0 , vy 0) denotes a motion vector CP1, (vx 1, vy 1) CP2 represents a motion vector, (vx 2, vy 2) represents the motion vector of CP3, (vx 3, vy 3) denotes the motion vector of CP4.
- ⁇ CP1, CP3 ⁇ conversion ⁇ CP1, CP2 ⁇ or ⁇ CP1, CP2, CP3 ⁇ can be realized by the following formula (14):
- the 6-parameter affine motion model combined with ⁇ CP1, CP2, CP4 ⁇ , ⁇ CP2, CP3, CP4 ⁇ , ⁇ CP1, CP3, CP4 ⁇ is converted into a control point ⁇ CP1, CP2, CP3 ⁇ to represent it.
- the conversion method is to substitute the motion vector of the control point and its coordinate information into formula (4) to obtain the model parameters, and then substitute the coordinate information of ⁇ CP1, CP2, CP3 ⁇ into formula (5) to obtain its motion vector.
- the conversion can be performed according to the following formulas (22)-(24), where W represents the width of the current block and H represents the height of the current block.
- (vx 0 , vy 0 ) Indicates a motion vector of CP1, (vx 1 , vy 1 ) indicates a motion vector of CP2, (vx 2 , vy 2 ) indicates a motion vector of CP3, and (vx 3 , vy 3 ) indicates a motion vector of CP4.
- a candidate motion vector list for the AMVP mode based on the affine motion model is constructed.
- the candidate motion vector list of the AMVP mode based on the affine motion model may be referred to as a control point motion vector predictor candidate list (control point motion vector predictor list), and the motion vector prediction value of each control point A motion vector including 2 (4-parameter affine motion model) control points or a motion vector including 3 (6-parameter affine motion model) control points.
- control point motion vector prediction value candidate list is pruned and sorted according to a specific rule, and it can be truncated or filled to a specific number.
- the motion vector prediction value of each control point in the control point motion vector prediction value candidate list is used to obtain the motion vector of each sub motion compensation unit in the current coding block by formulas (3) / (5), and then each The pixel value of the corresponding position in the reference frame pointed by the motion vector of each sub-motion compensation unit is used as its prediction value to perform motion compensation using an affine motion model. Calculate the average value of the difference between the original value and the predicted value of each pixel in the current coding block, select the control point motion vector prediction value corresponding to the smallest average value as the optimal control point motion vector prediction value, and use it as the current encoding Predicted motion vector for 2/3 control points. An index number indicating the position of the control point motion vector prediction value in the control point motion vector prediction value candidate list is encoded into a code stream and sent to the decoder.
- the index number is parsed, and the control point motion vector predictor (CPMVP) is determined from the control point motion vector prediction value candidate list according to the index number.
- CPMVP control point motion vector predictor
- a motion search is performed within a certain search range using the control point motion vector prediction value as a search starting point to obtain control point motion vectors (CPMV).
- CPMV control point motion vectors
- the difference between the control point motion vector and the control point motion vector prediction (CPMVD) is passed to the decoding end.
- control point motion vector difference value is analyzed and added to the control point motion vector prediction value to obtain the control point motion vector.
- a control point motion vector fusion candidate list is constructed.
- control point motion vector fusion candidate list is pruned and sorted according to a specific rule, and it can be truncated or filled to a specific number.
- each control point motion vector in the fusion candidate list is used to obtain each sub-motion compensation unit in the current coding block (the size divided by pixels or a specific method is N 1 ⁇ N) by formulas (3) / (5). 2 pixel blocks), and then obtain the pixel value of the position in the reference frame pointed to by the motion vector of each sub-motion compensation unit, and use it as the prediction value to perform affine motion compensation. Calculate the average value of the difference between the original value and the predicted value of each pixel in the current coding block, and select the control point motion vector corresponding to the smallest average value of the difference as the motion vector of 2/3 control points . An index number indicating the position of the control point motion vector in the candidate list is encoded into the code stream and sent to the decoder.
- the index number is parsed, and the control point motion vector (CPMV) is determined from the control point motion vector fusion candidate list according to the index number.
- CPMV control point motion vector
- At least one (a) of a, b, or c can be expressed as: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or multiple .
- a syntax element may be used to signal the inter prediction mode.
- x0, y0 represent the coordinates of the current block in the video image.
- the condition for adopting the merge mode based on the affine motion model may be that the width and height of the current block are both greater than or equal to 8.
- the syntax element affine_merge_flag [x0] [y0] may be used to indicate whether a merge mode based on an affine motion model is adopted for the current block.
- the type (slice_type) of the slice where the current block is located is P type or B type.
- affine_merge_flag [x0] [y0] 1 indicates that the merge mode based on the affine motion model is used for the current block
- affine_merge_flag [x0] [y0] 0 indicates that the merge mode based on the affine motion model is not used for the current block. You can use the merge mode of the leveling motion model.
- the syntax element affine_inter_flag [x0] [y0] may be used to indicate whether the affine motion model based AMVP mode is used for the current block when the current block is a P-type or B-type slice.
- allowAffineInter 1 indicates that the AMVP mode based on the affine motion model is adopted for the current block
- the AMVP mode of the translational motion model may be adopted.
- affine_type_flag [x0] [y0] can be used to indicate whether a 6-parameter affine motion model is used for motion compensation for the current block when the current block is a P-type or B-type slice.
- affine_type_flag [x0] [y0] 0, indicating that the 6-parameter affine motion model is not used for motion compensation for the current block, and only 4-parameter affine motion model can be used for motion compensation;
- affine_type_flag [x0] [y0] 1, indicating that A 6-parameter affine motion model is used for motion compensation for the current block.
- MotionModelIdc [x0] [y0] 1 indicates that a 4-parameter affine motion model is used
- the variable MaxNumMergeCand is used to indicate the maximum list length and indicates the maximum length of the constructed candidate motion vector list.
- inter_pred_idc [x0] [y0] is used to indicate the prediction direction.
- PRED_L1 is used to indicate backward prediction.
- num_ref_idx_l0_active_minus1 indicates the number of reference frames in the forward reference frame list
- ref_idx_l0 [x0] [y0] indicates the forward reference frame index value of the current block.
- mvd_coding (x0, y0,0,0) indicates the first motion vector difference.
- mvp_l0_flag [x0] [y0] indicates the forward MVP candidate list index value.
- PRED_L0 indicates forward prediction.
- num_ref_idx_l1_active_minus1 indicates the number of reference frames in the backward reference frame list.
- ref_idx_l1 [x0] [y0] indicates the backward reference frame index value of the current block
- mvp_l1_flag [x0] [y0] indicates the backward MVP candidate list index value.
- ae (v) represents a syntax element using context-based adaptive binary coding (cabac) coding.
- Step 601 Parse the code stream according to the syntax structure shown in Table 1 to determine the inter prediction mode of the current block.
- step 602a is performed.
- step 602b is performed.
- Step 602a Construct a candidate motion vector list corresponding to the AMVP mode based on the affine motion model, and execute step 603a.
- the candidate control point motion vector of the current block is derived to add the candidate motion vector list.
- the candidate motion vector list may include a two-tuple list (the current coding block is a 4-parameter affine motion model) or a three-tuple list.
- the two-tuple list includes one or more two-tuples for constructing a 4-parameter affine motion model.
- the triple list includes one or more triples for constructing a 6-parameter affine motion model.
- the candidate motion vector binary / triple list is pruned and sorted according to a specific rule, and it can be truncated or filled to a specific number.
- A1 The process of constructing a candidate motion vector list using the inherited control motion vector prediction method will be described.
- FIG. 4 Take FIG. 4 as an example. For example, in the order of A1-> B1-> B0-> A0-> B2 in FIG. 4, traverse the neighboring position blocks around the current block to find the affine coding block where the neighboring position blocks are located. To obtain the control point motion information of the affine coding block, and then construct a motion model based on the control point motion information of the affine coding block, and derive the candidate control point motion information of the current block. Specifically, reference may be made to the related description in the 3) inherited control point motion vector prediction method, which is not repeated here.
- the affine decoding block is an affine coding block that is predicted using an affine motion model during the encoding phase.
- the motion vectors of the upper left and upper right control points of the current block are derived according to the 4 affine motion model formulas (6) and (7), respectively. .
- the motion vectors of the three control points of the adjacent affine decoding block are obtained, for example, in FIG. 4, the motion vector value (vx4 of the upper left control point (x4, y4)) , vy4) and the motion vector value (vx5, vy5) of the upper right control point (x5, y5) and the motion vector (vx6, vy6) of the lower left vertex (x6, y6).
- the 6-parameter motion model formulas (8), (9), and (10) are deduced to obtain the top left, top right, and bottom left controls of the current block, respectively. Dot motion vector.
- the motion vectors of the three control points of the adjacent affine decoding block are obtained, for example, in FIG. 4, the upper left control point (x4, y4 )
- the 6-parameter affine motion model consisting of 3 control points of adjacent affine decoding blocks is derived from the corresponding formulas (8), (9), and (10) of the 6-parameter affine motion model.
- the affine motion model used by the adjacent affine decoding block is a 4-parameter affine motion model, then obtain the motion vectors of the two control points of the affine decoding block: the motion vector value (vx4 of the upper left control point (x4, y4)) , vy4) and the motion vector value (vx5, vy5) of the upper right control point (x5, y5).
- the 4-parameter affine motion model consisting of 2 control points of adjacent affine decoding blocks is used to derive the motion vectors of the upper-left and upper-right 2 control points of the current block according to the 4-parameter affine motion model formulas (6) and (7), respectively. .
- A2 The process of constructing a candidate motion vector list using a control method for controlling a motion vector prediction will be described.
- the affine motion model used by the current decoding block is a 4-parameter affine motion model (that is, MotionModelIdc is 1), and the motion information of the neighboring coded blocks around the current coding block is used to determine the upper left vertex and the upper right vertex of the current coding block Motion vector.
- the constructed control point motion vector prediction method 1 or the constructed control point motion vector prediction method 2 may be used to construct a candidate motion vector list. For specific methods, refer to the descriptions in 4) and 5) above, which will not be repeated here. .
- the current decoding block affine motion model is a 6-parameter affine motion model (that is, MotionModelIdc is 2), and the motion information of the coded blocks adjacent to the current coding block is used to determine the upper left vertex and upper right vertex and the lower left of the current coding block. Vertices motion vector.
- the constructed control point motion vector prediction method 1 or the constructed control point motion vector prediction method 2 may be used to construct a candidate motion vector list. For specific methods, refer to the descriptions in 4) and 5) above, which will not be repeated here. .
- control point motion information can also be applied to this application, and details are not described herein.
- Step 603a Parse the code stream, determine the optimal control point motion vector prediction value, and execute step 604a.
- affine motion model used in the current decoding block is a 4-parameter affine motion model (MotionModelIdc is 1)
- the index number is parsed, and the optimal motion vector prediction of 2 control points is determined from the candidate motion vector list according to the index number value.
- the index number is mvp_l0_flag or mvp_l1_flag.
- the affine motion model used in the current decoding block is a 6-parameter affine motion model (MotionModelIdc is 2)
- the index number is parsed, and the optimal motion vector prediction of 3 control points is determined from the candidate motion vector list according to the index number value.
- Step 604a Parse the code stream and determine the motion vector of the control point.
- the affine motion model used in the current decoding block is a 4-parameter affine motion model (MotionModelIdc is 1), and the motion vector difference between the two control points of the current block is decoded from the code stream.
- the vector difference value and the motion vector prediction value obtain the motion vector value of the control point.
- the forward prediction is taken as an example.
- the motion vector difference between the two control points is mvd_coding (x0, y0,0,0) and mvd_coding (x0, y0,0,1).
- the motion vector difference values of the upper left position control point and the upper right position control point are obtained from the code stream, and are respectively added to the motion vector prediction value to obtain the motion vectors of the upper left position control point and the upper right position control point of the current block. value.
- the current decoding block affine motion model is a 6-parameter affine motion model (MotionModelIdc is 2)
- the motion vector differences of the three control points of the current block are obtained by decoding from the code stream, and the motion vector values of the control points are obtained according to the motion vector difference value and the motion vector prediction value of each control point.
- the forward prediction is taken as an example.
- the motion vector differences of the three control points are mvd_coding (x0, y0,0,0), mvd_coding (x0, y0,0,1), and mvd_coding (x0, y0,0,2).
- the motion vector difference values of the upper left control point, the upper right control point, and the lower left control point are obtained from the code stream and are respectively added to the motion vector prediction values to obtain the upper left control point, upper right control point, and lower left control of the current block.
- the point's motion vector value is obtained from the code stream and are respectively added to the motion vector prediction values to obtain the upper left control point, upper right control point, and lower left control of the current block.
- Step 602b Construct a motion information candidate list based on the merge mode of the affine motion model.
- the inherited control point motion vector prediction method and / or the constructed control point motion vector prediction method may be used to construct a motion information candidate list based on an affine motion model fusion mode.
- the motion information candidate list is pruned and sorted according to a specific rule, and it can be truncated or filled to a specific number.
- candidate control point motion information of the current block is derived and added to the motion information candidate list.
- A1, B1, B0, A0, and B2 in FIG. 3 it traverses the surrounding neighboring block of the current block, finds the affine coding block at that position, and obtains the control point motion information of the affine coding block, and then passes through it.
- the motion model derives candidate control point motion information for the current block.
- the candidate control point motion information is added to the candidate list; otherwise, the motion information in the candidate motion vector list is traversed in order to check whether there is motion in the candidate motion vector list with the candidate control point. The same sports information. If the candidate motion vector list does not have the same motion information as the candidate control point motion information, the candidate control point motion information is added to the candidate motion vector list.
- judging whether the two candidate motion information are the same requires judging whether their forward and backward reference frames and the horizontal and vertical components of each forward and backward motion vector are the same. Only when all the above elements are different, the two motion information are considered different.
- MaxNumMrgCand is a positive integer, such as 1, 2, 3, 4, 5, etc., the following uses 5 as an example to describe, no further details
- D2 Use the constructed control point motion vector prediction method to derive candidate control point motion information for the current block and add the motion information candidate list, as shown in FIG. 6B.
- step 601c motion information of each control point of the current block is acquired. Please refer to the control point motion vector prediction method 2 constructed in 5), step 501, which is not repeated here.
- step 602c the movement information of the control points is combined to obtain the structured movement information of the control points.
- step 501 in FIG. 5B which is not described herein again.
- Step 603c adding the constructed control point motion information to the candidate motion vector list.
- the combinations are traversed in a preset order to obtain a valid combination as the candidate control point motion information. If the candidate motion vector list is empty at this time, the candidate The control point motion information is added to the candidate motion vector list; otherwise, the motion information in the candidate motion vector list is sequentially traversed to check whether there is motion information in the candidate motion vector list that is the same as the candidate control point motion information. If the candidate motion vector list does not have the same motion information as the candidate control point motion information, the candidate control point motion information is added to the candidate motion vector list.
- a preset sequence is as follows: Affine (CP1, CP2, CP3)-> Affine (CP1, CP2, CP4)-> Affine (CP1, CP3, CP4)-> Affine (CP2, CP3, CP4)-> Affine (CP2, CP3, CP4) -> Affine (CP1, CP2)-> Affine (CP1, CP3)-> Affine (CP2, CP3)-> Affine (CP1, CP4)-> Affine (CP2, CP4)-> Affine (CP3, CP4), total 10 combinations.
- control point motion information corresponding to a combination is not available, the combination is considered to be unavailable. If a combination is available, determine the reference frame index of the combination (when two control points, the smallest reference frame index is selected as the reference frame index of the combination; when it is greater than two control points, the reference frame index with the most occurrences is selected first. If there are as many occurrences of multiple reference frame indexes, the one with the smallest reference frame index is selected as the combined reference frame index), and the motion vector of the control point is scaled. If the motion information of all control points after scaling is consistent, the combination is illegal.
- the embodiments of the present application may also fill the candidate motion vector list. For example, after the traversal process described above, the length of the candidate motion vector list is less than the maximum list length MaxNumMrgCand, and the candidate motion vector list may be filled. Until the length of the list is equal to MaxNumMrgCand.
- It can be filled by a method of supplementing zero motion vectors, or by a method of combining and weighted average of motion information of existing candidates in an existing list. It should be noted that other methods for obtaining candidate motion vector list filling can also be applied to the present application, and details are not described herein.
- Step S603b Parse the code stream to determine the optimal control point motion information.
- Step 604b Obtain a motion vector value of each sub-block in the current block according to the optimal control point motion information and the affine motion model adopted by the current decoding block.
- the preset position pixels in the motion compensation unit can be used Point motion information to represent the motion information of all pixels in the motion compensation unit. Assuming the size of the motion compensation unit is MxN, the preset position pixels can be the center point of the motion compensation unit (M / 2, N / 2), the upper left pixel (0, 0), and the upper right pixel (M-1,0 ), Or pixels at other locations.
- M / 2, N / 2 the center point of the motion compensation unit
- V 0 represents the motion vector of the upper left control point
- V 1 represents the motion vector of the upper right control point.
- Each small box represents a motion compensation unit.
- the coordinates of the center point of the motion compensation unit relative to the top left pixel of the current affine decoding block are calculated using formula (25), where i is the i-th motion compensation unit in the horizontal direction (from left to right), and j is the j-th vertical direction.
- Motion compensation units (from top to bottom), (x (i, j) , y (i, j) ) represents the pixel of the (i, j) th center of the motion compensation unit relative to the upper left control point of the current affine decoding block coordinate of.
- the affine motion model used in the current affine decoding block is a 6-parameter affine motion model
- substitute (x (i, j) , y (i, j) ) into the 6-parameter affine motion model formula (26) and obtain each
- the motion vectors of the center points of each motion compensation unit are used as the motion vectors (vx (i, j) , vy (i, j) ) of all pixels in the motion compensation unit.
- the affine motion model used in the current affine decoding block is a 4 affine motion model
- the motion vector of the center point of the motion compensation unit is used as the motion vector (vx (i, j) , vy (i, j) ) of all pixels in the motion compensation unit.
- Step 605b For each sub-block, perform motion compensation according to the determined motion vector value of the sub-block to obtain a pixel prediction value of the sub-block.
- steps 606a and 605b are used to perform motion compensation of the sub-block.
- a sub-block is divided into 4x4, that is, each 4x4 unit uses a different motion vector for motion compensation.
- the smaller the motion compensation unit the larger the number of reference pixels to be read when the average pixel performs motion compensation, and the higher the complexity of the interpolation operation.
- the total number of reference pixels required is (M + T-1) * (N + T-1) * K
- the average number of read pixels is (M + T-1) * (N + T-1) * K / M / N.
- T is the number of taps of the interpolation filter, such as 8, 4, 2, etc.
- the interpolation filter tap is 8
- the average number of reference pixel reads for 4x4 units, 8x4 units, and 8x8 units for unidirectional and bidirectional prediction can be calculated, as shown in Table 3.
- M in MxN indicates that the width of the sub-block is M pixels
- N indicates that the height of the sub-block is N pixels. It should be understood that both M and N are 2 n , and n is a positive integer. It should be noted that UNI in Table 3 indicates unidirectional prediction, and BI indicates bidirectional prediction.
- embodiments of the present application provide an image prediction method and device, which are used to reduce the complexity of motion compensation in the prior art while taking into account the prediction efficiency.
- different sub-block division methods are selected according to the prediction direction of the current image block.
- the method and the device are based on the same inventive concept. Since the principle of the method and the device for solving the problem is similar, the implementation of the device and the method can be referred to each other, and duplicated details will not be repeated.
- FIG. 7A is a schematic flowchart of an image prediction method 700 according to an embodiment of the present application. It should be noted that the image prediction method 700 is applicable to both the inter prediction of decoded video images and the inter prediction of encoded video images. The method may be performed by a video encoder (such as video encoder 20) or having For an electronic device with a video encoding function, the method 700 may include the following steps:
- Step S701 acquiring a motion vector of a control point of a current image block (for example, a current decoding block, a current affine image block, or a current affine decoding block);
- a current image block for example, a current decoding block, a current affine image block, or a current affine decoding block
- a motion vector of a control point of a current affine decoding block may be obtained according to the following steps.
- Step 1 Determine the predicted value of the motion vector of the control point of the current affine decoding block
- the affine motion model adopted by the current decoding block is a 4-parameter affine motion model, and the motion information of the decoded blocks adjacent to the current decoded block is used to determine the predicted values of the motion vectors of the upper left vertex and the upper right vertex of the current decoded block. Specifically, in the order of A2, B2, and B3 in FIG. 5A, the surrounding neighboring position blocks of the current block are traversed to find the motion vector in the same prediction direction as the current decoded block.
- the motion vector is scaled as the current The predicted value of the motion vector of the top left vertex control point of the affine decoded block; if no motion vector is found in the adjacent positions in the traversal space, the zero motion vector is used as the predicted value of the motion vector of the top left vertex control point of the current affine decoded block. .
- the surrounding neighboring position blocks of the current block are traversed to find the motion vector in the same prediction direction as the current decoded block. If found, the motion vector is scaled as the current affine.
- the predicted value of the motion vector of the top right vertex control point of the decoded block if no motion vector is found in the traversal of adjacent positions in the spatial domain, a zero motion vector is used as the predicted value of the motion vector of the top right vertex control point of the current affine decoded block.
- the affine motion model adopted by the current decoding block is a 6-parameter affine motion model.
- the motion information of the upper left vertex, upper right vertex, and lower left vertex of the current decoding block is determined by using the motion information of the decoded blocks adjacent to the current decoding block. Predictive value. Specifically, according to the sequence of A2, B2, and B3 in FIG. 5A, the surrounding neighboring position blocks of the current block are traversed to find the motion vector in the same prediction direction as the current decoded block.
- the motion vector is scaled as the current The predicted value of the motion vector of the top left vertex control point of the affine decoded block; if no motion vector is found in the adjacent positions in the traversal space, the zero motion vector is used as the predicted value of the motion vector of the top left vertex control point of the current affine decoded block. .
- the surrounding neighboring position blocks of the current block are traversed to find the motion vector in the same prediction direction as the current decoded block. If found, the motion vector is scaled as the current affine.
- the predicted value of the motion vector of the top right vertex control point of the decoded block if no motion vector is found in the traversal of adjacent positions in the spatial domain, a zero motion vector is used as the predicted value of the motion vector of the top right vertex control point of the current affine decoded block.
- the surrounding neighboring position blocks of the current block are searched to find the motion vector in the same prediction direction as the current decoded block. If found, the motion vector is scaled and used as the current affine decoded block.
- the predicted value of the motion vector of the lower left vertex control point if no motion vector is found during traversal of adjacent positions in the spatial domain, the zero motion vector is used as the predicted value of the motion vector of the lower left vertex control point of the current affine decoding block.
- Step 2 Parse the code stream and determine the motion vector of the control point.
- the affine motion model used in the current decoding block is a 4-parameter affine motion model
- the motion vector differences of the two control points of the current block are decoded from the code stream, and the motion vector differences and motion vectors of each control point are obtained respectively.
- the predicted value obtains the motion vector of the control point.
- the motion vector difference between the upper-left vertex control point and the upper-right vertex control point is decoded from the code stream, and respectively added to the corresponding motion vector prediction value to obtain the values of the upper-left vertex control point and the upper-right vertex control point of the current block.
- Motion vector value is decoded from the code stream, and respectively added to the corresponding motion vector prediction value to obtain the values of the upper-left vertex control point and the upper-right vertex control point of the current block.
- the affine motion model of the current decoding block is a 6-parameter affine motion model
- the motion vector differences of the three control points of the current block are decoded from the code stream, and obtained according to the motion vector difference values of the control points and the motion vector prediction values. Control point motion vector.
- the motion vector difference values of the upper-left vertex control point, the upper-right vertex control point, and the lower-left vertex control point are decoded from the code stream, and are respectively added to the corresponding motion vector prediction values to obtain the upper-left vertex control point, The motion vector values for the upper right vertex control point and the lower left vertex control point.
- the method of obtaining the motion vector of the control point of the current image block (such as the current decoding block, the current affine image block, or the current affine decoding block) is not limited in this article, and other acquisition methods may also be applicable to this method. Application will not be repeated here.
- Step S703 Use an affine transformation model to obtain the motion vector of each sub-block in the current image block according to the motion vector (for example, a plurality of affine control points) of the current image block using an affine transformation model.
- the size of the sub-block is determined based on the prediction direction of the current image block;
- FIG. 7B illustrates a 4 ⁇ 4 sub-block (also referred to as a motion compensation unit)
- FIG. 7C illustrates an 8 ⁇ 8 sub-block (also referred to as a motion compensation unit).
- the corresponding sub-block also referred to as a motion compensation unit
- the center point of is represented by a triangle.
- Step S705 Perform motion compensation according to a motion vector value of each sub-block in the current image block to obtain a pixel prediction value of each sub-block.
- step S703 in a specific implementation manner of the embodiment of the present application, in step S703:
- the size of the subblock in the current image block is UxV; or, if the prediction direction of the current image block is unidirectional prediction, the subblock in the current image block is The size is MxN, where U, M represents the width of the sub-block, V, N represents the height of the sub-block, and U, V, M, N are all 2n, and n is a positive integer.
- M 4 and N is 4. Accordingly, U is 8 and V is 8.
- the bi-predicted sub-blocks are divided into UxV.
- M is an integer such as 4, 8, or 16 and N is an integer such as 4, 8, or 16.
- the affine decoding block uses unidirectional prediction or bidirectional prediction, which is determined by the syntax element inter_pred_idc.
- the affine decoding block uses unidirectional prediction or bidirectional prediction, and is determined by affine_merge_idx.
- the prediction direction of the current affine decoding block is the same as the candidate motion information indicated by affine_merge_idx.
- the size of the decoding unit does not satisfy the use conditions of the affine decoding block, it is not necessary to parse affine-related syntax elements, such as affine_inter_flag and affine_merge_flag in Table 1.
- a one-way prediction sub-block partitioning method may also be adopted.
- the method of the present application can also be used in other sub-block division modes, such as the ATMVP mode.
- the size of its motion compensation unit is set to 4x4, and if the affine decoding block is bidirectional prediction, its motion compensation is set.
- the size of the unit is set to 8x4.
- an affine mode is allowed.
- the width W of the decoded block is less than 16, if the prediction direction of the affine decoded block is bidirectional, it is modified to be unidirectional. For example, discarding backward predicted motion information and converting it into forward prediction; or discarding forward predicted motion information and converting it into backward prediction.
- the code stream can also be restricted, so that when the width W of the affine decoded block W ⁇ 16, the flag bits for bidirectional prediction need not be parsed.
- the size of its motion compensation unit is set to 4x4
- its motion compensation unit is set to Set the size to 4x8.
- an affine mode is allowed.
- the height H of the decoded block is less than 16, if the prediction direction of the affine decoded block is bidirectional, it is modified to be unidirectional. For example, discarding backward predicted motion information and converting it into forward prediction; or discarding forward predicted motion information and converting it into backward prediction.
- code stream may also be limited, so that when the height H of the affine decoded block H ⁇ 16, it is not necessary to parse the flag bits for bidirectional prediction.
- the affine decoding block is unidirectional prediction
- the size of its motion compensation unit is set to 4x4
- the affine decoding block is bidirectional prediction
- the affine decoding block is based on the affine decoding block.
- Size for adaptive division The division can be one of the following three ways:
- an affine mode is allowed.
- the width W of the decoding block is equal to 8 and the height H is equal to 8
- the prediction direction of the affine decoding block is bidirectional, it is modified to be unidirectional. For example, discarding backward predicted motion information and converting it into forward prediction; or discarding forward predicted motion information and converting it into backward prediction.
- the code stream can also be limited, so that when the width W of the affine decoded block is equal to 8 and the height H is equal to 8, it is not necessary to analyze the flag bits for bidirectional prediction.
- the current decoding block in the prior art which is divided into MxN (that is, 4x4) sub-blocks, that is, each MxN (that is, 4x4) sub-blocks use corresponding motion vectors for motion compensation.
- the current The size of the sub-blocks in the decoded block is determined based on the prediction direction of the current decoded block; for example, if the current decoded block is unidirectional prediction, the size of the sub-block of the current decoded block is 4x4; if the current decoded block is bi-directionally predicted , Then the size of the sub-block of the current decoding block is 8x8.
- the size of the subblocks (or motion compensation units) of certain image blocks in the embodiments of the present application is relatively larger than the subblocks (or motion compensation units) in the prior art.
- the average number of read reference pixels required for each pixel for motion compensation is relatively small, and the computational complexity of interpolation is relatively low. Therefore, the embodiment of the present application reduces the motion compensation to some extent while taking into account the prediction efficiency. Complexity, which improves codec performance.
- FIG. 7D is a flowchart illustrating a process 1200 of a decoding method according to an embodiment of the present application.
- the process 1200 may be performed by the video decoder 30, and specifically, may be performed by an inter prediction unit of the video decoder 30 and an entropy decoding unit (also referred to as an entropy decoder).
- the process 1200 is described as a series of steps or operations. It should be understood that the process 700 may be performed in various orders and / or concurrently, and is not limited to the execution order shown in FIG. 7D. Assuming a video data stream with multiple video frames is using a video decoder, the process shown in FIG. 7D is described as follows:
- the size of the sub-block in the current affine decoding block is determined based on the prediction direction of the current decoding block (such as unidirectional prediction or bidirectional prediction), or the size of the sub-block. It is determined based on the prediction direction of the current decoded block and the size of the current decoded block. The following process will not be repeated, and can refer to the previous embodiment.
- Step S1201 The video decoder determines an inter prediction mode of the currently decoded block.
- the inter prediction mode may be an advanced motion vector prediction (Advanced Vector Prediction (AMVP) mode) or may be a merge mode.
- AMVP Advanced Vector Prediction
- steps S1211-S1216 are performed.
- steps S1221-S1225 are performed.
- Step S1211 The video decoder constructs a candidate motion vector prediction value MVP list.
- the video decoder uses an inter prediction unit (also referred to as an inter prediction module) to construct a candidate motion vector prediction value MVP list (also referred to as a candidate motion vector list).
- an inter prediction unit also referred to as an inter prediction module
- a candidate motion vector prediction value MVP list also referred to as a candidate motion vector list.
- the candidate motion vector prediction value MVP list can be a triplet candidate motion vector prediction value MVP list or a binary tuple candidate motion vector prediction value MVP List; the above two methods are as follows:
- Method 1 A motion vector prediction method based on a motion model is used to construct a candidate motion vector prediction value MVP list.
- all or part of neighboring blocks of the current decoding block are traversed in a predetermined order to determine the neighboring affine decoding blocks, and the number of the determined neighboring affine decoding blocks may be one or more.
- the neighboring blocks A, B, C, D, and E shown in FIG. 7A may be traversed in order to determine neighboring affine decoding blocks among the neighboring blocks A, B, C, D, and E.
- the inter prediction unit determines at least one set of candidate motion vector prediction values (each set of candidate motion vector prediction values is a two-tuple or three-tuple) according to at least one neighboring affine decoding block, and an adjacent affine is used below
- the decoding block is introduced as an example.
- the adjacent affine decoding block is called the first adjacent affine decoding block, as follows:
- a first affine model is determined according to a motion vector of a control point of a first adjacent affine decoding block, and then a motion vector of a control point of the current decoding block is predicted according to the first affine model.
- the method of predicting the motion vector of the control point of the current decoding block based on the motion vector of the control point of the first adjacent affine decoding block is also different, so the following description will be made on a case-by-case basis.
- the parameter model of the current decoding block is a 4-parameter affine transformation model:
- the first phase is obtained
- the motion vectors of the bottom two control points of the adjacent affine decoding block For example, the position coordinates (x 6 , y 6 ) and motion vectors (vx 6 , vy) of the lower left control point of the first adjacent affine decoding block can be obtained. 6), and a lower right position of the control point coordinates (x 7, y 7) and motion vector values (vx 7, vy 7) (step S1201).
- a first affine model is formed according to the motion vectors and coordinate positions of the bottom two control points of the first adjacent affine decoding block (the first affine model obtained at this time is a 4-parameter affine model) (step S1202) .
- the motion vector of the control point of the current decoding block is predicted according to the first affine model. For example, the position coordinates of the upper left control point and the position of the upper right control point of the current decoding block may be brought into the first affine model, respectively. , Thereby predicting the motion vector of the upper left control point and the motion vector of the upper right control point of the current decoding block, as shown in formulas (1) and (2) (step S1203).
- (x 0 , y 0 ) are the coordinates of the upper left control point of the current decoded block, and (x 1 , y 1 ) are the coordinates of the upper right control point of the current decoded block; in addition, ( vx 0 , vy 0 ) is a motion vector of the upper left control point of the predicted current decoded block, and (vx 1 , vy 1 ) is a motion vector of the upper right control point of the predicted current decoded block.
- the position coordinates (x 6 , y 6 ) of the lower left control point of the first adjacent affine decoding block and the position coordinates (x 7 , y 7 ) of the lower right control point are both based on the The position coordinates (x 4 , y 4 ) of the upper left control point of the first adjacent affine decoding block are calculated, where the position coordinates (x 6 , y) of the lower left control point of the first adjacent affine decoding block 6 ) is (x 4 , y 4 + cuH), and the position coordinates (x 7 , y 7 ) of the lower right control point of the first adjacent affine decoding block is (x 4 + cuW, y 4 + cuH) , CuW is the width of the first neighboring affine decoding block, and cuH is the height of the first neighboring affine decoding block.
- the motion vector of the lower left control point of the first neighboring affine decoding block is A motion vector of a lower left sub-block of the first adjacent affine decoding block
- a motion vector of a lower right control point of the first adjacent affine decoding block is a Motion vector.
- the first neighboring affine decoding block is located in a Coding Tree Unit (CTU) above the current decoding block and the first neighboring affine decoding block is a six-parameter affine decoding block, it is not based on the first phase
- the neighboring affine decoding block generates a candidate motion vector prediction value of a control point of the current block.
- CTU Coding Tree Unit
- the manner of predicting the motion vector of the control point of the current decoding block is not limited here.
- an optional determination method is also exemplified below:
- the position coordinates and motion vectors of the three control points of the first adjacent affine decoding block may be obtained, for example, the position coordinates (x 4 , y 4 ) and motion vector values (vx 4 , vy 4 ) of the upper left control point, an upper right position coordinates of the control points (x 5, y 5) and a motion vector value (vx 5, vy 5), the position coordinates of the lower left control point (x 6, y 6) and the motion vector (vx 6, vy 6).
- a 6-parameter affine is formed according to the position coordinates and motion vectors of the three control points of the first adjacent affine decoding block
- the position coordinates (x 0 , y 0 ) of the upper left control point and the position coordinates (x 1 , y 1 ) of the upper right control point of the current decoding block are substituted into a 6-parameter affine model to predict the motion vector of the upper left control point of the current decoding block And the motion vector of the upper right control point, as shown in formulas (4) and (5).
- (vx 0 , vy 0 ) is the motion vector of the upper left control point of the predicted current decoded block
- (vx 1 , vy 1 ) is the predicted upper right control point of the current decoded block. Motion vector.
- the parameter model of the current decoding block is a 6-parameter affine transformation model.
- the derivation method can be:
- the first adjacent affine decoding block is located above the CTU of the current decoding block and the first adjacent affine decoding block is a four-parameter affine decoding block, then the two lowermost sides of the first adjacent affine decoding block
- the position coordinates and motion vectors of the control points for example, the position coordinates (x 6 , y 6 ) and motion vectors (vx 6 , vy 6 ) of the lower left control point of the first adjacent affine decoding block can be obtained, and the lower right control
- a first affine model is formed according to the motion vectors of the two control points at the bottom of the first adjacent affine decoding block (the first affine model obtained at this time is a 4-parameter affine model).
- the motion vector of the control point of the current decoded block is predicted according to the first affine model. For example, the position coordinates of the upper left control point, the position of the upper right control point, and the position of the lower left control point of the current decoded block may be brought in respectively.
- To the first affine model thereby predicting the motion vector of the upper left control point, the motion vector of the upper right control point, and the motion of the lower left control point of the current decoding block, as shown in formulas (1), (2), (3) Show.
- Formulas (1) and (2) have been described above.
- (x 0 , y 0 ) are the coordinates of the upper-left control point of the current decoding block
- (x 1 , y 1 ) is the coordinate of the upper right control point of the current decoded block
- (x 2 , y 2 ) is the coordinate of the lower left control point of the current decoded block
- (vx 0 , vy 0 ) is the predicted upper left control of the current decoded block
- the motion vector of the point (vx 1 , vy 1 ) is the motion vector of the predicted upper right control point of the current decoded block
- (vx 2 , vy 2 ) is the predicted motion vector of the right and left lower control point of the current decoded block.
- the first neighboring affine decoding block is located in a Coding Tree Unit (CTU) above the current decoding block and the first neighboring affine decoding block is a six-parameter affine decoding block, it is not based on the first phase
- the neighboring affine decoding block generates a candidate motion vector prediction value of a control point of the current block.
- CTU Coding Tree Unit
- the manner of predicting the motion vector of the control point of the current decoding block is not limited here.
- an optional determination method is also exemplified below:
- the position coordinates and motion vectors of the three control points of the first adjacent affine decoding block may be obtained, for example, the position coordinates (x 4 , y 4 ) and motion vector values (vx 4 , vy 4 ) of the upper left control point, an upper right position coordinates of the control points (x 5, y 5) and a motion vector value (vx 5, vy 5), the position coordinates of the lower left control point (x 6, y 6) and the motion vector (vx 6, vy 6).
- a 6-parameter affine is formed according to the position coordinates and motion vectors of the three control points of the first adjacent affine decoding block
- the affine model predicts the motion vector of the upper left control point, the motion vector of the upper right control point, and the motion vector of the lower left control point of the current decoding block, as shown in formulas (4), (5), and (6).
- Formulas (4) and (5) have been described previously.
- (vx 0 , vy 0 ) is the predicted motion vector of the upper-left control point of the current decoding block
- (vx 1 , vy 1 ) are the motion vectors of the predicted upper right control point of the current decoded block
- (vx 2 , vy 2 ) are the predicted motion vectors of the lower left control point of the current decoded block.
- Method 2 A motion vector prediction method based on the combination of control points is used to construct a candidate motion vector prediction value MVP list.
- the parameter model of the current decoding block does not have the same way of constructing the candidate motion vector prediction value MVP list at the same time, which will be described below.
- the parameter model of the current decoding block is a 4-parameter affine transformation model.
- the derivation method can be:
- the motion information of the upper left vertex and the upper right vertex of the current decoded block is estimated using the motion information of the decoded blocks adjacent to the current decoded block.
- the motion vector of the upper left vertex adjacent to the decoded block A and / or B and / or C block is used as the candidate motion vector of the motion vector of the upper left vertex of the current decoded block;
- the motion vector of the decoding block D and / or E block is used as a candidate motion vector of the motion vector of the upper right vertex of the current decoding block.
- a candidate motion vector of the upper left vertex and a candidate motion vector of the upper right vertex are combined to obtain a set of candidate motion vector prediction values. Multiple records obtained by combining in this combination manner can form a candidate motion vector prediction value MVP list.
- the current decoding block parameter model is a 6-parameter affine transformation model.
- the derivation method can be:
- the motion information of the upper left vertex and the upper right vertex of the current decoded block is estimated using the motion information of the decoded blocks adjacent to the current decoded block.
- the motion vector of the upper left vertex adjacent to the decoded block A and / or B and / or C block is used as the candidate motion vector of the motion vector of the upper left vertex of the current decoded block;
- the motion vector of the decoded block D and / or E block is used as the candidate motion vector of the motion vector of the upper right vertex of the current decoded block;
- the motion vector of the decoded block F and / or G block adjacent to the lower left vertex is used as the lower left vertex of the current decoded block Candidate motion vector.
- a set of candidate motion vector prediction values can be obtained by combining the above candidate motion vector of the upper left vertex, the candidate motion vector of the upper right vertex, and the candidate motion vector of the lower left vertex. Multiple sets of candidate motion vectors obtained by combining in this combination manner
- the prediction value may constitute a candidate motion vector prediction value MVP list.
- the candidate motion vector prediction value MVP list can only be constructed by using the candidate motion vector prediction value predicted by the first method, or the candidate motion vector prediction value MVP can be constructed by only using the candidate motion vector prediction value obtained by the second method prediction.
- the candidate motion vector prediction value obtained by the prediction method 1 and the candidate motion vector prediction value obtained by the method 2 prediction can be used to jointly construct a candidate motion vector prediction value MVP list.
- the candidate motion vector prediction value MVP list can be pruned and sorted according to a pre-configured rule, and then truncated or filled to a specific number.
- the candidate motion vector prediction value MVP list When each group of candidate motion vector prediction values in the candidate motion vector prediction value MVP list includes motion vector prediction values of three control points, the candidate motion vector prediction value MVP list may be called a triple list; when the candidate motion vector When each group of candidate motion vector prediction values in the prediction value MVP list includes motion vector prediction values of two control points, the candidate motion vector prediction value MVP list may be referred to as a two-tuple list.
- Step S1212 The video decoder parses the bitstream to obtain an index and a motion vector difference MVD.
- the video decoder may parse the bitstream through an entropy decoding unit, and the index is used to indicate a target candidate motion vector group of the current decoding block, where the target candidate motion vector represents a motion vector prediction value of a set of control points of the current decoding block.
- Step S1213 The video decoder determines a target motion vector group from the candidate motion vector prediction value MVP list according to the index.
- the video decoder determines the target candidate motion vector group determined from the candidate MVP list according to the index as the optimal candidate motion vector prediction value (optionally, when the length of the candidate motion vector prediction value MVP list is 1) (You do not need to parse the bitstream to get the index, you can directly determine the target motion vector group.) The following briefly introduces the optimal candidate motion vector prediction value.
- the optimal motion vector prediction value of 2 control points is selected from the candidate motion vector prediction value MVP list established above; for example, the video decoder from The index number is parsed in the bitstream, and the optimal motion vector prediction value of 2 control points is determined from the candidate motion vector prediction value MVP list of the binary group according to the index number, and each group of candidate motions in the candidate motion vector prediction value MVP list
- the vector prediction values correspond to their respective index numbers.
- the optimal motion vector prediction value of 3 control points is selected from the candidate motion vector prediction value MVP list established above; for example, the video decoder from The index number is parsed in the code stream, and the optimal motion vector prediction value of the three control points is determined from the triplet candidate motion vector prediction value MVP list according to the index number.
- the candidate motion vector prediction value MVP list includes each candidate motion The vector prediction values correspond to their respective index numbers.
- Step S1214 The video decoder determines the motion vector of the control point of the current decoding block according to the target candidate motion vector group and the motion vector difference MVD parsed from the code stream.
- the motion vector difference values of the two control points of the current decoding block are obtained by decoding from the code stream, respectively according to the motion vector difference values of the control points and the index indication.
- a new candidate motion vector group For example, the motion vector difference MVD of the upper left control point and the motion vector difference MVD of the upper right control point are decoded from the code stream, and are respectively added to the motion vectors of the upper left control point and the upper right control point in the target candidate motion vector group to thereby A new candidate motion vector group is obtained. Therefore, the new candidate motion vector group includes new motion vector values of the upper left control point and the upper right control point of the current decoding block.
- the motion vector value of the third control point may also be obtained by using a 4-parameter affine transformation model based on the motion vector values of the two control points of the current decoding block in the new candidate motion vector group. For example, the motion vector (vx 0 , vy 0 ) of the upper left control point and the motion vector (vx 1 , vy 1 ) of the upper right control point of the current decoded block are obtained, and then the formula (7) is used to obtain the lower left control point (x of the current decoded block). 2 , y 2 ) 's motion vector (vx 2 , vy 2 ).
- (x 0 , y 0 ) is the position coordinates of the upper left control point
- (x 1 , y 1 ) is the position coordinates of the upper right control point
- W is the width of the current decoded block
- H is the height of the current decoded block.
- the motion vector difference values of the three control points of the current decoding block are obtained by decoding from the code stream, respectively according to the motion vector difference values of each control point and the MVD and the index indication To obtain a new candidate motion vector group.
- the motion vector difference MVD of the upper left control point, the motion vector difference MVD of the upper right control point, and the motion vector difference of the lower left control point are obtained from the code stream, and are respectively compared with the upper left control point, The motion vectors of the upper right control point and the lower left control point are added to obtain a new candidate motion vector group. Therefore, the new candidate motion vector group includes the motion vector values of the upper left control point, the upper right control point, and the lower left control point of the current decoding block.
- Step S1215 The video decoder uses the affine transformation model to obtain the motion vector value of each sub-block in the current decoding block according to the motion vector value of the control point of the current decoding block determined above.
- the size of the sub-block is based on the current The prediction direction of the image block is determined.
- the new candidate motion vector group obtained based on the target candidate motion vector group and the MVD includes two (upper left control points and upper right control points) or three control points (for example, upper left control point, upper right control point, and lower left control). Dot) motion vector.
- the motion information of pixels at preset positions in the motion compensation unit can be used to represent the motion of all pixels in the motion compensation unit information.
- the preset position pixels can be the center point of the motion compensation unit (M / 2, N / 2), the upper left pixel (0, 0), and the upper right pixel ( M-1,0), or other locations.
- FIG. 7C illustrates a 4x4 motion compensation unit
- FIG. 7D illustrates an 8x8 motion compensation unit.
- the coordinates of the center point of the motion compensation unit relative to the top left pixel of the current decoding block are calculated using formula (8-1), where i is the i-th motion compensation unit in the horizontal direction (from left to right), and j is the j-th vertical direction.
- Motion compensation units (from top to bottom), (x (i, j) , y (i, j) ) represents the coordinates of the (i, j) th center of the motion compensation unit relative to the pixel at the upper left control point of the current decoding block .
- the current decoding block is a 6-parameter decoding block
- a motion vector of one or more sub-blocks of the current decoding block is obtained based on the target candidate motion vector group
- the lower boundary of the current decoding block is The lower boundary of the CTU where the current decoding block is located coincides
- the motion vector of the sub-block in the lower left corner of the current decoding block is a 6-parameter affine model constructed from the three control points and the current decoding block.
- the position coordinates (0, H) of the lower left corner are calculated
- the motion vector of the subblock in the lower right corner of the current decoding block is a 6-parameter affine model constructed according to the three control points and the right of the current decoding block.
- the position coordinates (W, H) of the lower corner are calculated. For example, by substituting the position coordinates (0, H) of the lower left corner of the current decoded block into the 6-parameter affine model, the motion vector of the lower left corner sub-block of the current decoded block (instead of the The center point coordinate is substituted into the affine model for calculation), and the position coordinates (W, H) of the lower right corner of the current decoded block are substituted into the 6-parameter affine model to obtain the motion vector of the lower right corner of the current decoded block. (Instead of substituting the coordinates of the center point of the subblock in the lower right corner into the affine model for calculation).
- the motion vector of the lower left control point and the lower right control point of the current decoding block are used (for example, the subsequent other blocks are constructed based on the motion vector of the lower left control point and the lower right control point of the current block).
- the candidate motion vector prediction value MVP list of the other blocks uses an accurate value instead of an estimated value.
- W is the width of the current decoded block
- H is the height of the current decoded block.
- the current decoding block is a 4-parameter decoding block
- the motion vectors of one or more sub-blocks of the current decoding block are obtained based on the target candidate motion vector group, if the lower boundary of the current decoding block and The lower boundary of the CTU where the current decoding block is located coincides, and the motion vector of the lower left sub-block of the current decoding block is a 4-parameter affine model constructed according to the two control points and the current decoding block's The position coordinates (0, H) of the lower left corner are calculated, and the motion vector of the sub-block in the lower right corner of the current decoding block is a 4-parameter affine model constructed according to the two control points and the right of the current decoding block.
- the position coordinates (W, H) of the lower corner are calculated. For example, by substituting the position coordinates (0, H) of the lower left corner of the current decoded block into the 4-parameter affine model, the motion vector of the lower left sub-block of the current decoded block (instead of the The center point coordinates are substituted into the affine model for calculation), and the position coordinates (W, H) of the lower right corner of the current decoding block are substituted into the four-parameter affine model to obtain the motion vector of the subblock in the lower right corner of the current decoding block. (Instead of substituting the coordinates of the center point of the subblock in the lower right corner into the affine model for calculation).
- the motion vector of the lower left control point and the lower right control point of the current decoding block are used (for example, the subsequent other blocks are constructed based on the motion vector of the lower left control point and the lower right control point of the current block).
- the candidate motion vector prediction value MVP list of the other blocks uses an accurate value instead of an estimated value.
- W is the width of the current decoded block
- H is the height of the current decoded block.
- Step S1216 The video decoder performs motion compensation according to the motion vector value of each sub-block in the current decoding block to obtain the pixel prediction value of each sub-block. For example, the motion vector and reference frame index value of each sub-block are used in The corresponding sub-block is found in the reference frame, and interpolation filtering is performed to obtain the pixel prediction value of each sub-block.
- Step S1221 The video decoder constructs a candidate motion information list.
- the video decoder constructs a candidate motion information list (also referred to as a candidate motion vector list) through an inter prediction unit (also referred to as an inter prediction module), which may be constructed in one of two ways provided below. Or it can be constructed by a combination of two methods, and the candidate motion information list constructed is a triplet candidate motion information list; the above two methods are specifically as follows:
- Method 1 A motion vector prediction method based on a motion model is used to construct a candidate motion information list.
- all or part of neighboring blocks of the current decoding block are traversed in a predetermined order to determine the neighboring affine decoding blocks, and the number of the determined neighboring affine decoding blocks may be one or more.
- the neighboring blocks A, B, C, D, and E shown in FIG. 7A may be traversed in order to determine neighboring affine decoding blocks among the neighboring blocks A, B, C, D, and E.
- the inter prediction unit determines a set of candidate motion vectors according to each adjacent affine decoding block (each set of candidate motion vectors is a two-tuple or three-tuple).
- the following takes an adjacent affine decoding block as an example.
- the adjacent affine decoding block is called the first adjacent affine decoding block, as follows:
- a first affine model is determined according to a motion vector of a control point of a first adjacent affine decoding block, and then a motion vector of a control point of the current decoding block is predicted according to the first affine model, which is specifically described as follows:
- first adjacent affine decoding block is located above the CTU of the current decoding block and the first adjacent affine decoding block is a four-parameter affine decoding block, two lower-most two sides of the first adjacent affine decoding block are obtained.
- Position coordinates and motion vectors of the control points for example, the position coordinates (x 6 , y 6 ) and motion vectors (vx 6 , vy 6 ) of the lower left control point of the first adjacent affine decoding block can be obtained, and the lower right The position coordinates (x 7 , y 7 ) and motion vector values (vx 7 , vy 7 ) of the control point.
- a first affine model is formed according to the motion vectors of the two control points at the bottom of the first adjacent affine decoding block (the first affine model obtained at this time is a 4-parameter affine model).
- the motion vector of the control point of the current decoding block is predicted according to the first affine model.
- the position coordinates of the upper left control point, the position coordinates of the upper right control point, and the position of the lower left control point of the current decoding block can be predicted.
- the coordinates are brought into the first affine model respectively, thereby predicting the motion vector of the upper left control point, the motion vector of the upper right control point, and the motion of the lower left control point of the current decoded block to form a candidate motion vector triplet and adding candidate motions.
- the information list is shown in formulas (1), (2), and (3).
- the motion vector of the control point of the current decoding block is predicted according to the first affine model.
- the position coordinates of the upper left control point and the position coordinates of the upper right control point of the current decoding block may be brought into the first An affine model, so as to predict the motion vector of the upper left control point and the motion vector of the upper right control point of the current decoding block, form a candidate motion vector two-tuple, and add the candidate motion information list, as shown in formulas (1) and (2). Show.
- (X 2 , y 2 ) are the coordinates of the lower left control point of the current decoding block; in addition, (vx 0 , vy 0 ) is the predicted motion vector of the upper left control point of the current decoding block, and (vx 1 , vy 1 ) is The predicted motion vector of the upper right control point of the current decoded block, (vx 2 , vy 2 ) is the predicted motion vector of the lower left control point of the current decoded block.
- the first neighboring affine decoding block is located in a Coding Tree Unit (CTU) above the current decoding block and the first neighboring affine decoding block is a six-parameter affine decoding block, it is not based on the first phase
- the neighboring affine decoding block generates a candidate motion vector prediction value of a control point of the current block.
- CTU Coding Tree Unit
- the manner of predicting the motion vector of the control point of the current decoding block is not limited here.
- an optional determination method is also exemplified below:
- the position coordinates and motion vectors of the three control points of the first adjacent affine decoding block may be obtained, for example, the position coordinates (x 4 , y 4 ) and motion vector values (vx 4 , vy 4 ) of the upper left control point, an upper right position coordinates of the control points (x 5, y 5) and a motion vector value (vx 5, vy 5), the position coordinates of the lower left control point (x 6, y 6) and the motion vector (vx 6, vy 6).
- a 6-parameter affine is formed according to the position coordinates and motion vectors of the three control points of the first adjacent affine decoding block
- the affine model predicts the motion vector of the upper left control point, the motion vector of the upper right control point, and the motion vector of the lower left control point of the current decoding block, as shown in formulas (4), (5), and (6).
- Formulas (4) and (5) have been described previously.
- (vx 0 , vy 0 ) is the predicted motion vector of the upper-left control point of the current decoding block
- (vx 1 , vy 1 ) are the motion vectors of the predicted upper right control point of the current decoded block
- (vx 2 , vy 2 ) are the predicted motion vectors of the lower left control point of the current decoded block.
- Manner 2 A motion vector prediction method based on a combination of control points is used to construct a candidate motion information list.
- Option A The following two examples are exemplified as Option A and Option B:
- Solution A Combine the motion information of the two control points of the current decoding block to construct a 4-parameter affine transformation model.
- the combination of the two control points is ⁇ CP1, CP4 ⁇ , ⁇ CP2, CP3 ⁇ , ⁇ CP1, CP2 ⁇ , ⁇ CP2, CP4 ⁇ , ⁇ CP1, CP3 ⁇ , ⁇ CP3, CP4 ⁇ .
- Affine CP1, CP2
- control points can also be converted into a control point at the same position.
- the four-parameter affine transformation model obtained by combining ⁇ CP1, CP4 ⁇ , ⁇ CP2, CP3 ⁇ , ⁇ CP2, CP4 ⁇ , ⁇ CP1, CP3 ⁇ , ⁇ CP3, CP4 ⁇ into a control point ⁇ CP1, CP2 ⁇ Or ⁇ CP1, CP2, CP3 ⁇ .
- the conversion method is to substitute the motion vector of the control point and its coordinate information into formula (9-1) to obtain the model parameters, and then substitute the coordinate information of ⁇ CP1, CP2 ⁇ to obtain its motion vector, which is used as a set of candidate motion vector predictions. value.
- a 1 , a 2 , a 3 , and a 4 are parameters in the parameter model, and (x, y) represents position coordinates.
- a set of motion vector prediction values represented by the upper left control point and the upper right control point can also be converted according to the following formula, and added to the candidate motion information list:
- ⁇ CP1, CP3 ⁇ is converted to ⁇ CP1, CP2, CP3 ⁇ 's formula (9-3):
- ⁇ CP1, CP4 ⁇ is converted to ⁇ CP1, CP2, CP3 ⁇ 's formula (11):
- ⁇ CP3, CP4 ⁇ is converted to ⁇ CP1, CP2, CP3 ⁇ 's formula (13):
- Solution B Combine the motion information of the three control points of the current decoding block to build a 6-parameter affine transformation model.
- the combination of the three control points is ⁇ CP1, CP2, CP4 ⁇ , ⁇ CP1, CP2, CP3 ⁇ , ⁇ CP2, CP3, CP4 ⁇ , ⁇ CP1, CP3, CP4 ⁇ .
- Affine CP1, CP2, CP3
- control points can also be converted into a control point at the same position.
- the 6-parameter affine transformation model of ⁇ CP1, CP2, CP4 ⁇ , ⁇ CP2, CP3, CP4 ⁇ , ⁇ CP1, CP3, CP4 ⁇ is converted into a control point ⁇ CP1, CP2, CP3 ⁇ to represent it.
- the transformation method is to substitute the motion vector of the control point and its coordinate information into formula (14) to obtain the model parameters, and then substitute the coordinate information of ⁇ CP1, CP2, CP3 ⁇ to obtain its motion vector, which is used as a set of candidate motion vector predictions. value.
- a 1 , a 2 , a 3 , a 4 , a 5 , and a 6 are parameters in the parameter model, and (x, y) represents position coordinates.
- a set of motion vector prediction values represented by the upper left control point, the upper right control point, and the lower left control point can also be converted according to the following formula, and added to the candidate motion information list:
- the candidate motion information list may be constructed only by using the candidate motion vector prediction values predicted by the first method, or only the candidate motion vector prediction values obtained by the second prediction method may be used to construct the candidate motion information list.
- a candidate motion vector prediction value obtained by the first prediction and a candidate motion vector prediction value obtained by the second prediction are used to jointly construct a candidate motion information list.
- the candidate motion information list can be pruned and sorted according to pre-configured rules, and then truncated or filled to a specific number.
- the candidate motion information list When each group of candidate motion vector prediction values in the candidate motion information list includes motion vector prediction values of three control points, the candidate motion information list may be referred to as a triple list; when each group in the candidate motion information list When the candidate motion vector prediction value includes motion vector prediction values of two control points, the candidate motion information list may be referred to as a two-tuple list.
- Step S1222 The video decoder parses the bitstream to obtain an index.
- the video decoder may parse the bitstream through an entropy decoding unit, and the index is used to indicate a target candidate motion vector group of the current decoding block, where the target candidate motion vector represents a motion vector prediction value of a set of control points of the current decoding block.
- Step S1223 The video decoder determines a target motion vector group from the candidate motion information list according to the index.
- the target decoder motion vector group determined from the candidate motion information list according to the index is used as the optimal candidate motion vector prediction value (optionally, when the length of the candidate motion information list is 1, no You need to parse the bitstream to get the index, and you can directly determine the target motion vector group), specifically the optimal motion vector prediction value of 2 or 3 control points; for example, the video decoder parses the index number from the bitstream, and then The index number determines the optimal motion vector prediction value of 2 or 3 control points from the candidate motion information list.
- Each group of candidate motion vector prediction values in the candidate motion information list corresponds to its own index number.
- Step S1224 The video decoder uses the parameter affine transformation model to obtain the motion vector value of each sub-block in the current decoding block according to the motion vector value of the control point of the current decoding block determined above, where the size of the sub-block is based on the current The prediction direction of the image block is determined.
- the motion information of pixels at preset positions in the motion compensation unit can be used to represent the motion of all pixels in the motion compensation unit information.
- the preset position pixels can be the center point of the motion compensation unit (M / 2, N / 2), the upper left pixel (0, 0), and the upper right pixel ( M-1,0), or other locations.
- FIG. 7B illustrates a 4x4 motion compensation unit
- FIG. 7C illustrates an 8x8 motion compensation unit.
- the coordinates of the center point of the motion compensation unit relative to the top left pixel of the current decoding block are calculated using formula (5), where i is the i-th motion compensation unit in the horizontal direction (from left to right), and j is the j-th motion in the vertical direction. Compensation unit (from top to bottom), (x (i, j) , y (i, j) ) represents the coordinates of the (i, j) th center of the motion compensation unit relative to the pixel at the upper left control point of the current decoding block.
- the current decoding block is a 6-parameter decoding block
- a motion vector of one or more sub-blocks of the current decoding block is obtained based on the target candidate motion vector group
- the lower boundary of the current decoding block is The lower boundary of the CTU where the current decoding block is located coincides
- the motion vector of the sub-block in the lower left corner of the current decoding block is a 6-parameter affine model constructed from the three control points and the current decoding block.
- the position coordinates (0, H) of the lower left corner are calculated
- the motion vector of the subblock in the lower right corner of the current decoding block is a 6-parameter affine model constructed according to the three control points and the right of the current decoding block.
- the position coordinates (W, H) of the lower corner are calculated. For example, by substituting the position coordinates (0, H) of the lower left corner of the current decoded block into the 6-parameter affine model, the motion vector of the lower left corner sub-block of the current decoded block (instead of the The center point coordinate is substituted into the affine model for calculation), and the position coordinates (W, H) of the lower right corner of the current decoded block are substituted into the 6-parameter affine model to obtain the motion vector of the lower right corner of the current decoded block. (Instead of substituting the coordinates of the center point of the subblock in the lower right corner into the affine model for calculation).
- the motion vector of the lower left control point and the lower right control point of the current decoding block are used (for example, the subsequent other blocks are constructed based on the motion vector of the lower left control point and the lower right control point of the current block).
- the candidate motion information list of the other blocks uses accurate values instead of estimated values.
- W is the width of the current decoded block
- H is the height of the current decoded block.
- the current decoding block is a 4-parameter decoding block
- the motion vectors of one or more sub-blocks of the current decoding block are obtained based on the target candidate motion vector group, if the lower boundary of the current decoding block and The lower boundary of the CTU where the current decoding block is located coincides, and the motion vector of the lower left sub-block of the current decoding block is a 4-parameter affine model constructed according to the two control points and the current decoding block's The position coordinates (0, H) of the lower left corner are calculated, and the motion vector of the sub-block in the lower right corner of the current decoding block is a 4-parameter affine model constructed according to the two control points and the right of the current decoding block.
- the position coordinates (W, H) of the lower corner are calculated. For example, by substituting the position coordinates (0, H) of the lower left corner of the current decoded block into the 4-parameter affine model, the motion vector of the lower left sub-block of the current decoded block (instead of the The center point coordinates are substituted into the affine model for calculation), and the position coordinates (W, H) of the lower right corner of the current decoding block are substituted into the four-parameter affine model to obtain the motion vector of the subblock in the lower right corner of the current decoding block. (Instead of substituting the coordinates of the center point of the subblock in the lower right corner into the affine model for calculation).
- the motion vector of the lower left control point and the lower right control point of the current decoding block are used (for example, the subsequent other blocks are constructed based on the motion vector of the lower left control point and the lower right control point of the current block).
- the candidate motion information list of the other blocks uses accurate values instead of estimated values.
- W is the width of the current decoded block
- H is the height of the current decoded block.
- Step S1225 The video decoder performs motion compensation according to the motion vector value of each sub-block in the current decoding block to obtain the pixel prediction value of each sub-block. Specifically, according to one or more sub-blocks of the current decoding block The motion vector, the reference frame index and the prediction direction indicated by the index, to obtain a pixel prediction value of the current decoded block.
- S1211 and / or S1212 may be optional steps, for example, it does not involve the construction of a list and does not involve parsing an index from a code stream.
- S1221 and / or S1222 may be optional steps, for example, it does not involve the construction of a list and does not involve parsing the index from the code stream.
- FIG. 8A is a schematic block diagram of an image prediction apparatus 800 according to an embodiment of the present application. It should be noted that the image prediction device 800 is applicable to both inter prediction of decoded video images and inter prediction of encoded video images. It should be understood that the image prediction device 800 here may correspond to the frame in FIG. 2A The inter prediction module 211 may correspond to the motion compensation module 322 in FIG. 2C. The image prediction device 800 may include:
- An obtaining unit 801 configured to obtain a motion vector of a control point of a current image block (current affine image block);
- the inter prediction processing unit 802 is configured to perform inter prediction on the current image block.
- the inter prediction process includes: according to a motion vector (a motion vector group) of a control point (affine control point) of the current image block. )
- An affine transformation model is used to obtain the motion vector of each sub-block in the current image block, where the size of the sub-block is determined based on the prediction direction of the current image block; Motion compensation is performed to obtain the pixel prediction value of each sub-block.
- the inter prediction processing unit 802 is configured to obtain an motion vector of each sub-block in the current image block by using an affine transformation model according to a motion vector of a control point of the current image block.
- the image block is unidirectional prediction, and the size of the subblock of the current image block is 4x4; or, if the current image block is bidirectional prediction, the size of the subblock of the current image block is 8x8; according to the motion of each subblock in the current image block
- the vector is motion compensated to obtain the pixel prediction value of each sub-block.
- the size of the subblock in the current image block is UxV; or,
- the size of the subblock in the current image block is MxN
- U M represents the width of the sub-block
- V represents the height of the sub-block
- U, V, M, N are all 2n
- n is a positive integer
- M 4 and N is 4.
- U 8 and V is 8.
- the obtaining unit 801 is specifically configured to:
- the motion vector of the control point of the current image block is determined according to the target candidate motion vector prediction value group and the motion vector difference MVD parsed from the code stream.
- the prediction direction indication information is used to indicate unidirectional prediction or bidirectional prediction, and the prediction direction indication information is obtained by parsing or deriving from the code stream. of.
- the obtaining unit 801 is specifically configured to:
- target candidate motion information is determined from a candidate motion information list, where the target candidate motion information includes at least one target candidate motion vector group, and the target candidate motion vector group is used as a control point of the current image block. Motion vector.
- the prediction direction of the current image block is bidirectional prediction
- the target candidate motion information corresponding to the index in the candidate motion information list includes a corresponding A first target candidate motion vector group in the first reference frame list, and a second target candidate motion vector group corresponding to the second reference frame list;
- the prediction direction of the current image block is one-way prediction, wherein the target candidate motion information corresponding to the index in the candidate motion information list includes a first target candidate motion vector group corresponding to a first reference frame list, or The target candidate motion information corresponding to the index in the candidate motion information list includes a second target candidate motion vector group corresponding to a second reference frame list.
- the inter prediction processing unit is specifically configured to obtain an affine transformation model according to a motion vector of a control point of the current image block; Position coordinate information of each sub-block in the image block and the affine transformation model to obtain a motion vector of each sub-block in the current image block.
- the size of the sub-blocks of the current image block is determined in consideration of the inter-frame direction of the current image block.
- the size of the image is 4 * 4; if the prediction mode of the current coded image block is bidirectional prediction, and the size of the sub-block is 8 * 8; in this way, a balance between the complexity of the motion compensation and the prediction efficiency is achieved, which reduces the current The complexity of motion compensation in the technology, while taking into account the prediction efficiency, thereby improving the performance of encoding and decoding.
- each module in the image prediction apparatus is a functional body that implements various execution steps included in the image prediction method of the present application, that is, it has the steps to implement the image prediction method of the present application.
- the main body of the expansion and deformation of these steps please refer to the description of the image prediction method in this article for details. For the sake of brevity, this article will not repeat them.
- FIG. 8B is a schematic block diagram of an implementation manner of an encoding device or a decoding device (hereinafter referred to as a decoding device 1000) according to an embodiment of the present application.
- the decoding device 1000 may include a processor 1010, a memory 1030, and a bus system 1050.
- the processor and the memory are connected through a bus system, the memory is used to store instructions, and the processor is used to execute the instructions stored in the memory.
- the memory of the encoding device stores program code, and the processor may call the program code stored in the memory to perform various video encoding or decoding methods described in this application, especially the image prediction method in this application. To avoid repetition, it will not be described in detail here.
- the processor 1010 may be a Central Processing Unit (“CPU”), and the processor 1010 may also be another general-purpose processor, a digital signal processor (DSP), or a dedicated integration. Circuits (ASICs), off-the-shelf programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.
- a general-purpose processor may be a microprocessor or the processor may be any conventional processor or the like.
- the memory 1030 may include a read only memory (ROM) device or a random access memory (RAM) device. Any other suitable type of storage device may also be used as the memory 1030.
- the memory 1030 may include code and data 1031 accessed by the processor 1010 using the bus 1050.
- the memory 1030 may further include an operating system 1033 and an application program 1035.
- the application program 1035 includes at least one program that allows the processor 1010 to perform the video encoding or decoding method (especially the encoding method or the decoding method described in this application).
- the application program 1035 may include applications 1 to N, which further includes a video encoding or decoding application (referred to as a video decoding application) that executes the video encoding or decoding method described in this application.
- the bus system 1050 may include a power bus, a control bus, and a status signal bus in addition to a data bus. However, for the sake of clarity, various buses are marked as the bus system 1050 in the figure.
- the decoding device 1000 may further include one or more output devices, such as a display 1070.
- the display 1070 may be a tactile display that incorporates the display with a tactile unit operatively sensing a touch input.
- the display 1070 may be connected to the processor 1010 via a bus 1050.
- FIG. 9 is an explanatory diagram of an example of a video encoding system 1100 including the encoder 20 of FIG. 2A and / or the decoder 30 of FIG. 2C according to an exemplary embodiment.
- the system 1100 may implement a combination of various techniques of the present application.
- the video encoding system 1100 may include an imaging device 1101, a video encoder 20, a video decoder 30 (and / or a video encoder implemented by the logic circuit 1107 of the processing unit 1106), an antenna 1102, One or more processors 1103, one or more memories 1104, and / or a display device 1105.
- the imaging device 1101, the antenna 1102, the processing unit 1106, the logic circuit 1107, the video encoder 20, the video decoder 30, the processor 1103, the memory 1104, and / or the display device 1105 can communicate with each other.
- video encoding system 1100 is shown with video encoder 20 and video decoder 30, in different examples, video encoding system 1100 may include only video encoder 20 or only video decoder 30.
- the video encoding system 1100 may include an antenna 1102.
- the antenna 1102 may be used to transmit or receive an encoded bit stream of video data.
- the video encoding system 1100 may include a display device 1105.
- the display device 1105 may be used to present video data.
- the logic circuit 1107 may be implemented by the processing unit 1106.
- the processing unit 1106 may include application-specific integrated circuit (ASIC) logic, a graphics processor, a general-purpose processor, and the like.
- the video encoding system 1100 may also include an optional processor 1103, which may similarly include application-specific integrated circuit (ASIC) logic, a graphics processor, a general-purpose processor, and the like.
- ASIC application-specific integrated circuit
- the logic circuit 1107 may be implemented by hardware, such as dedicated hardware for video encoding, and the processor 1103 may be implemented by general software, operating system, and the like.
- the memory 1104 may be any type of memory, such as volatile memory (for example, Static Random Access Memory (SRAM), Dynamic Random Access Memory (DRAM), etc.) or non-volatile memory Memory (for example, flash memory, etc.).
- the memory 1104 may be implemented by a cache memory.
- the logic circuit 1107 may access the memory 1104 (eg, for implementing an image buffer).
- the logic circuit 1107 and / or the processing unit 1106 may include a memory (eg, a buffer, etc.) for implementing an image buffer or the like.
- video encoder 20 implemented by a logic circuit may include an image buffer (eg, implemented by processing unit 1106 or memory 1104) and a graphics processing unit (eg, implemented by processing unit 1106).
- the graphics processing unit may be communicatively coupled to the image buffer.
- the graphics processing unit may include a video encoder 20 implemented by a logic circuit 1107 to implement the various modules discussed with reference to FIG. 2A and / or any other encoder system or subsystem described herein.
- Logic circuits can be used to perform various operations discussed herein.
- Video decoder 30 may be implemented in a similar manner through logic circuit 1107 to implement the various modules discussed with reference to decoder 200 of FIG. 2B and / or any other decoder system or subsystem described herein.
- video decoder 30 implemented by a logic circuit may include an image buffer (implemented by processing unit 1106 or memory 1104) and a graphics processing unit (eg, implemented by processing unit 1106).
- the graphics processing unit may be communicatively coupled to the image buffer.
- the graphics processing unit may include a video decoder 30 implemented by a logic circuit 1107 to implement the various modules discussed with reference to FIG. 2B and / or any other decoder system or subsystem described herein.
- the antenna 1102 of the video encoding system 1100 may be used to receive an encoded bit stream of video data.
- the encoded bitstream may contain data, indicators, index values, mode selection data, etc. related to encoded video frames discussed herein, such as data related to coded segmentation (e.g., transform coefficients or quantized transform coefficients) , (As discussed) optional indicators, and / or data defining code partitions).
- the video encoding system 1100 may also include a video decoder 30 coupled to the antenna 1102 and used to decode the encoded bitstream.
- the display device 1105 is used to present a video frame.
- the description order of the steps does not represent the execution order of the steps, and it is feasible to perform according to the above description order, and it is also possible to perform without the above description order.
- the above step S1211 may be performed after step S1212, or may be performed before step S1212; the above step S1221 may be performed after step S1222, or may be performed before step S1222; the remaining steps are not illustrated here one by one.
- Computer-readable media may include computer-readable storage media, which corresponds to tangible media, such as data storage media, or communication media including any medium that facilitates transfer of a computer program from one place to another (e.g., according to a communication protocol) .
- computer-readable media may generally correspond to (1) tangible computer-readable storage media that is non-transitory, or (2) a communication medium such as a signal or carrier wave.
- a data storage medium may be any available medium that can be accessed by one or more computers or one or more processors to retrieve instructions, code, and / or data structures used to implement the techniques described in this application.
- the computer program product may include a computer-readable medium.
- such computer-readable storage media may include RAM, ROM, EEPROM, CD-ROM or other optical disk storage devices, magnetic disk storage devices or other magnetic storage devices, flash memory, or may be used to store instructions or data structures Any form of desired program code and any other medium accessible by a computer.
- any connection is properly termed a computer-readable medium.
- a coaxial cable is used to transmit instructions from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave. Wire, fiber optic cable, twisted pair, DSL or wireless technologies such as infrared, radio and microwave are included in the definition of media.
- the computer-readable storage media and data storage media do not include connections, carrier waves, signals, or other temporary media, but are instead directed to non-transitory tangible storage media.
- magnetic disks and compact discs include compact discs (CDs), laser discs, optical discs, digital versatile discs (DVDs), and Blu-ray discs, where disks typically reproduce data magnetically, and optical discs use lasers to reproduce optically data. Combinations of the above should also be included within the scope of computer-readable media.
- DSPs digital signal processors
- ASICs application specific integrated circuits
- FPGAs field programmable logic arrays
- processor may refer to any of the aforementioned structures or any other structure suitable for implementing the techniques described herein.
- functions described by the various illustrative logical blocks, modules, and steps described herein may be provided within dedicated hardware and / or software modules configured for encoding and decoding, or Into the combined codec.
- the techniques can be fully implemented in one or more circuits or logic elements.
- the techniques of this application may be implemented in a wide variety of devices or devices, including a wireless handset, an integrated circuit (IC), or a group of ICs (eg, a chipset).
- IC integrated circuit
- Various components, modules, or units are described in this application to emphasize functional aspects of the apparatus for performing the disclosed techniques, but do not necessarily need to be implemented by different hardware units.
- the various units may be combined in a codec hardware unit in combination with suitable software and / or firmware, or through interoperable hardware units (including one or more processors as described above) provide.
Landscapes
- Engineering & Computer Science (AREA)
- Multimedia (AREA)
- Signal Processing (AREA)
- Compression Or Coding Systems Of Tv Signals (AREA)
Abstract
本申请实施例公开了图像预测方法、装置以及相应的编码器和解码器。考虑当前图像块的预测方向来确定当前图像块的子块的尺寸,比如,如果当前编码图像块的预测模式为单向预测,子块的尺寸为4*4;如果当前编码图像块的预测模式为双向预测,子块的尺寸为8*8;这样的话,在运动补偿的复杂度与预测效率两者之间得到平衡,即降低现有技术中运动补偿的复杂度的同时,兼顾预测效率,从而提高编解码性能。
Description
本申请涉及视频编解码技术领域,尤其涉及图像预测方法、装置以及相应的编码器和解码器。
视频编码(视频编码和解码)广泛用于数字视频应用,例如广播数字电视、互联网和移动网络上的视频传播、视频聊天和视频会议等实时会话应用、DVD和蓝光光盘、视频内容采集和编辑系统以及可携式摄像机的安全应用。
随着1990年H.261标准中基于块的混合型视频编码方式的发展,新的视频编码技术和工具得到发展并为新的视频编码标准形成基础。其它视频编码标准包括MPEG-1视频、MPEG-2视频、ITU-T H.262/MPEG-2、ITU-T H.263、ITU-T H.264/MPEG-4第10部分高级视频编码(Advanced Video Coding,AVC)、ITU-T H.265/高效视频编码(HighEfficiency Video Coding,HEVC)以及此类标准的扩展,例如可扩展性和/或3D(three-dimensional)扩展。随着视频创建和使用变得越来越广泛,视频流量成为通信网络和数据存储的最大负担。因此大多数视频编码标准的目标之一是相较之前的标准,在不牺牲图片质量的前提下减少比特率。即使最新的高效视频编码(HighEfficiency video coding,HEVC)可以在不牺牲图片质量的前提下比AVC大约多压缩视频一倍,仍然亟需新技术相对HEVC进一步压缩视频。
发明内容
本申请实施例提供一种图像预测方法、装置及相应的编码器和编码器,一定程度上降低复杂性,从而提高编解码性能。
第一方面,本申请实施例提供一种图像预测方法,应当理解的是,本申请实施例的方法的执行主体可以是视频解码器或具有视频解码功能的电子设备,或者本申请实施例的方法的执行主体可以是视频编码器或具有视频编码功能的电子设备,该方法可以包括:
获取当前图像块(例如当前仿射图像块)的控制点的运动矢量;
根据所述当前图像块的控制点(例如仿射控制点)的运动矢量(例如运动矢量组)采用仿射变换模型获得当前图像块中每个子块的运动矢量,其中所述子块的尺寸是基于当前图像块的预测方向而确定的;
根据该当前图像块中每个子块的运动矢量值进行运动补偿,以得到每个子块的像素预测值。
应当理解的是,每个子块的像素预测值得到后,可以进一步修正像素预测值,也可以不做修正。当当前图像块的多个子块的像素预测值得到后,就得到了当前图像块的像素预测值。
在第一方面的某些实现方式中,如果所述当前图像块的预测方向为双向预测,所述当前图像块中子块的尺寸为UxV;或者,
如果所述当前图像块的预测方向为单向预测,所述当前图像块中子块的尺寸为MxN,
其中U,M表示所述子块的宽,V,N表示所述子块的高,以及U,V,M,N均为2n,n为正整数。
在第一方面的某些实现方式中,U>=M,V>=N,且U和V不能同时等于M和N。
在第一方面的一种实现方式中,U=2M,V=2N。
在第一方面的一种具体实现方式中,M为4,N为4。
在第一方面的一种具体实现方式中,U为8,V为8。
在第一方面的某些实现方式中,如果所述当前图像块的尺寸为宽度W>=2U,高度H>=2V,所述方法还包括:从码流中解析得到仿射相关的语法元素(例如affine_inter_flag,affine_merge_flag)。
在第一方面的某些实现方式中,所述获取当前图像块(当前仿射图像块)的控制点的运动矢量,包括:
接收从码流中解析得到的索引(例如候选索引)和运动矢量差值MVD;
根据所述索引,从候选运动矢量预测值MVP列表中确定目标候选MVP组;
根据目标候选MVP组和从码流中解析出的运动矢量差值MVD确定当前图像块的控制点的运动矢量。应当理解的是,本申请不限于此,例如,也可以不依据索引,而是解码端通过其他方法得到目标候选MVP组。
在第一方面的某些实现方式中,所述当前图像块的预测方向是通过如下方式得出的:
预测方向指示信息用于指示单向预测或双向预测,其中所述预测方向指示信息是从码流中解析或推导得到的。
在第一方面的某些实现方式中,所述获取当前图像块(例如当前仿射图像块)的控制点的运动矢量,包括:
接收从码流中解析得到的索引(例如候选索引);
根据所述索引,从候选运动信息列表中确定目标候选运动信息,其中所述目标候选运动信息包括至少一个目标候选运动矢量组(例如对应于第一参考帧列表(例如L0)的目标候选运动矢量组和/或对应于第二参考帧列表(例如L1)的目标候选运动矢量),所述目标候选运动矢量组作为当前图像块的控制点的运动矢量。
在第一方面的某些实现方式中,所述当前图像块的预测方向是通过如下方式得出的:
所述当前图像块的预测方向是双向预测,其中所述候选运动信息列表中与索引(例如候选索引)对应的目标候选运动信息包括对应于第一参考帧列表的第一目标候选运动矢量组(例如第一方向的目标候选运动矢量组),和,对应于第二参考帧列表的第二目标候选运动矢量组(例如第二方向的目标候选运动矢量组);
所述当前图像块的预测方向是单向预测,其中所述候选运动矢量预测值MVP列表中与索引(例如候选索引)对应的目标候选运动信息包括对应于第一参考帧列表的第一目标候选运动矢量组(例如第一方向的目标候选运动矢量组),
或者,所述候选运动矢量预测值MVP列表中与索引(例如候选索引)对应的目标候选运动信息包括:对应于第二参考帧列表的第二目标候选运动矢量组(例如第二方向的目标候选运动矢量组)。
在第一方面的某些实现方式中,当当前图像块的尺寸满足W>=16和H>=16时,允许采用仿射模式。应当理解的是,允许采用仿射模式还可以考虑其他条件,这里的“当当前图像块的尺寸满足W>=16和H>=16”,不应理解为“仅当当当前图像块的尺寸满足W>=16和H>=16”。
在第一方面的某些实现方式中,其特征在于,所述根据所述获取的当前图像块的控制点的运动矢量值采用仿射变换模型获得当前图像块中每个子块的运动矢量值,包括:
根据所述当前图像块的控制点的运动矢量值,得到仿射变换模型;
根据当前图像块中每个子块的位置坐标信息以及所述仿射变换模型,获得当前图像块中每个子块的运动矢量。
可见,相对于现有技术中当前解码块被划分为MxN(即4x4)的子块,即每个MxN(即4x4)的子块采用对应的运动矢量进行运动补偿,本申请实施例中,当前解码块中的子块的尺寸是基于当前解码块的预测方向而确定的;例如,若当前解码块为单向预测,则当前解码块的子块的尺寸为4x4;若当前解码块为双向预测,则当前解码块的子块的尺寸为8x8。从图像的整体来看,本申请实施例的某些图像块的子块(或者运动补偿单元)的尺寸相对于现有技术中的子块(或者运动补偿单元)相对较大些,这样的话,平均每个像素进行运动补偿时所需要的读取的参考像素个数相对较少、插值的运算复杂度相对较低,从而本申请实施例在兼顾预测效率的同时,一定程度上降低了运动补偿的复杂度,从而提高了编解码性能。
第二方面,本申请实施例提供一种图像预测方法,应当理解的是,本申请实施例的方法的执行主体可以是视频解码器或具有视频解码功能的电子设备,或者本申请实施例的方法的执行主体可以是视频编码器或具有视频编码功能的电子设备,该方法可以包括:
获取当前图像块(当前仿射图像块)的控制点的运动矢量;
根据所述当前图像块的控制点(仿射控制点)的运动矢量(运动矢量组)采用仿射变换模型获得当前图像块中每个子块的运动矢量,其中,若当前图像块(当前仿射图像块)为单向预测,将当前图像块的子块的尺寸设置为4x4;或者,若当前图像块为双向预测,将当前图像块的子块的尺寸设置为8x8。
根据该当前图像块中每个子块的运动矢量值进行运动补偿,以得到每个子块的像 素预测值。
可见,本申请实施例中考虑当前图像块的预测方向来确定当前图像块的子块的尺寸,比如,如果当前编码图像块的预测模式为单向预测,子块的尺寸为4*4;如果当前编码图像块的预测模式为双向预测,子块的尺寸为8*8;这样的话,在运动补偿的复杂度与预测效率两者之间得到平衡,即降低现有技术中运动补偿的复杂度的同时,兼顾预测效率,从而提高编解码性能。
第三方面,本申请实施例提供一种图像预测装置,包括用于实施第一方面的任意一种方法的若干个功能单元。举例来说,图像预测装置可以包括:
获取单元,用于获取当前图像块的控制点的运动矢量;
帧间预测处理单元,用于根据所述当前图像块的控制点的运动矢量采用仿射变换模型获得当前图像块中每个子块的运动矢量,其中所述子块的尺寸是基于当前图像块的预测方向而确定的;根据该当前图像块中每个子块的运动矢量进行运动补偿,以得到每个子块的像素预测值。
第四方面,本申请实施例提供一种图像预测装置,包括用于实施第一方面或第二方面的任意一种方法的若干个功能单元。举例来说,图像预测装置可以包括:
获取单元,用于获取当前图像块的控制点的运动矢量;
帧间预测处理单元,用于根据所述当前图像块的控制点的运动矢量采用仿射变换模型获得当前图像块中每个子块的运动矢量,其中,若当前图像块为单向预测,当前图像块的子块的尺寸为4x4;或者,若当前图像块为双向预测,当前图像块的子块的尺寸为8x8;根据该当前图像块中每个子块的运动矢量进行运动补偿,以得到每个子块的像素预测值。
第五方面,本申请提供了一种解码方法,包括:视频解码器确定当前解码块(可以具体为当前仿射解码块)的帧间方向;解析码流,以得到索引和运动矢量差值MVD;根据所述索引,从候选运动矢量预测值MVP列表(例如,仿射变换候选运动矢量列表)中确定目标运动矢量组(亦可称为目标候选运动矢量组),所述目标运动矢量组表示当前解码块的一组控制点的运动矢量预测值;根据目标候选运动矢量组和从码流中解析出的运动矢量差值MVD确定当前解码块的控制点的运动矢量;根据所述确定出的当前解码块的控制点的运动矢量值采用仿射变换模型获得当前解码块中每个子块的运动矢量值,其中所述子块的尺寸是基于当前解码块的预测方向而确定的,或者所述 子块的尺寸是基于当前解码块的预测方向和所述当前解码块的尺寸而确定的;以及,根据该当前解码块中每个子块的运动矢量值进行运动补偿,以得到每个子块的像素预测值。
可见,相对于现有技术中当前解码块被划分为MxN(即4x4)的子块,即每个MxN(即4x4)的子块采用对应的运动矢量进行运动补偿,本申请实施例中,当前解码块中的子块的尺寸是基于当前解码块的预测方向而确定的,或者所述子块的尺寸是基于当前解码块的预测方向和所述当前解码块的尺寸而确定的;例如,若当前解码块为AMVP模式的单向预测,则将当前解码块的子块的尺寸设置为4x4;若当前解码块为AMVP模式的双向预测,则将当前解码块的子块的尺寸设置为8x4或4x8。从图像的整体来看,本申请实施例的子块(或者运动补偿单元)的尺寸相对于现有技术中的子块(或者运动补偿单元)相对较大些,这样的话,平均每个像素进行运动补偿时所需要的读取的参考像素个数相对较少、插值的运算复杂度相对较低,从而本申请实施例在兼顾预测效率的同时,一定程度上降低了运动补偿的复杂度,从而提高了编解码性能。
在第五方面的某些实现方式下,如果所述当前解码块的帧间预测模式为AMVP模式的双向预测模式,则所述当前解码块中子块的尺寸为UxV;
如果所述当前解码块的帧间预测模式为AMVP模式的单向预测模式,则所述当前解码块中子块的尺寸为MxN。
在第五方面的某些实现方式下,如果所述当前解码块的帧间预测模式为AMVP模式的双向预测模式且所述当前解码块的尺寸为宽度W>=2U,高度H>=2V,则所述当前解码块中子块的尺寸为UxV;
如果所述当前解码块的帧间预测模式为AMVP模式的单向预测模式,则所述当前解码块中子块的尺寸为MxN。
在第五方面的某些实现方式下,若当前仿射解码块进行单向预测的子块划分为MxN,则前仿射解码块进行双向预测的子块划分为UxV;其中,U>=M,V>=N,且U和V不能同时等于M和N。
在第五方面的某些实现方式下,U=2M,N=V,或U=M,V=2N,或U=2M,V=2N。
在第五方面的某些实现方式下,M为4、8、16等整数,N为4、8、16等整数。
在第五方面的某些实现方式下,如果所述当前解码块的尺寸为宽度W>=2U,高度H>=2V,所述解析码流还包括:从码流中解析得到仿射相关的语法元素(例如affine_inter_flag,affine_merge_flag)。
在第五方面的某些实现方式下,如果所述当前解码块的帧间预测模式为AMVP模式的双向预测模式且所述当前解码块的尺寸不满足“宽度W>=2U,高度H>=2V”,则确定所述当前解码块的帧间预测模式为AMVP模式的单向预测模式,且当前解码块的子块的尺寸为MxN。
在第五方面的某些实现方式下,若当前仿射解码块为单向预测,则将当前仿射解码块的子块的尺寸设置为4x4;若当前仿射解码块为双向预测,则将当前仿射解码块的子块的尺寸设置为8x4。
在第五方面的某些实现方式下,当当前解码块的尺寸满足W>=16和H>=8时,允许采用仿射模式。
在第五方面的某些实现方式下,当当前解码块的尺寸满足W>=8和H>=8时,允许采用仿射模式;
当当前解码块的宽度W<16时,若该仿射解码块的预测方向为双向,则将所述当前解码块的帧间预测模式修改为单向预测模式。
在第五方面的某些实现方式下,所述将所述当前解码块的帧间预测模式修改为单向预测模式,包括:丢弃后向预测的运动信息,将其转换为前向预测;或,丢弃前向预测的运动信息,将其转换为后向预测。
在第五方面的某些实现方式下,当当前仿射解码块的宽度W<16时,不需要解析双向预测的标志位。
在第五方面的某些实现方式下,若该当前仿射解码块为单向预测,则将当前仿射解码块的子块的尺寸设置为4x4;若该当前仿射解码块为双向预测,则将当前仿射解码块的子块的尺寸设置为4x8。
在第五方面的某些实现方式下,只有当解码块的尺寸满足W>=8和H>=16时,允许采用仿射模式。
在第五方面的某些实现方式下,只有当解码块的尺寸满足W>=8和H>=8时,允许采用仿射模式;当解码块的高度H<16时,若该仿射解码块的预测方向为双向,则将所述当前解码块的帧间预测模式修改为单向预测模式。
在第五方面的某些实现方式下,若该当前仿射解码块为单向预测,则将该当前仿射解码块的子块的尺寸设置为4x4;若该当前仿射解码块为双向预测,则根据该仿射解码块的尺寸进行自适应划分;划分方式可以为以下三种方式的其中任意一种:
1)如果该仿射解码块宽度W大于等于H,则将该当前仿射解码块的子块的尺寸设 置为8x4,如果该仿射解码块宽度W小于H,则将该当前仿射解码块的子块的尺寸设置为4x8;或者
2)如果该仿射解码块宽度W大于H,则将该当前仿射解码块的子块的尺寸设置为8x4,如果该仿射解码块宽度W小于等于H,则将该当前仿射解码块的子块的尺寸设置为4x8;或者
3)如果该仿射解码块宽度W大于H,则将该当前仿射解码块的子块的尺寸设置为8x4;如果该当前仿射解码块的子块的宽度W小于H,则将该当前仿射解码块的子块的尺寸设置为4x8;如果该仿射解码块宽度W等于H,则将该当前仿射解码块的子块的尺寸设置为8x8。
在第五方面的某些实现方式下,当解码块的尺寸满足W>=8和H>=8,并且W不等于8,H不等于8时,允许采用仿射模式。
在第五方面的某些实现方式下,当解码块的尺寸满足W>=8和H>=8时,允许采用仿射模式;当解码块的宽度W等于8,高度H等于8时,若该仿射解码块的预测方向为双向,则将所述当前解码块的帧间预测模式修改为单向预测模式。
在第五方面的某些实现方式下,还可以通过对码流进行限制,使得当仿射解码块的宽度W等于8,高度H等于8时,不需要解析双向预测的标志位。
在第五方面的某些实现方式下,候选运动矢量预测值MVP列表(例如,仿射变换候选运动矢量列表)中可能只有一个候选运动矢量组,也可能有多个候选运动矢量组,其中,每个候选运动矢量组可以为一个运动矢量二元组或者运动矢量三元组。可选的,当候选运动矢量预测值MVP列表(例如,仿射变换候选运动矢量列表)的长度为1时,不需要解析码流得到索引,直接可以确定目标运动矢量组。
在第五方面的某些实现方式下,在先进运动矢量预测AMVP模式下,所述确定当前解码块(例如当前仿射解码块)的帧间预测模式,包括:解析码流,以得到帧间预测相关的一个或多个语法元素,其中所述一个或多个语法元素用于指示当前解码块采用AMVP模式,以及用于指示当前解码块采用单向预测或者双向预测。
在第五方面的某些实现方式下,所述根据所述确定出的当前解码块的控制点的运动矢量值采用仿射变换模型获得当前解码块中每个子块的运动矢量值,包括:基于仿射变换模型得到所述当前解码块的一个或多个子块的运动矢量(例如,将该一个或者多个子块的中心点的坐标代入该仿射变换模型,从而得到一个或多个子块的运动矢量),其中,所述仿射变换模型是基于所述当前解码块的一组控制点的位置坐标和所述 当前解码块的一组控制点的运动矢量确定的,换言之,所述仿射变换模型的模型参数是基于所述当前解码块的一组控制点的位置坐标和所述当前解码块的一组控制点的运动矢量确定的。
第六方面,本申请提供了另一种解码方法,包括:视频解码器确定当前解码块的预测方向;解析码流,以得到索引;视频解码器根据所述索引,从候选运动信息列表中确定目标运动矢量组;视频解码器根据确定出的当前解码块的控制点的运动矢量值采用参数仿射变换模型获得当前解码块中每个子块的运动矢量值,其中所述子块的尺寸是基于当前解码块的帧间方向而确定的,或者所述子块的尺寸是基于当前解码块的帧间方向和所述当前解码块的尺寸而确定的;根据该当前解码块中每个子块的运动矢量值进行运动补偿,以得到每个子块的像素预测值。
在第六方面的某些实现方式下,在融合merge模式下,所述确定当前解码块(例如当前仿射解码块)的预测方向,包括:基于所述索引(例如affine_merge_idx)确定当前解码块采用单向预测或者双向预测,其中当前解码块的预测方向与所述索引(例如affine_merge_idx)指示的候选运动信息的预测方向相同;
可见,相对于现有技术中当前解码块被划分为MxN(即4x4)的子块,即每个MxN(即4x4)的子块采用对应的运动矢量进行运动补偿,本申请实施例中,当前解码块中的子块的尺寸是基于当前解码块的预测方向而确定的,或者所述子块的尺寸是基于当前解码块的预测方向和所述当前解码块的尺寸而确定的;例如,若当前解码块为merge模式的单向预测,则将当前解码块的子块的尺寸设置为4x4;若当前解码块为merge模式的双向预测,则将当前解码块的子块的尺寸设置为8x4或4x8。从图像的整体来看,本申请实施例的某些图像块的子块(或者运动补偿单元)的尺寸相对于现有技术中的子块(或者运动补偿单元)的尺寸相对较大些,这样的话,平均每个像素进行运动补偿时所需要的读取的参考像素个数相对较少、插值的运算复杂度相对较低,从而本申请实施例在兼顾预测效率的同时,一定程度上降低了运动补偿的复杂度,从而提高了编解码性能。
在第六方面的某些实现方式下,如果所述当前解码块的帧间预测模式为MERGE模式的双向预测模式,则所述当前解码块中子块的尺寸为UxV;
如果所述当前解码块的帧间预测模式为MERGE模式的单向预测模式,则所述当前解码块中子块的尺寸为MxN。
在第六方面的某些实现方式下,如果所述当前解码块的帧间预测模式为MERGE模式的双向预测模式且所述当前解码块的尺寸为宽度W>=2U,高度H>=2V,则所述当前解码块中子块的尺寸为UxV;
如果所述当前解码块的帧间预测模式为MERGE模式的单向预测模式,则所述当前解码块中子块的尺寸为MxN。
在第六方面的某些实现方式下,若仿射解码块进行单向预测的子块划分为MxN,则双向预测的子块划分为UxV;其中,U>=M,V>=N,且U和V不能同时等于M和N。
在第六方面的某些实现方式下,U=2M,N=V,或U=M,V=2N,或U=2M,V=2N,
在第六方面的某些实现方式下,M为4、8、16等整数,N为4、8、16等整数。
在第六方面的某些实现方式下,如果所述当前解码块的尺寸为宽度W>=2U,高度H>=2V,所述解析码流包括:从码流中解析得到仿射相关的语法元素(例如affine_inter_flag,affine_merge_flag)。
在第六方面的某些实现方式下,如果所述当前解码块的帧间预测模式为MERGE模式的双向预测模式且所述当前解码块的尺寸不满足“宽度W>=2U,高度H>=2V”,则确定所述当前解码块的帧间预测模式为MERGE模式的单向预测模式,且当前解码块的子块的尺寸为MxN。
在第六方面的某些实现方式下,若当前仿射解码块为单向预测,则将当前仿射解码块的基本运动补偿单元的尺寸设置为4x4,若当前仿射解码块为单向预测,则将当前仿射解码块的运动补偿单元的尺寸设置为8x4。
在第六方面的某些实现方式下,当当前解码块的尺寸满足W>=16和H>=8时,允许采用仿射模式。
在第六方面的某些实现方式下,当当前解码块的尺寸满足W>=8和H>=8时,允许采用仿射模式;
当当前解码块的宽度W<16时,若该仿射解码块的预测方向为双向,则将所述当前解码块的帧间预测模式修改为单向预测模式。
在第六方面的某些实现方式下,所述将所述当前解码块的帧间预测模式修改为单向预测模式,包括:丢弃后向预测的运动信息,将其转换为前向预测;或,丢弃前向预测的运动信息,将其转换为后向预测。
在第六方面的某些实现方式下,当当前仿射解码块的宽度W<16时,不需要解析双向预测的标志位。
在第六方面的某些实现方式下,若该当前仿射解码块为单向预测,则将当前仿射解码块的运动补偿单元的尺寸设置为4x4,若该当前仿射解码块为双向预测,则将其运动补偿单元的尺寸设置为4x8。
在第六方面的某些实现方式下,当解码块的尺寸满足W>=8和H>=16时,允许采用仿射模式。
在第六方面的某些实现方式下,当解码块的尺寸满足W>=8和H>=8时,允许采用仿射模式;当解码块的高度H<16时,若该仿射解码块的预测方向为双向,则将所述当前解码块的帧间预测模式修改为单向预测模式。
在第六方面的某些实现方式下,若该仿射解码块为单向预测,则将其运动补偿单元的尺寸设置为4x4,若该仿射解码块为双向预测,则根据该仿射解码块的尺寸进行自适应划分;划分方式可以为以下三种方式的其中任意一种:
1)如果该仿射解码块宽度W大于等于H,则将其运动补偿单元的尺寸设置为8x4,如果该仿射解码块宽度W小于H,则将其运动补偿单元的尺寸设置为4x8;或者
2)如果该仿射解码块宽度W大于H,则将其运动补偿单元的尺寸设置为8x4,如果该仿射解码块宽度W小于等于H,则将其运动补偿单元的尺寸设置为4x8;或者
3)如果该仿射解码块宽度W大于H,则将其运动补偿单元的尺寸设置为8x4,如果该仿射解码块宽度W小于H,则将其运动补偿单元的尺寸设置为4x8,如果该仿射解码块宽度W等于H,则将其运动补偿单元的尺寸设置为8x8。
在第六方面的某些实现方式下,当解码块的尺寸满足W>=8和H>=8,并且W不等于8,H不等于8时,允许采用仿射模式。
在第六方面的某些实现方式下,当解码块的尺寸满足W>=8和H>=8时,允许采用仿射模式;而当解码块的宽度W等于8,高度H等于8时,若该仿射解码块的预测方向为双向,则将所述当前解码块的帧间预测模式修改为单向预测模式。
在第六方面的某些实现方式下,所述方法还可以通过对码流进行限制,使得当仿射解码块的宽度W等于8,高度H等于8时,不需要解析双向预测的标志位。
第七方面,本申请实施例提供一种视频解码器,包括用于实施前述方面的任意一种方法的若干个功能单元。举例来说,视频解码器可以包括:
熵解码单元,用于解析码流,以得到索引和运动矢量差值MVD;
帧间预测单元,用于确定当前解码块的预测方向;根据所述索引,从候选运动矢量预测值MVP列表中确定目标候选MVP组;根据目标候选MVP组和从码流中解析出的 运动矢量差值MVD确定当前解码块的控制点的运动矢量;根据所述确定出的当前解码块的控制点的运动矢量值采用仿射变换模型获得当前解码块中每个子块的运动矢量值,其中所述子块的尺寸是基于当前解码块的预测方向而确定的,或者所述子块的尺寸是基于当前解码块的预测方向和所述当前解码块的尺寸而确定的;根据该当前解码块中每个子块的运动矢量值进行运动补偿,以得到每个子块的像素预测值。
第八方面,本申请实施例提供另一种视频解码器,包括用于实施第前述方面的任意一种方法的若干个功能单元。举例来说,视频解码器可以包括:
熵解码单元,用于解析码流,以得到索引;
帧间预测单元,用于确定当前解码块的预测方向;根据所述索引,从候选运动信息列表中确定目标运动矢量组;根据确定出的当前解码块的控制点的运动矢量值采用参数仿射变换模型获得当前解码块中每个子块的运动矢量值,其中所述子块的尺寸是基于当前解码块的预测方向而确定的,或者所述子块的尺寸是基于当前解码块的预测方向和所述当前解码块的尺寸而确定的;根据该当前解码块中每个子块的运动矢量值进行运动补偿,以得到每个子块的像素预测值。
根据本申请第一方面的方法可由根据本申请第三方面的装置执行。根据本申请第一方面的方法的其它特征和实现方式直接取决于根据本申请第三方面的装置的功能性及其不同实现方式。
根据本申请第二方面的方法可由根据本申请第四方面的装置执行。根据本申请第二方面的方法的其它特征和实现方式直接取决于根据本申请第四方面的装置的功能性及其不同实现方式。
第九方面,本申请涉及解码视频流的装置,包含处理器和存储器。所述存储器存储指令,所述指令使得所述处理器执行根据第一方面或第二方面的方法。
第十方面,本申请涉及编码视频流的装置,包含处理器和存储器。所述存储器存储指令,所述指令使得所述处理器执行根据第一方面或第二方面的方法。
第十一方面,提出计算机可读存储介质,其上储存有指令,所述指令执行时,使得一个或多个处理器编码视频数据。所述指令使得所述一个或多个处理器执行根据第一或第二方面或第一或第二方面任何可能实施例的方法。
第十二方面,本申请涉及包括程序代码的计算机程序,所述程序代码在计算机上运行时执行根据第一或第二方面或第一或第二方面任何可能实施例的方法。
第十三方面,本申请实施例提供一种用于解码视频数据的设备,所述设备包括:
存储器,用于存储码流形式的视频数据;
视频解码器,用于解析码流,以得到索引和运动矢量差值MVD;根据所述索引,从候选运动矢量预测值MVP列表中确定目标运动矢量组,所述目标运动矢量组表示当前解码块的一组控制点的运动矢量预测值;根据目标运动矢量组和从码流中解析出的运动矢量差值MVD确定当前解码块的控制点的运动矢量;以及,还用于确定当前解码块的预测方向;根据所述确定出的当前解码块的控制点的运动矢量值采用仿射变换模型获得当前解码块中每个子块(例如一个或多个子块)的运动矢量值,其中所述子块的尺寸是基于当前解码块的预测方向而确定的,或者所述子块的尺寸是基于当前解码块的预测方向和所述当前解码块的尺寸而确定的;根据该当前解码块中每个子块(例如一个或多个子块)的运动矢量值进行运动补偿,以得到每个子块的像素预测值。换言之,即基于所述当前解码块的一个或多个子块的运动矢量,预测得到所述当前解码块的像素预测值。
第十四方面,本申请实施例提供一种用于解码视频数据的设备,所述设备包括:
存储器,用于存储码流形式的视频数据;
视频解码器,用于解析码流,以得到索引;根据所述索引,从候选运动信息列表中确定目标运动矢量组,所述目标运动矢量组表示当前解码块的一组控制点的运动矢量;以及,还用于确定当前解码块的预测方向;根据确定出的当前解码块的控制点的运动矢量值采用参数仿射变换模型获得当前解码块中每个子块(例如一个或多个子块)的运动矢量值,其中所述子块的尺寸是基于当前解码块的预测方向而确定的,或者所述子块的尺寸是基于当前解码块的预测方向和所述当前解码块的尺寸而确定的;根据该当前解码块中每个子块(例如一个或多个子块)的运动矢量值进行运动补偿,以得到每个子块的像素预测值;换言之,即基于所述当前解码块的一个或多个子块的运动矢量,预测得到所述当前解码块的像素预测值。
应当理解的是,本申请的第二至第十四方面与本申请的第一方面的技术方案相应,各方面及对应的可行实施方式所取得的有益效果参见第一方面,不再赘述。
在附图及以下说明中阐述一个或多个实施例的细节。
为了更清楚地说明本申请实施例或背景技术中的技术方案,下面将对本申请实施例或背景技术中所需要使用的附图进行说明。
图1为本申请实施例中所描述的实施方式中视频编码及解码系统的框图;
图2A为本申请实施例中所描述的实施方式中视频编码器的框图;
图2B为本申请实施例所描述的实施方式中帧间预测示意图;
图2C为本申请实施例中所描述的实施方式中视频解码器的框图;
图3为本申请实施例中所描述的实施方式中运动信息候选位置示意图;
图4为本申请实施例中所描述的实施方式中继承的控制点运动矢量预测示意图;
图5A为本申请实施例中所描述的实施方式中构造的控制点运动矢量预测示意图;
图5B为本申请实施例中所描述的实施方式中将控制点的运动信息进行组合,得到构造的控制点运动信息的流程示意图;
图6A为本申请实施例中所描述的实施方式中解码方法的流程图;
图6B为本申请实施例中所描述的实施方式中构造候选运动矢量列表示意图;
图6C为本申请实施例中所描述的实施方式中子块(亦称为运动补偿单元)的一种示意图;
图7A为本申请实施例中提供的一种图像预测方法的流程示意图;
图7B是本申请实施例提供的一种子块(亦称为运动补偿单元)的结构示意图;
图7C是本申请实施例提供的又一种子块(亦称为运动补偿单元)的结构示意图;
图7D为本申请实施例中提供的一种解码方法的流程示意图;
图8A是本身实施例提供的一种图像预测装置的结构示意图;
图8B是本身实施例提供的一种编码设备或解码设备的结构示意图;
图9是根据一示例性实施例的包含图2A的编码器20和/或图2C的解码器30的视频编码系统1100。
以下如果没有关于相同参考符号的具体注释,相同的参考符号是指相同或至少功能上等效的特征。
本申请实施例提供的视频图像预测方案可以应用于视频图像的编码或者解码中。 图1为本申请实施例中视频编码及解码系统10的一种示意性框图。如图1所示,系统10包含源装置11和目的装置12,源装置11产生编码视频数据并发送给目的装置12,目的装置12用于接收编码视频数据,并对编码视频数据进行解码并显示。源装置11及目的装置12可包括广泛范围的装置中的任一种,包含桌上型计算机、笔记型计算机、平板计算机、机顶盒、例如所谓的“智能”电话的电话手机、所谓的“智能”触控板、电视、摄影机、显示装置、数字媒体播放器、视频游戏控制台、视频流式传输装置等。
本申请实施例提供的图像块的图像预测方案可以应用于视频图像的编码或者解码中。图1为本申请实施例中视频编码及解码系统10的一种示意性框图。如图1所示,系统10包含源装置11和目的装置12,源装置11产生编码视频数据并发送给目的装置12,目的装置12用于接收编码视频数据,并对编码视频数据进行解码并显示。源装置11及目的装置12可包括广泛范围的装置中的任一种,包含桌上型计算机、笔记型计算机、平板计算机、机顶盒、例如所谓的“智能”电话的电话手机、所谓的“智能”触控板、电视、摄影机、显示装置、数字媒体播放器、视频游戏控制台、视频流式传输装置等。
目的装置12可经由链路16接收待解码的经编码视频数据。链路16可包括能够将经编码视频数据从源装置11传递到目的装置12的任何类型的媒体或装置。在一个可能的实现方式中,链路16可包括使源装置11能够实时将经编码视频数据直接传输到目的装置12的通信媒体。可根据通信标准(例如,无线通信协议)调制经编码视频数据且将其传输到目的装置12。通信媒体可包括任何无线或有线通信媒体,例如射频频谱或一个或多个物理传输线。通信媒体可形成基于包的网络(例如,局域网、广域网或因特网的全球网络)的部分。通信媒体可包含路由器、交换器、基站或可有用于促进从源装置11到目的装置12的通信的任何其它设备。
替代地,视频编码及解码系统10还包括存储装置,可将经编码数据从输出接口14输出到存储装置。类似地,可由输入接口15从存储装置存取经编码数据。存储装置可包含多种分散式或本地存取的数据存储媒体中的任一者,例如,硬盘驱动器、蓝光光盘、DVD、CD-ROM、快闪存储器、易失性或非易失性存储器或用于存储经编码视频数据的任何其它合适的数字存储媒体。在另一可行的实施方式中,存储装置可对应于文件服务器或可保持由源装置11产生的经编码视频的另一中间存储装置。目的装置12可经由流式传输或下载从存储装置存取所存储视频数据。文件服务器可为能够存储经编码视频数据且将此经编码视频数据传输到目的装置12的任何类型的服务器。可行的实施方式文件服务器包含网站服务器、文件传送协议服务器、网络附接存储装置或本地磁盘机。目的装置12可经由包含因特网连接的任何标准数据连接存取经编码视频数据。此数据连接可包含适合于存取存储于文件服务器上的经编码视频数据的无线信道(例如,Wi-Fi连接)、有线连接(例如,缆线调制解调器等)或两者的组合。经编码视频数据从存储装置的传输可为流式传输、下载传输或两者的组合。
本申请的技术不必限于无线应用或设定。技术可应用于视频解码以支持多种多媒体应用中的任一个,例如,空中电视广播、有线电视传输、卫星电视传输、流式传输视频传输(例如,经由因特网)、编码数字视频以用于存储于数据存储媒体上、解码存储于数据存储媒体上的数字视频或其它应用。在一些可能的实现方式中,系统10可经 配置以支持单向或双向视频传输以支持例如视频流式传输、视频播放、视频广播和/或视频电话的应用。
在图1的可能的实现方式中,源装置11可以包括视频源13、视频编码器20及输出接口14。在一些应用中,输出接口14可包括调制器/解调制器(调制解调器)和/或传输器。在源装置11中,视频源13可包括例如以下各种的源设备:视频捕获装置(例如,摄像机)、含有先前捕获的视频的存档、用以从视频内容提供者接收视频的视频馈入接口,和/或用于产生计算机图形数据作为源视频的计算机图形系统,或这些源的组合。作为一种可能的实现方式,如果视频源13为摄像机,那么源装置11及目的装置12可形成所谓的摄影机电话或视频电话。本申请中所描述的技术可示例性地适用于视频解码,且可适用于无线和/或有线应用。
可由视频编码器20来编码捕获、预捕获或计算产生的视频。经编码视频数据可经由源装置11的输出接口14直接传输到目的装置12。经编码视频数据也可(或替代地)存储到存储装置上以供稍后由目的装置12或其它装置存取以用于解码和/或播放。
目的装置12包含输入接口15、视频解码器30及显示装置17。在一些应用中,输入接口15可包含接收器和/或调制解调器。目的装置12的输入接口15经由链路16接收经编码视频数据。经由链路16传达或提供于存储装置上的经编码视频数据可包含由视频编码器20产生以供视频解码器30的视频解码器使用以解码视频数据的多种语法元素。这些语法元素可与在通信媒体上传输、存储于存储媒体上或存储于文件服务器上的经编码视频数据包含在一起。
显示装置17可与目的装置12集成或在目的装置12外部。在一些可能的实现方式中,目的装置12可包含集成显示装置且也经配置以与外部显示装置接口连接。在其它可能的实现方式中,目的装置12可为显示装置。一般来说,显示装置17向用户显示经解码视频数据,且可包括多种显示装置中的任一个,例如液晶显示器、等离子显示器、有机发光二极管显示器或另一类型的显示装置。
视频编码器20及视频解码器30可根据例如目前在开发中的下一代视频编解码压缩标准(H.266)操作且可遵照H.266测试模型(JEM)。替代地,视频编码器20及视频解码器30可根据例如ITU-TH.265标准,也称为高效率视频解码标准,或者,ITU-TH.264标准的其它专属或工业标准或这些标准的扩展而操作,ITU-TH.264标准替代地被称为MPEG-4第10部分,也称高级视频编码(advanced video coding,AVC)。然而,本申请的技术不限于任何特定解码标准。视频压缩标准的其它可能的实现方式包含MPEG-2和ITU-TH.263。
尽管未在图1中展示,但在一些方面中,视频编码器20及视频解码器30可各自与音频编码器及解码器集成,且可包含适当多路复用器-多路分用器(MUX-DEMUX)单元或其它硬件及软件以处置共同数据流或单独数据流中的音频及视频两者的编码。如果适用,那么在一些可行的实施方式中,MUX-DEMUX单元可遵照ITUH.223多路复用器协议或例如用户数据报协议(UDP)的其它协议。
视频编码器20及视频解码器30各自可实施为多种合适编码器电路中的任一者,例如,一个或多个微处理器、数字信号处理器(digital signal processing,DSP)、专 用集成电路(application specific integrated circuit,ASIC)、现场可编程门阵列(field-programmable gate array,FPGA)、离散逻辑、软件、硬件、固件或其任何组合。在技术部分地以软件实施时,装置可将软件的指令存储于合适的非暂时性计算机可读媒体中且使用一个或多个处理器以硬件执行指令,以执行本申请的技术。视频编码器20及视频解码器30中的每一者可包含于一个或多个编码器或解码器中,其中的任一者可在相应装置中集成为组合式编码器/解码器(CODEC)的部分。
JCT-VC开发了H.265(HEVC)标准。HEVC标准化基于称作HEVC测试模型(HM)的视频解码装置的演进模型。H.265的最新标准文档可从http://www.itu.int/rec/T-REC-H.265获得,最新版本的标准文档为H.265(12/16),该标准文档以全文引用的方式并入本文中。HM假设视频解码装置相对于ITU-TH.264/AVC的现有算法具有若干额外能力。
JVET致力于开发H.266标准。H.266标准化的过程基于称作H.266测试模型的视频解码装置的演进模型。H.266的算法描述可从http://phenix.int-evry.fr/jvet获得,其中最新的算法描述包含于JVET-F1001-v2中,该算法描述文档以全文引用的方式并入本文中。同时,可从https://jvet.hhi.fraunhofer.de/svn/svn_HMJEMSoftware/获得JEM测试模型的参考软件,同样以全文引用的方式并入本文中。
一般来说,HM的工作模型描述可将视频帧或图像划分成包含亮度及色度样本两者的树块或最大编码单元(largest coding unit,LCU)的序列,LCU也被称为CTU。树块具有与H.264标准的宏块类似的目的。条带包含按解码次序的数个连续树块。可将视频帧或图像分割成一个或多个条带。可根据四叉树将每一树块分裂成编码单元。例如,可将作为四叉树的根节点的树块分裂成四个子节点,且每一子节点可又为母节点且被分裂成另外四个子节点。作为四叉树的叶节点的最终不可分裂的子节点包括解码节点,例如,经解码图像块。与经解码码流相关联的语法数据可定义树块可分裂的最大次数,且也可定义解码节点的最小大小。
编码单元包含解码节点及预测单元(prediction unit,PU)以及与解码节点相关联的变换单元(transform unit,TU)。CU的大小对应于解码节点的大小且形状必须为正方形。CU的大小的范围可为8×8像素直到最大64×64像素或更大的树块的大小。每一CU可含有一个或多个PU及一个或多个TU。例如,与CU相关联的语法数据可描述将CU分割成一个或多个PU的情形。分割模式在CU是被跳过或经直接模式编码、帧内预测模式编码或帧间预测模式编码的情形之间可为不同的。PU可经分割成形状为非正方形。例如,与CU相关联的语法数据也可描述根据四叉树将CU分割成一个或多个TU的情形。TU的形状可为正方形或非正方形。
HEVC标准允许根据TU进行变换,TU对于不同CU来说可为不同的。TU通常基于针对经分割LCU定义的给定CU内的PU的大小而设定大小,但情况可能并非总是如此。TU的大小通常与PU相同或小于PU。在一些可行的实施方式中,可使用称作“残余四叉树”(residual qualtree,RQT)的四叉树结构将对应于CU的残余样本再分成较小单元。RQT的叶节点可被称作TU。可变换与TU相关联的像素差值以产生变换系数,变换系数可被量化。
一般来说,PU包含与预测过程有关的数据。例如,在PU经帧内模式编码时,PU可包含描述PU的帧内预测模式的数据。作为另一可行的实施方式,在PU经帧间模式编码时,PU可包含界定PU的运动矢量的数据。例如,界定PU的运动矢量的数据可描述运动矢量的水平分量、运动矢量的垂直分量、运动矢量的分辨率(例如,四分之一像素精确度或八分之一像素精确度)、运动矢量所指向的参考图像,和/或运动矢量的参考图像列表(例如,列表0、列表1或列表C)。
一般来说,TU使用变换及量化过程。具有一个或多个PU的给定CU也可包含一个或多个TU。在预测之后,视频编码器20可计算对应于PU的残余值。残余值包括像素差值,像素差值可变换成变换系数、经量化且使用TU扫描以产生串行化变换系数以用于熵解码。本申请通常使用术语“图像块”来指CU的解码节点。在一些特定应用中,本申请也可使用术语“图像块”来指包含解码节点以及PU及TU的树块,例如,LCU或CU。
所述视频编码器20编码视频数据。视频数据可包括一个或多个图片。视频编码器20可产生码流,所述码流以比特流的形式包含了视频数据的编码信息。所述编码信息可以包含编码图片数据及相关联数据。相关联数据可包含序列参数集(sequence paramater set,SPS)、图片参数集(picture paramater set,PPS)及其它语法结构。SPS可含有应用于零个或多个序列的参数。SPS描述的是编码视频序列(coded video sequence,CVS)一般特性的高层参数,序列参数集SPS包含该CVS中所有条带(slice)需要的信息。PPS可含有应用于零个或多个图片的参数。语法结构是指码流中以指定次序排列的零个或多个语法元素的集合。
作为一种可行的实施方式,HM支持各种PU大小的预测。假定特定CU的大小为2N×2N,HM支持2N×2N或N×N的PU大小的帧内预测,及2N×2N、2N×N、N×2N或N×N的对称PU大小的帧间预测。HM也支持2N×nU、2N×nD、nL×2N及nR×2N的PU大小的帧间预测的不对称分割。在不对称分割中,CU的一方向未分割,而另一方向分割成25%及75%。对应于25%区段的CU的部分由“n”后跟着“上(Up)”、“下(Down)”、“左(Left)”或“右(Right)”的指示来指示。因此,例如,“2N×nU”指水平分割的2N×2NCU,其中2N×0.5NPU在上部且2N×1.5NPU在底部。
在本申请中,“N×N”与“N乘N”可互换使用以指依照垂直维度及水平维度的图像块的像素尺寸,例如,16×16像素或16乘16像素。一般来说,16×16块将在垂直方向上具有16个像素(y=16),且在水平方向上具有16个像素(x=16)。同样地,N×N块一般在垂直方向上具有N个像素,且在水平方向上具有N个像素,其中N表示非负整数值。可将块中的像素排列成行及列。此外,块未必需要在水平方向上与在垂直方向上具有相同数目个像素。例如,块可包括N×M个像素,其中M未必等于N。
在使用CU的PU的帧内预测性或帧间预测性解码之后,视频编码器20可计算CU的TU的残余数据。PU可包括空间域(也称作像素域)中的像素数据,且TU可包括在将变换(例如,离散余弦变换(discrete cosine transform,DCT)、整数变换、小波变换或概念上类似的变换)应用于残余视频数据之后变换域中的系数。残余数据可对应于未经编码图像的像素与对应于PU的预测值之间的像素差。视频编码器20可形成包含CU的残余数据的TU,且接着变换TU以产生CU的变换系数。
JEM模型对视频图像的编码结构进行了进一步的改进,具体的,被称为“四叉树结合二叉树”(QTBT)的块编码结构被引入进来。QTBT结构摒弃了HEVC中的CU,PU,TU等概念,支持更灵活的CU划分形状,一个CU可以正方形,也可以是长方形。一个CTU首先进行四叉树划分,该四叉树的叶节点进一步进行二叉树划分。同时,在二叉树划分中存在两种划分模式,对称水平分割和对称竖直分割。二叉树的叶节点被称为CU,JEM的CU在预测和变换的过程中都不可以被进一步划分,也就是说JEM的CU,PU,TU具有相同的块大小。在现阶段的JEM中,CTU的最大尺寸为256×256亮度像素。
图2A为本申请实施例中视频编码器20的一种示意性框图。
如图2A所示,视频编码器20可以包括:预测模块21、求和器22、变换模块23、量化模块24和熵编码模块25。在一种示例下,预测模块21可以包括帧间预测模块211和帧内预测模块212,本申请实施例对预测模块21的内部结构不作限定。可选的,对于混合架构的视频编码器,视频编码器20也可以包括反量化模块26、反变换模块27和求和器28。
在图2A的一种可行的实施方式下,视频编码器20还可以包括存储模块29,应当理解的是,存储模块29也可以设置在视频编码器20之外。
在另一种可行的实施方式下,视频编码器20还可以包括滤波器(图2A中未示意)以对图像块的边界进行滤波从而从经重建的视频图像中去除伪影。在需要时,滤波器对求和器28的输出进行滤波。
可选地,视频编码器20还可以包括分割单元(图2A中未示意)。视频编码器20接收视频数据,且分割单元将视频数据分割成图像块。此分割也可包含分割成条带、图像块或其它较大单元,以及(例如)根据LCU及CU的四叉树结构进行图像块分割。视频编码器20示例性地说明编码在待编码的视频条带内的图像块的组件。一般来说,条带可划分成多个图像块(且可能划分成称作图像块的集合)。条带的类型包括I(主要用于帧内图像编码)、P(用于帧间前向参考预测图像编码)、B(用于帧间双向参考预测图像编码)。
预测模块21用于对当前需要处理的图像块进行帧内或者帧间预测得到当前块的预测值(本申请中可以称为预测信息)。在本申请实施例中,当前需要处理的图像块,可以简称为待处理块,也可以简称为当前图像块,也可以简称为当前块。还可以将在编码阶段,当前需要处理的图像块,简称为当前编码块,将在解码阶段,当前需要处理的图像块,简称为当前解码块或者当前译码块。
具体的,预测模块21包括的帧间预测模块211针对当前块进行帧间预测得到帧间预测值。帧内预测模块212针对当前块进行帧内预测得到帧内预测值。帧间预测模块211在已重建的图像中,为当前图像中的当前块寻找匹配的参考块,将参考块中的像素点的像素值作为当前块中像素点的像素值的预测信息或者预测值(以下不再区分信息和值),此过程称为运动估计(Motion estimation,ME)(如图2B所示),并传输当前块的运动信息。
需要说明的是,图像块的运动信息包括了预测方向的指示信息(通常为前向预测、 后向预测或者双向预测),一个或两个指向参考块的运动矢量(Motion vector,MV),以及参考块所在图像的指示信息(通常记为参考帧索引,Reference index)。
前向预测是指当前块从前向参考图像集合中选择一个参考图像获取参考块。后向预测是指当前块从后向参考图像集合中选择一个参考图像获取参考块。双向预测是指从前向和后向参考图像集合中各选择一个参考图像获取参考块。当使用双向预测方法时,当前块会存在两个参考块,每个参考块各自需要通过运动矢量和参考帧索引进行指示,然后根据两个参考块内像素点的像素值确定当前块内像素点像素值的预测值。
运动估计过程需要为当前块在参考图像中尝试多个参考块,最终使用哪一个或者哪几个参考块用作预测则使用率失真优化(Rate-distortion optimization,RDO)或者其他方法确定。
在预测模块21经由帧间预测或帧内预测产生当前块的预测值之后,视频编码器20通过从当前块减去预测值而形成残差信息。变换模块23用于对残差信息进行变换。变换模块23使用例如离散余弦变换(discrete cosine transformation,DCT)或概念上类似的变换(例如,离散正弦变换DST)将残差信息变换成残差变换系数。变换模块23可将所得残差变换系数发送到量化模块24。量化模块24对残差变换系数进行量化以进一步减小码率。在一些可行的实施方式中,量化模块24可接着执行包含经量化变换系数的矩阵的扫描。替代地,熵编码模块25可执行扫描。
在量化之后,熵编码模块25可熵编码经量化的残差变换系数得到码流。例如,熵编码模块25可执行上下文自适应性可变长度解码(CAVLC)、上下文自适应性二进制算术解码(CABAC)、基于语法的上下文自适应性二进制算术解码(SBAC)、概率区间分割熵(PIPE)解码或另一熵编码方法或技术。在通过熵编码模块25进行熵编码之后,可将经编码码流传输到视频解码器30或存档以供稍后传输或由视频解码器30检索。
反量化模块26及反变换模块27分别应用反量化及反变换,以在像素域中重构建残差块以供稍后用作参考图像的参考块。求和器28将经重构建得到的残差信息与通过预测模块21所产生的预测值相加以产生重建块,并将重建块作为参考块以供存储于存储模块29中。这些参考块可由预测模块21用作参考块以帧间或者帧内预测后续视频帧或图像中的块。
应当理解的是,视频编码器20的其它的结构变化可用于编码视频流。例如,对于某些图像块或者图像帧,视频编码器20可以直接地量化残差信息而不需要经变换模块23处理,相应地也不需要经反变换模块27处理;或者,对于某些图像块或者图像帧,视频编码器20没有产生残差信息,相应地不需要经变换模块23、量化模块24、反量化模块26和反变换模块27处理;或者,视频编码器20可以将经重构图像块作为参考块直接地进行存储而不需要经滤波器单元处理;或者,视频编码器20中量化模块24和反量化模块26可以合并在一起;或者,视频编码器20中变换模块23和反变换模块27可以合并在一起;或者,求和器22和求和器28可以合并在一起。
图2C为本申请实施例中视频解码器30的一种示意性框图。
如图2C所示,视频解码器30可以包括熵解码模块31、预测模块32、反量化模块 34、反变换模块35和重建模块36。在一种示例下,预测模块32可以包括运动补偿模块322和帧内预测模块321,本申请实施例对此不作限定。
在一种可行的实施方式中,视频解码器30还可以包括存储模块33。应当理解的是,存储模块33也可以设置在视频解码器30之外。在一些可行的实施方式中,视频解码器30可执行与关于来自图2A的视频编码器20描述的编码流程的示例性地互逆的解码流程。
在解码过程期间,视频解码器30从视频编码器20接收码流。视频解码器30接收到的码流连续通过熵解码模块31、反量化模块34以及反变换模块35分别进行熵解码、反量化、反变换之后得到残差信息。根据码流确定针对当前块使用的帧内预测还是帧间预测。若是帧内预测,则预测模块32中的帧内预测模块321利用当前块周围已重建块的参考像素的像素值按照所使用的帧内预测方法构建预测信息。如果是帧间预测,则运动补偿模块322需要解析出运动信息,并使用所解析出的运动信息在已重建的图像块中确定参考块,并将参考块内像素点的像素值作为预测信息(此过程称为运动补偿(motion compensation,MC))。重建模块36使用预测信息加上残差信息便可以得到重建信息。
如上文所注明,本申请示例性地涉及帧间解码。因而,本申请的特定技术可由运动补偿模块322来执行。在其它可行的实施方式中,视频解码器30的一个或多个其它单元可另外或替代地负责执行本申请的技术。
下面首先对本申请中涉及到的概念进行描述。
1)帧间预测模式
在HEVC中,使用两种帧间预测模式,分别为先进的运动矢量预测(advanced motion vector prediction,AMVP)模式和融合(merge)模式。
对于AMVP模式,先遍历当前块空域或者时域相邻的已编码块(记为邻块),根据各个邻块的运动信息构建候选运动矢量列表(也可以称为运动信息候选列表),然后通过率失真代价从候选运动矢量列表中确定最优的运动矢量,将率失真代价最小的候选运动信息作为当前块的运动矢量预测值(motion vector predictor,MVP)。其中,邻块的位置及其遍历顺序都是预先定义好的。率失真代价由公式(1)计算获得,其中,J表示率失真代价RD Cost,SAD为使用候选运动矢量预测值进行运动估计后得到的预测像素值与原始像素值之间的绝对误差和(sum of absolute differences,SAD),R表示码率,λ表示拉格朗日乘子。编码端将选择的运动矢量预测值在候选运动矢量列表中的索引值和参考帧索引值传递到解码端。进一步地,在MVP为中心的邻域内进行运动搜索获得当前块实际的运动矢量,编码端将MVP与实际运动矢量之间的差值(motion vector difference)传递到解码端。
J=SAD+λR (1)
对于Merge模式,先通过当前块空域或者时域相邻的已编码块的运动信息,构建候选运动矢量列表,然后通过计算率失真代价从候选运动矢量列表中确定最优的运动信息作为当前块的运动信息,再将最优的运动信息在候选运动矢量列表中位置的索引值 (记为merge index,下同)传递到解码端。当前块空域和时域候选运动信息如图3所示,空域候选运动信息来自于空间相邻的5个块(A0,A1,B0,B1和B2),若相邻块不可得(相邻块不存在或者相邻块未编码或者相邻块采用的预测模式不为帧间预测模式),则该相邻块的运动信息不加入候选运动矢量列表。当前块的时域候选运动信息根据参考帧和当前帧的图序计数(picture order count,POC)对参考帧中对应位置块的MV进行缩放后获得。首先判断参考帧中T位置的块是否可得,若不可得则选择C位置的块。
与AMVP模式类似,Merge模式的邻块的位置及其遍历顺序也是预先定义好的,且邻块的位置及其遍历顺序在不同模式下可能不同。
可以看到,在AMVP模式和Merge模式中,都需要维护一个候选运动矢量列表。每次向候选列表中加入新的运动信息之前都会先检查列表中是否已经存在相同的运动信息,如果存在则不会将该运动信息加入列表中。我们将这个检查过程称为候选运动矢量列表的修剪。列表修剪是为了防止列表中出现相同的运动信息,避免冗余的率失真代价计算。
在HEVC的帧间预测中,编码块内的所有像素都采用了相同的运动信息,然后根据运动信息进行运动补偿,得到编码块的像素的预测值。然而在编码块内,并不是所有的像素都有相同的运动特性,采用相同的运动信息可能会导致运动补偿预测的不准确,进而增加了残差信息。
现有的视频编码标准使用基于平动运动模型的块匹配运动估计,并且假设块中所有像素点的运动一致。但是由于在现实世界中,运动多种多样,存在很多非平动运动的物体,如旋转的的物体,在不同方向旋转的过山车,投放的烟花和电影中的一些特技动作,特别是在UGC场景中的运动物体,对它们的编码,如果采用当前编码标准中的基于平动运动模型的块运动补偿技术,编码效率会受到很大的影响,因此,产生了非平动运动模型,比如仿射运动模型,以便进一步提高编码效率。
基于此,根据运动模型的不同,AMVP模式可以分为基于平动模型的AMVP模式以及基于非平动模型的AMVP模式;Merge模式可以分为基于平动模型的Merge模式和基于非平动运动模型的Merge模式。
2)非平动运动模型
非平动运动模型预测指在编解码端使用相同的运动模型推导出当前块内每一个子运动补偿单元的运动信息,根据子运动补偿单元的运动信息进行运动补偿,得到预测块,从而提高预测效率。常用的非平动运动模型有4参数仿射运动模型或者6参数仿射运动模型。
其中,本申请实施例中涉及到的子运动补偿单元可以是一个像素点或按照特定方法划分的大小为N
1×N
2的像素块,其中,N
1和N
2均为正整数,N
1可以等于N
2,也可以不等于N
2。
4参数仿射运动模型如公式(2)所示:
4参数仿射运动模型可以通过两个像素点的运动矢量及其相对于当前块左上顶点像素的坐标来表示,将用于表示运动模型参数的像素点称为控制点。若采用左上顶点 (0,0)和右上顶点(W,0)像素点作为控制点,则先确定当前块左上顶点和右上顶点控制点的运动矢量(vx0,vy0)和(vx1,vy1),然后根据公式(3)得到当前块中每一个子运动补偿单元的运动信息,其中(x,y)为子运动补偿单元相对于当前块左上顶点像素的坐标,W为当前块的宽。
6参数仿射运动模型如公式(4)所示:
6参数仿射运动模型可以通过三个像素点的运动矢量及其相对于当前块左上顶点像素的坐标来表示。若采用左上顶点(0,0)、右上顶点(W,0)和左下顶点(0,H)像素点作为控制点,则先确定当前块左上顶点、右上顶点和左下顶点控制点的运动矢量分别为(vx0,vy0)和(vx1,vy1)和(vx2,vy2),然后根据公式(5)得到当前块中每一个子运动补偿单元的运动信息,其中(x,y)为子运动补偿单元相对于当前块的左上顶点像素的坐标,W和H分别为当前块的宽和高。
采用仿射运动模型进行预测的编码块称为仿射编码块。
通常的,可以使用基于仿射运动模型的先进运动矢量预测(Advanced Motion Vector Prediction,AMVP)模式或者基于仿射运动模型的融合(Merge)模式,获得仿射编码块的控制点的运动信息。
当前编码块的控制点的运动信息可以通过继承的控制点运动矢量预测方法或者构造的控制点运动矢量预测方法得到。
3)继承的控制点运动矢量预测方法
继承的控制点运动矢量预测方法,是指利用相邻已编码的仿射编码块的运动模型,确定当前块的候选的控制点运动矢量。
以图3所示的当前块为例,按照设定的顺序,比如A1->B1->B0->A0->B2的顺序遍历当前块周围的相邻位置块,找到该当前块的相邻位置块所在的仿射编码块,获得该仿射编码块的控制点运动信息,进而通过仿射编码块的控制点运动信息构造的运动模型,推导出当前块的控制点运动矢量(用于Merge模式)或者控制点的运动矢量预测值(用于AMVP模式)。A1->B1->B0->A0->B2仅作为一种示例,其它组合的顺序也适用于本申请。另外,相邻位置块不仅限于A1、B1、B0、A0、B2。
相邻位置块可以为一个像素点,按照特定方法划分的预设大小的像素块,比如可以为一个4x4的像素块,也可以为一个4x2的像素块,也可以为其他大小的像素块,不作限定。
下面以A1为例描述确定过程,其他情况以此类推:
如图4所示,若A1所在的编码块为4参数仿射编码块,则获得该仿射编码块左上顶点(x4,y4)的运动矢量(vx4,vy4)、右上顶点(x5,y5)的运动矢量(vx5,vy5);利用公式(6)计算获得当前仿射编码块左上顶点(x0,y0)的运动矢量(vx0,vy0),利用公式 (7)计算获得当前仿射编码块右上顶点(x1,y1)的运动矢量(vx1,vy1)。
通过如上基于A1所在的仿射编码块获得的当前块的左上顶点(x0,y0)的运动矢量(vx0,vy0)、右上顶点(x1,y1)的运动矢量(vx1,vy1)的组合为当前块的候选的控制点运动矢量。
若A1所在的编码块为6参数仿射编码块,则获得该仿射编码块左上顶点(x4,y4)的运动矢量(vx4,vy4)、右上顶点(x5,y5)的运动矢量(vx5,vy5)、左下顶点(x6,y6)的运动矢量(vx6,vy6);利用公式(8)计算获得当前块左上顶点(x0,y0)的运动矢量(vx0,vy0),利用公式(9)计算获得当前块右上顶点(x1,y1)的运动矢量(vx1,vy1)、利用公式(10)计算获得当前块左下顶点(x2,y2)的运动矢量(vx2,vy2)。
通过如上基于A1所在的仿射编码块获得的当前块的左上顶点(x0,y0)的运动矢量(vx0,vy0)、右上顶点(x1,y1)的运动矢量(vx1,vy1)、当前块左下顶点(x2,y2)的运动矢量(vx2,vy2)的组合为当前块的候选的控制点运动矢量。
需要说明的是,其他运动模型、候选位置、查找遍历顺序也可以适用于本申请,本申请实施例对此不做赘述。
需要说明的是,采用其他控制点来表示相邻和当前编码块的运动模型的方法也可以适用于本申请,此处不做赘述。
4)构造的控制点运动矢量(constructed control point motion vectors)预测方法1:
构造的控制点运动矢量预测方法,是指将当前块的控制点周边邻近的已编码块的运动矢量进行组合,作为当前仿射编码块的控制点的运动矢量,而不需要考虑周边邻近的已编码块是否为仿射编码块。
利用当前编码块周边邻近的已编码块的运动信息确定当前块左上顶点和右上顶点的运动矢量。以图5A所示为例对构造的控制点运动矢量预测方法进行描述。需要说明的是,图5A仅作为一种示例。
如图5A所示,利用左上顶点相邻已编码块A2,B2和B3块的运动矢量,作为当前块左上顶点的运动矢量的候选运动矢量;利用右上顶点相邻已编码块B1和B0块的运动矢 量,作为当前块右上顶点的运动矢量的候选运动矢量。将上述左上顶点和右上顶点的候选运动矢量进行组合,构成多个二元组,二元组包括的两个已编码块的运动矢量可以作为当前块的候选的控制点运动矢量,参见如下公式(11A)所示:
{v
A2,v
B1},{v
A2,v
B0},{v
B2,v
B1},{v
B2,v
B0},{v
B3,v
B1},{v
B3,v
B0} (11A);
其中,v
A2表示A2的运动矢量,v
B1表示B1的运动矢量,v
B0表示B0的运动矢量,v
B2表示B2的运动矢量,v
B3表示B3的运动矢量。
如图5A所示,利用左上顶点相邻已编码块A2,B2和B3块的运动矢量,作为当前块左上顶点的运动矢量的候选运动矢量;利用右上顶点相邻已编码块B1和B0块的运动矢量,作为当前块右上顶点的运动矢量的候选运动矢量,利用坐下顶点相邻已编码块A0、A1的运动矢量作为当前块左下顶点的运动矢量的候选运动矢量。将上述左上顶点、右上顶点以及左下顶点的候选运动矢量进行组合,构成三元组,三元组包括的三个已编码块的运动矢量可以作为当前块的候选的控制点运动矢量,参见如下公式(11B)、
(11C)所示:
{v
A2,v
B1,v
A0},{v
A2,v
B0,v
A0},{v
B2,v
B1,v
A0},{v
B2,v
B0,v
A0},{v
B3,v
B1,v
A0},{v
B3,v
B0,v
A0}(11B)
{v
A2,v
B1,v
A1},{v
A2,v
B0,v
A1},{v
B2,v
B1,v
A1},{v
B2,v
B0,v
A1},{v
B3,v
B1,v
A1},{v
B3,v
B0,v
A1}(11C);
其中,v
A2表示A2的运动矢量,v
B1表示B1的运动矢量,v
B0表示B0的运动矢量,v
B2表示B2的运动矢量,v
B3表示B3的运动矢量,v
A0表示A0的运动矢量,v
A1表示A1的运动矢量。
需要说明的是,其他控制点运动矢量的组合的方法也可适用于本申请,此处不做赘述。
需要说明的是,采用其他控制点来表示相邻和当前编码块的运动模型的方法也可以适用于本申请,此处不做赘述。
5)构造的控制点运动矢量(constructed control point motion vectors)预测方法2,参见图5B所示。
步骤501,获取当前块的各个控制点的运动信息。
以图5A所示为例,CPk(k=1,2,3,4)表示第k个控制点。A0,A1,A2,B0,B1,B2和B3为当前块的空域相邻位置,用于预测CP1、CP2或CP3;T为当前块的时域相邻位置,用于预测CP4。
设,CP1,CP2,CP3和CP4的坐标分别为(0,0),(W,0),(H,0)和(W,H),其中W和H为当前块的宽度和高度。
对于每个控制点,其运动信息按照以下顺序获得:
(1)对于CP1,检查顺序为B2->A2->B3,如果B2可得,则采用B2的运动信息。否则,检测A2,B3。若三个位置的运动信息均不可得,则无法获得CP1的运动信息。
(2)对于CP2,检查顺序为B0->B1;如果B0可得,则CP2采用B0的运动信息。否则,检测B1。若两个位置的运动信息均不可得,则无法获得CP2的运动信息。
(3)对于CP3,检测顺序为A0->A1。
(4)对于CP4,采用T的运动信息。
此处X可得表示包括X(X为A0,A1,A2,B0,B1,B2,B3或T)位置的块已经编码并且采用帧间预测模式;否则,X位置不可得。
需要说明的是,其他获得控制点的运动信息的方法也可适用于本申请,此处不做赘述。
步骤502,将控制点的运动信息进行组合,得到构造的控制点运动信息。
将两个控制点的运动信息进行组合构成二元组,用来构建4参数仿射运动模型。两个控制点的组合方式可以为{CP1,CP4},{CP2,CP3},{CP1,CP2},{CP2,CP4},{CP1,CP3},{CP3,CP4}。例如,采用CP1和CP2控制点组成的二元组构建的4参数仿射运动模型,可以记作Affine(CP1,CP2)。
将三个控制点的运动信息进行组合构成三元组,用来构建6参数仿射运动模型。三个控制点的组合方式可以为{CP1,CP2,CP4},{CP1,CP2,CP3},{CP2,CP3,CP4},{CP1,CP3,CP4}。例如,采用CP1、CP2和CP3控制点构成的三元组构建的6参数仿射运动模型,可以记作Affine(CP1,CP2,CP3)。
将四个控制点的运动信息进行组合构成的四元组,用来构建8参数双线性模型。采用CP1、CP2、CP3和CP4控制点构成的四元组构建的8参数双线性模型,记做Bilinear(CP1,CP2,CP3,CP4)。
本申请实施例中,为了描述方便,将由两个控制点(或者两个已编码块)的运动信息组合简称为二元组,将三个控制点(或者三个已编码块)的运动信息组合简称为三元组,将四个控制点(或者四个已编码块)的运动信息组合简称为四元组。
按照预置的顺序遍历这些模型,若组合模型对应的某个控制点的运动信息不可得,则认为该模型不可得;否则,确定该模型的参考帧索引,并将控制点的运动矢量进行缩放,若缩放后的所有控制点的运动信息一致,则该模型不合法。若确定构建该模型的控制点的运动信息均可得,并且模型合法,则将该构建该模型的控制点的运动信息加入运动信息候选列表中。
控制点的运动矢量缩放的方法如公式(12)所示:
其中,CurPoc表示当前帧的POC号,DesPoc表示当前块的参考帧的POC号,SrcPoc表示控制点的参考帧的POC号,MV
s表示缩放得到的运动矢量,MV表示控制点的运动矢量。
需要说明的是,亦可将不同控制点的组合转换为同一位置的控制点。
例如将{CP1,CP4},{CP2,CP3},{CP2,CP4},{CP1,CP3},{CP3,CP4}组合得到的4参数仿射运动模型转换为通过{CP1,CP2}或{CP1,CP2,CP3}来表示。转换方法为将控制点的运动矢量及其坐标信息,代入公式(2),得到模型参数,再将{CP1,CP2}的坐标信息代入公式(3),得到其运动矢量。
更直接地,可以按照以下公式(13)-(21)来进行转换,其中,W表示当前块的宽度,H表示当前块的高度,公式(13)-(21)中,(vx
0,vy
0)表示CP1的运动矢量,(vx
1,vy
1)表示CP2的运动矢量,(vx
2,vy
2)表示CP3的运动矢量,(vx
3,vy
3)表示CP4的运动矢量。
{CP1,CP2}转换为{CP1,CP2,CP3}可以通过如下公式(13)实现,即{CP1,CP2,CP3}中CP3的运动矢量可以通过公式(13)来确定:
{CP1,CP3}转换{CP1,CP2}或{CP1,CP2,CP3}可以通过如下公式(14)实现:
{CP2,CP3}转换为{CP1,CP2}或{CP1,CP2,CP3}可以通过如下公式(15)实现:
{CP1,CP4}转换为{CP1,CP2}或{CP1,CP2,CP3}可以通过如下公式(16)或者(17)实现:
{CP2,CP4}转换为{CP1,CP2}可以通过如下公式(18)实现,{CP2,CP4}转换为{CP1,CP2,CP3}可以通过公式(18)和(19)实现:
{CP3,CP4}转换为{CP1,CP2}可以通过如下公式(20)实现,{CP3,CP4}转换为{CP1,CP2,CP3}可以通过如下公式(20)和(21)实现:
例如将{CP1,CP2,CP4},{CP2,CP3,CP4},{CP1,CP3,CP4}组合的6参数仿射运动模型转换为控制点{CP1,CP2,CP3}来表示。转换方法为将控制点的运动矢量及其坐标信息,代入公式(4),得到模型参数,再将{CP1,CP2,CP3}的坐标信息代入公式(5),得到其运动矢量。
更直接地,可以按照以下公式(22)-(24)进行转换,其中,W表示当前块的宽度,H表示当前块的高度,公式(22)-(24)中,(vx
0,vy
0)表示CP1的运动矢量,(vx
1,vy
1)表示CP2的运动矢量,(vx
2,vy
2)表示CP3的运动矢量,(vx
3,vy
3)表示CP4的运动矢量。
{CP1,CP2,CP4}转换为{CP1,CP2,CP3}可以通过公式(22)实现:
{CP2,CP3,CP4}转换为{CP1,CP2,CP3}可以通过公式(23)实现:
{CP1,CP3,CP4}转换为{CP1,CP2,CP3}可以通过公式(24)实现:
6)基于仿射运动模型的先进运动矢量预测模式(Affine AMVP mode):
(1)构建候选运动矢量列表
利用继承的控制点运动矢量预测方法和/或构造的控制点运动矢量预测方法,构建基于仿射运动模型的AMVP模式的候选运动矢量列表。在本申请实施例中可以将基于仿射运动模型的AMVP模式的候选运动矢量列表称为控制点运动矢量预测值候选列表(control point motion vectors predictor candidate list),每个控制点的运动矢量预测值包括2个(4参数仿射运动模型)控制点的运动矢量或者包括3个(6参数仿射运动模型)控制点的运动矢量。
可选的,将控制点运动矢量预测值候选列表根据特定的规则进行剪枝和排序,并可将其截断或填充至特定的个数。
(2)确定最优的控制点运动矢量预测值
在编码端,利用控制点运动矢量预测值候选列表中的每个控制点运动矢量预测值,通过公式(3)/(5)获得当前编码块中每个子运动补偿单元的运动矢量,进而得到每个子运动补偿单元的运动矢量所指向的参考帧中对应位置的像素值,作为其预测值,进行采用仿射运动模型的运动补偿。计算当前编码块中每个像素点的原始值和预测值之间差值的平均值,选择最小平均值对应的控制点运动矢量预测值为最优的控制点运动矢量预测值,并作为当前编码块2个/3个控制点的运动矢量预测值。将表示该控制点运动矢量预测值在控制点运动矢量预测值候选列表中位置的索引号编码入码流发送给解码器。
在解码端,解析索引号,根据索引号从控制点运动矢量预测值候选列表中确定控制点运动矢量预测值(control point motion vectors predictor,CPMVP)。
(3)确定控制点的运动矢量
在编码端,以控制点运动矢量预测值作为搜索起始点在一定搜索范围内进行运动搜索获得控制点运动矢量(control point motion vectors,CPMV)。并将控制点运动矢量与控制点运动矢量预测值之间的差值(control point motion vectors differences,CPMVD)传递到解码端。
在解码端,解析控制点运动矢量差值,与控制点运动矢量预测值相加,得到控制点运动矢量。
7)仿射融合模式(Affine Merge mode):
利用继承的控制点运动矢量预测方法和/或构造的控制点运动矢量预测方法,构建控制点运动矢量融合候选列表(control point motion vectors merge candidate list)。
可选的,将控制点运动矢量融合候选列表根据特定的规则进行剪枝和排序,并可 将其截断或填充至特定的个数。
在编码端,利用融合候选列表中的每个控制点运动矢量,通过公式(3)/(5)获得当前编码块中每个子运动补偿单元(像素点或特定方法划分的大小为N
1×N
2的像素块)的运动矢量,进而得到每个子运动补偿单元的运动矢量所指向的参考帧中位置的像素值,作为其预测值,进行仿射运动补偿。计算当前编码块中每个像素点的原始值和预测值之间差值的平均值,选择差值的平均值最小对应的控制点运动矢量作为当前编码块2个/3个控制点的运动矢量。将表示该控制点运动矢量在候选列表中位置的索引号编码入码流发送给解码器。
在解码端,解析索引号,根据索引号从控制点运动矢量融合候选列表中确定控制点运动矢量(control point motion vectors,CPMV)。
另外,需要说明的是,本申请中,“至少一个”是指一个或者多个,“多个”是指两个或两个以上。“和/或”,描述关联对象的关联关系,表示可以存在三种关系,例如,A和/或B,可以表示:单独存在A,同时存在A和B,单独存在B的情况,其中A,B可以是单数或者复数。字符“/”一般表示前后关联对象是一种“或”的关系。“以下至少一项(个)”或其类似表达,是指的这些项中的任意组合,包括单项(个)或复数项(个)的任意组合。例如,a,b,或c中的至少一项(个),可以表示:a,b,c,a-b,a-c,b-c,或a-b-c,其中a,b,c可以是单个,也可以是多个。
在本申请中,当使用帧间预测模式来解码当前块时,可使用语法元素来用信号发送帧间预测模式。
目前所采用的解析当前块采用的帧间预测模式的部分语法结构,可以参见表1所示。需要说明的是,语法结构中的语法元素还可以通过其他标识来表示,本申请对此不作具体限定。
表1
语法元素merge_flag[x0][y0]可用于指示针对当前块是否采用融合模式。比如,当merge_flag[x0][y0]=1时,指示针对当前块采用融合模式,当merge_flag[x0][y0]=0时,指示针对当前块不采用融合模式。x0,y0表示当前块在视频图像的坐标。
变量allowAffineMerge可用于指示当前块是否满足采用基于仿射运动模型的merge模式的条件。比如allowAffineMerge=0,指示不满足采用基于仿射运动模型的 merge模式的条件,allowAffineMerge=1,指示满足采用基于仿射运动模型的merge模式的条件。采用基于仿射运动模型的merge模式的条件可以是:当前块的宽和高中均大于或者等于8。cbWidth表示当前块的宽,cbHeight表示当前块的高,即,当cbWidth<8或cbHeight<8时,allowAffineMerge=0,当cbWidth>=8且cbHeight>=8时,allowAffineMerge=1。
变量allowAffineInter可用于指示当前块是否满足采用基于仿射运动模型的AMVP模式的条件。比如allowAffineInter=0,指示不满足采用基于仿射运动模型的AMVP模式的条件,allowAffineInter=1,指示满足采用基于仿射运动模型的AMVP模式的条件。采用基于仿射运动模型的AMVP模式的条件可以是:当前块的宽和高中均大于或者等于16。即,当cbWidth<16或cbHeight<16时,allowAffineInter=0,当cbWidth>=16且cbHeight>=16时,allowAffineInter=1。
语法元素affine_merge_flag[x0][y0]可用于指示针对当前块是否采用基于仿射运动模型的merge模式。当前块所在条带的类型(slice_type)为P型或者B型。比如,affine_merge_flag[x0][y0]=1,指示针对当前块采用基于仿射运动模型的merge模式,affine_merge_flag[x0][y0]=0,指示针对当前块不采用基于仿射运动模型的merge模式,可以采用平运运动模型的merge模式。
语法元素affine_inter_flag[x0][y0]可用于指示在当前块所在条带为P型条带或者B型条带时,针对当前块是否采用基于仿射运动模型的AMVP模式。比如,allowAffineInter=1,指示针对当前块采用基于仿射运动模型的AMVP模式,allowAffineInter=0,指示针对当前块不采用基于仿射运动模型的AMVP模式,可以采用平动运动模型的AMVP模式。
语法元素affine_type_flag[x0][y0]可以用于指示:在当前块所在条带为P型条带或者B型条带时,针对当前块是否采用6参数仿射运动模型进行运动补偿。affine_type_flag[x0][y0]=0,指示针对当前块不采用6参数仿射运动模型进行运动补偿,可以仅采用4参数仿射运动模型进行运动补偿;affine_type_flag[x0][y0]=1,指示针对当前块采用6参数仿射运动模型进行运动补偿。
参见表2所示,MotionModelIdc[x0][y0]=1,指示采用4参数仿射运动模型,MotionModelIdc[x0][y0]=2,指示采用6参数仿射运动模型,MotionModelIdc[x0][y0]=0指示采用平动运动模型。
表2
变量MaxNumMergeCand用于表示最大列表长度,指示构造的候选运动矢量列表的最大长度。inter_pred_idc[x0][y0]用于指示预测方向。PRED_L1用于指示后向预测。 num_ref_idx_l0_active_minus1指示前向参考帧列表的参考帧个数,ref_idx_l0[x0][y0]指示当前块的前向参考帧索引值。mvd_coding(x0,y0,0,0)指示第一个运动矢量差。mvp_l0_flag[x0][y0]指示前向MVP候选列表索引值。PRED_L0指示前向预测。num_ref_idx_l1_active_minus1指示后向参考帧列表的参考帧个数。ref_idx_l1[x0][y0]指示当前块的后向参考帧索引值,mvp_l1_flag[x0][y0]表示后向MVP候选列表索引值。
表1中,ae(v)表示采用基于自适应二元算术编码(context-based adaptive binary arithmetic coding,cabac)编码的语法元素。
下面针对帧间预测流程进行详细描述。参见图6A所示。
步骤601:按照表1所示的语法结构,解析码流,确定当前块的帧间预测模式。
若确定当前块的帧间预测模式为基于仿射运动模型的AMVP模式,执行步骤602a。
即,语法元素merge_flag=0且affine_inter_flag=1,指示当前块的帧间预测模式为基于仿射运动模型的AMVP模式。
若确定当前块的帧间预测模式为基于仿射运动模型的融合(merge)模式,执行步骤602b。
即,语法元素merge_flag=1且affine_merge_flag=1,指示当前块的帧间预测模式为基于仿射运动模型的merge模式。
步骤602a:构造基于仿射运动模型的AMVP模式对应的候选运动矢量列表,执行步骤603a。
利用继承的控制点运动矢量预测方法和/或构造的控制点运动矢量预测方法,推导得到当前块的候选的控制点运动矢量,来加入候选运动矢量列表。
候选运动矢量列表可以包括二元组列表(当前编码块为4参数仿射运动模型)或三元组列表。二元组列表中包括一个或者多个用于构造4参数仿射运动模型的二元组。三元组列表中包括一个或者多个用于构造6参数仿射运动模型的三元组。
可选的,将候选运动矢量二元组/三元组列表根据特定的规则进行剪枝和排序,并可将其截断或填充至特定的个数。
A1:针对利用继承的控制运动矢量预测方式,构造候选运动矢量列表的流程进行说明。
以图4所示为例,比如,按照图4中A1->B1->B0->A0->B2的顺序遍历当前块周围的相邻位置块,找到相邻位置块所在的仿射编码块,获得该仿射编码块的控制点运动信息,进而通过仿射编码块的控制点运动信息构造运动模型,推导出当前块的候选的控制点运动信息。具体的,可以参见前面3)继承的控制点运动矢量预测方法中的相关描述,此处不再赘述。
示例性地,在当前块采用的仿射运动模型为4参数仿射运动模型(即,MotionModelIdc=1),若相邻仿射解码块为4参数仿射运动模型,则获取该仿射解码块2个控制点的运动矢量:左上控制点(x4,y4)的运动矢量值(vx4,vy4)和右上控制点(x5,y5)的运动矢量值(vx5,vy5)。仿射解码块为在编码阶段采用仿射运动模型进行 预测的仿射编码块。
利用相邻仿射解码块的2个控制点组成的4参数仿射运动模型,按照4仿射运动模型公式(6)、(7)分别推导得到当前块左上、右上2个控制点的运动矢量。
若相邻仿射解码块为6参数仿射运动模型,则获取相邻仿射解码块3个控制点的运动矢量,比如图4中,左上控制点(x4,y4)的运动矢量值(vx4,vy4)和右上控制点(x5,y5)的运动矢量值(vx5,vy5)和左下顶点(x6,y6)的运动矢量(vx6,vy6)。
利用相邻仿射解码块3个控制点组成的6参数仿射运动模型,按照6参数运动模型公式(8)、(9)、(10)分别推导得到当前块左上、右上和左下3个控制点的运动矢量。
示例性地,当前解码块的仿射运动模型是6参数仿射运动模型(即MotionModelIdc=2),
若相邻仿射解码块采用的仿射运动模型为6参数仿射运动模型,则获取相邻仿射解码块的3个控制点的运动矢量,比如图4中,左上控制点(x4,y4)的运动矢量值(vx4,vy4)和右上控制点(x5,y5)的运动矢量值(vx5,vy5)和左下顶点(x6,y6)的运动矢量(vx6,vy6)。
利用相邻仿射解码块3个控制点组成的6参数仿射运动模型,按照6参数仿射运动模型对应的公式(8)、(9)、(10)分别推导得到当前块左上、右上、左下3个控制点的运动矢量。
若相邻仿射解码块采用的仿射运动模型为4参数仿射运动模型,则获取该仿射解码块2个控制点的运动矢量:左上控制点(x4,y4)的运动矢量值(vx4,vy4)和右上控制点(x5,y5)的运动矢量值(vx5,vy5)。
利用相邻仿射解码块2个控制点组成的4参数仿射运动模型,按照4参数仿射运动模型公式(6)、(7)分别推导得到当前块左上、右上2个控制点的运动矢量。
需要说明的是,其他运动模型、候选位置、查找顺序也可以适用于本申请,在此不做赘述。需要说明的是,采用其他控制点来表示相邻和当前编码块的运动模型的方法也可以适用于本申请,在此不做赘述。
A2:针对利用构造的控制运动矢量预测方式,构造候选运动矢量列表的流程进行说明。
示例性的,当前解码块采用的仿射运动模型是4参数仿射运动模型(即,MotionModelIdc为1),利用当前编码块周边邻近的已编码块的运动信息确定当前编码块左上顶点和右上顶点的运动矢量。具体可以采用构造的控制点运动矢量预测方式1,或者采用构造的控制点运动矢量预测方法2,来构造候选运动矢量列表,具体方式参见上述4)和5)中的描述,此处不再赘述。
示例性的,当前解码块仿射运动模型是6参数仿射运动模型(即,MotionModelIdc为2),利用当前编码块周边邻近的已编码块的运动信息确定当前编码块左上顶点和右上顶点以及左下顶点的运动矢量。具体可以采用构造的控制点运动矢量预测方式1,或者采用构造的控制点运动矢量预测方法2,来构造候选运动矢量列表,具体方式参见上述4)和5)中的描述,此处不再赘述。
需要说明的是,其他控制点运动信息组合方式也可以适用于本申请,在此不做赘 述。
步骤603a:解析码流,确定最优的控制点运动矢量预测值,执行步骤604a。
B1,若当前解码块采用的仿射运动模型是4参数仿射运动模型(MotionModelIdc为1),则解析索引号,根据索引号从候选运动矢量列表中确定2个控制点的最优运动矢量预测值。
示例性的,索引号为mvp_l0_flag或mvp_l1_flag。
B2,若当前解码块采用的仿射运动模型是6参数仿射运动模型(MotionModelIdc为2),则解析索引号,根据索引号从候选运动矢量列表中确定3个控制点的最优运动矢量预测值。
步骤604a:解析码流,确定控制点的运动矢量。
C1,当前解码块采用的仿射运动模型是4参数仿射运动模型(MotionModelIdc为1),从码流中解码得到当前块的2个控制点的运动矢量差值,分别根据各控制点的运动矢量差值和运动矢量预测值获得控制点的运动矢量值。以前向预测为例,2个控制点的运动矢量差分别为mvd_coding(x0,y0,0,0)和mvd_coding(x0,y0,0,1)。
示例性的,从码流中解码得到左上位置控制点和右上位置控制点的运动矢量差值,并分别与运动矢量预测值相加,得到当前块左上位置控制点和右上位置控制点的运动矢量值。
C2,当前解码块仿射运动模型是6参数仿射运动模型(MotionModelIdc为2)
从码流中解码得到当前块的3个控制点的运动矢量差,分别根据各控制点的运动矢量差值和运动矢量预测值获得控制点的运动矢量值。以前向预测为例,3个控制点的运动矢量差分别为mvd_coding(x0,y0,0,0)和mvd_coding(x0,y0,0,1)、mvd_coding(x0,y0,0,2)。
示例性的,从码流中解码得到左上控制点、右上控制点和左下控制点的运动矢量差值,并分别与运动矢量预测值相加,得到当前块左上控制点、右上控制点和左下控制点的运动矢量值。
步骤602b:构造基于仿射运动模型的merge模式的运动信息候选列表。
具体的,可以利用继承的控制点运动矢量预测方法和/或构造的控制点运动矢量预测方法,构造基于仿射运动模型的融合模式的运动信息候选列表。
可选的,将运动信息候选列表根据特定的规则进行剪枝和排序,并可将其截断或填充至特定的个数。
D1:针对利用继承的控制运动矢量预测方式,构造候选运动矢量列表的流程进行说明。
利用继承的控制点运动矢量预测方法,推导得到当前块的候选的控制点运动信息,加入运动信息候选列表。
按照图3中A1,B1,B0,A0,B2的顺序遍历当前块的周边相邻位置块,找到该位置所在的仿射编码块,获得该仿射编码块的控制点运动信息,进而通过其运动模型,推导出当前块的候选的控制点运动信息。
如果此时候选运动矢量列表为空,则将该候选的控制点运动信息加入候选列表;否则依次遍历候选运动矢量列表中的运动信息,检查候选运动矢量列表中是否存在与该候选的控制点运动信息相同的运动信息。如果候选运动矢量列表中不存在与该候选的控制点运动信息相同的运动信息,则将该候选的控制点运动信息加入候选运动矢量列表。
其中,判断两个候选运动信息是否相同需要依次判断它们的前后向参考帧、以及各个前后向运动矢量的水平和竖直分量是否相同。只有当以上所有元素都不相同时才认为这两个运动信息是不同的。
如果候选运动矢量列表中的运动信息个数达到最大列表长度MaxNumMrgCand(MaxNumMrgCand为正整数,如1,2,3,4,5等,以下以5为例进行描述,不再赘述),则候选列表构建完毕,否则遍历下一个相邻位置块。
D2:利用构造的控制点运动矢量预测方法,推导得到当前块的候选的控制点运动信息,加入运动信息候选列表,参见图6B所示。
步骤601c,获取当前块的各个控制点的运动信息。可以参见5)中构造的控制点运动矢量预测方式2中,步骤501,此处不再赘述。
步骤602c,将控制点的运动信息进行组合,得到构造的控制点运动信息,可以参见图5B中步骤501,此处不再赘述。
步骤603c,将构造的控制点运动信息加入候选运动矢量列表。
若此时候选列表的长度小于最大列表长度MaxNumMrgCand,则按照预置的顺序遍历这些组合,得到合法的组合作为候选的控制点运动信息,如果此时候选运动矢量列表为空,则将该候选的控制点运动信息加入候选运动矢量列表;否则依次遍历候选运动矢量列表中的运动信息,检查候选运动矢量列表中是否存在与该候选的控制点运动信息相同的运动信息。如果候选运动矢量列表中不存在与该候选的控制点运动信息相同的运动信息,则将该候选的控制点运动信息加入候选运动矢量列表。
示例性的,一种预置的顺序如下:Affine(CP1,CP2,CP3)->Affine(CP1,CP2,CP4)->Affine(CP1,CP3,CP4)->Affine(CP2,CP3,CP4)->Affine(CP1,CP2)->Affine(CP1,CP3)->Affine(CP2,CP3)->Affine(CP1,CP4)->Affine(CP2,CP4)->Affine(CP3,CP4),总共10种组合。
若组合对应的控制点运动信息不可得,则认为该组合不可得。若组合可得,确定该组合的参考帧索引(两个控制点时,选择参考帧索引最小的作为该组合的参考帧索引;大于两个控制点时,先选择出现次数最多的参考帧索引,若有多个参考帧索引的出现次数一样多,则选择参考帧索引最小的作为该组合的参考帧索引),并将控制点的运动矢量进行缩放。若缩放后的所有控制点的运动信息一致,则该组合不合法。
可选地,本申请实施例还可以针对候选运动矢量列表进行填充,比如,经过上述遍历过程后,此时候选运动矢量列表的长度小于最大列表长度MaxNumMrgCand,则可以对候选运动矢量列表进行填充,直到列表的长度等于MaxNumMrgCand。
可以通过补充零运动矢量的方法进行填充,或者通过将现有列表中已存在的候选的运动信息进行组合、加权平均的方法进行填充。需要说明的是,其他获得候选运动 矢量列表填充的方法也可适用于本申请,在此不做赘述。
步骤S603b:解析码流,确定最优的控制点运动信息。
解析索引号,根据索引号从候选运动矢量列表中确定最优的控制点运动信息。
步骤604b:根据最优的控制点运动信息以及当前解码块采用的仿射运动模型获得当前块中每个子块的运动矢量值。
对于当前仿射解码块的每一个子块(一个子块也可以等效为一个运动补偿单元,子块的宽和高小于当前块的宽和高),可采用运动补偿单元中预设位置像素点的运动信息来表示该运动补偿单元内所有像素点的运动信息。假设运动补偿单元的尺寸为MxN,则预设位置像素点可以为运动补偿单元中心点(M/2,N/2)、左上像素点(0,0),右上像素点(M-1,0),或其他位置的像素点。以下以运动补偿单元中心点为例说明,参见图6C所示。图6C中V
0表示左上控制点的运动矢量,V
1表示右上控制点的运动矢量。每个小方框表示一个运动补偿单元。
运动补偿单元中心点相对于当前仿射解码块左上顶点像素的坐标使用公式(25)计算得到,其中i为水平方向第i个运动补偿单元(从左到右),j为竖直方向第j个运动补偿单元(从上到下),(x
(i,j),y
(i,j))表示第(i,j)个运动补偿单元中心点相对于当前仿射解码块左上控制点像素的坐标。
若当前仿射解码块采用的仿射运动模型为6参数仿射运动模型,将(x
(i,j),y
(i,j))代入6参数仿射运动模型公式(26),获得每个运动补偿单元中心点的运动矢量,作为该运动补偿单元内所有像素点的运动矢量(vx
(i,j),vy
(i,j))。
若当前仿射解码块采用的仿射运动模型为4仿射运动模型,将(x
(i,j),y
(i,j))代入4参数仿射运动模型公式(27),获得每个运动补偿单元中心点的运动矢量,作为该运动补偿单元内所有像素点的运动矢量(vx
(i,j),vy
(i,j))。
步骤605b:针对每个子块根据确定的子块的运动矢量值进行运动补偿得到该子块的像素预测值。
现有技术中通过步骤605a、604b得到了每个子块的运动矢量值后,再利用步骤606a、605b进行子块的运动补偿。在现有技术中,为了提高预测效率,会将子块划分为4x4,即每个4x4单元采用不同的运动矢量进行运动补偿。然而,运动补偿单元越小,平均每个像素进行运动补偿时所需要的读取的参考像素个数越多、插值的运算复杂度越高。进行一个MxN单元的运动补偿,所需要的总的参考像素个数为(M+T-1)*(N+T-1)*K,平均读取像素数为(M+T-1)*(N+T-1)*K/M/N。其中T为插值滤 波器的抽头数,如8,4,2等,K与预测方向相关,若为单向预测,则K=1,若为双向预测,则K=2。按照该计算方法,可以分别计算得到当插值滤波器抽头时为8时,4x4单元,8x4单元和8x8单元进行单向和双向预测的平均参考像素读取个数,如表3所示。需要说明的是,MxN中的M表示子块的宽width为M像素,N表示子块的高height为N像素,应当理解的是,M和N均为2
n,n为正整数。需要说明的是,表三中的UNI表示单向预测,BI表示双向预测。
表三
基于此,本申请实施例提供了一种图像预测方法及装置,用以降低现有技术中运动补偿的复杂度,同时兼顾预测效率。本申请中根据当前图像块的预测方向,选择不同的子块划分的方法。其中,方法和装置是基于同一发明构思的,由于方法及装置解决问题的原理相似,因此装置与方法的实施可以相互参见,重复之处不再赘述。
图7A为本申请实施例中的图像预测方法700的一种示意性流程图。需要说明的是,图像预测方法700既适用于解码视频图像的帧间预测,也适用于编码视频图像的帧间预测,该方法的执行主体可以是视频编码器(例如视频编码器20)或具有视频编码功能的电子设备,该方法700可以包括以下步骤:
步骤S701,获取当前图像块(例如当前解码块,当前仿射图像块或者当前仿射解码块)的控制点的运动矢量;
在一种实现方式下,可以按照以下步骤获取当前仿射解码块的控制点的运动矢量。
步骤1:确定当前仿射解码块的控制点的运动矢量的预测值;
示例性的,当前解码块采用的仿射运动模型是4参数仿射运动模型,利用当前解码块周边邻近的已解码块的运动信息确定当前解码块左上顶点和右上顶点的运动矢量的预测值。具体地,按照图5A中A2、B2、B3的顺序遍历当前块的周边相邻位置块, 查找与当前解码块预测方向相同的运动矢量,若找到,则将该运动矢量进行缩放后,作为当前仿射解码块的左上顶点控制点的运动矢量的预测值;若遍历空域相邻位置均未找到运动矢量,则采用零运动矢量作为当前仿射解码块的左上顶点控制点的运动矢量的预测值。同样地,按照图5A中B0、B1的顺序遍历当前块的周边相邻位置块,查找与当前解码块预测方向相同的运动矢量,若找到,则将该运动矢量进行缩放后,作为当前仿射解码块的右上顶点控制点的运动矢量的预测值;若遍历空域相邻位置均未找到运动矢量,则采用零运动矢量作为当前仿射解码块的右上顶点控制点的运动矢量的预测值。
需要说明的是,其他相邻块位置、相邻块的遍历顺序也可以适用于本申请,在此不做赘述。
示例性的,当前解码块采用的仿射运动模型是6参数仿射运动模型,利用当前解码块周边邻近的已解码块的运动信息确定当前解码块左上顶点、右上顶点和左下顶点的运动矢量的预测值。具体地,按照图5A中A2、B2、B3的顺序遍历当前块的周边相邻位置块,查找与当前解码块预测方向相同的运动矢量,若找到,则将该运动矢量进行缩放后,作为当前仿射解码块的左上顶点控制点的运动矢量的预测值;若遍历空域相邻位置均未找到运动矢量,则采用零运动矢量作为当前仿射解码块的左上顶点控制点的运动矢量的预测值。同样地,按照图5A中B0、B1的顺序遍历当前块的周边相邻位置块,查找与当前解码块预测方向相同的运动矢量,若找到,则将该运动矢量进行缩放后,作为当前仿射解码块的右上顶点控制点的运动矢量的预测值;若遍历空域相邻位置均未找到运动矢量,则采用零运动矢量作为当前仿射解码块的右上顶点控制点的运动矢量的预测值。按照图5A中A0、A1的顺序遍历当前块的周边相邻位置块,查找与当前解码块预测方向相同的运动矢量,若找到,则将该运动矢量进行缩放后,作为当前仿射解码块的左下顶点控制点的运动矢量的预测值;若遍历空域相邻位置均未找到运动矢量,则采用零运动矢量作为当前仿射解码块的左下顶点控制点的运动矢量的预测值。
需要说明的是,其他相邻块位置、相邻块的遍历顺序也可以适用于本申请,在此不做赘述。
步骤2:解析码流,确定控制点的运动矢量。
若当前解码块采用的仿射运动模型是4参数仿射运动模型,从码流中解码得到当前块的2个控制点的运动矢量差值,分别根据各控制点的运动矢量差值和运动矢量预测值获得控制点的运动矢量。
示例性的,从码流中解码得到左上顶点控制点和右上顶点控制点的运动矢量差值,并分别与对应的运动矢量预测值相加,得到当前块左上顶点控制点和右上顶点控制点的运动矢量值。
若当前解码块仿射运动模型是6参数仿射运动模型,从码流中解码得到当前块的3个控制点的运动矢量差,分别根据各控制点的运动矢量差值和运动矢量预测值获得控制点的运动矢量。
示例性的,从码流中解码得到左上顶点控制点、右上顶点控制点和左下顶点控制点的运动矢量差值,并分别与对应的运动矢量预测值相加,得到当前块左上顶点控制点、右上顶点控制点和左下顶点控制点的运动矢量值。
需要说明的是,本文中对获取当前图像块(例如当前解码块,当前仿射图像块或者当前仿射解码块)的控制点的运动矢量的方法不做限制,其他获取方法也可以适用于本申请,在此不做赘述。
步骤S703,根据所述当前图像块的控制点(例如多个仿射控制点)的运动矢量(例如运动矢量组)采用仿射变换模型获得当前图像块中每个子块的运动矢量,其中所述子块的尺寸是基于当前图像块的预测方向而确定的;
图7B示意了4x4的子块(亦称为运动补偿单元),图7C示意了8x8的子块(亦称为运动补偿单元),作为一种示例,相应子块(亦称为运动补偿单元)的中心点通过三角形来表示。
步骤S705,根据该当前图像块中每个子块的运动矢量值进行运动补偿,以得到每个子块的像素预测值。
相应地,在本申请实施例的具体实现方式下,步骤S703中:
如果所述当前图像块的预测方向为双向预测,所述当前图像块中子块的尺寸为UxV;或者,如果所述当前图像块的预测方向为单向预测,所述当前图像块中子块的尺寸为MxN,其中U,M表示所述子块的宽,V,N表示所述子块的高,以及U,V,M,N均为2n,n为正整数。在一种具体实现方式下,U=2M,V=2N。例如,M为4,N为4。相应地,U为8,V为8。
该方法700的具体实施方式如下:
若仿射解码块进行单向预测的子块划分为MxN,则令双向预测的子块划分为UxV。其中,U>=M,V>=N,且U和V不能同时等于M和N。更具体地,可以令U=2M,N=V,或U=M,V=2N,或U=2M,V=2N。M为4、8、16等整数,N为4、8、16等整数。在AMVP模式下,仿射解码块采用单向预测或者双向预测,通过语法元素inter_pred_idc确定。在merge模式下,仿射解码块采用单向预测或者双向预测,通过affine_merge_idx确定,当前仿射解码块的预测方向与affine_merge_idx指示的候选运动信息相同。
可选的,为了使得仿射解码块能够在水平方向划分为至少2个子块,在竖直方向划分为至少2个子块,可以采用双向预测的子块划分方式,对仿射解码块的使用条件进行限制。即,仿射解码块的允许使用尺寸为宽度W>=2U,高度H>=2V。当解码单元的尺寸不满足仿射解码块的使用条件时,不需要解析仿射相关的语法元素,如表一中的affine_inter_flag,affine_merge_flag。
可选的,为了使得仿射解码块能够在水平方向划分为至少2个子块,在竖直方向划分为至少2个子块,还可以采用单向预测的子块划分方式,对仿射解码块的使用条件进 行限制,并且当双向仿射解码块无法通过双向子块的划分方式进行划分时,则将该仿射解码块强制设置为单向预测。即,仿射解码块的允许使用尺寸为宽度W>=2M,高度H>=2N。当解码单元的尺寸不满足仿射解码块的使用条件时,不需要解析仿射相关的语法元素,如表一中的affine_inter_flag,affine_merge_flag。当双向仿射解码块的宽度W<2U或高度H<2V,则将其强制设置为单向预测。
需要说明的是,本申请的方法也可以用于其他子块划分模式,如ATMVP模式等。
具体地,在本申请的一个实施例中,若该仿射解码块为单向预测,则将其运动补偿单元的尺寸设置为4x4,若该仿射解码块为双向预测,则将其运动补偿单元的尺寸设置为8x4。
在一种实现方式下,当当前图像块(下文称为当前解码块)的尺寸满足W>=16和H>=16时,允许采用仿射模式。
可选的,当前解码块的尺寸满足W>=16和H>=8时,允许采用仿射模式。
可选的,当前解码块的尺寸满足W>=8和H>=8时,允许采用仿射模式。而当解码块的宽度W<16时,若该仿射解码块的预测方向为双向,则将其修改为单向。例如,丢弃后向预测的运动信息,将其转换为前向预测;或丢弃前向预测的运动信息,将其转换为后向预测。
需要说明的是,还可以通过对码流进行限制,使得当仿射解码块的宽度W<16时,不需要解析双向预测的标志位。
在本申请的另一个实施例中,若该仿射解码块为单向预测,则将其运动补偿单元的尺寸设置为4x4,若该仿射解码块为双向预测,则将其运动补偿单元的尺寸设置为4x8。
可选的,当解码块的尺寸满足W>=8和H>=16时,允许采用仿射模式。
可选的,当解码块的尺寸满足W>=8和H>=8时,允许采用仿射模式。而当解码块的高度H<16时,若该仿射解码块的预测方向为双向,则将其修改为单向。例如,丢弃后向预测的运动信息,将其转换为前向预测;或丢弃前向预测的运动信息,将其转换为后向预测。
需要说明的是,还可以通过对码流进行限制,使得当仿射解码块的高度H<16时,不需要解析双向预测的标志位。
在本申请的另一个实施例中,若该仿射解码块为单向预测,则将其运动补偿单元的尺寸设置为4x4,若该仿射解码块为双向预测,则根据该仿射解码块的尺寸进行自适应划分。划分方式可以为以下三种方式的其中一种:
1)如果该仿射解码块宽度W大于等于H,则将其运动补偿单元的尺寸设置为8x4,如果该仿射解码块宽度W小于H,则将其运动补偿单元的尺寸设置为4x8。
2)如果该仿射解码块宽度W大于H,则将其运动补偿单元的尺寸设置为8x4,如果 该仿射解码块宽度W小于等于H,则将其运动补偿单元的尺寸设置为4x8。
3)如果该仿射解码块宽度W大于H,则将其运动补偿单元的尺寸设置为8x4,如果该仿射解码块宽度W小于H,则将其运动补偿单元的尺寸设置为4x8,如果该仿射解码块宽度W等于H,则将其运动补偿单元的尺寸设置为8x8。
可选的,当解码块的尺寸满足W>=8和H>=8,并且W不等于8,H不等于8时,允许采用仿射模式。
可选的,当解码块的尺寸满足W>=8和H>=8时,允许采用仿射模式。而当解码块的宽度W等于8,高度H等于8时,若该仿射解码块的预测方向为双向,则将其修改为单向。例如,丢弃后向预测的运动信息,将其转换为前向预测;或丢弃前向预测的运动信息,将其转换为后向预测。
需要说明的是,还可以通过对码流进行限制,使得当仿射解码块的宽度W等于8,高度H等于8时,不需要解析双向预测的标志位。
可见,相对于现有技术中当前解码块被划分为MxN(即4x4)的子块,即每个MxN(即4x4)的子块采用对应的运动矢量进行运动补偿,本申请实施例中,当前解码块中的子块的尺寸是基于当前解码块的预测方向而确定的;例如,若当前解码块为单向预测,则当前解码块的子块的尺寸为4x4;若当前解码块为双向预测,则当前解码块的子块的尺寸为8x8。从图像的整体来看,本申请实施例的某些图像块的子块(或者运动补偿单元)的尺寸相对于现有技术中的子块(或者运动补偿单元)相对较大些,这样的话,平均每个像素进行运动补偿时所需要的读取的参考像素个数相对较少、插值的运算复杂度相对较低,从而本申请实施例在兼顾预测效率的同时,一定程度上降低了运动补偿的复杂度,从而提高了编解码性能。
图7D是示出根据本申请一种实施例的解码方法的过程1200的流程图。过程1200可由视频解码器30执行,具体的,可以由视频解码器30的帧间预测单元,以及熵解码单元(也称熵解码器)来执行。过程1200描述为一系列的步骤或操作,应当理解的是,过程700可以以各种顺序执行和/或同时发生,不限于图7D所示的执行顺序。假设具有多个视频帧的视频数据流正在使用视频解码器,图7D所示流程,相关描述如下:
应当理解的是,本申请实施例中,当前仿射解码块中的子块的尺寸是基于当前解码块的预测方向(例如单向预测或双向预测)而确定的,或者所述子块的尺寸是基于当前解码块的预测方向和所述当前解码块的尺寸而确定的。下文流程不再赘述,可以参见前面实施例。
步骤S1201:视频解码器确定当前解码块的帧间预测模式。
具体的,帧间预测模式可能为先进的运动矢量预测(Advanced Motion Vector Prediction,AMVP)模式,也可能为融合(merge)模式。
若确定出当前解码块的帧间预测模式为AMVP模式,则执行步骤S1211-S1216。
若确定出当前解码块的帧间预测模式为merge模式,则执行步骤S1221-S1225。
AMVP模式:
步骤S1211:视频解码器构建候选运动矢量预测值MVP列表。
具体地,视频解码器通过帧间预测单元(也称帧间预测模块)来构建候选运动矢量预测值MVP列表(也称候选运动矢量列表),可以采用如下提供的两种方式中的一种方式来构建,或者采用两种方式结合的形式来构建,构建的候选运动矢量预测值MVP列表可以为三元组的候选运动矢量预测值MVP列表,也可以为二元组的候选运动矢量预测值MVP列表;以上两种方式具体如下:
方式一,采用基于运动模型的运动矢量预测方法构建候选运动矢量预测值MVP列表。
首先,按照预先规定的顺序遍历当前解码块的全部或部分相邻块,从而确定其中的相邻仿射解码块,确定出的相邻仿射解码块的数量可能为一个也可能为多个。例如,可以依次遍历图7A所示的相邻块A、B、C、D、E,以确定出相邻块A、B、C、D、E中的相邻仿射解码块。该帧间预测单元至少会根据一个相邻仿射解码块确定一组候选运动矢量预测值(每一组候选运动矢量预测值为一个二元组或者三元组),下面以一个相邻仿射解码块为例进行介绍,为了便于描述称该一个相邻仿射解码块为第一相邻仿射解码块,具体如下:
根据第一相邻仿射解码块的控制点的运动矢量确定第一仿射模型,进而根据第一仿射模型预测该当前解码块的控制点的运动矢量。当前解码块的参数模型不同时,基于第一相邻仿射解码块的控制点的运动矢量预测当前解码块的控制点的运动矢量的方式也不同,因此下面分情况进行描述。
A、当前解码块的参数模型为4参数仿射变换模型:
若第一相邻仿射解码块位于当前解码块上方的编码树单元(Coding Tree Unit,CTU)且所述第一相邻仿射解码块为四参数仿射解码块,则获取该第一相邻仿射解码块最下侧两个控制点的运动矢量,例如,可以获取该第一相邻仿射解码块左下控制点的位置坐标(x
6,y
6)和运动矢量(vx
6,vy
6),以及右下控制点的位置坐标(x
7,y
7)和运动矢量值(vx
7,vy
7)(步骤S1201)。
根据该第一相邻仿射解码块最下侧两个控制点的运动矢量和坐标位置组成第一仿射模型(这时得到的第一仿射模型为4参数仿射模型)(步骤S1202)。
根据该第一仿射模型预测当前解码块的控制点的运动矢量,例如,可以将该当前解码块的左上控制点的位置坐标和右上控制点的位置坐标分别带入到该第一仿射模型,从而预测出当前解码块的左上控制点的运动矢量、右上控制点的运动矢量,具体如公式(1)、(2)所示(步骤S1203)。
在公式(1)、(2)中,(x
0,y
0)为当前解码块的左上控制点的坐标,(x
1,y
1)为当前解码块的右上控制点的坐标;另外,(vx
0,vy
0)为预测的当前解码块的左上控制点的运动矢量,(vx
1,vy
1)为预测的当前解码块的右上控制点的运动矢量。
可选的,所述第一相邻仿射解码块的左下控制点的位置坐标(x
6,y
6)和所述右下控制点的位置坐标(x
7,y
7)均为根据所述第一相邻仿射解码块的左上控制点的位置坐标(x
4,y
4)计算得到的,其中,所述第一相邻仿射解码块的左下控制点的位置坐标(x
6,y
6)为(x
4,y
4+cuH),所述第一相邻仿射解码块的右下控制点的位置坐标(x
7,y
7)为(x
4+cuW,y
4+cuH),cuW为所述第一相邻仿解码块的宽度,cuH为所述第一相邻仿射解码块的高度;另外,所述第一相邻仿射解码块的左下控制点的运动矢量为所述第一相邻仿射解码块的左下子块的运动矢量,所述第一相邻仿射解码块的右下控制点的运动矢量为第一相邻仿射解码块的右下子块的运动矢量。可以看出,第一相邻仿射解码块的左下控制点的位置坐标和所述右下控制点的位置坐标均是推导得到的,而不是从内存中读取得到的,因此采用该方法能够进一步减少内存的读取,提高了解码性能。作为另外一种可选方案,也可以在内存中预选存储左下控制点和右下控制点的位置坐标,后续要用的时候从内存中读取。
若第一相邻仿射解码块位于当前解码块上方的编码树单元(Coding Tree Unit,CTU)且所述第一相邻仿射解码块为六参数仿射解码块,则不基于第一相邻仿射解码块生成当前块的控制点的候选运动矢量预测值。
若该第一相邻仿射解码块不位于当前解码块的上方CTU,则预测当前解码块的控制点的运动矢量的方式此处不作限定。但是为了便于理解,下面也例举一种可选的确定方式:
可以获取该第一相邻仿射解码块的三个控制点的位置坐标和运动矢量,例如,左上控制点的位置坐标(x
4,y
4)和运动矢量值(vx
4,vy
4)、右上控制点的位置坐标(x
5,y
5)和运动矢量值(vx
5,vy
5)、左下控制点的位置坐标(x
6,y
6)和运动矢量(vx
6,vy
6)。
根据该第一相邻仿射解码块的三个控制点的位置坐标和运动矢量组成6参数仿射
模型。
将该当前解码块的左上控制点的位置坐标(x
0,y
0)和右上控制点的位置坐标(x
1,y
1)代入6参数仿射模型预测当前解码块的左上控制点的运动矢量,以及右上控制点的运动矢量,具体如公式(4)、(5)所示。
在公式(4)、(5)中,(vx
0,vy
0)为预测的当前解码块的左上控制点的运动矢量,(vx
1,vy
1)为预测的当前解码块的右上控制点的运动矢量。
B、当前解码块的参数模型为6参数仿射变换模型,推导的方式可以为:
若该第一相邻仿射解码块位于当前解码块的上方CTU且第一相邻仿射解码块为四参数仿射解码块,则获取该第一相邻仿射解码块最下侧两个控制点的位置坐标和运动矢量,例如,可以获取该第一相邻仿射解码块左下控制点的位置坐标(x
6,y
6)和运动矢量(vx
6,vy
6),以及右下控制点的位置坐标(x
7,y
7)和运动矢量值(vx
7,vy
7)。
根据该第一相邻仿射解码块最下侧两个控制点的运动矢量组成第一仿射模型(此时得到的第一仿射模型为一个4参数仿射模型)。
根据该第一仿射模型预测当前解码块的控制点的运动矢量,例如,可以将该当前解码块的左上控制点的位置坐标、右上控制点的位置坐标、左下控制点的位置坐标分别带入到该第一仿射模型,从而预测当前解码块的左上控制点的运动矢量、右上控制点的运动矢量和左下控制点的运动矢量,具体如公式(1)、(2)、(3)所示。
公式(1)、(2)以上已有描述,在公式(1)、(2)、(3)中,(x
0,y
0)为当前解码块的左上控制点的坐标,(x
1,y
1)为当前解码块的右上控制点的坐标,(x
2,y
2)为当前解码块的左下控制点的坐标;另外,(vx
0,vy
0)为预测的当前解码块的左上控制点的运动矢量,(vx
1,vy
1)为预测的当前解码块的右上控制点的运动矢量,(vx
2,vy
2)为预测的当前解码块的右左下控制点的运动矢量。
若第一相邻仿射解码块位于当前解码块上方的编码树单元(Coding Tree Unit,CTU)且所述第一相邻仿射解码块为六参数仿射解码块,则不基于第一相邻仿射解码块生成当前块的控制点的候选运动矢量预测值。
若该第一相邻仿射解码块不位于当前解码块的上方CTU,则预测当前解码块的控制点的运动矢量的方式此处不作限定。但是为了便于理解,下面也例举一种可选的确定方式:
可以获取该第一相邻仿射解码块的三个控制点的位置坐标和运动矢量,例如,左上控制点的位置坐标(x
4,y
4)和运动矢量值(vx
4,vy
4)、右上控制点的位置坐标(x
5,y
5)和运动矢量值(vx
5,vy
5)、左下控制点的位置坐标(x
6,y
6)和运动矢量(vx
6,vy
6)。
根据该第一相邻仿射解码块的三个控制点的位置坐标和运动矢量组成6参数仿射
模型。
将该当前解码块的左上控制点的位置坐标(x
0,y
0)、右上控制点的位置坐标(x
1,y
1)和左下控制点的位置坐标(x
2,y
2)代入6参数仿射模型预测当前解码块的左上控制点的运动矢量,右上控制点的运动矢量,及左下控制点的运动矢量,如公式(4)、(5)、(6)所示。
公式(4)、(5)前面已有描述,在公式(4)、(5)、(6)中,(vx
0,vy
0)为预测 的当前解码块的左上控制点的运动矢量,(vx
1,vy
1)为预测的当前解码块的右上控制点的运动矢量,(vx
2,vy
2)为预测的当前解码块的左下控制点的运动矢量。
方式二,采用基于控制点组合的运动矢量预测方法构建候选运动矢量预测值MVP列表。
当前解码块的参数模型不同时构建候选运动矢量预测值MVP列表的方式也不同,下面展开描述。
A、当前解码块的参数模型为4参数仿射变换模型,推导的方式可以为:
利用当前解码块周边邻近的已解码块的运动信息预估当前解码块左上顶点和右上顶点的运动矢量。如图7B所示:首先,利用左上顶点相邻已解码块A和/或B和/或C块的运动矢量,作为当前解码块左上顶点的运动矢量的候选运动矢量;利用右上顶点相邻已解码块D和/或E块的运动矢量,作为当前解码块右上顶点的运动矢量的候选运动矢量。将上述左上顶点的一个候选运动矢量和右上顶点的一个候选运动矢量进行组合可得到一组候选运动矢量预测值,按照这种组合方式组合得到的多条记录可以构成候选运动矢量预测值MVP列表。
B、当前解码块参数模型是6参数仿射变换模型,推导的方式可以为:
利用当前解码块周边邻近的已解码块的运动信息预估当前解码块左上顶点和右上顶点的运动矢量。如图7B所示:首先,利用左上顶点相邻已解码块A和/或B和/或C块的运动矢量,作为当前解码块左上顶点的运动矢量的候选运动矢量;利用右上顶点相邻已解码块D和/或E块的运动矢量,作为当前解码块右上顶点的运动矢量的候选运动矢量;利用左下顶点相邻已解码块F和/或G块的运动矢量,作为当前解码块左下顶点的运动矢量的候选运动矢量。将上述左上顶点的一个候选运动矢量、右上顶点的一个候选运动矢量和左下顶点的一个候选运动矢量进行组合可得到一组候选运动矢量预测值,按照这种组合方式组合得到的多组候选运动矢量预测值可以构成候选运动矢量预测值MVP列表。
需要说明的是,可仅采用方式一预测得到的候选运动矢量预测值来构建候选运动矢量预测值MVP列表,也可仅采用方式二预测得到的候选运动矢量预测值来构建候选运动矢量预测值MVP列表,还可采用方式一预测得到的候选运动矢量预测值和方式二预测得到的候选运动矢量预测值来共同构建候选运动矢量预测值MVP列表。另外,还可将候选运动矢量预测值MVP列表按照预先配置的规则进行剪枝和排序,然后将其截断或填充至特定个数。当候选运动矢量预测值MVP列表中的每一组候选运动矢量预测值包括三个控制点的运动矢量预测值时,可称该候选运动矢量预测值MVP列表为三元组列表;当候选运动矢量预测值MVP列表中的每一组候选运动矢量预测值包括两个控制点的运动矢量预测值时,可称该候选运动矢量预测值MVP列表为二元组列表。
步骤S1212:视频解码器解析码流,以得到索引和运动矢量差值MVD。
具体地,视频解码器可以通过熵解码单元解析码流,该索引用于指示当前解码块的目标候选运动矢量组,该目标候选运动矢量表示当前解码块的一组控制点的运动矢量预测值。
步骤S1213:视频解码器根据所述索引,从候选运动矢量预测值MVP列表中确定目 标运动矢量组。
具体地,视频解码器根据该索引从候选MVP列表中确定出的目标候选运动矢量组用于作为最优候选运动矢量预测值(可选的,当候选运动矢量预测值MVP列表的长度为1时,不需要解析码流得到索引,直接可以确定目标运动矢量组),下面对该最优候选运动矢量预测值进行简单介绍。
若当前解码块的参数模型是4参数仿射变换模型,那么从以上建立的候选运动矢量预测值MVP列表中选择的是2个控制点的最优运动矢量预测值;例如,该视频解码器从码流中解析索引号,再根据索引号从二元组的候选运动矢量预测值MVP列表中确定2个控制点的最优运动矢量预测值,该候选运动矢量预测值MVP列表中每组候选运动矢量预测值各自对应有各自的索引号。
若当前解码块的参数模型是6参数仿射变换模型,那么从以上建立的候选运动矢量预测值MVP列表中选择的是3个控制点的最优运动矢量预测值;例如,该视频解码器从码流中解析索引号,再根据索引号从三元组的候选运动矢量预测值MVP列表中确定3个控制点的最优运动矢量预测值,该候选运动矢量预测值MVP列表中每组候选运动矢量预测值各自对应有各自的索引号。
步骤S1214:视频解码器根据目标候选运动矢量组和从码流中解析出的运动矢量差值MVD确定当前解码块的控制点的运动矢量。
若当前解码块的参数模型是4参数仿射变换模型,那么从码流中解码得到当前解码块2个控制点的运动矢量差值,分别根据各控制点的运动矢量差值和所述索引指示的目标候选运动矢量组获得新的候选运动矢量组。例如,从码流中解码得到左上控制点的运动矢量差值MVD和右上控制点的运动矢量差值MVD,并分别与目标候选运动矢量组中左上控制点和右上控制点的运动矢量相加从而得到新的候选运动矢量组,因此,该新的候选运动矢量组包括当前解码块左上控制点和右上控制点的新的运动矢量值。
可选的,还可以根据新的候选运动矢量组中当前解码块2个控制点的运动矢量值,采用4参数仿射变换模型获得第3个控制点的运动矢量值。例如,获得当前解码块左上控制点的运动矢量(vx
0,vy
0)和右上控制点的运动矢量(vx
1,vy
1),然后利用公式(7)计算获得当前解码块左下控制点(x
2,y
2)的运动矢量(vx
2,vy
2)。
其中,(x
0,y
0)为左上控制点的位置坐标,(x
1,y
1)为右上控制点的位置坐标,W为当前解码块的宽,H为当前解码块的高。
若当前解码块参数模型是6参数仿射变换模型,那么从码流中解码得到当前解码块3个控制点的运动矢量差值,分别根据各控制点的运动矢量差值MVD和所述索引指示的目标候选运动矢量组获得新的候选运动矢量组。例如,从码流中解码得到左上控制点的运动矢量差值MVD、右上控制点的运动矢量差值MVD和左下控制点的运动矢量差值,并分别与目标候选运动矢量组中左上控制点、右上控制点、左下控制点的运动矢量相加从而得到新的候选运动矢量组,因此该新的候选运动矢量组包括当前解码块 左上控制点、右上控制点和左下控制点的运动矢量值。
步骤S1215:视频解码器根据以上确定出的当前解码块的控制点的运动矢量值采用仿射变换模型获得当前解码块中每个子块的运动矢量值,其中,所述子块的尺寸是基于当前图像块的预测方向而确定的。
具体地,基于目标候选运动矢量组和MVD得到的新的候选运动矢量组中包括两个(左上控制点和右上控制点)或者三个控制点(例如,左上控制点、右上控制点和左下控制点)的运动矢量。对于当前解码块的每一个子块(一个子块也可以等效为一个运动补偿单元),可采用运动补偿单元中预设位置像素点的运动信息来表示该运动补偿单元内所有像素点的运动信息。假设运动补偿单元的尺寸为MxN(M小于等于当前解码块的宽度W,N小于等于当前解码块的高度H,其中M、N、W、H为正整数,通常为2的幂次方,如4、8、16、32、64、128等),则预设位置像素点可以为运动补偿单元中心点(M/2,N/2)、左上像素点(0,0),右上像素点(M-1,0),或其他位置的像素点。图7C示意了4x4的运动补偿单元,图7D示意了8x8的运动补偿单元。
运动补偿单元中心点相对于当前解码块左上顶点像素的坐标使用公式(8-1)计算得到,其中i为水平方向第i个运动补偿单元(从左到右),j为竖直方向第j个运动补偿单元(从上到下),(x
(i,j),y
(i,j))表示第(i,j)个运动补偿单元中心点相对于当前解码块左上控制点像素的坐标。再根据当前解码块的仿射模型类型(6参数或4参数),将(x
(i,j),y
(i,j))代入6参数仿射模型公式(8-2)或者将(x
(i,j),y
(i,j))代入4参数仿射模型公式(8-3),获得每个运动补偿单元中心点的运动信息,作为该运动补偿单元内所有像素点的运动矢量(vx
(i,j),vy
(i,j))。
可选的,当前解码块为6参数解码块时,在基于所述目标候选运动矢量组得到所述当前解码块的一个或多个子块的运动矢量时,若所述当前解码块的下边界与所述当前解码块所在的CTU的下边界重合,则所述当前解码块的左下角的子块的运动矢量为根据所述三个控制点构造的6参数仿射模型和所述当前解码块的左下角的位置坐标(0,H)计算得到,所述当前解码块的右下角的子块的运动矢量为根据所述三个控制点构造的6参数仿射模型和所述当前解码块的右下角的位置坐标(W,H)计算得到。例如,将当前解码块的左下角的位置坐标(0,H)代入到该6参数仿射模型即可得到当前解码块的左下角的子块的运动矢量(而不是将左下角的子块的中心点坐标代入该仿射模型进行计算),将当前解码块的右下角的位置坐标(W,H)代入到该6参数仿射模型即 可得到当前解码块的右下角的子块的运动矢量(而不是将右下角的子块的中心点坐标代入该仿射模型进行计算)。这样一来,该当前解码块的左下控制点的运动矢量和右下控制点的运动矢量被用到时(例如,后续其他块基于该当前块的左下控制点和右下控制点的运动矢量构建该其他块的候选运动矢量预测值MVP列表),用到的是准确的值而不是估算值。其中,W为该当前解码块的宽,H为该当前解码块的高。
可选的,当前解码块为4参数解码块时,在基于所述目标候选运动矢量组得到所述当前解码块的一个或多个子块的运动矢量时,若所述当前解码块的下边界与所述当前解码块所在的CTU的下边界重合,则所述当前解码块的左下角的子块的运动矢量为根据所述两个控制点构造的4参数仿射模型和所述当前解码块的左下角的位置坐标(0,H)计算得到,所述当前解码块的右下角的子块的运动矢量为根据所述两个控制点构造的4参数仿射模型和所述当前解码块的右下角的位置坐标(W,H)计算得到。例如,将当前解码块的左下角的位置坐标(0,H)代入到该4参数仿射模型即可得到当前解码块的左下角的子块的运动矢量(而不是将左下角的子块的中心点坐标代入该仿射模型进行计算),将当前解码块的右下角的位置坐标(W,H)代入到该四参数仿射模型即可得到当前解码块的右下角的子块的运动矢量(而不是将右下角的子块的中心点坐标代入该仿射模型进行计算)。这样一来,该当前解码块的左下控制点的运动矢量和右下控制点的运动矢量被用到时(例如,后续其他块基于该当前块的左下控制点和右下控制点的运动矢量构建该其他块的候选运动矢量预测值MVP列表),用到的是准确的值而不是估算值。其中,W为该当前解码块的宽,H为该当前解码块的高。
步骤S1216:该视频解码器根据该当前解码块中每个子块的运动矢量值进行运动补偿,以得到每个子块的像素预测值,例如,通过每个子块的运动矢量和参考帧索引值,在参考帧中找到对应的子块,进行插值滤波,得到每个子块的像素预测值。
Merge模式:
步骤S1221:视频解码器构建候选运动信息列表。
具体地,视频解码器通过帧间预测单元(也称帧间预测模块)来构建候选运动信息列表(也称候选运动矢量列表),可以采用如下提供的两种方式中的一种方式来构建,或者采用两种方式结合的形式来构建,构建的候选运动信息列表为三元组的候选运动信息列表;以上两种方式具体如下:
方式一,采用基于运动模型的运动矢量预测方法构建候选运动信息列表。
首先,按照预先规定的顺序遍历当前解码块的全部或部分相邻块,从而确定其中的相邻仿射解码块,确定出的相邻仿射解码块的数量可能为一个也可能为多个。例如,可以依次遍历图7A所示的相邻块A、B、C、D、E,以确定出相邻块A、B、C、D、E中的相邻仿射解码块。该帧间预测单元会根据每个相邻仿射解码块确定一组候选运动矢量(每一组候选运动矢量为一个二元组或者三元组),下面以一个相邻仿射解码块为例进行介绍,为了便于描述称该一个相邻仿射解码块为第一相邻仿射解码块,具体如下:
根据第一相邻仿射解码块的控制点的运动矢量确定第一仿射模型,进而根据第一仿射模型预测该当前解码块的控制点的运动矢量,具体描述如下:
若该第一相邻仿射解码块位于当前解码块的上方CTU且该第一相邻仿射解码块为四参数仿射解码块,则获取该第一相邻仿射解码块最下侧两个控制点的位置坐标和运动矢量,例如,可以获取该第一相邻仿射解码块左下控制点的位置坐标(x
6,y
6)和运动矢量(vx
6,vy
6),以及右下控制点的位置坐标(x
7,y
7)和运动矢量值(vx
7,vy
7)。
根据该第一相邻仿射解码块最下侧两个控制点的运动矢量组成第一仿射模型(此时得到的第一仿射模型为一个4参数仿射模型)。
可选的,根据该第一仿射模型预测当前解码块的控制点的运动矢量,例如,可以将该当前解码块的左上控制点的位置坐标、右上控制点的位置坐标、左下控制点的位置坐标分别带入到该第一仿射模型,从而预测当前解码块的左上控制点的运动矢量、右上控制点的运动矢量和左下控制点的运动矢量,组成候选运动矢量三元组,加入候选运动信息列表,具体如公式(1)、(2)、(3)所示。
可选的,根据该第一仿射模型预测当前解码块的控制点的运动矢量,例如,可以将该当前解码块的左上控制点的位置坐标、右上控制点的位置坐标分别带入到该第一仿射模型,从而预测当前解码块的左上控制点的运动矢量和右上控制点的运动矢量,组成候选运动矢量二元组,加入候选运动信息列表,具体如公式(1)、(2)所示。
在公式(1)、(2)、(3)中,(x
0,y
0)为当前解码块的左上控制点的坐标,(x
1,y
1)为当前解码块的右上控制点的坐标,(x
2,y
2)为当前解码块的左下控制点的坐标;另外,(vx
0,vy
0)为预测的当前解码块的左上控制点的运动矢量,(vx
1,vy
1)为预测的当前解码块的右上控制点的运动矢量,(vx
2,vy
2)为预测的当前解码块的左下控制点的运动矢量。
若第一相邻仿射解码块位于当前解码块上方的编码树单元(Coding Tree Unit,CTU)且所述第一相邻仿射解码块为六参数仿射解码块,则不基于第一相邻仿射解码块生成当前块的控制点的候选运动矢量预测值。
若该第一相邻仿射解码块不位于当前解码块的上方CTU,则预测当前解码块的控制点的运动矢量的方式此处不作限定。但是为了便于理解,下面也例举一种可选的确定方式:
可以获取该第一相邻仿射解码块的三个控制点的位置坐标和运动矢量,例如,左上控制点的位置坐标(x
4,y
4)和运动矢量值(vx
4,vy
4)、右上控制点的位置坐标(x
5,y
5)和运动矢量值(vx
5,vy
5)、左下控制点的位置坐标(x
6,y
6)和运动矢量(vx
6,vy
6)。
根据该第一相邻仿射解码块的三个控制点的位置坐标和运动矢量组成6参数仿射
模型。
将该当前解码块的左上控制点的位置坐标(x
0,y
0)、右上控制点的位置坐标(x
1,y
1)和左下控制点的位置坐标(x
2,y
2)代入6参数仿射模型预测当前解码块的左上控制点的运动矢量,右上控制点的运动矢量,及左下控制点的运动矢量,如公式(4)、(5)、(6)所示。
公式(4)、(5)前面已有描述,在公式(4)、(5)、(6)中,(vx
0,vy
0)为预测的当前解码块的左上控制点的运动矢量,(vx
1,vy
1)为预测的当前解码块的右上控制点的运动矢量,(vx
2,vy
2)为预测的当前解码块的左下控制点的运动矢量。
方式二,采用基于控制点组合的运动矢量预测方法构建候选运动信息列表。
下面例举两种方案,分别表示为方案A和方案B:
方案A:将当前解码块的2个控制点的运动信息进行组合,用来构建4参数仿射变换模型。2个控制点的组合方式为{CP1,CP4},{CP2,CP3},{CP1,CP2},{CP2,CP4},{CP1,CP3},{CP3,CP4}。例如,采用CP1和CP2控制点构建的4参数仿射变换模型,记做Affine(CP1,CP2)。
需要说明的是,亦可将不同控制点的组合转换为同一位置的控制点。例如:将{CP1,CP4},{CP2,CP3},{CP2,CP4},{CP1,CP3},{CP3,CP4}组合得到的4参数仿射变换模型转换为控制点{CP1,CP2}或{CP1,CP2,CP3}来表示。转换方法为将控制点的运动矢量及其坐标信息,代入公式(9-1),得到模型参数,再将{CP1,CP2}的坐标信息代入,得到其运动矢量,作为一组候选运动矢量预测值。
在公式(9-1)中,a
1,a
2,a
3,and a
4均为参数模型中的参数,(x,y)表示位置坐标。
更直接地,也可以按照以下公式进行转换得到以左上控制点、右上控制点表示的一组运动矢量预测值,并加入候选运动信息列表:
{CP1,CP2}转换得到{CP1,CP2,CP3}的公式(9-2):
{CP1,CP3}转换得到{CP1,CP2,CP3}的公式(9-3):
{CP2,CP3}转换得到{CP1,CP2,CP3}的公式(10):
{CP1,CP4}转换得到{CP1,CP2,CP3}的公式(11):
{CP2,CP4}转换得到{CP1,CP2,CP3}的公式(12):
{CP3,CP4}转换得到{CP1,CP2,CP3}的公式(13):
方案B:将当前解码块的3个控制点的运动信息进行组合,用来构建6参数仿射变换模型。3个控制点的组合方式为{CP1,CP2,CP4},{CP1,CP2,CP3},{CP2,CP3,CP4},{CP1,CP3,CP4}。例如,采用CP1、CP2和CP3控制点构建的6参数仿射变换模型,记做Affine(CP1,CP2,CP3)。
要说明的是,亦可将不同控制点的组合转换为同一位置的控制点。例如:将{CP1,CP2,CP4},{CP2,CP3,CP4},{CP1,CP3,CP4}组合的6参数仿射变换模型转换为控制点{CP1,CP2,CP3}来表示。转换方法为将控制点的运动矢量及其坐标信息,代入公式(14),得到模型参数,再将{CP1,CP2,CP3}的坐标信息代入,得到其运动矢量,作为一组候选运动矢量预测值。
在公式(14)中,a
1,a
2,a
3,a
4,a
5,a
6为参数模型中的参数,(x,y)表示位置坐标。
更直接地,也可以按照以下公式进行转换得到以左上控制点、右上控制点、左下控制点表示的一组运动矢量预测值,并加入候选运动信息列表:
{CP1,CP2,CP4}转换得到{CP1,CP2,CP3}的公式(15):
{CP2,CP3,CP4}转换得到{CP1,CP2,CP3}的公式(16):
{CP1,CP3,CP4}转换得到{CP1,CP2,CP3}的公式(17):
需要说明的是,可仅采用方式一预测得到的候选运动矢量预测值来构建候选运动信息列表,也可仅采用方式二预测得到的候选运动矢量预测值来构建候选运动信息列表,还可采用方式一预测得到的候选运动矢量预测值和方式二预测得到的候选运动矢量预测值来共同构建候选运动信息列表。另外,还可将候选运动信息列表按照预先配置的规则进行剪枝和排序,然后将其截断或填充至特定个数。当候选运动信息列表中的每一组候选运动矢量预测值包括三个控制点的运动矢量预测值时,可称该候选运动信息列表为三元组列表;当候选运动信息列表中的每一组候选运动矢量预测值包括两个控制点的运动矢量预测值时,可称该候选运动信息列表为二元组列表。
步骤S1222:视频解码器解析码流,以得到索引。
具体地,视频解码器可以通过熵解码单元解析码流,该索引用于指示当前解码块的目标候选运动矢量组,该目标候选运动矢量表示当前解码块的一组控制点的运动矢量预测值。
步骤S1223:视频解码器根据所述索引,从候选运动信息列表中确定目标运动矢量组。
具体地,频解码器根据该索引从候选运动信息列表中确定出的目标候选运动矢量组用于作为最优候选运动矢量预测值(可选的,当候选运动信息列表的长度为1时,不需要解析码流得到索引,直接可以确定目标运动矢量组),具体来说是2个或3个控制点的最优运动矢量预测值;例如,视频解码器从码流中解析索引号,再根据索引号从候选运动信息列表中确定2个或3个控制点的最优运动矢量预测值,候选运动信息列表中每组候选运动矢量预测值各自对应有各自的索引号。
步骤S1224:视频解码器根据以上确定出的当前解码块的控制点的运动矢量值采用参数仿射变换模型获得当前解码块中每个子块的运动矢量值,其中所述子块的尺寸是基于当前图像块的预测方向而确定的。
具体地,即目标候选运动矢量组中包括的两个(左上控制点和右上控制点)或者三个控制点(例如,左上控制点、右上控制点和左下控制点)的运动矢量。对于当前解码块的每一个子块(一个子块也可以等效为一个运动补偿单元),可采用运动补偿单元中预设位置像素点的运动信息来表示该运动补偿单元内所有像素点的运动信息。假设运动补偿单元的尺寸为MxN(M小于等于当前解码块的宽度W,N小于等于当前解码块的高度H,其中M、N、W、H为正整数,通常为2的幂次方,如4、8、16、32、64、128等),则预设位置像素点可以为运动补偿单元中心点(M/2,N/2)、左上像素点(0,0),右上像素点(M-1,0),或其他位置的像素点。图7B示意了4x4的运动补偿单元,图7C示意了8x8的运动补偿单元。
运动补偿单元中心点相对于当前解码块左上顶点像素的坐标使用公式(5)计算得到,其中i为水平方向第i个运动补偿单元(从左到右),j为竖直方向第j个运动补偿单元(从上到下),(x
(i,j),y
(i,j))表示第(i,j)个运动补偿单元中心点相对于当前解码块左上控制点像素的坐标。再根据当前解码块的仿射模型类型(6参数或4参数),将(x
(i,j),y
(i,j))代入6参数仿射模型公式(25)或者将(x
(i,j),y
(i,j))代入4参数仿射模型公式(27),获得每个运动补偿单元中心点的运动信息,作为该运动补偿单元内所有像素点的运动矢量(vx
(i,j),vy
(i,j))。
可选的,当前解码块为6参数解码块时,在基于所述目标候选运动矢量组得到所述当前解码块的一个或多个子块的运动矢量时,若所述当前解码块的下边界与所述当前解码块所在的CTU的下边界重合,则所述当前解码块的左下角的子块的运动矢量为根据所述三个控制点构造的6参数仿射模型和所述当前解码块的左下角的位置坐标(0,H)计算得到,所述当前解码块的右下角的子块的运动矢量为根据所述三个控制点构造的6参数仿射模型和所述当前解码块的右下角的位置坐标(W,H)计算得到。例如,将当前解码块的左下角的位置坐标(0,H)代入到该6参数仿射模型即可得到当前解码块的左下角的子块的运动矢量(而不是将左下角的子块的中心点坐标代入该仿射模 型进行计算),将当前解码块的右下角的位置坐标(W,H)代入到该6参数仿射模型即可得到当前解码块的右下角的子块的运动矢量(而不是将右下角的子块的中心点坐标代入该仿射模型进行计算)。这样一来,该当前解码块的左下控制点的运动矢量和右下控制点的运动矢量被用到时(例如,后续其他块基于该当前块的左下控制点和右下控制点的运动矢量构建该其他块的候选运动信息列表),用到的是准确的值而不是估算值。其中,W为该当前解码块的宽,H为该当前解码块的高。
可选的,当前解码块为4参数解码块时,在基于所述目标候选运动矢量组得到所述当前解码块的一个或多个子块的运动矢量时,若所述当前解码块的下边界与所述当前解码块所在的CTU的下边界重合,则所述当前解码块的左下角的子块的运动矢量为根据所述两个控制点构造的4参数仿射模型和所述当前解码块的左下角的位置坐标(0,H)计算得到,所述当前解码块的右下角的子块的运动矢量为根据所述两个控制点构造的4参数仿射模型和所述当前解码块的右下角的位置坐标(W,H)计算得到。例如,将当前解码块的左下角的位置坐标(0,H)代入到该4参数仿射模型即可得到当前解码块的左下角的子块的运动矢量(而不是将左下角的子块的中心点坐标代入该仿射模型进行计算),将当前解码块的右下角的位置坐标(W,H)代入到该四参数仿射模型即可得到当前解码块的右下角的子块的运动矢量(而不是将右下角的子块的中心点坐标代入该仿射模型进行计算)。这样一来,该当前解码块的左下控制点的运动矢量和右下控制点的运动矢量被用到时(例如,后续其他块基于该当前块的左下控制点和右下控制点的运动矢量构建该其他块的候选运动信息列表),用到的是准确的值而不是估算值。其中,W为该当前解码块的宽,H为该当前解码块的高。
步骤S1225:该视频解码器根据该当前解码块中每个子块的运动矢量值进行运动补偿,以得到每个子块的像素预测值,具体来说,根据所述当前解码块的一个或多个子块的运动矢量,以及所述索引指示的参考帧索引和预测方向,预测得到所述当前解码块的像素预测值。
应当理解的是,以上方法流程的步骤中,步骤的描述顺序并不代表步骤的执行顺序,按照以上的描述顺序来执行是可行的,不按照以上的描述顺序来执行也是可行的。
应当理解的是,以上方法流程的步骤中,S1211和/或S1212可以是可选的步骤,例如,不涉及列表的构建,不涉及从码流中解析索引。S1221和/或S1222可以是可选的步骤,例如,不涉及列表的构建,不涉及从码流中解析索引。
图8A为本申请实施例中的图像预测装置800的一种示意性框图。需要说明的是,图像预测装置800既适用于解码视频图像的帧间预测,也适用于编码视频图像的帧间预测,应当理解的是,这里的图像预测装置800可以对应于图2A中的帧间预测模块211,或者可以对应于图2C中的运动补偿模块322,该图像预测装置800可以包括:
获取单元801,用于获取当前图像块(当前仿射图像块)的控制点的运动矢量;
帧间预测处理单元802,用于对所述当前图像块执行帧间预测,所述帧间预测过程包括:根据所述当前图像块的控制点(仿射控制点)的运动矢量(运动矢量组)采用仿射变换模型获得当前图像块中每个子块的运动矢量,其中所述子块的尺寸是基于 当前图像块的预测方向而确定的;根据该当前图像块中每个子块的运动矢量值进行运动补偿,以得到每个子块的像素预测值。
应当理解的是,当所述图像块的多个子块的像素预测值得到,相应的所述图像块的像素预测值就得到了。
在一种可行的实施方式中,帧间预测处理单元802用于根据所述当前图像块的控制点的运动矢量采用仿射变换模型获得当前图像块中每个子块的运动矢量,其中,若当前图像块为单向预测,当前图像块的子块的尺寸为4x4;或者,若当前图像块为双向预测,当前图像块的子块的尺寸为8x8;根据该当前图像块中每个子块的运动矢量进行运动补偿,以得到每个子块的像素预测值。
本申请实施例的图像预测装置中,在一些可行的实施方式中,如果所述当前图像块的预测方向为双向预测,所述当前图像块中子块的尺寸为UxV;或者,
如果所述当前图像块的预测方向为单向预测,所述当前图像块中子块的尺寸为MxN,
其中U,M表示所述子块的宽,V,N表示所述子块的高,以及U,V,M,N均为2n,n为正整数。
本申请实施例的图像预测装置中,在一些可行的实施方式中,U>=M,V>=N,且U和V不能同时等于M和N。
本申请实施例的图像预测装置中,在一些可行的实施方式中,U=2M,V=2N。例如,M为4,N为4。例如,U为8,V为8。
本申请实施例的图像预测装置中,在一些可行的实施方式中,所述获取单元801具体用于:
接收从码流中解析得到的索引和运动矢量差值MVD;
根据所述索引,从候选运动矢量预测值MVP列表中确定目标候选运动矢量预测值组;
根据目标候选运动矢量预测值组和从码流中解析出的运动矢量差值MVD确定当前图像块的控制点的运动矢量。
本申请实施例的图像预测装置中,在一些可行的实施方式中,预测方向指示信息用于指示单向预测或双向预测,其中所述预测方向指示信息是从所述码流中解析或推导得到的。
本申请实施例的图像预测装置中,在一些可行的实施方式中,所述获取单元801具体用于:
接收从码流中解析得到的索引;
根据所述索引,从候选运动信息列表中确定目标候选运动信息,其中所述目标候选运动信息包括至少一个目标候选运动矢量组,所述目标候选运动矢量组作为所述当前图像块的控制点的运动矢量。
本申请实施例的图像预测装置中,在一些可行的实施方式中,所述当前图像块的预测方向是双向预测,其中所述候选运动信息列表中与所述索引对应的目标候选运动信息包括对应于第一参考帧列表的第一目标候选运动矢量组,和,对应于第二参考帧列表的第二目标候选运动矢量组;
所述当前图像块的预测方向是单向预测,其中所述候选运动信息列表中与所述索引对应的目标候选运动信息包括:对应于第一参考帧列表的第一目标候选运动矢量组,或者,所述候选运动信息列表中与所述索引对应的目标候选运动信息包括:对应于第二参考帧列表的第二目标候选运动矢量组。
本申请实施例的图像预测装置中,在一些可行的实施方式中,当当前图像块的尺寸满足W>=16和H>=16时,允许采用仿射模式。
本申请实施例的图像预测装置中,在一些可行的实施方式中,所述帧间预测处理单元,具体用于根据所述当前图像块的控制点的运动矢量,得到仿射变换模型;根据当前图像块中每个子块的位置坐标信息以及所述仿射变换模型,获得当前图像块中每个子块的运动矢量。
由上可见,本申请实施例的图像预测装置中,考虑当前图像块的帧间方向来确定当前图像块的子块的尺寸,比如,如果当前编码图像块的预测模式为单向预测,子块的尺寸为4*4;如果当前编码图像块的预测模式为双向预测,子块的尺寸为8*8;这样的话,在运动补偿的复杂度与预测效率两者之间得到平衡,即降低现有技术中运动补偿的复杂度的同时,兼顾预测效率,从而提高编解码性能。
需要说明的是,本申请实施例的图像预测装置中的各个模块为实现本申请图像预测方法中所包含的各种执行步骤的功能主体,即具备实现完整实现本申请图像预测方法中的各个步骤以及这些步骤的扩展及变形的功能主体,具体请参见本文中对图像预测方法的介绍,为简洁起见,本文将不再赘述。
图8B为本申请实施例的编码设备或解码设备(简称为译码设备1000)的一种实现方式的示意性框图。其中,译码设备1000可以包括处理器1010、存储器1030和总线系统1050。其中,处理器和存储器通过总线系统相连,该存储器用于存储指令,该处理器用于执行该存储器存储的指令。编码设备的存储器存储程序代码,且处理器可以调用存储器中存储的程序代码执行本申请描述的各种视频编码或解码方法,尤其是本申请的图像预测方法。为避免重复,这里不再详细描述。
在本申请实施例中,该处理器1010可以是中央处理单元(Central Processing Unit,简称为“CPU”),该处理器1010还可以是其他通用处理器、数字信号处理器(DSP)、专用集成电路(ASIC)、现成可编程门阵列(FPGA)或者其他可编程逻辑器件、 分立门或者晶体管逻辑器件、分立硬件组件等。通用处理器可以是微处理器或者该处理器也可以是任何常规的处理器等。
该存储器1030可以包括只读存储器(ROM)设备或者随机存取存储器(RAM)设备。任何其他适宜类型的存储设备也可以用作存储器1030。存储器1030可以包括由处理器1010使用总线1050访问的代码和数据1031。存储器1030可以进一步包括操作系统1033和应用程序1035,该应用程序1035包括允许处理器1010执行本申请描述的视频编码或解码方法(尤其是本申请描述的编码方法或解码方法)的至少一个程序。例如,应用程序1035可以包括应用1至N,其进一步包括执行在本申请描述的视频编码或解码方法的视频编码或解码应用(简称视频译码应用)。
该总线系统1050除包括数据总线之外,还可以包括电源总线、控制总线和状态信号总线等。但是为了清楚说明起见,在图中将各种总线都标为总线系统1050。
可选的,译码设备1000还可以包括一个或多个输出设备,诸如显示器1070。在一个示例中,显示器1070可以是触感显示器,其将显示器与可操作地感测触摸输入的触感单元合并。显示器1070可以经由总线1050连接到处理器1010。
图9是根据一示例性实施例的包含图2A的编码器20和/或图2C的解码器30的视频编码系统1100的实例的说明图。系统1100可以实现本申请的各种技术的组合。在所说明的实施方式中,视频编码系统1100可以包含成像设备1101、视频编码器20、视频解码器30(和/或藉由处理单元1106的逻辑电路1107实施的视频编码器)、天线1102、一个或多个处理器1103、一个或多个存储器1104和/或显示设备1105。
如图所示,成像设备1101、天线1102、处理单元1106、逻辑电路1107、视频编码器20、视频解码器30、处理器1103、存储器1104和/或显示设备1105能够互相通信。如所论述,虽然用视频编码器20和视频解码器30绘示视频编码系统1100,但在不同实例中,视频编码系统1100可以只包含视频编码器20或只包含视频解码器30。
在一些实例中,如图所示,视频编码系统1100可以包含天线1102。例如,天线1102可以用于传输或接收视频数据的经编码比特流。另外,在一些实例中,视频编码系统1100可以包含显示设备1105。显示设备1105可以用于呈现视频数据。在一些实例中,如图所示,逻辑电路1107可以通过处理单元1106实施。处理单元1106可以包含专用集成电路(application-specific integrated circuit,ASIC)逻辑、图形处理器、通用处理器等。视频编码系统1100也可以包含可选处理器1103,该可选处理器1103类似地可以包含专用集成电路(application-specific integrated circuit,ASIC)逻辑、图形处理器、通用处理器等。在一些实例中,逻辑电路1107可以通过硬件实施,如视频编码专用硬件等,处理器1103可以通过通用软件、操作系统等实施。另外,存储器1104可以是任何类型的存储器,例如易失性存储器(例如,静态随机存取存储器(Static Random Access Memory,SRAM)、动态随机存储器(Dynamic Random Access Memory,DRAM)等)或非易失性存储器(例如,闪存等)等。在非限制性实例中,存储器1104可以由超速缓存内存实施。在一些实例中,逻辑电路1107可以访问存储器1104(例如用于实施图像缓冲器)。在其它实例中,逻辑电路1107和 /或处理单元1106可以包含存储器(例如,缓存等)用于实施图像缓冲器等。
在一些实例中,通过逻辑电路实施的视频编码器20可以包含(例如,通过处理单元1106或存储器1104实施的)图像缓冲器和(例如,通过处理单元1106实施的)图形处理单元。图形处理单元可以通信耦合至图像缓冲器。图形处理单元可以包含通过逻辑电路1107实施的视频编码器20,以实施参照图2A和/或本文中所描述的任何其它编码器系统或子系统所论述的各种模块。逻辑电路可以用于执行本文所论述的各种操作。
视频解码器30可以以类似方式通过逻辑电路1107实施,以实施参照图2B的解码器200和/或本文中所描述的任何其它解码器系统或子系统所论述的各种模块。在一些实例中,逻辑电路实施的视频解码器30可以包含(通过处理单元1106或存储器1104实施的)图像缓冲器和(例如,通过处理单元1106实施的)图形处理单元。图形处理单元可以通信耦合至图像缓冲器。图形处理单元可以包含通过逻辑电路1107实施的视频解码器30,以实施参照图2B和/或本文中所描述的任何其它解码器系统或子系统所论述的各种模块。
在一些实例中,视频编码系统1100的天线1102可以用于接收视频数据的经编码比特流。如所论述,经编码比特流可以包含本文所论述的与编码视频帧相关的数据、指示符、索引值、模式选择数据等,例如与编码分割相关的数据(例如,变换系数或经量化变换系数,(如所论述的)可选指示符,和/或定义编码分割的数据)。视频编码系统1100还可包含耦合至天线1102并用于解码经编码比特流的视频解码器30。显示设备1105用于呈现视频帧。
以上方法流程的步骤中,步骤的描述顺序并不代表步骤的执行顺序,按照以上的描述顺序来执行是可行的,不按照以上的描述顺序来执行也是可行的。例如上述步骤S1211可以在步骤S1212之后执行,也可以在步骤S1212之前执行;上述步骤S1221可以在步骤S1222之后执行,也可以在步骤S1222之前执行;其余步骤此处不再一一举例。
本领域技术人员能够领会,结合本文公开描述的各种说明性逻辑框、模块和算法步骤所描述的功能可以硬件、软件、固件或其任何组合来实施。如果以软件来实施,那么各种说明性逻辑框、模块、和步骤描述的功能可作为一或多个指令或代码在计算机可读媒体上存储或传输,且由基于硬件的处理单元执行。计算机可读媒体可包含计算机可读存储媒体,其对应于有形媒体,例如数据存储媒体,或包括任何促进将计算机程序从一处传送到另一处的媒体(例如,根据通信协议)的通信媒体。以此方式,计算机可读媒体大体上可对应于(1)非暂时性的有形计算机可读存储媒体,或(2)通信媒体,例如信号或载波。数据存储媒体可为可由一或多个计算机或一或多个处理器存取以检索用于实施本申请中描述的技术的指令、代码和/或数据结构的任何可用媒体。计算机程序产品可包含计算机可读媒体。
作为实例而非限制,此类计算机可读存储媒体可包括RAM、ROM、EEPROM、CD-ROM或其它光盘存储装置、磁盘存储装置或其它磁性存储装置、快闪存储器或可用来存储指令或数据结构的形式的所要程序代码并且可由计算机存取的任何其它媒体。并且, 任何连接被恰当地称作计算机可读媒体。举例来说,如果使用同轴缆线、光纤缆线、双绞线、数字订户线(DSL)或例如红外线、无线电和微波等无线技术从网站、服务器或其它远程源传输指令,那么同轴缆线、光纤缆线、双绞线、DSL或例如红外线、无线电和微波等无线技术包含在媒体的定义中。但是,应理解,所述计算机可读存储媒体和数据存储媒体并不包括连接、载波、信号或其它暂时媒体,而是实际上针对于非暂时性有形存储媒体。如本文中所使用,磁盘和光盘包含压缩光盘(CD)、激光光盘、光学光盘、数字多功能光盘(DVD)和蓝光光盘,其中磁盘通常以磁性方式再现数据,而光盘利用激光以光学方式再现数据。以上各项的组合也应包含在计算机可读媒体的范围内。
可通过例如一或多个数字信号处理器(DSP)、通用微处理器、专用集成电路(ASIC)、现场可编程逻辑阵列(FPGA)或其它等效集成或离散逻辑电路等一或多个处理器来执行指令。因此,如本文中所使用的术语“处理器”可指前述结构或适合于实施本文中所描述的技术的任一其它结构中的任一者。另外,在一些方面中,本文中所描述的各种说明性逻辑框、模块、和步骤所描述的功能可以提供于经配置以用于编码和解码的专用硬件和/或软件模块内,或者并入在组合编解码器中。而且,所述技术可完全实施于一或多个电路或逻辑元件中。
本申请的技术可在各种各样的装置或设备中实施,包含无线手持机、集成电路(IC)或一组IC(例如,芯片组)。本申请中描述各种组件、模块或单元是为了强调用于执行所揭示的技术的装置的功能方面,但未必需要由不同硬件单元实现。实际上,如上文所描述,各种单元可结合合适的软件和/或固件组合在编码解码器硬件单元中,或者通过互操作硬件单元(包含如上文所描述的一或多个处理器)来提供。
以上所述,仅为本申请示例性的具体实施方式,但本申请的保护范围并不局限于此,任何熟悉本技术领域的技术人员在本申请揭露的技术范围内,可轻易想到的变化或替换,都应涵盖在本申请的保护范围之内。因此,本申请的保护范围应该以权利要求的保护范围为准。
Claims (29)
- 一种图像预测方法,其特征在于,包括:获取当前图像块的控制点的运动矢量;根据所述当前图像块的控制点的运动矢量采用仿射变换模型获得当前图像块中每个子块的运动矢量,其中所述子块的尺寸是基于当前图像块的预测方向而确定的;根据该当前图像块中每个子块的运动矢量值进行运动补偿,以得到每个子块的像素预测值。
- 根据权利要求1所述的方法,其特征在于,如果所述当前图像块的预测方向为双向预测,所述当前图像块中子块的尺寸为UxV;或者,如果所述当前图像块的预测方向为单向预测,所述当前图像块中子块的尺寸为MxN,其中U,M表示所述子块的宽,V,N表示所述子块的高,以及U,V,M,N均为2 n,n为正整数。
- 根据权利要求2所述的方法,其特征在于,U>=M,V>=N,且U和V不能同时等于M和N。
- 根据权利要求2或3所述的方法,其特征在于,U=2M,V=2N。
- 根据权利要求2或3或4所述的方法,其特征在于,M为4,N为4。
- 根据权利要求2或3或4或5所述的方法,其特征在于,U为8,V为8。
- 根据权利要求1至6任一项所述的方法,所述获取当前图像块的控制点的运动矢量,包括:接收从码流中解析得到的索引和运动矢量差值MVD;根据所述索引,从候选运动矢量预测值MVP列表中确定目标候选运动矢量组;根据目标候选运动矢量组和从码流中解析出的运动矢量差值MVD确定当前图像块的控制点的运动矢量。
- 根据权利要求7所述的方法,其特征在于,预测方向指示信息用于指示单向预测或双向预测,其中所述预测方向指示信息是从所述码流中解析或推导得到的。
- 根据权利要求1至6任一项所述的方法,所述获取当前图像块的控制点的运动矢量,包括:接收从码流中解析得到的索引;根据所述索引,从候选运动信息列表中确定目标候选运动信息,其中所述目标候选运动信息包括至少一个目标候选运动矢量组,所述目标候选运动矢量组作为所述当前图像块的控制点的运动矢量。
- 根据权利要求9所述的方法,其特征在于,所述当前图像块的预测方向是双向预测,其中所述候选运动信息列表中与所述索引对应的目标候选运动信息包括对应于第一参考帧列表的第一目标候选运动矢量组,和,对应于第二参考帧列表的第二目标候选运动矢量组;或者,所述当前图像块的预测方向是单向预测,其中所述候选运动信息列表中与所述索引对应的目标候选运动信息包括:对应于第一参考帧列表的第一目标候选运动矢量组,或者,所述候选运动信息列表中与所述索引对应的目标候选运动信息包括:对应于第二参考帧列表的第二目标候选运动矢量组。
- 根据权利要求1至10任一项所述的方法,其特征在于,当当前图像块的尺寸满足W>=16和H>=16时,允许采用仿射模式。
- 根据权利要求1至11任一项所述的方法,其特征在于,所述根据所述获取的当前图像块的控制点的运动矢量值采用仿射变换模型获得当前图像块中每个子块的运动矢量,包括:根据所述当前图像块的控制点的运动矢量,得到仿射变换模型;根据当前图像块中每个子块的位置坐标信息以及所述仿射变换模型,获得当前图像块中每个子块的运动矢量。
- 一种图像预测方法,其特征在于,包括:获取当前图像块的控制点的运动矢量;根据所述当前图像块的控制点的运动矢量采用仿射变换模型获得当前图像块中每个子块的运动矢量,其中,若当前图像块为单向预测,当前图像块的子块的尺寸为4x4;或者,若当前图像块为双向预测,当前图像块的子块的尺寸为8x8。根据该当前图像块中每个子块的运动矢量值进行运动补偿,以得到每个子块的像素预测值。
- 一种图像预测装置,其特征在于,包括:获取单元,用于获取当前图像块的控制点的运动矢量;帧间预测处理单元,用于根据所述当前图像块的控制点的运动矢量采用仿射变换模型获得当前图像块中每个子块的运动矢量,其中所述子块的尺寸是基于当前图像块的预测方向而确定的;根据该当前图像块中每个子块的运动矢量进行运动补偿,以得到每个子块的像素预测值。
- 根据权利要求14所述的装置,其特征在于,如果所述当前图像块的预测方向为双向预测,所述当前图像块中子块的尺寸为UxV;或者,如果所述当前图像块的预测方向为单向预测,所述当前图像块中子块的尺寸为MxN,其中U,M表示所述子块的宽,V,N表示所述子块的高,以及U,V,M,N均为2 n,n为正整数。
- 根据权利要求15所述的装置,其特征在于,U>=M,V>=N,且U和V不能同时等于M和N。
- 根据权利要求15或16所述的装置,其特征在于,U=2M,V=2N。
- 根据权利要求15或16或17所述的装置,其特征在于,M为4,N为4。
- 根据权利要求15或16或17或18所述的装置,其特征在于,U为8,V为8。
- 根据权利要求14至19任一项所述的装置,所述获取单元具体用于: 接收从码流中解析得到的索引和运动矢量差值MVD;根据所述索引,从候选运动矢量预测值MVP列表中确定目标候选运动矢量预测值组;根据目标候选运动矢量预测值组和从码流中解析出的运动矢量差值MVD确定当前图像块的控制点的运动矢量。
- 根据权利要求20所述的装置,其特征在于,预测方向指示信息用于指示单向预测或双向预测,其中所述预测方向指示信息是从所述码流中解析或推导得到的。
- 根据权利要求14至19任一项所述的装置,所述获取单元具体用于:接收从码流中解析得到的索引;根据所述索引,从候选运动信息列表中确定目标候选运动信息,其中所述目标候选运动信息包括至少一个目标候选运动矢量组,所述目标候选运动矢量组作为所述当前图像块的控制点的运动矢量。
- 根据权利要求22所述的装置,其特征在于,所述当前图像块的预测方向是双向预测,其中所述候选运动信息列表中与所述索引对应的目标候选运动信息包括对应于第一参考帧列表的第一目标候选运动矢量组,和,对应于第二参考帧列表的第二目标候选运动矢量组;所述当前图像块的预测方向是单向预测,其中所述候选运动信息列表中与所述索引对应的目标候选运动信息包括:对应于第一参考帧列表的第一目标候选运动矢量组,或者,所述候选运动信息列表中与所述索引对应的目标候选运动信息包括:对应于第二参考帧列表的第二目标候选运动矢量组。
- 根据权利要求14至23任一项所述的装置,其特征在于,当当前图像块的尺寸满足W>=16和H>=16时,允许采用仿射模式。
- 根据权利要求14至24任一项所述的装置,其特征在于,所述帧间预测处理单元,具体用于根据所述当前图像块的控制点的运动矢量,得到仿射变换模型;根据当前图像块中每个子块的位置坐标信息以及所述仿射变换模型,获得当前图像块中每个子块的运动矢量。
- 一种图像预测装置,其特征在于,包括:获取单元,用于获取当前图像块的控制点的运动矢量;帧间预测处理单元,用于根据所述当前图像块的控制点的运动矢量采用仿射变换模型获得当前图像块中每个子块的运动矢量,其中,若当前图像块为单向预测,当前图像块的子块的尺寸为4x4;或者,若当前图像块为双向预测,当前图像块的子块的尺寸为8x8;根据该当前图像块中每个子块的运动矢量进行运动补偿,以得到每个子块的像素预测值。
- 一种视频编码器,其特征在于,包括:如权利要求14至26任一项所述的图像预测装置,其中所述图像预测装置用于基于当前图像块中每个子块的像素预测值,得到当前图像块的预测像素值;重建模块,用于根据所述当前图像块的预测像素值重建所述当前图像块。
- 一种视频解码器,其特征在于,包括:如权利要求14至26任一项所述的图像预测装置,其中所述图像预测装置用于基于当前图像块中每个子块的像素预测值,得到当前图像块的预测像素值;重建模块,用于根据所述当前图像块的预测像素值重建所述当前图像块。
- 一种视频解码设备,包括:相互耦合的非易失性存储器和处理器,所述处理器调用存储在所述存储器中的程序代码以执行如权利要求1-13任一项所描述的方法。
Priority Applications (3)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN201980062412.9A CN112740663B (zh) | 2018-09-24 | 2019-09-24 | 图像预测方法、装置以及相应的编码器和解码器 |
| EP19866037.5A EP3855733A4 (en) | 2018-09-24 | 2019-09-24 | IMAGE PREDICTION PROCESS AND DEVICE, AS WELL AS CORRESPONDING ENCODER AND DECODER |
| US17/209,962 US20210211715A1 (en) | 2018-09-24 | 2021-03-23 | Picture Prediction Method and Apparatus, and Corresponding Encoder and Decoder |
Applications Claiming Priority (4)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US201862735856P | 2018-09-24 | 2018-09-24 | |
| US62/735,856 | 2018-09-24 | ||
| US201862736458P | 2018-09-25 | 2018-09-25 | |
| US62/736,458 | 2018-09-25 |
Related Child Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| US17/209,962 Continuation US20210211715A1 (en) | 2018-09-24 | 2021-03-23 | Picture Prediction Method and Apparatus, and Corresponding Encoder and Decoder |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2020063599A1 true WO2020063599A1 (zh) | 2020-04-02 |
Family
ID=69952473
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2019/107614 Ceased WO2020063599A1 (zh) | 2018-09-24 | 2019-09-24 | 图像预测方法、装置以及相应的编码器和解码器 |
Country Status (4)
| Country | Link |
|---|---|
| US (1) | US20210211715A1 (zh) |
| EP (1) | EP3855733A4 (zh) |
| CN (1) | CN112740663B (zh) |
| WO (1) | WO2020063599A1 (zh) |
Families Citing this family (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| KR102672759B1 (ko) * | 2017-09-28 | 2024-06-05 | 삼성전자주식회사 | 부호화 방법 및 그 장치, 복호화 방법 및 그 장치 |
| EP4742674A2 (en) * | 2018-06-29 | 2026-05-13 | InterDigital VC Holdings, Inc. | Adaptive control point selection for affine motion model based video coding |
| US12003744B2 (en) * | 2021-10-25 | 2024-06-04 | Sensormatic Electronics, LLC | Hierarchical surveilance video compression repository |
Citations (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN103828373A (zh) * | 2011-10-05 | 2014-05-28 | 松下电器产业株式会社 | 图像编码方法、图像编码装置、图像解码方法、图像解码装置及图像编解码装置 |
| CN104363451A (zh) * | 2014-10-27 | 2015-02-18 | 华为技术有限公司 | 图像预测方法及相关装置 |
| CN104980762A (zh) * | 2012-02-08 | 2015-10-14 | 高通股份有限公司 | B切片中的预测单元限于单向帧间预测 |
| CN107113424A (zh) * | 2014-11-18 | 2017-08-29 | 联发科技股份有限公司 | 基于来自单向预测的运动矢量和合并候选的双向预测视频编码方法 |
Family Cites Families (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN106303543B (zh) * | 2015-05-15 | 2018-10-30 | 华为技术有限公司 | 视频图像编码和解码的方法、编码设备和解码设备 |
| US10681370B2 (en) * | 2016-12-29 | 2020-06-09 | Qualcomm Incorporated | Motion vector generation for affine motion model for video coding |
-
2019
- 2019-09-24 EP EP19866037.5A patent/EP3855733A4/en not_active Withdrawn
- 2019-09-24 CN CN201980062412.9A patent/CN112740663B/zh active Active
- 2019-09-24 WO PCT/CN2019/107614 patent/WO2020063599A1/zh not_active Ceased
-
2021
- 2021-03-23 US US17/209,962 patent/US20210211715A1/en not_active Abandoned
Patent Citations (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN103828373A (zh) * | 2011-10-05 | 2014-05-28 | 松下电器产业株式会社 | 图像编码方法、图像编码装置、图像解码方法、图像解码装置及图像编解码装置 |
| CN104980762A (zh) * | 2012-02-08 | 2015-10-14 | 高通股份有限公司 | B切片中的预测单元限于单向帧间预测 |
| CN104363451A (zh) * | 2014-10-27 | 2015-02-18 | 华为技术有限公司 | 图像预测方法及相关装置 |
| CN107113424A (zh) * | 2014-11-18 | 2017-08-29 | 联发科技股份有限公司 | 基于来自单向预测的运动矢量和合并候选的双向预测视频编码方法 |
Non-Patent Citations (2)
| Title |
|---|
| See also references of EP3855733A4 * |
| WANG, YANG: "CE4.2.12 Affine merge mode", JOINT VIDEO EXPERTS TEAM (JVET) OF ITU-T SG 16 WP 3 AND ISO/IEC JTC 1/ SC 29/WG 11 11TH MEETING, no. JVET-K0355-v1, 10 July 2018 (2018-07-10), Ljubljana, SI, XP030198925 * |
Also Published As
| Publication number | Publication date |
|---|---|
| US20210211715A1 (en) | 2021-07-08 |
| CN112740663A (zh) | 2021-04-30 |
| CN112740663B (zh) | 2022-06-14 |
| EP3855733A1 (en) | 2021-07-28 |
| EP3855733A4 (en) | 2021-12-08 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| JP7712977B2 (ja) | 画像予測の方法および装置、ならびにコーデック | |
| JP7693895B2 (ja) | ビデオ復号化方法及びビデオ・デコーダ | |
| TWI759389B (zh) | 用於視訊寫碼之低複雜度符號預測 | |
| JP7279154B2 (ja) | アフィン動きモデルに基づく動きベクトル予測方法および装置 | |
| JP7743595B2 (ja) | 動きベクトル予測方法及び関連する装置 | |
| US20220345739A1 (en) | Video picture prediction method and apparatus | |
| KR102494762B1 (ko) | 비디오 데이터 인터 예측 방법 및 장치 | |
| TW201924345A (zh) | 寫碼用於視頻寫碼之仿射預測移動資訊 | |
| WO2019154424A1 (zh) | 视频解码方法、视频解码器以及电子设备 | |
| CN112740663B (zh) | 图像预测方法、装置以及相应的编码器和解码器 | |
| KR20240100392A (ko) | 비디오 코딩에서 아핀 병합 모드에 대한 후보 도출 | |
| KR102566569B1 (ko) | 인터 예측 방법 및 장치, 비디오 인코더 및 비디오 디코더 | |
| RU2783337C2 (ru) | Способ декодирования видео и видеодекодер |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 19866037 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| ENP | Entry into the national phase |
Ref document number: 2019866037 Country of ref document: EP Effective date: 20210422 |







































