WO1999041912A2 - Method and arrangement for video coding - Google Patents

Method and arrangement for video coding Download PDF

Info

Publication number
WO1999041912A2
WO1999041912A2 PCT/IB1999/000151 IB9900151W WO9941912A2 WO 1999041912 A2 WO1999041912 A2 WO 1999041912A2 IB 9900151 W IB9900151 W IB 9900151W WO 9941912 A2 WO9941912 A2 WO 9941912A2
Authority
WO
WIPO (PCT)
Prior art keywords
block
coefficients
prediction block
candidate prediction
candidate
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/IB1999/000151
Other languages
French (fr)
Other versions
WO1999041912A3 (en
Inventor
Richard P. Kleihorst
Fabrice Cabrera
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Koninklijke Philips NV
Philips AB
Philips Svenska AB
Original Assignee
Koninklijke Philips Electronics NV
Philips AB
Philips Svenska AB
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Koninklijke Philips Electronics NV, Philips AB, Philips Svenska AB filed Critical Koninklijke Philips Electronics NV
Priority to DE69901525T priority Critical patent/DE69901525T2/en
Priority to JP54124199A priority patent/JP2001520846A/en
Priority to EP99900596A priority patent/EP0976251B1/en
Priority to KR1019997009376A priority patent/KR100584495B1/en
Publication of WO1999041912A2 publication Critical patent/WO1999041912A2/en
Publication of WO1999041912A3 publication Critical patent/WO1999041912A3/en
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Classifications

    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N19/00Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
    • H04N19/60Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using transform coding
    • H04N19/61Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using transform coding in combination with predictive coding
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N19/00Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
    • H04N19/50Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding
    • H04N19/503Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding involving temporal prediction
    • H04N19/51Motion estimation or motion compensation

Definitions

  • the invention relates to a method of coding video pictures, comprising the steps of transforming picture blocks of a video picture into transform coefficients, and encoding the difference between the transform coefficients of each transformed picture block and corresponding transform coefficients of a prediction block which is obtained by applying motion estimation to a previously encoded picture.
  • the invention also relates to an arrangement for carrying out the method.
  • Video encoders operating in accordance with international coding standards such as MPEG and H.261 are hybrid encoders, in which transform coding and predictive coding are combined.
  • the predictive coding which includes motion estimation and compensation, is usually carried out in the pixel domain, and the transform coding (DCT) is applied to the predictively coded signal.
  • DCT transform coding
  • an inverse DCT is required in the prediction loop.
  • research effort is being spent on carrying out the predictive coding in the transform domain.
  • the picture is transformed first and the predictive coding is then applied to the transform coefficients.
  • the inverse DCT can now be dispensed with.
  • a drawback of the prior-art hybrid video encoder is that motion estimation and compensation in the transform domain are more complex than in the pixel domain because the calculation of a prediction block involves a large number of matrix multiplications.
  • motion estimation is a severe problem.
  • Motion estimation comprises the steps of calculating a plurality of candidate prediction blocks, comparing each candidate prediction block with the transformed picture block, and selecting one of the candidate prediction blocks to be the prediction block. With each plurality of candidate prediction blocks, said large number of calculations is involved.
  • the method in accordance with the invention is characterized in that the steps of calculating a candidate prediction block and comparing the candidate prediction block with the transformed picture block are carried out for a subset of the transform coefficients of said candidate prediction block and transformed picture block.
  • the invention is based on the recognition that the majority of bits for encoding a picture block is usually spent on a few transform coefficients and that it is beneficial to restrict the search for the most-resembling prediction block to these coefficients.
  • the number of multiplications is considerably reduced thereby, in particular in coding standards such as MPEG where motion compensation is applied to 16*16 macroblocks comprising four 8*8 luminance blocks. Note that the eventually selected prediction block for the motion compensation will still be calculated completely. Picture quality is thus not affected.
  • the subset of transform coefficients comprises the DC coefficient and a predetermined number of adjacent AC coefficients, for example, the coefficients constituting a 2*2 or 3*3 submatrix in the upper left corner of an 8*8 block.
  • the submatrix may extend in the horizontal direction (for example, 2*3 or 2*4) because most of the motion in normal video scenes is found to be horizontal. The number of multiplications is even further reduced when only coefficients of a single row or single column are calculated.
  • the coefficients are chosen in dependence upon the represented by the candidate prediction block and expressed by its associated candidate motion vector. If the candidate motion vector represents horizontal motion, a submatrix is chosen which substantially extends in the horizontal direction. If the candidate motion vector represents vertical motion, the submatrix chosen substantially extends in the vertical direction.
  • Fig. 1 shows a hybrid encoder carrying out motion estimation in the transform domain in accordance with the invention.
  • Fig. 2 shows two phases of a moving object in the pixel domain.
  • Fig. 3 shows an example of the relative position of a prediction block with respect to previously encoded blocks in the transform domain.
  • Figs. 4A-4C and 5A-5D show examples of subsets of coefficients selected for carrying out motion estimation.
  • Fig. 6 shows a flow chart illustrating the operations of a motion estimation and compensation circuit which is shown in Fig. 1.
  • Fig. 7 shows an example of the relative position of a prediction macroblock block with respect to previously encoded blocks in the transform domain.
  • Fig. 1 shows a hybrid encoder carrying out motion estimation in the transform domain in accordance with the invention.
  • the encoder receives a video picture signal in the form of picture blocks X of 8*8 pixels. Each pixel block is transformed by a discrete cosine transform circuit 1 into a block Y of 8*8 DCT coefficients.
  • a subtracter 2 the coefficients of a prediction block ⁇ are subtracted from the corresponding coefficients of block Y to obtain a block of 8*8 difference coefficients.
  • the difference coefficients are quantized in a quantizer 3 and the quantized difference coefficients are then encoded into variable-length code words by a variable-length encoder 4.
  • the variable-length code words constitute the digital output bitstream of the hybrid encoder. This bitstream is applied to a transmission channel or stored on a storage medium (not shown).
  • the quantized difference coefficients are subjected to inverse quantization in an inverse quantizer 5, added to the current prediction block by an adder 6, and written into in a frame memory 7 to update the previously encoded picture stored therein.
  • the picture blocks stored in frame memory 6 are denoted Z. Note that they are stored in the transform domain.
  • a motion estimator and motion compensator circuit 8 receives the current picture block and searches in the frame memory 7 the prediction block
  • the position of the prediction block thus found is identified by a motion vector mv which defines the position of the prediction block with respect to the current block position.
  • the motion vector mv is applied to the variable-length encoder 4 and is part of the output bitstream.
  • FIG. 2 shows two phases of a moving object 10 in the pixel domain.
  • a picture segment 11 comprising twelve 8*8 pixel blocks of the picture to be coded and a corresponding picture segment 12 of the previously encoded picture stored in a frame memory are shown.
  • Reference numeral 13 denotes an 8*8 pixel block X to be coded and reference numeral 14 denotes the most- resembling prediction block X in the frame memory.
  • the motion vector mv associated with the prediction block X is also shown. Once this motion vector has been found by motion estimation, the prediction block can easily be obtained by reading the relevant pixels of prediction block X from the frame memory.
  • the hybrid decoder shown in Fig. 1 stores the previously encoded picture in the transform domain.
  • the coefficients of the prediction block ⁇ must now be calculated from coefficient blocks Z stored in the memory. This is shown by
  • Fig. 3 which shows that the coefficients of a prediction block ⁇ are to be calculated from four underlying coefficient blocks Z,, Z 2 , Z 3 and Z 4 .
  • the prediction block follows from the equation:
  • H, and W are transform matrices representing the relative position of the block ⁇
  • H, and W have one of the formats where I is an identity matrix and 0 is a matrix of zeroes.
  • I is an identity matrix
  • 0 is a matrix of zeroes.
  • motion estimation in the transform domain requires equation (1) to be calculated several times. This requires a large number of multiplications.
  • the calculation of each candidate prediction block in accordance with equation (1) is carried out for a restricted number of coefficients only. This reduces the number of multiplications so dramatically that a one-chip MPEG2 encoder operating in the transform domain becomes feasible. Needless to say that this a great step in the field of digital video processing.
  • the coefficients to be calculated for the purpose of motion estimation can be fixedly chosen. Preferably, they include the DC coefficient and a few AC coefficients representing lower spatial frequencies. That is because these coefficients are responsible for the majority of bits in the output signal and are thus most worthwhile to be minimized in the difference block (output of subtracter 2 in Fig. 1).
  • Figs. 4A-4C show examples of preferred choices. In each example, the upper-left coefficient is the DC coefficient. In Fig. 4A, the coefficients constitute a 3*3 (or any other size) submatrix. In Fig. 4B, the selected coefficients are the first coefficients of a zigzag scanned series. This choice is advantageous in view of already available circuits for zigzag scanning coefficient blocks. In Fig. 4C, the submatrix extends in the horizontal direction. The latter choice benefits from the property that most of the motion found in plain video scenes is horizontal motion.
  • the coefficients to be calculated for the purpose of motion estimation are adaptively chosen in dependence upon the candidate motion vectors which are being evaluated. If the candidate motion vectors represent horizontal motion (which in the 3DRS algorithm is the case if substantially horizontal motion has been found previously), a submatrix is chosen which extends in the horizontal direction. Examples thereof are shown in Figs. 5 A and 5B. If the candidate motion vectors represent vertical motion, the submatrix extends in the vertical direction as shown in Figs. 5C and 5D. The examples shown in Figs.
  • 5B and 5D are particularly advantageous because the calculation on of the coefficients of a single row or column drastically contributes to saving the number of multiplications (8 coefficients constituting a single row can more economically be computed than 8 coefficients constituting a 2*4 submatrix). It is further noted that the AC coefficients chosen are not necessarily contiguous to the DC coefficient. Because small amounts of motion become manifest as higher spatial frequencies in the difference block (output of subtracter 2 in Fig. 1), it may also be advantageous to evaluate AC coefficients representing higher spatial frequencies for small amounts of motion.
  • the motion estimation and motion compensation circuit 8 (Fig. 1) will now be described. Because it has to successively carry out a plurality of identical operations each involving a number of calculations, the circuit is preferably implemented as a microprocessor-controlled arithmetic unit.
  • the architecture of such a unit is well-known to those skilled in the art and therefore not shown in more detail.
  • the processor determines which subset of DCT coefficients are taken into account for the motion estimation process.
  • the set may be fixedly chosen (cf. Figs. 4A-4C) or adaptively selected in dependence upon the candidate motion vectors (cf. Figs. 5A-5D).
  • the microprocessor selects one of the candidate motion vectors mv k . Then, in a step 22, the processor retrieves the four coefficient blocks underlying the candidate prediction block associated with this motion vector from the frame memory. In accordance with equation (1), said four blocks are denoted Z,..Z 4 .
  • a step 23 the relevant coefficients of the candidate prediction block are calculated in accordance with equation (1).
  • the coefficients are denoted y u v .
  • the microprocessor compares the set of coefficients y u v with the corresponding coefficients y u v of the current input block.
  • the result thereof is stored as a resemblance indicator R k for the current candidate motion vector mv k .
  • the resemblance indicator is indicative of the mean absolute difference between input block and candidate prediction block in the transform domain:
  • w u v is a weighting factor indicating the relative relevance of the respective coefficient.
  • the quantizer matrix provides appropriate weighting factors.
  • all weighting factors w u v may be set to one.
  • the motion vector mv k of the best-matching prediction block just found is again selected in a step 30.
  • the four underlying blocks X ..7 ⁇ are re-read from the frame memory.
  • all coefficients of the prediction block Y are calculated in accordance with equation (1). Note the difference of this step 32 with the alleged similar step 23 in the motion estimation process.
  • all the coefficients are calculated whereas the step 23 applies to a few coefficients only.
  • the block ⁇ is applied to the subtracter 2 (see Fig. 1) of the hybrid encoder.
  • the candidate prediction blocks have the same size as the underlying previously encoded blocks.
  • H.263 provide an option ("advanced motion estimation") to transmit a motion vector for each 8*8 block.
  • the described embodiment can advantageously be implemented in a hybrid encoder in accordance with such a standard.
  • Other coding standards such as MPEG, transmit one motion vector for each macroblock which includes four 8*8 blocks.
  • the process of motion estimation and compensation is slightly different, as will now be explained with reference to Fig. 7.
  • Fig. 7 shows a (candidate) prediction macroblock 40 comprising four 8*8 blocks ⁇ ⁇ -Y .
  • the macroblock now covers nine previously encoded coefficient blocks Z r Z 9 .
  • each 8*8 block is individually calculated from its four underlying blocks in accordance with equation (1).
  • the motion estimation and compensation process further proceeds as described above with reference to Fig. 6.
  • all the coefficients (a total of 4*64 in total) are calculated (step 32).
  • For the purpose of motion estimation only the selected coefficients are calculated (step 23), and the corresponding coefficients are taken from the input macroblock for determining the resemblance indicator (step 25).
  • a hybrid video encoder which carries out motion estimation and compensation (8) in the transform domain.
  • the calculation of a prediction block ( ⁇ ) from previously encoded blocks (Z) stored in the transform-domain frame memory (7) requires a large number of multiplications. This applies particularly to the motion estimation algorithm.
  • only a few DCT coefficients of candidate prediction blocks are calculated, for example, the DC coefficient and some AC coefficients.
  • the AC coefficients are adaptively selected in dependence upon the motion vector which is being considered.
  • the calculated coefficients of the candidate prediction blocks and the corresponding coefficients of the current input picture block (Y) are then compared to identify the best-matching prediction block.

Landscapes

  • Engineering & Computer Science (AREA)
  • Multimedia (AREA)
  • Signal Processing (AREA)
  • Compression Or Coding Systems Of Tv Signals (AREA)
  • Compression, Expansion, Code Conversion, And Decoders (AREA)

Abstract

A hybrid video encoder which carries out motion estimation and compensation (8) in the transform domain. In such an encoder, the calculation of a prediction block (Ŷ) from previously encoded blocks (Z) stored in the transform-domain frame memory (7) requires a large number of multiplications. This applies particularly to the motion estimation algorithm. In accordance with the invention, only a few DCT coefficients of candidate prediction blocks are calculated, for example, the DC coefficient and some AC coefficients. In a preferred embodiment, the AC coefficients are adaptively selected in dependence upon the motion vector which is being considered. The calculated coefficients of the candidate prediction blocks and the corresponding coefficients of the current input picture block (Y) are then compared to identify the best-matching prediction block.

Description

Method and arrangement for video coding.
FIELD OF THE INVENTION
The invention relates to a method of coding video pictures, comprising the steps of transforming picture blocks of a video picture into transform coefficients, and encoding the difference between the transform coefficients of each transformed picture block and corresponding transform coefficients of a prediction block which is obtained by applying motion estimation to a previously encoded picture. The invention also relates to an arrangement for carrying out the method.
BACKGROUND OF THE INVENTION
Video encoders operating in accordance with international coding standards such as MPEG and H.261 are hybrid encoders, in which transform coding and predictive coding are combined. The predictive coding, which includes motion estimation and compensation, is usually carried out in the pixel domain, and the transform coding (DCT) is applied to the predictively coded signal. In such encoders, an inverse DCT is required in the prediction loop. Nowadays, research effort is being spent on carrying out the predictive coding in the transform domain. In such an embodiment of the hybrid encoder, the picture is transformed first and the predictive coding is then applied to the transform coefficients. The inverse DCT can now be dispensed with. A schematic diagram of such a hybrid video encoder has been published by the University of Maryland on the Internet (http : //dspserv . eng . umd . edu : 80.. \ ~ koc\/project/Video_Coding . html) .
A drawback of the prior-art hybrid video encoder is that motion estimation and compensation in the transform domain are more complex than in the pixel domain because the calculation of a prediction block involves a large number of matrix multiplications. In particular, motion estimation is a severe problem. Motion estimation comprises the steps of calculating a plurality of candidate prediction blocks, comparing each candidate prediction block with the transformed picture block, and selecting one of the candidate prediction blocks to be the prediction block. With each plurality of candidate prediction blocks, said large number of calculations is involved. OBJECT AND SUMMARY OF THE INVENTION
It is an object of the invention to provide a method of video encoding, alleviating the above mentioned drawbacks.
To achieve these and other objects, the method in accordance with the invention is characterized in that the steps of calculating a candidate prediction block and comparing the candidate prediction block with the transformed picture block are carried out for a subset of the transform coefficients of said candidate prediction block and transformed picture block.
The invention is based on the recognition that the majority of bits for encoding a picture block is usually spent on a few transform coefficients and that it is beneficial to restrict the search for the most-resembling prediction block to these coefficients. The number of multiplications is considerably reduced thereby, in particular in coding standards such as MPEG where motion compensation is applied to 16*16 macroblocks comprising four 8*8 luminance blocks. Note that the eventually selected prediction block for the motion compensation will still be calculated completely. Picture quality is thus not affected.
Advantageously, the subset of transform coefficients comprises the DC coefficient and a predetermined number of adjacent AC coefficients, for example, the coefficients constituting a 2*2 or 3*3 submatrix in the upper left corner of an 8*8 block. The submatrix may extend in the horizontal direction (for example, 2*3 or 2*4) because most of the motion in normal video scenes is found to be horizontal. The number of multiplications is even further reduced when only coefficients of a single row or single column are calculated.
In a preferred embodiment, the coefficients are chosen in dependence upon the represented by the candidate prediction block and expressed by its associated candidate motion vector. If the candidate motion vector represents horizontal motion, a submatrix is chosen which substantially extends in the horizontal direction. If the candidate motion vector represents vertical motion, the submatrix chosen substantially extends in the vertical direction.
BRIEF DESCRIPTION OF THE DRAWINGS
Fig. 1 shows a hybrid encoder carrying out motion estimation in the transform domain in accordance with the invention.
Fig. 2 shows two phases of a moving object in the pixel domain. Fig. 3 shows an example of the relative position of a prediction block with respect to previously encoded blocks in the transform domain.
Figs. 4A-4C and 5A-5D show examples of subsets of coefficients selected for carrying out motion estimation. Fig. 6 shows a flow chart illustrating the operations of a motion estimation and compensation circuit which is shown in Fig. 1.
Fig. 7 shows an example of the relative position of a prediction macroblock block with respect to previously encoded blocks in the transform domain.
DESCRIPTION OF EMBODIMENTS
Fig. 1 shows a hybrid encoder carrying out motion estimation in the transform domain in accordance with the invention. The encoder receives a video picture signal in the form of picture blocks X of 8*8 pixels. Each pixel block is transformed by a discrete cosine transform circuit 1 into a block Y of 8*8 DCT coefficients. In a subtracter 2, the coefficients of a prediction block Ϋ are subtracted from the corresponding coefficients of block Y to obtain a block of 8*8 difference coefficients. The difference coefficients are quantized in a quantizer 3 and the quantized difference coefficients are then encoded into variable-length code words by a variable-length encoder 4. The variable-length code words constitute the digital output bitstream of the hybrid encoder. This bitstream is applied to a transmission channel or stored on a storage medium (not shown).
In a prediction loop, the quantized difference coefficients are subjected to inverse quantization in an inverse quantizer 5, added to the current prediction block by an adder 6, and written into in a frame memory 7 to update the previously encoded picture stored therein. The picture blocks stored in frame memory 6 are denoted Z. Note that they are stored in the transform domain. A motion estimator and motion compensator circuit 8 receives the current picture block and searches in the frame memory 7 the prediction block
Y which most resembles said picture block. The position of the prediction block thus found is identified by a motion vector mv which defines the position of the prediction block with respect to the current block position. The motion vector mv is applied to the variable-length encoder 4 and is part of the output bitstream.
In the past, motion compensation used to be carried out in the pixel domain. This is shown by way of example in Fig. 2 which shows two phases of a moving object 10 in the pixel domain. In the Figure, a picture segment 11 comprising twelve 8*8 pixel blocks of the picture to be coded and a corresponding picture segment 12 of the previously encoded picture stored in a frame memory are shown. Reference numeral 13 denotes an 8*8 pixel block X to be coded and reference numeral 14 denotes the most- resembling prediction block X in the frame memory. Also shown is the motion vector mv associated with the prediction block X . Once this motion vector has been found by motion estimation, the prediction block can easily be obtained by reading the relevant pixels of prediction block X from the frame memory.
In contrast therewith, the hybrid decoder shown in Fig. 1 stores the previously encoded picture in the transform domain. The coefficients of the prediction block Ϋ must now be calculated from coefficient blocks Z stored in the memory. This is shown by
way of example in Fig. 3 which shows that the coefficients of a prediction block Ϋ are to be calculated from four underlying coefficient blocks Z,, Z2, Z3 and Z4. In mathematical notation, the prediction block follows from the equation:
Ϋ ' ∑ H W, (1)
1=1
where H, and W, are transform matrices representing the relative position of the block Ϋ
with respect to the underlying blocks. H, and W, have one of the formats
Figure imgf000006_0001
where I is an identity matrix and 0 is a matrix of zeroes. A theoretical background of equation (1) can be found, inter alia, in United States Patent US-A-5,408,274.
The above calculation of the coefficients of the prediction block presumes knowledge of the matrices H, and W„ i.e. knowledge of the motion vector mv. It is the purpose of a motion estimator (included in circuit 8 in Fig. 1) to determine the motion vector of the prediction block which most resembles the current input block. Motion estimation is a very complicated process, even in the pixel domain, and much research effort has been spent to find cost-effective and well-performing motion estimation algorithms. In fact, motion estimation is an iterative process of applying motion compensation to a plurality of candidate prediction blocks and selecting an optimal prediction block from said candidates. Hereinafter, an algorithm known as 3-Dimensional Recursive Search (3DRS) block matching algorithm will be referred to by way of example. This algorithm is disclosed, inter alia, in Applicant's United States Patent US-A-5, 072,293. A pleasant property of the 3DRS motion estimation algorithm is that only a restricted number of candidate motion vectors are to be evaluated.
In view of the recognition that motion estimation is similar to repeatedly applying motion compensation to a plurality of candidate prediction blocks, motion estimation in the transform domain requires equation (1) to be calculated several times. This requires a large number of multiplications. In accordance with the invention, the calculation of each candidate prediction block in accordance with equation (1) is carried out for a restricted number of coefficients only. This reduces the number of multiplications so dramatically that a one-chip MPEG2 encoder operating in the transform domain becomes feasible. Needless to say that this a great step in the field of digital video processing.
The coefficients to be calculated for the purpose of motion estimation can be fixedly chosen. Preferably, they include the DC coefficient and a few AC coefficients representing lower spatial frequencies. That is because these coefficients are responsible for the majority of bits in the output signal and are thus most worthwhile to be minimized in the difference block (output of subtracter 2 in Fig. 1). Figs. 4A-4C show examples of preferred choices. In each example, the upper-left coefficient is the DC coefficient. In Fig. 4A, the coefficients constitute a 3*3 (or any other size) submatrix. In Fig. 4B, the selected coefficients are the first coefficients of a zigzag scanned series. This choice is advantageous in view of already available circuits for zigzag scanning coefficient blocks. In Fig. 4C, the submatrix extends in the horizontal direction. The latter choice benefits from the property that most of the motion found in plain video scenes is horizontal motion.
In a preferred embodiment of the invention, the coefficients to be calculated for the purpose of motion estimation are adaptively chosen in dependence upon the candidate motion vectors which are being evaluated. If the candidate motion vectors represent horizontal motion (which in the 3DRS algorithm is the case if substantially horizontal motion has been found previously), a submatrix is chosen which extends in the horizontal direction. Examples thereof are shown in Figs. 5 A and 5B. If the candidate motion vectors represent vertical motion, the submatrix extends in the vertical direction as shown in Figs. 5C and 5D. The examples shown in Figs. 5B and 5D are particularly advantageous because the calculation on of the coefficients of a single row or column drastically contributes to saving the number of multiplications (8 coefficients constituting a single row can more economically be computed than 8 coefficients constituting a 2*4 submatrix). It is further noted that the AC coefficients chosen are not necessarily contiguous to the DC coefficient. Because small amounts of motion become manifest as higher spatial frequencies in the difference block (output of subtracter 2 in Fig. 1), it may also be advantageous to evaluate AC coefficients representing higher spatial frequencies for small amounts of motion.
The operation of the motion estimation and motion compensation circuit 8 (Fig. 1) will now be described. Because it has to successively carry out a plurality of identical operations each involving a number of calculations, the circuit is preferably implemented as a microprocessor-controlled arithmetic unit. The architecture of such a unit is well-known to those skilled in the art and therefore not shown in more detail. For the purpose of disclosing the invention, Fig. 6 shows a flow chart of operations which are carried out by the microprocessor controlling the arithmetic unit. It is assumed that the 3DRS motion estimation algorithm is used and that the candidate motion vectors mvk (k= l..K) are known in advance.
In a first step 20, the processor determines which subset of DCT coefficients are taken into account for the motion estimation process. As already discussed above, the set may be fixedly chosen (cf. Figs. 4A-4C) or adaptively selected in dependence upon the candidate motion vectors (cf. Figs. 5A-5D). In the step 20, the coefficients are expressed in terms of indexes (u=0..7, v=0..7) indicating their positions in the coefficient matrix.
In a step 21, the microprocessor selects one of the candidate motion vectors mvk. Then, in a step 22, the processor retrieves the four coefficient blocks underlying the candidate prediction block associated with this motion vector from the frame memory. In accordance with equation (1), said four blocks are denoted Z,..Z4.
In a step 23, the relevant coefficients of the candidate prediction block are calculated in accordance with equation (1). The coefficients are denoted yu v . Note that only the coefficients yu v are calculated having an index (u,v) to be taken into account as has been determined in the step 20. In a step 24, the microprocessor compares the set of coefficients yu v with the corresponding coefficients yu v of the current input block. The result thereof is stored as a resemblance indicator Rk for the current candidate motion vector mvk. In this example, the resemblance indicator is indicative of the mean absolute difference between input block and candidate prediction block in the transform domain:
Figure imgf000008_0001
where wu v is a weighting factor indicating the relative relevance of the respective coefficient. In video compression schemes such as MPEG, the quantizer matrix provides appropriate weighting factors. In a simple embodiment, all weighting factors wu v may be set to one. In a step 25, it is tested whether all candidate motion vectors mvk (k= l ..K) have been processed in this manner. As long as this is not the case, the steps 21-24 are repeated for the other candidate motion vectors mvk. If all candidate motion vector have been processed, the motion estimation process proceeds with a step 26 in which it is determined which of the stored resemblance indicators Rk (k= l..K) has the smallest value. The best- matching one of the candidate prediction blocks has now been identified and the motion estimation process is finished.
The arithmetic unit now proceeds with motion compensation. To this end, the motion vector mvk of the best-matching prediction block just found is again selected in a step 30. In a step 31, the four underlying blocks X..7^ are re-read from the frame memory. In a step 32, all coefficients of the prediction block Y are calculated in accordance with equation (1). Note the difference of this step 32 with the alleged similar step 23 in the motion estimation process. In the step 32, all the coefficients are calculated whereas the step 23 applies to a few coefficients only. The block Ϋ is applied to the subtracter 2 (see Fig. 1) of the hybrid encoder.
In the above described embodiment, the candidate prediction blocks have the same size as the underlying previously encoded blocks. Some coding standards such as
H.263 provide an option ("advanced motion estimation") to transmit a motion vector for each 8*8 block. The described embodiment can advantageously be implemented in a hybrid encoder in accordance with such a standard. Other coding standards, such as MPEG, transmit one motion vector for each macroblock which includes four 8*8 blocks. For such coding standards, the process of motion estimation and compensation is slightly different, as will now be explained with reference to Fig. 7.
Fig. 7 shows a (candidate) prediction macroblock 40 comprising four 8*8 blocks Ϋχ-Y . The macroblock now covers nine previously encoded coefficient blocks ZrZ9.
To calculate the coefficients of the macroblock, each 8*8 block is individually calculated from its four underlying blocks in accordance with equation (1). The motion estimation and compensation process further proceeds as described above with reference to Fig. 6. For the purpose of motion compensation, all the coefficients (a total of 4*64 in total) are calculated (step 32). For the purpose of motion estimation, only the selected coefficients are calculated (step 23), and the corresponding coefficients are taken from the input macroblock for determining the resemblance indicator (step 25).
In summary, a hybrid video encoder is disclosed which carries out motion estimation and compensation (8) in the transform domain. In such an encoder, the calculation of a prediction block (Ϋ) from previously encoded blocks (Z) stored in the transform-domain frame memory (7) requires a large number of multiplications. This applies particularly to the motion estimation algorithm. In accordance with the invention, only a few DCT coefficients of candidate prediction blocks are calculated, for example, the DC coefficient and some AC coefficients. In a preferred embodiment, the AC coefficients are adaptively selected in dependence upon the motion vector which is being considered. The calculated coefficients of the candidate prediction blocks and the corresponding coefficients of the current input picture block (Y) are then compared to identify the best-matching prediction block.

Claims

Claims
1. A method of coding video pictures, comprising the steps of transforming (1) picture blocks of a video picture into transform coefficients, and encoding (2-8) the difference (2) between the transform coefficients of each transformed picture block (Y) and corresponding transform coefficients of a prediction block ( Ϋ) which is obtained by applying motion estimation (8) to a previously encoded picture, the motion estimation comprising the steps of: calculating (23) a plurality of candidate prediction blocks, comparing (24) each candidate prediction block with the transformed picture block, and - selecting (26) one of the candidate prediction blocks to be the prediction block, characterized in that the steps of calculating (23) a candidate prediction block and comparing
(24) the candidate prediction block with the transformed picture block are carried out for a subset of the transform coefficients of said candidate prediction block and transformed picture block. 2. The method as claimed in claim 1, wherein the subset of transform coefficients comprises the DC coefficient and a predetermined number of AC coefficients.
3. The method as claimed in claim 1 or 2, wherein the subset of transform coefficients is adaptively chosen in dependence upon a candidate motion vector associated with the candidate prediction block. 4. The method as claimed in claim 3, wherein the subset of transform coefficients comprises the first N coefficients of the first M rows of the candidate prediction block, N and M being chosen in dependence upon the horizontal and vertical component, respectively, of the candidate motion vector.
5. The method as claimed in claim 4, wherein N = l for candidate motion vectors having a substantial vertical component and M = l for candidate motion vectors having a substantial horizontal component.
6. The method as claimed in claim 1, wherein said step of comparing (24) each candidate prediction block with the transformed prediction block includes comparing spatially corresponding coefficients, predetermined weighting factors (wu v) being applied to said spatially corresponding coefficients. 7. The method as claimed in claim 1, wherein the candidate prediction blocks are chosen in accordance with the 3-dimensional recursive block matching algorithm. 8. A block-based hybrid video encoder, comprising a transform encoder circuit
(1) for transforming picture blocks of a video picture into transform coefficients, and a predictive encoder (2-8) for encoding the difference (2) between each transformed picture block (Y) and a prediction block (Ϋ) which is obtained by applying motion estimation (8) to a previously encoded picture (Z), the predictive encoder including a motion estimator (8) arranged to: calculate (23) a plurality of candidate prediction blocks, compare (24) each candidate prediction block with the transformed picture block, and - select (26) one of the candidate prediction blocks to be the prediction block, characterized in that the motion estimator is arranged to carry out said calculation (23) of the candidate prediction block and comparison (24) of the candidate prediction block with the transformed picture block for a subset of transform coefficients of said candidate prediction block and transformed picture block.
PCT/IB1999/000151 1998-02-13 1999-01-28 Method and arrangement for video coding Ceased WO1999041912A2 (en)

Priority Applications (4)

Application Number Priority Date Filing Date Title
DE69901525T DE69901525T2 (en) 1998-02-13 1999-01-28 METHOD AND DEVICE FOR VIDEO CODING
JP54124199A JP2001520846A (en) 1998-02-13 1999-01-28 Method and apparatus for encoding a video image
EP99900596A EP0976251B1 (en) 1998-02-13 1999-01-28 Method and arrangement for video coding
KR1019997009376A KR100584495B1 (en) 1998-02-13 1999-01-28 Video picture coding apparatus and method

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
EP98200449 1998-02-13
EP98200449.1 1998-02-13

Publications (2)

Publication Number Publication Date
WO1999041912A2 true WO1999041912A2 (en) 1999-08-19
WO1999041912A3 WO1999041912A3 (en) 1999-11-11

Family

ID=8233387

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/IB1999/000151 Ceased WO1999041912A2 (en) 1998-02-13 1999-01-28 Method and arrangement for video coding

Country Status (7)

Country Link
US (1) US6498815B2 (en)
EP (1) EP0976251B1 (en)
JP (1) JP2001520846A (en)
KR (1) KR100584495B1 (en)
CN (1) CN1166215C (en)
DE (1) DE69901525T2 (en)
WO (1) WO1999041912A2 (en)

Cited By (9)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
WO2002067590A1 (en) * 2001-02-16 2002-08-29 Hantro Products Oy Video encoding of still images
WO2003043342A1 (en) * 2001-11-13 2003-05-22 Hantro Products Oy Method, apparatus and computer for encoding successive images
WO2008133898A1 (en) * 2007-04-23 2008-11-06 Aptina Imaging Corporation Compressed domain image summation apparatus, systems, and methods
US7542614B2 (en) 2007-03-14 2009-06-02 Micron Technology, Inc. Image feature identification and motion compensation apparatus, systems, and methods
US8891626B1 (en) 2011-04-05 2014-11-18 Google Inc. Center of motion for encoding motion fields
US8908767B1 (en) 2012-02-09 2014-12-09 Google Inc. Temporal motion vector prediction
US9094689B2 (en) 2011-07-01 2015-07-28 Google Technology Holdings LLC Motion vector prediction design simplification
US10986361B2 (en) 2013-08-23 2021-04-20 Google Llc Video coding using reference motion vectors
US11317101B2 (en) 2012-06-12 2022-04-26 Google Inc. Inter frame candidate selection for a video encoder

Families Citing this family (27)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US6128047A (en) * 1998-05-20 2000-10-03 Sony Corporation Motion estimation process and system using sparse search block-matching and integral projection
US6697427B1 (en) * 1998-11-03 2004-02-24 Pts Corporation Methods and apparatus for improved motion estimation for video encoding
WO2000034920A1 (en) * 1998-12-07 2000-06-15 Koninklijke Philips Electronics N.V. Motion vector estimation
US6741648B2 (en) * 2000-11-10 2004-05-25 Nokia Corporation Apparatus, and associated method, for selecting an encoding rate by which to encode video frames of a video sequence
KR100490395B1 (en) * 2001-10-29 2005-05-17 삼성전자주식회사 Motion vector estimation method and apparatus thereof
US7742525B1 (en) * 2002-07-14 2010-06-22 Apple Inc. Adaptive motion estimation
EP1480170A1 (en) * 2003-05-20 2004-11-24 Mitsubishi Electric Information Technology Centre Europe B.V. Method and apparatus for processing images
US8111752B2 (en) * 2004-06-27 2012-02-07 Apple Inc. Encoding mode pruning during video encoding
US7792188B2 (en) * 2004-06-27 2010-09-07 Apple Inc. Selecting encoding types and predictive modes for encoding video data
US20050286777A1 (en) * 2004-06-27 2005-12-29 Roger Kumar Encoding and decoding images
US7593465B2 (en) * 2004-09-27 2009-09-22 Lsi Corporation Method for video coding artifacts concealment
KR100694093B1 (en) * 2005-02-18 2007-03-12 삼성전자주식회사 Coefficient Prediction Apparatus and Method for Image Block
US8488684B2 (en) * 2008-09-17 2013-07-16 Qualcomm Incorporated Methods and systems for hybrid MIMO decoding
KR101675116B1 (en) 2009-08-06 2016-11-10 삼성전자 주식회사 Method and apparatus for encoding video, and method and apparatus for decoding video
US20110206132A1 (en) * 2010-02-19 2011-08-25 Lazar Bivolarsky Data Compression for Video
US9313526B2 (en) 2010-02-19 2016-04-12 Skype Data compression for video
US9078009B2 (en) 2010-02-19 2015-07-07 Skype Data compression for video utilizing non-translational motion information
US9609342B2 (en) 2010-02-19 2017-03-28 Skype Compression for frames of a video signal using selected candidate blocks
US9819358B2 (en) 2010-02-19 2017-11-14 Skype Entropy encoding based on observed frequency
GB2487200A (en) 2011-01-12 2012-07-18 Canon Kk Video encoding and decoding with improved error resilience
GB2491589B (en) 2011-06-06 2015-12-16 Canon Kk Method and device for encoding a sequence of images and method and device for decoding a sequence of image
CN103733625B (en) * 2011-06-14 2017-04-05 三星电子株式会社 Method for decoding motion vector
KR101616010B1 (en) 2011-11-04 2016-05-17 구글 테크놀로지 홀딩스 엘엘씨 Motion vector scaling for non-uniform motion vector grid
CN105027560A (en) * 2012-01-21 2015-11-04 摩托罗拉移动有限责任公司 Method of determining binary codewords for transform coefficients
US9172970B1 (en) 2012-05-29 2015-10-27 Google Inc. Inter frame candidate selection for a video encoder
US9503746B2 (en) 2012-10-08 2016-11-22 Google Inc. Determine reference motion vectors
US9313493B1 (en) 2013-06-27 2016-04-12 Google Inc. Advanced motion estimation

Family Cites Families (10)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US4245248A (en) * 1979-04-04 1981-01-13 Bell Telephone Laboratories, Incorporated Motion estimation and encoding of video signals in the transform domain
US4704628A (en) * 1986-07-16 1987-11-03 Compression Labs, Inc. Combined intraframe and interframe transform coding system
SE457402B (en) * 1987-02-20 1988-12-19 Harald Brusewitz PROCEDURE AND DEVICE FOR CODING AND DECODING IMAGE INFORMATION
US5072293A (en) 1989-08-29 1991-12-10 U.S. Philips Corporation Method of estimating motion in a picture signal
DE4118571A1 (en) * 1991-06-06 1992-12-10 Philips Patentverwaltung DEVICE FOR CONTROLLING THE QUANTIZER OF A HYBRID ENCODER
US5408274A (en) * 1993-03-11 1995-04-18 The Regents Of The University Of California Method and apparatus for compositing compressed video data
US5712809A (en) * 1994-10-31 1998-01-27 Vivo Software, Inc. Method and apparatus for performing fast reduced coefficient discrete cosine transforms
PL181416B1 (en) * 1995-09-12 2001-07-31 Philips Electronics Nv Mixed signal/model image encoding and decoding
US5790686A (en) * 1995-09-19 1998-08-04 University Of Maryland At College Park DCT-based motion estimation method
CA2310602C (en) * 1997-11-14 2009-05-19 Analysis & Technology, Inc. Apparatus and method for compressing video information

Cited By (12)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
WO2002067590A1 (en) * 2001-02-16 2002-08-29 Hantro Products Oy Video encoding of still images
WO2003043342A1 (en) * 2001-11-13 2003-05-22 Hantro Products Oy Method, apparatus and computer for encoding successive images
US7542614B2 (en) 2007-03-14 2009-06-02 Micron Technology, Inc. Image feature identification and motion compensation apparatus, systems, and methods
US7599566B2 (en) 2007-03-14 2009-10-06 Aptina Imaging Corporation Image feature identification and motion compensation apparatus, systems, and methods
US7924316B2 (en) 2007-03-14 2011-04-12 Aptina Imaging Corporation Image feature identification and motion compensation apparatus, systems, and methods
WO2008133898A1 (en) * 2007-04-23 2008-11-06 Aptina Imaging Corporation Compressed domain image summation apparatus, systems, and methods
US7920746B2 (en) 2007-04-23 2011-04-05 Aptina Imaging Corporation Compressed domain image summation apparatus, systems, and methods
US8891626B1 (en) 2011-04-05 2014-11-18 Google Inc. Center of motion for encoding motion fields
US9094689B2 (en) 2011-07-01 2015-07-28 Google Technology Holdings LLC Motion vector prediction design simplification
US8908767B1 (en) 2012-02-09 2014-12-09 Google Inc. Temporal motion vector prediction
US11317101B2 (en) 2012-06-12 2022-04-26 Google Inc. Inter frame candidate selection for a video encoder
US10986361B2 (en) 2013-08-23 2021-04-20 Google Llc Video coding using reference motion vectors

Also Published As

Publication number Publication date
JP2001520846A (en) 2001-10-30
US6498815B2 (en) 2002-12-24
CN1256049A (en) 2000-06-07
EP0976251B1 (en) 2002-05-22
DE69901525T2 (en) 2003-01-09
WO1999041912A3 (en) 1999-11-11
KR20010006292A (en) 2001-01-26
KR100584495B1 (en) 2006-06-02
US20020085630A1 (en) 2002-07-04
DE69901525D1 (en) 2002-06-27
CN1166215C (en) 2004-09-08
EP0976251A2 (en) 2000-02-02

Similar Documents

Publication Publication Date Title
EP0976251B1 (en) Method and arrangement for video coding
US8023562B2 (en) Real-time video coding/decoding
US7116830B2 (en) Spatial extrapolation of pixel values in intraframe video coding and decoding
JP2636622B2 (en) Video signal encoding method and decoding method, and video signal encoding apparatus and decoding apparatus
US6289049B1 (en) Method for coding motion vector in moving picture
KR950009699B1 (en) Motion vector detection method and apparatus
EP0637894B1 (en) Apparatus and method for detecting motion vectors to half-pixel accuracy
US5532747A (en) Method for effectuating half-pixel motion compensation in decoding an image signal
EP2168382B1 (en) Method for processing images and the corresponding electronic device
JP3836559B2 (en) Motion estimation method in digital video encoder
US6735345B2 (en) Efficient macroblock header coding for video compression
US20060008008A1 (en) Method of multi-resolution based motion estimation and recording medium storing program to implement the method
EP1629437A2 (en) Hybrid digital video compression
WO2002054777A1 (en) Mpeg-2 down-sampled video generation
KR20070027119A (en) Improved motion estimation method, video encoding method and apparatus using the method
WO2011080807A1 (en) Moving picture coding device and moving picture decoding device
US5689312A (en) Block matching motion estimation method
KR100727988B1 (en) Method and apparatus for predicting DC coefficient in transform domain
US20080107183A1 (en) Method and apparatus for detecting zero coefficients
KR100196828B1 (en) Method for selecting motion vector in image encoder
JPH08509583A (en) Differential encoding and decoding method and related circuit
JP2003189306A (en) Process and device for decoding video data coded according to mpeg standard
Choi et al. Adaptive image quantization using total variation classification
KR100617177B1 (en) Motion estimation method
JP2000513914A (en) Method and apparatus for encoding and decoding digital images

Legal Events

Date Code Title Description
WWE Wipo information: entry into national phase

Ref document number: 99800128.7

Country of ref document: CN

AK Designated states

Kind code of ref document: A2

Designated state(s): CN JP KR

AL Designated countries for regional patents

Kind code of ref document: A2

Designated state(s): AT BE CH CY DE DK ES FI FR GB GR IE IT LU MC NL PT SE

WWE Wipo information: entry into national phase

Ref document number: 1999900596

Country of ref document: EP

ENP Entry into the national phase

Ref document number: 1999 541241

Country of ref document: JP

Kind code of ref document: A

WWE Wipo information: entry into national phase

Ref document number: 1019997009376

Country of ref document: KR

121 Ep: the epo has been informed by wipo that ep was designated in this application
AK Designated states

Kind code of ref document: A3

Designated state(s): CN JP KR

AL Designated countries for regional patents

Kind code of ref document: A3

Designated state(s): AT BE CH CY DE DK ES FI FR GB GR IE IT LU MC NL PT SE

WWP Wipo information: published in national office

Ref document number: 1999900596

Country of ref document: EP

WWP Wipo information: published in national office

Ref document number: 1019997009376

Country of ref document: KR

WWG Wipo information: grant in national office

Ref document number: 1999900596

Country of ref document: EP

WWG Wipo information: grant in national office

Ref document number: 1019997009376

Country of ref document: KR