Motion prediction in video coding

Patent No. US9877037 (titled "Motion prediction in video coding") on Apr 18, 2017. The application was issued on Jan 23, 2018.

What is this patent about?

’037 is related to the field of video compression and motion compensation, specifically addressing how to manage pixel data precision during the reconstruction of frames that rely on multiple reference sources. In modern video codecs, bi-directional prediction improves efficiency by averaging multiple blocks, but the standard process of rounding these values at each intermediate step often introduces cumulative errors that degrade the final image quality.

The underlying idea behind ’037 is to mitigate rounding error accumulation by maintaining high-precision intermediate values throughout the multi-directional prediction process. Rather than rounding each individual reference block prediction down to the standard bit-depth immediately after interpolation, the system keeps these values in an expanded bit-depth format, combines them, and only performs a single precision reduction at the very end of the calculation chain.

The claims of ’037 focus on a method and apparatus that identifies bi-predicted or multi-predicted blocks and performs sub-pixel interpolation to generate multiple prediction signals at a second, higher precision. These high-precision signals are then merged into a combined prediction, which is subsequently downshifted to the original bit-depth of the video representation using a right-shift operation.

In practice, the invention works by utilizing a larger bit-depth in the registers during the filtering and summation stages of motion compensation. For example, if the input video uses 8-bit pixels, the interpolation filters generate 16-bit intermediate results; these 16-bit values are summed together, and only after this combination is the result shifted back to 8-bit accuracy, ensuring that the final pixel value is more mathematically representative of the source material.

This approach differs from prior solutions that required signaling a specific rounding offset in the bitstream or alternating rounding directions between frames to cancel out errors. By deferring the precision reduction until after the signals are combined, the invention simplifies the codec architecture and removes the need for extra signaling overhead while simultaneously improving the fidelity of the reconstructed video signal.

How does this patent fit in bigger picture?

Technical Landscape

In the early 2010s when ’037 was filed, video compression systems were typically implemented using hybrid coding architectures that relied on motion-compensated inter-prediction to reduce temporal redundancy. At a time when hardware and software constraints made memory bandwidth and computational complexity non-trivial, systems commonly relied on rounding intermediate prediction values to a standard bit-depth immediately after interpolation to minimize the data width of internal buffers. Consequently, when combining multiple prediction signals for bi-directional or multi-directional frames, the cumulative effect of successive rounding operations often introduced significant quantization errors and drift, which necessitated the inclusion of rounding direction indicators or complex offset signaling within the bitstream to maintain visual fidelity.

Prosecution Position

The disclosed invention represents a technical advancement in video coding efficiency and precision through an architectural shift in the prediction pipeline. By maintaining individual prediction signals at a higher bit-depth precision throughout the interpolation and combination stages, the system delays precision reduction until after the multi-directional signals have been merged. This integration of high-precision intermediate processing overcomes the technical constraint of rounding-error accumulation inherent in multi-stage prediction. The resulting technical effect is a reduction in quantization noise and the elimination of the need for rounding-offset signaling in the bitstream, thereby improving both the objective quality of the reconstructed video and the overall compression efficiency.

Claims

This patent contains 26 claims, with claims 1, 8, 15, 19, and 23 being independent. The independent claims focus on a method, apparatus, and computer program product for video encoding or decoding that manages pixel precision during multi-reference block prediction by obtaining high-precision interpolated predictions, combining them, and subsequently reducing the precision through right-bit shifting. The dependent claims further specify the use of integer samples, the application of specific rounding offsets, the reduction of precision to intermediate levels, and the application of these techniques to bi-directional or multidirectional block types.

Key Claim Terms New

Definitions of key terms used in the patent claims.

Term (Source)Support for SpecificationInterpretation
Combined prediction
(Claim 1, Claim 8, Claim 15, Claim 19, Claim 23)
The accuracy of the bi-directional or multidirectional prediction signal can then be downshifted to an appropriate accuracy for post processing purposes. This invention may keep the motion compensated prediction signal of each one of the predictions at highest precision possible after interpolation and perform the rounding to the bit-depth range of the video signal after both prediction signals are added.A multi-directional prediction signal (such as bi-directional) formed by merging two or more high-precision prediction signals before they are downshifted to the final bit-depth.
First precision
(Claim 1, Claim 8, Claim 15, Claim 19, Claim 23)
The first precision indicates the number of bits needed to represent values of pixels. The accuracy of the bi-directional or multidirectional prediction signal can then be downshifted to an appropriate accuracy for post processing purposes. This invention may keep the motion compensated prediction signal of each one of the predictions at highest precision possible after interpolation and perform the rounding to the bit-depth range of the video signal after both prediction signals are added.The bit-depth or number of bits used to represent the pixel values of the original video signal or reference blocks, typically representing the standard accuracy of the video representation.
Obtain a first prediction by interpolation
(Claim 1, Claim 8, Claim 15, Claim 19, Claim 23)
This invention may keep the motion compensated prediction signal of each one of the predictions at highest precision possible after interpolation and perform the rounding to the bit-depth range of the video signal after both prediction signals are added. The encoder may discard some information in the original video sequence in order to represent the video in a more compact form. Prediction signals are maintained in higher accuracy until the prediction signals have been combined.The process of generating a predicted pixel value from a reference block using sub-pixel estimation, where the resulting value is kept at a higher bit-depth than the source pixels.
Second precision
(Claim 1, Claim 8, Claim 15, Claim 19, Claim 23)
According to some embodiments of the invention prediction signals are maintained in a higher precision during the prediction calculation and the precision is reduced after the two or more prediction signals have been combined with each other. In some example embodiments prediction signals are maintained in higher accuracy until the prediction signals have been combined to obtain the bi-directional or multidirectional prediction signal. The second precision indicates the number of bits needed to represent values of said first prediction and values of said second prediction.An intermediate, higher bit-depth used during the interpolation and prediction calculation phases to maintain accuracy and reduce rounding errors before final combination.
Shifting bits of the combined prediction to the right
(Claim 1, Claim 8, Claim 15, Claim 19, Claim 23)
The accuracy of the bi-directional or multidirectional prediction signal can then be downshifted to an appropriate accuracy for post processing purposes. According to some embodiments of the invention prediction signals are maintained in a higher precision during the prediction calculation and the precision is reduced after the two or more prediction signals have been combined with each other.A bitwise operation used to decrease the numerical precision of the combined prediction signal, effectively performing a division by a power of two to return the value to the first precision.

Litigation Cases New

US Latest litigation cases involving this patent.

Case NumberFiling DateTitle
1:25-cv-00523Apr 7, 2025Nokia Technologies Oy V. Acer Inc.
1:23-cv-01237Oct 31, 2023Nokia Technologies Oy V. Hp, Inc.
1:23-cv-01236Oct 31, 2023Nokia Technologies Oy V. Amazon.Com, Inc.

Patent Family

Patent Family

File Wrapper

The dossier documents provide a comprehensive record of the patent's prosecution history - including filings, correspondence, and decisions made by patent offices - and are crucial for understanding the patent's legal journey and any challenges it may have faced during examination.

  • Get instant alerts for new documents

US9877037

SEP
Application Number
US15490469A
Filing Date
Apr 18, 2017
Status
Granted
Publication Date
Jan 23, 2018
External Links
Slate, USPTO , Google Patents