Spatially enhanced transform coding

Patent No. US8077991 (titled "Spatially enhanced transform coding") on Apr 10, 2008. The application was issued on Dec 13, 2011.

What is this patent about?

’991 is related to the field of digital video compression and, more specifically, to the efficient encoding of prediction error signals. In standard hybrid video codecs, image blocks are predicted using temporal or spatial methods, and the resulting residual—the difference between the original and the predicted pixels—is typically compressed using frequency-domain transforms like the Discrete Cosine Transform. While these transforms excel at packing energy for highly correlated data, they often struggle with localized irregularities or high-frequency noise, leading to coding inefficiencies.

The underlying idea behind ’991 is to treat the prediction error signal as a composite of two distinct types of data: well-correlated components suitable for frequency transforms and localized outlier values that are better handled in the spatial domain. Rather than forcing a transform to account for sharp, isolated pixel differences—which would scatter energy across many coefficients—the invention isolates these outliers. By coding the bulk of the signal through a transform and the remaining localized errors through direct spatial quantization, the system achieves a higher coding gain than either method could provide alone.

The claims of ’991 focus on a hybrid coding architecture that generates a prediction error signal by joining a transform-coded representation with a spatially-coded representation of the same difference signal. The independent claims cover both the encoding process—where a difference signal is decomposed into these two components—and the corresponding decoding process, which reconstructs the image block by summing the decoded transform information, the decoded spatial samples, and the original prediction.

In a practical implementation, the encoder identifies specific pixels where the prediction error is significantly higher than the surrounding neighborhood. These outliers are temporarily replaced with interpolated values to create a smooth, modified signal that the transform can compress highly efficiently. The difference between the original outlier and its smoothed version is then captured as a spatial sample, ensuring that the fine details or sharp edges are not lost or blurred by the quantization of frequency coefficients.

This approach differs from prior solutions by moving away from a strict 'either-or' choice between transform and spatial domains for a given block. Instead of selecting one mode for an entire macroblock, the invention allows concurrent coding of both types within the same data unit. This dual-path reconstruction enables the codec to maintain high fidelity for textures and edges while leveraging the energy compaction of transforms for the rest of the signal, effectively reducing the total bitrate required for high-quality video.

How does this patent fit in bigger picture?

Technical Landscape

In the mid-2000s when ’991 was filed, digital video compression was typically implemented using hybrid coding architectures that relied on a sequential pipeline of prediction followed by transform coding. At a time when systems commonly relied on Discrete Cosine Transforms (DCT) to decorrelate prediction error signals, the efficiency of the compression was heavily dependent on the signal's statistical correlation with fixed basis functions. When hardware and software constraints made the processing of non-correlated residuals non-trivial, standard practices dictated a binary choice in the coding path: either representing the entire residual block in the frequency domain or, less commonly, entirely in the spatial domain. This architectural rigidity meant that high-frequency textures, sensor noise, or edge information that did not align with transform basis functions often resulted in suboptimal bitrate allocation or significant quantization artifacts.

Prosecution Position

The disclosed invention represents a meaningful technical advancement through an architectural shift from mutually exclusive coding domains to a concurrent, dual-domain representation of prediction errors. By constructing the prediction error signal for a single image block as a weighted sum of both transform basis functions and quantized spatial samples, the system enables the simultaneous capture of well-correlated signal energy and stochastic or high-frequency components that resist traditional transform packing. This integration overcomes the technical constraint of coding efficiency degradation in advanced motion-compensated systems where residuals become increasingly decorrelated. The resulting technical effect is a more flexible and precise reconstruction of the prediction error, allowing the encoder to optimize the bitstream by utilizing frequency-domain coefficients for global block structure while employing spatial-domain samples for localized irregularities, thereby improving overall compression fidelity without significantly increasing decoder complexity.

Claims

This patent contains 39 claims, with independent claims 1, 13, 22, 31, and 39 focusing on methods and apparatuses for encoding and decoding data by combining transform coding and spatial coding to represent prediction error signals. The independent claims specifically address calculating difference signals, substituting outlier values, and joining transform and spatial representations for transmission or storage, as well as the corresponding processes for receiving and reconstructing the original data blocks. The dependent claims serve to further define the technical implementation by specifying outlier replacement techniques, dequantization relationships, the use of discrete orthogonal transforms, and the signaling of missing coefficients or samples within the data blocks.

Key Claim Terms New

Definitions of key terms used in the patent claims.

Term (Source)Support for SpecificationInterpretation
Coded prediction error signal
(Claim 1, Claim 22, Claim 31, Claim 39)
Various embodiments of the present invention provide a system and method for representing the prediction error signal as a weighted sum of different basis functions of a selected transform and quantized spatial samples. The coded prediction error signal includes a plurality of transform coefficients and a plurality of spatial samples. The decoded transform information, the decoded spatial information, and a reconstructed prediction of the block of data are then added, thereby forming a decoded representation of the block of data.A composite data representation of a block's prediction error that simultaneously includes both transform coefficients and spatial samples.
Difference signal
(Claim 1, Claim 39)
The second phase involves coding the prediction error, i.e. the difference between the predicted block of pixels and the original block of pixels. As a consequence of these inter-picture and intra-picture prediction techniques, a residual/error signal is formed by removing the predicated picture frame from the original. This prediction error signal is then typically block transform coded.A signal representing the pixel-by-pixel difference between a predicted block of data (generated via motion compensation or spatial mechanisms) and the original input block of pixels.
Outlier values
(Claim 39)
Substituting, by a processor, outlier values in the difference signal with non-outlier values to create a modified prediction error signal. [Note: The specification describes the general process of handling components not well correlated with basis functions, such as noise or edges, which correspond to the outlier substitution logic in the claims.]Specific values within the difference signal that deviate significantly from the expected distribution, which are substituted to create a modified signal for more efficient coding.
Spatial coding
(Claim 1, Claim 13, Claim 39)
Various embodiments of the present invention allow for the efficient spatial representation of those components of the prediction error signal of the same image block that are not well correlated with the basis functions of the applied transform. These components may include certain types of sensor noise, high frequency texture and edge information. The prediction error signal for a single image block is constructed using both transform basis functions and spatial samples (i.e., pixel values).A process of representing components of a signal directly in the spatial pixel domain, specifically used for signal elements not well correlated with transform basis functions.
Transform coding
(Claim 1, Claim 13, Claim 39)
Transform coding of the prediction error signal in video or image compression system typically comprises DCT-based linear transform, quantization of the transformed DCT coefficients, and context based entropy coding of the quantized coefficients. This allows for the utilization of those selected transform basis functions that give good overall representation of the image block with minimal amount of transform coefficients. The basis functions of the selected transform may comprise an orthogonal set of basis vectors, or the basis functions may not comprise an orthogonal set.A process of representing a signal as a weighted sum of basis functions of a selected transform (such as DCT) to reduce spatial redundancies by packing energy into transform coefficients.

Litigation Cases New

US Latest litigation cases involving this patent.

Case NumberFiling DateTitle
0:24-cv-04269Nov 25, 2024Element Television Company, Llc V. Nokia Corporation
1:23-cv-01237Oct 31, 2023Nokia Technologies Oy V. Hp, Inc.
1:23-cv-01236Oct 31, 2023Nokia Technologies Oy V. Amazon.Com, Inc.

Patent Family

Patent Family

File Wrapper

The dossier documents provide a comprehensive record of the patent's prosecution history - including filings, correspondence, and decisions made by patent offices - and are crucial for understanding the patent's legal journey and any challenges it may have faced during examination.

  • Get instant alerts for new documents

US8077991

SEP
Application Number
US12101019A
Filing Date
Apr 10, 2008
Status
Granted
Publication Date
Dec 13, 2011
External Links
Slate, USPTO , Google Patents