Grouping of image frames in video coding

Patent No. US7894521 (titled "Grouping of image frames in video coding") on Nov 29, 2002. The application was issued on Feb 22, 2011.

What is this patent about?

’521 is related to the field of scalable video coding and streaming. It addresses the challenges of managing compressed video sequences where frames are interdependent, particularly when a streaming server or network element needs to adjust the bit rate by removing certain frames without fully decoding the stream. The background context involves the use of temporal prediction, where frames like P-frames and B-frames rely on reference images, making it difficult to scale the video without causing decoding errors or synchronization issues.

The underlying idea behind ’521 is the organization of video frames into logically grouped sub-sequences that explicitly signal their interdependencies. Rather than treating a video as a linear string of frames or simple hierarchical layers, the invention treats groups of frames as independent units that carry metadata identifying which other sub-sequences they rely on for prediction. This insight allows a network node to prune the video stream by removing entire sub-sequences that are not needed for the reconstruction of the remaining parts, effectively enabling bit rate control without the overhead of full video parsing.

The claims of ’521 focus on a method and apparatus for encoding and decoding a video sequence by defining at least two distinct sub-sequences. The first sub-sequence contains independent frames, such as I-frames, while the second sub-sequence contains predicted frames that rely on the first. Crucially, the independent claims require the determination and encoding of explicit dependency data into the video sequence, which identifies the relationship between the second sub-sequence and the first, separate from standard picture type information.

In practice, the invention works by assigning each frame a unique identifier that combines a scalability layer number, a sub-sequence ID, and an image number. This metadata is preferably placed in the header fields of the transport protocol or the video bitstream itself. When a streaming server detects a drop in available bandwidth, it can examine these headers to identify which sub-sequences can be safely discarded. Because the dependencies are clearly mapped, the server can ensure that if a reference sub-sequence is removed, all dependent sub-sequences are also removed, preventing the decoder from attempting to process broken links.

This approach differs from prior solutions by moving the intelligence of frame dependency from the deep video payload to the accessible metadata layer. Traditional methods often required a server to buffer and parse the video for a long period to detect dependencies, or they suffered from numbering discontinuities when frames were removed. By using sub-sequence-specific numbering and explicit dependency signaling, the invention allows for seamless bit rate adjustment and easier random access, as the decoder can immediately identify the start of a valid, decodable group of frames even after a stream has been modified.

How does this patent fit in bigger picture?

Technical Landscape

In the early 2000s when ’521 was filed, multimedia streaming was typically implemented using a server-client architecture where data playback began before the complete file transfer, often constrained by fluctuating network bandwidth and limited terminal storage. At a time when systems commonly relied on fixed Group of Pictures (GOP) structures, video compression was achieved by reducing spatial and temporal redundancy through Intra-frames and motion-compensated Inter-frames. However, hardware and software constraints made the dynamic adjustment of bit rates non-trivial, as existing scalable coding methods required intensive decoding and buffering to identify inter-frame dependencies, particularly when reference picture selection or non-sequential image numbering was employed.

Prosecution Position

The disclosed invention represents a technical advancement by introducing an architectural shift in how scalable video sequences are organized and signaled. By grouping video frames into sub-sequences and explicitly determining dependency data—where each sub-sequence identifies the specific upper-layer or peer frames required for its prediction—the system enables a more flexible coding hierarchy. This integration of dependency identifiers into the bit stream allows streaming servers to perform bit-rate scaling and sub-sequence removal without the need for full decoding or complex parsing. Furthermore, the unique identifier combining scalability layer, sub-sequence ID, and image number overcomes the technical constraint of image numbering discontinuities, facilitating the seamless insertion of external video sequences and improving error detection during transmission.

Claims

The patent contains a total of 24 claims, with claims 1, 13, 14, 20, 21, 22, and 23 serving as the independent claims. These independent claims are generally focused on methods, encoding and decoding hardware, and computer program products for managing scalable video sequences by defining sub-sequences of different frame formats and encoding or determining the specific dependencies between these sub-sequences. The dependent claims generally serve to provide additional technical details regarding scalability layers, unique frame identification through identifiers like layer numbers and image numbers, buffer memory management for handling intentional image number discontinuities, and the handling of overlapping sub-sequences during the decoding process.

Key Claim Terms New

Definitions of key terms used in the patent claims.

Term (Source)Support for SpecificationInterpretation
Dependency
(Claim 1, Claim 13, Claim 14, Claim 20, Claim 21, Claim 22)
A sub-sequence will comprise information of all sub-sequences that have been directly used for predicting the image frames comprised by the sub-sequence in question. This information is signalled in the bit stream of the video sequence, preferably separate from the actual image information. The streaming server is capable of deducing the inter-dependencies of the different sub-sequences directly from the identifiers and their inter-dependencies.Signalled information in the bit stream that identifies which sub-sequences or frames were used as references for predicting the image frames of a specific sub-sequence.
First frame format
(Claim 1, Claim 13, Claim 14, Claim 20, Claim 21, Claim 22, Claim 23)
Such frames are called INTRA-frames, or I-frames. A video sequence always comprises some compressed image frames the image information of which has not been determined using motion-compensated temporal prediction. The video frames conforming to the first frame format being independent of other video frames, i.e. they are typically I-frames.A coding format for video frames that are independent of other video frames, typically corresponding to INTRA-frames or I-frames which do not use motion-compensated temporal prediction.
Scalable, compressed video sequence
(Claim 1, Claim 13, Claim 14, Claim 20, Claim 21, Claim 22, Claim 23)
To allow for flexible streaming of video files, many video coding systems employ scalable coding in which some elements or element groups of a video sequence can be removed without affecting the reconstruction of other parts of the video sequence. Scalability is typically implemented by grouping the image frames into a number of hierarchical layers. The base layer of each group of pictures GOP thus comprises one I-frame and a necessary number of P-frames.A video bit stream coded in a hierarchical manner (e.g., base and enhancement layers) allowing parts of the sequence to be removed to adjust bit rate without preventing the decoding of the remaining parts.
Second frame format
(Claim 1, Claim 13, Claim 14, Claim 20, Claim 21, Claim 22, Claim 23)
Correspondingly, motion-compensated video sequence image frames predicted from previous image frames, are called INTER-frames, or P-frames (Predicted). The video frames according to the second frame format being frames predicted from at least one other video frame, for example P-frames. In addition, many video compression methods employ bi-directionally predicted B-frames (Bi-directional).A coding format for video frames predicted from at least one other video frame, such as INTER-frames (P-frames) or bi-directionally predicted frames (B-frames).
Sub-sequence
(Claim 1, Claim 13, Claim 14, Claim 20, Claim 21, Claim 22, Claim 23)
A video sequence comprises a first sub-sequence formed therein, at least part of the sub-sequence being formed by coding video frames (I-frames) of the at least first frame format (I-frames), and at least a second sub-sequence, at least part of which is formed by coding video frames of the at least second frame format (e.g. P-frames). This allows the streaming server, for example, to adjust the bit rate advantageously, without video sequence decoding, parsing and buffering, because the streaming server is capable of deducing the inter-dependencies of the different sub-sequences directly from the identifiers and their inter-dependencies. Consequently, an essential aspect of the present invention is to determine the sub-sequences each sub-sequence is dependent on.A distinct grouping of video frames within a video sequence, where frames within the group may have specific temporal dependencies or be independently decodable as a unit.

Litigation Cases New

US Latest litigation cases involving this patent.

Case NumberFiling DateTitle
0:24-cv-04269Nov 25, 2024Element Television Company, Llc V. Nokia Corporation
1:23-cv-01236Oct 31, 2023Nokia Technologies Oy V. Amazon.Com, Inc.

Patent Family

Patent Family

File Wrapper

The dossier documents provide a comprehensive record of the patent's prosecution history - including filings, correspondence, and decisions made by patent offices - and are crucial for understanding the patent's legal journey and any challenges it may have faced during examination.

  • Get instant alerts for new documents

US7894521

SEP
Application Number
US10306942A
Filing Date
Nov 29, 2002
Status
Expired
Publication Date
Feb 22, 2011
External Links
Slate, USPTO , Google Patents