Patent No. US7894521 (titled "Grouping of image frames in video coding") on Nov 29, 2002. The application was issued on Feb 22, 2011.
’521 is related to the field of scalable video coding and streaming. It addresses the challenges of managing compressed video sequences where frames are interdependent, particularly when a streaming server or network element needs to adjust the bit rate by removing certain frames without fully decoding the stream. The background context involves the use of temporal prediction, where frames like P-frames and B-frames rely on reference images, making it difficult to scale the video without causing decoding errors or synchronization issues.
The underlying idea behind ’521 is the organization of video frames into logically grouped sub-sequences that explicitly signal their interdependencies. Rather than treating a video as a linear string of frames or simple hierarchical layers, the invention treats groups of frames as independent units that carry metadata identifying which other sub-sequences they rely on for prediction. This insight allows a network node to prune the video stream by removing entire sub-sequences that are not needed for the reconstruction of the remaining parts, effectively enabling bit rate control without the overhead of full video parsing.
The claims of ’521 focus on a method and apparatus for encoding and decoding a video sequence by defining at least two distinct sub-sequences. The first sub-sequence contains independent frames, such as I-frames, while the second sub-sequence contains predicted frames that rely on the first. Crucially, the independent claims require the determination and encoding of explicit dependency data into the video sequence, which identifies the relationship between the second sub-sequence and the first, separate from standard picture type information.
In practice, the invention works by assigning each frame a unique identifier that combines a scalability layer number, a sub-sequence ID, and an image number. This metadata is preferably placed in the header fields of the transport protocol or the video bitstream itself. When a streaming server detects a drop in available bandwidth, it can examine these headers to identify which sub-sequences can be safely discarded. Because the dependencies are clearly mapped, the server can ensure that if a reference sub-sequence is removed, all dependent sub-sequences are also removed, preventing the decoder from attempting to process broken links.
This approach differs from prior solutions by moving the intelligence of frame dependency from the deep video payload to the accessible metadata layer. Traditional methods often required a server to buffer and parse the video for a long period to detect dependencies, or they suffered from numbering discontinuities when frames were removed. By using sub-sequence-specific numbering and explicit dependency signaling, the invention allows for seamless bit rate adjustment and easier random access, as the decoder can immediately identify the start of a valid, decodable group of frames even after a stream has been modified.
In the early 2000s when ’521 was filed, multimedia streaming was typically implemented using a server-client architecture where data playback began before the complete file transfer, often constrained by fluctuating network bandwidth and limited terminal storage. At a time when systems commonly relied on fixed Group of Pictures (GOP) structures, video compression was achieved by reducing spatial and temporal redundancy through Intra-frames and motion-compensated Inter-frames. However, hardware and software constraints made the dynamic adjustment of bit rates non-trivial, as existing scalable coding methods required intensive decoding and buffering to identify inter-frame dependencies, particularly when reference picture selection or non-sequential image numbering was employed.
The disclosed invention represents a technical advancement by introducing an architectural shift in how scalable video sequences are organized and signaled. By grouping video frames into sub-sequences and explicitly determining dependency data—where each sub-sequence identifies the specific upper-layer or peer frames required for its prediction—the system enables a more flexible coding hierarchy. This integration of dependency identifiers into the bit stream allows streaming servers to perform bit-rate scaling and sub-sequence removal without the need for full decoding or complex parsing. Furthermore, the unique identifier combining scalability layer, sub-sequence ID, and image number overcomes the technical constraint of image numbering discontinuities, facilitating the seamless insertion of external video sequences and improving error detection during transmission.
The patent contains a total of 24 claims, with claims 1, 13, 14, 20, 21, 22, and 23 serving as the independent claims. These independent claims are generally focused on methods, encoding and decoding hardware, and computer program products for managing scalable video sequences by defining sub-sequences of different frame formats and encoding or determining the specific dependencies between these sub-sequences. The dependent claims generally serve to provide additional technical details regarding scalability layers, unique frame identification through identifiers like layer numbers and image numbers, buffer memory management for handling intentional image number discontinuities, and the handling of overlapping sub-sequences during the decoding process.
Definitions of key terms used in the patent claims.
US Latest litigation cases involving this patent.

The dossier documents provide a comprehensive record of the patent's prosecution history - including filings, correspondence, and decisions made by patent offices - and are crucial for understanding the patent's legal journey and any challenges it may have faced during examination.
Get instant alerts for new documents