Patent No. US8050321 (titled "Grouping of image frames in video coding") on Jan 25, 2006. The application was issued on Nov 1, 2011.
’321 is related to the field of digital video compression and streaming, specifically addressing the challenges of random access and error recovery within complex prediction hierarchies. In modern video codecs, temporal redundancy is reduced by predicting frames from previously decoded references; however, this creates interdependencies that make it difficult for a decoder to join a stream mid-way or recover from packet loss without decoding the entire preceding sequence.
The underlying idea behind ’321 is the explicit signaling of an initiation picture that marks the start of a strictly independent sequence of frames. By resetting the frame numbering scheme and providing a specific indicator at this boundary, the system ensures that no subsequent frames in that sequence will ever reference a picture that appeared before the initiation point. This creates a clean break in the dependency chain, allowing the decoder to flush its buffer and start fresh without risking the propagation of errors from missing prior data.
The claims of ’321 focus on a method and apparatus for encoding and decoding an independent sequence by embedding a specific indication of the first image frame in the decoding order. The invention specifically requires encoding identifier values according to a numbering scheme and then resetting the identifier value (typically to zero) for that first image frame. This mechanism ensures that all motion-compensated temporal prediction references within the sequence are self-contained and do not point to frames outside the defined boundary.
In practice, this is implemented by adding a separate flag or marker within the slice header of the video bitstream. When a user attempts to jump to a new position in a video file or joins a live broadcast, the decoder looks for this flag to identify a safe entry point. Once detected, the decoder treats the initiation picture as the new temporal anchor, effectively treating all previously stored reference frames as unusable for reference and clearing them from the buffer memory to maintain synchronization with the encoder.
This approach differs from prior methods by providing a robust way to handle reference picture selection where frames might otherwise reference much older data. Unlike standard GOP structures that might have hidden dependencies, this method uses the identifier reset to provide an unambiguous signal to the decoder. It allows for more flexible streaming and trick-play modes, such as fast-forwarding or seeking, by giving the decoder a clear roadmap of where independent sub-sequences begin and end within a continuous stream.
In the early 2000s when ’321 was filed, video streaming systems were typically implemented using motion-compensated temporal prediction where sequences were organized into rigid Groups of Pictures (GOPs). At a time when systems commonly relied on fixed anchor frames like I-frames to reset decoding dependencies, the introduction of flexible reference picture selection created scenarios where groups of pictures were no longer temporally independent. During this era, hardware and software constraints made the detection of random-access points non-trivial, as decoders often had to parse and buffer extensive segments of the bitstream to identify dependencies and avoid decoding errors caused by missing reference frames.
The disclosed invention achieves a technical advancement by integrating an explicit initiation picture indicator within a video sequence to define an independently decodable sub-sequence. This architectural shift allows the decoder to identify a specific frame as a functional reset point, regardless of whether it follows a traditional I-frame or utilizes complex reference picture selection. By signaling this initiation picture—often through a header flag or a reset numbering scheme—the system enables the technical effect of immediate buffer clearance of prior, unusable reference frames. This capability overcomes the constraint of inter-sequence dependency, facilitating efficient random access, error recovery, and seamless splicing of disparate video streams without requiring full bitstream re-parsing.
The patent contains a total of 11 claims, with claims 1, 4, 7, 8, 10, and 11 serving as the independent claims. These independent claims focus on methods, encoders, decoders, and computer program products for processing video sequences by identifying independent sequences of image frames, managing motion-compensated temporal prediction references, and resetting frame identifier values according to a specific numbering scheme. The dependent claims serve to provide additional technical details regarding the placement of flags within slice headers and the encoding of specific identifier values for the independent sequences.
Definitions of key terms used in the patent claims.
US Latest litigation cases involving this patent.

The dossier documents provide a comprehensive record of the patent's prosecution history - including filings, correspondence, and decisions made by patent offices - and are crucial for understanding the patent's legal journey and any challenges it may have faced during examination.
Get instant alerts for new documents