Grouping of image frames in video coding

Patent No. US8050321 (titled "Grouping of image frames in video coding") on Jan 25, 2006. The application was issued on Nov 1, 2011.

What is this patent about?

’321 is related to the field of digital video compression and streaming, specifically addressing the challenges of random access and error recovery within complex prediction hierarchies. In modern video codecs, temporal redundancy is reduced by predicting frames from previously decoded references; however, this creates interdependencies that make it difficult for a decoder to join a stream mid-way or recover from packet loss without decoding the entire preceding sequence.

The underlying idea behind ’321 is the explicit signaling of an initiation picture that marks the start of a strictly independent sequence of frames. By resetting the frame numbering scheme and providing a specific indicator at this boundary, the system ensures that no subsequent frames in that sequence will ever reference a picture that appeared before the initiation point. This creates a clean break in the dependency chain, allowing the decoder to flush its buffer and start fresh without risking the propagation of errors from missing prior data.

The claims of ’321 focus on a method and apparatus for encoding and decoding an independent sequence by embedding a specific indication of the first image frame in the decoding order. The invention specifically requires encoding identifier values according to a numbering scheme and then resetting the identifier value (typically to zero) for that first image frame. This mechanism ensures that all motion-compensated temporal prediction references within the sequence are self-contained and do not point to frames outside the defined boundary.

In practice, this is implemented by adding a separate flag or marker within the slice header of the video bitstream. When a user attempts to jump to a new position in a video file or joins a live broadcast, the decoder looks for this flag to identify a safe entry point. Once detected, the decoder treats the initiation picture as the new temporal anchor, effectively treating all previously stored reference frames as unusable for reference and clearing them from the buffer memory to maintain synchronization with the encoder.

This approach differs from prior methods by providing a robust way to handle reference picture selection where frames might otherwise reference much older data. Unlike standard GOP structures that might have hidden dependencies, this method uses the identifier reset to provide an unambiguous signal to the decoder. It allows for more flexible streaming and trick-play modes, such as fast-forwarding or seeking, by giving the decoder a clear roadmap of where independent sub-sequences begin and end within a continuous stream.

How does this patent fit in bigger picture?

Technical Landscape

In the early 2000s when ’321 was filed, video streaming systems were typically implemented using motion-compensated temporal prediction where sequences were organized into rigid Groups of Pictures (GOPs). At a time when systems commonly relied on fixed anchor frames like I-frames to reset decoding dependencies, the introduction of flexible reference picture selection created scenarios where groups of pictures were no longer temporally independent. During this era, hardware and software constraints made the detection of random-access points non-trivial, as decoders often had to parse and buffer extensive segments of the bitstream to identify dependencies and avoid decoding errors caused by missing reference frames.

Prosecution Position

The disclosed invention achieves a technical advancement by integrating an explicit initiation picture indicator within a video sequence to define an independently decodable sub-sequence. This architectural shift allows the decoder to identify a specific frame as a functional reset point, regardless of whether it follows a traditional I-frame or utilizes complex reference picture selection. By signaling this initiation picture—often through a header flag or a reset numbering scheme—the system enables the technical effect of immediate buffer clearance of prior, unusable reference frames. This capability overcomes the constraint of inter-sequence dependency, facilitating efficient random access, error recovery, and seamless splicing of disparate video streams without requiring full bitstream re-parsing.

Claims

The patent contains a total of 11 claims, with claims 1, 4, 7, 8, 10, and 11 serving as the independent claims. These independent claims focus on methods, encoders, decoders, and computer program products for processing video sequences by identifying independent sequences of image frames, managing motion-compensated temporal prediction references, and resetting frame identifier values according to a specific numbering scheme. The dependent claims serve to provide additional technical details regarding the placement of flags within slice headers and the encoding of specific identifier values for the independent sequences.

Key Claim Terms New

Definitions of key terms used in the patent claims.

Term (Source)Support for SpecificationInterpretation
Independent sequence
(Claim 1, Claim 4, Claim 7, Claim 8, Claim 10, Claim 11)
The invention is based on the idea of encoding a video sequence comprising an independent sequence of image frames, wherein at least one reference picture is predictable from at least one previous image frame that is earlier than the previous reference image frame in decoding order. After decoding the initiation picture, all following pictures of the independently decodable sequence can be decoded without prediction from any picture decoded prior to said initiation picture. This enables the decoder to start the decoding process without any prediction from any prior picture.A group of image frames where all motion-compensated temporal prediction references refer only to frames within that same group, allowing it to be decoded without information from frames outside the sequence.
Indication
(Claim 1, Claim 4, Claim 7, Claim 8, Claim 10, Claim 11)
An indication of at least one image frame is encoded into the video sequence, which indicated image frame is the first picture, in decoding order, of the independent sequence. According to an embodiment, the indication is encoded into the video sequence as a separate flag included in the header of a slice. This provides the decoder with the information about the first picture of an independently decodable sequence.A signal or flag encoded into the video bitstream that identifies a specific image frame as the starting point (initiation picture) of an independently decodable sequence.
Motion-compensated temporal prediction references
(Claim 1, Claim 4, Claim 7, Claim 8, Claim 10, Claim 11)
The contents of some (typically most) of the image frames in a video sequence are predicted from other frames in the sequence by tracking changes in specific objects or areas in successive image frames. A video sequence always comprises some compressed image frames the image information of which has not been determined using motion-compensated temporal prediction. Temporal prediction is typically carried out in video coding methods block- or macroblock-specifically, instead of image-frame-specifically.Data used to predict the contents of an image frame from other frames in a sequence by tracking changes in specific objects or areas over time.
Numbering scheme
(Claim 1, Claim 4, Claim 7, Claim 8, Claim 10, Claim 11)
According to an embodiment, identifier values for pictures are encoded according to a numbering scheme, and the identifier value for the indicated first picture of an independent sequence is reset, preferably to zero. The frames are typically numbered according to an arithmetical series. This enables identification of a picture boundary between two back-to-back initiation pictures by looking at the sub-sequence number of the initiation pictures.A system for assigning sequential identifier values to image frames to maintain decoding and display order.
Resetting the identifier value
(Claim 1, Claim 4, Claim 7, Claim 8, Claim 10, Claim 11)
According to an embodiment, identifier values for pictures are encoded according to a numbering scheme, and the identifier value for the indicated first picture of an independent sequence is reset, preferably to zero. This allows the decoder to detect the start of an independently decodable sequence. It also enables identification of a picture boundary between two back-to-back initiation pictures.The act of restarting the frame numbering sequence (typically to zero) at the first frame of an independent sequence to signal a new decodable boundary.

Litigation Cases New

US Latest litigation cases involving this patent.

Case NumberFiling DateTitle
1:25-cv-00523Apr 7, 2025Nokia Technologies Oy V. Acer Inc.
0:24-cv-04269Nov 25, 2024Element Television Company, Llc V. Nokia Corporation
1:23-cv-01237Oct 31, 2023Nokia Technologies Oy V. Hp, Inc.
1:23-cv-01236Oct 31, 2023Nokia Technologies Oy V. Amazon.Com, Inc.

Patent Family

Patent Family

File Wrapper

The dossier documents provide a comprehensive record of the patent's prosecution history - including filings, correspondence, and decisions made by patent offices - and are crucial for understanding the patent's legal journey and any challenges it may have faced during examination.

  • Get instant alerts for new documents

US8050321

SEP
Application Number
US11338934A
Filing Date
Jan 25, 2006
Status
Granted
Publication Date
Nov 1, 2011
External Links
Slate, USPTO , Google Patents