Video to data

Patent No. US10218954 (titled "Video to data") on Feb 7, 2015. The application was issued on Feb 26, 2019.

What is this patent about?

’954 is related to the field of automated video analysis and content generation. It addresses the technical challenge of extracting meaningful, context-aware metadata from video streams by simultaneously processing both auditory and visual data. The background context involves improving the accuracy of image recognition and speech transcription, which often suffer from false positives or a lack of semantic understanding when processed in isolation.

The underlying idea behind ’954 is the use of cross-modal synchronization to validate and refine data extracted from a video. By processing audio and image frames in parallel across a distributed architecture, the system uses the contextual topic derived from one stream (such as audio transcripts) to weight the probability of accuracy for objects identified in the other stream (the image frames). This creates a feedback loop where the audio context helps resolve visual ambiguities, such as distinguishing between a basketball and a soccer ball based on the surrounding commentary.

The claims of ’954 focus on a method and system that decompose a video into parallel processing pipelines for audio-to-text conversion and object identification. The independent claims specifically protect the mechanism of assigning a probability of accuracy to identified objects based on extracted contextual topics. This cross-referenced data is then used to generate new content—such as contextual text, images, or animations—which are integrated back into the original media to produce a content-rich video.

In practice, the invention functions by segmenting video and audio into small chunks that are handled by a distributed computation engine. While the audio pipeline performs speech recognition and keyword detection, the image pipeline performs edge detection and pattern recognition to identify faces, logos, or specific objects. The system then performs a temporal sync of these findings, using natural language processing to ensure that the metadata applied to a specific frame range is semantically consistent with what is being heard at that exact timestamp.

This approach differs from prior solutions by moving beyond simple keyword tagging or standalone image search. Instead of treating background elements as noise, the system uses them to build a probabilistic model of the scene's context. By leveraging the linear sequencing of video frames rather than static images, the invention can identify complex actions and reduce false positives, ultimately allowing for more precise advertisement targeting, automated closed captioning, and enhanced video search optimization.

How does this patent fit in bigger picture?

Technical Landscape

In the early 2010s when ’954 was filed, the extraction of metadata from multimedia content was typically implemented using isolated processing pipelines for distinct data types. At a time when systems commonly relied on independent audio-to-text transcription or basic reverse image searching to index video files, these processes often functioned as silos, where linguistic data from audio tracks remained decoupled from visual data within the frames. Hardware and software constraints made the real-time semantic synthesis of these disparate streams non-trivial, as computational resources were often focused on primary foreground object identification while treating background visual elements and non-speech audio as negligible noise.

Prosecution Position

The disclosed invention addresses the technical problem of semantic gaps and contextual inaccuracies in automated video indexing by implementing a multi-modal cross-referencing architecture. The solution involves the parallel segmentation and conversion of both image and audio data into text, followed by a cross-referencing operation that validates visual descriptors against topics extracted from the audio stream. This architectural shift enables the generation of high-confidence contextual descriptions and the identification of non-primary background elements that were previously filtered as noise. By integrating natural language processing with adjustable sensitivity thresholds for image detection, the system achieves a more granular and semantically coherent metadata set, enabling precise temporal placement of content-relevant information and recommendations.

Claims

The patent contains a total of 18 claims, with claims 1 and 13 serving as the independent claims. These independent claims focus on a method and a system for generating content-rich video data by processing audio and image files in parallel, identifying objects, and cross-referencing extracted text and topics to determine contextual accuracy. The dependent claims serve to further specify technical implementations such as natural language processing, image segmentation, motion detection, metadata generation, and the targeted placement of advertisements based on the extracted video data.

Key Claim Terms New

Definitions of key terms used in the patent claims.

Term (Source)Support for SpecificationInterpretation
Content-rich video
(Claim 1, Claim 13)
The present invention is generally directed to a method to generate data from video content, such as text and/or image-related information. The video text can be, for example, a context description of the video. An aspect of the method can include... generating video text based on a result of the cross-referencing.An enhanced video output that incorporates supplemental generated elements such as contextual text, images, or animations derived from the analysis of the original video's audio and visual data.
Contextual topic
(Claim 1, Claim 13)
The text from the image can be generated by identifying context, a symbol, a brand, a feature, an object, and/or a topic in the image of the video. The text can describe themes, identification of objects or other information of interest. The computational thresholds for identification of an object, face, etc. can be altered according to a then stated need or desire for non-primary, background, obstructed and/or grainy type images.A thematic subject or category derived from image files used to provide context for identifying specific objects within those images.
Cross-referencing
(Claim 1, Claim 13)
An aspect of the method can include... cross-referencing the text generated from the image of the video and the topics extracted from audio associated with the video, and generating video text based on a result of the cross-referencing. In some embodiments, natural language processing can be applied to the generation of text from an image of the video, converting audio associated with the video to text, or both.The process of comparing and correlating text derived from audio with data and topics derived from images to determine a unified context for the video.
Probability of accuracy
(Claim 1, Claim 13)
The size of the text can be adjusted by a ranking or scoring function that, for example, can adjust the text size based on confidence in the description or relevance to a search inquiry. To fine-tune the amount of object noise cluttering a data set, it can be useful to provide a user with an option to dial image detection sensitivity. Identification of only certain clearly identifiable faces or large unobstructed objects or band logos can be required with all other image noise disregarded or filtered.A confidence score or ranking assigned to an identified object based on its relevance to or consistency with the determined contextual topics.
Processing... in parallel
(Claim 1)
The text from the image can be generated by first segmenting images of the video, and then converting the segments of images to text in parallel. The text from the audio can be generated by first segmenting images of the audio, and then converting the segments of images to text in parallel.The simultaneous execution of audio and image data conversion tasks across multiple processors to improve efficiency.

Litigation Cases New

US Latest litigation cases involving this patent.

Case NumberFiling DateTitle
4:25-cv-01487Feb 13, 2025Cellular South Inc V. Google, Llc
6:24-cv-00245May 9, 2024Cellular South Inc V. Google, Llc

Patent Family

Patent Family

File Wrapper

The dossier documents provide a comprehensive record of the patent's prosecution history - including filings, correspondence, and decisions made by patent offices - and are crucial for understanding the patent's legal journey and any challenges it may have faced during examination.

  • Get instant alerts for new documents

US10218954

Application Number
US14910698A
Filing Date
Feb 7, 2015
Publication Date
Feb 26, 2019
External Links
Slate, USPTO , Google Patents