Patent No. US10218954 (titled "Video to data") on Feb 7, 2015. The application was issued on Feb 26, 2019.
’954 is related to the field of automated video analysis and content generation. It addresses the technical challenge of extracting meaningful, context-aware metadata from video streams by simultaneously processing both auditory and visual data. The background context involves improving the accuracy of image recognition and speech transcription, which often suffer from false positives or a lack of semantic understanding when processed in isolation.
The underlying idea behind ’954 is the use of cross-modal synchronization to validate and refine data extracted from a video. By processing audio and image frames in parallel across a distributed architecture, the system uses the contextual topic derived from one stream (such as audio transcripts) to weight the probability of accuracy for objects identified in the other stream (the image frames). This creates a feedback loop where the audio context helps resolve visual ambiguities, such as distinguishing between a basketball and a soccer ball based on the surrounding commentary.
The claims of ’954 focus on a method and system that decompose a video into parallel processing pipelines for audio-to-text conversion and object identification. The independent claims specifically protect the mechanism of assigning a probability of accuracy to identified objects based on extracted contextual topics. This cross-referenced data is then used to generate new content—such as contextual text, images, or animations—which are integrated back into the original media to produce a content-rich video.
In practice, the invention functions by segmenting video and audio into small chunks that are handled by a distributed computation engine. While the audio pipeline performs speech recognition and keyword detection, the image pipeline performs edge detection and pattern recognition to identify faces, logos, or specific objects. The system then performs a temporal sync of these findings, using natural language processing to ensure that the metadata applied to a specific frame range is semantically consistent with what is being heard at that exact timestamp.
This approach differs from prior solutions by moving beyond simple keyword tagging or standalone image search. Instead of treating background elements as noise, the system uses them to build a probabilistic model of the scene's context. By leveraging the linear sequencing of video frames rather than static images, the invention can identify complex actions and reduce false positives, ultimately allowing for more precise advertisement targeting, automated closed captioning, and enhanced video search optimization.
In the early 2010s when ’954 was filed, the extraction of metadata from multimedia content was typically implemented using isolated processing pipelines for distinct data types. At a time when systems commonly relied on independent audio-to-text transcription or basic reverse image searching to index video files, these processes often functioned as silos, where linguistic data from audio tracks remained decoupled from visual data within the frames. Hardware and software constraints made the real-time semantic synthesis of these disparate streams non-trivial, as computational resources were often focused on primary foreground object identification while treating background visual elements and non-speech audio as negligible noise.
The disclosed invention addresses the technical problem of semantic gaps and contextual inaccuracies in automated video indexing by implementing a multi-modal cross-referencing architecture. The solution involves the parallel segmentation and conversion of both image and audio data into text, followed by a cross-referencing operation that validates visual descriptors against topics extracted from the audio stream. This architectural shift enables the generation of high-confidence contextual descriptions and the identification of non-primary background elements that were previously filtered as noise. By integrating natural language processing with adjustable sensitivity thresholds for image detection, the system achieves a more granular and semantically coherent metadata set, enabling precise temporal placement of content-relevant information and recommendations.
The patent contains a total of 18 claims, with claims 1 and 13 serving as the independent claims. These independent claims focus on a method and a system for generating content-rich video data by processing audio and image files in parallel, identifying objects, and cross-referencing extracted text and topics to determine contextual accuracy. The dependent claims serve to further specify technical implementations such as natural language processing, image segmentation, motion detection, metadata generation, and the targeted placement of advertisements based on the extracted video data.
Definitions of key terms used in the patent claims.
US Latest litigation cases involving this patent.

The dossier documents provide a comprehensive record of the patent's prosecution history - including filings, correspondence, and decisions made by patent offices - and are crucial for understanding the patent's legal journey and any challenges it may have faced during examination.
Get instant alerts for new documents