Patent No. US9940972 (titled "Video to data") on Feb 7, 2014. The application was issued on Apr 10, 2018.
’972 is related to the field of automated video analysis and content generation. It addresses the technical challenge of efficiently extracting meaningful semantic information from large video files, which typically requires significant computational resources to process both visual and auditory streams simultaneously.
The underlying idea behind ’972 is the parallelized synchronization of disparate data streams to create a unified semantic map of a video. By splitting a video into discrete audio and image segments and distributing them across a cluster of processors, the system can perform object recognition and speech-to-text transcription in near real-time, subsequently cross-referencing these outputs to verify and weight the most relevant topics.
The claims of ’972 focus on a method and system that decomposes a video into audio and image files for parallel processing, specifically extracting and identifying objects to derive topical meta-data. The independent claims cover the specific workflow of generating video text through the cross-referencing of transcribed audio and identified visual objects, which is then used to dynamically inject new text, images, or animations back into the video stream.
In practice, the invention functions by segmenting audio at spectrum thresholds to identify quiet periods and dividing image frames into chunks for distributed processing. This allows the system to perform color edge detection and object template matching alongside natural language processing of the dialogue. By comparing the visual symbols found—such as a specific brand or object—with the topics extracted from the audio, the system assigns higher weights to confirmed themes, ensuring the generated meta-data accurately reflects the video's context.
This approach differs from prior art by moving beyond simple keyword transcription or basic image search. Instead of treating audio and video as isolated tracks, ’972 utilizes a multi-layered architecture to fuse visual and auditory semantics into a single data file. This enables sophisticated downstream applications, such as placing contextually relevant advertisements at precise timestamps or optimizing video search engines through deep, automated indexing of the actual content within the frames.
In the early 2010s when ’972 was filed, video analysis was typically implemented using siloed processing pipelines where visual data and audio data were treated as independent streams. At a time when systems commonly relied on isolated optical character recognition or basic speech-to-text transcription, the extraction of high-level semantic meaning was often limited by the lack of integration between these disparate data types. Hardware and software constraints made the real-time synchronization and cross-referencing of multi-modal inputs non-trivial, often resulting in metadata that captured either the visual or the auditory context, but rarely a unified conceptual understanding of the video content.
The disclosed invention represents a technical advancement through the architectural integration of parallelized image segmentation and audio processing to generate contextual video metadata. By cross-referencing topics extracted from audio streams with symbolic and feature-based data derived from image segments, the system overcomes the limitation of semantic gaps inherent in single-source analysis. This multi-modal approach enables the generation of highly specific context descriptions and the precise temporal placement of content, such as advertisements, based on a unified interpretation of both visual symbols and linguistic topics.
This patent contains 20 claims, with claims 1 and 17 serving as the independent claims. The independent claims focus on a method and a system for generating video data by processing audio and image files in parallel to extract objects, convert audio to text, and generate topical metadata through semantic analysis to enhance the video with new text, images, or animations. The dependent claims serve to provide additional technical details regarding natural language processing, specific object and symbol recognition such as brand logos or faces, motion detection, audio-visual segmentation techniques, and the generation of targeted advertisements for placement within the video.
Definitions of key terms used in the patent claims.
US Latest litigation cases involving this patent.

The dossier documents provide a comprehensive record of the patent's prosecution history - including filings, correspondence, and decisions made by patent offices - and are crucial for understanding the patent's legal journey and any challenges it may have faced during examination.
Get instant alerts for new documents