Video to data

Patent No. US9940972 (titled "Video to data") on Feb 7, 2014. The application was issued on Apr 10, 2018.

What is this patent about?

’972 is related to the field of automated video analysis and content generation. It addresses the technical challenge of efficiently extracting meaningful semantic information from large video files, which typically requires significant computational resources to process both visual and auditory streams simultaneously.

The underlying idea behind ’972 is the parallelized synchronization of disparate data streams to create a unified semantic map of a video. By splitting a video into discrete audio and image segments and distributing them across a cluster of processors, the system can perform object recognition and speech-to-text transcription in near real-time, subsequently cross-referencing these outputs to verify and weight the most relevant topics.

The claims of ’972 focus on a method and system that decomposes a video into audio and image files for parallel processing, specifically extracting and identifying objects to derive topical meta-data. The independent claims cover the specific workflow of generating video text through the cross-referencing of transcribed audio and identified visual objects, which is then used to dynamically inject new text, images, or animations back into the video stream.

In practice, the invention functions by segmenting audio at spectrum thresholds to identify quiet periods and dividing image frames into chunks for distributed processing. This allows the system to perform color edge detection and object template matching alongside natural language processing of the dialogue. By comparing the visual symbols found—such as a specific brand or object—with the topics extracted from the audio, the system assigns higher weights to confirmed themes, ensuring the generated meta-data accurately reflects the video's context.

This approach differs from prior art by moving beyond simple keyword transcription or basic image search. Instead of treating audio and video as isolated tracks, ’972 utilizes a multi-layered architecture to fuse visual and auditory semantics into a single data file. This enables sophisticated downstream applications, such as placing contextually relevant advertisements at precise timestamps or optimizing video search engines through deep, automated indexing of the actual content within the frames.

How does this patent fit in bigger picture?

Technical Landscape

In the early 2010s when ’972 was filed, video analysis was typically implemented using siloed processing pipelines where visual data and audio data were treated as independent streams. At a time when systems commonly relied on isolated optical character recognition or basic speech-to-text transcription, the extraction of high-level semantic meaning was often limited by the lack of integration between these disparate data types. Hardware and software constraints made the real-time synchronization and cross-referencing of multi-modal inputs non-trivial, often resulting in metadata that captured either the visual or the auditory context, but rarely a unified conceptual understanding of the video content.

Prosecution Position

The disclosed invention represents a technical advancement through the architectural integration of parallelized image segmentation and audio processing to generate contextual video metadata. By cross-referencing topics extracted from audio streams with symbolic and feature-based data derived from image segments, the system overcomes the limitation of semantic gaps inherent in single-source analysis. This multi-modal approach enables the generation of highly specific context descriptions and the precise temporal placement of content, such as advertisements, based on a unified interpretation of both visual symbols and linguistic topics.

Claims

This patent contains 20 claims, with claims 1 and 17 serving as the independent claims. The independent claims focus on a method and a system for generating video data by processing audio and image files in parallel to extract objects, convert audio to text, and generate topical metadata through semantic analysis to enhance the video with new text, images, or animations. The dependent claims serve to provide additional technical details regarding natural language processing, specific object and symbol recognition such as brand logos or faces, motion detection, audio-visual segmentation techniques, and the generation of targeted advertisements for placement within the video.

Key Claim Terms New

Definitions of key terms used in the patent claims.

Term (Source)Support for SpecificationInterpretation
Cross-referencing
(Claim 1, Claim 17)
An aspect of the method can include... cross-referencing the text generated from the image of the video and the topics extracted from audio associated with the video, and generating video text based on a result of the cross-referencing.The act of comparing and correlating text derived from audio with data derived from images to determine or verify the specific topics present in the video.
Segmenting image files
(Claim 17)
In yet other embodiments, the text from the image can be generated by first segmenting images of the video, and then converting the segments of images to text in parallel. The text from the audio can be generated by first segmenting images of the audio, and then converting the segments of images to text in parallel.The process of dividing video image data into discrete parts or sections to enable simultaneous processing across multiple computing resources.
Semantic information
(Claim 1, Claim 17)
Audio-to-text, however, lacks semantic and contextual language understanding. In some embodiments, natural language processing can be applied to the generation of text from an image of the video, converting audio associated with the video to text, or both.Meaning-based data derived from the contextual understanding of identified objects and audio, rather than just raw data or transcriptions.
Topical meta-data
(Claim 1)
The video text can be, for example, a context description of the video. An aspect of the method can include generating text from an image of the video, converting audio associated with the video to text, extracting topics from the text converted from the audio, cross-referencing the text generated from the image of the video and the topics extracted from audio associated with the video, and generating video text based on a result of the cross-referencing.A descriptive data set that characterizes the content of a video, created by synthesizing semantic information derived from both identified visual objects and audio content.
Video data
(Claim 1, Claim 17)
The text from the image of the video can be generated by identifying context, a symbol, a brand, a feature, an object, and/or a topic in the image of the video. The video content can be text, context, symbols, brands, features, objects, and/or topics related to or found in the video.Textual or descriptive information generated from the visual components of a video, such as identified symbols, brands, or features.

Litigation Cases New

US Latest litigation cases involving this patent.

Case NumberFiling DateTitle
4:25-cv-01487Feb 13, 2025Cellular South Inc V. Google, Llc
6:24-cv-00245May 9, 2024Cellular South Inc V. Google, Llc

Patent Family

Patent Family

File Wrapper

The dossier documents provide a comprehensive record of the patent's prosecution history - including filings, correspondence, and decisions made by patent offices - and are crucial for understanding the patent's legal journey and any challenges it may have faced during examination.

  • Get instant alerts for new documents

US9940972

Application Number
US14175741A
Filing Date
Feb 7, 2014
Publication Date
Apr 10, 2018
External Links
Slate, USPTO , Google Patents