Information infrastructure management data processing tools for processing data flow with distribution controls

Patent No. US10250639 (titled "Information infrastructure management data processing tools for processing data flow with distribution controls") on Jan 15, 2015. The application was issued on Apr 2, 2019.

What is this patent about?

’639 is related to the field of distributed computing and information infrastructure management, specifically focusing on the automated organization, sanitization, and protection of data. It addresses the risks associated with unstructured and semi-structured data—such as emails and word processing documents—which often contain sensitive intellectual property or private information that is difficult to track compared to structured databases. The invention provides a framework for identifying these critical data elements and managing their lifecycle through configurable filtering and segmented storage.

The underlying idea behind ’639 is to transform a linear, vulnerable data stream into a formless, granular architecture by separating high-value content from its original context. Instead of relying on traditional perimeter defenses like firewalls, the system deconstructs documents into their atomic parts—words, images, or data objects—and disperses them across multiple distributed stores. By breaking the context of the original file and replacing sensitive elements with placeholders, the invention ensures that an intruder gaining access to one part of the system finds only fragmented, non-sensical information.

The claims of ’639 focus on a multi-stage method for extracting and classifying data throughput based on sensitivity levels and taxonomic categories. The independent claims cover the identification of sensitive or select content using a plurality of filters, followed by the extraction of this content into specific data stores corresponding to its sensitivity level. A key aspect of the claims is the application of classification tags to the extracted data, which then drive automated downstream processes such as data mining, supplemental searches, and structured data transfers to predetermined storage locations.

In practice, the system functions as a dynamic gatekeeper that sanitizes data inputs by separating sensitive content from remainder data to create a 'sanitized' version of the original file. This remainder data is then subjected to inference filtering, which uses content, contextual, and taxonomic analysis to ensure that no sensitive information can be reconstructed or deduced by unauthorized users. The implementation allows for a 'rolling' exposure of data, where layers of the original document are only reassembled and displayed to a user who provides the specific security clearances required for each granular piece.

This approach differs from prior solutions by moving away from static classification labels that can be manipulated by attackers. Instead, it employs granular data control to physically and logically isolate information, making the network 'formless' and significantly harder to target. While traditional encryption protects a whole file, this invention allows for the independent management of data segments, enabling an enterprise to share unclassified portions of a document while keeping the strategic 'dots' hidden in vaulted, distributed repositories until they are needed for authorized reconstruction.

How does this patent fit in bigger picture?

Technical Landscape

In the late 2000s when ’639 was filed, enterprise information management was typically implemented using monolithic security architectures that focused on perimeter defense, such as firewalls, to protect structured databases. At a time when systems commonly relied on basic keyword indexing for search and retrieval, the vast majority of corporate data remained in unstructured or semi-structured formats—like word processing documents and emails—which were difficult to classify or secure programmatically. Hardware and software constraints of this era made the real-time semantic analysis and granular tracking of data across distributed storage environments non-trivial, often leading to a reliance on manual document labeling that could not keep pace with the rapid growth of digital data streams.

Prosecution Position

The disclosed invention represents a meaningful technical advancement through the integration of dynamic, adaptive categorical filters—including content-based, contextual, and taxonomic filters—that automatically identify and extract 'select content' from unstructured data inputs. This architectural shift moves away from static perimeter security toward a granular data control model where sensitive information is physically and logically separated from remainder data and stored in distributed, categorical data stores. The technical effect achieved is a transformation of data that enables automated policy enforcement—such as data destruction, archiving, or controlled distribution—based on the specific sensitivity and context of the content rather than its file location. This overcomes the constraint of information exposure by allowing for the controlled release of security-sensitive extracts only to authorized users while maintaining the utility of the non-sensitive remainder data.

Claims

This patent contains 18 claims, with claims 1 and 16 serving as the independent claims. The independent claims focus on methods for processing and sanitizing data throughput in distributed computing systems by identifying, extracting, and storing sensitive or select content based on sensitivity levels and taxonomic classifications. The dependent claims further specify the role of these processes by detailing access and release controls, encryption protocols, risk assessment measurements, and the application of specific content, contextual, and taxonomic filters to various data types.

Key Claim Terms New

Definitions of key terms used in the patent claims.

Term (Source)Support for SpecificationInterpretation
Classification tags
(Claim 1)
The information classification system automatically categorizes information in unstructured information files and labels the same. These tags or labels change based on results of content inferencing penetration testing. One axis of a data matrix may represent category groups like names, locations, and social security numbers.Metadata labels or markers generated by filters and associated with extracted data to enable structured processing, such as data mining, transfer, or specific security handling.
Inferencing
(Claim 16)
The objective is to locate hidden data and to infer data therefrom that is identifiable and relevant. The system implements an inference process to verify if anything can be inferenced from the sanitized data stream. It focuses on 'Unknown' data not identified by the initial set of filters.A process of analyzing sanitized or residual data using filters to identify hidden relationships, context, or potential security threats that were not caught by initial extractions.
Remainder data
(Claim 16)
A granular data stream is defined as the extract and/or remainder data filtered from an original data stream. The system deals with a situation where a whole data asset was already parsed and split into a remainder and extracts. Remainder data is stored in the distributed computer system.The residual portion of an original data stream that exists after sensitive or select content has been identified and extracted.
Select content
(Claim 1, Claim 16)
Select content (SC) is represented by one or more predetermined words, characters, images, data elements or data objects. The system employs a dynamic, adaptive filter to enhance SC collection and classification systems to organize such SC. It is stored in select content data stores apart from security sensitive content.Specific data elements of interest to an enterprise, defined by predetermined characters or objects, which are identified, classified, and managed separately from sensitive content to enhance organizational data value.
Sensitive content
(Claim 1, Claim 16)
The invention relates to identifying sensitive-secret or select data content and extracting key content. It includes securing secret or security sensitive sec-con data in the enterprise computer system. This content is extracted from a data input to obtain extracted security sensitive data for a corresponding security level and remainder data.Confidential or secret data within a data stream, represented by specific words, characters, images, or objects, which is categorized into multiple sensitivity levels for protected storage and controlled access.
Sensitivity levels
(Claim 1, Claim 16)
Sensitivity levels can be applied based upon time issues, competitor or size of company, or type of product. The system selects items with a high sensitivity level tagging, such as Top Secret, for specific filtering. Security levels of a document or data stream are upgraded or downgraded based on the results of inference tests.Hierarchical or orthogonal security classifications (e.g., Top Secret, Secret) assigned to specific words, characters, or data objects within a data stream to control access and storage.
Supplemental data search process
(Claim 1)
Search results are filtered again to find new keywords so another search will take place. The user has the ability to set the system for a continuous, non-stop cycle of filtering keywords and feeding them to search engines. This creates a growing tree of digital data streams around a content target.An automated, iterative search operation that uses filtered keywords or results from a primary search to trigger subsequent search cycles to refine or expand data.
Taxonomic category filter
(Claim 1, Claim 16)
Taxonomic classification filters are operatively coupled over a communications network to process data inputs. The system analyzes every word, character, icon, and image and categorizes them based on a predetermined rule set. One axis of a data matrix may represent category groups like names, locations, and social security numbers.A classification engine that organizes data into hierarchical structures or categories (such as names, locations, or social security numbers) and generates associated tags.

Litigation Cases New

US Latest litigation cases involving this patent.

Case NumberFiling DateTitle
2:25-cv-00595Apr 18, 2025Digitaldoors, Inc. v. SouthPoint Bank
8:25-cv-00002Jan 1, 2025Digital Doors, Inc. v. Sandy Spring Bank

Patent Family

Patent Family

File Wrapper

The dossier documents provide a comprehensive record of the patent's prosecution history - including filings, correspondence, and decisions made by patent offices - and are crucial for understanding the patent's legal journey and any challenges it may have faced during examination.

  • Get instant alerts for new documents

US10250639

Application Number
US14597345A
Filing Date
Jan 15, 2015
Publication Date
Apr 2, 2019
External Links
Slate, USPTO , Google Patents