Patent No. US10250639 (titled "Information infrastructure management data processing tools for processing data flow with distribution controls") on Jan 15, 2015. The application was issued on Apr 2, 2019.
’639 is related to the field of distributed computing and information infrastructure management, specifically focusing on the automated organization, sanitization, and protection of data. It addresses the risks associated with unstructured and semi-structured data—such as emails and word processing documents—which often contain sensitive intellectual property or private information that is difficult to track compared to structured databases. The invention provides a framework for identifying these critical data elements and managing their lifecycle through configurable filtering and segmented storage.
The underlying idea behind ’639 is to transform a linear, vulnerable data stream into a formless, granular architecture by separating high-value content from its original context. Instead of relying on traditional perimeter defenses like firewalls, the system deconstructs documents into their atomic parts—words, images, or data objects—and disperses them across multiple distributed stores. By breaking the context of the original file and replacing sensitive elements with placeholders, the invention ensures that an intruder gaining access to one part of the system finds only fragmented, non-sensical information.
The claims of ’639 focus on a multi-stage method for extracting and classifying data throughput based on sensitivity levels and taxonomic categories. The independent claims cover the identification of sensitive or select content using a plurality of filters, followed by the extraction of this content into specific data stores corresponding to its sensitivity level. A key aspect of the claims is the application of classification tags to the extracted data, which then drive automated downstream processes such as data mining, supplemental searches, and structured data transfers to predetermined storage locations.
In practice, the system functions as a dynamic gatekeeper that sanitizes data inputs by separating sensitive content from remainder data to create a 'sanitized' version of the original file. This remainder data is then subjected to inference filtering, which uses content, contextual, and taxonomic analysis to ensure that no sensitive information can be reconstructed or deduced by unauthorized users. The implementation allows for a 'rolling' exposure of data, where layers of the original document are only reassembled and displayed to a user who provides the specific security clearances required for each granular piece.
This approach differs from prior solutions by moving away from static classification labels that can be manipulated by attackers. Instead, it employs granular data control to physically and logically isolate information, making the network 'formless' and significantly harder to target. While traditional encryption protects a whole file, this invention allows for the independent management of data segments, enabling an enterprise to share unclassified portions of a document while keeping the strategic 'dots' hidden in vaulted, distributed repositories until they are needed for authorized reconstruction.
In the late 2000s when ’639 was filed, enterprise information management was typically implemented using monolithic security architectures that focused on perimeter defense, such as firewalls, to protect structured databases. At a time when systems commonly relied on basic keyword indexing for search and retrieval, the vast majority of corporate data remained in unstructured or semi-structured formats—like word processing documents and emails—which were difficult to classify or secure programmatically. Hardware and software constraints of this era made the real-time semantic analysis and granular tracking of data across distributed storage environments non-trivial, often leading to a reliance on manual document labeling that could not keep pace with the rapid growth of digital data streams.
The disclosed invention represents a meaningful technical advancement through the integration of dynamic, adaptive categorical filters—including content-based, contextual, and taxonomic filters—that automatically identify and extract 'select content' from unstructured data inputs. This architectural shift moves away from static perimeter security toward a granular data control model where sensitive information is physically and logically separated from remainder data and stored in distributed, categorical data stores. The technical effect achieved is a transformation of data that enables automated policy enforcement—such as data destruction, archiving, or controlled distribution—based on the specific sensitivity and context of the content rather than its file location. This overcomes the constraint of information exposure by allowing for the controlled release of security-sensitive extracts only to authorized users while maintaining the utility of the non-sensitive remainder data.
This patent contains 18 claims, with claims 1 and 16 serving as the independent claims. The independent claims focus on methods for processing and sanitizing data throughput in distributed computing systems by identifying, extracting, and storing sensitive or select content based on sensitivity levels and taxonomic classifications. The dependent claims further specify the role of these processes by detailing access and release controls, encryption protocols, risk assessment measurements, and the application of specific content, contextual, and taxonomic filters to various data types.
Definitions of key terms used in the patent claims.
US Latest litigation cases involving this patent.

The dossier documents provide a comprehensive record of the patent's prosecution history - including filings, correspondence, and decisions made by patent offices - and are crucial for understanding the patent's legal journey and any challenges it may have faced during examination.
Get instant alerts for new documents