Patent No. US9015301 (titled "Information infrastructure management tools with extractor, secure storage, content analysis and classification and method therefor") on May 9, 2007. The application was issued on Apr 21, 2015.
’301 is related to the field of information management and data processing within distributed computing systems. It specifically addresses the challenges of identifying, classifying, and securing sensitive or high-value information—referred to as select content—that often resides within unstructured or semi-structured data formats across an enterprise network.
The underlying idea behind ’301 is the use of an adaptive filtering framework that transforms raw data into organized, actionable intelligence by separating the essence of a document from its common container. By employing a combination of content-based, contextual, and taxonomic filters, the system creates an aggregated select content repository. This repository does not just store isolated keywords but captures the broader conceptual and relational meaning of the data, allowing the system to learn and refine its ability to handle future information based on previously identified patterns.
The claims of ’301 focus on a method and system for the automated management of data through the activation of categorical filters that trigger specific enterprise-defined actions. The independent claims describe a process where a data input is screened to extract select content and its associated taxonomic or contextual metadata, which is then stored in dedicated data stores. Once this intelligence is aggregated, the system applies a linked data process—such as extraction, archiving, or distribution limiting—to subsequent data inputs based on the results of the initial filtering and the security levels assigned to that content.
In practice, the invention functions as a dynamic gatekeeper that can be triggered manually or automatically by time, system conditions, or specific events like a security breach. When a document or data stream enters the system, it is deconstructed into granular elements. The system then uses its accumulated knowledge to decide whether to copy, extract, or destroy specific parts of the data. This granular data control ensures that sensitive information is physically or logically isolated from the remainder of the document, preventing unauthorized access while maintaining the utility of the non-sensitive portions.
This approach differs from prior solutions that rely heavily on static classification labels or perimeter-based security like firewalls. Traditional systems are often vulnerable if a label is tampered with or a boundary is breached. In contrast, ’301 implements formlessness by dispersing granular data segments across multiple distributed stores. Because the system manages the actual content and its context rather than just the file metadata, it can provide multi-level security where different users see different versions of the same document, reconstructed in real-time according to their specific clearance levels.
In the mid-2000s when ’301 was filed, enterprise information management was typically implemented using monolithic security architectures that focused on perimeter defense, such as firewalls, to protect internal assets. At a time when systems commonly relied on basic keyword indexing and manual document classification, the handling of unstructured and semi-structured data—such as emails and word processing files—was often non-trivial due to the lack of automated semantic or taxonomic analysis. During this era, hardware and software constraints made the continuous, real-time monitoring and granular manipulation of data across distributed storage nodes difficult to achieve, often resulting in security vulnerabilities from internal actors or portable media.
The disclosed invention represents a meaningful technical advancement by shifting from static perimeter security to a dynamic, granular data management architecture. It addresses the technical problem of protecting sensitive information within unstructured data streams by integrating an automated inference engine with a multi-layered filtering system. This structural solution enables the decomposition of data into 'granular data streams'—comprising extracted sensitive content and remainder data—which are then dispersed across distributed storage nodes. The technical effect achieved is a transformation of data that prevents unauthorized reconstruction and mitigates inference attacks, enabling secure information sharing and automated policy enforcement across an open enterprise ecosystem.
This patent contains a total of 124 claims, with claims 1, 25, 48, 71, 77, 94, 101, and 118 serving as the independent claims. The independent claims focus on methods, computer-readable media, and distributed systems for organizing and processing enterprise data by utilizing categorical filters to identify select content and automatically trigger specific data processes such as copying, extracting, archiving, or distributing based on security levels. The dependent claims generally serve to refine these processes by specifying enterprise policy types, detailing data deconstruction and tagging techniques, defining specific storage media and access controls, and outlining server-client configurations for executing the filtering and processing tasks.
Definitions of key terms used in the patent claims.
US Latest litigation cases involving this patent.

The dossier documents provide a comprehensive record of the patent's prosecution history - including filings, correspondence, and decisions made by patent offices - and are crucial for understanding the patent's legal journey and any challenges it may have faced during examination.
Get instant alerts for new documents