Information infrastructure management tools with extractor, secure storage, content analysis and classification and method therefor

Patent No. US9015301 (titled "Information infrastructure management tools with extractor, secure storage, content analysis and classification and method therefor") on May 9, 2007. The application was issued on Apr 21, 2015.

What is this patent about?

’301 is related to the field of information management and data processing within distributed computing systems. It specifically addresses the challenges of identifying, classifying, and securing sensitive or high-value information—referred to as select content—that often resides within unstructured or semi-structured data formats across an enterprise network.

The underlying idea behind ’301 is the use of an adaptive filtering framework that transforms raw data into organized, actionable intelligence by separating the essence of a document from its common container. By employing a combination of content-based, contextual, and taxonomic filters, the system creates an aggregated select content repository. This repository does not just store isolated keywords but captures the broader conceptual and relational meaning of the data, allowing the system to learn and refine its ability to handle future information based on previously identified patterns.

The claims of ’301 focus on a method and system for the automated management of data through the activation of categorical filters that trigger specific enterprise-defined actions. The independent claims describe a process where a data input is screened to extract select content and its associated taxonomic or contextual metadata, which is then stored in dedicated data stores. Once this intelligence is aggregated, the system applies a linked data process—such as extraction, archiving, or distribution limiting—to subsequent data inputs based on the results of the initial filtering and the security levels assigned to that content.

In practice, the invention functions as a dynamic gatekeeper that can be triggered manually or automatically by time, system conditions, or specific events like a security breach. When a document or data stream enters the system, it is deconstructed into granular elements. The system then uses its accumulated knowledge to decide whether to copy, extract, or destroy specific parts of the data. This granular data control ensures that sensitive information is physically or logically isolated from the remainder of the document, preventing unauthorized access while maintaining the utility of the non-sensitive portions.

This approach differs from prior solutions that rely heavily on static classification labels or perimeter-based security like firewalls. Traditional systems are often vulnerable if a label is tampered with or a boundary is breached. In contrast, ’301 implements formlessness by dispersing granular data segments across multiple distributed stores. Because the system manages the actual content and its context rather than just the file metadata, it can provide multi-level security where different users see different versions of the same document, reconstructed in real-time according to their specific clearance levels.

How does this patent fit in bigger picture?

Technical Landscape

In the mid-2000s when ’301 was filed, enterprise information management was typically implemented using monolithic security architectures that focused on perimeter defense, such as firewalls, to protect internal assets. At a time when systems commonly relied on basic keyword indexing and manual document classification, the handling of unstructured and semi-structured data—such as emails and word processing files—was often non-trivial due to the lack of automated semantic or taxonomic analysis. During this era, hardware and software constraints made the continuous, real-time monitoring and granular manipulation of data across distributed storage nodes difficult to achieve, often resulting in security vulnerabilities from internal actors or portable media.

Prosecution Position

The disclosed invention represents a meaningful technical advancement by shifting from static perimeter security to a dynamic, granular data management architecture. It addresses the technical problem of protecting sensitive information within unstructured data streams by integrating an automated inference engine with a multi-layered filtering system. This structural solution enables the decomposition of data into 'granular data streams'—comprising extracted sensitive content and remainder data—which are then dispersed across distributed storage nodes. The technical effect achieved is a transformation of data that prevents unauthorized reconstruction and mitigates inference attacks, enabling secure information sharing and automated policy enforcement across an open enterprise ecosystem.

Claims

This patent contains a total of 124 claims, with claims 1, 25, 48, 71, 77, 94, 101, and 118 serving as the independent claims. The independent claims focus on methods, computer-readable media, and distributed systems for organizing and processing enterprise data by utilizing categorical filters to identify select content and automatically trigger specific data processes such as copying, extracting, archiving, or distributing based on security levels. The dependent claims generally serve to refine these processes by specifying enterprise policy types, detailing data deconstruction and tagging techniques, defining specific storage media and access controls, and outlining server-client configurations for executing the filtering and processing tasks.

Key Claim Terms New

Definitions of key terms used in the patent claims.

Term (Source)Support for SpecificationInterpretation
Aggregated select content
(Claim 1, Claim 25, Claim 48, Claim 71, Claim 77, Claim 94, Claim 101, Claim 118)
A data input is processed through at least one activated categorical filter to obtain select content, and contextually associated select content and taxonomically associated select content as aggregated select content. The aggregated select content is stored in the corresponding select content data store. The system applies the associated data process to a further data input based upon a result of that further data being processed by the activated categorical filter utilizing the aggregated select content data.A collection of data resulting from a multi-layered filtering process that combines direct select content with its contextually and taxonomically related counterparts to form a comprehensive data set for policy application.
Contextual filters
(Claim 1, Claim 48, Claim 77, Claim 101)
The range defines an area within a document; data stream within that area key words for search will be located, selected and fed to search engines. Formulae and macros define ranges with informational content (contextual algorithms which link content), as well as indicate purpose and intent of the process as well as the target data. The system locates documents, which the are related to the initial document, each other by context or concept.Filters that identify and extract data based on the surrounding environment, related concepts, or algorithms that link content together rather than just matching specific keywords.
Data destruction process
(Claim 25, Claim 48, Claim 71, Claim 77, Claim 94, Claim 101, Claim 118)
A data process from the group of data processes including a copy process, a data extract process, a data archive process, a data distribution process and a data destruction process is associated with the activated categorical filter. It is a further object of the present invention to assist in data processing or manipulation including processes such as... data destruction (a document retention process). The system ensures that the information is properly handled, distributed, retained, deleted (document retention) and otherwise managed.A specific enterprise data management action, often part of a document retention policy, that ensures the permanent removal or deletion of data based on the results of categorical filtering.
Distribution security level
(Claim 1)
The enterprise designated filters screen data for enterprise policies such as a level of service policy, customer privacy policy, and document or data retention policy. A filter based on selecting a specific security level (Top Secret tagged content or Secret tagged content) may be used. The controlled release of corresponding extracted security sensitive data from the respective extract stores with the associated security clearances for corresponding security levels is permitted by the system.A classification or permission setting (e.g., Top Secret, Secret) assigned to select content that dictates how and to whom the data can be distributed or accessed.
Select content
(Claim 1, Claim 25, Claim 48, Claim 71, Claim 77, Claim 94, Claim 101, Claim 118)
The select content is represented by one or more predetermined words, characters, images, data elements or data objects. It includes tools for securing secret or security sensitive sec-con data in the enterprise computer system and to locate, identify and secure select content SC which may be of interest or importance to the enterprise. The system and process translates the sec-con or SC data and then stores the same in certain locations or secure stores.Specific data elements (words, characters, images, or objects) identified as important to an enterprise, often representing confidential information, intellectual property, or security-sensitive data that requires specialized handling.
Taxonomic classification filters
(Claim 1, Claim 48, Claim 77, Claim 101)
Semantic analysis, key word tagging and classification categorization (taxonomic analysis) should be conducted. Content classification with taxonomic classes occurs with tagging for formatting with bold, indexing, and paragraph marking, explicit element tagging for HTML and XML or database and spreadsheet table, field, ranges, row, and column designations. One axis may represent category groups (names, locations, and social security numbers).Filters that categorize information based on a hierarchical or systematic classification system (taxonomic analysis) to organize unstructured data into defined groups such as names, locations, or specific sensitive categories.

Litigation Cases New

US Latest litigation cases involving this patent.

Case NumberFiling DateTitle
2:25-cv-00595Apr 18, 2025Digitaldoors, Inc. v. SouthPoint Bank
8:25-cv-00002Jan 1, 2025Digital Doors, Inc. v. Sandy Spring Bank

Patent Family

Patent Family

File Wrapper

The dossier documents provide a comprehensive record of the patent's prosecution history - including filings, correspondence, and decisions made by patent offices - and are crucial for understanding the patent's legal journey and any challenges it may have faced during examination.

  • Get instant alerts for new documents

US9015301

Application Number
US11746440A
Filing Date
May 9, 2007
Publication Date
Apr 21, 2015
External Links
Slate, USPTO , Google Patents