Patent No. US9116809 (titled "Memory heaps in a memory model for a unified computing system") on Dec 21, 2012. The application was issued on Aug 25, 2015.
’809 is related to the field of heterogeneous computing architectures, specifically addressing the challenges of memory management in systems where a central processing unit and an accelerated processing device, such as a GPU, operate in a unified environment. Traditionally, these processors maintained separate address spaces, forcing programmers to manually move data between them, which created significant overhead and complexity. The invention seeks to streamline this by providing a more fluid, integrated approach to how these diverse processing units access and share physical memory resources.
The underlying idea behind ’809 is the implementation of a shared memory address space that utilizes specialized memory heaps to abstract the underlying physical hardware. Instead of requiring explicit data transfers, the system uses a mapper to direct memory operations to specific virtual pools—or heaps—based on a shared address. This allows pointers to be passed between the CPU and the accelerator without modification, effectively treating the disparate memory controllers and physical RAM modules as a single, cohesive resource with defined access properties.
The claims of ’809 focus on a method and system for mapping memory operations to one of a plurality of memory heaps based on a shared memory address (SMA). The independent claims highlight a mechanism where a mapper receives an address reference and determines the appropriate heap, which in turn dictates how the operation interacts with physical memory resources. Crucially, the claims specify that different memory attributes or access permissions can be applied to these heaps depending on which processor originated the request, allowing for fine-grained control over coherency and visibility.
In practice, the invention works by categorizing memory into functional types such as local, global, scratch, or coherent heaps. For instance, a coherent memory heap allows both the CPU and the accelerator to access the same data, with the system automatically handling the translation through an IOMMU or a system memory manager. Other heaps might be private to specific work-groups or replicated across threads to act as an extension of processor registers. This hierarchy ensures that high-performance graphics memory and high-capacity system memory can be utilized simultaneously and efficiently.
This approach differentiates itself from prior solutions by eliminating the need for the programmer to explicitly marshal data between separate device memories. By using a unified memory model where the mapping result is provided directly to the requesting processor, the system hides the complexity of physical memory location and hardware-specific protocols. Unlike conventional systems that treat the GPU as a peripheral with a disconnected memory space, this invention enables the accelerator to participate in the system's paging and protection schemes as a first-class citizen.
In the early 2010s when ’809 was filed, heterogeneous computing environments were typically implemented using discrete processing units that maintained isolated memory domains. At a time when central processing units and accelerated processing devices commonly relied on separate physical or logical address spaces, software developers were required to manually manage data movement and explicit memory marshalling between these distinct regions. Hardware and software constraints made the fluid sharing of data non-trivial, as the lack of a unified memory architecture necessitated complex synchronization and overhead-intensive copying operations to ensure data consistency across different instruction set architectures.
The disclosed invention achieves a technical advancement by establishing a unified memory architecture that integrates disparate processing clients into a single memory space with common access and storage properties. By implementing a structural solution that maps memory operations from multiple processors to specific memory heaps within a unified framework, the architecture overcomes the constraint of manual memory management. This shift enables a capability for seamless data sharing between central and accelerated processing units, reducing the computational overhead associated with data duplication and providing a consistent programming model across heterogeneous hardware components.
This patent contains 25 claims, with claims 1, 12, and 20 serving as the independent claims. The independent claims focus on a method, system, and computer-readable medium for allocating memory in a multi-processor environment by mapping memory operations to specific memory heaps based on shared memory addresses or address references. The dependent claims provide additional detail regarding physical memory resource associations, accessibility restrictions for specific processors or applications, the use of page tables, and the specific types of processors involved, such as CPUs and APDs.
Definitions of key terms used in the patent claims.
US Latest litigation cases involving this patent.

The dossier documents provide a comprehensive record of the patent's prosecution history - including filings, correspondence, and decisions made by patent offices - and are crucial for understanding the patent's legal journey and any challenges it may have faced during examination.
Get instant alerts for new documents