Patent No. US10671349 (titled "Accelerated mathematical engine") on Sep 20, 2017. The application was issued on Jun 2, 2020.
’349 is related to the field of accelerated mathematical engines, specifically hardware architectures designed to optimize the heavy computational load of convolution and deconvolution operations in neural networks. Traditional general-purpose processors often struggle with these tasks due to the overhead of constant memory fetching and the sequential nature of scalar arithmetic. The invention addresses these bottlenecks by providing a specialized hardware structure that maximizes data reuse and parallel execution.
The underlying idea behind ’349 is a two-dimensional systolic-like array of sub-circuits that performs massive parallel dot-product calculations in a single synchronized flow. By vectorizing image data and weights into specific formats that match the physical dimensions of the processor array, the system can execute thousands of multiplications and additions simultaneously. This approach minimizes the energy-intensive process of reading from SRAM by allowing a single piece of fetched data to be shared across multiple computational elements within the grid.
The claims of ’349 focus on a matrix processor architecture comprising a two-dimensional array of sub-circuits, where each sub-circuit integrates an ALU, an accumulator, and a shadow register. The independent claims specify a structure where a first input circuit feeds image data operands along one dimension while a second input circuit feeds weight operands along another. This configuration allows the hardware to convolve multiple input regions with multiple filters by cycling data through the array and capturing results without interrupting the next calculation cycle.
In practice, the invention utilizes in-line formatters to transform multi-dimensional image tensors and filter kernels into linear vectors that align with the rows and columns of the processor. A critical mechanical advantage is the use of the shadow register within each cell; this allows the accumulated result of a convolution to be stored and shifted out to the output array while the ALU immediately begins working on the next set of operands. This overlapping of computation and data transfer ensures that the mathematical engine remains at peak utilization, producing a new result every clock cycle.
This hardware-centric approach differs from prior solutions that rely on software-based matrix manipulation or standard GPU architectures that require frequent intermediate storage of partial products. By implementing booth encoders shared across rows and utilizing a clocked data-shifting mechanism, the ’349 patent reduces the managerial overhead of data movement. The result is a specialized engine that achieves higher throughput and lower power consumption by treating the convolution not as a series of discrete instructions, but as a continuous, parallelized flow of data through a dedicated physical grid.
In the late 2010s when ’349 was filed, high-throughput mathematical processing was typically implemented using general-purpose computing architectures that relied on sequential software-driven matrix manipulations. At a time when systems commonly relied on iterative bit-by-bit addition and shifting operations rather than dedicated hardware-level matrix acceleration, the execution of complex convolutions required extensive use of general arithmetic units for intermediate address generation and data transpositions. Hardware constraints made the real-time processing of large data sets non-trivial, as conventional designs frequently encountered bottlenecks caused by the repeated cycles of storing and fetching intermediate results from memory to complete multi-stage operations.
The disclosed architecture represents a technical advancement by integrating an accelerated mathematical engine specifically designed to execute complex convolution operations through direct matrix multiplication. This architectural shift overcomes the latency constraints of general-purpose processors by reducing the reliance on software-mediated data manipulation steps and the associated overhead of intermediate result storage. By implementing a hardware-level solution for matrix operations, the system achieves a higher rate of calculation and improved computational efficiency, enabling the processing of large data volumes without the performance bottlenecks inherent in traditional arithmetic logic unit cycles.
This patent includes a total of 21 claims, with independent claims 1, 8, and 17 establishing the foundational aspects of the technology. The independent claims focus on a matrix processor, a system, and a method for accelerating neural network convolutions using a two-dimensional array of sub-circuits equipped with arithmetic logic units, accumulators, and shadow registers to perform dot-product calculations between input data regions and filter weights. The dependent claims serve to provide specific technical refinements, such as defining the use of multiply-and-add circuits, Booth encoders, state machines for identifying reusable data, specific register configurations for data throughput, and post-processing steps like non-linear functions or pooling within the convolution layer.
Definitions of key terms used in the patent claims.
US Latest litigation cases involving this patent.

The dossier documents provide a comprehensive record of the patent's prosecution history - including filings, correspondence, and decisions made by patent offices - and are crucial for understanding the patent's legal journey and any challenges it may have faced during examination.
Get instant alerts for new documents