Accelerated mathematical engine

Patent No. US10671349 (titled "Accelerated mathematical engine") on Sep 20, 2017. The application was issued on Jun 2, 2020.

What is this patent about?

’349 is related to the field of accelerated mathematical engines, specifically hardware architectures designed to optimize the heavy computational load of convolution and deconvolution operations in neural networks. Traditional general-purpose processors often struggle with these tasks due to the overhead of constant memory fetching and the sequential nature of scalar arithmetic. The invention addresses these bottlenecks by providing a specialized hardware structure that maximizes data reuse and parallel execution.

The underlying idea behind ’349 is a two-dimensional systolic-like array of sub-circuits that performs massive parallel dot-product calculations in a single synchronized flow. By vectorizing image data and weights into specific formats that match the physical dimensions of the processor array, the system can execute thousands of multiplications and additions simultaneously. This approach minimizes the energy-intensive process of reading from SRAM by allowing a single piece of fetched data to be shared across multiple computational elements within the grid.

The claims of ’349 focus on a matrix processor architecture comprising a two-dimensional array of sub-circuits, where each sub-circuit integrates an ALU, an accumulator, and a shadow register. The independent claims specify a structure where a first input circuit feeds image data operands along one dimension while a second input circuit feeds weight operands along another. This configuration allows the hardware to convolve multiple input regions with multiple filters by cycling data through the array and capturing results without interrupting the next calculation cycle.

In practice, the invention utilizes in-line formatters to transform multi-dimensional image tensors and filter kernels into linear vectors that align with the rows and columns of the processor. A critical mechanical advantage is the use of the shadow register within each cell; this allows the accumulated result of a convolution to be stored and shifted out to the output array while the ALU immediately begins working on the next set of operands. This overlapping of computation and data transfer ensures that the mathematical engine remains at peak utilization, producing a new result every clock cycle.

This hardware-centric approach differs from prior solutions that rely on software-based matrix manipulation or standard GPU architectures that require frequent intermediate storage of partial products. By implementing booth encoders shared across rows and utilizing a clocked data-shifting mechanism, the ’349 patent reduces the managerial overhead of data movement. The result is a specialized engine that achieves higher throughput and lower power consumption by treating the convolution not as a series of discrete instructions, but as a continuous, parallelized flow of data through a dedicated physical grid.

How does this patent fit in bigger picture?

Technical Landscape

In the late 2010s when ’349 was filed, high-throughput mathematical processing was typically implemented using general-purpose computing architectures that relied on sequential software-driven matrix manipulations. At a time when systems commonly relied on iterative bit-by-bit addition and shifting operations rather than dedicated hardware-level matrix acceleration, the execution of complex convolutions required extensive use of general arithmetic units for intermediate address generation and data transpositions. Hardware constraints made the real-time processing of large data sets non-trivial, as conventional designs frequently encountered bottlenecks caused by the repeated cycles of storing and fetching intermediate results from memory to complete multi-stage operations.

Prosecution Position

The disclosed architecture represents a technical advancement by integrating an accelerated mathematical engine specifically designed to execute complex convolution operations through direct matrix multiplication. This architectural shift overcomes the latency constraints of general-purpose processors by reducing the reliance on software-mediated data manipulation steps and the associated overhead of intermediate result storage. By implementing a hardware-level solution for matrix operations, the system achieves a higher rate of calculation and improved computational efficiency, enabling the processing of large data volumes without the performance bottlenecks inherent in traditional arithmetic logic unit cycles.

Claims

This patent includes a total of 21 claims, with independent claims 1, 8, and 17 establishing the foundational aspects of the technology. The independent claims focus on a matrix processor, a system, and a method for accelerating neural network convolutions using a two-dimensional array of sub-circuits equipped with arithmetic logic units, accumulators, and shadow registers to perform dot-product calculations between input data regions and filter weights. The dependent claims serve to provide specific technical refinements, such as defining the use of multiply-and-add circuits, Booth encoders, state machines for identifying reusable data, specific register configurations for data throughput, and post-processing steps like non-linear functions or pooling within the convolution layer.

Key Claim Terms New

Definitions of key terms used in the patent claims.

Term (Source)Support for SpecificationInterpretation
Dot product calculation
(Claim 1, Claim 8, Claim 17)
The sub-circuits are coupled to receive N and M operands to perform a dot product calculation, such that an input region is convolved with a filter. The plurality of sub-circuits are configured to perform dot-multiplications to generate dot-products. These calculations are used to generate a result in the output array, accelerating convolutions in a neural network.An arithmetic operation performed by sub-circuits involving the multiplication of corresponding entries of two sequences of operands (e.g., input data and filter weights) and the summation of those products.
Filter
(Claim 1, Claim 8, Claim 17)
The M operands represent first weight values of a first filter of a plurality of filters. The sub-circuits are configured to receive second weight values of a second filter from the second input circuit. The method involves convolving input regions with a first filter and subsequently convolving a second filter with the first input region.A set of weight values (M operands) representing a kernel in a convolutional neural network that is applied to input data to produce a convolution result.
Input region
(Claim 1, Claim 8, Claim 17)
The N operands represent a first input region of a plurality of input regions obtained from image data. The system is configured such that the first input region is convolved with a first filter, and remaining input regions are subsequently received and convolved. This allows the engine to operate on large amounts of data by breaking image data into manageable regions for matrix multiplication.A specific subset or block of image data operands extracted from a larger data matrix to be processed by the matrix engine.
Shadow register
(Claim 1, Claim 8, Claim 17)
The specification describes sub-circuits within a two-dimensional array that comprise an arithmetic logic unit, an accumulator, and a shadow register. These components are coupled to perform dot product calculations of N and M operands. The shadow register facilitates the handling of weight values from different filters during the convolution process.A storage element within a sub-circuit of a matrix processor used to hold or buffer operands, such as weight values, allowing the sub-circuit to perform calculations on one set of data while potentially preparing or maintaining another.
Two-dimensional array
(Claim 1, Claim 8)
The matrix processor includes a first input circuit arranged in a first dimension and a second input circuit arranged in a second dimension of a two-dimensional array. The sub-circuits are coupled within this two-dimensional array to perform dot product calculations. This structure allows for the acceleration of convolutions by processing input regions and filter weights across the array dimensions.A hardware architecture of a matrix processor organized into rows and columns (dimensions) containing sub-circuits for parallel processing of operands.

Litigation Cases New

US Latest litigation cases involving this patent.

Case NumberFiling DateTitle
2:25-cv-00742Jul 23, 2025Perceptive Automata Llc V. Tesla, Inc.

Patent Family

Patent Family

File Wrapper

The dossier documents provide a comprehensive record of the patent's prosecution history - including filings, correspondence, and decisions made by patent offices - and are crucial for understanding the patent's legal journey and any challenges it may have faced during examination.

  • Get instant alerts for new documents

US10671349

Application Number
US15710433A
Filing Date
Sep 20, 2017
Publication Date
Jun 2, 2020
External Links
Slate, USPTO , Google Patents