Patent No. US10894974 (titled "Systems and methods to detect rare mutations and copy number variation") on Aug 4, 2017. The application was issued on Jan 19, 2021.
’974 is related to the field of molecular biology and bioinformatics, specifically focusing on the high-sensitivity detection of rare genetic mutations and copy number variations (CNVs) in cell-free DNA (cfDNA). The technology addresses the challenge of identifying low-frequency disease markers, such as those found in cancer or fetal abnormalities, which are often masked by the inherent noise and distortion introduced during the sequencing process.
The underlying idea behind ’974 is to maximize the conversion of limited biological samples into a digital library while using a non-unique tagging strategy to track individual molecules. By achieving high conversion efficiency (at least 20%) and using a specific molar excess of barcodes, the system ensures that even rare fragments are captured. The insight lies in using a relatively small set of barcodes that, when combined with the unique start and stop positions of naturally fragmented cfDNA, allow for the reconstruction of the original parent molecule’s sequence through consensus collapsing.
The claims of ’974 focus on a method for preparing and enriching cfDNA libraries by non-uniquely tagging double-stranded molecules with a limited set of 2 to 10,000 distinct barcodes. The process requires a ligation reaction using more than a 10× molar excess of tags to ensure high efficiency. Crucially, the claims specify that the number of different barcodes used is fewer than the number of DNA molecules mapping to a specific position, relying on the diversity of the fragments themselves to distinguish between parent polynucleotides.
In practice, the invention works by ligating barcodes to both ends of cfDNA fragments extracted from bodily fluids like blood or urine. These tagged molecules are then amplified to create a large pool of progeny. Because the system achieves high conversion, it can work with very low input amounts, such as less than 100 ng of DNA. The amplified pool is then selectively enriched for specific genomic regions of interest, such as actionable oncogenes or tumor suppressor genes, before being sequenced at high depth.
This approach differs from prior methods by moving away from the need for billions of unique identifiers to track every molecule. Instead, it leverages the stochastic fragmentation of cfDNA as a secondary identifier. By collapsing multiple sequence reads of the same family into a single consensus sequence, the system effectively filters out PCR-induced errors and sequencing artifacts. This allows for the detection of mutations at frequencies as low as 0.1%, providing a much clearer signal for early cancer detection and treatment monitoring than traditional analog sequencing.
In the early 2010s when ’974 was filed, the analysis of extracellular polynucleotides was typically implemented using sequencing protocols where the inherent error rates of the sequencing platform often exceeded the frequency of rare genetic variants. At a time when systems commonly relied on raw read counting for quantification, distinguishing true somatic mutations from stochastic sequencing noise was difficult, particularly when working with low-input volumes of cell-free DNA. Furthermore, hardware and software constraints made the high-fidelity reconstruction of original molecular templates non-trivial, as standard bioinformatics pipelines often lacked the means to identify and collapse PCR duplicates into accurate consensus sequences without losing the sensitivity required for sub-chromosomal resolution.
The disclosed invention represents a meaningful technical advancement through the integration of molecular tagging with a computational collapsing architecture that generates high-fidelity consensus sequences from amplified progeny reads. This architectural shift enables the system to overcome the technical constraint of sequencing-induced noise by using unique identifiers and mapping coordinates to group reads into families, allowing for the detection of rare mutations and copy number variations at sensitivities as low as 0.1%. The capability enabled by this method allows for the simultaneous quantification of fractional genetic alterations and rare variants from low-concentration samples, providing a robust framework for monitoring disease progression and tumor evolution through non-invasive bodily fluid analysis.
This patent contains 29 claims, with claims 1 and 18 serving as the independent claims. The independent claims focus on methods for preparing and generating sequencing libraries from cell-free nucleic acid molecules by non-uniquely tagging them with molecular barcodes through high-efficiency ligation, followed by amplification and selective enrichment of specific genomic regions. The dependent claims serve to specify various technical parameters and applications, including the types of bodily samples used, specific ranges for barcode diversity and molar excess, ligation efficiency thresholds, and the subsequent sequencing steps for detecting somatic variants associated with cancer.
Definitions of key terms used in the patent claims.
US Latest litigation cases involving this patent.

The dossier documents provide a comprehensive record of the patent's prosecution history - including filings, correspondence, and decisions made by patent offices - and are crucial for understanding the patent's legal journey and any challenges it may have faced during examination.
Get instant alerts for new documents