Systems and methods to detect rare mutations and copy number variation

Patent No. US10894974 (titled "Systems and methods to detect rare mutations and copy number variation") on Aug 4, 2017. The application was issued on Jan 19, 2021.

What is this patent about?

’974 is related to the field of molecular biology and bioinformatics, specifically focusing on the high-sensitivity detection of rare genetic mutations and copy number variations (CNVs) in cell-free DNA (cfDNA). The technology addresses the challenge of identifying low-frequency disease markers, such as those found in cancer or fetal abnormalities, which are often masked by the inherent noise and distortion introduced during the sequencing process.

The underlying idea behind ’974 is to maximize the conversion of limited biological samples into a digital library while using a non-unique tagging strategy to track individual molecules. By achieving high conversion efficiency (at least 20%) and using a specific molar excess of barcodes, the system ensures that even rare fragments are captured. The insight lies in using a relatively small set of barcodes that, when combined with the unique start and stop positions of naturally fragmented cfDNA, allow for the reconstruction of the original parent molecule’s sequence through consensus collapsing.

The claims of ’974 focus on a method for preparing and enriching cfDNA libraries by non-uniquely tagging double-stranded molecules with a limited set of 2 to 10,000 distinct barcodes. The process requires a ligation reaction using more than a 10× molar excess of tags to ensure high efficiency. Crucially, the claims specify that the number of different barcodes used is fewer than the number of DNA molecules mapping to a specific position, relying on the diversity of the fragments themselves to distinguish between parent polynucleotides.

In practice, the invention works by ligating barcodes to both ends of cfDNA fragments extracted from bodily fluids like blood or urine. These tagged molecules are then amplified to create a large pool of progeny. Because the system achieves high conversion, it can work with very low input amounts, such as less than 100 ng of DNA. The amplified pool is then selectively enriched for specific genomic regions of interest, such as actionable oncogenes or tumor suppressor genes, before being sequenced at high depth.

This approach differs from prior methods by moving away from the need for billions of unique identifiers to track every molecule. Instead, it leverages the stochastic fragmentation of cfDNA as a secondary identifier. By collapsing multiple sequence reads of the same family into a single consensus sequence, the system effectively filters out PCR-induced errors and sequencing artifacts. This allows for the detection of mutations at frequencies as low as 0.1%, providing a much clearer signal for early cancer detection and treatment monitoring than traditional analog sequencing.

How does this patent fit in bigger picture?

Technical Landscape

In the early 2010s when ’974 was filed, the analysis of extracellular polynucleotides was typically implemented using sequencing protocols where the inherent error rates of the sequencing platform often exceeded the frequency of rare genetic variants. At a time when systems commonly relied on raw read counting for quantification, distinguishing true somatic mutations from stochastic sequencing noise was difficult, particularly when working with low-input volumes of cell-free DNA. Furthermore, hardware and software constraints made the high-fidelity reconstruction of original molecular templates non-trivial, as standard bioinformatics pipelines often lacked the means to identify and collapse PCR duplicates into accurate consensus sequences without losing the sensitivity required for sub-chromosomal resolution.

Prosecution Position

The disclosed invention represents a meaningful technical advancement through the integration of molecular tagging with a computational collapsing architecture that generates high-fidelity consensus sequences from amplified progeny reads. This architectural shift enables the system to overcome the technical constraint of sequencing-induced noise by using unique identifiers and mapping coordinates to group reads into families, allowing for the detection of rare mutations and copy number variations at sensitivities as low as 0.1%. The capability enabled by this method allows for the simultaneous quantification of fractional genetic alterations and rare variants from low-concentration samples, providing a robust framework for monitoring disease progression and tumor evolution through non-invasive bodily fluid analysis.

Claims

This patent contains 29 claims, with claims 1 and 18 serving as the independent claims. The independent claims focus on methods for preparing and generating sequencing libraries from cell-free nucleic acid molecules by non-uniquely tagging them with molecular barcodes through high-efficiency ligation, followed by amplification and selective enrichment of specific genomic regions. The dependent claims serve to specify various technical parameters and applications, including the types of bodily samples used, specific ranges for barcode diversity and molar excess, ligation efficiency thresholds, and the subsequent sequencing steps for detecting somatic variants associated with cancer.

Key Claim Terms New

Definitions of key terms used in the patent claims.

Term (Source)Support for SpecificationInterpretation
Amplified progeny polynucleotides
(Claim 1, Claim 18)
Amplifying the tagged parent polynucleotides in the set produces a corresponding set of amplified progeny polynucleotides. Sequencing a subset of the set of amplified progeny polynucleotides produces a set of sequencing reads. These reads are then collapsed to generate consensus sequences corresponding to a unique polynucleotide among the set of tagged parent polynucleotides.The collection of daughter molecules generated through the replication (amplification) of the original parent polynucleotides that were tagged with barcodes.
Mappable base position
(Claim 1, Claim 18)
Mapping sequence reads derived from the sequencing onto a reference sequence involves identifying a subset of mapped sequence reads that align with a variant of the reference sequence at each mappable base position. For each mappable base position, a ratio of mapped sequence reads including a variant to the total sequence reads is calculated. This enables the determination of potential rare variants or mutations at specific genomic locations.A specific location or coordinate within a reference genome where a sequence read can be uniquely or significantly aligned.
Molecular barcodes
(Claim 1, Claim 18)
The barcode is a polynucleotide, which may further comprise random sequence or a fixed or semi-random set of oligonucleotides that in combination with the diversity of molecules sequenced from a select region enables identification of unique molecules. Barcodes can be at least a 3, 5, 10, 15, 20, 25, 30, 35, 40, 45, or 50 mer base pairs in length. They are attached to extracellular polynucleotides or fragments thereof prior to any amplification or enrichment step.Polynucleotide sequences, which may include random, fixed, or semi-random oligonucleotides, used to identify unique molecules or families of molecules when combined with other sequence properties like start/stop positions.
Non-uniquely tagging
(Claim 1, Claim 18)
In some embodiments, each tagged parent polynucleotide in the set is uniquely tagged. In other embodiments, the tags are non-unique. The plurality of polynucleotides is tagged with a number of different molecular barcodes ranging from at least 2 to fewer than a number of cell-free nucleic acid molecules of the subset that map to the mappable base position.A process of attaching molecular barcodes to a population of polynucleotides where the number of distinct barcode sequences is fewer than the number of individual polynucleotide molecules that map to the same genomic position, such that multiple different parent molecules may share the same barcode.
Selectively enriching
(Claim 1, Claim 18)
The methods of the disclosure may comprise selectively enriching regions from the subject's genome or transcriptome prior to sequencing. This can be achieved through selective amplification of sequences, selective amplification of tagged parent polynucleotides, or selective sequence capture of amplified progeny polynucleotides. Enrichment targets regions such as oncogenes, tumor suppressor genes, promoters, or regulatory sequence elements.The process of specifically isolating or increasing the relative concentration of particular genomic regions of interest from a larger pool of polynucleotides prior to sequencing.

Litigation Cases New

US Latest litigation cases involving this patent.

Case NumberFiling DateTitle
1:22-cv-00334Mar 17, 2022Illumina, Inc. v. Guardant Health, Inc. et al

Patent Family

Patent Family

File Wrapper

The dossier documents provide a comprehensive record of the patent's prosecution history - including filings, correspondence, and decisions made by patent offices - and are crucial for understanding the patent's legal journey and any challenges it may have faced during examination.

  • Get instant alerts for new documents

US10894974

Application Number
US15669779A
Filing Date
Aug 4, 2017
Publication Date
Jan 19, 2021
External Links
Slate, USPTO , Google Patents