Systems and methods to detect rare mutations and copy number variation

Patent No. US10876152 (titled "Systems and methods to detect rare mutations and copy number variation") on Apr 19, 2019. The application was issued on Dec 29, 2020.

What is this patent about?

’152 is related to the field of molecular diagnostics and bioinformatics, specifically focusing on the high-sensitivity detection of genetic aberrations in cell-free DNA (cfDNA). The technology addresses the challenge of identifying rare somatic mutations and copy number variations that are often masked by the inherent noise and distortion found in standard next-generation sequencing workflows. By analyzing extracellular polynucleotides isolated from bodily fluids like blood, the system provides a non-invasive means to monitor disease states such as cancer.

The underlying idea behind ’152 is to maximize the conversion of limited starting genetic material into a digital library where individual parent molecules can be accurately tracked and reconstructed. By using a high molar excess of molecular barcodes to tag both ends of cfDNA fragments, the invention ensures that even a small amount of input DNA (1–100 ng) is represented with high fidelity. This approach allows the system to distinguish between true biological variants and artifacts introduced during PCR amplification or sequencing by grouping reads into families derived from the same original molecule.

The claims of ’152 focus on a method for detecting somatic genetic variants by contacting a specific mass of cfDNA with at least an 80× molar excess of molecular barcodes to ensure a conversion efficiency of at least 20%. The process involves attaching barcodes to both ends of the cfDNA fragments, amplifying these tagged molecules, and then sequencing the resulting progeny. The independent claims specifically highlight the use of these barcodes to group sequence reads into families, which are then used to identify single nucleotide variations, indels, copy number variations, or gene fusions with high sensitivity.

In practice, the invention functions by implementing a high-diversity library preparation that captures a vast majority of the starting molecules, which is critical when only a few mutated fragments may exist in a 10 mL blood sample. After sequencing, the bioinformatics pipeline collapses the raw reads into consensus sequences based on the unique or non-unique barcode combinations and the genomic coordinates of the fragments. This collapsing mechanism effectively filters out random sequencing errors, as a true mutation will appear across multiple progeny in a family, whereas an error will typically appear in only one.

This method differs from prior approaches by significantly improving the conversion efficiency of the library preparation and utilizing a dual-tagging strategy to overcome the limitations of low-abundance templates. While traditional sequencing often loses rare signals due to inefficient ligation or high background noise, this system uses digital sequencing to provide a near-perfect representation of the original sample. By normalizing read counts and variant frequencies across predefined genomic regions, the technology enables the detection of cancer-related changes at frequencies as low as 0.1%, providing a robust tool for serial monitoring of tumor burden and treatment response.

How does this patent fit in bigger picture?

Technical Landscape

In the mid-2010s when ’152 was filed, the analysis of cell-free DNA was typically implemented using high-throughput sequencing platforms that were subject to inherent per-base error rates and amplification biases. At a time when systems commonly relied on raw read counting or basic alignment for genetic testing, distinguishing rare somatic mutations from sequencing noise was technically challenging. Furthermore, when hardware and software constraints made the high-fidelity reconstruction of low-concentration polynucleotides non-trivial, standard bioinformatics pipelines often struggled to provide the sensitivity required for sub-chromosomal copy number variation or rare variant detection in samples with low haploid genome equivalents.

Prosecution Position

The disclosed invention represents a meaningful technical advancement through the integration of molecular tagging and computational collapsing to generate high-fidelity consensus sequences from fragmented extracellular polynucleotides. This architectural shift moves beyond simple read mapping by grouping sequence reads into families derived from the same parent molecule, enabling the identification of unique molecules and the filtering of stochastic errors. The technical effect achieved is a significant increase in sensitivity, allowing for the detection of rare mutations and copy number variations at frequencies lower than the per-base sequencing error rate. This capability overcomes the technical constraint of signal-to-noise ratios in liquid biopsies, enabling the simultaneous quantification of genetic alterations and the characterization of disease heterogeneity from minimal starting material.

Claims

The patent contains a total of 32 claims, with claims 1 and 24 being the independent claims. These independent claims focus on a method for detecting somatic genetic variants in cell-free DNA from a human blood sample by using a high molar excess of molecular barcodes to tag DNA molecules, followed by amplification, sequencing, and mapping to identify variations such as single nucleotide variations, indels, copy number variations, or gene fusions. The dependent claims serve to specify various technical parameters and procedural refinements, including specific tagging efficiencies, ligation techniques, barcode sequence lengths, enrichment for cancer-associated target regions, the generation of consensus sequences from grouped sequence families, and the production of clinical reports or treatment recommendations based on the detected mutation profiles.

Key Claim Terms New

Definitions of key terms used in the patent claims.

Term (Source)Support for SpecificationInterpretation
Families
(Claim 24)
Collapsing comprises grouping sequences reads sequenced from amplified progeny polynucleotides into families, each family amplified from the same tagged parent polynucleotide. A consensus sequence is then determined based on sequence reads in a family. This process allows for the inferring of a quantitative measure of unique molecules in the set.Groups of sequencing reads that are identified as having originated from the same initial tagged parent polynucleotide based on shared barcode sequences and/or mapping coordinates.
Molecular barcodes
(Claim 1, Claim 24)
The barcode is a polynucleotide, which may further comprise random sequence or a fixed or semi-random set of oligonucleotides. Barcodes enable identification of unique molecules and can be at least a 3, 5, 10, 15, 20, 25, 30, 35, 40, 45, or 50mer base pairs in length. They are attached to extracellular polynucleotides or fragments thereof prior to any amplification or enrichment step.Short polynucleotide sequences, ranging from 2 to 1,000,000 different sequences and 5-20 base pairs in length, used to uniquely or non-uniquely label individual nucleic acid fragments.
Progeny polynucleotides
(Claim 1, Claim 24)
The method comprises amplifying the tagged parent polynucleotides in the set to produce a corresponding set of amplified progeny polynucleotides. Sequencing is then performed on a subset of the set of amplified progeny polynucleotides to produce a set of sequencing reads. Collapsing these reads generates consensus sequences corresponding to a unique polynucleotide among the set of tagged parent polynucleotides.The collection of duplicated nucleic acid molecules generated by amplifying the initial tagged parent polynucleotides.
Somatic genetic variants
(Claim 1, Claim 24)
Abnormal conditions may be selected from the group consisting of mutations, rare mutations, single nucleotide variants, indels, copy number variations, transversions, translocations, gene fusions, and gene amplifications. The prevalence/concentration of each rare variant identified in the subject is reported and quantified simultaneously. These variants are identified by comparing normalized numbers to a control or reference sample.Non-inherited genetic alterations, including SNVs, indels, CNVs, and gene fusions, detected in cell-free DNA that may indicate the presence of an abnormal condition such as cancer.
Tagged parent polynucleotides
(Claim 1, Claim 24)
The method comprises converting initial starting genetic material into the tagged parent polynucleotides. In some embodiments, the barcode is a polynucleotide, which may further comprise random sequence or a fixed or semi-random set of oligonucleotides. This enables identification of unique molecules in combination with the diversity of molecules sequenced from a select region.Original cfDNA molecules from a sample that have been modified by the attachment of molecular barcodes or adapters to their ends to enable tracking and identification of unique starting molecules during subsequent processing.

Litigation Cases New

US Latest litigation cases involving this patent.

Case NumberFiling DateTitle
1:22-cv-00334Mar 17, 2022Illumina, Inc. v. Guardant Health, Inc. et al

Patent Family

Patent Family

File Wrapper

The dossier documents provide a comprehensive record of the patent's prosecution history - including filings, correspondence, and decisions made by patent offices - and are crucial for understanding the patent's legal journey and any challenges it may have faced during examination.

  • Get instant alerts for new documents

US10876152

Application Number
US16389680A
Filing Date
Apr 19, 2019
Publication Date
Dec 29, 2020
External Links
Slate, USPTO , Google Patents