Systems and methods to detect rare mutations and copy number variation

Patent No. US10995376 (titled "Systems and methods to detect rare mutations and copy number variation") on Jan 11, 2021. The application was issued on May 4, 2021.

What is this patent about?

’376 is related to the field of genetic analysis and molecular diagnostics, specifically focusing on the detection of rare mutations and copy number variations within cell-free polynucleotides. The technology addresses the inherent challenges of identifying low-frequency genetic signals, such as those derived from tumors or fetuses, which are often obscured by the high background noise and distortion typical of standard next-generation sequencing workflows.

The underlying idea behind ’376 is to utilize a digital sequencing approach that employs molecular tagging to distinguish true genetic variants from artifacts introduced during laboratory processing. By assigning identifiers to individual parent molecules before amplification, the system can collapse multiple sequencing reads into high-fidelity consensus sequences, effectively filtering out errors and providing a more accurate quantitative measure of the original genetic material.

The claims of ’376 focus on a method for preparing DNA libraries for sequencing by attaching a specific number of molecular barcodes to a population of DNA molecules. The number of unique barcodes, denoted as n, is strategically calibrated to be at least 2 but no more than 100,000 times the mean number of expected duplicate molecules that share identical start and stop positions, ensuring that cognate fragments can be uniquely identified even in samples containing up to 300,000 haploid human genome equivalents.

In practice, the invention works by ligating these molecular barcodes to cell-free DNA fragments, followed by amplification and selective enrichment of specific genomic regions of interest, such as actionable oncogenes. This targeted approach allows for the deep sequencing of progeny polynucleotides, which are then grouped into families based on their barcodes and genomic coordinates to reconstruct the sequence of the original parent molecule with high sensitivity.

This methodology differs from prior approaches by optimizing the conversion efficiency of the library preparation and using a mathematically defined diversity of tags to handle the natural overlap of fragmented DNA. By focusing on the relationship between the number of tags and the expected number of duplicate molecules, the system achieves a sensitivity threshold as low as 0.1%, enabling the detection of rare variants that would otherwise be lost in the baseline noise of conventional analog sequencing.

How does this patent fit in bigger picture?

Technical Landscape

In the early 2010s when ’376 was filed, the analysis of genetic material from bodily fluids was typically implemented using standard sequencing protocols where the detection of rare variants was limited by the inherent error rates of the sequencing hardware. At a time when systems commonly relied on raw read depth to quantify genetic alterations, distinguishing low-frequency somatic mutations from background noise was non-trivial due to the accumulation of stochastic errors during library preparation and amplification. Furthermore, computational architectures for processing cell-free DNA often utilized uniform mapping and counting methods that did not account for the specific molecular diversity or the unique fragmentation patterns of extracellular polynucleotides.

Prosecution Position

The disclosed invention represents a meaningful technical advancement through the integration of molecular tagging and computational collapsing to generate high-fidelity consensus sequences from cell-free DNA. This architectural shift moves beyond simple read counting by utilizing unique identifiers and sequence start/stop positions to group progeny reads into families, thereby enabling the identification of the original parent polynucleotides. This capability enables the detection of rare mutations and copy number variations with a sensitivity exceeding the per-base error rate of the sequencing platform. By normalizing these consensus-derived measures against reference controls and applying probabilistic modeling, the system overcomes the technical constraint of distinguishing true biological variants from artifacts introduced during amplification and sequencing.

Claims

The patent contains a total of 20 claims, with claims 1 and 20 serving as the independent claims. These independent claims focus on a method for preparing DNA molecules for sequencing by attaching a specific range of molecular barcodes based on expected duplicate molecule counts, followed by amplification and selective enrichment of target regions. The dependent claims serve to further define the process by specifying the biological sources of the DNA, the physical properties and quantities of the genetic material, the specific configurations and lengths of the molecular barcodes, and the technical parameters for the amplification and enrichment stages.

Key Claim Terms New

Definitions of key terms used in the patent claims.

Term (Source)Support for SpecificationInterpretation
Haploid human genome equivalents
(Claim 20)
The disclosure provides for a method comprising providing a sample comprising between 100 and 100,000 haploid human genome equivalents of cell free DNA (cfDNA) polynucleotides. In some embodiments, the composition comprises between 1000 and 50,000 haploid human genome equivalents of cfDNA polynucleotides. This measure is used to determine the appropriate number of unique identifiers for tagging.A quantitative measure of the amount of DNA in a sample, where one equivalent represents the amount of DNA contained in a single human haploid germ cell.
Molecular barcodes
(Claim 1, Claim 20)
In some embodiments, the barcode is a polynucleotide, which may further comprise random sequence or a fixed or semi-random set of oligonucleotides that in combination with the diversity of molecules sequenced from a select region enables identification of unique molecules. The barcodes comprise oligonucleotides is at least a 3, 5, 10, 15, 20 25, 30, 35, 40, 45, or 50mer base pairs in length. Barcodes can be attached to the extracellular polynucleotides or fragments thereof prior to any amplification or enrichment step.Polynucleotide sequences, such as random sequences or fixed sets of oligonucleotides, attached to DNA fragments to enable identification of unique molecules and distinguish them from amplification duplicates.
Progeny polynucleotides
(Claim 1, Claim 20)
Amplifying the tagged parent polynucleotides in the set produces a corresponding set of amplified progeny polynucleotides. Sequencing a subset of the set of amplified progeny polynucleotides produces a set of sequencing reads. Collapsing the set of sequencing reads generates a set of consensus sequences, each consensus sequence corresponding to a unique polynucleotide among the set of tagged parent polynucleotides.The collection of polynucleotide molecules generated by amplifying the original tagged parent molecules, representing copies of the starting genetic material.
Selectively enriching
(Claim 1, Claim 20)
In some embodiments, the methods of the disclosure also comprise a step of selectively enriching regions from the subject's genome or transcriptome prior to sequencing. Enrichment can be performed by selective amplification of sequences, selective amplification of tagged parent polynucleotides, or selective sequence capture of amplified progeny polynucleotides. This allows for multiplex sequencing on specific regions of interest.The process of specifically isolating or amplifying targeted regions of the genome or transcriptome from a larger population of polynucleotides prior to sequencing.
Tagged parent polynucleotides
(Claim 1, Claim 20)
The method comprises providing at least one set of tagged parent polynucleotides, and for each set of tagged parent polynucleotides, amplifying the tagged parent polynucleotides in the set to produce a corresponding set of amplified progeny polynucleotides. The disclosure also provides for a method comprising converting initial starting genetic material into the tagged parent polynucleotides. The generation of consensus sequences is based on information from the tag and/or at least one of sequence information at the beginning (start) region of the sequence read, the end (stop) regions of the sequence read and the length of the sequence read.The original extracellular DNA molecules from a sample after they have been ligated or otherwise attached to molecular barcodes but prior to amplification.

Litigation Cases New

US Latest litigation cases involving this patent.

Case NumberFiling DateTitle
1:22-cv-00334Mar 17, 2022Illumina, Inc. v. Guardant Health, Inc. et al

Patent Family

Patent Family

File Wrapper

The dossier documents provide a comprehensive record of the patent's prosecution history - including filings, correspondence, and decisions made by patent offices - and are crucial for understanding the patent's legal journey and any challenges it may have faced during examination.

  • Get instant alerts for new documents

US10995376

Application Number
US17146359A
Filing Date
Jan 11, 2021
Publication Date
May 4, 2021
External Links
Slate, USPTO , Google Patents