Systems and methods to detect rare mutations and copy number variation

Patent No. US10822663 (titled "Systems and methods to detect rare mutations and copy number variation") on Oct 4, 2019. The application was issued on Nov 3, 2020.

What is this patent about?

’663 is related to the field of molecular diagnostics and bioinformatics, specifically focusing on the high-sensitivity detection of genetic aberrations in cell-free polynucleotides. The technology addresses the challenge of identifying rare mutations and copy number variations (CNVs) that are often masked by the inherent noise and distortion produced during PCR amplification and high-throughput sequencing. By analyzing DNA fragments shed into bodily fluids, the system provides a non-invasive means to monitor disease states such as cancer or fetal abnormalities.

The underlying idea behind ’663 is the use of a digital sequencing framework that treats polynucleotide sequencing as a communication theory problem, where the original molecule is a message and sequencing artifacts are noise. The key inventive insight lies in a non-unique tagging strategy combined with physical fragment characteristics—specifically start and stop positions—to track progeny molecules back to their original parent strands. This allows the system to collapse multiple sequencing reads into high-fidelity consensus sequences, effectively filtering out stochastic errors and amplification biases that would otherwise obscure low-frequency somatic variants.

The claims of ’663 focus on a method for detecting somatic genetic variants by ligating adapters containing molecular barcodes to both ends of cell-free DNA (cfDNA) molecules. Crucially, the method employs a specific tagging ratio where the number of unique barcodes is at least two but fewer than the total number of cfDNA fragments mapping to a specific genomic position. The independent claims further specify the use of these barcodes in conjunction with mapping coordinates (start and stop positions) to group sequencing reads into families, enabling the detection of single nucleotide variants, indels, gene fusions, and copy number variations.

In practice, the invention works by extracting cfDNA from a subject's blood or other bodily fluid and converting these fragments into a tagged library with high efficiency. After amplification and optional enrichment for cancer-related target regions, the resulting progeny molecules are sequenced. The bioinformatics pipeline then organizes these reads into familial groups based on their barcodes and genomic alignment. By comparing members within each family, the system can distinguish a true biological mutation—which would appear across the majority of family members—from a sequencing error that appears only sporadically.

This approach differs from prior methods by moving beyond simple quality filtering of individual reads to a structured signal decoding process at the family level. While traditional sequencing often requires large amounts of input DNA to overcome low conversion rates, this invention utilizes optimized ligation and digital collapsing to achieve high sensitivity even with limited samples. By normalizing read counts and utilizing consensus sequences, the system provides a more accurate quantitative measure of genetic material, allowing for the detection of rare variants at frequencies as low as 0.1% or less.

How does this patent fit in bigger picture?

Technical Landscape

In the early 2010s when ’663 was filed, the analysis of cell-free nucleic acids was typically implemented using standard sequencing protocols where the detection of rare genetic variants was limited by the inherent error rates of sequencing platforms. At a time when systems commonly relied on high-depth raw read counting rather than molecular barcoding to quantify genetic material, distinguishing true somatic mutations from stochastic sequencing noise was technically difficult. Furthermore, when hardware or software constraints made the high-fidelity reconstruction of original template molecules non-trivial, the sensitivity of liquid biopsy assays was often insufficient to detect low-frequency alterations against the background of healthy genomic DNA.

Prosecution Position

The disclosed invention represents a meaningful technical advancement through the integration of molecular tagging and computational collapsing to generate high-accuracy consensus sequences from fragmented extracellular polynucleotides. By attaching barcodes to parent polynucleotides prior to amplification and subsequently grouping progeny reads into families, the architecture enables the systematic identification and removal of amplification biases and sequencing errors. This structural approach achieves a significant increase in sensitivity, allowing for the detection of rare mutations and copy number variations at frequencies below the raw error rate of the sequencing platform. The solution overcomes the technical constraint of signal-to-noise ratios in cell-free DNA analysis, enabling the simultaneous quantification of genetic heterogeneity and rare variants from a single bodily sample.

Claims

This patent contains 30 claims, with claims 1 and 21 serving as the independent claims. The independent claims focus on a method for detecting somatic genetic variants by non-uniquely tagging cell-free DNA molecules with molecular barcodes, followed by amplification, sequencing, and analysis to identify mutations such as single nucleotide variants or copy number variations. The dependent claims further specify technical parameters such as barcode length, the use of consensus sequences and family grouping for error reduction, specific enrichment for cancer-associated target regions, and the application of these methods for tumor profiling and cancer diagnosis.

Key Claim Terms New

Definitions of key terms used in the patent claims.

Term (Source)Support for SpecificationInterpretation
Consensus sequences
(Claim 21)
Collapsing the set of sequencing reads generates a set of consensus sequences, each consensus sequence corresponding to a unique polynucleotide among the set of tagged parent polynucleotides. Collapsing comprises detecting and/or correcting errors, nicks or lesions present in the sense or anti-sense strand of the tagged parent polynucleotides or amplified progeny polynucleotides. The generation of consensus sequences is based on information from the tag and/or sequence information at the start/stop regions and length of the read.A representative sequence derived from a family of reads that filters out errors, nicks, or lesions by comparing multiple progeny reads from the same parent molecule.
Families
(Claim 21)
Collapsing comprises grouping sequences reads sequenced from amplified progeny polynucleotides into families, each family amplified from the same tagged parent polynucleotide. A consensus sequence is then determined based on sequence reads in a family. The method may involve assigning, for each family, a confidence score for each of a plurality of calls, taking into consideration a frequency of the call among members of the family.Groups of sequencing reads that are determined to have originated from the same original parent polynucleotide molecule based on shared identifiers.
Mappable base position
(Claim 1, Claim 21)
The method includes identifying a subset of mapped sequence reads that align with a variant of the reference sequence at each mappable base position. For each mappable base position, a ratio is calculated of the number of mapped sequence reads that include a variant to the number of total sequence reads. In some embodiments, the method comprises providing a plurality of sets of tagged parent polynucleotides, wherein each set is mappable to a different mappable position in the reference sequence.A specific location or coordinate within a reference genome where a sequence read can be uniquely or significantly aligned.
Molecular barcodes
(Claim 1, Claim 21)
In some embodiments, the barcode is a polynucleotide, which may further comprise random sequence or a fixed or semi-random set of oligonucleotides. The barcodes enable identification of unique molecules and may be at least a 3, 5, 10, 15, 20, 25, 30, 35, 40, 45, or 50 mer base pairs in length. Barcodes are attached to the extracellular polynucleotides or fragments thereof prior to any amplification or enrichment step.Polynucleotide sequences (tags) used to identify and track individual parent molecules and their amplified progeny, which may consist of random sequences, fixed sets, or semi-random oligonucleotides.
Non-uniquely tagging
(Claim 1, Claim 21)
In some embodiments, each barcode attached to extracellular polynucleotides or fragments thereof prior to sequencing is not unique. The disclosure also provides for a method comprising detecting genetic variation in non-uniquely tagged initial starting genetic material. The generation of consensus sequences can be based on information from the tag in combination with start/stop regions and length of the sequence read to enable identification of unique molecules.A process of attaching molecular barcodes to a population of polynucleotides where the number of available distinct barcode sequences is less than the number of individual polynucleotide molecules in a subset mapping to the same genomic position, such that multiple different parent molecules may share the same barcode.

Litigation Cases New

US Latest litigation cases involving this patent.

Case NumberFiling DateTitle
1:22-cv-00334Mar 17, 2022Illumina, Inc. v. Guardant Health, Inc. et al

Patent Family

Patent Family

File Wrapper

The dossier documents provide a comprehensive record of the patent's prosecution history - including filings, correspondence, and decisions made by patent offices - and are crucial for understanding the patent's legal journey and any challenges it may have faced during examination.

  • Get instant alerts for new documents

US10822663

Application Number
US16593633A
Filing Date
Oct 4, 2019
Publication Date
Nov 3, 2020
External Links
Slate, USPTO , Google Patents