Methods and systems for detecting genetic variants

Patent No. US10889858 (titled "Methods and systems for detecting genetic variants") on Dec 13, 2019. The application was issued on Jan 12, 2021.

What is this patent about?

’858 is related to the field of molecular biology and genetic diagnostics, specifically focusing on the high-sensitivity detection of genetic variants in cell-free DNA (cfDNA). The technology addresses the challenges of analyzing heterogeneous genomic samples, such as blood plasma containing circulating tumor DNA, where rare mutations or copy number variations must be distinguished from sequencing noise and amplification artifacts.

The underlying idea behind ’858 is that the accuracy of genetic quantification can be significantly improved by tracking the recovery of both strands of a double-stranded DNA molecule. By using duplex tagging to uniquely identify the original Watson and Crick strands, the system can distinguish between true biological variants present on both strands and errors introduced during the lab process. Furthermore, the invention utilizes the ratio of detected strand pairs to single strands to mathematically infer the number of molecules that were present in the original sample but lost during sequencing.

The claims of ’858 focus on a method for analyzing cfDNA by ligating library adaptors to double-stranded fragments using a high molar excess—specifically more than 10×—to ensure high conversion efficiency. The process involves amplifying these tagged molecules and using the molecular barcodes to sort the resulting sequence reads into families. These families are then categorized based on whether they represent both strands of the original duplex or only a single detected strand, allowing for a more precise reconstruction of the starting sample composition.

In practice, the invention achieves high efficiency by ensuring that at least 20% of the cfDNA population is tagged at both ends. Once sequenced, a programmed computer processor maps the reads to a reference genome and groups them into families using a combination of the barcode sequence and the genomic start/stop positions of the fragments. This multi-layered identification allows the system to collapse redundant reads into a single consensus sequence, effectively filtering out the 'stray' errors that typically plague deep sequencing assays.

This approach differs from prior methods by moving beyond simple redundancy reduction to a more sophisticated statistical estimation of unseen molecules. While traditional techniques only count the molecules they successfully sequence, ’858 accounts for the highly variable recovery rates across different genomic regions. By incorporating these 'unseen' counts into the final analysis, the method provides a much more accurate measure of copy number variation and achieves the extreme specificity required to detect rare cancer-associated mutations at frequencies below 1%.

How does this patent fit in bigger picture?

Technical Landscape

In the early 2010s when ’858 was filed, the detection of rare genetic variants in heterogeneous genomic samples was typically implemented using massively parallel sequencing of libraries prepared from cell-free DNA. At a time when these systems commonly relied on bioinformatics to estimate copy number variations based solely on the count of successfully sequenced molecules, technical constraints made it non-trivial to account for the high variability of molecules that were converted during sample preparation but remained unsequenced. When hardware and software constraints limited the ability to distinguish between sequencing artifacts and true biological variants, engineering practices generally focused on increasing raw read depth rather than modeling the underlying molecular population dynamics of the original double-stranded DNA fragments.

Prosecution Position

The disclosed invention represents a meaningful technical advancement by integrating a molecular tagging architecture that enables the estimation of unobserved DNA molecules to improve the accuracy of genetic quantification. By utilizing a library of molecular barcodes to differently tag the Watson and Crick strands of double-stranded DNA fragments, the system enables an architectural shift from simple read counting to a statistical inference model based on the recovery of paired versus single-stranded reads. This capability allows for the calculation of the probability of detection and the subsequent inference of the number of unseen fragments, effectively overcoming the technical constraint of sampling bias in sequencing libraries. The resulting technical effect is a significant increase in the sensitivity and specificity of rare variant detection, achieving high-fidelity consensus reads that can distinguish true genetic alterations from amplification or sequencing errors.

Claims

The patent includes a total of 29 claims, with claims 1 and 16 serving as the independent claims. These independent claims focus on methods for analyzing double-stranded cell-free DNA by using a high molar excess of barcoded adaptors to tag molecules and subsequently identifying whether one or both strands of the original DNA molecule are represented in the resulting sequence data. The dependent claims serve to specify various sample types, ligation techniques, barcode configurations, and target genomic regions, while also detailing computational steps for mapping sequences, grouping them into families, and estimating the quantity of original DNA molecules at specific genetic loci.

Key Claim Terms New

Definitions of key terms used in the patent claims.

Term (Source)Support for SpecificationInterpretation
Families
(Claim 16)
The grouping comprises classifying the plurality of sequence reads into families by identifying (i) different molecular barcodes coupled to the plurality of polynucleotide molecules and (ii) similarities between the plurality of sequence reads. Each family includes a plurality of nucleic acid sequences that are associated with a different combination of molecular barcodes and similar or identical sequence reads. Comparison of sequence reads of molecules derived from a single original molecule can be analyzed so as to determine the original, or “consensus” sequence.Groups of sequence reads that are determined to have originated from the same unique parent polynucleotide molecule based on shared barcode and positional information.
Library adaptors
(Claim 1, Claim 16)
The set of library adaptors can comprise a plurality of polynucleotide molecules with molecular barcodes, wherein the plurality of polynucleotide molecules are less than or equal to 80 nucleotide bases in length. In some embodiments, each of the library adapters is Y-shaped, bubble shaped or hairpin shaped. The library adaptors do not include a complete sequencer motif.Oligonucleotide molecules, often Y-shaped, bubble-shaped, or hairpin-shaped, that contain molecular barcodes and are designed to be ligated to the ends of target DNA fragments.
Molecular barcode
(Claim 1, Claim 16)
The library adaptors can comprise a plurality of polynucleotide molecules with molecular barcodes, wherein the molecular barcodes are at least 4 nucleotide bases in length. If tagged correctly, each original Watson and Crick side of the input double-stranded DNA molecule can be differently tagged and identified by the sequencer and subsequent bioinformatics. The molecular barcodes are different from one another and have an edit distance of at least 1 between one another.A unique or non-unique polynucleotide sequence used to tag individual DNA fragments, allowing for the identification and tracking of reads derived from the same original parent molecule.
Molecular barcodes
(Claim 1, Claim 16)
The set of library adaptors can comprise plurality of polynucleotide molecules with molecular barcodes. If tagged correctly, each original Watson and Crick (i.e., strand) side of the input double-stranded DNA molecule can be differently tagged and identified by the sequencer and subsequent bioinformatics. The molecular barcodes are at least 4 nucleotide bases in length and have an edit distance of at least 1 between one another.Unique or non-unique polynucleotide sequences used to tag individual DNA fragments, allowing for the identification of reads originating from the same parent molecule and the differentiation between Watson and Crick strands.
Tagged parent polynucleotides
(Claim 1, Claim 16)
Input double-stranded deoxyribonucleic acid (DNA) can be converted by a process that tags both halves of the individual double-stranded molecule. This can be performed using a variety of techniques, including ligation of hairpin, bubble, or forked adapters. The method comprises tagging the original DNA fragments in a single reaction using a library of a plurality of different tags.The original double-stranded DNA fragments from the sample after they have been ligated to library adaptors containing molecular barcodes, serving as the templates for subsequent amplification.
Watson strand and a Crick strand
(Claim 1, Claim 16)
If tagged correctly, each original Watson and Crick (i.e., strand) side of the input double-stranded DNA molecule can be differently tagged and identified by the sequencer and subsequent bioinformatics. For all molecules in a particular region, counts of molecules where both Watson and Crick sides were recovered (“Pairs”) versus those where only one half was recovered (“Singlets”) can be recorded. The number of unseen molecules can be estimated based on the number of Pairs and Singlets detected.The two individual complementary strands that constitute a single double-stranded DNA molecule.

Litigation Cases New

US Latest litigation cases involving this patent.

Case NumberFiling DateTitle
1:22-cv-00334Mar 17, 2022Illumina, Inc. v. Guardant Health, Inc. et al

Patent Family

Patent Family

File Wrapper

The dossier documents provide a comprehensive record of the patent's prosecution history - including filings, correspondence, and decisions made by patent offices - and are crucial for understanding the patent's legal journey and any challenges it may have faced during examination.

  • Get instant alerts for new documents

US10889858

Application Number
US16714579A
Filing Date
Dec 13, 2019
Publication Date
Jan 12, 2021
External Links
Slate, USPTO , Google Patents