Methods and systems for detecting genetic variants

Patent No. US11118221 (titled "Methods and systems for detecting genetic variants") on Jan 7, 2020. The application was issued on Sep 14, 2021.

What is this patent about?

’221 is related to the field of genetic analysis and bioinformatics, specifically focusing on the detection of rare genetic variants such as copy number variations (CNV) and single nucleotide variants (SNV) within cell-free DNA (cfDNA) samples. In clinical diagnostics, identifying these mutations is often hampered by the low concentration of target molecules and the noise introduced during library preparation and sequencing. The technology addresses the need for high-sensitivity monitoring of disease states, such as cancer, by improving the accuracy of molecule counting and error correction in heterogeneous polynucleotide populations.

The underlying idea behind ’221 is that the true count of original DNA fragments in a sample can be mathematically inferred by tracking the recovery of complementary strands through the use of duplex tagging. By labeling the Watson and Crick strands of a double-stranded molecule so they can be distinguished after amplification, the system can categorize sequence reads into pairs (where both strands are recovered) and singlets (where only one is recovered). This distribution allows for a statistical estimation of the unseen molecules—those that were present in the original sample but failed to be sequenced—thereby correcting for sampling bias and improving the quantitative accuracy of genetic assays.

The claims of ’221 focus on a method for processing cfDNA by attaching duplex tags containing molecular barcodes to both ends of the fragments and subsequently generating consensus sequences. The independent claims specifically cover the reduction of redundancy by sorting sequence reads into paired and unpaired reads based on whether both complementary strands of an original parent molecule were successfully detected. Furthermore, the claims describe a non-unique tagging approach where the number of distinct barcodes is intentionally smaller than the number of DNA fragments at a specific locus, using the combination of barcodes and endogenous sequence information to uniquely identify parent molecules.

In practice, the invention works by ligating specialized adapters to cfDNA fragments at high conversion efficiencies, often exceeding 50%. These adapters contain degenerate or semi-degenerate barcodes that survive the amplification process, allowing a computer processor to group progeny reads into families derived from the same original molecule. By comparing the reads within these families, the system performs consensus calling, which filters out stochastic errors introduced by polymerases or sequencers. If a mutation appears in both the Watson and Crick strands of a paired read, the confidence that the variant is a true biological mutation rather than an artifact is significantly increased.

This approach differs from prior methods by moving beyond simple unique identification to a more sophisticated statistical reconstruction of the original sample composition. While traditional barcoding helps identify duplicates, ’221 uses the relationship between paired and unpaired strands to calculate a more precise total molecule count at specific genomic loci. This enables the detection of rare variants at concentrations below 1% with a specificity greater than 99.9%, providing a robust framework for monitoring tumor burden and identifying drug-resistance mutations in patients undergoing targeted therapy.

How does this patent fit in bigger picture?

Technical Landscape

In the early 2010s when ’221 was filed, the detection of rare genetic variants in heterogeneous samples was typically implemented using massively parallel sequencing of genomic libraries. At a time when these systems commonly relied on bioinformatics to analyze only the molecules successfully converted and sequenced, technical constraints made it non-trivial to account for the stochastic loss of nucleic acids during library preparation. Because standard architectures lacked a mechanism to track the recovery of individual DNA strands, the count of converted but unsequenced molecules remained an unknown variable, which often resulted in highly variable sensitivity across different genomic regions and limited the accuracy of copy number variation and rare variant quantification.

Prosecution Position

The disclosed invention represents a meaningful technical advancement through an architectural shift in library preparation that enables the estimation of unsequenced polynucleotides. By integrating a tagging system that utilizes a library of molecular barcodes to differently label the Watson and Crick strands of double-stranded DNA fragments, the system enables the classification of sequence reads into paired and unpaired families. This structural approach allows for the application of probabilistic models to infer the number of unseen molecules based on the ratio of detected pairs to singlets. The resulting technical effect is a significant reduction in sequencing noise and a substantial increase in detection specificity, overcoming the constraint of sampling bias to enable the quantification of rare DNA variants at concentrations below 1% with greater than 99.9% specificity.

Claims

This patent contains 30 claims, with claims 1 and 18 serving as the independent claims. The independent claims focus on methods for analyzing cell-free DNA by tagging molecules with duplex barcodes, amplifying and sequencing these molecules, and then using the barcode information to track redundancy and generate consensus sequences from both paired and unpaired reads to quantify original genetic material. The dependent claims serve to specify sample types such as cancer-derived DNA, define technical parameters for barcode length and ligation efficiency, identify specific target genes for enrichment, and detail the computational processing steps for mapping reads and estimating quantitative measures of the parent polynucleotides.

Key Claim Terms New

Definitions of key terms used in the patent claims.

Term (Source)Support for SpecificationInterpretation
Consensus sequences
(Claim 1)
Methods disclosed herein can comprise collapsing, e.g., generating a consensus sequence by comparing multiple sequence reads. For example, sequence reads generated from a single original polynucleotide can be used to generate a consensus sequence of that original polynucleotide. Comparison of sequence reads of molecules derived from a single original molecule, including those that have sequence variants, can be analyzed so as to determine the original, or 'consensus' sequence.A representative nucleotide sequence generated by comparing and merging multiple sequence reads derived from the same original parent polynucleotide to filter out errors.
Duplex tags
(Claim 1, Claim 18)
Input double-stranded deoxyribonucleic acid (DNA) can be converted by a process that tags both halves of the individual double-stranded molecule, in some cases differently. If tagged correctly, each original Watson and Crick (i.e., strand) side of the input double-stranded DNA molecule can be differently tagged and identified by the sequencer and subsequent bioinformatics. The duplex tags are not sequencing adaptors.A set of adapters used to tag double-stranded DNA such that the first and second complementary strands of an original molecule are identified and can be distinguished as a pair.
Molecular barcodes
(Claim 1, Claim 18)
The set of library adaptors can comprise plurality of polynucleotide molecules with molecular barcodes, wherein the plurality of polynucleotide molecules are less than or equal to 80 nucleotide bases in length, wherein the molecular barcodes are at least 4 nucleotide bases in length. The molecular barcodes are different from one another and have an edit distance of at least 1 between one another. In some embodiments, the given molecular barcode is a randomer.Distinct nucleotide sequences within a tag used to identify and track individual parent polynucleotides and their progeny through amplification and sequencing.
Non-uniquely tagging
(Claim 18)
In any of the embodiments herein, there are more polynucleotides (e.g., cfDNA fragments) to be tagged than there are different molecular barcodes such that the tagging is not unique. The double-stranded cfDNA molecules that map to a mappable base position of a reference sequence are tagged with a number of different molecular barcodes ranging from at least 2 to fewer than a number of the double-stranded cfDNA molecules. This allows for non-unique tagging where the combination of barcode and fragment end sequence identifies the molecule.A tagging scheme where the number of available distinct molecular barcodes is smaller than the number of DNA fragments, resulting in multiple fragments sharing the same barcode.
Paired reads
(Claim 1, Claim 18)
For all molecules in a particular region, counts of molecules where both Watson and Crick sides were recovered (“Pairs”) versus those where only one half was recovered (“Singlets”) can be recorded. Each paired read corresponds to sequence reads generated from a first tagged strand and a second differently tagged complementary strand derived from a double-stranded polynucleotide molecule in the set. The number of unseen molecules can be estimated based on the number of Pairs and Singlets detected.Sequence reads where both the Watson and Crick strands of an original double-stranded DNA molecule have been recovered and identified.

Litigation Cases New

US Latest litigation cases involving this patent.

Case NumberFiling DateTitle
1:22-cv-00334Mar 17, 2022Illumina, Inc. v. Guardant Health, Inc. et al

Patent Family

Patent Family

File Wrapper

The dossier documents provide a comprehensive record of the patent's prosecution history - including filings, correspondence, and decisions made by patent offices - and are crucial for understanding the patent's legal journey and any challenges it may have faced during examination.

  • Get instant alerts for new documents

US11118221

Application Number
US16672267A
Filing Date
Jan 7, 2020
Publication Date
Sep 14, 2021
External Links
Slate, USPTO , Google Patents