Methods and systems for detecting genetic variants

Patent No. US10801063 (titled "Methods and systems for detecting genetic variants") on Oct 14, 2019. The application was issued on Oct 13, 2020.

What is this patent about?

’063 is related to the field of genetic analysis, specifically the detection of rare genetic variants and copy number variations (CNV) in cell-free DNA (cfDNA) samples. In liquid biopsy applications, such as cancer monitoring, the ability to accurately quantify DNA fragments is often hampered by the loss of molecules during library preparation and sequencing, as well as errors introduced during PCR amplification.

The underlying idea behind ’063 is that the number of original DNA molecules in a sample can be more accurately estimated by tracking the recovery of complementary strands from individual double-stranded fragments. By using duplex tagging to distinguish the Watson and Crick strands, the system can identify which molecules were recovered as pairs, which were recovered as single strands, and—crucially—statistically infer the number of molecules that were entirely lost during the process.

The claims of ’063 focus on a method for classifying sequence data by non-uniquely tagging double-stranded cfDNA with a significant molar excess of barcoded adapters. The process involves grouping mapped sequencing reads into families based on their molecular barcodes and genomic start/stop positions to generate consensus sequences. These consensus sequences are then categorized as either paired (representing both strands of the original duplex) or unpaired (representing only one strand).

In practice, the invention utilizes a library of adapters where the number of unique barcodes is intentionally smaller than the number of DNA fragments, relying on the combination of the barcode and the fragment endpoints to uniquely identify parent molecules. This high-efficiency ligation ensures that a significant portion of the sample is tagged at both ends, allowing the bioinformatics pipeline to collapse redundant reads into high-fidelity consensus sequences while filtering out artifacts that appear on only one strand.

This approach differentiates itself from prior methods by moving beyond simple error correction to address the problem of stochastic sampling loss. By quantifying the ratio of paired to unpaired strands at specific genetic loci, the method corrects for local sequencing biases and missing data. This statistical inference provides a more robust foundation for detecting minute changes in copy number and identifying rare somatic mutations with a specificity exceeding 99.9%.

How does this patent fit in bigger picture?

Technical Landscape

In the early 2010s when ’063 was filed, genomic analysis of heterogeneous samples was typically implemented using massively parallel sequencing of libraries where original nucleic acid fragments were converted into sequenceable forms via standard adapter ligation. At a time when systems commonly relied on bioinformatics to estimate copy number variation based primarily on the count of successfully sequenced reads, hardware and software constraints made it non-trivial to account for the stochastic loss of molecules during sample preparation. Engineering practices in this era generally accepted that a significant portion of the starting genetic material would remain unobserved, leading to inherent sensitivity limits when attempting to detect rare genetic variants or precise copy number changes in complex mixtures like cell-free DNA.

Prosecution Position

The disclosed invention represents a meaningful technical advancement through an architectural shift in how molecular redundancy and recovery are tracked to improve quantification accuracy. By utilizing a library of molecular barcodes to tag both strands of double-stranded DNA fragments in a single reaction, the system enables the differentiation between recovered pairs (both strands sequenced) and singlets (only one strand sequenced). This structural approach allows for the statistical inference of 'unseen' molecules that were converted but not sequenced, overcoming the technical constraint of sampling bias. The integration of this molecular counting method with high-fidelity amplification and consensus sequencing enables the detection of rare genetic variants with a specificity exceeding 99.9%, a capability enabled by correcting for the variable recovery rates across different genomic loci.

Claims

This patent contains a total of 28 claims, with claims 1 and 15 serving as the independent claims. The independent claims focus on methods for classifying consensus or unique sequencing reads derived from double-stranded cell-free DNA by utilizing high-molar excess adapter tagging with molecular barcodes to distinguish between paired Watson-Crick strands and unpaired single strands. The dependent claims serve to specify technical parameters such as sample types, DNA input amounts, barcode library sizes, ligation efficiencies, target cancer-related gene panels, and computational methods for quantifying both detected and undetected DNA molecules at specific genomic loci.

Key Claim Terms New

Definitions of key terms used in the patent claims.

Term (Source)Support for SpecificationInterpretation
Consensus sequence
(Claim 1)
Methods disclosed herein can comprise collapsing, e.g., generating a consensus sequence by comparing multiple sequence reads. Sequence reads generated from a single original polynucleotide can be used to generate a consensus sequence of that original polynucleotide. Comparison of sequence reads of molecules derived from a single original molecule can be analyzed so as to determine the original, or 'consensus' sequence, often using phylogenetic analysis or voting.A single representative nucleotide sequence derived by comparing and merging multiple sequence reads that originated from the same parent polynucleotide strand to filter out amplification and sequencing errors.
Molecular barcode
(Claim 1, Claim 15)
Each of the tags comprises a molecular barcode, which can be a randomer or selected from a collection of different molecular barcodes. The molecular barcodes are at least 4 nucleotide bases in length and have an edit distance of at least 1 between one another. They are used to group sequence reads into families, where each family comprises sequence reads from one of the template polynucleotides.A specific polynucleotide sequence within an adapter used to identify and distinguish individual nucleic acid molecules or strands.
Non-uniquely tagging
(Claim 1)
In any of the embodiments herein, there are more polynucleotides (e.g., cfDNA fragments) to be tagged than there are different molecular barcodes such that the tagging is not unique. The library adaptors provide tagged DNA fragments having from 2 to 1000 different combinations of molecular barcodes. This allows for non-unique tagging where the combination of the barcode and the fragment's genomic coordinates (start/end positions) are used to distinguish molecules.A process of attaching adapters to DNA fragments where the number of available unique molecular barcodes is fewer than the number of DNA fragments mapping to a specific genomic position, resulting in multiple different parent polynucleotides sharing the same barcode sequence.
Paired consensus sequences
(Claim 1)
For all molecules in a particular region, counts of molecules where both Watson and Crick sides were recovered (“Pairs”) versus those where only one half was recovered (“Singlets”) can be recorded. Each paired read corresponds to sequence reads generated from a first tagged strand and a second differently tagged complementary strand derived from a double-stranded polynucleotide molecule. This allows for the estimation of unseen molecules based on the number of Pairs and Singlets detected.Consensus sequences derived from sequence reads that represent both the Watson and the Crick strands of the same original double-stranded DNA molecule.
Unpaired consensus sequences
(Claim 1)
Unpaired reads represent a first tagged strand having no second differently tagged complementary strand derived from a double-stranded polynucleotide molecule represented among the sequence reads. These are also referred to as 'Singlets' where only one half of the original input double-stranded DNA molecule was recovered. The ratio of paired versus unpaired strands is used to infer the quantitative measure of individual DNA molecules for which neither strand was detected.Consensus sequences derived from sequence reads that represent only one of the two strands (either Watson or Crick) of the original double-stranded DNA molecule.

Litigation Cases New

US Latest litigation cases involving this patent.

Case NumberFiling DateTitle
1:22-cv-00334Mar 17, 2022Illumina, Inc. v. Guardant Health, Inc. et al

Patent Family

Patent Family

File Wrapper

The dossier documents provide a comprehensive record of the patent's prosecution history - including filings, correspondence, and decisions made by patent offices - and are crucial for understanding the patent's legal journey and any challenges it may have faced during examination.

  • Get instant alerts for new documents

US10801063

Application Number
US16601168A
Filing Date
Oct 14, 2019
Publication Date
Oct 13, 2020
External Links
Slate, USPTO , Google Patents