Methods and systems for detecting genetic variants

Patent No. US11149307 (titled "Methods and systems for detecting genetic variants") on Feb 4, 2021. The application was issued on Oct 19, 2021.

What is this patent about?

’307 is related to the field of genetic analysis, specifically the detection and quantification of rare genetic variants such as copy number variations (CNV) and single nucleotide variants (SNV) within cell-free DNA (cfDNA) samples. The technology addresses the inherent limitations of liquid biopsies, where the low concentration of tumor-derived DNA and the inefficiencies of sequencing workflows often lead to inaccurate molecule counts and reduced diagnostic sensitivity.

The underlying idea behind ’307 is that the true number of DNA fragments in a sample can be mathematically inferred by tracking the recovery of individual strands from double-stranded molecules. By using a non-unique tagging strategy—where a limited set of barcodes is combined with endogenous fragment sequences—the system can distinguish between molecules where both strands were sequenced (pairs) and those where only one was captured (singlets). This statistical distribution allows the system to calculate the number of unseen molecules that were present in the original sample but lost during the library preparation or sequencing process.

The claims of ’307 focus on a method for determining the number of cfDNA molecules by employing a high-efficiency ligation step using more than a 10× excess of adapters relative to the DNA population. The independent claims specify that adapters containing molecular barcodes are ligated to both ends of the fragments, achieving at least a 20% conversion efficiency. The method then utilizes the resulting sequence reads and barcodes to map fragments to a reference genome and quantify the original molecule population at specific loci based on the detection of one or both strands.

In practice, the invention works by tagging double-stranded polynucleotides with duplex adapters that allow the bioinformatics pipeline to differentiate between the Watson and Crick strands. After amplification and sequencing, the system groups reads into families based on their barcodes and genomic start/stop positions. By analyzing the ratio of paired reads to unpaired reads, the system applies a probabilistic model (such as a binomial distribution) to correct for sampling bias and library loss, providing a more accurate representation of the genetic landscape than simple read counting.

This approach differentiates itself from prior methods by moving beyond simple error correction to address the problem of stochastic sampling loss. While traditional molecular barcoding focuses on eliminating polymerase-induced artifacts, ’307 provides a mechanism to account for the molecules that never make it to the sequencer at all. This statistical inference is critical for detecting minute changes in copy number and rare somatic mutations, ensuring that clinical decisions are based on the actual molecular burden in the patient's blood rather than sequencing artifacts.

How does this patent fit in bigger picture?

Technical Landscape

In the early 2010s when ’307 was filed, genomic analysis was typically implemented using massively parallel sequencing of libraries where original nucleic acid fragments were converted into sequenceable forms through standard ligation. At a time when systems commonly relied on bioinformatics to analyze only the successfully sequenced reads rather than accounting for the total population of converted molecules, hardware and software constraints made the accurate estimation of unseen or unrecovered genetic material non-trivial. During this era, technical practices for quantifying rare genetic variants or copy number variations were often limited by the stochastic nature of sampling and the inherent noise introduced during library preparation and amplification, which frequently obscured low-frequency signals in heterogeneous samples.

Prosecution Position

The disclosed invention represents a technical advancement by providing an architectural shift in how molecular redundancy and strand-specific information are utilized to quantify nucleic acids. By integrating a tagging system that identifies both the Watson and Crick strands of a double-stranded DNA molecule, the method enables the differentiation between paired reads, unpaired reads, and unseen molecules. This structural approach allows for the statistical inference of the total number of original molecules at a locus, including those not directly detected by the sequencer. This capability overcomes the technical constraint of sampling bias and variable conversion efficiency, achieving a significant increase in sensitivity and specificity for detecting rare genetic alterations in heterogeneous populations, such as cell-free DNA.

Claims

This patent contains 26 claims, with claims 1, 13, and 15 serving as the independent claims. These independent claims focus on methods for quantifying cell-free DNA molecules in a sample by using a high excess of barcode-containing adapters, amplifying and sequencing the tagged molecules, and utilizing sequence information or strand-detection metrics to determine the original molecule count. The dependent claims serve to specify sample types such as blood or plasma, define barcode characteristics and adapter attachment efficiencies, incorporate steps for genomic enrichment and redundancy tracking, and provide further detail on calculating quantitative measures for single-stranded, double-stranded, or undetected DNA fragments, particularly for detecting low-concentration circulating tumor DNA.

Key Claim Terms New

Definitions of key terms used in the patent claims.

Term (Source)Support for SpecificationInterpretation
Amplified progeny polynucleotides
(Claim 1, Claim 13, Claim 15)
Subjecting the tagged polynucleotide fragments to nucleic acid amplification reactions under conditions that yield amplified polynucleotide fragments as amplification products of the tagged polynucleotide fragments. Comparison of sequence reads of molecules derived from a single original molecule, including those that have sequence variants, can be analyzed so as to determine the original, or 'consensus' sequence. Iterative rounds of amplification can introduce errors into progeny polynucleotides.The collection of daughter molecules generated through high-fidelity replication (e.g., PCR) of the tagged parent DNA fragments.
Both strands are detected
(Claim 13)
For all molecules in a particular region, counts of molecules where both Watson and Crick sides were recovered (“Pairs”) versus those where only one half was recovered (“Singlets”) can be recorded. The nucleotide call from a double-strand amplified fragment has a probability of p^2. This information is used to infer the count of molecules that were converted but not sequenced.A state where sequence reads are recovered and identified for both the Watson and the Crick strands of a single original double-stranded DNA fragment, often referred to as a 'Pair'.
Mappable base position
(Claim 1, Claim 13, Claim 15)
The method comprises mapping the consensus sequence to the target genome. The population of cfDNA molecules comprises a subset of cfDNA molecules that map to a mappable base position of a reference sequence. This allows for the quantification of molecules at specific genetic loci.A specific coordinate or location within a reference genome to which sequence reads can be uniquely assigned based on their nucleotide sequence.
Molecular barcodes
(Claim 1, Claim 13, Claim 15)
The library adaptors can comprise a plurality of polynucleotide molecules with molecular barcodes, wherein the molecular barcodes are at least 4 nucleotide bases in length. In some embodiments, the given molecular barcode is a randomer. The number of unseen molecules can be estimated based on the number of Pairs and Singlets detected using these barcodes.Short, distinct polynucleotide sequences used to uniquely or non-uniquely identify individual DNA fragments and their strands, enabling the grouping of sequence reads into families derived from the same original parent molecule.
Only one strand is detected
(Claim 13)
For sequence readouts from a single-strand, assuming that one of the 2 strands is seen, and the other is unseen, the probability of seeing one strand is “p”, but the probability of missing the other strand is (1-p). These are recorded as 'Singlets'. The number of unseen molecules can be estimated based on the number of Pairs and Singlets detected.A state where a sequence read is recovered and identified for only one of the two complementary strands (either Watson or Crick) of an original double-stranded DNA fragment, referred to as a 'Singlet'.
Quantitative measure
(Claim 13)
The number of unseen molecules can be estimated based on the number of Pairs and Singlets detected. For sequence readouts from double-strands, the nucleotide call from a double-strand amplified fragment has a probability of p*p=p^2. For sequence readouts from a single-strand, the nucleotide call from a single-strand amplified fragment has a probability 2*p*(1-p).A calculated estimate of the number of original DNA molecules, including those detected as pairs, those detected as singlets, and those inferred to be present but not sequenced.
Tagged parent polynucleotides
(Claim 1, Claim 13, Claim 15)
Input double-stranded deoxyribonucleic acid (DNA) can be converted by a process that tags both halves of the individual double-stranded molecule, in some cases differently. If tagged correctly, each original Watson and Crick (i.e., strand) side of the input double-stranded DNA molecule can be differently tagged and identified. These tagged molecules are then amplified to produce progeny polynucleotides.Original double-stranded DNA fragments from a sample that have been ligated to adapters containing molecular barcodes at both ends, serving as the template for subsequent amplification and sequencing.

Litigation Cases New

US Latest litigation cases involving this patent.

Case NumberFiling DateTitle
1:22-cv-00334Mar 17, 2022Illumina, Inc. v. Guardant Health, Inc. et al

Patent Family

Patent Family

File Wrapper

The dossier documents provide a comprehensive record of the patent's prosecution history - including filings, correspondence, and decisions made by patent offices - and are crucial for understanding the patent's legal journey and any challenges it may have faced during examination.

  • Get instant alerts for new documents

US11149307

Application Number
US17167974A
Filing Date
Feb 4, 2021
Publication Date
Oct 19, 2021
External Links
Slate, USPTO , Google Patents