Methods and systems for detecting genetic variants

Patent No. US9920366 (titled "Methods and systems for detecting genetic variants") on Sep 22, 2015. The application was issued on Mar 20, 2018.

What is this patent about?

’366 is related to the field of molecular biology and medical diagnostics, specifically the detection and quantification of genetic variants such as copy number variations (CNV) and single nucleotide variants (SNV). The technology is particularly useful for analyzing cell-free DNA (cfDNA) in bodily fluids to monitor diseases like cancer, where rare genetic alterations may be present at extremely low concentrations.

The underlying idea behind ’366 is that standard sequencing often fails to capture all molecules in a sample, leading to inaccurate quantification of genetic material. By using duplex tags to differently label the Watson and Crick strands of double-stranded DNA, the system can track which original molecules were fully recovered as pairs, which were partially recovered as singlets, and—crucially—statistically infer the number of molecules that were completely missed by the sequencer.

The claims of ’366 focus on a method for detecting double-stranded DNA by tagging molecules with a significant excess of molecular barcodes to ensure high conversion efficiency. The process involves grouping raw reads into families based on barcode and fragment end sequences, collapsing these into consensus reads, and then using the observed counts of paired and unpaired strands to calculate a quantitative measure of the total original molecules, including those for which neither strand was detected.

In practice, the invention utilizes a library of adapters, such as Y-shaped or bubble-shaped tags, to ensure that complementary strands are uniquely identifiable after amplification. By comparing the sequence of one strand with its complement, the system can filter out artifacts introduced during PCR or sequencing. If a variant appears in both strands of a duplex pair, it is called with high confidence, whereas variants appearing in only one strand are flagged as potential errors.

This approach differs from prior methods by moving beyond simple redundancy reduction to address the problem of unseen molecules. While traditional bioinformatics can reduce noise for molecules that are successfully sequenced, ’366 uses the ratio of paired to unpaired reads to correct for sampling bias across different genomic loci. This statistical correction allows for much higher sensitivity and specificity, enabling the detection of rare mutations at concentrations below 1% with greater than 99.9% specificity.

How does this patent fit in bigger picture?

Technical Landscape

In the early 2010s when ’366 was filed, genomic analysis was typically implemented using massively parallel sequencing of libraries where nucleic acid fragments were converted into sequenceable forms via standard adapter ligation. At a time when systems commonly relied on basic bioinformatics to estimate copy number variations from raw sequence counts, technical constraints in sample preparation and stochastic sampling made it non-trivial to account for molecules that were successfully converted but failed to be sequenced. During this era, engineering limitations in library preparation often resulted in a significant loss of quantitative accuracy because the underlying distribution of unsequenced molecules remained unknown, leading to high variability in sensitivity across different genomic regions.

Prosecution Position

The disclosed invention represents a meaningful technical advancement by introducing an architectural shift in library preparation and data processing that enables the estimation of unsequenced DNA molecules. By utilizing a library of molecular barcodes to tag both strands of double-stranded DNA fragments in a single reaction, the system enables the classification of sequence reads into paired strands (where both sides are recovered) and singlets (where only one side is recovered). This structural approach allows for the mathematical inference of the unseen molecule population based on the ratio of detected pairs to singlets, effectively overcoming the technical constraint of sampling bias. The resulting integration of physical duplex tagging with statistical redundancy tracking achieves a technical effect of higher sensitivity and specificity in detecting rare genetic variants and copy number variations in heterogeneous samples like cell-free DNA.

Claims

The patent contains a total of 21 claims, with claim 1 being the sole independent claim. This independent claim focuses on a method for detecting double-stranded DNA molecules in a biological sample by using a high excess of duplex tags to label complementary strands, enriching for specific genetic loci, and utilizing quantitative measures of detected paired, unpaired, and undetected strands to analyze the sample. The dependent claims serve to specify the use of cell-free nucleic acids, define the sorting and calculation of paired versus unpaired reads, detail the structural characteristics of the molecular barcodes and adaptors, and provide for the determination of total molecule counts and copy number variations.

Key Claim Terms New

Definitions of key terms used in the patent claims.

Term (Source)Support for SpecificationInterpretation
Consensus sequence reads
(Claim 1)
Comparison of sequence reads of molecules derived from a single original molecule, including those that have sequence variants, can be analyzed so as to determine the original, or 'consensus' sequence. This can be done phylogenetically. Such methods include, for example, linear or non-linear methods of building consensus sequences such as voting, averaging, statistical, maximum a posteriori or maximum likelihood detection.A representative nucleotide sequence derived from a family of raw reads, generated by analyzing redundant reads to filter out errors introduced during amplification or sequencing.
Duplex tags
(Claim 1)
Input double-stranded deoxyribonucleic acid (DNA) can be converted by a process that tags both halves of the individual double-stranded molecule, in some cases differently. If tagged correctly, each original Watson and Crick (i.e., strand) side of the input double-stranded DNA molecule can be differently tagged and identified by the sequencer and subsequent bioinformatics. This can be performed using a variety of techniques, including ligation of hairpin, bubble, or forked adapters or other adaptors having double-stranded and single stranded segments.A set of adapters used to label double-stranded DNA such that the two complementary strands of a single DNA molecule receive different, distinguishable identifiers, allowing for the subsequent identification of both the Watson and Crick strands.
Families
(Claim 1)
Sequence reads generated from a single original polynucleotide can be used to generate a consensus sequence of that original polynucleotide. The grouping comprises classifying the plurality of sequence reads into families by identifying (i) different molecular barcodes coupled to the plurality of polynucleotide molecules and (ii) similarities between the plurality of sequence reads. Each family includes a plurality of nucleic acid sequences that are associated with a different combination of molecular barcodes and similar or identical sequence reads.Groups of raw sequence reads that are determined to have originated from the same original parent polynucleotide strand based on shared barcode information and genomic alignment coordinates.
Molecular barcodes
(Claim 1)
The set of library adaptors can comprise plurality of polynucleotide molecules with molecular barcodes, wherein the plurality of polynucleotide molecules are less than or equal to 80 nucleotide bases in length, wherein the molecular barcodes are at least 4 nucleotide bases in length. The molecular barcodes are different from one another and have an edit distance of at least 1 between one another. In some embodiments, the given molecular barcode is a randomer.Distinct polynucleotide sequences within library adapters, often at least 4 bases long with a specific edit distance from one another, used to uniquely or non-uniquely identify individual parent polynucleotide fragments.
Selectively enriching
(Claim 1)
In some embodiments, the method further comprises selectively enriching a subset of the tagged DNA fragments. In some embodiments, the method further comprises, after enriching, amplifying the enriched tagged DNA fragments in the presence of sequencing adaptors comprising primers. Selective enrichment can involve separating polynucleotide molecules comprising one or more given sequences from the amplified polynucleotide molecules.The process of isolating or amplifying a specific subset of tagged DNA fragments that correspond to particular genomic regions or loci of interest from a larger population of DNA.
Third quantitative measure
(Claim 1)
The number of unseen molecules can be estimated based on the number of Pairs and Singlets detected. For all molecules in a particular region, counts of molecules where both Watson and Crick sides were recovered (“Pairs”) versus those where only one half was recovered (“Singlets”) can be recorded. If the formula is solved for N (or N3), then N3 (or N) will be inferred, where N3 represents the number of unseen fragments.An inferred estimate of the number of original double-stranded DNA fragments for which neither the Watson nor the Crick strand was successfully sequenced, calculated using the counts of recovered pairs and single strands.

Litigation Cases New

US Latest litigation cases involving this patent.

Case NumberFiling DateTitle
1:22-cv-00334Mar 17, 2022Illumina, Inc. v. Guardant Health, Inc. et al

Patent Family

Patent Family

File Wrapper

The dossier documents provide a comprehensive record of the patent's prosecution history - including filings, correspondence, and decisions made by patent offices - and are crucial for understanding the patent's legal journey and any challenges it may have faced during examination.

  • Get instant alerts for new documents

US9920366

Application Number
US14861989A
Filing Date
Sep 22, 2015
Publication Date
Mar 20, 2018
External Links
Slate, USPTO , Google Patents