Systems and methods to detect rare mutations and copy number variation

Patent No. US10837063 (titled "Systems and methods to detect rare mutations and copy number variation") on Nov 30, 2017. The application was issued on Nov 17, 2020.

What is this patent about?

’063 is related to the field of molecular diagnostics and bioinformatics, specifically focusing on the high-sensitivity detection of genetic aberrations in cell-free DNA (cfDNA). The technology addresses the challenge of identifying rare somatic mutations and copy number variations that are often masked by the inherent noise and distortion of standard next-generation sequencing workflows. By analyzing extracellular polynucleotides found in bodily fluids like blood or plasma, the system provides a non-invasive means of monitoring cancer progression and treatment efficacy.

The underlying idea behind ’063 is to treat the sequencing process as a communication channel where the original DNA molecules are the message and sequencing artifacts are noise. To overcome this noise, the invention utilizes a digital strategy that involves tagging individual DNA fragments with molecular barcodes and using the unique start and stop positions of these fragments to track them. By grouping progeny molecules into families derived from the same original parent molecule, the system can collapse multiple reads into a single consensus sequence, effectively filtering out random errors introduced during PCR amplification or the sequencing run itself.

The claims of ’063 focus on a method for detecting somatic genetic variants by non-uniquely tagging cfDNA molecules with a limited set of molecular barcodes. The process requires that the DNA fragments be flanked on both ends by these barcodes, creating a pool of non-uniquely tagged parent polynucleotides. The method specifically relies on a combination of the barcode sequence and the specific genomic coordinates—the beginning and end base positions—to group sequencing reads into families. This multi-layered identification allows the system to distinguish between true biological variants and technical artifacts across various mutation types, including single nucleotide variants and gene fusions.

In practice, the invention achieves high sensitivity by ensuring that the number of unique barcodes used is significantly smaller than the total number of DNA fragments that map to a specific genomic location. This non-unique tagging approach simplifies the molecular biology required while still providing enough diversity, when combined with fragment length and alignment data, to uniquely identify the original parent molecules. Once the reads are grouped into families, the system applies statistical or probabilistic models to determine the frequency of variants, allowing for the detection of mutations that occur at frequencies as low as 0.1% or less.

This approach differs from prior methods by maximizing the conversion efficiency of the library preparation, ensuring that rare tumor-derived fragments are not lost before they reach the sequencer. While traditional sequencing often requires large amounts of input DNA to overcome losses, this method is optimized for low-input samples typical of liquid biopsies. By shifting the focus from simple read counting to the analysis of molecular families, the technology provides a much clearer picture of the tumor's genetic landscape, enabling real-time monitoring of disease evolution and the emergence of drug resistance.

How does this patent fit in bigger picture?

Technical Landscape

In the early 2010s when ’063 was filed, the detection of rare genetic alterations in cell-free DNA was typically implemented using high-throughput sequencing at a time when inherent sequencing error rates and amplification biases often masked low-frequency variants. Systems commonly relied on standard mapping and counting protocols rather than molecular barcoding for error suppression, and hardware constraints made the high-fidelity reconstruction of original parent molecules from fragmented, low-input samples non-trivial.

Prosecution Position

The disclosed invention represents a technical advancement through the integration of molecular tagging with a consensus-based collapsing architecture to distinguish true biological variants from technical noise. By grouping sequencing reads into families derived from the same parent polynucleotide and applying quality-weighted filtering, the system enables an architectural shift from simple read counting to high-sensitivity detection of rare mutations and copy number variations. This capability allows for the accurate quantification of genetic heterogeneity in cell-free samples, overcoming the technical constraint of the per-base sequencing error rate.

Claims

The patent contains a total of 12 claims, with claim 1 serving as the sole independent claim. This independent claim focuses on a method for detecting somatic genetic variants in cell-free DNA by utilizing a specific non-unique molecular barcoding strategy, where the number of distinct barcodes is fewer than the number of DNA molecules at a given genomic position, followed by amplification, sequencing, and family-based grouping. The dependent claims serve to further define the process by specifying sample types and quantities, barcode library parameters, ligation techniques, consensus sequence generation, and methods for determining base call frequencies or enriching for specific cancer-related target regions.

Key Claim Terms New

Definitions of key terms used in the patent claims.

Term (Source)Support for SpecificationInterpretation
Families
(Claim 1)
Collapsing comprises: i. grouping sequences reads sequenced from amplified progeny polynucleotides into families, each family amplified from the same tagged parent polynucleotide; and ii. determining a consensus sequence based on sequence reads in a family. In some embodiments, collapsing comprises detecting and/or correcting errors, nicks or lesions present in the sense or anti-sense strand of the tagged parent polynucleotides or amplified progeny polynucleotides.Groups of sequencing reads that are determined to have originated from the same original parent polynucleotide molecule based on shared identifiers.
Mappable base position
(Claim 1)
The method comprises identifying a subset of mapped sequence reads that align with a variant of the reference sequence at each mappable base position. For each mappable base position, a ratio is calculated of (a) a number of mapped sequence reads that include a variant as compared to the reference sequence, to (b) a number of total sequence reads for each mappable base position. Ratios or frequency of variance for each mappable base position are normalized to determine potential rare variant(s) or other genetic alteration(s).A specific location or coordinate within a reference genome to which a sequence read can be uniquely or reliably aligned.
Molecular barcodes
(Claim 1)
In some embodiments, the barcode is a polynucleotide, which may further comprise random sequence or a fixed or semi-random set of oligonucleotides that in combination with the diversity of molecules sequenced from a select region enables identification of unique molecules. The barcodes comprise oligonucleotides is at least a 3, 5, 10, 15, 20 25, 30, 35, 40, 45, or 50 mer base pairs in length. Further, the methods of the disclosure comprise attaching one or more barcodes to the extracellular polynucleotides or fragments thereof prior to any amplification or enrichment step.Polynucleotide sequences, which may be random, fixed, or semi-random, attached to DNA fragments to enable the identification and tracking of unique molecules and their amplified progeny.
Non-uniquely tagged parent polynucleotides
(Claim 1)
In some embodiments, each tagged parent polynucleotide in the set is uniquely tagged. In other embodiments, the tags are non-unique. The disclosure also provides for a method comprising detecting genetic variation in non-uniquely tagged initial starting genetic material with a sensitivity of at least 5%, at least 1%, at least 0.5%, at least 0.1% or at least 0.05%.A collection of original cfDNA molecules from a sample that have been attached to barcodes in a manner where the number of distinct barcode identifiers is less than the total number of cfDNA molecules mapping to a specific genomic location, resulting in multiple different parent molecules sharing the same barcode.
Somatic genetic variant
(Claim 1)
In some embodiments, the variant is a nucleotide variant, single base substitution, or small indel, transversion, translocation, inversion, deletion, truncation or gene truncation about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15 or 20 nucleotides in length. The genetic alteration is copy number variation or one or more rare mutations. Disorders that are caused by rare genetic alterations (e.g., sequence variants) or changes in epigenetic markers, such as cancer and partial or complete aneuploidy, may be detected or more accurately characterized with DNA sequence information.A non-inherited genetic alteration, such as an SNV, CNV, indel, or gene fusion, present in the cfDNA of a subject, often associated with cancer.

Litigation Cases New

US Latest litigation cases involving this patent.

Case NumberFiling DateTitle
1:22-cv-00334Mar 17, 2022Illumina, Inc. v. Guardant Health, Inc. et al

Patent Family

Patent Family

File Wrapper

The dossier documents provide a comprehensive record of the patent's prosecution history - including filings, correspondence, and decisions made by patent offices - and are crucial for understanding the patent's legal journey and any challenges it may have faced during examination.

  • Get instant alerts for new documents

US10837063

Application Number
US15828099A
Filing Date
Nov 30, 2017
Publication Date
Nov 17, 2020
External Links
Slate, USPTO , Google Patents