Systems and methods to detect rare mutations and copy number variation

Patent No. US10738364 (titled "Systems and methods to detect rare mutations and copy number variation") on Feb 15, 2019. The application was issued on Aug 11, 2020.

What is this patent about?

’364 is related to the field of molecular diagnostics and bioinformatics, specifically focusing on the high-sensitivity detection of genetic aberrations in cell-free polynucleotides. The technology addresses the inherent challenges of analyzing cell-free DNA (cfDNA), where tumor-derived or fetal-derived genetic material is often present at extremely low concentrations compared to background germline DNA. By integrating advanced sample preparation with computational error-correction, the system aims to distinguish true biological mutations from the noise introduced by PCR amplification and sequencing artifacts.

The underlying idea behind ’364 is the use of a digital sequencing framework that treats individual polynucleotide molecules as discrete information packets to be tracked and verified. The core inventive insight involves non-unique tagging combined with physical molecule characteristics—specifically the start and stop positions of fragments—to create a unique identity for every parent molecule in a sample. This allows the system to collapse multiple sequencing reads into a single high-fidelity consensus sequence, effectively filtering out stochastic errors and providing a precise count of original molecules to detect both rare mutations and structural variations.

The claims of ’364 focus on a specialized system for detecting somatic genetic variants in cfDNA from human blood samples using a combination of molecular barcoding and fragment geometry. The system utilizes a nucleic acid sequencer to process cfDNA molecules that have been joined at both ends with molecular barcodes from a predefined set, where the number of available barcodes is intentionally fewer than the number of molecules mapping to a specific genomic position. The claimed processor then identifies unique parent molecules by cross-referencing these barcode sequences with the specific beginning and end base positions where the reads align to a human reference genome.

In practice, the invention functions by grouping sequencing reads into familial sets derived from the same original parent molecule. The system specifically filters for fragments with a length of 140 to 180 nucleotides, a characteristic size range for cfDNA, to ensure the analysis targets the most relevant biological material. By comparing the sequences within these familial groups, the system can detect a broad spectrum of somatic variants, including single nucleotide variants (SNVs), copy number variations (CNVs), indels, and gene fusions, even when these variants occur at frequencies below the nominal error rate of the sequencing platform.

This approach differs from prior methods by maximizing conversion efficiency and utilizing a hybrid identification strategy that does not rely solely on unique barcodes. Traditional sequencing often loses a significant percentage of the starting material during library preparation, but this system is designed to capture and sequence the vast majority of molecules in a 10 mL blood draw. By leveraging the natural diversity of fragment breakpoints alongside a limited set of barcodes, the technology achieves the sensitivity required for early-stage cancer monitoring and treatment adjustment without the prohibitive complexity of billions of unique tags.

How does this patent fit in bigger picture?

Technical Landscape

In the early 2010s when ’364 was filed, the detection of rare genetic alterations in cell-free DNA was typically implemented using standard sequencing protocols where the inherent error rate of the sequencing platform often exceeded the frequency of the target mutations. At a time when systems commonly relied on high-depth raw sequencing to identify variants, distinguishing true somatic mutations from stochastic noise or amplification artifacts was non-trivial. Furthermore, when hardware or software constraints made the accurate quantification of copy number variations from fragmented extracellular polynucleotides difficult, computational methods were limited by representational biases and the lack of robust molecular tracking to ensure the fidelity of the starting genetic material.

Prosecution Position

The disclosed invention represents a meaningful technical advancement through the integration of molecular tagging and computational collapsing to generate high-fidelity consensus sequences from fragmented cell-free polynucleotides. By attaching barcodes to parent molecules prior to amplification and subsequently grouping progeny reads into families, the architecture enables the suppression of amplification and sequencing errors, allowing for the detection of rare variants at frequencies as low as 0.1%. This structural approach overcomes the technical constraint of sequencing noise and enables the simultaneous quantification of copy number variations and rare mutations by normalizing unique molecular counts across predefined genomic regions, providing a comprehensive genetic profile from a non-invasive bodily sample.

Claims

The patent contains a total of 24 claims, with claim 1 serving as the sole independent claim. This independent claim focuses on a system for detecting somatic genetic variants in cell-free DNA for cancer testing, utilizing a nucleic acid sequencer and a processor programmed to align sequencing reads from non-uniquely tagged molecules and identify specific genetic variations based on molecule length and barcode associations. The dependent claims serve to specify various sequencing technologies, computer hardware configurations, network communication methods, user interface displays, and specific parameters for molecular barcodes and data processing algorithms.

Key Claim Terms New

Definitions of key terms used in the patent claims.

Term (Source)Support for SpecificationInterpretation
Beginning base positions and end base positions
(Claim 1)
In some embodiments, sequence reads of unique identity may be detected based on sequence information at the beginning (start) and end (stop) regions of the sequence read and the length of the sequence read. In other embodiments sequence molecules of unique identity are detected based on sequence information at the beginning (start) and end (stop) regions of the sequence read, the length of the sequence read and attachment of a barcode. The generation of consensus sequences is based on information from the tag and/or at least one of sequence information at the beginning (start) region of the sequence read, the end (stop) regions of the sequence read and the length of the sequence read.The specific genomic coordinates corresponding to the start and stop points of a cfDNA fragment, used in conjunction with barcodes and fragment length to identify unique parent molecules.
Consensus sequences
(Claim 1 (via 'associate... set of mapped sequencing reads'))
Collapsing the set of sequencing reads generates a set of consensus sequences, each corresponding to a unique polynucleotide among the set of tagged parent polynucleotides. Collapsing comprises grouping sequences reads sequenced from amplified progeny polynucleotides into families, each family amplified from the same tagged parent polynucleotide. A consensus sequence is then determined based on sequence reads in a family.A representative sequence generated by collapsing multiple sequencing reads that originated from the same parent polynucleotide molecule to reduce noise and sequencing errors.
Mappable base position
(Claim 1)
The method includes identifying a subset of mapped sequence reads that align with a variant of the reference sequence at each mappable base position. For each mappable base position, a ratio is calculated of (a) a number of mapped sequence reads that include a variant as compared to the reference sequence, to (b) a number of total sequence reads for each mappable base position. In some embodiments, the method comprises providing a plurality of sets of tagged parent polynucleotides, wherein each set is mappable to a different mappable position in the reference sequence.A specific coordinate or location within a reference genome where a sequence read can be uniquely or reliably aligned.
Molecular barcodes
(Claim 1)
In some embodiments, the barcode is a polynucleotide, which may further comprise random sequence or a fixed or semi-random set of oligonucleotides that in combination with the diversity of molecules sequenced from a select region enables identification of unique molecules. The barcodes comprise oligonucleotides at least a 3, 5, 10, 15, 20 25, 30, 35, 40, 45, or 50 mer base pairs in length. Further, the methods of the disclosure comprise attaching one or more barcodes to the extracellular polynucleotides or fragments thereof prior to any amplification or enrichment step.Polynucleotide sequences, which may be random, fixed, or semi-random, joined to the ends of cfDNA molecules to enable the identification and tracking of unique parent molecules during sequencing and analysis.
Non-uniquely tagged cfDNA
(Claim 1)
In some embodiments, each tagged parent polynucleotide in the set is uniquely tagged. In other embodiments, the tags are non-unique. The disclosure also provides for a method comprising detecting genetic variation in non-uniquely tagged initial starting genetic material with a sensitivity of at least 5%, at least 1%, at least 0.5%, at least 0.1% or at least 0.05%.Cell-free DNA molecules that have been attached to molecular barcodes where the number of distinct barcode sequences available is fewer than the number of individual cfDNA molecules mapping to a specific genomic location, necessitating the use of additional identifiers to distinguish molecules.
Somatic genetic variants
(Claim 1)
In some embodiments, the abnormal condition is selected from the group consisting of mutations, rare mutations, indels, copy number variations, transversions, translocations, inversion, deletions, aneuploidy, partial aneuploidy, polyploidy, chromosomal instability, chromosomal structure alterations, gene fusions, chromosome fusions, gene truncations, gene amplification, gene duplications, chromosomal lesions, DNA lesions, abnormal changes in nucleic acid chemical modifications, abnormal changes in epigenetic patterns, abnormal changes in nucleic acid methylation infection and cancer. The prevalence/concentration of each rare variant identified in the subject is reported and quantified simultaneously. Somatic mutations were successfully detected down to 0.1% sensitivity.Non-inherited genetic alterations, including single nucleotide variants, copy number variations, indels, and gene fusions, detected in cell-free DNA typically associated with cancer or other abnormal conditions.

Litigation Cases New

US Latest litigation cases involving this patent.

Case NumberFiling DateTitle
1:22-cv-00334Mar 17, 2022Illumina, Inc. v. Guardant Health, Inc. et al

Patent Family

Patent Family

File Wrapper

The dossier documents provide a comprehensive record of the patent's prosecution history - including filings, correspondence, and decisions made by patent offices - and are crucial for understanding the patent's legal journey and any challenges it may have faced during examination.

  • Get instant alerts for new documents

US10738364

Application Number
US16277724A
Filing Date
Feb 15, 2019
Publication Date
Aug 11, 2020
External Links
Slate, USPTO , Google Patents