Systems and methods to detect rare mutations and copy number variation

Patent No. US10704085 (titled "Systems and methods to detect rare mutations and copy number variation") on Sep 18, 2019. The application was issued on Jul 7, 2020.

What is this patent about?

’085 is related to the field of molecular biology and medical diagnostics, specifically the detection and quantification of rare genetic alterations in cell-free polynucleotides. The technology addresses the challenge of identifying low-frequency mutations and copy number variations (CNVs) within a background of normal germline DNA, which is critical for early cancer detection, disease monitoring, and personalized therapy selection.

The underlying idea behind ’085 is to overcome the inherent noise and distortion of next-generation sequencing by using a high-efficiency molecular tagging and consensus-building strategy. By attaching unique or non-unique identifiers to individual parent molecules before amplification, the system can track progeny reads back to their original source, allowing the software to distinguish between true biological variants and errors introduced during PCR or the sequencing process itself.

The claims of ’085 focus on a method for generating a genetic profile of a tumor by ligating molecular barcodes to both ends of double-stranded cell-free DNA (cfDNA) molecules. The process requires a high-efficiency conversion where at least 20% of the starting molecules are tagged, utilizing a significant molar excess of barcodes. These tagged molecules are then amplified, selectively enriched for cancer-associated target regions, and sequenced to produce reads that are grouped into families based on their barcodes and genomic start/stop positions.

In practice, the invention functions by collapsing these families of sequence reads into consensus sequences. By comparing multiple reads derived from the same parent molecule, the system filters out random sequencing artifacts, effectively increasing the sensitivity of the assay to detect somatic variants at frequencies as low as 0.1%. This digital approach allows for the simultaneous quantification of multiple genetic variants, providing a comprehensive snapshot of a tumor’s genetic heterogeneity from a simple blood draw.

This methodology differs from prior approaches by prioritizing high conversion efficiency and family-based error correction over simple depth of coverage. Traditional sequencing often loses the majority of starting fragments during library preparation, making it difficult to detect rare molecules; however, this invention ensures that a vast majority of the haploid genome equivalents are represented. By using molecular barcodes and consensus sequences to eliminate amplification bias, the system provides a much clearer signal for identifying somatic genetic variants that would otherwise be masked by technical noise.

How does this patent fit in bigger picture?

Technical Landscape

In the mid-2010s when ’085 was filed, the analysis of cell-free nucleic acids for diagnostic purposes was typically implemented using high-throughput sequencing platforms that were subject to inherent per-base error rates. At a time when systems commonly relied on standard mapping and quantification protocols, distinguishing rare genetic alterations from stochastic sequencing noise was non-trivial due to the low concentration of target variants within a high background of wild-type DNA. Furthermore, hardware and software constraints made the high-resolution detection of sub-chromosomal copy number variations difficult, as computational pipelines often lacked the sensitivity to normalize representational biases introduced during library preparation and amplification.

Prosecution Position

The disclosed invention represents a technical advancement through the integration of molecular tagging and computational collapsing techniques to improve the fidelity of cell-free DNA analysis. By attaching barcodes to parent polynucleotides prior to amplification and subsequently collapsing the resulting progeny reads into consensus sequences, the architecture enables the identification and removal of errors introduced during the sequencing process. This structural approach achieves a significant technical effect by enabling the detection of rare mutations and copy number variations with a sensitivity exceeding the raw error rate of the sequencing platform. The advancement is characterized by an architectural shift from simple read counting to a consensus-based modeling system that overcomes the technical constraints of signal-to-noise ratios in liquid biopsy applications.

Claims

This patent contains 30 claims, with claims 1 and 16 serving as the independent claims. The independent claims focus on methods for generating genetic profiles of tumors and quantifying somatic genetic variants from cell-free DNA in bodily fluids by utilizing high-efficiency ligation with a significant molar excess of molecular barcodes to tag, amplify, and sequence nucleic acid molecules. The dependent claims serve to specify technical parameters such as sample concentrations, barcode set sizes, and ligation efficiencies, while also detailing downstream data processing steps like sequence alignment, family grouping, consensus sequence generation, and the identification of specific variant types to inform medical treatment regimens.

Key Claim Terms New

Definitions of key terms used in the patent claims.

Term (Source)Support for SpecificationInterpretation
Consensus sequences
(Claim 1 (referenced via 'families' and 'quantifying'), Claim 16 (referenced via 'detecting' and 'quantifying'))
Collapsing the set of sequencing reads to generate a set of consensus sequences, each consensus sequence corresponding to a unique polynucleotide among the set of tagged parent polynucleotides. The generation of consensus sequences is based on information from the tag and/or at least one of sequence information at the beginning (start) region of the sequence read, the end (stop) regions of the sequence read and the length of the sequence read. Analyzing comprises detecting the tumor marker in the set of consensus sequences.A representative sequence derived by collapsing multiple sequencing reads from the same family (progeny of the same parent molecule) to reduce noise and sequencing errors.
Families
(Claim 1)
Collapsing comprises grouping sequences reads sequenced from amplified progeny polynucleotides into families, each family amplified from the same tagged parent polynucleotide. A consensus sequence is determined based on sequence reads in a family. Assigning, for each family, a confidence score for each of a plurality of calls takes into consideration a frequency of the call among members of the family.Groups of sequencing reads that are determined to have originated from the same original tagged parent polynucleotide based on barcode and alignment coordinates.
Genetic profile
(Claim 1)
The disclosure also provides for a method of characterizing the heterogeneity of an abnormal condition in a subject, the method comprising generating a genetic profile of extracellular polynucleotides in the subject. The genetic profile comprises a plurality of data resulting from copy number variation and/or other rare mutation (e.g., genetic alteration) analyses. In some embodiments, the prevalence/concentration of each rare variant identified in the subject is reported and quantified simultaneously.A comprehensive data set or report characterizing the heterogeneity of an abnormal condition based on detected copy number variations and rare mutations.
Molecular barcodes
(Claim 1, Claim 16)
In some embodiments, the barcode is a polynucleotide, which may further comprise random sequence or a fixed or semi-random set of oligonucleotides that in combination with the diversity of molecules sequenced from a select region enables identification of unique molecules. Barcodes may be at least a 3, 5, 10, 15, 20 25, 30, 35, 40, 45, or 50 mer base pairs in length. The disclosure provides for attaching one or more barcodes to the extracellular polynucleotides or fragments thereof prior to any amplification or enrichment step.Polynucleotide sequences, which may be random, fixed, or semi-random, used to tag parent molecules to enable identification of unique molecules and distinguish them from amplification progeny.
Progeny polynucleotides
(Claim 1, Claim 16)
Amplifying the tagged parent polynucleotides in the set produces a corresponding set of amplified progeny polynucleotides. Sequencing a subset of the set of amplified progeny polynucleotides produces a set of sequencing reads. Collapsing comprises grouping sequences reads sequenced from amplified progeny polynucleotides into families, each family amplified from the same tagged parent polynucleotide.The collection of polynucleotide molecules generated by amplifying the tagged parent polynucleotides, sharing the same barcode and sequence information as the parent.
Selectively enriching
(Claim 1)
In some embodiments, the methods of the disclosure also comprise a step of selectively enriching regions from the subject's genome or transcriptome prior to sequencing. Enriching the set of amplified progeny polynucleotides for polynucleotides mapping to one or more selected reference sequences may be done by: (i) selective amplification of sequences from initial starting genetic material converted to tagged parent polynucleotides; (ii) selective amplification of tagged parent polynucleotides; (iii) selective sequence capture of amplified progeny polynucleotides; or (iv) selective sequence capture of initial starting genetic material.The process of targeting and increasing the relative concentration of specific genomic regions of interest (e.g., cancer-associated loci) prior to sequencing.
Somatic genetic variants
(Claim 1, Claim 16)
The background of normal cfDNA was able to successfully detect somatic mutations down to 0.1% sensitivity. Abnormal conditions may be selected from the group consisting of mutations, rare mutations, single nucleotide variants, indels, copy number variations, transversions, translocations, inversion, deletions, aneuploidy, and cancer. The prevalence/concentration of each rare variant identified in the subject is reported and quantified simultaneously.Non-inherited genetic alterations, such as mutations or copy number variations, typically associated with a tumor or abnormal condition rather than the germline.
Start base position
(Claim 1)
In some embodiments, sequence reads of unique identity may be detected based on sequence information at the beginning (start) and end (stop) regions of the sequence read and the length of the sequence read. In other embodiments sequence molecules of unique identity are detected based on sequence information at the beginning (start) and end (stop) regions of the sequence read, the length of the sequence read and attachment of a barcode. Duplicate polynucleotides have the same start and stop positions.The specific genomic coordinate where a sequence read begins aligning to a reference sequence, used as a unique identifier for the original molecule.
Tagged parent polynucleotides
(Claim 1, Claim 16)
The disclosure provides for a method comprising providing at least one set of tagged parent polynucleotides and amplifying the tagged parent polynucleotides in the set to produce a corresponding set of amplified progeny polynucleotides. Each consensus sequence corresponds to a unique polynucleotide among the set of tagged parent polynucleotides. The method may comprise converting initial starting genetic material into the tagged parent polynucleotides.The initial double-stranded cell-free DNA molecules from a sample after they have been attached to molecular barcodes but prior to amplification.

Litigation Cases New

US Latest litigation cases involving this patent.

Case NumberFiling DateTitle
1:22-cv-00334Mar 17, 2022Illumina, Inc. v. Guardant Health, Inc. et al

Patent Family

Patent Family

File Wrapper

The dossier documents provide a comprehensive record of the patent's prosecution history - including filings, correspondence, and decisions made by patent offices - and are crucial for understanding the patent's legal journey and any challenges it may have faced during examination.

  • Get instant alerts for new documents

US10704085

Application Number
US16575079A
Filing Date
Sep 18, 2019
Publication Date
Jul 7, 2020
External Links
Slate, USPTO , Google Patents