Systems and methods to detect rare mutations and copy number variation

Patent No. US9902992 (titled "Systems and methods to detect rare mutations and copy number variation") on Mar 21, 2016. The application was issued on Feb 27, 2018.

What is this patent about?

’992 is related to the field of molecular diagnostics and bioinformatics, specifically focusing on the high-sensitivity detection of genetic aberrations within cell-free DNA (cfDNA). The technology addresses the challenge of identifying rare mutations and copy number variations that are often obscured by the inherent noise and distortion produced during standard sequencing library preparation and amplification processes.

The underlying idea behind ’992 is to treat polynucleotide sequencing as a communication theory problem, where the original DNA molecules are messages that must be decoded from a noisy signal. By utilizing a high-efficiency tagging process and grouping sequence reads into families derived from the same original molecule, the system can distinguish true biological variants from artifacts introduced by PCR or sequencing errors through the generation of consensus sequences.

The claims of ’992 focus on a method for detecting a diverse set of genetic aberrations by ligating barcoded adaptors to both ends of cfDNA molecules with high efficiency. The process requires using a significant molar excess of adaptors to ensure at least 20% of the initial molecules are tagged, followed by amplifying and sequencing these molecules to produce multiple reads for each original parent polynucleotide.

In practice, the invention works by mapping these reads to a reference human genome and organizing them into families based on their unique barcode sequences. By collapsing these families, the system yields a highly accurate base call for each original molecule at specific genetic loci. This error-reduction mechanism allows for the simultaneous detection of multiple mutation types, including single base substitutions, indels, gene fusions, and copy number variations, from a single bodily sample.

This approach differs from prior methods by significantly increasing the conversion efficiency of the library preparation, ensuring that rare molecules are not lost before sequencing begins. While traditional sequencing often struggles with a high background error rate, this method uses familial grouping to filter out stochastic noise, enabling the detection of mutations at frequencies as low as 0.1%, which is critical for early cancer detection and monitoring treatment response.

How does this patent fit in bigger picture?

Technical Landscape

In the early 2010s when ’992 was filed, the analysis of cell-free DNA for diagnostic purposes was typically implemented using high-throughput sequencing platforms that were subject to inherent per-base error rates. At a time when systems commonly relied on standard mapping and quantification of raw sequence reads, distinguishing rare genetic alterations from stochastic sequencing noise was non-trivial. Engineering constraints in library preparation and amplification often introduced representational biases, making the precise detection of low-frequency somatic mutations or subtle copy number variations difficult when working with the limited quantities of fragmented polynucleotides typically found in bodily fluids.

Prosecution Position

The disclosed invention represents a technical advancement through the integration of molecular tagging and computational collapsing to improve the fidelity of cell-free polynucleotide analysis. By attaching barcodes to parent polynucleotides prior to amplification and subsequently collapsing the resulting progeny reads into consensus sequences, the architecture enables the differentiation of true biological variants from errors introduced during sequencing or amplification. This structural approach overcomes the technical constraint of high background noise, enabling the detection of rare mutations and copy number variations with a sensitivity exceeding the raw error rate of the sequencing platform. The system achieves a significant capability shift by providing a statistical framework for normalizing read counts across predefined genomic regions, allowing for the characterization of genetic heterogeneity and disease progression from non-invasive samples.

Claims

The patent contains a total of 33 claims, with claim 1 serving as the sole independent claim. This independent claim focuses on a method for detecting multiple types of genetic aberrations in cell-free DNA by utilizing high-efficiency barcode tagging, molar excess ligation, and family-based sequence read collapsing to identify variations such as base substitutions, copy number changes, indels, and gene fusions. The dependent claims serve to specify technical parameters and variations of the process, including specific DNA input amounts, barcode lengths and types, enrichment strategies for particular genes, sequencing depth requirements, and the specific combinations or sensitivities of the genetic aberrations being detected.

Key Claim Terms New

Definitions of key terms used in the patent claims.

Term (Source)Support for SpecificationInterpretation
Barcodes
(Claim 1)
In some embodiments, the barcode is a polynucleotide, which may further comprise random sequence or a fixed or semi-random set of oligonucleotides that in combination with the diversity of molecules sequenced from a select region enables identification of unique molecules. The barcodes comprise oligonucleotides at least a 3, 5, 10, 15, 20 25, 30, 35, 40, 45, or 50mer base pairs in length. In some embodiments, each barcode attached to extracellular polynucleotides or fragments thereof prior to sequencing is unique.Polynucleotide sequences, which may be random, fixed, or semi-random, used in combination with other molecule properties to uniquely identify and track individual parent DNA molecules.
Barcode sequence
(Claim 1)
In some embodiments, the barcode is a polynucleotide, which may further comprise random sequence or a fixed or semi-random set of oligonucleotides that in combination with the diversity of molecules sequenced from a select region enables identification of unique molecules. The barcodes comprise oligonucleotides at least a 3, 5, 10, 15, 20 25, 30, 35, 40, 45, or 50mer base pairs in length. In some embodiments, each barcode attached to extracellular polynucleotides or fragments thereof prior to sequencing is unique.A specific polynucleotide sequence within a tag used to identify and group sequence reads that originated from the same initial DNA molecule.
Collapsing
(Claim 1)
Collapsing the set of sequencing reads to generate a set of consensus sequences, each consensus sequence corresponding to a unique polynucleotide among the set of tagged parent polynucleotides. Collapsing comprising detecting and/or correcting errors, nicks or lesions present in the sense or anti-sense strand of the tagged parent polynucleotides or amplified progeny polynucleotides. Decoding [collapsing] comprises grouping sequence reads of amplified molecules amplified from each of the at least one individual polynucleotide molecules.A computational process of reducing multiple sequencing reads within a family into a single consensus sequence or base call to reduce noise and sequencing errors.
Copy number variation
(Claim 1)
Determining a copy number variation in one or more of the predefined regions by (i) normalizing the number of reads in the predefined regions to each other and/or the number of unique barcodes in the predefined regions to each other; and (ii) comparing the normalized numbers obtained in step (i) to normalized numbers obtained from a control sample. In some embodiments, the percent of sequences having copy number variation in said bodily sample is determined by calculating the percentage of predefined regions with an amount of polynucleotides above or below a predetermined threshold.A genetic aberration characterized by an increase or decrease in the number of copies of a specific genomic region compared to a reference or control.
Families
(Claim 1)
Collapsing comprises: i. grouping sequences reads sequenced from amplified progeny polynucleotides into families, each family amplified from the same tagged parent polynucleotide; and ii. determining a consensus sequence based on sequence reads in a family. In some embodiments, the method further comprises: a) determining a quantitative measure of unique families; and b) based on (1) the quantitative measure of unique families and (2) the quantitative measure of sequence reads in each group, inferring a measure of unique tagged parent polynucleotides in the set.Groups of sequencing reads that are determined to have originated from the same original tagged parent polynucleotide based on shared barcode sequences and/or mapping coordinates.
Genetic aberrations
(Claim 1)
In some embodiments, bodily fluids are drawn from a subject suspected of having an abnormal condition which may be selected from the group consisting of, mutations, rare mutations, single nucleotide variants, indels, copy number variations, transversions, translocations, inversion, deletions, aneuploidy, partial aneuploidy, polyploidy, chromosomal instability, chromosomal structure alterations, gene fusions, chromosome fusions, gene truncations, gene amplification, gene duplications, chromosomal lesions, DNA lesions, abnormal changes in nucleic acid chemical modifications, abnormal changes in epigenetic patterns, abnormal changes in nucleic acid methylation infection and cancer.Specific alterations in the DNA sequence or structure, including single base substitutions, copy number variations, insertions, deletions, or gene fusions, often associated with cancer or fetal abnormalities.
Tagged parent polynucleotides
(Claim 1)
The disclosure provides for a method comprising: a. providing at least one set of tagged parent polynucleotides, and for each set of tagged parent polynucleotides; b. amplifying the tagged parent polynucleotides in the set to produce a corresponding set of amplified progeny polynucleotides. In some embodiments, the method further comprises converting initial starting genetic material into the tagged parent polynucleotides. Converting comprises any of blunt-end ligation, sticky end ligation, molecular inversion probes, PCR, ligation-based PCR, single strand ligation and single strand circularization.Original cell-free DNA molecules from a subject that have been modified by the ligation of barcode-containing adaptors to both ends, serving as the unique templates for subsequent amplification.

Litigation Cases New

US Latest litigation cases involving this patent.

Case NumberFiling DateTitle
1:25-cv-00263Mar 6, 2025Cold Spring Harbor Laboratory V. Guardant Health, Inc.
1:22-cv-00334Mar 17, 2022Illumina, Inc. v. Guardant Health, Inc. et al

Patent Family

Patent Family

File Wrapper

The dossier documents provide a comprehensive record of the patent's prosecution history - including filings, correspondence, and decisions made by patent offices - and are crucial for understanding the patent's legal journey and any challenges it may have faced during examination.

  • Get instant alerts for new documents

US9902992

Application Number
US15076565A
Filing Date
Mar 21, 2016
Publication Date
Feb 27, 2018
External Links
Slate, USPTO , Google Patents