Systems and methods to detect rare mutations and copy number variation

Patent No. US9598731 (titled "Systems and methods to detect rare mutations and copy number variation") on May 14, 2015. The application was issued on Mar 21, 2017.

What is this patent about?

’731 is related to the field of genetic diagnostics and bioinformatics, specifically focusing on the high-sensitivity detection of rare mutations and copy number variations within cell-free polynucleotides. The technology addresses the inherent challenges of analyzing circulating tumor DNA (ctDNA), where disease-related genetic signals are often buried under significant background noise from healthy germline DNA and artifacts introduced during the sequencing process.

The underlying idea behind ’731 is to treat the sequencing process as a communication channel prone to noise and distortion, which can be overcome by a sophisticated molecular tagging and collapsing strategy. By attaching identifiers to individual DNA fragments before amplification, the system can distinguish between true biological variants and errors introduced by PCR or the sequencer itself, effectively creating a high-fidelity digital representation of the original sample.

The claims of ’731 focus on a method for quantifying single nucleotide variant tumor markers by utilizing a specific non-unique tagging architecture on at least 10 ng of cell-free DNA. The process involves grouping sequence reads into families based on a combination of barcode sequences and the physical properties of the fragment—such as start/stop positions and length—to derive a consensus sequence for each original parent molecule.

In practice, the invention works by converting fragmented cell-free DNA into a library where each molecule is labeled with a barcode of at least five nucleotides. After these molecules are amplified and sequenced, the system uses the consensus sequences to filter out random errors; if a mutation appears in only one read of a family, it is discarded as noise, whereas mutations present across the family are retained as high-confidence calls. This allows for the detection of rare variants at frequencies as low as 0.1%, which is often below the raw error rate of standard sequencing platforms.

This approach differentiates itself from prior methods by maximizing conversion efficiency and utilizing a hybrid identification strategy that does not require billions of unique barcodes to track individual molecules. By combining relatively small sets of barcodes with the natural diversity of fragment endpoints, the technology achieves the sensitivity required for early cancer detection and longitudinal monitoring of treatment efficacy without the prohibitive costs or technical overhead of traditional molecular counting techniques.

How does this patent fit in bigger picture?

Technical Landscape

In the early 2010s when ’731 was filed, the analysis of cell-free nucleic acids was typically implemented using sequencing protocols where the inherent error rates of the sequencing platforms often exceeded the frequency of rare genetic variants. At a time when systems commonly relied on raw read counts and basic alignment algorithms to quantify genetic material, distinguishing true somatic mutations from stochastic sequencing noise was non-trivial. Furthermore, when hardware and software constraints made the high-fidelity reconstruction of low-input genetic samples difficult, standard practices often struggled to provide the sensitivity required to detect sub-chromosomal copy number variations or rare single-nucleotide variants within the high background of healthy genomic DNA.

Prosecution Position

The disclosed invention represents a technical advancement through an architectural shift in how sequence data is processed, utilizing a molecular tagging and collapsing strategy to overcome the limitations of sequencing noise. By attaching barcodes to parent polynucleotides prior to amplification and subsequently collapsing the resulting progeny reads into consensus sequences, the system enables the identification of unique starting molecules and the filtering of errors introduced during library preparation or sequencing. This integration of molecular barcoding with statistical normalization across predefined genomic regions enables a significant increase in sensitivity, allowing for the simultaneous detection of rare mutations and fractional copy number variations at levels as low as 0.1%.

Claims

The patent contains a total of 17 claims, with claim 1 serving as the sole independent claim. This independent claim focuses on a method for quantifying single nucleotide variant tumor markers in cell-free DNA by utilizing non-unique barcode tagging, sequence grouping into families based on molecular identifiers and read characteristics, and the generation of consensus sequences to identify specific genetic variants at reported tumor loci. The dependent claims serve to further define the process by specifying additional genetic alterations to be detected, detailing the physical parameters of the DNA samples and barcodes, outlining specific bioinformatics filtering and grouping criteria, and describing methods for normalizing data or calculating variant ratios.

Key Claim Terms New

Definitions of key terms used in the patent claims.

Term (Source)Support for SpecificationInterpretation
Cell-free DNA
(Claim 1)
Cell free DNA (“cfDNA”) has been known in the art for decades, and may contain genetic aberrations associated with a particular disease. One approach may include the monitoring of a sample derived from cell free nucleic acids, a population of polynucleotides that can be found in different types of bodily fluids. In some embodiments, extracellular polynucleotides comprise DNA.Extracellular DNA molecules found circulating in bodily fluids, such as blood or plasma, often originating from tumor cells or healthy cells.
Consensus sequences
(Claim 1)
Collapsing the set of sequencing reads to generate a set of consensus sequences, each consensus sequence corresponding to a unique polynucleotide among the set of tagged parent polynucleotides. The generation of consensus sequences is based on information from the tag and/or at least one of (i) sequence information at the beginning (start) region of the sequence read, (ii) the end (stop) regions of the sequence read and (iii) the length of the sequence read. Decoding reduces noise and/or distortion in the message, where noise comprises incorrect nucleotide calls.A representative nucleotide sequence derived by comparing multiple sequencing reads within a family to filter out errors and identify the true sequence of the original parent molecule.
Families
(Claim 1)
Collapsing comprises: (i) grouping sequences reads sequenced from amplified progeny polynucleotides into families, each family amplified from the same tagged parent polynucleotide; and (ii) determining a consensus sequence based on sequence reads in a family. The method further comprises determining a quantitative measure of unique families. Based on the quantitative measure of unique families and the quantitative measure of sequence reads in each group, a measure of unique tagged parent polynucleotides in the set is inferred.Groups of sequencing reads that are determined to have originated from the same original parent polynucleotide molecule based on shared identifiers.
Non-uniquely tagged parent polynucleotides
(Claim 1)
In some embodiments, each barcode attached to extracellular polynucleotides or fragments thereof prior to sequencing is not unique. The barcode is a polynucleotide, which may further comprise random sequence or a fixed or semi-random set of oligonucleotides that in combination with the diversity of molecules sequenced from a select region enables identification of unique molecules. The disclosure also provides for a method comprising detecting genetic variation in non-uniquely tagged initial starting genetic material.Initial DNA molecules from a sample that have been attached to barcodes where the number of available unique barcode sequences is fewer than the number of DNA molecules, resulting in multiple different parent molecules sharing the same barcode sequence.
Single nucleotide variant
(Claim 1)
In some embodiments the variant is a nucleotide variant, single base substitution, or small indel, transversion, translocation, inversion, deletion, truncation or gene truncation about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15 or 20 nucleotides in length. Disorders that are caused by rare genetic alterations (e.g., sequence variants) or changes in epigenetic markers, such as cancer and partial or complete aneuploidy, may be detected. The disclosure also provides for a method for detecting a rare mutation in a cell-free or substantially cell free sample.A genetic alteration consisting of a single base substitution at a specific mappable position in the genome compared to a reference sequence.
Single nucleotide variant tumor markers
(Claim 1)
Disorders that are caused by rare genetic alterations (e.g., sequence variants) or changes in epigenetic markers, such as cancer, may be detected or more accurately characterized with DNA sequence information. In some embodiments the variant is a nucleotide variant, single base substitution, or small indel. The reference sequence is the locus of a tumor marker, and analyzing comprises detecting the tumor marker in the set of consensus sequences.Specific genetic alterations consisting of a single base change at a particular genomic locus that are associated with the presence of a tumor.

Litigation Cases New

US Latest litigation cases involving this patent.

Case NumberFiling DateTitle
1:22-cv-00334Mar 17, 2022Illumina, Inc. v. Guardant Health, Inc. et al

Patent Family

Patent Family

File Wrapper

The dossier documents provide a comprehensive record of the patent's prosecution history - including filings, correspondence, and decisions made by patent offices - and are crucial for understanding the patent's legal journey and any challenges it may have faced during examination.

  • Get instant alerts for new documents

US9598731

Application Number
US14712754A
Filing Date
May 14, 2015
Publication Date
Mar 21, 2017
External Links
Slate, USPTO , Google Patents