Systems and methods to detect rare mutations and copy number variation

Patent No. US9834822 (titled "Systems and methods to detect rare mutations and copy number variation") on Apr 20, 2017. The application was issued on Dec 5, 2017.

What is this patent about?

’822 is related to the field of molecular diagnostics and bioinformatics, specifically focusing on the detection of rare genetic mutations and copy number variations (CNVs) within cell-free polynucleotides. The technology addresses the challenge of identifying low-abundance genetic signals, such as those derived from tumors or fetuses, which are often masked by the high noise and error rates inherent in standard next-generation sequencing workflows.

The underlying idea behind ’822 is to treat the sequencing process as a communication channel prone to noise and distortion, and to apply a systematic encoding and decoding strategy to recover the original genetic message. By attaching tags to fragmented DNA before amplification, the system creates a way to trace multiple sequence reads back to their single parent molecule, allowing the system to distinguish between true biological variants and artifacts introduced during PCR or sequencing.

The claims of ’822 focus on a method that converts a population of cell-free DNA into non-uniquely tagged parent polynucleotides using identifier sequences like barcodes. The process involves amplifying these tagged molecules, sequencing the resulting progeny, and then grouping the sequence reads into families based on shared barcodes and identical genomic start and stop positions. These families are then collapsed into a single consensus base call at specific genetic loci to determine the true frequency of mutations.

In practice, the invention functions by utilizing the diversity of the naturally occurring fragments in a sample alongside a limited set of barcodes to achieve unique identification of molecules. By collapsing reads into consensus sequences, the system effectively eliminates random errors; for instance, if a mutation appears in only one read of a family but not others, it is discarded as noise. This allows for the detection of rare variants with a sensitivity as low as 0.1%, even when the input DNA is limited to less than 100 nanograms.

This approach differs from prior methods by significantly increasing conversion efficiency—the percentage of original molecules successfully tagged and sequenced—and by using family-based filtering rather than simple quality score thresholds. While traditional sequencing might mistake a 1% error rate for a real mutation, this method uses the redundancy of amplified families to provide a high-fidelity genetic profile, enabling non-invasive monitoring of cancer progression, therapy response, and residual disease.

How does this patent fit in bigger picture?

Technical Landscape

In the early 2010s when ’822 was filed, the analysis of cell-free nucleic acids was typically implemented using standard sequencing protocols where the detection of rare genetic variants was limited by the inherent error rates of the sequencing platforms. At a time when systems commonly relied on high-depth raw read counts to distinguish signal from noise, the identification of low-frequency mutations or subtle copy number variations was often confounded by stochastic errors introduced during library preparation and amplification. Furthermore, when hardware and software constraints made the precise tracking of individual starting molecules non-trivial, bioinformatic pipelines generally processed sequence data in aggregate, which limited the ability to achieve sub-chromosomal resolution or high sensitivity in samples with low concentrations of target genetic material.

Prosecution Position

The disclosed invention represents a meaningful technical advancement through the integration of molecular tagging with a systematic computational collapsing architecture to enhance the sensitivity of genetic analysis. By attaching barcodes to parent polynucleotides prior to amplification and subsequently collapsing the resulting progeny reads into consensus sequences, the system effectively filters out stochastic errors and amplification biases that otherwise obscure rare mutations. This architectural shift enables the detection of variants at frequencies lower than the raw error rate of the sequencing platform itself. The technical effect achieved is a high-fidelity genetic profile that allows for the simultaneous quantification of copy number variations and rare sequence variants from limited starting material, overcoming the constraint of signal-to-noise ratios in cell-free DNA diagnostics.

Claims

This patent contains a total of 20 claims, with claim 1 serving as the sole independent claim. The independent claim focuses on a method for analyzing cell-free DNA through a process of non-unique tagging, amplification, and sequencing, followed by grouping reads into families based on identifiers and genomic positions to collapse data and determine base frequencies at specific genetic loci. The dependent claims serve to further define the technical parameters of the process, including specific conversion efficiencies, the number of unique identifiers used, the types of genetic variants detected, and the inclusion of specific cancer-related gene panels or bioinformatics filtering techniques.

Key Claim Terms New

Definitions of key terms used in the patent claims.

Term (Source)Support for SpecificationInterpretation
Collapsing
(Claim 1)
Collapsing comprising detecting and/or correcting errors, nicks or lesions present in the sense or anti-sense strand of the tagged parent polynucleotides or amplified progeny polynucleotides. The method comprises collapsing the set of sequencing reads to generate a set of consensus sequences, each consensus sequence corresponding to a unique polynucleotide among the set of tagged parent polynucleotides.A computational process of reducing multiple sequencing reads from the same original molecule into a single consensus sequence or base call to filter out errors.
Families
(Claim 1)
Collapsing comprises: i. grouping sequence reads sequences from amplified progeny polynucleotides into families, each family amplified from the same tagged parent polynucleotide; and ii. determining a consensus sequence based on sequence reads in a family.Groups of sequencing reads that are determined to have originated from the same original parent polynucleotide molecule based on shared identifiers and mapping coordinates.
Identifier sequence
(Claim 1)
In some embodiments, the barcode is a polynucleotide, which may further comprise random sequence or a fixed or semi-random set of oligonucleotides that in combination with the diversity of molecules sequenced from a select region enables identification of unique molecules. The barcodes comprise oligonucleotides at least a 3, 5, 10, 15, 20, 25, 30, 35, 40, 45, or 50mer base pairs in length.A polynucleotide sequence, such as a barcode, used to label a DNA molecule to enable its identification and the tracking of its amplified progeny.
Non-uniquely tagged parent polynucleotides
(Claim 1)
In some embodiments, each tagged parent polynucleotide in the set is uniquely tagged. In other embodiments, the tags are non-unique. The generation of consensus sequences is based on information from the tag and/or at least one of sequence information at the beginning (start) region of the sequence read, the end (stop) regions of the sequence read and the length of the sequence read.A collection of original DNA molecules from a sample that have been attached to identifier sequences (barcodes) where the barcodes themselves are not necessarily unique across the entire population, but become unique when combined with other molecular properties.
Start and stop positions
(Claim 1)
In some embodiments, sequence reads of unique identity may be detected based on sequence information at the beginning (start) and end (stop) regions of the sequence read and the length of the sequence read. In other embodiments sequences molecules of unique identity are detected based on sequence information at the beginning (start) and end (stop) regions of the sequence read, the length of the sequence read and attachment of a barcode.The specific genomic coordinates where a DNA fragment begins and ends, used as a physical signature to distinguish unique molecules.

Litigation Cases New

US Latest litigation cases involving this patent.

Case NumberFiling DateTitle
1:22-cv-00334Mar 17, 2022Illumina, Inc. v. Guardant Health, Inc. et al

Patent Family

Patent Family

File Wrapper

The dossier documents provide a comprehensive record of the patent's prosecution history - including filings, correspondence, and decisions made by patent offices - and are crucial for understanding the patent's legal journey and any challenges it may have faced during examination.

  • Get instant alerts for new documents

US9834822

Application Number
US15492659A
Filing Date
Apr 20, 2017
Publication Date
Dec 5, 2017
External Links
Slate, USPTO , Google Patents