Systems and methods to detect rare mutations and copy number variation

Patent No. US10947600 (titled "Systems and methods to detect rare mutations and copy number variation") on Oct 12, 2020. The application was issued on Mar 16, 2021.

What is this patent about?

’600 is related to the field of genetic analysis and bioinformatics, specifically focusing on the high-sensitivity detection of rare mutations and copy number variations (CNVs) within cell-free DNA (cfDNA). The technology addresses the challenge of identifying minute quantities of tumor-derived or fetal genetic material that co-circulate with a vast majority of healthy germline DNA in bodily fluids like blood plasma.

The underlying idea behind ’600 is to treat the sequencing process as a noisy communication channel and apply a digital-inspired error-correction strategy to recover the original genetic message. By attaching molecular barcodes to individual fragments before amplification, the system can track all progeny molecules back to their single parent strand. This allows the system to distinguish between true biological variants and the inevitable noise introduced by PCR amplification or sequencing errors, effectively collapsing multiple reads into a single, high-fidelity consensus sequence.

The claims of ’600 focus on a specific method for preparing cfDNA libraries by ligating adapters containing a calculated number of molecular barcodes. The number of unique barcodes, *n*, is mathematically constrained based on the expected number of duplicate molecules—those naturally sharing the same start and stop positions—to ensure that biological duplicates can be distinguished from one another. The process involves tagging these parent molecules, performing universal amplification, and then selectively enriching the progeny for specific genomic regions of interest.

In practice, the invention works by first bottlenecking or quantifying the initial genetic material to determine the diversity of the sample. By using a specific ratio of barcodes to the mean number of expected duplicate fragments, the method ensures a high probability that any two identical fragments from different cells receive different tags. After sequencing, the bioinformatics pipeline groups reads into families based on these barcodes and their genomic coordinates, allowing the system to filter out stochastic errors that appear in only a subset of a family’s reads.

This approach differs from prior methods by significantly increasing the conversion efficiency of the library preparation and using digital collapsing to bypass the standard error rates of sequencing platforms. While traditional sequencing might struggle to detect mutations below a 1% frequency due to background noise, this method can reliably identify rare variants at frequencies as low as 0.1%. This sensitivity is achieved by shifting the focus from individual read quality to the collective evidence provided by an entire family of molecules derived from a single original template.

How does this patent fit in bigger picture?

Technical Landscape

In the early 2010s when ’600 was filed, the detection of rare genetic alterations in cell-free DNA was typically implemented using standard sequencing protocols where the inherent error rates of the sequencing platform often exceeded the frequency of the target mutations. At a time when systems commonly relied on high-depth raw sequencing reads to identify variants, distinguishing true biological signals from stochastic noise and amplification artifacts was limited by the lack of molecular tracking mechanisms. Furthermore, when hardware and software constraints made the high-resolution analysis of sub-chromosomal copy number variations non-trivial, the field was restricted by the difficulty of normalizing representational biases across the genome in low-input samples.

Prosecution Position

The disclosed invention represents a meaningful technical advancement through the integration of molecular tagging and computational collapsing techniques to generate high-fidelity consensus sequences from fragmented extracellular polynucleotides. This architectural shift enables the suppression of sequencing and amplification errors, allowing for the detection of rare somatic mutations with a sensitivity as low as 0.1% even in samples with limited starting material. By quantifying unique parent molecules through the use of barcodes and start/stop coordinates, the system overcomes the technical constraint of representational bias, enabling simultaneous and accurate reporting of both sequence variants and fractional copy number variations across the genome.

Claims

This patent contains 30 claims, with claims 1 and 16 serving as the independent claims. These independent claims focus on a method for preparing cell-free DNA molecules for sequencing by attaching a specific range of molecular barcodes based on expected duplicate molecule counts, followed by amplification and selective enrichment of specific genomic regions. The dependent claims serve to further define the process by specifying sample sources such as blood or urine from cancer patients, narrowing the barcode quantity and length, detailing ligation techniques, and identifying specific genomic targets like oncogenes or tumor suppressor genes.

Key Claim Terms New

Definitions of key terms used in the patent claims.

Term (Source)Support for SpecificationInterpretation
Duplicate molecules
(Claim 1, Claim 16)
Determining z, wherein z is a measure of central tendency (e.g., mean, median or mode) of expected number of duplicate polynucleotides starting at any position in the genome, wherein duplicate polynucleotides have the same start and stop positions. Sequence reads of unique identity may be detected based on sequence information at the beginning (start) and end (stop) regions of the sequence read and the length of the sequence read.Multiple distinct polynucleotide molecules in a sample that happen to share the exact same genomic start and stop positions.
Molecular barcodes
(Claim 1, Claim 16)
In some embodiments, the barcode is a polynucleotide, which may further comprise random sequence or a fixed or semi-random set of oligonucleotides that in combination with the diversity of molecules sequenced from a select region enables identification of unique molecules. The barcodes comprise oligonucleotides is at least a 3, 5, 10, 15, 20 25, 30, 35, 40, 45, or 50 mer base pairs in length. Barcodes can be attached to extracellular polynucleotides or fragments thereof prior to any amplification or enrichment step.Polynucleotide sequences, which may include random or fixed sets of oligonucleotides, used to tag individual parent molecules to enable identification of unique molecules and distinguish them from duplicates.
Progeny polynucleotides
(Claim 1, Claim 16)
Amplifying the tagged parent polynucleotides in the set to produce a corresponding set of amplified progeny polynucleotides. Sequencing a subset of the set of amplified progeny polynucleotides produces a set of sequencing reads. Collapsing comprises grouping sequences reads sequenced from amplified progeny polynucleotides into families, each family amplified from the same tagged parent polynucleotide.The collection of polynucleotide molecules produced by amplifying the original tagged parent molecules, typically representing the library to be sequenced.
Selectively enriching
(Claim 1)
In some embodiments, the method comprises enriching the set of amplified progeny polynucleotides for polynucleotides mapping to one or more selected reference sequences by: (i) selective amplification of sequences from initial starting genetic material converted to tagged parent polynucleotides; (ii) selective amplification of tagged parent polynucleotides; (iii) selective sequence capture of amplified progeny polynucleotides; or (iv) selective sequence capture of initial starting genetic material.The process of increasing the relative concentration of specific genomic regions of interest from the total population of polynucleotides, often through sequence capture or targeted amplification.
Tagged parent polynucleotides
(Claim 1, Claim 16)
The method further comprises converting initial starting genetic material into the tagged parent polynucleotides. In some embodiments, the initial starting genetic material is cell-free nucleic acid. Each consensus sequence corresponds to a unique polynucleotide among the set of tagged parent polynucleotides.The initial starting genetic material, such as cell-free DNA fragments, after they have been ligated to adapters containing molecular barcodes but prior to amplification.

Litigation Cases New

US Latest litigation cases involving this patent.

Case NumberFiling DateTitle
1:22-cv-00334Mar 17, 2022Illumina, Inc. v. Guardant Health, Inc. et al

Patent Family

Patent Family

File Wrapper

The dossier documents provide a comprehensive record of the patent's prosecution history - including filings, correspondence, and decisions made by patent offices - and are crucial for understanding the patent's legal journey and any challenges it may have faced during examination.

  • Get instant alerts for new documents

US10947600

Application Number
US17068710A
Filing Date
Oct 12, 2020
Publication Date
Mar 16, 2021
External Links
Slate, USPTO , Google Patents