Systems and methods to detect rare mutations and copy number variation

Patent No. US11001899 (titled "Systems and methods to detect rare mutations and copy number variation") on Jan 19, 2021. The application was issued on May 11, 2021.

What is this patent about?

’899 is related to the field of genetic analysis and bioinformatics, specifically focusing on the detection of rare mutations and copy number variations within cell-free polynucleotides. In clinical diagnostics, identifying low-frequency genetic aberrations—such as those shed by tumors into the bloodstream—is often hindered by the inherent noise and distortion introduced during PCR amplification and high-throughput sequencing. The background context involves the challenge of distinguishing true somatic variants from technical artifacts when the signal from diseased tissue is extremely dilute compared to healthy germline DNA.

The underlying idea behind ’899 is the use of a controlled tagging strategy to enable the reconstruction of original parent molecules from a pool of noisy sequencing reads. By applying a specific number of molecular barcodes relative to the expected number of duplicate fragments (those sharing identical start and stop positions), the system ensures that individual molecules can be uniquely identified and tracked. This allows the software to collapse multiple progeny reads into a single consensus sequence, effectively filtering out random errors and providing a high-fidelity digital representation of the initial genetic material.

The claims of ’899 focus on a method for detecting cancer by converting a sample of polynucleotides, such as cell-free DNA, into a population of tagged parent molecules using a defined range of barcodes. Specifically, the independent claims require tagging the molecules with n different barcodes, where n is at least 2 but no more than 10,000*z, with z representing the mean number of duplicate molecules having identical start and stop positions. This mathematical constraint on barcode diversity is central to the claimed process of amplifying and sequencing these tagged molecules to identify the presence or absence of malignancy.

In practice, the invention works by extracting cell-free DNA from a bodily fluid and attaching these molecular barcodes before any amplification occurs. By sampling the amplified progeny at a sufficient depth—often 5- to 10-fold coverage per parent molecule—the system can group reads into families derived from the same ancestor. This familial grouping allows the bioinformatics pipeline to assign confidence scores to base calls, ensuring that a mutation is only reported if it appears consistently across the family, rather than appearing as a one-off error from the sequencer.

This approach differs from prior methods by optimizing the conversion efficiency of the library preparation and using the relationship between fragment start/stop positions and barcode diversity to resolve molecular identity. While traditional sequencing might mask a 0.1% tumor signal within a 2% baseline error rate, this method reduces the effective error rate by several orders of magnitude. By focusing on the quantitative measure of unique families rather than raw read counts, the invention also provides a more accurate assessment of copy number variations, negating the effects of uneven amplification bias.

How does this patent fit in bigger picture?

Technical Landscape

In the early 2010s when ’899 was filed, the analysis of cell-free nucleic acids was typically implemented using high-throughput sequencing platforms that were subject to inherent per-base error rates and amplification biases. At a time when genetic variant detection commonly relied on raw read counting rather than molecular tracking, distinguishing true low-frequency somatic mutations from stochastic sequencing noise was non-trivial. Furthermore, system architectures for copy number analysis often struggled with representational biases introduced during library preparation, particularly when working with the low input volumes characteristic of circulating polynucleotide samples.

Prosecution Position

The disclosed invention represents a technical advancement through the integration of molecular tagging and computational collapsing to generate high-fidelity consensus sequences from fragmented extracellular polynucleotides. By utilizing unique or non-unique identifiers in combination with sequence-specific markers like start/stop coordinates and fragment length, the architecture enables the filtering of sequencing artifacts and the normalization of amplification biases. This structural approach allows for the detection of rare mutations and copy number variations with a sensitivity exceeding the baseline error rate of the sequencing hardware, overcoming the technical constraint of signal-to-noise ratios in liquid biopsy diagnostics.

Claims

This patent contains 30 claims, with claims 1 and 27 serving as the independent claims. The independent claims are directed to methods for detecting the presence or absence of cancer in a subject by tagging polynucleotides or cell-free DNA with a specific range of molecular barcodes based on expected duplicate molecules, followed by amplification, sequencing, and data analysis. The dependent claims serve to further define the process by specifying sample types such as blood or tissue, detailing barcode lengths and quantities, outlining enrichment and mapping techniques, and identifying specific genetic or epigenetic markers like single nucleotide variations and copy number variations to guide cancer treatment decisions.

Key Claim Terms New

Definitions of key terms used in the patent claims.

Term (Source)Support for SpecificationInterpretation
Duplicate molecules
(Claim 1, Claim 27)
The method involves determining z, wherein z is a measure of central tendency (e.g., mean, median or mode) of expected number of duplicate polynucleotides starting at any position in the genome. Duplicate polynucleotides have the same start and stop positions. This measure is used to determine the number of unique identifiers (n) needed for tagging.Extracellular polynucleotide fragments that naturally share the exact same genomic start and stop coordinates, occurring independently in the original sample rather than as a result of amplification.
Haploid human genome equivalents
(Claim 27)
The disclosure provides for a method comprising providing a sample comprising between 100 and 100,000 haploid human genome equivalents of cell free DNA (cfDNA) polynucleotides. In some embodiments, the composition comprises between 1000 and 50,000 haploid human genome equivalents. This unit quantifies the initial starting genetic material prior to tagging.A quantitative measure of the amount of genomic DNA in a sample, where one equivalent represents the amount of DNA contained in a single human haploid nucleus.
Molecular barcodes
(Claim 1, Claim 27)
The barcode is a polynucleotide, which may further comprise random sequence or a fixed or semi-random set of oligonucleotides. In combination with the diversity of molecules sequenced from a select region, it enables identification of unique molecules. Barcodes may be at least a 3, 5, 10, 15, 20, 25, 30, 35, 40, 45, or 50 mer base pairs in length.Polynucleotide identifiers, which may be unique or non-unique, used in combination with endogenous sequence properties (like start/stop positions) to distinguish individual molecules in a sample.
Start and stop positions
(Claim 1, Claim 27)
Sequence reads of unique identity may be detected based on sequence information at the beginning (start) and end (stop) regions of the sequence read and the length of the sequence read. In other embodiments, sequence molecules of unique identity are detected based on sequence information at the start and stop regions, the length of the sequence read, and attachment of a barcode. This information helps in collapsing sequencing reads to generate consensus sequences.The specific genomic coordinates corresponding to the beginning and end of a fragmented polynucleotide, used as endogenous identifiers to distinguish unique molecules.
Tagged parent polynucleotides
(Claim 1, Claim 27)
The disclosure provides for a method comprising providing at least one set of tagged parent polynucleotides and amplifying them to produce a corresponding set of amplified progeny polynucleotides. The generation of consensus sequences is based on information from the tag and/or sequence information at the start and stop regions and the length of the sequence read. This enables identification of unique molecules from a select region.A population of original polynucleotide molecules from a sample that have been modified by the attachment of molecular barcodes to enable identification of unique molecules and their progeny after amplification.

Litigation Cases New

US Latest litigation cases involving this patent.

Case NumberFiling DateTitle
1:22-cv-00334Mar 17, 2022Illumina, Inc. v. Guardant Health, Inc. et al

Patent Family

Patent Family

File Wrapper

The dossier documents provide a comprehensive record of the patent's prosecution history - including filings, correspondence, and decisions made by patent offices - and are crucial for understanding the patent's legal journey and any challenges it may have faced during examination.

  • Get instant alerts for new documents

US11001899

Application Number
US17152529A
Filing Date
Jan 19, 2021
Publication Date
May 11, 2021
External Links
Slate, USPTO , Google Patents