Systems and methods to detect rare mutations and copy number variation

Patent No. US10876171 (titled "Systems and methods to detect rare mutations and copy number variation") on May 27, 2020. The application was issued on Dec 29, 2020.

What is this patent about?

’171 is related to the field of genetic analysis and bioinformatics, specifically focusing on the detection of rare mutations and copy number variations within cell-free polynucleotides. The technology addresses the significant challenge of identifying low-frequency somatic variants, such as those shed by tumors into the bloodstream, which are typically obscured by the inherent noise and distortion produced during PCR amplification and high-throughput sequencing processes.

The underlying idea behind ’171 is to treat the sequencing process as a noisy communication channel and apply a digital signaling approach to recover the original genetic message. By attaching molecular barcodes to both ends of individual DNA fragments before amplification, the system creates a traceable link between progeny reads and their unique parent molecules. This allows the system to distinguish between true biological variants and random technical errors by collapsing groups of reads into high-fidelity consensus sequences.

The claims of ’171 focus on a method for detecting somatic genetic variants by tagging cfDNA molecules with a specific number of molecular barcodes relative to the expected number of duplicate fragments. The independent claims require tagging both ends of the DNA fragments and define a barcode diversity (n) that is scaled to a measure of central tendency (z) of fragments sharing the same start and stop positions. This statistical bottlenecking ensures that duplicate molecules can be uniquely identified and grouped into families for accurate variant calling.

In practice, the invention works by mapping the barcoded sequencing reads to a reference genome and grouping them into families based on their unique barcodes and their genomic coordinates. By analyzing the base calls across all members of a family, the system can filter out errors that appear in only a subset of the progeny. This error-correction mechanism enables the detection of mutations at frequencies as low as 0.1%, which is often below the baseline error rate of standard sequencing platforms.

This approach differs from prior methods by optimizing the conversion efficiency of the library preparation and using the relationship between barcode diversity and fragment redundancy to maximize sensitivity. While traditional sequencing often loses rare signals during sample prep or mistakes them for noise, this method uses probabilistic modeling and family-based collapsing to provide a near-perfect representation of the original sample. This allows for non-invasive monitoring of disease progression and treatment efficacy through a simple blood draw.

How does this patent fit in bigger picture?

Technical Landscape

In the early 2010s when ’171 was filed, the analysis of cell-free nucleic acids was typically implemented using sequencing protocols where the inherent error rates of the sequencing platforms often exceeded the frequency of rare genetic variants. At a time when systems commonly relied on standard mapping and quantification of raw sequence reads, distinguishing true somatic mutations from stochastic sequencing noise or amplification artifacts was technically limited. Furthermore, when hardware and software constraints made the high-fidelity reconstruction of original molecular templates non-trivial, the detection of low-frequency sub-chromosomal alterations and copy number variations required significant computational overhead and often lacked the sensitivity necessary for early-stage disease monitoring.

Prosecution Position

The disclosed invention represents a meaningful technical advancement through the integration of molecular tagging and computational collapsing to generate high-confidence consensus sequences from fragmented extracellular polynucleotides. This architectural shift moves beyond simple read counting by utilizing unique identifiers and start/stop coordinates to group progeny reads into families, thereby enabling the identification of unique parent molecules and the filtering of errors introduced during amplification or sequencing. The technical effect achieved is a significant increase in sensitivity, allowing for the detection of rare mutations and copy number variations at frequencies as low as 0.1%, effectively overcoming the technical constraint of sequencing noise in the analysis of low-input, cell-free DNA samples.

Claims

The patent contains a total of 20 claims, with claims 1 and 10 serving as the independent claims. These independent claims focus on a method for detecting somatic genetic variants by tagging cell-free DNA molecules with specific molecular barcodes at both ends, followed by amplification, sequencing, and analysis based on those barcodes and genomic positions. The dependent claims serve to further define the process by specifying sample types like blood or plasma, setting parameters for barcode length and quantity, detailing enrichment for cancer-related target regions, and refining the statistical methods used for base calling and variant detection.

Key Claim Terms New

Definitions of key terms used in the patent claims.

Term (Source)Support for SpecificationInterpretation
Duplicate molecules
(Claim 1, Claim 10)
Duplicate polynucleotides have the same start and stop positions. The variable 'z' is a measure of central tendency (e.g., mean, median or mode) of the expected number of duplicate polynucleotides starting at any position in the genome. This measure helps determine the number of unique identifiers needed for tagging.Naturally occurring or amplified polynucleotides that share identical genomic coordinates, specifically the same start and stop base positions.
Expected number of duplicate molecules
(Claim 1, Claim 10)
The disclosure provides a method comprising: providing a sample comprising a plurality of human haploid genome equivalents of fragmented polynucleotides; determining z, wherein z is a measure of central tendency (e.g., mean, median or mode) of expected number of duplicate polynucleotides starting at any position in the genome, wherein duplicate polynucleotides have the same start and stop positions. This value is used to determine the number of unique identifiers (n) needed for tagging.A measure of central tendency (z) representing the number of distinct cfDNA molecules in a sample that naturally share the same genomic start and stop positions.
Families
(Claim 10)
Collapsing comprises grouping sequences reads sequenced from amplified progeny polynucleotides into families, each family amplified from the same tagged parent polynucleotide. A consensus sequence is then determined based on sequence reads in a family. This allows for the detection of mutations at a frequency less than the error rate introduced at the amplifying step.Groups of sequencing reads that are determined to have originated from the same original tagged parent polynucleotide based on shared barcodes and genomic coordinates.
Molecular barcodes
(Claim 1, Claim 10)
In some embodiments, the barcode is a polynucleotide, which may further comprise random sequence or a fixed or semi-random set of oligonucleotides. In combination with the diversity of molecules sequenced from a select region, it enables identification of unique molecules. Barcodes can be at least a 3, 5, 10, 15, 20, 25, 30, 35, 40, 45, or 50 mer base pairs in length.Polynucleotide sequences, which may be random, fixed, or semi-random, used to uniquely identify individual parent molecules when combined with other molecular properties like start/stop positions.
Somatic genetic variants
(Claim 1, Claim 10)
Disorders that are caused by rare genetic alterations (e.g., sequence variants) or changes in epigenetic markers, such as cancer, may be detected with DNA sequence information. These include mutations, rare mutations, single nucleotide variants, indels, transversions, translocations, inversion, and deletions. The methods can detect somatic mutations down to 0.1% sensitivity.Non-inherited genetic alterations, such as rare mutations or sequence variations, detected in cell-free DNA that differ from a reference genome.
Tagged parent polynucleotides
(Claim 1, Claim 10)
The methods of the disclosure comprise attaching one or more barcodes to the extracellular polynucleotides or fragments thereof prior to any amplification or enrichment step. This converts initial starting genetic material into the tagged parent polynucleotides. Each consensus sequence generated later corresponds to a unique polynucleotide among the set of tagged parent polynucleotides.The initial cell-free DNA molecules from a sample that have been attached to barcodes at both ends prior to any amplification or enrichment steps.

Litigation Cases New

US Latest litigation cases involving this patent.

Case NumberFiling DateTitle
1:22-cv-00334Mar 17, 2022Illumina, Inc. v. Guardant Health, Inc. et al

Patent Family

Patent Family

File Wrapper

The dossier documents provide a comprehensive record of the patent's prosecution history - including filings, correspondence, and decisions made by patent offices - and are crucial for understanding the patent's legal journey and any challenges it may have faced during examination.

  • Get instant alerts for new documents

US10876171

Application Number
US16885079A
Filing Date
May 27, 2020
Publication Date
Dec 29, 2020
External Links
Slate, USPTO , Google Patents