HK1237376A1 - Universal blocking oligo system and improved hybridization capture methods for multiplexed capture reactions - Google Patents

Universal blocking oligo system and improved hybridization capture methods for multiplexed capture reactions Download PDF

Info

Publication number
HK1237376A1
HK1237376A1 HK17111448.6A HK17111448A HK1237376A1 HK 1237376 A1 HK1237376 A1 HK 1237376A1 HK 17111448 A HK17111448 A HK 17111448A HK 1237376 A1 HK1237376 A1 HK 1237376A1
Authority
HK
Hong Kong
Prior art keywords
nucleic acid
nucleic acids
block
library
composition
Prior art date
Application number
HK17111448.6A
Other languages
Chinese (zh)
Inventor
E‧奥利瓦瑞斯
Original Assignee
因维蒂公司
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by 因维蒂公司 filed Critical 因维蒂公司
Publication of HK1237376A1 publication Critical patent/HK1237376A1/en

Links

Description

Universal blocking oligomer systems for multiple capture reactions and improved methods of hybrid capture
Related patent application
This patent application claims the benefit OF U.S. provisional patent application No.62/062612 entitled "Universal BLOCKING OLIGOSYSTEM FOR MULTIPLE LEXED CAPTURE REACTIONS" filed 10/2015, entitled Eric Olivares, and labeled attorney docket No. 055911-. The foregoing patent application is incorporated by reference herein in its entirety, including all texts, tables and figures.
Technical Field
The present technology relates in part to compositions and methods for nucleic acid manipulation, analysis, and high-throughput sequencing.
Background
Genetic information of living organisms (e.g., animals, plants, microorganisms, viruses) is encoded in deoxyribonucleic acid (DNA) or ribonucleic acid (RNA). Genetic information is a series of nucleotides or modified nucleotides that represent the primary structure of a nucleic acid. The nucleic acid content (e.g., DNA) of an organism is often referred to as the genome. In humans, the entire genome typically contains about 30,000 genes located on twenty-four (24) chromosomes. Most genes encode specific proteins that, after being expressed by transcription and translation, fulfill specific biochemical functions within a living cell.
Many medical conditions are caused by one or more genetic variations within the genome. Some genetic variations may predispose an individual to or develop any of a number of diseases such as diabetes, arteriosclerosis, obesity, various autoimmune diseases, and cancer (e.g., colorectal, breast, ovarian, lung). Such genetic diseases may be caused by the addition, substitution, insertion or deletion of one or more nucleotides within the genome.
Genetic variations can be identified by multiplex analysis of nucleic acid mixtures, typically obtained from multiple sources, for example by using next generation sequencing techniques. Such multiplex assays typically involve extensive nucleic acid manipulation prior to analysis, involving many different steps that are not conducive to high throughput processing. Furthermore, existing nucleic acid manipulation methods are often expensive, time consuming, and often suffer from significant drawbacks that can lead to sample contamination. The compositions and methods herein provide significant improvements over existing nucleic acid manipulation and analysis techniques, which are more conducive to high-throughput automation, more cost-effective, less time-consuming, and/or provide less risk of contamination.
Disclosure of Invention
In some aspects, presented herein is a composition for massively parallel nucleic acid sequencing, comprising a) a library of nucleic acids comprising a plurality of library inserts, wherein each nucleic acid of the library comprises (i) at least one library insert obtained from one of four or more samples, (ii) a first non-natural nucleic acid, and (iii) a second non-natural nucleic acid, wherein the first non-natural nucleic acid and the second non-natural nucleic acid are located on opposite sides of the at least one library insert, and the first non-natural nucleic acid comprises a first distinguishable nucleic acid barcode and the second non-natural nucleic acid comprises a second distinguishable nucleic acid barcode, wherein the first and second distinguishable nucleic acid barcodes are unique to one of the four or more samples; and b) four U-block nucleic acids, wherein (i) first and second U-block nucleic acids are configured to hybridize to the first non-natural nucleic acid on opposite sides of the first distinguishable nucleic acid barcode, and (ii) third and fourth U-block nucleic acids are configured to hybridize to the second non-natural nucleic acid on opposite sides of the second distinguishable nucleic acid barcode, and (iii) each of the U-block nucleic acids does not substantially hybridize to a portion of the first or second distinguishable nucleic acid barcode. In certain aspects, the library of nucleic acids comprises at least eight distinguishable nucleic acid barcodes.
In some aspects, the composition further comprises one or more capture nucleic acids, wherein (i) the capture nucleic acids comprise a member of a binding pair, and (ii) the capture nucleic acids are each configured to specifically hybridize to a subset of the one or more library inserts.
In certain embodiments, also presented herein is a method of analyzing a library of nucleic acids, comprising a) obtaining a library of nucleic acids comprising a plurality of library inserts, wherein each nucleic acid of the library comprises (i) at least one library insert obtained from one of four or more samples, (ii) a first non-natural nucleic acid, and (iii) a second non-natural nucleic acid, wherein the first non-natural nucleic acid and the second non-natural nucleic acid are located on opposite sides of the at least one library insert, and the first non-natural nucleic acid comprises a first distinguishable nucleic acid barcode and the second non-natural nucleic acid comprises a second distinguishable nucleic acid barcode, wherein the first and second distinguishable nucleic acid barcodes are unique to one of the four or more samples; b) contacting the library of nucleic acids with four U-block nucleic acids, wherein (i) first and second U-block nucleic acids are configured to hybridize to the first non-natural nucleic acid on opposite sides of the first distinguishable nucleic acid barcode, and (ii) third and fourth U-block nucleic acids are configured to hybridize to the second non-natural nucleic acid on opposite sides of the second distinguishable nucleic acid barcode, and (iii) each of the U-block nucleic acids is not substantially hybridized to a portion of the first or second distinguishable nucleic acid barcode; and c) contacting the library of nucleic acids with one or more capture nucleic acids, each of the one or more capture nucleic acids comprising a first member of a binding pair, wherein the one or more capture nucleic acids are configured to specifically hybridize to a subset of the nucleic acids of the library; d) capturing said captured nucleic acids, thereby providing captured nucleic acids comprising a subset of the nucleic acids of the library; e) contacting the captured nucleic acids with a primer set under amplification conditions, thereby providing amplicons; and f) analyzing the amplicons.
In certain aspects, analyzing comprises providing sequence reads. In some aspects, sequence reads can be obtained by methods including massively parallel sequencing and/or double-ended sequencing.
In certain aspects with respect to the compositions and methods herein, the non-natural nucleic acid comprises a universal nucleic acid. In some aspects, the nucleic acids of the library comprise four or more, or ten or more barcode nucleic acids. In some aspects, each library insert comprises one or two barcode sequences. In certain aspects, a U-block nucleic acid comprises a length of 10 to 40 nucleotides. In certain aspects, a U-block nucleic acid comprises a length of 10 to 20 nucleotides. In some aspects, the U-block nucleic acid comprises a locked nucleic acid and/or a bridged nucleic acid. In certain aspects, the U-block comprises a melting temperature of about 65 ℃ to about 90 ℃. In certain aspects, the U-block nucleic acid comprises a melting temperature of at least 65 ℃ or at least 75 ℃. In some aspects, the U-block nucleic acid does not comprise degenerate (degenerate) nucleotide bases. In some aspects the U-block nucleic acid does not comprise 3-nitropyrrole, 5-nitroindole, inosine, 2' -deoxyinosine, an analog, a derivative, or a combination thereof.
In some aspects, provided herein is a method of analyzing a nucleic acid library, comprising a) obtaining a library of nucleic acids comprising a first set of amplicons, wherein each amplicon comprises a first non-natural nucleic acid and a second non-natural nucleic acid, one or more distinguishable identifiers, and a library insert obtained from one of the one or more samples, wherein the library insert is located between the first and second non-natural nucleic acids; b) preparing a mixture comprising contacting nucleic acids of the library with one or more blocker nucleic acids and a capture nucleic acid, wherein (i) the one or more blocker nucleic acids are configured to specifically hybridize to the first and second non-native nucleic acids, (ii) the capture nucleic acid comprises a first member of a binding pair, and (iii) the capture nucleic acid is configured to specifically hybridize to a subset of the first set of amplicons, c) purifying the mixture to provide a purified nucleic acid, wherein the purified nucleic acid comprises the nucleic acids of the library, the one or more blocker nucleic acids and the capture nucleic acid, d) hybridizing the purified nucleic acid under hybridization conditions, e) capturing the capture nucleic acid to provide a captured nucleic acid, f) contacting the captured nucleic acid with a set of primers under amplification conditions to provide a second set of amplicons, and g) analyzing the second set of amplicons. In some aspects, the amplification conditions comprise a thermostable polymerase and/or polymerase chain reaction. In certain aspects, the preparing in (b) comprises contacting the nucleic acids of the library with competitor nucleic acids. In some embodiments, the capture nucleic acid is configured to specifically hybridize to a portion of the library insert. In certain embodiments, the one or more blocking nucleic acids are configured to specifically hybridize to a portion of the first non-natural nucleic acid and/or the second non-natural nucleic acid. In certain embodiments, the one or more blocking nucleic acids comprise locked nucleic acids and/or bridged nucleic acids.
In some aspects, a capture nucleic acid comprising a first member of a binding pair is configured to specifically hybridize to a portion of an exon, an intron, a portion of a selected chromosome, and/or a region of DNA comprising a genetic variation (e.g., a duplication, a polymorphism). In some embodiments, the first member of the binding pair comprises biotin, an antigen, a hapten, an antibody, or a portion thereof. In some aspects, the capturing in (e) comprises contacting the mixture with a second member of the binding pair. In some aspects, the second member of the binding pair comprises avidin, protein a, protein G, an antibody, or a binding portion thereof. In certain embodiments, the second member of the binding pair comprises a matrix. In some embodiments, the matrix comprises a magnetic compound. In some embodiments, the matrix comprises beads. In some embodiments, the matrix comprises polystyrene, polycarbonate, agarose gel, or agarose. In some embodiments, the matrix comprises a metal.
In certain embodiments, the hybridization conditions comprise denaturation. In certain embodiments, the hybridization conditions comprise hybridizing the captured nucleic acids to portions of one or more of the first set of amplicons. In certain embodiments, the hybridization conditions comprise incubating the captured nucleic acids at a temperature of about 25 ℃ to about 70 ℃. In certain embodiments, the hybridization conditions comprise incubating the captured nucleic acids at a temperature of about 35 ℃ to about 60 ℃. In certain embodiments, the hybridization conditions comprise incubation for an amount of time of about 1 hour to 24 hours or about 12 hours to 20 hours. In certain embodiments, the hybridization conditions do not include a polymerase. In some embodiments, the hybridizing in (d) comprises contacting the mixture with a hybridization buffer. In some embodiments, the hybridization in (d) comprises sequential steps of (i) contacting the mixture with a hybridization buffer, (ii) denaturing, and (iii) hybridizing.
In some aspects, analyzing comprises providing sequence reads. Sometimes sequence reads are obtained by methods that include next generation sequencing (e.g., massively parallel sequencing). Sometimes sequencing reads are obtained by methods that include paired-end sequencing.
In certain embodiments, the first non-native nucleic acid comprises at least one nucleic acid barcode. In certain embodiments, the second non-native nucleic acid comprises at least one nucleic acid barcode.
In certain embodiments, the methods claimed herein do not include a drying step. In some embodiments, the method does not comprise a denaturation step prior to (c). In some embodiments, the method does not include a denaturation step prior to (d). In certain embodiments, the method does not include heating to a temperature greater than 80 ℃ prior to (d). In certain embodiments, the method does not include heating to a temperature greater than 90 ℃ prior to (d). In some embodiments, prior to (e), the mixture is not immobilized on a substrate of a flow cell or array. In some embodiments, the purifying in (c) does not include adding a second member of the binding pair configured to bind the first member of the binding pair.
In some aspects, the sample may be obtained from one or more species, one or more tissues, one or more mammals, or one or more human subjects.
Certain embodiments are further described in the following specification, examples, claims and drawings.
Drawings
The drawings illustrate embodiments of the present technology and are not intended to be limiting. The drawings are not to scale and, in some instances, various aspects may be shown exaggerated or enlarged to facilitate an understanding of the detailed description for clarity and ease of illustration.
FIG. 1 shows an embodiment of a blocking method comprising four U-block nucleic acids (A ', C', E ', and G'). FIG. 1 shows representative nucleic acids of a library (Z) comprising library inserts (D) and distinguishable nucleic acid barcodes (B and F), wherein a plurality of different inserts and different distinguishable barcodes are present in a plurality of nucleic acids of the library. FIG. 1 shows U-blocking nucleic acids (A ', C', E ', and G'), each configured to specifically hybridize to a portion of the non-native nucleic acid (A, C, E and G) shown, and the U-blocking nucleic acid hybridizes adjacent to a nucleic acid barcode (B or F).
Detailed Description
Next Generation Sequencing (NGS) allows nucleic acids to be sequenced and analyzed genome-wide by methods that are faster and less expensive than traditional sequencing methods. The methods and compositions herein provide improvements in advanced sequencing technologies that can be used to locate and identify genetic variations and/or related diseases and disorders. In some embodiments, provided herein are methods that include, in part, the manipulation and preparation of nucleic acid cocktails for NGS.
Sequencing applications using genomic nucleic acid as a target material typically require the selection of a nucleic acid target of interest from a highly complex mixture. The quality of the sequencing work is often dependent on the efficiency of the selection process, which in turn depends on how much the nucleic acid target can be enriched relative to the non-target sequence. Selection and enrichment of nucleic acid libraries sometimes includes capture of adaptor-ligated inserts (e.g., genomic DNA inserts) by hybrid capture methods.
Most next generation sequencing library molecules contain non-native sequences (e.g., adaptor nucleic acids, barcode sequences, primer binding sites, and universal sequences) that enable subsequent sequencing. During the hybrid capture reaction, the non-native sequences can anneal to each other, resulting in contamination of the enriched nucleic acid pool. Most of these unwanted sequences are often due to undesired hybridization events between portions of the terminal adaptor sequences ligated to the library inserts. Sometimes, multiple library inserts can non-specifically anneal to each other through their terminal adaptors, resulting in a "daisy chain" of otherwise unwanted DNA fragments that are ligated and separated together.
One way to reduce the so-called "daisy chain" effect is to use a blocking nucleic acid that is intended to hybridize to a large portion of the adapter sequence. For traditional methods, blocking nucleic acids are required on each side of the adaptors, and each blocking nucleic acid contains a perfect complementary match to the adaptor sequence (including barcode sequences (e.g., index sequences)) contained in each adaptor. For high-throughput multiplex sequencing methods, multiple libraries are typically mixed, each library consisting of a different adapter sequence and a different barcode sequence. For such multiplex methods, it is necessary to synthesize multiple sets of classical blocking nucleic acids, each specific for an adaptor of each library. This approach is cumbersome and expensive, and requires the manufacture of many different, relatively long oligonucleotides, which prevents efficient and cost-effective automation of the library preparation and sequencing processes.
In some aspects, provided herein are new and improved compositions and methods for reducing unwanted capture events. In some embodiments, provided herein are novel U-block (i.e., universal blocking) nucleic acids. Compositions comprising the novel U-block nucleic acids provided herein and methods utilizing the U-block nucleic acids provided herein are less costly, increase efficiency and workflow, and are more amenable to automation than traditional methods.
In addition, traditional applications of hybrid capture methods typically involve combining a pool of adaptor-ligated library inserts or amplicons thereof with C0t-1DNA and blocking oligonucleotides, followed by a drying step. The drying step is typically performed in vacuum, which is time consuming and is performed in an open system, which provides a high risk of cross-contamination between samples. After drying, the samples were denatured and then annealed for several days. Biotinylated capture oligonucleotides (e.g., "baits") are then added and the hybridized nucleic acids are typically pulled down with avidin-coated beads. The pool of retained nucleic acids is then eluted from the beads and introduced into an automated sequencing process. The above steps are inefficient and time consuming, are not amenable to automation and can lead to cross contamination.
In some aspects, provided herein are improved methods for manipulating and preparing nucleic acid libraries for analysis (e.g., for high-throughput sequencing) that do not require extended incubation times and/or drying steps.
Test subject
The subject may be any living or non-living organism, including but not limited to a human, non-human animal, plant, bacteria, fungus, virus, or protist. The subject may be of any age (e.g., embryo, fetus, infant, child, adult). The subject may be of any gender (e.g., male, female, or a combination thereof). The subject may be pregnant. The subject may be a patient (e.g., a human patient).
Sample (I)
Provided herein are methods and compositions for analyzing a sample. A sample (e.g., a sample comprising nucleic acids) can be obtained from a suitable subject. The sample may be isolated or obtained directly from the subject or portion thereof. In some embodiments, the sample is obtained indirectly from the individual or medical professional. The sample may be any specimen isolated or obtained from a subject or portion thereof. The sample may be any specimen isolated or obtained from a plurality of subjects. Non-limiting examples of specimens include fluids or tissues from a subject, including but not limited to blood or blood products (e.g., serum, plasma, platelets, buffy coat, etc.), cord blood, chorionic villi, amniotic fluid, cerebrospinal fluid, spinal fluid, lavage fluid (e.g., of the lungs, stomach, peritoneum, ducts, ears, arthroscope), biopsy samples, celocentesis (celocentesis) samples, cells (blood cells, lymphocytes, placental cells, stem cells, bone marrow-derived cells, embryos, or fetal cells) or portions thereof (e.g., mitochondria, nuclei, extracts, etc.), urine, stool, sputum, saliva, nasal mucus, prostatic fluid, lavage fluid, semen, lymph, bile, tears, sweat, breast milk, breast fluid, etc., or combinations thereof. The fluid or tissue sample from which the nucleic acids are extracted may be acellular (e.g., cell-free). Non-limiting examples of tissues include organ tissue (e.g., liver, kidney, lung, thymus, adrenal gland, skin, bladder, reproductive organs, intestine, colon, spleen, brain, or portions thereof), epithelial tissue, hair follicles, catheters, blood vessels, bone, eye, nose, mouth, throat, ear, nail, or the like, portions thereof, or combinations thereof. The sample can include normal, healthy, diseased (e.g., infected), and/or cancerous (e.g., cancerous cells) cells or tissues. The sample obtained from the subject can comprise cells or cellular material (e.g., nucleic acids) of a variety of organisms (e.g., viral nucleic acids, fetal nucleic acids, bacterial nucleic acids, parasitic nucleic acids).
In some embodiments, the sample comprises nucleic acids or fragments thereof. The sample may comprise nucleic acids obtained from one or more subjects. In some embodiments, the sample comprises nucleic acids obtained from a single subject. In one embodiment, the sample comprises a mixture of nucleic acids. The mixture of nucleic acids can comprise two or more nucleic acid species having different nucleotide sequences, different fragment lengths, different sources (e.g., genomic origin, cell or tissue origin, subject origin, etc., or a combination thereof), or a combination thereof. The sample may comprise synthetic nucleic acids.
Nucleic acids
The term "nucleic acid" refers to one or more nucleic acids (e.g., a group or subgroup of nucleic acids) from any composition of, for example, DNA (e.g., complementary DNA (cdna), genomic DNA (gdna), etc.), RNA (e.g., messenger RNA (mrna), short inhibitory RNA (sirna), ribosomal RNA (rrna), tRNA, microRNA), and/or DNA or RNA analogs (e.g., containing base analogs, sugar analogs, and/or non-natural backbones, etc.), RNA/DNA hybrids, and Polyamide Nucleic Acids (PNA), all of which can be in single-stranded or double-stranded form, and unless otherwise specified, can include known analogs of natural nucleotides, which can function in a manner similar to naturally occurring nucleotides. Unless specifically limited, the term includes nucleic acids comprising deoxyribonucleotides, ribonucleotides, and known analogs of natural nucleotides. Nucleic acids may include as equivalents, derivatives, or variants thereof, suitable analogs of RNA or DNA synthesized from nucleotide analogs, single-stranded ("sense" or "antisense", "positive" or "negative" strand, "forward" or "reverse" reading frame), and double-stranded polynucleotides. The nucleic acid may be single-stranded or double-stranded. The nucleic acid can have any length of 2 or more, 3 or more, 4 or more, or 5 or more contiguous nucleotides. Nucleic acids may comprise a specific 5 'to 3' order of nucleotides known in the art as sequences (e.g., nucleic acid sequences).
The nucleic acid may be naturally occurring and/or may be synthesized, replicated, or altered by the human hand. For example, the nucleic acid may be an amplicon. The nucleic acid may be from a nucleic acid library, such as a gDNA, cDNA or RNA library. The nucleic acid may be synthetic (e.g. chemical synthesis) or produced (e.g. by in vitro polymerase extension, e.g. by amplification, e.g. by PCR). The nucleic acid may be or may be derived from a plasmid, phage, virus, Autonomously Replicating Sequence (ARS), centromere, artificial chromosome, or other nucleic acid that, in some embodiments, is capable of replicating or being replicated in vitro or in a host cell, nucleus or cytoplasm of a cell. Nucleic acids (e.g., libraries of nucleic acids) can comprise nucleic acids from one sample or from two or more samples (e.g., from 1 or more, 2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 11 or more, 12 or more, 13 or more, 14 or more, 15 or more, 16 or more, 17 or more, 18 or more, 19 or more, or 20 or more samples). The nucleic acids provided for the processes or methods described herein can comprise nucleic acids from 1 to 1000, 1 to 500, 1 to 200, 1 to 100, 1 to 50, 1 to 20, or 1 to 10 samples.
The term "gene" refers to segments of DNA involved in the production of polypeptide chains and may include regions preceding and following the coding region (leader and trailer) involved in transcription/translation and regulation of transcription/translation of the gene product, as well as intervening sequences (introns) between the individual coding segments (exons). Due to genetic variations in the gene sequence (e.g., mutations in the coding and non-coding portions of the gene), the gene may not necessarily produce peptides or may produce truncated or non-functional proteins. Genes, whether functional or non-functional, can generally be identified by homology to genes in a reference genome.
Oligonucleotides are relatively short nucleic acids. The oligonucleotide may be about 2 to 150, 2 to 100, 2 to 50, or 2 to about 35 nucleic acids in length. In some embodiments, the oligonucleotide is single stranded. In certain embodiments, the oligonucleotide is a primer. The primer is generally configured to hybridize to a selected complementary nucleic acid and is configured to be extended by a polymerase upon hybridization.
Nucleic acid isolation and purification
Nucleic acids can be derived, isolated, extracted, purified, or partially purified from one or more subjects, one or more samples, or one or more sources using suitable methods known in the art. Nucleic acids can be isolated, extracted and/or purified using any suitable method.
The term "isolated" as used herein refers to a nucleic acid that is removed from its original environment (e.g., the natural environment if it is naturally occurring; or the host cell if it is exogenously expressed) and thus altered from its original environment by human intervention (e.g., "by hand"). The term "isolated nucleic acid" as used herein may refer to nucleic acid removed from a subject (e.g., a human subject). An isolated nucleic acid may have fewer non-nucleic acid components (e.g., proteins, lipids) than the amount of components present in the source sample. Compositions comprising isolated nucleic acids may be about 50% to greater than 99% free of non-nucleic acid components. A composition comprising an isolated nucleic acid may be free of about 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or greater than 99% of non-nucleic acid components. The term "purified" as used herein can refer to a nucleic acid, provided that it contains less non-nucleic acid components (e.g., proteins, lipids, carbohydrates, salts, buffers, detergents, etc., or combinations thereof) than are present prior to subjecting the nucleic acid to a purification procedure. A composition comprising purified nucleic acid can be free of at least about 60%, 70%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or greater than 99% of other non-nucleic acid components. A composition comprising purified nucleic acid may comprise at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or greater than 99% of the total nucleic acid present in a sample prior to administration of a purification method.
In some embodiments, a purification mixture (e.g., a purification mixture of nucleic acids) provides purified nucleic acids. In certain embodiments, a mixture comprising library nucleic acids, blocker nucleic acids, capture nucleic acids, competitor nucleic acids, and/or combinations thereof is purified to provide purified nucleic acids. Nucleic acid purification sometimes comprises DNA purification columns or DNA purification beads. Various nucleic acid purification columns, resins, matrices and kits are known in the art. Any suitable nucleic acid purification method, resin, bead, matrix or reagent can be used with the methods herein. For example, a nucleic acid purification method can include binding (e.g., non-covalent binding) nucleic acids to a suitable cation exchange resin (e.g., cationic beads) comprising a metal, pulling down the bound nucleic acid complexes by use of a magnet, and then eluting the bound nucleic acids by addition of a low salt buffer. For example, IN certain embodiments, nucleic acid purification includes the use of AMPureXP magnetic beads (Beckman Coulter, inc., Indianapolis IN, (USA)), and the like. The nucleic acid purification methods used herein are generally modified to optimally recover short nucleic acids (e.g., blocking nucleic acids) as well as library inserts (e.g., adaptor-ligated inserts and amplicons thereof). In certain embodiments, the nucleic acid purification methods herein are altered to optimally recover nucleic acids having an average or absolute length of from about 5 to about 1000 nucleotides, from 5 to about 800 nucleotides, or from 5 to about 500 nucleotides. In certain embodiments, the nucleic acid purification methods used herein use a ratio of nucleic acid binding resin (e.g., nucleic acid binding beads, nucleic acid binding matrix, e.g., 100% slurry) to nucleic acid containing mixture of 1.8:1, 1.9:1, 2:1, 2.1:1, 2.2:1, 2.3:1, 2.4:1, 2.5:1, 2.6:1, 2.7:1, 2.8:1, 2.9:1, or 3:1 (vol: vol).
In some embodiments, the purification process comprises a washing step. In some embodiments, the purification process comprises an elution step.
In some embodiments, the purification process described herein does not include the use of a drying step, vacuum (e.g., flash vacuum), and/or lyophilization. Such a process results in a high risk of cross-contamination. In some embodiments, although the capture nucleic acids of the mixture typically comprise a member of a binding pair, the purification processes described herein do not comprise the use of a second member of a binding pair. In some embodiments, the purification method is performed in the absence of a hybridization buffer. In some embodiments, the purification process is performed in the absence of added calcium or magnesium salts. In some embodiments, the purification method is performed in the absence of a detergent (e.g., SDS), Ficoll, BAS, and/or polyvinylpyrrolidone.
Hybridization of
Substantially complementary single-stranded nucleic acids can hybridize to each other under hybridization conditions, thereby forming a partially or fully double-stranded nucleic acid. In some embodiments, all or a portion of a nucleic acid sequence may be substantially complementary to another nucleic acid sequence. As used herein, "substantially complementary" refers to nucleotide sequences that can hybridize to each other under suitable hybridization conditions. Hybridization conditions can be varied to tolerate different amounts of sequence mismatches within substantially complementary nucleic acids. Substantially complementary portions of nucleic acids that can hybridize to each other can be 75% or more, 76% or more, 77% or more, 78% or more, 79% or more, 80% or more, 81% or more, 82% or more, 83% or more, 84% or more, 85% or more, 86% or more, 87% or more, 88% or more, 89% or more, 90% or more, 91% or more, 92% or more, 93% or more, 94% or more, 95% or more, 96% or more, 97% or more, 98% or more, or 99% or more complementary to each other. In some embodiments, substantially complementary portions of nucleic acids that can hybridize to each other are 100% complementary. Nucleic acids, or portions thereof, that are configured to hybridize to each other typically comprise nucleic acid sequences that are substantially complementary to each other.
As used herein, "specifically hybridizes" refers to preferential hybridization under hybridization conditions, wherein two nucleic acids that are substantially complementary, or portions thereof, hybridize to each other, but not to other nucleic acids that are not substantially complementary to either of the two nucleic acids. For example, specific hybridization includes hybridization of a capture nucleic acid to a portion of a target amplicon that is substantially complementary to the capture nucleic acid. In some embodiments, nucleic acids or portions thereof configured to specifically hybridize are generally about 80% or more, 81% or more, 82% or more, 83% or more, 84% or more, 85% or more, 86% or more, 87% or more, 88% or more, 89% or more, 90% or more, 91% or more, 92% or more, 93% or more, 94% or more, 95% or more, 96% or more, 97% or more, 98% or more, 99% or more, or 100% complementary to each other over contiguous portions of the nucleic acid sequence. Specific hybridization differs by about 2-fold or more, typically about 10-fold or more, sometimes about 100-fold or more, 1000-fold or more, 10,000-fold or more, 100,000-fold or more, or 1,000,000-fold or more, from non-specific hybridization interactions (e.g., two nucleic acids that are not configured for specific hybridization, e.g., 80% or less, 70% or less, 60% or less, or 50% or less complementarity of the two nucleic acids).
In some embodiments, the methods described herein comprise hybridizing nucleic acids under hybridization conditions. Conditions that favor hybridization of substantially complementary nucleic acids are referred to herein as hybridization conditions. Methods for altering the stringency of hybridization conditions are well known in the art.
Hybridization conditions can be determined and/or adjusted depending on the nature of the nucleic acid used in the assay. Methods for optimizing hybridization conditions are well known in the art and can be found in Current Protocols in Molecular Biology, John Wiley & Sons, N.Y.,6.3.1-6.3.6 (1989). Nucleic acid sequence content (e.g., GC content, degree of mismatch) and/or length can sometimes affect hybridization of substantially complementary nucleic acids. Hybridization conditions generally include parameters that can be adjusted to optimally anneal two or more substantially complementary nucleic acids of interest. Non-limiting examples of such adjustable parameters include temperature, monovalent or divalent ion and/or cation concentration (e.g., Mg concentration), buffer concentration, phosphate concentration, glycerol concentration, DMSO concentration, nucleic acid concentration, and the like, or combinations thereof. Depending on the degree of mismatch between the substantially complementary nucleic acids, hybridization conditions can be adjusted to affect annealing and/or to select specific hybridization of selected nucleic acids (e.g., oligonucleotides or primers with known or predicted melting temperatures).
Hybridization conditions generally include heating or cooling a sample comprising nucleic acids to a suitable temperature. Suitable temperatures for hybridization are sometimes about 0 ℃ to 80 ℃, about 25 ℃ to 70 ℃, about 30 ℃ to 70 ℃, about 35 ℃ to 70 ℃, about 40 ℃ to 70 ℃, about 35 ℃ to 65 ℃, about 35 ℃ to 60 ℃, about 35 ℃ to 55 ℃, or about 40 ℃ to 50 ℃. In some embodiments, hybridization conditions include cooling the sample to a temperature of about 40 ℃, 41 ℃, 42 ℃, 43 ℃, 44 ℃, 45 ℃, 46 ℃, 47 ℃, 48 ℃, 49 ℃, 50 ℃, 51 ℃, 52 ℃, 53 ℃, 54 ℃, or about 55 ℃. In some embodiments, hybridizing a purified nucleic acid mixture under hybridization conditions comprises a denaturation step followed by a hybridization step (e.g., incubation for a period of time at a temperature suitable for hybridization). Hybridization conditions sometimes include denaturing a mixture of nucleic acids and then immediately cooling (e.g., rapidly lowering the temperature) the mixture to a suitable hybridization temperature.
In certain embodiments, the hybridization conditions comprise a denaturation process. Denaturation typically involves raising the temperature of a sample comprising nucleic acids (e.g., by heating) to a temperature at or above the melting point of one or more double-stranded nucleic acids in the sample. In some embodiments, denaturing comprises raising the temperature of the sample from about 70 ℃ to about 120 ℃, about 85 ℃ to about 105 ℃, about 90 ℃ to about 105 ℃, or about 95 ℃ to about 105 ℃. In some embodiments, denaturing comprises raising the temperature of the sample to about 70 ℃ or more, about 75 ℃ or more, about 80 ℃ or more, about 85 ℃ or more, or to about 90 ℃, about 91 ℃, about 92 ℃, about 93 ℃, about 94 ℃, about 95 ℃, about 96 ℃, about 97 ℃, about 98 ℃, about 99 ℃, about 100 ℃, about 101 ℃, about 102 ℃, about 103 ℃, about 104 ℃, or to about 105 ℃. In some embodiments, the nucleic acid is denatured at the desired denaturation temperature for about 1 second to about 30 minutes or longer, about 15 seconds to about 30 minutes, about 30 seconds to about 30 minutes, about 1 minute to about 20 minutes, about 1 minute to about 15 minutes, or about 5 minutes to about 10 minutes.
Hybridization conditions typically comprise an incubation period in which the sample is held at a desired hybridization temperature for a period of time. Traditional methods of hybridizing library nucleic acids, blocking and/or capturing nucleic acids require hybridization times in excess of 48 hours. While any suitable conditions may be used for hybridization, certain methods provided herein provide for significantly reduced hybridization times. In certain embodiments, the hybridization conditions comprise incubating the mixture at the desired hybridization temperature for about 1 minute to about 24 hours, about 5 minutes to about 24 hours, about 10 minutes to about 24 hours, about 15 minutes to about 24 hours, about 30 minutes to about 24 hours, about 1 hour to about 24 hours, about 2 hours to about 24 hours, about 8 hours to about 24 hours, about 10 hours to about 20 hours, or about 12 hours to about 20 hours. In some embodiments, the hybridization conditions comprise incubating the mixture at the desired hybridization temperature for about 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 hours. In some embodiments, the hybridization conditions comprise incubating the sample at the desired hybridization temperature for no longer than about 48 hours, no longer than about 24 hours, or no longer than about 18 hours.
In some embodiments, hybridizing comprises contacting the nucleic acid with a hybridization buffer. In some embodiments, the hybridization conditions comprise a mixture of nucleic acids with a suitable hybridization buffer. Hybridization buffers are known in the art and are commercially available. The hybridization buffer may comprise detergents (e.g., SDS), Ficoll, glycerol, BAS, polyvinylpyrrolidone, dextran glycerol, divalent cations (e.g., calcium and/or magnesium), monovalent cations, phosphates, and the like, or combinations thereof.
In some embodiments, the methods described herein do not include a denaturation step prior to the step of hybridizing the purified nucleic acid under hybridization conditions.
In certain embodiments, the hybridization conditions do not include a polymerase. In some embodiments, the hybridization conditions do not include polymerase chain reaction. In certain embodiments, after capturing the nucleic acids of the mixture, the mixture does not comprise a polymerase.
Amplification of
Nucleic acids can be amplified by suitable methods. The term "amplifying" as used herein refers to a method of subjecting a target nucleic acid in a sample to linear or exponential generation of an amplicon nucleic acid having the same or substantially the same (e.g., substantially the same) nucleotide sequence as the target nucleic acid or fragment thereof. In some embodiments, the amplification reaction comprises a suitable thermostable polymerase. Thermostable polymers are known in the art and are stable for extended periods of time at temperatures above 80 ℃ when compared to common polymerases found in most mammals. In certain embodiments, the term "amplifying" refers to a method comprising Polymerase Chain Reaction (PCR). Conditions which facilitate amplification (i.e., amplification conditions) are well known and typically include at least a suitable polymerase, a suitable template, a suitable primer or primer set, suitable nucleotides (e.g., dntps), suitable buffers, and the application of suitable annealing, hybridization, and/or extension times and temperatures. In certain embodiments, the amplification product (e.g., amplicon) may contain one or more additional and/or different nucleotides than the template sequence or portion thereof from which the amplicon was generated (e.g., the primer may contain "additional" nucleotides).
Nucleic acids can be amplified by thermal cycling methods or by isothermal amplification methods. In some embodiments, a rolling circle amplification method is used. In some embodiments, amplification occurs on a solid support (e.g., within a flow cell) in which nucleic acids, nucleic acid libraries, or portions thereof are immobilized. In certain sequencing methods, a library of nucleic acids is added to a flow cell and immobilized by hybridization to an anchor under suitable conditions. This type of nucleic acid amplification is commonly referred to as solid phase amplification. In some embodiments of solid phase amplification, all or part of the amplification product is synthesized by extension from the immobilized primer. The solid phase amplification reaction is similar to standard solution phase amplification except that at least one of the amplification oligonucleotides (e.g., primers) is immobilized on a solid support.
In some embodiments, solid phase amplification comprises a nucleic acid amplification reaction comprising only one oligonucleotide primer immobilized to a surface. In certain embodiments, the solid phase amplification comprises a plurality of different immobilized oligonucleotide primer species. In some embodiments, solid phase amplification may comprise a nucleic acid amplification reaction comprising one oligonucleotide primer immobilized on a solid surface and a second, different oligonucleotide primer in solution. A variety of different types of immobilized or solution-based primers can be used. Non-limiting examples of solid phase nucleic acid amplification reactions include interfacial amplification, bridge amplification, emulsion PCR, WildFire amplification (e.g., U.S. patent publication US20130012399), and the like, or combinations thereof.
Nucleic acid libraries
In certain embodiments, a nucleic acid library (e.g., a library of nucleic acids) is a collection or subset of total grnas, RNAs, or crnas obtained from one or more subjects. The nucleic acid library may comprise single-stranded and/or double-stranded nucleic acids. Nucleic acid libraries are typically generated from one or more samples and comprise endogenous or native nucleic acids of one or more subjects or organisms from which the samples were obtained. A nucleic acid library typically comprises a plurality of nucleic acids or nucleic acid fragments endogenous or native to one or more organisms from which a sample is obtained. Such endogenous or native nucleic acids are sometimes referred to herein as library inserts. Thus, a plurality of nucleic acids may refer to 103To 1020Nucleic acid(s). In some embodiments, a plurality of nucleic acid fingers 103Or more, 104Or more, 105Or more, 106Or more, 107Or more, 108Or more, 109Or more, 1010Or more, or 1012Or more nucleic acids (or inserts). The library insert may be a fragment of genomic DNA (e.g., a genomic DNA library), RNA (e.g., an RNA library), or cDNA (e.g., a cDNA library). Library inserts may comprise full-length genes, cDNAs, introns, exons, untranslated regions (e.g., promoters, enhancers, regulatory sequences, etc.) genesA portion thereof, or a combination thereof. Libraries of nucleic acids obtained from any source or subject typically contain 1000 or more, 10,000 or more, or 100,000 or more library inserts, which are different and often distinguishable from each other. In some embodiments, a library of nucleic acids comprises a plurality of library inserts in which the nucleic acids are prepared, assembled, and/or altered for a particular process, non-limiting examples of which include immobilization on a solid phase (e.g., a solid support such as a flow cell, bead), enrichment, amplification, cloning, detection, and/or for nucleic acid sequencing. In some embodiments, the library of nucleic acids comprises one or more library inserts obtained from one or more samples (e.g., one or more subjects, one or more tissues, one or more species, or a combination thereof). In some embodiments, each nucleic acid of the library comprises at least one library insert (e.g., one, two, three, or more inserts), one or more non-native nucleic acids comprising one or more nucleic acid barcodes. In certain embodiments, the nucleic acid library is prepared prior to or during a sequencing process. Nucleic acid libraries (e.g., sequencing libraries) can be prepared by suitable methods known in the art. Nucleic acid libraries can be prepared by targeted or non-targeted preparation processes.
In some embodiments, the library of nucleic acids is altered to comprise one or more non-natural nucleic acids, typically of known composition (e.g., synthetic nucleic acids, heterologous nucleic acids). In some embodiments, each nucleic acid of the library comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 or more non-natural nucleic acids. In some embodiments, each nucleic acid of the library comprises one or two non-native nucleic acids. The non-natural nucleic acid may be added at a suitable location, for example, on one end, the other end, or both ends of the library insert. In some embodiments, the nucleic acids of the library comprise non-natural nucleic acids on opposite ends of the library insert. For example, the nucleic acids of the library may comprise non-natural nucleic acids located at the 5 'end and/or the 3' end of the library insert. In some embodiments, the non-natural nucleic acid is covalently bound to the library insert, for example, by a suitable phosphodiester bond. In some embodiments, the nucleic acids of the library comprise a library insert, a first non-natural nucleic acid, and a second non-natural nucleic acid, wherein the first and second non-natural nucleic acids are located on opposite sides (e.g., 5 'and 3' sides) of the library insert.
In certain embodiments, the non-natural nucleic acid is not native and/or endogenous to the one or more subjects or organisms from which the nucleic acid library was obtained. Non-natural nucleic acids that are not native and/or endogenous to the subject or organism from which the nucleic acid library insert is obtained typically do not comprise genomic DNA or RNA (e.g., cDNA) derived from the subject or organism.
In certain embodiments, the non-natural nucleic acid comprises a suitable exogenous nucleic acid and/or a suitable synthetic nucleic acid. Non-limiting examples of synthetic nucleic acids include a distinguishable identifier, a nucleic acid barcode (e.g., a distinguishable nucleic acid barcode), a capture nucleic acid, a sequence tag, a random nucleic acid sequence, an adaptor nucleic acid, a restriction enzyme site, an overhang, a promoter, an enhancer, an origin of replication, a stem loop, a primer binding site, an oligonucleotide annealing site, a suitable integration site (e.g., a transposon, a viral integration site), one or more modified nucleotides, and the like, portions thereof, or combinations thereof. The non-natural nucleic acids may have the same or different nucleic acid sequences. In some embodiments, the non-natural nucleic acid is configured to hybridize to one or more capture nucleic acids, blocking nucleic acids (e.g., U-block nucleic acids), or primers.
In certain embodiments, the nucleic acid library is prepared by a ligation-based library preparation method. In some embodiments, the non-native nucleic acid is added to the nucleic acids of the library by a ligation-based library preparation method. Ligation-based library preparation methods typically use one or more adaptors. Adapters are typically synthetic nucleic acids (e.g., made by human hand) comprising a nucleic acid sequence that is not endogenous to or present in the organism from which the library was derived. In some embodiments, the non-native nucleic acid comprises an adaptor. In certain embodiments, the adapters can be used to prepare library inserts for analysis (e.g., single read sequencing, paired-end sequencing, and multiplex sequencing). In some embodiments, nucleic acid library preparation comprises ligating one or more adaptors to a plurality of library inserts. The adapter can be a relatively short double-stranded or single-stranded oligonucleotide (e.g., about 2 to about 10, 2 to about 30, 2 to about 50, or 2 to about 100 nucleic acids or more) that can include, for example, a distinguishable identifier, a distinguishable nucleic acid barcode, and/or one or more members of a binding pair. Adapters are typically located 5 'and/or 3' to the library insert. Adapters are typically located on the sides of the library insert. In some embodiments, only one strand of a double-stranded adaptor is introduced into the nucleic acids of the library (e.g., ligated to the library insert). Sometimes both strands of a double stranded adaptor are introduced into the library. In some embodiments, the single stranded nucleic acids of the library comprise a5 'adaptor and a 3' adaptor, wherein the sequences of the 5 'and 3' adaptors are substantially different and/or not substantially complementary.
In certain embodiments, the adapter comprises a portion that is substantially complementary to the flow cell anchor, which is sometimes used to immobilize the nucleic acid library to a solid support, such as the inner surface of a flow cell. In certain embodiments, the adapter comprises a portion that is substantially complementary to the amplification primer or the sequencing primer, which may be the same, different, or overlapping portions of the adapter. In some embodiments, at least one adaptor (e.g., a5 'or 3' adaptor) of the nucleic acids of the library comprises a distinguishable identifier (e.g., a distinguishable barcode sequence). In some embodiments, both strands of the double-stranded adaptor comprise a nucleic acid barcode, wherein the nucleic acid barcode of the first strand of the adaptor is substantially complementary to the nucleic acid barcode of the second strand of the adaptor.
In certain embodiments, two or more non-native nucleic acids comprise substantially the same portion. Portions of substantially identical non-natural nucleic acids are sometimes referred to as universal nucleic acids (e.g., universal nucleic acid sequences). In some embodiments, the non-natural nucleic acid comprises one or more universal nucleic acids. In certain embodiments, the non-natural nucleic acid comprises a first universal nucleic acid, a second universal nucleic acid, and a nucleic acid barcode, wherein the barcode is located between the first and second universal nucleic acids. Such non-native nucleic acids may be located 5 'and/or 3' to the library insert. For example, in some embodiments, the nucleic acids of the library can comprise a first universal nucleic acid, a first nucleic acid barcode, and a second universal nucleic acid located 5 'to the library insert, and a third universal nucleic acid, a second barcode, and a fourth universal nucleic acid located 3' to the library insert. In some embodiments, the nucleic acids of the library can comprise a first non-natural nucleic acid and a first nucleic acid barcode located 5 'to the library insert and a second non-natural nucleic acid and a second barcode located 3' to the library insert. In certain embodiments, the non-native nucleic acid is designed and/or configured to comprise a universal sequence flanking the distinguishable barcode. In some embodiments, the U-block nucleic acids herein are designed and/or configured to specifically hybridize to a universal nucleic acid sequence.
In some embodiments, the nucleic acid library or portion thereof is amplified (e.g., by PCR-based methods). In some embodiments, the sequencing method comprises amplification of a nucleic acid library. The library of nucleic acids may be amplified before or after immobilization on a solid support (e.g., a solid support in a flow cell). In some embodiments, the nucleic acid library comprises amplicons. The amplicons of the nucleic acid library may be single stranded or double stranded. In some embodiments, the amplicons of a nucleic acid library comprise library inserts and adapter sequences (e.g., library inserts flanked by adapter sequences). In some embodiments, the nucleic acid library comprises a plurality of amplicons, sometimes referred to herein as a library of amplicons.
In certain embodiments, each amplicon of the amplicon library comprises a library insert and one or more non-native nucleic acids. For example, in some embodiments, the amplicon comprises the library insert and 1, 2, 3, 4, 10 or more, or 50 or more non-natural nucleic acids. Amplicons typically comprise library inserts positioned between one or more 5 'non-natural nucleic acids and one or more 3' non-natural nucleic acids. In certain embodiments, the amplicon comprises 1, 2, 3, 4, or 5 non-natural nucleic acids located 5 'to the library insert and 1, 2, 3, 4, or 5 non-natural nucleic acids located 3' to the library insert. In some embodiments, the amplicons comprise one or more distinguishable barcodes located 5 'to the library insert and one or more distinguishable barcodes located 3' to the library insert. In certain embodiments, the amplicons comprise 1, 2, 3, 4, or 5 distinguishable barcodes located 5 'to the library insert and 1, 2, 3, 4, or 5 distinguishable barcodes located 3' to the library insert.
In some embodiments, the amplicons of the library comprise 1 or more substantially identical non-natural nucleic acids. Substantially identical nucleic acids are at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 100% identical in nucleic acid sequence. Similarly, two or more nucleic acids that are substantially identical refers to two or more nucleic acids that comprise substantially identical nucleic acid sequences.
Substantially different nucleic acids refer to two or more nucleic acid sequences that are less than 50%, less than 60%, less than 70%, less than 80%, or less than 85% identical in nucleic acid sequence. Two or more nucleic acids that are substantially different refers to two or more nucleic acids that comprise substantially different nucleic acid sequences.
In some embodiments, the methods herein comprise obtaining a library of nucleic acids comprising one or more nucleic acid inserts obtained from one or more samples. In some embodiments, the nucleic acid library can be obtained by generating the library as described herein or by methods known in the art. In certain embodiments, the nucleic acid library is obtained from a third party, wherein the third party generates the nucleic acid library. In some embodiments, the nucleic acid library is purchased from a supplier. In some embodiments, the methods herein comprise obtaining a library of nucleic acids comprising one or more library inserts obtained from one or more samples, wherein each nucleic acid of the library comprises at least one nucleic acid insert, a first non-natural nucleic acid, a second non-natural nucleic acid, and one or more nucleic acid barcodes, wherein the first non-natural nucleic acid is located 5 'to the at least one library insert and the second non-natural nucleic acid is located 3' to the library insert.
Distinguishable identifier
In some embodiments, the nucleic acid comprises one or more distinguishable identifiers. A distinguishable identifier can be introduced or linked (e.g., covalently, noncovalently, irreversibly, or reversibly linked) to a nucleic acid (e.g., a polynucleotide) that allows for detection and/or identification of the nucleic acid comprising the identifier. In some embodiments, the distinguishable identifier is introduced or linked to the nucleic acid (e.g., by a polymerase) prior to or during the sequencing method. Any suitable distinguishable identifier and/or detectable identifier may be used in the compositions or methods described herein. In certain embodiments, the distinguishable identifier can be directly or indirectly linked to (e.g., bound to) a nucleic acid-related (e.g., bound to). For example, the distinguishable identifier may be covalently or non-covalently bound to the nucleic acid. In some embodiments, the distinguishable identifier is bound or associated with a binding agent or member of a binding pair that binds covalently or non-covalently to the nucleic acid. In some embodiments, the distinguishable identifier is reversibly associated with the nucleic acid. In certain embodiments, the distinguishable identifier reversibly associated with the nucleic acid can be removed from the nucleic acid using a suitable method (e.g., by increasing salt concentration, denaturing, washing, adding a suitable solvent, and/or by heating).
In some embodiments, 1 or more, 2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 20 or more, 30 or more, or 50 or more distinguishable identifiers are used in the methods described herein (e.g., nucleic acid detection, analysis, and/or sequencing methods).
In some embodiments, the distinguishable identifier is a tag. In some embodiments, the nucleic acid comprises a detectable label, non-limiting examples of which include radioactive labels(e.g., isotopes), metal tags, fluorescent tags, chromophores, chemiluminescent tags, electro-chemiluminescent tags (e.g., origin)TM) A phosphorescent label, a quencher (e.g., a fluorophore quencher), a Fluorescence Resonance Energy Transfer (FRET) pair (e.g., a donor and an acceptor), a dye, a protein (e.g., an enzyme (e.g., alkaline phosphatase and horseradish peroxidase), an antibody, an antigen or portion thereof, a linker, a member of a binding pair), an enzyme substrate, a small molecule (e.g., biotin, avidin), a mass label, a quantum dot, a nanoparticle, and the like, or combinations thereof. Suitable fluorophores can be used as labels. The luminescent tags can be detected and/or quantified by a variety of suitable techniques, such as flow cytometry, quantitative polymerase chain reaction (qPCR), gel electrophoresis, gene chip analysis, microarray, mass spectrometry, cellular fluorescence analysis, fluorescence microscopy, confocal laser scanning microscopy, laser scanning cell counting, affinity chromatography, manual batch mode separation, electric field suspension, sequencing, and the like, and combinations thereof.
In some embodiments, the distinguishable identifier is a nucleic acid barcode.
Nucleic acid barcodes
In some embodiments, the non-natural nucleic acid comprises one or more distinguishable nucleic acid barcodes (e.g., index nucleotides, sequence tags, or "barcode" nucleotides). Nucleic acid barcodes are typically nucleic acids having a particular sequence that are introduced or appended to (e.g., correlated with) a particular nucleic acid or subset of nucleic acids of a sample to track and/or identify a particular nucleic acid or subset of nucleic acids in a mixture of nucleic acids. In certain embodiments, a distinguishable nucleic acid barcode comprises a distinguishable sequence of nucleotides that can be used as an identifier to allow for unambiguous identification of one or more nucleic acids (e.g., a subset of nucleic acids) in a sample, method, or assay. Distinguishable nucleic acid barcodes are typically configured to allow unambiguous identification of the source or identity of the nucleic acid with which the barcode is associated. In some embodiments, distinguishable nucleic acid barcodes (e.g., barcodes) may allow for the identification of the source of a particular nucleic acid in a mixture of nucleic acids obtained from different sources. In some embodiments, the distinguishable nucleic acid barcodes are configured to allow unambiguous identification of the source or identity of the nucleic acid associated with the barcode. For example, in certain embodiments, the distinguishable nucleic acid barcodes are specific and/or unique to a sample, sample source, library of nucleic acids obtained from the same subject or tissue, a particular genus or subset of nucleic acids, a particular species of nucleic acid, nucleic acids from the same chromosome, or the like, or a combination thereof. In some embodiments, a nucleic acid comprising an insert derived from a sample, subject, or tissue comprises a nucleic acid barcode specific and unique to the sample, subject, or tissue, thereby allowing unambiguous identification of nucleic acids derived from different samples, subjects, or tissues and/or inserts from nucleic acids. Thus, a distinguishable nucleic acid barcode that is unique to a sample, subject, or tissue is typically distinguishable and distinct from other nucleic acid barcodes in a nucleic acid mixture. In some embodiments, a distinguishable nucleic acid barcode having uniqueness is different and/or distinguishable from other barcodes in a composition comprising one or more samples derived from one or more sources (e.g., a library of nucleic acids derived from different samples or sources). In some embodiments, a distinguishable nucleic acid barcode that is unique to a sample, subject, or tissue is associated with (or is included in) a nucleic acid derived from the same sample, subject, tissue, or a specific subset thereof. Thus, in some embodiments, nucleic acids derived from the same sample, subject, or tissue typically comprise at least one distinguishable nucleic acid barcode having the same sequence, which is associated with each nucleic acid of the same sample, subject, or tissue.
In some embodiments, the distinguishable barcode comprises a distinguishable and/or unique sequence having 4 to 10, 4 to 15, 4 to 20, 4-50, or 20 or more contiguous nucleotides. The two distinguishable nucleic acid barcodes may differ in sequence by 1, 2, 3, 4, 5 or more nucleotides. Thus, in certain embodiments, two nucleic acid barcodes that are different and/or distinguishable may be up to 99% identical and comprise nucleic acid sequences that differ by at least one nucleotide. In some embodiments, the distinguishable barcode comprises a distinguishable and/or unique sequence having 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 or more contiguous nucleotides. In some embodiments, the distinguishable barcode comprises a distinguishable and/or unique sequence of no more than 10, no more than 15, or no more than 20 contiguous nucleotides. Distinguishable nucleic acid barcodes typically comprise a first end and a second end. The end of the distinguishable nucleic acid barcode sequence can be identified as the most 5 'or the most 3' nucleotide base of the distinguishable nucleic acid barcode sequence. For example, in some embodiments, the first end can be identified as the most 5 'nucleotide base of the distinguishable nucleic acid barcode sequence, and the second end can be identified as the most 3' nucleotide base of the distinguishable nucleic acid barcode sequence. The first end and the second end are typically located at opposite ends of the distinguishable nucleic acid barcode. In some embodiments, any two or more distinguishable nucleic acid barcodes in a library can be distinguished and/or identified by a nucleic acid sequencing method. In some embodiments, two or more distinguishable nucleic acid barcodes may be distinguished and/or identified by hybridization methods.
In some embodiments, a library of nucleic acids obtained from a plurality of sources or samples comprises a plurality of distinguishable nucleic acid barcodes. In some embodiments, each distinguishable nucleic acid barcode of a library can be used to identify the source of each nucleic acid of a mixed library. For example, a library of nucleic acids obtained from a plurality of sources may comprise a first library of nucleic acids obtained from a first subject comprising first and/or second distinguishable nucleic acid barcodes, a second library of nucleic acids obtained from a second subject comprising third and/or fourth distinguishable nucleic acid barcodes, a third library of nucleic acids obtained from a third subject comprising fifth and/or sixth distinguishable nucleic acid barcodes, a fourth library of nucleic acids obtained from a fourth subject comprising seventh and/or eighth distinguishable nucleic acid barcodes, and so forth. In some embodiments, each nucleic acid of the library obtained from a single source comprises 1, 2, 3, or 4 distinguishable nucleic acid barcodes, wherein each distinguishable nucleic acid barcode comprises a different sequence and wherein each distinguishable nucleic acid barcode redundantly identifies the same source. In some embodiments, a nucleic acid library comprising a plurality of library inserts obtained from a plurality of samples comprises at least 8, at least 10, at least 15, or at least 20 distinguishable nucleic acid barcodes. In some embodiments, a nucleic acid library, e.g., a nucleic acid library comprising a plurality of library inserts obtained from a plurality of samples or sources, comprises 10 or more, 20 or more, 50 or more, or 100 or more distinguishable nucleic acid barcodes.
In some embodiments, the non-native nucleic acid comprises one or more distinguishable nucleic acid barcodes. In certain embodiments, the non-native nucleic acid comprises one or two distinguishable nucleic acid barcodes. In certain embodiments, the non-natural nucleic acid does not comprise a distinguishable nucleic acid barcode. In some embodiments, each nucleic acid of the library comprises (i) a library insert, (ii) a first non-natural nucleic acid, (iii) a second non-natural nucleic acid, and (iv) a distinguishable nucleic acid barcode, wherein the first non-natural nucleic acid and the second non-natural nucleic acid are located on opposite sides of the library insert, and either the first non-natural nucleic acid or the second non-natural nucleic acid comprises the distinguishable nucleic acid barcode. In some embodiments, each nucleic acid of the library comprises (i) a library insert, (ii) a first non-natural nucleic acid, (iii) a second non-natural nucleic acid, (iv) a first distinguishable nucleic acid barcode and (v) a second distinguishable nucleic acid barcode, wherein the first non-natural nucleic acid and the second non-natural nucleic acid are located on opposite sides of the library insert, the first non-natural nucleic acid comprises the first distinguishable nucleic acid barcode and the second non-natural nucleic acid comprises the second distinguishable nucleic acid barcode.
Non-natural nucleic acids comprising a distinguishable nucleic acid barcode typically comprise one or two portions adjacent to one or both ends of the distinguishable nucleic acid barcode. In some embodiments, the portion of the non-natural nucleic acid adjacent to the end of the distinguishable nucleic acid barcode sequence does not comprise any nucleotides of the distinguishable nucleic acid barcode sequence. In some embodiments, the portion of the non-natural nucleic acid adjacent to the end of the distinguishable nucleic acid barcode sequence comprises 1, 2, or 3 consecutive nucleotides of the distinguishable nucleic acid barcode sequence, wherein 1, 2, or 3 consecutive nucleotides are located at the end of the distinguishable nucleic acid barcode sequence. The portion of the non-natural nucleic acid adjacent to the end of the distinguishable nucleic acid barcode typically comprises 5 to 75 nucleotides, 5 to 50 nucleotides, 5 to 45 nucleotides, 5 to 40 nucleotides, 5 to 35 nucleotides, 5 to 30 nucleotides, 5 to 25 nucleotides, 5 to 20 nucleotides, or 5 to 15 nucleotides located 5 'or 3' to one end of the distinguishable nucleic acid barcode. In some embodiments, the portion of the non-native nucleic acid proximal to the end of the distinguishable nucleic acid barcode is located 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 nucleotides from the end of the distinguishable nucleic acid barcode. In some embodiments, the portion of the non-native nucleic acid adjacent to the end of the distinguishable nucleic acid barcode sequence overlaps with the end of the distinguishable nucleic acid barcode sequence by 1, 2, or 3 nucleotides.
Combination pair
In some embodiments, the compositions or methods described herein comprise one or more binding pairs. In certain embodiments, the nucleic acid comprises one or more members of a binding pair. In some embodiments, a binding pair comprises at least two members (e.g., molecules) that are non-covalently bound to each other. The members of a binding pair typically bind specifically to each other. The members of a binding pair typically bind reversibly to each other, for example wherein the association of the two members of the binding pair can be disassociated by a suitable method. Any suitable binding pair or member thereof may be used in the compositions or methods described herein. Non-limiting examples of binding pairs include complementary nucleic acids, antibodies/antigens, antibodies/antibodies, antibodies/antibody fragments, antibodies/antibody receptors, antibodies/protein a or protein G, haptens/anti-haptens, sulfhydryl/maleimide, sulfhydryl/haloacetyl derivatives, amine/isocyanurates (isotriocy), amine/succinimidyl esters, amine/sulfonyl halides, biotin/avidin, biotin/streptavidin, folate/folate binding protein, receptors/ligands, vitamin B12/intrinsic factor, analogs thereof, derivatives thereof, binding portions thereof, and the like, or combinations thereof. In some embodiments, the binding pair comprises a metal or magnetic material and a magnet. Non-limiting examples of members of a binding pair include an antibody, an antibody fragment, a reduced antibody, a chemically-modified antibody, an antibody receptor, an antigen, a hapten, an anti-hapten, a peptide, a protein, a nucleic acid (e.g., double-stranded DNA (dsDNA), single-stranded DNA (ssDNA), or RNA), a nucleotide analog or derivative (e.g., bromodeoxyuridine (BrdU), an alkyl moiety (e.g., methyl moiety on methylated DNA or methylated histone), an alkanoyl (alkonyl) moiety (e.g., acetyl of acetylated protein (e.g., acetylated histone)), an alkanoic acid or alkanoate moiety (e.g., fatty acid), a glycerol moiety (e.g., lipid), a phosphoryl moiety, a glycosyl moiety, a ubiquitin moiety, a lectin, an aptamer, a receptor, a ligand, a metal ion, avidin, neutravidin, biotin, B12, an intrinsic factor, an analog thereof, a peptide, derivatives thereof, binding moieties thereof, and the like, or combinations thereof. For example, in certain embodiments, one member of the binding pair comprises biotin or an analog or derivative thereof, and the other member of the pair comprises avidin or an analog or derivative thereof. In another example, in certain embodiments, one member of the binding pair comprises a suitable metal (e.g., a matrix comprising a metal, metal nanoparticles, iron) and the other member comprises a magnet.
Joint
In some embodiments, the distinguishable identifier and/or the member of the binding pair is indirectly associated or bound to the nucleic acid via a linker. In certain embodiments, the distinguishable identifier is indirectly associated or bound to a member of the binding pair via a linker. The linker may provide a mechanism for covalently linking the distinguishable identifier and/or members of the binding pair to the nucleic acid or to each other. Any suitable linker may be used in the compositions or methods described herein. Non-limiting examples of suitable linkers include: silane, thiol, phosphonic acid, polyethylene glycol (PEG). Methods of linking two or more molecules using linkers are well known in the art and are sometimes referred to as "crosslinking". Non-limiting examples of crosslinking include amines reacted with N-hydroxysuccinimide (NHS) esters, imido esters, pentafluorophenyl (PFP) ester, hydroxymethylphosphine, ethylene oxide, or any other carbonyl compound; a carboxyl group which reacts with carbodiimide; a thiol group that reacts with maleimide, haloacetyl, pyridyl disulfide, and/or vinyl sulfone; an aldehyde reacted with hydrazine; any non-selective group that reacts with a bis-aziridine (diazirine) and/or an aryl azide; hydroxyl groups reactive with isocyanates; hydroxylamine reacted with carbonyl compounds, and the like, and combinations thereof.
Nucleic acid sequencing
In certain embodiments, nucleic acids (e.g., amplicons; nucleic acids of a library; isolated, purified, and/or captured nucleic acids) are analyzed by methods that include nucleic acid sequencing. In some embodiments, the nucleic acid can be sequenced. In some embodiments, a complete or substantially complete sequence, and sometimes a partial sequence, is obtained.
Any suitable method of sequencing nucleic acids may be used, non-limiting examples of which include Maxim & Gilbert, chain termination methods, sequencing-by-synthesis, sequencing-by-ligation, mass spectrometry, microscope-based techniques, and the like, or combinations thereof. In some embodiments, a first generation technology, e.g., a Sanger sequencing method including automated Sanger sequencing methods, including microfluidic Sanger sequencing, can be used in the methods provided herein. In some embodiments, sequencing techniques including the use of nucleic acid imaging techniques, such as Transmission Electron Microscopy (TEM) and Atomic Force Microscopy (AFM), may be used. In some embodiments, a high throughput sequencing method is used. High throughput sequencing methods typically involve clonally amplified DNA templates or single DNA molecules sequenced in a massively parallel fashion, sometimes within a flow cell. Next generation (e.g., second and third generation) sequencing technologies capable of sequencing DNA in a massively parallel manner can be used in the methods described herein, and are collectively referred to herein as "massively parallel sequencing" (MPS) or "massively parallel nucleic acid sequencing. In some embodiments, the MPS sequencing method utilizes a targeted approach, wherein sequence reads are generated from a particular chromosome, gene, or region of interest. The particular chromosome, gene, or region of interest is sometimes referred to herein as a targeted genomic region. In certain embodiments, a non-targeted approach is used, wherein most or all of the nucleic acid fragments in the sample are randomly sequenced, amplified, and/or captured.
MPS sequencing sometimes utilizes sequencing by synthesis and certain imaging processes. Nucleic acid sequencing techniques that can be used in the methods described herein are sequencing by synthesis and reversible terminator-based sequencing (e.g., Genome Analyzer of Illumina; Genome Analyzer II; HISEQ 2000; HISEQ 2500(Illumina, San Diego CA)). Using this technique, millions of nucleic acid (e.g., DNA) fragments can be sequenced in parallel. In one example of this type of sequencing technique, a flow cell is used that houses an optically clear slide, on the surface of which 8 to 16 separate lanes are bound oligonucleotide anchors (e.g., adapter primers).
In some embodiments, sequencing-by-synthesis comprises iteratively adding (e.g., by covalent addition) nucleotides to a primer or a preexisting nucleic acid strand in a template-directed manner. Each iterative addition of nucleotides is detected and the process is repeated multiple times until the sequence of the nucleic acid strand is obtained. The length of the obtained sequence depends in part on the number of addition and detection steps performed. In some embodiments of sequencing-by-synthesis, one, two, three, or more nucleotides of the same type (e.g., A, G, C or T) are added and detected in a round of nucleotides. The nucleotides may be added by any suitable method (e.g. enzymatically or chemically). For example, in some embodiments, a polymerase or ligase adds nucleotides to a primer or a preexisting nucleic acid strand in a template-directed manner. In some embodiments of sequencing-by-synthesis, different types of nucleotides, nucleotide analogs, and/or identifiers are used. In some embodiments, reversible terminators and/or removable (e.g., cleavable) identifiers are used. In some embodiments, fluorescently labeled nucleotides and/or nucleotide analogs are used. In certain embodiments, sequencing-by-synthesis comprises cleavage (e.g., cleavage and removal of an identifier) and/or washing steps. In some embodiments, the incorporation of one or more nucleotides is detected by a suitable method described herein or known in the art, non-limiting examples of which include any suitable imaging device or machine, suitable camera, digital camera, CCD (charge coupled device) based imaging device (e.g., CDD camera), CMOS (complementary metal oxide silicon) based imaging device (e.g., CMOS camera), photodiode (e.g., photomultiplier tube), electron microscope, field effect transistor (e.g., DNA field effect transistor), ISFET ion sensor (e.g., CHEMFET sensor), and the like, or combinations thereof. Other sequencing methods that can be used to perform the methods herein include digital PCR and sequencing by hybridization.
A suitable MPS method, system, or technology platform for performing the methods described herein can be used to obtain nucleic acid sequencing reads. Non-limiting examples of MPS platforms include Illumina/Solex/HiSeq (e.g., Illumina's Genome Analyzer; Genome Analyzer II; HISEQ 2000; HISEQ), SOLID, Roche/454, PACBIO and/or SMRT, Helicos tube Single Molecule Sequencing, Ion Torque and Ionsemicondurctor based Sequencing (e.g., developed by Life Technologies), WildFire, 5500xl W and/or 5500xl W Genetic Analyzer based Technologies (e.g., developed and sold by Life Technologies, U.S. patent publication No. 201US30012399), Polony Sequencing, pyrophosphoric acid, Massively Parallel Signal Sequencing (MPSS), RNA (RNAS) polymerase (RNAP) Sequencing, Laser GeneCommen platform, chemical nanopore arrays (EMS) based Sequencing transistor arrays, nano-based Sequencing by field effect microscopy, such as developed by microscope, and the like.
Other sequencing methods that can be used to perform the methods herein include digital PCR and sequencing by hybridization. Digital polymerase chain reaction (digital PCR or dPCR) can be used to directly identify and quantify nucleic acids in a sample. In some embodiments, digital PCR may be performed in an emulsion. For example, individual nucleic acids are isolated, e.g., in a microfluidic chamber device, and each nucleic acid is amplified separately by PCR. Nucleic acids can be separated such that there is no more than one nucleic acid per well. In some embodiments, different probes may be used to distinguish between various alleles (e.g., fetal alleles and maternal alleles). Alleles can be enumerated to determine copy number.
In certain embodiments, sequencing by hybridization may be used. The method comprises contacting a plurality of polynucleotide sequences with a plurality of polynucleotide probes, wherein each of the plurality of polynucleotide probes may optionally be tethered to a substrate. In some embodiments, the substrate can be a flat surface with an array of known nucleotide sequences. The pattern of hybridization to the array can be used to determine the polynucleotide sequences present in the sample. In some embodiments, each probe is tethered to a bead, such as a magnetic bead or the like. Hybridization to beads can be identified and used to identify a plurality of polynucleotide sequences in a sample.
In some embodiments, chromosome-specific sequencing is performed. In some embodiments, chromosome-specific sequencing is performed using DANSR (digital analysis of selected regions). Digital analysis of selected regions enables simultaneous quantification of hundreds of loci by DNA-dependent linkage of two locus-specific oligonucleotides via an intermediate "bridge" oligonucleotide to form a PCR template. In some embodiments, chromosome-specific sequencing is performed by generating a library enriched for chromosome-specific sequences. In some embodiments, sequence reads are obtained only for the selected chromosome set.
Competitor nucleic acids
Competitor nucleic acids are typically added prior to the hybridization process to reduce unwanted and non-specific hybridization events. The competitor nucleic acid may comprise a repeat nucleic acid. Repeated endogenous nucleic acids, such as Alu sequences or LINE sequences, are typically present in a nucleic acid library. Sometimes endogenous repeat nucleic acids can hybridize to each other, resulting in contamination of the captured hybridization mixture. This type of contamination can be reduced in part by adding excess exogenous competitor nucleic acid prior to hybridization. Any suitable competitor nucleic acid may be used in the compositions or methods described herein. In some embodiments, the competitor nucleic acid comprises C0t-1DNA, which can bind Alu, LINE, and other repeats present in the nucleic acid library. The C0t-1DNA may be obtained from a suitable organism and may comprise a mixture of nucleic acids from different organisms. The C0t-1DNA may be obtained from a suitable tissue of an organism. C0t-1DNA sometimes comprises nucleic acid isolated from placenta.
Blocking nucleic acids
In some aspects, provided herein are improved blocking methods and compositions. Typically, through hybridization events, unwanted nucleic acids contaminate the enriched nucleic acid pool after the hybrid capture process is complete. Most of the unwanted sequences are sometimes due to undesired hybridization events between the same portions of the terminal adaptor sequences (e.g., barcodes, portions complementary to the flow cell anchors and or primers, etc.) of the adaptor-ligated library inserts. Sometimes, unwanted library inserts can anneal to each other through their terminal adaptors, resulting in a "daisy chain" of otherwise unwanted DNA fragments ligated and separated together. In this way, capture of a single desired fragment can result in a large number of undesired fragments, which reduces the overall efficiency of the enrichment process.
In some embodiments, the so-called "daisy chain" effect and other unwanted hybridization events can be reduced by using blocking nucleic acids that are directed to hybridize to portions of the non-natural nucleic acids (e.g., adapter sequences) of the library. Blocking oligonucleotides are known in the art and are often configured to bind to and block hybridization between barcode sequences of a library. Conventional blocking oligonucleotides are typically 50 or more nucleic acids in length and are configured to hybridize to a barcode sequence and to synthetic nucleic acid portions flanking each side of the nucleic acid barcode. Conventional blocker oligonucleotides are therefore relatively long (e.g., > 50 nucleotides) so that they can anneal to non-native sequences flanking the barcode region and ensure high melting temperatures between the blocker oligonucleotides and their target sequences. In some embodiments, the methods herein use one or more blocking nucleic acids (e.g., traditional blocking nucleic acids).
In some embodiments, for multiplex sequencing methods in which multiple adaptor-ligated libraries are mixed, multiple sets of blocking nucleic acids must be synthesized, each set being specific for many different adaptors of each library. Thus, high throughput multiplex sequencing reactions typically involve 8, 10, 15, or 20 or more different barcode sequences, which require the synthesis of 8, 10, 15, or 20 or more blocking oligonucleotides, each configured to bind to each of the unique barcode sequences. This strategy is expensive and time consuming because many different blocking oligonucleotides of relatively long length must be designed and synthesized for multiplex sequencing of mixed libraries.
In some embodiments, the blocking nucleic acid is a U-block (U-block) nucleic acid. In some embodiments, the compositions or methods herein comprise a U-block (e.g., universal block) nucleic acid. The U-block nucleic acids of the compositions can be substantially identical or substantially different. U-block nucleic acids have several advantages over traditional blocking oligomers. First, the U-block nucleic acid does not substantially hybridize to a nucleic acid barcode sequence, nor is the U-block nucleic acid configured to hybridize to a nucleic acid barcode sequence. Thus, manipulation, capture, and multiplex sequencing of complex nucleic acid libraries containing 8, 10, 15, or 20 or more different barcode sequences does not require U-block nucleic acids specific for each unique barcode sequence. Also, the U-block nucleic acids are relatively short compared to conventional block oligonucleotides, making nucleic acid synthesis of the U-block nucleic acids more economical. In certain embodiments, a U-block nucleic acid provided herein has a nominal, average, or absolute length of 45 nucleotides or less, 40 nucleotides or less, 35 nucleotides or less, 30 nucleotides or less, 25 nucleotides or less, 20 nucleotides or less, 15 nucleotides or less, or 10 nucleotides or less. In some embodiments, a U-block nucleic acid has a nominal, average, mean, or absolute length of 8 to about 40, 8 to about 35, 8 to about 30, 8 to about 25, or 8 to about 20 nucleotides. The U-block nucleic acid may comprise any suitable nucleic acid, nucleotide or nucleotide analogue. In some embodiments, the U-block nucleic acid is synthetic (e.g., synthesized by the human hand). The U-block nucleic acid may be an oligonucleotide.
U-block nucleic acids are typically configured to block unwanted hybridization and/or subsequent amplification of non-natural portions of a nucleic acid library. In certain embodiments, the U-block nucleic acid is configured to hybridize to a synthetic nucleic acid region (non-natural nucleic acid region) flanking the barcode nucleic acid. Thus, the compositions herein typically comprise at least two U-block nucleic acids configured to hybridize to opposite sides of a distinguishable nucleic acid barcode. The region of synthetic nucleic acid flanking the barcode nucleic acid sequence is typically relatively short (e.g., 45 nucleotides or less, 40 nucleotides or less, 35 nucleotides or less, 30 nucleotides or less, 25 nucleotides or less, 20 nucleotides or less, 15 nucleotides or less, or 10 nucleotides or less) and thus provides a relatively short segment of nucleic acid for hybridization of the U-block nucleic acid. The early prototypes of U-block nucleic acids developed by the inventors herein consist exclusively of standard nucleotide bases and are inefficient at blocking unwanted hybridization events. By altering some or all of the nucleobases of the U-block nucleic acid, the efficiency and blocking capacity can be greatly improved. Thus, in some embodiments, a U-block nucleic acid is configured to comprise a higher melting temperature, in part by comprising a non-standard or altered nucleobase that increases the Tm of the U-block nucleic acid. In some embodiments, the U-block nucleic acid comprises a Tm that is higher than the Tm of an unmodified U-block nucleic acid of similar sequence consisting of standard nucleotides selected from the group consisting of guanine, cytosine, thymine, adenine and uracil. Any suitable change can be used to increase the Tm of the U-block nucleic acid. In some embodiments, U-block nucleic acids comprise altered nucleotides, nucleotide analogs, and/or altered nucleotide linkages, non-limiting examples of which include locked nucleic acids (LNAs, e.g., bicyclic nucleic acids), bridged nucleic acids (BNAs, e.g., constrained nucleic acids), C5-modified pyrimidine bases (e.g., 5-methyl-dC, propynyl pyrimidines, etc.), and alternative backbone chemistries, such as Peptide Nucleic Acids (PNAs), morpholino compounds, and the like, or combinations thereof. In some embodiments, the bridge nucleic acid is an altered RNA nucleotide. Any suitable BNA can be used in the compositions or methods described herein. In certain embodiments, the BNA monomer can comprise a five-, six-, or even seven-membered bridging structure. Non-limiting examples of new generation BNA monomers include 2', 4' -BNANC [ NH ], 2', 4' -BNANC [ NMe ], and 2', 4' -BNANC [ NBn ]. Non-base modifying agents may also be introduced into the U-block nucleic acid to increase Tm (or binding affinity), non-limiting examples of which include Minor Groove Binders (MGBs), spermines, G-clips, Uaq anthraquinone caps, and the like, or combinations thereof. More than one type of Tm enhancing change may be used in a U-block nucleic acid, e.g., a BNA nucleotide monomer in combination with a terminal MGB group. Many methods of increasing the Tm of complementary nucleic acids are known to those of skill in the art, and the use of all such variations is considered within the scope of the present invention. In some embodiments, a U-block nucleic acid comprises a melting temperature (Tm) of at least 40 ℃, at least 45 ℃, at least 50 ℃, at least 55 ℃, at least 60 ℃, at least 65 ℃, at least 70 ℃, at least 75 ℃, or at least 80 ℃.
In certain embodiments, the U-block nucleic acid is configured to specifically hybridize to one or more non-native nucleic acids of the library. Non-natural nucleic acids typically comprise synthetic nucleic acid sequences. In some embodiments, the non-natural nucleic acid does not comprise genomic DNA, a gene, mRNA, cDNA, or a portion thereof. In some embodiments, the U-block nucleic acid is configured to specifically hybridize to one or more amplicons of the library, wherein the one or more amplicons comprise a synthetic nucleic acid (e.g., one or more adaptor sequences, capture sequences, or primer binding sites). In some embodiments, the U-block nucleic acid is configured to specifically hybridize to one or more adaptors or portions thereof. In certain embodiments, the U-block nucleic acid is not configured to hybridize to a library insert. In certain embodiments, the U-block nucleic acid does not substantially hybridize to the nucleic acid insert. In certain embodiments, the U-block nucleic acid is not configured to hybridize to a nucleic acid barcode and is not complementary to a majority of the barcode sequence. In certain embodiments, the U-block nucleic acid does not substantially hybridize to the nucleic acid barcode or any portion of the nucleic acid barcode.
U-block nucleic acids typically comprise or consist of a nucleic acid sequence that is substantially complementary to a non-natural nucleic acid. U-block nucleic acids are sometimes configured to specifically hybridize to one or more non-native nucleic acids of a nucleic acid library. In some embodiments, each nucleic acid of the library comprises one or more (e.g., 1, 2, 3, 4, or more) non-natural nucleic acids, wherein the non-natural nucleic acids are common (e.g., common) to each nucleic acid of the library. For example, a nucleic acid library can be generated from two or more samples, wherein the nucleic acids derived from each sample comprise a unique distinguishable barcode sequence incorporated into an adapter sequence, and wherein the adapter comprises one or more (e.g., 1, 2, 3, 4, or more) identical non-natural nucleic acids (e.g., synthetic nucleic acids, universal nucleic acid sequences). Non-natural nucleic acids that are substantially identical and common between the nucleic acids of the library are sometimes referred to as universal nucleic acids. The U-block nucleic acid is typically substantially complementary to and configured to specifically hybridize with a universal nucleic acid sequence (e.g., an adaptor or portion thereof).
In some embodiments, the U-block nucleic acid is configured to block extension by a polymerase. In some embodiments, the U-block nucleic acid is configured to block extension of the U-block nucleic acid by a polymerase. For example, the U-block nucleic acid may comprise a suitable 3 'chain terminator (e.g., a 2',3 'dideoxynucleotide) or a suitable functional group that prevents a polymerase from extending the 3' end of the U-block nucleic acid (e.g., by forming a phosphodiester bond). Thus, a U-block nucleic acid configured to block extension by a polymerase typically comprises a suitable 3 'chain terminator (e.g., a 2',3 'dideoxynucleotide) or a suitable functional group that prevents the polymerase from extending the 3' end of the U-block nucleic acid. Thus, in certain embodiments, the U-block nucleic acid is not a nucleic acid primer suitable for amplification (e.g., PCR) or extension by a polymerase.
In some embodiments, the compositions herein comprise one or more nucleic acid libraries (e.g., a plurality of library inserts) derived from a plurality of samples and 2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, or 10 or more uniquely distinguishable barcodes. In certain embodiments, such compositions comprise no more than 4, sometimes no more than 8U-block nucleic acids, each configured to specifically hybridize to a non-native nucleic acid of each nucleic acid of the library, or a portion thereof. In some embodiments, the compositions herein comprise 1, 2, 3, 4, 5, 6, 7, or 8U-block nucleic acids and/or no more than 1, 2, 3, 4, 5, 6, 7, or 8U-block nucleic acids. In certain embodiments, the compositions herein comprise 2 to 4, 2 to 6, 2 to 8, or 4 to 8U-block nucleic acids. In some embodiments, the compositions herein comprise 1, 2, 3, 4, 5, 6, 7, or 8U-block nucleic acids, wherein (i) each U-block nucleic acid is substantially complementary to a portion of the first and/or second non-natural nucleic acid, (ii) at least one of the U-block nucleic acids is configured to hybridize adjacent to a first end of the distinguishable nucleic acid barcode, (iii) at least one of the U-block nucleic acids is configured to hybridize adjacent to a second end of the distinguishable nucleic acid barcode, and (iv) each U-block nucleic acid is not substantially hybridized to any portion of the distinguishable nucleic acid barcode.
In some embodiments, the U-block nucleic acid is configured to hybridize adjacent to an end of a distinguishable nucleic acid barcode. In some embodiments, a U-block nucleic acid configured to hybridize adjacent to an end of a distinguishable nucleic acid barcode refers to a U-block nucleic acid that is substantially complementary to a portion of a non-natural nucleic acid that is located adjacent to the distinguishable nucleic acid barcode. U-block nucleic acids are typically configured to hybridize to non-natural nucleic acids of a library comprising distinguishable nucleic acid barcodes, wherein the U-block nucleic acids are configured to hybridize to opposite sides of a barcode sequence. Thus, when a U-block nucleic acid hybridizes to a non-natural nucleic acid of a library comprising distinguishable nucleic acid barcodes, the hybridized U-block nucleic acid is flanked by distinguishable nucleic acid barcodes on both sides of the barcode. In certain embodiments, the U-block nucleic acid does not substantially hybridize to any portion of the barcode sequence. In some embodiments, the hybridized U-block nucleic acid may overlap the barcode sequence by 1, 2, 3, 4, 5, or 6 nucleotides. Thus, in certain embodiments, U-block nucleic acids that do not substantially hybridize to a barcode sequence can hybridize to a small portion (e.g., 6 nucleotides or less) of the barcode sequence.
In some embodiments, the compositions herein comprise up to four U-block nucleic acids. The U-block nucleic acids of the compositions can be identical. For example, where a composition comprises four U-block nucleic acids, the nucleic acid sequences of each of the U-block nucleic acids can be substantially identical (e.g., substantially identical). In some embodiments, the composition comprises four U-block nucleic acids, wherein the nucleic acid sequences of each of the U-block nucleic acids are substantially different.
Capture of nucleic acids
In some embodiments, the compositions or methods herein comprise a capture nucleic acid. The capture nucleic acid typically comprises a nucleic acid portion. In some embodiments, the capture nucleic acid is configured to specifically hybridize to a target nucleic acid, wherein the capture nucleic acid and its hybridized target can be captured by a suitable technique, such as a pull-down method. Any suitable capture nucleic acid or set of capture nucleic acids can be used in the methods or compositions herein. In some embodiments, the capture nucleic acid is an oligonucleotide.
The capture nucleic acid is typically bound (e.g., covalently or non-covalently) directly or indirectly to a member of a suitable binding pair. In certain embodiments, the capture nucleic acid comprises a member of a suitable binding pair. In some embodiments, the member of the binding pair is bound to the capture nucleic acid via a linker.
The capture nucleic acid comprising the member of the binding pair can be captured together with the nucleic acid target to which it hybridizes using a suitable capture method (e.g., captured by a capture method). The process of capturing nucleic acid (e.g., by a capture method) typically provides captured nucleic acid (e.g., captured nucleic acid). The terms "captured" and "enriched" as used herein may refer to a nucleic acid or subset of nucleic acids, provided that it contains fewer species of nucleic acids than in the sample prior to the capture method. A composition comprising captured nucleic acids may be free of about 50%, 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or greater than 99% of other (e.g., unwanted) nucleic acid species. In some embodiments, the captured nucleic acid comprises an enriched nucleic acid. The enriched nucleic acid may comprise an amount of one or more nucleic acid species enriched by at least 1.2-fold, 1.3-fold, 1.4-fold, 1.5-fold, 2-fold, 5-fold, 10-fold, 100-fold, or 1000-fold compared to the amount of one or more nucleic acid species prior to application of the enrichment or capture method. Non-limiting examples of capture methods include pull-down methods (e.g., gravity pull-down methods, centrifugal pull-down methods, magnetic pull-down methods), immunoprecipitation, and various column purification methods. Isolation and/or purification of captured nucleic acids typically involves association of a first member of a binding pair (e.g., binding to the captured nucleic acids) with a second member of the binding pair (e.g., binding to a matrix). The capture nucleic acid is sometimes bound to a first member of a binding pair that is configured to be securely bound and/or associated with a second member, where the second member is sometimes bound to a suitable substrate. In some embodiments, the capture nucleic acid is bound to a first member of a binding pair comprising a matrix. In some embodiments, the capture nucleic acid is indirectly associated with the matrix (e.g., via a linker or an intermediate binding molecule, such as an antibody). In some embodiments, the capture nucleic acid is bound to a first member of a binding pair comprising a magnetic substrate (e.g., a magnetic bead, a ferrous bead), and the second member of the binding pair is a magnet. The magnet and magnetic material may be used to capture (e.g., pull down) the captured nucleic acid and its bound target.
In some embodiments, a nucleic acid and/or amplicon (e.g., a nucleic acid and/or amplicon of a library) is contacted with a capture nucleic acid. In certain embodiments, the method comprises contacting a mixture comprising the nucleic acids of the library, the blocking nucleic acid, and optionally one or more competitor nucleic acids with the capture nucleic acid. Capture nucleic acids and methods of making capture nucleic acids are known in the art. In certain embodiments, the capture nucleic acids are used to capture a subset of nucleic acids of interest from a library comprising a plurality of nucleic acids of the library insert obtained from 2 or more, 10 or more, 50 or more, or 100 or more subjects, samples, or sources.
The capture nucleic acid can be a nucleic acid configured to specifically hybridize to a portion of one or more nucleic acids (e.g., a selected nucleic acid, a target nucleic acid, a nucleic acid of interest) of the library. In some embodiments, a capture nucleic acid configured to specifically hybridize to a portion of one or more nucleic acids of a library is substantially complementary to a suitable portion of a nucleic acid of the library. In some embodiments, the capture nucleic acid is configured to specifically hybridize to a portion of a non-native nucleic acid or a portion thereof. In some embodiments, the capture nucleic acid is configured to specifically hybridize to a portion of the adapter. In some embodiments, the capture nucleic acid is configured to specifically hybridize to a portion of one or more library inserts. The capture nucleic acid can be configured to hybridize to a promoter, enhancer, intron, exon, poly a segment, poly T segment, any suitable translational or transcriptional control sequence, and the like, or combinations thereof. The set of capture nucleic acids can be configured to specifically hybridize to a subset of genes (e.g., a set of genes of a chromosome, e.g., a set of genes expressing a family of enzymes) or any subset of nucleic acids of a library. In some embodiments, the capture nucleic acid set is configured to hybridize at or near one or more genomic regions suspected of containing a genetic variation (e.g., a deletion, an insertion, a SNP, etc.). The nucleic acid configured to specifically hybridize to the second nucleic acid is typically substantially complementary to the second nucleic acid. The capture nucleic acid is typically substantially complementary to a portion of one or more nucleic acid targets (e.g., a subset of the nucleic acids of the library).
In some embodiments, the method comprises preparing a mixture comprising or consisting essentially of a library of nucleic acids, a blocker nucleic acid, a capture nucleic acid, and optionally a competitor nucleic acid, wherein the mixture is subjected to a capture method. For example, in some embodiments, streptavidin beads are added, and the nucleic acids of the mixture (e.g., hybridized to the capture nucleic acids) associated with the capture nucleic acids are recovered (e.g., by centrifugation, filtration, immunoprecipitation, and/or by magnetic precipitation). In some embodiments, a washing step is used (e.g., by washing with ethanol). The captured nucleic acids (e.g., hybridized targets) can be eluted from the captured nucleic acids using a suitable method.
In some embodiments, certain methods or steps of methods are performed in the absence of a binding pair (e.g., two members of a binding pair). For example, the mixture typically comprises the capture nucleic acid and a first member of a binding pair (e.g., bound to the capture nucleic acid), and the mixture does not comprise a second member of the binding pair (e.g., the second member is configured to specifically bind to the first member). In some embodiments, the nucleic acids of the library are contacted with a blocking nucleic acid, a capture nucleic acid comprising a first member of a binding pair, and/or a competitor nucleic acid in the absence of a second member of the binding pair. In some embodiments, a mixture comprising the nucleic acids of the library, the blocking nucleic acids, the capture nucleic acids comprising the first member of the binding pair, and/or the competitor nucleic acids is purified in the absence of the second member of the binding pair. In some embodiments, a mixture comprising the nucleic acids of the library, the blocking nucleic acids, the capture nucleic acids comprising the first member of the binding pair, and/or the competitor nucleic acids hybridizes under suitable hybridization conditions in the absence of the second member of the binding pair.
In some embodiments, the capture nucleic acid is directly or indirectly bound (e.g., covalently or non-covalently bound) to a suitable matrix. In certain embodiments, the members of the binding pair are directly or indirectly bound (e.g., covalently or non-covalently bound) to a suitable matrix. Any suitable matrix may be used. In certain embodiments, the substrate comprises a surface (e.g., a surface of a flow cell, a surface of a tube, a surface of a chip), such as a metal surface (e.g., steel, gold, silver, aluminum, silicon, and copper). In some embodiments, the substrate (e.g., substrate surface) is coated and/or comprises functional groups and/or inert materials. In certain embodiments, the substrate comprises, for example, a bead, chip, capillary, plate, membrane, wafer (e.g., silicon wafer), comb, or needle. In some embodiments, the matrix comprises beads and/or nanoparticles. The matrix may be made of a suitable material, non-limiting examples of which include plastics or suitable polymers (e.g., polycarbonate, poly (vinyl alcohol), poly (divinylbenzene), polystyrene, polyamide, polyester, polyvinylidene fluoride (PVDF), polyethylene, polyurethane, polypropylene, etc.), borosilicate, glass, nylon, Wang resin, Merrifield resin, metals (e.g., iron, metal alloys), sepharose, agarose, polyacrylamide, dextran, cellulose, etc., or combinations thereof. In some embodiments, the matrix comprises a magnetic material (e.g., iron, nickel, cobalt, platinum, aluminum, etc.). In certain embodiments, the substrate comprises magnetic beads (e.g.,hematite, AMPureXP. Magnets can be used to purify and/or capture nucleic acids bound to certain substrates (e.g., substrates comprising metal or magnetic materials).
In certain embodiments, the capture nucleic acid and the nucleic acid that specifically hybridizes to the capture nucleic acid (e.g., the target nucleic acid) are captured by a suitable capture method, thereby providing a captured nucleic acid. Processes that include capture typically provide enriched nucleic acids (e.g., nucleic acids enriched for target nucleic acids). The captured nucleic acids are typically enriched for one or more species of nucleic acid. The captured nucleic acids can comprise one or more of a capture nucleic acid, a nucleic acid that specifically hybridizes to a capture nucleic acid (e.g., a captured nucleic acid, a target nucleic acid, an enriched nucleic acid of a library), a member of a binding pair, and/or a substrate. The nucleic acid (target nucleic acid) specifically hybridizing with the capture nucleic acid can be recovered by an appropriate method. For example, in some embodiments, nucleic acids enriched by a capture method can be isolated by denaturation (e.g., by heating) or by increasing or decreasing the salt concentration of a mixture comprising the captured nucleic acids. In some embodiments, the capture method comprises recovering the enriched nucleic acid using a filter, membrane, or column. In certain embodiments, gravity and/or centrifugation is used to capture nucleic acids and/or recover enriched nucleic acids. In certain embodiments, the capture method does not use centrifugation. In certain embodiments, the magnet is used to capture and/or recover the enriched nucleic acid. The nucleic acid bound to the matrix can generally be washed by a suitable method to remove unbound or unwanted material. In certain embodiments, the stringency of the wash solution can be adjusted by suitable means.
The captured nucleic acids, nucleic acids that specifically hybridize to the captured nucleic acids, enriched nucleic acids, and/or amplicons thereof can be analyzed by suitable methods, which can include, for example, processes including nucleic acid sequencing or mass spectrometry.
Fixing
In certain embodiments, the nucleic acid is immobilized (e.g., on a substrate). The nucleic acid may be immobilized by any suitable method. The nucleic acid may be immobilized directly or indirectly to a suitable matrix or material. The nucleic acid immobilized to the substrate may be covalently or non-covalently bound to the substrate. In certain embodiments, the nucleic acid is reversibly immobilized to a matrix and can be dissociated or removed from the matrix using a suitable method. In some embodiments, the nucleic acid comprising a first member of a binding pair is immobilized to a substrate by binding to a second member of the binding pair bound to or associated with the substrate. The nucleic acid may be non-specifically immobilized to a matrix, for example, wherein the matrix comprises anions (e.g., anion exchange groups, e.g., positively charged functional groups). In certain embodiments, the nucleic acid is immobilized to the substrate by magnetic attraction. For example, the nucleic acid may comprise a magnetic material (e.g., a metal), and the nucleic acid may be immobilized to a substrate (e.g., a portion of a tube, a surface) through the use of a magnet. In certain embodiments, the magnet is not in solution and/or is not in direct contact with a nucleic acid comprising a magnetic material. In certain embodiments, the magnet is in solution and in direct contact with a magnetic material associated with the nucleic acid. In some embodiments, the nucleic acid is immobilized by the formation of covalent bonds, for example, by crosslinking functional groups of the nucleic acid (e.g., functional groups of a nucleic acid analog) to a matrix comprising reactive groups. In some embodiments, a nucleic acid (e.g., amplicon, nucleic acid of a library, target nucleic acid) is immobilized to a substrate by hybridization (e.g., specific hybridization, annealing) with another nucleic acid bound (e.g., covalently or non-covalently) to the substrate.
The nucleic acid may be immobilized at any suitable step of the methods described herein. In some embodiments, the nucleic acid is immobilized to a flow cell (e.g., a flow cell surface) or an array (e.g., a chip). In certain embodiments, the methods described herein do not comprise immobilizing nucleic acids to a flow cell or chip. In certain embodiments, the methods described herein do not include immobilizing the nucleic acid to the flow cell or chip until after the capture method. In certain embodiments, the methods described herein do not comprise immobilizing the nucleic acids to a flow cell or chip prior to analyzing the nucleic acids (e.g., by nucleic acid sequencing). In some embodiments, nucleic acids of a library (e.g., amplicons of a library) are ligated to adapters, contacted with blocking nucleic acids, contacted with capture nucleic acids, contacted with competitor nucleic acids, purified and/or hybridized (e.g., subjected to denaturation and annealing processes), and are not immobilized to a flow cell, array or chip prior to or during any or all of the above processes. In certain embodiments, the nucleic acids of the library (e.g., amplicons of the library) are ligated to adapters, contacted with blocking nucleic acids, contacted with capture nucleic acids, contacted with competitor nucleic acids, purified and/or hybridized (e.g., subjected to denaturation and annealing processes) in the absence of a flow cell, array or chip.
Other nucleic acid methods
In some embodiments, the methods herein comprise preparing a mixture for subsequent hybridization and/or capture. In certain embodiments, preparing the mixture comprises contacting the nucleic acids (e.g., amplicons) of one or more nucleic acid libraries with one or more blocking nucleic acids, one or more capture nucleic acids, and/or one or more competitor nucleic acids. In some embodiments, the mixture is prepared prior to denaturation or hybridization. The mixture prepared by the methods described herein can comprise one or more nucleic acid libraries (e.g., amplicons), one or more blocking nucleic acids, one or more capture nucleic acids, and/or one or more competitor nucleic acids. The mixture prepared by the methods described herein may comprise or consist essentially of a library of nucleic acids, a blocking nucleic acid, and a capture nucleic acid. In certain embodiments, the prepared mixture comprises or consists essentially of a library of nucleic acids, a blocker nucleic acid, a capture nucleic acid, and a competitor nucleic acid. The prepared mixture consisting essentially of nucleic acids may contain water, EDTA, PEG, NaCl, and/or buffers (e.g., Tris or HEPES). In some embodiments, the mixture consisting essentially of nucleic acids does not contain a hybridization buffer. In certain embodiments, the method of preparing a mixture of nucleic acids does not include adding a hybridization buffer. In certain embodiments, the prepared mixture consisting essentially of nucleic acids does not contain calcium or magnesium salts. In some embodiments, the mixture of nucleic acids prepared does not contain calcium or magnesium salts, detergents (e.g., SDS), Ficoll, BSA, or polyvinylpyrrolidone. In some embodiments, the mixture consisting essentially of nucleic acids is free of detergents (e.g., SDS), Ficoll, BSA, or polyvinylpyrrolidone.
In some embodiments, the mixture is prepared as described herein, and the mixture is not heated until after the nucleic acids of the mixture are purified. For example, the mixture is not heated to a temperature greater than about 70 ℃, 75 ℃, 80 ℃, 85 ℃, 90 ℃, or greater than 95 ℃ prior to purifying the nucleic acid from the mixture and/or until after the nucleic acid of the mixture is purified.
Genetic variation and medical conditions
The presence or absence of a genetic variation can be determined using the compositions or methods described herein. Genetic variation is typically a specific genetic phenotype present in certain individuals, and often genetic variation is present in statistically significant subpopulations of individuals. In some embodiments, the genetic variation is a chromosomal abnormality (e.g., aneuploidy, replication of one or more chromosomes, loss of one or more chromosomes), a partial chromosomal abnormality or mosaicism (e.g., loss or gain of one or more segments of a chromosome), a translocation, or an inversion. Non-limiting examples of genetic variations include one or more deletions, duplications, insertions, mutations, polymorphisms (e.g., single nucleotide polymorphisms), fusions, repeats (e.g., short tandem repeats), and the like, and combinations thereof. The insertions, repeats, deletions, duplications, mutations, or polymorphisms can be of any length, and in some embodiments, are from about 1 base or base pair (bp) to about 250 megabases (Mb) in length. In some embodiments, the insertion, duplication, deletion, duplication, mutation, or polymorphism is from about 1 base or base pair (bp) to about 50,000 kilobases (kb) in length (e.g., about 10bp, 50bp, 100bp, 500bp, 1kb, 5kb, 10kb, 50kb, 100kb, 500kb, 1000kb, 5000kb, or 10,000kb in length).
In certain embodiments, a genetic variation determined for a subject to be present or absent is sometimes associated with a medical condition. Non-limiting examples of medical conditions include those associated with intellectual impairment (e.g., Down's syndrome), abnormal cell proliferation (e.g., cancer), non-Hodgkin's lymphoma, myelodysplastic syndrome, Williams 'syndrome, Langer-Giedion syndrome, Aliffel's syndrome, Retreore's syndrome, Jacob syndrome, retinoblastoma, Smith-Magenis, Edward's syndrome, papillary renal cell carcinoma, DiGeorge's syndrome, Angelman's syndrome, cat eye syndrome, familial adenomatous polyposis, Miller-Dick syndrome, the presence of microbial nucleic acids (e.g., viral, bacterial, fungal, yeast), and preeclampsia.
Examples
The following examples illustrate certain embodiments and do not limit the present technology.
Example 1: workflow advances in hybrid Capture workflow for efficiency, cost reduction, and data quality improvement
The new and improved methods exemplified herein provide the following advantages.
1. Vacuum and heat based concentration of the reaction mixture was avoided by using zero volume concentration of magnetic beads.
2. DNA decoy concentration for increased reaction kinetics
3. Short nucleic acid concentrations for workflow improvement
4. Denaturation of the completed hybridization reaction
5. Efficient and automatable streptavidin bead manipulation
The conventional protocol for the hybrid capture assay (e.g., website pdf files accessed at http:// www.nimblegen.com/products/lit/06588786001_ SeqCapeZlibrary SR _ UGuide _ v4p2. pdf: Roche-NimbleGen SeqCapeEZ Library SR users' Guide, 8/20/2014) requires a number of manipulations that are absolutely incompatible with high throughput genetic testing methods. The approach provided herein results in significant improvements in efficiency and data quality while reducing costs.
The setup of a traditional hybridization reaction involves combining the three components (DNA library, blocking oligonucleotide and human Cot-1DNA) and then performing the desired evaporation step (known to those skilled in the art as "Speed-vacuum" (Speed-vac)) to remove all liquid by centrifuging the open tube simultaneously under vacuum and heat. An evaporation step is required, in part because the hybridization buffer must be added at high concentrations, and additional dilution of the first three components is unacceptable. The evaporation step is particularly slow (> 1 hour), cross-contamination is prone to occur and is generally incompatible with high throughput processing.
For a description of the novel hybridization method, see table 1 below.
TABLE 1
The first aspect of the novel method described herein introduces the use of magnetic beads capable of non-specifically binding nucleic acids to rapidly concentrate the nucleic acid material onto the magnetic bead "pellet" which is then resuspended in hybridization buffer at a suitable concentration. This same improved method maintains the same amount of bait in a reduced final volume, which increases reaction kinetics while maintaining the concentration of all components.
A second aspect of the novel methods described herein incorporates biotinylated DNA capture baits into the concentration reaction.
A third aspect of the novel method involves varying the purification conditions to ensure efficient purification of short single-stranded DNA decoys and short single-stranded blocking oligomers (in addition to the longer double-stranded library and Cot-1 DNA). Concentration was performed using AMPureXP beads at a modified ratio of 2 parts beads to 1 part reaction (manufacturer's recommendation of 1.8:1), which resulted in more efficient binding of short nucleic acids (typically removed at a ratio of 1.8: 1).
A fourth aspect of the novel process involves denaturing the reaction mixture in its intact form in which all components are present. Traditional protocols suggest that DNA, Cot-1 and blocking oligos are denatured at 95 ℃ and then the plate is opened or opened to allow additional transfer of biotinylated DNA decoys. The results obtained from the novel process presented herein demonstrate that, in addition to including the bait in the concentration reaction (second aspect above), the entire reaction mixture including the bait can be denatured together at 95 ℃ and then immediately lowered to the hybridization temperature (47 ℃). Surprisingly, this method does not result in a loss of overall blocking efficiency and does not result in a reduction in capture hybridization efficiency. In contrast, this approach results in an increased yield of captured library nucleic acids. This provides a surprising workflow improvement as it eliminates multiple interactions with inefficient thermocyclers and eliminates the high risk of cross-contamination caused by opening patient sample plates.
The fifth aspect of the novel method further improves the capture of library nucleic acids using streptavidin capture beads and biotinylated baits. Traditional protocols require removal of the supernatant from the streptavidin beads, leaving a pellet of beads that needs to be resuspended using a hybridization reaction (10-15 μ l) containing the library nucleic acids. This approach has proven problematic and difficult due to the large number of streptavidin beads that are resuspended in a relatively small volume. To overcome the problems associated with the conventional methods and to enable automation, small particles of streptavidin beads are resuspended in 10. mu.l of hybridization buffer (e.g., under vigorous vortexing and pipetting), and then added to 10-15. mu.l of hybridization reaction. This maintains the buffer composition of the final reaction and enables complete automation.
In summary, the novel methods presented herein are completed in less time than required by conventional protocols (e.g., at least 2 days shorter), require fewer steps and less sample processing, unexpectedly produce higher quality data (percentage of on targets > 90%, relative to 70-80% using conventional protocols), allow for complete and efficient automation, and provide safer patient DNA sample processing. In addition, DNA libraries are not subjected to long periods of heat, such as from speed-vacuum drying and long hybridization times, which can lead to degradation of the library nucleic acids.
Example 2: examples of the embodiments
The following examples illustrate certain embodiments and do not limit the present technology.
A1. A composition for massively parallel nucleic acid sequencing, comprising:
a) a library of nucleic acids comprising a plurality of library inserts obtained from one or more samples and at least eight distinguishable nucleic acid barcodes, each nucleic acid barcode comprising a first end and a second end, and each nucleic acid of the library comprising two of (i) at least one library insert, (ii) a first non-natural nucleic acid, (iii) a second non-natural nucleic acid, and (iv) no more than the distinguishable nucleic acid barcodes, wherein
The first non-natural nucleic acid and the second non-natural nucleic acid are located on opposite sides of at least one library insert, and
the first non-natural nucleic acid and/or the second non-natural nucleic acid comprises no more than two distinguishable nucleic acid barcodes; and
b) no more than four U-block nucleic acids, wherein (i) each of the U-block nucleic acids is substantially complementary to a portion of the first and/or second non-natural nucleic acids, (ii) at least one of the U-block nucleic acids is configured to hybridize adjacent a first end of each of the distinguishable nucleic acid barcodes, (iii) at least one of the U-block nucleic acids is configured to hybridize adjacent a second end of each of the distinguishable nucleic acid barcodes, and (iv) each of the U-block nucleic acids is not substantially hybridized to a portion of at least eight of the distinguishable nucleic acid barcodes.
A2. The composition of embodiment a1, comprising one or more capture nucleic acids, wherein
(i) The capture nucleic acid comprises a member of a binding pair, and
(ii) the capture nucleic acids are each configured to specifically hybridize to a subset of the one or more library inserts.
A3. The composition of embodiment a1 or a2, wherein the one or more samples are obtained from one or more species.
A3.1. The composition of any one of embodiments a1 to A3, comprising four or more samples.
A3.2. The composition of any one of embodiments a1 to A3, comprising eight or more samples.
A4. The composition of any one of embodiments A3 to a3.2, wherein the first and second non-native nucleic acids are not endogenous to the genome of the one or more species.
A5. The composition of any one of embodiments a1 to a4, wherein the one or more samples are obtained from one or more tissues.
A6. The composition of any one of embodiments a1 to a4, wherein the one or more samples are obtained from one or more mammals.
A7. The composition of embodiment a6, wherein the one or more mammals are humans.
A8. The composition of embodiment a6 or a7, wherein the first and second non-native nucleic acids are not endogenous to the genome of the one or more mammals.
A9. The composition of any one of embodiments a1 to A8, comprising ten or more distinguishable nucleic acid barcodes.
A10. The composition of any one of embodiments a1 to a9, wherein the one or more library inserts are obtained from eight or more samples.
A11. The composition of any one of embodiments a1 to a10, wherein each nucleic acid of the library comprises two of the distinguishable nucleic acid barcodes.
A12. The composition of embodiment a11, wherein the first non-natural nucleic acid comprises a first distinguishable nucleic acid barcode and the second non-natural nucleic acid comprises a second distinguishable nucleic acid barcode.
A13. The composition of embodiment a12, wherein the U-block nucleic acids are each configured to block extension by a polymerase.
A14. The composition of any one of embodiments a1 to a13, wherein the first and second non-natural nucleic acids are synthetic nucleic acids.
A15. The composition of any one of embodiments a1 to a14, wherein the first and second non-native nucleic acids are substantially not identical.
A16. The composition of any one of embodiments a1 to a15, wherein the first and second non-native nucleic acids comprise adaptor nucleic acids.
A17. The composition of any one of embodiments a1 to a16, wherein one to four U-block nucleic acids comprise a length of 10 to 40 nucleotides.
A18. The composition of any one of embodiments a1 to a17, wherein no more than four U-block nucleic acids comprise a length of 10 to 30 nucleotides.
A19. The composition of any one of embodiments a1 to a18, wherein no more than four U-block nucleic acids comprise a length of 10 to 20 nucleotides.
A20. The composition of any one of embodiments a1 to a19, wherein no more than four U-block nucleic acids comprise locked nucleic acids.
A21. The composition of any one of embodiments a1 to a20, wherein no more than four U-block nucleic acids comprise a bridging nucleic acid.
A22. The composition of any one of embodiments a1 to a21, wherein no more than four U-block nucleic acids comprise a melting temperature of about 65 ℃ to about 90 ℃.
A23. The composition of any one of embodiments a1 to a22, wherein no more than four U-block nucleic acids comprise a melting temperature of at least about 65 ℃.
A24. The composition of any one of embodiments a1 to a22, wherein no more than four U-block nucleic acids comprise a melting temperature of at least about 75 ℃.
A25. The composition of any one of embodiments a1 to a24, wherein the composition comprises four U-block nucleic acids.
A26. The composition of any one of embodiments a1 to a25, wherein the first non-natural nucleic acid comprises one of at least eight distinguishable nucleic acid barcodes, a portion substantially complementary to a first U-block nucleic acid, and a portion substantially complementary to a second U-block nucleic acid, wherein the first U-block nucleic acid is configured to hybridize adjacent a first end of one distinguishable nucleic acid barcode and the second U-block nucleic acid is configured to hybridize adjacent a second end of one distinguishable nucleic acid barcode.
A27. The composition of any one of embodiments a1 to a26, wherein the second non-natural nucleic acid comprises one of at least eight distinguishable nucleic acid barcodes, a portion substantially complementary to a third U-block nucleic acid, and a portion substantially complementary to a fourth U-block nucleic acid, wherein the third U-block nucleic acid is configured to hybridize adjacent a first end of one distinguishable nucleic acid barcode and the fourth U-block nucleic acid is configured to hybridize adjacent a second end of one distinguishable nucleic acid barcode.
A28. The composition of any one of embodiments a1 to a27, wherein no more than four U-block nucleic acids are substantially non-complementary to at least eight distinguishable nucleic acid barcodes.
A29. The composition of any one of embodiments a1 to a28, comprising a competitor nucleic acid.
A29.1 the composition of embodiment a29, wherein the competitor nucleic acid comprises placental nucleic acid.
A30. The composition of embodiment a29, wherein the competitor nucleic acid comprises a repeat nucleic acid.
A31. The composition of embodiment a30, wherein the repetitive nucleic acid is human.
A32. The composition of embodiment a30, wherein the competitor nucleic acid comprises at least 60% of the repeat nucleic acid.
A33. The composition of embodiment a30, wherein the competitor nucleic acid comprises a synthetic repeat nucleic acid.
A34. The composition of embodiment a29, wherein the competitor nucleic acid comprises a C0t-1 nucleic acid.
A35. The composition of any one of embodiments a2 to a34, wherein the member of the binding pair comprises a biotin, an antigen, a hapten, an antibody or a portion thereof.
A36. The composition of embodiment a35, wherein the member of the binding pair comprises biotin.
A37. The composition of embodiment a35, wherein the member of the binding pair comprises a DNA binding protein recognition sequence or a portion thereof.
A38. The composition of any one of embodiments a1 to a37, wherein no more than four U-block nucleic acids are single stranded.
A39. The composition of any one of embodiments a1 to a38, wherein no more than four U-block nucleic acids comprise a chain terminator.
A40. The composition of any one of embodiments a1 to a39, wherein no more than four U-block nucleic acids comprise an inverted repeat.
A41. The composition of any one of embodiments a1 to a40, wherein the library of nucleic acids comprises single stranded nucleic acids.
A42. The composition of any one of embodiments a1 to a41, wherein the library of nucleic acids comprises amplicons.
A43. The composition of any one of embodiments a1 to a42, wherein the plurality of library inserts comprise genomic nucleic acid.
A44. The composition of any one of embodiments a1 to a43, wherein no more than four U-block nucleic acids comprise degenerate nucleotide bases.
A45. The composition of embodiment a44, wherein the degenerate nucleotide base is 3-nitropyrrole, 5-nitroindole, an analog or derivative thereof.
A46. The composition of embodiment a45, wherein the degenerate nucleotide base is inosine, 2' -deoxyinosine, an analog or derivative thereof.
A47. The composition of any one of embodiments a1 to a46, comprising first, second, third, and fourth U-block nucleic acids, wherein the first and second U-block nucleic acids are substantially complementary to a portion of the first non-natural nucleic acid and the third and fourth U-block nucleic acids are substantially complementary to a portion of the second non-natural nucleic acid.
A48. The composition of any one of embodiments a1 to a47, wherein no more than four U-block nucleic acids comprise substantially different nucleic acid sequences.
B1. A method of analyzing a nucleic acid library, comprising:
a) obtaining a library of nucleic acids comprising a plurality of library inserts obtained from one or more samples and at least eight distinguishable nucleic acid barcodes, each nucleic acid barcode comprising a first end and a second end, and each nucleic acid of the library comprising no more than two of (i) at least one library insert, (ii) a first non-natural nucleic acid, (iii) a second non-natural nucleic acid, and (iv) a distinguishable nucleic acid barcode, wherein
The first non-natural nucleic acid and the second non-natural nucleic acid are located on opposite sides of at least one library insert, and
the first non-natural nucleic acid and/or the second non-natural nucleic acid comprises no more than two distinguishable nucleic acid barcodes;
b) preparing a first mixture comprising contacting a library of nucleic acids with no more than four U-block nucleic acids, wherein each U-block nucleic acid is substantially complementary to a portion of the first and/or second non-natural nucleic acids, at least one of the U-block nucleic acids configured to hybridize adjacent a first end of each of the distinguishable nucleic acid barcodes and at least one of the U-block nucleic acids configured to hybridize adjacent a second end of each of the distinguishable nucleic acid barcodes;
c) preparing a second mixture comprising contacting the first mixture with one or more capture nucleic acids, wherein
(i) The capture nucleic acid comprises a first member of a binding pair, and
(ii) the capture nucleic acids are each configured to specifically hybridize to a subset of the one or more library inserts;
d) contacting the second mixture with a second member of the binding pair, thereby providing an isolated nucleic acid;
e) contacting the isolated nucleic acid with a primer set under amplification conditions, thereby providing an amplicon; and
f) the amplicons were analyzed.
B2. The method of embodiment B1, wherein the one or more samples are obtained from one or more species.
B3. The method of embodiment B2, wherein the library insert is obtained from four or more samples.
B4. The method of embodiment B2, wherein the library inserts are obtained from eight or more samples.
B5. The method of any one of embodiments B1 to B4, wherein the one or more samples are obtained from one or more tissues.
B6. The method of any one of embodiments B1 to B5, wherein the one or more samples are obtained from one or more mammals.
B7. The method of embodiment B6, wherein the one or more mammals are humans.
B8. The method of embodiment B6 or B7, wherein the first and second non-native nucleic acids are not endogenous to the genome of the one or more mammals.
B9. The method of any one of embodiments B1 to B8, the library of inserts comprising ten or more distinguishable nucleic acid barcodes.
B10. The method of any one of embodiments B1 to B9, wherein the first non-natural nucleic acid comprises a first distinguishable nucleic acid barcode and the second non-natural nucleic acid comprises a second distinguishable nucleic acid barcode.
B11. The method of any one of embodiments B1 to B10, wherein each nucleic acid of the library comprises two of the distinguishable nucleic acid barcodes.
B12. The method of any one of embodiments B1 to B11, wherein the U-block nucleic acids are each configured to block extension by a polymerase.
B13. The method of any one of embodiments B1 to B12, wherein the first and second non-natural nucleic acids are synthetic nucleic acids.
B14. The method of any one of embodiments B1 to B13, wherein the first and second non-native nucleic acids are substantially not identical.
B15. The method of any one of embodiments B1 to B14, wherein the first and second non-native nucleic acids comprise adaptor nucleic acids.
B16. The method of any one of embodiments B1 to B15, wherein the one to four U-block nucleic acids comprise a length of 10 to 40 nucleotides.
B17. The method of any one of embodiments B1 to B16, wherein no more than four U-block nucleic acids comprise a length of 10 to 30 nucleotides.
B18. The method of any one of embodiments B1 to B17, wherein no more than four U-block nucleic acids comprise a length of 10 to 20 nucleotides.
B19. The method of any one of embodiments B1 to B18, wherein no more than four U-block nucleic acids comprise locked nucleic acids.
B20. The method of any one of embodiments B1 to B19, wherein no more than four U-block nucleic acids comprise a bridging nucleic acid.
B21. The method of any one of embodiments B1 to B20, wherein no more than four U-block nucleic acids comprise a melting temperature of at least about 65 ℃.
B22. The method of any one of embodiments B1 to B21, wherein no more than four U-block nucleic acids comprise a melting temperature of at least about 75 ℃.
B23. The method of any one of embodiments B1 to B22, wherein no more than four U-block nucleic acids comprise a melting temperature of about 65 ℃ to about 90 ℃.
B24. The method of any one of embodiments B1 to B23, wherein the first non-natural nucleic acid comprises one of at least eight distinguishable nucleic acid barcodes, a portion substantially complementary to a first U-block nucleic acid, and a portion substantially complementary to a second U-block nucleic acid, wherein the first U-block nucleic acid is configured to hybridize adjacent a first end of one distinguishable nucleic acid barcode and the second U-block nucleic acid is configured to hybridize adjacent a second end of one distinguishable nucleic acid barcode.
B25. The method of any one of embodiments B1 to B24, wherein the second non-natural nucleic acid comprises one of at least eight distinguishable nucleic acid barcodes, a portion substantially complementary to a third U-block nucleic acid, and a portion substantially complementary to a fourth U-block nucleic acid, wherein the third U-block nucleic acid is configured to hybridize adjacent a first end of one distinguishable nucleic acid barcode and the fourth U-block nucleic acid is configured to hybridize adjacent a second end of one distinguishable nucleic acid barcode.
B26. The method of any one of embodiments B1 to B25, wherein no more than four U-block nucleic acids are substantially non-complementary to at least eight distinguishable nucleic acid barcodes.
B27. The method of embodiment B26, wherein no more than four U-block nucleic acids do not substantially hybridize to at least eight distinguishable nucleic acid barcode portions.
B28. The method of any one of embodiments B1 to B27, comprising contacting the first mixture with a competitor nucleic acid prior to c).
B28.1 the method of embodiment B28, wherein the competitor nucleic acid comprises placental nucleic acid.
B29. The method of embodiment B28 or B28.1, wherein the competitor nucleic acid comprises a repeat nucleic acid.
B30. The method of embodiment B29, wherein the repetitive nucleic acid is human.
B31. The method of embodiment B29, wherein the competitor nucleic acid comprises at least 60% of the repeat nucleic acid.
B32. The method of embodiment B1, wherein the competitor nucleic acid comprises a synthetic nucleic acid.
B33. The method of embodiment B1, wherein the competitor nucleic acid comprises a C0t-1 nucleic acid.
B34. The method of embodiment B1, wherein the first member of the binding pair comprises biotin, an antigen, a hapten, an antibody or a portion thereof.
B35. The method of embodiment B34, wherein the first member of the binding pair comprises biotin.
B36. The method of embodiment B1, wherein the first member of the binding pair comprises a DNA binding protein recognition sequence or a portion thereof.
B37. The method of any one of embodiments B1 to B36, wherein no more than four U-block nucleic acids are single stranded.
B38. The method of any one of embodiments B1 to B37, wherein no more than four U-block nucleic acids comprise a chain terminator.
B39. The method of any one of embodiments B1 to B38, wherein no more than four U-block nucleic acids comprise an inverted repeat.
B40. The method of any one of embodiments B1 to B39, wherein the library of nucleic acids comprises single-stranded nucleic acids.
B41. The method of any one of embodiments B1 to B40, wherein the library of nucleic acids comprises amplicons.
B42. The method of any one of embodiments B1 to B41, wherein the plurality of library inserts comprise genomic nucleic acid.
B43. The method of any one of embodiments B1 to B42, wherein no more than four U-block nucleic acids comprise degenerate nucleotide bases.
B44. The method of embodiment B43, wherein the degenerate nucleotide base is 3-nitropyrrole, 5-nitroindole, an analog or derivative thereof.
B45. The method of embodiment B43, wherein the degenerate nucleotide base is inosine, 2' -deoxyinosine, an analog or derivative thereof.
B46. The method of embodiment B1, wherein the first member of the binding pair comprises biotin, an antigen, a hapten, an antibody or a portion thereof.
B47. The method of embodiment B1, wherein the first member of the binding pair comprises biotin.
B48. The method of embodiment B1, wherein the first member of the binding pair comprises a DNA binding protein recognition sequence or a portion thereof.
B49. The method of any one of embodiments B1 to B48, comprising hybridizing the isolated nucleic acid under hybridizing conditions.
B50. The method of embodiment B1, wherein the amplification conditions comprise a thermostable polymerase.
B51. The method of embodiment B1, wherein the amplification conditions comprise polymerase chain reaction.
B52. The method of embodiment B1, wherein the capture nucleic acid is configured to specifically hybridize to a portion of an exon.
B53. The method of embodiment B1, wherein the capture nucleic acid is configured to specifically hybridize to a portion of a chromosome.
B54. The method of embodiment B1, wherein the second member of the binding pair comprises avidin, protein a, protein G, an antibody or binding portion thereof.
B55. The method of embodiment B1, wherein the second member of the binding pair comprises avidin, or a portion thereof.
B56. The method of embodiment B1, wherein the second member of the binding pair comprises a matrix.
B57. The method of embodiment B56, wherein the matrix comprises a magnetic compound.
B58. The method of embodiment B56, wherein the matrix comprises beads.
B59. The method of embodiment B56, wherein the matrix comprises polystyrene, polycarbonate, or agarose.
B60. The method of embodiment B56, wherein the substrate comprises magnetic beads.
B61. The method of embodiment B1, wherein the contacting in (d) comprises centrifugation.
B62. The method of embodiment B1, wherein the contacting in (d) comprises using a magnet.
B63. The method of embodiment B1, wherein analyzing comprises providing sequence reads.
B64. The method of embodiment B63, wherein the sequence reads are obtained by a method comprising massively parallel sequencing.
B65. The method of embodiment B63, wherein the sequence reads are obtained by a method comprising paired-end sequencing.
B66. The method of any one of embodiments B1 to B65, wherein the no more than four U-block nucleic acids comprise a first, second, third and fourth U-block nucleic acid, wherein the first and second U-block nucleic acids are substantially complementary to a portion of the first non-natural nucleic acid and the third and fourth U-block nucleic acids are substantially complementary to a portion of the second non-natural nucleic acid.
B67. The method of any one of embodiments B1 to B66, wherein no more than four U-block nucleic acids comprise substantially different nucleic acid sequences.
C1. A method of analyzing a nucleic acid library comprising
a) Obtaining a library of nucleic acids comprising a first set of amplicons, wherein each amplicon comprises a first non-natural nucleic acid and a second non-natural nucleic acid, one or more distinguishable identifiers, and a library insert obtained from one of the one or more samples, wherein the library insert is located between the first and second non-natural nucleic acids;
b) preparing a mixture comprising contacting a library of nucleic acids with one or more blocking nucleic acids and a capture nucleic acid, wherein
(i) One or more blocking nucleic acids are configured to specifically hybridize to portions of the first and second non-native nucleic acids,
(ii) the capture nucleic acid comprises a first member of a binding pair, and
(iii) the capture nucleic acid is configured to specifically hybridize to a subset of the first set of amplicons;
c) purifying the mixture, thereby providing a purified nucleic acid, wherein the purified nucleic acid comprises a library of nucleic acids, one or more blocking nucleic acids, and a capture nucleic acid;
d) hybridizing the purified nucleic acid under hybridization conditions;
e) capturing the captured nucleic acid, thereby providing a captured nucleic acid;
f) contacting the captured nucleic acids with a set of primers under amplification conditions, thereby providing a second set of amplicons; and
g) the second set of amplicons is analyzed.
C2. The method of embodiment C1, wherein the one or more samples are obtained from a human.
C3. The method of embodiment C2, wherein the first nucleic acid and the second nucleic acid are not endogenous to a human.
C4. The method of any one of embodiments C1 to C3, wherein the one or more samples are obtained from a tissue selected from the group consisting of breast tissue, colon tissue, pancreatic tissue, placenta, or epithelial tissue.
C5. The method of any one of embodiments C1 to C3, wherein the one or more samples are obtained from blood.
C6. The method of embodiment C5, wherein the one or more samples are obtained from circulating blood cells.
C7. The method of any one of embodiments C2 to C6, wherein the human is a fetus.
C8. The method of embodiment C5, wherein the one or more samples comprise circulating cell-free nucleic acid.
C9. The method of any one of embodiments C1 to C8, wherein the amplification conditions comprise a thermostable polymerase.
C10. The method of any one of embodiments C1 to C9, wherein the amplification conditions comprise polymerase chain reaction.
C11. The method of any one of embodiments C1 to C10, wherein the preparing in (b) comprises contacting the nucleic acids of the library with competitor nucleic acids.
C12. The method of embodiment C11, wherein the competitor nucleic acid comprises placental nucleic acid.
C13. The method of embodiment C11, wherein the competitor nucleic acid comprises a repeat nucleic acid.
C14. The method of embodiment C13, wherein the repeated nucleic acids are derived from a human.
C15. The method of embodiment C13, wherein the competitor nucleic acid comprises at least 60% or more of the repeat nucleic acid.
C16. The method of embodiment C11, wherein the competitor nucleic acid comprises a synthetic nucleic acid.
C17. The method of embodiment C11, wherein the competitor nucleic acid comprises a C0t-1 nucleic acid.
C18. The method of any one of embodiments C1 to C17, wherein the capture nucleic acid is configured to specifically hybridize to a portion of the library insert.
C19. The method of any one of embodiments C1 to C18, wherein the one or more blocking nucleic acids are configured to specifically hybridize to a portion of the first non-natural nucleic acid and/or the second non-natural nucleic acid.
C20. The method of any one of embodiments C1 to C18, wherein the one or more blocking nucleic acids are configured to prevent extension of the blocking nucleic acid by a polymerase.
C21. The method of any one of embodiments C1 to C20, wherein the one or more blocking nucleic acids comprise a chain terminator.
C22. The method of any one of embodiments C1 to C21, wherein the one or more blocking nucleic acids comprise an inverted repeat.
C23. The method of any one of embodiments C1 to C22, wherein the capture nucleic acid is configured to specifically hybridize to a portion of an exon.
C24. The method of any one of embodiments C1 to C23, wherein the capture nucleic acid is configured to specifically hybridize to a portion of a chromosome.
C25. The method of embodiment C24, wherein the capture nucleic acid is configured to specifically hybridize to a portion of the library insert comprising the genetic variation.
C26. The method of any one of embodiments C1 to C26, wherein the first member of the binding pair comprises biotin, an antigen, a hapten, an antibody or a portion thereof.
C27. The method of embodiment C26, wherein the first member of the binding pair comprises biotin.
C28. The method of any one of embodiments C1 to C27, wherein the first member of the binding pair comprises a CNC binding protein recognition sequence or portion thereof.
C29. The method of any one of embodiments C1 to C28, wherein the capturing in (e) comprises contacting the mixture with a second member of the binding pair.
C30. The method of embodiment C29, wherein the second member of the binding pair comprises avidin, protein a, protein G, an antibody or binding portion thereof.
C31. The method of embodiment C30, wherein the second member of the binding pair comprises a biotin protein, or portion thereof.
C32. The method of any one of embodiments C29 to C31, wherein the second member of the binding pair comprises a matrix.
C33. The method of embodiment C32, wherein the matrix comprises a magnetic compound.
C34. The method of embodiment C32, wherein the matrix comprises beads.
C35. The method of embodiment C32, wherein the matrix comprises polystyrene, polycarbonate, or agarose.
C36. The method of embodiment C32, wherein the substrate comprises a metal.
C37. The method of any one of embodiments C1 to C36, wherein the capturing in (e) comprises recovering the captured nucleic acids by a method comprising centrifugation.
C38. The method of any one of embodiments C1 to C37, wherein the capturing in (e) comprises recovering the captured nucleic acids by a method comprising the use of a magnet.
C39. The method of any one of embodiments C1 to C38, wherein the hybridization conditions comprise denaturation.
C40. The method of any one of embodiments C1 to C39, wherein the hybridizing in (d) comprises hybridizing the captured nucleic acids to portions of one or more of the first set of amplicons.
C41. The method of any one of embodiments C1 to C40, wherein the hybridization conditions comprise incubating the captured nucleic acids at a temperature of about 25 ℃ to about 70 ℃.
C42. The method of embodiment C41, wherein the incubating is at a temperature of about 35 ℃ to about 60 ℃.
C43. The method of any one of embodiments C41 to C42, wherein the incubation is for an amount of time of about 1 hour to about 24 hours.
C44. The method of any one of embodiments C41 to C43, wherein the incubation is for an amount of time of about 12 hours to about 20 hours.
C45. The method of any one of embodiments C1 to C44, wherein the hybridizing in (d) comprises contacting the mixture with a hybridization buffer.
C46. The method of any one of embodiments C1 to C45, wherein the hybridizing in (d) comprises sequential steps of (i) contacting the mixture with a hybridization buffer, (ii) denaturing, and (iii) hybridizing.
C47. The method of any one of embodiments C1 to C46, wherein the method does not comprise a drying step.
C48. The method of any one of embodiments C1 to C47, wherein the hybridization conditions do not include a polymerase.
C49. The method of any one of embodiments C1 to C48, wherein analyzing comprises providing sequence reads.
C50. The method of embodiment C49, wherein the sequence reads are obtained by a method comprising massively parallel sequencing.
C51. The method of embodiment C49, wherein the sequence reads are obtained by a method comprising paired-end sequencing.
C52. The method of any one of embodiments C1 to C51, wherein the first non-native nucleic acid comprises one or more distinguishable identifiers.
C53. The method of any one of embodiments C1 to C52, wherein the second non-native nucleic acid comprises one or more distinguishable identifiers.
C54. The method of any one of embodiments C1 to C53, wherein the one or more distinguishable identifiers comprise a nucleic acid barcode.
C55. The method of any one of embodiments C1 to C54, wherein the one or more blocking nucleic acids comprise locked nucleic acids.
C56. The method of any one of embodiments C1 to C55, wherein the one or more blocking nucleic acids comprise a bridging nucleic acid.
C57. The method of any one of embodiments C1 to C56, wherein the method does not include a denaturation step prior to (C).
C58. The method of any one of embodiments C1 to C57, wherein the method does not include a denaturation step prior to (d).
C59. The process of any one of embodiments C1 to C58, wherein the process does not include heating to a temperature greater than 80 ℃ prior to (d).
C60. The process of any one of embodiments C1 to C59, wherein the process does not include heating to a temperature greater than 90 ℃ prior to (d).
C61. The method of any one of embodiments C1 to C60, wherein the captured nucleic acids comprise a subset of the nucleic acids of the library.
C62. The method of any one of embodiments C1 to C61, wherein the one or more samples comprise 10 or more samples.
C63. The method of any one of embodiments C1 to C62, wherein the one or more distinguishable identifiers consists of 10 or more distinguishable identifiers.
C64. The method of any one of embodiments C1 to C63, wherein the first nucleic acid and the second nucleic acid comprise synthetic nucleic acids.
C65. The method of any one of embodiments C1 to C64, wherein the library insert comprises a portion of genomic nucleic acid.
C66. The method of any one of embodiments C1 to C65, wherein the purifying in (C) does not comprise adding a second member of the binding pair configured to bind to the first member of the binding pair.
C67. The method of embodiment C1, wherein the purifying in (C) comprises a method of non-specifically binding nucleic acids to a matrix.
C68. The method of any one of embodiments C1 to C67, wherein the purifying in (C) comprises using an anion exchange resin.
C69. The method of any one of embodiments C1 to C68, wherein the capturing in (C) comprises adding a second member of a binding pair configured to bind to a first member of the binding pair.
C70. The method of any one of embodiments C1 to C69, wherein prior to (e), the mixture is not immobilized on a substrate of a flow cell or array.
D1. A method of analyzing a genomic DNA library, comprising:
a) obtaining a genomic DNA library comprising a first set of single-stranded amplicons, wherein each amplicon comprises a first non-natural nucleic acid and a second non-natural nucleic acid, one or two nucleic acid barcodes, and a library insert obtained from the genome of one of ten or more personal subjects, wherein the library insert is located between the first and second non-natural nucleic acids, and wherein the first set of amplicons comprises a plurality of library inserts from ten or more personal subjects;
b) preparing a mixture comprising contacting the first set of amplicons with one to four blocking nucleic acids, C0t-1DNA, and a capture nucleic acid, wherein
(i) One to four blocking nucleic acids configured to specifically hybridize to the first and/or second non-native nucleic acid,
(ii) one to four blocking nucleic acids comprise locked nucleic acids and comprise a length of 10 to 30 nucleotides,
(iii) the capture nucleic acid comprises biotin, and
(iv) the capture nucleic acid is configured to specifically hybridize to a subset of the plurality of library inserts;
c) contacting the mixture with magnetic beads comprising a non-specific nucleic acid binding matrix, thereby providing purified nucleic acids, wherein the purified nucleic acids comprise a first set of amplicons, one to four blocker nucleic acids, C0t-1DNA, and capture nucleic acids;
d) hybridizing the purified nucleic acid, wherein hybridizing comprises the sequential steps of (i) contacting the mixture with a hybridization buffer, (ii) heating the purified nucleic acid to at least 95 ℃ for about 10 minutes, and (iii) hybridizing the purified nucleic acid by incubating at about 40 ℃ to about 50 ℃ for 12 to 20 hours;
e) capturing the captured nucleic acid, wherein capturing comprises contacting the purified nucleic acid with avidin-coated magnetic beads configured to specifically bind to the captured nucleic acid and immobilizing the captured nucleic acid using a magnet, thereby providing captured nucleic acid;
f) contacting the captured nucleic acids with a set of primers under amplification conditions, thereby providing a second set of amplicons; and
g) obtaining sequence reads from the second set of amplicons by a method comprising paired-end sequencing, wherein the method does not comprise a drying step.
Each patent, patent application, publication, and document cited herein is hereby incorporated by reference in its entirety. Citation of the above patents, patent applications, publications and documents is not an admission that any of the foregoing is pertinent prior art, nor does it constitute any admission as to the contents or date of such publications or documents.
Modifications may be made to the foregoing without departing from the basic aspects of the present technology. While the technology has been described in detail with reference to one or more specific embodiments, those skilled in the art will recognize that changes may be made to the embodiments specifically disclosed in this application, but that such modifications and improvements are within the scope and spirit of the technology.
The techniques illustratively described herein suitably may be practiced in the absence of any element which is not specifically disclosed herein. Thus, in each instance herein, any of the terms "comprising," "consisting essentially of … …," and "consisting of … …" can be substituted with either of the other two terms. The terms and expressions which have been employed are used as terms of description and not of limitation, and there is no intention in the use of such terms and expressions of excluding any equivalents of the features shown and described or portions thereof, and various modifications are possible within the scope of the technology claimed. The terms "a" or "an" may refer to one or more of the elements that it modifies (e.g., "the agent" may refer to one or more agents), unless the context clearly dictates otherwise, any one of the elements or more than one of the elements. As used herein, the term "about" refers to a value within 10% of the underlying parameter (i.e., plus or minus 10%), and the use of the term "about" at the beginning of a string of values modifies each value (e.g., "about 1, 2, and 3" refers to about 1, about 2, and about 3). For example, a weight of "about 100 grams" may include a weight between 90 grams and 110 grams. Further, when a list of values is described herein (e.g., about 50%, 60%, 70%, 80%, 85%, or 86%), the list includes all intermediate and fractional values thereof (e.g., 54%, 85.4%). Thus, it should be understood that although the present technology has been specifically disclosed by representative embodiments and optional features, modification and variation of the concepts herein disclosed may be resorted to by those skilled in the art, and that such modifications and variations are considered to be within the scope of this technology.
Certain embodiments of the present technology are set forth in the following claims.

Claims (111)

1. A composition for massively parallel nucleic acid sequencing, comprising:
a) a library of nucleic acids comprising a plurality of library inserts, wherein each nucleic acid of the library comprises (i) at least one library insert obtained from one of four or more samples, (ii) a first non-natural nucleic acid, and (iii) a second non-natural nucleic acid, wherein the first non-natural nucleic acid and the second non-natural nucleic acid are located on opposite sides of the at least one library insert, and the first non-natural nucleic acid comprises a first distinguishable nucleic acid barcode and the second non-natural nucleic acid comprises a second distinguishable nucleic acid barcode, wherein the first and second distinguishable nucleic acid barcodes are unique to one of the four or more samples; and
b) four U-block nucleic acids, wherein (i) first and second U-block nucleic acids are configured to hybridize to the first non-natural nucleic acid on opposite sides of the first distinguishable nucleic acid barcode, and (ii) third and fourth U-block nucleic acids are configured to hybridize to the second non-natural nucleic acid on opposite sides of the second distinguishable nucleic acid barcode, and (iii) each of the U-block nucleic acids does not substantially hybridize to a portion of the first or second distinguishable nucleic acid barcode.
2. The composition of claim 1, wherein the library of nucleic acids comprises at least eight distinguishable nucleic acid barcodes.
3. The composition of claim 2, wherein each of the at least eight distinguishable nucleic acid barcodes is present on a different nucleic acid of the library.
4. The composition of any one of claims 1 to 3, wherein the first and second U-block nucleic acids are substantially complementary to a portion of the first non-natural nucleic acid and the third and fourth U-block nucleic acids are substantially complementary to a portion of the second non-natural nucleic acid.
5. The composition of any one of claims 1 to 4, comprising no more than four U-block nucleic acids.
6. The composition of any one of claims 1 to 5, wherein the four U-block nucleic acids comprise substantially identical nucleic acid sequences.
7. The composition of any one of claims 1 to 5, wherein the four U-block nucleic acids comprise substantially different nucleic acid sequences.
8. The composition of any one of claims 1 to 7, comprising one or more capture nucleic acids, wherein
(i) The capture nucleic acid comprises a member of a binding pair; and
(ii) the capture nucleic acids are each configured to specifically hybridize to a subset of the nucleic acids of the library.
9. The composition of claim 8, wherein the capture nucleic acids are each configured to specifically hybridize to a subset of the plurality of inserts.
10. The composition of any one of claims 1 to 9, wherein the four or more samples are obtained from one or more species.
11. The composition of any one of claims 1 to 10, wherein the four or more samples comprise eight or more samples.
12. The composition of claim 10 or 11, wherein the first and second non-native nucleic acids are not endogenous to the genome of the one or more species.
13. The composition of any one of claims 10 to 12, wherein the first and second non-native nucleic acids are not present in the genome of the one or more species.
14. The composition of any one of claims 1 to 13, wherein the four or more samples are obtained from four or more tissues.
15. The composition of any one of claims 1 to 14, wherein the four or more samples are obtained from four or more different mammals.
16. The composition of claim 15, wherein the one or more mammals are humans.
17. The composition of claim 15 or 16, wherein the first and second non-native nucleic acids are not endogenous to the genomes of the four or more different mammals.
18. The composition of any one of claims 1 to 17, wherein the library of nucleic acids comprises ten or more distinguishable nucleic acid barcodes.
19. The composition of any one of claims 1 to 18, wherein the first and second distinguishable nucleic acid barcodes have 100% identical nucleic acid sequences.
20. The composition of any one of claims 1 to 19, wherein the first and second distinguishable nucleic acid barcodes have different and distinguishable nucleic acid sequences.
21. The composition of any one of claims 1 to 20, wherein the U-block nucleic acids are each configured to block extension by a polymerase.
22. The composition of any one of claims 1 to 21, wherein the first and second non-natural nucleic acids are synthetic nucleic acids.
23. The composition of any one of claims 1 to 22, wherein the first and second non-native nucleic acids are substantially non-identical.
24. The composition of any one of claims 1 to 23, wherein the first and second non-native nucleic acids comprise adaptor nucleic acids.
25. The composition of any one of claims 1 to 24, wherein the four U-block nucleic acids each comprise a length of 10 to 40 nucleotides.
26. The composition of any one of claims 1 to 25, wherein each of the four U-block nucleic acids comprises a length of 10 to 30 nucleotides.
27. The composition of any one of claims 1 to 26, wherein the four U-block nucleic acids each comprise a length of 10 to 20 nucleotides.
28. The composition of any one of claims 1 to 27, wherein each of the four U-block nucleic acids comprises a locked nucleic acid.
29. The composition of any one of claims 1 to 28, wherein each of the four U-block nucleic acids comprises a bridging nucleic acid.
30. The composition of any one of claims 1 to 29, wherein each of the four U-block nucleic acids comprises a melting temperature of about 65 ℃ to about 90 ℃.
31. The composition of any one of claims 1 to 30, wherein each of the four U-block nucleic acids comprises a melting temperature of at least 65 ℃.
32. The composition of any one of claims 1 to 31, wherein each of the four U-block nucleic acids comprises a melting temperature of at least 75 ℃.
33. The composition of any one of claims 1 to 32, wherein the first U-block nucleic acid and the second U-block nucleic acid are substantially identical and different from the third U-block nucleic acid and the fourth U-block nucleic acid.
34. The composition of any one of claims 1 to 32, wherein the first, second, third and fourth U-block nucleic acids are substantially identical.
35. The composition of any one of claims 1 to 34, wherein the first, second, third and fourth U-block nucleic acids are substantially different.
36. The composition of any one of claims 2 to 35, wherein the four U-block nucleic acids are not substantially complementary to the at least eight distinguishable nucleic acid barcodes.
37. The composition of any one of claims 1 to 36, wherein the four U-block nucleic acids do not substantially hybridize to a distinguishable nucleic acid barcode.
38. The composition of any one of claims 1 to 37, wherein the four U-block nucleic acids do not hybridize to any portion of a distinguishable nucleic acid barcode.
39. The composition of any one of claims 1 to 38, comprising a competitor nucleic acid.
40. The composition of any one of claims 8 to 39, wherein the member of the binding pair comprises biotin, an antigen, a hapten, an antibody or a portion thereof.
41. The composition of claim 40, wherein the member of the binding pair comprises biotin.
42. The composition of any one of claims 1 to 41, wherein the four U-block nucleic acids are single stranded.
43. The composition of any one of claims 1 to 42, wherein the four U-block nucleic acids comprise a chain terminator.
44. The composition of any one of claims 1 to 43, wherein the four U-block nucleic acids comprise inverted repeats.
45. The composition of any one of claims 1 to 44, wherein the library of nucleic acids comprises single stranded nucleic acids.
46. The composition of any one of claims 1 to 45, wherein the library of nucleic acids comprises amplicons.
47. The composition of any one of claims 1 to 46, wherein the plurality of library inserts comprise genomic nucleic acid.
48. The composition of any one of claims 1 to 47, wherein the four U-block nucleic acids do not comprise degenerate nucleotide bases.
49. The composition of claim 48, wherein the degenerate nucleotide base comprises a 3-nitropyrrole, a 5-nitroindole, an analog or derivative thereof.
50. The composition of claim 48, wherein the degenerate nucleotide base comprises inosine, 2' -deoxyinosine, an analog or derivative thereof.
51. A method of analyzing a nucleic acid library, comprising:
a) obtaining a library of nucleic acids comprising a plurality of library inserts, wherein each nucleic acid of the library comprises (i) at least one library insert obtained from one of four or more samples, (ii) a first non-natural nucleic acid, and (iii) a second non-natural nucleic acid, wherein the first non-natural nucleic acid and the second non-natural nucleic acid are located on opposite sides of the at least one library insert, and the first non-natural nucleic acid comprises a first distinguishable nucleic acid barcode and the second non-natural nucleic acid comprises a second distinguishable nucleic acid barcode, wherein the first and second distinguishable nucleic acid barcodes are unique to one of the four or more samples;
b) contacting the library of nucleic acids with four U-block nucleic acids, wherein (i) first and second U-block nucleic acids are configured to hybridize to the first non-natural nucleic acid on opposite sides of the first distinguishable nucleic acid barcode, and (ii) third and fourth U-block nucleic acids are configured to hybridize to the second non-natural nucleic acid on opposite sides of the second distinguishable nucleic acid barcode, and (iii) each of the U-block nucleic acids is not substantially hybridized to a portion of the first or second distinguishable nucleic acid barcode; and
c) contacting the library of nucleic acids with one or more capture nucleic acids each comprising a first member of a binding pair, wherein the one or more capture nucleic acids are configured to specifically hybridize to a subset of the nucleic acids of the library;
d) capturing said captured nucleic acids, thereby providing captured nucleic acids comprising a subset of the nucleic acids of the library;
e) contacting the captured nucleic acids with a primer set under amplification conditions, thereby providing amplicons; and
f) analyzing the amplicons.
52. The method of claim 51, wherein the amplification conditions comprise a thermostable polymerase.
53. The method of claim 52, wherein the amplification conditions comprise polymerase chain reaction.
54. The method of any one of claims 51 to 53, comprising contacting the nucleic acids of the library with competitor nucleic acids.
55. The method of any one of claims 51 to 54, wherein the first member of the binding pair comprises biotin, an antigen, a hapten, an antibody or a portion thereof.
56. The method of any one of claims 51 to 55, wherein the capturing in (d) comprises contacting the mixture with a second member of a binding pair.
57. The method of claim 56, wherein the second member of the binding pair comprises avidin, protein A, protein G, an antibody, or a binding portion thereof.
58. The method of claim 56, wherein the second member of the binding pair comprises avidin, or a portion thereof.
59. The method of any one of claims 56 to 58, wherein the second member of the binding pair comprises a matrix.
60. The method of claim 59, wherein the substrate comprises a magnetic compound.
61. The method of claim 59, wherein the substrate comprises a bead.
62. The method of claim 59, wherein the matrix comprises polystyrene, polycarbonate, or agarose.
63. The method of claim 59, wherein the substrate comprises a metal.
64. The method of any one of claims 51 to 63, wherein the analyzing comprises providing sequence reads.
65. The method of claim 64, wherein the sequence reads are obtained by a method comprising massively parallel sequencing.
66. The method of claim 65, wherein the sequence reads are obtained by a method comprising paired-end sequencing.
67. The method of any one of claims 51-66, wherein the library of nucleic acids comprises at least eight distinguishable nucleic acid barcodes.
68. The method of claim 67, wherein each of the at least eight distinguishable nucleic acid barcodes is present on a different nucleic acid of the library.
69. The method of any one of claims 51 to 68, wherein the first and second U-block nucleic acids are substantially complementary to a portion of the first non-natural nucleic acid and the third and fourth U-block nucleic acids are substantially complementary to a portion of the second non-natural nucleic acid.
70. The method of any one of claims 51 to 69, wherein the library is contacted with no more than four U-bolck nucleic acids.
71. The method of any one of claims 51-70, wherein the four U-block nucleic acids comprise substantially the same nucleic acid sequence.
72. The method of any one of claims 51-70, wherein the four U-block nucleic acids comprise substantially different nucleic acid sequences.
73. The method of any one of claims 51-72, wherein the capture nucleic acid is configured to specifically hybridize to a subset of the plurality of inserts.
74. The method of any one of claims 51 to 73, wherein the four or more samples are obtained from one or more species.
75. The method of any one of claims 51-74, wherein the four or more samples comprise eight or more samples.
76. The method of claims 74-75, wherein the first and second non-native nucleic acids are not endogenous to the genome of the one or more species.
77. The method of any one of claims 74-76, wherein the first and second non-native nucleic acids are not present in the genome of the one or more species.
78. The method of any one of claims 51-77, wherein the four or more samples are obtained from four or more tissues.
79. The method of any one of claims 51-79, wherein the four or more samples are obtained from four or more different mammals.
80. The method of claim 79, wherein the one or more mammals are humans.
81. The method of claim 79 or 80, wherein the first and second non-native nucleic acids are not endogenous to the genomes of the four or more different mammals.
82. The method of any one of claims 51-81, wherein the library of nucleic acids comprises ten or more distinguishable nucleic acid barcodes.
83. The method of any one of claims 51 to 82, wherein the first and second distinguishable nucleic acid barcodes have nucleic acid sequences that are 100% identical.
84. The method of any one of claims 51-82, wherein the first and second distinguishable nucleic acid barcodes have different and distinguishable nucleic acid sequences.
85. The method of any one of claims 51-84, wherein the U-block nucleic acids are each configured to block extension by a polymerase.
86. The method of any one of claims 51-85, wherein the first and second non-natural nucleic acids are synthetic nucleic acids.
87. The method of any one of claims 51-86, wherein the first and second non-native nucleic acids are substantially non-identical.
88. The method of any one of claims 51-87, wherein the first and second non-native nucleic acids comprise adaptor nucleic acids.
89. The method of any one of claims 51-88, wherein the four U-block nucleic acids each comprise a length of 10 to 40 nucleotides.
90. The method of any one of claims 51-89, wherein the four U-block nucleic acids each comprise a length of 10 to 30 nucleotides.
91. The method of any one of claims 51 to 90, wherein the four U-block nucleic acids each comprise a length of 10 to 20 nucleotides.
92. The method of any one of claims 51-91, wherein each of the four U-block nucleic acids comprises a locked nucleic acid.
93. The method of any one of claims 51-92, wherein each of the four U-block nucleic acids comprises a bridging nucleic acid.
94. The method of any one of claims 51-93, wherein each of the four U-block nucleic acids comprises a melting temperature of about 65 ℃ to about 90 ℃.
95. The method of any one of claims 51-93, wherein each of the four U-block nucleic acids comprises a melting temperature of at least 65 ℃.
96. The method of any one of claims 51-93, wherein each of the four U-block nucleic acids comprises a melting temperature of at least 75 ℃.
97. The method of any one of claims 51 to 94, wherein the first U-block nucleic acid and the second U-block nucleic acid are substantially identical and are different from the third and fourth U-block nucleic acids.
98. The method of any one of claims 51-96, wherein the first, second, third and fourth U-block nucleic acids are substantially identical.
99. The method of any one of claims 51-96, wherein the first, second, third, and fourth U-block nucleic acids are substantially different.
100. The method of any one of claims 67 to 99, wherein the four U-block nucleic acids are not substantially complementary to the at least eight distinguishable nucleic acid barcodes.
101. The method of any one of claims 51 to 100, wherein the four U-block nucleic acids do not substantially hybridize to a distinguishable nucleic acid barcode.
102. The method of any one of claims 51 to 101, wherein the four U-block nucleic acids do not hybridize to any portion of a distinguishable nucleic acid barcode.
103. The method of any one of claims 51 to 102, comprising competitor nucleic acids.
104. The method of any one of claims 51 to 103, wherein the four U-block nucleic acids are single stranded.
105. The method of any one of claims 51-104, wherein at least one of the four U-block nucleic acids comprises a chain terminator.
106. The method of any one of claims 51-105, wherein at least one of the four U-block nucleic acids comprises an inverted repeat.
107. The method of any one of claims 51 to 106, wherein the library of nucleic acids comprises single stranded nucleic acids.
108. The method of any one of claims 51-107, wherein the plurality of library inserts comprise genomic nucleic acid.
109. The method of any one of claims 51-108, wherein the four U-block nucleic acids do not comprise degenerate nucleotide bases.
110. The method of claim 109, wherein the degenerate nucleotide base comprises 3-nitropyrrole, 5-nitroindole, an analog or derivative thereof.
111. The method of claim 109, wherein the degenerate nucleotide base comprises inosine, 2' -deoxyinosine, an analog or derivative thereof.
HK17111448.6A 2014-10-10 2015-10-08 Universal blocking oligo system and improved hybridization capture methods for multiplexed capture reactions HK1237376A1 (en)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
US62/062,616 2014-10-10
US62/062,612 2014-10-10

Publications (1)

Publication Number Publication Date
HK1237376A1 true HK1237376A1 (en) 2018-04-13

Family

ID=

Similar Documents

Publication Publication Date Title
CN107109485B (en) Universal blocking oligomer system and improved hybrid capture method for multiple capture reactions
JP7532455B2 (en) Dislocations that maintain continuity
JP7651470B2 (en) Methods and compositions for analyzing nucleic acids
EP3207134B1 (en) Contiguity preserving transposition
Wang et al. Recent advances and application of whole genome amplification in molecular diagnosis and medicine
US20250059589A1 (en) Sample preparation for nucleic acid amplification
EP3699289A1 (en) Sample preparation for nucleic acid amplification
CN105658813B (en) Chromosome conformation capture method comprising selection and enrichment steps
CN110914449A (en) Construction of sequencing library
HK1237376A1 (en) Universal blocking oligo system and improved hybridization capture methods for multiplexed capture reactions
JP2022552155A (en) New method
US20070148636A1 (en) Method, compositions and kits for preparation of nucleic acids
HK1235082A1 (en) Contiguity preserving transposition
HK1235082B (en) Contiguity preserving transposition