EP4627111A2 - Séquençage métagénomique à lecture longue rentable - Google Patents
Séquençage métagénomique à lecture longue rentableInfo
- Publication number
- EP4627111A2 EP4627111A2 EP24725039.2A EP24725039A EP4627111A2 EP 4627111 A2 EP4627111 A2 EP 4627111A2 EP 24725039 A EP24725039 A EP 24725039A EP 4627111 A2 EP4627111 A2 EP 4627111A2
- Authority
- EP
- European Patent Office
- Prior art keywords
- methylation
- seq
- motif
- methylated
- strand oligo
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Classifications
-
- C—CHEMISTRY; METALLURGY
- C12—BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
- C12Q—MEASURING OR TESTING PROCESSES INVOLVING ENZYMES, NUCLEIC ACIDS OR MICROORGANISMS; COMPOSITIONS OR TEST PAPERS THEREFOR; PROCESSES OF PREPARING SUCH COMPOSITIONS; CONDITION-RESPONSIVE CONTROL IN MICROBIOLOGICAL OR ENZYMOLOGICAL PROCESSES
- C12Q1/00—Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions
- C12Q1/68—Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions involving nucleic acids
- C12Q1/6806—Preparing nucleic acids for analysis, e.g. for polymerase chain reaction [PCR] assay
-
- C—CHEMISTRY; METALLURGY
- C12—BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
- C12Q—MEASURING OR TESTING PROCESSES INVOLVING ENZYMES, NUCLEIC ACIDS OR MICROORGANISMS; COMPOSITIONS OR TEST PAPERS THEREFOR; PROCESSES OF PREPARING SUCH COMPOSITIONS; CONDITION-RESPONSIVE CONTROL IN MICROBIOLOGICAL OR ENZYMOLOGICAL PROCESSES
- C12Q2600/00—Oligonucleotides characterized by their use
- C12Q2600/154—Methylation markers
Definitions
- the present inventive concept is directed to methods and compositions for selectively amplifying a nucleic acid of one or more prokaryotic organisms of interest in a microbiome sample.
- Metagenomics have enabled the comprehensive and culture-independent study of microbiomes. However, for many applications, only certain taxa are of interest in a microbiome sample, for example, pathogenic bacteria, beneficial microbes, or taxa of low relative abundance. It is highly desirable that a method can effectively enrich bacterial taxa of interest directly from the microbiome.
- the present disclosure is based on, in part, the suppressing discovery' that natural self vs. non-self genome differentiation as provided by bacterial DNA methylation can be exploited to rationally choose methylation-sensitive restriction enzymes (REs), individually or in combination, to deplete host DNA and most background microbial DNA while enriching bacterial taxa of interest. Accordingly, the present disclosure herein provides for a new method of selectively enriching any bacterial target of interest in a sample of mixed bacterial taxa.
- REs methylation-sensitive restriction enzymes
- a method of selectively amplifying a nucleic acid of one or more prokaryotic organisms of interest (target organisms) in a microbiome sample comprises: (a) obtaining or having obtained a microbiome sample comprising a plurality of prokaryotic-organisms, wherein the microbiome sample comprises genomic DNA (gDNA) from one or more target organisms (target gDNA) and gDNA from one or more prokaryotic organisms that are not of interest (background organism gDNA); (b) isolating gDNA from the microbiome sample to obtain target gDNA fragments and background organism gDNA fragments; (c) selecting one or more methylation-sensitive restriction enzymes based on the DNA methylome (a collection of DNA methylation sequence motifs) of the one or more target organisms wherein the one or more methylation restriction enzymes target motifs that are primarily methylated in the target gDNA (methylated
- the methods provided herein further comprise an end repair and/or a dA-tailing step prior to the linking of the universal PCR adapters.
- the universal PCR adaptor can comprise an upper strand oligo and a lower strand oligo, wherein at least 10 nucleotides at the 3’ end of the upper strand oligo are complementary to and hybridize with at least 10 nucleotides at the 5’ end of the lower strand oligo to form a hybridized region and at least 20 nucleotides at the 5’ end of the upper strand oligo are not complementary' and do not hybridize to at least 20 nucleotides at the 3’ end of the lower strand oligo, and wherein the hybridized region does not comprise a nonmethylated motif targeted by the one or more restriction enzyme.
- the universal PCR adaptor further comprises a *T at the 3’ end of the upper strand oligo and/or is phosphorylated at the 5’ end of the lower strand oligo.
- the upper strand oligo may have a nucleotide sequence comprising any one of SEQ ID NOs: 1 and 3-18 and the lower strand oligo may have a nucleotide sequence comprising any one of SEQ ID NOs: 2 and 19-34.
- the upper strand oligo may have a nucleotide sequence comprising SEQ ID NO: 35 and/or the lower strand oligo has a nucleotide sequence comprising SEQ ID NO: 36; or the upper strand oligo has a nucleotide sequence comprising SEQ ID NO: 37 and/or the lower strand oligo has a nucleotide sequence comprising SEQ ID NO: 38.
- each upper-strand oligo of the universal primers herein can comprise a primer binding site complementary’ to a forward barcoded primer, and each lower- strand oligo comprises a primer binding site complementary’ to a reverse barcoded primer.
- the methods may further comprise deconvolving genomes of prokaryotic organisms in the microbiome sample to identify a methylation profile of the target prokaryotic organism, wherein deconvolving genomes comprises: A) obtaining a microbiome sample comprising a plurality of prokaryotic-organisms; B) sequencing nucleic acids of the prokaryotic organisms using single-molecule long reads sequencing technology, wherein the sequencing comprises the step of identifying methylated nucleotides, and at least one of the steps of: i. sequencing single molecule reads of nucleic acids; ii.
- step (D) determining nucleic acid methylation profiles of the assembled contigs or the single molecule reads in the microbiome sample based on motifs identified in step (D); F) separating the assembled contigs and/or the single molecule reads into bins corresponding to distinct prokaryotic organisms based on the methylation profiles of step (E); and G) assembling the bins of step (F), thereby obtaining assembled genomes of the distinct bacterial organisms in the microbiome sample, thereby deconvolving genomes of the prokaryotic organisms in the microbiome sample.
- the method further comprises the step of combining the methylation profiles of step (E) with other sequence features of the nucleic acids of the prokaryotic organisms in the microbiome sample prior to separating the assembled contigs and/or the single molecule reads into bins.
- the target prokaryotic organism(s) comprises one or more of an E. coll strain.
- Akkermansia muciniphila. Alistipes fmegoldii and Bacteroides ovatus. Bifidobacterium longum, Ruminococcus bromii, Eubacterium rectale, Bacteroides uniformis D7, Dorea longicatena and Barnesiella intestinihominis, Eubacterium sp. CAG:248, Alistipes onderdonkii, Enterocloster clostridioformis, Bacteroides fragilis, Gordonibacter pamelaea,
- Adlercreutzia sp. 8CFCBH1 Adlercreutzia equolifaciens, or any combination thereof.
- the target prokaryotic organism(s) are selected from specific pathogenic bacteria, beneficial bacteria, probiotics, or prokaryotic organisms with low relative abundance in a microbiome sample as determined by standard metagenomic sequencing.
- the target prokaryotic organism(s) comprising a methylated CTS6mAG motif may comprise one or more species of Alistipes finegoldii and/or Bacteroides ovatus
- the target prokaryotic organism(s) comprising a methylated RG6mATCY motif may comprise one or more species of Bifidobacterium longum, Ruminococcus bromii. Eubacterium rectale and Bacteroides uniformis
- the target prokaryotic organism(s) comprising a methylated G6mATGC motif may comprise one or more species of Dorea longicatena.
- the target prokary otic organism(s) comprising a methylated G6mATGC motif and methylated CTNN6mAG motif may comprise one or more species of Barnesiella intestinihominis.
- any of the foregoing or related methods may further comprise sequencing the nucleic acid of the amplified target prokary otic organism(s).
- the sequencing is performed using a single-molecule long read real time (SMRT) technology, nanopore sequencing technology 7 or other short read sequencing technologies including but not limited to Illumina sequencing platforms.
- SMRT single-molecule long read real time
- nanopore sequencing technology 7 or other short read sequencing technologies including but not limited to Illumina sequencing platforms.
- the microbiome sample may 7 be obtained from soil, air, water, sediment, oil, and combinations thereof.
- the microbiome sample is obtained from water selected from marine water, fresh water, and rainwater.
- the microbiome sample is obtained from a subject selected from a protozoa, an animal, or a plant.
- the subject is a mammal.
- the subject is human.
- the subject is an infant.
- the adaptor comprises an upper strand oligo and a lower strand oligo, wherein at least 10 nucleotides at the 3’ end of the upper strand oligo are complementary to and hybridize with at least 10 nucleotides at the 5’ end of the lower strand oligo to form a hybridized region and at least 20 nucleotides at the 5’ end of the upper strand oligo are not complementary and do not hybridize to at least 20 nucleotides at the 3’ end of the lower strand oligo, wherein the hybridized region does not comprise a motif targeted by a methylation sensitive restriction enzyme or does comprise a methylated nucleotide in a motif targeted by the methylation sensitive restriction enzyme.
- the universal adaptor provided herein further comprises a *T at the 3’ end of the upper strand oligo and/or a phosphorylated base at the 5’ end of the lower strand oligo.
- a subgroup of bacteria of interest can be enriched (circled by the red dashed line), by depleting the rest of the bacteria (the gray circle).
- mEnrich-seq allow researchers to examine different subgroups of bacteria in the microbiome sample (circled by individual yellow or blue dashed lines). mEnrich-seq can also be tailored with the use of multiple REs to more specifically enrich bacterial taxa with multiple methylation motifs (as illustrated by both the yellow and blue scissors, enriching the bacterium shared between blue and yellow dashed line).
- FIG. 2D depicts motif frequency on the mEnrich-seq reads that were mapped to N. gonorrhoeae. In replicates 1 and 2, 94.66% and 92.26% of reads that mapped to N. gonorrhoeae do not have any GATC sites (blue), respectively.
- FIG. 2F depicts scatter plots for three urine samples (UTI-1 to -3) showing the percentage of reads mapped to E. coli, human and other taxa from the input (x-axis) vs. mEnrich-seq (y-axis). Across all the three samples, E. coli reads are significantly enriched in mEnrich-seq.
- FIG. 3A-3D depict the enrichment of A. muciniphila from fecal samples by mEnrich-seq
- FIG. 3A shows scatter plots for three fecal samples (GUT-1 to -3) showing the percentage of reads mapped to A. muciniphila and other taxa from the input (x- axis) and mEnnch-seq (y-axis).
- top 50 abundant species, in input or mEnrich-seq are shown.
- the size of the dots represents the fold change by mEnrich-seq over the input.
- Significant increases in read % are observed for A.
- FIG. 3B shows fold enrichment of reads that map to A. muciniphila from mEnrich-seq over the input for GUT-1 to -3.
- FIG. 3C shows A. muciniphila genome coverage by sequencing reads of GUT-1 to -3. Reads from mEnrich-seq (orange) and the input standard metagenomic sequencing (blue) were mapped to the genome of corresponding A. muciniphila isolates (sequenced and assembled beforehand for evaluation purpose), respectively.
- FIG. 3D shows analysis of RAATTY motif frequency (per kilobase) for the taxa enriched by mEnrich-seq (XapI as RE) in GUT-1 to -3.
- X-axis RAATTY motif frequency on the reference genomes of enriched taxa;
- Y-axis read-level RAATTY frequency in mEnrich-seq.
- the green dashed line circles A. muciniphila. which is enriched due to its methylated RA6mATTY motif sites.
- the red dashed line circles the bacterial taxa enriched due to the sparse RAATTY frequency on their genomes.
- FIG. 4A-4J depict enrichment sequencing of bacterial taxa based on de novo discovered methylation motifs, wherein: FIG. 4A depicts a workflow for rational design of mEnrich-seq based on de novo motif discovery.
- An initial metagenomic sequencing was conducted to estimate the relative abundance of different taxa. The goal is to design mEnrich- seq to enrich taxa with relatively low abundance. To do so, de novo methylation motif discovery' was performed from the initial metagenomic sequencing data. Importantly, motif discovery' does not need fully assembled genomes, because a partially assembled contig can already inform us about the methylation motifs from a strain.
- FIG. 4B depicts fold enrichment of six bacterial strains with relatively low abundance; enrichment by mEnrich-seq based on de novo discovered methylation motifs.
- One of the strain B. intestinihominis was further ennched using multiple REs simultaneously to provide greater enrichment (85-fold) than a single RE (30-fold), illustrating the flexible design of mEnrich-seq when multiple methylation motifs are discovered on the strain of interest.
- FIG. 4C depicts scatter plots showing the percentage of reads mapped to bacteria from the input (x-axis) and mEnrich-seq using the RE (BspCNI) targeting the CTCAG motif (y-axis).
- the size of the dots represents the fold change by mEnrich-seq over the input.
- FIG. 4D depicts analysis of the CTCAG motif frequency (per kilobase) for the taxa enriched by mEnrich-seq (BspCNI as RE). X-axis.
- FIG. 4E depicts data as described for FIG. 4C but for the RE targeting the RGATCY moti.
- FIG. 4F depicts data as described for FIG. 4D but for the RE targeting the RGATCY motif.
- FIG. 4G depicts data as described for FIG. 4C but for the RE targeting the GATGC motif.
- FIG. 4H depicts data as described for FIG.
- FIG. 41 depicts data as described for FIG. 4C but for the multiple REs targeting the GATGC motif and CTNNAG motif simultaneously.
- FIG. 4J shows that mEnrich-seq improved genome assembly compared to standard metagenomic sequencing. The left shows the completeness (%) of genome assembly; blue bars show the input and orange bars show the mEnrich-seq; * completeness less than 1%. On the right, contig N50 length is shown in Mb. For each targeted species, reads from input and mEnrich-seq were normalized by random subsampling to make sure reads originated from the targeted species have the same yield for genome assembly comparison.
- FIG. 5A-5E depict experimental data assessing the broad applicability of mEnrich- seq, wherein: FIG. 5A depicts a histogram for the number of methylation motifs in each bacterial genome across the 4,601 bacterial methylomes mapped to date (as of 07/18/2022, REBASE). X-axis, number of motifs per bacterium; Y -axis, number of bacterial species having a given number of motif.
- FIG. 5B depicts a histogram for the number of non-degenerate bases (excluding 'N’s) per motif across 4683 motifs. For example, both GATC and CTNNAG have 4 non-degenerate bases, FIG.
- FIG. 5C depicts a phylogenic tree containing 4,601 bacterial strains with mapped methylomes to date.
- the bacterial strains that can be targeted (enriched by depleting the background metagenomic DNA) by at least one commercially available REs are highlighted with colors, demonstrating broad distribution across the tree.
- Number of motifs recognized by commercially available REs (and motif length) for a given bacteria are color coded as detailed in the panel at the bottom right
- FIG. 5D depicts a histogram for number of bacteria targeted (enriched by depleting the background metagenomic DNA) by each commercially available RE (please note that a single motif can be targeted by multiple restriction enzymes, namely, isoschizomers).
- FIG. 5E shows taxonomic distribution of the seven most prevalent methylation motifs.
- the taxonomic information is indicated by color-coding branches by order (top 27 orders are color coded). The distribution of top seven motifs is highlighted by colored squares around the phylogenetic tree.
- FIG. 6A-6C depict mEnrich-seq of urine samples using Nanopore Flongle flow cells and wherein: FIG. 6A shows three urine samples were sequenced as the Input: UTI-1 contained mostly bacterial DNA, while UTI-2 and UTI-3 were dominated by human DNA, FIG. 6B depicts scatter plots for six urine samples (UTI-1 to -6) showing the percentage of reads mapped to E. coli, human and other taxa from the input (x-axis) vs. mEnrich-seq (y- axis). The UTI-1 to -3 samples demonstrate consistent patterns as shown in FIG.
- FIG. 12A-12B depict agarose gels of enriched low-abundance bacterial taxa from the adult fecal sample based on de novo discovered methylation motifs, wherein: FIG. 12A depicts gels for the adult fecal sample digested by different REs. Left, the Mfll-digested the adult fecal sample DNA (targeting RGATCY motif) shows the band at -5- lOkb; the BspCNI (CTC AG as recognition site, part of the CTSAG motif)- digested sample shows a faint band at -5- lOkb (based upon the DNA ladder).
- targeting RGATCY motif shows the band at -5- lOkb
- BspCNI CTC AG as recognition site, part of the CTSAG motif
- FIG. 12B depicts gels for PCR bands amplified by the barcoded primers, using RE-digested adult fecal DNA (from Fig. 7a) as the template.
- the PCR bands of RE(s)-digested DNA are shown around 5—10 kb.
- FIG. 13 depicts a plot showing de novo methylation motif discovery from an initial shallow' sequencing run (400k PacBio Sequel II CCS reads). Many methylation motifs were discovered from SMRT sequencing data, even those from bacterial genomes with -10% completeness. The relative abundance of each bacterial genome is included in the x-axis label in red color. 4-mer motifs (those with four non-degenerate bases) and similarly 5-mer and 6- mer motifs are color coded [0050]
- FIG. 14A-14B shows via a direct comparison that the mEnrich-seq design has better fold enrichment than REMoDE.
- FIG. 14A-14B shows via a direct comparison that the mEnrich-seq design has better fold enrichment than REMoDE.
- FIG. 15A-15B show that mEnrich-seq design has better fold enrichment than REMoDE when they are compared with the same, real microbiome samples with DNA fragmentation and damages.
- FIG. 15A depicts agarose gel electrophoresis tested once on genomic DNA isolated from Mock-3 (illustrating the ideal high molecular weight gDNA from cell culture) and two infant fecal samples (representing DNA quality’ in real applications).
- FIG. 15B depicts data comparing the fold enrichment of A. muciniphila from two infant fecal samples using mEnrich-seq, REMoDE, or REMoDE with DNA repair beforehand (see Supplementary’ Information). As DNA repair is part of mEnrich-seq by default, this step is also applied to REMoDE for a fair comparison.
- FIG. 16 depicts an evaluation of the benefit of adapter ligation before vs. after RE- digestion using the Mock-3.
- the bar graph demonstrates that adapter ligation before RE- digestion indeed has better fold ennchment.
- FIG. 17 shows that an alternative adapter has comparable enrichment efficiency as the original adapter.
- the fold enrichment of A. muciniphila (using Xapl as RE) was tested on Mock-3 with either the original adapter (having an upper and loyver oligonucleotide corresponding to SEQ ID NOs: 35 and 36, respectively) or the alternative adapter (having an upper and loyver oligonucleotide corresponding to SEQ ID NOs: 37 and 38, respectively).
- FIG. 18 shows that GC content does not introduce systematic bias on read coverage across the E. coll genome analyzed by mEnrich-seq. Analysis was performed on Mock-1 and Mock-2 (E. coll relative abundance as 0.91% and 1.20%, see Supplementary Information). For a fair comparison, the same yield of data from both input and mEnrich-seq were used in data analysis and figures.
- Each dot represents a 5 kb genomic region, colored by GC content.
- X- axis number of input reads mapped to each 5 kb genomic region;
- Y-axis number of mEnrich- seq reads mapped to each 5 kb genomic region. To ease visualization, a random value was added on x-values to avoid overlapping dots.
- FIG. 19 shows that GC content does not introduce systematic bias on read coverage across the A. muciniphila genome by mEnrich-seq.
- the same type of scatter plots (as the ones above for E. coli) made with sequencing data from three human gut samples (A. muciniphila relative abundance: 0.48%, 0.52% and 1.52%, respectively) to characterize the enrichment of A. muciniphila. Consistent with the Circos plot for the same samples, read depths were capped at the 95% quantile to ease visualization.
- the present disclosure is based, at least in part, on the discovery that natural self vs. non-self genome differentiation as provided by bacterial DNA methylation can be exploited to rationally choose methylation-sensitive restriction enzymes (REs), individually or in combination, to deplete host DNA and most background microbial DNA while enriching bacterial taxa of interest. Accordingly, the present disclosure herein provides for a new method of selectively enriching any bacterial target of interest in a sample of mixed bacterial taxa.
- REs methylation-sensitive restriction enzymes
- the term “about.” can mean relative to the recited value, e.g . amount, dose, temperature, time, percentage, etc., ⁇ 10%, ⁇ 9%, ⁇ 8%, ⁇ 7%, ⁇ 6%, ⁇ 5%, ⁇ 4%, ⁇ 3%, ⁇ 2%, or ⁇ 1%.
- nucleic acid' refers to deoxyribonucleic acids (DNA) or ribonucleic acids (RNA) and polymers thereof in either single- or double-stranded form. Unless specifically limited, the term encompasses nucleic acids containing known analogues of natural nucleotides that have similar binding properties as the reference nucleic acid and are metabolized in a manner similar to naturally occurring nucleotides. Unless otherwise indicated, a particular nucleic acid sequence also implicitly encompasses conservatively modified variants thereof (e.g.. degenerate codon substitutions), alleles, orthologs, SNPs, and complementary sequences as well as the sequence explicitly indicated.
- DNA deoxyribonucleic acids
- RNA ribonucleic acids
- degenerate codon substitutions may be achieved by generating sequences in which the third position of one or more selected (or all) codons is substituted with mixed-base and/or deoxyinosine residues (Batzer et al., Nucleic Acid Res. 19:5081 (1991); Ohtsuka et al., J. Biol. Chem. 260:2605-2608 (1985); and Rossolmi et ak.Afo/. Cell. Probes 8:91-98 (1994)).
- a method of selectively amplifying a nucleic acid of one or more prokaryotic organisms of interest (target organisms) in a microbiome sample comprises: (a) obtaining a microbiome sample comprising a plurality of prokaryotic-organisms, wherein the microbiome sample comprises genomic DNA (gDNA) from one or more target organisms (target gDNA) and gDNA from one or more prokaryotic organisms that are not of interest (background organism gDNA); (b) isolating gDNA from the microbiome sample to obtain target gDNA fragments and background organism gDNA fragments; (c) selecting one or more methylationsensitive restriction enzymes based on the DNA methylome (a collection of DNA methylation sequence motifs) of the one or more target organisms wherein the one or more methylation restriction enzymes target motifs that are primarily methylated in the target gDNA (methylated motifs) and
- the methods here are, in part, directed to methods that improve upon other methods of using methylation sensitive restriction enzymes to enrich target nucleic acid populations, by ligating a set of universal adaptors to the target (and non-target) DNA before any digestion occurs.
- the universal adaptors also referred to herein as “universal PCR adaptors’'
- Suitable universal PCR adaptors are described further in Section (III) below.
- These universal adaptors may be targeted by suitable PCR primers after digestion has occurred to selectively amplify and/or sequence any un-digested DNA (i.e., any DNA fragments still containing universal adaptors at both the 5‘ and 3‘ ends).
- Suitable PCR primers can target certain primer binding motifs in the universal PCR adaptors as described further in Section (III) below.
- the methylated motif comprises a nucleic acid sequence motif having methylation on one or more nucleotides, wherein the methylation optionally comprises N4-methylcytosine (4mC), N6-methyladenine (6mA) or 5- Methylcytosine (5mC), and wherein the methylated motif optionally comprises GATC, RAATTY where R is A or G and Y is C or T, or CTSAG where S is G or C, RGATCY where R is A or G and Y is C or T, GATGC, CTNNAG and each A is N6-methyladenine (6mA).
- the methylation optionally comprises N4-methylcytosine (4mC), N6-methyladenine (6mA) or 5- Methylcytosine (5mC)
- the methylated motif optionally comprises GATC, RAATTY where R is A or G and Y is C or T, or CTSAG where S is G or C, RGATCY where R is A or G and Y
- the non-methylated motif comprises a nucleic acid sequence motif that is recognized and cleaved by a restriction enzyme, and wherein the non-methylated motif optionally comprises GATC, RAATTY where R is A or G and Y is C or T, or CTSAG where S is G or C, RGATCY where R is A or G and Y is C or T, GATGC, CTNNAG and each A is non-methylated.
- the one or more methylation sensitive restriction enzymes may comprise a class of natural and/or artificial restriction enzymes that are sensitive to DNA methylation, and optionally comprise methylation-sensitive restriction enzy mes DpnII, XapI, Apol, BspCNI, BseMII, Mfll, BstX2I, BstYI, Mfll, XhoII, SfaNI, AflII, Acul. Bpml, BpuEI, PstI, any methylation-sensitive isoschizomer thereof, or a combination of any thereof.
- the methylation-sensitive restriction enzyme comprises a restriction enzyme whose cleavage ability is blocked by the methylated motif sites.
- the method comprises identifying a methylation profile of the target prokaryotic organism and selecting the one or more restriction enzy mes based on the methylation profile of the target prokaryotic organism.
- the method further comprises deconvoluting genomes of prokaryotic organisms in the microbiome sample to identify a methylation profile of the target prokaryotic organism, wherein deconvoluting genomes comprises: A) obtaining a microbiome sample comprising a plurality of prokary otic- organisms; B) sequencing nucleic acids of the prokaryotic organisms using single-molecule long reads sequencing technology, wherein the sequencing comprises the step of identifying methylated nucleotides, and at least one of the steps of: i. sequencing single molecule reads of nucleic acids; ii.
- step (D) determining nucleic acid methylation profiles of the assembled contigs or the single molecule reads in the microbiome sample based on motifs identified in step (D); F) separating the assembled contigs and/or the single molecule reads into bins corresponding to distinct prokaryotic organisms based on the methylation profiles of step (E); G) assembling the bins of step (F), thereby ⁇ obtaining assembled genomes of the distinct bacterial organisms in the microbiome sample, thereby deconvoluting genomes of the prokaryotic organisms in the microbiome sample.
- the method of deconvoluting genomes further comprises the step of combining the methylation profiles of step (E) with other sequence features of the nucleic acids of the prokaryotic organisms in the microbiome sample prior to separating the assembled contigs and/or the single molecule reads into bins.
- Additional methods useful to deconvolve genomes to derive at ideal combinations of methylation sensitive restriction enzymes or that may be used in combination with any of the methods and compositions provided herein are described in US20200160936A1, US20220254446A1, US20220230704A1, which are each incorporated herein by reference in their entirety.
- the target prokary otic organism may comprise one or multiple of the following an E. coli strain, Akkermansia muciniphila, Alistipes fmegoldii and Bacteroides ovatus. Bifidobacterium longum.. Ruminococcus bromii, Eubacterium rectale, Bacteroides uniformis D7, Dorea longicatena and Barnesiella intestinihominis, Eubacterium sp.
- the target prokaryotic organisms may be selected from specific pathogenic bacteria, beneficial bacteria, probiotics, or prokaryotic organisms with low relative abundance in a microbiome sample as determined by standard metagenomic sequencing.
- the target prokaryotic organism, methylation motif and restriction enzyme may be selected from the following groups: (i) the target prokaryotic organism(s) comprise a methylated G6mATC motif, and the one or more methylation sensitive restriction enzyme comprises methylation sensitive DpnII and/or a methylation sensitive isoschizomer thereof; (ii) the target prokaryotic organism(s) comprise a methylated RA6mATTY motif and the one or more methylation sensitive restriction enzyme comprises methylation sensitive XapI , methylation sensitive Apol, and/or a methylation sensitive isoschizomer thereof, or any combination thereof; (iii) the target prokaryotic organism(s) comprise a methylated CTS6mAG motif and the one or more methylation sensitive restriction enzyme comprises methylation sensitive BspCNL methylation sensitive BseMII, a methylation sensitive isoschizomer thereof, or any combination thereof; (iv)
- the target prokaryotic organism(s) comprising a methylated G6mATC motif comprise Escherichia coll, Klebsiella pneumoniae, one or more species of Gammaproteobacteria class or any combination thereof; (ii) the target prokaryotic organism(s) comprising a methylated RA6mATTY motif comprise one or more species of Akkermansia muciniphila, Campylobacter spp., Acinetobacter spp., Spirochaeta spp., Treponema spp.
- the target prokaryotic organism(s) comprising a methylated CTS6mAG motif comprise one or more species ofAlistipes finegoldii and/or Bacteroides ovatus;
- the target prokaryotic organism(s) comprising a methylated RG6mATCY motif comprise one or more species of Bifidobacterium longum, Ruminococcus bromii, Eubacterium rectale and Bacteroides uniformis;
- the target prokaryotic organism(s) comprising a methylated G6mATGC motif comprise one or more species of Dorea longicatena, Barnesiella intestinihominis and Eubacterium sp.
- CAG:248 and/or the target prokaryotic organism(s) comprising a methylated G6mATGC motif and methylated CTNN6mAG motif comprise one or more species of Barnesiella intestinihominis.
- the method may further comprise a DNA shearing step, and wherein shearing the gDNA optionally comprises g-TUBEs and the gDNA fragments from target and background organisms have a length of about 10 kb.
- the method further comprises an end repair and dA-tailing step prior to the linking of said universal PCR adapter.
- the method may comprise sequencing the nucleic acid of the amplified target prokaryotic organism(s).
- the sequencing is performed using a single-molecule long read real time (SMRT) technology, nanopore sequencing technology or other short read sequencing technologies including but not limited to Illumina sequencing platforms.
- SMRT single-molecule long read real time
- the microbiome sample is obtained from soil, air, water, sediment, oil, and combinations thereof.
- the microbiome sample may be obtained from water selected from marine water, fresh water, and rainwater.
- the microbiome sample may be obtained from a subject selected from a protozoa, an animal, or a plant.
- the subject may be a mammal.
- the subject is a human.
- the subject may be an infant.
- the sample obtained from a mammal e.g., human
- compositions 111.
- a universal adaptor for use in any of the methods described herein.
- the universal adaptors can comprise about 50-100 bp double stranded short DNA fragments with an overhang at the 3’ end of the top strand, which facilitates ligation to all DNA molecules through T/A ligation.
- the universal adaptors will not include any recognition sequences in the adapter sequence. Suitable universal adaptors may be “Y- shaped” or “bell-shaped”, or may come in other structures.
- the universal PCR adaptor comprises an upper strand oligo and a lower strand oligo, wherein at least 10 nucleotides at the 3’ end of the upper strand oligo are complementary to and hybridize with at least 10 nucleotides at the 5’ end of the lower strand oligo to form a hy bridized region and at least 20 nucleotides at the 5’ end of the upper strand oligo are not complementary and do not hybridize to at least 20 nucleotides at the 3’ end of the lower strand oligo, and wherein the hybridized region does not comprise a nonmethylated motif targeted by the one or more restriction enzyme.
- the universal adaptor can further comprise a *T at the 3‘ end of the upper strand oligo, wherein *T comprises a phosphorothioated thiamine base having a phosphorothioate nucleotide modification, wherein the phosphorothioate modification renders an inter-nucleotide linkage formed by the universal adaptor resistant to nuclease degradation.
- the universal adaptor can be phosphory lated at the 5’ end, particularly on the lower strand oligonucleotide. Other modifications to prevent nuclease degradation are known in the art and are contemplated herein.
- the non-hybridized region of the universal PCR adaptor can comprise or consist of SEQ ID NO: 1 (for the upper oligo) and/or SEQ ID NO: 2 (for the lower oligo).
- Each of the upper and lower oligo may comprise of an additional 10 to 25 nucleotides that hybridize and that, when hybridized, do not comprise a non-methylated motif targeted by the one or more restriction enzyme.
- the upper and lower oligos comprise the same number of additional nucleotides that hybridize.
- the upper strand may comprise a nucleotide sequence of any one of SEQ ID NOs: 3-18 and the lower strand oligo has a nucleotide sequence comprising any one of SEQ ID NOs: 19-34.
- the upper strand oligo has a nucleotide sequence comprising SEQ ID NO: 35 and/or the lower strand oligo has a nucleotide sequence comprising SEQ ID NO: 36.
- the upper strand oligo has a nucleotide sequence comprising SEQ ID NO: 37 and/or the lower strand oligo has a nucleotide sequence comprising SEQ ID NO: 38.
- SEQ ID NOs: 1 -38 are provided in the Table 1 below, where the phosphorothioate linkage is indicated by an asterisk (*) and the 5’phosphorylation modification is indicated by a /5Phos/.
- each upper-strand oligo comprises a primer binding site complementary to a forward barcoded primer
- each lower-strand oligo comprises a primer binding site complementary to reverse barcoded primer
- a set of barcoded PCR primers is provided, where the barcoded PCR primers target any portion of the upper or lower oligo of the universal adaptor.
- the barcoded PCR primer may target a nucleic acid sequence comprising at least a portion of a nucleic acid sequence comprising or consisting of SEQ ID NOs 1 or 2, which are shown again in Table 2 for ease of reference.
- the term "‘target” means that the primer hybridizes to a nucleic acid comprising the nucleic acid target sequence to enable PCR amplification.
- the primer hybridizes to the 5’ strand (e g., upper oligo of universal adaptor) comprising the target sequence or to the 3 ' strand complementary to the 5 ’ strand. In other aspects, the primer hybridizes to the 3‘ strand (e.g., lower oligo of universal adaptor) comprising the target sequence or to the 5’ strand complementary to the 3’ strand.
- the 5’ strand e g., upper oligo of universal adaptor
- the primer hybridizes to the 3‘ strand (e.g., lower oligo of universal adaptor) comprising the target sequence or to the 5’ strand complementary to the 3’ strand.
- kits comprising (a) a set of universal PCR adapters and (b) one or methylation sensitive restriction enzymes.
- the kit further comprises (c) a set of barcoded PCR primers complementary to at least a portion of the universal PCR adapters.
- the one or more methylation sensitive restriction enzymes comprise restriction enzymes blocked by the methylated motifs in one or more target prokaryotic organisms, and wherein comprise but not limited to DpnII, XapI, Apol, BspCNI, BseMII, Mill, BstX2I, BstYI, Mill, XhoII, SfaNI, AflII, Acul, Bpml, BpuEI, PstI or a combination of any thereof.
- each of the universal PCR adaptors may comprise an upper strand oligo and a lower strand oligo, wherein at least 10 nucleotides at the 3’ end of the upper strand oligo are complementary to and hybridize with at least 10 nucleotides at the 5’ end of the lower strand oligo to form a hybridized region and at least 20 nucleotides at the 5’ end of the upper strand oligo are not complementary and do not hybridize to at least 20 nucleotides at the 3’ end of the lower strand oligo, wherein the hybridized region and wherein the hybridized region does not comprise a non-methylated motif targeted by the one or more restriction enzyme.
- each universal PCR adapter in the kit further comprising a *T at the 3’ end of the upper strand oligo, wherein *T comprises a phosphorothioated thiamine base having a phosphorothioate nucleotide modification, wherein the phosphorothioate modification renders an inter-nucleotide linkage formed by the Y shaped adaptor resistant to nuclease degradation.
- the universal adaptor can be phosphorylated at the 5’ end, particularly on the lower strand oligonucleotide. Other modifications to prevent nuclease degradation are known in the art and are contemplated herein.
- the upper strand oligo can have a nucleotide sequence comprising any one of SEQ ID NOs: 3-18 and/or the lower strand oligo can have a nucleotide sequence comprising any one of SEQ ID NOs: 19-34.
- the upper strand oligo can have a nucleotide sequence comprising SEQ ID NO: 35 and/or the lower strand oligo can have a nucleotide sequence comprising SEQ ID NO: 36.
- the upper strand oligo can have a nucleotide sequence comprising SEQ ID NO: 37 and/or the lower strand oligo can have a nucleotide sequence comprising SEQ ID NO: 38.
- the set of barcoded PCR primers may target any nucleic acid of the upper or lower oligo of the universal adaptor.
- the barcoded PCR primer may target a nucleic acid sequence comprising at least a portion of a nucleic acid sequence comprising or consisting of: SEQ ID NO: 1 or 2.
- the term “target” means that the primer hybridizes to a nucleic acid comprising the nucleic acid target sequence to enable PCR amplification.
- the forward primer hybridizes to the sense strand of the target sequence comprising upper oligo of universal adaptor (i.e., any one of SEQ ID NOs: 3- 18).
- the reverse primer hybridizes to the anti-sense strand of the target sequence comprising lower oligo of universal adaptor (i.e., any one of SEQ ID NOs: 19-34).
- Kits may optionally provide additional components such as buffers and interpretive information.
- the kit includes a container and a label or package insert(s) on or associated with the container.
- the invention provides articles of manufacture comprising contents of the kits described above.
- Metagenomics have enabled the comprehensive and culture-independent study of microbiomes. However, for many applications, only certain taxa are of interest in a microbiome sample, for example, pathogenic bacteria, beneficial microbes, or taxa of low relative abundance. It is highly desirable that a method can effectively enrich bacterial taxa of interest directly from the microbiome. To address this critical need. mEnrich-seq, a method that can enrich taxa of interest from metagenomic DNA before sequencing, was developed. The core idea is to exploit the natural self vs.
- non-self genome differentiation provided by bacterial DNA methylation and rationally choose methylation-sensitive restriction enzymes (REs), individually or in combination, to deplete host DNA and most background microbial DNA while enriching bacterial taxa of interest.
- This core idea is integrated with library preparation procedures in a way that only non-digested DNA libraries are sequenced. In-depth evaluations of mEnrich-seq were performed using both synthetic and real microbiome samples. The use of mEnrich-seq is demonstrated in several applications to enrich (up to 117-fold) genomic DNA of pathogenic or beneficial bacteria from urine and fecal samples, including several species that are hard to culture or of low abundance.
- mEnrich-seq provides microbiome researchers with a versatile and cost-effective approach for selective sequencing of diverse taxa of interest directly from the microbiome.
- Host genomic DNA can be eliminated using chemicals such as saponin to selectively lyse mammalian cells, which enables the efficient sequencing of pathogens in biospecimens that are dominated by host gDNA (for example, upper respiratory tract infection).
- Host genetic material can also be filtered using DNA binding proteins that specifically recognize 5-methylcytosine (5mC), abundant in mammalian genomes.
- 5mC 5-methylcytosine
- Flow cytometry has also been used to segregate complex microbiomes into small subsets, which can be sequenced and computationally analyzed separately.
- Adaptive sampling in the Nanopore platform can examine the first few hundred base pairs, during real-time sequencing, against a collection of reference sequences to reject the DNA molecules that meet certain criteria, for example, host DNA or a highly abundant bacterium for which enough data have been already collected.
- This strategy by adaptive sampling can reduce further consumption of sequencing yield by highly abundant species; however, its enrichment efficiency for specific taxa of interest is relatively modest, and insufficient to differentiate between bacterial taxa with highly similar genomes in complex microbiome samples.
- a method is described herein that takes advantage of bacterial DNA methylation in microbiomes, namely bacterial epigenomes.
- bacterial DNA methylation in the bacterial kingdom, there are three main forms of DNA methylation: N6-methyladenine (6mA), N4-methylcytosine (4mC) and 5mC, which are catalyzed by methyltransferases (MTases) that apply methyl groups in a highly sequence-specific manner, that is. some sequence motifs of a genome are nearly 100% methylated, while most of the genome is not methylated.
- the bacterial methylome has three fundamental properties.
- DNA methylation is present in nearly all bacteria (more than 95%).
- all genetic contents chromosomes, plasmids
- the methylation motifs are highly variable between different species and different strains of the same species. Based on these three properties, bacterial DNA methylation naturally differentiates self from nonself DNA, which serves as the foundation of restrictionmodification systems, and has been exploited as natural epigenetic barcodes to group highly similar species and strains in metagenomic analyses, namely methylation binning.
- E. coli K-12 MG1655 was obtained from American Type Culture Collection (ATCC, cat# 47076, VA. USA) and cultured with Miller Luria Broth (LB) medium, and on LB agar plates (Thermo Fisher Scientific, cat# BP1426-500, and cat# BP1425-500, respectively, MA, USA) at 37°C based on ATCC recommended protocols.
- N. gonorrhoeas MS 11 was obtained from ATCC (cat# BAA- 1833) and cultured based on protocols in Dillard, 2011 (Dillard, J. P. Genetic manipulation of Neisseria gonorrhoeae. Curr Protoc Microbiol 23, 4A-2 (2011)). Briefly, cells were grown in Gonococcal Base medium liquid (GCBL) with Kellogg’s supplements I&II, and on Chocolate II Agar plates (BD, cat# 22126, MD, USA) at 37°C in 5%CO 2 for 24 h.
- GCBL Gonococcal Base medium liquid
- Kellogg Kellogg
- UTI Urinary tract infection
- DNA from the human LCL cell line GM24149 E. coli MG1655, N. gonorrhoeae MS11, and urine samples was isolated with the DNeasy Blood and Tissue Kit (Qiagen, cat# 69506, MD, USA) based upon the manufacturer’s instructions, including the optional RNase A treatment provided with the kit.
- DNeasy Blood and Tissue Kit Qiagen, cat# 69506, MD, USA
- 2xl0 6 cells were collected from cell culture by centrifugation for 5 min at 300 x g. The cell pellets were resuspended in 180 pl PBS and 20 pl Proteinase K. After adding 200 pl Buffer AL, the mixture was incubated at 56°C for 30 min. The rest of extraction steps were performed according to manufacturer’s instructions.
- each 2 ml aliquot of urine was first centrifuged at 5000 x g and 4°C for 30 min, then the pellets were collected in a new tube.
- the urine pellets were resuspended with 180 pl of the bacterial enzymatic lysis buffer (20 mM Tris pH 8.0, 2mM EDTA, 1% SDS and 20 mg/ml lysozy me) and incubated at 37°C for 30 min.
- 20 pl of Proteinase K and 200 pl Buffer AL was added to the mixture, and it was incubated at 56°C for 30 min.
- the rest of the extraction steps were followed per manufacturer instructions with one modification, i.e., the DNA was eluted with 50 pl AE in the final step to concentrate the DNA.
- Fecal DNA was extracted with a QIAamp PowerFecal Pro DNA Kit (Qiagen, 51804) according to the manufacturer instructions. Briefly, 100 mg of fecal sample, 750 pl of PowerBead Solution and 60 pl of Solution Cl were added to the Bead Tube (provided). After incubation at 65°C for 10 min. the sample was homogenized with the TissueLyser II machine (Qiagen) for 10 min at 25 Hz speed (5 min each side). The homogenized samples were centrifuged to collect the supernatant at 15,000 xg for 1 min. The subsequent column-based extraction steps were performed as described in the QIAamp PowerFecal Pro DNA Kit protocol.
- the isolates were cultured from the urine samples with confirmed E. coli cases. Briefly, 10 pl of urine was spread by disposable inoculating loops (Thermo Fisher Scientific, cat# 22170201) quantitatively onto BD BBL MacConkey agar plates (Becton, Dickinson, and Company (BD), NJ, USA) and incubated for 24 h at 37°C aerobically. For each urine sample, a representative colony of E. coli was collected for DNA purification.
- Step 1 DNA shearing. Shearing was conducted to establish a uniform size of DNA inserts for better ligation efficiency, to maintain highly consistent molanty estimations for downstream loading into flow cells, and to reduce pore clogging.
- gDNA was sheared to around lOkb with g-TUBEs (Covaris, cat.520079, MA, USA) using an Eppendorf 5424 centrifuge at 5000 rpm for 3 min at room temperature. The g-TUBE was then inverted and centrifuged for another 3 min to collect the sheared DNA. The DNA was concentrated using 0.5x volume Ampure XP beads and eluted with 48 ul nuclease-free (NF) water.
- NF nuclease-free
- Step 2 End Repair and Barcode Adapter Ligation.
- the fragmented DNA was applied to ’End-prep” and “Ligation of Barcode Adapter” steps according to the manufacturer’s instructions.
- the “End-prep” was performed with 48 pl of sheared DNA in the reaction mixture using NEBNext Ultra II End Repair/dA-Tailing Module and NEBNext FFPE DNA Repair Mix (New England Biolabs (NEB), cat# E7546, M6630, respectively, MA, USA) with incubation at 20°C for 5 min and 65°C for 5 min. After purification with lx volume of AMPure XP beads, the end-repaired DNA was eluted in 30 pl nuclease-free water.
- the end-repaired DNA was ligated with 20 pl Barcode Adapter (ONT, cat# EXP-PBC096, blue cap) and Blunt/TA Ligase Master Mix (NEB, cat# M0367) for lOmin at room temperature, which attached the universal PCR handle to all of the DNA molecules.
- the Barcode-Adapter- ligated-sample was then purified with lx volume of Ampure XP beads and eluted in 26 pl NF water.
- 20 ng of the PCR Adapter ligated DNA was split into another tube as sample “Input”, while the rest of the sample was used for the restriction enzyme digestion step of mEnrich-seq.
- Step 3 Restriction enzyme (RE) digestion.
- RE digestion for mEnrich-seq was conducted following the manufacturer’s instructions.
- 10 U of enzyme (NEB, cat # R0543) was used in a 50 pl-reaction containing 5 pl of NEBuffer r3.1 (10X, NEB #B7203), 26 pl of adapter-ligated DNA and NF water to 50 pl. The reaction system was incubated at 37 °C for 30 min.
- BspCNI NEB, cat# R0624S
- 10 U of the enzyme was used in the 50 pl -reach on containing 5 pl of rCutSmart Buffer (10X, NEB # B6004S) for 30 min- incubation at 37 °C.
- the DNA was firstly digested with 20 U SfaNI in NEBuffer r3.1 for 30 min at 37 °C, then the digestion system was purified with 1.8X AMpure beads. The eluted DNA (20 pl) was further digested for 30 min at 37 °C with a 5 pl-mixture of Aflll, Acul, Bpml, Bpuel and Pstl-HF (1 pl of each).
- Step 4 Gel purification.
- the RE-digested products were loaded into a 1.5% agarose gel (Thermo Scientific, cat# R0492, MA, USA) pre-stained with SYBR Safe DNA gel stain (Invitrogen, cat# S33102, MA USA) for gel-based size selection.
- DNA ladder (NEB, cat#N3238S) was used as a comparison for size estimation and each gel was visualized under UV light on a ChemiDoc XRS+ System (Bio-Rad, CA, USA).
- the DNA band at 5 to lOkb size (as the DNA ladder indicates) was collected for extraction with a NucleoSpin Gel and PCR Clean-up purification kit following the manufacturer's instructions (Macherey-Nagel, cat# 740609.250, Germany), and was eluted with 15 pl NF water.
- Step 5 PCR amplification.
- the gel-purified DNA was amplified with the barcoded primers from the barcoding expansion pack mentioned above (EXP-PBC096, white caps, ONT).
- EXP-PBC096, white caps, ONT for PrimeSTAR GXL Premix (2x, Takara Bio, cat# R051A, Japan), a 50 pl -reaction system was prepared as following: 25 pl of Premix, 1 pl of barcoded primer (lOuM, white tube), 15 pl purified DNA and 9 pl water.
- the PCR reactions were pre-heated at 98°C for 1 min followed by 15-18 cycles of 98°C for 15 s, annealing 62°C for 15 s and extension at 68°C for 8 min, with a final extension at 68°C for 10 min on the ABI Veriti thermal cycler.
- the PCR products were purified with 0.5x volume of Ampure XP beads. Before moving to library prep, the purified PCR products were checked with an agarose gel for a distinct band at 5-10 kb in size (FIGS. 9A-9B, 10A-10B, 11A-11B, and 12A-12B).
- Step 6 Library prep and sequencing.
- the purified barcoded PCR products were pooled together for "DNA repair’ and “end-prep” step. Briefly. 47 pl of pooled PCR products were incubated at 20°C for 15min and 65°C for 15min, in a 60-reaction system containing 3.5 pl of NEBNext FFPE DNA Repair Buffer, 2 pl of NEBNext FFPE DNA Repair Mix, Ultra II End-prep reaction buffer 3.5 pl and Ultra II End-prep enzyme mix 3 pl.
- the purified DNA was ligated to Adapter Mix (ONT, SQK-LSK109) in the presence of Ligation Buffer (ONT, SQK- LSK109) and NEBNext Quick T4 Ligase (NEB, cat# E6056) for 30 min at room temperature. After beads purification, the adapter ligated DNA is ready for sequencing on ONT flow cells.
- the barcoded PCR products from ‘Step 5’ will firstly be purified and adjusted to shorter insert size (e.g., by sonication). Then, the fragmented DNA molecules are ready for users to prepare the library 7 with suitable kits compatible with various Illumina platforms. Compatible barcoding kits are required for a good sequencing result.
- the Barcode adapter sequence from the mEnrich-seq reads can be trimmed using computational methods for further analysis.
- Barcode adapter As the Barcode adapter ligation step is prior to the RE digestion during the library preparation, the DNA from targeted taxa will not be amplified by PCR if the barcode adapter is cut during the RE digestion step. Therefore, it’s critical to avoid a chosen RE cutting the sequence of the Barcode adapter or the junction between the adapter and metagenomic gDNA. It is suggested to check the sequence of the Barcode adapter (particularly the sequences that are reverse complementary to each other) for any possible RE motif site. If there’s a recognition sequence for the chosen RE, the bases can be substituted to avoid the cut. The sequences are as follows: top sequence: SEQ ID NO: 35 and bottom sequence: SEQ ID NO: 36.
- the top and bottom sequence are added to the annealing buffer (10 mM Tris, pH 7.5, 50 mM NaCl and 1 mM EDTA) to a concentration of 10 pM and annealed with the program: 95 °C for 2 min, -0. 1 C/sec for each of 800 s ( ⁇ 13 min) and hold at 4°C.
- the barcoded adapter is now ready to use and compatible with the original protocol.
- RE can be found through the Enzyme Finder (enzymefmder.neb.com) based on the chosen target motifs. For the chosen RE, it’s important to check its methylation sensitivity result provided by the manufacturer in REBASE (rebase.neb.com/rebase/rebase), which is commonly tested with modified oligos, and even better to re-confirm the sensitivity upon purchase.
- the PCR programs were as follows: pre-incubation at 95°C for 3 min, amplification for 14 cycles at 95°C for 15 s, 62°C for 15 s, and extension at 65°C for 10 min, with a final extension at 65°C for 10 min.
- the PCR products were checked with an agarose gel for successful amplification (bands at ⁇ 5 to 10 kb, FIGS. 9A-9B, 10A-10B, 11A-11B. and 12A-12B).
- the purified barcoded PCR products were used for “repair and end-prep” step in the ONT protocol.
- the library was quantified by Qubit and ready for sequencing.
- Two pairs of primers specific to tet(A) sequence were designed using Primer- BLAST as follows: te ⁇ -Fl: 5’- GAAACCCAACAGACCCCTGA-3’ (SEQ ID NO: 39), and tet(A)-R 5’- CCACCCGTTCCACGTTGTTA-3’(SEQ ID NO: 40); tet(A)-F2: 5’- TGAAACCCAACAGACCCCTG-3’(SEQ ID NO: 41), and tet(A)-R2 5’- CCCTGACGTTCCTCATCCAC-3’ (SEQ ID NO: 42) (Integrated DNA Technologies, IA, USA; 25 nmol, standard desalting).
- the tet(A) was amplified directly from the gDNA of UTI- 1 using the Q5 High-Fidelity 2X PCR Master Mix (NEB, cat #M0492S) with the following conditions: 50 ng of gDNA per reaction; 1 cycle at 95°C for 30s; 33 cycles of denature 95°C for 15s, annealing at 60 °C for 15s, extension at 72 °C for Imin; a final extension cycle at 72 °C for 5 min.
- PCR products were checked for amplification specificity on an agarose gel. Main bands of the PCR products were gel purified with NucleoSpin Gel and PCR Clean-up purification kit.
- the purified PCR products were sequenced by the Sanger method using tet(A)- F1 and tet(A)-V2 as forward primers._The two PCR products are shown in Table 3 below. Nucleotides denoted with an “n” in SEQ ID NO: 43 and 44 were undetermined by the Sanger method and may be any base (a,t,c, or g).
- the fold enrichment of E. coli was calculated based on E. coh reads from mEnrich-seq and Input samples that were mapped to the mock reference using minimap2 (v2.24-rl 122) with -x map-ont.
- the fold enrichment is defined as: target strain abundance (%) in enrichment sequencing versus target strain abundance (%) in the input sample.
- target strain abundance (%) in enrichment sequencing versus target strain abundance (%) in the input sample For genome completeness, the coverage of the E. coli was calculated using read depths across lOObp bins across the E. coli genome. For Circos plots. Nanopore reads from paired Input and mEnrich-seq samples were normalized to the same yield by random subsampling using seqtk (vl.2-r94) with 2-pass mode.
- rRNA ribosomal RNA
- muciniphila using minimap2 (v2.24-rl 122) with -x map-ont.
- sliding windows with intervals of 2.5kb and step sizes of 1.25kb were defined across the A. muciniphila isolate reference using bedtools (v2.29.2) with -w 2500 -s 1250.
- Reads mapped to defined windows were counted using bedtools multicov.
- read depths were capped at the 95% depth quantile of the whole genome for Circos plots drawing using circos (vO.69-6).
- RAATTY motif analysis For the species used for motif analysis, all genome references except A. muciniphila were retrieved fromNCBI, including C. aerofaciens (RefSeq genome sequence GCF_002736145.1), E. lenta (GCF_021378605.1), B. bifidum (GCF 000273525.1), B. longum (GCF_000196555.1), B. breve (GCF_001281425.1). The matched A. muciniphila reference was obtained from the isolated strain as previously described. Nanopore reads from the mEnrich-seq samples were mapped to the reference of each species using minimap2 (v2.24-rl 122) with -x map-ont.
- the raw SMRT reads pre-processing was performed with SMRTlink v8.0 (Pacific Biosciences).
- Raw multiplexed files were first de-multiplexed with lima.
- Circular Consensus reads (CCS reads) were generated using CCS software (v4.0.0) with --min-rq 0.99 based on subreads.
- CCS reads were assembled using metaFlye (v2.9) with -pacbio-hifi. The assembly was then binned and refined with the metaWRAP Binning and Bin refmement module 10 , which could combine the results from three metagenomic binning software - MaxBin2, metaBAT2, and CONCOCT. The completeness and the contamination level was checked using CheckM (v.1.1.3).
- the metagenomic profile computing and graphics were performed by R (v.4.2. 1).
- the genome assemblies from bacteria isolates from a previous study5 along with de novo assembled metagenomic contigs were used as the reference database for the following nanopore reads assignments and motif analyses.
- Nanopore reads w ere first mapped to the reference database described above using minimap2 (v2.24-rl 122) with -x map-ont. To reduce the possibility of low-quality mis-mapped reads from other species, non-primary and supplementary alignments were not included for reads assignment using samtools (vl.15.1) with view -F 2308. Other reads were then classified using Kraken2 (v2.0.8-beta) with the k2_pluspf_20200919 database.
- REBASE is the database constantly updated with bacterial methylomes mapped to date. For each bacterial methylome, a list of methylation motifs was reported (rebase.neb.com/rebase/rebase.html). All 4-mer to 6-mer methylation motifs were extracted and their corresponding restriction endonucleases (REs) from 4,644 bacterial methylomes (as of 7/18/2022). Among 4,644 PacBio organisms, the 1,358 most complete genomes were selected as representatives when multiple strains from the same species are present. The protein sequences from the 1,358 species level genomes were extracted to reconstruct a phylogenetic tree using PhyloPhlAn (v3.0). Visualization and annotation of the phylogenetic tree was performed by ITOL (itol.embl.de/).
- REMoDE libraries were prepared as previously described by Enam et al., (Restriction Endonuclease-Based Modification-Dependent Enrichment (REMoDE) of DNA for Metagenomic Sequencing. Appl Environ Microbiol 89, e01670-22 (2023), which is incorporated herein by reference in its entirety ). Briefly, 100 ng of DNA was used for the reaction mixture of 40 pL containing 10X Tango buffer and 1 pL of XapI (Fisher Scientific, ER1381). The mixture was incubated at 37°C for 30 mm.
- T5 exonuclease (NEB, #M0663) was added to each reaction mixture and incubated for 5 min at 37°C, after which the reaction w as immediately quenched with 8 pL 66 mM EDTA.
- REMoDE-repair For the library preparation of REMoDE with DNA repair (“REMoDE-repair’’), the DNA was first treated with the NEBNext FFPE DNA Repair Mix at 20°C for 15 min, purified with IX AMPure beads, before following the protocol described above (starting at XapI digestion).
- NGS libraries were prepared using Nextera XT (Illumina, cat # FC-131-1024) library preparation kit. The amplified libraries were resolved on an agarose gel and quantified with Qubit HS dsDNA kit. The libraries were pooled and sequenced on an Illumina MiSeq sequencer using a MiSeq reagent kit v2 (Illumina, cat# MS-102-2002; 300 cycles, paired-end). [0142] Data analysis.
- the paired-end short-read data of Input, REMoDE and mEnrich were each mapped to the reference using bwa with default parameters, and then reads mapped to each species were counted.
- 1 ng of DNA was diluted to 3 pL and incubated with 1 pL of Fragmentation Mix at 30° C for 1 min and then at 80° C for 1 min.
- the tagmented DNA was added to PCR mix containing 20 pL nuclease-free water, 25 pL LongAmp Taq 2X master mix, 1 pL Rapid Barcode Primers (RLB01-12A, at 10 pM), followed by 18 cycles of amplification (18 cycles of 95 °C for 15 s; 56 °C for 15 s; 65 °C for 6 min; 1 cycle of extension at 65 °C for 6 min).
- PCR products were purified with 0.6X AMPure Beads and eluted with 10 pl elution buffer (10 mM Tris-HCl pH 8.0 with 50 mM NaCl). 50-100 ftnoles of library pool was ready for sequencing after ligation with 1 pL Rapid Adapter for 5 min at room temperature.
- mEnrich-Dl the original mEnrich-seq
- mEnrich-D2 the standard ONT library without enrichment
- Input standard ONT library without enrichment
- Example 3 Methylation-guided enrichment for bacterial taxa of interest
- the core idea of enriching the gDNA of bacterial taxa of interest from the microbiome is to distinguish gDNA from different taxa based on their distinct characteristics.
- a methylation-guided enrichment sequencing of bacterial taxa of interest from microbiomes is shown and named mEnrich-seq (in which ‘m’ stands for methylation and seq for sequencing, FIG. 1A).
- m stands for methylation and seq for sequencing
- the gDNA of the bacterial taxa of interest (including its chromosome and mobile genetic elements, for example, plasmids) is left intact, enriched over the background, which can be sequenced by various sequencing platforms (FIG. 1A). More specifically, g-TUBEs are first used to shear the DNA size to around 10 kb, which facilitate the second step where DNA is ligated to barcode adapters. In the third step, cognate restriction enzymes (RE) corresponding to the methylation motif in the taxon of interest are used to cut the background DNA libraries into smaller fragments while preserving the targeted DNA libraries as the long fragments.
- RE restriction enzymes
- the choice of fragmentation size in the shearing step increases the chance that most gDNA molecules have at least one nonmethylated RE site that can be cut: 5-10 kb fragments expected to have 20-40 sites of a 4-mer motif, 5-10 sites of a 5-mer motif and 1-2.5 sites of a 6-mer motif.
- gel-based size selection is performed, which largely removes the digested shorter DNA fragments.
- step 2 barcoded primers are used to anchor the PCR adapter to perform amplification of the long target DNA fragments that have been protected from RE digestion (i.e., those that are methylated at restriction sites).
- step 3 the DNA-to-adapter ligation step
- step 3 only nondigested DNA libraries can be sequenced. This design minimizes allocation of sequencing throughput to background DNA fragments that are still long after RE digestion due to sparse distribution of the RE recognition motif(s) in certain genomes or genomic regions.
- mEnrich-seq is versatile along three dimensions.
- RE(s) can be tailored to enrich bacterial taxa at different taxonomic levels, which can be helpful depending on specific research goals (FIG. IB).
- the methylation-sensitive RE DpnII (which cuts doublestranded DNA (dsDNA) at nonmethylated GATC sites, but is sensitive to methylated G6mATC) was chosen for mEnrich-seq to selectively enrich E. coli gDNA while digesting N. gonorrhoeae (only 35 of the 2,434 GATC sites are methylated owing to the overlapping with GGTG6mA methylation motif) and human gDNA with hardly detectable methylated GATC sites.
- the library was prepared for Nanopore long-read sequencing and generated 200 Mb of data (mean read length of roughly 5 kb Table 5).
- the mock samples were also sequenced without mEnrich-seq protocol, labeled as ‘input’ (standard metagenomics).
- Table 5 summarizes the Nanopore sequencing dataset for the Mock-1 and Mock-2. Data were produced by Flongle flow cell. The read-length difference is due to the gel-based size selection step in the gel-based size selection step in mEnrich-seq.
- mEnrich-seq greatly increased the proportion (%) of reads mapped to the E. coli genome from 0.91 to 70.75% and from 1 .20 to 71 .41% in the two mock replicates, respectively (FIG. 2A): roughly 70-fold enrichment (FIG. 2B).
- mEnrich-seq reads mapping showed that the E. coli genome can be covered without systematic bias for both mock samples (more than 99.96% of genome covered; FIG. 2C). Also, 3.98 and 4.12% of reads from mEnrich-seq are mapped to N. gonorrhoeae.
- GATC sites are located within genomic regions with simple sequence repeats (Example 2, above) that tend to form secondary structure, making it less accessible to DpnII digestion.
- This characterization illustrates that mEnrich-seq not only enriches bacterial taxa of interest, but also carries some by-product reads that are either due to methylation at (partially) overlapping motif sites in background taxa or depleted restriction motifs in a subset of gDNA molecules.
- the highlight of mEnrich-seq is the enrichment efficiency compared to standard metagenomic sequencing despite by-product reads.
- mEnrich- seq was further tested on three urine samples (UTI-1-3) from patients with urinary tract infection (UTI) (E. coli positive urine culture, defined as more than or equal to 100.000 colony forming units per ml; see Example 2, above).
- UTI urinary tract infection
- the three UTI samples were sequenced with mEnrich-seq (with DpnII as RE) coupled with Nanopore sequencing.
- mEnrich-seq 1.16— 1.58 GB of sequencing data was generated with an average of 233,000 reads per sample (mean read length roughly 6 kb, Tables 6A-6B, below).
- coli reads was further confirmed by mapping each read manually to the National Center for Biotechnology (NCBI) Nucleotide database, and further validated by PCR (FIG. 7B) and Sanger sequencing (as described in Example 2, above). This observation is consistent with the increasing recognition that a culture-independent approach may provide additional information that can complement standard urine culture in monitoring pathogen genomes in UTI and other infectious diseases.
- NCBI National Center for Biotechnology
- FIG. 7B Sanger sequencing
- Tables 6A-6B show the summary of the Nanopore sequencing dataset for urine samples.
- Table 6A illustrates the input, isolates and mEnrich-seq of sample UTI-1 to -3 as first sequenced with a MinlON flow cell.
- Table 6B illustrates the three urine samples (UTI- 1 to -3) together with the additional three UTI urine samples (UT1-4 to -6) as sequenced on a Flongle flow cell for validation purposes.
- the efficient enrichment of E. coll genomes by mEnrich-seq can facilitate the culture-free study of E. coli genomes from urine microbiome with better sensitivity (examining E. coli with lower relative abundance) than standard metagenomics.
- mEnrich-seq can facilitate a more cost-effective examination of urine samples in the study of E. coli genomes.
- 6mA at GATC sites are also conserved in most Gammaproteobacteria including many enteric pathogens, which makes mEnrich-seq with DpnII applicable to enrich additional enteric pathogens and commensal bacteria as are described in the following Examples.
- RA6mATTY mediated by a 6mA MTase, AmuORF1905P, is one of the five methylated motifs in A. muciniphila (ATCC BAA-835) according to REBASE (The Restriction Enzyme Database) (FIG. 8A).
- mEnrich-seq was applied with XapI to enrich the muciniphila genome from three infant fecal samples (GUT-1-3) (Table 7. below), and observed a 20.0-fold increase of reads mapped to A. muciniphila in GUT-1 (from 0.48 to 9.59%). a 27.2-fold increase in GUT-2 (from 0.52 to 14.16%) and an 18.3-fold increase in GUT-3 (from 1.52 to 27.88%) (FIG. 3A-3B).
- the RAATTY frequency is comparable between the mEnrich-seq reads and RAATTY frequency of the reference genome (FIG. 3D), which is consistent with the enrichment due to methylated RA6mATTY motif on its genome.
- the frequency of RAATTY in the genomes of B. bifidum, C. aerofaciens and E. lenta is lower, which further decreased among mEnrich-seq reads that mapped to these three species (FIG. 3D).
- These three genomes with sparse frequency of RAATTY sites were enriched as by-products because of the more efficient depletion of most of the other background genomes with dense yet nonmethylated RAATTY sites.
- mEnrich-seq with XapI as the methylation-sensitive RE can efficiently enrich 4.
- muciniphila highly conserved RA6mATTY methylation across all the isolated strains examined
- mEnrich-seq can be a useful tool to study the genomes of A. muciniphila in a culture-independent, sensitive and cost-effective way, which may facilitate larger-scale association studies with different human diseases.
- mEnrich-seq is demonstrated when the methylation motifs of target bacteria of interest are known a priori.
- mEnrich- seq is used to enrich gDNA from low-abundance bacteria from a complex microbiome sample based on methylation motifs discovered de novo.
- GUT-5 an adult fecal microbiome sample (GUT-5) was built that has been recently well characterized with comprehensive cultured isolates.
- a pilot standard metagenomic sequencing of this adult fecal sample was analyzed (Nanopore sequencing, 7.08 Gb; further details in Example 2 above).
- Table 8 provides the summary of this Nanopore sequencing data for low- abundance bacterial taxa enrichment with de novo discovered methylation motifs.
- the top portion summanzes data from an adult fecal sample sequenced with the MiNION flow cell for the initial standard metagenomic profiling.
- the bottom portion shows data for the low abundance bacterial taxa enrichment by mEnrich-seq, sequenced with a Flongle flow cell.
- ovatus D3 was also enriched (4.3-fold, from 2.97 to 12.94%), which has a CTCAG frequency similar to A. finegoldii H10 and B. ovatus C6, reflecting that B. ovatus D3 has methylation at CTCAG sites that resist BspCNI digestion (FIG. 4D).
- Bifidobacterium longum F12 and B. longum DIO have the same methylation motif RG6mATCY.
- mEnrich-seq with Mfll achieved an 8.7-fold enrichment for B. longum F 12 (from 1.17 to 10.17%) and a sevenfold enrichment for B. longum D10 (from 0.32 to 2.24%) (FIG. 4E).
- enrichment was observed for Ruminococcus bromii (4.5-fold, from 2.58 to 11.69%), Eubacterium rectale G9 (3.8-fold, from 2.15 to 8.20%) and Bacteroides uniformis D7 (2.6-fold, from 2.69 to 6.89%; FIG. 4E).
- Further examination of RGATCY motif frequency supported that R. bromii, E. rectale G9 and B uniformis D7 were all enriched owing to methylation at RG6mATCY sites on their genomes (FIG. 4F).
- B. intestinihominis owing to the methylated G6mATGC motif on their genomes (FIG. 4H).
- SfaNI targeting GATGC
- this Example shows that mEnrich-seq can be tailored based on the de novo discovered methylation motifs from adult fecal microbiome samples to enrich bacterial strains with relatively low abundance in the sample.
- the mEnrich-seq strategy reduces the overall cost of recovering genomes from species with modest-to-low abundance, whose genomes would otherwise be cost prohibitive to resolve using standard metagenomics.
- mEnrich-seq significantly reduces sequencing reads from background bacteria, it can simplify the assembly of the enriched bacterial genomes.
- methylomes were compared with a list of 584 methylation-sensitive REs targeting 209 motifs that are commercially available (see Table 16, at end of the Examples) and found that 3,130 (68.03%) of the 4,601 strains have at least one (26.62% have two or more) methylation motif(s) that can be targeted (enriched by depleting the background metagenomic DNA) by commercially available RE(s).
- 3,130 strains When projected onto a phylogenetic tree (see Example 2, above), these 3,130 strains have broad taxonomic diversity (FIG. 5C), representing 54.78% of the species examined in this analysis.
- This analysis focused on REs with recognition motifs that have four, five or six nondegenerate bases, as these REs are expected to be readily applicable considering the expected frequency of their target motifs across bacterial genomes: roughly 256 bp for 4-mers, roughly 1 kb for 5-mers and roughly 4 kb for 6-mers. As researchers are actively developing more advanced DNA extraction methods that preserve high molecular weight gDNA from microbiome samples, it is expected that mEnrich-seq may expand to REs with additional target motifs with lower frequency.
- a methylation motif can be conserved at different taxonomic levels. Some methylation motifs are conserved at the family level. For example, the G6mATC motif by the Dam DNA methyltransferase is highly conserved in many families of the Gammaproteobacteria class, which includes many of the enteric pathogens (for example, E. coli, Salmonella enterica, Vibrio cholerae and so on) as well as some nonpathogenic species across many families such as Shewanellaceae and Moraxellaceae and so on. Some methylation motifs are conserved at the genus level.
- REs have very broad utility as their target motifs are methylated in a large number of bacterial genomes (FIG. 5D-5E). Taking DpnII (GATC) and XapI (RAATTY) as examples, the previous examples show their utilities with mEnrich-seq in two specific applications enriching for E. coli and A. muciniphila, respectively.
- mEnrich-seq with DpnII can help enrich many bacteria in the Gammaproteobacteria class in which G6mATC methylation is highly conserved
- mEnrich-seq with XapI can help enrich other bacteria such as Campylobacter spp., Acinetobacter spp., Spirochaeta spp., Treponema spp. and Brachyspira spp.
- the discriminative power of methylation motifs in practical applications is with respect to the bacteria in a specific microbiome sample, not among all the genomes used in the phylogenetic analysis.
- mEnrich-seq The core idea of mEnrich-seq is essentially two-fold: (i) methylation-guided digestions/enrichment followed by (ii) size selection.
- mEnrich-seq there are multiple ways to implement mEnrich-seq.
- the design described throughout the previous Examples performs adapter ligation before digestion ('ABD method’, uses gel purification for size selection, and utilizes the adapter sequence for subsequent PCR.
- AAD method adapter ligation after digestion
- Tn5-based adapter tagmentation T5 exonuclease to preferentially deplete shorter DNA fragments after digestion as described in a recent work called REMoDE (Enam, S. U.
- This library was the “Ligation-after-RE library” and it was directly compared to a standard mEnrich-seq library (500-100 ng as input) that was generated for evaluation of the adaptor ligation before RE digestion (“Ligation-before-RE”).
- the ONT data of Ligation-before-RE or Ligation-after-RE were each mapped to the reference using minimap2 (v2.24-rl !22) with -x map-ont and reads mapped to each species were counted.
- FIG. 16 and summarized in Table 11, below, adaptor-ligation before RE resulted in greater enrichment than adaptor-ligation after RE.
- FIGS 14A-14B, 15A-15B and 16 highlight the advantages of adapter ligation before digestion, it is worth noting that a single adapter sequence is not compatible with all the REs, because some enzymes can digest the adapter.
- the adapter sequence used in the previous Examples e.g., having an upper oligonucleotide sequence of SEQ ID NO: 35 and a lower oligonucleotide sequence of SEQ ID NO: 36, as described in Example 2, above
- 556 of the 584 (95%) commercially available REs is compatible with 556 of the 584 (95%) commercially available REs.
- one additional adapter was synthesized and found to be compatible with the remaining 28 commercially available REs.
- This second “alternative” adaptor contained a top oligo having a sequence of SEQ ID NO: 37 and a bottom oligo having a sequence of SEQ ID NO: 38. So, all the 584 commercially available REs are compatible with at least one of the two adapter oligos (FIG. 17).
- Example 9 Complementarity between mEnrich-seq and adaptive sampling.
- Adaptive sampling is effective in reducing the allocation of sequencing yield to host gDNA or highly abundant bacteria.
- desired taxa are indirectly enriched by the rejection of reads from highly abundant species, thus the efficiency enrichment by adaptive sampling from latest studies is still modest and insufficient for low-abundant bacteria.
- mEnrich-seq has the advantage of more efficiently depleting the vast background gDNA before sequencing, hence achieving high fold enrichments.
- Adaptive sampling can also select for specific sequences of interest (such as a desired genome): however, it does so by requiring a pre-defined genome sequence hence it will likely miss genes unique to a specific strain yet to be determined.
- mEnrich-seq has the advantage of enriching both known and unknown genes in a strain of interest, including mobile genetic elements, because all the genetic contents within a bacterial cell share the same methylation motifs. So, mEnrich-seq is complementary to the adaptive sampling method on the Nanopore sequencing platform.
- Bacterial DNA methylation provides a natural way for differentiating diverse bacterial taxa from each other. Although it has been used for metagenomic binning, it is only applicable as a post hoc analysis, after single molecule real-time (SMRT) sequencing or Nanopore sequencing. By contrast, mEnrich-seq enables the use of bacterial DNA methylation to enrich certain taxa of interest before sequencing, which allows microbiome researchers to tailor sequencing strategy to best serve their goals.
- SMRT single molecule real-time
- mEnrich-seq is directly applicable and is expected to facilitate many applications.
- the Examples herein demonstrate the use of mEnrich-seq to enrich E. coli from urine samples and to enrich A. muciniphila from fecal samples.
- they also demonstrate the use of mEnrich-seq along with de novo methylation discovers’ to enrich several low- abundance bacteria from a complex fecal microbiome sample.
- mEnrich-seq is broadly applicable and versatile. Based on a meta-analysis of the 4,601 bacterial methylomes mapped to date, it is estimated that roughly 68% of bacterial genomes can be targeted by mEnrich-seq with at least one RE, representing 54.78% of the species examined. In practice, mEnrich-seq may be applicable to a broader diversity of bacterial genomes, because epigenomes mapped to date represent only a small fraction of the bacterial diversity. mEnrich-seq is also very versatile in that it can be used along with one or multiple RE(s) (FIG. IB).
- mEnrich-seq provides a flexible way to dissect a microbiome sample when it is coupled with different REs, the choice of which can be tailored based on taxa of interest in a particular application.
- mEnrich-seq is compatible with different long-read and short-sequencing platforms (FIGs. 14A-14B and FIGs. 15A- 15B).
- the core idea of mEnrich-seq is essentially twofold: (1) methylation-guided digestions and/or enrichment follow ed by (2) size selection. Along w ith this core idea, there are multiple ways to implement mEnrich-seq.
- mEnrich-seq also has some limitations.
- mEnrich-seq is its efficacy in the enrichment of taxa of interest over standard metagenomics.
- mEnrich-seq is primarily applicable to enrich bacteria, archaea and dsDNA phages, but not applicable to single-stranded (ssDNA) phages, RNA phages and parasites, because they are not substrates of restriction-modification systems. It is worth noting a recent work leveraged the size differences between host and bacterial cells and used mechanical stress to effectively deplete host cells, which could be integrated with mEnrich-seq to study targeted taxa in host-rich samples.
- mEnrich-seq is a versatile and cost-effective approach for enrichment sequencing of diverse microbial taxa of interest directly from the microbiome. It provides microbiome researchers an additional approach to more effectively study complex microbiomes.
- Tables 14A-14B The MTase mediating RAATTY methylation is conserved across 1 12 genomes of A. muciniphila isolates.
- the amino acid sequence of the MTase (methylating RAATTY) is used as a query to blast against the 112 reference genomes.
- the prevalence of the MTase in 112 Akkermansia isolates is 100%, when using the following filter criteria for the BLAST hits: bitscores > 100 & identity > 80% & alignment length > 300.
- Table 15 The list of methylation motifs de novo discovered from the adult fecal metagenomic sample, de novo methylation discovery from microbiome samples was performed using existing tools (see Example 2). 112 methylation motifs were found across 34 bins (from 26 species) with lower abundance (Relative abundance profiled by Nanopore is indicated below). The bold motifs were the 4-mer or 5-mer motifs selected as the targets for illustrative purposes.
Landscapes
- Chemical & Material Sciences (AREA)
- Organic Chemistry (AREA)
- Life Sciences & Earth Sciences (AREA)
- Analytical Chemistry (AREA)
- Zoology (AREA)
- Wood Science & Technology (AREA)
- Proteomics, Peptides & Aminoacids (AREA)
- Health & Medical Sciences (AREA)
- Engineering & Computer Science (AREA)
- Biophysics (AREA)
- Immunology (AREA)
- Microbiology (AREA)
- Molecular Biology (AREA)
- Biotechnology (AREA)
- Physics & Mathematics (AREA)
- Chemical Kinetics & Catalysis (AREA)
- Biochemistry (AREA)
- Bioinformatics & Cheminformatics (AREA)
- General Engineering & Computer Science (AREA)
- General Health & Medical Sciences (AREA)
- Genetics & Genomics (AREA)
- Measuring Or Testing Involving Enzymes Or Micro-Organisms (AREA)
Abstract
L'invention propose des procédés d'amplification sélective d'un acide nucléique d'un ou de plusieurs organismes procaryotes d'intérêt (organismes cibles) dans un échantillon de microbiome. Les procédés comprennent généralement le marquage de fragments d'ADN génomique dans un échantillon, puis la soumission de fragments d'ADN marqués à une digestion d'enzyme de restriction à l'aide d'enzymes de restriction sensibles à la méthylation choisies pour enrichir sélectivement des acides nucléiques du ou des organismes procaryotes d'intérêt dans l'échantillon de microbiome.
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US202263382616P | 2022-11-07 | 2022-11-07 | |
| PCT/US2024/010215 WO2024103078A2 (fr) | 2022-11-07 | 2024-01-03 | Séquençage métagénomique à lecture longue rentable |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| EP4627111A2 true EP4627111A2 (fr) | 2025-10-08 |
Family
ID=91033510
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| EP24725039.2A Pending EP4627111A2 (fr) | 2022-11-07 | 2024-01-03 | Séquençage métagénomique à lecture longue rentable |
Country Status (2)
| Country | Link |
|---|---|
| EP (1) | EP4627111A2 (fr) |
| WO (1) | WO2024103078A2 (fr) |
Family Cites Families (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP4189495B2 (ja) * | 2004-08-02 | 2008-12-03 | 国立大学法人群馬大学 | ゲノムdnaのメチル化検出方法 |
| CA3131514A1 (fr) * | 2019-02-25 | 2020-09-03 | Twist Bioscience Corporation | Compositions et procedes de sequencage de nouvelle generation |
| JP2020191833A (ja) * | 2019-05-29 | 2020-12-03 | 森永乳業株式会社 | 細菌集団中の標的細菌のdnaを選択的に増幅する方法 |
| US20230212684A1 (en) * | 2020-05-05 | 2023-07-06 | The Board Of Trustees Of The Leland Stanford Junior University | Cell-free dna biomarkers and their use in diagnosis, monitoring response to therapy, and selection of therapy for prostate cancer |
| US20220135965A1 (en) * | 2020-10-26 | 2022-05-05 | Twist Bioscience Corporation | Libraries for next generation sequencing |
-
2024
- 2024-01-03 EP EP24725039.2A patent/EP4627111A2/fr active Pending
- 2024-01-03 WO PCT/US2024/010215 patent/WO2024103078A2/fr not_active Ceased
Also Published As
| Publication number | Publication date |
|---|---|
| WO2024103078A8 (fr) | 2024-06-20 |
| WO2024103078A2 (fr) | 2024-05-16 |
| WO2024103078A3 (fr) | 2024-07-18 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US20250109426A1 (en) | Compositions and methods for targeted depletion, enrichment, and partitioning of nucleic acids using crispr/cas system proteins | |
| Parkinson et al. | Preparation of high-quality next-generation sequencing libraries from picogram quantities of target DNA | |
| US9850525B2 (en) | CAS9-based isothermal method of detection of specific DNA sequence | |
| US12416003B2 (en) | Methods and compositions for enrichment of target polynucleotides | |
| JP7436493B2 (ja) | 次世代シーケンスにおいて低サンプルインプットを扱うための正規化対照 | |
| AU2018256358B2 (en) | Nucleic acid characteristics as guides for sequence assembly | |
| US20230279385A1 (en) | Sequence-Specific Targeted Transposition and Selection and Sorting of Nucleic Acids | |
| JP5116481B2 (ja) | シトシンの化学修飾により微生物核酸を簡素化するための方法 | |
| WO2023148235A1 (fr) | Procédés d'enrichissement d'acides nucléiques | |
| JP2024541111A (ja) | シーケンシングライブラリ構築方法と使用 | |
| US20230068726A1 (en) | Transposon systems for genome editing | |
| EP4627111A2 (fr) | Séquençage métagénomique à lecture longue rentable | |
| Stocks | Transposon Mediated Genetic Modification of Gram-positive Bacteria |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: THE INTERNATIONAL PUBLICATION HAS BEEN MADE |
|
| PUAI | Public reference made under article 153(3) epc to a published international application that has entered the european phase |
Free format text: ORIGINAL CODE: 0009012 |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: REQUEST FOR EXAMINATION WAS MADE |
|
| 17P | Request for examination filed |
Effective date: 20250822 |
|
| AK | Designated contracting states |
Kind code of ref document: A2 Designated state(s): AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC ME MK MT NL NO PL PT RO RS SE SI SK SM TR |