WO2025007023A2 - Compositions, constructions, vecteurs viraux végétaux et procédés d'édition de gènes végétaux - Google Patents

Compositions, constructions, vecteurs viraux végétaux et procédés d'édition de gènes végétaux Download PDF

Info

Publication number
WO2025007023A2
WO2025007023A2 PCT/US2024/036209 US2024036209W WO2025007023A2 WO 2025007023 A2 WO2025007023 A2 WO 2025007023A2 US 2024036209 W US2024036209 W US 2024036209W WO 2025007023 A2 WO2025007023 A2 WO 2025007023A2
Authority
WO
WIPO (PCT)
Prior art keywords
tnpb
sequence
plant
editing
seq
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/US2024/036209
Other languages
English (en)
Other versions
WO2025007023A3 (fr
Inventor
Steven Erik JACOBSEN
Zheng Li
Trevor John WEISS
Maris KAMALU
Jasmine AMERASEKERA
Jennifer DOUDNA
Benjamin ADLER
Honglue SHI
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
University of California Berkeley
University of California San Diego UCSD
Original Assignee
University of California Berkeley
University of California San Diego UCSD
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by University of California Berkeley, University of California San Diego UCSD filed Critical University of California Berkeley
Priority to EP24833079.7A priority Critical patent/EP4735609A2/fr
Publication of WO2025007023A2 publication Critical patent/WO2025007023A2/fr
Publication of WO2025007023A3 publication Critical patent/WO2025007023A3/fr
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Classifications

    • CCHEMISTRY; METALLURGY
    • C12BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12NMICROORGANISMS OR ENZYMES; COMPOSITIONS THEREOF; PROPAGATING, PRESERVING, OR MAINTAINING MICROORGANISMS; MUTATION OR GENETIC ENGINEERING; CULTURE MEDIA
    • C12N9/00Enzymes; Proenzymes; Compositions thereof; Processes for preparing, activating, inhibiting, separating or purifying enzymes
    • C12N9/14Hydrolases (3)
    • C12N9/16Hydrolases (3) acting on ester bonds (3.1)
    • C12N9/22Ribonucleases [RNase]; Deoxyribonucleases [DNase]
    • CCHEMISTRY; METALLURGY
    • C12BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12NMICROORGANISMS OR ENZYMES; COMPOSITIONS THEREOF; PROPAGATING, PRESERVING, OR MAINTAINING MICROORGANISMS; MUTATION OR GENETIC ENGINEERING; CULTURE MEDIA
    • C12N15/00Mutation or genetic engineering; DNA or RNA concerning genetic engineering, vectors, e.g. plasmids, or their isolation, preparation or purification; Use of hosts therefor
    • C12N15/09Recombinant DNA-technology
    • C12N15/63Introduction of foreign genetic material using vectors; Vectors; Use of hosts therefor; Regulation of expression
    • C12N15/79Vectors or expression systems specially adapted for eukaryotic hosts
    • C12N15/82Vectors or expression systems specially adapted for eukaryotic hosts for plant cells, e.g. plant artificial chromosomes (PACs)
    • C12N15/8201Methods for introducing genetic material into plant cells, e.g. DNA, RNA, stable or transient incorporation, tissue culture methods adapted for transformation
    • C12N15/8213Targeted insertion of genes into the plant genome by homologous recombination

Definitions

  • a composition in one aspect, can include a TnpB bacterial enzy me.
  • the composition can also include a polynucleotide including a first portion adapted to bind to at least a portion of the TnpB bacterial enzyme and a second portion for binding to a plant genomic target.
  • a construct in another aspect, can include a plant promotor operably connected to a polynucleotide encoding for a TnpB bacterial enzyme.
  • a plant viral vector is provided.
  • the plant viral vector can include a construct that can include a plant promotor operably connected to a polynucleotide encoding for a TnpB bacterial enzyme.
  • a plant cell in another aspect, can include a construct that can include a plant promotor operably connected to a polynucleotide encoding for a TnpB bacterial enzyme, or a plant viral vector that includes the construct.
  • a method for gene editing a plant can include introducing a construct or a plant viral vector into one or more plant cells.
  • the construct can include a plant promotor operably connected to a polynucleotide encoding for a TnpB bacterial enzyme.
  • the plan viral vector can include the construct.
  • FIGS. 1 A and IB The '‘DOTS’’ pipeline for detecting bona-fide RNA- guided TnpB systems in metagenomes.
  • FIG. IB A phylogenetic tree of TnpBs from metagenomes (red) as well as a public IS database, ISfinder (black), with the corresponding amino acid length shown in the bar graph below. Distinct TAM sequences and small-sized TnpB enzymes from metagenomes are also highlighted in red circles and black arrows, respectively.
  • FIGS. 2A-2C Representative loci architectures for TnpB-associated systems characterized in this study.
  • FIG. 2B The Web logo of the TAMs for ISBagOl and ISXfaOl TnpB systems.
  • FIG. 2C Comparable E. coli interference activities between Casl2a, ISBagOl and ISXfaOl, when the target site is flanked by a PAM/TAM at 26°C.
  • FIGS. 3A-3C Schematics of the constructs used to test gene editing activity by ISBagOl.
  • Top panel ISBagOl coRNA with a 20bp spacer targeting the AtPDS3 gene were driven by the AtU6-26 promoter and followed by the HDV ribozyme.
  • 2X indicates that the construct also contains a second similar AtU6-26 driven coRNA cassette with a second gRNA sequence that is not shown.
  • Zea mays codon optimized ISBagOl coding sequence was expressed in a separate transcription cassette driven by the UBQ10 gene promoter and an rbcS- E9 terminator.
  • FIG. 3B Left panel, editing efficiency of ISBagOl at the AtPDS3 gene at regions targeted by four different guide RNAs in Arabidopsis protoplasts, with ISBagOl protein coding sequence and coRNA coding sequence in split cassettes.
  • FIG. 3C Examples of detailed editing profiles by ISBagOl in Arabidopsis protoplasts detected by amplicon sequencing.
  • Editing events are shown on the left as: (position where the editing starts): (number of nucleotides of)D(deletion).
  • position 0 is between the 18th and 19th nucleotides of the guide RNA sequence, such that the 18th nucleotide is position -1, the 19th nucleotide is position +1.
  • AtPDS3 gRNAl reference sequence is SEQ ID NO: 610, amplicons in descending order are SEQ ID NOs: 61 1 -620.
  • AtPDS3 gRNA18 reference sequence is SEQ ID NO: 621 , amplicons in descending order are SEQ ID NOs: 622-625.
  • FIG. 4A and 4B Schematics of a bacteria selection assay.
  • FIG. 4A The E.coli cell harbors the selection plasmid with the inducible toxic gene induced by Arabinose. The cell can survive in the absence of Arabinose but not in the presence of Arabinose.
  • FIG. 4B Transforming another plasmid encoding TnpB and coRNA targeting the toxic gene will result in survival of the cell.
  • FIGS. 5A-5C The distribution of amino acid lengths of Cas9, Casl2, IscB and TnpB endonucleases.
  • FIG. 5B Representative loci architecture of CRISPR-Cas systems versus TnpB-associated transposon systems
  • FIG. 5C Molecular mechanism of DNA targeting by CRISPR-Cas enzymes and TnpB-coRNA.
  • FIG. 6 An example of the split expression cassette to express the coRNA and the TnpB as described in Example 8.
  • the UBQ10 promoter is used to drive expression of the TnpB, followed by the rbcS-E9 terminator.
  • Two SV40 NLS signals and two FLAG tags are located at the 3’ end, however these sequences may also be tested at the 5’ end of the TnpB.
  • the coRNA is being driven by the U6 promoter, with an HDV ribozyme sequence downstream of the spacer; however, a variety of promoters and RNA processing sequences may be tested.
  • FIG. 7A and 7B An example schematic of the TnpB and coRNA architecture commonly found in nature, where the coRNA overlaps with the TnpB at the 3’ end of the TnpB sequence.
  • FIG. 7B An example of a single expression cassette to express the TnpB and coRNA.
  • the UBQ10 promoter is used to drive expression of the TnpB-coRNA sequence, and terminated by the rbcS-E9 terminator.
  • the HDV ribozyme sequence downstream of the spacer is depicted, however a variety of RNA processing systems such as tRNA-Gly may be used.
  • FIG. 8 An example of the single expression cassette to express the TnpB and coRNA, followed by a tandem repeat of the coRNA and spacer sequences.
  • the UBQ10 promoter is used to drive expression of the TnpB-coRNA sequence, and terminated by the rbcS-E9 terminator.
  • the HDV ribozyme sequence upstream of the terminator is depicted, however a variety of RNA processing systems such as tRNA-Gly may be used.
  • FIGS. 9A-9D TRV2 schematic depicting the RNA virus sequence.
  • the red box indicates the part of the virus where TnpB and coRNA will most likely be cloned.
  • the light gray boxes indicate the promoter and terminator used to drive initial expression of the RNA virus.
  • the light blue boxes indicate the RNA virus.
  • FIG. 9A TnpB-ojRNA single expression cassette cloned into the “cargo’’ site.
  • FIG. 9B TnpB and coRNA split expression design.
  • FIG. 9C TRV2 schematic showing an example of a RNA processing sequence, HDV ribozyme.
  • FIG. 9D TnpB and tandem coRNA cloned into the “cargo” design.
  • FIG. 10A-10C ISYmul, ISDra2 and ISAaml editing frequency.
  • the target sites are plotted along the X-axis and the editing efficiency (percent indel reads (%)) are plotted on the Y-axis.
  • additional transfections w ere performed to replicate initial results, as indicated by “experiment_2” and “experiment_3” along the X-axis. Each dot indicates a single transfection.
  • the standard error of the mean (SEM) was calculated for each target site.
  • FIGS. 11 Editing frequency for TnpBs capable of editing an average greater than 0.1% at least one target site.
  • the target sites are plotted along the X-axis and the editing efficiency (percent indel reads (%)) are plotted on the Y-axis. Each dot indicates a single transfection. The standard error of the mean (SEM) was calculated for each target site.
  • FIG. 12A-12D Small RNA-Seq and structure analysis of ISBagOl and ISYmul ®RNA sequences.
  • FIG. 12A and FIG. 12B Small RNA-Seq results for ISYmul and ISBagOl with the read count plotted along the Y-axis, and the TnpB and coRNA positions in the bacterial expression vectors plotted along the X-axis.
  • ISYmul SEQ ID NO: 626) and ISBagOl (SEQ ID NO: 627).
  • FIG. 13A and 13B ISBagOl ®RNA variants for improved editing efficiency.
  • FIG. 13A coRNA (WT, SEQ ID NO: 628) variant structure for 1.1 (SEQ ID NO: 629) and 1.2 (SEQ ID NO: 630).
  • FIG. 13B ISBagOl gl was used in this experiment. The target sites are plotted along the X-axis and the editing efficiency (percent indel reads (%)) are plotted on the Y-axis.
  • the X-axis label gl_native indicates the TnpB with coRNA overlap Native DNA sequence; gl_variant_l. l and gl_variant_1.2 are novel roRNA variants. Each dot indicates a single transfection.
  • the standard error of the mean (SEM) was calculated for each target site. P value was calculated using a two-tailed t-Test assuming equal variances.
  • FIG. 14 Editing frequency for ISYmul using TRV vectors for genome editing in Arabidopsis protoplast cells.
  • the ISYmul TRV2 cargo configurations are plotted along the X- axis and the editing efficiency (percent indel reads (%)) are plotted on the Y-axis.
  • the vertical dashed line separates experiment one and two. Each dot indicates a single transfection.
  • the standard error mean (SEM) was calculated for each target site.
  • FIG. 15A-15G Gene editing outcomes in transgenic T1 plants expressing ISYmul or ISDra2 TnpB.
  • FIG. 15A-15D The genotypes are plotted along the X-axis and the editing efficiency (percent indel reads (%)) are plotted on the Y-axis. Each dot indicates an independent T1 plant. The standard error of the mean (SEM) was calculated for each target site. The number above each bar indicates the average editing frequency.
  • WT is an abbreviation for wild type
  • rdr6 is an abbreviation for the mutant in the RNA-DEP ENDENT RNA POLYMERASE 6 gene.
  • ISDra2 g!2 WT reference sequence is SEQ ID NO: 638, amplicons in descending order are SEQ ID NO 639-640.
  • ISYmul g!2 WT reference sequence is SEQ ID NO: 641, amplicons in descending order are SEQ ID NOs: 642-661.
  • FIG. 16A-16E TRV delivery of ISYmul for editing in whole plants.
  • FIG. 16A and 16B Violin plots depicting the genotypes along the X-axis and the editing efficiency (percent indel reads (%)) on the Y-axis. Each dot indicates an independent plant. The standard error of the mean (SEM) was calculated for each target site as indicated by the blue dot and vertical line. WT is an abbreviation for wild type, and rdr6 is an abbreviation for mutant in RNA-DEPENDENT RNA POLYMERASE 6 gene.
  • FIG. 16C and 16D DNA repair indel profile for ISYmul g2. The indel type is listed on the left.
  • the read counts for each indel are listed on the right.
  • the TAM is identified by the red box, and the target site is located by the black box in the Reference sequence.
  • ISYmul -g2-tRNA reference sequence is SEQ ID NO: 662, amplicons in descending order are SEQ ID NOs: 663-672.
  • ISYmul-g2-HDV-tRNA reference sequence is SEQ ID NO: 673, amplicons in descending order are SEQ ID NOs: 674- 686.
  • FIGS. 17A-17E Testing activities of novel TnpBs in bacteria.
  • FIG. 17A Cassette of dual plasmids for plasmid interference assay.
  • FIG. 17B Normalized CFU (Norm. CFU) of five identified novel TnpBs. All the interference assays were performed at 26°C.
  • FIG. 17C TAM sequences of TnpB_30 from the TAM depletion assay.
  • FIG. 17D Small RNA-seq confirmed expression of the coRNA of TnpB_30.
  • FIG. 17E Secondary structures of TnpB_30 coRNA (SEQ ID NO 687).
  • FIG. 18 TnpB_30 editing efficiency m Arabidopsis protoplast cells.
  • the target sites are plotted along the X-axis and the editing efficiency (percent indel reads (%)) is plotted on the Y-axis. Each dot indicates a single transfection. The standard error of the mean (SEM) was calculated for each target site.
  • FIGS. 19A-19D Spacer length optimization for ISDra2, ISYmul. ISBagOl, and TnpB_30.
  • the spacer lengths (nucleotides) are plotted along the X-axis and the editing efficiency (percent indel reads (%)) is plotted on the Y-axis. Each dot indicates a single transfection. The standard error of the mean (SEM) was calculated for each target site.
  • FIG. 19A ISDra2.
  • FIG. 19B ISYmul.
  • FIG. 19C ISBagOl.
  • FIG. 9D TnpB_30.
  • FIG. 20A-20C Comparison of ISYmul expression using the single or split expression cassette designs.
  • FIG. 20A Plasmid design for the single and split plasmids. Green arrow boxes symbolize promoters, red boxes symbolize terminators, and the black arrows indicate the orientation of the expression cassette.
  • FIG. 20B Small RNA-seq results for ISYmul with the read count plotted along the Y-axis, and the TnpB and wRNA positions in the bactenal expression vectors plotted along the X-axis.
  • TnpB and wRNA single and split expression configurations are plotted along the X-axis and the editing efficiency (percent indel reads (%)) is plotted on the Y-axis. Each dot indicates a single transfection. The standard error of the mean (SEM) w as calculated for each target site.
  • FIG. 21 ISYmul wRNA variants for improved editing efficiency.
  • A ISYmul g2 was used in this experiment. The coRNA variants are plotted along the X-axis and the editing efficiencies (percent indel reads (%)) are plotted on the Y-axis.
  • the X-axis label WT vO.O indicates the TnpB with coRNA overlap native DNA sequence; all other variants along the X- axis are novel ISYmul wRN A variants. Each dot indicates a single transfection. The standard error of the mean (SEM) w as calculated for each coRNA design tested.
  • FIG. 22 Representative ISYmul wRNA variants at stem-loop 2.
  • the stem-loop 2 are highlighted in the secondary structures of ISYmul wRNA variants. Representative in RNA variants of stem-loop 2 are compared to the wild-type sequences. The topologies of the RNA two-way junctions are labeled within the region of interest. Stem-loop 2 (top) is SEQ ID NO: 688, WT is SEQ ID NO: 689, v3.2 is SEQ ID NO: 690, v3. 16 is SEQ ID NO: 691, v2. 1 is SEQ ID NO: 692.
  • FIG. 23A and 23B ISYmul somatic editing in Tl transgenic plants.
  • ISYmul g2 and ISYmul gl2 were used in this experiment.
  • Panels A and B display box and whisker plots. Each dot indicates a single T1 transgenic plant.
  • the room and HS treatments stand for room temperature and heat shock plant growth conditions, respectively.
  • the genotypes are plotted along the X-axis and the editing efficiencies (percent indel reads (%)) are plotted on the Y- axis.
  • WT is an abbreviation for wild type
  • rdr6 is an abbreviation for mutant in RNA- DEP ENDENT RNA POLYMERASE 6 gene.
  • FIG. 23 A ISYmul g2 Tl.
  • FIG. 23B ISYmul g!2 Tl.
  • FIG. 24A and 24B ISDra2 somatic editing in T1 transgenic plants.
  • ISDra2 g9 and ISDra2 g!2 were used in this expenment.
  • Panels A and B display box and whisker plots. Each dot indicates a single Tl transgenic plant.
  • the room and HS treatments stand for room temperature and heat shock plant growth conditions, respectively.
  • the genotypes are plotted along the X-axis and the editing efficiencies (percent indel reads (%)) are plotted on the Y- axis.
  • WT is an abbreviation for wild type
  • rdr6 is an abbreviation for mutant in RNA- DEPENDENT RNA POLYMERASE 6 gene.
  • FIG. 24A ISDra2 g9 Tl.
  • FIG. 24B ISDra2 g!2 Tl.
  • FIG. 25A-25D TRV delivery of ISYmul for editing in whole plants.
  • FIG. 25A Schematic of the TRV1 and TRV2 plasmids.
  • FIG. 25B and 25C The TRV cargo configurations are plotted along the X-axis and the editing efficiency (percent indel reads (%)) is plotted on the Y -axis. Each dot indicates a single plant. The standard error of the mean (SEM) was calculated for each target site. WT samples are black, rdr6 samples are blue, and ku70 samples are red.
  • WT is an abbreviation for wild type, and rdr6 is an abbreviation for mutant in RNA-DEPENDENT RNA POLYMERASE 6 gene.
  • FIG. 25B ISYmul g2.
  • FIG. 25C ISYmul gl2.
  • FIG. 25D DNA repair indel profile for ISYmul g!2 using TRV to deliver ISYmul -gl2-HDV-tRNA-Isoleucine. The top six most common indel types are listed on the left. The read counts for each indel are listed on the right.
  • the TAM is identified by the red box, and the target site is located by the black box in the reference sequence.
  • FIG. 26A-26C Heritability of edits generated with ISYmul g2 encoded on the TRV vector.
  • FIG. 26 A Picture of a plant that underwent TRV delivery’ using ISYmul -g2-HDV- tRNA-Isoleucine. The white sectors, indicated by the yellow arrows, indicate biallelic mutations in the PDS3 gene. The 54.54% in the upper left comer is the editing frequency determined by random tissue sampling of three leaves distal from the TRV delivery’ location.
  • FIG. 26B Image of progeny seedlings from the plant in Figure 26A, depicting two albino seedlings.
  • FIG. 26C Sanger sequencing trace file screenshot for one of the albino plants in FIG. 26B (top sequence, SEQ ID NO: 700, bottom sequence SEQ ID NO: 701).
  • FIG. 27A-27F Testing the impact of wRNA processing on the activities of TnpB in bacteria.
  • FIG. 27A Different cassettes of TnpB and wRNA expression in bacteria.
  • FIG. 27B Normalized CFU of ISDra2, ISBagOl. ISYmul and ISAaml TnpB expressed in cassette (a).
  • FIG. 27C Normalized CFU of ISDra2, ISBagOl, ISYmul and ISAaml TnpB expressed in cassette (b).
  • FIG. 27D Normalized CFU of ISBagOl expressed in cassette (c) and (d).
  • FIG. 27E Normalized CFU of ISYmul expressed in cassette (c) and (e).
  • FIG. 27F Normalized CFU of ISYmul expressed in cassette (f) and (g). All the interference assays were performed in 26°C. Error bars represent the standard error of the mean (SEM) of three replicas.
  • RNA-guided genome editors e.g. CRISPR-Cas enzymes
  • tissue culture transformation approaches Altpeter et al., 2016; Chen et al., 2019; Nash & Voytas, 2021.
  • Tissue culture methods require considerable time, resources, technical expertise, and can cause unintended changes to the genome and epigenome (Altpeter et al., 2016).
  • tissue culture is restricted in its application, as only a limited number species and genotypes are amenable to this technique (Altpeter et al., 2016).
  • tissue culture transformation such as de novo induction of meristems (Maher et al., 2020), the ‘‘graft-mobile'’ gene editing system(Y ang et al., 2023), and viral delivery’ of editing reagents (Ellison et al.. 2020; Ghoshal et al.. 2020; Liu et al., 2022; Nagalakshmi et al.. 2022).
  • These techniques hint at the potential to overcome this bottleneck; however, they are limited by the need for pre-established transgenic lines, they require highly experienced personnel, are likely difficult to scale, and appear to be species and genotype dependent.
  • a second, slightly more efficient way that works in some crops, is to introduce CRISPR protein and guide RNAs directly into plant cells during the tissue culture process, followed by plant regeneration and backcrossing to eliminate epigenetic variation induced bytissue culture.
  • this method also requires difficult tissue culture procedures that are costly, time consuming, and that only work in some crop species.
  • Plant viruses are ideal vectors for delivering editing reagents into plants given their natural ability' to amplify and spread throughout the plant.
  • many plant viruses for example the broad host range RNA virus Tobacco rattle virus (TRV). are able to infect almost all parts of the plant, and yet the virus is rarely transmitted into future sexual generations because viral genome does not integrate into the plant genome (Bradamante et al., 2021 ).
  • TRV RNA virus Tobacco rattle virus
  • sgRNA single guide RNA
  • TRV vectors can induce highly efficient biallelic gene editing or epigenome editing in plants (Ellison et al., 2020; Ghoshal et al., 2020; Nagalakshmi et al., 2022).
  • delivery of both the Cas nuclease and the sgRNA has not yet been possible with TRV or other viruses as their low cargo capacity- (-1.5-2 kb RNA) limits the size of the Cas nuclease to be fewer than 400 amino acids (aa).
  • the target DNA binding and cleavage of both systems involve the recognition of a protospacer adjacent motif (PAM) or transposon adjacent motif (TAM) and the formation of an J?-loop structure betw een the target DNA and guide RNA (FIG. 5C) (Nakagawa et al., 2023; Sasnauskas et al., 2023; Schuler et al., 2022).
  • PAM protospacer adjacent motif
  • TAM transposon adjacent motif
  • these ancestral proteins typically form a complex with a relatively large (-200 nt) guide RNA called OMEGA (for obligate mobile element-guided activity) RNA (coRNA) (Altae-Tran et al., 2021) which is encoded at either the left or right boundary- of the transposons (FIG. 5B).
  • OMEGA for obligate mobile element-guided activity
  • coRNA coRNA
  • the TAM sequences of these ancestral RNA-guided systems are directly encoded on one end of the transposon boundary, providing a catalog of ‘ TAM” sequences from nature (FIG. 5B).
  • TnpB enzymes bear significantly larger diversities (Altae-Tran et al., 2021) and higher genome editing efficiencies (Altae-Tran et al., 2021; Karvelis et al., 2021) and the more compact size of TnpB offers advantages for the viral delivery underlying our strategy (FIG. 5A).
  • the present disclosure herein relates, in part, to one or more small TnpB type bacterial enzymes that are less than about 500 amino acids in length and are capable of editing plant cells.
  • An advantage of this small enzyme is that it is small enough to be cloned in a plant viral vector, such as for example TRV. This allows for an easier method to edit plant genomes that does not require tissue culture or plant transformation.
  • the terms “include” and “including” have the same meaning as the terms “comprise” and “comprising.”
  • the terms “comprise” and “comprising” should be interpreted as being “open” transitional terms that permit the inclusion of additional components further to those components recited in the claims.
  • the terms “consist” and “consisting of’ should be interpreted as being “closed” transitional terms that do not permit the inclusion of additional components other than the components recited in the claims.
  • the term “consisting essentially of’ should be interpreted to be partially closed and allowing the inclusion only of additional components that do not fundamentally alter the nature of the claimed subj ect matter.
  • the modal verb “may” refers to the preferred use or selection of one or more options or choices among the several described embodiments or features contained within the same. Where no options or choices are disclosed regarding a particular embodiment or feature contained in the same, the modal verb “may” refers to an affirmative act regarding howto make or use an aspect of a described embodiment or feature contained in the same, or a definitive decision to use a specific skill regarding a described embodiment or feature contained in the same. In this latter context, the modal verb “may” has the same meaning and connotation as the auxiliary verb “can.”
  • nucleic acids or polypeptide sequences refer to two or more sequences or subsequences that are the same or have a specified percentage of amino acid residues or nucleotides that are the same (i.e., 50%, 55%, 60%, 65%, 70%. 75%. 80%.
  • nucleic acid or polypeptide sequence e.g., 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more identity over a specified region, e.g., of an entire nucleic acid or polypeptide sequence or individual portions or domains of a nucleic acid or polypeptide), when compared and aligned for maximum correspondence over a comparison window, or designated region as measured using one of the following sequence comparison algorithms or by manual alignment and visual inspection. Such sequences are then said to be “substantially identical.” This definition also refers to the complement of a test sequence, in the context of nucleic acids.
  • the identify exists over a region that is about or at least about 5, 10, 15, 20, 50, 100, or 1000, amino acids in length, to about, less than about, or at least about 220, 100 or 1000 amino acids or nucleotides in length.
  • the identity exists over a region that is at least about 5, 10, 15, or 16 amino acids in length to about 100, about 20 to about 75, about 30 to about 50 amino acids or nucleotides in length.
  • sequence comparison typically one sequence acts as a reference sequence, to which test sequences are compared.
  • test and reference sequences are entered into a computer, subsequence coordinates are designated, if necessary', and sequence algorithm program parameters are designated.
  • sequence algorithm program parameters Preferably, default program parameters can be used, or alternative parameters can be designated.
  • sequence comparison algorithm then calculates the percent sequence identities for the test sequences relative to the reference sequence, based on the program parameters.
  • BLAST and BLAST 2.0 are described in Altschul et al., Nuc. Acids Res. 25:3389-3402 (1977) and Altschul et al., J. Mol. Biol. 215:403-410 (1990), respectively.
  • the software for performing BLAST analyses is publicly available through the website of the National Center for Biotechnology Information (NCBI).
  • NCBI National Center for Biotechnology Information
  • BLAST and BLAST 2.0 are used, with the parameters described herein, to determine percent sequence identify for the nucleic acids and proteins.
  • a BLAST algorithm involves first identifying high scoring sequence pairs (HSPs) by identify ing short words of length W in the query sequence, which either match or satisfy' some positive-valued threshold score T when aligned wi th a word of the same length in a database sequence.
  • T is referred to as the neighborhood word score threshold (Altschul et al., supra).
  • these initial neighborhood word hits act as seeds for initiating searches to find longer HSPs containing them.
  • the word hits are extended in both directions along each sequence for as far as the cumulative alignment score can be increased.
  • cumulative scores are calculated using, for nucleotide sequences, the parameters M (reward score for a pair of matching residues; always >0) and N (penalty' score for mismatching residues; always ⁇ 0).
  • a scoring matrix is used to calculate the cumulative score.
  • extension of the word hits in each direction are halted when: the cumulative alignment score falls off by the quantify X from its maximum achieved value; the cumulative score goes to zero or below, due to the accumulation of one or more negative-scoring residue alignments; or the end of either sequence is reached.
  • the BLAST algorithm parameters W. T, and X determine the sensitivity and speed of the alignment.
  • the NCBI BLASTN or BLASTP program is used to align sequences.
  • the BLASTN or BLASTP program uses the defaults used by the NCBI.
  • the BLASTN program (for nucleotide sequences) uses as defaults: a word size (W) of 28; an expectation threshold (E) of 10; max matches in a query range set to 0; match/mismatch scores of 1. -2; linear gap costs; the filter for low complexity regions used; and mask for lookup table only used.
  • the BLASTP program (for amino acid sequences) uses as defaults: a word size (W) of 3; an expectation threshold (E) of 10; max matches in a query range set to 0; the BLOSUM62 matrix (see Henikoff & Henikoff, Proc. Natl. Acad. Sci. USA 89: 10915 (1992)); gap costs of existence: 11 and extension: 1; and conditional compositional score matrix adjustment.
  • polypeptide “peptide” and “protein” are used interchangeably herein to refer to a polymer of amino acid residues.
  • the terms apply to amino acid polymers in which one or more amino acid residue is an artificial chemical mimetic of a corresponding naturally occurring amino acid, as well as to naturally occurring amino acid polymers and non-naturally occurring amino acid polymer.
  • amino acid refers to naturally occurring and synthetic amino acids, as well as amino acid analogs and amino acid mimetics that function in a manner similar to the naturally occurring amino acids.
  • Naturally occurring amino acids are those encoded by the genetic code, as well as those amino acids that are later modified, e.g., hydroxyproline, y- carboxyglutamate, and O-phosphoserine.
  • Amino acid analogs refers to compounds that have the same basic chemical structure as a naturally occurring amino acid, i.e., an a carbon that is bound to a hydrogen, a carboxyl group, an amino group, and an R group, e.g., homoserine, norleucine, methionine sulfoxide, methionine methyl sulfonium. Such analogs have modified R groups (e g., norleucine) or modified peptide backbones, but retain the same basic chemical structure as a naturally occurring amino acid.
  • Amino acid mimetics refers to chemical compounds that have a structure that is different from the general chemical structure of an amino acid, but that functions in a manner similar to a naturally occurring amino acid.
  • Amino acids may be referred to herein by either their commonly known three letter symbols or by the one-letter symbols recommended by the IUPAC-IUB Biochemical Nomenclature Commission. Nucleotides, likewise, may be referred to by their commonly accepted single-letter codes.
  • “Conservatively modified variants” applies to both amino acid and nucleic acid sequences. With respect to particular nucleic acid sequences, conservatively modified variants refers to those nucleic acids which encode identical or essentially identical amino acid sequences, or where the nucleic acid does not encode an amino acid sequence, to essentially identical sequences. Because of the degeneracy of the genetic code, a large number of functionally identical nucleic acids encode any given protein.
  • the codons GCA, GCC, GCG and GCU all encode the amino acid alanine.
  • the codon can be altered to any of the corresponding codons described without altering the encoded polypeptide.
  • Such nucleic acid variations are “silent variations,” which are one species of conservatively modified variations. Every nucleic acid sequence herein which encodes a polypeptide also describes every possible silent variation of the nucleic acid.
  • each codon in a nucleic acid can be modified to yield a functionally identical molecule. Accordingly, each silent variation of a nucleic acid which encodes a polypeptide is implicit in each described sequence with respect to the expression product, but not with respect to actual probe sequences.
  • amino acid sequences one of skill will recognize that individual substitutions to a peptide, polypeptide, or protein sequence which alters a single amino acid is a “conservatively modified variant” where the alteration results in the substitution of an amino acid with a chemically similar amino acid. Conservative substitution tables providing functionally similar amino acids are well known in the art. Such conservatively modified variants are in addition to and do not exclude polymorphic variants, interspecies homologs, and alleles.
  • the following eight groups each contain amino acids that are conservative substitutions for one another: 1) Alanine (A), Glycine (G); 2) Aspartic acid (D). Glutamic acid (E); 3) Asparagine (N), Glutamine (Q); 4) Arginine (R), Lysine (K); 5) Isoleucine (I), Leucine (L), Methionine (M), Valine (V); 6) Phenylalanine (F), Tyrosine (Y), Try ptophan (W); 7) Serine (S), Threonine (T); and 8) Cysteine (C), Methionine (M) (see, e.g., Creighton. Proteins (1984)).
  • polynucleotide refers to a nucleotide, oligonucleotide, polynucleotide (which terms may be used interchangeably), or any fragment thereof. These phrases also refer to DNA or RNA of genomic, natural, or synthetic origin (which may be single-stranded or double-stranded and may represent the sense or the antisense strand).
  • nucleic acid and oligonucleotide as used herein, may refer to polydeoxyribonucleotides (containing 2-deoxy-D-ribose).
  • polyribonucleotides containing D- ribose
  • polynucleotide that is an N glycoside of a purine or pyrimidine base.
  • nucleic acid oligonucleotide
  • polynucleotide polynucleotide
  • an oligonucleotide also can comprise nucleotide analogs in which the base, sugar, or phosphate backbone is modified as well as non-purine or non-pyrimidine nucleotide analogs.
  • Oligonucleotides can be prepared by any suitable method, including direct chemical synthesis by a method such as the phosphotriester method of Narang et al., 1979, Meth. Enzymol. 68:90-99; the phosphodiester method of Brown etal., (919, Meth. Enzymol.
  • percent identity refers to the percentage of residue matches between at least two polynucleotide sequences aligned using a standardized algorithm. Such an algorithm may insert, in a standardized and reproducible way, gaps in the sequences being compared in order to optimize alignment between two sequences, and therefore achieve a more meaningful comparison of the two sequences. Percent identity' for a nucleic acid sequence may be determined as understood in the art. (See, e.g., U.S. Patent No. 7,396,664, which is incorporated herein by reference in its entirety).
  • NCBI National Center for Biotechnology Information
  • BLAST Basic Local Alignment Search Tool
  • NCBI National Center for Biotechnology Information
  • the BLAST software suite includes various sequence analysis programs including “blastn,” that is used to align a known polynucleotide sequence with other polynucleotide sequences from a variety of databases.
  • blastn a tool that is used to align a known polynucleotide sequence with other polynucleotide sequences from a variety of databases.
  • BLAST 2 Sequences also available is a tool called “BLAST 2 Sequences” that is used for direct pairwise comparison of two nucleotide sequences. “BLAST 2 Sequences” can be accessed and used interactively at the NCBI website.
  • percent identity may be measured over the length of an entire defined polynucleotide sequence, for example, as defined by a particular SEQ ID number, or may be measured over a shorter length, for example, over the length of a fragment taken from a larger, defined sequence, for instance, a fragment of at least 20, at least 30, at least 40, at least 50, at least 70, at least 100, or at least 200 contiguous nucleotides.
  • Such lengths are exemplary only, and it is understood that any fragment length supported by the sequences shown herein, in the tables, figures, or Sequence Listing, may be used to describe a length over which percentage identity may be measured.
  • variant may be defined as a nucleic acid sequence having at least 50% sequence identity to the particular nucleic acid sequence over a certain length of one of the nucleic acid sequences using blastn with the “BLAST 2 Sequences” tool available at the National Center for Biotechnology Information’s website.
  • BLAST 2 Sequences available at the National Center for Biotechnology Information
  • Such a pair of nucleic acids may show, for example, at least 60%, at least 70%. at least 80%. at least 85%. at least 90%. at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% or greater sequence identity over a certain defined length.
  • Nucleic acid sequences that do not show a high degree of identity' may nevertheless encode similar amino acid sequences due to the degeneracy of the genetic code where multiple codons may encode for a single amino acid. It is understood that changes in a nucleic acid sequence can be made using this degeneracy to produce multiple nucleic acid sequences that all encode substantially the same protein.
  • polynucleotide sequences as contemplated herein may encode a protein and may be codon-optimized for expression in a particular host. In the art, codon usage frequency tables have been prepared for a number of host organisms including humans, mouse, rat, pig, E. coli, plants, and other host cells.
  • a “recombinant nucleic acid” is a sequence that is not naturally occurring or has a sequence that is made by an artificial combination of two or more otherwise separated segments of sequence. This artificial combination is often accomplished by chemical synthesis or, more commonly, by the artificial manipulation of isolated segments of nucleic acids, e.g. , by genetic engineering techniques known in the art.
  • the term recombinant includes nucleic acids that have been altered solely by addition, substitution, or deletion of a portion of the nucleic acid.
  • a recombinant nucleic acid may include a nucleic acid sequence operably linked to a promoter sequence. Such a recombinant nucleic acid may be part of a vector that is used, for example, to transform a cell.
  • nucleic acids disclosed herein may be '‘substantially isolated or purified.”
  • the term “substantially isolated or purified” refers to a nucleic acid that is removed from its natural environment, and is at least 60% free, preferably at least 75% free, and more preferably at least 90% free, even more preferably at least 95% free from other components with which it is naturally associated.
  • the polynucleotide sequences contemplated herein may be present in expression vectors.
  • the vectors may comprise a polynucleotide encoding an ORF of a protein operably linked to a promoter.
  • “Operably linked” refers to the situation in w hich a first nucleic acid sequence is placed in a functional relationship with a second nucleic acid sequence.
  • a promoter is operably linked to a coding sequence if the promoter affects the transcription or expression of the coding sequence.
  • Operably linked DNA sequences may be in close proximity or contiguous and, where necessary to join two protein coding regions, in the same reading frame.
  • Vectors contemplated herein may comprise a heterologous promoter operably linked to a polynucleotide that encodes a protein.
  • a “heterologous promoter” refers to a promoter that is not the native or endogenous promoter for the protein or RNA that is being expressed.
  • expression refers to the process by which a polynucleotide is transcribed from a DNA template (such as into mRNA or another RNA transcript) and/or the process by which a transcribed mRNA is subsequently translated into peptides, polypeptides, or proteins.
  • Transcripts and encoded polypeptides may be collectively referred to as "gene product.
  • vector refers to some means by which nucleic acid (e.g, DNA) can be introduced into a host organism or host tissue.
  • nucleic acid e.g, DNA
  • vectors including plasmid vector, bacteriophage vectors, cosmid vectors, bacterial vectors, and viral vectors.
  • a “vector” may refer to a recombinant nucleic acid that has been engineered to express a heterologous polypeptide (e.g, the fusion proteins disclosed herein).
  • the recombinant nucleic acid typically includes c/.s- acting elements for expression of the heterologous polypeptide.
  • compositions Compositions, Constructs, Viral Vectors, and Plant Cells
  • compositions for plant gene editing can include a TnpB bacterial enzyme.
  • the TnpB bacterial enzyme can be from any bacterial TnpB system.
  • the TnpB bacterial enzy me can have a length of less than about 500 amino acids, less than about 450 amino acids, less than about 420 amino acids, or less than about 400 amino acids.
  • the TnpB can be from, or derived from, Brevibacillus agri.
  • a TnpB derived from Brevibacillus agri refers to a variant of the TnpB of Brevibacilhis agri.
  • the TnpB can have an amino acid sequence that is 70 % or more, 75 % or more. 80 % or more, 85 % or more, 90 % or more, 95 % or more, or 99 % or more identical to the amino acid sequence of the TnpB of Brevibacillus agri.
  • the TnpB can have an amino acid sequence that is 70 % or more, 75 % or more, 80 % or more, 85 % or more, 90 % or more, 95 % or more, or 99 % or more identical to the amino acid sequence of SEQ ID NO: 5.
  • the TnpB can have the amino acid sequence of SEQ ID NO: 5. It should be understood that a variant of the TnpB of Brevibacillus agri can be a synthetically designed variant or can be a TnpB that is of similar sequence but from another species or genus of bacteria.
  • the TnpB can be from, or derived from, Xylella fastidiosa.
  • a TnpB derived from Xylella fastidiosa refers to a variant of the TnpB of Xylella fastidiosa.
  • the TnpB can have an amino acid sequence that is 70 % or more, 75 % or more, 80 % or more, 85 % or more, 90 % or more, 95 % or more, or 99 % or more identical to the amino acid sequence of the TnpB of Xylella fastidiosa.
  • the TnpB can have an amino acid sequence that is 70 % or more, 75 % or more, 80 % or more, 85 % or more. 90 % or more, 95 % or more, or 99 % or more identical to the amino acid sequence of SEQ ID NO: 10.
  • the TnpB can have the amino acid sequence of SEQ ID NO: 10. It should be understood that a variant of the TnpB of Xylella fastidiosa can be a synthetically designed variant or can be a TnpB that is of similar sequence but from another species or genus of bacteria.
  • the TnpB can have an amino acid sequence that is 70 % or more, 75 % or more, 80 % or more, 85 % or more, 90 % or more, 95 % or more, or 99 % or more, or 100 % identical to one or more of the amino acid sequences of SEQ ID NOs: 46, 47, 48, 49, and 702.
  • the TnpB can comprise a peptide tag.
  • the peptide tag can include any suitable peptide tag known in the art.
  • the peptide tag can be a FLAG tag.
  • the peptide tag(s) is located at the N-terminus of the TnpB protein, while in others, the peptide tag(s) is located at the C-terminus of the TnpB protein, while in still further aspects, the peptide tag(s) is located within the interior portion of the TnpB, or any combination thereof.
  • the compositions disclosed herein can also include a polynucleotide.
  • the polynucleotide can include a first portion adapted to bind to at least a portion of the TnpB and a second portion for binding to a plant genomic target.
  • the polynucleotide can be RNA.
  • the polynucleotide can comprise an coRNA.
  • the coRNA is about 100-400 nucleotides in length, or about 200-300 nucleotides in length.
  • a DNA sequence that is 70 % or more. 75 % or more, 80 % or more, 85 % or more.
  • the polynucleotide can include a spacer element.
  • the space element can be any length suitable for use in plant gene editing using the systems and methods disclosed herein.
  • the first portion of the polynucleotide and the second portion of the polynucleotide can also be present as separate polynucleotides.
  • the separate polynucleotides may include sequences to allow for joining or hybridization in a plant cell.
  • constructs are disclosed herein.
  • the constructs can include a plant promoter operably connected to a polynucleotide encoding for a TnpB.
  • the TnpB encoded for in the polynucleotide can exhibit any or all of the properties discussed above with respect to TnpB.
  • the polynucleotide encoding for a TnpB can include a polynucleotide sequence that is 70 % or more, 75 % or more, 80 % or more, 85 % or more, 90 % or more, 95 % or more, or 99 % or more identical to the polynucleotide sequence of one or more of SEQ ID NOs: 1-4, 6-9, or 34-45.
  • the polynucleotide encoding for a TnpB can have the polynucleotide sequence of one or more of SEQ ID NOs: 1-4, 6-9, or 34- 45.
  • the polynucleotide encoding for a TnpB can be codon optimized for expression in plant cells. In one or more aspects, the polynucleotide encoding for a TnpB can be Zea mays codon optimized.
  • the construct can include a plant promoter operably connected to the polynucleotide encoding for a TnpB.
  • the plant promoter can be any suitable promoter for expressing the TnpB in the plant cell of interest.
  • the plant promoter can include a UBQ10 gene promoter.
  • the polynucleotide can further encode for an coRNA having a first portion adapted to bind to at least a portion of the TnpB bacterial enzyme and a second portion for binding to a plant genomic target.
  • the coRNA can include any or all of the features and parameters discussed above with respect to coRNA.
  • the polynucleotide encoding for TnpB and the coRNA can utilize the same promoter and/or be part of the same cassette.
  • a sequence of the coRNA and a sequence of the TnpB bacterial enzyme can at least partly overlap in a cassette or construct.
  • a polynucleotide that encodes for an coRNA and a TnpB that at least partly overlap can have a sequence that is 70 % or more, 75 % or more. 80 % or more, 85 % or more, 90 % or more. 95 % or more, or 99 % or more identical to the polynucleotide sequence of one or more of SEQ ID NOs: 140-147, 150-170, 431-433, or 475-481.
  • a sequence of the coRNA and a sequence of the TnpB bacterial enzyme can be present as a split cassette and under the control of separate promoters.
  • the polynucleotide encoding for an coRNA can be a separate polynucleotide and/or construct than the one encoding for the TnpB.
  • plant viral vectors are disclosed.
  • the plant viral vector can include one or more of the constructs mentioned above.
  • the plant viral vector can be any suitable vector for use in plant gene editing.
  • the plant viral vector can be, or can be derived from, the Tobacco rattle virus (TRV).
  • plant cells are disclosed.
  • the plant cells can include one or more of the constructs mentioned above and/or the plant viral vectors mentioned above.
  • the constructs and plant viral vectors can include any or all of the respective parameters and properties described above.
  • the cell can be a cell of a major agricultural plant, e.g., Barley, Beans (Dry Edible), Canola, Com, Cotton (Pima), Cotton (Upland). Flaxseed, Hay (Alfalfa). Hay (Non-Alfalfa), Oats, Peanuts, Rice, Sorghum, Soybeans, Sugarbeets. Sugarcane.
  • the cell is a cell of a vegetable crops which include but are not limited to.
  • alfalfa sprouts aloe leaves, arrow root, arrowhead, artichokes, asparagus, bamboo shoots, banana flowers, bean sprouts, beans, beet tops, beets, bittermelon, bok choy, broccoli, broccoli rabe (rappini), brussels sprouts, cabbage, cabbage sprouts, cactus leaf (nopales), calabaza, cardoon, carrots, cauliflower, celery, chayote, Chinese artichoke (crosnes), Chinese cabbage, Chinese celery, Chinese chives, choy sum, chrysanthemum leaves (tung ho), collard greens, com stalks, com-sweet, cucumbers, daikon, dandelion greens, dasheen, dau mue (pea tips), donqua (winter melon), eggplant, endive, escarole, fiddle head fems, field cress, frisee, gai choy (chinese mustard),
  • j erusalem artichokes, jicama kale greens, kohlrabi, lamb's quarters (quilete), lettuce (bibb), lettuce (boston), lettuce (boston red), lettuce (green leaf), lettuce (iceberg), lettuce (lolla rossa), lettuce (oak leaf - green), lettuce (oak leaf - red), lettuce (processed), lettuce (red leaf), lettuce (romaine), lettuce (ruby romaine), lettuce (russian red mustard), linkok, lo bok, long beans, lotus root, mache, maguey (agave) leaves, malanga. mesculin mix. mizuna, moap (smooth luffa).
  • moo moo, moqua (fuzzy squash), mushrooms, mustard, nagaimo, okra, ong choy, onions green, opo (long squash), ornamental com, ornamental gourds, parsley, parsnips, peas, peppers (bell type), peppers, pumpkins, radicchio, radish sprouts, radishes, rape greens, rape greens, rhubarb, romaine (baby red), rutabagas, salicomia (sea bean), sinqua (angled/ridged luffa), spinach, squash, straw bales, sugarcane, sweet potatoes, swiss chard, tamarindo, taro, taro leaf, taro shoots, tatsoi, tepeguaje (guaje), tindora, tomatillos, tomatoes, tomatoes (cherry), tomatoes (grape type), tomatoes (plum type), tumeric, turnip tops green
  • any plant cell and/or type of plant may be used in the present disclosure so long as it remains viable after being transformed with a sequence of nucleic acids.
  • the plant cell is not adversely affected by the transduction of the necessary nucleic acid sequences, the subsequent expression of the proteins or the resulting intermediates.
  • methods for gene editing in a plant are disclosed.
  • the methods can include introducing one or more of the constructs mentioned above and/or one or more of the plant viral vectors mentioned above into one or more plant cells.
  • the contracts and/or plant viral vectors can be introduced into the plant cells using any suitable techniques know n by those of skill in the art.
  • the methods for gene editing in a plant can be utilized on any desired plant genus or species. In certain aspects, the methods for gene editing can be performed on any agricultural crop or other plant.
  • Example 1 Detection of bona-fide RNA-guided TnpB systems in metagenomes
  • TnpB associated transposons In order to determine the coRNA architecture in functional RNA-guided TnpB systems and the TAM sequences, the boundaries of the TnpB-associated transposons need to be precisely determined. TnpB associated transposons are moving from one location to another and the genomic regions flanking the transposon ends therefore are different. The exact region of the system can be determined by common strategies such as multiple genome alignments (Altae-Tran et al., 2021).
  • the read mapping approach has several advantages over multiple genome alignments: (1) it requires only one metagenomic sample; (2) the species or strains representing the genome do not need to be dominant in the sample; and (3) the analysis can essentially be applied to any assembled sequences, including mobile genetic elements which typically require detection in multiple habitats or samples.
  • the "DOTS ” pipeline we applied the "DOTS ” pipeline to identify a total of 96 putative RNA-guided IS605 TnpB systems from only 500 ggKbase metagenomes (FIG. IB).
  • RNA-guided IS605 TnpB systems bear diverse TAM sequences, including CTAT, TCAC and ACAG, which are not present in public IS databases such as ISfinder, have a range of molecular sizes (336-478 aa) (FIG. IB), and have not been reported previously.
  • TnpB ⁇ 420 aa
  • NBI National Center for Biotechnology Information
  • TnpB-associated transposons are all from organisms in ambient environments, such as soil and plant-associated microbes, making them ideal candidates for plant editing.
  • Preliminary . results show that two TnpB systems have robust RNA-guided activities in E.coli'. one is a hypercompact IS200/IS605 TnpB (359 aa) from Brevibacillus agri (ISBagOl).
  • IS607 TnpB (398 aa) from Xylella fastidiosa (ISXfaOl) which is a plant pathogen (FIGS. 2A-2C).
  • ISXfaOl Xylella fastidiosa
  • FIGS. 2A-2C RNA-guided DNA cleavage activities of TnpB from IS607 transposons have yet to be reported, and our preliminary results suggests that beyond IS200/IS605, it should be fruitful to mine TnpBs from additional IS607s.
  • Table 1 includes the construct sequences for the constructs used in Example 3.
  • Table 2 lists the spacer sequences of ISBagOl that were tested.
  • Table 1 Construct Sequences for Example 3
  • Table 2 spacer sequences of ISBagOl that were tested in Example 3
  • Example 4 Identification and testing of tiny TnpB systems in plant cells
  • new RNA-guided TnpB systems will be identified from IS605 (with IS605 TnpA), IS607 (with IS607 TnpA), and IS 1341 (without TnpA) transposons using the “DOTS’” approach described above.
  • Thousands of new RNA-guided TnpB systems in the range of 300-400 amino acids in size will be selected, manually curated and subjected to subsequent sanity checks: (1) the catalytic domain is not broken (2) structure prediction algorithms (e.g., Alphafold2) can produce a reasonable 3D folding of the TnpB protein and (3) the RE of the transposons comprise putative coRNA-like secondary structures by RNA folding programs (e.g.. RNAfold).
  • RNA folding programs e.g.. RNAfold
  • TnpB TnpB
  • Tests in bacteria will be run at 26 °C to screen for nucleases that are active at temperatures that are optimal for most crop plants. All active nucleases will then be tested in Arabidopsis and maize protoplasts to quickly test if these enzymes work in plant cells, and to screen for ones that work in both monocots and dicots. Enzymes whose catalytic activity is optimal around 23-28 °C will be prioritized as these are the temperatures that most crop plants are grown. This will be tested by running parallel protoplast editing assays at different temperatures.
  • Protoplasts will be transfected with the small Cas systems encoded on plasmids, and editing outcomes will be assayed by deep amplicon sequencing to determine cutting efficiency, similar to the methods used above to test ISBagOl.
  • Prioritized Cas systems will initially be tested to determine if (1) they are compatible with a single expression cassette format containing both the TnpB protein and oiRNA (FIG. 3A). (2) an N terminal or C terminal nuclear localization sequence shows activity, (3) if codon optimization is necessary.
  • RNA-guided TnpB systems that show good activity in protoplasts will be tested in whole plant settings.
  • the constructs which are used to test TnpB activities in protoplasts are binary vectors, thus, they can be used directly for Agrobacterium-mediated transformation of the Arabidopsis plants.
  • the Arabidopsis protoplast assays will target the AtPDS3 gene, which provides a convenient readout for editing in whole plants.
  • TnpB systems and the best performing guide RNA sequences will be transformed into Arabidopsis plants and the TO plants used for transformation, as well as the T1 transgenic plants will be cultured at the optimum temperature determined for the specific TnpB system being tested.
  • the somatic editing efficiency of target regions in transgenic T1 plants will be evaluated by amplicon sequencing.
  • sy stems with high efficiency will be screened for white sectors caused by the loss of function of the AtPDSS gene, as seen previously for CasO (Li et al., 2023).
  • T2 populations of T1 plants with relatively high somatic editing efficiency will be screened for albino seedlings and tested for transgene segregation. Fully albino seedlings in the T2 population in which the transgene has been segregated away (null segregants) indicate that the mutation in the AtPDS3 gene has been meiotically inherited.
  • Example 6 Testing tiny TnpB systems in plant viral systems
  • RNA-guided TnpB systems will be initially encoded into tobacco rattle virus (TRV) derived viral vectors together with the isoleucine tRNA sequence which was shown to be the most effective RNA mobility sequence in our previous study (Nagalakshmi et al., 2022).
  • Arabidopsis and Nicotiana benthamiana plants will be infected in order to test for both somatic and germline editing, using protocols established in our previous work (Ghoshal et al., 2020; Nagalakshmi et al.. 2022).
  • AtPDS3 guide RNA sequences For Arabidopsis, the best performing AtPDS3 guide RNA sequences will be used so that we can observe white sectors on the initially infected plants, and also screen for albino seedling in the next generation which will be indicative of germline transmission of edits (Li et al., 2023).
  • Example 7 Optimizing the efficiency of tiny TnpB systems
  • Example 8 Testing vecotrs for TnpB-mediated editing in plants [00109] This Example describes exemplary' experimental guidelines for constructing and testing vectors expressing the TnpB protein and the associated coRNA. These vectors will be used to modify plant genomic DNA in protoplast cells, somatic whole plant cells, and germ cells.
  • TnpBs capable of targeted mutagenesis in bacteria and human cells were selected for testing in plants: ISDra2, ISYmul, ISAaml, and ISDgelO (Xiang et al. 2023; Karvelis et al. 2021).
  • a library' of spacer sequences targeting the Arabidopsis PDS3 gene (AT4G14210), will be designed according to their Transposon associated motif (TAM) sequence (Table 3).
  • TAM Transposon associated motif
  • Table 3 the gRNA sequences are listed as DNA sequences that are the equivalent RNA sequences (except the RNA will have the T replaced with a U).
  • a similar approach will be used to test for editing capabilities in other species, such as maize.
  • TnpBs as described herein may be used to target DNA modifications in the plant genome
  • a variety of constructs will be created. Based on reports published in human cells, we will test a split system for expression of the TnpB and coRNA sequences (Xiang et al. 2023; Karvelis et al. 2021). Here the TnpB and coRNA will be expressed by a separate promoter and terminator. The TnpB sequence will be cloned into a binary vector with NLS sequence(s) at the 5 ? and/or 3’ end of the TnpB sequence.
  • the TnpB will be expressed using a Pol II promoter and terminator such as UBQ10 and rbcS E9, respectively (Table 20 below).
  • a Pol II promoter and terminator such as UBQ10 and rbcS E9, respectively (Table 20 below).
  • the IV2 intron (Table 20), of various copy numbers, will be cloned into the TnpB coding sequence to test for enhanced editing.
  • the coRNA sequence will be cloned into a separate intermediate vector driven by either a Pol III promoter (such as AtU6) or a Pol II promoter (such as UBQ10) (Table 20).
  • the spacer sequences will then be cloned into the intermediate coRNA vector using a ty pel I restriction enzymes such as PaqCI.
  • RNA processing sequences such as Hammerhead ribozyme. HDV ribozyme, or tRNA-Gly will be tested for improved editing (Table 20).
  • Table 20 Once the binary vector containing the TnpB sequence and intermediate vector for each of the coRNA-spacer is created, the coRNA expression cassette from the intermediate vector will get assembled into the TnpB-containing binary' vector (FIG. 6).
  • TnpBs capable of editing further optimization will be done, such as testing various coRNA sequence lengths and spacer sequence lengths.
  • TnpBs The ability of TnpBs to edit whole plants and faithfully transmit those edits to the next generation is crucial for the creation of novel stable genoty pes.
  • TnpBs To evaluate the TnpBs for somatic editing in stable transgenic plants, and transmission of germline edits to transgene-free progenies, the TnpBs along with promoter, terminator, and RNA processing combinations capable of generating targeted DNA modifications in protoplasts will be transformed into Arabidopsis using the floral dip technique (Zhang et al. 2006).
  • T1 seeds will be harvested, subjected to selection to identify transgenic plants, and quantified for editing with somatic tissue such as leaves of T1 plants using deep amplicon next generation sequencing (NGS).
  • NGS deep amplicon next generation sequencing
  • we suspect highly active germ-cell promoters may improve germline editing and heritability of mutations.
  • additional pol II promoters such as RPS5A, YAO1 and EC1.2 promoters.
  • Heritability of edits which disrupts the PDS3 gene function can be easily identified due to the bleached phenotype of biallelic mutations in the PDS3 (AT4G14210) gene being targeted for editing in the T2 generation.
  • Albino plants in T2 generation in which the TnpB transgene has already segregated away indicates the heritability of the edits in the PDS3 gene. These plants will also be subjected to Sanger sequencing to confirm the edits in the PDS3 gene.
  • Example 9 Single expression cassettes for improved editing efficiency
  • This Example describes exemplary' experimental guidelines for constructing and testing vectors expressing the TnpB protein and the associated coRNA in a single expression cassette. These vectors may be used to modify plant genomic DNA in protoplasts, whole plants and to generate heritable edits in genomic sequences using similar methods as described in Example 8. We expect these results to translate to eukaryotic species other than those being tested here.
  • the single expression cassette will work as efficiently or better than the split systems previously described.
  • we will clone the TnpB and overlapping cuRN A into a single expression cassette (see informal sequence listing below - SEQ ID NOs: 140-147) on a binary vector.
  • the single expression cassette will contain a promoter (such as UBQ10), TnpB coding sequence, coRNA sequence, spacer sequence, RNA processing sequence, and terminator, as depicted in (FIG. 7B). Further optimization of TnpB and coRNA expression will be performed using various promoters, terminators, and RNA processing sequences such as tRNA-Gly, Hammerhead ribozyme and HDV ribozyme (Table 20).
  • This Example describes exemplary experimental guidelines for constructing and testing vectors expressing the TnpB protein and the associated coRNA with a tandem repeat of the coRNA in a single expression cassette. These vectors may be used to modify genomic DNA in protoplasts, whole plants and to generate heritable edits in genomic sequences using similar methods as described in Example 8.
  • TnpBs have been shown to self-process their own coRNA (Nety et al. 2023), we will test tandem repeats of the coRNA for improved coRNA processing, editing, and multiplexing capabilities; an example with one tandem coRNA repeat is depicted in FIG. 8, however the copy number may vary.
  • Example 11 Constructing and testing viral vectors for TnpB-mediated somatic and germ cell editing
  • This Example describes exemplary' experimental guidelines for constructing and testing vectors expressing the TnpB protein and the associated coRNA in viral vectors such as Tobacco Rattle Virus (TRY).
  • TRY Tobacco Rattle Virus
  • Delivery' of viral vectors for genome editing has the potential to expedite the creation of novel genetic diversity, and at the same time avoid the process of generating transgenic plants.
  • viruses harboring the RNA- guided endonuclease and the gRNA (such as Cas9 and Casl2a) cannot fit in single viral vectors such as TRV.
  • TRV is a bipartite positive-sense single stranded RNA virus consisting of two components: TRV1 and TRV2 (E. E. Ellison, Chamness, and Voytas 2021).
  • TRV1 contains the RNA-dependent RNA polymerase (RdRP) which replicates both TRV1 and TRV2 and expresses sub genomic RNAs of other coding regions.
  • TRV1 also encodes the movement protein (MP), responsible for enabling cell-to-cell movement, and the viral suppressor of RNA silencing (VSR), responsible for suppression of the host immune response.
  • TRV2 contains the coat protein (CP) and can be modified for expression of heterologous sequences.
  • the native TRV2 contains coding sequences to enable transmission between plants, which are removed in the vector. Heterologous sequences are expressed from a sub genomic promoter.
  • TRV1 and TRV2 vectors will be delivered to Arabi dopsis plants using previously described approaches such as syringe infiltration of leaves, agro-pri eking, and agro-flooding methods (Nagalakshmi et al. 2022).
  • Somatic editing will be quantified using methods similar to those described in Example 8. Briefly, two approaches will be used: deep amplicon NGS, and visually via the photobleaching phenotype caused by a bi-allelic frameshift mutation in the AtPDS3 gene (AT4G14210).
  • Heritable germline editing will also be assayed in the offspring generations using methods similar to those described in Example 8.
  • Example 12 Utilizing TnpBs for genome editing in plant cells
  • This Example describes experimental guidelines for constructing and testing vectors expressing the TnpB protein and the associated coRNA. These vectors were used to modify plant genomic DNA m ' Arabidopsis mesophyll protoplast cells. This example is related to the prophetic Example 8, and the experiments were performed as described in that example. [00142] Creating vectors for TnpB-mediated genome editing
  • TnpBs capable of targeted mutagenesis in bacteria and human cells were selected for testing in plants: ISDra2, ISYmul, and ISAaml (Xiang et al. 2023; Karvelis et al. 2021).
  • the TnpB protein coding sequences have been separated from the DNA sequence encoding the coRNA, resulting in separate expression cassettes.
  • the coRNA overlaps with or is directly downstream of the TnpB sequence, as revealed by metagenomic data in Example 1 and publicly available genomic data (FIG. 7A).
  • TnpB systems encoded in the single expression cassette format.
  • the native DNA sequence encoding the TnpB and the coRNA was cloned into a single expression cassette, driven by the UBQ10 promoter, followed by the desired spacer sequence, HDV ribozyme RNA processing sequence, and rbcS-E9 terminator on a binary' vector, as depicted in FIG. 7B (Tables 4 and 5 below).
  • gRNA sequences targeting the Arabidopsis PDS3 gene (AT4G14210) were designed according to their Transposon associated motif (TAM) sequence (Table 6).
  • This Example describes experimental guidelines for testing vectors expressing the
  • TnpB protein and the associated oRNA were used to modify plant genomic DNA in protoplast cells.
  • 21 novel TnpBs were identified using the pipeline described in Example 1. Using the same approach described above, the 21 TnpBs targeting the PDS3 gene (AT4G14210) were cloned into individual binary' vectors (Table 8 (SEQ ID NOs: 150-170) and Table 9 (SEQ ID NOs: 171-358). Each vector was transfected into Arabidopsis mesophyll protoplast cells. Protoplast cells were subjected to a 37°C heat shock treatment for 2 hours at 16 hours post transfection. At 48 hours post transfection, protoplasts were harvested for genomic DNA extraction and analysis for targeted mutagenesis using NGS.
  • TnpB_20, TnpB_25, TnpB_15 and TnpB_21 reached up to an editing frequency of 0.53%, 0.57%, 0.83, and 1.36% at an individual target site, respectively. (FIG. 11, Table 7).
  • Example 14 - mRNA engineering for improved TnpB editing efficiency
  • This Example describes exemplary experimental guidelines for testing vectors expressing the TnpB protein and the associated CDRNA.
  • RNA engineering of the ISDra2 OJRNA is a strategy' for improving the TnpBs ability to edit genomic DNA in mammalian cells (Li et al. 2024).
  • TnpB and ⁇ RNA are co-expressed as a single transcript where sometimes there is an overlapping region between the C-terminus of TnpB coding sequences (CDS) and the 5 ’-end of the ⁇ RNA.
  • CDS TnpB coding sequences
  • the ISDra2 ⁇ RNA scaffold sequence overlaps with the TnpB CDS except for 7 nucleotides.
  • Such overlap between the CDS and ⁇ RNA introduces additional complexity and difficulty 7 for RNA engineering, i.e.
  • RNA-seq and RNA folding analysis on ISYmul and ISBagOl ⁇ RNA suggests a three stem-loop (127-nt) and a five stem-loop (186-nt) ⁇ RNA architecture, respectively (FIG. 12).
  • the three stem-loop architectures of ISYmul ⁇ RNA resembles the ISDra2 ⁇ RNA (Sasnauskas et al. 2023; Nakagawa et al. 2023), in which the apical loop of stem-loop 1 forms pseudoknot interactions w ith the 3’-end of the ⁇ RNA scaffold.
  • the ⁇ RNA of ISBagOl bears two additional stem-loops at the 5’-end forming an RNA structure comprised of five stem-loops, where the pseudoknot interactions are still preserved at similar positions.
  • CRISPR-Casl2f which has been structurally characterized and engineered
  • RNA engineering for TnpB and CRISPR- Casl2f
  • stem truncation it has been shown that trimming or deleting RNA stems that are in structurally disordered regions and/or not interacting with proteins can improve the editing activities for ISDra2 TnpB and CRISPR- CasI2f (Nakagawa et al. 2023; Li et al. 2024; Kim et al. 2021).
  • Each TRV2 plasmid was co-transfected into Arabidopsis mesophyll protoplast cells together with the TRV 1 plasmid.
  • Protoplast cells were subjected to a 37°C heat shock treatment for 2 hours at 16 hours post transfection.
  • protoplasts were harvested for genomic DNA extraction and analysis for targeted mutagenesis using NGS. Detectable editing was observed in all three configurations tested, with a similar trend in both experiments one and two.
  • An average 0.15% and 1.09% editing using the ISYmul-g2- tRNA-Isoleucine configuration, and 0.84% and 1.61% editing using the ISYmul -g2-tandem- 200bp-wRNA-tRNA-Isoleucine configuration FIG.
  • a screen will be performed using the same ISYmul - g2 -tandem ®RN A-tRNA-Isoleucine design, but with varying lengths for the second ⁇ RNA, ranging from 100-250 nucleotides in increments of ten nucleotides (Table 12). Based on ISYmul small RNA-seq data in E.coli, it is suspected that 127 nucleotides may be an optimal length for the second coRN A. which will also be tested (FIG. 12A). Once the range with highest editing efficiencies is found, all ⁇ RNA sequence lengths in that range to identity the optimal length of the second ⁇ RNA sequence will be tested.
  • This Example describes exemplary experimental guidelines for constructing vectors to express the TnpB protein and the associated ⁇ RNA in transgenic plants by Agrobacterium transformation, as well as testing the corresponding editing outcomes in these transgenic plants. This example is related to Examples 5, 8, and 9.
  • the native DNA sequence encoding the TnpB and the ®RNA were cloned into a single expression cassette (SEQ ID NOs: 140 and 143), driven by the UBQ10 promoter (SEQ ID NO: 132), followed by the desired spacer sequence (SEQ ID NOs: 55, 65. 121, 123), HDV ribozyme RNA processing sequence (SEQ ID NO: 137), and rbcS-E9 terminator (SEQ ID NO: 135) on a binary vector, as depicted in Figure 7B (Tables 4, 5, and 6) and used floral dipped according to Zhang et al. (Zhang et al. 2006).
  • T1 seeds were harvested and subjected to hygromycin selection on MS plates to select for transgenic plants. Seedlings that passed selection were then transplanted to soil and grown at room temperature for three weeks.
  • leaf tissue samples were randomly collected from each transgenic plant and editing efficiencies were quantified by amplicon sequencing using Illumina next generation sequencing (NGS). Editing in leaf tissue for ISDRa2 (g9 and gl 2) and ISYmul (g2 and gl 2) was observed.
  • ISDra2 g9 demonstrated an average editing frequency in wild type (WT) of 0.32% and in rdr6 of 0.04% (FIG. 15A).
  • ISDra2 g!2 demonstrated an average editing frequency of 0.28% and 0.37% in wild type and rdr6. respectively (FIG. 15B).
  • ISYmul g2 we observed an editing frequency in wild type of 0.14%, and in rdr6 of 4.59% (FIG. 15C).
  • ISYmul g!2 demonstrated a very high editing frequency 7 in wildtype (48.96%) and in rdr6 (79.73%) (FIG. 15D).
  • analysis of the repair profiles for ISDra2 g9 and ISYmul g!2 revealed deletion-dominant repair outcomes (FIGS. 15E-15G).
  • Example 17 TnpB-mediated editing via delivery of a TRY vector to whole plants
  • This Example describes exemplary experimental guidelines for constructing vectors to express the TnpB protein and the associated coRNA via the Tobacco Rattle Virus (TRV) delivery 7 . These vectors were used to edit plant DNA in whole plants.
  • TRV Tobacco Rattle Virus
  • TRV vectors for TnpB-mediated genome editing of Arctbidopsis were designed. Two different viral vector configurations were cloned into the TRV “cargo” site for testing: 1. ISYmul-g2-tRNA-Isoleucine (SEQ ID NO: 456) and 2. ISYmul-g2- HDV-tRNA-Isoleucine (SEQ ID NO: 457), as shown in FIGS. 9A and 9C, respectively (Table 12).
  • Each TRV2 plasmid was co-delivered to wild type or rdr6 Arctbidopsis seedlings together with the TRV1 plasmid at 9 days post germination using the agroflood method (Nagalakshmi et al. 2022). After four days of agroflood co-culture, seedlings were transplanted to soil and grown for 18 days. Next, leaf tissue samples were randomly sampled from each plant for genomic DNA extraction and analysis for targeted mutagenesis using Illumina sequencing. Editing was detected in both vector configurations tested.
  • FIG. 16A An average 0.03% and 0.12% editing efficiency using the ISYmul -g2-tRNA ⁇ Isoleucine configuration in the wild type and rdr6 samples, respectively (FIG. 16A) was observed. Notably, editing efficiency up to 1.26% in the rdr6 genotype (FIG. 16A) was observed. Additionally, an average of 0.29% and 0.28% using the ISYmul-g2-HDV-tRNA-Isoleucine configuration in wild type and rdr6, respectively (FIG. 16B) was observed. Notably, moderately high editing efficiency up to 5.54% and 8.93% in the wild type and rdr6 samples, respectively (FIG. 16B) was observed.
  • This Example describes exemplary experimental guidelines for constructing vectors to express the TnpB protein and the associated coRNA in bacteria. These vectors were used to test the activities of novel TnpB proteins in bacteria.
  • TnpBs were identified using either the DART approach outlined in Examples 1 and 13 or a protein homology -based search, including seven new examples of TnpBs (Table 13 below). A subset of these TnpBs was confirmed for their plasmid interference activities in bacteria.
  • the native DNA sequence encoding the TnpB and the wRNA was cloned into a single expression cassette (bearing a CmR gene) driven by the TetR promoter (FIG. 17A), followed by the desired spacer sequence targeting another vector bearing an AmpR gene (FIG. 17A).
  • Both vectors were co-transformed into E.coli, which were then subjected to serial dilution and grown on plates with either the single antibiotic (Cam) or double antibiotic (Cam and Amp). The experiments were performed at 26°C.
  • the activities of TnpB were calculated by the ratio of colony -forming units (normalized CFU) betw een the single antibiotic plate and the double antibiotic plate.
  • An active TnpB will target the AmpR vector resulting in plasmid interference and reduced normalized CFU.
  • the normalized CFU for each TnpB were also compared under conditions with and without the predicted TAM sequences on the AmpR vector.
  • TnpB_5 TnpB_15, TnpB_19, TnpB_21 and TnpB_30
  • TnpB_30 TnpB_30
  • the TAM of TnpB_30 was further confirmed by a TAM depletion assay as the “AGGAG” motif (FIG. 17C).
  • Small RNA-seq of TnpB_30 revealed a usRNA scaffold size of 138 nucleotides, consisting of a three stem-loop architecture (Fig. 17E).
  • Example 19 Identifying novel TnpBs for genome editing in plant cells
  • This Example describes exemplary' experimental guidelines for testing vectors expressing the TnpB protein and the associated coRNA. These vectors were used to modify plant genomic DNA in protoplast cells. It is expected that these results translate to eukaryotic species other than those being tested in this Example.
  • TnpB 30 the most active novel TnpB in bacteria in Example 18, generated edits at an average 0.45% across nine target sites, reaching up to and average of 1.2% at target site g2 (FIG. 18, Table 15).
  • Example 20 The spacer length for ISYmul, ISDra2, ISBagOl. and TnpB 30 can impact editing efficacy
  • This Example describes exemplary experimental guidelines for testing vectors expressing the TnpB protein and the associated wRNA. These vectors were used to modify plant genomic DNA in protoplast cells. It is expected that these results translate to eukaryotic species other than those being tested in this Example.
  • Each vector was transfected into Arabidopsis mesophyll protoplast cells.
  • Protoplast cells were subjected to a 37°C heat shock treatment for 2 hours at 16 hours post transfection.
  • protoplasts were harvested for genomic DNA extraction and targeted mutagenesis analysis using amplicon sequencing. Analysis of the editing frequencies revealed very little difference between the spacer lengths for ISDra2 (FIG. 19A).
  • ISYmul displayed a preference for spacer lengths of 16 and 19 nucleotides, with 16 being the optimal length (FIG. 19B).
  • ISBagOl appeared to have a strong preference for a 16 nucleotide spacer length (FIG. 19C).
  • TnpB_30 showed increased editing for spacer lengths greater than 15 nucleotides, with 19 nucleotides demonstrating the highest editing efficiency (FIG. 19D). In summary, these data indicate the spacer length can influence TnpB-mediated genome editing efficacy.
  • This Example describes exemplary experimental guidelines for constructing and testing vectors expressing the TnpB protein and the associated wRNA in a single expression cassette, or a split expression cassette design. These vectors may be used to modify plant genomic DNA in protoplasts, whole plants and to generate heritable edits in genomic sequences using similar methods as described in other Examples in this patent. It is expected that these results translate to eukaryotic species, and TnpBs, other than those being tested in this Example. This Example is related to Example 9.
  • the single expression cassette contained a pol II promoter (such as UBQ 10 (SEQ ID NO: 132)), TnpB coding sequence, wRNA sequence, spacer sequence, HDV ribozyme RNA processing sequence (SEQ ID NO: 137). and terminator (such as rbcS-E9 (SEQ ID NO: 135)) (FIG. 20A).
  • a pol II promoter such as UBQ 10 (SEQ ID NO: 132)
  • TnpB coding sequence such as TnpB coding sequence
  • wRNA sequence such as wRNA sequence
  • spacer sequence such as HDV ribozyme RNA processing sequence
  • HDV ribozyme RNA processing sequence SEQ ID NO: 137
  • terminator such as rbcS-E9 (SEQ ID NO: 135)
  • Example 22 ISYmul wRNA engineering for improved plant genome editing
  • This Example describes exemplary experimental guidelines for testing vectors expressing the TnpB protein and the associated wRNA. This Example is related to Example 14. It is expected that these results translate to TnpBs and eukaryotic species other than those being tested in this Example.
  • RNA engineering of the ISDra2 wRNA is a strategy for improving the TnpB’s ability to edit genomic DNA in mammalian cells (Li et al. 2024).
  • the TnpB and wRNA are co-expressed as a single transcript where sometimes there is an overlapping region between the C-terminus of TnpB coding sequences (CDS) and the 5’-end of the wRNA.
  • CDS TnpB coding sequences
  • the ISDra2 wRNA scaffold sequence overlaps with the TnpB CDS except for 7 nucleotides.
  • Such overlap between the CDS and UJRNA introduces additional complexity and difficulty for RNA engineering, i.e. changing usRNA sequences can also interfere with the CDS.
  • RNA secondary structure analysis revealed that v3.2 features an SISI rather than an S2S1 two-way junction topology at lower stem-loop 2; v3.16 features an S2S0 rather than an SI SO two-way junction topology at the middle of stem-loop 2; and v2. 1 features stem truncation at upper stem-loop 2 (FIG. 22). Since modulation of RNA two-way junctions in stem-loop 2, as seen with v3.2, enhances genome editing activities the most, we hypothesize that maintaining similar topologies in stem-loop 2 are critical for ISYmul TnpB activity. Consequently, 9 additional variants based on the ISYmul oiRNA variant v3.2 with a new mutation (Table 18) were designed.
  • Example 14 we plan to combine coRNA variant mutations to test for improved editing. Stacked coRNA variants using combinations of v3.2, v3.16, v3.4, v5.2, and/or v2.1 to evaluate enhanced ISYmul editing capabilities (Table 18) have been designed. It is anticipated that further exploration of coRNA engineering variants and stacking will significantly improve ISYmul genome editing efficacy. [00191] Example 23 - Highly efficient ISYmul-mediated somatic editing in transgenic Arabidopsis
  • This Example describes exemplary experimental guidelines for constructing vectors to express the TnpB protein and the associated oiRNA in transgenic plants by Agrobacterium transformation, as well as testing the corresponding editing outcomes in these transgenic plants. This Example is related to Example 16.
  • the native DNA sequence encoding the TnpB and the rnRNA were cloned into a single expression cassette (SEQ ID NO: 140), driven by the UBQ 10 promoter (SEQ ID NO: 132), followed by the desired spacer sequence (SEQ ID NOs: 55 and 65), HDV ribozy me RNA processing sequence (SEQ ID NO: 137), and rbcS-E9 terminator (SEQ ID NO: 135) on a binary vector, as depicted in FIG. 7B, and used for floral dipped according to Zhang et al. 2006. Floral dip was performed using wild type (WT) and RNA-DEPENDENT RNA POLYMERASE 6 (rdr6) genotypes, as rdr6 has been shown to have reduced transgene silencing.
  • WT wild type
  • rdr6 RNA-DEPENDENT RNA POLYMERASE 6
  • ISYmul g2 we observed an average editing frequency in WT of 1.56%. and 2.49% in rdr6 for the room temperature samples (FIG. 23 A).
  • ISYmul gl2 demonstrated a high editing frequency in WT (44.86%) and in rdr6 (75.55%) for the room temperature grown plants (FIG. 23B).
  • Analysis of editing efficiency in the heat shock treatment samples revealed increased editing at ISYmul g2 target site in both WT and rdr6, averaging 9.81% and 32.47%. respectively (FIG. 23 A).
  • the ISYmul g!2 plants that underwent the heat shock treatment showed increased editing in WT, averaging 63.73%; whereas the rdr6 samples saw no difference in editing efficiency, averaging 75.43% (FIG. 23B).
  • Example 24 - Efficient ISDra2-mediated somatic editing in transgenic Arabidopsis requires heat shock treatment.
  • This Example describes exemplary experimental guidelines for constructing vectors to express the TnpB protein and the associated wRNA in transgenic plants by Agrobacterium transformation, as well as testing the corresponding editing outcomes in these transgenic plants. This Example is related to Examples 16.
  • the native DNA sequence encoding the TnpB and the wRNA were cloned into a single expression cassette (SEQ ID NO: 143), driven by the UBQ IO promoter (SEQ ID NO: 132), followed by the desired spacer sequence (SEQ ID NOs: 121 and 123), EIDV ribozy me RNA processing sequence (SEQ ID NO: 137), and rbcS-E9 terminator (SEQ ID NO: 135) on a binary vector, as depicted in FIG. 7B 7B, and used for floral dipped according to Zhang et al. 2006. Floral dip was performed using wild type (WT) and RNA-DEPENDENT RNA POLYMERASE 6 (rdrti) genotypes, as rdr6 has been shown to have reduced transgene silencing.
  • T1 seeds were harvested and subjected to hygromycin selection on Vz MS plates to select for transgenic plants. Seedlings that passed selection were then transplanted to soil and grown at room temperature for one week. After one week, room temperature plants continued to grow at room temperature for two more weeks; however, plants that underwent heat shock treatment were exposed to 8 hours of heat exposure at 37°C every day for 5 days, followed by 2 days of recovery at room temperature. This heat shock regime lasted for two weeks.
  • leaf tissue samples were randomly collected from each transgenic plant and editing efficiencies were quantified by amplicon sequencing using Illumina next generation sequencing (NGS). Editing in leaf tissue for ISDra2 g9 and ISDra2 gl2 was detected.
  • NGS next generation sequencing
  • Example 25 Somatic editing of whole plants using TRV vectors expressing the
  • This Example describes exemplary experimental guidelines for constructing vectors to express the TnpB protein and the associated wRNA via Tobacco Rattle Virus (TRV) delivery'. These vectors were used to edit plant DNA in whole plants.
  • TRV Tobacco Rattle Virus
  • TRV vectors for TnpB-mediated genome editing in Arabidopsis were designed targeting ISYmul g2 target site.
  • Three different viral vector configurations were cloned into the TRV "cargo" site for testing: (1) ISYmul -g2-tRNA-Isoleucine (SEQ ID NO: 437), (2) ISYmul-g2-HDV-tRNA-Isoleucine (SEQ ID NO: 438), and (3) ISYmul -g2-tandem- 200bp-wRNA-tRNA-Isoleucine (SEQ ID NO: 439) (FIG. 25A).
  • Each TRV2 plasmid was co-delivered with the TRV1 plasmid to wild type (WT), ku70, or RNA-DEPENDENT RNA POLYMERASE 6 (rdr 6) Arabidopsis seedlings at 9 days post germination using the agroflood method (Nagalakshmi et al. 2022).
  • Ku70 is a nonhomologous end joining DNA repair mutant. We hypothesized the ku 70 genetic background may show increased editing efficiency.
  • TRV vectors to target the ISYmul gl2 target site.
  • the ISYmul-gl2-HDV -tRNA-Isoleucine TRV vector was used (Table 19).
  • the TRV2 plasmid was co-delivered with TRV1 to WT or rdr6 Arabidopsis seedlings using the same agroflood methods described above (Nagalakshmi et al. 2022). Additionally, the plant growth and heat shock treatment was the same as described above.
  • three leaf tissue samples distal from the TRV delivery location were randomly collected and pooled from each individual plant for genomic DNA extraction and analysis for targeted mutagenesis using Illumina amplicon sequencing. Editing was detected in both genotypes tested.
  • Example 26 Heritable editing using TRV vectors expressing the ISYmul TnpB and uiRNA in Arabidopsis
  • This Example describes exemplary experimental guidelines for constructing vectors to express the TnpB protein and the associated wRNA via Tobacco Rattle Virus (TRV) delivery. These vectors were used to edit plant DNA in whole plants. This Example is related to the previously filed Example 17.
  • TRV Tobacco Rattle Virus
  • Example 27 Efficient plasmid interference activities of TnpBs with matured O)RNA
  • This Example describes exemplary experimental guidelines for constructing vectors to express the TnpB protein and the associated wRNA in bacteria. These vectors were used to test the activities of TnpB proteins with various mechanisms of wRNA expression and processing.
  • TnpB and wRNA in bacteria were created as depicted in (FIG. 27 A): (a) A single expression cassette for TnpB and wRN A with the 3’-end of the guide uncapped; (b) A single expression cassette for TnpB and wRNA with a 16-nt guide region with the 3’-end capped by HDV ribozyme; (c) A split expression cassette for TnpB and UJRNA, with the wRNA scaffold length inferred from small RNA-seq experiments, a 16- nucleotide guide region with the 3’ -end capped by HDV ribozyme; (d) A split expression cassette for TnpB and uiRNA, with the wRNA scaffold truncated by 52 nucleotides at the 5’- end, a 16-nucleotide guide region with the 3'-end capped by HDV ribozyme; (e) A split expression cassette for TnpB
  • TnpB was expressed from ISDra2, ISBagOl, ISYmul, and ISAaml using cassette (a). Strong plasmid interference activities were observed for ISDra2 and ISBagOl but not for ISYmul and ISAaml (FIG. 27B). When the HDV ribozyme was introduced at the 3’-end of the guide in cassette (b), the activities of ISYmul and ISAaml were restored (FIG. 27C), indicating that the lack of 3 '-end processing was responsible for their weak activities in cassette (a).
  • cassette (f) Utilizing these insights into coRNA processing, a single cassette (f) was designed where the 16-nucleotide guide region was capped by an intact 127-nucleotide coRNA scaffold from ISYmul. This design demonstrated efficient plasmid interference activity (FIG. 27F). However, capping the 16-nucleotide guide region with an extended coRNA scaffold resulted in diminished activity (FIG. 27F), suggesting inefficient 3 ’-end processing. The design of cassette (I) can potentially be leveraged for multiplex editing by alternating the intact 127-nucleotide coRNA scaffold and guide region.
  • TnpB Transposon-associated TnpB is a programmable RNA-guided DNA endonuclease. Nature, 599(7886), 692-696.
  • TnpB structure reveals minimal functional core of Casl2 nuclease family. Nature, 616(7956), 384- 389.
  • Table 4 Vectors for TnpB-mediated genome editing
  • Example 12 Additional Sequences for Example 12 (In Table 6. the gRNA sequences are listed as DNA sequences that are the equivalent RNA sequences (except the RNA will have the T replaced with a U)).
  • Table 8 Various TnpB sequences associated with Example 13 Table 9: TAM Sequences and gRNA Sequences for Selected TnpBs (In Table 9, the gRNA sequences are listed as DNA sequences that are the equivalent RNA sequences (except the RNA will have the T replaced with a U)
  • Example 14 Sequences Associated with Example 14 (In Table 10, the coRNA sequences are listed as DNA sequences that are the equivalent RNA sequences (except the RNA will have the T replaced with a U))
  • Table 11 Additional Sequences Associated with Example 14
  • Table 14 Selected TAM and gRNA sequences (In Table 14, the gRNA sequences are listed as DNA sequences that are the equivalent RNA sequences (except the RNA will have the T replaced with a U)
  • Table 15 Editing Efficeincy of Selected TnpBs - Example 19
  • Table 16 Various spacer sequences for selected TnpBs associated with Example 20
  • Table 17 associated with
  • Table 18 Selected coRNA sequences associated with Example 22 (In Table 18, the coRNA sequences are listed as DNA sequences that are the equivalent RNA sequences (except the RNA will have the T replaced with a U))
  • ISDgelO TnpB amino acid sequence SEQ ID NO: 48
  • ISDra2 TnpB amino acid sequence SEQ ID NO: 49
  • ISDgelO coRNA DNA sequence SEQ ID NO: 52
  • TnpB30 amino acid sequence (SEQ ID NO: 702)

Landscapes

  • Health & Medical Sciences (AREA)
  • Life Sciences & Earth Sciences (AREA)
  • Genetics & Genomics (AREA)
  • Engineering & Computer Science (AREA)
  • Chemical & Material Sciences (AREA)
  • Bioinformatics & Cheminformatics (AREA)
  • Biomedical Technology (AREA)
  • Biotechnology (AREA)
  • Organic Chemistry (AREA)
  • Zoology (AREA)
  • Wood Science & Technology (AREA)
  • Molecular Biology (AREA)
  • General Engineering & Computer Science (AREA)
  • Microbiology (AREA)
  • Biochemistry (AREA)
  • General Health & Medical Sciences (AREA)
  • Plant Pathology (AREA)
  • Biophysics (AREA)
  • Physics & Mathematics (AREA)
  • Cell Biology (AREA)
  • Medicinal Chemistry (AREA)
  • Micro-Organisms Or Cultivation Processes Thereof (AREA)
  • Breeding Of Plants And Reproduction By Means Of Culturing (AREA)

Abstract

L'invention concerne des compositions, des constructions, des vecteurs viraux et des procédés d'édition de gènes végétaux. Les compositions peuvent comprendre une enzyme bactérienne TnpB et un polynucléotide ayant une première partie conçue pour se lier à au moins une partie de l'enzyme bactérienne TnpB et une seconde partie conçue pour se lier à une cible génomique végétale.
PCT/US2024/036209 2023-06-29 2024-06-28 Compositions, constructions, vecteurs viraux végétaux et procédés d'édition de gènes végétaux Ceased WO2025007023A2 (fr)

Priority Applications (1)

Application Number Priority Date Filing Date Title
EP24833079.7A EP4735609A2 (fr) 2023-06-29 2024-06-28 Compositions, constructions, vecteurs viraux végétaux et procédés d'édition de gènes végétaux

Applications Claiming Priority (6)

Application Number Priority Date Filing Date Title
US202363511070P 2023-06-29 2023-06-29
US63/511,070 2023-06-29
US202363520258P 2023-08-17 2023-08-17
US63/520,258 2023-08-17
US202463566666P 2024-03-18 2024-03-18
US63/566,666 2024-03-18

Publications (2)

Publication Number Publication Date
WO2025007023A2 true WO2025007023A2 (fr) 2025-01-02
WO2025007023A3 WO2025007023A3 (fr) 2025-03-27

Family

ID=93940080

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/US2024/036209 Ceased WO2025007023A2 (fr) 2023-06-29 2024-06-28 Compositions, constructions, vecteurs viraux végétaux et procédés d'édition de gènes végétaux

Country Status (2)

Country Link
EP (1) EP4735609A2 (fr)
WO (1) WO2025007023A2 (fr)

Family Cites Families (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
EP3500967A1 (fr) * 2016-08-17 2019-06-26 The Broad Institute, Inc. Méthodes d'identification de systèmes crispr-cas de classe 2
WO2023097224A1 (fr) * 2021-11-23 2023-06-01 The Broad Institute, Inc. Nucléases isrb reprogrammables et leurs utilisations
WO2024047552A1 (fr) * 2022-08-31 2024-03-07 Indian Council Of Agricultural Research Systèmes et procédés d'édition ciblée de génome dans des plantes
EP4602156A2 (fr) * 2022-10-11 2025-08-20 The Trustees of Columbia University in the City of New York Compositions, méthodes et systèmes de modification d'adn

Also Published As

Publication number Publication date
EP4735609A2 (fr) 2026-05-06
WO2025007023A3 (fr) 2025-03-27

Similar Documents

Publication Publication Date Title
Li et al. Highly efficient heritable genome editing in wheat using an RNA virus and bypassing tissue culture
Razzaq et al. Modern trends in plant genome editing: an inclusive review of the CRISPR/Cas9 toolbox
Sedeek et al. Plant genome engineering for targeted improvement of crop traits
Pyott et al. Engineering of CRISPR/Cas9‐mediated potyvirus resistance in transgene‐free Arabidopsis plants
Yu et al. CRISPR/Cas9-induced targeted mutagenesis and gene replacement to generate long-shelf life tomato lines
CA2964796C (fr) Procedes et compositions pour l'edition genomique multiplex a guidage arn et autres technologies d'arn
Kaur et al. Genome editing: a promising approach for achieving abiotic stress tolerance in plants
Prado et al. CRISPR technology towards genome editing of the perennial and semi-perennial crops citrus, coffee and sugarcane
Cao et al. Genome-wide identification of cytosine-5 DNA methyltransferases and demethylases in Solanum lycopersicum
US12385053B2 (en) Genomic alteration of plant germline
Khan et al. Genome editing in cotton: challenges and opportunities
EP4461813A2 (fr) Compositions et procédés pour augmenter la durée de conservation de banane
US20180116141A1 (en) Haploid induction
Arnal et al. A restorer‐of‐fertility like pentatricopeptide repeat gene directs ribonucleolytic processing within the coding sequence of rps3‐rpl16 and orf240a mitochondrial transcripts in A rabidopsis thaliana
Sattar et al. CRISPR/Cas9: a new genome editing tool to accelerate cotton (Gossypium spp.) breeding
Uranga Virus-induced genome editing: Methods and applications in plant breeding
Das et al. Genome editing of rice by CRISPR-Cas: end-to-end pipeline for crop improvement
JP2018529348A (ja) 半数体及びそれに続く倍加半数体植物の産生方法
WO2025007023A2 (fr) Compositions, constructions, vecteurs viraux végétaux et procédés d'édition de gènes végétaux
Liang et al. An efficient targeted mutagenesis system using CRISPR/Cas in monocotyledons
WO2025076141A1 (fr) Administration virale de grna au greffon
CN121195071A (zh) 乙酰辅酶a羧化酶突变体
US20210348177A1 (en) Generation of heritably gene-edited plants without tissue culture
CA3153420A1 (fr) Modification genetique de plantes
Phogat et al. Recent Advances in CRISPR-Cas for Climate Resilient Agriculture in Cereals

Legal Events

Date Code Title Description
WWE Wipo information: entry into national phase

Ref document number: 2024833079

Country of ref document: EP

NENP Non-entry into the national phase

Ref country code: DE

ENP Entry into the national phase

Ref document number: 2024833079

Country of ref document: EP

Effective date: 20260129

ENP Entry into the national phase

Ref document number: 2024833079

Country of ref document: EP

Effective date: 20260129

121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 24833079

Country of ref document: EP

Kind code of ref document: A2