WO2025004002A2 - Traitement de la maladie de pompe - Google Patents

Traitement de la maladie de pompe Download PDF

Info

Publication number
WO2025004002A2
WO2025004002A2 PCT/IB2024/056362 IB2024056362W WO2025004002A2 WO 2025004002 A2 WO2025004002 A2 WO 2025004002A2 IB 2024056362 W IB2024056362 W IB 2024056362W WO 2025004002 A2 WO2025004002 A2 WO 2025004002A2
Authority
WO
WIPO (PCT)
Prior art keywords
seq
gaa
sequence
polypeptide
variant protein
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/IB2024/056362
Other languages
English (en)
Other versions
WO2025004002A3 (fr
Inventor
Ruby BOYANAPALLI
Simon Moore
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Takeda Pharmaceutical Co Ltd
Original Assignee
Takeda Pharmaceutical Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Takeda Pharmaceutical Co Ltd filed Critical Takeda Pharmaceutical Co Ltd
Publication of WO2025004002A2 publication Critical patent/WO2025004002A2/fr
Publication of WO2025004002A3 publication Critical patent/WO2025004002A3/fr
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Classifications

    • A—HUMAN NECESSITIES
    • A61—MEDICAL OR VETERINARY SCIENCE; HYGIENE
    • A61P—SPECIFIC THERAPEUTIC ACTIVITY OF CHEMICAL COMPOUNDS OR MEDICINAL PREPARATIONS
    • A61P21/00—Drugs for disorders of the muscular or neuromuscular system
    • A—HUMAN NECESSITIES
    • A01—AGRICULTURE; FORESTRY; ANIMAL HUSBANDRY; HUNTING; TRAPPING; FISHING
    • A01K—ANIMAL HUSBANDRY; AVICULTURE; APICULTURE; PISCICULTURE; FISHING; REARING OR BREEDING ANIMALS, NOT OTHERWISE PROVIDED FOR; NEW BREEDS OF ANIMALS
    • A01K67/00—Rearing or breeding animals, not otherwise provided for; New or modified breeds of animals
    • A01K67/027—New or modified breeds of vertebrates
    • A01K67/0275—Genetically modified vertebrates, e.g. transgenic
    • A—HUMAN NECESSITIES
    • A61—MEDICAL OR VETERINARY SCIENCE; HYGIENE
    • A61K—PREPARATIONS FOR MEDICAL, DENTAL OR TOILETRY PURPOSES
    • A61K48/00—Medicinal preparations containing genetic material which is inserted into cells of the living body to treat genetic diseases; Gene therapy
    • A61K48/005—Medicinal preparations containing genetic material which is inserted into cells of the living body to treat genetic diseases; Gene therapy characterised by an aspect of the 'active' part of the composition delivered, i.e. the nucleic acid delivered
    • A—HUMAN NECESSITIES
    • A61—MEDICAL OR VETERINARY SCIENCE; HYGIENE
    • A61K—PREPARATIONS FOR MEDICAL, DENTAL OR TOILETRY PURPOSES
    • A61K48/00—Medicinal preparations containing genetic material which is inserted into cells of the living body to treat genetic diseases; Gene therapy
    • A61K48/005—Medicinal preparations containing genetic material which is inserted into cells of the living body to treat genetic diseases; Gene therapy characterised by an aspect of the 'active' part of the composition delivered, i.e. the nucleic acid delivered
    • A61K48/0058—Nucleic acids adapted for tissue specific expression, e.g. having tissue specific promoters as part of a contruct
    • A—HUMAN NECESSITIES
    • A61—MEDICAL OR VETERINARY SCIENCE; HYGIENE
    • A61K—PREPARATIONS FOR MEDICAL, DENTAL OR TOILETRY PURPOSES
    • A61K48/00—Medicinal preparations containing genetic material which is inserted into cells of the living body to treat genetic diseases; Gene therapy
    • A61K48/005—Medicinal preparations containing genetic material which is inserted into cells of the living body to treat genetic diseases; Gene therapy characterised by an aspect of the 'active' part of the composition delivered, i.e. the nucleic acid delivered
    • A61K48/0066—Manipulation of the nucleic acid to modify its expression pattern, e.g. enhance its duration of expression, achieved by the presence of particular introns in the delivered nucleic acid
    • A—HUMAN NECESSITIES
    • A61—MEDICAL OR VETERINARY SCIENCE; HYGIENE
    • A61P—SPECIFIC THERAPEUTIC ACTIVITY OF CHEMICAL COMPOUNDS OR MEDICINAL PREPARATIONS
    • A61P43/00—Drugs for specific purposes, not provided for in groups A61P1/00-A61P41/00
    • C—CHEMISTRY; METALLURGY
    • C12—BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12N—MICROORGANISMS OR ENZYMES; COMPOSITIONS THEREOF; PROPAGATING, PRESERVING, OR MAINTAINING MICROORGANISMS; MUTATION OR GENETIC ENGINEERING; CULTURE MEDIA
    • C12N15/00—Mutation or genetic engineering; DNA or RNA concerning genetic engineering, vectors, e.g. plasmids, or their isolation, preparation or purification; Use of hosts therefor
    • C12N15/09—Recombinant DNA-technology
    • C12N15/63—Introduction of foreign genetic material using vectors; Vectors; Use of hosts therefor; Regulation of expression
    • C12N15/79—Vectors or expression systems specially adapted for eukaryotic hosts
    • C12N15/85—Vectors or expression systems specially adapted for eukaryotic hosts for animal cells
    • C—CHEMISTRY; METALLURGY
    • C12—BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12N—MICROORGANISMS OR ENZYMES; COMPOSITIONS THEREOF; PROPAGATING, PRESERVING, OR MAINTAINING MICROORGANISMS; MUTATION OR GENETIC ENGINEERING; CULTURE MEDIA
    • C12N15/00—Mutation or genetic engineering; DNA or RNA concerning genetic engineering, vectors, e.g. plasmids, or their isolation, preparation or purification; Use of hosts therefor
    • C12N15/09—Recombinant DNA-technology
    • C12N15/63—Introduction of foreign genetic material using vectors; Vectors; Use of hosts therefor; Regulation of expression
    • C12N15/79—Vectors or expression systems specially adapted for eukaryotic hosts
    • C12N15/85—Vectors or expression systems specially adapted for eukaryotic hosts for animal cells
    • C12N15/86—Viral vectors
    • C—CHEMISTRY; METALLURGY
    • C12—BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12N—MICROORGANISMS OR ENZYMES; COMPOSITIONS THEREOF; PROPAGATING, PRESERVING, OR MAINTAINING MICROORGANISMS; MUTATION OR GENETIC ENGINEERING; CULTURE MEDIA
    • C12N9/00—Enzymes; Proenzymes; Compositions thereof; Processes for preparing, activating, inhibiting, separating or purifying enzymes
    • C12N9/14—Hydrolases (3)
    • C12N9/24—Hydrolases (3) acting on glycosyl compounds (3.2)
    • C12N9/2402—Hydrolases (3) acting on glycosyl compounds (3.2) hydrolysing O- and S- glycosyl compounds (3.2.1)
    • C12N9/2405—Glucanases
    • C12N9/2408—Glucanases acting on alpha -1,4-glucosidic bonds
    • C—CHEMISTRY; METALLURGY
    • C12—BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12N—MICROORGANISMS OR ENZYMES; COMPOSITIONS THEREOF; PROPAGATING, PRESERVING, OR MAINTAINING MICROORGANISMS; MUTATION OR GENETIC ENGINEERING; CULTURE MEDIA
    • C12N9/00—Enzymes; Proenzymes; Compositions thereof; Processes for preparing, activating, inhibiting, separating or purifying enzymes
    • C12N9/14—Hydrolases (3)
    • C12N9/24—Hydrolases (3) acting on glycosyl compounds (3.2)
    • C12N9/2402—Hydrolases (3) acting on glycosyl compounds (3.2) hydrolysing O- and S- glycosyl compounds (3.2.1)
    • C12N9/2405—Glucanases
    • C12N9/2434—Glucanases acting on beta-1,4-glucosidic bonds
    • C12N9/2445—Beta-glucosidase (3.2.1.21)
    • A—HUMAN NECESSITIES
    • A01—AGRICULTURE; FORESTRY; ANIMAL HUSBANDRY; HUNTING; TRAPPING; FISHING
    • A01K—ANIMAL HUSBANDRY; AVICULTURE; APICULTURE; PISCICULTURE; FISHING; REARING OR BREEDING ANIMALS, NOT OTHERWISE PROVIDED FOR; NEW BREEDS OF ANIMALS
    • A01K2217/00—Genetically modified animals
    • A01K2217/07—Animals genetically altered by homologous recombination
    • A01K2217/075—Animals genetically altered by homologous recombination inducing loss of function, i.e. knock out
    • A—HUMAN NECESSITIES
    • A01—AGRICULTURE; FORESTRY; ANIMAL HUSBANDRY; HUNTING; TRAPPING; FISHING
    • A01K—ANIMAL HUSBANDRY; AVICULTURE; APICULTURE; PISCICULTURE; FISHING; REARING OR BREEDING ANIMALS, NOT OTHERWISE PROVIDED FOR; NEW BREEDS OF ANIMALS
    • A01K2227/00—Animals characterised by species
    • A01K2227/10—Mammal
    • A01K2227/105—Murine
    • A—HUMAN NECESSITIES
    • A01—AGRICULTURE; FORESTRY; ANIMAL HUSBANDRY; HUNTING; TRAPPING; FISHING
    • A01K—ANIMAL HUSBANDRY; AVICULTURE; APICULTURE; PISCICULTURE; FISHING; REARING OR BREEDING ANIMALS, NOT OTHERWISE PROVIDED FOR; NEW BREEDS OF ANIMALS
    • A01K2267/00—Animals characterised by purpose
    • A01K2267/03—Animal model, e.g. for test or diseases
    • A01K2267/0306—Animal model for genetic diseases
    • C—CHEMISTRY; METALLURGY
    • C12—BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12N—MICROORGANISMS OR ENZYMES; COMPOSITIONS THEREOF; PROPAGATING, PRESERVING, OR MAINTAINING MICROORGANISMS; MUTATION OR GENETIC ENGINEERING; CULTURE MEDIA
    • C12N2750/00—MICROORGANISMS OR ENZYMES; COMPOSITIONS THEREOF; PROPAGATING, PRESERVING, OR MAINTAINING MICROORGANISMS; MUTATION OR GENETIC ENGINEERING; CULTURE MEDIA ssDNA viruses
    • C12N2750/00011—Details
    • C12N2750/14011—Parvoviridae
    • C12N2750/14111—Dependovirus, e.g. adenoassociated viruses
    • C12N2750/14141—Use of virus, viral particle or viral elements as a vector
    • C12N2750/14143—Use of virus, viral particle or viral elements as a vector viral genome or elements thereof as genetic vector
    • C—CHEMISTRY; METALLURGY
    • C12—BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12N—MICROORGANISMS OR ENZYMES; COMPOSITIONS THEREOF; PROPAGATING, PRESERVING, OR MAINTAINING MICROORGANISMS; MUTATION OR GENETIC ENGINEERING; CULTURE MEDIA
    • C12N2750/00—MICROORGANISMS OR ENZYMES; COMPOSITIONS THEREOF; PROPAGATING, PRESERVING, OR MAINTAINING MICROORGANISMS; MUTATION OR GENETIC ENGINEERING; CULTURE MEDIA ssDNA viruses
    • C12N2750/00011—Details
    • C12N2750/14011—Parvoviridae
    • C12N2750/14111—Dependovirus, e.g. adenoassociated viruses
    • C12N2750/14141—Use of virus, viral particle or viral elements as a vector
    • C12N2750/14145—Special targeting system for viral vectors
    • C—CHEMISTRY; METALLURGY
    • C12—BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12N—MICROORGANISMS OR ENZYMES; COMPOSITIONS THEREOF; PROPAGATING, PRESERVING, OR MAINTAINING MICROORGANISMS; MUTATION OR GENETIC ENGINEERING; CULTURE MEDIA
    • C12N2800/00—Nucleic acids vectors
    • C12N2800/22—Vectors comprising a coding region that has been codon optimised for expression in a respective host
    • C—CHEMISTRY; METALLURGY
    • C12—BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12N—MICROORGANISMS OR ENZYMES; COMPOSITIONS THEREOF; PROPAGATING, PRESERVING, OR MAINTAINING MICROORGANISMS; MUTATION OR GENETIC ENGINEERING; CULTURE MEDIA
    • C12N2830/00—Vector systems having a special element relevant for transcription
    • C12N2830/008—Vector systems having a special element relevant for transcription cell type or tissue specific enhancer/promoter combination
    • C—CHEMISTRY; METALLURGY
    • C12—BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12N—MICROORGANISMS OR ENZYMES; COMPOSITIONS THEREOF; PROPAGATING, PRESERVING, OR MAINTAINING MICROORGANISMS; MUTATION OR GENETIC ENGINEERING; CULTURE MEDIA
    • C12N2830/00—Vector systems having a special element relevant for transcription
    • C12N2830/15—Vector systems having a special element relevant for transcription chimeric enhancer/promoter combination
    • C—CHEMISTRY; METALLURGY
    • C12—BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12N—MICROORGANISMS OR ENZYMES; COMPOSITIONS THEREOF; PROPAGATING, PRESERVING, OR MAINTAINING MICROORGANISMS; MUTATION OR GENETIC ENGINEERING; CULTURE MEDIA
    • C12N2830/00—Vector systems having a special element relevant for transcription
    • C12N2830/42—Vector systems having a special element relevant for transcription being an intron or intervening sequence for splicing and/or stability of RNA
    • C—CHEMISTRY; METALLURGY
    • C12—BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12N—MICROORGANISMS OR ENZYMES; COMPOSITIONS THEREOF; PROPAGATING, PRESERVING, OR MAINTAINING MICROORGANISMS; MUTATION OR GENETIC ENGINEERING; CULTURE MEDIA
    • C12N2830/00—Vector systems having a special element relevant for transcription
    • C12N2830/50—Vector systems having a special element relevant for transcription regulating RNA stability, not being an intron, e.g. poly A signal

Definitions

  • the invention relates to acid alpha-glucosidase (GAA) gene therapy.
  • Acid alpha-glucosidase is an enzyme that is responsible for the critical degradation of glycogen in lysosomes of cells. Loss of its activity leads to progressive intralysosomal accumulation of undegraded glycogen and lysosomal distention.
  • Pompe disease is caused by mutations and reduced activity of the GAA gene (gad).
  • PD can be broadly classified into infantile-onset (IOPD) or late-onset (LOPD) PD.
  • IOPD patients have under 1% GAA activity, develop cardiomegaly, muscle weakness with hypotonia, hepatomegaly, breathing problems and die within the first year of life if left not treated.
  • LOPD patients have at least 1% GAA activity and manifest a less severe phenotype but present with progressive limb muscle weakness and respiratory insufficiency.
  • ERT enzyme replacement therapy
  • Lumizyme® marketed as Myozyme® outside of the United States; Sanofi Genzyme
  • Pompe disease remains a devastating illness.
  • ERT enzyme replacement therapy
  • LOPD is heterogenous and the severity falls along a spectrum, many patients still lose independent mobility and/or require ventilator support as their symptoms progress. Patients reach a clinical plateau within 2-3 years of treatment and some show a decline over time (Harfouche (2020) J. Patient Rep. Outcomes 4(1): 83).
  • ERT therapy The primary deficiencies of ERT therapy are: (1) detrimental immune responses including neutralizing antibodies against recombinant GAA enzyme, especially in cross-reactive immune-material (CRIM) negative patients; (2) poor uptake of the GAA enzyme by muscle cells from circulation; (3) limited availability of the GAA enzyme in circulation (85% taken up by the liver); (4) reduced stability of the GAA enzyme at neutral pH; and (4) progressive endosomal dysfunction reducing efficacy of the endogenous enzyme delivery to lysosomes.
  • Other complications associated with ERT include infusion site reactions and the requirement of bi-weekly or even weekly (in severe cases) infusions. These deficiencies of ERT translate to a poor quality of life, indicating a sustained unmet need for these patients.
  • ERTs are being developed to address some of these problems and to improve the current standard of care (SOC).
  • SOC standard of care
  • Such strategies are focused on improving uptake and bioavailability of GAA into muscles and include: (1) development of a GAA with high mannose 6-phosphate (M6P) content to improve uptake from circulation; (2) chimeric GAA variants with synthetic uptake domains; (3) administering beta-2 agonists to upregulate the expression of the cation-independent M6P receptor (CI-MPR) to improve cellular uptake (Farah et al. (2014) FASEB J. 28(5):2272-2280); and (4) combining ERT with pharmacological chaperones to improve GAA enzyme stability in plasma (Okumiya et al.
  • M6P mannose 6-phosphate
  • CI-MPR cation-independent M6P receptor
  • glycogen accumulates in virtually all tissues of PD patients, the clinical manifestations are predominantly observed in the skeletal, cardiac, and respiratory muscles.
  • the major unmet needs in PD are due to limited availability of GAA enzyme in the respiratory and deep skeletal muscles that is a consequence of not only low levels of circulating GAA enzyme and poor GAA enzyme uptake into muscle cells, but also of reduced enzyme stability and exaggerated immune responses.
  • GAA gene therapy has the potential to have a lasting therapeutic effect on patients suffering from PD by delivering continuous, high exposure of the missing enzyme to the affected tissues.
  • GT GAA gene therapy
  • the present disclosure relates to a gene therapy that addresses the shortcomings of ERTs and competitor gene therapy (GT) candidates, thereby providing a transformative therapy for Pompe disease (PD) patients.
  • GT competitor gene therapy
  • the present therapy targets hard to treat muscle tissues by combining (1) an engineered AAV capsid that efficiently delivers the genetic payload to muscles and can be dosed into environmentally seropositive patients, (2) transcriptional elements that drive maximal expression of GAA in muscles, and (3) an engineered GAA protein variant that has enhanced stability, uptake, catalytic activity, and a reduced immunogenic profile.
  • the present GAA GT efficiently delivers and provides a high, continuous exposure of missing GAA enzyme to GAA-deficient muscle tissues by optimizing tissue-specific delivery, expression and function of the GAA enzyme and as a result will significantly improve patient quality of life and extend the lifespan of Pompe disease patients.
  • the present disclosure concerns methods and compositions to alleviate the primary unmet need of Pompe patients that leads to morbidity and mortality by achieving therapeutic GAA protein levels in the lysosomes of diaphragm, cardiac, skeletal, and smooth muscle cells.
  • the muscle-tropic AAV9 capsid influences delivery (z.e., transduction) of the recombinant genome comprising the engineered GAA transgene to the targeted muscle tissue.
  • the muscle-specific promoter and enhancer elements drive strong expression of the GAA transgene in muscle cells.
  • the engineered GAA protein is processed and trafficked to cellular lysosomes where it can act to break down glycogen.
  • a portion of the engineered GAA that is expressed in transduced cells is secreted and taken up by surrounding non-transduced cells. This local cross-correction within muscle allows for better treatment of the difficult-to-reach deep skeletal muscle tissues.
  • the therapeutic mechanism of action does not require transport into the serum and passive uptake by muscle cells, a major limitation of liver-directed GT approaches.
  • the disclosure provides a nucleic acid encoding an acid alphaglucosidase (GAA) protein, the nucleic acid comprising a first polynucleotide sequence that is at least 95%, at least 96%, at least 97% or at least 98% identical to CO3-MP-46-NA (SEQ ID NO:36).
  • the first polynucleotide sequence is at least 99% identical to CO3-MP-46-NA (SEQ ID NO:36).
  • the first polynucleotide sequence is at least 99.5% identical to CO3-MP-46-NA (SEQ ID NO:36).
  • the disclosure provides a nucleic acid encoding an acid alphaglucosidase (GAA) protein, the nucleic acid comprising a first polynucleotide sequence that is at least 95%, at least 96%, at least 97% or at least 98% identical to CO3-MP-WT-NA (SEQ ID NO:34).
  • GAA acid alphaglucosidase
  • the first polynucleotide sequence is at least 99% identical to CO3-MP-WT-NA (SEQ ID NO:34).
  • the first polynucleotide sequence is at least 99.5% identical to CO3-MP-WT-NA (SEQ ID NO:34).
  • the nucleic acid further comprises a second polynucleotide sequence that is at least 95%, at least 96%, at least 97%, or at least 98% identical to CO3-PP- WT-NA (SEQ ID NO: 38).
  • the second polynucleotide sequence is at least 99% identical to CO3-PP-WT-NA (SEQ ID NO:38).
  • the second polynucleotide sequence is CO3-PP-WT-NA (SEQ ID NO: 38).
  • the nucleic acid further comprises a second polynucleotide sequence that is at least 95%, at least 96%, at least 97%, or at least 98% identical to CO3-PP- 46-NA (SEQ ID NO:40).
  • the second polynucleotide sequence is at least 99% identical to CO3-PP-46-NA (SEQ ID NO:40).
  • the second polynucleotide sequence is CO3-PP-46-NA (SEQ ID NO:40).
  • the nucleic acid further comprises a third polynucleotide sequence that is at least 95%, at least 96%, at least 97%, or at least 98% identical to CO3-SP- WT-NA (SEQ ID NO:42).
  • the third polynucleotide sequence is CO3-SP-WT-NA (SEQ ID NO:42).
  • the encoded GAA protein comprises a first polypeptide sequence that is at least 95%, at least 96%, at least 97% or at least 98% identical to MP-46- AA (SEQ ID NO:37).
  • the first polypeptide sequence is at least 99% identical to MP-46-AA (SEQ ID NO:37).
  • the first polypeptide sequence is at least 99.5% identical to MP-46-AA (SEQ ID NO:37).
  • the encoded GAA protein comprises an amino acid substitution selected from the group consisting of L32W, L36S, L37T, P47Q, Q58V, A70L, P86E, D95E, L117D, S143Q, T151I, T158S, T274N, R275K, A445G, T494E, E530V, N535R, L577T, L650G, S676D, L678H, T700G, A719H, A758P, A820E, Q838K, L868F, L879E, R891H, Q902G, V921R, and S940A, numbered relative to the full-length wild-type GAA protein sequence of FL-WT-AA (SEQ ID NO:2).
  • the encoded GAA protein comprises a first polypeptide sequence that is at least 95%, at least 96%, at least 97% or at least 98% identical to MP-WT- AA (SEQ ID NO:35).
  • the first polypeptide sequence is at least 99% identical to MP-WT-AA (SEQ ID NO: 35).
  • the first polypeptide sequence is at least 99.5% identical to MP-WT-AA (SEQ ID NO: 35).
  • the encoded GAA protein further comprises a second polypeptide sequence that is at least 95% identical to PP-WT-AA (SEQ ID NO:39).
  • the second polypeptide sequence is at least 97% identical to PP-WT-AA (SEQ ID NO:39).
  • the second polypeptide sequence is PP-WT-AA (SEQ ID NO:39).
  • the encoded GAA protein further comprises a third polypeptide sequence that is at least 95% identical to SP-WT-AA (SEQ ID NO:43). [0038] In some embodiments, the third polypeptide sequence is SP-WT-AA (SEQ ID NO:43).
  • the encoded GAA protein further comprises an amino acid substitution selected from the group consisting of L117D, L313V, A497D, S620D, S735W, L744M, and G850S, numbered relative to the full-length wild-type GAA protein sequence of FL-WT-AA (SEQ ID NO:2).
  • the encoded GAA protein further comprises an amino acid substitution selected from the group consisting of L32W, L36S, L37T, P47Q, Q58V, A70L, P86E, D95E, S143Q, T158S, T274N, R275K, A445G, T494E, E530V, N535R, L577T, L678T, T700G, A719H, A758P, A820E, Q838K, L868F, L879E, R891H, Q902G, V921R, and S940A, numbered relative to the full-length wild-type GAA protein sequence of FL-WT- AA (SEQ ID NO:2).
  • the encoded GAA protein further comprises an amino acid substitution selected from the group consisting of L650G, L650S L650T, L650E, L650Y, L650F, S676D, and L678H, numbered relative to the full-length wild-type GAA protein sequence of FL-WT-AA (SEQ ID NO:2).
  • the encoded GAA protein further comprises an amino acid substitution selected from the group consisting of T 15 II, L650G, S676D, and L678H, numbered relative to the full-length wild-type GAA protein sequence of FL-WT-AA (SEQ ID N0:2).
  • the polynucleotide sequence encoding the GAA protein comprises no more than five CpG dinucleotides.
  • the polynucleotide sequence encoding the GAA protein comprises a GC content of from 55% to 60%.
  • the first polynucleotide sequence is CO3-MP-46-NA (SEQ ID NO:36).
  • polynucleotide sequence of CO3-FL-46-NA (SEQ ID NO:32).
  • the polynucleotide sequence of CO3-FL-46-dNA (SEQ ID NO: 103).
  • the first polynucleotide sequence is CO3-MP-WT-NA (SEQ ID NO:34).
  • polynucleotide sequence of CO3-FL-WT-NA (SEQ ID NO:31).
  • the present disclosure provides an expression cassette comprising a nucleic acid encoding an acid alpha-glucosidase (GAA) and at least one regulatory nucleic acid sequence operably linked to the sequence encoding the GAA protein.
  • GAA acid alpha-glucosidase
  • the at least one regulatory nucleic acid sequence is selected from the group consisting of a promoter, an enhancer, an intron, a post-transcriptional regulatory element, an inverted terminal repeat (ITR), a polyadenylation (poly A) sequence, and a combination thereof.
  • the at least one regulatory nucleic acid sequence comprises a promoter.
  • the promoter is a muscle-specific promoter.
  • the muscle-specific promoter comprises a polynucleotide sequence that is at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SPc512_NA (SEQ ID NO:46).
  • the muscle-specific promoter comprises the polynucleotide sequence of SPc512_NA (SEQ ID NO:46).
  • the muscle-specific promoter comprises a polynucleotide sequence that is at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to 77sDesmin_NA (SEQ ID NO:47).
  • the muscle-specific promoter comprises the polynucleotide sequence of 77sDesmin_NA (SEQ ID NO:47).
  • the at least one regulatory nucleic acid sequence comprises an enhancer.
  • the enhancer is a muscle-specific enhancer.
  • the muscle-specific enhancer comprises a polynucleotide sequence that is at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to Dph-CRE04_NA (SEQ ID NO:48). [0061] In some embodiments, the muscle-specific enhancer comprises the polynucleotide sequence of Dph-CRE04_NA (SEQ ID NO:48).
  • the muscle-specific enhancer comprises a polynucleotide sequence that is at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to sk-SH4_NA (SEQ ID NO:49).
  • the muscle-specific enhancer comprises the polynucleotide sequence of sk-SH4_NA (SEQ ID NO:49).
  • the at least one regulatory nucleic acid sequence comprises an intron.
  • the intron comprises a polynucleotide sequence that is at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to MVM_NA (SEQ ID NO:50).
  • the intron comprises the polynucleotide sequence of MVM NA (SEQ ID NO: 50).
  • the present disclosure provides a mammalian expression vector comprising an expression cassette comprising a nucleic acid encoding an acid alphaglucosidase (GAA) and at least one regulatory nucleic acid sequence operably linked to the sequence encoding the GAA protein.
  • GAA acid alphaglucosidase
  • the mammalian expression vector comprises an adeno- associated virus (AAV) vector.
  • AAV adeno- associated virus
  • the AAV vector comprises an AAV8 or AAV9 capsid polypeptide encapsidating the expression cassette.
  • the AAV vector comprises an engineered capsid polypeptide encapsidating the expression cassette.
  • the present disclosure provides a host cell comprising a nucleic acid encoding an acid alpha-glucosidase (GAA) protein.
  • GAA acid alpha-glucosidase
  • the host cell comprises an expression cassette comprising a nucleic acid encoding an acid alpha-glucosidase (GAA) and at least one regulatory nucleic acid sequence operably linked to the sequence encoding the GAA protein.
  • GAA acid alpha-glucosidase
  • the host cell further comprises a nucleic acid encoding an AAV capsid polypeptide.
  • the AAV capsid polypeptide is an AAV8 or AAV9 capsid polypeptide.
  • the host cell further comprises a nucleic acid encoding a viral helper gene selected from the group consisting of E4, E2a, and VA.
  • the present disclosure provides a recombinant acid alpha-glucosidase (GAA) variant protein, wherein the GAA variant protein comprises an amino acid substitution selected from the group consisting of L117D, T151I, L313V, L650G, L650S, L650T, L650E, L650Y, L650F, S676D, L678H, and L868F, numbered relative to the full- length wild-type GAA protein sequence of FL-WT-AA (SEQ ID NO:2).
  • GAA acid alpha-glucosidase
  • the recombinant GAA variant protein comprises an amino acid substitution selected from the group consisting of L32W, L36S, L37T, P47Q, Q58V, A70L, P86E, D95E, S143Q, T158S, T274N, R275K, A445G, T494E, E530V, N535R, L577T, L678T, T700G, A719H, A758P, A820E, Q838K, L868F, L879E, R891H, Q902G, V921R, and S940A, numbered relative to the full-length wild-type GAA protein sequence of FL-WT-AA (SEQ ID NO:2).
  • the recombinant GAA variant protein comprises an amino acid substitution selected from the group consisting of T 15 II, L650G, S676D, and L678H, numbered relative to the full-length wild-type GAA protein sequence of FL-WT-AA (SEQ ID N0:2).
  • the recombinant GAA variant protein comprises a first polypeptide sequence that is at least 95%, at least 96%, at least 97% or at least 98% identical to MP-46-AA (SEQ ID NO:37).
  • the first polypeptide sequence is at least 99% identical to MP-46-AA (SEQ ID NO: 37).
  • the first polypeptide sequence is at least 99.5% identical to MP-46-AA (SEQ ID NO:37).
  • the first polypeptide sequence is MP-46-AA (SEQ ID NO:37).
  • the recombinant GAA variant protein further comprises a second polypeptide sequence that is at least 95% identical to PP-WT-AA (SEQ ID NO:39).
  • the second polypeptide sequence is at least 97% identical to PP-WT-AA (SEQ ID NO:39).
  • the second polypeptide sequence is PP-WT-AA (SEQ ID NO:39).
  • the recombinant GAA variant protein further comprises a third polypeptide sequence that is at least 95% identical to SP-WT-AA (SEQ ID NO:43).
  • the third polypeptide sequence is SP-WT-AA (SEQ ID NO:43).
  • the recombinant GAA variant protein comprises the polypeptide sequence of FL-46-AA (SEQ ID NO:33).
  • the present disclosure provides a recombinant acid alpha-glucosidase (GAA) variant protein, wherein the GAA variant protein comprises a first polypeptide sequence that is at least 95%, at least 96%, at least 97% or at least 98% identical to MP-46- AA (SEQ ID NO:37) and wherein the GAA variant comprises one or more variant amino acids selected from the group consisting of L32W, L36S, L37T, P47Q, Q58V, A70L, P86E, D95E, L117D, S143Q, T151I, T158S, T274N, R275K, A445G, T494E, E530V, N535R, L577T, L650G, S676D, L678H, T700G, A719H, A758P, A820E, Q838K, L868F, L879E, R891H, Q902G, V921R, and S
  • the first polypeptide sequence is at least 99% identical to MP-46-AA (SEQ ID NO:37).
  • the first polypeptide sequence is at least 99.5% identical to MP-46-AA (SEQ ID NO:37).
  • the recombinant GAA variant protein comprises a tryptophan residue at position 32 (relative to SEQ ID NO:2).
  • the recombinant GAA variant protein comprises a serine residue at position 36 (relative to SEQ ID NO:2).
  • the recombinant GAA variant protein comprises a threonine residue at position 37 (relative to SEQ ID NO:2). [0095] In some embodiments, the recombinant GAA variant protein comprises a glutamine residue at position 47 (relative to SEQ ID NO:2).
  • the recombinant GAA variant protein comprises a valine residue at position 58 (relative to SEQ ID NO:2).
  • the recombinant GAA variant protein comprises a leucine residue at position 70 (relative to SEQ ID NO:2).
  • the recombinant GAA variant protein comprises a glutamic acid residue at position 86 (relative to SEQ ID NO:2).
  • the recombinant GAA variant protein comprises a glutamic acid residue at position 95 (relative to SEQ ID NO:2).
  • the recombinant GAA variant protein comprises an aspartic acid residue at position 117 (relative to SEQ ID NO:2).
  • the recombinant GAA variant protein comprises a glutamine residue at position 143 (relative to SEQ ID NO:2).
  • the recombinant GAA variant protein comprises an isoleucine residue at position 151 (relative to SEQ ID NO:2).
  • the recombinant GAA variant protein comprises a serine residue at position 158 (relative to SEQ ID NO:2).
  • the recombinant GAA variant protein comprises an asparagine residue at position 274 (relative to SEQ ID NO:2).
  • the recombinant GAA variant protein comprises a lysine residue at position 275 (relative to SEQ ID NO:2).
  • the recombinant GAA variant protein comprises a glycine residue at position 445 (relative to SEQ ID NO:2).
  • the recombinant GAA variant protein comprises a glutamic acid residue at position 494 (relative to SEQ ID NO:2).
  • the recombinant GAA variant protein comprises a valine residue at position 530 (relative to SEQ ID NO:2). [00109] In some embodiments, the recombinant GAA variant protein comprises an arginine residue at position 535 (relative to SEQ ID NO:2).
  • the recombinant GAA variant protein comprises a threonine residue at position 577 (relative to SEQ ID NO:2).
  • the recombinant GAA variant protein comprises a glycine residue at position 650 (relative to SEQ ID NO:2).
  • the recombinant GAA variant protein comprises an aspartic acid residue at position 676 (relative to SEQ ID NO:2).
  • the recombinant GAA variant protein comprises a histidine residue at position 678 (relative to SEQ ID NO:2).
  • the recombinant GAA variant protein comprises a glycine residue at position 700 (relative to SEQ ID NO:2).
  • the recombinant GAA variant protein comprises a histidine residue at position 719 (relative to SEQ ID NO:2).
  • the recombinant GAA variant protein comprises a proline residue at position 758 (relative to SEQ ID NO:2).
  • the recombinant GAA variant protein comprises a glutamic acid residue at position 820 (relative to SEQ ID NO:2).
  • the recombinant GAA variant protein comprises a lysine residue at position 838 (relative to SEQ ID NO:2).
  • the recombinant GAA variant protein comprises a phenylalanine residue at position 868 (relative to SEQ ID NO:2).
  • the recombinant GAA variant protein comprises a glutamic acid residue at position 879 (relative to SEQ ID NO:2).
  • the recombinant GAA variant protein comprises a histidine residue at position 891 (relative to SEQ ID NO:2).
  • the recombinant GAA variant protein comprises a glycine residue at position 902 (relative to SEQ ID NO:2). [00123] In some embodiments, the recombinant GAA variant protein comprises an arginine residue at position 921 (relative to SEQ ID NO:2).
  • the recombinant GAA variant protein comprises an alanine residue at position 940 (relative to SEQ ID NO:2).
  • the first polypeptide sequence is MP-46-AA (SEQ ID NO:37).
  • the recombinant GAA variant protein further comprises a second polypeptide sequence that is at least 95% identical to PP-WT-AA (SEQ ID NO:39).
  • the second polypeptide sequence is at least 97% identical to PP-WT-AA (SEQ ID NO:39).
  • the second polypeptide sequence is PP-WT-AA (SEQ ID NO:39).
  • the recombinant GAA variant protein further comprises a third polypeptide sequence that is at least 95% identical to SP-WT-AA (SEQ ID NO:43).
  • the third polypeptide sequence is SP-WT-AA (SEQ ID NO:43).
  • the recombinant GAA variant protein comprises the polypeptide sequence of FL-46-AA (SEQ ID NO:33).
  • the present disclosure provides a recombinant acid alphaglucosidase (GAA) variant protein, wherein the GAA variant protein comprises a set of amino acid substitutions, numbered relative to the full-length wild-type GAA protein sequence of FL-WT-AA (SEQ ID NO:2), selected from the group consisting of a) T151I, L650G, S676D, and L678H, b) L650S, S676D, and L678H, c) L650T, S676D, and L678H, d) L650E, S676D, and L678H, e) L650Y, S676D, and L678H, f) L650F, S676D, and L678H, g) L650G, S676D, and L678H, and h) S676D, and L678H.
  • GAA acid alphaglucosidase
  • the recombinant GAA variant protein comprises a first polypeptide sequence that is at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% identical to amino acid residues 70-952 of FL-6-AA (SEQ ID NO: 14).
  • the recombinant GAA variant protein comprises a first polypeptide sequence that is at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% identical to amino acid residues 70-952 of FL-7-AA (SEQ ID NO: 16).
  • the recombinant GAA variant protein comprises a first polypeptide sequence that is at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% identical to amino acid residues 70-952 of FL-8-AA (SEQ ID NO: 18).
  • the recombinant GAA variant protein comprises a first polypeptide sequence that is at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% identical to amino acid residues 70-952 of FL-9-AA (SEQ ID NO:20).
  • the recombinant GAA variant protein comprises a first polypeptide sequence that is at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% identical to amino acid residues 70-952 of FL-10-AA (SEQ ID NO:22).
  • the recombinant GAA variant protein comprises a first polypeptide sequence that is at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% identical to amino acid residues 70-952 of FL-11-AA (SEQ ID NO:24).
  • the recombinant GAA variant protein comprises a first polypeptide sequence that is at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% identical to amino acid residues 70-952 of FL-12-AA (SEQ ID NO:26).
  • the recombinant GAA variant protein comprises a first polypeptide sequence that is at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% identical to amino acid residues 70-952 of FL-13-AA (SEQ ID NO:28).
  • the recombinant GAA variant protein further comprises a second polypeptide sequence that is at least 95%, at least 97%, or 100% identical to PP-WT- AA (SEQ ID NO:39).
  • the recombinant GAA variant protein further comprises a third polypeptide sequence that is at least 95% or 100% identical to SP-WT-AA (SEQ ID NO:43).
  • the present disclosure provides a nucleic acid encoding the recombinant GAA variant protein.
  • the present disclosure provides an expression cassette comprising a nucleic acid encoding the recombinant GAA variant protein and at least one regulatory nucleic acid sequence operably linked to the sequence encoding the GAA protein.
  • the at least one regulatory nucleic acid sequence is selected from the group consisting of a promoter, an enhancer, an intron, a post-transcriptional regulatory element, an inverted terminal repeat (ITR), a polyadenylation (poly A) sequence, and a combination thereof.
  • the at least one regulatory nucleic acid sequence comprises a promoter.
  • the promoter is a muscle-specific promoter.
  • the muscle-specific promoter comprises a polynucleotide sequence that is at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SPc512_NA (SEQ ID NO:46).
  • the muscle-specific promoter comprises the polynucleotide sequence of SPc512_NA (SEQ ID NO:46).
  • the muscle-specific promoter comprises a polynucleotide sequence that is at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to sk-SH4_NA (SEQ ID NO:49).
  • the muscle-specific promoter comprises the polynucleotide sequence of sk-SH4_NA (SEQ ID NO:49).
  • the at least one regulatory nucleic acid sequence comprises an enhancer.
  • the enhancer is a muscle-specific enhancer.
  • the muscle-specific enhancer comprises a polynucleotide sequence that is at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to Dph-CRE04_NA (SEQ ID NO:48).
  • the muscle-specific enhancer comprises the polynucleotide sequence of Dph-CRE04_NA (SEQ ID NO:48).
  • the muscle-specific enhancer comprises a polynucleotide sequence that is at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to hDesmin NA (SEQ ID NO:47).
  • the muscle-specific enhancer comprises the polynucleotide sequence of hDesmin_NA (SEQ ID NO:47).
  • the at least one regulatory nucleic acid sequence comprises an intron.
  • the intron comprises a polynucleotide sequence that is at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to MVM_NA (SEQ ID NO:50).
  • the intron comprises the polynucleotide sequence of MVM NA (SEQ ID NO: 50).
  • the present disclosure provides a mammalian expression vector comprising an expression cassette comprising a nucleic acid encoding the recombinant GAA variant protein and at least one regulatory nucleic acid sequence operably linked to the sequence encoding the GAA protein.
  • the mammalian expression vector comprising an adeno- associated virus (AAV) vector.
  • AAV adeno- associated virus
  • the AAV vector comprises an AAV8 or AAV9 capsid polypeptide encapsidating the expression cassette.
  • the AAV vector comprises an engineered capsid polypeptide encapsidating the expression cassette.
  • the present disclosure provides a host cell comprising a nucleic acid encoding the recombinant GAA variant protein.
  • the present disclosure provides a host cell comprising a nucleic acid encoding the recombinant GAA variant protein and at least one regulatory nucleic acid sequence operably linked to the sequence encoding the GAA protein.
  • the host cell further comprises a nucleic acid encoding an AAV capsid polypeptide.
  • the AAV capsid polypeptide is an AAV8 or AAV9 capsid polypeptide.
  • the host cell further comprises a nucleic acid encoding a viral helper gene selected from the group consisting of E4, E2a, and VA.
  • the present disclosure provides a pharmaceutical composition
  • a mammalian expression vector comprising an expression cassette comprising a nucleic acid encoding the recombinant GAA variant protein and at least one regulatory nucleic acid sequence operably linked to the sequence encoding the GAA protein, and a pharmaceutically acceptable carrier.
  • the pharmaceutical composition is for the treatment of Pompe disease.
  • the present disclosure provides a method for treating Pompe disease in a subject in need thereof, the method comprising administering to the subject a therapeutically effective amount of the composition.
  • the composition is administered by epicardial injection, intravenous injection, intramuscular injection, intraperitoneal injection, intracardiac injection, intracardiac catheterization, direct intramyocardial injection, transvascular administration, antegrade intracoronary injection, retrograde injection, transendomyocardial injection, or molecular cardiac surgery with recirculating delivery (MCARD).
  • epicardial injection intravenous injection, intramuscular injection, intraperitoneal injection, intracardiac injection, intracardiac catheterization, direct intramyocardial injection, transvascular administration, antegrade intracoronary injection, retrograde injection, transendomyocardial injection, or molecular cardiac surgery with recirculating delivery (MCARD).
  • the present disclosure provides a method of treating Pompe disease in a subject in need thereof, the method comprising administering to the subject a therapeutically effective amount of a recombinant GAA variant protein, wherein the recombinant GAA variant protein comprises a first polypeptide sequence that is at least 95%, at least 96%, at least 97% or at least 98% identical to MP-46-AA (SEQ ID NO: 37) and wherein the recombinant GAA variant comprises one or more variant amino acids selected from the group consisting of L32W, L36S, L37T, P47Q, Q58V, A70L, P86E, D95E, LI 17D, S143Q, T151I, T158S, T274N, R275K, A445G, T494E, E530V, N535R, L577T, L650G, S676D, L678H, T700G, A719H, A758P, A820E, Q838K
  • the method comprises administering to the subject a therapeutically effective amount of a recombinant GAA variant protein, wherein the recombinant GAA variant protein comprises an amino acid substitution selected from the group consisting of L117D, T151I, L313V, L650G, L650S, L650T, L650E, L650Y, L650F, S676D, L678H, and/or L868F, numbered relative to the full-length wild-type GAA protein sequence of FL-WT-AA (SEQ ID NO:2).
  • the method comprises administering to the subject a therapeutically effective amount of a recombinant GAA variant protein, wherein the recombinant GAA variant protein comprises an amino acid substitution selected from the group consisting of L32W, L36S, L37T, P47Q, Q58V, A70L, P86E, D95E, S143Q, T158S, T274N, R275K, A445G, T494E, E530V, N535R, L577T, L678T, T700G, A719H, A758P, A820E, Q838K, L868F, L879E, R891H, Q902G, V921R, and S940A, numbered relative to the full-length wild-type GAA protein sequence of FL-WT-AA (SEQ ID NO:2).
  • the method comprises administering to the subject a therapeutically effective amount of a recombinant GAA variant protein, wherein the recombinant GAA variant protein comprises an amino acid substitution selected from the group consisting of T151I, L650G, S676D, and L678H, numbered relative to the full-length wild-type GAA protein sequence of FL-WT-AA (SEQ ID NO:2).
  • the method comprises administering to the subject a therapeutically effective amount of a recombinant GAA variant protein, wherein the recombinant GAA variant protein comprises a first polypeptide sequence that is at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to MP -46- AA (SEQ ID NO:37).
  • the method comprises administering to the subject a therapeutically effective amount of a recombinant GAA variant protein, wherein the recombinant GAA variant protein comprises a first polypeptide sequence that is MP-46-AA (SEQ ID NO: 37).
  • the method comprises administering to the subject a therapeutically effective amount of a recombinant GAA variant protein, wherein the recombinant GAA variant protein further comprises a second polypeptide sequence that is at least 95% or at least 97% identical to PP-WT-AA (SEQ ID NO:39).
  • the method comprises administering to the subject a therapeutically effective amount of a recombinant GAA variant protein, wherein the recombinant GAA variant protein comprises a second polypeptide sequence that is PP-WT- AA (SEQ ID NO:39).
  • the method comprises administering to the subject a therapeutically effective amount of a recombinant GAA variant protein, wherein the recombinant GAA variant protein further comprises a third polypeptide sequence that is at least 95% identical to SP-WT-AA (SEQ ID NO:43).
  • the method comprises administering to the subject a therapeutically effective amount of a recombinant GAA variant protein comprises a third polypeptide sequence that is SP-WT-AA (SEQ ID NO:43).
  • the method comprises administering to the subject a therapeutically effective amount of a recombinant GAA variant protein, wherein the recombinant GAA variant protein comprises the polypeptide sequence of FL-46-AA (SEQ ID NO:33).
  • the present disclosure provides a method of treating Pompe disease in a subject in need thereof, the method comprising administering to the subject a therapeutically effective amount of a recombinant GAA variant protein, wherein the recombinant GAA variant protein comprises a first polypeptide sequence that is at least 95%, at least 96%, at least 97% or at least 98% identical to MP-46-AA (SEQ ID NO: 37) and wherein the GAA variant comprises one or more variant amino acids selected from the group consisting of L32W, L36S, L37T, P47Q, Q58V, A70L, P86E, D95E, L117D, S143Q, T151I, T158S, T274N, R275K, A445G, T494E, E530V, N535R, L577T, L650G, S676D, L678H, T700G, A719H, A758P, A820E, Q838K, L868F
  • the method comprises administering to the subject a therapeutically effective amount of a recombinant GAA variant protein, wherein the first polypeptide sequence is at least 99% or at least 99.5% identical to MP-46-AA (SEQ ID NO:37).
  • the method comprises administering to the subject a therapeutically effective amount of a recombinant GAA variant protein, wherein the recombinant GAA variant protein comprises a tryptophan residue at position 32, a serine residue at position 36, a threonine residue at position 37, a glutamine residue at position 47, a valine residue at position 58, a leucine residue at position 70, a glutamic acid residue at position 86, a glutamic acid residue at position 95, an aspartic acid residue at position 117, a glutamine residue at position 143, an isoleucine residue at position 151, a serine residue at position 158, an asparagine residue at position 274, a lysine residue at position 275, a glycine residue at position 445, a glutamic acid residue at position 494, a valine residue at position 530, an arginine residue at position 535, a threonine residue at position 577
  • the method comprises administering a recombinant GAA variant protein wherein the recombinant GAA variant protein the first polypeptide sequence is MP-46-AA (SEQ ID NO:37).
  • the method comprises administering a recombinant GAA variant protein wherein the recombinant GAA variant protein further comprises a second polypeptide sequence that is at least 95% or at least 97% identical to PP-WT-AA (SEQ ID NO:39).
  • the method comprises administering a recombinant GAA variant protein, wherein the recombinant GAA variant protein, and wherein the second polypeptide sequence is PP-WT-AA (SEQ ID NO:39).
  • the method comprises administering a recombinant GAA variant protein, wherein the recombinant GAA variant protein further comprises a third polypeptide sequence that is at least 95% identical to SP-WT-AA (SEQ ID NO:43).
  • the method comprises administering to the subject a therapeutically effective amount of a recombinant GAA variant protein, wherein the third polypeptide sequence is SP-WT-AA (SEQ ID NO:43).
  • the method comprises administering a recombinant GAA variant protein comprises the polypeptide sequence of FL-46-AA (SEQ ID NO:33).
  • the present disclosure provides a method of treating Pompe disease in a subject in need thereof, comprising administering to the subject a therapeutically effective amount of a recombinant GAA variant protein, wherein the recombinant GAA protein comprises a set of amino acid substitutions, numbered relative to the full-length wildtype GAA protein sequence of FL-WT-AA (SEQ ID NO:2), selected from the group consisting of: a) T151I, L650G, S676D, and L678H, b) L650S, S676D, and L678H, c) L650T, S676D, and L678H, d) L650E, S676D, and L678H, e) L650Y, S676D, and L678H, f) L650F, S676D, and L678H, g) L650G, S676D, and L678H, and h) S676D, and L678H
  • the method comprises administering to the subject a therapeutically effective amount of a recombinant GAA variant protein, wherein the recombinant GAA variant protein comprises a first polypeptide sequence that is at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% identical to amino acid residues 70-952 of FL-6-AA (SEQ ID NO: 14).
  • the method comprises administering to the subject a therapeutically effective amount of a recombinant GAA variant protein, wherein the recombinant GAA variant protein comprises a first polypeptide sequence that is at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% identical to amino acid residues 70-952 of FL-7-AA (SEQ ID NO: 16).
  • the method comprises administering to the subject a therapeutically effective amount of a recombinant GAA variant protein, wherein the recombinant GAA variant protein comprises a first polypeptide sequence that is at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% identical to amino acid residues 70-952 of FL-8-AA (SEQ ID NO: 18).
  • the method comprises administering to the subject a therapeutically effective amount of a recombinant GAA variant protein, wherein the recombinant GAA variant protein comprises a first polypeptide sequence that is at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% identical to amino acid residues 70-952 of FL-9-AA (SEQ ID NO:20).
  • the method comprises administering to the subject a therapeutically effective amount of a recombinant GAA variant protein, wherein the recombinant GAA variant protein comprises a first polypeptide sequence that is at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% identical to amino acid residues 70-952 of FL-10-AA (SEQ ID NO:22).
  • the method comprises administering to the subject a therapeutically effective amount of a recombinant GAA variant protein, wherein the recombinant GAA variant protein comprises a first polypeptide sequence that is at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% identical to amino acid residues 70-952 of FL-11-AA (SEQ ID NO:24).
  • the method comprises administering to the subject a therapeutically effective amount of a recombinant GAA variant protein, wherein the recombinant GAA variant protein comprises a first polypeptide sequence that is at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% identical to amino acid residues 70-952 of FL-12-AA (SEQ ID NO:26).
  • the method comprises administering to the subject a therapeutically effective amount of a recombinant GAA variant protein, wherein the recombinant GAA variant protein comprises a first polypeptide sequence that is at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% identical to amino acid residues 70-952 of FL-13-AA (SEQ ID NO:28).
  • the method comprises administering to the subject a therapeutically effective amount of a recombinant GAA variant protein, wherein the recombinant GAA variant protein further comprises a second polypeptide sequence that is at least 95%, at least 97%, or 100% identical to PP-WT-AA (SEQ ID NO:39).
  • the method comprises administering to the subject a therapeutically effective amount of a recombinant GAA variant protein, wherein the recombinant GAA variant protein further comprises a third polypeptide sequence that is at least 95% or 100% identical to SP-WT-AA (SEQ ID NO:43).
  • the present disclosure provides a use of a recombinant GAA variant protein for a method of treating Pompe disease in a subject in need thereof, wherein the recombinant GAA variant protein comprises a first polypeptide sequence that is at least 95%, at least 96%, at least 97% or at least 98% identical to MP-46-AA (SEQ ID NO: 37) and wherein the recombinant GAA variant comprises one or more variant amino acids selected from the group consisting of L32W, L36S, L37T, P47Q, Q58V, A70L, P86E, D95E, LI 17D, S143Q, T151I, T158S, T274N, R275K, A445G, T494E, E530V, N535R, L577T, L650G, S676D, L678H, T700G, A719H, A758P, A820E, Q838K, L868F, L879E, R
  • the present disclosure provides a use of a recombinant GAA variant protein for a method of treating Pompe disease in a subject in need thereof, wherein the recombinant GAA variant protein comprises a first polypeptide sequence that is at least 95%, at least 96%, at least 97% or at least 98% identical to MP-46-AA (SEQ ID NO: 37) and wherein the GAA variant comprises one or more variant amino acids selected from the group consisting of L32W, L36S, L37T, P47Q, Q58V, A70L, P86E, D95E, L117D, S143Q, T151I, T158S, T274N, R275K, A445G, T494E, E530V, N535R, L577T, L650G, S676D, L678H, T700G, A719H, A758P, A820E, Q838K, L868F, L879E, R891H, Q
  • the present disclosure provides a use of a recombinant GAA variant protein for a method of treating Pompe disease in a subject in need thereof, wherein the recombinant GAA protein comprises a set of amino acid substitutions, numbered relative to the full-length wild-type GAA protein sequence of FL-WT-AA (SEQ ID NO:2), selected from the group consisting of: a) T151I, L650G, S676D, and L678H, b) L650S, S676D, and L678H, c) L650T, S676D, and L678H, d) L650E, S676D, and L678H, e) L650Y, S676D, and L678H, f) L650F, S676D, and L678H, g) L650G, S676D, and L678H, and h) S676D, and L678H.
  • SEQ ID NO:2 FL-WT-AA
  • the present disclosure provides a use of a recombinant GAA variant protein for the manufacture of a medicament for treatment of Pompe disease in a subject, wherein the recombinant GAA variant protein comprises a first polypeptide sequence that is at least 95%, at least 96%, at least 97% or at least 98% identical to MP-46-AA (SEQ ID NO:37) and wherein the recombinant GAA variant comprises one or more variant amino acids selected from the group consisting of L32W, L36S, L37T, P47Q, Q58V, A70L, P86E, D95E, L117D, S143Q, T151I, T158S, T274N, R275K, A445G, T494E, E530V, N535R, L577T, L650G, S676D, L678H, T700G, A719H, A758P, A820E, Q838K, L868F, L879E,
  • the present disclosure provides a use of a recombinant GAA variant protein for the manufacture of a medicament for treatment of Pompe disease in a subject, wherein the recombinant GAA variant protein comprises a first polypeptide sequence that is at least 95%, at least 96%, at least 97% or at least 98% identical to MP-46-AA (SEQ ID NO: 37) and wherein the GAA variant comprises one or more variant amino acids selected from the group consisting of L32W, L36S, L37T, P47Q, Q58V, A70L, P86E, D95E, LI 17D, S143Q, T151I, T158S, T274N, R275K, A445G, T494E, E530V, N535R, L577T, L650G, S676D, L678H, T700G, A719H, A758P, A820E, Q838K, L868F, L879E, R891H
  • the present disclosure provides a use of a recombinant GAA variant protein for the manufacture of a medicament for treatment of Pompe disease in a subject, wherein the recombinant GAA protein comprises a set of amino acid substitutions, numbered relative to the full-length wild-type GAA protein sequence of FL-WT-AA (SEQ ID NO:2), selected from the group consisting of: a) T151I, L650G, S676D, and L678H, b) L650S, S676D, and L678H, c) L650T, S676D, and L678H, d) L650E, S676D, and L678H, e) L650Y, S676D, and L678H, f) L650F, S676D, and L678H, g) L650G, S676D, and L678H, and h) S676D, and L678H.
  • SEQ ID NO:2 FL-WT-AA
  • FIG 1A, FIG. IB, FIG. 1C, FIG. ID, FIG. IE, FIG. 1G, and FIG. 1H collectively illustrate (A) a vector comprising a representative GAA transgene driven by a muscle specific promoter and enhancer packaged within an AAV9 capsid.
  • the following naming convention will be used to describe each vector:
  • AAV9.Dph-CRE04.SPc512.WT hGAA GAA protein expression measured using GAA activity in the heart, quadriceps, and triceps of GAA' /_ mice dosed with 3xl0 13 vg/kg AAV9 capsid vectors comprising constructs containing various enhancer elements (SK-SH4) and/or Desmin promoter and WT GAA using a 4-MUG assay.
  • gaa +!+ wild-type (WT) vehicle-treated mouse
  • gacr ⁇ gaa knock-out (KO) vehicle-treated mouse
  • E-G PAS staining of heart tissue from GAA KO mice after dosing with a 3xl0 13 vg/kg AAV9.SKSH4. Desmin.
  • WT hGAA compared to control gaa +/+ nd gaa' 1 ' mice; and (H) western blot analysis of heart tissue protein from GAA KO mice that received 3xl0 13 vg/kg AAV9 capsid vector comprising constructs containing WT hGAA, with or without an enhancer, or a buffer group showing trafficking and lysosomal processing of human GAA (lane 1 — WT hGAA; lanes 2 and 10 — molecular weight markers; lanes 3 and 4 — heart protein from GAA KO (gacC’) mice treated with vectors comprising AAV9. desmin.
  • FIG. 2A and FIG. 2B collectively illustrate (A) GAA activity (using 4-MUG assay) and (B) percent glycogen reduction in heart, quadriceps, triceps, gastrocnemius, diaphragm and aorta of GAA KO mice after dosing with 3xl0 13 vg/kg AAV9 capsid vectors comprising constructs with or without enhancer elements.
  • FIG. 3 illustrates the engineering approach for making GAA variants with certain desired properties and an optimal GAA protein.
  • FIG. 4A, FIG. 4B, FIG. 4C, FIG. 4D, FIG. 4E, FIG. 4F, and FIG. 4G collectively illustrate (A) the codon optimized GAA constructs tested, containing 5’ and 3’ inverted terminal repeats (ITRs) from AAV2, the Sk-SH4 enhancer, the human desmin promoter, the Minute Virus of Mice (MVM) intron, an SV40 polyadenylation (poly A) signal (SV40pA), and a DNA ID tag, for expressing WT human GAA (AAV9.Sk-SH4. desmin. WT hGAA), CO1 GAA (AAV9.Sk-SH4. desmin. CO 1 GAA), CO2GAA (AAV9.Sk-
  • WT hGAA AAV9.Sk-SH4.desmin.C01 GAA
  • CO2 AAV9.Sk-SH4.desmin.CO2 GAA
  • CO3 AAV9.Sk-SH4.desmin.CO3 GAA.
  • FIG. 5A and FIG. 5B collectively illustrate the kinetic activity of human GAA (WT hGAA) protein variants 1-4 (Var 1, Var 2, Var 3, and Var 4) as compared to WT hGAA using (A) 4-MUG and (B) glycogen as the substrate.
  • Purified variant GAA proteins 1-4 or WT hGAA were incubated with various concentrations of 4-MUG and the concentration of the fluorescent 4-MU product formed over time was quantified.
  • FIG. 6A, FIG. 6B, and FIG. 6C collectively illustrate a comparison of immunogenic profiles of WT hGAA compared to GAA variant 4 (Var 4) using a major histocompatibility (MHC)-associated peptide proteomics (MAPPs) assay.
  • MHC major histocompatibility
  • MAPPs major histocompatibility-associated peptide proteomics
  • FIG. 7A, FIG. 7B, FIG. 7C, FIG. 7D, FIG. 7E, FIG. 7F, FIG. 7G, FIG. 7H, and FIG. 71 collectively illustrate (A) a vector comprising the GAA Var 4 transgene driven by the Dph-CRE04 enhancer and the Spc512 muscle specific promoter packaged within an AAV9 capsid; (B) GAA activity using a 4-MUG assay in heart tissue lysates; (C) reduction in glycogen levels in heart tissue lysates of 12-week-old GAA KO mice dosed for 5 or 12 weeks with 2.4xl0 13 vg/kg AAV9 capsid vectors; (D) terminal heart weight; and (E) left ventricular (LV) mass of 12-week-old GAA KO mice dosed for 12 weeks with 2.4xl0 13 vg/kg AAV9 capsid constructs; and (F-I) echocardiography images of 12-week-old GAA KO mice dosed for 12 weeks with
  • FIG. 8A, FIG. 8B, FIG. 8C, FIG. 8D, and FIG. 8E collectively illustrate (A) GAA activity using 4-MUG as the substrate; (B) reduction in glycogen level in quadriceps muscle of 12-week-old GAA KO mice dosed for 5 or 12 weeks with 2.4xl0 13 vg/kg AAV9 capsid constructs; (C) images (i) GAA protein IHC staining; (ii) glycogen by PAS staining;
  • FIG. 9 illustrates the activity of GAA variants (Var 6, Var 7, Var 8, Var 9, Var 10, Var 11, Var 12, and Var 13) as compared to controls (WT hGAA, control plasmid, and no plasmid) in C2C12 gaa' 1 ' mouse muscle cells following transfection of the cells with plasmids and activity measured in cell extracts using the 4-MUG assay.
  • FIG. 10A, FIG. 10B, and FIG. 10C collectively illustrate (A) kinetic activity using 4-MUG as the substrate and (B-C) activity on glycogen of GAA variant 6 (Var 6) compared to WT hGAA.
  • FIG. HA, FIG. 11B, FIG. 11C, and FIG. HD collectively illustrate (A) kinetic activity using 4-MUG as the substrate, (B) stability, (C) thermostability, and (D) thermochemical stability of GAA variants (Var 4, Var 4+6), as compared to WT hGAA.
  • FIG. 12A and FIG. 12B collectively illustrate the uptake of GAA variants (Var 4, Var 4+6), by gaa 1 ' mouse muscle cells. Purified proteins were incubated with cells for 24 hours and the activity of GAA measured in cell lysates using the (A) 4-MUG assay and (B) the immunogenic profile of GAA variants (Var 4, Var 4+6), as compared to WT hGAA using the MAPPs assay. [00223] FIG. 13A, FIG. 13B, and FIG.
  • FIG. 13C collectively illustrate the GAA activity in (A) heart tissue lysates, (B) quadriceps tissue lysates, and (C) diaphragm tissue lysates of GAA KO mice that received either vehicle or a vector expressing WT hGAA or GAA Var 4+6 at either 3xl0 12 , IxlO 13 , or 3xl0 13 vg/kg vg/kg 12 weeks post dosing using 4-MUG as the substrate.
  • FIG. 14A, FIG. 14B, and FIG. 14C collectively illustrate reduction in glycogen levels in (A) heart muscle tissue lysates, (B) quadriceps muscle tissue lysates, and (C) diaphragm muscle tissue lysates of GAA KO mice that received either buffer or vector expressing WT hGAA or GAA Var 4+6 at either 3xl0 12 , IxlO 13 , or 3xl0 13 vg/kg 12 weeks post dosing.
  • FIG. 15A, FIG. 15B, and FIG. 15C collectively illustrate (A) the increase in GAA activity using 4-MUG as the substrate, (B) the reduction in glycogen in the brain, and (C) rescue of muscle function of GAA KO mice that received either buffer or vector expressing WT hGAA or GAA Var 4+6 at either 3xl0 12 , IxlO 13 , or 3xl0 13 vg/kg 12 weeks post dosing.
  • FIG. 16A, FIG. 16B, FIG. 16C, and FIG. 16D collectively illustrate (A) a vector comprising the GAA Var 4+6 transgene or WT hGAA transgene driven by the Dph-CRE04 enhancer and the Spc512 muscle-specific promoter packaged within an AAV9 capsid compared to the AT845-like construct comprising natural WT hGAA transgene (not codon optimized) driven by the MCK muscle specific promoter packaged within an AAV8 capsid; and GAA activity in (B) heart muscle tissue, (C) quadriceps muscle tissue, and (D) diaphragm muscle tissue of GAA KO mice that received either buffer or [AAV9.Dph- CRE04.SPc512.WT hGAA], [AAV9.Dph-CRE04.SPc512.GAA Var 4+6] or [AAV8.MCK Enh.MCK pro.WT hGAA] at either 3xl0 12 , Ix
  • FIG. 17A, FIG. 17B, and FIG. 17C collectively illustrate glycogen reduction in (A) heart muscle tissue, (B) quadriceps muscle tissue, and (C) diaphragm muscle tissue of GAA KO mice that received either buffer, [AAV9.Dph-CRE04.SPc512.WT hGAA], [AAV9.Dph-CRE04.SPc512.GAA Var 4+6] or [AAV8.MCK Enh.MCK pro.WT hGAA] at either 3xl0 12 , IxlO 13 , or 3xl0 13 vg/kg vg/kg 12 weeks post dosing. [00228] FIG.
  • FIG. 19A, FIG. 19B, FIG. 19C, FIG. 19D, FIG. 19E, FIG. 19F, FIG. 19G, FIG. 19H, FIG. 191, FIG. 19J, FIG. 19K, FIG. 19L, FIG. 19M, and FIG. 19N collectively illustrate SEQ ID NOs: 1-10 and 13-30, representing numerous human GAA polynucleotides and protein variants.
  • FIG. 20A, FIG. 20B, and FIG. 20C collectively illustrate SEQ ID NOs: 31, 60, and 62, representing codon-altered human GAA polynucleotides and protein variants.
  • FIG. 21A, FIG. 21B, and FIG. 21C collectively illustrate SEQ ID NOs: 32, 34, 35, and 37-45, representing polynucleotides and polypeptides associated with human GAA protein variants.
  • FIG. 22H collectively illustrate SEQ ID NOs: 46-54 and 85-102, representing polynucleotides and polypeptides of promoters, enhancers, and mammalian expression vectors.
  • FIG. 23 A, and FIG. 23B illustrate SEQ ID NO: 103; representing the polynucleotide associated with a human GAA protein variant (CO3-FL-46-dNA). (SEQ ID NO: 108).
  • FIG. 24 illustrate SEQ ID NO: 104; representing the polynucleotide associated with the vector expressing GAA Var 4+6.
  • FIG. 25 illustrates the list of IUPAC degenerate nucleotide codes.
  • FIG. 26A and FIG. 26B illustrate SEQ ID NO: 105; representing the polynucleotide associated with a human GAA protein variant (COl-FL-46-dNA). (SEQ ID NO: 109).
  • FIG. 27A and FIG. 27B illustrate SEQ ID NO: 106; representing the polynucleotide associated with a human GAA protein variant (CO2-FL-46-dNA). (SEQ ID NO: 110).
  • FIG. 28A and FIG. 28B illustrate a polynucleotide associated with a human GAA protein variant (CO3-MP-46-dNA). Figures discloses SEQ ID NOS 111-112, respectively, in order of appearance.
  • FIG. 29A and FIG. 29B illustrate a polynucleotide associated with a human GAA protein variant (COl-MP-46-dNA).
  • Figures discloses SEQ ID NOS 113-114, respectively, in order of appearance.
  • FIG. 30A and FIG. 30B illustrate a polynucleotide associated with a human GAA protein variant (CO2-MP-46-dNA).
  • FIG. 31 illustrates a dose dependent increase in the number of viral genomes that remained constant for a six-month period
  • FIG. 32 illustrates a dose dependent increase in GAA activity in the heart tissue following injection.
  • FIG. 33 illustrates a correction of glycogen accumulation observed in heart following injection.
  • FIG. 34 illustrates a dose dependent increase in GAA activity in the quadriceps tissue following injection.
  • FIG. 35 a correctio illustrates n of glycogen accumulation observed in quadriceps tissue following injection.
  • FIG. 36 illustrates a dose dependent increase in GAA activity in the diaphragm tissue following injection.
  • FIG. 37 illustrates a correction of glycogen accumulation observed in diaphragm tissue following injection.
  • ERT regimens including: (1) detrimental immune responses including neutralizing antibodies against recombinant GAA enzyme, especially in cross-reactive immune-material (CRIM) negative patients; (2) poor uptake of the GAA enzyme by muscle cells from circulation; (3) limited availability of the GAA enzyme in circulation (85% taken up by the liver); (4) reduced stability of the GAA enzyme at neutral pH; and (4) progressive endosomal dysfunction reducing efficacy of the endogenous enzyme delivery to lysosomes.
  • AAV therapy for Pompe disease while promising, have to this point encountered similar difficulties in the case of liver-directed therapies (e.g., low serum stability, poor uptake and inefficient deep tissue distribution).
  • AAV8 therapies with muscle-specific expression have been forced to rely upon dangerously high doses.
  • compositions herein described include novel muscle- directed expression vectors expressing optimized GAA variants.
  • these constructs perform better functionally (e.g., demonstrating greater uptake in muscle tissue, restoration of GAA activity and glycogen reduction).
  • these aforementioned effects were surprisingly accomplished at dramatically lower doses than reported for other muscle- directed gene therapy approaches.
  • GAA Var 4+6 administered at IxlO 13 vg/kg exhibited greater glycogen muscle reduction than AAV8.MCKenh.MCKpro.ET hGAA at a 3-fold higher dose (see FIG. 17A-FIG. 17C).
  • Alpha-glucosidase and “GAA” are used interchangeably and refer a protein with glucosidase activity for hydrolyzing terminal, nonreducing ( 4)-linked a-D-glucose residues in polysaccharides with release of D-glucose (e.g., active GAA, also referred to herein as “GAA mature polypeptide,” “GAA MP,” or simply “MP”) or a protein precursor thereof (e.g., a pro-protein or a pre-pro-protein, often referred to as pGAA and ppGAA), e.g., as measured by quantification of glucose release from glycogen following incubation with the GAA polypeptide.
  • active GAA also referred to herein as “GAA mature polypeptide,” “GAA MP,” or simply “MP”
  • MP a protein precursor thereof
  • pGAA and ppGAA e.g., as measured by quantification of glucose release from glycogen following incubation with the GAA polypeptide.
  • GAA is translated as an inactive, single-chain polypeptide that includes a signal peptide and a propeptide, often referred to as a GAA pre-pro-protein.
  • the GAA pre-pro- protein undergoes post-translational processing to form an active GAA protein. This processing includes removal (e.g., by cleavage) of the signal peptide, followed by removal (e.g., by cleavage) of the propeptide, to form a mature GAA polypeptide.
  • polynucleotides encoding the wild-type human GAA encode for an inactive single-chain polypeptide (e.g., a pre-pro-protein; amino acids 1-952 of GAA-FL- WT-AA (SEQ ID NO:2)) that undergoes post-translational processing to form an active GAA protein.
  • a pre-pro-protein e.g., amino acids 1-952 of GAA-FL- WT-AA (SEQ ID NO:2)
  • the GAA pre-pro-protein is first cleaved with a signal peptidase to release the encoded signal peptide (amino acids 1-27 of GAA-FL-WT-AA (SEQ ID NO:2)), forming a GAA precursor (amino acids 28-952 of SEQ ID NO:2; 110 kDa).
  • the GAA precursor is cleaved by additional proteases to release a first associated polypeptide of 19.4 kD (amino acids 792-952 of SEQ ID NO:2), a second associated polypeptide of 3.9 kD (amino acids 78-113 of SEQ ID NO:2), and a third associated polypeptide of 10.4 kDa (amino acids 122-200 of SEQ ID NO:2), forming a mature GAA (amino acids 203-782 of SEQ ID NO:2; 70 kDa).
  • the “MP” designation refers to a precursor polypeptide that includes the first associated polypeptide, the second associated polypeptide, the third associated polypeptide, and the mature GAA polypeptide.
  • GAA polypeptide As used herein, the terms “GAA polypeptide”, “GAA protein”, and “GAA variant protein” refer to a polypeptide having GAA glucosidase activity under particular conditions, e.g., as measured by quantification of glucose release from glycogen or using the artificial substrate 4-MUG following incubation with the GAA polypeptide.
  • GAA polypeptides include precursor polypeptides (e.g., GAA pre-pro-polypeptides and pro-polypeptides) which, when activated by the post-translational processing described above, become active GAA polypeptides with GAA glucosidase activity, as well as the active GAA polypeptides (e.g., GAA-MP) themselves.
  • a human GAA polypeptide refers to a polypeptide that includes an amino acid sequence with high sequence identity (e.g., at least 85%, 90%, 95%, 99%, or more) to the portion of the wild-type human GAA polypeptide that includes the mature GAA polypeptide, GAA-MP- AA (SEQ ID NO: 35) or to the portions of the disclosed variant GAA polypeptides (variants 1-4, 6-13 and variant 4+6), shown in FIG. 19B-FIG. 19N.
  • an amino acid sequence with high sequence identity e.g., at least 85%, 90%, 95%, 99%, or more
  • GAA polypeptides with one or more of the amino acid substitutions L32W, L36S, L37T, P47Q, Q58V, A70L, P86E, D95E, L117D, S143Q, T151I, T158S, T274N, R275K, A445G, T494E, E530V, N535R, L577T, L650G, L650S, L650T, L650E, L650Y, L650F, S676D, L678H, L678T, T700G, A719H, A758P, A820E, Q838K, L868F, L879E, R891H, Q902G, V921R, and S940A, relative to the wild-type human GAA polypeptide, present in variant 4+6 described herein.
  • Non-limiting examples of wild-type GAA polypeptides include human GAA polypeptides (e.g., GenBank accession nos. NP 000143.2 (GAA-FL-WT-AA (SEQ ID NO:2)) and UniProt accession no. P10253), and natural variants thereof; bovine GAA (e.g., UniProt accession no. Q9MYM4); murine GAA (e.g., UniProt accession no. P70699); rat GAA (e.g., UniProt accession no. Q6P7A9), and natural variants thereof; and other mammalian GAA homologues (e.g., chimpanzee, ape, hamster, guinea pig, etc.).
  • human GAA polypeptides e.g., GenBank accession nos. NP 000143.2 (GAA-FL-WT-AA (SEQ ID NO:2)
  • UniProt accession no. P10253 UniProt accession no
  • GAA polypeptide includes natural variants and artificial constructs.
  • GAA encompasses any natural variants, alternative sequences, isoforms, or mutant proteins that retain some basal GAA glucosidase activity (e.g., at least 5%, 10%, 25%, 50%, 75%, or more of the corresponding wild-type activity as assayed), including one or more variant amino acids found in the human population, such as S46P, C103G, C103R, C127F, R190H, Y191C, L208P, P217L, G219R, R224P, R224Q, R224W, T234K, T234R, A237, S251L, S254L, E262K, P266S, P285R, P285S, L291F, L291P, Y292C, G293R, L299R, H308L, H308P, G309
  • NCBI National Center for Biotechnology Information
  • GAA protein (z.e., as translated with a signal peptide and propeptide) can include one or more variants, with the Variant 4+6 finding particular use in some embodiments.
  • This is referred to as “GAA-FL-Var46-AA” (SEQ ID NO:30) with the nucleic acid sequence being referred to herein as “GAA-FL-Var46-NA.”
  • codon-optimized sequences CO1-FL-WT-AA, CO2-FL-WT-AA, and CO3-FL-WT-AA also encode the full-length GAA protein.
  • specifically included in the definition of GAA is all such variants exemplified herein.
  • GAA amino acids refers to the corresponding amino acid in the full-length, wild-type human GAA pre-pro-polypeptide sequence (GAA-FL-WT-AA), presented as SEQ ID NO:2 in FIG. 19A.
  • the recited amino acid number refers to the analogous (e.g., structurally or functionally equivalent) and/or homologous (e.g., evolutionarily conserved in the primary amino acid sequence) amino acid in the full-length, wild-type GAA pre-pro-polypeptide sequence.
  • an LI 17D amino acid substitution refers to an L to D substitution at position 117 of the full- length, wild-type human GAA pre-pro-peptide sequence (GAA-FL-WT-AA (SEQ ID NO:2)), as well as an L to D substitution at position 48 of the mature, wild-type GAA singlechain polypeptide (GAA-MP-WT-AA (SEQ ID NO:35)). Both of these nomenclatures describe the same L to D amino acid substitution, in different GAA polypeptides.
  • GAA polynucleotide refers to a polynucleotide encoding a GAA polypeptide having GAA glucosidase activity under particular conditions, e.g., as measured by quantification of glucose release from glycogen following incubation with the GAA polypeptide.
  • GAA polynucleotides include polynucleotides encoding GAA precursor polypeptides, including GAA pre-pro-polypeptides, GAA pro-polypeptides, and mature, single-chain GAA polypeptides.
  • GAA polynucleotides encoding a GAA polypeptide that includes one or more of the amino acid substitutions L32W, L36S, L37T, P47Q, Q58V, A70L, P86E, D95E, L117D, S143Q, T151I, T158S, T274N, R275K, A445G, T494E, E530V, N535R, L577T, L650G, L650S, L650T, L650E, L650Y, L650F, S676D, L678H, L678T, T700G, A719H, A758P, A820E, Q838K, L868F, L879E, R891H, Q902G, V921R, and S940A, relative to the wild-type human GAA polypeptide, present in variant 4+6 described herein.
  • a human GAA polynucleotide refers to a polynucleotide that encodes a polypeptide that includes an amino acid sequence with high sequence identity (e.g., at least 85%, 90%, 95%, 99%, or more) to the portion of the wild-type human GAA polypeptide that includes the mature GAA polypeptide, GAA-MP-AA (SEQ ID NO: 35) or to the portions of the disclosed variant GAA polypeptides (variants 1-4, 6-13 and variant 4+6), shown in FIG. 19B-FIG. 19N
  • GAA polynucleotides can include regulatory elements, such as promoters, enhancers, terminators, polyadenylation sequences, and introns, as well viral packaging elements, such as inverted terminal repeats (“ITRs”), and/or other elements that support replication of the polynucleotide in a non-viral host cell, e.g., a replicon supporting propagation of the polynucleotide, e.g., in a bacterial, yeast, or mammalian host cell.
  • regulatory elements such as promoters, enhancers, terminators, polyadenylation sequences, and introns
  • viral packaging elements such as inverted terminal repeats (“ITRs”), and/or other elements that support replication of the polynucleotide in a non-viral host cell, e.g., a replicon supporting propagation of the polynucleotide, e.g., in a bacterial, yeast, or mammalian host cell.
  • codon-altered GAA polynucleotides Of particular use in the present disclosure are codon-altered GAA polynucleotides. As described herein, the codon-altered GAA polynucleotides provide increased expression of transgenic GAA in vivo, as compared to the level of GAA expression provided by a natively- coded GAA construct (e.g., a polynucleotide encoding the same GAA amino acid sequence using the wild-type human codons).
  • a natively- coded GAA construct e.g., a polynucleotide encoding the same GAA amino acid sequence using the wild-type human codons.
  • the term “increased expression” refers to an increased level of transgenic GAA protein in a tissue (e.g., a muscular tissue) of an animal administered the codon-altered polynucleotide encoding GAA, as compared to the level of transgenic GAA protein in the same tissue of an animal administered a natively-coded GAA construct. Increased expression of the protein leads to an increase in GAA activity; thus, increased expression leads to increased activity.
  • increased expression refers to at least 25% greater transgenic GAA polypeptide in a tissue of an animal administered the codon-altered GAA polynucleotide, as compared to the level of transgenic GAA polypeptide in the same tissue of an animal administered a natively-coded GAA polynucleotide.
  • increased expression refers to an effect generated by the alteration of the codon sequence, rather than hyperactivity caused by an underlying amino acid substitution. That is, the expression level obtained from a codon-optimized sequence encoding a GAA variant described herein is compared relative to the expression level obtained from a natively-coded GAA variant protein.
  • increased expression refers to at least 50% greater, at least 75% greater, at least 100% greater, at least 3 -fold greater, at least 4-fold greater, at least 5-fold greater, at least 6-fold greater, at least 7-fold greater, at least 8-fold greater, at least 9-fold greater, at least 10-fold greater, at least 15-fold greater, at least 20-fold greater, at least 25-fold greater, at least 30-fold greater, at least 40-fold greater, at least 50- fold greater, at least 60-fold greater, at least 70-fold greater, at least 80-fold greater, at least 90-fold greater, at least 100-fold greater, at least 125-fold greater, at least 150-fold greater, at least 175-fold greater, at least 200-fold greater, at least 225-fold greater, or at least 250-fold greater transgenic GAA polypeptide in a tissue of an animal administered the codon-altered GAA polynucleotide, as compared to the level of transgenic GAA polypeptide in the same tissue of an animal administered a
  • GAA activity or “GAA glucosidase activity” herein is meant the ability to hydrolyzing terminal, non-reducing ( 4)-linked a-D-glucose residues in polysaccharides with release of D-glucose.
  • the activity levels can be measured using any GAA activity known in the art.
  • An exemplary assay for determining GAA activity by quantification of glucose release from glycogen following incubation with the GAA polypeptide.
  • the therapeutic potential of a GAA polynucleotide composition is evaluated by the increase in GAA activity in a tissue of an animal administered a GAA polynucleotide, e.g., instead of or in addition to increased GAA expression in the tissue.
  • increased GAA activity refers to a greater increase in GAA activity in a tissue of an animal administered a codon-altered GAA polynucleotide, relative to a baseline GAA activity in the tissue of the animal prior to administration of the codon-altered GAA polynucleotide, as compared to the increase in GAA activity in the same tissue of an animal administered a natively-coded GAA polynucleotide, relative to a baseline GAA activity in the tissue of the animal prior to administration of the natively-coded GAA polynucleotide.
  • increased GAA activity refers to at least a 25% greater increase in GAA activity in a tissue of an animal administered the codon-altered GAA polynucleotide, relative to a baseline level of GAA activity in the tissue of the animal prior to administration of the codon-altered GAA polynucleotide, as compared to the increase in the level GAA activity in the blood of an animal administered a natively-coded GAA polynucleotide, relative to the baseline level of GAA activity in the animal prior to administration of the natively-coded GAA polynucleotide.
  • increased GAA activity refers to at least 50% greater, at least 75% greater, at least 100% greater, at least 3 -fold greater, at least 4-fold greater, at least 5-fold greater, at least 6-fold greater, at least 7-fold greater, at least 8-fold greater, at least 9-fold greater, at least 10-fold greater, at least 15-fold greater, at least 20-fold greater, at least 25-fold greater, at least 30-fold greater, at least 40-fold greater, at least 50-fold greater, at least 60-fold greater, at least 70-fold greater, at least 80-fold greater, at least 90-fold greater, at least 100-fold greater, at least 125-fold greater, at least 150-fold greater, at least 175-fold greater, at least 200-fold greater, at least 225-fold greater, or at least 250-fold greater increase in GAA activity in a tissue of an animal administered the codon-altered GAA polynucleotide, relative to a baseline level of GAA activity in the tissue of the animal prior to administration of the codon-
  • the GAA amino acid numbering system is dependent on whether the GAA pre-pro-peptide (e.g., amino acids 1-69 of the full-length, wild-type human GAA sequence, inclusive of the signal peptide and pro-peptide) is included. Where the pre- pro-peptide is included, the numbering is referred to as “pre-pro-peptide inclusive” or “PPI”. Where the pre-pro-peptide is not included, the numbering is referred to as “pre-pro-peptide exclusive” or “PPE .” For example, LI 17D is PPI numbering for the same amino acid substitution as L48D, in PPE numbering. Similarly, the GAA amino acid numbering is also dependent upon the size of the of the signal peptide and/or propeptide in the particular GAA polypeptide.
  • the GAA amino acid numbering is also dependent upon the size of the of the signal peptide and/or propeptide in the particular GAA polypeptide.
  • GAA gene therapy includes any therapeutic approach of providing an exogenous nucleic acid encoding GAA to a patient to relieve, diminish, or prevent the reoccurrence of one or more symptoms (e.g., clinical factors) associated with a GAA deficiency (e.g., Pompe disease).
  • the term encompasses administering any compound, drug, procedure, or regimen comprising a nucleic acid encoding a GAA molecule, including any modified form of GAA (e.g., a GAA 4+6 variant), for maintaining or improving the health of an individual with a GAA deficiency (e.g., Pompe disease).
  • a therapeutically effective amount or dose or “therapeutically sufficient amount or dose” or “effective or sufficient amount or dose” refer to a dose that produces therapeutic effects for which it is administered.
  • a therapeutically effective amount of a drug useful for treating Pompe disease can be the amount that is capable of preventing or relieving one or more symptoms associated with Pompe disease.
  • a therapeutically effective treatment results in a decrease in the severity of musculoskeletal ailments (e.g., limb-girdle muscle weakness (LGMW)) in a subject.
  • LGMW limb-girdle muscle weakness
  • the term “gene” refers to the segment of a DNA molecule that codes for a polypeptide chain (e.g., the coding region).
  • a gene is positioned by regions immediately preceding, following, and/or intervening the coding region that are involved in producing the polypeptide chain (e.g., regulatory elements such as a promoter, enhancer, polyadenylation sequence, 5' -untranslated region, 3 ' -untranslated region, or intron).
  • regulatory elements refers to nucleotide sequences, such as promoters, enhancers, terminators, polyadenylation sequences, introns, etc., that provide for the expression of a coding sequence in a cell.
  • promoter element refers to a nucleotide sequence that assists with controlling expression of a coding sequence. Generally, promoter elements are located 5' of the translation start site of a gene. However, in certain embodiments, a promoter element may be located within an intron sequence, or 3' of the coding sequence.
  • a promoter useful for a gene therapy vector is derived from the native gene of the target protein (e.g., a GAA promoter). In some embodiments, a promoter useful for a gene therapy vector is specific for expression in a particular cell or tissue of the target organism (e.g, a muscle-specific promoter).
  • one of a plurality of well characterized promoter elements is used in a gene therapy vector described herein.
  • well-characterized promoter elements include the CMV early promoter, the (3 -actin promoter, and the methyl CpG binding protein 2 (MeCP2) promoter.
  • the promoter is a constitutive promoter, which drives substantially constant expression of the target protein.
  • the promoter is an inducible promoter, which drives expression of the target protein in response to a particular stimulus (e.g, exposure to a particular treatment or agent).
  • an “MVM intron” refers to an intron sequence derived from minute virus of mice having high sequence identity to SEQ ID NO:50.
  • MVM intron For further information on the MVM intron itself, see Haut and Pintel, J Virol. 72(3): 1834-43 (1998), and use of the MVM intron in AAV gene therapy vectors, see Wu Z et al., Mol Then, 16(2):280-9 (2008), both of which are hereby incorporated by reference.
  • operably linked refers to the relationship between a first reference nucleotide sequence (e.g., a gene) and a second nucleotide sequence (e.g., a regulatory control element) that allows the second nucleotide sequence to affect one or more properties associated with the first reference nucleotide sequence (e.g., a transcription rate).
  • a regulatory control element is operably linked to a GAA transgene when the regulatory element is positioned within a gene therapy vector such that it exerts an effect (e.g., a promotive or tissue selective affect) on transcription of the GAA transgene.
  • a vector refers to any nucleic acid construct used to transfer a GAA nucleic acid into a host cell.
  • a vector includes a replicon, which functions to replicate the nucleic acid construct.
  • vectors useful for gene therapy include plasmids, phages, cosmids, artificial chromosomes, and viruses, which function as autonomous units of replication in vivo.
  • a vector is a viral vector for introducing a GAA nucleic acid into the host cell.
  • Many modified eukaryotic viruses useful for gene therapy are known in the art. For example, adeno-associated viruses (AAVs) are particularly well suited for use in human gene therapy because humans are a natural host for the virus, the native viruses are not known to contribute to any diseases, and the viruses illicit a mild immune response.
  • AAVs adeno-associated viruses
  • GAA viral vector refers to a recombinant virus comprising a GAA polynucleotide, encoding a GAA polypeptide, which is sufficient for expression of the GAA polypeptide when introduced into a suitable animal host (e.g., a human).
  • a suitable animal host e.g., a human
  • recombinant viruses in which a codon-altered GAA polynucleotide, which encodes a GAA polypeptide, has been inserted into the genome of the virus.
  • GAA viral vectors are recombinant viruses in which the native genome of the virus has been replaced with a GAA polynucleotide, which encodes a GAA polypeptide. Included within the definition of GAA viral vectors are recombinant viruses comprising a GAA polynucleotide which encodes a GAA polypeptide with one or more of the amino acid substitutions L32W, L36S, L37T, P47Q, Q58V, A70L, P86E, D95E, L117D, S143Q, T151I, T158S, T274N, R275K, A445G, T494E, E530V, N535R, L577T, L650G, L650S, L650T, L650E, L650Y, L650F, S676D, L678H, L678T, T700G, A719H, A758P, A820E, Q838K, L868F, L
  • GAA viral particle refers to a viral particle encapsidating a GAA polynucleotide, encoding a GAA polypeptide, which is specific for expression of the GAA polypeptide when introduced into a suitable animal host (e.g., a human).
  • suitable animal host e.g., a human
  • recombinant viral particles encapsidating a genome in which a codon-altered GAA polynucleotide, which encodes a GAA polypeptide, has been inserted.
  • GAA viral particles are recombinant viral particles encapsidating a GAA polynucleotide, which encodes a GAA polypeptide, which replaces the native genome of the virus. Included within the definition of GAA viral particles are recombinant viral particles encapsidating a GAA polynucleotide which encodes a GAA polypeptide with one or more of the amino acid substitutions L32W, L36S, L37T, P47Q, Q58V, A70L, P86E, D95E, L117D, S143Q, T151I, T158S, T274N, R275K, A445G, T494E, E530V, N535R, L577T, L650G, L650S, L650T, L650E, L650Y, L650F, S676D, L678H, L678T, T700G, A719H, A758P, A820E, Q
  • AAV adeno-associated virus
  • ssDNA linear single-stranded DNA
  • kb kilobases
  • AAVs are not currently known to cause disease and cause only a very mild immune response.
  • Gene therapy vectors using AAVs can “infect” or transduce both dividing and quiescent cells and persist in an extrachromosomal state without integrating into the genome of the host cell or integrating the genome at a low frequency.
  • the wt genome comprises inverted terminal repeats (ITRs) at both ends of the DNA strand, and two open reading frames (ORFs): rep and cap.
  • the former is composed of four overlapping genes encoding Rep proteins required for the AAV life cycle, and the latter contains overlapping nucleotide sequences of capsid proteins: VP1, VP2 and VP3, which interact to form a capsid with icosahedral symmetry.
  • ITRs seem to be the only sequences required in cis next to the therapeutic gene: structural (cap) and packaging (rep) proteins can be delivered in trans. With this assumption many methods were established for efficient production of recombinant AAV (AAV), or engineered AAV, vectors containing a heterologous sequence, e.g., a reporter or nucleic acid encoding a therapeutic gene product.
  • AAV type 1 AAV1
  • AAV type 2 AAV2
  • AAV type 3 AAV3
  • AAV type 4 AAV4
  • AAV type 5 AAV5
  • AAV type 6 AAV6
  • AAV type 7 AAV7
  • AAV8 AAV8
  • AAV type 9 AAV9 viruses
  • viruses e.g., encapsidating a GAA polynucleotide
  • viruses formed by one or more variant AAV capsid proteins e.g., encapsidating a GAA polynucleotide.
  • viruses formed using a non-naturally occurring, engineered capsid protein e.g., recombinant viruses and viral particles formed using a non-naturally occurring, engineered capsid protein.
  • capsid protein refers to an expression product of a cap nucleic acid from an AAV serotype that forms a protein shell for an AAV virus, such as a wt capsid protein from serotypes 1, 6, 8, or 9; or a protein that shares at least 50% (alternatively at least 75, 80, 85, 90, 95, 96, 97, 98, 99%, or 99.5%) amino acid sequence identity with a wt capsid protein and displays a functional activity of a wt capsid protein.
  • a “functional activity” of a protein is any activity associated with the physiological function of the protein, whether in vitro, ex vivo, or in vivo.
  • functional activities of an AAV capsid protein may include its ability to form a capsid, evade host antibodies, recognize, and enter a cell, deliver DNA genome to the nucleus and transcription of its DNA genome.
  • the capsid protein is a variant of the wt capsid protein with an altered functional activity such as tissue transduction or tissue tropism, e.g., into or to muscle, respectively.
  • tissue transduction or tissue tropism e.g., into or to muscle, respectively.
  • encapsidates or “packaged” means encloses or surrounds a gene or virus in a protein shell or capsid.
  • the capsid polypeptides described herein refer to the VP1 form of the capsid polypeptide. However, it will be appreciated that these VP1 sequences also define the sequences of the VP2 form and the VP3 form.
  • CpG refers to a cytosine-guanine dinucleotide along a single strand of DNA, with the “p” representing the phosphate linkage between the two.
  • CpG island refers to a region within a polynucleotide having a statistically elevated density of CpG dinucleotides.
  • a region of a polynucleotide is a CpG island if, over a 200-base pair window: (i) the region has GC content of greater than 50%, and (ii) the ratio of observed CpG dinucleotides per expected CpG dinucleotides is at least 0.6, as defined by the relationship:
  • nucleic acid refers to deoxyribonucleotides or ribonucleotides and polymers thereof in either single- or double-stranded form and complements thereof.
  • the term encompasses nucleic acids containing known nucleotide analogs or modified backbone residues or linkages, which are synthetic, naturally occurring, and non-naturally occurring, which have similar binding properties as the reference nucleic acid, and which are metabolized in a manner similar to the reference nucleotides.
  • Examples of such analogs include, without limitation, phosphorothioates, phosphoramidates, methyl phosphonates, chiral-methyl phosphonates, 2-O-methyl ribonucleotides, and peptide-nucleic acids (PNAs).
  • PNAs peptide-nucleic acids
  • nucleic acid compositions herein is meant any molecule or formulation of a molecule that includes a GAA polynucleotide, encoding a GAA polynucleotide. Included within the definition of nucleic acid compositions are GAA polynucleotides, aqueous solutions of GAA polynucleotides, viral particles encapsidating a GAA polynucleotide, and aqueous formulations of viral particles encapsidating a GAA polynucleotide.
  • a nucleic acid composition, as disclosed herein, includes a codon-altered GAA gene, that encodes a GAA polypeptide.
  • amino acid refers to naturally occurring and non-natural amino acids, including amino acid analogs and amino acid mimetics that function in a manner similar to the naturally occurring amino acids.
  • Naturally occurring amino acids include those encoded by the genetic code, as well as those amino acids that are later modified, e.g., hydroxyproline, y-carboxyglutamate, and O-phosphoserine.
  • Naturally occurring amino acids can include, e.g., D- and L-amino acids.
  • amino acid sequences one of ordinary skill in the art will recognize that individual substitutions, deletions or additions to a nucleic acid or peptide sequence that alters, adds or deletes a single amino acid or a small percentage of amino acids in the encoded sequence is a “conservatively modified variant” where the alteration results in the substitution of an amino acid with a chemically similar amino acid. Conservative substitution tables providing functionally similar amino acids are well known in the art. Such conservatively modified variants are in addition to and do not exclude polymorphic variants, interspecies homologs, and alleles of the disclosure.
  • nucleic acids or peptide sequences refer to two or more sequences or subsequences that are the same or have a specified percentage of amino acid residues or nucleotides that are the same (i.e., about 60% identity, preferably 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or higher identity over a specified region, when compared and aligned for maximum correspondence over a comparison window or designated region) as measured using a BLAST or BLAST 2.0 sequence comparison algorithms with default parameters described below, or by manual alignment and visual inspection.
  • sequence identity and/or similarity is determined using standard techniques known in the art, including, but not limited to, the local sequence identity algorithm of Smith & Waterman, Adv. Appl. Math., 2:482 (1981), by the sequence identity alignment algorithm of Needleman & Wunsch, J. Mol. Biol., 48:443 (1970), by the search for similarity method of Pearson & Lipman, Proc. Natl. Acad. Sci.
  • PILEUP creates a multiple sequence alignment from a group of related sequences using progressive, pair wise alignments. It may also plot a tree showing the clustering relationships used to create the alignment. PILEUP uses a simplification of the progressive alignment method of Feng & Doolittle, J. Mol. Evol. 35:351-360 (1987); the method is similar to that described by Higgins & Sharp CABIOS 5: 151-153 (1989), both incorporated by reference.
  • Useful PILEUP parameters including a default gap weight of 3.00, a default gap length weight of 0.10, and weighted end gaps.
  • Another example of a useful algorithm is the BLAST algorithm, described in: Altschul et al., J. Mol. Biol. 215, 403-410, (1990); Altschul et al., Nucleic Acids Res.
  • a particularly useful BLAST program is the WU-BLAST-2 program which was obtained from Altschul et al., Methods in Enzymology, 266:460-480 (1996); http://blast.wusd/edu/blast/ README.html], WU-BLAST-2 uses several search parameters, most of which are set to the default values.
  • the HSP S and HSP S2 parameters are dynamic values and are established by the program itself depending upon the composition of the particular sequence and composition of the particular database against which the sequence of interest is being searched; however, the values may be adjusted to increase sensitivity.
  • Gapped BLAST uses BLOSUM- 62 substitution scores; threshold T parameter set to 9; the two-hit method to trigger ungapped extensions; charges gap lengths of k a cost of 10+k; Xu set to 16, and Xg set to 40 for database search stage and to 67 for the output stage of the algorithms. Gapped alignments are triggered by a score corresponding to ⁇ 22 bits.
  • a % amino acid sequence identity value is determined by the number of matching identical residues divided by the total number of residues of the “longer” sequence in the aligned region.
  • the “longer” sequence is the one having the most actual residues in the aligned region (gaps introduced by WU-Blast-2 to maximize the alignment score are ignored).
  • “percent (%) nucleic acid sequence identity” with respect to the coding sequence of the polypeptides identified is defined as the percentage of nucleotide residues in a candidate sequence that are identical with the nucleotide residues in the coding sequence of the cell cycle protein.
  • a preferred method utilizes the BLASTN module of WU- BL AST-2 set to the default parameters, with overlap span and overlap fraction set to 1 and 0.125, respectively.
  • the alignment may include the introduction of gaps in the sequences to be aligned.
  • sequences which contain either more or fewer amino acids than the protein encoded by the wild-type GAA sequence of (SEQ ID NO:2) it is understood that in one embodiment, the percentage of sequence identity will be determined based on the number of identical amino acids or nucleotides in relation to the total number of amino acids or nucleotides.
  • sequence identity of sequences shorter than SEQ ID NO:2 will be determined using the number of nucleotides in the shorter sequence, in one embodiment. In percent identity calculations relative weight is not assigned to various manifestations of sequence variation, such as, insertions, deletions, substitutions, etc.
  • Percent sequence identity may be calculated, for example, by dividing the number of matching identical residues by the total number of residues of the “shorter” sequence in the aligned region and multiplying by 100. The “longer” sequence is the one having the most actual residues in the aligned region.
  • allelic variants refers to polymorphic forms of a gene at a particular genetic locus, as well as cDNAs derived from mRNA transcripts of the genes, and the polypeptides encoded by them.
  • the term “preferred mammalian codon” refers a subset of codons from among the set of codons encoding an amino acid that are most frequently used in proteins expressed in mammalian cells as chosen from the following list: Gly (GGC, GGG); Glu (GAG); Asp (GAC); Vai (GTG, GTC); Ala (GCC, GCT); Ser (AGC, TCC); Lys (AAG); Asn (AAC); Met (ATG); He (ATC); Thr (ACC); Trp (TGG); Cys (TGC); Tyr (TAT, TAC); Leu (CTG); Phe (TTC); Arg (CGC, AGG, AGA); Gin (CAG); His (CAC); and Pro (CCC).
  • the term “codon-altered” or “codon-optimized” refers to a polynucleotide sequence encoding a polypeptide (e.g., a GAA protein), where at least one codon of the native polynucleotide encoding the polypeptide has been changed to improve a property of the polynucleotide sequence.
  • the improved property promotes increased transcription of mRNA coding for the polypeptide, increased stability of the mRNA (e.g., improved mRNA half-life), increased translation of the polypeptide, and/or increased packaging of the polynucleotide within the vector.
  • Non-limiting examples of alterations that can be used to achieve the improved properties include changing the usage and/or distribution of codons for particular amino acids, adjusting global and/or local GC content, removing AT -rich sequences, removing repeated sequence elements, adjusting global and/or local CpG dinucleotide content, removing cryptic regulatory elements (e.g., TATA box and CCAAT box elements), removing of intron/exon splice sites, improving regulatory sequences (e.g., introduction of a Kozak consensus sequence), and removing sequence elements capable of forming secondary structure (e.g., stem-loops) in the transcribed mRNA.
  • cryptic regulatory elements e.g., TATA box and CCAAT box elements
  • intron/exon splice sites e.g., introduction of a Kozak consensus sequence
  • improving regulatory sequences e.g., introduction of a Kozak consensus sequence
  • sequence elements capable of forming secondary structure e.g., stem-
  • CO-number refers to codon altered polynucleotides encoding GAA polypeptides and/or the encoded polypeptides, including variants.
  • C03-FL refers to the Full Length codon altered CO3 polynucleotide sequence or amino acid sequence (sometimes referred to herein as “CO3-WT-FL-AA” for the Amino Acid sequence and “CO3-FL-NA” for the Nucleic Acid sequence) encoded by the CO3 polynucleotide sequence.
  • amino acid sequences will be identical, as the amino acid sequences are not altered by the codon optimization.
  • sequence constructs of the disclosure include, but are not limited to, CO1-FL-WT-NA, CO1- FL-46-NA, CO1-FL-1-NA, CO1-FL-2-NA, CO1-FL-3-NA, CO1-FL-4-NA, CO1-FL-5-NA, CO1-FL-6-NA, CO1-FL-7-NA, CO1-FL-8-NA, CO1-FL-9-NA, CO1-FL-10-NA, CO1-FL- 11-NA, CO1-FL-12-NA, CO1-FL-13-NA, CO2-FL-WT-NA, CO2-FL-46-NA, CO2-FL-1- NA, CO2-FL-2-NA, CO2-FL-3-NA, CO2-FL-4-NA, CO2-FL-5-NA, CO2-FL-6-NA, CO2- FL-7-NA, CO2-FL-8-NA, CO2-FL-9-NA, CO2-FL-10-NA, CO2-FL-11-NA, CO2-FL-12- NA, CO2-FL-13-NA, CO3-FL-WT
  • muscle-specific expression refers to the preferential or predominant in vivo expression of a particular gene (e.g., a codon-altered, transgenic GAA gene) in musculoskeletal tissue, as compared to in other tissues.
  • muscle-specific expression means that at least 50% of all expression of the particular gene occurs within hepatic tissues of a subject.
  • muscle-specific expression means that at least 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 99%, or 100% of all expression of the particular gene occurs within musculoskeletal tissues of a subject.
  • a muscle-specific regulatory element is a regulatory element that drives musclespecific expression of a gene in musculoskeletal tissue.
  • the present disclosure provides codon-altered polynucleotides encoding a GAA polypeptide, e.g., a wild-type or variant GAA polypeptide. These codon-altered polynucleotides provide markedly improved expression of GAA glucosidase activity in vivo, as demonstrated in Example 3. Specifically, Applicants have achieved these advantages through the discovery of several codon-altered polynucleotide schemas, referred to herein as CO1, CO2, and CO3, for encoding a GAA polypeptide.
  • a codon-altered polynucleotide provided herein has a nucleotide sequence with high sequence identity to CO1-FL-WT-NA (SEQ ID NO: 60), CO2- FL-WT-NA (SEQ ID NO: 62), or CO3-FL-WT-NA (SEQ ID NO: 31) encoding a human GAA pre-pro-polypeptide.
  • the wild-type human GAA gene encodes a pre-pro-polypeptide having a 27 amino acid signal peptide (1-27 of SEQ ID NO:2) and an 42 amino acid pro-peptide (aa 28- 69 of SEQ ID NO:2), which are cleaved from the encoded polypeptide prior to activation of GAA.
  • signal peptides and/or pro-peptides may be mutated, replaced by signal peptides and/or pro-peptides from other genes or other organisms, or completely removed, without affecting the sequence of the mature polypeptide left after the signal and pro-peptide are removed by cellular processing.
  • a codon-altered polynucleotide provided herein has a nucleotide sequence with high sequence identity to CO1-FL-WT-NA (SEQ ID NO:60), CO2-FL-WT- NA (SEQ ID NO: 62), or CO3 -FL-WT-NA (SEQ ID NO: 31), e.g, where the wild-type human GAA signal peptide and/or propeptide has been modified or replaced with an alternative signal peptide and/or propeptide.
  • the improved expression of GAA glucosidase activity provided by the CO1, CO2, and CO3 polynucleotide sequences are further improved when placed in operable communication with a muscle-specific regulatory control element, such as Dph-CRE04_NA (SEQ ID NO: 48) or Csk-SH4-NA (SEQ ID NO: 49).
  • a muscle-specific regulatory control element such as Dph-CRE04_NA (SEQ ID NO: 48) or Csk-SH4-NA (SEQ ID NO: 49).
  • the disclosure provides polynucleotides having a codon-altered GAA polynucleotide that is operably linked to a muscle-specific regulatory control element.
  • the disclosure provides codon-altered polynucleotides encoding a GAA polypeptide containing one or more known amino acid substitution.
  • GAA variant polypeptides having advantageous properties e.g., improved specific activity, improved thermostability, and/or reduced immunogenicity
  • GAA variants 1-4, 6-13 and variant 4+6 e.g., GAA variants 1-4, 6-13 and variant 4+6.
  • the disclosure provides codon-altered polynucleotides encoding a GAA polypeptide containing one or more amino acid substitutions present in any one of GAA variants 1-4, 6-13 or variant 4+6.
  • the codon-altered polynucleotide encodes a GAA polypeptide having any combination of the amino acid substitutions present in the 4+6 variant: L32W, L36S, L37T, P47Q, Q58V, A70L, P86E, D95E, L117D, S143Q, T151I, T158S, T274N, R275K, A445G, T494E, E530V, N535R, L577T, L650G, L650S, L650T, L650E, L650Y, L650F, S676D, L678H, L678T, T700G, A719H, A758P, A820E, Q838K, L868F, L879E, R891H, Q902G, V921R, and S940A.
  • the codon-altered polynucleotide encodes a GAA polypeptide having all of the amino acid substitutions
  • GC content of human genes varies widely, from less than 25% to greater than 90%. However, in general, human genes with higher GC contents are expressed at higher levels. For example, Kudla et al. (PLoS Biol., 4(6):80 (2006)) demonstrate that increasing a gene’s GC content increases expression of the encoded polypeptide, primarily by increasing transcription and effecting a higher steady state level of the mRNA transcript. Generally, the desired GC content of a codon-optimized gene construct is thought to be equal or greater than 60%. However, native AAV genomes have GC contents of around 56%.
  • the codon-altered polynucleotides provided herein have a CG content that more closely matches the GC content of native AAV virions (e.g., around 56% GC), which is lower than the preferred CG contents of polynucleotides that are conventionally codon-optimized for expression in mammalian cells (e.g., at or above 60% GC).
  • CO1-FL-WT-NA (SEQ ID NO:60) has a GC content of about 63.7%
  • CO2-FL-WT-NA (SEQ ID NO:62) has a GC content of about 59.1%
  • CO3-FL-WT-NA (SEQ ID NO:31) has a GC content of about 57.5%.
  • the overall GC content of a codon-altered polynucleotide encoding a GAA polypeptide is no more than 60%. In some embodiments, the overall GC content of a codon-altered polynucleotide encoding a GAA polypeptide is no more than 59%. In some embodiments, the overall GC content of a codon-altered polynucleotide encoding a GAA polypeptide is no more than 58%.
  • the overall GC content of a codon-altered polynucleotide encoding a GAA polypeptide is no more than 57%. In some embodiments, the overall GC content of a codon- altered polynucleotide encoding a GAA polypeptide is no more than 56%. In some embodiments, the overall GC content of a codon-altered polynucleotide encoding a GAA polypeptide is no more than 55%. [00310] In some embodiments, the overall GC content of a codon-altered polynucleotide encoding a GAA polypeptide is from 55% to 60%.
  • the overall GC content of a codon-altered polynucleotide encoding a GAA polypeptide is from 56% to 60%. In some embodiments, the overall GC content of a codon-altered polynucleotide encoding a GAA polypeptide is from 57% to 60%. In some embodiments, the overall GC content of a codon-altered polynucleotide encoding a GAA polypeptide is from 58% to 60%. In some embodiments, the overall GC content of a codon-altered polynucleotide encoding a GAA polypeptide is from 59% to 60%.
  • the overall GC content of a codon- altered polynucleotide encoding a GAA polypeptide is from 55% to 59%. In some embodiments, the overall GC content of a codon-altered polynucleotide encoding a GAA polypeptide is from 56% to 59%. In some embodiments, the overall GC content of a codon- altered polynucleotide encoding a GAA polypeptide is from 57% to 59%. In some embodiments, the overall GC content of a codon-altered polynucleotide encoding a GAA polypeptide is from 58% to 59%.
  • the overall GC content of a codon- altered polynucleotide encoding a GAA polypeptide is from 55% to 58%. In some embodiments, the overall GC content of a codon-altered polynucleotide encoding a GAA polypeptide is from 56% to 58%. In some embodiments, the overall GC content of a codon- altered polynucleotide encoding a GAA polypeptide is from 57% to 58%. In some embodiments, the overall GC content of a codon-altered polynucleotide encoding a GAA polypeptide is from 55% to 57%. In some embodiments, the overall GC content of a codon- altered polynucleotide encoding a GAA polypeptide is from 56% to 57%.
  • the overall GC content of a codon-altered polynucleotide encoding a GAA polypeptide is 57.5 ⁇ 0.5%. In some embodiments, the overall GC content of a codon-altered polynucleotide encoding a GAA polypeptide is 57.5 ⁇ 0.4%. In some embodiments, the overall GC content of a codon-altered polynucleotide encoding a GAA polypeptide is 57.5 ⁇ 0.3%. In some embodiments, the overall GC content of a codon-altered polynucleotide encoding a GAA polypeptide is 57.5 ⁇ 0.2%.
  • the overall GC content of a codon-altered polynucleotide encoding a GAA polypeptide is 57.5 ⁇ 0.1%. In some embodiments, the overall GC content of a codon-altered polynucleotide encoding a GAA polypeptide is 57.5%.
  • CpG dinucleotides z.e., a cytosine nucleotide followed by a guanine nucleotide
  • CpG-depleted AAV vectors evade immune detection in mice, under certain circumstances (Faust et al., J. Clin. Invest. 2013; 123, 2994-3001).
  • the wildtype GAA coding sequence (SEQ ID NO: 1) contains over 120 CpG dinucleotides.
  • the codon-altered polynucleotides provided herein are codon-altered to reduce the number of CpG dinucleotides in the GAA coding sequence.
  • CO3-FL-WT-NA (SEQ ID NO:31) has no CpG dinucleotides
  • CO1- FL-WT-NA SEQ ID NO: 60
  • CO2-FL-WT-NA (SEQ ID NO: 62) has no CpG dinucleotides.
  • a sequence of a codon-altered polynucleotide encoding a GAA polypeptide has less than 20 CpG dinucleotides.
  • a sequence of a codon-altered polynucleotide encoding a GAA polypeptide has less than 15 CpG dinucleotides.
  • a sequence of a codon-altered polynucleotide encoding a GAA polypeptide has less than 12 CpG dinucleotides.
  • a sequence of a codon-altered polynucleotide encoding a GAA polypeptide has less than 10 CpG dinucleotides. In some embodiments, a sequence of a codon-altered polynucleotide encoding a GAA polypeptide has less than 5 CpG dinucleotides. In some embodiments, a sequence of a codon-altered polynucleotide encoding a GAA polypeptide has less than 3 CpG dinucleotides. In some embodiments, a sequence of a codon-altered polynucleotide encoding a GAA polypeptide has no CpG dinucleotides.
  • sequence of a codon-altered polynucleotide encoding a GAA polypeptide has no more than 18, 17, 16, 15, 14, 13, 12, 11, 10, 9, 8, 7, 6, 5, 4, 3, 2, 1, or no CpG dinucleotides.
  • a nucleic acid composition provided herein includes a GAA polynucleotide (e.g., a codon-altered polynucleotide) encoding a GAA polypeptide, where the GAA polynucleotide includes a nucleotide sequence having high sequence identity to all or a portion of the CO1 codon-optimized sequence.
  • GAA polynucleotide e.g., a codon-altered polynucleotide
  • the GAA polynucleotide includes a sequence having high sequence identity to the portion of the CO1 codon-optimized sequence that encodes for the mature GAA polypeptide. Accordingly, in some embodiments, the sequence of the codon- altered polynucleotide has at least 95% identity to C01-MP-WT-NA (SEQ ID NO:63). In a specific embodiment, the sequence of the codon-altered polynucleotide has at least 96% identity to C01-MP-WT-NA (SEQ ID NO:63). In a specific embodiment, the sequence of the codon-altered polynucleotide has at least 97% identity to C01-MP-WT-NA (SEQ ID NO:63).
  • the sequence of the codon-altered polynucleotide has at least 98% identity to C01-MP-WT-NA (SEQ ID NO:63). In a specific embodiment, the sequence of the codon-altered polynucleotide has at least 99% identity to C01-MP-WT-NA (SEQ ID NO:63). In a specific embodiment, the sequence of the codon-altered polynucleotide has at least 99.5% identity to C01-MP-WT-NA (SEQ ID NO:63). In a specific embodiment, the sequence of the codon-altered polynucleotide has at least 99.9% identity to C01-MP-WT- NA (SEQ ID NO:63).
  • the sequence of the codon-altered polynucleotide is C01-MP-WT-NA (SEQ ID NO:63).
  • SEQ ID NO:63 the sequence identity between a GAA polypeptide and the portion of the CO1 codon-optimized sequence that encodes for the mature GAA polypeptide. That is, the GAA polynucleotide may also encode for a signal peptide, a pro-peptide, and/or a purification/detection tag, but the sequence comparison should not include these sequences.
  • a GAA polynucleotide having high sequence identity to CO1- MP-WT-NA further includes a polynucleotide sequence encoding a GAA signal peptide having the amino acid sequence of SP-WT-AA (SEQ ID NO:43).
  • the GAA signal polynucleotide has a nucleic acid sequence that is at least 90%, 95%, 96%, 97%, 98%, or 100% identical to CO1-SP-WT-NA (SEQ ID NO:70).
  • a GAA polynucleotide having high sequence identity to CO1- MP-WT-NA further includes a polynucleotide sequence encoding a GAA propeptide having the amino acid sequence of PP-WT-AA (SEQ ID NO:39).
  • the GAA pro-peptide polynucleotide has a nucleic acid sequence that is at least 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to CO1-PP-WT-NA (SEQ ID NO:71).
  • the GAA pro-peptide polynucleotide has a nucleic acid sequence that is at least 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to CO1-PP-46-NA (SEQ ID NO:72).
  • a GAA polynucleotide having high sequence identity to CO1- MP-WT-NA further includes a polynucleotide sequence encoding a GAA propeptide having the amino acid sequence of PP-46-AA (SEQ ID NO:41).
  • the GAA pro-peptide polynucleotide has a nucleic acid sequence that is at least 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to CO1-PP-46-NA (SEQ ID NO:72).
  • the GAA polynucleotide includes a sequence having high sequence identity to the entirety of the CO1 codon-optimized sequence, encoding for the GAA pre-pro-polypeptide. Accordingly, in some embodiments, the sequence of the codon- altered polynucleotide has at least 95% identity to CO1-FL-WT-NA (SEQ ID NO:60). In a specific embodiment, the sequence of the codon-altered polynucleotide has at least 96% identity to CO1-FL-WT-NA (SEQ ID NO:60).
  • the sequence of the codon-altered polynucleotide has at least 97% identity to CO1-FL-WT-NA (SEQ ID NO: 60). In a specific embodiment, the sequence of the codon-altered polynucleotide has at least 98% identity to CO1-FL-WT-NA (SEQ ID NO:60). In a specific embodiment, the sequence of the codon-altered polynucleotide has at least 99% identity to CO1-FL-WT-NA (SEQ ID NO: 60). In a specific embodiment, the sequence of the codon-altered polynucleotide has at least 99.5% identity to CO1-FL-WT-NA (SEQ ID NO:60).
  • sequence of the codon-altered polynucleotide has at least 99.9% identity to CO1-FL-WT-NA (SEQ ID NO:60). In another specific embodiment, the sequence of the codon-altered polynucleotide is CO1-FL-WT-NA (SEQ ID NO:60).
  • the GAA polypeptide encoded by a GAA polynucleotide having high sequence identity to the CO1 codon-optimized sequence includes an amino acid sequencing having high sequence identity to the human wild-type mature GAA polypeptide (MP-WT-AA; SEQ ID NO:35). Accordingly, in some embodiments, the encoded GAA polypeptide has a sequence that is at least 90% identical to MP-WT-AA (SEQ ID NO:35). In some embodiments, the encoded GAA polypeptide has a sequence that is at least 95% identical to MP-WT-AA (SEQ ID NO:35).
  • the encoded GAA polypeptide has a sequence that is at least 96% identical to MP-WT-AA (SEQ ID NO:35). In some embodiments, the encoded GAA polypeptide has a sequence that is at least 97% identical to MP-WT-AA (SEQ ID NO:35). In some embodiments, the encoded GAA polypeptide has a sequence that is at least 98% identical to MP-WT-AA (SEQ ID NO:35). In some embodiments, the encoded GAA polypeptide has a sequence that is at least 99% identical to MP-WT-AA (SEQ ID NO:35).
  • the encoded GAA polypeptide has a sequence that is at least 99.5% identical to MP-WT-AA (SEQ ID NO:35). In some embodiments, the encoded GAA polypeptide has a sequence that is at least 99.8% identical to MP-WT-AA (SEQ ID NO:35). In some embodiments, the encoded GAA polypeptide has a sequence that is at least 99.8% identical to MP-WT-AA (SEQ ID NO:35). In some embodiments, the encoded GAA polypeptide has a sequence identical to MP-WT- AA (SEQ ID NO:35).
  • the GAA polypeptide may also include a signal peptide, a pro-peptide, and/or a purification/detection tag, but the sequence comparison should not include these sequences.
  • the GAA polypeptide encoded by a GAA polynucleotide having high sequence identity to the CO1 codon-optimized sequence includes an amino acid sequencing having high sequence identity to the human wild-type GAA pre-pro-polypeptide (FL-WT-AA; SEQ ID NO:2). Accordingly, in some embodiments, the encoded GAA polypeptide has a sequence that is at least 90% identical to FL-WT-AA (SEQ ID NO:2). In some embodiments, the encoded GAA polypeptide has a sequence that is at least 95% identical to FL-WT-AA (SEQ ID NO:2).
  • the encoded GAA polypeptide has a sequence that is at least 96% identical to FL-WT-AA (SEQ ID NO:2). In some embodiments, the encoded GAA polypeptide has a sequence that is at least 97% identical to FL-WT-AA (SEQ ID NO:2). In some embodiments, the encoded GAA polypeptide has a sequence that is at least 98% identical to FL-WT-AA (SEQ ID NO:2). In some embodiments, the encoded GAA polypeptide has a sequence that is at least 99% identical to FL-WT-AA (SEQ ID NO:2). In some embodiments, the encoded GAA polypeptide has a sequence that is at least 99.5% identical to FL-WT-AA (SEQ ID NO:2).
  • the encoded GAA polypeptide has a sequence that is at least 99.8% identical to FL-WT-AA (SEQ ID NO:2). In some embodiments, the encoded GAA polypeptide has a sequence that is at least 99.8% identical to FL-WT-AA (SEQ ID NO:2). In some embodiments, the encoded GAA polypeptide has a sequence identical to FL-WT-AA (SEQ ID NO:2).
  • the GAA polypeptide may also include a purification/detection tag, but the sequence comparison should not include these sequences.
  • the GAA polypeptide encoded by a GAA polynucleotide having high sequence identity to the CO1 codon-optimized sequence includes one or more known amino acid substitutions, e.g., one or more amino acid substitutions described in U.S. Patent Application Publication No. 2021/0189365, the content of which is incorporated herein by reference in its entirety.
  • the GAA polypeptide encoded by a GAA polynucleotide having high sequence identity to the CO1 codon-optimized sequence includes one or more amino acid substitutions present in one of GAA variants 1-4 described herein.
  • the GAA polypeptide encoded by a GAA polynucleotide having high sequence identity to the CO1 codon-optimized sequence includes one or more amino acid substitutions present in one of GAA variants 6-13 described herein.
  • the GAA polypeptide encoded by a GAA polynucleotide having high sequence identity to the CO1 codon-optimized sequence includes one or more amino acid substitutions present in GAA variant 4+6: L32W, L36S, L37T, P47Q, Q58V, A70L, P86E, D95E, L117D, S143Q, T151I, T158S, T274N, R275K, A445G, T494E, E530V, N535R, L577T, L650G, L650S, L650T, L650E, L650Y, L650F, S676D, L678H, L678T, T700G, A719H, A758P, A820E, Q838K, L868F, L879E, R891H, Q902G, V921R, and S940A.
  • the sequence of the codon-altered polynucleotide has at least 95% identity to the portion of a codon-optimized GAA polynucleotide encoding the GAA variant 46 mature polypeptide of the GAA (CO1-MP-46- NA; SEQ ID NO:66). In a specific embodiment, the sequence of the codon-altered polynucleotide has at least 96% identity to CO1-MP-46-NA (SEQ ID NO: 66). In a specific embodiment, the sequence of the codon-altered polynucleotide has at least 97% identity to CO1-MP-46-NA (SEQ ID NO:66).
  • the sequence of the codon- altered polynucleotide has at least 98% identity to CO1-MP-46-NA (SEQ ID NO:66). In a specific embodiment, the sequence of the codon-altered polynucleotide has at least 99% identity to CO1-MP-46-NA (SEQ ID NO:66). In a specific embodiment, the sequence of the codon-altered polynucleotide has at least 99.5% identity to CO1-MP-46-NA (SEQ ID NO:66). In a specific embodiment, the sequence of the codon-altered polynucleotide has at least 99.9% identity to CO1-MP-46-NA (SEQ ID NO:66).
  • the sequence of the codon-altered polynucleotide is CO1-MP-46-NA (SEQ ID NO:66).
  • SEQ ID NO:66 When determining the sequence identity between a GAA polypeptide and the portion of the CO1 codon-optimized sequence that encodes for the mature GAA polypeptide, only the portions of the sequence encoding the mature polypeptide should be considered. That is, the GAA polypeptide may also encode for a signal peptide, a pro-peptide, and/or a purification/detection tag, but the sequence comparison should not include these sequences.
  • a GAA polynucleotide having high sequence identity to CO1- MP-46-NA further includes a polynucleotide sequence encoding a GAA signal peptide having the amino acid sequence of SP-WT-AA (SEQ ID NO:43).
  • the GAA signal polynucleotide has a nucleic acid sequence that is at least 90%, 95%, 96%, 97%, 98%, or 100% identical to CO1-SP-WT-NA (SEQ ID NO:70).
  • a GAA polynucleotide having high sequence identity to CO1- MP-46-NA further includes a polynucleotide sequence encoding a GAA propeptide having the amino acid sequence of PP-WT-AA (SEQ ID NO:39).
  • the GAA pro-peptide polynucleotide has a nucleic acid sequence that is at least 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to CO1-PP-WT-NA (SEQ ID NO:71).
  • a GAA polynucleotide having high sequence identity to CO1- MP-46-NA further includes a polynucleotide sequence encoding a GAA propeptide having the amino acid sequence of PP-46-AA (SEQ ID NO:41).
  • the GAA pro-peptide polynucleotide has a nucleic acid sequence that is at least 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to CO1-PP-46-NA (SEQ ID NO:72).
  • the GAA polynucleotide includes a sequence having high sequence identity to the entirety of a codon-optimized GAA polynucleotide encoding the variant 46 GAA pre-pro-polypeptide. Accordingly, in some embodiments, the sequence of the codon-altered polynucleotide has at least 95% identity to CO1-FL-46-NA (SEQ ID NO:65). In a specific embodiment, the sequence of the codon-altered polynucleotide has at least 96% identity to CO1-FL-46-NA (SEQ ID NO: 65).
  • the sequence of the codon-altered polynucleotide has at least 97% identity to CO1-FL-46-NA (SEQ ID NO:65). In a specific embodiment, the sequence of the codon-altered polynucleotide has at least 98% identity to CO1-FL-46-NA (SEQ ID NO:65). In a specific embodiment, the sequence of the codon-altered polynucleotide has at least 99% identity to CO1-FL-46-NA (SEQ ID NO:65). In a specific embodiment, the sequence of the codon-altered polynucleotide has at least 99.5% identity to CO1-FL-46-NA (SEQ ID NO:65).
  • the sequence of the codon-altered polynucleotide has at least 99.9% identity to CO1-FL-46- NA (SEQ ID NO:65). In another specific embodiment, the sequence of the codon-altered polynucleotide is CO1-FL-46-NA (SEQ ID NO:65).
  • the GAA polypeptide encoded by a GAA polynucleotide having high sequence identity to the CO1 codon-optimized sequence includes an amino acid sequencing having high sequence identity to the human variant 4+6 mature GAA polypeptide (MP-46-AA; SEQ ID NO: 37).
  • the encoded GAA polypeptide has a sequence that is at least 90% identical to MP-46-AA (SEQ ID NO:37). In some embodiments, the encoded GAA polypeptide has a sequence that is at least 95% identical to MP-46-AA (SEQ ID NO:37). In some embodiments, the encoded GAA polypeptide has a sequence that is at least 96% identical to MP-46-AA (SEQ ID NO:37). In some embodiments, the encoded GAA polypeptide has a sequence that is at least 97% identical to MP-46-AA (SEQ ID NO:37). In some embodiments, the encoded GAA polypeptide has a sequence that is at least 98% identical to MP-46-AA (SEQ ID NO:37).
  • the encoded GAA polypeptide has a sequence that is at least 99% identical to MP-46-AA (SEQ ID NO:37). In some embodiments, the encoded GAA polypeptide has a sequence that is at least 99.5% identical to MP-46-AA (SEQ ID NO:37). In some embodiments, the encoded GAA polypeptide has a sequence that is at least 99.8% identical to MP-46-AA (SEQ ID NO:37). In some embodiments, the encoded GAA polypeptide has a sequence that is at least 99.8% identical to MP-46-AA (SEQ ID NO:37). In some embodiments, the encoded GAA polypeptide has a sequence identical to MP-46-AA (SEQ ID NO:37).
  • the GAA polypeptide may also include a signal peptide, a pro-peptide, and/or a purification/detection tag, but the sequence comparison should not include these sequences.
  • the GAA polypeptide encoded by a GAA polynucleotide having high sequence identity to the CO1 codon-optimized sequence includes an amino acid sequencing having high sequence identity to the human variant 4+6 GAA pre-pro- polypeptide (FL-46-AA; SEQ ID NO:30). Accordingly, in some embodiments, the encoded GAA polypeptide has a sequence that is at least 90% identical to FL-46- AA (SEQ ID NO:30). In some embodiments, the encoded GAA polypeptide has a sequence that is at least 95% identical to FL-46-AA (SEQ ID NO: 30).
  • the encoded GAA polypeptide has a sequence that is at least 96% identical to FL-46-AA (SEQ ID NO:30). In some embodiments, the encoded GAA polypeptide has a sequence that is at least 97% identical to FL-46-AA (SEQ ID NO:30). In some embodiments, the encoded GAA polypeptide has a sequence that is at least 98% identical to FL-46-AA (SEQ ID NO:30). In some embodiments, the encoded GAA polypeptide has a sequence that is at least 99% identical to FL-46-AA (SEQ ID NO:30). In some embodiments, the encoded GAA polypeptide has a sequence that is at least 99.5% identical to FL-46-AA (SEQ ID NO:30).
  • the encoded GAA polypeptide has a sequence that is at least 99.8% identical to FL-46-AA (SEQ ID NO:30). In some embodiments, the encoded GAA polypeptide has a sequence that is at least 99.8% identical to FL-46-AA (SEQ ID NO:30). In some embodiments, the encoded GAA polypeptide has a sequence identical to FL-46-AA (SEQ ID NO:30). When determining the sequence identity between an encoded GAA polypeptide and the GAA pre-pro-polypeptide, only the portions of the sequence corresponding to the pre-pro-polypeptide should be considered. That is, the GAA polypeptide may also include a purification/detection tag, but the sequence comparison should not include these sequences.
  • the nucleotide sequence of the GAA polynucleotide having high sequence identity to a CO1 codon-optimized sequence (e.g., SEQ ID NO:60, 63, 65, or 66) has a reduced GC content, as compared to the wild-type GAA coding sequence SEQ ID NO: 1, as described above. Accordingly, in some embodiments, the sequence of the codon- altered polynucleotide having high sequence identity to a CO1 codon-optimized sequence has a GC content of no more than 66%.
  • the sequence of the codon-altered polynucleotide having high sequence identity to a CO1 codon-optimized sequence has a GC content of no more than 63.5%. In some embodiments, the sequence of the codon-altered polynucleotide having high sequence identity to a CO1 codon-optimized sequence has a GC content of no more than 65%, no more than 64%, no more than 63%, no more than 62%, or no more than 61%.
  • the sequence of the codon-altered polynucleotide having high sequence identity to a CO1 codon-optimized sequence has a GC content of from 61% to 66%. In some embodiments, the sequence of the codon-altered polynucleotide having high sequence identity to a CO1 codon-optimized sequence has a GC content of from 62% to 66%, from 63% to 66%, from 64% to 66%, from 65% to 66%, from 61% to 65%, from 62% to 65%, from 63% to 65%, from 64% to 65%, from 61% to 64%, from 62% to 64%, from 63% to 64%, from 61% to 63%, from 62% to 63%, or from 61% to 62%.
  • the sequence of the codon-altered polynucleotide having high sequence identity to a CO1 codon-optimized sequence has a GC content of 63.5% ⁇ 1.0. In some embodiments, the sequence of the codon-altered polynucleotide having high sequence identity to a CO1 codon-optimized sequence has a GC content of 63.5% ⁇ 0.8. In some embodiments, the sequence of the codon-altered polynucleotide having high sequence identity to a CO1 codon-optimized sequence has a GC content of 63.5% ⁇ 0.6.
  • the sequence of the codon-altered polynucleotide having high sequence identity to a CO1 codon-optimized sequence has a GC content of 63.5% ⁇ 0.5. In some embodiments, the sequence of the codon-altered polynucleotide having high sequence identity to a CO1 codon-optimized sequence has a GC content of 63.5% ⁇ 0.4. In some embodiments, the sequence of the codon-altered polynucleotide having high sequence identity to a CO1 codon-optimized sequence has a GC content of 63.5% ⁇ 0.3.
  • the sequence of the codon-altered polynucleotide having high sequence identity to a CO1 codon-optimized sequence has a GC content of 63.5% ⁇ 0.2. In some embodiments, the sequence of the codon-altered polynucleotide having high sequence identity to a CO1 codon-optimized sequence has a GC content of 63.5% ⁇ 0.1. In some embodiments, the sequence of the codon-altered polynucleotide having high sequence identity to a CO1 codon-optimized sequence has a GC content of 63.5%.
  • the nucleotide sequence of the GAA polynucleotide having high sequence identity to a CO1 codon-optimized sequence (e.g., SEQ ID NO:60, 63, 65, or 66) has a reduced number of CpG dinucleotides, as compared to the wild-type GAA coding sequence SEQ ID NO: 1, as described above. Accordingly, in some embodiments, the sequence of the codon-altered polynucleotide having high sequence identity to a CO1 codon- optimized sequence has no more than 15 CpG dinucleotides.
  • the sequence of the codon-altered polynucleotide having high sequence identity to a CO1 codon- optimized sequence has no more than 10 CpG dinucleotides. In some embodiments, the sequence of the codon-altered polynucleotide having high sequence identity to a CO1 codon- optimized sequence has no more than 5 CpG dinucleotides. In some embodiments, the sequence of the codon-altered polynucleotide having high sequence identity to a CO1 codon- optimized sequence has no more than 4 CpG dinucleotides.
  • the sequence of the codon-altered polynucleotide having high sequence identity to a CO1 codon- optimized sequence has no more than 3 CpG dinucleotides. In some embodiments, the sequence of the codon-altered polynucleotide having high sequence identity to a CO1 codon- optimized sequence has no more than 2 CpG dinucleotides. In some embodiments, the sequence of the codon-altered polynucleotide having high sequence identity to a CO1 codon- optimized sequence has no more than 1 CpG dinucleotide. In some embodiments, the sequence of the codon-altered polynucleotide having high sequence identity to a CO1 codon- optimized sequence has no CpG dinucleotides.
  • a nucleic acid composition provided herein includes a GAA polynucleotide (e.g., a codon-altered polynucleotide) encoding a GAA polypeptide, where the GAA polynucleotide includes a nucleotide sequence having high sequence identity to all or a portion of the CO2 codon-optimized sequence.
  • GAA polynucleotide e.g., a codon-altered polynucleotide
  • the GAA polynucleotide includes a sequence having high sequence identity to the portion of the CO2 codon-optimized sequence that encodes for the mature GAA polypeptide. Accordingly, in some embodiments, the sequence of the codon- altered polynucleotide has at least 95% identity to CO2-MP-WT-NA (SEQ ID NO:64). In a specific embodiment, the sequence of the codon-altered polynucleotide has at least 96% identity to CO2-MP-WT-NA (SEQ ID NO:64). In a specific embodiment, the sequence of the codon-altered polynucleotide has at least 97% identity to CO2-MP-WT-NA (SEQ ID NO:64).
  • the sequence of the codon-altered polynucleotide has at least 98% identity to CO2-MP-WT-NA (SEQ ID NO:64). In a specific embodiment, the sequence of the codon-altered polynucleotide has at least 99% identity to CO2-MP-WT-NA (SEQ ID NO:64). In a specific embodiment, the sequence of the codon-altered polynucleotide has at least 99.5% identity to CO2-MP-WT-NA (SEQ ID NO:64). In a specific embodiment, the sequence of the codon-altered polynucleotide has at least 99.9% identity to CO2-MP-WT- NA (SEQ ID NO:64).
  • the sequence of the codon-altered polynucleotide is C02-MP-WT-NA (SEQ ID NO:64).
  • SEQ ID NO:64 When determining the sequence identity between a GAA polypeptide and the portion of the CO2 codon-optimized sequence that encodes for the mature GAA polypeptide, only the portions of the sequence encoding the mature polypeptide should be considered. That is, the GAA polynucleotide may also encode for a signal peptide, a pro-peptide, and/or a purification/detection tag, but the sequence comparison should not include these sequences.
  • a GAA polynucleotide having high sequence identity to CO2- MP-WT-NA further includes a polynucleotide sequence encoding a GAA signal peptide having the amino acid sequence of SP-WT-AA (SEQ ID NO:43).
  • the GAA signal polynucleotide has a nucleic acid sequence that is at least 90%, 95%, 96%, 97%, 98%, or 100% identical to CO2-SP-WT-NA (SEQ ID NO:73).
  • a GAA polynucleotide having high sequence identity to CO2- MP-WT-NA further includes a polynucleotide sequence encoding a GAA propeptide having the amino acid sequence of PP-WT-AA (SEQ ID NO:39).
  • the GAA pro-peptide polynucleotide has a nucleic acid sequence that is at least 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to CO2-PP-WT-NA (SEQ ID NO:74).
  • the GAA pro-peptide polynucleotide has a nucleic acid sequence that is at least 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to CO2-PP-46-NA (SEQ ID NO:75).
  • a GAA polynucleotide having high sequence identity to CO2- MP-WT-NA further includes a polynucleotide sequence encoding a GAA propeptide having the amino acid sequence of PP-46-AA (SEQ ID NO:41).
  • the GAA pro-peptide polynucleotide has a nucleic acid sequence that is at least 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to CO2-PP-46-NA (SEQ ID NO:75).
  • the GAA polynucleotide includes a sequence having high sequence identity to the entirety of the CO2 codon-optimized sequence, encoding for the GAA pre-pro-polypeptide. Accordingly, in some embodiments, the sequence of the codon- altered polynucleotide has at least 95% identity to CO2-FL-WT-NA (SEQ ID NO:62). In a specific embodiment, the sequence of the codon-altered polynucleotide has at least 96% identity to CO2-FL-WT-NA (SEQ ID NO:62).
  • the sequence of the codon-altered polynucleotide has at least 97% identity to CO2-FL-WT-NA (SEQ ID NO: 62). In a specific embodiment, the sequence of the codon-altered polynucleotide has at least 98% identity to CO2-FL-WT-NA (SEQ ID NO:62). In a specific embodiment, the sequence of the codon-altered polynucleotide has at least 99% identity to CO2-FL-WT-NA (SEQ ID NO: 62). In a specific embodiment, the sequence of the codon-altered polynucleotide has at least 99.5% identity to CO2-FL-WT-NA (SEQ ID NO:62).
  • sequence of the codon-altered polynucleotide has at least 99.9% identity to CO2-FL-WT-NA (SEQ ID NO:62). In another specific embodiment, the sequence of the codon-altered polynucleotide is CO2-FL-WT-NA (SEQ ID NO:62).
  • the GAA polypeptide encoded by a GAA polynucleotide having high sequence identity to the CO2 codon-optimized sequence includes an amino acid sequencing having high sequence identity to the human wild-type mature GAA polypeptide (MP-WT-AA; SEQ ID NO:35). Accordingly, in some embodiments, the encoded GAA polypeptide has a sequence that is at least 90% identical to MP-WT-AA (SEQ ID NO:35). In some embodiments, the encoded GAA polypeptide has a sequence that is at least 95% identical to MP-WT-AA (SEQ ID NO:35).
  • the encoded GAA polypeptide has a sequence that is at least 96% identical to MP-WT-AA (SEQ ID NO:35). In some embodiments, the encoded GAA polypeptide has a sequence that is at least 97% identical to MP-WT-AA (SEQ ID NO:35). In some embodiments, the encoded GAA polypeptide has a sequence that is at least 98% identical to MP-WT-AA (SEQ ID NO:35). In some embodiments, the encoded GAA polypeptide has a sequence that is at least 99% identical to MP-WT-AA (SEQ ID NO:35).
  • the encoded GAA polypeptide has a sequence that is at least 99.5% identical to MP-WT-AA (SEQ ID NO:35). In some embodiments, the encoded GAA polypeptide has a sequence that is at least 99.8% identical to MP-WT-AA (SEQ ID NO:35). In some embodiments, the encoded GAA polypeptide has a sequence that is at least 99.8% identical to MP-WT-AA (SEQ ID NO:35). In some embodiments, the encoded GAA polypeptide has a sequence identical to MP-WT- AA (SEQ ID NO:35).
  • the GAA polypeptide may also include a signal peptide, a pro-peptide, and/or a purification/detection tag, but the sequence comparison should not include these sequences.
  • the GAA polypeptide encoded by a GAA polynucleotide having high sequence identity to the CO2 codon-optimized sequence includes an amino acid sequencing having high sequence identity to the human wild-type GAA pre-pro-polypeptide (FL-WT-AA; SEQ ID NO:2). Accordingly, in some embodiments, the encoded GAA polypeptide has a sequence that is at least 90% identical to FL-WT-AA (SEQ ID NO:2). In some embodiments, the encoded GAA polypeptide has a sequence that is at least 95% identical to FL-WT-AA (SEQ ID NO:2).
  • the encoded GAA polypeptide has a sequence that is at least 96% identical to FL-WT-AA (SEQ ID NO:2). In some embodiments, the encoded GAA polypeptide has a sequence that is at least 97% identical to FL-WT-AA (SEQ ID NO:2). In some embodiments, the encoded GAA polypeptide has a sequence that is at least 98% identical to FL-WT-AA (SEQ ID NO:2). In some embodiments, the encoded GAA polypeptide has a sequence that is at least 99% identical to FL-WT-AA (SEQ ID NO:2). In some embodiments, the encoded GAA polypeptide has a sequence that is at least 99.5% identical to FL-WT-AA (SEQ ID NO:2).
  • the encoded GAA polypeptide has a sequence that is at least 99.8% identical to FL-WT-AA (SEQ ID NO:2). In some embodiments, the encoded GAA polypeptide has a sequence that is at least 99.8% identical to FL-WT-AA (SEQ ID NO:2). In some embodiments, the encoded GAA polypeptide has a sequence identical to FL-WT-AA (SEQ ID NO:2).
  • the GAA polypeptide may also include a purification/detection tag, but the sequence comparison should not include these sequences.
  • the GAA polypeptide encoded by a GAA polynucleotide having high sequence identity to the CO2 codon-optimized sequence includes one or more known amino acid substitutions, e.g., one or more amino acid substitutions described in U.S. Patent Application Publication No. 2021/0189365, the content of which is incorporated herein by reference in its entirety.
  • the GAA polypeptide encoded by a GAA polynucleotide having high sequence identity to the CO2 codon-optimized sequence includes one or more amino acid substitutions present in one of GAA variants 1-4 described herein.
  • the GAA polypeptide encoded by a GAA polynucleotide having high sequence identity to the CO2 codon-optimized sequence includes one or more amino acid substitutions present in one of GAA variants 6-13 described herein.
  • the GAA polypeptide encoded by a GAA polynucleotide having high sequence identity to the CO2 codon-optimized sequence includes one or more amino acid substitutions present in GAA variant 4+6: L32W, L36S, L37T, P47Q, Q58V, A70L, P86E, D95E, L117D, S143Q, T151I, T158S, T274N, R275K, A445G, T494E, E530V, N535R, L577T, L650G, L650S, L650T, L650E, L650Y, L650F, S676D, L678H, L678T, T700G, A719H, A758P, A820E, Q838K, L868F, L879E, R891H, Q902G, V921R, and S940A.
  • the sequence of the codon-altered polynucleotide has at least 95% identity to the portion of a codon-optimized GAA polynucleotide encoding the GAA variant 46 mature polypeptide of the GAA (CO2-MP-46- NA; SEQ ID NO:68).
  • the sequence of the codon-altered polynucleotide has at least 96% identity to CO2-MP-46-NA (SEQ ID NO: 68).
  • the sequence of the codon-altered polynucleotide has at least 97% identity to CO2-MP-46-NA (SEQ ID NO:68).
  • the sequence of the codon- altered polynucleotide has at least 98% identity to CO2-MP-46-NA (SEQ ID NO:68). In a specific embodiment, the sequence of the codon-altered polynucleotide has at least 99% identity to CO2-MP-46-NA (SEQ ID NO:68). In a specific embodiment, the sequence of the codon-altered polynucleotide has at least 99.5% identity to CO2-MP-46-NA (SEQ ID NO:68). In a specific embodiment, the sequence of the codon-altered polynucleotide has at least 99.9% identity to CO2-MP-46-NA (SEQ ID NO:68).
  • the sequence of the codon-altered polynucleotide is CO2-MP-46-NA (SEQ ID NO:68).
  • SEQ ID NO:68 When determining the sequence identity between a GAA polypeptide and the portion of the CO2 codon-optimized sequence that encodes for the mature GAA polypeptide, only the portions of the sequence encoding the mature polypeptide should be considered. That is, the GAA polypeptide may also encode for a signal peptide, a pro-peptide, and/or a purification/detection tag, but the sequence comparison should not include these sequences.
  • a GAA polynucleotide having high sequence identity to CO2- MP-46-NA further includes a polynucleotide sequence encoding a GAA signal peptide having the amino acid sequence of SP-WT-AA (SEQ ID NO:43).
  • the GAA signal polynucleotide has a nucleic acid sequence that is at least 90%, 95%, 96%, 97%, 98%, or 100% identical to CO2-SP-WT-NA (SEQ ID NO:73).
  • a GAA polynucleotide having high sequence identity to CO2- MP-46-NA further includes a polynucleotide sequence encoding a GAA propeptide having the amino acid sequence of PP-WT-AA (SEQ ID NO:39).
  • the GAA pro-peptide polynucleotide has a nucleic acid sequence that is at least 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to CO2-PP-WT-NA (SEQ ID NO:74).
  • a GAA polynucleotide having high sequence identity to CO2- MP-46-NA further includes a polynucleotide sequence encoding a GAA propeptide having the amino acid sequence of PP-46-AA (SEQ ID NO:41).
  • the GAA pro-peptide polynucleotide has a nucleic acid sequence that is at least 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to CO2-PP-46-NA (SEQ ID NO:75).
  • the GAA polynucleotide includes a sequence having high sequence identity to the entirety of a codon-optimized GAA polynucleotide encoding the variant 46 GAA pre-pro-polypeptide. Accordingly, in some embodiments, the sequence of the codon-altered polynucleotide has at least 95% identity to CO2-FL-46-NA (SEQ ID NO:67). In a specific embodiment, the sequence of the codon-altered polynucleotide has at least 96% identity to CO2-FL-46-NA (SEQ ID NO: 67).
  • the sequence of the codon-altered polynucleotide has at least 97% identity to CO2-FL-46-NA (SEQ ID NO:67). In a specific embodiment, the sequence of the codon-altered polynucleotide has at least 98% identity to CO2-FL-46-NA (SEQ ID NO:67). In a specific embodiment, the sequence of the codon-altered polynucleotide has at least 99% identity to CO2-FL-46-NA (SEQ ID NO:67). In a specific embodiment, the sequence of the codon-altered polynucleotide has at least 99.5% identity to CO2-FL-46-NA (SEQ ID NO:67).
  • sequence of the codon-altered polynucleotide has at least 99.9% identity to CO2-FL-46- NA (SEQ ID NO:67). In another specific embodiment, the sequence of the codon-altered polynucleotide is CO2-FL-46-NA (SEQ ID NO:67).
  • the GAA polypeptide encoded by a GAA polynucleotide having high sequence identity to the CO2 codon-optimized sequence includes an amino acid sequencing having high sequence identity to the human variant 4+6 mature GAA polypeptide (MP-46-AA; SEQ ID NO: 37). Accordingly, in some embodiments, the encoded GAA polypeptide has a sequence that is at least 90% identical to MP-46-AA (SEQ ID NO:37). In some embodiments, the encoded GAA polypeptide has a sequence that is at least 95% identical to MP-46-AA (SEQ ID NO:37).
  • the encoded GAA polypeptide has a sequence that is at least 96% identical to MP-46-AA (SEQ ID NO:37). In some embodiments, the encoded GAA polypeptide has a sequence that is at least 97% identical to MP-46-AA (SEQ ID NO:37). In some embodiments, the encoded GAA polypeptide has a sequence that is at least 98% identical to MP-46-AA (SEQ ID NO:37). In some embodiments, the encoded GAA polypeptide has a sequence that is at least 99% identical to MP-46-AA (SEQ ID NO:37). In some embodiments, the encoded GAA polypeptide has a sequence that is at least 99.5% identical to MP-46-AA (SEQ ID NO:37).
  • the encoded GAA polypeptide has a sequence that is at least 99.8% identical to MP-46-AA (SEQ ID NO:37). In some embodiments, the encoded GAA polypeptide has a sequence that is at least 99.8% identical to MP-46-AA (SEQ ID NO:37). In some embodiments, the encoded GAA polypeptide has a sequence identical to MP-46-AA (SEQ ID NO:37). When determining the sequence identity between an encoded GAA polypeptide and the mature GAA polypeptide, only the portions of the sequence corresponding to the mature polypeptide should be considered.
  • the GAA polypeptide may also include a signal peptide, a pro-peptide, and/or a purification/detection tag, but the sequence comparison should not include these sequences.
  • the GAA polypeptide encoded by a GAA polynucleotide having high sequence identity to the CO2 codon-optimized sequence includes an amino acid sequencing having high sequence identity to the human variant 4+6 GAA pre-pro- polypeptide (FL-46-AA; SEQ ID NO:30). Accordingly, in some embodiments, the encoded GAA polypeptide has a sequence that is at least 90% identical to FL-46- AA (SEQ ID NO:30).
  • the encoded GAA polypeptide has a sequence that is at least 95% identical to FL-46-AA (SEQ ID NO: 30). In some embodiments, the encoded GAA polypeptide has a sequence that is at least 96% identical to FL-46-AA (SEQ ID NO:30). In some embodiments, the encoded GAA polypeptide has a sequence that is at least 97% identical to FL-46-AA (SEQ ID NO:30). In some embodiments, the encoded GAA polypeptide has a sequence that is at least 98% identical to FL-46-AA (SEQ ID NO:30). In some embodiments, the encoded GAA polypeptide has a sequence that is at least 99% identical to FL-46-AA (SEQ ID NO:30).
  • the encoded GAA polypeptide has a sequence that is at least 99.5% identical to FL-46-AA (SEQ ID NO:30). In some embodiments, the encoded GAA polypeptide has a sequence that is at least 99.8% identical to FL-46-AA (SEQ ID NO:30). In some embodiments, the encoded GAA polypeptide has a sequence that is at least 99.8% identical to FL-46-AA (SEQ ID NO:30). In some embodiments, the encoded GAA polypeptide has a sequence identical to FL-46-AA (SEQ ID NO:30).
  • the GAA polypeptide may also include a purification/detection tag, but the sequence comparison should not include these sequences.
  • the nucleotide sequence of the GAA polynucleotide having high sequence identity to a CO2 codon-optimized sequence (e.g., SEQ ID NO: 62, 64, 67, or 68) has a reduced GC content, as compared to the wild-type GAA coding sequence SEQ ID NO: 1, as described above. Accordingly, in some embodiments, the sequence of the codon- altered polynucleotide having high sequence identity to a CO2 codon-optimized sequence has a GC content of no more than 61.5%.
  • the sequence of the codon- altered polynucleotide having high sequence identity to a CO2 codon-optimized sequence has a GC content of no more than 59%. In some embodiments, the sequence of the codon-altered polynucleotide having high sequence identity to a CO2 codon-optimized sequence has a GC content of no more than 60.5%, no more than 59.5%, no more than 58.5%, no more than 57.5%, or no more than 56.5%.
  • the sequence of the codon-altered polynucleotide having high sequence identity to a CO2 codon-optimized sequence has a GC content of from 56.5% to 61.5%. In some embodiments, the sequence of the codon-altered polynucleotide having high sequence identity to a CO2 codon-optimized sequence has a GC content of from 57.5% to 61.5%, from 58.5% to 61.5%, from 59.5% to 61.5%, from 60.5% to 61.5%, from 56.5% to 60.5%, from 57.5% to 60.5%, from 58.5% to 60.5%, from 59.5% to 60.5%, from 56.5% to 59.5%, from 57.5% to 59.5%, from 58.5% to 59.5%, from 56.5% to 58.5%, from 57.5% to 58.5%, from 57.5% to 58.5%, or from 56.5% to 57.5%.5%.
  • the sequence of the codon-altered polynucleotide having high sequence identity to a CO2 codon-optimized sequence has a GC content of 59% ⁇ 1.0. In some embodiments, the sequence of the codon-altered polynucleotide having high sequence identity to a CO2 codon-optimized sequence has a GC content of 59% ⁇ 0.8. In some embodiments, the sequence of the codon-altered polynucleotide having high sequence identity to a CO2 codon-optimized sequence has a GC content of 59% ⁇ 0.6.
  • the sequence of the codon-altered polynucleotide having high sequence identity to a CO2 codon-optimized sequence has a GC content of 59% ⁇ 0.5. In some embodiments, the sequence of the codon-altered polynucleotide having high sequence identity to a CO2 codon-optimized sequence has a GC content of 59% ⁇ 0.4. In some embodiments, the sequence of the codon-altered polynucleotide having high sequence identity to a CO2 codon-optimized sequence has a GC content of 59% ⁇ 0.3.
  • the sequence of the codon-altered polynucleotide having high sequence identity to a CO2 codon-optimized sequence has a GC content of 59% ⁇ 0.2. In some embodiments, the sequence of the codon-altered polynucleotide having high sequence identity to a CO2 codon-optimized sequence has a GC content of 59% ⁇ 0.1. In some embodiments, the sequence of the codon-altered polynucleotide having high sequence identity to a CO2 codon-optimized sequence has a GC content of 59%.
  • the nucleotide sequence of the GAA polynucleotide having high sequence identity to a CO2 codon-optimized sequence has a reduced number of CpG dinucleotides, as compared to the wild-type GAA coding sequence SEQ ID NO: 1, as described above. Accordingly, in some embodiments, the sequence of the codon-altered polynucleotide having high sequence identity to a CO2 codon- optimized sequence has no more than 15 CpG dinucleotides.
  • the sequence of the codon-altered polynucleotide having high sequence identity to a CO2 codon- optimized sequence has no more than 10 CpG dinucleotides. In some embodiments, the sequence of the codon-altered polynucleotide having high sequence identity to a CO2 codon- optimized sequence has no more than 5 CpG dinucleotides. In some embodiments, the sequence of the codon-altered polynucleotide having high sequence identity to a CO2 codon- optimized sequence has no more than 4 CpG dinucleotides.
  • the sequence of the codon-altered polynucleotide having high sequence identity to a CO2 codon- optimized sequence has no more than 3 CpG dinucleotides. In some embodiments, the sequence of the codon-altered polynucleotide having high sequence identity to a CO2 codon- optimized sequence has no more than 2 CpG dinucleotides. In some embodiments, the sequence of the codon-altered polynucleotide having high sequence identity to a CO2 codon- optimized sequence has no more than 1 CpG dinucleotide. In some embodiments, the sequence of the codon-altered polynucleotide having high sequence identity to a CO2 codon- optimized sequence has no CpG dinucleotides.
  • a nucleic acid composition provided herein includes a GAA polynucleotide (e.g., a codon-altered polynucleotide) encoding a GAA polypeptide, where the GAA polynucleotide includes a nucleotide sequence having high sequence identity to all or a portion of the CO3 codon-optimized sequence.
  • GAA polynucleotide e.g., a codon-altered polynucleotide
  • the GAA polynucleotide includes a sequence having high sequence identity to the portion of the CO3 codon-optimized sequence that encodes for the mature GAA polypeptide. Accordingly, in some embodiments, the sequence of the codon- altered polynucleotide has at least 95% identity to CO3-MP-WT-NA (SEQ ID NO:34). In a specific embodiment, the sequence of the codon-altered polynucleotide has at least 96% identity to CO3-MP-WT-NA (SEQ ID NO:34).
  • the sequence of the codon-altered polynucleotide has at least 97% identity to CO3-MP-WT-NA (SEQ ID NO:34). In a specific embodiment, the sequence of the codon-altered polynucleotide has at least 98% identity to CO3-MP-WT-NA (SEQ ID NO:34). In a specific embodiment, the sequence of the codon-altered polynucleotide has at least 99% identity to CO3-MP-WT-NA (SEQ ID NO:34). In a specific embodiment, the sequence of the codon-altered polynucleotide has at least 99.5% identity to CO3-MP-WT-NA (SEQ ID NO:34).
  • the sequence of the codon-altered polynucleotide has at least 99.9% identity to CO3-MP-WT- NA (SEQ ID NO:34). In another specific embodiment, the sequence of the codon-altered polynucleotide is C03-MP-WT-NA (SEQ ID NO:34).
  • the GAA polynucleotide may also encode for a signal peptide, a pro-peptide, and/or a purification/detection tag, but the sequence comparison should not include these sequences.
  • a GAA polynucleotide having high sequence identity to (203- MP-WT-NA (SEQ ID NO:34) further includes a polynucleotide sequence encoding a GAA signal peptide having the amino acid sequence of SP-WT-AA (SEQ ID NO:43).
  • the GAA signal polynucleotide has a nucleic acid sequence that is at least 90%, 95%, 96%, 97%, 98%, or 100% identical to CO3-SP-WT-NA (SEQ ID NO:42).
  • a GAA polynucleotide having high sequence identity to (203- MP-WT-NA (SEQ ID NO:34) further includes a polynucleotide sequence encoding a GAA propeptide having the amino acid sequence of PP-WT-AA (SEQ ID NO:39).
  • the GAA pro-peptide polynucleotide has a nucleic acid sequence that is at least 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to CO3-PP-WT-NA (SEQ ID NO:38).
  • the GAA pro-peptide polynucleotide has a nucleic acid sequence that is at least 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to CO3-PP-46-NA (SEQ ID NO:40).
  • a GAA polynucleotide having high sequence identity to (203- MP-WT-NA (SEQ ID NO:34) further includes a polynucleotide sequence encoding a GAA propeptide having the amino acid sequence of PP-46-AA (SEQ ID NO:41).
  • the GAA pro-peptide polynucleotide has a nucleic acid sequence that is at least 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to CO3-PP-46-NA (SEQ ID NO:40).
  • the GAA polynucleotide includes a sequence having high sequence identity to the entirety of the CO3 codon-optimized sequence, encoding for the GAA pre-pro-polypeptide. Accordingly, in some embodiments, the sequence of the codon- altered polynucleotide has at least 95% identity to CO3-FL-WT-NA (SEQ ID NO:31). In a specific embodiment, the sequence of the codon-altered polynucleotide has at least 96% identity to CO3-FL-WT-NA (SEQ ID NO:31).
  • the sequence of the codon-altered polynucleotide has at least 97% identity to CO3-FL-WT-NA (SEQ ID NO:31). In a specific embodiment, the sequence of the codon-altered polynucleotide has at least 98% identity to CO3-FL-WT-NA (SEQ ID NO:31). In a specific embodiment, the sequence of the codon-altered polynucleotide has at least 99% identity to C03-FL-WT-NA (SEQ ID NO:31). In a specific embodiment, the sequence of the codon-altered polynucleotide has at least 99.5% identity to C03-FL-WT-NA (SEQ ID NO:31).
  • sequence of the codon-altered polynucleotide has at least 99.9% identity to C03-FL-WT-NA (SEQ ID NO:31). In another specific embodiment, the sequence of the codon-altered polynucleotide is C03-FL-WT-NA (SEQ ID NO:31).
  • the GAA polypeptide encoded by a GAA polynucleotide having high sequence identity to the CO3 codon-optimized sequence includes an amino acid sequencing having high sequence identity to the human wild-type mature GAA polypeptide (MP-WT-AA; SEQ ID NO:35). Accordingly, in some embodiments, the encoded GAA polypeptide has a sequence that is at least 90% identical to MP-WT-AA (SEQ ID NO:35). In some embodiments, the encoded GAA polypeptide has a sequence that is at least 95% identical to MP-WT-AA (SEQ ID NO:35).
  • the encoded GAA polypeptide has a sequence that is at least 96% identical to MP-WT-AA (SEQ ID NO:35). In some embodiments, the encoded GAA polypeptide has a sequence that is at least 97% identical to MP-WT-AA (SEQ ID NO:35). In some embodiments, the encoded GAA polypeptide has a sequence that is at least 98% identical to MP-WT-AA (SEQ ID NO:35). In some embodiments, the encoded GAA polypeptide has a sequence that is at least 99% identical to MP-WT-AA (SEQ ID NO:35).
  • the encoded GAA polypeptide has a sequence that is at least 99.5% identical to MP-WT-AA (SEQ ID NO:35). In some embodiments, the encoded GAA polypeptide has a sequence that is at least 99.8% identical to MP-WT-AA (SEQ ID NO:35). In some embodiments, the encoded GAA polypeptide has a sequence that is at least 99.8% identical to MP-WT-AA (SEQ ID NO:35). In some embodiments, the encoded GAA polypeptide has a sequence identical to MP-WT- AA (SEQ ID NO:35).
  • the GAA polypeptide may also include a signal peptide, a pro-peptide, and/or a purification/detection tag, but the sequence comparison should not include these sequences.
  • the GAA polypeptide encoded by a GAA polynucleotide having high sequence identity to the CO3 codon-optimized sequence includes an amino acid sequencing having high sequence identity to the human wild-type GAA pre-pro-polypeptide (FL-WT-AA; SEQ ID NO:2). Accordingly, in some embodiments, the encoded GAA polypeptide has a sequence that is at least 90% identical to FL-WT-AA (SEQ ID NO:2). In some embodiments, the encoded GAA polypeptide has a sequence that is at least 95% identical to FL-WT-AA (SEQ ID NO:2).
  • the encoded GAA polypeptide has a sequence that is at least 96% identical to FL-WT-AA (SEQ ID NO:2). In some embodiments, the encoded GAA polypeptide has a sequence that is at least 97% identical to FL-WT-AA (SEQ ID NO:2). In some embodiments, the encoded GAA polypeptide has a sequence that is at least 98% identical to FL-WT-AA (SEQ ID NO:2). In some embodiments, the encoded GAA polypeptide has a sequence that is at least 99% identical to FL-WT-AA (SEQ ID NO:2). In some embodiments, the encoded GAA polypeptide has a sequence that is at least 99.5% identical to FL-WT-AA (SEQ ID NO:2).
  • the encoded GAA polypeptide has a sequence that is at least 99.8% identical to FL-WT-AA (SEQ ID NO:2). In some embodiments, the encoded GAA polypeptide has a sequence that is at least 99.8% identical to FL-WT-AA (SEQ ID NO:2). In some embodiments, the encoded GAA polypeptide has a sequence identical to FL-WT-AA (SEQ ID NO:2).
  • the GAA polypeptide may also include a purification/detection tag, but the sequence comparison should not include these sequences.
  • the GAA polypeptide encoded by a GAA polynucleotide having high sequence identity to the CO3 codon-optimized sequence includes one or more known amino acid substitutions, e.g., one or more amino acid substitutions described in U.S. Patent Application Publication No. 2021/0189365, the content of which is incorporated herein by reference in its entirety.
  • the GAA polypeptide encoded by a GAA polynucleotide having high sequence identity to the CO3 codon-optimized sequence includes one or more amino acid substitutions present in one of GAA variants 1-4 described herein.
  • the GAA polypeptide encoded by a GAA polynucleotide having high sequence identity to the CO3 codon-optimized sequence includes one or more amino acid substitutions present in one of GAA variants 6-13 described herein.
  • the GAA polypeptide encoded by a GAA polynucleotide having high sequence identity to the CO3 codon-optimized sequence includes one or more amino acid substitutions present in GAA variant 4+6: L32W, L36S, L37T, P47Q, Q58V, A70L, P86E, D95E, L117D, S143Q, T151I, T158S, T274N, R275K, A445G, T494E, E530V, N535R, L577T, L650G, L650S, L650T, L650E, L650Y, L650F, S676D, L678H, L678T, T700G, A719H, A758P, A820E, Q838K, L868F, L879E, R891H, Q902G, V921R, and S940A.
  • the sequence of the codon-altered polynucleotide has at least 95% identity to the portion of a codon-optimized GAA polynucleotide encoding the GAA variant 46 mature polypeptide of the GAA (CO3-MP-46- NA; SEQ ID NO:36).
  • the sequence of the codon-altered polynucleotide has at least 96% identity to CO3-MP-46-NA (SEQ ID NO: 36).
  • the sequence of the codon-altered polynucleotide has at least 97% identity to CO3-MP-46-NA (SEQ ID NO:36).
  • the sequence of the codon- altered polynucleotide has at least 98% identity to CO3-MP-46-NA (SEQ ID NO:36). In a specific embodiment, the sequence of the codon-altered polynucleotide has at least 99% identity to CO3-MP-46-NA (SEQ ID NO:36). In a specific embodiment, the sequence of the codon-altered polynucleotide has at least 99.5% identity to CO3-MP-46-NA (SEQ ID NO: 36). In a specific embodiment, the sequence of the codon-altered polynucleotide has at least 99.9% identity to CO3-MP-46-NA (SEQ ID NO:36).
  • the sequence of the codon-altered polynucleotide is CO3-MP-46-NA (SEQ ID NO:36).
  • SEQ ID NO:36 the sequence identity between a GAA polypeptide and the portion of the CO3 codon-optimized sequence that encodes for the mature GAA polypeptide. That is, the GAA polypeptide may also encode for a signal peptide, a pro-peptide, and/or a purification/detection tag, but the sequence comparison should not include these sequences.
  • a GAA polynucleotide having high sequence identity to CO3- MP-46-NA further includes a polynucleotide sequence encoding a GAA signal peptide having the amino acid sequence of SP-WT-AA (SEQ ID NO:43).
  • the GAA signal polynucleotide has a nucleic acid sequence that is at least 90%, 95%, 96%, 97%, 98%, or 100% identical to CO3-SP-WT-NA (SEQ ID NO:42).
  • a GAA polynucleotide having high sequence identity to CO3- MP-46-NA (SEQ ID NO: 36) further includes a polynucleotide sequence encoding a GAA propeptide having the amino acid sequence of PP-WT-AA (SEQ ID NO:39).
  • the GAA pro-peptide polynucleotide has a nucleic acid sequence that is at least 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to CO3-PP-WT-NA (SEQ ID NO:38).
  • a GAA polynucleotide having high sequence identity to CO3- MP-46-NA (SEQ ID NO: 36) further includes a polynucleotide sequence encoding a GAA propeptide having the amino acid sequence of PP-46-AA (SEQ ID NO:41).
  • the GAA pro-peptide polynucleotide has a nucleic acid sequence that is at least 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to CO3-PP-46-NA (SEQ ID NO:40).
  • the GAA polynucleotide includes a sequence having high sequence identity to the entirety of a codon-optimized GAA polynucleotide encoding the variant 46 GAA pre-pro-polypeptide. Accordingly, in some embodiments, the sequence of the codon-altered polynucleotide has at least 95% identity to CO3-FL-46-NA (SEQ ID NO:69). In a specific embodiment, the sequence of the codon-altered polynucleotide has at least 96% identity to CO3-FL-46-NA (SEQ ID NO: 69).
  • the sequence of the codon-altered polynucleotide has at least 97% identity to CO3-FL-46-NA (SEQ ID NO:69). In a specific embodiment, the sequence of the codon-altered polynucleotide has at least 98% identity to CO3-FL-46-NA (SEQ ID NO:69). In a specific embodiment, the sequence of the codon-altered polynucleotide has at least 99% identity to CO3-FL-46-NA (SEQ ID NO:69). In a specific embodiment, the sequence of the codon-altered polynucleotide has at least 99.5% identity to CO3-FL-46-NA (SEQ ID NO:69).
  • sequence of the codon-altered polynucleotide has at least 99.9% identity to CO3-FL-46- NA (SEQ ID NO:69). In another specific embodiment, the sequence of the codon-altered polynucleotide is CO3-FL-46-NA (SEQ ID NO:69).
  • the GAA polypeptide encoded by a GAA polynucleotide having high sequence identity to the CO3 codon-optimized sequence includes an amino acid sequencing having high sequence identity to the human variant 4+6 mature GAA polypeptide (MP-46-AA; SEQ ID NO: 37). Accordingly, in some embodiments, the encoded GAA polypeptide has a sequence that is at least 90% identical to MP-46-AA (SEQ ID NO:37). In some embodiments, the encoded GAA polypeptide has a sequence that is at least 95% identical to MP-46-AA (SEQ ID NO:37).
  • the encoded GAA polypeptide has a sequence that is at least 96% identical to MP-46-AA (SEQ ID NO:37). In some embodiments, the encoded GAA polypeptide has a sequence that is at least 97% identical to MP-46-AA (SEQ ID NO:37). In some embodiments, the encoded GAA polypeptide has a sequence that is at least 98% identical to MP-46-AA (SEQ ID NO:37). In some embodiments, the encoded GAA polypeptide has a sequence that is at least 99% identical to MP-46-AA (SEQ ID NO:37). In some embodiments, the encoded GAA polypeptide has a sequence that is at least 99.5% identical to MP-46-AA (SEQ ID NO:37).
  • the encoded GAA polypeptide has a sequence that is at least 99.8% identical to MP-46-AA (SEQ ID NO:37). In some embodiments, the encoded GAA polypeptide has a sequence that is at least 99.8% identical to MP-46-AA (SEQ ID NO:37). In some embodiments, the encoded GAA polypeptide has a sequence identical to MP-46-AA (SEQ ID NO:37).
  • the GAA polypeptide may also include a signal peptide, a pro-peptide, and/or a purification/detection tag, but the sequence comparison should not include these sequences.
  • the GAA polypeptide encoded by a GAA polynucleotide having high sequence identity to the CO3 codon-optimized sequence includes an amino acid sequencing having high sequence identity to the human variant 4+6 GAA pre-pro- polypeptide (FL-46-AA; SEQ ID NO:30). Accordingly, in some embodiments, the encoded GAA polypeptide has a sequence that is at least 90% identical to FL-46- AA (SEQ ID NO:30). In some embodiments, the encoded GAA polypeptide has a sequence that is at least 95% identical to FL-46-AA (SEQ ID NO: 30).
  • the encoded GAA polypeptide has a sequence that is at least 96% identical to FL-46-AA (SEQ ID NO:30). In some embodiments, the encoded GAA polypeptide has a sequence that is at least 97% identical to FL-46-AA (SEQ ID NO:30). In some embodiments, the encoded GAA polypeptide has a sequence that is at least 98% identical to FL-46-AA (SEQ ID NO:30). In some embodiments, the encoded GAA polypeptide has a sequence that is at least 99% identical to FL-46-AA (SEQ ID NO:30). In some embodiments, the encoded GAA polypeptide has a sequence that is at least 99.5% identical to FL-46-AA (SEQ ID NO:30).
  • the encoded GAA polypeptide has a sequence that is at least 99.8% identical to FL-46-AA (SEQ ID NO:30). In some embodiments, the encoded GAA polypeptide has a sequence that is at least 99.8% identical to FL-46-AA (SEQ ID NO:30). In some embodiments, the encoded GAA polypeptide has a sequence identical to FL-46-AA (SEQ ID NO:30). When determining the sequence identity between an encoded GAA polypeptide and the GAA pre-pro-polypeptide, only the portions of the sequence corresponding to the pre-pro-polypeptide should be considered. That is, the GAA polypeptide may also include a purification/detection tag, but the sequence comparison should not include these sequences.
  • the nucleotide sequence of the GAA polynucleotide having high sequence identity to a CO3 codon-optimized sequence (e.g., SEQ ID NO:31, 34, 69, or 36) has a reduced GC content, as compared to the wild-type GAA coding sequence SEQ ID NO: 1, as described above. Accordingly, in some embodiments, the sequence of the codon- altered polynucleotide having high sequence identity to a CO3 codon-optimized sequence has a GC content of no more than 60%.
  • the sequence of the codon-altered polynucleotide having high sequence identity to a CO3 codon-optimized sequence has a GC content of no more than 57.5%. In some embodiments, the sequence of the codon-altered polynucleotide having high sequence identity to a CO3 codon-optimized sequence has a GC content of no more than 59%, no more than 58%, no more than 57%, no more than 56%, or no more than 55%.
  • the sequence of the codon-altered polynucleotide having high sequence identity to a CO3 codon-optimized sequence has a GC content of from 55% to 60%. In some embodiments, the sequence of the codon-altered polynucleotide having high sequence identity to a CO3 codon-optimized sequence has a GC content of from 56% to 60%, from 57% to 60%, from 58% to 60%, from 59% to 60%, from 55% to 59%, from 56% to 59%, from 57% to 59%, from 58% to 59%, from 55% to 58%, from 56% to 58%, from 57% to 58%, from 55% to 57%, from 56% to 57%, or from 55% to 56%.
  • the sequence of the codon-altered polynucleotide having high sequence identity to a CO3 codon-optimized sequence has a GC content of 57.5% ⁇ 1.0. In some embodiments, the sequence of the codon-altered polynucleotide having high sequence identity to a CO3 codon-optimized sequence has a GC content of 57.5% ⁇ 0.8. In some embodiments, the sequence of the codon-altered polynucleotide having high sequence identity to a CO3 codon-optimized sequence has a GC content of 57.5% ⁇ 0.6.
  • the sequence of the codon-altered polynucleotide having high sequence identity to a CO3 codon-optimized sequence has a GC content of 57.5% ⁇ 0.5. In some embodiments, the sequence of the codon-altered polynucleotide having high sequence identity to a CO3 codon-optimized sequence has a GC content of 57.5% ⁇ 0.4. In some embodiments, the sequence of the codon-altered polynucleotide having high sequence identity to a CO3 codon-optimized sequence has a GC content of 57.5% ⁇ 0.3.
  • the sequence of the codon-altered polynucleotide having high sequence identity to a CO3 codon-optimized sequence has a GC content of 57.5% ⁇ 0.2. In some embodiments, the sequence of the codon-altered polynucleotide having high sequence identity to a CO3 codon-optimized sequence has a GC content of 57.5% ⁇ 0.1. In some embodiments, the sequence of the codon-altered polynucleotide having high sequence identity to a CO3 codon-optimized sequence has a GC content of 57.5%.
  • the nucleotide sequence of the GAA polynucleotide having high sequence identity to a CO3 codon-optimized sequence has a reduced number of CpG dinucleotides, as compared to the wild-type GAA coding sequence SEQ ID NO: 1, as described above. Accordingly, in some embodiments, the sequence of the codon-altered polynucleotide having high sequence identity to a CO3 codon- optimized sequence has no more than 15 CpG dinucleotides.
  • the sequence of the codon-altered polynucleotide having high sequence identity to a CO3 codon- optimized sequence has no more than 10 CpG dinucleotides. In some embodiments, the sequence of the codon-altered polynucleotide having high sequence identity to a CO3 codon- optimized sequence has no more than 5 CpG dinucleotides. In some embodiments, the sequence of the codon-altered polynucleotide having high sequence identity to a CO3 codon- optimized sequence has no more than 4 CpG dinucleotides.
  • the sequence of the codon-altered polynucleotide having high sequence identity to a CO3 codon- optimized sequence has no more than 3 CpG dinucleotides. In some embodiments, the sequence of the codon-altered polynucleotide having high sequence identity to a CO3 codon- optimized sequence has no more than 2 CpG dinucleotides. In some embodiments, the sequence of the codon-altered polynucleotide having high sequence identity to a CO3 codon- optimized sequence has no more than 1 CpG dinucleotide. In some embodiments, the sequence of the codon-altered polynucleotide having high sequence identity to a CO3 codon- optimized sequence has no CpG dinucleotides.
  • an expression cassette for expressing a GAA polynucleotide as disclosed herein, e.g., a codon-altered GAA polynucleotide.
  • an expression cassette comprises one or more nucleic acids encoding a GAA protein and at least one regulatory nucleic acid sequence operably linked to the sequence encoding the GAA protein.
  • the at least one regulatory nucleic acid sequence is selected from the group consisting of a promoter, an enhancer, an intron, a post- transcriptional regulatory element, an inverted terminal repeat (ITR), a polyadenylation (poly A) sequence, and a combination thereof.
  • the at least one regulatory nucleic acid sequence comprises a promoter.
  • the promoter is a muscle-specific promoter.
  • the muscle-specific promoter comprises a polynucleotide sequence that is at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to SPc512_NA (SEQ ID NO:46).
  • the muscle-specific promoter comprises the polynucleotide sequence of SPc512_NA (SEQ ID NO:46).
  • the muscle-specific promoter comprises a polynucleotide sequence that is at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to HsDesmin_NA (SEQ ID NO:47). In some embodiments, the muscle-specific promoter comprises the polynucleotide sequence of HsDesmin NA (SEQ ID NO:47).
  • the at least one regulatory nucleic acid sequence comprises an enhancer.
  • the enhancer is a muscle-specific enhancer.
  • the muscle-specific enhancer comprises a polynucleotide sequence that is at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to Dph- CRE04 NA (SEQ ID NO:48).
  • the muscle-specific enhancer comprises the polynucleotide sequence of Dph-CRE04_NA (SEQ ID NO:48).
  • the muscle-specific enhancer comprises a polynucleotide sequence that is at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to Csk-SH4_NA (SEQ ID NO:49). In some embodiments, the muscle-specific enhancer comprises the polynucleotide sequence of sk-SH4_NA (SEQ ID NO:49).
  • the at least one regulatory nucleic acid sequence comprises an intron.
  • the intron comprises a polynucleotide sequence that is at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to MVM_NA (SEQ ID NO:50).
  • the intron comprises the polynucleotide sequence of MVM NA (SEQ ID NO:50).
  • the codon-altered polynucleotides and associated expression cassettes described herein are integrated into expression vectors.
  • expression vectors include viral vectors (e.g., vectors suitable for gene therapy), plasmid vectors, bacteriophage vectors, cosmids, phagemids, artificial chromosomes, and the like.
  • the gene therapy vector is an adeno-associated virus (AAV) based gene therapy vector.
  • AAV systems have been described previously and are generally well known in the art (Kelleher and Vos, Biotechniques, 17(6): 1110-17 (1994); Cotten et al., P.N.A.S. U.S.A., 89(13):6094-98 (1992); Curiel, Nat Immun, 13(2-3): 141-64 (1994);
  • the expression cassette is a mammalian expression vector.
  • the mammalian expression vector comprises an adeno-associated virus (AAV) vector.
  • the AAV vector comprises an AAV8 or AAV9 capsid polypeptide encapsidating the expression cassette.
  • the AAV vector comprises an AAV vector (e.g., AAVMYO or myoAAV) engineered with enhanced tropism to muscle (El Andari et al., Sci Adv 8(38):eabn4704 (2022); Tabebordbar Cell 184(19): 4919-4938 e22 (2021).
  • the codon-altered polynucleotides described herein are integrated into a viral gene therapy vector.
  • viral vectors include: retrovirus, e.g., Moloney murine leukemia virus (MMLV), Harvey murine sarcoma virus, murine mammary tumor virus, and Rous sarcoma virus; adenoviruses, adeno-associated viruses; SV40-type viruses; polyomaviruses; Epstein-Barr viruses; papilloma viruses; herpes viruses; vaccinia viruses; and polio viruses.
  • retrovirus e.g., Moloney murine leukemia virus (MMLV), Harvey murine sarcoma virus, murine mammary tumor virus, and Rous sarcoma virus
  • adenoviruses adeno-associated viruses
  • SV40-type viruses polyomaviruses
  • Epstein-Barr viruses Epstein-Barr viruses
  • papilloma viruses herpes viruses
  • the gene therapy vector is a retrovirus, and particularly a replication-deficient retrovirus.
  • Protocols for the production of replication-deficient retroviruses are known in the art. For review, see Kriegler, M., Gene Transfer and Expression, A Laboratory Manual, W.H. Freeman Co., New York (1990) and Murry, E. J., Methods in Molecular Biology, Vol. 7, Humana Press, Inc., Cliffton, N.J. (1991).
  • the codon-altered polynucleotides described herein are integrated into a retroviral expression vector.
  • These systems have been described previously, and are generally well known in the art (Mann et al., Cell, 33: 153-159, 1983; Nicolas and Rubinstein, In: Vectors: A survey of molecular cloning vectors and their uses, Rodriguez and Denhardt, eds., Stoneham: Butterworth, pp. 494-513, 1988; Temin, In: Gene Transfer, Kucherlapati (ed.), New York: Plenum Press, pp. 149-188, 1986).
  • the retroviral vector is a lentiviral vector (see, for example, Naldini et al., Science, 272(5259):263-267 , 1996; Zufferey et al., Nat Biotechnol, 15(9):871-875, 1997; Blomer et al., J Virol., 71(9):6641-6649, 1997; U.S. Pat. Nos. 6,013,516 and 5,994,136).
  • the codon-altered polynucleotides described herein can be administered to a subject by a non-viral method.
  • naked DNA can be administered into a cell by electroporation, sonoporation, particle bombarment, or hydrodyamic delivery.
  • DNA can also be encapsulated or coupled with polymers, e.g., liposomes, polysomes, polypleses, dendrimers, and administered to the subject as a complex.
  • DNA can be coupled to inorganic nanoparticles, e.g., gold, silica, iron oxide, or calcium phosphate particles, or attached to cell-penetrating peptides for delivery to cells in vivo.
  • Codon-altered GAA coding polynucleotides can also be incorporated into artificial chromosomes, such as Artificial Chromosome Expression (ACEs) (see, e.g., Lindenbaum et al., Nucleic Acids Res., 32(21):el72 (2004)) and mammalian artificial chromosomes (MACs).
  • ACEs Artificial Chromosome Expression
  • MACs mammalian artificial chromosomes
  • a wide variety of vectors can be used for the expression of a GAA polypeptide from a codon-altered polypeptide in cell culture, including eukaryotic and prokaryotic expression vectors.
  • a plasmid vector is contemplated for use in expressing a GAA polypeptide in cell culture.
  • plasmid vectors containing replicon and control sequences which are derived from species compatible with the host cell are used in connection with these hosts.
  • the vector can carry a replication site, as well as marking sequences which are capable of providing phenotypic selection in transformed cells.
  • the plasmid will include the codon-altered polynucleotide encoding the GAA polypeptide, operably linked to one or more control sequences, for example, a promoter.
  • Non-limiting examples of vectors for prokaryotic expression include plasmids such as pRSET, pET, pBAD, etc., wherein the promoters used in prokaryotic expression vectors include lac, trc, trp, recA, araBAD, etc.
  • vectors for eukaryotic expression include: (i) for expression in yeast, vectors such as pAO, pPIC, pYES, pMET, using promoters such as A0X1, GAP, GALI, AUG1, etc; (ii) for expression in insect cells, vectors such as pMT, pAc5, pIB, pMIB, pBAC, etc., using promoters such as PH, plO, MT, Ac5, OpIE2, gp64, polh, etc., and (iii) for expression in mammalian cells, vectors such as pSVL, pCMV, pRc/RSV, pcDNA3, pBPV, etc., and vectors derived from viral systems such as vaccinia virus, adeno-associated viruses, herpes viruses, retroviruses, etc., using promoters such as CMV, SV40, EF-1, UbC, RSV, ADV, BPV, and
  • the disclosure provides an AAV gene therapy vector that includes a codon-altered GAA polynucleotide, as described herein, internal terminal repeat (ITR) sequences on the 5’ and 3’ ends of the vector, one or more promoter and/or enhancer sequences operably linked to the GAA polynucleotide, and a poly-adenylation signal following the 3’ end of the GAA polynucleotide sequence.
  • ITR internal terminal repeat
  • promoter and/or enhancer sequences include one or more copies of a muscle-specific regulatory control element.
  • the codon-altered GAA polynucleotides and viral vectors described herein are produced according to conventional methods for nucleic acid amplification and vector production.
  • Two predominant platforms have developed for large-scale production of recombinant AAV vectors. The first platform is based on replication in mammalian cells, while the second is based on replication in invertebrate cells.
  • the first platform is based on replication in mammalian cells, while the second is based on replication in invertebrate cells.
  • the disclosure provides methods for producing an adeno-associated virus (AAV) particle.
  • the methods include introducing a codon- altered GAA polynucleotide construct having high nucleotide sequence identity (e.g., at least 95%, 96%, 97%, 98%, 99%, 99.5%, 99.9%, or 100%) to one of a CO1, CO2, or CO3 sequence, as described herein, into a host cell where the polynucleotide construct is competent for replication in the host cell.
  • the host cell is a mammalian host cell e.g., an HEK, CHO, or BHK cell. In a specific embodiment, the host cell is an HEK 293 cell. In some embodiments, the host cell is an invertebrate cell, e.g, an insect cell. In a specific embodiment, the host cell is an SF9 cell.
  • the present disclosure provides expression constructs such as helper plasmids (e.g., non- AAV expression constructs) comprising a nucleic acid that encodes one or more of the AAV capsid polypeptides described herein.
  • helper plasmids e.g., non- AAV expression constructs
  • Such plasmids are useful as expression constructs for producing AAV capsid polypeptides or proteins or to transfect cells (e.g., as part of a triple transfection) in the preparation of engineered AAV vectors.
  • AAV vectors could be produced using herpes virus, baculovirus, stable genetically engineered cell lines, or any other method known in the art (Dobrowsky et al. (2021) Curr. Opinion Biomed. Engin. 20: 100353, the disclosure of which is hereby incorporated herein by reference in its entirety).
  • the capsid helper plasmid may comprise one or more nucleic acid sequences to regulate expression of the AAV capsid polypeptide.
  • the sequences include but are not limited to, a promoter, an enhancer, an intron, a post-transcriptional regulatory sequence, a polyadenylation (poly A) signal, or any combination thereof, which are operably linked to the nucleic acid sequences that encode the AAV capsid polypeptide.
  • the promoter may be a heterologous promoter, a tissue-specific promoter, a cellspecific promoter, a constitutive promoter, an inducible promoter, a hybrid promoter, or any combination thereof.
  • the capsid helper plasmid of the present disclosure comprises at least one promoter capable of expressing, or directed to primarily express, the nucleic acid segment in a suitable host cell e.g., a muscle cell) into which the engineered capsid helper plasmid can be transfected.
  • Exemplary promoters include, but are not limited to, a ubiquitous promoter, a CMV promoter, a P-actin promoter, a muscle-specific promoter, a Desmin promoter, an SPc5-12 promoter, an MCK-based promoter an insulin promoter, an enolase promoter, a BDNF promoter, an NGF promoter, an EGF promoter, a growth factor promoter, an axon-specific promoter, a dendrite-specific promoter, a brain-specific promoter, a hippocampal-specific promoter, a kidney-specific promoter, a retinal- specific promoter, an elafin promoter, a cytokine promoter, an interferon promoter, a growth factor promoter, an al -antitrypsin promoter, a brain cell-specific promoter, a neural cell- specific promoter, a central nervous system cell-specific promoter, a peripheral nervous system cell-specific promoter, an interleukin promoter,
  • Exemplary enhancer sequences include, but are not limited to, one or more selected from the group consisting of a CMV enhancer, a muscle-specific enhancer, a synthetic enhancer, a liver-specific enhancer, a vascular-specific enhancer, a brain-specific enhancer, a neural cell-specific enhancer, a lung-specific enhancer, a kidney-specific enhancer, a pancreas-specific enhancer, retinal-specific enhancer, and an islet cell-specific enhancer.
  • Exemplary post-transcriptional regulatory sequences include a woodchuck hepatitis post-transcription regulatory element (WPRE)), one or more ribosome entry sites (IRES), one or more polyadenylation (poly A) signal sequences, or any combination thereof.
  • WPRE woodchuck hepatitis post-transcription regulatory element
  • IVS ribosome entry sites
  • poly A polyadenylation
  • a polyA signal may be an artificial polyA.
  • suitable polyA sequences include, e.g., bovine growth hormone, SV40, rabbit beta globin, and TK polyA, amongst others.
  • the capsid helper plasmid described herein may contain other appropriate transcription initiation, termination, and efficient RNA processing signals.
  • Such sequences include splicing, inducible expression control elements, regulatory elements that enhance expression, sequences that stabilize cytoplasmic mRNA, sequences that enhance translation efficiency (e.g., Kozak consensus sequence), sequences that enhance protein stability, and when desired, sequences that enhance secretion of the encoded product.
  • a Kozak sequence is included.
  • the present disclosure provides GAA polypeptide variants that have advantageous properties relative to wild-type GAA polypeptides.
  • GAA variants 1-4 a first series of variant GAA polypeptides (GAA variants 1-4) were identified having specific activities that are approximately 50% to 250% higher than the human wild-type GAA polypeptide. Further, these variants demonstrated increased myoblast uptake and improved stability following plasma challenge.
  • the amino acid sequences for GAA variants 1-4, as well as non-codon altered polynucleotides encoding the same, are provided in FIG. 19B-FIG. 19E.
  • GAA variants 6-13 a second series of variant GAA polypeptides (GAA variants 6-13) were identified, several of which demonstrated significantly increased catalytic activity relative to the human wild-type GAA polypeptide.
  • GAA variant 4 demonstrated approximately 5.5-fold higher activity than the human wild-type GAA polypeptide, as shown in FIG. 9, and improved kinetic parameters, as shown in FIG. 10A.
  • the amino acid sequences for GAA variants 6-13, as well as non-codon altered polynucleotides encoding the same, are provided in FIG 19F-FIG. 19M.
  • GAA variant 4+6 a second generation GAA variant polypeptide, referred to herein as GAA variant 4+6, was constructed by combining the amino acid substitutions in the GAA variant 4 polypeptide and the GAA variant 6 polypeptide: L32W, L36S, L37T, P47Q, Q58V, A70L, P86E, D95E, L117D, S143Q, T151I, T158S, T274N, R275K, A445G, T494E, E530V, N535R, L577T, L650G, L650S, L650T, L650E, L650Y, L650F, S676D, L678H, L678T, T700G, A719H, A758P, A820E, Q838K, L868F, L879E, R891H, Q902G, V921R, and S940A.
  • FIG. 11A-FIG. 11B GAA variant 4+
  • the disclosure provides variant GAA polypeptides having high sequence identity to the variant 4+6 GAA pre-pro-polypeptide (GAA-FL-46- AA; SEQ ID NO:30) and/or the variant 4+6 GAA mature polypeptide (GAA-MP-46-AA; SEQ ID NO:37).
  • the GAA variant protein comprises a first polypeptide sequence that is at least 95%, at least 96%, at least 97% or at least 98% identical to MP-46-AA (SEQ ID NO:37).
  • the first polypeptide sequence is at least 99% identical to MP-46-AA (SEQ ID NO:37).
  • the first polypeptide sequence is at least 99.5% identical to MP-46-AA (SEQ ID NO:37).
  • the first polypeptide sequence is MP-46-AA (SEQ ID NO:37).
  • the GAA variant protein comprises a first polypeptide sequence that is at least 95%, at least 96%, at least 97% or at least 98% identical to MP-46- AA (SEQ ID NO:37) and comprises one or more variant amino acids selected from the group consisting of L32W, L36S, L37T, P47Q, Q58V, A70L, P86E, D95E, L117D, S143Q, T151I, T158S, T274N, R275K, A445G, T494E, E530V, N535R, L577T, L650G, S676D, L678H, T700G, A719H, A758P, A820E, Q838K, L868F, L879E, R891H, Q902G, V921R, and S940A.
  • the first polypeptide sequence is at least 99% identical to MP-46-AA (SEQ ID NO:37
  • the GAA variant protein further comprises a second polypeptide sequence that is at least 95% identical to PP-WT-AA (SEQ ID NO:39). In some embodiments, the second polypeptide sequence is at least 97% identical to PP-WT-AA (SEQ ID NO:39). In some embodiments, the second polypeptide sequence is PP-WT-AA (SEQ ID NO: 39). [00408] In some embodiments, the recombinant GAA variant protein further comprises a third polypeptide sequence that is at least 95% identical to SP-WT-AA (SEQ ID NO:43). In some embodiments, the third polypeptide sequence is SP-WT-AA (SEQ ID NO:43). In some embodiments, the recombinant GAA variant protein comprises the polypeptide sequence of FL-46-AA (SEQ ID NO:33).
  • a GAA variant protein comprises an amino acid substitution selected from the group consisting of L117D, T151I, L313V, L650G, L650S, L650T, L650E, L650Y, L650F, S676D, L678H, and L868F, numbered relative to the full- length wild-type GAA protein sequence of FL-WT-AA (SEQ ID NO:2).
  • a GAA variant protein comprises an amino acid substitution selected from the group consisting of L32W, L36S, L37T, P47Q, Q58V, A70L, P86E, D95E, S143Q, T158S, T274N, R275K, A445G, T494E, E530V, N535R, L577T, L678T, T700G, A719H, A758P, A820E, Q838K, L868F, L879E, R891H, Q902G, V921R, and S940A, numbered relative to the full-length wild-type GAA protein sequence of FL-WT- AA (SEQ ID NO:2).
  • a GAA variant protein comprises an amino acid substitution selected from the group consisting of T 15 II, L650G, S676D, and L678H, numbered relative to the full-length wild-type GAA protein sequence of FL-WT-AA (SEQ ID N0:2).
  • the recombinant GAA variant protein comprises a tryptophan residue at position 32 (relative to SEQ ID NO:2). In some embodiments, the recombinant GAA variant protein comprises a serine residue at position 36 (relative to SEQ ID NO:2). In some embodiments, the recombinant GAA variant protein comprises a threonine residue at position 37 (relative to SEQ ID NO:2). In some embodiments, the recombinant GAA variant protein comprises a glutamine residue at position 47 (relative to SEQ ID NO:2). In some embodiments, the recombinant GAA variant protein comprises a valine residue at position 58 (relative to SEQ ID NO:2).
  • the recombinant GAA variant protein comprises a leucine residue at position 70 (relative to SEQ ID NO:2). In some embodiments, the recombinant GAA variant protein comprises a glutamic acid residue at position 86 (relative to SEQ ID NO:2). In some embodiments, the recombinant GAA variant protein comprises a glutamic acid residue at position 95 (relative to SEQ ID NO:2). In some embodiments, the recombinant GAA variant protein comprises an aspartic acid residue at position 117 (relative to SEQ ID NO:2). In some embodiments, the recombinant GAA variant protein comprises a glutamine residue at position 143 (relative to SEQ ID NO:2).
  • the recombinant GAA variant protein comprises an isoleucine residue at position 151 (relative to SEQ ID NO:2). In some embodiments, the recombinant GAA variant protein comprises a serine residue at position 158 (relative to SEQ ID NO:2). In some embodiments, the recombinant GAA variant protein comprises an asparagine residue at position 274 (relative to SEQ ID NO:2). In some embodiments, the recombinant GAA variant protein comprises a lysine residue at position 275 (relative to SEQ ID NO:2). In some embodiments, the recombinant GAA variant protein comprises a glycine residue at position 445 (relative to SEQ ID NO:2).
  • the recombinant GAA variant protein comprises a glutamic acid residue at position 494 (relative to SEQ ID NO:2). In some embodiments, the recombinant GAA variant protein comprises a valine residue at position 530 (relative to SEQ ID NO:2). In some embodiments, the recombinant GAA variant protein comprises an arginine residue at position 535 (relative to SEQ ID NO:2). In some embodiments, the recombinant GAA variant protein comprises a threonine residue at position 577 (relative to SEQ ID NO:2). In some embodiments, the recombinant GAA variant protein comprises a glycine residue at position 650 (relative to SEQ ID NO:2).
  • the recombinant GAA variant protein comprises an aspartic acid residue at position 676 (relative to SEQ ID NO:2). In some embodiments, the recombinant GAA variant protein comprises a histidine residue at position 678 (relative to SEQ ID NO:2). In some embodiments, the recombinant GAA variant protein comprises a glycine residue at position 700 (relative to SEQ ID NO:2). In some embodiments, the recombinant GAA variant protein comprises a histidine residue at position 719 (relative to SEQ ID NO:2). In some embodiments, the recombinant GAA variant protein comprises a proline residue at position 758 (relative to SEQ ID NO:2).
  • the recombinant GAA variant protein comprises a glutamic acid residue at position 820 (relative to SEQ ID NO:2). In some embodiments, the recombinant GAA variant protein comprises a lysine residue at position 838 (relative to SEQ ID NO:2). In some embodiments, the recombinant GAA variant protein comprises a phenylalanine residue at position 868 (relative to SEQ ID NO:2). In some embodiments, the recombinant GAA variant protein comprises a glutamic acid residue at position 879 (relative to SEQ ID NO:2). In some embodiments, the recombinant GAA variant protein comprises a histidine residue at position 891 (relative to SEQ ID NO:2).
  • the recombinant GAA variant protein comprises a glycine residue at position 902 (relative to SEQ ID NO:2). In some embodiments, the recombinant GAA variant protein comprises an arginine residue at position 921 (relative to SEQ ID NO:2). In some embodiments, the recombinant GAA variant protein comprises an alanine residue at position 940 (relative to SEQ ID NO:2).
  • the recombinant GAA variant protein further comprises a second polypeptide sequence that is at least 95% identical to PP-WT-AA (SEQ ID NO:39).
  • the first polypeptide sequence is MP-46-AA (SEQ ID NO:37).
  • the second polypeptide sequence is at least 97% identical to PP-WT-AA (SEQ ID NO:39).
  • the second polypeptide sequence is PP-WT-AA (SEQ ID NO: 39).
  • the recombinant GAA variant protein comprises the polypeptide sequence of FL-46-AA (SEQ ID NO:33).
  • the recombinant GAA variant protein a set of amino acid substitutions, numbered relative to the full-length wild-type GAA protein sequence of FL-WT-AA (SEQ ID NO:2), selected from the group consisting of: a) T151I, L650G, S676D, and L678H, b) L650S, S676D, and L678H, c) L650T, S676D, and L678H, d) L650E, S676D, and L678H, e) L650Y, S676D, and L678H, f) L650F, S676D, and L678H, g) L650G, S676D, and L678H, and h) S676D, and L678H.
  • the recombinant GAA variant protein comprises a first polypeptide sequence that is at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% identical to amino acid residues 70-952 of FL-6-AA (SEQ ID NO: 14).
  • the recombinant GAA variant protein comprises a first polypeptide sequence that is at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% identical to amino acid residues 70-952 of FL-7-AA (SEQ ID NO: 16).
  • the recombinant GAA variant protein comprises a first polypeptide sequence that is at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% identical to amino acid residues 70-952 of FL-8-AA (SEQ ID NO: 18).
  • the recombinant GAA variant protein comprises a first polypeptide sequence that is at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% identical to amino acid residues 70-952 of FL-9-AA (SEQ ID NO:20).
  • the recombinant GAA variant protein comprises a first polypeptide sequence that is at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% identical to amino acid residues 70-952 of FL-10-AA (SEQ ID NO:22).
  • the recombinant GAA variant protein comprises a first polypeptide sequence that is at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% identical to amino acid residues 70-952 of FL-11-AA (SEQ ID NO:24).
  • the recombinant GAA variant protein comprises a first polypeptide sequence that is at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% identical to amino acid residues 70-952 of FL-12-AA (SEQ ID NO:26).
  • the recombinant GAA variant protein comprises a first polypeptide sequence that is at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% identical to amino acid residues 70-952 of FL-13-AA (SEQ ID NO:28).
  • the recombinant GAA variant protein further comprises a second polypeptide sequence that is at least 95%, at least 97%, or 100% identical to PP-WT- AA (SEQ ID NO: 39).
  • FIG. 19A-FIG. 30B provide the nucleotide sequences of various components of the vectors tested herein.
  • GAA KO a suitable animal model of Pompe disease. This model was generated by insertion of a neomycin cassette into exon 6 of the mouse gaa gene, thereby creating a functional knockout (KO) of the gaa genes. GAA KO recapitulates critical features of both the infantile and the adult forms of PD at a pace suitable for the evaluation of gene therapy (GT).
  • GT gene therapy
  • Glycogen accumulation in cardiac and skeletal muscles can be detected as early as 3 weeks of age, resembling IOPD, and reduction in the number of myofibrils and signs of damaged muscle structure, impaired autophagic flux in skeletal muscle, mild cardiac defects and muscle weakness leading to locomotor defects which develop by 8-9 months all resemble LOPD (reviewed in Geel et al. (2007) Mol. Genet. Metab. 92(4):299-307).
  • LOPD Reviewed in Geel et al. (2007) Mol. Genet. Metab. 92(4):299-307.
  • these secondary tissues will be characterized and monitored in the mice experiments to determine if this disease phenotype is replicated in the mouse model.
  • AAV vector preparations comprising test AAV transgene constructs or controls were prepared for injection by dilution in vehicle (1.5 mM KH2PO4, 2.7mM KCI, 8.1 mM Na2HPO4, 136.9mM NaCl, 0.001% Pluronic F-68), and doses were administered through a single intravenous administration of vector or buffer only into the tail vein at 3xl0 12 vg/kg, IxlO 13 vg/kg or 3xl0 13 vg/kg, as indicated. Clinical and mortality observations were conducted daily post dosing until the end of each study.
  • mice were anesthetized with isoflurane, euthanized, and necropsied. All animals were perfused with 0.9% saline and total wet weight of whole heart, brain, quadriceps, triceps, and gastrocnemius was recorded. Each tissue was further dissected into 20-30 mg pieces. Samples were transferred to separate ceramic Precellys bead tubes (P000916-LYSK0-A; Bertin, Rockville, MD) and snap-frozen in liquid nitrogen. Samples were transferred to dry ice and stored at -80°C.
  • Vector Copy Number DNA was extracted and purified from tissue homogenates using a MagMAX kit (Thermofisher, Los Angeles, California) according to manufacturer instructions. Vector copy number was quantified using a digital polymerase chain reaction (dPCR) quantification assay with primers and probes designed on a proprietary DNA sequence, using a linearized vector plasmid as the reference standard. Each 12 ml dPCR reaction contained 2-200 ng sample genomic DNA (gDNA) that was run using a Qiacuity 4 instrument (Qiagen, Germantown MD). Vector genome copy numbers (VGCNs) were normalized to microgram of DNA used in the dPCR reaction.
  • dPCR digital polymerase chain reaction
  • Enzymatic reactions were set up using 10 pL of sample (cell lysate or tissue homogenate) diluted appropriately and 75 pL of 4- methylumbelliferyl-a-D-glucopyranoside (4MU-a-Gal; also referred to herein as “4-MUG”; Burlington, MA) substrate, in black 96-well plates (PerkinElmer, Waltham, MA).
  • the reaction mixture was incubated at 37°C for 1 hour and then stopped by adding 150 pL of stop buffer (133 mM Glycine, 83 mM Sodium Carbonate, pH 10).
  • a standard curve (0-4.25 nmol/mL) was used to measure released fluorescent 4-methylumbelliferone (4-MU) from the individual reaction mixture using the Spectramax M3 reader (PerkinElmer, Waltham, MA) at 460 nm (emission) and 360 nm (excitation).
  • the protein concentration of the clarified supernatant was quantified using a BradfordUltra assay (AbCam, Waltham, MA).
  • the released 4-MU concentration was divided by the sample protein concentration, and activity was reported as nanomoles per hour per milligram protein.
  • H&E haematoxylin and eosin
  • Muscle cryosections were cut (4-5 Im), placed onto slides, and air-dried for 20 minutes.
  • H&E and Periodic Acid Schiff (PAS) stains were performed according to established protocols using kits from Biovision (Waltham, MA).
  • IHC GAA immunohistochemistry
  • samples were incubated overnight at 4°C with an anti-GAA rabbit antibody (Sigma/HPA029126, St.
  • MAPPs Major Histocompatibility - Associated Peptide Proteomics
  • PBMCs Human peripheral blood mononuclear cells
  • monocyte-derived dendritic cells To prepare monocyte-derived dendritic cells (MoDCs), fresh PBMC from 20 healthy donors were used and CD14+ cells (monocytes) were isolated using RoboSepTM negative human monocyte isolation kits and a RoboSepTM cell isolation instrument (StemCell Technologies, Cambridge, UK) according to the manufacturer’s instructions. Monocytes were re-suspended in MoDC culture medium (RPMI 1640 supplemented with 10% FBS, 50gm 2-ME, 2mM L-Glutamine (all from ThermoFisher Scientific, Loughborough, UK), IL-4 (Peprotech, London, UK), and GM-CSF (Peprotech) and plated in tissue culture flasks.
  • MoDC culture medium RPMI 1640 supplemented with 10% FBS, 50gm 2-ME, 2mM L-Glutamine (all from ThermoFisher Scientific, Loughborough, UK), IL-4 (Peprotech, London, UK), and
  • HLA-DR/peptide complexes were purified from the cell lysate by immunoprecipitation using magnetic beads (Promega, Southampton, UK) coated with anti-HLA-DR antibody (BioLegend, London, UK) overnight at 4 °C.
  • Peptides bound to HLA-DR were eluted under acidic conditions (3% MeCN, 0.2% TFA; ThermoFisher Scientific Waltham, MA) and purified by solid phase extraction using Oasis® HLB pElution plates (Waters, Ellsmere Port, UK). Peptides were freeze-dried using a 5301 vacuum concentrator (Eppendorf, Stevenage, UK) and stored at -80 °C until analyzed by mass spectrometry (MS).
  • Peptides were identified using the Sequest algorithm, built in the Proteome Discoverer software v2.1 (ThermoFisher Scientific, Waltham, MA) against a proprietary database and the sequences of the test samples determined. Once the final list of identified peptides was completed, the sequence heatmaps were generated using MATLAB (MathWorks®, Cambridge, UK).
  • C2C12 myoblast cells were seeded into 24-well tissue culture plates (Coming,
  • Lysate GAA activity of duplicate wells was measured with the GAA 4-methylumbelliferyl-a- D-glucopyranoside (4-MUG) assay (described above) and normalized to lysate protein concentration as determined using the bicinchoninic acid protein assay (Pierce, Appleton, Wisconsin). Thermal Stability
  • Thermal protein unfolding was monitored using a Prometheus NT.48 instrument (NanoTemper Technologies, Miinchen, Germany). For each condition, 50 pl of a 1 mg/ml protein solution was prepared, and 20 pl of sample was filled into 3 low volume differential scanning fluorimetry (nanoDSF) Grade Standard Capillaries (NanoTemper Technologies, Miinchen, Germany), respectively, and loaded into the instrument. Thermal unfolding of the proteins was monitored in a 1 °C/minute thermal ramp from 25 °C to 95 °C. Tm values were determined automatically by the PR control software.
  • An 8M stock of GuHCl was prepared by mixing 7.64 g of GuHCl with 4.21 ml assay buffer (50 M sodium phosphate buffer, pH 7.5, 150 mM NaCl) and the pH adjusted to pH 7.5 using 1 M Tris, pH 8.0.
  • 9 M stock of urea was prepared freshly by mixing 5.41 g urea with 5.9 ml lx assay buffer.
  • One microliter of concentrated protein stock (final concentration 0.5 mg/ml) was added to 30 p.l of a series of denaturant concentrations (0.25- 6M) and the mixture was incubated for 1 hour and 16 hours at 25°C.
  • Trans-thoracic echocardiography was performed on mice that were anaesthetized with isoflurane.
  • Parasternal long axis and short axis images were obtained using an MX550S probe attached to Vevo3100 (FujiFilms, Visulsonics, Ontario, Canada). Images were acquired when the rectal temperature was between 36-38 °C and respiratory rate was 40-120 breaths per minute. Images were analyzed by Vevo Lab 5.5.0 (FujiFilms, Visulsonics, Ontario, Canada). An average of three heart beats were used for analysis.
  • An accelerating rotarod assay was used to determine neuro-motor coordination of the Pompe GAA KO mice.
  • the assay was performed on a rotarod apparatus (Model 47650; Ugo Basile, Italy) that was set to accelerate from 5-40 rpm over 5 minutes. Latency and the speed at fall were recorded from a total of 6 trials (3 trials/day). Animals were habituated and trained before the test trials.
  • Example 2 Establishing Optimal Tissue Expression For Acid Alpha-Glucosidase Gene Delivery Relevant To Pompe Disease Patients Using Mouse Models
  • a comparison of GAA protein activity and glycogen levels in key muscle tissues in mice after treatment with engineered AAV9 vectors comprising a GAA transgene in the presence of a muscle-specific promoter (SPc512) in the presence or absence of either a single muscle-specific enhancer (CSk-SH5 or Dph-CRE04) or both the CSK-SH5 and the Dph- CRE04 muscle-specific enhancers in two orientations was performed and compared to control gaa +l+ (GAA WT) and gaa 1 ' (GAA KO) mice (FIG. 2A and FIG. 2B).
  • GAA KO mice were dosed once with 3xl0 13 vg/kg of a AAV9.GAA vector only differing in their enhancer-promoter elements (z.e., all vectors comprised an AAV9 capsid and encoded the same codon optimized wild-type (WT) human GAA transgene CO3 - see Example 3). After a 5 week incubation, the animals were sacrificed, their muscle tissues harvested, and GAA enzymatic activity determined according to the 4-MUG assay described in Example 1.
  • GAA activity were observed in heart, diaphragm, quadriceps, gastrocnemius, triceps and aorta of GAA KO mice treated with vectors that included enhancer elements as compared to animals that were dosed with vectors that did not have any enhancer elements added to the muscle specific SPc512 promoter (FIG. 2A).
  • CSk-SH5 increased GAA activity in heart, quadriceps, triceps, gastrocnemius, diaphragm and aorta
  • Dph- CRE04 increased GAA activity in heart quadriceps, triceps, gastrocnemius and diaphragm.
  • FIG. 3 shows the strategy that was used to engineer a GAA transgene that had improved catalytic activity, stability, muscle cell uptake, and immunogenic profile, beginning with a codon optimized GAA protein coding sequence, followed by two separate campaigns that identified optimal amino acid substitutions.
  • Example 3 describes the identification of a preferred codon-optimized human GAA that provided enhanced gene expression.
  • Example 4 describes Campaign 1 in which certain amino acid substitutions were introduced into the preferred codon-optimized GAA amino acid sequence that improved stability, uptake into muscle, and reduced immunogenicity of the GAA protein and a preferred variant of GAA was identified.
  • Example 5 describes Campaign 2, in which certain amino acid substitutions were introduced into the preferred GAA variant of Campaign 1 that improved catalytic activity.
  • Example 6 shows additional in vivo data for the final GAA variant that contained all the preferred amino acid substitutions, which showed significantly enhanced activity, stability, muscle cell uptake, and immunogenic profile.
  • Example 3 Codon Optimization of WT GAA Nucleotide Sequence to Increase Expression
  • the human WT GAA nucleotide sequence was codon optimized in order to improve expression in human muscle cells, while reducing the immuno - stimulatory CpG content.
  • Three codon variants with reduced CpGs (named CO1, CO2, and CO3) were inserted into an expression cassette comprising a muscle-specific Sk-SH4 enhancer and the muscle specific desmin promoter [AAV9. Sk-SH4. desmin. COGAA] (FIG. 4A) and tested for expression at a dose of 3xl0 13 vg/kg in GAA KO mice.
  • the CO3 codon optimized variant demonstrated increased GAA expression and activity at or slightly higher than the WT hGAA sequence and resulted in a similar or more efficient reduction in the levels of glycogen in heart, diaphragm, and quadriceps muscle (FIG. 4B-FIG. 4D and FIG. 4E-FIG. 4G, respectively).
  • Significant clearance of glycogen in the heart muscle of GAA KO mice dosed with vectors comprising CO3 was also observed in tissue sections using PAS staining. Further, a 76 kDa mature form of GAA was detected in tissue lysates confirming correct GAA processing in the lysosomes. While these results were encouraging, insufficient clearance of glycogen observed in the skeletal and respiratory muscles suggested that further improvements to the GAA protein were required to increase its efficacy.
  • Example 4 Engineering GAA To Identify Therapeutically Improved Variants With Enhanced Stability, Muscle Uptake, And Reduced Immunogenicity (Campaign 1)
  • GAA variant proteins 1-4 (Var 1, Var 2, Var 3, and Var 4) with increased stability and uptake by muscle cells, as well as significantly reduced immunogenic profile compared to WT hGAA were obtained from Codexis, Inc. (WO2021127457). These variant proteins were identified using Codexis’ proprietary CodeEvolver® technology platform. Table 3 provides the amino acid sequence of each of these proteins as well as a consensus sequence.
  • Table 4 summarizes certain key characteristics of variant human GAA proteins 1- 4 (Var 1, Var 2, Var 3, and Var 4) from Campaign 1.
  • Column 2 shows the degree of uptake of the GAA variants compared to WT hGAA by cultured C2C12 GAA KO mouse myoblasts after overnight incubation with the purified variant GAA proteins or the WT hGAA control.
  • Column 3 shows activity of the GAA variants 1-4 or the WT hGAA control as measured in cell lysates using the 4-MUG assay as described in Example 1.
  • Column 4 shows the stability following incubation of the purified variant GAA proteins or WT hGAA control in human plasma for 3 days at 37°C using the 4-MUG assay and compared to unchallenged protein.
  • GAA variants 1-4 showed an increase in uptake by cultured C2C12 GAA KO mouse myoblasts after overnight incubation with purified GAA variant proteins or WT hGAA control, as described in Example 1 (Table 4, column 2).
  • GAA variants 1-4 all had increased kinetic activity compared to WT hGAA with the order of greatest activity being Var 3 > Var 4 > Vari > Var 2 > WT GAA (Table 4, column 3 and FIG. 5A).
  • FIG. 5B shows enzymatic activity of GAA variants 1-4 when incubated with the natural substrate, glycogen, for 1 hour.
  • Var 2 and Var 4 had similar activity to WT hGAA, whereas Var 1 had about double the activity of WT hGAA and Var 3 had greater than 3x the activity of WT hGAA with the order of activity being Var 3 > Var 1 > Var 4 > Var 2 > WT.
  • Var 4 had the highest level of stability compared to WT hGAA.
  • GAA variant 4 protein and WT hGAA protein were also subjected to a MAPPs analysis, an in vitro assay that identifies immunogenic peptides presented by dendrictic cells (DC) on MHCII molecules to provide an indication of the immunogenicity of the GAA variant proteins (see Example 1).
  • MAPPs assays are a useful tool for interrogating clinical immunogenicity root causes by determining the natural presentation of peptides and confirming the relevance of T cell epitopes that have been identified via peptide library T cell epitope mapping approaches (Karie et al. (2020)).
  • FIG. 6A and FIG. 6C show exemplary immunogenic profiles for GAA variant 4 (Var 4).
  • GAA Var 4 Based on the above the data, the preferred combination of improvement to muscle uptake, enzymatic activity, stability, and reduced immunogenic profile was observed for GAA Var 4.
  • the amino acid changes from GAA Var 4 were introduced into the nucleotide sequence encoding the WT GAA (SEQ ID NO: 1).
  • AAV9 vectors expressing either WT hGAA or GAA Var 4 under the control of the SPp512 promoter and Dph-CRE04 enhancer element were constructed and assessed at 5 weeks and 12 weeks in GAA KO mice following i.v. injection at a dose of 2.4xl0 13 vg/kg. At 12 weeks post dosing supraphy si ologi cal GAA activity in the heart (FIG.
  • urinary Glc4 a biomarker that is associated with Pompe disease, also showed a normalization to WT levels in GAA KO mice dosed with vector expressing GAA Var 4 (FIG. 8E)
  • the hGAA variant proteins 1-4 were designed to have increased uptake and stability as well as reduce immunogenicity, the mutations to the WT GAA did not specifically address enhanced activity.
  • Campaign 2 was specifically designed to identify mutations in the WT GAA amino acid sequence that would enhance GAA activity and is described in Example 5.
  • GAA variant 6 (Var 6) was further evaluated in vitro for catalytic activity using kinetic assays using either the synthetic substrate 4-MUG (FIG. 10A) or the natural substrate glycogen (FIG. 10B). A 3.5x fold improvement in the activity on lOmg/ml or lOOmg/ml glycogen substrate was observed for GAA Var 6 compared to the WT GAA (FIG. 10B).
  • Example 6 GAA With Enhanced Activity, Stability, Uptake, and Reduced Immunogenicity (Combining Campaign 1 and Campaign 2 hGAA Variants)
  • GAA Var 4+6 was first evaluated for time-dependent catalytic activity using the 4-MUG substrate compared to GAA Var 4.
  • the GAA Var4+6 had an increased activity that was 4-fold higher than the WT hGAA as well as a 1.5 fold increase in activity over the GAA Var 4 (FIG. 11 A).
  • GAA Var 4+6 demonstated significantly higher stability in both PBS as well as serum at RT after 24, 48 or 72 hours incubation compared not only to the WT hGAA but also to GAA Var 4 (FIG. 11B).
  • GAA Var 4+6 demonstrated a higher stability when tested at pH 5.2 and 7.4 (FIG. 11C), as well as higher thermal and chemical stability (FIG.
  • GAA Var 4+6 also demonstated an increase in uptake of this protein by mouse muscle cells, when purified GAA Var 4+6 protein was incubated for 24 hours and activity measure in cell lysates compared to both WT hGAA and GAA Var 4 (FIG. 12A).
  • the immunogenic profile of GAA Var 4+6 as assessed using MAPPs assays demonstrated a reduced display of MHCII peptides using human PBMC derived dendritic cells (FIG. 12B).
  • Example 7 Treatment of Pompe Disease By Administration Of Vectors Containing An Engineered Optimal GAA Transgene And Tissue Specific Expression
  • the finalized lead GAA variant 4+6 (Var 4+6) was combined with the Dph- CRE04 enhancer and SPc512 promoter, vectorized using an AAV9 capsid, and assessed following dosing of GAA KO mice.
  • a dose dependent increase in GAA activity (0.75-50 fold of WT mice) was observed when GAA KO mice were dosed with the vectors expressing both the WT GAA or GAA Var 4+6 in heart tissue lysates (FIG. 13A). More importantly, a significant increase in activity, approximately 5-20 fold higher, at all doses tested was observed with the GAA Var 4+6 compared to WT GAA (FIG.
  • the present gene therapy has a combination of the following properties to allow for better targetting to the muscles: (1) AAV9, a vector known to have higher tropism to muscles (Inagaki et.al., 2006), (2) a muscle-specific enhancer element that significantly increases the expression of the GAA protein transgene, (3) a muscle-specific promoter; and (4) an engineered hGAA variant that has been codon-optimized and designed to have increased stability, muscle cell uptake, reduced immunogenicity, and increased enzymatic activity.
  • GAA KO mice were treated with an AAV9 vector expressing either the WT hGAA protein or GAA Var 4+6 using the Dph-CRE04 musclespecific enhancer element and the SPc512 muscle promoter and compared them to GAA KO mice treated with the AAV8.MCK.WT hGAA vector.
  • FIG. 16B a 5-fold and 43 -fold increase in GAA activity in the heart of GAA KO mice treated with vector expressing the WT GAA and GAA Var 4+6, respectively, was seen compared to GAA KO mice treated with the AT845-like vector at 3xl0 13 vg/kg.
  • FIG. 16C and FIG. 16D also show a similar trend of the increased GAA activity in the deeper skeletal muscles such as the quadriceps and diaphragm, respectively, at much lower dose in GAA KO mice dosed with the GAA Var 4+6 vector compared to the AAV8.MCK.WT hGAA vector.
  • the increased GAA activity with the GAA Var 4+6 vector was associated with a greater than 90% glycogen reduction at a dose of IxlO 13 vg/kg (3-fold lower dosing) compared to a similar level of reduction with the AAV8.MCK.WT hGAA vector at 3xl0 13 vg/kg in heart quadriceps, and diaphragm muscle (FIG. 17A-FIG. 17C). Only a 28% reduction in glycogen was observed in the quadriceps with the AAV8.MCK.WT hGAA vector dosed at 3xl0 13 vg/kg, for example (FIG. 17B).
  • GAA KO mice were injected either 3xl0 12 or IxlO 13 3xl0 13 vg/kg of the AAV9 vector expressing either the GAA Var 4+6 using the Dph-CRE04 muscle-specific enhancer element and the SPc512 muscle promoter at either at 12 weeks or 24 weeks old. Mice were sacrificed at either 5 weeks, 3 months or 6 months after vector injection to determine the durability of transgene expression over time and the associated efficacy. A dose dependent increase in the number of viral genomes was observed in the heart tissue of dosed GAA KO mice that remained constant for the entire 6 month period and at all time points tested (FIG. 31).
  • the dose dependent increase in vector genomes translated into a dose and time dependent increase in GAA activity in the heart as well as the quadriceps and diaphragm muscles (FIG. 32-FIG. 37).
  • the dose and time dependent transduction was comparable between the younger (12 week old) as well older (24 week old) GAA KO mice with the time dependent increase in GAA activity across all tissues tested irrespective of the age at dosing (FIG. 32-FIG. 37). This is in contrast to observations from other groups who have developed GTs for Pompe (Puzzo e.al., 2017; Collela, Eggers ).

Landscapes

  • Health & Medical Sciences (AREA)
  • Life Sciences & Earth Sciences (AREA)
  • Chemical & Material Sciences (AREA)
  • Engineering & Computer Science (AREA)
  • Genetics & Genomics (AREA)
  • General Health & Medical Sciences (AREA)
  • Organic Chemistry (AREA)
  • Biotechnology (AREA)
  • Bioinformatics & Cheminformatics (AREA)
  • Zoology (AREA)
  • Wood Science & Technology (AREA)
  • Molecular Biology (AREA)
  • Biomedical Technology (AREA)
  • Medicinal Chemistry (AREA)
  • General Engineering & Computer Science (AREA)
  • Veterinary Medicine (AREA)
  • Animal Behavior & Ethology (AREA)
  • Biochemistry (AREA)
  • Public Health (AREA)
  • Pharmacology & Pharmacy (AREA)
  • Microbiology (AREA)
  • Epidemiology (AREA)
  • General Chemical & Material Sciences (AREA)
  • Chemical Kinetics & Catalysis (AREA)
  • Biophysics (AREA)
  • Physics & Mathematics (AREA)
  • Environmental Sciences (AREA)
  • Nuclear Medicine, Radiotherapy & Molecular Imaging (AREA)
  • Plant Pathology (AREA)
  • Virology (AREA)
  • Animal Husbandry (AREA)
  • Biodiversity & Conservation Biology (AREA)
  • Neurology (AREA)
  • Orthopedic Medicine & Surgery (AREA)
  • Physical Education & Sports Medicine (AREA)
  • Micro-Organisms Or Cultivation Processes Thereof (AREA)
  • Enzymes And Modification Thereof (AREA)
  • Medicines That Contain Protein Lipid Enzymes And Other Medicines (AREA)

Abstract

L'invention concerne des polypeptides d'alpha-glucosidase acide (GAA) variants, des polynucléotides optimisés par codons codant pour GAA, ainsi que des méthodes et des constructions de thérapie génique à GAA.
PCT/IB2024/056362 2023-06-30 2024-06-28 Traitement de la maladie de pompe Ceased WO2025004002A2 (fr)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
US202363511367P 2023-06-30 2023-06-30
US63/511,367 2023-06-30

Publications (2)

Publication Number Publication Date
WO2025004002A2 true WO2025004002A2 (fr) 2025-01-02
WO2025004002A3 WO2025004002A3 (fr) 2025-02-13

Family

ID=91853495

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/IB2024/056362 Ceased WO2025004002A2 (fr) 2023-06-30 2024-06-28 Traitement de la maladie de pompe

Country Status (1)

Country Link
WO (1) WO2025004002A2 (fr)

Citations (6)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US4797368A (en) 1985-03-15 1989-01-10 The United States Of America As Represented By The Department Of Health And Human Services Adeno-associated virus as eukaryotic expression vector
US5139941A (en) 1985-10-31 1992-08-18 University Of Florida Research Foundation, Inc. AAV transduction vectors
US5994136A (en) 1997-12-12 1999-11-30 Cell Genesys, Inc. Method and means for producing high titer, safe, recombinant lentivirus vectors
US6013516A (en) 1995-10-06 2000-01-11 The Salk Institute For Biological Studies Vector and method of use for nucleic acid delivery to non-dividing cells
WO2019153009A1 (fr) 2018-02-05 2019-08-08 Audentes Therapeutics, Inc. Éléments de régulation de la transcription et utilisations associées
WO2021127457A1 (fr) 2019-12-20 2021-06-24 Codexis, Inc. Variants d'alpha-glucosidase acide modifiés

Family Cites Families (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
WO2003092598A2 (fr) * 2002-04-30 2003-11-13 University Of Florida Traitement contre la maladie de pompe
CA3230004A1 (fr) * 2021-08-25 2023-03-02 Yunxiang Zhu Particules d'aav comprenant une proteine capsidique tropique du foie et une alpha-glucosidase acide (gaa) et leur utilisation pour traiter la maladie de pompe

Patent Citations (7)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US4797368A (en) 1985-03-15 1989-01-10 The United States Of America As Represented By The Department Of Health And Human Services Adeno-associated virus as eukaryotic expression vector
US5139941A (en) 1985-10-31 1992-08-18 University Of Florida Research Foundation, Inc. AAV transduction vectors
US6013516A (en) 1995-10-06 2000-01-11 The Salk Institute For Biological Studies Vector and method of use for nucleic acid delivery to non-dividing cells
US5994136A (en) 1997-12-12 1999-11-30 Cell Genesys, Inc. Method and means for producing high titer, safe, recombinant lentivirus vectors
WO2019153009A1 (fr) 2018-02-05 2019-08-08 Audentes Therapeutics, Inc. Éléments de régulation de la transcription et utilisations associées
WO2021127457A1 (fr) 2019-12-20 2021-06-24 Codexis, Inc. Variants d'alpha-glucosidase acide modifiés
US20210189365A1 (en) 2019-12-20 2021-06-24 Codexis, Inc. Engineered acid alpha-glucosidase variants

Non-Patent Citations (59)

* Cited by examiner, † Cited by third party
Title
"CLONING AND EXPRESSION VECTORS FOR GENE FUNCTION ANALYSIS", 2001, BIOTECHNIQUES PRESS
"CONTROLLED DRUG BIOAVAILABILITY, DRUG PRODUCT DESIGN AND PERFORMANCE", 1984, WILEY & SONS
"CURRENT PROTOCOLS IN MOLECULAR BIOLOGY", 1993, JOHN WILEY & SONS
"GenBank", Database accession no. NP_000143.2
"Macromolecule Sequencing and Synthesis, Selected Methods and Applications", 1988, ALAN R. LISS, INC, article "Current Methods in Sequence Comparison and Analysis", pages: 127 - 149
"MOLECULAR CLONING: A LABORATORY MANUAL", 1989, COLD SPRING HARBOR LABORATORY PRESS
"UniProt", Database accession no. Q6P7A9
ALTSCHUL ET AL., J. MOL. BIOL., vol. 215, 1990, pages 403 - 410
ALTSCHUL ET AL., METHODS IN ENZYMOLOGY, vol. 266, 1996, pages 460 - 480
ALTSCHUL ET AL., NUCL. ACIDS RES., vol. 25, pages 3389 - 3402
ALTSCHUL ET AL., NUCLEIC ACIDS RES., vol. 25, 1997, pages 3389 - 3402
ASOKAN A ET AL., MOL. THER., vol. 20, no. 4, 2012, pages 443 - 455
BLOMER ET AL., J VIROL., vol. 71, no. 9, 1997, pages 6641 - 6649
COTTEN ET AL., P.N.A.S. U.S.A., vol. 89, no. 13, 1992, pages 6094 - 98
CURIEL, NAT IMMUN, vol. 13, no. 2-3, 1994, pages 141 - 64
DAYABERNS, CLIN. MICROBIOL. REV., vol. 21, no. 4, 2008, pages 583 - 593
DEVEREUX ET AL., NUCL. ACID RES., vol. 12, 1984, pages 387 - 395
DIAZ-MANERA, J. ET AL., THE LANCET. NEUROLOGY, vol. 20, no. 12, 2021, pages 1027 - 1037
DOBROWSKY ET AL., CURR. OPINION BIOMED. ENGIN., vol. 20, 2021, pages 100353
EL ANDARI ET AL., SCI ADV, vol. 8, no. 38, 2022, pages eabn4704
FARAH ET AL., FASEB J., vol. 28, no. 5, 2014, pages 2272 - 2280
FAUST ET AL., J. CLIN. INVEST., vol. 123, 2013, pages 2994 - 3001
FENGDOOLITTLE, J. MOL. EVOL., vol. 35, 1987, pages 351 - 360
GARDINER-GARDEN M. ET AL., J MOL BIOL., vol. 196, no. 2, 1987, pages 261 - 82
GEEL ET AL., MOL. GENET. METAB., vol. 92, no. 4, 2007, pages 49057 - 307
GIEGE ET AL.: "SHORT PROTOCOLS IN MOLECULAR BIOLOGY", 1999, OXFORD UNIVERSITY PRESS, article "CRYSTALLIZATION OF NUCLEIC ACIDS AND PROTEINS"
GRAY ET AL., HUMAN GENE THERAPY, vol. 22, 2011, pages 1143 - 53
GUINER ET AL., NAT. COMM., vol. 8, 2017, pages 16106
HAMMERLING ET AL.: "MONOCLONAL ANTIBODIES AND T-CELL HYBRIDOMAS", 1981, ELSEVIER
HARFOUCHE, J. PATIENT REP. OUTCOMES, vol. 4, no. 1, 2020, pages 83
HAUTPINTEL, J VIROL., vol. 72, no. 3, 1998, pages 1834 - 43
HIGGINSSHARP, CABIOS, vol. 5, 1989, pages 151 - 153
KABAT ET AL.: "FROM GENES TO CLONES: INTRODUCTION TO GENE TECHNOLOGY", 1987, NATIONAL INSTITUTES OF HEALTH
KARLIN ET AL., PROC. NATL. ACAD. SCI. U.S.A., vol. 90, 1993, pages 5873 - 5787
KELLEHER AND VOS, BIOTECHNIQUES, vol. 17, no. 6, 1994, pages 1110 - 17
KISHNANI, P.S., GENETICS IN MEDICINE, vol. 25, no. 2, 2023, pages 100328
KRIEGLER, M.: "GENE TRANSFER AND EXPRESSION, A LABORATORY MANUAL", 1990, W.H. FREEMAN CO.
KUDLA ET AL., PLOS BIOL., vol. 4, no. 6, 2006, pages 80
LAZAR, EUR. HEART J., vol. 38, no. 30, 2017, pages 2333 - 2342
LINDENBAUM ET AL., NUCLEIC ACIDS RES., vol. 32, no. 21, 2004, pages e172
MANN ET AL., CELL, vol. 33, 1983, pages 153 - 159
MCCALL ET AL., J. SMOOTH MUSCLE RES., vol. 54, no. 0, 2018, pages 100 - 118
MORELAND ET AL.: "Species-specific differences in the processing of acid α-glucosidase are due to the amino acid identity at position 201", GENE, vol. 491, 2012, pages 25 - 30, XP055028473, DOI: 10.1016/j.gene.2011.09.011
MURRY, E. J.: "SEQUENCES OF PROTEINS OF IMMUNOLOGICAL INTEREST", vol. 7, 1991, U.S. DEPARTMENT OF HEALTH AND HUMAN SERVICES
MUZYCZKA, CURR TOP MICROBIOL IMMUNOL, vol. 158, 1992, pages 97 - 129
NALDINI ET AL., SCIENCE, vol. 272, no. 5259, 1996, pages 263 - 267
NEEDLEMANWUNSCH, J. MOL. BIOL., vol. 48, 1970, pages 443
NICOLASRUBINSTEIN: "Vectors: A survey of molecular cloning vectors and their uses", 1988, COLD SPRING HARBOR LABORATORY PRESS, pages: 494 - 513
PEARSONLIPMAN, PROC. NATL. ACAD. SCI. U.S.A., vol. 85, 1988, pages 2444
PÉREZ-LUZDIAZ-NIDO, J BIOMED BIOTECHNOL., 2010, pages 2010
ROIG-ZAMBONI V ET AL.: "Structure of human lysosomal acid a-glucosidase-a guide for the treatment of Pompe disease", NAT COMMUN., vol. 8, no. 1, 2017, pages 1111, XP055915132, DOI: 10.1038/s41467-017-01263-3
SMITHWATERMAN, ADV. APPL. MATH., vol. 2, 1981, pages 482
SPALDING, CELL, vol. 122, no. 1, 2005, pages 133 - 43
TABEBORDBAR, CELL, vol. 184, no. 19, 2021, pages 4919 - 4938
TEMIN: "Gene Transfer", 1986, PLENUM PRESS, pages: 149 - 188
VANDENDRIESSCHE ET AL., BLOOD, vol. 114, 2009, pages 2009 - 2010
WU Z ET AL., MOL THER., vol. 16, no. 2, 2008, pages 280 - 9
XU ET AL., JCI INSIGHT, vol. 4, no. 5, 2019, pages e125358
ZUFFEREY ET AL., NAT BIOTECHNOL, vol. 15, no. 9, 1997, pages 871 - 875

Also Published As

Publication number Publication date
WO2025004002A3 (fr) 2025-02-13

Similar Documents

Publication Publication Date Title
US20220396813A1 (en) Recombinase compositions and methods of use
JP2024038518A (ja) 常染色体優性疾患のための遺伝子編集
Viecelli et al. Treatment of phenylketonuria using minicircle‐based naked‐DNA gene transfer to murine liver
US20250313859A1 (en) Compositions useful in treatment of metachromatic leukodystrophy
EP3519569B1 (fr) Vecteurs viraux recombinés adéno-associés pour le traitement de la mucopolysaccharidose
US20250049955A1 (en) Compositons and methods for the treatment of neurological disorders related to glucosylceramidase beta deficiency
CN117321213A (zh) 具有优选表达水平的腺相关病毒组合物
EP3562494A1 (fr) Thérapie génique pour le traitement de la phénylcétonurie
CN119613504A (zh) 腺相关病毒变异衣壳和其使用方法
JP2018520646A (ja) ファブリー病の遺伝子治療
EP4433601A2 (fr) Compositions et méthodes de traitement de la sclérose latérale amyotrophique et de troubles associés à la moelle épinière
KR20220004696A (ko) 폼페병의 치료에 유용한 조성물
IL293972A (en) Preparations for the treatment of Friedrich's ataxia
CN111718947B (zh) 用于治疗ⅲa或ⅲb型粘多糖贮积症的腺相关病毒载体及用途
TW202229560A (zh) 治療法布瑞氏症之組成物及方法
TW202449167A (zh) 用於治療肌肉萎縮性脊髓側索硬化症之組成物及方法
JP2023526923A (ja) ポンペ病の治療に有用な組成物
US20250161493A1 (en) Compositions and methods for in vivo nuclease-mediated treatment of ornithine transcarbamylase (otc) deficiency
WO2024231820A1 (fr) Traitement de la maladie de pompe
CN121844054A (zh) 具有增加的脑富集的腺相关病毒组合物
KR20230003554A (ko) 낮은 전사 활성을 갖는 프로모터를 사용하여 뉴클레아제 발현 및 표적-외 활성을 감소시키기 위한 조성물 및 방법
US20250099618A1 (en) Recombinant tert-encoding viral genomes and vectors
KR20250156211A (ko) 글루코실세라미다제 베타 1 결핍증과 관련된 신경 장애의 치료를 위한 조성물 및 방법
WO2023034966A1 (fr) Compositions et procédés d'utilisation de celles-ci pour traiter les troubles associés à la thymosine βeta 4
CN118574933A (zh) 用于治疗鸟氨酸氨甲酰转移酶(otc)缺乏症的方法

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 24739722

Country of ref document: EP

Kind code of ref document: A2

NENP Non-entry into the national phase

Ref country code: DE