WO2024173367A2 - Polypeptides recombinés pour la biosynthèse d'acide ursodésoxycholique (udca) - Google Patents
Polypeptides recombinés pour la biosynthèse d'acide ursodésoxycholique (udca) Download PDFInfo
- Publication number
- WO2024173367A2 WO2024173367A2 PCT/US2024/015560 US2024015560W WO2024173367A2 WO 2024173367 A2 WO2024173367 A2 WO 2024173367A2 US 2024015560 W US2024015560 W US 2024015560W WO 2024173367 A2 WO2024173367 A2 WO 2024173367A2
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- polypeptide
- seq
- sequence
- amino acid
- cyp
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Classifications
-
- C—CHEMISTRY; METALLURGY
- C12—BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
- C12Y—ENZYMES
- C12Y114/00—Oxidoreductases acting on paired donors, with incorporation or reduction of molecular oxygen (1.14)
- C12Y114/14—Oxidoreductases acting on paired donors, with incorporation or reduction of molecular oxygen (1.14) with reduced flavin or flavoprotein as one donor, and incorporation of one atom of oxygen (1.14.14)
- C12Y114/14001—Unspecific monooxygenase (1.14.14.1)
-
- C—CHEMISTRY; METALLURGY
- C12—BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
- C12N—MICROORGANISMS OR ENZYMES; COMPOSITIONS THEREOF; PROPAGATING, PRESERVING, OR MAINTAINING MICROORGANISMS; MUTATION OR GENETIC ENGINEERING; CULTURE MEDIA
- C12N9/00—Enzymes; Proenzymes; Compositions thereof; Processes for preparing, activating, inhibiting, separating or purifying enzymes
- C12N9/0004—Oxidoreductases (1.)
- C12N9/0012—Oxidoreductases (1.) acting on nitrogen containing compounds as donors (1.4, 1.5, 1.6, 1.7)
- C12N9/0036—Oxidoreductases (1.) acting on nitrogen containing compounds as donors (1.4, 1.5, 1.6, 1.7) acting on NADH or NADPH (1.6)
- C12N9/0038—Oxidoreductases (1.) acting on nitrogen containing compounds as donors (1.4, 1.5, 1.6, 1.7) acting on NADH or NADPH (1.6) with a heme protein as acceptor (1.6.2)
- C12N9/0042—NADPH-cytochrome P450 reductase (1.6.2.4)
-
- C—CHEMISTRY; METALLURGY
- C12—BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
- C12N—MICROORGANISMS OR ENZYMES; COMPOSITIONS THEREOF; PROPAGATING, PRESERVING, OR MAINTAINING MICROORGANISMS; MUTATION OR GENETIC ENGINEERING; CULTURE MEDIA
- C12N9/00—Enzymes; Proenzymes; Compositions thereof; Processes for preparing, activating, inhibiting, separating or purifying enzymes
- C12N9/0004—Oxidoreductases (1.)
- C12N9/0071—Oxidoreductases (1.) acting on paired donors with incorporation of molecular oxygen (1.14)
-
- C—CHEMISTRY; METALLURGY
- C12—BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
- C12P—FERMENTATION OR ENZYME-USING PROCESSES TO SYNTHESISE A DESIRED CHEMICAL COMPOUND OR COMPOSITION OR TO SEPARATE OPTICAL ISOMERS FROM A RACEMIC MIXTURE
- C12P33/00—Preparation of steroids
- C12P33/06—Hydroxylating
-
- C—CHEMISTRY; METALLURGY
- C12—BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
- C12Y—ENZYMES
- C12Y106/00—Oxidoreductases acting on NADH or NADPH (1.6)
- C12Y106/02—Oxidoreductases acting on NADH or NADPH (1.6) with a heme protein as acceptor (1.6.2)
- C12Y106/02004—NADPH-hemoprotein reductase (1.6.2.4), i.e. NADP-cytochrome P450-reductase
-
- C—CHEMISTRY; METALLURGY
- C12—BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
- C12R—INDEXING SCHEME ASSOCIATED WITH SUBCLASSES C12C - C12Q, RELATING TO MICROORGANISMS
- C12R2001/00—Microorganisms ; Processes using microorganisms
- C12R2001/645—Fungi ; Processes using fungi
- C12R2001/84—Pichia
Definitions
- FIELD FIELD
- the present disclosure relates to recombinant polypeptides with cytochrome P450 (CYP) activity capable of selectively catalyzing the conversion of precursor compounds, LCA or 3-KCA, to the product compounds, UDCA or 3-KUDCA, and the engineering of these polypeptides for heterologous expression in yeast allowing for improved fermentative bioproduction of these product compounds.
- CYP cytochrome P450
- REFERENCE TO SEQUENCE LISTING [0003] The official copy of the Sequence Listing is submitted concurrently with the specification via USPTO Patent Center as an WIPO Standard ST.26 formatted XML file with file name “13421-020WO1.xml”, a creation date of February 13, 2024, and a size of 1,355,138 bytes.
- Ursodeoxycholic acid is a secondary bile acid that has high therapeutic value in the treatment of cholestatic liver diseases. UDCA also has cytoprotective and anti-inflammatory properties and is being pursued as a potential therapeutic to treat neurological, cardiovascular, and inflammatory bowel diseases.
- UDCA is primarily harvested from animal bile (black bears, cattle, poultry). UDCA can also be produced via a costly chemical synthesis route from cholic acid that has 5-steps including a Wolff-Kishner ketone reduction, and an epimerization at C7 to produce UDCA.
- WO2022115710A1 describes a recombinant yeast transformed with a naturally occurring cytochrome P450 monooxygenase from Gibberella zeae that is capable of carrying out the biosynthesis of UDCA from LCA, or the structurally related bioconversion of 3-KCA to 3-KUDCA. [0005] There remains a need for a more efficient recombinant cell systems for the commercially viable biosynthetic production UDCA.
- the present disclosure relates generally to recombinant polypeptides with cytochrome P450 (CYP) activity capable of selectively catalyzing the conversion of precursor compounds, LCA or 3-KCA, to the product compounds, UDCA or 3-KUDCA, and the engineering of these polypeptides for heterologous expression in yeast allowing for improved fermentative bioproduction of these product compounds.
- CYP cytochrome P450
- the present disclosure provides a recombinant host cell comprising a heterologous nucleic acid encoding a polypeptide with CYP activity capable of converting LCA to UDCA and/or converting 3-KCA to 3-KUDCA, wherein the polypeptide with CYP activity comprises an amino acid sequence at least 90% sequence identity to SEQ ID NO: 2 and an amino acid difference relative to SEQ ID NO: 2 at one or more positions selected from L7, S31, Q54, V56, K70, K76, I120, A143, L147, A153, N160, H170, G175, E176, V210, V261, T307, N316, Q320, G325, T333, E338, V364, L365, P401, E402, S408, Q424, L444, F452, K479, and M485.
- the polypeptide with CYP activity comprises an amino acid sequence at least 90% sequence identity to SEQ ID NO: 2 and an amino acid difference relative to SEQ ID NO: 2 at one or more
- the amino acid differences are selected from L7R, S31V, Q54K, V56G, K70L, K70N, K76H, I120A, A143G, A143R, L147R, L147W, A153Q, A153R, N160H, H170V, G175E, E176D, V210R, V261E, T307A, N316S, Q320M, G325E, T333V, E338Q, V364L, L365A, P401D, E402V, S408R, Q424V, L444F, F452D, K479R, and M485T.
- the polypeptide amino acid sequence comprises a combination of amino acid differences selected from: S31V, Q54K, A143G, T307A, G325E ⁇ ⁇ S31V, A143G, G175E, P401D, K498R S31V, A143G, L147C, G175E, P401D 1V P12 A14 R [0009]
- the heterologous nucleic acid encoding the polypeptide further comprises a silent mutation selected from R328 (CGT > CGC), A368 (GCT > GCC), and V476 (GTT > GTC).
- the heterologous nucleic acid comprises: (a) a sequence of at least 80% identity to a sequence selected from the group consisting of SEQ ID NO: 5, 7, 9, 11, 13, 15, 17, 19, 21, 23, 25, 27, 29, 31, 33, 35, 37, 39, 41, 43, 45, 47, 49, 51, 53, 55, 57, 1185, 1187, 1189, 1191, 1193, 1195, 1197, 1199, 1201, 1203, ⁇ 3 ⁇ 1205, 1207, 1209, 1211, 1213, 1215, 1217, 1219, 1243, 1245, 1247, 1249, 1251, 1253, 1255, 1257, 1259, 1261, 1263, 1265, 1267, 1269, 1271, and 1275; or (b) a codon degenerate sequence of a sequence selected from the group consisting of SEQ ID NO: 5, 7, 9, 11, 13, 15, 17, 19, 21, 23, 25, 27, 29, 31, 33, 35, 37, 39, 41
- the polypeptide with CYP activity comprises an amino acid sequence having at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% sequence identity to SEQ ID NO: 2.
- the polypeptide with CYP activity comprises an amino acid sequence selected from SEQ ID NO: 6, 8, 10, 12, 14, 16, 18, 20, 22, 24, 26, 28, 30, 32, 34, 36, 38, 40, 42, 44, 46, 48, 50, 52, 54, 56, 58, 1186, 1188, 1190, 1192, 1194, 1196, 1198, 1200, 1202, 1204, 1206, 1208, 1210, 1212, 1214, 1216, 1218, 1220, 1244, 1246, 1248, 1250, 1252, 1254, 1256, 1258, 1260, 1262, 1264, 1266, 1268, 1270, and 1272; optionally, wherein the polypeptide is truncated by 1-31 amino acids at its N-terminus.
- the heterologous nucleic acid encodes a polypeptide with CPR activity; optionally, wherein the polypeptide with CPR activity comprises an amino acid sequence of at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% sequence identity to a sequence selected from SEQ ID NO: 4 and 1278.
- the heterologous nucleic acid encodes a polypeptide with CYP activity capable of converting LCA to UDCA and/or converting 3-KCA to 3-KUDCA fused via a linker to a second polypeptide second polypeptide with CPR activity; optionally, wherein the fusion polypeptide comprises an amino acid sequence of SEQ ID NO: 1280
- the heterologous nucleic acid is integrated into a site in the host cell genome. In at least one embodiment, the heterologous nucleic acid is integrated at one or more sites in the genome; optionally, integrated at three or more sites.
- the heterologous nucleic acid is under the control of a promoter system selected from pGal1/10, and pCAT1:pFDH1.
- the source organism of the recombinant host cell is selected from Saccharomyces cerevisiae, Pichia pastoris, Yarrowia lipolytica, and Escherichia coli.
- the recombinant host cell is ⁇ 4 ⁇ Saccharomyces cerevisiae and the heterologous nucleic acid is integrated at one or more sites in the genome selected from X-4, XI-2, XII-4, and X-2.
- the heterologous nucleic acid is integrated at three or more sites; optionally, integrated at four or more sites.
- the recombinant host cell is Pichia pastoris and the heterologous nucleic acid is integrated at one or more sites in the genome selected from AOX1, Int6, Int15, and HIS4.
- the heterologous nucleic acid is integrated at three or more sites; optionally, integrated at four or more sites.
- the present disclosure also provides a method for producing UDCA or 3-KUDCA comprising: (a) culturing a recombinant host cell of the present disclosure in a suitable medium comprising LCA and/or 3-KCA; and (b) recovering the produced UDCA and/or 3-KUDCA.
- the present disclosure also provides a recombinant polypeptide with CYP activity capable of converting LCA to UDCA and/or converting 3-KCA to 3-KUDCA, wherein the polypeptide comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 2 and an amino acid difference relative to SEQ ID NO: 2 at one or more positions selected from L7, S31, Q54, V56, K70, K76, I120, A143, L147, A153, N160, H170, G175, E176, V210, V261, T307, N316, Q320, G325, T333, E338, V364, L365, P401, E402, S408, Q424, L444, F452, K479, and M485.
- the polypeptide comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 2 and an amino acid difference relative to SEQ ID NO: 2 at one or more positions selected from L7, S31, Q54, V56, K70, K76, I120, A
- the amino acid differences are selected from L7R, S31V, Q54K, V56G, K70L, K70N, K76H, I120A, A143G, A143R, L147R, L147W, A153Q, A153R, N160H, H170V, G175E, E176D, V210R, V261E, T307A, N316S, Q320M, G325E, T333V, E338Q, V364L, L365A, P401D, E402V, S408R, Q424V, L444F, F452D, K479R, and M485T.
- the polypeptide amino acid sequence comprises a combination of amino acid differences selected from: S31V, Q54K, A143G, T307A, G325E S31V Q54K A143G T307A G325E L365A ⁇ 5 ⁇ S31V, Q54K, A143G, A153Q, T307A, G325E S31V, Q54K, A143G, A153P, T307A, G325E S31V Q54K A143G T307A G325E Q424V [ p yp p , p yp p p o acid sequence having at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at
- the polypeptide comprises an amino acid sequence selected from SEQ ID NO: 6, 8, 10, 12, 14, 16, 18, 20, 22, 24, 26, 28, 30, 32, 34, 36, 38, 40, 42, 44, 46, 48, 50, 52, 54, 56, 58, 1186, 1188, 1190, 1192, 1194, 1196, 1198, 1200, 1202, 1204, 1206, 1208, 1210, 1212, 1214, 1216, 1218, 1220, 1244, 1246, 1248, 1250, 1252, 1254, 1256, 1258, 1260, 1262, 1264, 1266, 1268, 1270, 1272, and 1276.
- the recombinant polypeptide with CYP activity capable of converting LCA to UDCA and/or converting 3-KCA to 3-KUDCA of the present disclosure comprises an N-terminal truncation of from 1 to 31 amino acids as compared to SEQ ID NO: 2.
- the polypeptide comprises an N-terminal truncation of from 1 to 31 amino acids of an amino acid sequence selected from SEQ ID NO: 6, 8, 10, 12, 14, 16, 18, 20, 22, 24, 26, 28, 30, 32, 34, 36, 38, 40, 42, 44, 46, 48, 50, 52, 54, 56, 58, 1186, 1188, 1190, ⁇ 6 ⁇ 1192, 1194, 1196, 1198, 1200, 1202, 1204, 1206, 1208, 1210, 1212, 1214, 1216, 1218, 1220, 1244, 1246, 1248, 1250, 1252, 1254, 1256, 1258, 1260, 1262, 1264, 1266, 1268, 1270, and 1272.
- the polypeptide comprises an amino acid of SEQ ID NO: 1276.
- the recombinant polypeptide with CYP activity capable of converting LCA to UDCA and/or converting 3-KCA to 3-KUDCA of the present disclosure is fused via a linker to a second polypeptide.
- the second polypeptide has CPR activity.
- the second polypeptide with CPR activity comprises an amino acid sequence having at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% sequence identity to a sequence selected from SEQ ID NO: 4 and 1278.
- the recombinant polypeptide with CYP activity capable of converting LCA to UDCA and/or converting 3-KCA to 3-KUDCA is fused via a linker to a second polypeptide second polypeptide with CPR activity; optionally, wherein the fusion polypeptide comprises an amino acid sequence of SEQ ID NO: 1280.
- the present disclosure also provides a polynucleotide encoding a polypeptide of the present disclosure.
- the polynucleotide sequence comprises: (a) a sequence of at least 80% identity to a sequence selected from the group consisting of SEQ ID NO: 5, 7, 9, 11, 13, 15, 17, 19, 21, 23, 25, 27, 29, 31, 33, 35, 37, 39, 41, 43, 45, 47, 49, 51, 53, 55, 57, 1185, 1187, 1189, 1191, 1193, 1195, 1197, 1199, 1201, 1203, 1205, 1207, 1209, 1211, 1213, 1215, 1217, 1219, 1243, 1245, 1247, 1249, 1251, 1253, 1255, 1257, 1259, 1261, 1263, 1265, 1267, 1269, 1271, and 1275; or (b) a codon degenerate sequence of a sequence selected from the group consisting of SEQ ID NO: 5, 7, 9, 11, 13, 15, 17, 19, 21, 23, 25, 27, 29, 31, 33, 35, 37, 39, 41, 43, 45, 47, 49, 51, 53, 55, 57
- the present disclosure also provides an expression vector comprising a polynucleotide of the present disclosure.
- the expression vector comprises a control sequence.
- the present disclosure also provides a host cell comprising a polynucleotide or an expression vector of the present disclosure.
- the present disclosure provides a method for preparing a compound of structural formula (I) ⁇ 7 ⁇ wherein, R 1 is hydrogen, substituted C1-C20 alkyl, or an optionally is a hydroxyl, or an oxo group; the method comprising contacting under suitable reactions conditions a compound of structural formula (II) wherein, R 1 is substituted C1-C20 alkyl, or an optionally or a is a hydroxyl, or an oxo group; and a recombinant polypeptide of the present disclosure.
- R 2 is hydrogen, and (a) the compound of structure formula (I) is UDCA and the compound of structural formula (II) is LCA; or (b) the compound of structure formula (I) is 3- KUDCA and the compound of structural formula (II) is 3-KCA.
- R 2 is a hydroxyl, and the compound of structure formula (I) is UCA and the compound of structural formula (II) is DCA.
- FIG.1 depicts Schemes A and B used for PCR primer synthesis of DNA fragments used in building strains of Saccharomyces cerevisiae with CPR and CYP genes under control of pGal1/10 bidirectional promoter as described in Example 1.
- UDCA refers to the compound ursodeoxycholic acid having the chemical structure shown as compound 1a in Table 1 (below).
- UDCA derivatives include, but are not limited to, structural analogs of UDCA, such as 7 ⁇ -hydroxy-3-oxo-5 ⁇ -cholanoic acid (3-KUDCA), compound 1b, and the other exemplary compounds shown below in Table 1 (below).
- 3-KUDCA 7 ⁇ -hydroxy-3-oxo-5 ⁇ -cholanoic acid
- UDCA precursor refers to a compound capable of being converted into UDCA, or a UDCA derivative, by an enzyme capable producing UDCA alone, or in combination with another enzyme or a non-enzymatic chemical reaction.
- UDCA precursors as referenced in the present disclosure include, but are not limited to, the exemplary compounds summarized in Table 2 (below).
- TABLE 2 Exemplary UDCA and UDCA derivative precursor compounds Abbrev. ⁇ 11 ⁇ Lithocholic acid LCA [ ] ac y as use ee ee s o e caay c ac y o a cyoc o e monooxygenase enzyme.
- CPR activity refers to the catalytic activity of a cytochrome P450 reductase enzyme.
- “Conversion” as used herein refers to the enzymatic conversion of a substrate(s) to a corresponding product(s). “Percent conversion” refers to the percent of the substrate that is converted to the product within a period of time under specified conditions. Thus, the “enzymatic activity” or “activity” of an enzymatic conversion can be expressed as “percent conversion” of the substrate to the product.
- “Substrate” as used herein in the context of an enzyme mediated process refers to the compound or molecule acted on by the enzyme.
- “Product” as used herein in the context of an enzyme mediated process refers to the compound or molecule resulting from the activity of the enzyme.
- “Host cell” as used herein refers to a cell capable of being functionally modified with recombinant nucleic acids and functioning to express recombinant products, including polypeptides and compounds produced by activity of the polypeptides.
- the nucleic acid may be wholly comprised ribonucleosides (e.g., RNA), wholly comprised of 2'-deoxyribonucleotides (e.g., DNA) or mixtures of ribo- and 2'-deoxyribonucleosides.
- nucleoside units of the nucleic acid can be linked together via phosphodiester linkages (e.g., as in naturally occurring nucleic acids), or the nucleic acid can include one or more non-natural linkages (e.g., phosphorothioester linkage).
- Nucleic acid or polynucleotide is intended to include single- stranded or double-stranded molecules, or molecules having both single-stranded regions and double-stranded regions.
- Nucleic acid or polynucleotide is intended to include molecules composed of the naturally occurring nucleobases (i.e., adenine, guanine, uracil, thymine, and cytosine), or molecules comprising that include one or more modified and/or synthetic nucleobases, such as, for example, inosine, xanthine, hypoxanthine, etc.
- nucleobases i.e., adenine, guanine, uracil, thymine, and cytosine
- Protein “Protein,” “polypeptide,” and “peptide” are used herein interchangeably to denote a polymer of at least two amino acids covalently linked by an amide bond, regardless of length or post-translational modification (e.g., glycosylation, phosphorylation, lipidation, myristilation, ubiquitination, etc.).
- protein or “polypeptide” or “peptide” polymer can include D- and L-amino acids, and mixtures of D- and L-amino acids.
- “Naturally-occurring” or “wild-type” as used herein refers to the form as found in nature.
- a naturally occurring nucleic acid sequence is the sequence present in an organism that can be isolated from a source in nature, and which has not been intentionally modified by human manipulation.
- “Recombinant,” “engineered,” or “non-naturally occurring” when used herein with reference to, e.g., a cell, nucleic acid, or polypeptide refers to a material, or a material corresponding to the natural or native form of the material, that has been modified in a manner that would not otherwise exist in nature, or is identical thereto but is produced or derived from synthetic materials and/or by manipulation using recombinant techniques.
- Non-limiting examples include, among others, recombinant cells expressing genes that are not found within the native (non-recombinant) form of the cell or express native genes that are otherwise expressed at a different level.
- “Nucleic acid derived from” as used herein refers to a nucleic acid having a sequence at least substantially identical to a sequence of found in naturally in an organism. For example, cDNA molecules prepared by reverse transcription of mRNA isolated from an organism, or nucleic acid molecules prepared synthetically to have a sequence at least substantially identical to, or which hybridizes to a sequence at least substantially identical to a nucleic sequence found in an organism.
- Coding sequence refers to that portion of a nucleic acid (e.g., a gene) that encodes an amino acid sequence of a protein.
- a nucleic acid e.g., a gene
- Heterologous nucleic acid refers to any polynucleotide that is introduced into a host cell by laboratory techniques and includes polynucleotides that are removed from a host cell, subjected to laboratory manipulation, and then reintroduced into a host cell.
- Codon degenerate describes a nucleotide sequence that has one or more different codons relative to the reference nucleotide sequence, but which encodes a polypeptide that is identical to the polypeptide encoded by a reference nucleotide sequence.
- the different codons between the nucleotide sequence and the reference nucleotide sequence are called “synonyms” or “synonymous” codons in that they use different triplets of nucleotides to encode the same amino acid in a polypeptide.
- Codon optimized refers to changes in the codons of the polynucleotide encoding a protein to those preferentially used in a particular organism such that the encoded protein is efficiently expressed in the organism of interest.
- the genetic code is degenerate in that most amino acids are represented by several different “synonymous” codons, it is well known that codon usage by particular organisms is nonrandom and biased towards particular codon triplets. This codon usage bias may be higher in reference to a given gene, genes of common function or ancestral origin, highly expressed proteins versus low copy number proteins, and the aggregate protein coding regions of an organism's genome.
- the polynucleotides encoding the imine reductase enzymes may be codon optimized for optimal production from the host organism selected for expression.
- “Preferred, optimal, high codon usage bias codons” refers to codons that are used at higher frequency in the protein coding regions than other codons that code for the same amino acid.
- the preferred codons may be determined in relation to codon usage in a single gene, a set of genes of common function or origin, highly expressed genes, the codon frequency in the aggregate protein coding regions of the whole organism, codon frequency in the aggregate protein coding regions of related organisms, or combinations thereof. Codons whose frequency increases with the level of gene expression are typically optimal codons for expression.
- codon frequency e.g., codon usage, relative synonymous codon usage
- codon preference in specific organisms, including multivariate analysis, for example, using cluster analysis or correspondence analysis, and the effective number of codons used in a gene (see GCG CodonPreference, Genetics Computer Group Wisconsin Package; CodonW, John Peden, University of Nottingham; McInerney, J. O, 1998, Bioinformatics 14:372-73; Stenico et al., 1994, Nucleic Acids Res.222437-46; Wright, F., 1990, Gene 87:23-29).
- Codon usage tables are available for a growing list of organisms (see for example, Wada et al., 1992, Nucleic Acids Res.20:2111-2118; Nakamura et al., 2000, Nucl. Acids Res.28:292; Duret, et al., supra; Henaut and Danchin, "Escherichia coli and Salmonella,” 1996, Neidhardt, et al. Eds., ASM Press, Washington D.C., p.2047-2066.
- the data source for obtaining codon usage may rely on any available nucleotide sequence capable of coding for a ⁇ 14 ⁇ protein.
- nucleic acid sequences actually known to encode expressed proteins e.g., complete protein coding sequences-CDS
- expressed sequence tags e.g., expressed sequence tags
- genomic sequences see for example, Mount, D., Bioinformatics: Sequence and Genome Analysis, Chapter 8, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, N.Y., 2001; Uberbacher, E. C., 1996, Methods Enzymol.266:259-281; Tiwari et al., 1997, Comput. Appl. Biosci.13:263-270).
- Control sequence refers to all sequences, which are necessary or advantageous for the expression of a polynucleotide and/or polypeptide as used in the present disclosure.
- Each control sequence may be native or foreign to the nucleic acid sequence encoding a polypeptide.
- control sequences include, but are not limited to, a leader, a promoter, a polyadenylation sequence, a pro-peptide sequence, a signal peptide sequence, and a transcription terminator.
- control sequences typically include a promoter, and transcriptional and translational stop signals.
- control sequences may be provided with linkers for the purpose of introducing specific restriction sites facilitating ligation of the control sequences with the coding region of the nucleic acid sequence encoding a polypeptide.
- “Operably linked” as used herein refers to a configuration in which a control sequence is appropriately placed (e.g., in a functional relationship) at a position relative to a polynucleotide sequence or polypeptide sequence of interest such that the control sequence directs or regulates the expression of the sequence of interest.
- Promoter sequence refers to a nucleic acid sequence that is recognized by a host cell for expression of a polynucleotide of interest, such as a coding sequence.
- the promoter sequence contains transcriptional control sequences, which mediate the expression of a polynucleotide of interest.
- the promoter may be any nucleic acid sequence which shows transcriptional activity in the host cell of choice including mutant, truncated, and hybrid promoters, and may be obtained from genes encoding extracellular or intracellular polypeptides either homologous or heterologous to the host cell.
- Percentage of sequence identity “percent sequence identity,” “percentage homology,” or “percent homology” are used interchangeably herein to refer to values quantifying comparisons of the sequences of polynucleotides or polypeptides, and are determined by comparing two optimally aligned sequences over a comparison window, wherein the portion of the polynucleotide or polypeptide sequence in the comparison window may comprise additions or deletions (or gaps) as compared to the reference sequence for optimal alignment of the two sequences.
- the percentage values may be calculated by determining the number of positions at which the identical nucleic acid base or amino acid residue occurs in both sequences to yield the number of matched positions, dividing the number of matched positions by the total number of positions in the window of comparison and multiplying the result by 100 to yield the percentage of sequence identity.
- the percentage may be calculated by determining the number of positions at which either the identical nucleic acid base or amino ⁇ 15 ⁇ acid residue occurs in both sequences or a nucleic acid base or amino acid residue is aligned with a gap to yield the number of matched positions, dividing the number of matched positions by the total number of positions in the window of comparison and multiplying the result by 100 to yield the percentage of sequence identity.
- This algorithm involves first identifying high scoring sequence pairs (HSPs) by identifying short words of length W in the query sequence, which either match or satisfy some positive-valued threshold score T when aligned with a word of the same length in a database sequence.
- T is referred to as, the neighborhood word score threshold (Altschul et al, supra).
- M forward score for a pair of matching residues; always >0
- N penalty score for mismatching residues; always ⁇ 0).
- a scoring matrix is used to calculate the cumulative score. Extension of the word hits in each direction are halted when: the cumulative alignment score falls off by the quantity X from its maximum achieved value; the cumulative score goes to zero or below, due to the accumulation of one or more negative- scoring residue alignments; or the end of either sequence is reached.
- the BLAST algorithm parameters W, T, and X determine the sensitivity and speed of the alignment.
- the BLASTP program uses as defaults a wordlength (W) of 3, an expectation (E) of 10, and the BLOSUM62 scoring matrix (see Henikoff and Henikoff, 1989, Proc Natl Acad Sci USA 89:10915).
- W wordlength
- E expectation
- BLOSUM62 scoring matrix see Henikoff and Henikoff, 1989, Proc Natl Acad Sci USA 89:10915.
- Exemplary determination of sequence alignment and % sequence identity can ⁇ 16 ⁇ employ the BESTFIT or GAP programs in the GCG Wisconsin Software package (Accelrys, Madison Wis.), using default parameters provided.
- “Reference sequence” refers to a defined sequence used as a basis for a sequence comparison.
- a reference sequence may be a subset of a larger sequence, for example, a segment of a full-length nucleic acid or polypeptide sequence.
- a reference sequence typically is at least 20 nucleotide or amino acid residue units in length but can also be the full length of the nucleic acid or polypeptide. Since two polynucleotides or polypeptides may each (1) comprise a sequence (i.e., a portion of the complete sequence) that is similar between the two sequences, and (2) may further comprise a sequence that is divergent between the two sequences, sequence comparisons between two (or more) polynucleotides or polypeptide are typically performed by comparing sequences of the two polynucleotides or polypeptides over a “comparison window” to identify and compare local regions of sequence similarity.
- Comparison window refers to a conceptual segment of at least about 20 contiguous nucleotide positions or amino acids residues wherein a sequence may be compared to a reference sequence of at least 20 contiguous nucleotides or amino acids and wherein the portion of the sequence in the comparison window may comprise additions or deletions (or gaps) of 20 percent or less as compared to the reference sequence (which does not comprise additions or deletions) for optimal alignment of the two sequences.
- “Substantial identity” or “substantially identical” refers to a polynucleotide or polypeptide sequence that has at least 70% sequence identity, at least 80% sequence identity, at least 85% sequence identity, at least 90% sequence identity, at least 95 % sequence identity, or at least 99% sequence identity, as compared to a reference sequence over a comparison window of at least 20 nucleoside or amino acid residue positions, frequently over a window of at least 30-50 positions, wherein the percentage of sequence identity is calculated by comparing the reference sequence to a sequence that includes deletions or additions which total 20 percent or less of the reference sequence over the window of comparison.
- “Corresponding to,” “reference to,” or “relative to” when used in the context of the numbering of a given amino acid or polynucleotide sequence refers to the numbering of the residues of a specified reference sequence when the given amino acid or polynucleotide sequence is compared to the reference sequence.
- the residue number or residue position of a given polymer is designated with respect to the reference sequence rather than by the actual numerical position of the residue within the given amino acid or polynucleotide sequence.
- a given amino acid sequence such as that of an engineered imine reductase, can be aligned to a reference sequence by introducing gaps to optimize residue matches between the two sequences.
- isolated as used herein in reference to a molecule means that the molecule (e.g., UDCA, polynucleotide, polypeptide) is substantially separated from other compounds that naturally accompany it, e.g., protein, lipids, and polynucleotides.
- the term embraces nucleic acids which have been removed or purified from their naturally occurring environment or expression system (e.g., host cell or in vitro synthesis).
- substantially pure refers to a composition in which a desired molecule is the predominant species present (i.e., on a molar or weight basis it is more abundant than any other individual macromolecular species in the composition) and is generally a substantially purified composition when the object species comprises at least about 50 percent of the macromolecular species present by mole or % weight.
- “Recovered” as used herein in relation to an enzyme, protein, or compound refers to a more or less pure form of the enzyme, protein, or compound.
- Engineered Genes Encoding Recombinant Polypeptides with CYP Activity provides engineered genes encoding a polypeptide with cytochrome P450 (CYP) monooxygenase activity capable of converting LCA to UDCA and/or capable of converting 3-KCA to 3-KUDCA.
- the engineered CYP genes are derived from a naturally occurring parent CYP gene isolated from Gibberella zeae that encodes the 513 amino acid CYP polypeptide of SEQ ID NO: 2.
- the 7-position hydroxylating monooxygenase activity of the recombinant CYP polypeptides encoded by the engineered genes of the present disclosure is capable of catalyzing the conversion of LCA (compound 2a) to UDCA (compound 1a) as shown by the upper reaction illustrated in Scheme 1 (below).
- Scheme 1 [0080]
- the CYP activity of the polypeptides encoded by the engineered genes of the present disclosure is capable of catalyzing the conversion of 3-KCA (compound 2b) to 3- KUDCA (compound 1b) as shown by the lower reaction illustrated in Scheme 1.
- the 3-KUDCA product can then be converted to UDCA by a further keto reduction reaction carried out by a ketoreductase (KRED) enzyme.
- KRED ketoreductase
- the CYP activity of the polypeptides encoded by the engineered genes of the present disclosure is capable of catalyzing the 7-position hydroxylation of DCA (compound ⁇ 19 ⁇ 2c) to UCA (compound 1e) and/or CA (compound 1f) as shown by the reaction illustrated in Scheme 2.
- Scheme 2 DCA to UCA/CA conversion of Scheme 2) can be provided by the 264 amino acid CPR from Gibberella zeae of SEQ ID NO: 4.
- the 264 amino acid CPR polypeptide of SEQ ID NO: 4 can be encoded by a yeast codon-optimized 795 nucleotide sequence of SEQ ID NO: 3.
- the SAND122 strain is capable of converting LCA to UDCA and/or converting 3-KCA to 3- KUDCA.
- a host cell e.g., S. cerevisiae or P. pastoris
- a heterologous nucleic acid comprising an engineered CYP gene of the present disclosure
- a gene encoding a CPR polypeptide e.g., SEQ ID NO: 4
- the product UDCA is produced by the host cell in greater yield relative to a comparable recombinant host cell integrated with the parent gene encoding the wild-type CYP polypeptide of SEQ ID NO: 2.
- the enhanced yield of the UDCA and/or 3-KUDCA biosynthetic product is correlated with the one or more amino acid residue differences in recombinant polypeptides of the present disclosure, as compared to the amino acid sequence of wild-type CYP from G. zeae of SEQ ID NO: 2 from which the engineered polypeptide sequences are derived.
- Exemplary engineered CYP genes and encoded recombinant polypeptides with CYP activity that exhibit the unexpected and surprising technical effect of comparable or increased UDCA yield when integrated in a recombinant host cell are summarized in Table 3 below (as well as in the following Examples and the accompanying Sequence Listing).
- the recombinant polypeptides have one or more amino acid residue differences as compared to SEQ ID NO: 2 at an amino acid position selected from L7, S31, Q54, V56, K70, K76, I120, A143, L147, A153, N160, H170, G175, E176, V210, V261, T307, N316, Q320, G325, T333, E338, V364, L365, P401, E402, S408, Q424, L444, F452, K479, and M485.
- the recombinant polypeptides have one or more amino acid residue differences as compared to SEQ ID NO: 2 selected from L7R, S31V, Q54K, V56G, K70L, K70N, K76H, I120A, A143G, A143R, L147R, L147W, A153Q, A153R, N160H, H170V, G175E, E176D, V210R, V261E, T307A, N316S, Q320M, G325E, T333V, E338Q, V364L, L365A, P401D, E402V, S408R, Q424V, L444F, F452D, K479R, and M485T.
- SEQ ID NO: 2 selected from L7R, S31V, Q54K, V56G, K70L, K70N, K76H, I120A, A143G, A143R, L147R, L147W, A153Q, A153R
- residue differences relative to SEQ ID NO: 2 at residue positions associated with increased CYP activity can be used in various combinations to form recombinant CYP polypeptides having desirable functional characteristics when integrated in a recombinant host cell, for example increased yield of a product compound, such as UDCA.
- Some exemplary combinations of amino acid differences include those combinations found in the exemplary polypeptides of Table 3 and elsewhere herein.
- the present disclosure provides a recombinant polypeptide having increased CYP activity, a set of amino acid residue differences as compared to SEQ ID NO: 2 selected from: S31V, Q54K, A143G, T307A, G325E ⁇ 37 ⁇ S31V, Q54K, A143R, L147R, T307A, G325E, P401D S31V, V57F, L147W, P401D S31V L147W P401D [0091]
- the present disclosure provides a range of recombinant polypeptides having CYP activity, wherein the polypeptide comprises an amino acid sequence comprising one or more of the amino acid differences or sets of amino acid differences (relative to SEQ ID NO: 2) disclosed in any one of SEQ ID NO: 6, 8, 10, 12, 14, 16, 18,
- a recombinant polypeptide of the present disclosure having CYP activity can have an amino acid sequence comprising one or more of the amino acid differences or sets of amino acid differences (relative to SEQ ID NO: 2) disclosed in any one of SEQ ID NO: 6, 8, 10, 12, 14, 16, 18, 20, 22, 24, 26, 28, 30, 32, 34, 36, 38, 40, 42, 44, 46, 48, 50, 52, 54, 56, 58, 1186, 1188, 1190, 1192, 1194, 1196, 1198, 1200, 1202, 1204, 1206, 1208, 1210, 1212, 1214, 1216, 1218, 1220, 1244, 1246, 1248, 1250, 1252, 1254, 1256, 1258, 1260, 1262, 1264, 1266, 1268, 1270, 1272, and 1276, and additionally have 1-2, 1-3, 1-4, 1-5, 1-6, 1-7, 1-8, 1-9, 1-10, 1-11, 1-12, 1-14, 1
- the number of differences can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 14, 15, 16, 18, 20, 22, 24, 26, 30, 35, 40, 45, 50, 55, or 60 residue differences at the other residue positions.
- any of the engineered CYP polypeptides disclosed herein can further comprise other residue differences relative to the reference polypeptide of SEQ ID NO: 2 at other residue positions.
- Residue differences at these other residue positions can provide for additional variations in the amino acid sequence without adversely affecting the ability of the recombinant polypeptide to carry out the desired biocatalytic conversion (e.g., conversion of LCA (compound 2a) to UDCA (compound 1a)).
- the recombinant polypeptides can have additionally 1-2, 1-3, 1-4, 1-5, 1-6, 1-7, 1-8, 1-9, 1-10, 1-11, 1-12, 1-14, 1-15, 1-16, 1-18, 1-20, 1-22, 1-24, 1-26, 1-30, 1-35, 1-40 residue differences at other amino acid residue positions as compared to SEQ ID NO: 2.
- the number of differences can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 14, 15, 16, 18, 20, 22, 24, 26, 30, 35, and 40 residue differences at other residue positions.
- the residue difference at these other positions can include conservative changes or non-conservative changes.
- the residue differences can comprise conservative substitutions and non-conservative substitutions as compared to the reference polypeptide of SEQ ID NO: 2.
- Other engineered modifications of the CYP polypeptides contemplated by the present disclosure include modification of the amino acid sequence at either its N- or C- terminus by truncation or fusion.
- the CYP polypeptides (or CPR polypeptides) of the present disclosure can be further engineered by truncation at the N- and/or C-terminus of the polypeptide chain to provide a truncated version of the polypeptide that retains CYP activity.
- trCYP207 This truncated version of CYP207, referred to as trCYP207 (SEQ ID NO: 1276) retains CYP activity a gene encoding it is heterologously expressed in a recombinant Pichia host cell. Accordingly, in at least one embodiment any of the engineered CYP polypeptides disclosed herein can include an engineered version of the polypeptide (and the encoding gene), wherein from 1-31 N-terminal amino acids are truncated.
- such a recombinant polypeptide with an N-terminal truncation and CYP activity capable of converting LCA to UDCA and/or converting 3-KCA to 3- KUDCA comprises an N-terminal truncation of 1 to 31 amino acids as compared to SEQ ID NO: 2.
- polypeptide with CYP activity can comprise an N-terminal truncation of from 1 to 31 amino acids of engineered recombinant polypeptide comprising a sequence of any of SEQ ID NO: 6, 8, 10, 12, 14, 16, 18, 20, 22, 24, 26, 28, 30, 32, 34, 36, 38, 40, 42, 44, 46, 48, 50, 52, 54, 56, 58, 1186, 1188, 1190, 1192, 1194, 1196, 1198, 1200, 1202, 1204, 1206, 1208, 1210, 1212, 1214, 1216, 1218, 1220, 1244, 1246, 1248, 1250, 1252, 1254, 1256, 1258, 1260, 1262, 1264, 1266, 1268, 1270, and 1272.
- the N-terminal truncated polypeptide with CYP activity comprises an amino acid of SEQ ID NO: 1276.
- the engineered CYP (and CPR) polypeptides of the present disclosure can be modified via fusion to other polypeptides.
- an engineered recombinant gene encoding a polypeptide having CYP activity of the present disclosure is fused to a gene encoding a recombinant polypeptide having CPR activity, whereby the fusion gene when expressed heterologously in a recombinant host cell expression a fusion polypeptide having both CYP activity and CPR activity.
- engineered CYP-CPR fusion polypeptides can be used in the biosynthetic processes for preparing UDCA.
- the recombinant polypeptide with CYP activity is fused via a polypeptide linker to a second polypeptide with CPR activity.
- the polypeptide with CYP activity used in the fusion is an 1-31 amino acid N-terminal truncated version of an engineered polypeptide comprising a sequence of at least 80%, 90%, 95%, or 99% identity to any of SEQ ID NO: 6, 8, 10, 12, 14, 16, 18, 20, 22, 24, 26, 28, 30, 32, 34, 36, 38, 40, 42, 44, 46, 48, 50, 52, 54, 56, 58, 1186, 1188, 1190, 1192, 1194, 1196, 1198, 1200, 1202, 1204, 1206, 1208, 1210, 1212, 1214, 1216, 1218, 1220, 1244, 1246, 1248, 1250, 1252, 1254, 1256, 1258, 1260, 1262, 1264, 1266, 1268, 1270, and 1272.
- the second polypeptide with CPR activity comprises an amino acid sequence having at least 80%, 90%, 95%, or 99% sequence identity to a sequence selected from SEQ ID NO: 4 and 1278.
- the recombinant fusion polypeptide with CYP activity and CPR activity is capable of converting LCA to UDCA and/or converting 3-KCA to 3-KUDCA is fused and comprises an amino acid sequence of at least 80%, 90%, 95%, or 99% sequence identity to SEQ ID NO: 1280.
- the engineered polypeptides of the present disclosure can be fused to polypeptides such as antibody tags (e.g., myc epitope), purification sequences (e.g., His tags for binding to metals), and cell localization signals (e.g., secretion signals).
- polypeptides such as antibody tags (e.g., myc epitope), purification sequences (e.g., His tags for binding to metals), and cell localization signals (e.g., secretion signals).
- the recombinant polypeptides described herein are not restricted to the genetically encoded amino acids.
- the polypeptides described herein may be comprised, either in whole or in part, of naturally occurring and/or synthetic non-encoded amino acids.
- the present disclosure provides polynucleotides encoding the recombinant polypeptides having CYP activity and increased activity and/or yield as described herein.
- the polynucleotide encoding a recombinant polypeptide having CYP activity comprises an amino acid sequence that is at least about 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more identical to the polypeptide sequence of SEQ ID NO: 2.
- the polynucleotide encodes a recombinant polypeptide comprising an amino acid sequence that has the percent identity described above and has one or more amino acid residue differences as compared to SEQ ID NO: 2 described elsewhere herein.
- the polynucleotide has a sequence encoding a recombinant polypeptide which polynucleotide sequence has one or more neutral codon differences relative to SEQ ID NO: 1, which codon differences do not encode an amino acid difference but result in increased yield of the desired product compound (e.g., UDCA) produced by a recombinant host cell in which the polynucleotide sequence is integrated.
- the desired product compound e.g., UDCA
- the polypeptide is encoded by a polynucleotide sequence having at least 80% identity to SEQ ID NO: 1, and at least one neutral codon difference as compared to SEQ ID NO: 1 at a position encoding an amino acid residue selected from R328, A368, and V476; optionally, wherein the ⁇ 41 ⁇ neutral codon difference is selected from: R328 (CGT > CGC), A368 (GCT > GCC), and V476 (GTT > GTC).
- the polynucleotides encoding the recombinant polypeptides having CYP activity and increased activity and/or yield as described herein can include a combination of one or more codon differences relative to SEQ ID NO: 1, wherein at least one of the codon differences encodes an amino acid difference as compared to SEQ ID NO: 2 and at least one codon difference is a neutral codon difference that does not encode an amino acid difference as compared to SEQ ID NO: 2 Accordingly, in at least one embodiment, the present disclosure provides a polynucleotide sequence encoding a recombinant polypeptide having CYP activity, wherein the polynucleotide sequence comprises a combination of a codon differences encoding an amino acid difference and a neutral codon difference selected from: R328 (CGT > CGC), A368 (GCT > GCC), and V476 (GTT > GTC).
- the polynucleotide comprises a sequence encoding an exemplary recombinant polypeptide having CYP activity as disclosed in Table 3 and the accompanying Sequence Listing.
- the polynucleotide comprises a sequence of at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, or at least 99% identity to a sequence selected from the group consisting of SEQ ID NO: 5, 7, 9, 11, 13, 15, 17, 19, 21, 23, 25, 27, 29, 31, 33, 35, 37, 39, 41, 43, 45, 47, 49, 51, 53, 55, 57, 1185, 1187, 1189, 1191, 1193, 1195, 1197, 1199, 1201, 1203, 1205, 1207, 1209, 1211, 1213, 1215, 1217, 1219, 1243, 1245, 1247, 1249, 1251, 1253, 1255, 1257, 1259, 1261, 1263, 1265, 1267,
- the polynucleotide comprises a codon degenerate sequence of a sequence selected from the group consisting of SEQ ID NO: 5, 7, 9, 11, 13, 15, 17, 19, 21, 23, 25, 27, 29, 31, 33, 35, 37, 39, 41, 43, 45, 47, 49, 51, 53, 55, 57, 1185, 1187, 1189, 1191, 1193, 1195, 1197, 1199, 1201, 1203, 1205, 1207, 1209, 1211, 1213, 1215, 1217, 1219, 1243, 1245, 1247, 1249, 1251, 1253, 1255, 1257, 1259, 1261, 1263, 1265, 1267, 1269, 1271, and 1275.
- the polynucleotide sequences encoding the recombinant polypeptides of the present disclosure may be operatively linked to one or more heterologous regulatory sequences that control gene expression to create a recombinant polynucleotide capable of expressing the polypeptide.
- Expression constructs containing a heterologous polynucleotide encoding the recombinant polypeptide can be introduced into appropriate host cells to express the corresponding polypeptide. Because of the knowledge of the codons corresponding to the various amino acids, availability of a protein sequence provides a description of all the polynucleotides capable of encoding the subject.
- the present disclosure specifically contemplates each and every possible variation of polynucleotides that could be made by selecting combinations based on the possible codon choices, and all such variations are to be considered specifically disclosed for any polypeptide disclosed herein, including the amino acid sequences presented in Table 3 and the accompanying Sequence Listing.
- the codons can be selected to fit the host cell in which the protein is being produced. For example, preferred codons used in bacteria are used to express the gene in bacteria; preferred codons used in yeast are used for expression in yeast; and preferred codons used in mammals are used for expression in mammalian cells.
- codon optimized polynucleotides encoding the recombinant polypeptide may contain preferred codons at about 40%, 50%, 60%, 70%, 80%, or greater than 90% of codon positions of the full-length coding region.
- the present disclosure also provides an expression vector comprising a polynucleotide encoding a recombinant polypeptide having CYP activity and increased thermostability, and one or more expression regulating regions such as a promoter, a terminator, a replication origin, or the like, depending on the type of hosts into which they are to be introduced.
- the various nucleic acid and control sequences described above may be joined together to produce a recombinant expression vector which may include one or more convenient restriction sites to allow for insertion or substitution of the nucleic acid sequence encoding the recombinant polypeptide at such sites.
- a polynucleotide sequence of the present disclosure may be expressed by inserting the nucleic acid sequence or a nucleic acid construct comprising the sequence into an appropriate vector for expression.
- the coding sequence is located in the vector so that the coding sequence is operably linked with the appropriate control sequences for expression.
- the recombinant expression vector may be any vector (e.g., a plasmid or virus), which can be conveniently subjected to recombinant DNA procedures and can bring about the expression of the polynucleotide sequence.
- the choice of the vector will typically depend on the compatibility of the vector with the host cell into which the vector is to be introduced.
- the vectors may be linear or closed circular plasmids.
- the expression vector may be an autonomously replicating vector, i.e., a vector that exists as an extrachromosomal entity, the replication of which is independent of chromosomal replication, e.g., a plasmid, an extrachromosomal element, a mini-chromosome, or an artificial chromosome.
- the vector may contain any means for assuring self-replication.
- the vector may be one which, when introduced into the host cell, is integrated into the genome, and replicated together with the chromosome(s) into which it has been integrated.
- the expression vector further comprises one or more selectable markers, which permit easy selection of transformed cells.
- the engineered genes of the present disclosure that encode recombinant polypeptides with CYP activity can be incorporated in recombinant host cells to enable in vivo biosynthesis of compounds cholic acid compounds, such as UDCA and UDCA derivative compounds, that require a cytochrome P450 monooxygenase catalyzed reaction.
- These recombinant host cells comprise a polynucleotide or expression vector that encodes the recombinant polypeptide with CYP activity, wherein the polynucleotide is operatively linked to one or more control sequences for expression of the polypeptide in the host cell.
- Host cells for use in expressing recombinant genes encoding the polypeptides with CYP activity of the present disclosure are well known in the art and include but are not limited to, bacterial cells, such as E.
- coli or fungal cells, such as Saccharomyces cerevisiae or Pichia pastoris, insect cells, such as Drosophila S2 and Spodoptera Sf9, animal cells, such as CHO, COS, BHK, 293, and plant cells.
- insect cells such as Drosophila S2 and Spodoptera Sf9
- animal cells such as CHO, COS, BHK, 293, and plant cells.
- Appropriate mediums and growth conditions for culturing the recombinant host cells so that they express the polypeptide with CYP activity are well known in the art.
- the recombinant host cells can comprise heterologous nucleic acids encoding not only polypeptides with CYP and CPR activity capable of converting LCA to UDCA (or 3-KCA to 3- KUDCA) but also other enzymes capable of producing other precursor compounds, such as precursors for the substrates, LCA or 3-KCA.
- nucleic acid sequences encoding pathway enzymes for producing such precursor compounds are known in the art and can readily be used in accordance with the present disclosure.
- the nucleic acid sequence encoding the enzymes which form a part of the pathway further include one or more additional nucleic acid sequences, for example, a nucleic acid sequence controlling expression of the enzymes which form a part of the biosynthetic pathway, and these one or more additional nucleic acid sequences together with the nucleic acid sequence encoding the polypeptides with CPR and/or CYP activity can be considered a heterologous nucleic acid sequence.
- a variety of techniques and methodologies are available and well known in the art for introducing heterologous nucleic acid sequences, such as nucleic acid sequences encoding the enzymes (e.g., CPR and CYP), into a host cell so as to attain expression the host cell.
- heterologous nucleic acids can include integration of the nucleic acids into specific loci in the genome of a host cell via CRISPR-Cas9 and other techniques, some of which are demonstrated in the Examples herein.
- CRISPR-Cas9 CRISPR-Cas9
- heterologous genes and their locus of integration in a recombinant host cell’s genome can result in improved biosynthetic production of a desired product, such as UDCA.
- a heterologous nucleic acid encoding a polypeptide with CYP activity can be integrated in a host cell’s genome in 1, 2, 3, 4, or more copies.
- the heterologous nucleic acid encoding the recombinant polypeptide having CYP activity can be integrated in the host cell’s genome at one or more loci, including but not limited to the well- known genomic loci in Saccharomyces cerevisiae of X-2, X-4, XI-2, XII-4, NDE1, XII-5, Gal80, and ROQ1, or the loci in Pichia pastoris of AOX1, Int6, Int15, and HIS4.
- the heterologous nucleic acids encoding the recombinant enzymes with CPR activity and CYP activity, and any other pathway enzymes will further comprise transcriptional promoters capable of controlling expression of the enzymes in the recombinant host cell.
- the transcriptional promoters are selected to be compatible with the host cell, so that promoters obtained from bacterial cells are used when a bacterial host cell is selected in accordance herewith, while a fungal promoter is used when a fungal host cell is selected, a plant promoter is used when a plant cell is selected, and so on.
- Promoters useful in the recombinant host cells of the present disclosure may be constitutive or inducible, provided such promoters are operable in the host cells. Promoters that may be used to control expression in fungal host cells, such as Saccharomyces cerevisiae and Pichia pastoris, are well known in the art and include, but are not limited to inducible promoters, such as a Gal1 promoter or Gal10 promoter, a constitutive promoter, such as an alcohol dehydrogenase (ADH) promoter, a glyceraldehyde-3-phosphate dehydrogenase (GPD) promoter, or an S. pombe Nmt, or ADH promoter.
- inducible promoters such as a Gal1 promoter or Gal10 promoter
- a constitutive promoter such as an alcohol dehydrogenase (ADH) promoter, a glyceraldehyde-3-phosphate dehydrogenase (GPD) promoter, or an S.
- Exemplary promoters that may be used to control expression in bacterial cells can include the Escherichia coli promoters lac, tac, trc, trp or the T7 promoter.
- Exemplary promoters that may be used to control expression in plant cells include, for example, a Cauliflower Mosaic Virus 35S promoter (Odell et al. (1985) Nature 313:810-812), a ubiquitin promoter (U.S. Pat. No.5,510,474; Christensen et al. (1989)), or a rice actin promoter (McElroy et al. (1990) Plant Cell 2:163-171).
- Exemplary promoters that can be used in mammalian cells include, a viral promoter such as an SV40 promoter or a metallothionine promoter. All of these host cell promoters are well known by and readily available to one of ordinary skill in the art. Further nucleic acid control elements useful for controlling expression in a recombinant host cell can include transcriptional terminators, enhancers, and the like, all of which may be used with the heterologous nucleic acids incorporate in the recombinant host cells of the present disclosure. [0114] A wide variety of techniques are well known in the art for linking transcriptional promoters and other control elements to heterologous nucleic acid sequences encoding pathway genes for biosynthesis of UDCA or 3-KUDCA.
- the heterologous nucleic acid sequences of the present disclosure comprise a promoter capable of controlling expression in a host cell, wherein the promoter is linked to a nucleic acid sequence encoding a recombinant polypeptide having CYP activity of the present disclosure, and as necessary, other enzymes constituting a pathway for production of a UDCA precursor, UDCA, and/or UDCA derivative.
- This heterologous nucleic acid sequence can be integrated into a recombinant expression vector which ensures good expression in the desired host cell, wherein the expression vector is suitable for expression in a host cell, meaning that the recombinant expression vector comprises the heterologous nucleic acid sequence linked to any genetic elements required to achieve expression in the host cell.
- Genetic elements that may be included in the expression vector in this regard include a transcriptional termination region, one or more nucleic acid sequences encoding marker genes, one or more origins of replication, and the like.
- the expression vector further comprises genetic elements required for the integration of the vector or a portion thereof in the host cell's genome.
- an expression vector comprising a heterologous nucleic acid of the present disclosure may further contain a marker gene.
- Marker genes useful in accordance with the present disclosure include any genes that allow the distinction of transformed cells from non-transformed cells, including all selectable and screenable marker genes.
- a marker gene may be a resistance marker such as an antibiotic resistance marker against, for example, kanamycin or ampicillin.
- Screenable markers that may be employed to identify transformants through visual inspection include ⁇ -glucuronidase (GUS) (U.S. Pat. Nos.5,268,463 and 5,599,670) and green fluorescent protein (GFP) (Niedz et al., 1995, Plant Cell Rep., 14: 403).
- GUS ⁇ -glucuronidase
- GFP green fluorescent protein
- the present disclosure also provides of a method for producing UDCA or 3-KUDCA, wherein a heterologous nucleic acid encoding a recombinant polypeptide having CYP activity (e.g., an exemplary engineered polypeptide of Table 3) can be introduced into a recombinant host cell.
- the recombinant host cell can then be used for production of the polypeptide or incorporated in a biocatalytic process that utilized the CYP activity of the recombinant polypeptide expressed by the host cell for the catalytic conversion of a substrate, e.g., the conversion of LCA to UDCA.
- the recombinant host cell can further comprise a pathway of enzymes capable of producing a compound precursor (e.g., LCA) which can act as a substrate for the recombinant polypeptides with CPR and CYP activity.
- a recombinant host cell comprising a heterologous nucleic acid encoding a recombinant polypeptide having CYP activity of the present disclosure can provide improved biosynthesis of a desired product compound (e.g., UDCA or UDCA derivative) in terms of titer, yield, and production rate, due to the improved characteristics of the expressed ⁇ 46 ⁇ CYP activity in the cell associated with the amino acid and codon differences engineered in the gene.
- the present disclosure provides a method for producing UDCA and/or 3-KUDCA comprising: (a) culturing a recombinant host cell of the present disclosure in a suitable medium comprising LCA and/or 3-KCA; and (b) recovering the produced UDCA and/or 3-KUDCA.
- the recombinant polypeptides with CYP activity of the present disclosure can be incorporated in any biosynthesis method requiring a CYP catalyzed biocatalytic step, whether in vivo or in vitro.
- the recombinant polypeptides having CYP activity can be used in a method for preparing a compound of structural formula (I) [0119]
- the present disclosure provides a method for preparing a compound of structural formula (I) wherein, R 1 is substituted C1-C20 alkyl, or an optionally or a is a hydroxyl, or an oxo group.
- the method comprises contacting a recombinant polypeptide having CYP activity of the present disclosure (e.g., an exemplary polypeptide of Table 3) under suitable reactions conditions with a compound of structural formula (II) 1 wherein, R is a an substituted C1-C20 alkyl, or an optionally substituted aryl; R 2 is hydrogen, or a hydroxyl; and R 3 is a hydroxyl, or an oxo group.
- a recombinant polypeptide having CYP activity of the present disclosure e.g., an exemplary polypeptide of Table 3
- R is a an substituted C1-C20 alkyl, or an optionally substituted aryl
- R 2 is hydrogen, or a hydroxyl
- R 3 is a hydroxyl, or an oxo group.
- Exemplary conversions of UDCA precursor compounds of structural formula (II) to UDCA and related compounds of structural formula (I) that are catalyzed by the recombinant polypeptides having CYP activity of the present disclosure can include: (1) conversion of LCA to UDCA; (2) the conversion of 3-KCA to 3-KUDCA; (3) the conversion of DCA to UCA. Accordingly, in at least one embodiment of the biosynthesis method for conversion a UDCA precursor compound of structural formula (II) to a UDCA compound of structural formula (I), R 2 is hydrogen, the compound of structural formula (II) is LCA, and the compound of structure formula (I) is UDCA.
- R 2 is hydrogen, the compound of structural formula (II) is 3-KCA and the compound of structure formula (I) is 3-KUDCA. In at least one embodiment, R 2 is a hydroxyl, the compound of structure formula (II) is DCA, and the compound of structural formula (II) is DCA.
- the present disclosure also contemplates that the methods for biocatalytic conversion of a UDCA precursor compound of structural formula (II) to a UDCA compound of structural formula (I) using an recombinant polypeptide having CYP activity of the present disclosure can comprise additional chemical or biocatalytic steps carried out on the product compound of structural formula (II), including steps of product compound work-up, extraction, isolation, purification, and/or crystallization, each of which can be carried out under a range of conditions.
- Suitable reaction conditions for the biosynthesis of compounds such as UDCA, 3- KUDCA, and UCA are known in the art and can be used with the recombinant polypeptides having CYP activity of the present disclosure.
- suitable reaction conditions for the exemplary polypeptides of the present disclosure can be determined using routine techniques known in the art for optimizing biocatalytic reactions. It is contemplated that various ranges of suitable reaction conditions with the recombinant polypeptides of the present disclosure, including but not limited to ranges of pH, temperature, buffer, solvent system, substrate loading, polypeptide loading, co-substrate or co-factor loading, atmosphere, and reaction time. Suitable reaction conditions can be readily determined and optimized for particular reactions by routine experimentation that includes, but is not limited to, contacting the recombinant polypeptide and substrate under experimental reaction conditions of concentration, pH, temperature, solvent conditions, and detecting the production of the desired compound of structural formula (I).
- the suitable reaction conditions comprise a reaction solution of ⁇ pH 7-8, a temperature of 25C to 37C; optionally, the reaction conditions comprise a reaction solution of ⁇ pH 7 and a temperature of ⁇ 30C. In at least one embodiment, the reaction solution is allowed to incubate at a temperature of 25C to 37C for a reaction time of at least 1, 6, 12, 24, or 48 hours, before the amount of reaction product is determined.
- Example 1 Genomic Integration of Heterologous CYP/CPR Gene Pairs into Saccharomyces cerevisiae Host Cells
- This example illustrates the preparation of recombinant yeast host cell strains (Saccharomyces cerevisiae), with genomically integrated copies of a pair of heterologous genes encoding a cytochrome P450 (CYP) and a cytochrome reductase (CPR) from Gibberella zeae.
- CYP cytochrome P450
- CPR cytochrome reductase
- SAND122 expresses a pair of genes from Gibberella zeae that encode a polypeptide with cytochrome P450 (CYP) activity (SEQ ID NO: 2) and a polypeptide with cytochrome P450 reductase (CPR) activity (SEQ ID NO: 4) and has been shown to carry out the conversion of LCA to UDCA and 3-KCA to 3- KUDCA.
- CYP cytochrome P450
- CPR cytochrome P450 reductase
- S. cerevisiae of SEQ ID NO: 2 and the CPR polypeptide from G. zeae of SEQ ID NO: 4 were codon-optimized for optimal expression in S. cerevisiae as SEQ ID NO: 1, and SEQ ID NO: 3, respectively.
- These codon-optimized genes were integrated into the S. cerevisiae genome under the bidirectional pGal1/10 promoter system (SEQ ID NO: 59) and ADH1t (SEQ ID NO: 60) and PGKt terminators (SEQ ID NO: 61) to generate single copy strains (SH010 and SH013) as follows.
- Nucleic acids with the codon-optimized CYP gene of SEQ ID NO: 1, and the codon optimized CPR gene of SEQ ID NO: 3 were synthesized as fragments (Twist Bioscience, Inc., South San Francisco, California), which were further PCR amplified to incorporate ⁇ 20-30 bp overhang sequences homologous to the bidirectional pGal1/10 promoter and the respective terminator sequences on the 5’ or 3’ ends of the respective gene sequences. Locus specific PCR amplicons were also amplified with ⁇ 20-30 bp homologous overhangs to the terminator regions to enable efficient integration into the S. cerevisiae X-4 locus.
- Fragment C containing the CYP gene was amplified using the forward primer P5 and the reverse primer P6.
- Fragment D containing the PGK1t terminator was amplified using the forward primer P7 and the reverse primer P8.
- Fragment E containing the X-4 downstream homology region was amplified using the forward primer P9 and the reverse primer P10.
- the five fragments A-E were assembled by NEBuilder® HiFi Assembly and amplified as a single linear DONOR fragment using the following rescue forward primer P11 and the rescue reverse primer P12.
- the final assembled 4680 bp linear DONOR DNA was gel extracted and integrated as a knock-in using CRISPR- Cas9 into the X-4 locus of the yeast strain CENPK2-1D.
- This SH010 strain was used as the parent strain to introduce a gal80 knockout to enable induction of the bidirectional promoter pGal1/10 in the absence of glucose.
- the Gal80 knockout DONOR DNA was amplified using the forward rescue primer P13 and the reverse rescue primer P14 and integrated as a knock-in using CRISPR-Cas9 to disrupt the native gal80 gene in SH010 resulting in the ⁇ Gal80, single copy CYP-CPR strain SH013.
- PCR primers summarized in Table 4 also were used to amplify the above-mentioned parts as follows and illustrated by Scheme B of FIG.1.
- the 5’ upstream homology arm (Fragment A) was amplified using the forward primer P17 and the reverse primer P18 and the 3’ downstream homology arm (Fragment C) was amplified using the forward primer P19 and the reverse primer P20.
- 5’ upstream homology arm (Fragment A) was amplified using the forward primer P21 and the reverse primer P22 and the 3’ downstream homology arm (Fragment C) was amplified using the forward primer P23 and the reverse primer P24.
- the XI-2 DONOR DNA was transformed and integrated into XI-2 locus as a knock-in using CRISPR-Cas9 into the SH013 ⁇ 51 ⁇ (single copy) strain to generate the double copy strain SHP021.
- the XII-4 DONOR was transformed and integrated into XII-4 locus as a knock-in using CRISPR-Cas9 into the SHP021 (double copy) strain to generate the triple copy strain SH020.
- the X-2 DONOR was transformed and integrated into X-2 locus as a knock-in using CRISPR-Cas9 into the SH020 (triple copy) strain to generate the quadruple copy strain SHP025.
- cerevisiae strains were picked into 500 mL baffled conical flasks containing 2 x YPD media (100 mL working volume). The flasks were incubated at 30 o C for 48 h with shaking at 250 rpm at 85 % humidity. After this time, a 1% final concentration solution of galactose (2.5 mL of a 40 % stock solution) was added and the flask was further incubated for 4 hours to induce protein production. After this induction period, the biomass was harvested via centrifugation and the supernatant removed.
- LC/MS sample preparation The bioconversion reaction was extracted and diluted with MeOH for sample preparation as described above. The prepared samples were loaded onto an Agilent 6470 Triple Quadrupole LC/MS and the compounds of interest were detected using MS/MS in SIM mode. The compounds were quantified relative to calibration curves which were prepared using serial dilutions of stock solutions of 3-KCA and 3-KUDCA.
- LC/MS instrumentation and parameters LC/MS system: Agilent 6470 Triple Quadrupole LC/MS; Column: Agilent EclipsePlusC18 column 1.8 um, 3.0 x 50 mm; Mobile phase: Solvent A: H 2 O with 0.1% formic acid; Solvent B: Acetonitrile with 0.1% formic acid; Gradient: 0 - 0.2 min isocratic 40% B, 0.2 - 2.2 min gradient to 85% B, 2.2 - 2.8 min isocratic 85% B; Mode: MS/MS in SIM mode (monitoring at 389.3 for 3-KUDCA and 373.2 for 3-KCA); Fragmentor: 90 volts.
- the URA3 marker was integrated into the X-4 site (Easy-Clone 2.0) of a yeast strain under the pGal1 promoter within the pGal1/10 promoter system (SEQ ID NO: 59) and PGKt1 terminator (SEQ ID NO: 61), along with the codon- optimized polynucleotide sequence of SEQ ID NO: 3, which encodes the CPR polypeptide of SEQ ID NO: 4, under the pGal10 promoter within the pGal1/10 promoter system (SEQ ID NO: 59) and ADH1t terminator (SEQ ID NO: 60).
- the resulting strain SHEVP003 was used as the negative control and screening host for CYP SSM library integration. Libraries were plated on selective media containing 5’-fluoroorotic acid, which selects for URA3 negative cells, indicating successful integration of DONOR DNA into the SHEVP003 screening strain. SH013 strain was ⁇ 53 ⁇ used as a control strain during screening of the SSM library strains for fold-improvement calculations with respect to conversion of 3-KCA to 3-KUDCA as described below.
- Genomic DNA from the SH013 strain was used as the template to generate two PCR products: (1) a first PCR product (Fragment A), which does not harbor any degenerate codons, and (2) a second PCR product (Fragment B), which has sequence overlap with the Fragment A, and is amplified harboring one NNK degenerate codon only.
- Primers spanning the full 513 codons of the CYP polypeptide were designed according to standard site-saturation mutagenesis protocols and used for amplification of Fragments A and B and overlap extension.
- Fragment A was amplified using a single forward primer P34 (SEQ ID NO: 95) and a series of 513 reverse primers designed according to the location of the desired mutagenesis site in CYP.
- the 513 reverse primers for Fragment A listed in Table 6A (below) and are provided in the accompanying Sequence Listing as SEQ ID NOs: 102-614.
- cerevisiae strains were picked into 96-well plates containing 2 x YPD media (300 ⁇ L per well) using a QPix TM 420 colony picking system. The plates were incubated at 30 o C for 24 h with shaking at 250 rpm at 85 % humidity. After this time, the strains were sub-cultured (40 ⁇ L inoculation volume per strain) into a second 96-well plate containing 2 x YPD media (1.2 mL per well) using an Agilent Bravo automated liquid handling platform. The plates were incubated at 30 o C for 48 h with shaking at 250 rpm at 85 % humidity.
- Example 2 The same screening strain was used to integrate the resulting libraries as in Example 2 (SHEVP003), however the SHS036 strain was used as the parent control strain in order to determine fold-improvement in conversion of 3-KCA to 3-KUDCA as described below.
- a semi-synthetic approach was used to construct the first set of combinatorial libraries. Genomic DNA from the SHS036 strain was used as the template to generate a full-length PCR product using primer pair of P34 (SEQ ID NO: 95) and P37 (SEQ ID NO: 98) while incorporating uracil using a dNTP mix comprising of the following deoxyribonucleotides: dATP, dGTP, dCTP, dTTP, dUTP.
- the resulting PCR product was gel purified and digested with Uracil-DNA Glycosylase and Endonuclease IV at 37 C for 2 hours, followed by enzyme denaturation at 94 C for two minutes, to generate a pool of fragments in the range of 50-100 bases. These fragments were further combined with differing ratios of pools of the synthesized oligonucleotide primers (each oligo up to 55 bases in length and encoding one or more amino acid change) in several individual assembly PCR reactions using forward primer P5 (SEQ ID NO: 66) and reverse primer P6 (SEQ ID NO: 67) to reassemble the full-length PCR product as Fragment B and incorporate mutagenic amino acid changes within each pool randomly.
- forward primer P5 SEQ ID NO: 66
- reverse primer P6 SEQ ID NO: 67
- Fragment A was amplified using the forward primer P37 (SEQ ID NO: 98) and reverse primer P4 (SEQ ID NO: 65) and Fragment C was amplified using the forward primer P7 (SEQ ID NO: 68) and the reverse primer P38 (SEQ ID NO: 99).
- Fragments A, B, and C were assembled by NEBuilder® HiFi Assembly and amplified as a single linear DONOR using the following forward rescue primer P35 (SEQ ID NO: 96) and reverse primer P36 (SEQ ID NO: 97). The assembled PCR products were then pooled together, and gel purified to provide a combinatorial library of linear donor DNA.
- SAND121 A recombinant strain of P. pastoris, denoted as “SAND121” has been described previously in WO2022115710A1. SAND121 expresses a pair of genes encoding a cytochrome P450 (CYP) and a reductase (CPR) from Gibberella zeae and has been shown to carry out the conversion of LCA to UDCA and 3-KCA to 3-KUDCA.
- CYP cytochrome P450
- CPR reductase
- yeast codon-optimized polynucleotide sequences of SEQ ID NO: 1 and 3 encoding the CYP polypeptide of SEQ ID NO: 2 and the CPR polypeptide of SEQ ID NO: 4, respectively, were integrated into the P. pastoris genome under the bidirectional pCAT1:pFDH1 promoter system (SEQ ID NO: 1002) and tDAS1 and tDAS2 terminators SEQ ID NO: 1003 and SEQ ID NO: 1004, respectively, to generate a single copy strain (SHP026) as follows.
- a single plasmid (wbplasmid147; SEQ ID NO: 1059) containing a zeocin resistance marker, a pUC origin of replication, a HIS4 homologous sequence (SEQ ID NO: 1005), and a cassette comprising of tDAS1:CPR001:pCAT1:pFDH1:CYP001:tDAS2 was assembled.
- Sequences were amplified using PCR using the following primers: wboligos5777, wboligos5792, wboligos5793, wboligos5796, wboligos5797, wboligos5798, wboligos5800, wboligos5801, wboligos5802, wboligos5803, wboligos5793, wboligos5794, wboligos5795, and wboligos5776.
- the sequences of the PCR primers described in the following P. pastoris strain building examples by their Primer Name are listed in Table 10 below and the accompanying Sequence Listing.
- the restriction enzyme cut site found in the HIS4 homologous sequence linearized the plasmid, allowing for homologous integration at the HIS4 locus in Pichia competent cells, which were prepared from the wild-type BG-10 strain using Thermo- ⁇ 80 ⁇ Fisher Pichia EasyComp Transformation Kit (catalog no. K173001) according to manufacturer’s instructions.5 microliters of linearized plasmid were used to transform 50 ⁇ L of BG-10 competent cells and plated on YPDS media plates containing zeocin. After 3 days of growth, individual colonies were confirmed for integration of the linearized plasmid using PCR and genomic DNA isolated from individual colonies.
- a recombinant Pichia host cell strain (SHP029) with two copies of a gene encoding CYP (SEQ ID NO: 2) and one copy of a gene encoding CPR (SEQ ID NO: 4) was also constructed as follows.
- a single plasmid (wbplasmid171; SEQ ID NO: 1061) containing an ampicillin resistance marker, a G418 resistance marker, a pUC origin of replication, promoter pAOX1 (SEQ ID NO: 1051), CYP (SEQ ID NO: 2), and terminator tDAS2 (SEQ ID NO: 1004) was assembled using the method described above.
- Sequences were amplified using PCR with the following primers: wboligos5979, wboligos6014, wboligos6013, wboligos6015, wboligos6016, wboligos5921, wboligos5922, wboligos5802, wboligos5803, and wboligos5978.
- a single plasmid was digested with SacI and further purified using Zymo Research’s DNA Clean and Concentrator Kit (catalog no. D4014).
- the restriction enzyme cut site found in pAOX1, linearized the plasmid, allowing for homologous integration at the AOX1 locus.
- Five microliters of linearized plasmid were used to transform 50 ⁇ L of SHP026 competent cells and plated on YPD media plates containing G418. After 3 days of growth, individual colonies were confirmed for integration of the linearized plasmid using PCR and genomic DNA from each individual colony.
- a recombinant Pichia host cell strain (SHP030) with one copy of a gene encoding the S31V mutant of CYP (SEQ ID NO: 6) and one copy of the gene encoding CPR (SEQ ID NO: 4) was also constructed as follows.
- a single plasmid (wbplasmid160; SEQ ID NO: 1060) containing a zeocin resistance marker, a pUC origin of replication, a HIS4 homologous sequence (SEQ ID NO: 1005), and a cassette comprising tDAS1:CPR:pCAT1:pFDH1:CYP100:tDAS2 was assembled as described above.
- wbplasmid147 shares the same sequence as wbplasmid160 except for the CYP CDS, inverse PCR was used to amplify the vector excluding the CYP CDS using primers wboligos5800 and wboligos5803.
- CYP (SEQ ID NO: 2) was amplified from SHS036 gDNA using primers wboligos5801 and wboligos5802. Following plasmid recovery and sequence confirmation, a single plasmid was digested with BamHI and further purified using Zymo Research’s DNA Clean and Concentrator Kit (catalog no. D4014). The restriction enzyme cut site, found in the HIS4 homologous sequence, linearized the plasmid, allowing for homologous integration at the HIS4 locus. Five microliters of linearized plasmid were used to transform 50 ⁇ L of BG-10 competent cells and plated on YPDS media plates containing zeocin.
- a recombinant Pichia host cell strain (SHP033) with two copies of the gene encoding CYP (SEQ ID NO: 2) and one copy of a gene encoding CPR (SEQ ID NO: 4) was also constructed as follows.
- a single plasmid (wbplasmid172; SEQ ID NO: 1062) containing an ampicillin resistance marker, a G418 resistance marker, a pUC origin of replication, promoter pAOX1 (SEQ ID NO:1051), gene encoding CYP100 (SEQ ID NO: 6), and terminator tDAS2 (SEQ ID NO: 1004) was assembled with the method described above.
- wbplasmid171 shares the same sequence as wbplasmid172 except for the CYP CDS, inverse PCR was used to amplify the vector excluding the CYP CDS using primers wboligos5921 and wboligos5803.
- CYP was amplified from SHS036 gDNA using primers wboligos5922 and wboligos5802. Following plasmid recovery and sequence confirmation as described above, a single plasmid was digested with SacI and further purified using Zymo Research’s DNA Clean and Concentrator Kit (catalog no. D4014). The restriction enzyme cut site, found in pAOX1, linearized the plasmid, allowing for homologous integration at the AOX1 locus. Five microliters of linearized plasmid were used to transform 50 ⁇ L of SHP030 competent cells and plated on YPD media plates containing G418.
- a recombinant Pichia host cell strain (SHP034) with three copies of a gene encoding the S31V CYP (SEQ ID NO: 6) and two copies of the gene encoding CPR (SEQ ID NO: 4) was also constructed as follows.
- a single plasmid (wbplasmid173; SEQ ID NO:1063) containing a hygromycin resistance marker, an ampicillin resistance marker, a pUC origin of replication, an Int6 homologous sequence (SEQ ID NO: 1052), and a cassette comprising terminator tPMP20 (SEQ ID NO: 1053), CPR (SEQ ID NO: 4), bidirectional promoter pCAT1:FDH1 (SEQ ID NO: 1054), CYP100 (SEQ ID NO: 6), and terminator tFLD1 (SEQ ID NO: 1055), was assembled as described above.
- Sequences were amplified using PCR using the following primers: wboligos5983, wboligos5984, wboligos5985, wboligos5986, wboligos5987, wboligos5988, wboligos5989, wboligos5990, wboligos5991, and wboligos5992.
- a single plasmid was digested with BsgI and further purified using Zymo Research’s DNA Clean and Concentrator Kit (catalog no. D4014).
- the restriction enzyme cut site found in the Int6 homologous sequence, linearized the plasmid, allowing for homologous integration at its corresponding locus (an intergenic region).
- Five microliters of linearized plasmid were used to transform 50 ⁇ L of SHP033 competent cells and plated on YPD media plates containing hygromycin. After 3 days of growth, individual colonies were confirmed for integration of the linearized plasmid using PCR and genomic DNA from each individual colony.
- a recombinant Pichia host cell strain (SHP035) with four copies of the gene encoding S31V CYP (SEQ ID NO: 6) and three copies of the gene encoding CPR (SEQ ID NO: 4) was also constructed as follows.
- a single plasmid (wbplasmid174; SEQ ID NO:1064) containing a ⁇ 82 ⁇ nourseothricin resistance marker, an ampicillin resistance marker, a pUC origin of replication, an Int15 homologous sequence (SEQ ID NO: 1056), and a cassette comprising of terminator tFBA2 (SEQ ID NO: 1057), gene encoding CPR (SEQ ID NO:4), bidirectional promoter pCAT1:FDH1 (SEQ ID: 1053), gene encoding CYP100 (SEQ ID NO: 6), and terminator tADH2 (SEQ ID NO: 1058), was assembled as described above.
- Sequences were amplified using PCR using the following primers: wboligos5993, wboligos5994, wboligos5995, wboligos5996, wboligos5997, wboligos5998, wboligos5999, wboligos6000, wboligos6001, and wboligos6002.
- a single plasmid was digested with SalI and further purified using Zymo Research’s DNA Clean and Concentrator Kit (catalog no. D4014).
- the restriction enzyme cut site found in the Int15 homologous sequence, linearized the plasmid, allowing for homologous integration at its corresponding locus (an intergenic region).
- Five microliters of linearized wbplasmid173 and 5 microliters of linearized wbplasmid174 were used to co-transform 50 ⁇ L of SHP033 competent cells and plated on YPD media plates containing hygromycin and nourseothricin. After 3 days of growth, individual colonies were confirmed for integration of the linearized plasmids using PCR and genomic DNA from each individual colony.
- a recombinant Pichia host cell strain (SHP038) with one copy of a polynucleotide of SEQ ID NO: 31 encoding the CYP105 mutant (S31V, Q54K, A143G, T307A, G325E) of SEQ ID NO: 32 identified in Example 3 and one copy of the gene of SEQ ID NO: 3 encoding CPR (SEQ ID NO: 4) was also constructed as follows.
- a single plasmid (wbplasmid193; SEQ ID NO: 1065) containing a zeocin resistance marker, a pUC origin of replication, a HIS4 homologous sequence (SEQ ID NO: 1005), and a cassette comprising of tDAS1:CPR001:pCAT1:pFDH1:CYP105:tDAS2 was assembled as described above.
- wbplasmid147 shares the same sequence as wbplasmid193 except for the CYP CDS, inverse PCR was used to amplify the vector minus the CYP CDS using primers wboligos5800 and wboligos5803.
- CYP105 (SEQ ID NO: 32) was amplified from SHS042 gDNA using primers wboligos5801 and wboligos5802. Following plasmid recovery and sequence confirmation as described above, a single plasmid was digested with BbvcI and further purified using Zymo Research’s DNA Clean and Concentrator Kit (catalog no. D4014). The restriction enzyme cut site, found in the HIS4 homologous sequence, linearized the plasmid, allowing for homologous integration at the HIS4 locus. Five microliters of linearized plasmid were used to transform 50 ⁇ L of BG-10 competent cells and plated on YPDS media plates containing zeocin.
- a recombinant Pichia host cell strain (SHP052) with one copy of a gene encoding CYP105 mutant (S31V, Q54K, A143G, T307A, G325E) of SEQ ID NO: 32 identified in Example 3 was also constructed as follows.
- a single plasmid (wbplasmid202; SEQ ID NO:1066) containing a geneticin resistance marker, an ampicillin resistance marker, a pUC origin of ⁇ 83 ⁇ replication, and a cassette comprising of pAOX1:CYP105:tDAS2 was assembled as described above.
- wbplasmid202 shares the same sequence as wbplasmid160 except for the CYP100 CDS, inverse PCR was used to amplify the vector minus the CYP100 CDS using primers wboligos5803 and wboligos5921.
- CYP105 SEQ ID NO: 32
- a single plasmid was digested with SacI and further purified using Zymo Research’s DNA Clean and Concentrator Kit (catalog no. D4014).
- the restriction enzyme cut site found in the Int6 homologous sequence, linearized the plasmid, allowing for homologous integration at the Int6 locus (intergenic region).
- Five microliters of linearized plasmid were used to transform 50 ⁇ L of BG-10 competent cells and plated on YPDS media plates containing geneticin. After 3 days of growth, individual colonies were confirmed for integration of the linearized plasmid using PCR and genomic DNA from each individual colony.
- a recombinant Pichia host cell strain (SHP053) with two copies of the gene encoding CYP105 (SEQ ID NO: 32) identified in Example 3 and one copy of the gene encoding CPR (SEQ ID NO: 4) was also constructed as follows.
- Plasmid wbplasmid202 was assembled, digested, and purified as described above. Five microliters of linearized plasmid were used to transform 50 ⁇ L of SHP038 competent cells and plated on YPDS media plates containing geneticin. After 3 days of growth, individual colonies were confirmed for integration of the linearized plasmid using PCR and genomic DNA from each individual colony. [0185] The strain build lineage for all nine of the above-described P. pastoris strains is summarized in the chart of FIG.3. [0186] B.
- Screening of strains for 3-KCA to 3-KUDCA conversion was carried out according to the following assay: individual colonies of recombinant P. pastoris strains were picked into 500 mL baffled conical flasks containing BMGY media (100 mL working volume). The flasks were incubated at 30 o C for 72 h with shaking at 250 rpm at 85 % humidity and were supplemented with glycerol (2 % final concentration) every 24 hours. After this time, the biomass was harvested via centrifugation and the supernatant was discarded.
- the biomass was then resuspended in BMMY and incubated at 30 o C with shaking at 250 rpm at 85 % humidity for 6 hours to induce protein production. After this time, the biomass was harvested via centrifugation and the supernatant removed. A portion of the resulting biomass (2 g) was resuspended in 10 mL of bioconversion buffer (295-10000 mg/L 3-KCA, 0.1 M phosphate buffer, 2 % MeOH, 2 mM 5-ALA, pH 9) and the flask was incubated at 30 o C for 24 h with shaking at 250 rpm at 85 % humidity.
- bioconversion buffer (295-10000 mg/L 3-KCA, 0.1 M phosphate buffer, 2 % MeOH, 2 mM 5-ALA, pH 9
- SSM Site Saturation Mutagenesis
- the heterologous nucleic acid sequence from the single copy strain SHS058 build which includes the codon-optimized polynucleotide sequence of SEQ ID NO:31 encoding the CYP polypeptide of SEQ ID NO: 32, expressed under the pGal1 promoter within the pGal1/10 promoter system (SEQ ID NO: 59) and PGK1t terminator (SEQ ID NO: 61), was used to design SSM oligonucleotide to generate SSM libraries at positions spanning the entire polypeptide sequence.
- the URA3 marker was integrated into the X-4 site (Easy-Clone 2.0) of a yeast strain under the pGal1 ⁇ 85 ⁇ promoter within the pGal1/10 promoter system (SEQ ID NO: 59) and PGKt1 terminator (SEQ ID NO: 61), along with the codon-optimized polynucleotide sequence of SEQ ID NO: 3, which encodes the CPR polypeptide of SEQ ID NO: 4, under the pGal10 promoter within the pGal1/10 promoter system (SEQ ID NO: 59) and ADH1t terminator (SEQ ID NO: 60).
- the resulting strain SHEVP003 was used as the negative control and screening host for CYP SSM library integration. Libraries were plated on selective media containing 5’-fluoroorotic acid, which selects for URA3 negative cells, indicating successful integration of DONOR DNA into the SHEVP003 screening strain. SHS058 strain was used as a control strain during screening of the SSM library strains for fold-improvement calculations with respect to conversion of 3-KCA to 3-KUDCA as described below.
- Genomic DNA from the SHS058 strain was used as the template to generate two PCR products: (1) a first PCR product (Fragment A), which does not harbor any degenerate codons, and (2) a second PCR product (Fragment B), which has sequence overlap with the Fragment A, and is amplified harboring one NNK degenerate codon only.
- Primers spanning the full 513 codons of the CYP polypeptide were designed according to standard site-saturation mutagenesis protocols and used for amplification of Fragments A and B and overlap extension.
- Fragment A was amplified using a single forward primer P34 (SEQ ID NO: 95) and a series of 513 reverse primers designed according to the location of the desired mutagenesis site in CYP.
- the 513 reverse primers for Fragment A listed in Table 6A (above) and are provided in the accompanying Sequence Listing as SEQ ID NOs: 102-614. Additional reverse primers were designed as to not perturb the codon mutations in the backbone of CYP polypeptide SEQ ID NO: 32.
- the reverse primers for Fragment A are provided in Table 12A below and are provided in the accompanying Sequence Listing as SEQ ID NOs: 1067-1114.
- the 513 forward primers for Fragment B are listed in Table 6B and are provided in the accompanying Sequence Listing as SEQ ID NOs: 615-984. Additional forward primers were designed as to not perturb the codon mutations in the backbone of CYP polypeptide SEQ ID NO: 32.
- the forward primers for Fragment B are provided in Table 12B below and are provided in the accompanying Sequence Listing as SEQ ID NOs: 1115-1184.
- the plates were incubated at 30 o C for 24 h with shaking at 250 rpm at 85 % humidity. After this time, the strains were sub-cultured (40 ⁇ L inoculation volume per strain) into a second 96-well plate containing 1 x YPD media (300 ⁇ L per well) using an Agilent Bravo automated liquid handling platform. The plates were incubated at 30 o C for 48 h with shaking at 250 rpm at 85 % humidity. After this time, a solution of galactose (40 % stock solution) was added using the Bravo (1 % final concentration) and the plates were further incubated for 4 hours to induce protein production. After this time, the biomass was harvested via centrifugation and the supernatant removed.
- the resulting biomass was resuspended in 300 ⁇ L of bioconversion buffer (590 mg/L 3-KCA, 0.1 M phosphate buffer, pH 9) using the Agilent Bravo automated liquid handling platform and the plates were incubated at 30 o C for 24 h with shaking at 250 rpm at 85 % humidity. After this time, MeOH (300 ⁇ L) was added using the Bravo and the plates were incubated at 30 o C with shaking at 250 rpm at 85 % humidity for a further 30 minutes. The biomass was then separated by centrifugation and the resulting supernatant diluted with MeOH (1000 x total dilution) using the Agilent Bravo automated liquid handling platform before analysis using an Agilent 6470 Triple Quadrupole LC/MS.
- bioconversion buffer 590 mg/L 3-KCA, 0.1 M phosphate buffer, pH 9
- LC/MS sample preparation The bioconversion reaction was extracted and diluted with MeOH for sample preparation as described above. The prepared samples were loaded onto an Agilent 6470 Triple Quadrupole LC/MS and the compounds of interest were detected using MS/MS in SIM mode. The compounds were quantified relative to calibration curves which were prepared using serial dilutions of stock solutions of 3-KCA and 3-KUDCA.
- LC/MS instrumentation and parameters LC/MS system: Agilent 6470 Triple Quadrupole LC/MS; Column: Agilent EclipsePlusC18 column 1.8 um, 3.0 x 50 mm; Mobile phase: Solvent A: H 2 O with 0.1% formic acid; Solvent B: Acetonitrile with 0.1% formic acid; ⁇ 89 ⁇ Gradient: 0 - 0.2 min isocratic 40% B, 0.2 - 2.2 min gradient to 85% B, 2.2 - 2.8 min isocratic 85% B; Mode: MS/MS in SIM mode (monitoring at 389.3 for 3-KUDCA and 373.2 for 3-KCA); Fragmentor: 90 volts.
- This example illustrates the preparation of further combinatorial libraries of CYP genes based on the parent strain SHS088 which carries a gene of SEQ ID NO: 1187 encoding the CYP207 (SEQ ID NO: 1188) which has the 6 amino acid differences S31V, Q54K, A143G, T307A, G325E, V364L relative to the wild-type CYP of SEQ ID NO: 2, identified through SSM optimization of CYP105 for improved bioconversion of 3-KCA to 3-KUDCA described in Example 5.
- SHS088 which carries a gene of SEQ ID NO: 1187 encoding the CYP207 (SEQ ID NO: 1188) which has the 6 amino acid differences S31V, Q54K, A143G, T307A, G325E, V364L relative to the wild-type CYP of SEQ ID NO: 2, identified through SSM optimization of CYP105 for improved bioconversion of 3-KCA to 3-KUDCA described in Example 5.
- Genomic DNA from the SHS088 strain was used as the template to generate a full-length PCR product using primer pair of P34 (SEQ ID NO: 95) and P37 (SEQ ID NO: 98), while incorporating uracil using a dNTP mix comprising of the following deoxyribonucleotides: dATP, dGTP, dCTP, dTTP, dUTP.
- the resulting PCR product was gel purified and digested with Uracil-DNA Glycosylase and Endonuclease IV at 37 C for 2 hours, followed by enzyme denaturation at 94 C for two minutes, to generate a pool of fragments in the range of 50-100 bases. These fragments were further combined with differing ratios of pools of the synthesized oligonucleotide primers (each oligo up to 55 bases in length and encoding one or more amino acid change) in several individual assembly PCR reactions using forward primer P5 (SEQ ID NO: 66) and reverse primer P6 (SEQ ID NO: 67) to reassemble the full-length PCR product as Fragment B and incorporate mutagenic amino acid changes within each pool randomly.
- forward primer P5 SEQ ID NO: 66
- reverse primer P6 SEQ ID NO: 67
- Fragment A was amplified using the forward primer P37 (SEQ ID NO: 98) and reverse primer P4 (SEQ ID NO: 65) and Fragment C was amplified using the forward primer P7 (SEQ ID NO: 68) and the reverse primer P38 (SEQ ID NO: 99).
- Fragments A, B, and C were assembled by NEBuilder® HiFi Assembly and amplified as a single linear DONOR using the following forward rescue primer P35 (SEQ ID NO: 96) and reverse primer P36 (SEQ ID NO: 97). The assembled PCR products were then pooled together, and gel purified to provide a combinatorial library of linear donor DNA.
- A. P. pastoris strain building [0224] A recombinant Pichia host cell strain (SHP070) with one copy of the gene encoding CYP110 (SEQ ID NO: 1185) was also constructed as follows. A single plasmid, wbplasmid226 (SEQ ID NO: 1273) containing a geneticin resistance marker, an ampicillin resistance marker, a pUC origin of replication, and a cassette comprising of pAOX1:CYP110:tDAS2 was assembled as described above.
- the wbplasmid226 shares the same sequence as wbplasmid202 (SEQ ID NO: 1066) except for the CYP105 coding sequence (SEQ ID NO: 31) was replaced with the CYP110 coding sequence of SEQ ID NO: 1185.
- Inverse PCR was used to amplify the vector minus the CYP100 coding sequence using primers wboligos5803 and wboligos5921.
- the gene encoding CYP110 (SEQ ID NO: 1185) was amplified from SHS062 gDNA using primers wboligos5922 and wboligos5802.
- CYP110 in SHS062 was expressed under the pGal1 promoter within the pGal1/10 promoter system (SEQ ID NO: 59) and PGK1t terminator (SEQ ID NO: 61).
- SEQ ID NO: 59 pGal1/10 promoter system
- PGK1t terminator SEQ ID NO: 61.
- a single plasmid was digested with SacI and further purified using Zymo Research’s DNA Clean and Concentrator Kit (catalog no. D4014).
- the restriction enzyme cut site found in pAOX1, linearized the plasmid, allowing for homologous integration at AOX1 locus.
- Five microliters of linearized plasmid were used to transform 50 ⁇ L of BG-11 competent cells and plated on YPD media plates containing geneticin.
- a recombinant Pichia host cell strain (SHP079) with one copy of the gene encoding CYP207 (SEQ ID NO: 1187) was also constructed as follows.
- a single plasmid, wbplasmid231 (SEQ ID NO: 1274), containing a geneticin resistance marker, an ampicillin resistance marker, ⁇ 93 ⁇ a pUC origin of replication, and a cassette comprising of pAOX1:CYP207:tDAS2 was assembled as described above.
- the wbplasmid231 shares the same sequence as wbplasmid202 (SEQ ID NO: 1066) except for the CYP105 coding sequence (SEQ ID NO: 31) was replaced with the CYP207 coding sequence of SEQ ID NO: 1187.
- Inverse PCR was used to amplify the vector minus the CYP100 coding sequence using primers wboligos5803 and wboligos5921.
- the gene encoding CYP207 (SEQ ID NO: 1187) was amplified from SHS088 gDNA using primers wboligos5922 and wboligos5802.
- a recombinant Pichia host cell strain (SHP083) with one copy of the gene encoding CYP110 (SEQ ID NO: 1185) was also constructed as follows.
- the restriction enzyme cut site found in pAOX1, linearized the plasmid, allowing for homologous integration at AOX1 locus.
- LC/MS instrumentation and parameters LC/MS system: Agilent 6470 Triple Quadrupole LC/MS; Column: Agilent EclipsePlusC18 column 1.8 um, 3.0 x 50 mm; Mobile phase: Solvent A: H 2 O with 0.1% formic acid; Solvent B: Acetonitrile with 0.1% formic acid; Gradient: 0 - 0.2 min isocratic 40% B, 0.2 - 2.2 min gradient to 85% B, 2.2 - 2.8 min isocratic 85% B; Mode: MS/MS in SIM mode (monitoring at 391.5 for DCA and 407.5 for UCA and CA); Fragmentor: 90 volts.
- Example 8 Design and Expression of a Heterologous CYP-CPR Fusion Polypeptide in Pichia pastoris Host Cells
- This example illustrates the design and expression of a heterologous gene encoding a fusion of a cytochrome P450 (CYP) from Gibberella zeae and an endogenous cytochrome reductase (CPR) from Pichia.
- CYP cytochrome P450
- CPR endogenous cytochrome reductase
- the resulting cytosolic fused enzyme with CYP and CPR activities demonstrated an ability to carry out the bioconversion of Lithocholic acid (LCA) to Ursodeoxycholic acid (UDCA) and/or 3-Keto-lithocholic acid (3-KCA) to 3-Keto-ursodeoxycholic acid (3-KUDCA).
- LCA Lithocholic acid
- UDCA Ursodeoxycholic acid
- 3-KCA 3-Keto-lithocholic acid
- 3-Keto-ursodeoxycholic acid 3-Keto-ursodeoxycholic acid
- a polynucleotide of SEQ ID NO: 1275 was synthesized that encodes an N-terminal 31 amino acid truncation of CYP207 named trCYP207 (SEQ ID NO:1276) fused via short polypeptide linker (aa sequence: ARA) to the N-terminus of a truncated version of the Pichia pastoris endogenous reductase, trNCP1 (SEQ ID NO: 1278).
- the resulting cytosolic fused enzyme encoded by the fusion gene of SEQ ID NO: 1279 has the amino acid sequence of SEQ ID NO: 1280.
- Example 9 Design of Recombinant Pichia pastoris Host Cells that Overexpress Endogenous NCP1 Gene with CPR Activity
- This example illustrates the preparation of a recombinant Pichia host cell that heterologously expresses the engineered CYP110 polypeptide (SEQ ID NO: 1186) with CYP activity and overexpresses the endogenous cytochrome reductase NCP1 gene with CPR activity from Pichia.
Landscapes
- Chemical & Material Sciences (AREA)
- Life Sciences & Earth Sciences (AREA)
- Organic Chemistry (AREA)
- Health & Medical Sciences (AREA)
- Zoology (AREA)
- Engineering & Computer Science (AREA)
- Wood Science & Technology (AREA)
- Genetics & Genomics (AREA)
- Bioinformatics & Cheminformatics (AREA)
- Biochemistry (AREA)
- General Health & Medical Sciences (AREA)
- General Engineering & Computer Science (AREA)
- Molecular Biology (AREA)
- Biotechnology (AREA)
- Microbiology (AREA)
- Medicinal Chemistry (AREA)
- Biomedical Technology (AREA)
- Chemical Kinetics & Catalysis (AREA)
- General Chemical & Material Sciences (AREA)
- Micro-Organisms Or Cultivation Processes Thereof (AREA)
- Preparation Of Compounds By Using Micro-Organisms (AREA)
- Peptides Or Proteins (AREA)
Abstract
Priority Applications (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| EP24757530.1A EP4665352A2 (fr) | 2023-02-14 | 2024-02-13 | Polypeptides recombinés pour la biosynthèse d'acide ursodésoxycholique (udca) |
| CN202480016475.1A CN120826230A (zh) | 2023-02-14 | 2024-02-13 | 用于熊去氧胆酸(udca)生物合成的重组多肽 |
Applications Claiming Priority (4)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US202363484926P | 2023-02-14 | 2023-02-14 | |
| US63/484,926 | 2023-02-14 | ||
| US202363586338P | 2023-09-28 | 2023-09-28 | |
| US63/586,338 | 2023-09-28 |
Publications (2)
| Publication Number | Publication Date |
|---|---|
| WO2024173367A2 true WO2024173367A2 (fr) | 2024-08-22 |
| WO2024173367A3 WO2024173367A3 (fr) | 2025-01-16 |
Family
ID=92420651
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/US2024/015560 Ceased WO2024173367A2 (fr) | 2023-02-14 | 2024-02-13 | Polypeptides recombinés pour la biosynthèse d'acide ursodésoxycholique (udca) |
Country Status (3)
| Country | Link |
|---|---|
| EP (1) | EP4665352A2 (fr) |
| CN (1) | CN120826230A (fr) |
| WO (1) | WO2024173367A2 (fr) |
Family Cites Families (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| IT201900017411A1 (it) * | 2019-09-27 | 2021-03-27 | Ice S P A | Processo di preparazione di acido ursodesossicolico |
| CA3201311A1 (fr) * | 2020-11-30 | 2022-06-02 | Sandhill One, Llc | Procedes enzymatiques pour la conversion de lca et 3-kca en udca et 3-kudca |
-
2024
- 2024-02-13 WO PCT/US2024/015560 patent/WO2024173367A2/fr not_active Ceased
- 2024-02-13 CN CN202480016475.1A patent/CN120826230A/zh active Pending
- 2024-02-13 EP EP24757530.1A patent/EP4665352A2/fr active Pending
Also Published As
| Publication number | Publication date |
|---|---|
| EP4665352A2 (fr) | 2025-12-24 |
| CN120826230A (zh) | 2025-10-21 |
| WO2024173367A3 (fr) | 2025-01-16 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| CA2985641A1 (fr) | Constructions d'expression et procedes de modification genetique de la levure methylotrophique | |
| CN104854245A (zh) | 通过代谢工程生产麦角硫因 | |
| US9957497B2 (en) | Hydrocarbon synthase gene and use thereof | |
| US20240191214A1 (en) | Recombinant prenyltransferase polypeptides engineered for enhanced biosynthesis of cannabinoids | |
| AU2006300688B2 (en) | Yeast and method of producing L-lactic acid | |
| CN115851810A (zh) | 酿酒酵母从头合成柚皮素工程菌株及其构建方法与应用 | |
| CN114341350B (zh) | 用于制备熊脱氧胆酸的方法 | |
| WO2024229191A1 (fr) | Levure génétiquement modifiée à production d'érythritol accrue | |
| KR20250129721A (ko) | 효소, 살리드로사이드를 생산하는 균주 및 생산 방법 | |
| WO2009070822A2 (fr) | Cellule de pichia pastoris de recombinaison | |
| US20240401019A1 (en) | Recombinant olivetolic acid cyclase polypeptides engineered for enhanced biosynthesis of cannabinoids | |
| CN120624387A (zh) | 一种用于催化合成d-手性肌醇的酶制剂及d-手性肌醇的制备方法 | |
| WO2023004336A1 (fr) | Levure génétiquement modifiée produisant de l'acide lactique | |
| EP4665352A2 (fr) | Polypeptides recombinés pour la biosynthèse d'acide ursodésoxycholique (udca) | |
| CN115873881A (zh) | 一种产1,3-丁二醇的基因工程菌及其应用 | |
| JP7847424B2 (ja) | 耐酸性酵母遺伝子ベースの合成プロモーター | |
| CN117586933A (zh) | 一种利用甲醇培养重组盐单胞菌的方法及其应用 | |
| WO2025049529A1 (fr) | Cellules hôtes recombinantes, polypeptides et procédés de biosynthèse d'hydrocortisone | |
| WO2024183112A1 (fr) | Procédé d'expression de d-aminoacide oxydase et son utilisation | |
| WO2024196925A2 (fr) | Polypeptides recombinants pour la biosynthèse de thymohydroquinone (thq) | |
| WO2026072711A1 (fr) | Polypeptides modifiés et cellules hôtes pour la biosynthèse de la thymohydroquinone (thq) | |
| US20250270598A1 (en) | Recombinant polypeptides with prenyltransferase activity for biosynthesis of cannabinoids and hop compounds | |
| KR20220039887A (ko) | 메탄 및 자일로스를 동시 대사하는 메탄자화균의 개발 및 이를 이용한 시노린 생산방법 | |
| WO2023069921A1 (fr) | Polypeptides de thca synthase recombinants modifiés pour une biosynthèse améliorée de cannabinoïdes | |
| CN119662581B (zh) | 羰基还原酶BsCR突变体、工程菌及合成克唑替尼手性中间体的应用 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 24757530 Country of ref document: EP Kind code of ref document: A2 |
|
| WWE | Wipo information: entry into national phase |
Ref document number: 202480016475.1 Country of ref document: CN |
|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 24757530 Country of ref document: EP Kind code of ref document: A2 |
|
| WWE | Wipo information: entry into national phase |
Ref document number: 202517083814 Country of ref document: IN |
|
| WWE | Wipo information: entry into national phase |
Ref document number: 2024757530 Country of ref document: EP |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| WWP | Wipo information: published in national office |
Ref document number: 202480016475.1 Country of ref document: CN |
|
| ENP | Entry into the national phase |
Ref document number: 2024757530 Country of ref document: EP Effective date: 20250915 |
|
| WWP | Wipo information: published in national office |
Ref document number: 202517083814 Country of ref document: IN |
|
| WWP | Wipo information: published in national office |
Ref document number: 2024757530 Country of ref document: EP |