WO2007071996A2 - Sequences de signal de translocation twin-arginine (tat) de streptomyces - Google Patents

Sequences de signal de translocation twin-arginine (tat) de streptomyces Download PDF

Info

Publication number
WO2007071996A2
WO2007071996A2 PCT/GB2006/004816 GB2006004816W WO2007071996A2 WO 2007071996 A2 WO2007071996 A2 WO 2007071996A2 GB 2006004816 W GB2006004816 W GB 2006004816W WO 2007071996 A2 WO2007071996 A2 WO 2007071996A2
Authority
WO
WIPO (PCT)
Prior art keywords
tat
polypeptide
signal peptide
amino acid
protein
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/GB2006/004816
Other languages
English (en)
Other versions
WO2007071996A3 (fr
Inventor
Tracy Palmer
David Andrew Widdick
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
University of East Anglia
Original Assignee
University of East Anglia
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by University of East Anglia filed Critical University of East Anglia
Priority to US12/086,836 priority Critical patent/US20090221038A1/en
Priority to EP06831424A priority patent/EP1973932A2/fr
Publication of WO2007071996A2 publication Critical patent/WO2007071996A2/fr
Publication of WO2007071996A3 publication Critical patent/WO2007071996A3/fr
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Classifications

    • CCHEMISTRY; METALLURGY
    • C07ORGANIC CHEMISTRY
    • C07KPEPTIDES
    • C07K14/00Peptides having more than 20 amino acids; Gastrins; Somatostatins; Melanotropins; Derivatives thereof
    • C07K14/195Peptides having more than 20 amino acids; Gastrins; Somatostatins; Melanotropins; Derivatives thereof from bacteria
    • C07K14/36Peptides having more than 20 amino acids; Gastrins; Somatostatins; Melanotropins; Derivatives thereof from bacteria from Actinomyces; from Streptomyces (G)
    • CCHEMISTRY; METALLURGY
    • C07ORGANIC CHEMISTRY
    • C07KPEPTIDES
    • C07K2319/00Fusion polypeptide

Definitions

  • TWIN-ARGININE TRANSLOCATION STREPTOMYCES SIGNAL SEQUENCES
  • This invention relates to Twin-arginine translocation (Tat) signal peptides, which have been identified in the cell wall associated fraction of the gram-positive soil bacterium Streptomyces coelicolor.
  • the invention further relates to fusion polypeptides comprising a TAT signal peptide and a heterologous polypeptide.
  • the Tat system is involved in the transport of pre- folded protein substrates (Thomas et al., (2001) MoI. Microbiol 39:47 - 53). Proteins are targeted to the Tat pathway by possession of N-terminal tripartite signal peptides.
  • the signal peptides include a conserved twin-arginine motif in the N-region of Tat signal peptide. The motif has been defined as R-R-x- ⁇ - ⁇ , where ⁇ represents a hydrophobic amino acid.
  • the Tat pathway comprises the three-membrane protein TatA, TatB and TatC.
  • a fourth protein TatE forms a minor component of the Tat machinery and has a similar function to TatA.
  • Studies by several groups suggest that the major role of the Tat pathway is in the translocation of redox proteins that integrate their cofactors in the cytoplasm. Other more recent studies indicate that the Tat pathway may play a broader role in protein secretion (Ochsner et al., (2001) Proc. Natl. Acad. Sci. USA 99:8312 - 8317). Because of the ability to secrete pre-folded protein substrates, the Tat pathway represents a significant mechanism for secreting a high level of heterologous proteins.
  • Tat substrates in organisms other than Bacillus subtilits and E. coli have been based predominantly in in silico analysis of genome sequences using programs trained to recognize specific features of tat targeting sequences. While these programs are useful tools for identifying candidate Tat substrates encoded within bacterial genomes experimental verification of the in silico predicted sequences other than phenotype analysis has been lacking. Many of these predicted Tat substrates are in microorganisms of the Streptomyces genes such as S. coelicolor and S. avermitilis. Streptomyces are gram-positive spore forming microorganisms which produce a range of diverse and important secondary metabolites, including many commercially available antibiotics.
  • Streptomyces are important in the field of biotechnology because they are prolific protein secretors. Prior in silico predictions estimate between 145 and 189 proteins from S. coelicolor may be Tat dependent.
  • the present invention demonstrates the verification and importance of the Tat secretory pathway in Streptomyces and further is directed to Tat signal peptides which comprise a motif outside of the previous identified Tat signal peptide motif. Further the Tat signal peptides according to the invention maybe useful in the secretion of heterologous proteins in Streptomyces.
  • novel Tat signal peptides and methods of using the novel Tat signal peptides for producing heterologous polypeptides in a bacterial host cell as described in the appended claims.
  • the invention is directed to novel Twin Arginine Translocation (TAT) signal peptides.
  • TAT Twin Arginine Translocation
  • the invention is directed to an isolated signal peptide comprising the TAT signal sequence motif (X "1 )RR(X +2 )(X +3 )(X +4 ), wherein R is arginine, X "1 is amino acid M, H, A, P, K, R, N, T, G, S, D, Q or E; X +2 Js amino acid A, P, K, R, N, T, G, S, D, Q or E; X +3 is I 1 W, F, L, V, Y, M, C, H, A, P, N or T; and X +4 is Q, I, L, V, M or F, and wherein said motif is not within the first 35N' terminal residues of the amino acid sequence of the polypeptide.
  • R is arginine
  • X "1 is amino acid M, H, A, P, K, R, N, T, G, S, D, Q or E
  • the invention is directed to an isolated TAT signal peptide comprising the sequence motif (X "1 ) RR(X +2 ) (X +3 ) (X +4 ), wherein R is arginine, X "1 is amino acid H, A, P, K, R, N, T, G, S, D, Q E or L; X +2 is A, P, K, R, N, T, G, S, D, Q or E; X +3 is I 1 W, F, L, V, Y, M, C, H, A, P, N or T; and X +4 is T 1 G or A, and wherein the motif is within the first 35N' terminal residues of the amino acid sequence of the polypeptide.
  • R is arginine
  • X "1 is amino acid H, A, P, K, R, N, T, G, S, D, Q E or L
  • X +2 is A, P, K, R, N, T, G, S, D, Q or E
  • the TAT signal peptide comprises a sequence motif (X "1 ) RR(X +2 ) (X +3 ) (X +4 ) that is within the first 35N' terminal residues, wherein when X "1 is H then X +4 is A.
  • the TAT signal peptide comprises a sequence motif (X "1 )RR(X +2 )(X +3 )(X +4 ) that is within the first 35N' terminal residues, wherein when X "1 is L then X +4 is G.
  • the invention is directed to an isolated TAT signal peptide comprising the sequence motif (X "1 )RR(X +2 )(X +3 ) (X +4 ), wherein RR represents two adjacent arginine residue and X designates positions restrict to other selected amino acids: X "1 is M, H, A, P, K, R, N, T, G, S, D, Q, or E; X +2 is a polar amino acid residue; and X +3 and X +4 are non-polar amino acid residues, and wherein the motif is not within the first 35N' terminal residues of the amino acid sequence of the polypeptide.
  • the TAT signal polypeptide is a variant of the polypeptide having the TAT motif that is not within the first 35N' terminal residues of the amino acid sequence of the signal peptide.
  • TAT signal peptides comprising the amino acid sequences of TAT signal peptides of proteins SCO2286 (SEQ ID NO: 218), SCO3790long (SEQ ID NO: 227), and SCO6580long (SEQ ID NO: 241), SCO1590 (SEQ ID NO: 211), SCO1824 (SEQ ID NO: 213), SCO6580short (SEQ ID NO: 182), or SCO3790short (SEQ ID NO: 122).
  • the invention is directed to an isolated polynucleotide sequence comprising a polynucleotide sequence encoding a TAT signal peptide of the invention.
  • the isolated polynucleotide sequence is a nucleic acid molecule comprising a first nucleotide sequence encoding a TAT signal sequence encompassed by the invention operably linked to a second nucleotide sequence encoding a heterologous polypeptide.
  • the invention is directed to a nucleic acid molecule comprising a first nucleotide sequence encoding a TAT signal sequence encompassed by the invention operably linked to a second nucleotide sequence encoding a homologous polypeptide.
  • the invention is directed to an expression vector comprising a first nucleotide sequence encoding a TAT signal sequence encompassed by the invention operably linked to a second nucleotide sequence encoding a heterologous polypeptide.
  • the invention is directed to a bacterial host cell host cell that is genetically transformed with the recombinant expression vector encompassed by the invention.
  • the invention is directed to a fusion polypeptide comprising a TAT secretion signal peptide encompassed by the invention and a heterologous polypeptide.
  • the fusion polypeptide comprises a TAT signal peptide that is naturally expressed in Streptomyces.
  • the heterologous polypeptide is an enzyme, growth factor or hormone.
  • the enzyme may be a protease, a carbohydrase, such as amylases, cellulases, xylanases, and lipases; an isomerase, such as racemases, epimerases, tautomerases, or mutases; a transferase; a glucoamylase; a kinase, an amidase, an esterase, or an oxidase.
  • the heterologous polypeptide is not naturally associated with any secretion signal peptide.
  • the invention is directed to a method for producing a heterologous polypeptide comprising culturing host cells in culture medium under conditions suitable for producing said polypeptide, said host cells comprising an expression vector comprising a first nucleotide sequence encoding a TAT signal peptide encompassed by the invention operatively linked to a second nucleotide sequence encoding a heterologous polypeptide, and producing said heterologous polypeptide.
  • the method uses a TAT signal peptide of the invention that comprises the sequence motif (X "1 )RR(X +2 )(X +3 )(X +4 ), wherein R is arginine, X "1 is amino acid M, H, A, P, K, R, N, T, G, S, D, Q or E; X +2 is amino acid A, P, K, R, N, T, G, S, D, Q or E; X +3 is I 1 W, F, L, V, Y, M, C, H, A, P 1 N or T; and X +4 is Q, I, L, V, M or F, and wherein said motif is not within the first 35N' terminal residues of the amino acid sequence of the polypeptide.
  • R is arginine
  • X "1 is amino acid M, H, A, P, K, R, N, T, G, S, D, Q or E
  • X +2 is amino acid A, P, K, R, N, T
  • the heterologous polypeptide that is produced by the method of the invention includes a TAT signal peptide that comprises the sequence motif (X- 1 )RR(X +2 )(X +3 )(X +4 ), wherein R is arginine, X "1 is amino acid H, A; P, K, R, N, T, G, S, D, Q E or L; X +2 is A, P, K, R, N, T, G, S, D, Q or E; X +3 is I, W, F, L, V, Y, M, C, H, A, P, N or T; and X +4 is T, G or A, and wherein the motif is within the first 35N' terminal residues of the amino acid sequence of the polypeptide.
  • R is arginine
  • X "1 is amino acid H, A
  • X +2 is A, P, K, R, N, T, G
  • the step of producing the heterologous polypeptide comprises recovering the polypeptide from the culture medium.
  • the host cell is a prokaryotic cell.
  • the host cell is a Streptomyces bacterial cell.
  • the host cell is a S. coelicolor or an S. lividans bacterial ceil.
  • the method of the invention produces a heterologous polypeptide that can be an enzyme a growth factor or a hormone.
  • the invention is directed to a method for secreting a heterologous protein from a host microorganism comprising operably ligating a nucleotide acid sequence encoding the heterologous protein to a TAT signal sequence encompassed by the invention and inserting the ligated pair into an expression vector in a host microorganism, expressing the heterologous protein under the control of the TAT signal and secreting the heterologous protein from the microorganism by the TAT expression pathway.
  • the expressed heterologous protein is secreted in its correctly -folded active form.
  • the host organism is a Streptomyces strain, for example a S. lividans strain.
  • the invention provides for a method of identifying TAT signal peptides of polypeptides secreted in a microorganism comprising identifying a TAT signal peptide of a secreted polypeptide and validating the ability of said signal peptide to secrete a biological functional polypeptide.
  • testing the validity of the TAT signal peptide comprises expressing a fusion polypeptide in a host microorganism, comprising the TAT signal sequence operably linked to a heterologous polypeptide and testing the biological activity of the expressed heterologous polypeptide.
  • the heterologous polypeptide is agarase.
  • FIG. 1 A-B The tatC mutant of S. coelicolor has pleiotropic phenotypes. Colonies of the S. coelicolor wildtype (left hand plate) and ⁇ tatC strain TP4 (right hand plate) cultured on either MS (A), or R5 (B), media.
  • FIG. 2 A-B illustrates a 2-dimensional gel analysis of extracellular protein fractions isolated from the S. coelicolor M145 wild type (A) and a ⁇ tatC mutant derivative (B). Strains were cultured on R5, and proteins associated with the cell wall were separated in the first dimension by isoelectric focusing (pH gradient 4 - 7) and in the second dimension by SDS PAGE.
  • Protein spots that are circled represent proteins that migrated at specific positions in the extracellular fractions from the wild type strain but were absent from the corresponding position in extracellular fractions of the ⁇ tatC strain. These proteins spots were identified by in gel tryptic digest followed by mass spectrometry and identities of the proteins are indicated by the SCO number.
  • FIG. 3 A-D illustrates a 2-dimensional gel analysis of extracellular fractions isolated from strains grown on Complete medium (CM) (A and B) and Mannitol Soya (MS) (C and D). Proteins from the wild type S coelicolor M145 strain are shown in A and C, and proteins from the tatC mutant strains are shown in panels B and D. Protein spots that are circled represent proteins that migrated at specific positions in the extracellular fractions from the wild type strain but were absent from the corresponding position in extracellular fractions of the ⁇ tatC strain. These proteins spots were identified by in gel tryptic digest followed by mass spectrometry and identities of the proteins are indicated by the SCO number.
  • CM Complete medium
  • MS Mannitol Soya
  • FIG. 4 A-C Agarase is a Tat substrate.
  • A) the agarase signal peptide with the consecutive arginine residues of the twin-arginine motif are highlighted in bold and underlined.
  • B) S. coelicolor strain M145 (WT) and TP1 (_ ⁇ afC::Apra R ) were grown on MM-C minimal medium for 5 days and stained with Lugol solution.
  • FIG. 5 illustrates the export of agarase activity mediated by S. coelicolor signal peptides some of which are encompassed by the instant invention.
  • the signal peptide is fused to the mature sequence of agarase.
  • the y-axis shows the signal peptides from a range of cell wall associated S. coelicolor proteins (listed in Tables 3 and 4) gives the % agarase activity of the secreted protein for each fusion protein when compared with agarase including it native signal peptide (set at 100%).
  • Each construct carries the same native agarase promoter and ribosome-binding site and therefore, the activity is a measure of the efficientcy of export directed by each particular signal peptide.
  • the assay was verified using negative controls, signal peptides that do not posses twin-arginines in the signal peptide and were strongly detected in the extracellular fractions of both the wild type (M145) and ⁇ tatC mutant strain by MudPIT analysis. None of these signal peptides mediated extracellular agarase activity.
  • the signal peptides from the following proteins were also tested and were found to be negative in this assay: SCO0432, SCO0474, SCO0494, SCO0930, SCP1230, SCO1396, SCO1824, ACO1948, ACO1968, SCO2226, SCO2383, SCO2446, SCO2591 , SCO2637, SCO2786, SACO2821 , SCO3456, SCO4010, SCO4142, SCO4152, SCO4884, SCO4885, SCO5074, SCO5113. SCO5447, SCO5461 , SCO6009, SCO6644, SCO6738, and SCO7399.
  • the symbol * annotates that two versions of the annotated signal peptides (designated long and short), representing two alternative start sites, were tested (see Tables 3 and 4).
  • FIG. 6 illustrates the Tat-dependent export of agarase mediated by S. coelicolor signal peptides.
  • the signal peptide of each designated SCO protein, fused in frame with the mature sequence of agarase was expressed in two S. lividans tat+ strains (1326, lower left on each plate, or 10-164, lower right of each plate), or in the 10- 164 isogenic tatC strain (top of each plate). Strains were cultured on minimal medium containing 0.5% glucose for 5 days and stained with lugol solution. (Top Left) plate (DagA) corresponds to agarase expressed from an identical construct with its native signal peptide. Note that although expression of agarase in S.
  • nucleic acids are written left to right in 5 1 to 3' orientation; amino acid sequences are written left to right in amino to carboxy orientation, respectively.
  • polypeptide refers to a compound made up of a single chain of amino acid residues linked by peptide bonds.
  • protein as used herein may be synonymous with the term “polypeptide” or may refer, in addition, to a complex of two or more polypeptides.
  • fusion polypeptide or'Tat fusion polypeptide as used herein refers to a Tat signal peptide linked to the protein of interest. It is understood that a protein of interest refers to a heterologous protein that is operably linked to the Tat signal peptide of the invention.
  • a “signal peptide” as used herein refers to an amino-terminal extension on a protein to be secreted. Nearly all secreted proteins use an amino-terminal protein extension which plays a crucial role in the targeting to and translocation of precursor proteins across the membrane and which is proteolytically removed by a signal peptidase during or immediately following membrane transfer.
  • a “Tat signal peptide” refers to a N-terminally extended sequence which includes two consecutive arginine residues and which functions in the secretion of proteins in prefolded confirmation.
  • a “Tat signal peptide” may be interchangeably referred to as “Tat peptide” or “Tat polypeptide”.
  • a “tagged Tat fusion polypeptide” herein refers to a fusion polypeptide, which comprises a Tat signal peptide and a heterologous peptide, to which a tag sequence can be linked and used to identify transformants and/or to facilitate the purification of recombinant Tat fusion polypeptides.
  • a "protein of interest” or “polypeptide of interest” refers to the protein to be expressed and secreted by the host cell.
  • the protein of interest may be any protein which up until now has been considered for expression in prokaryotes.
  • the protein of interest may be either homologous or heterologous to the host. In the first case over expression should be read as expression above normal levels in said host. In the latter case basically any expression is of course over expression.
  • heterologous protein refers to a protein or polypeptide that does not naturally occur in a host cell.
  • heterologous proteins include enzymes such as hydrolases including proteases, cellulases, amylases, other carbohydrases, and lipases; isomerases such as racemases, epimerases, tautomerases, or mutases; transferases, kinases and phophatases.
  • the heterologous gene may encode therapeutically significant proteins or peptides, such as growth factors, cytokines, ligands, receptors and inhibitors, as well as vaccines and antibodies.
  • the gene may encode commercially important industrial proteins or peptides, such as proteases, carbohydrases such as amylases and glucoamylases, cellulases, oxidases and lipases.
  • the gene of interest may be a naturally occurring gene, a mutated gene or a synthetic gene.
  • homologous protein refers to a protein or polypeptide native or naturally occurring in a host cell.
  • the invention includes host cells producing the homologous protein via recombinant DNA technology.
  • the present invention encompasses a host cell having a deletion or interruption of the nucleic acid encoding the naturally occurring homologous protein, such as a protease, and having nucleic acid encoding the homologous protein re- introduced in a recombinant form.
  • the host cell produces the homologous protein.
  • polynucleotide or "nucleic acid molecule” includes RNA, DNA and cDNA molecules. As used herein, the term refers to a polymeric form of nucleotides of any length, either ribonucleotides or deoxyribonucleotides. This term refers only to the primary structure of the molecule and thus includes double- and single-stranded DNA and RNA.
  • modifications for example, labels which are known in the art, methylation, "caps", substitution of one or more of the naturally occurring nucleotides with an analog, internucleotide modifications such as, for example, those with uncharged linkages (e.g., methyl phosphonates, phosphotriesters, phosphoamidates, carbamates, etc.) and with charged linkages (e.g., phosphorothioates, phosphorodithioates, etc.), those containing pendant moieties, such as, for example proteins (including for e.g., nucleases, toxins, antibodies, signal peptides, poly-L-lysine, etc.), those with intercalators (e.g., acridine, psoralen, etc.), those containing chelates (e.g., metals, radioactive metals, boron, oxidative metals, etc.), those containing alkylators, those with modified linkages (e.g., alkylators, those with
  • nucleic acid segments provided by this invention may be assembled from fragments of the genome and short oligonucleotide linkers, or from a series of oligonucleotides, or from individual nucleotides, to provide a synthetic nucleic acid which is capable of being expressed in a recombinant transcriptional unit comprising regulatory elements derived from a microbial or viral operon, or a eukaryotic gene.
  • a “heterologous” nucleic acid construct or sequence has a portion of the sequence which is not native to the cell in which it is expressed.
  • Heterologous, with respect to a control sequence refers to a control sequence (i.e. promoter or enhancer) that does not function in nature to regulate the same gene the expression of which it is currently regulating.
  • heterologous nucleic acid sequences are not endogenous to the cell or part of the genome in which they are present, and have been added to the cell, by infection, transfection, microinjection, electroporation, or the like.
  • a “heterologous” nucleic acid construct may contain a control sequence/DNA coding sequence combination that is the same as, or different from a control sequence/DNA coding sequence combination found in the native cell.
  • an "vector” refers to a nucleic acid construct designed for transfer between different host cells.
  • An "expression vector” refers to a vector that has the ability to incorporate and express heterologous DNA fragments in a foreign cell. Many prokaryotic and eukaryotic expression vectors are commercially available. Selection of appropriate expression vectors is within the knowledge of those having skill in the art. Accordingly, an "expression cassette” or “expression vector” is a nucleic acid construct generated recombinantly or synthetically, with a series of specified nucleic acid elements that permit transcription of a particular nucleic acid in a target cell.
  • the recombinant expression cassette can be incorporated into a plasmid, chromosome, mitochondrial DNA, plastid DNA, virus, or nucleic acid fragment.
  • the recombinant expression cassette portion of an expression vector includes, among other sequences, a nucleic acid sequence to be transcribed and a promoter.
  • plasmid refers to a circular double-stranded (ds) DNA construct used as a cloning vector, and which forms an extrachromosomal self-replicating genetic element in many bacteria and some eukaryotes.
  • selectable marker-encoding nucleotide sequence refers to a nucleotide sequence which is capable of expression in mammalian cells and where expression of the selectable marker confers to cells containing the expressed gene the ability to grow in the presence of a corresponding selective agent.
  • promoter refers to a nucleic acid sequence that functions to direct transcription of a downstream gene.
  • the promoter will generally be appropriate to the host cell in which the target gene is being expressed.
  • the promoter together with other transcriptional and translational regulatory nucleic acid sequences are necessary to express a given gene.
  • control sequences also termed “control sequences”
  • the transcriptional and translational regulatory sequences include, but are not limited to, promoter sequences, ribosomal binding sites, transcriptional start and stop sequences, translational start and stop sequences, and enhancer or activator sequences.
  • Chimeric gene or “heterologous nucleic acid construct”, as defined herein refers to a non-native gene (i.e., one that has been introduced into a host) that may be composed of parts of different genes, including regulatory elements.
  • a chimeric gene construct for transformation of a host cell is typically composed of a transcriptional regulatory region (promoter) operably linked to a heterologous protein coding sequence, or, in a selectable marker chimeric gene, to a selectable marker gene encoding a protein conferring antibiotic resistance to transformed cells.
  • a typical chimeric gene of the present invention for transformation into a host cell, includes a transcriptional regulatory region that is constitutive or inducible, a signal peptide coding sequence, a protein coding sequence, and a terminator sequence.
  • a chimeric gene construct may also include a second DNA sequence encoding a signal peptide if secretion of the target protein is desired.
  • a nucleic acid is "operably linked" when it is placed into a functional relationship with another nucleic acid sequence.
  • DNA encoding a secretory leader is operably linked to DNA for a polypeptide if it is expressed as a preprotein that participates in the secretion of the polypeptide; a promoter or enhancer is operably linked to a coding sequence if it affects the transcription of the sequence; or a ribosome binding site is operably linked to a coding sequence if it is positioned so as to facilitate translation.
  • "operably linked" means that the DNA sequences being linked are contiguous, and, in the case of a secretory leader, contiguous and in reading phase. However, enhancers do not have to be contiguous. Linking is accomplished by ligation at convenient restriction sites. If such sites do not exist, the synthetic oligonucleotide adaptors or linkers are used in accordance with conventional practice.
  • the term "gene” means the segment of DNA involved in producing a polypeptide chain, that may or may not include regions preceding and following the coding region, e.g. 5' untranslated (5' UTR) or “leader” sequences and 3' UTR or “trailer” sequences, as well as intervening sequences (introns) between individual coding segments (exons).
  • 5' UTR 5' untranslated
  • leader leader
  • 3' UTR or “trailer” sequences as well as intervening sequences (introns) between individual coding segments (exons).
  • recombinant includes reference to a cell or vector, that has been modified by the introduction of a heterologous nucleic acid sequence or that the cell is derived from a cell so modified.
  • recombinant cells express genes that are not found in identical form within the native (non-recombinant) form of the cell or express native genes that are otherwise abnormally expressed, under expressed or not expressed at all as a result of deliberate human intervention.
  • the terms “transformed”, “stably transformed” or “transgenic” with reference to a cell means the cell has a non-native (heterologous) nucleic acid sequence integrated into its genome or as an episomal plasmid that is maintained through two or more generations.
  • the term "expression” refers to the process by which a polypeptide is produced based on the nucleic acid sequence of a gene.
  • the process includes both transcription and translation.
  • the term “introduced” in the context of inserting a nucleic acid sequence into a cell means “transfection”, or “transformation” or “transduction” and includes reference to the incorporation of a nucleic acid sequence into a eukaryotic or prokaryotic cell where the nucleic acid sequence may be incorporated into the genome of the cell (for example, chromosome, plasmid, plastid, or mitochondrial DNA), converted into an autonomous replicon, or transiently expressed (for example, transfected mRNA).
  • isolated or purified refer to a nucleic acid or polypeptide that is removed from at least one component with which it is naturally associated.
  • substantially equivalent can refer both to nucleotide and amino acid sequences, for example a mutant sequence, that varies from a reference sequence by one or more substitutions, deletions, or additions, the net effect of which does not result in an adverse functional dissimilarity between the reference and subject sequences.
  • a substantially equivalent sequence varies from one of those listed herein by no more than about 20% (i.e., the number of individual residue substitutions, additions, and/or deletions in a substantially equivalent sequence, as compared to the corresponding reference sequence, divided by the total number of residues in the substantially equivalent sequence is about 0.2 or less).
  • Such a sequence is said to have 80% sequence identity to the listed sequence.
  • a substantially equivalent, e.g., mutant, sequence of the invention varies from a listed sequence by no more than 10% (90% sequence identity); in a variation of this embodiment, by no more than 5% (95% sequence identity); and in a further variation of this embodiment, by no more than 2% (98% sequence identity).
  • Substantially equivalent, e.g., mutant, amino acid sequences according to the invention generally have at least 95% sequence identity with a listed amino acid sequence, whereas substantially equivalent nucleotide sequence of the invention can have lower percent sequence identities, taking into account, for example, the redundancy or degeneracy of the genetic code.
  • sequences having substantially equivalent biological activity and substantially equivalent expression characteristics are considered substantially equivalent.
  • activity or "biological activity” refers to an activity associated with a particular protein, such as enzymatic activity. Biological activity refers to any activity that would normally be attributed to that protein by one skilled in the art.
  • alanine A
  • arginine R
  • asparagine N
  • aspartic acid D
  • cysteine C
  • glutamic acid E
  • glutamine Q
  • G histidine
  • isoleucine I
  • leucine L
  • lysine K
  • methionine M
  • phenylalanine F
  • proline P
  • serine S
  • threonine T
  • tryptophan W
  • tyrosine Y
  • valine V
  • TATFIND refers to a Tat substrate recognition program developed to detect putative Tat substrates in bacterial genomes.
  • the program is based on the position and sequences of the Tat motif (WO 03/079007).
  • the motif as disclosed in WO 03/079007 is within the first 35 amino acids of a protein sequence, (X- 1 )R°R +1 (X +2 )(X +3 ) (X +4 ), wherein X "1 is M, H, A 1 P, K, R, N, T, G, S, D, Q or E; R 0 R +1 represent the twin-arginines; X +2 is A, P, K, R, N, T, G, S, D, Q or E; X +3 is I 1 W, F, L, V, Y, M, C, H, A, P, N or T (positively charged residues being excluded from this position) and X +4 is Q, I 1 L 1 V, M or F.
  • TatP refers to a Tat substrate recognition program that recognizes the same Tat motif as TATFIND.
  • the TatP program partially uses a neural network as well as a rule based classification used in the TATFIND program and reference is made to Bendtsen, J. D. et al., (2005), BMC Bioinformatics 6:167 - 175.
  • the present invention provides novel gram-positive microorganism secretion factors and methods that can be used in microorganisms to provide protein secretion and the production of proteins in secreted form.
  • Tat Nucleic Acids and Amino Acids The invention is based on the discovery of novel Tat signal peptides.
  • the invention comprises isolated Tat signal peptides that include, but are not limited to, a polypeptide comprising the amino acid sequence set forth as SEQ ID NO: 17, 20, 23 26 29 32 35, 38, 41 , 44, 47, 50,53, 56, 59, 62, 65, 68, 71 , 74, 77, 80, 83, 86, 89, 92, 95, 98, 101 , 104, 107, 110, 113 116, 119, 122, 125, 128 134, 134, 137, 140, 143, 146, 149, 152, 155, 158, 161 , 164, 167, 170, 173, 176, 179, 182, 185, 188, 191 , 194, 197, 200, 203, and 204-253.
  • a polypeptide comprising the amino acid sequence set forth as SEQ ID NO: 17, 20, 23 26 29 32 35, 38, 41 , 44, 47, 50,53, 56, 59, 62
  • the Tat signal peptides include polypeptides comprising an amino acid sequence encoded by any one nucleotide sequence generated by amplifying a nucleic acid using polymerase chain reaction (PCR) primer pairs having SEQ ID NO: 15 and 16, 18 and 19, 21 and 22, 24 and 25, 27 and 28, 30 and 31, 33 and 34, 36 and 37, 39 and 40, 42 and 43, 45 and 46, 48 and 49, 51 and 52, 54 and 55, 57 and 58, 60 and 61 , 63 and 64, 66 and 67, 69 and 70, 72 and 73, 75 and 76, 78 and 79, 81 and 82, 84 and 85, 87 and 88, 90 and 91 , 93 and 94, 96 and 97, 99 and 100, 102 and 103, 105 and 106, 108 and 109, 111 and 112, 114 and 115, 117 and 118, 120 and 121 , 123 and 124, 126 and 127, 129 and 130, 132 and 133,
  • the present invention is directed to certain novel Tat signal peptides that comprise sequences which include the motif (X "1 )RR(X +2 )(X +3 ) (X +4 ), wherein RR represents two adjacent arginine residue and X designates positions restrict to other selected amino acids: X "1 is M, H, A, P, K, R, N, T, G, S, D, Q, or E; X +2 is A, P, K, R, N, T, G, S, D, Q or E; X +3 is I 1 W, F, L, V 1 Y, M, C, H, A, P, N or T; and X +4 is Q, I, L, V, M or F and wherein the motif is not within the first 35N' terminal residues of the amino acid sequence of the polypeptide.
  • X "1 is M, H, A, P, K, R, N, T, G, S, D, Q, or E
  • X +2 is A, P, K, R
  • Tat signal sequences of the invention include the motif (X '1 )RR(X +2 )(X +3 ) (X +4 ), wherein RR represents two adjacent arginine residue and X designates positions restrict to other selected amino acids: X "1 is M 1 H 1 A, P 1 K, R 1 N 1 T 1 G, S 1 D, Q, or E; X +2 is a polar amino acid residue; and X +3 and X +4 are non-polar amino acid residues, and wherein the motif is not within the first 35N' terminal residues of the amino acid sequence of the polypeptide.
  • non-polar amino acids are alanine, leucine, isoleucine, valine, proline, phenylalanine, tryptophan, and methionine; and polar amino acids include glycine, serine, threonine, cysteine, tyrosine, asparagine, glutamine, arginine, lysine, histidine, aspartic acid and glutamic acid.
  • Tat signal peptides encompassed by the invention include, but are not limited to the Tat signal peptides of proteins SCO2286 - MTPANHQAPTSAPSPAPSQSSHAPELRAAARSLGRRRFLTVTGAAAALAFAVNLPAAGTA
  • RAAAWTVAAAGAGAVGVAGAPSAQAA SEQ ID NO: 227).
  • Tat signal peptide comprises the Tat signal peptide of protein SCO6580 long - MTPFTDSSRTDA
  • the invention contemplates isolated variants of the Tat signal peptides of the invention.
  • the variant of a Tat signal peptide is encoded by an alternative start site that leads to the translation of a shorter Tat signal peptide.
  • Variant Tat signal peptides include, but are not limited to the Tat signal peptide of protein SCO3790, designated SCO3790short (SEQ ID NO: 122), and the Tat signal peptide of protein
  • SCO6580 designated SCO6580 short - (SEQ ID NO: 182).
  • the invention also contemplates Tat signal peptides that comprise the motif (X "1 )RR(X +2 )(X +3 ) (X +4 ), wherein RR represents two adjacent arginine residue and X designates positions restrict to other selected amino acids: X "1 is H, A, P, K, R, N, T, G, S,
  • X +2 is A, P, K, R, N, T, G, S, D, Q or E
  • X +3 is I, W, F, L, V, Y, M, C, H, A, P, N or T
  • X +4 is T, G or A and wherein the motif is within the first 35N' terminal residues of the amino acid sequence of the polypeptide.
  • X "1 is S then X +4 will be T; in other embodiments when X '1 is H then X +4 will be A, and in other embodiments when X '1 is L then X +4 will be G.
  • the primary amino acid sequence of some of the Tat signal peptides dictates that the amino acid residues at positions X +3 and X +4 are non-polar amino acid residues.
  • the invention comprises Tat signal sequences that include the motif (X "1 ) RR(X +2 ) (X +3 ) (X +4 ), wherein RR represents two adjacent arginine residue and X designates positions restrict to other selected amino acids: X '1 is H, A, P, K, R, N, T, G, S, D, Q E or L; X +2 is A, P, K, R, N, T, G, S, D, Q or E; and X +3 and X +4 are non-polar amino acid residues.
  • Preferred Tat signal peptides include, but are not limited to the Tat signal sequence of protein SCO1590 - MGGVSRRAFTVAALSAFTLVPEASAA (SEQ ID NO: 211) and Tat signal sequence of protein SCO1824 - MTAPLSR HRRALAIPAGLAVAASLAFLP GTPAAATPAAEAA (SEQ ID NO: 213).
  • the invention also provides biologically active variants of any of the amino acid sequences set forth as SEQ ID NO: 17, 20, 23 26 29 32 35, 38, 41 , 44, 47, 50,53, 56, 59, 62, 65, 68, 71 , 74, 77, 80, 83, 86, 89, 92, 95, 98, 101 , 104, 107, 110, 113 116, 119, 122, 125, 128 131 , 134, 137, 140, 143, 146, 149, 152, 155, 158, 161 , 164, 167, 170, 173, 176, 179, 182, 185, 188, 191, 194, 197, 200, 203, and 204-253, and substantial equivalents thereof.
  • Substantial equivalent Tat signal peptides have at least about 80%, or 85%, more typically at least about 90%, 91%, 92%, 93%, or 94% and even more typically at least about 95%, 96%, 97%, 98% or 99% amino acid identity, and that retain biological activity.
  • the biological activity of a signal peptide refers to the ability of the signal peptide to translocate the polypeptide to the extracellular space. Fragments of the Tat peptides of the present invention which are capable of exhibiting biological activity are also encompassed by the present invention.
  • variants when made in reference to a polypeptide refers to any polypeptide differing from naturally occurring polypeptides by amino acid insertions, deletions, and substitutions, or combinations thereof, which can be created using, e g., recombinant DNA techniques.
  • a variant of a polypeptide can also refer to a shortened version of a polypeptide that is generated by use of an alternative start codon as a translation initiation site.
  • the alternative codon can be an in-frame or an out-of-frame start codon. Examples of variants of polypeptides that are translated from alternative start codons include but are not limited to the TAT signal polypeptides of SEQ ID NO: 122 and SEQ ID NO: 182.
  • amino acid substitutions are the result of replacing one amino acid with another amino acid having similar structural and/or chemical properties, i.e., conservative amino acid replacements.
  • Constant amino acid substitutions may be made on the basis of similarity in polarity, charge, solubility, hydrophobicity, hydrophilicity, and/or the amphipathic nature of the residues involved.
  • nonpolar (hydrophobic) amino acids include alanine, leucine, isoleucine, valine, proline, phenylalanine, tryptophan, and methionine; polar neutral amino acids include glycine, serine, threonine, cysteine, tyrosine, asparagine, and glutamine; positively charged (basic) amino acids include arginine, lysine, and histidine; and negatively charged (acidic) amino acids include aspartic acid and glutamic acid.
  • “Insertions” or “deletions” are typically in the range of about 1 to 5 amino acids. The variation allowed may be experimentally determined by systematically making insertions, deletions, or substitutions of amino acids in a polypeptide molecule using recombinant DNA techniques and assaying the resulting recombinant variants for activity.
  • insertions, deletions or non-conservative alterations can be engineered to produce altered polypeptides.
  • Such alterations can, for example, alter one or more of the biological functions or biochemical characteristics of the polypeptides of the invention.
  • such alterations may change polypeptide characteristics such as ligand-binding affinities, interchain affinities, or degradation/turnover rate.
  • such alterations can be selected so as to generate polypeptides that are better suited for expression, scale up and the like in the host cells chosen for expression.
  • the present invention also provides isolated Tat peptides encoded by the nucleic acids/polynculeotides of the present invention or by degenerate variants of the nucleic acids of the invention.
  • degenerate variants is intended nucleotide fragments that differ from a nucleic acid of the invention by nucleotide sequence but, due to the degeneracy of the genetic code, encode an identical polypeptide sequence.
  • a fusion protein or fusion polypeptide comprises a Tat signal peptide operatively linked to a polypeptide/protein of interest.
  • the term "operatively linked" is intended to indicate that the Tat signal polypeptide and the polypeptide of interest are fused in-frame with one another.
  • the Tat signal peptides are fused to the N-terminal end of the peptide of interest.
  • Polypeptides of interest include homologous or heterologous polypeptides.
  • Polypeptides of interest include full-length polypeptides that are naturally synthesized with a signal peptide, the mature form of the full-length polypeptides, and polypetides that naturally lack a signal peptide.
  • the fusion polypeptides thus comprise a Tat signal sequence that includes the motif (X "1 )RR(X +2 )(X +3 ) (X +4 ), wherein RR represents two adjacent arginine residue and X designates positions restrict to other selected amino acids:
  • X "1 is M, H, A, P, K, R, N, T, G, S, D, Q 1 or E;
  • X +2 is A, P 1 K, R 1 N, T 1 G, S 1 D, Q or E;
  • X +3 is I 1 W, F 1 L 1 V, Y 1 M, C 1 H 1 A 1 P, N or T; and
  • X +4 is Q 1 1 1 L, V 1 M or F and wherein the motif is not within the first 35N'
  • Tat signal sequences comprised by the fusion polypepitde can include the motif (X "1 )RR(X +2 )(X +3 ) (X + "), wherein RR represents two adjacent arginine residue and X designates positions restrict to other selected amino acids: X "1 is M, H 1 A 1 P 1 K 1 R 1 N 1 T, G 1 S 1 D 1 Q 1 or E; X +2 is a polar amino acid residue; and X +3 and X +4 are non-polar amino acid residues, and wherein the motif is not within the first 35N 1 terminal residues of the amino acid sequence of the polypeptide.
  • non-polar amino acids are alanine, leucine, isoleucine, valine, proline, phenylalanine, tryptophan, and methionine; and polar amino acids include glycine, serine, threonine, cysteine, tyrosine, asparagine, glutamine, arginine, lysine, histidine, aspartic acid and glutamic acid.
  • Some preferred fusion polypepitdes comprise Tat signal peptides that include, but are not limited to the Tat signal peptides of proteins SCO2286 - MTPANHQAPTSAPSPAPSQSSHAPELRAAARSLGRRRFLTVTGAAAALAFAVNLPAAGTA SAA (SEQ ID NO: 218) and SCO3790 long -
  • Tat signal peptide comprises the Tat signal peptide of protein SCO6580 long - MTPFTDSSRTDA GTDPSADGPGESLRRALGVNRRRFLSTCTAVAAGAVAAPVFGASPALAH (SEQ ID NO: 241).
  • the fusion polypeptides of the invention comprise variants of the novel Tat signal peptides.
  • the fusion polypeptide comprises a variant of a Tat signal peptide that is encoded by an alternative start site that leads to the translation of a shorter Tat signal peptide.
  • Variant Tat signal peptides include, but are not limited to the Tat signal peptide of protein SCO3790, designated SCO3790short (SEQ ID NO: 122), and the Tat signal peptide of protein SCO6580, designated SCO6580 short - (SEQ ID NO: 182).
  • the invention also contemplates fusion polypepitdes that comprise Tat signal peptides that include the motif (X "1 ) RR(X +2 ) (X +3 ) (X +4 ), wherein RR represents two adjacent arginine residue and X designates positions restrict to other selected amino acids:
  • X '1 is H 1 A, P, K, R, N, T, G, S, D, Q E or L;
  • X +2 is A, P, K, R, N 1 T, G, S, D, Q or E;
  • X +3 is I, W, F, L, V, Y, M, C, H, A, P, N or T; and
  • X +4 is T, G or A and wherein the motif is within the first 35N' terminal residues of the amino acid sequence of the polypeptide.
  • X +4 when X "1 is S then X +4 will be T; in other embodiments when X "1 is H then X +4 will be A, and in other embodiments when X "1 is L then X +4 will be G.
  • the primary amino acid sequence of some of the Tat signal peptides dictates that the amino acid residues at positions X +3 and X +4 are non-polar amino acid residues.
  • the invention comprises Tat signal sequences that include the motif (X "1 ) RR(X +2 ) (X +3 ) (X +4 ), wherein RR represents two adjacent arginine residue and X designates positions restrict to other selected amino acids: X "1 is H, A, P, K 1 R, N, T, G, S, D, Q E or L; X +2 is A, P, K, R, N, T, G 1 S, D, Q or E; and X +3 and X +4 are non-polar amino acid residues.
  • Preferred fusion polypepitdes comprise Tat signal peptides that include, but are not limited to the Tat signal sequence of protein SCO1590 - MGGVSRRAFTVAALSAFTLVPEASAA (SEQ ID NO: 211) and Tat signal sequence of protein SCO1824 - MTAPLSR HRRALAIPAGLAVAASLAFLP GTPAAATPAAEAA (SEQ ID NO: 213).
  • the fusion polypeptide comprises a Tat signal peptide that is the secretory leader sequence of polypeptides that are naturally expressed by Streptomyces that is operably linked to a heterologous polypeptide or protein of interest.
  • the fusion polypeptide comprises a Tat signal peptide and a heterologous polypeptide such as an enzyme, a growth factor or a hormone.
  • Enzymes include but are not limited to protease, a carbohydrase, such as amylases, cellulases, xylanases, and lipases; an isomerase, such as racemases, epimerases, tautomerases, or mutases; a transferase; a glucoamylase; a kinase, an amidase, an esterase, or an oxidase.
  • protease a carbohydrase, such as amylases, cellulases, xylanases, and lipases
  • an isomerase such as racemases, epimerases, tautomerases, or mutases
  • a transferase such as racemases, epimerases, tautomerases, or mutases
  • a transferase such as racemases, epimerases, tautomerases, or mutases
  • the protein of interest may be an enzyme such as a carbohydrase, such as an ⁇ - amylase, an alkaline ⁇ -amylase, a ⁇ -amylase, a cellulase; a dextranase, an ⁇ -glucosidase, an ⁇ -galactosidase, a glucoamylase, a hemicellulase, a pentosanase, a xylanase, an invertase, a lactase, a naringanase, a pectinase or a pullulanase; a protease such as an acid protease, an alkali protease, bromelain, ficin, a neutral protease, papain, pepsin, a peptidase, rennet, rennin, chymosin, subtilisin, thermolysin
  • the protein may be an aminopeptidase, a carboxypeptidase, a chitinase, a cutinase, a deoxyribonuclease, an ⁇ -galactosidase, a ⁇ -galactosidase, a ⁇ -glucosidase, a laccase, a mannosidase, a mutanase, a pectinolytic enzyme, a polyphenoloxidase, ribonuclease or transglutaminase, for example.
  • the enzyme may be a wild-type enzyme or a variant of a wild-type enzyme.
  • the enzyme may also be a hybrid enzyme, which comprises at least two fragments from different enzymes, for example a catalytic domain of one enzyme and a starch binding domain of a different enzyme or two fragments each fragment comprising part of the catalytic domain of the enzymes.
  • the fusion polypeptide of the invention comprises a Tat signal peptide, as recited herein, and an enzyme that is a protease, a carbohydrase, an isomerase, a glucoamylase, a kinase, an amidase, an esterase, or an oxidase.
  • the fusion polypeptide caomprises a Tat signal peptide and a heterologous poplypeptide that is not naturally associated with a secretion signal peptide.
  • the fusion polypeptide comprises a heterologous polypeptide that may be a therapeutic protein (i.e., a protein having a therapeutic biological activity).
  • Suitable therapeutic proteins include: erythropoietin, cytokines such as interferon- ⁇ , interferon- ⁇ , interferon- ⁇ , interferon-o, and granulocyte-CSF, GM-CSF, coagulation factors such as factor VIII, factor IX, and human protein C, antithrombin III, thrombin, soluble IgE receptor ⁇ -chain, IgG, IgG fragments, IgG fusions, IgM, IgA, interleukins, urokinase, chymase, and urea trypsin inhibitor, IGF-binding protein, epidermal growth factor, growth hormone-releasing factor, annexin V fusion protein, angiostatin, vascular endothelial growth factor-2, myeloid progenitor inhibitory factor-1 , osteoprotegerin, ⁇ -1-antitrypsin, ⁇ -feto proteins, DNase II, kringle 3 of human plasminogen, gluco
  • the invention encompasses the polynucleotides that encode the fusion polypeptides.
  • the invention is directed to isolated polynucleotides that comprise a polynucleotide sequence encoding a Tat signal peptide, as recited above, that is operably linked to a second nucleotide sequence encoding a heterologous polypeptide.
  • the Tat polynucleotides that encode the Tat polypeptides of the invention include the sequence information of the nucleic acid sequences generated by PCR using the primer pairs having SEQ ID NO: 15 and 16, 18 and 19, 21 and 22, 24 and 25, 27 and 28, 30 and 31 , 33 and 34, 36 and 37, 39 and 40, 42 and 43, 45 and 46, 48 and 49, 51 and 52, 54 and 55, 57 and 58, 60 and 61 , 63 and 64, 66 and 67, 69 and 70, 72 and 73, 75 and 76, 78 and 79, 81 and 82, 84 and 85, 87 and 88, 90 and 91 , 93 and 94, 96 and 97, 99 and 100, 102 and 103, 105 and 106, 108 and 109, 111 and 112, 114 and 115, 117 and 118, 120 and 121 , 123 and 124, 126 and 127, 129 and 130, 132 and 133, 135 and
  • the polynucleotides of the present invention also include, but are not limited to a polynucleotide comprising the protein coding sequence of the Tat signal peptides having SEQ ID NO: 254 -312.
  • the polynucleotides of the present invention also include polynucleotides that hybridize under stringent conditions to the complement of any of the polynucleotides of the invention , a polynucleotide encoding any one of the Tat signal peptides of SEQ ID NO: 17, 20, 23 26 29 32 35, 38, 41 , 44, 47, 50,53, 56, 59, 62, 65, 68, 71 , 74, 77, 80, 83, 86, 89, 92, 95, 98, 101 , 104, 107, 110, 113 116, 119, 122, 125, 128 131 , 134, 137, 140, 143, 146, 149, 152, 155, 158, 161
  • allelic variant denotes any of two or more alternative forms of a gene occupying the same chromosomal locus. Allelic variation arises naturally through mutation, and may result in polymorphism within populations. Gene mutations can be silent (no change in the encoded polypeptide) or may encode polypeptides having altered amino acid sequences.
  • An allelic variant of a polypeptide is a polypeptide encoded by an allelic variant of a gene.
  • the invention also encompasses polynucleotide fragments of the nucleic acid sequences of the invention.
  • the polynucleotide fragments can be used in various hybridization procedures or microarray procedures to identify or amplify identical or related parts of mRNA or DNA molecules from which the Tat signal peptides of the invention can be derived.
  • polynucleotides of the invention additionally include the complement of any of the polynucleotides recited above.
  • the polynucleotides of the invention also provide polynucleotides including nucleotide sequences that are substantially equivalent to the Tat polynucleotides recited above.
  • Polynucleotides according to the invention can have, e.g., at least about 65%, at least about 70%, at least about 75%, at least about 80%, more typically at least about 90%, and even more typically at least about 95%, sequence identity to a polynucleotide recited above.
  • the invention also provides the complement of such polynucleotides.
  • the polynucleotide can be DNA (genomic, cDNA, amplified, or synthetic) or RNA. Methods and algorithms for obtaining such polynucleotides are well known to those of skill in the art and can include, for example, methods for determining hybridization conditions which can routinely isolate polynucleotides of the desired sequence identities.
  • a Tat fusion polypeptide of the invention can be produced by standard recombinant DNA techniques. For example, DNA fragments coding for the different polypeptide sequences are ligated together in-frame in accordance with conventional techniques, e.g., by employing blunt-ended or stagger-ended termini for ligation, restriction enzyme digestion to provide for appropriate termini, filling-in of cohesive ends as appropriate, alkaline phosphatase treatment to avoid undesirable joining, and enzymatic ligation.
  • the fusion gene can be synthesized by conventional techniques including automated DNA synthesizers.
  • PCR amplification of gene fragments can be carried out using anchor primers that give rise to complementary overhangs between two consecutive gene fragments that can subsequently be annealed and reamplified to generate a chimeric gene sequence (see, e.g., Ausubel, et al. (eds.) CURRENT PROTOCOLS IN MOLECULAR BIOLOGY, John Wiley & Sons, 1992).
  • a Tat signal peptide-encoding nucleic acid can be cloned into such an expression vector such that the fusion moiety e.g. polypeptide of interest, is linked in-frame to the Tat signal peptide.
  • An expression vector comprising a Tat signal peptide -encoding polynucleotide can be any vector capable of expressing the polynucleotide encoding Tat signal peptide or Tat-fusion polypeptide in a selected host organism, and the choice of vector will depend on the host cell into which the expression vector is introduced.
  • the invention encompasses expression vectors that comprise a first nucleotide sequence encoding a Tat signal peptide, as recited herein, operably linked to a second nucleotide sequence encoding a heterologous polypeptide.
  • the expression vector typically includes the components of a cloning vector, such as, for example, an element that permits autonomous replication of the vector in the selected host organism and one or more phenotypically detectable markers for selection purposes.
  • the expression vector normally comprises control nucleotide sequences encoding a promoter, operator, ribosome binding site, translation initiation signal and optionally, a repressor gene or one or more activator genes.
  • control nucleotide sequences encoding a promoter, operator, ribosome binding site, translation initiation signal and optionally, a repressor gene or one or more activator genes.
  • the nucleic acid sequence the modified enzyme is operably linked to the control sequences in proper manner with respect to expression.
  • a polynucleotide in a vector is operably linked to a control sequence that is capable of providing for the expression of the coding sequence by the host cell, i.e. the vector is an expression vector.
  • the control sequences may be modified, for example, by the addition of further transcriptional regulatory elements to make the level of transcription directed by the control sequences more responsive to transcriptional modulators.
  • the control sequences may in particular comprise promoters.
  • the nucleic acid sequence encoding for the Tat signal peptide or the Tat signal peptide fusion polypepitde is operably combined with a suitable promoter sequence.
  • the promoter can be any DNA sequence having transcription activity in the host organism of choice and can be derived from genes that are homologous or heterologous to the host organism.
  • suitable promoters for directing the transcription of the modified nucleotide sequence, such as modified enzyme nucleic acids, in a bacterial host include the promoter of the Streptomyces coelicolor agarase gene dagA promoters, the promoters of the Bacillus licheniformis .alpha.-amylase gene (amyL), the aprE promoter of Bacillus subtilis, the promoters of the Bacillus stearothermophilus maltogenic amylase gene (amyM, the promoters of the Bacillus amyloliquefaciens .alpha.-amylase gene (amyQ), the promoters of the Bacillus subtilis xylA and xylB genes and a promoter derived from a Lactococcus sp. -derived promoter including the P170 promoter.
  • the promoter of the Streptomyces coelicolor agarase gene dagA promoters the promoters of the
  • the Tat fusion polypeptide may, in addition, can comprise a tag sequence that is fused to the C-terminus of the Tat fusion polypeptide to generate a tagged Tat fusion polypeptide.
  • tag sequences can be used to identify transformants and/or to facilitate the purification of recombinant Tat fusion polypeptides.
  • the Tat fusion polypeptide it may be expressed to contain a tag such as those of maltose binding protein (MBP), glutathione-S-transferase (GST) or thioredoxin (TRX), or as a His tag.
  • Kits for expression and purification of such fusion proteins are commercially available from New England BioLab (Beverly, Mass.), Pharmacia (Piscataway, N.J.) and Invitrogen, respectively.
  • the Tat fusion polypeptide can also be tagged with an epitope and subsequently purified by using a specific antibody directed to such epitope.
  • One such epitope (“FLAG. RTM.) is commercially available from Kodak (New Haven, Conn.).
  • Another tag that can be used in the invention is the c-myc tag, as is described in the examples.
  • This invention further provides expression vectors comprising at least a fragment of the polynucleotides set forth above and host cells or organisms transformed with these expression vectors.
  • Useful vectors include plasmids, cosmids, lambda phage derivatives, phagemids, and the like, that are well known in the art.
  • the invention also provides a vector including a polynucleotide of the invention and a host cell containing the polynucleotide.
  • the vector contains an origin of replication functional in at least one organism, convenient restriction endonuclease sites, and a selectable marker for the host cell.
  • the present invention further includes novel expression vectors comprising promoter elements operatively linked to polynucleotide sequences encoding a Tat fusion protein/polypeptide that comprises a Tat signal peptide of the invention and a protein of interest.
  • the invention encompasses plasmids pTDW92, pTDW73, pTDW121 , pTDW102, pTDW119, pTDW74, pTDW75, pTDW48, pTDW76, pTDW103, pTDW118, pTDW77, pTDW49, pTDW51 , pTDW52, pTDW78, pTDW91 , pTDW79, pTDW104, pTDW90, pTDW89, pTDW88, pTDW80, pTDW53, pTDW106, pTDW81, pTDW82, pTDW83, pTDW84, pTDW
  • the expression host strain for heterologously expressed protein will be a Streptomyces strain.
  • the genus Streptomyces includes all members known to those skilled in the art, including but not limited to S. coelicolor, S. lividans, In some embodiments, the expression host strain is S. coelicolor. in other embodiments, the expression host strain is S. lividans.
  • the genetic elements and tools required for heterologous expression of proteins is known in the art for Streptomyces, including expression vectors, promoters, as well as fermentation protocols (Glibert et al., (1995) Crit. Rev. Biotechnology 15:13 - 39).
  • suitable bacterial host organisms are gram positive bacterial species such as Bacillaceae including Bacillus subtilis, Bacillus Hcheniformis, Bacillus lentus, Bacillus brevis, Bacillus stearothermophilus, Bacillus alkalophilus, Bacillus amyloliquefaciens, Bacillus coagulans, Bacillus lautus, Bacillus megate ⁇ um and Bacillus thuhngiensis, and other Streptomyces species such as Streptomyces murinus, S. rubiginosus, S. griseus, S. avermitilis, lactic acid bacterial species including Lactococcus spp.
  • Bacillaceae including Bacillus subtilis, Bacillus Hcheniformis, Bacillus lentus, Bacillus brevis, Bacillus stearothermophilus, Bacillus alkalophilus, Bacillus amyloliquefaciens, Bacillus coagulans, Bacillus lautus, Bacill
  • a DNA sequence encoding a Tat signal sequence is linked to the 5' end of the polynucleotide encoding the protein of interest, such that the signal sequence directs the secretion of the polypeptide sequence via the Tat pathway.
  • Hopwood et al. (1985) Genetic Manipulation of Streptomyces: Laboratory Manual The John lnnes Foundation, Norwich, UK, Fernandez- Abalos et al., (2003) Microbiol. 149:1623 - 1632; Connell, N.D. (2001) Curr. Opin.
  • the present invention further provides host cells genetically transformed with the vectors of the invention, which may be, for example, an expression vector that contains the Tat signal polynucleotides and/or the Tat fusion polynucleotides of the invention.
  • the invention encompasses a bacterial host cell transformed with an expression vector, as recited above.
  • the vector may be, for example, in the form of a plasmid, a viral particle, a phage etc.
  • the host cells and transformed cells can be cultured in conventional nutrient media modified as appropriate for activating promoters, selecting transformants or amplifying Tat signal and/or Tat fusion polynucleotides.
  • the specific culture conditions, such as temperature, pH and the like will be apparent to those skilled in the art.
  • the present invention still further provides host cells genetically engineered to express the polynucleotides of the invention, wherein such polynucleotides are in operative association with a regulatory sequence heterologous to the host cell which drives expression of the polynucleotides in the cell.
  • the host cell can be a higher eukaryotic host cell, such as a mammalian cell, or a lower eukaryotic host cell, such as a yeast cell.
  • the host cell is a prokaryotic cell, such as a bacterial cell.
  • the bacterial host cells may be from gram positive or gram negative bacteria.
  • the invention encompasses host cells from gram positive bacteria.
  • a number of types of gram positive cells that may act as suitable host cells for expression of the Tat signal peptides and/or Tat fusion polypeptides include, for example, the Strep ⁇ omyces, Bacillus and Lactococcus species recited herein.
  • the bacteria generally used are Bacillus subtilis, S. lividans.
  • the polypeptides of the invention are expressed in S. lividans cells. If the protein is made in bacteria, it may be necessary to modify the protein produced therein, for example by phosphorylation or glycosylation of the appropriate sites, in order to obtain a biologically active protein. Such covalent attachments may be accomplished using known chemical or enzymatic methods.
  • marker gene expression suggests that a gene of interest is also present, its presence and expression should be confirmed.
  • heterologous protein such as agarase
  • recombinant cells containing the insert can be identified by the absence of marker gene function.
  • a marker gene can be placed in tandem with nucleic acid encoding the secretion factor under the control of a single promoter. Expression of the marker gene in response to induction or selection usually indicates expression of the secretion factor as well.
  • host cells which contain the coding sequence for a secretion factor and express the protein may be identified by a variety of procedures known to those of skill in the art. These procedures include, but are not limited to, DNA-DNA or DNA-RNA hybridization and protein bioassay or immunoassay techniques which include membrane- based, solution-based, or chip-based technologies for the detection and/or quantification of the nucleic acid or protein.
  • Means for determining the levels of secretion of a heterologous or homologous protein in a gram-positive host cell and detecting secreted proteins include, using either polyclonal or monoclonal antibodies specific for the protein. Examples include enzyme- linked immunosorbent assay (ELISA), radioimmunoassay (RIA) and fluorescent activated cell sorting (FACS). These and other assays are described, among other places, in Hampton R et al (1990, Serological Methods, a Laboratory Manual, APS Press, St Paul MN) and Maddox DE et al (1983, J Exp Med 158:1211).
  • ELISA enzyme- linked immunosorbent assay
  • RIA radioimmunoassay
  • FACS fluorescent activated cell sorting
  • Means for producing labeled hybridization or PCR probes for detecting specific polynucleotide sequences include oligo labeling, nick translation, end-labeling or PCR amplification using a labeled nucleotide.
  • the nucleotide sequence, or any portion of it may be cloned into a vector for the production of an mRNA probe.
  • Such vectors are known in the art, are commercially available, and may be used to synthesize RNA probes in vitro by addition of an appropriate RNA polymerase such as T7, T3 or SP6 and labeled nucleotides.
  • reporter molecules or labels include those radionuclides, enzymes, fluorescent, chemiluminescent, or chromogenic agents as well as substrates, cofactors, inhibitors, magnetic particles and the like.
  • Patents teaching the use of such labels include US Patents 3,817,837; 3,850,752; 3,939,350; 3,996,345; 4,277,437; 4,275,149 and 4,366,241.
  • recombinant immunoglobulins may be produced as shown in US Patent No. 4,816,567 and incorporated herein by reference.
  • an enzyme reporter assay for identifying a secreted polypeptide is described herein. While the present invention describes an agarase reporter assay for determining the presence of secreted polypeptides, it is understood that any assay that uses any one reporter described above can be used to identify the secreted polypetides.
  • the enzymatic activity of a secreted TAT fusion polypeptide can be determined by contacting the secreted polypeptide with a substrate and detecting the reaction product produced by the action of the enzyme on the substrate. In addition, verifying that the correct enzyme or other reporter molecule is secreted can be accomplished by performing mass spectroscopy of the secreted protein.
  • Host cells transformed with polynucleotide sequences encoding heterologous or homologous protein may be cultured under conditions suitable for the expression and recovery of the encoded protein from cell culture.
  • "Recovering" a polypeptide from a culture medium refers to collecting the polypeptide in the culture medium into which it was secreted by the host cell.
  • the polypeptide can also be recovered from a lysate prepared from the host cells and further purified.
  • a secreted polypeptide may be recovered from the cell wall fraction prepared according to methods known in the art.
  • One skilled in the art can readily follow known methods for isolating polypeptides and proteins in order to obtain one of the isolated polypeptides or proteins of the present invention.
  • the invention also encompasses fusion polypeptides that comprise a Tat signal peptide and a heterologous peptide and a polypeptide domain that will facilitate purification of soluble proteins (Kroll DJ et al (1993) DNA Cell Biol 12:441-53).
  • purification facilitating domains include, but are not limited to, metal chelating peptides such as histidine-tryptophan modules that allow purification on immobilized metals (Porath J (1992) Protein Expr Purif 3:263-281), protein A domains that allow purification on immobilized immunoglobulin, and the domain utilized in the FLAGS extension/affinity purification system (Immunex Corp, Seattle WA).
  • a cleavable linker sequence such as Factor XA or enterokinase (Invitrogen, San Diego CA) between the purification domain and the heterologous protein can be used to facilitate purification.
  • the invention is directed to a method for producing a heterologous polypeptide comprising culturing host cells in culture medium under conditions suitable for producing the heterologous polypeptide, wherein the host cells contain an expression vector that comprises a first nucleotide sequence encoding a TAT signal peptide encompassed by the invention operatively linked to a second nucleotide sequence encoding a heterologous polypeptide, and producing the heterologous polypeptide.
  • the method uses a TAT signal peptide of the invention that comprises the sequence motif (X "1 )RR(X +2 )(X +3 )(X +4 ), wherein R is arginine, X "1 is amino acid M, H, A, P, K, R, N, T, G, S, D, Q or E; X +2 is amino acid A, P, K, R, N, T, G, S, D, Q or E; X +3 is I, W, F, L, V, Y, M, C, H, A, P, N or T; and X +4 is Q, I, L, V, M or F, and wherein the motif is not within the first 35N' terminal residues of the amino acid sequence of the polypeptide.
  • R is arginine
  • X "1 is amino acid M, H, A, P, K, R, N, T, G, S, D, Q or E
  • X +2 is amino acid A, P, K, R, N, T
  • the heterologous polypeptide that is produced by the method of the invention includes a TAT signal peptide that comprises the sequence motif (X- 1 )RR(X +2 )(X +3 )(X +4 ), wherein R is arginine, X "1 is amino acid H, A, P, K, R, N, T, G, S, D, Q E or L; X +2 is A 1 P, K, R, N 1 T 1 G, S 1 D, Q or E; X +3 is I 1 W, F 1 L, V, Y, M, C 1 H 1 A 1 P 1 N or T; and X +4 is T, G or A 1 and wherein the motif is within the first 35N' terminal residues of the amino acid sequence of the polypeptide.
  • the step of producing the heterologous polypeptide comprises recovering the polypeptide from the culture medium.
  • the host cell is a prokaryotic cell. In other embodiments, the host cell is a Streptomyces bacterial cell. In yet other embodiments, the host cell is a S. coelicolor or an S. lividans bacterial cell. In further embodiments of this aspect, the method of the invention produces a heterologous polypeptide that can be an enzyme a growth factor or a hormone.
  • any signal sequence that directs a nascent polypeptide into a secretory pathway may be used in the present invention. It is to be understood that the any one of the newly identified signal sequences are encompassed by the invention.
  • the signal peptides that are specifically contemplated by the instant invention included the Tat- dependent signal peptides. Specific examples include but are not limited to the Tat signal peptides listed in Tables 2, 3, and 6.
  • AtatC mutant S. coelicor strains To determine the impact of Tat-mediated protein export in S. coelicolor, marked mutations in each of the tatA, tatB and tatC genes and umarked in-frame deletions of tatB and tatC were constructed.
  • S. coelicolor tat deletion strains were prepared as follows. Antibiotic-marked and unmarked in-frame deletions were constructed using the method of Datsenko and Wanner (Datsenko et al., (2000) Proc. Natl. Acad. Sci. USA 97:6640 - 6645) modified as described by Gust et al., (Gust et al., (2003) Proc. Natl. Acad. Sci. USA 100:1541 - 1546).
  • Cosmid 141 which carries the tatA and tatC genes, and cosmid P8, which carries the tatB gene were mutated with apramycin replacement cassettes prepared for apramycin replacement cassettes prepared for tatA, tatB, and tatC genes by using primers listed in Table 1.
  • the mutated cosmids were transferred by mating into S. coelicolor 1 MI 45, and single-crossover recombinants were selected for on MS medium containing apramycin and kanamycin (Berks et al., (2003) Adv Microbiol Physiol 47:187-254).
  • Double crossover recombinants subsequently were selected by several rounds of growth on nonselective media followed by selection for colonies that were apramycin-resistant and kanamycin sensitive. Loss of the tat genes in each strain was confirmed by PCR, and the strains were designated TP1 ( ⁇ te£C::ApraR), TP2 ( ⁇ f ⁇ 64::ApraR), and TP3 ( ⁇ tatB::ApraR). Markerless strains were constructed by the method of Gust et al. (2), loss of the apramycin-resistance cassette was confirmed by PCR, and the strains were designated TP4 ( ⁇ tatC) and TP5 ( ⁇ tatB).
  • the apramycin resistance cassette from plJ773 was amplified using the tatA, tatB or tatC primers, which resulted in a PCR fragments where the final 39bp at each end corresponded to the flanking regions of the target genes.
  • the PCR products were transformed by electroporation into strain BW25113/plJ790 harboring either cosmid 141, which carries the tatA and tatC genes, orP8, which carries the tatB gene, and transformants were selected for apramycin resistance.
  • the mutated cosmids were extracted, transformed into E.
  • Marker-less strains were constructed by transforming either cosmid 141 harboring the ⁇ f ⁇ £C::Apra R allele or cosmid P8 harboring the ⁇ faf ⁇ ::Apra R allele into E. coli strain DH5 ⁇ carrying plasmid pBT340, inducing expression of the FLP recombinase by growth at 42°C and subsequent testing for colonies- that were apramycin sensitive.
  • the cosmids carrying the unmarked alleles were then transformed into E. coli strain ET12567/pUZ8002 and mated into the appropriate S. coelicolor ApraR-marked tat deletion strain with selection for kanamycin resistance.
  • Kanamycin resistant colonies were grown for several rounds without selection and colonies that were sensitive to both apramycin and kanamycin were identified.
  • the loss of the apramycin resistance cassette was confirmed by PCR and the strains were designated TP4 ( ⁇ tatC) and TP5 ( ⁇ tatB).
  • the gross phenotype of all of the coelicolor tat mutants were similar to those observed for the tatB and tatC mutants of S. lividans (Schaerlaekens et al., (2001) J Bacteriol 183:6727-6732; (Schaerlaekens et al., (2004) Microbiology 150:21-31; and Figure
  • the primer sequences used to assemble a reporter construct based on the promoter and structural gene for agarase, dagA are listed in table 1. Initially a PCR product (amplified with primers Agarase-F and AgaraseL eadnde covering the dagA promoter was digested with Hindlll and EcoRI and cloned into similarly digested pBluescriptll (Stratagene).
  • a second PCR product corresponding to the AadA streptomycin resistance gene of plJ778 plus its promoter (amplified using primers AadAFnde and AadARbam with plJ778 as the template), was digested with Ndel and BamHI and cloned into the above construct that had been pre-digested with the same enzymes.
  • a third PCR fragment (amplified with primers AgaraseLEADBAM and AgaraseRxba) covering the region encoding the DagA mature protein sequence was digested with BamHI and Xbal and cloned into this to give plasmid pTDW45.
  • Plasmid pTDW45 carries a fragment that corresponds to the dagA promoter and region of dagA encoding the mature protein sequence separated by aadA gene and flanked by BgIII sites.
  • the BgIII fragment from pTDW45 was cloned into pSET152 (Bierman et al., (1992) Gene 116:43 - 49) that has been previously digested with BamHI to give pTDW46.
  • This shuttle vector can replicate in E. coli and be mated into S. coelicolor where it integrates into the chromosome site-specifically.
  • Plasmid pU6902dagA carries a C-terminally myc eptitope-tagged derivative of agarase under control of the thiostrepton-inducible promoter PtipA.
  • the Agarase gene was amplified using primers kddagaf and kddagamyc (Table 1) digested with Ndel and BamHI and cloned into similarly digested plJ6902 (Huang et al., (2005) MoI Microbiol 58:1276- 1287).
  • the plJ6902dagA PtipA-dagAmyc allele was transferred to the ⁇ C31 site on the chromosome of M145 and TP4 as described in Bierman et al., supra.
  • Protein were prepared for 2D gel analysis or MUDPIT by following the method of (Hesketh et al., (2002) MoI. Microbiol. 46:917 - 932; Yu et al., (2003) Anal. Chem. 75:6023 -6028; and Perkins et al., (1999) Electrophoresis 20:3551 - 3567) S. coelicolor strains M145 or TP1 tatC were grown by inoculating 10 6 spores onto sterile cellophane disks placed on the surface of complete medium (CM), R5 or mannitol soya (MS) media.
  • CM complete medium
  • MS mannitol soya
  • the protein samples were resuspended in IEF sample-buffer (8M urea, 0.5% CHAPS, 0.2% DTT, 0.5% IPG buffer pH 4 - 7 (Amersham Biosciences), 0.002% bromophenol blue) and protein concentration was determined using the Biorad DC protein assay.
  • the proteins in the samples were separated in the first dimension by their differing isoelectric points. 500 - 1000 ⁇ g samples of protein were loaded onto Amersham Biosciences 18cm lmmobline Drystrip pH 4 - 7 or pH 6 - 11 iso-electric focusing gels and resolved by electrophoresis for 33 kVh using an Amersham Biosciences Ettan IPGphor iso- electric focusing unit.
  • Foci ussed strips were treated as described by Hesketh et al., (2002) supra and then separated in the second dimension by SDS-PAGE using Amersham Biosciences DALT 12.5% gels. The gels were stained using a colloidal Coomassie Blue stain and scanned with a Proxpress Proteomic Imaging System scanner for later comparison. Protein spots of interest were subsequently excised from the gel, digested with trypsin and identified by MALDI-TOF peptide mass fingerprint analysis as described previously (Hesketh et al., (2002) supra).
  • Typical 2D gel electrophoretograms for the cell wall-associated proteome of these strains cultured on R5 medium are shown in Figure 2 and for MS and CM media in Figure 3. Significant differences in the staining patterns were observed between the two strains, and proteins that were present in M 145 but absent from the iatC strain clearly represented candidates for Tat substrates. Putative Tat-targeted proteins were identified by MALDI-TOF mass spectrometry (after in-gel digestion with trypsin); several of them are marked in Figs. 2 and 3.
  • Electrophoresis 21:935-948 cell lysis, as well as substrate modification (see below) and up-regulation of Sec substrates, also may contribute to the presence of additional extracellular proteins in the AtatC strain that are absent in M145 (Table 4).
  • additional extracellular proteins in the AtatC strain that are absent in M145 (Table 4).
  • Several of the exported proteins were detected as multiple spots in the M145 sample, possibly as a result of posttranslational modification or proteolysis, and it was not uncommon for one or more of the additional spots for a particular protein to be absent from cell wall fractions of the AtatC strain.
  • Proteins are listed by SCO number and the signal peptide sequences are given in the right column. Twin-arginine dipeptides, where present, are shown in bold. Where multiple arginine dipeptides are present, only the most plausible one is marked. Method of observation indicates which technique was used to identify a particular protein (2D is two- dimensional gel electrophoresis; MudPIT is multidimensional protein identification technology).
  • TatFind and/or TatP provided for SCO numbers indicates whether a particular protein has been identified as a putative Tat substrate by either of the two prediction programs TatFind 1.4 or TatPlO (information regarding the programs can be found on the world wide web at signalfind.org and cbs.dtu.dk/services/TatP-1.0, respectively).
  • Method of observation indicates which technique was used to identify a particular protein (2D is two-dimensional gel electrophoresis; MudPIT is multidimensional protein identification technology. Bracketed after MudPIT is indicated "tatC if the protein was found in the cell-wall fraction of the 'tatC mutant strain only, and 'mututal' if it was found in the cell-wall fractions of both the M145 and the 'tatC mutant strains.
  • the annotation TatFind and/or TatP provided for SCO numbers indicates whether a particular protein has been identified as a putative Tat substrate by either of the two prediction programs TatFind or TatP.
  • This protein migrated in a unique position on two-dimensional gel analysis of the extracellular fraction from the M145 strain. However it was detected in the tatC mutant strain by MudPIT analysis.
  • TrIeSe proteins have also been identified in the extracellular fraction of S. coelicolor grown in liquid medium (Kim et al., (2005) J Bacteriol 187: 2957-2966).
  • Example 5
  • RapiGest was denatured prior to MudPIT mass spectrometry by the addition of 40 ⁇ l of 500 mMHCI to give a concentration between 30 - 50 mM (pH ⁇ 2) followed by incubation at 37 0 C for 45 min.
  • the cloudy samples containing hydrolyzed detergent were centrifuged at 13,000 rpm for 10 min in a BioFuge TM (Heraeus) and the supematants carefully removed for chromatographic separation.
  • Samples were loaded onto a biphasic column comprising a strong cation exchange phase (SCX) and a reverse phase. Peptides were eluted stepwise from the SCX phase by using increasing concentrations of salt onto the reverse phase. A reverse-phase gradient then was generated and peptides were eluted into a Q-ToF2 mass spectrometer (Micromass, Manchester, U.K.) The data from each reverse-phase gradient was combined and searched by using MASCOT (Matrix Science, London, U.K. (Perkins et al., Electrophoresis 20:3551-3567 (1999)). The details for the analysis of the samples are as follows.
  • a 75 ⁇ m PicoFrot capillary (NewObjective, Inc.) was packed first with 90 mm of Symmetry C18 300A reverse-phase material (Waters, Ltd, UK) followed by 30 mm of PartiSphere strong cation exchange material (Whatman) using a pressure packing device.
  • the resulting biphasic micro capillary column was equilibrated to 5% acetonitrile/0.1% formic acid. After loading the sample the column was mounted on the Z-spray ion source of a Q-ToF2 mass spectrometer (Micromass, Manchester, UK) and inline with a capillary HPLC system (CapLC, Water Ltd., UK).
  • a fully automated 9-step chromatography run was carried out, with the mass spectrometer operating in data- dependent modes during each reverse phase elution.
  • the three buffer solutions used for chromatography were 0.1 % formic acid (buffer A), 100% acetonitrile/0.1% formic acid (buffer b) and 50OmM ammonium acetate/0.1% formic acid (buffer c). Elution was performed using increasing concentrations of buffer C followed by a reverse-phase gradient. Buffer C concentrations were 0, 10, 25, 50, 75, 100, 200, 300 and 50OmM.
  • MS/MS collision-induced dissociation
  • Protein scores were derived from peptide ion scores as a non-probabilistic basis for ranking proteins.
  • 2D is two-dimensional gel electrophoresis; MudPIT is multidimensional protein identification technology. Bracketed after MudPIT is indicated "tatC if the protein was found in the cell-wall fraction of the tatC mutant strain only, 'M145' if the protein was found in the cell-wall fraction of the M145 strain only and 'mutual' if it was found in the cell-wall fractions of both the M145 and the tatC mutant strains.
  • TlIeSe proteins have also been identified in the extracellular fraction of S. coelicolor grown in liquid medium (Kim DW, Chater K, Lee KJ, Hesketh, A (2005) J Bacteriol 187: 2957-2966).
  • Example 6 A Tat Transport Assay Based on Agarase
  • Sec-dependent signal peptides are not normally recognized by the Tat machinery, and Tat-targeted proteins usually are folded and, therefore, are not normally compatible for Sec-dependent export.
  • a reporter-based assay for Tat transport was designed to address directly whether the group of 43 putative Tat substrates were indeed synthesized with bona fide Tat signal peptides.
  • Embedding of S. coelicolor colonies into agar is a phenomenon associated with secretion of the enzyme agarase, which degrades agar to smaller oligosaccharides. As the agar is broken down around the colony, this causes the colony to sink down into the medium.
  • Agarase is encoded by the dagA gene, and the protein product bears an N- terminal signal peptide containing an apparent twin-arginine motif
  • MVNRRDLIKWSAVALGAGAGLAGPAPAAHAIAD SEQ ID NO: 313
  • Figure 4A To determine whether agarase activity could be used in a reporter assay to test the ability of a signal peptide to secrete a protein via the TAT pathway, Applicants tested whether agarase is a TAT-dependent substrate. a) First, extracellular agarase activity of wild type M145 S. coelicor was compared to that in a tatC mutant (TP1 ( ⁇ fafC::Apra R ).
  • S. coelicor strains M145 and TP1 were grown on MM-C minimal medium for 5 days and were stained with lugol solution.
  • the dagA gene was placed under control of the tipA thistrepton-inducible promoter and incorporated onto the chromosome in a single copy at the ⁇ C31 attachment site. After integration of the construct into S. coelicolor strain M145 (wild type) and TP4 ( ⁇ tatC), harboring plJ6902-c/ao// ⁇ in a single copy, thiostrepton was added to induce expression of agarase. Transcription of dagA in S.
  • the TAT targeting signal peptide of DagA was swapped for the Sec-targeting signal peptides of three S. coelicolor proteins: SCO3053, SCO5660, and SCO6199, and the agarase activity determined as described above.
  • DagA can be used as a reporter exclusively for TAT-mediated protein secretion.
  • Example 7 ldenfication of bona fide TAT-exported proteins and signal peptides using the TAT Transport Assay Based on Agarase.
  • the two proteomic techniques above identified putative (a total of 43) proteins that could have been transported by the Tat pathway (Table 3). To identify whether these proteins contained bona fide Tat-targeting signal peptides, the reporter-transport assay based on agarase activity was used. Fusions of the signal peptides of each of the putative TAT-targeted proteins to mature agarase were constructed and expressed in a nonagarase-producing host strain (Streptomyces lividans 1326) and assessed for the production of agarase activity using the lugol-based plate test.
  • Tat dependency of Streptomyces coelicolor agarase was demonstrated by growing strain M145 and TP4 harboring either plL6902 or plJ6902dagA on minimal medium containing glucose, which strongly represses expression of the native agarase. Approximately 10 4 spores of each strain were streaked or spotted onto minimal medium containing glucose, additionally supplemented with apramycin and thiostrepton (to induce expression of the myc-tagged agarase) and grown at 37°C for 72 hours. And then stained with Lugol solution (Sigma) for 45 min. Plasmids encoding signal peptide-agarase fusion proteins (listed in Table 2) were mated into S.
  • MM-C minimal Medium (MM-C) (10g Agar, 1g (NH 4 ) 2 S ⁇ 4 , 0.5g K 2 HPO, 0.2g FeSO 4 JH 2 O in 1L) lacking a carbon source other than agar.
  • GIMP open source software distributed under a GNU general public license at ⁇ www.gimp.org>
  • Tat-dependent signal peptides lacking consecutive arginine residues have been reported (e.g., Ignatova et al., (2002) 291 :146-149).
  • a total of seven exported proteins were identified in the cell wall fraction of M145 that did not contain obvious twin-arginine motifs in their signal peptides (Table 3).
  • the agarase test was applied to six of these seven signal peptides, but none were found to direct export of active agarase.
  • Tat-dependent proteins in S. coelicolor. This number represents 30% of all of the exported proteins that we detected in the cell wall fraction of the taf strain and clearly demonstrates that the Tat pathway is a major protein translocation route.
  • the agarase reporter assay developed here and used to validate possible Tat-targeting signal peptides is particularly powerful because it is not only facile and rapid but also semiquantitative and, therefore, provides a ready measure of transport efficiency for a particular Tat-targeting signal peptide. It is also anticipated that the agarase reporter system will facilitate the exploitation of the Tat pathway for heterologous protein production.
  • Tat-exported proteins listed in Table 6 represent a broad spectrum of functional classes: several are predicted to be involved in phosphate and carbohydrate metabolism, nutrient transport, and lipid metabolism. However, unlike the well characterized E. coli system, only 3 of the 27 Tat substrates detected here are likely to be cofactor-containing and, remarkably, 2 of these (SCO6272 and SCO6281) are associated with a type I modular polyketide synthase gene luster, indicating a role in secondary metabolism and, therefore, are not expected a priori to be exported. Substrates of the twin-arginine translocation pathway have hitherto not been found associated with a secondary metabolite gene cluster.
  • Tat substrates identified here include two proteins involved in peptidoglycan metabolism, SCO0736 and SCO1172, the latter being a probable cell wall amidase.
  • SCO1172 As a Tat substrate may well account for the fragility of the Streptomyces tat mutants.
  • Tat secretome This likely would enable a further proteomic analysis of the "Tat secretome” in this organism and the assignment of Tat-targeted proteins not associated with the cell wall.

Landscapes

  • Chemical & Material Sciences (AREA)
  • Organic Chemistry (AREA)
  • Health & Medical Sciences (AREA)
  • General Health & Medical Sciences (AREA)
  • Biochemistry (AREA)
  • Biophysics (AREA)
  • Life Sciences & Earth Sciences (AREA)
  • Genetics & Genomics (AREA)
  • Medicinal Chemistry (AREA)
  • Molecular Biology (AREA)
  • Proteomics, Peptides & Aminoacids (AREA)
  • Gastroenterology & Hepatology (AREA)
  • Micro-Organisms Or Cultivation Processes Thereof (AREA)
  • Peptides Or Proteins (AREA)
  • Preparation Of Compounds By Using Micro-Organisms (AREA)

Abstract

La présente invention concerne de nouveaux polypeptides de signal Tat et des procédés pour utiliser les polypeptides de signal Tat pour produire des polypeptides hétérologues. La présente invention concerne en outre un nouvel essai indicateur pour déterminer l'activité biologique des protéines sécrétées.
PCT/GB2006/004816 2005-12-20 2006-12-20 Sequences de signal de translocation twin-arginine (tat) de streptomyces Ceased WO2007071996A2 (fr)

Priority Applications (2)

Application Number Priority Date Filing Date Title
US12/086,836 US20090221038A1 (en) 2005-12-20 2006-12-20 Twin-Arginine Translocation (TAT) Streptomyces Signal Sequences
EP06831424A EP1973932A2 (fr) 2005-12-20 2006-12-20 Sequences de signal de translocation twin-arginine (tat) de streptomyces

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
US75187205P 2005-12-20 2005-12-20
US60/751,872 2005-12-20

Publications (2)

Publication Number Publication Date
WO2007071996A2 true WO2007071996A2 (fr) 2007-06-28
WO2007071996A3 WO2007071996A3 (fr) 2007-11-08

Family

ID=38189013

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/GB2006/004816 Ceased WO2007071996A2 (fr) 2005-12-20 2006-12-20 Sequences de signal de translocation twin-arginine (tat) de streptomyces

Country Status (3)

Country Link
US (1) US20090221038A1 (fr)
EP (1) EP1973932A2 (fr)
WO (1) WO2007071996A2 (fr)

Cited By (6)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
WO2009147015A1 (fr) * 2008-05-29 2009-12-10 Henkel Ag & Co. Kgaa Micro-organisme à sécrétions optimisées
WO2011135370A1 (fr) 2010-04-30 2011-11-03 University Of East Anglia Sécrétion bactérienne
EP2426199A2 (fr) 2006-10-20 2012-03-07 Danisco US Inc. Oxydases de polyol
US20130045871A1 (en) * 2010-03-18 2013-02-21 Cornell University Engineering correctly folded antibodies using inner membrane display of twin-arginine translocation intermediates
US8383568B2 (en) 2005-10-21 2013-02-26 Danisco Us Inc. Polyol oxidases
CN111499688A (zh) * 2020-04-14 2020-08-07 江南大学 一种信号肽及其在生产α-淀粉酶中的应用

Families Citing this family (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN109504690B (zh) * 2019-01-16 2022-06-21 河南省商业科学研究所有限责任公司 参与调控链霉菌act产量的胁迫、信号和转录调控基因
WO2025175008A1 (fr) * 2024-02-14 2025-08-21 The Texas A&M University System Nouveaux peptides antimicrobiens dans une colle de fixation de tiques et protéines immunomodulatrices de tiques

Non-Patent Citations (3)

* Cited by examiner, † Cited by third party
Title
BENDTSEN JANNICK DYRLOV ET AL: "Prediction of twin-arginine signal peptides" BMC BIOINFORMATICS, BIOMED CENTRAL, LONDON, GB, vol. 6, no. 1, 2 July 2005 (2005-07-02), page 167, XP021000762 ISSN: 1471-2105 *
DILKS KIERAN ET AL: "Prokaryotic utilization of the twin-arginine translocation pathway: a genomic survey." JOURNAL OF BACTERIOLOGY FEB 2003, vol. 185, no. 4, February 2003 (2003-02), pages 1478-1483, XP002445523 ISSN: 0021-9193 *
PALMER T ET AL: "Export of complex cofactor-containing proteins by the bacterial Tat pathway" TRENDS IN MICROBIOLOGY, ELSEVIER SCIENCE LTD., KIDLINGTON, GB, vol. 13, no. 4, April 2005 (2005-04), pages 175-180, XP004842094 ISSN: 0966-842X *

Cited By (7)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US8383568B2 (en) 2005-10-21 2013-02-26 Danisco Us Inc. Polyol oxidases
EP2426199A2 (fr) 2006-10-20 2012-03-07 Danisco US Inc. Oxydases de polyol
WO2009147015A1 (fr) * 2008-05-29 2009-12-10 Henkel Ag & Co. Kgaa Micro-organisme à sécrétions optimisées
US20130045871A1 (en) * 2010-03-18 2013-02-21 Cornell University Engineering correctly folded antibodies using inner membrane display of twin-arginine translocation intermediates
WO2011135370A1 (fr) 2010-04-30 2011-11-03 University Of East Anglia Sécrétion bactérienne
CN111499688A (zh) * 2020-04-14 2020-08-07 江南大学 一种信号肽及其在生产α-淀粉酶中的应用
CN111499688B (zh) * 2020-04-14 2021-05-04 江南大学 一种信号肽及其在生产α-淀粉酶中的应用

Also Published As

Publication number Publication date
WO2007071996A3 (fr) 2007-11-08
US20090221038A1 (en) 2009-09-03
EP1973932A2 (fr) 2008-10-01

Similar Documents

Publication Publication Date Title
Müller et al. The Tat pathway in bacteria and chloroplasts
CA2793978C (fr) Expression de proteines de toxines recombinantes en forte quantite
Calvio et al. Swarming differentiation and swimming motility in Bacillus subtilis are controlled by swrA, a newly identified dicistronic operon
Yamamoto et al. Localization of the vegetative cell wall hydrolases LytC, LytE, and LytF on the Bacillus subtilis cell surface and stability of these enzymes to cell wall-bound or extracellular proteases
EP2440573B1 (fr) Souche de bacillus permettant une production accrue de protéine
Nakano et al. Multiple pathways of Spx (YjbD) proteolysis in Bacillus subtilis
US20080248525A1 (en) Twin-Arginine translocation in bacillus
JP2016501550A (ja) Crm197に関する方法及び組成物
JP2023123512A (ja) 組み換えerwiniaアスパラギナーゼの製造のための方法
EP0941349B1 (fr) Procedes de production de polypeptides dans des mutants de cellules de bacilles induisant la surfactine
JP2016063844A (ja) ブレビバチルス属細菌を用いたプロテインa様蛋白質の生産方法
CN101657538A (zh) 在芽孢杆菌属细胞中获得遗传感受态的方法
EP1973932A2 (fr) Sequences de signal de translocation twin-arginine (tat) de streptomyces
JP7016552B2 (ja) 組換えタンパク質の分泌を増加させる方法
Kodama et al. Bacillus subtilis AprX involved in degradation of a heterologous protein during the late stationary growth phase
CA2295878C (fr) Augmentation de la production de proteines dans des micro-organismes gram-positifs
CA2296689C (fr) Augmentation de la production de proteines dans les micro-organismes gram-positifs
CA2507307C (fr) Augmentation de la production de proteines secg dans le bacillus subtilis
WO2011135370A1 (fr) Sécrétion bactérienne
US6630328B2 (en) Increasing production of proteins in gram-positive microorganisms
Readnour et al. Evolution of Gram+ Streptococcus pyogenes has maximized efficiency of the Sortase A cleavage site
Brockmeier New strategies to optimize the secretion capacity for heterologous proteins in Bacillus subtilis
KR102084054B1 (ko) 재조합 단백질의 분비를 증가시키는 방법
US6506579B1 (en) Increasing production of proteins in gram-positive microorganisms using SecG
Berger et al. Inactivation of genes encoding extracellular proteases in Bacillus 2 halodurans BhFC01 and the impact on its modified flagellin type III 3 secretion pathway towards improving peptide expression 4

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application
WWE Wipo information: entry into national phase

Ref document number: 12086836

Country of ref document: US

NENP Non-entry into the national phase

Ref country code: DE

WWE Wipo information: entry into national phase

Ref document number: 2006831424

Country of ref document: EP

WWP Wipo information: published in national office

Ref document number: 2006831424

Country of ref document: EP