WO2005011576A2 - Chromatographie numerique - Google Patents

Chromatographie numerique Download PDF

Info

Publication number
WO2005011576A2
WO2005011576A2 PCT/US2004/023908 US2004023908W WO2005011576A2 WO 2005011576 A2 WO2005011576 A2 WO 2005011576A2 US 2004023908 W US2004023908 W US 2004023908W WO 2005011576 A2 WO2005011576 A2 WO 2005011576A2
Authority
WO
WIPO (PCT)
Prior art keywords
peptides
proteins
sets
mixture
selecting
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/US2004/023908
Other languages
English (en)
Other versions
WO2005011576A3 (fr
Inventor
Fred E. Regnier
Jiri Ademec
Xiang Zhang
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Beyond Genomics Inc
Original Assignee
Beyond Genomics Inc
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Beyond Genomics Inc filed Critical Beyond Genomics Inc
Publication of WO2005011576A2 publication Critical patent/WO2005011576A2/fr
Publication of WO2005011576A3 publication Critical patent/WO2005011576A3/fr
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Classifications

    • GPHYSICS
    • G01MEASURING; TESTING
    • G01NINVESTIGATING OR ANALYSING MATERIALS BY DETERMINING THEIR CHEMICAL OR PHYSICAL PROPERTIES
    • G01N33/00Investigating or analysing materials by specific methods not covered by groups G01N1/00 - G01N31/00
    • G01N33/48Biological material, e.g. blood, urine; Haemocytometers
    • G01N33/50Chemical analysis of biological material, e.g. blood, urine; Testing involving biospecific ligand binding methods; Immunological testing
    • G01N33/68Chemical analysis of biological material, e.g. blood, urine; Testing involving biospecific ligand binding methods; Immunological testing involving proteins, peptides or amino acids
    • G01N33/6803General methods of protein analysis not limited to specific proteins or families of proteins
    • G01N33/6848Methods of protein analysis involving mass spectrometry
    • G01N33/6851Methods of protein analysis involving laser desorption ionisation mass spectrometry
    • GPHYSICS
    • G01MEASURING; TESTING
    • G01NINVESTIGATING OR ANALYSING MATERIALS BY DETERMINING THEIR CHEMICAL OR PHYSICAL PROPERTIES
    • G01N33/00Investigating or analysing materials by specific methods not covered by groups G01N1/00 - G01N31/00
    • G01N33/48Biological material, e.g. blood, urine; Haemocytometers
    • G01N33/50Chemical analysis of biological material, e.g. blood, urine; Testing involving biospecific ligand binding methods; Immunological testing
    • G01N33/68Chemical analysis of biological material, e.g. blood, urine; Testing involving biospecific ligand binding methods; Immunological testing involving proteins, peptides or amino acids
    • G01N33/6803General methods of protein analysis not limited to specific proteins or families of proteins
    • G01N33/6842Proteomic analysis of subsets of protein mixtures with reduced complexity, e.g. membrane proteins, phosphoproteins, organelle proteins
    • GPHYSICS
    • G01MEASURING; TESTING
    • G01NINVESTIGATING OR ANALYSING MATERIALS BY DETERMINING THEIR CHEMICAL OR PHYSICAL PROPERTIES
    • G01N33/00Investigating or analysing materials by specific methods not covered by groups G01N1/00 - G01N31/00
    • G01N33/48Biological material, e.g. blood, urine; Haemocytometers
    • G01N33/50Chemical analysis of biological material, e.g. blood, urine; Testing involving biospecific ligand binding methods; Immunological testing
    • G01N33/68Chemical analysis of biological material, e.g. blood, urine; Testing involving biospecific ligand binding methods; Immunological testing involving proteins, peptides or amino acids
    • G01N33/6803General methods of protein analysis not limited to specific proteins or families of proteins
    • G01N33/6848Methods of protein analysis involving mass spectrometry

Definitions

  • proteomics is a field of science relating to the study of the proteome.
  • the proteome refers to the entire compliment of proteins in a given sample. Sample complexity, however, can vary from a few thousand protems to tens of thousands of proteins or more. Accordingly, proteomics has accelerated the need for high throughput screening and characterization of complex biological samples, such as, e.g., serum.
  • Applications of proteomics include investigating differences in the abundance of certain proteins between two or more sample sets (e.g., disease vs. control), and protein state differences, such as, post-translation modification, splice variants, amino acid polymorphisms, etc.
  • a separation stage precedes the measurement stage of a proteome analysis.
  • Two of the most widely used separation strategies are two dimensional gel (“2-D gel”) electrophoresis and multi-dimensional liquid chromatography ("MDLC"). These strategies can be performed on either intact proteins, proteins which have been subsequently modified in the laboratory, or proteolytic fragments of proteins known as peptides. Peptides can be produced by a number of methods, such as, e.g., by chemical, enzymatic, or physical means.
  • the measurement stage of a proteome analysis typically involves a single mass spectrometric analysis or multiple mass spectrometric analyses such as tandem mass spectrometric analysis, often referred to as MS/MS.
  • Protein identification today is frequently based on the comparison of experimentally derived tryptic peptide sequences with peptide sequences generated by an in silico analysis of genomic data. For example, in a typical MDLC approach, proteins are converted to a mixture of tryptic peptides that are then fractionated chromatographically and analyzed by mass spectrometry, typically by sequencing peptides in the tandem mode of analysis (MS/MS).
  • MS/MS tandem mode of analysis
  • Analysis in the MS/MS mode is typically achieved by selecting the molecular ion of a peptide (often referred to as "the parent ion") with a first mass spectrometer (often referred to as the first dimension of mass spectrometry) and directing the parent ion into a collision cell were it collides with an inert gas such as helium.
  • the parent ion is fragmented in the collision cell to a series of fragment ions (often referred to as "daughter ions”), among which are a ladder of ions with sequentially decreasing numbers of amino acids. Two prominent sets of ions are produced. One is a sequence ladder with amino acid deletions from the C-terminal end of the peptide.
  • the daughter ions are then typically directed into a second mass spectrometer (often referred to as the second dimension of mass spectrometry) to resolve the fragmentation pattern of the parent ion, i.e., the selected peptide.
  • the problem in digesting proteins and then analyzing the peptide cleavage fragments is that too many peptides are generated for direct mass spectral analysis.
  • a mammalian cell at any given time may express as many as 20,000 proteins, and a tryptic digest, for example, generally yields an average of 50 tryptic peptides per protein. As a result, approximately 1,000,000 peptides would be generated from the mammalian cell sample.
  • each fraction from, e.g., a reversed phase chromatography ("RPC") column capable of producing 100 peaks would contain an average of 10,000 peptides. This is far too many peptides for current mass spectrometers. Going to a higher resolution RPC column with a peak capacity of 200 wouldn't solve the problem; the fractions from the RPC column would still contain an average of 5,000 peptides each. Further, when sample complexity increases to the point that multiple peptides have the same molecular weight (isobaric peptides), a series of problems begin to occur when the number of peptides exceeds the analytic capacity of the analysis system.
  • RPC reversed phase chromatography
  • the dynamic range of the analysis system can be significantly limited (for example, in the case of poorly fractionated tryptic digests of human serum, the dynamic range can be less than two orders of magnitude).
  • the MUDPIT approach is designed to capture all proteins on a cation exchange chromatography column, the column is then eluted with step gradients of increasing salt concentration to produce a series of fractions. These fractions are further fractionated by gradient elution RPC.
  • the typical objective in the MUDPIT approach is to analyze and determine the sequence of as many of the peptides in the tryptic digest of the proteome as possible. For example, in the mammalian cell example above, if 10 fractions were selected and each transferred to an RPC column with a peak capacity of 200, the RPC fractions would contain approximately between 500-1,000 peptides each. This number of peptides can still exceed the analytical capacity of current mass spectrometry systems.
  • the present invention provides methods and systems that facilitate the high throughput screening and characterization of mixtures of proteins by providing novel peptide selection approaches.
  • the invention provides methods for analyzing a mixture of proteins utilizing novel peptide selection approaches.
  • the methods generate a plurality of peptides from a mixture of proteins.
  • the methods select sets of peptides from the plurality of peptides and sequence at least a portion of the selected peptides to produce data sets of, for example, sequence information. These data sets are then used to determine, for example, the likelihood of the presence of two or more proteins in the mixture of proteins.
  • the peptide selection approach of the invention selects peptides such that two or more sets of peptides are each individually non-representative of the protein mixture. In one embodiment, peptides are selected such that two or more sets of peptides are each individually representative of less than about 75% of the proteins in the protein mixture, hi another embodiment, peptides are selected such that two or more sets of peptides are each individually representative of less than about 70% but greater than about 30% of the proteins in the protein mixture.
  • the peptide selection approach of the invention selects peptides such that two or more sets of peptides contains less than about 15% of the peptides that are known or predicted to be found in the plurality of peptides generated from the mixture of proteins, i one embodiment, the number of peptides in each of the two or more sets of peptides is in the range from about 0.1% to about 1% of the number of peptides known or predicted to be found in the plurality of peptides.
  • the peptide selection approach of the invention selects peptides such that the number of peptides in two or more sets of peptides is less than about a peptide number capacity of the sequencing method used to sequence selected peptides to produce data sets.
  • the sequencing method comprises a MALDI-MS method having a peptide number capacity in the range from about 100 to about 200.
  • the peptide sequencing method comprises an ESI-MS method having a peptide number capacity in the range from about 30 to about 60.
  • the sequencing method comprises a RPC-MS method having a peptide number capacity in the range from about 3,000 to about 40,000.
  • the sequencing method comprises a RPC-(MALDI-MS) method having a peptide number capacity in the range from about 10,000 to about 40,000. In another embodiment, the sequencing method comprises a RPC-(ESI-MS) method having a peptide number capacity in the range from about 3,000 to about 12,000. In one embodiment, the number of peptides in each of the two or more sets of peptides is less than about 150. In another embodiment, the number of peptides in each of the two or more sets of peptides is less than about 50.
  • the methods of the present invention utilize a peptide selection approach that comprises at least three different chromatographic selection techniques used in series to produce the sets of peptides from which data sets are derived to determine the likelihood of the presence of two or more proteins in the protein mixture.
  • one chromatographic selection technique selects peptides in the plurality of peptides based on peptide acidity, another selects peptides based on the presence or absence of cysteine, and yet another selects peptides based on the presence or absence of hisitdine.
  • peptides are selected through a plurality of selection techniques that select peptides on the basis of, for example, the presence or absence of specific amino acids, the average number of specific amino acids, the ratio of specific types of amino acids (such as, e.g., the ratio of the number of acidic to basic amino acids), posttranslational modifications (such as, e.g., glycosylation and phosphorylation), or combinations of thereof.
  • the selection techniques produce fractioned samples suitable for sequencing by, for example, a RPC-MS method. It is to be understood that in the present invention the sets of selected peptides need not contain a set of signature peptides that are representative of substantially all the proteins in the proteome or mixture of proteins being analyzed.
  • a preferred selection technique comprises a LC affinity capture technique that captures Z+ components in an input analyte and allows Z- components to flow-through, where Z+ components are peptides containing Z (where Z, for example, is an amino acid) and Z- components are peptides not containing Z.
  • Such selection techniques are often referred to as orthogonal selection techniques by analogy to the concept of orthogonal vectors in mathematics.
  • chromatographic selection techniques are also preferred that produce a small number of peaks.
  • a preferred chromatographic selection technique produces only two peaks, each peak corresponding to a set of peptides.
  • the first peak contains the peptides not retained by the chromatography column
  • the second peak contains the peptides initially retained by the column and later desorped from the column by changing, for example, the mobile phase.
  • the step of selecting sets of peptides further includes using a reversed phase chromatography (RPC) technique.
  • the step of sequencing includes ionizing peptides in a set of peptides and selecting ionized peptides that have a mass-to-charge ratio in a first mass-to-charge ratio range (the parent ions) with a first dimension of mass spectrometry (such as, e.g., a time-of-flight mass spectrometer (TOF-MS), or quadrapole mass spectrometer), fragmenting the parent ions to produce daughter ions, and generating with a second dimension of mass spectrometry a data set comprising a fragmentation spectrum of the daughter ions.
  • a first dimension of mass spectrometry such as, e.g., a time-of-flight mass spectrometer (TOF-MS), or quadrapole mass spectrometer
  • the step of sequencing includes ionizing peptides in a set of peptides and selecting ionized peptides that have an ion mobility in a first ion range of mobilities (the parent ions) with a first dimension of mass spectrometry (such as, e.g., an ion mobility spectrometer (IMS)), fragmenting the parent ions to produce daughter ions, and generating with a second dimension of mass spectrometry a data set comprising a fragmentation spectrum of the daughter ions.
  • a first dimension of mass spectrometry such as, e.g., an ion mobility spectrometer (IMS)
  • Suitable ionization techniques for ionizing peptides include, but are not limited, electrospray ionization (ESI), matrix-assisted laser desoprtion/ionization (MALDI), fast atom bombardment (FAB), electron ionization (El), chemical ionization, desorption chemical ionization (DCI), Negative-ion chemical ionization, field desorption (FD), field ionization (FI), Secondary Ion Mass Spectrometry (SIMS), and Atmospheric Pressure Chemical Ionization (APCI).
  • the step of sequencing includes subjecting a set of peptides to reverse phase chromatography to further fractionate the set of peptides prior to analysis by mass spectrometry.
  • the step of sequencing comprises loading a set of peptides into an RPC column to produce a series of fractions from the set of peptides and analyzing the fractions as they elute with a mass spectrometry instrument to produce a data set for the set of peptides.
  • the step of determining the likelihood of the presence of two or more proteins in a mixture of proteins includes comparing at least two or more of the data sets of the sequencing step to at least one of a deoxyribonucleic acid ("DNA”) database and a protein database.
  • the data sets are compared to peptide sequences derived from a DNA library.
  • the DNA sequences from either a DNA or Expressed Sequence Tag (“EST") database are converted into a protein and subsequently tryptic peptide sequences in silico and then compared with the sequence of experimentally derived peptides.
  • EST Expressed Sequence Tag
  • the data sets are compared to peptide sequences derived from a protein database.
  • tryptic peptide sequences are generated in silico and then compared with the experimentally derived peptides sequences.
  • the data sets are compared to experimentally determined tryptic peptide sequences of the proteins in a protein database.
  • the step of determining the likelihood of the presence of two or more proteins in a mixture of proteins preferably includes selecting two or more data sets such that in combination the sequence information of the selected data sets is representative of greater than about 80% of the proteins in the mixture of proteins.
  • preferably data sets are selected such that in combination the sequence information of the selected data sets is representative of greater than about 90% of the proteins in the mixture of proteins.
  • data set refers to the spectrometric data associated with one or more multi-dimensional mass spectrometric measurements.
  • a data set can comprise one or more parent ion fragmentation spectra.
  • data set includes both raw spectrometric data and data that has been processed (e.g., to remove noise, to remove baseline, to detect peaks, to normalize peaks, etc.).
  • the present invention provides an article of manufacture where the functionality of a method of the present invention is embedded as computer-readable instructions on a computer-readable medium, such as, but not limited to, a floppy disk, a hard disk, an optical disk, a magnetic tape, a PROM, an EPROM, CD-ROM, or DVD- ROM.
  • Figure 1 is a flow diagram of analyzing a mixture of proteins in accordance with various embodiments of the present invention
  • Figure 2 is a graphical representation of a relationship between the number of peptides identifiable by an analytical system and the total number of peptides in a sample
  • Figure 3 is a flow diagram of selecting proteins from a protein mixture in accordance with various embodiments of the present invention
  • Figures 4A and 4B are flow diagrams of selecting sets of peptides from a plurality of peptides according to various embodiments of the present invention
  • Figures 5A-5C are plots of the average number of protems containing certain amino acids or combinations of certain amino acids as a function of protein molecular weight based on an in silico analysis of the E.
  • Figures 5D-5G are plots of the average number of tryptic peptides containing certain amino acids or combinations of certain amino acids as a function of protein molecular weight based on an in silico analysis of the E. Coli genome;
  • Figures 6A and 6B are schematic representations of the series of chromatographic selection techniques of Example 1 used for selecting sets of peptides according one embodiment of the present invention;
  • Figure 7 A is a schematic representations of the series of chromatographic selection techniques of Example 2 used for selecting sets of peptides from a plurality of peptides generated from human serum albumin ("HSA");
  • Figure 7B is a plot of molecular weight versus predicted retention time for all the peptides of Figure 7 A analyzed together;
  • Figure 8 A is a schematic representations of the series of chromatographic selection techniques of Example 3 used for selecting sets of peptides from a plurality of peptides generated from human serum albumin ("HSA") and pregnancy zone protein (“PZP”);
  • FIG. 1 a flow chart of various embodiments of a method of analyzing a mixture of proteins according to the present invention are shown.
  • a mixture of protems 101 is subjected to a peptide generating step 103 to produce a plurality of peptides 104.
  • Suitable techniques for generating peptides from proteins include any sequence specific cleavage process. Generation of peptides from proteins can be achieved by chemical, enzymatic or physical means, including, for example, sonication or shearing.
  • a protease enzyme is used, such as trypsin, chymotrypsin, papain, gluc-C, serine proteases, thiol proteases, enoproteinase Arg C, enoproteinase Asp N, and enoproteinase Lys C, and pepsin.
  • chemical agents such as cyanogen bromide can be used to effect proteolysis.
  • the proteolytic agent can be immobilized in or on a support, or can be free in solution.
  • a preferred technique for generating a plurality of peptides from the protein mixture includes enzymatic hydrolysis of peptide bonds with trypsin to produce the plurality of peptides.
  • the plurality of peptides 104 is then subjected to a selecting step 105 to select sets of peptides from the plurality of peptides 104. At least a portion of the peptides in two or more sets of peptides are subjected to a sequencing step 107 to produce a data set for the associated set of peptides. Two or more of the data sets are then used in a determining step 109 to determine the likelihood of the presence of two or more proteins in the protein mixture 101.
  • the steps of selecting sets of peptides and sequencing peptides are preferably performed utilizing a LC-MS approach 110 that combines the selecting step 105 and sequencing steps 107.
  • RPC 112 is used to separate peptides in the sets of peptides to provide fractions for the sequencing using a mass spectrometer method.
  • RPC can be used to separate peptides in a set of peptides to provide fractions for subsequent analysis by ESI-MS or MALDI-MS.
  • the methods of the present invention further include a step of selecting a set of proteins 114 from a biological sample 116 to generate the mixture of proteins 101.
  • a wide variety of biological samples are suitable for use in the present invention, including, but not limited to, cells, body fluids (such as, e.g., serum, plasma, urine, lymph, cerebrospinal fluid, amniotic fluid, synovial fluid, sebum, and saliva), tissue, and whole organism homogenates.
  • Proteins may be selected from biological samples, for example, to remove contaminants from the sample, to remove abundant proteins not of interest, or to produce a plurality of protein mixtures for analysis.
  • Suitable protein selection techniques include, but are not limited to, size- exclusion chromatography, ion exchange chromatography, hydrophobic interaction chromatography, gel-filtration chromatography, metal affinity chromatography, dye- based affinity chromatography, antibody based affinity chromatography, and combinations thereof.
  • cibicron blue can be used to select proteins with andenosine 5 -monophosphate ("AMP") binding sites.
  • the step of selecting peptides to produce sets of peptides comprises subjecting the plurality of peptides to a series of selection techniques until at least two or more of the sets of peptides output from the selection techniques meet the analysis sample criteria of an embodiment of the present invention for a set of peptides.
  • the selecting step 105 selects peptides such that two or more sets of peptides meet the analysis sample criteria that a set of peptides is individually non-representative of the protein mixture 101.
  • the selecting step 105 selects peptides such that two or more sets of peptides meet the analysis sample criteria that a set of peptides is individually representative of less than about 25% of the proteins in the protein mixture 101. In another aspect, the selecting step 105 selects peptides such that two or more sets of peptides meet the analysis sample criteria that a set of peptides contain less than about 25% of the peptides that are known or predicted to be found in the plurality of peptides 104.
  • the selecting step 105 selects peptides such that the number of peptides in two or more sets of peptides meets the analysis sample criteria that the number of peptides is less than about a peptide number capacity of the peptide sequencing method of the sequencing step 107.
  • the selection techniques separate peptides in an input analyte into sets of peptides on the basis of the posttranslational modifications and/or amino acids present or absent from the peptides in the input analyte.
  • one set of peptides may contain components of the input analyte that flow-through a column under an initial set of conditions (e.g., temperature, eluent pH); while the another set of peptides may contain the analyte components that were retained under the initial set of conditions, but which are later eluted under different conditions.
  • the combined sets of peptides produced by a selection technique do not need to contain all of the components of the input analyte and a selection technique can produce more than two sets of peptides.
  • the fractions that elute within different retention time ranges in a chromatographic selection technique can comprise a set of peptides (e.g., one set may comprise eluates with a retention time in the range of 1 min. to 5 min., another set those eluates with a retention time in the range of 5 min. to 15 min., yet another set those eluates with a retention time in the range of 15 min. to 40 min., etc.).
  • the methods of the present invention do not require that each set of peptides produced by a selection technique meet the analysis sample criteria of an embodiment of the present invention for a set of peptides.
  • a plurality of peptides can be subject to a first selection technique that produces a first set of peptides and a second set of peptides, neither of which meet the analysis sample criteria of an embodiment of the present invention for a set of peptides.
  • the first set of peptides can serve as the input analyte for a second selection technique, which produces a third set of peptides and a fourth set of peptides.
  • the third and fourth sets of peptides may meet the analysis sample criteria of an embodiment of the present invention for a set of peptides, neither is required to do so.
  • sets of peptides are subjected selection techniques until two or more of the sets of peptides meet the analysis sample criteria of an embodiment of the present invention for a set of peptides.
  • Peptides selected by a selection technique can be described by the posttranslational modifications and/or amino acids substantially present and/or absent from the selected peptides.
  • a set of peptides is representative of about x% of the proteins in a protein mixture when about x% of the proteins in the protein mixture are known or predicted to have the same combination of posttranslational modification and/or amino acid presences and absences as the selected peptides in the set of peptides.
  • a plurality of peptides is generated from a protein mixture and an input analyte comprising the plurality of peptides is subjected to a histidine selection technique to produce: (a) a first set of peptides that contains peptides having substantially no histidines ("His") but that does not substantially contain peptides having one or more His; and (b) a second set of peptides that contains peptides having one or more His but that does not substantially contain peptides having no His.
  • His histidines
  • the first set of peptides is representative of about l-x% of the proteins in the protein mixture (i.e., the fraction of proteins know or predicted to have no His) and the second set of peptides is representative of about x% of the proteins in the protein mixture.
  • the first and second sets of peptides can each be subjected to a lectin based selection technique for selecting glycopeptides.
  • the glycopeptide selection technique produces: (al) a third set of peptides that contains peptides having substantially no His and substantially no glycosylation but that does not substantially contain peptides having one or more His or glycosylation; and (a2) a fourth set of peptides that contains peptides having substantially no His but which have glycosylation, but the fourth set of peptides does not substantially contain peptides having one or more His or which have no glycosylation.
  • the glycopeptide selection technique produces: (bl) a fifth set of peptides that contains peptides having one or more His and substantially no glycosylation but that does not substantially contain peptides having no His or peptides having glycosylation; and (b2) a sixth set of peptides that contains peptides having one or more His and glycosylation but that does not substantially contain peptides having no His or no glycosylation.
  • the third set of peptides is representative of about (l-x%)(l-z%) of the proteins in the protein mixture
  • the fourth set of peptides is representative of about (l-x%)(z%) of the proteins in the protein mixture
  • the fifth set of peptides is representative of about (x%)(l-z%) of the proteins in the protein mixture
  • the sixth set of peptides is representative of about (x%)(z%) of the proteins in the protein mixture; assuming there is no correlation between the presence or absence of His and the presence or absence of glycosylation.
  • databases predict that about 90% of the proteins in yeast, and about 76% of the proteins in humans, have at least one peptide with both cysteine and histidine; but of the peptides produced by a tryptic digest, only about 3% of these peptides in yeast, and about 6% of these peptides in humans, contain both cysteine and histidine.
  • a set of peptides produced by selecting peptides that contain both Cys and His could be said, for example, to be representative of about 90% of the proteins in a mixture of yeast proteins, or representative of about 76% of the proteins in a mixture of human proteins.
  • the peptide selection approach of the present invention selects peptides such that the number of peptides in two or more sets of peptides is less than about a peptide number capacity of the sequencing method used to sequence selected peptides and produce data sets.
  • Analytical systems that are designed to analyze multiple components generally have a capacity given by a maximum number of components which they can accommodate in a single analysis.
  • the term "peptide number capacity" refers to the maximum number of peptides that can be accommodated in a single scan of a sequencing method. Where the sequencing method is, for example, a mass spectrometry method, the peptide number capacity, P nc , can be determined from the relation,
  • P ms is peak capacity of the mass spectrometer in a single scan (at the resolution at which the peptide is to be sequenced and the scan range used in sequencing the peptide), and k is the number of isotope peaks and multiple charge state peaks appearing in the mass spectrum from each peptide.
  • the peak capacity, P ms - of a mass spectrometer would be 2000 for an instrument with unit mass resolution and set to scan over a 2000 amu range, such as, e.g., the useful peptide mass range from 500-2500 amu.
  • Increases in resolution and scan range increase the mass spectrometer peak capacity, P ⁇ , s , which can be expressed as,
  • P ms ⁇ m r, (2) where ⁇ m is the scan range and r the mass resolution. Values of k can vary with the type of mass spectrometer. For example, a cluster of 3-4 isotope peaks is often seen in
  • the sequencing method comprises utilizing an ESI-MS instrument with a peptide number capacity in the range from about 40 to about 50.
  • the sequencing method comprises utilizing a MALDI-MS instrument with a peptide number capacity in the range from about 100 to about 150.
  • Suitable mass spectrometers include, for example, time-of-flight, quadrupole, RF multipole, magnetic sector, and electrostatic mass spectrometers.
  • the sequencing method is a chromatography-mass spectrometry method (such as, for example, RPC-MS) the peptide number capacity, P nc , can be determined from the relation, where P ⁇ and k are as described previously, and P cll is the peak capacity of the chromatography system, such as, e.g., a RPC.
  • the sequencing method comprises utilizing an RPC-MS sequencing method having a peptide number capacity in the range from about 8,000 to about 10,000 (such as, for example, using a RPC column with a peak capacity of about 200 and an ESI-MS instrument).
  • the sequencing technique comprises utilizing an RPC-MS sequencing technique having a peptide number capacity in the range from about 10,000 to about 20,000 (such as, for example, using a RPC column with a peak capacity of about 200 and a MALDI-MS instrument).
  • the peak capacity of the chromatography system, P C is the combined peak capacity of the system.
  • P cll ⁇ is the peak capacity in the first dimension
  • Prp C is the peak capacity of the reversed phase chromatography column
  • Pj is the peak capacity of the i th dimension of chromatography for an n dimensional chromatography system.
  • the maximum number of peptides that a mass spectrometer or chromatography-mass spectrometry system can accommodate is typically less than the peptide number capacity P nc because of isobaric peptides even though these isobaric peptides are of different sequence.
  • isobaric peptides are typically present because the probability of isobaric peptides increases with the total number of peptides, N sa , in a sample of peptides.
  • Figure 2 is a plot 200 of the number of components identifiable, N l , by a chromatography-mass spectrometry (such as, e.g., a RPC-MS system) or a mass spectrometry system as a function of with the number of peptides, N sa , in the peptide sample submitted to the system.
  • a chromatography-mass spectrometry such as, e.g., a RPC-MS system
  • mass spectrometry system a mass spectrometry system as a function of with the number of peptides, N sa , in the peptide sample submitted to the system.
  • the number of components identified, Ni d increases with the number of peptides, N sa , in the peptide sample, as illustrated by Region A 202 of Figure 2.
  • the rate at which Ni increases begins to decrease and the number of identifiable components approaches a maximum, C max , as illustrated by Region B 204 of Figure 2.
  • C max represents the number of isobaric peptides in the peptide sample.
  • Ni d the number of peptides in a sample that can be identified
  • N sa the total number of peptides in the sample
  • Ni d C max N sa /(J+N sa ) (6)
  • N sa increases, so does the number of isobaric peptides and the relationship between N d and N sa is no longer linear, as illustrated by Region B 204.
  • Region B 204 increasing numbers of peptides are being identified (Ni d is increasing) but some are not because of mass overlap. As N sa continues to increase the number of isobaric peptides and mass overlap increases and the number of components being identified, Ni d , begins to decrease, as illustrated by Region C 206.
  • the maximum number of peptides that can be accommodated in a single scan and still be in Region B 204 of Figure 2 (i.e., the experimentally achievable peptide number capacity) with an ESI-MS instrument is typically in the range from about 40 to about 50 peptides, and is typically in the range from about 100 to about 150 with a MALDI-MS instrument.
  • the methods of the present invention further include a step of selecting a set of proteins from a biological sample to generate a mixture of proteins.
  • a wide variety of techniques can be used to select a set of proteins from a biological sample including, but not limited to size-exclusion chromatography, ion exchange chromatography, hydrophobic interaction chromatography, gel-filtration chromatography, metal affinity chromatography, dye-based affinity chromatography, antibody based affinity chromatography, and combinations thereof.
  • a set of proteins is selected from a sample using a series of chromatographic selection techniques. For example, some of the difficulties associated with analyzing the proteome of human blood are the high concentrations of human serum albumin ("HSA”) and transferrins.
  • HSA human serum albumin
  • FIG. 3 a schematic flow diagram 300 of one embodiment of removing one or more abundant proteins from a human blood sample to produce one or more mixtures of proteins is illustrated.
  • the human blood sample 302 is subjected to a series of affinity LC selection techniques 304 including an affinity column to capture an abundant protein (Protein A, which can be any protein) 305, an anti-HSA antibody affinity column to capture HSA 307 and an anti-transferrin antibody affinity column to capture transferrin 309.
  • the resultant eluate 310 is further fractionated through, for example, a mixed mode ion exchange protein separation technique 312.
  • the resultant eluate 310 is input to an anion exchange (“AEX") column 313 to produce a first set of proteins 314 from the components captured by and later eluted from the AEX column, and a second set of proteins 316 from the flow-through.
  • AEX anion exchange
  • the second set of proteins 316 is input to a cation exchange ("CEX") column 317 to produce a third set of proteins 318 from the flow-through and a fourth set of proteins 319 from the components captured by and later eluted from the CEX column.
  • the first set of proteins 314 (comprising primarily positively charged proteins), the third set of proteins 318 (comprising primarily neutral proteins), and the fourth set of proteins 319 (comprising primarily negatively charged proteins) are preferably desalted using reverse-phase columns 321, 323, 325 to produce protein mixtures 326, 327, 328 for analysis.
  • the peptide selecting step comprises a series of chromatographic selection techniques which are coupled in tandem to produce sets of peptides.
  • each chromatographic selection technique selects peptides on the basis of the presence or absence of one or more of the following: (1) specific amino acids; (2) the average number of specific amino acids; (3) the ratio of specific types of amino acids (such as, e.g., the ratio of the number of acidic to basic amino acids); and (4) posttranslational modifications (such as, e.g., glycosylation and phosphorylation). Selection of specific amino acids will generally be of broadest utility in studying protein expression and protein transport.
  • One approach to specific amino acid based selection is to select peptides with a typically low abundance amino acid such as, for example, cysteine and histidine.
  • One concept in specific amino selection is that the peptides containing the specific amino acid(s) can serve as signature peptides for the parent protein from which they were derived.
  • the selection of specific amino acids will generally be of less utility in studying post-translational modifications, mutations, and protein degradation unless the particular amino acid(s) being selected is contained or predicted to be contained in a peptide being selected.
  • a wide variety of chromatographic selection techniques are suitable for use in the present invention, including, but not limited to, liquid-chromatography, electrophoresis, and gas-chromatography.
  • suitable chromatographic selection techniques include liquid-chromatography separation techniques such as affinity selection of low abundance amino acids (such as, e.g., cysteine and histidine), anion exchange to select acidic peptides, and strong anion exchange in tandem with strong cation exchange to select neutral peptides.
  • suitable chromatographic selection techniques also include liquid-chromatography separation techniques directed towards posttranslational modifications such as lectin based selection for selecting glycopeptides, and avidin based selection for selecting phosphopetides.
  • Cysteine Selection separate peptides on the basis of whether they contain one or more cysteines ("Cys").
  • DNA databases indicate that in yeast, about 92% of all proteins and about 9% of the peptides produced from these proteins by a tryptic digest contain or are predicted to contain Cys. DNA databases indicate that in humans, about 89% of all proteins and about 17% of the peptides produced from these proteins by a tryptic digest contain or are predicted to contain Cys.
  • One preferred approach to cysteine selection uses an affinity selection approach, hi one embodiment, an isotope coded affinity tag is used to select peptides containing one or more Cys.
  • Cys containing peptides are labeled with a biotin affinity group derivatized with a sulfhydryl-specific iodoacetate moiety and a preferably having an isotope coded linker (e.g., linkers with either eight hydrogens 1H or eight deuteriums 2 H).
  • a biotin affinity group derivatized with a sulfhydryl-specific iodoacetate moiety and a preferably having an isotope coded linker (e.g., linkers with either eight hydrogens 1H or eight deuteriums 2 H).
  • a reagent suitable for isotope coded affinity tagging of Cys-containing peptides includes the ICATTM reagent available from Applied Biosystems, Inc., Foster City, California, hi one embodiment of cysteine selection, peptides containing one or more Cys are initially retained on an avidin affinity column and peptides with no Cys are eluted to produce a set of peptides. hi this embodiment, after substantially all the Cys absent peptides have been eluted, the column is then washed to elute peptides containing one or more Cys to produce another set of peptides.
  • Suitable chromatographic selection techniques also include the use, for example, of chemical modification to increase or decrease the number of peptides retained by a chromatographic column.
  • the number of Cys containing peptides retained by an avidin affinity column can be reduced through a sulfhydryl derivatization process.
  • cysteine based selection to first reduce proteins and then derivatize them with a biotinylated alkylating agent. When this approach is reversed so that proteins are first derivatized with a biotinylated alkylating agent and then reduced, only free sulfhydryl groups on the proteins are biotinylated.
  • cysteine based selection is a covalent chromatography procedure for the selection of Cys-containing peptides such as has been reported by S. Wang, X. Zhang, and F.E. Regnier in Journal of Chromatography A, vol. 949, pp. 153- 160 (2002). In this approach, Cys-containing peptides were selected by covalent chromatography using thiol disulfide exchange.
  • Histidine selection techniques separate peptides on the basis of whether they contain one or more histidine (“His"). According to DNA and protein databases, His should occur in about 97% of all proteins in yeast, and constitute about 15% of the peptides in a tryptic digest. In humans, databases indicate that His should occur in about 83% of all proteins and in about 17% of all peptides in a tryptic digest.
  • Histidine selection uses an affinity selection approach. In one embodiment, immobilized affinity metal chromatography (“IMAC”) is used. In one approach, an IMAC column carrying iminodiacetate complexed with copper II (“Cu +2 " or "Cu(II)”) is used to capture His-containing peptides.
  • IMAC immobilized affinity metal chromatography
  • an IMAC column can also capture Cys, tryptophan (“Trp”) and N-terminal amino groups
  • steps such as, for example, the reduction, alkylation, and acylation steps in Global Isotope Stable Tagging ("GIST") be used to enhance the selectivity of Cu(II) loaded IMAC columns for His-containing peptides.
  • Alkylation of cysteine can greatly reduce its binding to Cu(II) loaded MAC columns, whereas, acylation of primary amine groups in peptides reduces cooperative, electrostatic binding to carboxyl groups in IMAC columns. This latter reduction in cooperative electrostatic binding is a significant factor in reducing tryptophan binding.
  • reduction and alkylation steps can be used to substantially eliminate Cys binding to Cu(IT) loaded IMAC columns.
  • acylation can be used to introduce acyl groups that substantially block the binding of Tip and N-terminal amino groups to Cu(II) loaded IMAC columns.
  • Other types of binding in Cu (IT) loaded IMAC columns arise from chelation on His and hydrophobic interactions.
  • the number of histidine containing peptides selected by a Cu( ⁇ )-IMAC column can be reduced by acylation as well. Free amine groups work cooperatively with the binding histidine on Cu(II) to affect adsorption.
  • Acidic, Basic and Neutral Peptide Selections Selection techniques based on the average number of specific amino acids or the ratio of specific types of amino acids include, but are not limited to, selection by anion exchange chromatography to select acidic peptides, selection by Ga +3 -IMAC to select acidic peptides and selection of neutral peptides by a combination of strong anion exchange ("SAX") and strong cation exchange (“SCX”) chromatography.
  • SAX strong anion exchange
  • SCX strong cation exchange
  • E. Coli the genome predicts that about 99.7% of all proteins contain either one glutamate or aspartate. Recalling that substantially all tryptic peptides contain a carboxyl group, substantially all Asp and/or Glu containing peptides produced by a tryptic digest will contain two carboxyl groups. Examination of the E. Coli genome also predicts that about 98.0% of all protems have at least two Asp and/or Glu, about 92.2% have at least three Asp and/or Glu, and about 78.9% have at least four Asp and/or Glu.
  • acidic peptides are selected using an anion exchange chromatography selection process.
  • a strong anion exchange column and a 10 mM buffer at pH 7.5 is used to capture acidic peptides.
  • substantially none of the peptides with only one acidic amino acid are captured, whereas substantially all of the peptides with three or more acidic amino acids (Asp or Glu) will be captured, hi one approach, rather than trying to further fractionate the mixture of captured peptides with varying numbers of acidic amino acids, captured peptides are eluted in a single step with the initial buffer made up to 500 mM in NaCI.
  • the anion exchange chromatography process increases or decreases the number of peptides selected by a SAX column by manipulating the selectivity of the sorbent with mobile phase additives. For example, since SAX columns select peptides based on net charge, peptides that are more anionic will be more strongly retained and require higher ionic mobile phases for elution. In one approach, the ionic strength of the mobile phase is increased to decrease the number of acidic peptides retained.
  • the number of acidic amino acids necessary for a peptide to be retained by the SAX column also increases. Since the probability of acidic amino acid occurrence decreases with increasing numbers of acidic amino acids there are fewer of these peptides and the number of peptides captured will decrease.
  • the primary amine groups in the peptides of the input analyte are acylated causing them to become more anionic and resulting in more peptides being retained by the SAX column, hi yet another approach, the pH of the mobile phase is decreased to increase the number of peptides retained by the SAX column.
  • the selecting step includes using a selection technique based on the ratio of specific types of amino acids that separates neutral peptides from acidic and basic peptides.
  • a strong anion exchange column and a strong cation exchange column are coupled in series and eluted with a low ionic strength buffer at a fixed pH to elute a set of peptides in the flow-through that are near net neutral in charge.
  • Peptides that are either strongly anionic or strongly cationic will be captured and can be later eluted to produce another set of peptides.
  • the highest stringency of selection can be achieved when the buffer concentration is at or below 10 mM; and more peptides will be eluted as the ionic strength of the mobile phase is increased.
  • acylation of primary amine groups of the peptides in an input analyte will cause these peptides to be less acidic.
  • Such acylation results in different peptides being selected by the SCX-SAX tandem column set, the selection criteria being established by a number of parameters such as the pH of the buffer being used, and increases the stringency of the neutral peptide selection, e.g., more peptides will be retained by the SCX-SAX tandem column set.
  • the selecting step includes using a selection technique that selects peptides based on posttranslational modifications (such as, e.g., glycosylation and phosphorylation).
  • Posttranslational modifications are known to be involved in, for example, cellular signal transduction and modulation of protein function, transport and lifetime.
  • addition of carbohydrate groups to certain amino acids can affect protein function, and can be associated with a metabolic or disease state of a cell.
  • reversible phosphorylation is known to be involved in the regulation of cellular signal transduction, gene expression, metabolism, and cell growth.
  • the posttranslational modifications that result from glycosylation can be complex, resulting in a single amino acid sequence being present in multiple glycoforms.
  • saccharides can form both heterogeneous and homogeneous carbohydrate structures
  • oligosaccharides can assemble stepwise at serine and threonine residues to form O-linked glycosylation sites
  • oligosaccharides can preassemble and attach at asparagine residues to form N-linked glycosylation sites.
  • peptides are selected based on their glycosylation using a lectin based selection technique.
  • Lectins are plant proteins that bind to carbohydrate groups with differing specificity.
  • glycopeptides are initially retained on a lectin column and peptides with substantially no glycosylation are eluted to produce a set of peptides.
  • the glycopeptides are then eluted with a suitable displacing agent to produce another set of peptides.
  • lectins are commercially available, some of which are listed in
  • Suitable displacing agents include, for example, methyl- ⁇ -D-mannopyranoside for a concanavalin A column, L- fucose for AAA or LTA (Tetragonolobus purpureas) column, D-mannose for VFA column, lactose for ECorA column, sialic acid for MMA column, and N- acetylglucosamine for a BSI-A4 or BSI-B4 column.
  • Phosphorylation is a posttranslational modification found in roughly a third of all mammalian cellular proteins.
  • Phosphorylated serine accounts for about 90% of these modifications, phosphorylated threonine for about 10% and phosphorylated tyrosine for about 0.1%.
  • peptides are selected based on their phosphorylation using an IMAC selection technique.
  • IMAC columns for capturing phosphopeptides include, but are not limited to, columns carrying nitrilotriacetate and/or iminodiacetate complexed with Fe 3+ , Ga 3+ , Al 3+ , or Zr 3+ .
  • an IMAC column carrying an iminodiacetate complexed with Ga 3+ is preferred.
  • IMAC columns carrying nitrilotriacetate complexed with Fe 3+ or Ga 3+ are preferred because such columns facilitate the direct ionization of phosphopeptides.
  • peptide cleavage products are loaded onto the TMAC column under acidic conditions (pH 2.0-3.5), unbound nonphosphopeptides are selected by removal from the column with an acidic wash solution to produce one set of peptides, and phosphopeptides are selected by elution under alkaline conditions ( ⁇ pH 10) to produce another set of peptides.
  • Ga complexed IMAC columns are preferable for selection of peptides with multiple phosphate groups, and Fe 3+ complexed IMAC columns are preferable for selection of monophosphorylated peptides because multiphosphorylated peptides are strongly bound.
  • peptides are selected based on their phosphorylation using a technique that uses site-specific chemical modification of the phosphoamino acids followed by introduction of an affinity ligand.
  • peptides are subjected to strongly alkaline conditions to chemically modify the phosphoamino acids such that phosphate groups on serine residues undergo ⁇ -elimination to form a dehydroalanyl residue, and those on threonine resdiues yield dehudroamini-2-butyric acid. Both of these elimination products react readily with ethanedithiol.
  • the remaining free thiol group of ethanedithiol can then, for example, be linked to a biotin affinity tag for affinity-based selection on an avidin column.
  • peptides are selected based on their phosphorylation using a phosphopeptide biotinylation procedure that introduces an isotope-coded affinity tag.
  • the ethanedithiol reagent is with labeled with either four hydrogen or four deuterium atoms.
  • the free sulfhydryl of the ethanedithiol addition product is then tagged with iodoacetyl-PEO-biotin for affinity based-selection of peptides.
  • Selection techniques can be combined in a wide variety of ways for selecting sets of peptides in accordance with the present invention.
  • selection protocol refers to a series of selection techniques.
  • the plurality of peptides (the original analyte) is subjected to a series of selection techniques where the sets of peptides output from a selection technique serve as the input analyte for the next technique in the series.
  • one or more analytes of a set of peptides can be labeled (e.g., with an affinity tag, isotope, etc.) or derivatized prior to their use as an input analyte for a subsequent selection technique.
  • the selecting step selects peptides on the basis of the presence or absence of two or more of the following: (1) specific amino acids; (2) the average number of specific amino acids; (3) the ratio of specific types of amino acids (such as, e.g., the ratio of the number of acidic to basic amino acids); and (4) posttranslational modifications (such as, e.g., glycosylation and phosphorylation).
  • the selecting step further includes a RPC technique.
  • the flow diagrams 400, 420 of Figures 4 A and 4B illustrate various embodiments of a step of selecting sets of peptides from a plurality of peptides 401, 421.
  • Figures 4 A and 4B illustrate a series of three selection techniques, it is to be understood that a series of four or more selection techniques can be used to select peptides to produce sets of peptides in accordance with the present invention.
  • the plurality of peptides (the original analyte) 401, 421 is subjected to a series of selection techniques where the sets of peptides of one technique serve as the input analyte for the next technique in the series.
  • One or more peptides of a set of peptides can be labeled (e.g., with an affinity tag, isotope, etc.) or derivatized prior to their use as an input analyte for a subsequent selection technique.
  • a set of peptides of one selection technique can serve as the input analyte for another selection technique.
  • the first set of peptides 403 serves as the input analyte for the second selection technique, M b , 405a, which produces a third set of peptides 406 and a fourth set of peptides 407 from the first set of peptides 403.
  • the third set of peptides 406 and fourth set of peptides 407 that serve as input analytes to a third selection technique, M c , 410a, 410b.
  • a plurality of peptides 401 (the original analyte) is subjected to a series of selection techniques until at least two or more of the sets of peptides meet the analysis sample criteria of an embodiment of the present invention for a set of peptides.
  • the plurality of peptides is subjected to a first selection technique, M a , 402 to produce a first set of peptides 403 and a second set of peptides 404.
  • the first set of peptides 403 comprises components of the original analyte that are not substantially present in the second set of peptides 404.
  • the first selection technique 402 comprises a strong anion exchange LC technique such that the flow-through comprises peptides with less than three aspartic acids (“Asp” or “D") and/or glutamic acids ("Glu” or “E") and the initially retained peptides contain three or more D and/or E.
  • the first set of peptides 403 comprises flow-through and will not substantially contain peptides with three or more D and/or E
  • the second set of peptides 404 comprises retained analyte and will not substantially contain peptides with less than three D and/or E.
  • the first set of peptides may nevertheless contain some peptides with three or more D and/or E, and the second set of peptides some peptides with less than three D and/or E, because of, for example, well known thermodynamic phenomena that create a distribution of input analyte components between mobile and stationary phases.
  • the first and second sets of peptides 403, 404 are each subjected to a second selection technique, M b , 405a, 405b to produce a third set of peptides 406 and a fourth set of peptides 407 from the first set of peptides 403, and to produce a fifth set of peptides 408 and a sixth set of peptides 409 from the second set of peptides 404.
  • the second selection technique, M b , 405a, 405b comprises a histidine LC selection technique that separates peptides on the basis of whether they contain one or more histidines ("His"), hi one embodiment, peptides containing one or more His are initially retained by a LC column of the second selection technique 405a, 405b and peptides with substantially no His are eluted to produce the third set of peptides 406 from the first set of peptides 403, and the fifth set of peptides 408 from the second set of peptides 404.
  • His histidine LC selection technique
  • the column is then washed to elute peptides containing one or more His to produce the fourth set of peptides 407 from the first set of peptides 403, and the sixth set of peptides 409 from the second set of peptides 404.
  • the first and second sets of peptides can be run through the same LC column, or, more preferably, separate LC columns can be used for the first and second sets of peptides to perform the second selection technique. The use of separate columns, for example, allows for processing the first and second sets of peptides substantially in parallel.
  • a set of peptides is then sequenced to produce a data set for each sequenced set.
  • a set of peptides is subjected to a RPC technique prior to sequencing with, for example, a mass spectrometry instrument.
  • sets of peptides meeting the analysis sample criteria of an embodiment of the present invention for a set of peptides are subjected to additional selection techniques before sequencing.
  • one or more of the third, fourth, fifth and sixth sets of peptides are also subjected to a third selection technique, M c , 410a, 410b, 410c, 41 Od to produce, respectively, seventh and eighth sets of peptides 411, 412, ninth and tenth sets of peptides 413, 414, eleventh and twelfth sets of peptides 415, 416, and thirteenth and fourteenth sets of peptides 417, 418.
  • the third selection technique, M c , 410a, 410b, 410c, 41 Od comprises a cysteine LC selection technique that separates peptides on the basis of whether they contain one or more Cys.
  • all the sets of peptides of the seventh to fourteenth sets of peptides that meet the analysis sample criteria of an embodiment of the present invention for a set of peptides are then sequenced to produce a data set for each sequenced set of peptides.
  • the third, fourth, fifth and sixth sets of peptides can also be run through the same LC column, or, more preferably, separate LC columns can be used for the different sets of peptides to perform the third selection technique. The use of separate columns, for example, allows for processing the third, fourth, fifth and sixth sets of peptides substantially in parallel.
  • a plurality of peptides 421 (the original analyte) is subjected to a series of selection techniques to produce two or more of the sets of peptides that meet the analysis sample criteria of an embodiment of the present invention for a set of peptides.
  • the plurality of peptides is subjected to a first selection technique 423 to produce a first set of peptides 424 and a second set of peptides 426.
  • the first set of peptides 424 comprises components of the original analyte that are not substantially present in the second set of peptides 426.
  • the first selection technique 423 is an LC technique employing a strong anion exchange column such that the flow-through comprises peptides with less than three D and/or E and the initially retained peptides contain three or more D and/or E.
  • the first set of peptides 424 comprises flow- through and will not substantially contain peptides with three or more D and/or E
  • the second set of peptides 426 comprises retained analyte and will not substantially contain peptides with less than three D and/or E.
  • the first set of peptides may nevertheless contain some peptides with three or more D and/or E, and the second set of peptides some peptides with less than three D and/or E, because of, for example, well known thermodynamic phenomena that create a distribution of input analyte components between mobile and stationary phases.
  • the first set of peptides 424 and the second sets of peptides 426 are each subjected to a second selection technique 427a, 427b to produce a third set of peptides 428 and a fourth set of peptides 430 from the first set of peptides 424, and to produce a fifth set of peptides 432 and a sixth set of peptides 434 from the second set of peptides 426.
  • the second selection technique 427a, 427b comprises a LC histidine selection technique that separates peptides on the basis of whether they contain one or more histidines ("His").
  • peptides containing one or more H are initially retained by a LC column of the second selection technique 427a, 427b and peptides with substantially n His are eluted to produce the third set of peptides 428 from the first set of peptides 424, and the fifth set of peptides 432 from the second set of peptides 426.
  • the column is then washed to elute peptides containing one or more His to produce the fourth set of peptides 430 from the first set of peptides 424, and the sixth set of peptides 434 from the second set of peptides 426.
  • the first and second sets of peptides can be run through the same selection device (e.g., GC column, LC column, etc.) or, more preferably, separate selection devices (e.g., separate GC columns, LC columns, etc.) can be used for the first and second sets of peptides to perform the second selection technique.
  • separate selection devices e.g., separate GC columns, LC columns, etc.
  • the use of separate columns, for example, allows for processing the first and second sets of peptides substantially in parallel.
  • the fourth set of peptides 430 is then subjected to a third selection technique 435 to produce a seventh set of peptides 436 and an eighth set of peptides 438 from the fourth set of peptides 430.
  • the third selection technique 435 comprises a LC cysteine selection technique that separates peptides on the basis of whether they contain one or more Cys.
  • two or more the third 428, fifth 432, sixth 434, seventh 436 and eighth 438 sets of peptides that meet the analysis sample criteria of an embodiment of the present invention for a set of peptides are then sequenced to produce a data set for each sequenced set of peptides.
  • Peptide Sequencing A wide variety of methods and instruments are suitable for sequencing peptides in accordance with the present invention. Preferred methods and instruments include, but are not limited to, sequencing by mullti-dimensional mass spectrometry with an ESI- MS or MALDI-MS instrument.
  • Suitable mass spectrometers for an ESI-MS or MALDI- MS instrument include, for example, time-of-flight, quadrupole, RF multipole, magnetic sector, and electrostatic mass spectrometers.
  • the step of sequencing includes subjecting a set of peptides to reverse phase chromatography prior to analysis by mass spectrometry.
  • the step of sequencing comprises loading a set of peptides into an RPC column to produce a series of fractions from the set of peptides and analyzing the fractions as they elute with a mass spectrometry instrument to produce one ore more data sets for the set of peptides.
  • the step of sequencing comprises ionizing at least a portion of the peptides in a set of peptides with the ionization source of a mass spectrometer and selecting ionized peptides that have a mass-to-charge ratio (the parent ions) in a first mass-to-charge ratio range with a first dimension of mass spectrometry, fragmenting the parent ions to produce daughter ions, and generating with a second dimension of mass spectrometry a data set comprising a mass spectrum of the daughter ions.
  • Suitable techniques for fragmenting ions include, but are not limited to: collision induced dissociation ("CID”) with an inert gas collision partner (such as, e.g., helium, argon and xenon); thermally induced dissociation ("TID”), radiation induced dissociation, electron capture dissociation (“ECD”), multiphoton dissociation “MPD”), photon induced dissociation (“PID”); and surface induced dissociation ("SID”) with a chemically modified surface.
  • the parent ions are fragmented in a collision cell to a series of fragment ions, among which are a ladder of ions with sequentially decreasing numbers of amino acids.
  • the fragmentation can occur anywhere along the peptide, a spectrum of mass-to-charge ratios is generated. Typically, two prominent sets of ions are observed in the fragmentation spectrum. One set is a sequence ladder with amino acid deletions from the C-terminal end of the peptide (often referred to as the y series), while the other set is a sequence ladder with amino acid deletions from the N-terminal end (often referred to as the b series). Complete or partial amino acid sequence information for the parent ions is then obtained by interpretation of the fragmentation spectra. As the different amino acids within a peptide each have different masses, the fragmentation spectrum of a peptide is usually characteristic of the peptide sequence.
  • the data sets include the masses and sequences of peptides which are used for database searches to identify proteins, hi one embodiment, determining the likelihood of the presence of two or more proteins in the mixture of proteins comprises identifying the proteins present in the mixture of proteins. In another embodiment, determining the likelihood of the presence of two or more proteins comprises identifying potential candidate proteins present in the mixture of proteins and ranking the potential candidate proteins. The ranking of proteins can be based, for example, on the number and extent to which a protein candidate meets the matching criteria of the database search.
  • determining the likelihood of the presence of two or more proteins comprises assigning a probability that the two or more proteins are present in the mixture of proteins.
  • the step of determimng the likelihood of the presence of two or more proteins in a mixture preferably includes selecting two or more data sets such that in combination the sequence information of the selected data sets is representative of greater than about 80% of the proteins in the mixture of proteins, hi another embodiment, data sets are selected such that in combination the sequence information of the selected data sets is representative of greater than about 90% of the proteins in the mixture of proteins.
  • a data set is representative of about x% of the proteins in a protein mixture when about x% of the proteins in the protein mixture are known or predicted to have the same combination of posttranslational modification and/or amino acid presences and absences as the selected peptides in the set of peptides from which the data set was produced.
  • a combination of data sets is representative of about x% of the proteins in a protein mixture when about x% of the proteins in the protein mixture are known or predicted to have the same combination of posttranslational modification and/or amino acid presences and absences as the selected peptides in the combined sets of peptides from which the combination of data sets was produced.
  • protein identification, protein candidate determination, and assignment of a probability that a protein is present are based on the comparison of peptide sequence information from one or more data sets with: (1) peptide sequences generated by an in silico analysis of genomic data (such as, e.g., a DNA or EST database); (2) peptide sequences generated by an in silico digestion of proteins in a protein database; and/or (3) experimentally determined peptide sequences.
  • genomic data such as, e.g., a DNA or EST database
  • peptide sequences generated by an in silico digestion of proteins in a protein database and/or (3) experimentally determined peptide sequences.
  • One advantage of searching an EST database is that EST data comes from mRNA. For example, it is typically possible to identify the protein parent from the sequence information from a single peptide because almost all peptides obtained from a tryptic digest are unique to their parent protein.
  • peptides that are unique to their parent are often referred to signature peptides because they may be used as a signature for their protein parent.
  • a target protein can be identified with as little as a two or three amino acid sequence from a single peptide when this sequence information is combined with mass data and the specificity of the digest to generate, for example, a peptide sequence tag.
  • Specific sequence tags can facilitate the identification proteins from, for example, incomplete databases, error-prone databases, and facilitate the analysis of novel proteins. It is also possible that a peptide appears in more than one protein. When a peptide appears in more than one protein, these proteins are almost always related and such peptides are often referred to as familial peptides.
  • the fact that two proteins carry the same peptide is generally of biological significance and therefore important.
  • the relationship between related proteins can be, for example, that they have a common biological function, that they are splice variants carrying the same exon, or that they are mutation based variants of the same protein.
  • identification of the parent protein from which the peptide is derived typically depends on finding another peptide in the digest that is unique to the protein in question.
  • the step of determining the likelihood of the presence of two or more proteins in a mixture of proteins includes comparing at least two or more of the data sets of the sequencing step to at least one of a deoxyribonucleic acid ("DNA”) database and a protein database, h one embodiment, the data sets are compared to peptide sequences derived from a DNA library. For example, the DNA sequences from either a DNA or Expressed Sequence Tag (“EST”) database are converted into a protein and subsequently tryptic peptide sequences in silico and then compared with the peptide sequence information obtained from one or more data sets.
  • DNA deoxyribonucleic acid
  • protein database e.g., a protein database
  • the data sets are compared to peptide sequences derived from a protein database, hi one version, tryptic peptide sequences are generated in silico and then compared with the peptide sequence information obtained from one or more data sets. In another version, the data sets are compared to experimentally determined tryptic peptide sequences of the proteins in the protein database.
  • the present invention provides increased throughput for proteomic studies of mixtures of proteins
  • the methods of the present invention increase the throughput of peptide sequencing by increasing the efficiency of the sequencing component of the analysis.
  • the methods increase the efficiency of mass spectral analysis by producing sets of peptides such that the number of peptides does not exceed the peptide number capacity of the sequencing method.
  • the method selects a combination of sets of peptides where the combination of sets of peptides is representative of greater than about 80% of the proteins in the mixture of proteins, and sequences the combination of sets of peptides in less than about 10 hours.
  • the method selects a combination of sets of peptides where the combination of sets of peptides is representative of greater than about 90% of the proteins in the mixture of proteins, and sequences the combination of sets of peptides in less than about 4 hours.
  • the method sequences all the sets of peptides in the combination of sets of peptides, such that data sets are produced for all the sets of peptides in the combination in a time that does not exceed four times the time required for an RPC separation of one of the sets of peptides in the combination.
  • each set of peptides selected for the combination of sets of peptides is individually representative of less than about 75% of the proteins in the mixture of proteins.
  • each set of peptides selected for the combination of sets of peptides is individually representative of between about 30% to about 70% of the proteins in the mixture of proteins.
  • the method selects two or more sets of peptides wherein the number of peptides in the two or more sets of peptides is such that the sequencing method is capable of at least about 50% of the peptides in the two sets of peptides in less than about 180 minutes.
  • the method selects two or more sets of peptides wherein the number of peptides in the two or more sets of peptides is such that a MALDI-MS sequencing method is capable of sequencing at least about 50% of the peptides in the two sets of peptides in less than about 90 minutes.
  • the method selects two or more sets of peptides wherein the number of peptides in the two or more sets of peptides is such that an ESI-MS sequencing method is capable of sequencing at least about 50% of the peptides in the two sets of peptides in less than about 180 minutes, i yet another embodiment, the method selects two or more sets of peptides wherein the number of peptides in the two or more sets of peptides is such that a RPC-MS sequencing method is capable of sequencing at least about 50% of the peptides in the two sets of peptides in less than about 180 minutes.
  • the functionality of the methods described above may be implemented as computer-readable instructions on a general purpose computer.
  • the computer maybe separate from, detachable from, or integrated into a chromatography- mass spectrometry system.
  • the computer-readable instructions may be written in any one of a number of high-level languages, such as, for example, FORTRAN, PASCAL, C, C++, or BASIC. Further, the computer-readable instructions may be written in a script, macro, or functionality embedded in commercially available software, such as EXCEL or VISUAL BASIC. Additionally, the computer-readable instructions could be implemented in an assembly language directed to a microprocessor resident on a computer. For example, the computer-readable instructions could be implemented in Intel 80x86 assembly language if it were configured to run on an IBM PC or PC clone.
  • the computer-readable instructions be embedded on an article of manufacture including, but not limited to, a computer-readable program medium such as, for example, a floppy disk, a hard disk, an optical disk, a magnetic tape, a PROM, an EPROM, or CD-ROM.
  • a computer-readable program medium such as, for example, a floppy disk, a hard disk, an optical disk, a magnetic tape, a PROM, an EPROM, or CD-ROM.
  • EXAMPLE 1 Peptide Selection of Tryptic Digest of E. Coli Proteins
  • Example 1 provides an example of an analysis of a mixture of E. Coli proteins.
  • the mixture of E. Coli proteins was predicted by an in silico analysis of the E. Coli genome.
  • An in silco analysis of the E. Coli genome F. R. Blattner, G. Plunkett, 3rd, C. A. Bloch, N. T. Perna, V. Burland, M. Riley, J. Collado-Vides, J. D. Glasner, C. K. Rode, G. F. Mayhew, J. Gregor, N. W. Davis, H. A. Kirkpatrick, M. A. Goeden, D. J. Rose, B.
  • Figures 5A-5G provide further information on the average number of peptides and proteins containing certain amino acids or combinations of certain amino acids as a function of protein molecular weight based on an in silico analysis of the E. Coli genome.
  • Figures 5A-5C provide information on the average number of proteins containing certain amino acids as a function of protein molecular weight, where the x-axis in
  • Figures 5A-5C is in kiloDaltons (kDa).
  • Figure 5A is a plot 501 of the average number of Cys contaimng proteins 502 (filled square symbols) as a function of protein molecular weight
  • Figure 5B is a plot 503 of the average number of His containing proteins 504 (filled triangle symbols) as a function of protein molecular weight
  • Figure 5C is a plot 505 of the average number of His containing proteins 506 (filled diamond symbols) as a function of protein molecular weight.
  • Figures 5D-5G provide information on the average number of peptides containing certain amino acids or combinations of certain amino acids as a function of protein molecular weight, where the x-axis in Figures 5D-5G is also in kiloDaltons (kDa).
  • Figure 5D is a plot 507 that compares the average number of Cys containing tryptic peptides 508 (filled square symbols), the average number of tryptic peptides containing both Cys and one or more of Asp or Glu 509 (filled triangle symbols), and the average number of tryptic peptides containing both Cys and three or more of Asp or Glu 510 (filled diamond symbols) as a function of protein molecular weight.
  • Figure 5E is a plot 511 that compares the average number of His containing tryptic peptides 512 (filled triangle symbols), the average number of Cys containing tryptic peptides 513 (filled square symbols), and the average number of tryptic peptides containing both Cys and His 514 (filled circle symbols) as a function of protein molecular weight.
  • Figure 5F is a plot 515 that compares the average number of His containing tryptic peptides 516 (filled triangle symbols), the average number of tryptic peptides containing both His and one or more of Asp or Glu 517 (filled circle symbols), and the average number of tryptic peptides containing both His and three or more of Asp or Glu 518 (filled diamond symbols) as a function of protein molecular weight.
  • Figure 5G is a plot 519 that compares the average number of tryptic peptides containing one or more of Asp or Glu 520 (filled triangle symbols), the average number of Cys contaimng tryptic peptides 521 (filled square symbols), and the average number of tryptic peptides containing both Cys and one or more of Asp or Glu 522 (filled diamond symbols) as a function of protein molecular weight.
  • An example of operating selectivity columns in tandem to select sets of peptides in accordance with embodiments of the present invention from the plurality of peptides generated by the above in silico tryptic digest of E. Coli proteins is illustrated by the schematic diagrams 600, 601 in Figures 6 A and 6B.
  • cysteine- containing peptides are considered to be derivatized with biotin and thus selectable with an avidin affinity column.
  • the number of peptides selected by various combinations of avidin selection of cysteine peptides, histidine selection with Cu(II)-IMAC, and acidic amino acid selection with a strong anion exchange ("SAX") column is based on the above in silico analysis of the E. Coli genome.
  • selection techniques are run in tandem to form sets of peptides some of which serve as samples Si- S 6 for subsequent submittal to a RPC-MS system for sequencing by a RPC-MS sequencing method.
  • Absolute numbers of peptides are used in this example to give an idea of the relative numbers of peptides selected and to provide, for example, examples of N sa and the percentage of proteins in the protein mixture represented by the various sets of peptides. Although absolute numbers of peptides selected will vary between organisms, relative selection and percent coverage of the proteome will generally be independent of the organism.
  • the 4,289 proteins predicted from the E. Coli genome (the protein mixture 602 in this example) are subjected to an in silico tryptic digest to generate a plurality of peptides (the original analyte) 603 containing 130,000 peptides.
  • the plurality of peptides 603 is first subjected to a LC selection technique 604 to select peptides based on their acidity by employing a strong anion exchange ("SAX") column such that the flow-through 605 comprises peptides with less than three D and/or ⁇ and the initially retained peptides contain three or more D and/or ⁇ .
  • SAX strong anion exchange
  • the flow-through from the SAX column comprises the first set of peptides 605 and the initially retained peptides are eluted to produce the second set of peptides 606.
  • the first 605 and second 606 sets of peptides are each then subjected to a second
  • LC selection technique 607a, 607b to select peptides based on the presence or absence of His by employing a Cu(II)-IMAC column.
  • the flow-through 609 of these Cu(II)- MAC column comprises peptides that do not contain His, whereas the retained peptides 608 each contain one or more His.
  • the peptides in the first set of peptides that are retained by the first Cu(II)-IMAC column 607a are eluted to produce a fourth set of peptides 610 having 16,040 peptides which together are representative of 3,915 proteins in the protein mixture.
  • the flow-through of the first Cu(II)-IMAC column 607a produces a fifth set of peptides 611 that is subjected to further LC selection 612a to select peptides based on the presence or absence of Cys by employing an avidin affinity column.
  • the peptides retained by this avidin affinity column 612a are then eluted to produce a sixth set of peptides 613 having 6,101 peptides which together are representative of 2,864 proteins in the protein mixture.
  • at least a portion of the peptides of the fourth set of peptides 610 form the first sample St and at least a portion of the peptides of the sixth set of peptides 613 form the second sample S .
  • Cu( ⁇ )IMAC column 607b are eluted to produce a seventh set of peptides 614 having 7,677 peptides which together are representative of 3,135 proteins in the protein mixture.
  • the flow-through of the second Cu(II)IMAC column 607b produces an eighth set of peptides 615 having 11,537 peptides which together are representative of 3,537 proteins in the protein mixture.
  • at least a portion of the peptides of the seventh set of peptides 614 form the third sample S 3 and at least a portion of the peptides of the eighth set of peptides 615 form a fourth sample S .
  • a portion of the eluate that forms the eighth set of peptides 615 is subjected to further LC selection 612b to select peptides from the eighth set of peptides 615 based on the presence or absence of Cys using an avidin affinity column.
  • the peptides retained 616 by this avidin affinity column 612b are eluted to produce a ninth set of peptides 617 having 2,115 peptides which together are representative of 1,527 proteins in the protein mixture.
  • the flow-through 618 of this avidin affinity column 612b produces a tenth set of peptides 619 having 9,421 peptides which together are representative of 3,266 proteins in the protein mixture.
  • At least a portion of the peptides of the ninth set of peptides 617 then fo ⁇ n the fifth sample S 5 and at least a portion of the peptides of the tenth set of peptides 615 form the sixth sample S 6 .
  • FIGs 6A and 6B in addition to the specific selection of peptides containing certain amino acids, there is value in removing classes of peptides. For example, using the SAX column peptides with multiple acidic amino acids can typically be removed from samples in a few minutes. As illustrated, samples Si and S come from the flow-through from the SAX column.
  • Table 2 provides for each sample of Example 1, and certain combinations of these samples, a summary of: the number of peptides in the sample or combination of samples; the percentage of the plurality of peptides in the sample or combination of samples; the number of proteins in the protein mixture represented by the sample or combination of samples; and the percentage of the proteins in the protein mixture represented by in the sample or combination of samples.
  • Table 2 illustrates that a variety of combinations of sets of peptides and a variety of combinations of their corresponding data sets can be selected that are representative of greater than about 80% and greater than about 90% of the proteins in the protein mixture 602. It is seen that the combination of samples S 3 and S 4 (i.e., the seventh and eight sets of peptides) is representative of about 92% of the proteins in the mixture of proteins (e.g., the selection techniques that produce samples S 3 and S 4 result in sets of peptides that together contain signature peptides from 92% of the E. Coli proteome of this example).
  • Table 2 illustrates that the peptide separation approach of the present invention can provide a greater than 90% reduction in original analyte (the plurality of peptides 603) complexity.
  • sample S 2 retained only 4.7% of the original 130,000 peptides from the tryptic digest whereas samples S 3 , S 4 and S 6 retained 5.9%, 8.9%, and 7.3%, respectively.
  • Complexity can be reduced even further, for example, a set of peptides selected by a serial grouping of SAX, avidin, and Cu(II)IMAC columns, to select peptides with three or more Asp or Glu, one or more His, and one or more Cys, would contain only 1% of the peptides in the original analyte.
  • Figures 7 A, 8 A and 9 A provide schematic representations 700, 800, 900 of the series of chromatographic selection techniques used to select peptides in Examples 2, 3 and 4, respectively.
  • the chromatographic selection techniques used in Examples 2-4 are the same.
  • the original analyte of the example is first subjected to a LC selection technique 704 to select peptides based on their acidity by employing a SAX column such that the flow-through comprises peptides with less than three D and or E and the initially retained peptides contain three or more D and/or E.
  • the initially retained peptides and the flow-through of the first LC selection technique 704 are then subjected a second LC selection technique 706a, 706b to select peptides based on the presence or absence of His by employing a Cu( ⁇ )-IMAC column.
  • Both the initially retained peptides and the flow-through of the second selection techniques 706a, 706b are then subjected to a third LC selection technique 708a, 708b, 708c, 708d to select peptides based on the presence or absence of Cys by employing an avidin affinity column.
  • Examples 2-4 all cysteine-containing peptides are considered to be labeled with a biotin affinity group derivatized with iodoacetate acid and thus selectable with avidin affinity columns.
  • the third LC selection techniques produce eight sets of peptides in each of Examples 2-4 that each form one of samples S ⁇ -S 8 for submittal to an RPC-MS sequencing method. Although the contents of these eight sets of peptides can vary between Examples 2-4, they share the same group of peptide selection criteria across these examples. For example, the peptides of sample Si share the selection criteria of peptides with less than three Asp or Glu, but with one or more His, and with one or more Cys. Table 3 summarizes the peptide selection criteria met by samples S ⁇ -S 8 in Examples 2-4. TABLE 3
  • Figure 7B, 8B, 8C, and 9B-J Figure provide plots of molecular weight in atomic mass units (amu) versus predicted retention time (% ACN) on a RPC column for the plurality of peptides of Examples 2-4, or individual samples S ⁇ -S 8 .
  • the retention times were predicted based on the retention coefficients of Chabanet, C, Yvon, M., Journal of Chromatography, vol. 599, pp. 211-225 (1992). Although the predicted retention times do not exactly match experimentally determined retention times, the correlation between experimental and predicted retention times is greater than 0.9. In addition, in Examples 2-4, the predicted retention times do not consider the contribution of modified Cys to peptide retention.
  • EXAMPLE 2 Peptide Selection of Tryptic Digest of HSA
  • Example 2 provides an example of an analysis of a plurality of peptides 702 (the original analyte) generated by an in silico tryptic digest of human serum albumin ("HSA") 701.
  • HSA human serum albumin
  • the in silico tryptic digest of HSA yielded 45 peptides.
  • Table 4 provides for each sample S ⁇ -S 8 of Example 2 a summary of: the number of peptides in the sample or combination of samples; the percentage of the plurality of peptides in the sample or combination of samples; the number of proteins in the protein mixture represented by the sample or combination of samples; and the percentage of the proteins in the protein mixture represented by in the sample or combination of samples.
  • Figure 7B is a plot 750 of molecular weight in atomic mass units (amu) versus predicted retention time (% ACN) on a RPC column for an input analyte comprising all 45 peptides of the plurality of peptides 702 loaded at once.
  • the various retention times for peptides of various molecular weights are represented by the filled circles 752.
  • the amino acid sequence for HSA (Sequence I.D. No. 1) 754 is given in the upper right corner of the plot 750.
  • EXAMPLE 3 Peptide Selection of Tryptic Digest of HSA and PZP
  • Example 3 provides an example of an analysis of a plurality of peptides 802 generated by an in silico tryptic digest of a mixture of human proteins consisting of HSA and pregnancy zone protein ("PZP") 801. The mixture of proteins was assumed to contain less than 0.5 mg% of PZP and 3500-5500 mg% of HSA. The in silico tryptic digest of the HSA and PZP mixture yielded 111 peptides. The eight sets of peptides produced 821, 822, 823, 824, 825, 826, 827, 828, serve as samples Si- S 8 for subsequent submittal to a sequencing method.
  • Table 5 provides for each sample Si-Ss of Example 3 a summary of: the number of peptides in the sample or combination of samples; the percentage of the plurality of peptides in the sample or combination of samples; the number of proteins in the protein mixture represented by the sample or combination of samples; and the percentage of the proteins in the protein mixture represented by in the sample or combination of samples.
  • Figure 8B is a plot 830 of molecular weight in atomic mass units (amu) versus predicted retention time (% ACN) on a RPC column for an input analyte comprising all 111 peptides of the plurality of peptides 802 loaded at once.
  • the retention times for the peptides of PZP are represented by the filled circles 832, and those for the peptides of HSA by unfilled circles 834.
  • Figure 8C is a plot 850 of molecular weight in atomic mass units (amu) versus predicted retention time (% ACN) on a RPC column for an input analyte comprising sample S 3 of Example 3 823.
  • amino acid sequences for the peptides of sample S 3 are given in the upper right corner of the plot for both sequences of the peptides of PZP 852 (Sequence I.D. Nos. 2-9) and those for the peptides of HSA 854 (Sequence ID. Nos. 10-15).
  • EXAMPLE 4 Peptide Selection of Tryptic Digest of Protein Mixture
  • Example 4 provides an example of an analysis of a plurality of peptides 902 generated by an in silico tryptic digest of a mixture of human proteins consisting of five abundant human proteins and PZP 901.
  • the five abundant proteins were HSA, alpha- 1 - antitrypsin, transfettin, fibrinogen, and immunoglobulin-G.
  • the mixture of proteins was assumed to contain equal quantities of PZP and of the five abundant proteins.
  • the in silico tryptic digest of this mixture of proteins yielded 277 peptides.
  • Table 6 provides for each sample S ⁇ -S 8 of Example 4 a summary of: the number of peptides in the sample or combination of samples; the percentage of the plurality of peptides in the sample or combination of samples; the number of proteins in the protein mixture represented by the sample or combination of samples; and the percentage of the proteins in the protein mixture represented by in the sample or combination of samples.
  • Figure 9B is a plot 930 of molecular weight in atomic mass units (amu) versus predicted retention time (% ACN) on a RPC column for an input analyte comprising all 277 peptides of the plurality of peptides 902 loaded at once.
  • the retention times for the peptides of PZP are represented by the filled circles 932, and those for the peptides of the other proteins by unfilled circles 934.
  • Figures 9C-9 J provide predicted RPC retention time plots for each of samples
  • FIGS 9C-9J The retention times for the peptides of PZP are represented by the filled circles 932, and those for the peptides of the other proteins by unfilled circles 934.
  • Figure 9C is a plot 940 of predicted retention time for an input analyte comprising sample Si 921 and provides the amino acid sequences of the peptides of PZP 942 (Sequence ID. Nos. 16-20) in sample Si and those for the other proteins 944 (Sequence I.D. Nos. 21-34) in sample Si.
  • Figure 9D is a plot 945 of predicted retention time for an input analyte comprising sample S 2 922 and provides the amino acid sequences of the peptides of PZP 946 (Sequence ID. Nos. 35-43) in sample S 2 and those for the other proteins 948 (Sequence ID. Nos. 44-67) in sample S 2 .
  • Figure 9E is a plot 950 of predicted retention time for an input analyte comprising sample S 3 923 and provides the amino acid sequences of the peptides of PZP 952 (Sequence ID. Nos. 2-9) in sample S 3 and those for the other proteins 954 (Sequence ID. Nos. 10-15 and 68-86) in sample S 3 .
  • Figure 9F is a plot 955 of predicted retention time for an input analyte comprising sample S 4 924 and provides the amino acid sequences of the peptides of PZP 956 (Sequence ID. Nos. 87-122) in sample S and those for the other proteins 954 (Sequence ID. Nos. 123-220) in sample S 4 .
  • Figure 9G is a plot 960 of predicted retention time for an input analyte comprising sample S 5 925 and provides the amino acid sequences of the peptides of PZP 962 (Sequence ID. Nos. 221-222) in sample S 5 and those for the other proteins 964 (Sequence I.D. Nos. 223-230) in sample S 5 .
  • Figure 9H is a plot 965 of predicted retention time for an input analyte comprising sample S 6 926 and provides the amino acid sequences of the peptides of PZP 966 (Sequence ID. Nos. 231) in sample S 6 and those for the other proteins 968 (Sequence ID. Nos. 232-241) in sample S 6 .
  • Figure 91 is a plot 970 of predicted retention time for an input analyte comprising sample S 7 927 and provides the amino acid sequences of the peptides of PZP 972 (Sequence I.D. Nos. 242-243) in sample S 7 and those for the other proteins 974 (Sequence I.D. Nos. 244-258) in sample S 7 .
  • Figure 9J is a plot 975 of predicted retention time for an input analyte comprising sample S 8 928 and provides the amino acid sequences of the peptides of PZP 976 (Sequence I.D. Nos. 259-262) in sample S 8 and those for the other proteins 978 (Sequence ID. Nos. 263-277) in sample S 8 .
  • the claims should not be read as limited to the described order or elements unless stated to that effect. While the invention has been particularly shown and described with reference to specific illustrative embodiments, it should be understood that various changes in form and detail may be made without departing from the spirit and scope of the invention as defined by the appended claims. By way of example, any of the disclosed features may be combined with any of the other disclosed features to analyze a mixture of proteins in accordance with the invention. Therefore, all embodiments that come within the scope and spirit of the following claims and equivalents thereto are claimed as the invention.

Landscapes

  • Life Sciences & Earth Sciences (AREA)
  • Health & Medical Sciences (AREA)
  • Engineering & Computer Science (AREA)
  • Molecular Biology (AREA)
  • Physics & Mathematics (AREA)
  • Bioinformatics & Cheminformatics (AREA)
  • Bioinformatics & Computational Biology (AREA)
  • Hematology (AREA)
  • Chemical & Material Sciences (AREA)
  • Biomedical Technology (AREA)
  • Urology & Nephrology (AREA)
  • Immunology (AREA)
  • Biophysics (AREA)
  • Spectroscopy & Molecular Physics (AREA)
  • Analytical Chemistry (AREA)
  • Microbiology (AREA)
  • Biotechnology (AREA)
  • Proteomics, Peptides & Aminoacids (AREA)
  • Food Science & Technology (AREA)
  • Medicinal Chemistry (AREA)
  • Cell Biology (AREA)
  • Biochemistry (AREA)
  • General Health & Medical Sciences (AREA)
  • General Physics & Mathematics (AREA)
  • Pathology (AREA)
  • Optics & Photonics (AREA)
  • Investigating Or Analysing Biological Materials (AREA)
  • Other Investigation Or Analysis Of Materials By Electrical Means (AREA)

Abstract

La présente invention se rapporte à des techniques de sélection de peptides, qui facilitent le criblage et la caractérisation à haut débit de mélanges de protéines. Dans divers aspects, l'invention concerne des procédés permettant d'analyser un mélange de protéines. Dans divers modes de réalisation, les procédés selon l'invention permettent : de générer une pluralité de peptides à partir d'un mélange de protéines ; de sélectionner des jeux de peptides parmi la pluralité de peptides ; d'ordonner au moins une partie des peptides sélectionnés pour produire des jeux de données contenant des informations de séquences ; et de déterminer la probabilité que deux protéines ou plus soient présentes dans le mélange de protéines. Dans un mode de réalisation, la technique de sélection de peptides consiste à sélectionner des peptides de façon qu'au moins deux jeux de peptides soient chacun séparément non représentatifs du mélange de protéines. Dans un autre mode de réalisation, des peptides sont sélectionnés de façon qu'au moins deux jeux de peptides soient chacun séparément représentatifs de moins de 70 % environ mais de plus de 30 % environ des protéines contenues dans le mélange de protéines.
PCT/US2004/023908 2003-07-25 2004-07-23 Chromatographie numerique Ceased WO2005011576A2 (fr)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
US49014903P 2003-07-25 2003-07-25
US60/490,149 2003-07-25

Publications (2)

Publication Number Publication Date
WO2005011576A2 true WO2005011576A2 (fr) 2005-02-10
WO2005011576A3 WO2005011576A3 (fr) 2005-05-06

Family

ID=34115364

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/US2004/023908 Ceased WO2005011576A2 (fr) 2003-07-25 2004-07-23 Chromatographie numerique

Country Status (1)

Country Link
WO (1) WO2005011576A2 (fr)

Cited By (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20150000232A1 (en) * 2013-06-28 2015-01-01 Illinois Tool Works Inc. Three-phase portable airborne component extractor with rotational direction control
CN118221801A (zh) * 2024-05-22 2024-06-21 山东省食品药品检验研究院 一组人纤维蛋白原特征多肽及其应用

Family Cites Families (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US5538897A (en) * 1994-03-14 1996-07-23 University Of Washington Use of mass spectrometry fragmentation patterns of peptides to identify amino acid sequences in databases

Cited By (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20150000232A1 (en) * 2013-06-28 2015-01-01 Illinois Tool Works Inc. Three-phase portable airborne component extractor with rotational direction control
CN118221801A (zh) * 2024-05-22 2024-06-21 山东省食品药品检验研究院 一组人纤维蛋白原特征多肽及其应用

Also Published As

Publication number Publication date
WO2005011576A3 (fr) 2005-05-06

Similar Documents

Publication Publication Date Title
James Protein identification in the post-genome era: the rapid rise of proteomics
Tholey et al. Top-down proteomics for the analysis of proteolytic events-Methods, applications and perspectives
US8030089B2 (en) Method of analyzing differential expression of proteins in proteomes by mass spectrometry
Quadroni et al. Proteomics and automation
Chen et al. Application of LC/MS to proteomics studies: current status and future prospects
Swanson et al. The continuing evolution of shotgun proteomics
US20150160234A1 (en) Rapid and Quantitative Proteome Analysis and Related Methods
Mo et al. Analytical aspects of mass spectrometry and proteomics
van Schaick et al. Computer-aided gradient optimization of hydrophilic interaction liquid chromatographic separations of intact proteins and protein glycoforms
Regnier et al. Multidimensional chromatography and the signature peptide approach to proteomics
Qian et al. High-throughput proteomics using Fourier transform ion cyclotron resonance mass spectrometry
Chiou et al. Clinical proteomics: current status, challenges, and future perspectives
Paša‐Tolić et al. Gene expression profiling using advanced mass spectrometric approaches
Wittke et al. Differential polypeptide display: the search for the elusive target
Mesmin et al. Complexity reduction of clinical samples for routine mass spectrometric analysis
Schrattenholz Proteomics: how to control highly dynamic patterns of millions of molecules and interpret changes correctly?
Patterson Protein identification and characterization by mass spectrometry
EP1469314B1 (fr) Méthode de spectrométrie de masse
Patrie et al. Top-down mass spectrometry for protein molecular diagnostics, structure analysis, and biomarker discovery
Leinenbach et al. Proteome analysis of sorangium cellulosum employing 2D-HPLC-MS/MS and improved database searching strategies for CID and ETD fragment spectra
EP1393061A1 (fr) Procedes d'analyse de proteines en phase multiples
Wu et al. A computational approach for the identification of site-specific protein glycosylations through ion-trap mass spectrometry
Meyers et al. Protein identification and profiling with mass spectrometry.
CA2616888C (fr) Procede de spectrometrie de masse
Zybailov et al. Mass spectrometry-based methods of proteome analysis

Legal Events

Date Code Title Description
AK Designated states

Kind code of ref document: A2

Designated state(s): AE AG AL AM AT AU AZ BA BB BG BR BW BY BZ CA CH CN CO CR CU CZ DE DK DM DZ EC EE EG ES FI GB GD GE GH GM HR HU ID IL IN IS JP KE KG KP KR KZ LC LK LR LS LT LU LV MA MD MG MK MN MW MX MZ NA NI NO NZ OM PG PH PL PT RO RU SC SD SE SG SK SL SY TJ TM TN TR TT TZ UA UG US UZ VC VN YU ZA ZM ZW

AL Designated countries for regional patents

Kind code of ref document: A2

Designated state(s): BW GH GM KE LS MW MZ NA SD SL SZ TZ UG ZM ZW AM AZ BY KG KZ MD RU TJ TM AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HU IE IT LU MC NL PL PT RO SE SI SK TR BF BJ CF CG CI CM GA GN GQ GW ML MR NE SN TD TG

121 Ep: the epo has been informed by wipo that ep was designated in this application
122 Ep: pct application non-entry in european phase