WO2000060507A2 - Analyzing molecule and protein diversity - Google Patents

Analyzing molecule and protein diversity Download PDF

Info

Publication number
WO2000060507A2
WO2000060507A2 PCT/US2000/008777 US0008777W WO0060507A2 WO 2000060507 A2 WO2000060507 A2 WO 2000060507A2 US 0008777 W US0008777 W US 0008777W WO 0060507 A2 WO0060507 A2 WO 0060507A2
Authority
WO
WIPO (PCT)
Prior art keywords
molecules
ofthe
theoretical
diversity
target surfaces
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/US2000/008777
Other languages
French (fr)
Other versions
WO2000060507A3 (en
Inventor
Edward A. Wintner
Ciamac C. Moallemi
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Neogenesis Inc
Original Assignee
Neogenesis Inc
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Neogenesis Inc filed Critical Neogenesis Inc
Priority to JP2000609930A priority Critical patent/JP2002541560A/en
Priority to EP00925889A priority patent/EP1203330A2/en
Priority to CA002369570A priority patent/CA2369570A1/en
Priority to AU44511/00A priority patent/AU4451100A/en
Publication of WO2000060507A2 publication Critical patent/WO2000060507A2/en
Publication of WO2000060507A3 publication Critical patent/WO2000060507A3/en
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Classifications

    • GPHYSICS
    • G16INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
    • G16CCOMPUTATIONAL CHEMISTRY; CHEMOINFORMATICS; COMPUTATIONAL MATERIALS SCIENCE
    • G16C20/00Chemoinformatics, i.e. ICT specially adapted for the handling of physicochemical or structural data of chemical particles, elements, compounds or mixtures
    • G16C20/50Molecular design, e.g. of drugs
    • GPHYSICS
    • G16INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
    • G16BBIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
    • G16B15/00ICT specially adapted for analysing two-dimensional [2D] or three-dimensional [3D] molecular structures, e.g. structural or functional relations or structure alignment
    • GPHYSICS
    • G16INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
    • G16CCOMPUTATIONAL CHEMISTRY; CHEMOINFORMATICS; COMPUTATIONAL MATERIALS SCIENCE
    • G16C20/00Chemoinformatics, i.e. ICT specially adapted for the handling of physicochemical or structural data of chemical particles, elements, compounds or mixtures
    • G16C20/40Searching chemical structures or physicochemical data

Definitions

  • This invention relates to analyzing molecule and protein diversity.
  • Combinatorial chemistry allows the creation of unprecedented numbers of organic compounds.
  • the rational synthesis of millions of small organic molecules is now achievable in a matter of days.
  • small we mean a molecule having fewer than 1500 Daltons, where a Dalton is defined as 1/12 of the weight of a carbon 12 atom or roughly the weight of a hydrogen atom.
  • the utility of small molecules as drugs depends in part on molecular complementarity: how well the molecules fit and/or stick to chemically active sites (often in the form of depressions in a protein) on the surface of a cell, on the surface of an intracellular organelle, or on a cytosolic protein.
  • the potential molecular complementarity of a small molecule is in large part determined by two factors: 1) The shape of the molecule, meaning the total Van Der Waals (VDW) surface of a given conformation of the molecule and how it follows or does not follow the VDW surface of the target site of interest.
  • VDW Van Der Waals
  • the shape complementarity of molecule and target are largely responsible for energetic forces such as the displacement of water (the "hydrophobic effect") and so called “Van Der Waals” or “London dispersion” forces. 2)
  • the types of potential energetic interactions (such as hydrogen bonding, percentage ionic bonding, proximity of polarizable moieties) found at various places on the molecule, and on the manner, order, and spatial orientation in which the energetically interactive portions ofthe molecule are connected to each other and presented in the presence of the target of interest.
  • each of these factors is measured or calculated in only a relative way with respect to a specific set of molecules or with respect to specific protein surfaces being examined.
  • a general comparison of potential molecular complementarity between two sets of molecules requires doing calculations or experiments using an absolute or fixed frame of reference. For example, if company 1 has tested a molecule set A for affinity to a set of target surfaces X found in cancerous cells, and company 2 has tested a molecule set B against a set of target surfaces Y found in nervous tissue, the value of molecule set B with respect to the target surfaces X of company 1 is not apparent because each of the affinity evaluations has been performed against a different standard. Also, without a full molecular calculation of set A against target surfaces Y, it is not apparent whether potential chemically active portions of the target surface Y could be bonded by molecules in set A. Finally, even if all calculations of sets A and B vs.
  • Implementations ofthe invention provide an absolute or fixed frame of reference for comparing sets of molecules and protein surfaces.
  • molecules in terms of their complementarity or attraction to a fully enumerated basis set of theoretical protein surfaces within given parameters, it is possible to measure the diversity of a set of molecules against a non relative standard, i.e., an absolute measurement.
  • This allows for an efficient comparison of different sets of molecules, for example, sets of drugs, and enables meaningful categorization of classes of molecules against a standard set of surfaces.
  • this allows for the detection of theoretical protein surfaces to which no molecules in a set are complementary, thus enhancing the ability of a chemist to design novel molecules that supplement the deficiencies ofthe original set.
  • real world protein surfaces By defining real world protein surfaces in terms of their similarity to a basis set of theoretical protein surfaces, it is similarly possible to categorize real world sets of protein surfaces against a standard set of theoretical surfaces. This allows for improved classification of proteins into similar target classes by the similarity of their surface sites.
  • the invention features a computer-based method in which a set of constraints on possible target surfaces is defined, and a fully enumerated set of theoretical target surfaces under the defined constraints is also defined, such that each surface has a defined, continuous volume and a defined, continuous surface area.
  • One or more sets of objects are mapped to the fully enumerated set of theoretical target surfaces to define corresponding subsets of the fully enumerated set of theoretical target surfaces.
  • An aspect of diversity ofthe objects is analyzed based on degrees of similarities and differences among the corresponding subsets.
  • Implementations ofthe invention may include one or more ofthe following features.
  • the target surfaces may include negative space target surfaces.
  • the objects may include positive space object surfaces associated with different molecules.
  • the objects may be mapped by defining corresponding subsets ofthe fully enumerated set of negative space theoretical target surfaces to which positive space object surfaces of conformations of molecules are complementary.
  • the aspect of diversity that is analyzed may be the difference or similarity between the molecules which map to those negative space theoretical target surfaces.
  • the objects may include negative space object surfaces associated with different proteins, and the objects may be mapped by defining corresponding subsets ofthe fully enumerated set of negative space theoretical target surfaces to which negative space object surfaces of protein pockets are similar.
  • the aspect of diversity that is analyzed may be the difference or similarity between protein pockets which map to those negative space theoretical target surfaces.
  • the objects may include positive space object surfaces associated with different molecules and negative space object surfaces associated with different proteins. In the case of molecules, the objects may be mapped by defining corresponding subsets ofthe fully enumerated set of negative space theoretical target surfaces to which positive space object surfaces of conformations of molecules are complementary.
  • the objects may be mapped by defining corresponding subsets of the fully enumerated set of negative space theoretical target surfaces to which negative space object surfaces of protein pockets are similar.
  • the aspect of diversity that is analyzed may be the difference or similarity of the molecules which map to those negative space theoretical target surfaces to the protein pockets which map to those negative space theoretical target surfaces.
  • the theoretical target surfaces and the objects may be polyhedrons, e.g., cubes, all ofthe same size and shape.
  • the set of all theoretical target surfaces defines a diversity space within which the diversity of objects can be measured by mapping those objects to the diversity space. Regions of the diversity space to which no objects map may be identified, and molecules may be designed that occupy at least one ofthe unfilled theoretical target surfaces ofthe diversity space. Complementarity may be associated with binding affinities of positive space object surfaces of conformations of molecules to negative space theoretical target surfaces.
  • the constraints may include volume, associations of each of a number of sites of the target surface with a preselected molecular property drawn from a larger set of possible molecular properties, including hydrophobic, polarizable, H-bond acceptor, H-bond donor, H-bond donor/acceptor, potentially positively charged, and potentially negatively charged. Fewer than all of the sites ofthe target surface may each be associated with a different one ofthe molecular properties and all of the other sites of the target surface may be associated with a common molecular property, such as slightly hydrophobic.
  • the degrees of similarities or differences may involve functional properties associated with the corresponding subsets of the fully enumerated set of theoretical target surfaces or shape properties associated with the corresponding subsets ofthe fully enumerated set of theoretical target surfaces.
  • Each of the objects may be defined by quantizing molecules into polyhedrons. Each of a fixed set of orientations of each conformation of each of the objects may be fitted to each of the target surfaces, and each ofthe fittings may be scored.
  • the constraints may include a resolution ofthe polyhedrons, e.g., 4.24 Angstroms, or maximum and minimum numbers of polyhedrons. Each of the polyhedrons may share a common interface with another of the polyhedrons. The constraints may also include the absence of any occluded volumes greater than a given user-defined parameter.
  • the target surfaces may be defined conceptually as having been carved out of a flat surface.
  • the invention features categorizing existing molecules based on negative space target surfaces to which conformations ofthe molecules are complementary, and designing novel molecules that are complementary to negative space target surfaces to which no conformations of the existing molecular are complementary.
  • the invention features a method of creating novel molecules to be tested as ligands for proteins.
  • proteins are categorized based on target surfaces to which their pockets of known structure map, and novel molecules are designed that are complementary to the negative space target surfaces to which the protein pockets map.
  • the invention features a computer programmed to determine the chemical similarity of different molecules.
  • the program approximates the surface shape of each one of a plurality of molecules of interest by linking a series of cubes, each cube having a dimension R, the locations of the cubes being determined by the calculated electron probability density of the individual one of the molecules of interest, each cube sharing at least one of its six faces with another cube, such that there is a specific number of linked cubes which varies for each individual one of the plurality of molecules of interest.
  • the chemical reactivity of each individual one of the plurality of molecules of interest is approximated by assigning each cube of each individual one of the plurality of molecules of interest, no more than one functionality value from a plurality of M different chemical functionality values.
  • V is approximated by subtracting a number V R ⁇ cubes of dimension R from a surface, wherein each of the cube spaces shares at least one face with another cube space and wherein N of the cube spaces has one of a plurality of M different chemical functionality values.
  • An attraction value K is calculated for each one of the plurality of molecules of interest to the chemically active surface.
  • a list of overall attraction values to the chemically active surface is calculated.
  • Implementations ofthe invention may include one or more ofthe following features.
  • the calculation of the attraction value K may be performed on a plurality of different predetermined chemically active surfaces, and a matrix of overall attractive values of each molecule of interest to each of the different surfaces may be calculated.
  • the molecules of interest may include organic molecules.
  • the chemically active surface having a plurality of predetermined active chemical locations may be calculated to correspond to the shape of an actual protein surface structure.
  • the molecules of interest may be organic molecules of 1500 Daltons or less.
  • the chemically active surface having a plurality of predetermined active chemical locations may be compared to an actual protein surface to calculate a similarity value ofthe actual protein surface to the predetermined active chemical locations.
  • the predetermined chemically active surfaces may be compared to a plurality of actual protein surfaces and a matrix of similarity values may be calculated.
  • the cube spaces subtracted from the surface may be calculated to approximate the electron probability density of at least one of a plurality of depressions in known protein surface structures.
  • the N sites of chemical functionality may be calculated to approximate the location and type of chemical functionality of actual depressions in known protein structures.
  • Figure 1 shows a (CH 2 ) n chain encapsulated by 4.24 A cubic units.
  • Figure 2 shows examples of surfaces allowed and disallowed by the non- occlusion parameter in a theoretical target surface generation algorithm. Gray shading represents the opening ofthe theoretical surface. A is allowed. B is disallowed due to two occluded negative space cubes (marked X).
  • Figure 3 shows a theoretical target surface of 13 negative space cubes and four sites of specific molecular property interaction: hydrophobic (white), polarizable (purple), H-bond accepting (green), and H-bond donating (orange). Blue shading indicates the opening of the theoretical surface.
  • Figure 4 shows a "quantized” representation (Q-file) of one conformation of molecule 6a superimposed on its atomic structure (ball and stick and spacefilling model). Molecular property characteristics ofthe Q-file are hydrophobic (white quanta), polarizable (purple quanta), H-bond accepting (green quanta), and negatively charged (red quanta).
  • Figure 5 shows test molecules.
  • Figure 6 shows a ranking of molecules by QCSD similarity scores.
  • Figure 7 illustrates examination of a theoretical target surface common to molecules 8c (top) and 8a (bottom). Blue shading indicates opening ofthe theoretical surface. Specific points of complementarity on the theoretical target surface are hydrophobic (white) polarizable (purple) and H-bond donating (orange). Superimposition ofthe original molecular conformations onto the theoretical target surface demonstrates that the extra phenyl substituent of 8c protrudes from the opening ofthe theoretical surface and is not involved in complementarity to the surface.
  • Figure 8 illustrates examination of a theoretical target surface common to molecules la (A) and 5a (B). Blue shading indicates opening of the theoretical surface.
  • Figure 9 illustrates ranking of molecules in Figure 5 by Tanimoto similarity score of 2D UNITY finge ⁇ rints.
  • Figure 10 shows a QSCD plot of all ofthe theoretical surface shapes covered by all ofthe conformations of all of the molecules (blue dots) in Figure 5. The total volume of the cube encompasses all 49,268,918 theoretical surface shapes as listed in Table 1. Red dots show two exemplary theoretical surface shapes (a, b) not covered by any ofthe molecules in Fig. 5. Axes used are functions of opening area, opening length/width, and depth per opening quantum.
  • Figure 11 shows a map ofthe 20 compounds in Fig. 5 (blue dots) in a representative BCUT three-axis diversity space.
  • BCUT axes used are, respectively: 1) BCUT HACCEPT S LNVDIST 050 R H, 2) BCUT HDONOR S LNVDIST 030 R H, and 3) BCUT TAB POLAR S LNVDIST 300 R L. Red dot shows an unfilled coordinate of diversity space, at (7.54, 7.25, 6.82). The information contained in this BCUT coordinate does not reveal information about the shape of a molecule which might be able to fill this position in diversity space.
  • Figure 12 shows use of QSCD to design complementary combinatorial libraries to unmatched theoretical target surfaces. Many conceivable libraries of a given shape and functionality may be designed to fill a given unmet diversity need.
  • Figure 13 shows two sample surfaces.
  • Figure 14 illustrates the quantization process.
  • Figure 15 shows a legend for symbols used in functionality rule diagrams.
  • Figure 16 shows functionality rules for Potential Negative Charge
  • Figure 17 shows functionality rules for Potential Positive Charge Functionality. The structures are searched for in order.
  • Figure 18 shows a functionality rule for Hydrogen Bond Donor/ Acceptor Functionality.
  • Figure 19 shows a functionality rule for Hydrogen Bond Donor Functionality.
  • Figure 20 shows a functionality rule for Hydrogen Bond Acceptor Functionality. The structures are searched for in order.
  • Figure 21 shows a functionality rule for Polarizable Functionality.
  • Figure 25 shows an example Core Molecule Mi to fill Central set Ci.
  • Figure 28 shows an Example Core Molecule Mi to fill Central set Ci.
  • Figure 30 shows a protein quantization process.
  • molecular diversity can be defined as the measure, based on biological criteria, of the difference or similarity between small molecules.
  • biological criteria of the difference or similarity between small molecules.
  • Each ofthe existing methods of calculating biologically relevant diversity of small molecules defines slightly different criteria for molecular comparison, and thus a different configuration of diversity space as a whole. Examples include low dimensional diversity space such as BCUT metrics, high dimensional diversity space such as Chem-X/ChemDiverse multiple point pharmacophores, and empirical biological diversity space such as affinity finge ⁇ rinting.
  • affinity finge ⁇ rinting Another example of current diversity methods is affinity finge ⁇ rinting, in which molecules are empirically assayed against a panel of 10-20 actual proteins selected to be promiscuous in their ability to bind small molecules. Position in molecular diversity space is assigned through the resulting string of IC 50 binding values, and these affinity finge ⁇ rints provide unprecedented ability to group similarly active compounds in diversity space. However, because the actual mode of binding in any assay is not inco ⁇ orated in the resulting IC 5 0 value, the mapping of molecules to the selected protein panel is an irreversible transformation. Thus, an empty coordinate in affinity finge ⁇ rinting diversity space (an "unmatched" string of IC50S to a given protein panel) cannot be back- translated into a 3D molecular template.
  • a similar affinity finge ⁇ rinting diversity method has been put into practice using a panel of computational surfaces of real- world protein pockets and a modified form ofthe DOCK program. While this method shows similar promise in its ability to detect pharmacological similarity, it is, like its empirical affinity finge ⁇ rinting counte ⁇ art, an irreversible mapping. Thus, for the most part, current methods are able to successfully identify compounds of the same pharmacological class as being similar and compounds of different pharmacological classes as being different. Given a starting pharmacophore from known ligands and/or the target site of a target crystal structure, such methods interface well with the design of complementary combinatorial libraries.
  • QSCD which defines a molecule numerically by a mapping that describes its complementarity K to every distinct theoretical protein surface of resolution R not exceeding volume V with N sites of M types of chemical functionality
  • P MN - K is defined as an algorithm that takes into account the molecular shape and chemical functionality of both the given molecule and the given theoretical protein surface. From this definition, it follows that a comparison of two molecules will yield a numerical difference that is representative of their complementarities: to the extent that two molecules each have complementarity for the same theoretical protein surfaces, the molecules are similar; to the extent that two molecules have no complementarity to common theoretical protein surfaces, the molecules are dissimilar.
  • each theoretical surface to be formed by successively carving c ubic units out of an initially flat surface. These cubic units represent "negative space" that a potential ligand could occupy. Given cubic units with sides of length R (the resolution ofthe model), we use at most V/R 3 negative space cubes to describe each theoretical target surface. Others have previously employed cubic units to successfully approximate complementarity between small molecules and individual protein surfaces. The size of a negative space cube is directly related to the resolution and type of diversity data which the user desires as output.
  • the size ofthe negative space cube In choosing the size ofthe negative space cube, one motivation is to maximize negative space cube size such that the difference of a single cube in a surface is highly differentiating in terms of molecular recognition (i.e., every surface is orthogonal to every other surface). At the same time, enough information must be retained in each negative space cube to predict shape and functional complementarity at a ligand/surface interface. The former constraint minimizes overlap of diversity information while the latter constraint maximizes precision of diversity information. Together, the competing constraints result in a basis unit for the enumeration of theoretical target surfaces that minimizes the number of negative space cubes needed to accurately model diversity for a given volume V.
  • 4.24 A is the approximate VDW "cross-section" of a (CH2)n chain; a series of 4.24 A units neatly encapsulates a (CH2)n chain in its ground state conformation as shown in Figure 1.
  • the basis set for diversity can be a set of theoretical target surfaces comprised of all possible shape combinations of 6 to 14 negative space cubes of resolution 4.24 A (negative volume between 460 and 1070 cubic A) subject to the following rules: Surfaces are created by successively "carving out" negative space cubes from a flat block of infinite width and depth (the theoretical target). All negative space cubes of a given surface must share at least one face with another negative space cube of the surface, and all must be part of a single, contiguous negative surface. No negative space cubes may be occluded in the +Z axis ofthe infinite surface block; that is, there may be no solid surface between any negative space cube and the surface plane ofthe infinite block.
  • the surface A is allowed, but the surface B is disallowed. Surfaces duplicating a previous surface with respect to rotation in the X-Y plane are discarded.
  • the occlusion rule provides a compromise between complete coverage of topological possibilities and acceptable computational speed. This compromise was made based on the topological assumption that occlusions of 4.24 A or more are infrequent in small molecule/target interactions, and that their omission would thus have only a small effect on predicting diversity of binding affinities of small molecules. Applying the rules yields 49,268,918 unique negative surface shapes including chiral opposites. Covering a negative volume between 460 and 1070 cubic A, these surface shapes are deemed sufficient to examine diversity of most small molecules. For instance, examining a previously published reference set of pharmaceutically relevant compounds (a filtered Comprehensive Medicinal Chemistry or CMC database), 5049 out of 5120 compounds (98.6%) have a volume of 1070 cubic A or less.
  • each negative space cube is assigned a molecular property characteristic Pm that represents the dominant molecular environment which any atoms that are placed within that negative space will experience. Properties used are PI hydrophobic, P2 polarizable (includes aromatics), P3 H-bond acceptor, P4 H-bond donor, P5 H-bond donor/acceptor, P6 potentially positively charged (basic), and P7 potentially negatively charged (acidic). These seven types of molecular environments are assumed to represent a minimal basis set of factors that contributes to the electrostatic/VDW complementarity of a ligand and a target surface.
  • Table 1 Numerical breakdown ofthe total number of theoretical target surfaces created using the algorithm given in the text. Surfaces consist of 6-14 negative space cubes and 4 sites of 7 possible molecular property characteristics. Number of functionally different surfaces per surface shape varies for infrequent cases in which a given shape has an axis of symmetry, so actual number of unique surfaces is slightly less than (# surface shapes) * 7 4* N!/((N-4)! * 4!).
  • the small molecules must be formatted in a similar frame of reference, for instance by quantizing them into positive space cubes ("quanta") of resolution 4.24 A according to the following process (illustrated in Fig. 4):
  • a set of up to 100 minimized energy conformations within user-defined parameters is created.
  • Tripos Multisearch modeling is used, and all conformations within 10 kcal of the lowest energy conformation found are accepted.
  • a 4.24 A 3D grid of cubes (quanta) is aligned on top of the 3D structure using the molecule's principle axes of rotation (calculated with all atoms having mass 1).
  • Order of dominance is from P 7 to Pi, in order of maximum complementarity score obtainable by a given characteristic as shown in Table 2:
  • Table 2 Relative magnitudes of parameters used in calculating molecular property interactions between negative space cubes (theoretical target surfaces) and positive space cubes (quantized molecules). Magnitudes (listed from highest to lowest): +++, ++, +, 0, -, --, — .
  • Minimum % of VDW radius parameter allows for a user-defined protrusion beyond the surface of a quantum cube, adding a measure of topological "flexibility" to the quantization process. A user defined 32% was found to be especially good.
  • the total number of 4.24 A quanta that have been assigned a property characteristic is counted.
  • the grid alignment is shifted per user-defined parameters and the process is repeated until all shift combinations have been searched.
  • a "Q-file” (3D configuration of property-assigned quanta) is saved that has the lowest number of quanta in and is closest to the principle alignment.
  • a typical Q-file (molecule 6a) is shown in Fig. 4, superimposed upon its co ⁇ esponding conformation. The process of optimization of quantization parameters is described later.
  • each quantized conformation can be mapped into the diversity space defined by the set of 1.1 x 10 14 theoretical target surfaces. In general, the following process is used. For each quantized conformation of each molecule, each of its 24 possible X/Y/Z rotations (6 faces * 4 rotations per face) is fit to each ofthe 49,268,918 available surface shapes.
  • a score is generated for the complementarity ofthe given conformation to each theoretical target surface of a given shape from based on user-defined parameters. (The process of optimization ofthe complementarity parameters is described later.)
  • the following complementarity parameters can be used: a) A negative parameter for each rotatable bond of the conformation. b) If conformational energies are calculated, a negative parameter for the energy ofthe conformation above the lowest energy conformation from that molecule. c) A positive parameter for the hydrophobic energy gained by removing "water” from any hydrophobic (Pi) or polarizable (P 2 ) surface face of either the conformation or the theoretical surface.
  • the implementations ofthe invention resolve the problem to a framework bounded by 24 possible fitting orientations and a finite number of translations. This approximation allows three-dimensional diversity computation on a scale that is applicable to very large sets of molecules.
  • QSCD makes many approximations of molecular recognition. As explained, these include cubic units of 4.24 A resolution, gross approximations of surface contact area, exactly 4 points of 7 finite types of molecular property characteristics, static theoretical surfaces, and a limited set (up to 100) of low energy conformers. Thus, the final complementarity scores are not presumed to give precise binding energies for any individual match of conformation to target surface. However, taken over all conformations of a molecule and across an enumerated set of theoretical target surfaces, the scoring system is statistically relevant as explained below. Model Validation
  • Table 4 Tabulation of surface shapes and total number of theoretical target surfaces complementary to each molecule in Fig. 5.
  • mappings were scored in similarity from 0 to 1000 based on a function of the number of theoretical surfaces in common:
  • the first term in this equation gives a percentage measure (0-100) of shape similarity between molecules A and B, while the second term gives a measure from 0-10 of functional similarity per given shape overlap.
  • the scoring constant ⁇ in the equation above adjusts the influence of functionality on scoring.
  • Fig. 6 shows a plot of all 190 pairings ranked by similarity score. Circles show "heterogeneous" pairs of expected dissimilarity (e.g. 2a, 6b), while squares show “homogeneous” pairs of expected similarity (e.g. 2a, 2b). Clearly, the QCSD model ranks homogeneous pairs almost exclusively higher than heterogeneous pairs; all 1 pharmacologically similar pairs fell within the top 20 scores out of 190. All homogeneous scores were ranked above 25, while the median score in this experiment was 2.8, showing good "signal to noise.” The QCSD model is thus a valid predictor of target binding similarity among these molecules.
  • the pairings also reveal further validation. As might be expected from their relative rigidity (low number of accessible conformations) and structural similarity, the highest scoring pairs are 2a/2b, 8a/8b, and 8d/8e. Furthermore, examination ofthe pairings of 8c with 8a,b,d,e (triangles in Fig. 6, yellow in Table 5) yields scores that are within the top 20% of the pairing experiment but which are generally lower that the "homogeneous" pairs. This makes sense from a target-binding point of view, considering that one face of 8c contains a large molecular difference (an extra phenyl substituent).
  • Fig. 7 shows one such case of a surface common to both 8a and 8c; the protruding phenyl substituent plays no role in complementarity.
  • Fig. 7 shows one such case of a surface common to both 8a and 8c; the protruding phenyl substituent plays no role in complementarity.
  • FIG. 8 depicts one such case between la and 5a; conformations of la and 5a are displayed that were found in the QCSD model to be complementary to the same surface (Fig. 8A, 8B).
  • 3D overlays (Fig. 8C) confirm correlation of general shape and 4 points of functionality, although they also make clear the limits of resolution of complementarity information using 4.24 A units.
  • the surface in question can detect general shape and functional similarity, but by does not provide a basis to predict atom-for-atom overlap between molecules.
  • Fig. 9 and Table 6, contained in Figure 23, show the same set of 20 molecules ranked by Tanimoto similarity of standard 2D UNITY fingerprints (see discussion below).
  • the data demonstrate that the 2D model is equally capable of predicting pharmacologically similar pairs; UNITY ranks similarity between ATI and AT2 subtype binders much higher than our QCSD model, although it finds unusually high similarity between 8a and 8c.
  • 2D fingerprint descriptors have been found effective in clustering pharmacologically similar compounds, and are widely used in determining molecular diversity of existing structures.
  • the QCSD model determines not only diversity of existing structures, but also the structure of non-existing diversity. Given theoretical surface shapes for which no complements exist in a general screening library, QSCD allows the design of molecules to fill the given diversity void.
  • the QSCD basis set is created through a reversible process. Although some information resolution may be lost in fixing the parameters of a cube's size and functional scope, information content is retained in either direction.
  • a single molecular conformation and orientation corresponds to a defined pattern in QSCD space
  • a single point in QSCD space corresponds to a unique 3D shape with a defined 3D array of functionality.
  • unoccupied points in QSCD space directly define the molecular shapes and functionalities which those molecules do not cover.
  • a set of detailed 3D molecular templates is immediately available for the creation of novel molecules.
  • Fig. 10 shows a plot of all of the theoretical surface shapes covered by all ofthe conformations of all ofthe molecules used in the example implementation (see Fig. 5).
  • the total volume ofthe cube in Fig. 10 encompasses all 49,268,918 theoretical surface shapes as listed in Table 1.
  • many theoretical surface shapes are "unfilled” by the set of compounds shown in Fig. 5.
  • searching for molecules or libraries to enhance the diversity ofthe given set of compounds the chemist is presented with a set of actual 3D templates into which new compound libraries may be designed.
  • mapping the same set of compounds in a "non-reversible" diversity space would also display a set of coordinates to which the molecules map, there would be no way to visualize the 3D shape of any point that was not filled by one of the compounds in the set.
  • the coordinates specified for an unfilled point leave the chemist with a set of normalized eigenvalues. While these may give an idea of relative abundance of a given functionality (e.g. H-Bond Donor) at this point in diversity space, the coordinates give no hint of what shape or class of molecules might fill that diversity void.
  • the above example shows how QSCD is a reversible diversity model with respect to molecular shape.
  • QSCD makes possible the contemplation of a "complete" library of screening molecules at a given resolution.
  • the model thus offers a theoretical and practical answer to the problem of generating lead structures for genomic targets of unknown structure and function.
  • UNITY 2D fingerprints (Unity 4.0, Tripos Ine, 1699 S. Hanley Rd., St. Louis, MO, 63144) were generated on an R10000 Silicon Graphics workstation. Pairwise Tanimoto coefficients were computed as described by Dixon and Koehler.
  • QSCD software for molecule quantization, mapping of Q-files, and surface complementarity display was developed using the Java programming language (JDK 1.2) and the Java3D graphics API (version 1.1) on Intel-based workstations.
  • Theoretical target surfaces were stored and indexed using an Oracle 7.3.3 database.
  • Parameters for theoretical target surface generation/molecular quantization and parameters for complementarity mapping/scoring were alternately optimized in three successive rounds as described below.
  • the parameters used for theoretical target surface generation and the closely related parameters for quantization of small molecules into quantized files (Q-files) were optimized in the context ofthe algorithms mentioned above.
  • Parameters were iteratively optimized by varying a given parameter and then quantizing training molecules other than those in Fig. 5. Training molecules used were taken from in house structures and two published SAR sets.
  • mapping/scoring molecular conformations to theoretical target surfaces were optimized in the context of the algorithm stated above. Parameters were iteratively optimized by varying a given parameter and then mapping a constant set of training molecules (see above) to a constant set of theoretical target surfaces, using the most cu ⁇ ent surface generation and quantization parameters. Diversity pairing scores were generated for all training molecules, and parameters were chosen which accurately predicted known homogeneous heterogeneous pairs and which maximized "signal to noise" of homogeneous scores over heterogeneous scores.
  • the minimum overlap requirement was set to either 9 quanta or N-2 quanta of a conformation of N quanta. This range allows large conformations to fit partially into a theoretical surface (protruding volume must be at the mouth of the surface) while also allowing smaller conformations to be considered for complementarity. It excludes large conformations which do not overlap at least 9 quanta.
  • Approximate computational speeds of typical QSCD operations are as follows on a single Pentium III 500 MHz workstation: Generation ofthe basis set of theoretical target surface used in the study required 17 min.; this data was stored for access by subsequent QSCD functions. Quantization of 100 conformations of a given molecule into 100 Q-files required 250 seconds. Complementarity mapping of 100 Q-files onto the basis set of theoretical target surfaces used in the study required 40 seconds. Algorithm for Designing Molecules for Unfilled Target Surfaces
  • Appendix E describes an algorithm for quantization of protein surfaces.
  • Appendix F describes an algorithm for comparing protein surfaces to determine a degree of similarity or dissimilarity. The following algorithm generates a set of files T of quantized protein binding surfaces which together represent the available surface of a given protein binding site.
  • Protein binding site see l, below
  • Transinc translational increment in angstroms
  • Rot # rotations (odd integer)
  • Rotvar rotational variance (%)
  • CN 94025) which minimally consists of: a) A calculated probable electron density surface ofthe binding site b) A list of all known atom types in the molecule with their coordinates and atomic radii c) A list of known connectivities of all atoms with the type of bond connecting each atom
  • step 7 At a given grid nexus place a cube with the center of one face tangent to the probable electron density surface ofthe protein binding site 6. If the cube contains a protein atom coordinate or any atomic radii of protein atoms protrude into the trial cube by more than Tol % of their atomic radius, then step 7 ' ., otherwise step 8.
  • a trial cube contains a protein atom coordinate, or any atomic radii of protein atoms protrude into the trial cube by more than TolA of their atomic radius, or if the cube does not intersect with a volume ofthe convex hull equal to at least R 3 * TolB (see 3. above), then the trial cube is removed. Otherwise the trial cube becomes a set cube (with unchecked faces).
  • step 9 Designate all cubes as negative space cubes fully enclosed except as detailed below. Designate the layer of cubes which is a) pe ⁇ endicular to the line pe ⁇ endicular to the probable electron density surface at the grid nexus being examined b) farthest from the grid nexus as negative space cubes which are open at their faces farthest from the grid nexus and pe ⁇ endicular to the line pe ⁇ endicular to the probable electron density surface at the grid nexus.
  • Types of functionality M may include but are not limited to:
  • Appendix G describes an algorithm for determining the complementarity of a library of molecules to a set of protein surfaces.
  • novel molecules for testing as ligands for proteins
  • novel molecules can be designed based on complementarity to negative space cube targets to which a set of protein pockets map. The following outline describes the steps for doing so:
  • Appendix H contains example parameter values useful in connection with the algorithms described in Appendixes A through G.
  • a 4.24 A cube was found to be the largest predictive unit size of diversity measure for our criteria of designing general screening libraries. For example, both 4.48 and 4.00 A units gave poorer prediction of homogeneous/heterogeneous pairs than the pairings of Fig. 6 (4.24 A units). This is likely due to the fact that most organic small molecules are themselves quantized by a limited basis set: the VDW radii of H, C, N, O and a few other atoms (see for example Fig.l). If there is no constraint on size of cubic units, however (i.e., if there is no attempt to maximize orthogonality of theoretical target surfaces), other unit measures of diversity can be found. A unit of 2.12 A should also provide effective diversity information but at a much higher resolution.
  • size of diversity space in terms of unique molecular points. In other words, what is the minimum set of molecules needed to fully cover a given diversity space. This calculation is dependent on two factors: the resolution stipulated in the model (e.g., what amount of molecular change is recognized as different) and the maximum values of each dimension ofthe model's basis axes. In the model of QSCD discussed above, resolution is fixed by cubic units of 4.24 A, and maximum values are fixed at 14 units (molecular volume of 1070 cubic A) and 4 points of 7 types of molecular property characteristics. As describe above, the result is a set of 1.1 * 10 14 unique molecular points.
  • Table 7 Summation of binding energies for an interaction of an average complementary molecule/theoretical target surface pair in the context ofthe QSCD model used herein.
  • An average complementary theoretical target surface is also assumed to have 60% non-polar exposed faces. Constants used in the table are taken from Ajay and Murcko.
  • the resolution used to calculate diversity translates roughly to nanomolar binding conditions for an average molecule/target surface pair.
  • a general screening library guaranteed to contain at least one nanomolar binder to any given target of interest would thus number at least 24 million molecules. This is a large number and will be attenuated by the fact that some molecules have significantly more than 100 conformations available to them.
  • the QSCD model suggests that if, in the near future, combinatorial chemistry and high-throughput screening are to generate initial hits primarily in the nanomolar rather than micromolar range, then the field must continue to focus its efforts on the development of numerically competent synthesis and screening technologies.
  • a surface opening O is a set of lattice squares in 7? denoted by their corners
  • the area of a surface opening is ed to be the number of lattice squares it contains. Surface openings are considered to be unique upto translations and rotations ofthe x-y plane.
  • a surface shape 5 is a set of "negative space" cubes represented as lattice cubes in Z 3 denoted by their comers:
  • the volume of a surface shape is defined to be the number of lattice cubes it contains. Surface shapes are considered to be unique up to translations and rotations of the x-y plane.
  • shape(0, d) ⁇ (x, y, z) ⁇ (x,y) e 0, d(x, y) > -z)
  • the function d specifies the "depth" ofthe surface shape at each opening point. Sample surface shapes and openings can be seen in Figure 13.
  • a theoretical surface consists of a surface shape where the cubes in the surface shape are each associated with functionality.
  • the set T of seven specific types of characteristic functionality is used:
  • a functionality map / : S — » defines the assignment. By default, all cubes are assigned functionality % , and all possibilties are considered where upto n/ of the cubes are given one ofthe functionalities T ⁇ - .
  • N ⁇ to be a set containing the only the surface openmg with a smgle square at (0, 0)
  • T 5 define T 5 to the set of all possible openings obtamed by addmg a single square to O adjacent to a square already present in O
  • Algorithm A.2 ⁇ PENlNGFl TER(C, ⁇ t , M nc , M c ) filter a set of surface openings O using area-threshold parameter At, max-non-central parameter M nc , and max- contiguous parameter M c .
  • Each atom in the molecule is assigned a functionality based on its type and connectivity
  • Each cube is assigned a functionality based on the atoms that it contains
  • a map / M — T assigning functionalities to all of the atoms
  • the map is defined by the using a set of rules to match molecular substructures based on extended atom types (as generated by T ⁇ pos, for example) and bondmg patems that encapsulate each functionality type
  • the algo ⁇ thm keeps track of atoms that are excluded from matching a lower p ⁇ o ⁇ ty rule because they have already been matched in a higher p ⁇ o ⁇ ty rule, where T ⁇ has the highest p ⁇ onty and T ⁇ the lowest.
  • Atoms not matching any rule are assigned functionality T ⁇ . No atoms are assigned functionality T %
  • the functionality rules used can be seen as follows:
  • Algorithm B.l ATOMFUNCTIONAL AP( ⁇ ) assign functionality to the atoms in molecular structure -Vf and determine which atoms are excluded from quantizafton
  • a conformation is a mapping c ⁇ M. — ⁇ E 3 of the molecular structure into three dimensional space.
  • OpenEye Omega software Open Eye Scientific Software Inc., 335c Winische Way, Santa Fe, NM, 87501
  • upto n c representation conformations for each molecule are generated within given rule-based energy parameters.
  • a coordinate frame T (R. t) I 3 ⁇ R 3 is a ⁇ gid motion ofthe space defined by a rotation R S S0 3 (R) and a translation t 6 R 3 that transforms a point p by the rule
  • a lattice on R 3 is implicitly defined by q £ Z 3 ⁇ [r ⁇ . rg + r) x [r ⁇ y , rg y + r) x [r ⁇ r 2 , rg 2 + r) C R 3
  • the base coordinate frame is generated from a conformation of molecular structure M.
  • a subset . of the atoms in M. are selected via F RAMEATO S These are atoms m or near ⁇ ng structures close to the center of the conformation
  • a ⁇ ng atom is defined to be an atom that contains at least one bond which, if removed, would not result in the molecule being disconnected If there are an insufficient number of ⁇ ng atoms, all atoms sufficently close to the center ofthe conformation are used
  • the base coordinate frame is calculated in B ASEFRAME.
  • the x-axis of the base coordinate frame is defined to to be the solution to the optimization problem max ⁇ x ⁇ (c(a) — p) ⁇ ⁇ M where
  • the base coordinate frame then simply involves cente ⁇ ng the conformation by translating p to the origin, and using the new x, y, and z
  • Algorithm B.2 FRAMEAT0MS(J . c, 7 , r m , q ⁇ select a set of atoms to use for generating the coordmate frame from a molecular structure M. and a conformation c.
  • the following parameters are used, ⁇ ng factor 7 , ring minimum r m , radius factor q ⁇ .
  • Algorithm B.3 RlNGGR ⁇ UP( ⁇ TZ, 77) calculate the set of atoms that can be reached starting at atom a and crossmg over at most 77 atoms that are not in the set of ⁇ ng atoms TZ
  • the lattice defined by a coordinate frame places corners of cubes at points all of whose coordinates are integer multiples of r. Given a particular conformation, it may be better to shift the lattice by a length of ⁇ /2 in a particular direction, recentering the lattice cubes.
  • I ⁇ - ⁇ c f ⁇ M ⁇ ] set ( to be the .th lowest element in the set ⁇
  • j to be the index of the least member of the set ⁇ (n ⁇ , d z , d Vt i , d x .i ⁇ , where comparisons are done using a dictionary ordering (that is, compare the first component, if case of equality compare the second component, etc.)
  • the base coordinate frame is not necessanly optimal for quantization, so a set of "close * ' frames are also exammed.
  • a set of "close * ' frames are also exammed.
  • R x to be the coordmate frame corresponding to rotation about the x-axis by ⁇ v r (2 ⁇ x + 1 — n r )/(4n r ) radians
  • R y to be the coordmate frame corresponding to rotation about the y- axis by ⁇ v r (2 ⁇ y + 1 — n r )/(4n r ) radians
  • R z to be the coordmate frame corresponding to rotation about the --axis by ⁇ v r (2 ⁇ z + 1 — n r )/(4n r ) radians
  • T x (p) p + (rv t (2 Jx + 1 - n r )/(2n P ).0.0)
  • T y (p) p + (Q, rv t (2j y + 1 - n r )/(2n r ), 0)
  • cubification is the process of determining which lattice cubes are filled by the conformation.
  • the set of lattice cubes is constructed by taking any cube in which an atom center in the conformation directly falls and also cubes which are sufficiently close to the van Der Waals sphere of an atom.
  • Algorithm B.7 CUBIFY(- , c, T, r, t): quantize the conformation c of the molecular structure M into a set of cubes defined on the coordinate frame T usmg parameters: coordinate frame T, resolution r, tolerance t.
  • d to be the distance of the point in R 3 m the cube defined by q that is closest to p
  • Algorithm B.8 Q ⁇ Am ⁇ Z ⁇ .(M,C,r,t,rf,r m ,q r ,Cf,C t ,n t , ⁇ t ,n r , ⁇ r ) quantize the set of conformations C for the molecular structure M with parameters resolution r, tolerance t, nng factor 77, ⁇ ng minimum r m , radius factor q r , cente ⁇ ng fraction c/, cente ⁇ ng tolerance c t , number of translations n t , translational vanance v t , number of rotattons n ⁇ , rotational va ⁇ ance ⁇ r , pola ⁇ zable minimum p m .
  • Algorithm Cl FITSURFACES( , /,, E c ,T b , ) calculate all surfaces with functionality that are complementary to the quantizated conformation Q with functionality map /,, conformational energy E c , and rt, rotatable bonds using the followmg parameters minimum surface openmg area A, maximum surface volume V, area- threshold A t , max-non-central M nc , max-contiguous M c , max-extrustion M e , number of points of characte ⁇ stic functionality /, minimum energy £—, deliberately, minimum fit quanta q mtn , minimum slackness s min , maximum slackness s mol , maximum protrusion levels p max , translational-rotational-vibrational entropy E trVs rotatable bond coefficient c r , hydrophobic energy coefficient C h , hydrophobic surface energy coefficient c s , potential function
  • TZ to be the set of 24 lattice rotations
  • Algo ⁇ thm C.3 DETECTSURFACES(S C , A V. A t , M nc , M c , M e ) detect additional surfaces by adding cubes to 5- subject to parameters minimum surface opening area A, maximum surface volume V, area-threshold A t , max-non-central M nc , max- contiguous ⁇ / c , max-extrustion M e
  • V to be the set of all possible openmgs obtained by adding a single square to O adjacent to a square already present m O
  • Complementarity energy between a quantization of a molecular conformation and a a cubic theoretical surface is the sum of several components:
  • a library is a set of molecular structures Given a library, the set of complementary theoretical surfaces is defined as the union of all surface shape/functionality pairs complementary to any quantized conformation of any molecule in the library
  • the algo ⁇ thm LlBRARYCOMPARE calculates a score proportional to the similarity of the two hbra ⁇ es
  • the score is calculated by representmg each library as its set of complementary theoretical surfaces, and usmg the S IMILARITYSCORE p ⁇ mitive to determine the similanty or dissimila ⁇ ty two sets of theoretical surfaces If the molecular bra ⁇ es each contam only one molecule, then the algonthm calculates a score proprotional to the similanty of the two molecules
  • Algorithm D.l LIBRARYCOMPARE(_ ! , _ 2 ) calculate a similanty score between 0 and 1000 for two molecular hbra ⁇ es
  • the target sites of a protein surface are quantized into the same negative space cubic representation used by theoretical surfaces. This allows the following analyses:
  • the protein quantization process is accomplished in the following steps, as depicted in Figure 30
  • a protein surface is generated from the 3D structure
  • a protein surface is a set of mangles defining the surface ofthe protein that is accessible to water molecules (known as the Connolly surface).
  • Michael Connolly's MSRoll software is an example of a package that can generate a protein surface suitable for this purpose
  • Subsets of the surface which are target sites likely for the binding of small molecules are detected. This can be accomplished, for example, by looking for highly concave regions.
  • Michael Connolly's MSForm software is an example of a package that can measure surface curvature and detect pockets suitable for this purpose
  • Each target site is quantized mto a set of negative space cubes with associated functionalities using the protem function map and the algo ⁇ thm T ARGETSlTE- QUANTIZE.
  • the unde ⁇ ngly process is very similar to the algonthm Q UANTIZE, APPENDIX E PROTEIN QUANTIZATION
  • Each set of quantized negative space cubes with functionality is convened to a set of theoretical surfaces satisfying the proper constraints (for example, no occluded cubes are allowed) using the algo ⁇ thm B UILDSURFACES
  • Algo ⁇ thm E.l npj S r calculate a negative space cubic representation ofthe target site defined by the t ⁇ angles in set T, with v as a normal vector pointing out ofthe target site, / as a functionality map for the entire protein, using parameters resolution r, lattice density n, lattice van Der Waals radius r v , lattice tolerance t;, number of lattice neighbors p , buffer distance b, search radius s r , and additional parameters for subrountme calls (see below) as necessary
  • V C R 3 define a set of points V C R 3 to be the points on a lattice with coordinate frame 7) and cube side length r; such that p € V if p is contained in the target site and the closest triangle in T is at least distance b away from p
  • a target surface set is defined as the set of all theoretical target surfaces to which a set of known protein surfaces map.
  • the target surface set may compnse, for example, all of the surfaces mapped from one protem, all of the surfaces mapped from multiple proteins, or all ofthe surfaces mapped from specific sites on multiple proteins
  • the algo ⁇ thm P ROTEINCOMPARE calculates a score proportional to the similanty ofthe two sets of protein surfaces.
  • the score is calculated by representmg each protein surface set as its target surface set, and usmg the S IMILARITYSCORE p ⁇ mitive to determine the similanty or dissimilanty two sets of theoretical surfaces.
  • Algorithm F.l PROTEINCOMPARE ⁇ I , ⁇ ) calculate a similanty score between 0 and 1000 for two protein surface sets.
  • T to be the target surface set corresponding to V
  • the algorithm P ROTEINLIBRARYCOMPARE calculates a score proportional to the com- plementanty of a library of small molecules and a set of protem surfaces.
  • the score is calculated by representing the protein surface set as the theoretical surface set to which it is similar, the molecular library as the theoretical surface set to which it is complementary, and usmg the S IMILARITYSCORE pnmitive to determine the similanty or dissimilarity two sets of theoretical surfaces.
  • T p to be the target surface set corresponding to V
  • Rotational vanance ( ⁇ r ): 0.1 Polarizable minimum (p m ). 2 Minimum fit quanta (q m ⁇ n ): 9 Minimum slackness ( s m ⁇ n ): 2 Maximum slackness (s m ⁇ :r ): 0 Maximum protrusion levels (p ma x)'- 1 Minimum energy (E mm ): 8.0 kCal
  • Buffer distance (b) 0.5 Angstroms

Landscapes

  • Life Sciences & Earth Sciences (AREA)
  • Spectroscopy & Molecular Physics (AREA)
  • Physics & Mathematics (AREA)
  • Chemical & Material Sciences (AREA)
  • Engineering & Computer Science (AREA)
  • Bioinformatics & Cheminformatics (AREA)
  • General Health & Medical Sciences (AREA)
  • Theoretical Computer Science (AREA)
  • Crystallography & Structural Chemistry (AREA)
  • Health & Medical Sciences (AREA)
  • Bioinformatics & Computational Biology (AREA)
  • Biotechnology (AREA)
  • Evolutionary Biology (AREA)
  • Medical Informatics (AREA)
  • Biophysics (AREA)
  • Medicinal Chemistry (AREA)
  • Pharmacology & Pharmacy (AREA)
  • Computing Systems (AREA)
  • Investigating Or Analysing Biological Materials (AREA)
  • Peptides Or Proteins (AREA)

Abstract

A computer-based method in which a set of constraints is placed on possible target surfaces, and a fully enumerated set of theoretical target surfaces under the given constraints is created, such that each surface has a defined, continuous volume and a defined, continuous surface area. One or more sets of objects are mapped to the fully enumerated set of theoretical target surfaces to define corresponding subsets of the fully enumerated set of theoretical target surfaces. An aspect of diversity of the objects is analyzed based on degrees of similarities and differences among the corresponding subsets.

Description

ANALYZING MOLECULE AND PROTEIN DIVERSITY
This application claims priority from Provisional United States Patent Application Serial Number 60/127,486, filed April 2, 1999, and incorporated by reference.
This invention relates to analyzing molecule and protein diversity. Combinatorial chemistry allows the creation of unprecedented numbers of organic compounds. The rational synthesis of millions of small organic molecules is now achievable in a matter of days. There are estimated to be more than ten to the hundredth power of small molecules that could be synthesized using current methods. By "small", we mean a molecule having fewer than 1500 Daltons, where a Dalton is defined as 1/12 of the weight of a carbon 12 atom or roughly the weight of a hydrogen atom.
An important question is how can one create a set of molecules of such diversity as to contain at least one potent binder to any given target of interest? This question is central to drug discovery in an era that is characterized by a growing wealth of DNA sequence information and a relative dearth of corresponding target structures and their functions. In a world in which there are many more putative targets than can be studied by x-ray crystallography, multi- dimensional NMR, or other high resolution biophysical techniques, any attempt to generate biologically active ligands to targets of unknown structure will require general screening libraries: libraries of molecules that cover a high percentage of so-called "diversity space".
The utility of small molecules as drugs depends in part on molecular complementarity: how well the molecules fit and/or stick to chemically active sites (often in the form of depressions in a protein) on the surface of a cell, on the surface of an intracellular organelle, or on a cytosolic protein. The potential molecular complementarity of a small molecule is in large part determined by two factors: 1) The shape of the molecule, meaning the total Van Der Waals (VDW) surface of a given conformation of the molecule and how it follows or does not follow the VDW surface of the target site of interest. The shape complementarity of molecule and target are largely responsible for energetic forces such as the displacement of water (the "hydrophobic effect") and so called "Van Der Waals" or "London dispersion" forces. 2) The types of potential energetic interactions (such as hydrogen bonding, percentage ionic bonding, proximity of polarizable moieties) found at various places on the molecule, and on the manner, order, and spatial orientation in which the energetically interactive portions ofthe molecule are connected to each other and presented in the presence of the target of interest. Typically, each of these factors is measured or calculated in only a relative way with respect to a specific set of molecules or with respect to specific protein surfaces being examined. A general comparison of potential molecular complementarity between two sets of molecules, however, requires doing calculations or experiments using an absolute or fixed frame of reference. For example, if company 1 has tested a molecule set A for affinity to a set of target surfaces X found in cancerous cells, and company 2 has tested a molecule set B against a set of target surfaces Y found in nervous tissue, the value of molecule set B with respect to the target surfaces X of company 1 is not apparent because each of the affinity evaluations has been performed against a different standard. Also, without a full molecular calculation of set A against target surfaces Y, it is not apparent whether potential chemically active portions of the target surface Y could be bonded by molecules in set A. Finally, even if all calculations of sets A and B vs. surfaces X are performed, neither company will gain a measure ofthe "absolute diversity" of their molecule sets; that is, they will have no measure ofthe likelihood that either sets A or B will contain a molecule that has potential activity against any given target T. This is because the standard against which they have measured their molecules are only a small subset of potential target surfaces.
SUMMARY OF THE INVENTION Implementations ofthe invention provide an absolute or fixed frame of reference for comparing sets of molecules and protein surfaces. By defining molecules in terms of their complementarity or attraction to a fully enumerated basis set of theoretical protein surfaces within given parameters, it is possible to measure the diversity of a set of molecules against a non relative standard, i.e., an absolute measurement. This allows for an efficient comparison of different sets of molecules, for example, sets of drugs, and enables meaningful categorization of classes of molecules against a standard set of surfaces. Furthermore, this allows for the detection of theoretical protein surfaces to which no molecules in a set are complementary, thus enhancing the ability of a chemist to design novel molecules that supplement the deficiencies ofthe original set. By defining real world protein surfaces in terms of their similarity to a basis set of theoretical protein surfaces, it is similarly possible to categorize real world sets of protein surfaces against a standard set of theoretical surfaces. This allows for improved classification of proteins into similar target classes by the similarity of their surface sites. By evaluating an actual set of molecules against theoretical protein surfaces and further by evaluating a set of real world protein surfaces of interest against the same theoretical protein surfaces, it is possible to select classes of molecules that are likely to have substantial activity against the real world protein surfaces. This allows for improved molecule screening, for example, in drug research. By evaluating an actual set of molecules against theoretical protein surfaces and further by evaluating a set of real world protein surfaces of interest against the same theoretical protein surfaces, it is also possible to find actual protein surfaces to which no molecule in the set A has likely activity. This enhances the ability of a chemist to design molecules beyond those in set A that match the previously unmatched protein surfaces of interest, and may thereby be tested for pharmaceutical activity, thus supplementing deficiencies in the original set of molecules.
Thus, in general, in one aspect, the invention features a computer-based method in which a set of constraints on possible target surfaces is defined, and a fully enumerated set of theoretical target surfaces under the defined constraints is also defined, such that each surface has a defined, continuous volume and a defined, continuous surface area. One or more sets of objects are mapped to the fully enumerated set of theoretical target surfaces to define corresponding subsets of the fully enumerated set of theoretical target surfaces. An aspect of diversity ofthe objects is analyzed based on degrees of similarities and differences among the corresponding subsets.
Implementations ofthe invention may include one or more ofthe following features. The target surfaces may include negative space target surfaces. The objects may include positive space object surfaces associated with different molecules. The objects may be mapped by defining corresponding subsets ofthe fully enumerated set of negative space theoretical target surfaces to which positive space object surfaces of conformations of molecules are complementary. The aspect of diversity that is analyzed may be the difference or similarity between the molecules which map to those negative space theoretical target surfaces.
The objects may include negative space object surfaces associated with different proteins, and the objects may be mapped by defining corresponding subsets ofthe fully enumerated set of negative space theoretical target surfaces to which negative space object surfaces of protein pockets are similar. The aspect of diversity that is analyzed may be the difference or similarity between protein pockets which map to those negative space theoretical target surfaces. The objects may include positive space object surfaces associated with different molecules and negative space object surfaces associated with different proteins. In the case of molecules, the objects may be mapped by defining corresponding subsets ofthe fully enumerated set of negative space theoretical target surfaces to which positive space object surfaces of conformations of molecules are complementary. In the case of proteins, the objects may be mapped by defining corresponding subsets of the fully enumerated set of negative space theoretical target surfaces to which negative space object surfaces of protein pockets are similar. The aspect of diversity that is analyzed may be the difference or similarity of the molecules which map to those negative space theoretical target surfaces to the protein pockets which map to those negative space theoretical target surfaces.
The theoretical target surfaces and the objects may be polyhedrons, e.g., cubes, all ofthe same size and shape. The set of all theoretical target surfaces defines a diversity space within which the diversity of objects can be measured by mapping those objects to the diversity space. Regions of the diversity space to which no objects map may be identified, and molecules may be designed that occupy at least one ofthe unfilled theoretical target surfaces ofthe diversity space. Complementarity may be associated with binding affinities of positive space object surfaces of conformations of molecules to negative space theoretical target surfaces.
The constraints may include volume, associations of each of a number of sites of the target surface with a preselected molecular property drawn from a larger set of possible molecular properties, including hydrophobic, polarizable, H-bond acceptor, H-bond donor, H-bond donor/acceptor, potentially positively charged, and potentially negatively charged. Fewer than all of the sites ofthe target surface may each be associated with a different one ofthe molecular properties and all of the other sites of the target surface may be associated with a common molecular property, such as slightly hydrophobic. The degrees of similarities or differences may involve functional properties associated with the corresponding subsets of the fully enumerated set of theoretical target surfaces or shape properties associated with the corresponding subsets ofthe fully enumerated set of theoretical target surfaces. Each of the objects may be defined by quantizing molecules into polyhedrons. Each of a fixed set of orientations of each conformation of each of the objects may be fitted to each of the target surfaces, and each ofthe fittings may be scored.
The constraints may include a resolution ofthe polyhedrons, e.g., 4.24 Angstroms, or maximum and minimum numbers of polyhedrons. Each of the polyhedrons may share a common interface with another of the polyhedrons. The constraints may also include the absence of any occluded volumes greater than a given user-defined parameter. The target surfaces may be defined conceptually as having been carved out of a flat surface. In general, in another aspect, the invention features categorizing existing molecules based on negative space target surfaces to which conformations ofthe molecules are complementary, and designing novel molecules that are complementary to negative space target surfaces to which no conformations of the existing molecular are complementary. In general, in another aspect, the invention features a method of creating novel molecules to be tested as ligands for proteins. In the method, proteins are categorized based on target surfaces to which their pockets of known structure map, and novel molecules are designed that are complementary to the negative space target surfaces to which the protein pockets map. In general, in another aspect, the invention features a computer programmed to determine the chemical similarity of different molecules. The program approximates the surface shape of each one of a plurality of molecules of interest by linking a series of cubes, each cube having a dimension R, the locations of the cubes being determined by the calculated electron probability density of the individual one of the molecules of interest, each cube sharing at least one of its six faces with another cube, such that there is a specific number of linked cubes which varies for each individual one of the plurality of molecules of interest. The chemical reactivity of each individual one of the plurality of molecules of interest is approximated by assigning each cube of each individual one of the plurality of molecules of interest, no more than one functionality value from a plurality of M different chemical functionality values. The surface shape and chemical reactivity of a chemically active surface having a volume equal to
V is approximated by subtracting a number V R^ cubes of dimension R from a surface, wherein each of the cube spaces shares at least one face with another cube space and wherein N of the cube spaces has one of a plurality of M different chemical functionality values. An attraction value K is calculated for each one of the plurality of molecules of interest to the chemically active surface. A list of overall attraction values to the chemically active surface is calculated.
Implementations ofthe invention may include one or more ofthe following features. The calculation of the attraction value K may be performed on a plurality of different predetermined chemically active surfaces, and a matrix of overall attractive values of each molecule of interest to each of the different surfaces may be calculated. The molecules of interest may include organic molecules. The chemically active surface having a plurality of predetermined active chemical locations may be calculated to correspond to the shape of an actual protein surface structure. The molecules of interest may be organic molecules of 1500 Daltons or less. The chemically active surface having a plurality of predetermined active chemical locations may be compared to an actual protein surface to calculate a similarity value ofthe actual protein surface to the predetermined active chemical locations. The predetermined chemically active surfaces may be compared to a plurality of actual protein surfaces and a matrix of similarity values may be calculated. The cube spaces subtracted from the surface may be calculated to approximate the electron probability density of at least one of a plurality of depressions in known protein surface structures. The N sites of chemical functionality may be calculated to approximate the location and type of chemical functionality of actual depressions in known protein structures.
Other advantages and features will become apparent from the following description and from the claims.
DESCRIPTION
Figure 1 shows a (CH2)n chain encapsulated by 4.24 A cubic units. Figure 2 shows examples of surfaces allowed and disallowed by the non- occlusion parameter in a theoretical target surface generation algorithm. Gray shading represents the opening ofthe theoretical surface. A is allowed. B is disallowed due to two occluded negative space cubes (marked X). Figure 3 shows a theoretical target surface of 13 negative space cubes and four sites of specific molecular property interaction: hydrophobic (white), polarizable (purple), H-bond accepting (green), and H-bond donating (orange). Blue shading indicates the opening of the theoretical surface. Figure 4 shows a "quantized" representation (Q-file) of one conformation of molecule 6a superimposed on its atomic structure (ball and stick and spacefilling model). Molecular property characteristics ofthe Q-file are hydrophobic (white quanta), polarizable (purple quanta), H-bond accepting (green quanta), and negatively charged (red quanta). Figure 5 shows test molecules.
Figure 6 shows a ranking of molecules by QCSD similarity scores.
Figure 7 illustrates examination of a theoretical target surface common to molecules 8c (top) and 8a (bottom). Blue shading indicates opening ofthe theoretical surface. Specific points of complementarity on the theoretical target surface are hydrophobic (white) polarizable (purple) and H-bond donating (orange). Superimposition ofthe original molecular conformations onto the theoretical target surface demonstrates that the extra phenyl substituent of 8c protrudes from the opening ofthe theoretical surface and is not involved in complementarity to the surface. Figure 8 illustrates examination of a theoretical target surface common to molecules la (A) and 5a (B). Blue shading indicates opening of the theoretical surface. Specific points of surface complementarity found by QSCD are hydrophobic (white), polarizable (purple), and H-bond donating/accepting (yellow). Overlay plot (C) of the non-hydrogen backbones of la (orange) and 5a (green) indicate similar features. Aromatic/hydrophobic overlaps are shown in purple; H-Bond donating oxygens are in red. Overlay generated with Sybyl version 6.5 (Tripos Ine, 1699 S. Hanley Rd., St Louis, MO, 63144).
Figure 9 illustrates ranking of molecules in Figure 5 by Tanimoto similarity score of 2D UNITY fingeφrints. Figure 10 shows a QSCD plot of all ofthe theoretical surface shapes covered by all ofthe conformations of all of the molecules (blue dots) in Figure 5. The total volume of the cube encompasses all 49,268,918 theoretical surface shapes as listed in Table 1. Red dots show two exemplary theoretical surface shapes (a, b) not covered by any ofthe molecules in Fig. 5. Axes used are functions of opening area, opening length/width, and depth per opening quantum. Figure 11 shows a map ofthe 20 compounds in Fig. 5 (blue dots) in a representative BCUT three-axis diversity space. BCUT axes used are, respectively: 1) BCUT HACCEPT S LNVDIST 050 R H, 2) BCUT HDONOR S LNVDIST 030 R H, and 3) BCUT TAB POLAR S LNVDIST 300 R L. Red dot shows an unfilled coordinate of diversity space, at (7.54, 7.25, 6.82). The information contained in this BCUT coordinate does not reveal information about the shape of a molecule which might be able to fill this position in diversity space.
Figure 12 shows use of QSCD to design complementary combinatorial libraries to unmatched theoretical target surfaces. Many conceivable libraries of a given shape and functionality may be designed to fill a given unmet diversity need.
Figure 13 shows two sample surfaces. Figure 14 illustrates the quantization process.
Figure 15 shows a legend for symbols used in functionality rule diagrams. Figure 16 shows functionality rules for Potential Negative Charge
Functionality. The structures are searched for in order.
Figure 17 shows functionality rules for Potential Positive Charge Functionality. The structures are searched for in order.
Figure 18 shows a functionality rule for Hydrogen Bond Donor/ Acceptor Functionality.
Figure 19 shows a functionality rule for Hydrogen Bond Donor Functionality.
Figure 20 shows a functionality rule for Hydrogen Bond Acceptor Functionality. The structures are searched for in order. Figure 21 shows a functionality rule for Polarizable Functionality. Figure 22 shows Table 5, a ranking of molecules in Fig. 5 by QSCD diversity score. Blue = homogeneous pairs, yellow = +phenyl pairs (8c), green = ATI-AT2 pairs (3,4)
Figure 23 shows Table 6, a ranking of molecules in Fig. 5 by Tanimoto similarity score of 2D UNITY fingeφrints. Blue = homogeneous pairs, yellow = +phenyl pairs (8c), green = ATI-AT2 pairs (3,4)
Figure 24 shows an example subset of theoretical surfaces Ti containing 4 members and an example central set Ci for Ti (F = 1, E = 3) where the black face denotes a point of attachment Al on Ci: Figure 25 shows an example Core Molecule Mi to fill Central set Ci.
Figure 26 shows an example Library (Mi,B) where B = a set of amines. Figure 27 shows an example subset of target surfaces Ti containing 4 members and an Example Central set Ci for Ti (F = 1, E = 3), where the black face denotes a point of attachment Al on Ci: Figure 28 shows an Example Core Molecule Mi to fill Central set Ci.
Figure 29 shows an example Library L(Mi, B) where B = a set of amines. Figure 30 shows a protein quantization process.
We define diversity as the measure, based on pre-defined criteria, ofthe difference or similarity among all members of a set. In a pharmaceutical setting, molecular diversity can be defined as the measure, based on biological criteria, of the difference or similarity between small molecules. Each ofthe existing methods of calculating biologically relevant diversity of small molecules defines slightly different criteria for molecular comparison, and thus a different configuration of diversity space as a whole. Examples include low dimensional diversity space such as BCUT metrics, high dimensional diversity space such as Chem-X/ChemDiverse multiple point pharmacophores, and empirical biological diversity space such as affinity fingeφrinting. Many known ways of quantifying the diversity of molecules use molecular properties such as functionality and connectivity as a basis for categorization (see for instance Potter and Matter, J. Med. Chem., 1998, p. 478). For example, in the BCUT method used to generate four- to six-dimensional diversity space, molecules are broken down into matrices according to connectivity and molecular interaction properties. Coordinates in diversity space are assigned through the resulting eigenvalues of these matrices, leading to useful multi-dimensional plots of molecular diversity. However, because the use of eigenvalues is an irreversible transformation (different 3D shapes can map to the same eigenvalues), it follows that an empty coordinate in BCUT diversity space cannot be translated into a 3D template of a "missing molecule." Thus, while a model such as BCUT diversity is well validated as a tool for finding combinatorial matches to a lead compound or pharmacophore, it cannot be . directly used to populate the entire diversity space that it defines.
Similarly, in the popular Chem-X/ChemDiverse diversity package, molecules are broken down into all accessible three- or four-point pharmacophores of triangular or tetrahedral functionality distances. If the model is used to display molecular diversity, coordinates in diversity space are assigned through the resulting string of accessible three- or four-point pharmacophores; this method has been shown to be highly effective in classifying molecules by pharmacological similarity. However, the mapping of complex 3D shapes to a set of triangular or tetrahedral functionality distances is an irreversible transformation; empty three- or four-point pharmacophores in Chem-X derived diversity space cannot be translated into a 3D template of a complex shape. Since a set of coordinates in Chem-X is insufficient to define the shape of "missing molecules," Chem-X cannot be used to directly populate empty molecular diversity space.
Another example of current diversity methods is affinity fingeφrinting, in which molecules are empirically assayed against a panel of 10-20 actual proteins selected to be promiscuous in their ability to bind small molecules. Position in molecular diversity space is assigned through the resulting string of IC50 binding values, and these affinity fingeφrints provide unprecedented ability to group similarly active compounds in diversity space. However, because the actual mode of binding in any assay is not incoφorated in the resulting IC50 value, the mapping of molecules to the selected protein panel is an irreversible transformation. Thus, an empty coordinate in affinity fingeφrinting diversity space (an "unmatched" string of IC50S to a given protein panel) cannot be back- translated into a 3D molecular template. A similar affinity fingeφrinting diversity method has been put into practice using a panel of computational surfaces of real- world protein pockets and a modified form ofthe DOCK program. While this method shows similar promise in its ability to detect pharmacological similarity, it is, like its empirical affinity fingeφrinting counteφart, an irreversible mapping. Thus, for the most part, current methods are able to successfully identify compounds of the same pharmacological class as being similar and compounds of different pharmacological classes as being different. Given a starting pharmacophore from known ligands and/or the target site of a target crystal structure, such methods interface well with the design of complementary combinatorial libraries.
The design of combinatorial libraries to cover all of diversity space is a rather different problem, however. In this case, it is not enough to be able to compare existing molecules for differences or similarities. In addition to being able to place molecules relative to one another in diversity space, one must be able to point to an absolute area of diversity space not yet covered and from its coordinates design a novel set of compounds to fill that uncovered space.
In order to rationally and systematically fill diversity space, an informationally reversible diversity model is needed. This model must be formulated such that: members (in this case molecules) can be assigned to coordinates for similarity/dissimilarity comparison, and empty coordinates retain the information necessary to directly generate coordinate membership.
One good path for such a model is to use as coordinates the exact information that differentiates one member from another, without intervening, iπeversible transformations. To apply this reasoning to molecular diversity, it must first be asked: what are the criteria by which diversity of compounds is to be measured (what information differentiates one molecule from another). One of the most fundamental criterion in molecular drug discovery is the extent to which two molecules have similar or different binding affinities to a given target. With the assumption that similar binding affinity tracks with a molecule's complementarity to similar target surfaces, we have selected as our criterion for diversity complementarity to a fully enumerated set of theoretical target surfaces. Given the above definition of molecular diversity, it remains to provide parameters under which to define a biologically relevant basis set of enumerated theoretical target surfaces and to quantify molecular complementarity to a given theoretical target surface at a level which is both in accordance with known principles of molecular recognition and computationally applicable to millions of compounds. With a numerical determination of complementarity and a biologically relevant basis set of surfaces, molecular diversity space is thus absolutely established as the molecular complement to a fully enumerated set of theoretical target surfaces. We introduce the concept quantized surface complementarity diversity
(QSCD), which defines a molecule numerically by a mapping that describes its complementarity K to every distinct theoretical protein surface of resolution R not exceeding volume V with N sites of M types of chemical functionality PMN- K is defined as an algorithm that takes into account the molecular shape and chemical functionality of both the given molecule and the given theoretical protein surface. From this definition, it follows that a comparison of two molecules will yield a numerical difference that is representative of their complementarities: to the extent that two molecules each have complementarity for the same theoretical protein surfaces, the molecules are similar; to the extent that two molecules have no complementarity to common theoretical protein surfaces, the molecules are dissimilar. In other words, "similarity" between molecules is defined as the ability to complement the same theoretical protein surfaces and "difference" between molecules is defined as the ability to complement different theoretical protein surfaces. Because QSCD uses complementarity to theoretical protein surfaces as a basis for categorization, both 3-D shape and molecular functionality are taken into account. Also, because QSCD uses as its basis a complete set of theoretical protein surfaces (under R, V, and PMN), the method provides diversity information in a fixed frame of reference: coordinates in quantized surface complementarity diversity space are independent of any molecules or natural protein surfaces compared in that space. Because of this, not only can the diversity of disparate sets of molecules be compared without having to compare the molecules to each other directly, but the diversity of any set of actual natural protein surfaces can be examined through complementarity to the theoretical basis set. In addition, complementarity of molecules to a set of theoretical protein surfaces representative of actual natural protein surfaces can be examined in the context both of any other set of molecules and any other set of protein surfaces. Furthermore, given a set of molecules, it becomes immediately apparent what percentage of theoretical protein surfaces are covered by complementary molecules, thus giving a measure of the set's molecular diversity in the space defined by all potential surfaces under R, V, and PMN- Not only does this provide a measure of diversity in an absolute sense which is not relative to any historically biased set of surfaces or molecules, but it also makes clear a set of theoretical target surfaces to which no molecules in the initial set bind, allowing a straightforward design of novel molecules to supplement the initial set. Theoretical Target Surfaces
In one implementation, to generate a finite set of theoretical target surfaces that approximates all possible binding pockets with volume equal to or less than V, we consider each theoretical surface to be formed by successively carving c ubic units out of an initially flat surface. These cubic units represent "negative space" that a potential ligand could occupy. Given cubic units with sides of length R (the resolution ofthe model), we use at most V/R3 negative space cubes to describe each theoretical target surface. Others have previously employed cubic units to successfully approximate complementarity between small molecules and individual protein surfaces. The size of a negative space cube is directly related to the resolution and type of diversity data which the user desires as output. In choosing the size ofthe negative space cube, one motivation is to maximize negative space cube size such that the difference of a single cube in a surface is highly differentiating in terms of molecular recognition (i.e., every surface is orthogonal to every other surface). At the same time, enough information must be retained in each negative space cube to predict shape and functional complementarity at a ligand/surface interface. The former constraint minimizes overlap of diversity information while the latter constraint maximizes precision of diversity information. Together, the competing constraints result in a basis unit for the enumeration of theoretical target surfaces that minimizes the number of negative space cubes needed to accurately model diversity for a given volume V.
A resolution of 4.24 A negative space cubes was found by computer optimization of test molecules to provide an upper limit of cube size while still maintaining an acceptable level of molecular shape information. Interestingly, 4.24 A is the approximate VDW "cross-section" of a (CH2)n chain; a series of 4.24 A units neatly encapsulates a (CH2)n chain in its ground state conformation as shown in Figure 1.
In one implementation the basis set for diversity can be a set of theoretical target surfaces comprised of all possible shape combinations of 6 to 14 negative space cubes of resolution 4.24 A (negative volume between 460 and 1070 cubic A) subject to the following rules: Surfaces are created by successively "carving out" negative space cubes from a flat block of infinite width and depth (the theoretical target). All negative space cubes of a given surface must share at least one face with another negative space cube of the surface, and all must be part of a single, contiguous negative surface. No negative space cubes may be occluded in the +Z axis ofthe infinite surface block; that is, there may be no solid surface between any negative space cube and the surface plane ofthe infinite block. As shown in figure 2, the surface A is allowed, but the surface B is disallowed. Surfaces duplicating a previous surface with respect to rotation in the X-Y plane are discarded. The occlusion rule provides a compromise between complete coverage of topological possibilities and acceptable computational speed. This compromise was made based on the topological assumption that occlusions of 4.24 A or more are infrequent in small molecule/target interactions, and that their omission would thus have only a small effect on predicting diversity of binding affinities of small molecules. Applying the rules yields 49,268,918 unique negative surface shapes including chiral opposites. Covering a negative volume between 460 and 1070 cubic A, these surface shapes are deemed sufficient to examine diversity of most small molecules. For instance, examining a previously published reference set of pharmaceutically relevant compounds (a filtered Comprehensive Medicinal Chemistry or CMC database), 5049 out of 5120 compounds (98.6%) have a volume of 1070 cubic A or less.
Within each ofthe 49,268,918 unique negative surface shapes, each negative space cube is assigned a molecular property characteristic Pm that represents the dominant molecular environment which any atoms that are placed within that negative space will experience. Properties used are PI hydrophobic, P2 polarizable (includes aromatics), P3 H-bond acceptor, P4 H-bond donor, P5 H-bond donor/acceptor, P6 potentially positively charged (basic), and P7 potentially negatively charged (acidic). These seven types of molecular environments are assumed to represent a minimal basis set of factors that contributes to the electrostatic/VDW complementarity of a ligand and a target surface. In one implementation, four positions of particular molecular property PI -7 are assigned, leading to 74*N!/((N-4)!*4!) surfaces for each surface shape of N negative space cubes. All other (N-4) cubes not assigned a particular molecular property are given property P8, slightly hydrophobic. The latter assignment is based on an assumption that hydrophobic effects are, on average, the largest single component contributing to ligand/target interaction. In sum, the above process implies as a basis set for molecular diversity 1.1 * 1014 theoretical target surfaces of negative volume between 460 and 1070 cubic A and having four sites of specific molecular property characteristics PI -7. The numerical breakdown of these 110 trillion surfaces is listed in Table 1. Table 1 : Numerical breakdown ofthe total number of theoretical target surfaces created using the algorithm given in the text. Surfaces consist of 6-14 negative space cubes and 4 sites of 7 possible molecular property characteristics. Number of functionally different surfaces per surface shape varies for infrequent cases in which a given shape has an axis of symmetry, so actual number of unique surfaces is slightly less than (# surface shapes) * 74*N!/((N-4)!*4!).
Volume Number of Approx. number Exact number of (number unique of functionally unique surfaces
(N) of surface different surfaces negative shapes per unique space surface shape: cubes) 74*N!/((N-4)!*4!)
6 212 36,015 7,163,338
7 885 84,035 73,271,443
8 3,959 168,070 655,324,488
9 17,747 302,526 5,350,917,208
10 81,407 504,210 40,912,578,322
11 375,897 792,330 297,622,676,624
12 1,753,218 1,188,495 2,082,225,979,379
13 8,224,443 1,716,715 14,116,888,070,845
14 38,811,150 2,403,401 93,264,917,290,356
Total 6-14 49,268,918 109,808,653,272,003 One such surface is shown in Fig. 3.
A pseudocode description of an algorithm for determining the set of theoretical target surfaces is set forth in Appendix A. Molecular Quantization
To measure complementarity of small molecules to the basis set of theoretical target surfaces, the small molecules must be formatted in a similar frame of reference, for instance by quantizing them into positive space cubes ("quanta") of resolution 4.24 A according to the following process (illustrated in Fig. 4):
A set of up to 100 minimized energy conformations within user-defined parameters is created. In one implementation, Tripos Multisearch modeling is used, and all conformations within 10 kcal of the lowest energy conformation found are accepted. For each conformation, a 4.24 A 3D grid of cubes (quanta) is aligned on top of the 3D structure using the molecule's principle axes of rotation (calculated with all atoms having mass 1).
To all 4.24 A quanta which contain at least a user-defined % of the VDW radius of any atom, a dominant molecular property characteristic is assigned based on connectivity rules (e.g. R-[C=0]-0-H yields P , R-O-H yields P5; see definitions of Pi - P7 above). Order of dominance is from P7 to Pi, in order of maximum complementarity score obtainable by a given characteristic as shown in Table 2:
Table 2: Relative magnitudes of parameters used in calculating molecular property interactions between negative space cubes (theoretical target surfaces) and positive space cubes (quantized molecules). Magnitudes (listed from highest to lowest): +++, ++, +, 0, -, --, — .
Figure imgf000020_0001
Minimum % of VDW radius parameter allows for a user-defined protrusion beyond the surface of a quantum cube, adding a measure of topological "flexibility" to the quantization process. A user defined 32% was found to be especially good.
The total number of 4.24 A quanta that have been assigned a property characteristic is counted.
The grid alignment is shifted per user-defined parameters and the process is repeated until all shift combinations have been searched.
For each conformation in the original set, a "Q-file" (3D configuration of property-assigned quanta) is saved that has the lowest number of quanta in and is closest to the principle alignment. Thus, an average molecule in this implementation is represented by 100 Q-files, each file consisting of N positive space cubes or quanta of 4.24 A resolution having an assigned molecular property characteristic Pm (m=l-7). A typical Q-file (molecule 6a) is shown in Fig. 4, superimposed upon its coπesponding conformation. The process of optimization of quantization parameters is described later.
A pseudocode description of an algorithm for performing the quantization is set forth in Appendix B. Mapping Given molecules which have been rendered into sets of Q-files, each quantized conformation can be mapped into the diversity space defined by the set of 1.1 x 1014 theoretical target surfaces. In general, the following process is used. For each quantized conformation of each molecule, each of its 24 possible X/Y/Z rotations (6 faces * 4 rotations per face) is fit to each ofthe 49,268,918 available surface shapes. For a given conformation-to-surface shape fit, if at least a user-defined minimum number of negative and positive space cubes overlap (in one implementation either 9 quanta or N-2 quanta of a conformation of N quanta), and if no quanta ofthe conformation extend beyond the bounds ofthe surface shape except at the mouth ofthe surface shape, then the complementarity of the quantized conformation to all theoretical target surfaces of that shape is examined in detail as explained next. If the above conditions are not met, the next conformation is examined.
A score is generated for the complementarity ofthe given conformation to each theoretical target surface of a given shape from based on user-defined parameters. (The process of optimization ofthe complementarity parameters is described later.) The following complementarity parameters can be used: a) A negative parameter for each rotatable bond of the conformation. b) If conformational energies are calculated, a negative parameter for the energy ofthe conformation above the lowest energy conformation from that molecule. c) A positive parameter for the hydrophobic energy gained by removing "water" from any hydrophobic (Pi) or polarizable (P2) surface face of either the conformation or the theoretical surface. d) A positive parameter for the hydrophobic energy gained by removing "water" from any mildly hydrophobic (P ) surface face of the theoretical surface. e) A positive or negative molecular property interaction parameter for overlapping negative and positive space cubes as depicted in Table 2.
If and only if the resulting score meets a user-defined minimum, then the conformation (and thus the molecule it represents) is said to be complementary to the given theoretical target surface.
The computational advantage inherent in the process of molecule and surface quantization is realized in the speed of complementarity checking.
Whereas a traditional docking program must search a high-dimensional configuration space, the implementations ofthe invention resolve the problem to a framework bounded by 24 possible fitting orientations and a finite number of translations. This approximation allows three-dimensional diversity computation on a scale that is applicable to very large sets of molecules.
A pseudocode description of an algorithm for performing the mapping is set forth in Appendix C.
The above process results in a complementarity map that consists of a list of all theoretical target surfaces to which at least one conformation of a molecule is complementary. Comparison of these maps provides a novel method for measuring diversity of small molecules. We term the model on which this process is based quantized surface complementarity diversity (QSCD) because it calculates diversity by measuring complementarity to a quantized representation of theoretical target surfaces.
To maintain a computationally efficient complementarity scoring system,
QSCD makes many approximations of molecular recognition. As explained, these include cubic units of 4.24 A resolution, gross approximations of surface contact area, exactly 4 points of 7 finite types of molecular property characteristics, static theoretical surfaces, and a limited set (up to 100) of low energy conformers. Thus, the final complementarity scores are not presumed to give precise binding energies for any individual match of conformation to target surface. However, taken over all conformations of a molecule and across an enumerated set of theoretical target surfaces, the scoring system is statistically relevant as explained below. Model Validation
To test the validity ofthe QCSD model, i.e., its ability to predict the extent to which two molecules have similar or different binding affinities, eight sets of test molecules were analyzed (Fig. 5), seven of which were known to have binding affinities to seven distinct targets (in addition to a known overlap between sets 3 and 4). An eighth set with no known binding affinities was chosen with minor atomic and spatial changes to examine the sensitivity ofthe QCSD model at 4.24 A resolution. Known activities ofthe molecules in Fig. 5 are listed in Table 3 with references.
Table 3: Pharamacological activities ofthe molecules used in this study (see Fig. 5). a) Doherty et al. J. Med. Chem. 1995, 35, 1259-1263. b) Uehling et al. J. Med. Chem. 1995, 38, X 106-1118. c) Chang et al. J. Med. Chem. 1994, 37, 4464-4478. d) Chang et al. J. Med. Chem. 1993, 36, 2558-2568. e) Tsutsumi et al. J. Med. Chem. 1994 37, 3492-3502. f) Penning et al. J. Med. Chem. 1995, 35, 858-868. g) Cristalli et al. J. Med. Chem. 1995, 35, 1462-1472. h: numbers in parentheses indicate IC50 in ATI subtype assay of series 4. i: numbers in parentheses indicate IC50 in AT2 subtype assay of series 3.
Assay IC50 or Kj (nm) Ref.
1a Binding to Endothelin A Receptor 400 a
1 b Binding to Endothelin A Receptor 170 a
2a Inhibition of DNA fragmentation by 28 b Topoisomerase I
2b Inhibition of DNA fragmentation by 143 b Topoisomerase I
3a Binding to AT2 subtype of Angiotensin II 17 (0.45)h c Receptor
3b Binding to AT2 subtype of Angiotensin II 173 (31 )h c Receptor
4a Binding to AT1 subtype of Angiotensin II 0.85 d Receptor 4b Binding to AT1 subtype of Angiotensin II 1.4 Receptor
4c Binding to AT1 subtype of Angiotensin II 1.2
Receptor (23,000)'
5a Inhibition of Prolylendopeptidase Protease 5 Activity
5b Inhibition of Prolylendopeptidase Protease 10.3 Activity
6a Binding to Leukotriene B4 Receptor 320 f
6b Binding to Leukotriene B4 Receptor 3.2 f
7a Binding to A2A Type Adenosine Receptor 6.3 g
7b Binding to A2A Type Adenosine Receptor 41.3 g
8a none —
8b none —
8c none —
8d none —
8e none —
The bulk of these molecules have previously been used as part of an in- depth study validating molecular descriptor approaches for the prediction of molecular diversity within compound classes. This is a more stringent discrimination than the base criterion sought for the QCSD model, which seeks at a minimum to show accurate diversity prediction between compound classes.
Conformations of all 20 test molecules were "quantized" and then mapped onto the basis set of 1.1 * 1014 theoretical surfaces. Complementary surfaces are tabulated for each molecule in Table 4.
Table 4: Tabulation of surface shapes and total number of theoretical target surfaces complementary to each molecule in Fig. 5.
Complementary Complementary surface shapes surfaces (shape plus functionality)
1a 376 16,127,687
1b 379 9,086,768
2a 27 545,584
2b 27 416,210
3a 414 4,970,816
3b 315 813,024 4a 487 4,542,463
4b 479 12,388,826
4c 482 7,595,982
5a 337 2,080,523
5b 374 1,966,837
6a 220 192,067
6b 186 153,436
7a 362 298,927
7b 269 22,367
8a 45 5,561,654
8b 41 3,959,678
8c 333 17,324,247
8d 64 2,059,546
8e 87 1,343,811 average: 4,572,523
There are many ways to analyze the resulting set of complementarity mappings. Because in this case individual molecule comparisons were desired, each ofthe 20 mappings was compared pairwise for a total of 190 data points. Mappings were scored in similarity from 0 to 1000 based on a function of the number of theoretical surfaces in common:
Score = SS * FS = ShapeScore * FunctionahtyScore
SS = 100 * # theoretical target surface shapes common to A & B total # surface shapes complementary either to A or to B
Figure imgf000025_0001
The first term in this equation gives a percentage measure (0-100) of shape similarity between molecules A and B, while the second term gives a measure from 0-10 of functional similarity per given shape overlap. The complete scores are detailed in Table 5 contained in Figure 22. Using this scoring system, the maximum score obtainable by very rigid, structurally similar molecules is 1000. However, many molecules can only be sampled by an examination of up to 100 low energy conformations (an average molecule w/5+ rotatable bonds will have at least 35 = 243 conformations). Thus, for most molecules with more than 100 accessible conformations, similarity scores between 0-100 are observed. The scoring constant Φ in the equation above adjusts the influence of functionality on scoring. A value of 0.33 was found to be optimal (as discussed below), meaning that shape is the dominant criterion in our measure of diversity. A pseudocode description of an algorithm for determining either the similarity of two molecules or the similarity of two libraries of molecular structures is set forth in Appendix D.
Fig. 6 shows a plot of all 190 pairings ranked by similarity score. Circles show "heterogeneous" pairs of expected dissimilarity (e.g. 2a, 6b), while squares show "homogeneous" pairs of expected similarity (e.g. 2a, 2b). Clearly, the QCSD model ranks homogeneous pairs almost exclusively higher than heterogeneous pairs; all 1 pharmacologically similar pairs fell within the top 20 scores out of 190. All homogeneous scores were ranked above 25, while the median score in this experiment was 2.8, showing good "signal to noise." The QCSD model is thus a valid predictor of target binding similarity among these molecules.
The pairings also reveal further validation. As might be expected from their relative rigidity (low number of accessible conformations) and structural similarity, the highest scoring pairs are 2a/2b, 8a/8b, and 8d/8e. Furthermore, examination ofthe pairings of 8c with 8a,b,d,e (triangles in Fig. 6, yellow in Table 5) yields scores that are within the top 20% of the pairing experiment but which are generally lower that the "homogeneous" pairs. This makes sense from a target-binding point of view, considering that one face of 8c contains a large molecular difference (an extra phenyl substituent). To the extent that this face is not involved in complementarity to a target surface, the molecules are similar; to the extent that this face must be complementary for binding to occur, the molecules are quite different. Fig. 7 shows one such case of a surface common to both 8a and 8c; the protruding phenyl substituent plays no role in complementarity. As a rule, there is close similarity of shape and functionality for molecules which score 25 or higher. In addition to the pairings that one would expect, several other pairs scored between 25-35. When these molecules were examined by molecular modeling, significant overlaps were found, suggesting that these high scores are not just "noise" in the QCSD model. Fig. 8 depicts one such case between la and 5a; conformations of la and 5a are displayed that were found in the QCSD model to be complementary to the same surface (Fig. 8A, 8B). 3D overlays (Fig. 8C) confirm correlation of general shape and 4 points of functionality, although they also make clear the limits of resolution of complementarity information using 4.24 A units. As can be seen from Fig. 8, the surface in question can detect general shape and functional similarity, but by does not provide a basis to predict atom-for-atom overlap between molecules.
A final result comes from examination ofthe QCSD model's rankings of sets 3 and 4. While set 4 is known to bind exclusively to the ATI subtype of the Angiotensin II receptor, set 3 is known to bind to both the ATI subtype and the AT2 subtype. While the QCSD model found high similarity within sets 3 (score = 27) and 4 (avg. score = 54), it found an average similarity of 6.9 between 3a and set 4 and an average similarity of 3.3 between 3b and set 4 (diamonds in Fig. 6, green in Table 5 contained in Figure 22). Based on the QCSD model, one would therefore conclude that while sets 3 and 4 share a limited number of complementary theoretical surfaces, they are dissimilar with respect to the majority of theoretical target surfaces. This is in fact the case with the AT2 subtype of the Angiotensin II receptor, to which 4c binds 50,000 times more poorly than 3a (see Table 3). Advantages of the QCSD model
Having validated the basis set used for the QCSD model in the classification of molecular diversity, it must be noted that other models may do as well or better in detecting target binding similarity/ dissimilarity between molecules. For instance, Fig. 9 and Table 6, contained in Figure 23, show the same set of 20 molecules ranked by Tanimoto similarity of standard 2D UNITY fingerprints (see discussion below). The data demonstrate that the 2D model is equally capable of predicting pharmacologically similar pairs; UNITY ranks similarity between ATI and AT2 subtype binders much higher than our QCSD model, although it finds unusually high similarity between 8a and 8c. In general, such 2D fingerprint descriptors have been found effective in clustering pharmacologically similar compounds, and are widely used in determining molecular diversity of existing structures.
Among the advantages of the QCSD model is the value of its negative information: The QCSD model determines not only diversity of existing structures, but also the structure of non-existing diversity. Given theoretical surface shapes for which no complements exist in a general screening library, QSCD allows the design of molecules to fill the given diversity void.
As stipulated in its formulation, the QSCD basis set is created through a reversible process. Although some information resolution may be lost in fixing the parameters of a cube's size and functional scope, information content is retained in either direction. Just as a single molecular conformation and orientation corresponds to a defined pattern in QSCD space, likewise, a single point in QSCD space (within the limits of volume V, resolution R, and N sites of functionality Pm) corresponds to a unique 3D shape with a defined 3D array of functionality. Given any starting set of molecules, unoccupied points in QSCD space directly define the molecular shapes and functionalities which those molecules do not cover. Thus, a set of detailed 3D molecular templates (at the resolution of the QSCD model used) is immediately available for the creation of novel molecules.
As an example, Fig. 10 shows a plot of all of the theoretical surface shapes covered by all ofthe conformations of all ofthe molecules used in the example implementation (see Fig. 5). The total volume ofthe cube in Fig. 10 encompasses all 49,268,918 theoretical surface shapes as listed in Table 1. As can be seen from the plot and two expanded points, many theoretical surface shapes are "unfilled" by the set of compounds shown in Fig. 5. Thus, in searching for molecules or libraries to enhance the diversity ofthe given set of compounds, the chemist is presented with a set of actual 3D templates into which new compound libraries may be designed. In comparison, although mapping the same set of compounds in a "non-reversible" diversity space would also display a set of coordinates to which the molecules map, there would be no way to visualize the 3D shape of any point that was not filled by one of the compounds in the set. Using BCUT values for example (Fig. 11), the coordinates specified for an unfilled point leave the chemist with a set of normalized eigenvalues. While these may give an idea of relative abundance of a given functionality (e.g. H-Bond Donor) at this point in diversity space, the coordinates give no hint of what shape or class of molecules might fill that diversity void. The above example shows how QSCD is a reversible diversity model with respect to molecular shape. Within a given surface shape in QSCD, there are many combinations of functionality leading to many different theoretical surfaces. If a given library fills only a portion of theoretical surfaces of a given surface shape, by following the same process outlined above and in Fig. 10, unfilled surfaces of specific shape and functionality may be identified and filled with complementary libraries. By using data-mining algorithms to analyze and intersect the shape and functionality of unfilled surfaces, a minimal set of "missing" 3D combinatorial templates can be deduced from the QSCD mapping of a given set of general screening compounds. These templates represent the smallest number of combinatorial syntheses which need to be executed in order to fill out the diversity of the set of screening compounds. One such template is depicted in Fig. 12. In conjunction with the efficiency of core-based combinatorial chemistry, QSCD makes possible the contemplation of a "complete" library of screening molecules at a given resolution. The model thus offers a theoretical and practical answer to the problem of generating lead structures for genomic targets of unknown structure and function. Technical details of an implementation
For the example discussed above, molecular conformations were generated with Multisearch in Sybyl (version 6.5, Tripos Ine, 1699 S. Hanley Rd., St. Louis, MO, 63144) on an R10000 Silicon Graphics workstation.
Conformations were subsequently sorted by energy and conformations within 10 kcal of the lowest energy were accepted. Overlay plots of molecules (Fig. 8B) were also generated using Sybyl.
UNITY 2D fingerprints (Unity 4.0, Tripos Ine, 1699 S. Hanley Rd., St. Louis, MO, 63144) were generated on an R10000 Silicon Graphics workstation. Pairwise Tanimoto coefficients were computed as described by Dixon and Koehler.
QSCD software for molecule quantization, mapping of Q-files, and surface complementarity display was developed using the Java programming language (JDK 1.2) and the Java3D graphics API (version 1.1) on Intel-based workstations. Theoretical target surfaces were stored and indexed using an Oracle 7.3.3 database.
Parameters for theoretical target surface generation/molecular quantization and parameters for complementarity mapping/scoring were alternately optimized in three successive rounds as described below. The parameters used for theoretical target surface generation and the closely related parameters for quantization of small molecules into quantized files (Q-files) were optimized in the context ofthe algorithms mentioned above. Parameters were iteratively optimized by varying a given parameter and then quantizing training molecules other than those in Fig. 5. Training molecules used were taken from in house structures and two published SAR sets.
Concomitant with molecular quantization, an enumerated set of theoretical target surfaces was created with coπesponding parameters. Using the cuπent optimized complementarity/scoring parameters, molecules were then mapped to theoretical target surfaces and all diversity pairing scores generated as described in the text. Parameters were chosen which accurately predicted known homogeneous/heterogeneous pairs and which maximized "signal to noise" of homogeneous scores over heterogeneous scores.
The parameters used for mapping/scoring molecular conformations to theoretical target surfaces were optimized in the context of the algorithm stated above. Parameters were iteratively optimized by varying a given parameter and then mapping a constant set of training molecules (see above) to a constant set of theoretical target surfaces, using the most cuπent surface generation and quantization parameters. Diversity pairing scores were generated for all training molecules, and parameters were chosen which accurately predicted known homogeneous heterogeneous pairs and which maximized "signal to noise" of homogeneous scores over heterogeneous scores.
As mentioned earlier, for a given conformation-to-surface shape fit to be accepted, the minimum overlap requirement was set to either 9 quanta or N-2 quanta of a conformation of N quanta. This range allows large conformations to fit partially into a theoretical surface (protruding volume must be at the mouth of the surface) while also allowing smaller conformations to be considered for complementarity. It excludes large conformations which do not overlap at least 9 quanta.
Approximate computational speeds of typical QSCD operations are as follows on a single Pentium III 500 MHz workstation: Generation ofthe basis set of theoretical target surface used in the study required 17 min.; this data was stored for access by subsequent QSCD functions. Quantization of 100 conformations of a given molecule into 100 Q-files required 250 seconds. Complementarity mapping of 100 Q-files onto the basis set of theoretical target surfaces used in the study required 40 seconds. Algorithm for Designing Molecules for Unfilled Target Surfaces
The following algorithm could be used to design novel molecules based on complementarity to unfilled theoretical target surfaces that are not complementary to any existing molecular conformations.
1) Existing molecules are quantized and those negative space cube target surfaces to which their conformations are complementary are identified, (see
Appnedixes A, B, and C)
2) For a given set of existing molecules and a desired set of theoretical target surfaces with given shapes and functionalities, those theoretical target surfaces are identified to which no existing molecular conformations are complementary. Novel molecules are designed to be complementary to the above identified theoretical target surfaces as follows: a) Let the set of those theoretical target surfaces to which no existing molecular conformations are complementary = T. b) "Cluster T into subsets Tl-Tn such that for each Ti there is a central set of negative space cubes of given functionality Ci such that all target surfaces in Ti can be created by adding up to E additional negative space cubes of given functionality to Ci at up to each of F points of attachment on Ci where each point of attachment Aj (j = 1-F) is a single face of Ci to which zero, one or a set of up to E negative space cubes may be added. (See Fig. 24) c) For each Ci, design a core molecule Mi that fills but does not extend beyond the space defined by Ci within a tolerance limit TOL, the core molecule furthermore complementary to the functionality of Ci, and the core molecule furthermore containing at least one combinatorial site (defined as a reactive site that can be further functionalized and/or extended in a combinatorial step under given chemical conditions) that can project potential combinatorial building blocks through at least one plane Aj. (See Fig. 25) d) Computationally enumerate the combinatorial library L(Mi,B) defined by Mi and a set of building blocks B that are each no larger in volume than E negative space cubes (See Fig. 26). e) Quantize nc conformations of each molecule in L(Mi,B) and determine the set W of all theoretical surfaces to which any conformation of any molecule in L(Mi,B) is complementary (see Appendixes B and C). f) Compute the set of target surfaces M = W fl Tn g) If M contains an acceptable number of novel surfaces as defined by the user, then chemically synthesize the actual library L(Mi,B). Otherwise, choose a new Mi in step 3 for the given Ci and repeat until conditions in this step are met. h) Move on to the next Ci in step c. Protein Diversity
An extension of any diversity model based on an absolute frame of reference is that the same basis set may be used to classify actual proteins. By mapping onto the QSCD basis set all surfaces of volume V of a known protein, actual proteins can be compared and classified by their 3D binding sites. In addition to providing a diversity map of known protein binding sites within the universe of all theoretical protein surfaces under given parameters, the theoretical surfaces of QSCD may thus be used to correlate protein classes to complementary molecular core structures.
Appendix E describes an algorithm for quantization of protein surfaces. Appendix F describes an algorithm for comparing protein surfaces to determine a degree of similarity or dissimilarity. The following algorithm generates a set of files T of quantized protein binding surfaces which together represent the available surface of a given protein binding site.
Algorithm T(protein binding site, H, R, V, TolA, TolB, Transinc, Rot, Rotvar, PM)
Protein binding site: see l, below
H: resolution of scanning grid in angstroms, H<=R
R: resolution of cube = length of cube side in angstroms V: total volume of each surface in cubic angstroms
TolA: tolerance (%) of atomic radii
TolB: tolerance (%) of cube volume which must intersect convex hull
Transinc: translational increment in angstroms
Rot: # rotations (odd integer) Rotvar: rotational variance (%)
PM: For each of V/R3 cubes used, any of M types of chemical functionality
1. Take a protein binding site file (obtained by any number of commercially available methods, such as running a "Connolly surface search" on a standard "PDB file" (Michael L. Connolly, 1259 El Camino Real, #184, Menlo Park,
CN 94025) which minimally consists of: a) A calculated probable electron density surface ofthe binding site b) A list of all known atom types in the molecule with their coordinates and atomic radii c) A list of known connectivities of all atoms with the type of bond connecting each atom
2. Overlay flexible 2D square grid(s) of grid size HxH angstroms over all calculated probable electron density surfaces in the protein binding site
3. Calculate the convex hull of the set of points defined by the points of each grid nexus 4. Examine each grid nexus one at a time
5. At a given grid nexus place a cube with the center of one face tangent to the probable electron density surface ofthe protein binding site 6. If the cube contains a protein atom coordinate or any atomic radii of protein atoms protrude into the trial cube by more than Tol % of their atomic radius, then step 7 '., otherwise step 8.
7. Translate the cube Transinc angstroms away from the grid nexus on the probable electron density surface while keeping the center ofthe cube on a line perpendicular to the probable electron density surface at the the grid nexus. If the cube is now R angstroms away from the grid nexus, take the next nexus in 4., otherwise return to step 6.
8. Designate the cube as being a set cube with 5 unchecked faces and 1 checked face (the checked face being that which faces the grid nexus being examined). Designate the initial frame of reference to be the X, Y, and Z axes co-linear with the sides of the cube.
9. Rotate the unchecked cube about its center in all combinations (total of Rot combinations) of the following units: a) X rotations: any one of Rot units (degrees) from - Rotvar * 90/2 to
+Rotvar * 90/2 by Rotvar * 90/Rot b) Y rotations: any one of Rot units (degrees) from - Rotvar * 90/2 to +Rotvar * 90/2 by Rotvar * 90/Rot c) Z rotations: any one of Rot units (degrees) from - Rotvar * 90/2 to +Rotvar * 90/2 by Rotvar * 90/Rot
10. For a given rotated system from 9., place face contiguous trial cubes of side length R at all unchecked cube face(s), not to exceed the nearest integer to V/R3 total cubes. If addition of a new trial cube would exceed the nearest integer to V/R3 total cubes, proceed directly to 1 1 without adding further trial cubes
11. All unchecked cube faces (not including cube faces on trial cubes) become checked cube faces.
12. If a trial cube contains a protein atom coordinate, or any atomic radii of protein atoms protrude into the trial cube by more than TolA of their atomic radius, or if the cube does not intersect with a volume ofthe convex hull equal to at least R3 * TolB (see 3. above), then the trial cube is removed. Otherwise the trial cube becomes a set cube (with unchecked faces).
13. If the total number of set cubes becomes = the nearest integer to V/R3 (yielding a resulting cube combination of the nearest integer to V/R3 cubes), then go to step 15. 14. If no trial cubes in step 12 became set cubes and there are no unchecked cube faces remaining then start with the next rotated cube in step 9. If no rotations produce resulting cube combinations in step 13. then start with the next translation in step 7.
15. Ignore any remaining rotations from step 9. Designate all cubes as negative space cubes fully enclosed except as detailed below. Designate the layer of cubes which is a) peφendicular to the line peφendicular to the probable electron density surface at the grid nexus being examined b) farthest from the grid nexus as negative space cubes which are open at their faces farthest from the grid nexus and peφendicular to the line peφendicular to the probable electron density surface at the grid nexus.
16. Based on the types of proximal atoms in the protein suπounding each negative space cube, and based upon the bonds which these proximal atoms form and the other atoms to which these proximal atoms are bonded, ascribe to each negative space cube one and only one of M types of chemical functionality.
Types of functionality M may include but are not limited to:
Acidic regions,
Basic regions, regions of formal charge +1, regions of formal charge -1, regions of partial charge between +0.5 and +1 , regions of partial charge between -0.5 and -1, regions of partial charge between 0 and +0.5, regions of partial charge between 0 and -0.5, hydrophobic regions, polarizable regions, hydrogen bond donating regions hydrogen bond accepting regions hydrogen bond donating/accepting regions 17. Yield a "quantized" protein binding surface in terms ofthe nearest integer to V/R negative space cubes of side length R with any one of M types of functionality per negative space cube.
18. Return to step 4. for each grid nexus
19. Compare all quantized surfaces from 17. and remove any which are identical 20. Yield a set of files T of quantized protein binding surfaces which together represent the available surfaces ofthe given protein binding site
Comparing Molecules to Proteins
Appendix G describes an algorithm for determining the complementarity of a library of molecules to a set of protein surfaces.
Algorithm for Designing Molecules for a Set of Protein Target Surfaces In another procedure, novel molecules (for testing as ligands for proteins) can be designed based on complementarity to negative space cube targets to which a set of protein pockets map. The following outline describes the steps for doing so:
1) A set of existing protein pockets is quantized and those negative space cube target surfaces to which their quantizations map are identified. (See Appendixes B and C.)
2) Novel molecules are designed to be complementary to the above identified set of target surfaces as follows: a) Let the set of target surfaces = T. b) Cluster T into subsets TI -Tn such that for each Ti there is a central set of negative space cubes of given functionality Ci such that all target surfaces in Ti can be created by adding up to E additional negative space cubes of given functionality to Ci at up to each of F points of attachment on Ci where each point of attachment Aj (j = 1-F) is a single face of Ci to which zero, one or a set of up to E negative space cubes may be added. (See Fig. 27) c) For each Ci, design a core molecule Mi that fills but does not extend beyond the space defined by Ci within a tolerance limit TOL, the core molecule furthermore complementary to the functionality of Ci, and the core molecule furthermore containing at least one combinatorial site
(defined as a reactive site that can be further functionalized and/or extended in a combinatorial step under given chemical conditions) that can project potential combinatorial building blocks through at least one plane Aj. (See Fig. 28) d) Computationally enumerate the combinatorial library L(Mi,B) defined by Mi and a set of building blocks B that are each no larger in volume than E negative space cubes (See Fig. 29). e) Quantize nc conformations of each molecule in L(Mi,B) and determine the set W of all theoretical surfaces to which any conformation of any molecule in L(Mi,B) is complementary (See Appendixes B and C.) f) Compute the set of target surfaces M = W fl Tn g) If M contains an acceptable number of novel surfaces as defined by the user, then chemically synthesize the actual library L(Mi,B). Otherwise, choose a new Mi in step 3 for the given Ci and repeat until conditions in step 7 are met. h) Move on to the next Ci in step 3.
Parameter Values
Appendix H contains example parameter values useful in connection with the algorithms described in Appendixes A through G.
Further Extensions
As mentioned, a 4.24 A cube was found to be the largest predictive unit size of diversity measure for our criteria of designing general screening libraries. For example, both 4.48 and 4.00 A units gave poorer prediction of homogeneous/heterogeneous pairs than the pairings of Fig. 6 (4.24 A units). This is likely due to the fact that most organic small molecules are themselves quantized by a limited basis set: the VDW radii of H, C, N, O and a few other atoms (see for example Fig.l). If there is no constraint on size of cubic units, however (i.e., if there is no attempt to maximize orthogonality of theoretical target surfaces), other unit measures of diversity can be found. A unit of 2.12 A should also provide effective diversity information but at a much higher resolution. Such a "high-resolution" adaptation of QSCD brings with it numerical (and thus computational) challenges. 112 negative space cubes (14 x 8) are now required at the upper limit of theoretical target surface size, translating to exponentially greater numbers of theoretical target surfaces and, depending on the stringency of fitting parameters, coπespondingly greater numbers of surface fits per molecule as in Table 4. At this resolution, the assumption of no occlusions in theoretical target surfaces becomes less valid, and removal of this assumption increases computational complexity further. A corollary of any absolute diversity model is a prediction ofthe total
"size" of diversity space in terms of unique molecular points. In other words, what is the minimum set of molecules needed to fully cover a given diversity space. This calculation is dependent on two factors: the resolution stipulated in the model (e.g., what amount of molecular change is recognized as different) and the maximum values of each dimension ofthe model's basis axes. In the model of QSCD discussed above, resolution is fixed by cubic units of 4.24 A, and maximum values are fixed at 14 units (molecular volume of 1070 cubic A) and 4 points of 7 types of molecular property characteristics. As describe above, the result is a set of 1.1 * 1014 unique molecular points. Since, using the parameters of this study, an average molecule covers 4.6 million of the unique molecular points bounded by QSCD space (Table 4), the model predicts a minimum 1.1 * 1014 / 4.6 * 106 = 24 million molecules would be necessary to completely cover diversity space.
We estimate that an average complementary molecule in the context of the QSCD model has a ΔG of complementarity on the order of -11 kcal (Table 7).
Table 7. Summation of binding energies for an interaction of an average complementary molecule/theoretical target surface pair in the context ofthe QSCD model used herein. An average molecule is assumed to have a buried volume of 12 cubic quanta (= 915 cubic A at 4.24 A resolution), 36 exposed faces (4.24 A square), 21 non-polar exposed faces (60%), 10 rotatable bonds, 4 points of complementary electrostatic/VDW potential, and a conformational energy within 2 kcal/mol of ground state. An average complementary theoretical target surface is also assumed to have 60% non-polar exposed faces. Constants used in the table are taken from Ajay and Murcko.
Energetic contribution Average ΔG translational/vibrational entropic loss (constant) +9 kcal/mol constant +0.7 kcal/mol (RTln3) per rotatable +7 kcal/mol bond; assume rigid theoretical target surface
ΔΔG conformation from ground state +2 kcal/mol constant -0.03 kcal/mol per square A non-polar -23 kcal/mol buried surface = -0.54 kcal/mol per 4.24 A square non-polar buried face; total 21 molecular faces + 21 theoretical surface faces total interaction from Table 2 -6.0 kcal/mol (four complementary points)
sum of binding energies -11 kcal/mol binding affinity to nearest integer ( ÷1.363 ) 10"8 = 10 nanomolar
In other words, the resolution used to calculate diversity translates roughly to nanomolar binding conditions for an average molecule/target surface pair. Given that some 24 million molecules are needed to completely cover diversity space under these conditions, a general screening library guaranteed to contain at least one nanomolar binder to any given target of interest would thus number at least 24 million molecules. This is a large number and will be attenuated by the fact that some molecules have significantly more than 100 conformations available to them. However, the QSCD model suggests that if, in the near future, combinatorial chemistry and high-throughput screening are to generate initial hits primarily in the nanomolar rather than micromolar range, then the field must continue to focus its efforts on the development of numerically competent synthesis and screening technologies. By defining molecules in terms of their complementarity to a fully enumerated set of theoretical protein surfaces under given parameters, and by defining actual protein surfaces in terms of their similarity to the same set of theoretical protein surfaces, the model allows: 1. Numerical prediction of protein surface diversity (how many different types of possible protein surfaces exist under a given set of parameters).
2. Numerical prediction of how many and what type of molecules would be necessary to create a "universal molecular library" (a library which contains at least one complement [of a given score] to any given protein surface).
3. Comparison of the similarity or difference of molecules or sets of molecules based on their complementarity to theoretical protein surfaces. The frame of reference for such comparisons is fixed no matter how many or what types of molecules are involved.
4. Numerical prediction of the percent of protein surface diversity to which a given set of molecules is complementary.
5. Comparison ofthe similarity or difference of actual protein surfaces or sets of surfaces based on their similarity to theoretical protein surfaces. The frame of reference for such comparisons is fixed no matter how many or what types of protein surfaces are involved.
6. Numerical prediction of the percent of protein surface diversity to which a given set of actual protein surfaces is similar. 7. Prediction ofthe actual protein surfaces to which a given molecule is complementary. 8. Prediction of how many and what type of molecules would be necessary to create a "universal molecular library" against a given set of actual protein surfaces (a library which contains at least one complement [of a given score] to each actual protein surface).
This application discloses information from an article which has been accepted for publication in the peer-reviewed Journal of Medicinal Chemistry and is cuπently scheduled for publication in the May 18th issue. The article, listed by the Journal of Medicinal Chemistry as JM990504B, is titled "Quantized surface complementarity diversity (QSCD): A model based on small molecule- target complementarity," and is incoφorated herein by reference.
Other implementations are within the scope ofthe following claims.
Appendix A
Theoretical Surfaces
A surface opening O is a set of lattice squares in 7? denoted by their corners
0 = {(xt, yt) € Z2, ι = l .. . m}
By definition, a surface opening must be connected. For connectivity purposes, two points p, q € Z2 are neighbors iff
Figure imgf000042_0001
The area of a surface opening is denned to be the number of lattice squares it contains. Surface openings are considered to be unique upto translations and rotations ofthe x-y plane.
A surface shape 5 is a set of "negative space" cubes represented as lattice cubes in Z3 denoted by their comers:
S = {(x., 2/,, 2.) € l?, ι = l . .. n}
By definition, a surface shape must satisfy the following conditions-
• All cubes are below the x-y plane. That is, if (x, y , z ) € S, then z < 0.
• The surface shape is connected. For connectivity purposes, two cubes p, q € Z3 are neighbors iff
Figure imgf000042_0002
• There are no occlusions along the s-axis. That is, if (x, y, z) & S and z < 0, then (x. y, z + l) € S.
The volume of a surface shape is defined to be the number of lattice cubes it contains. Surface shapes are considered to be unique up to translations and rotations of the x-y plane.
Every surface shape specifies a unique opening: openιng(S) = {(x. y)\(x, y, z) € 5} APPENDIX A. THEOREΗCAL SURFACES
Further, given a surface opening O and a function d : O — > N, a unique surface shape with the given opening is defined: shape(0, d) = {(x, y, z)\(x,y) e 0, d(x, y) > -z)
The function d specifies the "depth" ofthe surface shape at each opening point. Sample surface shapes and openings can be seen in Figure 13.
A theoretical surface consists of a surface shape where the cubes in the surface shape are each associated with functionality. The set T of seven specific types of characteristic functionality is used:
• T\ : Potential Negative Charge
• ?ϊ. Potential Positive Charge
• T?,: Hydrogen Bond Donor/ Acceptor
• '- Hydrogen Bond Donor
• $: Hydrogen Bond Acceptor
• T§: Polarizable
• Hydrophobic
%: Slightly Hydrophobic
A functionality map / : S — » defines the assignment. By default, all cubes are assigned functionality %, and all possibilties are considered where upto n/ of the cubes are given one ofthe functionalities T\- .
Generating all surface shapes is accomplished by the following steps:
1. Generate all surface openings (S URFACEOPENINGS).
2. Filter the surface openings to remove openings that are unlikely to resemble surfaces found in nature (OPENINGFILTER).
3. From the set of filtered openings, generate all surfaces shapes whose associated openings are in the set (S URFACESHAPES).
4. Filter the set of surfaces shapes to remove shapes unlikely to resemble surfaces found in nature (SHAPEFILTER).
5. From the set of filtered surface shapes, add all possible combinations of functionality to define a set of theoretical surfaces (FUNCTIONA IZESURFACE).
APPENDIX A THEORETICAL SURFACES
Algorithm A.l SURFACEθPENlNGS(A). calculate all surface openings with area less than or equal to A
1 define Nι to be a set containing the only the surface openmg with a smgle square at (0, 0)
2 for i <- 2 to A do contain all openings of area t}
Figure imgf000044_0001
5 define T5 to the set of all possible openings obtamed by addmg a single square to O adjacent to a square already present in O
6 for all P € V do
7 if no rotation or translation of P is present in N then
8 add P to
9 end if ιo end for
11 end for
Figure imgf000044_0002
APPENDIX A. THEORETICAL SURFACES
Algorithm A.2 θPENlNGFl TER(C, Λt, Mnc, Mc) filter a set of surface openings O using area-threshold parameter At, max-non-central parameter Mnc, and max- contiguous parameter Mc.
1 O <- 0
2 for all O e O do has 2 or more
Figure imgf000045_0001
12 end if blocks of lattice squares in 0}
Figure imgf000045_0002
15 if (x + 1, y) e O and (x, y + 1) € O and (x + 1, y + 1) e O then
16 remove (x. y), (x + l, y), (x. y + l), (x + 1 y + 1) from O
17 end if
18 end for
19 define c to be the number of squares in the largest connected component left in 0
20 if area(O) < /nc and (area(O) < At or c < Mc) then
21 add O to 0
22 end if
23 end for
24 return (0)
APPENDIX A. THEOREΗCAL SURFACES
Algorithm A.3 SURFACESHAPES(C, V). calculate all surface shapes with openings in the set O and volume less than or equal to V.
1 for all O € O do
2 define d : O → N such that d(x, y ) = 1
3 So «— 0 {<So will contain all surfaces with openmg 0}
4 v <— area(O)
5 while v < V do
6 5 «- shape(0,d)
7 if .So contains no translations or rotations of S then
8 add S to S0
9 end if
10 for x < area(O) to area(O), y < area(O) to area(O) do
11 if (x,y) 6 O and v + 1 < V then
12 u *— υ + l,d(x,y) <— d(x,y) + 1
13 goto step s
14 else
15 υ*—υ~d(x,y) + l.d(x,y)*—l
16 end if
17 end for
18 goto step l {no more surfaces left with this opening }
19 end while
20 end for
21 U^\J0e0S0
22 return (U)
Algorithm A.4 SHAPEFILTER(IS, Λ/e): filter a set of surfaces <S using max-extrusion parameter Λ/e
1 <-0
2 for all 5 € 5 do
3 O <— openιng(S)
4 d <— depth(S) {corresponding depth function}
5 for all (x,y) € O do
6 forall (x.y) € {(x - l.y), (x + l.y), (x,y - l),(x y + l)}do
7 if (x, y) € O and d(x, y) > d(x, y) - Me then
8 goto step 5
9 end if
10 end for
11 goto step 2 {surface does not qualify }
12 end for
13 add 5 to S {surface does qualify}
14 end for
15 return («S) APPENDIX A. THEORETICAL SURFACES
S <π0{set of functionalized surfaces_}_ define V to be the set of all the (vola^s)) selections of nf cubes from S for all V e V do define / : S → T such that f(s) = FΆ if s V, f(s) = i for s € V number the elements of V as vι , .. . , υnf loop add (S, f) to S fo
Figure imgf000047_0001
ass gn vi to e the next higher functionality goto step 6 else
Figure imgf000047_0002
end if end for goto step 3 {no more functionality assigments left with the set V of cubes} end loop end for return (S)
Appendix B
Quantization
In the quantization process a molecule is reduced into a representation such that complementarity can be calculated against a set of theoretical target surfaces. Quantization takes place in the following steps-
1. Each atom in the molecule is assigned a functionality based on its type and connectivity
2 Three dimensional conformations of the molecular structure are generated.
3 Each conformation is converted mto "positive space" cubes based on the positions of its atoms.
4. Each cube is assigned a functionality based on the atoms that it contains
Finally, the quantized form of the molecule is defined to be the set of all selected conformation quantizations with functionalities assigned The quantization process is summarized in Figure 14
APPENDIX B QUANTIZAΗON
B.l Atomic Functionality Assignment
Given a molecular structure as a set of atoms M, a map / M — T is defined assigning functionalities to all of the atoms The map is defined by the using a set of rules to match molecular substructures based on extended atom types (as generated by Tπpos, for example) and bondmg patems that encapsulate each functionality type The algoπthm keeps track of atoms that are excluded from matching a lower pπoπty rule because they have already been matched in a higher pπoπty rule, where T\ has the highest pπonty and Tη the lowest. Atoms not matching any rule are assigned functionality Tγ. No atoms are assigned functionality T% The functionality rules used can be seen as follows:
• T\ Figure 16
• Ti. Figure 17
• z Figure 18
• T4 Figure 19
• T5. Figure 20
• T&: Figure 21
Within each functionality type, the functionality rules are search sequentially in the order listed in Figures 16-21. Figure 15 provides a legend for the symbols used in the functionality rule diagrams
Algorithm B.l ATOMFUNCTIONAL AP(Λ^ ) assign functionality to the atoms in molecular structure -Vf and determine which atoms are excluded from quantizafton
1 c *— 0 {£c is the set of atoms that are excluded from consideration. }
2 q «— 0 [εq is the set of atoms that are excluded from the rest of the quantization process.}
3 define / M — ' T such that f(a) = Tγ for all atoms a £ M
4 for i <— 1 to 6 do
5 for each functionality rule for Tz considered in the proper order do
6 find all matches m M with consideration to £c
7 according to the rule, add appropnate atoms to £ and εη, and set / to T, for selected atoms
8 end for
9 end for
10 return (/, £,)
APPENDIX B. QUANTIZAΗON
B.2 Conformation Generation
A conformation is a mapping c M. — ♦ E3 of the molecular structure into three dimensional space. Using OpenEye Omega software (Open Eye Scientific Software Inc., 335c Winische Way, Santa Fe, NM, 87501), upto nc representation conformations for each molecule are generated within given rule-based energy parameters.
APPENDIX B. QUANTIZAΗON
B.3 Base Coordinate Frame
A coordinate frame T = (R. t) I3 → R3 is a πgid motion ofthe space defined by a rotation R S S03(R) and a translation t 6 R3 that transforms a point p by the rule
T(p) = Rp + t
Given a resolution r (the length of a side of a lattice cube) and a coordinate frame, a lattice on R3 is implicitly defined by q £ Z3 <→ [rς^. rg + r) x [rςy, rgy + r) x [rςr2, rg2 + r) C R3
The base coordinate frame is generated from a conformation of molecular structure M. First, a subset . of the atoms in M. are selected via F RAMEATO S These are atoms m or near πng structures close to the center of the conformation A πng atom is defined to be an atom that contains at least one bond which, if removed, would not result in the molecule being disconnected If there are an insufficient number of πng atoms, all atoms sufficently close to the center ofthe conformation are used
Then, the base coordinate frame is calculated in B ASEFRAME. The x-axis of the base coordinate frame is defined to to be the solution to the optimization problem max \xτ(c(a) — p)\ ζM where
P — —=- ^ (α) \M\ ^-1 is the center of the conformation. Due to the non-linear nature of this problem, it is solved using an approximate gradient descent method. Given x, the y-axis ofthe base coordinate frame is defined to be the solution of a second optimization problem
, ax. ∑ \yτ{c{a) - p)\
||y|| = l xry=0 ^ α€Λ-t
Given x and y, z is uniquely specified. The base coordinate frame then simply involves centeπng the conformation by translating p to the origin, and using the new x, y, and z
APPENDIX B. QUANTIZAΗON
Algorithm B.2 FRAMEAT0MS(J . c, 7 , rm, qτ select a set of atoms to use for generating the coordmate frame from a molecular structure M. and a conformation c. The following parameters are used, πng factor 7 , ring minimum rm, radius factor qτ.
1 71 <— 0 {71 will hold the set of πng atoms }
2 for all α € Λ do
3 if α is not a Hydrogen atom and α is m a πng then
4 add α to 72.
5 end if
6 end for
7 P *~ TM\ ∑αeΛi c(α) I*6 mathematical center of the conformation }
8 r <— ma α€Λi llc(α) p\\ {the maximum distance from the center of the conformation}
9 C <— 0 {a placeholder for atoms in rings we have already examined }
10 ig <— 0 {a set of nngs which are close to the center of the conformation }
11 for all o € 71 do
12 if \\c(a) — p\\ < rqr and a C then {if a πng atom is sufficiently close to the center of the conformation, add all members of its ring group }
13 g <— RlNGGROUP(a, 72., r/)
14 add Q to Tlg
15 C i- c u g
16 end if
17 end for
18 V <— {α € M such that||c(α) — p\\ < rqτ} {a default set of atoms close to the center, to be used if we haven't collected enough ring atoms }
19 if TZg = 0 then {if there are no πng groups, use the default }
20 return (V)
21 end if
22 if m xgg-ft, \Q\ < rm then {if the biggest πng group is smaller that rm, use the default}
23 return (2?)
24 end if
25 Bg <— UρeTC ιsi >6 ^ {a set °f "D1g" πng groups having at least 6 atoms }
26 M <- 0
27 for all a € UgeB ^ t'0 {use a^ "^lS" nnS group atoms and any neighboπng atoms}
28 add a and all atoms bonded to α to M, if not already present
29 end for
30 return (M) APPENDD B QUANTIZAΗON
Algorithm B.3 RlNGGRθUP(α TZ, 77) calculate the set of atoms that can be reached starting at atom a and crossmg over at most 77 atoms that are not in the set of πng atoms TZ
1 d«-0
2 Th «- {a}, Tn «- 0
3 ε «- {a}, 9 «- {a}
4 while d < 77 do
5 ifTΛ = 0 hen
6 d<-d+l
Figure imgf000053_0001
8 else
9 ά<-pop(7^)
10 if ά is not a Hydrogen atom and ά i ε then
11 for all atoms 0 bonded to ά do
12 if 0 is not a Hydrogen atom and 0 £ then
13 add 0 to £
14 ifo € 72. then
15 add 0 to .7
16 add olo h
17 else
18 add 0 to Tn
19 end if
20 end if
21 end for
22 end if
23 end if
24 end while
25 return (Q)
APPENDIX B. QUANTIZAΗON
Figure imgf000054_0001
5 v «- u(α)/||«(o)||
6 loop
7 * *- ∑o M sιgn(vτu(o))u(o)
8 ««-a/||β|| _
9 for all o € M do
10 if sιgn(sτu(o)) φ 0 and sιgn(sτu(o)) Φ sιgn(wτu(o)) then
11 v *— s
12 goto step 6
13 end if
14 end for
15 v <— s
16 goto step 18
17 end loop
18 if Σαe, \vT )\ > ΣaeM |rff«(o)l hen
19 di «— υ
20 end if
21 end for
22 da -(0,0.0)
23 for all α € Λ4 do
24 υ *— t x (u(a) — d[u(a)d )
for all o e M such that sign (sτ (o)) φ 0
Figure imgf000054_0002
29 'f ΣaeM lsT"(°)l > ∑aeM I «(")l then
30 a <— s
31 end if
32 end if
33 end for
34 R <— [dι,d2, ! x d2]T
35 return (R, -Rp) APPENDLXB. QUANTIZAΗON
B.4 Lattice Centering
By our definition above, the lattice defined by a coordinate frame places corners of cubes at points all of whose coordinates are integer multiples of r. Given a particular conformation, it may be better to shift the lattice by a length of τ/2 in a particular direction, recentering the lattice cubes.
Algorithm B.5 CENTERFRAME(T, M, c, Cf, ct): recenter the coordinate frame T for the set of atoms M and conformation c using parameters: centering fraction c/, resolution r, centering tolerance c*.
I <- \cf\M\] set ( to be the .th lowest element in the set {|T(c(α)) |, α € M} set yι to be the /th lowest element in the set {\T(c(a))y\, a € M} set zι to be the Zth lowest element in the set {|T(c(α))r|, a € M}
Λ ^ {a e M such that \T(c(a))x\ < xι, \T(c(a))y\ < yι, |T(c(α))z < } define
Figure imgf000055_0001
for i ; +— 1 to 8 do define the coordinate frame Ti(p) = p + vt
Qi <- Cυmτγ(M, c, Ti θ T,r, ct)
Figure imgf000055_0002
end for define j to be the index of the least member of the set { (n{ , dz , dVti , dx.i }, where comparisons are done using a dictionary ordering (that is, compare the first component, if case of equality compare the second component, etc.)
16: return (Tj)
APPENDIX B. QUANTIZAΗON
B.5 Lattice Perturbations
The base coordinate frame is not necessanly optimal for quantization, so a set of "close*' frames are also exammed. By parameterizing the space of coordmate frames S03(R) x R3 into 3 rotational dimensions and 3 translational dimensions and taking nτ equal steps in each rotational dimension and nt equal steps m each translational dimension, a set of n n perturbation frames is built. These can later be composed with the base coordinate frame to give the desired set.
Algorithm B.6 PERTURBATIONFRAMES(Π4, υt, nr, υr)- generate a set of perturbatton coordmates frames using parameters: resolution r, number of translations nt, transla¬
_ ti_on_al_ v_aπ_ance vt, number of rotations nr, rotational vaπance vr
2 for ιx «— 0 to nr — 1 do
3 define Rx to be the coordmate frame corresponding to rotation about the x-axis by πvr(2ιx + 1 — nr)/(4nr) radians
4 for iy *— 0 to nτ — 1 do
5 define Ry to be the coordmate frame corresponding to rotation about the y- axis by πvr(2ιy + 1 — nr)/(4nr) radians
6 for ιz *— 0 to nr — 1 do
7 define Rz to be the coordmate frame corresponding to rotation about the --axis by πvr(2ιz + 1 — nr)/(4nr) radians
8 for x <— 0 to nt — 1 do
9 define the coordinate frame
Tx(p) = p + (rvt(2Jx + 1 - nr)/(2nP).0.0)
10 for jy <— 0 to nt — 1 do
11 define the coordinate frame
Ty(p) = p + (Q, rvt (2jy + 1 - nr)/(2nr), 0)
12 for jz *— 0 to nt - 1 do
13 define the coordinate frame rB(p) = p + (0, 0. r.t(2js + l - nr)/(2nr))
14 add Tz o Ty o Tx o Rz o Ry θ Rx to T
15 end for
16 end for
17 end for
18 end for
19 end for
20 end for
21 return (T) APPENDIX B. QUANTIZAΗON
B.6 Cubification
Given a conformation and a lattice (defined by a coordinate frame and resolution), cubification is the process of determining which lattice cubes are filled by the conformation. The set of lattice cubes is constructed by taking any cube in which an atom center in the conformation directly falls and also cubes which are sufficiently close to the van Der Waals sphere of an atom.
Algorithm B.7 CUBIFY(- , c, T, r, t): quantize the conformation c of the molecular structure M into a set of cubes defined on the coordinate frame T usmg parameters: coordinate frame T, resolution r, tolerance t.
Q <— ® {Q will contain the set of cubes in the quantization }
2 define a function d : M —* R that takes each atom to its van Der Waals radius
3 for all a _ M do
4 p <— T(c(α)){the transformed position}
5 Q «— ( 1-Pι/rJ , -P-JrJ , LP*/ΓJ ) {the integer coordinates of the cube into which the atom falls }
6 if q Q then
7 add q to Q
8 end if
9 define Q to be the set of cubes neighboring q:
Q +- {q € Z3 such that τaax. (\qx - qs\, \qy - qv\, \qz - qz \) = l }
10 for all q € Q do
11 define d to be the distance of the point in R3 m the cube defined by q that is closest to p
Figure imgf000057_0001
12 if d - d(α) < tr and q then
13 add q to Q {add a neighboπng cube if it is too close to the van Der Waals sphere of the atom}
14 end if
15 end for
16 end for
17 return (Q)
APPENDIX B QUANTIZAΗON
B.7 Entire Process
Given a set of conformations, m Q UANTIZE we see the entire quantization algoπthm For each conformation
1 a base frame and centering frame are calculated
2 perturbations of the base frame are used in order to find the lattice that results first in the smallest number of cubes and second with the least distance from the base frame
3 using the atomic functionality map, functionalities are assigned to the cubes in the minimal quantization
APPENDIX B. QUANTIZAΗON
Algorithm B.8 QυAm\ZΕ.(M,C,r,t,rf,rm,qr,Cf,Ct,ntt,nrr) quantize the set of conformations C for the molecular structure M with parameters resolution r, tolerance t, nng factor 77, πng minimum rm, radius factor qr, centeπng fraction c/, centeπng tolerance ct, number of translations nt, translational vanance vt, number of rotattons nτ, rotational vaπance υr, polaπzable minimum pm.
1 (fs£q) *- ATOMFUNCTIONMAP(JM) {build a functionality map and a list of atoms to be excluded from quantization }
2 S <— 0 {will hold the resulting quantizations }
3 for all c e C do
4 M <— FRAMEATOMS(. -εq,c,rf,rm,qr)
5 Tb <— BASEFRAME(J , c) {base coordmate frame }
6 Tc <- CENTERFRAME(T, M - εq, c, cf,ct) {lattice centering }
7 qm <— oo{smallest number of cubes seen so far }
8 dm *— 00 {smallest transform distance seen so far }
9 7" <— PERTURBATlθNFRAMES(nt, υt, τιr, ur) {set of perturbation frames }
10 for all Tp € T do
11 T <- Tc o Tp o T6 {the total coordmate transform }
12 Q <— CUBIFY(Λ1 - εq,c,T,τ,t) {cubes given this transform}
13 q *— I Q| {number of cubes}
1 d <- Σo6Λ(-ε, IITP(C(°)) - c(o) II {transform distance}
15 if |Q| <9mor(|Q|-gmandd< m)then
16 <?m<-<7,dm<-d,em<-Q,Tm-r
17 end if
18 end for
19 define /_ Sm→J as /m(ς) = T7
20 for all α € - - £, do
21 q - (LTm(c(α))_/rJ, LTm(c( ))y/rJ, \Tm(c(a))z/r\) {q is the the cube α falls into}
22 if f(a) has higher pπoπty than fm{q) then
23 if f(a) φ T6 then fm(q) - /(α)
25 else if at least pm atoms with T functionality are in q then
26 /m(?) - /(α)
27 end if
28 end if
29 end for
30 add(Qm,/m)to5
31 end for
32 return (<S) Appendix C
Surface Complementarity
In order to view molecules in the space of theoretical surfaces, we must establish complementarity between quantized conformations and theoretical surfaces. This is done in FITSURFACES as follows:
1. all 24 possible orientations of each quantizaed conformation are considered
2. for each orientation, the quantized conformation is shifted down and below the plane
3. a set of surfaces which fit each shifted conformation are detected
4. functionalities for each surface are computed such that the binding energy between the surface and the quantized conformation are favorable
APPENDIX C SURFACE COMPLEMENTARITY
Algorithm Cl FITSURFACES( , /,, Ec,Tb, ) calculate all surfaces with functionality that are complementary to the quantizated conformation Q with functionality map /,, conformational energy Ec, and rt, rotatable bonds using the followmg parameters minimum surface openmg area A, maximum surface volume V, area- threshold At, max-non-central Mnc, max-contiguous Mc, max-extrustion Me, number of points of characteπstic functionality /, minimum energy £—,„, minimum fit quanta qmtn, minimum slackness smin, maximum slackness smol, maximum protrusion levels pmax, translational-rotational-vibrational entropy EtrVs rotatable bond coefficient cr, hydrophobic energy coefficient Ch, hydrophobic surface energy coefficient cs, potential function P.
1 ιS <— 0{wιll hold the resulting surfaces }
2 define TZ to be the set of 24 lattice rotations
3 for all R e TZ do
4 Q <— {Rq,q £ Q} {rotate the quantization by R}
5 tx +- (max_eQ qx + mm_efi qx)/2\
6 ty *- [{ma.xqeQ qy + mιn-6Q ςy)/2j
7 t.+-mmqeQqz
8 Q*- {{qx - tx,qy - ty,qz - tz),q e Q] {recenter over the x-y plane}
9 for all d +— 0 to max 6C qz do
10 Qd <— {{qx,qy z — d),q £ Q such that qz < d} {shift the quantization down and only keep cubes on or under the x-y plane} U if |Qdl < πnn(<j— ,„, |Q|) — smιn then {skip if not enough cubes are below the plane}
12 goto step 9
13 end if
14 if max eo 9* > Pm i then {skip if too many layers of cubes are stickmg out of the plane }
15 goto step 9
16 end if
17 Sd <- CθRESURFACE( d)
18 Ad <- m (smax + \Sd\, A),Vd *- mιn(smαl + |5d| V)
19 for all S € DETECTSURFACES(5d, Ad, Vd, At.Mnc, Mc, Me) do
20 for all (5/ /-) £ FUNCTIONALIZESURFACE(5.TI/) do
21 ifENERGY(S ,/s,Q,,/,,£c,'r6,r,£trt.cr ch,cs,P)
22 if no translation or rotation of (5/, /-) is in <S then
23 add(S ,/-) toS
24 end if
25 end if
26 end for 2 end for
28 end for
29 end for
30 return (S) APPENDIX C. SURFACE COMPLEMENTARITY
Algorithm C.2 CORESURFACE( Q): calculate the minimal surface fitting a quantized conformation.
O «— ©{surface opening} for all σ € do add (qx, qy) to O, if not already present end for define a depth function d : O — ► N such that d(x, y) — 1 for all q € Q do
<*(«*, g <- max(d(gx, 9W), 1 - qz) end for return (shape(C d))
APPENDIX C SURFACE COMPLEMENTARITY
Algoπthm C.3 DETECTSURFACES(SC, A V. At, Mnc, Mc, Me) detect additional surfaces by adding cubes to 5- subject to parameters minimum surface opening area A, maximum surface volume V, area-threshold At, max-non-central Mnc, max- contiguous Λ/c, max-extrustion Me
1 S *— 0{the detected surfaces }
2 Sn «- {Sc}
3 while Sn Φ 0 do
4 Sm <- pop(<Sn), O «- openrng(Sm ), dm <- depth(S)
5 if O is connected and area(O) < A then {add surfaces obtamed by addmg cubes below Sm)
6 d <— dm, v <— volume(S-)
7 while t < V do
8 S <— shape(0, d)
9 add 5 to 5
10 for x < area(O) to area(O), y * area(O) to area(O) do i i if (x, y) 6 O and υ + 1 < V then
12 υ «— υ + l, d(x, y) *— d(x, y) + 1
13 goto step 7
14 else
15 v *- υ - d(x, y) + dm(x, y), d(x, y) *- dm(x, y)
16 end if
17 end for
18 goto step 2θ{no more surfaces left }
19 end while
20 end if
21 if area(O) < A and volume(Sm) < V then {consider surfaces made by enlarging the opening}
22 define V to be the set of all possible openmgs obtained by adding a single square to O adjacent to a square already present m O
23 for all P e V do
24 define dp P — » N such that dp(x y) = dm{x, y) for (x y) € O and dp(x, y) = 1 otherwise
25 add shape(P, dp) to Sn
26 end for
27 end if
28 end while
29 0 <- OPENlNGFlLTER({openmg(S), S € S}. At, Mnc, Mc)
30 for all S e S do {filter surfaces based on openings}
31 if openιng(S) t O then
32 delete S from S
33 end if
34 end for
35 S «— SHAPEFILTER(5, e){filter surfaces based on shape}
36 return (S) APPENDIX C SURFACE COMPLEMENTARITY
Cl Energy
Complementarity energy between a quantization of a molecular conformation and a a cubic theoretical surface is the sum of several components:
• translational-vibration-rotational entropy (a constant)
• a conformational energy term, representing the energy of this conformation relative to the minimal energy conformation (calculated by the conformational generator)
• a term proportional to the number of rotatable bonds
• potential energy, the sum of energy due to interactions between the functionalities of overlapping negative space surface cubes and positive space quantization cubes (represented as a function P : T x T → R U {— oo})
• hydrophobic energy due to the exclusion of water from cubes in the surface, proportional to the surface area from which water is excluded
APPENDIX C. SURFACE COMPLEMENTARITY
Algorithm C.4 ENERGY(5, fs, Q, fq, Ec, Tb,τ, _trυ, cr, ch, cs, P): calculate the com- plementaπty energy between a surface S with functionality map fs and a quantized conformation Q with functionality map /,, conformational energy Ec, and τ-(, rotatable bonds using parameters: resolution r, translational-rotational-vibrational entropy _trt,, rotatable bond coefficient cr, hydrophobic energy coefficient Ch, hydrophobic surface energy coefficient cs, potential function P.
1. _-. «— crr„ {energy due to rotatable bonds }
2. _p <— 0 {energy due to potential interactions }
3 for all s 65 do
4 ifs G Qthen
5 -p^-„ + P(/,(«), /.(*))
6 end if
7 end for
8: _/, «— 0 {energy due to hydrophobicity } 9 for all s e 5 D do 10 define
T *- { (sx + l,Sy,sz),(sx-l,sy,Sz),
(SX,Sy + l,SZ),(SX,Sy ~ 1,5.), (SX,Sy,SZ - l) } li α <— r2 |T n 5| {the area of contact between the surface and the quantized conformation at this point }
12 if fs(s) = s then {hydrophobic energy due to slightly hydrophobic surface cubes}
13 Eh <— Eh+ chcsa
14 else if f,(s) € {T6, γ} then {hydrophobic energy due to non-polar surface cubes}
15 Eh <— E + Ch
16 end if
17 if fq(s) € {-^,^7} then {hydrophobic energy due to non-polar quantized conformation cubes }
18 Eh+-Eh + cha
19 end if
20 end for
21 return (Etrv + Eτ + Ep + Eh - Ec)
Appendix D
Molecular Library Comparison
A library is a set of molecular structures Given a library, the set of complementary theoretical surfaces is defined as the union of all surface shape/functionality pairs complementary to any quantized conformation of any molecule in the library
The algoπthm LlBRARYCOMPARE calculates a score proportional to the similarity of the two hbraπes The score is calculated by representmg each library as its set of complementary theoretical surfaces, and usmg the S IMILARITYSCORE pπmitive to determine the similanty or dissimilaπty two sets of theoretical surfaces If the molecular braπes each contam only one molecule, then the algonthm calculates a score proprotional to the similanty of the two molecules
Algorithm D.l LIBRARYCOMPARE(_!, _2) calculate a similanty score between 0 and 1000 for two molecular hbraπes
1 define 7 to be the theoretical surfaces complementary to _ι
2 define T to be the theoretical surfaces complementary to £2
3 return (SIMILARITYSCORE^, T2))
APPENDIX D MOLECULAR LIBRARY COMPARISON
Algorithm D.2 SIMILIARITΎSCORE(TI,72). calculate a similanty score between 0 and 1000 for two sets of theoretical surfaces i define sets of complementary surface shapes i «- {S such that for some /, (5. /) 6 Tx } S *- {S such that for some f,(S.f) e T2}
For purposes of comparmg two theoretical surfaces or surface shapes, note that translations and x-y plane rotations of a surface are considered identical to the oπginal surface
2 ss <- [_ι n S21 / |-?ι U S21 {shape score }
3 Sf <— 0{ functionality score }
4 for all S e Si n S2 do
5 define complementary functionalities for this surface shape i - {/ such that (S,/)€7"ι} F2 «- {/ such that (5, /) € T2}
6 define sets of "active" cubes, that is cubes with non-default functionality
Qx — {Q C Z3 such mat 3/ € ι,V96 Q,/(ςr) .F8} 2 <- (Q C Z3 such that 3/ e F2, Vg € Q, /(?) ≠ T8} Qx — {Q C Z3 such that 3/ 6 Fi n F2 V9 € Q, J(q) φ Ts) β/*-s + |Q,|/(|5ιn32|mιn(|βι|,|Qa|))
8 end for
9 return (1000s-s/)
Appendix E
Protein Quantization
In a process complementary to the quantization of a molecular conformation, the target sites of a protein surface are quantized into the same negative space cubic representation used by theoretical surfaces. This allows the following analyses:
• compaπson between a set of known protein target sites and the set of all possible theoretical surfaces within given parameters of volume, shape, and functionality
• compaπson between two different sets of known protein target sites.
• compaπson between a set of known protein target sites and a set of theoretical surfaces to which a given set of molecules is complementary
The protein quantization process is accomplished in the following steps, as depicted in Figure 30
1 A 3D crystal structure ofthe protem is examined and a functionality map is built for the protein atoms using the algoπthm A TOMFUNCTIONALMAP
2 A protein surface is generated from the 3D structure A protein surface is a set of mangles defining the surface ofthe protein that is accessible to water molecules (known as the Connolly surface). Michael Connolly's MSRoll software is an example of a package that can generate a protein surface suitable for this purpose
3 Subsets of the surface which are target sites likely for the binding of small molecules are detected. This can be accomplished, for example, by looking for highly concave regions. Michael Connolly's MSForm software is an example of a package that can measure surface curvature and detect pockets suitable for this purpose
4. Each target site is isolated and examined individually.
5. Each target site is quantized mto a set of negative space cubes with associated functionalities using the protem function map and the algoπthm T ARGETSlTE- QUANTIZE. The undeπngly process is very similar to the algonthm Q UANTIZE, APPENDIX E PROTEIN QUANTIZATION
using an imaginary molecule that is defined by atoms centered at a set of points with a given radius that fill the target site
Each set of quantized negative space cubes with functionality is convened to a set of theoretical surfaces satisfying the proper constraints (for example, no occluded cubes are allowed) using the algoπthm B UILDSURFACES
Algoπthm E.l
Figure imgf000069_0001
npj Sr, ) calculate a negative space cubic representation ofthe target site defined by the tπangles in set T, with v as a normal vector pointing out ofthe target site, / as a functionality map for the entire protein, using parameters resolution r, lattice density n, lattice van Der Waals radius rv, lattice tolerance t;, number of lattice neighbors p, buffer distance b, search radius sr, and additional parameters for subrountme calls (see below) as necessary
1 define a coordinate frame T; with oπgm at the center ofthe target site, z axis m the direction of v, x and y axes determined by the longest side ofthe pocket
2 define a set of points V C R3 to be the points on a lattice with coordinate frame 7) and cube side length r; such that p € V if p is contained in the target site and the closest triangle in T is at least distance b away from p
3 remove from V all points who do not have at least np neighbonng points also in V, where each pomt has neighbors consisting ofthe 26 points on the lattice offset by at most one cube from the pomt in question
4 ifP is disconnected, consider each connected component seperately in the following steps
5 define a "molecular" structure M with conformation c by imagining V to be a set of atoms with van Der Waals radius r v
6 define a base coordinate frame T„ using the algonthm B ASEFRAME on M and c
7 define a centeπng coordinate frame Tc using the algoπthm C ENTERFRAME on M and c
8 define a set of perturbation frames using PERTURBATIONFRAMES
9 as in QUANTIZE, by examining all perturbation frames, find the total frame which minimizes the number of negative space cubes of resolution r (calculated using CUBIFY with tolerance ti) and the total transformation distance
10 denote the above set of negative space cubes by Q
11 redefine the coordinate system for the negative space cubes in Q, such that the z axis is the direction with the greatest component in the v direction, and the maximum z value is 0
12 build a functionality map fq Q -→ T by, for each q Q, finding the highest pπoπty functionality associated by / with an atom the protein withm search radius sr to the center of q, or assigning Tβ if there are no such atoms
13 return (Q, fq)
APPENDIX E PROTEIN QUANTIZATION
Algorithm E.2 BulLDSURFACES(QJ9, 0, 7i/ ) return the set of theoretical surfaces satifymg proper constraints descnbed by the negative space cubes Q with functionality /,, using parameters- maximum occlusions m0, number of functionality points n/, additional parameters for subroutine calls (see below) as necessary
1 define O <— 0 and d . I? -→ Z such that d(x, y) = 0{ future surface opening and depths}
2 for all q € Q do {build an openmg and depth function }
3 if (qx, qy) $. O then
4 add (qx, qy) to O, d(qx, qy) <- max(d{qx, qy), 1 - <?_)
5 end if
6 end for
7 oc «— REMOVEOCCLUDEDCUBES( Q, 0, d)
8 if oc > τι0 then
9 return (0){too many occluded cubes } 10 end if l i if O is disconnected then
12 define O to be the set of connected components, 5 <— 0
13 for all Oc € O do {generate a set of surfaces for each connected component }
14 Qc «— {q € Q such that (qx, qy) e Oc}
15 5 *- 5 U BUILDSURFACES(QC, /,)
16 end for
17 return (<S)
18 end if
19 if O does not satisfy filteπng rules imposed by OPENINGFILTER then
20 return (0)
21 end if
22 S «— shape(ϋ, d)
23 if S does not satisfy filteπng rules imposed by SHAPEFILTER then
24 return (0)
25 end if
26 S <- 0
27 Qα <— {q € Q such that fq(q) Jcg} {set of negative space cubes with non- default functionality }
28 for all Q C Qα such that |Q| = nf do
29 define fs S → T such that f,(x, y, z) = fq(x, y, z ) for (x, y) £ Qα, fa (x, y. z) = . s otherwise, add (S, fa) to 5
30 end for
31 return (<S) APPENDIX E. PROTEM QUANTIZAΗON
Algorithm E_ REMOVEOCCLUDEDCUBES(Q, 0,d): remove from Q cubes which are occluded, adjusting openmg O and depths d, re m the number of occluded cubes removed.
1 oc *— 0{count of occluded cubes }
2 for all (x. y) € O do
3 for z < 1 to 1 — d(x,y) do
4 if (x,y,z) € Qand (x,y,z + 1) g Qthen
5 oc «— oc + 1
6 remove (x, t/, 2) from
7 end if
8 end for
9 Zxy <- {qz,{x,y,qz) € 2}
10 if_ly=0then
11 delete (x, y) from O, d(x, y) «- 0
12 else
13 d(x, y) ι- m{l- qz,qz € Zxy}
14 end if
15 end for
16 return (oc)
Appendix F
Protein Surface Comparison
A target surface set is defined as the set of all theoretical target surfaces to which a set of known protein surfaces map. The target surface set may compnse, for example, all of the surfaces mapped from one protem, all of the surfaces mapped from multiple proteins, or all ofthe surfaces mapped from specific sites on multiple proteins
The algoπthm P ROTEINCOMPARE calculates a score proportional to the similanty ofthe two sets of protein surfaces. The score is calculated by representmg each protein surface set as its target surface set, and usmg the S IMILARITYSCORE pπmitive to determine the similanty or dissimilanty two sets of theoretical surfaces.
Algorithm F.l PROTEINCOMPARE^I, ^) calculate a similanty score between 0 and 1000 for two protein surface sets.
1 define 7ϊ to be the target surface set corresponding to Vι
2 define T to be the target surface set corresponding to V
3 return (SIMILARITYSCORE(TI, T2))
Appendix G
Protein/Library Comparison
The algorithm P ROTEINLIBRARYCOMPARE calculates a score proportional to the com- plementanty of a library of small molecules and a set of protem surfaces. The score is calculated by representing the protein surface set as the theoretical surface set to which it is similar, the molecular library as the theoretical surface set to which it is complementary, and usmg the S IMILARITYSCORE pnmitive to determine the similanty or dissimilarity two sets of theoretical surfaces.
Algorithm G.l PROTEINLIBRARYCOMPARE(£, 7'): calculate a complementaπty score between 0 and 1000 for a molecular library £ and a protein surface set V.
1 define 7j to be the theoretical surfaces complementary to £
2 define Tp to be the target surface set corresponding to V
3 return (SlMlLARlTYSC0RE(7ϊ, 7p))
Appendix H
Parameter Values
The values of parameters used are as follows:
• Maximum openmg area ( A): 15
• Maximum surface volume ( V): 18
• Area threshold: (At): 8
• Maximum non-central openmg squares ( nc): 5
• Maximum contiguous openmg squares ( Mc). 3
• Maximum surface extrusions ( Me): 3
• Number of surface cubes of specific functionality (n/): 4
• Maximum number of conformations per molecule: 300
• Resolution (r): 4.24 Angstroms
• Tolerance ( i): 0.32
• Ring factor (77): 0
• Ring minimum (rm): 13
• Radius factor (ςrr): 0.75
• Centering tolerance (ct): 0.1
• Centeπng fraction (c ): 0.75
• Number of translations (nt): 5
• Translational vaπance ( v.): 0.2
• Number of rotations (nr): 5 APPENDIX H. PARAMETER VALUES
Rotational vanance ( υr): 0.1 Polarizable minimum (pm). 2 Minimum fit quanta (qmιn): 9 Minimum slackness ( smιn): 2 Maximum slackness (smα:r): 0 Maximum protrusion levels (pmax)'- 1 Minimum energy (Emm): 8.0 kCal
Translational-rotation-vibrational entropy ( Etrv)- — 9 0 kCal Rotatable bond coefficient (rc)- -0.7 kCal/bond Hydrophobic energy coefficient (c>,): 0.025 kCal Angstrom2 Hydrophobic surface energy coefficient (cs): 0.8 Potential function ( P): kCal
Surface
Molecule
Figure imgf000075_0001
Figure imgf000075_0002
Lattice density ( n): 1.5 Angstroms
Lattice van Der Waals radius ( r„): 0 75 Angstroms
Lattice tolerance ( i;): 0.2
Buffer distance (b): 0.5 Angstroms
Number of lattice neighbors ( p): 3
Search radius (sr): 2.22 Angstroms
Maximum number of occluded cubes ( τn0): 1

Claims

1. A computer-based method comprising defining a set of constraints on possible target surfaces, defining a fully enumerated set of theoretical target surfaces under the defined constraints, such that each surface has a defined, continuous volume and a defined, continuous surface area, mapping one or more sets of objects to the fully enumerated set of theoretical target surfaces to define corresponding subsets of the fully enumerated set of theoretical target surfaces, and analyzing an aspect of diversity of the objects based on degrees of similarities and differences among the corresponding subsets.
2. The method of claim 1 in which the target surfaces comprise negative space target surfaces.
3. The method of claim 1 in which the objects comprise positive space object surfaces associated with different molecules.
4. The method of claim 2 in which the objects comprise positive space object surfaces associated with different molecules and in which the objects are mapped by defining corresponding subsets ofthe fully enumerated set of negative space theoretical target surfaces to which positive space object surfaces of conformations of molecules are complementary, and the aspect of diversity that is analyzed is the difference or similarity between the molecules which map to those negative space theoretical target surfaces.
5. The method of claim 1 in which the objects comprise negative space object surfaces associated with different proteins.
6. The method of claim 2 in which the objects comprise negative space object surfaces associated with different proteins and in which the objects are mapped by defining corresponding subsets of the fully enumerated set of negative space theoretical target surfaces to which negative space object surfaces of protein pockets are similar, and the aspect of diversity that is analyzed is the difference or similarity between protem pockets which map to those negative space theoretical target surfaces.
7. The method of claim 1 in which the objects comprise positive space object surfaces associated with different molecules and negative space object surfaces associated with different proteins.
8. The method of claim 2 in which the objects comprise positive space object surfaces associated with different molecules and negative space object surfaces associated with different proteins and in which, in the case of molecules, the objects are mapped by defining corresponding subsets of the fully enumerated set of negative space theoretical target surfaces to which positive space object surfaces of conformations of molecules are complementary, in the case of proteins, the objects are mapped by defining corresponding subsets ofthe fully enumerated set of negative space theoretical target surfaces to which negative space object surfaces of protein pockets are similar, and the aspect of diversity that is analyzed is the difference or similarity ofthe molecules which map to those negative space theoretical target surfaces to the protein pockets which map to those negative space theoretical target surfaces.
9. The method of claim 1 in which the theoretical target surfaces comprise polyhedrons.
10. The method of claim 1 in which the objects comprise polyhedrons.
11. The method of claim 9 or 10 in which the polyhedrons comprise cubes.
12. The method of claim 9 or 10 in which the polyhedrons are all ofthe same size and shape.
13. The method of claim 1 in which the set of all theoretical target surfaces defines a diversity space within which the diversity of objects can be measured by mapping those objects to the diversity space.
14. The method of claim 13 also including identifying regions ofthe diversity space to which no objects map.
15. The method of claim 14 also including designing molecules that occupy at least one of the unfilled theoretical target surfaces of the diversity space.
16. The method of claims 4 or 8 in which complementarity is associated with binding affinities of positive space object surfaces of conformations of molecules to negative space theoretical target surfaces.
17. The method of claim 1 in which the constraints comprise volume.
18. The method of claim 1 in which the constraints comprise associating each of a number of sites of the target surface with a preselected molecular property.
19. The method of claim 18 in which each of the preselected molecular properties is drawn from a larger set of possible molecular properties.
20. The method of claim 18 in which the preselected molecular properties include hydrophobic, polarizable, H-bond acceptor, H-bond donor, H-bond donor/acceptor, potentially positively charged, and potentially negatively charged.
21. The method of claim 18 in which fewer than all of the sites of the target surface are each associated with a different one of the molecular properties and all ofthe other sites ofthe target surface are associated with a common molecular property.
22. The method of claim 21 in which the common molecular property comprises slightly hydrophobic.
23. The method of claim 1 in which the degrees of similarities or differences comprise functional properties associated with the corresponding subsets ofthe fully enumerated set of theoretical target surfaces.
24. The method of claim 1 in which the degrees of similarities or differences comprise shape properties associated with the corresponding subsets of the fully enumerated set of theoretical target surfaces.
25. The method of claim 1 further comprising defining each of the objects by quantizing molecules into polyhedrons.
26. The method of claim 1 also including fitting each of a fixed set of orientations of each conformation of each of the objects to each of the target surfaces.
27. The method of claim 26 further comprising scoring each ofthe fittings.
28. The method of claim 9 in which the constraints comprise a resolution of the polyhedrons.
29. The method of claim 28 in which the resolution is 4.24 Angstroms.
30. The method of claim 9 in which the constraints comprise maximum and minimum numbers of polyhedrons.
31. The method of claim 9 in which each ofthe polyhedrons shares a common interface with another ofthe polyhedrons.
32. The method of claim 1 in which each ofthe target surfaces has no occlusions of volume greater than a given parameter.
33. The method of claim 1 in which the target surfaces are defined conceptually as having been carved out of a flat surface.
34. A method comprising categorizing existing molecules based on negative space target surfaces to which conformations of the molecules are complementary, and designing novel molecules that are complementary to negative space target surfaces to which no conformations ofthe existing molecular are complementary.
35. A method of creating novel molecules to be tested as ligands for proteins, comprising categorizing proteins based on target surfaces to which their pockets of known structure map, and designing novel molecules that are complementary to the negative space target surfaces to which the protein pockets map.
36. A computer programmed to determine the chemical similarity of different molecules, the program comprising approximating the surface shape of each one of a plurality of molecules of interest by linking a series of cubes, each cube having a dimension R, the locations ofthe cubes being determined by the calculated electron probability density ofthe individual one of the molecules of interest, each cube sharing at least one of its six faces with another cube, such that there is a specific number of linked cubes which varies for each individual one of the plurality of molecules of interest; approximating the chemical reactivity of each individual one ofthe plurality of molecules of interest by assigning each cube of each individual one of the plurality of molecules of interest, no more than one functionality value from a plurality of M different chemical functionality values; approximating the surface shape and chemical reactivity of a chemically active surface having a volume equal to V by subtracting a number V/R.3 cubes of dimension R from a surface, wherein each of the cube spaces shares at least one face with another cube space and wherein N of the cube spaces has one of a plurality of M different chemical functionality values; calculating an attraction value K for each one of the plurality of molecules of interest to the chemically active surface; and calculating a list of overall attraction values to the chemically active surface.
37. The computer of claim 36, wherein further the calculation of the attraction value K is performed on a plurality of different predetermined chemically active surfaces, and a matrix of overall attractive values of each molecule of interest to each of the different surfaces is calculated.
38. The computer of claim 36, wherein the plurality of molecules of interest includes organic molecules.
39. The computer of claim 38, wherein further the chemically active surface having a plurality of predetermined active chemical locations is calculated to correspond to the shape of an actual protein surface structure.
40. The computer of claim 36, wherein further the molecules of interest are organic molecules of 1500 Daltons or less.
41. The computer of claim 36 wherein further the chemically active surface having a plurality of predetermined active chemical locations is compared to an actual protein surface to calculate a similarity value ofthe actual protein surface to the predetermined active chemical locations.
42. The computer of claim 41 wherein further a plurality of predetermined chemically active surfaces are compared to a plurality of actual protein surfaces and a matrix of similarity values is calculated.
43. The computer of claim 42 wherein further the cube spaces subtracted from the surface are calculated to approximate the electron probability density of at least one of a plurality of depressions in known protein surface structures.
44. The computer of claim 42 wherein further the N sites of chemical functionality are calculated to approximate the location and type of chemical functionality of actual depressions in known protein structures.
PCT/US2000/008777 1999-04-02 2000-03-31 Analyzing molecule and protein diversity Ceased WO2000060507A2 (en)

Priority Applications (4)

Application Number Priority Date Filing Date Title
JP2000609930A JP2002541560A (en) 1999-04-02 2000-03-31 Analysis of molecular and protein diversity
EP00925889A EP1203330A2 (en) 1999-04-02 2000-03-31 Analyzing molecule and protein diversity
CA002369570A CA2369570A1 (en) 1999-04-02 2000-03-31 Analyzing molecule and protein diversity
AU44511/00A AU4451100A (en) 1999-04-02 2000-03-31 Analyzing molecule and protein diversity

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
US12748699P 1999-04-02 1999-04-02
US60/127,486 1999-04-02

Publications (2)

Publication Number Publication Date
WO2000060507A2 true WO2000060507A2 (en) 2000-10-12
WO2000060507A3 WO2000060507A3 (en) 2001-04-12

Family

ID=22430392

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/US2000/008777 Ceased WO2000060507A2 (en) 1999-04-02 2000-03-31 Analyzing molecule and protein diversity

Country Status (5)

Country Link
EP (1) EP1203330A2 (en)
JP (1) JP2002541560A (en)
AU (1) AU4451100A (en)
CA (1) CA2369570A1 (en)
WO (1) WO2000060507A2 (en)

Cited By (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
WO2001098457A3 (en) * 2000-06-16 2003-02-27 Neogenesis Pharmaceuticals Inc Surface model of a protein
EP2581849A1 (en) * 2002-07-24 2013-04-17 Keddem Bio-Science Ltd. Drug discovery method
CN110390997A (en) * 2019-07-17 2019-10-29 成都火石创造科技有限公司 A kind of chemical molecular formula joining method
CN113421610A (en) * 2021-07-01 2021-09-21 北京望石智慧科技有限公司 Molecular superimposition conformation determination method and device and storage medium

Families Citing this family (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN110010199B (en) * 2019-03-27 2021-01-01 华中师范大学 An analytical method for identifying protein-specific drug-binding pockets

Non-Patent Citations (4)

* Cited by examiner, † Cited by third party
Title
BURKHARD P ET AL: "An example of a protein ligand found by database mining: Description of the docking method and its verification by a 2.3 A X-ray structure of a thrombin-ligand complex." JOURNAL OF MOLECULAR BIOLOGY, vol. 277, no. 2, 27 March 1998 (1998-03-27), pages 449-466, XP000974379 ISSN: 0022-2836 *
JIANG F ET AL: "SOFT DOCKING MATCHING OF MOLECULAR SURFACE CUBES" JOURNAL OF MOLECULAR BIOLOGY, vol. 219, no. 1, 1991, pages 79-102, XP000972336 ISSN: 0022-2836 *
JOHN MOUNT ET AL: "Icepick: A flexible surface-based system for molecular diversity" JOURNAL OF MEDICINAL CHEMISTRY, [Online] vol. 42, no. 1, 1999, pages 60-66, XP002156348 Retrieved from the Internet: <URL:http://pubs3.acs.org/s97is.vts> [retrieved on 2000-12-28] *
PARKS CAMDEN A ET AL: "The measurement of molecular diversity by receptor site interaction simulation." JOURNAL OF COMPUTER-AIDED MOLECULAR DESIGN, vol. 12, no. 5, 1998, pages 441-449, XP000972384 ISSN: 0920-654X *

Cited By (7)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
WO2001098457A3 (en) * 2000-06-16 2003-02-27 Neogenesis Pharmaceuticals Inc Surface model of a protein
EP2581849A1 (en) * 2002-07-24 2013-04-17 Keddem Bio-Science Ltd. Drug discovery method
US9405885B2 (en) 2002-07-24 2016-08-02 Keddem Bioscience Ltd. Drug discovery method
CN110390997A (en) * 2019-07-17 2019-10-29 成都火石创造科技有限公司 A kind of chemical molecular formula joining method
CN110390997B (en) * 2019-07-17 2023-05-30 成都火石创造科技有限公司 Chemical molecular formula splicing method
CN113421610A (en) * 2021-07-01 2021-09-21 北京望石智慧科技有限公司 Molecular superimposition conformation determination method and device and storage medium
CN113421610B (en) * 2021-07-01 2023-10-20 北京望石智慧科技有限公司 A method, device and storage medium for determining molecular superposition conformation

Also Published As

Publication number Publication date
CA2369570A1 (en) 2000-10-12
JP2002541560A (en) 2002-12-03
WO2000060507A3 (en) 2001-04-12
AU4451100A (en) 2000-10-23
EP1203330A2 (en) 2002-05-08

Similar Documents

Publication Publication Date Title
US6240374B1 (en) Further method of creating and rapidly searching a virtual library of potential molecules using validated molecular structural descriptors
García-Sosa et al. WaterScore: a novel method for distinguishing between bound and displaceable water molecules in the crystal structure of the binding site of protein-ligand complexes
US5025388A (en) Comparative molecular field analysis (CoMFA)
US20070134662A1 (en) Structural interaction fingerprint
Bergman Statistical methods for pharmaceutical research planning
US8374837B2 (en) Descriptors of three-dimensional objects, uses thereof and a method to generate the same
US7110888B1 (en) Method for determining a shape space for a set of molecules using minimal metric distances
Andersson et al. Mapping of ligand‐binding cavities in proteins
EP1203330A2 (en) Analyzing molecule and protein diversity
Cottrell et al. Incorporating partial matches within multiobjective pharmacophore identification
MEZEY Local-shape analysis of macromolecular electron densities
Thareja et al. Computational tools in cheminformatics
Shahlaei et al. Virtual screening based on pharmacophore model followed by docking simulation studies in search of potential inhibitors for p38 map kinase
US8165818B2 (en) Method and apparatus for searching molecular structure databases
Good 3D molecular similarity indices and their application in QSAR studies
EP1862927B1 (en) Descriptors of three-dimensional objects, uses thereof and a method to generate the same
CA2321303C (en) Method for determining a shape space for a set of molecules using minimal metric distances
Rarey et al. Docking and scoring for structure‐based drug design
AU2008202475A1 (en) Descriptors of three-dimensional objects, uses thereof and a method to generate the same
WO2009146735A1 (en) Descriptors of three-dimensional objects, uses thereof and a method to generate the same
US6470305B1 (en) Chemical analysis by morphological similarity
CA2633179A1 (en) Descriptors of three-dimensional objects, uses thereof and a method to generate the same
HK1115654B (en) Descriptors of three-dimensional objects, uses thereof and a method to generate the same
Treado Computational Studies of Packing and Jamming in Biological Systems
US20070005258A1 (en) Identification of ligands for macromolecules

Legal Events

Date Code Title Description
AK Designated states

Kind code of ref document: A2

Designated state(s): AE AG AL AM AT AU AZ BA BB BG BR BY CA CH CN CR CU CZ DE DK DM DZ EE ES FI GB GD GE GH GM HR HU ID IL IN IS JP KE KG KP KR KZ LC LK LR LS LT LU LV MA MD MG MK MN MW MX NO NZ PL PT RO RU SD SE SG SI SK SL TJ TM TR TT TZ UA UG US UZ VN YU ZA ZW

AL Designated countries for regional patents

Kind code of ref document: A2

Designated state(s): GH GM KE LS MW SD SL SZ TZ UG ZW AM AZ BY KG KZ MD RU TJ TM AT BE CH CY DE DK ES FI FR GB GR IE IT LU MC NL PT SE BF BJ CF CG CI CM GA GN GW ML MR NE SN TD TG

121 Ep: the epo has been informed by wipo that ep was designated in this application
DFPE Request for preliminary examination filed prior to expiration of 19th month from priority date (pct application filed before 20040101)
AK Designated states

Kind code of ref document: A3

Designated state(s): AE AG AL AM AT AU AZ BA BB BG BR BY CA CH CN CR CU CZ DE DK DM DZ EE ES FI GB GD GE GH GM HR HU ID IL IN IS JP KE KG KP KR KZ LC LK LR LS LT LU LV MA MD MG MK MN MW MX NO NZ PL PT RO RU SD SE SG SI SK SL TJ TM TR TT TZ UA UG US UZ VN YU ZA ZW

AL Designated countries for regional patents

Kind code of ref document: A3

Designated state(s): GH GM KE LS MW SD SL SZ TZ UG ZW AM AZ BY KG KZ MD RU TJ TM AT BE CH CY DE DK ES FI FR GB GR IE IT LU MC NL PT SE BF BJ CF CG CI CM GA GN GW ML MR NE SN TD TG

ENP Entry into the national phase

Ref country code: CA

Ref document number: 2369570

Kind code of ref document: A

Format of ref document f/p: F

Ref country code: JP

Ref document number: 2000 609930

Kind code of ref document: A

Format of ref document f/p: F

WWE Wipo information: entry into national phase

Ref document number: 2000925889

Country of ref document: EP

REG Reference to national code

Ref country code: DE

Ref legal event code: 8642

ENP Entry into the national phase

Ref document number: 2369570

Country of ref document: CA

WWP Wipo information: published in national office

Ref document number: 2000925889

Country of ref document: EP

WWW Wipo information: withdrawn in national office

Ref document number: 2000925889

Country of ref document: EP