A hybrid clustering of protein binding sites.

FEBS J

Protein Information Technology Group, Department of Computer Science, Eötvös University, Budapest, Hungary.

Published: March 2010

The Protein Data Bank contains the description of approximately 27 000 protein-ligand binding sites. Most of the ligands at these sites are biologically active small molecules, affecting the biological function of the protein. The classification of their binding sites may lead to relevant results in drug discovery and design. Clusters of similar binding sites were created here by a hybrid, sequence and spatial structure-based approach, using the OPTICS clustering algorithm. A dissimilarity measure was defined: a distance function on the amino acid sequences of the binding sites. All the binding sites were clustered in the Protein Data Bank according to this distance function, and it was found that the clusters characterized well the Enzyme Commission numbers of the entries. The results, carefully color coded by the Enzyme Commission numbers of the proteins, containing the 20 967 binding sites clustered, are available as html files in three parts at http://pitgroup.org/seqclust/.

Download full-text PDF

Source
http://dx.doi.org/10.1111/j.1742-4658.2010.07578.xDOI Listing

Publication Analysis

Top Keywords

binding sites
28
sites
8
protein data
8
data bank
8
distance function
8
sites clustered
8
enzyme commission
8
commission numbers
8
binding
7
hybrid clustering
4

Similar Publications

Want AI Summaries of new PubMed Abstracts delivered to your In-box?

Enter search terms and have AI summaries delivered each week - change queries or unsubscribe any time!