CLUSS: Clustering of protein sequences based on a new similarity measure

A Krause; Abdellali Kelil; AJ Enright; Alain Fleury; C Notredame; D Higgins; ELL Sonnhammer; F Titgemeyer; G Reinert; G Yona; H Lodish; IV Tetko; J Felsenstein; J Heringa; J Rocha; JD Thompson; JD Thompson; JH Ward; JH Ward; JS Varré; K Katoh; K Sjölander; K Sjölander; M Ike; M Kimura; MO Dayhoff; MY Leung; N Côté; N Wicker; P Pipenbacher; R Jothi; RC Edgar; RC Edgar; RO Duda; Ryszard Brzezinski; S Fanning; S Henikoff; S Karlin; S Karlin; S Karlin; S Vinga; SF Altschul; SF Altschul; Shengrui Wang; T Fukamizo; T Ishimizu; V Batagelj

CLUSS: Clustering of protein sequences based on a new similarity measure

Authors: A Krause
Abdellali Kelil
AJ Enright
Alain Fleury
C Notredame
D Higgins
ELL Sonnhammer
F Titgemeyer
G Reinert
G Yona
H Lodish
IV Tetko
J Felsenstein
J Heringa
J Rocha
JD Thompson
JD Thompson
JH Ward
JH Ward
JS Varré
K Katoh
K Sjölander
K Sjölander
M Ike
M Kimura
MO Dayhoff
MY Leung
N Côté
N Wicker
P Pipenbacher
R Jothi
RC Edgar
RC Edgar
RO Duda
Ryszard Brzezinski
S Fanning
S Henikoff
S Karlin
S Karlin
S Karlin
S Vinga
SF Altschul
SF Altschul
Shengrui Wang
T Fukamizo
T Ishimizu
V Batagelj
Publication date: 1 January 2007
Publisher: BioMed Central
Doi

Abstract

Abstract Background The rapid burgeoning of available protein data makes the use of clustering within families of proteins increasingly important. The challenge is to identify subfamilies of evolutionarily related sequences. This identification reveals phylogenetic relationships, which provide prior knowledge to help researchers understand biological phenomena. A good evolutionary model is essential to achieve a clustering that reflects the biological reality, and an accurate estimate of protein sequence similarity is crucial to the building of such a model. Most existing algorithms estimate this similarity using techniques that are not necessarily biologically plausible, especially for hard-to-align sequences such as proteins with different domain structures, which cause many difficulties for the alignment-dependent algorithms. In this paper, we propose a novel similarity measure based on matching amino acid subsequences. This measure, named SMS for Substitution Matching Similarity, is especially designed for application to non-aligned protein sequences. It allows us to develop a new alignment-free algorithm, named CLUSS, for clustering protein families. To the best of our knowledge, this is the first alignment-free algorithm for clustering protein sequences. Unlike other clustering algorithms, CLUSS is effective on both alignable and non-alignable protein families. In the rest of the paper, we use the term "<it>phylogenetic</it>" in the sense of "<it>relatedness of biological functions</it>". Results To show the effectiveness of CLUSS, we performed an extensive clustering on COG database. To demonstrate its ability to deal with hard-to-align sequences, we tested it on the GH2 family. In addition, we carried out experimental comparisons of CLUSS with a variety of mainstream algorithms. These comparisons were made on hard-to-align and easy-to-align protein sequences. The results of these experiments show the superiority of CLUSS in yielding clusters of proteins with similar functional activity. Conclusion We have developed an effective method and tool for clustering protein sequences to meet the needs of biologists in terms of phylogenetic analysis and prediction of biological functions. Compared to existing clustering methods, CLUSS more accurately highlights the functional characteristics of the clustered families. It provides biologists with a new and plausible instrument for the analysis of protein sequences, especially those that cause problems for the alignment-dependent algorithms.</p

Similar works

Full text

Open in the Core reader

Download PDF

Available Versions

Crossref

Last time updated on 01/04/2019

Springer - Publisher Connector

Last time updated on 28/04/2017

Directory of Open Access Journals

oai:doaj.org/article:2ebebf1bb...

Last time updated on 18/12/2014

Springer - Publisher Connector

Last time updated on 05/06/2019