Remote Homology Detection of Protein Sequences

Comin, Matteo; Verzotto, Davide

research

Remote Homology Detection of Protein Sequences

Authors: Matteo Comin
Davide Verzotto
Publication date: 1 January 2010
Publisher: Dagstuhl Seminar Proceedings. 10231 - Structure Discovery in Biology: Motifs, Networks & Phylogenies
Doi

Abstract

The classification of protein sequences using string kernels provides valuable insights for protein function prediction. Almost all string kernels are based on patterns that are not independent, and therefore the associated scores are obtained using a set of redundant features. In this talk we will discuss how a class of patterns, called Irredundant, is specifically designed to address this issue. Loosely speaking the set of Irredundant patterns is the smallest class of independent patterns that can describe all patterns in a string. We present a classification method based on the statistics of these patterns, named Irredundant Class. Results on benchmark data show that Irredundant Class outperforms most of the string kernel methods previously proposed, and it achieves results as good as the current state-of-the-art methods with a fewer number of patterns. Unfortunately we show that the information carried by the irredundant patterns can not be easily interpreted, thus alternative notions are needed

Similar works

Full text

Open in the Core reader

Download PDF

Available Versions

Dagstuhl Research Online Publication Server

oai:drops-oai.dagstuhl.de:2741

Last time updated on 17/11/2016