Search CORE

Protein subcellular localization prediction for Gram-negative bacteria using amino acid subalphabets and a combination of multiple support vector machines

Author: Krishnan Arun
Li Kuo-Bin
Sung Wing-Kin
Wang Jiren
Publication venue: BioMed Central
Publication date: 01/01/2005
Field of study

BACKGROUND: Predicting the subcellular localization of proteins is important for determining the function of proteins. Previous works focused on predicting protein localization in Gram-negative bacteria obtained good results. However, these methods had relatively low accuracies for the localization of extracellular proteins. This paper studies ways to improve the accuracy for predicting extracellular localization in Gram-negative bacteria. RESULTS: We have developed a system for predicting the subcellular localization of proteins for Gram-negative bacteria based on amino acid subalphabets and a combination of multiple support vector machines. The recall of the extracellular site and overall recall of our predictor reach 86.0% and 89.8%, respectively, in 5-fold cross-validation. To the best of our knowledge, these are the most accurate results for predicting subcellular localization in Gram-negative bacteria. CONCLUSION: Clustering 20 amino acids into a few groups by the proposed greedy algorithm provides a new way to extract features from protein sequences to cover more adjacent amino acids and hence reduce the dimensionality of the input vector of protein features. It was observed that a good amino acid grouping leads to an increase in prediction performance. Furthermore, a proper choice of a subset of complementary support vector machines constructed by different features of proteins maximizes the prediction accuracy

Springer

Directory of Open Access Journals

ScholarBank@NUS

SChloro: directing Viridiplantae proteins to six chloroplastic sub-compartments

Author: Casadio Rita
Fariselli Piero
Martelli Pier Luigi
Savojardo Castrense
Publication venue: 'Oxford University Press (OUP)'
Publication date: 22/10/2016
Field of study

Motivation: Chloroplasts are organelles found in plants and involved in several important cell processes. Similarly to other compartments in the cell, chloroplasts have an internal structure comprising several sub-compartments, where different proteins are targeted to perform their functions. Given the relation between protein function and localization, the availability of effective computational tools to predict protein sub-organelle localizations is crucial for large-scale functional studies. Results: In this paper we present SChloro, a novel machine-learning approach to predict protein sub-chloroplastic localization, based on targeting signal detection and membrane protein information. The proposed approach performs multi-label predictions discriminating six chloroplastic sub-compartments that include inner membrane, outer membrane, stroma, thylakoid lumen, plastoglobule and thylakoid membrane. In comparative benchmarks, the proposed method outperforms current state-of-the-art methods in both single-and multi-compartment predictions, with an overall multi-label accuracy of 74%. The results demonstrate the relevance of the approach that is eligible as a good candidate for integration into more general large-scale annotation pipelines of protein subcellular localization

Archivio istituzionale della ricerca - Alma Mater Studiorum Università di Bologna

Archivio istituzionale della ricerca - Università di Padova

Institutional Research Information System University of Turin

Minimalist Ensemble Algorithms for Genome-Wide Protein Localization Prediction

Author: Hu Jianjun
Lin J.-R.
Liu R.
Mondal A. M.
Publication venue: Scholar Commons
Publication date: 01/01/2012
Field of study

Background Computational prediction of protein subcellular localization can greatly help to elucidate its functions. Despite the existence of dozens of protein localization prediction algorithms, the prediction accuracy and coverage are still low. Several ensemble algorithms have been proposed to improve the prediction performance, which usually include as many as 10 or more individual localization algorithms. However, their performance is still limited by the running complexity and redundancy among individual prediction algorithms. Results This paper proposed a novel method for rational design of minimalist ensemble algorithms for practical genome-wide protein subcellular localization prediction. The algorithm is based on combining a feature selection based filter and a logistic regression classifier. Using a novel concept of contribution scores, we analyzed issues of algorithm redundancy, consensus mistakes, and algorithm complementarity in designing ensemble algorithms. We applied the proposed minimalist logistic regression (LR) ensemble algorithm to two genome-wide datasets of Yeast and Human and compared its performance with current ensemble algorithms. Experimental results showed that the minimalist ensemble algorithm can achieve high prediction accuracy with only 1/3 to 1/2 of individual predictors of current ensemble algorithms, which greatly reduces computational complexity and running time. It was found that the high performance ensemble algorithms are usually composed of the predictors that together cover most of available features. Compared to the best individual predictor, our ensemble algorithm improved the prediction accuracy from AUC score of 0.558 to 0.707 for the Yeast dataset and from 0.628 to 0.646 for the Human dataset. Compared with popular weighted voting based ensemble algorithms, our classifier-based ensemble algorithms achieved much better performance without suffering from inclusion of too many individual predictors. Conclusions We proposed a method for rational design of minimalist ensemble algorithms using feature selection and classifiers. The proposed minimalist ensemble algorithm based on logistic regression can achieve equal or better prediction performance while using only half or one-third of individual predictors compared to other ensemble algorithms. The results also suggested that meta-predictors that take advantage of a variety of features by combining individual predictors tend to achieve the best performance. The LR ensemble server and related benchmark datasets are available at http://mleg.cse.sc.edu/LRensemble/cgi-bin/predict.cgi

Scholar Commons - Institutional Repository of the University of South Carolina

FGsub: Fusarium graminearum protein subcellular localizations predicted from primary structures

Author: A Höglund
A Pierleoni
A Reinhardt
AC Christina
Chenglei Sun
CJ Shin
FG Priest
H Chen
H Nakashima
J Cedano
J Liu
J Wang
JL Gardy
JM Chang
JW Bennett
K Nakai
KC Chou
KJ Park
KY Lee
Luonan Chen
M Bhasin
MS Scott
O Emanuelsson
P Garga
P Horton
R Nair
RS Goswami
SJ Hua
T Tamura
TU Consortium
U Guldener
Weihua Tang
WZ Li
Xing-Ming Zhao
XM Zhao
XM Zhao
Y Cai
Y Huang
Publication venue: BioMed Central
Publication date: 01/01/2010
Field of study

Digital Repository @ Iowa State University (ISU)

Machine learning methods for omics data integration

Author: Zhou Wengang
Publication venue: Iowa State University Digital Repository
Publication date: 01/01/2011
Field of study

High-throughput technologies produce genome-scale transcriptomic and metabolomic (omics) datasets that allow for the system-level studies of complex biological processes. The limitation lies in the small number of samples versus the larger number of features represented in these datasets. Machine learning methods can help integrate these large-scale omics datasets and identify key features from each dataset. A novel class dependent feature selection method integrates the F statistic, maximum relevance binary particle swarm optimization (MRBPSO), and class dependent multi-category classification (CDMC) system. A set of highly differentially expressed genes are pre-selected using the F statistic as a filter for each dataset. MRBPSO and CDMC function as a wrapper to select desirable feature subsets for each class and classify the samples using those chosen class-dependent feature subsets. The results indicate that the class-dependent approaches can effectively identify unique biomarkers for each cancer type and improve classification accuracy compared to class independent feature selection methods. The integration of transcriptomics and metabolomics data is based on a classification framework. Compared to principal component analysis and non-negative matrix factorization based integration approaches, our proposed method achieves 20-30% higher prediction accuracies on Arabidopsis tissue development data. Metabolite-predictive genes and gene-predictive metabolites are selected from transcriptomic and metabolomic data respectively. The constructed gene-metabolite correlation network can infer the functions of unknown genes and metabolites. Tissue-specific genes and metabolites are identified by the class-dependent feature selection method. Evidence from subcellular locations, gene ontology, and biochemical pathways support the involvement of these entities in different developmental stages and tissues in Arabidopsis

The Trypanosoma brucei MitoCarta and its regulation and splicing pattern during development

Author: Astrid Chanfon
Audic
Bannai
Benne
Bertrand
Besteiro
Bhasin
Bochud-Allemann
Borst
Brown
Burges
Chaudhuri
Claros
Cui
Daniel Nilsson
de Almeida
Dubchak
Eisenhaber
Eisenhaber
Emanuelsson
Ferguson
Folsch
Guda
Guda
Guo
Halic
Hashimi
Herrmann
Horton
Horton
Horvath
Hua
Huang
Huinan Wang
Juan Cui
Kanehisa
Kapila Gunasekera
Kumar
Lee
Li
Long
Lu
Matthews
Michels
Mokranjac
Nair
Nakai
Nilsson
Pagliarini
Panigrahi
Park
Perocchi
Petsalaki
Priest
Priest
Priest
Prilusky
Pusnik
Reinhardt
Sabatini
Simpson
Sloof
Small
Sutton
Tasker
Tetaud
Torsten Ochsenreiter
Uboldi
Vassella
Vickerman
von Heijne
Xiaobai Zhang
Xiaofeng Song
Xie
Ying Xu
Publication venue: Oxford University Press
Publication date: 01/01/2010
Field of study

It has long been known that trypanosomes regulate mitochondrial biogenesis during the life cycle of the parasite; however, the mitochondrial protein inventory (MitoCarta) and its regulation remain unknown. We present a novel computational method for genome-wide prediction of mitochondrial proteins using a support vector machine-based classifier with ∼90% prediction accuracy. Using this method, we predicted the mitochondrial localization of 468 proteins with high confidence and have experimentally verified the localization of a subset of these proteins. We then applied a recently developed parallel sequencing technology to determine the expression profiles and the splicing patterns of a total of 1065 predicted MitoCarta transcripts during the development of the parasite, and showed that 435 of the transcripts significantly changed their expressions while 630 remain unchanged in any of the three life stages analyzed. Furthermore, we identified 298 alternatively splicing events, a small subset of which could lead to dual localization of the corresponding proteins

Bern Open Repository and Information System (BORIS)

The Trypanosoma \u3ci\u3ebrucei\u3c/i\u3e MitoCarta and its regulation and splicing pattern during development

Author: Chanfon Astrid
Cui Juan
Gunasekera Kapila
Nilsson Daniel
Ochsenreiter Torsten
Song Xiaofeng
Wang Huinan
Xu Ying
Zhang Xiaobai
Publication venue: DigitalCommons@University of Nebraska - Lincoln
Publication date: 01/01/2010
Field of study

It has long been known that trypanosomes regulate mitochondrial biogenesis during the life cycle of the parasite; however, the mitochondrial protein inventory (MitoCarta) and its regulation remain unknown. We present a novel computational method for genome-wide prediction of mitochondrial proteins using a support vector machine-based classifier with ~90% prediction accuracy. Using this method, we predicted the mitochondrial localization of 468 proteins with high confidence and have experimentally verified the localization of a subset of these proteins. We then applied a recently developed parallel sequencing technology to determine the expression profiles and the splicing patterns of a total of 1065 predicted MitoCarta transcripts during the development of the parasite, and showed that 435 of the transcripts significantly changed their expressions while 630 remain unchanged in any of the three life stages analyzed. Furthermore, we identified 298 alternatively splicing events, a small subset of which could lead to dual localization of the corresponding proteins

DigitalCommons@University of Nebraska

Many Local Pattern Texture Features: Which Is Better for Image-Based Multilabel Human Protein Subcellular Localization Classification?

Author: Fan Yang
Hong-Bin Shen
Ying-Ying Xu
Publication venue: 'Hindawi Limited'
Publication date: 01/01/2014
Field of study

Human protein subcellular location prediction can provide critical knowledge for understanding a protein’s function. Since significant progress has been made on digital microscopy, automated image-based protein subcellular location classification is urgently needed. In this paper, we aim to investigate more representative image features that can be effectively used for dealing with the multilabel subcellular image samples. We prepared a large multilabel immunohistochemistry (IHC) image benchmark from the Human Protein Atlas database and tested the performance of different local texture features, including completed local binary pattern, local tetra pattern, and the standard local binary pattern feature. According to our experimental results from binary relevance multilabel machine learning models, the completed local binary pattern, and local tetra pattern are more discriminative for describing IHC images when compared to the traditional local binary pattern descriptor. The combination of these two novel local pattern features and the conventional global texture features is also studied. The enhanced performance of final binary relevance classification model trained on the combined feature space demonstrates that different features are complementary to each other and thus capable of improving the accuracy of classification