Search CORE

Queensland University of Technology ePrints Archive

The Australian National University

University of Queensland eSpace

SUMOhydro: A Novel Method for the Prediction of Sumoylation Sites Based on Hydrophobic Properties

Author: Chen Yong-Zi
Chen Zhen
Gong Yu-Ai
Ying Guoguang
Publication venue: Public Library of Science
Publication date: 01/01/2012
Field of study

Sumoylation is one of the most essential mechanisms of reversible protein post-translational modifications and is a crucial biochemical process in the regulation of a variety of important biological functions. Sumoylation is also closely involved in various human diseases. The accurate computational identification of sumoylation sites in protein sequences aids in experimental design and mechanistic research in cellular biology. In this study, we introduced amino acid hydrophobicity as a parameter into a traditional binary encoding scheme and developed a novel sumoylation site prediction tool termed SUMOhydro. With the assistance of a support vector machine, the proposed method was trained and tested using a stringent non-redundant sumoylation dataset. In a leave-one-out cross-validation, the proposed method yielded an excellent performance with a correlation coefficient, specificity, sensitivity and accuracy equal to 0.690, 98.6%, 71.1% and 97.5%, respectively. In addition, SUMOhydro has been benchmarked against previously described predictors based on an independent dataset, thereby suggesting that the introduction of hydrophobicity as an additional parameter could assist in the prediction of sumoylation sites. Currently, SUMOhydro is freely accessible at http://protein.cau.edu.cn/others/SUMOhydro/

CiteSeerX

FigShare

SVM-based prediction of linear B-cell epitopes using Bayes Feature Extraction

Author: Kam Yiu-Wing
Ng Lisa FP
Simarmata Diane
Tong Joo Chuan
Wee Lawrence JK
Publication venue: BioMed Central
Publication date: 01/01/2010
Field of study

10.1186/1471-2164-11-S4-S21BMC Genomics11SUPPL. 4S2

Springer - Publisher Connector

ScholarBank@NUS

Prediction of Ubiquitination Sites by Using the Composition of k-Spaced Amino Acid Pairs

Author: A Catic
A Hershko
AL Chernorudskiy
AL Hitchcock
AL Schwartz
Chuan Wang
CM Pickart
CR Ingrell
CW Tung
E Tomlinson
Franca Fraternali
H Li
J Herrmann
J Peng
J Shao
J Song
J Song
J Xu
JN Si
K Chen
K Chen
K Chen
K Chen
K Haglund
L Hicke
M Gribskov
P Radivojac
Ren-Xiang Yan
RM Centor
RX Yan
S Kawashima
V Neduva
VN Vapnik
WC Lee
XB Wang
XG Yang
XG Yang
Xiao-Feng Wang
Y Cai
Y Xue
Y Xue
Yong-Zi Chen
YR Tang
YZ Chen
Zhen Chen
Ziding Zhang
Publication venue: Public Library of Science
Publication date: 29/07/2011
Field of study

As one of the most important reversible protein post-translation modifications, ubiquitination has been reported to be involved in lots of biological processes and closely implicated with various diseases. To fully decipher the molecular mechanisms of ubiquitination-related biological processes, an initial but crucial step is the recognition of ubiquitylated substrates and the corresponding ubiquitination sites. Here, a new bioinformatics tool named CKSAAP_UbSite was developed to predict ubiquitination sites from protein sequences. With the assistance of Support Vector Machine (SVM), the highlight of CKSAAP_UbSite is to employ the composition of k-spaced amino acid pairs surrounding a query site (i.e. any lysine in a query sequence) as input. When trained and tested in the dataset of yeast ubiquitination sites (Radivojac et al, Proteins, 2010, 78: 365–380), a 100-fold cross-validation on a 1∶1 ratio of positive and negative samples revealed that the accuracy and MCC of CKSAAP_UbSite reached 73.40% and 0.4694, respectively. The proposed CKSAAP_UbSite has also been intensively benchmarked to exhibit better performance than some existing predictors, suggesting that it can be served as a useful tool to the community. Currently, CKSAAP_UbSite is freely accessible at http://protein.cau.edu.cn/cksaap_ubsite/. Moreover, we also found that the sequence patterns around ubiquitination sites are not conserved across different species. To ensure a reasonable prediction performance, the application of the current CKSAAP_UbSite should be limited to the proteome of yeast

Extraction of consensus protein patterns in regions containing non-proline cis peptide bonds and their functional assessment

Author: Exarchos Konstantinos P
Exarchos Themis P
Fotiadis Dimitrios I
Papaloukas Costas
Rigas Georgios
Publication venue: BioMed Central
Publication date: 01/01/2011
Field of study

Abstract Background In peptides and proteins, only a small percentile of peptide bonds adopts the <it>cis </it>configuration. Especially in the case of amide peptide bonds, the amount of <it>cis </it>conformations is quite limited thus hampering systematic studies, until recently. However, lately the emerging population of databases with more 3D structures of proteins has produced a considerable number of sequences containing non-proline <it>cis </it>formations (<it>cis</it>-nonPro). Results In our work, we extract regular expression-type patterns that are descriptive of regions surrounding the <it>cis</it>-nonPro formations. For this purpose, three types of pattern discovery are performed: i) exact pattern discovery, ii) pattern discovery using a chemical equivalency set, and iii) pattern discovery using a structural equivalency set. Afterwards, using each pattern as predicate, we search the Eukaryotic Linear Motif (ELM) resource to identify potential functional implications of regions with <it>cis</it>-nonPro peptide bonds. The patterns extracted from each type of pattern discovery are further employed, in order to formulate a pattern-based classifier, which is used to discriminate between <it>cis</it>-nonPro and <it>trans</it>-nonPro formations. Conclusions In terms of functional implications, we observe a significant association of <it>cis</it>-nonPro peptide bonds towards ligand/binding functionalities. As for the pattern-based classification scheme, the highest results were obtained using the structural equivalency set, which yielded 70% accuracy, 77% sensitivity and 63% specificity.</p

Springer - Publisher Connector

Prodepth: Predict Residue Depth by Support Vector Regression Approach from Protein Sequences Only

Author: A Pintar
A Pintar
A Pintar
A Schlessinger
A Schlessinger
A Schlessinger
A Schlessinger
A Shrake
AG Murzin
AR Kinjo
AR Kinjo
Ashley M. Buckle
B Lee
B Rost
B Rost
B Rost
C Chothia
CK Smith
D Baker
D Varrazzo
D Xie
DT Jones
DT Jones
E Schmitt
EM Marcotte
F Ferre
G Pollastri
Geoffrey I. Webb
GP Raghava
H Chen
H Zhang
H Zhou
Hao Tan
HM Berman
J Cheng
J Cheng
J Qiu
J Song
J Song
J Song
J Song
J Wan
James C. Whisstock
JC Whisstock
Jiangning Song
JJ Ward
JM Chandonia
JU Bowie
K Bajaj
K Chen
K Vlahovicek
Khalid Mahmood
L Kurgan
LA Kurgan
M Connolly
M Kumar
M Lee
M Stout
ME Lacombe-Harvey
MK Kalita
MN Nguyen
O Schueler-Furman
P Radivojac
RG Coleman
Ruby H. P. Law
S Ahmad
S Chakravarty
S Liu
S Miller
Sean David Mooney
SF Altschul
T Hamelryck
T Ishida
T Joachims
T Noguchi
Tatsuya Akutsu
TL Blundell
V Vapnik
V Vapnik
W Kabsch
W Liu
W Zhang
WL DeLano
X Wang
Y Bromberg
Y Kalidas
Y Ofran
Y Ofran
Z Yuan
Z Yuan
ZX Wang
Publication venue: Public Library of Science
Publication date: 01/01/2009
Field of study

Residue depth (RD) is a solvent exposure measure that complements the information provided by conventional accessible surface area (ASA) and describes to what extent a residue is buried in the protein structure space. Previous studies have established that RD is correlated with several protein properties, such as protein stability, residue conservation and amino acid types. Accurate prediction of RD has many potentially important applications in the field of structural bioinformatics, for example, facilitating the identification of functionally important residues, or residues in the folding nucleus, or enzyme active sites from sequence information. In this work, we introduce an efficient approach that uses support vector regression to quantify the relationship between RD and protein sequence. We systematically investigated eight different sequence encoding schemes including both local and global sequence characteristics and examined their respective prediction performances. For the objective evaluation of our approach, we used 5-fold cross-validation to assess the prediction accuracies and showed that the overall best performance could be achieved with a correlation coefficient (CC) of 0.71 between the observed and predicted RD values and a root mean square error (RMSE) of 1.74, after incorporating the relevant multiple sequence features. The results suggest that residue depth could be reliably predicted solely from protein primary sequences: local sequence environments are the major determinants, while global sequence features could influence the prediction performance marginally. We highlight two examples as a comparison in order to illustrate the applicability of this approach. We also discuss the potential implications of this new structural parameter in the field of protein structure prediction and homology modeling. This method might prove to be a powerful tool for sequence analysis

CiteSeerX

University of Melbourne Institutional Repository

: Protein Long Local Structure Prediction

Author: Altschul
Baeten
Benros
Benros
Blundell
Boeckmann
Bowie
Bystroff
Bystroff
de Bakker
de Brevern
de Brevern
de Brevern
de Brevern
de Brevern
de Brevern
de Brevern
Dong
Doppelt
Dudev
Eddy
Etchebest
Etchebest
Fawcett
Fiser
Fitzkee
Fourrier
Hastie
Hazout
Ihaka
Jauch
Joachims
Jones
Karchin
Kohonen
Kuang
Kuang
Lewis
Lin
Mittelman
Murzin
Noble
Noguchi
Offmann
Pei
Rangwala
Rohl
Rooman
Sander
Sawada
Song
Soto
Tyagi
Tyagi
Tyagi
Ward
Xiang
Yang
Zhang
Zhou
Zhu
Publication venue: 'Wiley'
Publication date: 14/01/2009
Field of study

International audienceA relevant and accurate description of three-dimensional (3D) protein structures can be achieved by characterizing recurrent local structures. In a previous study, we developed a library of 120 3D structural prototypes encompassing all known 11-residues long local protein structures and ensuring a good quality of structural approximation. A local structure prediction method was also proposed. Here, overlapping properties of local protein structures in global ones are taken into account to characterize frequent local networks. At the same time, we propose a new long local structure prediction strategy which involves the use of evolutionary information coupled with Support Vector Machines (SVMs). Our prediction is evaluated by a stringent geometrical assessment. Every local structure prediction with a Calpha RMSD less than 2.5 A from the true local structure is considered as correct. A global prediction rate of 63.1% is then reached, corresponding to an improvement of 7.7 points compared with the previous strategy. In the same way, the prediction of 88.33% of the 120 structural classes is improved with 8.65% mean gain. 85.33% of proteins have better prediction results with a 9.43% average gain. An analysis of prediction rate per local network also supports the global improvement and gives insights into the potential of our method for predicting super local structures. Moreover, a confidence index for the direct estimation of prediction quality is proposed. Finally, our method is proved to be very competitive with cutting-edge strategies encompassing three categories of local structure predictions. Proteins 2009. (c) 2009 Wiley-Liss, Inc

TANGLE: Two-Level Support Vector Regression Approach for Protein Backbone Torsion Angle Prediction from Primary Sequences

Author: A Schlessinger
A Schlessinger
A Schlessinger
AG de Brevern
B Rost
B Rost
B Rost
B Xue
C Bystroff
C Haynes
C Mooney
C Zhang
C Zheng
Christian Schönbach
D Xie
DT Jones
E Faraggi
E Faraggi
G Helles
Geoffrey I. Webb
GN Ramachandran
GP Raghava
H Zhang
H Zhang
Hao Tan
HJ Dyson
HS Kang
J Cheng
J Gao
J Gsponer
J Song
J Song
J Song
J Song
J Song
J Song
Jiangning Song
JJ Ward
JS Chauhan
K Chen
K Chen
K Chen
L Chen
L Kurgan
M Kumar
Mingjun Wang
MJ Mizianty
MJ Rooman
MJ Wood
MJ Wood
MK Kalita
MN Nguyen
MN Nguyen
MV Berjanskii
O Dor
O Dor
O Zimmermann
P Chen
P Kountouris
P Kountouris
P Sliz
PC Chen
R Gaudet
R Karchin
R Kuang
R Verma
S Ahmad
S Ahmad
S Liang
S Qiu
S Wu
S Wu
SF Altschul
T Ishida
T Zhang
T Zhang
Tatsuya Akutsu
V Vapnik
V Vapnik
W Kabsch
W Liu
W Zhang
X Miao
X Wang
XY Pan
Y Ofran
Y Ofran
YM Huang
Z Markovic-Housley
Z Yuan
Z Yuan
Z Yuan
Z Yuan
Publication venue: Public Library of Science
Publication date: 01/01/2012
Field of study

Protein backbone torsion angles (Phi) and (Psi) involve two rotation angles rotating around the Cα-N bond (Phi) and the Cα-C bond (Psi). Due to the planarity of the linked rigid peptide bonds, these two angles can essentially determine the backbone geometry of proteins. Accordingly, the accurate prediction of protein backbone torsion angle from sequence information can assist the prediction of protein structures. In this study, we develop a new approach called TANGLE (Torsion ANGLE predictor) to predict the protein backbone torsion angles from amino acid sequences. TANGLE uses a two-level support vector regression approach to perform real-value torsion angle prediction using a variety of features derived from amino acid sequences, including the evolutionary profiles in the form of position-specific scoring matrices, predicted secondary structure, solvent accessibility and natively disordered region as well as other global sequence features. When evaluated based on a large benchmark dataset of 1,526 non-homologous proteins, the mean absolute errors (MAEs) of the Phi and Psi angle prediction are 27.8° and 44.6°, respectively, which are 1% and 3% respectively lower than that using one of the state-of-the-art prediction tools ANGLOR. Moreover, the prediction of TANGLE is significantly better than a random predictor that was built on the amino acid-specific basis, with the p-value<1.46e-147 and 7.97e-150, respectively by the Wilcoxon signed rank test. As a complementary approach to the current torsion angle prediction algorithms, TANGLE should prove useful in predicting protein structural properties and assisting protein fold recognition by applying the predicted torsion angles as useful restraints. TANGLE is freely accessible at http://sunflower.kuicr.kyoto-u.ac.jp/~sjn/TANGLE/