Search CORE

Open Access LMU

Identifying and classifying biomedical perturbations in text

Author: Beisswanger
Bundschus
Cokol
Cokol
Evans
Friedman
Friedman
Hermjakob
Kanehisa
M. E. Crawford
P. M. Roberts
Providence
Pyysalo
R. Rodriguez-Esteban
Rzhetsky
Tsai
Publication venue: Oxford University Press
Publication date: 01/01/2008
Field of study

Molecular perturbations provide a powerful toolset for biomedical researchers to scrutinize the contributions of individual molecules in biological systems. Perturbations qualify the context of experimental results and, despite their diversity, share properties in different dimensions in ways that can be formalized. We propose a formal framework to describe and classify perturbations that allows accumulation of knowledge in order to inform the process of biomedical scientific experimentation and target analysis. We apply this framework to develop a novel algorithm for automatic detection and characterization of perturbations in text and show its relevance in the study of gene–phenotype associations and protein–protein interactions in diabetes and cancer. Analyzing perturbations introduces a novel view of the multivariate landscape of biological systems

CiteSeerX

Anticipating annotations and emerging trends in biomedical literature

Author: Bernd Wachmann
Dmitriy Fradkin
Fabian Mörchen
Julien Etienne
Markus Bundschus
Mathäus Dejori
Publication venue: 'Association for Computing Machinery (ACM)'
Publication date: 01/01/2008
Field of study

The BioJournalMonitor is a decision support system for the analysis of trends and topics in the biomedical literature. Its main goal is to identify potential diagnostic and therapeu-tic biomarkers for specific diseases. Several data sources are continuously integrated to provide the user with up-to-date information on current research in this field. State-of-the-art text mining technologies are deployed to provide added value on top of the original content, including named en-tity detection, relation extraction, classification, clustering, ranking, summarization, and visualization. We present two novel technologies that are related to the analysis of tem-poral dynamics of text archives and associated ontologies. Currently, the MeSH ontology is used to annotate the sci-entific articles entering the PubMed database with medical terms. Both the maintenance of the ontology as well as the annotation of new articles is performed largely manually. We describe how probabilistic topic models can be used to anno-tate recent articles with the most likely MeSH terms. This provides our users with a competitive advantage because, when searching for MeSH terms, articles are found long be-fore they are manually annotated. We further present a study on how to predict the inclusion of new terms in the MeSH ontology. The results suggest that early prediction of emerging trends is possible. The trend ranking functions are deployed in our system to enable interactive searches for the hottest new trends relating to a disease

CiteSeerX

Use of ontology structure and Bayesian models to aid the crowdsourcing of ICD-11 sanctioning rules

Author: Bui
Bundschus
Chapman
Coden
Cornet
Csongor Nyulas
Doan
Hahn
Hearst
Koller
Leroy
Liu
Mark A. Musen
Mortensen
Mortensen
Nadkarni
Navas
Rebholz-Schuhmann
Rector
Robert J.G. Chalmers
Rodríguez-González
Rogers
Samson W. Tu
Schulz
Sutton
Tania Tudorache
Tudorache
Wang
Whitehall
Yun Lou
Zeng
Publication venue: 'Elsevier BV'
Publication date
Field of study

GenCLiP: a software program for clustering gene lists by literature profiling and constructing gene co-occurrence networks related to custom keywords

Author: AA Schaffer
BT Alako
C Plake
C Rodriguez-Penagos
D Chaussabel
D Lee
EG Cerami
G Karakiulakis
H Kim
Hui-Yong Tian
Jin Zhao
K Fundel
Kai-Tai Yao
KJ Bussey
LJ Jensen
M Bundschus
M Suderman
MB Eisen
N Daraselia
P Shannon
R Hammamieh
R Hoffmann
R Rubinstein
RT Tsai
S Li
T Ide
TK Jenssen
VK Gajendran
Yi-Bo Zhou
Z Huang
ZF Hu
Zhen-Fu Hu
Zhong-Xi Huang
Publication venue: BioMed Central
Publication date: 01/07/2008
Field of study

Abstract Background Biomedical researchers often want to explore pathogenesis and pathways regulated by abnormally expressed genes, such as those identified by microarray analyses. Literature mining is an important way to assist in this task. Many literature mining tools are now available. However, few of them allows the user to make manual adjustments to zero in on what he/she wants to know in particular. Results We present our software program, GenCLiP (Gene Cluster with Literature Profiles), which is based on the methods presented by Chaussabel and Sher (<it>Genome Biol </it>2002, 3(10):RESEARCH0055) that search gene lists to identify functional clusters of genes based on up-to-date literature profiling. Four features were added to this previously described method: the ability to 1) manually curate keywords extracted from the literature, 2) search genes and gene co-occurrence networks related to custom keywords, 3) compare analyzed gene results with negative and positive controls generated by GenCLiP, and 4) calculate probabilities that the resulting genes and gene networks are randomly related. In this paper, we show with a set of differentially expressed genes between keloids and normal control, how implementation of functions in GenCLiP successfully identified keywords related to the pathogenesis of keloids and unknown gene pathways involved in the pathogenesis of keloids. Conclusion With regard to the identification of disease-susceptibility genes, GenCLiP allows one to quickly acquire a primary pathogenesis profile and identify pathways involving abnormally expressed genes not previously associated with the disease.</p

A context-blocks model for identifying clinical relationships in patient records

Author: A Névéol
A Roberts
AK McCallum
AM Cohen
AR Aronson
Aurélie Névéol
C Friedman
ES Chen
F Leitner
H Shatkay
H Xu
J Aberdeen
J Björne
J Lafferty
L Smith
L Tanabe
M Bundschus
M Craven
M Krallinger
N Ponomareva
O Uzuner
O Uzuner
R Harpaz
R Islamaj Doğan
R Islamaj Doğan
Rezarta Islamaj Doğan
SM Meystre
SV Pakhomov
TC Rindflesch
TC Rindflesch
X Wang
X Wang
X Wang
Zhiyong Lu
Publication venue: BioMed Central
Publication date: 01/01/2011
Field of study

Public Library of Science (PLOS)

Gene-Disease Network Analysis Reveals Functional Modules in Mendelian, Complex and Environmental Diseases

Author: A Bauer-Mehren
A Bauer-Mehren
A Bauer-Mehren
A Fernández
A Hamosh
A López García De Lomana
A-L Barabasi
A-L Barabási
AD D'Andrea
AH Smith
AJ Enright
Anna Bauer-Mehren
C Gabor
C Margadant
C The UniProt
C-Y Yang
CJ Mattingly
CR Scriver
CT Butts
D Botstein
D-T Bau
DHEW Huberts
E Cerami
E Ravasz
F Scaldaferri
Ferran Sanz
G Lima-Mendez
HY Chiou
I Celik
J Freudenberg
J Lim
JA Kennedy
JA Mitchell
JN Hirschhorn
Jv Reeuwijk
K-I Goh
KM Dipple
L De Luca
LA Garraway
Laura I. Furlong
LH Hartwell
M Argos
M Bundschus
M Cokol
M Melkoniemi
M Oti
MA van Driel
MA Yildirim
Markus Bundschus
MEJ Newman
MG Kann
Michael Rautschka
Miguel A. Mayer
MJ Thun
MP Snead
N Przulj
NA Zaghloul
NN Ahmad
R Rubinstein
R Sankaranarayanan
R Sharan
Raya Khanin
RB Altman
RH Duerr
S Ananiadou
S Carreira
S Chavali
S Jones
S Park
S Suthram
S van Dongen
SA Navarro Silvera
SI Berger
T Tsuda
TE Klein
V Radosavljević
Y Li
Z Lu
Publication venue: Public Library of Science
Publication date: 01/01/2011
Field of study

Scientists have been trying to understand the molecular mechanisms of diseases to design preventive and therapeutic strategies for a long time. For some diseases, it has become evident that it is not enough to obtain a catalogue of the disease-related genes but to uncover how disruptions of molecular networks in the cell give rise to disease phenotypes. Moreover, with the unprecedented wealth of information available, even obtaining such catalogue is extremely difficult. We developed a comprehensive gene-disease association database by integrating associations from several sources that cover different biomedical aspects of diseases. In particular, we focus on the current knowledge of human genetic diseases including mendelian, complex and environmental diseases. To assess the concept of modularity of human diseases, we performed a systematic study of the emergent properties of human gene-disease networks by means of network topology and functional annotation analysis. The results indicate a highly shared genetic origin of human diseases and show that for most diseases, including mendelian, complex and environmental diseases, functional modules exist. Moreover, a core set of biological pathways is found to be associated with most human diseases. We obtained similar results when studying clusters of diseases, suggesting that related diseases might arise due to dysfunction of common biological processes in the cell. For the first time, we include mendelian, complex and environmental diseases in an integrated gene-disease association database and show that the concept of modularity applies for all of them. We furthermore provide a functional analysis of disease-related modules providing important new biological insights, which might not be discovered when considering each of the gene-disease association repositories independently. Hence, we present a suitable framework for the study of how genetic and environmental factors, such as drugs, contribute to diseases. The gene-disease networks used in this study and part of the analysis are available at http://ibi.imim.es/DisGeNET/DisGeNETweb.html#Download

Open Access LMU

UPF Digital Repository

HypertenGene: extracting key hypertension genes from biomedical literature with position and automatically-generated template features

Author: A Rzhetsky
AK Ramani
B Rosario
C Blaschke
Chi-Hsin Huang
F Sha
H-W Chun
Hong-Jie Dai
HW Chun
J Lafferty
J Nocedal
J Xiao
JN Darroch
K Becker
K Hirohata
M Bundschus
M Craven
M Masseroli
M Shimbo
N Kambhatla
P Ruch
Po-Ting Lai
R Bunescu
R Weissberg
RC Bunescu
Richard Tzong-Han Tsai
RT Tsai
RTK Lin
T Ono
T Rindflesch
TC Rindflesch
TF Smith
TH Tsai
Wen-Harn Pan
Wen-Lian Hsu
Y Yamamoto
Yen-Ching Chang
Yue-Yang Bow
Z GuoDong
Publication venue: BioMed Central
Publication date: 01/01/2009
Field of study

Abstract Background The genetic factors leading to hypertension have been extensively studied, and large numbers of research papers have been published on the subject. One of hypertension researchers' primary research tasks is to locate key hypertension-related genes in abstracts. However, gathering such information with existing tools is not easy: (1) Searching for articles often returns far too many hits to browse through. (2) The search results do not highlight the hypertension-related genes discovered in the abstract. (3) Even though some text mining services mark up gene names in the abstract, the key genes investigated in a paper are still not distinguished from other genes. To facilitate the information gathering process for hypertension researchers, one solution would be to extract the key hypertension-related genes in each abstract. Three major tasks are involved in the construction of this system: (1) gene and hypertension named entity recognition, (2) section categorization, and (3) gene-hypertension relation extraction. Results We first compare the retrieval performance achieved by individually adding template features and position features to the baseline system. Then, the combination of both is examined. We found that using position features can almost double the original AUC score (0.8140vs.0.4936) of the baseline system. However, adding template features only results in marginal improvement (0.0197). Including both improves AUC to 0.8184, indicating that these two sets of features are complementary, and do not have overlapping effects. We then examine the performance in a different domain--diabetes, and the result shows a satisfactory AUC of 0.83. Conclusion Our approach successfully exploits template features to recognize true hypertension-related gene mentions and position features to distinguish key genes from other related genes. Templates are automatically generated and checked by biologists to minimize labor costs. Our approach integrates the advantages of machine learning models and pattern matching. To the best of our knowledge, this the first systematic study of extracting hypertension-related genes and the first attempt to create a hypertension-gene relation corpus based on the GAD database. Furthermore, our paper proposes and tests novel features for extracting key hypertension genes, such as relative position, section, and template features, which could also be applied to key-gene extraction for other diseases.</p