32,288 research outputs found

    Elephant Search with Deep Learning for Microarray Data Analysis

    Full text link
    Even though there is a plethora of research in Microarray gene expression data analysis, still, it poses challenges for researchers to effectively and efficiently analyze the large yet complex expression of genes. The feature (gene) selection method is of paramount importance for understanding the differences in biological and non-biological variation between samples. In order to address this problem, a novel elephant search (ES) based optimization is proposed to select best gene expressions from the large volume of microarray data. Further, a promising machine learning method is envisioned to leverage such high dimensional and complex microarray dataset for extracting hidden patterns inside to make a meaningful prediction and most accurate classification. In particular, stochastic gradient descent based Deep learning (DL) with softmax activation function is then used on the reduced features (genes) for better classification of different samples according to their gene expression levels. The experiments are carried out on nine most popular Cancer microarray gene selection datasets, obtained from UCI machine learning repository. The empirical results obtained by the proposed elephant search based deep learning (ESDL) approach are compared with most recent published article for its suitability in future Bioinformatics research.Comment: 12 pages, 5 Tabl

    Machine Learning and Integrative Analysis of Biomedical Big Data.

    Get PDF
    Recent developments in high-throughput technologies have accelerated the accumulation of massive amounts of omics data from multiple sources: genome, epigenome, transcriptome, proteome, metabolome, etc. Traditionally, data from each source (e.g., genome) is analyzed in isolation using statistical and machine learning (ML) methods. Integrative analysis of multi-omics and clinical data is key to new biomedical discoveries and advancements in precision medicine. However, data integration poses new computational challenges as well as exacerbates the ones associated with single-omics studies. Specialized computational approaches are required to effectively and efficiently perform integrative analysis of biomedical data acquired from diverse modalities. In this review, we discuss state-of-the-art ML-based approaches for tackling five specific computational challenges associated with integrative analysis: curse of dimensionality, data heterogeneity, missing data, class imbalance and scalability issues

    Artificial Neural Network Inference (ANNI): A Study on Gene-Gene Interaction for Biomarkers in Childhood Sarcomas

    Get PDF
    Objective: To model the potential interaction between previously identified biomarkers in children sarcomas using artificial neural network inference (ANNI). Method: To concisely demonstrate the biological interactions between correlated genes in an interaction network map, only 2 types of sarcomas in the children small round blue cell tumors (SRBCTs) dataset are discussed in this paper. A backpropagation neural network was used to model the potential interaction between genes. The prediction weights and signal directions were used to model the strengths of the interaction signals and the direction of the interaction link between genes. The ANN model was validated using Monte Carlo cross-validation to minimize the risk of over-fitting and to optimize generalization ability of the model. Results: Strong connection links on certain genes (TNNT1 and FNDC5 in rhabdomyosarcoma (RMS); FCGRT and OLFM1 in Ewing’s sarcoma (EWS)) suggested their potency as central hubs in the interconnection of genes with different functionalities. The results showed that the RMS patients in this dataset are likely to be congenital and at low risk of cardiomyopathy development. The EWS patients are likely to be complicated by EWS-FLI fusion and deficiency in various signaling pathways, including Wnt, Fas/Rho and intracellular oxygen. Conclusions: The ANN network inference approach and the examination of identified genes in the published literature within the context of the disease highlights the substantial influence of certain genes in sarcomas

    Allele specific expression analysis identifies regulatory variation associated with stress-related genes in the Mexican highland maize landrace Palomero Toluqueño.

    Get PDF
    BackgroundGene regulatory variation has been proposed to play an important role in the adaptation of plants to environmental stress. In the central highlands of Mexico, farmer selection has generated a unique group of maize landraces adapted to the challenges of the highland niche. In this study, gene expression in Mexican highland maize and a reference maize breeding line were compared to identify evidence of regulatory variation in stress-related genes. It was hypothesised that local adaptation in Mexican highland maize would be associated with a transcriptional signature observable even under benign conditions.MethodsAllele specific expression analysis was performed using the seedling-leaf transcriptome of an F1 individual generated from the cross between the highland adapted Mexican landrace Palomero Toluqueño and the reference line B73, grown under benign conditions. Results were compared with a published dataset describing the transcriptional response of B73 seedlings to cold, heat, salt and UV treatments.ResultsA total of 2,386 genes were identified to show allele specific expression. Of these, 277 showed an expression difference between Palomero Toluqueño and B73 alleles under benign conditions that anticipated the response of B73 cold, heat, salt and/or UV treatments, and, as such, were considered to display a prior stress response. Prior stress response candidates included genes associated with plant hormone signaling and a number of transcription factors. Construction of a gene co-expression network revealed further signaling and stress-related genes to be among the potential targets of the transcription factors candidates.DiscussionPrior activation of responses may represent the best strategy when stresses are severe but predictable. Expression differences observed here between Palomero Toluqueño and B73 alleles indicate the presence of cis-acting regulatory variation linked to stress-related genes in Palomero Toluqueño. Considered alongside gene annotation and population data, allele specific expression analysis of plants grown under benign conditions provides an attractive strategy to identify functional variation potentially linked to local adaptation

    Automated data integration for developmental biological research

    Get PDF
    In an era exploding with genome-scale data, a major challenge for developmental biologists is how to extract significant clues from these publicly available data to benefit our studies of individual genes, and how to use them to improve our understanding of development at a systems level. Several studies have successfully demonstrated new approaches to classic developmental questions by computationally integrating various genome-wide data sets. Such computational approaches have shown great potential for facilitating research: instead of testing 20,000 genes, researchers might test 200 to the same effect. We discuss the nature and state of this art as it applies to developmental research
    • …
    corecore